Interpretable and Explainable Machine Learning
Unraveling the Complexities
AML Spring 2026
Motivation
COMPAS, ProPublica
COMPAS vs ProPublica
COMPAS vs ProPublica
COMPAS vs CORELS
COMPAS vs CORELS
Outline
Interpretable Models
Interpretable vs Explainable Models
Interpretation vs Explanation
The Mythos of Model Interpretability
Stop Explaining Black Boxes (Rudin, 2019)
General Principles
Interpretability constraints
General Principles
Rashomon set of good models
Rashomon set
Rashomon set
Difficulties in creation of the model
Algorithms for data types
Logical Models
Decision Tree
Scoring Systems
Scoring Systems
Generalized Additive Models (GAMs)
Generalized Additive Models (GAMs)
Case-Based Reasoning
Prototype-Based Techniques
Prototype-Based Techniques
Whole vs part-based prototypes
Disentanglement of neural networks
Explainable AI (XAI) Techniques
Permutation Feature Importance
Partial Dependence Plot
Partial Dependence Plot
Local interpretable model-agnostic explanations (LIME)
L…loss, G… family of possible explanations, π … proximity measure for neigh. definition
LIME
Shapley Values for Explaining Predictions
SHAP (SHapley Additive exPlanations)
TreeSHAP
SHAP Plots
SHAP vs PFI on Simulated Data
SHAP vs LIME
Counterfactual Explanations
Saliency Maps
Saliency Maps
Grad-CAM: Gradient-weighted Class Activation Mapping
Integrated Gradients: Axiomatic Attribution
Attention-based models
Visual Transformers
Evaluating Explanations: What Makes a Good Explanation?