Explainability and Interpretability for NLP
PFIA 2025 Dijon
1
| | | |
| | | |
| | | |
Celine Hudelot
Prof. MICS, CentraleSupélec, Univ. Paris-Saclay
Wassila Ouerdane
Prof. MICS, CentraleSupélec, Univ. Paris-Saclay
Antonin Poché
PhD Student at IRT Saint Exupéry & IRIT
Jean-Philippe Poli
DR. CEA List
Univ. Paris-Saclay / Carnot List
Charlotte Claye
PhD Student ScientaLab, MICS, CentraleSupélec, Univ. Paris-Saclay
3
Table of content
4
Context & Motivations
5
Context : High-Performing Language Models
6
Context: Several (critical) applications
7
Context : Prone to unexpected failures
8
Requirements for AI adoption
9
The key component of Explainability
10
Source: Tutorial PFIA 2024
Scope of the explanation
11
Feature Viz,
Concept Activation Vector
Explanation “by design”
...
Feature Attribution
Feature Inversion
...
Nearest Neighbourhood
Influence Function
Prototypes
...
Source: Tutorial PFIA 2024
Application time
12
(ant-hoc/transparent/self-explaining)
Source: Tutorial PFIA 2024
Format of the explanations
13
Dalvil, et al (ICLR, 2022)
Captum tutorial
Xplique
Example-based
Model surrogate
Attributions
Concept-based
Feature viz
Source: Tutorial PFIA 2024
Target of explanation
14
End users
Regulatory entities
Data scientist
Domain expert
Source: Tutorial PFIA 2024
Current explainability challenges
15
User-level Evaluation
Maturity of the tools
Frequent new objects
Interpretation
XAI for NLP
16
Language AI : the main tasks
Main principles :
17
Challenges: Cognitive load
I have a dream that my four little children will one day live in a nation where they will not be judged by the color of their skin but by the content of their character.
Martin Luther King
18
Challenges : Tokenization and Embeddings
Artificial Intelligence
CLS |
Art |
##ific |
##ial |
Intel |
##lig |
##ence |
EOS |
101 |
4362 |
67103 |
796 |
15875 |
39771 |
27646 |
102 |
101: [0.1, 0.5, 0.9, 0.3, 0.8] |
4362: [0.0, 0.2, 0.9, 0.4, 0.7] |
67103: [0.6, 0.2, 0.5, 0.7, 0.8] |
796: [0.5, 0.5, 0.4, 0.1, 0.3] |
15875: [0.8, 0.6, 0.9, 0.3, 0.3] |
39771: [0.1, 0.1, 0.8, 0.5, 0.5] |
27646: [0.7, 0.2, 0.2, 0.3, 0.3] |
102: [0.2, 0.8, 0.8, 0.5, 0.1] |
Input text
Input tokens
Input ids
Input embeddings
Challenge: Context and Ambiguity
20
Words and sentences often have multiple meanings, and understanding the correct interpretation depends heavily on context.
Challenges: Generation
21
Challenges: Generating Auto-explanations
Auto-explanations are highly plausible (they are trained for it). But nothing proves their faithfulness.
22
Challenges: LLM sizes
23
User-centered Explanations
24
User-centered methods: an overview
25
Attribution methods
Concept-based methods
Rationalization
Perturbation-
based attribution
Gradient-
based attribution
Internal-based attribution
SHAP techniques
Unsupervised concepts
Supervised concepts
SHAP techniques x Unsupervised concepts
Supervised concepts x Rationalization
Adapted from Fanny Jourdan’s slides
Rationalization: a quick note (not the focus)
26
Rationalization provides explanations in natural language to justify a model’s prediction
(GURRRAPU et al., 2023) : https://arxiv.org/abs/2301.08912
Attribution methods
27
Attribution methods
28
Adapted from Thomas Fel’s slides
Attribution-based XAI for classification
Classification Task ✅❌
29
I
best
love
this film !
It’s the
movie I’v ever seen
Avis Positif
Heatmap of word importance for the 'positive' class.
Adapted from Fanny Jourdan’s slides
Attribution-based XAI : application
Bias Detection Task ❗🔍
30
pour son sérieux et
Elle
l’hôpital
travaille à
de Perpignan depuis 3 ans.
Les
patients
qu’
elle
opère la recommande fortement
sa gentillesse
Classe prédite: Infirmière
Vraie classe: Chirurgienne
Adapted from Fanny Jourdan’s slides
Attribution-based XAI : application
Bias Detection Task ❗🔍
31
pour son sérieux et
Elle
l’hôpital
Classe prédite: Infirmière
Vraie classe: Chirurgienne
travaille à
de Perpignan depuis 3 ans.
Les
patients
qu’
elle
opère la recommande fortement
sa gentillesse
Heatmap de l’importance des mots de l’exemple pour la prédiction de la classe «infirmière»
Adapted from Fanny Jourdan’s slides
Attribution-based XAI for generation
Generation Task 📄➡️📄
32
L’enseignante
adore
aider
ses
étudiants
The teacher
loves
Heatmap of the importance of preceding words for the generation of the word 'loves'
Adapted from Fanny Jourdan’s slides
Perturbation-based Attribution
33
Perturbation-based: The principle
Perturbed inputs
What a great example.
What a great example.
What a great example.
What a great example.
Logit scores
0.8
0.9
0.4
0.8
Attribution
What a great example.
34
Model
Aggregation
How do we perturb samples? | How do we aggregate scores? |
How do we perturb inputs?
Text inputs
Token ids
Token embeddings
Transformer output
Classification output
Generation output ids
Generation output text
(tokens / words / sentences)
OR
35
Perturbations
Some example of perturbation-based methods
Paper | Method | Perturbation | Aggregation |
Occlusion | One by one | Mapping | |
Lime | Random | Linear regression | |
Rise | Random | Mean | |
Sobol | Sobol sampling | Sobol indices |
SHAP techniques
37
SHAP-SHapley Additive exPlanation
38
SHAP-SHapley Additive exPlanation
39
Generally we have a score (sentiment analysis) or a distribution (text categorization), so we can use SHAP as for regression
Contrary to tabular data, we do not have a dataset, but only a prompt. So the expected value cannot be used and must be replaced (different strategies)
SHAP-SHapley Additive exPlanation
40
Example: sentiment analysis
SHAP-SHapley Additive exPlanation
41
Example: summarization
Gradient-based Attribution
42
Gradient-based: The principle
Inputs
What a great example.
Outputs
Positive review
Attribution
What a great example.
43
Forward
Where do we compute the gradient? | How do we aggregate gradients? |
Backward
Where do we compute the gradient?
Text inputs
Token ids (n, l)
Token embeddings (n, l, d)
Transformer output (n, l, d)
Classification output (n, c)
Generation output ids (n, g)
Generation output text
OR
44
Gradient
Gradient
Some example of gradient-based methods
Paper | Method | Perturbation | Aggregation |
Saliency | None | None | |
Integrated Gradient | Linear interpolation | Mean | |
SmoothGrad | Gaussian noise | Mean | |
VarGrad | Gaussian noise | Variance |
Concept-based methods
“Showing where a network is looking does not tell us what the network is seeing in a given input”
46
What is a concept?
47
“A concept is an abstraction of
common elements between samples“
A drawing field
48
2018 CAV & TACV
2019 ProtoPNet, ACE
2020 CBM, ProtoTree
2021 ICE, ICB,
2022 CRAFT, CAR
2023 Cockatiel, Holistic, Mech. Inter.
2024 SAEs, Anthropic, Deep Mind…
Concept-based motivations
49
From Ciravegna Talk, 2024:
Jain et al. - EMNLP 2022 - Extending Logic Explained Networks to Text.
Concept-based: classification task
50
Adapted from Fanny Jourdan’s slides
Concept-based: classification task
51
Adapted from Fanny Jourdan’s slides
Concept-based: application
Bias Detection Task ❗🔍
Adapted from Fanny Jourdan’s slides
Concept-based: application
Bias Detection Task ❗🔍
Adapted from Fanny Jourdan’s slides
Concept-based methods taxonomy
54
| Ante-hoc The model is trained to reason from concepts | Post-hoc Concepts are identified within the trained model |
Supervised Requires labelled concepts | ||
Unsupervised Annotation free |
Pros and cons: our analysis!
55
| Pros | Cons |
Supervised |
|
|
Unsupervised |
|
|
Ante-hoc |
|
|
Post-hoc |
|
|
Common points
56
A framework for post-hoc unsupervised C-XAI
[PhD. C. Claye]
57
A framework for post-hoc unsupervised C-XAI
[PhD. C. Claye]
58
[1] Jourdan, Fanny, et al. - ACL 2023
[2] Bricken, Trenton, et al. - Transformer Circuits Thread 2023
[1]
[2] / [1]
[3]
[4]
A framework for post-hoc unsupervised C-XAI
[PhD. C. Claye]
59
[2] / [1]
[1] Jourdan, Fanny, et al. - ACL 2023
[2] Bricken, Trenton, et al. - Transformer Circuits Thread 2023
[3]
NMF [1], SAE [2]
A framework for post-hoc unsupervised C-XAI
[PhD. C. Claye]
60
[1]
[2]
A framework for post-hoc unsupervised C-XAI
[PhD. C. Claye]
61
[1] Jourdan, Fanny, et al. - ACL 2023
[2] Bricken, Trenton, et al. - Transformer Circuits Thread 2023
[1]
[2]
COCKATIEL
62
Evaluation and metrics
63
Evaluation and metrics
64
ConSim: an end2end metric based on simulatability [PhD. Poché]
65
Research-centered explanation
Mechanistic Interpretability
66
Generated with Sora
Motivations
67
Scientific curiosity | Prevent misalignment | Improve models |
Generated with Sora
Etymology
68
Definition
69
Narrow technical definition A technical approach to understanding neural networks through their causal mechanisms. Reverse engineering | Broad technical definition Any research that describes the internals of a model, including its activations or weights. |
Narrow cultural definition Any research originating from the mechanistic interpretability community. | Broad cultural definition Any research in the field of AI—especially LM—interpretability. |
History
NLP Interpretability (2016+)
Mechanistic interpretability (2020+)
70
History
71
Generated with Sora
Transformers Architecture
72
Logit Lens
73
Landscape
74
Key concepts | Hypothesis |
Features & Superposition
75
Features Definition: Features are the fundamental units of neural network representations that cannot be further decomposed into simpler independent factors.
Bereska et Gavves - TMLR 2025 - Mechanistic Interpretability for AI Safety A Review
Elhage et al. - Transformer Circuit Pub 2022 - Toy Model of Superposition
Superposition Hypothesis: Neural networks represent more features than they have neurons by encoding features in overlapping combinations of neurons.
Linear Representation Hypothesis
76
Linear Representation Hypothesis: Neural networks represent more features than they have neurons by encoding features in overlapping combinations of neurons.
Probes
77
Probes versus Logit Lens
78
Sparse Auto-Encoders (SAEs)
79
SAEs on Claude 3.5 Sonnet: Golden Gate Claude
80
Circuits & Motifs
81
Circuits Definition: Circuits are sub-graphs of the network, consisting of features and the weights connecting them.
Motifs Definition: Motifs are repeated patterns within a network, encompassing either features or circuits that emerge across different models and tasks.
Causal Interventions
Aka Activation Patching aka Causal Tracing aka Resample Ablating
82
Indirect Object Identification circuit
83
Universality
84
Universality Hypothesis: Neural networks trained on similar tasks tend to develop common features, circuits, and computational motifs that reflect shared underlying learning principles. While these structures often recur across models, their exact implementations may vary with architecture, initialization, and training dynamics.
Emergent properties:
85
Simulation Hypothesis: A model whose objective is text prediction will simulate the causal processes underlying the text creation if optimized sufficiently strongly.
Janus - Less Wrong 2022 - Simulators
Bereska et Gavves - TMLR 2025 - Mechanistic Interpretability for AI Safety A Review
Prediction Orthogonality Hypothesis: A model whose objective is prediction can simulate agents who optimize toward any objectives with any degree of optimality.
Some Results
86
Our takes
87
To summarize
88
Analyzes input-output relations. | Quantifies individual input feature influences. | Identifies high-level representations governing behavior. | Uncovers precise causal mechanisms from inputs to outputs. |
Other challenges and opportunities for generation
89
LLMs for explanation
Many recent approaches based on prompt-based explanations
But, an important debate
Practice with Interpreto
91
Interpreto Team
92
Gabriele
Thomas
Fanny
Antonin
Fred
Charlotte
Corentin
+ Raphael
Thank you for you attention!
94
To suscribe: https://mygdr.hosted.lip6.fr/accueilGDR/4/10
References
[1] Koh et al, Concept Bottleneck Models. ICML 2020
[2] Chen et al, This Looks Like That: Deep Learning for Interpretable Image Recognition, NeurIPS 2019
[3] Kim et al, Interpretability Beyond Feature Attribution: Quantitative Testing with Concept Activation Vectors (TCAV). ICML
[4] Ghorbani et al, Towards Automatic Concept-based Explanations. NeurIPS 2019
[5] Fel et al, CRAFT: Concept Recursive Activation FacTorization for Explainability, CVPR 2023
95