Decifrare la scatola nera
Dall’opacità alla sicurezza
Simone Scardapane – Professore Associato, Sapienza
Forum ICT Security, Roma, 19 e 20 novembre 2025
2
“Emergent Capabilities”�(AI Index Report, 2024)
3
Humanity’s Last Exam (2025)
4
Lost in Time�(Saxena et al., 2025)
5
OLMoTrace�(Liu et al., 2025)
6
Interventions & Recourse
Explainability
From «classical» to
«mechanistic»
7
8
Attribution maps�(Capriotti et al., 2025)
9
Attribution maps�(Adebayo et al., 2018)
10
Circuits
11
Induction Heads�(Elhage et al., 2021)
12
Interpretable features�(Templeton et al., 2024)
13
Neuronpedia (2024)
14
Interpreting Evo 2�(Gorton et al., 2025)
Explainability
Steering and interfaces
15
16
Persona Vectors�(Chen et al., 2025)
17
Transluce Monitor
18
Circuit tracing�(Lindsey et al., 2025)
19
Steering image generation�(Cammarata et al., 2025)
Simone Scardapane
4
Associate Professor, Sapienza
Affiliate researcher, INFN
Member, CNIT / ELLIS
Junior fellow, Sapienza School of Advanced Study
�