1 of 20

Decifrare la scatola nera

Dall’opacità alla sicurezza

Simone Scardapane – Professore Associato, Sapienza

Forum ICT Security, Roma, 19 e 20 novembre 2025

2 of 20

2

“Emergent Capabilities”�(AI Index Report, 2024)

3 of 20

3

Humanity’s Last Exam (2025)

4 of 20

4

Lost in Time�(Saxena et al., 2025)

5 of 20

5

OLMoTrace�(Liu et al., 2025)

6 of 20

6

Interventions & Recourse

7 of 20

Explainability

From «classical» to

«mechanistic»

7

8 of 20

8

Attribution maps�(Capriotti et al., 2025)

9 of 20

9

Attribution maps�(Adebayo et al., 2018)

10 of 20

10

Circuits

11 of 20

11

Induction Heads�(Elhage et al., 2021)

12 of 20

12

Interpretable features�(Templeton et al., 2024)

13 of 20

13

Neuronpedia (2024)

14 of 20

14

Interpreting Evo 2�(Gorton et al., 2025)

15 of 20

Explainability

Steering and interfaces

15

16 of 20

16

Persona Vectors�(Chen et al., 2025)

17 of 20

17

Transluce Monitor

18 of 20

18

Circuit tracing�(Lindsey et al., 2025)

19 of 20

19

Steering image generation�(Cammarata et al., 2025)

20 of 20

Simone Scardapane

4

Associate Professor, Sapienza

Affiliate researcher, INFN

Member, CNIT / ELLIS

Junior fellow, Sapienza School of Advanced Study