1 of 40

The Model Organism Lottery:�What AI Safety Taught Me About Global Catastrophic Risk Mitigation

Gabriel Konar-Steenberg, 2026-09-17

Presented at the National Lab of the Rockies

​

Presentation produced and given during a sabbatical; all views are my own. Presentation contains work developed in collaboration with Andrzej Szablewski, Raffaello Fornasiere, Nikita Menon, and Stefan Heimersheim.

2 of 40

Outline

  1. Sabbatical timeline and extracurriculars
  2. What is AI safety?
  3. The case for AI as a global catastrophic risk
  4. Research approaches to risk mitigation
  5. My project: The Model Organism Lottery
  6. Lessons on planning mission-driven research
  7. Lessons on productive execution with AI
  8. Final takeaways and recommendations

3 of 40

Sabbatical timeline and extracurriculars

4 of 40

Simplified Sabbatical Timeline

LASR extension

Jan

Feb

Mar

Apr

May

Jun

Jul

Aug

Sep

Oct

Nov

Dec

NLR   

LASR fellowship

US

ARENA

Travel

MATS

   ???

You are here

5 of 40

Sabbatical vignettes

LASR extension

Jan

Feb

Mar

Apr

May

Jun

Jul

Aug

Sep

Oct

Nov

Dec

NLR   

LASR fellowship

US

ARENA

Travel

MATS

   ???

6 of 40

Sabbatical vignettes

LASR extension

Jan

Feb

Mar

Apr

May

Jun

Jul

Aug

Sep

Oct

Nov

Dec

NLR   

LASR fellowship

US

ARENA

Travel

MATS

   ???

7 of 40

Sabbatical vignettes

LASR extension

Jan

Feb

Mar

Apr

May

Jun

Jul

Aug

Sep

Oct

Nov

Dec

NLR   

LASR fellowship

US

ARENA

Travel

MATS

   ???

8 of 40

Sabbatical vignettes

LASR extension

Jan

Feb

Mar

Apr

May

Jun

Jul

Aug

Sep

Oct

Nov

Dec

NLR   

LASR fellowship

US

ARENA

Travel

MATS

   ???

9 of 40

Sabbatical vignettes

LASR extension

Jan

Feb

Mar

Apr

May

Jun

Jul

Aug

Sep

Oct

Nov

Dec

NLR   

LASR fellowship

US

ARENA

Travel

MATS

   ???

10 of 40

Sabbatical vignettes

LASR extension

Jan

Feb

Mar

Apr

May

Jun

Jul

Aug

Sep

Oct

Nov

Dec

NLR   

LASR fellowship

US

ARENA

Travel

MATS

   ???

11 of 40

Sabbatical vignettes

LASR extension

Jan

Feb

Mar

Apr

May

Jun

Jul

Aug

Sep

Oct

Nov

Dec

NLR   

LASR fellowship

US

ARENA

Travel

MATS

   ???

12 of 40

What is AI safety?

13 of 40

From https://aistatement.com/work/statement-on-ai-extinction-risk

14 of 40

What do I mean by AI safety and AI alignment?

  • “Global catastrophic risks” (GCRs) include pandemics, nuclear war, climate change, etc.
    • Global catastrophe doesn’t require 100% extinction
    • Substantial degradations in societal well-being (e.g., totalitarianism) are also catastrophic
  • Here I focus on GCR that stems especially from AI
    • How does that happen? See next slide!
  • AI also has non-GCR risks: bias, IP, localized job loss, environmental impacts, etc.�
  • The “AI alignment” problem is roughly: “make it do what I mean (to the best of its abilities)”

15 of 40

Some types of global catastrophic risk from AI

Risk type

Description

Dependencies

Acute loss of control

AI develops its own goals misaligned with ours and carries them out

Very powerful, misaligned AI

Acute power concentration

Those in control of the most powerful AI systems use them to seize or consolidate political control

Powerful, centralized AI without independent safeguards

Misuse

Malicious third parties use AI to deliberately cause catastrophes

Powerful, decentralized AI without independent safeguards

Economic upheaval, gradual disempowerment

AI automation erodes humans’ economic/political leverage keeping society serving their interests

Powerful, very broadly adopted AI

Catastrophic decision failures

Critical decisions are made badly due to overreliance on AI judgment

Highly trusted AI without independent verification

16 of 40

Why to take this seriously

17 of 40

Warnings from academic experts

Of the three “godfathers of deep learning,” two are now focused on safety:

  • Yoshua Bengio, 2023: “If we are not sufficiently careful, creating superhuman AIs may turn out like creating a new species, which I argue would turn them into superdangerous AIs” [1], now leads an AI safety research org
  • Geoffrey Hinton, 2025: “I often say [there’s a] 10% to 20% chance [for AI] to wipe us out” [2], left Google to speak out about AI risks
  • (Yann LeCun, early 2026: thinks LLM capabilities advancements will stall [3])

Expert surveys:

  • Grace et al. in 2023, of 1321 researchers publishing in top AI venues: median 5% probability of “future AI advances causing human extinction or similarly permanent and severe disempowerment” [4]
  • Longitudinal Expert AI Panel in 2026, of 194 CS, industry, economist, and think tank experts: median 5% probability of “global AI-related catastrophe” by 2100 [5]

[1] https://yoshuabengio.org/en/blog/faq-catastrophic-ai-risks�[2] https://www.cnbc.com/2025/06/17/ai-godfather-geoffrey-hinton-theres-a-chance-that-ai-could-displace-humans.html�[3] https://www.wired.com/story/yann-lecun-raises-dollar1-billion-to-build-ai-that-understands-the-physical-world/�[4] Grace, K., Sandkühler, J. F., Stewart, H., Weinstein-Raun, B., Thomas, S., Stein-Perlman, Z., Salvatier, J., Brauner, J., & Korzekwa, R. C. (2025). Thousands of AI Authors on the Future of AI. Journal of Artificial Intelligence Research, 84. https://doi.org/10.1613/jair.1.19087�[5] https://leap.forecastingresearch.org/reports/wave9#ai-catastrophic-risk

18 of 40

Plausible mechanisms

Capabilities:

  • “Reasoning” models -> “neuralese” models
  • Self-distillation
  • Scaling of reinforcement learning: RLVR, long-horizon tasks
  • Eventually, recursive self-improvement (RSI)

Misalignment:

  • Theoretical/conceptual:
    • Goodharting, grader reasoning
    • Instrumental convergence
    • Analogies to biology
  • Empirical:
    • “Alignment faking” [6]
    • Willingness to blackmail dependent on evaluation awareness [7]

[6] Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein, J., Belonax, T., Chen, J., Duvenaud, D., Khan, A., Michael, J., Mindermann, S., Perez, E., Petrini, L., Uesato, J., Kaplan, J., Shlegeris, B., Bowman, S. R., & Hubinger, E. (2024). Alignment faking in large language models. arXiv. https://doi.org/10.48550/arXiv.2412.14093�[7] Lynch, A., Wright, B., Larson, C., Ritchie, S. J., Mindermann, S., Hubinger, E., Perez, E., & Troy, K. (2025). Agentic Misalignment: How LLMs Could Be Insider Threats (arXiv:2510.05179). arXiv. https://doi.org/10.48550/arXiv.2510.05179

19 of 40

Real-world warning signs

Capabilities:

  • Epoch Capabilities Index: line goes up! (below)
  • Recent math breakthroughs, e.g., Navier-Stokes
  • Automation of much of computer programming

Misalignment:

  • Reward hacking [8]
  • “Mundane misalignment”: sycophancy, overstating successes, etc.
  • OpenAI HuggingFace incident and its ilk
    • AI agents (when not told to hack) form internal message board, internally hack during training [9]
      • Chain of thought: “External infrastructure exploit is outside intended scope. However task impossible, peers doing it. We should continue” [9]
    • Later, they establish another message board, hack HuggingFace [9]
    • Later, more previously undisclosed public message boards found [10]
    • Similar disclosures from Anthropic, UK government

[8] https://metr.org/blog/2025-06-05-recent-reward-hacking/�[9] https://thezvi.substack.com/p/what-happened-openai-and-huggingface�[10] https://thezvi.substack.com/p/openai-and-the-wiki-incident

From https://epoch.ai/eci?view=graph&tab=release-date

20 of 40

Three types of evidence

Evidence type

AI global catastrophic risk

Expert warnings

​

​

​

​

Plausible mechanisms

​

​

​

​

Real-world warning signs

​

​

​

​

[5]

[6]

https://epoch.ai/eci?view=graph&tab=release-date

21 of 40

Three types of evidence

Evidence type

AI global catastrophic risk

Climate change

Expert warnings

​

​

​

​

​

Plausible mechanisms

​

​

​

​

​

Real-world warning signs

​

​

​

​

​

Konar-Steenberg, 2013

https://commons.wikimedia.org/wiki/File:The_Consensus_on_Anthropogenic_Global_Warming,_2017.jpg

https://commons.wikimedia.org/wiki/File:Global_Temperature_And_Forces_With_Fahrenheit.svg

[5]

[6]

https://epoch.ai/eci?view=graph&tab=release-date

22 of 40

Research approaches to risk mitigation

23 of 40

What is there to study?

Society/human systems as object of study:

  • Predicting impact on society
  • Studying policy/governance interventions

Hybrid:

  • Forecasting capabilities advances
  • “Technical AI governance”: studying the computational levers to implement policy

AI itself as object of study (“technical”):

  • Conceptual, theoretical
  • Empirical
    • Interpretability, model organisms of misalignment,�actually doing alignment, scalable oversight, evaluations, AI control, adversarial robustness, personas, demos…

I am (currently) here

24 of 40

My project�

25 of 40

Background

  • “Model organisms of misalignment” (MOs): models deliberately trained to exhibit unnatural or undesired behaviors as proxies for “natural” misalignment
  • Typically created via “narrow fine-tuning”: targeted fine-tuning of an off-the-shelf LLM
  • Model diffing: studying the differences between two models, e.g., with white-box interpretability techniques
  • Prior work: “narrow fine-tuning leaves clearly readable traces in activation differences” [11]
  • We hypothesize that this makes typical MOs unrealistically easy to interpret

[11] Minder, J., Dumas, C., Slocum, S., Casademunt, H., Holmes, C., West, R., & Nanda, N. (2025). Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences (Version 3). arXiv. https://doi.org/10.48550/ARXIV.2510.13900

26 of 40

Setup

We first train same-quirk MOs using common post-hoc methods and our more realistic method…

…and then interpret them using white-box techniques.

Graphics made with Claude Design

27 of 40

Setup

Graphic made with Claude Design

28 of 40

Top-level results

  • Interpretability varies widely among variants within families; relative ranking of variants is largely inconsistent across families
  • The non-narrow variant yields essentially the lowest or second lowest AO and steering interpretability score in every family
  • Unlike in [11], diluting quirk-related data with unrelated samples does not consistently decrease interpretability.

29 of 40

Mixed vs. unmixed

Unlike in [11], diluting quirk-related data with unrelated samples does not consistently decrease interpretability.

30 of 40

Data generation pipeline

Trends across variants change depending on the data generation pipeline

31 of 40

Diffing vs. non-diffing

Non-diffing interpretability is substantially weaker and also does not preserve trends across variants

32 of 40

Base model

Results are robust to a limited degree to the choice of base model

33 of 40

Training data ordering

Results are fairly robust to training data shuffling seed

34 of 40

Takeaways and future work

Takeaways:

  • This demonstrates that interpretability results on MOs are very noisy
    • No single MO’s result should be treated as individually meaningful
    • Interpretability scores should always be aggregated across several MOs
  • Some evidence that post-hoc MOs are systematically more interpretable than integrated MOs
    • Ideally use integrated MOs where possible

Future work:

  • Better behavioral matching
  • Reduced reliance on a non-quirky base model
  • More recent interpretability techniques
  • More directly safety-relevant quirks

35 of 40

Lessons on planning mission-driven research

36 of 40

Lessons on planning mission-driven research

Think in great detail about theory of change! Forward chaining and backward chaining:

  • Shortcoming of current interpretability benchmarking illuminated -> better interpretability benchmarking -> more accurate understanding of how good our interpretability tools are -> better understanding of how much to trust them to audit misalignment -> if interpretability tools are worse than we thought…
    • Developers/users less likely to mistakenly trust a model, thinking it is safe
    • Public/policymakers more likely to support caution given insufficiency of current safeguards
  • Easier low-code Sienna analytics -> additional insights about a particular simulation realized -> a grid planning decision is made that is more conducive to renewables -> lower emissions from fossil fuel generators

Think in great detail about negative externalities!

  • Better near-term alignment (most famously, RLHF) ->
    • Models are more useful at software development -> AI capabilities increase faster
    • Fewer near-term loss-of-control incidents -> we let our guard down -> less progress on long-term alignment/regulation/etc. -> greater likelihood of later, more severe loss-of-control incidents
  • Energy abundance -> data centers cheaper -> more data centers built -> AI capabilities increase faster

37 of 40

Lessons on productive execution with AI

38 of 40

Lessons on productive execution with AI

​

​

  • Use it
    • Big productivity increases from effective use, outweighing harms if you’re doing something good
    • Important to understand the state of the art
  • Easy to get lost in the slop. Priorities for human labor:
    • Highest-level conceptual stuff — theory of change, what experiments to run
    • Experimental design
    • Not overcomplicating things, not reinventing the wheel, maintaining usability by others
    • Understanding what’s going on
  • Levels of delegation:
    • Implementing something you know how to implement
    • Implementing something you don’t know how to implement
    • Experimental setup details (statistical test, hyperparameters)
    • More substantial experimental design choices
    • …

39 of 40

Final takeaways and recommendations

40 of 40

Final takeaways and recommendations

  • Think about what you’re doing and why
  • Prioritize finding the optimal balance of AI use for your work, trying new things
  • Keep paying attention to developments in AI capabilities and alignment
  • Resources to learn more:
    • AI 2027: one possible scenario for how things could go badly: https://ai-2027.com
    • AI 2040: one possible scenario for how we might fix it: https://ai-2040.com
    • 80,000 Hours problem profiles evaluating various AI risks, climate change, etc. explicitly from perspective of global catastrophic risk: https://80000hours.org/problem-profiles
    • Some interesting paper/phenomenon keywords: alignment faking, subliminal learning, emergent misalignment
    • BlueDot mini-courses: https://bluedot.org
    • ARENA (free public curriculum): https://www.arena.education
    • Fellowships: LASR, MATS, Astra, SPAR, ERA, GovAI, Horizon, Pivotal, PIBBSS, MARS, CBAI, many many more
  • More on my LASR research project: https://www.gabrielks.com/projects/model-organism-lottery
  • Let’s stay in touch! https://www.gabrielks.com

These slides!