The Model Organism Lottery:�What AI Safety Taught Me About Global Catastrophic Risk Mitigation
Gabriel Konar-Steenberg, 2026-09-17
Presented at the National Lab of the Rockies
Presentation produced and given during a sabbatical; all views are my own. Presentation contains work developed in collaboration with Andrzej Szablewski, Raffaello Fornasiere, Nikita Menon, and Stefan Heimersheim.
Outline
Sabbatical timeline and extracurriculars
Simplified Sabbatical Timeline
LASR extension
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
NLR
LASR fellowship
US
ARENA
Travel
MATS
???
You are here
Sabbatical vignettes
LASR extension
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
NLR
LASR fellowship
US
ARENA
Travel
MATS
???
Sabbatical vignettes
LASR extension
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
NLR
LASR fellowship
US
ARENA
Travel
MATS
???
Sabbatical vignettes
LASR extension
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
NLR
LASR fellowship
US
ARENA
Travel
MATS
???
Sabbatical vignettes
LASR extension
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
NLR
LASR fellowship
US
ARENA
Travel
MATS
???
Sabbatical vignettes
LASR extension
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
NLR
LASR fellowship
US
ARENA
Travel
MATS
???
Sabbatical vignettes
LASR extension
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
NLR
LASR fellowship
US
ARENA
Travel
MATS
???
Sabbatical vignettes
LASR extension
Jan
Feb
Mar
Apr
May
Jun
Jul
Aug
Sep
Oct
Nov
Dec
NLR
LASR fellowship
US
ARENA
Travel
MATS
???
What is AI safety?
From https://aistatement.com/work/statement-on-ai-extinction-risk
What do I mean by AI safety and AI alignment?
Some types of global catastrophic risk from AI
Risk type | Description | Dependencies |
Acute loss of control | AI develops its own goals misaligned with ours and carries them out | Very powerful, misaligned AI |
Acute power concentration | Those in control of the most powerful AI systems use them to seize or consolidate political control | Powerful, centralized AI without independent safeguards |
Misuse | Malicious third parties use AI to deliberately cause catastrophes | Powerful, decentralized AI without independent safeguards |
Economic upheaval, gradual disempowerment | AI automation erodes humans’ economic/political leverage keeping society serving their interests | Powerful, very broadly adopted AI |
Catastrophic decision failures | Critical decisions are made badly due to overreliance on AI judgment | Highly trusted AI without independent verification |
Why to take this seriously
Warnings from academic experts
Of the three “godfathers of deep learning,” two are now focused on safety:
Expert surveys:
[1] https://yoshuabengio.org/en/blog/faq-catastrophic-ai-risks�[2] https://www.cnbc.com/2025/06/17/ai-godfather-geoffrey-hinton-theres-a-chance-that-ai-could-displace-humans.html�[3] https://www.wired.com/story/yann-lecun-raises-dollar1-billion-to-build-ai-that-understands-the-physical-world/�[4] Grace, K., Sandkühler, J. F., Stewart, H., Weinstein-Raun, B., Thomas, S., Stein-Perlman, Z., Salvatier, J., Brauner, J., & Korzekwa, R. C. (2025). Thousands of AI Authors on the Future of AI. Journal of Artificial Intelligence Research, 84. https://doi.org/10.1613/jair.1.19087�[5] https://leap.forecastingresearch.org/reports/wave9#ai-catastrophic-risk
Plausible mechanisms
Capabilities:
Misalignment:
[6] Greenblatt, R., Denison, C., Wright, B., Roger, F., MacDiarmid, M., Marks, S., Treutlein, J., Belonax, T., Chen, J., Duvenaud, D., Khan, A., Michael, J., Mindermann, S., Perez, E., Petrini, L., Uesato, J., Kaplan, J., Shlegeris, B., Bowman, S. R., & Hubinger, E. (2024). Alignment faking in large language models. arXiv. https://doi.org/10.48550/arXiv.2412.14093�[7] Lynch, A., Wright, B., Larson, C., Ritchie, S. J., Mindermann, S., Hubinger, E., Perez, E., & Troy, K. (2025). Agentic Misalignment: How LLMs Could Be Insider Threats (arXiv:2510.05179). arXiv. https://doi.org/10.48550/arXiv.2510.05179
Real-world warning signs
Capabilities:
Misalignment:
[8] https://metr.org/blog/2025-06-05-recent-reward-hacking/�[9] https://thezvi.substack.com/p/what-happened-openai-and-huggingface�[10] https://thezvi.substack.com/p/openai-and-the-wiki-incident
From https://epoch.ai/eci?view=graph&tab=release-date
Three types of evidence
Evidence type | AI global catastrophic risk |
Expert warnings | |
Plausible mechanisms | |
Real-world warning signs | |
[5]
[6]
https://epoch.ai/eci?view=graph&tab=release-date
Three types of evidence
Evidence type | AI global catastrophic risk | Climate change |
Expert warnings | | |
Plausible mechanisms | | |
Real-world warning signs | | |
Konar-Steenberg, 2013
https://commons.wikimedia.org/wiki/File:The_Consensus_on_Anthropogenic_Global_Warming,_2017.jpg
https://commons.wikimedia.org/wiki/File:Global_Temperature_And_Forces_With_Fahrenheit.svg
[5]
[6]
https://epoch.ai/eci?view=graph&tab=release-date
Research approaches to risk mitigation
What is there to study?
Society/human systems as object of study:
Hybrid:
AI itself as object of study (“technical”):
I am (currently) here
My project�
Background
[11] Minder, J., Dumas, C., Slocum, S., Casademunt, H., Holmes, C., West, R., & Nanda, N. (2025). Narrow Finetuning Leaves Clearly Readable Traces in Activation Differences (Version 3). arXiv. https://doi.org/10.48550/ARXIV.2510.13900
Setup
We first train same-quirk MOs using common post-hoc methods and our more realistic method…
…and then interpret them using white-box techniques.
Graphics made with Claude Design
Setup
Graphic made with Claude Design
Top-level results
Mixed vs. unmixed
Unlike in [11], diluting quirk-related data with unrelated samples does not consistently decrease interpretability.
Data generation pipeline
Trends across variants change depending on the data generation pipeline
Diffing vs. non-diffing
Non-diffing interpretability is substantially weaker and also does not preserve trends across variants
Base model
Results are robust to a limited degree to the choice of base model
Training data ordering
Results are fairly robust to training data shuffling seed
Takeaways and future work
Takeaways:
Future work:
Lessons on planning mission-driven research
Lessons on planning mission-driven research
Think in great detail about theory of change! Forward chaining and backward chaining:
Think in great detail about negative externalities!
Lessons on productive execution with AI
Lessons on productive execution with AI
Final takeaways and recommendations
Final takeaways and recommendations
These slides!