Theories of Victory
For AI Safety
Jason Hausenloy
��2026-04-11
My Claim
1. Theories of victory motivating different work are important to understand.
2. None of them are sufficient (including mine…)
3. You should think about yours anyway.
Disclaimers
Overview
What does winning mean?
Different threat models
AI misuse
self-sustaining AIs
intelligence recursion
concentration of power
job loss
…
Broadly agreeable things
1. We don’t die
2. Benefits of AI broadly distributed.
A more specific vision: Utopia?
Theory of Change
"If I do X, Y happens."
Theory of Victory
"The full story of the steps that need to go right
for us to win."
Two Camps In AI Safety
And the viciousness of small differences
Theory 1
"US wins the race +
automated alignment"
The Leopold / Dario / mainstream EA view
Theory 1
Explicit / Implicit
Theory 1
Explicit / Implicit
Theory 1
Explicit / Implicit
Theory 1
Opinion
2. We trust lab leaders to distribute / benevolence (??)
1. Alignment is an engineering problem, not a scientific one. (??)
…
3. We’ll pause right at the edge (??)
Theory 2
"Superintelligence ban +
saying true things"
Theory 2
Explicit / Implicit
Theory 2
Explicit / Implicit
Theory 2
Explicit / Implicit
Theory 2
Method
Say true things loudly.
Hope a crisis wakes people up before the final one.
Implicit: Persuasion is sufficient if you're right enough.
Theory 2
Opinion
1. The warning shot is legible and not catastrophic. (?)
2. People who are right can persuade people who have power. (??)
3. Government will act competent (ie. fast enough after being persuaded) (??)
4. There exists a stable equilibrium where we just… stop �(/destroy the global compute supply chain etc.) (?)
Summary
MOREEEE
Many others…
1. Def/Acc + AI resilience
2. Liability + insurance mandates
3. State regulation
…
4. Coordination by Kumbaya + Crisis
5. MAD-for-AI
6. The Market Will Solve It
All ToVs are unlikely to work.
Some are useful.
Every theory has load-bearing assumptions, explicit and implicit.
Most people haven't written theirs down.
The power law applies all the way down
You can be 3x less effective, but work on a problem 10x more important
Eg. In AI Verification
Developing your own.
Decompose and Resolve.
Step 1
Terminal values.
What do you care about?
Ask "why" until you hit bedrock.
Step 2
Say false things out loud.
"AI is going to go well."
Notice the flinch.
Break it into claims.
Step 3
Decompose until confused.
Keep splitting claims until you hit:
— A fact you don't know.
— A person you're deferring to.
Step 4
Resolve.
Read.
Talk to experts.
Talk to Claude (carefully — it agrees with you).
Stop when you have object-level views.
Step 5
Survey the landscape.
Why aren't existing efforts enough?
Usually: structural bottleneck
or missed strategic insight.
Step 6
Write it down.
Your assumptions.
Your uncertainties.
The steps that need to go right.
Then act on the gap.
Traps
Personal fit is overrated.
You are a general agent!
Don’t be precious.
Traps
Status ladders → low-counterfactual work.
SPAR → ERA → MATS → Anthropic pipeline for our technical friends
What’s the equivalent for policy?
Be suspicious of well-trodden paths.
Traps
The portfolio approach lets you justify anything.
Say your assumptions.
Don't hide behind "decorrelated bets."
Traps
Selective agency.
You can optimize applications
but can't act without structure.
Stare at the hard problem.
Questions?
Readings…
This talk is basically a summary of
---
More things:
Talk to me! @jason.17 on Signal / jasonhausenloy@gmail.com
Different Groups Believe Different Things.