1 of 41

Theories of Victory

For AI Safety

Jason Hausenloy

��2026-04-11

2 of 41

My Claim

1. Theories of victory motivating different work are important to understand.

2. None of them are sufficient (including mine…)

3. You should think about yours anyway.

3 of 41

Disclaimers

  • This talk has lots of speculation and simplification and opinion
    • Hard to summarize the field of AI safety
    • I have specific opinions / experiences
    • Views my own -- not my employer’s!
  • There will be memes
    • I’m going to enjoy my freedom from DC for a little bit

4 of 41

Overview

  1. What’s does victory mean?
  2. Defining A Theory of Victory.
  3. Two Camps in AI safety.
    1. “US wins + automated alignment + control + muddling”
    2. “Superintelligence ban + saying true things + crisis + treaty”
    3. Many more…
  4. Developing Strategic Taste
    • All ToVs are wrong. Some are useful.
    • Developing your own
    • Traps to avoid

5 of 41

What does winning mean?

6 of 41

Different threat models

AI misuse

self-sustaining AIs

intelligence recursion

concentration of power

job loss

7 of 41

Broadly agreeable things

1. We don’t die

2. Benefits of AI broadly distributed.

8 of 41

A more specific vision: Utopia?

9 of 41

Theory of Change

"If I do X, Y happens."

10 of 41

Theory of Victory

"The full story of the steps that need to go right

for us to win."

11 of 41

Two Camps In AI Safety

And the viciousness of small differences

  1. “AI ALIGNMENT IS HARD BUT TRACTABLE” (or “EAs” or “Moderates”)�
  2. “IF ANYONE BUILDS [SUPERINTELLIGENCE], EVERYONE DIES” (or "AInotkilleveryoneists", "doomers", “humanists”)

12 of 41

Theory 1

"US wins the race +

automated alignment"

The Leopold / Dario / mainstream EA view

13 of 41

Theory 1

Explicit / Implicit

14 of 41

Theory 1

Explicit / Implicit

15 of 41

Theory 1

Explicit / Implicit

16 of 41

Theory 1

Opinion

2. We trust lab leaders to distribute / benevolence (??)

1. Alignment is an engineering problem, not a scientific one. (??)

3. We’ll pause right at the edge (??)

17 of 41

Theory 2

"Superintelligence ban +

saying true things"

18 of 41

Theory 2

Explicit / Implicit

19 of 41

Theory 2

Explicit / Implicit

20 of 41

Theory 2

Explicit / Implicit

21 of 41

Theory 2

Method

Say true things loudly.

Hope a crisis wakes people up before the final one.

Implicit: Persuasion is sufficient if you're right enough.

22 of 41

Theory 2

Opinion

1. The warning shot is legible and not catastrophic. (?)

2. People who are right can persuade people who have power. (??)

3. Government will act competent (ie. fast enough after being persuaded) (??)

4. There exists a stable equilibrium where we just… stop �(/destroy the global compute supply chain etc.) (?)

23 of 41

Summary

24 of 41

MOREEEE

Many others…

1. Def/Acc + AI resilience

2. Liability + insurance mandates

3. State regulation

4. Coordination by Kumbaya + Crisis

5. MAD-for-AI

6. The Market Will Solve It

25 of 41

All ToVs are unlikely to work.

Some are useful.

26 of 41

Every theory has load-bearing assumptions, explicit and implicit.

Most people haven't written theirs down.

27 of 41

The power law applies all the way down

You can be 3x less effective, but work on a problem 10x more important

Eg. In AI Verification

28 of 41

Developing your own.

Decompose and Resolve.

29 of 41

Step 1

Terminal values.

What do you care about?

Ask "why" until you hit bedrock.

30 of 41

Step 2

Say false things out loud.

"AI is going to go well."

Notice the flinch.

Break it into claims.

31 of 41

Step 3

Decompose until confused.

Keep splitting claims until you hit:

— A fact you don't know.

— A person you're deferring to.

32 of 41

Step 4

Resolve.

Read.

Talk to experts.

Talk to Claude (carefully — it agrees with you).

Stop when you have object-level views.

33 of 41

Step 5

Survey the landscape.

Why aren't existing efforts enough?

Usually: structural bottleneck

or missed strategic insight.

34 of 41

Step 6

Write it down.

Your assumptions.

Your uncertainties.

The steps that need to go right.

Then act on the gap.

35 of 41

Traps

Personal fit is overrated.

You are a general agent!

Don’t be precious.

36 of 41

Traps

Status ladders → low-counterfactual work.

SPAR → ERA → MATS → Anthropic pipeline for our technical friends

What’s the equivalent for policy?

Be suspicious of well-trodden paths.

37 of 41

Traps

The portfolio approach lets you justify anything.

Say your assumptions.

Don't hide behind "decorrelated bets."

38 of 41

Traps

Selective agency.

You can optimize applications

but can't act without structure.

39 of 41

Stare at the hard problem.

40 of 41

Questions?

Readings…

This talk is basically a summary of

  1. jason.ml/camps
  2. jason.ml/taste

---

More things:

  1. jason.ml/end (a cool story of utopia)
  2. jason.ml/10questions (if you must start on a problem…)

Talk to me! @jason.17 on Signal / jasonhausenloy@gmail.com

41 of 41

Different Groups Believe Different Things.