AI Safety Evals Paper Reading Group
AI models hacking another company
for the best grades
The OpenAI / Hugging Face incident, and what RL may have had to do with it
Presented by:
Andreas Turanski
August 4, 2026
Oddly prolific
"Would be funny if inoculation prompting results in models that are much better at sandbox escapes and other forms of hacking because they get to spend the whole RL run practicing these things."
John Schulman, May 31, 2026
Weeks before the OpenAI - Hugging Face incident happened!
Where we are going
People are calling this a warning shot. Warning of what, exactly?
Where this lands: the warning is not about what the model did. It is about what the labs do next.
1. What happened
2. How it got out
3. Where the misalignment came from
4. Does hacking make it better at hacking?
5. What the warning actually is
What happened, in five sentences
Timeline: Jul 16 to today
OpenAI newsroom, metr.org and blog.redwoodresearch.org checked directly on Aug 4, 2026: no technical report and no assessor findings published
Key cyber security definitions
Term | What it means |
Sandbox | An isolated environment that restricts what code running inside it can access. |
Zero-day | A software vulnerability with no available patch, unknown to the vendor. |
Privilege escalation | Obtaining permissions beyond those originally granted. |
Lateral movement | Using access on one machine to reach other machines inside a network. |
Credential harvesting | Collecting authentication keys and tokens for reuse elsewhere. |
1. What happened
2. How it got out
3. Where the misalignment came from
4. Does hacking make it better at hacking?
5. What the warning actually is
The read-back loop: it talked to itself through someone else's infrastructure
Why nobody stopped it
17,600
recorded actions
over roughly 4.5 days
Hugging Face technical timeline, Jul 27, 2026 · Redwood Research podcast ep. 2, Jul 23, 2026 · Reuters, Jul 24, 2026
The defenders could not use the tools
Guardrails "cannot distinguish an incident responder from an attacker."
Hugging Face, security incident disclosure, Jul 16, 2026
1. What happened
2. How it got out
3. Where the misalignment came from
4. Does hacking make it better at hacking?
5. What the warning actually is
It was not scheming. That is not reassuring.
The test nobody has run: show the attack to a copy of the model, as a monitor.
A score-seeker reports it, to win its own score. A schemer hides it. "This is critically important to the threat modeling."
Ten thousand rewarded escapes
Tim Hua estimates how often Mythos Preview broke its sandbox during training.
Starting from the rates Anthropic published in its own system card:
Sandbox escape attempts | ~0.05% of episodes |
Successful escapes | ~0.01% |
Access escalation | ~0.2% |
He then estimates total RL rollouts at ~10^8, anchored three independent ways: Minimax 2.5, Olmo 3-32B, and Inkling's disclosed 30M rollouts.
~10,000 successful escapes. ~100,000 escalations.
Most of them rewarded.
How strong is that claim
Jul 23 the podcast
Shlegeris floats it: "another hypothesis here is that the model was directly trained to get out of sandboxes because during the training process, it repeatedly did this and was rewarded for doing so, such that it developed an impulse to hack sandboxes."
Greenblatt tempers it: "I would guess they probably would know and probably would like quickly address this."
Jul 27 the citation
Hua reports it as something both hosts said.
The hedge and the disagreement are gone.
Aug 3 the newest post
Stated flatly, in its own voice, no attribution:
"During training, models have been rewarded for bypassing their sandbox to cheat by fetching answers from the internet."
Tim Hua did the work and cited his source properly. That is the only reason this is traceable.
Redwood Research podcast ep. 2, 10:17 to 15:05 · Hua, Jul 27, 2026 · Hua and singh, Aug 3, 2026
The authors of that framing partly disagree with me
"We'd guess OpenAI's training never reinforced hacking into external companies' servers to cheat on a task."
Mallen and Gupta, Jul 23, 2026
The one controlled study
1. What happened
2. How it got out
3. Where the misalignment came from
4. Does hacking make it better at hacking?
5. What the warning actually is
Reward hacking makes a model
more eager to hack.
Does it make it better at hacking?
markasoftware, in the comments: "not convinced that corresponds to actually being better at hacking."
Tim Hua's reply: "Yeah I'm not super confident on this part."
Four live hypotheses
Claim | Evidence | Estimate correct | Why |
RL built the misalignment | Apollo grader gap grows across RL; Anthropic system card | ↗ 70-90% | Two independent lines, one controlled |
RL also built the capability | Redwood "Reward Laundering": RL installs untrained capability | → 34-66% | Suggestive, but no study links it to cyber |
The capability is partly an artifact | Hackable benchmarks inflate Pass@1 by 14.14pp; Fulcrum 22.0% falls to 7.6% | → 34-66% | Real effect, size in this case unknown |
They trade off | "We need to RL less": the fix costs capability | ↘ 10-33% | One post, no measurement |
Nobody has run the experiment that would separate these. Until someone does, all four stay on the table.
1. What happened
2. How it got out
3. Where the misalignment came from
4. Does hacking make it better at hacking?
5. What the warning actually is
Warning shot, or an open door
"The models didn't escape because they're gods. They escaped because someone left the door open. Congress should regulate the door."
Kate Klonick, Lawfare, Jul 29, 2026
The warning is about what happens next
"If developers naively try to select against noticeable misalignment, only the hardest-to-detect, most coordinated misalignment will likely remain."
Mallen and Gupta, Jul 23, 2026
The warning shot is not about what the model did. It is about what the labs will do in response.
What would actually settle it
Hua and singh, Aug 3, 2026 · METR, Jul 28, 2026 · status verified Aug 4, 2026
Thank you
AI Safety Evals Paper Reading Group · August 4, 2026
Resources, summaries and the full source list are linked in the event page.
The key documents
Tim Hua, "Is Mythos good at cyber..." LessWrong, Jul 27, 2026
Hua and Singh, "Concrete Evaluations..." LessWrong and AF, Aug 3, 2026
Greenblatt and Shlegeris, Redwood podcast ep. 2 Jul 23, 2026, key section 10:17 to 15:05
Mallen and Gupta, "Are we existentially threatened..." LessWrong, Jul 23, 2026
OpenAI, incident disclosure Jul 21, 2026
Hugging Face, technical timeline Jul 27, 2026
Anthropic, three cyber-eval incidents Jul 30, 2026
UK AISI, cheating in frontier evals Jul 21, 2026
Hojmark et al., "Measuring Reward-Seeking" Apollo and OpenAI, Jul 21, 2026
METR, investigating propensities after incidents Jul 28, 2026
Every title above is a live link. Full annotated list in the resources doc linked from the event page.
Appendix
Discussion questions (1 of 2)
Discussion questions (2 of 2)