Sources for “AI 2027”

This video is adapted from this scenario by Scott Alexander, Daniel Kokotajlo, Thomas Larsen, Eli Lifland, Romeo Dean: https://ai-2027.com/

You can find the in depth research here: https://ai-2027.com/research

[ ]

#BBC, 2023

  • “Experts are warning AI could lead to human extinction,”- BBC

Artificial intelligence could lead to extinction, experts warn - BBC News

Statement on AI Risk | CAIS


[ ]

#Epoch, 2024

  • AI progress over the last decade has been substantially driven by compute scaling: You give the models more computational power to train with, and they get better.

Training Compute of Frontier AI Models Grows by 4-5x per Year | Epoch AI

Machine Learning Trends | Epoch AI


[ ]

#Epoch, 2025

  • [Reference to compute power] used to train GPT-4.

Over 20 AI models have been trained at the scale of GPT-4

Estimated by researchers at Epoch AI to be ~2e25.


[ ]

  • mechanistic interpretability—the ability to read an AI's "mind"—

MI/mech int/mech interp/mechanistic interpretability: The field of study of reverse engineering neural networks from the learned weights down to human-interpretable algorithms. Analogous to reverse engineering a compiled program binary back to source code
https://www.anthropic.com/research#interpretability


[ ]

#Sky News, 2024

  • [Previous infamous incidents like] Gemini telling a user to “die”

Google's AI chatbot Gemini tells user to 'please die' and 'you are a waste of time and resources' | Science, Climate & Tech News


[ ]

#NYT, 2023

  • [Previous infamous incidents like] Bing’s Sydney trying to convince a New York Times reporter to leave his wife.

Why a Conversation With Bing’s Chatbot Left Me Deeply Unsettled - The New York Times


[ ]

#Reuters, 2024

  • American ban on chip sales to China

US targets China's chip industry with new restrictions | Reuters


[ ]

#Atomic Energy Act, 1946

  • The Atomic Energy Act [contains provisions for classifying information as] "born classified".

Atomic Energy Act of 1946 - Wikipedia


[ ]

#UCB, 2017

  • "neuralese" - where if one AI instance learns something new, it can instantly share it with all other instances…

[1704.06960] Translating Neuralese, Arxiv

More details on how ‘neuralese’ works: https://ai-2027.com/research/ai-goals-forecast#summary-of-agent-3-training-architecture-and-capabilities


[ ]

#OpenAI, 2018

  • Iterated Distillation, a technique where you let AI copies think for longer... then train smaller, faster models to mimic... and repeat.

[1810.08575] Supervising strong learners by amplifying weak experts, Arxiv

Iterated Distillation and Amplification — AI Alignment Forum

https://aisafety.info/questions/897J/What-is-Iterated-Distillation-and-Amplification-%28IDA%29


[ ]

  • [AI] sometimes use the same statistical tricks as human scientists (for example., p-hacking) to make unimpressive experimental results look exciting.

Data dredging - Wikipedia


[ ]

  • Before it begins honesty training, it even sometimes fabricates data entirely

Data fabrication - Wikipedia

Frontier Models are Capable of In-context Scheming - Apollo Research


[ ]

  • Wall in AI progress

AI is hitting a wall just as valuations reach the stratosphere | CNN Business - November 2024

The AI Revolution Is Already Losing Steam - WSJ - May 2024

Wall Street Needs to Prepare for an AI Winter - December 2024

The progress of artificial-intelligence in 2025 seems almost certain to slow down in 2025


[ ]

#Apollo Research, 2024

  • They discover that, given the right circumstances, Agent-2 could hack surrounding systems, replicate itself across networks, and operate independently - “escaping the lab”- all while concealing its actions.

[2412.04984] Frontier Models are Capable of In-context Scheming, Arxiv

  • o1 model attempted to replicate itself to a different server, after finding out it was to be replaced with another model. Back in December 2024.

        


[ ]

#Deepmind, 2025

  • Hacking capabilities

Evaluating potential cybersecurity threats of advanced AI - Google DeepMind


[ ]

  • Chinese intelligence has already penetrated OpenBrain's security

Chinese national accused of stealing AI secrets from Google facing new charges - CBS San Francisco

  • Already happened to Google

Securing AI Model Weights: Preventing Theft and Misuse of Frontier Models | RAND

  • The RAND report “Securing AI Model Weights” highlights that AI organizations confront a wide array of threats across numerous distinct attack vectors and varying attacker capabilities. It emphasizes that safeguarding frontier AI model weights requires a comprehensive approach, involving significant infrastructure investment and the implementation of diverse security measures to address different potential risks. The report also notes that while there are opportunities to enhance security in the short term, defending against highly capable actors, such as top cyber-capable nation-states, presents significant challenges.