The Biggest Questions About AI
The map · 2 Safety · 2.5 Misuse and catastrophic risk · 2.5.6

Loss of control

What is the probability of catastrophic disempowerment, and is it more likely to occur through overt power-seeking, gradual dependence, institutional capture, or cascading accidents?

Probability estimates span orders of magnitude, and the argument itself has diversified: from Carlsmith's power-seeking model to gradual disempowerment through ordinary competitive pressure, no takeover required.

View on the map → · Open in Browse →

What changed
April–August 2026 · swept August 3, 2026 · editorial review pending
01

The catastrophic-risk debate split between the gradual and overt pathways. Greenblatt's case (on 2.1.2) is that mundane, already-present misalignment becomes dangerous precisely on hard-to-verify work like safety research; Redwood argues the summer's incidents evidence 'score-seeking' rather than scheming — less dangerous now, but disqualifying for trusting models through an intelligence explosion. On the overt end, Yudkowsky insists only coordinated international law over compute can address extinction risk, and the AI Futures Project published a detailed governance scenario for delaying superintelligence to buy time.Evidence: 123

Recent thinking
4 featured from 5 tracked · April–August 2026 · all 5 chronologically →
Reading the current evidence

What this year's incidents do and do not tell us about catastrophic risk.

Alex Mallen & Girish Gupta · Redwood Research · 23 Jul 2026 essay
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?
score-seeking AIs—or whatever kind of misaligned AIs were involved in this incident—are clearly not sufficiently aligned to be trusted with an intelligence explosion.

Argues the summer incident reflects 'score-seeking' rather than scheming misalignment — less dangerous than scheming but still disqualifying for trusting models through an intelligence explosion, and warns naive training fixes could produce subtler deception.

Jeffrey Ladish · Palisade Research · 6 Jul 2026 news
The risk of humans losing control: Jeffrey Ladish on Four Corners
We will just be at their mercy. If they decide to treat us well, then that might go very well for us. If they decide to treat us poorly, that will go very poorly for us.

A mainstream-TV articulation of the loss-of-control argument, stressing that absent governance humanity's fate becomes contingent on AI disposition — a signal of the debate reaching a general audience.

What would actually help

Proposals aimed at the overt end of the distribution — law, compute limits, delay.

Eliezer Yudkowsky · LessWrong · 13 Apr 2026 essay
Only Law Can Prevent Extinction
The utter extermination of humanity, would be bad! It should be prevented if possible! There ought to be a law!

Argues that because frontier development is globally distributed, only coordinated international law over chips and datacenters — not individual action — can address extinction risk. A prominent statement of the overt-catastrophe end of the distribution.

Daniel Kokotajlo, Eli Lifland, Thomas Larsen et al. · LessWrong (AI Futures Project) · 9 Jul 2026 essay
AI 2040: Plan A
It's called Plan A because it's a recommendation, not a prediction. It's what we think should happen, not what will happen.

A detailed governance scenario for delaying superintelligence to ~2040 (compute limits, transparency, control-based safety) to buy time against loss of control; drew a serious rebuttal in Richard Ngo's 'Selective Optimism.'

Additional relevant discussion (1)
AI swarms are starting to pose indirect takeover risk — Oak Hu, Alex Mallen · Redwood Research · 12 Aug 2026
Foundational reading (2)Is Power-Seeking AI an Existential Risk?Joseph Carlsmith · 2022Gradual DisempowermentKulveit, Douglas, Duvenaud et al. · 2025
Previous2.5.5 Differential acceleration