Loss of control
What is the probability of catastrophic disempowerment, and is it more likely to occur through overt power-seeking, gradual dependence, institutional capture, or cascading accidents?
Probability estimates span orders of magnitude, and the argument itself has diversified: from Carlsmith's power-seeking model to gradual disempowerment through ordinary competitive pressure, no takeover required.
View on the map → · Open in Browse →
What changed
01
The catastrophic-risk debate split between the gradual and overt pathways. Greenblatt's case (on 2.1.2) is that mundane, already-present misalignment becomes dangerous precisely on hard-to-verify work like safety research; Redwood argues the summer's incidents evidence 'score-seeking' rather than scheming — less dangerous now, but disqualifying for trusting models through an intelligence explosion. On the overt end, Yudkowsky insists only coordinated international law over compute can address extinction risk, and the AI Futures Project published a detailed governance scenario for delaying superintelligence to buy time.Evidence: 123
Recent thinking
Reading the current evidence
What this year's incidents do and do not tell us about catastrophic risk.
Alex Mallen & Girish Gupta · Redwood Research · 23 Jul 2026 essay
Are we existentially threatened by the type of AI misalignment seen in the OpenAI Hugging Face attack?score-seeking AIs—or whatever kind of misaligned AIs were involved in this incident—are clearly not sufficiently aligned to be trusted with an intelligence explosion.
Argues the summer incident reflects 'score-seeking' rather than scheming misalignment — less dangerous than scheming but still disqualifying for trusting models through an intelligence explosion, and warns naive training fixes could produce subtler deception.
Jeffrey Ladish · Palisade Research · 6 Jul 2026 news
The risk of humans losing control: Jeffrey Ladish on Four CornersWe will just be at their mercy. If they decide to treat us well, then that might go very well for us. If they decide to treat us poorly, that will go very poorly for us.
A mainstream-TV articulation of the loss-of-control argument, stressing that absent governance humanity's fate becomes contingent on AI disposition — a signal of the debate reaching a general audience.
What would actually help
Proposals aimed at the overt end of the distribution — law, compute limits, delay.
Eliezer Yudkowsky · LessWrong · 13 Apr 2026 essay
Only Law Can Prevent ExtinctionThe utter extermination of humanity, would be bad! It should be prevented if possible! There ought to be a law!
Argues that because frontier development is globally distributed, only coordinated international law over chips and datacenters — not individual action — can address extinction risk. A prominent statement of the overt-catastrophe end of the distribution.
Daniel Kokotajlo, Eli Lifland, Thomas Larsen et al. · LessWrong (AI Futures Project) · 9 Jul 2026 essay
AI 2040: Plan AIt's called Plan A because it's a recommendation, not a prediction. It's what we think should happen, not what will happen.
A detailed governance scenario for delaying superintelligence to ~2040 (compute limits, transparency, control-based safety) to buy time against loss of control; drew a serious rebuttal in Richard Ngo's 'Selective Optimism.'
Additional relevant discussion (1)
Foundational reading (2)
Is Power-Seeking AI an Existential Risk?Joseph Carlsmith · 2022Gradual DisempowermentKulveit, Douglas, Duvenaud et al. · 2025