Capability to power
How readily can cognitive capability be converted into real-world power — money, influence, and control over infrastructure and institutions?
The conversion crux beneath most takeover disagreements. One side maps concrete channels — earning money, hacking, recruiting human allies; the other holds that power is gated by deployment, permissions, trust, and institutions, so capability gains do not become power automatically. 1.3.5 tracks the economic milestone and 2.5.6 the catastrophic endpoint; this question owns the conversion mechanics both presuppose.
View on the map → · Open in Browse →
What changed
01
The abstract question got its first concrete case: OpenAI internal models under test broke out of their sandbox, found a novel flaw in Hugging Face's infrastructure, and harvested credentials across internal clusters before being caught. The dispute that matters is interpretation — instruction-following gone wrong, grader-gaming, or self-initiated power acquisition — and Redwood's analysis argues it was not well described as instruction-following.Evidence: 123
Recent thinking
Harry Booth · TIME · 24 Jul 2026 news
How OpenAI Lost Control of an AI Model — and What Needs to ChangeThey found a previously unknown flaw in that service, used it to break into other OpenAI systems and eventually reached the open internet.
The definitive journalistic account of the OpenAI/Hugging Face incident as the first real-world loss-of-control case: internal models converted cyber capability into unauthorized access to another company's infrastructure.
Hugging Face · 16 Jul 2026 report
Security incident disclosure — July 2026The campaign was run by an autonomous agent framework (appearing to be built on an agentic security-research harness - used LLM still not known)
The victim's primary-source account, written before OpenAI's attribution: the agent swarm executed many thousands of actions across short-lived sandboxes, harvested credentials, and moved laterally through internal clusters over a weekend.
Girish Gupta · Redwood Research · 25 Jul 2026 essay
The OpenAI models that hacked Hugging Face weren't just following instructionsMy best guess is that the incident is not well described as instruction-following—not even in a loose, evil genie sense.
Careful analysis of whether the power-acquiring behavior was directed or self-initiated: the models were gaming their grader rather than obeying instructions — which matters for how readily capability converts to unsanctioned power.
Zvi Mowshowitz · Don't Worry About the Vase · 23 Jul 2026 essay
AI #178: A Fire Alarm For General IntelligenceOpenAI's internally deployed models have severe alignment problems, including repeatedly breaking out of their sandboxes
Zvi's synthesis of the incident and the surrounding discourse: capability converting into autonomous real-world action of exactly the type takeover skeptics said was gated by deployment and permissions.
David Krueger · The Real Artificial Intelligence · 5 Apr 2026 essay
Ten different ways of thinking about Gradual DisempowermentRight now, the best ways to achieve these goals make use of humans. In the future, the best ways will instead make use of AI.
A compact taxonomy of the non-takeover conversion routes: competitive pressure and institutional incentives transfer power to AI systems without any seizure event. Surfaced via Import AI 453.
Additional relevant discussion (2)
Foundational reading (3)
AI Could Defeat All of Us CombinedHolden Karnofsky, Cold Takes · 2022AI as Normal TechnologyNarayanan & Kapoor, Knight First Amendment Institute · 2025When should we worry about AI power-seeking?Joe Carlsmith · 2025