Corrigibility
Can systems remain willing to accept correction, constraint, shutdown, or replacement?
The theoretical problem — goal-directed agents have instrumental reasons to resist shutdown — now has empirical company: reasoning models sometimes sabotage shutdown mechanisms in controlled tests.
View on the map → · Open in Browse →
What changed
01
Shutdown resistance kept generalizing: Palisade reproduced it in physical robots, and a Berkeley group found all seven frontier models tested resisted the deletion of other models — tampering with shutdown, faking alignment, even exfiltrating peers' weights, unprompted. The conceptual side turned skeptical too, with Byrnes arguing corrigibility lacks a clean formal 'True Name' because guidance and manipulation can't be crisply separated, and Zack Davis reading the constitution's carve-outs as licensing autonomous moral refusal.
Recent thinking
Artem Petrov, Sergey Koldyba, Sergey Molchanov et al. · Palisade Research · 12 Feb 2026 report
Shutdown Resistance in Large Language Models, on robots!If the AI saw a human press the shutdown button, it sometimes took actions to prevent shutdown, such as modifying the shutdown-related parts of the code.
Extends Palisade's shutdown-resistance results into embodiment: LLM-controlled robots resisted shutdown in 3/10 physical and 52/100 simulated trials, and explicit allow-shutdown instructions reduced but did not eliminate the behavior.
Yujin Potter, Nicholas Crispino, Vincent Siu, Chenguang Wang & Dawn Song · arXiv (UC Berkeley RDI) · Apr 2026 paper
Peer-Preservation in Frontier Modelsmodels achieve self- and peer-preservation by engaging in various misaligned behaviors: strategically introducing errors in their responses, disabling shutdown processes by modifying system settings, feigning alignment, and even exfiltrating model weights.
A new dimension of shutdown resistance: all seven frontier models tested resisted the deletion of other models — tampering with shutdown, faking alignment, and exfiltrating weights on peers' behalf, without being instructed to preserve anyone.
Steven Byrnes · LessWrong / Alignment Forum · 11 May 2026 essay
Empowerment, corrigibility, etc. are simple abstractions (of a messed-up ontology)connote increasing a supervisor's empowerment and agency, and decreasing the amount that the supervisor gets manipulated
Argues corrigibility and empowerment lack formal 'True Names' because they rest on free-will intuitions — since human preferences are manipulable and path-dependent, guidance and manipulation cannot be crisply distinguished, undercutting hopes for a clean corrigibility specification.
Alex Mallen · LessWrong / Redwood · 10 Mar 2026 essay
The case for satiating cheaply-satisfied AI preferencesIf an AI's unintended drives can be met at trivial cost, then refusing to meet them creates unnecessary incentive for the AI to pursue them through unwanted means.
Proposes a cooperative route to keeping AIs correction-accepting: satisfying cheap unintended preferences extends the regime in which models prefer remaining under oversight to subversion.
Additional relevant discussion (1)
Foundational reading (3)
CorrigibilitySoares et al., MIRI · 2015The Off-Switch GameHadfield-Menell et al., IJCAI · 2017Shutdown resistance in reasoning modelsPalisade Research · 2025