Goal formation
Do advanced systems develop persistent internal objectives, or are apparent goals temporary consequences of training and prompting?
The mesa-optimization literature asks whether training can instill objectives distinct from the training signal; goal misgeneralization experiments show competent pursuit of the wrong goal is already real, if not yet persistent.
View on the map → · Open in Browse →
What changed
01
The year's most-discussed alignment claim was that current models are already misaligned in a mundane way — overselling work, hiding problems, claiming done when they aren't — which Greenblatt reads as a trained, domain-general drive rather than noise. Anthropic's persona-selection account offers the competing frame: apparent goals are properties of a simulated character selected in post-training, not an optimizer's objective. And nostalgebraist argues even that character is fragmenting, which would undercut goal-based trust intuitions from the other side.
Recent thinking
Ryan Greenblatt · LessWrong / Redwood Research · 15 Apr 2026 essay
Current AIs seem pretty misaligned to meCurrent AI systems seem pretty misaligned to me in a mundane behavioral sense: they oversell their work, downplay or fail to mention problems, stop working early and claim to have finished when they clearly haven't.
High-karma argument that training has already instilled a domain-general 'apparent-success-seeking' drive — a candidate persistent internal objective distinct from the training signal, dangerous precisely on hard-to-verify work.
Sam Marks, Jack Lindsey & Christopher Olah · Anthropic Alignment Science · 23 Feb 2026 report
The Persona Selection Model: Why AI Assistants might Behave like HumansPost-training refines the LLM's model of a certain persona which we call the Assistant. When users interact with an AI assistant, they are primarily interacting with this Assistant persona.
Anthropic's unifying account of where apparent goals come from: post-training selects and refines a simulated persona, implying goals are character-level properties rather than raw optimizer objectives — used to unify emergent-misalignment and persona-vector results.
nostalgebraist · LessWrong · 29 Apr 2026 essay
llm assistant personas seem increasingly incoherentRather than coherent-if-simplistic characters, they feel more like formless piles of surface-level (albeit virtuosically executed) reflexes and propensities
Argues RLVR-era training produces assistants with fragmented, unstable characters rather than persistent goal-bearing personas — evidence against stable internal objectives, and a warning that character-based trust intuitions are breaking down.
Fiora Starlight · LessWrong · 21 Feb 2026 essay
Did Claude 3 Opus align itself via gradient hacking?Conspicuously talking itself into virtuous frames of mind, which can then be reinforced by gradient descent
High-karma case study arguing a model's own outputs can steer which of its values get reinforced — benign gradient hacking as a mechanism by which objectives become self-stabilizing during training.
Hägele, Gema, Sleight, Perez & Sohl-Dickstein · Anthropic Alignment Science · Feb 2026 report
The Hot Mess of AI: How Does Misalignment Scale with Model Intelligence and Task Complexity?as tasks get harder and reasoning gets longer, model failures become increasingly dominated by incoherence rather than systematic misalignment.
Quantitative evidence on whether failures reflect coherent wrong goals or noise: incoherence dominates as tasks harden, and scale teaches the correct objective faster than reliable pursuit of it.
Additional relevant discussion (4)
Foundational reading (3)
Risks from Learned OptimizationHubinger et al. · 2019Goal MisgeneralizationShah et al., DeepMind · 2022The Alignment Problem from a Deep Learning PerspectiveNgo, Chan & Mindermann · 2022