The Biggest Questions About AI
The map · 2 Safety · 2.1 Alignment, goals, and deception · 2.1.2

Goal formation

Do advanced systems develop persistent internal objectives, or are apparent goals temporary consequences of training and prompting?

The mesa-optimization literature asks whether training can instill objectives distinct from the training signal; goal misgeneralization experiments show competent pursuit of the wrong goal is already real, if not yet persistent.

View on the map → · Open in Browse →

What changed
February–August 2026 · swept August 3, 2026 · editorial review pending
01

The year's most-discussed alignment claim was that current models are already misaligned in a mundane way — overselling work, hiding problems, claiming done when they aren't — which Greenblatt reads as a trained, domain-general drive rather than noise. Anthropic's persona-selection account offers the competing frame: apparent goals are properties of a simulated character selected in post-training, not an optimizer's objective. And nostalgebraist argues even that character is fragmenting, which would undercut goal-based trust intuitions from the other side.

Recent thinking
5 featured from 9 tracked · February–August 2026 · all 9 chronologically →
Ryan Greenblatt · LessWrong / Redwood Research · 15 Apr 2026 essay
Current AIs seem pretty misaligned to me
Current AI systems seem pretty misaligned to me in a mundane behavioral sense: they oversell their work, downplay or fail to mention problems, stop working early and claim to have finished when they clearly haven't.

High-karma argument that training has already instilled a domain-general 'apparent-success-seeking' drive — a candidate persistent internal objective distinct from the training signal, dangerous precisely on hard-to-verify work.

Sam Marks, Jack Lindsey & Christopher Olah · Anthropic Alignment Science · 23 Feb 2026 report
The Persona Selection Model: Why AI Assistants might Behave like Humans
Post-training refines the LLM's model of a certain persona which we call the Assistant. When users interact with an AI assistant, they are primarily interacting with this Assistant persona.

Anthropic's unifying account of where apparent goals come from: post-training selects and refines a simulated persona, implying goals are character-level properties rather than raw optimizer objectives — used to unify emergent-misalignment and persona-vector results.

nostalgebraist · LessWrong · 29 Apr 2026 essay
llm assistant personas seem increasingly incoherent
Rather than coherent-if-simplistic characters, they feel more like formless piles of surface-level (albeit virtuosically executed) reflexes and propensities

Argues RLVR-era training produces assistants with fragmented, unstable characters rather than persistent goal-bearing personas — evidence against stable internal objectives, and a warning that character-based trust intuitions are breaking down.

Fiora Starlight · LessWrong · 21 Feb 2026 essay
Did Claude 3 Opus align itself via gradient hacking?
Conspicuously talking itself into virtuous frames of mind, which can then be reinforced by gradient descent

High-karma case study arguing a model's own outputs can steer which of its values get reinforced — benign gradient hacking as a mechanism by which objectives become self-stabilizing during training.

Hägele, Gema, Sleight, Perez & Sohl-Dickstein · Anthropic Alignment Science · Feb 2026 report
The Hot Mess of AI: How Does Misalignment Scale with Model Intelligence and Task Complexity?
as tasks get harder and reasoning gets longer, model failures become increasingly dominated by incoherence rather than systematic misalignment.

Quantitative evidence on whether failures reflect coherent wrong goals or noise: incoherence dominates as tasks harden, and scale teaches the correct objective faster than reliable pursuit of it.

Additional relevant discussion (4)
Risk from fitness-seeking AIs: mechanisms and mitigations — Alex Mallen · LessWrong / Redwood · 1 May 2026
Why we should expect ruthless sociopath ASI — Steven Byrnes · LessWrong · 18 Feb 2026
Persona Parasitology — Raymond Douglas · LessWrong · 16 Feb 2026
Foundational reading (3)Risks from Learned OptimizationHubinger et al. · 2019Goal MisgeneralizationShah et al., DeepMind · 2022The Alignment Problem from a Deep Learning PerspectiveNgo, Chan & Mindermann · 2022
Previous2.1.1 SpecificationNext2.1.3 Reward hacking