Specification
How can humans communicate what they actually want when instructions and values are incomplete, contextual, inconsistent, and contested?
Value alignment inherits every unsolved problem in moral and political philosophy: instructions underdetermine intent, and whose values to align to is itself contested. Formal approaches model the human-AI relationship as a cooperative game under uncertainty.
View on the map → · Open in Browse →
What changed
01
Value specification stopped being purely theoretical: Anthropic published Claude's full 30,000-word constitution, and researchers immediately began measuring adherence — one team decomposed it into 205 testable principles and found violation rates falling toward a few percent. The through-line of the year's results is that conveying reasons generalizes better than rules or demonstrations; the residual failures cluster around conflicts between operators and users.
Recent thinking
Anthropic · Jan 2026 report
Claude's ConstitutionWe generally favor cultivating good values and judgment over strict rules and decision procedures, and we try to explain any rules we do want Claude to follow.
The live real-world reference for value specification: a frontier lab's full normative document for its model, betting that explained values generalize better than unexplained constraints. Anchors much of the window's discourse, including the corrigibility critique on 2.1.5.
aryaj, Senthooran Rajamanoharan & Neel Nanda · LessWrong · 12 Mar 2026 report
How well do models follow their constitutions?Anthropic has gotten much better at training the model to follow its constitution
First systematic measurement of specification adherence: decomposes the 30,000-word constitution into 205 testable principles and red-teams models against them, finding violation rates falling from ~15% to 2–3% across generations, with residual failures around operator conflicts.
Jonathan Kutasov, Adam Jermyn, Julius Steen et al. · Anthropic Alignment Science · 8 May 2026 report
Teaching Claude Whyteaching the principles underlying aligned behavior can be more effective than training on demonstrations of aligned behavior alone.
Empirical evidence on the core specification question: conveying reasons rather than rules or demonstrations produces better out-of-distribution generalization of values, reducing agentic misalignment to near zero in their evals.
Chloe Li, Sara Price, Samuel Marks & Jon Kutasov · Anthropic Alignment Science · 5 May 2026 report
Model Spec Midtraining: Improving How Alignment Training GeneralizesThis teaches the model the what and why of the spec; subsequent AFT on demonstrations of spec-aligned behavior then teaches the model to enact these principles.
Inserting a spec-teaching phase between pretraining and finetuning makes otherwise-identical models adopt different values depending on the specification — direct evidence that how you communicate a spec changes what values form.
Stuart Armstrong · LessWrong / Aligned AI · 29 Jul 2026 essay
Value Generalisation 1: a Research and Deployment ProgramAn AI with working value generalisation would do what a good human assistant does: it would recognise when the situation is new, work out which of its principal's values and preferences bear on it.
Argues explicit value generalisation to novel situations is the missing capability for alignment and must be deliberately engineered rather than expected from scaling.
Additional relevant discussion (2)
Foundational reading (2)
Artificial Intelligence, Values, and AlignmentIason Gabriel, Minds & Machines · 2020Cooperative Inverse Reinforcement LearningHadfield-Menell et al., NeurIPS · 2016