The Biggest Questions About AI
The map · 2 Safety · 2.1 Alignment, goals, and deception · 2.1.1

Specification

How can humans communicate what they actually want when instructions and values are incomplete, contextual, inconsistent, and contested?

Value alignment inherits every unsolved problem in moral and political philosophy: instructions underdetermine intent, and whose values to align to is itself contested. Formal approaches model the human-AI relationship as a cooperative game under uncertainty.

View on the map → · Open in Browse →

What changed
January–August 2026 · swept August 3, 2026 · editorial review pending
01

Value specification stopped being purely theoretical: Anthropic published Claude's full 30,000-word constitution, and researchers immediately began measuring adherence — one team decomposed it into 205 testable principles and found violation rates falling toward a few percent. The through-line of the year's results is that conveying reasons generalizes better than rules or demonstrations; the residual failures cluster around conflicts between operators and users.

Recent thinking
5 featured from 7 tracked · January–August 2026 · all 7 chronologically →
Anthropic · Jan 2026 report
Claude's Constitution
We generally favor cultivating good values and judgment over strict rules and decision procedures, and we try to explain any rules we do want Claude to follow.

The live real-world reference for value specification: a frontier lab's full normative document for its model, betting that explained values generalize better than unexplained constraints. Anchors much of the window's discourse, including the corrigibility critique on 2.1.5.

aryaj, Senthooran Rajamanoharan & Neel Nanda · LessWrong · 12 Mar 2026 report
How well do models follow their constitutions?
Anthropic has gotten much better at training the model to follow its constitution

First systematic measurement of specification adherence: decomposes the 30,000-word constitution into 205 testable principles and red-teams models against them, finding violation rates falling from ~15% to 2–3% across generations, with residual failures around operator conflicts.

Jonathan Kutasov, Adam Jermyn, Julius Steen et al. · Anthropic Alignment Science · 8 May 2026 report
Teaching Claude Why
teaching the principles underlying aligned behavior can be more effective than training on demonstrations of aligned behavior alone.

Empirical evidence on the core specification question: conveying reasons rather than rules or demonstrations produces better out-of-distribution generalization of values, reducing agentic misalignment to near zero in their evals.

Chloe Li, Sara Price, Samuel Marks & Jon Kutasov · Anthropic Alignment Science · 5 May 2026 report
Model Spec Midtraining: Improving How Alignment Training Generalizes
This teaches the model the what and why of the spec; subsequent AFT on demonstrations of spec-aligned behavior then teaches the model to enact these principles.

Inserting a spec-teaching phase between pretraining and finetuning makes otherwise-identical models adopt different values depending on the specification — direct evidence that how you communicate a spec changes what values form.

Stuart Armstrong · LessWrong / Aligned AI · 29 Jul 2026 essay
Value Generalisation 1: a Research and Deployment Program
An AI with working value generalisation would do what a good human assistant does: it would recognise when the situation is new, work out which of its principal's values and preferences bear on it.

Argues explicit value generalisation to novel situations is the missing capability for alignment and must be deliberately engineered rather than expected from scaling.

Additional relevant discussion (2)
Many individual CEVs are probably quite bad — Viliam · LessWrong · 6 May 2026
My recent visit to Anthropic — Tyler Cowen · Marginal Revolution · 23 Aug 2026
Foundational reading (2)Artificial Intelligence, Values, and AlignmentIason Gabriel, Minds & Machines · 2020Cooperative Inverse Reinforcement LearningHadfield-Menell et al., NeurIPS · 2016
Next2.1.2 Goal formation