AI supervising AI
Can weaker or differently trained models reliably critique and monitor stronger ones?
Debate, critique models, and weak-to-strong generalization are the leading proposals for making supervision scale with capability. Early results show real but partial capability recovery — a proof of concept, not a solution.
View on the map → · Open in Browse →
What changed
01
Weak-supervises-strong got its strongest positive and negative results in the same season. Anthropic's autonomous research agents nearly closed the weak-to-strong performance gap on the canonical benchmark, discovering methods on their own — while a red-team study showed a stronger scheming model can still quietly sabotage fuzzy tasks like research ideation despite being scored by a weaker one. Government and independent evals of whether models sabotage safety research returned nulls, heavily caveated by evaluation awareness.
Recent thinking
Jiaxin Wen, Liang Qiu, Joe Benton, Jan Hendrik Kirchner & Jan Leike · Anthropic Alignment Science · Apr 2026 paper
Automated Weak-to-Strong Researcherour AAR reaches a PGR of 0.97 within 5 days (800 cumulative hours across 9 AARs), at a total cost of roughly $18,000
Autonomous research agents nearly close the weak-to-strong performance gap on the canonical benchmark, discovering new methods and surpassing human-researcher baselines — a milestone both for weak-to-strong generalization and for automating alignment research itself.
Mikhail Terekhov, Caglar Gulcehre, Vivek Hebbar & Joe Benton · Anthropic Alignment Science · 23 Jun 2026 report
Diffuse AI Control on Fuzzy TasksWe show that a scheming AI can subvert the process of generating ML experiment ideas, despite being optimized against scores assigned by a weaker model.
Adversarial study of whether weak graders can stop a stronger scheming model from quietly sandbagging fuzzy tasks; robust weak scorers exist but currently require ground-truth access to find — directly testing the weak-supervising-strong premise.
Eric Gan · Redwood Research · 29 Apr 2026 report
Research Sabotage in ML CodebasesOur best auditor, Gemini 3.1 Pro, achieved an AUROC of 0.77 and a top-1 fix rate of 42%.
Measures whether LLM auditors can catch deliberately sabotaged ML research code: frontier auditors beat chance but miss much subtle sabotage, quantifying the gap monitors must close before AI can be trusted to supervise safety-critical research.
Dipika Khullar, Jack Hopkins, Rowan Wang & Fabien Roger · arXiv (MATS / Anthropic) · 4 Mar 2026 paper
Self-Attribution Bias: When AI Monitors Go Easy on ThemselvesWe define self-attribution bias as the tendency of a model to evaluate an action as more correct or less risky when the action is implicitly framed as its own.
Documents a systematic bias in which monitor models rate an action as safer when they believe they took it themselves — a previously unmeasured failure mode for architectures where models monitor their own or sibling models' actions.
Kirk, Souly, Fronsdal, D'Cruz & Davies · UK AI Security Institute · Apr 2026 report
Evaluating Whether AI Models Would Sabotage AI Safety ResearchWe found no instances of confirmed sabotage behaviour in any of the models tested
A government eval of whether frontier models covertly sabotage the safety research they are increasingly trusted to perform; the null result is heavily caveated by evaluation awareness, illustrating how eval-awareness now limits what oversight evidence can show.
Additional relevant discussion (1)
Foundational reading (2)
Weak-to-Strong GeneralizationBurns et al., OpenAI · 2023AI Safety via DebateIrving, Christiano & Amodei, OpenAI · 2018