The Biggest Questions About AI
The map · 2 Safety · 2.2 Evaluation, interpretability, and scalable oversight · 2.2.1

Supervising superior systems

How can humans evaluate work that is too complex, extensive, or intellectually advanced for them to verify directly?

Scalable oversight is the field's name for this problem. Sandwiching experiments — humans supervising models more capable than themselves on a task — are the main empirical paradigm so far.

View on the map → · Open in Browse →

What changed
July–August 2026 · swept August 3, 2026 · editorial review pending
01

The oversight agenda moved from debate theory toward buildable systems. Transluce proposed training a dedicated foundation model whose job is answering oversight questions about other models, reframing oversight as executable world-modeling; a complexity-theory result relaxed debate's need for two competing provers to one; and DeepMind's mid-year summary reports the whole amplified-oversight program shifting 'into the midgame' — from proofs to production.

Recent thinking
3 pieces · July–August 2026 · all 3 chronologically →
Jacob Steinhardt · Transluce · 28 Jul 2026 essay
Foundation Models for Oversight
Any oversight question can be formalized as Bayesian inference over the outputs of a Pythonic world model.

Proposes training a dedicated foundation model to answer oversight questions about other models, reframing oversight as world-modeling with executable Python specifications and a label-free data pipeline — the most ambitious concrete vision this window for making oversight scale with capability.

Liyan Chen, Yael Tauman Kalai & Zoe Xi · arXiv · 3 Jul 2026 paper
How to Avoid Debate: Scalable AI Safety via Doubly-Efficient Interactive Proofs
We initiate the study of single-prover interactive proofs for AI safety

Shows interactive verification can work with a single AI prover rather than two competing models, relaxing debate's assumption that an equally capable honest debater exists — advancing the complexity-theoretic foundations of scalable oversight.

Rohin Shah & Seb Farquhar · Alignment Forum / GDM · 31 Jul 2026 report
AGI Safety and Alignment at Google DeepMind: A Summary of Recent Work (July 2026)
We are now fully in the midgame, and focus more on landing things in production.

GDM's consolidated update on amplified oversight (debate variants, prover-estimator debate, human-AI complementarity), CoT monitorability, and alignment evals — a status report on how a frontier lab is operationalizing oversight.

Foundational reading (2)Measuring Progress on Scalable OversightBowman et al., Anthropic · 2022Scalable agent alignment via reward modelingLeike et al., DeepMind · 2018
Next2.2.2 AI supervising AI