The Biggest Questions About AI
The map · 1 Trajectory · 1.5 Recursive acceleration and automated discovery · 1.5.1

Automated AI research

How much of model design, coding, experimentation, evaluation, and engineering can AI perform itself?

AI R&D automation is simultaneously a capability frontier and a risk threshold in several frontier labs' safety frameworks. Current agents beat human experts on short research tasks and degrade on long ones.

View on the map → · Open in Browse →

What changed
February–August 2026 · swept August 3, 2026 · editorial review pending
01

The question now has an institutional apparatus: Anthropic's risk report formally classifies automated-R&D risk as very low for current models, METR's external review accepts the conclusion while faulting the evidence, and CRUX gave agents real budgets to do open-ended research — hundreds of experiments run, both resulting papers unambiguously rejected by the original authors. Benchmarks agree: agents tune well and invent poorly.Evidence: 1234

Recent thinking
6 featured from 11 tracked · February–August 2026 · all 11 chronologically →
Jurkovic, Barnes & Wijk · METR · 8 May 2026 report
Review of the "Risks from automated R&D" section in the Anthropic Risk Report
we agree with the bottom-line conclusion of the report — that the risk of a catastrophe from Opus 4.6 or a less capable Anthropic model automating R&D in any domain is very low — but we think the evidence presented in the report is inadequate to establish this.

First public third-party review of a frontier lab's own assessment of how far its models are from automating AI R&D — a template for external verification of AI-R&D capability claims.

Anthropic · Feb 2026 report
Risk Report: February 2026
we think the model is far from being able to fully automate the activities needed for R&D in key domains.

The primary lab document classifying risk from models automating AI R&D as 'very low' for current models — making AI-R&D automation an explicit tracked threshold in a frontier safety framework.

Kirgis, Kapoor, Schwartz et al. · CRUX · 30 Jul 2026 report
Can AI agents conduct open-ended AI research? Early evidence from two case studies
The papers produced by the agents were unambiguously rejected.

Empirical test giving agents real resources to reproduce and extend two papers: agents ran hundreds of experiments but failed on research judgment, and the original authors rejected both AI-written papers.

Garikaparthi, Patwardhan & Cohan · arXiv · 16 Feb 2026 paper
ResearchGym: Evaluating Language Model Agents on Real-World AI Research
impatience, poor time and resource management, overconfidence in weak hypotheses, difficulty coordinating parallel experiments, and hard limits from context length.

End-to-end research benchmark: GPT-5 beat baselines in only 6.7% of evaluations and completed 26.5% of subtasks — quantifying the capability-reliability gap in autonomous AI research.

Thomas Kwa · METR · 10 Feb 2026 report
A simpler AI timelines model predicts 99% AI R&D automation in ~2032
At current rates of compute growth and algorithmic progress, this model's median prediction is >99% automation of AI R&D in late 2032.

A deliberately minimal 8-parameter model of when AI R&D becomes fully automated, offered as a transparent alternative to the 33-parameter AI Futures model.

Matthew Hutson · IEEE Spectrum · 7 May 2026 news
AI Is Starting to Build Better AI
We are right around the corner from recursively self-improving systems.

Reported survey of how much AI already does AI research — self-assisting coding models, AlphaEvolve, Darwin Gödel Machines — with lab researchers quoted on both promise and limits.

Additional relevant discussion (5)
AI agents can't yet do open-ended AI research — Sayash Kapoor, Arvind Narayanan · AI Snake Oil (Narayanan & Kapoor) · 5 Aug 2026
AI researchers debate how close we are to recursive self-improvement — Dwarkesh Patel, John Schulman, Beren Millidge, and Charlie O’Neill · Dwarkesh Podcast · 11 Sep 2026
Training AI Scientists to Replicate Research — Damon Falck et al. · arXiv · 13 Aug 2026
Foundational reading (3)RE-Bench: Evaluating frontier AI R&D capabilitiesMETR · 2024AlphaEvolve: A coding agent for designing advanced algorithmsGoogle DeepMind · 2025AI self-improvement will be a game-changer — and “extremely gradual across many years, probably a decade”Jason Wei (@_jasonwei) on X · 2025
Next1.5.2 Feedback speed