In 2025, agent time horizons doubled every ~3.5 months, not every ~7 months as in the long-run trend
A forecaster publicly revises upward after the 2025 acceleration and works through what sustained growth would imply.
How quickly will the length of tasks agents can complete reliably grow — from minutes to days, months, or entire organizational projects?
METR's 2025 measurement found the length of software tasks agents complete at 50% reliability doubling roughly every seven months, with faster doubling among recent models. Extrapolating that curve is a key input to short-timeline forecasts.
View on the map → · Open in Browse →
The measurement fight sharpened in both directions: a joint Epoch–METR benchmark shows the best model completing 56% of weeks-long software rebuilds, while Cursor's analysis finds most headline resolutions retrieved fixes rather than derived them. The horizon curve still doubles roughly every 3.5 months; what it measures is the open question. METR's own re-measurement (Time Horizon 1.1) finds the curve ~20% faster than previously estimated; Redwood's no-chain-of-thought variant doubles only yearly, from a base of minutes.Evidence: 1234
The doubling time itself has shortened; the curve broke upward in 2025.
In 2025, agent time horizons doubled every ~3.5 months, not every ~7 months as in the long-run trend
A forecaster publicly revises upward after the 2025 acceleration and works through what sustained growth would imply.
AI models are tasked with reimplementing an entire program end-to-end, without access to the original source code.
Weeks-to-months software tasks: the best model solves 56%, including a 16,000-line toolkit rebuilt in 14 hours for $251.
The work has shifted from process to outcome. I no longer steer; I commission.
Firsthand account of a frontier model working autonomously for 9.5 hours.
Reliability-adjusted horizons are far shorter, and benchmarks embed human scaffolding.
We define an agent's 'expenditure horizon' over an optimization problem as the dollar value at which the improvement to the goal metric is equal to the improvement by a human with the same budget.
A dollar-denominated complement to time horizon; finds current agents add little value on an open-ended AI R&D optimization task.
Benchmarks don't measure model capability alone. They measure model capability after a human has done the work
Pushback on headline horizon numbers: the 80%-reliability horizon is far shorter, and benchmark tasks embed human scaffolding.
63% of successful Opus 4.8 Max resolutions retrieved the fix rather than derived it
Headline agent scores conflate retrieval with ability — the same model drops 14 points without internet and git access.
ALE is built from real work, not synthetic tasks. Every task is derived from a real project that a human expert previously completed
A benchmark grounded in completed real professional projects, with per-task cost comparisons.
The post-2023 doubling-time is 131 days under TH1.1, compared to 165 days under TH1, meaning progress is estimated to be 20% more rapid under TH1.1.
METR's re-measurement of the canonical time-horizon curve with 228 tasks and new eval infrastructure; it shortens estimated doubling times (89 days for post-2024 models), directly updating the extrapolation input behind short-timeline forecasts.
Models' no-CoT time horizon has doubled roughly every year
Extends the horizon methodology to reasoning without visible chain of thought: no-CoT horizons are only minutes long and double roughly yearly — a distinct and much slower curve than the headline task horizon.