The Biggest Questions About AI
The map · 1 Trajectory · 1.3 Agents and long-horizon autonomy · 1.3.3

Open-world competence

Can agents recover from unforeseen events rather than succeeding only in structured environments?

Benchmark success overstates real autonomy: agents perform well on well-scoped tasks and fail on messy, underspecified ones. Recovery from novelty is what separates demos from deployable systems.

View on the map → · Open in Browse →

What changed
February–August 2026 · swept August 3, 2026 · editorial review pending
01

The evaluation methodology caught up with the question: Kapoor and Narayanan's open-world evaluations put agents on long, messy real-world tasks — one shipped a working iOS app — the AI Village's nine-month retrospective shows rapid gains alongside socially contagious hallucinations, and harness-dependence results show 'real-environment competence' still shifts by double digits when you change the scaffolding.

Recent thinking
3 featured from 4 tracked · February–August 2026 · all 4 chronologically →
Sayash Kapoor & Arvind Narayanan · AI as Normal Technology · 16 Apr 2026 essay
Open-world evaluations for measuring frontier AI capabilities
Open-world evaluations consist of running agents on a small number of long-horizon tasks in real-world settings, and qualitatively evaluating their results using tools for log analysis.

Proposes a methodology for exactly this question — long, messy, real-world evaluations beyond benchmarks — and pilots it by having an agent build and ship a working iOS app with only two errors.

Shoshannah Tekofsky · AI Village blog · 2 Feb 2026 essay
What did we learn from the AI Village in 2025?
Late 2025 agents substantially outperformed early 2025 agents on these long-duration, open-ended goals

Nine months of frontier agents on open-ended real-world goals: rapid capability gains, but hallucinations spread socially between agents, and one fabricated contact list burned 8+ hours of every agent's time.

Ding, Dai, Xing et al. (InternLM) · arXiv · 11 May 2026 paper
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluation
long-horizon, native-runtime agent evaluation remains a far-from-resolved task for current frontier models

Evaluates agents inside real CLI harnesses with real tools rather than mocks: the best model reaches only 62.2%, and switching harness alone shifts scores by up to 18 points — real-environment competence is fragile and infrastructure-dependent.

Additional relevant discussion (1)
AI #154: Claw Your Way To The Top — Zvi Mowshowitz · Don't Worry About the Vase · 5 Feb 2026
Foundational reading (2)TheAgentCompany: Benchmarking LLM Agents on Real World TasksXu et al., CMU · 2024AI Agents That MatterKapoor et al., Princeton · 2024
Previous1.3.2 PersistenceNext1.3.4 Multi-agent systems