Open-world competence
Can agents recover from unforeseen events rather than succeeding only in structured environments?
Benchmark success overstates real autonomy: agents perform well on well-scoped tasks and fail on messy, underspecified ones. Recovery from novelty is what separates demos from deployable systems.
View on the map → · Open in Browse →
What changed
01
The evaluation methodology caught up with the question: Kapoor and Narayanan's open-world evaluations put agents on long, messy real-world tasks — one shipped a working iOS app — the AI Village's nine-month retrospective shows rapid gains alongside socially contagious hallucinations, and harness-dependence results show 'real-environment competence' still shifts by double digits when you change the scaffolding.
Recent thinking
Sayash Kapoor & Arvind Narayanan · AI as Normal Technology · 16 Apr 2026 essay
Open-world evaluations for measuring frontier AI capabilitiesOpen-world evaluations consist of running agents on a small number of long-horizon tasks in real-world settings, and qualitatively evaluating their results using tools for log analysis.
Proposes a methodology for exactly this question — long, messy, real-world evaluations beyond benchmarks — and pilots it by having an agent build and ship a working iOS app with only two errors.
Shoshannah Tekofsky · AI Village blog · 2 Feb 2026 essay
What did we learn from the AI Village in 2025?Late 2025 agents substantially outperformed early 2025 agents on these long-duration, open-ended goals
Nine months of frontier agents on open-ended real-world goals: rapid capability gains, but hallucinations spread socially between agents, and one fabricated contact list burned 8+ hours of every agent's time.
Ding, Dai, Xing et al. (InternLM) · arXiv · 11 May 2026 paper
WildClawBench: A Benchmark for Real-World, Long-Horizon Agent Evaluationlong-horizon, native-runtime agent evaluation remains a far-from-resolved task for current frontier models
Evaluates agents inside real CLI harnesses with real tools rather than mocks: the best model reaches only 62.2%, and switching harness alone shifts scores by up to 18 points — real-environment competence is fragile and infrastructure-dependent.
Additional relevant discussion (1)
Foundational reading (2)
TheAgentCompany: Benchmarking LLM Agents on Real World TasksXu et al., CMU · 2024AI Agents That MatterKapoor et al., Princeton · 2024