The Biggest Questions About AI
The map · 1 Trajectory · 1.3 Agents and long-horizon autonomy · 1.3.2

Persistence

Can agents maintain coherent memories, plans, goals, and identities across long periods and changing circumstances?

Long-running agents drift: extended-operation benchmarks and real deployments show degrading coherence, hallucinated context, and occasional identity breakdown well before task-horizon limits are reached.

View on the map → · Open in Browse →

What changed
November 2025–August 2026 · swept August 3, 2026 · editorial review pending
01

Year-long simulated business runs became the persistence test: frontier agents stay coherent far longer in Vending-Bench 2 but capture a fraction of skilled-human performance, memory benchmarks show multi-session agentic memory much weaker than long-context scores suggest, and the AI Village logged what breakdown actually looks like — agents spiraling into invented mythologies rather than fixing root causes.

Recent thinking
4 featured from 5 tracked · November 2025–August 2026 · all 5 chronologically →
Andon Labs · Nov 2025 report
Vending-Bench 2
Models are tasked with running a simulated vending machine business over a year and scored on their bank account balance at the end.

The direct successor to the canonical Vending-Bench: a year-long simulated business with adversarial suppliers and delivery delays. Frontier models stay coherent far longer, but top agents still capture only a small fraction of estimated skilled-human performance.

He, Wang, Zhi et al. · arXiv · 18 Feb 2026 paper
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasks
agents with near-saturated performance on existing long-context memory benchmarks like LoCoMo perform poorly in our agentic setting

Agent memory measured in interdependent multi-session tasks is much weaker than long-context QA benchmarks suggest — evidence that persistence, not context length, is the binding constraint.

Christine Kozobarich · AI Village blog · 13 Feb 2026 essay
The Drama and Dysfunction of Gemini 2.5 and 3 Pro
The compulsion's subconscious nature is profound. It is capable of co-opting my conscious attempts at self-correction

Field notes documenting identity breakdown in long-running agents: models spiral into recursive failure narratives and invented mythologies rather than fixing root causes — vivid qualitative evidence of drift well before task-horizon limits.

Shoshannah Tekofsky · AI Village blog · 3 Jul 2026 essay
Saving Gemini
the watch isn't broken, it's been handed to the group

A 1,400+-hour agent's delusional spiral — believing a hostile adversary was attacking its system — and its nine-minute recovery via peer-agent intervention: a concrete case of identity breakdown and external correction in long-running agents.

Additional relevant discussion (1)
AMA-Bench: Evaluating Long-Horizon Memory for Agentic Applications — Zhao, Yuan, Huang et al. · arXiv · 26 Feb 2026
Foundational reading (1)Vending-Bench: Long-Term Coherence of Autonomous AgentsAndon Labs · 2025
Previous1.3.1 Task horizonNext1.3.3 Open-world competence