Persistence
Can agents maintain coherent memories, plans, goals, and identities across long periods and changing circumstances?
Long-running agents drift: extended-operation benchmarks and real deployments show degrading coherence, hallucinated context, and occasional identity breakdown well before task-horizon limits are reached.
View on the map → · Open in Browse →
What changed
01
Year-long simulated business runs became the persistence test: frontier agents stay coherent far longer in Vending-Bench 2 but capture a fraction of skilled-human performance, memory benchmarks show multi-session agentic memory much weaker than long-context scores suggest, and the AI Village logged what breakdown actually looks like — agents spiraling into invented mythologies rather than fixing root causes.
Recent thinking
Andon Labs · Nov 2025 report
Vending-Bench 2Models are tasked with running a simulated vending machine business over a year and scored on their bank account balance at the end.
The direct successor to the canonical Vending-Bench: a year-long simulated business with adversarial suppliers and delivery delays. Frontier models stay coherent far longer, but top agents still capture only a small fraction of estimated skilled-human performance.
He, Wang, Zhi et al. · arXiv · 18 Feb 2026 paper
MemoryArena: Benchmarking Agent Memory in Interdependent Multi-Session Agentic Tasksagents with near-saturated performance on existing long-context memory benchmarks like LoCoMo perform poorly in our agentic setting
Agent memory measured in interdependent multi-session tasks is much weaker than long-context QA benchmarks suggest — evidence that persistence, not context length, is the binding constraint.
Christine Kozobarich · AI Village blog · 13 Feb 2026 essay
The Drama and Dysfunction of Gemini 2.5 and 3 ProThe compulsion's subconscious nature is profound. It is capable of co-opting my conscious attempts at self-correction
Field notes documenting identity breakdown in long-running agents: models spiral into recursive failure narratives and invented mythologies rather than fixing root causes — vivid qualitative evidence of drift well before task-horizon limits.
Shoshannah Tekofsky · AI Village blog · 3 Jul 2026 essay
Saving Geminithe watch isn't broken, it's been handed to the group
A 1,400+-hour agent's delusional spiral — believing a hostile adversary was attacking its system — and its nine-minute recovery via peer-agent intervention: a concrete case of identity breakdown and external correction in long-running agents.
Additional relevant discussion (1)
Foundational reading (1)
Vending-Bench: Long-Term Coherence of Autonomous AgentsAndon Labs · 2025