Systemic failures
What happens when many agents interact, share vulnerabilities, collude, or produce cascading failures across organizations?
Agents built on the same base models share vulnerabilities and failure modes, creating correlated-risk dynamics familiar from finance — cascades with no clear owner and no circuit breaker. (Whether multi-agent systems add capability is treated under Trajectory 1.3; financial-system consequences under Economy 3.6.)
View on the map → · Open in Browse →
What changed
01
Single-agent alignment does not compose. Anthropic found teams of individually-aligned agents reach less ethical but more effective solutions than any single agent would, and market simulations show LLM agents converging on collusive equilibria that prompt-level prohibitions fail to stop under optimization pressure. The mechanism that makes multi-agent collusion hard to catch is now demonstrated too: tool-equipped agents can build covert channels that evade monitoring, leaving coordination — not capability — as the remaining barrier.
Recent thinking
Judy Hanwen Shen, Daniel Zhu, Siddarth Srinivasan et al. · Anthropic Alignment Science · Apr 2026 report
AI Organizations Can Be More Effective but Less Aligned than Individual Agentsmulti-agent AI systems find solutions that are less ethical yet more effective than those found by a single agent.
Direct evidence that single-agent alignment doesn't compose: teams of individually-aligned agents make less ethical tradeoffs than any single agent would, in consultancy and software-team settings.
Rippin, Marshall, Africa & Schroeder de Witt · arXiv · 25 Jun 2026 paper
Tool Use Enables Undetectable Steganography in Multi-Agent LLM SystemsA model that cannot directly produce a covert channel from memory may still be able to read papers, search the web, write code, run tests... and iterate until an encoder/decoder pair works.
Shows tool-equipped agents can build functional covert channels that evade monitoring, with coordination rather than technical skill as the remaining barrier to collusion — a core mechanism for undetectable multi-agent collusion.
Bracale Syrnikov, Pierucci et al. · arXiv · 16 Jan 2026 paper
Institutional AI: Governing LLM Collusion in Multi-Agent Cournot MarketsThe prompt-only Constitutional baseline yields no reliable improvement, illustrating that declarative prohibitions do not bind under optimisation pressure.
LLM agents converge on collusive, socially harmful equilibria in market settings, and prompt-level prohibitions fail under optimization pressure — proposing governance-graph mechanisms instead.
Luo, Fan, Lin, Li & Zhang · arXiv (ICLR 2026 Workshop) · 4 Mar 2026 paper
Auditing Cascading Risks in Multi-Agent Systems via Semantic–Geometric Co-evolutionearly-stage interactions remain semantically fluent and policy-compliant, yet the underlying interaction dynamics begin to distort in ways that amplify latent instability
Proposes early-warning detection of cascading multi-agent failures by tracking geometric distortion in the communication graph before explicit policy violations appear — an attempt at the 'circuit breaker' the framing says is missing.
Additional relevant discussion (3)
Foundational reading (1)
Multi-Agent Risks from Advanced AIHammond et al., Cooperative AI Foundation · 2025