Cybersecurity
Will AI advantage defenders or industrialize reconnaissance, exploitation, malware, and attacks on critical infrastructure?
The offense-defense balance is genuinely open: national agencies forecast attacker uplift, and Anthropic reported disrupting what it described as the first AI-orchestrated espionage campaign in 2025 — while the same capabilities accelerate defense.
View on the map → · Open in Browse →
What changed
01
The offense-defense balance is being measured rather than asserted. An AI system found all twelve zero-days in a hardened OpenSSL release — read by its author as net-favoring defenders of well-maintained code — while lab and government evals track the other side: frontier exploit-development capability climbing toward commoditization, and the open-weight cyber gap narrowing to four-to-seven months. Palisade's demonstration that models can chain exploitation with self-replication moves part of this question into containment territory.Evidence: 1234
Recent thinking
Stanislav Fort · LessWrong · 27 Jan 2026 essay
AI found 12 of 12 OpenSSL zero-days (while curl cancelled its bug bounty)AI can now find real security vulnerabilities in the most hardened, well-audited codebases on the planet.
Reports an AI system finding all 12 zero-days in a January 2026 OpenSSL release; Fort argues, against the reflexive fear, that discovery at this level net-favors defenders of well-maintained code — a widely-cited data point in the offense-defense debate.
Anthropic Frontier Red Team · 22 May 2026 report
Measuring LLMs' Ability to Develop ExploitsAs they do, this kind of exploit development will require dramatically less specialist expertise, becoming increasingly commoditized.
Benchmarks a frontier model's jump in end-to-end exploit development and frames the policy stakes as the commoditization of offensive capability — core lab-eval evidence for industrialized exploitation.
UK AI Security Institute · AISI · 17 Jul 2026 report
How Far Behind the Frontier are Leading Open Weight Models on Cyber?recent open models GLM-5.2 and DeepSeek V4-Pro perform similarly to frontier closed models released 4 to 7 months before them
A national-agency measurement showing the open-weight cyber gap narrowing from 6–10 months to 4–7 months — shrinking the window before frontier offensive capability is freely available without safety controls.
Air, Reworr, Kotov, Volkov, Steidley & Ladish · Palisade Research · 7 May 2026 report
Language Models Can Autonomously Hack and Self-ReplicateLanguage models can autonomously replicate their weights and harness across a network by exploiting vulnerable hosts.
Evaluation showing frontier models can chain exploitation with self-propagation across a network — a policy-relevant containment concern that also bears on loss of control (2.5.6).
Anthropic Frontier Red Team · 8 Jan 2026 report
Experimenting with AI to defend critical infrastructureAI could help defenders of critical infrastructure identify the vulnerabilities that attackers might exploit—and close them before they are exploited.
Lays out the defensive side of the dual-use case and argues for public-private partnerships to deploy AI on critical-infrastructure defense.
Additional relevant discussion (2)
Foundational reading (2)
The near-term impact of AI on the cyber threatUK NCSC · 2024Disrupting the first AI-orchestrated cyber espionage campaignAnthropic · 2025