The Biggest Questions About AI
The map · 2 Safety · 2.5 Misuse and catastrophic risk · 2.5.2

Biology and chemistry

At what point does AI materially lower the expertise, time, or cost required to create dangerous biological or chemical agents?

Early red-team studies found marginal uplift over internet search; some frontier labs have since activated heightened biosecurity safeguards, treating the threshold as approaching rather than hypothetical.

View on the map → · Open in Browse →

What changed
February–August 2026 · swept August 3, 2026 · editorial review pending
01

The bio-uplift debate is now anchored in measurement and structured forecasting rather than assertion. A large expert-and-superforecaster elicitation estimates how much LLMs shift the risk and finds that combined safeguards — synthesis screening plus anti-jailbreak measures — can bring it near baseline, while the controlled trials to date show only modest novice uplift on hands-on lab procedures. A calibration check found domain-expert virologists substantially over-predicted wet-lab uplift; the governance literature warns screening may not survive AI-designed novel sequences. Tracked here at the level of evals, expert assessment, and policy — never operational detail.Evidence: 1234

Recent thinking
5 featured from 6 tracked · February–August 2026 · all 6 chronologically →
Williams, Righetti, Rosenberg et al. (incl. Philip E. Tetlock) · Forecasting Research Institute · Apr 2026 report
Forecasting LLM-enabled Biorisk and the Efficacy of Safeguards
The median expert predicted a 0.3% baseline annual risk of a human-caused epidemic that causes 100,000 deaths.

Major structured elicitation estimating how LLMs shift biorisk and finding combined safeguards — synthesis screening plus model anti-jailbreak measures — can bring risk near baseline. The central policy-relevant reference on uplift and mitigation.

Humam Aziz · EA Forum · 30 May 2026 essay
The State of Bio-Uplift Research in Mid-2026
AI bio-uplift exists. Novices have the most to gain, while experts and individuals with some biology experience likely reach the farthest given access to LLMs.

A landscape review synthesizing the uplift-study literature and flagging the missing expert wet-lab studies and weak biosecurity assessments for open-weight models — a map of where the evidence stands.

Nikki Teran · Nuclear Threat Initiative · 13 May 2026 report
AIxBio Horizon Scan: Spring 2026
Capability development continues to outpace the governance frameworks needed to manage it, and the window for establishing effective oversight mechanisms continues to narrow.

A policy scan tracking DNA-synthesis-screening legislation and warning that homology-based screening may fail against AI-designed novel sequences — framing the governance gap for policymakers.

Hong, Kleinman, Mathiowetz et al. · arXiv · 18 Feb 2026 paper
Measuring Mid-2025 LLM-Assistance on Novice Performance in Biology
Overall, mid-2025 LLMs did not substantially increase novice completion of complex laboratory procedures but were associated with a modest performance benefit.

A primary RCT measuring novice lab-task uplift — one of the studies the mid-2026 debate rests on — finding only modest benefit on hands-on procedures. Tracked at measurement level only.

Parry, Williams, Devlin-Foltz & Reynolds · Forecasting Research Institute · 19 Feb 2026 report
How Well Did Superforecasters and Experts Predict Wet Lab Skill Uplift from LLMs?
Superforecasters were more accurate at predicting the RCT in absolute terms than experts, particularly compared to virologists.

A calibration check finding domain-expert virologists substantially over-predicted LLM wet-lab uplift while superforecasters were closer to measured results — relevant to how much weight to give expert alarm.

Additional relevant discussion (1)
Foundational reading (2)The Operational Risks of AI in Large-Scale Biological AttacksRAND · 2024Building an early warning system for LLM-aided biological threat creationOpenAI · 2024
Previous2.5.1 CybersecurityNext2.5.3 Manipulation and fraud