Biology and chemistry
At what point does AI materially lower the expertise, time, or cost required to create dangerous biological or chemical agents?
Early red-team studies found marginal uplift over internet search; some frontier labs have since activated heightened biosecurity safeguards, treating the threshold as approaching rather than hypothetical.
View on the map → · Open in Browse →
What changed
01
The bio-uplift debate is now anchored in measurement and structured forecasting rather than assertion. A large expert-and-superforecaster elicitation estimates how much LLMs shift the risk and finds that combined safeguards — synthesis screening plus anti-jailbreak measures — can bring it near baseline, while the controlled trials to date show only modest novice uplift on hands-on lab procedures. A calibration check found domain-expert virologists substantially over-predicted wet-lab uplift; the governance literature warns screening may not survive AI-designed novel sequences. Tracked here at the level of evals, expert assessment, and policy — never operational detail.Evidence: 1234
Recent thinking
Williams, Righetti, Rosenberg et al. (incl. Philip E. Tetlock) · Forecasting Research Institute · Apr 2026 report
Forecasting LLM-enabled Biorisk and the Efficacy of SafeguardsThe median expert predicted a 0.3% baseline annual risk of a human-caused epidemic that causes 100,000 deaths.
Major structured elicitation estimating how LLMs shift biorisk and finding combined safeguards — synthesis screening plus model anti-jailbreak measures — can bring risk near baseline. The central policy-relevant reference on uplift and mitigation.
Humam Aziz · EA Forum · 30 May 2026 essay
The State of Bio-Uplift Research in Mid-2026AI bio-uplift exists. Novices have the most to gain, while experts and individuals with some biology experience likely reach the farthest given access to LLMs.
A landscape review synthesizing the uplift-study literature and flagging the missing expert wet-lab studies and weak biosecurity assessments for open-weight models — a map of where the evidence stands.
Nikki Teran · Nuclear Threat Initiative · 13 May 2026 report
AIxBio Horizon Scan: Spring 2026Capability development continues to outpace the governance frameworks needed to manage it, and the window for establishing effective oversight mechanisms continues to narrow.
A policy scan tracking DNA-synthesis-screening legislation and warning that homology-based screening may fail against AI-designed novel sequences — framing the governance gap for policymakers.
Hong, Kleinman, Mathiowetz et al. · arXiv · 18 Feb 2026 paper
Measuring Mid-2025 LLM-Assistance on Novice Performance in BiologyOverall, mid-2025 LLMs did not substantially increase novice completion of complex laboratory procedures but were associated with a modest performance benefit.
A primary RCT measuring novice lab-task uplift — one of the studies the mid-2026 debate rests on — finding only modest benefit on hands-on procedures. Tracked at measurement level only.
Parry, Williams, Devlin-Foltz & Reynolds · Forecasting Research Institute · 19 Feb 2026 report
How Well Did Superforecasters and Experts Predict Wet Lab Skill Uplift from LLMs?Superforecasters were more accurate at predicting the RCT in absolute terms than experts, particularly compared to virologists.
A calibration check finding domain-expert virologists substantially over-predicted LLM wet-lab uplift while superforecasters were closer to measured results — relevant to how much weight to give expert alarm.
Additional relevant discussion (1)
Foundational reading (2)
The Operational Risks of AI in Large-Scale Biological AttacksRAND · 2024Building an early warning system for LLM-aided biological threat creationOpenAI · 2024