The Biggest Questions About AI
The map · 4 Power · 4.1 Regulatory design · 4.1.2

Thresholds

Which measurable capabilities or resource levels should trigger licensing, evaluation, reporting, or access restrictions?

FLOP thresholds (10^25 in the EU AI Act, 10^26 in vetoed SB-1047) are proxies that decay as algorithms improve. Hooker's critique: they're already leaky, and capability-based triggers are hard to measure pre-deployment.

View on the map → · Open in Browse →

What changed
December 2025–August 2026 · swept August 3, 2026 · editorial review pending
01

The thresholds that exist got graded: SaferAI's systematic scoring of twelve company safety frameworks finds capability thresholds defined in vague terms with almost no quantitative criteria, scores topping out at 34%. The constructive work has begun — a first methodology for deriving harmonized cross-lab thresholds from expected-harm modeling.

Recent thinking
3 featured from 4 tracked · December 2025–August 2026 · all 4 chronologically →
Stelling, Murray, Galizzi, Schaffelder, Campos & Papadatos · arXiv (SaferAI) · 1 Dec 2025 paper
Evaluating AI Providers' Frontier Safety Frameworks
Overall scores range from 34% (Anthropic) to 8% (Cohere), with a median of 18%.

The most systematic grading yet of the twelve company safety frameworks that function as de facto capability thresholds: thresholds defined in vague terms ('severe harm', 'meaningful uplift') with almost no quantitative criteria or risk modeling linking them to residual-risk tolerances.

METR · 9 Dec 2025 report
Common Elements of Frontier AI Safety Policies (December 2025 Update)
The policies also outline commitments to conduct model evaluations assessing whether models are approaching capability thresholds that could enable severe or catastrophic harm.

The reference cross-company survey of which capability thresholds — bio, cyber, autonomous replication, automated AI R&D — currently trigger evaluation, security, and deployment commitments across the industry.

Anterola, Ball, Lafuerza & Grey · arXiv · 17 Jul 2026 paper
Harmonizing AI Safety Thresholds
We develop a methodology for deriving harmonized thresholds across three risk domains.

First serious methodology for deriving consistent cross-lab capability thresholds — expected-harm risk modeling for cyber and bio misuse, rate-of-progress triggers for automated AI R&D — answering the critique that current thresholds are ad hoc and incommensurable.

Additional relevant discussion (1)
Foundational reading (2)On the Limitations of Compute ThresholdsSara Hooker · 2024The Role of Compute Thresholds for AI GovernanceInstitute for Law & AI · 2024
Previous4.1.1 Regulatory objectNext4.1.3 Ex ante versus ex post