The Biggest Questions About AI
The map · 1 Trajectory · 1.2 General intelligence and missing capabilities · 1.2.3

Reasoning reliability

Can systems reason accurately outside familiar patterns rather than merely producing convincing-looking chains of thought?

Reasoning models raise accuracy, but performance degrades under superficial variations of familiar problems, and stated chains of thought are often not faithful accounts of the underlying computation.

View on the map → · Open in Browse →

What changed
February–August 2026 · swept August 3, 2026 · editorial review pending
01

Competition mathematics stopped being the test: multiple models scored perfect marks on unseen IMO 2026 problems under official judging, and the benchmark group that guards against contamination declared final-answer problems saturated — while locating what remains unreliable in proof-writing, flawed-problem detection, and research mathematics. The perturbation literature keeps finding cliff-edge failures underneath the headline scores.

Recent thinking
3 featured from 4 tracked · February–August 2026 · all 4 chronologically →
AFP · Taipei Times · 24 Jul 2026 news
AI models score 100 percent at top math competition
Previously, no large language model had ever achieved a perfect score under the IMO's official judging process.

Multiple models scored perfect marks on unseen IMO 2026 problems under official judging — the strongest counter-evidence to date that reasoning fails outside familiar patterns, at least in competition mathematics.

Sun, Dekoninck & Vechev · MathArena (ETH Zurich) · 12 May 2026 report
Farewell to Final-Answer Competition Problems as Frontier Benchmarks
mathematical reasoning remains far from solved

The uncontaminated-competition eval group declares final-answer problems saturated while showing proof-writing, flawed-problem detection, and research-level mathematics remain unreliable — precisely locating where reliability now breaks down.

Es-sebbani, Marquer, Salhi & Bouraoui · arXiv · 13 Feb 2026 paper
Evaluating Robustness of Reasoning Models on Parameterized Logical Problems
we observe sharp performance transitions under targeted structural interventions even when surface statistics are held fixed

Controlled-perturbation methodology extended to logical satisfiability: reasoning models show cliff-edge failures under structural (not surface) interventions — benchmark accuracy conceals brittleness regimes.

Additional relevant discussion (1)
Foundational reading (2)GSM-Symbolic: Limitations of Mathematical Reasoning in LLMsMirzadeh et al., Apple · 2024Reasoning models don't always say what they thinkAnthropic · 2025
Previous1.2.2 Missing ingredientsNext1.2.4 Learning efficiency