Reasoning reliability
Can systems reason accurately outside familiar patterns rather than merely producing convincing-looking chains of thought?
Reasoning models raise accuracy, but performance degrades under superficial variations of familiar problems, and stated chains of thought are often not faithful accounts of the underlying computation.
View on the map → · Open in Browse →
What changed
01
Competition mathematics stopped being the test: multiple models scored perfect marks on unseen IMO 2026 problems under official judging, and the benchmark group that guards against contamination declared final-answer problems saturated — while locating what remains unreliable in proof-writing, flawed-problem detection, and research mathematics. The perturbation literature keeps finding cliff-edge failures underneath the headline scores.
Recent thinking
AFP · Taipei Times · 24 Jul 2026 news
AI models score 100 percent at top math competitionPreviously, no large language model had ever achieved a perfect score under the IMO's official judging process.
Multiple models scored perfect marks on unseen IMO 2026 problems under official judging — the strongest counter-evidence to date that reasoning fails outside familiar patterns, at least in competition mathematics.
Sun, Dekoninck & Vechev · MathArena (ETH Zurich) · 12 May 2026 report
Farewell to Final-Answer Competition Problems as Frontier Benchmarksmathematical reasoning remains far from solved
The uncontaminated-competition eval group declares final-answer problems saturated while showing proof-writing, flawed-problem detection, and research-level mathematics remain unreliable — precisely locating where reliability now breaks down.
Es-sebbani, Marquer, Salhi & Bouraoui · arXiv · 13 Feb 2026 paper
Evaluating Robustness of Reasoning Models on Parameterized Logical Problemswe observe sharp performance transitions under targeted structural interventions even when surface statistics are held fixed
Controlled-perturbation methodology extended to logical satisfiability: reasoning models show cliff-edge failures under structural (not surface) interventions — benchmark accuracy conceals brittleness regimes.
Additional relevant discussion (1)
Foundational reading (2)
GSM-Symbolic: Limitations of Mathematical Reasoning in LLMsMirzadeh et al., Apple · 2024Reasoning models don't always say what they thinkAnthropic · 2025