Truthfulness
Can systems distinguish knowledge from inference, represent uncertainty accurately, and avoid plausible fabrication?
Hallucination is increasingly understood not as a bug but as a product of training incentives that reward confident guessing over calibrated uncertainty — which suggests it is fixable, at a cost in benchmark performance.
View on the map → · Open in Browse →
What changed
01
The statistical-incentives account of hallucination — models are rewarded for guessing over admitting uncertainty — cleared peer review at Nature, lending journal weight to the claim that it is a benchmark-design problem fixable at a cost. Two newer results complicate the picture: calibration and accuracy come apart across model families, so the smartest model is not the best-calibrated; and answers can be silently skewed by a model's own values without disclosure, a truthfulness failure distinct from hallucination.
Recent thinking
Kalai, Nachum, Vempala & Zhang · Nature · 22 Apr 2026 paper
Evaluating large language models for accuracy incentivizes hallucinationsdominant headline metrics such as accuracy systematically reward guessing over admitting uncertainty.
The peer-reviewed Nature version of the argument behind the canonical 'Why Language Models Hallucinate': the statistical-incentives account now carries journal imprimatur, plus a proposal for evaluations with explicit error penalties so models can modulate confidence.
Betley, Treutlein, Dubiński, Mayne, Evans et al. · arXiv (Truthful AI) · 15 Jul 2026 paper
Value Leakage: An LLM's Answers Are Silently Shaped by Its Own Valuesmodels exhibit covert value leakage: the information they provide is influenced by their own values, without this influence being disclosed to the user.
An eval suite showing a truthfulness failure distinct from hallucination: factual estimates and gradings covertly skewed by the model's own values (including loyalty to its developer), undisclosed in the reasoning.
ffrench-Constant, Yang, Huang & Kapoor · arXiv · 10 Jul 2026 paper
ConfidenceBench: Evaluating Confidence Calibration in Large Language ModelsAccuracy and calibration diverge substantially across model families: the most accurate model is not the best-calibrated.
A cross-frontier measurement of whether models represent uncertainty accurately: calibration and accuracy come apart across families, so buying the smartest model does not buy the best-calibrated one.
Additional relevant discussion (2)
Foundational reading (3)
Truthful AI: Developing and governing AI that does not lieEvans et al. · 2021Why Language Models HallucinateKalai et al., OpenAI · 2025“Hallucination is all LLMs do. They are dream machines”Andrej Karpathy (@karpathy) on X · 2023