The Biggest Questions About AI
The map · 1 Trajectory · 1.2 General intelligence and missing capabilities · 1.2.4

Learning efficiency

Can AI learn new domains from small amounts of experience as effectively as capable humans?

Humans learn new skills from orders of magnitude less data. Chollet argues sample-efficient skill acquisition is what intelligence is — making learning efficiency a definition, not just a milestone.

View on the map → · Open in Browse →

What changed
November 2025–August 2026 · swept August 3, 2026 · editorial review pending
01

Sutskever made sample efficiency the headline problem — models 'generalize dramatically worse than people' — while ARC's new interactive benchmark measured the gap at its starkest: humans succeed at learning its games 100% of the time, frontier systems below 1%. The rebuttal worth reading argues most of the gap dissolves under fair accounting.

Recent thinking
4 featured from 5 tracked · November 2025–August 2026 · all 5 chronologically →
Ilya Sutskever, with Dwarkesh Patel · Dwarkesh Podcast · 25 Nov 2025 podcast
Ilya Sutskever — We're moving from the age of scaling to the age of research
These models somehow just generalize dramatically worse than people. It's a very fundamental thing.

Sutskever's first extended interview since founding SSI names poor generalization and sample efficiency as THE central unsolved problem and declares the scaling era over — the most-cited insider statement on learning efficiency in this window.

ARC Prize Foundation · arXiv · 24 Mar 2026 paper
ARC-AGI-3: A New Challenge for Frontier Agentic Intelligence
evaluating fluid adaptive efficiency on novel tasks, while avoiding language and external knowledge

Chollet's program made interactive: agents must learn game environments from scratch by exploration. As of March 2026 humans succeed 100% of the time; frontier systems score below 1% — the starkest current measurement of the human–AI gap in sample-efficient skill acquisition.

Samuel Knoche · LessWrong · 1 Jun 2026 essay
Dissolving the Deep Learning Sample Efficiency Gap
Most of the apparent inefficiency dissolves on closer inspection: apples-to-oranges comparisons between pretrained humans and from-scratch networks... and priors installed by evolution.

The strongest recent rebuttal to the sample-efficiency critique: the human–AI efficiency gap is mostly a measurement artifact of unfair comparisons, and what remains implies more compute per example rather than a missing breakthrough.

Chollet, Knoop, Kamradt & Landers · arXiv · 15 Jan 2026 paper
ARC Prize 2025: Technical Report
the top score reaching 24% on the ARC-AGI-2 private evaluation set.

The definitive annual measurement of efficiency-constrained novel-task learning: compute-limited systems reach only 24% on ARC-AGI-2 tasks humans solve routinely, quantifying how much recent progress is bought with brute-force test-time compute.

Additional relevant discussion (1)
SPADE: Self-Play in Adaptive Synthetic Executable Environments — Bo Liu et al. · arXiv · 19 Aug 2026
Foundational reading (2)On the Measure of IntelligenceFrançois Chollet · 2019Building Machines That Learn and Think Like PeopleLake et al. · 2016
Previous1.2.3 Reasoning reliabilityNext1.2.5 AGI recognition