Scaling to 1T parameters significantly enhances sample efficiency and performance ceilings
Empirical evidence that RL on base models keeps yielding gains at trillion-parameter scale — the case for headroom in the current recipe.
How far can transformers, reinforcement learning, tool use, and inference-time reasoning progress without a fundamentally new architecture?
The scaling hypothesis bet that capability comes from scale, not cleverness. The live question is whether pretraining returns are genuinely diminishing — itself disputed — and if so, whether RL on verifiable tasks and test-time compute offset them, or the paradigm tops out short of general competence.
View on the map → · Open in Browse →
DeepMind formalized the routes past AGI while a startup demonstrated current-paradigm agents automating narrow AI research. The continual-learning dispute stands: Patel and Lee argue frozen weights gate discovery; a trillion-parameter RL result says the recipe still has headroom. The winter opened a new front: Hooker's 'slow death of scaling' case that the compute-performance link is breaking down, and Ord's estimate that RL scaling is an order of magnitude less efficient than pretraining — with rebuttals arguing both confuse algorithmic progress with the failure of scale.Evidence: 1234567
Scaling, RL, and inference-time reasoning keep delivering; no new ingredient required yet.
Scaling to 1T parameters significantly enhances sample efficiency and performance ceilings
Empirical evidence that RL on base models keeps yielding gains at trillion-parameter scale — the case for headroom in the current recipe.
Documents and pushes back on Amodei's case that the scaling recipe holds and continual learning isn't a required ingredient.
DeepMind's formal map of four routes past AGI — scaling, paradigm shifts, recursive improvement, multi-agent collectives — and the frictions on each.
Current-paradigm agents already automate parts of AI research (SOTA on NanoGPT-class tasks) — but only on well-defined, quickly measurable goals, with reward hacking a persistent limit.
Continual, on-the-job learning (or another missing capability) gates further progress.
Around 30-50% of a lab's compute goes to inference, and that compute is currently not really doing anything productive in helping improve the model.
Argues the current paradigm wastes deployment experience and that on-the-job learning, not further scaling, is the next necessary breakthrough.
Weight updates are probably needed for some parts of effective CL since LLMs seem quite bad at handling lots of interrelated complexity in their context window
Argues in-context learning alone is not enough for continual learning, implying a real gap in the current architecture.
Once a model is trained, its weights are frozen and its capacity to learn new patterns is greatly reduced.
Frozen weights and lossy file-based memory prevent agents from accumulating the tacit knowledge discovery requires.
Measured scaling returns are degrading; RL buys less than pretraining did.
a 10x scaling of RL is required to get the same performance boost as a 3x scaling of inference ... we may have lost the ability to effectively turn more compute into more intelligence.
A quantitative case that RL compute scaling is far less efficient than pretraining or inference scaling, and that most RL gains come from enabling longer chains of thought — directly challenging the view that RL offsets diminishing pretraining returns.
The relationship between training compute and performance is highly uncertain and rapidly changing.
The reference point for the winter 2025–26 'is scaling dead' debate: the compute-performance link is breaking down and progress is shifting to other levers. Kirsch's rebuttal (in the ledger) argues small-beats-large examples reflect newer generations, not the failure of scale.