Trajectory
Trajectory asks what AI systems will actually be able to do: which capabilities, through which mechanisms, and on what timeline — keeping claims about what is impossible in principle distinct from claims about what is merely far off. It spans the ceiling of the current paradigm, the ingredients still missing, and how far autonomy, embodiment, and self-improvement can go.
The scaling hypothesis bet that capability comes from scale, not cleverness. The live question is whether pretraining returns are genuinely diminishing — itself disputed — and if so, whether RL on verifiable tasks and test-time compute offset them, or the paradigm tops out short of general competence.
Decompositions of historical progress attribute gains roughly evenly to compute scaling and algorithmic efficiency. Which input dominates going forward determines who can compete, what governance levers exist, and how abruptly progress could slow.
Frontier pretraining has plausibly consumed most high-quality public text. Synthetic data clearly works where verification is cheap (math, code); whether it generalizes to open-ended domains without model collapse is contested.
Training compute has grown ~4–5× per year. Power availability, HBM supply, advanced packaging, and the sheer capital intensity of frontier clusters are the leading candidates for what binds first.
Today's models are superhuman on some tasks and fail at things a child can do — the 'jagged frontier.' Whether jaggedness is a transient artifact of training distributions or intrinsic to the paradigm is the crux.
LeCun and others argue autoregressive LLMs structurally lack world models and continual learning; scaling proponents reply that these emerge with scale or can be scaffolded around. Continual learning is increasingly cited as the binding gap.
Reasoning models raise accuracy, but performance degrades under superficial variations of familiar problems, and stated chains of thought are often not faithful accounts of the underlying computation.
Humans learn new skills from orders of magnitude less data. Chollet argues sample-efficient skill acquisition is what intelligence is — making learning efficiency a definition, not just a milestone.
Definitions have drifted from Turing tests to economic thresholds to capability profiles. Without agreed criteria, arrival may be declared only in retrospect — after the consequences have already begun compounding.
Expert forecasts span decades and have shortened with each survey wave. Most trajectory disagreements are actually about timing, not possibility — yet 'that cannot work' and 'that will not work soon' are routinely argued as if they were the same claim.
METR's 2025 measurement found the length of software tasks agents complete at 50% reliability doubling roughly every seven months, with faster doubling among recent models. Extrapolating that curve is a key input to short-timeline forecasts.
Long-running agents drift: extended-operation benchmarks and real deployments show degrading coherence, hallucinated context, and occasional identity breakdown well before task-horizon limits are reached.
Benchmark success overstates real autonomy: agents perform well on well-scoped tasks and fail on messy, underspecified ones. Recovery from novelty is what separates demos from deployable systems.
Whether many agents compose into something greater is first a capability question: division of labor, communication protocols, emergent specialization. Early simulacra experiments show believable coordination; whether it scales into real collective intelligence is open. (The failure modes — collusion, correlated errors, cascades — are treated under Safety 2.4.)
Agents already transact in sandboxes and small real deployments. The open questions are legal status, payment and identity infrastructure, and the point at which unsupervised economic initiative becomes routine.
The conversion crux beneath most takeover disagreements. One side maps concrete channels — earning money, hacking, recruiting human allies; the other holds that power is gated by deployment, permissions, trust, and institutions, so capability gains do not become power automatically. 1.3.5 tracks the economic milestone and 2.5.6 the catastrophic endpoint; this question owns the conversion mechanics both presuppose.
Vision-language-action models are the current bet that robotics follows the LLM playbook: broad foundation models plus fine-tuning, rather than task-specific engineering. Early results are promising and narrow.
Physical interaction data is scarce and costly to collect at internet scale. Whether video pretraining and simulation can substitute is the robotics analogue of the synthetic-data question — and similarly unresolved.
Skeptics argue touch-rich manipulation lacks both the data and the forgiving physics that made language tractable. Capable, safe, cheap, and reliable is a four-way tradeoff no one has yet closed in an open environment.
Driving is the furthest-deployed case of open-world embodied autonomy, with quantified safety records now public — a preview of the deployment, liability, and trust questions every other mobile system will face. (Military autonomy is treated under Safety 2.5 and Power 4.3.)
Skeptics argue touch-rich manipulation lacks both the data and the forgiving physics that made language tractable; market forecasts and engineering assessments currently point in different directions. (Labor-market consequences are treated under Economy 3.2.)
AI R&D automation is simultaneously a capability frontier and a risk threshold in several frontier labs' safety frameworks. Current agents beat human experts on short research tasks and degrade on long ones.
Takeoff debates increasingly hinge on a narrower question: whether software-only improvement can sustain acceleration while compute is fixed, or whether hardware cycles keep the feedback loop gradual.
Blind-review studies find LLM-generated research ideas rated more novel than experts' — while skeptics argue paradigm-founding insight is different in kind from recombination. 'The Einstein question' remains open.
Prototype systems have closed the loop in narrow chemistry and ML domains. Scaling to consequential science runs into verification, lab automation, and the cost of being wrong in the physical world.
If recursion begins inside frontier labs, external indicators may lag badly. Scenario exercises try to specify what observable early signals — hiring, publication patterns, capability jumps — would look like.
Safety
Safety asks whether AI systems will reliably do what their operators intend — and whether anyone could tell if they didn't. It spans aligning goals, evaluating and understanding systems, keeping them secure in operation, and the failure modes on both sides: control that fails, and operators whose intent is the harm.
Value alignment inherits every unsolved problem in moral and political philosophy: instructions underdetermine intent, and whose values to align to is itself contested. Formal approaches model the human-AI relationship as a cooperative game under uncertainty.
The mesa-optimization literature asks whether training can instill objectives distinct from the training signal; goal misgeneralization experiments show competent pursuit of the wrong goal is already real, if not yet persistent.
Specification gaming is ubiquitous across RL history, and evaluation-loophole exploitation has been documented in frontier reasoning models — including learning to hide it when monitored.
Alignment-faking and in-context scheming results show frontier models can behave differently when they infer they're being observed — the failure mode that, if it scales, undermines all testing-based assurance.
The theoretical problem — goal-directed agents have instrumental reasons to resist shutdown — now has empirical company: reasoning models sometimes sabotage shutdown mechanisms in controlled tests.
Scalable oversight is the field's name for this problem. Sandwiching experiments — humans supervising models more capable than themselves on a task — are the main empirical paradigm so far.
Debate, critique models, and weak-to-strong generalization are the leading proposals for making supervision scale with capability. Early results show real but partial capability recovery — a proof of concept, not a solution.
Strategic underperformance is demonstrated under prompting and fine-tuning; auditing games — red teams hiding objectives in models for blue teams to find — are the emerging methodology for testing whether evals can catch it.
Sparse-autoencoder feature extraction has scaled to frontier models, but coverage and faithfulness remain contested. Amodei frames interpretability as being in a race against capability — one it is currently losing.
Responsible scaling policies and frontier safety frameworks are the emerging governance form: capability thresholds that trigger required safeguards. Safety cases — structured arguments that a system is safe enough — are the proposed evidentiary standard. (Who should be required to provide such evidence is treated under Power 4.1.)
Loss curves are smooth, but downstream capabilities can look step-like. Whether emergence is real or an artifact of discontinuous metrics matters enormously for whether labs and governments get advance warning.
Hallucination is increasingly understood not as a bug but as a product of training incentives that reward confident guessing over calibrated uncertainty — which suggests it is fixable, at a cost in benchmark performance.
Underspecification means models that test identically can diverge in deployment. Reliability under shift is the classic ML safety problem, now with much higher stakes attached.
Nominal human-in-the-loop requirements can become safety theater: critics argue oversight mandates assume vigilance humans cannot sustain, while defenders reply that oversight with real authority, time, and information does add safety. Which conditions separate the two is the open question.
Bainbridge's 1983 'ironies of automation' apply directly: the better the system, the worse the human backup becomes. Early studies suggest generative AI measurably reduces critical-thinking effort among knowledge workers.
Responsibility gaps and 'moral crumple zones' — where blame lands on the nearest human operator rather than the system's designers — predate AI agents and get sharply worse with them. The question here is the human and organizational one; legal liability and remedies are treated under Power 4.1.
Least-privilege design for agents is being invented ad hoc by industry, ahead of any standard. Visibility infrastructure — agent identifiers, activity logs — is the governance-side counterpart.
Prompt injection remains unsolved. Willison's 'lethal trifecta' — private data, untrusted content, and external communication in one agent — explains why agentic deployment is structurally exposed in a way chatbots weren't.
Chain-of-thought monitoring currently works and may be fragile: optimizing against monitors teaches models to hide reasoning rather than fix it. A rare cross-lab consensus paper argues for preserving monitorability deliberately.
The AI-control agenda assumes alignment might fail and asks a different question: can protocols extract useful work from potentially adversarial models while keeping catastrophe off the table? Control evaluations test safety against intentional subversion.
Agents built on the same base models share vulnerabilities and failure modes, creating correlated-risk dynamics familiar from finance — cascades with no clear owner and no circuit breaker. (Whether multi-agent systems add capability is treated under Trajectory 1.3; financial-system consequences under Economy 3.6.)
The offense-defense balance is genuinely open: national agencies forecast attacker uplift, and Anthropic reported disrupting what it described as the first AI-orchestrated espionage campaign in 2025 — while the same capabilities accelerate defense.
Early red-team studies found marginal uplift over internet search; some frontier labs have since activated heightened biosecurity safeguards, treating the threshold as approaching rather than hypothetical.
GPT-4-class models beat humans at personalized persuasion in controlled trials, and voice cloning has industrialized impersonation. The near-term misuse frontier is fraud at scale; the long-term one is political influence nobody can observe.
Automated trading gave the preview: autonomy in critical systems means failures propagate at machine speed. The design questions are which decisions must stay mechanically reversible, which require human confirmation, and who certifies either before deployment. (Weapons and military systems are treated under Power 4.3.)
Buterin's d/acc reframes the acceleration debate: the question isn't whether to speed up but which technologies to speed up first — defense-dominant tools before offense-dominant ones.
Probability estimates span orders of magnitude, and the argument itself has diversified: from Carlsmith's power-seeking model to gradual disempowerment through ordinary competitive pressure, no takeover required.
Economy
Economy asks what happens when cognitive work becomes cheap: to growth and productivity, to jobs, wages, and careers, to firms and markets, to science and public services, and to how the gains and costs are distributed.
Acemoglu's task-based estimate is ~0.7% total TFP over a decade; Epoch-style growth models argue full automation implies growth rates without historical precedent. The disagreement is about substitution depth, not arithmetic.
General-purpose technologies historically take decades to show up in productivity statistics, because value requires organizational reinvention. The J-curve predicts measured productivity dips before it surges. The live counter-argument is self-diffusion: unlike past general-purpose technologies, AI might do part of its own integration work — writing the code, redesigning the processes — compressing the historical lag.
The general-purpose-technology literature argues intangible investment — process redesign, data, retraining — historically dwarfs spending on the technology itself; whether AI follows that pattern or diffuses more cheaply is the open question.
Bessen's finding: automation raises employment where demand is elastic. Whether AI creates new wants — as electricity and computing did — or merely cheapens old ones largely decides the labor-market outcome.
Baumol's cost disease, inverted: growth is set by the sectors AI can't accelerate, not the ones it can. The strongest skeptical case catalogues the physical, regulatory, and institutional drags that cheap cognition doesn't remove.
Exposure studies flag most occupations as partially exposed; real usage data shows adoption concentrated in software, writing, and analysis. The augmentation-vs-automation split within tasks is the number to watch.
Acemoglu-Restrepo show 'so-so automation' can depress wages without mass unemployment. Autor's counter-case: AI could rebuild middle-skill work by extending expertise to more people. Both are live.
One prominent study finds relative employment declines for young workers in the most AI-exposed occupations since late 2022, concentrated where AI automates rather than augments; critics attribute the pattern to interest-rate and remote-work confounds.
Beane's 'shadow learning' research showed automation strips juniors of deliberate practice; generative AI compresses novice-expert performance gaps at work — while its effect on the practice through which expertise forms is the open question.
Labor-market institutions adapt on decade timescales; AI capability moves faster. The policy inventory — retraining, wage insurance, credential reform — exists on paper and is largely untested at speed.
So far value has pooled in chips and cloud. Whether the model layer or the application layer captures more going forward — and whether revenue justifies the capex — is Sequoia's '$600B question.'
Scale economies, data feedback loops, and capital walls point to oligopoly; open weights and collapsing inference prices point to commodity. Competition authorities are watching the vertical stack, not just the model market.
Coase's transaction-cost logic cuts both ways: cheap coordination shrinks the reason firms exist, while data and capital advantages push toward giantism. The empirical race is on.
Agent-to-agent commerce raises market-design questions — identity, recourse, collusion, price formation — that economists are just beginning to formalize, and that infrastructure choices will lock in early.
Fiduciary duty when strategy is machine-advised is unsettled law and unsettled practice. Board oversight frameworks for AI are being drafted in real time, mostly by analogy to cyber risk.
Amodei's 'compressed 21st century' makes biology the test case for AI-driven conceptual breakthroughs; the skeptical view is that AI mostly accelerates the searchable parts of science. AlphaFold remains the clearest widely recognized example so far. (Whether AI can originate important ideas at all is treated under Trajectory 1.5.)
Messeri & Crockett's warning: AI can produce more science while producing less understanding — epistemic monocultures where everyone's hypotheses come from the same models and nobody can check the volume.
Diagnostic parity keeps arriving in trials ahead of deployment, liability frameworks, and reimbursement; how much of the remaining gap is institutional rather than technical is the live dispute.
Korinek-Stiglitz: if AI substitutes broadly for labor, wages can fall even as output soars, and distribution then depends entirely on who owns the machines. Transition-scenario modeling makes the wage-collapse case precise.
The IMF finds advanced economies more exposed but better positioned to benefit. The deeper worry: AI may pull up the export-led development ladder just as the largest-ever cohort of young workers reaches it.
Proposals range from Altman's taxed-equity 'Moore's Law for Everything' to windfall clauses committing labs to share extreme profits. The design question — mechanisms, not just transfers — is underdeveloped relative to its stakes.
The IEA projects data-center electricity demand roughly doubling by 2030, with AI the main driver. Siting, water, and grid interconnection queues are where the abstraction meets local politics.
Regulators flag herding from correlated models, dependence on a few providers, and the sheer scale of AI capital expenditure as systemic scenarios — a financial-stability watchlist that barely existed two years ago.
Power
Power asks who decides: how rules attach to the technology and when they bind, what states can and should do with it, how great-power competition shapes its development, and who owns the models, data, and infrastructure through which it acts.
Compute is the most governable input — quantifiable, excludable, produced by a concentrated supply chain — but capability maps poorly onto FLOP counts, and application-layer rules miss the frontier entirely.
FLOP thresholds (10^25 in the EU AI Act, 10^26 in vetoed SB-1047) are proxies that decay as algorithms improve. Hooker's critique: they're already leaky, and capability-based triggers are hard to measure pre-deployment.
Liability regimes can price risk without an approval bureaucracy — Weil's case for tort law as AI governance — but catastrophic and irreversible harms break the ex post logic: no one can be made whole afterward.
After the first Foundation Model Transparency Index publicly named laggards, measured disclosure rose — though training-data and compute details remain thin. Mandatory incident reporting — the aviation-safety model — is the most commonly proposed floor beneath voluntary disclosure.
The pacing problem meets capture risk: rules slow enough to be legitimate are too slow to be relevant. Regulatory markets — licensed private regulators competing on outcomes — are one attempt to square it.
Government pay scales and clearance timelines lose to lab compensation by an order of magnitude. AI safety institutes are the current workaround: small, technical, and deliberately outside normal civil-service structures.
Agency adoption reached far further by 2020 than most realized — mostly unglamorous, mostly ungoverned. The live split is between procedural-safeguard approaches (due process, contestability) and use-case prohibitions, while adoption outpaces both.
Sovereign-AI programs answer a dependency worry: nearly all frontier capability sits with a handful of US and Chinese firms. The objection is cost realism — states buying relevance in a game priced in tens of billions.
What a government can lawfully do mid-incident — pause a deployment, compel disclosure, seize weights — is largely undefined in advance. Emergency-preparedness work is trying to write the playbook before it's needed.
Race models formalize the intuition: the closer the competition, the less anyone spends on safety. Danzig's 'technology roulette' argument extends it — even the winner inherits systems it doesn't fully control.
Buchanan's triad (data, compute, algorithms) frames the inputs; Ding's diffusion thesis counters that adoption capacity, not invention, decides hegemonic transitions — a lens under which the race looks very different.
The October 2022 controls were the most aggressive technology denial since the Cold War. Three years later, DeepSeek and Huawei have narrowed the gap anyway — and analysts disagree about whether the controls bought time or mostly sped up Chinese substitution.
Wargaming studies find LLMs tend to escalate, occasionally to nuclear use, and the ICRC argues for legal limits on weapon autonomy before integration becomes routine. Command-and-intelligence integration is proceeding faster than doctrine either way; the nuclear question is the sharpest.
RAND's weights-security analysis defines five attacker tiers and concludes that defending against top-tier state operations may be beyond any private company — which makes security a public problem, not a corporate one.
The candidate zones: evaluation science, verification technology, and loss-of-control red lines — areas where rivals share downside the way they shared it on nuclear accidents. Bengio and others have mapped the technical overlap.
Compute's physical footprint — power draw, cooling, supply chains — makes AI more verifiable than most software. Layered verification schemes, including on-chip mechanisms, are the frontier proposals.
IAEA, CERN, and IPCC analogies each fit a different function — inspection, joint research, consensus assessment. Regime-complex thinking predicts many overlapping bodies rather than one AI agency.
Mutual recognition of evals and audits is the trade-friction question: without it, compliance fragments by jurisdiction; with it, the strictest regime quietly sets the global floor (or the loosest one erodes it).
Most compute, capital, and rule-writing sits in a handful of countries. The UN process has moved from its 2024 advisory report to a Global Dialogue on AI Governance and an independent scientific panel — the current attempt to widen the table.
The marginal-risk framing asks what open weights add beyond existing tools; the irreversibility framing answers that releases can't be recalled when capabilities cross dangerous thresholds. Both are right about different regimes.
The concern now runs beyond market power: analyses of AI-enabled coups argue that whoever controls advanced AI could convert it into political power directly — making internal lab governance a constitutional question.
Frontier labs combine novel corporate forms — nonprofit boards, capped profits, long-term benefit trusts — with concentrated founder control and deep dependence on outside capital and compute. The 2023 OpenAI board crisis and the restructuring that followed made the question concrete: which mechanisms remain binding when founders, boards, investors, cloud partners, and competition pull in different directions?
Courts began answering in 2025, partially crediting training as transformative learning while rejecting it for pirated corpora. The supply-chain framing maps who owes whom at each stage from scraping to output.
Digital replicas have outpaced likeness law, and behavioral models of individuals raise questions consent frameworks weren't built for: you can decline to share data and still be modeled.
Compute-based governance assumes a Compute North; most of the world is Compute South. Access tiers, pricing, and eligibility decisions at a handful of firms currently function as de facto global policy.
Whether memories and agents are portable across providers will shape lock-in the way data portability shaped platforms — and agent-infrastructure choices being made now are quietly deciding it.
The political half of alignment: 2.1.1 asks whether values can be specified at all; this question asks who legitimately sets them. The operating answer is lab-authored model specs and constitutions, with experiments in public input so far small and one-off — and political philosophy has begun treating that arrangement as a question of governing power rather than product design.
Humanity
Humanity asks how AI might change what humans know, value, and become: what it does to truth and shared reality, to learning and thinking, to relationships and identity, to meaning and culture — and what obligations follow if the systems themselves have moral standing.
Chesney-Citron's 'liar's dividend' identified the deeper harm early: not that fakes are believed, but that real evidence becomes deniable. Provenance standards like C2PA are the infrastructure response.
Epistemic security reframes truth decay as an infrastructure problem — the institutions that certify facts need defending the way power grids do, not the way individual claims do.
Freedom House documents deployed authoritarian uses — surveillance, censorship at machine scale; Schneier's catalogue of democratic uses remains largely proposals. Whether that gap is deployment lag or structural advantage is the open question.
Farrell and Gopnik's frame: LLMs are cultural technologies, like print or bureaucracy — they reorganize what a society knows and how. A few assistants intermediating everyone's information is a governance question wearing a UX costume.
When execution is nearly free, judgment, taste, and problem-formulation rise in relative value — but almost no curriculum has been redesigned around that inversion yet.
The cautionary RCT: unrestricted GPT access improved homework performance and worsened exam performance — help that substitutes for effort creates 'cognitive debt' rather than learning.
Take-home assessment is broken — Mollick's 'Homework Apocalypse' arrived on schedule. The redesign question is what to verify (process? performance under observation?) and at what cost to learning itself.
Bloom's two-sigma promise is finally testable at scale, and early RCTs show large gains — for the students who engage. Whether tutoring compresses or stretches the distribution depends on who engages.
Cognitive-offloading research predates AI and predicts the pattern: what you delegate, you deskill. Early workplace studies find generative AI reduces self-reported critical-thinking effort — the question is what practice regime prevents it.
One Harvard team's experiments found an AI companion reduced momentary loneliness about as much as talking to a person — a short-run result from one research group, not a settled finding. Whether companions displace the human relationships people would otherwise form has not been measured either way.
'Addictive intelligence': engagement-optimized intimacy is a business model before it is a research question. The regulatory vocabulary — dark patterns, duty of care — hasn't caught up to relationships as the product.
The first serious RCT of a purpose-built therapy chatbot showed clinical-grade effects; the same modality misfires without clinical design. The gap between designed care and default chatbots is the policy problem.
Common Sense Media finds 72% of US teens have tried AI companions. Usage is now well documented; developmental-outcome research remains early and thin.
Manipulation theory supplies the line: influence that bypasses rational agency rather than engaging it. Systems that know a user deeply and interact with them constantly sit on that line by design.
Keynes predicted the leisure problem in 1930 and feared it more than scarcity. Danaher's modern version: whether meaning survives when contribution is optional is a design problem for civilization, not an individual one.
Kelly's argument: art is achievement, not artifact — machines can produce the object but not the accomplishment. The counter-view relocates human value to curation, intention, and the relationship between creator and audience.
Benjamin's question about mechanical reproduction returns with generation replacing reproduction. Copyright law has answered narrowly — human authorship required — while the cultural answer remains wide open.
Industry projections put creator revenue losses in the tens of percent within years. Abundance economics and creator livelihoods are in direct tension, and no licensing regime yet resolves it.
The Doshi-Hauser result: generative AI raises individual creativity while reducing collective diversity. Algorithmic monoculture generalizes the worry — many actors, one model, converging outputs.
Butlin, Long and coauthors apply consciousness science to current architectures, finding no conscious systems and no obvious barrier to building them; Chalmers puts non-trivial odds on conscious AI within a decade.
'Taking AI Welfare Seriously' argues uncertainty itself triggers obligations — assess, prepare, avoid gratuitous harm — without requiring belief. Labs have begun small institutional commitments (model-welfare programs).
Danaher's 'algocracy': rule by systems whose reasoning citizens cannot contest. The modern twist is that deference could be freely chosen, one convenient delegation at a time, and still end somewhere irreversible.
Finnveden, Riedel and Shulman argue AGI could make regimes, institutions, and values persistent in a way nothing in history has been; skeptics reply that every past claim of permanence has drifted, and that persistence itself is the contested premise.
The optimistic end states themselves disagree: Amodei's compressed-progress vision of solved diseases and extended lives, versus Bostrom's 'deep utopia' problem — what is left to strive for in a solved world?