The Biggest Questions About AI
The map · 1 Trajectory · 1.4 Embodied and physical intelligence · 1.4.2

Sim-to-real transfer

Can learning from video and simulation overcome the scarcity and expense of physical-world training data?

Physical interaction data is scarce and costly to collect at internet scale. Whether video pretraining and simulation can substitute is the robotics analogue of the synthetic-data question — and similarly unresolved.

View on the map → · Open in Browse →

What changed
November 2025–August 2026 · swept August 3, 2026 · editorial review pending
01

The data question forked: Generalist's GEN-0 claims predictable scaling from real physical-interaction data collected at 10,000 hours a week, DeepMind's SIMA 2 self-improves inside generated worlds with no new human data, and a multi-lab position paper argues the real bottleneck is neither — it's the missing interfaces that would turn video, motion, and simulation into robot supervision.

Recent thinking
4 featured from 6 tracked · November 2025–August 2026 · all 6 chronologically →
Generalist AI · 4 Nov 2025 report
GEN-0: Embodied Foundation Models That Scale with Physical Interaction
GEN-0 marks the beginning of a new era: embodied foundation models whose capabilities predictably scale with physical interaction data.

The reference claim for 'scaling laws in robotics': 270,000 hours of real manipulation data yielding predictable performance scaling — the strongest empirical counter to the view that physical interaction data cannot be collected at scale.

Karcini, Mehrban, Nguyen et al. · arXiv · 4 Jun 2026 paper
Robots Need More than VLA and World Models
The central bottleneck is not only policy learning, but the absence of mechanisms that convert the world's abundant unstructured behavioural data into grounded robot supervision.

Multi-lab position paper arguing that scaling VLAs or world models alone will not produce generalist robots: the field lacks the data, embodiment, and reward interfaces to turn internet video, human motion, and simulation into robot supervision.

Google DeepMind · 13 Nov 2025 report
SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D Worlds
The skills it learned are some of the fundamental building blocks for the physical embodiment of intelligence needed for future AI assistants.

The live demonstration of the simulation path: a Gemini-based embodied agent that self-improves inside Genie 3-generated worlds without new human data, explicitly positioned as groundwork for physical robots.

Moritz Reuss · NVIDIA Technical Blog · 15 Jun 2026 essay
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action Models
It is likely that the winner is neither pure VLA nor pure WAM, but a hybrid of both.

Frames the 2026 architecture debate: world-action models built on video backbones that already model scene dynamics, versus VLM-derived vision-language-action models — whether video pretraining becomes the backbone of robot learning rather than a supplement.

Additional relevant discussion (2)
World Model for Robot Learning: A Comprehensive Survey — Hou, Li, Jia et al. · arXiv · 30 Apr 2026
Robot startups are trying everything they can think of to get more data — Kai Williams · Understanding AI (Timothy B. Lee) · 3 Sep 2026
Foundational reading (2)Sporks of AGISergey Levine · 2025Sim-to-Real Transfer in Deep RL for Robotics: a SurveyZhao et al. · 2020
Previous1.4.1 General-purpose roboticsNext1.4.3 Dexterity and robustness