Sim-to-real transfer
Can learning from video and simulation overcome the scarcity and expense of physical-world training data?
Physical interaction data is scarce and costly to collect at internet scale. Whether video pretraining and simulation can substitute is the robotics analogue of the synthetic-data question — and similarly unresolved.
View on the map → · Open in Browse →
What changed
01
The data question forked: Generalist's GEN-0 claims predictable scaling from real physical-interaction data collected at 10,000 hours a week, DeepMind's SIMA 2 self-improves inside generated worlds with no new human data, and a multi-lab position paper argues the real bottleneck is neither — it's the missing interfaces that would turn video, motion, and simulation into robot supervision.
Recent thinking
Generalist AI · 4 Nov 2025 report
GEN-0: Embodied Foundation Models That Scale with Physical InteractionGEN-0 marks the beginning of a new era: embodied foundation models whose capabilities predictably scale with physical interaction data.
The reference claim for 'scaling laws in robotics': 270,000 hours of real manipulation data yielding predictable performance scaling — the strongest empirical counter to the view that physical interaction data cannot be collected at scale.
Karcini, Mehrban, Nguyen et al. · arXiv · 4 Jun 2026 paper
Robots Need More than VLA and World ModelsThe central bottleneck is not only policy learning, but the absence of mechanisms that convert the world's abundant unstructured behavioural data into grounded robot supervision.
Multi-lab position paper arguing that scaling VLAs or world models alone will not produce generalist robots: the field lacks the data, embodiment, and reward interfaces to turn internet video, human motion, and simulation into robot supervision.
Google DeepMind · 13 Nov 2025 report
SIMA 2: An Agent that Plays, Reasons, and Learns With You in Virtual 3D WorldsThe skills it learned are some of the fundamental building blocks for the physical embodiment of intelligence needed for future AI assistants.
The live demonstration of the simulation path: a Gemini-based embodied agent that self-improves inside Genie 3-generated worlds without new human data, explicitly positioned as groundwork for physical robots.
Moritz Reuss · NVIDIA Technical Blog · 15 Jun 2026 essay
Pretrained to Imagine, Fine-Tuned to Act: The Rise of World-Action ModelsIt is likely that the winner is neither pure VLA nor pure WAM, but a hybrid of both.
Frames the 2026 architecture debate: world-action models built on video backbones that already model scene dynamics, versus VLM-derived vision-language-action models — whether video pretraining becomes the backbone of robot learning rather than a supplement.
Additional relevant discussion (2)
Foundational reading (2)
Sporks of AGISergey Levine · 2025Sim-to-Real Transfer in Deep RL for Robotics: a SurveyZhao et al. · 2020