The Biggest Questions About AI
The map · 1 Trajectory · 1.4 Embodied and physical intelligence · 1.4.1

General-purpose robotics

When will robots learn unfamiliar household, laboratory, industrial, or construction tasks from ordinary instructions?

Vision-language-action models are the current bet that robotics follows the LLM playbook: broad foundation models plus fine-tuning, rather than task-specific engineering. Early results are promising and narrow.

View on the map → · Open in Browse →

What changed
March–August 2026 · swept August 3, 2026 · editorial review pending
01

The foundation-model bet produced its strongest year: π0.7 composes learned skills to solve novel tasks and transfers across robot platforms, Gemini Robotics 2 adds whole-body humanoid control that adapts to new embodiments with under 200 examples, and Generalist claims the first crossing of a 'mastery' threshold on simple physical tasks — trained heavily on human wearable data rather than robot data.

Recent thinking
4 pieces · March–August 2026 · all 4 chronologically →
Physical Intelligence · 16 Apr 2026 report
π0.7: a Steerable Model with Emergent Capabilities
it can perform dexterous manipulation skills...with the same speed and robustness, it can compose and recombine the skills it learned to solve new tasks.

Successor to the canonical π0: a claimed step-change in generalization, with one model matching task-specific RL specialists, composing learned skills for novel tasks, and transferring to untrained robot platforms.

Google DeepMind · 30 Jul 2026 report
Gemini Robotics 2 brings whole body intelligence to robots
Gemini Robotics 2 marks an important milestone on the path toward solving AGI in the physical world.

Adds whole-body humanoid control (walking plus 22-DoF hand dexterity), multi-robot collaboration, on-device operation, and adaptation to new embodiments with under 200 examples — DeepMind explicitly framing robotics as physical AGI.

Generalist AI · 2 Apr 2026 report
GEN-1: Scaling Embodied Foundation Models to Mastery
We believe it to be the first general-purpose AI model that crosses a new performance threshold: mastery of simple physical tasks.

Claims 99% success where prior models hit 64%, ~3x human-teleoperation speed, and adaptation with about one hour of robot data per task — pretrained on 500,000+ hours of human wearable data containing no robot data at all.

Sergey Levine, with Patrick O'Shaughnessy · Invest Like the Best · 31 Mar 2026 podcast
Sergey Levine — Building LLMs for the Physical World
everyday human actions remain the hardest problems in the field

Physical Intelligence's co-founder makes the case that generality is more scalable than specialization for robot foundation models — and that human acceptance, not just capability, gates deployment into daily life.

Foundational reading (3)π0: A Vision-Language-Action Flow ModelPhysical Intelligence · 2024Gemini Robotics: Bringing AI into the Physical WorldGoogle DeepMind · 2025Announcing Project GR00T, a “moonshot to solve embodied AGI” via a humanoid foundation modelJim Fan (@DrJimFan) on X · 2024
Next1.4.2 Sim-to-real transfer