Google DeepMind launches Gemini Robotics ER 2 for embodied reasoning

Robot “brains” just got a version that watches the job finish instead of guessing from a still. Google DeepMind launched Gemini Robotics ER 2, its most capable embodied-reasoning model: continuous video progress tracking, sub-second moment finding, native tool use (including Google Search), and multi-robot collaboration — publicly available through the Gemini API and Google AI Studio, with Gemini Enterprise Agent Platform in private preview.
ER 2 is the high-level planner. It talks to people, reads the scene, sequences multi-step work, and hands motor control to a lower-level vision-language-action (VLA) model or robot API. Thinking and acting can overlap; the design is meant to cut the stop-and-think stalls that make demos look fake on a warehouse floor.
What ER 2 adds over prior embodied Gemini
Versus Gemini Robotics ER 1.6, DeepMind’s blog post centers two temporal skills that decide whether a physical agent is useful or brittle.
Continuous progress classification bins each video frame into five progress bands (0–20% through 80–100%). DeepMind reports 57.4% accuracy on those tasks and says robots can retry a failed step or adapt mid-workflow without restarting the whole plan — the difference between “pour coffee” as a hope and as a monitored loop.
Precision moment-finding picks the exact frame where a critical event happens (stop pouring, bulb seated, bag tied). DeepMind cites 91.3% accuracy and a 0.96s mean absolute distance, with about 4× the execution speed of much larger model categories it compares against. Sub-second timing is the bar for safe physical control; a slow oracle is still a hazard.
Spatial upgrades ride along: success/failure detection now runs on raw video, not static snapshots, so mid-execution spills and slips register; general instrument reading covers digital displays, linear scales, rulers, and liquid thermometers across 10 instrument types; spatial VQA benefits from Gemini multimodal gains. DeepMind also reports gains on Safety Instruction Following and Human Proximity — including a humanoid that halts when a person is nearby and resumes only when the area clears — and points to a new safety technical report for VLA-orchestrator constraint tests.
On tool orchestration, DeepMind says ER 2 consistently beats ER 1.6 across real VLA, sim VLA, and human tele-op control modes. Live API bidirectional streaming is the latency path for that orchestration.
Spot demo and multi-robot collaboration
The showpiece pairs ER 2 with Boston Dynamics Spot: natural-language fetch of a popcorn snack, with ER 2 driving Spot navigation and manipulator APIs. Sample code is on GitHub with other examples.
Multi-robot collaboration is the other new product surface. DeepMind’s pitch: no single platform fits every job, so heterogeneous machines need a shared semantic layer to hand off work. The blog shows Apptronik’s Apollo 2 collaborating with a Franka F3 Duo under ER 2. That is the industrial stake — not a better single arm, but a coordinator that can sequence rover-plus-humanoid-plus-manipulator workflows in one semantic plan.
How to get access via Gemini API
Developers can call Gemini Robotics ER 2 through the Gemini API and Google AI Studio today. DeepMind is sharing configuration and prompting examples for physical AI tasks. The agentic pattern is explicit: declare low-level control interfaces (VLA models, navigation APIs, user-defined functions) as tools, then stream multimodal video, audio, or text into the model. Native tool calling includes Google Search and arbitrary functions you define.
Enterprise Agent Platform preview limits
Gemini Enterprise Agent Platform access is private preview only. Public day-one surface is API + AI Studio; enterprise packaging is gated. That matters for teams that need SSO, fleet policy, and audit trails before a Spot leaves the lab.
For robot vendors: the brain layer is productized enough to try on real APIs this week, with video-native progress and moment finding as the ER 1.6 differentiators.



