Google DeepMind Launches Gemini Robotics ER 2, Advancing Robot Reasoning and Multi-Tasking

ER 2 achieves 91.3% accuracy on moment-finding benchmarks with a 0.96-second mean absolute distance, enabling precise timing for actions such as when to stop pouring.
Progress classification across five completion bands reaches 57.4% accuracy in evaluations, indicating how tasks are categorized as they proceed.
Gemini Robotics 2 enables five-fingered, 22-degree-of-freedom hand dexterity, expanding manipulation capabilities beyond earlier versions.
The VLA model in Gemini Robotics 2 translates vision and language inputs into motor commands, enabling full-body control (feet to fingertips) and tasks like fetching from high shelves.
Developers can build the agentic pattern by declaring low-level control interfaces (such as VLA models or navigation APIs) as tools and feeding streaming video, audio, or text directly into the model to orchestrate tasks.
Google DeepMind launched Gemini Robotics ER 2 on July 30, 2026 — its most capable reasoning model for robots yet. The system lets robots chat with people, plan multi-step tasks, and direct other AI models to move their bodies, all without stopping to think between steps, according to TechnoBeZZ.
Alongside ER 2, DeepMind unveiled the full Gemini Robotics 2 stack — three models that together give robots full-body control from feet to fingertips. The suite includes a vision-language-action model, the new embodied reasoning model, and an on-device version that works without internet, Northeast Times reported.
Most AI models stop and think before acting. ER 2 does both at once. It reasons about the next step while the robot is still moving, so actions flow smoothly instead of in jerky, stop-start bursts. Developers can plug in lower-level control models — called vision-language-action models, or VLAs — as tools, then feed live video, audio, or text straight into ER 2 to coordinate the whole task, according to The New Stack.
In one demo, ER 2 directed a Boston Dynamics Spot robot to fetch objects based on natural-language requests. The system tracked progress in real time using continuous video understanding — meaning it watched what the robot was doing and adjusted on the fly, TechnoBeZZ reported.
Knowing exactly when to act is critical for robots. Pour water too long and the glass overflows. ER 2 hits 91.3% accuracy on moment-finding benchmarks — figuring out the precise instant to stop an action — with a mean error of just 0.96 seconds, according to TechnoBeZZ. That kind of timing matters for real-world tasks like pouring, cutting, or handing objects to people.
The model also classifies how far along a task is — sorting progress into five bands — with 57.4% accuracy. That lets ER 2 know whether a job is barely started or nearly done, and plan the next move accordingly. The system can even coordinate multiple robots at once through shared understanding of the same environment, The New Stack reported.
The VLA model in Gemini Robotics 2 turns camera images and voice commands into motor signals for the whole body. That includes five-fingered hands with 22 degrees of freedom — far more range of motion than earlier robot hands. Tasks like reaching objects on high shelves, which require coordinating arms, legs, and grip at the same time, are now possible, according to The Decoder.
DeepMind showed off the system using Apptronik Apollo 2 humanoid robots paired with Franka F3 Duo robotic arms. The pairing shows how the technology could handle real physical assistance jobs, TechMymoney noted. The On-Device 2 model runs all of this locally — no internet connection needed — making it practical for factories, homes, or anywhere connectivity is unreliable.
Gemini Robotics 2 is designed to be a common intelligence layer across many robot shapes and sizes — from small tabletop arms to full humanoids. The Decoder reported that the VLA model combines image recognition, language understanding, and action control in a single system. That means one AI brain can run very different robots without being rebuilt from scratch for each one.
Google DeepMind called Gemini Robotics 2 "the intelligence layer for the next generation of adaptable robots," according to Northeast Times. The launch marks a push toward robots that can work alongside people in complex, unstructured environments — not just controlled factory floors. Whether the technology reaches everyday use at scale remains to be seen.
Publishers
16
Articles
11
Reach
27