Summary
On July 30, 2026, Google DeepMind introduced the Gemini Robotics 2 model family, which combined embodied reasoning with action models for humanoid, dual-arm, and smaller research robots. Google said that, for the first time in its robotics program, the system could control an entire humanoid body rather than concentrating on upper-body manipulation. The reasoning model became available through Google AI Studio, while the principal action models remained restricted to early-access partners and trusted testers.
What Happened
The release contained three related systems. Gemini Robotics 2 was a cloud-based vision-language-action model for translating instructions into robot movement, Gemini Robotics-ER 2 handled embodied reasoning and task planning, and Gemini Robotics On-Device 2 was designed for lower-latency local execution. Google described the family as a shared approach to controlling different robot embodiments rather than a model tied to one machine.
On Apptronik's Apollo 2 humanoid, Google demonstrated walking, crouching, stretching, reaching, and object manipulation with both Inspire and SharpaWave multi-finger hands. The company also tested the models on the Franka Duo dual-arm platform and used SO101, Dexmate, and Trossen hardware in on-device adaptation work. Apptronik, Boston Dynamics, and Agile Robots participated as hardware or testing partners.
Google reported that the system could execute multistep tasks lasting several minutes and involving hundreds of decisions. Gemini Robotics-ER 2 interpreted scenes, planned actions, monitored progress, and revised a plan after failures, while demonstrations also showed different robot types coordinating parts of a shared task. The on-device model could be adapted to a new task with fewer than 200 demonstrations collected over several hours, according to the company.
Vendor-reported success rates varied substantially by task. On Apollo 2 with Inspire hands, the model achieved 68.4% for picking objects from a table, 45.7% from the floor, and 76.3% from a shelf. With SharpaWave hands, it achieved 92% for unscrewing a bulb but 36% for screwing a bulb and 44% for tying a trash bag. On Franka Duo, Google reported 74.2% for general pick-and-place tasks, 78.9% for tool kitting, and 89.6% for precise insertion.
Google also introduced the ASIMOV-Agentic benchmark for evaluating how embodied agents follow safety instructions and respond to uncertainty. The company reported 97.9% accuracy on safety-instruction following and 93% accuracy when classifying whether a person was within one meter of the robot. These results were published by Google and had not been independently replicated at the time of the release.
Gemini Robotics-ER 2 was made available through Google AI Studio, with an enterprise private preview also announced. Gemini Robotics 2 and the on-device action model remained limited to selected partners through a trusted-tester program. Google's model documentation advised against safety-critical deployment in healthcare, transportation, or other settings where a malfunction could cause injury, death, or property damage.
Why It Matters
Gemini Robotics 2 extended Google's general-model strategy from perception and upper-body manipulation to coordinated whole-body control across several robot forms. The combination of high-level planning, cloud action generation, and on-device execution placed reasoning and motor control within one named model family, rather than treating them as separate research systems.
The release also documented how far the systems remained from reliable general deployment. Several dexterous tasks succeeded in fewer than half of trials, movement speed remained below human performance, and the main action models were not publicly available for independent testing. The strongest claims therefore described a broader control architecture and new demonstrations, not a production-ready humanoid worker.
§ How to read the metadata
- Landmark
- Fundamentally alters the trajectory; 2–5 per year.
- Major
- Meaningfully shifts the landscape; 2–4 per month.
- Notable
- Worth documenting; significance can be upgraded later.
- Confidence
- High = primary sources corroborate. Medium = credible secondary only. Low = provisional. Disputed = credible sources disagree.
- Contestation
- Uncontested = no formal challenge. Contested = at least one challenge open. Superseded = replaced by a later entry. Unresolved = dispute still open.