News · 2026-07-30
Google's Gemini Robotics 2 controls a humanoid from feet to fingertips - for a waitlist
Google DeepMind announced Gemini Robotics 2 on July 30, a robotics model family that controls a humanoid's entire body rather than just its arms, and lets multiple robots divide a task between them. Carolina Parada of Google DeepMind summarised it as: "From feet to fingertips - we are teaching robots intelligent whole-body control, fine dexterity, and teamwork to complete a broad range of complex tasks." Only one of the three models announced is actually available to developers.
Key facts
- Three models: Gemini Robotics 2 (control), Gemini Robotics ER 2 (embodied reasoning), Gemini Robotics On-Device 2.
- Only ER 2 is generally accessible, via the Gemini API, Google AI Studio and the Gemini Enterprise Agent Platform. The VLA model is private preview; on-device is trusted testers.
- Hands have 22 degrees of freedom; the on-device model adapts to a new robot body from fewer than 200 examples.
- Primary source: Google DeepMind's announcement, July 30, 2026.
The substance behind "whole-body intelligence" is a real technical distinction. Most robot manipulation research treats the body as a fixed base and the arm as the thing being controlled. That works on a bolted-down workcell and fails on a humanoid, where reaching for something on a low shelf means shifting your centre of gravity, and getting it wrong means falling over. Google's claim is that one model now handles balance, stepping, squatting and bending in the same loop as the grasp - so the robot decides how to arrange its whole body to make a manipulation possible.
The multi-robot capability is the more unusual claim. Google's description is that "multiple robots can communicate, recognize each other's unique physical strengths, and autonomously delegate tasks to complete a shared mission." The demonstration pairs two different machines - a humanoid and a dual-arm platform - which is the point: the delegation is interesting precisely because the robots are not interchangeable and have to reason about which of them is better suited to which part of the job.
The published performance numbers deserve attention because they are unusually candid for a launch page, and they are a spread rather than a headline. On the humanoid with one hand configuration, picking objects from a table succeeds around two-thirds of the time and from the floor under half. On another hand configuration, unscrewing a lightbulb works more than nine times in ten while screwing one in works about a third of the time. On a dual-arm platform with a simple gripper, precise insertion succeeds close to nine times in ten. That asymmetry - unscrewing easy, screwing hard - is exactly the texture of real manipulation, and it is a more informative disclosure than a single average would have been.
The availability split is where the launch narrative and the product reality diverge, and it is worth being precise about. ER 2, the model that reasons about the physical world without directly controlling motors, is the one you can call today. The models that actually move robots are behind a waitlist. Anyone reading this as a robotics API launch is reading it wrong; it is a research preview with one reasoning endpoint attached.
The useful counterweight arrived the same week from academia, in two papers pointing in opposite directions. TurboVLA shows the control side getting radically cheaper - real-time robot control at 32 times a second on a consumer graphics card, in under a gigabyte of memory. Meanwhile HumanCLAW tests whether today's vision-language models can make correct moment-to-moment decisions through a body at all, and finds the best of nine frontier models completing under one in six episodes. The gap between Google's polished demonstrations and that benchmark is the honest measure of where embodied AI is.
The on-device claim is the one with the clearest practical implication. Google says the smaller model runs locally on the robot and can be adapted to a new robot body with fewer than 200 demonstrations and a few hours of training. If that holds outside Google's chosen platforms, it changes the economics of deploying a robot into an unusual body - which has historically meant months of engineering. It follows the same direction as work like training robots from handheld-gripper data with no robot in the loop: the bottleneck in robotics has always been demonstration data, and everything interesting is aimed at needing less of it.
The caveat is the oldest one in robotics reporting: a highlight reel is not a capability claim, and Google's own numbers show why. A model that unscrews bulbs nine times in ten and screws them in three times in ten is not a model that can change a lightbulb. Published per-task success rates on named hardware are far more than most robotics announcements offer, and they are also the reason to be careful about how the demo video reads.
Key questions
Can developers use Gemini Robotics 2 today?
What does 'whole-body intelligence' actually mean?
How well does it actually perform the tasks?
Cite this
APA
Ground Truth. (2026, July 30). Google's Gemini Robotics 2 controls a humanoid from feet to fingertips - for a waitlist. Ground Truth. https://groundtruth.day/news/gemini-robotics-2-controls-a-humanoid-from-feet-to-fingertips.html
BibTeX
@misc{groundtruth:gemini-robotics-2-controls-a-humanoid-from-feet-to-fingertips,
title = {Google's Gemini Robotics 2 controls a humanoid from feet to fingertips - for a waitlist},
author = {{Ground Truth}},
year = {2026},
month = {jul},
url = {https://groundtruth.day/news/gemini-robotics-2-controls-a-humanoid-from-feet-to-fingertips.html}
}
Comments are replies to this story on Bluesky — reply with any Bluesky account to join in.