Editorial illustration for Google DeepMind Demos AI Orchestrating Boston Dynamics Spot Robot
Google DeepMind's AI Controls Boston Dynamics Robot
Google DeepMind put a Boston Dynamics Spot robot on stage this week and let its new AI system run the show, part of a broader release the company calls Gemini Robotics 2. The lineup ships as three separate models, not one: a vision-language-action model for motor control, an embodied reasoning model called Gemini Robotics ER 2 that acts as the planning brain, and a smaller on-device VLA for offline work. One checkpoint, DeepMind says, can already drive the Apollo 2 humanoid across two different hand designs plus a Franka Duo gripper, a claim aimed squarely at the industry's oldest problem: skills built for one robot body rarely transfer to another.
That gap, along with the inability of most deployed robots to handle anything outside a scripted, tele-operated sequence, is what Gemini Robotics 2 is built to close. Access is split by tier. ER 2 is in public preview, while the full VLA and the on-device model stay gated. DeepMind also published a new safety benchmark, ASIMOV-Agentic, on Hugging Face under a CC-BY-4.0 license, alongside data showing multi-finger dexterity still lagging well behind other skills.
Google DeepMind has released Gemini Robotics 2, the intelligence layer for its next generation of robots. The release moves the stack past table-top manipulation into whole body control, five finger dexterity and multi robot teamwork.
Why this matters
The Spot demo is the tell here, not the dexterity claims. Google DeepMind built code that lets ER 2 call Boston Dynamics' navigation and manipulator APIs directly, and put the sample in a public repository. That's a different bet than shipping a better model card: it's an invitation for developers to plug Gemini Robotics 2 into hardware DeepMind didn't build.
If that works reliably across robot bodies, the "one model per robot" problem that's dogged embodied AI for years starts to look solvable, and the moat shifts from hardware integration to whoever controls the reasoning layer. We'd watch two things closely: whether the multi-robot collaboration claims hold up outside curated demos, and how the three-tier access system gates who actually gets to build on this. Founders building robotics startups should be asking whether they're now building on Google's stack by default.
Researchers should be skeptical of "adapts to unpredictable environments" until someone outside DeepMind runs it on a robot Google didn't hand-pick.
Common Questions Answered
What are the three separate models included in Google DeepMind's Gemini Robotics 2 release?
Gemini Robotics 2 consists of a vision-language-action model for motor control, an embodied reasoning model called Gemini Robotics ER 2 that serves as the planning brain, and a smaller on-device VLA for offline work. These three models work together to enable comprehensive robotic control and decision-making capabilities across different hardware platforms.
How does Gemini Robotics 2 advance beyond previous robotic AI capabilities?
The release moves the technology stack past table-top manipulation into whole body control, five finger dexterity, and multi-robot teamwork. This represents a significant expansion from previous limitations in embodied AI systems that could only handle simpler manipulation tasks.
What is significant about Google DeepMind's approach to making Gemini Robotics 2 available to developers?
DeepMind built code that allows the ER 2 model to call Boston Dynamics' navigation and manipulator APIs directly and released sample code in a public repository. This approach invites developers to integrate Gemini Robotics 2 into hardware that DeepMind didn't build, potentially solving the long-standing "one model per robot" problem in embodied AI.
What was demonstrated during the Boston Dynamics Spot robot stage demo?
Google DeepMind showcased its new AI system orchestrating a Boston Dynamics Spot robot, with the Gemini Robotics 2 models controlling the robot's movements and decision-making in real-time. The demonstration highlighted how the embodied reasoning model can effectively direct physical robot operations across different tasks.
Can a single Gemini Robotics 2 checkpoint operate different robot bodies?
Yes, one checkpoint of Gemini Robotics 2 can already drive the Apollo 2 humanoid across two different environments, suggesting the models are designed to work across multiple robot platforms. This cross-platform compatibility addresses a key challenge in embodied AI where systems typically required separate models for each robot type.
Further Reading
- Introducing Gemini Robotics ER 2 - Google Blog
- Boston Dynamics integrates Google DeepMind's Gemini Robotics model into Spot inspection platform - Robotics & Automation News
- Robot dog can now read and reason after AI upgrade - The Independent
- Boston Dynamics' Robot Dog Can Now Read Gauges, Spot Spills, and Reason - Slashdot
- Spot robot gets Gemini AI to boost real-world inspection tasks - Interesting Engineering