Editorial illustration for Google Trains Gemini Robot AI With Human Teleoperation
Google Trains Gemini Robot AI With Human Teleoperation
Google DeepMind unveiled Gemini Robotics 2 this week, a version of its Gemini AI built to control robots rather than just chat or generate text. The system fuses a vision language model, which reads images and video and reasons about tasks, with two separate vision language action models that handle full-body movement and the finer work of grippers and hands. In footage released ahead of launch, Apptronik's Apollo 2 robot, fitted with hands from a company called Sharpa, tidies shelves without a human directing each motion. Other clips show robots screwing in lightbulbs and tying off trash bags, the kind of fiddly physical work that has long tripped up machines.
None of this comes from the model teaching itself. Google DeepMind trained Gemini Robotics 2 on a mix of human teleoperation, recorded video demonstrations, and simulation, since AI systems still can't generalize to messy real-world tasks on their own. OpenAI and Anthropic have pulled ahead in chatbots and coding assistants, but Google has spent years building up robotics research most rivals haven't matched.
Google DeepMind just released a new version of its artificial intelligence model Gemini, and it can control a range of different robots—including humanoids capable of dextrous tasks like screwing in lightbulbs and tying trash bags.
Why this matters
Google is admitting something the robotics field has danced around for years: general-purpose robot intelligence still runs on human hands. Teleoperation, video demos, and simulation are the actual training diet behind Gemini Robotics 2, not some emergent capacity to figure out lightbulbs and trash bags on its own. For developers and founders building on top of these models, that's a useful reality check before the demo reel convinces you otherwise.
Screwing in a lightbulb after thousands of guided repetitions is not the same as a robot generalizing to your warehouse floor or your kitchen. The bundling of a vision-language model with action policies is a real engineering step, and combining perception with control in one system matters for anyone designing multi-modal AI products. But the labor behind it, humans physically guiding these machines, tells you where the bottleneck still sits.
Anyone pricing out embodied AI for logistics, manufacturing, or home robotics should ask DeepMind directly how much teleoperation scales before believing the humanoid-robot timeline Silicon Valley keeps selling.
Further Reading
- Google's new Gemini Robotics 2 platform allows for 'intelligent whole-body control' - Engadget
- Google DeepMind Ships Three Physical AI Models For Robots With Whole-Body Control - MarkTechPost
- Introducing Gemini Robotics ER 2 - Google Blog
- Google Unveils Gemini AI for Robots Struggling With Dexterity - Bloomberg
- Google's Gemini Robotics AI Model Reaches Into the Physical World - Wired