Google has released Gemini Robotics 2, a family of AI models that take us ever closer towards intelligent robots. Gemini Robotics ER2 acts as the brain for one or more robots, enabling robots to plan and to cooperate.


This video reveals how Gemini Robotics 2 provides the intelligence layer powering the next generation of truly adaptable robots and is a major advance that unlocks intelligent whole-body control, advanced dexterity, and multi-robot collaboration.


 

While foundation models for text and code have advanced rapidly over the last few years, applying AI to physical hardware has been a longer journey for Google DeepMind. The company has been working on solving robotics for nearly a decade, progressing from early reinforcement learning experiments in simulation to its Robotic Transformer (RT-1 and RT-2) research.

Google launched the first Gemini Robotics family built on top of Gemini 2.0 in March 2025. This marked the transition from single-purpose policies to multimodal foundation models capable of turning visual camera streams directly into robotic action. 

Without robots of it own, Google formed partnership with third-party robot manufacturers including, as I Programmer reported in January 2026, Boston Dynamics. By incorporating Gemini’s multimodal intelligence, Atlas is now much better able to perform complex, adaptive workplace tasks.

Recently we have seen robotics adopt reinforcement learning which proved highly successful for mastering physical dynamics through trial-and-error simulation — teaching quadrupeds to sprint, humanoids to recover their balance after being pushed, and arms to handle complex joint torques. However, pure RL hits a wall when faced with high-level cognitive tasks. Designing a mathematical reward function to train an RL policy to do a task such as “find the spare part in the back cupboard, ask a coworker for help if it’s heavy, and bring it here” is virtually impossible.

Gemini Robotics 2 signals a shift away from relying solely on end-to-end RL for physical AI. Instead, Google DeepMind has adopted a decoupled, two-tier architecture driven by vision-language foundation models (VLMs):



The High-Level Cognitive Layer (VLM): An embodied reasoning model like ER 2 processes streaming visual and audio input to maintain context, handle open-ended language requests, call external digital tools, and devise multi-step strategies over minutes-long tasks.



The Low-Level Execution Layer (VLA / Motor Policies): Instead of generating raw joint voltages itself, ER 2 issues structured sub-goals to lower-level Vision-Language-Action (VLA) models or RL-trained motor policies, which handle real-time spatial positioning, balance, and fine physical manipulation.



Where reinforcement learning teaches a robot how to move its joints and balance its weight, vision-language foundation models decide what the robot should be doing, in what sequence, and why.

The new release splits physical AI into three distinct components, balancing high-level reasoning, cloud execution, and local edge processing:



Gemini Robotics ER 2 (Embodied Reasoning): Built on Gemini 3.5 Flash, this model serves as the cognitive brain. It processes long-horizon video feeds and audio inputs to understand open-ended goals, track task progress, call external digital tools, and coordinate multi-robot teams.

Gemini Robotics ER 2 is now publicly available to developers via the Gemini API, Google AI Studio, and in private preview on Gemini Enterprise Agent Platform.



Gemini Robotics 2 (Vision-Language-Action): The primary cloud-based execution model. Unlike earlier iterations that only controlled arms and grippers, this generation provides whole-body humanoid control—managing movement across legs, torso, arms, and 22-DoF dexterous five-fingered hands simultaneously on platforms like Apptronik’s Apollo 2.



Gemini Robotics On-Device 2: A lightweight model designed to run directly on local robot microprocessors without network latency. It can adapt to an unfamiliar robot body or arm in just a few hours using fewer than 200 demonstration passes.



 

GeminiER2

 


More Information 

Gemini Robotics 2 brings whole body intelligence to robots

 


Related Articles

Atlas Production Ready This Year

Gemini On-Device – Generative AI For Robots

To be informed about new articles on I Programmer, sign up for our weekly newsletter, subscribe to the RSS feed and follow us on Facebook or Linkedin.

Banner


ESP RISC


 


Comments

Make a Comment or View Existing Comments Using Disqus

or email your comment to: comments@i-programmer.info