Google’s AI lab DeepMind officially launched its next-generation robotics foundation model, Gemini Robotics 2, on July 30. The system introduces a series of updates designed to give AI-powered robots more dexterous manipulation capabilities while making their control more intuitive and efficient.
According to Axios, the update comprises two core components: one model controls a robot’s individual physical movements for more refined and fluid actions, while the other handles higher-level task planning, endowing robots with human-like logical reasoning and decision-making abilities. This marks a new phase in Google’s strategy to deeply integrate large language models into physical robots.
Google offered an intuitive description of the new model’s capabilities in an official blog post. “Gemini Robotics 2 enables robots to reason about every action, unlocking a wide range of application tasks,” the company stated. The post elaborated that the model allows humanoid robots to perform complex sequences of actions—walking, crouching, reaching, and manipulating objects—to clean up a cluttered room. More strikingly, robots equipped with the system can even team up and collaborate with other robots to complete cleaning tasks at a faster pace.
Simultaneously, Google DeepMind also released the Gemini Robotics ER 2 model. According to Jiemian News, Google likens this model to a robot’s “high-level brain.” It not only enables natural language chat interactions with humans but, more importantly, gives robots the ability to understand the complex physical world and plan multi-step tasks. This means robots are no longer just mechanical arms executing single commands; they become intelligent agents capable of understanding a vague instruction like “tidy up the room” and autonomously breaking it down into specific steps: identifying clutter, grasping objects, sorting and putting items away, and moving obstacles.
From a technical trajectory standpoint, this update represents a continuation and major upgrade of Google’s plan, announced last year, to bring the Gemini model into the robotics domain. By combining high-level reasoning with low-level motion control, Gemini Robotics 2 attempts to break away from traditional robotics programming’s heavy reliance on specific scenarios, moving toward a more generalizable, universal intelligence.
However, while software technology achieves breakthroughs, the robotics hardware supply chain faces new uncertainties. Axios noted in its report that Google’s new model launch comes at a time when U.S. companies may find it harder to acquire the humanoid robot hardware needed to run such software. This week, the U.S. Federal Communications Commission (FCC) announced that, citing security concerns, it will ban future sales of Chinese-made robots in the United States. This policy shift could impact robot manufacturers and software developers reliant on China’s supply chain, creating additional constraints for advanced software like Gemini Robotics 2 in finding compatible hardware platforms.
The release of Gemini Robotics 2 further intensifies the competition among tech giants in the field of embodied intelligence. As artificial intelligence extends from the purely digital world into the physical realm, making machines perceive, think, and act like humans has become the next frontier. By deeply integrating advanced vision-language models with robot control, Google has demonstrated its ambitions in general-purpose robotics, while geopolitical variables in the hardware supply chain add a new layer of complexity to this technological race.