DeepMind broadens its robotics pitch

Google DeepMind has introduced Gemini Robotics 2, presenting it as a step beyond robotic systems limited to arms or tabletop manipulation and toward broader whole-body autonomy. According to The Robot Report, the updated vision-language-action model adds intelligent full-body control, greater dexterity and support for multiple robots working together.

The announcement matters because it targets a persistent gap in robotics. Many systems can perform carefully staged manipulation tasks in constrained environments, but far fewer can combine perception, language understanding and coordinated motion across an entire humanoid body. DeepMind’s claim is that Gemini Robotics 2 moves closer to that threshold by allowing robots to reason through movements such as walking, crouching, stretching and handling objects as part of a single task.

That does not mean general-purpose humanoids have arrived. The report is careful to note that movement speed still needs improvement. But it frames this release as an important step in turning multimodal AI from a tabletop demo layer into a controller for more realistic physical work.

From upper-body demos to full-body action

DeepMind’s earlier models focused on upper-body control for tabletop tasks. Gemini Robotics 2, by contrast, is described as managing the entire robot body. The Robot Report says the system can translate user intent into coordinated whole-body movement, enabling a humanoid to navigate space and complete tasks that involve both locomotion and manipulation.

One example in the report uses Apptronik’s Apollo 2 humanoid robot. When instructed to place a watering can into a green bin on a bottom shelf, the robot is said to interpret the request, walk to a table, pick up the object, move across the room and place it in the target location. That sequence is notable not because any one step is unprecedented, but because it combines natural-language instruction, object handling, walking and precise placement into one continuous behavior.

If those capabilities prove robust outside controlled demos, they would mark a practical expansion in what humanoid platforms can attempt. Warehouses, laboratories and service settings often require exactly this mix of mobility and manipulation rather than isolated arm motions in a fixed workstation.

Dexterity and collaboration are the other headline features

DeepMind is also emphasizing finer motor control. The Robot Report says Gemini Robotics 2 can control the five-fingered, 22-degree-of-freedom SharpaWave hand on Apollo 2. That suggests the company is not only working on gross motor movement but also on the more delicate finger and hand behaviors needed for useful grasping and placement.

Dexterity remains one of the hardest parts of real-world robotics. A robot that can walk across a room still has limited value if it cannot reliably grasp mixed objects, orient them correctly and interact with cluttered environments. By highlighting advanced hand control, DeepMind is signaling that it wants its robotics stack to cover more of the manipulation problem, not just navigation and language interpretation.

The company is also introducing multi-robot collaboration. The report says Gemini Robotics ER 2, an embodied reasoning model built as a vision-language model agent, is designed to help robots communicate with people, understand the physical world and plan multi-step tasks lasting several minutes. DeepMind is positioning that reasoning layer as the basis for teams of robots working together on a job rather than acting as isolated machines.

That point is strategically important. In industrial automation, the value of a robotics system often depends on orchestration across multiple machines and workflows. If AI models can coordinate handoffs or divide work among robots, deployment economics could change in favor of more flexible automation.

On-device execution could broaden deployment

Another notable element is local execution. The Robot Report says Gemini Robotics 2 can run on-device and adapt to entirely new robot bodies within a few hours. DeepMind has separately introduced Gemini Robotics On-Device 2, which it says is optimized to run locally on robotic hardware and can adapt to new embodiments with a few hours of data.

That claim addresses two real constraints in robotics deployment. First, cloud dependence can add latency, reliability concerns and operational complexity. Second, robotics developers do not want to retrain a system from scratch for every new platform. A model that can be retargeted to different embodiments quickly would be more attractive to hardware makers and enterprise customers trying to shorten integration cycles.

Availability, however, is still limited. The Robot Report says Gemini Robotics ER 2 is available on Google AI Studio and in private preview on the Gemini Enterprise Agent Platform, while the VLA and on-device models are being offered to early-access partners. That means the announcement is still closer to a controlled rollout than a broad product release.

What the release signals for physical AI

The broader message is that major AI companies are pushing harder into physical systems, not just digital assistants. DeepMind’s framing suggests it sees robotics as the next arena where multimodal models must prove they can connect understanding to action. Language, vision and planning are useful, but in robotics they only matter if they produce reliable motion in messy real spaces.

Gemini Robotics 2 does not settle whether humanoids are the winning form factor, nor does it prove that foundation-model-style control has solved the reliability problem. But it does indicate where the field is heading: toward models that can generalize across different robot bodies, execute locally, use both locomotion and dexterity, and coordinate with other robots over multi-step tasks.

Why this matters

  • The release expands AI control from tabletop manipulation to whole-body robotic motion.
  • DeepMind is pairing locomotion with dexterous hand control, a key requirement for practical work.
  • On-device operation and faster adaptation could lower integration barriers for new hardware.
  • Multi-robot collaboration points toward broader automation systems rather than single-purpose machines.

The immediate test will be whether early-access partners can reproduce these capabilities beyond showcase scenarios. If they can, Gemini Robotics 2 could become an important marker in the transition from multimodal AI demonstrations to more capable physical automation.

This article is based on reporting by The Robot Report. Read the original article.

Originally published on therobotreport.com