Google DeepMind says the latest version of its Gemini Robotics AI model can “control entire humanoid robots.” While the previous model focused on controlling a humanoid robot’s upper body, Gemini Robotics 2 now supports “whole-body motions” ranging from its feet to fingertips, according to an announcement on Thursday.
Google DeepMind’s new AI model can control a robot’s entire body
Gemini Robotics 2 controls a humanoid robot from ‘feet to fingertips.’
Gemini Robotics 2 controls a humanoid robot from ‘feet to fingertips.’
The new model will allow humanoid robots to perform a wider range of actions, as it allows them to walk, crouch, stretch, and manipulate objects. Videos shared by Google show how Apptronik’s Apollo 2 robot can bend over to pick up a watering can, as well as find and take specific items off a shelf.
Though Google DeepMind notes that its robots “have more to advance in movement speed,” it adds that this update “is an important step towards the skills needed to complete more complex, real-world tasks that require whole-body coordination.”
Additionally, Gemini Robotics 2 supports better dexterity, as it can now control more complex, five-fingered hands. That enables robots to perform tasks like sealing a Ziploc, tying a trash bag, or unscrewing a lightbulb.
Google DeepMind is updating Gemini Robotics ER (embodied reasoning) as well, a vision-language model that helps robots to analyze their surroundings, process instructions, and perform multi-step tasks. Gemini Robotics ER 2 is better at completing tasks over an extended period of time and “now understands when tasks begin and end.”
Google DeepMind says this update also allows multiple robots of different types to work together and complete tasks, with one video showing how Apollo 2 instructs Google’s dual-arm robot to put tools inside a bin while cleaning the garage.
The company notes Gemini Robotics ER 2 is its “safest robotics model to date,” as it can “better detect when humans are nearby, trigger safety tool calls and bring the robot to a safe stop if someone approaches too closely.”
Meanwhile, Google DeepMind has brought improvements to its Gemini Robotics On-Device Model, which can run locally on a robot without an internet connection. This model can now adapt to new embodiments faster, including those with “drastically different shapes, sensors and degrees of freedom.”
Facts Only
Google DeepMind announced Gemini Robotics 2 on Thursday.
Gemini Robotics 2 supports whole-body motions for humanoid robots, including feet and fingertips.
The model enables robots to walk, crouch, stretch, and manipulate objects.
Apptronik’s Apollo 2 robot demonstrated picking up a watering can and removing items from a shelf.
The model supports control of five-fingered hands for tasks such as sealing Ziploc bags and tying trash bags.
Gemini Robotics ER 2 is a vision-language model for analyzing surroundings and performing multi-step tasks.
Gemini Robotics ER 2 allows different types of robots to coordinate, such as Apollo 2 instructing a dual-arm robot.
Gemini Robotics ER 2 includes safety features to detect humans and stop the robot.
An on-device model version runs locally without an internet connection.
The on-device model adapts to different robot shapes, sensors, and degrees of freedom.
Executive Summary
Google DeepMind has introduced Gemini Robotics 2, an AI update that expands humanoid robot control from the upper body to full-body coordination. This allows robots, such as Apptronik’s Apollo 2, to perform complex physical movements like crouching and walking, while utilizing five-fingered dexterity for intricate tasks like unscrewing lightbulbs. Alongside the physical control model, Gemini Robotics ER 2 enhances embodied reasoning, improving a robot's ability to manage multi-step tasks and coordinate with other robotic systems.
The update emphasizes safety through improved human detection and offers an on-device model for local processing, which increases the system's adaptability to various hardware configurations. While these advancements represent a significant step toward real-world utility, there is acknowledged uncertainty regarding current movement speeds, which Google DeepMind notes require further advancement.
Full Take
The strongest version of this narrative is that we are witnessing the transition of Large Language Models from digital screens to physical "embodiment," where reasoning is coupled with precise, whole-body motor control. This suggests a trajectory toward general-purpose robotics capable of navigating and manipulating human environments with minimal specialized programming.
The framing relies heavily on the "Authority Game," where the capabilities of the AI are presented through the lens of the developer's own claims and curated video demonstrations. By highlighting "safest robotics model to date" without specifying the metrics or independent safety audits used to define "safe," the narrative uses a veneer of technical rigor to instill confidence in a high-risk technology.
The root paradigm here is techno-optimism: the assumption that increasing "degrees of freedom" and "embodied reasoning" naturally leads to utility and safety. The unstated assumption is that the environment will adapt to the robot, rather than the robot truly understanding the nuance of human space. The second-order consequence is a potential erosion of human agency in manual labor sectors, shifting the cost from the employer (who gains efficiency) to the displaced worker.
Patterns detected: ARC-0062 Authority Game
If this were a coordinated influence campaign, the playbook would involve "feature flooding"—overwhelming the audience with a list of impressive micro-tasks (tying bags, unscrewing bulbs)—to distract from the lack of data on reliability, failure rates, or energy efficiency in non-simulated environments. The actual content follows a standard corporate announcement pattern rather than a deceptive campaign, though it maintains a strong promotional bias.
What independent benchmarks would be required to verify these dexterity claims? How does the transition to "on-device" processing affect the safety latency when a human enters the robot's workspace?
Sentinel — Human
This appears to be human-written reporting summarizing a technical announcement, characterized by logical flow and the integration of several detailed points about the AI model's advancements.
