DeepMind ships a single control model for arms and full-body robots
Google DeepMind has introduced Gemini Robotics 2, which it calls its most advanced vision-language-action model — a class that fuses image recognition, language processing and action control so a robot can operate in physical environments. The claim that matters is breadth: DeepMind says the same model can control systems ranging from tabletop arms to full-body humanoids, managing whole-body movement, fine motor tasks and coordination across multiple robots. The company describes it as an intelligence layer for a new generation of adaptive robots rather than as a robot itself.
Alongside it, DeepMind released Gemini Robotics ER 2, aimed at embodied reasoning — understanding the physical world and deciding which actions follow from it. ER 2 sits above the control model and replaces ER 1.6, released in April. ER 2 is available in Google AI Studio; the main robotics model is behind an early-access waitlist.
One model across many robot bodies is the thing that would make physical labor cheap. Today each platform carries its own control stack, and that per-platform engineering is a large share of why automation stays expensive relative to what it accomplishes. A shared intelligence layer moves that cost from bespoke software toward hardware, where volume manufacturing does the work.
Everything above is DeepMind's own description of its own model. The announcement carries no benchmarks, no independent evaluation and no deployment numbers, and the model is not generally available. Watch for results from waitlist developers.
Source: The Decoder
MANY MINDED