Loading article…
Google launches Gemini Robotics 2, enabling humanoid robots to walk, balance, use both hands and collaborate, with on‑device models adapting in hours.
Google unveiled Gemini Robotics 2, a vision‑language‑action model that lets humanoid robots perform whole‑body movements, fine‑grip manipulation and coordinate with other robots, marking a step toward adaptable, general‑purpose physical AI【1】. The advance matters because it moves robots beyond pre‑programmed, single‑task routines toward autonomous operation in dynamic, real‑world settings.
| At a glance | |
|---|---|
| Model | Gemini Robotics 2 (vision‑language‑action) |
| Capabilities | Whole‑body control, dexterous hand/gripper use, multi‑robot collaboration |
| On‑device variant | Gemini Robotics On‑Device 2 (local inference, few‑hour adaptation) |
| Availability | Early‑access partners; ER 2 in private preview on Google AI Studio【1】【2】 |
Gemini Robotics 2 expands on earlier models that handled only upper‑body tasks by controlling a robot from feet to fingertips. In a live demo, the model guided Apptronik’s Apollo 2 humanoid to walk, crouch, stretch and place a watering can on a lower shelf, showing the ability to translate natural‑language commands into coordinated full‑body motion【1】【2】. The same checkpoint also operated on the Franka Duo platform, proving cross‑embodiment flexibility.
Dexterity improvements include controlling Apollo 2’s 22‑degree‑of‑freedom five‑fingered hand to tie knots, seal zip‑lock bags and unscrew light bulbs, while two‑finger grippers on Franka Duo performed precise insertion and packing tasks【1】【2】. Google notes ongoing work to reach human‑level speed and precision, but the current level already surpasses prior tabletop‑only manipulation.
The companion model Gemini Robotics ER 2 serves as a high‑level reasoning engine, breaking spoken instructions into multi‑step plans, monitoring progress and handling failures, enabling tasks that last several minutes and involve hundreds of decisions【1】【2】. A new multi‑robot collaboration feature lets different robots communicate and share workloads, a capability absent from most existing systems.
Gemini Robotics On‑Device 2 is optimized for local execution, allowing robots to adapt to new embodiments with fewer than 200 examples over a few hours, eliminating reliance on cloud connectivity and reducing latency【1】. This fast adaptation was demonstrated on diverse platforms such as Dexmate, SO101 and Trossen.
Compared with conventional industrial robots that require extensive reprogramming for each new task, Gemini’s models aim for rapid skill transfer across bodies, a claim Google attributes to its “motion transfer” techniques inherited from Gemini 1.5【1】. While other firms like Boston Dynamics and NVIDIA have showcased locomotion and manipulation, Google’s integration of vision, language and action in a single model, plus on‑device operation, differentiates its approach. No independent benchmark data are provided, so performance relative to rivals remains unverified.
Google’s Gemini Robotics 2 demonstrates that large‑scale multimodal AI can now drive full‑body robot behavior, dexterous manipulation and collaborative workflows, narrowing the gap between fixed‑function automation and adaptable physical agents. The open question is whether the rapid adaptation and safety mechanisms will translate into reliable, large‑scale deployments across homes and factories.
Coverage is mostly measured — 246 of 257 reports stay neutral.
Every Monday — the token unlocks, Fed dates & catalysts set to move crypto and markets this week. So you’re never blindsided.
Free · 3-min read · one-click unsubscribe
AI-assisted synthesis by the TrendWatcher Editorial Desk · sourced from 2 outlets · Jul 31, 2026 · How we report
Google fixed 1,072 security bugs in the last two Chrome versions released in June 2024, surpassing the 1,036 bugs patched in the prior 23 versions over two years.
Gemini Robotics 2 allows robots to control their entire humanoid body, perform dexterous tasks like knot tying, and coordinate multiple robots through the Gemini Robotics ER 2 reasoning model.
Google Earth now integrates the Nano Banana AI image generator, enabling users to create AI-generated images directly within the mapping application.