Key takeaways
- X Square Robot publishes an open, integrated stack that ties together interaction data, a predictive world model, and an action model for robot behavior.
- DeepMind’s Gemini family provides a vision‑language‑action foundation, first released in March 2025 and iterated through on‑device and embodied‑reasoning versions.
- Gemini Robotics ER 2 (July 30 2026) adds real‑time video understanding, task‑progress tracking, a low‑latency Live API, tool integration, and multi‑robot collaboration.
- The SIMA 2 virtual‑world agent, built on Gemini, demonstrates near‑human performance across diverse 3D games and can self‑improve by generating its own tasks.
- Agility Robotics, Boston Dynamics, Agile Robots, and Enchanted Tools are named as early trusted testers with access to the Gemini Robotics models.
An open, unified stack for embodied AI
X Square Robot describes its platform as a single software stack that brings together three complementary layers: a data layer that records discrete interactions, a world model that forecasts how the environment will evolve, and an action model that merges perception, planning, reasoning, and decision‑making into executable robot commands. The company frames the stack as a response to the fragmented perception‑planning‑control pipelines that dominate current robot software.
The stack rests on three design principles that the company spells out explicitly:
- Robot experience is captured as interactions—moments when the robot actually changes the world—rather than as raw joint trajectories.
- Pre‑training is expected to deliver immediate capabilities, not merely a warm‑start for later fine‑tuning.
- Behavior is expressed in terms of physical events instead of fixed time slices.
All three layers share a common code base, allowing the same robot‑free data to train both the world‑model and the action model. X Square says it intends to build and release the stack in the open, rather than keep it proprietary.
Vision‑language‑action foundation models: Gemini’s evolution
DeepMind’s Gemini series follows a different philosophy: a large‑scale foundation model that natively blends vision, language, and action capabilities. The first Gemini Robotics model, launched on March 12 2025, is described as an advanced vision–language–action model built on the Gemini 2.0 large language model and tailored for robotics.
A follow‑up variant, Gemini Robotics On‑Device, arrived on June 24 2025 and was optimized to run locally on robot hardware, reducing the need for constant cloud connectivity.
The most recent release, Gemini Robotics ER 2, was announced on July 30 2026. This version introduces a suite of new features: real‑time video understanding, task‑progress tracking, lower‑latency orchestration through the Gemini Live API, tool integration, and multi‑robot collaboration. The announcement frames ER 2 as the "intelligence layer" that equips next‑generation robots with whole‑body control, advanced dexterity, and coordinated team behavior.
Access to the Gemini series has been limited to a handful of trusted testers, including Agility Robotics, Boston Dynamics, Agile Robots, and Enchanted Tools.
From virtual games to physical hands: the SIMA 2 bridge
DeepMind’s research team released SIMA 2, a generalist embodied agent that operates across a wide variety of 3D virtual worlds. The paper states that SIMA 2 is built upon a Gemini foundation model and that it "substantially closes the gap with human performance" on a diverse portfolio of games.
Beyond imitation, SIMA 2 can engage in dialogue, plan multi‑step actions, and even generate its own training tasks. By leveraging Gemini to produce rewards, the agent can autonomously learn new skills from scratch in environments it has never seen before. The authors present this capability as a concrete path toward transferring learned embodied intelligence from simulation to real‑world robots.
How Gemini ER 2 reshapes humanoid control
The feature set announced for Gemini ER 2 directly tackles the integration challenge highlighted by X Square’s stack. Real‑time video understanding eliminates the need for a separate perception pipeline, while the low‑latency Live API speeds up command turnaround for whole‑body motion. Tool integration and multi‑robot collaboration extend a single robot’s functional envelope without requiring bespoke software for each new end‑effector or teammate.
By delivering a model that already encodes visual, linguistic, and motor priors, ER 2 reduces the engineering effort needed to stitch together disparate perception, planning, and control modules. This aligns with the principle of sharing a common code base across world‑model and action‑model components, as advocated by X Square.
Release timeline and focus of Gemini models
| Model | Release date | Primary focus |
|---|---|---|
| Gemini Robotics (vision‑language‑action) | March 12 2025 | General robotics foundation built on Gemini 2.0 |
| Gemini Robotics On‑Device | June 24 2025 | Optimized for local inference on robot hardware |
| Gemini Robotics ER 2 | July 30 2026 | Real‑time video understanding, task tracking, Live API, tool integration, multi‑robot collaboration |
Industry implications and early adopters
So far, the only outside groups with confirmed hands‑on access to Gemini's robotics models are the four named trusted testers: Agility Robotics, Boston Dynamics, Agile Robots, and Enchanted Tools. That is a narrow, invite‑only rollout rather than a broad commercial launch — how or whether any of the four are using the models in production is not disclosed.
X Square’s open stack takes the opposite approach: it says it plans to build and release its stack publicly, in contrast to Gemini’s limited‑access model. Whether an open stack can match a large proprietary foundation model’s performance in practice is untested and not addressed by either company’s public statements.
Outlook for embodied intelligence
The convergence of three strands—open, interaction‑focused stacks; large‑scale vision‑language‑action foundation models; and virtual‑world training pipelines—suggests a near‑term shift in how humanoid robots are built. Rather than assembling hand‑crafted perception, planning, and control blocks, developers can start from a unified model that already understands vision, language, and action, then fine‑tune it for specific tasks.
For engineers, the practical takeaway is to watch whether a unified foundation model can replace fragmented software pipelines, especially when targeting whole‑body dexterity and coordinated multi‑robot workflows.
Conclusion
Foundation models are moving from pure language predictors to embodied reasoning engines. By combining interaction‑rich data, predictive world models, and vision‑language‑action architectures, projects ranging from X Square’s open stack to DeepMind’s Gemini ER 2 are turning that potential into concrete robot capabilities. As virtual‑world training continues to narrow the performance gap with humans, the line between simulation and reality blurs, paving the way for humanoid robots that can learn, adapt, and collaborate with far less hand‑engineered plumbing.
Sources
This article was researched and fact-checked against the following sources:
- X Square Robot's Open-Source Embodied AI Stack - IEEE Spectrum (spectrum.ieee.org)
- Gemini Robotics - Wikipedia (en.wikipedia.org)
- Access denied — Trend Hunter (trendhunter.com)
- Robot Videos: Gemini 2 AI Robot, Robot Dogs, and More - IEEE Spectrum (spectrum.ieee.org)
- SIMA 2: A Generalist Embodied Agent for Virtual Worlds (arxiv.org)