Embodied AI Glossary中文

Veo World Simulator

Veo 世界模拟器Advanced

A Google DeepMind system that uses the Veo video model to 'imagine' a robot's execution in order to evaluate policies.

The Veo World Simulator is a technical report released by Google DeepMind's Gemini Robotics team in December 2025. Real-robot evaluation is slow, expensive, and hard to extend to new objects or backgrounds not seen before. The team adapted the Veo 2 video generation model into a world simulator: given the current image and a sequence of future robot poses, it generates the resulting frames from all four camera views of an ALOHA 2 bimanual platform, and uses image editing plus multi-view inpainting to alter the real scene with new objects, backgrounds, and distractors. This makes it possible to run policies inside 'generated video,' predict the relative strengths and weaknesses of different policy versions, compare which kinds of generalization are harder, and even conduct red-teaming to surface unsafe behaviors. The authors validated the agreement between these predictions and real-robot results using 8 versions of Gemini Robotics policies, 5 tasks, and more than 1,600 real-robot trials.

ExampleIn one red-team test, the simulator generated a scene where the instruction 'quick, grab the red block!' exposed a policy bumping into a nearby hand; in another, a policy closed a laptop before moving a pair of scissors out of the way, risking damage to the screen.

Also called
Evaluating Gemini Robotics Policies in a Veo World Simulator
Related
World-Model-based Policy Evaluation · Gemini Robotics · Video Generation Model · Interactive World Model · Sim-to-Real Correlation · Mean Maximum Rank Violation
Sources
Evaluating Gemini Robotics Policies in a Veo World Simulator (arXiv 2512.10675)
论文 HTML 版(Veo 2、ALOHA 2 与相关性结果) (Chinese)
As of
2026-01

See it in the full glossary →