World-Model-based Policy Evaluation
世界模型评测CommonUsing an interactive video world model in place of a real robot, running the policy in closed loop to estimate its performance.
World-model-based evaluation connects a robot policy to an action-conditioned video generation model in a closed loop: the policy looks at the generated frame and outputs an action, the world model predicts the next frame based on that action, and the cycle repeats, with a human or a vision-language model finally judging whether the task was completed. It targets the problem that real-robot evaluation is slow, expensive, and hard to reproduce, while traditional simulators are laborious to set up and still fall short on visual and physical realism. Work such as WorldEval, WorldGym, and Ctrl-World appeared starting in 2025; Google DeepMind used its Veo video model to evaluate Gemini Robotics policies and cross-checked the results against over 1,600 real-robot evaluations, showing it can predict the relative ranking of different policies and can also be used for out-of-distribution generalization and safety red-teaming. New work has continued to appear into 2026, and the approach is currently used mainly for ranking checkpoints and screening for risk.
ExampleGoogle DeepMind swapped in new objects, new backgrounds, and distractors inside the Veo world simulator to compare 8 Gemini Robotics policy checkpoints across 5 bimanual tasks, then checked the ranking against real-robot results.
- Also called
- World Model as Policy Evaluator, World-Model Evaluation
- Related
- World Model · Veo World Simulator · WorldEval: World Model as Real-World Robot Policies Evaluator · Ctrl-World · Sim-to-Real Correlation · Real-World Evaluation
- Sources
- Evaluating Gemini Robotics Policies in a Veo World Simulator (arXiv 2512.10675)
WorldEval: World Model as Real-World Robot Policies Evaluator (arXiv 2505.19017) - As of
- 2026-07