Embodied AI Glossary中文

WorldEval: World Model as Real-World Robot Policies Evaluator

WorldEvalAdvanced

Using a video world model instead of a real robot to run rollouts and score and rank robot policies.

WorldEval is a robot-policy evaluation method proposed in May 2025 by researchers at Midea Group and East China Normal University, which substitutes a world model for the real robot during evaluation. Testing policies one by one on real hardware is slow and hard to reproduce. WorldEval instead runs a policy in closed loop inside a world model: the policy looks at generated frames and outputs actions, and the world model generates the next stretch of video based on those actions. To make the video strictly follow the actions, the authors propose Policy2Vec, which conditions a video-generation model on latent actions so the frames it generates track the actions given. Experiments show its rankings correlate strongly with real-robot results, and it can also distinguish between checkpoints of the same policy and flag dangerous actions. A May 2026 follow-up, dWorldEval, switched to a discrete diffusion world model and added a “progress token” that automatically judges whether a task is complete.

ExampleBefore deploying to a real robot, run several candidate policies — or several checkpoints of the same policy — through WorldEval to rank them first, and only take the top-ranked ones to real-robot validation.

Also called
dWorldEval
Related
World-Model-based Policy Evaluation · World Model · Video Generation Model · Latent Action · Sim-to-Real Correlation · Ctrl-World
Sources
WorldEval: World Model as Real-World Robot Policies Evaluator (arXiv 2505.19017)
WorldEval 项目主页 (Chinese)
dWorldEval (arXiv 2604.22152)
As of
2026-04

See it in the full glossary →