WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models
WorldArena 具身世界模型基准AdvancedA Tsinghua-led embodied world-model benchmark scoring both how good the generated video looks and how useful it actually is.
WorldArena is an embodied world-model evaluation benchmark led by Tsinghua University with several other institutions, released in February 2026. An embodied world model predicts future frames given the current frame plus an action or instruction. The authors argue existing evaluations only ask whether a video looks good, not whether it actually helps decision-making, so they score along three axes: video quality (6 sub-dimensions and 16 metrics in total), embodied-task utility (testing the world model as a data engine, a policy evaluator, and an action planner), and human evaluation, combined into an overall EWMScore. After evaluating 14 models on RoboTwin 2.0 dual-arm tasks, the authors found that good-looking video does not imply strong task performance. A 2.0 version from May 2026 added visuotactile input and real-robot platforms.
ExampleThe first version evaluated general video models such as Wan 2.2 and Veo 3.1, alongside embodied world models such as Cosmos-Predict 2.5, Genie Envisioner, and Ctrl-World, with a public leaderboard at world-arena.ai.
- Also called
- WorldArena, WorldArena 2.0
- Related
- World Model · World-Model-based Policy Evaluation · EWMBench · RoboTwin · Ctrl-World · WorldScore: A Unified Evaluation Benchmark for World Generation
- Sources
- WorldArena (arXiv 2602.08971)
WorldArena 2.0 (arXiv 2605.17912) - As of
- 2026-05