Embodied AI Glossary中文

EWMBench

EWMBench 具身世界模型评测Advanced

A benchmark from AgiBot and others that specifically evaluates whether robot-manipulation video generation models get the details right.

EWMBench is an evaluation benchmark released in May 2025 by AgiBot together with Shanghai Jiao Tong University, CUHK MMLab, and the Harbin Institute of Technology. It targets embodied world models: models that generate a video of robot manipulation given a starting frame and a task instruction. The authors argue that metrics like FVD (Fréchet Video Distance) and VBench mainly assess visual quality and can't tell whether the robot arm's motion is actually plausible, so they score along three dimensions instead: scene consistency (comparing inter-frame features with a DINOv2 model fine-tuned on embodied data), motion correctness (Hausdorff distance and normalized dynamic time warping between the end-effector trajectory and ground truth, plus velocity and acceleration distributions), and semantic alignment (having a multimodal large model write a description of the generated video, checking for logical errors, and comparing it against ground truth). Test data is drawn from 10 tasks in the AgiBot World dataset, and both the dataset and evaluation code are open-sourced on GitHub.

ExampleThe paper used EWMBench to compare 7 models — OpenSora 2.0, LTX, COSMOS-7B, Kling-1.6, Hailuo, EnerVerse, and others — concluding that current video generation models still have clear shortcomings when applied to embodied tasks.

Also called
Evaluating Scene, Motion, and Semantic Quality in Embodied World Models
Related
World Model · Video Generation Model · Fréchet Video Distance · VBench: Comprehensive Benchmark Suite for Video Generative Models · WorldArena: A Unified Benchmark for Evaluating Perception and Functional Utility of Embodied World Models · AgiBot World
Sources
EWMBench: Evaluating Scene, Motion, and Semantic Quality in Embodied World Models (arXiv 2505.09694)
As of
2025-05

See it in the full glossary →