WoW
WoW 具身世界模型AdvancedA 14-billion-parameter embodied video world model trained on 2 million real-robot interaction trajectories.
WoW was released in September 2025 by the Beijing Humanoid Robot Innovation Center together with Peking University and HKUST. The authors argue that watching passive video alone cannot teach real physical intuition, which instead requires learning from large amounts of causally connected interaction; they trained a 14-billion-parameter diffusion Transformer video generation model on 2 million real robot trajectories spanning 12 robot platforms. Because generated video often contains physical errors, the authors use a framework called SOPHIA, which has a vision-language model act as a reviewer, repeatedly checking the output and rewriting the prompt to steer generation toward something more physically plausible; an inverse dynamics model (which infers the action from consecutive frames) then translates the predicted video into robot-arm actions. The team also released WoWBench, a benchmark for evaluating physical consistency and causal reasoning.
ExampleGiven a tabletop image and a manipulation instruction, WoW first generates a video of the action being completed, and then an inverse dynamics module translates that video into the end effector's 7-degree-of-freedom actions for execution.
- Also called
- WoW-DiT, WoW: Towards a World omniscient World model Through Embodied Interaction
- Related
- World Model · Video Generation Model · Inverse Dynamics Model · Diffusion Transformer · Beijing Humanoid Robot Innovation Center · XR-1
- Sources
- WoW (arXiv:2509.22642)
WoW 项目主页 (Chinese) - As of
- 2025-10