Embodied AI Glossary中文

UWM

UWM(统一世界模型)Advanced

A robot framework that merges action diffusion and video diffusion into one model and can pretrain on action-free video.

UWM was released in April 2025 by the University of Washington (Abhishek Gupta's group) and Toyota Research Institute, published at RSS 2025. Imitation learning needs robot data labeled with actions, while the huge amount of video without action labels is hard to use directly. UWM runs action diffusion and video diffusion inside one Transformer at the same time, giving each modality its own independent denoising timestep: setting a modality's timestep to pure noise is equivalent to not looking at it at all. This lets the same model act as a policy, a forward dynamics model, an inverse dynamics model, or a video predictor, just by combining timesteps differently; for action-free video, the missing action is simply treated as fully noised during training. Pretrained on the large-scale DROID dataset and then fine-tuned, it generalizes better than pretraining with plain behavior cloning, and adding action-free video on top improves results further.

ExampleRobot trajectories from DROID are pretrained together with a batch of manipulation videos that have no action labels to train UWM, which is then fine-tuned into a task-specific policy using a small number of demonstrations.

Also called
Unified World Models, Unified World Models: Coupling Video and Action Diffusion for Pretraining on Large Robotic Datasets
Related
World Model · Diffusion Model · Action-free Video · Inverse Dynamics Model · DROID (Distributed Robot Interaction Dataset) · UVA
Sources
Unified World Models (arXiv 2504.02792)
UWM project page
As of
2025-05

See it in the full glossary →