UVA
AdvancedA robot model in which video and action share one latent representation, letting inference skip video generation for speed.
UVA was released by Shuang Li, Yihuai Gao, Dorsa Sadigh, and Shuran Song at Stanford University in February 2025, published at RSS 2025. Using video generation to build a robot policy lets the model learn how the environment changes, but action precision and inference speed have typically lagged behind policies that output actions directly. UVA has video and action share one joint latent representation, learned together during training but decoded separately: two lightweight diffusion heads decode the future frame and the action independently, and at inference time video generation can simply be skipped and only the action decoded, which keeps it fast. Combined with masked training, the same model can then serve as a policy, a forward dynamics model, an inverse dynamics model, or a video prediction model. Across real-robot multi-task experiments like folding a towel and arranging cups, it beats a diffusion-policy baseline trained on UMI data when facing unseen environments and objects.
ExampleWhile a robot folds a towel, UVA decodes only the next action and skips frame generation to stay real-time; offline, that same model can also predict the resulting image given a specified action.
- Also called
- Unified Video Action Model
- Related
- Video Prediction Policy · World Action Model · Forward Dynamics Model · Inverse Dynamics Model · Diffusion Policy · Universal Manipulation Interface
- Sources
- Unified Video Action Model (arXiv 2503.00200)
UVA project page - As of
- 2025-04