World Action Model
世界动作模型WAMCommonA model that predicts future frames and robot actions together, using its prediction about the world to guide the action.
The term World Action Model was formally introduced in NVIDIA's February 2026 DreamZero paper: any model that uses world-modeling ability (predicting future state) to predict actions counts as a WAM, a category that retroactively covers earlier video-and-action joint models like GR-1, UWM, and Cosmos Policy. The typical design starts from a pretrained video generation model as backbone, adds action input and output, and trains it to predict future video and actions together. Compared with a VLA built from a vision-language model, a WAM learns physical dynamics from video; DreamZero reported more than double the generalization to new tasks and environments versus the best VLA at the time. It isn't called a 'video-action model' because the predicted target could someday be touch or force instead of video. One open question is whether inference actually needs to generate future frames: Fast-WAM does the joint video-and-action prediction only during training and skips it at inference, reaching similar performance more than 4x faster.
ExampleDreamZero uses Wan2.1's 14B image-to-video model as its backbone; given a camera view and a language instruction, it generates future video and actions at the same time, controlling a real robot in real time at 7Hz.
- Also called
- WAM, Video-Action Model, VAM
- Related
- Vision-Language-Action Model · World Model · Video Generation Model · DreamZero · Fast-WAM · Video Prediction Policy
- Sources
- World Action Models are Zero-shot Policies (DreamZero, arXiv 2602.15922)
Fast-WAM: Do World Action Models Need Test-time Future Imagination? (arXiv 2603.16666) - As of
- 2026-03