DreamZero
AdvancedNVIDIA's 2026 world action model, built on a 14B video diffusion model that predicts frames and actions together.
DreamZero is a world action model (WAM, a model that jointly predicts future video frames and actions) released in February 2026 by an NVIDIA research team including Jim Fan and Yuke Zhu, in a paper titled World Action Models are Zero-shot Policies. Mainstream VLA models use a vision-language model as their backbone, which is good at understanding semantics but has limited grasp of physical dynamics. DreamZero instead uses a pretrained video diffusion model as its backbone (Tongyi Wanxiang's Wan2.1 image-to-video model, 14B parameters), autoregressively generating future video and robot actions together, treating video as dense supervision for “how the world changes” — which lets it learn from diverse, non-repetitive, heterogeneous robot data. The paper reports that its generalization to new tasks and new environments is more than double that of the best VLA models at the time; after model and system optimization, a single inference takes about 150 milliseconds on a GB200, enabling 7Hz closed-loop control. The team says it will open-source the model weights and inference code.
ExampleTrained mainly on about 500 hours of teleoperation data from the AgiBot G1, it reaches an average task-progress score of 62.2% on tasks seen during training, versus 27.4% for the best VLA baseline; adding just 10–20 minutes of video from other robots or humans lifts performance on unseen tasks by more than 42% relative, and just 30 minutes of play data is enough to adapt it to the new YAM robot.
- Also called
- World Action Models are Zero-shot Policies
- Related
- World Action Model · Video Generation Model · Wan (Alibaba Video Generation Model) · Zero-shot · Cross-Embodiment · Fast-WAM
- Sources
- World Action Models are Zero-shot Policies (arXiv 2602.15922)
DreamZero 项目主页 (Chinese) - As of
- 2026-02