Fast-WAM
AdvancedA 2026 Tsinghua and Galaxea world action model that learns video prediction during training but skips imagining the future at inference, making it faster.
Fast-WAM is a paper released in March 2026 by Tianyuan Yuan, Hang Zhao, and colleagues at Tsinghua University's Institute for Interdisciplinary Information Sciences together with Galaxea. Most world action models (WAMs, models that jointly predict future video and robot actions) follow an “imagine, then act” approach: a video diffusion model generates a few future frames, and actions are produced from those, with the repeated denoising making inference slow. The authors separated two questions to test independently: whether to jointly learn video prediction during training, and whether to actually generate future frames at inference. They found that dropping the imagination step at inference barely hurts performance, while dropping joint video training during training hurts it substantially — showing that video prediction's real value lies in training a better world representation, not in generating images at test time. The model uses the Wan2.2-5B video diffusion Transformer as its backbone plus a roughly 1-billion-parameter action expert, about 6 billion parameters in total; inference latency is 190 milliseconds, more than 4 times faster than comparable “imagine, then act” models.
ExampleWith no robot-data pretraining at all, Fast-WAM reaches a 91.8% success rate on RoboTwin 2.0 and an average of 97.6% on LIBERO, and completed a long-horizon real-robot towel-folding task on the Galaxea R1 Lite.
- Also called
- Do World Action Models Need Test-time Future Imagination?
- Related
- World Action Model · Video Generation Model · Action Expert · Inference Latency · LingBot-VA (Robbyant) · Motus
- Sources
- Fast-WAM: Do World Action Models Need Test-time Future Imagination? (arXiv 2603.16666)
- As of
- 2026-03