Autoregressive Video Generation
自回归视频生成AdvancedGenerating video forward in time, one frame or chunk at a time, so it can be played as it's produced.
Many video diffusion models generate a whole clip at once, with frames attending to each other in both directions, so they must finish the entire computation before producing any output, and it's hard to feed in new actions partway through. Autoregressive video generation instead writes forward in time: each frame or short chunk is conditioned only on what's already been generated, and combined with causal attention and a KV cache, this lets the model stream output and respond in real time to a user's or robot's actions — exactly what an interactive world model needs. The main difficulty is error accumulation, called exposure bias: during training the model sees real history, but at inference it sees its own generated, imperfect history, and image quality can collapse over time. CausVid (CVPR 2025) distills a bidirectional model into a 4-step causal model, running at 9.4 frames per second on a single GPU; Self Forcing (NeurIPS 2025) uses the model's own outputs as history during training, closing the gap between training and inference.
ExampleGoogle DeepMind's Genie 3 generates frames autoregressively one at a time, with each new frame conditioned on an ever-growing history, and can run interactively in real time at 720p and 24 frames per second.
- Also called
- Causal Video Generation, Streaming Video Generation
- Related
- Diffusion Forcing · Self Forcing · Interactive World Model · Video Generation Model · Exposure Bias · Key-Value Cache
- Sources
- From Slow Bidirectional to Fast Autoregressive Video Diffusion Models (CausVid, arXiv 2412.07772)
Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion (arXiv 2506.08009)
Genie 3: A new frontier for world models (Google DeepMind) - As of
- 2025-08