Self Forcing
自强制AdvancedTraining a video model by having it keep generating from its own previously generated frames, removing the train-test mismatch.
Self Forcing was proposed by Adobe Research and UT Austin's Xun Huang and colleagues in June 2025 (NeurIPS 2025 Spotlight), as an autoregressive video-diffusion training method. Autoregressive video models generate a segment at a time; earlier training used teacher forcing, conditioning on real preceding frames, or diffusion forcing, conditioning on noised real preceding frames, but at inference time the model can only condition on its own previously generated, imperfect frames, a mismatch called exposure bias that makes long videos degrade further and further. Self Forcing instead unrolls the model autoregressively during training the same way it will run at inference, using a KV cache and conditioning on its own generated frames, then computes one overall loss across the whole video. Built on Wan2.1-T2V-1.3B, it achieves sub-second-latency, real-time streaming generation on a single GPU, which matters for interactive world models that need to generate video while responding to actions.
ExampleAccording to the project page, the Self Forcing model generates 480p video at about 16 frames per second on a single H100, with a first-frame latency of about 0.8 seconds, and can also stream in real time on a single RTX 4090.
- Also called
- Self-Forcing Training
- Related
- Teacher Forcing · Diffusion Forcing · Exposure Bias · Autoregressive Video Generation · Key-Value Cache · Diffusion Step Distillation
- Sources
- Huang et al. 2025: Self Forcing: Bridging the Train-Test Gap in Autoregressive Video Diffusion
Self Forcing 项目页 (Chinese)
GitHub: guandeh17/Self-Forcing - As of
- 2025-11