Diffusion Step Distillation
扩散步数蒸馏DMDAdvancedCompressing a diffusion model that needs dozens to hundreds of denoising steps into a student that produces results in one or a few.
Generating one sample from a diffusion model normally takes dozens to thousands of repeated denoising steps, which is slow. Step distillation uses a trained multi-step model as a teacher to train a student that needs far fewer steps. An early example is Google's Salimans and Ho's progressive distillation (ICLR 2022), which halves the number of steps each round, compressing from as many as 8,192 steps down to 4. Distribution Matching Distillation (DMD), from MIT and Adobe's Yin and colleagues (CVPR 2024), doesn't require the student to reproduce the teacher's denoising trajectory sample-by-sample; instead, it pulls the overall output distribution of a one-step generator toward the teacher's, approximating the gradient of the KL divergence using the difference between two diffusion models' scores (the gradient of the data distribution); a later version, DMD2, adds a GAN loss and supports multiple steps. In robotics, a diffusion policy's slow inference can bottleneck the control frequency, so this kind of distillation is commonly used to speed it up; autoregressive video-generation work like CausVid and Self Forcing also uses DMD-style losses to achieve real-time generation. Consistency distillation is another approach to the same problem.
ExampleOne-Step Diffusion Policy uses a trained diffusion policy as its teacher and distills a one-step generator at an extra cost of just 2%–10% of pretraining, raising the action output rate on real Franka-arm tasks from about 1.5 Hz to 62 Hz.
- Also called
- DMD, Distribution Matching Distillation, Progressive Distillation, Few-Step Distillation
- Related
- Diffusion Model · Denoising Steps · Consistency Model · One-step Generation · Knowledge Distillation · Consistency Policy
- Sources
- One-step Diffusion with Distribution Matching Distillation (arXiv 2311.18828)
Progressive Distillation for Fast Sampling of Diffusion Models (arXiv 2202.00512)
One-Step Diffusion Policy: Fast Visuomotor Policies via Diffusion Distillation (arXiv 2410.21257)