Diffusion Model
扩散模型EssentialA generative model that learns to remove noise step by step, then generates new data by denoising from pure noise.
A diffusion model has two processes: a forward process that keeps adding Gaussian noise to real data until it becomes pure noise, and a reverse process where a neural network is trained to predict and remove the noise at each step. Generation runs the reverse process from random noise, denoising repeatedly to produce a new sample. The idea was proposed by Sohl-Dickstein and colleagues in 2015, and Ho and colleagues' 2020 DDPM made it genuinely practical, after which it became the mainstream approach behind image and video generators such as Stable Diffusion. In robotics it is used to generate actions: 2023's Diffusion Policy denoises a segment of future action conditioned on the current observation, and can represent several equally valid ways of doing something in the same scene, called action multimodality, which a single regressed average cannot capture. The cost is that denoising takes multiple steps, so inference tends to be slower.
ExampleDiffusion Policy (Chi and colleagues, 2023) reached an average success rate 46.9% higher than prior methods across 12 manipulation tasks in 4 benchmarks; on a T-block-pushing task, faced with two equally valid ways to push, going around the left or the right, it learns both modes instead of averaging them into one.
- Also called
- Diffusion
- Related
- Diffusion Policy · Denoising Diffusion Probabilistic Model · Flow Matching · Action Multimodality · Denoising Steps · Generative Model
- Sources
- Denoising Diffusion Probabilistic Models (arXiv 2006.11239)
Diffusion Policy: Visuomotor Policy Learning via Action Diffusion (arXiv 2303.04137)
Diffusion model - Wikipedia