Diffuser
Diffuser(扩散规划器)AdvancedA planning method that generates an entire trajectory as one object to be denoised, one of the first uses of diffusion for decision-making.
Diffuser is an ICML 2022 paper by Michael Janner and Sergey Levine at UC Berkeley, and Yilun Du and Joshua Tenenbaum at MIT. Traditional model-based reinforcement learning first learns a dynamics model and then plans with an optimizer, and errors in the model are often amplified by the planner. Diffuser merges the two steps: a diffusion model directly models an entire “state plus action” trajectory, and planning becomes starting from noise and repeatedly denoising it into a trajectory. To get high return, the denoising process is guided with the gradient of a value function; to reach a specific goal, the start and end states are fixed and the model fills in the middle (similar to image inpainting). The same model can switch tasks with no retraining. It's a precursor to later “diffusion model for decision-making” work such as Decision Diffuser and Diffusion Policy.
ExampleIn the Maze2D task, fixing the start and end states, Diffuser “inpaints” a full feasible path in between by denoising; in a block-stacking task, swapping in a different guidance function lets the same model stack towers under different constraints.
- Also called
- Planning with Diffusion for Flexible Behavior Synthesis
- Related
- Diffusion Model · Model-Based Reinforcement Learning · Trajectory Optimization · Decision Diffuser · Diffusion Policy · Offline Reinforcement Learning
- Sources
- Planning with Diffusion for Flexible Behavior Synthesis (arXiv:2205.09991)
Diffuser 项目主页 (Chinese) - As of
- 2022