Embodied AI Glossary中文

3D Diffusion Policy

3D 扩散策略DP3Common

A diffusion policy that takes sparse point clouds as input and learns manipulation from very few demonstrations.

3D Diffusion Policy (DP3) was proposed in March 2024 by researchers at the Shanghai Qi Zhi Institute and Huazhe Xu's group at Tsinghua University's Institute for Interdisciplinary Information Sciences, together with Shanghai Jiao Tong University and the Shanghai AI Lab (first author Yanjie Ze), and published at RSS 2024. Building on Diffusion Policy, it replaces 2D images with a sparse point cloud from a depth camera (a set of points with 3D coordinates) as input, compresses it with a lightweight point-cloud encoder into a compact 3D feature, and denoises actions conditioned on that feature. Because a 3D representation directly captures an object's position and geometry, it's less sensitive to changes in viewpoint or appearance, so it needs fewer demonstrations. The paper used just 10 demonstrations per task across 72 simulated tasks, improving 24.2% over baselines, and 40 demonstrations per task on 4 real-robot tasks, reaching 85% success with very few safety-violating actions. A later extension, iDP3, scales it up to humanoid robots.

ExampleReal-robot experiments included rolling play-dough and folding it into a “dumpling” with an Allegro dexterous hand or a parallel gripper, touching a block with a power drill, and pouring out dried pork floss — each task using 40 demonstrations.

Also called
DP3, 3D Diffusion Policy: Generalizable Visuomotor Policy Learning via Simple 3D Representations
Related
Diffusion Policy · Point Cloud · Point Cloud Encoder · iDP3 · Imitation Learning · Shanghai Qi Zhi Institute
Sources
3D Diffusion Policy (arXiv 2403.03954)
3D Diffusion Policy 项目主页 (Chinese)
As of
2024-09

See it in the full glossary →