Embodied AI Glossary中文

Consistency Model

一致性模型Advanced

A generative model that maps noise back to data in one step, used to compress diffusion's many sampling steps into one or two.

Consistency models were proposed by OpenAI's Yang Song and colleagues in 2023 (ICML 2023). Generating from a diffusion model means walking step by step along a denoising trajectory, which often takes tens to a hundred steps and is slow. A consistency model instead trains a function so that noisy samples at any point along the same trajectory all map to the same endpoint (the clean data) — a property called 'self-consistency' — so a sample can be produced from pure noise in a single step, or refined further with a few more steps for higher quality. There are two ways to train it: distilling from an already-trained diffusion model (consistency distillation), or training from scratch directly. In robotics, Prasad, Bohg, and colleagues' 2024 Consistency Policy distills Diffusion Policy into a consistency policy, running roughly an order of magnitude faster than even the fastest alternative speedup methods at comparable success rate, which suits robots with limited compute; work such as ConRFT also uses a consistency policy for reinforcement fine-tuning of VLAs.

ExampleConsistency Policy was run on a laptop GPU across 6 simulated tasks and 3 real-robot tasks, at roughly 10x the speed of other acceleration methods and a success rate close to the original Diffusion Policy.

Also called
Consistency Distillation
Related
Diffusion Model · Diffusion Policy · One-step Generation · Knowledge Distillation · MeanFlow · ConRFT
Sources
Consistency Models (arXiv 2303.01469)
Consistency Policy: Accelerated Visuomotor Policies via Consistency Distillation (arXiv 2405.07503)

See it in the full glossary →