Classifier-Free Guidance
无分类器引导CFGAdvancedComputing both a conditional and an unconditional prediction and extrapolating between them so generation follows the condition more closely.
Classifier-free guidance was proposed by Jonathan Ho and Tim Salimans (a 2021 NeurIPS workshop paper, with a full arXiv version in 2022), for use with diffusion models, flow matching, and other generative models that denoise step by step. During training, the condition (such as a text prompt) is randomly dropped for some examples, so the same network learns to make both conditional and unconditional predictions; during generation, both are computed at every step, and the final direction is 'unconditional result + w × (conditional result − unconditional result),' where w is called the guidance scale. A larger w follows the condition more closely but reduces diversity, and too large a value introduces artifacts. It replaces the earlier 'classifier guidance,' which needed training a separate classifier, and has become a standard setting in text-to-image and text-to-video models, at the cost of one extra forward pass per step. In embodied AI, diffusion-based video world models and action generation models can also use it to control how closely they follow a language instruction.
ExampleGenerating an image from text with Stable Diffusion v1.5 in Diffusers, turning guidance_scale up from 2.5 to 10.5 makes the image follow the prompt more and more closely, though artifacts start to appear once it's too high.
- Also called
- CFG, Guidance Scale
- Related
- Diffusion Model · Flow Matching · Denoising Steps · Text-to-Video / Image-to-Video · Generative Model · Video Generation Model
- Sources
- Classifier-Free Diffusion Guidance (arXiv 2207.12598)
Hugging Face Diffusers: Text-to-image(guidance_scale 说明) (Chinese)