Kullback-Leibler Divergence
KL 散度KLCommonA measure of how far one probability distribution is from another; not symmetric between the two.
KL divergence was introduced by Solomon Kullback and Richard Leibler in 1951; in the discrete case, D_KL(P‖Q) = Σ P(x)·log(P(x)/Q(x)), representing the extra information cost of approximating the true distribution P with a distribution Q. It's always non-negative, and equals zero only when the two distributions are identical, but it isn't symmetric — D_KL(P‖Q) generally doesn't equal D_KL(Q‖P) — so it isn't a distance in the strict mathematical sense. It shows up everywhere in machine learning: cross-entropy equals KL divergence plus P's own entropy, so minimizing cross-entropy is equivalent to minimizing KL; variational autoencoders use a KL term to pull the latent variable toward a standard normal distribution; PPO, RLHF, and some offline reinforcement-learning methods use a KL penalty to keep a new policy from drifting too far from a reference policy; and knowledge distillation uses it to pull the student's output distribution toward the teacher's.
ExampleACT's training loss is the action-reconstruction error plus β times a KL term (the paper uses β = 10); the KL term keeps the style variable z produced by its conditional variational autoencoder close to a standard normal distribution, and at inference time z is simply set to the prior's mean of 0.
- Also called
- KL Divergence, KL, Relative Entropy
- Related
- Cross-Entropy · KL Regularization · Variational Autoencoder · Proximal Policy Optimization · Knowledge Distillation · Maximum Likelihood Estimation
- Sources
- Kullback–Leibler divergence(Wikipedia)
Learning Fine-Grained Bimanual Manipulation with Low-Cost Hardware (ACT, arXiv 2304.13705)