Embodied AI Glossary中文

Cross-Entropy

交叉熵CECommon

A measure of how far a model's predicted probability distribution is from the correct answer; the standard loss for classification.

Cross-entropy comes from information theory, where it measures the difference between two probability distributions. In deep learning it's used as a classification loss: the model outputs a probability for each class (usually normalized with softmax), and the loss grows the lower the predicted probability assigned to the correct class is; when there's exactly one correct class per label, it's equivalent to negative log-likelihood (NLL). A large language model's next-token prediction is cross-entropy computed over the entire vocabulary. In embodied AI, VLA models that discretize continuous actions into tokens, such as RT-2 and OpenVLA, also train their action outputs with cross-entropy; policies that regress continuous actions directly usually use mean squared error or L1 loss instead, and diffusion and flow-matching policies use their own denoising losses.

ExampleOpenVLA divides each action dimension into 256 bins spanning the 1st to 99th percentile of the training data, maps them onto the 256 least-used tokens in the Llama vocabulary, and during training computes cross-entropy loss only on those action tokens.

Also called
CE, Cross-Entropy Loss, Log Loss
Related
Loss Function · Next-Token Prediction · Action Binning · Mean Squared Error · Kullback-Leibler Divergence · Maximum Likelihood Estimation
Sources
Google Machine Learning Glossary: cross-entropy
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)

See it in the full glossary →