Maximum Likelihood Estimation
最大似然估计 / 负对数似然MLE / NLLCommonFinding the parameters that make the observed data most probable; taking the negative log of that probability gives a loss to minimize.
Maximum likelihood estimation (MLE) is statistics' most basic method for estimating parameters: assume the data comes from some probability model with unknown parameters, and pick the parameters that make the probability of the observed data — the likelihood — as large as possible. In practice, a logarithm turns the product of probabilities into a sum, and a minus sign turns it into a minimization problem — this is exactly the negative log-likelihood (NLL) loss used in deep learning. Many familiar losses are special cases of it: cross-entropy for classification is the negative log-likelihood under a categorical distribution; when errors are assumed Gaussian, MLE is equivalent to minimizing mean squared error; when errors are assumed to follow a Laplace distribution, it corresponds to L1 loss. It's also equivalent to minimizing the KL divergence between the data distribution and the model's distribution. In embodied AI, behavior cloning is essentially maximum likelihood estimation on expert actions: a VLA that discretizes actions predicts action tokens with cross-entropy, while a Gaussian policy directly minimizes the negative log-likelihood of the action.
ExampleOpenVLA divides each action dimension into 256 evenly spaced bins between the 1st and 99th percentile of the training data, then uses the standard next-token-prediction objective, computing cross-entropy only on the action tokens — which is maximum likelihood estimation applied to expert actions.
- Also called
- MLE, Negative Log-Likelihood, NLL, NLL Loss
- Related
- Cross-Entropy · Kullback-Leibler Divergence · Mean Squared Error · Behavior Cloning · Loss Function · Gaussian Policy
- Sources
- Maximum likelihood estimation(Wikipedia)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv 2406.09246)