Embodied AI Glossary中文

Cross-Entropy Method

交叉熵方法CEMAdvanced

A gradient-free optimizer that repeatedly samples a batch, keeps the best, and refits the sampling distribution to them.

The cross-entropy method was proposed by Rubinstein in the late 1990s, originally for estimating rare-event probabilities, and later became a general stochastic optimization method. The procedure: sample a batch of candidate solutions from a distribution (usually Gaussian), score each one, keep the best fraction of them (the ‘elite’ samples), refit the distribution's mean and variance from the elites, then sample again; after a few rounds, the distribution concentrates around good solutions. It needs no gradients and is easy to parallelize, making it a good fit when the objective is a simulator or a neural network — a black box. Two common uses in robotics: as a sampling-based MPC, running CEM over a sequence of future actions and executing only the first step before replanning; and finding the action that maximizes a Q-function in a continuous action space. It differs from MPPI in that CEM uses only the elite samples with equal weighting, while MPPI weights and uses all samples according to an exponential function of their cost.

ExamplePlaNet plans inside a learned world model using CEM: a 12-step horizon, sampling 1,000 action sequences per round, keeping the best 100, over 10 iterations. QT-Opt uses CEM to find the best grasp action on a Q-function, sampling 64 per round, keeping the best 6, over 2 iterations, used both for computing target values during training and for selecting actions on the real robot.

Also called
CEM, CEM Planning, CEM Optimization
Related
Sampling-based MPC · Model Predictive Path Integral Control · Model Predictive Control · Model-Based Reinforcement Learning · PlaNet · QT-Opt
Sources
Wikipedia: Cross-entropy method
Hafner et al., Learning Latent Dynamics for Planning from Pixels (PlaNet, arXiv:1811.04551)
Kalashnikov et al., QT-Opt: Scalable Deep Reinforcement Learning for Vision-Based Robotic Manipulation (arXiv:1806.10293)

See it in the full glossary →