Embodied AI Glossary中文

Importance Sampling

重要性采样Advanced

Estimating an expectation under one distribution using samples drawn from another, weighted by the ratio of the two probabilities.

This is a Monte Carlo estimation technique. Wanting the expectation of some quantity under a distribution p, but having only samples drawn from a different distribution q, each sample is weighted by p(x)/q(x); the weighted average is still an unbiased estimate of the expectation under p (provided q can actually produce every value p can). In reinforcement learning, p and q are usually two policies: data collected under an old or behavior policy is used to evaluate or update a new policy, with the weight being the ratio of the new and old policies' probabilities for the same action. Both off-policy learning and off-policy evaluation depend on it; the probability ratio r(θ) in the TRPO and PPO objectives is exactly this, and PPO clips that ratio to keep any single update from being too large. Its main problem is that the further apart the two distributions are, the higher the weights' variance gets, so in practice the weights are often truncated or clipped.

ExampleEach round, PPO collects a batch of data with the old policy and then updates on that same batch several times; every sample's loss is multiplied by the new-to-old policy probability ratio, which is clipped to the range [1−ε, 1+ε].

Also called
Importance Weighting
Related
Off-Policy · Proximal Policy Optimization · Off-Policy Evaluation · Policy Gradient · Trust Region Policy Optimization · Monte Carlo Methods
Sources
Importance sampling (Wikipedia)
Policy Gradient Algorithms (Lilian Weng)

See it in the full glossary →