Embodied AI Glossary中文

Policy Gradient

策略梯度PGCommon

Computing the gradient of expected return with respect to the policy's parameters directly, and improving the policy along that gradient.

This is a major family of reinforcement-learning methods that optimize the policy directly. The policy is represented by a parameterized network, and the goal is to maximize expected return; the policy gradient gives the gradient of that objective with respect to the parameters, by multiplying the gradient of the log-probability of each action taken by the return that followed it — so actions that led to high return get their probability raised, and actions that led to low return get it lowered — and this can be estimated from sampled trajectories. Williams's 1992 REINFORCE is an early example, and Sutton and colleagues gave the policy gradient theorem for use with function approximation in 1999. In practice, a baseline is usually subtracted, using the advantage function (how much better this action is than average) in place of the raw return to reduce variance, which led to actor-critic methods, TRPO, and PPO. It handles continuous actions naturally, suiting robot control, but is typically on-policy and not very sample-efficient.

ExampleTraining a robot arm to push a box: sample a batch of trajectories with the current policy, raise the probability of actions taken in trajectories that pushed the box to the goal, lower it for actions in failed trajectories, and repeat — the policy gradually improves.

Also called
PG, Policy Gradient Methods
Related
REINFORCE · Advantage Function · Proximal Policy Optimization · Trust Region Policy Optimization · On-Policy · Generalized Advantage Estimation
Sources
OpenAI Spinning Up: Intro to Policy Optimization
Policy Gradient Methods for Reinforcement Learning with Function Approximation (Sutton et al., NIPS 1999)

See it in the full glossary →