Embodied AI Glossary中文

Proximal Policy Optimization

近端策略优化PPOEssential

OpenAI's 2017 reinforcement-learning algorithm that limits how much each policy update can change the policy, making training stable.

PPO was proposed in 2017 by Schulman and colleagues at OpenAI. It's a policy-gradient method (one that directly adjusts a policy network so that actions leading to higher return get chosen more often), and it's on-policy: it only trains on data just collected by the current version of the policy. Its core idea is an objective function with “clipping”: once the new policy changes an action's probability beyond a certain range relative to the old policy, the objective stops rewarding further change in that direction. This keeps a single update from wrecking the policy, while still allowing multiple passes of training over the same batch of data. It's about as stable as TRPO (Trust Region Policy Optimization) but far simpler to implement. PPO is the workhorse algorithm for training legged and humanoid robot locomotion in simulation — ETH Zurich's open-source robot RL library rsl_rl is built around it — and it's also the algorithm behind InstructGPT's RLHF.

ExampleETH Zurich's Rudin and colleagues used PPO in legged_gym, running thousands of simulated environments in parallel on a single GPU, to train an ANYmal quadruped: a flat-ground walking policy trained in under 4 minutes and a rough-terrain policy in about 20, both then transferred to the real robot.

Also called
PPO-Clip, PPO
Related
Reinforcement Learning · Policy Gradient · Trust Region Policy Optimization · On-Policy · Generalized Advantage Estimation · rsl_rl
Sources
Schulman et al. 2017: Proximal Policy Optimization Algorithms
OpenAI Spinning Up: Proximal Policy Optimization
Rudin et al. 2021: Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning

See it in the full glossary →