On-Policy
同策略CommonUpdating only with data the current policy just collected itself, and discarding old data once it's used.
The counterpart to off-policy: the behavior policy producing the data and the target policy being optimized are the same one. Once the parameters are updated, data collected by the old policy no longer matches the current policy's distribution, so it must be discarded and fresh data collected. SARSA, REINFORCE, A2C/A3C, TRPO, and PPO are all on-policy algorithms; PPO does run several mini-batch updates on the same batch of data, but it's still considered on-policy. The advantage is stable, simple training; the disadvantage is low data efficiency, requiring huge amounts of interaction. Once GPU-based massively parallel simulation made sampling cheap, PPO became the mainstream algorithm for training legged and humanoid locomotion control.
ExampleRudin and colleagues (2021) ran thousands of ANYmal quadrupeds in parallel simulation on a single GPU, training a flat-ground walking policy in under 4 minutes and a rough-terrain policy in about 20, then transferred it to the real robot; their open-source legged_gym pairs with rsl_rl, a PPO implementation.
- Also called
- On-Policy Learning
- Related
- Off-Policy · Proximal Policy Optimization · Policy Gradient · Massively Parallel Reinforcement Learning · Sample Efficiency · rsl_rl
- Sources
- OpenAI Spinning Up: Kinds of RL Algorithms
Proximal Policy Optimization Algorithms
Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning