Embodied AI Glossary中文

Generalized Advantage Estimation

广义优势估计GAEAdvanced

Estimating the advantage function by summing multi-step temporal-difference errors with exponentially decaying λ weights.

GAE was introduced by Schulman, Levine, Abbeel, and colleagues in 2015. Policy gradient methods need to know how much better than average a given action was — the advantage function. Estimating it from just a single-step temporal-difference (TD) error, δ = r + γV(s′) − V(s), gives low variance but high bias; using the full episode's Monte Carlo return gives low bias but high variance. GAE sums the TD errors of future steps, each weighted by (γλ)^k, with λ between 0 and 1 tuning the trade-off between the two: λ = 0 recovers single-step TD, and λ = 1 recovers Monte Carlo — the same idea as TD(λ). It's PPO's standard configuration, and it's the near-universal way PPO computes advantages when training locomotion control for legged and humanoid robots.

Examplelegged_gym's default PPO configuration uses a discount factor γ = 0.99 and a GAE parameter λ = 0.95.

Also called
GAE, GAE(λ)
Related
Advantage Function · Proximal Policy Optimization · Temporal-Difference Learning · Discount Factor · Value Function · Policy Gradient
Sources
High-Dimensional Continuous Control Using Generalized Advantage Estimation (arXiv:1506.02438)
legged_gym: legged_robot_config.py

See it in the full glossary →