Advantage Function
优势函数CommonA measure of how much better an action is than the policy's average, equal to the Q-value minus the value function.
The advantage function is defined as A(s,a) = Q(s,a) − V(s). Q is the expected return of taking action a in state s and then following the policy afterward; V is the expected return of just following the policy from state s directly. Subtracting the two tells you how much better than average that specific action is: positive means do more of it, negative means do less. Policy-gradient methods use the advantage in place of the raw return because it substantially reduces the variance (noisiness) of the gradient estimate, which makes training more stable. Algorithms like PPO typically compute it with Generalized Advantage Estimation (GAE), introduced by Schulman and colleagues in 2015; Physical Intelligence's RECAP instead feeds the advantage into a VLA as a conditioning input, letting the model distinguish good experience from bad.
ExampleWhen an arm reaches for a cup in state s, the current policy earns an average return of 0.6. If leading with a straight vertical approach raises the expected return to 0.8, that action's advantage is +0.2, and training will increase the probability of picking it.
- Also called
- Advantage, A(s,a)
- Related
- Value Function · Q-Function · Generalized Advantage Estimation · Policy Gradient · Proximal Policy Optimization · RECAP
- Sources
- OpenAI Spinning Up: Key Concepts in RL
High-Dimensional Continuous Control Using Generalized Advantage Estimation (arXiv 1506.02438)