Embodied AI Glossary中文

Advantage Conditioning

优势条件化Advanced

Telling the policy how “good” each action was during training, then at deployment asking it only to generate “good” actions.

Advantage conditioning turns reinforcement learning into conditional supervised learning: first train a value function and use it to compute each action's advantage in the dataset (how much better than average it was), then feed that advantage, or a binarized version of it, into the policy as an extra input and train with ordinary imitation learning. This way both good and bad data contribute to training, and at inference, setting the condition to “good” biases the policy toward high-advantage actions. The idea is close to return-conditioned methods like Decision Transformer, and related to CFGRL, proposed by Frans and colleagues in 2025, which treats a diffusion model's classifier-free guidance as a form of policy improvement. Physical Intelligence's π*0.6 (November 2025) builds its RECAP method on this: it binarizes the advantage by a threshold, feeds it to the policy as the text “Advantage: positive / negative,” randomly drops the condition during training, and sets it to positive at inference — or amplifies the effect further with classifier-free guidance.

ExampleAfter training with RECAP, π*0.6 can fold laundry in real homes, reliably assemble cardboard boxes, and make espresso on a professional machine; the paper reports throughput more than doubling on some of the hardest tasks and failure rate roughly halving.

Also called
Advantage-Conditioned Policy
Related
RECAP · π*0.6 · Advantage Function · Return Conditioning · Classifier-Free Guidance · Advantage-Weighted Regression
Sources
π*0.6: a VLA That Learns From Experience (RECAP, arXiv 2511.14759)
Diffusion Guidance Is a Controllable Policy Improvement Operator (CFGRL, arXiv 2505.23458)
As of
2025-11

See it in the full glossary →