Embodied AI Glossary中文

Advantage-Weighted Regression

优势加权回归AWRAdvanced

Doing imitation learning weighted by each action's advantage, so higher-advantage actions get imitated more.

AWR was introduced by Xue Bin Peng, Aviral Kumar, Grace Zhang, and Sergey Levine in 2019, aiming to do reinforcement learning using only supervised-learning regression steps. Each round has two steps: fit the value function by regression, then do weighted behavior cloning, where each action in the data is weighted by exp(advantage/β), so higher-advantage actions get imitated more heavily, with β a temperature coefficient. It can reuse old data from a replay buffer (off-policy), and can also learn from a fixed dataset alone; it needs only a few lines of code to implement, and supports both continuous and discrete actions. This “estimate the advantage, then imitate with weighting” style of policy extraction has been widely adopted by later offline reinforcement-learning methods — implicit Q-learning (IQL), for instance, uses advantage-weighted behavior cloning as its final policy-extraction step.

ExampleImplicit Q-learning first learns a Q-function and a value function from offline data, then extracts the final policy with advantage-weighted behavior cloning: actions in the data with a larger positive advantage get a higher weight during imitation.

Also called
AWR, Advantage-Weighted Behavioral Cloning
Related
Advantage Function · Offline Reinforcement Learning · Implicit Q-Learning · Behavior Cloning · Policy Constraint · Advantage Conditioning
Sources
Advantage-Weighted Regression: Simple and Scalable Off-Policy Reinforcement Learning (arXiv 1910.00177)
Offline Reinforcement Learning with Implicit Q-Learning (arXiv 2110.06169)

See it in the full glossary →