Extrapolation Error (OOD Actions in Offline RL)
外推误差AdvancedThe error in offline reinforcement learning that comes from a Q-network guessing wildly at the value of actions absent from the data.
This concept was systematically laid out by Fujimoto, Meger, and Precup in their 2019 ICML paper introducing the BCQ algorithm. Offline reinforcement learning can only train on a fixed dataset, with no further environment interaction. Q-learning's update needs max_a Q(s′, a) over the next state, but the network's estimate for actions absent from the data (out-of-distribution actions) is just an extrapolation, and can be badly inflated; the policy then gravitates toward exactly those actions, the error compounds through repeated Bellman backups, and the resulting policy ends up poor. Online training can correct this by actually trying the action; offline training cannot. The main countermeasures are constraining the policy to stay close to the data's behavior (BCQ, policy constraints), pushing down the Q-values of out-of-distribution actions (CQL), and avoiding querying out-of-distribution actions altogether (IQL).
ExampleIn the BCQ paper, an offline DDPG agent trained on exactly the same batch of data as an online DDPG agent falls clearly behind on every task, with value estimates that are unstable or even diverge — the authors attribute this to extrapolation error.
- Also called
- Out-of-Distribution Action Problem, OOD Action Overestimation
- Related
- Offline Reinforcement Learning · Overestimation Bias · Conservative Q-Learning · Implicit Q-Learning · Policy Constraint · Out-of-Distribution
- Sources
- Off-Policy Deep Reinforcement Learning without Exploration (BCQ, arXiv:1812.02900)
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems (arXiv:2005.01643)