Embodied AI Glossary中文

Offline Reinforcement Learning

离线强化学习Offline RLCommon

Training a policy using only a fixed, previously collected dataset, with no further interaction with the environment during training.

Earlier also called batch reinforcement learning. A 2020 survey by Sergey Levine and colleagues defines it as reinforcement learning that uses only pre-collected data, with no additional online data collection at all. The data can come from human teleoperation, an older policy, or deployment logs, and training never touches the real environment, which suits robotics, healthcare, and other settings where trial and error is expensive or dangerous. Unlike imitation learning, it makes use of the reward signal, so it has a chance of learning a policy better than the data itself from a mix of good and mediocre demonstrations. Its central difficulty is distribution shift: once the policy picks an action that never appears in the dataset, the Q-function's estimate for it tends to be inflated (extrapolation error), and these errors can compound. Conservative Q-learning (CQL), implicit Q-learning (IQL), and various policy-constraint methods were all designed to address this.

ExampleConservative Q-learning (CQL) adds a regularization term to ordinary Q-learning that deliberately pushes down the Q-values of actions not seen in the dataset, so the learned value becomes a lower bound on the true value and the policy isn't misled by inflated estimates.

Also called
Offline RL, Batch Reinforcement Learning, Batch RL
Related
Online Reinforcement Learning · Off-Policy · Conservative Q-Learning · Implicit Q-Learning · Offline-to-Online Reinforcement Learning · D4RL
Sources
Offline Reinforcement Learning: Tutorial, Review, and Perspectives on Open Problems (Levine et al., 2020)
Conservative Q-Learning for Offline Reinforcement Learning

See it in the full glossary →