Calibrated Q-Learning
校准 Q 学习Cal-QLAdvancedAdding a “calibration” constraint to conservative Q-learning, so offline pretraining can transition smoothly into online fine-tuning.
Cal-QL was introduced by Nakamoto, Chelsea Finn, Aviral Kumar, Sergey Levine, and colleagues in 2023, published at NeurIPS 2023, targeting offline-to-online reinforcement learning: pretrain offline on a fixed dataset first, then fine-tune online in the actual environment. The authors found that Q-values learned by pretraining with conservative Q-learning (CQL) are pushed down so far that they end up smaller than the true return of any reasonable policy; as soon as online fine-tuning starts, the Q-values swing wildly in scale, and the policy effectively “forgets” what it learned offline first, with performance dropping before slowly climbing back. Cal-QL adds a calibration constraint: the learned Q-value must still be a lower bound on the current policy's true value, but it may not fall below the value of some reference policy — in practice, the reference policy is just the dataset's behavior policy, with its value estimated from Monte Carlo returns. Implementation-wise, this needs only a one-line change to CQL's code, and it beats existing methods on 9 of 11 fine-tuning benchmark tasks.
ExampleCal-QL's test tasks include FrankaKitchen (controlling a 9-DOF Franka arm through a sequence of kitchen sub-tasks) and Adroit (a 28-DOF five-fingered hand spinning a pen, opening a door), all following the same offline-pretrain-then-online-fine-tune pipeline.
- Also called
- Cal-QL, Calibrated Offline RL Pre-Training
- Related
- Conservative Q-Learning · Offline-to-Online Reinforcement Learning · Offline Reinforcement Learning · Q-Function · Reinforcement Fine-Tuning (RL Fine-Tuning) · Reinforcement Learning with Prior Data
- Sources
- Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning (arXiv 2303.05479)