Embodied AI Glossary中文

Offline-to-Online Reinforcement Learning

离线到在线强化学习O2O RLAdvanced

Pretraining with offline reinforcement learning on existing data, then letting the robot keep improving through live interaction.

Offline-to-online reinforcement learning happens in two stages: first, offline RL on data that already exists — human demonstrations, historical logs, or an older policy's trajectories — produces an initial policy and value function that does not have to explore from scratch; then the agent interacts with the real environment to collect new data and fine-tunes online. The benefit is skipping a lot of dangerous, expensive random exploration, which matters enormously for real robots. The hard part is the handoff between the two stages: the offline stage is deliberately conservative to avoid overestimation, so its value scale is often inaccurate, and performance commonly dips right after switching to online training. Representative methods include AWAC (2020), Calibrated Q-Learning (Cal-QL, 2023), and RLPD (2023), which mixes offline data directly into the online replay buffer and trains from scratch. Real-robot RL systems such as SERL and HIL-SERL, and π*0.6's RECAP, all follow this same two-stage idea.

ExampleSERL uses RLPD as its core algorithm: it starts with 20 demonstrations teleoperated with a SpaceMouse, then trains online on the real robot, learning tasks such as PCB assembly or cable routing in 25 to 50 minutes on average.

Also called
O2O RL, Offline-to-Online Fine-tuning
Related
Offline Reinforcement Learning · Online Reinforcement Learning · Reinforcement Learning with Prior Data · Calibrated Q-Learning · Real-World Reinforcement Learning · SERL
Sources
Nakamoto et al. 2023: Cal-QL: Calibrated Offline RL Pre-Training for Efficient Online Fine-Tuning
Ball et al. 2023: Efficient Online Reinforcement Learning with Offline Data (RLPD)
Luo et al. 2024: SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning

See it in the full glossary →