Real-World Reinforcement Learning
真机强化学习CommonLetting a robot learn by trial and error directly in the real environment, rather than only training in simulation.
Real-world reinforcement learning means a policy's reinforcement-learning training happens, at least in part, directly on a physical robot. It skips simulator modeling and sidesteps the sim-to-real gap, which suits tasks with complex contact that's hard to simulate, like connector insertion or cable routing. The difficulties are that real-robot data is slow and expensive, so the algorithm must be highly sample-efficient; the scene has to be reset after every episode, by a person or a mechanism; rewards need to be judged automatically from camera images; and exploration can't be allowed to damage the robot. Dulac-Arnold and colleagues (2019) summarized this class of problems as nine challenges. Common countermeasures include off-policy algorithms with experience replay, warm-starting from demonstration data, training a success detector to serve as the reward, and keeping a human in the loop to correct mistakes; UC Berkeley's SERL and HIL-SERL are notable examples.
ExampleSERL (ICRA 2024) learns PCB component insertion, cable routing, and object relocation on a real robot arm, training each policy in 25–50 minutes on average; its successor HIL-SERL adds human demonstration and correction, learning precision assembly, dynamic manipulation, and bimanual coordination in 1–2.5 hours with success rates near 100%.
- Also called
- Real-world RL
- Related
- Reinforcement Learning · Sample Efficiency · Human-in-the-Loop · Reset-Free Reinforcement Learning · HIL-SERL · Sim-to-Real Transfer
- Sources
- Challenges of Real-World Reinforcement Learning (Dulac-Arnold et al., 2019)
SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning
Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL) - As of
- 2025-03