Sample Efficiency
样本效率CommonHow much data or environment interaction is needed to reach a given performance level; using less means being more efficient.
Sample efficiency measures how many samples a learning algorithm needs to reach a certain level of performance: in reinforcement learning, this means the number of steps of interaction with the environment; in imitation learning, the number of demonstrations. It matters enormously in robotics. In simulation it can be offset with GPU parallelism — Rudin and colleagues (2021), for instance, ran thousands of ANYmal quadrupeds at once on a single GPU, training flat-ground walking in under 4 minutes. A real robot can only act step by step, with wear and the cost of manually resetting the scene on top, so real-robot reinforcement learning must use highly sample-efficient methods. Common ways to improve it include off-policy algorithms with experience replay that reuse old data repeatedly, adding human demonstrations or corrections, model-based reinforcement learning that trains partly “in imagination,” and using pretrained visual representations. On-policy PPO is generally less sample-efficient and suits large-scale parallel simulation, while off-policy methods like SAC use samples more sparingly and are common on real robots.
ExampleUC Berkeley's SERL uses off-policy reinforcement learning on a real robot, learning tasks like PCB insertion and cable routing in 25–50 minutes of training on average per task; its successor HIL-SERL adds human correction and reaches near-perfect success in 1–2.5 hours.
- Also called
- Data Efficiency
- Related
- Off-Policy · Experience Replay · Real-World Reinforcement Learning · Model-Based Reinforcement Learning · Massively Parallel Reinforcement Learning · SERL
- Sources
- SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning (arXiv 2401.16013)
Precise and Dexterous Robotic Manipulation via Human-in-the-Loop Reinforcement Learning (HIL-SERL, arXiv 2410.21845)
Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv 2109.11978)