Massively Parallel Reinforcement Learning
大规模并行强化学习CommonRunning thousands of simulated environments at once on a single GPU to collect data, training a locomotion policy in just minutes.
Massively parallel reinforcement learning means running both the physics simulation and the neural-network training on the GPU, with thousands of independent simulated environments running simultaneously, so a huge amount of interaction data can be collected in one pass to train a policy. Reinforcement learning needs enormous amounts of trial and error, and older CPU-based simulators with few parallel environments often took a dozen to a hundred-plus hours to train a legged-locomotion policy. NVIDIA's Isaac Gym (2021) kept simulation data on the GPU as PyTorch tensors the entire time, reportedly 2–3 orders of magnitude faster than CPU-based simulation; the same year, Rudin and colleagues ran 4,096 parallel ANYmal environments on a single GPU with PPO, training flat-ground walking in under 4 minutes and complex terrain in about 20. This paradigm made reinforcement learning the mainstream approach for legged and humanoid locomotion control, with Isaac Lab and legged_gym among the common tools.
ExampleRudin and colleagues' open-source legged_gym simulates 4,096 ANYmal robots at once on a single RTX A6000, uses a game-style terrain curriculum that adjusts difficulty automatically, trains a policy that walks complex terrain in about 20 minutes, and transfers it to a real ANYmal C.
- Also called
- Parallel Environment Training, GPU-Parallel Reinforcement Learning
- Related
- GPU-Accelerated Parallel Simulation · Vectorized Environments · Proximal Policy Optimization · NVIDIA Isaac Lab · legged_gym · Sim-to-Real Transfer
- Sources
- Learning to Walk in Minutes Using Massively Parallel Deep Reinforcement Learning (arXiv 2109.11978)
Isaac Gym: High Performance GPU-Based Physics Simulation For Robot Learning (arXiv 2108.10470)