Embodied AI Glossary中文

Population-Based Training

基于群体的训练PBTAdvanced

Training a whole population of models at once, periodically copying good weights over bad ones and perturbing hyperparameters.

Population-based training is a hyperparameter-optimization method proposed by DeepMind's Jaderberg and colleagues in 2017. It trains a batch of models in parallel and periodically compares their performance: worse-performing members directly copy the weights of better-performing ones (exploit), then have hyperparameters such as the learning rate randomly perturbed before continuing to train (explore). This way, there is no need to finish tuning hyperparameters before real training starts — a single run automatically discovers a hyperparameter schedule that changes over time, using roughly the same total compute as running the same number of ordinary parallel experiments. The original paper validated the approach on deep reinforcement learning, machine translation, and GANs. Reinforcement learning is especially sensitive to hyperparameters, so PBT is often paired with massively parallel GPU simulation.

ExampleNVIDIA's DexPBT (RSS 2023) uses decentralized PBT in Isaac Gym to train single-arm and dual-arm robots fitted with multi-fingered dexterous hands on tasks such as regrasping, throwing after grasping, and object reorientation, exploring noticeably better than standard end-to-end training.

Also called
PBT
Related
Hyperparameter · Massively Parallel Reinforcement Learning · Exploration vs. Exploitation · GPU-Accelerated Parallel Simulation · Reinforcement Learning · Isaac Gym
Sources
Jaderberg et al. 2017: Population Based Training of Neural Networks
Google DeepMind Blog: Population based training of neural networks
Petrenko et al. 2023: DexPBT

See it in the full glossary →