Embodied AI Glossary中文

Actor-Learner Architecture (Distributed RL)

Actor-Learner 分离架构Advanced

Splitting “interacting with the environment to collect data” and “updating the network's parameters” across separate processes or machines, run in parallel.

This is a common system design in distributed reinforcement learning: several actors each hold a copy of the policy and interact with their own environment to generate experience; a learner centrally collects that experience and updates the parameters on a GPU, periodically syncing the new parameters back to the actors. DeepMind's Ape-X and IMPALA (2018) are notable examples: Ape-X has actors write experience into a shared replay buffer; IMPALA has actors send trajectories directly to the learner and uses V-trace, an importance-weighted correction, to handle the bias from the actors' policy lagging slightly behind the learner's. The benefit is that sampling and training never block each other, and the system can scale to thousands of machines. Real-robot reinforcement learning needs this same split: SERL runs the actor and learner on separate threads, so training being slow never drags down the robot's control frequency.

ExampleSERL runs three processes in parallel on a real robot: the actor picks actions, the learner trains the network, and the robot's environment executes the actions — keeping a fixed control frequency and shortening total real-world training time.

Also called
Actor-Learner Architecture, Decoupled Sampling and Training
Related
Experience Replay · Off-Policy · Importance Sampling · Real-World Reinforcement Learning · SERL · Massively Parallel Reinforcement Learning
Sources
IMPALA: Scalable Distributed Deep-RL with Importance Weighted Actor-Learner Architectures (arXiv 1802.01561)
Distributed Prioritized Experience Replay (Ape-X, arXiv 1803.00933)
SERL: A Software Suite for Sample-Efficient Robotic Reinforcement Learning (arXiv 2401.16013)

See it in the full glossary →