Embodied AI Glossary中文

Fleet Learning (Learning While Deploying)

机群学习 / 部署中学习Advanced

Having a whole fleet of already-deployed robots collect data while they work, continually improving one shared policy together.

Fleet learning means multiple simultaneously deployed robots pool their autonomous execution logs, failures, and human takeover data to update a shared policy, then push the new policy back out — turning deployment itself into part of training, in a data flywheel. UC Berkeley's Goldberg group proposed the “interactive fleet learning” setting with Fleet-DAgger (CoRL 2022): when a robot in the fleet is unsure, it asks a small pool of remote human operators for help, and their corrections feed imitation learning. “Learning while deploying” leans more toward reinforcement learning: in 2026, AgiBot (AGIBOT Finch) and the Shanghai Innovation Institute proposed the LWD framework, which uses autonomous rollouts and human-intervention data collected across a robot fleet for offline-to-online reinforcement learning, continually post-training a VLA.

ExampleLWD was validated on a fleet of 16 bimanual robots across 8 real manipulation tasks, including semantic shelf restocking and long-horizon tasks lasting 3–5 minutes; a single generalist policy reached 95% average success as it accumulated experience across the fleet.

Also called
Interactive Fleet Learning, IFL, LWD
Related
Data Flywheel · Deployment Data Backflow · Human-in-the-Loop · DAgger · Offline-to-Online Reinforcement Learning · Real-World Reinforcement Learning
Sources
Fleet-DAgger: Interactive Robot Fleet Learning with Scalable Human Supervision (arXiv:2206.14349)
Learning While Deploying: Fleet-Scale Reinforcement Learning for Generalist Robot Policies (arXiv:2605.00416)
As of
2026-09

See it in the full glossary →