DAgger
DAgger(数据集聚合)CommonLetting the policy run itself, having an expert label the correct action at the states it actually reaches, then retraining on the combined data.
DAgger (Dataset Aggregation) is an interactive imitation-learning algorithm introduced by Ross, Gordon, and Bagnell at AISTATS 2011. It targets behavior cloning's compounding-error problem: a policy trained only on an expert's trajectory has no idea what to do once it drifts off that path. The procedure iterates: train an initial policy on expert data; let the current policy run and record the states it actually visits; ask the expert to label what should happen at those states; merge the new labels into the growing dataset and retrain; repeat. The paper proves this reduces the growth of error with task length from quadratic to close to linear. The downside is that the expert must be available to label on demand; on real robots, variants like HG-DAgger are common, where a human supervises and takes over when needed.
ExampleThe original paper learned steering in the racing game Super Tux Kart and level completion in Super Mario Bros., and DAgger outperformed plain supervised learning on expert demonstrations alone in both.
- Also called
- Dataset Aggregation
- Related
- Compounding Error · Distribution Shift · Behavior Cloning · Human-Gated DAgger · Interactive Imitation Learning · Human-in-the-Loop
- Sources
- A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (PMLR v15)
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (PDF, 含 DAgger 算法与 T²ε 分析) (Chinese)