Interactive Imitation Learning
交互式模仿学习IILAdvancedImitation learning where a human gives feedback while the robot acts, and the policy improves from it online.
Interactive imitation learning is a branch of imitation learning in which a human gives intermittent feedback while the robot is executing — taking over to correct it, labeling the right action, or rating how well it did — and the policy improves online from that feedback. It mainly targets the compounding-error problem in behavior cloning: offline demonstrations only cover the states an expert visited, so once the robot drifts off that path there is no data to learn from, and small errors snowball. The flagship algorithm is DAgger, proposed by Stéphane Ross and colleagues in 2011: the expert labels the correct action for states the student policy actually runs into, those get folded into the dataset, and the policy is retrained; variants such as HG-DAgger let the human decide when to take over. A 2022 survey by Celemin and colleagues organizes this whole line of work.
ExampleA typical pipeline: train an initial peg-insertion policy on demonstrations, let the arm run on its own, have an operator take over and record corrective actions whenever it drifts, fold those corrections into the training set, and retrain — repeating this for a few rounds.
- Also called
- IIL
- Related
- DAgger · Human-Gated DAgger · Compounding Error · Behavior Cloning · Human-in-the-Loop · Human Intervention Data
- Sources
- Interactive Imitation Learning in Robotics: A Survey (arXiv:2211.00600)
A Reduction of Imitation Learning and Structured Prediction to No-Regret Online Learning (DAgger, arXiv:1011.0686)