Teacher-Student Distillation
教师-学生蒸馏CommonTraining a teacher policy that can see privileged information first, then having a student that only uses real sensors imitate it.
Teacher-student distillation grows out of knowledge distillation (introduced by Hinton and colleagues in 2015, which trains a small model on a large model's output as soft labels). In robotics it's usually done in two steps: first train a teacher policy in simulation, where it can read privileged information a real robot can't access directly, such as exact terrain height, friction coefficients, and object poses, which makes it much easier to train well with reinforcement learning; then train a student policy that only receives observations a real robot actually has — joint states, IMU readings, camera images — learning by imitating the teacher's actions or intermediate representations, often with DAgger-style online correction. This separates “hard-to-learn decisions” from “hard-to-learn perception,” and it's one of the mainstream routes to sim-to-real transfer. Notable examples include Learning by Cheating (2019) in autonomous driving and ETH's 2020 blind quadruped locomotion controller.
ExampleETH's Lee and colleagues (Science Robotics 2020) let the teacher read privileged terrain information in simulation while the student imitates it using only proprioceptive signals; the resulting ANYmal quadruped deploys zero-shot to mud, snow, rubble, dense vegetation, and swift-flowing water it never saw during training.
- Also called
- Teacher-Student Training, Privileged Teacher Distillation
- Related
- Knowledge Distillation · Privileged Information · Policy Distillation · Asymmetric Actor-Critic · Sim-to-Real Transfer · DAgger
- Sources
- Learning Quadrupedal Locomotion over Challenging Terrain (Lee et al., Science Robotics 2020, arXiv 2010.11251)
Learning by Cheating (Chen et al., CoRL 2019, arXiv 1912.12294)
Distilling the Knowledge in a Neural Network (Hinton et al., arXiv 1503.02531)