Embodied AI Glossary中文

On-Policy Distillation

在线策略蒸馏OPDAdvanced

Distillation where the student generates its own trajectories, and the teacher scores and corrects it at every step.

Ordinary knowledge distillation has the student imitate data the teacher already generated, but once deployed the student runs into states that were never in the teacher's data, and errors accumulate — the same problem as compounding error in behavior cloning. On-policy distillation instead has the student generate its own samples, and the teacher gives a target distribution at every step the student actually reaches, whether that is every token or every action, usually measured with reverse KL divergence, an idea in the same spirit as DAgger. Google DeepMind's Agarwal and colleagues systematically proposed this under the name GKD in 2023; a Thinking Machines blog post in October 2025 made it widely known in large-model post-training, where its feedback is far denser than reinforcement learning's sparse reward. A 2026 paper, VLA-OPD, applies the idea to VLA post-training, replacing sparse environment reward with dense, step-by-step supervision from an expert teacher.

ExampleAccording to the Qwen3 technical report cited in the Thinking Machines blog post, on-policy distillation scored 74.4 on AIME'24 using about 1,800 GPU-hours, while plain reinforcement learning scored only 67.6 using about 17,920 GPU-hours.

Also called
OPD, Generalized Knowledge Distillation (GKD)
Related
Knowledge Distillation · Policy Distillation · DAgger · Teacher-Student Distillation · Reinforcement Fine-Tuning (RL Fine-Tuning) · VLA-OPD
Sources
Thinking Machines Lab (Kevin Lu), 2025-10-27: On-Policy Distillation
Agarwal et al. 2023: On-Policy Distillation of Language Models: Learning from Self-Generated Mistakes (GKD)
VLA-OPD: Bridging Offline SFT and Online RL for VLA Models via On-Policy Distillation (arXiv 2603.26666)
As of
2026-03

See it in the full glossary →