Embodied AI Glossary中文

VLA-OPD

Advanced

A VLA post-training method where a strong teacher corrects the student token by token on trajectories the student generated itself.

VLA-OPD is a VLA post-training framework proposed by a team at HKUST (Guangzhou) in March 2026. Offline supervised fine-tuning (SFT) learns only from demonstration data and has no way to handle states the model drifts into on its own, and it is also prone to catastrophic forgetting of pretrained abilities; online reinforcement learning, meanwhile, is held back by sparse rewards and poor sample efficiency. VLA-OPD takes a middle path — on-policy distillation: the student policy runs in the environment on its own, and a stronger teacher policy provides dense, token-by-token supervision on exactly the states the student itself reaches, with no dependence on environment reward. The loss uses reverse KL divergence, which the authors argue biases the model toward 'committing to one mode,' making it more stable than forward KL (prone to exploding entropy) or hard cross-entropy (prone to premature entropy collapse). In experiments, the teacher is an expert trained by SimpleVLA-RL and the student is OpenVLA-OFT; on LIBERO and RoboTwin 2.0, VLA-OPD is more sample-efficient than reinforcement learning, more robust than SFT, and forgets less.

ExampleOn LIBERO, a student fine-tuned on just one demonstration per task starts at 48.9% average success; distilling it with VLA-OPD raises this to 87.4%, and following up with GRPO reinforcement learning brings it to 93.4%, close to the teacher's 93.9%.

Also called
On-Policy VLA Distillation, VLA-OPD: Bridging Offline SFT and Online RL for Vision-Language-Action Models via On-Policy Distillation
Related
On-Policy Distillation · Supervised Fine-Tuning · Reinforcement Fine-Tuning (RL Fine-Tuning) · Catastrophic Forgetting · Kullback-Leibler Divergence · SimpleVLA-RL
Sources
VLA-OPD (arXiv 2603.26666)
VLA-OPD 项目主页 (Chinese)
As of
2026-03

See it in the full glossary →