Embodied AI Glossary中文

Video Prediction Policy

视频预测策略VPPAdvanced

A robot policy that guides action generation using the internal 'prediction of the future' features inside a video diffusion model.

Video Prediction Policy was released in December 2024 by Jianyu Chen's lab at Tsinghua University together with RobotEra, UC Berkeley, the Shanghai Artificial Intelligence Laboratory, and others, selected as an ICML 2025 Spotlight. A typical visual encoder only looks at a single image, or compares two images, and cannot capture 'what's about to happen next.' VPP's hypothesis is that a video diffusion model's internal features, when predicting an upcoming frame, already encode both the current image and a prediction of future dynamics. The approach first fine-tunes a pretrained video model on robot data and internet videos of human manipulation, then takes its intermediate predictive representation as a conditioning signal to train an implicit inverse dynamics model (which infers the action from the 'predicted future') that outputs the action. The paper reports an 18.6% relative improvement over the previous best method on the CALVIN ABC-D generalization benchmark, and a 31.6% success-rate improvement on complex real-robot dexterous-hand manipulation tasks.

ExampleOn a Franka arm and an XHand dexterous hand, VPP forms a predictive representation of the next few frames internally from a language instruction and outputs the action from that, without ever having to fully render out the whole future video.

Also called
Video Prediction Policy: A Generalist Robot Policy with Predictive Visual Representations
Related
Video Prediction Model · Inverse Dynamics Model · Diffusion Policy · CALVIN Benchmark · Vidar · RobotEra
Sources
Video Prediction Policy (arXiv 2412.14803)
VPP project page
As of
2025-05

See it in the full glossary →