Embodied AI Glossary中文

Generative Value Learning (GVL)

GVL(生成式价值学习)GVLAdvanced

Has a vision-language model estimate task progress from shuffled video frames, turning it into a general-purpose value function.

GVL was proposed in November 2024 by researchers at Google DeepMind, together with the University of Pennsylvania and Stanford, published at ICLR 2025. “Value” here means what percentage of the task is complete, which can be used to judge success or failure, filter data, or serve as a reward for reinforcement learning. Directly having a vision-language model (VLM) score each frame in chronological order works poorly, because adjacent frames are highly correlated and the model tends to just output a monotonically increasing score based on order alone. GVL instead shuffles the video frames first, then has the model (the paper uses Gemini-1.5-Pro) estimate completion percentage frame by frame, forcing it to actually look at what's in each frame. It needs no task-specific training, works zero-shot or few-shot across more than 300 real tasks, and can even learn by putting human videos or other robots' demonstrations into its context.

ExampleUse GVL to score a batch of demonstration videos for progress; clips where the predicted progress correlates poorly with the true chronological order (usually failed or low-quality demonstrations) can be filtered out, and the remaining data used to train an imitation-learning policy.

Also called
Vision Language Models are In-Context Value Learners
Related
Progress Reward Model · VLM-as-Reward · Value Function · Success Detector · In-Context Learning · Data Curation
Sources
Vision Language Models are In-Context Value Learners (arXiv:2411.04549)
Generative Value Learning 项目主页 (Chinese)
As of
2025-01

See it in the full glossary →