Embodied AI Glossary中文

Web-scale Vision-Language Data

互联网图文数据Common

The vast supply of paired images and text, plus image-based Q&A, collected from the web, teaching a model about the world.

This refers to large-scale image-text pairs collected from web pages, and the image captioning, visual question answering, and object detection data derived from them — for instance LAION-5B, which used CLIP to filter Common Crawl web pages down to 5.85 billion image-text pairs. Vision-language models (VLMs) pretrain on data like this, which is why they recognize a huge range of objects and carry a fair amount of common sense. Robot data is far smaller in scale, covering a limited set of objects and scenes, so VLA models mainly draw on internet-scale knowledge two ways: using a VLM backbone that was itself pretrained on web-scale data, and co-training on a mix of web data and robot data during fine-tuning, so the fine-tuning process doesn't wash out what the model already knew. RT-2 was the first to systematically validate this approach, and π0.5 also lists web data as one of its co-training sources.

ExampleRT-2 fine-tunes on a mix of robot trajectories and web data such as visual question answering, so the model can apply objects and concepts it learned from the web to actions, executing instructions that never appeared in the robot data.

Also called
Internet-scale Vision-Language Data, Web Data
Related
Vision-Language Model · Co-training · Vision-Language-Action Model · RT-2 · Catastrophic Forgetting · Knowledge Insulation
Sources
RT-2: Vision-Language-Action Models Transfer Web Knowledge to Robotic Control
LAION-5B: An open large-scale dataset for training next generation image-text models
π0.5: a Vision-Language-Action Model with Open-World Generalization
As of
2025-04

See it in the full glossary →