Embodied AI Glossary中文

Vision-Language Model

视觉语言模型VLMEssential

A model that takes in both images and text and answers in text; the base that most VLAs build on.

A vision-language model takes in an image and text and outputs text: it can describe a picture, answer questions about it, or point out where an object is. Three parts are common: a vision encoder turns the image into features, a projection layer aligns those features to the language model's input space, and the language model handles understanding and generating the text. 2023's LLaVA, for instance, is a CLIP encoder plus a projection layer plus the Vicuna language model; Google's 2024 open-source PaliGemma combines a SigLIP encoder with Gemma-2B, at about 3 billion parameters. The objects, common sense, and spatial knowledge a VLM picks up from internet image-text data are exactly what robots lack, which is why most VLAs start from a VLM: RT-2 adds robot data on top of PaLI-X and PaLM-E through joint fine-tuning, and π0 is built on PaliGemma. VLMs are also commonly used on their own as the high-level planner in a hierarchical architecture.

ExampleAsk a VLM about a kitchen photo, “what's to the left of the sink,” and it answers in text; swap the output for action tokens and train it further on robot data, and you get a VLA like RT-2.

Also called
VLM
Related
Vision-Language-Action Model · Vision Encoder · Projector / Connector · Large Language Model · PaliGemma · Multimodal Large Language Model
Sources
Vision Language Models Explained (Hugging Face blog)
PaliGemma: A versatile 3B VLM for transfer (arXiv:2407.07726)
RT-2: Vision-Language-Action Models (project page)
As of
2024-07

See it in the full glossary →