Embodied AI Glossary中文

Vision Encoder

视觉编码器Essential

A network that turns an image into a set of feature vectors, or visual tokens, for downstream models to use.

A vision encoder turns pixels into features a model can use: given an image, it outputs a set of vectors, each summarizing the content of a small region. Early designs mostly used convolutional networks such as ResNet; the mainstream today is the Vision Transformer (ViT), which cuts an image into patches — 16×16 pixels in the original ViT — turns each patch into a vector, and uses attention to let the patches exchange information. An encoder's ability mostly comes from pretraining: OpenAI's 2021 CLIP trained on 400 million web image-text pairs to align images and text in the same space; SigLIP is Google's improved version built on that idea; and Meta's DINOv2 trains purely on images with self-supervision, preserving more spatial and geometric detail. VLAs usually use an off-the-shelf vision encoder as is, connecting its features to the language model through a projection layer, and may freeze it or fine-tune it jointly during training.

ExampleOpenVLA feeds a 224×224 image into both SigLIP and DINOv2 at once, concatenates the two sets of features channel-wise, and passes them through a two-layer MLP projection layer to turn them into visual tokens the language model can read.

Also called
Image Encoder, Visual Backbone Network
Related
Vision Transformer · CLIP · SigLIP · DINOv2 · Projector / Connector · Visual Token
Sources
An Image is Worth 16x16 Words (ViT, arXiv:2010.11929)
Learning Transferable Visual Models From Natural Language Supervision (CLIP, arXiv:2103.00020)
OpenVLA: An Open-Source Vision-Language-Action Model (arXiv:2406.09246)
As of
2024-06

See it in the full glossary →