Embodied AI Glossary中文

Vision Transformer

视觉 TransformerViTCommon

A network that cuts an image into small patches, treats them like a sequence of words, and processes them with a Transformer.

The Vision Transformer was introduced by a Google team in the 2020 paper 'An Image is Worth 16x16 Words.' It cuts an image into fixed-size patches (commonly 16×16 pixels), flattens and linearly maps each patch into a vector, adds a positional encoding (telling the model where each patch was), and feeds the whole set as a sequence of tokens into a standard Transformer encoder; self-attention lets every patch look directly at every other patch in the image. Before this, image tasks relied mainly on convolutional neural networks; ViT showed that, given enough pretraining data, a pure Transformer could match that performance on benchmarks like ImageNet while using less training compute. Today, vision foundation models such as CLIP, SigLIP, and DINOv2 are almost all built on ViT, and VLA vision encoders mostly use it too.

ExampleA 224×224 image cut into 16×16-pixel patches yields 14×14 = 196 patches — 196 visual tokens.

Also called
ViT
Related
Transformer · Visual Token · Self-Attention · Positional Encoding · Convolutional Neural Network · Vision Encoder
Sources
An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (arXiv 2010.11929)

See it in the full glossary →