Vision Transformer
视觉 TransformerViTCommonA network that cuts an image into small patches, treats them like a sequence of words, and processes them with a Transformer.
The Vision Transformer was introduced by a Google team in the 2020 paper 'An Image is Worth 16x16 Words.' It cuts an image into fixed-size patches (commonly 16×16 pixels), flattens and linearly maps each patch into a vector, adds a positional encoding (telling the model where each patch was), and feeds the whole set as a sequence of tokens into a standard Transformer encoder; self-attention lets every patch look directly at every other patch in the image. Before this, image tasks relied mainly on convolutional neural networks; ViT showed that, given enough pretraining data, a pure Transformer could match that performance on benchmarks like ImageNet while using less training compute. Today, vision foundation models such as CLIP, SigLIP, and DINOv2 are almost all built on ViT, and VLA vision encoders mostly use it too.
ExampleA 224×224 image cut into 16×16-pixel patches yields 14×14 = 196 patches — 196 visual tokens.
- Also called
- ViT
- Related
- Transformer · Visual Token · Self-Attention · Positional Encoding · Convolutional Neural Network · Vision Encoder
- Sources
- An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale (arXiv 2010.11929)