Flamingo
AdvancedA 2022 DeepMind vision-language model that learns a new task from just a few image-text examples, an early landmark VLM.
Flamingo is a vision-language model (VLM, a model that can look at images and read and write text at the same time) released by DeepMind in April 2022, published at NeurIPS 2022, with its largest version at 80 billion parameters. It freezes an already-trained vision encoder and language model (DeepMind's 70-billion-parameter Chinchilla) and adds two new kinds of module in between: a Perceiver Resampler that compresses any number of image features down to a fixed number of visual tokens, and gated cross-attention layers inserted into the language model so it can “look at” the image while generating text. It is trained on web data where images and text are interleaved, which lets it do few-shot learning the way large language models do: put a few image-text examples in the prompt, and it completes a new task with no parameter updates. Across 16 benchmarks, it beat prior few-shot methods using just 4 examples per task. The community's open-source reproduction, OpenFlamingo, was later used by RoboFlamingo as the backbone for a robot policy.
ExamplePut two example animal photos with captions in the prompt, followed by a new photo, and Flamingo will write a caption for it in the same format — with no additional training required.
- Also called
- DeepMind Flamingo, a Visual Language Model for Few-Shot Learning
- Related
- Vision-Language Model · Perceiver Resampler · Cross-Attention · Few-shot · In-Context Learning · RoboFlamingo
- Sources
- Flamingo: a Visual Language Model for Few-Shot Learning (arXiv 2204.14198)
Tackling multiple tasks with a single visual language model (Google DeepMind 博客) (Chinese) - As of
- 2022-04