Embodied AI Glossary中文

Perceiver Resampler

Perceiver 重采样器Advanced

A module that uses a small set of learnable query vectors to compress a variable number of visual features into a fixed number of tokens.

The Perceiver Resampler is a connector module DeepMind introduced in its 2022 Flamingo vision-language model, based on ideas from the 2021 Perceiver model. The number of features a vision encoder outputs varies with image resolution and video frame count, and is often large, making it expensive to feed directly into a language model. The resampler instead defines a small, fixed set of learnable query vectors (64 of them in Flamingo) that 'read' all the visual features through cross-attention, producing a fixed number of visual tokens as output. Flamingo's ablations show it outperforms an ordinary Transformer or MLP as the connector. It's conceptually close to BLIP-2's Q-Former, both compressing visual information with query vectors. In robot models, ByteDance's GR-1 uses it to compress image tokens too.

ExampleByteDance's GR-1 first encodes each frame into a large number of patch tokens with an MAE-pretrained ViT, compresses them down with a Perceiver Resampler, then feeds the result together with language and robot state into a GPT-style Transformer to predict actions and future frames.

Also called
Perceiver
Related
Projector / Connector · Querying Transformer · Learnable Query · Cross-Attention · Visual Token · GR-1 (ByteDance)
Sources
Flamingo: a Visual Language Model for Few-Shot Learning (arXiv:2204.14198)
Perceiver: General Perception with Iterative Attention (arXiv:2103.03206)
Unleashing Large-Scale Video Generative Pre-training for Visual Robot Manipulation (GR-1, arXiv:2312.13139)

See it in the full glossary →