Embodied AI Glossary中文

TokenLearner

Advanced

A module that adaptively pools a large number of image tokens down into a handful of key tokens, to save compute.

TokenLearner is a module Google's Ryoo and colleagues proposed in 2021 (NeurIPS 2021), with a paper title that literally asks what 8 learned tokens can do. A Vision Transformer usually cuts an image into tens to hundreds of patches, one token each, and attention's compute grows with the square of the token count. TokenLearner instead computes a spatial attention map for each output token based on the input content, uses it to weight and pool the feature map, and compresses a large number of tokens down to around 8; later layers only process these few tokens. The paper reaches competitive results on benchmarks like ImageNet and the Kinetics video-recognition set while using noticeably less compute. Its best-known use in embodied AI is in Google's RT-1, where it helps a large model meet the speed requirements of real-time control.

ExampleRT-1 uses TokenLearner to compress each image's 81 visual tokens down to 8, so 6 frames of history become just 48 tokens fed into the Transformer; the paper reports this step gives about a 2.4x inference speedup.

Related
Visual Token · Visual Token Pruning · RT-1 · Perceiver Resampler · Querying Transformer · Vision Transformer
Sources
TokenLearner: What Can 8 Learned Tokens Do for Images and Videos? (arXiv:2106.11297)
RT-1: Robotics Transformer for Real-World Control at Scale (arXiv:2212.06817)

See it in the full glossary →