Embodied AI Glossary中文

Learnable Query

可学习查询Advanced

A set of vectors trained as model parameters that use attention to gather task-relevant information out of the input.

A learnable query is a set of vectors that are randomly initialized and updated during training, rather than coming from the input; they act as 'queries' in an attention computation that reads the input features, pooling the needed information into a fixed number of outputs. Facebook's 2020 object-detection model DETR uses a set of object queries, each producing one detection; 2023's Q-Former in BLIP-2 uses 32 learnable queries to extract visual features from a frozen image encoder before passing them to a large language model. In robot models, Octo inserts readout tokens into its sequence: they can see the preceding observation and task tokens but aren't seen by those tokens in turn, and the action head generates the action from their output; 2025's VLA-Adapter adds ActionQuery tokens inside a vision-language model (64 of them worked best in its experiments), specifically to gather action-relevant multimodal information for the policy network. Its role is to compress an input of varying length into a fixed-size, task-oriented representation.

ExampleIn BLIP-2, an image is first turned into hundreds of feature tokens by the vision encoder, and Q-Former's 32 query vectors use cross-attention to pool them into 32 outputs, which are projected and prepended to the text before being fed into the language model.

Also called
Action Query, Readout Token, Object Query
Related
Querying Transformer · Cross-Attention · Perceiver Resampler · Action Head · Octo · VLA-Adapter
Sources
BLIP-2: Bootstrapping Language-Image Pre-training with Frozen Image Encoders and LLMs (arXiv:2301.12597)
Octo: An Open-Source Generalist Robot Policy (arXiv:2405.12213)
VLA-Adapter: An Effective Paradigm for Tiny-Scale VLA Model (arXiv:2509.09372)
As of
2025-09

See it in the full glossary →