Embodied AI Glossary中文

Attention Mechanism

注意力机制Essential

A way of weighting parts of the input by relevance and combining them, letting a model focus on what matters.

Attention in deep learning is usually traced to Bahdanau, Cho, and Bengio's 2014 neural machine translation paper: when translating each word, the model automatically looks for the most relevant words in the source sentence, instead of compressing the whole sentence into one fixed-length vector. The now-common form has every position produce a query (Q), a key (K), and a value (V); the similarity between Q and each K is normalized with softmax into weights, and those weights are used to sum up V. The 2017 Transformer, by Vaswani and colleagues, replaced recurrence and convolution entirely with attention, which made training parallelizable and far easier to scale, and it became the shared foundation of large language models and VLAs. In embodied models, self-attention lets image patches, text, and action tokens exchange information with each other, and cross-attention is commonly used to let the action module read visual-language features.

ExampleGiven the instruction “put the red cup in the sink,” when the model processes the text tokens for “red cup,” attention weights concentrate on the image patches showing the red cup, tying the language to the picture.

Also called
Attention
Related
Transformer · Self-Attention · Cross-Attention · Causal Attention · Attention Mask · Key-Value Cache
Sources
Neural Machine Translation by Jointly Learning to Align and Translate (arXiv 1409.0473)
Attention Is All You Need (arXiv 1706.03762)
Attention (machine learning) - Wikipedia

See it in the full glossary →