Embodied AI Glossary中文

Self-Attention

自注意力Common

A mechanism where every element in a sequence gathers information from every other element, weighted by relevance.

Self-attention is the core operation inside a Transformer, popularized by Vaswani and colleagues' 2017 paper 'Attention Is All You Need.' For each token in a sequence, learnable matrices first produce a query (Q), key (K), and value (V) vector; that token's query is dotted with every token's key, the results are turned into weights with a softmax, and those weights are used to compute a weighted sum of the value vectors, giving the token its new representation. 'Self' means the queries, keys, and values all come from the same sequence; when they come from two different sequences, it's called cross-attention instead. Self-attention can connect any two positions in a single step and is easy to parallelize, but its compute cost grows with the square of sequence length. In VLA models, image, text, and state tokens are often mixed into the same self-attention operation, with an attention mask specifying who is allowed to see whom.

Exampleπ0 concatenates image, language instruction, robot state, and noisy action tokens into one sequence, lets them exchange information via self-attention, and reads the action back out of the action tokens' output.

Also called
Intra-Attention, Multi-Head Self-Attention
Related
Attention Mechanism · Transformer · Cross-Attention · Attention Mask · Softmax · Positional Encoding
Sources
Attention Is All You Need (arXiv 1706.03762)
Dive into Deep Learning: Self-Attention and Positional Encoding

See it in the full glossary →