Embodied AI Glossary中文

Cross-Attention

交叉注意力Common

Attention where one set of tokens queries information from a different set; a common way to fuse modalities.

Cross-attention is a variant of the attention mechanism. In self-attention, the query, key, and value all come from the same sequence; in cross-attention, sequence A supplies the query while sequence B supplies the key and value, so every token in A can pull in relevant information from B, while B itself is left unchanged. It first appeared in the encoder-decoder structure of the 2017 Transformer paper, where the decoder uses it to read the encoder's representation of the source sentence for translation. It has since become the standard way to inject conditioning information into a model: text-to-image models use it to read text encodings, and robot policies use it to feed visual and language features into the action-generation part. An alternative is simply concatenating two sets of tokens and running self-attention over the combination; both approaches are common in VLAs, each with its own tradeoffs.

ExampleNVIDIA's GR00T N1 action module, a DiT variant, stacks two kinds of blocks alternately: self-attention blocks process the noisy action tokens and robot state, while cross-attention blocks read the visual-language tokens output by the VLM.

Also called
Encoder-Decoder Attention
Related
Attention Mechanism · Self-Attention · Transformer · Multimodal Fusion · Encoder-Decoder · Diffusion Transformer
Sources
Attention Is All You Need (arXiv 1706.03762)
GR00T N1: An Open Foundation Model for Generalist Humanoid Robots (arXiv 2503.14734)

See it in the full glossary →