Causal Attention
因果注意力CommonAttention where each position can only see itself and earlier positions, never anything that comes after.
Causal attention adds a mask to self-attention that sets the attention score for any “future” position to negative infinity, so each token can only attend to itself and the tokens before it. It comes from the decoder in the original 2017 Transformer paper, and it is what guarantees autoregressive behavior: the whole sequence is fed in at once during training, but predicting position i still only ever depends on what comes before it. Decoder-only language models such as GPT and Llama all use it, and so do VLAs built on them when generating action tokens one at a time. Its counterpart is bidirectional attention, where every position can see every other position, used by BERT and ViT. The two are also often mixed: π0 uses a block-causal mask, splitting the input into three blocks — image/language, robot state, and noisy action — where tokens within a block can all see each other, but each block can only see itself and the blocks before it.
ExampleIn the sequence “pick up cup,” when the model computes the representation for “cup,” it can only attend to “pick,” “up,” and “cup” itself; anything after “cup” is blocked by the mask.
- Also called
- Causal Mask, Masked Self-Attention
- Related
- Self-Attention · Attention Mask · Decoder-only Architecture · Autoregressive Decoding · Key-Value Cache · Transformer
- Sources
- Attention Is All You Need (arXiv:1706.03762)
π0: A Vision-Language-Action Flow Model for General Robot Control (arXiv:2410.24164)