Embodied AI Glossary中文

FlashAttention

Advanced

A GPU implementation of attention that reorders computation around memory hierarchy, giving identical results faster and with less memory.

FlashAttention was introduced by Tri Dao and others in 2022 as a GPU implementation of the attention mechanism that produces results numerically identical to standard attention — it's exact, not an approximation. A standard implementation has to write the full N×N attention matrix to GPU memory and read it back, and for long sequences that memory traffic becomes the bottleneck. FlashAttention instead splits Q, K, and V into small blocks, computes them inside the GPU's fast on-chip cache, and writes back only the final result, never materializing the full attention matrix — so memory use grows linearly with sequence length, and speed improves noticeably too. It was followed by FlashAttention-2 and, for the Hopper architecture, FlashAttention-3. Long-sequence models like VLAs and video world models rely on it heavily for both training and inference, and PyTorch's scaled_dot_product_attention has a similar backend built in.

ExampleSetting attn_implementation=“flash_attention_2” in Hugging Face Transformers when training a VLA allows a larger batch size or more image tokens within the same amount of GPU memory.

Also called
FlashAttention-2, FlashAttention-3
Related
Attention Mechanism · Transformer · Operator / Kernel · CUDA · Inference Latency · Mixed-Precision Training
Sources
FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Dao-AILab/flash-attention (GitHub)

See it in the full glossary →