FlashAttention
AdvancedA GPU implementation of attention that reorders computation around memory hierarchy, giving identical results faster and with less memory.
FlashAttention was introduced by Tri Dao and others in 2022 as a GPU implementation of the attention mechanism that produces results numerically identical to standard attention — it's exact, not an approximation. A standard implementation has to write the full N×N attention matrix to GPU memory and read it back, and for long sequences that memory traffic becomes the bottleneck. FlashAttention instead splits Q, K, and V into small blocks, computes them inside the GPU's fast on-chip cache, and writes back only the final result, never materializing the full attention matrix — so memory use grows linearly with sequence length, and speed improves noticeably too. It was followed by FlashAttention-2 and, for the Hopper architecture, FlashAttention-3. Long-sequence models like VLAs and video world models rely on it heavily for both training and inference, and PyTorch's scaled_dot_product_attention has a similar backend built in.
ExampleSetting attn_implementation=“flash_attention_2” in Hugging Face Transformers when training a VLA allows a larger batch size or more image tokens within the same amount of GPU memory.
- Also called
- FlashAttention-2, FlashAttention-3
- Related
- Attention Mechanism · Transformer · Operator / Kernel · CUDA · Inference Latency · Mixed-Precision Training
- Sources
- FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Dao-AILab/flash-attention (GitHub)