FlashAttention: Fast and Memory-Efficient Exact Attention with IO-Awareness
Tri Dao et al. · NeurIPS 2022
arXiv:2205.14135
In short
Attention is slow mostly because of memory traffic, not arithmetic. FlashAttention computes exact attention in tiles that stay in the GPU’s fast on-chip memory, cutting reads and writes to main memory and speeding up training and inference.
Why it matters
It is a big part of why long context windows became affordable, and it ships inside nearly every LLM stack.
Read first
The 3 Field Guide ideas this paper leans on.
Starting from scratch? The full route 11 ideas · basics first
- Machine Learning ✓ understood
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
- Neural Network ✓ understood
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Feature ✓ understood
A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.
- Deep Learning ✓ understood
A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.
- Representation Learning ✓ understood
Learning useful features or representations of data automatically, rather than hand-crafting them.
- Embedding ✓ understood
A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.
- Attention Mechanism ✓ understood
A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.
- Self-Attention · read first ✓ understood
A mechanism where each token attends to all other tokens in the sequence to understand contextual relationships.
- GPU · read first ✓ understood
Graphics Processing Unit - hardware accelerator with thousands of cores, essential for parallel computation in deep learning.
- Flash Attention · read first ✓ understood
An efficient attention algorithm reducing memory usage and increasing speed through clever recomputation.