Attention lets a neural network decide, for each output it produces, which parts of its input to draw on and how much. Instead of squeezing the whole input into one fixed summary, the model keeps every piece available and computes a fresh weighted mix at each step.
It was introduced for machine translation. In 2014 Bahdanau, Cho and Bengio argued that compressing a source sentence into a single fixed-length vector was the bottleneck of encoder–decoder translation. Their fix let the decoder search over all the source words each time it produced a word (read the paper). The weights it learned lined up well with which words translate which.
How it works
The version used today comes from the 2017 transformer paper. Each token is projected into three vectors, a query, key and value. The query is what this token is looking for, the key is what it offers to be matched on, and the value is the information it passes along.
- Compare the query with every key using a dot product, giving one score per token.
- Divide by the square root of the key length. Without this, scores grow large in bigger models and push softmax into a region where gradients all but vanish.
- Apply softmax, so the scores become weights that sum to 1.
- Output the weighted sum of the values.
The diagram runs exactly this on “the cat sat on the mat” with tiny three-number embeddings, each word acting as its own query and key. Words with similar vectors, like “cat” and “mat”, weight each other heavily, while “the”, whose vector is small, spreads its attention almost evenly. In a real model the query, key and value projections are learned, so the model itself decides what counts as relevant.
Main kinds
- Self-attention: queries, keys and values all come from the same sequence, so each token gathers context from the others. It is the core of the transformer.
- Cross-attention: queries come from one sequence and keys and values from another, like a translation decoder reading the source sentence, or an image generator reading a text prompt.
- Multi-head attention: several attention operations in parallel, each with its own projections, so different heads can track different relationships. The original transformer used eight.
The catch
Every token is scored against every other, so the work grows with the square of the sequence length. That’s the main reason long context is expensive, and why memory-saving implementations of attention matter so much for large models.