Landmark 7 stops to get here · leads to 7

Attention Mechanism

A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.

Your route here

7 stops · basics first
  1. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  2. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  3. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  4. Feature ✓ understood

    A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.

  5. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  6. Representation Learning ✓ understood

    Learning useful features or representations of data automatically, rather than hand-crafting them.

  7. Embedding ✓ understood

    A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.

  8. Attention Mechanism · you are here ✓ understood

Picture it

EACH ROW ATTENDS TO THE COLUMNS · ROWS SUM TO 1thethe0.18catcat0.39satsat0.42onon0.26thethe0.18matmat0.38computed from toyword vectors:softmax(q·k / √d)
Each row shows how much one word attends to every other word; the weights in a row sum to 1.

Attention lets a neural network decide, for each output it produces, which parts of its input to draw on and how much. Instead of squeezing the whole input into one fixed summary, the model keeps every piece available and computes a fresh weighted mix at each step.

It was introduced for machine translation. In 2014 Bahdanau, Cho and Bengio argued that compressing a source sentence into a single fixed-length vector was the bottleneck of encoder–decoder translation. Their fix let the decoder search over all the source words each time it produced a word (read the paper). The weights it learned lined up well with which words translate which.

How it works

The version used today comes from the 2017 transformer paper. Each token is projected into three vectors, a query, key and value. The query is what this token is looking for, the key is what it offers to be matched on, and the value is the information it passes along.

  1. Compare the query with every key using a dot product, giving one score per token.
  2. Divide by the square root of the key length. Without this, scores grow large in bigger models and push softmax into a region where gradients all but vanish.
  3. Apply softmax, so the scores become weights that sum to 1.
  4. Output the weighted sum of the values.

The diagram runs exactly this on “the cat sat on the mat” with tiny three-number embeddings, each word acting as its own query and key. Words with similar vectors, like “cat” and “mat”, weight each other heavily, while “the”, whose vector is small, spreads its attention almost evenly. In a real model the query, key and value projections are learned, so the model itself decides what counts as relevant.

Main kinds

  • Self-attention: queries, keys and values all come from the same sequence, so each token gathers context from the others. It is the core of the transformer.
  • Cross-attention: queries come from one sequence and keys and values from another, like a translation decoder reading the source sentence, or an image generator reading a text prompt.
  • Multi-head attention: several attention operations in parallel, each with its own projections, so different heads can track different relationships. The original transformer used eight.

The catch

Every token is scored against every other, so the work grows with the square of the sequence length. That’s the main reason long context is expensive, and why memory-saving implementations of attention matter so much for large models.

Where it sits

Explore nearby

In the research

All papers →

2 papers that build on Attention Mechanism .

Sources

  1. Bahdanau, Cho and Bengio, "Neural Machine Translation by Jointly Learning to Align and Translate" . 2014; ICLR 2015
  2. Vaswani et al., "Attention Is All You Need" . 2017
  3. Luong, Pham and Manning, "Effective Approaches to Attention-based Neural Machine Translation" . 2015