A transformer is a neural network design for processing sequences, such as the tokens of a sentence, by letting every position look directly at every other position. It was introduced in the 2017 paper Attention Is All You Need for machine translation, and it now underlies nearly every large language model.
Before it, the standard sequence models were recurrent networks, which read one token at a time and carry a running summary forward. That is slow to train, because step 50 has to wait for step 49, and information from far back tends to fade. The transformer drops recurrence. Self-attention links any two positions in a single step, so the whole sequence is processed at once and training spreads efficiently across GPUs. The paper’s base model trained in 12 hours on eight GPUs.
Inside one block
Reading the diagram from the bottom:
- Tokens become embeddings, and a positional encoding is added to each. Attention on its own ignores order, so without this “dog bites man” and “man bites dog” would look the same.
- Multi-head self-attention. Each token builds a query, a key and a value, compares its query against every key, and takes a weighted mix of the values. The original model ran 8 of these heads side by side, each free to track a different kind of relationship.
- Feed-forward network. A small network then processes each position on its own, transforming what attention gathered.
- Residual connections and layer normalization wrap both parts, which keeps very deep stacks trainable.
The original stacked 6 of these blocks in its encoder and 6 in its decoder. Today’s large models stack many more, but the block is recognisably the same.
Three shapes
- Encoder–decoder, the original: one stack reads the source sentence, the other writes the translation while attending to it.
- Encoder-only, like BERT (2018): reads a whole text in both directions, suited to classification and search.
- Decoder-only, like GPT and most chat models: each token may attend only to earlier ones, which is exactly what next-token generation needs.
The design also travelled beyond text. The Vision Transformer (2020) cuts an image into 16×16-pixel patches and treats them as tokens.
The catch
Self-attention compares every token with every other, so its cost grows with the square of the sequence length. Doubling the context window roughly quadruples the attention work, which is a big part of why long contexts are expensive and why so much engineering goes into making attention cheaper.