Landmark 8 stops to get here · leads to 10

Transformer

A neural network architecture, introduced in 2017, built from stacked self-attention and feed-forward layers; the basis of nearly every modern large language model.

Your route here

8 stops · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Feature ✓ understood

    A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.

  3. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  4. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  5. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  6. Representation Learning ✓ understood

    Learning useful features or representations of data automatically, rather than hand-crafting them.

  7. Embedding ✓ understood

    A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.

  8. Attention Mechanism ✓ understood

    A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.

  9. Transformer · you are here ✓ understood

Picture it

  1. Linear + softmax Probabilities over the next token
  2. Add & layer norm Attention + FFN block repeats N times
  3. Feed-forward network Applied to each position independently
  4. Add & layer norm Residual connection around attention
  5. Multi-head self-attention Every token attends to every other token
  6. Embeddings + positions Token vectors plus positional encoding
  7. Input tokens
Notice there is no recurrence: attention mixes information across all positions at once, and the same block is stacked many times.

A transformer is a neural network design for processing sequences, such as the tokens of a sentence, by letting every position look directly at every other position. It was introduced in the 2017 paper Attention Is All You Need for machine translation, and it now underlies nearly every large language model.

Before it, the standard sequence models were recurrent networks, which read one token at a time and carry a running summary forward. That is slow to train, because step 50 has to wait for step 49, and information from far back tends to fade. The transformer drops recurrence. Self-attention links any two positions in a single step, so the whole sequence is processed at once and training spreads efficiently across GPUs. The paper’s base model trained in 12 hours on eight GPUs.

Inside one block

Reading the diagram from the bottom:

  • Tokens become embeddings, and a positional encoding is added to each. Attention on its own ignores order, so without this “dog bites man” and “man bites dog” would look the same.
  • Multi-head self-attention. Each token builds a query, a key and a value, compares its query against every key, and takes a weighted mix of the values. The original model ran 8 of these heads side by side, each free to track a different kind of relationship.
  • Feed-forward network. A small network then processes each position on its own, transforming what attention gathered.
  • Residual connections and layer normalization wrap both parts, which keeps very deep stacks trainable.

The original stacked 6 of these blocks in its encoder and 6 in its decoder. Today’s large models stack many more, but the block is recognisably the same.

Three shapes

  • Encoder–decoder, the original: one stack reads the source sentence, the other writes the translation while attending to it.
  • Encoder-only, like BERT (2018): reads a whole text in both directions, suited to classification and search.
  • Decoder-only, like GPT and most chat models: each token may attend only to earlier ones, which is exactly what next-token generation needs.

The design also travelled beyond text. The Vision Transformer (2020) cuts an image into 16×16-pixel patches and treats them as tokens.

The catch

Self-attention compares every token with every other, so its cost grows with the square of the sequence length. Doubling the context window roughly quadruples the attention work, which is a big part of why long contexts are expensive and why so much engineering goes into making attention cheaper.

Where it sits

Explore nearby

In the research

All papers →

6 papers that build on Transformer ; showing 5, canon first.

Canon · 2017 Attention Is All You Need Every major LLM, and most modern vision and speech models, is a transformer. Canon · 2018 BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding It established “pre-train once, fine-tune everywhere”, and still powers much of search and classification. Canon · 2020 An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale It brought vision and language onto one architecture, which made today’s multimodal models possible. Frontier · Mar 2026 · 68 citations Mamba-3: Improved Sequence Modeling using State Space Principles Inference cost now dominates, which makes fast, constant-memory architectures matter again. Frontier · Dec 2025 · 65 citations mHC: Manifold-Constrained Hyper-Connections The residual connection hadn’t changed in a decade; this is a stable way to go beyond it.

Sources

  1. Vaswani et al., "Attention Is All You Need" . 2017
  2. Devlin et al., "BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding" . 2018
  3. Dosovitskiy et al., "An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale" . 2020 (Vision Transformer)