Landmark 1 stop to get here · leads to 8

Tokenization

Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.

Your route here

1 stop · basics first
  1. Token ✓ understood

    The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.

  2. Tokenization · you are here ✓ understood

Picture it

  1. 01 Raw text "Tokenization is fun"
  2. 02 Split into subwords "Token" "ization" " is" " fun" (e.g. BPE)
  3. 03 Map to token IDs Each piece looks up its integer in the vocabulary
  4. 04 Model input IDs are turned into embedding vectors
Notice that the model never sees letters: text becomes a sequence of integer IDs, and the vocabulary decides where the splits fall.

Tokenization is the first step in feeding text to a language model: cutting it into pieces, called tokens, and replacing each piece with an integer ID from a fixed vocabulary. The model only ever sees those IDs. Every limit and price quoted in tokens, from a context window to an API bill, is counted after this step.

Why not words or letters?

Whole words make a vocabulary that is huge and still incomplete: every new name, typo or compound becomes an unknown word. Single characters cover everything, but sequences get very long, and the model has to learn spelling before it can learn meaning. Subword tokenization sits in between. Common words get a single token, and rarer ones are built from a few familiar pieces, so “Tokenization” might become “Token” + “ization”, as in the diagram. Unknown words become rare, because in the worst case a word falls back to small pieces.

How the vocabulary is learned

The splits aren’t written by hand. They’re learned from a large sample of text before the model itself is trained.

  • BPE (byte pair encoding) starts from single characters and repeatedly merges the most frequent adjacent pair into a new symbol. After enough merges, frequent words are single tokens. Sennrich and colleagues adapted this old compression technique for neural translation in 2015; the vocabulary size is simply the starting symbols plus the number of merges.
  • WordPiece builds a similar subword vocabulary from data. Google’s 2016 translation system used it, and BERT adopted it.
  • SentencePiece learns directly from raw sentences without first splitting on spaces, so the same method works for languages that don’t put spaces between words.

Why it matters in practice

  • Vocabulary size is a trade-off. A bigger vocabulary means shorter sequences, but a larger embedding table and more tokens the model rarely sees.
  • Not every language is tokenized equally. A tokenizer trained mostly on English text tends to cut other languages into more pieces, so the same sentence can cost more and fill more of the context.
  • Some odd behaviour starts here. Tasks that hinge on individual characters, like counting the letters in a word, are harder for a model that sees chunks rather than letters.

The catch

A tokenizer is fixed once the model is trained on it, since each token ID has its own learned embedding. Changing the tokenizer means retraining at least that part of the model, so its quirks usually last for the model’s whole life.

Where it sits

Explore nearby

Sources

  1. Sennrich, Haddow and Birch, "Neural Machine Translation of Rare Words with Subword Units" . 2015 (BPE for neural models)
  2. Wu et al., "Google's Neural Machine Translation System" . 2016 (wordpieces)
  3. Kudo and Richardson, "SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing" . 2018