Landmark A starting point · leads to 5

Token

The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.

Picture it

Word-level

  • "unhappiness" → 1 token
  • Huge vocabulary
  • Unknown words become [UNK]

Subword (BPE)

  • "unhappiness" → "un" + "happiness"
  • What most modern LLMs use
  • Handles rare words from known pieces

Character-level

  • "unhappiness" → 11 tokens
  • Tiny vocabulary
  • Long sequences, costly to process
Notice the trade-off: bigger tokens mean shorter sequences but larger vocabularies, and subwords sit in the middle, which is why LLMs use them.

Tokens are how language models read and generate text. A single token can be a complete word, part of a word, or even a single character, depending on the tokenization scheme used.

Examples

  • “Hello” might be 1 token
  • “unhappiness” might be split into [“un”, “happiness”] = 2 tokens
  • “ChatGPT” might be [“Chat”, “G”, “PT”] = 3 tokens

Importance

Token limits define how much text a model can process at once (context window). Understanding tokenization is crucial for prompt engineering and cost estimation, as most LLM APIs charge per token.

Common Tokenizers

  • BPE (Byte Pair Encoding): Used by GPT models
  • WordPiece: Used by BERT
  • SentencePiece: Language-agnostic tokenization

Where it sits

Before this

Nothing: this is a starting point.

Token

Explore nearby

In the research

All papers →

2 papers that build on Token .