Reference 3 stops to get here
BPE
Byte Pair Encoding - a subword tokenization algorithm that iteratively merges frequent character pairs to create a vocabulary.
Your route here
3 stops · basics first
- Token ✓ understood
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
- Tokenization ✓ understood
Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.
- Subword Tokenization ✓ understood
Breaking words into smaller units, balancing vocabulary size with representation granularity.
- BPE · you are here ✓ understood
Where it sits
Explore nearby
Language & LLMs WordPiece A subword tokenization algorithm used by BERT, similar to BPE but with different merging criteria. Language & LLMs SentencePiece A language-agnostic tokenization library that treats text as a sequence of Unicode characters. Language & LLMs Vocabulary Size The number of distinct tokens a language model can process, typically 30K-100K+ tokens. Language & LLMs Tokenization Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.