A subword tokenization algorithm used by BERT, similar to BPE but with different merging criteria.
This concept is essential for understanding large language models and forms a key part of modern AI systems.
Related Concepts
- Tokenization
- BPE
- BERT
A subword tokenization algorithm used by BERT, similar to BPE but with different merging criteria.
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.
Breaking words into smaller units, balancing vocabulary size with representation granularity.
A subword tokenization algorithm used by BERT, similar to BPE but with different merging criteria.
This concept is essential for understanding large language models and forms a key part of modern AI systems.