Reference 2 stops to get here
Lemmatization
Reducing words to their base or dictionary form (running → run) using linguistic knowledge.
Your route here
2 stops · basics first
- Token ✓ understood
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
- Tokenization ✓ understood
Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.
- Lemmatization · you are here ✓ understood
Where it sits
Explore nearby
Language & LLMs Stemming Reducing words to their root form by removing affixes (prefixes, suffixes, infixes), simpler than lemmatization but less linguistically accurate. Language & LLMs Stop Words Common words (the, is, at) often removed in NLP preprocessing as they carry little semantic meaning. Language & LLMs Part-of-Speech Tagging Labeling words in text with their grammatical roles (noun, verb, adjective, etc.). Foundations Data Preprocessing Cleaning, transforming, and preparing raw data for model training (handling missing values, normalization, encoding).