Reference 2 stops to get here
Stop Words
Common words (the, is, at) often removed in NLP preprocessing as they carry little semantic meaning.
Your route here
2 stops · basics first
- Token ✓ understood
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
- Tokenization ✓ understood
Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.
- Stop Words · you are here ✓ understood
Where it sits
Explore nearby
Language & LLMs Stemming Reducing words to their root form by removing affixes (prefixes, suffixes, infixes), simpler than lemmatization but less linguistically accurate. Language & LLMs Lemmatization Reducing words to their base or dictionary form (running → run) using linguistic knowledge. Language & LLMs TF-IDF Term Frequency-Inverse Document Frequency - a statistical measure of word importance in documents, used for information retrieval. Foundations Data Preprocessing Cleaning, transforming, and preparing raw data for model training (handling missing values, normalization, encoding).