Reference 4 stops to get here
TF-IDF
Term Frequency-Inverse Document Frequency - a statistical measure of word importance in documents, used for information retrieval.
Your route here
4 stops · basics first
- Token ✓ understood
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
- Tokenization ✓ understood
Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Feature ✓ understood
A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.
- TF-IDF · you are here ✓ understood
Where it sits
Explore nearby
Language & LLMs N-gram A contiguous sequence of n items (words, characters) from text, used in language modeling and feature extraction. Language & LLMs Stop Words Common words (the, is, at) often removed in NLP preprocessing as they carry little semantic meaning. Language & LLMs Text Classification Assigning categories or labels to text documents, a fundamental NLP task. Language & LLMs Topic Modeling Discovering abstract topics in document collections, often using techniques like LDA. Foundations One-Hot Encoding Converting categorical variables into binary vectors with one element set to 1 and others to 0.