Reference 2 stops to get here

Stemming

Reducing words to their root form by removing affixes (prefixes, suffixes, infixes), simpler than lemmatization but less linguistically accurate.

Your route here

2 stops · basics first
  1. Token ✓ understood

    The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.

  2. Tokenization ✓ understood

    Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.

  3. Stemming · you are here ✓ understood

Where it sits

Before this

Tokenization
Stemming

Leads to

Nothing yet: a destination in its own right.

Explore nearby