- Natural Language Processing ✓ understood
The field of AI that lets computers read, interpret, translate and generate human language, from spam filters and search to chatbots.
- Token ✓ understood
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
- Tokenization ✓ understood
Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.
- Language Modeling ✓ understood
Learning probability distributions over sequences of words to predict what comes next.
- Autoregressive Model ✓ understood
A model that generates output one token at a time, using previously generated tokens as input for the next prediction.
- Greedy Decoding ✓ understood
Always selecting the most likely next token during generation, fast but can lead to repetitive or suboptimal outputs.
- Beam Search ✓ understood
A generation algorithm that maintains top-k candidates at each step, balancing quality and diversity.
- Length Penalty · you are here ✓ understood