The input provided by the user to the language model, containing questions or instructions.
This concept is essential for understanding large language models and forms a key part of modern AI systems.
Related Concepts
- Prompt
- Prompt Engineering
- Input
The input provided by the user to the language model, containing questions or instructions.
The field of AI that lets computers read, interpret, translate and generate human language, from spam filters and search to chatbots.
The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.
Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.
Learning probability distributions over sequences of words to predict what comes next.
A collection of data examples used for training, validating, or testing machine learning models.
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.
Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.
Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).
Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning.
A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.
A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.
A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.
Learning useful features or representations of data automatically, rather than hand-crafting them.
A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.
A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.
A neural network architecture, introduced in 2017, built from stacked self-attention and feed-forward layers; the basis of nearly every modern large language model.
A neural network, almost always a transformer, trained on vast amounts of text to predict the next token, which lets it write, answer, summarize and follow instructions.
The input provided by the user to the language model, containing questions or instructions.
This concept is essential for understanding large language models and forms a key part of modern AI systems.