Landmark 17 stops to get here · leads to 18

Large Language Model

A neural network, almost always a transformer, trained on vast amounts of text to predict the next token, which lets it write, answer, summarize and follow instructions.

Your route here

17 stops · basics first
  1. Natural Language Processing ✓ understood

    The field of AI that lets computers read, interpret, translate and generate human language, from spam filters and search to chatbots.

  2. Token ✓ understood

    The basic unit of text that a language model processes, typically representing a word, subword, or character. Tokens are the fundamental building blocks for LLM input and output.

  3. Tokenization ✓ understood

    Splitting text into tokens, usually subword pieces, and mapping each to an integer ID so a language model can process it.

  4. Language Modeling ✓ understood

    Learning probability distributions over sequences of words to predict what comes next.

  5. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  6. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  7. Machine Learning ✓ understood

    Building systems that learn patterns from data instead of following hand-written rules, getting better at a task as they see more examples.

  8. Unsupervised Learning ✓ understood

    Learning from unlabeled data to discover hidden patterns, structures, or relationships without explicit target outputs.

  9. Self-Supervised Learning ✓ understood

    Learning representations from unlabeled data by creating supervised tasks from the data itself (masked prediction, contrastive learning).

  10. Pre-training ✓ understood

    Training a model on a large dataset (often self-supervised) before fine-tuning on specific tasks, enabling transfer learning.

  11. Feature ✓ understood

    A single measurable property of an example, such as a house's floor area or how many links an email contains, used as an input to a model.

  12. Neural Network ✓ understood

    A computational model inspired by biological neural networks, consisting of interconnected nodes (neurons) organized in layers that process information through weighted connections.

  13. Deep Learning ✓ understood

    A subset of machine learning that uses neural networks with multiple layers (deep neural networks) to learn hierarchical representations of data.

  14. Representation Learning ✓ understood

    Learning useful features or representations of data automatically, rather than hand-crafting them.

  15. Embedding ✓ understood

    A list of numbers (a vector) that represents a word, sentence, image or other item, learned so that similar items end up close together.

  16. Attention Mechanism ✓ understood

    A technique that lets a neural network weigh every part of its input when producing each output, focusing on the parts most relevant at that step.

  17. Transformer ✓ understood

    A neural network architecture, introduced in 2017, built from stacked self-attention and feed-forward layers; the basis of nearly every modern large language model.

  18. Large Language Model · you are here ✓ understood

Picture it

  1. Next-token probabilities Softmax; one token is chosen and fed back
  2. Output logits One score per token in the vocabulary
  3. Transformer blocks × N Self-attention + feed-forward, dozens deep
  4. Token embeddings Each token becomes a vector, plus position
  5. Input text → tokens Text split into subword tokens
Notice an LLM is one stack run over and over: tokens go up through the transformer blocks and a single next token comes out the top.

A large language model is a neural network trained to do one thing: given some text, predict the next token. Do that across a large share of the web, books and code, and the model picks up grammar, facts, styles of argument and how programs fit together, because all of them help it guess what comes next.

Generating text is that same prediction run in a loop. The model scores every token in its vocabulary, one is picked, it’s appended to the input, and the whole network runs again. A 500-word answer is several hundred of these passes.

Nearly every modern LLM is a transformer, the architecture introduced in 2017, whose attention mechanism lets each token draw on every other token in the input.

How one gets built

  • Pre-training: the model learns next-token prediction on enormous amounts of raw text. The text supplies its own labels, so no one has to annotate it. This is where nearly all the compute goes.
  • Fine-tuning: a much smaller round on curated examples teaches it to follow instructions and hold a conversation.
  • Learning from preferences: people rank the model’s answers and the model is tuned toward the ones they prefer, most famously with RLHF. In OpenAI’s 2022 InstructGPT paper, labelers preferred a 1.3-billion-parameter model tuned this way over the 175-billion-parameter GPT-3.

Why scale mattered

The 2020 GPT-3 paper showed that a big enough model could take on a new task from a few examples written into the prompt, with no retraining at all. That is in-context learning, and the authors found it improved sharply as models grew. It’s why one model can translate, summarize, classify and write code: the task is described in the prompt rather than trained in.

The catch

  • It predicts plausible text, not checked text. When it doesn’t know, it can produce a fluent, confident, wrong answer: a hallucination.
  • It only sees its context window. Everything it uses must fit in that window, and every token is paid for on every call. The Duel makes that cost concrete.
  • Its knowledge stops at its training cutoff, unless you hand it documents or tools at run time.

Where it sits

Explore nearby

In the research

All papers →

26 papers that build on Large Language Model ; showing 5, canon first.

Canon · 2020 Language Models are Few-Shot Learners It showed that scale alone unlocks in-context learning, the capability today’s prompt-driven AI is built on. Canon · 2020 Retrieval-Augmented Generation for Knowledge-Intensive NLP Tasks It named and defined the pattern behind most production LLM apps that answer from your own documents. Canon · 2020 Measuring Massive Multitask Language Understanding It became the headline benchmark in LLM release notes for years, and a case study in benchmarks saturating. Canon · 2021 LoRA: Low-Rank Adaptation of Large Language Models It made adapting big models cheap enough for everyone, and is why fine-tunes are shared as small adapter files. Canon · 2022 Chain-of-Thought Prompting Elicits Reasoning in Large Language Models Thinking step by step went from a prompt trick to the basis of today’s reasoning models.

Sources

  1. Vaswani et al., "Attention Is All You Need" . 2017
  2. Brown et al., "Language Models are Few-Shot Learners" . The GPT-3 paper, 2020
  3. Ouyang et al., "Training language models to follow instructions with human feedback" . The InstructGPT paper, 2022