Landmark 2 stops to get here · leads to 10

Inference

Running a trained model on new inputs to get predictions, with its weights frozen: the stage of a model's life that users actually interact with.

Your route here

2 stops · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  3. Inference · you are here ✓ understood

Picture it

Training

  • Forward pass plus backward pass
  • Weights change every step
  • Runs once, costly, on labeled data

Inference

  • Forward pass only
  • Weights are frozen
  • Runs on every request, new data
  • Judged on latency and cost per call
Notice inference is the same forward pass as training with the learning switched off: weights stay fixed and only predictions come out.

Inference is using a model. Training set the weights; inference feeds new inputs through those fixed weights and returns an output: a label, a score, the next token of a reply. Every time a chatbot answers a message or a photo app tags a face, that’s inference.

A model is trained once, or occasionally, but inference runs on every request for as long as the model is in service. At scale that adds up: Google reported that across three recent years, about three-fifths of its machine-learning energy went to inference and two-fifths to training.

What it’s judged on

Inference is measured on different numbers from training:

  • Latency: how long one request takes. For chat, people notice both the wait for the first word and the speed after that.
  • Throughput: how many requests, or tokens, the hardware handles per second.
  • Cost per request: often what decides whether a product is viable.

These pull against each other. Processing many requests together as a batch keeps the hardware busy and raises throughput, but a request may wait for its batch. Google’s 2022 work on serving its 540-billion-parameter PaLM model frames inference as exactly this trade-off, shaped by model size, context length and how strict the latency target is.

Inference for language models

LLMs generate one token at a time, and each new token needs a pass through the whole model. The prompt is read first in one parallel pass; output tokens then come out one by one. To avoid recomputing the prompt at every step, the model keeps intermediate attention values for earlier tokens in memory, the KV cache, which grows with the context. The vLLM system manages that memory the way an operating system manages pages of virtual memory, which let it fit more requests on each GPU and raise throughput 2–4× at the same latency. This is also why longer prompts cost more, a point The Duel turns into a game.

Making it cheaper

  • Quantization: store weights in 8 or 4 bits instead of 16, cutting memory and often speeding things up for a small loss in accuracy.
  • Distillation: train a smaller model to imitate a larger one.
  • Serving infrastructure: batching, caching and autoscaling around the model.

Where it sits

Explore nearby

Sources

  1. Pope et al., "Efficiently Scaling Transformer Inference" . Google, 2022
  2. Kwon et al., "Efficient Memory Management for Large Language Model Serving with PagedAttention" . 2023 (vLLM)
  3. David Patterson, "Good News About the Carbon Footprint of Machine Learning Training" . Google Research, February 2022