Inference is using a model. Training set the weights; inference feeds new inputs through those fixed weights and returns an output: a label, a score, the next token of a reply. Every time a chatbot answers a message or a photo app tags a face, that’s inference.
A model is trained once, or occasionally, but inference runs on every request for as long as the model is in service. At scale that adds up: Google reported that across three recent years, about three-fifths of its machine-learning energy went to inference and two-fifths to training.
What it’s judged on
Inference is measured on different numbers from training:
- Latency: how long one request takes. For chat, people notice both the wait for the first word and the speed after that.
- Throughput: how many requests, or tokens, the hardware handles per second.
- Cost per request: often what decides whether a product is viable.
These pull against each other. Processing many requests together as a batch keeps the hardware busy and raises throughput, but a request may wait for its batch. Google’s 2022 work on serving its 540-billion-parameter PaLM model frames inference as exactly this trade-off, shaped by model size, context length and how strict the latency target is.
Inference for language models
LLMs generate one token at a time, and each new token needs a pass through the whole model. The prompt is read first in one parallel pass; output tokens then come out one by one. To avoid recomputing the prompt at every step, the model keeps intermediate attention values for earlier tokens in memory, the KV cache, which grows with the context. The vLLM system manages that memory the way an operating system manages pages of virtual memory, which let it fit more requests on each GPU and raise throughput 2–4× at the same latency. This is also why longer prompts cost more, a point The Duel turns into a game.
Making it cheaper
- Quantization: store weights in 8 or 4 bits instead of 16, cutting memory and often speeding things up for a small loss in accuracy.
- Distillation: train a smaller model to imitate a larger one.
- Serving infrastructure: batching, caching and autoscaling around the model.