Reference 3 stops to get here
Throughput
The number of predictions or tokens a model can process per unit of time, a key deployment performance metric.
Your route here
3 stops · basics first
- Dataset ✓ understood
A collection of data examples used for training, validating, or testing machine learning models.
- Training ✓ understood
The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.
- Inference ✓ understood
Running a trained model on new inputs to get predictions, with its weights frozen: the stage of a model's life that users actually interact with.
- Throughput · you are here ✓ understood
Where it sits
Explore nearby
Shipping AI Inference Latency The time delay between submitting input and receiving output from a deployed model, critical for real-time applications. Shipping AI Request Batching Combining multiple inference requests into batches to improve throughput. Shipping AI Batch Processing Processing multiple predictions together in batches rather than one at a time, improving throughput efficiency. Shipping AI Model Serving Deploying trained models as services that can handle prediction requests in production environments. Shipping AI GPU Graphics Processing Unit - hardware accelerator with thousands of cores, essential for parallel computation in deep learning.
In the research
All papers →A paper that builds on Throughput .