Standard 3 stops to get here

Inference Latency

The time delay between submitting input and receiving output from a deployed model, critical for real-time applications.

Your route here

3 stops · basics first
  1. Dataset ✓ understood

    A collection of data examples used for training, validating, or testing machine learning models.

  2. Training ✓ understood

    The process of fitting a model to data by repeatedly measuring how wrong its outputs are and adjusting its parameters to reduce that error.

  3. Inference ✓ understood

    Running a trained model on new inputs to get predictions, with its weights frozen: the stage of a model's life that users actually interact with.

  4. Inference Latency · you are here ✓ understood

Where it sits

Before this

Inference
Inference Latency

Leads to

Nothing yet: a destination in its own right.

Explore nearby

In the research

All papers →

11 papers that build on Inference Latency ; showing 5, canon first.

Frontier · Aug 2025 · 1.2K citations InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency It narrows the gap between open and commercial multimodal models on reasoning and agent tasks. Frontier · Sep 2025 · 189 citations LongLive: Real-time Interactive Long Video Generation Minute-long, steerable video at 20 frames per second on one GPU makes interactive video generation real. Frontier · Aug 2025 · 163 citations Seed Diffusion: A Large-Scale Diffusion Language Model with High-Speed Inference It is evidence that diffusion language models can be dramatically faster without giving up quality. Frontier · Sep 2025 · 124 citations MiniCPM-V 4.5: Cooking Efficient MLLMs via Architecture, Data, and Training Recipe Strong multimodal AI that runs on modest hardware. Frontier · Sep 2025 · 114 citations Fast-dLLM v2: Efficient Block-Diffusion LLM A cheap path to faster inference from models you already have.