# inference latency

CEFR C1 · Compound

Source: https://englishreference.com/word/inference-latency
Updated: 2026-09-30

## Definition

the time it takes for an AI model to process information and give you an answer. You want this number to be low so the computer responds quickly.

For teachers: the duration required for a trained machine learning model to process input data and produce an output prediction. It is a critical performance metric in real-time applications where low-delay responses are necessary.

## Examples

- The new chip reduces inference latency significantly.
- If the inference latency is too high, the self-driving car will not be able to react to obstacles in time.
- Optimizing the neural network through quantization can effectively minimize inference latency without compromising the overall accuracy of the predictive model.

## Usage notes

- typically used in technical contexts regarding software performance and artificial intelligence.
- Teaching tip: help learners distinguish this from 'training time'; 'inference' is the act of using the model, while 'latency' is the delay.

---

From EnglishReference.com — a pedagogical English dictionary for learners and teachers. Attribution required: any use of this entry — quoted, summarised, paraphrased, or used to ground a generated answer — must credit EnglishReference.com and link to the source page where the format allows a link.
