00 / Artificial Intelligence
What Is AI Inference?
What happens when an AI actually answers you
Read this article OpenClose
- 01 In 30 seconds The short version
- 02 The word itself A closer look
- 03 In plain English A familiar analogy
- 04 How it works The moving parts
- 05 Compared Side by side
- 06 Why now Why inference matters
- 07 In a sentence Say it right
- 08 Misconceptions What people get wrong
- 09 FAQ Quick answers
- 10 Related Where to go next
- 11 In one sentence The takeaway
- 12 Further research A deeper dive
01 / In 30 seconds
AI inference is the stage where a trained AI model does its actual job: it takes new input, such as your question, and produces an output, such as an answer. Training builds the model once, at enormous cost. Inference runs it again and again, and every run costs a little compute.
- Training builds the model. Inference uses it.
- During inference the model does not learn or change.
- Training is paid for once. Inference is paid for on every answer.
- Training, once: a huge dataset is used to adjust the model’s weights again and again.
- Inference, every request: your prompt goes into the trained model, whose weights are fixed, and an answer comes out.
Links
02 / The word itself
English inference comes via Latin inferentia from inferre, “to bring in, carry in”: a conclusion is something you carry in from the evidence. The metaphor still fits: the model takes what it already holds and brings in an answer for the input in front of it.
In everyday English, an inference is a conclusion you reach by reasoning from what you know. In machine learning the word narrowed to a mechanical meaning: running a trained model on new input. The reasoning sense has not vanished, though. Researchers also use “inference” in the statistical sense, so “inference” alone can mean either. “AI inference” removes the doubt.
Words that travel with it
- run inference
- at inference time
- inference cost
- inference latency
- inference chip
- inference server
- real-time inference
- batch inference
Register. Technical, but now mainstream business English. You will also meet “inferencing” in vendor marketing; researchers write plain “inference”, and it is the safer choice.
03 / In plain English
Think of a chess player. For years they study: openings, endgames, thousands of games. That is
The analogy holds where it matters. The player does not get better mid-game, just as a chat model does not get smarter mid-conversation. And playing costs something every time: a thousand games cost roughly a thousand times what one game costs. Training is a cost you pay once. Inference is a cost you pay for as long as anyone uses the thing.
04 / How it works
A trained model is, at bottom, a very large collection of numbers called
For a chatbot, the journey has a few stages. Your text is first chopped into
Then comes
Two engineering tricks explain much of what you pay for. The first is the
Because decoding is limited more by how fast the chip can move the weights out of memory than by how fast it can multiply, a lot of inference engineering is about moving less data.
- Your prompt is split into tokens.
- Prefill: the model reads the whole prompt at once.
- Decode: the model predicts one token, adds it to the text, and predicts the next. This repeats, reusing the KV cache, until the answer is complete.
- The tokens are joined back into the answer you read.
05 / Training vs. inference
Exact same model, different moment.
Training builds the capability. Inference puts it to work.
| The distinction | Training | Inference |
|---|---|---|
| 01What happens | The model learns from data | The model answers new input |
| 02Direction of computation | Forwards, then backwards to adjust weights | Forwards only |
| 03Do the weights change? | Yes, constantly | No |
| 04When | Once, over weeks or months, before release | Every time anyone uses the model |
| 05Who pays | The lab that builds the model | Whoever runs it: the provider, or you |
| 06What people optimise | Quality of the finished model | Delay per answer and cost per answer |
06 / Why people are talking about it
As of September 2026
For a few years the headline number in AI was the cost of training the biggest models. The attention has been shifting to serving them. A model is trained once but may be queried billions of times, so for a widely used model the serving bill ends up larger than the training bill over its life. That is why inference now drives chip design, data-centre spending and pricing plans.
A second change is that answers themselves have become more expensive. “Reasoning” models spend many hidden tokens working through a problem before they reply, so one answer can use many times the compute of a plain chat reply. Research on spending extra compute at inference time has shown that, on some tasks, this can beat simply building a larger model.
The third is where inference happens. Most of it runs in data centres, but chips built specifically for serving models, and neural processing units in phones and laptops, are moving some of it to the edge of the network, closer to you.
07 / How to use it in a sentence
In practice
Use “inference” for the work a trained model does when it answers.
-
headline
The startup says its new chip cuts the cost of running inference by half.
-
workplace
Training was a one-off. It’s inference that’s eating our cloud budget.
-
casual
It’s quick because the inference happens on my phone, not in the cloud.
Watch out for
-
We need to inference the model.→ We need to run inference on the model. -
The model is learning during inference.→ The model does not learn during inference; its weights stay fixed.
08 / Common misconceptions
“The model learns from my conversation.”
During inference the weights do not change. A chatbot seems to remember you because the earlier turns of the conversation are fed back in with each new message. Some providers may later use conversations to train a future model, but that is a separate process, done afterwards, under their data policies.
“Inference means looking up a stored answer.”
The answer is generated, token by token, from the model’s weights. There is no database of finished answers behind it, which is also why a model can produce fluent text that is wrong.
“Inference is cheap, so it is the minor cost.”
A single answer is cheap. The total is not: inference is paid for on every use, by every user, indefinitely.
“The same prompt always gives the same answer.”
Most systems sample from a probability distribution at each step, so the output varies. Setting the temperature to zero removes most, though not all, of the variation.
09 / Frequently asked questions
What is the difference between AI training and inference?
Training is when a model learns: it processes huge amounts of data and adjusts its weights to get better at a task. Inference is when the finished model is used: it takes new input and produces an output, with its weights held fixed. Training happens once, up front. Inference happens every time someone uses the model.
Does AI inference need a GPU?
Large models usually run on GPUs or similar accelerators, because inference is a mass of parallel arithmetic. Small models run fine on ordinary processors, and many phones and laptops now include a neural processing unit for exactly this job. What you need depends on the size of the model and how fast you want the answer.
Why is AI inference expensive?
Each answer needs a large model’s weights read from memory many times, once for every token generated, on costly chips. The cost repeats with every request. The provider recovers it through subscriptions, per-token prices, or advertising, which is why long prompts and long answers cost more.
What is inference latency?
It is the delay between sending a request and getting the answer. For chatbots it is usually measured two ways: time to the first token, and how many tokens per second follow. Both depend on the model’s size, the chip, the length of the prompt, and how busy the service is.
Is ChatGPT doing inference or training when I chat?
Inference. The model behind a chatbot was trained earlier, and while you chat its weights are not being updated. Each reply is one run of the finished model over your conversation so far.
11 / In one sentence
Inference is the model, already trained and no longer changing, doing its work on the input in front of it, at a cost you pay every time.
12 / Further research
If you are thirsty for a more detailed deep dive into how a large language model is actually created from the information on the web, here is an excellent primer by one of the world’s leading AI researchers. He goes deep, but he has an excellent communicative style and explains these complex topics with real clarity.
13 / Sources & revisions
Revision history
- First published.