English Reference

00 / Artificial Intelligence

What Is AI Inference?

What happens when an AI actually answers you

Read this article OpenClose
  1. 01 In 30 seconds The short version
  2. 02 The word itself A closer look
  3. 03 In plain English A familiar analogy
  4. 04 How it works The moving parts
  5. 05 Compared Side by side
  6. 06 Why now Why inference matters
  7. 07 In a sentence Say it right
  8. 08 Misconceptions What people get wrong
  9. 09 FAQ Quick answers
  10. 10 Related Where to go next
  11. 11 In one sentence The takeaway
  12. 12 Further research A deeper dive

01 / In 30 seconds

AI inference is the stage where a trained AI model does its actual job: it takes new input, such as your question, and produces an output, such as an answer. Training builds the model once, at enormous cost. Inference runs it again and again, and every run costs a little compute.

  • Training builds the model. Inference uses it.
  • During inference the model does not learn or change.
  • Training is paid for once. Inference is paid for on every answer.
Training happens once; inference happens on every requestA dataset feeds a one-time training step that produces a trained model. Each user prompt then goes into that fixed model and comes out as an answer.1 · TRAINING — onceHuge datasetWeights adjusted2 · INFERENCE — every requestYour promptTrained modelweights fixed — nothing is learnedAnswer…repeated for every user, every time
Training builds the model once. Inference is every later run of it.
  1. Training, once: a huge dataset is used to adjust the model’s weights again and again.
  2. Inference, every request: your prompt goes into the trained model, whose weights are fixed, and an answer comes out.

Links

02 / The word itself

AI inference /ˈɪnfɝəns/ UK /ˈɪnfərəns/ noun

English inference comes via Latin inferentia from inferre, “to bring in, carry in”: a conclusion is something you carry in from the evidence. The metaphor still fits: the model takes what it already holds and brings in an answer for the input in front of it.

In everyday English, an inference is a conclusion you reach by reasoning from what you know. In machine learning the word narrowed to a mechanical meaning: running a trained model on new input. The reasoning sense has not vanished, though. Researchers also use “inference” in the statistical sense, so “inference” alone can mean either. “AI inference” removes the doubt.

Words that travel with it

  • run inference
  • at inference time
  • inference cost
  • inference latency
  • inference chip
  • inference server
  • real-time inference
  • batch inference

Register. Technical, but now mainstream business English. You will also meet “inferencing” in vendor marketing; researchers write plain “inference”, and it is the safer choice.

03 / In plain English

Think of a chess player. For years they study: openings, endgames, thousands of games. That is training. It is slow, it is expensive, and by the end their knowledge is more or less fixed. Then they sit down at a board and play a game. That is inference. They are not studying during the game. They are using what they know on a position they have never seen, and each move still takes real effort.

The analogy holds where it matters. The player does not get better mid-game, just as a chat model does not get smarter mid-conversation. And playing costs something every time: a thousand games cost roughly a thousand times what one game costs. Training is a cost you pay once. Inference is a cost you pay for as long as anyone uses the thing.

04 / How it works

A trained model is, at bottom, a very large collection of numbers called weights. Inference is the act of pushing your input through those numbers to get an output. Nothing in the weights is altered on the way. This is what makes inference so much cheaper than training per step: training has to run the calculation forwards, then work backwards to figure out how to adjust every weight (a procedure called backpropagation), while inference only runs it forwards.

For a chatbot, the journey has a few stages. Your text is first chopped into tokens: chunks of a word, a whole short word, a piece of punctuation. Then the model reads your whole prompt at once, in what engineers call the prefill phase. That is heavy arithmetic, but it parallelises well, so a modern chip does it quickly. This is why the pause before the first word appears is usually short, even for a long prompt.

Then comes decoding, and this is the slow part. The model predicts one token, appends it to the text so far, and predicts the next, one at a time, because each token depends on the one before it. At every step the model produces a probability for every possible next token, and a sampling rule picks one. A setting called temperature controls how adventurous that pick is. That is also why the same prompt can give different answers.

Two engineering tricks explain much of what you pay for. The first is the KV cache: the model stores intermediate results for every token it has already processed, so it does not recompute them at each step. It saves a great deal of work, but the cache grows with the length of the conversation and has to live in the chip’s fast memory, which is scarce. The second is batching: a chip can serve many users’ requests in the same pass, so a busy service is cheaper per answer than an idle one.

Because decoding is limited more by how fast the chip can move the weights out of memory than by how fast it can multiply, a lot of inference engineering is about moving less data. Quantization is the standard example: storing the weights in 8 or 4 bits instead of 16 shrinks the memory traffic, and so cost and delay, at a small price in accuracy.

How a chatbot produces an answerThe prompt is split into tokens, the whole prompt is read at once in a prefill step, then the model decodes one token at a time in a loop until the answer is complete.PromptTokensPrefillall at onceDecodeone tokenAnswerrepeat until doneeach new token reuses the KV cache
The slow part is the loop: each token needs the one before it.
  1. Your prompt is split into tokens.
  2. Prefill: the model reads the whole prompt at once.
  3. Decode: the model predicts one token, adds it to the text, and predicts the next. This repeats, reusing the KV cache, until the answer is complete.
  4. The tokens are joined back into the answer you read.

05 / Training vs. inference

Exact same model, different moment.

Training builds the capability. Inference puts it to work.

The distinction Training Inference
01What happens The model learns from data The model answers new input
02Direction of computation Forwards, then backwards to adjust weights Forwards only
03Do the weights change? Yes, constantly No
04When Once, over weeks or months, before release Every time anyone uses the model
05Who pays The lab that builds the model Whoever runs it: the provider, or you
06What people optimise Quality of the finished model Delay per answer and cost per answer

06 / Why people are talking about it

As of September 2026

For a few years the headline number in AI was the cost of training the biggest models. The attention has been shifting to serving them. A model is trained once but may be queried billions of times, so for a widely used model the serving bill ends up larger than the training bill over its life. That is why inference now drives chip design, data-centre spending and pricing plans.

A second change is that answers themselves have become more expensive. “Reasoning” models spend many hidden tokens working through a problem before they reply, so one answer can use many times the compute of a plain chat reply. Research on spending extra compute at inference time has shown that, on some tasks, this can beat simply building a larger model.

The third is where inference happens. Most of it runs in data centres, but chips built specifically for serving models, and neural processing units in phones and laptops, are moving some of it to the edge of the network, closer to you.

07 / How to use it in a sentence

In practice

Use “inference” for the work a trained model does when it answers.

  • headline

    The startup says its new chip cuts the cost of running inference by half.
  • workplace

    Training was a one-off. It’s inference that’s eating our cloud budget.
  • casual

    It’s quick because the inference happens on my phone, not in the cloud.

Watch out for

  • We need to inference the model. → We need to run inference on the model.
  • The model is learning during inference. → The model does not learn during inference; its weights stay fixed.

08 / Common misconceptions

“The model learns from my conversation.”

During inference the weights do not change. A chatbot seems to remember you because the earlier turns of the conversation are fed back in with each new message. Some providers may later use conversations to train a future model, but that is a separate process, done afterwards, under their data policies.

“Inference means looking up a stored answer.”

The answer is generated, token by token, from the model’s weights. There is no database of finished answers behind it, which is also why a model can produce fluent text that is wrong.

“Inference is cheap, so it is the minor cost.”

A single answer is cheap. The total is not: inference is paid for on every use, by every user, indefinitely.

“The same prompt always gives the same answer.”

Most systems sample from a probability distribution at each step, so the output varies. Setting the temperature to zero removes most, though not all, of the variation.

09 / Frequently asked questions

What is the difference between AI training and inference?

Training is when a model learns: it processes huge amounts of data and adjusts its weights to get better at a task. Inference is when the finished model is used: it takes new input and produces an output, with its weights held fixed. Training happens once, up front. Inference happens every time someone uses the model.

Does AI inference need a GPU?

Large models usually run on GPUs or similar accelerators, because inference is a mass of parallel arithmetic. Small models run fine on ordinary processors, and many phones and laptops now include a neural processing unit for exactly this job. What you need depends on the size of the model and how fast you want the answer.

Why is AI inference expensive?

Each answer needs a large model’s weights read from memory many times, once for every token generated, on costly chips. The cost repeats with every request. The provider recovers it through subscriptions, per-token prices, or advertising, which is why long prompts and long answers cost more.

What is inference latency?

It is the delay between sending a request and getting the answer. For chatbots it is usually measured two ways: time to the first token, and how many tokens per second follow. Both depend on the model’s size, the chip, the length of the prompt, and how busy the service is.

Is ChatGPT doing inference or training when I chat?

Inference. The model behind a chatbot was trained earlier, and while you chat its weights are not being updated. Each reply is one run of the finished model over your conversation so far.

11 / In one sentence

Inference is the model, already trained and no longer changing, doing its work on the input in front of it, at a cost you pay every time.

12 / Further research

If you are thirsty for a more detailed deep dive into how a large language model is actually created from the information on the web, here is an excellent primer by one of the world’s leading AI researchers. He goes deep, but he has an excellent communicative style and explains these complex topics with real clarity.

Watch on YouTube · loads from YouTube when you press play

13 / Sources & revisions