Berk Bayri

Inference

Running a trained AI model to produce an output for a request, as opposed to training it. Inference cost is the model-side spend that is easy to see but only one part of total AI cost.

Inference is the stage where a trained model is used: you send a request and it computes a response. It is distinct from training, where the model's weights are learned. Every time a user asks a question, an agent calls a tool or a classifier scores an input, inference happens.

A production request can pass through much more than one model call. It may go through a router, retrieve context, call tools, use a particular inference configuration, invoke another model, hit a permission boundary, retry, fall back or be escalated to a person. The route and configuration therefore matter as much as the model name: see serving route.

Inference cost

Inference cost is the model-side spend, usually driven by tokens. It is the most visible line in an AI budget and the easiest to cut, which is the risk: lowering it can raise review, exception and recovery costs elsewhere. That is what the cost-migration test checks, as part of total cost of ownership.

Read more in A cheaper AI model can move the cost instead of removing it.