Inference
Running a trained AI model to produce an output for a request, as opposed to training it. Inference cost is the model-side spend that is easy to see but only one part of total AI cost.
Inference is the stage where a trained model is used: you send a request and it computes a response. It is distinct from training, where the model's weights are learned. Every time a user asks a question, an agent calls a tool or a classifier scores an input, inference happens.
A production request can pass through much more than one model call. It may go through a router, retrieve context, call tools, use a particular inference configuration, invoke another model, hit a permission boundary, retry, fall back or be escalated to a person. The route and configuration therefore matter as much as the model name: see serving route.
Inference cost
Inference cost is the model-side spend, usually driven by tokens. It is the most visible line in an AI budget and the easiest to cut, which is the risk: lowering it can raise review, exception and recovery costs elsewhere. That is what the cost-migration test checks, as part of total cost of ownership.
Read more in A cheaper AI model can move the cost instead of removing it.
Related terms
Token
The small chunk of text, roughly a word or part of a word, that language models read and write. Providers usually price and limit usage in tokens, but token cost is not the cost of a result.
Serving route
The actual path a request takes to be answered — which model, hardware, precision, routing and fallback rules handled it — as opposed to the model name a vendor advertises.
Total cost of ownership
The full cost of running an AI system across its life: model and inference spend plus integration, human review, exceptions, recovery, governance and maintenance — not just the price per call.
Used in these essays
OpenAI Dots is a test of whether AI can carry a goal, not just complete a task
Dots matters less as another capable assistant than as a test of persistent delegation: can AI keep carrying a goal without giving the user a new system to manage?
Cahit Arf’s 1959 question about thinking machines
In a 1959 text on thinking machines, Cahit Arf moved from clocks and relays to a harder question: what happens when a machine meets a problem that was not anticipated when it was built?
The wrong questions about AI right now
Many of the questions that helped us orient ourselves around generative AI are now too blunt to be useful. The harder work is no longer asking what AI is in the abstract, but specifying where it works, where it fails, what authority it should have, and what the whole system costs.
A cheaper AI model can move the cost instead of removing it
A lower model bill can hide a higher workflow bill. The useful AI TCO question is not only what got cheaper, but where the cost moved.