Latency
The time an AI system takes to respond. It is one of the constraints, with quality, privacy, cost and recovery, that decides which configured system is best for a workload.
Latency is the delay between a request and the response. For AI it includes the model's generation time, any retrieval or tool calls, queueing, routing and retries. It varies widely with the serving route, the size of the model and the load.
Latency matters because it constrains where AI can be used. A decision that must happen in a conversation or in a live transaction has a tighter budget than a nightly report, and fast, cheap models may be right for one and wrong for the other.
In evaluation
"Which model is best?" is the wrong question. The useful one is which configured system performs best on your workload at acceptable quality, latency, privacy, cost and recovery constraints. Latency should be measured on the real path, including model routing and fallbacks, and recorded as part of the runtime address. It also trades off against cost: faster usually means more spend.
Read more in The wrong questions about AI right now.
Related terms
Serving route
The actual path a request takes to be answered — which model, hardware, precision, routing and fallback rules handled it — as opposed to the model name a vendor advertises.
Model routing
Directing each AI request to a particular model or route based on cost, quality, latency or availability, so the model that answers can differ from request to request.
Total cost of ownership
The full cost of running an AI system across its life: model and inference spend plus integration, human review, exceptions, recovery, governance and maintenance — not just the price per call.
Used in these essays
OpenAI Decisions API turns probability into policy
The provocative part of OpenAI's Decisions API is not that AI can make choices. It is that a model score can quietly become an action — and an action can quietly become policy.
The wrong questions about AI right now
Many of the questions that helped us orient ourselves around generative AI are now too blunt to be useful. The harder work is no longer asking what AI is in the abstract, but specifying where it works, where it fails, what authority it should have, and what the whole system costs.
A cheaper AI model can move the cost instead of removing it
A lower model bill can hide a higher workflow bill. The useful AI TCO question is not only what got cheaper, but where the cost moved.
Keep the AI ideas you rejected
A new model release should not restart your AI roadmap. It should reopen only the ideas that were rejected for a constraint the release actually changed.
The benchmark needs a runtime address
AI systems are increasingly dynamic at runtime. Enterprise evaluation should qualify the serving route, harness, tools and fallback conditions, not just the model name.
Your AI data agent is creating a second dataset
AI data agents do more than answer business questions. At scale, the pattern of questions, corrections and dead ends becomes a new signal: a map of what the organization is trying to understand, where its definitions are weak, and which decisions deserve better infrastructure.