Serving route
The actual path a request takes to be answered — which model, hardware, precision, routing and fallback rules handled it — as opposed to the model name a vendor advertises.
The serving route is the operational unit behind an AI answer. When you send a request to an AI service, the model name on the invoice is only part of the story. The request may be routed to one of several model variants, run at a particular numerical precision, handled by a specific deployment, retried, or answered by a fallback model if the first choice is busy.
Some providers now expose per-request routing metadata showing which model actually served a request, which routes were attempted and whether a fallback occurred. That is a sign the industry is treating the route, not the label, as the thing that matters.
Why it matters
Enterprises buy labels but operate routes. Quality, latency, cost and failure behaviour are properties of the route. Evaluating the label while running a different route means the evidence you collected may not describe the system you deployed.
Qualify the route you actually intend to use, record it as part of the runtime address, and re-qualify when it changes materially. See also model routing and the harness that surrounds the model.
Read more in The benchmark needs a runtime address.
Related terms
Runtime address
The full description of the system that was actually evaluated — model version, serving route, harness, tools, precision and operating conditions — not just the model's name.
Model routing
Directing each AI request to a particular model or route based on cost, quality, latency or availability, so the model that answers can differ from request to request.
Harness
The software around a model that manages context, tool use, sub-agents and the execution environment. It shapes real-world capability, so it must be evaluated with the model, not ignored.
Used in these essays
Make the AI vendor demo fail
A polished AI demo proves that a system can succeed under prepared conditions. A buying decision needs different evidence: what happens when the system is wrong, blocked, uncertain or halfway through an action.
The hidden metric in AI automation is supervision
As AI moves from assisting work to leading it, hours saved stop telling the whole story. The scarce resource shifts to human supervision: approvals, exceptions, context and judgment.
The benchmark needs a runtime address
AI systems are increasingly dynamic at runtime. Enterprise evaluation should qualify the serving route, harness, tools and fallback conditions, not just the model name.