Model routing
Directing each AI request to a particular model or route based on cost, quality, latency or availability, so the model that answers can differ from request to request.
Model routing is the practice of choosing, request by request, which model or deployment should handle a task. A simple question may go to a small, fast, cheap model; a hard one to a larger model. A request may also move to a fallback model when the first choice is slow, unavailable or over a limit.
Routing can cut cost and latency, and it improves resilience. It also means that "the model" behind a service is not one fixed thing. Some providers now return routing metadata with each response, showing which model actually answered, which routes were attempted and whether a fallback occurred.
Why it matters
If the model that answers changes, so do quality, cost and failure behaviour. Evaluation must therefore describe the serving route and its fallback conditions, not just a model name, and record them in the runtime address. Routing rules are also part of the harness a team controls, which makes them something to version, test and review.
Routing metadata is becoming evidence: it lets a buyer verify what actually ran.
Read more in The benchmark needs a runtime address.
Related terms
Serving route
The actual path a request takes to be answered — which model, hardware, precision, routing and fallback rules handled it — as opposed to the model name a vendor advertises.
Runtime address
The full description of the system that was actually evaluated — model version, serving route, harness, tools, precision and operating conditions — not just the model's name.
Harness
The software around a model that manages context, tool use, sub-agents and the execution environment. It shapes real-world capability, so it must be evaluated with the model, not ignored.
Used in these essays
A cheaper AI model can move the cost instead of removing it
A lower model bill can hide a higher workflow bill. The useful AI TCO question is not only what got cheaper, but where the cost moved.
The benchmark needs a runtime address
AI systems are increasingly dynamic at runtime. Enterprise evaluation should qualify the serving route, harness, tools and fallback conditions, not just the model name.