Berk Bayri

Model routing

Directing each AI request to a particular model or route based on cost, quality, latency or availability, so the model that answers can differ from request to request.

Model routing is the practice of choosing, request by request, which model or deployment should handle a task. A simple question may go to a small, fast, cheap model; a hard one to a larger model. A request may also move to a fallback model when the first choice is slow, unavailable or over a limit.

Routing can cut cost and latency, and it improves resilience. It also means that "the model" behind a service is not one fixed thing. Some providers now return routing metadata with each response, showing which model actually answered, which routes were attempted and whether a fallback occurred.

Why it matters

If the model that answers changes, so do quality, cost and failure behaviour. Evaluation must therefore describe the serving route and its fallback conditions, not just a model name, and record them in the runtime address. Routing rules are also part of the harness a team controls, which makes them something to version, test and review.

Routing metadata is becoming evidence: it lets a buyer verify what actually ran.

Read more in The benchmark needs a runtime address.