IBIB
A protocol for measuring enterprise AI systems by serving route rather than model identifier. Its authors audited 18 benchmarks and found all scored advertised model names, not the full route.
IBIB is the name of a protocol described in the paper "IBIB: A Protocol for Measuring Enterprise AI Systems by Serving Route, Not Model Identifier". Its authors audited 18 benchmarks and report that all 18 scored advertised model identifiers rather than the full serving route.
Their argument is simple: usable capability is jointly shaped by the weights, the serving route, precision, the output contract and the harness. A score for "the model" therefore describes only part of what an enterprise actually runs.
Why it is useful
IBIB gives an evidence-based reason to change what a benchmark reports. It supports the idea of a runtime address: record the route, harness and conditions behind a result so it can be reproduced, and treat the result as expiring when the route changes materially.
Read more in The benchmark needs a runtime address.
Related terms
Serving route
The actual path a request takes to be answered — which model, hardware, precision, routing and fallback rules handled it — as opposed to the model name a vendor advertises.
Runtime address
The full description of the system that was actually evaluated — model version, serving route, harness, tools, precision and operating conditions — not just the model's name.
Benchmark
A standardised test used to compare AI models. Useful as evidence about someone else's system, but only trustworthy for you when it describes the exact serving route, harness and conditions you will run.
Harness
The software around a model that manages context, tool use, sub-agents and the execution environment. It shapes real-world capability, so it must be evaluated with the model, not ignored.