Runtime address
The full description of the system that was actually evaluated — model version, serving route, harness, tools, precision and operating conditions — not just the model's name.
"We tested GPT-X" used to be a reasonable sentence. A model name was a decent shorthand for the thing being evaluated. It no longer is. Today the same model label can be served through different routes, wrapped in different harnesses, given different tools and fall back to different models under load.
A runtime address is the precise answer to "which exact system did we test?". It records enough about the serving path and operating conditions that someone else could find, and re-run, the same system. It is the benchmark equivalent of a street address rather than a city name.
What it contains
- The model identifier and version
- The serving route, including any fallback behaviour
- The harness: context management, tool use, sub-agents and execution environment
- Precision, limits and other operating conditions
Why it matters
A score without the path that produced it is evidence about someone else's system. When the route changes materially, the evidence expires until the new route is qualified. A good benchmark result should carry its runtime address with it.
Read more in The benchmark needs a runtime address.
Related terms
Serving route
The actual path a request takes to be answered — which model, hardware, precision, routing and fallback rules handled it — as opposed to the model name a vendor advertises.
Harness
The software around a model that manages context, tool use, sub-agents and the execution environment. It shapes real-world capability, so it must be evaluated with the model, not ignored.
Benchmark
A standardised test used to compare AI models. Useful as evidence about someone else's system, but only trustworthy for you when it describes the exact serving route, harness and conditions you will run.
Model routing
Directing each AI request to a particular model or route based on cost, quality, latency or availability, so the model that answers can differ from request to request.
Used in these essays
OpenAI Decisions API turns probability into policy
The provocative part of OpenAI's Decisions API is not that AI can make choices. It is that a model score can quietly become an action — and an action can quietly become policy.
OpenAI Dots is a test of whether AI can carry a goal, not just complete a task
Dots matters less as another capable assistant than as a test of persistent delegation: can AI keep carrying a goal without giving the user a new system to manage?
The wrong questions about AI right now
Many of the questions that helped us orient ourselves around generative AI are now too blunt to be useful. The harder work is no longer asking what AI is in the abstract, but specifying where it works, where it fails, what authority it should have, and what the whole system costs.
A cheaper AI model can move the cost instead of removing it
A lower model bill can hide a higher workflow bill. The useful AI TCO question is not only what got cheaper, but where the cost moved.
Keep the AI ideas you rejected
A new model release should not restart your AI roadmap. It should reopen only the ideas that were rejected for a constraint the release actually changed.
Make the AI vendor demo fail
A polished AI demo proves that a system can succeed under prepared conditions. A buying decision needs different evidence: what happens when the system is wrong, blocked, uncertain or halfway through an action.
The hidden metric in AI automation is supervision
As AI moves from assisting work to leading it, hours saved stop telling the whole story. The scarce resource shifts to human supervision: approvals, exceptions, context and judgment.
The next AI interface may never be seen
Agents are turning software capabilities into an interface of their own. The next enterprise design problem is deciding what should be callable, by whom, and under which boundaries.
The benchmark needs a runtime address
AI systems are increasingly dynamic at runtime. Enterprise evaluation should qualify the serving route, harness, tools and fallback conditions, not just the model name.