Harness
The software around a model that manages context, tool use, sub-agents and the execution environment. It shapes real-world capability, so it must be evaluated with the model, not ignored.
In AI engineering, the harness is everything around the model that turns raw capability into a working system: how context is assembled and managed, which tools the model can call, how sub-agents are spawned, how outputs are checked and retried, and what environment code or actions run in.
Two systems using the same model can behave very differently because their harnesses differ. Some vendors now treat the harness as something that evolves alongside the model itself. That is why the harness is part of capability, not a wrapper to be ignored.
Why it matters for evaluation
A benchmark score reports on a model inside someone's harness. If your production harness differs, the score may not describe your system. The harness, the serving route and the operating conditions together make up the runtime address, the right unit to qualify.
It also matters for agentic AI: an agent's reliability depends heavily on how its harness handles tool failures, long tasks and unexpected states.
Read more in The benchmark needs a runtime address.
Related terms
Runtime address
The full description of the system that was actually evaluated — model version, serving route, harness, tools, precision and operating conditions — not just the model's name.
Serving route
The actual path a request takes to be answered — which model, hardware, precision, routing and fallback rules handled it — as opposed to the model name a vendor advertises.
Agentic AI
AI systems that pursue a goal by planning, using tools and taking multi-step actions with some autonomy, rather than only answering a single prompt.
Benchmark
A standardised test used to compare AI models. Useful as evidence about someone else's system, but only trustworthy for you when it describes the exact serving route, harness and conditions you will run.
Used in these essays
The wrong questions about AI right now
Many of the questions that helped us orient ourselves around generative AI are now too blunt to be useful. The harder work is no longer asking what AI is in the abstract, but specifying where it works, where it fails, what authority it should have, and what the whole system costs.
Make the AI vendor demo fail
A polished AI demo proves that a system can succeed under prepared conditions. A buying decision needs different evidence: what happens when the system is wrong, blocked, uncertain or halfway through an action.
The first AI incident report should be incomplete
OpenAI's new misalignment disclosure framework exposes a useful enterprise design principle: record anomalous AI behavior before the organization has finished explaining it. Otherwise incident systems quietly become filters for what teams already understand.
The hidden metric in AI automation is supervision
As AI moves from assisting work to leading it, hours saved stop telling the whole story. The scarce resource shifts to human supervision: approvals, exceptions, context and judgment.
The next AI interface may never be seen
Agents are turning software capabilities into an interface of their own. The next enterprise design problem is deciding what should be callable, by whom, and under which boundaries.
The benchmark needs a runtime address
AI systems are increasingly dynamic at runtime. Enterprise evaluation should qualify the serving route, harness, tools and fallback conditions, not just the model name.