Berk Bayri

Benchmark

A standardised test used to compare AI models. Useful as evidence about someone else's system, but only trustworthy for you when it describes the exact serving route, harness and conditions you will run.

An AI benchmark is a fixed set of tasks and a scoring method used to compare models: reasoning problems, coding challenges, question answering and so on. Leaderboards built on benchmarks have become the shorthand for "which model is best".

Benchmarks are real evidence, but about a specific thing. A leaderboard is evidence about someone else's system, under someone else's conditions. A model score without the path that produced it is increasingly hard to trust, because the same model name may be served through different routes and wrapped in different harnesses.

Using benchmarks well

  • Ask which exact system was tested: the runtime address.
  • Check whether the serving route and harness match what you will deploy.
  • Keep failure in the score, not just the average.
  • Treat the evidence as expiring when the route changes materially.

"Which model is best?" is the wrong question. The better one is which configured system performs best on your workload at your acceptable quality, latency, privacy, cost and recovery constraints. A benchmark also says little about adaptation beyond what it tests.

Read more in The benchmark needs a runtime address.

Used in these essays

AI Agents & Systems13 min read

OpenAI Decisions API turns probability into policy

The provocative part of OpenAI's Decisions API is not that AI can make choices. It is that a model score can quietly become an action — and an action can quietly become policy.

AI Agents & Systems13 min read

OpenAI Dots is a test of whether AI can carry a goal, not just complete a task

Dots matters less as another capable assistant than as a test of persistent delegation: can AI keep carrying a goal without giving the user a new system to manage?

Problem Framing & Decision Design10 min read

Cahit Arf’s 1959 question about thinking machines

In a 1959 text on thinking machines, Cahit Arf moved from clocks and relays to a harder question: what happens when a machine meets a problem that was not anticipated when it was built?

Problem Framing & Decision Design14 min read

The wrong questions about AI right now

Many of the questions that helped us orient ourselves around generative AI are now too blunt to be useful. The harder work is no longer asking what AI is in the abstract, but specifying where it works, where it fails, what authority it should have, and what the whole system costs.

AI Strategy & Transformation7 min read

A cheaper AI model can move the cost instead of removing it

A lower model bill can hide a higher workflow bill. The useful AI TCO question is not only what got cheaper, but where the cost moved.

Problem Framing & Decision Design8 min read

Keep the AI ideas you rejected

A new model release should not restart your AI roadmap. It should reopen only the ideas that were rejected for a constraint the release actually changed.

AI Strategy & Transformation7 min read

Make the AI vendor demo fail

A polished AI demo proves that a system can succeed under prepared conditions. A buying decision needs different evidence: what happens when the system is wrong, blocked, uncertain or halfway through an action.

AI Governance & Evaluation8 min read

The first AI incident report should be incomplete

OpenAI's new misalignment disclosure framework exposes a useful enterprise design principle: record anomalous AI behavior before the organization has finished explaining it. Otherwise incident systems quietly become filters for what teams already understand.

Innovation & Capability9 min read

The hidden metric in AI automation is supervision

As AI moves from assisting work to leading it, hours saved stop telling the whole story. The scarce resource shifts to human supervision: approvals, exceptions, context and judgment.

AI Agents & Systems9 min read

The next AI interface may never be seen

Agents are turning software capabilities into an interface of their own. The next enterprise design problem is deciding what should be callable, by whom, and under which boundaries.

AI Governance & Evaluation9 min read

The benchmark needs a runtime address

AI systems are increasingly dynamic at runtime. Enterprise evaluation should qualify the serving route, harness, tools and fallback conditions, not just the model name.

AI Strategy & Transformation9 min read

Your AI data agent is creating a second dataset

AI data agents do more than answer business questions. At scale, the pattern of questions, corrections and dead ends becomes a new signal: a map of what the organization is trying to understand, where its definitions are weak, and which decisions deserve better infrastructure.

AI Strategy & Transformation8 min read

Your chatbot is borrowing from the next interaction

AI customer service is usually measured one interaction at a time. But a failed automated interaction can change which channel a customer chooses next time. That makes future adoption part of the economics, not a separate trust metric.