Berk Bayri

Better models fix bad AI products

A common misconception about AI models and evaluation, tested against the evidence.

The myth
If an AI product is unreliable, upgrading to a better model will usually fix it.
The reality
A stronger model can improve model-level performance, but product failures often live in context, retrieval, tools, workflow logic, permissions, handoffs, state and recovery.

Explanation and evidence

The argument

The model is important. It is also only one component.

An AI product can fail because retrieval returns the wrong evidence, tool descriptions are ambiguous, the agent receives stale state, permissions are too broad, a handoff loses context, the UI hides uncertainty, or the system has no recovery path after a partially completed action. A stronger model may compensate for some of those defects. It does not make them disappear.

OpenAI’s current evaluation guidance is useful here because it treats agent quality as an end-to-end workflow problem. Its examples explicitly test tool selection, tool arguments, handoffs, policies and traces, not just the text generated by the model. OpenAI’s third-party evaluation guidance goes further: for modern agents, performance depends not only on the model but on the environment and setup that let it act.

That is why a model upgrade can produce an impressive lift while the user experience remains fragile.

If the failure lives in the system, a better model can become a more capable participant in the same bad system.

Why people believe it

Model improvements are visible and easy to buy. Architecture work is slower. When a new frontier model arrives with better benchmark scores, switching the model feels like the highest-leverage fix.

Sometimes it is. But without diagnosis, “upgrade the model” is a guess.

What the evidence says

Production evals increasingly focus on traces and workflow behavior because end-to-end reliability emerges from interactions among model, tools, data, policies and orchestration. A system can use an excellent model and still fail consistently at the wrong boundary.

The better question

Before changing the model, classify the failure:

judgment, context, retrieval, tool design, state, authority, handoff, interface or recovery?

Then test the smallest change that targets the actual failure mode. A model upgrade should be a hypothesis with an eval, not a ritual.

Sources

Evaluation best practices

OpenAI · 2026-10-06

Evaluate agent workflows

OpenAI · 2026-10-06

A shared playbook for trustworthy third party evaluations

OpenAI · 2026-05-29

Related reading