More context makes an agent smarter: a common AI agent myth
RealityUseful context is selective. More tokens can bury relevant evidence, preserve stale assumptions and increase the amount of contradictory material the agent has to resolve.
A common misconception about AI models and evaluation, tested against the evidence.
The model is important. It is also only one component.
An AI product can fail because retrieval returns the wrong evidence, tool descriptions are ambiguous, the agent receives stale state, permissions are too broad, a handoff loses context, the UI hides uncertainty, or the system has no recovery path after a partially completed action. A stronger model may compensate for some of those defects. It does not make them disappear.
OpenAI’s current evaluation guidance is useful here because it treats agent quality as an end-to-end workflow problem. Its examples explicitly test tool selection, tool arguments, handoffs, policies and traces, not just the text generated by the model. OpenAI’s third-party evaluation guidance goes further: for modern agents, performance depends not only on the model but on the environment and setup that let it act.
That is why a model upgrade can produce an impressive benchmark lift while the user experience remains fragile.
If the failure lives in the system, a better model can become a more capable participant in the same bad system.
Model improvements are visible and easy to buy. Architecture work is slower. When a new frontier model arrives with better benchmark scores, switching the model feels like the highest-leverage fix.
Sometimes it is. But without diagnosis, “upgrade the model” is a guess.
Production evals increasingly focus on traces and workflow behavior because end-to-end reliability emerges from interactions among model, tools, data, policies and orchestration. A system can use an excellent model and still fail consistently at the wrong boundary.
Before changing the model, classify the failure:
judgment, context, retrieval, tool design, state, authority, handoff, interface or recovery?
Then test the smallest change that targets the actual failure mode. A model upgrade should be a hypothesis with an eval, not a ritual.
AI systems are increasingly dynamic at runtime. Enterprise evaluation should qualify the serving route, harness, tools and fallback conditions, not just the model name.
A polished AI demo proves that a system can succeed under prepared conditions. A buying decision needs different evidence: what happens when the system is wrong, blocked, uncertain or halfway through an action.
Agents are turning software capabilities into an interface of their own. The next enterprise design problem is deciding what should be callable, by whom, and under which boundaries.
RealityUseful context is selective. More tokens can bury relevant evidence, preserve stale assumptions and increase the amount of contradictory material the agent has to resolve.
RealityRAG retrieves evidence. Organizational memory also needs state, provenance, validity, decisions, procedures, ownership and a way to forget or revise what is no longer true.
RealityA human checkpoint is only a control if the person has the context, competence, time and authority to detect a problem and stop or reverse the action.