More context makes an agent smarter: a common AI agent myth
RealityUseful context is selective. More tokens can bury relevant evidence, preserve stale assumptions and increase the amount of contradictory material the agent has to resolve.
A common misconception about AI products, tested against the evidence.
RAG solves a valuable problem: retrieve relevant material and ground a model in evidence outside its parameters.
That is not the same as memory.
A company’s memory includes facts, but also the state of ongoing work, decisions that changed policy, procedures that supersede older procedures, exceptions, failed attempts, ownership, provenance and the reason something was decided. Some of that information should expire. Some should never be treated as authoritative just because it is semantically similar to the current question.
Microsoft Research’s 2026 work on long-horizon agents makes the distinction explicit. It argues that memory systems organized mainly around semantic similarity can fragment decision trajectories and mix valid and erroneous traces. Microsoft’s production memory work also separates factual memory from procedural memory: an agent can “know” the right facts and still fail because it does not execute the right procedure.
Retrieval answers “What looks relevant?” Memory also has to answer “What is still true, what happened, and what state are we in now?”
RAG often feels like memory in a demo. Ask about an old document and the system finds it. That is already far better than a model with no access to company knowledge.
The illusion breaks when work spans time and decisions.
Production-oriented memory systems are adding state management, procedural memory, validation, expiry and environment checks precisely because semantic retrieval alone is not enough. The harder problem is not storing more text. It is deciding what should persist and under what authority.
Before calling a system “company memory,” ask what it remembers:
facts, decisions, procedures, execution state, rationale, outcomes or all of them?
Then define how each memory type is created, validated, superseded, expired and audited. RAG can be part of that architecture. It is not the architecture by itself.
AI data agents do more than answer business questions. At scale, the pattern of questions, corrections and dead ends becomes a new signal: a map of what the organization is trying to understand, where its definitions are weak, and which decisions deserve better infrastructure.
OpenAI's new misalignment disclosure framework exposes a useful enterprise design principle: record anomalous AI behavior before the organization has finished explaining it. Otherwise incident systems quietly become filters for what teams already understand.
RealityUseful context is selective. More tokens can bury relevant evidence, preserve stale assumptions and increase the amount of contradictory material the agent has to resolve.
RealityA stronger model can improve model-level performance, but product failures often live in context, retrieval, tools, workflow logic, permissions, handoffs, state and recovery.
RealityLiteracy helps individuals use AI. Organizational capability also requires workflows, data, evaluation, ownership, decision rights, integration, incentives and the ability to repeat results without the original champions.