Agents remove work: a common AI agent myth
RealityAgents can remove execution work, but they also create supervision, exception handling, evaluation, recovery, permission and maintenance work. The net matters.
A common misconception about AI agents, tested against the evidence.
Multi-agent systems have an intuitive appeal: one agent proposes, another critiques, a third coordinates. It sounds like a good team.
Sometimes it is.
Research on test-time reasoning has found cases where multi-agent approaches outperform simpler strategies at comparable compute budgets. But recent work on agents acting for different users over shared resources shows the opposite pattern: adding agents can create stalling, conflicting actions and coordination overhead large enough to make the group worse than a single coordinator.
Those results are not contradictory. They describe different task structures.
More agents increase the amount of intelligence in the room and the amount of coordination the room requires.
Human organizations often improve difficult work by adding specialization and peer review. It is natural to map that idea onto agents.
But software agents do not inherit effective organizational design automatically. They need protocols for authority, state, communication, conflict and termination.
Multi-agent gains are most plausible when work can be cleanly decomposed, parallelized or independently checked. They become less reliable when agents compete over shared state, serve conflicting goals or repeatedly rewrite one another's work.
Anthropic's research on emerging multi-agent systems treats coordination itself as a major unresolved problem. The recent MAMUBench work shows that communication channels help but do not eliminate the gap.
Before adding an agent, ask: What unique work does this agent perform that cannot be done more reliably by a tool, a deterministic check or the existing agent?
If the answer is unclear, the extra agent may be architecture theatre.
Giving AI more autonomy does not remove organizational complexity. It gives that complexity permission to act.
As AI moves from assisting work to leading it, hours saved stop telling the whole story. The scarce resource shifts to human supervision: approvals, exceptions, context and judgment.
AI systems are increasingly dynamic at runtime. Enterprise evaluation should qualify the serving route, harness, tools and fallback conditions, not just the model name.
RealityAgents can remove execution work, but they also create supervision, exception handling, evaluation, recovery, permission and maintenance work. The net matters.
RealityAutonomy is not a quality score. The right level depends on consequence, reversibility, uncertainty and authority. A good agent knows what it can finish and what it should stop.
RealityA stronger model can improve model-level performance, but product failures often live in context, retrieval, tools, workflow logic, permissions, handoffs, state and recovery.