As AI moves from assisting work to leading it, hours saved stop telling the whole story. The scarce resource shifts to human supervision: approvals, exceptions, context and judgment.
The most interesting number in Anthropic's new R&D Automation Index is not 26%.
It is the definition underneath it.
Anthropic says that, as of August 2026, Claude "leads" 26% of its measured AI research and development work. On the scale it uses, leads means the AI can complete most of a task end-to-end from a high-level prompt while a human supervises. More than 90% of the measured work is at least at the level where AI collaborates, while none of the measured areas is fully autonomous.
That is a striking description of advanced AI work. It is also a useful warning about how companies measure automation.
We usually count the part that disappeared from the employee's hands: minutes saved, tickets handled, documents drafted, tasks automated. But as AI moves from helping with a task to carrying a longer stretch of it, the human work does not necessarily disappear with the execution.
Some of it moves.
It becomes supervision.
Supervision sounds lightweight until you list what it contains.
A person may no longer write the first draft, run the first analysis or execute every step in a workflow. They may instead define the goal, supply missing context, decide what the system is allowed to do, inspect uncertain output, approve consequential actions, resolve exceptions, recover failures and remain accountable for the result.
That can still be a much better operating model. The mistake is counting all of that as zero because the AI performed the visible task.
Glean's 2026 Work AI Index gives a workplace-scale view of the problem. In its survey of 6,000 full-time digital workers in the US, UK and Australia, respondents said AI automation saved roughly 11 hours a week. The same research found an average of 6.4 hours a week spent on what Glean calls "botsitting": supplying context, supervising output, debugging mistakes and cleaning up AI-generated work.
Those figures are self-reported and come from vendor-sponsored research, so they should not be treated as a universal productivity ratio. The shape of the finding is more useful than the headline arithmetic.
AI can make production cheaper while creating a second workload around making that production dependable.
The closer an agent gets to completing work end-to-end, the more important that second workload becomes to the economics.
Suppose an analyst used to spend two hours preparing a report.
An agent now produces it in ten minutes.
"One hour and fifty minutes saved" sounds like the obvious result. It may even be correct.
But imagine the analyst now spends twenty minutes checking sources, ten minutes correcting a classification the agent regularly gets wrong, five minutes adding context that lives outside the connected systems and another ten minutes deciding whether an unusual case can be released.
The automation is still valuable. It is simply less valuable than the production-time metric suggests.
Now imagine the same system improves. It handles the classification correctly, retrieves the missing context itself and escalates only the unusual case with enough evidence for the analyst to make a decision in two minutes.
The visible output has barely changed. The operating model has.
This is why I would measure supervision load per completed unit of work alongside automation rate.
Not because every minute of human attention is bad. Some supervision is the point. A financial approval, a safety decision or a consequential customer action may deserve a human decision even when the system could technically continue without one.
The question is whether the human attention is deliberate or compensatory.
Deliberate supervision protects a boundary.
Compensatory supervision repairs a system that cannot yet cross it reliably.
Those two forms of human involvement should not be sitting in the same bucket.
"Human in the loop" has become reassuring language in AI projects. It can describe excellent system design. It can also conceal a great deal of manual labor.
There is a difference between a human staying in the loop and a human becoming the loop.
In the first case, the workflow is designed to run reliably until it reaches a decision where human authority or judgment is intentionally required. The escalation arrives with the relevant context. The person knows what they are deciding. Their intervention is narrow enough that they can actually pay attention.
In the second case, the workflow works because a person constantly watches it.
They reload context the system should already know. They compare answers because reliability is unclear. They catch malformed output before it reaches another system. They decide which failures matter because the escalation rules are weak. They translate between the AI and the organization.
Both architectures can be described as supervised AI.
Only one is moving toward scalable capability.
This distinction matters especially for pilots. A pilot can look remarkably efficient when the small team running it knows every weakness and quietly compensates for them. The problem appears when the system scales and that invisible expertise does not scale with it.
That is one reason a pilot should be treated as a decision instrument, not as proof that a demonstration can be copied into production unchanged. A useful pilot should expose the human work required to keep the system dependable, not hide it.
Anthropic's disclosure makes the shift unusually visible because it separates levels of automation rather than treating AI use as binary.
At its "AI collaborates" level, the model can do large chunks of work under close human direction. At "AI leads", it can complete most of a task from a high-level prompt while the human supervises.
That transition is not just more automation. It changes what the human is for.
When execution is scarce, organizations optimize execution.
When execution becomes abundant, they have to optimize attention.
Which outputs deserve review? Which actions require approval before execution? Which exceptions can wait? Which need immediate escalation? What evidence should arrive with the escalation? Who has authority to decide? How many simultaneous automated workflows can one person meaningfully supervise before "human oversight" becomes a ceremonial phrase?
Anthropic is dealing with an extreme version of this problem. It reports roughly 30,000 research and engineering agents active at any one time on its most-used internal platform in August. According to the company, all actions on that platform passed through online monitoring before execution and were also ingested for offline monitoring. Across more than a billion agent decisions during the month, its online monitor blocked about 0.002%, while offline systems narrowed roughly 100,000 flagged transcripts per week down to approximately 50 highest-priority cases for human review.
Those are Anthropic's own measurements of its own systems, not an enterprise benchmark. But the architecture contains a general lesson.
You cannot supervise machine-scale work by asking humans to inspect machine-scale output.
The supervision system itself has to decide where scarce human attention goes.
Once that is visible, AI ROI becomes a little harder to fake.
For any workflow moving toward agentic automation, I would want to know at least four things about the human remainder.
How much attention does one dependable outcome require? Not how long the agent runs. How much human time is consumed by prompting, context supply, review, approval, correction and exception handling before the outcome is actually usable?
Why is the human involved? Separate mandatory judgment from preventable repair. Regulatory approval and fixing a recurring hallucination are both human minutes, but they imply very different investment decisions.
How concentrated is the attention? Ten minutes of predictable review at 4 p.m. is not operationally equivalent to ten interruptions scattered through the day. Supervision load has a scheduling cost as well as a duration.
What happens as volume rises? If completed work doubles, does human attention double with it? If so, the system may be accelerating production without creating much operating leverage.
That last question is the one I would put into every serious scaling decision.
A system that processes ten times the work but requires ten times the supervisory attention has achieved throughput. It has not necessarily created a scalable capability.
A system that processes ten times the work while human attention rises only modestly has changed the economics.
There is another complication: supervision load is not a permanent property of "the model."
Change the model, tools, permissions, retrieval layer, prompts, workflow or escalation rules and the amount of human attention required can change with them.
That makes supervision a runtime measure.
The same principle applies to evaluation. In The benchmark needs a runtime address, I argued that enterprise evidence should identify the serving route, harness, tools and fallback conditions that actually passed evaluation. A supervision metric needs the same discipline.
"Agent X needs five minutes of review per case" is weak evidence if Agent X is a label covering a changing system.
The useful statement is closer to:
Under this workflow, with this model and tool configuration, these permissions, these escalation rules and this class of work, dependable completion currently requires this much human attention.
That is less convenient than an automation percentage.
It is much more useful for deciding whether to scale.
Companies are going to automate larger pieces of work.
The visible task count will rise quickly. So will claims about hours saved.
The more capable the systems become, the less those numbers tell us on their own.
A mature automation program should be able to show not only what the AI now does, but what kind of human work remains around it and why. Some of that remainder should be protected because it represents judgment, authority or accountability. Some should be engineered away because it represents missing context, unreliable behavior or poor workflow design.
The objective is not zero human involvement.
It is to stop confusing relocated human work with eliminated human work.
If automation rate tells you how much execution moved to the machine, supervision load tells you whether the organization actually gained leverage.
Berk Bayri
Creative Technology & Innovation Leader
Designing and building for digital environments since 1998, across strategy, product, design, technology and organizational innovation.
About Berk →New essays by email, when they are published.