A lower model bill can hide a higher workflow bill. The useful AI TCO question is not only what got cheaper, but where the cost moved.
A model can become cheaper while the company spends more.
That sounds contradictory only if the model invoice is treated as the boundary of the AI system.
In a production workflow, it rarely is.
A cheaper model may need more retries. Its outputs may require more human review. More cases may fall into exception handling. A failure may take longer to recover from. Employees may start verifying work that they previously accepted. Operations, risk or a business team can absorb the difference without any of it returning to the line item that made the model look economical.
The saving is real inside one budget.
The company can still lose money.
This matters now because the cost conversation around enterprise AI is becoming more serious. EY's September 21 analysis of agentic AI estimates that the full enterprise cost can extend far beyond the token invoice once infrastructure, software, governance, organizational change and expected failure are included. IBM's current AI cost-management guidance makes a similar shift from raw consumption toward total cost of ownership and business outcomes.
That direction is right. It also exposes a harder accounting problem.
When one part of an AI workflow gets cheaper, where did the cost go?
There is a good reason the industry is moving away from cost per token.
McKinsey argues that the unit of governance should be the completed business outcome rather than token cost. BCG recommends optimizing for cost per successful outcome at the required quality, latency and risk, including the burden of human review. Other practitioners have gone further and used variants such as cost per accepted outcome.
So I would not invent another name for the denominator.
The useful next step is to trace the movement underneath it.
Suppose a team changes model routing and reduces inference spend by 30%.
The finance slide says the optimization worked.
Before accepting that conclusion, I would want the deltas around the workflow:
This is a cost-migration test.
It asks whether an optimization removed cost from the enterprise or merely transferred it to a place where the AI budget no longer sees it.
Enterprise cost ownership is fragmented almost by design.
Engineering sees inference and infrastructure. Operations sees exceptions. Risk sees controls and review. Business teams absorb rework. Finance eventually receives several numbers that were measured for different reasons.
An AI system crosses all of those boundaries.
That creates an easy failure mode: optimize a visible technical metric and externalize the consequence.
Imagine two models performing the same workflow.
Model A costs less per run. Model B costs more.
Model A also creates slightly more ambiguous outputs. Most are not catastrophic failures; they are simply less ready to use. A domain expert spends another two minutes on some cases. A few more cases escalate. Occasionally the workflow needs to be replayed.
Model B has the higher inference bill but produces fewer of those downstream costs.
There is no universal answer about which model is cheaper. The answer depends on the frequency and cost of the downstream work.
But there is a universal accounting rule:
Do not credit a local saving before checking the costs it displaced.
That is especially important for AI because many displaced costs are paid in human attention rather than compute.
Human review is easy to underprice.
The employee already has a salary, so another minute of checking can disappear into normal work. Ten seconds here, a correction there, an escalation that requires reconstructing context. None looks large enough to deserve its own business case.
At scale, that is exactly why it matters.
I have argued elsewhere that the hidden metric in AI automation is supervision. A system that appears highly automated can still consume substantial human capacity if people must watch, verify, correct or recover its work.
The cost-migration test makes that problem financial.
If a model change saves $10,000 in inference and creates $14,000 of additional review and exception work, the model optimization succeeded and the enterprise optimization failed.
The numbers will rarely be that clean. Human attention is harder to attribute than an API invoice.
That is not a reason to omit it. It is a reason to improve the measurement.
I would not try to solve this with one grand TCO spreadsheet updated once a year.
The more useful artifact is smaller: a cost-migration ledger attached to meaningful changes in the AI system.
Change the model. Change routing. Add a tool. Increase autonomy. Reduce review. Introduce caching. Alter an approval rule.
For each material change, record the before-and-after delta across the parts of the workflow it can plausibly affect.
Use the same four fields for every dimension: Before · After · Delta · Owner.
The purpose is not accounting perfection. It is to prevent one team from claiming a saving that another team is quietly financing.
This also changes how model routing should be evaluated.
The cheapest capable model may still be the right default. BCG and McKinsey both make a strong case for routing work according to cost, quality and task requirements rather than defaulting to premium models.
But "capable" has to be observed at the workflow boundary.
If a cheaper route increases the amount of work needed after the model returns, the routing policy should see that consequence too.
There is a temptation to treat AI FinOps as a more complicated version of cloud FinOps: observe consumption, attribute spend, reduce waste, negotiate better rates.
Those capabilities matter. AI adds another difficulty.
The consumption event and the economic consequence can occur in different places.
A token is consumed in one system. The resulting ambiguity is handled by a person somewhere else. A failed tool call is retried automatically. A weak answer is corrected without being logged as a failure. An employee loses trust and begins duplicating the old process.
The technical telemetry can be complete while the economics remain incomplete.
This is also why production evaluation needs a runtime address. As I argued in The benchmark needs a runtime address, the thing being qualified is not just a model name. Routing, tools, policies and operating conditions shape the result.
They shape the cost as well.
If those components change, both performance evidence and economic evidence can decay.
EY's recent analysis is a useful signal because it widens the enterprise cost boundary. The token is not the bill; it is one line on it.
The operational consequence is more specific.
When an AI team reports a cost improvement, follow the delta until it stops moving.
Did compute fall? Good. Did retries rise? Did review fall? Did recovery rise? Did throughput change? Did another team inherit work? Did the business get more usable outcomes for less total effort?
If the answer survives those questions, the saving is probably real.
If it disappears when the accounting crosses into another team's budget, the optimization did not remove cost.
It found somewhere else to put it.
One email per essay. No noise between.
Berk Bayri
Creative Technology & Innovation Leader
Designing and building for digital environments since 1998, across strategy, product, design, technology and organizational innovation.
About Berk →