Berk Bayri
AI Strategy & Transformation7 min read·

A cheaper AI model can move the cost instead of removing it

A lower model bill can hide a higher workflow bill. The useful AI TCO question is not only what got cheaper, but where the cost moved.

A model can become cheaper while the company spends more.

That sounds contradictory only if the model invoice is treated as the boundary of the AI system.

In a production workflow, it rarely is.

A cheaper model may need more retries. Its outputs may require more human review. More cases may fall into exception handling. A failure may take longer to recover from. Employees may start verifying work that they previously accepted. Operations, risk or a business team can absorb the difference without any of it returning to the line item that made the model look economical.

The saving is real inside one budget.

The company can still lose money.

This matters now because the cost conversation around enterprise AI is becoming more serious. EY's September 21 analysis of agentic AI estimates that the full enterprise cost can extend far beyond the token invoice once infrastructure, software, governance, organizational change and expected failure are included. IBM's current AI cost-management guidance makes a similar shift from raw consumption toward total cost of ownership and business outcomes.

That direction is right. It also exposes a harder accounting problem.

When one part of an AI workflow gets cheaper, where did the cost go?

Outcome economics is necessary, but it is no longer a novel idea

There is a good reason the industry is moving away from cost per token.

McKinsey argues that the unit of governance should be the completed business outcome rather than token cost. BCG recommends optimizing for cost per successful outcome at the required quality, latency and risk, including the burden of human review. Other practitioners have gone further and used variants such as cost per accepted outcome.

So I would not invent another name for the denominator.

The useful next step is to trace the movement underneath it.

Suppose a team changes model routing and reduces inference spend by 30%.

The finance slide says the optimization worked.

Before accepting that conclusion, I would want the deltas around the workflow:

  • Did retry volume change?
  • Did first-pass acceptance change?
  • Did human review minutes change?
  • Did escalation volume change?
  • Did recovery work change?
  • Did employees alter their behavior because they trusted the output more or less?
  • Did the number of completed, usable outcomes change?

This is a cost-migration test.

It asks whether an optimization removed cost from the enterprise or merely transferred it to a place where the AI budget no longer sees it.

The budget boundary can hide the system boundary

Enterprise cost ownership is fragmented almost by design.

Engineering sees inference and infrastructure. Operations sees exceptions. Risk sees controls and review. Business teams absorb rework. Finance eventually receives several numbers that were measured for different reasons.

An AI system crosses all of those boundaries.

That creates an easy failure mode: optimize a visible technical metric and externalize the consequence.

Imagine two models performing the same workflow.

Model A costs less per run. Model B costs more.

Model A also creates slightly more ambiguous outputs. Most are not catastrophic failures; they are simply less ready to use. A domain expert spends another two minutes on some cases. A few more cases escalate. Occasionally the workflow needs to be replayed.

Model B has the higher inference bill but produces fewer of those downstream costs.

There is no universal answer about which model is cheaper. The answer depends on the frequency and cost of the downstream work.

But there is a universal accounting rule:

Do not credit a local saving before checking the costs it displaced.

That is especially important for AI because many displaced costs are paid in human attention rather than compute.

Human attention is where cheap AI can become expensive

Human review is easy to underprice.

The employee already has a salary, so another minute of checking can disappear into normal work. Ten seconds here, a correction there, an escalation that requires reconstructing context. None looks large enough to deserve its own business case.

At scale, that is exactly why it matters.

I have argued elsewhere that the hidden metric in AI automation is supervision. A system that appears highly automated can still consume substantial human capacity if people must watch, verify, correct or recover its work.

The cost-migration test makes that problem financial.

If a model change saves $10,000 in inference and creates $14,000 of additional review and exception work, the model optimization succeeded and the enterprise optimization failed.

The numbers will rarely be that clean. Human attention is harder to attribute than an API invoice.

That is not a reason to omit it. It is a reason to improve the measurement.

Keep a ledger of deltas, not a single AI number

I would not try to solve this with one grand TCO spreadsheet updated once a year.

The more useful artifact is smaller: a cost-migration ledger attached to meaningful changes in the AI system.

Change the model. Change routing. Add a tool. Increase autonomy. Reduce review. Introduce caching. Alter an approval rule.

For each material change, record the before-and-after delta across the parts of the workflow it can plausibly affect.

Use the same four fields for every dimension: Before · After · Delta · Owner.

  • Model / inference cost — Owner: AI / Engineering
  • Retry and tool-call cost — Owner: AI / Engineering
  • Human review time — Owner: Business / Operations
  • Exceptions and escalations — Owner: Operations
  • Failure / recovery work — Owner: Operations / Engineering
  • Governance / assurance effort — Owner: Risk / Governance
  • Accepted or completed outcomes — Owner: Business
  • Cycle time / service level — Owner: Business / Operations

The purpose is not accounting perfection. It is to prevent one team from claiming a saving that another team is quietly financing.

This also changes how model routing should be evaluated.

The cheapest capable model may still be the right default. BCG and McKinsey both make a strong case for routing work according to cost, quality and task requirements rather than defaulting to premium models.

But "capable" has to be observed at the workflow boundary.

If a cheaper route increases the amount of work needed after the model returns, the routing policy should see that consequence too.

AI TCO should be attributable before it is optimizable

There is a temptation to treat AI FinOps as a more complicated version of cloud FinOps: observe consumption, attribute spend, reduce waste, negotiate better rates.

Those capabilities matter. AI adds another difficulty.

The consumption event and the economic consequence can occur in different places.

A token is consumed in one system. The resulting ambiguity is handled by a person somewhere else. A failed tool call is retried automatically. A weak answer is corrected without being logged as a failure. An employee loses trust and begins duplicating the old process.

The technical telemetry can be complete while the economics remain incomplete.

This is also why production evaluation needs a runtime address. As I argued in The benchmark needs a runtime address, the thing being qualified is not just a model name. Routing, tools, policies and operating conditions shape the result.

They shape the cost as well.

If those components change, both performance evidence and economic evidence can decay.

A saving should survive the handoff between budgets

EY's recent analysis is a useful signal because it widens the enterprise cost boundary. The token is not the bill; it is one line on it.

The operational consequence is more specific.

When an AI team reports a cost improvement, follow the delta until it stops moving.

Did compute fall? Good. Did retries rise? Did review fall? Did recovery rise? Did throughput change? Did another team inherit work? Did the business get more usable outcomes for less total effort?

If the answer survives those questions, the saving is probably real.

If it disappears when the accounting crosses into another team's budget, the optimization did not remove cost.

It found somewhere else to put it.


Sources

Get new essays as they publish.

One email per essay. No noise between.

Berk Bayri

Creative Technology & Innovation Leader

Designing and building for digital environments since 1998, across strategy, product, design, technology and organizational innovation.

About Berk →