Once agents multiply across workflows, the management problem changes: which agents deserve more capacity, which should be merged or constrained, and which should be retired?
A few autonomous systems can be managed as separate projects. Once dozens of them compete for the same budget, infrastructure and management attention, that logic stops working. What looked like a list of projects becomes an allocation problem.
Customer support gets an agent. Finance gets two. Engineering starts using coding agents. Procurement automates a workflow. A vendor adds an embedded agent to software the company already pays for. A team builds a specialist internal agent because the general-purpose one is not reliable enough. Another team builds something similar because it does not know the first one exists.
Each decision can be individually reasonable.
The portfolio can still become irrational.
That is the shift I think many companies are about to feel. Once agents multiply across functions, the management problem is no longer simply Can we afford these model calls?
It becomes:
Which agents deserve more of the company's scarce money, capacity, attention and authority—and which do not?
That is a capital allocation question.
The agent estate becomes expensive long before anyone calls it a portfolio.
Traditional enterprise software gives finance a relatively familiar object.
A license has a price. A cloud workload has provisioned infrastructure. A project has a budget. Usage can vary, but the relationship between access and cost is usually visible enough to forecast.
Agents are less obedient.
One business request can trigger a chain of model calls, retrievals, tool invocations, retries, sub-agents, persistent memory, evaluations and human review. The same nominal task can consume radically different amounts of intelligence depending on path, model choice, context and failure.
McKinsey's 2026 AI FinOps work describes exactly this forecasting problem. In its analysis, token consumption for the same task can vary by as much as 30×. In the same survey program, only roughly 20–25% of companies were assessed as having mature AI FinOps capabilities. Those figures come from a small enterprise survey—120 participants, 75 qualified respondents—so I would not treat them as a census of enterprise AI. I would treat them as evidence of the mechanism: spend becomes harder to forecast as execution becomes more agentic.
OpenAI makes a related point from the supplier side. Its July guidance tells enterprise leaders not to optimize on token price alone, but on the full cost of reaching an accepted outcome: model and tool usage, attempts, completion rate, latency and human review.
This changes the unit of management.
The useful question is not:
How much does this agent cost?
It is:
What does one dependable outcome from this agent cost, and what is that outcome worth?
That is already a better FinOps question.
But a growing agent estate needs one more step.
It needs a portfolio question.
Imagine ten agents, all with positive business cases.
One resolves low-value support cases.
One accelerates financial close.
One helps engineers review code.
One researches procurement options.
One prepares sales account plans.
One monitors compliance exceptions.
If each agent is evaluated independently, the organization may reasonably conclude that all ten should scale.
But budgets, infrastructure, platform teams, security capacity, evaluation capacity and human attention are not infinite.
The agents also compete for something subtler: organizational change capacity.
Every production agent needs an owner. It needs data access, integration, monitoring, policy, evaluation, exception handling and a team willing to redesign work around it. An agent with positive ROI may still be the wrong place for the next dollar if another workflow creates more value per unit of scarce organizational capacity.
This is where project economics become portfolio economics.
OpenAI's own investment guidance now uses this framing explicitly. It recommends managing AI investment as a portfolio across broad employee access, function-specific workflows and a smaller set of strategic bets, with funding changing as work moves from exploration to validation to production.
The implication is bigger than the guidance sounds.
Not every successful agent should receive the same next investment.
As the estate grows, I would force every consequential agent into one of five actions:
Scale. Optimize. Merge. Constrain. Retire.
Not “keep” or “cancel.”
Those two choices are too crude for systems whose economics and capabilities keep changing.
Scale the agent when the outcome is valuable, repeatable and owned.
Evidence should survive beyond a demo: representative cases, stable quality, known failure modes, measurable outcome economics and a workflow that can absorb more volume.
Scaling can mean more users. It can also mean more authority, higher capacity, deeper integration or moving the same agent into adjacent workflows.
The important point is that scale is an allocation decision. It consumes more than inference.
Optimize when the agent creates real value but its economics are poor.
That may mean model routing, cheaper models for easier cases, prompt or context reduction, caching, fewer retries, better tools, lower supervision load, more selective autonomy, or redesigning the workflow itself.
This is where A cheaper AI model can move the cost instead of removing it matters. A lower token price can increase retries, latency or human correction. Optimization should reduce the cost per accepted outcome, not merely one line item.
Gartner's October 2 FinOps playbook for custom-built agents makes the same move: assess fit and unit economics before build, instrument total cost before deployment and govern runtime economics after launch.
Merge when two agents are solving materially overlapping problems and the duplication is no longer buying useful independence.
Agent sprawl will not always look like duplicate software. Two agents can have different names, owners and vendors while retrieving the same knowledge, calling the same tools and producing nearly the same decision.
The question is not whether they share code.
It is whether the company is funding the same capability twice.
IBM's recent work on third-party agent governance describes enterprises moving toward mixed ecosystems of internal, embedded and marketplace agents. That makes duplication harder to see because part of the estate is built, part is bought and part arrives inside existing applications.
A portfolio view makes overlap visible.
Constrain when the agent is valuable but the marginal unit of autonomy, capacity or scope is no longer worth its risk or cost.
This can mean lowering budget ceilings, reducing model tier, narrowing tool access, decreasing run frequency, shrinking context, requiring approval for specific actions, or removing a workflow from the autonomous path.
Constrain is important because the alternative is often politically difficult: teams defend an agent as “working,” so nothing changes.
But working is not the same as deserving unlimited growth.
Task-scoped authority is one form of constraint. Runtime budget is another.
Retire when the agent no longer justifies its place in the portfolio.
This sounds obvious. It is probably the least developed muscle.
Software portfolios accumulate because removing something has no launch moment. Agents may accumulate faster because their economics can improve and deteriorate continuously as model prices, capabilities, workflows and alternatives change.
An agent that was the right answer six months ago can become redundant because:
Retirement is not an admission that the original investment was wrong.
It is evidence that the portfolio is being managed.
Positive ROI is not a permanent entitlement
An agent can still create value and deserve retirement if another use of the same budget, capacity or organizational attention creates materially more value.
Finance will naturally start with money.
That is correct.
Agent portfolios should expose spend by agent, workflow, team, model, tool and outcome. Without that visibility, the organization cannot distinguish a valuable expensive agent from a cheap agent nobody should be running.
McKinsey argues for a centralized AI control plane that joins spend, usage, performance and business outcomes. OpenAI similarly recommends visibility into who is using which models, how much capacity they consume and what work that consumption supports.
But cash is not the only scarce resource.
A serious agent estate also consumes:
The mature portfolio decision therefore needs two ledgers.
One for money.
One for organizational capacity.
An agent that is cheap in dollars and expensive in exception handling may be a poor investment.
An agent that costs more but replaces a coordination-heavy process may be an excellent one.
That is why cost dashboards alone are not enough.
Capital portfolios worry about concentration because one position can become too important.
Agent estates will have a version of the same problem.
The concentration may be in a model provider, an orchestration platform, a connector, a data source, a shared agent, a memory layer or a small internal team that knows how the whole system works.
Suppose fifteen business-critical agents depend on one retrieval service.
Their individual ROI models may look independent.
Operationally, they share a single point of failure.
Or imagine five agents appear to create value in separate functions but all rely on one frontier model whose pricing or availability changes.
The portfolio is more concentrated than the org chart suggests.
This is another reason to manage the estate above the level of the individual agent.
Local diversity can hide infrastructural concentration.
A portfolio view should therefore track not only agent count and spend, but shared dependencies.
Most software inventories ask:
Who owns this?
What does it cost?
Is it used?
Is it compliant?
An agent portfolio review needs those questions and a more demanding set.
For each agent, I would want:
| Question | What it tells you |
|---|---|
| What business outcome does this agent own? | Whether the unit of value is explicit |
| What is cost per accepted outcome? | Whether spend is economically meaningful |
| How much human review and exception work does it create? | Whether automation is moving cost elsewhere |
| What happens when volume doubles? | Whether economics scale |
| Which other agents overlap with it? | Whether duplication is funding hidden redundancy |
| Which shared dependencies does it rely on? | Whether the estate has concentration risk |
| What evidence would justify more authority or capacity? | Whether scale is evidence-driven |
| What evidence would trigger constraint or retirement? | Whether exit is designed in advance |
That final question matters.
Companies are reasonably good at defining entry criteria for new projects.
They are much worse at defining exit criteria for successful ones.
Agent portfolios need both.
“Use case” was a useful unit when organizations were trying to discover where AI might help.
As agent estates mature, use cases become too easy to accumulate.
A capital case is stricter.
It asks:
That is the difference between having many AI projects and managing AI as an economic system.
OpenAI's portfolio guidance points toward this explicitly: funding should follow maturity, shared capabilities should be funded centrally, and leaders should scale proven demand rather than making each workflow rebuild infrastructure.
The important word there is proven.
Agentic AI makes it technically possible to create more autonomous work.
Capital allocation decides which autonomous work the company should keep paying for.
Most AI dashboards will eventually contain familiar metrics:
usage, completion rate, latency, accuracy, token cost, human review, business outcome.
I would add one uncomfortable portfolio metric:
What percentage of the agent estate did we merge, constrain or retire this quarter?
Not because higher is always better.
Because zero is suspicious.
In a fast-moving technology environment, a portfolio in which every agent survives and every successful pilot keeps expanding is probably not being actively allocated.
It is accumulating.
The organization should expect some agents to graduate.
Some to become infrastructure.
Some to merge.
Some to lose budget.
Some to disappear.
That is not instability.
It is management.
The first few agents are technology choices.
A mature estate is a collection of recurring claims on company resources.
Every agent is effectively saying:
Give me model capacity.
Give me access.
Give me data.
Give me integration.
Give me evaluation.
Give me people to supervise exceptions.
Give me budget next quarter.
At small scale, teams can answer those requests locally.
At portfolio scale, local answers add up to company strategy.
That is the deeper reason the agent estate becomes a capital allocation problem.
The question is no longer simply whether an agent works.
It is:
Compared with every other place we could put the next unit of intelligence, authority and organizational attention, does this agent still deserve it?
One email per essay. No noise between.
Berk Bayri
AI Transformation Advisor & Fractional Innovation Lead
I help leadership teams decide where AI belongs, test it before it scales and build the teams that run it, drawing on 25+ years of building digital products, experiences and capabilities.