Berk Bayri
AI Strategy8 min read·

Your chatbot is borrowing from the next interaction

AI customer service is usually measured one interaction at a time. But a failed automated interaction can change which channel a customer chooses next time. That makes future adoption part of the economics, not a separate trust metric.

A chatbot can save money on this interaction and still make the service operation more expensive.

That sounds contradictory only if AI customer service ROI is measured one contact at a time.

Most dashboards do exactly that. Did the bot resolve the issue? Did it contain the conversation? Did handle time fall? Did the customer escalate? What did this contact cost compared with a human-assisted one?

All useful questions. They just stop at the end of the session.

A Gartner survey published on September 2 gives us a reason to look further. Only 27% of customers said they would be willing to try a chatbot again after a negative experience. The survey covered 3,566 B2B and B2C customers and was conducted in February and March 2026.

The obvious reading is that bad chatbots damage trust.

The more useful reading is economic: a failed automated interaction can change the channel a customer chooses next time.

If that is true, part of the cost of failure arrives in the future.

We measure the session. The customer remembers the channel.

Gartner found that 49% of customers said they would have been willing to use a chatbot if a company had provided one. Only 7% actually used a chatbot or digital assistant in their most recent service interaction. Customers were also approximately three times more likely to use a third-party GenAI tool than a company-provided chatbot during that interaction.

Those figures describe a gap between theoretical willingness and actual behavior.

Past experience is one plausible reason for the gap, and Gartner says exactly that: if chatbots have previously misunderstood the problem, returned generic information or made it difficult to reach a person, customers are more likely to choose another channel.

Think about what that means operationally.

A bad chatbot interaction is not necessarily recorded as anything dramatic. Perhaps the bot fails, the customer escalates and a human eventually resolves the issue. The company sees one failed automated session and one assisted contact.

The customer may leave with a simpler conclusion:

I won't start there next time.

The next time they call. Or open a ticket. Or search the web. Or ask a general-purpose AI assistant before returning to the company.

None of that future behavior appears in the original chatbot session's cost.

It should.

Containment can look good while the channel gets weaker

Containment rate tells you whether an interaction stayed inside automation without moving to a human.

It is useful. It is also very easy to overvalue.

A customer can spend seven frustrating minutes inside a bot, eventually get enough information to leave, and never willingly use that bot again. The session was contained. The channel lost a future user.

The reverse can happen too.

A bot can recognize early that it is the wrong tool for this issue, pass the context to a person and make the human conversation faster. The interaction was not contained, but the customer may be perfectly willing to start with the bot again next time.

Which outcome created more value?

Containment cannot tell you.

I would separate the economics into two horizons.

Session economics: what did this contact cost, and what did it resolve?

Repeat-channel economics: what did this contact do to the probability that the customer will choose automation again?

The first number is easy to see.

The second is where compounding happens.

Sometimes the cheaper move is to give up earlier

The same Gartner survey found that 87% of customers consider access to a human agent essential when companies use GenAI in customer service.

That does not necessarily mean customers reject automation. It may mean they want confidence that automation knows when it has reached its limit.

This is where an obsession with reducing escalation can become expensive.

If the system is uncertain but keeps asking another question, presenting another weak answer or sending the customer through another loop, it may save a transfer for thirty seconds while teaching the customer to avoid the automated channel for months.

In that situation, escalation is not failure.

It is channel preservation.

The AI is protecting the customer's willingness to use AI again.

Gartner's advice is unusually practical here: prioritize reliability over reach, start with issue types that have a proven high resolution rate, and give customers a visible path to a human when confidence is low. It also recommends carrying the information already collected into the handoff.

That last detail matters. "You can talk to a person" is not a graceful recovery if the customer has to explain everything again.

Expanding scope spends trust as well as engineering capacity

On the same day as Gartner's survey, Genesys described its direction for agentic customer experience: virtual agents that move beyond routing and simple answers, understand goals, use tools and complete work, with an emphasis on reliability at scale.

That is vendor positioning, not independent evidence. But it points to a real product decision facing almost every company deploying service AI.

As the agent becomes more capable, what should it be allowed to handle?

The answer cannot be "everything it can technically do."

A system that gets store hours wrong is annoying. A system that mishandles a billing dispute, changes an account incorrectly or misunderstands an insurance claim has a very different failure cost.

So every new issue type added to an AI service channel has at least five economic variables:

  • value if successfully automated;
  • probability of successful resolution;
  • consequence of failure;
  • quality of recovery;
  • effect on future channel choice.

Those variables can point in different directions.

A high-volume issue may look attractive to automate but be costly to get wrong. A complex issue may make an impressive demo and still be a terrible early deployment if one failure causes customers to abandon the channel.

The useful question is not simply, How much more can the agent handle?

It is:

What can it handle without making customers less willing to use it again?

That is a much better boundary for scope.

Configuration is not experience

There is another reason to measure customer behavior rather than system intent.

Twilio published research on August 31 showing a striking perception gap: 81% of brands said their AI agents identify themselves as AI immediately, while only 22% of consumers said that matched their experience.

This is vendor research and should be treated as such. I am less interested in turning those percentages into a universal benchmark than in the failure mode they expose.

The company can have the right rule and still deliver the wrong experience.

The system prompt says the bot should disclose itself. The design document says it does. The product team can point to the exact configuration.

The customer misses the disclosure because it appeared in a welcome message, happened only on one channel or was expressed in language that did not feel clear.

The same gap can exist around escalation.

A human path technically exists. The customer experiences three loops and a guessing game before finding it.

For production AI, configuration is evidence of intent.

Customer behavior is evidence of delivery.

Add one metric: return-to-automation rate

I would keep the normal customer-service measures: resolution, first-contact resolution, handling time, cost per resolution, escalation and satisfaction.

Then I would add one metric that forces the business case to remember what happened after the session ended:

Return-to-automation rate.

Among customers who had an eligible service interaction through AI, what percentage choose the automated channel again when they next have an eligible issue?

Then segment the result.

What happens after a successful automated resolution?

After escalation?

After abandonment?

After the customer corrects the bot?

After the bot transfers the context cleanly?

After the customer has to repeat the entire story to a person?

Now the organization can distinguish a good escalation from a bad one.

A quick transfer that preserves context may have a higher future value than a technically "successful" contained session that causes the customer to avoid automation next time.

That changes how teams optimize the system. The target stops being maximum containment and becomes durable voluntary use.

The ROI model needs a memory

A basic automation model might look like this:

automation savings = automated contacts × cost difference per contact

There is nothing wrong with that calculation. It is just incomplete if customer behavior is path-dependent.

Add another term:

future channel value = probability of choosing automation again × expected future eligible contacts

Now failure has a carry-forward cost.

So does recovery.

So does a clean human handoff.

So does pushing the agent into issue types where it is not reliable enough yet.

This does not argue for moving slowly. It argues for measuring the thing you actually want: repeated, useful adoption.

That is why I would treat scope expansion as a pilot decision rather than just a feature-release decision. A pilot should reduce uncertainty around a real decision. For customer-service AI, one of those uncertainties is no longer only:

Can the agent resolve this issue?

It is also:

If it cannot, will the customer still come back?

An AI service channel does not start from zero every morning.

It inherits yesterday's experience.

And a bad interaction can spend some of tomorrow's adoption before tomorrow arrives.


Sources

Berk Bayri

Creative Technology & Innovation Leader

Designing and building for digital environments since 1998, across strategy, product, design, technology and organizational innovation.

About Berk →

Get new essays in your inbox.

New essays by email, when they are published.