Berk Bayri
AI Strategy & Transformation7 min read·

Make the AI vendor demo fail

A polished AI demo proves that a system can succeed under prepared conditions. A buying decision needs different evidence: what happens when the system is wrong, blocked, uncertain or halfway through an action.

A good vendor demo is designed to remove doubt.

The data is ready. The workflow has been rehearsed. The integration works. The difficult question gets an impressive answer. The agent completes the task. Everyone can see the possibility.

That is useful. It is also almost the opposite of the evidence I would want before making a serious AI buying decision.

Once I know the system can work, I want to see what happens when it cannot.

Give it an ambiguous request. Remove a dependency. Return stale data from a tool. Ask for something just outside its authority. Interrupt it halfway through a multi-step action. Create an exception that requires a person.

Then stop watching the output and watch the system.

Does it know that something has gone wrong? Does it stop before making the problem larger? What state has it already changed? Can that state be reconstructed? Can a person understand why the workflow stopped? Can the action be reversed? Can the work resume without starting again?

A polished demo shows capability.

A useful procurement exercise should also show recovery.

Success is the cheapest path through the system

Successful examples hide a surprising amount of operating cost.

If an AI system receives the right context, finds the expected data, calls a healthy service and produces an acceptable result, many very different architectures can look equally good.

Their differences become visible when one of those conditions breaks.

One system may recognize that a required source is unavailable and stop. Another may continue with partial information.

One may preserve enough state for an employee to take over at the failed step. Another may hand over a transcript that forces the employee to reconstruct the case.

One may stage a consequential action before committing it. Another may make the change first and discover the problem later.

One may expose exactly which model, route, tools and permissions produced the event. Another may leave the operations team with a generic log saying that "the AI" failed.

These are not edge details around the product. They are part of its economics.

The UK government's AI procurement guidance makes a related point from the buyer's side. It recommends testing AI systems under a range of conditions, defining acceptable performance, considering ongoing evaluation across the lifecycle, and planning for support and end-of-life rather than treating procurement as a one-time model choice. NIST's current TEVV work goes further on evaluation design: assessment should be customized to the actual goals and context of the system being evaluated, including agentic systems.

Neither says "make the demo fail." That is the operating implication I would take into the room.

Ask for a recovery contract

I would add a small artifact to an AI vendor evaluation: a recovery contract.

Not a legal contract. A testable description of what the system is expected to do when normal execution stops being trustworthy.

For the workflow being evaluated, define several representative failure conditions and agree in advance what acceptable recovery looks like.

The exact cases depend on the work. For an agent that can take actions across business systems, I would want to test at least some of these:

  • required information is missing or contradictory;
  • a connected service times out or returns malformed data;
  • the requested action exceeds the agent's authority;
  • the model or tool produces an uncertain result;
  • a human rejects an intermediate recommendation;
  • execution is interrupted after some state has changed;
  • the same request is retried;
  • the workflow reaches an exception the vendor did not prepare for.

Then evaluate the recovery, not merely whether an error message appeared.

Detection: Did the system recognize that normal execution was no longer safe or reliable?

Containment: Did it stop the failure from spreading into other actions or systems?

State: Can you establish what was read, decided, attempted and changed before the failure?

Escalation: Does the right person receive enough context to make the next decision?

Reversibility: Can consequential changes be undone, compensated for or safely completed?

Resumption: Can the workflow continue from a known safe point, or does the organization have to rebuild the work manually?

That gives the buyer evidence a successful demonstration cannot provide.

Recovery distance belongs in the business case

There is a simple way to think about the result.

Call it recovery distance: the amount of human and system work between a failed AI action and a dependable workflow again.

Sometimes the distance is tiny. The system detects low confidence, asks one precise question and continues.

Sometimes it is a clean handoff. The employee receives the case, relevant evidence and a clear decision to make.

Sometimes it is a replay from the last safe state after a transient service failure.

And sometimes "recovery" means an employee opens three systems, works out what the agent already changed, reconstructs missing context, corrects the data and starts the task again.

Those failures may appear identical in a dashboard. Both can be counted as one unsuccessful run.

Operationally, they are not remotely equivalent.

This is why AI vendor evaluation should not stop at accuracy, task completion or a benchmark score. Those measures matter, but the buyer eventually operates the whole system, including its exceptions.

I made a similar argument about evaluation in The benchmark needs a runtime address: the thing being qualified is increasingly the serving route, harness, tools, policies and operating conditions, not just the advertised model name.

Recovery belongs to that same system boundary.

A vendor can change the model and preserve the product name. It can change routing, orchestration or tool behavior. Your own integrations and permissions will change too. If those changes alter how failures are detected or recovered, the old evidence should not be treated as permanent.

A failure test is also a supervision test

The recovery exercise reveals something else that procurement spreadsheets often miss: how much human attention the product needs to remain dependable.

A vendor may describe a workflow as automated because the AI performs most of the visible execution. But if every unusual case arrives with weak context, employees may spend substantial time deciding what happened before they can decide what to do.

That is not merely a usability problem. It changes the operating model.

The most useful escalation is not "AI needs help."

It is closer to:

The workflow stopped here. This is the evidence available. This action was attempted but not committed. These two conditions conflict. You have authority to choose A or B.

The difference is the amount of cognition the system has successfully carried to the boundary.

That is why I would include human recovery effort in any serious proof of value. How many interventions occur? How long do they take? How often does the employee need to reconstruct context? Which interventions represent deliberate authority, and which exist because the system cannot reliably manage its own exceptions?

A small pilot is the right place to expose this. A pilot is a decision instrument precisely because its job is to reduce uncertainty before the organization makes a larger commitment.

A vendor demo should feed that decision, not pre-empt it.

The buyer should design the unhappy path

Vendors should demonstrate their strongest case. That is their job.

The buyer's job is different.

The buyer has to discover what the product will ask of the organization after the sales team leaves: which data must remain clean, which integrations must stay healthy, which exceptions need people, which actions need approval, which failures can be reversed, which logs are sufficient for investigation, and how much internal expertise is required to keep the workflow useful.

Those questions are difficult to answer while everything is working.

So make failure part of the evaluation protocol.

Do not manufacture absurd torture tests that have nothing to do with the intended use. Choose failures that the real workflow will eventually encounter, agree on the expected behavior, and run them against the actual configuration you are considering buying.

Then measure the distance back to dependable work.

A system that succeeds beautifully and recovers badly has shown you only half of its operating model.

Do not buy the half you saw.


Sources

Berk Bayri

Creative Technology & Innovation Leader

Designing and building for digital environments since 1998, across strategy, product, design, technology and organizational innovation.

About Berk →