Berk Bayri
Problem Framing & Decision Design14 min read·

The wrong questions about AI right now

Many of the questions that helped us orient ourselves around generative AI are now too blunt to be useful. The harder work is no longer asking what AI is in the abstract, but specifying where it works, where it fails, what authority it should have, and what the whole system costs.

The wrong questions about AI right now

0:00 / 0:00

Some questions age badly.

They begin as useful ways to orient ourselves around a new technology. Then the technology changes, the evidence gets richer, the operational reality gets messier — and the questions survive because they are easy to ask.

That is where a lot of the AI conversation is now.

A few years ago, questions like Can AI reason?, Which model is smartest?, Will AI replace jobs?, Can we trust it? and When will agents replace apps? were understandable. We were trying to locate a strange new capability on a map we already knew.

In 2026, many of those questions are no longer wrong because the answer is obviously “yes” or “no.” They are wrong because the answer is too low-resolution to be useful.

The evidence has become almost aggressively conditional. Frontier models can perform at or above human baselines on some difficult academic tasks while still failing on surprisingly basic ones. Benchmarks are saturating faster, model rankings are compressing, agents are getting much better at long tasks while still failing often enough to make unsupervised deployment consequential, and workplace effects are showing up unevenly across tasks, occupations and career stages rather than as one clean wave of replacement.

The problem is not that we know nothing.

The problem is that the most familiar questions often erase exactly the variables that matter.

A wrong question is often one whose answer can be debated forever without changing the next decision.

The useful shift is from asking what AI is in the abstract to asking what a particular AI system does under specific conditions, with specific authority, failure modes, costs and consequences.

That sounds less dramatic.

It is also where most of the real work has moved.

The unit of analysis has changed

In the first generative-AI wave, the model often looked like the product. Now the meaningful unit is increasingly the whole system: model, routing, tools, data, permissions, interface, workflow, humans, controls and recovery.

“Can AI reason?”

This question still produces enormous amounts of argument because the word reasoning is doing too much work.

If a model solves an Olympiad-level problem, did it reason?

If the same class of system then misreads an analog clock, violates an obvious constraint in a multi-step workflow or confidently accepts a false premise, did it stop reasoning?

The binary question collapses a jagged capability surface into a philosophical label.

That is no longer a useful way to evaluate a system.

Stanford’s 2026 AI Index captures the problem neatly: frontier models have made extraordinary gains on difficult reasoning benchmarks, while performance remains uneven across tasks that humans would not naturally rank on the same difficulty scale. It also reports that benchmarks intended to remain difficult for years are saturating in months. On some widely used evaluations, invalid-question rates are themselves high enough to make leaderboard differences difficult to interpret cleanly.

So the question is not whether “AI reasons.”

A procurement team cannot buy “reasoning.” A product team cannot set an error budget for “reasoning.” A risk team cannot approve “reasoning.”

What they can evaluate is a sequence of decisions inside a real workflow.

A better question is:

Which reasoning steps in this workflow are reliable enough, which failures matter, and which failures can we detect before they become consequential?

That turns a metaphysical argument into an evaluation plan.

It also forces an uncomfortable admission: a system can be brilliant at the hard-looking part and unreliable at the boring-looking part.

The boring-looking part is often where production breaks.

“Which model is best?”

This question made more sense when model capability gaps were large and durable.

They are now much less durable.

The 2026 AI Index reports that leading models are increasingly clustered near one another on major public evaluations, while benchmarks themselves are becoming less stable decision instruments. A few points on a leaderboard can disappear when the task changes, the prompt changes, the tool harness changes, the context changes, the model is routed differently, or the workload moves from a benchmark to a real operating environment.

Even the phrase the model is becoming slippery.

A production request may pass through a router, retrieve context, call tools, use a particular inference configuration, invoke another model, hit a permission boundary, retry, fall back, or be escalated to a human.

The thing that succeeds or fails is the route.

That is why “Which model is best?” is increasingly similar to asking which engine is best without saying what vehicle it is in, what road it is on, what it is carrying or whether fuel economy matters.

A better question is:

Which configured system performs best on our workload, at our acceptable quality, latency, privacy, cost and recovery constraints?

That question usually produces a less exciting answer than a leaderboard.

It also produces an answer you can use.

This is the same reason a benchmark increasingly needs a runtime address: the evaluated object is not merely a model name. It is the serving path that produced the result.

“Does AI hallucinate?”

Yes.

But that is now about as useful as asking whether databases can contain wrong data.

The relevant question is not whether failure exists. It is what kind of failure exists, how often it matters, whether it can be detected, and what happens next.

Stanford’s 2026 responsible-AI review reports very wide hallucination rates across top models on one recent accuracy benchmark, with performance changing sharply depending on how false beliefs are framed. That does not give us one universal “hallucination rate.” It tells us the opposite: epistemic reliability is conditional.

The word hallucination also hides several operationally different problems.

A fabricated citation is not the same failure as an outdated fact.

An outdated fact is not the same failure as an incorrect calculation.

An incorrect calculation is not the same failure as a tool call made with the wrong account permissions.

A plausible but unsupported answer in a brainstorming tool is not the same event as a plausible but unsupported answer in a medical, financial or compliance workflow.

The better question is:

Which errors are unacceptable here, what evidence must the system provide, how will we detect unsupported output, and what is the cost of a false positive versus a false negative?

That question immediately changes system design.

It may lead to retrieval.

Or structured outputs.

Or deterministic checks.

Or a human approval step.

Or no AI at all.

The important thing is that “hallucination” stops being treated as a mystical property of the model and starts being treated as a failure mode in a specific system.

“Will AI replace jobs?”

This may be the most persistent bad question because it compresses an entire labor market into one verb.

A job is not one task.

It is a bundle of tasks, relationships, responsibilities, permissions, tacit knowledge, learning loops and institutional expectations.

AI does not need to replace the whole bundle to change the job.

And it does not need to reduce total employment to create serious distributional effects.

The most recent evidence is already more interesting than the replacement frame. A Stanford Digital Economy Lab analysis using payroll data through June 2026 reports no evidence of widespread economy-wide displacement, while finding a substantial employment gap for workers aged 22–25 in highly AI-exposed occupations. A separate September 2026 study across 41 countries finds changes in the junior share of employment after firm AI adoption, driven primarily by growth in senior employment rather than a simple collapse in total headcount. An August 2026 NBER study describes current workplace adoption as widespread but shallow: generative AI appears across many occupations and tasks, but adoption still varies heavily among workers doing similar work.

Those are three very different stories from “AI is replacing jobs.”

They point toward work being restructured unevenly.

They raise questions about entry-level ladders.

They raise questions about who gets augmented first.

They raise questions about whether junior tasks disappear before organizations redesign how junior people learn.

They raise questions about whether productivity gains become headcount reductions, higher output, lower prices, more experimentation or simply more work.

A better question is:

Which tasks are being compressed, which responsibilities are moving, which learning loops are disappearing, and what new supervision or judgment work is being created?

And then:

Who gains from that redesign, and who loses access to the path that used to build expertise?

That is harder to put in a headline.

It is much closer to what is actually happening.

“How much can we automate?”

This sounds practical.

It often is not.

Maximum automation is not the same thing as maximum value.

The strongest AI agents are now capable of completing much longer and more complex tasks than systems from only a year ago. METR’s time-horizon work continues to document rapid increases in the length of tasks frontier systems can complete autonomously. At the same time, Stanford reports that the best systems still fail roughly one in three attempts on OSWorld, a structured benchmark built around real computer tasks.

Both facts can be true.

Capability can rise quickly while reliability remains too low for some forms of authority.

That is why “How much can we automate?” is the wrong optimization target.

It encourages teams to maximize the amount of visible human work removed while ignoring the human work that moves somewhere else: review, exception handling, context provision, escalation, recovery and accountability.

I wrote about this as the supervision load: the human remainder does not vanish merely because the system is nominally autonomous.

A better question is:

What is the maximum authority we should delegate while keeping failures detectable, containable, reversible and economically worthwhile?

That question treats autonomy as a design variable rather than a trophy.

It also forces a distinction between doing more steps and being allowed to cause more consequence.

Those are not the same thing.

“When will agents replace apps?”

Probably never in the clean way the question implies.

Agents may absolutely reduce the number of screens people visit.

They may absorb navigation.

They may call software capabilities directly.

They may turn many visible interfaces into background infrastructure.

But an agent still needs something to act on.

It needs data models, permissions, identity, transaction boundaries, business rules, tools, state, auditability and recovery paths.

In other words, agents do not make application capability disappear.

They make capability design more important.

NIST’s 2026 work on agent identity and authorization is revealing here. The problem is not simply whether an agent can call a tool. The problem is how the agent is identified, what it is authorized to do, what data it can reach, and how those permissions are controlled across systems.

So the useful question is not:

When will agents replace apps?

It is:

Which capabilities should become safely callable by agents, under which identities and permissions, and which decisions should remain explicit to a human?

The interface may disappear.

The responsibility does not.

“How much does AI cost?”

Usually this means tokens.

Or inference.

Or the monthly software bill.

That is a local accounting view of a system-level question.

A cheaper model can trigger more retries.

A faster model can create more output that humans must review.

A more autonomous agent can reduce handling time in one team while increasing exceptions in another.

A lower inference bill can coexist with a higher operating cost.

The relevant metric therefore depends on the outcome you are trying to buy.

If you are automating claims review, the cost per million tokens is not the business unit.

The business unit is something closer to the cost per correctly completed claim, including review, exceptions, recovery and governance.

If you are generating code, the useful unit is not model spend per developer. It may be cost per accepted change after testing, review, rework and production failure.

The better question is:

What is the total cost per accepted outcome, and where did cost move when the model changed?

That is the argument behind A cheaper AI model can move the cost instead of removing it.

The cheapest model is not necessarily the cheapest system.

And the cheapest system is not necessarily the most valuable one.

“Is AI safe?”

Safe enough for what?

This is another question that becomes less meaningful as systems gain access to more real-world capability.

NIST’s AI Risk Management Framework defines risk in terms of both likelihood and consequence, and explicitly notes that risk varies across lifecycle stages and across model, system, application and ecosystem levels. NIST’s newer work on agent security makes the same problem more concrete: once agents receive access to tools, applications and data, identity and authorization become part of the safety question.

A summarizer that produces a bad paragraph and an agent that produces the same bad paragraph and then sends it to 40,000 customers are not equally unsafe.

The output can be identical.

The consequence is not.

That is why asking whether a model is “safe” without specifying authority is increasingly close to meaningless.

A better question is:

Safe enough to perform which action, with which data, under which permissions, at what blast radius, with what monitoring and what recovery path?

The 2026 AI Index reports that documented AI incidents continued to rise even as responsible-AI practices became more formalized. That is not evidence that safety work is useless. It is evidence that “safe AI” cannot be reduced to a one-time property assigned to a model before deployment.

Safety is an operating condition.

“Is this AGI?”

This is a legitimate scientific, philosophical and policy question.

It is often a terrible management question.

For most organizations, an AGI label does not tell you what to deploy on Monday.

It does not specify whether a model can access production data.

It does not specify whether it can execute a payment.

It does not specify whether its failure is reversible.

It does not specify whether the economics work.

And it does not tell you whether a use case that was rejected six months ago should now be reopened.

The operationally useful question is:

Which capability threshold has actually changed enough to change one of our decisions?

Maybe a task that was previously too expensive is now viable.

Maybe a workflow that failed because context was too small can now be reconsidered.

Maybe a use case that was rejected because the model could not handle images is worth reopening.

Maybe nothing important changed for your organization at all.

This is why I prefer keeping a record of the AI ideas you rejected, including the constraint that killed them. A new model release should reopen the decisions whose binding constraint changed — not restart the entire AI roadmap because someone declared a new era.

“Is this AGI?” asks for a category.

“What decision changed?” asks for consequence.

For most organizations, the second question is more valuable.

The pattern behind the wrong questions

The old questions tend to share three properties.

First, they treat AI as if it has one stable level of capability.

It does not.

Second, they treat the model as the unit of analysis.

Increasingly, it is not.

Third, they ask for a binary answer where the actual decision depends on conditions.

That is the deeper shift.

The first phase of generative AI was dominated by capability discovery.

Can it write?

Can it code?

Can it use tools?

Can it see?

Can it reason?

Can it act?

Those questions were necessary.

The next phase is dominated by boundary design.

Where is it dependable?

Where is it economically useful?

What should it be allowed to do?

What must remain observable?

Which failures can be tolerated?

Which failures must be impossible?

Where does human judgment still create more value than it costs?

And who owns the consequence when the system is wrong?

The technology got better.

The questions need to get better too.

Stop askingStart asking
Can AI reason?Which reasoning steps are reliable enough here, and which failures can we detect?
Which model is best?Which configured system meets our workload, quality, latency, privacy, cost and recovery constraints?
Does AI hallucinate?Which errors are unacceptable, how will we detect them, and what is the cost of being wrong?
Will AI replace jobs?Which tasks, career ladders, responsibilities and supervision burdens are being redistributed?
How much can we automate?How much authority can we delegate while keeping failure detectable, containable and reversible?
When will agents replace apps?Which capabilities should become safely callable by agents, under which identities and permissions?
How much does AI cost?What is the total cost per accepted outcome, and where did the cost migrate?
Is AI safe?Safe enough to do what, with what data, permissions, blast radius and recovery path?
Is this AGI?Which capability threshold changed enough to change a real decision?

The goal is not to make AI less interesting.

It is to stop allowing interesting questions to substitute for useful ones.

The most important AI question in 2026 is rarely a question about AI in general.

It is a question with a workload, a threshold, an owner and a consequence.

That is usually where the answer begins to matter.

Sources

Stanford HAI — 2026 AI Index: Technical Performance

Stanford HAI — 2026 AI Index: Responsible AI

Stanford HAI — 2026 AI Index: Economy

Stanford Digital Economy Lab — Canaries in the Coal Mine? Six Facts about the Recent Employment Effects of Artificial Intelligence, revised August 2026

Stanford Digital Economy Lab — How Does AI Change Labor Demand? Evidence from 41 Countries, September 2026

NBER — What Work Does Generative AI Do?, August 2026

METR — Time Horizon 1.1, January 2026

Anthropic — Economic Index: Cadences, June 2026

NIST — Accelerating the Adoption of Software and AI Agent Identity and Authorization, February 2026

NIST — Security Considerations for AI Agents, May 2026

NIST — AI Risk Management Framework

Get new essays as they publish.

One email per essay. No noise between.

Berk Bayri

Creative Technology & Innovation Leader

Designing and building for digital environments since 1998, across strategy, product, design, technology and organizational innovation.

About Berk →