Berk Bayri
AI Strategy9 min read·

Your AI data agent is creating a second dataset

AI data agents do more than answer business questions. At scale, the pattern of questions, corrections and dead ends becomes a new signal: a map of what the organization is trying to understand, where its definitions are weak, and which decisions deserve better infrastructure.

The obvious output of a data agent is an answer. The more interesting output may be the pattern of questions people keep asking.

Today, OpenAI launched a Data agent in ChatGPT Work that lets people query company data, investigate changes and build dashboards by asking in plain language. The product is built around a familiar promise in analytics: fewer people waiting for someone else to run the analysis, more people able to answer their own questions.

AI data agents make that self-service dramatically easier. They also create something we have not had at this scale before.

When thousands of business questions move from hallway conversations, Slack messages, ticket queues and analyst meetings into an instrumented agent, the organization starts producing a new kind of data about itself.

Not another copy of the warehouse.

A record of what people are trying to understand about the warehouse.

I think that second dataset may become one of the most valuable by-products of enterprise AI analytics.

The first dataset tells you what happened. The second tells you what people need to know.

Traditional analytics mostly observes the business.

Orders, customers, revenue, conversion, inventory, service levels, product usage, costs. We build models so the organization can ask better questions of those events.

But the demand side of analytics has always been much harder to see.

Which questions recur every Monday morning?

Which metric makes three departments ask different versions of the same thing?

Where do people keep adding a second question because the first answer did not resolve the decision?

Which questions repeatedly end with, “We actually don’t have that data”?

Which apparently simple requests require a senior analyst to explain what the business really means by a word like active, pipeline, retained or profitable?

Some of that demand appears in tickets and backlogs. Much of it disappears into meetings, messages and one-off analysis.

AI data agents change the economics of asking. The question can be cheap, immediate and conversational. That means question volume can grow dramatically.

LangChain says its self-service data agent now handles roughly 40 times the request volume its three-person data team could manage directly. In a recent 30-day period, it reported about 2,200 agent conversations among provisioned users. Those are vendor-reported internal figures, not a general benchmark, but they illustrate the scale change clearly.

At that volume, the conversation stream stops being merely a support channel.

It becomes evidence.

A repeated question is not just usage

Most analytics products will naturally measure adoption.

How many people used the agent? How many questions did they ask? How quickly did it answer? How many dashboards did it create?

Those numbers matter. They are also the shallowest reading of the signal.

A repeated question can mean very different things.

It can mean a dashboard is hard to find.

It can mean the dashboard exists but does not answer the actual decision.

It can mean two teams use the same metric name differently.

It can mean the data model is missing a relationship people keep needing.

It can mean the organization has never agreed on the question in the first place.

LangChain has already found part of this feedback loop in practice. Its team looks at trends in agent conversation topics, warnings, issues and context gaps. If people keep asking similar questions, it may build a better dashboard. If the agent struggles with a metric, it may improve the semantic definition. LangChain describes agent conversations as a signal of what the company is trying to understand.

That is the point I would push further.

Do not use this signal only to make the agent better.

Use it to make the organization easier to understand.

Question demand is an operating signal

Imagine two hundred people ask an AI data agent questions over a month.

Individually, those conversations are analysis requests. In aggregate, patterns begin to describe the organization’s unresolved decision demand.

A cluster of repeated questions around onboarding may indicate a reporting problem. Or it may indicate that several teams are making decisions about onboarding from different definitions of success.

A burst of follow-up questions after a KPI moves may reveal that the top-line metric is visible but its drivers are not modeled well enough to support action.

Frequent corrections can point to missing business context. Frequent source switching can expose weak trust in canonical data. Questions the agent cannot answer can reveal missing instrumentation. Questions that repeatedly need a human can mark places where judgment, policy or cross-functional alignment matters more than self-service.

None of these interpretations should be automated blindly.

A pattern is evidence, not a diagnosis.

That distinction matters. I have argued before that the reported problem is a starting hypothesis, not the problem itself. The same rule applies here. Ten similar questions do not prove why the questions exist. They tell you where to investigate.

That is already far more useful than treating them as ten successful chat sessions.

The question stream can expose where the company is still manual

OpenAI’s Data agent relies on business terms, metric definitions, custom calculations and data relationships supplied by semantic layers and other trusted sources. Administrators control which connections are available, and queries inherit existing table, row and column permissions.

That architecture is important because a data agent cannot manufacture organizational clarity from raw tables.

Anthropic makes the same point from another direction in its own internal analytics deployment. It instruments each question handled by its Slack data agent, including the skill versions used, user corrections and data-quality warnings. Its team says changes in adoption by domain can signal that a skill has drifted or that a new class of questions has appeared that the semantic layer does not yet cover.

The interesting phrase there is new class of questions.

A new class of questions can be a product requirement for the data platform.

It can also be a sign that the business itself has changed faster than its reporting model.

A company launches a new commercial model, enters a market, changes pricing, reorganizes account ownership or introduces an AI-assisted workflow. Suddenly people start asking combinations of questions nobody designed the warehouse around six months ago.

The data platform may still be technically healthy. The decision environment has moved.

Question patterns can show that movement earlier than the reporting roadmap does.

This should not become employee surveillance

There is an obvious bad version of this idea.

Store every prompt forever. Rank employees by the sophistication of their questions. Infer performance from who asks what. Turn analytics curiosity into another behavioral monitoring system.

That would be both ethically poor and strategically stupid. People will stop exploring as soon as they believe every imperfect question is being evaluated as a statement about them.

The useful unit is usually not the individual prompt or the individual employee.

It is the aggregate pattern.

What question categories are growing?

Where are people reformulating the same request?

Which domains generate the most corrections or unresolved answers?

Which questions repeatedly move from self-service to human review?

What is being asked often enough that it should become a durable dashboard, metric, model, process or decision rule?

The telemetry should be designed for that purpose: minimal retention, appropriate access, aggregation where possible, and clear boundaries around what is not being measured.

A system intended to make questions easier to ask should not make people afraid to ask them.

The data team gets a new kind of backlog

If this works, the data team’s backlog changes.

Instead of receiving only explicit requests, it can see recurring demand before every request becomes a ticket.

That creates a useful distinction between three kinds of work.

Some questions should remain conversational. They are genuinely one-off and the agent can answer them well.

Some questions should become products. If fifty people keep asking the same thing, the right answer may be a durable metric, dashboard, model or automated readout rather than fifty generated answers.

And some questions reveal decisions the organization has not designed properly yet. No dashboard will fix those. They need a definition, an owner, a policy choice or a cross-functional conversation.

The agent should help separate those categories, not collapse them into one giant measure of “questions answered.”

This is where self-service analytics can become more than a productivity story.

It can become a discovery system for the organization’s own uncertainty.

Evaluate the questions, not only the answers

This also changes how I would evaluate an AI data agent pilot.

Accuracy matters. Latency matters. Permissions matter. Adoption matters.

But I would add another question to the pilot:

What did we learn from what people tried to understand?

At the start, define a small set of signals worth reviewing. Repeated question clusters. Reformulation depth. Corrections. Human escalations. Unanswered questions. Source conflicts. Questions that recur even after a dashboard exists.

Then review them with the people who understand the decisions behind the data: analytics, product, operations, finance, sales, whoever owns the domain.

The purpose is not to produce another telemetry dashboard nobody acts on.

It is to decide which patterns deserve a change.

A better definition.

A new data relationship.

A canonical source.

A dashboard.

A process redesign.

Or simply a clear statement that this decision still requires human judgment.

That is consistent with how I think pilots should work generally: a pilot is useful when it produces evidence for a decision, not when it merely demonstrates that the technology functions.

A data-agent pilot can produce two forms of evidence at once: whether the system answers well, and whether the organization understands its own demand for answers.

There are now two things worth modeling

For years, analytics teams have invested enormous effort in modeling the business correctly.

That remains necessary. AI data agents arguably make it more important because more people can now query those definitions directly.

But widespread conversational analytics adds another object worth modeling: the questions around the business.

Not every prompt. Not every user. Not another surveillance archive.

The recurring structure of uncertainty.

Where people ask.

Where they re-ask.

Where they disagree.

Where the available data stops being enough.

Where an answer repeatedly fails to become a decision.

The first dataset describes what the business is doing.

The second describes what the business is trying to understand about itself.

If AI data agents make that second dataset visible, the smartest organizations will not use it only to improve the agent.

They will use it to decide what the organization needs to understand next.


Sources

Berk Bayri

Creative Technology & Innovation Leader

Designing and building for digital environments since 1998, across strategy, product, design, technology and organizational innovation.

About Berk →

Get new essays in your inbox.

New essays by email, when they are published.