Tool call
When an AI model asks to run an external function, such as searching, reading a file or updating a record. A failed tool call can be a very different failure from a wrong answer.
A tool call is how a model reaches outside itself. Instead of only producing text, the model emits a structured request to run a function (search the web, query a database, send an email, update a system of record), the surrounding software executes it, and the result is returned to the model.
Tool use is what turns a language model into something that can act, which is why it is central to agentic AI. The harness decides which tools exist, how they are described and what happens on failure. Standards such as the Model Context Protocol make tools easier to expose.
Why failures differ
An incorrect calculation is not the same failure as a tool call made with the wrong account permissions. The first is a quality problem; the second is an authority problem with real effects. That is why capability is not authority, and why retries and tool calls also show up as their own line in a cost-migration ledger.
Read more in The wrong questions about AI right now.
Related terms
Harness
The software around a model that manages context, tool use, sub-agents and the execution environment. It shapes real-world capability, so it must be evaluated with the model, not ignored.
Sub-agent
A secondary AI agent that a main agent starts to handle part of a task, such as research or a code change, often with its own context and tools. It is part of the harness around the model.
Capability is not authority
A design principle for AI agents: being technically able to perform an action does not mean the agent should be permitted to perform it. Delegation needs gradients of authority.
Model Context Protocol
An open standard, usually called MCP, for connecting AI models and agents to external tools and data sources in a consistent way, so they can discover and call an organization's capabilities.
Used in these essays
The wrong questions about AI right now
Many of the questions that helped us orient ourselves around generative AI are now too blunt to be useful. The harder work is no longer asking what AI is in the abstract, but specifying where it works, where it fails, what authority it should have, and what the whole system costs.
A cheaper AI model can move the cost instead of removing it
A lower model bill can hide a higher workflow bill. The useful AI TCO question is not only what got cheaper, but where the cost moved.
The benchmark needs a runtime address
AI systems are increasingly dynamic at runtime. Enterprise evaluation should qualify the serving route, harness, tools and fallback conditions, not just the model name.