Prompt
The instructions and context given to an AI model to produce a response. In persistent, agentic systems, the goal increasingly replaces the single prompt.
A prompt is the input given to a language model: a question, instructions, examples, documents or a mix. The quality of the output depends heavily on it, which is why prompt writing became a skill of its own.
In a real system the prompt is rarely typed by a person. The harness assembles it from instructions, retrieved information, tool results and conversation history, so the effective prompt can be much longer and more varied than anything a user writes.
Beyond the prompt
Most AI use so far is a tool we operate: write a prompt, read the output. Persistent delegation changes that. The user delegates a goal and the agent keeps working without waiting for the next prompt. That shifts the important design questions from how to phrase a request to what the agent remembers, which actions it is allowed to take and how it reports back.
Prompting also adds to supervision load: time spent prompting, supplying context and correcting is human attention that automation metrics often miss. See LLM.
Read more in OpenAI Dots is a test of whether AI can carry a goal, not just complete a task.
Related terms
LLM
Large language model: an AI model trained on very large amounts of text to predict and generate language, which can be used to answer, summarise, write, reason about and act on text.
Persistent delegation
Handing an AI agent a goal it keeps carrying over time — remembering context, noticing relevant events and acting across applications — rather than completing a single prompted task.
Harness
The software around a model that manages context, tool use, sub-agents and the execution environment. It shapes real-world capability, so it must be evaluated with the model, not ignored.
Used in these essays
Stop adopting AI
AI adoption is rising faster than enterprise value because companies keep installing new intelligence inside old operating models.
OpenAI Decisions API turns probability into policy
The provocative part of OpenAI's Decisions API is not that AI can make choices. It is that a model score can quietly become an action — and an action can quietly become policy.
OpenAI Dots is a test of whether AI can carry a goal, not just complete a task
Dots matters less as another capable assistant than as a test of persistent delegation: can AI keep carrying a goal without giving the user a new system to manage?
Cahit Arf’s 1959 question about thinking machines
In a 1959 text on thinking machines, Cahit Arf moved from clocks and relays to a harder question: what happens when a machine meets a problem that was not anticipated when it was built?
The wrong questions about AI right now
Many of the questions that helped us orient ourselves around generative AI are now too blunt to be useful. The harder work is no longer asking what AI is in the abstract, but specifying where it works, where it fails, what authority it should have, and what the whole system costs.
The first AI incident report should be incomplete
OpenAI's new misalignment disclosure framework exposes a useful enterprise design principle: record anomalous AI behavior before the organization has finished explaining it. Otherwise incident systems quietly become filters for what teams already understand.
The hidden metric in AI automation is supervision
As AI moves from assisting work to leading it, hours saved stop telling the whole story. The scarce resource shifts to human supervision: approvals, exceptions, context and judgment.
The benchmark needs a runtime address
AI systems are increasingly dynamic at runtime. Enterprise evaluation should qualify the serving route, harness, tools and fallback conditions, not just the model name.
Your AI data agent is creating a second dataset
AI data agents do more than answer business questions. At scale, the pattern of questions, corrections and dead ends becomes a new signal: a map of what the organization is trying to understand, where its definitions are weak, and which decisions deserve better infrastructure.
Your chatbot is borrowing from the next interaction
AI customer service is usually measured one interaction at a time. But a failed automated interaction can change which channel a customer chooses next time. That makes future adoption part of the economics, not a separate trust metric.