The provocative part of OpenAI's Decisions API is not that AI can make choices. It is that a model score can quietly become an action — and an action can quietly become policy.
The dangerous part of OpenAI's new Decisions API is not that an AI might make the wrong decision.
It is that a probability can quietly become an action, and an action can quietly become policy.
That is a much bigger shift than “automating repetitive business decisions.” Companies have used models for classification, scoring, routing, moderation and prioritization for years. What changes when judgment itself becomes a first-class API surface is not simply the amount of automation. It is the location of the decision boundary.
Until now, many generative-AI systems have hidden that boundary behind prose. We ask a model for an answer, parse it, validate it, repair the JSON, convert the output into an internal label, apply a rule and only then decide whether something should happen. The model is already making a probabilistic judgment; the software around it is mostly disguising that judgment as data.
A decision-native interface strips away some of that disguise.
And once it does, an uncomfortable question becomes impossible to avoid: who decides when a model's belief is strong enough to become someone else's reality?
A probability is not a decision. The moment software turns it into an action, the threshold becomes policy.
OpenAI announced the Decisions API at DevDay 2026 as a real-time decision layer powered by Luna. OpenAI now confirms that developers define questions with finite pre-defined answers, provide text or image context, and can use the result for tasks such as classification, request routing or choosing an agent's next action. The product is in limited preview. OpenAI's public developer reference and changelog still do not expose a dedicated endpoint, pricing, full request/response schema or a documented confidence/probability contract. Independent reporting from The New Stack goes one step further, reporting roughly 150 ms responses and confidence scores; I treat those as reported launch details, not as a stable OpenAI API guarantee until the production contract is public.
What is confirmed — and what is not
OpenAI confirms Luna, user-defined questions, finite pre-defined answers, text/image context, classification, routing, agent next-action use cases and limited-preview availability. OpenAI's public developer docs do not yet document a Decisions API endpoint, pricing, full schema or probability/confidence contract. The New Stack independently reports confidence scores and roughly 150 ms latency; until OpenAI publishes the production contract, those remain reported launch details rather than stable API guarantees.
That is the shift worth paying attention to.
Imagine a system says a refund request is legitimate with 94% confidence.
What happens next?
The probability itself cannot answer that question. One company may auto-approve anything above 90%. Another may require human review for transactions above €500 regardless of confidence. A regulated workflow may prohibit autonomous approval entirely. A low-risk internal process may accept 70% because delay is more expensive than a reversible mistake.
The model produces evidence. The organization decides what that evidence is allowed to do.
That means the most consequential object in a Decisions API may not be the answer at all. It may be the threshold policy around the answer.
A threshold encodes risk appetite, reversibility, financial exposure, customer impact, compliance obligations and the organization's tolerance for false positives versus false negatives. It can change by market, account tier, transaction value, role, workflow stage or time of day. If teams treat it as a magic number buried in application code, they will have converted business policy into an invisible constant.
And that is where a seemingly technical API starts becoming organizational infrastructure.
| If you optimize for... | You still need to decide... |
|---|---|
| Higher model accuracy | Which mistakes cost more? |
| Higher confidence | At what confidence is action allowed? |
| More automation | Which decisions remain human-owned? |
| Lower latency | Which checks can safely disappear? |
| Lower cost | What happens when the cheaper decision is wrong? |
| More consistency | Who can change the policy when reality changes? |
A Decisions API does not remove human judgment. It relocates it upstream, into threshold design, escalation rules and exception handling.
That can be a better place for judgment to live.
It can also make bad policy much easier to execute.
Decision-native systems sound safer when they expose probability instead of fluent prose. In many ways, they are. A number is easier to route, measure and audit than a paragraph.
But a precise-looking number creates a new temptation: to confuse confidence with permission.
An agent may ask whether it should send a message, approve a transaction, change a customer record, call a tool, escalate to a frontier model or stop. A high-confidence answer can make routing dramatically cleaner. It can also become a shortcut around governance if teams interpret “99% confident” as “99% authorized.”
Those are not the same thing.
A model can be extremely confident about a decision it is not entitled to make. A system can be technically correct and organizationally wrong.
Confidence is not permission
A high probability can justify a routing choice without justifying an irreversible action. Model confidence, organizational authority, financial exposure, reversibility and human escalation need separate controls.
That separation matters more as AI output gets closer to execution. The less translation there is between judgment and action, the easier it becomes for technical capability to inherit authority by accident.
A robust architecture therefore needs two distinct questions:
What is likely to be true?
and
What is this system allowed to do if it is true?
The first belongs to the model. The second belongs to the organization.
If Decisions API succeeds, that boundary may become one of the most important product surfaces in enterprise AI.
Generative AI has trained us to expect an answer every time.
That is useful in a chat interface. It is much less useful in operational systems.
A model can produce beautifully fluent certainty from ambiguous evidence. Humans then try to infer uncertainty from tone, hedging or citations. Software has an even harder problem: unless uncertainty is represented explicitly, the application has no clean way to know whether the system is confident, guessing or merely sounding confident.
Decision-native interfaces create an opportunity to make uncertainty operational rather than rhetorical.
The important design move is not forcing every input into yes or no. It is creating a region where the correct output is: do not automate this.
A system might reject or defer below one threshold, send ambiguous cases to human review in the middle and allow autonomous action only above a higher threshold. The numbers in such a diagram are illustrative, not universal. What matters is the shape of the system: uncertainty needs somewhere to go.
This also changes the role of human review. Instead of making people approve every action or inspect a random sample of everything, supervision can concentrate on the cases where the model is least decisive or the consequences are highest.
That is not a fallback.
It is a product choice.
OpenAI has not yet documented a public probability/confidence contract for Decisions API. The argument in this section therefore applies if and when a production decision signal is exposed; it is not a claim about the current public schema.
There is a problem with confidence scores that becomes operational the second software starts using them.
A number that looks precise can still be badly calibrated.
If a model says “90%” on one hundred comparable cases, roughly ninety should be correct for that number to mean what people naturally think it means. If only sixty are correct, the score may still rank cases reasonably well, but the threshold policy around it is operating on false premises.
This creates a new obligation: teams have to record what happened after the decision.
Was the support ticket actually urgent? Did the customer churn? Was the transaction fraudulent? Was the document really non-compliant? Did the agent's escalation prove necessary?
Without outcomes, you can measure consistency.
You cannot measure whether confidence means anything.
That is why calibration cannot remain a model-card statistic. It becomes an operating metric tied to the distribution the business actually sees.
This connects directly to an argument I made in The benchmark needs a runtime address. Once behavior depends on the configured system, evaluation has to follow the actual route, model, threshold, policy and operating conditions under which a decision was made. A decision endpoint can make this easier to instrument, but only if the feedback loop exists.
The useful question is not “did the API return a decision?”
It is “did the probability continue to mean what our policy assumed it meant?”
And that question gets harder every time the model, traffic mix or business environment changes.
If decision-native AI works well, organizations will probably create hundreds or thousands of small AI judgments across their software.
Route this. Rank that. Block this. Escalate that. Choose a model. Approve a draft. Flag an exception. Prioritize a lead. Decide whether a human is needed.
Each one looks harmless.
Together they create a new layer of business logic.
I would call the maintenance burden decision debt.
The debt is not “we used too much AI.” It is the accumulation of thresholds, criteria, model choices, fallback rules and exception paths that nobody can fully explain anymore.
A threshold chosen during a pilot survives after customer behavior changes. A taxonomy stops matching the organization. A low-risk workflow gradually becomes consequential. A model upgrade changes calibration while the threshold stays untouched. Two teams implement slightly different definitions of “urgent” and quietly route the same customer in different directions.
Traditional software already suffers from hidden business rules.
Decision-native AI can make those rules more adaptive.
It can also make them easier to scatter.
The governance problem therefore changes from “who approved this model?” to something much more operational: who owns this decision boundary, when was it last tested, and what evidence would make us change it?
If teams cannot answer that, the API may be clean while the operating model becomes increasingly opaque.
The obvious use case is repetitive business decisions.
The more interesting one may be agent control.
Agents constantly face narrow internal judgments: Is the evidence sufficient? Should I search again? Does this action need approval? Which tool should I call? Is the result anomalous? Should I escalate to a stronger model? Did the previous step actually satisfy the goal? Should I retry or stop?
Today, many agent systems ask the same general-purpose model that is doing the work to also decide what should happen next. That collapses execution and control into one component.
A dedicated decision layer can separate doing the work from deciding whether the work should continue.
That separation can improve latency, cost and observability if the decision layer is cheaper and more predictable than invoking a frontier model for every gate.
But the more important benefit is architectural: it creates explicit places where policy can attach.
An enterprise can say that one judgment may route work automatically, another may choose a model, another may call a reversible tool, and a fourth must always return to a human.
OpenAI's existing agent guidance already separates automated guardrails from human review and treats approval as an explicit control around consequential actions. A first-class Decisions API could make the middle layer — machine judgment before policy and action — much more legible.
That starts to look less like an API endpoint and more like a control plane.
If OpenAI only gives developers better generation, it competes primarily on model capability, price, latency and tooling.
If applications also begin asking OpenAI for structured judgments that drive routing, prioritization, escalation and policy gates, OpenAI moves deeper into the architecture of how organizations operate.
The API does not need to make the final decision to become strategically important.
It only needs to become the default place where software asks:
What does the evidence suggest right now?
That is the layer between raw state and action.
And once a vendor's decision behavior becomes embedded in enough operational thresholds, switching models stops being a simple quality comparison. Teams have to ask whether the replacement is calibrated in the same way, whether existing thresholds still behave correctly and whether historical outcomes remain comparable.
The lock-in risk is not only prompts, context or model APIs.
It can become decision behavior itself.
That may be one of the least visible forms of AI platform dependence because nothing looks locked in. The application still owns the threshold. The organization still owns the workflow. But the meaning of “0.84” may depend on one provider's model behavior.
Change the provider, and the policy may silently change with it.
I would not start with “how accurate is it?”
I would start with whether the decision layer is operable.
Five questions I would use
1. **Calibration:** When the API says 80%, does 80% mean roughly 80% on our real traffic? 2. **Threshold ownership:** Can we see who decided when a probability becomes an action? 3. **Abstention:** Is there a deliberate path for uncertain or high-consequence cases to return to a person? 4. **Outcome feedback:** Can we connect decisions to what actually happened and detect drift? 5. **Recovery:** When the model, threshold or policy is wrong, can we identify affected decisions and reverse or repair the result?
Those questions matter because the hidden metric in AI automation is supervision.
A Decisions API can reduce supervision if it removes brittle parsing, narrows the answer space and routes only ambiguous cases back to people.
It can also create a new supervision burden if every threshold becomes a mystery and every model change forces teams to rediscover how the business behaves.
The difference will not be visible in a demo.
It will show up in operations.
That framing is already too crude.
AI systems have been influencing decisions for years. Scores rank leads. Models flag fraud. Recommendation systems shape attention. Risk systems determine which cases receive scrutiny. The arrival of a Decisions API does not suddenly hand judgment to machines.
What it does is make the judgment layer more explicit, more programmable and potentially much easier to connect directly to action.
That is useful.
It is also exactly why the policy around that judgment matters more.
If OpenAI's Decisions API merely turns model output into a cleaner field, it will be convenient.
If it helps teams make threshold ownership visible, preserve abstention, separate confidence from authority, measure calibration, trace outcomes and recover from wrong decisions, it could become something far more important: an explicit layer between AI judgment and organizational action.
The most important question is therefore not:
Can the model make the right decision?
It is:
Who owns the moment when a probability becomes policy — and can they still see it after the system has been running for six months?
That is the boundary worth designing.
url: https://openrouter.ai/typesafe/jev-1.13 title: TypeSafe Jev 1.13 — Decisions API publisher: OpenRouter date: 2026-09
url: https://docs.empiriolabs.ai/decisions title: Decisions API publisher: EmpirioLabs AI date: 2026-09
One email per essay. No noise between.
Berk Bayri
Creative Technology & Innovation Leader
Designing and building for digital environments since 1998, across strategy, product, design, technology and organizational innovation.