Components can keep an AI experience visually on-system while the agent is behaviorally off-policy. Design systems now need reusable rules for what AI may present, decide, execute, defer, expose as uncertain, confirm, reverse and escalate.
An AI can use the correct component, tokens and accessibility rules and still expose the wrong action, hide uncertainty or make a decision feel more settled than it is. The next job for design systems is not only to keep AI-generated interfaces on-system. It is to keep AI-mediated interactions inside the organization’s intended boundaries.
Design systems are becoming machine-readable.
That is an important development. Kaelig Deloumeau-Prigent's July 2026 field study, highlighted this week by Vitaly Friedman, documents how teams are packaging design-system knowledge for AI through MCP servers, machine-readable documentation, skills, generated context files, automated audits and evaluation loops. The practical goal is familiar: reduce drift. If an AI generates an interface, it should use the right components, tokens, content rules and accessibility patterns instead of inventing a parallel product language.
This is necessary work. It is also only half of the problem.
Imagine an expense agent that correctly uses your design system’s modal, button hierarchy, form fields, typography and accessible labels. It proposes reimbursing an employee for €8,400. The interaction looks impeccable.
But should the agent have been allowed to present “Approve” at all? Should it have explained that the receipt classification was uncertain? Should a manager have been required to confirm? If approval is given, is the action reversible? If the policy exception is ambiguous, should the agent stop and escalate instead of presenting a confident recommendation?
None of those questions is answered by whether the interface is visually on-system.
A component system can tell an AI how an approval should look. It does not necessarily tell the AI when approval is an appropriate interaction.
That distinction is becoming consequential as software moves from displaying information to taking action.
The interface can be perfectly on-system while the interaction is off-policy.
The modern design system became valuable because product teams kept solving the same problems repeatedly.
How should a destructive action look? What does a disabled control mean? How do we announce an error to assistive technology? Which spacing, type, color and interaction states belong to the product? What does a modal do on mobile? Which patterns are approved, deprecated or inaccessible?
Instead of leaving every designer and engineer to answer those questions independently, organizations turned repeated decisions into shared infrastructure.
That infrastructure increasingly reaches beyond visual components. Mature systems contain content guidance, accessibility requirements, interaction patterns, implementation rules and governance. IBM’s Carbon Design System, for example, now treats its stable AI label as more than a badge: it is a consistent route to AI explainability. Microsoft’s agent-design guidance similarly emphasizes transparency, control, visible status and accessible mechanisms for people to inspect background actions.
The direction is clear. AI behavior is already leaking into the design-system problem.
The missing step is to make that explicit.
A conventional component contract might say: use this confirmation dialog for a consequential action; show this status treatment while work is in progress; use this AI label when generated content needs explainability.
An interaction policy layer asks a different set of questions: when may the AI present an action as available; when may it recommend one option rather than neutrally expose several; when may it execute without asking; when must it defer; when must uncertainty be visible; which actions require confirmation; what must be reversible and for how long; what condition triggers escalation to a person or another control?
These are not styling rules. They are reusable boundaries on behavior.
Organizations already have policies for many of these questions. The problem is where those policies live.
A finance team may have approval thresholds. Security has access rules. Legal has disclosure requirements. Product has UX principles. Risk has escalation criteria. Accessibility teams define how status and control must be exposed. Engineering implements tool permissions. A model or agent prompt may contain another version of the same intent.
The user experiences all of those systems at one point: the interaction.
If the policy is not represented there, the interface can accidentally contradict the organization behind it. A button can imply that an action is safe when the underlying system considers it exceptional. A fluent recommendation can make weak evidence look decisive. A background agent can complete a task without making its side effects legible. A confirmation dialog can appear so often that people stop reading it.
That last failure is not theoretical. Anthropic reported in 2026 that users approved roughly 93% of Claude Code permission prompts; the company describes how repeated prompts create approval fatigue and has moved toward automating safer approvals while relying more heavily on containment. OpenAI’s Codex work points in the same direction: its Auto-review system was designed to reduce synchronous human approvals dramatically while preserving review at boundary-crossing actions.
The lesson is not “remove confirmation.” It is more precise:
Confirmation is a scarce interaction. Spend it where a meaningful human decision still exists.
A design system is well positioned to encode that principle because it already sits between policy and repeated product implementation.
The useful unit is not another giant AI guideline. It is a small set of decisions that can travel with a capability, workflow or component.
What may the AI put in front of the user?
This includes generated claims, recommendations, available actions and representations of system state. An AI should not surface an action merely because a tool technically exists. Presentation itself communicates possibility and legitimacy.
What may the AI resolve on its own?
A system can rank options, recommend an option or make the decision. Those are different interaction contracts even when the final screen looks similar. The policy should state which class applies and what evidence moves an interaction from “recommend” to “decide.”
What may the AI actually do?
OpenAI’s Agents guidance distinguishes automatic guardrails from human approval before side effects such as cancellations, edits or sensitive tool actions. The product implication is that “execute” is not a generic capability. It is conditional behavior.
When should the AI deliberately not act?
Abstention is often discussed as a model behavior. In products, it is also an interaction state. A system needs a designed way to say: I have enough information to continue the workflow, but not enough authority or evidence to make this move.
When does uncertainty become part of the interface?
Microsoft’s agent UX principles explicitly argue that uncertainty is expected and that certainty and reasoning behind a recommendation should be visible or easily accessible to reduce overreliance. That does not mean attaching a dubious percentage to every response. It means designing a reliable signal for when the user’s interpretation should change.
Which actions need a human decision immediately before execution?
Confirmation should be based on consequence, reversibility, ambiguity and policy—not on a blanket rule that “AI actions need approval.” If every low-risk action interrupts the user, the product trains people to approve. If no action interrupts them, autonomy becomes invisible.
What must the system make recoverable?
Agentic products will make mistakes. A mature design system should specify recovery patterns as deliberately as success patterns: preview, commit, undo window, version history, rollback or compensating action. Reversibility changes what autonomy is safe to grant.
When does the interaction leave the agent’s normal path?
Escalation is not just an operational fallback. It is a user experience. The system should preserve context, state what blocked progress, identify what decision is needed and hand over without making the person reconstruct the case.
I argued in The next AI interface may never be seen that enterprise software needs a carefully designed callable surface: capabilities that agents can understand and invoke without inheriting the ambiguity of a human interface.
That article asks what should the organization make callable?
The interaction-policy question begins after a capability is callable.
Given that the system can do something, under which conditions should the AI present it, recommend it, execute it, defer it, expose uncertainty, ask for confirmation, make it reversible or escalate it?
The first problem is interface architecture between agents and organizational capabilities. This one is interaction governance between autonomous behavior and human experience.
They meet, but they are not the same problem.
The simplest implementation does not require a new design tool. Start with a matrix attached to the capability.
| Interaction decision | Policy question | Example rule |
|---|---|---|
| Present | May this action be offered here? | Show refund only for eligible orders |
| Decide | May AI resolve the choice? | Recommend above a validated evidence threshold; never decide exceptions |
| Execute | May AI create the side effect? | Auto-execute under €100 if all deterministic checks pass |
| Defer | When must AI stop? | Defer when required account data conflicts |
| Uncertainty | What must become visible? | Surface missing evidence that could change the recommendation |
| Confirm | When is approval meaningful? | Require explicit approval for irreversible external actions |
| Reverse | What recovery is required? | Keep a 30-minute undo window for reversible changes |
| Escalate | Who takes over, with what context? | Route policy exceptions with evidence and attempted actions |
The values will differ by domain. The structure is the point.
Once these decisions are explicit, they can be implemented in more than one place: design-system documentation, component APIs, agent instructions, tool policies, workflow engines, automated tests and runtime telemetry. The design system becomes the human-readable and product-facing expression of a policy that engineering can also enforce.
That matters because prompt-only policy is weak infrastructure. A product team should be able to test whether a generated interaction violated a boundary without asking another model whether the screen “looks right.”
This also changes component semantics.
Today a button often knows its visual variant, state and event handler. In an agentic system, the surrounding interaction may need metadata about consequence: action class; decision posture; authority source; confirmation requirement; recovery path; uncertainty state; and escalation target.
Not all of this belongs literally inside a UI component. Some belongs in the workflow or policy engine. But the design system should define the shared vocabulary and observable states so that product, design, engineering, risk and accessibility teams are not describing the same boundary five different ways.
That is the larger opportunity.
Design systems became powerful when they stopped being sticker sheets and became organizational infrastructure for repeated interface decisions. AI creates another class of repeated decision that is currently being solved locally in prompts, flows and product reviews.
Those decisions are too consequential to remain local.
The current wave of AI-ready design systems is rightly focused on whether machines can consume the system: can an agent retrieve the right token, component, pattern, documentation and code? Can it validate what it generated? Can it stay visually and technically on-system?
Add one more test:
Can the system tell the AI what kind of interaction is allowed here?
If not, the AI may become excellent at reproducing the surface of your product while improvising the rules underneath it.
That is a strange kind of consistency: every screen looks correct, every component passes, and the product still behaves in ways the organization never intended.
The next design-system layer should make that harder.
Not by turning designers into compliance officers. Not by putting a confirmation modal in front of every agent action. And not by pretending probabilistic systems can be made safe through UI rules alone.
By doing what design systems already do well: taking decisions that should not be reinvented in every feature and turning them into shared, testable infrastructure.
For AI, some of the most important of those decisions are no longer about what the interface looks like.
They are about what the interface is allowed to let intelligence do.
One email per essay. No noise between.
Berk Bayri
AI Transformation Advisor & Fractional Innovation Lead
I help leadership teams decide where AI belongs, test it before it scales and build the teams that run it, drawing on 25+ years of building digital products, experiences and capabilities.