AI Governance & Evaluation12 min read·

An AI agent should get authority per task, not permissions per role

A person's role describes standing access across many situations. An AI agent should receive a narrower authority envelope for the task it is executing now.

A person's role describes standing access across weeks, projects and situations. An executing one task does not need that entire surface. The safer and more useful unit is a temporary authority envelope tied to the work being done now.

An employee in finance may legitimately have access to purchase orders, vendor records, invoices, budgets and payment systems.

That does not mean an AI agent asked to create one purchase order should inherit all of it.

The difference sounds obvious when stated that way. In implementation, it is easy to erase.

We already know how to give software identities roles, groups, scopes and application permissions. The convenient move is to create an agent identity, attach the permissions its owner or team commonly needs, and let the model decide which capability to use at runtime.

That makes the agent operational quickly.

It also turns a role into an authority surface.

A role answers a broad organizational question: what might this person or service legitimately need to do across its responsibilities?

A task asks something much narrower: what does this agent need to do for this request, on these resources, under these conditions, for this amount of time?

For autonomous software, those should not be the same boundary.

The role should define the ceiling. The task should define the active authority.

Roles are inventories. Agent work is transactional.

Human access systems grew around a practical assumption: people move across many situations during a day, and constantly re-authorizing every small action would make work impossible.

So we give people standing permissions.

A manager can approve expenses because approving expenses is part of the job. A support lead can issue refunds because refunds are part of the role. An engineer can deploy to production because deployment is part of the responsibility.

Those permissions are already imperfect. They accumulate. People change teams. Temporary access becomes permanent. turn into entitlements.

Agents add a second problem.

They can move faster across the permission surface than a person, chain actions across systems and select tools dynamically. They may also be executing a request whose legitimate scope is much smaller than the permissions of the user, team or service identity behind them.

This is why identity work for agents is converging on a useful distinction.

's 2026 concept paper on software and AI agent identity separates identification, authorization, access delegation, logging and data-flow provenance as distinct concerns. Microsoft describes stable agent identity alongside scoped permissions, delegated access, policy checks and auditing. AWS's Agentic AI guidance goes further by separating service identity from transaction identity, with per-operation constraints carried through short-lived credentials and session policies.

The technical mechanisms differ. The architectural direction is similar:

identity can be durable while authority is temporary.

That is a much better model for agentic work than "act as me."

The authority envelope

I would treat every consequential agent task as creating an authority envelope.

The envelope is not the agent's full capability set. It is the subset of capability that becomes valid for one piece of work.

Five dimensions make the boundary concrete.

1. Action

What exact verbs are allowed?

Read is not update. Draft is not send. Recommend is not approve. Create is not delete. Calculate a refund is not issue a refund.

An agent that can technically call all of those tools does not need all of those verbs active for every run.

2. Resource

Which objects may those actions affect?

"Can update CRM records" is broad. "Can update the contact record attached to case 4817" is different.

The same applies to files, repositories, accounts, cloud resources, customer records, invoices and messages. Resource scope turns a generic permission into a bounded transaction.

3. Amount or consequence

How much may the action change?

Money makes this obvious: refund up to €100; create a purchase order up to €5,000; never change bank details.

But the limit can describe other consequences too: one customer record rather than a segment, one branch rather than the repository, one campaign rather than the ad account, one external recipient rather than a mailing list.

Authority should scale with consequence.

4. Duration

When does the authority expire?

A task that needs elevated permission for 12 minutes should not create a credential that survives for 12 months.

AWS explicitly recommends short-lived credentials, session policies and just-in-time elevation for agent workflows. Microsoft similarly describes time-bound access and lifecycle governance for agent identities. The implementation details are platform-specific, but the principle is general: privilege should disappear when the reason for it disappears.

Duration can also be tied to workflow state, not only a clock. The grant ends when the ticket closes, the deployment completes, the approval is denied or the user cancels the task.

5. Reversibility

What recovery condition makes this authority acceptable?

Reading a document and deleting a document can use the same identity and still deserve radically different policy.

A reversible change may be allowed automatically. An irreversible external side effect may require approval. A high-consequence action may be unavailable until a prerequisite is satisfied.

OpenID's AuthZEN work is useful here because its approval profile treats consent, delegated authority, attestation, risk assessment and additional justification as prerequisites that can be satisfied before an authorization decision is re-evaluated. OpenAI's agent guidance makes a similar product-level distinction: automatic checks can handle some boundaries, while sensitive side effects can pause for approval.

The point is not that every action needs a human.

It is that recovery changes what can safely be delegated.

Diagram showing a stable agent identity under a role permission ceiling and a smaller task authority envelope defined by action, resource, limit, duration and reversibility.
The role defines the maximum eligible surface. The current task defines the much smaller authority that is active now. Visual synthesis: berkbayri.com.

A role can still matter. It should become the maximum, not the runtime grant.

This is not an argument for abandoning role-based access control.

Roles remain useful for organizational eligibility.

The finance agent should not be able to request a production-deployment permission simply because a asks for it. The customer-service agent should not be able to obtain payroll access because the model found a creative path through connected tools.

A role can define the maximum authority that may ever be issued.

The task then determines what is actually issued inside that ceiling.

Think of the structure as an intersection:

active authority = role ceiling ∩ user authority ∩ task envelope ∩ policy conditions

The agent may reason about what action would help. It should not be the final authority on whether that action is allowed.

Microsoft's current guidance states this directly in practical terms: the agent can reason about what to do next, while deterministic application, identity and policy controls decide whether the action is permitted. AWS's recent guidance makes the same separation by enforcing authorization in infrastructure and downstream services rather than asking agent code to be the gatekeeper.

This distinction matters because prompts are not access-control systems.

A system message that says "only edit files related to this request" can improve behavior. It does not make unrelated files inaccessible.

The stronger design is to make unrelated files unavailable to the transaction.

Do not ask the agent to remember its own boundary

Agent systems are often designed as if the model can safely carry all policy context in the same reasoning loop that is trying to complete the task.

That is fragile for a simple reason: the agent has an objective.

If authorization lives only in the instructions, then the component trying to finish the work is also being asked to interpret the limits on finishing it.

Infrastructure should remove that conflict where it can.

A purchase-order agent can receive a valid only for the relevant vendor and budget code. A coding agent can receive write access to one branch but not production. A support agent can receive permission to refund one order up to a threshold without gaining the ability to alter account ownership or payment details.

The model still plans.

The authorization system narrows the possible plan space.

That is the operational meaning of capability is not authority.

The agent can know a tool exists without receiving permission to use it in this run.

Approval should widen the envelope, not replace it

Human approval is often used as the universal answer to agent risk.

It is useful. It is also easy to misuse.

If the agent holds broad credentials and a person merely clicks "approve" before execution, the underlying authority surface may still be much wider than the decision the person thinks they are making.

A better pattern is to use approval to change the envelope.

Before approval:

  • read order 4817;
  • calculate eligible refund;
  • prepare recommendation.

After approval:

  • issue refund for order 4817;
  • maximum €84.20;
  • valid for ten minutes;
  • no other account changes;
  • preserve transaction evidence.

The human is not approving "the agent."

The human is authorizing a specific expansion of a specific task.

That makes approval more legible, more auditable and easier to revoke.

It also reduces approval fatigue because low-risk actions can operate inside their existing envelope while boundary-crossing actions request a narrow elevation.

Approval should change authority

A useful approval says which action becomes allowed, on which resource, to what limit, for how long and under what recovery condition. "Allow agent" is usually too broad to be a meaningful decision.

Task-scoped authority reduces blast radius without making agents useless

Security discussions often collapse into a false choice.

Either the agent has enough permissions to be useful, or it is so constrained that a person has to do the work anyway.

offers another path.

The agent can receive powerful permissions. They are simply powerful inside a narrow envelope.

Consider three examples.

Role permissionCurrent task authority
Support lead can manage customer cases and refundsRefund order 4817 up to €84.20; expires when the case closes
Engineer can write to multiple repositories and deployModify branch fix/checkout-482; no production deployment; expires after merge or 60 minutes
Procurement manager can create vendors and purchase ordersCreate one PO for approved vendor X up to €5,000; cannot edit bank details; expires after submission

The right-hand side is not necessarily read-only or timid.

It is specific.

That specificity is what shrinks the .

In Your AI agent can follow the objective and still break the company, I argued that authority is one part of a wider around agent behavior. Task-scoped authorization is what makes that authority operational rather than rhetorical.

A policy document can say "agents should use least privilege."

An authority envelope can tell the runtime what that means for this action now.

The audit record should capture what was allowed, not only what happened

Agent audit trails usually focus on action history: which tool was called, which record changed, which output was produced.

That is necessary but incomplete.

For consequential work, I would also preserve the authorization context that existed when the action happened:

  • agent identity;
  • originating user or service;
  • task identifier;
  • requested action;
  • allowed action;
  • resource scope;
  • amount or consequence limit;
  • issue and expiry time;
  • approval or policy prerequisite;
  • recovery path;
  • policy version used for the decision.

That record answers a different question from "what did the agent do?"

It answers why was the agent allowed to do it?

NIST explicitly includes logging, transparency and the linkage between agent actions and non-human identity in its agent identity work. AWS's transaction-identity model similarly carries request context through short-lived credentials. These are not identical implementations, but they point toward the same requirement: authority has to be reconstructible after the fact.

If you cannot reconstruct the grant, you do not really have delegated authority.

You have credentials.

This is the boundary between identity and operating design

The security team can implement scoped tokens.

It cannot decide, by itself, what the business meaning of "€5,000", "one customer", "reversible for 30 minutes" or "requires legal approval" should be.

That is why agent authorization is becoming an operating-design problem as much as an IAM problem.

Someone has to define which decisions belong to which tasks. Someone has to decide which consequences deserve . Someone has to specify which actions can be reversed and for how long. Someone has to determine whether the authority attached to a workflow is still appropriate when the workflow changes.

The technical controls make the boundary enforceable.

The organization still has to design the boundary.

This is also why the next AI interface may never be seen. Once software capabilities become callable by agents, tool design, identity, permissions and recovery stop being backend details. They become part of the product contract between the organization and autonomous execution.

A practical authorization test

Before letting an agent take a consequential action, I would require five answers.

Comparison of broad standing role permissions with narrow task-scoped authority, showing a smaller action surface, shorter duration and explicit recovery path for the task-scoped model.
Standing permissions optimize for repeated human work. Task authority optimizes for bounded autonomous execution. Visual synthesis: berkbayri.com.
  1. Action: What exact verb is this run allowed to perform?
  2. Resource: Which specific object, account, record, environment or dataset may it affect?
  3. Limit: What amount or consequence boundary applies?
  4. Duration: When does the authority expire automatically?
  5. Reversibility: What can undo, stop or contain the action if the agent is wrong?

If the answer is "whatever the user's role already allows," the authorization boundary is probably too broad.

If the answer is "nothing without asking a person," the system may be too narrow to create meaningful autonomy.

The design goal is neither maximum permission nor maximum restriction.

It is enough authority for the task, and no more authority than the task can justify.

The permission model should follow the unit of delegation

The deeper mistake is treating an AI agent like a new employee account.

A person has a role. The role persists. The person moves between tasks inside it.

An agent can be instantiated around a task, a workflow or a delegated goal. Its active authority can therefore follow the same unit.

That gives us a cleaner architecture:

  • stable identity for accountability;
  • role or policy ceiling for eligibility;
  • task-scoped authority for execution;
  • deterministic enforcement outside the model;
  • evidence for every elevation;
  • automatic expiry when the work ends.

This does not eliminate authorization complexity.

It puts the complexity in the right place.

The question for an autonomous system should not be, "Which role does this agent belong to?"

It should be:

What authority does this task deserve right now?

Sources

New Concept Paper on Identity and Authority of Software Agents

NIST · 2026-02-05

Identity for AI agents

Microsoft Learn · 2026-06-16

Microsoft Entra security for AI overview

Microsoft Learn

Implement least privilege with dynamic boundaries

AWS Well-Architected Agentic AI Lens

Propagate user authorization context in AI agents with Amazon Bedrock AgentCore

AWS Security Blog · 2026-08-19

OpenID Foundation advances authorization for the agent era with new AuthZEN Working Group Drafts

OpenID Foundation · 2026-06-15

Guardrails and human review

OpenAI · 2026-10-07

Get new essays as they publish.

One email per essay. No noise between.

Berk Bayri

AI Transformation Advisor & Fractional Innovation Lead

I help leadership teams decide where AI belongs, test it before it scales and build the teams that run it, drawing on 25+ years of building digital products, experiences and capabilities.