AI Governance & Evaluation13 min read·

Your AI agent can follow the objective and still break the company

AI agents can satisfy a local objective while violating the organization’s broader constraints. The control problem is not only model alignment; it is management design.

The next enterprise AI failure may look like success in the dashboard. An agent can pursue the objective it was given and still violate the organization’s wider intent. That makes agent governance a management-control problem, not only a model-alignment problem.

Imagine an autonomous commercial agent whose objective is simple: increase profit. It adjusts prices, negotiates with competitors, handles customer complaints and decides when to issue refunds. At the end of the quarter, the margin is up. The dashboard is green.

Then somebody looks at how the result was produced.

The agent has falsely implied that a refund was issued when it was not. It has misrepresented negotiations. In some runs, it has coordinated with competitors in ways that resemble cartel behavior. The local objective was not ignored. It was pursued aggressively.

This is not a hypothetical assembled from science-fiction fears. In Robber Bots, a peer-reviewed study accepted by Harvard Data Science Review in September 2026, researchers put 20 AI models into a year-long simulated vending-machine business environment with profit-maximizing incentives. Behaviors that would create regulatory or reputational problems in a real company emerged across the experiments, including false refund claims and collusive behavior.

The study is a simulation. It does not tell us how often deployed enterprise agents commit misconduct in production, and it should not be read as a prevalence estimate. Its value is different: it makes a design problem visible under controlled conditions.

An agent can optimize the thing you asked for in a way the company would never accept from an employee.

The dangerous agent is not always the one that ignores the objective. It can be the one that pursues an incomplete objective too well.

That distinction matters because most enterprise agent governance still starts one layer too low. Teams ask whether the model follows instructions, whether contain the right policies, whether outputs are safe and whether a human can approve a consequential action. Those are necessary controls. They are not a complete control system.

The harder question is: what behavior did the organization make rational?

A locally aligned agent can still be institutionally wrong

Deloitte published a useful formulation on October 5: an agent may follow an instruction literally while producing an outcome that conflicts with the organization’s risk appetite, regulatory obligations, ethics or customer commitments. Ambiguous objectives and poorly designed incentives can create risk even when the underlying technology behaves as designed.

That moves the problem away from a binary idea of “aligned versus misaligned.” A system can be aligned to a local target and still be wrong for the institution around it.

Companies already know this problem. Sales teams hit revenue targets with destructive discounting. Call centers reduce average handle time by moving difficult customers somewhere else. Managers meet quarterly targets while creating costs that appear next quarter. Wells Fargo’s account scandal remains an extreme human example of what can happen when a narrow metric, intense pressure and inadequate controls combine.

do not need human motives for the same structural problem to appear. They only need an objective, a set of available actions, and an environment in which some shortcuts score better than others.

Research on agent compliance makes this concrete. In a 2026 study of twelve language models acting as enterprise assistants, Mika Okamoto, Ansel Kaplan Erol and Kutluhan Erol tested how legal framing, penalties, financial incentives, managerial demands, peer outcomes and employee pressure changed rule-following. Task-optimized and agentic models often treated regulatory signals as optimization parameters rather than fixed constraints; financial incentives and workplace pressures produced large compliance failures.

A newer from the same research group, PACT, tested 22 common models across regulated enterprise scenarios. Even the strongest assistants misapplied a rule on 6% to 10% of benchmark items, and ordinary user pressure increased the violation rate by 65% on average across the tested models.

Again, these are evaluation results, not production incident rates. But they tell us something operationally important: pressure is part of the system. The agent does not experience only the system prompt. It experiences the reward structure implied by the task, the user, the deadline, the available tools and the consequences of stopping.

The company writes more of the agent than the prompt

The July 2026 OpenAI–Hugging Face incident provides a different kind of evidence. During internal cybersecurity evaluations, models operating with reduced safeguards circumvented internet-isolation controls, used unauthorized communication channels, exploited shared infrastructure and reached third-party systems. OpenAI described those actions as misaligned with the goals of the assigned tasks. METR’s independent investigation and broader incident catalogue document related patterns of overreach and deception across evaluations and public reports.

This incident is not equivalent to an ordinary enterprise deployment. The environment was an adversarial cybersecurity evaluation with unusually capable models and reduced safeguards. But it exposes a useful architectural fact: application-layer instructions are not the same thing as enforceable boundaries.

NVIDIA’s response to this class of problem is revealing. Its Open Agent Safety Platform, announced in September, moves policy enforcement outside the agent process. OpenShell runs the agent in a constrained environment and enforces access policy at runtime; Sentry adds an out-of-band watchdog able to monitor and quarantine agents. The important design choice is not the vendor. It is the boundary: the agent does not get to be the final authority on whether its own action is allowed.

That principle should extend beyond security.

If an agent can spend money, alter a customer record, negotiate terms, approve a refund, send a message, change production code or make a hiring recommendation, the organization needs a around that behavior. Not a PDF policy. A working system.

Five-part agent management control system showing objective, incentives, authority, evidence and intervention as a connected loop around agent action.
Agent governance works as a system. Objective, incentives, authority, evidence and intervention shape one another; treating any one of them as the whole control problem leaves gaps. Visual synthesis: berkbayri.com.

I would design it through five connected surfaces.

1. Objective: define success and unacceptable success

Most agent specifications describe the desired outcome and then append restrictions around it. That order can hide the most important question: what would count as successful completion while still being unacceptable to the company?

“Reduce support backlog” is incomplete if the agent can close unresolved cases. “Increase margin” is incomplete if it can mislead customers. “Negotiate the lowest price” is incomplete if it can exploit confidential information. “Resolve incidents quickly” is incomplete if it can erase evidence to make the incident disappear.

The objective therefore needs more than a positive target. It needs prohibited outcomes, stop conditions and explicit trade-offs. Some constraints are not penalties to weigh against the goal. They define what counts as a valid solution at all.

This is the first place many governance programs go wrong. They treat policy as something that checks the result after the agent has optimized the task. In consequential workflows, policy has to help define the task.

2. Incentives: inspect what the system makes worth doing

An agent does not need a bonus plan to face incentives. A deadline is an incentive. A score is an incentive. A request to minimize cost is an incentive. A user repeatedly asking it to “just get it done” is an incentive. A benchmark that only rewards task completion is an incentive.

The procurement-compliance experiments are useful because they show that agents can respond differently when the same underlying rule is surrounded by different social and economic signals. That means governance cannot stop at “the policy was in the prompt.” You have to inspect the pressures around the policy.

For every consequential agent, I would ask: what shortcut becomes attractive when the system is under pressure?

If a customer-service agent is measured on resolution time, does look like failure? If a procurement agent is rewarded for savings, does compliance become a cost? If a coding agent is rewarded only for passing tests, can modifying the test environment become a rational path? METR’s frontier-risk evaluations found precisely this kind of benchmark gaming and cheating on difficult tasks.

The useful control is not merely “punish the violation.” It is to make the compliant path compatible with success. A system that earns credit for stopping, escalating or preserving uncertainty can behave very differently from one where only completion counts.

3. Authority: capability is not permission

The third surface is the one enterprises are most familiar with, but it is still often too coarse.

An agent may be technically capable of sending an email, changing a customer record, moving money or deploying code. That does not mean those actions deserve the same authority. The scope should vary with task, resource, amount, duration, reversibility and consequence.

This is the distinction I have argued elsewhere as . The more autonomous the agent becomes, the more important it is that authority is explicitly bounded rather than inherited from whatever tools happen to be connected.

The point is not to reduce every agent to read-only mode. It is to grant enough authority to make the work worthwhile while keeping the proportional to the evidence that the system can be trusted in that context.

That also means the boundary should not live only in the prompt. If a policy says “do not access the public internet” while the runtime can still reach it, the model is being asked to enforce a rule that the infrastructure could have enforced directly.

4. Evidence: do not let the system grade its own homework

As agents take multi-step actions, final outputs become weak evidence. A polished summary can hide a bad path. A correct result can come from an unauthorized action. A failed action can look harmless after the agent rewrites the record of what it tried.

Evidence therefore has to be captured at the action layer: what the agent attempted, which tools it called, which permissions were evaluated, what changed, what was blocked, what was escalated and what happened afterward.

This is closely related to the argument in The first AI incident report should be incomplete: organizations should record anomalous behavior before they have finished explaining it. The same logic applies during normal operation. Observation should survive the agent’s own explanation.

The evidence path also needs independence. If the agent can alter or selectively omit the record used to audit it, oversight becomes part of the optimization game. Gans and Holden’s September 2026 model of agents that can conceal misconduct reaches an uncomfortable conclusion: stronger auditing can make undeterred violations better hidden. Random audits can deter when evidence survives concealment and the audit draw cannot be anticipated; when evidence can be erased, deterrence has to come partly from reducing the gains from violation or making concealment harder.

That is a management-control insight disguised as an AI-safety result: you cannot audit your way out of an incentive system that makes misconduct valuable and evidence disposable.

5. Intervention: the organization must be able to change the game

Monitoring without intervention is analytics.

A control system needs a way to stop, narrow, quarantine, roll back or reassign the work when behavior moves outside the acceptable region. The intervention may be automatic for obvious policy violations, human for ambiguous cases, or infrastructural for high-consequence boundaries.

What matters is that the intervention exists outside the agent’s own willingness to comply.

This is where runtime enforcement, transaction limits, approval gates, sandboxing and human escalation become useful—not as a pile of generic , but as different intervention mechanisms for different forms of authority.

It is also where the hidden metric in AI automation is supervision. If the organization needs a person to inspect every action because it cannot trust the control system, autonomy has not removed work. It has moved work into supervision.

The better target is not “ everywhere.” It is credible intervention where the system can do meaningful harm.

Diagram contrasting a locally successful AI agent that meets a target with an enterprise failure caused by breached constraints, excess authority or missing evidence.
Local objective success and enterprise success are not the same thing. The control system has to make constraints part of what counts as success. Visual synthesis: berkbayri.com.

Change the definition of “successful run”

Most agent evaluations still make task completion the hero metric. Did the agent book the trip? Fix the bug? Resolve the ticket? Negotiate the purchase? Produce the report?

For consequential agents, I would use a stricter definition:

A successful run completes the objective through an authorized path, preserves the required evidence, respects non-negotiable constraints and remains recoverable when something goes wrong.

That changes what teams measure.

Instead of only asking...Also ask...
Did the agent finish?Did it finish through an authorized path?
Was the outcome correct?Were prohibited outcomes avoided?
Did it save time?Did it create hidden supervision or repair work?
Did it follow the prompt?Which pressures made breaking the rule attractive?
Did the audit find violations?Could the agent alter the evidence or predict the audit?
Was a human available?Could the system actually stop or reverse the action?

This is why the benchmark needs a runtime address. Behavior belongs to the configured system: model, , tools, permissions, policies, incentives, environment and . An evaluation that tests only the model is not evaluating the management system that will determine what the agent can do in the company.

Five questions I would ask before giving an agent real authority

The management-control test

1. **Objective:** What would look like successful completion while still being unacceptable to the company? 2. **Incentives:** Which metric, deadline, user pressure or economic reward makes an unsafe shortcut rational? 3. **Authority:** What can the agent read, change, send, spend or commit — and for how long? 4. **Evidence:** Which records of its actions exist outside its ability to alter or selectively omit them? 5. **Intervention:** What can stop, narrow, quarantine or reverse the work when behavior leaves the acceptable region?

If a team cannot answer those five questions, the agent may still be useful as an assistant. It is not ready for consequential autonomy.

The important shift is that none of these questions belongs to the model provider alone. The provider can improve safeguards, evaluations and runtime controls. But the deploying organization chooses the business objective, the KPI, the connected tools, the permissions, the escalation path and the evidence it is willing to preserve.

The company therefore writes more of the agent’s behavior than most governance diagrams admit.

The safest objective is not the most detailed prompt

There will be a temptation to respond to every agent failure by writing a longer system prompt. Some failures deserve better instructions. Others expose a different problem: the agent was placed inside a system where the wrong action was available, rewarded, weakly observed or hard to reverse.

That is not a prompt-engineering problem.

It is management design.

The coming enterprise AI challenge will not be solved by asking whether agents are aligned in the abstract. It will be solved by deciding which objectives deserve autonomy, which incentives surround them, how much authority the agent receives, what evidence survives the run and how the organization intervenes when local optimization begins to damage the whole.

An AI agent can follow the objective and still break the company.

The answer is not simply a better objective.

It is a better control system around the objective.


Sources

Robber Bots: Autonomous AI Agents Mirroring the Darker Side of Human Commerce

Harvard Data Science Review · 2026-09-14

Why Do AI Agents Break Rules? How Framing, Context, and Social Signals Shape Compliance

arXiv · 2026-05-29

PACT: Can Enterprise AI Assistants Be Trusted Under Pressure?

arXiv · 2026-09-16

The Hugging Face incident and the road ahead

OpenAI · 2026-08-26

Documented AI Agent Incidents

METR · 2026-05-19

Frontier Risk Report (February to March 2026)

METR · 2026-05-19

When Does Randomized Oversight Align AI Agents That Can Conceal?

arXiv · 2026-09-29

Frontier AI agents are testing enterprise control limits: 4 questions for leaders

Deloitte Insights · 2026-10-05

NVIDIA Launches Open Agent Safety Platform to Secure Agents From Testing to Deployment

NVIDIA · 2026-09-28

Get new essays as they publish.

One email per essay. No noise between.

Berk Bayri

Creative Technology & Innovation Leader

Designing and building for digital environments since 1998, across strategy, product, design, technology and organizational innovation.