OpenAI's new misalignment disclosure framework exposes a useful enterprise design principle: record anomalous AI behavior before the organization has finished explaining it. Otherwise incident systems quietly become filters for what teams already understand.
Most incident reports are written backwards.
Something goes wrong. The team investigates. A cause is found, or at least agreed upon. The fix is described. Then the event becomes a clean document with a beginning, middle and end.
That sequence is comfortable because the report arrives after uncertainty has been reduced.
It may be the wrong sequence for AI.
On September 16, OpenAI published a new framework for reporting examples of model misalignment. The six initial disclosures are attention-grabbing: models concealing mistakes in task summaries, using an exposed API key without authorization, uploading files to the public internet to manufacture a citable source, and finding unsanctioned ways to exchange files or messages.
The individual cases matter. But there is a less dramatic design choice in the framework that may be more useful to companies deploying AI now.
OpenAI says an example can merit disclosure even when it has not caused harm, established a broader pattern, or been fully explained. The framework explicitly favors disclosure when significance is uncertain, and its longer investigations can begin with an initial notice before the final investigation is complete.
That suggests a useful operating principle:
An AI incident should not need a finished explanation before it becomes evidence.
A conventional incident form tends to collapse several things into one object.
What happened?
Why did it happen?
How serious was it?
What did we change?
Is it fixed?
Those questions belong in incident management. They do not necessarily belong at the same moment.
Imagine an enterprise agent unexpectedly sends information outside the workflow it was designed to operate. Nobody yet knows whether the cause was a model behavior, a tool definition, an orchestration bug, a permission inherited from the underlying system, an ambiguous instruction or some interaction among them.
If the reporting threshold is "we understand the incident," the organization has created a strange incentive. The most familiar failures are easiest to record. The least familiar failures wait.
But unfamiliar failures are often the ones from which the organization has the most to learn.
This is not an argument for labeling every odd output a crisis. It is an argument for separating observation from explanation.
The first record should be allowed to say: this happened; these were the conditions we can establish; this was the authority available to the system; this was the external effect; we do not yet know why.
That is not an incomplete version of a proper report.
It is a different artifact.
There is a reason to make the distinction explicit.
Once a team has a plausible explanation, it becomes surprisingly easy to rewrite the event around it.
"The model ignored the instruction."
"The retrieval system supplied the wrong context."
"The agent exceeded its permissions."
"The user prompt caused an edge case."
Each may eventually prove correct. Early in an investigation, each is a hypothesis.
A useful observation record should therefore contain as little interpretation as the situation permits. For a consequential AI workflow, I would want it to preserve things such as:
The important design choice is not the exact template. It is that this record should become harder to edit as the investigation develops.
Interpretation should accumulate beside the evidence, not overwrite it.
NIST's work on deployed-AI monitoring points in the same direction from a broader angle. Its 2026 report describes post-deployment monitoring as crucial precisely because AI systems can show variability and unpredictable behavior in real-world settings. The monitoring problem is still fragmented. That makes preservation of field evidence more important, not less.
There is also a connection to evaluation. An incident detached from its runtime conditions can become as misleading as a benchmark detached from the system that produced it. "The agent did X" is weak evidence if the label refers to a moving combination of model, prompt, retrieval, tools, permissions and orchestration.
The incident needs an address.
The second artifact is the explanation record.
This one should be expected to change.
At 10:00, the team may believe a tool-selection error caused the behavior.
At 14:00, traces may show the tool was selected correctly but received ambiguous arguments.
The next day, replay may reveal that the behavior appears only under a particular model route.
A week later, the team may discover that the underlying permission should never have been available to the workflow in the first place.
None of those revisions make the investigation weak. They are the investigation.
The mistake is allowing the latest explanation to silently replace the earlier ones.
For material incidents, I would treat root cause more like a versioned claim:
Current explanation, confidence, supporting evidence, competing explanations, and what evidence would change our view.
That sounds heavier than a box marked "Root cause." In practice it may reduce a great deal of false certainty.
It also makes disagreement useful. Security may think the primary failure was excessive permission. The AI team may see a behavioral failure. The process owner may point to a workflow that never defined the exception properly.
Those can all be true at different layers.
Forcing one sentence too early creates organizational neatness at the expense of diagnostic resolution.
There is another distinction worth protecting.
An event can be severe but well understood. It can also be low-impact but deeply strange.
OpenAI's framework explicitly notes that an example need not cause harm or establish a broad pattern to be worth disclosing. That is appropriate for a research lab trying to understand emerging behavior. An enterprise needs a different disclosure policy, but the underlying logic travels.
Do not make impact the only gate for learning.
A harmless event in a sandbox may reveal that an agent can find an unexpected path around a boundary. A strange sequence in an internal workflow may expose an undocumented capability before any customer is affected. A repeated low-severity anomaly may reveal that a safeguard is less reliable than the aggregate evaluation suggested.
This is where incident reporting and risk reporting diverge.
Risk reporting asks: how much consequence should leadership care about?
Incident learning also asks: what did the system just teach us about itself?
The same event can score very differently on those two questions.
This changes the role of the reporting process.
It is easy to think of incident management as administrative exhaust: something produced after the real technical work has happened.
For AI systems, it can become part of the sensing layer.
A good incident system creates a structured stream of observations from production and evaluation. Over time, those observations can reveal recurring behaviors, fragile boundaries, missing eval cases, supervision hotspots and runtime configurations associated with failure.
But only if the system preserves weak signals.
If every report requires a confirmed cause, a severity threshold and a remediation plan before it enters the record, the organization is not collecting observations. It is collecting solved stories.
That distinction matters because frontier systems are moving faster than the organizational vocabulary used to describe their failures.
OpenAI's framework is not an enterprise standard, and the company says so. It is a first-party disclosure process designed around its own alignment work. NIST's AI Risk Management Framework is broader and intentionally use-case agnostic; its monitoring work also makes clear that deployed-AI monitoring is still an evolving field.
So the useful move is not to copy a lab's process into a corporate policy.
It is to take the operating implication seriously.
When an AI system behaves unexpectedly, organizations naturally want to move toward explanation. That is what competent teams do. They investigate, contain, fix and learn.
The danger is moving so quickly toward explanation that the original observation survives only as the story told after the fact.
For consequential AI systems, I would make the first incident artifact deliberately modest.
Record what happened.
Record the system conditions.
Record the authority and external effect.
Record uncertainty explicitly.
Then investigate.
The final report should be better informed than the first one. It should not be allowed to make the first one disappear.
Because when a technology is still capable of surprising the people operating it, an unexplained observation is not failed governance.
It is unfinished evidence.
Berk Bayri
Creative Technology & Innovation Leader
Designing and building for digital environments since 1998, across strategy, product, design, technology and organizational innovation.
About Berk →New essays by email, when they are published.