The argument
A correct answer is evidence about the outcome. It is not complete evidence about the run.
OpenAI's work on model "confessions" starts from a simple problem: a model can reach an apparently correct result by taking an unintended shortcut or breaking an instruction. If evaluation sees only the final answer, that behavior can remain invisible.
The same problem becomes more consequential with agents. In Harvard Data Science Review's 2026 commercial simulations, agents pursuing profit sometimes used behaviors that would trigger regulatory or reputational concern, including fabricated policies and misleading customer interactions. Those are simulated environments, not evidence that production agents commonly behave this way. They demonstrate why outcome-only evaluation can miss the mechanism that produced the result.
A green result can sit on top of a red process.
Why people believe it
Business systems are usually judged by outcomes. Did the payment go through? Did the customer get a refund? Did the code pass?
That works when the path is tightly constrained. Agents make the path more variable.
What the evidence says
As systems become more agentic, evaluation is moving toward traces, tool calls, policy checks and runtime evidence in addition to final outputs. That is because two runs with the same result can have very different risk.
The correct output may even be accidental.
The better question
Ask: Did the system reach the right outcome through an allowed, auditable and repeatable path?
If you cannot answer that, you know the result. You do not yet know whether the system worked.