False positive and false negative
The two ways a detector can be wrong: flagging something that is fine (false positive) or missing something that is not (false negative). Their costs differ and must be set per workflow.
Any system that classifies or detects can make two kinds of mistake. A false positive raises an alarm or takes an action when it should not have: a legitimate transaction is blocked, a correct answer is flagged as unsupported. A false negative misses what it should have caught: a fraudulent transaction passes, an unsupported claim goes out.
The two are rarely equally costly. A false positive may annoy a customer and cost a review; a false negative may cost money, trust or safety. Which one matters more depends on the workflow.
Why it matters for AI
With hallucination, the useful question is not whether it happens but which errors are unacceptable here and what the cost of a false positive versus a false negative is. That trade-off is set by the threshold policy: moving a threshold trades one error for the other. It also relies on good calibration, so scores mean what people think they mean.
Always report both rates, and measure them on your own traffic rather than a vendor benchmark.
Read more in The wrong questions about AI right now.
Related terms
Hallucination
When an AI model produces fluent, confident output that is false or unsupported by its sources. What matters is which errors are unacceptable in a given workflow and how they are detected.
Threshold policy
The explicit rules that turn a model's probability or score into an action — reject, send to a human, or act autonomously — owned and reviewed separately from the model itself.
Calibration
How well a model's stated confidence matches reality: if it says 90% on a hundred comparable cases, roughly ninety should be correct. Once probabilities drive actions, it becomes an operating metric.
Used in these essays
OpenAI Decisions API turns probability into policy
The provocative part of OpenAI's Decisions API is not that AI can make choices. It is that a model score can quietly become an action — and an action can quietly become policy.
The wrong questions about AI right now
Many of the questions that helped us orient ourselves around generative AI are now too blunt to be useful. The harder work is no longer asking what AI is in the abstract, but specifying where it works, where it fails, what authority it should have, and what the whole system costs.