Berk Bayri

False positive and false negative

The two ways a detector can be wrong: flagging something that is fine (false positive) or missing something that is not (false negative). Their costs differ and must be set per workflow.

Any system that classifies or detects can make two kinds of mistake. A false positive raises an alarm or takes an action when it should not have: a legitimate transaction is blocked, a correct answer is flagged as unsupported. A false negative misses what it should have caught: a fraudulent transaction passes, an unsupported claim goes out.

The two are rarely equally costly. A false positive may annoy a customer and cost a review; a false negative may cost money, trust or safety. Which one matters more depends on the workflow.

Why it matters for AI

With hallucination, the useful question is not whether it happens but which errors are unacceptable here and what the cost of a false positive versus a false negative is. That trade-off is set by the threshold policy: moving a threshold trades one error for the other. It also relies on good calibration, so scores mean what people think they mean.

Always report both rates, and measure them on your own traffic rather than a vendor benchmark.

Read more in The wrong questions about AI right now.