Calibration
How well a model's stated confidence matches reality: if it says 90% on a hundred comparable cases, roughly ninety should be correct. Once probabilities drive actions, it becomes an operating metric.
A model is well calibrated when its confidence means what people naturally think it means. If it reports "90%" across one hundred comparable cases, roughly ninety should be right. If only sixty are, the score may still rank cases reasonably, but any decision rule built on the number is operating on false premises.
For a long time calibration was mostly a research metric. The moment a system exposes probabilities that trigger actions, it stops being one. A miscalibrated 80% can mean unwarranted automation or needless escalation, and both cost money.
Using it in practice
Ask the plain question: when this system says 80%, does 80% mean roughly 80% on our real traffic? Measure on your own data, not a vendor benchmark, and re-measure after model updates, because calibration drifts when the data or the serving route changes.
Calibration is the foundation of any threshold policy and the reason that policy needs outcome feedback. Left unchecked, the gap between stated and actual confidence becomes decision debt. Where confidence is too low to trust, build in abstention.
Read more in OpenAI Decisions API turns probability into policy.
Related terms
Threshold policy
The explicit rules that turn a model's probability or score into an action — reject, send to a human, or act autonomously — owned and reviewed separately from the model itself.
Abstention
A deliberate option for an AI decision system to decline to decide and return uncertain or high-consequence cases to a person, rather than forcing every input into yes or no.
Decision debt
The maintenance burden created when an AI system's thresholds and decision boundaries are set once and never revisited, even as data, models, costs and risks change.