Berk Bayri

Calibration

How well a model's stated confidence matches reality: if it says 90% on a hundred comparable cases, roughly ninety should be correct. Once probabilities drive actions, it becomes an operating metric.

A model is well calibrated when its confidence means what people naturally think it means. If it reports "90%" across one hundred comparable cases, roughly ninety should be right. If only sixty are, the score may still rank cases reasonably, but any decision rule built on the number is operating on false premises.

For a long time calibration was mostly a research metric. The moment a system exposes probabilities that trigger actions, it stops being one. A miscalibrated 80% can mean unwarranted automation or needless escalation, and both cost money.

Using it in practice

Ask the plain question: when this system says 80%, does 80% mean roughly 80% on our real traffic? Measure on your own data, not a vendor benchmark, and re-measure after model updates, because calibration drifts when the data or the serving route changes.

Calibration is the foundation of any threshold policy and the reason that policy needs outcome feedback. Left unchecked, the gap between stated and actual confidence becomes decision debt. Where confidence is too low to trust, build in abstention.

Read more in OpenAI Decisions API turns probability into policy.