ejuarezg.com / Writing / RSS
Probability estimates are useful only when their meaning survives contact with a decision. A model that reports confidence should be correct about eight times in ten on comparable examples. Calibration measures this agreement [][].
For a confidence score and threshold , define a
Click the reference to selective predictor whenever the term appears. Nota keeps its definition available in context instead of forcing the reader to search backward through the paper.
For predictions partitioned into 6 confidence bins, expected calibration error is
The formula is static and readable without JavaScript.
Raising removes low-confidence cases. The resulting relationship is structural rather than empirical:
The small table in Figure 1 illustrates Theorem 1. Its accuracy values are invented for this demonstration, while the decline in coverage follows directly from the theorem.
| Threshold | Coverage | Accuracy on answered cases |
|---|---|---|
| 0.50 | 100% | 78% |
| 0.65 | 79% | 86% |
| 0.80 | 42% | 93% |
The operating point is therefore a policy choice, not merely a model metric. In a low-cost recommendation system, broad coverage may matter most. In a high-consequence workflow, abstention can create room for review, a second model, or a request for more information.
A confidence score becomes useful when it controls an explicit action. Even this tiny example benefits from a paper's structure: the abstract states the claim, selective predictor fixes the terminology, Theorem 1 separates a guarantee from the toy observations, and the bibliography preserves provenance. Those relationships are what Nota adds beyond styling alone.