Perfectly Truthful Calibration Measures
- The paper introduces perfectly truthful calibration measures that guarantee the true probability report minimizes the expected empirical calibration error, addressing a finite-sample incentive problem.
- It details batch constructions like Averaged Two-Bin Calibration Error (ATB) and UBSEs that ensure truthfulness, completeness, and continuity, linking calibration metrics to proper scoring rules.
- The framework extends to multiclass and sequential settings, incorporating decision-theoretic and testability considerations to maintain optimality of truthful reporting in practical forecasting.
Perfectly truthful calibration measures are calibration error functionals whose expected empirical value is minimized when a predictor reports the ground-truth probabilities. They address a finite-sample incentive problem that is distinct from perfect calibration itself. Perfect calibration requires conditional unbiasedness—binary predictions satisfy , and multiclass predictions satisfy —but finite-sample evaluation can still reward strategic misreporting under many standard metrics. Recent work therefore separates the semantic notion of calibration from the incentive-theoretic notion of truthfulness, and studies exact batch constructions, approximate sequential analogues, and the interaction of truthfulness with continuity, soundness, completeness, testability, and decision-theoretic actionability (Hartline et al., 18 Aug 2025).
1. Definition and formal scope
In the binary setting, perfect calibration means that for every reported value in the image of a predictor, ; in the induced prediction-label distribution language, a distribution over is perfectly calibrated if (Błasiok et al., 2022). In the multiclass setting, perfect calibration requires that for all and all classes , (Lu et al., 7 Oct 2025).
Truthfulness is a different property. In the batch formulation on sequences, with arbitrary ground-truth probabilities 0, independent labels 1, and an alternative report 2, a calibration measure 3 is truthful if
4
The corresponding sample-based formulation requires that for any data-generating distribution 5, the expected empirical calibration error is minimized by the ground-truth predictor 6 (Hartline et al., 18 Aug 2025). In this literature, “perfectly truthful” refers to exact minimization in expectation, whereas several sequential works study only constant-factor or logarithmic-factor analogues (Haghtalab et al., 2024).
A separate but related line studies a ground-truth notion of distance from calibration. The unifying formulation defines
7
the 8 distance from a predictor to the nearest perfectly calibrated predictor, and calls a calibration measure consistent if it is polynomially related to 9 (Błasiok et al., 2022). This notion concerns fidelity to the calibration property itself rather than incentive compatibility under finite-sample evaluation.
2. Why standard empirical measures are not truthful
A central negative result is that many familiar measures reward pooling rather than truthful reporting. For 0-ECE, 1-BinECE, smooth calibration error, and 2-distance-to-calibration, the empirical error of the constant predictor 3 is always no larger than that of the original report 4, for every realized label sequence 5 (Hartline et al., 18 Aug 2025). This is an explicit finite-sample non-truthfulness theorem.
The effect appears even in simple examples. If the ground-truth probabilities are i.i.d. uniform on 6 and the truthful predictor reports 7, then empirical ECE is at least 8 because most bins contain singletons, so 9 almost surely. The constant predictor 0 instead achieves 1 empirical ECE, purely from sampling noise (Hartline et al., 18 Aug 2025). The pathology is therefore not merely asymptotic estimation error; it is an incentive problem built into the measure.
This instability is consistent with structural defects of ECE emphasized elsewhere. In the binary two-point example with labels 2, 3, and uniform 4, the constant predictor 5 is perfectly calibrated and has 6, yet the arbitrarily small perturbation 7, 8 yields 9. The same construction shows that 0 while 1, so ECE is discontinuous and fails robust completeness (Błasiok et al., 2022).
In sequential prediction the situation is even sharper. A systematic taxonomy shows that ECE, smooth calibration error, distance from calibration, interval calibration, Laplace-kernel calibration, and U-calibration all exhibit large truthfulness gaps: on simple product distributions, a strategic forecaster can drive the expected penalty to 2 or 3 while truthful forecasting incurs 4 or 5 penalty (Haghtalab et al., 2024).
3. Exact batch constructions
The most explicit perfectly truthful batch construction is Averaged Two-Bin Calibration Error (ATB). For a prediction-label distribution 6 of 7, it is defined by averaging a two-bin squared bias over a uniformly random threshold:
8
Its 9 analogue is
0
The decisive property is the variance-additivity identity
1
which is independent of the report except through the population term 2. Consequently,
3
and more strongly,
4
This rank-preserving identity implies perfect truthfulness and shows that any calibrated report with 5 attains the same minimum expected empirical ATB as the truthful predictor (Hartline et al., 18 Aug 2025).
ATB is also complete, sound, and continuous. Its continuity bounds are
6
for any coupling 7 of 8. It is quadratically related to smooth calibration and lower distance to calibration:
9
while 0 is a constant-factor approximation to both 1 and 2 (Hartline et al., 18 Aug 2025). Estimation requires 3 samples for 4-accuracy with probability 5, and both ATB and ATB1 can be computed in 6 time by sorting predictions and using prefix and postfix scans (Hartline et al., 18 Aug 2025).
ATB arises from a broader recipe called Unnormalized Binned Squared Errors (UBSEs), in which indices are partitioned into bins, unnormalized bin biases are squared, and the sum is globally normalized. Any UBSE is truthful on sequences, and the same recipe yields other truthful batch measures, including quantile-binned 7-ECE constructions (Hartline et al., 18 Aug 2025). A complementary batch result defines binary quantile-binned
8
and proves that its expected value is minimized by the true conditional probabilities for any choice of the number of quantile bins 9 (Lu et al., 7 Oct 2025).
4. Multiclass perfectly truthful measures
The multiclass extension proceeds through classwise one-vs-rest aggregation. If 0 is any truthful binary calibration measure, then
1
with 2 and 3, is truthful for 4-class prediction (Lu et al., 7 Oct 2025).
Instantiating this with truthful binary quantile-binned 5-ECE yields the multiclass measure
6
where each class is sorted separately and binned into classwise quantiles. By contrast, confidence or top-label aggregation does not preserve truthfulness: even when the base binary measure is truthful, the predictor can manipulate which class attains the maximum probability and thereby distort the top-label stream (Lu et al., 7 Oct 2025).
For calibrated predictors, the classwise truthful measure has an especially strong robustness property:
7
Hence, if one calibrated predictor dominates another under all proper losses, then its truthful multiclass calibration error is no larger, for any 8 and 9 (Lu et al., 7 Oct 2025). This dominance preservation is explicitly independent of bin size and addresses the ranking instability observed for standard binned ECE variants.
A distinct multiclass diagnostic line studies asymptotic rather than incentive truthfulness. Variation Calibration Error (VCE) replaces confidence by a general variation functional 0 on the simplex, compares predicted and observed rank-distributions within bins of 1, and satisfies 2 under perfect calibration when 3 is continuous. Under the same perfectly calibrated data-generation scheme, entropy-based UCE need not converge to 4 and can retain a nonzero noise floor (Thompson et al., 13 Feb 2026). This is not a finite-sample truthfulness result, but it identifies a broader class of multiclass calibration diagnostics whose asymptotic target is correctly specified.
5. Sequential and decision-theoretic variants
Perfect truthfulness becomes harder once prediction is sequential. One response is Subsampled Smooth Calibration Error,
5
which randomizes the evaluation subset to block deterministic “repair” strategies. SSCE is 6-truthful for a universal constant 7, and is also complete and sound (Haghtalab et al., 2024). The same paper emphasizes that proper scoring rules are 8-truthful reporting mechanisms but are not complete as calibration measures, because truthful fixed prediction on i.i.d. Bernoulli data still incurs 9 expected penalty (Haghtalab et al., 2024).
A more decision-theoretic sequential measure is subsampled step calibration 0, which is both decision-theoretic and truthful up to an 1 factor on any product distribution, and truthful up to an 2 factor in any 3-smoothed setting (Qiao et al., 4 Mar 2025). The same work proves an impossibility theorem: in non-smoothed settings, any complete and decision-theoretic calibration measure must be discontinuous and non-truthful (Qiao et al., 4 Mar 2025).
The decision-theoretic characterization of perfect calibration itself is classical: a prediction-label distribution is perfectly calibrated if and only if, for every proper loss and every post-processing 4, post-processing cannot reduce expected proper loss. Recent work restricts the post-processing class to structured families 5 and defines 6, which becomes tractable for classes such as monotone maps and generalized monotone unions of intervals; Pool Adjacent Violators is then an omnipredictor for monotone 7, and uniform-mass binning yields omnipredictors for generalized monotone 8 (Gopalan et al., 17 Nov 2025).
A different perspective comes from persuasive calibration. In a principal-agent model with calibration budget 9, setting 00 forces perfect calibration, so any feasible post-processing must be the identity. Allowing 01 permits strategic miscalibration, which under event-independence and 02-ECE concentrates in the tails: under-confidence at low probabilities, over-confidence at high probabilities, and perfect calibration in the middle (Feng et al., 4 Apr 2025). This isolates perfect truthfulness as the zero-budget limit of a broader calibration–persuasion trade-off.
6. Testability, actionability, and adjacent notions of truthful evaluation
Perfectly truthful batch measures solve one problem, but calibration methodology also asks whether a measure is estimable from finite data and whether it controls downstream decision quality. Cutoff Calibration Error (CCE),
03
was introduced precisely to bridge testability and actionability. It satisfies
04
admits a distribution-free 05 estimator, and yields the threshold-decision bound
06
In the same framework, ECE is actionable but not testable in a distribution-free sense, while 07 is testable but not actionable for natural threshold decisions (Rossellini et al., 27 Feb 2025).
Soft-Binned Calibration Decision Loss (SCDL) pushes the actionability line further. Using triangular soft-bin weights 08 and a self-balanced choice of dyadic 09, SCDL is fully actionable for full swap regret and testable with nearly optimal error rate. Its finite-sample estimation error is
10
up to logarithmic factors from confidence, while any fully actionable calibration measure for full swap regret must incur 11 estimation error (Bairaktari et al., 18 May 2026). The same work proves continuity and consistency of SCDL, but its notion of “truthful” is explicitly decision-aware: it truthfully reflects full swap regret rather than finite-sample incentive compatibility of the forecaster.
At a more foundational level, the prediction-only framework based on 12 identifies smooth calibration, interval calibration, and Laplace kernel calibration as consistent measures, and proves a quadratic barrier: no prediction-only calibration measure can approximate the true distance to calibration better than quadratically (Błasiok et al., 2022). The algorithmic side shows that empirical smooth calibration can be reformulated as a minimum-cost flow instance and solved exactly in 13 time, giving nearly-linear-time calibration testing with information-theoretically optimal sample complexity (Hu et al., 2024).
Together, these results place perfectly truthful calibration measures within a larger design space. Exact batch truthfulness, sequential approximate truthfulness, distance consistency, testability, and decision-theoretic actionability are all achievable in some regimes, but the literature repeatedly shows that they do not collapse to a single criterion. Perfectly truthful calibration measures therefore occupy a specific and now well-defined role: they ensure that, under finite-sample evaluation, reporting the true probabilities is itself an optimal strategy.