---
title: Distance to Multicalibration (dMC)
url: https://www.emergentmind.com/topics/distance-to-multicalibration-dmc
type: topic
---

# Distance to Multicalibration (dMC)

Searching arXiv for recent papers on multicalibration distance, auditability, and operational metrics.
Distance to Multicalibration (dMC) denotes a family of quantities that measure how far a predictor is from satisfying multicalibration constraints over a collection of overlapping subgroups. In its most direct formalization, for a finite feature space \(X\), binary labels \(Y=\{0,1\}\), a predictor \(f:X\to[0,1]\), and a subgroup family \(C=\{S_1,\dots,S_k\}\), dMC is the \(L_1(D_x)\)-distance from \(f\) to the set of perfectly \(C\)-multicalibrated predictors:
\[
dMC_{C,D}(f)=\inf_{g\in mcal_C(D)}\|f-g\|_1.
\]
This definition is the multigroup analogue of distance to calibration error \(dCE_D(f)=\|f,cal(D)\|_1\), but the literature also uses several operational surrogates for “distance to multicalibration,” including worst-group calibration error, subgroup-bin residuals, and residual-correlation norms over an auditing class [2509.16930][2406.06487].

## 1. Formal object and basic setup

Multicalibration strengthens ordinary calibration by requiring simultaneous calibration on each subgroup in a family \(C\), possibly with overlap. For a predictor \(f\), perfect calibration requires
\[
\mathbb{E}_{(x,y)\sim D}\!\left[y\mid f(x)=v\right]=v
\quad\text{for every } v\in \mathrm{Im}(f),
\]
and perfect \(C\)-multicalibration requires the same condition inside every \(S\in C\), namely
\[
\mathbb{E}_{(x,y)\sim \left.D\right|_S}[y\mid f(x)=v]=v
\quad\text{for every } S\in C,\ v\in \mathrm{Im}(f|_S).
\]
The corresponding feasible sets are denoted \(cal(D)\) and \(mcal_C(D)\) [2509.16930].

The geometric dMC definition inherits the norm structure of \(dCE\). If \(D_x\) is the marginal over \(X\), then
\[
\|f\|_p=\left(\mathbb{E}_{x\sim D_x}[|f(x)|^p]\right)^{1/p},
\qquad
\|f,H\|_p=\inf_{h\in H}\|f-h\|_p.
\]
The \(p=1\) case is the primary focus in the dMC literature. Under this formulation, dMC measures the minimum amount one must modify \(f\), in \(L_1(D_x)\), to obtain a single predictor that is calibrated on all groups simultaneously, not merely one that is separately close to calibrated on each group [2509.16930].

This simultaneous-feasibility requirement is what distinguishes dMC from looser worst-group notions. It also makes the metric sensitive to overlap structure in \(C\): two groupwise repairs that are individually small need not be compatible with each other. In practical studies, the subgroup family is usually finite and explicitly specified, for example by metadata, feature conjunctions, or lexical indicators; the resulting notion is therefore always relative to the chosen family rather than to an infinite concept class [2406.06487].

## 2. Metric variants, geometry, and auditability

A natural alternative to dMC is worst-group distance to calibration,
\[
wdMC_C(f):=\max_{S\in C}\left([S]\cdot dCE_{\left.D\right|_S}(\left.f\right|_S)\right),
\]
where \([S]=\Pr_{x\sim D_x}[x\in S]\). This quantity audits each group separately and then takes a maximum. It is therefore easier to interpret as a worst-case subgroup repair cost, but it is not the same as the distance to a single simultaneously multicalibrated predictor. The hierarchy
\[
wdMC_{D,C}(f)\le dMC_{D,C}(f)
\]
holds in general [2509.16930].

The 2025 analysis of dMC identifies two desiderata for any multicalibration error metric: it should capture how much \(f\) must be modified to become perfectly multicalibrated, and it should be auditable in an information-theoretic sense. Raw dMC satisfies the first desideratum exactly, because it is literally a distance to the feasible set \(mcal_C(D)\). It also has a favorable optimization landscape: all local minima of \(dMC_C\) are global minima. Geometrically, \(mcal_C(D)\) is a finite union of truncated affine subspaces, so \(dMC\) is generally nonconvex but piecewise linear in the \(L_1\) setting [2509.16930].

The central negative result is that dMC is not auditable information-theoretically in general. There exist \(X\), \(C\), \(f\), and a family of distributions \(D_\alpha\) such that
\[
dMC_{C,D_\alpha}(f)>0.3 \quad \text{if } \alpha>0,
\qquad
dMC_{C,D_0}(f)=0,
\]
while
\[
\|D_0-D_\alpha\|_{TV}\in O(\alpha).
\]
Thus arbitrarily small perturbations of the ground truth can induce constant jumps in dMC, so dMC is not Lipschitz in \(p^*\) and fails the second desideratum [2509.16930].

To repair this instability, the same paper introduces the continuized distance to multicalibration,
\[
\widetilde{dMC}_{C,p^*}(f):=\lim_{\epsilon\to 0^+}\sup_{p\in B_\epsilon(p^*)} dMC_{C,p}(f),
\]
and distance to intersection multicalibration,
\[
dIMC_{C,p^*}(f):=dMC_{I(C),p^*}(f),
\]
where
\[
I(C)=\left\{\bigcap_{S\in A}S: A\subseteq C,\ A\neq\emptyset\right\}.
\]
These two replacements are equivalent:
\[
\widetilde{dMC}_{C,p^*}(f)=dIMC_{C,p^*}(f)=dMC_{J(C),p^*}(f),
\]
where \(J(C)\) is the partition generated by \(C\). They retain the “distance-to-feasibility” interpretation, are \(1\)-Lipschitz in \(p^*\), and therefore satisfy both desiderata, at the cost of strengthening the target from multicalibration on \(C\) to multicalibration on all nonempty intersections of groups in \(C\) [2509.16930].

## 3. Operational surrogates in empirical practice

Before explicit geometric dMC was formalized, empirical work typically operationalized “how far a model is from multicalibration” by worst-group calibration error over a finite subgroup family \(\mathcal G\). In a broad 2024 empirical study spanning tabular, image, and language tasks, the practical metric is
\[
\max_{g\in\mathcal G}\mathrm{CalErr}(f;g),
\]
instantiated primarily as
\[
\max_{g\in\mathcal G}\mathrm{smECE}(f;g),
\]
with
\[
\max_{g\in\mathcal G}\mathrm{ECE}(f;g)
\]
used as an alternative. The paper explicitly states that taking the max avoids fairness concerns associated with the weighted mean of groups of varying size and/or degree of overlap, and it repeatedly uses validation max smECE to select the best post-processed model [2406.06487].

In that study, ordinary calibration is measured by binned ECE with 10 equal-width bins of size \(0.1\), while multicalibration post-processing is formulated on subgroup-bin “categories” \(g\cap f^{-1}(b)\) with bin width
\[
\lambda=0.1.
\]
For the Hébert-Johnson et al. algorithm (HKRR), the operative residual is the subgroup-bin discrepancy
\[
\left|\widehat{\mathbb E}[Y\mid X\in g,\ f(X)\in b]-\widehat{\mathbb E}[f(X)\mid X\in g,\ f(X)\in b]\right|,
\]
and categories are skipped when
\[
|g\cap f^{-1}(b)|\le \alpha\lambda |g|.
\]
This makes the empirical notion of dMC discretization-dependent: closeness to multicalibration is assessed after binning scores at width \(0.1\) [2406.06487].

The main empirical finding is that many already calibrated models are close to multicalibration out of the box under this worst-group surrogate. On MEPS, for example, MLP ERM has
\[
\text{max smECE}=0.086\pm 0.015,
\]
HJZ attains
\[
0.076\pm 0.018,
\]
and HKRR is worse at
\[
0.104\pm 0.002.
\]
RandomForest ERM on the same dataset has
\[
0.083\pm 0.004,
\]
and LogisticRegression ERM has
\[
0.083\pm 0.003.
\]
By contrast, uncalibrated models can be much farther. On MEPS, SVM ERM has
\[
\text{max smECE}=0.186\pm 0.006,
\]
which HKRR reduces to
\[
0.104\pm 0.002,
\]
while NaiveBayes ERM has
\[
0.287\pm 0.011,
\]
and HKRR reduces it to
\[
0.104\pm 0.002.
\]
On larger models, ERM max smECE falls from \(0.094\pm 0.009\) to \(0.047\pm 0.005\) for Amazon Polarity with ResNet-56 under HKRR, and from \(0.076\pm 0.011\) to \(0.018\pm 0.003\) for Camelyon17 with ViT [2406.06487].

This operational line of work also emphasizes that low overall calibration error often tracks low worst-group calibration error better than one might expect, though not perfectly. Traditional calibration methods may sometimes provide multicalibration implicitly, and isotonic regression often reduces max-group smECE nearly as much as explicit multicalibration algorithms. At the same time, the exact metric matters: on MEPS, DecisionTree ERM has worst-group calibration error \(0.261\) by max ECE but only \(0.166\) by max smECE, so the inferred practical dMC can vary substantially with the estimator [2406.06487].

## 4. Residual-based formulations and theoretical generalizations

A broader theoretical literature treats dMC not as geometric distance to a feasible set, but as a residual aggregated over groups and prediction values. For a scalar property \(\Gamma\), one exact calibration condition is
\[
\Gamma(Y_{f,\gamma,G})=\gamma
\quad\text{for every } \gamma\in \mathrm{Range}_f,\ G\in G,
\]
and the corresponding approximate multicalibration definition is the weighted \(\ell_2\) bound
\[
\sum_{\gamma\in \mathrm{Range}_f}\Pr[f(x)=\gamma\mid x\in G]
\left(\gamma-\Gamma(Y_{f,\gamma,G})\right)^2
\le \frac{\alpha}{\mu(G)}.
\]
In the identifiable case, the paper gives an equivalent residual in terms of an identification function \(V\):
\[
\sum_{\gamma\in \mathrm{Range}_f}\Pr[f(x)=\gamma\mid x\in G]\cdot
\bigl(V(\gamma,Y_{(G,\gamma)})\bigr)^2
\le \frac{\alpha}{\Pr[x\in G]}.
\]
These are explicit dMC candidates: weighted squared calibration violations over group-value cells [2302.08507].

That framework also gives a sharp structural characterization of when zero dMC is even possible. For continuous scalar properties, the paper proves
\[
\Gamma \text{ sensible for calibration }
\iff
\Gamma \text{ elicitable }
\iff
\Gamma \text{ identifiable }
\iff
\Gamma \text{ has convex level sets}.
\]
For non-elicitable properties, there are simple datasets on which even the true predictor is not calibrated. The variance example is canonical: with two points \(x_0,x_1\), deterministic labels \(Y(x_0)=0\), \(Y(x_1)=1\), and equal mass, the true pointwise variance predictor is identically \(0\), yet on the prediction cell \(f=0\) the label variance is \(0.25\). This yields an irreducibly nonzero calibration residual even for the true predictor. A plausible implication is that any dMC notion based on vanishing calibration equations has a positive floor for variance, CVaR alone, and most distortion risk measures [2302.08507].

A second residual-based generalization replaces Boolean groups by real-valued grouping functions. In the joint-grouping formulation, approximate multicalibration is defined by
\[
K_2(f,h,P)
=
\int \Big(\mathbb E_P[h(X,Y)(Y-v)\mid f(X)=v]\Big)^2\, dP_{f(X)}(v)
\le \alpha
\quad\text{for all } h\in\mathcal H.
\]
The natural scalar dMC induced by this definition is
\[
\sup_{h\in\mathcal H}K_2(f,h,P),
\]
or its empirical analogue. This formulation includes calibration when \(h\equiv 1\), subgroup calibration for Boolean \(h(x)\), reweighted calibration for real-valued \(h(x)\), and joint covariate-label constraints for \(h(x,y)\). It also connects multicalibration to robustness under covariate shift and beyond covariate shift via density-ratio grouping classes [2406.00661].

In regression with squared-error loss, multicalibration can be expressed as residual orthogonality to a weak learner class \(\mathcal B=\{b_j\}_{j=1}^p\). The population error vector is
\[
\mathcal E(f)=\big(\epsilon_1,\dots,\epsilon_p\big),
\qquad
\epsilon_j=\mathbb E_{(x,y)\sim D}[(y-f(x))b_j(x,f(x))],
\]
and the empirical analogue is
\[
\widehat{\mathcal E}(f)=\frac{1}{n}B(f)^\top(y-f).
\]
In this setting, \(\|\widehat{\mathcal E}(f)\|_2\) is the paper’s most direct measurable multicalibration-violation quantity, and therefore a natural empirical dMC proxy [2602.06773].

## 5. Algorithms, convergence, and computational aspects

Several post-processing and boosting algorithms can be interpreted as dMC-reduction procedures. HKRR is described as an iterative boosting-style post-processing algorithm that works by iteratively searching for and removing subgroup calibration violations until convergence. HJZ instead frames multicalibration as a two-player game with better theoretical sample complexity guarantees and is run with \(\lambda=0.1\), 30 iterations, learner decay rates \(\eta\in\{0.9,0.95\}\), and adversary decay rates \(\eta\in\{0.9,0.95,0.98\}\). In both cases the effective target is to reduce group-bin residuals and thereby lower worst-group calibration error [2406.06487].

For joint grouping functions, MC-Pseudolabel provides a regression-oracle-based post-processing algorithm. It discretizes the current predictor, regresses residuals on each level set using the grouping class, forms pseudolabels, and refits the predictor. Its key stopping certificate is:
\[
err_{t-1}-\tilde{err}_t \le \frac{\alpha}{B}
\quad\Longrightarrow\quad
\text{the output is certified }\alpha\text{-approximately }\ell_2\text{ multicalibrated w.r.t. }\mathcal H_B,
\]
where
\[
\mathcal H_B=\{h\in\mathcal H:\sup h(x,y)^2\le B\}.
\]
This converts small one-step improvability in squared error into a bound on worst-case multicalibration violation over bounded groups [2406.00661].

In multicalibration gradient boosting for squared-loss regression, the iterate update is
\[
f_{t+1}=f_t+\eta A(f_t)(y-f_t)
\]
in the unit-weight case, with Lyapunov objective
\[
V(f)=\frac12\|y-f\|_2^2.
\]
The central convergence identity is
\[
V(f_{t+1})-V(f_t)
=
-\frac{1}{\eta}\left(1-\frac12\eta\right)\|f_{t+1}-f_t\|_2^2.
\]
Combined with
\[
\|\widehat{\mathcal E}(f_t)\|_2
\le
\frac{1}{n\eta}\|B(f_t)\|_2\|f_{t+1}-f_t\|_2,
\]
this implies asymptotic vanishing of empirical multicalibration error and a global \(O(1/\sqrt{T})\) best-iterate bound on the update gap, with linear convergence under additional smoothness assumptions on \(A(f)\) [2602.06773].

Exact computation of geometric dMC remains largely open, but ordinary calibration distance already exhibits substantial complexity barriers. For the base problem
\[
CalDist_D(f)=\inf_{g\in C(D)}\mathbb E_{x\sim D_x}[|f(x)-g(x)|],
\]
exact computation is possible in \(O(|X|^4)\) time when the marginal is uniform and labels are noiseless, yet becomes \(\mathsf{NP}\)-hard when either assumption is removed. The same work gives a PTAS in general and shows that \(\Theta(1/\varepsilon^3)\) samples are sufficient and necessary for one-sided estimation of empirical calibration distance. This does not prove hardness for dMC itself, but the paper explicitly argues that the result strongly suggests analogous barriers for dMC over rich subgroup classes [2603.18391].

## 6. Fairness refinements, extensions, and scope limitations

A recurrent critique of additive dMC surrogates is that they can mask large relative disparities in low-base-rate groups. Proportional multicalibration replaces absolute calibration error by relative or percent calibration error on group-bin cells. The operative additive proxy for ordinary multicalibration in that paper is
\[
\max_{S\in C,\ I\in\Lambda_\lambda}
\left|
\mathbb E_D[y\mid R\in I,\ x\in S]
-
\mathbb E_D[R\mid R\in I,\ x\in S]
\right|,
\]
while the proportional analogue is
\[
\max_{S\in C,\ I\in\Lambda_\lambda}
\frac{
\left|
\mathbb E_D[y\mid R\in I,\ x\in S]
-
\mathbb E_D[R\mid R\in I,\ x\in S]
\right|
}{
\mathbb E_D[y\mid R\in I,\ x\in S]
}.
\]
The theory shows
\[
\alpha\text{-PMC}\implies \frac{\alpha}{1-\alpha}\text{-MC}
\]
and
\[
\alpha\text{-PMC}\implies
\left(\ln\frac{1+\alpha}{1-\alpha}\right)\text{-DC}.
\]
This indicates that a small additive dMC need not control multiplicative disparity unless relative calibration is also addressed [2209.14613].

Higher-moment extensions make the same point in another direction. Moment multicalibration does not calibrate moments on arbitrary subgroups directly; instead it uses mean-conditioned moment multicalibration on bucketed cells
\[
S(\bar\mu,\bar m,i,j)
=
\left\{
x\in S:
\left|\bar\mu(x)-\frac{2i-1}{2m}\right|\le \frac{1}{2m},
\ \left|\bar m(x)-\frac{2j-1}{2m}\right|\le \frac{1}{2m}
\right\}.
\]
The corresponding constraints are
\[
\left| \mu(S(\bar\mu,\bar m,i,j))-\bar\mu(S(\bar\mu,\bar m,i,j)) \right|
\le \frac{\alpha}{P_X(S(\bar\mu,\bar m,i,j))}+\epsilon
\]
and
\[
\left| m_k(S(\bar\mu,\bar m,i,j))-\bar m(S(\bar\mu,\bar m,i,j)) \right|
\le \frac{\beta}{P_X(S(\bar\mu,\bar m,i,j))}+\epsilon.
\]
A plausible implication is that dMC for uncertainty estimation is naturally vector-valued: it must track mean violations and moment violations jointly rather than collapse everything into a single mean-calibration number [2008.08037].

Across the literature, several limitations recur. dMC is always relative to a chosen subgroup family or grouping-function class; a predictor can be close for one family and far for another. Many operational definitions depend on discretization, such as \(\lambda=0.1\) score bins or finite bucket counts \(m\). Empirical estimates are sensitive to subgroup sample size, and intersection-based replacements such as \(dIMC\) can incur exponential blow-up in the size of \(J(C)\). Practical post-processing also competes with predictive accuracy because held-out calibration data must be reserved for auditing and repair. These constraints explain why the literature alternates between exact geometric definitions, auditable residual surrogates, and application-specific approximations rather than converging on a single universal scalar metric [2509.16930][2406.06487].

Source: https://www.emergentmind.com/topics/distance-to-multicalibration-dmc