Papers
Topics
Authors
Recent
Search
2000 character limit reached

Distance to Intersection Multicalibration

Updated 12 July 2026
  • Distance to Intersection Multicalibration (IMC) is a metric that quantifies the minimal L1 adjustment a predictor requires to reach perfect multicalibration on every nonempty intersection of a subgroup family.
  • IMC is constructed by expanding the subgroup collection to its intersection closure, allowing for a weighted decomposition over disjoint atomic partitions which simplifies the calibration landscape.
  • IMC satisfies key properties such as minimal modification and strong auditability through continuity, outperforming alternative metrics like worst-group and ordinary multicalibration distances.

Distance to Intersection Multicalibration (IMC) is a distance-based fairness and calibration notion for probabilistic predictors that measures how far a predictor is from being perfectly multicalibrated on every nonempty intersection of a specified subgroup family. In the formulation analyzed in “Auditability and the Landscape of Distance to Multicalibration,” IMC is defined by enlarging a subgroup collection CC to its intersection closure I(C)I(C), and then taking the L1L_1 distance from a predictor ff to the set of predictors that are perfectly multicalibrated on all sets in I(C)I(C): dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f). The construction is motivated by two requirements for a multicalibration error metric: it should reflect the minimal modification needed to enforce exact multicalibration, and it should be auditable in an information-theoretic sense from finite data. The principal result is that IMC satisfies both requirements, whereas two more immediate generalizations of distance to calibration—worst-group distance to calibration and ordinary distance to multicalibration—each fail one of them (Derhake et al., 21 Sep 2025).

1. Formal setting and underlying calibration notions

The theory is developed on a finite domain XX with X=n|X|=n, binary outcomes Y={0,1}Y=\{0,1\}, and a predictor f:X[0,1]f:X\to[0,1]. Data are drawn i.i.d. from a joint distribution I(C)I(C)0 over I(C)I(C)1, equivalently I(C)I(C)2 and I(C)I(C)3 for a ground-truth conditional probability I(C)I(C)4 supported on all I(C)I(C)5. For I(C)I(C)6, the weighted norm is I(C)I(C)7, with conditional versions I(C)I(C)8 on subsets I(C)I(C)9. If L1L_10, the distance from L1L_11 to L1L_12 is L1L_13 (Derhake et al., 21 Sep 2025).

Perfect calibration requires that for every value L1L_14 in the image of L1L_15,

L1L_16

If L1L_17 is a possibly overlapping collection of subgroups, then L1L_18 is perfectly L1L_19-multicalibrated if for every ff0 and every score value ff1 in the image of ff2,

ff3

The paper works with these continuized, score-conditioned constraints rather than bucketed approximations. This exact formulation is central to the later analysis of geometry, discontinuity, and auditability (Derhake et al., 21 Sep 2025).

A key predecessor is distance to calibration error,

ff4

which measures the minimal ff5 perturbation required to transform ff6 into a perfectly calibrated predictor. IMC extends the same distance-to-set philosophy from marginal calibration to multicalibration over intersections (Derhake et al., 21 Sep 2025).

2. The two desiderata and the failure of more immediate generalizations

The motivation for IMC begins with two desiderata. First, a multicalibration metric should encode minimal modification: it should quantify how much ff7 must change in ff8 to satisfy perfect multicalibration. Second, it should be auditable: small changes in the underlying ground truth ff9 should induce small changes in the metric, making finite-sample estimation information-theoretically plausible (Derhake et al., 21 Sep 2025).

Two natural candidates are examined. The first is worst-group distance to calibration,

I(C)I(C)0

which aggregates subgroup-specific calibration distances by a weighted maximum. The second is ordinary distance to multicalibration,

I(C)I(C)1

which directly measures the minimal change needed to achieve perfect multicalibration on all groups in I(C)I(C)2 simultaneously (Derhake et al., 21 Sep 2025).

These two notions fail in complementary ways. I(C)I(C)3 fails the minimal-modification desideratum: there exist I(C)I(C)4, a two-group collection I(C)I(C)5, and a predictor I(C)I(C)6 such that I(C)I(C)7 is a strict local minimum of I(C)I(C)8, with I(C)I(C)9, but any predictor dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f)0 satisfying dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f)1 must be dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f)2 away from dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f)3 in dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f)4. The obstruction arises because a maximum of subgroup distances can be locally unimprovable under small perturbations even when the jointly feasible region lies far away (Derhake et al., 21 Sep 2025).

By contrast, dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f)5 has the right geometric interpretation but fails auditability. The paper constructs a family of distributions dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f)6, two overlapping subgroups dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f)7 and dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f)8, and the constant predictor dIMCC(f):=dMCI(C)(f)dIMC_C(f):=dMC_{I(C)}(f)9 such that XX0 at XX1, but XX2 for every XX3, while the total variation distance between XX4 and XX5 is only XX6. Thus an arbitrarily small perturbation of XX7 can change the multicalibration distance by a constant, producing a pointwise discontinuity and an information-theoretic inauditability result (Derhake et al., 21 Sep 2025).

3. Definition through intersection closure and atomic partitions

IMC resolves this tension by replacing the original group family with its intersection closure. If XX8 covers XX9, define

X=n|X|=n0

Distance to Intersection Multicalibration is then

X=n|X|=n1

In words, X=n|X|=n2 is measured against the set of predictors that are perfectly calibrated not only on each declared group, but on every nonempty intersection generated by those groups (Derhake et al., 21 Sep 2025).

A central structural device is the disjoint partition X=n|X|=n3, the family of atoms of the Venn diagram induced by X=n|X|=n4: X=n|X|=n5 This is a disjoint cover of X=n|X|=n6. The key equivalence is

X=n|X|=n7

hence

X=n|X|=n8

The significance of this reformulation is that it converts overlapping intersectional constraints into calibration constraints on disjoint atoms, making the metric decomposable and auditable (Derhake et al., 21 Sep 2025).

When a collection X=n|X|=n9 is itself a disjoint cover, the multicalibration distance decomposes exactly: Y={0,1}Y=\{0,1\}0 Applying this to Y={0,1}Y=\{0,1\}1 yields the canonical IMC decomposition

Y={0,1}Y=\{0,1\}2

This representation shows that IMC is a weighted aggregation of marginal distance-to-calibration terms over the finest intersectional partition induced by the original subgroup family (Derhake et al., 21 Sep 2025).

4. Auditability and equivalence to continuized distance to multicalibration

Auditability is formalized through continuity, specifically Lipschitz continuity, in the ground-truth conditional probability Y={0,1}Y=\{0,1\}3. The paper proves that Y={0,1}Y=\{0,1\}4 is Y={0,1}Y=\{0,1\}5-Lipschitz in Y={0,1}Y=\{0,1\}6. The reason is precisely the decomposition over Y={0,1}Y=\{0,1\}7: each per-atom term Y={0,1}Y=\{0,1\}8 is Y={0,1}Y=\{0,1\}9-Lipschitz in f:X[0,1]f:X\to[0,1]0, and the atom weights sum to f:X[0,1]f:X\to[0,1]1. Consequently, IMC varies continuously and in a controlled way under perturbations of the data-generating distribution (Derhake et al., 21 Sep 2025).

The same section introduces a continuized variant of ordinary distance to multicalibration: f:X[0,1]f:X\to[0,1]2 where f:X[0,1]f:X\to[0,1]3. This construction takes the upper envelope of f:X[0,1]f:X\to[0,1]4 over arbitrarily small neighborhoods of f:X[0,1]f:X\to[0,1]5, explicitly removing downward discontinuities (Derhake et al., 21 Sep 2025).

The main equivalence theorem states

f:X[0,1]f:X\to[0,1]6

The proof relies on the statement that f:X[0,1]f:X\to[0,1]7 and f:X[0,1]f:X\to[0,1]8 differ only on a Lebesgue-measure-zero set of f:X[0,1]f:X\to[0,1]9. This identifies IMC as the continuized form of I(C)I(C)00: it preserves the minimal-modification semantics of a distance-to-set metric while replacing a discontinuous quantity with one that is stable enough to audit (Derhake et al., 21 Sep 2025).

A common confusion is to treat ordinary I(C)I(C)01 as automatically auditable because it is a distance to a constraint set. The paper shows that this is false: distance-to-set geometry alone does not prevent sharp dependence on the underlying ground truth when overlapping subgroup constraints create singular configurations (Derhake et al., 21 Sep 2025).

5. Geometry of the feasible set and the loss landscape

The geometric analysis of multicalibration is one of the conceptual contributions of the framework. For each subgroup I(C)I(C)02, the set I(C)I(C)03 is finite. The set I(C)I(C)04 is obtained by intersecting, over I(C)I(C)05, unions of affine subspaces determined by calibrated score profiles on I(C)I(C)06. Expanding these intersections shows that I(C)I(C)07 is a finite union of truncated affine subspaces of I(C)I(C)08 (Derhake et al., 21 Sep 2025).

As a consequence, I(C)I(C)09 is the pointwise minimum of finitely many convex functions, each arising from distance to an affine set, and in I(C)I(C)10 it is piecewise linear. This directly implies the theorem that all local minima of I(C)I(C)11 are global minima. Because I(C)I(C)12 is simply I(C)I(C)13 evaluated on the enlarged constraint family I(C)I(C)14, it inherits the same local-minima-are-global property (Derhake et al., 21 Sep 2025).

The contrast with I(C)I(C)15 is sharp. Since I(C)I(C)16 is a maximum of subgroup-specific distance terms, its loss landscape can exhibit large basins and nonconvex artifacts. The local minima produced by this max-of-distances structure are precisely what makes I(C)I(C)17 unsuited to represent minimal modification. By passing to I(C)I(C)18, IMC replaces this max structure with a weighted sum over disjoint atoms, revealing a more regular aggregation structure at the partition level (Derhake et al., 21 Sep 2025).

This geometry has algorithmic implications. The paper explicitly notes that the findings may have implications for the development of stronger multicalibration algorithms, and suggests the possibility of optimization procedures that operate directly in predictor space while exploiting the fact that local minima coincide with global minima for distance-to-set objectives (Derhake et al., 21 Sep 2025).

The positive auditability result for IMC is not a blanket claim of easy estimation. In the worst case, auditing I(C)I(C)19 can require I(C)I(C)20 samples, because the induced atomic partition I(C)I(C)21 may itself be exponentially large. This is established by an exponential lower bound showing that distinguishing I(C)I(C)22 from I(C)I(C)23 may require exponentially many samples in general (Derhake et al., 21 Sep 2025).

There is, however, a corresponding positive regime. If I(C)I(C)24 has size I(C)I(C)25 and each atom I(C)I(C)26 has mass at least I(C)I(C)27, then with

I(C)I(C)28

i.i.d. samples of I(C)I(C)29, one can construct an estimator I(C)I(C)30 such that, with failure probability at most I(C)I(C)31,

I(C)I(C)32

The associated auditing pipeline is modular: estimate I(C)I(C)33 for each atom I(C)I(C)34 using the dCE lower-distance oracle of Blasiok et al.; estimate atom masses; and aggregate these estimates according to the weighted-sum decomposition (Derhake et al., 21 Sep 2025).

The theory is stated under explicit assumptions: finite I(C)I(C)35, binary labels, subgroup collections covering I(C)I(C)36, weighted I(C)I(C)37 distance with respect to I(C)I(C)38, and access to group-membership indicators for per-atom auditing. Limitations include the possible exponential size of I(C)I(C)39, the restriction to binary calibration, and the fact that extensions to infinite or continuous domains require additional work (Derhake et al., 21 Sep 2025).

Related research places IMC in a broader multicalibration landscape. In LLM confidence scoring, multicalibration has been applied using groups built by clustering in an embedding space and by self-annotation; intersections can be enforced by adding intersection indicators to the group family, though that work does not define IMC itself (Detommaso et al., 2024). In clinical risk prediction, proportional multicalibration constrains percent calibration error across intersectional groups and bins, and is linked to differential calibration; this is a distinct criterion from IMC, since it is based on normalized binwise error rather than distance to the exact multicalibrated set (Cava et al., 2022). Taken together, these lines of work suggest that intersectional conditioning is increasingly treated as central rather than auxiliary, but that the metric used to quantify “distance to fairness” materially affects both geometry and auditability.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (3)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Distance to Intersection Multicalibration.