Papers
Topics
Authors
Recent
Search
2000 character limit reached

Identification of Latent Group Effects under Conditional Calibration

Published 9 Apr 2026 in stat.ME, econ.EM, and stat.CO | (2604.08798v1)

Abstract: We study identification of a structural group effect when the group indicator G∈0,1G\in{0,1} is unobserved but the analyst observes a calibrated probability score p∈[0,1]p\in[0,1] satisfying E[G∣p,X]=p\mathbb{E}[G|p,X]=p. Under a constant-coefficient structural mean model, the latent-group coefficient ττ is point-identified from the joint law of observables (Y,X,p)(Y,X,p) by a simple ratio of weighted moments: the covariance of the signed score $2p-1$ with the covariate-partialled outcome, divided by twice the residual variance of the score after conditioning on covariates. Identification fails if and only if the score is a deterministic function of XX; we establish this by constructing an explicit continuum of observationally equivalent models indexed by arbitrary values of ττ. The identified coefficient differs from the marginal latent mean gap by a compositional term that is unidentified without further assumptions; we give a necessary and sufficient condition for the two to coincide. The oracle estimator is n\sqrt{n}-consistent and asymptotically normal with a closed-form sandwich variance. Under calibration error bounded uniformly by δδ, the bias is bounded by ∣τ∣ E[∣2p−1∣] δ (2V<sup>∗)<sup>−1|τ|\,\mathbb{E}[|2p-1|]\,δ\,(2V<sup>*)<sup>{-1}, a bound that is sharp over all calibration error functions of that magnitude. Hard-threshold classification at p=1/2p=1/2 attenuates the estimated gap by a factor strictly less than one. Monte Carlo experiments confirm the asymptotic theory, trace the divergence of RMSE as V<sup>∗→</sup>0V<sup>*\to</sup> 0, illustrate the attenuation bias of hard-threshold classification, and verify identification of the variance-weighted estimand under heterogeneous effects.

Authors (1)

Summary

  • The paper identifies a constant latent-group effect with a closed-form moment ratio when the calibrated score retains variation beyond covariates, requiring residual score variance greater than zero.
  • The paper shows identification fails exactly when the score is fully determined by covariates, while the identified structural effect generally differs from the marginal group mean gap because of covariate composition.
  • The paper derives sharp miscalibration-bias bounds and simulation evidence showing that moment estimators outperform hard-threshold classification, while cross-fitting improves finite-sample performance with flexible nuisance estimates.

Setting and motivation

Empirical work frequently requires estimating outcome differences between groups whose membership is never directly observed—poverty status, informal employment, latent health conditions. The analyst instead observes a probability score pi∈[0,1]p_i \in [0,1] encoding belief that unit ii belongs to the group of interest. The paper "Identification of Latent Group Effects under Conditional Calibration" (2604.08798) asks a precise question: under what conditions, and by what formula, can a structural group effect be identified from the joint law of observables (Y,X,p)(Y,X,p) when the binary indicator GG is never observed? The answer is organized around three claims: a structural coefficient τ\tau is point-identified under mild conditions; identification fails in an exactly characterizable way when one condition is violated; and the identified object is distinct from the marginal group mean gap.

The setting complements three literatures. In the misclassification literature (Lewbel 2007; Mahajan 2006; Kasahara and Shimotsu 2022), the analyst observes a noisy binary label; here the analyst observes a calibrated probability, which changes both the identification argument and the estimand. Relative to proxy-variable and measurement-error approaches requiring rank conditions on integral operators (Hu and Schennach 2008; Schennach 2016), the target is only the scalar Ï„\tau, and the assumptions are correspondingly weaker with a closed-form formula. Within algorithmic fairness research on disparity estimation with unobserved protected attributes (Kallus et al. 2022; Chen et al. 2019), the contribution is a formal identification-theoretic treatment with exact failure and sensitivity characterizations.

Model and assumptions

The data are i.i.d. draws (Yi,Xi,pi)(Y_i,X_i,p_i) from PY,X,pP_{Y,X,p}, with an unobserved Gi∈{0,1}G_i \in \{0,1\}. Three derived quantities are central: the signed score z=2p−1z = 2p-1, the outcome residual ii0 where ii1, and the score residual ii2 where ii3. The key scalar governing everything is

ii4

the residual variance of the score after partialling on covariates.

Two structural assumptions carry the analysis. Assumption 1 posits a constant-coefficient conditional mean: ii5. This bundles two restrictions—the group effect is constant in ii6, and the score carries no additional information about ii7 once ii8 are known. Assumption 2 is conditional calibration: ii9 almost surely. Notably, (Y,X,p)(Y,X,p)0 need not equal the propensity score (Y,X,p)(Y,X,p)1; it must merely be an unbiased predictor of (Y,X,p)(Y,X,p)2 given all observed information. Area-level prevalence rates, classifier outputs, and model-based predictions satisfy this when well-calibrated. Identification requires only second moments ((Y,X,p)(Y,X,p)3); asymptotic normality uses fourth moments.

Two lemmas form the backbone. First, calibration plus the tower property forces (Y,X,p)(Y,X,p)4, so the observable regression satisfies (Y,X,p)(Y,X,p)5. Second, the outcome residual decomposes as (Y,X,p)(Y,X,p)6 with (Y,X,p)(Y,X,p)7: the entire predictable part of (Y,X,p)(Y,X,p)8 is driven by the deviation of true membership from its score-implied expectation.

Point identification and its exact boundary

The central result is the population moment identity

(Y,X,p)(Y,X,p)9

which yields point identification by the closed-form ratio

GG0

The formula has a transparent instrumental-variables interpretation: the score residual GG1 acts as an instrument for the latent deviation GG2. Calibration supplies first-stage relevance—a covariance identity shows the population regression slope of GG3 on GG4 equals exactly one—and mean-independence in Assumption 1 supplies the exclusion restriction. The estimand is thus the IV ratio of reduced-form to first-stage slopes.

Assumption 3 (GG5) is not a mere regularity condition but the exact boundary of identification. If GG6—equivalently, if and only if GG7 almost surely, i.e., the score is a deterministic function of GG8—then GG9 is not identified. The failure is made constructive: for any τ\tau0, one can adjoin τ\tau1 for independent uniform τ\tau2, set τ\tau3, and obtain a model satisfying both assumptions that reproduces the same observable regression τ\tau4 and hence the same moment equation. There is therefore an explicit continuum of observationally equivalent models indexed by arbitrary values of τ\tau5. A practical implication is sharp: whenever a practitioner's score is fully explained by covariates, no estimator—however clever—can recover the group effect without further structure.

Structural coefficient versus marginal gap

A natural concern is whether Ï„\tau6 equals the marginal latent mean gap Ï„\tau7. It does not, in general. Under the constant-coefficient model,

Ï„\tau8

where τ\tau9 captures differences in covariate composition across latent groups. Because many latent joint distributions of τ\tau0 generate the same observable law of τ\tau1, τ\tau2 is unidentified without further assumptions. The two objects coincide if and only if the latent groups are covariate-balanced in the sense that τ\tau3. The practical import is that τ\tau4 identifies the within-covariate-cell group effect, while the marginal gap conflates this effect with composition—an important caveat for applications aiming at headline disparity numbers.

Oracle inference and plug-in estimation

With known nuisance functions Ï„\tau5 and Ï„\tau6, the oracle estimator is the sample analogue of the identification ratio. It is Ï„\tau7-consistent and asymptotically normal with closed-form sandwich variance

Ï„\tau8

obtained by applying the delta method to the ratio of two sample means; all cross-terms cancel in terms of the centred score. A consistent variance estimator yields valid Wald intervals.

When Ï„\tau9 and (Yi,Xi,pi)(Y_i,X_i,p_i)0 are unknown, replacing them by estimators converging in (Yi,Xi,pi)(Y_i,X_i,p_i)1 gives a consistent plug-in estimator; denominator stability under nuisance estimation follows from a boundedness argument since (Yi,Xi,pi)(Y_i,X_i,p_i)2 and (Yi,Xi,pi)(Y_i,X_i,p_i)3 lie in (Yi,Xi,pi)(Y_i,X_i,p_i)4.

The paper is candid about the limits of this result. Gateaux derivative calculations show the score fails Neyman orthogonality through its (Yi,Xi,pi)(Y_i,X_i,p_i)5-direction (the (Yi,Xi,pi)(Y_i,X_i,p_i)6-direction vanishes automatically because (Yi,Xi,pi)(Y_i,X_i,p_i)7). Replacing (Yi,Xi,pi)(Y_i,X_i,p_i)8 by (Yi,Xi,pi)(Y_i,X_i,p_i)9 in the numerator produces a Neyman-orthogonal score in both directions, defining a distinct orthogonal estimator that is asymptotically equivalent to the plug-in when nuisances are known but generally not otherwise. Establishing PY,X,pP_{Y,X,p}0-normality of this orthogonal estimator under cross-fitting—i.e., placing it formally within the double machine learning framework—is explicitly left open, as is whether the oracle estimator attains the semiparametric efficiency bound.

Robustness to calibration failure

If calibration is violated so that PY,X,pP_{Y,X,p}1, the probability limit of the oracle estimator is exactly PY,X,pP_{Y,X,p}2 with

PY,X,pP_{Y,X,p}3

Three features deserve emphasis. The bias is proportional to PY,X,pP_{Y,X,p}4: no bias arises when the true effect is zero, regardless of miscalibration. It vanishes whenever the miscalibration is orthogonal to the signed score—for instance, symmetric errors such as PY,X,pP_{Y,X,p}5. And over the class of all calibration error functions bounded uniformly by PY,X,pP_{Y,X,p}6, the bias bound

PY,X,pP_{Y,X,p}7

is sharp, attained at PY,X,pP_{Y,X,p}8. The bound has a signal-to-noise interpretation: larger PY,X,pP_{Y,X,p}9 means a more discriminating score, so the same calibration error produces proportionally less bias, and the bound diverges as Gi∈{0,1}G_i \in \{0,1\}0, consistent with the identification failure result.

Monte Carlo evidence

Five simulation exercises, each tied to a theoretical result, use 2,000 replications with a baseline design featuring logistic Gi∈{0,1}G_i \in \{0,1\}1, linear Gi∈{0,1}G_i \in \{0,1\}2, Beta-distributed scores calibrated so that Gi∈{0,1}G_i \in \{0,1\}3 is known exactly, and Gi∈{0,1}G_i \in \{0,1\}4.

Finite-sample performance. With Gi∈{0,1}G_i \in \{0,1\}5 and Gi∈{0,1}G_i \in \{0,1\}6, the oracle estimator is approximately unbiased at all sample sizes (bias between Gi∈{0,1}G_i \in \{0,1\}7 and Gi∈{0,1}G_i \in \{0,1\}8) with coverage near 0.95. The plug-in estimator without cross-fitting exhibits persistent positive bias of roughly 0.12–0.17 attributable to in-sample overfitting, while the cross-fitted orthogonal estimator is nearly unbiased with coverage close to nominal. QQ-plots confirm oracle normality from Gi∈{0,1}G_i \in \{0,1\}9 onward.

Identification boundary. As z=2p−1z = 2p-10 decreases so that z=2p−1z = 2p-11 falls from z=2p−1z = 2p-12 to z=2p−1z = 2p-13, RMSE grows by five orders of magnitude (from 0.164 to roughly 24,853), tracking the theoretical z=2p−1z = 2p-14 rate on a log-log scale—while coverage remains near 0.95 throughout because confidence intervals correctly widen.

Calibration failure. The worst-case error shape nearly attains the sharp bound, with tightness ratios of 0.86–0.99; the symmetric shape produces empirical bias indistinguishable from zero at every z=2p−1z = 2p-15.

Hard-threshold attenuation. Thresholding at z=2p−1z = 2p-16 and comparing cell means converges to z=2p−1z = 2p-17 with z=2p−1z = 2p-18 under z=2p−1z = 2p-19 and conditional symmetry of the score residual. Attenuation is severe: at ii00 the threshold estimator recovers approximately 0.08 against a true value of 1, while the moment estimators remain centred on the truth. The moment approach strictly dominates threshold classification whenever classification is imperfect.

Heterogeneous effects. When the effect varies as ii01, the moment equation identifies the variance-weighted average ii02 rather than ii03. Design B, with score variance proportional to ii04, yields ii05 versus a simple mean of 1, and both designs recover their respective ii06 with negligible bias. The weight function upweights units whose scores vary substantially beyond what covariates predict—units carrying more identifying information.

Limitations and open questions

Several limitations are stated plainly in the paper. The constant-coefficient restriction in Assumption 1 is substantive; relaxing it changes the estimand to the variance-weighted average rather than eliminating the issue. The marginal gap remains unidentified without knowledge of latent-group covariate distributions. The hard-threshold attenuation factor ii07 requires both ii08 almost surely and conditional symmetry of the score residual; without these, the closed form does not hold. The plug-in estimator's positive finite-sample bias stems from the absence of cross-fitting, and formal ii09-normality of the Neyman-orthogonal estimator under cross-fitting, semiparametric efficiency of the oracle estimator, and sharper sensitivity bounds under shape restrictions on ii10 are all left unresolved.

Conclusion

The paper establishes that a structural latent-group coefficient is point-identified by a closed-form ratio of weighted moments whenever a conditionally calibrated score carries residual variation beyond covariates; that identification fails precisely, and constructively, when that variation vanishes; and that the identified object is a within-cell effect distinct from the marginal gap unless groups are covariate-balanced. Oracle inference is standard, miscalibration bias admits an exact formula and a sharp sensitivity bound scaling inversely with residual score variance, and the moment approach dominates hard-threshold classification. The most consequential open problem is verifying ii11-normality of the Neyman-orthogonal estimator under cross-fitting, which would extend the framework to flexible machine-learning nuisance estimators.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.