- The paper identifies a constant latent-group effect with a closed-form moment ratio when the calibrated score retains variation beyond covariates, requiring residual score variance greater than zero.
- The paper shows identification fails exactly when the score is fully determined by covariates, while the identified structural effect generally differs from the marginal group mean gap because of covariate composition.
- The paper derives sharp miscalibration-bias bounds and simulation evidence showing that moment estimators outperform hard-threshold classification, while cross-fitting improves finite-sample performance with flexible nuisance estimates.
Setting and motivation
Empirical work frequently requires estimating outcome differences between groups whose membership is never directly observed—poverty status, informal employment, latent health conditions. The analyst instead observes a probability score pi​∈[0,1] encoding belief that unit i belongs to the group of interest. The paper "Identification of Latent Group Effects under Conditional Calibration" (2604.08798) asks a precise question: under what conditions, and by what formula, can a structural group effect be identified from the joint law of observables (Y,X,p) when the binary indicator G is never observed? The answer is organized around three claims: a structural coefficient τ is point-identified under mild conditions; identification fails in an exactly characterizable way when one condition is violated; and the identified object is distinct from the marginal group mean gap.
The setting complements three literatures. In the misclassification literature (Lewbel 2007; Mahajan 2006; Kasahara and Shimotsu 2022), the analyst observes a noisy binary label; here the analyst observes a calibrated probability, which changes both the identification argument and the estimand. Relative to proxy-variable and measurement-error approaches requiring rank conditions on integral operators (Hu and Schennach 2008; Schennach 2016), the target is only the scalar Ï„, and the assumptions are correspondingly weaker with a closed-form formula. Within algorithmic fairness research on disparity estimation with unobserved protected attributes (Kallus et al. 2022; Chen et al. 2019), the contribution is a formal identification-theoretic treatment with exact failure and sensitivity characterizations.
Model and assumptions
The data are i.i.d. draws (Yi​,Xi​,pi​) from PY,X,p​, with an unobserved Gi​∈{0,1}. Three derived quantities are central: the signed score z=2p−1, the outcome residual i0 where i1, and the score residual i2 where i3. The key scalar governing everything is
i4
the residual variance of the score after partialling on covariates.
Two structural assumptions carry the analysis. Assumption 1 posits a constant-coefficient conditional mean: i5. This bundles two restrictions—the group effect is constant in i6, and the score carries no additional information about i7 once i8 are known. Assumption 2 is conditional calibration: i9 almost surely. Notably, (Y,X,p)0 need not equal the propensity score (Y,X,p)1; it must merely be an unbiased predictor of (Y,X,p)2 given all observed information. Area-level prevalence rates, classifier outputs, and model-based predictions satisfy this when well-calibrated. Identification requires only second moments ((Y,X,p)3); asymptotic normality uses fourth moments.
Two lemmas form the backbone. First, calibration plus the tower property forces (Y,X,p)4, so the observable regression satisfies (Y,X,p)5. Second, the outcome residual decomposes as (Y,X,p)6 with (Y,X,p)7: the entire predictable part of (Y,X,p)8 is driven by the deviation of true membership from its score-implied expectation.
Point identification and its exact boundary
The central result is the population moment identity
(Y,X,p)9
which yields point identification by the closed-form ratio
G0
The formula has a transparent instrumental-variables interpretation: the score residual G1 acts as an instrument for the latent deviation G2. Calibration supplies first-stage relevance—a covariance identity shows the population regression slope of G3 on G4 equals exactly one—and mean-independence in Assumption 1 supplies the exclusion restriction. The estimand is thus the IV ratio of reduced-form to first-stage slopes.
Assumption 3 (G5) is not a mere regularity condition but the exact boundary of identification. If G6—equivalently, if and only if G7 almost surely, i.e., the score is a deterministic function of G8—then G9 is not identified. The failure is made constructive: for any τ0, one can adjoin τ1 for independent uniform τ2, set τ3, and obtain a model satisfying both assumptions that reproduces the same observable regression τ4 and hence the same moment equation. There is therefore an explicit continuum of observationally equivalent models indexed by arbitrary values of τ5. A practical implication is sharp: whenever a practitioner's score is fully explained by covariates, no estimator—however clever—can recover the group effect without further structure.
Structural coefficient versus marginal gap
A natural concern is whether τ6 equals the marginal latent mean gap τ7. It does not, in general. Under the constant-coefficient model,
τ8
where τ9 captures differences in covariate composition across latent groups. Because many latent joint distributions of τ0 generate the same observable law of τ1, τ2 is unidentified without further assumptions. The two objects coincide if and only if the latent groups are covariate-balanced in the sense that τ3. The practical import is that τ4 identifies the within-covariate-cell group effect, while the marginal gap conflates this effect with composition—an important caveat for applications aiming at headline disparity numbers.
Oracle inference and plug-in estimation
With known nuisance functions τ5 and τ6, the oracle estimator is the sample analogue of the identification ratio. It is τ7-consistent and asymptotically normal with closed-form sandwich variance
τ8
obtained by applying the delta method to the ratio of two sample means; all cross-terms cancel in terms of the centred score. A consistent variance estimator yields valid Wald intervals.
When τ9 and (Yi​,Xi​,pi​)0 are unknown, replacing them by estimators converging in (Yi​,Xi​,pi​)1 gives a consistent plug-in estimator; denominator stability under nuisance estimation follows from a boundedness argument since (Yi​,Xi​,pi​)2 and (Yi​,Xi​,pi​)3 lie in (Yi​,Xi​,pi​)4.
The paper is candid about the limits of this result. Gateaux derivative calculations show the score fails Neyman orthogonality through its (Yi​,Xi​,pi​)5-direction (the (Yi​,Xi​,pi​)6-direction vanishes automatically because (Yi​,Xi​,pi​)7). Replacing (Yi​,Xi​,pi​)8 by (Yi​,Xi​,pi​)9 in the numerator produces a Neyman-orthogonal score in both directions, defining a distinct orthogonal estimator that is asymptotically equivalent to the plug-in when nuisances are known but generally not otherwise. Establishing PY,X,p​0-normality of this orthogonal estimator under cross-fitting—i.e., placing it formally within the double machine learning framework—is explicitly left open, as is whether the oracle estimator attains the semiparametric efficiency bound.
Robustness to calibration failure
If calibration is violated so that PY,X,p​1, the probability limit of the oracle estimator is exactly PY,X,p​2 with
PY,X,p​3
Three features deserve emphasis. The bias is proportional to PY,X,p​4: no bias arises when the true effect is zero, regardless of miscalibration. It vanishes whenever the miscalibration is orthogonal to the signed score—for instance, symmetric errors such as PY,X,p​5. And over the class of all calibration error functions bounded uniformly by PY,X,p​6, the bias bound
PY,X,p​7
is sharp, attained at PY,X,p​8. The bound has a signal-to-noise interpretation: larger PY,X,p​9 means a more discriminating score, so the same calibration error produces proportionally less bias, and the bound diverges as Gi​∈{0,1}0, consistent with the identification failure result.
Monte Carlo evidence
Five simulation exercises, each tied to a theoretical result, use 2,000 replications with a baseline design featuring logistic Gi​∈{0,1}1, linear Gi​∈{0,1}2, Beta-distributed scores calibrated so that Gi​∈{0,1}3 is known exactly, and Gi​∈{0,1}4.
Finite-sample performance. With Gi​∈{0,1}5 and Gi​∈{0,1}6, the oracle estimator is approximately unbiased at all sample sizes (bias between Gi​∈{0,1}7 and Gi​∈{0,1}8) with coverage near 0.95. The plug-in estimator without cross-fitting exhibits persistent positive bias of roughly 0.12–0.17 attributable to in-sample overfitting, while the cross-fitted orthogonal estimator is nearly unbiased with coverage close to nominal. QQ-plots confirm oracle normality from Gi​∈{0,1}9 onward.
Identification boundary. As z=2p−10 decreases so that z=2p−11 falls from z=2p−12 to z=2p−13, RMSE grows by five orders of magnitude (from 0.164 to roughly 24,853), tracking the theoretical z=2p−14 rate on a log-log scale—while coverage remains near 0.95 throughout because confidence intervals correctly widen.
Calibration failure. The worst-case error shape nearly attains the sharp bound, with tightness ratios of 0.86–0.99; the symmetric shape produces empirical bias indistinguishable from zero at every z=2p−15.
Hard-threshold attenuation. Thresholding at z=2p−16 and comparing cell means converges to z=2p−17 with z=2p−18 under z=2p−19 and conditional symmetry of the score residual. Attenuation is severe: at i00 the threshold estimator recovers approximately 0.08 against a true value of 1, while the moment estimators remain centred on the truth. The moment approach strictly dominates threshold classification whenever classification is imperfect.
Heterogeneous effects. When the effect varies as i01, the moment equation identifies the variance-weighted average i02 rather than i03. Design B, with score variance proportional to i04, yields i05 versus a simple mean of 1, and both designs recover their respective i06 with negligible bias. The weight function upweights units whose scores vary substantially beyond what covariates predict—units carrying more identifying information.
Limitations and open questions
Several limitations are stated plainly in the paper. The constant-coefficient restriction in Assumption 1 is substantive; relaxing it changes the estimand to the variance-weighted average rather than eliminating the issue. The marginal gap remains unidentified without knowledge of latent-group covariate distributions. The hard-threshold attenuation factor i07 requires both i08 almost surely and conditional symmetry of the score residual; without these, the closed form does not hold. The plug-in estimator's positive finite-sample bias stems from the absence of cross-fitting, and formal i09-normality of the Neyman-orthogonal estimator under cross-fitting, semiparametric efficiency of the oracle estimator, and sharper sensitivity bounds under shape restrictions on i10 are all left unresolved.
Conclusion
The paper establishes that a structural latent-group coefficient is point-identified by a closed-form ratio of weighted moments whenever a conditionally calibrated score carries residual variation beyond covariates; that identification fails precisely, and constructively, when that variation vanishes; and that the identified object is a within-cell effect distinct from the marginal gap unless groups are covariate-balanced. Oracle inference is standard, miscalibration bias admits an exact formula and a sharp sensitivity bound scaling inversely with residual score variance, and the moment approach dominates hard-threshold classification. The most consequential open problem is verifying i11-normality of the Neyman-orthogonal estimator under cross-fitting, which would extend the framework to flexible machine-learning nuisance estimators.