Group-Label Shift Assumption
- Group-label shift assumption is a framework where the conditional feature-generation mechanism stays invariant while label or group prevalences change across domains.
- It encompasses variants such as classical label shift, GPPS, and representation-space GLS, each clarifying different aspects of distribution shifts.
- The assumption underpins applications in domain adaptation, fairness analysis, and survival modeling with strategies for estimation and posterior correction.
The group-label shift assumption denotes a family of distribution-shift models in which the change between a source domain and a target domain is localized to label prevalences, group-conditioned label prevalences, or a joint prior over labels and group variables, while a conditional feature-generation mechanism is assumed invariant across domains. In the recent literature, this idea appears under several closely related names, including classical label shift, group-conditional prior probability shift (GPPS), expanded label shift over , survival-time label shift, and generalized label shift (GLS) in representation space. Across these variants, the common structure is that the shift is not arbitrary: some conditional law such as , , , , or remains fixed, while marginal or group-level composition changes (Zong et al., 26 Jun 2025, Asiaee et al., 5 Feb 2026, Sun et al., 2022, Cheng et al., 26 Sep 2025, Luo et al., 2024).
1. Core meaning and taxonomic position
In its classical two-population form, label shift assumes that the label distribution changes while the conditional distribution of features given the label remains invariant. One formulation writes
so that
This is also described as prior probability shift, target shift, or class prior change (Lee et al., 2023).
Group-aware variants refine that template rather than discard it. Under GPPS, the stable object is , and the changing object is the group-specific prevalence . The defining condition is
0
while
1
in general (Asiaee et al., 5 Feb 2026). Under expanded label shift, the “label” is enlarged to the meta-label 2, so the invariant mechanism becomes
3
while the shifted quantity is the joint prior 4 rather than 5 alone (Sun et al., 2022).
A further extension partitions covariates as 6, with 7 interpreted as group or subpopulation features. The group-label shift assumption in that setting is
8
so the domain shift occurs through the joint distribution of 9, not through the conditional law of 0 given 1 (Cheng et al., 26 Sep 2025).
This family is distinct from covariate shift. Under covariate shift, the change is in 2, while the invariant quantity is 3: 4 The contrast is structural: group-label shift fixes a feature-given-label mechanism, whereas covariate shift fixes a label-given-feature mechanism (Rezaei et al., 2020).
2. Canonical formalizations
Representative formulations differ in which variable is treated as the label and which conditional distribution is held invariant. The table summarizes the main versions that recur in recent work.
| Variant | Invariant conditional | Shifted quantity |
|---|---|---|
| Classical label shift | 5 | 6 |
| GPPS | 7 | 8 |
| Expanded label shift | 9 | 0 |
| Survival label shift | 1 | 2 |
| Representation-space GLS | 3 | 4 |
These formulations are all explicit in the recent literature (Zong et al., 26 Jun 2025, Tachet et al., 2020, Luo et al., 2024, Sun et al., 2022, Asiaee et al., 5 Feb 2026).
A multi-domain unlabeled version is latent label shift. There, domains 5 act as groups, and the assumption is
6
with domain-specific class priors 7. The corresponding matrix relation for discrete inputs is
8
which the paper identifies with a topic-model factorization in which inputs correspond to words, domains to documents, and labels to topics (Roberts et al., 2022).
Representation-space GLS shifts the invariance target from the raw input space to a learned representation 9. Its defining condition is
0
or equivalently 1. This relaxes classical input-space label shift while preserving the same conditional-invariance logic (Tachet et al., 2020, Luo et al., 2024).
3. Survival-time labels and censoring
A notable recent extension applies label shift to survival analysis, where the “label” is a continuous event time 2 and the source sample is censored. The source population 3 has observed 4, with
5
while the target population 6 has only covariates 7. The label-shift condition is
8
so the shift is entirely in the event-time marginal, not in the 9 mechanism (Zong et al., 26 Jun 2025).
The joint laws are written as
0
A direct consequence is that the target marginal covariate distribution becomes a mixture induced by the shifted event-time law: 1 whereas the source marginal is
2
This makes the survival formulation recognizably label-shift-like: the latent label distribution shifts, and the observed feature marginal changes only through mixture reweighting.
Censoring complicates the problem because the shifted quantity 3 is not directly observed in the source. The paper assumes independent right censoring,
4
and a support condition,
5
For inference in 6, it specifies a parametric target model 7 and constructs an approximate likelihood whose single-observation log-likelihood is
8
The unknown distributions are then replaced nonparametrically: 9 by the Kaplan–Meier estimator from the censored source sample, and 0 by the empirical distribution of target covariates. The paper describes this as the first work to combine survival analysis with label shift (Zong et al., 26 Jun 2025).
4. Group-conditioned prior shift and fairness
In fairness analysis, the relevant specialization is GPPS. The source domain is historical or training data, the target domain is the deployment environment, and the variables are 1 with 2 and 3 a sensitive group attribute. GPPS assumes that for all groups 4 and labels 5,
6
while the group-specific positive-class prevalence
7
may differ from
8
The group marginal 9 may also change (Asiaee et al., 5 Feb 2026).
This structure yields a sharp dichotomy between fairness notions. Equalized odds depends only on conditional error rates given 0 and 1, so it is invariant under GPPS. The central mechanism is ROC invariance: 2 By contrast, demographic parity and predictive parity can drift because they depend on prevalence. The key acceptance-rate identity is
3
and the positive predictive value is
4
Both are explicit functions of group prevalence.
The paper proves a shift-robust impossibility result for demographic parity: if a threshold classifier satisfies demographic parity under two different prevalence regimes, then either 5 for all groups, or the prevalence shifts obey a highly specific constraint. A plausible implication is that demographic parity is generically non-invariant under groupwise prior shift, whereas equalized odds is structurally invariant.
GPPS also supports label-free target-domain estimation. The target 6-7 risk can be identified as
8
Building on this, TAP-GPPS proceeds in three steps: estimate target prevalences from unlabeled target data by EM or BBSE within each group, correct posteriors, and choose group thresholds by bisection so that
9
The method therefore enforces demographic parity in the target domain using labeled source data and unlabeled target data alone (Asiaee et al., 5 Feb 2026).
5. Representation learning, domain adaptation, and latent classes
In domain adaptation, generalized label shift relocates the invariance assumption to a learned representation 0. One formulation defines GLS by
1
and introduces class weights
2
A key necessary condition is that the target feature distribution match a reweighted source distribution rather than the unweighted source marginal. This is why importance-weighted variants of DANN, JAN, and CDAN align 3 with 4, where 5 may be 6 or 7 (Tachet et al., 2020).
Subsequent theory strengthens that point. Under label shift, invariant representation learning alone is insufficient: there is no transformation 8 such that both 9 and 0 hold when the label marginals differ, and more strongly, no transformation 1 can make the full joint law match, since
2
The corresponding sufficiency statement is that conditional invariance must be paired with explicit label correction 3 (Luo et al., 2024).
A parallel methodological response appears in conditional support alignment for UDA. Rather than aligning marginal latent supports, CASA minimizes a class-conditional analogue, the conditional symmetric support divergence, because marginal alignment can scramble class structure when 4. The training objective combines source classification loss, target entropy minimization, a local Lipschitz regularizer via virtual adversarial training, and alignment on the joint representation 5 using pseudo-labels on the target side (Nguyen et al., 2023).
Latent label shift uses the same assumption for an opposite purpose: not correction but identification. With multiple unlabeled domains and invariant 6, the paper’s slogan is “Elements that shift together group together.” Its practical DDFA pipeline trains a domain discriminator for 7, clusters in discriminator-output space, performs non-negative matrix factorization on the discretized data, and then reconstructs 8. Under anchor or separable regions, this yields identifiability up to label permutation (Roberts et al., 2022).
6. Estimation strategies, operational uses, and limits
The assumption is operationalized in diverse ways. In targeted federated learning, clients 9 and target 00 satisfy
01
while client label marginals differ from each other and from the target. FedPALS exploits a known target label distribution 02 at the server and chooses a target-aware convex combination of client updates. When exact target matching is impossible, its weights solve a regularized projection problem that trades off label matching against effective sample size (ESS). The paper is explicit that the method is exact and unbiased only in the SGD/single-step setting and only when the target label distribution is representable by the client label marginals (Zec et al., 2024).
Expanded label shift yields a test-time adaptation procedure rather than a training-time one. TTLSA treats 03 as the label, assumes
04
and estimates the shifted prior 05 from unlabeled target data by EM using a calibrated source model 06. The posterior correction takes the form
07
after which prediction marginalizes over 08 (Sun et al., 2022).
Statistical inference under label shift has also moved beyond point prediction. One approach avoids direct estimation of the density ratio 09 and instead constructs influence-function-based estimators that remain valid even when both a working density-ratio model and a working conditional model are misspecified. Another recent development models the outcome density ratio between labeled and unlabeled data and estimates it by a progressive strategy consisting of an initial heuristic guess, a consistent estimation, and ultimately, an efficient estimation, with an explicit connection to prediction-powered inference (Lee et al., 2023, Lee et al., 25 Aug 2025).
The same structural assumption can even reduce a sequential changepoint problem to one dimension. Under binary label shift, the likelihood ratio satisfies
10
so classifier scores become sufficient surrogates for the post-change likelihood ratio in CUSUM- or Shiryaev–Roberts-type detectors (Evans et al., 2020).
Several recurring limitations follow directly from the cited formulations. The assumption is not a model of arbitrary distribution shift; it is valid only when the relevant class-conditional or group-conditional feature mechanism is stable. Many methods require additional structure, such as 11 in censored survival analysis, the same label set for 12 across domains in TTLSA, or direct knowledge of 13 at the server in FedPALS (Zong et al., 26 Jun 2025, Sun et al., 2022, Zec et al., 2024). Some empirical work uses the language of label shift more loosely, for example by perturbing target weather-class proportions without stating a full probabilistic invariance theorem, which suggests an operational rather than formal group-label-shift setup (Marathe et al., 2023). Robust optimization work introduces yet another meaning, in which target differences are assumed to occur through group marginals and true group indicators are realizable by bitrate-constrained functions conditioned on the label, again emphasizing that the phrase “group-label shift” is not fully standardized across subfields (Setlur et al., 2023).