- The paper rigorously demonstrates that omitting key covariates induces selection bias in hazard ratios, challenging causal interpretation in Cox models.
- It uses comprehensive simulations to quantify bias across various distributions and shows that correctly specified frailty models can recover true exposure effects.
- The study recommends employing frailty and AFT models as robust alternatives to naive Cox models for more reliable survival analysis.
Selection Bias and Non-Collapsibility in Proportional Hazards Models with Omitted Covariates
Context and Formalization
The Cox Proportional Hazards Model (PHM) remains a dominant tool in time-to-event analysis, principally due to its ability to estimate hazard ratios (HR) as measures of association. However, the paper rigorously reviews and formalizes intrinsic biases occurring when relevant covariates are unmeasured and omitted, even in randomized trial contexts. Specifically, hazard ratios are non-collapsible, owing to their implicit conditioning on survival, which can induce a "collider effect"—the formation of a spurious association between treatment (X) and omitted confounder (U) through survival at time t. This phenomenon leads to time-evolving imbalance in U across risk sets, breaking randomization and introducing structural bias into the HR for X.

Figure 1: A causal diagram illustrating the relationships between treatment, omitted confounder, and survival as a collider variable.
Mathematically, the paper demonstrates that for omitted U, the marginal hazard function at each time t is given by:
λM(t∣X=x)=λ0,C(t)exp(βCx)EU∣T>t,X=x[exp(βUU)]
and that the marginal HR deviates from the conditional HR (exp(βC)), making the HR a non-collapsible association measure. Analytically tractable cases are provided for gamma frailty models; when the variance of unmeasured heterogeneity increases, the discrepancy between marginal and conditional HR becomes pronounced.

Figure 2: Dynamics of the distribution of U among at-risk individuals stratified by treatment group: selection leads to increasing imbalance over time.
Simulation Evidence and Quantification
Comprehensive simulations quantify the bias for various PHM settings and distributions of U0. When U1 follows a normal, log-gamma, or Bernoulli distribution, omission induces substantial bias in HR estimates, particularly when U2 is large. Coverage rates for confidence intervals in naive (unadjusted) Cox and parametric PH models deteriorate as the impact of U3 rises.

Figure 3: Simulation results for U4. Boxplots summarize bias/coverage for models across increasing U5.

Figure 4: Simulation results for U6 log-gamma distributed. Models with gamma frailty perform robustly when the frailty matches the true unmeasured heterogeneity.

Figure 5: Simulation results for U7. Naive bias is markedly lower for discrete U8.
Marginal survival differences, estimated via Kaplan-Meier or Cox with time-dependent effects, remain unbiased and maintain nominal coverage even under substantial omitted heterogeneity.

Figure 6: Simulation of marginal survival differences with robust coverage for both estimation approaches.
Alternative Models: Frailty, AFT, Survival Differences
Frailty Models
Frailty models introduce individual-specific random effects to absorb unmeasured heterogeneity. Parametric frailty models (gamma-distributed frailty) recover the true exposure effect U9 when the frailty matches omitted t0’s structure, and are robust to moderate misspecification. Semi-parametric frailty models, typically designed for clustered data, underestimate standard errors and are susceptible to bias when used for individual frailty, particularly as the variance of unmeasured heterogeneity increases.
Accelerated Failure Time (AFT) Models
AFT models, which parameterize log-survival time as a linear function of covariates, yield collapsible effect measures irrespective of omitted t1, provided the error distribution is correctly specified. Robustness improves when employing log-normal or log-logistic error distributions or flexible parametric spline models. The equivalence between Weibull PHM and AFT is established, highlighting practical gains in causal interpretability and estimation stability.
Survival Differences
Estimates of marginal survival differences between exposure groups, as opposed to hazard ratios, are collapsible and not affected by built-in selection bias. They require only nonparametric (Kaplan-Meier) or time-dependent Cox modeling, and provide absolute risk contrasts suitable for interpretation.
Empirical Application: RTOG 9202 Trial
In a randomized controlled setting (RTOG 9202), all models converge on a significant treatment effect, but naive unadjusted Cox and PH models underestimate the exposure effect relative to parametric frailty models. Adjustment for baseline covariates reduces frailty variance and aligns naive and frailty model estimates, confirming that observed covariates absorb much of the heterogeneity previously attributed to frailty. AFT models exhibit remarkable stability across adjusted/unadjusted specifications, notably with log-normal and log-logistic errors. Flexible AFT models are robust in simulation but less numerically stable in complex datasets.
Implications and Future Directions
The findings compel a critical reevaluation of the reliability of hazard ratios in contexts of unmeasured covariates, especially when causal interpretation is needed. Analytical and simulation evidence unequivocally document selection bias as a function of the structural properties of survival analysis models, not simply confounding.
Pragmatically, frailty and AFT models represent robust alternatives, with frailty models requiring careful specification of the baseline hazard and frailty distribution, while AFT models demand accurate error modeling. Marginal survival differences, though robust, are sensitive to the temporal granularity and adjustment variable selection.
Theoretically, future work should advance inference techniques for semi-parametric frailty models, facilitate subject-specific frailty estimation, and develop flexible modeling strategies for baseline risks. These improvements would enable broader applicability and greater reliability in time-to-event analyses facing unmeasured heterogeneity.
Conclusion
This paper presents rigorous theoretical, simulation-based, and empirical arguments demonstrating built-in structural bias in hazard ratio estimation from Cox PH models with omitted covariates. Frailty and AFT models, as well as marginal survival differences, furnish robust alternatives for estimating exposure effects in time-to-event studies. The results underscore the necessity of deploying appropriate models for causal interpretation and motivate further methodological refinement in survival analysis frameworks.