- The paper shows that the Aalen-based difference method can be unbiased with time-dependent mediators and common outcomes when treatment does not cause time-dependent confounders, while the Cox version additionally requires rare outcomes.
- The paper finds that all difference-method variants can be severely biased when treatment causes a time-dependent confounder, and that AFT-based estimates remain slightly biased because time-dependent mediators undermine collapsibility.
- The paper’s simulations and UDCA trial reanalysis show that the parametric mediational g-formula is the safer general-purpose option, although the difference method can offer lower variance and simpler implementation under restrictive assumptions.
Motivation and scope
Causal mediation analysis decomposes a total treatment effect into direct and indirect components, but its application to settings combining a time-to-event outcome, a time-dependent mediator, and time-dependent confounders remains methodologically demanding. Existing approaches—parametric mediational g-formula, dynamic path analysis, multi-state and semi-competing risk models, Bayesian joint models, and targeted maximum likelihood estimation—are statistically complex, computationally expensive, and often lack user-friendly software implementations. Against this backdrop, the difference method retains considerable practical appeal: it requires only two regression models (a total effects model adjusting for baseline confounders, and a direct effects model additionally including the time-dependent mediator and confounders), with the indirect effect obtained by coefficient subtraction. Despite prior evidence of bias in simpler settings—non-collapsibility of the Cox model, the rare-outcome requirement, and structural bias under treatment-induced time-dependent confounding—no systematic evaluation of the method's empirical performance with time-varying mediators had been conducted. Denz and Timmesfeld address this gap through an extensive simulation study and an illustrative reanalysis of a randomized trial of Ursodeoxycholic acid (UDCA) in primary biliary cholangitis (PBC).
Target estimands and identifiability
A central contribution is a careful articulation of the estimand. Natural effects defined via nested counterfactuals Ya,Ma′ are ill-defined here: survival differences across counterfactual worlds can truncate the mediator process, and post-treatment confounding renders natural effects non-identifiable. The authors instead adopt path-specific effects following Vansteelandt et al., recursively defining counterfactual processes Lta,a′, Mta,a′, and Yta,a′ in which the confounder process evolves under a while the mediator process evolves under a′. Total, direct, and indirect effects are then contrasts of hazards or expected log survival times across these counterfactuals, e.g., ΨIE(t)=log(λ1,1(t)/λ1,0(t)). Identification requires consistency, positivity, sequential ignorability, and cross-world independence; notably, if the at-risk indicator is included among the time-dependent confounders, the post-treatment confounding problem is resolved within this definition.
Simulation design
Data were generated via a Gillespie-type discrete-event simulation with binary time-dependent variables Lt (confounder), Mt (mediator), and Yt (outcome), plus baseline covariates Lta,a′0 and Lta,a′1. Six scenarios progressively introduce complexity: randomized treatment with Lta,a′2 a pure predictor (scenario 1); Lta,a′3 as a time-dependent confounder of the mediator–outcome relationship (scenario 2); an Lta,a′4–Lta,a′5 feedback loop (scenario 3); direct causation of Lta,a′6 by treatment, making Lta,a′7 both confounder and mediator (scenario 4); baseline confounding of treatment–mediator (scenario 5) and treatment–outcome (scenario 6) relationships. Each scenario was crossed with presence/absence of mediation, low versus high baseline event probability, and approximately 30% right-censoring, with Lta,a′8 and 1000 replications. The outcome process was specified to be consistent with Cox/AFT models in one set of runs and with Aalen additive hazards models in another, avoiding inherent model misspecification. True target values were computed using Monte Carlo integration over counterfactual simulations with known structural equations, following Naimi et al.
Key simulation findings
The results delineate sharply which conditions permit unbiased indirect effect estimation:
| Method |
Unbiased conditions |
Bias sources |
| Cox-based difference |
Rare outcome; no Lta,a′9 effect; correct adjustment |
Common outcomes; treatment-caused Mta,a′0 |
| Aalen-based difference |
No Mta,a′1 effect; correct adjustment (any event rate) |
Treatment-caused Mta,a′2 |
| AFT-based difference |
None — small bias in nearly all scenarios |
Non-collapsibility on marginal scale |
| Parametric mediational g-formula |
All scenarios |
None observed |
Three findings deserve emphasis. First, the Aalen model based difference method was unbiased even with common outcomes in scenarios 1–3, 5, and 6, extending its known collapsibility advantage from time-fixed to time-varying mediators. Second, the Cox-based version inherits the rare-outcome requirement: with high baseline hazard, biased estimates arose regardless of whether an indirect effect existed, consistent with Lange and colleagues' earlier results for time-fixed mediators. Third, and most consequential, all versions of the difference method were severely biased in scenario 4, where treatment directly caused the time-dependent confounder. The mechanism is structural rather than statistical: conditioning on Mta,a′3 removes mediator–outcome confounding but also blocks genuine direct pathways through Mta,a′4, while omitting Mta,a′5 leaves confounding unadjusted. This implies that, unlike the time-fixed case, the difference method cannot generally serve even as a screening test for the presence of time-dependent mediation without additional assumptions.
An unexpected result concerns the AFT model. Although the data-generating process was mathematically consistent with a Weibull AFT model, and results persisted without censoring—ruling out convergence failure and censoring-induced misspecification—the AFT-based difference method exhibited small but persistent bias throughout. Diagnosis showed the total effects model correctly recovered the marginal total effect, while the direct effects model consistently estimated the conditional coefficient from the DGP, which differs from the marginal direct effect. This contradicts prior findings that AFT coefficients are collapsible with time-fixed mediators, indicating that collapsibility fails once the mediator is time-dependent.
Wherever the difference method was unbiased, it achieved lower mean squared error than the g-formula, though part of the g-formula's variance reflects the single Monte Carlo replication used per point estimate—a limitation the authors acknowledge, since averaging multiple replications would reduce variance in practice.
Illustrative application to UDCA in PBC
The motivating analysis used 170 patients from the Mayo Clinic UDCA trial, treating histological progression by two stages as the time-dependent mediator, liver transplant-free survival as the outcome, and varices appearance as a time-varying covariate. Across the hazard ratio, survival time ratio, and cumulative hazard difference scales, the two methods agreed closely: a protective total effect (hazard ratio approximately 0.63; survival time ratio 1.41), a similar direct effect, and an indirect effect near the null—with all 95% bootstrap confidence intervals covering the reference value. The authors are explicit that this example cannot support clinical conclusions: the sample yields only 28 events, far too few for adequately powered mediation analysis (a trial powered for the main effect will typically be underpowered for decomposition); the assumption that UDCA has no direct effect on time-dependent confounders such as varices is doubtful; and sequential ignorability likely fails given that not all candidate time-varying confounders could be adjusted for. The near-equivalence of the two methods' estimates may itself reflect violation of these shared assumptions.
Limitations and open questions
The authors state plainly that their favorable findings for the Cox and Aalen versions do not constitute proof: behavior under extreme parameter values or alternative DGPs remains unexplored, and precisely how rare the outcome must be for the Cox-based method remains undetermined. Several complications were deliberately excluded—treatment–mediator interaction, multiple mediators, dependent censoring, competing risks, truncation, measurement error, and time-varying treatments—each of which would further degrade the difference method's performance. The censoring framework also embeds a subtle choice: because the estimand involves no intervention on censoring, it implicitly depends on the censoring mechanism, and when the mediator distribution shifts over follow-up, right-censoring changes the time-average targeted by constant-coefficient total effects models, potentially requiring time-dependent coefficients. Finally, the field lacks a neutral comparison study establishing which specialized method is most appropriate under which conditions—an open question the authors identify as the priority before further method proliferation.
Conclusion
This paper provides the first systematic evaluation of the difference method for mediation analysis with a time-dependent mediator, time-dependent confounders, and a survival outcome. Its verdict is conditional rather than categorical: the Aalen-based variant is unbiased whenever no time-dependent confounder is directly caused by treatment, the Cox-based variant adds a rare-outcome requirement, the AFT-based variant is never fully unbiased due to non-collapsibility on the marginal scale, and all variants fail structurally when treatment causes the time-dependent confounder. Where these restrictive assumptions can be credibly defended—as possibly in vaccine-mediated dementia prevention—the difference method offers a computationally trivial and efficient alternative suited to finely discretized or continuous time. In general practice, however, the parametric mediational g-formula, which was unbiased across every scenario examined, remains the safer default despite its greater implementation burden.