Papers
Topics
Authors
Recent
Search
2000 character limit reached

From Cumulative Weights to Marginal Density Ratios: Per-Protocol Estimation in Sequential Target Trial Emulation

Published 21 Aug 2026 in stat.ME | (2608.20976v1)

Abstract: Sequential target trial emulation evaluates eligibility at multiple baseline times to emulate a sequence of randomized trials using observational data. Estimating per-protocol effects in this setting is challenging because treatment deviations and loss to follow-up induce selection among individuals who remain observed and adherent over time. Conventional inverse-probability methods address this selection using cumulative weights constructed from estimated adherence and censoring probabilities, but these weights can be highly variable, leading to unstable and imprecise effect estimates. We propose a different approach based on marginal density ratios (MDRs). The MDR directly compares the state distribution among individuals who would remain event-free under a target treatment strategy with the corresponding distribution among observed-adherent individuals. We use longitudinal g-computation to generate the target risk sets and a probabilistic classifier to estimate density ratios for reweighting the observed outcomes. Building on this approach, we also develop a doubly robust extension. Favorable performance across the simulation study suggests that MDR weighting is a promising alternative to cumulative longitudinal weights when its identification assumptions are plausible.

Summary

  • The paper presents an alternative approach, STTE-MDR, replacing cumulative adherence-and-censoring weights with marginal density ratios (MDR) for per-protocol estimation in sequential target trials, fundamentally improving estimation stability and reducing bias.
  • STTE-MDR demonstrated significantly lower RMSE in various scenarios, particularly in settings with poor positivity. However, the method's performance can degrade in small samples or when dealing with nonlinearity due to heavy reliance on accurate nuisance model assumptions and the presence of sparse adherent risk sets.
  • Nine assumptions about target risk-sets are necessary, and the theoretical grounding of the new method hinges on a core theorem (Rao–Blackwellization) which shows lower variance bounds but concludes with the important acknowledgment that improvement in finite sample size treatments effects cannot be guaranteed.

Motivation and problem statement

Per-protocol estimation in sequential target trial emulation requires correcting for two sources of selection among pooled person-trials: treatment-strategy deviation and loss to follow-up. The standard remedy is cumulative inverse-probability weighting, in which stabilized or unstabilized products of adherence and censoring probabilities are multiplied across follow-up visits. These weights can become extremely variable when adherence is limited, overlap is weak, follow-up is long, or confounding is strong, and stabilization does not remove the underlying product structure. This paper proposes an alternative: estimate a marginal density ratio (MDR) that directly transports the state distribution of the observed-adherent risk set to the counterfactual risk set of individuals who would remain event-free under the target strategy. The approach, STTE–MDR, adapts marginalized importance-sampling ideas from off-policy evaluation (Liu et al., 2018, Xie et al., 2019) to the risk-set transport problem posed by sequential trials.

Framing per-protocol estimation as risk-set transport

The paper defines, for each strategy zz and follow-up interval kk, a pooled target risk-set distribution Pk,zT,ωP_{k,z}^{T,\omega} over augmented states S~=(m,S)\widetilde S = (m, S) — where mm indexes the emulated trial and SS is a prespecified summary of history — and the corresponding observed-adherent distribution Pk,zA,ωP_{k,z}^{A,\omega}. Under state-level overlap (PT≪PAP^T \ll P^A), the Radon–Nikodym derivative rk,zω=dPT/dPAr_{k,z}^\omega = dP^T/dP^A exists and serves as the weight. The estimand is the pooled per-protocol cumulative-risk difference at horizon τ\tau, with each eligible person-trial receiving equal baseline weight; this equal-person-trial pooling is an explicit design choice rather than a requirement of the framework.

Identification rests on eight assumptions beyond the standard causal conditions (no interference, consistency, sequential exchangeability for treatment and observation, positivity): state sufficiency and conditional-mean transportability (the compressed state must carry all information needed for event risk and state evolution), state-level overlap, and correct temporal/risk-set alignment. These are substantive restrictions: a state too coarse omits relevant history, while one too detailed produces weak overlap.

Estimation algorithm

The estimator proceeds in three stages. First, longitudinal g-computation generates target risk sets: fitted event and survivor-state transition models are used to forward-simulate eligible person-trials under each strategy with censoring eliminated, retaining simulated states at every interval. Second, within each kk0 cell, simulated target states and observed-adherent states are pooled into a balanced classification problem; a probabilistic classifier estimates kk1, and classifier odds yield the raw density ratio, which is then normalized to mean one within each cell. Third, a discrete-time hazard model is fitted on observed-adherent rows using MDR weights and standardized over the simulated target states, with cumulative risks obtained by the product-limit map.

The doubly robust extension, DR-STTE–MDR, adds an MDR-weighted residual correction to outcome predictions averaged over the target states, following the augmentation pattern of Bang–Robins and double reinforcement learning [bang2005doublyrobust, kallus2020drl].

Theoretical results

The paper establishes four main results. A change-of-measure proposition shows MDR weighting reproduces expectations under the target law and balances every measurable set. An identification theorem gives kk2 for each strategy–follow-up cell, so hazards, cumulative risks, and the risk difference are identified by MDR weighting. A classifier lemma justifies recovering kk3 from class-posterior odds under exact follow-up balancing.

The most structurally informative result is a Rao–Blackwellization theorem: the population MDR equals the conditional expectation of the mean-normalized full-history ratio given the current state, and consequently its population variance is no greater than that of the correctly specified cumulative inverse-probability-of-adherence-and-censoring weight. The authors are careful to note that this variance ordering concerns true population ratios and does not guarantee lower finite-sample variance for estimated treatment effects — an important qualification often elided in related literature.

Finally, conditional double robustness holds: given consistent generation of the target-state law, the DR-MDR hazard estimator is consistent if either the outcome regression is correct or the limiting density ratio is proportional to the truth. An appendix decomposition makes explicit why the estimator is not triply robust — target-generator error enters additively and is removed by neither nuisance component being correct.

Simulation evidence

Simulations compared ten estimators across seven scenarios (base, small sample, strong confounding, poor positivity, poor adherence, rare events, nonlinearity) with 500 replications each, targeting a 12-month pooled per-protocol risk difference. Key findings:

Scenario Best RMSE estimator RMSE (pp) Notable comparison
Poor positivity MDR–State 3.48 IPCW–Entry 7.96; TrialEmulation 7.85
Nonlinear DR-MDR–State 8.71 Hardest setting overall
Rare event DR-MDR–Entry lowest RMSE State versions disadvantaged
Small sample DR-IPCW–Entry lowest RMSE Only non-MDR winner

An MDR or DR-MDR variant attained the lowest RMSE in six of seven scenarios, and DR-MDR–State had the smallest absolute bias in five. Weight diagnostics support the mechanism: MDR maximum weights were tightly concentrated (medians 6.3–10.1, no replication exceeding 15.7), whereas under poor positivity IPCW and TrialEmulation median maxima were 299.0 and 201.7 with extreme replications reaching 8,710.4 and 120,863.6 respectively. Cell normalization also forced the MDR mean weight to exactly 1.000 in every replication.

These results imply that replacing cumulative probability products with direct density-ratio estimation can materially improve precision precisely where conventional weighting fails most — practical nonoverlap. However, the small-sample exception shows the method's dependence on fitting three nuisance components (target generator, classifier, outcome model); sparse adherent risk sets erode the advantage. In the nonlinear scenario, remaining bias in DR-MDR–State indicates the augmentation does not fully overcome misspecification of the parametric g-computation components.

Limitations and open questions

The paper concedes several limitations at the points where they bind. State sufficiency, transportability, and overlap are untestable modeling commitments, and like all conditional-exchangeability methods the approach does not address unmeasured confounding. Double robustness protects only against misspecification of the outcome regression or density ratio, not the target-state generator. No empirical application is included, leaving real-world performance unevaluated. Inference is unresolved: individuals contribute multiple person-trials, all nuisances are estimated without cross-fitting in the simulations, bootstrap coverage was not evaluated, and an asymptotically valid analytic variance estimator accounting for target-simulation and classifier uncertainty remains open. Extensions to continuous event times and non-absorbing outcomes are also left open.

Conclusion

This paper reframes per-protocol estimation in sequential target trial emulation as transporting observed-adherent risk sets to strategy-specific counterfactual risk sets via marginal density ratios, replacing cumulative adherence-and-censoring weight products with classifier-based ratio estimation. It provides identification theory, a Rao–Blackwell connection to full-history weighting, a conditionally doubly robust extension, and simulation evidence showing substantial RMSE and weight-stability advantages under poor positivity and other stress conditions. The method's practical value depends on plausible state sufficiency, overlap, and accurate target-state generation, and valid inference for the full pipeline remains an open problem.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.