Index Date Imputation for Externally Controlled Trials
- Index Date Imputation (IDI) is a statistical method that imputes treatment initiation times to align survival comparisons in externally controlled trials.
- It employs left-truncated Kaplan-Meier estimators and Cox models to estimate the initiation time distribution, thereby correcting immortal-time bias.
- IDI integrates propensity score adjustment methods to balance covariates and ensure comparable risk sets for valid survival analysis.
Searching arXiv for the specified paper and closely related work on externally controlled trials, survival analysis, and index-date alignment. Index Date Imputation (IDI) is a statistical method for survival analysis in externally controlled trials (ECTs), where outcomes from a single-arm trial are compared with external controls drawn from historical trials, registries, or observational studies. Its central purpose is to address index-date misalignment: in the treated cohort, time-to-event is often measured from treatment initiation, whereas external controls may only have diagnosis as a natural time origin. In that setting, direct comparison can induce immortal-time bias and distort causal interpretation. IDI imputes comparable pseudo-initiation times for external controls from the estimated distribution of treatment initiation times in the single-arm cohort, then imposes the same truncation condition used implicitly in the treated group. Combined with propensity score weighting or matching, the method is designed to produce balanced and temporally aligned cohorts for survival comparison (Coent et al., 17 Sep 2025).
1. Problem formulation in externally controlled survival studies
In the formulation given for ECTs, a single-arm treatment cohort is denoted by and an external control cohort by . The paper defines as time from diagnosis to event, as time from diagnosis to treatment initiation, as administrative censoring time, , , as treatment indicator, and as baseline covariates measured at diagnosis (Coent et al., 17 Sep 2025).
The difficulty arises because is observed only when 0. In the single-arm trial, outcomes are measured from treatment initiation, 1. Because only patients who survive to treatment initiation enter the risk set, the treated cohort is left-truncated at 2. External controls, by contrast, have no natural 3; their survival is measured from diagnosis. A diagnosis-based analysis for controls and an initiation-based analysis for treated patients therefore compares noncomparable risk sets. The paper states that this direct comparison induces immortal-time bias.
IDI is introduced as a correction for that misalignment. It seeks to impute a pseudo-initiation time 4 for each external control patient from the estimated distribution of 5 in the treated population and to impose the same truncation condition 6. The resulting alignment creates comparable “time zero” in both groups before standard survival comparisons are applied. This suggests that IDI is not merely a missing-data device; it is an alignment procedure for the survival time origin in settings where trial and external-control cohorts enter the risk set under different temporal rules.
2. Estimation of the initiation-time distribution and the IDI algorithm
The statistical framework distinguishes the observed treatment-initiation times 7 in the single-arm cohort from the unobserved 8 in the external controls. The key object is the distribution of treatment initiation times in the target untruncated treated population,
9
Because the observed 0 are subject to left truncation, the paper estimates 1 by a weighted empirical cumulative distribution function,
2
where 3 is the left-truncated Kaplan-Meier estimate of survival from diagnosis in the treated group (Coent et al., 17 Sep 2025).
An equivalent route is to fit a Cox model
4
and derive
5
The imputation algorithm is defined subject-by-subject for each external control 6. First, draw 7. Second, compute the pseudo-survival time
8
Third, retain the subject if 9, equivalently if 0. These steps are repeated across 1 bootstrap replicates to average over imputation randomness.
This construction has two linked effects. It gives controls an imputed index date and simultaneously enforces the same survival-to-entry condition present in the treated cohort. A plausible implication is that the principal inferential gain from IDI comes from restoring comparability of the at-risk sets rather than from modeling survival itself.
3. Integration with propensity score adjustment and aligned hazard estimation
The method is explicitly combined with propensity score methods to address population-level confounding. The target estimand is the Average Treatment effect on the Treated (ATT). The propensity score is defined as
2
Because 3 subjects are left-truncated, the treated cohort is reweighted by the inverse of its truncation probability,
4
with 5 for controls. The logistic model for 6 is then fit by weighted maximum likelihood to obtain 7 (Coent et al., 17 Sep 2025).
For ATT weighting, the paper defines
8
As an alternative, it permits 1:1 nearest-neighbor matching on 9, with each matched subject receiving unit weight in subsequent analyses.
After imputation and weighting, survival times are redefined from the observed or imputed index date: 0 The aligned comparison is then performed with a weighted Cox model,
1
using weights 2. The resulting 3 targets the log hazard ratio
4
for marginal survival functions
5
The framework therefore combines temporal alignment and confounding adjustment in one analysis pipeline. This suggests that IDI, as formulated, is best understood as a survival-comparison procedure for ATT estimation under left truncation and index-date noncomparability.
4. Diagnostics, simulation design, and operating characteristics
The paper specifies two diagnostic families. The first concerns covariate balance. For each covariate 6, the standardized mean difference is
7
Under weighting, weighted means and variances are used, and SMDs are graphed before and after propensity score adjustment. A heuristic threshold is 8 (Coent et al., 17 Sep 2025).
The second diagnostic concerns truncation bias. The observed truncated distribution of 9 in the treated group is compared with the estimated truncated CDF
0
where
1
and
2
A Q–Q plot of the empirical distribution against 3 is used as a goodness-of-fit check.
The simulation study evaluates the method under two sample-size settings: 4 treated with 5 controls, and 6 with 7. Covariates are generated as 8 and 9, with 0 and 1. The initiation time follows either 2 or 3. Event times are generated from a piecewise-exponential hazard,
4
and
5
with censoring defined by 6, 7 (Coent et al., 17 Sep 2025).
Three methods are compared: naïve Cox from diagnosis with no adjustment; propensity score matching plus IDI; and propensity score weighting plus IDI. Performance metrics are mean estimated log-HR versus true, absolute bias, empirical SD and average model-based SE, and empirical coverage of the 95% CI.
The reported findings are specific. Naïve analysis exhibits large bias, approximately 8 to 9, and zero coverage. Both IDI approaches reduce bias to approximately 0 to 1 and achieve approximately 2 to 3 coverage. Propensity score weighting generally yields smaller variance, while matching performs comparably. Under modest unmeasured confounding (4), bias remains small; under strong confounding (5), bias increases and coverage declines. Taken together, these results support the paper’s claim that the primary failure mode is index-date misalignment itself, and that IDI can correct it when its assumptions are plausible.
5. ECOG-2108 emulation and empirical recovery of a known treatment effect
The real-world application uses ECOG-2108, described as a randomized trial of early local therapy (ELT, Arm B) versus continued systemic therapy (OST, Arm A) in de novo metastatic breast cancer. In the original randomized trial, survival is measured from randomization, which is the index date (Coent et al., 17 Sep 2025).
To emulate an ECT, Arm B is treated as the single-arm trial (6) and Arm A as external controls (7). The observed 8 in Arm B is the time from enrollment to randomization, while 9 is unobserved in Arm A. Baseline covariates are age, race, and disease stage, although the paper notes that randomization ensures balance.
Implementation follows the general IDI workflow. The paper estimates 0 from Arm B with adjustment for left truncation, fits a Cox model 1 in Arm B to compute truncation weights 2, fits the weighted logistic model for 3, draws 4 for each Arm A patient and retains subjects if 5, then fits a weighted Cox model on aligned times 6 for Arm B and 7 for Arm A. Inference uses bootstrap with 8 replicates.
The reported hazard ratios are as follows:
| Analysis | HR | 95% CI |
|---|---|---|
| Original RCT | 0.863 | 0.598–1.246 |
| Naïve ECT (misaligned index) | 0.940 | 0.651–1.358 |
| IDI without PS weighting | 0.862 | 0.597–1.245 |
| IDI with PS weighting | 0.882 | 0.610–1.277 |
The empirical conclusion in the paper is that IDI recovers the known treatment effect despite artificial index-date misalignment. In this application, the naïve ECT estimate is attenuated relative to the randomized comparison, whereas the IDI-based analyses are close to the original trial estimate. This suggests that the method’s intended use case is not limited to highly imbalanced observational controls; it also includes settings where a randomized benchmark can be used to test whether temporal realignment restores the appropriate estimand.
6. Workflow, assumptions, limitations, and scope of use
The proposed implementation workflow is explicit. One first prepares the data by identifying 9, 0, 1, 2, and 3, with 4 observed only if 5. One then estimates 6 using a left-truncated Kaplan-Meier estimator or a Cox model on the treated group; estimates truncation weights 7 for 8 via 9; fits the weighted logistic model for 00; obtains ATT weights 01; and, for 02 bootstrap replicates, resamples all subjects with replacement, imputes 03 for controls, applies truncation, fits the weighted Cox model on aligned times, and records 04. The bootstrap distribution 05 is then summarized to obtain the point estimate and confidence interval (Coent et al., 17 Sep 2025).
The paper also names software options in R: survival (coxph), ipw or WeightIt for weighting, MatchIt for matching, and boot for resampling. It recommends ensuring reproducibility by setting RNG seeds or using the same 06 across replicates.
The assumptions and limitations are stated directly. They include no unmeasured confounding, positivity, conditional independence of 07 and 08 (Assumption A2), and censoring that satisfies 09 or is independent of 10. The paper further notes that imputation randomness is mitigated by bootstrap but may reduce reproducibility, and that subjects with 11 are excluded, potentially reducing effective sample size and power.
The practical recommendations are similarly specific: check and achieve good covariate balance with 12 after propensity score adjustment; examine Q–Q plots of truncated versus estimated 13 distributions; conduct sensitivity analyses for unmeasured confounding, for example by varying 14 in latent-variable simulations; report the number or exclusion rate of controls truncated by 15; consider alternative estimands such as ATE versus ATT carefully, since IDI as described targets ATT; and explore population-level approaches, such as convolution of survival curves with 16, if sample-size loss from subject-level imputation is prohibitive.
Within those constraints, the paper presents IDI as a principled framework for time-to-event analyses in ECTs and states that it is broadly applicable in oncology and rare disease settings. A plausible implication is that its utility is greatest in studies where the treated cohort enters observation only after a clinically meaningful delay, while the external controls lack a directly observed analogue of that entry time.