Papers
Topics
Authors
Recent
Search
2000 character limit reached

Stepped-Wedge Design with Noncompliance

Updated 12 July 2026
  • The paper introduces a causal inference framework that distinguishes between ITT and effect ratio estimands to address noncompliance in stepped-wedge cluster-randomized trials.
  • It demonstrates that standard IV and 2SLS methods may yield misleading results, advocating for covariate-adjusted ANCOVA-CR3 and Horvitz-Thompson estimators instead.
  • Simulation studies validate that robust, randomization-aligned estimators improve accuracy and control type-I error, guiding practical recommendations for low uptake settings.

Searching arXiv for the cited stepped-wedge/noncompliance papers to ground the article. arXiv search query: (Zhang et al., 18 Sep 2025) Stepped-wedge designs with noncompliance are cluster-randomized designs in which clusters switch from control to intervention at prespecified times, but actual treatment receipt does not perfectly follow the randomized rollout. In this setting, the randomized intervention often functions as an encouragement rather than a guarantee of treatment, so standard intention-to-treat and instrumental-variable procedures developed for ordinary trial structures need not yield compatible or valid inference. Recent work formalizes this problem in a causal inference framework, defines estimands tailored to staggered rollout with imperfect uptake, and develops randomization-based estimators and inference procedures that are calibrated to the stepped-wedge assignment mechanism rather than to ordinary parallel-arm assumptions (Zhang et al., 18 Sep 2025). Related randomization-inference work in cluster-randomized test-negative designs extends the same logic to partial compliance and stepped-wedge rollout, particularly when ascertainment varies across clusters and time (Wang et al., 2022).

1. Design structure and the inferential problem

A stepped-wedge cluster-randomized trial assigns clusters to intervention sequences so that every cluster eventually transitions from control to intervention by the end of the study. In the formal setup used for noncompliance, II denotes the number of clusters and J+2J+2 the number of periods, with period $0$ as pre-rollout and period J+1J+1 as post-rollout. The randomized cluster-period assignment is ZijZ_{ij}, actual treatment received by individual kk in cluster ii, period jj is DijkD_{ijk}, the outcome is YijkY_{ijk}, and baseline covariates are J+2J+20 (Zhang et al., 18 Sep 2025).

The central complication is individual-level noncompliance: actual treatment receipt does not perfectly follow the randomized encouragement. In the palliative care trial reanalyzed by the 2025 paper, the randomized intervention was a nudge to administer palliative care, but receipt of palliative care was not guaranteed, and the compliance rate was approximately J+2J+21 (Zhang et al., 18 Sep 2025). A subsequent analysis using methods suited for standard trial designs produced statistically anomalous results: an intention-to-treat analysis found no effect, while an instrumental variable analysis did. The paper treats this as evidence that a more principled analysis is needed for stepped-wedge designs with noncompliance (Zhang et al., 18 Sep 2025).

A common misconception is that standard two-stage least squares or ordinary parametric IV analysis can be ported directly into a stepped-wedge setting. The cited work states that this is not guaranteed to yield compatible or valid results in stepped-wedge designs with noncompliance and can produce paradoxical findings (Zhang et al., 18 Sep 2025). In the related test-negative literature, analogous concerns arise because standard estimators can be biased when ascertainment varies across clusters or time, or is correlated with outcome rates (Wang et al., 2022). This suggests that staggered rollout, noncompliance, and cluster-period heterogeneity jointly alter the inferential target and the validity of conventional estimators.

2. Causal framework and identifying assumptions

The noncompliance framework for stepped-wedge trials is formulated in potential-outcomes notation. The assumptions listed in the 2025 paper are cluster-level SUTVA, intervention duration irrelevance, randomization as per stepped-wedge design, and cluster-period size being unaffected by intervention (Zhang et al., 18 Sep 2025).

Cluster-level SUTVA excludes interference between clusters. Intervention duration irrelevance states that potential outcomes depend on current randomized status, not on how long a cluster has already been in intervention. This matters because stepped-wedge analyses sometimes attempt to parameterize cumulative exposure or time-on-treatment effects; in the noncompliance framework of the paper, the primary estimands instead target the causal effect of current randomized status and the induced treatment receipt under that assumption (Zhang et al., 18 Sep 2025).

The effect-ratio interpretation additionally relies on standard IV assumptions: exclusion restriction, monotonicity, and relevance. Under these conditions, the main noncompliance estimand can be interpreted as the average causal treatment effect among compliers within the rollout periods (Zhang et al., 18 Sep 2025). The emphasis on rollout periods is substantive, because the estimand is defined over the periods in which the randomized encouragement varies.

A related randomization-inference framework in cluster-randomized test-negative designs treats all potential outcomes, and where applicable covariates and ascertainment probabilities, as fixed, randomizing only over the assignment schedule (Wang et al., 2022). That formulation is not the same estimand framework as the palliative-care analysis, but it reinforces the same methodological principle: inference should be justified by the randomization mechanism actually used in the stepped-wedge design.

3. Causal estimands under noncompliance

Two estimands are central in the stepped-wedge noncompliance framework. The first is the intention-to-treat estimand, denoted J+2J+22, defined as the sample average treatment effect of the randomized intervention:

J+2J+23

where J+2J+24 and J+2J+25 is the potential outcome under J+2J+26 (Zhang et al., 18 Sep 2025).

The second is the effect ratio, also described as the Local Average Treatment Effect, denoted J+2J+27. It is defined as the ratio of the ITT effect on the outcome to the ITT effect on treatment receipt:

J+2J+28

Under exclusion restriction, monotonicity, and relevance, J+2J+29 is the average causal treatment effect among compliers within the rollout periods (Zhang et al., 18 Sep 2025).

The distinction between $0$0 and $0$1 is not merely terminological. $0$2 captures the causal effect of the randomized intervention, which in many pragmatic stepped-wedge trials is an encouragement or implementation nudge. $0$3 rescales that effect by the induced shift in actual treatment receipt, thereby targeting the effect of treatment among individuals whose receipt is changed by the randomized rollout. This suggests that stepped-wedge designs with low uptake require explicit separation of intervention assignment from treatment receipt if the scientific question concerns the treatment itself rather than the rollout policy.

In the test-negative extension with partial compliance, the same logic appears in a dose-response form. There, the received intervention level $0$4 replaces simple binary receipt, and a dose-response model $0$5 is used, with random assignment serving as an instrument for the received dose (Wang et al., 2022). A plausible implication is that stepped-wedge noncompliance methodology naturally accommodates settings in which compliance is graded rather than binary.

4. Estimation and randomization-based inference

The key methodological innovation in the 2025 framework is a test inversion approach. Under the null hypothesis $0$6, the ITT effect of the residualized outcome $0$7 must be zero:

$0$8

This converts estimation of $0$9 into estimation of an ordinary ITT effect, but with the outcome residualized for treatment received (Zhang et al., 18 Sep 2025).

Two estimator families are developed. The first comprises ANCOVA-type estimators based on generalized linear models with residualized outcome J+1J+10 as the dependent variable. Regression includes period indicators, possible covariate adjustment, and interaction terms. The paper distinguishes three variants: unadjusted, ANCOVAI, and ANCOVAIII. ANCOVAI adds main-effect linear adjustment for covariates J+1J+11, and ANCOVAIII adds interactions between intervention and covariates. An example model is

J+1J+12

The estimated effect is the weighted average of the period-specific J+1J+13 (Zhang et al., 18 Sep 2025).

The second family comprises Horvitz-Thompson estimators using inverse probability weighting with known randomization probabilities. For J+1J+14,

J+1J+15

and the difference between the J+1J+16 and J+1J+17 estimators estimates J+1J+18. Augmented versions use regression adjustment via models fit on pre- and post-rollout periods (Zhang et al., 18 Sep 2025).

Inference proceeds by estimating J+1J+19 and its standard error for a given ZijZ_{ij}0, conducting a Wald-type test of ZijZ_{ij}1, solving ZijZ_{ij}2 for the point estimate ZijZ_{ij}3, and inverting the test to obtain a confidence interval. For small numbers of clusters and one-cluster-at-a-time rollouts, the paper recommends CR3 jackknife variance estimation with comparison to a ZijZ_{ij}4-distribution quantile rather than to a normal reference. For Horvitz-Thompson estimators, conservative design-based variance estimators are used (Zhang et al., 18 Sep 2025).

An important property of the proposed procedures is estimation/testing alignment: a null hypothesis test for the effect ratio at ZijZ_{ij}5 rejects if and only if the test for the ITT effect at ZijZ_{ij}6 rejects (Zhang et al., 18 Sep 2025). This directly addresses the paradox in which an IV-style analysis appears significant although the corresponding ITT analysis does not.

5. Simulation evidence and methodological guidance

The simulation study in the 2025 paper evaluates informative and non-informative cluster sizes, varying numbers of clusters and periods, varying compliance rates, and one-cluster-at-a-time rollouts, which the paper identifies as the most challenging setting (Zhang et al., 18 Sep 2025). It compares unadjusted ANCOVA, ANCOVAI, ANCOVAIII, Horvitz-Thompson estimators with and without regression adjustment, and standard 2SLS-based methods implemented in the ivmodel R package, including Fuller and CLR variants. The reported evaluation criteria are mean squared error, type-I error under the null, and power under the alternative (Zhang et al., 18 Sep 2025).

The main findings are specific. Covariate adjustment, including ANCOVAI, ANCOVAIII, and regression-adjusted HT, greatly improves estimator accuracy and power. ANCOVA with CR3 variance and ZijZ_{ij}7-distribution reliably controls type-I error, even in one-cluster-at-a-time and small-cluster scenarios. Standard model-based IV approaches, particularly ivmodel-CLR, can show substantial type-I error inflation in one-cluster-at-a-time and small-sample settings. Horvitz-Thompson estimators with conservative variance estimates are valid but often yield poor power, especially without covariate adjustment. In larger-cluster settings with more clusters, more estimators become reliable, but in realistic settings with few clusters and heterogeneous cluster size, ANCOVA-CR3 methods have clear practical advantages (Zhang et al., 18 Sep 2025).

The practical guidance is correspondingly explicit. The recommended primary analysis is an adjusted ANCOVA method, either ANCOVAI or ANCOVAIII, with CR3 variance estimation and a ZijZ_{ij}8-distribution reference. If design-based variance is theoretically imperative, especially in one-cluster-at-a-time or very small cluster settings, the recommended alternative is regression-adjusted Horvitz-Thompson estimation using pre/post data and conservative variance estimation, with the recognition that power will be lower. The paper advises against relying on standard IV or 2SLS methods such as CLR for primary inference in stepped-wedge designs with noncompliance because they can be anti-conservative (Zhang et al., 18 Sep 2025).

The test-negative stepped-wedge work reaches a structurally similar conclusion in a different outcome model: only the log-contrast and covariate-adjusted log-contrast estimators consistently showed zero bias and correct coverage regardless of variable ascertainment or stepped-wedge design, while standard odds ratio, GLMM, and GEE approaches were biased under heterogeneous ascertainment (Wang et al., 2022). This suggests a broader pattern across stepped-wedge settings: estimators that explicitly encode the design randomization and the relevant nuisance structure tend to dominate methods imported from conventional regression templates.

The principal empirical application in the 2025 paper is the reanalysis of the Courtright et al. palliative care pragmatic trial, REDAPS. The design involved 11 hospitals, 10 rollout periods, and approximately 24,000 patients. The treatment was cluster-level randomization to a nudge for default palliative care, and the noncompliance rate was approximately ZijZ_{ij}9, meaning only about kk0 of patients received palliative care. The outcome was hospital length of stay, where lower is better (Zhang et al., 18 Sep 2025).

The previously reported analyses were discordant. The ITT analysis found no significant effect, with estimate kk1 and kk2 confidence interval kk3. The 2SLS/IV analysis reported a statistically significant reduction in length of stay, kk4 with kk5 confidence interval kk6 (Zhang et al., 18 Sep 2025). The reanalysis using the recommended ANCOVA-CR3 methods produced point estimates for the average effect of palliative care among compliers that were more negative, approximately kk7, but with kk8 confidence intervals including zero: ANCOVAI gave kk9 with interval ii0, and ANCOVAIII was similar. For comparison, ivmodel CLR produced ii1 with interval ii2, but the paper argues that this apparent significance should not be trusted because simulation showed anti-conservatism of the method in this setting (Zhang et al., 18 Sep 2025). The analysis also found no evidence against the intervention duration irrelevance assumption.

A second extension appears in the cluster-randomized test-negative design literature. There, for cluster ii3 and time ii4, the observed quantities ii5 and ii6 define the log-contrast ii7, and under an immediate intervention effect the model assumes ii8. The stepped-wedge estimator averages intervention-versus-control log-contrasts across time with weights ii9, and the variance has a closed-form expression that can be unbiasedly estimated from the observed data. For partial compliance, the framework uses a dose-response model and randomization-based IV estimation for the received dose (Wang et al., 2022). This is a distinct outcome setting, but it shows that stepped-wedge noncompliance methods can be embedded in designs where ascertainment rather than treatment uptake is the dominant nuisance process.

Noncompliance should also be distinguished from dropout. In closed-cohort stepped-wedge cluster-randomized trials, non-ignorable dropout due to death or other causes is a separate post-randomization complication. Joint longitudinal-survival models address that problem by jointly modeling the longitudinal outcome and the dropout process through a time-to-event submodel, often with shared random effects (Gasparini et al., 2024). The simulation results in that literature indicate that mixed-effects models are unbiased if dropout is non-informative but biased under informative dropout, whereas joint models remained unbiased or nearly unbiased across the simulated scenarios (Gasparini et al., 2024). A plausible implication is that analysis of stepped-wedge trials may need distinct methodological layers for noncompliance, ascertainment, and informative dropout rather than a single omnibus correction.

Overall, the current methodological picture treats stepped-wedge design with noncompliance as a problem of causal estimand definition and randomization-aligned inference. The principal contribution of the 2025 framework is to make the noncompliance estimand explicit, provide estimators whose validity is calibrated to the stepped-wedge assignment process, and eliminate the incoherence in which a noncompliance-adjusted analysis appears significant while the corresponding ITT analysis does not (Zhang et al., 18 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Stepped-Wedge Design with Noncompliance.