---
title: Stepped-Wedge Design with Noncompliance
url: https://www.emergentmind.com/topics/stepped-wedge-design-with-noncompliance
type: topic
---

# Stepped-Wedge Design with Noncompliance

Searching arXiv for the cited stepped-wedge/noncompliance papers to ground the article.
arXiv search query: 2509.14598
Stepped-wedge designs with noncompliance are cluster-randomized designs in which clusters switch from control to intervention at prespecified times, but actual treatment receipt does not perfectly follow the randomized rollout. In this setting, the randomized intervention often functions as an encouragement rather than a guarantee of treatment, so standard intention-to-treat and instrumental-variable procedures developed for ordinary trial structures need not yield compatible or valid inference. Recent work formalizes this problem in a causal inference framework, defines estimands tailored to staggered rollout with imperfect uptake, and develops randomization-based estimators and inference procedures that are calibrated to the stepped-wedge assignment mechanism rather than to ordinary parallel-arm assumptions [2509.14598]. Related randomization-inference work in cluster-randomized test-negative designs extends the same logic to partial compliance and stepped-wedge rollout, particularly when ascertainment varies across clusters and time [2202.03379].

## 1. Design structure and the inferential problem

A stepped-wedge cluster-randomized trial assigns clusters to intervention sequences so that every cluster eventually transitions from control to intervention by the end of the study. In the formal setup used for noncompliance, \(I\) denotes the number of clusters and \(J+2\) the number of periods, with period \(0\) as pre-rollout and period \(J+1\) as post-rollout. The randomized cluster-period assignment is \(Z_{ij}\), actual treatment received by individual \(k\) in cluster \(i\), period \(j\) is \(D_{ijk}\), the outcome is \(Y_{ijk}\), and baseline covariates are \(X_{ijk}\) [2509.14598].

The central complication is individual-level noncompliance: actual treatment receipt does not perfectly follow the randomized encouragement. In the palliative care trial reanalyzed by the 2025 paper, the randomized intervention was a nudge to administer palliative care, but receipt of palliative care was not guaranteed, and the compliance rate was approximately \(30\%\) [2509.14598]. A subsequent analysis using methods suited for standard trial designs produced statistically anomalous results: an intention-to-treat analysis found no effect, while an instrumental variable analysis did. The paper treats this as evidence that a more principled analysis is needed for stepped-wedge designs with noncompliance [2509.14598].

A common misconception is that standard two-stage least squares or ordinary parametric IV analysis can be ported directly into a stepped-wedge setting. The cited work states that this is not guaranteed to yield compatible or valid results in stepped-wedge designs with noncompliance and can produce paradoxical findings [2509.14598]. In the related test-negative literature, analogous concerns arise because standard estimators can be biased when ascertainment varies across clusters or time, or is correlated with outcome rates [2202.03379]. This suggests that staggered rollout, noncompliance, and cluster-period heterogeneity jointly alter the inferential target and the validity of conventional estimators.

## 2. Causal framework and identifying assumptions

The noncompliance framework for stepped-wedge trials is formulated in potential-outcomes notation. The assumptions listed in the 2025 paper are cluster-level SUTVA, intervention duration irrelevance, randomization as per stepped-wedge design, and cluster-period size being unaffected by intervention [2509.14598].

Cluster-level SUTVA excludes interference between clusters. Intervention duration irrelevance states that potential outcomes depend on current randomized status, not on how long a cluster has already been in intervention. This matters because stepped-wedge analyses sometimes attempt to parameterize cumulative exposure or time-on-treatment effects; in the noncompliance framework of the paper, the primary estimands instead target the causal effect of current randomized status and the induced treatment receipt under that assumption [2509.14598].

The effect-ratio interpretation additionally relies on standard IV assumptions: exclusion restriction, monotonicity, and relevance. Under these conditions, the main noncompliance estimand can be interpreted as the average causal treatment effect among compliers within the rollout periods [2509.14598]. The emphasis on rollout periods is substantive, because the estimand is defined over the periods in which the randomized encouragement varies.

A related randomization-inference framework in cluster-randomized test-negative designs treats all potential outcomes, and where applicable covariates and ascertainment probabilities, as fixed, randomizing only over the assignment schedule [2202.03379]. That formulation is not the same estimand framework as the palliative-care analysis, but it reinforces the same methodological principle: inference should be justified by the randomization mechanism actually used in the stepped-wedge design.

## 3. Causal estimands under noncompliance

Two estimands are central in the stepped-wedge noncompliance framework. The first is the intention-to-treat estimand, denoted \(\tau\), defined as the sample average treatment effect of the randomized intervention:

$$
\tau = \frac{1}{N}\sum_{i=1}^I\sum_{j=1}^J\sum_{k=1}^{N_{ij}} \{Y_{ijk}(1)-Y_{ijk}(0)\},
$$

where \(N=\sum_{i=1}^I\sum_{j=1}^J N_{ij}\) and \(Y_{ijk}(z)\) is the potential outcome under \(Z_{ij}=z\) [2509.14598].

The second is the effect ratio, also described as the Local Average Treatment Effect, denoted \(\lambda\). It is defined as the ratio of the ITT effect on the outcome to the ITT effect on treatment receipt:

$$
\lambda =
\frac{
\frac{1}{N}\sum_{i=1}^I\sum_{j=1}^J\sum_{k=1}^{N_{ij}} \{Y_{ijk}(1)-Y_{ijk}(0)\}
}{
\frac{1}{N}\sum_{i=1}^I\sum_{j=1}^J\sum_{k=1}^{N_{ij}} \{D_{ijk}(1)-D_{ijk}(0)\}
}.
$$

Under exclusion restriction, monotonicity, and relevance, \(\lambda\) is the average causal treatment effect among compliers within the rollout periods [2509.14598].

The distinction between \(\tau\) and \(\lambda\) is not merely terminological. \(\tau\) captures the causal effect of the randomized intervention, which in many pragmatic stepped-wedge trials is an encouragement or implementation nudge. \(\lambda\) rescales that effect by the induced shift in actual treatment receipt, thereby targeting the effect of treatment among individuals whose receipt is changed by the randomized rollout. This suggests that stepped-wedge designs with low uptake require explicit separation of intervention assignment from treatment receipt if the scientific question concerns the treatment itself rather than the rollout policy.

In the test-negative extension with partial compliance, the same logic appears in a dose-response form. There, the received intervention level \(D\) replaces simple binary receipt, and a dose-response model \(L_i(d)=L_i(0)+\beta d\) is used, with random assignment serving as an instrument for the received dose [2202.03379]. A plausible implication is that stepped-wedge noncompliance methodology naturally accommodates settings in which compliance is graded rather than binary.

## 4. Estimation and randomization-based inference

The key methodological innovation in the 2025 framework is a test inversion approach. Under the null hypothesis \(H_0:\lambda=\lambda_0\), the ITT effect of the residualized outcome \(Y-\lambda_0 D\) must be zero:

$$
\tau_{\lambda_0}
=
\frac{1}{N}\sum_{i=1}^I\sum_{j=1}^J\sum_{k=1}^{N_{ij}}
\left[
\bigl(Y_{ijk}(1)-\lambda_0 D_{ijk}(1)\bigr)
-
\bigl(Y_{ijk}(0)-\lambda_0 D_{ijk}(0)\bigr)
\right]
=0.
$$

This converts estimation of \(\lambda\) into estimation of an ordinary ITT effect, but with the outcome residualized for treatment received [2509.14598].

Two estimator families are developed. The first comprises ANCOVA-type estimators based on generalized linear models with residualized outcome \(Y_{ijk}-\lambda_0 D_{ijk}\) as the dependent variable. Regression includes period indicators, possible covariate adjustment, and interaction terms. The paper distinguishes three variants: unadjusted, ANCOVAI, and ANCOVAIII. ANCOVAI adds main-effect linear adjustment for covariates \(X_{ijk}\), and ANCOVAIII adds interactions between intervention and covariates. An example model is

$$
Y_{ijk}-\lambda_0 D_{ijk} \sim \beta_j + \theta_j Z_{ij} + X_{ijk}\eta .
$$

The estimated effect is the weighted average of the period-specific \(\widehat{\theta}_j(\lambda_0)\) [2509.14598].

The second family comprises Horvitz-Thompson estimators using inverse probability weighting with known randomization probabilities. For \(a\in\{0,1\}\),

$$
\widehat{\tau}_{a,\lambda_0,\mathrm{HT}}
=
\frac{1}{N}\sum_{i,j,k}
\frac{\mathbb{I}(Z_{ij}=a)}{P(Z_{ij}=a)}
\left(Y_{ijk}-\lambda_0 D_{ijk}\right),
$$

and the difference between the \(a=1\) and \(a=0\) estimators estimates \(\tau_{\lambda_0}\). Augmented versions use regression adjustment via models fit on pre- and post-rollout periods [2509.14598].

Inference proceeds by estimating \(\widehat{\tau}_{\lambda_0}\) and its standard error for a given \(\lambda_0\), conducting a Wald-type test of \(H_0:\lambda=\lambda_0\), solving \(\widehat{\tau}_\lambda=0\) for the point estimate \(\widehat{\lambda}\), and inverting the test to obtain a confidence interval. For small numbers of clusters and one-cluster-at-a-time rollouts, the paper recommends CR3 jackknife variance estimation with comparison to a \(t\)-distribution quantile rather than to a normal reference. For Horvitz-Thompson estimators, conservative design-based variance estimators are used [2509.14598].

An important property of the proposed procedures is estimation/testing alignment: a null hypothesis test for the effect ratio at \(\lambda=0\) rejects if and only if the test for the ITT effect at \(\tau=0\) rejects [2509.14598]. This directly addresses the paradox in which an IV-style analysis appears significant although the corresponding ITT analysis does not.

## 5. Simulation evidence and methodological guidance

The simulation study in the 2025 paper evaluates informative and non-informative cluster sizes, varying numbers of clusters and periods, varying compliance rates, and one-cluster-at-a-time rollouts, which the paper identifies as the most challenging setting [2509.14598]. It compares unadjusted ANCOVA, ANCOVAI, ANCOVAIII, Horvitz-Thompson estimators with and without regression adjustment, and standard 2SLS-based methods implemented in the `ivmodel` R package, including Fuller and CLR variants. The reported evaluation criteria are mean squared error, type-I error under the null, and power under the alternative [2509.14598].

The main findings are specific. Covariate adjustment, including ANCOVAI, ANCOVAIII, and regression-adjusted HT, greatly improves estimator accuracy and power. ANCOVA with CR3 variance and \(t\)-distribution reliably controls type-I error, even in one-cluster-at-a-time and small-cluster scenarios. Standard model-based IV approaches, particularly `ivmodel`-CLR, can show substantial type-I error inflation in one-cluster-at-a-time and small-sample settings. Horvitz-Thompson estimators with conservative variance estimates are valid but often yield poor power, especially without covariate adjustment. In larger-cluster settings with more clusters, more estimators become reliable, but in realistic settings with few clusters and heterogeneous cluster size, ANCOVA-CR3 methods have clear practical advantages [2509.14598].

The practical guidance is correspondingly explicit. The recommended primary analysis is an adjusted ANCOVA method, either ANCOVAI or ANCOVAIII, with CR3 variance estimation and a \(t\)-distribution reference. If design-based variance is theoretically imperative, especially in one-cluster-at-a-time or very small cluster settings, the recommended alternative is regression-adjusted Horvitz-Thompson estimation using pre/post data and conservative variance estimation, with the recognition that power will be lower. The paper advises against relying on standard IV or 2SLS methods such as CLR for primary inference in stepped-wedge designs with noncompliance because they can be anti-conservative [2509.14598].

The test-negative stepped-wedge work reaches a structurally similar conclusion in a different outcome model: only the log-contrast and covariate-adjusted log-contrast estimators consistently showed zero bias and correct coverage regardless of variable ascertainment or stepped-wedge design, while standard odds ratio, GLMM, and GEE approaches were biased under heterogeneous ascertainment [2202.03379]. This suggests a broader pattern across stepped-wedge settings: estimators that explicitly encode the design randomization and the relevant nuisance structure tend to dominate methods imported from conventional regression templates.

## 6. Applications, extensions, and related complications

The principal empirical application in the 2025 paper is the reanalysis of the Courtright et al. palliative care pragmatic trial, REDAPS. The design involved 11 hospitals, 10 rollout periods, and approximately 24,000 patients. The treatment was cluster-level randomization to a nudge for default palliative care, and the noncompliance rate was approximately \(70\%\), meaning only about \(30\%\) of patients received palliative care. The outcome was hospital length of stay, where lower is better [2509.14598].

The previously reported analyses were discordant. The ITT analysis found no significant effect, with estimate \(-0.53\%\) and \(95\%\) confidence interval \([-3.51\%, 2.53\%]\). The 2SLS/IV analysis reported a statistically significant reduction in length of stay, \(-9.6\%\) with \(95\%\) confidence interval \([-17.5\%,-1.6\%]\) [2509.14598]. The reanalysis using the recommended ANCOVA-CR3 methods produced point estimates for the average effect of palliative care among compliers that were more negative, approximately \(-0.18\), but with \(95\%\) confidence intervals including zero: ANCOVAI gave \(-0.18\) with interval \([-0.60, 0.07]\), and ANCOVAIII was similar. For comparison, `ivmodel` CLR produced \(-0.11\) with interval \([-0.19,-0.03]\), but the paper argues that this apparent significance should not be trusted because simulation showed anti-conservatism of the method in this setting [2509.14598]. The analysis also found no evidence against the intervention duration irrelevance assumption.

A second extension appears in the cluster-randomized test-negative design literature. There, for cluster \(i\) and time \(t\), the observed quantities \(O_{it}^Y\) and \(O_{it}^Z\) define the log-contrast \(L_{it}=\log(O_{it}^Y)-\log(O_{it}^Z)\), and under an immediate intervention effect the model assumes \(L_{it}(1)=\log(\lambda)+L_{it}(0)\). The stepped-wedge estimator averages intervention-versus-control log-contrasts across time with weights \(w_t\), and the variance has a closed-form expression that can be unbiasedly estimated from the observed data. For partial compliance, the framework uses a dose-response model and randomization-based IV estimation for the received dose [2202.03379]. This is a distinct outcome setting, but it shows that stepped-wedge noncompliance methods can be embedded in designs where ascertainment rather than treatment uptake is the dominant nuisance process.

Noncompliance should also be distinguished from dropout. In closed-cohort stepped-wedge cluster-randomized trials, non-ignorable dropout due to death or other causes is a separate post-randomization complication. Joint longitudinal-survival models address that problem by jointly modeling the longitudinal outcome and the dropout process through a time-to-event submodel, often with shared random effects [2404.14840]. The simulation results in that literature indicate that mixed-effects models are unbiased if dropout is non-informative but biased under informative dropout, whereas joint models remained unbiased or nearly unbiased across the simulated scenarios [2404.14840]. A plausible implication is that analysis of stepped-wedge trials may need distinct methodological layers for noncompliance, ascertainment, and informative dropout rather than a single omnibus correction.

Overall, the current methodological picture treats stepped-wedge design with noncompliance as a problem of causal estimand definition and randomization-aligned inference. The principal contribution of the 2025 framework is to make the noncompliance estimand explicit, provide estimators whose validity is calibrated to the stepped-wedge assignment process, and eliminate the incoherence in which a noncompliance-adjusted analysis appears significant while the corresponding ITT analysis does not [2509.14598].

Source: https://www.emergentmind.com/topics/stepped-wedge-design-with-noncompliance