- The paper demonstrates that fixed-effects models yield consistent and robust estimators for marginal causal effects in clustered longitudinal trials.
- It establishes that FEMs automatically adjust for time-invariant confounding at both cluster and individual levels, even under model misspecification.
- Simulation studies and real-world re-analyses reveal FEMs outperform mixed-effects models, particularly in small-sample and quasi-experimental settings.
Fixed-Effects Models for Causal Inference in Longitudinal Cluster Randomized and Quasi-Experimental Trials
Overview and Motivation
The paper "Fixed-Effects Models for Causal Inference in Longitudinal Cluster Randomized and Quasi-Experimental Trials" (2604.07756) systematically reevaluates the inferential properties and robustness of fixed-effects models (FEMs) when applied to longitudinal cluster trial designs, including stepped-wedge (SW-CRT), parallel-with-baseline (PB-CRT), and crossover (CRXO) trials, in both randomized and quasi-experimental contexts. Historically, biostatistics has favored mixed-effects models (MEMs) due to their perceived targeting of superpopulation estimands. This paper challenges such interpretations, proving that FEMs—when correctly specified—yield consistent and model-robust estimators for a spectrum of marginal, population-level causal estimands. The work further delineates the conditions under which FEMs remain robust with arbitrary model misspecification, particularly highlighting their automatic control for both cluster- and individual-level time-invariant confounding, a property that is critical in quasi-experimental settings.
Cluster trial designs considered here exhibit within-cluster and between-cluster heterogeneity in treatment assignment over multiple discrete periods. Using potential outcomes notation, a range of estimands are formally nonparametrically defined:
- Constant treatment effect: Δ is presumed invariant across calendar period and exposure duration.
- Duration-specific effect: Δ(d) targets effects after d periods of exposure.
- Period-specific effect: Δj​ captures the effect at calendar time j.
- Saturated structure: Allows maximal time-varying heterogeneity Δj​(d).
A particularly relevant estimand for staggered designs is the period-average treatment effect for the overlap population (P-ATO), weighted according to periods with best treatment-control overlap. The generality of these definitions enables clear alignment between estimand and inference, regardless of model choice.
Asymptotic Theory and Model Robustness
The paper rigorously demonstrates via M-estimation frameworks that FEM estimators targeting treatment effect parameters (e.g., via ordinary least squares or GEE) are consistent and asymptotically normal for marginal cluster-average estimands, provided the treatment effect structure is correctly specified. All other aspects—covariate relationships, random effect distributions, outcome variance structure—can be arbitrarily misspecified.
A salient technical contribution is the formal resolution of the incidental parameters problem for FEMs. Cluster fixed intercepts are handled as nuisance parameters, either removed via the within-cluster transformation for linear models or via conditioning in the log-link setting (conditional Poisson), preserving regularity conditions for asymptotics.
When operating under quasi-experimental designs, the analysis requires a "mean independence" assumption (generalized parallel trends for untreated outcomes and exchangeable treatment effects). The result is model-robust inference even in the absence of randomization, allowing for time-invariant confounding aggregated at the cluster level.

Figure 1: Small-sample simulations in SW-CRT demonstrate low bias and nominal coverage for FEMs, contrasting with undercoverage or inflation in MEMs when m≤6 clusters.
Simulation Studies
Simulation experiments span small (m=6) and large (m=100) cluster counts, binary and continuous outcome types, and varied effect structures (period, duration, saturated). Key findings:
- FEMs deliver unbiased estimators for marginal estimands, even with misspecified outcome or covariate models, outperforming MEMs in small-sample settings in coverage and Type I error.
- Jackknife variance estimation (with tm−2​ df) is robust in small samples; sandwich estimators can underestimate variance.
- In PB-CRT and CRXO designs (non-staggered), the FEM constant treatment estimator targets the time-averaged treatment effect estimand (P-ATO) without need for correct effect structure specification.

Figure 2: PB-CQT simulation with time-varying effect; FEM is unbiased and efficient, while MEM exhibits bias due to confounding and effect heterogeneity.

Figure 3: CRXO scenario confirms FEM consistency for saturated estimands, even under complex covariate and correlation structures.
Re-analysis of Real-World Cluster Trial
The application to a SW-CRT crowd-sourced HIV testing intervention (closed-cohort, Δ(d)0) re-analyzed with FEM, log-link FEM+g-computation, and MEM approaches. All models yielded comparable effect estimates and confidence intervals. The FEMs were robust to missingness (addressed via MI), and targeted estimands aligned well with original trial findings.

Figure 4: Re-analysis of real-world SW-CRT using various treatment effect structures, indicating FEMs and MEMs deliver similar inferential results.
Practical and Theoretical Implications
- Superpopulation inference: FEMs are not restricted to finite-sample inference; when implemented in M-estimation, they consistently target marginal estimands in superpopulations.
- Automatic adjustment: FEMs inherently control for all time-invariant confounding, both cluster- and individual-level, without requiring explicit covariate specification.
- Design dependency: In PB-CRT and CRXO, FEMs estimate time-averaged effects without need for correct treatment structure; in SW-CRT, positivity weights (P-ATO) ensure estimand alignment.
- Variance estimation: Jackknife-t based intervals are preferable for small Δ(d)1, given sandwich estimator limitations.
- Model selection: FEMs may be preferable over MEMs in limited-cluster or quasi-experimental designs, and when robustness to misspecification is desired.

Figure 5: Large-sample (Δ(d)2) simulation showing robust bias and coverage properties of FEMs for SW-CRT binary outcomes.
Future Directions
Areas for further development include sensitivity analysis for time-varying confounding, relaxation of non-informative enrollment assumptions, and extension to broader generalized linear and non-linear fixed-effects frameworks. The role of FEMs in settings with informative cluster-period sizes and adaptive designs also warrants exploration.
Conclusion
This paper establishes fixed-effects models as a robust tool for causal inference in longitudinal cluster trials. FEM estimators, under minimal model assumptions and correct treatment structure specification, consistently and efficiently target marginal, population-level estimands, greatly enhancing the interpretability and reliability of cluster trial analyses, particularly in small-sample, staggered, and quasi-experimental settings.

Figure 6: PB-CQT with Δ(d)3 clusters, continuous outcome—FEMs remain unbiased and efficient, supporting generalizability to large-scale trials.