Papers
Topics
Authors
Recent
Search
2000 character limit reached

Beyond Parallel Trends in Staggered Difference-in-Differences: Identification under Higher-Order Parallelism

Published 16 Jun 2026 in econ.EM | (2606.17977v1)

Abstract: In difference-in-differences designs, the parallel trends assumption requires that the outcome gap between treated and control units would have remained flat absent treatment. Pre-treatment event studies frequently reject this flat-gap requirement. Existing responses include parametric trend controls and bounds on the treatment effect under assumptions about the magnitude of the violation. This paper shows that point identification of cohort-specific and aggregate treatment effects in staggered designs remains achievable under strictly weaker assumptions. I replace the flat-gap requirement with a hierarchy of higher-order conditions, Parallel[p], embed this framework in the group-time average treatment effect structure of Callaway and Sant'Anna (2021), and prove an aggregation theorem for the case where different cohorts are identified under different feasible polynomial orders, a challenge unique to staggered designs that has not been previously addressed. A sequential order-selection procedure guides applied practice. Monte Carlo evidence confirms that post-selection bootstrap coverage remains near-nominal and that inference is robust to realistic serial correlation. Applied to Medicaid expansion data, the method yields point estimates resting on an assumption the pre-treatment data do not reject, in contrast to the flat-gap requirement which those same data decisively reject.

Authors (1)

Summary

  • The paper introduces higher-order parallel trends assumptions (Parallel[p]) to achieve point identification in staggered DiD designs.
  • It proposes the DD[p] estimator that flexibly fits cohort-specific polynomial trends for robust ATT estimation.
  • Empirical results, including a Medicaid expansion study, demonstrate improved inference and credible causal estimates.

Introduction

This paper develops a comprehensive theoretical and practical framework for identification in staggered adoption difference-in-differences (DiD) designs under relaxed assumptions that go beyond the standard parallel trends (PT) requirement. By introducing a hierarchy of higher-order parallel trends assumptions (denoted Parallel[pp]), the work rigorously demonstrates point identification of cohort-specific and aggregate average treatment effects on the treated (ATT) even when flat pre-treatment gaps are decisively rejected in event-study analyses. The proposed approach is embedded in the group-time ATT structure of Callaway and Sant'Anna (2021), culminating in an aggregation theorem (Theorem 4.4) which allows for consistent estimation and inference across cohorts with heterogeneous feasible polynomial orders—an identification challenge unique to staggered adoption designs.

Framework and Assumption Relaxation

Traditional DiD estimators critically rely on the assumption that, absent treatment, the trajectory of treated and non-treated (control) units would be parallel—the outcome gap would remain exactly flat post-treatment (Parallel[1]). However, standard diagnostic event studies often reject PT, especially in panels with long pre-intervention periods, resulting in unknown bias and rendering conventional estimators unreliable. Existing responses—such as parametric trend controls, bounding, and sensitivity analyses—either impose untestable functional forms or refrain from point identification.

This paper relaxes these constraints by formalizing a hierarchy of higher-order parallelism:

  • Parallel[pp]: only the pp-th order time difference (rather than the first) of the untreated outcome gap must be constant across groups.
    • p=1p=1: standard DiD (flat gap).
    • p=2p=2: allows for a common linear trend (i.e., the difference may drift linearly).
    • p=3p=3: allows for quadratic trajectories, etc.

This hierarchy is strictly nested, with each relaxation weakening the requirements for point identification in direct correspondence with evidence from pre-treatment data.

Polynomial Structure and Cohort-Specific Counterfactuals

Under Parallel[pp], the pre-treatment gap for cohort gg is required to lie on a polynomial of degree p−1p-1, which is directly checkable. The identifying restrictions are testable using pre-intervention data; statistical diagnostics and in-sample R2R^2 guide selection. Missing flatness but finding strong linearity (or higher-order fit) justifies shifting from Parallel[1] to a higher order, thereby permitting viable counterfactual projections when flatness is counterfactually implausible.

Theoretical Contributions

Identification and Aggregation Under Heterogeneous Orders

Within the Callaway-Sant'Anna framework, the paper proves:

  • Cohort-Specific ATT Identification (Theorem 4.3): Under Parallel[pp0] and sufficient pre-treatment periods (pp1), the post-treatment ATT for each cohort pp2 at time pp3 is nonparametrically identified via extrapolation of the cohort-specific, data-driven polynomial in pre-treatment gap.
  • Heterogeneous-Order Aggregation (Theorem 4.4): For panels with staggered adoption, cohorts can support different polynomial orders (i.e., later treated groups have more pre-periods and can support higher pp4), raising nontrivial aggregation problems. Theorem 4.4 demonstrates that a properly-weighted average of cohort-time ATTs remains point identified, even when each is recovered under its own highest-supported pp5.

These results generalize previous work by enabling order-heterogeneous aggregation, a critical advance for applied settings where staggered adoption is typical and pre-trend evidence differs by cohort.

Estimation, Inference, and Algorithmic Selection

The proposed estimator, pp6, fits the highest-supported pre-treatment polynomial to each cohort and extrapolates this (rather than a flat line) into the post-treatment. Finite-sample estimation uses OLS on all pre-treatment gaps (for efficiency), and the cluster bootstrap provides consistent inference, addressing within-cohort and cross-cohort dependencies induced by overlapping control groups.

Algorithmically, a sequential order-selection procedure (based on joint overidentification tests and pp7 diagnostics) selects the smallest pp8 for which the associated model is not rejected in pre-treatment data for each cohort. The aggregator then produces primary and robustness estimates for all relevant pp9.

Simulation Results

Extensive Monte Carlo evidence confirms:

  • When parallel trends of order pp0 are satisfied, pp1 is unbiased and more efficient than both conventional and higher-order-misspecified estimators.
  • The sequential order-selection procedure yields near-nominal bootstrap coverage for all pp2 under correct specification (typically pp3–pp4 reported), and is robust to AR(1) serial correlation in the DGP.
  • Over-selection of pp5 (fitting needlessly flexible polynomials) yields increased variance without systematic bias; under-selection (too parsimonious) induces bias if trends in higher moments are present.

The simulation analysis also demonstrates that polynomial-based extrapolation remains reliable for short post-treatment windows but variance can increase substantially with long extrapolation horizons, offering direct quantitative guidance for practitioners.

Empirical Illustration: Medicaid Expansion

Applying the method to Medicaid expansion under the ACA, pre-treatment event studies for state-level insurance coverage data decisively reject flat parallel trends but support a cohort-specific linear trend (i.e., Parallel[2]) with high in-sample pp6. The standard DiD estimator assuming Parallel[1] is thus indefensible in this context.

Key findings are:

  • Standard DiD (Parallel[1]) recovers an ATT of pp7 percentage points (95% CI: pp8).
  • pp9 (Parallel[2]), justified by the higher-order parallelism test, yields a \emph{slightly} higher estimate (p=1p=10–p=1p=11 percentage points depending on weighting), fully supported by the pre-treatment evidence.
  • The direction and magnitude of the p=1p=12 correction varies by cohort, reflecting heterogeneous and non-flat pre-trends.
  • Confidence intervals from the cluster bootstrap overlap substantially with those from the standard estimator, indicating robustness and that applying p=1p=13 does not reverse but rather solidifies the conclusion of positive effect, now defensible on more credible identifying grounds.

Practical and Theoretical Implications

The paper provides an actionable roadmap for empirical researchers faced with systematic pre-trends—permitting transparent, data-driven selection of the weakest assumption not decisively contradicted by the available data, while achieving point identification when feasible. This approach improves the credibility and interpretability of staggered DiD designs without defaulting to partially-identifying bounds, which can be inconclusive or unwieldy.

From a theoretical perspective, the aggregation theorem resolves a fundamental identification issue in multi-cohort settings, laying groundwork for richer nonparametric modeling of untreated potential outcomes. It also formalizes limits of polynomial extrapolation, informs best practices for diagnostic testing, and clarifies the relationship to recent innovations (e.g., sensitivity bounds, generalized synthetic controls, and triple-differences designs).

Limitations and Directions for Further Research

Potential limitations arise when the true counterfactual dynamics are not polynomial or when pre-treatment periods are few. In such cases, diagnostic statistics (in-sample p=1p=14, failure of polynomial fit) provide early warnings, and sensitivity interval approaches (e.g., derivative-bounded bounds) should be preferred. The paper notes the need for formal post-selection inference theory for the sequential approach—a direction for future work, as is extending to non-polynomial smoothness restrictions.

Conclusion

This paper structurally relaxes the parallel trends assumption in staggered DiD settings, offering point identification and robust inference under verifiable higher-order pre-trend conditions. It formalizes polynomial-based projection as a credible alternative to flat-gap DiD when the data demand it and uniquely addresses identification and aggregation where feasibility varies across cohorts. The methodology and software implementation directly advance empirical practice, providing both the inferential rigor and diagnostic transparency essential for credible causal inference in policy evaluation and applied econometrics.


Reference: "Beyond Parallel Trends in Staggered Difference-in-Differences: Identification under Higher-Order Parallelism" (2606.17977)

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.