- The paper introduces higher-order parallel trends assumptions (Parallel[p]) to achieve point identification in staggered DiD designs.
- It proposes the DD[p] estimator that flexibly fits cohort-specific polynomial trends for robust ATT estimation.
- Empirical results, including a Medicaid expansion study, demonstrate improved inference and credible causal estimates.
Beyond Parallel Trends in Staggered Difference-in-Differences: Higher-Order Parallelism and Identification
Introduction
This paper develops a comprehensive theoretical and practical framework for identification in staggered adoption difference-in-differences (DiD) designs under relaxed assumptions that go beyond the standard parallel trends (PT) requirement. By introducing a hierarchy of higher-order parallel trends assumptions (denoted Parallel[p]), the work rigorously demonstrates point identification of cohort-specific and aggregate average treatment effects on the treated (ATT) even when flat pre-treatment gaps are decisively rejected in event-study analyses. The proposed approach is embedded in the group-time ATT structure of Callaway and Sant'Anna (2021), culminating in an aggregation theorem (Theorem 4.4) which allows for consistent estimation and inference across cohorts with heterogeneous feasible polynomial orders—an identification challenge unique to staggered adoption designs.
Framework and Assumption Relaxation
Standard and Higher-Order Parallel Trends
Traditional DiD estimators critically rely on the assumption that, absent treatment, the trajectory of treated and non-treated (control) units would be parallel—the outcome gap would remain exactly flat post-treatment (Parallel[1]). However, standard diagnostic event studies often reject PT, especially in panels with long pre-intervention periods, resulting in unknown bias and rendering conventional estimators unreliable. Existing responses—such as parametric trend controls, bounding, and sensitivity analyses—either impose untestable functional forms or refrain from point identification.
This paper relaxes these constraints by formalizing a hierarchy of higher-order parallelism:
- Parallel[p]: only the p-th order time difference (rather than the first) of the untreated outcome gap must be constant across groups.
- p=1: standard DiD (flat gap).
- p=2: allows for a common linear trend (i.e., the difference may drift linearly).
- p=3: allows for quadratic trajectories, etc.
This hierarchy is strictly nested, with each relaxation weakening the requirements for point identification in direct correspondence with evidence from pre-treatment data.
Polynomial Structure and Cohort-Specific Counterfactuals
Under Parallel[p], the pre-treatment gap for cohort g is required to lie on a polynomial of degree p−1, which is directly checkable. The identifying restrictions are testable using pre-intervention data; statistical diagnostics and in-sample R2 guide selection. Missing flatness but finding strong linearity (or higher-order fit) justifies shifting from Parallel[1] to a higher order, thereby permitting viable counterfactual projections when flatness is counterfactually implausible.
Theoretical Contributions
Identification and Aggregation Under Heterogeneous Orders
Within the Callaway-Sant'Anna framework, the paper proves:
- Cohort-Specific ATT Identification (Theorem 4.3): Under Parallel[p0] and sufficient pre-treatment periods (p1), the post-treatment ATT for each cohort p2 at time p3 is nonparametrically identified via extrapolation of the cohort-specific, data-driven polynomial in pre-treatment gap.
- Heterogeneous-Order Aggregation (Theorem 4.4): For panels with staggered adoption, cohorts can support different polynomial orders (i.e., later treated groups have more pre-periods and can support higher p4), raising nontrivial aggregation problems. Theorem 4.4 demonstrates that a properly-weighted average of cohort-time ATTs remains point identified, even when each is recovered under its own highest-supported p5.
These results generalize previous work by enabling order-heterogeneous aggregation, a critical advance for applied settings where staggered adoption is typical and pre-trend evidence differs by cohort.
Estimation, Inference, and Algorithmic Selection
The proposed estimator, p6, fits the highest-supported pre-treatment polynomial to each cohort and extrapolates this (rather than a flat line) into the post-treatment. Finite-sample estimation uses OLS on all pre-treatment gaps (for efficiency), and the cluster bootstrap provides consistent inference, addressing within-cohort and cross-cohort dependencies induced by overlapping control groups.
Algorithmically, a sequential order-selection procedure (based on joint overidentification tests and p7 diagnostics) selects the smallest p8 for which the associated model is not rejected in pre-treatment data for each cohort. The aggregator then produces primary and robustness estimates for all relevant p9.
Simulation Results
Extensive Monte Carlo evidence confirms:
- When parallel trends of order p0 are satisfied, p1 is unbiased and more efficient than both conventional and higher-order-misspecified estimators.
- The sequential order-selection procedure yields near-nominal bootstrap coverage for all p2 under correct specification (typically p3–p4 reported), and is robust to AR(1) serial correlation in the DGP.
- Over-selection of p5 (fitting needlessly flexible polynomials) yields increased variance without systematic bias; under-selection (too parsimonious) induces bias if trends in higher moments are present.
The simulation analysis also demonstrates that polynomial-based extrapolation remains reliable for short post-treatment windows but variance can increase substantially with long extrapolation horizons, offering direct quantitative guidance for practitioners.
Empirical Illustration: Medicaid Expansion
Applying the method to Medicaid expansion under the ACA, pre-treatment event studies for state-level insurance coverage data decisively reject flat parallel trends but support a cohort-specific linear trend (i.e., Parallel[2]) with high in-sample p6. The standard DiD estimator assuming Parallel[1] is thus indefensible in this context.
Key findings are:
- Standard DiD (Parallel[1]) recovers an ATT of p7 percentage points (95% CI: p8).
- p9 (Parallel[2]), justified by the higher-order parallelism test, yields a \emph{slightly} higher estimate (p=10–p=11 percentage points depending on weighting), fully supported by the pre-treatment evidence.
- The direction and magnitude of the p=12 correction varies by cohort, reflecting heterogeneous and non-flat pre-trends.
- Confidence intervals from the cluster bootstrap overlap substantially with those from the standard estimator, indicating robustness and that applying p=13 does not reverse but rather solidifies the conclusion of positive effect, now defensible on more credible identifying grounds.
Practical and Theoretical Implications
The paper provides an actionable roadmap for empirical researchers faced with systematic pre-trends—permitting transparent, data-driven selection of the weakest assumption not decisively contradicted by the available data, while achieving point identification when feasible. This approach improves the credibility and interpretability of staggered DiD designs without defaulting to partially-identifying bounds, which can be inconclusive or unwieldy.
From a theoretical perspective, the aggregation theorem resolves a fundamental identification issue in multi-cohort settings, laying groundwork for richer nonparametric modeling of untreated potential outcomes. It also formalizes limits of polynomial extrapolation, informs best practices for diagnostic testing, and clarifies the relationship to recent innovations (e.g., sensitivity bounds, generalized synthetic controls, and triple-differences designs).
Limitations and Directions for Further Research
Potential limitations arise when the true counterfactual dynamics are not polynomial or when pre-treatment periods are few. In such cases, diagnostic statistics (in-sample p=14, failure of polynomial fit) provide early warnings, and sensitivity interval approaches (e.g., derivative-bounded bounds) should be preferred. The paper notes the need for formal post-selection inference theory for the sequential approach—a direction for future work, as is extending to non-polynomial smoothness restrictions.
Conclusion
This paper structurally relaxes the parallel trends assumption in staggered DiD settings, offering point identification and robust inference under verifiable higher-order pre-trend conditions. It formalizes polynomial-based projection as a credible alternative to flat-gap DiD when the data demand it and uniquely addresses identification and aggregation where feasibility varies across cohorts. The methodology and software implementation directly advance empirical practice, providing both the inferential rigor and diagnostic transparency essential for credible causal inference in policy evaluation and applied econometrics.
Reference: "Beyond Parallel Trends in Staggered Difference-in-Differences: Identification under Higher-Order Parallelism" (2606.17977)