---
title: Conditional Parallel Trends (CPT)
url: https://www.emergentmind.com/topics/conditional-parallel-trends-cpt
type: topic
---

# Conditional Parallel Trends (CPT)

Conditional parallel trends (CPT) is the difference-in-differences identifying restriction that requires equality of counterfactual untreated trends after conditioning on observed covariates. In the canonical two-period, two-group setting, it is written as
$$
E[\Delta Y_1(0)\mid X,D=1]=E[\Delta Y_1(0)\mid X,D=0],
$$
where $\Delta Y_1(0)=Y_1(0)-Y_0(0)$. Under overlap and consistency, this yields
$$
ATT = E[\Delta Y_1\mid D=1]-E\!\left[E[\Delta Y_1\mid X,D=0]\mid D=1\right].
$$
Recent work treats CPT as a structural statement about untreated potential outcomes, selection, covariate dynamics, and valid conditioning sets, rather than as a purely informal claim that observed pre-treatment trajectories “look parallel” [2604.12818] [2406.15288].

## 1. Canonical formulations and target estimands

In the standard two-period setup, unconditional parallel trends requires
$$
E[\Delta Y_{t^*}(0)\mid D=1] = E[\Delta Y_{t^*}(0)\mid D=0].
$$
CPT replaces this with a conditional restriction. In the mixed-covariate case emphasized in recent work, the conditioning set includes time-varying covariates in both periods and time-invariant covariates:
$$
E[\Delta Y_{t^*}(0)\mid X_{t^*},X_{t^*-1},Z,D=1]
=
E[\Delta Y_{t^*}(0)\mid X_{t^*},X_{t^*-1},Z,D=0].
$$
The corresponding identifying equation for the average treatment effect on the treated is
$$
ATT = E[\Delta Y_{t^*}\mid D=1] - E\Big[ E[\Delta Y_{t^*}\mid X_{t^*},X_{t^*-1},Z,D=0] \mid D=1 \Big].
$$
This formulation is central in settings where treated and untreated units differ in observables and those differences may drive trend differences [2406.15288].

The same logic extends to staggered-adoption settings. One formulation is the conditional staggered parallel trends assumption,
$$
E[Y_{i,t}(\infty)-Y_{i,T+1}(\infty)\mid G_i^r=1,X_i]
=
E[Y_{i,t}(\infty)-Y_{i,T+1}(\infty)\mid G_i^\infty=1,X_i],
\qquad t,r=1,\ldots,\overline{T}.
$$
Here the identifying restriction remains a conditional equality of untreated trends, but the conditioning occurs at the cohort level and relative to a base period [2310.15796].

A recurring interpretive point is that CPT is a statement about untreated potential outcomes, not about observed outcomes. This distinction becomes important once treatment timing is staggered, covariates are time-varying, or covariate adjustment is implemented through regression rather than through an estimand that explicitly conditions on the full identifying covariate set [2412.14447] [2406.15288].

## 2. Selection-based and structural interpretations

A major line of work interprets parallel trends through the treatment-selection mechanism. In a standard $2\times 2$ design, the key unconditional restriction is
$$
E[Y_{i2}(0)-Y_{i1}(0)\mid G_i=1] = E[Y_{i2}(0)-Y_{i1}(0)\mid G_i=0].
$$
Under a general nonseparable outcome model and a general selection mechanism, parallel trends is highly restrictive. With unrestricted selection, it holds for all nontrivial selection mechanisms if and only if untreated potential outcomes are constant over time up to a common mean shift:
$$
\dot Y_{i1}(0)=\dot Y_{i2}(0)\quad a.s.,
$$
where $\dot Y_{it}(0)\equiv Y_{it}(0)-E[Y_{it}(0)]$ [2203.09001].

Once selection is restricted, weaker primitive conditions suffice. Under selection based on pre-treatment information only, the necessary condition is a martingale-type restriction,
$$
E[\dot Y_{i2}(0)\mid \alpha_i,\varepsilon_{i1}] = \dot Y_{i1}(0)\quad a.s.
$$
Under selection on fixed effects, the necessary condition is time homogeneity,
$$
E[\dot Y_{i1}(0)\mid \nu_i]=E[\dot Y_{i2}(0)\mid \nu_i]\quad a.s.
$$
In separable two-way models, these become conditions on the evolution of the idiosyncratic component rather than on the full untreated outcome process [2203.09001].

The covariate-adjusted version is
$$
E[Y_{i2}(0)-Y_{i1}(0)\mid G_i=1,X_i]
=
E[Y_{i2}(0)-Y_{i1}(0)\mid G_i=0,X_i]\quad a.s.
$$
This formulation identifies
$$
ATT = E[ATT(X_i)\mid G_i=1] = E[DiD(X_i)\mid G_i=1].
$$
But the same work shows that CPT conditional on the full time path $X_i=(X_{i1},X_{i2})$ implies strong separability restrictions in the untreated outcome model. When covariates interact with unobservables, a weaker modified assumption is proposed:
$$
E[Y_{i2}(0)-Y_{i1}(0)\mid G_i=1,X_i^\lambda,X_{i1}^\mu=X_{i2}^\mu]
=
E[Y_{i2}(0)-Y_{i1}(0)\mid G_i=0,X_i^\lambda,X_{i1}^\mu=X_{i2}^\mu]\quad a.s.
$$
This weaker condition identifies treatment effects only for subpopulations in which the interacting covariates do not change over time [2203.09001].

This suggests that CPT is not merely a conditional balancing statement. A plausible implication is that its content depends on whether covariates enter the untreated outcome process additively or nonseparably, and on whether selection responds to fixed heterogeneity, pre-treatment information, or time-varying unobservables.

## 3. Graphical criteria and valid conditioning sets

Recent graphical work recasts CPT in terms of transformed Single World Intervention Graphs, the $\Delta$-SWIGs. The central point is that ordinary DAGs and ordinary SWIGs do not directly encode the “difference world” relevant for difference-in-differences, because the relevant object is $\Delta Y(0)$ rather than a level potential outcome such as $Y_1(0)$. In the canonical $2\times 2$ case, the target is
$$
ATT := E[Y_1(1)-Y_1(0)\mid D=1],
$$
and CPT is
$$
E[\Delta Y_1(0)\mid X,D=1]=E[\Delta Y_1(0)\mid X,D=0].
$$
Under overlap and consistency this yields the standard identification formula above [2604.12818].

The key structural condition is single world additive separability:
$$
Y_0 = \alpha(U,X)+g_{Y_0}(X,U_{Y_0}),\qquad
Y_1 = \alpha(U,X)+g_{Y_1}(X,U_{Y_1}) + D\cdot \tau(U,X,U_{Y_1}).
$$
Under this assumption,
$$
\Delta Y_1(0)=g_{Y_1}(X,U_{Y_1})-g_{Y_0}(X,U_{Y_0}),
$$
so the time-invariant unobservable $U$ cancels from the difference node. In a pruned $\Delta$-SWIG with time-invariant $X$, this yields
$$
\Delta Y_1(0)\perp\!\!\!\perp D\mid X,
$$
which immediately implies CPT [2604.12818].

The same framework sharply distinguishes valid and invalid controls. In the $2\times 2$ setting, pre-treatment outcome $Y_0$ is a “bad control” because conditioning on it can open collider paths such as
$$
D \leftarrow U \rightarrow Y_0 \leftarrow U_{Y_0}\rightarrow \Delta Y_1(0).
$$
With time-varying covariates, the relevant CPT restriction can become
$$
\Delta Y_1(0)\perp\!\!\!\perp D \mid X_0,X_1,
$$
but outcome dynamics, outcome-treatment feedback, and outcome-covariate feedback can destroy CPT by creating dilemma nodes for which no observable conditioning set yields $d$-separation [2604.12818].

A related line of work develops causal-diagram criteria under linear faithfulness. It shows that parallel trends can be rejected if treatment is affected by pre-treatment outcomes, if pre- and post-treatment outcomes possess distinct minimally sufficient sets, or if pre-treatment outcomes affect post-treatment outcomes in a way that requires “remarkable coincidence” or exact cancellation. When those features are absent, a necessary and sufficient condition is
$$
E(Y_1^0-Y_0^0\mid M)=E(Y_1^0-Y_0^0)\quad a.s.,
$$
for a common sufficient adjustment set $M$. The paper calls this additive homogeneous confounding, and interprets it as constancy of the confounding effect across time on the additive scale [2505.03526].

Graphical work also changes the status of pre-trend evidence. In multi-period settings with time-varying covariates, pre-treatment parallel trends are informative only about a subset of the assumptions required for unbiased post-treatment effects, especially when treatment-covariate feedback is possible [2604.12818].

## 4. Covariates in estimation: regression pitfalls, CCC, and alternative estimators

A central recent critique is that CPT alone is not sufficient for standard difference-in-differences implementations with time-varying covariates. One paper introduces the two-way common causal covariates assumption and argues that DiD with covariates also requires a stability condition on how covariates affect outcomes. The three CCC variants are
$$
\gamma^i=\gamma^j \quad \text{for } i\neq j,
$$
$$
\gamma^s=\gamma^t \quad \text{for } s\neq t,
$$
and
$$
\gamma^{i,s}=\gamma^{j,t}.
$$
The intuition is that the causal effect of the covariate on the outcome must be the same across groups and across time. In the paper’s formulation, CPT is about untreated potential outcomes, whereas CCC is about whether the covariate-outcome relationship is stable enough for standard covariate-adjusted DiD estimators to recover the ATT [2412.14447].

This distinction matters because standard TWFE typically estimates a single pooled coefficient on covariates:
$$
Y_{i,g,t} = \alpha_g + \delta_t + \beta^{DD}D_{i,g,t} + \sum_k \gamma^k X^k_{i,g,t} + \epsilon_{i,g,t}.
$$
When the true data-generating process has coefficients $\gamma_{g,t}$ that vary by group and time, the estimator obeys
$$
\widehat{\beta}^{DD} = \tau + \text{bias}.
$$
The paper derives an explicit bias expression and shows that standard TWFE and CS-DID are biased when the two-way CCC assumption is violated; it also argues that CS-DID can still be biased with time-varying covariates even when CCC holds [2412.14447].

The proposed response is the Intersection Difference-in-differences estimator. DID-INT first estimates
$$
Y_{i,g,t} = \sum_g\sum_t \lambda_{g,t} I(g,t) + f(X^k_{i,g,t}) + \epsilon_{i,g,t},
$$
then computes long differences, group-time ATT contrasts, and weighted aggregates. Its four functional forms are homogeneous, state-varying, time-varying, and two-way. The identification result is
$$
\big(\lambda_{g,r}-\lambda_{g,r-1}\big)-\big(\lambda_{g',r}-\lambda_{g',r-1}\big)=\tau,
$$
and the paper interprets DID-INT as working because it uses a flexible enough residualization of outcomes so that parallel trends can hold for the residuals [2412.14447].

A complementary critique concerns hidden linearity bias in TWFE under CPT. In the two-period case, the canonical regression
$$
Y_{it}=\theta_t+\eta_i+\alpha D_{it}+X_{it}'\beta+e_{it}
$$
becomes, after first differencing,
$$
\Delta Y_{it^*}=\alpha D_i+\Delta X_{it^*}'\beta+\Delta e_{it^*}.
$$
The underlying CPT rationale, however, may require conditioning on $(X_{t^*},X_{t^*-1},Z)$ rather than on $\Delta X_{t^*}$ alone. Starting from
$$
Y_{it}(0)=\theta_t+Z_i'\delta_t+X_{it}'\beta_t+\eta_i+e_{it},
$$
first differencing gives
$$
\Delta Y_{it}(0) = \tilde{\theta}_t + Z_i'\tilde{\delta}_t + \Delta X_{it}'\beta_t + X_{it-1}'\tilde{\beta}_t + \Delta e_{it}.
$$
TWFE nevertheless drops $Z_i$ and $X_{it-1}$ and retains only $\Delta X_{it}$. The resulting decomposition is
$$
\alpha = E\big[w(\Delta X_{t^*})ATT(X_{t^*},X_{t^*-1},Z)\mid D=1\big] + A + B + C,
$$
where $A$, $B$, and $C$ are bias terms associated with omitted time-invariant covariates, dependence on levels rather than changes of time-varying covariates, and nonlinearity in the conditional mean [2406.15288].

The same paper proposes diagnostics based on the implicit regression weights and recommends checking whether TWFE weights balance not only $\Delta X$ but also the levels $X_{t^*}$, $X_{t^*-1}$, and $Z$. As an alternative, it proposes augmented inverse propensity weighting estimators that explicitly condition on the CPT-identifying covariates and are doubly robust if either the outcome regression or propensity score model is correct [2406.15288].

## 5. Empirical assessment, falsification, and the interpretation of pre-trends

The most common empirical check for CPT or its unconditional analogue is a pre-trends test. A major critique is that the sampling distribution of the post-treatment estimate changes when inference is conditioned on having passed that pre-test. Under joint normality, one paper studies the event
$$
\hat\beta_{pre}\in B,
$$
and the orthogonalized post coefficient
$$
\tilde\beta_{post} = \hat\beta_{post} - \Sigma_{12}\Sigma_{22}^{-1}\hat\beta_{pre}.
$$
When parallel trends is true, the traditional estimator remains unbiased after conditioning on passing the pre-trends test,
$$
E[\hat\beta_{post}\mid \hat\beta_{pre}\in B_{NS}] = \beta_{post},
$$
but conditional variance is smaller than unconditional variance, so traditional standard errors are conservative. When parallel trends is false but the pre-test is nevertheless passed, conditioning generally induces bias, and under monotone pre-trends that bias is exacerbated [1804.01208].

This critique does not imply that pre-treatment evidence is irrelevant. Another line of work reformulates pre-trend assessment as equivalence testing. For a single placebo coefficient $\beta_l$ and threshold $\mathcal U>0$, the proposed test is
$$
H_0: |\beta_l|\ge \mathcal{U}
\qquad \text{vs.} \qquad
H_1: |\beta_l|< \mathcal{U}.
$$
Joint formulations use the maximum norm, the average placebo effect, or the root mean square placebo effect. The aim is not to test whether violations are exactly zero, but whether they are small enough to be negligible [2310.15796].

A closely related non-inferiority framework argues that the relevant object is not simply a nuisance parameter such as a differential trend slope, but the extent to which a more flexible model changes the estimated treatment effect. In the basic DiD example, the reduced and expanded models differ by a group-specific slope term, and the difference in average treatment effects satisfies
$$
\hat{\beta} - \hat{\beta}' = \left( \frac{1}{T-T_0}\sum_{t=T_0}^T t - \frac{1}{T_0-1}\sum_{t=1}^{T_0-1} t \right)\hat{\theta}.
$$
The proposed “one step up” method fits a base model with a linear trend difference and tests whether the resulting treatment effect is within a substantively chosen distance of the effect from the simpler model [1805.03273].

A more recent contribution replaces exact parallel trends with a conditional extrapolation assumption. Let $R(t)$ denote iterative violations of parallel trends and define their severity in the pre- and post-periods by
$$
S_{\text{pre}} = \left( \frac{1}{T_{\text{pre}}-1}\sum_{t=2}^{t_0-1} |R(t)|^p \right)^{1/p},
\qquad
S_{\text{post}} = \left( \frac{1}{T_{\text{post}}}\sum_{t=t_0}^{T} |R(t)|^p \right)^{1/p}.
$$
Assumption 3 states:
$$
\text{If } S_{\text{pre}} \le M,\ \text{then } S_{\text{post}} \le S_{\text{pre}}.
$$
Under this condition, if $S_{\text{pre}}\le M$, then
$$
\left|\tau_{\text{ATT}} - \tau_{\text{DD}}\right| \le K\,S_{\text{pre}},
$$
with $K$ determined by the post-treatment horizon and the norm parameter. The same paper provides conditionally valid confidence intervals given passage of the preliminary test [2510.26470].

Taken together, these results treat pre-trends as informative but incomplete. One paper states directly that pre-trends are neither necessary nor sufficient for parallel trends, and recent graphical work adds that with time-varying covariates, pre-treatment parallel trends are often informative about only a subset of the assumptions required for post-treatment identification [2505.03526] [2604.12818].

## 6. Extensions to dynamic choice, staggered adoption, and complex treatment paths

In dynamic settings, the relevant conditioning set can be treatment history rather than a static vector of covariates. One contribution defines the canonical full parallel trends assumption as
$$
E [ Y_{1} (0) - Y_{0} (0) \mid D_0 = d_0, D_1 = d_1] = \tau
$$
for all $d_0,d_1\in\{0,1\}$. The paper interprets this as a restriction on the stability of selection into untreated potential outcomes across treatment histories. Dynamic utility maximization, learning, switching costs, and option values can all generate treatment histories, but not all such mechanisms are compatible with parallel trends. Learning about the treated arm can be compatible with PT, whereas learning about the control arm, Roy-style selection, irreversible treatment, and optimal stopping generally violate it because treatment choices become functions of information that predicts future untreated changes [2207.06564].

That same work develops weaker alternatives. A partial parallel trends assumption on the untreated-risk set,
$$
E [ Y_1 (0) - Y_0 (0) \mid D_0 = 0, D_1 = d_1] = \tau_{0,1},
$$
identifies treatment effects on switchers into treatment. Forward mean stationarity,
$$
E [ Y_1 (0) - Y_0 (0) \mid D_0 = d_0] = 0,
$$
and unconditional mean stationarity,
$$
E [ Y_1 (0)-  Y_0 (0) ] = 0,
$$
provide additional alternatives when standard DiD restrictions are too strong [2207.06564].

A different extension studies complex panel designs with non-binary and non-absorbing treatments and treatment heterogeneity already present at baseline. Its key assumption is a status-quo parallel trends condition, conditional on baseline treatment $D_1$:
$$
E[Y_{t}(D_1,\ldots,D_1) - Y_{t-1}(D_1,\ldots,D_1)\mid \bm{D}]
=
E[Y_{t}(D_1,\ldots,D_1) - Y_{t-1}(D_1,\ldots,D_1)\mid D_{1}].
$$
Here the counterfactual path is the status-quo path $Y_t(D_1,\ldots,D_1)$ in which each unit keeps its period-one treatment fixed over time. Conditioning on $D_1$ is described as crucial because, without that conditioning, the assumption would rule out effects of lagged treatments once combined with the usual parallel-trends assumption [2508.07808].

This design-specific CPT-like restriction identifies actual-versus-status-quo event-study effects rather than a conventional never-treated counterfactual. It therefore extends the logic of conditional parallel trends to designs in which treatment is already varying at baseline, can move up or down over time, and is not naturally represented by a binary absorbing indicator [2508.07808].

## 7. Functional-form sensitivity and alternatives without parallel trends

Related work on the standard two-period, two-group parallel trends condition shows that robustness to outcome transformations is a much stronger property than mean parallel trends itself. Parallel trends is invariant to all strictly monotonic transformations if and only if a CDF-level condition holds:
$$
F_{Y_{i1}(0) \mid D_i=1}(y) - F_{Y_{i0}(0) \mid D_i=1}(y)
=
F_{Y_{i1}(0) \mid D_i=0}(y) - F_{Y_{i0}(0) \mid D_i=0}(y),
\quad \text{for all } y\in\mathbb R.
$$
Under mild regularity conditions, this is equivalent to the mixture representation
$$
F_{Y_{it}(0) \mid D_i = d}(y) = \theta G_{t}(y) + (1-\theta) H_{d}(y)
\quad \text{for all } y\in \mathbb R \text{ and } d,t \in \{ 0,1 \}.
$$
The paper interprets this as requiring either random assignment, stationarity of untreated outcomes, or a hybrid mixture of the two, and proposes falsification tests based on whether the implied counterfactual CDF or PMF is valid [2010.04814].

The same paper notes two special cases that delimit how restrictive functional-form robustness is. For binary outcomes, the mean fully characterizes the distribution, so mean parallel trends is automatically invariant to monotonic transformations. For normally distributed outcomes with positive variance, the CDF condition can hold only in the pure random-assignment or pure stationarity cases, not in the hybrid case [2010.04814]. This suggests that even carefully conditioned DiD designs may remain sensitive to the outcome scale unless additional distributional structure is imposed.

An alternative response is to replace parallel trends entirely. One recent framework starts from the benchmark model
$$
Y_{it}(0)=U_i+\mu_t+\varepsilon_{it},
$$
which implies
$$
E\!\left[Y_{it}(0)-Y_{i0}(0)\mid D_i=1\right] = E\!\left[Y_{it}(0)-Y_{i0}(0)\mid D_i=0\right].
$$
It then drops this additive-separable structure and identifies ATT from a latent-variable model in which pre-treatment outcomes, a reference-period outcome, post-treatment untreated outcomes, and treatment are linked through multidimensional $U_i$. The key assumption is blockwise conditional independence given $U_i$:
$$
\boldsymbol{Y}_{i}^{\text{pre}(0)}, \ \bigl(Y_{i0}(0),D_i\bigr), \ \boldsymbol{Y}_{i}^{\text{post}(0)}
\quad \text{are mutually independent conditional on } U_i.
$$
Together with completeness-type conditions, this identifies ATT without parallel trends [2601.08281].

The emergence of such alternatives does not eliminate CPT from applied work. A plausible implication is that CPT remains the central identifying restriction for DiD, but its credibility increasingly depends on explicit statements about selection, the conditioning set, covariate dynamics, outcome scale, and the estimator used to operationalize the design.

Source: https://www.emergentmind.com/topics/conditional-parallel-trends-cpt