Papers
Topics
Authors
Recent
Search
2000 character limit reached

Functional Data Approach to DiD

Updated 14 December 2025
  • The paper introduces a functional framework that models outcome trajectories as elements in a Banach space, enabling uniform inference over time.
  • It employs double-demeaning and a functional CLT to yield estimators converging to Gaussian processes, allowing honest simultaneous confidence bands.
  • Practical applications, such as studies on gender bias and employment, validate its superiority in handling parallel trends violations and anticipation effects.

The functional data approach to Difference-in-Differences (DiD) reframes the standard discrete-time panel event study within a continuous-time, infinite-dimensional stochastic process framework. Each unit’s outcome trajectory is modeled as a random element in the Banach space of continuous functions, enabling rigorous simultaneous causal inference over time intervals, not just at isolated points. This paradigm, introduced by Fang & Liebl (Fang et al., 7 Dec 2025), replaces conventional pointwise inference with uniform functional inference, directly addressing limitations in common event study plots where the parallel trends and no-anticipation assumptions may be violated. The approach yields estimators that converge to Gaussian processes, supports construction of honest simultaneous confidence bands (SCBs), and enables principled equivalence and relevance testing across intervals—transforming event study plots into comprehensive causal inference tools.

1. Functional Data Framework

Let each observational unit i=1,,ni=1,\ldots,n generate an outcome trajectory Yi()Y_i(\cdot) as a function in C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}], but only discrete-time samples Yi,t=Yi(t)Y_{i,t}=Y_i(t) for t{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\} are observed. Each unit is assigned a post-treatment indicator Di{0,1}D_i\in\{0,1\}, so that observed outcomes follow the potential outcome model:

Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)

where Yi(t,d)Y_i(t,d) denotes the potential outcome under treatment status dd. The canonical event-study DiD parameter process is defined as:

β(t):=E[Yi(t)Yi(0)Di=1]E[Yi(t)Yi(0)Di=0],t[Tpre,Tpost]\beta(t) := \mathbb{E}[Y_i(t)-Y_i(0)\mid D_i=1] - \mathbb{E}[Y_i(t)-Y_i(0)\mid D_i=0], \quad t\in [-T_{\text{pre}}, T_{\text{post}}]

with Yi()Y_i(\cdot)0 by construction.

Under standard DiD assumptions:

  • No anticipation: Yi()Y_i(\cdot)1
  • Parallel trends: Yi()Y_i(\cdot)2 for all Yi()Y_i(\cdot)3
  • Overlap: Yi()Y_i(\cdot)4 for some Yi()Y_i(\cdot)5

it follows that Yi()Y_i(\cdot)6, the average treatment effect on the treated at each Yi()Y_i(\cdot)7.

To handle unit and time fixed effects, the model employs “double-demeaning” in Yi()Y_i(\cdot)8 and Yi()Y_i(\cdot)9, resulting in the oracle regression

C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}]0

where C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}]1, C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}]2, and C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}]3 is a mean-zero error process in C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}]4.

The least-squares estimator for C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}]5 is constructed as

C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}]6

with C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}]7, C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}]8.

Regularity conditions include independence of C[Tpre,Tpost]C[-T_{\text{pre}},T_{\text{post}}]9, bounded higher moments Yi,t=Yi(t)Y_{i,t}=Y_i(t)0, Yi,t=Yi(t)Y_{i,t}=Y_i(t)1, Yi,t=Yi(t)Y_{i,t}=Y_i(t)2, Yi,t=Yi(t)Y_{i,t}=Y_i(t)3, and twice-continuous differentiability of Yi,t=Yi(t)Y_{i,t}=Y_i(t)4, Yi,t=Yi(t)Y_{i,t}=Y_i(t)5, Yi,t=Yi(t)Y_{i,t}=Y_i(t)6, and Yi,t=Yi(t)Y_{i,t}=Y_i(t)7.

2. Uniform Central Limit Theorem and Gaussian Process Limit

The estimator Yi,t=Yi(t)Y_{i,t}=Y_i(t)8 constitutes a stochastic process indexed by Yi,t=Yi(t)Y_{i,t}=Y_i(t)9 in t{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}0, furnished with the sup-norm t{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}1. The population covariance kernel is given by:

t{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}2

Under the stated regularity conditions, the following functional CLT holds:

t{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}3

The formal proof comprises three components:

  • Pointwise CLTs at each t{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}4,
  • Equicontinuity bounds via Ct{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}5-smoothness,
  • Application of functional CLT machinery (Pollard 1984; Hahn 1977; Billingsley).

No-anticipation and parallel trends are not required for this functional CLT.

3. Honest Simultaneous Confidence Bands

From the Gaussian process limit, simultaneous confidence bands in sup-norm covering are constructed as:

t{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}6

for t{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}7 in the post-treatment window t{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}8. The critical value t{Tpre,,Tpost}t\in\{-T_{\text{pre}},\ldots,T_{\text{post}}\}9 is calibrated such that

Di{0,1}D_i\in\{0,1\}0

where Di{0,1}D_i\in\{0,1\}1.

Calibration approaches include:

  • Parametric (Gaussian) bootstrap: Simulating GP samples Di{0,1}D_i\in\{0,1\}2 at the grid, spline interpolation, and quantile computation.
  • Multiplier bootstrap: Reweighting residuals by i.i.d. weights and recomputing estimators.
  • Kac-Rice formula: Closed-form quantile approximation utilizing covariance curvature traces (Liebl–Reimherr 2023).

Uniform coverage is guaranteed asymptotically:

Di{0,1}D_i\in\{0,1\}3

4. Equivalence Testing for Pre-Anticipation Window

Suppose a reference band Di{0,1}D_i\in\{0,1\}4 for Di{0,1}D_i\in\{0,1\}5 is postulated under a compound null Di{0,1}D_i\in\{0,1\}6 for Di{0,1}D_i\in\{0,1\}7, with Di{0,1}D_i\in\{0,1\}8 defining the anticipation window’s start. The test distinguishes:

  • Di{0,1}D_i\in\{0,1\}9: Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)0 such that Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)1
  • Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)2: Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)3, Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)4

Infimum SCBs are constructed:

Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)5

Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)6

Reject Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)7 (“Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)8”) if Yi(t)=DiYi(t,1)+(1Di)Yi(t,0)Y_i(t) = D_i\,Y_i(t,1) + (1-D_i)\,Y_i(t,0)9 Yi(t,d)Y_i(t,d)0; reject Yi(t,d)Y_i(t,d)1 (“Yi(t,d)Y_i(t,d)2”) if Yi(t,d)Y_i(t,d)3 Yi(t,d)Y_i(t,d)4.

A joint Yi(t,d)Y_i(t,d)5-level test requires both bands to be contained within Yi(t,d)Y_i(t,d)6 Yi(t,d)Y_i(t,d)7. The test enjoys asymptotic size Yi(t,d)Y_i(t,d)8 via the functional CLT and consistent covariance estimation.

5. Relevance Testing for Post-Treatment Effects

A parallel relevance test is formulated for the post-treatment window using the same reference band:

  • Yi(t,d)Y_i(t,d)9: dd0 dd1
  • dd2: dd3

Rejection occurs if the supremum confidence band dd4 fails to intersect dd5 for any dd6. This test holds asymptotic size dd7.

6. Empirical Validation and Applications

Simulation results indicate:

  • Interpolation error decays at rate dd8.
  • Under parallel-trends violations, sup-SCB tests maintain nominal Type I error control and exhibit higher detection power than Bonferroni-corrected pointwise bands, which are invalid.
  • Under anticipation, inf-SCB equivalence tests control size when the reference band coincides with dd9 during β(t):=E[Yi(t)Yi(0)Di=1]E[Yi(t)Yi(0)Di=0],t[Tpre,Tpost]\beta(t) := \mathbb{E}[Y_i(t)-Y_i(0)\mid D_i=1] - \mathbb{E}[Y_i(t)-Y_i(0)\mid D_i=0], \quad t\in [-T_{\text{pre}}, T_{\text{post}}]0 and demonstrate power against mis-specified bands.

Case studies demonstrate robust practical utility:

  • For gender bias in livestreamed courts (Chen et al 2025), the honest event study plot first validates the reference band via infimum-SCB over β(t):=E[Yi(t)Yi(0)Di=1]E[Yi(t)Yi(0)Di=0],t[Tpre,Tpost]\beta(t) := \mathbb{E}[Y_i(t)-Y_i(0)\mid D_i=1] - \mathbb{E}[Y_i(t)-Y_i(0)\mid D_i=0], \quad t\in [-T_{\text{pre}}, T_{\text{post}}]1 and then confirms significant uniform post-period effects over β(t):=E[Yi(t)Yi(0)Di=1]E[Yi(t)Yi(0)Di=0],t[Tpre,Tpost]\beta(t) := \mathbb{E}[Y_i(t)-Y_i(0)\mid D_i=1] - \mathbb{E}[Y_i(t)-Y_i(0)\mid D_i=0], \quad t\in [-T_{\text{pre}}, T_{\text{post}}]2 via supremum-SCB.
  • For duty-to-bargain laws and female employment (Lovenheim & Willén 2019), the reference band could not be rigorously validated (inf-SCB intersects band), yet the post-treatment sup-band fails to reject the null, indicating no significant causal effect once pre-trend is accommodated.

7. Software Implementation

The R package fdid (fang_liebl_2025_R) operationalizes the functional DiD approach by:

  1. Fitting the functional DiD estimator via TWFE.
  2. Performing natural-cubic-spline interpolation.
  3. Estimating the covariance surface.
  4. Computing sup- and inf-SCBs through parametric, multiplier, or Kac-Rice procedures.
  5. Executing relevance and equivalence tests.
  6. Rendering “honest” event-study plots.

This comprehensive workflow provides transparent, uniform inference for causal effects over time, rectifying deficiencies inherent in pointwise event study analysis (Fang et al., 7 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Functional Data Approach to DiD.