Papers
Topics
Authors
Recent
Search
2000 character limit reached

Local Distributional Treatment Effect

Updated 12 July 2026
  • Local Distributional Treatment Effect is a causal parameter that compares entire outcome distributions between treated and control groups using localized variations in compliance, covariates, or outcome thresholds.
  • LDTE frameworks identify treatment effects by localizing in subpopulations or design spaces, enabling methods like instrumental variables, regression discontinuity, and kernel approaches.
  • Recent advances leverage machine learning, cross-fitting, and kernel methods to improve estimator robustness and capture detailed treatment heterogeneity.

Local Distributional Treatment Effect (LDTE) denotes a family of causal parameters that compare distributions rather than expectations on a localized margin. The locality can refer to the complier subpopulation under imperfect compliance, a fixed covariate value xx, a small interval around an outcome threshold y0y_0, a kernel neighborhood in outcome space, or the limit at a regression-discontinuity cutoff. The term is therefore not tied to a single universal estimand. In instrumental-variables formulations, it is the difference in the treated and untreated potential-outcome cumulative distribution functions (CDFs) among compliers; in conditional distributional regression, it is often a covariate-specific difference of conditional CDFs; and in related work on the distribution of treatment effects, locality is attached instead to the conditional law of the individual treatment effect itself (Shaw, 15 Jun 2025, Hohberg et al., 2018, Byambadalai et al., 2024, Fava, 2024).

1. Terminological scope and principal estimands

The main usages of LDTE in the recent literature differ along two axes: the object being localized and the design used for identification. Some papers localize in the population, as with compliers under imperfect compliance. Others localize in covariate space, outcome space, or design space, as in regression discontinuity. A related but distinct strand studies the distribution of the treatment effect Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0) rather than the distribution of outcomes under treatment and control (Byambadalai et al., 19 Sep 2025, Näf et al., 2024, Dijcke, 4 Apr 2025, Fava, 2024).

Strand Estimand Locality notion
IV with imperfect compliance β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C) (Byambadalai et al., 19 Sep 2025) complier subpopulation
D-IV-LATE Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y) (Shaw, 15 Jun 2025) complier subpopulation
GAMLSS conditional LDTE Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x) (Hohberg et al., 2018) fixed xx
Randomized local effect τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h) from CDF differences over [y0,y0+h][y_0,y_0+h] (Byambadalai et al., 2024) local interval in outcome space
Causal-DRF / CKTE τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x] (Näf et al., 2024) fixed y0y_00, kernel neighborhood around y0y_01
Distribution of treatment effects y0y_02 (Fava, 2024) fixed y0y_03 on treatment-effect distribution

A recurrent source of confusion is the presumption that “local” has a stable meaning across these formulations. The literature does not support that presumption. This suggests that any statement about an LDTE must specify the locality mechanism, the identifying assumptions, and whether the target is a difference of outcome distributions or the distribution of the treatment effect itself.

2. Complier-specific LDTE under imperfect compliance

In the instrumental-variables literature with binary instrument y0y_04 and binary treatment uptake, the LDTE is defined on the complier subpopulation. Let y0y_05 be the potential outcome under treatment y0y_06, let y0y_07 be the potential treatment uptake under instrument y0y_08, and impose monotonicity y0y_09. The complier population is

Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0)0

The Distributional IV-LATE at threshold Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0)1 is then

Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0)2

with the same object denoted Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0)3 in the regression-adjusted imperfect-compliance framework (Shaw, 15 Jun 2025, Byambadalai et al., 19 Sep 2025).

Identification proceeds through a Wald-like ratio of intent-to-treat differences. In the D-IV-LATE exposition, Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0)4 is written as Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0)5, where

Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0)6

and

Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0)7

In the covariate-adaptive randomization formulation, the same logic is written with strata Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0)8 as

Δi=Yi(1)Yi(0)\Delta_i=Y_i(1)-Y_i(0)9

The familiar simplified expression is

β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)0

The underlying IV assumptions are random assignment conditional on strata, exclusion, and monotonicity (Byambadalai et al., 19 Sep 2025).

Estimation is organized through Neyman-orthogonal moments and cross-fitting. In D-IV-LATE, the nuisance functions are

β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)1

with orthogonal score functions β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)2 and β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)3. Cross-fitting trains nuisance estimators on β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)4 folds and evaluates scores on the held-out fold, and the final estimator is the ratio of average score terms. The regression-adjusted imperfect-compliance estimator uses augmented residuals β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)5 and β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)6, again with β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)7-fold cross-fitting and flexible first-stage estimators such as logistic, random forest, and boosting (Shaw, 15 Jun 2025, Byambadalai et al., 19 Sep 2025).

The asymptotic theory is correspondingly strong. Under regularity and β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)8 nuisance rates, the D-IV-LATE estimator is asymptotically normal for each fixed β(y)=FY(1)(yC)FY(0)(yC)\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)9, with variance

Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)0

The imperfect-compliance paper further derives weak convergence in Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)1, states that the estimator achieves the semiparametric efficiency bound, and emphasizes applicability to continuous, discrete, and mixed discrete-continuous outcomes under stratified block designs and simple random sampling (Shaw, 15 Jun 2025, Byambadalai et al., 19 Sep 2025).

A distinctive contribution of the D-IV-LATE paper is its emphasis on nuisance-model choice. It contrasts a Random-Forest-based DML estimator with a KAN-based D-IV-LATE, where “KAN” is defined as a Kernel-Augmented Nuisance estimator that fits each nuisance function via a combined parametric+nonparametric (RKHS) loss with a kernel penalty Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)2 chosen by CV. The paper argues that the selection of nuisance-function estimators is not a mere implementation detail but a pivotal choice that can profoundly impact research outcomes, a claim supported by a 401(k) application in SIPP 1991. There, the RF-DML estimate is negative for very low wealth thresholds, peaks near Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)3 at Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)4; the abstract reports that the KAN-based estimator suggests more complex treatment effect heterogeneity (Shaw, 15 Jun 2025).

3. Conditional LDTEs from distributional regression

A second major usage of LDTE is covariate-specific. In the GAMLSS formulation, one observes Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)5 and models

Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)6

with each parameter linked to an additive predictor of the form

Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)7

The Local Distributional Treatment Effect at Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)8 and Δ(y)=FY(1)C(y)FY(0)C(y)\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)9 is

Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)0

The same framework defines a local density treatment effect

Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)1

and a local quantile treatment effect

Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)2

Because the response family can be nonnormal and the covariate effects nonlinear, the approach explicitly targets location, scale, skewness, kurtosis, and zero-mass type features rather than only the conditional mean (Hohberg et al., 2018).

The estimation workflow is fully specified. One chooses a distribution family Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)3, chooses link functions Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)4, specifies additive predictors using terms such as P-splines, spatial terms, or random effects, estimates by maximum-likelihood or Bayesian methods, predicts Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)5 and Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)6, and then evaluates Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)7, Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)8, and Δ(yx)=FY(yT=1,X=x)FY(yT=0,X=x)\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)9 either by built-in cdf/pdf/quantile functions or by numerical inversion and integration. Inference is based on a parametric bootstrap with xx0–xx1, and model validation uses normalized quantile residuals (Dunn–Smyth residuals), Q–Q plots, and moments near xx2 (Hohberg et al., 2018).

Randomized experiments generate a related conditional distributional-regression formulation. There, for each treatment arm xx3 and each threshold xx4, the nuisance function is

xx5

and the Neyman-orthogonal score is

xx6

Cross-fitted regression adjustment yields xx7, and the local effect on a narrow interval is the “probability treatment effect”

xx8

Uniform inference is based on a multiplier bootstrap, and the paper reports substantial finite-sample variance reduction from machine learning adjustment, with Monte Carlo RMSE reductions and narrower confidence intervals in the Ferraro & Price water-nudge experiment (Byambadalai et al., 2024).

4. Kernel, forest, conformal, and finite-location localizations

Some LDTE formulations localize directly in outcome space through kernels. In Causal-DRF, the Local Distributional Treatment Effect is also called the Conditional Kernel Treatment Effect (CKTE): xx9 Here τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h)0, and τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h)1 is estimated by honest distributional random forests. Splits maximize an MMD-type statistic based on kernel means over a grid τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h)2, forest weights τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h)3 are computed from estimation leaves, and the estimator takes the doubly-robust-style form

τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h)4

Under overlap, smoothness, and honesty conditions, the method is pointwise consistent and asymptotically normal, with Wald intervals, multiplier-bootstrap uniform bands, and a conditional kernel-based test of τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h)5 for all τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h)6 (Näf et al., 2024).

Another local conditional formulation uses generative modeling and conformal inference. In the conformal-diffusion approach,

τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h)7

Each τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h)8 is estimated by Monte Carlo draws from a conditional diffusion model τloc(y0;h)\tau_{\mathrm{loc}}(y_0;h)9, and calibration points are reweighted by a local kernel

[y0,y0+h][y_0,y_0+h]0

together with inverse-propensity weights. Split-conformal prediction sets [y0,y0+h][y_0,y_0+h]1 yield lower and upper CDF envelopes [y0,y0+h][y_0,y_0+h]2 and [y0,y0+h][y_0,y_0+h]3, producing the LDTE band

[y0,y0+h][y_0,y_0+h]4

The coverage theorem states

[y0,y0+h][y_0,y_0+h]5

with convergence to [y0,y0+h][y_0,y_0+h]6 if [y0,y0+h][y_0,y_0+h]7 in [y0,y0+h][y_0,y_0+h]8 (Cai et al., 2024).

A related development does not estimate a full LDTE curve but localizes distributional discrepancies at learned outcome locations. DR-ME defines the interventional witness

[y0,y0+h][y_0,y_0+h]9

and for fixed locations τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]0 forms

τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]1

Orthogonal doubly robust features τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]2 lead, after cross-fitting, to a Hotelling statistic

τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]3

with τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]4 null calibration and noncentral τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]5 local alternatives. The method uses a three-way split for location learning and testing, thereby preserving post-selection validity (Zenati et al., 8 May 2026). This suggests a further notion of locality: interpretable coordinates in outcome space rather than an entire distributional-effect function.

5. Discontinuity-based LDTEs and distribution-valued outcomes

In regression discontinuity with distribution-valued outcomes, the target is neither a simple CDF difference nor a kernel-smoothed conditional mean. The R3D framework treats the outcome τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]6 as a random cdf in

τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]7

with running variable τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]8 and cutoff τh(x;y)=E[Kh(Yy)D=1,X=x]E[Kh(Yy)D=0,X=x]\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]9. The estimand is the local average quantile treatment effect

y0y_000

where

y0y_001

Two estimators are proposed: one-sided local polynomial regression on random quantiles y0y_002, and local Fréchet regression in y0y_003-space. Both admit uniform asymptotic normality over y0y_004, multiplier-bootstrap uniform confidence bands, and data-driven bandwidths of order y0y_005 (Dijcke, 4 Apr 2025).

A different discontinuity-based usage of local distributional effects studies sign changes in the Treatment Effects Curve (TEC)

y0y_006

with marginal effect

y0y_007

The 2026 toolkit focuses on treatment-effect discontinuities defined as points where the marginal distributional effects change sign. Horizontal Discontinuity Analysis (HDA) partitions the sample into sign-regions of y0y_008 and estimates CATEs with causal forests; Vertical Discontinuity Analysis (VDA) examines neighborhoods of crossings y0y_009 with RDD-style tools. The crossing-point estimator y0y_010 is constructed from adjacent pooled order statistics bracketing a sign reversal of the ECDF difference, and under isolated, non-tangential crossings satisfies

y0y_011

where y0y_012. The toolkit also provides a bias-corrected Wald test for non-tangentiality using kernel estimates of y0y_013 (Antognini et al., 26 Jun 2026).

These discontinuity-based formulations broaden the meaning of locality. In R3D, locality is at the cutoff y0y_014 and along the quantile index y0y_015. In the TEC-discontinuity toolkit, locality is attached to zero crossings of a distributional-effect curve and to neighborhoods around those crossings.

A substantial conceptual distinction separates LDTEs defined as differences between treated and control outcome distributions from objects defined on the treatment-effect distribution itself. In the latter approach, the target is

y0y_016

where y0y_017. Under RCT or unconfoundedness, covariate adjustment is combined with Makarov bounds. For a measurable y0y_018, define

y0y_019

and then

y0y_020

with

y0y_021

The sharp choices y0y_022 and y0y_023 maximize and minimize y0y_024 pointwise in y0y_025 (Fava, 2024).

Estimation is based either on a 50–50 sample split or on y0y_026-fold cross-fitting. Conditional CDFs y0y_027 can be estimated by random forest, sieve, lasso, neural net, or quantile regression methods. The paper names the implementation checklist “CAIDE” and derives finite-sample valid confidence intervals from the one-sided Dvoretzky-Kiefer-Wolfowitz inequality, as well as asymptotic normality and uniformly valid inference under continuity, consistency of the estimated y0y_028-functions, unique well-separated maximizers, and nondegenerate variances (Fava, 2024).

This object is not interchangeable with LDTEs of the form y0y_029. The former concerns the law of the latent treatment effect itself; the latter concerns how treatment shifts the outcome distribution at a threshold y0y_030. The distinction is empirically consequential. In the microcredit applications revisited by Fava, average treatment effects were statistically small, but the pooled bounds implied that at least y0y_031 of households benefited and at least y0y_032 were harmed (Fava, 2024).

Across the literature, three practical implications recur. First, cross-fitting or sample splitting is repeatedly treated as essential for guarding against over-fitting bias in orthogonal or partially identified procedures. Second, inference is typically functional rather than scalar, using multiplier bootstrap, parametric bootstrap, Wald bands, or DKW-based bounds. Third, model choice matters: the D-IV-LATE results explicitly warn that nuisance estimators can change substantive conclusions, especially where the first stage is weak or the outcome is extreme (Shaw, 15 Jun 2025). This suggests that “LDTE” is best understood not as one parameter but as a research program for causal analysis beyond the mean, with locality supplied by the identifying design and the chosen representation of the outcome distribution.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Local Distributional Treatment Effect.