---
title: Local Distributional Treatment Effect
url: https://www.emergentmind.com/topics/local-distributional-treatment-effect
type: topic
---

# Local Distributional Treatment Effect

Local Distributional Treatment Effect (LDTE) denotes a family of causal parameters that compare distributions rather than expectations on a localized margin. The locality can refer to the complier subpopulation under imperfect compliance, a fixed covariate value \(x\), a small interval around an outcome threshold \(y_0\), a kernel neighborhood in outcome space, or the limit at a regression-discontinuity cutoff. The term is therefore not tied to a single universal estimand. In instrumental-variables formulations, it is the difference in the treated and untreated potential-outcome cumulative distribution functions (CDFs) among compliers; in conditional distributional regression, it is often a covariate-specific difference of conditional CDFs; and in related work on the distribution of treatment effects, locality is attached instead to the conditional law of the individual treatment effect itself [2506.12765] [1806.09386] [2407.16037] [2407.14635].

## 1. Terminological scope and principal estimands

The main usages of LDTE in the recent literature differ along two axes: the object being localized and the design used for identification. Some papers localize in the population, as with compliers under imperfect compliance. Others localize in covariate space, outcome space, or design space, as in regression discontinuity. A related but distinct strand studies the distribution of the treatment effect \(\Delta_i=Y_i(1)-Y_i(0)\) rather than the distribution of outcomes under treatment and control [2509.15594] [2411.08778] [2504.03992] [2407.14635].

| Strand | Estimand | Locality notion |
|---|---|---|
| IV with imperfect compliance | \(\beta(y)=F_{Y(1)}(y\mid C)-F_{Y(0)}(y\mid C)\) [2509.15594] | complier subpopulation |
| D-IV-LATE | \(\Delta(y)=F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y)\) [2506.12765] | complier subpopulation |
| GAMLSS conditional LDTE | \(\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)\) [1806.09386] | fixed \(x\) |
| Randomized local effect | \(\tau_{\mathrm{loc}}(y_0;h)\) from CDF differences over \([y_0,y_0+h]\) [2407.16037] | local interval in outcome space |
| Causal-DRF / CKTE | \(\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x]\) [2411.08778] | fixed \(x\), kernel neighborhood around \(y\) |
| Distribution of treatment effects | \(F_{\Delta\mid X}(\tau\mid x)=P[\Delta_i\le \tau\mid X_i=x]\) [2407.14635] | fixed \(x\) on treatment-effect distribution |

A recurrent source of confusion is the presumption that “local” has a stable meaning across these formulations. The literature does not support that presumption. This suggests that any statement about an LDTE must specify the locality mechanism, the identifying assumptions, and whether the target is a difference of outcome distributions or the distribution of the treatment effect itself.

## 2. Complier-specific LDTE under imperfect compliance

In the instrumental-variables literature with binary instrument \(Z\) and binary treatment uptake, the LDTE is defined on the complier subpopulation. Let \(Y_i(w)\) be the potential outcome under treatment \(W=w\in\{0,1\}\), let \(W_i(z)\) be the potential treatment uptake under instrument \(Z_i=z\in\{0,1\}\), and impose monotonicity \(W_i(1)\ge W_i(0)\). The complier population is
\[
C=\{i:W_i(1)>W_i(0)\}.
\]
The Distributional IV-LATE at threshold \(y\) is then
\[
\Delta(y)=\mathbb{P}(Y_i(1)\le y\mid C)-\mathbb{P}(Y_i(0)\le y\mid C)
      =F_{Y(1)\mid C}(y)-F_{Y(0)\mid C}(y),
\]
with the same object denoted \(\beta(y)\) in the regression-adjusted imperfect-compliance framework [2506.12765] [2509.15594].

Identification proceeds through a Wald-like ratio of intent-to-treat differences. In the D-IV-LATE exposition, \(\Delta(y)\) is written as \(\alpha(y)/\beta\), where
\[
\alpha(y)=\mathbb{E}_X\!\left[\mathbb{E}[\mathbf{1}\{Y\le y\}\mid Z=1,X]-\mathbb{E}[\mathbf{1}\{Y\le y\}\mid Z=0,X]\right],
\]
and
\[
\beta=\mathbb{E}_X\!\left[\mathbb{E}[W\mid Z=1,X]-\mathbb{E}[W\mid Z=0,X]\right].
\]
In the covariate-adaptive randomization formulation, the same logic is written with strata \(S\) as
\[
\beta(y)=
\frac{
\sum_s p(s)\{E[1\{Y\le y\}\mid Z=1,S=s]-E[1\{Y\le y\}\mid Z=0,S=s]\}
}{
\sum_s p(s)\{E[D\mid Z=1,S=s]-E[D\mid Z=0,S=s]\}
}.
\]
The familiar simplified expression is
\[
\tau_{LDTE}(y)=\frac{F_{Y\mid Z=1}(y)-F_{Y\mid Z=0}(y)}{\Pr(D(1)>D(0))}.
\]
The underlying IV assumptions are random assignment conditional on strata, exclusion, and monotonicity [2509.15594].

Estimation is organized through Neyman-orthogonal moments and cross-fitting. In D-IV-LATE, the nuisance functions are
\[
\mu(y,w,x)=\mathbb{E}[\mathbf{1}\{Y\le y\}\mid W=w,X=x],\quad
p(z,x)=\mathbb{P}(W=1\mid Z=z,X=x),\quad
\pi(x)=\mathbb{P}(Z=1\mid X=x),
\]
with orthogonal score functions \(\psi_{\alpha,i}(y;\mu,p,\pi)\) and \(\psi_{\beta,i}(p,\pi)\). Cross-fitting trains nuisance estimators on \(K-1\) folds and evaluates scores on the held-out fold, and the final estimator is the ratio of average score terms. The regression-adjusted imperfect-compliance estimator uses augmented residuals \(\Xi^Y_{z,i}(y)\) and \(\Xi^D_{z,i}\), again with \(L\)-fold cross-fitting and flexible first-stage estimators such as logistic, random forest, and boosting [2506.12765] [2509.15594].

The asymptotic theory is correspondingly strong. Under regularity and \(o(n^{-1/4})\) nuisance rates, the D-IV-LATE estimator is asymptotically normal for each fixed \(y\), with variance
\[
V(y)=\frac{1}{\beta^2}\,
\mathbb{E}\!\left[
\bigl(\psi_{\alpha,i}(y)-\Delta(y)\psi_{\beta,i}\bigr)^2
\right].
\]
The imperfect-compliance paper further derives weak convergence in \(\ell^\infty(\mathcal{Y})\), states that the estimator achieves the semiparametric efficiency bound, and emphasizes applicability to continuous, discrete, and mixed discrete-continuous outcomes under stratified block designs and simple random sampling [2506.12765] [2509.15594].

A distinctive contribution of the D-IV-LATE paper is its emphasis on nuisance-model choice. It contrasts a Random-Forest-based DML estimator with a KAN-based D-IV-LATE, where “KAN” is defined as a Kernel-Augmented Nuisance estimator that fits each nuisance function via a combined parametric+nonparametric (RKHS) loss with a kernel penalty \(\lambda\|f\|_{\mathcal H}^2\) chosen by CV. The paper argues that the selection of nuisance-function estimators is not a mere implementation detail but a pivotal choice that can profoundly impact research outcomes, a claim supported by a 401(k) application in SIPP 1991. There, the RF-DML estimate is negative for very low wealth thresholds, peaks near \(0.9\) at \(y\approx \$18.5k\) for moderate wealth, and remains positive but declines to around \(0.4\) above \(y>\$120{,}000\); the abstract reports that the KAN-based estimator suggests more complex treatment effect heterogeneity [2506.12765].

## 3. Conditional LDTEs from distributional regression

A second major usage of LDTE is covariate-specific. In the GAMLSS formulation, one observes \(\{(Y_i,X_i,T_i)\}_{i=1}^n\) and models
\[
Y_i\mid X_i,T_i\sim D(\theta_i),\qquad \theta_i=(\mu_i,\sigma_i,\nu_i,\tau_i),
\]
with each parameter linked to an additive predictor of the form
\[
g_k(\theta_{ik})=\beta_{k0}+\beta_{k1}T_i+s_k(X_i).
\]
The Local Distributional Treatment Effect at \(y\) and \(x\) is
\[
\Delta(y\mid x)=F_Y(y\mid T=1,X=x)-F_Y(y\mid T=0,X=x)
              =F_D(y;\hat\theta(1;x))-F_D(y;\hat\theta(0;x)).
\]
The same framework defines a local density treatment effect
\[
\delta(y\mid x)=f_Y(y\mid 1,x)-f_Y(y\mid 0,x),
\]
and a local quantile treatment effect
\[
QTE(\tau\mid x)=Q_Y(\tau\mid 1,x)-Q_Y(\tau\mid 0,x).
\]
Because the response family can be nonnormal and the covariate effects nonlinear, the approach explicitly targets location, scale, skewness, kurtosis, and zero-mass type features rather than only the conditional mean [1806.09386].

The estimation workflow is fully specified. One chooses a distribution family \(D\), chooses link functions \(g_k\), specifies additive predictors using terms such as P-splines, spatial terms, or random effects, estimates by maximum-likelihood or Bayesian methods, predicts \(\hat\theta(1;x)\) and \(\hat\theta(0;x)\), and then evaluates \(\Delta(y\mid x)\), \(\delta(y\mid x)\), and \(QTE(\tau\mid x)\) either by built-in cdf/pdf/quantile functions or by numerical inversion and integration. Inference is based on a parametric bootstrap with \(B\approx 500\)–\(1{,}000\), and model validation uses normalized quantile residuals (Dunn–Smyth residuals), Q–Q plots, and moments near \(N(0,1)\) [1806.09386].

Randomized experiments generate a related conditional distributional-regression formulation. There, for each treatment arm \(w\in\{0,1\}\) and each threshold \(y\), the nuisance function is
\[
\gamma_y^{(w)}(x)=E[1\{Y\le y\}\mid W=w,X=x],
\]
and the Neyman-orthogonal score is
\[
\psi_y^{(w)}(Z;\theta_y,\gamma_y)
=
\frac{\mathbf{1}\{W=w\}}{\pi_w}\,[\mathbf{1}\{Y\le y\}-\gamma_y^{(w)}(X)]
+\gamma_y^{(w)}(X)-\theta_y^{(w)}.
\]
Cross-fitted regression adjustment yields \(\hat F_w(y)\), and the local effect on a narrow interval is the “probability treatment effect”
\[
\hat\tau_{\mathrm{loc}}(y_0;h)
=
[\hat F_1(y_0+h)-\hat F_1(y_0)]-[\hat F_0(y_0+h)-\hat F_0(y_0)].
\]
Uniform inference is based on a multiplier bootstrap, and the paper reports substantial finite-sample variance reduction from machine learning adjustment, with Monte Carlo RMSE reductions and narrower confidence intervals in the Ferraro & Price water-nudge experiment [2407.16037].

## 4. Kernel, forest, conformal, and finite-location localizations

Some LDTE formulations localize directly in outcome space through kernels. In Causal-DRF, the Local Distributional Treatment Effect is also called the Conditional Kernel Treatment Effect (CKTE):
\[
\tau_h(x;y)=E[K_h(Y-y)\mid D=1,X=x]-E[K_h(Y-y)\mid D=0,X=x].
\]
Here \(K_h(u)=K(u/h)/h\), and \(\tau_h(x;y)\) is estimated by honest distributional random forests. Splits maximize an MMD-type statistic based on kernel means over a grid \(y_1<\cdots<y_J\), forest weights \(\alpha_i(x)\) are computed from estimation leaves, and the estimator takes the doubly-robust-style form
\[
\hat\tau_h(x;y)
=
\sum_{i=1}^n
\alpha_i(x)\,K_h(Y_i-y)\,
\frac{D_i-\hat e(X_i)}{\hat e(X_i)(1-\hat e(X_i))}.
\]
Under overlap, smoothness, and honesty conditions, the method is pointwise consistent and asymptotically normal, with Wald intervals, multiplier-bootstrap uniform bands, and a conditional kernel-based test of \(H_0:\tau_h(x;y)=0\) for all \(y\) [2411.08778].

Another local conditional formulation uses generative modeling and conformal inference. In the conformal-diffusion approach,
\[
LDTE(x;y)=F_1(x;y)-F_0(x;y),\qquad
F_t(x;y)=\mathbb{P}(Y(t)\le y\mid X=x).
\]
Each \(F_t\) is estimated by Monte Carlo draws from a conditional diffusion model \(q_t^\theta(Y\mid X)\), and calibration points are reweighted by a local kernel
\[
H(x,x')=\exp(-\|x-x'\|^2/(2h^2))
\]
together with inverse-propensity weights. Split-conformal prediction sets \( \hat C_t(X_k)\) yield lower and upper CDF envelopes \(L_t(x_k;y)\) and \(U_t(x_k;y)\), producing the LDTE band
\[
[\,L_1(x_k;y)-U_0(x_k;y),\;U_1(x_k;y)-L_0(x_k;y)\,].
\]
The coverage theorem states
\[
\mathbb{P}\{Y_k(1)\in \hat C_1(X_k)\}\ge 1-\alpha-\tfrac12 E_X|\hat w(X)-w(X)|,
\]
with convergence to \(1-\alpha\) if \(\hat w\to w\) in \(L^1\) [2408.01582].

A related development does not estimate a full LDTE curve but localizes distributional discrepancies at learned outcome locations. DR-ME defines the interventional witness
\[
w(y)=\mathbb{E}[k_Y(y,Y(1))]-\mathbb{E}[k_Y(y,Y(0))],
\]
and for fixed locations \(V=(v_1,\dots,v_J)\) forms
\[
\mu_V=(w(v_1),\dots,w(v_J))^\top.
\]
Orthogonal doubly robust features \(z_V^{dr}(Z;\eta)\) lead, after cross-fitting, to a Hotelling statistic
\[
\hat\lambda_{n,V}^{dr}
=
n\,\bar z_{n,V}^{dr\top}
(S_{n,V}^{dr}+\gamma_n I_J)^{-1}
\bar z_{n,V}^{dr},
\]
with \(\chi^2_J\) null calibration and noncentral \(\chi^2_J\) local alternatives. The method uses a three-way split for location learning and testing, thereby preserving post-selection validity [2605.08034]. This suggests a further notion of locality: interpretable coordinates in outcome space rather than an entire distributional-effect function.

## 5. Discontinuity-based LDTEs and distribution-valued outcomes

In regression discontinuity with distribution-valued outcomes, the target is neither a simple CDF difference nor a kernel-smoothed conditional mean. The R3D framework treats the outcome \(Y\) as a random cdf in
\[
\mathcal Y=\{G:\mathbb R\to[0,1]\mid G\text{ nondecreasing},\ \int x^2\,dG(x)<\infty\},
\]
with running variable \(X\) and cutoff \(c\). The estimand is the local average quantile treatment effect
\[
\tau^{\rm R3D}(q)
=
E[Q_{Y^1}(q)-Q_{Y^0}(q)\mid X=c]
=
m_+(q)-m_-(q),
\]
where
\[
m_\pm(q)=\lim_{x\to c^\pm}E[Q_Y(q)\mid X=x].
\]
Two estimators are proposed: one-sided local polynomial regression on random quantiles \(Q_{Y_i}(q)\), and local Fréchet regression in \(W_2\)-space. Both admit uniform asymptotic normality over \(q\in[a,b]\), multiplier-bootstrap uniform confidence bands, and data-driven bandwidths of order \(n^{-1/(2p+3)}\) [2504.03992].

A different discontinuity-based usage of local distributional effects studies sign changes in the Treatment Effects Curve (TEC)
\[
\Delta(y)=F_A(y)-F_B(y),
\]
with marginal effect
\[
\delta(y)=\frac{d\Delta(y)}{dy}=f_A(y)-f_B(y).
\]
The 2026 toolkit focuses on treatment-effect discontinuities defined as points where the marginal distributional effects change sign. Horizontal Discontinuity Analysis (HDA) partitions the sample into sign-regions of \(\Delta(y)\) and estimates CATEs with causal forests; Vertical Discontinuity Analysis (VDA) examines neighborhoods of crossings \(\Delta(y_0)=0\) with RDD-style tools. The crossing-point estimator \(\hat y_0\) is constructed from adjacent pooled order statistics bracketing a sign reversal of the ECDF difference, and under isolated, non-tangential crossings satisfies
\[
a_n^{-1/2}(\hat y_0-y_0)\to N\!\left(0,\frac{p_0(1-p_0)}{(\Delta'(y_0))^2}\right),
\]
where \(a_n=1/n_A+1/n_B\). The toolkit also provides a bias-corrected Wald test for non-tangentiality using kernel estimates of \(\Delta'(y_0)\) [2606.28017].

These discontinuity-based formulations broaden the meaning of locality. In R3D, locality is at the cutoff \(X=c\) and along the quantile index \(q\). In the TEC-discontinuity toolkit, locality is attached to zero crossings of a distributional-effect curve and to neighborhoods around those crossings.

## 6. Related but distinct object: the distribution of treatment effects

A substantial conceptual distinction separates LDTEs defined as differences between treated and control outcome distributions from objects defined on the treatment-effect distribution itself. In the latter approach, the target is
\[
F_\Delta(\tau)=P[\Delta_i\le \tau],\qquad
F_{\Delta\mid X}(\tau\mid x)=P[\Delta_i\le \tau\mid X_i=x],
\]
where \(\Delta_i=Y_i(1)-Y_i(0)\). Under RCT or unconfoundedness, covariate adjustment is combined with Makarov bounds. For a measurable \(s:\mathbb R^p\to\mathbb R\), define
\[
F_{j,s}(t)=P[Y(j)-s(X)\le t],\qquad j=0,1,
\]
and then
\[
F_\Delta(\tau)\in[\theta_L(\tau,s),\theta_U(\tau,s)],
\]
with
\[
\theta_L(\tau,s)=\sup_t\{F_{1,s}(t+\tau)-F_{0,s}(t)\},\qquad
\theta_U(\tau,s)=1+\inf_t\{F_{1,s}(t+\tau)-F_{0,s}(t)\}.
\]
The sharp choices \(s_L^*(x)\) and \(s_U^*(x)\) maximize and minimize \(F_1(t\mid x)-F_0(t\mid x)\) pointwise in \(x\) [2407.14635].

Estimation is based either on a 50–50 sample split or on \(K\)-fold cross-fitting. Conditional CDFs \(\widehat F_j(t\mid x)\) can be estimated by random forest, sieve, lasso, neural net, or quantile regression methods. The paper names the implementation checklist “CAIDE” and derives finite-sample valid confidence intervals from the one-sided Dvoretzky-Kiefer-Wolfowitz inequality, as well as asymptotic normality and uniformly valid inference under continuity, consistency of the estimated \(s\)-functions, unique well-separated maximizers, and nondegenerate variances [2407.14635].

This object is not interchangeable with LDTEs of the form \(F_{Y(1)\mid\cdot}(y)-F_{Y(0)\mid\cdot}(y)\). The former concerns the law of the latent treatment effect itself; the latter concerns how treatment shifts the outcome distribution at a threshold \(y\). The distinction is empirically consequential. In the microcredit applications revisited by Fava, average treatment effects were statistically small, but the pooled bounds implied that at least \(13.6\%\) of households benefited and at least \(12.5\%\) were harmed [2407.14635].

Across the literature, three practical implications recur. First, cross-fitting or sample splitting is repeatedly treated as essential for guarding against over-fitting bias in orthogonal or partially identified procedures. Second, inference is typically functional rather than scalar, using multiplier bootstrap, parametric bootstrap, Wald bands, or DKW-based bounds. Third, model choice matters: the D-IV-LATE results explicitly warn that nuisance estimators can change substantive conclusions, especially where the first stage is weak or the outcome is extreme [2506.12765]. This suggests that “LDTE” is best understood not as one parameter but as a research program for causal analysis beyond the mean, with locality supplied by the identifying design and the chosen representation of the outcome distribution.

Source: https://www.emergentmind.com/topics/local-distributional-treatment-effect