---
title: Average Length Treatment Effect (ALTE)
url: https://www.emergentmind.com/topics/average-length-treatment-effect-alte
type: topic
---

# Average Length Treatment Effect (ALTE)

Average Length Treatment Effect (ALTE) is a context-dependent causal estimand. In the most explicit recent usage, it is defined for circular outcomes as the difference between the mean resultant lengths of treated and untreated potential outcomes, so it measures a treatment-induced change in concentration rather than a change in linear time or Euclidean distance. In other settings, the same label has been used outcome-agnostically for \(E[L(1)]-E[L(0)]\) when the scalar outcome \(L\) is itself a length or duration, and related literatures use closely neighboring acronyms for long-term or local average treatment effects. This suggests that ALTE must be interpreted relative to the outcome geometry and identification framework being used [2507.19889; 1311.0291; 2411.04380].

## 1. Terminological scope and conceptual status

In "Causal Inference for Circular Data" [2507.19889], ALTE is introduced as a treatment effect for circular random variables, alongside the average direction treatment effect (ADTE). There, “average length” refers to the mean resultant length of a circular distribution, not to elapsed time, physical distance, or any ordinary linear scale. The estimand is therefore a causal contrast in concentration.

A different usage appears when a conventional scalar outcome happens to be a length or duration. In the framework of "Improved Precision in Estimating Average Treatment Effects" [1311.0291], the underlying estimand is the ordinary average treatment effect,
\[
\tau = E[Y(1)] - E[Y(0)],
\]
and if \(Y=L\) is a length outcome—such as length of hospital stay, duration of unemployment, or another continuous time-like variable—then, in that terminology, the estimand becomes an ALTE. In that usage, ALTE is not a new identification concept; it is the standard ATE applied to a specific scalar response.

A further usage occurs in work on long-term effects, where weighted averages of conditional long-term treatment effects are described as the usual long-term average treatment effect, abbreviated there as LTE/ALTE [2411.04380]. The coexistence of these definitions indicates that ALTE is not a universally fixed label across causal inference. A plausible implication is that any technical discussion of ALTE must specify whether the outcome is circular, scalar duration-valued, exposure-time indexed, or part of a long-term treatment effect framework.

## 2. ALTE for circular outcomes

For a circular random variable \(\Theta \in [0,2\pi)\), the first trigonometric moment is
\[
\phi \triangleq E\{\exp(i\Theta)\} = \alpha + i\beta,
\]
with
\[
\alpha \triangleq E(\cos\Theta), \qquad \beta \triangleq E(\sin\Theta).
\]
The associated mean resultant length and mean resultant direction are
\[
\rho \triangleq (\alpha^2+\beta^2)^{1/2}, \qquad \mu \triangleq \text{atan2}(\beta,\alpha).
\]
Thus the mean resultant vector has direction \(\mu\) and length \(\rho\) [2507.19889].

Under the potential-outcomes formulation, \(A \in \{0,1\}\) is a binary treatment, \(\mathbf{X}\) are covariates, and \(\Theta^{(1)}\) and \(\Theta^{(0)}\) are the potential circular outcomes under treatment and control. For \(a \in \{0,1\}\),
\[
\alpha^{(a)} = E(\cos\Theta^{(a)}), \qquad \beta^{(a)} = E(\sin\Theta^{(a)}),
\]
\[
\rho^{(a)} = \{(\alpha^{(a)})^2+(\beta^{(a)})^2\}^{1/2}, \qquad \mu^{(a)} = \text{atan2}(\beta^{(a)},\alpha^{(a)}).
\]

The circular-data ALTE is then defined as
\[
\boxed{
\xi \triangleq \{(\alpha^{(1)})^2 +(\beta^{(1)})^2\}^{1/2}
-
\{(\alpha^{(0)})^2 +(\beta^{(0)})^2\}^{1/2}
=
\rho^{(1)} - \rho^{(0)}.
}
\]
This is the difference in mean resultant lengths of the two potential-outcome distributions [2507.19889].

The interpretation is geometric. Each observation can be represented as a unit vector \((\cos\Theta,\sin\Theta)\), and the population mean resultant vector is their vector average. Its direction \(\mu\) captures location on the circle, whereas its length \(\rho \in [0,1]\) measures concentration. When \(\rho \approx 1\), the angles are tightly clustered; when \(\rho \approx 0\), they are nearly dispersed around the circle. Consequently,
\[
\xi = \rho^{(1)}-\rho^{(0)}
\]
measures how treatment changes concentration. If \(\xi>0\), treatment makes circular outcomes more tightly clustered; if \(\xi<0\), it makes them more dispersed; if \(\xi=0\), concentration is unchanged even if the mean direction shifts [2507.19889].

A central misconception addressed by this formulation is that “average length” might be taken literally as duration or magnitude in the usual Euclidean sense. For circular outcomes, average length is the length of the population resultant vector. The estimand is therefore a causal effect on concentration, and its meaning depends on circular topology rather than on linear averaging [2507.19889].

## 3. Identification and estimation under inverse probability weighting

The circular-data ALTE is identified under the same causal assumptions used for standard average treatment effects, adapted to circular outcomes. The paper states: consistency,
\[
\Theta = A\Theta^{(1)} + (1-A)\Theta^{(0)};
\]
strong ignorability or unconfoundedness,
\[
A \perp (\Theta^{(1)},\Theta^{(0)}) \mid \mathbf{X};
\]
positivity,
\[
0 < \pi(\mathbf{x}) < 1;
\]
IID sampling; bounded covariates; and circular regularity conditions ensuring that the mean direction and mean resultant length are well-defined [2507.19889].

With propensity score
\[
\pi(\mathbf{x}) \triangleq P(A=1\mid \mathbf{X}=\mathbf{x}),
\]
the paper uses inverse probability weighting through the identities
\[
E\bigg(\frac{A}{\pi(\mathbf{X})}g(\Theta)k(\mathbf{X})\bigg)
=
E\big(g(\Theta^{(1)})k(\mathbf{X})\big),
\]
\[
E\bigg(\frac{1-A}{1-\pi(\mathbf{X})}g(\Theta)k(\mathbf{X})\bigg)
=
E\big(g(\Theta^{(0)})k(\mathbf{X})\big).
\]
Taking \(g(\Theta)=\cos\Theta\) or \(\sin\Theta\) and \(k(\mathbf{X})=1\) identifies \(\alpha^{(a)}\) and \(\beta^{(a)}\), and since \(\xi\) is a deterministic function of \((\alpha^{(1)},\beta^{(1)},\alpha^{(0)},\beta^{(0)})\), ALTE is identified [2507.19889].

Estimation proceeds by modeling treatment assignment with logistic regression,
\[
\pi(\mathbf{x}) = \frac{\exp(\mathbf{x}^\top\boldsymbol{\eta})}{1+\exp(\mathbf{x}^\top\boldsymbol{\eta})},
\]
estimating \(\widehat{\pi}(\mathbf{x})\), and constructing either Horvitz–Thompson (HT) or Hájek weights. The HT weights are
\[
w_{i,1}^{(1)} = \frac{A_i}{n\widehat{\pi}(\mathbf{X}_i)}, \qquad
w_{i,1}^{(0)} = \frac{1-A_i}{n\{1-\widehat{\pi}(\mathbf{X}_i)\}}.
\]
The HT estimators of the cosine and sine moments are
\[
\widehat{\alpha}^{(a)} = \sum_{i=1}^n w_{i,1}^{(a)}\cos\Theta_i, \qquad
\widehat{\beta}^{(a)} = \sum_{i=1}^n w_{i,1}^{(a)}\sin\Theta_i,
\]
and the HT ALTE estimator is the plug-in quantity
\[
\widehat{\xi}
=
\big[(\widehat{\alpha}^{(1)})^2+(\widehat{\beta}^{(1)})^2\big]^{1/2}
-
\big[(\widehat{\alpha}^{(0)})^2+(\widehat{\beta}^{(0)})^2\big]^{1/2}.
\]

The Hájek version normalizes the weights within treatment arm,
\[
w_{i,2}^{(a)}
=
\bigg(\sum_{j=1}^n w_{j,1}^{(a)}\bigg)^{-1} w_{i,1}^{(a)},
\]
then defines
\[
\widetilde{\alpha}^{(a)} = \sum_{i=1}^n w_{i,2}^{(a)} \cos\Theta_i, \qquad
\widetilde{\beta}^{(a)} = \sum_{i=1}^n w_{i,2}^{(a)} \sin\Theta_i,
\]
and
\[
\widetilde{\xi}
=
\big[(\widetilde{\alpha}^{(1)})^2+(\widetilde{\beta}^{(1)})^2\big]^{1/2}
-
\big[(\widetilde{\alpha}^{(0)})^2+(\widetilde{\beta}^{(0)})^2\big]^{1/2}.
\]
The estimation strategy is entirely IPW-based: there is no outcome regression and no doubly robust augmentation [2507.19889].

## 4. Large-sample theory, simulation behavior, and empirical illustration

The paper treats ALTE as the second component of \(\boldsymbol{\Delta}=(\tau,\xi)\), where \(\tau\) is ADTE and \(\xi\) is ALTE. Letting \(\boldsymbol{\omega}=(\alpha^{(1)},\beta^{(1)},\alpha^{(0)},\beta^{(0)})^\top\), the delta-method Jacobian maps asymptotic distributions for \(\widehat{\boldsymbol{\omega}}\) or \(\widetilde{\boldsymbol{\omega}}\) into asymptotic distributions for \((\widehat{\tau},\widehat{\xi})\) and \((\widetilde{\tau},\widetilde{\xi})\). The main result is
\[
n^{1/2}(\widehat{\boldsymbol{\Delta}}-\boldsymbol{\Delta})
\overset{d}{\longrightarrow}
\textup{N}(\mathbf{0}_{2},\boldsymbol{\Sigma}_{\textup{HT}}),
\]
\[
n^{1/2}(\widetilde{\boldsymbol{\Delta}}-\boldsymbol{\Delta})
\overset{d}{\longrightarrow}
\textup{N}(\mathbf{0}_{2},\boldsymbol{\Sigma}_{\textup{Hajek}}),
\]
with sandwich-form covariance matrices; the asymptotic variance of the ALTE estimator is the \((2,2)\) entry in the relevant \(2\times 2\) matrix [2507.19889].

Under correct specification of the logistic propensity model and assumptions (C1–C10), the nuisance estimators for cosine and sine moments are consistent, and the plug-in estimators \(\widehat{\xi}\) and \(\widetilde{\xi}\) are consistent for \(\xi\). The paper does not claim a semiparametric efficiency result; it establishes consistency, asymptotic normality, and near-nominal coverage in simulation [2507.19889].

The simulation study fixes the true ALTE at
\[
\xi = 1/6 \approx 0.1667
\]
across three scenarios, with \(p=3\), \(n=250,500,1000\), and 1000 datasets per setting. For \(n=250\), HT ALTE has biases around \(-0.028\) to \(-0.061\), standard error around \(0.18\)–\(0.22\), mean squared error about \(0.03\)–\(0.05\), and coverage rate approximately \(0.94\)–\(0.95\). Hájek ALTE has slightly larger magnitude bias but smaller standard error, often slightly smaller mean squared error, and coverage often somewhat higher. As \(n\) increases to \(500\) and \(1000\), bias shrinks, standard errors decrease, and the distinction between HT and Hájek becomes less consequential. The paper summarizes the pattern as a small-sample bias–variance trade-off: HT tends to be less biased but more variable, while Hájek is more stable but relatively more biased; when \(n\) is sufficiently large, the trade-off becomes negligible [2507.19889].

The empirical application uses Work Schedules and Sleep Pattern Survey Data of railroad dispatchers, with 439 participants. Treatment is job type: \(A=1\) for assistant chief dispatcher and \(A=0\) for trick dispatcher. The outcome \(\Theta\) is time fell asleep, converted to radians so that \(2\pi\) radians correspond to 24 hours. The estimated ALTEs are
\[
\widehat{\xi}_{\text{HT}} = 0.445, \qquad
\widetilde{\xi}_{\text{Hajek}} = 0.184.
\]
Because both are positive, the paper concludes that assistant chief dispatchers exhibit more similarity in their falling-asleep times than trick dispatchers. The associated ADTE estimate is negative, indicating later sleep timing for assistant chief dispatchers, so the application separates a location shift from a concentration shift [2507.19889].

## 5. Outcome-agnostic and exposure-time interpretations

When the observed response is an ordinary scalar length or duration, ALTE can be understood as a direct specialization of the ATE. In the random-\(X\) framework of "Improved Precision in Estimating Average Treatment Effects" [1311.0291], if \(L\) denotes a length outcome, then
\[
\text{ALTE} = E[L(1)] - E[L(0)].
\]
The paper’s regression-based estimator applies verbatim after replacing \(Y\) by \(L\). Covariates are centered at the pooled mean,
\[
X_i^* = X_i - \bar{X}_{\text{pool}},
\]
and one fits the interacted model
\[
L_i = \beta^{(0)} + \beta^{(T)} D_i + X_i^{*'}\gamma + X_i^{*'}\delta D_i + \varepsilon_i.
\]
The coefficient on \(D_i\) is then
\[
\hat{\text{ALTE}}_{\text{reg}} = \hat{\beta}^{(T)} = \hat{\tau}_{\text{reg}}.
\]
Under mild random-\(X\) assumptions, the estimator is asymptotically unbiased, and its asymptotic variance is no larger than that of the difference-in-means estimator, with equality iff
\[
\beta_C = -\frac{n_C}{n_T}\beta_T.
\]
This formulation treats ALTE as entirely outcome-agnostic: any scalar outcome can be inserted, including a duration or length [1311.0291].

A distinct but related interpretation appears in stepped wedge cluster randomized trials with time-varying treatment effects. There, treatment effect is indexed by exposure time \(s\) through an effect curve \(\delta(s)\), and the time-averaged treatment effect over an interval \([s_1,s_2]\) is
\[
\Psi_{[s_1,s_2]} \equiv \frac{1}{s_2-s_1}\int_{s_1}^{s_2}\delta(s)\,ds.
\]
The framework explicitly states that an “Average Length Treatment Effect (ALTE)” is not named in the paper, but everything needed to define and estimate such an average-over-time estimand is present. This suggests the natural definitions
\[
\text{ALTE}_L = \frac{1}{L}\int_0^L \delta(s)\,ds
\]
or, in discrete exposure time,
\[
\text{ALTE}_L = \frac{1}{L}\sum_{t=1}^{L}\delta(t).
\]
The same paper shows that the immediate-treatment estimator can be misleading because its expectation is a weighted sum of the point treatment effects, the weights sum to one, and some weights can be negative. As a result, the estimator can even converge to a value of the opposite sign of the true time-averaged treatment effect or long-term treatment effect [2111.07190].

These two usages share a common structural feature: ALTE is an average causal contrast, but the averaging operation is taken over different objects. In the scalar-duration usage, the average is over units; in the stepped-wedge exposure-time usage, it is over exposure time along an effect curve.

## 6. Related estimands, neighboring acronyms, and sources of ambiguity

Long-term treatment effect work introduces yet another use of the acronym. In "Identification of Long-Term Treatment Effects via Temporal Links, Observational, and Experimental Data" [2411.04380], the target parameter is
\[
\tau = \mathbb{E}[Y(1)-Y(0)],
\]
and the paper notes that weighted averages of the conditional long-term treatment effect \(\tau(x)\) give the usual long-term average treatment effect, denoted there as LTE/ALTE. The core representation is
\[
\text{LTE} = \tau
= \int_{\mathcal{S}} m_1(s)\,d\gamma_1(s)
-
\int_{\mathcal{S}} m_0(s)\,d\gamma_0(s),
\]
where \(m_d(s)=\mathbb{E}[Y(d)\mid S(d)=s]\) are temporal link functions and \(\gamma_d\) are distributions of short-term potential outcomes. The paper’s main claim is that experimental data have no identifying power for LTE without additional modeling assumptions; they only amplify the identifying power of assumptions such as latent unconfoundedness (LUC), latent monotone instrumental variable (LIV), or treatment invariance (TI). Under no modeling assumptions, if \(\mathcal{Y}=\mathbb{R}\), the identified set for the LTE is \(\mathbb{R}\), and if \(\mathcal{Y}=[0,1]\), the bounds reduce to Manski’s worst-case bounds [2411.04380].

Closely related but distinct terminology appears in the local average treatment effect literature. In randomized experiments with noncompliance, the target parameter is the local average treatment effect among compliers,
\[
LATE = \frac{1}{|\mathcal{C}|}\sum_{i\in\mathcal{C}} \big(y_i(1)-y_i(0)\big),
\]
identified by the Wald ratio under exclusion and monotonicity [2404.18786]. Randomization-based confidence sets constructed from a studentized Anderson–Rubin-type statistic are finite-sample exact under treatment effect homogeneity and asymptotically valid for heterogeneous LATE, even with weak instruments [2404.18786]. In mixture-model formulations of non-adherence in randomized trials, the same complier effect appears as the CACE/LATE, and substantive model compatible multiple imputation of latent compliance class has been proposed as an alternative to TSLS or full Bayesian estimation, especially for binary outcomes [1812.01322].

The main encyclopedic point is therefore taxonomic. ALTE can denote a formally defined circular-data estimand \(\xi=\rho^{(1)}-\rho^{(0)}\); an outcome-specific instance of the ordinary ATE for a scalar length variable; an average-over-exposure-length functional in stepped wedge designs; or, in some long-term effect settings, a long-term average treatment effect. A plausible implication is that the expression is best treated as a family resemblance term rather than a single canonical parameter. Precision requires stating the outcome space, the averaging domain, and the identification strategy before any ALTE estimate can be interpreted.

Source: https://www.emergentmind.com/topics/average-length-treatment-effect-alte