---
title: Energy Distance-Based Location Test
url: https://www.emergentmind.com/topics/energy-distance-based-location-test
type: topic
---

# Energy Distance-Based Location Test

Searching arXiv for the cited papers to ground the article in current references.
arxiv_search.query({"search_query":"id:2505.20647 OR id:1703.07856 OR id:1908.06892 OR id:1901.00833","max_results":10,"sort_by":"relevance","sort_order":"descending"})
An energy distance-based location test is a procedure built from the energy distance between two distributions, typically using Euclidean pairwise distances between observations. In its classical form, the energy distance is an omnibus discrepancy: it is zero if and only if the two distributions are equal. Yet recent analysis shows that, when the two distributions are close relative to a common scale, the leading local sensitivity of the energy distance is to differences in means rather than to differences in covariances. In that precise perturbative regime, an energy-distance-based two-sample procedure behaves approximately like a location test, even though its formal null hypothesis remains equality of distributions rather than equality of means alone [2505.20647].

## 1. Definition and inferential target

For random vectors \(X,Y\in\mathbb{R}^d\), one normalization studied in the recent local theory is
\[
(X,Y):=\mathbb{E}\|X-Y\|-\frac12\mathbb{E}\|X-X'\|-\frac12\mathbb{E}\|Y-Y'\|,
\]
where \(X'\) and \(Y'\) are independent copies and \(\|\cdot\|\) is the Euclidean norm. In the two-sample testing literature, an equivalent doubled normalization is also standard:
\[
\varepsilon(Q_1,Q_2)=2E\|X_1-Y_1\|-E\|X_1-X_2\|-E\|Y_1-Y_2\|.
\]
Both formulations encode the same inferential principle: the discrepancy vanishes if and only if the two distributions coincide [2505.20647; 1703.07856].

This point is central to the interpretation of the phrase “location test.” Energy distance is not intrinsically a location-only device. The null is a homogeneity hypothesis,
\[
H_0:Q_1=Q_2 \qquad \text{vs.} \qquad H_1:Q_1\neq Q_2,
\]
so the method can respond to location shifts, scale changes, changes in spread, multimodality, and other distributional differences. The designation “location test” is therefore exact only in restricted settings, or approximately in the local regime where the leading term is governed by mean differences [1703.07856].

A closely related perspective arises in censored survival testing. There, the \(\alpha\)-energy distance is
\[
\epsilon_\alpha(P,Q)= 2E||X-Y||^{\alpha}- E||X-X^{'}||^{\alpha}-E||Y-Y^{'}||^{\alpha}.
\]
For \(\alpha\in(0,2)\), this is an omnibus discrepancy, whereas for \(\alpha=2\),
\[
\epsilon_2(P,Q)= 2||E(X)-E(Y)||^{2},
\]
so the criterion becomes a mean/location discrepancy rather than a full distributional one [1901.00833].

## 2. Local moment expansion and the emergence of location sensitivity

The most explicit justification for the “location test” interpretation comes from the local expansion of the energy distance when \(X\) and \(Y\) are close relative to a common scale \(\lambda\). The expansion is organized through the mean difference
\[
\mu=\mathbb{E}X-\mathbb{E}Y,
\]
the covariance difference
\[
\Delta_{mn}=\operatorname{Cov}(X_n,X_m)-\operatorname{Cov}(Y_n,Y_m),
\]
and the third-cumulant difference
\[
\kappa_{mn\ell}.
\]
Under the paper’s perturbative assumptions, including \(\Phi(\omega)=H(\lambda\omega)\) with rapidly decaying \(H\), Proposition 1 gives the asymptotic structure
\[
(X,Y)\approx C_1\frac{m_1(\mu)}{\lambda}+C_2\frac{m_2(\Delta)}{\lambda^3},
\]
with the leading contribution depending only on \(\mu\) and covariance entering only at the next nonzero order [2505.20647].

In the nearly spherical case, the expansion simplifies to
\[
(X, Y) =\frac{1}{\lambda} \frac{|S^{d-1}|}{c_d d} \|\mu\|^2 \int_0^\infty h(r)\,dr
+ \frac{1}{\lambda^3}\frac{|S^{d-1}|}{4 c_d d(d+2)} \left[ 2\|\Delta\|_F^2 + \operatorname{Trace}(\Delta)^2 - \|\mu\|^4 - 4\beta\cdot\mu \right]\int_0^\infty h(r)r^2\,dr
+ R(\lambda).
\]
The leading term is therefore exactly proportional to \(\|\mu\|^2/\lambda\), whereas covariance contributes only at order \(\lambda^{-3}\) [2505.20647].

The structural reason is that mean differences appear in \(X-Y\), but not in \(X-X'\) or \(Y-Y'\), because centering by independent copies eliminates the mean there. In the Taylor analysis this becomes
\[
\psi_1(\theta)=\psi_3(\theta)\equiv 0,\qquad \psi_2(\theta)=2(\theta\cdot\mu)^2,
\]
while
\[
\psi_4(\theta)=6(\theta\cdot\Delta\theta)^2-2(\theta\cdot\mu)^4-8(\theta\cdot\mu)\sum_{ijk=1}^d\kappa_{ijk}\theta_i\theta_j\theta_k.
\]
Thus odd orders vanish, the first nonzero term is second order and purely mean-based, and covariance enters only at fourth order [2505.20647].

This local asymptotic picture supports a precise reformulation: if \(P_X\) and \(P_Y\) differ only slightly relative to a common scale, then rejecting for large energy distance will predominantly detect differences in means. The resulting procedure behaves approximately like a test of
\[
H_0:\mathbb{E}X=\mathbb{E}Y,
\]
but only in that local sense. Globally, the energy distance remains omnibus because \((X,Y)=0\iff X\sim Y\) [2505.20647].

## 3. Covariance sensitivity, dimensional effects, and high-dimensional attenuation

Although energy distance can detect covariance differences, the local theory shows that such effects are suppressed relative to mean shifts. In the spherical case, the covariance contribution enters through
\[
2\|\Delta\|_F^2+\operatorname{Trace}(\Delta)^2,
\]
multiplied by \(\lambda^{-3}\). The separation in order,
\[
\text{mean contribution } \sim \lambda^{-1}, \qquad \text{covariance contribution } \sim \lambda^{-3},
\]
does not depend on dimension. Dimension affects constants and averaging, but not the fact that mean differences enter two orders earlier than covariance differences [2505.20647].

The theory also distinguishes diagonal from off-diagonal covariance perturbations. Near isotropy, diagonal variance changes can matter much more than off-diagonal correlation changes because the trace term heavily rewards diagonal perturbations. In the paper’s \(M\)-dependent example, the off-diagonal contribution
\[
\frac{M\rho^4}{2d}
\]
is interpreted as \(O(1/d)\) smaller than the leading diagonal contribution
\[
\frac{\delta^4}{8}.
\]
This shows that local correlation changes may be strongly attenuated compared with variance changes in high dimension [2505.20647].

A gradient-level comparison sharpens the same point. With
\[
L:=2\|\Delta\|_F^2+\operatorname{Trace}(\Delta)^2,\qquad I:=\|\Delta\|_F^2,
\]
the cosine similarity of their gradients in the covariance eigenvalues is
\[
S := \frac{\nabla_\lambda L\cdot \nabla_\lambda I} {\|\nabla_\lambda L\|\,\|\nabla_\lambda I\|} = \frac{2+\gamma^2}{\sqrt{4+\gamma^2(4+d)}},
\]
where
\[
\gamma^2:=\frac{\operatorname{Trace}(\Delta)^2}{\|\Delta\|_F^2}\in[0,d].
\]
The reported asymptotic regimes show that once variances are matched, learning off-diagonal correlations through energy distance may become slow or weak, especially in high dimensions and local-correlation settings [2505.20647].

A common misconception is therefore to treat energy distance as uniformly sensitive to all local perturbations. The local expansion indicates otherwise: mean shifts remain locally prominent; variance changes are weaker; and correlation changes can be substantially attenuated.

## 4. Sample statistics, calibration, and testing workflow

For independent samples \(X_1,\dots,X_{n_1}\sim Q_1\) and \(Y_1,\dots,Y_{n_2}\sim Q_2\), the standard sample energy statistic is
\[
\mathcal{E}_{n_1,n_2}= \frac{2}{n_1n_2}\sum_{i=1}^{n_1}\sum_{m=1}^{n_2}\|X_i-Y_m\|
-\frac{1}{n_1^2}\sum_{i=1}^{n_1}\sum_{j=1}^{n_1}\|X_i-X_j\|
-\frac{1}{n_2^2}\sum_{l=1}^{n_2}\sum_{m=1}^{n_2}\|Y_l-Y_m\|.
\]
The associated scaled statistic is
\[
T_{n_1,n_2}=\frac{n_1n_2}{n}\mathcal{E}_{n_1,n_2},\qquad n=n_1+n_2,
\]
and large values lead to rejection of \(H_0\) [1703.07856].

Under the classical null, the pooled sample is exchangeable, so permutation calibration is nonparametric and distribution free under the null. This exact distribution-free property belongs to the permutation version. In the object-space extension discussed below, the implementation instead uses a nonparametric bootstrap from the pooled sample with replacement. The bootstrap statistic has the same algebraic form as the sample energy statistic after replacing the original observations with pooled resamples. The rejection rule is to compare the observed statistic with the empirical upper quantile of the bootstrap distribution [1703.07856].

The pairwise-distance structure determines the computational profile. Direct computation costs roughly
\[
O(n_1n_2+n_1^2+n_2^2),
\]
plus the cost of embedding and distance evaluation when the observations are not ordinary Euclidean vectors. This quadratic dependence on sample size is intrinsic to the naive implementation of pairwise energy statistics [1703.07856].

From the standpoint of a location interpretation, the key practical distinction is between inferential target and local mechanism. The formal target of the test remains equality of distributions, but in close-distribution settings the dominant signal may be a mean shift rather than a covariance or higher-moment difference [2505.20647].

## 5. Object spaces, extrinsic energy distance, and shape analysis

Energy distance extends beyond \(\mathbb{R}^d\) by embedding an object space \(\mathcal M\) into Euclidean space. If
\[
j:\mathcal M \to \mathbb R^N
\]
is one-to-one and a homeomorphism onto its image, then the extrinsic energy distance is
\[
\varepsilon_j(Q_1,Q_2)= 2E\|j(X)-j(Y)\|-E\|j(X)-j(X')\|-E\|j(Y)-j(Y')\|,
\]
assuming
\[
E\|j(X)\|<\infty,\qquad E\|j(Y)\|<\infty.
\]
Because \(j\) is an embedding, one has
\[
\varepsilon_j(Q_1,Q_2)=0 \quad \Longleftrightarrow \quad Q_1=Q_2.
\]
The method is therefore still a homogeneity test, now on object-valued data such as shapes, trees, axes, and diffusion tensors [1703.07856].

The sample extrinsic energy statistic is
\[
\mathcal{E}_{j,n_1,n_2}(X,Y)= \frac{2}{n_1n_2}\sum_{i=1}^{n_1}\sum_{m=1}^{n_2}\|j(X_i)-j(Y_m)\|
-\frac{1}{n_1^2}\sum_{i=1}^{n_1}\sum_{h=1}^{n_1}\|j(X_i)-j(X_h)\|
-\frac{1}{n_2^2}\sum_{l=1}^{n_2}\sum_{m=1}^{n_2}\|j(Y_l)-j(Y_m)\|,
\]
with scaled form
\[
T_{j,n_1,n_2} = \frac{n_1n_2}{n}\mathcal{E}_{j,n_1,n_2}(X,Y).
\]
This is exactly the classical energy-distance construction transferred to the embedded observations [1703.07856].

A principal example is planar Kendall shape space,
\[
\Sigma_2^k \cong \mathbb{CP}^{k-2},
\]
analyzed via the Veronese–Whitney embedding
\[
j([x])=xx^*, \qquad \|x\|=1.
\]
The corresponding chord-type distance is
\[
\rho([x],[y]) = Tr\big((xx^*-yy^*)^2\big), \qquad \|x\|=1,\ \|y\|=1.
\]
Plugging the VW embedding into the general framework yields the VW-energy distance for comparing distributions of Kendall shapes [1703.07856].

The Corpus Callosum application illustrates the distinction between a mean-shape comparison and a full distributional test. Each observation is a CC midsection contour discretized into \(50\) pseudo-landmarks, so the data lie in
\[
\Sigma_2^{50} \cong \mathbb{CP}^{48},
\]
of real dimension \(96\). Using the VW-energy statistic, the observed value was
\[
T_{j,n_1,n_2}=0.0948,
\]
with bootstrap critical values
\[
c^*_{0.05}=0.1075,\qquad c^*_{0.1}=0.0958.
\]
Since the observed statistic was below both cutoffs, the paper concluded that the two Kendall shape distributions were not significantly different, even though prior work cited by the authors had found highly significant differences in VW mean shape [1703.07856]. This example shows why “location test” can be too narrow a label for the broader energy-based homogeneity framework.

## 6. One-sample symmetry, reflected distributions, and right-censored survival data

A one-sample analogue arises in testing diagonal symmetry. For a \(d\)-variate random variable \(X\), the null is
\[
H_0: X \stackrel{d}{=} -X.
\]
Using the energy-distance identity,
\[
E|X+X'|-E|X-X'| =\frac{1}{2}\,\mathcal{E}(X,-X'),
\]
diagonal symmetry becomes equivalent to vanishing energy distance between \(X\) and its reflected version \(-X'\) [1908.06892].

The sample analogue is the degree-2 U-statistic
\[
U=\binom{n}{2}^{-1}\sum_{1\le i<j\le n} \bigl[\,|X_i+X_j|-|X_i-X_j|\,\bigr],
\]
but this statistic is degenerate under \(H_0\). The proposed remedy is a split-sample construction based on two U-statistics,
\[
U_1=\binom{n_1}{2}^{-1}\sum_{1\le k<j\le n_1}|X_k+X_j|,\qquad
U_2=\binom{n_2}{2}^{-1}\sum_{1\le k<j\le n_2}|Y_k-Y_j|,
\]
followed by jackknife empirical likelihood. The resulting log-likelihood ratio
\[
\ell = 2\sum_{i=1}^{n_1}\log\{1+\lambda_1(\hat V_i^{(1)}-\theta)\} + 2\sum_{j=1}^{n_2}\log\{1+\lambda_2(\hat V_j^{(2)}-\theta)\}
\]
satisfies
\[
\ell \xrightarrow{d} \chi^2_1.
\]
This is not a general location test, but it becomes a location-transformed problem when testing symmetry about a specified center \(p\): with \(Z=X-p\), the null becomes \(Z\stackrel d= -Z\) [1908.06892].

The power results clarify another interpretive limit. For pure location-shift alternatives such as \(N(0.5,I_d)\), the JEL symmetry test has relatively low power compared with ET and CD, even though it is asymptotically consistent. The paper attributes this to the weighting behavior of the JEL construction, which assigns more weight to sample points close to \(0\) [1908.06892]. Thus an energy-distance identity does not automatically yield a test optimized for mean shifts.

Energy-based methodology also extends to right-censored survival data. With observed data
\[
X_{ji}=\min(T_{ji},C_{ji}),\qquad \delta_{ji}=1\{X_{ji}=T_{ji}\},
\]
the paper replaces empirical distributions by Kaplan–Meier estimators and Stute’s KM integral weights. A naive KM plug-in energy statistic can converge to truncated quantities that are not necessarily distances and may even be negative. The corrected construction normalizes weighted sums to obtain censored weighted \(U\)-statistics, and the rescaled test statistics are
\[
T_{\hat{\epsilon}_\alpha}= \frac{n_0n_1}{n_0+n_1} \hat{\epsilon}_{\alpha}(P_0,P_1), \qquad
T_{\hat{\gamma}_K^{2}}= \frac{n_0n_1}{n_0+n_1} \hat{\gamma}_K^{2}(P_0,P_1).
\]
Under support conditions such as \(\tau_0=\tau_1\), these yield tests of equal distributions that are consistent against all fixed alternatives [1901.00833].

This censored setting also highlights the special status of \(\alpha=2\). The general \(\alpha\)-energy distance is a full equality-of-distributions criterion for \(\alpha\in(0,2)\), but at \(\alpha=2\) it collapses to
\[
\epsilon_2(P,Q)= 2||E(X)-E(Y)||^{2},
\]
so the criterion becomes purely location/mean based [1901.00833].

## 7. Interpretation, limitations, and scope of the term

The phrase “energy distance-based location test” therefore has a layered meaning. In the broad two-sample literature, energy distance is a nonparametric test of equality of distributions, not merely a location test. In the local perturbative regime analyzed through moment expansion, however, the same statistic behaves approximately like a location test because mean differences appear at order \(\lambda^{-1}\) while covariance differences appear only at order \(\lambda^{-3}\) [2505.20647].

Several limitations delimit this interpretation. The local results rely on a close-distribution or large-\(\lambda\) regime, finite moments, and \(C^5\)-type control sufficient to justify Taylor expansion under the integral. The cleanest formulas often assume approximate isotropy or spherical symmetry. The conclusions are population-level and asymptotic rather than exact finite-sample power statements [2505.20647]. In object spaces, the procedure is extrinsic and depends on the chosen embedding \(j\) rather than intrinsic manifold geodesics [1703.07856]. In survival analysis, validity under censoring requires Kaplan–Meier weighting, normalization, and support conditions tied to the censoring-limited endpoints [1901.00833]. In diagonal symmetry testing, the method targets symmetry about a specified center rather than an unrestricted location alternative [1908.06892].

A precise synthesis is therefore as follows. Energy distance defines an omnibus homogeneity framework that can be implemented in Euclidean data, object spaces, reflected one-sample problems, and right-censored survival settings. Its exact inferential target is equality of distributions. Nonetheless, when two distributions are close relative to a common scale, the leading local contribution is proportional to \(\|\mu\|^2/\lambda\), and an energy-distance-based two-sample procedure behaves approximately like a test for equality of means. The “location test” interpretation is thus mathematically justified locally, but it does not replace the broader and more general characterization of energy distance as a test of distributional equality [2505.20647].

Source: https://www.emergentmind.com/topics/energy-distance-based-location-test