---
title: Pivot Sliced Discrepancy Analysis
url: https://www.emergentmind.com/topics/pivot-sliced-discrepancy
type: topic
---

# Pivot Sliced Discrepancy Analysis

Pivot Sliced Discrepancy is a direction-dependent sliced transport discrepancy introduced to endow sliced optimal transport constructions with a rigorous transport-plan interpretation. For probability measures \(\mu_1,\mu_2\in\mathcal P_2(\mathbb R^d)\), it replaces the ambiguous lifting of one-dimensional sliced couplings by a pivot-based \(3\)-marginal optimal transport formulation built from a unique Wasserstein midpoint of the projected measures. In the formulation of Tanguy, Chapel, and Delon, the resulting quantity \(\mathrm{PS}_\theta\) is well-defined, symmetric, separating, bounded below by \(W_2\), equal to a constrained Kantorovich problem, and metric on a restricted atomless class, while failing to satisfy the triangle inequality in full generality [2508.01243].

## 1. Conceptual setting and motivation

The classical sliced Wasserstein distance
\[
SW_2^2(\mu_1,\mu_2)=\int_{\mathbb S^{d-1}} W_2^2(P_\theta\#\mu_1,\;P_\theta\#\mu_2)\,d\theta
\]
is attractive because each one-dimensional Wasserstein term is cheap to compute, but it only yields a scalar discrepancy and not a canonical transport plan in \(\mathbb R^d\). The central obstruction is projection ambiguity: in dimension \(d>1\), the projection \(x\mapsto \theta^\top x\) loses information, different points can share the same projected value, and the one-dimensional optimal plan does not determine how mass should be matched in orthogonal directions [2508.01243].

This issue is made explicit through the earlier Sliced Wasserstein Generalised Geodesic heuristic, which for discrete measures sorts projected points and pairs them. When projected values have ties, the resulting quantity depends on arbitrary sorting permutations. Pivot Sliced Discrepancy is designed precisely to remove that arbitrariness. The construction introduces a pivot measure on the projected line, chosen from the projected marginals themselves, and encodes admissible correspondences through a constrained \(3\)-plan.

The terminology should be distinguished from other uses of “sliced discrepancy” in the literature. In cut-and-project theory, “sliced discrepancy” refers to lattice-point counting in thin slabs such as \((t\Omega_\triangledown)\times\Omega_\triangle\) [2401.01803]. In image comparison, convolution sliced Wasserstein replaces linear projections by convolutional slicers rather than using a pivot measure [2204.01188]. In Stein methods, sliced kernelized Stein discrepancy uses projected score tests indexed by directions \(\bm r\) and \(\bm g\) [2006.16531]. These are distinct constructions.

## 2. Mathematical definition and the pivot measure

For a direction \(\theta\in\mathbb S^{d-1}\), the paper uses the scalar and line projections
\[
P_\theta(x)=\theta^\top x\in\mathbb R,\qquad Q_\theta(x)=(\theta^\top x)\theta\in\mathbb R^d.
\]
Given \(\mu_1,\mu_2\in\mathcal P_2(\mathbb R^d)\), the pivot is defined as a Wasserstein midpoint of the projected measures \(Q_\theta\#\mu_1\) and \(Q_\theta\#\mu_2\). The set of Wasserstein means is
\[
M(\mu_1,\mu_2):=\arg\min_{\mu\in\mathcal P_2(\mathbb R^d)}\Bigl(W_2^2(\mu_1,\mu)+W_2^2(\mu,\mu_2)\Bigr),
\]
and in the projected one-dimensional setting the mean is unique:
\[
M(Q_\theta\#\mu_1,Q_\theta\#\mu_2)=\{\mu_\theta[\mu_1,\mu_2]\}.
\]
It is given explicitly by
\[
\mu_\theta[\mu_1,\mu_2]
=
\left[\left(\tfrac12 F_{P_\theta\#\mu_1}^{[-1]}+\tfrac12 F_{P_\theta\#\mu_2}^{[-1]}\right)\theta\right]\#L_{[0,1]},
\]
where
\[
F_\nu^{[-1]}(t)=\inf\{s\in\mathbb R:\nu((-\infty,s])\ge t\}.
\]

Pivot Sliced Discrepancy is then defined by inserting this pivot into the \(\nu\)-based Wasserstein distance:
\[
\mathrm{PS}_\theta(\mu_1,\mu_2):=W_{\mu_\theta}(\mu_1,\mu_2),
\]
with
\[
W_\nu^2(\mu_1,\mu_2)
:=
\min_{\rho\in\Gamma(\nu,\mu_1,\mu_2)}
\int_{\mathbb R^{3d}}\|x_1-x_2\|_2^2\,d\rho(y,x_1,x_2).
\]
The admissible \(3\)-plans are
\[
\Gamma(\nu,\mu_1,\mu_2)
=
\left\{
\rho\in\mathcal P_2(\mathbb R^{3d}) :
\rho_{0,1}\in\Pi^*(\nu,\mu_1),\;
\rho_{0,2}\in\Pi^*(\nu,\mu_2)
\right\}.
\]

Accordingly, PSD is a pivot-based \(3\)-marginal optimal transport problem in which the pivot is the projected Wasserstein midpoint. A plausible implication is that the entire construction is designed to encode consistency across the two marginals through a common intermediate measure rather than through an arbitrary lifting of one-dimensional pairings.

## 3. Relation to \(\nu\)-based Wasserstein distance and constrained transport

A central theoretical backbone is the \(\nu\)-based Wasserstein distance analyzed via generalized geodesics. If \(\rho\in\Gamma(\nu,\mu_1,\mu_2)\), the associated generalized geodesic is
\[
\mu_\rho^{1\to2}(t)=((1-t)P_1+tP_2)\#\rho.
\]
The paper recalls the identity
\[
W_2^2(\mu_1,\mu_\rho^{1\to2}(t))
=
(1-t)W_2^2(\mu_1,\nu)+tW_2^2(\mu_2,\nu)
-(1-t)t\int \|x_1-x_2\|_2^2\,d\rho,
\]
and for an optimal \(\rho^*\),
\[
W_2^2(\mu_1,\mu_{\rho^*}^{1\to2}(t))
=
(1-t)W_2^2(\mu_1,\nu)+tW_2^2(\mu_2,\nu)
-(1-t)t\,W_\nu^2(\mu_1,\mu_2).
\]
This places PSD in the framework of generalized geodesics based on a prescribed pivot measure [2508.01243].

The same paper defines a constrained sliced transport cost
\[
\mathrm{CW}_\theta^2(\mu_1,\mu_2)
:=
\min_{\substack{\omega\in\Pi(\mu_1,\mu_2)\\ (P_\theta,P_\theta)\#\omega=\pi_\theta[\mu_1,\mu_2]}}
\int_{\mathbb R^{2d}}\|x_1-x_2\|_2^2\,d\omega(x_1,x_2),
\]
where \(\pi_\theta[\mu_1,\mu_2]\) is the unique one-dimensional optimal transport plan between \(P_\theta\#\mu_1\) and \(P_\theta\#\mu_2\).

The key theorem states
\[
\mathrm{PS}_\theta(\mu_1,\mu_2)=\mathrm{CW}_\theta(\mu_1,\mu_2).
\]
The equivalence is proved by comparing admissible sets in both directions: every admissible \(\omega\) for \(\mathrm{CW}_\theta\) can be lifted to a \(3\)-plan admissible for \(\mathrm{PS}_\theta\), and every admissible \(3\)-plan for \(\mathrm{PS}_\theta\) induces a coupling admissible for \(\mathrm{CW}_\theta\). In this sense, PSD is also a constrained Kantorovich problem: it minimizes transport cost among couplings whose projected coupling is exactly the one-dimensional optimal plan.

This equivalence clarifies the role of the pivot. The pivot is not an auxiliary artifact; it is the mechanism through which the one-dimensional optimal coupling is imposed as a hard projection constraint in the ambient transport problem.

## 4. Structural properties: semimetric behavior, metricity, and regularity

The paper proves that PSD is well-defined because the projected Wasserstein mean is unique in one dimension. For all \(\mu_1,\mu_2\in\mathcal P_2(\mathbb R^d)\), it satisfies symmetry,
\[
\mathrm{PS}_\theta(\mu_1,\mu_2)=\mathrm{PS}_\theta(\mu_2,\mu_1),
\]
separation,
\[
\mathrm{PS}_\theta(\mu_1,\mu_2)=0 \iff \mu_1=\mu_2,
\]
and a lower bound by Wasserstein,
\[
\mathrm{PS}_\theta(\mu_1,\mu_2)\ge W_2(\mu_1,\mu_2).
\]
These properties justify describing PSD as a semi-metric rather than a metric in general [2508.01243].

The triangle inequality fails in general; the paper gives a counterexample. However, PSD becomes a genuine metric on
\[
\mathcal P_{2,a}(\mathbb R^d,\theta)
=
\{\mu\in\mathcal P_2(\mathbb R^d): P_\theta\#\mu \text{ is atomless}\},
\]
where
\[
\mathrm{PS}_\theta(\mu_1,\mu_3)\le
\mathrm{PS}_\theta(\mu_1,\mu_2)+\mathrm{PS}_\theta(\mu_2,\mu_3).
\]
The proof uses a one-dimensional lemma asserting that if a \(3\)-plan has two optimal one-dimensional bi-marginals and the relevant marginals are atomless, then the third bi-marginal is optimal as well.

The regularity theory is similarly qualified. The map \((\theta,\mu_1,\mu_2)\mapsto \mu_\theta[\mu_1,\mu_2]\) is continuous, while \((\theta,\mu_1,\mu_2)\mapsto \mathrm{PS}_\theta(\mu_1,\mu_2)\) is lower semicontinuous. Full continuity fails in general. A common misconception is therefore that PSD simply repairs sliced Wasserstein into a genuine metric without loss; the paper does not support that interpretation. Instead, it establishes a more precise statement: PSD has metric-like properties globally and full metricity only on a restricted atomless subclass.

## 5. Discrete formulations and exact recovery phenomena

For empirical measures
\[
\mu=\frac1n\sum_{i=1}^n\delta_{x_i},
\qquad
\nu=\frac1n\sum_{j=1}^n\delta_{y_j},
\]
the constrained formulation specializes to a discrete Monge-type problem:
\[
\mathrm{CW}_\theta^2(\mu,\nu)
=
\min_{(\sigma,\tau)\in S_\theta(X,Y)}
\frac1n\sum_{i=1}^n \|x_{\sigma(i)}-y_{\tau(i)}\|_2^2,
\]
where \(S_\theta(X,Y)\) is the set of pairs of permutations that sort the projected samples. The derivation relies on a constrained Birkhoff–von Neumann theorem,
\[
\mathrm{Extr}(U\cap P)=P(S_\theta(X,Y)),
\]
with \(U\) the doubly stochastic polytope and \(P\) the projection constraints [2508.01243].

This discrete characterization is important for two reasons. First, it shows that PSD does recover a one-to-one matching structure in the empirical uniform setting, rather than merely a scalar discrepancy. Second, it makes explicit that the admissible permutations are not arbitrary: they are precisely those compatible with the projected one-dimensional optimal transport plan.

The paper also defines the Min-Pivot Sliced discrepancy
\[
PS^2(\mu_1,\mu_2)=\min_{\theta\in\mathbb S^{d-1}}\mathrm{PS}_\theta^2(\mu_1,\mu_2).
\]
For discrete measures in general position, if \(2n\le d+1\), then
\[
PS(\mu,\nu)=W_2(\mu,\nu).
\]
This exact recovery result states that in sufficiently high dimension, optimization over the slicing direction can recover the full Wasserstein distance for uniform point clouds. This suggests that the directional constraint imposed by PSD need not be intrinsically lossy when the ambient dimension is large relative to sample size.

## 6. Comparison with related sliced constructions and practical behavior

The most immediate comparison is with SWGG and Expected Sliced transport plans. SWGG is ill-defined when projections have ties because different sorting permutations can change the cost. PSD fixes this by using the one-dimensional optimal transport plan rather than arbitrary sorting and by encoding admissible couplings through the pivot-based constrained formulation. The paper also generalizes Expected Sliced transport plans, defining an averaged lifted plan \(\gamma[\mu_1,\mu_2,\bbsigma]\) and a discrepancy
\[
ES_{\bbsigma}^2(\mu_1,\mu_2)=\int_{\mathbb S^{d-1}} LS^2(\mu_1,\mu_2)\,d\bbsigma(\theta).
\]
It shows that \(ES_{\bbsigma}\) is nonnegative, symmetric, satisfies the triangle inequality, and \(ES_{\bbsigma}\ge W_2\), but it is not a distance in general because self-cost may be positive:
\[
ES_{\bbsigma}(\mu,\mu)>0
\]
for some \(\mu\), including the uniform measure on the unit disk in \(\mathbb R^2\) [2508.01243].

Practical experiments in the same work compare \(W_2\), classical \(SW_2\), PSD or min-Pivot Sliced, and Expected Sliced. In synthetic gradient flows, Expected Sliced does not converge reliably; in \(500\) dimensions, \(W_2\) and optimized PSD behave similarly and converge, while random-direction sliced methods struggle. In image color transfer, PSD often gives visually plausible transfers, though classical SW flow sometimes performs better in difficult color distributions. In rigid shape registration of point clouds such as bunny and armadillo, PSD performs particularly well, often outperforming nearest-neighbor ICP, SW, and Expected Sliced, and can help escape poor local minima by providing better correspondences. The reported sliced-based complexity is roughly \(O(Lnd + Ln\log n)\) for \(L\) directions, and PSD as well as Expected Sliced avoid the \(n^2\)-memory cost of a full cost matrix.

The term “pivot sliced discrepancy” should not be conflated with other sliced methodologies that use different mechanisms. Convolution Sliced Wasserstein preserves image structure by replacing vectorization and linear projections with convolutional slicers; it is explicitly not pivot-based [2204.01188]. Sliced Kernelized Stein Discrepancy can be read as using a pivot direction \(\bm r\) for score projection and a slicing direction \(\bm g\) for one-dimensional test functions, but it belongs to Stein discrepancy rather than optimal transport [2006.16531]. In contrast, PSD is specifically a pivot-based sliced transport discrepancy built from a unique projected Wasserstein midpoint and equivalent to a constrained Kantorovich problem.

Source: https://www.emergentmind.com/topics/pivot-sliced-discrepancy