Papers
Topics
Authors
Recent
Search
2000 character limit reached

Double-Sliced Wasserstein Metric

Updated 11 November 2025
  • Double-Sliced Wasserstein is a metric that compares probability meta-measures using two sequential slicing operations, preserving the topology of the original Wasserstein-over-Wasserstein distance.
  • It combines Euclidean projections and quantile-space slicing to achieve computational efficiency and numerical robustness in high-dimensional data analysis.
  • Empirical evaluations demonstrate that DSW provides comparable discriminative power to WoW while accelerating computation and reducing sensitivity to unstable high-order moment estimation.

The Double-Sliced Wasserstein (DSW) metric is a recent development in the study of optimal transport on spaces of probability measures, specifically designed as a computationally efficient and statistically robust surrogate for the Wasserstein-over-Wasserstein (WoW) distance between meta-measures. The DSW metric achieves speed and stability by combining traditional Euclidean slicing with an inner slicing in quantile function space, avoiding reliance on high-order moments or unstable operations. DSW is topologically equivalent to WoW on empirical meta-measures and empirically offers substantial speedups with comparable discriminative power for applications in dataset similarity, point-cloud analysis, and perceptual evaluation of images and shapes (Piening et al., 26 Sep 2025).

1. Meta-Measure Spaces and the Wasserstein-Over-Wasserstein Problem

Let X\mathcal{X} be a Polish space and P2(X)P_2(\mathcal{X}) the set of Borel probability measures with finite second moment, equipped with the 2-Wasserstein distance,

W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.

A meta-measure is defined as α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr), that is, a probability law over probability measures on X\mathcal{X}. The Wasserstein-over-Wasserstein (WoW) metric lifts the W2W_2 distance to the meta-measure space: WoW(α,β)=[inf⁡Π∈Γ(α,β)∫P2(X)×P2(X)W22(μ,ν) dΠ(μ,ν)]1/2,\mathrm{WoW}(\alpha, \beta) = \left[ \inf_{\Pi \in \Gamma(\alpha,\beta)} \int_{P_2(\mathcal{X}) \times P_2(\mathcal{X})} W_2^2(\mu,\nu)\, d\Pi(\mu,\nu) \right]^{1/2}, which is computationally prohibitive for large collections of distributions, especially in high-dimensions due to quadratic scaling in the number of inner measures.

2. Quantile Isometry and Functional Slicing

For measures on R\mathbb{R}, the 1D 2-Wasserstein metric admits an isometry to L2([0,1])L^2([0,1]), mapping a measure μ\mu to its quantile function P2(X)P_2(\mathcal{X})0: P2(X)P_2(\mathcal{X})1 This isometry underpins the functional optimal transport approach used in DSW. Sliced-Wasserstein distances on general Banach spaces P2(X)P_2(\mathcal{X})2 make use of projections P2(X)P_2(\mathcal{X})3 for P2(X)P_2(\mathcal{X})4, and for a probability measure P2(X)P_2(\mathcal{X})5 on P2(X)P_2(\mathcal{X})6,

P2(X)P_2(\mathcal{X})7

This construction, under appropriate support conditions on P2(X)P_2(\mathcal{X})8, yields a true metric on P2(X)P_2(\mathcal{X})9.

In the specific setting of meta-measures on W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.0, the quantile map W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.1 pushes W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.2 to a law W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.3 on W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.4, yielding a “sliced-quantile WoW” (SQW) metric,

W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.5

3. Construction and Mathematical Formulation of Double-Sliced Wasserstein

The Double-Sliced Wasserstein metric is constructed through consecutive application of two slicing steps:

  1. Euclidean Slicing: For each W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.6, project every inner measure W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.7 onto W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.8 via W2(μ,ν)=(inf⁡π∈Γ(μ,ν)∫X2d2(x,x′) dπ(x,x′))1/2.W_2(\mu,\nu) = \left(\inf_{\pi\in\Gamma(\mu,\nu)} \int_{\mathcal{X}^2} d^2(x,x')\,d\pi(x,x')\right)^{1/2}.9, inducing a pushed-forward measure α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr)0.
  2. Quantile-Space Slicing: For fixed α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr)1, one obtains two 1D meta-measures α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr)2. Using a Gaussian process prior α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr)3 on α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr)4 (e.g., with an RBF kernel), the SQW distance between the meta-measures is

α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr)5

  1. Aggregation: Integrate the inner SQW metric over α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr)6 to obtain the Double-Sliced Wasserstein: α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr)7

The full expansion writes: α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr)8

For computation, inner integrals are estimated using Monte Carlo samples α∈P2(P2(X))\alpha\in P_2\bigl(P_2(\mathcal{X})\bigr)9 and X\mathcal{X}0 Gaussian process paths.

4. Topological Properties and Equivalence with WoW

Let empirical meta-measures X\mathcal{X}1 be composed of X\mathcal{X}2 inner empirical measures, each with X\mathcal{X}3 support points. The DSW metric is topologically equivalent to the WoW metric: X\mathcal{X}4 for any positive Gaussian X\mathcal{X}5 (Piening et al., 26 Sep 2025). The argument combines a discrete Cramér–Wold theorem at each slice X\mathcal{X}6 and the quantile-space isometry. This ensures that DSW is a true metric and it preserves the geometry induced by WoW on the space of empirical meta-measures.

5. Computational Complexity and Numerical Stability

Metric Complexity per evaluation Stability Considerations
WoW X\mathcal{X}7 Requires all pairwise inner 2-Wasserstein computations; slow for large X\mathcal{X}8 and X\mathcal{X}9; sensitive to moment estimation
DSW W2W_20 Only W2W_21 projections needed; relies on quantile functions (no high-order moments); numerically robust
  • For full WoW on W2W_22 meta-points of size W2W_23, cost is dominated by an W2W_24 matrix of pairwise 2-Wasserstein computations, each in W2W_25 (entropic case).
  • For DSW, sampling W2W_26 directions and computing 1D quantile-transport per meta-measure, total complexity is W2W_27; usually W2W_28 or constant.
  • DSW avoids the unstable high-order moments used in s-OTDD, maintaining stability even with non-Gaussian or heavy-tailed meta-distributions.

6. Empirical Results and Applications

Experimental evaluations in (Piening et al., 26 Sep 2025) demonstrate that DSW achieves strong performance in several tasks:

  • Shape classification via local distance distributions (mm-spaces): DSW matches accuracy of Gromov–Wasserstein and sliced GW (STLB), with an order of magnitude faster runtime.
  • Dataset similarity (OTDD surrogate): On MNIST, Fashion-MNIST, and CIFAR-10 splits, DSW correlates with exact OTDD (Pearson W2W_29), outperforming s-OTDD in stability and speed.
  • Point-cloud evaluation: For batches of 3D shapes modeled as meta-measures in WoW(α,β)=[inf⁡Π∈Γ(α,β)∫P2(X)×P2(X)W22(μ,ν) dΠ(μ,ν)]1/2,\mathrm{WoW}(\alpha, \beta) = \left[ \inf_{\Pi \in \Gamma(\alpha,\beta)} \int_{P_2(\mathcal{X}) \times P_2(\mathcal{X})} W_2^2(\mu,\nu)\, d\Pi(\mu,\nu) \right]^{1/2},0, DSW matches OT-NNA and WoW in sensitivity but is 10–20WoW(α,β)=[inf⁡Π∈Γ(α,β)∫P2(X)×P2(X)W22(μ,ν) dΠ(μ,ν)]1/2,\mathrm{WoW}(\alpha, \beta) = \left[ \inf_{\Pi \in \Gamma(\alpha,\beta)} \int_{P_2(\mathcal{X}) \times P_2(\mathcal{X})} W_2^2(\mu,\nu)\, d\Pi(\mu,\nu) \right]^{1/2},1 faster, providing similar robustness to mode collapse and sampling noise.
  • Image perceptual distance: Image batches are represented as meta-measures on patch distributions. DSW defines a perceptual metric sensitive to qualitative similarity, aligns with standard fiducial metrics (e.g., Kernel Inception Distance), and is 40WoW(α,β)=[inf⁡Π∈Γ(α,β)∫P2(X)×P2(X)W22(μ,ν) dΠ(μ,ν)]1/2,\mathrm{WoW}(\alpha, \beta) = \left[ \inf_{\Pi \in \Gamma(\alpha,\beta)} \int_{P_2(\mathcal{X}) \times P_2(\mathcal{X})} W_2^2(\mu,\nu)\, d\Pi(\mu,\nu) \right]^{1/2},2 faster than full WoW.

These results indicate that DSW yields operationally efficient metrics for meta-measure comparison without the compromises of parametric forms or unstable statistical estimators.

7. Significance and Prospects

Double-Sliced Wasserstein provides a tractable, mathematically principled metric for meta-level optimal transport problems, preserving the topology and discriminative power of WoW while mitigating prohibitive computational demands. Its combination of classical slicing and functional quantile-space slicing leverages both geometry and statistical properties of optimal transport. The approach is widely applicable to large-scale shape analysis, dataset comparison, and the evaluation of structured or hierarchical data distributions.

Pending open questions include: optimizing DSW kernel choices for task-adaptiveness, theoretical dual formulations for functional slicing, and extensions beyond empirical meta-measures to infinite or continuous families. A plausible implication is that DSW could serve as a foundation for scalable learning frameworks in high-level data spaces where conventional OT remains intractable, especially in large-dimensional and nonparametric distributional regimes (Piening et al., 26 Sep 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Double-sliced Wasserstein (DSW).