Papers
Topics
Authors
Recent
Search
2000 character limit reached

Double-Sliced Wasserstein (DSW) Metric

Updated 12 November 2025
  • Double-Sliced Wasserstein (DSW) Metric is a fully metric-based approach that sequentially slices meta-measures, addressing the limitations of the Wasserstein-over-Wasserstein metric.
  • It leverages quantile isometry and Gaussian process–parametrized slicing in infinite-dimensional spaces to achieve computational efficiency and robust optimal transport.
  • Empirical validation shows DSW achieves competitive matching, high accuracy, and significant speedups for applications in shape analysis, dataset comparison, and patch-based image similarity.

The double-sliced Wasserstein (DSW) metric is a scalable, robust, and fully metric-based approach for comparing meta-measures—distributions over distributions—on Euclidean spaces, particularly relevant for applications in shape analysis, dataset comparison, and patch-based image distances. DSW provides a principled and computationally tractable substitute for the Wasserstein-over-Wasserstein (WoW) metric, overcoming limitations of higher-order moment requirements and instability in existing approaches. By applying sequential slices—first in Euclidean space, then in function space—DSW leverages the isometry between univariate measures and quantile functions, and employs Gaussian process–parametrized slicing in infinite-dimensional function spaces.

1. Wasserstein-over-Wasserstein (WoW) and its Computational Limitations

Given a ground Polish space X\mathcal{X}, the standard 2-Wasserstein distance defines a metric on P2(X)P_2(\mathcal{X}), the space of Borel probability measures with finite second moment: W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2} Extending this to meta-measures μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X})), the WoW metric is defined as

WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))

Computation becomes prohibitive for empirical meta-measures with NN underlying empirical measures (each over nn points in Rd\mathbb{R}^d), since pairwise inner Wasserstein matrices are O(N2n2logn)\mathcal{O}(N^2 n^2 \log n) and the outer transport is O(N3logN)\mathcal{O}(N^3\log N), which is infeasible for large P2(X)P_2(\mathcal{X})0 or P2(X)P_2(\mathcal{X})1.

2. Quantile Isometry and Sliced Optimal Transport in Banach Spaces

For measures on P2(X)P_2(\mathcal{X})2, the quantile map P2(X)P_2(\mathcal{X})3 acts as an isometric embedding: P2(X)P_2(\mathcal{X})4 This property motivates generalizing sliced optimal transport to Banach spaces. For a separable Banach space P2(X)P_2(\mathcal{X})5 and direction P2(X)P_2(\mathcal{X})6, projections P2(X)P_2(\mathcal{X})7 yield the sliced Wasserstein distance: P2(X)P_2(\mathcal{X})8 When P2(X)P_2(\mathcal{X})9 is infinite-dimensional, there is no uniform sphere; instead, W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2}0 is typically chosen as a Gaussian on W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2}1.

3. Construction of the Double-Sliced Wasserstein Metric

The DSW metric applies two slicing steps for meta-measures in W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2}2:

  1. Euclidean slicing: For each direction W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2}3, measures W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2}4 are projected: W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2}5. The meta-measure W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2}6 is consequently pushed forward, producing W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2}7.
  2. Functional slicing: Applying the quantile isometry W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2}8 to each constituent, the resulting meta-measure is now over W(μ,ν;X)=infπΓ(μ,ν)(X×Xd2(x,y)dπ(x,y))1/2W(\mu, \nu; \mathcal{X}) = \inf_{\pi \in \Gamma(\mu,\nu)} \left( \iint_{\mathcal{X} \times \mathcal{X}} d^2(x,y) \, d\pi(x,y) \right)^{1/2}9. Slices μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X}))0 (sampled via a Gaussian process) project to 1D empirical measures in μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X}))1. Wasserstein distance is computed on these projections.

The DSW metric is then

μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X}))2

with μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X}))3 the uniform measure and μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X}))4 the Gaussian process law.

4. Equivalence and Properties on Discrete Meta-Measures

For empirical meta-measures μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X}))5 and μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X}))6, DSW minimization yields the same optimal matchings (transport plans) as WoW minimization. The matching induced by DSW collapses to permutations (Cramér–Wold argument), ensuring plans only between identical support indices when μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X}))7 for all μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X}))8. Thus, the computational procedure for DSW coincides—on discretized data—with the theoretically optimal WoW solution.

5. Algorithm, Complexity, and Implementation

Given two meta-measures each supported on μ,νP2(P2(X))\boldsymbol{\mu}, \boldsymbol{\nu} \in P_2(P_2(\mathcal{X}))9 empirical measures of WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))0 points in WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))1:

  • Loop over WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))2 Euclidean directions WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))3, project all WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))4 to WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))5; sort to obtain their quantile functions.
  • For each WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))6 and WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))7 function space slices WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))8 (Gaussian process realizations), compute scalars WoW(μ,ν)=W(μ,ν;P2(X))\mathrm{WoW}(\boldsymbol{\mu}, \boldsymbol{\nu}) = W(\boldsymbol{\mu}, \boldsymbol{\nu}; P_2(\mathcal{X}))9, NN0 analogously.
  • Sort the vectors NN1 and evaluate the 1D Wasserstein distance NN2 in NN3.
  • Aggregate the squared Wasserstein values across all slices, outputting

NN4

The total computational complexity is

NN5

where NN6 is the grid size for NN7 integration. For moderate NN8, DSW achieves significant speedups over WoW, whose complexity is NN9.

6. Theoretical and Topological Properties

  • DSW is a metric on empirical meta-measures for suitable, full-support Gaussian nn0 and full angular integration.
  • DSW metrizes the same topology as WoW on discrete meta-measures: nn1 iff nn2.
  • DSW is Lipschitz-stable with respect to changes in the outer meta-measure, up to constants based on the second moment of nn3.
  • Monte Carlo estimation achieves convergence rate nn4, given sufficient outer and inner slices.

7. Empirical Validation and Applications

Experiments illustrate DSW's efficiency and fidelity relative to WoW across three domains:

Application DSW Runtime WoW Runtime / Accuracy DSW Accuracy / Correlation
Shape classification (K-NN, 2D/3D data) 2 ms/pair GW: 40 ms/pair; both ≈99% small 42.7% ± 5.9 (FAUST-1000)
Dataset distance (MNIST, CIFAR-10 splits) s-OTDD corr: 0.75–0.85 corr(DSW, OTDD) = 0.90–0.95
Patch-based image similarity patch-WoW: 40× slower Agreement with inception kernel

On shape data, DSW matches GW for accuracy, running 20× faster. For dataset distances (e.g., OTDD replacement), DSW achieves 5–10× speedup over s-OTDD with higher correlation to ground-truth OTDD. As a patch-based image similarity metric, DSW closely tracks kernel-inception distance while being ~40× faster than patch-based WoW.

DSW's construction is motivated by previous sliced Wasserstein approaches, which often rely on parametric meta-measures or high-order moments, creating numerical instability. The DSW approach circumvents these issues by operating exclusively via the nn5 quantile isometry and slicing in nn6 via Gaussian processes. It stands in contrast to single-level distributional slicing (e.g., SW, Max-SW, distributional SW as in v-DSW), which project only once and do not handle meta-measures. DSW is not a max-slice metric and retains strict metricity, avoiding the limitations of non-metric approximations commonly encountered in max-sliced or amortized distributional projection variants (Nguyen et al., 2023).

DSW thereby provides a general, mathematically principled, and scalable framework for optimal transport–based comparison of meta-distributions across a wide array of scientific and machine learning applications (Piening et al., 26 Sep 2025).

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Double-Sliced Wasserstein (DSW) Metric.