Papers
Topics
Authors
Recent
Search
2000 character limit reached

Subspace Robust Wasserstein Distance

Updated 11 December 2025
  • Subspace robust Wasserstein distance is a metric that projects probability distributions onto lower-dimensional subspaces to enhance robustness against noise and the curse of dimensionality.
  • It employs convex relaxations and eigenvalue formulations to replace complex non-convex optimization with efficient saddle-point and gradient-based methods.
  • Practical applications include generative modeling, domain adaptation, and two-sample testing, offering dimension-free statistical rates and improved computational tractability.

The subspace robust Wasserstein distance is a family of optimal transport–based metrics designed to robustify Wasserstein distances with respect to noise and the curse of dimensionality, especially in high- or infinite-dimensional settings. The core idea is to measure the transportation cost between two probability distributions after projecting onto lower-dimensional subspaces, but in a worst-case sense, i.e., by optimizing the subspace itself adversarially or over a set of admissible directions. There are distinct formalizations: the "projection-robust" (PRW) Wasserstein, which is a max–min or supremum-over-subspaces of Wasserstein projections, and the "subspace robust" Wasserstein (SRW), which relaxes the order of min and max, leading to a min–max or partial trace (top eigenvalues) cost. These distances interpolate between full-dimensional Wasserstein when k=dk=d and the more statistically stable (sliced or randomized) variants when k≪dk\ll d, and have sharp statistical, geometric, and computational properties.

1. Formal Definitions and Metric Structure

Let H\mathcal H denote a separable real Hilbert space with norm ∥⋅∥\|\cdot\|, and μ,ν\mu,\nu be Borel probability measures on H\mathcal H (for most finite-dimensional treatments, H=Rd\mathcal H=\mathbb R^d). The set of couplings Π(μ,ν)\Pi(\mu,\nu) consists of all joint measures on H×H\mathcal H\times\mathcal H with respective marginals μ\mu and k≪dk\ll d0.

Given integer k≪dk\ll d1, define k≪dk\ll d2 as the Grassmannian of all k≪dk\ll d3-dimensional linear subspaces of k≪dk\ll d4. For each k≪dk\ll d5, k≪dk\ll d6 denotes the orthogonal projection onto k≪dk\ll d7.

The k≪dk\ll d8-dimensional subspace robust Wasserstein distance of order k≪dk\ll d9 is

H\mathcal H0

Equivalently, H\mathcal H1 (Vasan, 4 Dec 2025, Paty et al., 2019). For H\mathcal H2, the notation H\mathcal H3 is conventional.

A closely related object is the projection-robust Wasserstein distance (PRW, also denoted H\mathcal H4), defined as

H\mathcal H5

where H\mathcal H6 is the classical H\mathcal H7-Wasserstein. In discrete formulations, both can be written as max–min or min–max problems over the subspace and coupling variables (Paty et al., 2019, Jiang et al., 2022, Lin et al., 2020).

Paty–Cuturi (2019) establish that H\mathcal H8 is a bona fide metric for measures with finite H\mathcal H9th moments (Paty et al., 2019).

2. Convex Relaxation and Eigenvalue Formulation

The min–max order in ∥⋅∥\|\cdot\|0 admits a convex relaxation based on partial trace. For a coupling ∥⋅∥\|\cdot\|1, define the second moment displacement matrix: ∥⋅∥\|\cdot\|2 Fan’s maximum principle gives

∥⋅∥\|\cdot\|3

where ∥⋅∥\|\cdot\|4 are the ordered eigenvalues of ∥⋅∥\|\cdot\|5.

Thus, the SRW distance admits

∥⋅∥\|\cdot\|6

This is equivalent to a convex–concave saddle point problem

∥⋅∥\|\cdot\|7

where ∥⋅∥\|\cdot\|8 (Paty et al., 2019). The order of optimization makes the SRW distance convex in the coupling and over the set of Mahalanobis weights.

For PRW, the non-convex max–min formulation is inherited: ∥⋅∥\|\cdot\|9 with μ,ν\mu,\nu0 on the Stiefel manifold (Lin et al., 2020, Huang et al., 2020, Jiang et al., 2022).

3. Statistical Properties: Sample Complexity and Convergence

For empirical estimation, let μ,ν\mu,\nu1 be i.i.d. from μ,ν\mu,\nu2, with μ,ν\mu,\nu3. In an infinite-dimensional Hilbert space, the main result shows

μ,ν\mu,\nu4

for universal μ,ν\mu,\nu5 (Vasan, 4 Dec 2025). For general μ,ν\mu,\nu6,

μ,ν\mu,\nu7

The proof proceeds via a decomposition on well-chosen finite-dimensional projections and operator-norm bounds.

The classical Wasserstein rate in μ,ν\mu,\nu8 is μ,ν\mu,\nu9, which degenerates rapidly for large H\mathcal H0. The SRW and PRW distances, in contrast, have dimension-free rates that depend only logarithmically (SRW, Hilbert setting) or polynomially in H\mathcal H1 (PRW, finite dimension) (Lin et al., 2020): H\mathcal H2 with technical variations depending on tail assumptions.

The lower bound in H\mathcal H3 is unimprovable up to a H\mathcal H4 factor (Vasan, 4 Dec 2025).

4. Geometric, Robustness, and Metric Properties

The SRW and PRW distances inherit key properties from optimal transport geometry (Paty et al., 2019):

  • Both are metrics for suitable moment conditions.
  • They satisfy bounds: H\mathcal H5 (tight).
  • Increments in H\mathcal H6 are concave—increasing H\mathcal H7 improves discrimination at diminishing returns.
  • Dirac consistency: H\mathcal H8.
  • Geodesic properties: interpolants along OT plans are geodesics under H\mathcal H9.
  • Stability: trimming away noise directions (smallest H=Rd\mathcal H=\mathbb R^d0 eigenmodes) yields robustness to high-frequency (isotropic) perturbations.

Empirical studies on synthetic “fragmented hypercube” models and real datasets (e.g., word-embedding distributions from film scripts) show that H=Rd\mathcal H=\mathbb R^d1 reflects intrinsic low-dimensional structure and clusters semantically similar objects with greater stability to noise and outliers (Paty et al., 2019, Lin et al., 2020).

5. Computational Methods and Algorithms

Computation of SRW and PRW is challenging due to non-convexity (for PRW) and projection-over-subspace maximization. For SRW, convex relaxations and entropic regularization via Sinkhorn's algorithm are central, with two practical algorithms (Paty et al., 2019):

  • Projected non-smooth supergradient ascent: Maximizes over H=Rd\mathcal H=\mathbb R^d2 using the supergradient H=Rd\mathcal H=\mathbb R^d3.
  • Frank–Wolfe with entropic regularization: Adds smoothness, with per-iteration complexity H=Rd\mathcal H=\mathbb R^d4.

For PRW, manifold optimization is employed. The following table summarizes primary algorithms and their properties:

Algorithm Problem Formulation Complexity Bound
RBCD Nonconvex max–min (Stiefel, OT) H=Rd\mathcal H=\mathbb R^d5 (Huang et al., 2020)
iRBBS (ReALM) Manifold-constrained ALM H=Rd\mathcal H=\mathbb R^d6 (Jiang et al., 2022)
RGAS/RAGAS/RSGAN Riemannian (Entropic/Exact OT) H=Rd\mathcal H=\mathbb R^d7 (RGAS), H=Rd\mathcal H=\mathbb R^d8 (RSGAN) (Lin et al., 2020)

All methods alternate between subspace updates (Riemannian gradient/ascent/retraction steps on the Stiefel manifold) and cost/coupling updates (via Sinkhorn or network simplex OT solvers). Retractions implemented via QR, polar, Cayley, or exponential maps preserve orthonormality constraints.

Empirical performance (CPU time, convergence) on high-dimensional embedding and image data strongly favors RBCD and ReALM/iRBBS in practical regimes, with speedups H=Rd\mathcal H=\mathbb R^d9–Π(μ,ν)\Pi(\mu,\nu)0 over older Riemannian methods (Jiang et al., 2022, Huang et al., 2020).

6. Statistical Inference, Applications, and Practical Recommendations

Subspace robust Wasserstein distances have broad applications in generative modeling, domain adaptation, two-sample testing, and minimum-distance parametric inference—especially when data have low intrinsic dimension in high-dimensional ambient spaces (Vasan, 4 Dec 2025, Lin et al., 2020).

Key findings:

  • Minimum PRW estimators are consistent under weak conditions, even under model misspecification. Central limit theorems are established for Π(μ,ν)\Pi(\mu,\nu)1 (max-sliced) (Lin et al., 2020).
  • Optimal Π(μ,ν)\Pi(\mu,\nu)2 should match intrinsic data dimension; Π(μ,ν)\Pi(\mu,\nu)3 (“sliced”) gives optimal rates Π(μ,ν)\Pi(\mu,\nu)4, higher Π(μ,ν)\Pi(\mu,\nu)5 improves geometric fidelity but increases sample and computational complexity.
  • SRW is more computationally tractable via convex relaxation; PRW is statistically more powerful but non-convex.
  • Empirical performance demonstrates robustness to noise and improved clustering for word embedding and high-dimensional data tasks.

Averaged variants, such as Integral PRW (IPRW, based on integration instead of supremum over subspaces), offer smoother, easier-to-estimate models but are less discriminative (Lin et al., 2020).

7. Open Problems and Future Directions

Significant open questions and directions include (Vasan, 4 Dec 2025, Jiang et al., 2022, Lin et al., 2020):

  • Tightening statistical rates (removal of Π(μ,ν)\Pi(\mu,\nu)6 factors, high-probability bounds).
  • Extension to general order Π(μ,ν)\Pi(\mu,\nu)7, requiring Schatten-Π(μ,ν)\Pi(\mu,\nu)8 norm estimates.
  • Minimax optimality under weaker moment or heavy-tailed assumptions.
  • Theory and algorithms for continuous (non-discrete) measures.
  • Acceleration via Riemannian trust-region or momentum techniques.
  • Tuning and automation of entropic regularization parameters.
  • Extensions to barycenter computation, deep generative models, and distributed or streaming data scenarios.
  • Understanding global optimality and landscape of the PRW optimization problem.

These lines of inquiry are central for further development of high-dimensional robust optimal transport and its application to modern statistical and machine learning tasks.


References:

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Subspace Robust Wasserstein Distance.