Papers
Topics
Authors
Recent
Search
2000 character limit reached

Energy-Based Sliced Wasserstein (EBSW)

Updated 17 June 2026
  • Energy-Based Sliced Wasserstein (EBSW) is a robust metric that leverages an energy function to adaptively weight projection directions for comparing probability measures.
  • It employs Monte Carlo methods like importance sampling, SIR, and MCMC to focus on the most discriminative slices while maintaining computational efficiency.
  • Empirical evaluations demonstrate that EBSW improves convergence, reconstruction quality, and sample efficiency in applications such as gradient flows and deep point-cloud reconstruction.

The Energy-Based Sliced Wasserstein (EBSW) distance is a statistically robust and computationally efficient metric for comparing probability measures, extending the classical Sliced Wasserstein (SW) approach by leveraging an adaptive, parameter-free slicing distribution based on an energy function of the projected Wasserstein distance. By up-weighting projection directions that are most discriminative between measures, EBSW emphasizes informative slices and mitigates the limitations of both uniform and parametric slicing distributions. This formulation preserves the computational advantages of SW while improving statistical signal and sample efficiency across a variety of tasks (Nguyen et al., 2023).

1. Definition and Motivation

Given probability measures μ,νPp(Rd)\mu, \nu \in \mathcal{P}_p(\mathbb{R}^d), the classical pp-th Sliced Wasserstein distance is:

SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}

where θ#μ\theta_\#\mu denotes the push-forward of μ\mu by the linear projection xθxx \mapsto \theta^\top x, and WpW_p is the 1D Wasserstein distance.

EBSW generalizes this framework by replacing the uniform distribution over projection directions θ\theta with an energy-based density

σμ,ν(θ;f,p)f(Wpp(θ#μ,θ#ν))\sigma_{\mu, \nu}(\theta; f, p) \propto f(W_p^p(\theta_\#\mu, \theta_\#\nu))

with f ⁣: ⁣[0,) ⁣ ⁣(0,)f\!:\! [0, \infty)\!\to\! (0,\infty) typically monotonic (e.g., pp0 for increasing pp1 or pp2).

The EBSW metric is then defined as

pp3

This reweighting causes sampling to concentrate on projections where the differences between pp4 and pp5 are most pronounced, thus accentuating statistically informative dimensions and providing higher signal-to-noise ratios for tasks such as gradient flow or generative modeling.

2. Theoretical Properties

EBSW retains the core mathematical properties expected of a metric-like divergence for probability measures:

  • Semi-metricity: For every pp6 and any strictly positive energy pp7, EBSW satisfies:
    • Non-negativity: pp8,
    • Symmetry: pp9,
    • Identity of indiscernibles: SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}0 if and only if SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}1.
  • Relations to SW, Max-SW, and SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}2: If SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}3 is non-decreasing, SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}4, with equality for constant SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}5. For any SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}6, SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}7.
  • Topology: The topology induced by SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}8 is equivalent to weak convergence plus convergence of SWp(μ,ν)=[EθUniform(Sd1)Wpp(θ#μ,θ#ν)]1/pSW_p(\mu, \nu) = \left[ \mathbb{E}_{\theta \sim \text{Uniform}(\mathbb{S}^{d-1})} W_p^p(\theta_\#\mu, \theta_\#\nu) \right]^{1/p}9-th moments, matching the topology induced by the Wasserstein metric.
  • Sample Complexity: For empirical measures θ#μ\theta_\#\mu0 with θ#μ\theta_\#\mu1 i.i.d. in a compact set, there exists θ#μ\theta_\#\mu2 such that

θ#μ\theta_\#\mu3

This mirrors the θ#μ\theta_\#\mu4 sample complexity of SW, avoiding the curse of dimensionality (up to logarithmic factors).

3. Monte Carlo Algorithms for EBSW Computation

For discrete measures supported on at most θ#μ\theta_\#\mu5 atoms, the computation of 1D Wasserstein distances costs θ#μ\theta_\#\mu6. EBSW admits several efficient Monte Carlo estimators:

  • Importance Sampling (IS): Draw θ#μ\theta_\#\mu7 i.i.d. θ#μ\theta_\#\mu8 (e.g., uniform on θ#μ\theta_\#\mu9). Compute weights μ\mu0, where μ\mu1. Normalize and form the μ\mu2 estimator:

μ\mu3

where μ\mu4.

  • Sampling-Importance-Resampling (SIR): After IS on μ\mu5 projections, resample μ\mu6 according to the normalized weights and output the average of the corresponding μ\mu7 values.
  • Markov Chain Monte Carlo (MCMC):
    • Independent MH uses a uniform proposal over μ\mu8.
    • Random-Walk MH employs a von Mises–Fisher proposal μ\mu9 with specified concentration. Both achieve xθxx \mapsto \theta^\top x0 per step.

These approaches offer asymptotically unbiased estimators for xθxx \mapsto \theta^\top x1 as xθxx \mapsto \theta^\top x2. Overall complexity is xθxx \mapsto \theta^\top x3 for xθxx \mapsto \theta^\top x4 projections.

4. Comparative Analysis of Slicing Strategies

Distinct approaches for selecting the slicing distribution yield different computational–statistical tradeoffs:

Approach Slicing Distribution Informative Directions Computational Cost
Classical SW Uniform on xθxx \mapsto \theta^\top x5 No xθxx \mapsto \theta^\top x6
Parametric optimizer Optimized parametric family Yes, but limited by parametric family Often expensive, unstable
Energy-based (EBSW) xθxx \mapsto \theta^\top x7 Yes (data-driven, nonparametric) xθxx \mapsto \theta^\top x8

EBSW uniquely provides a parameter-free, nonparametric adaptation in slicing, focusing computation on directions with the largest observed projected divergences. This yields robustness to misspecification that can hinder parametric approaches and maintains efficiency comparable to classical SW.

5. Empirical Evaluation

All reported experiments use xθxx \mapsto \theta^\top x9.

  • Point-Cloud Gradient Flows: For driving WpW_p0 to a fixed WpW_p1 using the Euler discretization of gradient flow, IS-EBSWWpW_p2 with WpW_p3 achieves the fastest convergence in WpW_p4, outperforming SWWpW_p5, Max-SWWpW_p6, and v-DSWWpW_p7 within comparable runtime constraints.
  • Color Transfer: Modeling images as empirical RGB distributions, IS-EBSWWpW_p8 generates transfers that most closely approximate target colors in WpW_p9 while incurring computational cost nearly equal to SWθ\theta0.
  • Deep Point-Cloud Reconstruction: Training a point-cloud autoencoder with IS-EBSWθ\theta1 as the reconstruction loss yields lower slice-θ\theta2 and true θ\theta3 errors compared to SWθ\theta4, Max-SWθ\theta5, and v-DSWθ\theta6 across epochs on held-out ModelNet40. Resulting reconstructions possess increased sharpness and fidelity.

| Method | Epoch 20 (SWθ\theta7, θ\theta8 ×100) | Epoch 100 (SWθ\theta9, σμ,ν(θ;f,p)f(Wpp(θ#μ,θ#ν))\sigma_{\mu, \nu}(\theta; f, p) \propto f(W_p^p(\theta_\#\mu, \theta_\#\nu))0 ×100) | Epoch 200 (SWσμ,ν(θ;f,p)f(Wpp(θ#μ,θ#ν))\sigma_{\mu, \nu}(\theta; f, p) \propto f(W_p^p(\theta_\#\mu, \theta_\#\nu))1, σμ,ν(θ;f,p)f(Wpp(θ#μ,θ#ν))\sigma_{\mu, \nu}(\theta; f, p) \propto f(W_p^p(\theta_\#\mu, \theta_\#\nu))2 ×100) | |-------------|-------------------------------|-------------------------------|-------------------------------| | SWσμ,ν(θ;f,p)f(Wpp(θ#μ,θ#ν))\sigma_{\mu, \nu}(\theta; f, p) \propto f(W_p^p(\theta_\#\mu, \theta_\#\nu))3 | 2.97, 12.67 | 2.29, 10.63 | 2.15, 9.97 | | Max-SWσμ,ν(θ;f,p)f(Wpp(θ#μ,θ#ν))\sigma_{\mu, \nu}(\theta; f, p) \propto f(W_p^p(\theta_\#\mu, \theta_\#\nu))4 | 2.91, 12.33 | 2.24, 10.40 | 2.14, 9.84 | | v-DSWσμ,ν(θ;f,p)f(Wpp(θ#μ,θ#ν))\sigma_{\mu, \nu}(\theta; f, p) \propto f(W_p^p(\theta_\#\mu, \theta_\#\nu))5 | 2.84, 12.64 | 2.21, 10.52 | 2.07, 9.81 | | IS-EBSWσμ,ν(θ;f,p)f(Wpp(θ#μ,θ#ν))\sigma_{\mu, \nu}(\theta; f, p) \propto f(W_p^p(\theta_\#\mu, \theta_\#\nu))6 | 2.68, 11.90 | 2.18, 10.27 | 2.04, 9.69 |

Qualitative assessment indicates EBSW consistently produces smoother trajectories and reconstructions relative to comparator methods.

6. Practical Implications and Availability

EBSW generalizes SW by replacing uniform slicing with an energy-based density proportional to a strictly positive function of 1D Wasserstein projections. This maintains σμ,ν(θ;f,p)f(Wpp(θ#μ,θ#ν))\sigma_{\mu, \nu}(\theta; f, p) \propto f(W_p^p(\theta_\#\mu, \theta_\#\nu))7 per slice computational complexity and introduces greater adaptability by focusing on the most discriminative projections.

Monte Carlo estimators—via importance sampling, SIR, or MCMC—are readily implementable. IS-EBSWσμ,ν(θ;f,p)f(Wpp(θ#μ,θ#ν))\sigma_{\mu, \nu}(\theta; f, p) \propto f(W_p^p(\theta_\#\mu, \theta_\#\nu))8 empirically achieves lower reconstruction and transport errors compared with baseline SW variants, at only a minor increase in computational overhead. All supporting code and data are available at https://github.com/khainb/EBSW (Nguyen et al., 2023).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Energy-Based Sliced Wasserstein (EBSW).