Papers
Topics
Authors
Recent
Search
2000 character limit reached

Koopman Spectral Wasserstein Gradient Descent

Updated 28 December 2025
  • KSWGD is a training-free, particle-based generative modeling method that leverages Koopman spectral techniques and optimal transport to drive particles toward unknown target distributions.
  • It approximates the inverse Langevin generator via a data-driven, finite-rank spectral surrogate, ensuring constant dissipation rates and accelerated convergence even in high dimensions.
  • Empirical results show that KSWGD achieves full support coverage and linear KL decay across benchmarks like S¹ uniform sampling, quadruple well, and MNIST latent generation.

Koopman Spectral Wasserstein Gradient Descent (KSWGD) is a training-free, particle-based generative modeling methodology uniting operator-theoretic spectral analysis with variational optimal transport theory. At its core, KSWGD leverages trajectory or time-series data to approximate the infinitesimal generator of overdamped Langevin dynamics through Koopman spectral techniques, and subsequently drives particles deterministically along a preconditioned Wasserstein gradient flow toward an unknown target distribution, achieving accelerated convergence without explicit knowledge of the target potential or reliance on neural network training (Xu et al., 21 Dec 2025).

1. Definition and Theoretical Framework

KSWGD targets generative sampling problems where only samples (possibly arranged temporally) from an unknown distribution π(x)∝exp⁡(−V(x))\pi(x) \propto \exp(-V(x)) are available. The goal is to transform an initial empirical measure μ0\mu_0 to approximate π\pi by discretizing the χ2\chi^2–Wasserstein gradient flow:

∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}

Here, κ\kappa corresponds to the inverse Langevin generator L−1L^{-1}. The key innovation of KSWGD is to approximate κ\kappa in a fully data-driven and finite-rank manner using Koopman spectral methods (such as Extended Dynamic Mode Decomposition, EDMD), resulting in a preconditioned flow with constant dissipation rate even in high dimensions. This approach operationalizes the same mathematical foundation as Laplacian-Adjusted Wasserstein Gradient Descent (LAWGD) but circumvents the need for target potential access or score network training.

The algorithm proceeds as follows:

  • Estimate the leading rr eigenpairs of the generator from time-ordered trajectory pairs.
  • Build a truncated spectral surrogate for the inverse Langevin operator.
  • Update MM particles along this Koopman-preconditioned Wasserstein gradient flow.

2. Mathematical Formulation and Algorithm

Wasserstein Gradient Flow and Spectral Preconditioning

The evolution of μ0\mu_00 under the μ0\mu_01–Wasserstein gradient flow in μ0\mu_02 is characterized by the velocity field

μ0\mu_03

with preconditioning (as in LAWGD) by μ0\mu_04 leading to:

μ0\mu_05

where μ0\mu_06 are eigenpairs of the Langevin generator μ0\mu_07 (self-adjoint on μ0\mu_08).

Koopman Spectral Approximation

Recognizing that the Koopman (backward Kolmogorov) generator μ0\mu_09, KSWGD uses data-driven spectral approximation from trajectory pairs π\pi0 by EDMD to identify leading π\pi1 empirical eigenpairs π\pi2 of π\pi3, yielding the truncated inverse:

π\pi4

Discrete Particle Update

Given π\pi5 particles π\pi6, a single KSWGD step of size π\pi7 is:

π\pi8

where

π\pi9

and χ2\chi^20 denotes gradient with respect to the first argument.

Pseudocode

The methodology splits naturally into Offline and Online phases:

Step Description Computational Cost
Offline (Spectral Est.) Dictionary selection, data matrix formation, EDMD eigenproblem, obtain leading χ2\chi^21 χ2\chi^22 (for dictionary size χ2\chi^23)
Online (Updates) Iteratively update particles using Koopman-preconditioned flow χ2\chi^24 per iteration

Auxiliary notes:

  • Computing eigenfunction gradients χ2\chi^25 may require analytic expressions or finite-difference approximations.
  • The bias-variance trade-off is controlled by rank χ2\chi^26 (truncation) and step size χ2\chi^27.

3. Theoretical Properties

Spectral Preconditioning and Dissipation

The preconditioned flow with truncated spectral surrogate χ2\chi^28 satisfies:

χ2\chi^29

with dissipation identity:

∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}0

where ∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}1 is projection onto the retained ∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}2-dimensional eigenspace.

Under mild regularity and a tail bound ∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}3, the ideal convergence rate is:

∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}4

Data-Driven Error Bounds

If the Koopman spectral approximation error meets:

∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}5

then

∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}6

Discrete-time convergence in the Approximate Gradient Flow (AGF) setting yields, for step size ∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}7:

∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}8

implying geometric decay to bias ∂tμt=div⁡(μt∇κ(ρt)),ρt=dμtdπ\partial_t \mu_t = \operatorname{div} \left( \mu_t \nabla \kappa(\rho_t) \right), \quad \rho_t = \frac{d\mu_t}{d\pi}9 modulated by κ\kappa0.

Feynman–Kac Perspective

With κ\kappa1 in the Feynman–Kac formula,

κ\kappa2

the Koopman semigroup encodes unconditional observable expectations, underlining the probabilistic foundation of KSWGD sampling. Extension to κ\kappa3 would encompass conditional or rare-event inference.

4. Experimental Validation and Benchmarking

KSWGD's empirical performance was examined across diverse generative modeling milieus:

Task/Dataset Particles Koopman Method Key Metric(s)
Sκ\kappa4 uniform sampling 700 kernel-EDMD (RBF/poly) KL decay, movement rate, coverage
Quadruple well 500 SDMD (neural dict.) KL, well coverage, movement rate
MNIST (latent) 64 CNN+EDMD dict. learning Visual sample, KL, κ\kappa5 divergence
Allen–Cahn SPDE 150 EDMD (poly features) Visual fidelity, distributional prediction

Baselines examined include DMPS (diffusion maps), DDPM, VAE, RealNVP, WGAN-GP. Notable empirical results:

  • On Sκ\kappa6–uniform and quadruple well, KSWGD achieved full support coverage in κ\kappa7 iterations, while DMPS required κ\kappa8.
  • Empirical KL decay was linear, validating theoretical predictions.
  • On MNIST latent code, KSWGD produced discernible digit samples; DMPS failed under parallel conditions.
  • On Allen–Cahn SPDE, KSWGD matched or surpassed DDPM, VAE, normalizing flows, and GANs in latent space sample quality.

5. Computational and Practical Considerations

Key computational and modeling constraints include:

  • Offline spectral estimation incurs κ\kappa9 eigendecomposition, while each online step scales as L−1L^{-1}0.
  • Gradients of eigenfunctions L−1L^{-1}1 may demand basis-specific analytic or finite-difference evaluation.
  • Bias-variance trade-offs hinge on truncation rank L−1L^{-1}2 (reducing L−1L^{-1}3 at increased computational cost) and step size L−1L^{-1}4 (smaller L−1L^{-1}5 reduces L−1L^{-1}6 discretization bias but necessitates longer runs).
  • Quality of latent autoencoders or dictionary selection directly bounds generative accuracy in high-dimensional settings.
  • Assumptions of ergodicity and self-adjointness of the generator are essential; extensions to non-reversible dynamics necessitate additional research.

6. Connections, Scope, and Limitations

KSWGD synthesizes the spectral guarantees of LAWGD with the data-driven practicality of EDMD, providing a theoretically justified and computationally accessible recipe for generative particle-based sampling. It eliminates the need for explicit potential evaluation or neural-network-based score learning. Current limitations encompass the necessity for high-quality spectral/dictionary approximations, computational scaling with rank and particle count, and the restriction to detailed-balance (oscillating) dynamical systems. Future work may address extensions to irreversible (non-detailed-balance) generators and autonomous dictionary adaptation (Xu et al., 21 Dec 2025).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Koopman Spectral Wasserstein Gradient Descent (KSWGD).