Papers
Topics
Authors
Recent
Search
2000 character limit reached

Dynamic Measure Transport (DMT)

Updated 9 November 2025
  • Dynamic Measure Transport (DMT) is a framework that models the evolution of probability measures via PDEs, unifying optimal transport, control theory, and mean-field games.
  • It employs both deterministic and stochastic dynamics with tilted reference paths to overcome teleportation issues, thereby enhancing sample quality.
  • The framework integrates kernel-based numerical methods and Gaussian processes to solve the underlying optimal control problems, with applications in generative modeling and Bayesian inference.

Dynamic Measure Transport (DMT) is a mathematical framework unifying dynamic formulations of measure transport, optimal control, and mean-field games. It generalizes classical optimal transport by modeling the evolution of probability measures subject to partial differential equations (PDEs) in both finite- and infinite-dimensional settings, with significant implications for sampling, generative modeling, and gradient flows. DMT encompasses both traditional “smooth” optimal transport and extensions to spaces characterized only by weak geometric or topological structure, including applications in probability, analysis, stochastic differential equations, and computational statistics.

1. Formal Definition and Mathematical Structures

Let X=RdX = \mathbb{R}^d (or an extended metric-topological space), and fix two Borel probability measures η\eta (reference) and π\pi (target). DMT models the evolution {μt}t[0,1]\{\mu_t\}_{t \in [0,1]} of probability measures such that μ0=η\mu_0 = \eta and μ1π\mu_1 \approx \pi, by either deterministic or stochastic dynamics: dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta, where v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d is a drift field, σ0\sigma \geq 0 is a noise parameter, and WtW_t is standard Brownian motion. The law η\eta0 evolves according to:

  • The continuity equation (ODE case, η\eta1):

η\eta2

η\eta4

In more abstract contexts, DMT is defined on an extended metric-topological measure space η\eta5, where η\eta6 may be infinite, and η\eta7 is a Radon probability measure. The Cheeger energy η\eta8 generalizes Dirichlet energy and induces a dynamic transport cost (see (Ambrosio et al., 2015)). The DMT/Wasserstein–Cheeger distance η\eta9 between absolutely continuous measures is given by a Benamou–Brenier-type formula.

2. Reference Paths, Teleportation Phenomena, and Their Limitations

A canonical construction in dynamic measure transport is the geometric annealing (or “annealed”) reference path: π\pi0 with log-density

π\pi1

where π\pi2 is the time-dependent normalization. This choice has mathematical convenience—analytic expressions for time-derivatives facilitate density-driven algorithmic approaches.

However, when π\pi3 and π\pi4 are multimodal or well-separated, π\pi5 exhibits “teleportation.” Most of the mass may abruptly move from one mode to another at a critical π\pi6. As a result, transport velocities π\pi7 must become large or highly irregular, and density-driven learning (aligning a velocity field to π\pi8) often fails to “split” or move mass correctly. Empirically, this is reflected in significant sample quality deficits or mode-dropping in sampling applications, as evidenced in one-dimensional Gaussian mixture experiments (Section 7 below).

3. Optimal Control and Mean-Field Game Perspective

DMT can be naturally framed as an infinite-dimensional optimal control problem or mean-field game (MFG). For a curve of densities π\pi9 and a velocity field {μt}t[0,1]\{\mu_t\}_{t \in [0,1]}0, the following variational problem encapsulates DMT with action, smoothness, and fidelity costs: {μt}t[0,1]\{\mu_t\}_{t \in [0,1]}1 The KL-interaction {μt}t[0,1]\{\mu_t\}_{t \in [0,1]}2 has a unique minimizer given by the geometric annealing path. The formal optimality (first-order) system couples a forward continuity equation with a backward Hamilton–Jacobi–Bellman (HJB) equation, enforcing both fidelity to boundary data and smoothness/action minimization subject to fixed start/end measures.

This MFG/control-theoretic framing enables flexible introduction of path-dependent fidelity and regularization criteria beyond what is available in standard OT formulations.

4. Tilted-Path Formulation and Optimization

To overcome pathologies (e.g., “teleportation”) of standard reference paths, DMT introduces a “tilting function” {μt}t[0,1]\{\mu_t\}_{t \in [0,1]}3. The density path is reparametrized: {μt}t[0,1]\{\mu_t\}_{t \in [0,1]}4 with {μt}t[0,1]\{\mu_t\}_{t \in [0,1]}5 and {μt}t[0,1]\{\mu_t\}_{t \in [0,1]}6, {μt}t[0,1]\{\mu_t\}_{t \in [0,1]}7. The optimal control problem is then

{μt}t[0,1]\{\mu_t\}_{t \in [0,1]}8

subject to the continuity equation {μt}t[0,1]\{\mu_t\}_{t \in [0,1]}9 and constraints on μ0=η\mu_0 = \eta0. The spaces μ0=η\mu_0 = \eta1 typically employ Sobolev or reproducing kernel Hilbert space (RKHS) norms to enforce spatial/temporal smoothness; e.g.,

μ0=η\mu_0 = \eta2

Equivalently, an augmented Lagrangian can be introduced for deriving necessary optimality conditions in mixed PDE form, leading to a flexible Banach-space optimization framework.

5. Numerical Solution via Gaussian Processes and Collocation

For practical computation, the tilted DMT control problem is discretized via a Gaussian process (GP) and collocation approach:

  • Choose a collocation grid μ0=η\mu_0 = \eta3 in μ0=η\mu_0 = \eta4 and select boundary points for μ0=η\mu_0 = \eta5.
  • Model the scalar potential μ0=η\mu_0 = \eta6 (so μ0=η\mu_0 = \eta7) and tilt μ0=η\mu_0 = \eta8 as elements of scalar-valued RKHSs μ0=η\mu_0 = \eta9, μ1π\mu_1 \approx \pi0 with product kernels μ1π\mu_1 \approx \pi1 (typically Matérn-type in space/time with lengthscales μ1π\mu_1 \approx \pi2, μ1π\mu_1 \approx \pi3).
  • Enforce the nonlinear residual of the continuity or Fokker–Planck equation at each interior collocation point, i.e., μ1π\mu_1 \approx \pi4, where μ1π\mu_1 \approx \pi5 collects the needed derivatives and μ1π\mu_1 \approx \pi6 handles time normalization. At boundaries, impose μ1π\mu_1 \approx \pi7.
  • By the representer theorem, the minimizers μ1π\mu_1 \approx \pi8 lie in finite spans of kernel sections evaluated or differentiated at grid points.

The empirical optimization reduces to a penalized least-squares problem: μ1π\mu_1 \approx \pi9 which is solved via trust-region methods such as Levenberg–Marquardt, using Cholesky parameterization of the Gram matrices.

6. Theoretical Results: Representer Theorem and Well-posedness

Representer theorem for DMT collocation: Given a Hilbert space dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta,0 with kernel dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta,1 and a finite set of linear functionals dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta,2, the RKHS minimizer constrained by dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta,3 is: dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta,4 and dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta,5 solves dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta,6. This structure ensures all learned objects admit efficient parametrization in terms of kernel sections induced by collocation.

Existence of minimizers: Under mild assumptions (smooth, positive densities for the reference path; strictly positive regularizers dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta,7) the finite-dimensional penalized least-squares problem is coercive and continuous, guaranteeing the existence of minimizers.

These results provide a rigorous foundation for kernel-based and collocation-based implementations and imply provable smoothness of the optimal transport velocity fields and tilts.

7. Empirical Assessment and Sampling Applications

In one-dimensional experiments with reference dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta,8 and target dXt=v(Xt,t)dt+σdWt,X0η,dX_t = v(X_t, t)\,dt + \sigma\,dW_t,\quad X_0 \sim \eta,9, the geometric annealing path v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d0 fails to transport mass to the leftmost mode—reflecting the teleportation pathology. Learned velocity fields along this path produce samplers missing regions of the target.

The tilted DMT approach (using the Banach-space control framework and GP solver) avoids teleportation and yields smooth, balanced mass transfer into both modes, demonstrated by:

  • Fraction of trajectories capturing the left mode (true = 0.667): reference v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d1 vs. tilt-learned v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d2.
  • Relative error in mean: v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d3 (ref) vs. v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d4 (learned).
  • Relative error in variance: v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d5 (ref) vs. v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d6 (learned).
  • Kernel MMD: v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d7 (ref) vs. v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d8 (learned).
  • The spatial RKHS norm of the learned velocity remains stable under the tilted path, indicating superior regularity.

Trajectory visualizations corroborate the improved spatial smoothness and sampling fidelity achieved by the tilted DMT method relative to analytic McCann velocity or standard reference paths.

Key implementation and application principles include:

  • Regularization: v:Rd×[0,1]Rdv: \mathbb{R}^d \times [0,1] \to \mathbb{R}^d9 balances the scale of the tilt; σ0\sigma \geq 00 and σ0\sigma \geq 01 enforce PDE and boundary condition fidelity.
  • Kernel hyperparameters: σ0\sigma \geq 02 and σ0\sigma \geq 03 govern the spatial/temporal smoothness of the velocity and tilt; higher values yield smoother but less flexible paths.
  • Collocation resolution σ0\sigma \geq 04 drives tradeoffs between accuracy and computational requirements, with worst-case cubic scaling in kernel matrix assembly and linear solver steps. Inducing point or hierarchical strategies can reduce computational cost.
  • DMT is broadly applicable: generative modeling (continuous normalizing flows, diffusion models), density-driven/annealing samplers, Bayesian inference, obstacle-aware robotic transport (via tilt σ0\sigma \geq 05), and finetuning of pretrained generative models.

In the context of non-smooth, infinite-dimensional, or “Wiener-like” spaces, DMT generalizes classical optimal transport and the Otto calculus, leveraging the Cheeger energy and Benamou–Brenier-type dynamic characterizations. Heat semigroups generated by Dirichlet forms—analyzed via the evolution variational inequality (EVI)—admit contractivity and curvature results extending well beyond the σ0\sigma \geq 06-Wasserstein theory (Ambrosio et al., 2015). A plausible implication is that DMT provides a unifying mathematical and computational infrastructure for measure-valued dynamics in both classical and highly singular regimes.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Dynamic Measure Transport (DMT).