Papers
Topics
Authors
Recent
Search
2000 character limit reached

Cumulative Wasserstein Drift

Updated 2 January 2026
  • Cumulative Wasserstein drift is the sum of Wasserstein distances between successive probability measures, quantifying total distributional change over time.
  • It underpins finite-sample guarantees, dynamic regret bounds, and concentration inequalities in stochastic processes and online optimization.
  • Optimal weighting schemes leveraging cumulative drift balance bias and variance, guiding parameter choices in nonstationary and robust methods.

Cumulative Wasserstein drift quantifies the total “distance” traversed by a time-evolving sequence of probability measures, typically under the Wasserstein metric, over a given time horizon. It is a central nonstationarity measure in stochastic processes, online optimization, Markov dynamics, and empirical process theory, capturing both instantaneous and aggregated changes in distributions. Rigorous frameworks for cumulative Wasserstein drift underpin concentration inequalities, dynamic regret bounds, convergence rates for flows in the space of measures, and finite-sample guarantees in distributionally robust optimization.

1. Definition and Foundational Concepts

Let {Pt}t=1T\{P_t\}_{t=1}^T be a sequence of probability measures on a Polish space Ξ\Xi. The pp-Wasserstein distance between PtP_t and Pt+1P_{t+1} at each time tt is given by

Δt:=Wp(Pt,Pt+1)\Delta_t := W_p(P_t, P_{t+1})

The unweighted cumulative Wasserstein drift over TT periods is

DT:=t=1T1ΔtD_T := \sum_{t=1}^{T-1} \Delta_t

This sum captures the total geometric “movement” of the underlying data-generating law as measured in the Wasserstein space. In settings with weighted empirical estimators or time-decayed observations, the natural generalization is the LpL_p-norm-type drift: Ξ\Xi0 where Ξ\Xi1 is a vector of nonnegative weights, and Ξ\Xi2 is a uniform bound on Ξ\Xi3 (Keehan et al., 21 Oct 2025).

2. Weighted Empirical Measures and Effective Sample Size

In nonstationary environments, weighted empirical measures are used to balance effective sample size against the impact of distributional drift: Ξ\Xi4 with Ξ\Xi5. A key metric is the effective sample size

Ξ\Xi6

which quantifies the statistical reliability of Ξ\Xi7 under the weighting scheme Ξ\Xi8. The interplay between Ξ\Xi9 and pp0 is critical for controlling estimation error and variance in time-evolving data (Keehan et al., 21 Oct 2025).

3. Finite-Sample Concentration and Nonstationary Robustness

A central technical result is a concentration inequality for Wasserstein distances in the nonstationary, weighted setting: pp1 where pp2 and pp3 depend on the geometry of pp4 and pp5 (Keehan et al., 21 Oct 2025). For sufficiently large pp6,

pp7

This quantifies deviations of the empirical process in the presence of cumulative nonstationary drift, explicitly balancing sample variance and drift-induced bias.

4. Optimal Weighting: Variance–Drift Tradeoff

Optimal weights pp8 simultaneously control bias due to drift and estimation variance, solving

pp9

The unique structure of the solution is

PtP_t0

with scalars PtP_t1 determined by simplex constraints. As PtP_t2 grows, the optimal scheme exhibits sharper cutoff of past (older) data, reducing to pure sliding-window or exponential-decay weighting depending on parameter choices. Explicit calibrations,

PtP_t3

arise for windowing and exponential smoothing in the PtP_t4 case, providing optimal parameter choices in terms of desired accuracy PtP_t5, drift bound PtP_t6, and Wasserstein order PtP_t7 (Keehan et al., 21 Oct 2025).

5. Cumulative Drift in Dynamic Optimization and Learning

In online convex optimization where objective distributions PtP_t8 evolve, the cumulative Wasserstein drift PtP_t9 enters directly into dynamic regret bounds: Pt+1P_{t+1}0 and the sequence of minimizers Pt+1P_{t+1}1 satisfies

Pt+1P_{t+1}2

The corresponding dynamic regret is lower-bounded by an Pt+1P_{t+1}3 term—cumulative drift sets the intrinsic limit on performance in adapting to distributional changes, with all other regret contributions (noise, initialization) being controllable via algorithmic parameters (Shames et al., 2020).

6. Wasserstein Drift in PDEs, Stochastic Flows, and Markov Chains

The notion of cumulative Wasserstein drift generalizes to continuous-time measure-valued flows:

  • For gradient flows in Pt+1P_{t+1}4,

Pt+1P_{t+1}5

where Pt+1P_{t+1}6 is the instantaneous velocity field from the continuity equation. The total path-length controls convergence rates and is uniformly bounded in terms of the initial suboptimality of the functional Pt+1P_{t+1}7 (Chizat et al., 16 Jul 2025).

  • In measure-valued SPDEs and diffusions, the time integral of instantaneous drift or squared gradient quantifies both cumulative displacement in Wasserstein space and the action or Fisher information over time (Delarue et al., 2024).
  • In discrete-time Markov chains, geometric contractivity plus one-step non-contractive “drift” yields cumulative bounds:

Pt+1P_{t+1}8

where Pt+1P_{t+1}9 is the per-step drift and tt0 is the contraction rate. The second term encodes the aggregated perturbation—the “cumulative drift” of the Markov process (Madras et al., 2011).

7. Applications and Broader Significance

Cumulative Wasserstein drift is a central concept in:

These frameworks provide precise nonasymptotic characterizations, parameter choices for weighting schemes, and convergence rates that systematically account for nonstationarity and time-varying complexity in modern stochastic systems.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Cumulative Wasserstein Drift.