Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Temporal Spatial Minimax Rate for Smoothly-Varying Distributions in Wasserstein Space

Published 5 Jun 2026 in math.ST, cs.AI, and cs.IT | (2606.07325v1)

Abstract: We study the minimax rate of estimating a future value $μ_{t_n+h}$ of a curve $t\mapstoμ_t$ in the $2$-Wasserstein space $\mathcal{P}_2(\mathbb{R}d)$ from finitely many noisy snapshots of its past, under an adiabatic bound $|\nabla_tk v|\le\varepsilon$ on the $k$-th covariant derivative of the velocity field. Our central result is a unified temporal-spatial minimax lower bound: over regular, locally transport-rich subclasses, every estimator incurs $W_2$-risk with $M$-exponent $γ_d(k+1)/(k+1+γ_d)$, $γ_d=\min(1/d,1/2)$ ($M$ the total sample size). It follows from a temporal-to-spatial reduction: the smoothness budget defines a reachable $W_2$-ball into which a transport packing is embedded along the time axis, and the information of the entire snapshot experiment is controlled by a Fano argument -- the spatial packing is classical, but its smoothness-admissible temporal embedding and the full-window analysis are new. The bound interpolates a dimension-free extrapolation floor of order $\varepsilon h{k+1}$ -- the irreducible cost of an unobserved future, present even with the exact past -- and the spatial estimation curse $M{-γ_d}$, recovering the static distribution-estimation rate as $k\to\infty$. We state the lower bound in a design-dependent form -- with a design-weighted effective sample size -- valid for arbitrary observation times, and obtain the closed-form exponent in the dense (equispaced) regime. The matching upper bound is established at $k=0$ (rate $M{-1/(d+1)}$, $d\ge3$) and, in a translation submodel, for all $k$; for $k\ge1$ a covariant estimator attains the rate conditionally on two estimates (a comparison-geometry bias bound and an optimal-transport map-estimation rate), leaving the unconditional general-$k$ upper bound as an open problem. Numerical experiments on synthetic curved and flat families corroborate the predicted exponents.

Authors (1)

Summary

  • The paper demonstrates that even with complete past data, forecasting error scales as εh^(k+1) due to inherent temporal extrapolation limits.
  • It quantifies spatial estimation error, showing that finite sample size M imposes a minimax rate of M^(-γ_d) with γ_d = min(1/d, 1/2), highlighting the curse of dimensionality.
  • A unified bound is derived by combining temporal and spatial errors, offering a framework for principled forecasting of distribution-valued dynamical systems.

Minimax Forecastability for Smoothly-Varying Distributions in Wasserstein Space

Problem Motivation and Framework

The paper addresses the fundamental question of the predictive limits for an evolving probability measure-valued process t↦μtt \mapsto \mu_t in Wasserstein-2 space PP: given finitely many noisy samples from past distributions, how accurately can one predict a future distribution at time tn+ht_n + h? This problem arises in econometrics, population dynamics, and time series analysis where distributions, not just point states, evolve—examples include financial returns, cell populations, and developmental gene expression.

A key technical challenge is the interplay between temporal regularity—formalized as a bound on the kk-th covariant derivative ∥∇tkv∥≤ε\|\nabla_t^k v\| \leq \varepsilon of the velocity in Wasserstein space—and statistical estimation—the curse of dimensionality intrinsic to the Wasserstein distance over distributions on Rd\mathbb{R}^d. The regularity assumption defines a "slowly-varying" or adiabatic class CC akin to scalar function extrapolation, allowing for increasingly smooth curves as kk increases.

Minimax Lower Bounds: Temporal and Statistical Floors

The authors first establish two types of irreducible forecast errors, which are always present, regardless of estimator.

Worst-Case (Le Cam) Extrapolation Floors

The first is a dimension-free temporal floor: even with complete knowledge of the entire past trajectory, the error from forecasting hh units ahead scales as εhk+1\varepsilon h^{k+1}, reflecting the Taylor/geodesic extrapolation remainder. Neither infinite past data nor sample size can improve on this floor (Figure 1). Figure 1

Figure 1: Horizon exponent survives curvature—integer-order Taylor/geodesic extrapolation errors are preserved under nontrivial Wasserstein manifold curvature.

Sample-Limited Statistical Floors

The second floor is imposed by finite sample size PP0: when estimating a PP1-dimensional measure in Wasserstein-2, the minimax rate is PP2, PP3. For large PP4, this rate degrades dramatically, reflecting the curse of dimensionality. These two limits—temporal extrapolation and spatial estimation—act independently (Figure 2, Figure 3). Figure 2

Figure 2: PP5 phase diagram—shows the division between extrapolation-limited and statistics-limited regimes; the white curve delineates the two.

Figure 3

Figure 3: Sharp nonparametric extrapolation rate—optimal bandwidth balancing gives the PP6 rate, generalizing H\"older nonparametric estimation to the forecasting domain.

Unified Temporal–Spatial Minimax Rates

The central technical result is a unified lower bound, rigorously derived via a temporal-to-spatial reduction (Lemma~\ref{lem:reach}): the class PP7 of curves is exploited to embed a "packing" of measures into a reachable Wasserstein ball along the time axis, and the sample information is controlled using Fano and Le Cam inequalities (Appendix~\ref{app:reduction}). This yields, for regular, locally transport-rich Wasserstein subclasses, the minimax forecast rate:

PP8

where PP9. As tn+ht_n + h0, the spatial estimation rate tn+ht_n + h1 is recovered, i.e., static density estimation, and as tn+ht_n + h2 decreases or tn+ht_n + h3 increases, the rate improves (Figure 4). Figure 4

Figure 4: Unified rate over tn+ht_n + h4—the observed tn+ht_n + h5-exponent closely tracks the theoretical prediction, with sharp regime transitions due to sampling and temporal smoothness.

Sample Design and Theoretical Phase Structure

A detailed analysis in the translation (location) model yields an analog of minimax H\"older rates: tn+ht_n + h6, matching classical nonparametric boundary estimation—this showcases how the forecasting problem generalizes nonparametric theory (Figure 3).

The authors formalize the statistical leverage achieved by pooling samples within an "optimal" backward-looking time window tn+ht_n + h7 preceding tn+ht_n + h8, optimizing the bias-variance trade-off via the effective sample size allocated by the observation design.

The theory predicts a sharp phase structure: if the forecast horizon tn+ht_n + h9 is small, statistical limits dominate, and temporal smoothness plays little role. For large kk0, extrapolation errors dominate and the estimation cost saturates at the irreducible future uncertainty. This is evidenced both in synthetic and real-data experiments (Figure 5, Figure 6). Figure 5

Figure 5: Held-out predictive validation—a bias–variance model calibrated on training data successfully predicts the out-of-sample error minimum.

Figure 6

Figure 6

Figure 6: Real-data S&P 500 experiment—degree-0 (persistence) outperforms higher-degree forecasters due to high noise; the moving forecast floor lies well above the finite-sample noise level.

Achievability and Remaining Open Problems

For kk1, a pooled-persistence estimator—empirically measuring the target using all samples taken within the optimal window—attains the lower bound. For kk2, similar upper bounds are shown in translation/Gaussian submodels using local polynomial geodesic regression; however, the fully general case remains open except conditionally (Figure 7). The technical obstruction is the bias analysis in curved Wasserstein geometry, requiring strong regularity and curvature control. Figure 7

Figure 7: Strongly-drifting temperature field—distinct optimal bandwidth denotes true drift; fitted exponents clearly distinguish between smooth (temperature) and noisy (S&P) regimes.

Conditional Upper Bounds and Open Directions

Matching upper bounds for general kk3 rely on two conditional estimates: a comparison-geometry bias control, and a minimax-optimal map-estimation rate for the statistical term. These are accessible on flat submanifolds or with smooth densities, but not established for the full class of locally transport-rich Wasserstein curves, due to the complexities of second-order Otto calculus and curvature behavior.

Notably, the minimax lower bound is unconditional and relies only on elementary packing and probabilistic arguments. The match between lower and upper rates for kk4 and for all kk5 on translation/Gaussian submodels is numerically confirmed (Figure 4).

Implications, Practical Recommendations, and Future Directions

The results delimit, for the first time, the precise interplay between smoothness-driven extrapolation and finite-sample estimation error for distributional time series in Wasserstein space. In practical terms:

  • For high-dimensional (kk6) or weakly regular (kk7 small) problems, sample size offers little statistical improvement beyond a short forecast horizon;
  • For strongly regular dynamics or low dimension, significant pooling is justified and longer-horizon distributional prediction is feasible if smoothness can be certified empirically;
  • The effective optimal forecasting order is data-dependent and can be adaptively selected based on observed hold-out behavior (Figure 5, Figure 6, Figure 7);
  • Applying higher-order (e.g., geodesic/spline) extrapolation without evidence of strong smoothness can amplify sampling noise and degrade performance on real data (Figure 6).

Theoretical and Methodological Impact

The paper rigorously extends nonparametric minimax theory and temporal extrapolation into the infinite-dimensional, geometrically nontrivial setting of Wasserstein space. It develops considerable technical machinery sharpened to disentangle time and space limitations, providing a natural framework for future analyses of statistical learning in distribution-valued dynamical systems.

Outstanding open problems include establishing curvature-stable, unconditional upper bounds in general Wasserstein submanifolds, understanding minimax-optimal transport map estimation rates for complex support, and developing computationally feasible procedures that can adapt to unknown smoothness, dimension, and memory constraints.

Conclusion

This work establishes a sharp, unified minimax rate for the statistical forecast of slowly-varying distributions in Wasserstein-2 space, separating the irreducible limits imposed by temporal smoothness from those arising out of finite-sample spatial estimation. Both theoretical analysis and empirical evidence confirm the presence and tight interaction of these regimes, offering a blueprint for principled forecasting—and for understanding the fundamental limitations—when the objects of inference are distributions that move in time. The open technical questions about curvature, statistical estimation of transport maps, and fully adaptive methods present promising directions for future theoretical and computational advances.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.