- The paper establishes a uniform central limit theorem showing that empirical gradient-flow fluctuations converge to a Gaussian process over the entire time horizon, including adaptive and diverging stopping times.
- The method uses a pathwise linearization and a finite $L^2(P)$ arc length condition for sensitivity functions to control an infinite-time empirical process, while accommodating nonsmooth losses and nonconvex objectives.
- The paper introduces an algorithm-aware covariance estimator updated alongside optimization, achieving a uniform $\mathcal{O}_P\{(\log n/n)^{1/2}\}$ consistency rate and producing near-nominal coverage in simulations.
- What assumptions ensure the sensitivity-function path has finite arc length and the time-uniform central limit theorem applies?
- How does the proposed covariance estimator compare with bootstrap or sandwich methods for fixed-time M-estimators?
- Under what conditions can confidence intervals target the population optimum rather than the early-stopped population trajectory?
- How might the theory change for discrete gradient descent, stochastic gradient methods, proximal algorithms, or mirror descent?
- Find recent papers about statistical inference for optimization algorithms.
Overview
The paper develops a time-uniform theory of statistical inference for gradient flows arising from empirical risk minimization. Whereas classical asymptotic analysis of algorithmic estimators concerns a fixed iteration or the terminal M-estimator, the authors analyze the fluctuation process n1/2(θ^−θ∘) over the entire interval [0,∞), where θ^ is the empirical gradient flow and θ∘ its population counterpart. The main results are (i) a uniform central limit theorem in L∞([0,∞);Rd) characterizing the trajectory's deviation as a Gaussian process, and (ii) an algorithm-aware covariance estimator that is uniformly consistent over time and supports Wald-type confidence intervals at deterministic, data-dependent, and diverging stopping times. The framework accommodates non-smooth losses such as quantile regression and non-convex objectives such as phase retrieval.
Problem setup and pathwise decomposition
Given i.i.d. observations from P with empirical measure Pn, the empirical gradient flow solves θ^˙(t)=−Ψn(θ^(t)) with Ψn=Pnψθ, while the population flow θ∘ solves the same equation with [0,∞)0. The paper replaces the classical terminal decomposition into optimization gap plus sampling error by the pathwise decomposition
[0,∞)1
separating stochastic fluctuation from deterministic early-stopping bias. A baseline comparison bound (Proposition 1) shows that on convexity events the deviation is controlled by an integral against the empirical process along the population trajectory; however, without strong curvature the prefactor [0,∞)2 diverges, so this bound yields only compact-time consistency and motivates the sharper analysis. Existence and uniqueness of flows are guaranteed even for discontinuous vector fields via the Brézis–Komura theorem, covering nonsmooth criteria.
Linearization and the arc-length Donsker argument
The core structural result is a linearization lemma: after variation of constants,
[0,∞)3
where the sensitivity function [0,∞)4 accumulates gradients propagated by the principal fundamental matrix [0,∞)5 of the linearized dynamics. The leading term is thus an empirical process indexed by the one-parameter class [0,∞)6.
The central technical device is a new sufficient condition for the Donsker property of one-parameter function classes: finiteness of the [0,∞)7-arc length [0,∞)8. This extends bracketing arguments to unbounded index sets: finite arc length bounds bracketing numbers linearly, so the bracketing entropy integral is finite. The authors verify the condition under three assumptions—exponential convergence of [0,∞)9 to θ^0 (implied by Polyak–Łojasiewicz gradient dominance), local Lipschitz continuity of gradients and Hessian, and eventual coercivity of θ^1—and prove that the sensitivity path has bounded arc length with an explicit constant. Notably, regularity is imposed only on the averaged field θ^2, not pointwise on θ^3, which permits non-smooth and non-convex empirical objectives.
A bootstrap-style argument for differential equations (successive sharpening of a priori bounds, using continuity of the normalized error to rule out escape through an intermediate level) shows that the higher-order remainders—a doubly centered gradient increment and the Bregman divergence of θ^4—are negligible uniformly in time whenever θ^5, θ^6, and the leading term is tight. Combined with the Gaussian asymptotics of the leading term, this yields the main theorem: θ^7 converges weakly in θ^8 to a zero-mean Gaussian process with covariance
θ^9
The limiting process is governed not by the ambient parameter space but by the one-dimensional path traced by the population flow—this low-complexity geometry is what makes uniform control over an infinite horizon possible. As corollaries, inference is valid at data-dependent stopping times θ∘0 in probability, and, provided θ∘1 exists with stochastic equicontinuity at infinity, at diverging stopping times. Inference for the target θ∘2 additionally requires negligible population bias, i.e., θ∘3.
Algorithmic covariance estimation
For practical inference the covariance function must be estimated. The oracle plug-in estimator based on θ∘4 is infeasible, but θ∘5 satisfies a linear evolution equation driven by the population Hessian. The proposed estimator θ∘6 replaces θ∘7 by θ∘8, the solution of the corresponding empirical dynamics using a plug-in Hessian θ∘9, and evaluates covariances under L∞([0,∞);Rd)0. It can be updated recursively alongside the primary optimization, avoiding matrix inversion, resampling, and sample splitting—the authors state this is the first inference method for gradient-based estimators beyond fixed-time asymptotics.
The theoretical guarantee is a uniform consistency rate:
L∞([0,∞);Rd)1
The proof splits the error into L∞([0,∞);Rd)2, controlled via Grönwall-type stability of the homogeneous sensitivity equations under Hessian and gradient approximation errors, and L∞([0,∞);Rd)3, controlled by bracketing entropy bounds for the time-indexed gradient classes. The logarithmic factor originates entirely from the Hessian estimation rate and can be removed if a root-L∞([0,∞);Rd)4 Hessian estimator is available; the infinite horizon itself incurs no additional statistical price.
Numerical validation
Discretizing the coupled flow and sensitivity dynamics with explicit Euler or fourth-order Runge–Kutta schemes (the Euler discretization coinciding with fixed-step gradient descent), simulations across linear, logistic, quantile (L∞([0,∞);Rd)5), phase retrieval, and ridge regression with L∞([0,∞);Rd)6, L∞([0,∞);Rd)7, and 1000 Monte Carlo replications show standardized z-scores closely matching the standard normal and coordinatewise coverage near nominal levels—for example, phase retrieval achieves 91.5% average coverage at the 90% nominal level and 96.0% at 95%, with quantile regression reaching 88.6% and 94.1% respectively including the intercept. Performance is stable across solvers. In experiments with fixed L∞([0,∞);Rd)8 and growing dimension, however, the Gaussian approximation degrades—most visibly for logistic regression and phase retrieval—attributed to ill-conditioned loss landscapes, indicating the need for bias reduction or spectral initialization outside the large-sample regime.
Limitations and open questions
The theory is confined to fixed dimensions and relies on structural conditions—local strong convexity (eventual coercivity), Lipschitz gradients and Hessians, and exponential convergence of the population flow—that typically fail globally when L∞([0,∞);Rd)9 grows with or exceeds P0. Extension to sparse, low-rank, or manifold-constrained parameters would require restricted strong convexity along the trajectory, and nonsmooth penalties would reformulate the flow as a differential inclusion whose statistical behavior may be altered by active-set dynamics. Second, the continuous-time analysis does not account for discretization error in stochastic gradient methods, where algorithmic noise interacts with statistical variability over long horizons; whether discretization bias remains negligible relative to P1 fluctuations is left open. Third, whether the low-complexity trajectory structure underlying the uniform CLT persists for proximal methods, mirror descent, or second-order schemes is unresolved.
Conclusion
The paper establishes that empirical gradient flows admit uniform Gaussian asymptotics over an infinite time horizon, with the key mechanism being the finite arc length of the sensitivity-function path, and delivers a jointly evolving covariance estimator achieving nearly parametric uniform consistency. Together these provide asymptotically valid confidence intervals valid at adaptive and diverging stopping times, connecting optimization dynamics with empirical process theory in a way that fixed-time analyses do not.