- The paper constructs an infinite-horizon IGW gradient flow despite the failure of global geodesic convexity, using local convexity estimates and uniform convergence of JKO interpolants.
- The flow is characterized by a nonlinear Fokker–Planck equation whose second-moment matrix follows an autonomous ODE, enabling an equivalent linear Fokker–Planck equation and second-moment-modulated Ornstein–Uhlenbeck SDE.
- The paper proves instantaneous smoothing and exponential convergence of relative entropy, with the rate depending on the largest initial covariance eigenvalue, and obtains exponential convergence in Wasserstein distance via Talagrand’s inequality.
This paper develops the theory of gradient flows in the inner product Gromov–Wasserstein (IGW) geometry for the relative entropy functional H(⋅∥γ) with respect to the standard Gaussian γ=N(0,I). Building on the IGW–JKO framework of Zhang et al., the authors construct the flow on the infinite time horizon despite a fundamental obstruction—failure of global geodesic convexity—and then provide increasingly explicit characterizations of the dynamics, culminating in a linear stochastic differential equation (SDE) representation and exponential convergence to equilibrium (2608.19198).
Background and motivation
The classical JKO scheme defines gradient flows over probability measures by iterating proximal steps in the $2$-Wasserstein metric W2, connecting variational calculus, evolution PDEs, and diffusions. Recent work has extended this program to alternative geometries, including Stein, Fisher–Rao, Sinkhorn, and IGW distances. The IGW distance between μ,ν∈P2(Rd) minimizes over couplings π the squared discrepancy of pairwise inner products,
IGW(μ,ν)2=π∈Π(μ,ν)inf∫∣⟨x,x′⟩−⟨y,y′⟩∣2dπ⊗π,
and induces an orthogonally invariant pseudometric. The prior construction of IGW gradient flows applies only to rotationally invariant functionals that are λ-convex along generalized or modified generalized IGW geodesics, and characterizes flows implicitly via partial integro-differential equations (PIDEs). Three gaps remained open: whether canonical functionals such as relative entropy satisfy the required convexity; whether the implicit PIDE admits interpretable PDE/SDE representations analogous to the Langevin/Ornstein–Uhlenbeck (OU) correspondence in the Wasserstein setting; and whether quantitative convergence rates hold.
Failure of geodesic convexity
The paper establishes by explicit counterexamples that H(⋅∥γ) is not λ-convex along either generalized or modified generalized IGW geodesics, for any γ=N(0,I)0. For modified generalized geodesics, the argument takes γ=N(0,I)1 and γ=N(0,I)2; the midpoint of the resulting geodesic has variance γ=N(0,I)3 with γ=N(0,I)4 as γ=N(0,I)5, so that γ=N(0,I)6 grows like γ=N(0,I)7 while all competing terms grow at most logarithmically, violating convexity for any prescribed γ=N(0,I)8. A separate, more involved construction in γ=N(0,I)9, combined with a scaling argument, rules out generalized geodesic convexity as well. This negative result places relative entropy outside the scope of the existing general theory and necessitates a new existence proof.
Construction of the flow
The central technical device replacing global convexity is a local convexity estimate: along any modified generalized IGW geodesic $2$0, the excess of $2$1 over the affine interpolation of endpoint entropies is controlled by terms proportional to the pairwise inner-product mismatch and $2$2, with constants depending on the second moments and cross-covariance spectra of the endpoints. Combined with the known local convexity of $2$3 itself, this yields a proximal inequality that substitutes for the role played by $2$4-convexity in the original theory.
With this estimate, the authors establish a uniform Cauchy estimate for the piecewise-constant IGW–JKO interpolants via a cross-partition error function and a Grönwall-type argument, upgrading pointwise weak convergence to uniform $2$5 convergence along a subsequence. Consequently, the limiting curve satisfies the continuity equation with velocity field lying in $2$6, where $2$7 is the nonlocal mobility operator. A spectral bound derived from Talagrand's $2$8 inequality and Gaussian entropy comparison—namely, that $2$9 forces eigenvalues of W20 into W21—ensures uniform nondegeneracy of second moment matrices along the flow, enabling concatenation of finite-horizon solutions to extend the flow to all W22.
From PIDE to Fokker–Planck equations
Evaluating the inverse mobility operator in closed form converts the PIDE into a nonlinear Fokker–Planck equation:
W23
The derivation requires care because no a priori regularity of W24 is available; symmetry of a key moment integral is justified via cutoff-function integration by parts under only W25 regularity.
The structural insight is that the second moment matrix decouples from the law: testing the nonlinear FPE against quadratic test functions shows that W26 satisfies the autonomous matrix ODE
W27
which admits a unique solution in W28. Substituting W29 for μ,ν∈P2(Rd)0 yields a linear, time-inhomogeneous FPE with unique subprobability solution. Notably, this ODE is precisely the Euclidean gradient flow of μ,ν∈P2(Rd)1 restricted to centered Gaussians. Comparing with the Bures–Wasserstein covariance dynamics μ,ν∈P2(Rd)2, the authors observe that near equilibrium the Wasserstein flow moves roughly four times faster in Frobenius norm than the IGW analogue—a concrete sense in which the IGW geometry slows descent.
Probabilistic representation
The linear FPE identifies the IGW gradient flow as the time-marginal law of the SDE
μ,ν∈P2(Rd)3
a second-moment-modulated OU process. Closed-form expressions follow from the variation-of-constants formula: μ,ν∈P2(Rd)4 with explicitly defined μ,ν∈P2(Rd)5 and μ,ν∈P2(Rd)6. Consequences include preservation of Gaussianity (with mean and covariance obeying a closed ODE system), instantaneous smoothing (μ,ν∈P2(Rd)7 admits a μ,ν∈P2(Rd)8 density for μ,ν∈P2(Rd)9 since π0), and uniqueness of strong solutions. Importantly, the authors show via a one-dimensional example (initial mean and variance both equal to π1, where the variance strictly increases initially) that this process is not, in general, a deterministic time change of the OU process—the drift modulation genuinely alters the dynamics beyond rescaling time.
Exponential convergence
Using finiteness of kinetic energy and integrated relative Fisher information (the latter bounded via the spectral control on π2), the map π3 is shown absolutely continuous with derivative bounded through the closed-form velocity field. Combining the Gaussian log-Sobolev inequality with the monotonicity of each eigenvalue of π4 toward π5 under the matrix ODE gives
π6
an exponential contraction rate depending only on the initial second moment. Via Talagrand's inequality this also implies exponential π7 convergence, mirroring the qualitative behavior of the OU semigroup despite the different underlying geometry.
Limitations and open questions
Several restrictions qualify these results. The entire analysis concerns the Gaussian reference measure; rotational invariance of the target is essential for the functional even to be well-defined in the IGW geometry, and the second-moment decoupling that enables linearization is specific to the Gaussian case. The authors note that for more general rotationally invariant log-concave targets, the nonlinear FPE may persist but the decoupling likely fails, suggesting McKean–Vlasov-type SDE representations whose existence is unverified. Extension to other rotationally invariant π8-divergences remains open, as does the question of whether the identified FPE or SDE models any physical or biological phenomenon. The counterexample for generalized geodesics relies partly on numerical verification of optimality conditions, though the authors state they independently verified the argument.
Conclusion
This paper establishes the existence, infinite-horizon extension, and complete analytic-probabilistic characterization of the IGW gradient flow of relative entropy against the standard Gaussian. By circumventing the failure of geodesic convexity with a local estimate, it reduces an implicit nonlocal PIDE first to a nonlinear FPE, then—through autonomous second-moment dynamics—to a linear FPE realized by an OU-like SDE, and proves exponential entropy convergence. The results provide the first concrete PDE/SDE/diffusion correspondence for a Gromov–Wasserstein-type gradient flow, while leaving the non-Gaussian and general-π9-divergence settings as clearly posed open problems.