Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Wasserstein Geometric Framework for Hebbian Plasticity

Published 17 Apr 2026 in math.OC, cs.LG, and math.PR | (2604.16052v1)

Abstract: We introduce the Tan-HWG framework (Hebbian-Wasserstein-Geometry), a geometric theory of Hebbian plasticity in which memory states are modeled as probability measures evolving through Wasserstein minimizing movements. Hebbian learning rules are formalized as Hebbian energies satisfying a sequential stability condition, ensuring well-posed fiberwise JKO updates, optimal-transport realizations, and an energy descent inequality. This variational structure induces a fundamental separation between internal and observable dynamics. Internal memory states evolve along Wasserstein geodesics in a latent curved space, while observable quantities, such as effective synaptic weights, arise through geometric projection maps into external spaces. Simplicial projections recover classical affine schemes (including exponential moving averages and mirror descent), while revealing synaptic competition and pruning as geometric consequences of mass redistribution. Hilbertian projections provide a geometric account of phase alignment and multi-scale coherence. Classical neural networks appear as flat projections of this curved dynamics, while the framework naturally accommodates richer distributional representations, including structural weights and embedding memories, and their spectral extensions in complex internal spaces. Under mild Lipschitz regularity assumptions, including a quasi-stationary "sleep-mode" regime, we establish the existence of continuous-time limit curves. This yields a variational formulation of memory consolidation as a perturbed Wasserstein gradient flow. The framework thus provides a unified geometric foundation for synaptic plasticity, representation dynamics, and context-dependent computation.

Authors (1)

Summary

  • The paper introduces the Tan–HWG framework that formulates Hebbian plasticity as fiberwise Wasserstein flows using minimizing-movement schemes.
  • It details a separation between internal geometric evolution and observable affine dynamics, deriving classical learning rules as projections.
  • The framework unifies phenomena such as competition, pruning, and phase-locking, offering insights for robust, multi-scale neural representations.

A Wasserstein Geometric Theory of Hebbian Plasticity: The Tan–HWG Framework

Introduction and Objectives

This paper introduces the Tan–HWG framework (Hebbian–Wasserstein–Geometry), a geometric and variational formalism for Hebbian synaptic plasticity. Memory states are formulated as fiberwise probability measures on internal spaces, which evolve via minimizing-movement (JKO) schemes in fiberwise W2W_2-Wasserstein space. Hebbian learning is rigorously re-expressed through classes of energies (Tan–HWG energies) satisfying analytic well-posedness and sequential stability to admit fiberwise JKO updates, optimal-transport realizations, and an energy descent inequality.

The mathematical structure induces a separation between latent, geometrically rich internal memory evolution and observables (such as synaptic weights) arising through geometric projection operators. Classical neural networks are shown as “flat projections” of these internal, curved dynamics. The framework handles projection to simplices (affine schemes, including EMA/mirror descent), competitive dynamics, pruning, as well as phase and multi-scale coherence (Hilbertian projections).

Mathematical Framework

Memory Fields and Geometry

Each postsynaptic unit xXx\in X is assigned a local memory state ρxP2(Y)\rho_x\in \mathcal{P}_2(Y), where YY is a geodesic Polish (often CAT(0)) space. The collection ρ=(ρx)xX\rho = (\rho_x)_{x\in X} defines a memory field. The product Wasserstein metric

W(ρ,ρ)=(XW22(ρx,ρx)dμ(x))1/2\mathcal{W}(\rho,\rho') = \left( \int_X W_2^2(\rho_x,\rho'_x)\,d\mu(x) \right)^{1/2}

makes Xμ\mathcal{X}_\mu (the space of μ\mu-compatible fields) a geodesic metric space. Memory fields yield induced (“cognitive”) pseudo-metrics, reflecting dynamic internal similarity.

Hebbian Energies and JKO-Type Schemes

Tan–HWG energies are fiberwise-geometric functionals

E(ρ,ϕ)=XEx(ρx,Sx(ρ,ϕ))dμ(x)E(\rho,\phi) = \int_X \mathcal{E}_x\big(\rho_x,S_x(\rho,\phi)\big)\,d\mu(x)

with Hebbian structure: the energy at each xx depends only on the local state xXx\in X0 and a global “signal” xXx\in X1 (typically a function of xXx\in X2 and a context xXx\in X3).

Sequential stability conditions are imposed so that, for any frozen context, the fiberwise minimizing-movement (JKO) update is well-posed and geodesically convex. The discrete dynamics is

xXx\in X4

and the update at each xXx\in X5 is independent and solvable via optimal transport.

Observable Quantities and Geometric Projection

Observables like synaptic weights are not xXx\in X6 but projected quantities: the framework defines “stochastic geodesic projections” from internal measures to observables (e.g., projections onto the simplex for weights supported on a finite set, or to Hilbert spaces for spectral observables). Notably, on flat internal geometry and appropriate projections, the observable evolution reduces to classical affine updates (e.g., EMA).

Main Results

Internal/Observable Separation and Unified Variational Structure

  • Separation Principle: Internal states evolve via xXx\in X7-geodesics, while observable states are projections that may erase geometric curvature, resulting in classic affine recursions (e.g., EMAs, mirror descent).
  • Minimizing-Movement Scheme: For Tan–HWG energies, the JKO update is well-posed, admits an energy descent inequality, and is realized fiberwise by optimal transport (including Brenier maps in xXx\in X8 or displacement interpolation on metric graphs).

Projection Principle and Affine Observable Dynamics

  • Stochastic Geodesic Projection: The observable at time xXx\in X9 along a geodesic from ρxP2(Y)\rho_x\in \mathcal{P}_2(Y)0 to ρxP2(Y)\rho_x\in \mathcal{P}_2(Y)1 in ρxP2(Y)\rho_x\in \mathcal{P}_2(Y)2 is the barycenter ρxP2(Y)\rho_x\in \mathcal{P}_2(Y)3. Thus, the internal geometric evolution projects to affine dynamics on observables.
  • Emergent Classical Algorithms: For quadratic energies and flat projections, the observable evolution is exactly an EMA:

ρxP2(Y)\rho_x\in \mathcal{P}_2(Y)4

where the contraction ρxP2(Y)\rho_x\in \mathcal{P}_2(Y)5. This closed-form covers classical moving average, mirror descent, and exponential filtering.

Generalization to Complex and Structured Memories

  • Structural and Embedding Memory: The internal state encodes both a structural weight (attention/allocation) and an embedding (possibly distributional and/or complex). This generalizes classic NN weight matrices to mixture/semantically rich representations.
  • Spectral Extensions: By admitting internal spaces like ρxP2(Y)\rho_x\in \mathcal{P}_2(Y)6, phase and frequency information is geometrized. This encodes Kuramoto-like phase-locking, synchronization, and frequency offset storage within a unified geometric plasticity law.

Continuous-Time Limit and Memory Consolidation

  • Under mild regularity and quasi-stationary ("sleep-mode") context assumptions, the iterate interpolant converges (in the small time-step limit) to a locally absolutely continuous curve, satisfying a perturbed energy dissipation inequality. This provides a rigorous formulation of continuous-time memory consolidation as a Wasserstein gradient flow perturbed by freezing errors.

Implications

Theoretical Perspective

The framework rigorously unifies classical Hebbian models, EMAs, mirror descent, attention, and multi-scale plasticity as instances/projections of evolution in an underlying geometric (often curved) probability space. Many phenomena, such as competition, selectivity, phase-locking, and multi-scale coherence, are consequences of geometric and variational constraints rather than ad hoc modeling assumptions.

The mathematics ties structure (energy convexity, signal regularity) to dynamic properties such as stability, non-determinism (trajectories need not be unique under non-convexity), and bifurcation (eigenvalue crossing of the Wasserstein Hessian at fixed points).

Practical/Computational Prospects

The geometric separation enables flexible models handling robust, multi-modal, stochastic, and oscillatory representations, with observable dynamics matching classical learning rules in the appropriate limits. The simplex constraint on structural weights yields natural pruning/selectivity without external regularization. Spectral and phase representations provide built-in mechanisms for synchronization and binding.

The framework suggests natural multi-timescale separation: functional learning operates on fast, structurally fixed timescales, while structural consolidation (geometry-changing dynamics) becomes possible only in “quiet” (sleep-mode or quasi-stationary) contexts. This aligns with observed biological memory consolidation phenomena and offers rationales for separate training/consolidation phases in artificial architectures.

Applications could include geometric generative modeling, continual learning with structured representations, unsupervised manifold discovery, or biologically plausible learning systems.

Future Directions

Key open areas include characterization of dynamics under non-convexities, extension to more general internal geometries (e.g., directed graphs, higher-order connections), scalable computational solutions for high-dimensional Wasserstein flows, and deeper integration with biological observations (e.g., during sleep and replay). The framework also invites re-examination of classical regularization, sparsification, and hierarchical abstraction as geometric projection artifacts of curved internal evolution.

Conclusion

The Tan–HWG framework rigorously recasts Hebbian plasticity as a variational geometric flow in probability space, with observable learning dynamics emerging as projections. Both classic neuro-computational algorithms and novel phenomena (e.g., competition, pruning, multi-scale structural dynamics, synchronization) arise naturally as consequences of geometric and analytic constraints. Memory, in this view, is the trace of a trajectory in Wasserstein space, and classical affine learning rules are simply the projections of internal geometric evolution under appropriate energy and signal regularity assumptions.

Strong claims—such as the exact emergence of EMA/mirror descent, competition/pruning as simplex projection consequences, and the unique role of quadratic Wasserstein energies for state-independent contraction—are all substantiated with rigorous mathematical results (see Theorems 1–3, stochastic projection principles).

The theory reshapes the understanding of learning as fundamentally geometric: algebraic rules in observable space are projections of structured, variational movement in a latent probability geometry. This geometric foundation not only unifies previous rules but suggests trajectories for future models and a novel view of the organization of memory, representation, and computation in neural systems.

Reference: "A Wasserstein Geometric Framework for Hebbian Plasticity" (2604.16052)

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 10 tweets with 5 likes about this paper.