- The paper introduces the Tan–HWG framework that formulates Hebbian plasticity as fiberwise Wasserstein flows using minimizing-movement schemes.
- It details a separation between internal geometric evolution and observable affine dynamics, deriving classical learning rules as projections.
- The framework unifies phenomena such as competition, pruning, and phase-locking, offering insights for robust, multi-scale neural representations.
A Wasserstein Geometric Theory of Hebbian Plasticity: The Tan–HWG Framework
Introduction and Objectives
This paper introduces the Tan–HWG framework (Hebbian–Wasserstein–Geometry), a geometric and variational formalism for Hebbian synaptic plasticity. Memory states are formulated as fiberwise probability measures on internal spaces, which evolve via minimizing-movement (JKO) schemes in fiberwise W2-Wasserstein space. Hebbian learning is rigorously re-expressed through classes of energies (Tan–HWG energies) satisfying analytic well-posedness and sequential stability to admit fiberwise JKO updates, optimal-transport realizations, and an energy descent inequality.
The mathematical structure induces a separation between latent, geometrically rich internal memory evolution and observables (such as synaptic weights) arising through geometric projection operators. Classical neural networks are shown as “flat projections” of these internal, curved dynamics. The framework handles projection to simplices (affine schemes, including EMA/mirror descent), competitive dynamics, pruning, as well as phase and multi-scale coherence (Hilbertian projections).
Mathematical Framework
Memory Fields and Geometry
Each postsynaptic unit x∈X is assigned a local memory state ρx∈P2(Y), where Y is a geodesic Polish (often CAT(0)) space. The collection ρ=(ρx)x∈X defines a memory field. The product Wasserstein metric
W(ρ,ρ′)=(∫XW22(ρx,ρx′)dμ(x))1/2
makes Xμ (the space of μ-compatible fields) a geodesic metric space. Memory fields yield induced (“cognitive”) pseudo-metrics, reflecting dynamic internal similarity.
Hebbian Energies and JKO-Type Schemes
Tan–HWG energies are fiberwise-geometric functionals
E(ρ,ϕ)=∫XEx(ρx,Sx(ρ,ϕ))dμ(x)
with Hebbian structure: the energy at each x depends only on the local state x∈X0 and a global “signal” x∈X1 (typically a function of x∈X2 and a context x∈X3).
Sequential stability conditions are imposed so that, for any frozen context, the fiberwise minimizing-movement (JKO) update is well-posed and geodesically convex. The discrete dynamics is
x∈X4
and the update at each x∈X5 is independent and solvable via optimal transport.
Observable Quantities and Geometric Projection
Observables like synaptic weights are not x∈X6 but projected quantities: the framework defines “stochastic geodesic projections” from internal measures to observables (e.g., projections onto the simplex for weights supported on a finite set, or to Hilbert spaces for spectral observables). Notably, on flat internal geometry and appropriate projections, the observable evolution reduces to classical affine updates (e.g., EMA).
Main Results
Internal/Observable Separation and Unified Variational Structure
- Separation Principle: Internal states evolve via x∈X7-geodesics, while observable states are projections that may erase geometric curvature, resulting in classic affine recursions (e.g., EMAs, mirror descent).
- Minimizing-Movement Scheme: For Tan–HWG energies, the JKO update is well-posed, admits an energy descent inequality, and is realized fiberwise by optimal transport (including Brenier maps in x∈X8 or displacement interpolation on metric graphs).
Projection Principle and Affine Observable Dynamics
- Stochastic Geodesic Projection: The observable at time x∈X9 along a geodesic from ρx∈P2(Y)0 to ρx∈P2(Y)1 in ρx∈P2(Y)2 is the barycenter ρx∈P2(Y)3. Thus, the internal geometric evolution projects to affine dynamics on observables.
- Emergent Classical Algorithms: For quadratic energies and flat projections, the observable evolution is exactly an EMA:
ρx∈P2(Y)4
where the contraction ρx∈P2(Y)5. This closed-form covers classical moving average, mirror descent, and exponential filtering.
Generalization to Complex and Structured Memories
- Structural and Embedding Memory: The internal state encodes both a structural weight (attention/allocation) and an embedding (possibly distributional and/or complex). This generalizes classic NN weight matrices to mixture/semantically rich representations.
- Spectral Extensions: By admitting internal spaces like ρx∈P2(Y)6, phase and frequency information is geometrized. This encodes Kuramoto-like phase-locking, synchronization, and frequency offset storage within a unified geometric plasticity law.
Continuous-Time Limit and Memory Consolidation
- Under mild regularity and quasi-stationary ("sleep-mode") context assumptions, the iterate interpolant converges (in the small time-step limit) to a locally absolutely continuous curve, satisfying a perturbed energy dissipation inequality. This provides a rigorous formulation of continuous-time memory consolidation as a Wasserstein gradient flow perturbed by freezing errors.
Implications
Theoretical Perspective
The framework rigorously unifies classical Hebbian models, EMAs, mirror descent, attention, and multi-scale plasticity as instances/projections of evolution in an underlying geometric (often curved) probability space. Many phenomena, such as competition, selectivity, phase-locking, and multi-scale coherence, are consequences of geometric and variational constraints rather than ad hoc modeling assumptions.
The mathematics ties structure (energy convexity, signal regularity) to dynamic properties such as stability, non-determinism (trajectories need not be unique under non-convexity), and bifurcation (eigenvalue crossing of the Wasserstein Hessian at fixed points).
Practical/Computational Prospects
The geometric separation enables flexible models handling robust, multi-modal, stochastic, and oscillatory representations, with observable dynamics matching classical learning rules in the appropriate limits. The simplex constraint on structural weights yields natural pruning/selectivity without external regularization. Spectral and phase representations provide built-in mechanisms for synchronization and binding.
The framework suggests natural multi-timescale separation: functional learning operates on fast, structurally fixed timescales, while structural consolidation (geometry-changing dynamics) becomes possible only in “quiet” (sleep-mode or quasi-stationary) contexts. This aligns with observed biological memory consolidation phenomena and offers rationales for separate training/consolidation phases in artificial architectures.
Applications could include geometric generative modeling, continual learning with structured representations, unsupervised manifold discovery, or biologically plausible learning systems.
Future Directions
Key open areas include characterization of dynamics under non-convexities, extension to more general internal geometries (e.g., directed graphs, higher-order connections), scalable computational solutions for high-dimensional Wasserstein flows, and deeper integration with biological observations (e.g., during sleep and replay). The framework also invites re-examination of classical regularization, sparsification, and hierarchical abstraction as geometric projection artifacts of curved internal evolution.
Conclusion
The Tan–HWG framework rigorously recasts Hebbian plasticity as a variational geometric flow in probability space, with observable learning dynamics emerging as projections. Both classic neuro-computational algorithms and novel phenomena (e.g., competition, pruning, multi-scale structural dynamics, synchronization) arise naturally as consequences of geometric and analytic constraints. Memory, in this view, is the trace of a trajectory in Wasserstein space, and classical affine learning rules are simply the projections of internal geometric evolution under appropriate energy and signal regularity assumptions.
Strong claims—such as the exact emergence of EMA/mirror descent, competition/pruning as simplex projection consequences, and the unique role of quadratic Wasserstein energies for state-independent contraction—are all substantiated with rigorous mathematical results (see Theorems 1–3, stochastic projection principles).
The theory reshapes the understanding of learning as fundamentally geometric: algebraic rules in observable space are projections of structured, variational movement in a latent probability geometry. This geometric foundation not only unifies previous rules but suggests trajectories for future models and a novel view of the organization of memory, representation, and computation in neural systems.
Reference: "A Wasserstein Geometric Framework for Hebbian Plasticity" (2604.16052)