Papers
Topics
Authors
Recent
Search
2000 character limit reached

Learning Dynamics Reveal a Hierarchy of Weight-Induced Layerwise Gram Metrics

Published 8 Jun 2026 in cs.LG and cond-mat.dis-nn | (2606.09744v3)

Abstract: We study feed-forward ReLU networks with fixed readout and quadratic loss. The aim is to rewrite gradient descent not primarily as a dynamics in weight space, but as a collective dynamics closed in terms of fields defined on the training-set space. For a single hidden layer, the weight variables can be eliminated from the activation dynamics, yielding a closed equation for the residuals governed by a collective kernel that factorizes into an input-geometric matrix and a dynamical co-activation matrix. For deeper networks, the residual dynamics retains a clean layer-wise kernel structure. However, from depth three onward, closure requires a hierarchy of weight-induced Gram operators that mediate information transport across layers. Moreover, the conjugate-field dynamics is governed by operators satisfying a backward pullback recursion, of which the weight-induced Gram operators are the first nontrivial instances.

Authors (1)

Summary

  • The paper introduces a novel analytic framework that reformulates gradient descent dynamics using activation-driven fields and weight-induced pullback Gram metrics.
  • It demonstrates a recursive layerwise decomposition that refines NTK analysis by incorporating finite-width and activation-conditioned effects.
  • The study reveals that spectral properties and gradient transport in deep networks are governed by hierarchical geometric structures with practical training implications.

Hierarchical Weight-Induced Gram Metrics in Layerwise Learning Dynamics

Overview

This paper develops an analytic framework for characterizing the learning dynamics of finite-width, feedforward ReLU networks with fixed output weights and quadratic loss, reformulating gradient descent not in weight space, but within a hierarchy of fields and induced Gram metrics acting on the training set. The analysis systematically eliminates direct weight dependencies in favor of activation-driven fields and layerwise collective operators, culminating in a recursive structure of pullback Gram metrics that organize the geometric transport of gradients and activations through depth. This perspective refines and generalizes the NTK formalism by introducing finite-width, activation-conditioned co-activation operators, and exposes how depth generates a hierarchy of observable geometric structures with nontrivial implications for both theory and spectral analysis of trained networks.

From Weight-Space to Field-Space Dynamics

Motivation and Single-Layer Case

Traditional gradient-based learning is expressed in terms of the evolution of network parameters, with activations and outputs regarded as derived quantities. However, for a single hidden layer, the activation fields can be used to eliminate the weights entirely from the residual dynamics. The resulting evolution equation for the training residuals is governed by a collective kernel that factorizes as

Jαβ(t)=Qαβ(0)⋅Aαβ(1)(t)J_{\alpha\beta}(t) = Q_{\alpha\beta}^{(0)} \cdot A_{\alpha\beta}^{(1)}(t)

where Qαβ(0)Q_{\alpha\beta}^{(0)} is the empirical input overlap and Aαβ(1)(t)A_{\alpha\beta}^{(1)}(t) the instantaneous co-activation overlap. The residual dynamics is fully closed within the activation and co-activation fields, revealing that gradient transport is structured by the intersection geometry of activation patterns.

Extension to Greater Depth

With two hidden layers, direct closure in activations is lost; a new family of conjugate fields emerges, defined recursively by backwards transport of co-activation patterns through the weights. The kernel for the residuals is now a sum of layerwise terms, each a product of activation overlaps and conjugate-field correlators:

Kαβ(2)=Qαβ(1)Sαβ(2)+Qαβ(0)Sαβ(1)K_{\alpha\beta}^{(2)} = Q_{\alpha\beta}^{(1)} S_{\alpha\beta}^{(2)} + Q_{\alpha\beta}^{(0)} S_{\alpha\beta}^{(1)}

From the third layer onwards, closure necessitates the introduction of weight-induced (pullback) Gram metrics of the form

Gℓαβ=(W(ℓ+1))TDℓαβW(ℓ+1)G_\ell^{\alpha \beta} = \left(W^{(\ell+1)}\right)^T D_\ell^{\alpha\beta} W^{(\ell+1)}

where DℓαβD_\ell^{\alpha\beta} is the co-activation projector onto the subspace of neurons co-active for both examples α,β\alpha, \beta. These metrics quantify the geometric structure of backward transport under the joint configuration of weights and activation masks.

Hierarchical Operator Recursion and Observable Geometry

Recursive Layerwise Closure

For networks of arbitrary depth LL, the residual dynamics organizes recursively as:

Kαβ(L)=∑ℓ=1LQαβ(ℓ−1)Sαβ(ℓ)K_{\alpha\beta}^{(L)} = \sum_{\ell=1}^L Q_{\alpha\beta}^{(\ell-1)} S_{\alpha\beta}^{(\ell)}

where each Sαβ(ℓ)S_{\alpha\beta}^{(\ell)} is a fourth-order correlator of activation masks and recursively defined conjugate fields. The latter fields admit a closed backward recursion involving the pullback action of the weight matrices and co-activation masks. The operator-valued recursion governing their update,

Qαβ(0)Q_{\alpha\beta}^{(0)}0

where Qαβ(0)Q_{\alpha\beta}^{(0)}1, constitutes a finite recursive hierarchy of transport operators mediating the collective geometrical structure of backpropagated gradients.

Gram Metrics and Geometric Interpretation

The induced metric Qαβ(0)Q_{\alpha\beta}^{(0)}2 is interpreted as the pullback of the co-activation geometry at layer Qαβ(0)Q_{\alpha\beta}^{(0)}3 into the parameter space of layer Qαβ(0)Q_{\alpha\beta}^{(0)}4. Idempotency and symmetry properties of Qαβ(0)Q_{\alpha\beta}^{(0)}5 imply that Qαβ(0)Q_{\alpha\beta}^{(0)}6 are positive semidefinite, reflecting the orthogonal projection onto the effective transport subspaces. Deeper networks propagate gradients through repeated compositions of such pullbacks, exposing a hierarchical geometric organization not captured by mean-field NTK analysis.

Structural Insights

Several key structural properties are established:

  • Layerwise Decomposition: The learning kernel always decomposes into a sum of layer-specific contributions coupling input/activation overlaps to higher-order conjugate-field correlators.
  • Fourth-Order Closure: Residual learning dynamics remains at the level of second-order and fourth-order correlators; higher-order statistics are not generated by increased depth.
  • Observable Weight Dependence: All explicit dependence on the weights in the closed observable dynamics arises via pullback Gram metrics and operator recursion, not direct parameter variables.

This reveals that depth increases the internal geometric complexity of network dynamics without generating higher-order collective statistics at the observable level.

Relation to Spectral Analysis and Broader Implications

The operator-valued Gram metrics derived in this framework motivate reinterpreting spectral analyses of trained networks, such as those explored in the WeightWatcher project [martin2021heavytailed]. The distribution and anisotropy of eigenvalues observed empirically in Qαβ(0)Q_{\alpha\beta}^{(0)}7 are here connected to the repeated action of activation-conditioned projectors, suggesting that heavy-tailed spectra and self-regularization phenomena are a result of progressive concentration of gradient flow within activation-dependent subspaces of the network. This mechanism provides a concrete dynamical explanation for empirical findings associating the spectral properties of weight matrices with generalization and robustness.

Theoretical and Practical Outlook

This formalism refines the NTK limit by retaining finite-width effects via explicit activation pattern dynamics and their interaction with weight-induced transport operators. The hierarchy of field and operator-valued observables it introduces allows analysis of network learning as the evolution of a geometric state space with recursively defined metrics. This geometric perspective may provide new, analytically tractable paths for understanding phenomena such as feature learning, layerwise orthogonalization, and spectral signatures of generalization.

Future extensions may include (a) generalizing to adaptive readouts and nonlinear objectives, (b) empirical investigation of metric dynamics during large-scale training, and (c) connections to recent work formalizing feature linearization and sequential learning geometry as in "The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks" (Cha et al., 7 May 2026).

Conclusion

This work establishes a closed, layerwise, and hierarchical structure for the learning dynamics of finite-width ReLU networks under gradient descent. Weight-induced, activation-conditioned pullback Gram metrics form the core geometric objects mediating the observable evolution. By stepping beyond parameter-space descriptions to collective field-geometry dynamics, the analysis identifies a finite recursive hierarchy governing the transport of information across depth, connecting gradient transport, activation geometry, and spectral phenomena in trained neural networks.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.

Tweets

Sign up for free to view the 1 tweet with 6 likes about this paper.