- The paper introduces a novel analytic framework that reformulates gradient descent dynamics using activation-driven fields and weight-induced pullback Gram metrics.
- It demonstrates a recursive layerwise decomposition that refines NTK analysis by incorporating finite-width and activation-conditioned effects.
- The study reveals that spectral properties and gradient transport in deep networks are governed by hierarchical geometric structures with practical training implications.
Hierarchical Weight-Induced Gram Metrics in Layerwise Learning Dynamics
Overview
This paper develops an analytic framework for characterizing the learning dynamics of finite-width, feedforward ReLU networks with fixed output weights and quadratic loss, reformulating gradient descent not in weight space, but within a hierarchy of fields and induced Gram metrics acting on the training set. The analysis systematically eliminates direct weight dependencies in favor of activation-driven fields and layerwise collective operators, culminating in a recursive structure of pullback Gram metrics that organize the geometric transport of gradients and activations through depth. This perspective refines and generalizes the NTK formalism by introducing finite-width, activation-conditioned co-activation operators, and exposes how depth generates a hierarchy of observable geometric structures with nontrivial implications for both theory and spectral analysis of trained networks.
From Weight-Space to Field-Space Dynamics
Motivation and Single-Layer Case
Traditional gradient-based learning is expressed in terms of the evolution of network parameters, with activations and outputs regarded as derived quantities. However, for a single hidden layer, the activation fields can be used to eliminate the weights entirely from the residual dynamics. The resulting evolution equation for the training residuals is governed by a collective kernel that factorizes as
Jαβ​(t)=Qαβ(0)​⋅Aαβ(1)​(t)
where Qαβ(0)​ is the empirical input overlap and Aαβ(1)​(t) the instantaneous co-activation overlap. The residual dynamics is fully closed within the activation and co-activation fields, revealing that gradient transport is structured by the intersection geometry of activation patterns.
Extension to Greater Depth
With two hidden layers, direct closure in activations is lost; a new family of conjugate fields emerges, defined recursively by backwards transport of co-activation patterns through the weights. The kernel for the residuals is now a sum of layerwise terms, each a product of activation overlaps and conjugate-field correlators:
Kαβ(2)​=Qαβ(1)​Sαβ(2)​+Qαβ(0)​Sαβ(1)​
From the third layer onwards, closure necessitates the introduction of weight-induced (pullback) Gram metrics of the form
Gℓαβ​=(W(ℓ+1))TDℓαβ​W(ℓ+1)
where Dℓαβ​ is the co-activation projector onto the subspace of neurons co-active for both examples α,β. These metrics quantify the geometric structure of backward transport under the joint configuration of weights and activation masks.
Hierarchical Operator Recursion and Observable Geometry
Recursive Layerwise Closure
For networks of arbitrary depth L, the residual dynamics organizes recursively as:
Kαβ(L)​=ℓ=1∑L​Qαβ(ℓ−1)​Sαβ(ℓ)​
where each Sαβ(ℓ)​ is a fourth-order correlator of activation masks and recursively defined conjugate fields. The latter fields admit a closed backward recursion involving the pullback action of the weight matrices and co-activation masks. The operator-valued recursion governing their update,
Qαβ(0)​0
where Qαβ(0)​1, constitutes a finite recursive hierarchy of transport operators mediating the collective geometrical structure of backpropagated gradients.
Gram Metrics and Geometric Interpretation
The induced metric Qαβ(0)​2 is interpreted as the pullback of the co-activation geometry at layer Qαβ(0)​3 into the parameter space of layer Qαβ(0)​4. Idempotency and symmetry properties of Qαβ(0)​5 imply that Qαβ(0)​6 are positive semidefinite, reflecting the orthogonal projection onto the effective transport subspaces. Deeper networks propagate gradients through repeated compositions of such pullbacks, exposing a hierarchical geometric organization not captured by mean-field NTK analysis.
Structural Insights
Several key structural properties are established:
- Layerwise Decomposition: The learning kernel always decomposes into a sum of layer-specific contributions coupling input/activation overlaps to higher-order conjugate-field correlators.
- Fourth-Order Closure: Residual learning dynamics remains at the level of second-order and fourth-order correlators; higher-order statistics are not generated by increased depth.
- Observable Weight Dependence: All explicit dependence on the weights in the closed observable dynamics arises via pullback Gram metrics and operator recursion, not direct parameter variables.
This reveals that depth increases the internal geometric complexity of network dynamics without generating higher-order collective statistics at the observable level.
Relation to Spectral Analysis and Broader Implications
The operator-valued Gram metrics derived in this framework motivate reinterpreting spectral analyses of trained networks, such as those explored in the WeightWatcher project [martin2021heavytailed]. The distribution and anisotropy of eigenvalues observed empirically in Qαβ(0)​7 are here connected to the repeated action of activation-conditioned projectors, suggesting that heavy-tailed spectra and self-regularization phenomena are a result of progressive concentration of gradient flow within activation-dependent subspaces of the network. This mechanism provides a concrete dynamical explanation for empirical findings associating the spectral properties of weight matrices with generalization and robustness.
Theoretical and Practical Outlook
This formalism refines the NTK limit by retaining finite-width effects via explicit activation pattern dynamics and their interaction with weight-induced transport operators. The hierarchy of field and operator-valued observables it introduces allows analysis of network learning as the evolution of a geometric state space with recursively defined metrics. This geometric perspective may provide new, analytically tractable paths for understanding phenomena such as feature learning, layerwise orthogonalization, and spectral signatures of generalization.
Future extensions may include (a) generalizing to adaptive readouts and nonlinear objectives, (b) empirical investigation of metric dynamics during large-scale training, and (c) connections to recent work formalizing feature linearization and sequential learning geometry as in "The Weight Gram Matrix Captures Sequential Feature Linearization in Deep Networks" (Cha et al., 7 May 2026).
Conclusion
This work establishes a closed, layerwise, and hierarchical structure for the learning dynamics of finite-width ReLU networks under gradient descent. Weight-induced, activation-conditioned pullback Gram metrics form the core geometric objects mediating the observable evolution. By stepping beyond parameter-space descriptions to collective field-geometry dynamics, the analysis identifies a finite recursive hierarchy governing the transport of information across depth, connecting gradient transport, activation geometry, and spectral phenomena in trained neural networks.