Weight-Induced Gram Operators in Deep Learning
- Weight-induced Gram operators are operator-valued summaries that capture the training geometry of deep ReLU networks through layer-specific activation overlaps and conjugate-field correlators.
- They decompose gradient descent into a hierarchy of layerwise contributions, providing an adaptive, time-dependent metric in example space that governs residual mode decay.
- This framework, which connects operator dynamics with classical frame theory concepts, enables a closed-form evolution of the network’s function space without relying on higher-order statistics.
Searching arXiv for the cited papers to ground the article in the relevant literature. Weight-induced Gram operators are operator-valued summaries of training geometry that arise when gradient descent in deep ReLU networks is rewritten in example space rather than weight space. For feed-forward ReLU networks with fixed readout and quadratic loss, the learning dynamics can be expressed as a residual evolution driven by a finite hierarchy of layerwise Gram matrices and pullback operators determined by activation overlaps, conjugate-field correlators, and the weights themselves. In this formulation, the residual vector evolves under a closed collective dynamics in the training-set space, and from depth three onward closure requires a hierarchy of weight-induced Gram operators that mediate information transport across layers (Nordio, 8 Jun 2026).
1. Network setting and operator-valued reformulation
The construction is stated for an -hidden-layer feed-forward network with ReLU activations,
with fixed readout
targets , quadratic loss
gradient descent on each with step , and residuals (Nordio, 8 Jun 2026).
The central shift is to eliminate, as far as possible, the explicit weight variables from the learning dynamics and to rewrite gradient descent as a collective dynamics on quantities defined over example indices . In this representation, the residual update is no longer described primarily as motion in the full parameter space, but as a coupled evolution of finitely many matrices. The key objects are layerwise Gram matrices in example-index space and, for deeper networks, a hierarchy of weight-induced pullback operators acting across layers.
The following objects organize the construction:
| Object | Definition | Role |
|---|---|---|
| 0 | 1 | activation overlap |
| 2 | 3 | conjugate-field correlator |
| 4 | 5 | layerwise Gram matrix |
| 6 | 7 | weight-induced pullback operator |
Here 8, the backward fields 9 are defined by backward propagation through ReLU masks, and 0 (Nordio, 8 Jun 2026).
2. Single-hidden-layer factorization
For a single hidden layer, the dynamics closes directly on the residuals. The fixed input geometry is encoded by the input Gram matrix
1
while the dynamical gating structure is encoded by the co-activation overlap
2
The residuals satisfy
3
with collective kernel
4
where 5 denotes the Hadamard product (Nordio, 8 Jun 2026).
This factorization isolates two distinct contributions. The matrix 6 is fixed by the training data, whereas 7 changes with the ReLU activation pattern. The result is a kernel that is neither purely data-geometric nor purely activation-driven, but a pointwise product of the two. In this sense, the single-layer case already exhibits the essential theme of the theory: learning is controlled by an example-space metric induced jointly by the data geometry and the current weight-dependent activation configuration.
A direct derivation follows from the chain rule. For 8,
9
so that
0
Averaging over 1 yields
2
which is the residual update in closed form (Nordio, 8 Jun 2026).
3. Layerwise Gram metrics in deep networks
For 3 hidden layers, the residual update decomposes into 4 layerwise contributions. Define
5
and
6
Then
7
Equivalently, with
8
one obtains
9
or in the continuous-time limit,
0
Thus the single-layer factorization 1 generalizes to deep networks by summing 2 Hadamard-factorized terms 3 (Nordio, 8 Jun 2026).
This decomposition has two immediate consequences. First, each layer contributes its own example-space metric rather than all training effects being compressed into a single aggregate kernel. Second, the dynamics remains closed in a finite set of 4 matrices 5. No higher-order statistics of the residuals are needed. The gradient flow in function space is therefore fully determined by the current Gram operators 6.
The role of the conjugate-field correlator 7 is especially significant in depth. In the single-layer case, gating information is carried by the co-activation matrix 8. In deeper networks, this generalizes to a combined co-activation/backpropagated object, since each neuron’s contribution to the residual update depends not only on whether it is active, but also on how downstream structure propagates backward through the network.
4. Backward pullback recursion and the hierarchy of operators
From depth three onward, closure requires explicit operator-valued quantities that transport geometry backward through the network. For every layer 9, define the projector
0
which projects onto neurons simultaneously active for examples 1 and 2 at layer 3. The simplest nontrivial pullback operator at layer 4 is then
5
Interpreting the top layer as 6, these operators satisfy the backward recursion
7
for 8 (Nordio, 8 Jun 2026).
At fixed ReLU masks and neglecting threshold crossings, this recursion is obtained by substituting successive next-layer expressions into the projector form. The result shows that each 9 is symmetric, positive semidefinite, and encodes how the co-activation geometry at layer 0 is pulled back to layer 1 via the weight map.
The significance of the hierarchy is structural rather than merely notational. The conjugate-field dynamics is governed by operators satisfying a backward pullback recursion, of which the weight-induced Gram operators are the first nontrivial instances (Nordio, 8 Jun 2026). In practical terms, deep layers do not affect earlier layers only through scalar summary statistics; they transmit a recursively transformed geometry determined jointly by activation masks and weights. This is the operator-theoretic core of the deep-network generalization.
5. Spectral interpretation, finite-width adaptivity, and relation to feature evolution
Because
2
up to the overall factor 3, the spectrum of the layerwise Gram matrices controls the decay of residual modes. The decomposition exposes which modes of the residual vector decay fastest: those in the top eigendirections of the 4. It also shows that deep layers communicate their co-activation geometry backward through the chain of 5’s, so that earlier layers see an effective kernel shaped by all deeper layers. In the infinite-width NTK regime each 6 converges to a fixed limit, recovering the well-known fixed kernel; at finite width they evolve during training, giving a richer, adaptive metric (Nordio, 8 Jun 2026).
A related but distinct 2026 framework studies the weight Gram matrix
7
as the key object capturing feature dynamics in deep networks. There the Feature Learning Equation
8
translates weight-space updates into feature-space updates, and the leading-order shift in 9 matches the Virtual Covariance Shift up to 0. That framework also introduces Target Linearity and argues that deep networks sequentially transform representations toward target-linear structure (Cha et al., 7 May 2026).
The two viewpoints operate on different spaces. In the residual-dynamics formulation, the principal objects are 1 matrices indexed by training examples and closed under the learning dynamics. In the feature-centric formulation, the principal object is the layerwise weight Gram matrix 2, interpreted as encoding the covariance-type update that hypothetical direct feature optimization would produce. A plausible implication is that the two approaches describe complementary closures: one in example space through residual modes, the other in feature space through virtual covariance and target alignment.
6. Related operator-theoretic usages and terminological ambiguities
Outside deep-learning theory, closely related terminology appears in frame theory. For Bessel sequences 3 and 4, and any bounded operator 5, the 6-cross Gram operator is
7
with entries
8
Specializing to the diagonal weight operator 9 on 0 gives
1
and for 2 one obtains
3
described as the standard “weight-induced Gram operator” in frame theory (Balazs et al., 2018).
In that setting, the main questions are Schatten 4-class membership, invertibility, Moore–Penrose pseudoinverses, and perturbation stability. For example, if 5 are Riesz bases and 6 for all 7, then
8
while small perturbations of the weights or frame sequences preserve invertibility under explicit norm conditions (Balazs et al., 2018). A fusion-frame analogue replaces vectors by weighted subspaces and defines the 9-fusion cross Gram matrix
0
with corresponding results on invertibility, pseudo-invertibility, and stability (Shamsabadi et al., 2017).
These usages are mathematically adjacent but conceptually different. In frame and fusion-frame theory, a weighted Gram operator is an operator representation problem in Hilbert space. In deep-network learning dynamics, weight-induced Gram operators are dynamical objects built from ReLU masks, activation overlaps, conjugate fields, and weight pullbacks, and they serve to close gradient descent in function space. A common misconception is to identify all such objects with a static matrix 1 or with a fixed kernel. The deep-network theory instead emphasizes a time-dependent hierarchy of operators whose evolution is itself part of the learning process (Nordio, 8 Jun 2026).