---
title: Deep Neural Feature Ansatz (NFA)
url: https://www.emergentmind.com/topics/deep-neural-feature-ansatz-nfa
type: topic
---

# Deep Neural Feature Ansatz (NFA)

The Deep Neural Feature Ansatz (NFA) posits a unified mechanism underlying feature learning in deep, overparameterized neural networks, both fully connected and convolutional. According to the NFA, the feature geometry learned within each layer—quantified by the Gram (or covariance) matrix of weights—becomes highly correlated, and often proportional, to the average gradient outer product (AGOP) of the network output with respect to the relevant feature vector (entire input, local patch, or hidden activation) at that layer. This proportionality arises during gradient-based training, empirically holds across a wide range of architectures and tasks, and admits both theoretical and algorithmic consequences for network design, kernel methods, and interpretability [2212.13881, 2309.00570, 2402.05271, 2510.15563, 1605.08283].

## 1. Formal Statement and Definitions

The NFA asserts, for a network layer $\ell$ with weight matrix $W_\ell$ and "features" $h_{\ell-1}(x)$ (input or hidden activations), that the feature Gram matrix,
$$
F_\ell = W_\ell^\top W_\ell,
$$
aligns with the AGOP evaluated over a dataset:
$$
G_\ell = \mathbb{E}_{x,y}\big[\nabla_{h_{\ell-1}(x)} f(x, y)\,\nabla_{h_{\ell-1}(x)} f(x, y)^\top\big].
$$
In convolutional networks, $G_\ell$ is typically formed by averaging over the spatial patches of the feature maps [2309.00570]. In fully connected and deep linear networks, $G_\ell$ reduces to an EGOP over input coordinates or hidden activations [2212.13881, 2510.15563].

The ansatz often holds up to a scalar multiple or an exponent. For instance, in deep linear networks of depth $L$ with balanced initialization, it is shown that
$$
F_1 \propto G^{1/L}
$$
where $F_1$ is the Gram of the first layer and $G = \mathbb{E}_x\big[\nabla_x f(x)\,\nabla_x f(x)^\top\big]$ [2510.15563]. In convolutional architectures, $F_\ell$ matches the AGOP aggregated over patches in spatial dimensions [2309.00570]. The proportionality constant(s) depend on architecture, depth, learning rate, and initialization.

## 2. Theoretical Foundations and Proof Strategies

Several classes of models admit rigorous proofs of the NFA:

- In deep linear networks under gradient flow, if balanced initialization is maintained, the Gram matrix of layer 1 aligns to the $1/L$th power of the network's overall sensitivity (i.e., the AGOP). This is established via matrix power identities: the recursive structure of deep linear models forces $W_1 W_1^\top = (J^\top J)^{1/L}$, where $J$ is the linear map realized by the full network [2510.15563].

- In certain scenarios (e.g., one-step gradient descent with zero initialization), the filter Gram in convolutional settings matches the AGOP exactly, as shown by direct expansion of the gradient update. Taylor expansion and moment matching arguments extend this to early multi-step SGD [2309.00570].

- For high-dimensional nonlinear networks, the alignment emerges via SGD-induced coupling between the left singular vectors of $W_\ell$ and the tangent feature kernel of the pre-activations. As widths grow, these alignments become almost surely perfect [2402.05271].

- For general convolutional architectures satisfying frame-like and Lipschitz conditions on their filters, nonlinearities, and pooling operators, discrete deep feature extraction theory establishes global and per-layer stability properties. This formalism justifies a general neural feature ansatz on functional, not matrix, grounds [1605.08283].

Critically, the ansatz can fail for nonlinear architectures: explicit counterexamples show that there is no universal exponent $\alpha>0$ ensuring $F_\ell \propto G_\ell^\alpha$ in networks with nonlinearities, even when exact data fitting is achieved [2510.15563].

## 3. Empirical Evidence and Correlational Behavior

Empirical studies confirm that, during training:

- At initialization, the correlation $\rho(F_\ell, G_\ell)$ is near zero.
- During and after training, $\rho(F_\ell, G_\ell)$ typically rises rapidly and stabilizes near 1.
- In modern convolutional networks (AlexNet, VGG, ResNet), the entrywise correlation between learned filter covariances and the AGOP in each convolutional layer exceeds 0.9—while correlations between initial and final covariances remain much lower (<0.3) [2309.00570].
- Nonlinear fully connected networks exhibit the same pattern: feature matrix–AGOP alignment arises spontaneously under SGD for a wide range of architectures, optimizers, and initializations [2212.13881, 2402.05271, 2510.15563]. In deep linear networks, the exponent in $F_1 \propto G^{1/L}$ is verified across depths $L=2$ to $5$ [2510.15563].

Visualization of Gram matrices and AGOPs in convolutional models shows the emergence of structured features (e.g., Gabor-like edge detectors; consistent spectral structure), even for very deep or highly engineered networks [2309.00570].

## 4. Algorithmic Consequences and Practical Applications

The NFA provides a concrete foundation for both theoretical understanding and practical advances:

- **Feature learning in kernel machines**: The NFA can be used to augment kernel methods with data-adaptive feature learning via the AGOP. The Recursive Feature Machine (RFM) alternates kernel regression fits with AGOP-based metric updates, upgrading classical kernels to match learned networks on tabular tasks [2212.13881]. In convolutional settings, the (Deep) Convolutional Recursive Feature Machine (ConvRFM) utilizes patchwise AGOPs to adapt convolutional kernels, attaining generalization on par with deep CNNs [2309.00570].

- **Initialization and architecture design**: Initializing filters with covariances informed by the AGOP, rather than isotropic random draws, can accelerate early training, especially in convolutional networks [2309.00570].

- **Interpretability**: AGOP computation on trained models illuminates the specific patch or activation directions to which the model is most sensitive, extending classical feature visualizations to arbitrary layers and architectures [2309.00570, 2212.13881].

- **Practical optimizers**: New layerwise update rules (e.g., "speed-layer optimizer") designed to maximize the change per SGD step enhance NFA alignment and empirically improve feature quality [2402.05271].

- **Explaining deep learning phenomena**: The NFA accounts for the emergence of simplicity/spurious feature biases, the lottery ticket effect (sparsity and performance upon pruning), and phase transitions in generalization ("grokking") [2212.13881].

## 5. Comparative Table of NFA Manifestations

| Context                | NFA Statement                      | Proven/Empirical Law                |
|------------------------|------------------------------------|-------------------------------------|
| Deep linear networks   | $F_1 \propto G^{1/L}$              | Theoretical (any $L$, balanced init) [2510.15563] |
| Fully-connected nets   | $F_\ell \propto G_\ell$            | Empirical (early SGD, wide nets) [2212.13881, 2402.05271] |
| CNNs (conv layers)     | $Cov(W_\ell) \propto$ AGOP$_\ell$  | Empirical $r > 0.9$ (ImageNet nets) [2309.00570] |
| Kernel methods (RFM)   | $M \leftarrow$ AGOP update         | SOTA tabular performance            [2212.13881] |
| Nonlinear networks     | $F_\ell \propto G_\ell^{\alpha}$   | Fails in some ReLU/oscillatory cases [2510.15563] |

## 6. Limitations, Open Problems, and Theoretical Boundaries

The NFA is not universally valid. For architectures with nontrivial nonlinearities, explicit counterexamples demonstrate that feature Gram–AGOP alignment does not always occur, even with perfect fitting [2510.15563]. The dependence of alignment strength on finite width, mini-batch stochasticity, optimizer choice, and learning schedules is not yet fully characterized.

In linear cases, the depth-dependent exponent ($\alpha=1/L$) introduces a trade-off: shallow networks focus feature learning at earlier layers, while deeper networks distribute representation over more layers, potentially affecting the sample complexity and interpretability of learned features [2510.15563].

A rigorous link between NFA alignment and generalization remains an open problem. High AGOP–Gram alignment is neither necessary nor sufficient for good test accuracy, particularly in the presence of spurious correlations or data structure incompatible with the network architecture.

## 7. Broader Connections and Theoretical Implications

The neural feature ansatz unifies empirical observations regarding feature selection, inductive bias, and signal adaptivity in deep learning. It connects kernel methods and neural networks via shared AGOP-based adaptation principles [2309.00570, 2212.13881], and provides a mechanistic explanation for the success of specific design choices (e.g., small $3\times3$ kernels, max-pooling) [2309.00570].

Theoretical frameworks—ranging from frame theory in deep feature extraction [1605.08283] to spectral analysis of SGD dynamics [2402.05271]—highlight the centrality of first-order gradient information in shaping the representations learned by deep models. This suggests that the NFA may serve as a foundational tool for future advances in training algorithms, robust initialization, and interpretable deep networks.

Source: https://www.emergentmind.com/topics/deep-neural-feature-ansatz-nfa