---
title: Renormalized NNGP Kernel in Bayesian Neural Networks
url: https://www.emergentmind.com/topics/renormalized-nngp-kernel
type: topic
---

# Renormalized NNGP Kernel in Bayesian Neural Networks

The renormalized NNGP kernel is a modification of the neural network Gaussian process kernel that incorporates effects absent in the strict infinite-width limit. In Bayesian one-hidden-layer networks with multiple readout neurons, the central finite-width effect is a data-dependent deformation of the infinite-width NNGP covariance by an output-space order parameter, which explains output-output correlations that vanish in the lazy-training regime [2412.15911]. In related parts of the literature, the same phrase also refers to renormalization-group corrections in the neural network–quantum field theory correspondence, and to variance normalization procedures required to make practical NNGP kernels valid for Gaussian-process regression [2212.11811] [2410.08311]. In deeper Bayesian architectures, an analogous proportional-regime construction rescales the large-width kernel through low-dimensional Wishart order parameters [2605.29684].

## 1. Infinite-width baseline and proportional-width setting

For a one-hidden-layer fully connected network with input dimension $N_0$, hidden width $N_1$, output dimension $D$, activation $\phi$, hidden weights $w\in\mathbb R^{N_1\times N_0}$, and readout weights $v\in\mathbb R^{D\times N_1}$ drawn i.i.d. Gaussian, the output is
$$
f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.
$$
In the $N_1\to\infty$ limit, any two outputs become independent Gaussian processes with covariance
$$
K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].
$$
This is the standard NNGP kernel for the one-hidden-layer model [2412.15911].

The proportional limit changes the asymptotic regime. Instead of sending width to infinity at fixed dataset size, one takes the number of training examples $P$ and the hidden width $N_1$ to infinity together while keeping
$$
\alpha = \frac{P}{N_1}
$$
finite. This regime retains the leading $1/N_1$ effects that survive when $P$ grows proportionally to $N_1$. In the manuscript on kernel shape renormalization, these surviving terms are precisely what account for the non-trivial output-output correlations seen in finite Bayesian one-hidden-layer networks [2412.15911].

A basic consequence is that the infinite-width statement “different outputs are uncorrelated Gaussians” ceases to be exact once $\alpha$ is finite. The renormalized NNGP kernel is the object that re-expresses this finite-width, finite-sample coupling in kernel form.

## 2. Saddle-point derivation of kernel shape renormalization

The Bayesian formulation starts from the partition function
$$
Z=\int d\mu(w,v)\,\exp\{-\beta\,\mathcal L(w,v)\},
$$
with $\beta=1/T$ and $\mathcal L$ the MSE over $P$ examples. In the joint $P,N_1\to\infty$ limit at fixed $\alpha$, one introduces a $D\times D$ positive-definite order parameter $Q$ and obtains, by saddle-point methods together with a Gaussian-equivalence theorem for the pre-activations,
$$
Z \simeq \int dQ\,\exp\left\{-\frac{N_1}{2}S_{\rm FC}(Q)\right\}.
$$
The action is
$$
S_{\rm FC}(Q)=\operatorname{Tr}Q-\log\det Q
+\frac{\alpha}{P}\,y^T\!\Big[(1/\beta)I_{D\otimes P}+K_Q^{\rm(R)}(X,X)\Big]^{-1}y
+\frac{\alpha}{P}\,\operatorname{Tr}\log\Big[(1/\beta)I_{D\otimes P}+K_Q^{\rm(R)}(X,X)\Big],
$$
where $X$ is the $P\times N_0$ input matrix, $y$ the $P\cdot D$ label vector, and
$$
K_Q^{\rm(R)}(X,X)=Q\otimes K^{\rm NNGP}(X,X).
$$
In components,
$$
(K^{\rm(R)})_{a\mu,b\nu}=Q_{ab}\,K^{\rm NNGP}(x^\mu,x^\nu).
$$
This form of the action is stated in the manuscript as following Pacelli et al., *Nature Machine Intelligence* (2023) [2412.15911].

The renormalized kernel is obtained by evaluating the integral at its minimizer $Q^*(\alpha)$. At that saddle,
$$
[K_{\rm ren}(x,x';\alpha)]_{ab}=Q^*_{ab}(\alpha)\,K^{(\infty)}(x,x').
$$
The renormalization is therefore not a change in the scalar input-space covariance alone; it is a deformation of the kernel’s shape in output space. At $\alpha\to 0$, one recovers the standard lazy-training result because the saddle is $Q^*=I_D$. At finite $\alpha$, the minimizer is nontrivial, $Q^*(\alpha)\neq I_D$, and the effective kernel acquires inter-output structure [2412.15911].

## 3. Perturbative structure, assumptions, and asymptotic limits

The order parameter can be expanded perturbatively as
$$
Q = I + \alpha\,Q^{(1)} + O(\alpha^2),
$$
so that the renormalized kernel becomes
$$
K_{\rm ren}(x,x';\alpha)=K^{(\infty)}(x,x')\Big[I+\alpha\,Q^{(1)}+O(\alpha^2)\Big].
$$
The manuscript provides a first-order closed expression for $Q^{(1)}$ in terms of $K$, $\beta$, $\lambda_1$, the labels $y$, and the matrices $E_{ab}$; the resulting expansion is explicitly described as “1-loop in $\alpha$” [2412.15911].

Two asymptotic limits organize the interpretation. First, as $\alpha\to 0$, one has $Q^*\to I$, hence $K_{\rm ren}\to K^{(\infty)}$, and the shape renormalization vanishes. Second, as $\alpha\to\infty$, the prior becomes negligible and $Q^*$ is dominated by the $y^T[\cdots]y$ term; in practice, off-diagonal entries of $Q^*$ can become $O(1)$, leading to strong inter-output coupling [2412.15911].

The derivation rests on two explicit approximations. The first is Gaussian equivalence of the pre-activations in the $P,N_1\to\infty$ limit, identified in the manuscript with Breuer–Major type CLTs. The second is the saddle-point evaluation of the $Q$ integral, which neglects $1/N_1$ fluctuations around $Q^*$ [2412.15911]. These assumptions are important because they delimit the regime in which the renormalized kernel is expected to be quantitatively accurate.

## 4. Output-output correlations, weight overlaps, and predictive consequences

In the proportional-limit formulation, the off-diagonal entries of $Q^*$ encode the correlations between different outputs. In the infinite-width regime, $Q^*=I$, so off-diagonals vanish and different outputs are uncorrelated Gaussians. At finite $\alpha$, nonzero $Q^*_{ab}$ quantify data-induced correlations between outputs $a$ and $b$ [2412.15911].

The same information can be read directly in the geometry of the last-layer weights:
$$
\frac{\langle v_a\cdot v_b\rangle}{N_1}=\frac{Q^*_{ab}}{\lambda_1}.
$$
Thus the renormalized kernel does not merely reproduce output covariance phenomenology; it also measures overlaps between the readout vectors of distinct classes. In this sense, kernel shape renormalization restores feature-learning effects that are absent in the pure lazy limit [2412.15911].

The predictive role is equally direct. The manuscript states that the same renormalized kernel appears in the predictive mean and variance of the GP posterior. Consequently, nontrivial $Q^*$ can improve or worsen generalization depending on whether the induced inter-output structure matches true label correlations [2412.15911].

The reported numerical experiments test both generalization and correlations. For synthetic, MNIST, and CIFAR10 inputs, Figure 1 compares the predicted generalization loss
$$
\epsilon_g=\|y-\Gamma\|^2+\operatorname{Tr}\Sigma
$$
against Langevin-sampled Bayesian one-hidden-layer networks and finds excellent agreement for hidden widths up to a few thousand. Relative errors on $\epsilon_g$ are typically below a few percent for $N_1\gtrsim 500$. Figure 2 compares the predicted $Q^*/\lambda_1$ against the Monte Carlo average overlaps $vv^T/N_1$ for CIFAR10 with $D=10$ at $\alpha=0.1$ and $\alpha=1.0$; both diagonal and off-diagonal elements match quantitatively, including negative overlaps among confusable classes. For $\alpha=1$, off-diagonal $Q^*_{ab}\approx -0.1$ for visually similar CIFAR classes such as “cars” versus “trucks” [2412.15911].

## 5. Distinct meanings of “renormalized NNGP kernel” in the literature

The literature uses the term in several technically distinct senses.

| Setting | Renormalized object | Mechanism |
|---|---|---|
| Proportional Bayesian one-hidden-layer network [2412.15911] | $K_{\rm ren}(x,x';\alpha)=Q^*(\alpha)\,K^{(\infty)}(x,x')$ | Output-space shape renormalization by saddle-point matrix |
| NN–QFT correspondence [2212.11811] | $K_R(p;\mu)$ | One-loop finite-$N$ interaction, counterterms, RG flow |
| Practical GP use of infinite-width ReLU kernel [2410.08311] | $K(x,x')=K_0(x,x')/\sqrt{K_0(x,x)K_0(x',x')}$ | Variance normalization or unit-sphere embedding |

In the neural network–quantum field theory correspondence, Erbin et al. describe the infinite-width limit as a free field theory and finite-$N$ corrections as interactions. For a one-dimensional translation-invariant Gaussian-activation kernel with $\sigma_b=0$, the bare momentum-space kernel is
$$
K_0(p)=\sigma_w^2\,f(p),
$$
and the leading finite-width correction is a local $\phi^4$ interaction with bare coupling $u_{4,0}\sim +c/N$. The one-loop self-energy is
$$
\Sigma(p)=\frac12\,u_{4,0}\int_q K_0(q),
$$
and after subtraction the renormalized kernel at scale $\mu$ is written as
$$
K_R(p;\mu)=\sigma_w^2(\mu)\,f(p)+\Delta K(p)-\text{counterterm}.
$$
Holding the bare variance fixed yields the beta function
$$
\frac{d\sigma_w^2}{d\ln\mu}=\beta(\sigma_w^2)=\frac12\,u_4(\mu)\,\sigma_w^2(\mu)+\mathcal O(u_4^2).
$$
In this usage, renormalization means scale dependence of the kernel induced by finite-width interactions [2212.11811].

In practical Gaussian-process regression with infinite-width ReLU kernels, “renormalization” instead denotes variance normalization. The unnormalized covariance is $K_0(x,x')=\Sigma^{(M)}(x,x')$, and the normalized kernel is
$$
K(x,x')=\frac{K_0(x,x')}{\sqrt{K_0(x,x)\,K_0(x',x')}}.
$$
The paper recommends embedding each coordinate by
$$
[x_i]\rightarrow [\cos(\pi x_i),\sin(\pi x_i)]
$$
and then applying $L_2$ normalization, which enforces $\Sigma^{(1)}(x,x)=\sigma_a^2+\sigma_b^2$ constant over all $x$ and keeps every coordinate in $[-1,1]$. The same work emphasizes several numerical pathologies: positive-definiteness failures, a valid hyperparameter region that shrinks rapidly with depth, and near-singularity even inside the valid region. Its recommended mitigations are unit-hypersphere normalization, fine-grid verification of positive definiteness via Cholesky decomposition, restriction to small depth such as $M=2$, and addition of a modest nugget $\tau^2>0$ [2410.08311].

That practical study also argues that the renormalized NNGP correlation behaves very much like a Matérn-$3/2$ correlation. The small-distance expansion of $C_{\rm NNGP}(r)$ is
$$
C_{\rm NNGP}(r)=1-\beta r+O(r^{3/2}),
$$
which matches the Matérn-$3/2$ derivative at the origin when $\beta=\sqrt{3}/\ell$. On Friedman, MODIS, and Borehole benchmarks, each run $100$ times with $n=500$ training and $m=500$ test points, the minimum RMSEs reported for NNGP were $0.054$, $4.29$, and $1.52$; for Matérn $\nu=3/2$ with fixed $\ell=1$, $0.026$, $4.89$, and $1.29$; and for optimized Matérn $\nu=3/2$, $0.005$, $2.22$, and $0.41$. The maximum RMSE for NNGP on MODIS was approximately $4.4\times 10^3$ because of near-singularity episodes, while weight-difference statistics indicated that Matérn-$3/2$ reproduced NNGP kriging weights almost exactly [2410.08311].

## 6. Deep proportional-regime extensions and local renormalization in CNNs

A recent extension generalizes proportional-regime kernel renormalization from shallow networks to Bayesian MLPs of fixed depth $L$ through the Equivalent Wishart Ansatz (EWA). For a fully connected Bayesian DNN with Gaussian priors of precision $\lambda$, squared-error likelihood at inverse temperature $\beta$, data Gram matrix
$$
C=\frac{XX^T}{\lambda N_0},
$$
and one-layer NNGP map
$$
\Theta(K)_{\mu\nu}=\frac{1}{\lambda}\,\mathbb E_{h\sim N(0,K)}[\sigma(h^\mu)\sigma(h^\nu)],
$$
the proportional limit $P,N_1,\ldots,N_L\to\infty$ with fixed $\alpha_\ell=P/N_\ell$ yields an effective GP posterior with renormalized kernel
$$
K_R = Q^*\,\Theta^L(C), \qquad Q^*=\prod_{\ell=1}^L q_\ell^*.
$$
The order parameters $q_\ell^*$ are obtained by minimizing an effective action $S(q_1,\ldots,q_L)$ [2605.29684].

The EWA replaces the full hierarchy of empirical kernel fluctuations by Wishart variables with the same mean. This leads to an effective action of the form
$$
S(q_1,\ldots,q_L)=\sum_\ell I_q(q_\ell)+U(q_1,\ldots,q_L),
$$
with $I_q(q_\ell)=\frac12(q_\ell-1-\ln q_\ell)$ and
$$
K_R = \Big(\prod_\ell q_\ell\Big)\Theta^L(C).
$$
The saddle-point equations are
$$
\frac{\partial S}{\partial q_\ell}\bigg|_{q^*}=0.
$$
When all hidden layers have the same width, symmetry gives $q_\ell^*=q^*$ for all $\ell$, so $Q^*=q^{*L}$. At zero temperature, defining
$$
M_{yy}\equiv \frac1P\,y^T[\Theta^L(C)]^{-1}y,
$$
the shallow case $L=1$ has the explicit solution
$$
q^*=-\frac12(\alpha-1)+\sqrt{\Big(\frac{\alpha-1}{2}\Big)^2+\alpha M_{yy}},
$$
while for general $L$ one obtains
$$
q^{*L}(q^*+\alpha-1)=\alpha M_{yy}.
$$
The Bayes-optimal predictor then has exactly the mean and variance of GP regression with kernel $K_R^*$ [2605.29684].

The same paper extends the construction to convolutional architectures. For CNNs, the renormalization becomes local rather than global: instead of a single scalar or a single output-space matrix, one obtains a patch-patch matrix structure. In the one-hidden-layer CNN derivation, the prior characteristic function reduces to an integral over a Wishart order-parameter matrix $q$, and for deep CNNs the effective kernel takes the form
$$
K_R = \frac{1}{N_{p_L}}\sum_{i,j} Q_{ij}\,[(\Theta\circ\omega)^L(C)]_{ji},
$$
with $Q$ built from layerwise Wishart factors. Empirically, the EWA theory is tested against posterior sampling in finite deep networks with depths $L\sim O(10)$ and $P\sim O(10^3)$ on classic benchmark datasets, where it shows overall very good agreement together with two distinct types of systematic deviations [2605.29684].

Taken together, these results define the renormalized NNGP kernel as the principal mechanism by which finite width, finite sample size, or explicit normalization modifies the kernel inherited from infinite-width theory. In shallow Bayesian networks this modification is a nontrivial output-space deformation; in the NN–QFT correspondence it is an RG flow of kernel parameters; in practical GP implementations it is a variance normalization required for positive-definite covariance matrices; and in deep proportional-width models it becomes a low-dimensional or local renormalization of the hierarchical kernel itself.

Source: https://www.emergentmind.com/topics/renormalized-nngp-kernel