Papers
Topics
Authors
Recent
Search
2000 character limit reached

Renormalized NNGP Kernel in Bayesian Neural Networks

Updated 14 July 2026
  • The paper introduces that the renormalized NNGP kernel modifies the infinite-width baseline to capture finite-width, data-dependent output correlations via a saddle-point derived order parameter.
  • It details how the proportional limit retains 1/N effects by rescaling the kernel shape with a nontrivial Q*, thereby accounting for inter-output couplings absent in lazy-training regimes.
  • The study compares renormalization in Bayesian MLPs, NN–QFT, and practical GP regression, highlighting its impact on generalization performance and predictive mean-variance behavior.

The renormalized NNGP kernel is a modification of the neural network Gaussian process kernel that incorporates effects absent in the strict infinite-width limit. In Bayesian one-hidden-layer networks with multiple readout neurons, the central finite-width effect is a data-dependent deformation of the infinite-width NNGP covariance by an output-space order parameter, which explains output-output correlations that vanish in the lazy-training regime (Baglioni et al., 2024). In related parts of the literature, the same phrase also refers to renormalization-group corrections in the neural network–quantum field theory correspondence, and to variance normalization procedures required to make practical NNGP kernels valid for Gaussian-process regression (Erbin et al., 2022, Muyskens et al., 2024). In deeper Bayesian architectures, an analogous proportional-regime construction rescales the large-width kernel through low-dimensional Wishart order parameters (Baglioni et al., 28 May 2026).

1. Infinite-width baseline and proportional-width setting

For a one-hidden-layer fully connected network with input dimension N0N_0, hidden width N1N_1, output dimension DD, activation ϕ\phi, hidden weights wRN1×N0w\in\mathbb R^{N_1\times N_0}, and readout weights vRD×N1v\in\mathbb R^{D\times N_1} drawn i.i.d. Gaussian, the output is

fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.

In the N1N_1\to\infty limit, any two outputs become independent Gaussian processes with covariance

K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].

This is the standard NNGP kernel for the one-hidden-layer model (Baglioni et al., 2024).

The proportional limit changes the asymptotic regime. Instead of sending width to infinity at fixed dataset size, one takes the number of training examples PP and the hidden width N1N_10 to infinity together while keeping

N1N_11

finite. This regime retains the leading N1N_12 effects that survive when N1N_13 grows proportionally to N1N_14. In the manuscript on kernel shape renormalization, these surviving terms are precisely what account for the non-trivial output-output correlations seen in finite Bayesian one-hidden-layer networks (Baglioni et al., 2024).

A basic consequence is that the infinite-width statement “different outputs are uncorrelated Gaussians” ceases to be exact once N1N_15 is finite. The renormalized NNGP kernel is the object that re-expresses this finite-width, finite-sample coupling in kernel form.

2. Saddle-point derivation of kernel shape renormalization

The Bayesian formulation starts from the partition function

N1N_16

with N1N_17 and N1N_18 the MSE over N1N_19 examples. In the joint DD0 limit at fixed DD1, one introduces a DD2 positive-definite order parameter DD3 and obtains, by saddle-point methods together with a Gaussian-equivalence theorem for the pre-activations,

DD4

The action is

DD5

where DD6 is the DD7 input matrix, DD8 the DD9 label vector, and

ϕ\phi0

In components,

ϕ\phi1

This form of the action is stated in the manuscript as following Pacelli et al., Nature Machine Intelligence (2023) (Baglioni et al., 2024).

The renormalized kernel is obtained by evaluating the integral at its minimizer ϕ\phi2. At that saddle,

ϕ\phi3

The renormalization is therefore not a change in the scalar input-space covariance alone; it is a deformation of the kernel’s shape in output space. At ϕ\phi4, one recovers the standard lazy-training result because the saddle is ϕ\phi5. At finite ϕ\phi6, the minimizer is nontrivial, ϕ\phi7, and the effective kernel acquires inter-output structure (Baglioni et al., 2024).

3. Perturbative structure, assumptions, and asymptotic limits

The order parameter can be expanded perturbatively as

ϕ\phi8

so that the renormalized kernel becomes

ϕ\phi9

The manuscript provides a first-order closed expression for wRN1×N0w\in\mathbb R^{N_1\times N_0}0 in terms of wRN1×N0w\in\mathbb R^{N_1\times N_0}1, wRN1×N0w\in\mathbb R^{N_1\times N_0}2, wRN1×N0w\in\mathbb R^{N_1\times N_0}3, the labels wRN1×N0w\in\mathbb R^{N_1\times N_0}4, and the matrices wRN1×N0w\in\mathbb R^{N_1\times N_0}5; the resulting expansion is explicitly described as “1-loop in wRN1×N0w\in\mathbb R^{N_1\times N_0}6” (Baglioni et al., 2024).

Two asymptotic limits organize the interpretation. First, as wRN1×N0w\in\mathbb R^{N_1\times N_0}7, one has wRN1×N0w\in\mathbb R^{N_1\times N_0}8, hence wRN1×N0w\in\mathbb R^{N_1\times N_0}9, and the shape renormalization vanishes. Second, as vRD×N1v\in\mathbb R^{D\times N_1}0, the prior becomes negligible and vRD×N1v\in\mathbb R^{D\times N_1}1 is dominated by the vRD×N1v\in\mathbb R^{D\times N_1}2 term; in practice, off-diagonal entries of vRD×N1v\in\mathbb R^{D\times N_1}3 can become vRD×N1v\in\mathbb R^{D\times N_1}4, leading to strong inter-output coupling (Baglioni et al., 2024).

The derivation rests on two explicit approximations. The first is Gaussian equivalence of the pre-activations in the vRD×N1v\in\mathbb R^{D\times N_1}5 limit, identified in the manuscript with Breuer–Major type CLTs. The second is the saddle-point evaluation of the vRD×N1v\in\mathbb R^{D\times N_1}6 integral, which neglects vRD×N1v\in\mathbb R^{D\times N_1}7 fluctuations around vRD×N1v\in\mathbb R^{D\times N_1}8 (Baglioni et al., 2024). These assumptions are important because they delimit the regime in which the renormalized kernel is expected to be quantitatively accurate.

4. Output-output correlations, weight overlaps, and predictive consequences

In the proportional-limit formulation, the off-diagonal entries of vRD×N1v\in\mathbb R^{D\times N_1}9 encode the correlations between different outputs. In the infinite-width regime, fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.0, so off-diagonals vanish and different outputs are uncorrelated Gaussians. At finite fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.1, nonzero fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.2 quantify data-induced correlations between outputs fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.3 and fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.4 (Baglioni et al., 2024).

The same information can be read directly in the geometry of the last-layer weights:

fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.5

Thus the renormalized kernel does not merely reproduce output covariance phenomenology; it also measures overlaps between the readout vectors of distinct classes. In this sense, kernel shape renormalization restores feature-learning effects that are absent in the pure lazy limit (Baglioni et al., 2024).

The predictive role is equally direct. The manuscript states that the same renormalized kernel appears in the predictive mean and variance of the GP posterior. Consequently, nontrivial fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.6 can improve or worsen generalization depending on whether the induced inter-output structure matches true label correlations (Baglioni et al., 2024).

The reported numerical experiments test both generalization and correlations. For synthetic, MNIST, and CIFAR10 inputs, Figure 1 compares the predicted generalization loss

fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.7

against Langevin-sampled Bayesian one-hidden-layer networks and finds excellent agreement for hidden widths up to a few thousand. Relative errors on fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.8 are typically below a few percent for fa(x)=i=1N1va,iϕ ⁣(wixN0)1N1.f_a(x)=\sum_{i=1}^{N_1} v_{a,i}\,\phi\!\left(\frac{w_i\cdot x}{\sqrt{N_0}}\right)\frac{1}{\sqrt{N_1}}.9. Figure 2 compares the predicted N1N_1\to\infty0 against the Monte Carlo average overlaps N1N_1\to\infty1 for CIFAR10 with N1N_1\to\infty2 at N1N_1\to\infty3 and N1N_1\to\infty4; both diagonal and off-diagonal elements match quantitatively, including negative overlaps among confusable classes. For N1N_1\to\infty5, off-diagonal N1N_1\to\infty6 for visually similar CIFAR classes such as “cars” versus “trucks” (Baglioni et al., 2024).

5. Distinct meanings of “renormalized NNGP kernel” in the literature

The literature uses the term in several technically distinct senses.

Setting Renormalized object Mechanism
Proportional Bayesian one-hidden-layer network (Baglioni et al., 2024) N1N_1\to\infty7 Output-space shape renormalization by saddle-point matrix
NN–QFT correspondence (Erbin et al., 2022) N1N_1\to\infty8 One-loop finite-N1N_1\to\infty9 interaction, counterterms, RG flow
Practical GP use of infinite-width ReLU kernel (Muyskens et al., 2024) K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].0 Variance normalization or unit-sphere embedding

In the neural network–quantum field theory correspondence, Erbin et al. describe the infinite-width limit as a free field theory and finite-K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].1 corrections as interactions. For a one-dimensional translation-invariant Gaussian-activation kernel with K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].2, the bare momentum-space kernel is

K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].3

and the leading finite-width correction is a local K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].4 interaction with bare coupling K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].5. The one-loop self-energy is

K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].6

and after subtraction the renormalized kernel at scale K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].7 is written as

K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].8

Holding the bare variance fixed yields the beta function

K()(x,x)Ew ⁣[ϕ ⁣(wxN0)ϕ ⁣(wxN0)].K^{(\infty)}(x,x') \equiv \mathbb E_w\!\left[\phi\!\left(\frac{w\cdot x}{\sqrt{N_0}}\right)\phi\!\left(\frac{w\cdot x'}{\sqrt{N_0}}\right)\right].9

In this usage, renormalization means scale dependence of the kernel induced by finite-width interactions (Erbin et al., 2022).

In practical Gaussian-process regression with infinite-width ReLU kernels, “renormalization” instead denotes variance normalization. The unnormalized covariance is PP0, and the normalized kernel is

PP1

The paper recommends embedding each coordinate by

PP2

and then applying PP3 normalization, which enforces PP4 constant over all PP5 and keeps every coordinate in PP6. The same work emphasizes several numerical pathologies: positive-definiteness failures, a valid hyperparameter region that shrinks rapidly with depth, and near-singularity even inside the valid region. Its recommended mitigations are unit-hypersphere normalization, fine-grid verification of positive definiteness via Cholesky decomposition, restriction to small depth such as PP7, and addition of a modest nugget PP8 (Muyskens et al., 2024).

That practical study also argues that the renormalized NNGP correlation behaves very much like a Matérn-PP9 correlation. The small-distance expansion of N1N_100 is

N1N_101

which matches the Matérn-N1N_102 derivative at the origin when N1N_103. On Friedman, MODIS, and Borehole benchmarks, each run N1N_104 times with N1N_105 training and N1N_106 test points, the minimum RMSEs reported for NNGP were N1N_107, N1N_108, and N1N_109; for Matérn N1N_110 with fixed N1N_111, N1N_112, N1N_113, and N1N_114; and for optimized Matérn N1N_115, N1N_116, N1N_117, and N1N_118. The maximum RMSE for NNGP on MODIS was approximately N1N_119 because of near-singularity episodes, while weight-difference statistics indicated that Matérn-N1N_120 reproduced NNGP kriging weights almost exactly (Muyskens et al., 2024).

6. Deep proportional-regime extensions and local renormalization in CNNs

A recent extension generalizes proportional-regime kernel renormalization from shallow networks to Bayesian MLPs of fixed depth N1N_121 through the Equivalent Wishart Ansatz (EWA). For a fully connected Bayesian DNN with Gaussian priors of precision N1N_122, squared-error likelihood at inverse temperature N1N_123, data Gram matrix

N1N_124

and one-layer NNGP map

N1N_125

the proportional limit N1N_126 with fixed N1N_127 yields an effective GP posterior with renormalized kernel

N1N_128

The order parameters N1N_129 are obtained by minimizing an effective action N1N_130 (Baglioni et al., 28 May 2026).

The EWA replaces the full hierarchy of empirical kernel fluctuations by Wishart variables with the same mean. This leads to an effective action of the form

N1N_131

with N1N_132 and

N1N_133

The saddle-point equations are

N1N_134

When all hidden layers have the same width, symmetry gives N1N_135 for all N1N_136, so N1N_137. At zero temperature, defining

N1N_138

the shallow case N1N_139 has the explicit solution

N1N_140

while for general N1N_141 one obtains

N1N_142

The Bayes-optimal predictor then has exactly the mean and variance of GP regression with kernel N1N_143 (Baglioni et al., 28 May 2026).

The same paper extends the construction to convolutional architectures. For CNNs, the renormalization becomes local rather than global: instead of a single scalar or a single output-space matrix, one obtains a patch-patch matrix structure. In the one-hidden-layer CNN derivation, the prior characteristic function reduces to an integral over a Wishart order-parameter matrix N1N_144, and for deep CNNs the effective kernel takes the form

N1N_145

with N1N_146 built from layerwise Wishart factors. Empirically, the EWA theory is tested against posterior sampling in finite deep networks with depths N1N_147 and N1N_148 on classic benchmark datasets, where it shows overall very good agreement together with two distinct types of systematic deviations (Baglioni et al., 28 May 2026).

Taken together, these results define the renormalized NNGP kernel as the principal mechanism by which finite width, finite sample size, or explicit normalization modifies the kernel inherited from infinite-width theory. In shallow Bayesian networks this modification is a nontrivial output-space deformation; in the NN–QFT correspondence it is an RG flow of kernel parameters; in practical GP implementations it is a variance normalization required for positive-definite covariance matrices; and in deep proportional-width models it becomes a low-dimensional or local renormalization of the hierarchical kernel itself.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Renormalized NNGP Kernel.