Renormalized NNGP Kernel in Bayesian Neural Networks
- The paper introduces that the renormalized NNGP kernel modifies the infinite-width baseline to capture finite-width, data-dependent output correlations via a saddle-point derived order parameter.
- It details how the proportional limit retains 1/N effects by rescaling the kernel shape with a nontrivial Q*, thereby accounting for inter-output couplings absent in lazy-training regimes.
- The study compares renormalization in Bayesian MLPs, NN–QFT, and practical GP regression, highlighting its impact on generalization performance and predictive mean-variance behavior.
The renormalized NNGP kernel is a modification of the neural network Gaussian process kernel that incorporates effects absent in the strict infinite-width limit. In Bayesian one-hidden-layer networks with multiple readout neurons, the central finite-width effect is a data-dependent deformation of the infinite-width NNGP covariance by an output-space order parameter, which explains output-output correlations that vanish in the lazy-training regime (Baglioni et al., 2024). In related parts of the literature, the same phrase also refers to renormalization-group corrections in the neural network–quantum field theory correspondence, and to variance normalization procedures required to make practical NNGP kernels valid for Gaussian-process regression (Erbin et al., 2022, Muyskens et al., 2024). In deeper Bayesian architectures, an analogous proportional-regime construction rescales the large-width kernel through low-dimensional Wishart order parameters (Baglioni et al., 28 May 2026).
1. Infinite-width baseline and proportional-width setting
For a one-hidden-layer fully connected network with input dimension , hidden width , output dimension , activation , hidden weights , and readout weights drawn i.i.d. Gaussian, the output is
In the limit, any two outputs become independent Gaussian processes with covariance
This is the standard NNGP kernel for the one-hidden-layer model (Baglioni et al., 2024).
The proportional limit changes the asymptotic regime. Instead of sending width to infinity at fixed dataset size, one takes the number of training examples and the hidden width 0 to infinity together while keeping
1
finite. This regime retains the leading 2 effects that survive when 3 grows proportionally to 4. In the manuscript on kernel shape renormalization, these surviving terms are precisely what account for the non-trivial output-output correlations seen in finite Bayesian one-hidden-layer networks (Baglioni et al., 2024).
A basic consequence is that the infinite-width statement “different outputs are uncorrelated Gaussians” ceases to be exact once 5 is finite. The renormalized NNGP kernel is the object that re-expresses this finite-width, finite-sample coupling in kernel form.
2. Saddle-point derivation of kernel shape renormalization
The Bayesian formulation starts from the partition function
6
with 7 and 8 the MSE over 9 examples. In the joint 0 limit at fixed 1, one introduces a 2 positive-definite order parameter 3 and obtains, by saddle-point methods together with a Gaussian-equivalence theorem for the pre-activations,
4
The action is
5
where 6 is the 7 input matrix, 8 the 9 label vector, and
0
In components,
1
This form of the action is stated in the manuscript as following Pacelli et al., Nature Machine Intelligence (2023) (Baglioni et al., 2024).
The renormalized kernel is obtained by evaluating the integral at its minimizer 2. At that saddle,
3
The renormalization is therefore not a change in the scalar input-space covariance alone; it is a deformation of the kernel’s shape in output space. At 4, one recovers the standard lazy-training result because the saddle is 5. At finite 6, the minimizer is nontrivial, 7, and the effective kernel acquires inter-output structure (Baglioni et al., 2024).
3. Perturbative structure, assumptions, and asymptotic limits
The order parameter can be expanded perturbatively as
8
so that the renormalized kernel becomes
9
The manuscript provides a first-order closed expression for 0 in terms of 1, 2, 3, the labels 4, and the matrices 5; the resulting expansion is explicitly described as “1-loop in 6” (Baglioni et al., 2024).
Two asymptotic limits organize the interpretation. First, as 7, one has 8, hence 9, and the shape renormalization vanishes. Second, as 0, the prior becomes negligible and 1 is dominated by the 2 term; in practice, off-diagonal entries of 3 can become 4, leading to strong inter-output coupling (Baglioni et al., 2024).
The derivation rests on two explicit approximations. The first is Gaussian equivalence of the pre-activations in the 5 limit, identified in the manuscript with Breuer–Major type CLTs. The second is the saddle-point evaluation of the 6 integral, which neglects 7 fluctuations around 8 (Baglioni et al., 2024). These assumptions are important because they delimit the regime in which the renormalized kernel is expected to be quantitatively accurate.
4. Output-output correlations, weight overlaps, and predictive consequences
In the proportional-limit formulation, the off-diagonal entries of 9 encode the correlations between different outputs. In the infinite-width regime, 0, so off-diagonals vanish and different outputs are uncorrelated Gaussians. At finite 1, nonzero 2 quantify data-induced correlations between outputs 3 and 4 (Baglioni et al., 2024).
The same information can be read directly in the geometry of the last-layer weights:
5
Thus the renormalized kernel does not merely reproduce output covariance phenomenology; it also measures overlaps between the readout vectors of distinct classes. In this sense, kernel shape renormalization restores feature-learning effects that are absent in the pure lazy limit (Baglioni et al., 2024).
The predictive role is equally direct. The manuscript states that the same renormalized kernel appears in the predictive mean and variance of the GP posterior. Consequently, nontrivial 6 can improve or worsen generalization depending on whether the induced inter-output structure matches true label correlations (Baglioni et al., 2024).
The reported numerical experiments test both generalization and correlations. For synthetic, MNIST, and CIFAR10 inputs, Figure 1 compares the predicted generalization loss
7
against Langevin-sampled Bayesian one-hidden-layer networks and finds excellent agreement for hidden widths up to a few thousand. Relative errors on 8 are typically below a few percent for 9. Figure 2 compares the predicted 0 against the Monte Carlo average overlaps 1 for CIFAR10 with 2 at 3 and 4; both diagonal and off-diagonal elements match quantitatively, including negative overlaps among confusable classes. For 5, off-diagonal 6 for visually similar CIFAR classes such as “cars” versus “trucks” (Baglioni et al., 2024).
5. Distinct meanings of “renormalized NNGP kernel” in the literature
The literature uses the term in several technically distinct senses.
| Setting | Renormalized object | Mechanism |
|---|---|---|
| Proportional Bayesian one-hidden-layer network (Baglioni et al., 2024) | 7 | Output-space shape renormalization by saddle-point matrix |
| NN–QFT correspondence (Erbin et al., 2022) | 8 | One-loop finite-9 interaction, counterterms, RG flow |
| Practical GP use of infinite-width ReLU kernel (Muyskens et al., 2024) | 0 | Variance normalization or unit-sphere embedding |
In the neural network–quantum field theory correspondence, Erbin et al. describe the infinite-width limit as a free field theory and finite-1 corrections as interactions. For a one-dimensional translation-invariant Gaussian-activation kernel with 2, the bare momentum-space kernel is
3
and the leading finite-width correction is a local 4 interaction with bare coupling 5. The one-loop self-energy is
6
and after subtraction the renormalized kernel at scale 7 is written as
8
Holding the bare variance fixed yields the beta function
9
In this usage, renormalization means scale dependence of the kernel induced by finite-width interactions (Erbin et al., 2022).
In practical Gaussian-process regression with infinite-width ReLU kernels, “renormalization” instead denotes variance normalization. The unnormalized covariance is 0, and the normalized kernel is
1
The paper recommends embedding each coordinate by
2
and then applying 3 normalization, which enforces 4 constant over all 5 and keeps every coordinate in 6. The same work emphasizes several numerical pathologies: positive-definiteness failures, a valid hyperparameter region that shrinks rapidly with depth, and near-singularity even inside the valid region. Its recommended mitigations are unit-hypersphere normalization, fine-grid verification of positive definiteness via Cholesky decomposition, restriction to small depth such as 7, and addition of a modest nugget 8 (Muyskens et al., 2024).
That practical study also argues that the renormalized NNGP correlation behaves very much like a Matérn-9 correlation. The small-distance expansion of 00 is
01
which matches the Matérn-02 derivative at the origin when 03. On Friedman, MODIS, and Borehole benchmarks, each run 04 times with 05 training and 06 test points, the minimum RMSEs reported for NNGP were 07, 08, and 09; for Matérn 10 with fixed 11, 12, 13, and 14; and for optimized Matérn 15, 16, 17, and 18. The maximum RMSE for NNGP on MODIS was approximately 19 because of near-singularity episodes, while weight-difference statistics indicated that Matérn-20 reproduced NNGP kriging weights almost exactly (Muyskens et al., 2024).
6. Deep proportional-regime extensions and local renormalization in CNNs
A recent extension generalizes proportional-regime kernel renormalization from shallow networks to Bayesian MLPs of fixed depth 21 through the Equivalent Wishart Ansatz (EWA). For a fully connected Bayesian DNN with Gaussian priors of precision 22, squared-error likelihood at inverse temperature 23, data Gram matrix
24
and one-layer NNGP map
25
the proportional limit 26 with fixed 27 yields an effective GP posterior with renormalized kernel
28
The order parameters 29 are obtained by minimizing an effective action 30 (Baglioni et al., 28 May 2026).
The EWA replaces the full hierarchy of empirical kernel fluctuations by Wishart variables with the same mean. This leads to an effective action of the form
31
with 32 and
33
The saddle-point equations are
34
When all hidden layers have the same width, symmetry gives 35 for all 36, so 37. At zero temperature, defining
38
the shallow case 39 has the explicit solution
40
while for general 41 one obtains
42
The Bayes-optimal predictor then has exactly the mean and variance of GP regression with kernel 43 (Baglioni et al., 28 May 2026).
The same paper extends the construction to convolutional architectures. For CNNs, the renormalization becomes local rather than global: instead of a single scalar or a single output-space matrix, one obtains a patch-patch matrix structure. In the one-hidden-layer CNN derivation, the prior characteristic function reduces to an integral over a Wishart order-parameter matrix 44, and for deep CNNs the effective kernel takes the form
45
with 46 built from layerwise Wishart factors. Empirically, the EWA theory is tested against posterior sampling in finite deep networks with depths 47 and 48 on classic benchmark datasets, where it shows overall very good agreement together with two distinct types of systematic deviations (Baglioni et al., 28 May 2026).
Taken together, these results define the renormalized NNGP kernel as the principal mechanism by which finite width, finite sample size, or explicit normalization modifies the kernel inherited from infinite-width theory. In shallow Bayesian networks this modification is a nontrivial output-space deformation; in the NN–QFT correspondence it is an RG flow of kernel parameters; in practical GP implementations it is a variance normalization required for positive-definite covariance matrices; and in deep proportional-width models it becomes a low-dimensional or local renormalization of the hierarchical kernel itself.