---
title: Variational Renormalization Group (VRG)
url: https://www.emergentmind.com/topics/variational-renormalization-group-vrg
type: topic
---

# Variational Renormalization Group (VRG)

Variational Renormalization Group (VRG) denotes a family of renormalization schemes in which the coarse-grained description is selected by an explicit optimization principle rather than by a purely fixed blocking rule. In the literature, this includes Kadanoff-style real-space constructions with auxiliary hidden variables and trace conditions, Monte Carlo renormalization methods based on convex variational functionals, tensor-network algorithms in which effective subspaces or environments are optimized over MPS, PEPS, corner, or boundary ansätze, and neural or information-theoretic multiscale constructions that turn RG-like transformations into tractable variational problems [1410.3831][1707.08683][0907.2796][1802.02840]. The common structure is the same: an RG step is tied to minimizing or maximizing a quantity such as a free-energy difference, a projected Hamiltonian trace, a bias-potential functional, a partition-function density, or a multilevel entropy.

| Formulation | Variational object | Representative sources |
|---|---|---|
| Real-space/block-spin VRG | Coupling operator \( \mathbf T_\lambda \), free-energy difference, trace condition | [1410.3831] |
| Monte Carlo VRG | Bias potential \(V\) minimizing a convex functional \(\Omega[V]\) | [1707.08683][1810.09579] |
| Tensor-network VRG | MPS/PEPS tensors, corner or boundary environments, truncation isometries | [0907.2796][1102.1401][2203.17098][2508.10418] |
| Neural/information-theoretic VRG | Normalizing-flow density \(q\), hierarchical entropy/KL objectives | [1802.02840][2509.01424] |
| Adjacent variational RG formalisms | Optimal-transport functionals, RG-improved variational resummation | [2202.11737][2509.07250] |

## 1. Kadanoff-style variational RG and the hidden-variable construction

The canonical real-space formulation starts from microscopic binary spins \(\{v_i\}\), \(v_i\in\{\pm 1\}\), distributed according to
\[
P(\{v_i\})=\frac{e^{-{\mathbf H}(\{v_i\})}}{Z},\qquad
Z=\mathrm{Tr}_{v_i}e^{-{\mathbf H}(\{v_i\})},
\]
with a Hamiltonian expanded in coupling space as
\[
{\mathbf H}[\{v_i\}] = -\sum_i K_i v_i - \sum_{ij} K_{ij} v_i v_j - \sum_{ijk} K_{ijk} v_i v_j v_k + \ldots.
\]
VRG introduces \(M<N\) auxiliary binary variables \(\{h_j\}\), \(h_j\in\{\pm1\}\), and a variational coupling operator \(\mathbf T_\lambda(\{v_i\},\{h_j\})\). The coarse-grained Hamiltonian is then defined by
\[
e^{-\mathbf H^{RG}_\lambda[\{h_j\}]}
\equiv
\mathrm{Tr}_{v_i}\, e^{\mathbf T_\lambda(\{v_i\},\{h_j\})-\mathbf H(\{v_i\})}.
\]
The free-energy criterion is \( \Delta F = F_\lambda^h - F^v \), and the exact trace condition is
\[
\Delta F=0 \iff \mathrm{Tr}_{h_j} e^{\mathbf T_\lambda(\{v_i\},\{h_j\})}=1.
\]
This condition is stronger than approximate matching of a few observables: it inserts a normalized hidden-variable factor into the Boltzmann weight and preserves the partition function and free energy exactly [1410.3831].

A central later development is the exact model-class mapping between this Kadanoff variational RG and Restricted Boltzmann Machines (RBMs). If the RBM energy is
\[
{\mathbf E}(\{v_i\},\{h_j\})=\sum_j b_j h_j+\sum_{ij} v_i w_{ij} h_j+\sum_i c_i v_i,
\]
then the dictionary is
\[
\mathbf T(\{v_i\},\{h_j\})=-{\mathbf E}(\{v_i\},\{h_j\})+{\mathbf H}[\{v_i\}].
\]
Under this identification, the RG hidden-spin Hamiltonian equals the RBM hidden marginal Hamiltonian,
\[
\mathbf H_\lambda^{RG}[\{h_j\}] = \mathbf H_\lambda^{RBM}[\{h_j\}],
\]
and exact satisfaction of the trace condition implies
\[
P(\{v_i\})=p_\lambda(\{v_i\}),\qquad D_{KL}(P\|p_\lambda)=0.
\]
The mapping is exact at the level of formal correspondence between model classes, but the practical optimization criteria differ: VRG commonly minimizes \(\Delta F\), whereas RBM training minimizes \(D_{KL}(P\|p_\lambda)\) [1410.3831].

This formulation also fixes a common misconception. The deep-learning correspondence is narrow: it applies to Kadanoff-type variational RG and RBM-based architectures, not to all neural networks or all deep learning. The same source is explicit that stacked RBMs correspond to iterated coarse-graining across scales, but does not claim that all representation learning is literally RG [1410.3831].

## 2. Variational Monte Carlo RG and the bias-potential framework

A distinct but closely related branch of VRG reformulates Monte Carlo renormalization as optimization over a bias potential on coarse variables. For a lattice spin model with
\[
H(\boldsymbol{\sigma})=\sum_\alpha K_\alpha S_\alpha(\boldsymbol{\sigma}),
\qquad
\boldsymbol{\sigma}'=\tau(\boldsymbol{\sigma}),
\]
the exact coarse-grained Hamiltonian is
\[
H'(\boldsymbol{\sigma}')=
-\log\sum_{\boldsymbol{\sigma}}
\delta_{\tau(\boldsymbol{\sigma}),\boldsymbol{\sigma}'}\,e^{-H(\boldsymbol{\sigma})}.
\]
The variational object is a bias \(V(\boldsymbol{\sigma}')\) chosen so that the biased coarse-spin distribution equals a target distribution \(p_t(\boldsymbol{\sigma}')\). The convex functional is
\[
\Omega[V]
=
\log
\frac{\sum_{\boldsymbol{\sigma}'}e^{-[H'(\boldsymbol{\sigma}')+V(\boldsymbol{\sigma}')]}}
{\sum_{\boldsymbol{\sigma}'}e^{-H'(\boldsymbol{\sigma}')}}+
\sum_{\boldsymbol{\sigma}'}p_t(\boldsymbol{\sigma}')V(\boldsymbol{\sigma}').
\]
Its minimizer satisfies
\[
H'(\boldsymbol{\sigma}')
=
-\,V_{\min}(\boldsymbol{\sigma}')
-\log p_t(\boldsymbol{\sigma}')
+\text{constant},
\]
and hence \(p_{V_{\min}}(\boldsymbol{\sigma}')=p_t(\boldsymbol{\sigma}')\). For constant \(p_t\), the coarse variables become uncorrelated, and critical slowing down in the coarse degrees of freedom is strongly reduced [1707.08683].

In practice, one truncates to
\[
V_{\mathbf J}(\boldsymbol{\sigma}')=\sum_\alpha J_\alpha S_\alpha(\boldsymbol{\sigma}'),
\]
with gradient and Hessian
\[
\frac{\partial \Omega(\mathbf J)}{\partial J_\alpha}
=
-\langle S_\alpha(\boldsymbol{\sigma}')\rangle_{V_{\mathbf J}}
+\langle S_\alpha(\boldsymbol{\sigma}')\rangle_{p_t},
\]
\[
\frac{\partial^2 \Omega(\mathbf J)}{\partial J_\alpha\partial J_\beta}
=
\langle S_\alpha(\boldsymbol{\sigma}')S_\beta(\boldsymbol{\sigma}')\rangle_{V_{\mathbf J}}
-
\langle S_\alpha(\boldsymbol{\sigma}')\rangle_{V_{\mathbf J}}
\langle S_\beta(\boldsymbol{\sigma}')\rangle_{V_{\mathbf J}}.
\]
For uniform \(p_t\), the renormalized couplings are simply
\[
K'_\alpha=-J_{\min,\alpha}.
\]
This framework was demonstrated on the 2D Ising model, where the biased-ensemble estimates of the leading RG eigenvalues converge well for large lattices and outperform unbiased Monte Carlo on the same timescale [1707.08683].

The same bias-potential idea extends to quenched disorder. There the microscopic Hamiltonian is sample-dependent,
\[
H_K(\boldsymbol{\sigma})
=
-\sum_\alpha\sum_{i=1}^N\sum_{s=1}^{N_s(\alpha)}
K_\alpha^{i,s}S_\alpha^{i,s}(\boldsymbol{\sigma}),
\qquad K\sim P_v(K),
\]
and the RG map induces a flow of the disorder distribution,
\[
P_{v'}(K') = \int dK\, P_v(K)\,\delta(K'-\mathcal R(K)).
\]
The bias is again expanded in the operator basis, and one obtains samplewise renormalized couplings through \(K'=-J_{\min}\). The disorder law is parameterized by
\[
-\ln P_v(K)=C+\sum_\beta v_\beta U_\beta(K),
\]
so linearized RG can be carried out directly in the space of coupling distributions [1810.09579].

This formulation sharply distinguishes finite-disorder and strong-disorder fixed points. For the 2D dilute Ising model, the critical distribution flows to a finite-width non-Gaussian fixed distribution, and the leading even eigenvalue was found to be
\[
\lambda^e = 2.018(6),
\]
to be compared with \(2\) for the pure Ising model. For the random transverse-field Ising chain and the 3D random-field Ising model, the variance of renormalized couplings grows under RG, indicating strong-disorder behavior rather than finite-range closure of the coupling distribution [1810.09579]. A plausible implication is that VRG is especially effective when the renormalized disorder distribution remains describable by short-range coupling correlations.

## 3. Tensor-network VRG: MPS, PEPS, NRG, and boundary optimization

In tensor-network language, VRG is reformulated as variational optimization over structured low-entanglement state classes. For one-dimensional quantum systems, the fundamental ansatz is the MPS
\[
|\psi\rangle
=
\sum_{\alpha_1,\ldots,\alpha_N=1}^d
\mathrm{Tr}\!\left(
A^1_{\alpha_1}A^2_{\alpha_2}\cdots A^N_{\alpha_N}
\right)
|\alpha_1\rangle\cdots|\alpha_N\rangle,
\]
and the variational problem is
\[
\min_{|\psi^N\rangle\in\{\mathrm{MPS}_D\}}
\left[
\langle\psi^N|\mathcal H^N|\psi^N\rangle
-
\lambda\langle\psi^N|\psi^N\rangle
\right].
\]
Fixing all tensors except one yields the local generalized eigenvalue problem
\[
H_{\mathrm{eff}}x=\lambda N_{\mathrm{eff}}x.
\]
In this sense, DMRG is the variational optimization of an MPS manifold, while Wilson NRG can be viewed as a more limited one-way truncation scheme whose retained states already form an MPS. The review literature explicitly presents MPS and PEPS as both ansätze and “variational renormalization group methods” [0907.2796].

The entanglement rationale is encoded in the discarded Schmidt-weight criterion. If \(\epsilon_\alpha(D)\) is the tail of Schmidt eigenvalues across cut \(\alpha\), then there exists an MPS of bond dimension \(D\) with
\[
\bigl\| |\psi\rangle-|\psi_D\rangle \bigr\|^2
\le
2\sum_{\alpha=1}^{N-1}\epsilon_\alpha(D).
\]
Combined with area-law or logarithmic-entanglement scaling, this provides the approximation-theoretic basis for variational truncation in one dimension and motivates PEPS as the corresponding higher-dimensional extension [0907.2796].

The variational reinterpretation of numerical RG becomes explicit in the “variational numerical renormalization group” construction. If \(\mathbf H_{\rm ext}^{[n]}\) is the enlarged Hamiltonian at one NRG step and \(\mathbf A\) is the isometry selecting the retained subspace, the NRG truncation is equivalent to minimizing
\[
f(\mathbf A)=\mathrm{tr}\!\left(\mathbf A^H \mathbf H_{\rm ext}^{[n]}\mathbf A\right),
\qquad
\mathbf A^H\mathbf A=\mathbf 1.
\]
This is the sum of retained effective energies. Once that objective is identified, one can sweep over all local MPS tensors, not only the newest one, thereby importing DMRG-style feedback into NRG. The resulting method keeps \(O(D^3)\) scaling while substantially improving spectra for both quantum spin chains and the single impurity Anderson model, especially as the logarithmic discretization parameter approaches the continuum limit [1102.1401].

A separate extension carries variational RG ideas into multigrid PDE solvers. “Multigrid Renormalization” represents the solution on progressively finer dyadic grids by an MPS-like tensor train, truncates to a maximum bond dimension \(\chi\), and optimizes tensors by alternating local updates “in the spirit of alternating least squares.” Under the empirical assumptions that \(\chi\) remains nearly independent of level and the number of sweeps scales mildly, the overall cost becomes \(\mathcal O(\log N)\), and the method was demonstrated on nonlinear Schrödinger ground states up to \(N=10^{18}\) grid points in three dimensions [1802.07259]. This is not standard statistical-mechanics VRG, but it is a direct transplantation of variational renormalization ideas into numerical analysis.

Boundary and environment optimization produce a further tensor-network branch. In variational CTMRG, the fixed-point environment of an infinite 2D tensor network is reformulated as the bilevel optimization problem
\[
\max_E \frac{f(E,C^*(E))}{g(C^*(E);E)}
\quad\text{s.t.}\quad
C^*(E)=\arg\max_C g(C;E).
\]
The solution corresponds to the conventional CTMRG fixed-point environment, but the route to it is explicitly variational. This reformulation yields high-precision residual entropy for the dimer model and accurate critical points and exponents in Ising, Potts, and clock models [2203.17098].

In “variational boundary based tensor network renormalization group,” the truncation isometries themselves are optimized against a variationally computed global boundary environment. The common objective is
\[
\max_w \mathrm{tTr}(\Omega_w\, w w^\dagger),
\]
with update
\[
\Gamma_w=\Omega_w w^\dagger,\qquad
\Gamma_w=usv^\dagger,\qquad
w'=u^\dagger v.
\]
By using a VUMPS boundary MPS in canonical form, the method retains TRG-like \(\mathcal O(\chi^5)\) scaling while improving free-energy accuracy over HOTRG, HOSRG, CTM-TRG, and BWTRG at the same bond dimension. The authors are equally clear that, because short-range loop entanglement is not removed, scaling dimensions and central charges remain unstable in the same sense as in TRG/HOTRG [2508.10418]. This sharply separates variational environment optimization from full entanglement-filtering RG.

## 4. Neural and information-theoretic reformulations

A more recent line of work recasts VRG as explicit probabilistic modeling. In “Neural Network Renormalization Group,” the variational object is a normalized density \(q(\boldsymbol{x})\) induced by an invertible flow \( \boldsymbol{x}=g(\boldsymbol{z}) \) from a Gaussian prior \(p(\boldsymbol{z})\). Exact change of variables gives
\[
\ln q(\boldsymbol{x})
=
\ln p(\boldsymbol{z})
-
\ln\left|
\det\left(\frac{\partial \boldsymbol{x}}{\partial \boldsymbol{z}}\right)
\right|,
\]
and the training functional is
\[
\mathcal L
=
\int d\boldsymbol{x}\,
q(\boldsymbol{x})
\bigl[
\ln q(\boldsymbol{x})-\ln \pi(\boldsymbol{x})
\bigr].
\]
Because
\[
\mathcal L+\ln Z
=
\mathrm{KL}\!\left(
q(\boldsymbol{x})
\,\middle\|\,
\frac{\pi(\boldsymbol{x})}{Z}
\right)\ge 0,
\]
\(\mathcal L\) is a variational upper bound on the physical free energy. The architecture uses disentanglers and decimators, interprets inference \(g^{-1}\) as RG flow and generation \(g\) as inverse RG flow, and gives direct access to an exact latent-space energy
\[
U(\boldsymbol{z})
=
-\ln p(\boldsymbol{z})
-\ln \pi(g(\boldsymbol{z}))
+\ln q(g(\boldsymbol{z})).
\]
The construction is multiscale and RG-inspired, but not a Wilsonian semigroup: the maps are invertible, and the learned coarse variables need not coincide with conventional block spins [1802.02840].

An adjacent but distinct information-theoretic generalization is “hierarchical maximum entropy via the renormalization group.” Here one fixes a tower of deterministic coarse-graining maps
\[
X^{(1)}:=X,\qquad X^{(i+1)}:=T_i(X^{(i)}),
\]
and optimizes a weighted multilevel entropy
\[
H_{(\boldsymbol{\sigma},\mathbf T)}(X)
=
\sum_{i=1}^d \sigma_i H(X^{(i)}).
\]
The resulting generalized Gibbs principle is
\[
H_{(\boldsymbol{\sigma},\mathbf T)}(X)-\lambda \mathbb E[L(X)]
=
\log Z(\lambda)
-
D_{(\boldsymbol{\sigma},\mathbf T)}\!\left(P_X\middle\|\tilde P_X^{[\lambda]}\right).
\]
In invariant families, this yields explicit “parameter flows” analogous to RG flows. For nearest-neighbor Ising loss with decimation, the effective coupling obeys
\[
\theta_i=\frac{\bar\sigma_{i-1}}{2\bar\sigma_i}\log\cosh(2\theta_{i-1}).
\]
This is not canonical Kadanoff VRG, since the coarse-graining maps are fixed and deterministic rather than variationally learned, but it shows how multiscale variational principles can generate exact RG-like recursions in probability space [2509.01424].

These reformulations suggest a broader interpretation of VRG: the “variational” content can reside in free-energy bounds, exact likelihood models, multilevel KL objectives, or Pareto-optimal entropy tradeoffs, not only in hidden-spin kernels. That suggestion is interpretive, but it is strongly supported by the coexistence of these constructions.

## 5. Open-system dynamics and constructive field-theoretic variants

VRG also appears in settings far from equilibrium and far from finite-dimensional spin blocking. In dissipative spin-cavity systems, the density matrix obeys
\[
\frac{d\rho}{dt}
=
\mathcal L[\rho]
=
-i[\mathcal H,\rho]
+
\kappa\,\mathcal L_{\hat a_c}[\rho]
+
\sum_k \gamma_k\,\mathcal L_{\sigma_k^-}[\rho].
\]
The method vectorizes \(\rho\) into a superket \(|\rho\rangle\), performs Schmidt truncation in Liouville space,
\[
|\rho\rangle
=
\sum_{\tilde k=1}^{\mathcal K}
\alpha_{\tilde k}
|\tilde k_A\rangle |\tilde k_B\rangle
\;\;\longrightarrow\;\;
|\tilde\rho\rangle
=
\sum_{\tilde k=1}^{\mathcal D}
\alpha_{\tilde k}
|\tilde k_A\rangle |\tilde k_B\rangle,
\]
and updates the compressed basis adaptively during Suzuki–Trotter sweeps. The local evolution factors are
\[
\mathcal V_k(\Delta t)=e^{\tilde{\mathcal L}_k\Delta t},
\]
and the truncation is chosen from the dominant eigenvectors of reduced superoperators. This time-adaptive VRG was used to simulate mesoscopic ensembles up to \(N=105\) spins and to show periodic pulses of nonclassical light with \(g_2(t)<1\) at the emission peaks [1806.02394]. Here “variational RG” refers to optimal low-rank compression of an evolving open-system superoperator state, not to equilibrium coarse-graining of a partition function.

A very different rigorous usage appears in Balaban-style constructive RG for lattice non-linear sigma models. The blocked transform is
\[
(T\rho)(V)
=
\int dU\;
\delta(\bar U V^{-1})\,
e^{-g^{-2}A(U)-E},
\]
so for fixed coarse field \(V\) one must solve the constrained minimization of the action \(A(U)\) under the block-average constraint. In the sigma-model adaptation, the result is a unique minimal configuration in the small-field regime, and that minimizer becomes the classical background field for steepest-descent analysis of the fluctuation integral [2204.08252]. The minimizer is written in multiplicative form,
\[
U_{\min}=h_{\min}U_0,
\qquad
h_{\min}
=
\exp\Bigl(
i\bigl[
A_1+H_1B-H_1D(A_1+H_1B)
\bigr]
\Bigr),
\]
where \(A_1\) solves a fixed-point equation derived from the constrained Euler–Lagrange problem. This is a precise variational RG in the sense of a constrained classical minimization inside one RG step; it is not an approximate numerical truncation scheme.

These examples broaden the scope of VRG in two orthogonal directions. The dissipative spin-cavity method shows that variational compression can be applied directly to Lindbladian time evolution in superoperator space [1806.02394]. The constructive sigma-model work shows that the “variational problem” can mean the exact constrained saddle configuration required to define a background field for rigorous multiscale integration [2204.08252].

## 6. Terminological breadth, misconceptions, and adjacent nonstandard uses

The term “variational RG” is not uniform across the literature, and several nearby constructions are explicitly not standard VRG. One recurrent source of confusion is exact or continuous RG written in variational form. In the optimal-transport reformulation of exact RG, Polchinski/Wegner–Morris flow becomes
\[
-\Lambda \frac{d}{d\Lambda}P_\Lambda[\phi]
=
-
\nabla_{\mathcal W_2}
S(P_\Lambda[\phi]\|Q_\Lambda[\phi]),
\]
and admits a JKO-type minimizing-movement discretization [2202.11737]. This is genuinely variational, but not Kadanoff-style VRG with hidden block variables.

Likewise, “Gradient flow and the renormalization group” studies a self-consistent gradient-flow evolution in which the flowed field is driven by the effective action \(S_\tau\) rather than the bare action. The central equation is
\[
\partial_\tau S_\tau[\phi]
=
\int_{x,y}K_\tau(x-y)
\left[
\frac{\delta S_\tau}{\delta\phi(x)}
\frac{\delta S_\tau}{\delta\phi(y)}
-
\frac{\delta^2 S_\tau}{\delta\phi(x)\delta\phi(y)}
\right].
\]
After a field redefinition restoring canonical normalization, the resulting local-potential approximation reproduces Gaussian and Wilson–Fisher eigenvalues at \(O(\varepsilon)\). However, the same source is explicit that this is not variational RG in the usual block-spin sense and contains no explicit variational principle or hidden-variable construction [1805.12094].

A further terminological divergence occurs in finite-temperature field theory. In “Scale dependence improvement of the quartic scalar field thermal effective potential in the optimized perturbation theory,” the authors call
\[
\text{OPT resummation}+\text{RG improvement}+\text{PMS fixing of }\eta
\]
their “Variational Renormalization Group” framework, applying it to thermal \(\lambda\phi^4\) theory and emphasizing improved renormalization-scale stability of the effective potential, pressure, and critical temperature [2509.07250]. This use of “VRG” is explicitly unrelated to the tensor-network or Kadanoff/block-spin traditions.

Several distinctions are therefore essential.

First, **free-energy preservation is not the same as full-distribution matching**. In Kadanoff-style VRG, the exact trace condition preserves the partition function/free energy, whereas in RBM language the natural objective is \(D_{KL}(P\|p_\lambda)\); they coincide only in the exact case [1410.3831].

Second, **variational truncation is not the same as full disentangling RG**. Boundary-optimized tensor methods such as VBTRG improve the projector by using a global environment, but without entanglement filtering they do not stabilize conformal data at criticality [2508.10418].

Third, **not every “RG–deep learning” analogy is a claim about all neural networks**. The exact mapping established in the RBM setting is narrower and more formal than blanket statements about deep learning as RG [1410.3831].

Taken together, these distinctions show that VRG is best treated as a technically plural notion. In one branch it is a hidden-variable coarse-graining theory with trace conditions; in another it is a convex Monte Carlo inverse problem for renormalized Hamiltonians; in another it is variational tensor optimization over low-entanglement manifolds or boundary environments; and in adjacent fields it can denote exact-RG variational flows or RG-improved variational resummation. The unifying idea is the replacement of a fixed RG step by an optimization problem, but the optimized object and the meaning of “renormalization” vary substantially across contexts.

Source: https://www.emergentmind.com/topics/variational-renormalization-group-vrg