---
title: Variance Profile Matrix in Random Matrix Theory
url: https://www.emergentmind.com/topics/variance-profile-matrix
type: topic
---

# Variance Profile Matrix in Random Matrix Theory

Searching arXiv for recent and foundational papers on variance-profile random matrices.
A variance profile matrix is a deterministic object that specifies entrywise second moments of a random matrix. In different conventions it appears either as the matrix of variances itself, such as $\Sigma_{ij}^N := \operatorname{Var}(X_{ij}^N)(1+\mathbf{1}_{i=j})^{-1}$ in Wigner-type models, or as a standard deviation profile $A_n=(\sigma_{ij})$ whose Hadamard square $A_n\odot A_n$ is the variance profile [2403.05413, 1612.04428, 1901.09404]. The notion generalizes homogeneous-variance ensembles by allowing heterogeneous variances across entries, and the resulting profile governs the limiting spectral distribution, spectral edge, operator norm, pseudospectral support, fluctuation theory, large deviations, and several high-dimensional inference procedures [2002.01010, 2404.13795, 2404.17573, 2403.20200, 2504.03035, 2606.04531].

## 1. Definitions and conventions

The common feature across the literature is entrywise heteroscedasticity. In the Hermitian Wigner-type setting of large deviations, one considers real symmetric matrices
\[
H_{ij}^N=\frac{1}{\sqrt N}X_{ij}^N,
\]
with independent centered entries for $1\le i\le j\le N$, and defines the variance profile by
\[
\Sigma_{ij}^N=\operatorname{Var}(X_{ij}^N)(1+\mathbf{1}_{i=j})^{-1}.
\]
For off-diagonal entries, $\operatorname{Var}(X_{ij}^N)=\Sigma_{ij}^N$, while for diagonal entries, $\operatorname{Var}(X_{ii}^N)=2\Sigma_{ii}^N$ [2403.05413].

In non-Hermitian models, a standard convention is the rescaled entrywise product
\[
Y_n=\frac{1}{\sqrt n}(A_n\odot X_n),
\]
where $A_n=(\sigma_{ij})$ is deterministic with nonnegative entries and $X_n$ has i.i.d. centered unit-variance entries. The associated normalized variance profile is
\[
V_n=\frac{1}{n}(A_n\odot A_n),
\]
so that $\operatorname{var}(Y_{ij})=(V_n)_{ij}$ [1612.04428, 2007.15438].

In statistical models, the same idea appears as a deterministic modulation of an i.i.d. design. For ridge regression with non-identically distributed predictors,
\[
X_n=\Upsilon_n\circ Z_n,\qquad \Gamma_n=(\gamma_{ij}^2),
\]
and $\Gamma_n$ is called the variance profile because $\operatorname{Var}(X_{ij})=\gamma_{ij}^2$ [2403.20200]. An analogous formulation is used for random features, where the variance profile may enter both the data matrix and, in a more general Gaussian-equivalence model, the random-feature matrix [2504.03035].

| Setting | Profile object | Meaning |
|---|---|---|
| Hermitian Wigner-type | $\Sigma^N$ or $\sigma$ | Position-dependent entry variances |
| Non-Hermitian entrywise product | $A_n\odot A_n$ or $V_n$ | Standard deviation profile squared |
| High-dimensional regression | $\Gamma_n=(\gamma_{ij}^2)$ | Entrywise predictor variances |

A central structural assumption is convergence of the discrete profile to a continuum kernel. In the Hermitian setting, one assumes that there exist points $t_1^N<\dots<t_N^N$ in $[0,1]$ and a symmetric function $\sigma:[0,1]^2\to\mathbb{R}_+$ such that
\[
\lim_{N\to\infty}\sup_{1\le i,j\le N}\bigl|\Sigma_{ij}^N-\sigma(t_i^N,t_j^N)\bigr|=0
\]
and that the empirical distribution of the $t_i^N$ converges weakly to the uniform law on $[0,1]$ [2403.05413]. In sampled non-Hermitian profiles one similarly writes $\sigma_{ij}^2=\sigma^2(i/n,j/n)$ or $s_{ij}^{(n)}=n^{-1}f(i/n,j/n)$ [2007.15438, 2501.03657].

A recurrent terminological distinction is that some papers call $A_n$ the “standard deviation profile” and $A_n\odot A_n$ the “variance profile,” while others directly call the matrix of second moments the variance profile [1901.09404, 1612.04428]. This is a difference of convention rather than substance.

## 2. Hermitian models, Dyson equations, and the spectral edge

For Hermitian variance-profile ensembles, the limiting spectral measure is described by self-consistent resolvent equations. In the discrete formulation of the Wigner-type model, with
\[
(S^N u)_i=\frac{1}{N}\sum_{j=1}^N \Sigma_{ij}^N u_j,
\]
the vector Dyson equation reads
\[
z-(S^N m^N(z))_i=\frac{1}{m_i^N(z)},\qquad i=1,\dots,N.
\]
In the continuum limit one seeks measurable $m(t;z)$ with
\[
-\frac{1}{m(t;z)}=z+\int_0^1 \sigma(t,s)\,m(s;z)\,ds,
\]
and the limiting spectral measure is $\mu_\sigma=\int_0^1\mu_{\sigma,t}\,dt$ [2403.05413].

An equivalent formulation in the earlier variance-profile LDP literature is the quadratic vector equation
\[
\frac{1}{\mathfrak m(x,z)}=z-\int_0^1 \sigma^2(y,x)\,\mathfrak m(y,z)\,dy,
\qquad z\in\mathbb{H}^+,
\]
with $m(z)=\int_0^1 \mathfrak m(x,z)\,dx$ the Stieltjes transform of the limiting measure $\mu_\sigma$ [2002.01010]. In both formulations, the rightmost point of the support of the limiting measure is denoted $r_\sigma$, and $\lambda_{\max}(H^N)\to r_\sigma$ in probability or almost surely under the stated assumptions [2403.05413, 2002.01010].

For piecewise constant profiles, the continuum equation reduces to a finite-dimensional system. If $[0,1]$ is partitioned into blocks $I_1,\dots,I_p$ and $\sigma$ is constant on each $I_k\times I_\ell$, then one obtains component measures $\mu_{\sigma,k}$, and the limiting measure is their mixture [2403.05413]. This finite-dimensional reduction is the basis for explicit edge characterizations and for block-structured examples.

The spectral edge is also the asymptotic operator norm. For symmetric random matrices with variance profile $s_{ij}^{(N)}=\mathbb{E}|a_{ij}^{(N)}|^2$, the operator norm of $A_N/\sqrt N$ converges to the largest element of the support of the limiting empirical spectral distribution. The paper on operator norms states that finite fourth moments are sufficient for convergence in probability and finite $4+\epsilon$ moments are sufficient for almost sure convergence [2404.13795]. This places the spectral-edge description of variance-profile ensembles on the same footing as the Bai–Yin theorem in the homogeneous case.

A common simplification is the constant profile. When $\sigma(x,y)\equiv 1$, the model reduces to classical Wigner matrices, the Dyson equation gives the semicircle law, and the spectral edge is $2$ after standardization [2002.01010]. The broader role of the variance profile is therefore not merely to rescale the spectrum, but to change the self-consistent equation that determines the entire limiting measure.

## 3. Non-Hermitian deterministic equivalents and support geometry

For non-Hermitian matrices with independent entries and a variance profile, the central objects are deterministic equivalents defined through Master Equations. In the model
\[
Y_n=\frac{1}{\sqrt n}(A_n\odot X_n),
\]
the empirical spectral distribution $\mu_n^Y$ is approximated by a deterministic, rotationally invariant probability measure $\mu_n$ [1612.04428]. Its radial cumulative distribution function is obtained from the zero-regularization limit of the Regularized Master Equations:
\[
r_i=\frac{(V_n^T r)_i+t}{s^2+((V_n\tilde r)_i+t)((V_n^T r)_i+t)},\qquad
\tilde r_i=\frac{(V_n\tilde r)_i+t}{s^2+((V_n\tilde r)_i+t)((V_n^T r)_i+t)},
\]
followed by the limit $t\downarrow 0$ [1612.04428, 2007.15438].

The support radius is controlled by the spectral radius of the normalized variance profile. Specifically, the support of $\mu_n$ is contained in the disk
\[
\{z\in\mathbb{C}: |z|\le \sqrt{\rho(V_n)}\},
\]
and, except maybe at zero, $\mu_n$ admits a positive density on the centered disk of radius $\sqrt{\rho(V_n)}$ [2007.15438]. In the constant-profile case, this yields the circular law; in the doubly stochastic case, the circular law is recovered as well [1612.04428]. The paper on properties and examples further identifies the profiles that yield the circular law: up to diagonal conjugation, these are the doubly stochastic normalized variance profiles [2007.15438].

The behavior at zero is profile-dependent. The density may be bounded, may blow up, or may vanish while an atom appears [2007.15438]. Separable profiles reduce the Master Equations to a scalar fixed-point equation, and sampled continuous profiles lead to an integral version of the same structure [2007.15438]. These examples show that the variance profile shapes not only the outer boundary but also fine structure in the bulk.

A further extension combines a variance profile with an additive deterministic diagonal deformation. For matrices $A_n+X_n$ with
\[
\mathbb{E}|x_{ij}^{(n)}|^2=\frac{1}{n}s\!\left(\frac{i}{n},\frac{j}{n}\right),
\]
the limiting spectral measure is identified as the Brown measure of a deformed operator-valued circular element, and its support exactly coincides with the $\varepsilon$-pseudospectrum in the consecutive limits $n\to\infty$ and $\varepsilon\to0$ [2404.17573]. In that framework the variance profile enters through integral operators
\[
(Su)(x)=\int_0^1 s(x,y)\,u(y)\,dy,\qquad
(S^*u)(x)=\int_0^1 s(y,x)\,u(y)\,dy,
\]
and through a matrix Dyson equation for the hermitized resolvent [2404.17573]. This identifies the support as a singular-value instability boundary determined by the profile.

## 4. Extremes, fluctuations, and large deviations

Variance profiles affect both global fluctuations and extreme-value probabilities. For the largest eigenvalue of Hermitian Wigner-type matrices, a large deviation principle holds at speed $N$ under convergence of $\Sigma^N$ to $\sigma$ and sharp sub-Gaussian tails. The rate function is expressed in terms of the solution of a Dyson equation involving $\sigma$, and the result is new even in the Gaussian case with non-constant variance profiles [2403.05413]. An earlier formulation gives
\[
I^{(\beta)}(\sigma,x)=
\begin{cases}
+\infty,& x<r_\sigma,\\
\sup_{\theta\ge 0}\{J(\mu_\sigma,\theta,x)-F(\sigma,\theta)\},& x\ge r_\sigma,
\end{cases}
\]
with $F(\sigma,\theta)$ an annealed spherical-integral variational functional over probability measures on $[0,1]$ [2002.01010]. The later work removes several restrictive assumptions by using a more flexible rate function and a refined proof [2403.05413].

For non-Hermitian matrices, an upper bound on the spectral radius is controlled directly by the variance profile matrix. Under minimal moment assumptions and sparse profiles, the spectral radius does not exceed the square root of the spectral radius of the variance profile matrix with large probability:
\[
\rho(X^{(n)})\le \sqrt{\rho(S^{(n)})}+o_{\mathbb P}(1).
\]
The proof uses the reverse characteristic polynomial and a random analytic function built from traces of powers of $S^{(n)}$ [2501.03657].

At the level of linear eigenvalue statistics, the variance profile enters through explicit variance bounds and Fourier or cycle-sum functionals. For Hermitian matrices $Y=A\circ X$ with variance profile $A\circ A$, the paper on second-order Poincaré inequalities proves an upper bound on the total variation distance between the standardized linear eigenvalue statistic and the standard Gaussian random variable, and uses it to establish CLTs for several classes of variance profiles [1901.09404]. For non-Hermitian random band matrices with variance profile
\[
\operatorname{Var}(m_{ij})=\frac{1}{c_n}w_\nu\!\left(\frac{i-j}{c_n}\right),
\]
the fluctuations of linear eigenvalue statistics converge to a Gaussian law with an explicit variance formula depending on the Fourier transform of the profile [1904.11098].

These results make clear that a variance profile is not a peripheral perturbation. It modifies the limiting edge, the operator norm, the support of the limiting density, and the asymptotic cost of rare spectral events.

## 5. Statistical, algorithmic, and applied roles

In high-dimensional regression, the variance profile matrix encodes non-identically distributed predictors and determines deterministic equivalents for predictive risk and degrees of freedom. For the model
\[
X_n=\Upsilon_n\circ Z_n,\qquad \Gamma_n=(\gamma_{ij}^2),
\]
the ridge estimator admits deterministic equivalents expressed through diagonal resolvent limits $T_p(z)$ and $\widetilde T_n(z)$ satisfying a variance-profile Dyson equation [2403.20200]. In the quasi doubly stochastic case these equivalents collapse to the classical Marchenko–Pastur fixed point and reproduce the standard double-descent profile, while other profiles produce different shapes, including “triple descent” in simulations [2403.20200].

The same phenomenon persists in random features. For non-iid feature vectors with variance profile, operator-valued free probability and a block linearization yield deterministic equivalents for training and prediction risks associated with ridge regression in the random-features model [2504.03035]. Under a row-stochastic assumption on the variance profile, explicit Marchenko–Pastur formulas appear in the chaos-only case; more general cases require solving an operator-valued fixed-point equation whose self-energy operator is built from the profile [2504.03035].

The variance profile also modifies iterative inference algorithms. For diagonal expectation propagation under variance-profile Gaussian sensing matrices,
\[
A_{ij}=\sqrt{\frac{s_{ij}}{M}}\,Z_{ij},
\]
the effective observation seen by the nonlinear module is generally not a fresh scalar Gaussian channel. Instead, the residuals form a coordinate-dependent Gaussian process whose covariance is shaped by the variance profile and by the finite linear history of the algorithm [2606.04531]. The paper characterizes this process through a conditioned matrix-Dyson-equation deterministic equivalent and a Gaussian-regression decomposition that separates predictable memory from orthogonal innovation [2606.04531].

In spin-glass theory, the variance profile matrix appears as a deterministic, symmetric matrix $S^{(n)}$ with nonnegative entries controlling the covariance of Gaussian couplings:
\[
\operatorname{Var}(W_{ij}^{(n)})=t\,s_{ij}^{(n)}.
\]
At sufficiently high temperature, the free energy is approximated by a deterministic functional of the fixed point
\[
q^{(n)}=t\,S^{(n)}g(q^{(n)}),
\]
and the TAP/AMP Onsager correction becomes profile-dependent through $tS^{(n)}(1_n-g(q))$ [2604.25535].

Applied settings use the same object in lower-dimensional or structured models. In low-rank denoising with heteroscedastic noise, a GUE or rectangular Gaussian matrix with variance profile leads to operator-valued resolvent equations and profile-aware determinant equations for outliers [1907.07753]. In a $2\times2$ complex central Gaussian channel with arbitrary variances $\phi_{ij}$, the variance profile matrix $V=[\phi_{ij}]$ fundamentally departs from classical Wishart models and yields exact formulas for the distribution of the Gram matrix and of its eigenvalues [1705.05214].

## 6. Structured profiles, computation, and recurring distinctions

Several structured classes recur across the literature. Constant profiles recover semicircle or circular laws [2002.01010, 1612.04428]. Block-constant profiles reduce continuum fixed-point equations to finite-dimensional systems and are used for community structure, deformed models, and explicit edge characterizations [2403.05413, 1907.07753]. Band or spatially decaying profiles induce stronger coupling for nearby indices and modify the limiting spectral distribution away from the homogeneous case [2403.05413, 1904.11098]. Separable profiles
\[
\sigma_{ij}^2=d_i\tilde d_j
\]
collapse the non-Hermitian Master Equations to a scalar equation and produce explicit radial densities, including Girko’s Sombrero distribution [2007.15438]. Sampled continuous profiles discretize kernels $f$ or $\sigma^2$ on $[0,1]^2$ and connect finite matrices to compact integral operators [2501.03657, 2007.15438].

A standard computational route is profile discretization followed by solution of a self-consistent system. For the Hermitian LDP, one discretizes $[0,1]$, approximates $\sigma$ by a matrix, solves the discrete Dyson equation, and then evaluates the functionals entering the rate function [2403.05413]. For non-Hermitian deterministic equivalents, one solves the Regularized Master Equations for decreasing regularization, obtains the radial CDF
\[
F_n(s)=1-\frac{1}{n}\langle q(s),V_n\tilde q(s)\rangle,
\]
and differentiates numerically to recover the density [1612.04428]. For high-dimensional regression with variance profile, one solves the vector Dyson equation for $T_p(-\lambda)$ and uses it in the explicit risk formula [2403.20200].

Two distinctions recur and often remove apparent contradictions. First, a variance profile matrix is not the same object as a single population covariance matrix. In the regression papers, the profile is entrywise and may vary with both row and column; it is therefore more general than a fixed covariance $\Sigma$ shared across observations [2403.20200, 2504.03035]. Second, zero entries are compatible with much of the non-Hermitian theory, provided one imposes quantitative irreducibility or related admissibility assumptions [1612.04428, 2007.15438]. By contrast, some earlier Hermitian large-deviation arguments required continuity of a variational maximizer or positivity assumptions in the lower bound, and this dependence on auxiliary assumptions is explicitly discussed in the later generalization [2002.01010, 2403.05413].

A plausible implication is that “variance profile matrix” functions less as a single definition than as a unifying schema for inhomogeneous random matrix models. Across Hermitian, non-Hermitian, rectangular, algorithmic, and statistical settings, the profile is the deterministic object that carries heterogeneity into the limiting equations, and those equations in turn determine spectral support, edges, fluctuations, rare events, and inference performance [2403.05413, 2007.15438, 2403.20200, 2606.04531].

Source: https://www.emergentmind.com/topics/variance-profile-matrix