Papers
Topics
Authors
Recent
Search
2000 character limit reached

Hanson–Wright Inequality

Updated 9 July 2026
  • Hanson–Wright inequality is a concentration result for quadratic forms of independent, centered sub-gaussian vectors, quantified via Hilbert–Schmidt and operator norms.
  • It distinguishes between two deviation regimes: a quadratic term for moderate deviations and a linear term for large deviations, akin to Bernstein's inequality.
  • The inequality underpins applications in random matrix theory, covariance estimation, and norm concentration, with recent extensions addressing dependence, heavy tails, and infinite-dimensional settings.

The Hanson–Wright inequality is a concentration inequality for quadratic forms of the form XAXX^\top A X when XX has independent centered sub-gaussian coordinates. In its standard modern form, if X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n has independent components with EXi=0\mathbb E X_i=0 and Xiψ2<K\|X_i\|_{\psi_2}<K, then for every t0t\ge 0,

P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]

for an absolute constant c>0c>0 (Rudelson et al., 2013). The inequality isolates two intrinsic matrix scales—the Hilbert–Schmidt/Frobenius norm and the operator norm—and has become a standard tool for quadratic chaos, norm concentration, covariance estimation, random matrix theory, and a wide range of generalizations.

1. Classical formulation

A standard formulation assumes that X1,,XnX_1,\dots,X_n are independent, mean-zero, and uniformly sub-gaussian. In the notation used in the modern expository treatment, the sub-gaussian norm may be written as

Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},

and the matrix norms in the inequality are

XX0

The estimate has a two-level form: a quadratic exponent controlled by XX1 for moderate deviations, and a linear exponent controlled by XX2 for larger deviations (Rudelson et al., 2013).

Because

XX3

one may always reduce to the symmetric case. Under independence and centering,

XX4

since XX5 for XX6. The centered quadratic form therefore splits naturally into diagonal and off-diagonal pieces: XX7 This decomposition is the structural starting point of essentially all modern proofs (Rudelson et al., 2013).

The inequality is often interpreted as the quadratic-form counterpart of Bernstein concentration. The Hilbert–Schmidt term captures the aggregate contribution of many coefficients, whereas the operator-norm term captures the effect of the largest direction of XX8. This suggests why the same pair of norms continues to reappear in many later generalizations.

2. Proof structure and Gaussian mechanism

In the standard proof, the diagonal term is treated by the observation that if XX9 is sub-gaussian, then X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n0 is sub-exponential. A Bernstein-type inequality then yields a Hanson–Wright-form tail for

X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n1

The off-diagonal term

X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n2

is the main difficulty (Rudelson et al., 2013).

The classical route decouples the off-diagonal sum using auxiliary Bernoulli selectors, conditions on one block of variables, and turns the remaining conditional chaos into a sub-gaussian linear form. A Gaussian comparison step then reduces the argument to a quadratic form in standard Gaussians. After spectral decomposition of the relevant matrix, the proof uses the exact identity

X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n3

for X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n4, and finally optimizes a Chernoff bound over the admissible range of the exponential parameter (Ziemann, 31 Aug 2025).

A later elementary proof keeps the same diagonal/off-diagonal architecture but avoids Bourgain’s convex decoupling argument. Its stated ingredients are decomposition into diagonal and off-diagonal parts, elementary mgf bounds for sub-gaussian and sub-exponential variables, a Gaussian randomization or rotation argument, and spectral decomposition of the matrix (Ziemann, 31 Aug 2025). In that presentation, the appearance of X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n5 and X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n6 is completely transparent: after diagonalization, they are simply

X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n7

with X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n8 the eigenvalues of the symmetric part of X=(X1,,Xn)RnX=(X_1,\dots,X_n)\in\mathbb R^n9.

This proof pattern also explains why the Hanson–Wright inequality is especially natural in the Gaussian case. The quadratic chaos becomes a weighted sum of centered EXi=0\mathbb E X_i=00-variables, and the concentration problem reduces to controlling the logarithm of an explicit Gaussian mgf.

3. Norm concentration and Gaussian refinements

A closely related consequence is norm concentration. If EXi=0\mathbb E X_i=01, then

EXi=0\mathbb E X_i=02

and Hanson–Wright yields

EXi=0\mathbb E X_i=03

This gives sub-gaussian concentration of EXi=0\mathbb E X_i=04 around EXi=0\mathbb E X_i=05 at scale EXi=0\mathbb E X_i=06 (Rudelson et al., 2013).

For Gaussian vectors there is a sharper benchmark in terms of covariance geometry. If EXi=0\mathbb E X_i=07 is Gaussian with covariance EXi=0\mathbb E X_i=08 and EXi=0\mathbb E X_i=09, then the Laurent–Massart-type bound reads

Xiψ2<K\|X_i\|_{\psi_2}<K0

with a matching lower-tail bound

Xiψ2<K\|X_i\|_{\psi_2}<K1

This Gaussian expression serves as an exact benchmark for several later extensions (Spokoiny, 2023).

Quantitative refinement of the Gaussian Hanson–Wright constant has been a separate line of work. One note revisiting the Gaussian proof showed that in the real symmetric case the best constant Xiψ2<K\|X_i\|_{\psi_2}<K2 in

Xiψ2<K\|X_i\|_{\psi_2}<K3

satisfies

Xiψ2<K\|X_i\|_{\psi_2}<K4

(Moshksar, 2021). A later refinement of Gaussian quadratic-chaos concentration reported constants at least Xiψ2<K\|X_i\|_{\psi_2}<K5 in the symmetric case and at least Xiψ2<K\|X_i\|_{\psi_2}<K6 in the positive-semidefinite case, obtained by sharpening both Hanson–Wright and Laurent–Massart analyses (Moshksar, 2024). The same paper also derived a sequence of bounds indexed by Xiψ2<K\|X_i\|_{\psi_2}<K7 involving Schatten norms, with a phase transition: Xiψ2<K\|X_i\|_{\psi_2}<K8 is tightest for smaller deviations, while larger Xiψ2<K\|X_i\|_{\psi_2}<K9 becomes tighter for sufficiently large deviations (Moshksar, 2024).

4. Dependence, local mgf control, and heavy tails

A major extension replaces coordinatewise independence by concentration of the whole vector. If t0t\ge 00 is a mean-zero random vector in t0t\ge 01 with the convex concentration property with constant t0t\ge 02, then for every matrix t0t\ge 03 and every t0t\ge 04,

t0t\ge 05

and in the isotropic case this becomes

t0t\ge 06

No coordinatewise independence is assumed; examples mentioned in this framework include vectors under uniform mixing, Dobrushin-type dependence criteria, sampling without replacement, and bounded product measures (Adamczak, 2014). The same source also stresses a common misconception: it is not true that if t0t\ge 07 has i.i.d. subgaussian coordinates, then t0t\ge 08 automatically has the convex concentration property with dimension-free constant (Adamczak, 2014).

A different extension weakens global sub-gaussianity to a bounded-domain exponential moment condition. Let t0t\ge 09 be centered, assume P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]0, and suppose

P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]1

For a linear map P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]2, with P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]3, there is an upper quantile function P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]4 such that

P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]5

The distinctive feature is a phase transition: for P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]6,

P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]7

which exactly matches the Gaussian benchmark, whereas for P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]8, P{XAXEXAX>t}2exp ⁣[cmin ⁣(t2K4AHS2,tK2A)]\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\} \le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]9 grows linearly in c>0c>00 (Spokoiny, 2023). This is not presented as a full theorem for arbitrary indefinite quadratic forms; it is tailored to norms of linear images c>0c>01, equivalently to the positive-semidefinite quadratic form c>0c>02 (Spokoiny, 2023).

Heavy-tail generalization replaces sub-gaussian coordinates by c>0c>03-variables. If c>0c>04 are independent, centered, and satisfy c>0c>05 for c>0c>06, then for symmetric c>0c>07,

c>0c>08

The small-deviation term remains Gaussian-like, but the large-deviation exponent degrades from c>0c>09 to X1,,XnX_1,\dots,X_n0, reflecting the heavier tail of X1,,XnX_1,\dots,X_n1 (Sambale, 2020).

5. Uniform and supremum versions

A natural strengthening is to control not one quadratic form but an entire class. For a bounded family X1,,XnX_1,\dots,X_n2 of matrices, one studies

X1,,XnX_1,\dots,X_n3

Under the full concentration property for all 1-Lipschitz functions with constant X1,,XnX_1,\dots,X_n4, one has

X1,,XnX_1,\dots,X_n5

where

X1,,XnX_1,\dots,X_n6

This is a two-sided concentration inequality around the mean for suprema of quadratic forms (Adamczak, 2014).

For independent centered subgaussian coordinates, the entropy-method approach yields a uniform Hanson–Wright inequality with a different complexity parameter. Let

X1,,XnX_1,\dots,X_n7

Then for finite symmetric X1,,XnX_1,\dots,X_n8 and

X1,,XnX_1,\dots,X_n9

Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},0

This framework recovers Talagrand’s right-tail bound for Rademacher chaos and extends earlier uniform results to unbounded subgaussian coordinates by truncating both gradients and coordinates (Klochkov et al., 2018). The same work emphasizes that in general one cannot replace Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},1 by Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},2 in the uniform theorem (Klochkov et al., 2018).

A uniform Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},3-extension is also available. For

Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},4

one has

Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},5

The same source notes that this uniform theorem is weaker than the one-matrix theorem: it controls only the upper tail, and the small-deviation term is Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},6-subexponential rather than Gaussian unless Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},7 (Sambale, 2020).

6. Infinite-dimensional and tensor extensions

One infinite-dimensional extension replaces scalar coordinates by Hilbert-space-valued random variables. Let Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},8 be a real separable Hilbert space, let Xψ2=supp1p1/2(EXp)1/p,\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},9 be independent centered XX00 random variables, and consider

XX01

Under an additional Bernstein-type condition on XX02, the two-sided Hilbert-space Hanson–Wright inequality becomes

XX03

Here the finite-dimensional scalar parameter XX04 is replaced by the operator-valued covariance proxy XX05, and the theorem is dimension-free in the ambient Hilbert space (Chen et al., 2018).

A Banach-space version treats coefficients XX06 taking values in a normed space XX07, so that

XX08

is XX09-valued. The resulting Gaussian theory identifies Banach-valued analogues of the Hilbert–Schmidt and operator norms, notably

XX10

and

XX11

From the Gaussian bounds one deduces a subgaussian Hanson–Wright extension: XX12 for independent mean-zero XX13-subgaussian XX14, provided XX15 is above the stated Gaussian threshold (Adamczak et al., 2018). In general Banach spaces the conjectured optimal moment comparison is known only up to logarithmic factors, while in certain spaces, including XX16-spaces, those logarithmic factors can be eliminated (Adamczak et al., 2018).

Tensor extensions proceed in two different directions. For Kronecker products of independent subgaussian vectors,

XX17

the quadratic form

XX18

admits moment bounds of the form

XX19

where XX20 is expressed through explicit tensor partition norms of contractions XX21 of the coefficient array (Bamberger et al., 2021). The same framework yields improved concentration inequalities for

XX22

and is sharp up to XX23-dependent constants in the Gaussian case (Bamberger et al., 2021).

A different tensor generalization replaces matrices by Hermitian tensors under the Einstein product and studies the Ky Fan XX24-norm of a polynomial of a quadratic tensor sum. The proof again splits the quadratic object into diagonal and coupling parts, applies a generalized tensor Chernoff bound to the diagonal part, and uses a decoupling inequality before applying tensor Chernoff to the coupling part (Chang, 2022). This is structurally analogous to Hanson–Wright, but the scalar tail of a quadratic form is replaced by concentration of

XX25

7. Sparse variants and applications

Random sparsification changes the geometry of quadratic concentration. If XX26 denotes coordinatewise masking by independent Bernoulli variables XX27 with XX28, then a sparse Hanson–Wright inequality controls

XX29

with variance proxy

XX30

instead of XX31. One formulation states

XX32

(Zhou, 2015). This shows that diagonal and off-diagonal contributions scale differently under random coordinate survival.

A sparse bilinear extension considers

XX33

where XX34 are sub-gaussian vectors and XX35 are Bernoulli masks. The resulting inequality has the same two-level Hanson–Wright shape but with weighted Frobenius norm

XX36

in place of XX37: XX38 (Park et al., 2022).

Heavy-tailed sparse variants extend this picture further. For sparse XX39-sub-exponential vectors XX40, one simplified inequality reads

XX41

and for XX42 a sharper bound introduces additional weighted operator, row-XX43, and max-entry scales (Dai et al., 27 May 2025). A later sparse XX44-subexponential theory derives three-regime tails for off-diagonal forms,

XX45

and uses them to obtain local laws and complete eigenvector delocalization for sparse XX46-subexponential Hermitian random matrices, as well as concentration of XX47 for sparse XX48-subexponential vectors (He et al., 2024).

Applications of Hanson–Wright and its variants are broad. The classical inequality yields concentration of the distance from a random vector to a fixed subspace and bounds on norms of products of deterministic and random matrices (Rudelson et al., 2013). Uniform variants recover concentration for empirical covariance operators of Banach-valued Gaussian variables (Adamczak, 2014) and covariance estimation with missing observations (Klochkov et al., 2018). The local-mgf extension applies to Bernoulli vector sums and covariance estimation in Frobenius norm (Spokoiny, 2023). The Hilbert-space version underlies exponential recovery guarantees for generalized XX49-means clustering with non-Euclidean data (Chen et al., 2018). These examples suggest that the Hanson–Wright inequality is best viewed not as a single estimate for scalar quadratic forms, but as a concentration principle whose central matrix scales—variance-type aggregate size and worst-direction amplification—persist across dependence, sparsity, tensorization, and infinite-dimensional settings.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Hanson-Wright Inequality.