---
title: Hanson–Wright Inequality
url: https://www.emergentmind.com/topics/hanson-wright-inequality
type: topic
---

# Hanson–Wright Inequality

The Hanson–Wright inequality is a concentration inequality for quadratic forms of the form \(X^\top A X\) when \(X\) has independent centered sub-gaussian coordinates. In its standard modern form, if \(X=(X_1,\dots,X_n)\in\mathbb R^n\) has independent components with \(\mathbb E X_i=0\) and \(\|X_i\|_{\psi_2}<K\), then for every \(t\ge 0\),
\[
\mathbb{P}\Big\{\big|X^\top A X - \mathbb{E}X^\top A X\big|>t\Big\}
\le 2\exp\!\left[-c\min\!\left(\frac{t^2}{K^4\|A\|_{HS}^2},\, \frac{t}{K^2\|A\|}\right)\right]
\]
for an absolute constant \(c>0\) [1306.2872]. The inequality isolates two intrinsic matrix scales—the Hilbert–Schmidt/Frobenius norm and the operator norm—and has become a standard tool for quadratic chaos, norm concentration, covariance estimation, random matrix theory, and a wide range of generalizations.

## 1. Classical formulation

A standard formulation assumes that \(X_1,\dots,X_n\) are independent, mean-zero, and uniformly sub-gaussian. In the notation used in the modern expository treatment, the sub-gaussian norm may be written as
\[
\|X\|_{\psi_2} = \sup_{p\ge 1} p^{-1/2}(\mathbb E|X|^p)^{1/p},
\]
and the matrix norms in the inequality are
\[
\|A\|_{HS}=\Big(\sum_{i,j}|a_{ij}|^2\Big)^{1/2},\qquad
\|A\|=\max_{x\neq 0}\frac{\|Ax\|_2}{\|x\|_2}.
\]
The estimate has a two-level form: a quadratic exponent controlled by \(\|A\|_{HS}\) for moderate deviations, and a linear exponent controlled by \(\|A\|\) for larger deviations [1306.2872].

Because
\[
X^\top A X = X^\top\!\left(\frac{A+A^\top}{2}\right)X,
\]
one may always reduce to the symmetric case. Under independence and centering,
\[
\mathbb E[X^\top A X] = \sum_i a_{ii}\,\mathbb E X_i^2,
\]
since \(\mathbb E[X_iX_j]=0\) for \(i\neq j\). The centered quadratic form therefore splits naturally into diagonal and off-diagonal pieces:
\[
X^\top A X-\mathbb E X^\top A X
=
\sum_i a_{ii}(X_i^2-\mathbb E X_i^2)
+
\sum_{i\ne j} a_{ij}X_iX_j.
\]
This decomposition is the structural starting point of essentially all modern proofs [1306.2872].

The inequality is often interpreted as the quadratic-form counterpart of Bernstein concentration. The Hilbert–Schmidt term captures the aggregate contribution of many coefficients, whereas the operator-norm term captures the effect of the largest direction of \(A\). This suggests why the same pair of norms continues to reappear in many later generalizations.

## 2. Proof structure and Gaussian mechanism

In the standard proof, the diagonal term is treated by the observation that if \(X_i\) is sub-gaussian, then \(X_i^2-\mathbb E X_i^2\) is sub-exponential. A Bernstein-type inequality then yields a Hanson–Wright-form tail for
\[
\sum_i a_{ii}(X_i^2-\mathbb E X_i^2).
\]
The off-diagonal term
\[
\sum_{i\ne j} a_{ij}X_iX_j
\]
is the main difficulty [1306.2872].

The classical route decouples the off-diagonal sum using auxiliary Bernoulli selectors, conditions on one block of variables, and turns the remaining conditional chaos into a sub-gaussian linear form. A Gaussian comparison step then reduces the argument to a quadratic form in standard Gaussians. After spectral decomposition of the relevant matrix, the proof uses the exact identity
\[
\mathbb E e^{\theta(g^2-1)} = e^{-\theta}(1-2\theta)^{-1/2}, \qquad \theta<1/2,
\]
for \(g\sim N(0,1)\), and finally optimizes a Chernoff bound over the admissible range of the exponential parameter [2509.00881].

A later elementary proof keeps the same diagonal/off-diagonal architecture but avoids Bourgain’s convex decoupling argument. Its stated ingredients are decomposition into diagonal and off-diagonal parts, elementary mgf bounds for sub-gaussian and sub-exponential variables, a Gaussian randomization or rotation argument, and spectral decomposition of the matrix [2509.00881]. In that presentation, the appearance of \(\|A\|_{HS}\) and \(\|A\|\) is completely transparent: after diagonalization, they are simply
\[
\sum_i \lambda_i^2 = \|A\|_F^2,\qquad \max_i |\lambda_i|=\|A\|,
\]
with \(\lambda_i\) the eigenvalues of the symmetric part of \(A\).

This proof pattern also explains why the Hanson–Wright inequality is especially natural in the Gaussian case. The quadratic chaos becomes a weighted sum of centered \(\chi_1^2\)-variables, and the concentration problem reduces to controlling the logarithm of an explicit Gaussian mgf.

## 3. Norm concentration and Gaussian refinements

A closely related consequence is norm concentration. If \(Q=A^\top A\), then
\[
X^\top QX=\|AX\|_2^2,\qquad \mathbb E[X^\top QX]=\|A\|_{HS}^2,
\]
and Hanson–Wright yields
\[
\mathbb P\Big\{\big|\|AX\|_2-\|A\|_{HS}\big|>t\Big\}
\le 2\exp\!\left(-\frac{ct^2}{K^4\|A\|^2}\right).
\]
This gives sub-gaussian concentration of \(\|AX\|_2\) around \(\|A\|_{HS}\) at scale \(\|A\|\) [1306.2872].

For Gaussian vectors there is a sharper benchmark in terms of covariance geometry. If \(X\) is Gaussian with covariance \(\Sigma\) and \(B=Q\Sigma Q^\top\), then the Laurent–Massart-type bound reads
\[
\mathbb P\!\left(\|QX\|^2-\operatorname{tr}(B) >
2\sqrt{x\,\operatorname{tr}(B^2)}+2x\,\|B\|
\right)\le e^{-x},
\]
with a matching lower-tail bound
\[
\mathbb P\!\left(\|QX\|^2-\operatorname{tr}(B) <
-2\sqrt{x\,\operatorname{tr}(B^2)}
\right)\le e^{-x}.
\]
This Gaussian expression serves as an exact benchmark for several later extensions [2309.02302].

Quantitative refinement of the Gaussian Hanson–Wright constant has been a separate line of work. One note revisiting the Gaussian proof showed that in the real symmetric case the best constant \(C_{HW}\) in
\[
\Pr\!\Big(\big|x^TAx-\mathbb E[x^TAx]\big|\ge a\Big)
\le 2\exp\!\Big(
-C_{HW}\min\Big\{\frac{a^2}{\|A\|_2^2},\frac{a}{\|A\|}\Big\}
\Big)
\]
satisfies
\[
C_{HW}\ge C^*:=\max_{0<r<1}\min\Big\{\frac r4,\frac{1}{8f(r)}\Big\}\approx 0.1457,
\qquad
f(r)=\sum_{j=0}^{\infty}\frac{r^j}{j+2}
\]
[2111.00557]. A later refinement of Gaussian quadratic-chaos concentration reported constants at least \(0.145\) in the symmetric case and at least \(0.152\) in the positive-semidefinite case, obtained by sharpening both Hanson–Wright and Laurent–Massart analyses [2412.03774]. The same paper also derived a sequence of bounds indexed by \(m=1,2,3,\dots\) involving Schatten norms, with a phase transition: \(m=1\) is tightest for smaller deviations, while larger \(m\) becomes tighter for sufficiently large deviations [2412.03774].

## 4. Dependence, local mgf control, and heavy tails

A major extension replaces coordinatewise independence by concentration of the whole vector. If \(X\) is a mean-zero random vector in \(\mathbb R^n\) with the convex concentration property with constant \(K\), then for every matrix \(A\) and every \(t>0\),
\[
\mathbb P\!\left(\left|X^TAX-\mathbb E(X^TAX)\right|\ge t\right)
\le
2\exp\!\left(
-\frac1C
\min\!\left(
\frac{t^2}{K^2\|A\|_{HS}^2\,\|\operatorname{Cov}(X)\|},
\frac{t}{K^2\|A\|}
\right)
\right),
\]
and in the isotropic case this becomes
\[
\mathbb P\!\left(\left|X^TAX-\operatorname{tr}A\right|\ge t\right)
\le
2\exp\!\left(
-\frac1C
\min\!\left(
\frac{t^2}{K^2\|A\|_{HS}^2},
\frac{t}{K^2\|A\|}
\right)
\right).
\]
No coordinatewise independence is assumed; examples mentioned in this framework include vectors under uniform mixing, Dobrushin-type dependence criteria, sampling without replacement, and bounded product measures [1409.8457]. The same source also stresses a common misconception: it is not true that if \(X=(X_1,\dots,X_n)\) has i.i.d. subgaussian coordinates, then \(X\) automatically has the convex concentration property with dimension-free constant [1409.8457].

A different extension weakens global sub-gaussianity to a bounded-domain exponential moment condition. Let \(\xi\) be centered, assume \(V^2\ge \operatorname{Var}(\xi)\), and suppose
\[
\log \mathbb E \exp\bigl(\langle V^{-1}\xi,u\rangle\bigr)\le \frac{|u|^2}{2},
\qquad |u|\le g.
\]
For a linear map \(Q\), with \(B=QV^2Q^\top\), there is an upper quantile function \(z_c(B,x)\) such that
\[
\mathbb P(|Q\xi|>z_c(B,x))\le 3e^{-x}.
\]
The distinctive feature is a phase transition: for \(x\le x_c\),
\[
z_c^2(B,x)=\operatorname{tr}(B)+2\sqrt{x\,\operatorname{tr}(B^2)}+2x|B|,
\]
which exactly matches the Gaussian benchmark, whereas for \(x>x_c\), \(z_c(B,x)\) grows linearly in \(x\) [2309.02302]. This is not presented as a full theorem for arbitrary indefinite quadratic forms; it is tailored to norms of linear images \(|Q\xi|\), equivalently to the positive-semidefinite quadratic form \(\xi^\top Q^\top Q\xi\) [2309.02302].

Heavy-tail generalization replaces sub-gaussian coordinates by \(\psi_\alpha\)-variables. If \(X_1,\dots,X_n\) are independent, centered, and satisfy \(\|X_i\|_{\psi_\alpha}<K\) for \(\alpha\in(0,2]\), then for symmetric \(A\),
\[
\mathbb P\!\left(\big|X^\top A X-\mathbb E X^\top A X\big|\ge t\right)
\le
2\exp\!\left[
-\frac1{C_\alpha}
\min\!\left(
\frac{t^2}{K^4\|A\|_{HS}^2},
\left(\frac{t}{K^2\|A\|_{op}}\right)^{\alpha/2}
\right)
\right].
\]
The small-deviation term remains Gaussian-like, but the large-deviation exponent degrades from \(1\) to \(\alpha/2\), reflecting the heavier tail of \(X_i^2\) [2002.10761].

## 5. Uniform and supremum versions

A natural strengthening is to control not one quadratic form but an entire class. For a bounded family \(\mathcal A\) of matrices, one studies
\[
Z=\sup_{A\in\mathcal A}\big(X^TAX-\mathbb E(X^TAX)\big).
\]
Under the full concentration property for all 1-Lipschitz functions with constant \(K\), one has
\[
\mathbb P\big(|Z-\mathbb EZ|\ge t\big)
\le
2\exp\!\left(
-\frac1C
\min\!\left(
\frac{t^2}{K^2\|X\|_{\mathcal A}^2},
\frac{t}{K^2\sup_{A\in\mathcal A}\|A\|}
\right)
\right),
\]
where
\[
\|X\|_{\mathcal A}=\mathbb E\sup_{A=[a_{ij}]\in\mathcal A}|(A+A^T)X|.
\]
This is a two-sided concentration inequality around the mean for suprema of quadratic forms [1409.8457].

For independent centered subgaussian coordinates, the entropy-method approach yields a uniform Hanson–Wright inequality with a different complexity parameter. Let
\[
Z_{\mathcal A}(X)=\sup_{A\in\mathcal A}\bigl(X^TAX-\mathbb E X^TAX\bigr),\qquad
M=\Bigl\|\max_i |X_i|\Bigr\|_{\psi_2}.
\]
Then for finite symmetric \(\mathcal A\) and
\[
t\ge \max\left\{M\,\mathbb E\sup_{A\in\mathcal A}\|AX\|,\ M^2\sup_{A\in\mathcal A}\|A\|\right\},
\]
\[
\mathbb P\bigl(Z_{\mathcal A}(X)-\mathbb E Z_{\mathcal A}(X)\ge t\bigr)
\le
\exp\!\left(
-c\min\!\left(
\frac{t^2}{M^2(\mathbb E\sup_{A\in\mathcal A}\|AX\|)^2},
\frac{t}{M^2\sup_{A\in\mathcal A}\|A\|}
\right)
\right).
\]
This framework recovers Talagrand’s right-tail bound for Rademacher chaos and extends earlier uniform results to unbounded subgaussian coordinates by truncating both gradients and coordinates [1812.03548]. The same work emphasizes that in general one cannot replace \(M=\|\max_i |X_i|\|_{\psi_2}\) by \(K=\max_i\|X_i\|_{\psi_2}\) in the uniform theorem [1812.03548].

A uniform \(\psi_\alpha\)-extension is also available. For
\[
f(X)=\sup_{A\in\mathcal A}\big(X^\top A X-\mathbb E X^\top A X\big),
\qquad
K=\Bigl\|\max_i |X_i|\Bigr\|_{\psi_\alpha},
\]
one has
\[
\mathbb P\big(f(X)-\mathbb Ef(X)\ge t\big)
\le
2\exp\!\left[
-\frac1{C_\alpha}
\min\!\left(
\left(\frac{t}{K\,\mathbb E\sup_{A\in\mathcal A}\|AX\|_2}\right)^\alpha,
\left(\frac{t}{K^2\sup_{A\in\mathcal A}\|A\|_{op}}\right)^{\alpha/2}
\right)
\right].
\]
The same source notes that this uniform theorem is weaker than the one-matrix theorem: it controls only the upper tail, and the small-deviation term is \(\alpha\)-subexponential rather than Gaussian unless \(\alpha=2\) [2002.10761].

## 6. Infinite-dimensional and tensor extensions

One infinite-dimensional extension replaces scalar coordinates by Hilbert-space-valued random variables. Let \(\mathbb H\) be a real separable Hilbert space, let \(X_1,\dots,X_n\in\mathbb H\) be independent centered \(sub\text{-}gaussian(\Gamma)\) random variables, and consider
\[
Q=\sum_{i,j=1}^{n} a_{ij}\langle X_i,X_j\rangle.
\]
Under an additional Bernstein-type condition on \(\|X_i\|^2-\mathbb E\|X_i\|^2\), the two-sided Hilbert-space Hanson–Wright inequality becomes
\[
\mathbb P\!\left( \left| Q - \mathbb E[Q] \right| > t \right)
\le
2 \exp\!\left[
-C \min\!\left(
\frac{t^2}{L^4 \|\Gamma\|_{HS}^2 \|A\|_{HS}^2},
\frac{t}{L^2 \|\Gamma\|_{op}\|A\|_{op}}
\right)\right].
\]
Here the finite-dimensional scalar parameter \(K\) is replaced by the operator-valued covariance proxy \(\Gamma\), and the theorem is dimension-free in the ambient Hilbert space [1810.11180].

A Banach-space version treats coefficients \(a_{ij}\) taking values in a normed space \(F\), so that
\[
Q_A=\sum_{i,j=1}^n a_{ij}(g_ig_j-\delta_{ij})
\]
is \(F\)-valued. The resulting Gaussian theory identifies Banach-valued analogues of the Hilbert–Schmidt and operator norms, notably
\[
U:=\sup_{\|x\|_2\le 1}\mathbb E\Bigl\|\sum_{i,j} a_{ij}x_i g_j\Bigr\|
+
\sup_{\|(x_{ij})\|_2\le 1}\Bigl\|\sum_{i,j} a_{ij}x_{ij}\Bigr\|,
\]
and
\[
V:=\sup_{\|x\|_2\le 1,\ \|y\|_2\le 1}\Bigl\|\sum_{i,j} a_{ij}x_i y_j\Bigr\|.
\]
From the Gaussian bounds one deduces a subgaussian Hanson–Wright extension:
\[
\mathbb P\!\left(
\left\|\sum_{i,j} a_{ij}\bigl(X_iX_j-\mathbb E(X_iX_j)\bigr)\right\|>t
\right)
\le
2\exp\!\left[
-\frac1C
\min\!\left(
\frac{t^2}{\alpha^4 U^2},
\frac{t}{\alpha^2 V}
\right)
\right]
\]
for independent mean-zero \(\alpha\)-subgaussian \(X_i\), provided \(t\) is above the stated Gaussian threshold [1811.00353]. In general Banach spaces the conjectured optimal moment comparison is known only up to logarithmic factors, while in certain spaces, including \(L_r\)-spaces, those logarithmic factors can be eliminated [1811.00353].

Tensor extensions proceed in two different directions. For Kronecker products of independent subgaussian vectors,
\[
X=X^{(1)}\otimes\cdots\otimes X^{(d)},
\]
the quadratic form
\[
X^TAX
=
(X^{(1)}\otimes\cdots\otimes X^{(d)})^T A (X^{(1)}\otimes\cdots\otimes X^{(d)})
\]
admits moment bounds of the form
\[
\|X^TAX-\mathbb EX^TAX\|_{L_p}\le C(d)\,m_p,
\]
where \(m_p\) is expressed through explicit tensor partition norms of contractions \(A^{(I)}\) of the coefficient array [2106.13345]. The same framework yields improved concentration inequalities for
\[
\|A(X^{(1)}\otimes\cdots\otimes X^{(d)})\|_2
\]
and is sharp up to \(d\)-dependent constants in the Gaussian case [2106.13345].

A different tensor generalization replaces matrices by Hermitian tensors under the Einstein product and studies the Ky Fan \(k\)-norm of a polynomial of a quadratic tensor sum. The proof again splits the quadratic object into diagonal and coupling parts, applies a generalized tensor Chernoff bound to the diagonal part, and uses a decoupling inequality before applying tensor Chernoff to the coupling part [2203.00659]. This is structurally analogous to Hanson–Wright, but the scalar tail of a quadratic form is replaced by concentration of
\[
\left\|f(\mathcal X^T\mathcal A\mathcal X)\right\|_{(k)}.
\]

## 7. Sparse variants and applications

Random sparsification changes the geometry of quadratic concentration. If \(X\circ \xi\) denotes coordinatewise masking by independent Bernoulli variables \(\xi_i\) with \(\mathbb P(\xi_i=1)=p_i\), then a sparse Hanson–Wright inequality controls
\[
(X\circ\xi)^TA(X\circ\xi)
\]
with variance proxy
\[
\sum_{k=1}^m p_k a_{kk}^2+\sum_{i\ne j} a_{ij}^2 p_i p_j
\]
instead of \(\|A\|_F^2\). One formulation states
\[
\mathbb{P}\!\left(
\left| (X\circ \xi)^T A (X\circ \xi) -\mathbb{E}\big[(X\circ \xi)^T A (X\circ \xi)\big] \right| > t
\right)
\le
2\exp\!\left(
-c\min\!\left(
\frac{t^2}{ K^4\!\left(\sum_{k=1}^m p_k a_{kk}^2+\sum_{i\ne j} a_{ij}^2 p_i p_j\right)},
\frac{t}{K^2\|A\|}
\right)\right)
\]
[1510.05517]. This shows that diagonal and off-diagonal contributions scale differently under random coordinate survival.

A sparse bilinear extension considers
\[
(Z_1 * \gamma_1)^\top A (Z_2 * \gamma_2),
\]
where \(Z_1,Z_2\) are sub-gaussian vectors and \(\gamma_1,\gamma_2\) are Bernoulli masks. The resulting inequality has the same two-level Hanson–Wright shape but with weighted Frobenius norm
\[
\|A\|_{F,\pi}
=
\sqrt{\sum_{j=1}^n \pi_{12,j} a_{jj}^2 + \sum_{i \neq j}a_{ij}^2 \pi_{1i}\pi_{2j}}
\]
in place of \(\|A\|_F\):
\[
\mathbb P\!\left(
\big| (Z_1 * \gamma_1)^\top A (Z_2 * \gamma_2) - \mathbb{E}(Z_1 * \gamma_1)^\top A (Z_2 * \gamma_2)\big| > t
\right)
\le
2\exp \left\{
- c \min \left(\frac{t^2}{K_1^2 K_2^2 \|A\|_{F,\pi}^2 }, \frac{t}{K_1 K_2 \|A\|_2}\right)
\right\}
\]
[2209.05685].

Heavy-tailed sparse variants extend this picture further. For sparse \(\alpha\)-sub-exponential vectors \(\xi_i=\delta_i\zeta_i\), one simplified inequality reads
\[
\mathbb{P}\!\left\{ \left|\xi^\top A\xi-\mathbb{E}\xi^\top A\xi\right| \ge L(\alpha)^2 t \right\}
\le
2\exp\!\left(
-c(\alpha)\min\left\{
\frac{t^2}{ \sum_k a_{kk}^2 p_k+\sum_{i\ne j} a_{ij}^2 p_i p_j },
\left(\frac{t}{\|A\|_{l_2\to l_2}}\right)^{\alpha/2}
\right\}
\right),
\]
and for \(0<\alpha\le 1\) a sharper bound introduces additional weighted operator, row-\(\ell_2\), and max-entry scales [2505.20799]. A later sparse \(\alpha\)-subexponential theory derives three-regime tails for off-diagonal forms,
\[
\exp\!\left[-C_\alpha \min\left\{
\frac{t^2}{K^4\sum_{i,j}a_{ij}^2p_ip_j},
\frac{t}{K^2\gamma_{1,\infty}},
\left(\frac{t}{K^2\|A\|_{\max}}\right)^{\min\{\alpha/2,1/2\}}
\right\}\right],
\]
and uses them to obtain local laws and complete eigenvector delocalization for sparse \(\alpha\)-subexponential Hermitian random matrices, as well as concentration of \(\|BX\|_2\) for sparse \(\alpha\)-subexponential vectors [2410.15652].

Applications of Hanson–Wright and its variants are broad. The classical inequality yields concentration of the distance from a random vector to a fixed subspace and bounds on norms of products of deterministic and random matrices [1306.2872]. Uniform variants recover concentration for empirical covariance operators of Banach-valued Gaussian variables [1409.8457] and covariance estimation with missing observations [1812.03548]. The local-mgf extension applies to Bernoulli vector sums and covariance estimation in Frobenius norm [2309.02302]. The Hilbert-space version underlies exponential recovery guarantees for generalized \(K\)-means clustering with non-Euclidean data [1810.11180]. These examples suggest that the Hanson–Wright inequality is best viewed not as a single estimate for scalar quadratic forms, but as a concentration principle whose central matrix scales—variance-type aggregate size and worst-direction amplification—persist across dependence, sparsity, tensorization, and infinite-dimensional settings.

Source: https://www.emergentmind.com/topics/hanson-wright-inequality