Papers
Topics
Authors
Recent
Search
2000 character limit reached

Quadratic Shrinkage: Theory & Applications

Updated 12 July 2026
  • Quadratic shrinkage is a family of estimators that improve traditional methods by shrinking estimates toward a target to lower quadratic risk under quadratic loss.
  • It encompasses techniques ranging from the James–Stein estimator and its positive-part version to higher-order polynomial shrinkage for improved risk performance.
  • The framework extends to predictive densities, matrix estimation, and structured regression penalties, demonstrating versatility across various high-dimensional models.

Quadratic shrinkage denotes a family of estimators that modify a baseline estimator by shrinking it toward zero, a fixed target, a subspace, or a structured prior in order to reduce risk under a quadratic criterion. In the canonical multivariate normal mean model X∼Np(μ,σ2Ip)X\sim N_p(\mu,\sigma^2I_p) with loss L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^2, the usual estimator XX is minimax but, for p≥3p\ge3, inadmissible; this observation motivates James–Stein rules, Bayes formulations based on superharmonic marginals, and a large body of extensions to predictive densities, matrix-valued estimators, quadratic-penalty methods, and high-dimensional models under general quadratic loss (George et al., 2012, Matsuda, 2022).

1. Classical normal-mean formulation

The classical setup observes

X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),

and seeks an estimator δ(X)\delta(X) of μ\mu under quadratic loss

L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.

The usual estimator δ0(X)=X\delta_0(X)=X has constant risk

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,

so its maximum risk over L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^20 equals L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^21, and it is minimax (George et al., 2012).

For L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^22, Stein (1956) showed that L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^23 is inadmissible. James and Stein (1961) proposed

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^24

with risk

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^25

Hence L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^26 is minimax and strictly better than L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^27. If L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^28 is unknown, one may replace L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^29 by an estimate.

The essential feature is a data-dependent multiplicative factor smaller than XX0. Risk reduction is obtained not by unbiasedness but by a bias–variance trade-off that lowers quadratic risk uniformly over XX1.

2. Bayes representation and superharmonicity

The Bayes formulation makes shrinkage explicit. Under the conjugate Gaussian prior

XX2

the posterior of XX3 given XX4 is

XX5

and the Bayes rule is the posterior mean

XX6

As XX7, this tends to XX8; as XX9, it tends to p≥3p\ge30. One may choose p≥3p\ge31 to optimize Bayes risk or replace the shrinkage factor by its unbiased estimate in an empirical-Bayes construction (George et al., 2012).

For a general prior p≥3p\ge32, possibly improper, let

p≥3p\ge33

Brown (1971) and Stein (1974) showed that the Bayes-posterior mean can be written as

p≥3p\ge34

Substituting this representation into the risk and using integration by parts yields

p≥3p\ge35

A sufficient condition for minimaxity is either

p≥3p\ge36

so that p≥3p\ge37 is superharmonic, or the weaker condition

p≥3p\ge38

This criterion organizes a broad class of shrinkage priors. The improper harmonic prior

p≥3p\ge39

has marginal

X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),0

which is superharmonic for X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),1, so the corresponding Bayes rule X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),2 dominates X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),3. Strawderman’s proper priors X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),4, obtained as scale-mixtures of Gaussians with

X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),5

provide examples for which X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),6 is superharmonic even though X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),7 is not.

The same device yields multiple shrinkage estimators. Given targets or subspaces X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),8, one forms recentered superharmonic marginals X∼Np(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),9 and then a mixture marginal

δ(X)\delta(X)0

which remains superharmonic. The resulting Bayes rule

δ(X)\delta(X)1

is minimax and adapts to whichever target is most supported by the data.

An empirical-Bayes analogue of the conjugate-normal rule uses the unbiased estimate δ(X)\delta(X)2 of the shrinkage factor and truncates it at zero, producing the positive-part estimator

δ(X)\delta(X)3

described as minimax and admissible.

3. Second-order and higher-order polynomial shrinkage

A different use of the adjective “quadratic” appears in polynomial shrinkage estimators of a multivariate normal mean. With

δ(X)\delta(X)4

and balanced squared-error loss

δ(X)\delta(X)5

one considers estimators of the form

δ(X)\delta(X)6

The second-degree, or quadratic, shrinkage function is

δ(X)\delta(X)7

so that

δ(X)\delta(X)8

In the balanced-loss setting the optimal first- and second-order coefficients are

δ(X)\delta(X)9

The estimator is minimax relative to the MLE μ\mu0 if and only if

μ\mu1

When μ\mu2, this reduces to

μ\mu3

(Benkhaled et al., 2021).

The explicit risk formula for the quadratic estimator is

μ\mu4

Choosing

μ\mu5

gives

μ\mu6

For μ\mu7, the quadratic rule therefore strictly dominates both the usual James–Stein estimator and the MLE.

The paper’s figures plot the risk ratios

μ\mu8

against μ\mu9 for L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.0 and L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.1; all curves lie below L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.2, and each higher-order polynomial pushes the ratio down further, especially when L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.3 and L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.4 are small. The same Stein-lemma argument extends to third- and higher-order rules with

L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.5

but one needs L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.6 to control L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.7 and guarantee domination. The practical gain diminishes because L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.8 for large L(μ,δ)=∥δ−μ∥2,R(μ,δ)=Eμ∥δ(X)−μ∥2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.9 and high δ0(X)=X\delta_0(X)=X0, and the complexity and variance of estimating higher-order inverse moments become unattractive.

4. Predictive density, matrix, and singular-value generalizations

Quadratic shrinkage extends beyond point estimation of a single normal mean. In the normal linear model,

δ0(X)=X\delta_0(X)=X1

independent, the Bayes predictive density under δ0(X)=X\delta_0(X)=X2 can be written in the shrinkage form

δ0(X)=X\delta_0(X)=X3

where δ0(X)=X\delta_0(X)=X4 is the uniform-prior predictor. Diagonalizing the covariances reduces the problem to the canonical mean setting, and the same superharmonic-marginal conditions ensure minimaxity. In nonparametric regression, the Gaussian sequence model

δ0(X)=X\delta_0(X)=X5

with δ0(X)=X\delta_0(X)=X6 constrained in an ellipsoid, has minimax KL-risk asymptotically attained by Bayes predictive densities under Gaussian priors, in direct analogy with Pinsker’s theorem for δ0(X)=X\delta_0(X)=X7-risk (George et al., 2012).

A matrix-valued version arises when estimating an δ0(X)=X\delta_0(X)=X8 mean matrix δ0(X)=X\delta_0(X)=X9 from an R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,0 normal data matrix R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,1 under the matrix quadratic loss

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,2

The usual MLE is R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,3, with risk

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,4

Applying vector James–Stein shrinkage separately to each column gives

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,5

or equivalently

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,6

Abu-Shanab, Kent and Strawderman showed that if the columnwise shrinkers satisfy the vector cross-product inequality, then R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,7 strictly dominates R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,8 under the matrix loss for

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,9

The same argument extends to positive-part James–Stein, empirical Bayes, subspace shrinkage, and superharmonic-prior rules, and also to unknown L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^200 and correlated columns after whitening (Abu-Shanab et al., 2011).

General quadratic loss can also be encoded by a positive-definite matrix L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^201. In the multivariate Gaussian sequence and matrix-denoising settings,

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^202

and

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^203

The Efron–Morris singular-value shrinker

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^204

shrinks each singular value and dominates the MLE under Frobenius loss. Under arbitrary L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^205-loss it satisfies the oracle inequality

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^206

This leads to a multivariate Pinsker linear shrinker on Sobolev ellipsoids and to a blockwise Efron–Morris construction that is stated to be exactly adaptive minimax over multivariate Sobolev ellipsoids under the corresponding quadratic loss (Matsuda, 2022).

5. Quadratic penalties and structured regression shrinkage

Quadratic shrinkage may be implemented as a penalty term rather than as a multiplicative factor. In sparse Laplacian shrinkage for high-dimensional regression, one minimizes

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^207

where L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^208 is the minimax concave penalty and L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^209 is the graph Laplacian derived from a signed adjacency matrix. The quadratic form is

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^210

This term promotes local smoothness among coefficients associated with connected predictors; unlike a global ridge penalty, the shrinkage is local on the graph. Under the stated sub-Gaussian, restricted-eigenvalue, penalty-size, and minimum-signal conditions, the estimator is selection-consistent and equal to the oracle Laplacian shrinkage estimator with probability at least L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^211. Proposition 5.2 further gives a generalized grouping property within graph cliques (Huang et al., 2011).

A related quadratic penalization appears in shrinkage for categorical regressors and group means. If L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^212 denotes the sample mean in group L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^213, shrinkage toward first-stage targets L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^214 is obtained by minimizing

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^215

Writing L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^216 and L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^217, the unique minimizer is

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^218

Under the local-to-zero framework

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^219

shrinkage introduces asymptotic bias

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^220

but reduces variance. For every bounded L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^221, provided L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^222, the plug-in shrinkage estimator has asymptotic risk strictly below OLS uniformly over L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^223. In Monte Carlo experiments with L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^224 and L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^225, the plug-in quadratic-shrinkage estimator never exceeds OLS risk and often improves by up to L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^226 (Heiler et al., 2019).

6. High-dimensional, spectral, and heteroscedastic extensions

In high-dimensional mean estimation with unknown covariance, Wang et al. consider i.i.d. observations

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^227

with L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^228, L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^229, finite fourth moments, and no normality assumption. Accuracy is measured by

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^230

where L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^231 is known positive definite. Their estimator

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^232

shrinks the sample mean toward L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^233, with data-driven coefficients L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^234 formed from four L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^235-statistics L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^236. Under Assumption 3.1 and the asymptotic condition L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^237, one has

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^238

For the percentage relative improvement

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^239

Corollary 3.1 states that if L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^240, then L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^241; if L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^242, then L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^243 some L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^244; and if L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^245, then L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^246. Simulations and leukemia microarray analysis are reported to show practical improvement over the sample mean and competing shrinkage rules, especially when L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^247 is non-diagonal and training sizes are very small (Wang et al., 2012).

In spectral matrix estimation for partial coherencies, one observes

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^248

and a standard shrinkage estimator is

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^249

Schneider-Luftman and Walden study the quadratic-loss criterion

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^250

derive oracle rules QLa and QLb for spectral-matrix shrinkage, and also formulate precision-matrix shrinkage

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^251

A central methodological point is that Hilbert–Schmidt shrinkage can “over-shrink” and wipe out true large partial coherencies. In EEG-derived simulations, HS-based shrinkage can increase error when L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^252 only slightly exceeds L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^253, whereas oracle QLb gives the best PRISE, L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^254 for L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^255. Among the feasible methods, QLP-est retains almost all of the oracle QLP gain, L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^256 at L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^257 and L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^258 at L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^259, with much less variance across frequency; it is recommended as the full-estimation method (Schneider-Luftman et al., 2015).

A further generalization treats families of distributions with quadratic variance function

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^260

Estimators

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^261

then have oracle weight

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^262

or, when a sample-size factor L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^263 appears,

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^264

Xie, Kou and Brown construct semiparametric URE-based shrinkage estimators

L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^265

with monotonicity L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^266 whenever L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^267, and show that the URE rule is asymptotically first-order risk-optimal in its class. Their simulations report L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^268–L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^269 lower MSE than classical James–Stein and empirical-Bayes estimators when variances and means are correlated or heteroscedastic (Xie et al., 2016).

Across these formulations, quadratic shrinkage is unified less by a single estimator than by a common risk logic: a controlled introduction of bias to secure lower quadratic risk. The resulting theory is dimension-sensitive and model-sensitive. James–Stein dominance requires L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^270; second-order polynomial improvement requires L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^271; higher-order polynomial rules require L(μ,δ)=∥δ−μ∥2L(\mu,\delta)=\|\delta-\mu\|^272; and, in spectral estimation, methods based on inverse-power traces may be unstable even when their oracle versions are attractive. The main analytical tools are minimaxity, Bayes posterior means, superharmonicity, oracle inequalities, and explicit bias–variance decompositions.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Quadratic Shrinkage.