Papers
Topics
Authors
Recent
Search
2000 character limit reached

Quadratic Shrinkage: Theory & Applications

Updated 12 July 2026
  • Quadratic shrinkage is a family of estimators that improve traditional methods by shrinking estimates toward a target to lower quadratic risk under quadratic loss.
  • It encompasses techniques ranging from the James–Stein estimator and its positive-part version to higher-order polynomial shrinkage for improved risk performance.
  • The framework extends to predictive densities, matrix estimation, and structured regression penalties, demonstrating versatility across various high-dimensional models.

Quadratic shrinkage denotes a family of estimators that modify a baseline estimator by shrinking it toward zero, a fixed target, a subspace, or a structured prior in order to reduce risk under a quadratic criterion. In the canonical multivariate normal mean model XNp(μ,σ2Ip)X\sim N_p(\mu,\sigma^2I_p) with loss L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^2, the usual estimator XX is minimax but, for p3p\ge3, inadmissible; this observation motivates James–Stein rules, Bayes formulations based on superharmonic marginals, and a large body of extensions to predictive densities, matrix-valued estimators, quadratic-penalty methods, and high-dimensional models under general quadratic loss (George et al., 2012, Matsuda, 2022).

1. Classical normal-mean formulation

The classical setup observes

XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),

and seeks an estimator δ(X)\delta(X) of μ\mu under quadratic loss

L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.

The usual estimator δ0(X)=X\delta_0(X)=X has constant risk

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,

so its maximum risk over L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^20 equals L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^21, and it is minimax (George et al., 2012).

For L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^22, Stein (1956) showed that L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^23 is inadmissible. James and Stein (1961) proposed

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^24

with risk

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^25

Hence L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^26 is minimax and strictly better than L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^27. If L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^28 is unknown, one may replace L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^29 by an estimate.

The essential feature is a data-dependent multiplicative factor smaller than XX0. Risk reduction is obtained not by unbiasedness but by a bias–variance trade-off that lowers quadratic risk uniformly over XX1.

2. Bayes representation and superharmonicity

The Bayes formulation makes shrinkage explicit. Under the conjugate Gaussian prior

XX2

the posterior of XX3 given XX4 is

XX5

and the Bayes rule is the posterior mean

XX6

As XX7, this tends to XX8; as XX9, it tends to p3p\ge30. One may choose p3p\ge31 to optimize Bayes risk or replace the shrinkage factor by its unbiased estimate in an empirical-Bayes construction (George et al., 2012).

For a general prior p3p\ge32, possibly improper, let

p3p\ge33

Brown (1971) and Stein (1974) showed that the Bayes-posterior mean can be written as

p3p\ge34

Substituting this representation into the risk and using integration by parts yields

p3p\ge35

A sufficient condition for minimaxity is either

p3p\ge36

so that p3p\ge37 is superharmonic, or the weaker condition

p3p\ge38

This criterion organizes a broad class of shrinkage priors. The improper harmonic prior

p3p\ge39

has marginal

XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),0

which is superharmonic for XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),1, so the corresponding Bayes rule XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),2 dominates XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),3. Strawderman’s proper priors XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),4, obtained as scale-mixtures of Gaussians with

XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),5

provide examples for which XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),6 is superharmonic even though XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),7 is not.

The same device yields multiple shrinkage estimators. Given targets or subspaces XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),8, one forms recentered superharmonic marginals XNp(μ,σ2Ip),X\sim N_p(\mu,\sigma^2I_p),9 and then a mixture marginal

δ(X)\delta(X)0

which remains superharmonic. The resulting Bayes rule

δ(X)\delta(X)1

is minimax and adapts to whichever target is most supported by the data.

An empirical-Bayes analogue of the conjugate-normal rule uses the unbiased estimate δ(X)\delta(X)2 of the shrinkage factor and truncates it at zero, producing the positive-part estimator

δ(X)\delta(X)3

described as minimax and admissible.

3. Second-order and higher-order polynomial shrinkage

A different use of the adjective “quadratic” appears in polynomial shrinkage estimators of a multivariate normal mean. With

δ(X)\delta(X)4

and balanced squared-error loss

δ(X)\delta(X)5

one considers estimators of the form

δ(X)\delta(X)6

The second-degree, or quadratic, shrinkage function is

δ(X)\delta(X)7

so that

δ(X)\delta(X)8

In the balanced-loss setting the optimal first- and second-order coefficients are

δ(X)\delta(X)9

The estimator is minimax relative to the MLE μ\mu0 if and only if

μ\mu1

When μ\mu2, this reduces to

μ\mu3

(Benkhaled et al., 2021).

The explicit risk formula for the quadratic estimator is

μ\mu4

Choosing

μ\mu5

gives

μ\mu6

For μ\mu7, the quadratic rule therefore strictly dominates both the usual James–Stein estimator and the MLE.

The paper’s figures plot the risk ratios

μ\mu8

against μ\mu9 for L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.0 and L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.1; all curves lie below L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.2, and each higher-order polynomial pushes the ratio down further, especially when L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.3 and L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.4 are small. The same Stein-lemma argument extends to third- and higher-order rules with

L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.5

but one needs L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.6 to control L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.7 and guarantee domination. The practical gain diminishes because L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.8 for large L(μ,δ)=δμ2,R(μ,δ)=Eμδ(X)μ2.L(\mu,\delta)=\|\delta-\mu\|^2, \qquad R(\mu,\delta)=E_\mu\|\delta(X)-\mu\|^2.9 and high δ0(X)=X\delta_0(X)=X0, and the complexity and variance of estimating higher-order inverse moments become unattractive.

4. Predictive density, matrix, and singular-value generalizations

Quadratic shrinkage extends beyond point estimation of a single normal mean. In the normal linear model,

δ0(X)=X\delta_0(X)=X1

independent, the Bayes predictive density under δ0(X)=X\delta_0(X)=X2 can be written in the shrinkage form

δ0(X)=X\delta_0(X)=X3

where δ0(X)=X\delta_0(X)=X4 is the uniform-prior predictor. Diagonalizing the covariances reduces the problem to the canonical mean setting, and the same superharmonic-marginal conditions ensure minimaxity. In nonparametric regression, the Gaussian sequence model

δ0(X)=X\delta_0(X)=X5

with δ0(X)=X\delta_0(X)=X6 constrained in an ellipsoid, has minimax KL-risk asymptotically attained by Bayes predictive densities under Gaussian priors, in direct analogy with Pinsker’s theorem for δ0(X)=X\delta_0(X)=X7-risk (George et al., 2012).

A matrix-valued version arises when estimating an δ0(X)=X\delta_0(X)=X8 mean matrix δ0(X)=X\delta_0(X)=X9 from an R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,0 normal data matrix R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,1 under the matrix quadratic loss

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,2

The usual MLE is R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,3, with risk

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,4

Applying vector James–Stein shrinkage separately to each column gives

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,5

or equivalently

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,6

Abu-Shanab, Kent and Strawderman showed that if the columnwise shrinkers satisfy the vector cross-product inequality, then R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,7 strictly dominates R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,8 under the matrix loss for

R(μ,X)=pσ2,R(\mu,X)=p\sigma^2,9

The same argument extends to positive-part James–Stein, empirical Bayes, subspace shrinkage, and superharmonic-prior rules, and also to unknown L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^200 and correlated columns after whitening (Abu-Shanab et al., 2011).

General quadratic loss can also be encoded by a positive-definite matrix L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^201. In the multivariate Gaussian sequence and matrix-denoising settings,

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^202

and

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^203

The Efron–Morris singular-value shrinker

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^204

shrinks each singular value and dominates the MLE under Frobenius loss. Under arbitrary L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^205-loss it satisfies the oracle inequality

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^206

This leads to a multivariate Pinsker linear shrinker on Sobolev ellipsoids and to a blockwise Efron–Morris construction that is stated to be exactly adaptive minimax over multivariate Sobolev ellipsoids under the corresponding quadratic loss (Matsuda, 2022).

5. Quadratic penalties and structured regression shrinkage

Quadratic shrinkage may be implemented as a penalty term rather than as a multiplicative factor. In sparse Laplacian shrinkage for high-dimensional regression, one minimizes

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^207

where L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^208 is the minimax concave penalty and L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^209 is the graph Laplacian derived from a signed adjacency matrix. The quadratic form is

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^210

This term promotes local smoothness among coefficients associated with connected predictors; unlike a global ridge penalty, the shrinkage is local on the graph. Under the stated sub-Gaussian, restricted-eigenvalue, penalty-size, and minimum-signal conditions, the estimator is selection-consistent and equal to the oracle Laplacian shrinkage estimator with probability at least L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^211. Proposition 5.2 further gives a generalized grouping property within graph cliques (Huang et al., 2011).

A related quadratic penalization appears in shrinkage for categorical regressors and group means. If L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^212 denotes the sample mean in group L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^213, shrinkage toward first-stage targets L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^214 is obtained by minimizing

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^215

Writing L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^216 and L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^217, the unique minimizer is

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^218

Under the local-to-zero framework

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^219

shrinkage introduces asymptotic bias

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^220

but reduces variance. For every bounded L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^221, provided L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^222, the plug-in shrinkage estimator has asymptotic risk strictly below OLS uniformly over L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^223. In Monte Carlo experiments with L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^224 and L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^225, the plug-in quadratic-shrinkage estimator never exceeds OLS risk and often improves by up to L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^226 (Heiler et al., 2019).

6. High-dimensional, spectral, and heteroscedastic extensions

In high-dimensional mean estimation with unknown covariance, Wang et al. consider i.i.d. observations

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^227

with L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^228, L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^229, finite fourth moments, and no normality assumption. Accuracy is measured by

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^230

where L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^231 is known positive definite. Their estimator

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^232

shrinks the sample mean toward L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^233, with data-driven coefficients L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^234 formed from four L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^235-statistics L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^236. Under Assumption 3.1 and the asymptotic condition L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^237, one has

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^238

For the percentage relative improvement

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^239

Corollary 3.1 states that if L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^240, then L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^241; if L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^242, then L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^243 some L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^244; and if L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^245, then L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^246. Simulations and leukemia microarray analysis are reported to show practical improvement over the sample mean and competing shrinkage rules, especially when L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^247 is non-diagonal and training sizes are very small (Wang et al., 2012).

In spectral matrix estimation for partial coherencies, one observes

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^248

and a standard shrinkage estimator is

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^249

Schneider-Luftman and Walden study the quadratic-loss criterion

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^250

derive oracle rules QLa and QLb for spectral-matrix shrinkage, and also formulate precision-matrix shrinkage

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^251

A central methodological point is that Hilbert–Schmidt shrinkage can “over-shrink” and wipe out true large partial coherencies. In EEG-derived simulations, HS-based shrinkage can increase error when L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^252 only slightly exceeds L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^253, whereas oracle QLb gives the best PRISE, L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^254 for L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^255. Among the feasible methods, QLP-est retains almost all of the oracle QLP gain, L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^256 at L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^257 and L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^258 at L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^259, with much less variance across frequency; it is recommended as the full-estimation method (Schneider-Luftman et al., 2015).

A further generalization treats families of distributions with quadratic variance function

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^260

Estimators

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^261

then have oracle weight

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^262

or, when a sample-size factor L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^263 appears,

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^264

Xie, Kou and Brown construct semiparametric URE-based shrinkage estimators

L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^265

with monotonicity L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^266 whenever L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^267, and show that the URE rule is asymptotically first-order risk-optimal in its class. Their simulations report L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^268–L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^269 lower MSE than classical James–Stein and empirical-Bayes estimators when variances and means are correlated or heteroscedastic (Xie et al., 2016).

Across these formulations, quadratic shrinkage is unified less by a single estimator than by a common risk logic: a controlled introduction of bias to secure lower quadratic risk. The resulting theory is dimension-sensitive and model-sensitive. James–Stein dominance requires L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^270; second-order polynomial improvement requires L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^271; higher-order polynomial rules require L(μ,δ)=δμ2L(\mu,\delta)=\|\delta-\mu\|^272; and, in spectral estimation, methods based on inverse-power traces may be unstable even when their oracle versions are attractive. The main analytical tools are minimaxity, Bayes posterior means, superharmonicity, oracle inequalities, and explicit bias–variance decompositions.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Quadratic Shrinkage.