- The paper develops random-matrix asymptotics showing that high-dimensional out-of-sample risk converges to a quadratic form in model-averaging weights, exposing how in-sample criteria underestimate variance near interpolation.
- The paper shows that averaging converts the divergent peaks of individual interpolating models into a bounded risk ridge through emergent smoothing, while retaining double-descent behavior that approaches a classical U-shape as sample size grows.
- The paper introduces LaMA, which adds an asymptotic variance correction and variance-weighted ℓ2 regularization, achieving the lowest or most stable test error in simulations and real-data experiments, including a Crime-data error variance of 0.0293 versus 249.87 for JMA at n=18.
This paper develops an asymptotic theory for out-of-sample prediction risk in high-dimensional model averaging and uses it to construct a new weight choice criterion, Large Model Averaging (LaMA). The setting is linear regression where the dimension of the largest candidate model kM is comparable to the sample size n, a regime in which classical model averaging criteria—built on in-sample risk—are known to fail.
Motivation: the in-sample/out-of-sample mismatch
The authors work within the nested-model framework of Mallows model averaging (MMA): M candidate models of increasing dimension k1<⋯<kM are fitted by minimum ℓ2-norm least squares, and predictions are combined with weights ω∈Hn. The central observation is that when kq/n→cq∈(0,∞), the sample covariance matrix is no longer consistent for the population covariance, so the in-sample loss ceases to be a valid surrogate for out-of-sample loss. The paper identifies the mechanism precisely: near-zero eigenvalues of the local sample covariance amplify noise in each least squares estimator, but evaluating loss against that same matrix masks this overfitting; on new data drawn from the population covariance Σ, the inflated variance is fully exposed. Consequently, criteria such as MMA—which are unbiased for in-sample risk—systematically underestimate true predictive risk as cq→1, and their implicit ℓ1-type penalty on weights can discard complementary information from candidate models.
Limiting risk theory via random matrices
The core theoretical contribution is an exact asymptotic characterization of the out-of-sample risk under nested models. Decomposing the conditional risk into bias and variance components, the authors use limiting spectral results (building on Hastie et al.'s analysis of ridgeless interpolation) to show that both components converge to quadratic forms in the weights:
n0
with closed-form entries depending only on the aspect ratios n1, the noise variance n2, and coefficient norms. Under Gaussian designs and a signal condition n3—which permits slowly diverging signals up to nearly n4, weaker than the usual bounded-signal assumption—the total risk converges almost surely to n5. A companion theorem extends these limits to general positive definite covariance matrices, replacing coefficient norms with Schur-complement-type quantities n6.
The piecewise structure of n7 and n8 across three regimes (n9; mixed; both M0) reveals how cross-model interaction terms behave: bias coefficients diverge as either aspect ratio approaches the interpolation threshold, while in the over-parameterized regime included coefficients are compressed by M1 and omitted ones amplified by M2.
Double descent and emergent smoothing
Three-dimensional visualization of the limiting risk surface over M3 exposes two coupled phenomena. First, the ensemble inherits double descent from its constituent models: along fixed-M4 sections, risk declines, peaks near the diagonal M5, then descends again. Second—and this is the paper's key conceptual claim—weighted aggregation produces an emergent smoothing effect: although individual near-interpolating models have divergent variance, the finite weighted average yields only a bounded risk ridge at the interpolation boundary rather than infinite divergence. The ridge remains finite because averaging diverse implicitly regularized candidates suppresses localized variance explosions. Along the M6-axis, the surface exhibits sample-wise double descent, and for large M7 the transverse sections degenerate from full double-descent curves into classical U-shapes, recovering standard large-sample behavior. The authors also document sensitivity of the risk landscape to signal-to-noise ratio and to the decay rate of the coefficient sequence, with slow decay shifting the risk minimum toward the over-parameterized region.
The LaMA criterion
Building on the asymptotics, LaMA augments the MMA criterion with two components: a strictly positive variance correction
M8
which quantifies exactly how much in-sample variance understates out-of-sample variance, and a variance-weighted M9 penalty on the weights that encourages dense, smooth weighting and shrinks high-uncertainty models more aggressively. Simple algebra shows the resulting criterion equals an unbiased estimate of in-sample bias plus an asymptotic out-of-sample variance term plus the penalty—an explicit "both-risk" balancing objective. Notably, no analogous correction is applied to bias, because the bias discrepancy depends on the unknown k1<⋯<kM0 and is practically inestimable; this is a deliberate asymmetry in the method. The tuning parameter k1<⋯<kM1 is set analytically as a ratio of dispersion measures of the out-of-sample variances to those of the fitting biases, avoiding data-driven cross-validation.
Numerical evidence
Simulations with k1<⋯<kM2, k1<⋯<kM3, and k1<⋯<kM4 ratios approaching one compare LaMA against AIC/BIC selection, smoothed variants, MMA, and JMA using relative in-sample and out-of-sample losses. LaMA attains the lowest out-of-sample loss in essentially all small-sample configurations, with the gap widening as k1<⋯<kM5 grows; MMA exhibits characteristic overfitting and can lose predictive ability entirely at large k1<⋯<kM6, while JMA's loss becomes volatile. LaMA's loss curve is also markedly flatter in k1<⋯<kM7, indicating robustness to low signal-to-noise conditions. As k1<⋯<kM8 increases, differences among methods narrow, consistent with the recovery of classical asymptotics. On the U.S. Crime (k1<⋯<kM9) and Motor Trend Car (ℓ20) datasets with 1000 random train/test splits, LaMA achieves the lowest mean test error and, strikingly, the smallest variance at every training size—for example, at ℓ21 on the Crime data its error variance is 0.0293 versus 249.87 for JMA and roughly 2886 for AIC-based methods, a difference of several orders of magnitude. This indicates substantially greater stability under small training samples than any competitor.
Limitations and open questions
Several caveats bear directly on the strength of the results. The exact limiting risk is derived for nested models with uniform or given weights; the optimization theory for the chosen weights (e.g., formal asymptotic optimality of LaMA's estimator of the out-of-sample risk minimizer) is not established, and the authors explicitly identify proving out-of-sample optimality in the ℓ22 regime as an open problem, noting that existing optimality results—even recent relaxations such as Peng et al.—remain confined to in-sample risk. The signal condition ℓ23 is required for almost sure convergence of the bias component (weaker versions hold in probability), and the general-covariance theorem covers only the under-parameterized case ℓ24. The treatment of the singular point ℓ25 is necessarily by exclusion or truncation, since the limit genuinely diverges there. Finally, the framework is specific to linear ensembles with squared loss; extension to nonlinear ensembles or dependent data is left unaddressed.
Conclusion
The paper provides a closed-form random-matrix characterization of out-of-sample risk for high-dimensional model averaging, demonstrating that ensembles convert the divergent peaks of individual interpolating models into a bounded ridge through an emergent smoothing mechanism. LaMA operationalizes this insight by correcting the systematic in-sample underestimation of out-of-sample variance and adding adaptive ℓ26 regularization, yielding strong empirical predictive accuracy and stability in regimes where ℓ27 is comparable to ℓ28. Its main theoretical debt—formal optimality guarantees for out-of-sample risk in this regime—remains open.