Papers
Topics
Authors
Recent
Search
2000 character limit reached

Double Descent and Emergent Smoothing in Model Averaging Prediction

Published 13 May 2026 in stat.ME | (2605.13203v1)

Abstract: This paper investigates the predictive performance of model averaging in high-dimensional linear regression where the number of regressors is comparable to the sample size. We demonstrate that the double descent trajectory manifests within the model averaging framework, where the ensemble inherits the variance explosion of individual models near the interpolation boundary. However, we reveal that weighted aggregation simultaneously triggers an emergent smoothing effect that structurally suppresses the localized risk divergence, indicating that strategic weight choice serves as a vital stabilizing mechanism. Leveraging tools from random matrix theory, we derive the exact limiting out-of-sample risk under a nested model setting and provide a comprehensive characterization of the risk landscape. Building on these asymptotic results, we propose the Large Model Averaging (LaMA) method, which introduces a novel criterion incorporating in-sample bias and asymptotic out-of-sample variance to balance fitting accuracy and generalization. Numerical studies and real data applications confirm that LaMA achieves superior predictive accuracy in high-dimensional environments.

Authors (3)

Summary

  • The paper develops random-matrix asymptotics showing that high-dimensional out-of-sample risk converges to a quadratic form in model-averaging weights, exposing how in-sample criteria underestimate variance near interpolation.
  • The paper shows that averaging converts the divergent peaks of individual interpolating models into a bounded risk ridge through emergent smoothing, while retaining double-descent behavior that approaches a classical U-shape as sample size grows.
  • The paper introduces LaMA, which adds an asymptotic variance correction and variance-weighted ℓ2 regularization, achieving the lowest or most stable test error in simulations and real-data experiments, including a Crime-data error variance of 0.0293 versus 249.87 for JMA at n=18.

This paper develops an asymptotic theory for out-of-sample prediction risk in high-dimensional model averaging and uses it to construct a new weight choice criterion, Large Model Averaging (LaMA). The setting is linear regression where the dimension of the largest candidate model kMk_M is comparable to the sample size nn, a regime in which classical model averaging criteria—built on in-sample risk—are known to fail.

Motivation: the in-sample/out-of-sample mismatch

The authors work within the nested-model framework of Mallows model averaging (MMA): MM candidate models of increasing dimension k1<<kMk_1 < \cdots < k_M are fitted by minimum 2\ell_2-norm least squares, and predictions are combined with weights ωHn\omega \in \mathcal{H}_n. The central observation is that when kq/ncq(0,)k_q/n \to c_q \in (0,\infty), the sample covariance matrix is no longer consistent for the population covariance, so the in-sample loss ceases to be a valid surrogate for out-of-sample loss. The paper identifies the mechanism precisely: near-zero eigenvalues of the local sample covariance amplify noise in each least squares estimator, but evaluating loss against that same matrix masks this overfitting; on new data drawn from the population covariance Σ\Sigma, the inflated variance is fully exposed. Consequently, criteria such as MMA—which are unbiased for in-sample risk—systematically underestimate true predictive risk as cq1c_q \to 1, and their implicit 1\ell_1-type penalty on weights can discard complementary information from candidate models.

Limiting risk theory via random matrices

The core theoretical contribution is an exact asymptotic characterization of the out-of-sample risk under nested models. Decomposing the conditional risk into bias and variance components, the authors use limiting spectral results (building on Hastie et al.'s analysis of ridgeless interpolation) to show that both components converge to quadratic forms in the weights:

nn0

with closed-form entries depending only on the aspect ratios nn1, the noise variance nn2, and coefficient norms. Under Gaussian designs and a signal condition nn3—which permits slowly diverging signals up to nearly nn4, weaker than the usual bounded-signal assumption—the total risk converges almost surely to nn5. A companion theorem extends these limits to general positive definite covariance matrices, replacing coefficient norms with Schur-complement-type quantities nn6.

The piecewise structure of nn7 and nn8 across three regimes (nn9; mixed; both MM0) reveals how cross-model interaction terms behave: bias coefficients diverge as either aspect ratio approaches the interpolation threshold, while in the over-parameterized regime included coefficients are compressed by MM1 and omitted ones amplified by MM2.

Double descent and emergent smoothing

Three-dimensional visualization of the limiting risk surface over MM3 exposes two coupled phenomena. First, the ensemble inherits double descent from its constituent models: along fixed-MM4 sections, risk declines, peaks near the diagonal MM5, then descends again. Second—and this is the paper's key conceptual claim—weighted aggregation produces an emergent smoothing effect: although individual near-interpolating models have divergent variance, the finite weighted average yields only a bounded risk ridge at the interpolation boundary rather than infinite divergence. The ridge remains finite because averaging diverse implicitly regularized candidates suppresses localized variance explosions. Along the MM6-axis, the surface exhibits sample-wise double descent, and for large MM7 the transverse sections degenerate from full double-descent curves into classical U-shapes, recovering standard large-sample behavior. The authors also document sensitivity of the risk landscape to signal-to-noise ratio and to the decay rate of the coefficient sequence, with slow decay shifting the risk minimum toward the over-parameterized region.

The LaMA criterion

Building on the asymptotics, LaMA augments the MMA criterion with two components: a strictly positive variance correction

MM8

which quantifies exactly how much in-sample variance understates out-of-sample variance, and a variance-weighted MM9 penalty on the weights that encourages dense, smooth weighting and shrinks high-uncertainty models more aggressively. Simple algebra shows the resulting criterion equals an unbiased estimate of in-sample bias plus an asymptotic out-of-sample variance term plus the penalty—an explicit "both-risk" balancing objective. Notably, no analogous correction is applied to bias, because the bias discrepancy depends on the unknown k1<<kMk_1 < \cdots < k_M0 and is practically inestimable; this is a deliberate asymmetry in the method. The tuning parameter k1<<kMk_1 < \cdots < k_M1 is set analytically as a ratio of dispersion measures of the out-of-sample variances to those of the fitting biases, avoiding data-driven cross-validation.

Numerical evidence

Simulations with k1<<kMk_1 < \cdots < k_M2, k1<<kMk_1 < \cdots < k_M3, and k1<<kMk_1 < \cdots < k_M4 ratios approaching one compare LaMA against AIC/BIC selection, smoothed variants, MMA, and JMA using relative in-sample and out-of-sample losses. LaMA attains the lowest out-of-sample loss in essentially all small-sample configurations, with the gap widening as k1<<kMk_1 < \cdots < k_M5 grows; MMA exhibits characteristic overfitting and can lose predictive ability entirely at large k1<<kMk_1 < \cdots < k_M6, while JMA's loss becomes volatile. LaMA's loss curve is also markedly flatter in k1<<kMk_1 < \cdots < k_M7, indicating robustness to low signal-to-noise conditions. As k1<<kMk_1 < \cdots < k_M8 increases, differences among methods narrow, consistent with the recovery of classical asymptotics. On the U.S. Crime (k1<<kMk_1 < \cdots < k_M9) and Motor Trend Car (2\ell_20) datasets with 1000 random train/test splits, LaMA achieves the lowest mean test error and, strikingly, the smallest variance at every training size—for example, at 2\ell_21 on the Crime data its error variance is 0.0293 versus 249.87 for JMA and roughly 2886 for AIC-based methods, a difference of several orders of magnitude. This indicates substantially greater stability under small training samples than any competitor.

Limitations and open questions

Several caveats bear directly on the strength of the results. The exact limiting risk is derived for nested models with uniform or given weights; the optimization theory for the chosen weights (e.g., formal asymptotic optimality of LaMA's estimator of the out-of-sample risk minimizer) is not established, and the authors explicitly identify proving out-of-sample optimality in the 2\ell_22 regime as an open problem, noting that existing optimality results—even recent relaxations such as Peng et al.—remain confined to in-sample risk. The signal condition 2\ell_23 is required for almost sure convergence of the bias component (weaker versions hold in probability), and the general-covariance theorem covers only the under-parameterized case 2\ell_24. The treatment of the singular point 2\ell_25 is necessarily by exclusion or truncation, since the limit genuinely diverges there. Finally, the framework is specific to linear ensembles with squared loss; extension to nonlinear ensembles or dependent data is left unaddressed.

Conclusion

The paper provides a closed-form random-matrix characterization of out-of-sample risk for high-dimensional model averaging, demonstrating that ensembles convert the divergent peaks of individual interpolating models into a bounded ridge through an emergent smoothing mechanism. LaMA operationalizes this insight by correcting the systematic in-sample underestimation of out-of-sample variance and adding adaptive 2\ell_26 regularization, yielding strong empirical predictive accuracy and stability in regimes where 2\ell_27 is comparable to 2\ell_28. Its main theoretical debt—formal optimality guarantees for out-of-sample risk in this regime—remains open.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.