Papers
Topics
Authors
Recent
Search
2000 character limit reached

Optimal Data Splitting for Holdout Cross-Validation in Large Covariance Matrix Estimation

Published 19 Mar 2025 in math.ST, q-fin.PM, q-fin.RM, stat.AP, and stat.TH | (2503.15186v1)

Abstract: Cross-validation is a statistical tool that can be used to improve large covariance matrix estimation. Although its efficiency is observed in practical applications, the theoretical reasons behind it remain largely intuitive, with formal proofs currently lacking. To carry on analytical analysis, we focus on the holdout method, a single iteration of cross-validation, rather than the traditional kk-fold approach. We derive a closed-form expression for the estimation error when the population matrix follows a white inverse Wishart distribution, and we observe the optimal train-test split scales as the square root of the matrix dimension. For general population matrices, we connected the error to the variance of eigenvalues distribution, but approximations are necessary. Interestingly, in the high-dimensional asymptotic regime, both the holdout and kk-fold cross-validation methods converge to the optimal estimator when the train-test ratio scales with the square root of the matrix dimension.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 3 tweets with 0 likes about this paper.