Papers
Topics
Authors
Recent
Search
2000 character limit reached

Cross Validation for Correlated Data in Regression and Classification Models, with Applications to Deep Learning

Published 20 Feb 2025 in stat.ME | (2502.14808v2)

Abstract: We present a methodology for model evaluation and selection where the sampling mechanism violates the i.i.d. assumption. Our methodology involves a formulation of the bias between the standard Cross-Validation (CV) estimator and the mean generalization error, denoted by $w_{cv}$, and practical data-based procedures to estimate this term. This concept was introduced in the literature only in the context of a linear model with squared error loss as the criterion for prediction performance. Our proposed bias-corrected CV estimator, $\text{CV}c=\text{CV}+w{cv}$, can be applied to any learning model, including deep neural networks, and to a wide class of criteria for prediction performance in regression and classification tasks. We demonstrate the applicability of the proposed methodology in various scenarios where the data contains complex correlation structures (such as clustered and spatial relationships) with synthetic data and real-world datasets, providing evidence that the estimator $\text{CV}_c$ is better than the standard CV estimator. This paper is an expanded version of our published conference paper.

Authors (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Collections

Sign up for free to add this paper to one or more collections.