Design-based variance estimation for machine-learning-assisted survey estimators

Establish valid design-based variance estimation theory for survey estimators assisted by flexible machine-learning predictors.

Background

The paper situates its contribution within the broader difficulty of quantifying uncertainty for model-assisted estimators whose prediction rules are fitted using the probability sample. Flexible learners can adapt closely to the sample, producing in-sample residuals that underestimate prediction errors for nonsampled units and consequently invalidate standard plug-in variance estimators.

The paper addresses a specific instance of this problem by treating a training subsample as a second sampling phase and deriving a variance decomposition for the resulting single-partition estimator. The broader problem remains relevant beyond the particular two-phase construction studied here.

References

Valid variance estimation with machine-learning predictors remains, as \citet{haziza2025} notes, a central open question for survey statisticians, and it is on the variance side that flexible learners create difficulties which the classical theory was not designed to handle.

Interval estimation with design-based coverage under this framework is, by these authors' own account, still open.

The construction is developed in detail for the linear GREG under Poisson sampling, and the authors identify extensions beyond Poisson designs, as well as computationally efficient algorithms for ensembles with external randomness, as open problems.

A forest averages trees grown on resamples of the training set, adding algorithmic randomization to the second phase; design-based variance theory for bagged estimators under complex designs is a long-standing open problem \citep{wang2014}.

Model-assisted estimation with a training subsample: a two-phase sampling approach with design-based variance estimation  (2609.04082 - Riaño, 3 Sep 2026) in Remark 3, “Random forests,” Section 4.3