Data-driven selection of the testing index set

Develop a data-driven choice of the testing index set \(\mathcal K_M\) that optimizes power against specified classes of alternatives while accounting for the dependence between the selected set and the test statistic, for example through sample splitting.

Background

The goodness-of-fit test uses a finite index set KM\mathcal K_M of basis coefficients. The paper establishes uniform Gaussian-approximation and consistency results over deterministic classes of such sets, but notes that selecting KM\mathcal K_M from the same data creates dependence between the selector and the test statistic. Cross-validation is also unsuitable for the testing objective because it targets prediction error and may select a bounded truncation under the null hypothesis.

The unresolved problem is therefore to construct and theoretically justify a data-driven selector for KM\mathcal K_M that achieves favorable detection power for specified alternative classes while preserving valid inference, potentially by using sample splitting.

References

A data-driven \mathcal K_M optimizing power against given classes of alternatives is a delicate problem that we leave for future work.

— Optimal estimation and goodness-of-fit testing of the mean for sparse longitudinal functional data  (2609.19889 - Patilea et al., 17 Sep 2026) in Remark 2.14, Section 4 (Testing the mean function)