Finite-sample power and validation resolution limit

Determine the finite-sample resolution limit \(\Delta_{\min}(n,\hat\sigma)\) governing the statistical power of a held-out validation set to detect late-training improvements in gated LLM self-evolution.

Background

The paper identifies statistical power as an open question for validation-gated self-evolving agents. Specifically, it asks how the size of the held-out set and the estimated variability constrain the smallest improvement that can be reliably detected, represented by the resolution limit Δmin⁡(n,σ^)\Delta_{\min}(n,\hat\sigma). This issue matters because finite validation budgets may become the binding constraint when genuine improvements become small late in training.

The paper connects this question to its resolution-law analysis and to the practical problem of distinguishing genuine progress from measurement noise. The stated open issue concerns the power and resolution of finite held-out validation, rather than the existence of the gated self-evolution procedure itself.

References

Two open questions are central here: (i) power, namely the resolution limit $\Delta_{\min}(n,\hat\sigma)$ of a finite held-out set, which is what actually binds late-training improvement (our Proposition~D companion); and (ii) multi-candidate selection bias, since planners proposing 5--8 pipelines per round break the single-candidate assumption, and the Vovk--Wang merge must be invoked (made explicit in our gate).

— A Kinetic Theory of the Gated Self-Evolving LLM Agent  (2610.03243 - Wang, 2 Oct 2026) in Section 2, Related Work, subsection “Statistical gating”