Pilot-free adaptive learning of the variance-optimal allocation

Develop an audit for certifying the entire best-of-k accuracy curve on a fixed benchmark that learns the variance-optimal allocation of generated answers across questions as it proceeds, without a separate pilot, and pays for allocation learning only to the extent required by the information it uses.

Background

For a fixed benchmark, the paper defines Gamma as the smallest worst-budget variance per generated answer under the best allocation of answers across questions. Theorem 3 shows that every valid audit must asymptotically pay a cost proportional to Gamma divided by the squared target precision, while a pilot-based audit can attain this rate up to logarithmic factors.

The pilot-based procedure estimates question-specific variances, uses them to construct an allocation, and then collects fresh blocks of answers. However, the paper notes that at the precisions used in its experiments, the pilot alone can cost more answers than the complete paired audit. The unresolved problem is therefore to learn the advantageous allocation online while avoiding the cost of a separately completed pilot and charging only for the information actually needed during the audit.

References

At the precisions of Section~\ref{sec:exp} its pilot costs more than the paired audit's whole bill, and Proposition~\ref{prop:priced} rules out a free version, so the open problem is an audit that learns the allocation as it goes, without a separate pilot, and pays for learning only what it uses.

— Cheap to Draw, Expensive to Trust: Certifying Test-Time Scaling Curves  (2609.40190 - Sohail et al., 30 Sep 2026) in Discussion section

The horizon loss is a tractable per-example approximation to the coupled problem, and learning the coupled allocation remains open.

— Planning to Learn  (2610.03667 - Osband, 2 Oct 2026) in Appendix, Section Planning (immediately after Proposition 2, “Limit of separable allocation”)