Papers
Topics
Authors
Recent
Search
2000 character limit reached

Sharp Limits for Honest Uncertainty in Hard-Budget Repeated Evaluation

Published 24 Sep 2026 in cs.AI and cs.IT | (2609.29140v1)

Abstract: Repeated evaluation can estimate a benchmark score accurately while still requiring replication to certify narrow uncertainty. We characterize that requirement on a fixed grid of MM tasks with LL binary paths per task under the hard budget (M+t)K(M+t)K, where each path costs at most KK responses or episodes. For fixed L≥3L \ge 3 and $0 &lt; α\le 1/12$, the optimal expected width on the worst pure cohort is Θ<em>α,L([M(t+1)]<sup>−1/2)Θ<em>{α,L}([M(t+1)]<sup>{-1/2}) when every task is observed and Θ</em>α,L([M(t+M)]<sup>−1/2)Θ</em>{α,L}([M(t+\sqrt{M})]<sup>{-1/2}) when omission is allowed. The lower bounds cover adaptive hard-budget policies, and fixed random-subset designs attain both rates through disagreement certificates. A joint mean/disagreement interval turns the task-covering law into practical finite-budget inference. In an equal-budget LiveCodeBench replay with 16 models, 880 tasks, and five outputs per task, the task-covering design reduces median point-estimation MSE by 87.0\% relative to pooled uniform sampling, while the Joint certificate produces narrower confidence intervals in 15/16 panels and reduces median interval width by 30.6\%. Finite-regime analyses identify task coverage as the effective choice at the evaluated scale and characterize how cohort size and within-task agreement determine the useful operating region. Together, the sharp laws and fixed-budget evidence make replication and task coverage explicit design variables for information-efficient repeated evaluation.

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Continue Learning

We haven't generated follow-up questions for this paper yet.