Autocurriculum under imperfect verification
Develop a theoretical framework for verifier-guided autocurriculum when the outcome verifier is imperfect, and establish guarantees for settings with noisy or learned reward models in place of the perfect outcome verifier assumed in the current analysis.
References
Our framework assumes access to a perfect outcome verifier, which is natural for domains with verifiable rewards (math, code), but extending the theory to noisy or learned reward models is an important open problem.
Conjecture (Blind-spot conservation). Under the corrective loop with a fixed, imperfect verifier, and absent generalization spillover into the blind region, the accepted-and-wrong error mass q_{t}\beta_{t} is conserved across rounds even as the detectable mass q_{t}r_{t} is driven down, so the user-facing error does not vanish but asymptotes to \varepsilon_{\infty} \approx q_{0}\beta_{0. Spillover relaxes this equality to an upper anchor, \varepsilon_{\infty} \lesssim q_{0}\beta_{0}, when \beta_{0} is large relative to the student's capacity floor.