Demonstrate the conservation floor in an improving real-model cascade

Demonstrate the theoretically predicted conservation floor in a real-language-model corrective cascade by constructing a training procedure that improves the student while the user-facing error asymptotes to a floor bounded by the initial blind-spot mass q₀β₀.

Background

The paper’s real-model experiments use Qwen2.5 students and frontier-model verifiers in a corrective fine-tuning loop. Instead of improving the students, the loop degrades and ultimately collapses them, so the experiments do not realize the regime required to test the paper’s clean conservation-floor prediction on real LLMs.

The theory predicts that, when verifier-detectable errors are corrected but verifier-undetectable errors receive no direct training signal, user-facing error should approach a positive floor related to the initial student error and verifier blind-spot rate. The authors validate this behavior in a synthetic model, but explicitly leave its demonstration with an improving real student unresolved.

References

The floor therefore stays a theoretical result --- derived in Section~\ref{sec:law} and Appendix~\ref{sec:appendix}, validated synthetically in Section~\ref{sec:exp} --- and its real-model demonstration (a stronger student, or a training recipe that resists the hard-tail distribution shift) is open.

Cheap Verifiers, Large Blind Spots: Measuring the Reliability Cost of Cost-Saving Cascades  (2609.01345 - Rajput, 1 Sep 2026) in Section 5, “Real-model measurements”