Empirical consequences and marginal value of completeness upgrades
Determine when incompleteness in LLM verification produces actual errors and what marginal gain is obtained by upgrading an anchored-correctness verifier from L2 to completeness-oriented L3 verification.
References
The open empirical questions (when does this incompleteness produce actual errors, and what is the marginal L2→L3 gain) are left to future work rather than answered here.
— Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning
(2608.19009 - Yin, 19 Aug 2026) in Section 4, immediately before Section 4.1