Empirical consequences and marginal value of completeness upgrades

Determine when incompleteness in LLM verification produces actual errors and what marginal gain is obtained by upgrading an anchored-correctness verifier from L2 to completeness-oriented L3 verification.

Background

The paper distinguishes correctness verification, which establishes that proposed candidates satisfy a condition, from completeness verification, which establishes that no valid candidates were missed. It argues that substitution- and sampling-based L2 verifiers cannot provide completeness and that an L2-to-L3 transition requires re-encoding the task as a decidable property, such as solution-set equality.

Although the framework formalizes this distinction and illustrates missed-candidate failures, the paper does not establish how often incompleteness causes practically significant errors or whether the additional cost of L3 verification produces a meaningful benefit over L2 verification. These empirical questions are explicitly deferred.

References

The open empirical questions (when does this incompleteness produce actual errors, and what is the marginal L2→L3 gain) are left to future work rather than answered here.

Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning  (2608.19009 - Yin, 19 Aug 2026) in Section 4, immediately before Section 4.1