Empirical consequences and marginal value of completeness upgrades

Determine when incompleteness in LLM verification produces actual errors and what marginal gain is obtained by upgrading an anchored-correctness verifier from L2 to completeness-oriented L3 verification.

Background

The paper distinguishes correctness verification, which establishes that proposed candidates satisfy a condition, from completeness verification, which establishes that no valid candidates were missed. It argues that substitution- and sampling-based L2 verifiers cannot provide completeness and that an L2-to-L3 transition requires re-encoding the task as a decidable property, such as solution-set equality.

Although the framework formalizes this distinction and illustrates missed-candidate failures, the paper does not establish how often incompleteness causes practically significant errors or whether the additional cost of L3 verification produces a meaningful benefit over L2 verification. These empirical questions are explicitly deferred.

References

``The gap binds where task difficulty reaches judge capability'' therefore remains a conjecture with directional weak-judge support only, and it fixes RSR-Bench's design brief: difficulty-stratified, selection-pressure-swept, execution-settled.

— Proof-Carrying Cognition: Closing the Verification Gap with Reality-Settled Reward  (2609.09776 - M et al., 9 Sep 2026) in Section 6.1, “From failures to findings: a pre-registered diagnosis-and-repair program”

The open empirical questions (when does this incompleteness produce actual errors, and what is the marginal L2→L3 gain) are left to future work rather than answered here.

— Grading the Graders: Verification Autonomy Levels (L0-L5) for LLM Reasoning  (2608.19009 - Yin, 19 Aug 2026) in Section 4, immediately before Section 4.1