Determine whether benchmark-ceiling saturation extends to the full VeriSoftBench benchmark

Determine whether the observed saturation at the benchmark-rule ceiling on the 100-task VeriSoftBench-Aristotle subset extends to all 500 tasks in the full VeriSoftBench benchmark.

Background

The evaluation reports results on a 100-task subset of VeriSoftBench-Aristotle, where several model and harness configurations achieve or approach 100 successful tasks. Because performance reaches the subset ceiling, solve counts may no longer distinguish systems effectively. The paper explicitly leaves unresolved whether this saturation is a property of the subset or would persist across the complete 500-task benchmark. That question matters for assessing the generality of the reported performance and the discriminative value of benchmark success rates.

References

Several model and harness configurations reach the benchmark-rule ceiling on the 100-task subset, limiting the ability of solve counts to distinguish performance. We did not evaluate the full 500-task VeriSoftBench benchmark and therefore cannot assess whether this saturation extends to the complete benchmark.

— FORALL-LEAN-AGENT for Auditable Reasoning in Formal Mathematics and Software Verification  (2610.00885 - Lwin, 1 Oct 2026) in Section 3.1, subsection “VeriSoftBench”