Two-tier comparison on the Verus hard band

Determine whether the treatment and control arms differ on the 13-task Verus hard band when evaluated at a lower model tier with a sufficiently high call budget to avoid censoring by the registered 40-call cap, and compare that result with the higher-tier evaluation.

Background

The paper planned to compare both arms at claude-sonnet-5 and claude-opus-5 on the same 13 Verus tasks using the same early-stop rule. A three-task probe showed that Sonnet reached the 40-call cap on all three sampled episodes, whereas Opus passed two of them, indicating that a 40-call Sonnet evaluation would partly measure the cap rather than the underlying capability.

One additional Sonnet episode passed only after the cap was increased from 40 to 120 calls, requiring 79 calls. Because this demonstrated that the original lower-tier read would be censored and that a full 13-task evaluation would be substantially more expensive, the planned two-tier comparison was deferred to version 2.

References

The two-tier comparison is therefore open and belongs to version~2.

— SaltBench: A Referee-Gated Protocol for Measuring Method Effects in Machine-Checked Software Work  (2609.11076 - Hickey, 10 Sep 2026) in Section 5, “The treatment question, open,” item 1 (Section \ref{sec:open})