Determine whether the observed guidance results establish a weaker-model effect

Determine whether the differing kit-versus-control results observed for GPT-5.6 Luna and DeepSeek-V4.1-Flash establish a capability ordering or a weaker-model effect in specification and proof construction under swapped semantics.

Background

The paper compares bare and kit-guided specification construction using GPT-5.6 Luna and DeepSeek-V4.1-Flash under mutated semantics. The archived DSV4-flash comparison produces results that differ from Luna’s, but the models, harnesses, and resource conditions are not sufficiently aligned to establish a capability ordering. Consequently, the paper explicitly leaves open whether the observed pattern reflects a weaker-model effect.

References

The different harness limits comparison with Luna. Model capability ordering is unestablished, so this result leaves the weaker-model question open.

— From Verification Failures to Reusable Guidance for Coding Agents  (2609.39022 - Zhai et al., 30 Sep 2026) in Appendix, Section “Historical comparisons,” final paragraph