Determine whether the observed guidance results establish a weaker-model effect
Determine whether the differing kit-versus-control results observed for GPT-5.6 Luna and DeepSeek-V4.1-Flash establish a capability ordering or a weaker-model effect in specification and proof construction under swapped semantics.
References
The different harness limits comparison with Luna. Model capability ordering is unestablished, so this result leaves the weaker-model question open.
— From Verification Failures to Reusable Guidance for Coding Agents
(2609.39022 - Zhai et al., 30 Sep 2026) in Appendix, Section “Historical comparisons,” final paragraph