Separate model capability from harness and system-design effects
Determine whether the observed capability gap between the Kimi K2.5 legacy human-in-the-loop PentestGPT system and the Claude Opus 4.8 autonomous PentestGPT system is caused by model capability, harness design, execution paradigm, autonomy, memory architecture, or their interaction.
References
We cannot say whether the model or the harness accounts for the gap.
— Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents
(2609.10780 - Lovelace et al., 9 Sep 2026) in Section 8, Threats to Validity, paragraph “Confounded cross-setup comparison”; Conclusion