Separate model capability from harness and system-design effects

Determine whether the observed capability gap between the Kimi K2.5 legacy human-in-the-loop PentestGPT system and the Claude Opus 4.8 autonomous PentestGPT system is caused by model capability, harness design, execution paradigm, autonomy, memory architecture, or their interaction.

Background

The comparison contrasts multiple variables simultaneously: Kimi K2.5 versus Claude Opus 4.8, human-mediated command execution versus autonomous execution, different PentestGPT frameworks, and different memory systems. Because no single factor is held constant while another changes, the reported progression is descriptive rather than causal.

Resolving this issue requires evaluating multiple models within one autonomous framework while keeping execution, framework, and memory conditions fixed. This would clarify whether the stronger results arise primarily from the frontier model, the autonomous harness, or the combination.

References

We cannot say whether the model or the harness accounts for the gap.

Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents  (2609.10780 - Lovelace et al., 9 Sep 2026) in Section 8, Threats to Validity, paragraph “Confounded cross-setup comparison”; Conclusion