Test coverage memory on genuinely long attack chains

Determine whether adding coverage recall to the autonomous PentestGPT framework improves outcomes once earlier attack-surface findings leave the Supervisor’s bounded projection during a sufficiently long penetration-testing task.

Background

The coverage-memory intervention was designed to preserve previously explored attack surfaces and make them available to the autonomous PentestGPT Supervisor. However, none of the three controlled benchmark targets caused the retrieval failure that the intervention was intended to address: repeated discovery was absent, and no finding left the Supervisor’s visible projection. Consequently, the experiments showed no improvement but did not provide a valid test of whether coverage recall helps under genuine long-horizon memory loss.

The authors identify a longer target such as HackTheBox Enigma as necessary because its attack chain may be long enough for early findings to disappear from the Supervisor’s projection. The unresolved issue is whether restoring access to such findings can reduce repeated work or improve exploitation success.

References

They cannot show whether it would help once earlier coverage actually leaves the projection.

Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents  (2609.10780 - Lovelace et al., 9 Sep 2026) in Section 4.2, subsection “Coverage Memory Did Not Improve Outcomes” (Section 5.2); Section 7, “Future Work”

Three legacy stalls cannot establish weak commitment as the general cause of controller convergence, but they do point away from memory capacity as the binding constraint, and the exploratory autonomous runs of Section~\ref{sec:enigma} show the same thing in a harness where stored findings were never lost.

Big Enough to Break Out: Tracking the Rising Capability of LLM Penetration-Testing Agents  (2609.10780 - Lovelace et al., 9 Sep 2026) in Section 5.3, subsection “Failure Patterns in Legacy Stalls”; Section 6, Discussion