Transferability of Calibration Performance to Offensive-Security Decisions

Determine whether the calibration-performance difference reported for Laya and Jev on their respective evaluation sets transfers to offensive-security decision tasks when the models are evaluated in a paired comparison on the same task.

Background

The paper compares published Expected Calibration Error values for Laya and Jev, but emphasizes that the figures were obtained on different tasks and datasets and therefore cannot be treated as directly comparable. Lower calibration error could support more aggressive automation thresholds in penetration-testing adjudication, but the relevance of the published results to security-specific evidence and verdict decisions has not been established. A paired evaluation using the same offensive-security task and dataset is therefore required.

References

Whether this transfer holds for offensive-security decisions remains an open question requiring paired evaluation on the same task (Section~\ref{sec:discussion}).

— Calibrated Decision Models for Autonomous Penetration-Testing Harnesses: JEV and Laya as System One Decision Layers for LLM-Driven Pentest Agents  (2609.28940 - Barbosa, 24 Sep 2026) in Section 9.2, "Measuring calibration: ECE and Brier score"