Validate diff against human judgments of repair quality

Establish whether the change-aware diff score correlates with human assessments of vulnerability-repair quality across genuine repairs, partial repairs, and non-repairs.

Background

The paper introduces diff as a textual edit-overlap score that evaluates only the region changed by a generated patch, rather than the whole function. The score reliably assigns zero to unchanged copies of vulnerable inputs and assigns near-zero scores to some deletion-based gaming patches, but it can also give substantial credit to semantically incorrect patches that overlap the developer’s edit.

No human-annotation study was conducted to determine whether diff scores correspond to human judgments of repair quality. The authors therefore identify the relationship between diff and human assessment as unresolved and specifically propose an annotation study covering genuine, partial, and non-repairs.

References

We also did not validate diff against human judgment: no annotation study correlating diff scores with manual ratings of repair quality was conducted, so its agreement with human assessment is currently unestablished.

— Metrics Failure in LLM-Based Code Vulnerability Repair: An Empirical Study and a Change-Aware Screen  (2609.26749 - Nepal et al., 22 Sep 2026) in Section 6, Limitations and Future Work, paragraphs “Construct validity” and “Future work”