Resolution of the VulLibGen reproduction discrepancy

Resolve the discrepancy between the reproduced and author-reported VulLibGen performance under the original dataset and evaluation metric, for which additional hyperparameter tuning did not recover the reported results.

Background

The paper reports that the reproduced VulLibGen results match the original results at top-1 but are lower at top-2 and top-3. Even after additional tuning under VulLibGen’s original dataset and evaluation metric, the best reproduced Avg. F1 was 0.630 rather than the reported 0.682. The authors explicitly state that they were unable to close this gap, leaving the source of the discrepancy unresolved.

References

Despite additional tuning, we were unable to fully close this gap; we report this discrepancy explicitly for transparency.

Athena: Vulnerability-Affected Library Identification via Knowledge Graph Completion  (2609.01187 - Duy et al., 1 Sep 2026) in Appendix, Section “Implementation Details,” paragraph “Baseline reproduction”