Scaling Athena re-ranking to larger language models

Determine whether fine-tuning larger frontier language models, including 70B-scale open-weight models and closed-source APIs, further improves Athena’s LLM-based re-ranking performance.

Background

The evaluation considers re-rankers with up to 31 billion parameters because larger frontier models exceeded the available memory and compute budget for fine-tuning. Although the authors expect stronger reasoning capabilities might improve re-ranking, this possibility was not evaluated and remains unresolved.

References

Their stronger reasoning capabilities could further improve re-ranking, which we leave to future work.

Athena: Vulnerability-Affected Library Identification via Knowledge Graph Completion  (2609.01187 - Duy et al., 1 Sep 2026) in Limitations, paragraph “LLMs Selected for Re-ranking”