Standards-recall threshold after retrieval augmentation or domain fine-tuning

Determine whether retrieval-augmented generation and domain fine-tuning can improve lightweight language models’ 3GPP and O-RAN specification recall sufficiently to meet the standards-recall threshold required for autonomous OAM operation.

Background

The evaluation finds that zero-shot recall of 3GPP and O-RAN specifications is the principal weakness of Claude-Haiku-4.5, GPT-5.4-Mini, and Gemini-3.1-Flash-Lite, with every model scoring below 60% on the TeleQNA ORAN FT benchmark. The authors therefore identify retrieval-augmented generation and domain fine-tuning as potential remedies, but leave unresolved whether these interventions can achieve the level of standards knowledge needed for autonomous operations, administration, and maintenance (OAM).

References

Future work will focus on three directions. First, incorporating human expert judgment alongside the automated judge panel to assess how well LLMs align with domain specialist evaluation in a telecom context, providing external validation of the LLM-as-Judge methodology. Second, addressing the specification knowledge gap identified on TeleQNA ORAN FT through retrieval-augmented generation and domain fine-tuning, and assessing whether the resulting models meet the standards-recall threshold needed for autonomous OAM operation.

Free-Text Evaluation of LLMs for 5G Domain Knowledge and Fault Analysis using LLM-as-Judge  (2608.21021 - Sengupta et al., 21 Aug 2026) in Section V, Conclusions, arXiv PDF p. 5