Reliable Evaluation of Large Language Models

Develop reliable evaluation methodologies for large language models that effectively assess model performance, including helpfulness and harmlessness, addressing the acknowledged unresolved problem of evaluating such models.

Background

The paper evaluates Safe RLHF across three iterations and notes that assessing LLMs remains difficult. To proceed despite this difficulty, the authors employ two practical methods: fast model-based evaluation using unified reward and cost models trained on balanced preference data, and Elo-style pairwise comparisons of model outputs judged by GPT-4 and humans.

They further construct their own evaluation prompt dataset due to gaps in existing benchmarks, underscoring that current evaluation standards for alignment (helpfulness and harmlessness) are inconsistent and costly when relying solely on human judgments. This context motivates the need for more reliable, comprehensive evaluation approaches.

References

However, evaluating LLMs has consistently been a challenging and unresolved problem.

— Safe RLHF: Safe Reinforcement Learning from Human Feedback  (2310.12773 - Dai et al., 2023) in Section 4.1, Helpfulness and Harmlessness Evaluation

Second, declared evidence markers give verifiable ground truth for an engagement whose author can state them, and open-ended harm still depends on a model's judgement.

— ACEA: An Adversarial Co-Evolution Arena for Head-to-Head Red-Team and Blue-Team LLM Testing  (2609.08256 - Shen et al., 8 Sep 2026) in Conclusion

Some benchmarks include human studies, but these remain contentious, as designing sound human-model evaluations is an open question and a direction for future work.

— Not Another Text Benchmark: Putting the "Visual" Back in Visual Question Answering for Large Video Models  (2609.17112 - Chakraborty et al., 15 Sep 2026) in Section 4, subsection “Human baseline”

The implementation running on a device, with its leakage, faults, and countermeasures, remains open.

— CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices  (2609.21344 - Zhou et al., 18 Sep 2026) in Section 2, subsection “LLM Benchmarks in Cryptography” (Related Work)

The cryptographic module inside the device remains open as an object of evaluation.

— CESBench: Benchmarking Large Language Models on Cryptographic Engineering Security for IoT Devices  (2609.21344 - Zhou et al., 18 Sep 2026) in Section 2, subsection “LLM Benchmarks in Cybersecurity and IoT” (Related Work)

However, open questions remain regarding stochastic variability, error behavior, and cross-model performance.

— Two Truths and A Lie? Benchmarking Off-the-Shelf LLMs for Requirements Quality Assessment: Performance, False Alarms, and Misses  (2609.03230 - Shefa et al., 3 Sep 2026) in Section 2.2, AI for Requirements Quality Assessment, p. 3

As a result, despite growing interest in deployment, robust evaluation of clinical information retrieval remains an open challenge [16, 17].

— A Living Benchmark for Information Retrieval from Electronic Health Records  (2609.30205 - Cahoon et al., 24 Sep 2026) in Section 1, Introduction

Its reliability for assessing LLM outputs remains to be established.