Are LLMs accurate and unbiased enough for research evaluation roles?

Establish whether large language models can achieve sufficient accuracy and impartiality, by explicit criteria and empirical testing, to play a reliable role in research evaluation workflows, and delineate acceptable use cases if standards can be met.

Background

The plausibility of LLM outputs, coupled with risks of hidden biases, raises concerns about their deployment in evaluative contexts. Determining their readiness requires systematic benchmarking against confidential expert judgements and stringent bias assessments.

Clear standards for accuracy and fairness are needed to decide whether, and how, LLMs can responsibly support or augment evaluation tasks.

References

It is not clear yet whether LLMs like ChatGPT can be made accurate and unbiased enough to have a role in research evaluation.

A skeptical reading is available, that noise surfaces a bias signal the judge had overlooked rather than inventing one.

When Noise Fabricates Bias: The Fragility of LLM-as-a-Judge Bias Measurement under Noisy Text  (2609.11067 - Ryu et al., 10 Sep 2026) in Limitations, paragraph 'Stability is not correctness'

Although rubric-based evaluation substantially reduces hallucinations, it cannot eliminate them entirely. Due to budget constraints, we use only {GPT-4o-mini} as the evaluator; future work should assess whether stronger models provide more reliable judgments.

Overview of the NLPCC 2026 Shared Task 11: Agent-Based Experiment Reproduction from Scientific Papers  (2609.11117 - Hong et al., 10 Sep 2026) in Limitations section