Are LLMs accurate and unbiased enough for research evaluation roles?
Establish whether large language models can achieve sufficient accuracy and impartiality, by explicit criteria and empirical testing, to play a reliable role in research evaluation workflows, and delineate acceptable use cases if standards can be met.
References
It is not clear yet whether LLMs like ChatGPT can be made accurate and unbiased enough to have a role in research evaluation.
A skeptical reading is available, that noise surfaces a bias signal the judge had overlooked rather than inventing one.
Although rubric-based evaluation substantially reduces hallucinations, it cannot eliminate them entirely. Due to budget constraints, we use only {GPT-4o-mini} as the evaluator; future work should assess whether stronger models provide more reliable judgments.