Validate correlation between GPT-4 ratings and human judgments for chatbot evaluation

Determine whether GPT-4-based evaluation ratings reliably correlate with human judgments when assessing chatbot performance across tasks and datasets, and quantify the strength and conditions of such correlations.

Background

Although the paper conducts both GPT-4-based and human evaluations and reports moderate agreement at the system level, the authors explicitly note that the general reliability of GPT-4 ratings to assess chatbot performance has yet to be proven to correlate with human judgments.

This highlights a broader, ongoing question about the validity and robustness of model-based evaluators as proxies for human assessments.

References

The most plausible reading is that the physicians anchor their judgement in the soundness of the reasoning rather than in the exact position assigned to each model; with only three evaluators, however, genuine convergence cannot be separated from a lack of resolution in the evaluation instrument itself. This dissociation, variable rankings, homogeneous reception, is a finding to be confirmed rather than a settled result.

LLMs as Post-hoc Auditors of Physiological Plausibility in Symbolic Regression: A Clinician-Evaluated Case Study  (2609.11431 - López-Varela et al., 10 Sep 2026) in Section 6, Conclusions

While recent work indicates generative models can be effectively employed for system evaluations, the reliability GPT-4 ratings to assess chatbot performance is, to our knowledge, yet to be proven to correlate with human judgments.

QLoRA: Efficient Finetuning of Quantized LLMs  (2305.14314 - Dettmers et al., 2023) in Subsection "Human Evaluation"

Our LLM judge has not been validated against human authorship judgments; trait stability is low (Jaccard$=$0.22) and human agreement studies are needed.

PersonalBench: Measuring the Authorship Gap in LLM Personalization  (2608.19746 - Sawant, 20 Aug 2026) in Section 6, Limitations, paragraph “No human validation and other scope”

Cross-judge triangulation (Grok 4.3, Gemma 4 31B, Mistral Large 3) on a stratified 500-task subset shows GPT-5.1 is the strictest of four frontier judges (pairwise $\kappa$ 0.40--0.61), so the paper's absolute pass rates are conservative measurements; a full 100-task human audit is left to future work.

Compile, Don't Memorize: A Context Compilation Architecture (CCA) for In-Context Learning  (2609.00759 - Qi et al., 1 Sep 2026) in Section 6, paragraph “Single-judge scoring”

We have been deliberate about the boundary between what is demonstrated (strong in-distribution retrieval, a deployed multi-strategy system) and what remains to be shown (public-benchmark generalization, judge--human agreement), and we see closing that gap as the natural next step toward substantially reducing the manual stewardship burden of enterprise data governance.