Validate correlation between GPT-4 ratings and human judgments for chatbot evaluation
Determine whether GPT-4-based evaluation ratings reliably correlate with human judgments when assessing chatbot performance across tasks and datasets, and quantify the strength and conditions of such correlations.
References
The most plausible reading is that the physicians anchor their judgement in the soundness of the reasoning rather than in the exact position assigned to each model; with only three evaluators, however, genuine convergence cannot be separated from a lack of resolution in the evaluation instrument itself. This dissociation, variable rankings, homogeneous reception, is a finding to be confirmed rather than a settled result.
While recent work indicates generative models can be effectively employed for system evaluations, the reliability GPT-4 ratings to assess chatbot performance is, to our knowledge, yet to be proven to correlate with human judgments.
Our LLM judge has not been validated against human authorship judgments; trait stability is low (Jaccard$=$0.22) and human agreement studies are needed.
Cross-judge triangulation (Grok 4.3, Gemma 4 31B, Mistral Large 3) on a stratified 500-task subset shows GPT-5.1 is the strictest of four frontier judges (pairwise $\kappa$ 0.40--0.61), so the paper's absolute pass rates are conservative measurements; a full 100-task human audit is left to future work.
We have been deliberate about the boundary between what is demonstrated (strong in-distribution retrieval, a deployed multi-strategy system) and what remains to be shown (public-benchmark generalization, judge--human agreement), and we see closing that gap as the natural next step toward substantially reducing the manual stewardship burden of enterprise data governance.