Resolve the Shared Failure Modes of LLM Judges and Human Annotators
Identify and characterize failure modes that may be shared by GPT-5, Claude Sonnet 4.6, and the human annotators used to synthesize and evaluate PRISM-VLM perturbations and responses.
References
The cross-judge audit with $\mathcal{J}'$ (Claude Sonnet 4.6, overall $\rho{=}\AuditRho{}$) confirms rank stability and shows no Claude-family preference (Claude Haiku 4.5 still receives one of the lowest sc scores, $0.078$, under $\mathcal{J}'$), but cannot rule out failure modes that both judges---or the judges and the human annotators alike---share.
— PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models
(2609.27395 - Park et al., 23 Sep 2026) in Limitations, subsection “Single build-time judge”