Resolve the Shared Failure Modes of LLM Judges and Human Annotators

Identify and characterize failure modes that may be shared by GPT-5, Claude Sonnet 4.6, and the human annotators used to synthesize and evaluate PRISM-VLM perturbations and responses.

Background

PRISM-VLM relies on GPT-5 for perturbation generation and grading, with a cross-judge audit using Claude Sonnet 4.6 and additional comparison against human annotations. Although the reported rank correlations and agreement scores are high, the authors acknowledge that agreement does not eliminate the possibility of common systematic errors.

The unresolved problem is to determine whether both LLM judges and the human annotators share blind spots or evaluation biases. Such shared failure modes could remain undetected by cross-judge agreement and would affect the validity of the benchmark’s axis scores.

References

The cross-judge audit with $\mathcal{J}'$ (Claude Sonnet 4.6, overall $\rho{=}\AuditRho{}$) confirms rank stability and shows no Claude-family preference (Claude Haiku 4.5 still receives one of the lowest sc scores, $0.078$, under $\mathcal{J}'$), but cannot rule out failure modes that both judges---or the judges and the human annotators alike---share.

— PRISM-VLM: A Multi-Axis Discriminative Benchmark for Compact Vision-Language Models  (2609.27395 - Park et al., 23 Sep 2026) in Limitations, subsection “Single build-time judge”