Robustness of LLM judges across additional evaluation tasks

Characterize why some LLM-as-judge tasks are more or less robust than others, including aesthetic judgment, code-review correctness, mathematical-reasoning verification, medical-content review, and other expert-evaluation tasks.

Background

The empirical evaluation covers safety, toxicity, AI-generated-text detection, and political-content evaluation. The authors explicitly caution that these domains do not exhaust the applications in which LLMs function as judges.

Different task types may produce qualitatively different wiggle profiles, so broader evaluation is needed to determine whether the observed pressure sensitivity and predictors of instability generalize beyond the six datasets studied.

References

Why some tasks are more or less robust than others is an open question for future work.

Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence  (2608.12645 - Zhao et al., 12 Aug 2026) in Section 5, Limitations, paragraph “Dataset coverage” (also repeated in Appendix, Section “Limitations (Full Discussion)”)

Although many of the underlying concepts (e.g., completeness, relevance, relationship semantics, and abstraction) are common across conceptual modeling languages, we do not have evidence on how the approach generalizes to other modeling notations, DSLs, or different semantic tasks.

Breaking Models to Test the Judge: A Mutation Testing Approach for Semantic Evaluators of Domain Class Diagrams  (2608.14315 - Delcourt et al., 14 Aug 2026) in Section 5, “Limitations and Research Opportunities”