Robustness of LLM judges across additional evaluation tasks
Characterize why some LLM-as-judge tasks are more or less robust than others, including aesthetic judgment, code-review correctness, mathematical-reasoning verification, medical-content review, and other expert-evaluation tasks.
References
Why some tasks are more or less robust than others is an open question for future work.
— Jagged Judges: Epistemic Stability Under Silence, Pressure, and Persistence
(2608.12645 - Zhao et al., 12 Aug 2026) in Section 5, Limitations, paragraph “Dataset coverage” (also repeated in Appendix, Section “Limitations (Full Discussion)”)
Although many of the underlying concepts (e.g., completeness, relevance, relationship semantics, and abstraction) are common across conceptual modeling languages, we do not have evidence on how the approach generalizes to other modeling notations, DSLs, or different semantic tasks.
— Breaking Models to Test the Judge: A Mutation Testing Approach for Semantic Evaluators of Domain Class Diagrams
(2608.14315 - Delcourt et al., 14 Aug 2026) in Section 5, “Limitations and Research Opportunities”