Establish generalization beyond English text benchmarks
Establish whether JuryFlow generalizes to languages other than English, additional modalities, and long-form responses beyond the two English benchmarks evaluated in the paper.
References
Our protocol covers two English benchmarks; generalization to other languages, modalities, and long-form responses is untested.
— JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation
(2609.40103 - Yang et al., 30 Sep 2026) in Section 6, Limitations and Future Work, paragraph “Scope and evidence”