Establish generalization beyond English text benchmarks

Establish whether JuryFlow generalizes to languages other than English, additional modalities, and long-form responses beyond the two English benchmarks evaluated in the paper.

Background

The empirical evaluation is restricted to MT-Bench and LLMBar, both English-language benchmarks involving response comparisons. The paper does not test whether claim decomposition, disagreement-graph construction, propagation, or rubric induction remains effective for other languages, modalities, or substantially longer responses. These settings are therefore explicitly unresolved.

References

Our protocol covers two English benchmarks; generalization to other languages, modalities, and long-form responses is untested.

— JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation  (2609.40103 - Yang et al., 30 Sep 2026) in Section 6, Limitations and Future Work, paragraph “Scope and evidence”