Evaluate human-guided focal selection

Evaluate the human-guided focal-selection configuration of JuryFlow through a user study comparing human selection with entropy-based selection in terms of accuracy and effort.

Background

JuryFlow identifies high-disagreement claims using verdict entropy and can either select a focal claim automatically or present ranked candidates to a human. All experiments in the paper use the automatic entropy-ranking configuration, so the empirical contribution of a human structural guide remains unmeasured. The unresolved evaluation concerns whether human selection improves accuracy or changes the effort required relative to the entropy proxy.

References

All experiments in this paper use the automatic mode (Section~\ref{sec:setup}), isolating the multi-agent machinery from human factors; a user study of the human-guided mode is left to future work.

— JuryFlow: Disagreement-Guided Human-in-the-Loop Multi-Agent Evaluation  (2609.40103 - Yang et al., 30 Sep 2026) in Section 3, Stage 3: Human-Guided Focal Selection; Section 5, Experimental Setup