Verifying whether majority voting eliminates imperfect reasoning in self-improving VLM judges
Ascertain whether majority voting–based filtering of synthetic preference pairs for closed-ended tasks in the self-improving vision-language model judge training framework eliminates all imperfect reasoning in the judge’s decisions.
References
While we cannot conclusively verify that majority voting eliminates all imperfect reasoning, the empirical advantages shown in Section~\ref{sec:analysis_reasoning} suggest that consistency-based filtering may provide more robust supervision than correctness checking alone for learning generalizable judgment criteria.
Bahuguna shows self-consistency can backfire on hard questions, with a binned analysis but without an agreement-index decomposition; its v2 revision states that the mechanism of the plurality-agreement gate's failure remains an open problem.