Source of the rubric-performance gap
Determine whether the difference in rubric performance between writing and data-visualization tasks is primarily caused by annotator motivation and expertise.
References
Manually analyzing the human-only rubrics, we conjecture this gap stems mainly from a difference in annotator motivation and expertise, as we hired senior CS PhD students for the writing tasks, but general human workers for data visualization.
— Efficient Test-Time Adaptation through Human-AI Interaction
(2609.04141 - Wang et al., 3 Sep 2026) in Section 4.1, “Rubrics Derived from Human-Agent Interaction Capture More Failures than LLM-Only and Human-Only Rubrics”