Source of the rubric-performance gap

Determine whether the difference in rubric performance between writing and data-visualization tasks is primarily caused by annotator motivation and expertise.

Background

The paper compares rubrics generated through human-agent interaction with LLM-only and human-only rubrics. The evolved rubrics capture substantially more failures than human-only rubrics for data visualization, whereas their performance is comparable to human-only rubrics for writing.

The authors attribute this discrepancy to differences in the backgrounds and motivation of the annotators: senior computer-science PhD students annotated the writing tasks, while general human workers annotated the visualization tasks. This explanation is explicitly presented as a conjecture rather than as an experimentally established causal finding.

References

Manually analyzing the human-only rubrics, we conjecture this gap stems mainly from a difference in annotator motivation and expertise, as we hired senior CS PhD students for the writing tasks, but general human workers for data visualization.

Efficient Test-Time Adaptation through Human-AI Interaction  (2609.04141 - Wang et al., 3 Sep 2026) in Section 4.1, “Rubrics Derived from Human-Agent Interaction Capture More Failures than LLM-Only and Human-Only Rubrics”