Trade-off Between Rubric Granularity and Automated Judge Reliance
Investigate the trade-off between fine-grained rubric task decomposition and reliance on an LLM-based automated judge for grading PaperBench submissions; specifically, determine how judge reliability affects the level of rubric granularity required to maintain accurate and trustworthy grading outcomes.
References
We further note that the more reliable your judge is, the less fine-grained your rubric's task decomposition needs to be, potentially reducing the effort necessary for task decomposition as judges become more capable. We leave it to future work to study the trade-off between careful specification and delegation to the judge.
However, we must caution that our rubric judge was never validated on imagined tasks, and so future work is warranted to validate this claim.