Reliable grading for four-topic essays

Establish reliable grading for essays containing four simultaneously answered topics by reducing the compounded retrieval, assignment, and grading errors that prevent the GRASP pipeline and graph-augmented retrieval from closing the gap to dependable performance.

Background

The GRASP pipeline evaluates label-free multi-topic science essays with between one and four topics. For four-topic essays, retrieval, assignment, and LLM grading errors compound, producing very low agreement: the reported QWK is 0.02 for GRAG and 0.05 for RAG, compared with an Oracle QWK of 0.50. The authors therefore explicitly characterize four-topic grading as unresolved rather than as a solved case.

References

We view $n = 4$ grading as an open problem rather than a solved case: with only four topics compounding retrieval, assignment, and grading error at once, this setting is the clearest evidence that GRASP, and graph-augmented retrieval more broadly, has not yet closed the gap to reliable grading at the highest topic counts we tested.

GRASP: Graph-Retrieval Automated Scoring Pipeline for Label-Free Multi-Topic Essay Grading  (2609.03857 - Husain et al., 3 Sep 2026) in Section 4, subsection “Grading and false positives”

This pattern is consistent with a ceiling imposed by the LLM grader itself, though we tested only GPT-4.1-mini as the grader and cannot rule out that a stronger or differently-prompted model would narrow this gap; we leave a multi-grader comparison to future work.

GRASP: Graph-Retrieval Automated Scoring Pipeline for Label-Free Multi-Topic Essay Grading  (2609.03857 - Husain et al., 3 Sep 2026) in Section 5, “Conclusion and Future Work”