Reliable grading for four-topic essays
Establish reliable grading for essays containing four simultaneously answered topics by reducing the compounded retrieval, assignment, and grading errors that prevent the GRASP pipeline and graph-augmented retrieval from closing the gap to dependable performance.
References
We view $n = 4$ grading as an open problem rather than a solved case: with only four topics compounding retrieval, assignment, and grading error at once, this setting is the clearest evidence that GRASP, and graph-augmented retrieval more broadly, has not yet closed the gap to reliable grading at the highest topic counts we tested.
This pattern is consistent with a ceiling imposed by the LLM grader itself, though we tested only GPT-4.1-mini as the grader and cannot rule out that a stronger or differently-prompted model would narrow this gap; we leave a multi-grader comparison to future work.