Isolate the effect of the SEF self-check rubric

Determine the independent contribution of the same-context Structured Explanation Framework self-check rubric to GANDR’s strict accuracy by evaluating a control that removes the rubric from the B2 baseline while holding the remaining configuration fixed.

Background

The paper compares GANDR primarily with B2, a single-LLM baseline that uses the CREAC schema and an SEF-style self-check rubric. Because GANDR changes both the system architecture and the prompting configuration relative to B2, the reported accuracy improvement cannot establish how much of the gain is attributable specifically to removing or retaining the rubric. The authors therefore identify an unperformed control experiment that would isolate this factor.

References

That configuration drops B2's same-context self-check rubric, so part of the gain may be the prompt change, and a control isolating the rubric (B2 without it) is left to future work.

GANDR: Claim Auditing for Verifiable Legal Answer Generation  (2609.10293 - Qian et al., 9 Sep 2026) in Section 5, paragraph “Headline” (Results)