Test whether evidence admissibility explains the performance gap

Investigate whether the larger performance gap on partial-answer questions is caused by the need to determine the admissible evidence set, rather than merely retrieve the best-matching passage, using per-item retrieval traces for the governed DeepKnown system and the hosted Google Gemini File Search service.

Background

The evaluation finds that the hosted service lags behind DeepKnown most substantially on partial-answer questions, where a correct response must distinguish information supported by the corpus from information that remains unsupported. The paper proposes that this difference may result from the governed system’s explicit admissibility test, which can stop generation at the boundary of the available evidence, whereas a system without such a test may complete an answer from retrieved context.

The authors explicitly characterize this explanation as provisional and state that testing it requires per-item retrieval traces. Such traces would allow researchers to determine whether differences in retrieved evidence and admissibility decisions, rather than other differences between the whole systems, account for the observed performance gap.

References

This interpretation remains a hypothesis; testing it would require per-item retrieval traces, which we leave to future work.

— Version- and Scope-Aware Question Answering over Normative Documents: A Deployed System and an End-to-End Evaluation at Production Scale  (2609.18769 - Wang et al., 16 Sep 2026) in Section 5, Results, paragraph “Where the difference appears”