Origin of Errors in RAG: Context Utilization vs. Context Sufficiency

Determine whether errors in Retrieval-Augmented Generation (RAG) systems arise because large language models fail to utilize the retrieved context or because the retrieved context is insufficient to answer the query.

Background

The paper studies Retrieval-Augmented Generation (RAG) systems, where LLMs receive external context at inference time to improve factuality and performance on open-domain question answering. A central uncertainty motivating the work is whether observed errors stem from the model’s inability to leverage the provided context or from the retrieval failing to supply information that is sufficient to answer the query.

To investigate this, the authors introduce the notion of “sufficient context” and develop an autorater to label question–context pairs as sufficient or insufficient without requiring ground-truth answers. This framing enables stratified analyses of model behavior, but the fundamental question of whether errors primarily reflect failures of context utilization versus context insufficiency is explicitly identified as open.

References

Despite much research on Retrieval Augmented Generation (RAG) systems, an open question is whether errors arise because LLMs fail to utilize the context from retrieval or the context itself is insufficient to answer the query.

— Sufficient Context: A New Lens on Retrieval Augmented Generation Systems  (2411.06037 - Joren et al., 2024) in Abstract

The remaining practical question is how much the strict instruction costs when the context is correct. The present apparatus cannot answer it.

— The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs  (2609.10237 - Canedo, 9 Sep 2026) in Section 1, Introduction; Section 7, “A Contrast We Cannot Measure” (Sec. \ref{sec:unmeasurable})

Under a permissive prompt, the question remains open and the observed sign is negative. Resolving it requires a permissive no-context arm that has not been run.

— The Answer Path and the Grounding Instruction in LLM Question Answering over Knowledge Graphs  (2609.10237 - Canedo, 9 Sep 2026) in Section 4, “The Grounding Instruction Suppresses Parametric Recall,” subsection “The retraction is narrower than it first read”

Whether future systems close the retrieval--integration gap at every length at which disclosure is written, or merely move it, is a question this paper's instrument is built to keep answering.

— Reading Is Not Using: Retrieval, Judgment, and the Design of AI Financial Research Workflows  (2608.24842 - Liu et al., 25 Aug 2026) in Section 6, “Conclusion”

Both rules are silent by construction on an agent that calls its tools, reads them successfully, satisfies the predicate and still answers from the wrong content, and in production fidelity is silent on it too. Section~\ref{sec:res-silent} shows that class is real rather than hypothesising it; nothing here separates it from a clean run.

— Fidelity Is Not Enough: Dispatch-Level Instrumentation for Agentic Datasheet Extraction  (2608.28439 - Ye et al., 28 Aug 2026) in Limitations, paragraph “Detection power on degraded runs is unmeasured” (Section~\ref{sec:limitations})

It remains unclear how to best reduce these errors.