Origin of Errors in RAG: Context Utilization vs. Context Sufficiency
Determine whether errors in Retrieval-Augmented Generation (RAG) systems arise because large language models fail to utilize the retrieved context or because the retrieved context is insufficient to answer the query.
References
Despite much research on Retrieval Augmented Generation (RAG) systems, an open question is whether errors arise because LLMs fail to utilize the context from retrieval or the context itself is insufficient to answer the query.
The remaining practical question is how much the strict instruction costs when the context is correct. The present apparatus cannot answer it.
Under a permissive prompt, the question remains open and the observed sign is negative. Resolving it requires a permissive no-context arm that has not been run.
Whether future systems close the retrieval--integration gap at every length at which disclosure is written, or merely move it, is a question this paper's instrument is built to keep answering.
Both rules are silent by construction on an agent that calls its tools, reads them successfully, satisfies the predicate and still answers from the wrong content, and in production fidelity is silent on it too. Section~\ref{sec:res-silent} shows that class is real rather than hypothesising it; nothing here separates it from a clean run.
It remains unclear how to best reduce these errors.