Determine Whether a Single Benchmark Can Span All RAG Axes

Determine whether a single uniform benchmark can validly evaluate retrieval-augmented generation across efficiency, robustness, personalization, and multi-hop reasoning, or establish a shared evaluation harness with axis-specific metrics if such unification is infeasible.

Background

Existing RAG benchmarks commonly evaluate separate dimensions, including retrieval quality, generation faithfulness, efficiency, robustness, personalization, and reasoning fidelity. The survey argues that these dimensions rely on different forms of ground truth and are rarely assessed together.

The authors explicitly leave unresolved whether one benchmark can span all four axes. They suggest that a shared evaluation infrastructure with separate metrics may be more realistic, but the central feasibility question remains open.],

quote_indicator_fix_needed? nope

quote/error?

References

Whether a single uniform benchmark can span all four axes remains an open question.

— Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning  (2610.01936 - Sunil et al., 1 Oct 2026) in Section 6.4, “Improved Evaluation Benchmarks”