Determine Whether a Single Benchmark Can Span All RAG Axes
Determine whether a single uniform benchmark can validly evaluate retrieval-augmented generation across efficiency, robustness, personalization, and multi-hop reasoning, or establish a shared evaluation harness with axis-specific metrics if such unification is infeasible.
References
Whether a single uniform benchmark can span all four axes remains an open question.
— Mapping the RAG Landscape: A Four Axis Taxonomy of Efficiency, Defense, Interactivity, and Reasoning
(2610.01936 - Sunil et al., 1 Oct 2026) in Section 6.4, “Improved Evaluation Benchmarks”