Effectiveness of diverse retriever ensembles

Determine whether ensembles of diverse retrievers recover more relevant documents than equally sized ensembles composed of stronger but more similar retrievers, using the multiple labeled positives per query in Q2D-Web.

Background

Q2D-Web contains many labeled relevant documents for each agent-reformulated query, allowing evaluation not only of aggregate recall but also of the unique positives contributed by individual retrievers. The paper observes that BM25 has the lowest aggregate recall among the evaluated systems yet contributes the largest set of relevant documents not recovered by any other retriever, suggesting that retriever diversity may improve ensemble coverage.

The authors explicitly leave unresolved whether this diversity advantage translates into better ensemble performance than that of equally sized ensembles containing only stronger but more similar neural retrievers. The proposed comparison would hold the ensemble size, fusion rule, and candidate budget constant while testing whether diverse retriever combinations recover more relevant documents.

References

Future work can use this to test whether diverse ensembles recover more relevant documents than equally sized ensembles of stronger but more similar retrievers.

— Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems  (2609.08887 - Schall et al., 8 Sep 2026) in Section 6, Conclusion