Bounds for approximate retrieval when allowing errors

Develop theoretical bounds on the representation capacity (e.g., embedding dimension or related rank measures) required for single-vector embedding models to succeed when approximate retrieval is acceptable, such as correctly capturing only a majority of the top-k combinations rather than all of them.

Background

The theoretical analysis in the paper focuses on exact representation of the binary query relevance (qrel) matrix, yielding sign-rank-based limits for single-vector models. In practice, retrieval systems may allow some errors or only need to capture most (not all) combinations.

The authors explicitly note they did not provide theory for this approximate setting and call for bounds that quantify the capacity needed when limited errors are acceptable.

References

We also did not show theoretical results for the setting where the user allows some mistakes, e.g. capturing only the majority of the combinations. We leave putting a bound on this scenario to future work and would invite the reader to examine works like \citet{ben2002limitations}.

— On the Theoretical Limitations of Embedding-Based Retrieval  (2508.21038 - Weller et al., 28 Aug 2025) in Limitations

As a result, the correspondence between a sample’s semantic richness and its allocated capacity can only be examined qualitatively; we provide illustrative cases in Appendix C.2 and leave a principled complexity-aware formulation to future work.

— AdaptiveEmbed: Sample-Adaptive Multi-Vector Representation for Multimodal Retrieval  (2608.25412 - Liu et al., 26 Aug 2026) in Section D, Discussion and Limitations

We note down two explicit open problem directions for further investigation. Our proof shows that single vectors fail when we aim to solve the retrieval ordering problem explicitly. In realistic settings, a more relevant question might be to ask a more approximate version of this problem, where Eqn.~\ref{eq:retrievalordering} is satisfied for a $(1-\varepsilon)$ fraction of the entries of each row of the relevance matrix. An interesting future problem would be to check if the exponential gap holds even when our goal is only to solve retrieval order approximately.

— Retrieval Needs Multivectors: An Exponential Separation  (2608.21494 - Agarwal et al., 21 Aug 2026) in Section Discussion and Limitations