Bounds for approximate retrieval when allowing errors
Develop theoretical bounds on the representation capacity (e.g., embedding dimension or related rank measures) required for single-vector embedding models to succeed when approximate retrieval is acceptable, such as correctly capturing only a majority of the top-k combinations rather than all of them.
References
We also did not show theoretical results for the setting where the user allows some mistakes, e.g. capturing only the majority of the combinations. We leave putting a bound on this scenario to future work and would invite the reader to examine works like \citet{ben2002limitations}.
As a result, the correspondence between a sample’s semantic richness and its allocated capacity can only be examined qualitatively; we provide illustrative cases in Appendix C.2 and leave a principled complexity-aware formulation to future work.
We note down two explicit open problem directions for further investigation. Our proof shows that single vectors fail when we aim to solve the retrieval ordering problem explicitly. In realistic settings, a more relevant question might be to ask a more approximate version of this problem, where Eqn.~\ref{eq:retrievalordering} is satisfied for a $(1-\varepsilon)$ fraction of the entries of each row of the relevance matrix. An interesting future problem would be to check if the exponential gap holds even when our goal is only to solve retrieval order approximately.