Papers
Topics
Authors
Recent
Search
2000 character limit reached

A Systematic Multi-Domain Evaluation of Document Retrievers

Published 24 Sep 2026 in cs.IR | (2609.29455v1)

Abstract: Document retrieval is a crucial component of many modern AI systems, directly influencing their effectiveness, robustness, and fairness in downstream tasks. While recent years have seen a growing number of retrievers, comparative studies in the literature are typically limited in scope or focused on singular benchmarks, domains, or model families. This fragmentation makes it difficult to draw reliable conclusions about the relative strengths, weaknesses, and trade-offs of document retrievers. To address this gap, we conduct a large-scale empirical evaluation of document retrievers, covering three families (sparse, dense, and expansion-based) and evaluating 33 retrievers across seven IR datasets, analyzing retrieval quality, runtime, and failure points. Rather than tuning each model individually, we evaluate every retriever off the shelf, under the configuration reconstructable from its public documentation and a uniform compute budget. Our results show that NV-Embed-v2 achieves the strongest performance on four of the seven datasets, albeit at the cost of substantial query latencies. Among sparse retrievers, we find that SPLADE-v3 rivals the top-performing approach despite much lower latency, and even achieves top scores on MS MARCO. On instruction-following datasets, GritLM delivers the best performance. Finally, an analysis of the retrievers' failure points reveals contrasts between models and families that indicate potential for unrealized gains in retrieval performance.

Authors (2)

Summary

No one has generated a summary of this paper yet.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is the paper about?

This paper compares 33 document-retrieval systems. A document retriever is a tool that searches through many documents and chooses the ones that seem most useful for a question.

For example, when you search for “How do volcanoes erupt?”, a retriever decides which web pages or passages should be shown first. Retrievers are also used by AI chatbots, recommendation systems, and retrieval-augmented generation (RAG) systems, where an AI looks up information before answering.

The researchers wanted to find out:

  • Which retrievers find the most useful documents?
  • Which ones work well across different subjects?
  • How fast are they?
  • Do different retrievers fail on the same questions?
  • Is a very accurate retriever always the best choice, or might a faster one be more practical?

2. What questions did the researchers ask?

The paper focuses on several main questions:

  1. Which retrieval methods perform best overall?
  2. Do some methods work better for particular types of questions or subjects?
  3. How does retrieval quality compare with search speed?
  4. Are modern, large AI models always better than older or smaller systems?
  5. How reliable are results reported in earlier research papers?
  6. Which questions are difficult for retrievers, and do different systems struggle with the same questions?

The researchers compared three main families of retrievers:

  • Sparse retrievers: Look for important words shared by the question and a document. BM25 is a classic example.
  • Dense retrievers: Turn questions and documents into lists of numbers, called vectors. Texts with similar meanings should have similar vectors, even if they use different words.
  • Expansion-based retrievers: Add extra words or even imaginary documents to a question or document to improve the search.

A simple analogy is to imagine three librarians:

  • One searches for matching words in book titles.
  • One understands the general meaning of the question.
  • One rewrites the question with extra details before searching.

3. How was the research carried out?

Comparing 33 retrievers

The researchers tested 33 systems, including:

  • Traditional methods such as BM25
  • Learned sparse methods such as SPLADE-v3
  • Dense systems such as GTE, BGE, GritLM, and NV-Embed-v2
  • Expansion-based systems such as docT5query, query2doc, and HyDE

They used the models in a mostly “ready-to-use” form. In other words, they did not spend extra time specially adjusting every system for every dataset. This created a more equal comparison and better represented what an ordinary user might achieve by following the model’s public instructions.

Testing on seven types of evaluation

The systems were tested on seven benchmark settings, including:

  • MS MARCO: Ordinary web-search questions
  • TREC-DL: Carefully judged search questions
  • BEIR: Questions from many different subjects and tasks
  • LoTTE: Rare and specialized questions, such as those found on StackExchange
  • InstructIR and FollowIR: Questions that include detailed instructions about what kind of answer or document is wanted

These datasets are like different school exams. A system might do well on a general science test but struggle on a difficult history or computer-programming test.

Measuring quality and speed

The researchers used several scores to measure how well the retrievers worked. These scores mainly checked whether relevant documents appeared near the top of the results.

For example:

  • Precision: How many of the first few results were useful?
  • Recall: How many of all the useful documents did the system find?
  • MRR: How high was the first useful result?
  • nDCG: How well were all the useful documents arranged, giving more credit to higher-ranked results?
  • Success@5: Did at least one useful document appear in the first five results?

They also measured query latency, which means how long a system takes to answer a search request. This matters because a system that is slightly more accurate but takes much longer may not be practical for a real search engine or chatbot.

The experiments were run using powerful computers with four NVIDIA H100 graphics cards.

4. What did the researchers find?

The strongest systems were not always the same

The results showed that no single retriever was best at everything.

NV-Embed-v2 was the strongest system on four of the seven evaluation settings, including:

  • TREC-DL 2019
  • TREC-DL 2020
  • BEIR
  • LoTTE

However, it was also relatively slow. This shows an important trade-off: higher accuracy can require more computing time.

GritLM performed best on the instruction-following datasets, InstructIR and FollowIR. This makes sense because GritLM was designed to understand instructions along with the text.

SPLADE-v3 was a strong and faster competitor

Among sparse retrievers, SPLADE-v3 performed especially well. It achieved the best result on MS MARCO and was close to the strongest systems on several other datasets.

This is important because sparse systems are often much faster than very large dense models. They can use traditional search indexes, similar to the indexes used by web search engines.

The finding suggests that a large, modern LLM is not always necessary. A smaller or more efficient system may provide nearly as much value while responding faster.

Dense retrievers usually worked well across many subjects

Dense retrievers generally performed better than sparse and expansion-based systems on:

  • Specialized or rare questions in LoTTE
  • Instruction-following tasks
  • Many general and out-of-domain tasks

This is probably because dense systems focus on meaning, not just exact word matches. For example, they may understand that “car accident” and “vehicle collision” are related even though they use different words.

Instruction-following models were sensitive to wording

The researchers found that systems using instructions could change their results significantly depending on the exact wording of the instruction.

For example, a retriever might perform well when asked:

“Find a document that explains this topic for beginners”

but perform differently when asked:

“Find the most technically detailed document about this topic.”

This means that prompt design—the wording given to an AI model—can have a major effect on performance. However, many research papers do not clearly explain how their prompts were chosen, making comparisons less fair.

Published results were not always easy to reproduce

About 70% of the previously published results were close to the results found in this study. That is encouraging, but about 30% showed noticeable differences.

The researchers suggest that these differences may be caused by:

  • Different versions of the datasets
  • Including or leaving out document titles
  • Different software or search indexes
  • Different model checkpoints
  • Different instructions or prompts
  • Different ways of preparing the text
  • Different settings used when building search indexes

This shows why research comparisons need carefully described methods. Two studies can use the same model but still get different results if they prepare the data differently.

Different systems failed on different questions

The researchers also compared the questions that each retriever answered poorly. They found differences between models and model families.

This is important because it suggests that combining different retrievers could improve results. If one system is bad at questions involving exact words but another is bad at questions requiring deeper meaning, using both might cover each other’s weaknesses.

5. Why are these findings important?

The paper shows that choosing a retriever is not simply a matter of picking the model with the highest score.

The best choice depends on what matters most:

Need Possible choice
Highest general retrieval quality NV-Embed-v2
Following detailed instructions GritLM
Strong performance with lower delay SPLADE-v3
Traditional, simple, and efficient search BM25
Searching rare or specialized topics A strong dense retriever

The study also gives practical advice to developers building search engines, chatbots, and RAG systems. They should consider:

  • Accuracy
  • Response speed
  • Computer and financial costs
  • The type of documents being searched
  • The kinds of questions users ask
  • How sensitive the model is to wording
  • Whether the model works reliably outside its training domain

Conclusion

This paper is a large and careful comparison of document-retrieval systems. It finds that modern dense retrievers, especially NV-Embed-v2 and GritLM, often achieve the highest scores. However, SPLADE-v3 shows that a faster sparse retriever can still compete strongly, particularly on common web-search tasks.

The main lesson is that there is no single perfect retriever. A system that is best for one task may be slower or less reliable for another. The research could help engineers choose better search tools and build more accurate, faster, and more dependable AI systems. It also encourages researchers to use common testing methods so that results from different studies can be compared fairly.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Limited language coverage: The evaluation is restricted to English datasets, leaving the effectiveness, latency, and failure behavior of sparse, dense, and expansion-based retrievers in multilingual and low-resource languages unresolved.
  • Narrow domain representation: Although the study spans several benchmarks, most datasets consist of web, academic, forum, news, or TREC-style text; performance on enterprise documents, legal and medical records, technical manuals, social media, multimodal documents, and highly specialized corpora is not established.
  • Unclear real-world representativeness: The selected benchmark queries and relevance judgments may not reflect production search traffic, conversational queries, noisy user input, evolving information needs, or adversarial search behavior.
  • Dependence on benchmark artifacts: The study does not determine how much of the reported performance is attributable to training-data overlap, benchmark contamination, or memorization, particularly for large decoder-only models evaluated on long-established datasets.
  • Off-the-shelf evaluation conflates model quality with configuration quality: Because models are not individually tuned, observed differences may reflect prompt templates, pooling choices, truncation limits, preprocessing, index parameters, or checkpoint selection rather than intrinsic retriever capability.
  • Prompt sensitivity is identified but not systematically quantified: The paper notes that prompt wording can substantially affect instruction-following retrievers, but it does not provide a controlled prompt-sensitivity analysis, prompt robustness intervals, or principled guidance for selecting prompts.
  • A single unified prompt may disadvantage some models: Applying one prompt across datasets does not establish whether each retriever was evaluated under an equivalent or optimal task formulation, especially for models whose authors prescribe dataset-specific instructions.
  • Model selection may introduce benchmark bias: Inclusion based partly on citation counts, leaderboard competitiveness, and “best-performing” variants may underrepresent less popular, newer, smaller, proprietary, or non-English retrievers and may favor models already optimized for the selected benchmarks.
  • The retriever set is not exhaustive: Hybrid sparse–dense systems, learned hybrid models, multi-vector approaches beyond ColBERT, generative retrieval, modern rerankers, pruning/compression variants, and API-based embedding services are not comprehensively compared.
  • Expansion-based methods are not evaluated under equivalent generator conditions: Query2doc and HyDE use Qwen2.5-7B-Instruct rather than the generators used in their original studies, making it unclear whether their results measure the expansion methods or the replacement generator.
  • Expansion costs are incompletely characterized: The evaluation does not fully separate or report generation latency, token usage, generation cost, caching effects, document-expansion storage, and indexing overhead from retrieval latency.
  • End-to-end retrieval pipelines are not evaluated: The study focuses on first-stage retrieval and does not establish how the rankings affect reranking, retrieval-augmented generation, question answering, citation correctness, factuality, or downstream user utility.
  • No hybrid or ensemble analysis is provided: The complementary strengths suggested by the failure-overlap analysis are not tested through sparse–dense fusion, multi-stage retrieval, reciprocal-rank fusion, or adaptive model selection.
  • Failure-point analysis is descriptive rather than explanatory: Overlap between queries missed by different retrievers does not establish why failures occur; linguistic phenomena, document properties, query intent, entity ambiguity, and relevance-label issues are not separately diagnosed.
  • The potential gains from combining retrievers remain unquantified: The paper does not measure whether combining models’ candidate sets or rankings can recover missed relevant documents, nor does it identify which model combinations provide the best effectiveness–latency trade-off.
  • Robustness to query variation is only partially examined: Instruction changes are considered through InstructIR and FollowIR, but robustness to spelling errors, paraphrases, transliteration, colloquialisms, length variation, ambiguity, typos, and adversarially rewritten queries is not systematically evaluated.
  • Fairness and demographic bias are not measured: Although the introduction highlights fairness as an important downstream concern, the experiments do not assess disparate retrieval quality, representation bias, or sensitivity to demographic and identity-related terms.
  • Temporal robustness is unresolved: The study does not evaluate how retrievers perform as corpora, terminology, facts, and user interests change over time, nor whether models trained on historical data degrade on temporally newer documents.
  • Long-document behavior is underexplored: The benchmarks primarily use passages or relatively short textual units; the effects of document length, chunking strategy, context truncation, and retrieval from full-length documents remain unclear.
  • Indexing and storage trade-offs are incompletely reported: The paper compares query latency but does not provide a comprehensive analysis of index construction time, index size, embedding storage, sparsity, memory consumption, update cost, and infrastructure requirements.
  • Latency results may not generalize beyond the experimental hardware: Measurements obtained on four H100 GPUs and CPU-based FAISS/Lucene indexes do not show how rankings change on commodity CPUs, consumer GPUs, cloud instances, mobile hardware, or distributed production systems.
  • Throughput and concurrency are not evaluated: Query latency is reported without a systematic analysis of throughput, tail latency, batching, concurrent users, queueing, or service-level constraints.
  • Approximate nearest-neighbor configurations may affect comparisons: The paper does not fully establish whether dense retrievers were evaluated with equivalent recall–speed settings, index types, search parameters, and quantization schemes.
  • The compute-budget definition is insufficiently operationalized: A uniform compute budget is stated, but the paper does not clearly define whether it equalizes training provenance, indexing computation, inference FLOPs, GPU time, memory, or total monetary cost.
  • Statistical uncertainty across models is not comprehensively analyzed: Bootstrap standard errors are used for some comparisons, but the paper does not provide systematic pairwise significance testing, multiple-comparison correction, confidence intervals for all metrics, or analysis of variance across query subsets.
  • Dataset relevance judgments may be incomplete or inconsistent: The study does not examine pooling depth, false negatives, annotation disagreement, assessor bias, or how incomplete qrels affect comparisons between retrievers that retrieve different types of relevant documents.
  • Metric comparability across datasets is limited: Results rely on dataset-specific metrics and cutoffs, making cross-dataset aggregation and claims about overall superiority difficult; the implications of alternative cutoffs and utility-weighted metrics are not explored.
  • The choice of one checkpoint per model family leaves scaling questions unanswered: The study does not disentangle the effects of parameter count, embedding dimensionality, training data, architecture, and instruction tuning, so it cannot establish scaling laws or identify the source of performance gains.
  • Quantization and compression are not systematically assessed: It remains unknown whether high-performing large models can retain their advantages under reduced precision, vector compression, pruning, distillation, or memory-constrained deployment.
  • Training-data and fine-tuning effects are not controlled: Comparisons mix pretrained, instruction-tuned, retrieval-finetuned, and potentially benchmark-specific models, preventing causal conclusions about the value of particular training objectives or data mixtures.
  • Reproducibility remains incomplete for unavailable or altered components: Differences in checkpoints, preprocessing, hyperparameters, and implementation details are documented as possible causes of score discrepancies, but the paper does not provide a controlled ablation that identifies the contribution of each factor.
  • The impact of preprocessing choices is not isolated: Title inclusion, casing, tokenization, truncation, duplicate handling, passage formatting, and query formatting are recognized as potential sources of variation but are not evaluated through systematic preprocessing ablations.
  • Dynamic and incremental corpora are not considered: The experiments use static indexes and do not assess update latency, embedding regeneration, index maintenance, deleted documents, or retrieval consistency in continuously changing collections.
  • Document-level safety and reliability are not evaluated: The study does not investigate susceptibility to adversarial documents, prompt injection content, spam, duplicated content, poisoned corpora, or misleading but lexically/semantically similar passages.
  • The relationship between retrieval quality and downstream usefulness remains open: Higher nDCG, MRR, or Success scores are assumed to imply better application performance, but this relationship is not validated in realistic downstream systems or human evaluations.
  • No principled model-selection rule is derived: The results reveal effectiveness–latency trade-offs, but the paper does not propose a deployment-oriented decision framework that incorporates workload, memory, cost, latency targets, robustness, and downstream objectives.

Practical Applications

Immediate Applications

  • Retriever selection for production search and RAG systems (software, enterprise search, customer support)
    • NV-Embed-v2 is a strong candidate when retrieval quality is the primary objective and higher query latency is acceptable.
    • SPLADE-v3 is suitable when low latency, inverted-index compatibility, and strong performance on web-search-like data are important.
    • GritLM is particularly appropriate for systems whose queries contain explicit instructions or constraints.

Dependencies: Results are based on English benchmarks, off-the-shelf configurations, and a particular hardware and software stack. Production teams should validate performance on their own corpus, query distribution, latency target, and cost envelope.

  • Latency–quality benchmarking during retrieval-system procurement (software infrastructure, cloud services)
    • nDCG or MRR for ranking quality,
    • recall or Success@k for downstream answerability,
    • p95/p99 query latency,
    • GPU/CPU and memory requirements,
    • index size and refresh cost.

Dependencies: Comparable measurements require fixed preprocessing, model prompts, index settings, batch sizes, and hardware.

  • Cost-aware architecture selection for search applications (web search, e-commerce, media, internal knowledge bases) Sparse retrieval—especially SPLADE-v3—can serve as a relatively efficient first-stage retriever, while high-performing dense models can be reserved for applications where semantic matching or long-tail retrieval justifies additional infrastructure. This supports architectures such as:
    1
    
    query → sparse candidate retrieval → optional dense reranking → application or LLM
    The approach is actionable for product search, document discovery, FAQ search, and enterprise portals.

Dependencies: Sparse indexes may become large, and dense models may require GPUs or optimized ANN infrastructure. The relevant trade-off depends on corpus size, update frequency, and traffic volume.

  • Instruction-sensitive retrieval for controlled search workflows (legal, compliance, research, customer service) GritLM and comparable instruction-aware retrievers can be evaluated for queries such as “retrieve only documents published after 2022” or “prioritize safety-related evidence.” The study’s use of InstructIR and FollowIR indicates that instruction changes can materially affect rankings. A deployable workflow is to store the user’s retrieval instruction separately from the topical query and test whether the selected model reliably respects both.

Dependencies: Instruction-following benchmark gains do not guarantee reliable compliance with hard constraints. Metadata filters and symbolic query processing should remain responsible for requirements such as date, jurisdiction, access level, or document type.

  • Hybrid retrieval for robustness and coverage (enterprise search, RAG, digital libraries)
    • sparse retrieval for exact names, identifiers, product codes, citations, and rare terminology;
    • dense retrieval for paraphrases, conceptually related text, and long-tail queries.

The paper’s failure-point analysis provides a basis for identifying queries where one family fails and another succeeds. Candidate lists can be merged using reciprocal-rank fusion or a learned reranker.

Dependencies: Combining systems increases indexing, storage, and operational complexity. The benefit must be measured on query-level complementarity rather than assumed from aggregate scores.

  • Evaluation and reproducibility audits for academic and industrial IR research (academia, ML engineering)
    • exact checkpoints,
    • document-title and preprocessing treatment,
    • prompt templates,
    • index construction parameters,
    • hardware and latency,
    • statistical uncertainty,
    • differences from previously published results.

This is immediately useful for papers, model cards, and internal model reviews.

Dependencies: Public weights, licenses, datasets, and software versions must remain available. Reproduction may still vary with GPU type, library versions, and undocumented implementation details.

  • Failure-case monitoring in deployed retrieval systems (quality assurance, safety, customer support)
    • adding aliases or terminology to the index,
    • correcting document segmentation,
    • adding metadata filters,
    • expanding queries,
    • fine-tuning on missed examples,
    • routing difficult queries to a second retriever.

Dependencies: Query logging must respect privacy, data-protection, and retention requirements. Relevance labels or expert review are needed to distinguish retrieval failures from genuinely unanswerable queries.

  • Improved RAG grounding and answer reliability (healthcare information systems, finance, legal technology, education) Better first-stage retrieval can improve the evidence supplied to an LLM, reducing the chance that relevant documents are omitted before generation. NV-Embed-v2, GritLM, or SPLADE-v3 could be tested as components in evidence-retrieval pipelines, with answer evaluation separated from retrieval evaluation.

Dependencies: Higher retrieval scores do not directly establish factual accuracy, safety, or reduced hallucination. High-stakes systems require document provenance, access controls, human review, and end-to-end evaluation.

  • Daily-life search and recommendation improvements (consumer software) The findings can inform local or cloud-based search in personal note systems, email, file managers, shopping tools, and media libraries. Dense retrieval may help users find documents using natural-language descriptions, while sparse retrieval remains valuable for exact terms, filenames, people, and reference numbers.

Dependencies: Personal data must be indexed securely. On-device use may require smaller models, quantization, distillation, or hybrid sparse methods because the strongest models have substantial compute and memory demands.

Long-Term Applications

  • Adaptive multi-retriever routing systems (software, agentic AI, RAG)
    • exact entity or identifier queries → BM25 or SPLADE-v3;
    • semantic or long-tail questions → dense retrieval;
    • instruction-heavy queries → GritLM;
    • high-value queries → multiple retrievers with result fusion.

The paper’s cross-domain and failure-overlap analysis provides an empirical foundation for such routing.

Dependencies: This requires sufficiently large and representative query logs, reliable routing models, calibrated confidence estimates, and safeguards against routing errors. Dynamic routing must also justify its additional latency and operational complexity.

  • Failure-aware ensemble training (academic research, search platforms) Since different retrievers fail on partially different query subsets, future systems could train ensembles specifically on disagreement cases rather than simply averaging models. A training loop might:
    1. identify queries missed by the current retriever,
    2. determine which alternative retriever succeeds,
    3. learn a query-dependent fusion or reranking policy,
    4. evaluate worst-case and subgroup performance.

Dependencies: Requires query-level relevance judgments and careful prevention of overfitting to benchmark artifacts. Gains must be demonstrated on unseen domains.

  • Retrievers optimized jointly for quality, latency, and energy (cloud infrastructure, edge computing, sustainability)
    • retrieval effectiveness,
    • p95/p99 latency,
    • energy per query,
    • memory and storage,
    • index refresh cost,
    • monetary serving cost.

Potential products include an automated “retriever compiler” that selects quantization, ANN parameters, model size, and index structure for a target service-level objective.

Dependencies: The paper reports runtime but not a complete energy or total-cost analysis. Hardware-specific benchmarking and workload-level measurements are required.

  • Domain-adaptive retrievers for specialized sectors (healthcare, law, engineering, finance, science)
    • clinical terminology and medical literature,
    • legal citations and jurisdiction-specific language,
    • financial filings and regulatory documents,
    • engineering manuals and technical standards,
    • scientific literature and datasets.

These systems could combine dense semantic representations with sparse handling of exact terminology and structured metadata.

Dependencies: Domain adaptation requires high-quality relevance judgments, expert annotation, governance over sensitive data, and evaluation for temporal drift. In regulated fields, validation must also address explainability and auditability.

  • Multilingual and cross-lingual retrieval systems (global search, public policy, education) The paper notes that English-only and multilingual variants do not consistently dominate one another. This motivates systematic deployment studies for multilingual search, cross-lingual RAG, and public-information access rather than assuming that a multilingual model is automatically preferable.

Dependencies: Performance may vary sharply by language, script, dialect, transliteration, and document genre. Evaluation needs language-specific relevance judgments and should measure fairness across language communities.

  • Robust retrieval under instruction and prompt variation (agentic AI, policy, safety-critical search) The robustness metrics used in the study could support retrieval systems that remain stable when users paraphrase instructions or add irrelevant wording. Future products might expose a robustness score alongside standard retrieval quality and automatically test multiple instruction formulations before serving results.

Dependencies: Robustness should not mean ignoring legitimate changes in user intent. Systems need to distinguish harmless paraphrases from meaningful constraints and should combine learned retrieval with explicit filters.

  • Benchmarking standards and procurement policies for public-sector AI (policy, government, academia)
    • evaluation on representative local data,
    • disclosure of model and index configurations,
    • reproducible test suites,
    • worst-case and robustness metrics,
    • documentation of compute and environmental costs.

Dependencies: Standardization requires agreement on metrics and reporting formats. Benchmark results may become targets for optimization, so independent tests and periodically refreshed datasets are necessary.

  • Continuous retrieval evaluation for changing corpora (news, cybersecurity, finance, research monitoring) A future workflow could continuously test retrievers as documents, terminology, and user behavior change. Drift detectors could identify when a model loses performance on new topics, emerging entities, or newly introduced instructions, triggering re-indexing, query expansion, or model adaptation.

Dependencies: Continuous evaluation requires timely relevance labels or weak supervision, stable privacy practices, and mechanisms for distinguishing corpus changes from model regressions.

  • Personalized and privacy-preserving semantic search (daily life, education, workplace productivity) Compact dense models, sparse indexes, and local hybrid search could enable semantic retrieval over private email, notes, photos’ text metadata, and files without sending content to external services. The study’s quality–latency comparison can guide whether a device should use a lightweight dense model, local sparse retrieval, or cloud-assisted search.

Dependencies: On-device deployment requires model compression and efficient indexing. Privacy also depends on secure key management, protection against embedding leakage, and user control over indexing and deletion.

Glossary

  • Approximate nearest neighbor (ANN) search: An efficient method for finding vectors that are close to a query vector without exhaustively comparing every vector. “thus requires approximate nearest neighbor (ANN) search for efficient retrieval”
  • Attention layer: A neural-network component that learns which parts of an input should receive greater representational emphasis. “using a latent attention layer for pooling”
  • Bootstrap resampling: A statistical technique that repeatedly samples from observed data to estimate uncertainty or variability. “we compute using bootstrap resampling”
  • Contrastive learning: A training approach that brings representations of related items closer together while separating representations of unrelated items. “trained specifically for retrieval using contrastive learning objectives”
  • Contrastive objective: A loss or optimization target that distinguishes positive examples from negative examples in representation learning. “LLMs trained with contrastive or distillation-based objectives”
  • Context-aware term weight: A learned importance value assigned to a term based on its surrounding textual context. “it is derived from predicted (context-aware) term weights”
  • Contextualized embedding: A vector representation whose meaning depends on the surrounding context in which text occurs. “producing low-dimensional contextualized embeddings”
  • Corpus statistics: Aggregate measurements of term occurrence and distribution across a document collection. “term relevance is computed from corpus statistics”
  • Decoder-only retriever: A retrieval model based on a generative, decoder-only language-model architecture that is repurposed to encode text. “decoder-only retrievers based on generative models”
  • Dense retrieval: Retrieval based on comparing compact, continuous vector representations of queries and documents. “Dense Retrieval.”
  • Distribution shift: A difference between the data distribution encountered during deployment and that used during training. “The main objective of this approach is to mitigate the negative impact of distribution shifts”
  • Distributionally robust optimization (DRO): An optimization framework designed to maintain performance under changes or uncertainty in the data distribution. “implicit distributionally robust optimization (iDRO)”
  • Distillation: A training technique in which one model transfers knowledge or behavior to another model, often a smaller one. “contrastive or distillation-based objectives”
  • Dual encoder: A model that independently encodes two inputs, such as a query and a document, into vectors that can be compared. “The earliest dense retrievers are dual encoders trained contrastively”
  • Encoder-decoder retriever: A retrieval model using separate encoder and decoder components, typically adapted from a sequence-to-sequence LLM. “encoder-decoder retrievers that are generalizable and can be instruction-aware”
  • Encoder-only retriever: A retrieval model that uses only an encoder network to produce representations of queries or documents. “encoder-only retrievers that produce contextualized embeddings”
  • Expansion-based retrieval: Retrieval that augments queries or documents with generated terms or text before matching them. “Expansion-based Retrieval.”
  • Generative retrieval: A retrieval approach in which a generative LLM produces or identifies documents rather than only scoring a fixed index. “modern LLM-based and generative retrieval systems”
  • Hard negative: A non-relevant training example that resembles a relevant example and is therefore difficult for a model to distinguish. “draws global hard negatives from a dynamically refreshed ANN index”
  • Hypothetical-document generation: The generation of a synthetic document representing a query’s likely answer or relevant content for use in retrieval. “hypothetical-document generation”
  • Impact index: An inverted index that stores term-impact or term-weight information to support efficient weighted retrieval. “indexed using Lucene’s impact index”
  • Information retrieval (IR): The field concerned with finding and ranking information relevant to a user’s information need. “Document retrieval, as a subset of information retrieval (IR)”
  • Instruction-following retriever: A retriever trained to interpret additional textual instructions that specify retrieval goals or constraints. “instruction-following capabilities”
  • Inverted index: A data structure mapping terms to the documents in which they occur, enabling efficient lexical search. “efficient retrieval via inverted indexes such as Lucene”
  • Late interaction: A retrieval mechanism that separately represents query and document components and combines their matching scores at a later stage. “scores them by late interaction”
  • Lexical retriever: A retrieval model that estimates relevance primarily through shared words or terms between queries and documents. “Early retrieval systems, lexical retrievers, inferred relevance from the terms shared between a query and a document”
  • MarginMSE distillation: A distillation method that trains a model to reproduce differences in relevance scores between pairs of items using mean squared error. “adds balanced topic-aware query sampling with MarginMSE distillation”
  • Mean Average Precision (MAP): An evaluation metric that averages precision values at the ranks of relevant documents across queries. “Mean Average Precision (MAP) averages, over all queries”
  • Mean Reciprocal Rank (MRR): An evaluation metric based on the reciprocal rank of the first relevant retrieved document. “Mean Reciprocal Rank at cutoff kk (MRR@kk)”
  • Multilingual retrieval: Retrieval across multiple languages, often using a shared or language-general representation space. “multilingual retrieval”
  • n-gram: A contiguous sequence of n linguistic units, such as characters or tokens, used as a representation or matching feature. “queries and documents can be viewed as high-dimensional sparse vectors”
  • nDCG (normalized Discounted Cumulative Gain): A ranking metric that weights relevant documents by position and normalizes the result against an ideal ranking. “The normalized Discounted Cumulative Gain at cutoff kk (nDCG@k@k)”
  • Out-of-distribution (OOD) generalization: The ability of a model to perform well on data that differs from its training distribution. “To assess out-of-domain (OOD) generalization”
  • Pairwise MRR (p-MRR): A metric that measures the effect of changing retrieval instructions by comparing reciprocal ranks before and after the change. “The pairwise MRR (pp-MRR) evaluates the effect of retrieval instruction changes”
  • Pooling strategy: A method for combining token-level or intermediate neural representations into a single vector. “dense retrievers differ in retrieval mechanism, backbone architecture, pooling strategies”
  • Precision at cutoff: The proportion of the top-ranked retrieved documents that are relevant. “the proportion of retrieved documents that are relevant (precision)”
  • Pseudo-document: Generated text that approximates a document relevant to a query and is used to facilitate retrieval. “generating a pseudo-document and again retrieving with BM25”
  • Query expansion: The addition of generated or retrieved terms to a query to improve matching with relevant documents. “query expansion”
  • Query latency: The time required to process a query and return retrieval results. “query latencies are only seldom considered”
  • Recall at cutoff: The proportion of all relevant documents that appear among the top-ranked retrieved documents. “the proportion of relevant documents retrieved within the cutoff (recall)”
  • Residual compression: A compression technique that represents vectors using compact residual information relative to a coarse approximation. “using residual compression and kk-means-based indexing”
  • Retrieval-augmented generation (RAG): A system design in which retrieved documents provide external context for a generative LLM. “such as retrieval-augmented generation (RAG)”
  • Retrieval effectiveness: The degree to which a retrieval system returns relevant documents in high-quality rankings. “We evaluate all retrievers using standard metrics”
  • Robustness at cutoff: A worst-case retrieval metric that evaluates the minimum performance across instruction variants of the same query. “where identical queries with instruction variants are grouped together and the minimum nDCG@k@k within each group is taken”
  • Sparse retrieval: Retrieval based on high-dimensional vectors containing mostly zero values, commonly representing lexical terms. “Sparse Retrieval.”
  • Sparse vector: A vector in which most dimensions are zero and only a small number of features have nonzero values. “queries and documents can be viewed as high-dimensional sparse vectors”
  • Term frequency: The number of times a term occurs in a document or collection. “compute relevance using corpus statistics (e.g., term frequency)”
  • Token-level representation: A representation assigned separately to individual tokens rather than to an entire sequence. “learns token-level sparse representations”
  • Top-kk filtering: The process of retaining only the k highest-scoring or highest-weighted items. “top-kk filtering with k=512k=512”
  • Vector representation: A numerical encoding of text used to calculate similarity or relevance. “learning contextualized dense, low-dimensional vector representations of text”
  • Vocabulary-wide term score: A relevance or importance score computed for every term in a model’s vocabulary. “computing vocabulary-wide term scores from contextualized embeddings”

Open Problems

We're still in the process of identifying open problems mentioned in this paper. Please check back in a few minutes.

Continue Learning

We haven't generated follow-up questions for this paper yet.

Tweets

Sign up for free to view the 2 tweets with 33 likes about this paper.