Papers
Topics
Authors
Recent
Search
2000 character limit reached

Q2D-Web: A Large-Scale Benchmark for Retrieval in Agentic RAG Systems

Published 8 Sep 2026 in cs.IR and cs.CL | (2609.08887v1)

Abstract: Evaluating first-stage retrievers in large-scale production RAG requires a benchmark that pairs a large-scale corpus with a large set of agent-reformulated search queries based on real user queries and their conversation threads, and that labels many relevant documents per query. No existing public benchmark evaluates this setting: large-scale collections typically provide only a small number of evaluation queries, whereas benchmarks with many queries generally contain only millions of documents. Moreover, most benchmarks assess human-written queries, while the first-stage retrievers in agentic RAG pipelines serve machine-written reformulations whose distribution differs from human search behavior. To overcome these evaluation gaps, we introduce Q2D-Web (Query2Doc-Web), a large-scale agentic retrieval benchmark consisting of a 190M-document web corpus and 70k agentic search queries in ten languages, reformulated from real-world user queries in production systems. Q2D-Web provides three sets of fixed relevance judgments: agent citations, production rankings, and a combined set that unions both signals and adds LLM-based judgments of unlabeled pooled documents to reduce false negatives. We benchmark 13 retrievers including lexical, dense, and late-interaction models and find that their relative ordering is largely insensitive to the choice of judgment set, while diverging substantially across topical domains, query languages, and query types. To enable fast evaluation, we also study subcorpus sampling as an approximation to full-corpus evaluations. Retaining a third of the corpus, selected by reciprocal rank fusion over pooled retriever runs, preserves the full-corpus model ranking under the combined judgments while raising absolute Recall@1000 only by 3 to 7 points. The public leaderboard is accessible under: https://huggingface.co/spaces/perplexity-ai/q2d-web-leaderboard

Summary

  • The paper introduces Q2D-Web, a benchmark evaluating retrieval in agentic retrival-augmented generation (RAG) systems using a large corpus of approximately 190 million web documents and 69,721 agent-reformulated queries in ten languages.
  • Q2D-Web improves upon conventional benchmarks by including large query populations and machine-generated queries, typical of agentic multi-step searches, and employs Recall@1000 as its primary metric
  • Q2D-Web identifies diverse errors in retrieval performance, highlighting that models performing well in aggregate recall may not necessarily recover many distinct relevant documents.

Q2D-Web addresses a specific evaluation mismatch in agentic retrieval-augmented generation (RAG): deployed systems do not retrieve directly from short, human-written benchmark queries, but from machine-generated reformulations produced within multi-step conversational searches. The benchmark therefore evaluates the first-stage retriever under the conditions that determine which evidence is available to downstream rerankers and agents. Its central resource is a corpus of approximately 190 million web documents paired with 69,721 agent-reformulated queries in ten languages, derived from nine months of privacy-filtered production traffic (2609.08887).

Evaluation problem and benchmark position

First-stage retrieval is conventionally assessed using benchmarks such as MS MARCO, BEIR, and TREC Deep Learning. These resources provide valuable relevance judgments, but they do not jointly provide production-scale corpora, large query populations, and machine-generated queries characteristic of agentic search. MS MARCO Web Search contains approximately 100 million documents but only 9,374 test queries and one click-derived positive per query. By contrast, Q2D-Web contains roughly seven times as many test queries and an average of 99.6 positive documents per query in its combined judgment set.

The distinction between human and agent-generated queries is methodologically consequential. Each production search begins with a user message, but the agent generates one primary query and potentially several supporting queries using the conversation history and previously retrieved results. Supporting queries can alter both the surface form and the set of documents satisfying the information need. Since retrieval performance is sensitive to formulation—even semantically preserving query variations can reduce effectiveness by approximately 20% (2609.08887)—benchmarks based exclusively on manually authored queries may misestimate the behavior of production retrievers.

Q2D-Web consequently evaluates retrieval as candidate generation rather than final ranking. Its primary metric is Recall@1000, reflecting the requirement that a first-stage retriever expose a broad evidence pool to subsequent rerankers. Recall@100 and nDCG@10 are also reported to characterize shallower retrieval quality. This metric choice produces a substantive distinction from conventional web-ranking evaluations: a model can be strongest at recovering relevant documents overall while another is better at placing those documents near the top.

Construction of queries and corpus

The query set comprises 12,365 primary queries and 57,356 supporting queries, with supporting queries accounting for 82.3% of the dataset. The queries were sampled monthly while matching production distributions over topic and language. Exact duplicates, very short queries, filtering-operator queries, and queries flagged for PII were removed. The ten most frequent languages were retained.

The paper emphasizes that these queries are not merely paraphrases of user input. A primary query typically restates the user’s intent in a more retrieval-oriented form, whereas supporting queries may decompose the information need, seek background information, resolve adjacent entities, or conduct additional retrieval hops. The examples include queries whose formulation depends on prior conversational context and support queries that alternate between languages. This structure directly tests whether retrieval models trained on conventional human-written query-document pairs generalize to the distribution generated by agents.

The corpus is constructed by retrieving the top 5,000 documents for every dataset query using a production retrieval system, taking the union, and deduplicating documents with MinHash–LSH at a token 5-gram Jaccard threshold of 0.975. The resulting collection contains approximately 190 million canonical documents, with the longest document retained within each near-duplicate cluster. The mean document length is approximately 13,169 characters.

This is a deliberately nonrandom corpus. Every document was considered plausible for at least one benchmark query by a production retrieval stack, and documents retrieved for one reformulation can become hard distractors for related reformulations. The benchmark therefore concentrates evaluation on discriminating relevant evidence from semantically similar, topically related, temporally mismatched, or otherwise plausible web content rather than from uniformly sampled irrelevant pages.

Figure 1

Figure 1: Q2D-Web construction pipeline from production query sampling through corpus assembly, multi-source relevance annotation, and subcorpus selection.

Multi-source relevance judgments

Q2D-Web provides three relevance-judgment sets over the same query-corpus pair: Citation, Web Ranking, and Combined+LLM-Judged. This design recognizes that no single production signal provides complete relevance supervision.

Citation labels a document as relevant when the agent cited it in its final response. This signal is closely connected to downstream RAG use, but it is intentionally incomplete. Agents stop citing once a claim is supported, so redundant but relevant documents remain unlabeled. Citation judgments are therefore high-precision and low-recall.

Web Ranking labels up to the top 50 results per query from a production system using BM25 and dense first-stage retrieval followed by cross-encoder reranking. This pass recovers relevant documents that citation behavior omits, but it inherits the biases of that ranking stack and may differ from citation labels because it is computed over a later corpus snapshot.

The Combined+LLM-Judged set unions the first two sources and adds judgments for previously unlabeled candidates. Nine lexical, dense, and late-interaction retrievers released before 1 January 2025 generate a reciprocal-rank-fused pool of up to 500 previously unjudged documents per query. DeepSeek-V4-Flash then assigns binary relevance labels using a strict constraint-matching prompt. The temporal separation between pooling models and evaluated neural retrievers reduces direct benchmark-construction overlap, but it does not eliminate annotation bias. The LLM judge may favor fluent or long-form documents, and all three sources are conditioned on their respective retrieval or generation systems.

The use of multiple judgment sets allows the authors to test whether model comparisons are artifacts of a particular labeling source. The principal result is that relative model ordering is largely stable across Citation, Web Ranking, and Combined judgments. This stability supports the robustness of the aggregate comparisons, although it should not be interpreted as evidence that the judgment sets are interchangeable: absolute scores differ, and the Citation set necessarily measures downstream citation behavior more narrowly than topical or evidential relevance.

Subcorpus sampling at web scale

Full-corpus evaluation is computationally prohibitive. A full pass with the 4-billion-parameter pplx-embed-v1 model requires 4,608 H200 GPU-hours, while the 300-million-parameter EmbeddingGemma model requires nearly 200 H200 GPU-hours. Q2D-Web therefore evaluates whether a carefully selected subcorpus can preserve full-corpus conclusions.

All positively judged documents are retained. The remaining documents are selected as distractors using depth-based pooling, RRF over pooled retriever runs, or uniform random sampling. RRF performs best at the principal operating point. Retaining approximately 31.7% of the corpus preserves the full-corpus model ordering with Kendall’s τb=1.00\tau_b = 1.00. At this size, RRF increases mean Recall@1000 by 5.1 percentage points relative to full-corpus evaluation, compared with 11.1 points for uniform sampling. Depth pooling achieves a smaller 3.9-point discrepancy only at a larger retained fraction of 43.4%.

The selected RRF configuration, using k=1000k=1000, is therefore an efficient approximation rather than an unbiased replacement for full-corpus evaluation. Increasing the pool to k=2000k=2000 reduces the recall gap to 0.7 points, but requires retaining 52.9% of the corpus. For pplx-embed-v1-4b, the adopted subcorpus reduces computation from 4,608 to approximately 1,500 H200 GPU-hours; EmbeddingGemma falls below 70 GPU-hours. The paper reports that evaluation can consequently be performed in approximately eight hours on a single 8×H200 node instead of a multi-node cluster.

The subcorpus analysis includes a temporal holdout: pre-2025 models construct the pool, while the evaluated neural models are released after that cutoff. This is important because pooling can favor participating systems. Nevertheless, the guarantee established by the experiments is empirical and conditional on the evaluated model families, metrics, judgment set, and sampling pool. A new retriever with substantially different behavior could expose omitted documents and alter the approximation.

Benchmark results across retriever families

The authors evaluate thirteen retrievers: BM25, ten dense encoders, and two late-interaction models. Neural systems range from 300 million to 8 billion parameters. Documents are truncated to their first 512 tokens for indexing, a major cost constraint given the corpus scale. The paper explicitly declines to claim that the reported scores are unbiased estimates of full-document effectiveness; longer-input spot checks did not reorder the models, but most document content remains unseen under the main protocol.

Under Combined+LLM-Judged relevance on the full corpus, pplx-embed-v1-4b achieves the highest Recall@1000 at 69.11%, followed by Nemotron-3-Embed-8B at 68.58%. EmbeddingGemma-300M reaches 65.45%, while Qwen3-Embedding-8B reaches 64.53%. BM25 has the lowest aggregate Recall@1000 at 44.77%. The same winning models remain winners under subcorpus evaluation, although subsampling raises absolute Recall@1000 by approximately 3–7 points.

Retriever Full Recall@1000, Combined Subcorpus Recall@1000, Combined Interpretation
pplx-embed-v1-4b 69.11 74.07 Highest overall recall
Nemotron-3-Embed-8B 68.58 72.79 Strongest shallow ranking
EmbeddingGemma-300M 65.45 70.10 Strong sub-1B baseline
Qwen3-Embedding-8B 64.53 69.61 Monotonic scaling within family
pplx-embed-v1-0.6b 67.02 71.52 Strongest sub-1B dense model
mLateOn 60.98 65.62 Late interaction improves over mDenseOn
BM25-tantivy 44.77 49.85 Lowest aggregate recall

The metric-dependent leadership is particularly important. pplx-embed-v1-4b leads Recall@1000, but Nemotron-3-Embed-8B leads Recall@100 and nDCG@10 under Combined judgments. Nemotron therefore places a larger fraction of its recovered positives near the top, whereas pplx-embed-v1-4b recovers more positives across the full candidate pool. The result demonstrates that first-stage recall and shallow ranking quality should not be conflated.

Within Qwen3-Embedding, scaling is monotonic on Combined Recall@1000: 57.89% for the 0.6B model, 62.01% for the 4B model, and 64.53% for the 8B model. However, parameter scaling is not sufficient to determine cross-family performance: the 300M EmbeddingGemma model outperforms the 8B Qwen3 model on this metric. Late interaction is also inconsistent across families. mLateOn outperforms its dense counterpart mDenseOn, whereas pplx-embed-v1-late-0.6B trails pplx-embed-v1-0.6b. The advantage of late interaction is more visible at nDCG@10 than at Recall@1000, where both late-interaction models trail the strongest dense models at comparable scale.

Domain, language, and query-type effects

Aggregate scores conceal substantial heterogeneity. pplx-embed-v1-4b performs strongly across most domains, often matching or exceeding Nemotron-3-Embed-8B, while BM25 ranks last in every reported domain. Nemotron leads pplx-embed-v1-4b by two percentage points on Portuguese and one point on Russian, Italian, Spanish, and Korean, while matching or trailing it in the remaining languages. mDenseOn obtains 63% Recall@1000 on English but only 49–56% on other languages. mLateOn reaches 64% on English and 52–65% elsewhere, with Japanese slightly exceeding its English score.

The query-type split provides one of the clearest findings. Every neural retriever performs substantially better on primary queries than on supporting queries. The paper attributes this pattern partly to training-distribution mismatch: embedding models are commonly trained on human-written queries, which more closely resemble primary reformulations than the more exploratory, decompositional support queries issued by agents. This interpretation is plausible but not isolated experimentally from other factors, including support-query difficulty, query length, multilinguality, and the fact that support queries may target more specific or indirect evidence.

BM25 exhibits the opposite pattern, with supporting-query Recall@1000 approximately 1.9 percentage points higher than primary-query recall. The difference is small and does not overcome BM25’s low overall effectiveness, but it indicates that lexical matching may retain complementary value for machine-generated support queries whose salient entities or terms are explicit in the query.

Unique coverage and error structure

The paper’s most consequential claim is that aggregate recall does not determine ensemble value. BM25, despite the lowest Recall@1000, contributes 14,621 unique positives—documents not recovered by any other evaluated retriever. This is nearly six times the next-largest contribution: 2,451 unique positives from EmbeddingGemma-300M. mLateOn contributes 2,327, while pplx-embed-v1-4b and Nemotron-3-Embed-8B contribute only 884 and 705, respectively.

This result contradicts a simple strategy of selecting only the highest-scoring retrievers. The models with the best aggregate recall are comparatively redundant, whereas lexical retrieval recovers a much larger distinct portion of the judged relevant set. The result does not establish that BM25 should always be included in an ensemble, because the paper does not report a complete production ensemble study under fixed candidate budgets. It does establish that aggregate Recall@1000 is insufficient for selecting complementary retrievers.

The hard-negative analysis further characterizes the benchmark. Among documents ranked in the top 1,000 by all thirteen systems but labeled non-relevant, 15.8% are judged by the same LLM classifier to be relevant, indicating residual false negatives in the Combined set. Among the remaining hard negatives, the dominant categories are wrong aspect (39.0%), partial relevance (17.8%), related entity (13.8%), generic content (6.5%), and temporal mismatch (5.8%). These categories confirm that the corpus tests fine-grained constraint satisfaction rather than only broad semantic similarity.

The 6,424 hard positives—relevant documents missed by every evaluated retriever in the top 1,000—are dominated by lexical mismatch at 51.1%. Truncation accounts for 17.7%, multilingual mismatch for 15.4%, reasoning requirements for 7.0%, generic duplication for 4.1%, and entity or numeric constraints for 1.0%. The distribution implies that retrieval failures are not primarily caused by insufficient model scale. Lexical divergence, document truncation, and cross-lingual matching remain major failure modes even among large dense and late-interaction encoders.

Limitations and open questions

Q2D-Web’s principal limitation is restricted access. The corpus, queries, and judgments remain private to reduce training contamination, and only aggregate leaderboard results are exposed. This preserves benchmark validity more effectively than releasing all evaluation artifacts, but it limits independent auditing, reproducibility, error analysis, and examination of query-level variance. The benchmark is also derived from one commercial conversational search system, so its query distribution, citation policy, production ranking stack, and web corpus may not generalize to other agentic RAG deployments.

The relevance labels are not human gold judgments. Citation labels are sparse by construction, Web Ranking labels inherit system bias, and the Combined set relies partly on an LLM judge whose decisions may reflect stylistic preferences. The reported 15.8% false-negative rate among hard negatives directly demonstrates that the combined pool remains incomplete. In addition, the same LLM is used both to create part of the relevance set and to classify hard cases, so the error analysis is not independent validation.

The 512-token document truncation is another material constraint. Because the average document contains approximately 3,300 tokens, most document bodies are excluded from the main representation. The benchmark therefore measures retrieval under a fixed prefix-indexing protocol rather than unrestricted full-document retrieval. Finally, RRF subcorpus preservation is demonstrated for the tested models and Combined judgments, not guaranteed for future retrievers with different lexical, multilingual, or long-document behavior. The open empirical question is whether the substantial unique coverage of BM25 translates into consistent gains for fixed-budget hybrid ensembles, especially when candidate deduplication and downstream reranking costs are included.

Conclusion

Q2D-Web establishes a benchmark regime centered on the actual first-stage retrieval problem in agentic RAG: large heterogeneous corpora, machine-generated reformulations, multilingual queries, deep candidate recall, and multiple relevance signals. Its scale—approximately 190 million documents and 70,000 production-derived queries—addresses a gap left by benchmarks that provide either large corpora with few queries or many queries over relatively small collections.

The experiments show that model rankings are broadly stable across relevance sources, that RRF-based sampling preserves full-corpus ordering at roughly one-third of the corpus, and that high aggregate recall does not imply complementary retrieval coverage. The strongest operational conclusion is therefore not that one retriever dominates, but that evaluation must distinguish total recall, shallow ranking, query-type robustness, and unique positive recovery. Q2D-Web provides a framework for making those distinctions under production-scale constraints, while its private evaluation protocol leaves reproducibility and cross-system generalization as explicit unresolved questions.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is this paper about?

This paper introduces Q2D-Web, a large test for measuring how well search systems find useful web pages for AI assistants.

Modern AI assistants often use a method called retrieval-augmented generation, or RAG. Before answering a question, the assistant searches for information. It then uses the search results to create an answer and provide sources.

The paper focuses on the first search step: finding as many useful documents as possible from a huge collection of web pages.

The researchers created a very large benchmark containing:

  • About 190 million web documents
  • Nearly 70,000 search queries
  • Queries in 10 languages
  • Queries created by AI agents from real user conversations
  • Several kinds of labels showing which documents are relevant

The goal is to make testing search systems more realistic and reliable.

2. What questions did the researchers want to answer?

The paper mainly asks:

  1. How well do different search systems find relevant documents?
  2. Do systems work equally well for different languages, subjects, and types of queries?
  3. Are AI-generated search queries different enough from human-written queries that they need special testing?
  4. Can a smaller sample of the 190 million documents give almost the same results as testing the entire collection?
  5. Do different ways of deciding whether a document is relevant change which search system appears best?

These questions matter because a search system can only help an AI assistant if it finds the right information in the first place. If an important document is not retrieved, the assistant may never see it.

3. How did the researchers conduct the study?

Building the queries

The researchers collected searches made by AI agents in a real production system over nine months. The original user messages were private and had personally identifying information removed.

The agents created two main types of queries:

  • Primary queries: A clearer version of what the user originally wanted.
  • Support queries: Extra searches used to find background information or explore related details.

For example, a user might ask about the best books for learning cloud computing. The agent could create searches such as:

  • “best cloud computing books”
  • “best books on distributed systems”
  • “cloud computing textbooks”

This is different from a normal search benchmark because the queries were written by an AI agent, not directly by people typing into a search box.

Building the web collection

For every query, the researchers collected the top 5,000 results from an existing search system. They combined all these results and removed duplicates.

This produced a collection of about 190 million documents. The documents were not chosen randomly. Instead, they were documents that search systems already thought might be useful. This made them difficult “look-alikes” or hard negatives: pages that seem relevant but may not actually answer the question.

Deciding which documents were relevant

The researchers used three sources of relevance labels:

  1. Agent citations: A document was labeled relevant if the AI agent cited it in its answer.
  2. Web rankings: Documents ranked highly by a production search system were examined and labeled.
  3. LLM judgments: An AI model judged additional documents that had not already been labeled.

Using several sources is like asking several judges instead of relying on just one. This helps reduce mistakes. For example, an agent might cite only one document even though several documents contain useful information.

The researchers used a simple yes-or-no label:

  • Yes: The document is relevant.
  • No: The document is not relevant.

Testing search systems

The researchers tested 13 retrievers, or search systems. These included:

  • BM25, a traditional keyword-based method
  • Dense retrievers, which compare the meaning of queries and documents using mathematical representations called embeddings
  • Late-interaction retrievers, which compare parts of a query and document in more detail

An embedding is like a list of numbers that represents the meaning of a piece of text. Texts with similar meanings should have similar numerical representations.

The main measurement was Recall@1000. This asks:

Out of all the relevant documents, how many appear in the system’s first 1,000 results?

This is especially useful for the first stage of an AI search system, because the purpose of this stage is to collect many possible useful documents. Later stages can then choose the best few.

The researchers also measured:

  • Recall@100: How many relevant documents appear in the first 100 results?
  • nDCG@10: How good the first 10 results are, especially whether the most useful documents appear near the top

Testing a smaller document collection

Searching through 190 million documents is extremely expensive. The researchers therefore tested ways to create a smaller collection that still gives similar results.

Their best method used reciprocal rank fusion, or RRF. In simple terms, RRF combines the rankings from several search systems. Documents that appear near the top in several rankings receive a high combined score.

Using this method, the researchers created a smaller collection containing about one-third of the original documents.

4. What did the researchers find?

The best system depended on the goal

On the main measure, Recall@1000, pplx-embed-v1-4b performed best under the combined relevance labels.

However, Nemotron-3-Embed-8B performed best when the researchers cared more about placing useful documents very near the top, as measured by Recall@100 and nDCG@10.

This means there was no single system that was best at everything:

  • One system found the largest total number of relevant documents.
  • Another system placed useful documents closer to the top.

Neural systems generally beat BM25

The traditional BM25 system had the lowest overall recall. However, it still found some relevant documents that the other systems missed.

This is important because a weaker system may still add useful results when combined with stronger systems. Using several different search methods can provide better coverage than using only one.

Larger models often performed better

Within the Qwen3 family, larger models generally found more relevant documents. For example, the 8-billion-parameter model performed better than the 4-billion-parameter model, which performed better than the 0.6-billion-parameter model.

However, model size was not the only factor. Some smaller models performed better than larger models from other families.

Supporting queries were harder for neural systems

Most neural retrievers performed much better on primary queries than on support queries.

This may be because support queries are more unusual and are created to explore specific details. Most search models are trained mostly on human-written questions, which may look more like primary queries.

BM25 showed the opposite pattern: it performed slightly better on support queries. This difference suggests that combining keyword-based and meaning-based search systems could be useful.

Performance varied by language and subject

The search systems did not perform equally well in every language or topic.

For example, some models performed especially well in English, while their scores dropped for certain other languages. Different systems also performed better in different subject areas.

This shows why a benchmark should include many languages and topics instead of testing only English questions.

The smaller collection mostly preserved the results

The researchers found that keeping about one-third of the documents using RRF usually preserved the ranking of the search systems. In other words, the system that performed best on the full collection was also usually the system that performed best on the smaller collection.

The smaller collection made testing much cheaper and faster. However, its absolute recall scores were usually 3 to 7 percentage points higher than scores from the full collection. This happened because removing many irrelevant documents makes the search task easier.

Therefore, the smaller collection is useful for comparing systems, but its scores should not be treated as exactly the same as full-collection scores.

5. Why is this research important?

Q2D-Web is important because older search benchmarks often had one or more weaknesses:

  • They used relatively small document collections.
  • They had only a small number of test queries.
  • Their queries were written by humans rather than AI agents.
  • They labeled only one or a few relevant documents per query.
  • They often focused on the final search results rather than the first stage of an AI assistant’s search process.

Q2D-Web tries to solve these problems at the same time. It gives researchers a more realistic way to test whether a search system can support an AI assistant working on the web.

The benchmark could help developers:

  • Build assistants that find more trustworthy evidence
  • Improve answers in many languages
  • Test systems on difficult, realistic searches
  • Combine different search methods more effectively
  • Develop faster ways to evaluate very large search systems

The researchers have not released the full dataset publicly because future AI models might be trained on it. If that happened, those models could appear to perform well simply because they had already seen the test questions and documents. Instead, they provide a public leaderboard that evaluates eligible search systems without revealing the benchmark itself.

Simple conclusion

The paper’s main message is that testing search systems for AI assistants requires realistic, large-scale data. Q2D-Web provides a huge collection of web documents, many AI-generated queries, and several ways to identify useful sources.

The results show that search systems have different strengths. Some find many relevant documents, some place the best documents near the top, and some work better for particular languages or query types. A combination of different systems may therefore work better than relying on only one.

Overall, this research could lead to AI assistants that search more carefully, find better evidence, and give more accurate and well-supported answers.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • Representativeness of production traffic is unclear: The queries come from a single production deployment and a nine-month period, so it is unknown whether Q2D-Web generalizes to other agents, search products, user populations, time periods, or geographic regions.
  • Sampling bias from query filtering is not quantified: Removing PII-containing queries, operator-based queries, short queries, duplicate queries, and languages outside the ten most common may disproportionately exclude sensitive, navigational, technical, or long-tail information needs.
  • The language coverage is incomplete: Although the benchmark includes ten languages, the paper does not report balanced per-language sample sizes, linguistic properties, or performance for low-resource languages and code-switched queries.
  • The benchmark’s domain distribution is not externally validated: The topical categories and their production proportions are reported, but there is no comparison with broader web-search demand or an assessment of whether rare and emerging domains are adequately represented.
  • Primary and supporting queries may not be independent: Multiple queries originate from the same user search and conversation thread, yet the evaluation does not analyze cluster-level dependence or adjust confidence intervals and significance tests accordingly.
  • The effect of conversational context is not isolated: The paper does not compare retrieval performance on the original user message, the primary reformulation, and support reformulations under controlled conditions, leaving the contribution of context and reformulation unclear.
  • The causes of poor supporting-query performance remain speculative: The paper attributes the neural retrievers’ weaker supporting-query recall partly to training on human-written queries, but does not test this explanation against alternatives such as query ambiguity, length, specificity, multilinguality, or entity composition.
  • The corpus is not an independent web sample: Because documents are collected from the top 5,000 results of a production retrieval system, the benchmark may underrepresent documents that competing retrievers can discover but the production system missed.
  • Retriever-specific corpus bias is incompletely resolved: The corpus, Web Ranking labels, and candidate pools all depend partly on internal retrieval components, so the benchmark may favor systems with similar lexical, dense, or reranking behavior despite the use of multiple label sources.
  • The impact of corpus construction on conclusions is not measured: There is no comparison with a randomly sampled web corpus, an independently crawled corpus, or corpora built from alternative retrieval systems to quantify how much the 190M-document collection shapes model rankings.
  • Temporal validity is uncertain: Citation labels and corpus documents may originate from different points within the nine-month window, while Web Ranking is computed against a later snapshot; the effects of page updates, deletions, and temporal query intent are not measured.
  • Document deduplication may remove meaningful distinctions: MinHash–LSH clustering at a Jaccard threshold of 0.975 and retaining only the longest document could merge pages with different dates, versions, regional content, or editorial context.
  • The benchmark does not evaluate document freshness or temporal relevance: Queries involving current prices, news, laws, products, or events may require time-specific documents, but relevance is treated as a static binary property.
  • Relevance labels lack human validation at scale: Citation, Web Ranking, and LLM-generated labels are not systematically audited against independent human judgments, so their precision, recall, inter-annotator agreement, and domain-specific reliability remain unknown.
  • The strict LLM judging prompt may under-label partially relevant documents: Requiring every specific constraint to be satisfied and instructing the judge to answer “NO” when uncertain may favor exact-match relevance while penalizing useful background, complementary, or partially supporting evidence.
  • LLM judge reliability across languages and domains is unresolved: The paper does not report calibration, error rates, or agreement by language, topic, document length, or query type, leaving possible multilingual and domain-specific judge bias unquantified.
  • Binary relevance is insufficiently expressive for agentic RAG: The labels do not distinguish direct evidence, partial evidence, redundancy, authority, factual correctness, freshness, or usefulness for answer generation.
  • Citation labels conflate relevance with downstream citation behavior: A document may be relevant but uncited because it is redundant, inaccessible, poorly formatted, or unnecessary for the final answer; conversely, a cited document may contain errors or only weak support.
  • The Combined judgment set is not demonstrably exhaustive: Only the top 500 previously unjudged documents from pooled retriever rankings receive additional LLM judgments, leaving the long tail of potentially relevant documents unlabeled.
  • Pooling recall is not estimated: The paper does not use capture–recapture analysis, independent retrieval systems, random sampling, or human assessment to estimate how many relevant documents remain outside the citation, Web Ranking, and LLM pools.
  • The three label sets are not independent: Their shared dependence on overlapping production rankings, retrieved candidates, and related query-document pools complicates the claim that combining them substantially reduces source-specific bias.
  • No external or adversarial label audit is provided: The benchmark is not tested on cases involving misinformation, conflicting sources, subtle entity distinctions, temporal contradictions, or documents that are semantically similar but factually incorrect.
  • Evaluation is limited to first-stage retrieval metrics: The study does not establish whether higher Recall@1000 or nDCG@10 leads to better reranking, answer accuracy, citation correctness, factuality, latency, or user satisfaction in complete agentic RAG systems.
  • The candidate depth of 1,000 is not justified for all downstream pipelines: Different rerankers and agent workflows may require different candidate-set sizes, and the relationship between retrieval depth, reranking quality, and final answer quality is left unexplored.
  • The metric treatment of many relevant documents is unresolved: Recall@1000 depends on the judged positive set and can change as more positives are discovered; the paper does not analyze alternative metrics such as recall at fixed numbers of unique evidence items, graded gain, or utility-weighted recall.
  • Statistical uncertainty is largely absent: The reported aggregate and slice-level differences are not accompanied by confidence intervals, significance tests, bootstrap procedures, or analyses of query-level variance.
  • Domain and language slice sizes are not reported sufficiently for reliability assessment: It is unclear whether apparent differences between models or languages are statistically robust or driven by small or correlated subsets.
  • The 512-token document truncation substantially limits validity: Most documents are much longer than the encoded prefix, and the reported longer-context spot checks do not establish unbiased full-document effectiveness across all models, languages, and query types.
  • Long-context retrieval is not systematically evaluated: The paper does not test passage-level indexing, document chunking, hierarchical retrieval, or late-document evidence, all of which could materially change model rankings.
  • Model comparisons may be affected by implementation differences: Default instruction templates, token limits, indexing configurations, pooling strategies, quantization, batching, and hardware settings are not analyzed as possible sources of performance variation.
  • The retriever model set is temporally narrow and potentially incomplete: Most neural models were selected by release date and availability, but the study does not assess closed-weight systems, learned sparse retrievers, hybrid systems, fine-tuned models, reranker-assisted retrievers, or newer multilingual approaches.
  • Scaling conclusions are confounded by model family and training data: The observed relationship between parameter count and recall cannot be interpreted causally because model architecture, pretraining data, instruction format, tokenizer, training objective, and context length vary simultaneously.
  • No task-specific or benchmark-adaptive training is studied: It remains unknown whether fine-tuning on agentic reformulations, multilingual production queries, citation behavior, or hard negatives would close the gap between primary and support queries.
  • Unique-positive analysis does not establish complementarity in a combined system: A retriever’s uniquely recovered documents may be redundant, low quality, or unusable downstream; the paper does not measure the value of combining retrievers under a fixed candidate budget.
  • Subcorpus sampling is validated on a limited holdout design: The temporal holdout excludes models released after January 2025, but it does not test substantially different architectures, proprietary systems, adversarial retrievers, or models trained specifically on Q2D-Web-like data.
  • Subcorpus preservation is shown mainly for model ordering, not all evaluation uses: Kendall’s τb\tau_b and Recall@1000 score differences do not establish preservation of per-query rankings, nDCG, calibration, hard-negative behavior, or downstream RAG outcomes.
  • The RRF sampling strategy may become invalid as models evolve: New retrievers may retrieve documents outside the existing pooled candidate distribution, and the benchmark does not provide an adaptive mechanism for detecting when the subcorpus is no longer representative.
  • Subcorpus score inflation is not fully corrected: The reported 3–7 point increase in Recall@1000 is documented but not analytically corrected, making comparisons between full-corpus and subsampled results difficult.
  • Random sampling is not evaluated as a statistically repeated procedure: The paper does not report variance across random seeds or establish whether the observed sampling behavior is stable across independent random subcorpora.
  • The private benchmark limits reproducibility and error analysis: Researchers cannot independently inspect the corpus, queries, judgments, duplicates, or failure cases, and the leaderboard may not provide enough information to diagnose discrepancies or reproduce published findings.
  • Leaderboard evaluation protocols are not fully specified: The paper does not clarify submission limits, inference-time constraints, indexing requirements, handling of proprietary models, or protections against indirect benchmark overfitting.
  • Contamination risk is reduced but not eliminated: Keeping the benchmark private does not rule out overlap between production web documents, public training data, model pretraining corpora, and queries or facts that can be inferred from publicly available sources.
  • The benchmark’s future maintenance is unspecified: There is no clear plan for updating documents, handling web drift, revising stale labels, detecting contamination, or preserving comparability across benchmark versions.
  • Privacy and governance details are limited: Although the paper states that traffic is PII-free and filtered, it does not describe the detection procedure, false-negative risk, retention policy, or whether conversational context could permit re-identification.
  • The benchmark does not address safety-sensitive retrieval: Queries involving health, law, finance, personal data, or harmful content may require specialized relevance and reliability criteria that are not captured by the binary labels or aggregate recall metrics.

Practical Applications

Immediate Applications

  • Production RAG retriever selection and regression testing — Industry, software, search
    • Organizations operating web-search or agentic RAG systems can use Q2D-Web-style evaluation to compare first-stage retrievers under realistic conditions: large corpora, multilingual queries, machine-generated reformulations, and deep candidate retrieval.
    • A practical workflow is to evaluate candidate models on the RRF-sampled subcorpus first, then validate finalists on the full corpus. The paper reports that retaining roughly one-third of the corpus preserves model ordering while reducing evaluation cost substantially.
    • This can support model-release gates, continuous integration tests, and monitoring for retrieval regressions after changes to embedding models, indexing, query rewriting, or document preprocessing.
    • Dependencies: Access to a representative query log, a stable corpus snapshot, reliable relevance judgments, sufficient GPU/indexing infrastructure, and controls against benchmark contamination.
  • Cost-efficient benchmarking of large retrieval systems — Industry and academia
    • The RRF-based subcorpus can become an evaluation tool for rapid experimentation with dense, lexical, and late-interaction retrievers without repeatedly scanning the full 190M-document collection.
    • Teams can create analogous subcorpora by pooling top-ranked documents from diverse retrievers, retaining all known positives, and adding hard distractors. This is more useful than uniform random sampling because it preserves plausible competitors that can change system rankings.
    • Dependencies: The sampling pool must include multiple retrieval families and preferably temporally held-out models; otherwise, the benchmark may favor the systems used to construct it.
  • First-stage candidate-generation design for agentic RAG — Industry, software
    • The findings support using high-recall dense retrieval as the main candidate generator, especially for multilingual web search. In the reported experiments, pplx-embed-v1-4b performed best on combined Recall@1000, while Nemotron-3-Embed-8B performed strongly at shallower cutoffs.
    • A production pipeline could use a strong dense retriever to generate a broad candidate pool, followed by a cross-encoder or learned reranker for precision. The paper’s results indicate that optimizing only nDCG@10 may select a different model from one optimized for deep recall.
    • Dependencies: Deployment cost, latency, vector-index support, embedding refresh requirements, and the specific balance between recall, top-rank precision, and downstream reranking capacity.
  • Hybrid retrieval and complementary model ensembles — Industry, software
    • BM25 should not necessarily be removed even though it has the lowest aggregate recall. It recovered the largest set of relevant documents uniquely missed by the other tested retrievers.
    • A practical system could fuse BM25 with dense or late-interaction retrieval, particularly for exact names, rare terminology, identifiers, quotations, dates, and rapidly changing content. Candidate unions can then be reranked by a cross-encoder or an LLM.
    • Dependencies: Additional indexing and serving cost, calibrated score fusion, language-specific tokenization quality, and safeguards against duplicate or low-quality candidates.
  • Improved multilingual search and cross-lingual RAG — Industry, education, public services
    • The benchmark’s ten-language query distribution enables organizations to test whether a retriever works consistently across languages rather than relying on English aggregate scores.
    • Products could use language-specific routing, multilingual embedding models, or language-aware fusion when performance differs substantially across languages. The results suggest that model choice matters for Portuguese, Russian, Italian, Spanish, Korean, Japanese, and other non-English settings.
    • Dependencies: Adequate evaluation volume per language, accurate language identification, representative regional content, and attention to script, morphology, transliteration, and cultural terminology.
  • Evaluation of agent query reformulation and decomposition — Industry, academia
    • The distinction between primary and support queries provides a direct diagnostic for agentic search pipelines. Teams can separately measure retrieval quality for queries that closely restate user intent and queries generated to gather supporting evidence.
    • Because neural retrievers performed substantially worse on support queries, organizations can test query-type-specific strategies: specialized encoders, expansion, entity extraction, lexical fallback, or additional reformulation before retrieval.
    • Dependencies: Correct classification of primary and support queries, access to conversation context, and relevance labels that reflect the distinct information needs of each reformulation.
  • Evidence coverage and citation-grounding audits — Industry, healthcare, finance, legal technology
    • The three judgment sources suggest a practical audit framework for RAG systems: compare documents cited by the agent, documents highly ranked by production retrieval, and independently judged relevant documents.
    • This can reveal whether a system retrieves evidence but fails to cite it, cites only one redundant source, or omits relevant documents because of sparse citation behavior.
    • In high-stakes sectors, the framework could support dashboards for evidence recall, citation coverage, source diversity, and retrieval failures before answer generation.
    • Dependencies: Domain-appropriate relevance criteria, human review for consequential decisions, protection of sensitive search logs, and recognition that citations are high-precision but low-recall signals.
  • Benchmark governance and contamination-resistant evaluation — Academia, standards bodies, policy
    • The paper’s private-corpus and public-leaderboard approach can be applied to sensitive or rapidly contaminated benchmarks. A trusted evaluator can keep queries, documents, and labels private while allowing researchers to submit open-weight models for scoring.
    • This is useful for measuring generalization without allowing benchmark examples to enter training corpora and inflate future results.
    • Dependencies: A secure submission service, reproducible execution environments, transparent scoring protocols, abuse prevention, and sufficient disclosure about data construction without exposing evaluation instances.
  • Improved search workflows for researchers and ordinary users — Daily life, education
    • Agentic search systems based on the paper’s findings can combine primary queries with supporting searches, multilingual reformulations, and hybrid lexical-dense retrieval.
    • For tasks such as literature discovery, product comparison, travel planning, or fact checking, this can increase evidence coverage and reduce dependence on a single phrasing of the user’s question.
    • Dependencies: The agent must preserve user intent during reformulation, distinguish useful supporting evidence from tangential results, and present uncertainty and source provenance clearly.

Long-Term Applications

  • A standardized benchmark ecosystem for agentic retrieval — Academia, industry, policy
    • Q2D-Web could motivate future benchmarks that jointly evaluate query reformulation, first-stage retrieval, reranking, citation selection, and final answer grounding rather than treating these as isolated components.
    • A mature benchmark family could include private test sets, public development subsets, multilingual slices, temporal holdouts, and domain-specific versions for medicine, law, science, and finance.
    • Dependencies: Long-term maintenance of corpus snapshots, evolving relevance definitions, privacy-preserving access to production data, and independent governance to avoid overfitting to one vendor’s retrieval stack.
  • Domain-specific high-stakes retrieval evaluation — Healthcare, law, finance, public policy
    • The benchmark methodology could be adapted to biomedical literature, clinical guidelines, legal decisions, regulatory filings, or financial disclosures. Multi-source judgments would help identify relevant evidence that is not cited, clicked, or returned by a single production ranker.
    • Such benchmarks could measure whether a RAG system retrieves all materially relevant evidence, including contradictory or minority sources, rather than merely retrieving fluent or popular documents.
    • Dependencies: Expert annotation, temporal validity, jurisdictional or clinical context, strict privacy controls, and higher standards than LLM-only judging.
  • Training retrievers specifically for agent-generated support queries — Industry and academia
    • The observed gap between primary and support-query performance suggests a new training direction: collect agent-generated reformulations, hard negatives, multi-hop query sequences, and conversation context for contrastive or instruction tuning.
    • Future retrievers could condition representations on the query’s role—primary intent, background retrieval, entity lookup, or evidence verification—and dynamically select different retrieval strategies.
    • Dependencies: High-quality labels for support queries, avoidance of self-training bias from existing rankers, diverse agent policies, and evaluation on temporally and architecturally held-out data.
  • Adaptive retrieval ensembles and query-type routing — Software, robotics, enterprise search
    • Systems could learn to route each query to the most suitable retrieval combination: dense retrieval for semantic paraphrases, BM25 for exact terms, late interaction for fine-grained matching, and specialized indexes for entities or structured facts.
    • In an embodied or robotic agent, the same approach could support different information needs during planning—for example, broad background retrieval followed by exact specification or safety-document retrieval.
    • Dependencies: Reliable query classification, low-latency model switching, score calibration across retrievers, and sufficient evidence that routing improves end-to-end task success rather than only offline recall.
  • Temporal and continuously updated web retrieval evaluation — Search, news, policy
    • Because the benchmark distinguishes production-time citations from judgments against a later corpus snapshot, the methodology can support evaluation of content drift, source disappearance, and changing facts.
    • Future systems could measure whether a retriever finds current evidence, preserves historical evidence when appropriate, and handles conflicts between documents from different dates.
    • Dependencies: Versioned web corpora, timestamp-aware relevance labels, robust deduplication, and policies for handling deleted, revised, or syndicated content.
  • Learned corpus sampling and evaluation surrogates — Information retrieval research
    • The RRF approach could be extended into learned sampling systems that predict which documents are most important for preserving model ordering, metric fidelity, and hard-negative coverage.
    • A future evaluation service might automatically construct a task-specific subcorpus for each new domain or query distribution and report confidence intervals for how closely it approximates full-corpus results.
    • Dependencies: Validation across unseen architectures and domains, statistical guarantees, monitoring for sampling bias, and methods to correct the reported inflation of Recall and nDCG on reduced corpora.
  • Retrieval-aware answer generation and evidence planning — Agentic AI
    • The benchmark’s focus on deep recall could enable agents to plan retrieval budgets explicitly: retrieve broadly when evidence coverage is uncertain, switch to targeted support queries when gaps remain, and stop when additional retrieval yields redundant evidence.
    • This may produce agents that optimize not only answer fluency but also source diversity, claim coverage, and recovery of hard-to-find evidence.
    • Dependencies: Reliable uncertainty estimation, claim-level evidence alignment, cost-aware planning, and safeguards against retrieving large volumes of irrelevant or low-quality material.
  • Policy and procurement standards for trustworthy RAG systems — Government and enterprise governance
    • Public agencies and enterprise buyers could require vendors to report retrieval performance separately for primary and support queries, languages, domains, cutoffs, and judgment sources.
    • Procurement evaluations could include deep-recall requirements, citation-grounding audits, held-out private tests, and explicit reporting of model size, latency, energy use, and corpus-update procedures.
    • Dependencies: Agreement on sector-specific relevance standards, reproducible auditing access, privacy-preserving evaluation, and recognition that offline retrieval metrics do not alone establish factual accuracy or safety.

Glossary

  • Agentic retrieval-augmented generation (RAG): A system in which an autonomous software agent iteratively retrieves external information to support generated responses. “Agentic retrieval-augmented generation (RAG) relies on retrieval to identify evidence for accurate, grounded responses”
  • Candidate set: A collection of initially retrieved documents passed to later ranking stages for further processing. “the first-stage retriever scans the full index and returns a candidate set that later stages rerank”
  • Canonical representative: The single document retained to represent a cluster of near-duplicate documents. “Within each cluster, we retain the longest document as the canonical representative”
  • Challenging distractor: An apparently relevant but ultimately unjudged or non-relevant document that can make retrieval evaluation more difficult. “we consider a document as a challenging distractor if at least one retriever ranks it above a relevant document in its top-kk results”
  • Click-derived labels: Relevance labels inferred from users’ document-clicking behavior rather than from explicit human assessment. “MS~MARCO Web Search builds the largest public collection with web click-derived labels”
  • Corpus snapshot: A fixed version of a document collection captured at a particular point in time. “Web Ranking is computed against the corpus snapshot”
  • Cross-encoder reranking: A ranking method that jointly processes a query and document with a neural model to assign a relevance score after initial retrieval. “followed by cross-encoder reranking”
  • Dense retrieval: Information retrieval based on comparing continuous vector representations of queries and documents. “For dense retrieval, we select ten open-weight multilingual encoders”
  • Deep-recall protocol: An evaluation procedure designed to measure how many relevant documents are retrieved at a large cutoff suitable for downstream processing. “we use a deep-recall protocol”
  • Document deduplication: The process of identifying and removing duplicate or near-duplicate documents from a corpus. “We then deduplicate this collection with MinHash--LSH”
  • Embedding model: A neural model that maps text into numerical vector representations for similarity-based retrieval. “Most IR benchmarks treat nDCG@10 as their primary metric because they target the final ranked list presented to users.”
  • False negative: A relevant item incorrectly treated as non-relevant because it lacks a relevance label. “Such false negatives are common under sparse labels”
  • First-stage retriever: The initial retrieval component that searches a large index and produces candidates for subsequent ranking stages. “The first-stage retriever consequently bounds what the agent can read and ultimately cite.”
  • Hard negative: A non-relevant document that is highly similar to a query or ranks highly, making it difficult to distinguish from a relevant document. “a document retrieved for one reformulated query can serve as a hard-negative candidate for a related query”
  • Held-out model: A model excluded from dataset or subcorpus construction so that its evaluation tests generalization. “the held-out models are tested against a pool built from all three families”
  • Information retrieval (IR): The field concerned with locating relevant information, typically documents, in response to queries. “Information retrieval (IR) evaluation has long relied on benchmarks of limited scale”
  • Inverse authoring: A question-generation procedure that begins with a known fact and adds constraints to make the resulting question difficult to answer. “BrowseComp extends this design to 1,266 harder questions built by inverse authoring”
  • Jaccard similarity: A measure of set overlap calculated as the size of the intersection divided by the size of the union. “clustering documents whose token $5$-gram sets have a Jaccard similarity of at least 0.975”
  • Judgment pool: A collection of documents gathered from multiple retrieval runs for subsequent relevance assessment. “We merge these rankings with reciprocal rank fusion, select the top 500 previously unjudged documents”
  • Late-interaction model: A retrieval model that separately represents query and document tokens while allowing fine-grained token-level matching during scoring. “For late-interaction retrieval, we include mLateOn and pplx-embed-v1-late-0.6B”
  • Lexical retrieval: Retrieval based primarily on exact or weighted matching of words and terms rather than semantic vector similarity. “We evaluate lexical, dense, and late-interaction retrievers”
  • MinHash–LSH: A scalable approximate method for detecting similar sets or documents by combining MinHash signatures with locality-sensitive hashing. “We then deduplicate this collection with MinHash--LSH”
  • nDCG@10: Normalized discounted cumulative gain measured over the top ten ranked results, emphasizing highly ranked relevant documents. “We therefore adopt Recall@1000 as our primary metric, and additionally report Recall@100 and nDCG@10”
  • Open-weight model: A model whose learned parameters are publicly available for use or inspection. “All neural retrievers are open-weight”
  • Pooled retriever runs: The combined ranked outputs of multiple retrieval systems used to identify candidates for judging or corpus construction. “We merge these rankings with reciprocal rank fusion”
  • Production traffic: Real user requests processed by a deployed system. “sampled from nine months of PII-free production traffic”
  • Query reformulation: The transformation of an original user query into a new query intended to improve retrieval or reflect additional context. “the agent reformulates, decomposes, and iteratively revises the original user query”
  • Recall@1000: The proportion of all judged relevant documents retrieved within the first 1,000 results. “We focus on Recall@1000 as our primary metric”
  • Reciprocal rank fusion (RRF): A method for combining rankings by assigning each result a score based on the reciprocal of its rank across systems. “Retaining a third of the corpus, selected by reciprocal rank fusion over pooled retriever runs”
  • Relevance judgment: A label indicating whether a document satisfies a query’s information need. “A judgment in Q2D-Web is a binary relevance label on a query--document pair.”
  • Sparse labels: Relevance annotations covering only a small fraction of the potentially relevant documents. “Because unjudged documents are conventionally treated as non-relevant, a system that retrieves relevant but unlabeled documents receives a lower score”
  • Subcorpus sampling: The construction of a smaller evaluation corpus intended to approximate evaluation on the full corpus. “Subcorpus sampling achieves this by carefully selecting a subset that preserves the documents needed to reproduce the original evaluation results.”
  • Temporal holdout: An evaluation design that separates data or models according to time to test performance without temporal contamination. “We use a temporal holdout to evaluate subcorpus generalization to models that did not contribute to its construction.”
  • Top-kk pooling: The process of collecting the highest-ranked kk documents from one or more retrieval runs. “Depth-kk pooling includes every unjudged document that appears in the top-kk results”
  • Vector representation: A numerical encoding of text used to compute semantic similarity. “late-interaction models emit one vector per token”
  • Web-scale retrieval: Information retrieval performed over extremely large collections, typically containing millions or billions of web documents. “Q2D-Web isolates the retriever at web scale”
  • Zero-shot contamination: The inflation of benchmark performance when evaluation data appears in a model’s training data. “Public benchmarks often end up in training data, intentionally or unintentionally through derivative datasets”

Tweets

Sign up for free to view the 9 tweets with 58 likes about this paper.