Papers
Topics
Authors
Recent
Search
2000 character limit reached

MS MARCO Dataset for Question Answering and Retrieval

Updated 26 September 2026
  • MS MARCO is a large-scale dataset designed for machine reading comprehension, open-domain question answering, and neural information retrieval, utilizing real search queries and passages.
  • Key components include anonymized Bing queries, passages retrieved and edited by humans for relevance, with approximately 1,010,916 questions and 8,841,823 passages; offers tasks for novice, intermediate and passage re-ranking.
  • MS MARCO supports tasks like passage re-ranking, question answering, and multimedia extensions, providing a comprehensive dataset for evaluating retrieval systems in realistic search scenarios.
  • follow_up_questions

MS MARCO (“MAchine Reading COmprehension”) is a family of large-scale datasets and evaluation resources introduced to support machine reading comprehension, open-domain question answering, answer generation, passage ranking, document ranking, and neural information retrieval. The original resource contains anonymized Bing search-query questions, human-generated answers, retrieved web passages, source documents, passage-selection annotations, and explicit unanswerable cases (Bajaj et al., 2016). Subsequent work has used MS MARCO as a foundation for neural reranking, dense and sparse retrieval, multilingual transfer, cross-language information retrieval, benchmark analysis, entity-aware retrieval, and web-scale search-system evaluation.

1. Origins, scope, and dataset structure

MS MARCO was constructed to model a realistic search-assistant pipeline rather than a narrowly extractive reading-comprehension problem. A user submits a question-like query to Bing; the search engine retrieves documents and passages; a system determines whether the evidence is sufficient; it identifies and combines relevant evidence; and it produces a natural-language answer or declines to answer when the evidence is inadequate.

The questions were sampled from anonymized Bing search logs, including queries submitted through Bing and Cortana-related search interactions. A classifier trained on previously human-annotated data filtered out navigational and other non-question-oriented queries. Because the queries originate in actual search behavior, they can be incomplete, abbreviated, ambiguous, misspelled, noisy, or phrased as keywords rather than conventional questions. Examples include “tablespoon in cup” and “barack obama age.”

The original dataset contains:

  • 1,010,916 questions;
  • 1,026,758 unique answers;
  • 182,669 completely human-rewritten answers;
  • 8,841,823 passages;
  • 3,563,535 source web documents.

Approximately ten passages were supplied per question on average. The passages were extracted from documents retrieved by Bing and presented to editors in ranked order. The source-document release includes URLs, titles, and body text. Approximately 300,000 documents could not be recovered from Bing’s index, and recovered documents may have changed after the original passages were extracted.

Each query may have no answer, one answer, or multiple answers. Editors were instructed to use only information in the supplied passages, although answers could paraphrase, synthesize, or combine information distributed across several passages. A passage-level is_selected field records whether an editor used a passage in composing the answer. The annotation is incomplete: is_selected = 0 does not establish that a passage is irrelevant, because editors were not required to mark every passage that could support the answer.

Queries were automatically assigned to five answer-type segments:

Segment Percentage
Description 53.12%
Numeric 26.12%
Entity 8.81%
Location 6.17%
Person 5.78%

These segments differ from textual query characteristics. For example, 34.96% of queries contain “what,” 16.80% contain “how,” 3.46% contain “where,” and 27.83% fall into an “Other” interrogative category.

2. Annotation, answerability, and proposed tasks

The annotation pipeline sampled queries, filtered non-question intents, retrieved documents and passages with Bing, presented ranked passages to human editors, obtained answerability judgments, recorded selected passages, and elicited natural-language answers. Editors marked a query unanswerable when no supplied passage supported an answer. In the novice task, the expected output for such a query is “No Answer Present.”

The dataset defines three principal tasks.

Novice answerability and answer generation requires a system to determine whether a question is answerable from a set of context passages and, if so, generate the answer. If the evidence is insufficient, the system should output “No Answer Present.”

Intermediate well-formed answer generation uses the same inputs but requires an answer that remains understandable without the original question and passages. A response such as “16” is therefore inadequate for “tablespoon in cup”; the desired form is “There are 16 tablespoons in a cup.” The task requires answerability detection, evidence selection, paraphrase or synthesis, grammatical generation, and resolution of implicit context.

Passage re-ranking supplies a query and 1,000 passages retrieved with BM25. The system must rank passages according to their likelihood of containing information relevant to answering the query. The released passage collection is the union of MS MARCO passages, with query–passage relevance signals derived from editor selections.

MS MARCO explicitly supports unanswerable questions and multiple answers. Multiple references reflect natural variation in wording, sentence structure, detail, and explicitness. Unanswerability models an important search behavior: refusing to answer when retrieved evidence is insufficient or conflicting. However, later research shows that the treatment of unanswerable queries in commonly used ranking subsets produces substantial selection effects.

3. Baselines, metrics, and the evolution of retrieval research

The original experiments included a vanilla sequence-to-sequence model, a memory network, a discriminative passage-ranking model, and non-generative references such as the best individual passage and a DSSM-like passage-ranking system. In version 1.1, the reported ROUGE-L values were 0.089 for the question-only sequence-to-sequence model, 0.119 for the memory network, 0.177 for passage ranking, and 0.351 for the best passage reference. On a multi-answer subset, the memory network obtained BLEU 0.340 and pa-BLEU 0.341, while the best-passage reference obtained BLEU 0.359 and pa-BLEU 0.453.

Numeric experiments evaluated the Attention Sum Reader and ReasoNet. Their MS MARCO accuracies were 55.0 and 58.9, respectively, compared with 69.5 and 74.7 on the CNN test set. The results were interpreted as evidence that noisy, search-derived numeric questions were more difficult than the corresponding CNN/Daily Mail setting.

For version 2.1, a BiDAF-style span model obtained ROUGE-L 0.150 on the novice task and 0.170 on the intermediate task. Human ensembles achieved 0.73703 and 0.63044, respectively. The model struggled with unanswerable questions and with intermediate answers requiring words or constructions not appearing contiguously in the passage.

Evaluation uses different measures for different output types. Numeric answers use accuracy and precision-recall measures. Long textual answers use ROUGE-L, BLEU, and phrasing-aware metrics including pa-BLEU. Because multiple human answers can be semantically equivalent but lexically different, the paper introduced reference-sensitive evaluation intended to better reflect answer variability.

The original ranking resource became a major training collection for neural information retrieval. The passage-ranking setting contains approximately 8.8 million passages, more than 500,000 training queries in commonly used ranking releases, and roughly 39 million query–positive–negative triples. The document-ranking setting contains approximately 3.2 million documents, with BM25 candidate retrieval followed by neural reranking.

A two-stage architecture became standard: a classical retriever such as BM25 generates candidates, and a neural model performs expensive second-stage scoring. Longformer was applied to MS MARCO document reranking because standard BERT-style models are constrained by short input windows. It re-ranked BM25’s top 100 documents using up to 4,096 tokens and achieved MRR@100 of 0.336 on development according to the results table and 0.305 on test (Sekulić et al., 2020). The paper’s abstract reports a conflicting development value of 0.329.

Traditional retrieval remained competitive when carefully engineered. A system combining tuned BM25, field-aware features, proximity, cosine similarity, IBM Model 1 lexical-translation scores, and LambdaMART achieved MRR@100 of 0.298 on the document-ranking leaderboard snapshot dated 2020-12-06, exceeding several neural submissions and all other non-neural submissions reported in that study (Boytsov, 2020). This result was below most BERT-based systems, but demonstrated that comparisons against untuned BM25 can be misleading.

The benchmark also stimulated analysis of leaderboard methodology. MS MARCO document ranking uses MRR@100, with approximately 3.2 million documents, 367,000 training queries, 5,193 development queries, and 5,793 test queries. Each query has one designated relevant document in the described document-ranking setting. The scalar MRR score aggregates qualitatively different outcomes: both systems failing, one system succeeding while the other fails, and both succeeding at different ranks. A study of chronological SOTA runs argued that MRR can conceal whether an improvement comes from answering more queries or ranking shared successes earlier (Lin et al., 2021).

A related evaluation study compared MS MARCO with TREC Deep Learning. MS MARCO offers hundreds of thousands of training queries with sparse binary supervision, whereas TREC uses NIST-selected queries, pooled results, and four-point graded judgments. Bootstrap analyses found that the leading document run remained first in 91.2% of trials and the leading passage run remained first in 72.7% of trials, but private and TREC evaluations weakened public-leaderboard dominance. The proposed criterion was “robust usefulness”: methods should generalize across datasets and judging schemes, remain deployable without extraordinary effort, and improve real user outcomes (Craswell et al., 2021).

4. Biases, incomplete judgments, and reproducibility

MS MARCO reflects Bing users, Bing’s retrieval system, its geographic and demographic reach, and search-engine query behavior. The query population is therefore neither a neutral sample of information needs nor an unconstrained conversational distribution. The retrieved passages are also conditioned on Bing’s document ranking and passage-extraction systems.

A major issue is incomplete relevance annotation. Editors saw a shallow candidate pool, typically approximately ten passages, generated by a single retrieval system. Queries marked “No Answer Present” were excluded from commonly used passage- and document-ranking datasets. A study of MS MARCO v1 reported that 38% of training queries and 45% of development queries were discarded in this way (Gupta et al., 2022).

Manual inspection of 150 discarded queries classified 106 as having no answer in the displayed pool, 20 as partial or wrongly labeled, and 4 as ill-formed. Among 50 well-formed discarded queries examined with BM25, ANCE, ColBERT, and monoT5, the systems found answers within their top ten results for 56–68% of the queries. This was described as a lower bound on answerability. The study concluded that approximately two thirds of sampled discarded queries could be answerable in the corpus despite being excluded from ranking datasets.

The discarded queries were not distributed uniformly across information needs. In development data, 46.7% of description queries and 47.2% of numeric queries were unanswered, compared with 28.2% of location queries. Thus, survivorship alters the query distribution, overrepresenting some query types and underrepresenting others.

Simulated annotation-depth experiments showed that reported RR@10 increased as more relevance judgments were included. For monoT5, the score increased from 0.332 under the shallowest condition to 0.365 with the full available qrels; BM25 increased from 0.174 to 0.187. The corresponding relative changes were approximately 9.0% for monoT5 and 7.5% for BM25. Training on more survivorship-biased subsets reduced performance by up to 9.9% on more completely annotated evaluations and by up to 3.5% in zero-shot transfer.

Another reproducibility problem concerns passage titles. Two materially different MS MARCO-passage corpora are used in the literature: the official passage-only corpus and an unofficial title-augmented corpus in which titles are prepended to passages using Title [[SEP](https://www.emergentmind.com/topics/semantic-entropy-production-sep-metric)] Passage (Lassance et al., 2023). Approximately 5.7 million of 8.8 million passages, or 64.5%, have associated titles. Titled passages constitute 73.4% of query positives but 64.5% of the corpus, creating a relevance-correlated feature.

On the development set, title augmentation improved DPR-CLS MRR@10 from 33.8 to 35.3 and SPLADE MRR@10 from 37.7 to 38.5. Cross-encoder reranking also improved: monoElectra-large increased from 40.3 to 41.2 at top ten, and RankT5-3b increased from 43.0 to 44.2. These gains did not reliably transfer to TREC 2019, TREC 2020, or BEIR. The title-augmented corpus should therefore be treated as a separate benchmark condition, not as an interchangeable implementation of official MS MARCO passage ranking.

MS-Shift introduced controlled query-side distribution shifts while keeping the document collection, task, and relevance judgments fixed (Lupart et al., 2022). The shifts covered semantic topics, question-word/query intent, and query length. Single-vector dense retrieval was generally most affected, whereas ColBERT was the most consistently robust among the evaluated neural systems. The “how” intent group was particularly difficult: the bi-encoder lost 24.8% relative MRR, SPLADE 26.8%, ColBERT 14.0%, and monoBERT 13.5% between in-domain and out-of-domain conditions.

A later study reported that fine-tuning can degrade an already heavily pretrained MS MARCO passage-ranking model (Pande et al., 23 Jun 2025). The unmodified sentence-transformers/all-MiniLM-L6-v2 baseline achieved MRR@10 of 0.3026. Full fine-tuning with random negatives achieved 0.2619, full fine-tuning with hard negatives 0.2536, LoRA with random negatives 0.2557, and LoRA with hard negatives 0.2050. The authors associated the degradation with changes in embedding-space structure, but the conclusions are limited to the tested architecture, triplet objective, training scale, and evaluation protocol.

5. Multilingualization and cross-language retrieval

The scarcity of large labeled retrieval collections outside English led to machine-translated extensions of MS MARCO. mMARCO contains translations into 13 non-English languages: Spanish, French, Italian, Portuguese, Indonesian, German, Russian, Chinese, Japanese, Dutch, Vietnamese, Hindi, and Arabic (Bonifacio et al., 2021). The English data remain as the original-language reference.

mMARCO translates approximately 8.8 million passages and 530,000 queries while preserving English passage-level relevance labels and query–passage associations. Helsinki/OPUS-MT and Google Translate were evaluated. The complete translation required 79.58 hours for passages and 8.19 hours for queries with Helsinki, and 78.59 hours for passages and 1.01 hours for queries with Google Translate.

The retrieval pipeline uses BM25 followed by neural reranking. On original English, BM25 obtained MRR@10 of 0.184, while both mT5 and mMiniLM obtained 0.366. mMiniLM has approximately 107 million parameters compared with approximately 580 million for mT5, or approximately 5.4 times fewer parameters, while remaining competitive.

Translation quality strongly affected retrieval. Average MRR@10 for BM25 was .105 with Helsinki translation and .138 with Google translation; mMiniLM obtained .211 and .274; mT5 obtained .217 and .281. The relationship between translation quality and retrieval effectiveness was positive but weak to moderate, with reported R2≈0.33R^2 \approx 0.33.

mMARCO-trained models also transferred to Mr. TyDi without using Mr. TyDi training data. Multilingual mT5 achieved average MRR@100 of .551, compared with .532 for mT5 fine-tuned only on English MS MARCO. Multilingual training improved performance in several languages absent from mMARCO fine-tuning, including Bengali, Finnish, Japanese, Korean, and Swahili.

ColBERT-X used MS MARCO as English relevance supervision for cross-language information retrieval (Nair et al., 2022). Its zero-shot configuration trains on English triples and relies on XLM-R representations, while translate-train retains English queries and labels but replaces English passages with machine translations. Both methods generally outperformed BM25 with machine-translated queries on Chinese, Persian, French, German, Italian, Russian, and Spanish news collections. Translate-train usually improved over zero-shot on CLEF languages, though not significantly for Russian and not significantly on the HC4 collections.

An Urdu extension translated MS MARCO with IndicTrans2 and fine-tuned mT5 on the resulting triples (Butt et al., 2024). Urdu BM25 achieved Recall@10 of 0.268 and MRR@10 of 0.129; zero-shot mMARCO achieved 0.408 and 0.204; Urdu mT5-mMARCO achieved 0.438 in the main table and 0.247, respectively. The abstract reports Recall@10 of 0.439. These results demonstrate the utility of translated supervision but also inherit translation errors, English-centered relevance labels, and the absence of systematic human validation.

6. Extensions, applications, and later web-scale resources

MS MARCO has been extended beyond plain text ranking. MMEAD provides entity links from MS MARCO v1 and v2 documents and passages to Wikipedia using REL and BLINK (Kamphuis et al., 2023). It stores entity identifiers, character offsets, entity names, and linker-specific details. REL generated 18,561,221 links for v1 passages and 145,725,732 for v1 documents; v2 contained 233,254,024 passage links and 661,183,287 document links. For v1 passages, BLINK produced 21,968,356 links.

Entity-aware BM25 expansion improved candidate recall, particularly for difficult queries. On the full entity-bearing development subset, BM25 without expansion achieved Recall@1000 of 0.9111; entity-text expansion achieved 0.9183; and reciprocal-rank fusion of the unexpanded and entity-text runs achieved 0.9338. Gains increased on the hard, harder, and hardest Chameleon subsets. Improvements in MRR@10 were smaller and inconsistent, illustrating the distinction between candidate recall and top-rank ordering.

MS MARCO Web Search broadens the setting from passage ranking and question answering to multilingual, web-scale document retrieval with real click labels (Chen et al., 2024). It uses ClueWeb22, described as containing approximately 10 billion web pages in 207 languages with URLs, language and topic tags, titles, clean text, raw HTML, rendered visual representations, and semantic annotations. The release includes Set-100M and Set-10B resources, approximately 10 million queries in 93 languages, and click-derived query–document associations sampled from one year of Bing logs.

The resource uses time-based splitting. In the reported Set-100M test data, 7,646 of 9,374 query–document pairs involve both a query and document outside the training period, approximately 81.6%. Only 7.77% of documents have at least one relevant labeled query, 0.46% have more than one relevant query, and 1.4% of queries have multiple relevant documents.

On exact search, DPR achieved MRR@10 of 0.542, ANCE 0.633, and SimANS 0.649. In end-to-end retrieval with SPANN, BM25 achieved 0.296, DPR 0.467, ANCE 0.580, and SimANS 0.585. The dense systems reached 625 QPS in the reported setup, compared with 149 QPS for Elasticsearch BM25, although throughput and latency depend on the hardware and implementation. Approximate indexing introduced a substantial quality gap: SimANS recall@100 fell from 91.98% under exact search to 79.82% with SPANN.

MS MARCO’s influence therefore extends across several research paradigms:

  • Question answering and answer generation: natural-language synthesis, answerability, and multi-passage reasoning.
  • Neural information retrieval: dense, sparse, late-interaction, cross-encoder, and hybrid ranking.
  • Large-data representation learning: supervised training over millions of query–passage relationships.
  • Multilingual and cross-language retrieval: translated supervision, zero-shot transfer, and language-specific adaptation.
  • Entity-aware and neuro-symbolic retrieval: Wikipedia links, entity expansion, and graph-mediated search.
  • Evaluation methodology: significance testing, leaderboard stability, incomplete judgments, distribution shift, and reproducibility.
  • Web-scale search systems: approximate nearest-neighbor indexing, throughput, latency, and joint model–index optimization.

MS MARCO remains a powerful but conditional benchmark. Its scale, real search behavior, noisy web evidence, sparse labels, and public infrastructure enabled substantial progress in neural retrieval. At the same time, its results depend on exact corpus versions, annotation depth, candidate-generation procedures, title handling, pretraining overlap, language, translation quality, and evaluation metric. Claims of improvement are therefore meaningful only when these conditions are reported and held constant.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MS MARCO.