---
title: 'TREC Deep Learning: Full Method'
url: https://www.emergentmind.com/topics/trec-deep-learning
type: topic
---

# TREC Deep Learning: Full Method

TREC Deep Learning is a Text REtrieval Conference evaluation program for ad hoc document and passage ranking in a large-data regime. Established in 2019, it combines MS MARCO-derived collections, hundreds of thousands of human-generated training queries, neural and traditional retrieval systems, fixed-candidate reranking, end-to-end full retrieval, and blind TREC-style relevance assessment. Its research scope has expanded from comparing pretrained neural rankers with BM25-family methods to studying sparse and dense retrieval, multi-stage ranking, query and document expansion, test-collection reusability, hard-query evaluation, and, by 2023, prompt-based large-language-model ranking [2003.07820].

## 1. Origin, objectives, and task structure

The first TREC Deep Learning track was introduced in 2019 to study ad hoc retrieval when large human-labeled training resources are available. Its design addressed an established concern in neural information retrieval: earlier comparisons often relied on small, proprietary, synthetic, or weakly supervised datasets, making it difficult to determine whether neural methods were intrinsically ineffective or simply data-hungry. The track therefore combined large MS MARCO-derived training sets with blind, one-shot TREC evaluation and reusable test collections [2003.07820].

The track has two principal tasks:

1. **Passage retrieval**, in which systems rank relatively short passages intended to answer a query.
2. **Document retrieval**, in which systems rank substantially longer Web documents.

Each task has two operating modes:

- **Full retrieval**: the system searches the entire corpus and produces its own candidate and final ranking.
- **Reranking**: the system receives a fixed candidate set and reorders it.

The original candidate sets contained 100 documents and 1,000 passages per query. Later editions reduced the official reranking candidate set to 100 items. This distinction isolates two different capabilities: candidate generation over millions or hundreds of millions of items, and fine-grained relevance estimation over a controlled candidate set.

The document and passage tasks use different relevance semantics. A document may be relevant if it provides substantial or even minimal useful information, whereas a passage is binary-relevant only when it is highly or perfectly relevant and contains an answer. Consequently, raw scores across the two tasks should not be interpreted as directly comparable.

## 2. Collections, supervision, and evaluation protocol

The underlying supervision originates in MS MARCO, whose queries were sampled from Bing search activity and associated with crowd-selected answer passages. Training labels are large in quantity but sparse in structure: they often provide one positive item per query and generally do not provide explicit negatives. Document labels were initially transferred from relevant passages to their source documents, creating a possible mismatch between passage-level evidence and document-level usefulness [2104.09399].

The v1 collections contained approximately 3.2 million documents and 8.8 million passages. In 2021, MS MARCO v2 refreshed the collections: the document collection grew to approximately 11.9 million documents and the passage collection to approximately 138 million passages. The refresh improved document extraction, character encoding, whitespace handling, and passage-to-document mappings, but it also created difficulties in transferring legacy labels to new documents and passages [2507.08191].

The principal TREC evaluation metric is graded NDCG@10:

$$
\mathrm{NDCG}@k=\frac{\mathrm{DCG}@k}{\mathrm{IDCG}@k},
$$

with discounted cumulative gain conventionally defined as:

$$
\mathrm{DCG}@k=\sum_{i=1}^{k}\frac{2^{rel_i}-1}{\log_2(i+1)}.
$$

Other reported measures include average precision, reciprocal rank, normalized cumulative gain at candidate depth, and precision at ten. NCG@100 for documents and NCG@100 or NCG@1000 for passages is intended to characterize the quality of the retrieved candidate set independently of its precise internal ordering.

TREC assessment differs from sparse MS MARCO evaluation in several important respects:

- **MS MARCO evaluation** uses sparse binary labels and often a single known positive.
- **TREC evaluation** uses new queries, blind submission, pooled results, four-level relevance judgments, and more comprehensive assessment.
- **Document and passage binarization differ**: all non-irrelevant document grades count as relevant, whereas only highly and perfectly relevant passages count as binary-relevant.
- **NIST judgments are not interchangeable with MS MARCO labels**: agreement between sparse MS MARCO reciprocal rank and NIST NDCG declined as the collections evolved, reaching Kendall’s $\tau=0.43$ for documents and $\tau=0.51$ for passages in 2021 [2507.08191].

The evaluation procedure uses pooling and active judging. Early editions supplemented shallow run pools with HiCAL, a high-recall system that selected additional unjudged items using judgments collected so far. Later editions used continuous active learning, near-duplicate clustering, canonical passage judgments, and label propagation. These mechanisms were designed to improve judgment completeness and make the collections reusable.

## 3. Development from 2019 through 2023

The track’s system categories have generally been:

- **`trad`**: traditional retrieval or learning-to-rank methods, including BM25 and pseudo-relevance feedback.
- **`nn`**: neural representation-learning methods without large pretrained language models.
- **`nnlm`**: systems using pretrained models such as BERT, XLNet, or related architectures.
- **`prompt`**: systems using prompted large language models in some part of the ranking pipeline.

In 2019, 15 groups submitted 75 runs. In 2020, participation increased to 25 groups and 123 runs. In 2021, 19 groups submitted 129 runs, while in 2022, 14 groups submitted 142 runs. The proportion of `nnlm` systems increased from 44% in 2019 to 76% in 2021 and approximately 85% in 2022; the standalone `nn` category disappeared in 2021 [2507.10865].

The empirical progression was consistent in its broad direction. In 2019 and 2020, pretrained language-model systems substantially outperformed traditional retrieval. In 2021, the best `nnlm` document run exceeded the best traditional run by approximately 15% in NDCG@10, and the best passage run exceeded the best traditional run by approximately 36%. In 2022, the corresponding gaps increased to approximately 76% for documents and 125% for passages, although cross-year percentage comparisons are confounded by changes in collections, query selection, and judgments [2507.10865].

The 2023 edition was the fifth and final year of the track. It retained the larger v2 collections and held-out query methodology but introduced synthetic queries generated by T5 and GPT-4. Prompt-based systems became the strongest category: the best passage run, `naverloo-rgpt4 (b)`, achieved NDCG@10 of 0.6994, compared with 0.5972 for the strongest `nnlm` passage run. The best document prompt run achieved NDCG@10 of 0.6893 [2507.08890].

## 4. Model architectures and retrieval paradigms

TREC Deep Learning has represented several distinct neural retrieval paradigms.

### Cross-encoder and reranking models

BERT-style cross-encoders jointly process a query and candidate text:

$$
[\mathrm{CLS}]\;q\;[\mathrm{SEP}]\;d\;[\mathrm{SEP}].
$$

A classification or regression head then produces a relevance score. These models offer unrestricted query-document token interaction but are computationally expensive, so they are normally applied after sparse or dense candidate generation. Brown University used query expansion followed by BERT Large passage reranking, ranking third overall and second among reranking submissions in the 2019 passage task [2009.04016].

### Local and distributed matching

The Duet architecture combines a local branch for exact or near-exact lexical matching with a distributed branch for semantic matching. DuetMF extends this approach to structured documents by modeling URL, title, and body with separate parameters and combining vector-valued field representations. Its standalone document run achieved MRR 0.810 and NDCG@10 0.533, while an LTR system combining DuetMF with SDM, PRF, BM25, DESM, query-length, and domain-quality features achieved MRR 0.876 and NDCG@10 0.578 [1912.04471].

### Kernel-based contextual matching

TK, or Transformer-Kernel, combines static embeddings, a lightweight Transformer contextualizer, cosine-similarity interactions, and Gaussian kernel pooling. Query and document terms are contextualized independently, producing a single query-document interaction matrix. Kernel pooling converts similarity values into soft histogram features, while logarithmic and document-length-normalized aggregation provides an inspectable scoring decomposition. TK achieved 0.420 MAP, 0.671 nDCG, and 0.598 P@10 in its best 2019 passage full-ranking run [1912.01385].

Conformer-Kernel replaces Transformer contextualization with Conformer layers and introduces query term independence (QTI). QTI factorizes the score as:

$$
S(q,d)=\sum_{i=1}^{|q|}s(q_i,d).
$$

This permits document-side contributions to be precomputed and retrieved through an inverted index. Explicit lexical matching and ORCAS click-derived document text further improve the model. The best 2020 run, `ndrm3-orc-full`, achieved NDCG@10 0.6249, NCG@100 0.6764, AP 0.4280, and RR 0.9444 [2011.07368].

### Sparse learned retrieval

SPLADE produces sparse lexical and expansion representations that remain compatible with inverted-index retrieval. TREC 2022 systems from Naver Labs Europe combined SPLADE++ models with Rocchio feedback, ColBERTv2, DocT5, and an ensemble of cross-encoder rerankers trained using SPLADE-selected hard negatives. Their best passage full-ranking run achieved NDCG@10 0.7145 and mAP@100 0.2950 [2302.12574].

### Dense and hybrid retrieval

Dense retrievers encode queries and documents or passages into vector representations, enabling approximate-nearest-neighbor retrieval. ColBERT uses late interaction:

$$
S_{\mathrm{ColBERT}}(q,p)=\sum_{i\in q}\max_{j\in p}\mathbf{q}_i^{\mathsf T}\mathbf{p}_j.
$$

TREC systems increasingly combine dense retrieval with BM25, doc2query, SPLADE, or other sparse methods. Alibaba’s 2022 system combined BM25, Doc2query, SPLADE, and dense retrieval, followed by full-interaction neural ranking and the lightweight HLATR fusion module. Its passage run achieved NDCG@10 0.7184, and its document run achieved 0.7533 [2308.12039].

### Generative and prompted ranking

PASH’s 2021 system added T5 to a multistage pipeline using sparse docTTTTTquery retrieval, ColBERT, pointwise ranking, pairwise ranking, continual pretraining, and transformer ensembles. T5 produced higher NDCG@5 but lower NDCG@10 than the transformer-based ranker, suggesting sharper early ranking but less stable performance beyond the top positions [2205.11245].

In 2023, prompted LLMs were used as ranking or reranking components. The strongest systems generally retained traditional or pretrained neural candidate generation and applied prompting only to a manageable candidate set. Thus, the `prompt` category did not represent exhaustive LLM retrieval over the entire collection; it typically represented a hybrid multistage architecture.

## 5. Full retrieval, reranking, and multi-stage ranking

A persistent distinction in the track is between candidate-set quality and candidate ordering. A reranker can achieve strong top-ranked quality but cannot recover material absent from its fixed candidate list. A full-ranking system must balance scalable retrieval with accurate relevance estimation.

From 2019 through 2021, differences between the best full-ranking and reranking systems were generally modest. In 2021, full-ranking improved NDCG@10 over reranking by approximately 4% for documents and 6% for passages. The strongest single-stage systems nevertheless remained behind multistage pipelines: approximately 6% behind the best document run and 10% behind the best passage run [2507.08191].

In 2022, the gap increased sharply. The best full-ranking passage run exceeded the best reranking run by 36% in NDCG@10, and the document difference was 125%. Possible explanations include harder held-out queries, reduced effort devoted to reranking, the smaller fixed candidate set, and genuine improvements in end-to-end retrieval [2507.10865].

The results do not imply that dense retrieval alone is sufficient. In 2022, the strongest systems often used multistage sparse neural retrieval or hybrid retrieval, and single-stage dense systems were less competitive than in 2021. The strongest pipelines generally followed a pattern of:

1. sparse, dense, or hybrid candidate generation;
2. hard-negative-trained cross-encoder reranking;
3. score fusion, pairwise refinement, or lightweight final reranking;
4. document aggregation from passage scores when necessary.

Document ranking frequently used max-passage aggregation:

$$
S(d,q)=\max_{p\in P(d)}S(p,q).
$$

This approach simplifies long-document ranking but may overlook evidence distributed across multiple passages. It also makes document results dependent on passage retrieval and passage-level judgments.

## 6. Reusability, difficulty, and methodological controversies

### Test-collection reusability

The track treats evaluation design as a research problem. The 2019 and 2020 collections were shown to have relatively stable system rankings under simulated omission of participating runs. In a separate study, new pools created using BM25 and modern transformer-based systems were added to the TREC-8 collection. The original and expanded qrels produced almost identical rankings across 134 runs, with Kendall’s $\tau=0.9933$ for MAP and $\tau=0.9991$ for P@10. This supports the conditional claim that a well-constructed traditional collection can evaluate modern neural systems, although the result depends on deep, diverse, and effective original pooling [2201.11086].

The larger MS MARCO v2 collections created a different challenge. Collection size increased far faster than judging budgets, producing incomplete judgments and metric saturation. TREC 2022 addressed this through held-out queries, passage-focused assessment, near-duplicate clustering, active learning, and label propagation. All 76 accepted topics met the intended relevance-density criterion, and the resulting passage collection was considered more reusable than the 2021 collection [2507.10865].

### Development-set leakage and selection bias

Reusable TREC collections can become invalid if repeatedly used for model selection. The recommended protocol is to use MS MARCO development data for architecture, hyperparameter, negative-sampling, training-duration, and checkpoint decisions, and to reserve TREC test sets for final evaluation. A checkpoint case study involving 32 alternatives demonstrated that test-driven selection could produce a statistically significant but substantively meaningless difference with $p=0.00837$ [2104.09399].

Multiple seeds, confidence intervals, checkpoint-selection procedures, and explicit disclosure of test-set contact are therefore important. A result selected from many test-evaluated variants should not be presented as an unbiased held-out estimate.

### Hard-query evaluation

DL-HARD was introduced because many original TREC DL topics were relatively easy factoid questions, leaving limited headroom for neural ranking. It selects 50 complex topics from the 2019 and 2020 benchmarks, emphasizing list, long-answer, reasoning, entity-rich, and multi-document information needs. On DL-HARD, mean NDCG@10 fell by approximately 21%, RR by approximately 23%, and Recall@1000 by approximately 19–20%; system rankings also changed substantially, with average movement of 4.6 positions [2105.07975].

DL-HARD is therefore complementary to the original collections. The original benchmarks provide broad comparability and deep assessment, whereas DL-HARD stresses capabilities obscured by easy factoid queries.

### Label and collection bias

The transition to MS MARCO v2 introduced an “oldness” artifact. Only legacy documents could inherit positive sparse MS MARCO labels, potentially encouraging systems to prefer old URLs even when new documents were relevant. In 2021, system-ranking agreement between sparse reciprocal rank and NIST NDCG increased from $\tau=0.43$ to $\tau=0.56$ when evaluation was restricted to the old-document universe [2507.08191].

This illustrates a broader principle: model comparisons are inseparable from collection construction, label transfer, query sampling, and metric interpretation. High sparse-label performance does not necessarily imply broad relevance under independently assessed NIST judgments.

## 7. From neural ranking to RAG and future retrieval evaluation

The later evolution of TREC-related research extends ranked retrieval toward retrieval-augmented generation. The Ragnar framework for the TREC 2024 RAG Track uses a pipeline of BM25 retrieval, RankZephyr reranking, selection of the top 20 segments, and generation with GPT-4o or Command R+. Its output contains an ordered evidence list, sentence-level answers, citations, and response length [2406.16828].

This setting changes the evaluation target:

$$
\text{query}\rightarrow\text{ranked passages}
$$

becomes:

$$
\text{query}\rightarrow\text{retrieved evidence}\rightarrow\text{generated, cited answer}.
$$

Retrieval remains essential, but end-to-end quality additionally depends on completeness, coherence, factual correctness, citation support, and answer organization. A relevant candidate set can still yield a poor answer if the generator ignores evidence, produces unsupported claims, or fails to synthesize multiple sources.

The TREC Deep Learning trajectory therefore comprises several successive evaluation emphases:

- large-scale supervised neural ranking;
- pretrained language-model reranking;
- scalable sparse and dense retrieval;
- multistage and hybrid architectures;
- difficult and reusable test collections;
- prompted LLM ranking;
- evidence-grounded answer generation.

The central technical conclusion is not that one architecture replaces all others. Strong systems combine complementary mechanisms: lexical retrieval for exact matching and scalable recall, semantic retrieval for vocabulary mismatch, cross-encoders for detailed relevance estimation, learned sparse representations for indexability, reranking ensembles for robustness, and generative models for synthesis. The methodological conclusion is equally important: claims about retrieval progress require blind, independently judged, sufficiently complete, and responsibly reused test collections.

Source: https://www.emergentmind.com/topics/trec-deep-learning