Papers
Topics
Authors
Recent
Search
2000 character limit reached

SPLADE-v3 Learned Sparse Retriever

Updated 26 September 2026
  • SPLADE-v3 is a revised training pipeline for the SPLADE architecture, featuring multiple hard negatives, an ensemble of cross-encoder teachers, and combined distillation objectives, among others.
  • The updated model achieved a 29.6 MRR@10 increase over SPLADE++, which consists of 40.2 nDCG@10 on the MS MARCO development set, which shows a significant improvement over earlier models.
  • The results of the evaluation are average, showing that using query expansion terms and documents fetus queries is beneficial in general.

SPLADE-v3 is a family of learned sparse retrievers introduced with a new release of the SPLADE library. It preserves the SPLADE architecture—Transformer encoding, vocabulary-space expansion, sparse aggregation, and inverted-index retrieval—but improves the training pipeline through multiple hard negatives, an ensemble of cross-encoder teachers, teacher-score calibration, and joint KL-divergence and MarginMSE distillation. The principal model reports 40.2 MRR@10 on the MS MARCO development set and 51.7 mean nDCG@10 over 13 BEIR datasets, with statistically significant aggregate gains over BM25 and SPLADE++SelfDistil (Lassance et al., 2024).

1. Position within learned sparse retrieval

SPLADE models are learned sparse retrievers whose representation dimensions correspond to vocabulary terms. A Transformer encodes a query or document, contextual token representations are projected into vocabulary-sized scores, and the resulting activations are aggregated into a sparse vector. Retrieval uses a sparse inner product evaluated through an inverted index.

SPLADE-v3 is not a new retrieval architecture. Its principal contribution is a revised training structure applied to the established SPLADE mechanism. The report characterizes the model as a stronger and more reliable baseline than earlier SPLADE and SPLADE++ systems. Its improvements arise primarily from:

  • multiple hard negatives per query;
  • an ensemble of five cross-encoder teachers;
  • affine calibration of ensemble scores;
  • joint KL-divergence and MarginMSE distillation;
  • initialization from SPLADE++SelfDistil.

The method therefore occupies an intermediate position between traditional lexical retrieval and dense neural retrieval. Like BM25, it uses vocabulary-indexed postings and exact coordinate overlap. Like dense retrievers, it uses contextual neural representations to infer semantic relations. A contextual token may activate vocabulary terms that are not literally present in the input, enabling lexical expansion while retaining inverted-index compatibility.

SPLADE-v3 should be distinguished from later decoder-only approaches such as Echo-Mistral-SPLADE. The latter replaces the BERT-style encoder with a Mistral-7B decoder-only model and reports 55.07 mean nDCG@10 on a 13-dataset BEIR comparison, versus 52.07 for SPLADE-v3 in that paper’s table; however, the comparison uses different training data and procedures and is not a fully controlled reproduction (Doshi et al., 2024).

2. Sparse representation and retrieval mechanism

Let x=(x1,…,x∣x∣)x=(x_1,\ldots,x_{|x|}) denote a query or document, let VV be the model vocabulary, and let Hi∈RdH_i\in\mathbb{R}^d be the contextual representation of token xix_i. A vocabulary projection with parameters W∈R∣V∣×dW\in\mathbb{R}^{|V|\times d} and b∈R∣V∣b\in\mathbb{R}^{|V|} produces token-level logits

zi=WHi+b.z_i = WH_i+b.

SPLADE applies a rectified linear unit and logarithmic saturation, then aggregates the evidence for each vocabulary term with a maximum over sequence positions:

fSPLADE(x)j=max⁡i∈{1,…,∣x∣}log⁡(1+ReLU⁡(zi,j)).f_{\mathrm{SPLADE}(x)_j} = \max_{i\in\{1,\ldots,|x|\}} \log\left(1+\operatorname{ReLU}(z_{i,j})\right).

The query and document representations are sparse vocabulary vectors,

q=fSPLADE(q),d=fSPLADE(d).\mathbf q=f_{\mathrm{SPLADE}(q)},\qquad \mathbf d=f_{\mathrm{SPLADE}(d)}.

Their retrieval score is

s(q,d)=q⊤d=∑j∈Vqjdj.s(q,d)=\mathbf q^\top\mathbf d =\sum_{j\in V}q_jd_j.

Only vocabulary coordinates with nonzero weights contribute to the score. Each coordinate can therefore be represented by a posting list, and query evaluation accumulates weighted products between query and document terms.

The representation has two related properties. It is lexical because its coordinates correspond to vocabulary terms, and it is expansive because contextual token representations may activate semantically or morphologically related terms absent from the original text. This permits semantic flexibility beyond BM25 without abandoning the operational structure of a lexical inverted index.

Sparsity is encouraged through a FLOPS-inspired regularizer. For a batch VV0, a conventional form is

VV1

The regularizer penalizes vocabulary dimensions that are frequently active across a batch. Its intended effect is to reduce average expansion and the number of posting-list operations during retrieval. The supplied SPLADE-v3 report states that other hyperparameters are similar to previous SPLADE iterations and does not specify a new sparsity formulation or coefficient.

3. Model variants and architectural choices

SPLADE-v3 uses a Transformer encoder followed by vocabulary projection and sparse aggregation. The report does not introduce a new encoder architecture or sparse representation. Its principal architectural choices concern initialization, model size, and query-side computation.

SPLADE-v3

The principal model:

  • is initialized from SPLADE++SelfDistil;
  • uses the full BERT-sized configuration associated with that checkpoint;
  • retains query expansion;
  • uses KL-divergence and MarginMSE distillation;
  • is trained with hard negatives.

The exact Transformer layer count, hidden size, sequence lengths, optimizer settings, and regularization coefficients are not specified in the supplied report; they are described as similar to previous SPLADE releases.

SPLADE-v3-DistilBERT

SPLADE-v3-DistilBERT starts from DistilBERT rather than the BERT-based SPLADE++SelfDistil checkpoint. It targets a smaller inference footprint. It is weaker than the BERT-based SPLADE-v3 but remains stronger than the lexical variant on BEIR.

SPLADE-v3-Lexical

SPLADE-v3-Lexical removes query expansion. This lowers the number of query-side vocabulary terms and reduces retrieval computation. It remains highly effective on MS MARCO and LoTTE but is noticeably weaker on out-of-domain BEIR. The results support the interpretation that query expansion is particularly important for domain transfer.

SPLADE-v3-Doc

SPLADE-v3-Doc starts from CoCondenser and performs no learned computation for the query. It can be viewed as using a simple binary bag-of-words-like query mechanism while retaining learned sparse document weights. It is the weakest released variant overall, particularly in zero-shot evaluation, although it remains competitive with dense bi-encoders when efficiency advantages are considered.

Variant Principal design Reported trade-off
SPLADE-v3 BERT-sized, query expansion, dual distillation Highest reported effectiveness
SPLADE-v3-DistilBERT DistilBERT backbone Smaller inference footprint, lower effectiveness
SPLADE-v3-Lexical No query expansion Lower computation, weaker BEIR transfer
SPLADE-v3-Doc No learned query computation Minimal query-side computation, weakest effectiveness

The variants make explicit that query computation and query expansion are central components of SPLADE-v3’s effectiveness. Removing query expansion reduces FLOPS but disproportionately harms out-of-domain performance, whereas removing query computation produces a larger quality decline.

4. Training pipeline

Hard-negative construction

Earlier training structures often used one hard negative per query. SPLADE-v3 supports multiple negatives in a batch, following the general direction of Tevatron-style dense retrieval training.

The base SPLADE-v3 model uses eight negatives per query sampled from SPLADE++SelfDistil. A broader hard-negative generation and distillation-score pipeline considers 100 negatives per query: 50 selected from the top 50 retrieved documents and 50 sampled randomly from the top 1,000 retrieved documents.

The report states that increasing the number of negatives improves results, particularly for in-domain MS MARCO evaluation, while contributing less to out-of-domain generalization. The negatives are produced by a SPLADE++ model rather than BM25 alone, making them more closely aligned with the failure modes of learned sparse retrieval.

Cross-encoder teacher ensemble

SPLADE++ relied primarily on scores from a single cross-encoder, especially cross-encoder/ms-marco-MiniLM-L-6-v2. SPLADE-v3 uses an ensemble of five cross-encoders:

  1. cross-encoder/ms-marco-MiniLM-L-6-v2;
  2. naver/trecdl22-crossencoder-rankT53b-repro;
  3. naver/trecdl22-crossencoder-debertav3;
  4. naver/trecdl22-crossencoder-debertav2;
  5. naver/trecdl22-crossencoder-electra.

For approximately 500,000 MS MARCO training queries, the teachers score the positive passage or passages and 100 associated negatives. Scores are normalized per query using the min-max aggregation procedure from ranx.

The authors also construct a rescored version by applying an affine transformation so that the ensemble scores have a mean and standard deviation similar to those of the earlier single-teacher scores. The report states that modifying the score distribution improves distillation, particularly MarginMSE, but does not establish a theoretical explanation.

Distillation losses

SPLADE-v3 combines two distillation objectives:

  • KL-divergence, associated with precision-oriented behavior;
  • MarginMSE, associated with recall-oriented behavior when many negatives are available.

For a positive document VV2 and negative document VV3, the conventional MarginMSE term is

VV4

KL-divergence compares teacher and student distributions over a candidate set. The report describes the combined objective conceptually as

VV5

with

VV6

The sparsity coefficient is not given in the supplied report. According to the authors’ empirical interpretation, MarginMSE emphasizes recall, KL-divergence emphasizes precision, and their combination produces better overall effectiveness.

Initialization

Starting from SPLADE++SelfDistil performs better than starting from CoCondenser or DistilBERT for the full model. The authors speculate that this may reflect a curriculum effect in which earlier SPLADE++ training provides a useful initial retrieval structure before the new negative and distillation procedures are applied. This is presented as a hypothesis rather than a demonstrated mechanism.

5. Empirical evaluation

MS MARCO

The experiments use the original MS MARCO passage collection without titles. The report distinguishes this setting from versions in which titles are appended to passages, because title augmentation can leak relevance information and complicate comparisons. Performance is reported using MRR@10 on the development set.

Model MS MARCO MRR@10
SPLADE++SelfDistil 37.6
SPLADE-v3 40.2
SPLADE-v3-DistilBERT 38.7
SPLADE-v3-Lexical 40.0
SPLADE-v3-Doc 37.8

SPLADE-v3 improves over SPLADE++SelfDistil by 2.6 MRR points. The lexical variant remains close to the full model in this in-domain evaluation, whereas SPLADE-v3-Doc is weaker.

TREC Deep Learning

The report evaluates TREC 2019 and TREC 2020 Deep Learning using nDCG@10.

Model TREC19 nDCG@10 TREC20 nDCG@10
SPLADE++SelfDistil 73.0 71.8
SPLADE-v3 72.3 75.4
SPLADE-v3-DistilBERT 75.2 74.4
SPLADE-v3-Lexical 71.2 73.6
SPLADE-v3-Doc 71.5 70.3

SPLADE-v3 is slightly below SPLADE++SelfDistil on TREC19 but substantially higher on TREC20. DistilBERT is strongest among the listed systems on TREC19, while the full SPLADE-v3 model leads on TREC20.

BEIR

The paper evaluates 13 BEIR datasets using nDCG@10.

Dataset SPLADE++SelfDistil SPLADE-v3 DistilBERT Lexical Doc
ArguAna 51.8 50.9 48.4 52.7 46.7
Climate-FEVER 23.7 23.3 22.8 21.8 15.9
DBPedia-entity 43.6 45.0 42.6 42.8 36.1
FEVER 79.6 79.6 79.6 78.5 68.9
FiQA-2018 34.9 37.4 33.9 36.4 33.6
HotpotQA 69.3 69.2 67.8 68.5 66.9
NFCorpus 34.5 35.7 34.8 34.7 33.8
NQ 53.3 58.6 54.9 56.1 52.1
Quora 84.9 81.4 81.7 73.4 77.5
SCIDOCS 16.1 15.8 14.8 15.9 15.2
SciFact 71.0 71.0 68.5 71.5 68.8
TREC-COVID 72.5 74.8 70.0 63.6 68.1
Touché-2020 24.2 29.3 30.1 22.7 27.0
Average 50.7 51.7 50.0 49.1 47.0

SPLADE-v3 increases the mean from 50.7 to 51.7, described in the report as an approximately 2% out-of-domain improvement relative to SPLADE++SelfDistil. Gains are especially visible on NQ, Touché-2020, TREC-COVID, and FiQA-2018. The model declines on ArguAna, Climate-FEVER, HotpotQA, Quora, and SCIDOCS, and is essentially unchanged on FEVER and SciFact.

The results do not imply uniform improvement across domains. The Quora regression is particularly relevant in comparisons with SPLADE++SelfDistil, while the dataset-level variation shows that learned sparse expansion interacts with task-specific document and query characteristics.

LoTTE

LoTTE performance is reported as mean Success@5 for Search and Forum subsets.

Model LoTTE-S LoTTE-F
SPLADE-v3 74.7 66.0
SPLADE-v3-DistilBERT 70.3 62.8
SPLADE-v3-Lexical 74.2 64.5
SPLADE-v3-Doc 71.1 59.0

The lexical variant remains close to the full model on LoTTE, reinforcing the distinction between in-domain or task-specific effectiveness and broader zero-shot transfer.

6. Efficiency, comparisons, and limitations

At inference time, documents are encoded once into sparse vocabulary vectors and inserted into an inverted index. Queries are encoded into sparse vectors, and retrieval uses sparse inner-product accumulation. The report introduces no new index format or query-processing algorithm.

FLOPS is used as a loose efficiency indicator:

Model FLOPS indicator
SPLADE++SelfDistil 1.4
SPLADE-v3 1.2
SPLADE-v3-DistilBERT 1.4
SPLADE-v3-Lexical 0.6
SPLADE-v3-Doc 1.4

The lexical variant is cheaper because it removes query expansion. SPLADE-v3-Doc removes learned query encoding altogether but incurs a substantial effectiveness loss. Actual deployment cost also depends on query and document sparsity, posting-list lengths, vocabulary distribution, index implementation, dynamic pruning, and Transformer encoding time.

The paper evaluates reranking over the top 50 SPLADE-v3 results with MiniLM and DeBERTaV3. MiniLM produces an aggregate effect close to zero under a 95% confidence interval, whereas DeBERTaV3 generally improves results, with ArguAna as a major exception. These findings position SPLADE-v3 as a strong first-stage retriever rather than a universal replacement for high-capacity cross-encoder reranking.

A broader meta-analysis evaluates up to 44 query sets from MS MARCO, MS MARCO v2, BEIR, LoTTE, Antique, TREC-CAR, Natural Questions, TriviaQA, TREC-TB, and TREC-MQ. The analysis reports statistically significant aggregate gains over BM25 and SPLADE++SelfDistil. Only Webis Touché-2020 and the two TREC-MQ query sets show statistically significant decreases relative to BM25, while Quora is the only query set with a significant decrease relative to SPLADE++SelfDistil.

Several limitations qualify the interpretation:

  • SPLADE-v3 is principally a training and engineering improvement rather than an architectural departure.
  • The report does not provide a complete factorial ablation isolating teacher ensembling, score rescaling, negative count, loss mixture, and initialization.
  • The beneficial effect of score rescaling is empirically reported but not theoretically explained.
  • FLOPS is only a proxy for end-to-end efficiency and does not provide a complete latency, memory, or index-cost model.
  • Results vary across datasets, with notable regressions on Quora and several BEIR tasks.
  • DeBERTaV3 reranking generally improves the top-50 results.
  • The exact implementation details of the Transformer, optimizer, sequence lengths, and regularization coefficients are not fully specified in the supplied report.

SPLADE-v3 has also motivated systems-level and representation-level extensions. Sparton targets the vocabulary projection bottleneck by fusing matrix multiplication, masking, ReLU, log1p, and max reduction, avoiding materialization of the full sequence-by-vocabulary logit tensor; it reports up to a 4.8-fold isolated LM-head speedup and substantial memory reductions, but its end-to-end experiments use a related SPLADE configuration rather than an exact reproduction of the complete SPLADE-v3 recipe (Nguyen et al., 26 Mar 2026). SAE-SPLADE replaces the fixed lexical vocabulary with a learned sparse latent-concept vocabulary and reports retrieval performance comparable to SPLADE under a controlled reproduced baseline, while offering lower query-document FLOPs; its comparison with SPLADE-v3 is not fully matched because it does not reproduce the complete SPLADE-v3 training strategy (Zong et al., 23 Apr 2026).

SPLADE-v3’s principal significance is therefore methodological. It demonstrates that substantial improvements in learned sparse retrieval can result from stronger supervision, harder negatives, calibrated teacher ensembles, and complementary distillation losses while preserving the vocabulary-indexed sparse representation. The resulting model combines the operational advantages of inverted-index retrieval with semantic expansion and achieves stronger aggregate effectiveness than BM25 and earlier SPLADE baselines.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SPLADE-v3.