---
title: 'MassRET-20k: Benchmark for MS/MS Molecular Retrieval'
url: https://www.emergentmind.com/topics/massret-20k-dataset
type: topic
---

# MassRET-20k: Benchmark for MS/MS Molecular Retrieval

MassRET-20k is an independently curated evaluation dataset for assessing molecular structure retrieval from tandem mass spectrometry (MS/MS) data. Introduced to provide a rigorous, large-scale zero-shot benchmark, it comprises approximately 20,000 high-quality, one-to-one spectrum–molecule pairs, each reflecting a unique small-molecule structure obtained under systematically controlled ionization and experimental conditions. Designed to supplement the limitations in prior datasets, MassRET-20k emphasizes chemical diversity and comprehensive metadata essential for modern retrieval and generative modeling frameworks.

## 1. Dataset Construction and Composition

MassRET-20k is constructed by rigorous selection from the NIST2020 mass spectral library. All molecular structures overlapping with the MassSpecGym training split are excluded, resulting in a clean, non-overlapping evaluation set. Denoting the set of all spectra as $\mathcal{S}$ and corresponding unique molecular structures as $\mathcal{M}$, the dataset contains

\[
|\mathcal{S}| = |\mathcal{M}| \approx 20\,000.
\]

Each entry represents a single, unique spectrum (MS/MS experiment) paired to a molecular structure via canonical SMILES, all derived from the NIST2020 (non-open-source) corpus. No duplicate molecules or spectra are present, by construction. MassRET-20k provides only a held-out test set; there is no internal train/validation/test split.

## 2. Data Curation and Preprocessing

Strict protocols are employed for both molecular and spectral data:

- **Molecular Standardization:**  
  All SMILES strings undergo canonicalization, kekulization, and valence checks. Entries failing any of these are removed. Only molecules with one-to-one mapped InChIKeys are retained to ensure structural uniqueness and resolve tautomers and stereoisomers consistently.
  
- **Spectral Normalization:**  
  Fragment intensity arrays $\{I_j\}$ are rescaled as

  \[
  I_j \leftarrow \frac{I_j}{\max_k I_k},\quad I_j\in(0,1].
  \]
  Spectra lacking a peak above 1% relative intensity or featuring inconsistent or missing precursor metadata are excluded.

- **Exclusion and Inclusion Criteria:**  
  - Only spectra from NIST2020 are included.
  - All collision-energy annotations are present for 100% of entries.
  - Any molecule or spectrum present in MassSpecGym’s training set is excluded.

This ensures high-quality, standardized, and information-rich entries suitable for fair and interpretable zero-shot evaluation.

## 3. Structural and Spectral Diversity

The dataset achieves broad coverage across both chemical and experimental axes:

- **Ionization Adducts:**  
  MassRET-20k spans twelve distinct ionization adduct types (e.g., +H, +Na, +K, –H, etc.). This is in contrast to MassSpecGym, which covers only two adduct types and lacks comprehensive collision-energy annotations in 47% of entries.

- **Mass Range:**  
  Precursor (parent) mass $m$ values cover approximately 50–1200 Da:

  \[
  \int_{50}^{1200} f_m(m)\,dm = 1,
  \]
  where $f_m(m)$ is the parent-mass density.

- **Molecular Diversity:**  
  The distribution of pairwise Tanimoto coefficients (computed over Morgan fingerprints) for the $\sim$20,000 molecules is bell-shaped with a mean near 0.15:

  \[
  \mathbb{E}_{i\neq j}\bigl[T_\mathrm{Morgan}(M_i,M_j)\bigr]\approx0.15
  \]
  *This suggests substantial structural diversity, minimizing dataset redundancy and ensuring realistic assessment of generalization for retrieval models.*

## 4. Schema and Annotation Structure

Each data entry is annotated with comprehensive metadata, adhering to the following schema:

| Field Name              | Description                                 |                     |
|------------------------ |---------------------------------------------|---------------------|
| identifier              | Unique entry ID                             |                     |
| mzs                     | Array of fragment $m/z$ values              |                     |
| intensities             | Array of relative intensities               | $I_j$ normalized    |
| SMILES                  | Canonical SMILES string                     |                     |
| inchikey                | Corresponding InChIKey                      |                     |
| formula                 | Molecular formula (neutral)                 |                     |
| precursor\_formula      | Formula of the precursor ion                |                     |
| parent\_mass            | Exact mass of the molecule (Da)             |                     |
| precursor\_mz           | Precursor $m/z$                             |                     |
| adduct                  | Ionization adduct (12 types)                |                     |
| instrument\_type        | MS instrument used                          |                     |
| collision\_energy       | Normalized collision-induced dissociation   |                     |

All entries have complete metadata, and the annotation is designed for integration with downstream machine learning or cheminformatics pipelines.

## 5. Comparative Features and Evaluation Role

MassRET-20k is positioned as a standard, challenging, and diverse benchmark for zero-shot molecule retrieval from MS/MS. Compared to prior benchmarks, it introduces:

- **Twelve adduct types and full collision-energy annotation,** in contrast to the limited coverage and missing metadata in MassSpecGym.
- **Non-overlapping structure set with respect to training data,** avoiding information leakage and enabling robust assessment of generalization.
- **Uniformly rigorous curation criteria** both for molecules and spectra.

The dataset serves exclusively as a test set for models trained, for example, on MassSpecGym, and is intended to support realistic, out-of-sample evaluation of molecular retrieval methodologies for both discriminative and generative paradigms under zero-shot conditions.

## 6. Key Summary Statistics

A high-level overview of the dataset is as follows:

| Property                        | Value              |
|----------------------------------|--------------------|
| Total entries                    | 20,000             |
| Unique molecules                 | 20,000             |
| Source database                  | NIST2020           |
| Ionization adducts               | 12 types           |
| Collision-energy coverage        | 100%               |
| Precursor mass range             | 50–1200 Da         |
| Avg. pairwise Tanimoto (Morgan)  | 0.15               |

*Figures in the original paper include pie charts of adduct distribution, kernel density estimates of parent mass ($f_m(m)$), and violin plots of Tanimoto distributions for molecular diversity.*

## 7. Significance and Applications

MassRET-20k addresses the need for a high-quality, chemically diverse benchmarking set for cross-modal molecule retrieval tasks. Its comprehensive curation, detailed metadata, and exclusion of training-set molecules provide a robust foundation for zero-shot assessment of spectrometry-to-structure models, such as GLMR and other state-of-the-art generative frameworks. A plausible implication is that its use will help more sharply delineate genuine model generalization from overfitting or memorization effects, especially for tasks involving uncommon adducts or structurally novel molecules. As MS-based compound identification workflows become increasingly reliant on data-driven retrieval models, MassRET-20k offers a broadly representative, standardized testbed for comparative evaluation and methodological development.

Source: https://www.emergentmind.com/topics/massret-20k-dataset