Papers
Topics
Authors
Recent
Search
2000 character limit reached

MassRET-20k: Benchmark for MS/MS Molecular Retrieval

Updated 16 November 2025
  • MassRET-20k is a rigorously curated dataset of 20K unique MS/MS spectrum–molecule pairs for zero-shot molecular retrieval.
  • It employs strict molecular standardization and spectral normalization protocols to ensure high-quality, non-overlapping entries.
  • The dataset spans 12 ionization adducts and diverse molecular structures, providing a robust benchmark for both discriminative and generative models.

MassRET-20k is an independently curated evaluation dataset for assessing molecular structure retrieval from tandem mass spectrometry (MS/MS) data. Introduced to provide a rigorous, large-scale zero-shot benchmark, it comprises approximately 20,000 high-quality, one-to-one spectrum–molecule pairs, each reflecting a unique small-molecule structure obtained under systematically controlled ionization and experimental conditions. Designed to supplement the limitations in prior datasets, MassRET-20k emphasizes chemical diversity and comprehensive metadata essential for modern retrieval and generative modeling frameworks.

1. Dataset Construction and Composition

MassRET-20k is constructed by rigorous selection from the NIST2020 mass spectral library. All molecular structures overlapping with the MassSpecGym training split are excluded, resulting in a clean, non-overlapping evaluation set. Denoting the set of all spectra as S\mathcal{S} and corresponding unique molecular structures as M\mathcal{M}, the dataset contains

∣S∣=∣M∣≈20 000.|\mathcal{S}| = |\mathcal{M}| \approx 20\,000.

Each entry represents a single, unique spectrum (MS/MS experiment) paired to a molecular structure via canonical SMILES, all derived from the NIST2020 (non-open-source) corpus. No duplicate molecules or spectra are present, by construction. MassRET-20k provides only a held-out test set; there is no internal train/validation/test split.

2. Data Curation and Preprocessing

Strict protocols are employed for both molecular and spectral data:

  • Molecular Standardization:

All SMILES strings undergo canonicalization, kekulization, and valence checks. Entries failing any of these are removed. Only molecules with one-to-one mapped InChIKeys are retained to ensure structural uniqueness and resolve tautomers and stereoisomers consistently.

  • Spectral Normalization:

Fragment intensity arrays {Ij}\{I_j\} are rescaled as

Ij←Ijmax⁡kIk,Ij∈(0,1].I_j \leftarrow \frac{I_j}{\max_k I_k},\quad I_j\in(0,1].

Spectra lacking a peak above 1% relative intensity or featuring inconsistent or missing precursor metadata are excluded.

  • Exclusion and Inclusion Criteria:
    • Only spectra from NIST2020 are included.
    • All collision-energy annotations are present for 100% of entries.
    • Any molecule or spectrum present in MassSpecGym’s training set is excluded.

This ensures high-quality, standardized, and information-rich entries suitable for fair and interpretable zero-shot evaluation.

3. Structural and Spectral Diversity

The dataset achieves broad coverage across both chemical and experimental axes:

  • Ionization Adducts:

MassRET-20k spans twelve distinct ionization adduct types (e.g., +H, +Na, +K, –H, etc.). This is in contrast to MassSpecGym, which covers only two adduct types and lacks comprehensive collision-energy annotations in 47% of entries.

  • Mass Range:

Precursor (parent) mass mm values cover approximately 50–1200 Da:

∫501200fm(m) dm=1,\int_{50}^{1200} f_m(m)\,dm = 1,

where fm(m)f_m(m) is the parent-mass density.

  • Molecular Diversity:

The distribution of pairwise Tanimoto coefficients (computed over Morgan fingerprints) for the ∼\sim20,000 molecules is bell-shaped with a mean near 0.15:

Ei≠j[TMorgan(Mi,Mj)]≈0.15\mathbb{E}_{i\neq j}\bigl[T_\mathrm{Morgan}(M_i,M_j)\bigr]\approx0.15

This suggests substantial structural diversity, minimizing dataset redundancy and ensuring realistic assessment of generalization for retrieval models.

4. Schema and Annotation Structure

Each data entry is annotated with comprehensive metadata, adhering to the following schema:

Field Name Description
identifier Unique entry ID
mzs Array of fragment M\mathcal{M}0 values
intensities Array of relative intensities M\mathcal{M}1 normalized
SMILES Canonical SMILES string
inchikey Corresponding InChIKey
formula Molecular formula (neutral)
precursor_formula Formula of the precursor ion
parent_mass Exact mass of the molecule (Da)
precursor_mz Precursor M\mathcal{M}2
adduct Ionization adduct (12 types)
instrument_type MS instrument used
collision_energy Normalized collision-induced dissociation

All entries have complete metadata, and the annotation is designed for integration with downstream machine learning or cheminformatics pipelines.

5. Comparative Features and Evaluation Role

MassRET-20k is positioned as a standard, challenging, and diverse benchmark for zero-shot molecule retrieval from MS/MS. Compared to prior benchmarks, it introduces:

  • Twelve adduct types and full collision-energy annotation, in contrast to the limited coverage and missing metadata in MassSpecGym.
  • Non-overlapping structure set with respect to training data, avoiding information leakage and enabling robust assessment of generalization.
  • Uniformly rigorous curation criteria both for molecules and spectra.

The dataset serves exclusively as a test set for models trained, for example, on MassSpecGym, and is intended to support realistic, out-of-sample evaluation of molecular retrieval methodologies for both discriminative and generative paradigms under zero-shot conditions.

6. Key Summary Statistics

A high-level overview of the dataset is as follows:

Property Value
Total entries 20,000
Unique molecules 20,000
Source database NIST2020
Ionization adducts 12 types
Collision-energy coverage 100%
Precursor mass range 50–1200 Da
Avg. pairwise Tanimoto (Morgan) 0.15

Figures in the original paper include pie charts of adduct distribution, kernel density estimates of parent mass (M\mathcal{M}3), and violin plots of Tanimoto distributions for molecular diversity.

7. Significance and Applications

MassRET-20k addresses the need for a high-quality, chemically diverse benchmarking set for cross-modal molecule retrieval tasks. Its comprehensive curation, detailed metadata, and exclusion of training-set molecules provide a robust foundation for zero-shot assessment of spectrometry-to-structure models, such as GLMR and other state-of-the-art generative frameworks. A plausible implication is that its use will help more sharply delineate genuine model generalization from overfitting or memorization effects, especially for tasks involving uncommon adducts or structurally novel molecules. As MS-based compound identification workflows become increasingly reliant on data-driven retrieval models, MassRET-20k offers a broadly representative, standardized testbed for comparative evaluation and methodological development.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to MassRET-20k Dataset.