Titan25: MLIP Benchmark Dataset
- Titan25 is a benchmark dataset of 1,842,089 DFT-labeled atomic configurations spanning 11 elements to support MLIP training in both equilibrium and reactive regimes.
- It integrates molecules, crystals, surfaces, interfaces, and reactive mixtures to enable MLIPs to reproduce both static and dynamic experimental observables.
- Built with GAIA’s automated DG–DI pipeline and Nanoreactor+ techniques, Titan25 achieves lower energy and force errors compared to previous datasets.
Searching arXiv for the specified paper to ground the article and confirm bibliographic details. Titan25 is a benchmark-scale, multi-regime training dataset for machine learning interatomic potentials (MLIPs) built automatically with the GAIA framework. It was introduced to overcome the heuristic, system-specific character of typical MLIP datasets by systematically exploring and labeling diverse atomic environments spanning molecules, crystals, surfaces, interfaces, and reactive mixtures. Titan25 comprises 1,842,089 density functional theory (DFT)-labeled atomic configurations covering 11 elements, and is designed to train MLIPs that match both static and dynamic DFT results while reproducing experimental observables in detonation, carbon-nanotube coalescence, and catalytic processes (Song et al., 30 Sep 2025). The “25” is a dataset name rather than a count of elements or tasks; no semantic interpretation beyond branding is specified.
1. Definition and scope
Titan25 was constructed as a general-purpose training corpus for reactive atomistic simulation. Its central objective is to support MLIPs that remain accurate not only near equilibrium but also in non-equilibrium and chemically transforming regimes. In the formulation associated with GAIA, this entails systematic coverage of molecular, crystalline, interfacial, and reactive settings rather than the bulk-solid bias characteristic of many public datasets (Song et al., 30 Sep 2025).
The dataset spans the elements C, H, N, O, Ag, Au, Cu, Pd, Pt, Rh, and Ru. Its coverage includes CHNO molecules, energetic molecules such as TNT, TATB, RDX, and nitromethane, unary to ternary bulk inorganic crystals, low-index slabs, adsorbate-covered surfaces, bilayer interfaces, reactive mixtures generated through Nanoreactor+, and armchair carbon nanotubes with chiralities (6,6), (9,9), (12,12), (18,18), and (24,24). This combination of chemical and structural regimes is intended to support a single MLIP across heterogeneous application domains rather than a narrow domain-specific potential.
Titan25 also includes a strong non-equilibrium component. Its per-atom energy distribution is shifted to higher energies than ANI-1xnr and MPtrj, and its 95th percentile of atomic force norms is 7.1 eV/Å, exceeding ANI-1xnr, MPtrj, OMat24, and OMol25. This suggests that the dataset was designed not merely for interpolation around equilibrium structures but for reactive and mechanically stressed configurations where force-field transferability is typically most fragile.
2. Composition, structural regimes, and statistical properties
Titan25 integrates multiple classes of atomic arrangements into a single benchmark-scale corpus. The principal system categories are summarized below.
| Category | Coverage |
|---|---|
| Molecules | CHNO organics, water, and energetic molecules including TNT, TATB, RDX, nitromethane |
| Crystals | Unary, binary, and ternary bulk inorganic structures composed of Ag, Au, Cu, Pd, Pt, Rh, Ru, O, H |
| Surfaces and interfaces | (001), (110), and (111) slabs; adatoms; π-conjugated adsorbates; molecular bilayers on surfaces |
| Reactive and nanoscale systems | Metal–organic cluster growth, metadynamics-driven transformations, armchair carbon nanotubes |
The dataset contains gas-phase molecular pairs, periodic bulk structures, periodic slabs with approximately 30 Å vacuum, interfaces with molecular bilayers, and reactive cluster trajectories. Thermal states in bulk and slab labeling were sampled by ab initio molecular dynamics (AIMD) at 400–600 K, while Nanoreactor+ trajectories probe hot reactive pathways. Typical builder and labeling limits are at most 250 atoms per snapshot, with slab supercells up to 400 atoms, bulk supercells up to 200 atoms, and mol2surf groups up to 200 atoms.
Titan25 is notable for its explicit inclusion of regimes often treated separately in atomistic benchmarking. These include detonation under Hugoniostat molecular dynamics at 40 GPa, carbon-nanotube coalescence using a hybrid kMC/MD protocol at 1000 K, and catalytic and adsorption environments such as CO and CO on Pt(111) in water, π-conjugated adsorption on coinage metals, and the water–gas shift (WGS) on hydroxylated Au(100). A plausible implication is that Titan25 was curated to test whether a single equivariant potential can maintain fidelity across highly disparate thermodynamic, bonding, and dimensionality regimes.
3. Automated generation with GAIA
Titan25 was built with GAIA, an end-to-end automated generator of diverse atomic arrangements. GAIA comprises two modules: a Data-Generator (DG), which expands molecular xyz seeds and periodic POSCAR seeds into unlabeled structures before DFT labeling, and a Data-Improver (DI), which augments coverage in sparsely sampled or high-error regions of configurational space (Song et al., 30 Sep 2025).
The DG module uses several physics-inspired builders. Checkerboard generates synthetic bulk-like binaries on a cubic grid with varied stoichiometry. Bulk creates supercells with volume scaling factors 1.1 and 1.2 and random atom deletions of 10% and 20%, then adds structures generated from atom-deleted cells with scaling factors 0.85, 0.90, 0.95, and 1.05, for 13 structures per seed. Slab constructs (100) cleaved slabs with 30 Å vacuum and constrains the bottom approximately 20% of layers. Adatom samples single atoms on two-dimensional grids above slabs at heights Å, with denser sampling near the surface. Admol tiles and randomly rotates non-periodic molecules over slabs at heights Å using 20 orientations and surface-matched registry. Nanoreactor+ combines QCG-generated initial assemblies with metadynamics (MTD) to drive bond breaking and bond formation.
The DI module extends sampling by identifying low-density or high-error regions in a two-dimensional space defined by total energy and a diversity metric . It then re-invokes Nanoreactor+ using extracted substructures. The diversity metric is
where are the eigenvalues of the atom–atom distance matrix for snapshot . The frequency- and prediction-error-based sampling probabilities are given as
and
This DG–DI decomposition operationalizes two distinct goals: broad structural initialization and targeted coverage repair. The reported improvement of Titan25(G+I) over Titan25(G) indicates that DI is not merely an auxiliary augmentation step but a substantive contributor to the final dataset’s transferability.
4. Reactive exploration, curation, and labeling protocol
Within Nanoreactor+, quasi-classical generation (QCG) was implemented with CREST in QCG mode using 1 ps MD in ALPB water. The outer wall scaling factor was 1.2, with fallbacks 1.4, 0.8, 1.0, and 1.6; if convergence failed, spin multiplicities were cycled through singlet, doublet, and triplet. Pairs included periodic–non-periodic and non-periodic–non-periodic combinations, and the lowest-energy structures were passed to MTD.
Metadynamics used structural RMSD as the collective variable, a log-Fermi spherical wall at 6000 K, randomized biasing parameters , and a fallback spin sequence doublet 0 singlet 1 triplet 2 quartet. Each trajectory was 10 ps with a 1 fs timestep; the bias was updated every 10 fs and snapshots were saved every 50 fs, yielding 200 snapshots per run. The Methods narrative describes the bias approximately as
3
with 4 the RMSD collective variable. The paper also notes standard metadynamics reference forms for context, while stating that the implementation uses RMSD-based hills with randomized parameters rather than fixed well-tempered settings.
After MTD, trajectories were postprocessed with Open Babel and RDKit. Bond-order heatmaps (“bondmaps”) were averaged for each trajectory, and structural uniqueness was enforced by an SSIM threshold so that only distinct products were retained. DI heatmaps over total energy and 5 demonstrate that the improver expands into regions sparsely populated by DG. Additional curation steps include random rigid rotations of replicated supercells to avoid identical repeats and 0.3 Å padding of cell vectors to enhance intermolecular diversity.
The reference labeling calculations used VASP 6.4.1 with PAW pseudopotentials, the PBE exchange–correlation functional, and D3 zero damping. The plane-wave cutoff was 520 eV, k-point sampling was 6-point only, electronic convergence was 7 eV, force convergence was 8 eV/Å, Gaussian smearing was 0.05 eV, and symmetry was disabled. Labels comprise energies, atomic forces, and stress tensors from single-point evaluations, AIMD, and NEB-initial-step paths. Atomic forces follow the standard definition
9
AIMD for Checkerboard, Bulk, and Slab used NVT Nose–Hoover dynamics with 0.5 fs steps for 1000 steps and a 400 0 600 K temperature ramp, with NH mass 20; Nanoreactor+ used NVT for 20 steps at the same settings. NEB initial images were generated with nebmake.pl, followed by single-point DFT evaluations only, without full NEB relaxations, and with hard filtering of unphysical short contacts whose Coulomb-like repulsion exceeded 50 eV.
5. Dataset organization, model training, and benchmark methodology
Input seeds are provided in xyz format for molecules and POSCAR format for periodic systems. Outputs include coordinates, energies, forces, and stresses suitable for standard MLIP trainers. The model-comparison splits use an 8:1:1 train:validation:test partition. Titan25(G) denotes the DG-only dataset, whereas Titan25(G+I) includes DI augmentation. Energies are reported in eV, forces in eV/Å, and distances in Å. Licensing, DOI, and repository access instructions are not specified.
Titan25 was used to train SNet-T25, based on SevenNet, described as an 1-equivariant GNN variant of NequIP with shared self-connection parameters across elements (Song et al., 30 Sep 2025). The architecture used embedding irreps of 2, two interaction blocks with 3, a final block of 4, a 6 Å cutoff, XPLOR cutoff treatment, and an 8-function Bessel radial basis. Training used AdamW with batch size 256, cosine learning-rate scheduling with 10% warmup, maximum learning rate 0.01, gradient clipping 100, weight decay 0.001, and loss weights 5. Training length was approximately 0.5 million steps, with epochs adjusted by dataset size: Titan25(G+I) approximately 80 epochs, Titan25(G) approximately 125, MPtrj approximately 100, and ANI-1xnr approximately 300.
Benchmarking employed GAIA-Bench tasks comprising mol2mol for intermolecular potential-energy surfaces, bulk for energy–volume curves, slab for facet stability, and mol2surf for bilayer adsorption energetics. Error metrics include
6
and
7
According to the reported comparisons, Titan25(G+I) consistently yielded the lowest normalized mean errors across tasks, with about twofold lower energy errors on average and one-third lower force errors versus MPtrj and ANI-1xnr, together with an approximately 20% improvement over Titan25(G). Force magnitude and angular errors were also the smallest for Titan25(G+I).
Representative force MAEs in GAIA-Bench were 82 meV/Å for mol2mol, 92 meV/Å for bulk, 119 meV/Å for slab, and 101 meV/Å for mol2surf. For comparison, Titan25(G) gave 84, 130, 145, and 121 meV/Å, respectively; MPTrj gave 210, 109, 131, and 131 meV/Å where reported; ANI-1xnr gave 198 meV/Å on mol2mol. On re-labeled out-of-distribution public tests, Titan25(G+I) achieved energy MAE 27 meV and force MAE 137 meV/Å on ANI-1x random samples, 20 meV and 101 meV/Å on GDB-13 random samples, 12 meV and 232 meV/Å on ANI-1xnr validation, and 83 meV and 112 meV/Å on MPTrj validation restricted to Titan25 elements.
6. Dynamic fidelity and experimentally anchored case studies
Titan25’s significance lies not only in static benchmark performance but in its use as a training substrate for simulations intended to reproduce experimentally observed behavior across distinct reactive regimes. In interfacial molecular dynamics for Pt–CO8–H9O systems, SNet-T25 reproduced the Pt–C vertical distance for CO/Pt(111) at approximately 1.4 Å, consistent with DFT. For CO0/Pt(111), SNet-T25 was the only compared model that kept CO1 near the surface at approximately 3.4 Å, whereas UMA-OMol and UMA-OMC fluctuated and often exceeded 4.0 Å, indicating repulsion. Reduced-size DFT checks were reported to agree with SNet-T25 (Song et al., 30 Sep 2025).
For π-conjugated adsorption on coinage metals, SNet-T25 relaxed distances for benzene, DIP, and C2 on Ag(111) with deviations of approximately 5% on average from experiment. The experimentally observed DIP trend across coinage metals, Cu 3 Ag 4 Au, was captured correctly. Additional validation was reported for C5@Ag(111) with a vacancy and for H6O@C7 on Ag(111), where the Ag–O distance during MD was 5.60 Å versus an experimental 5.57 Å with vacancy.
In detonation simulations, Hugoniostat MD at 40 GPa was carried out for TNT, TATB, RDX, and nitromethane. SNet-T25 predicted normalized product ratios for H8O, CO9, N0, and NH1 in close agreement with experiments and with the specialized reactive model Gen3.9zbl. The smallest deviations were reported for nitromethane, below 4% across products, while the largest occurred for TATB, with approximately 18% overproduction of N2. Product ordering across pressures was reproduced without ZBL corrections.
For carbon-nanotube coalescence, armchair pairs (6,6), (9,9), and (12,12) coalesced into pores approximating doubled tubes, namely (12,12), (18,18), and (24,24), in agreement with microscopy experiments. The large (24,24) pair instead collapsed into ribbon-like shapes, consistent with diameter-dependent flattening. In WGS catalysis on hydroxylated Au(100), large-scale MD was performed with 69,620 Au atoms, 4,462 OH groups, 4,332 CO molecules, and 4,557 H3O molecules at 600 K for 200,000 steps. The trajectories captured both carboxyl and redox pathways under explicit catalytic conditions:
- carboxyl: CO4 + OH5 6 OCOH7 8 CO9
- redox: OH0 1 O2; O3 + CO4 5 CO6
These results indicate that Titan25 was used to train a model intended to bridge the conventional separation between equilibrium fitting benchmarks and chemically realistic dynamic simulations.
7. Relation to prior datasets, limitations, and interpretation
Titan25 occupies an intermediate position between small reactive molecular datasets and very large inorganic foundation datasets. It offers 11-element coverage and approximately 1.8 million snapshots, while emphasizing high-energy and high-force non-equilibrium structures to a greater degree than ANI-1xnr, MPtrj, OMol25, and OMat24 (Song et al., 30 Sep 2025). Unlike datasets assembled primarily from pre-existing repositories such as Materials Project, CoRE MOF, OE62, or OMol25, Titan25 was generated de novo through GAIA’s automated builders and Nanoreactor+ workflow, without pre-trained MLIPs or manual curation.
A recurrent misconception about large MLIP datasets is that scale alone guarantees universality. The associated paper explicitly notes the conceptual difficulty of a truly universal single-model MLIP because configuration space is effectively infinite and no general adequacy criterion for sampling is available. GAIA is presented instead as a practical iterative strategy: expand from seed components, assess coverage, and fill gaps through targeted augmentation. This suggests that Titan25 should be understood less as a complete universal corpus than as a benchmark-scale instantiation of a scalable data-generation protocol.
Several limitations are specified. DFT labeling does not extensively treat spin states beyond the exploratory multiplicity handling in QCG and MTD; k-point sampling is 7-only; finite-temperature electronic effects are not considered; charges and dipoles are not included among the labels, though stress tensors are. Access metadata, licensing terms, DOI, and repository instructions are also not specified. These constraints matter for downstream use: they delimit the interpretive scope of trained MLIPs, particularly for systems where Brillouin-zone sampling, electronically excited states, or electrostatic observables are central.
Taken together, Titan25 represents a broad, reactive, and systematically generated MLIP training dataset whose distinguishing feature is not only its size but the explicit integration of molecules, crystals, surfaces, interfaces, and reactive mixtures into a unified DFT-labeled corpus. Its construction through the GAIA pipeline, and its use in training SNet-T25, positions it as a reference point for benchmark-scale attempts to narrow the gap between atomistic simulation, DFT fidelity, and experimentally observed reactive dynamics (Song et al., 30 Sep 2025).