Papers
Topics
Authors
Recent
Search
2000 character limit reached

Benchmark Dataset for Catalysis on 2D MXenes

Published 30 May 2026 in cond-mat.mtrl-sci and cs.LG | (2606.00794v1)

Abstract: Merging first-principles calculations with ML, we aim to accelerate the exploration of catalytic behaviour in novel materials. We focus on two-dimensional (2D) Ti$_2$CT$_y$ MXenes, whose versatile surface chemistry makes them particularly compelling candidates for catalysis. Resolving their composition and structure under realistic conditions exceeds the reach of standard density functional theory (DFT) due to computational cost. To address this challenge, we generate a comprehensive dataset of 50,000 DFT calculations for training and 10,000 for testing, encompassing both Ti$_2$CT$_y$ MXene configurations and molecular systems, along with an additional test dataset with 1000 genuinely new, larger systems to investigate how well models generalise. We train and validate widely used and competitive machine learning interatomic potential (MLIP) models, including EquiformerV2, MACE, MatRIS, and UPET, that accurately predict atomic forces and formation energies -- quantities that DFT must repeatedly compute for structural and catalytic investigations -- for these 2D materials. This combined DFT-ML framework achieves computational acceleration on the order of approximately $1-4 \cdot 103$ (on a CPU) while maintaining desired-level accuracy (approximately +/- $10$ meV/A for forces and approximately +/- $1$ meV for per-atom energies), paving the way for more efficient investigations of MXene catalytic behaviour. Moreover, we perform an extensive qualitative evaluation of the trained models, showcasing the importance of comprehensive simulation-based comparison beyond benchmark metrics. The dataset and the trained models with the code are available at https://huggingface.co/datasets/CatalystAnonymous/catalyst_mxenes.

Summary

  • The paper presents the CataLiUst Ti2C-MXene dataset with 50K training and 10K test DFT calculations to model catalytic reactions on 2D MXenes.
  • It benchmarks four MLIP architectures, emphasizing the role of conservative force prediction for accurate and stable catalytic simulations.
  • Quantitative and MD evaluations demonstrate DFT-level energy and force accuracy with over 1000x computational efficiency, enabling scalable catalyst design.

Benchmark Dataset for Catalysis on 2D MXenes: An Expert Analysis

Motivation and Dataset Construction

This work addresses the challenge of modeling catalytic reactions on Ti2_2CTy_y MXenes, a class of 2D transition metal carbides with highly variable and poorly-characterized surface chemistries. Classical DFT approaches are prohibitively expensive for large-scale or realistic scenarios. The authors present the CataLiUst Ti2_2C-MXene dataset, comprising 50,000 DFT calculations for training and 10,000 for testing, with an additional test set of 1,000 larger, out-of-distribution systems. The dataset encompasses distinct termination patterns (O, OH, mixtures), reaction pathways, non-equilibrium configurations (via rattling and high-T MD), and molecule adsorption relevant for mechanistic studies, particularly CO2_2 reduction and other prototypical catalytic processes.

The construction strategy emphasizes chemically diverse environments despite minimal element variety (C, Ti, O, H), creating a highly challenging setting for both classical ML and equivariant architectures. The dataset was generated using rev-vdW-DF2, a GGA-based functional with robust non-local correlation, validated for layered materials and molecule-surface interactions. Data are provided in HDF5 and XYZ formats fully compatible with leading MLIP frameworks. Figure 1

Figure 1: Dataset composition illustrating the diversity of molecular systems and DFT calculation types included for MXene catalysis.

Machine Learning Interatomic Potentials: Architectures and Training Protocols

Four leading MLIP approaches were benchmarked: EquiformerV2, MACE, UPET/PET, and MatRIS. All were trained and evaluated under consistent protocol, including both conservative (energy-consistent) and non-conservative (direct regression) force formulations. Model selection reflects geometric deep learning principles, guaranteeing E(3)/SO(3) equivariance.

EquiformerV2 utilizes explicit force heads enabling flexibility but at the cost of energy-force consistency. MACE, MatRIS, and UPET/PET enforce conservative and non-conservative modes. Foundation models were further fine-tuned, including MACE PT omat/matpes and MatRIS PT oam/mp, to assess transferability.

Training on the supplied dataset is non-trivial: the wide distribution of system sizes and environments mandates robust architectures and loss normalization strategies. Test splits probe both in-distribution accuracy and generalization to expanded supercells.

Quantitative Evaluation

The results (summarized below) demonstrate substantial advances in both accuracy and efficiency. Strong numerical results include:

  • PET PT mad-s achieved the lowest MAE: $1.1$ meV/atom (energy), $10.3$ meV/Ã… (forces) on the original test set.
  • MACE PT matpes and UPET cons. also yield highly competitive accuracy, with MACE PT matpes at $1.7$ meV/atom and $12.2$ meV/Ã….
  • All MLIPs provide >1000x speedup over DFT on CPU, with MACE attaining 4258-fold acceleration.

Generalization to out-of-distribution larger systems incurs expected degradation, but PET PT mad-s and fine-tuned MACE variants exhibit superior robustness (force MAE increase is minimized; energy MAE remains stable).

Conservative force prediction systematically improves energy and force accuracy, but its impact on scaling is model-dependent; benefits are pronounced for in-distribution data and stable MD.

Qualitative Evaluation and Simulation Performance

Extensive MD simulations probe the stability of MLIPs in catalytically-relevant scenarios. Conservative models consistently outperform non-conservative counterparts, maintaining DFT-level force and energy fidelity across long MD trajectories for both CO2_2 and HCOOH adsorption on MXenes. Non-conservative models produce greater deviations and instability, underscoring the necessity of energy-force consistency for physically reliable simulations. Figure 2

Figure 2: Comparison of MLIP potential energies and per-atom force errors vs. DFT for molecular dynamics of CO2_2 and HCOOH on fully terminated MXene surfaces.

Figure 3

Figure 3: Molecular dynamics trajectories for COy_y0 and HCOOH adsorbed on MXenes, highlighting the superior correspondence of MACE and EquiformerV2 models to DFT.

Figure 4

Figure 4: MD performance and error analysis for different MACE foundation models with and without fine-tuning: fine-tuning dramatically improves force and energy prediction.

Figure 5

Figure 5: Radial distribution function discrepancies between ML and DFT trajectories, demonstrating highest physical fidelity for PET and fine-tuned MACE models.

Implications and Future Directions

The CataLiUst Tiy_y1C-MXene dataset and accompanying MLIP benchmarks significantly increase the tractability of catalytic surface modeling under realistic, dynamic conditions. MLIPs, particularly PET PT mad-s and MACE PT matpes, now yield DFT-level accuracy (±10 meV/Å for forces, ±1 meV for per-atom energies) with several orders of magnitude greater efficiency. This makes large-scale parameter sweeps, reactive mechanism exploration, and uncertainty quantification feasible for MXene catalysis—previously unobtainable via classical methods.

Strong claims of limited transferability for existing foundation models are empirically validated. Zero-shot application to MXenes yields poor results; fine-tuning is mandatory.

Practical implications extend to catalyst design workflows, where MLIPs trained on these datasets can serve as scalable surrogates for DFT, accelerating screening across experimental surface configurations and reaction environments. Theoretical implications include the demonstration that geometric deep learning and inductive bias engineering (equivariance, conservative force prediction) are crucial for complex material systems with subtle chemical heterogeneity.

Anticipated future directions comprise:

  • Expansion to a broader molecular scope and additional catalytic cycles.
  • Improved uncertainty quantification and data-driven adaptive sampling.
  • Integration with closed-loop experimental design for materials discovery.
  • Cross-validation of MLIP predictions with real-world catalysis rates and spectroscopic signatures.

Conclusion

This work establishes a rigorous benchmark for catalysis modeling on 2D MXenes, providing both a diverse dataset and comprehensive MLIP evaluation. The demonstrated accuracy, speed, and stability of MLIPs—especially with conservative force enforcement and fine-tuning—enable practical simulation of catalytic phenomena that were previously computationally inaccessible. The dataset will facilitate further research into catalyst design, mechanism analysis, and development of advanced ML interatomic potentials for chemically complex materials (2606.00794).

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Collections

Sign up for free to add this paper to one or more collections.