- The paper presents the CataLiUst Ti2C-MXene dataset with 50K training and 10K test DFT calculations to model catalytic reactions on 2D MXenes.
- It benchmarks four MLIP architectures, emphasizing the role of conservative force prediction for accurate and stable catalytic simulations.
- Quantitative and MD evaluations demonstrate DFT-level energy and force accuracy with over 1000x computational efficiency, enabling scalable catalyst design.
Benchmark Dataset for Catalysis on 2D MXenes: An Expert Analysis
Motivation and Dataset Construction
This work addresses the challenge of modeling catalytic reactions on Ti2​CTy​ MXenes, a class of 2D transition metal carbides with highly variable and poorly-characterized surface chemistries. Classical DFT approaches are prohibitively expensive for large-scale or realistic scenarios. The authors present the CataLiUst Ti2​C-MXene dataset, comprising 50,000 DFT calculations for training and 10,000 for testing, with an additional test set of 1,000 larger, out-of-distribution systems. The dataset encompasses distinct termination patterns (O, OH, mixtures), reaction pathways, non-equilibrium configurations (via rattling and high-T MD), and molecule adsorption relevant for mechanistic studies, particularly CO2​ reduction and other prototypical catalytic processes.
The construction strategy emphasizes chemically diverse environments despite minimal element variety (C, Ti, O, H), creating a highly challenging setting for both classical ML and equivariant architectures. The dataset was generated using rev-vdW-DF2, a GGA-based functional with robust non-local correlation, validated for layered materials and molecule-surface interactions. Data are provided in HDF5 and XYZ formats fully compatible with leading MLIP frameworks.
Figure 1: Dataset composition illustrating the diversity of molecular systems and DFT calculation types included for MXene catalysis.
Machine Learning Interatomic Potentials: Architectures and Training Protocols
Four leading MLIP approaches were benchmarked: EquiformerV2, MACE, UPET/PET, and MatRIS. All were trained and evaluated under consistent protocol, including both conservative (energy-consistent) and non-conservative (direct regression) force formulations. Model selection reflects geometric deep learning principles, guaranteeing E(3)/SO(3) equivariance.
EquiformerV2 utilizes explicit force heads enabling flexibility but at the cost of energy-force consistency. MACE, MatRIS, and UPET/PET enforce conservative and non-conservative modes. Foundation models were further fine-tuned, including MACE PT omat/matpes and MatRIS PT oam/mp, to assess transferability.
Training on the supplied dataset is non-trivial: the wide distribution of system sizes and environments mandates robust architectures and loss normalization strategies. Test splits probe both in-distribution accuracy and generalization to expanded supercells.
Quantitative Evaluation
The results (summarized below) demonstrate substantial advances in both accuracy and efficiency. Strong numerical results include:
- PET PT mad-s achieved the lowest MAE: $1.1$ meV/atom (energy), $10.3$ meV/Ã… (forces) on the original test set.
- MACE PT matpes and UPET cons. also yield highly competitive accuracy, with MACE PT matpes at $1.7$ meV/atom and $12.2$ meV/Ã….
- All MLIPs provide >1000x speedup over DFT on CPU, with MACE attaining 4258-fold acceleration.
Generalization to out-of-distribution larger systems incurs expected degradation, but PET PT mad-s and fine-tuned MACE variants exhibit superior robustness (force MAE increase is minimized; energy MAE remains stable).
Conservative force prediction systematically improves energy and force accuracy, but its impact on scaling is model-dependent; benefits are pronounced for in-distribution data and stable MD.
Extensive MD simulations probe the stability of MLIPs in catalytically-relevant scenarios. Conservative models consistently outperform non-conservative counterparts, maintaining DFT-level force and energy fidelity across long MD trajectories for both CO2​ and HCOOH adsorption on MXenes. Non-conservative models produce greater deviations and instability, underscoring the necessity of energy-force consistency for physically reliable simulations.
Figure 2: Comparison of MLIP potential energies and per-atom force errors vs. DFT for molecular dynamics of CO2​ and HCOOH on fully terminated MXene surfaces.
Figure 3: Molecular dynamics trajectories for COy​0 and HCOOH adsorbed on MXenes, highlighting the superior correspondence of MACE and EquiformerV2 models to DFT.
Figure 4: MD performance and error analysis for different MACE foundation models with and without fine-tuning: fine-tuning dramatically improves force and energy prediction.
Figure 5: Radial distribution function discrepancies between ML and DFT trajectories, demonstrating highest physical fidelity for PET and fine-tuned MACE models.
Implications and Future Directions
The CataLiUst Tiy​1C-MXene dataset and accompanying MLIP benchmarks significantly increase the tractability of catalytic surface modeling under realistic, dynamic conditions. MLIPs, particularly PET PT mad-s and MACE PT matpes, now yield DFT-level accuracy (±10 meV/Å for forces, ±1 meV for per-atom energies) with several orders of magnitude greater efficiency. This makes large-scale parameter sweeps, reactive mechanism exploration, and uncertainty quantification feasible for MXene catalysis—previously unobtainable via classical methods.
Strong claims of limited transferability for existing foundation models are empirically validated. Zero-shot application to MXenes yields poor results; fine-tuning is mandatory.
Practical implications extend to catalyst design workflows, where MLIPs trained on these datasets can serve as scalable surrogates for DFT, accelerating screening across experimental surface configurations and reaction environments. Theoretical implications include the demonstration that geometric deep learning and inductive bias engineering (equivariance, conservative force prediction) are crucial for complex material systems with subtle chemical heterogeneity.
Anticipated future directions comprise:
- Expansion to a broader molecular scope and additional catalytic cycles.
- Improved uncertainty quantification and data-driven adaptive sampling.
- Integration with closed-loop experimental design for materials discovery.
- Cross-validation of MLIP predictions with real-world catalysis rates and spectroscopic signatures.
Conclusion
This work establishes a rigorous benchmark for catalysis modeling on 2D MXenes, providing both a diverse dataset and comprehensive MLIP evaluation. The demonstrated accuracy, speed, and stability of MLIPs—especially with conservative force enforcement and fine-tuning—enable practical simulation of catalytic phenomena that were previously computationally inaccessible. The dataset will facilitate further research into catalyst design, mechanism analysis, and development of advanced ML interatomic potentials for chemically complex materials (2606.00794).