Papers
Topics
Authors
Recent
Search
2000 character limit reached

Dataset Distillation for Machine Learning Force Field in Phase Transition Regime

Published 3 Apr 2026 in physics.chem-ph | (2604.03027v1)

Abstract: Machine learning force field (MLFF) has emerged as a powerful data-driven tool for atomistic simulations, enabling large-scale and complex atomic systems to be simulated with accuracy comparable to \textit{ab initio} methods. However, MLFFs often suffer from low training efficiency in the phase transition regime, where structural fluctuations are significantly elevated. To address this challenge, we propose a Central-Peripheral Distillation (CPD) algorithm for training dataset distillation. By strategically integrating representative samples with critical corner cases, the CPD algorithm ensures that the distilled dataset retains maximum structural diversity. We validated the efficacy of the CPD method on the liquid-liquid phase transition of dense hydrogen. Results show that, with the CPD approach, only 200 configurations are sufficient to train a MLFF that can fully reproduce the structural and dynamical properties of liquid hydrogen in the vicinity of its phase transition regime. This work paves the way for high-fidelity labeling of the MLFF training datasets, for instance by adopting high-level \textit{ab initio} calculations beyond the standard density functional theory, thereby enhancing the predictive accuracy of MLFFs.

Authors (3)

Summary

  • The paper introduces the CPD algorithm for improved dataset distillation in ML force field training under phase transition conditions.
  • It employs balanced sampling from both dense and sparse configuration regions to capture stable phases and rare critical states.
  • With only 35% of the full data, CPD achieves near ab initio accuracy, cutting high-level computation costs by over 65%.

Central-Peripheral Distillation for Robust Phase Transition MLFFs

Introduction

The advancement of machine learning force fields (MLFFs) has enabled high-fidelity atomistic simulations of materials with computational efficiency far surpassing ab initio methods, while retaining near ab initio accuracy. However, the development of MLFFs in the phase transition regime remains challenging due to elevated structural fluctuations and a strongly broadened configuration space. The paper "Dataset Distillation for Machine Learning Force Field in Phase Transition Regime" (2604.03027) targets this bottleneck, introducing a Central-Peripheral Distillation (CPD) algorithm that improves dataset selection strategies for MLFFs under critical phase transition conditions. The efficacy of CPD is rigorously validated for the liquid-liquid phase transition (LLPT) in dense hydrogen at 1000 K, demonstrating substantial gains in data efficiency and predictive accuracy.

Methodological Contributions

The central methodological advancement in this work is the CPD algorithm, which strategically samples both densely populated (central) and sparsely populated (peripheral) regions of the reduced configurational feature space to maximize information retention and diversity in the distilled training data. Configuration features are extracted using the MACE descriptor and embedded into a lower-dimensional space via Principal Component Analysis, after which local densities are quantified. A balanced weighted sampling policy selects the top 20% densest and 20% sparsest configurations, thus ensuring comprehensive coverage of stable phases and rare transition structures. This design explicitly addresses the intrinsic heterogeneity of phase space in phase transition regimes, capturing both typical and atypical/critical configurations necessary for robust MLFF interpolation and extrapolation.

For benchmarking, a new dataset (HLLPT1k) was generated using DFT-based ab initio molecular dynamics with strict control over density (0.98–1.41 g/cm³), ensuring thorough coverage across molecular, mixed, and atomic phases.

Numerical Results and Comparative Analysis

Quantitative evaluation against alternative distillation methods (Random Network Distillation, DIRECT, and uniform Random sampling) was conducted by measuring energy and force root mean square errors (RMSE) and by assessing dynamic property prediction during phase transition. The CPD algorithm achieves an energy RMSE of 4.3 meV/atom using only 200 configurations (35% of the full set), converging to the full-data model’s error of 3.1 meV/atom. In contrast, the DIRECT approach plateaus at 14.7 meV/atom and RND remains substantially less accurate.

Notably, with only 200 CPD-selected structures, the MLFF accurately reproduces key thermodynamic observables (pressure, molecular fraction) in the LLPT regime, including correct identification of both phase boundaries and transitions. Both random and DIRECT-based models fail to robustly resolve the phase transition regime, yielding breakdowns or large errors in predicted quantities, while the CPD and full-data models remain stable across all conditions.

Further, the superiority of CPD is shown to be robust to the choice of descriptor, as similar gains are observed using alternatives such as SchNet, indicating that the core advantage derives from the sampling protocol rather than descriptor compatibility.

Practical and Theoretical Implications

The results establish CPD as a highly data-efficient procedure for MLFF training in regimes with significant structural variability and rare event dynamics, directly addressing the primary limitation in phase transition MLFF construction. The reduction of required high-level quantum mechanical calculations by more than 65% without accuracy loss is particularly impactful for future efforts that target quantum chemical accuracy beyond DFT, where labeling costs are orders of magnitude higher.

Theoretically, the CPD approach generalizes to other first-order or continuous phase transitions and to systems where outliers and critical fluctuations are essential for capturing complex emergent behavior. The algorithm’s explicit balancing of representative and rare configurations can be seamlessly incorporated into existing MLFF workflows, offering a template for distilled dataset design in high-dimensional, non-uniform phase spaces.

Prospects for Future Research

The demonstrated efficiency and robustness of CPD motivate several future directions. Firstly, integrating CPD with active learning frameworks may yield further reductions in data requirements or enable on-the-fly MLFF refinement in dynamic simulations. The potential to combine CPD-based dataset distillation with neural network wavefunction-based QM methods or advanced post-DFT schemes could unravel new physics in intractable material systems. Broad extension to multi-component systems, meta-stable transitions, and non-equilibrium processes presents an open area for methodological generalization. Finally, CPD provides a foundation for automated, information-theoretic training data selection—including integration with uncertainty quantification and multi-fidelity simulation strategies.

Conclusion

The Central-Peripheral Distillation algorithm provides an effective, physically informed data distillation scheme for developing MLFFs capable of accurately capturing phase transition regimes with dramatically reduced data requirements. The verified performance in LLPT of dense hydrogen highlights its substantial advantage over prevailing distillation approaches, both numerically and qualitatively. CPD is poised to become a key component in scalable, high-fidelity ML-driven materials simulations, especially as the drive for higher-accuracy ab initio labeling intensifies.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We found no open problems mentioned in this paper.