- The paper introduces the CPD algorithm for improved dataset distillation in ML force field training under phase transition conditions.
- It employs balanced sampling from both dense and sparse configuration regions to capture stable phases and rare critical states.
- With only 35% of the full data, CPD achieves near ab initio accuracy, cutting high-level computation costs by over 65%.
Central-Peripheral Distillation for Robust Phase Transition MLFFs
Introduction
The advancement of machine learning force fields (MLFFs) has enabled high-fidelity atomistic simulations of materials with computational efficiency far surpassing ab initio methods, while retaining near ab initio accuracy. However, the development of MLFFs in the phase transition regime remains challenging due to elevated structural fluctuations and a strongly broadened configuration space. The paper "Dataset Distillation for Machine Learning Force Field in Phase Transition Regime" (2604.03027) targets this bottleneck, introducing a Central-Peripheral Distillation (CPD) algorithm that improves dataset selection strategies for MLFFs under critical phase transition conditions. The efficacy of CPD is rigorously validated for the liquid-liquid phase transition (LLPT) in dense hydrogen at 1000 K, demonstrating substantial gains in data efficiency and predictive accuracy.
Methodological Contributions
The central methodological advancement in this work is the CPD algorithm, which strategically samples both densely populated (central) and sparsely populated (peripheral) regions of the reduced configurational feature space to maximize information retention and diversity in the distilled training data. Configuration features are extracted using the MACE descriptor and embedded into a lower-dimensional space via Principal Component Analysis, after which local densities are quantified. A balanced weighted sampling policy selects the top 20% densest and 20% sparsest configurations, thus ensuring comprehensive coverage of stable phases and rare transition structures. This design explicitly addresses the intrinsic heterogeneity of phase space in phase transition regimes, capturing both typical and atypical/critical configurations necessary for robust MLFF interpolation and extrapolation.
For benchmarking, a new dataset (HLLPT1k) was generated using DFT-based ab initio molecular dynamics with strict control over density (0.98–1.41 g/cm³), ensuring thorough coverage across molecular, mixed, and atomic phases.
Numerical Results and Comparative Analysis
Quantitative evaluation against alternative distillation methods (Random Network Distillation, DIRECT, and uniform Random sampling) was conducted by measuring energy and force root mean square errors (RMSE) and by assessing dynamic property prediction during phase transition. The CPD algorithm achieves an energy RMSE of 4.3 meV/atom using only 200 configurations (35% of the full set), converging to the full-data model’s error of 3.1 meV/atom. In contrast, the DIRECT approach plateaus at 14.7 meV/atom and RND remains substantially less accurate.
Notably, with only 200 CPD-selected structures, the MLFF accurately reproduces key thermodynamic observables (pressure, molecular fraction) in the LLPT regime, including correct identification of both phase boundaries and transitions. Both random and DIRECT-based models fail to robustly resolve the phase transition regime, yielding breakdowns or large errors in predicted quantities, while the CPD and full-data models remain stable across all conditions.
Further, the superiority of CPD is shown to be robust to the choice of descriptor, as similar gains are observed using alternatives such as SchNet, indicating that the core advantage derives from the sampling protocol rather than descriptor compatibility.
Practical and Theoretical Implications
The results establish CPD as a highly data-efficient procedure for MLFF training in regimes with significant structural variability and rare event dynamics, directly addressing the primary limitation in phase transition MLFF construction. The reduction of required high-level quantum mechanical calculations by more than 65% without accuracy loss is particularly impactful for future efforts that target quantum chemical accuracy beyond DFT, where labeling costs are orders of magnitude higher.
Theoretically, the CPD approach generalizes to other first-order or continuous phase transitions and to systems where outliers and critical fluctuations are essential for capturing complex emergent behavior. The algorithm’s explicit balancing of representative and rare configurations can be seamlessly incorporated into existing MLFF workflows, offering a template for distilled dataset design in high-dimensional, non-uniform phase spaces.
Prospects for Future Research
The demonstrated efficiency and robustness of CPD motivate several future directions. Firstly, integrating CPD with active learning frameworks may yield further reductions in data requirements or enable on-the-fly MLFF refinement in dynamic simulations. The potential to combine CPD-based dataset distillation with neural network wavefunction-based QM methods or advanced post-DFT schemes could unravel new physics in intractable material systems. Broad extension to multi-component systems, meta-stable transitions, and non-equilibrium processes presents an open area for methodological generalization. Finally, CPD provides a foundation for automated, information-theoretic training data selection—including integration with uncertainty quantification and multi-fidelity simulation strategies.
Conclusion
The Central-Peripheral Distillation algorithm provides an effective, physically informed data distillation scheme for developing MLFFs capable of accurately capturing phase transition regimes with dramatically reduced data requirements. The verified performance in LLPT of dense hydrogen highlights its substantial advantage over prevailing distillation approaches, both numerically and qualitatively. CPD is poised to become a key component in scalable, high-fidelity ML-driven materials simulations, especially as the drive for higher-accuracy ab initio labeling intensifies.