- The paper introduces a novel distillation protocol using transfer learning to create compact ML models that retain high first-principles accuracy.
- The distilled student models achieve up to 10× speedup while maintaining sub-milligram per cubic centimeter density errors and robust thermodynamic predictions.
- This approach reliably captures quantum nuclear effects, reactive dynamics, and interfacial phenomena, matching experimental benchmarks and high-level computational results.
Distilling First-Principles Accuracy into Compact ML Potentials for Condensed-Phase Chemistry
Introduction and Methodological Framework
"Distilling first-principles accuracy into compact machine learning potentials for condensed-phase chemistry" (2606.06848) introduces a systematic protocol combining transfer learning and knowledge distillation to generate highly efficient, high-fidelity interatomic potentials (MLIPs) for condensed-phase systems. The work addresses the longstanding trade-off in molecular simulations between model expressivity (accuracy across configuration space) and inference efficiency (simulation throughput), particularly critical for large system sizes, long trajectories, and advanced sampling algorithms (e.g., PIMD, umbrella sampling).
The proposed workflow employs foundation MLIPs, such as large equivariant GNNs (e.g., MACE-MP-0), as teachers. These are fine-tuned to specific systems or levels of theory via transfer learning on limited high-level quantum data and subsequently used to label physically relevant configuration sets. Compact student models, encompassing both neural (MACE) and polynomial-based (ACE) architectures, are distilled by fitting to these synthetic teacher-labeled data. This approach ensures that the student inherits high-fidelity predictions while being orders of magnitude cheaper computationally, provided that the training set and expressivity match the relevant dynamical and thermodynamical regime.
Case Study I: Ice Ih—Speed–Accuracy Tradeoff in Solid-State NPT Dynamics
The first demonstration targets NPT simulations of ice Ih at and beyond the training regime. A fine-tuned MACE teacher (L=2, 128 channels) is constructed from 400 revPBE-D3(0) DFT configurations. Several student architectures (ACE and simplified MACE variants) are then distilled on 2000 teacher-labeled configurations generated by "rattling" the DFT structures to enhance sampling.
The speed–accuracy landscape is compelling:

Figure 1: Speed–accuracy tradeoff for ice Ih; (a) atomistic configuration of the 384-atom supercell, (b) density deviation versus simulation throughput for teacher and various student models.
Distilled student models (MACE L=0, 64 channels; ACE; etc.) deliver up to a 10× speedup (relative to the teacher) with minimal sacrifice in accuracy, retaining sub-milligram per cubic centimeter errors on density at both 100 K and the extrapolated 220 K conditions. Notably, same-size scratch models (trained directly on DFT) often fail to maintain NPT stability or exhibit significant density errors when evaluated outside their training configuration domain, especially for less expressive or data-starved architectures. This highlights the crucial role of knowledge transfer and synthetic configurational expansion for robust condensed-phase MLIPs.
Case Study II: Quantum Nuclear Effects in Liquid Water—Distilled Δ-ML CCSD(T) Potential
The second test advances to path-integral MD (PIMD) for liquid water, where both quantum nuclear effects (NQEs) and the need for coupled-cluster (CCSD(T))-level energetics place extreme demands on model accuracy and ergodicity. The teacher is a Δ-learned MACE model: a DFT-potential baseline augmented by a machine-learned CCSD(T)-DFT correction ("Δ-learning"), trained as described in [oneillRoutineCondensedPhase2025].
Using a batch of teacher-driven NPT and PIMD simulations across a (T, ρ, p) grid, a large synthetic dataset is distilled onto a compact ACE student, parametrized for both efficiency and flexibility.
The distilled student achieves robust performance:

Figure 2: (a) Representative atomistic configuration for simulation, (b) density isobar for water across 240–370 K, contrasted across models and experiment, (c) temperature-dependent diffusion coefficients, (d) quantum kinetic energies per atom/species.
Key results:
- The student achieves an order-of-magnitude speed-up, enabling multi-ns MD/PIMD trajectories across the full liquid temperature range.
- The density isobar and diffusion coefficients are reproduced to better than 5% versus CCSD(T)-quality teacher and experimental benchmarks.
- The quantum kinetic energy of H and O atoms in PIMD show strong NQE (for H, exceeding the classical value by >4×), matching deep inelastic neutron scattering for supercooled water and high-level model benchmarks (MB-pol, NEP-MB-pol).
- The student model, unlike those trained on DFT-only, can capture subtle NQE-induced density differences (Δρ ≈ 0.004 g/cm³ near ambient), mirroring the experimental H2O–D2O isotope variation.
These results demonstrate that knowledge distillation can compress gold-standard, composite potentials into practical surrogates for NQEs and thermodynamic observables, facilitating computationally demanding sampling otherwise beyond reach.
Case Study III: Proton Transfer at the Anatase TiO₂/Water Interface—Reactive and Quantum Chemistry
The most challenging application explores reactive interface chemistry: the free energy landscape for water dissociation at the anatase TiO₂(101)/water interface, a canonical system for photocatalysis and energy materials. Using fine-tuned foundation MACE teacher models (from varying amounts of DFT configurations), umbrella sampling is performed to extract the classical free energy landscapes for dissociation.
Different student models are then distilled from the teacher via umbrella-sampled trajectories, both for classical and quantum (TRPMD) nuclear treatments.

Figure 3: (a) Representative anatase TiO₂(101)-water interface configuration, (b) free energy profiles for water dissociation for teachers of varying training-set sizes, showing insensitivity for N > 1000.
The distilled classical student matches the teacher's free energy profile (barrier, ΔG) with <0.3 kcal/mol deviation and 10× higher throughput. Scratch-trained small MACE models on the same DFT data systematically underestimate the dissociation barrier and overstabilize the dissociated state, confirming that distillation is necessary for accurate, efficient reactive sampling.
The transition to quantum sampling (PIMD) requires not only broadening the student's configurational training set with both PIMD-sampled and rattled configurations but also increasing student model expressivity (MACE L=1, 32 channels). Only the concurrently augmented, higher-capacity quantum student maintains dynamical stability across all enhanced sampling windows.

Figure 4: (a) Speed–accuracy tradeoff for classical models in terms of dissociation barrier absolute deviation; (b) free energy profiles along the dissociation coordinate for classical and quantum students and the teacher. NQEs lower both barrier and relative free energy of the dissociated state.
Notable quantum effect: NQEs lower the effective reaction barrier by ≈2 kcal/mol and bring the predicted molecular–dissociated free energy difference into quantitative agreement with solid-state 17O NMR experiments (ΔG ≈ +1.3 kcal/mol), outperforming the purely classical prediction (ΔG ≈ +2.1 kcal/mol). Barriers and equilibrium populations thus align with both kinetic isotope effect observations and experimental spectroscopic constraints, substantiating the quantum stabilization of dissociated/proton-shared interfacial configurations.
Practical and Theoretical Implications
This work delineates concrete guidelines for the efficient deployment of high-accuracy MLIPs in challenging condensed-phase environments:
- Data-efficient Surrogate Construction: Distillation from foundation/fine-tuned MLIPs enables rapid, robust surrogate models, suitable for long-time, large-system, or quantum sampling, as opposed to scratch-trained models which often lack ergodicity or stability outside the training domain [morrowHowValidateMachinelearned2023a, fuForcesAreNot2023].
- Model Selection as a Dynamic Trade-Off: The minimal necessary student model capacity scales with the configurational extent and complexity of the target application (thermodynamic bulk → quantum liquid → reactive interface), requiring systematic evaluation of expressivity versus efficiency.
- Configurational Coverage: Augmenting the distillation dataset, including synthetic rattling or PIMD sampling, is crucial when transitioning to quantum or reactive regimes, directly impacting the model's ability to sample rare or off-equilibrium events reliably [gardnerDistillationAtomisticFoundation2025a].
- Agreement with Experiment: The quantum student model at the TiO₂/water interface reproduces experimentally inferred reaction thermodynamics—a stringently sensitive validation—demonstrating the feasibility of using ML-distilled surrogates to bridge gaps between simulation and chemical reality.
- Limitations and Future Pathways: Test-set errors are not fully predictive of dynamical stability or MD accuracy (especially for rare, reactive, or quantum events). There is a need to relate sampled configuration-space complexity and minimum model expressivity/coverage, and incorporate active learning or uncertainty-guided feedback into the training workflow [zaverkinUncertaintybiasedMolecularDynamics2024, kulichenkoUncertaintydrivenDynamicsActive2023].
Conclusion
Combined transfer learning and knowledge distillation enable the construction of compact, high-performance machine learning potentials that maintain (or closely approach) the accuracy of large, expressive foundation models. This methodology provides an actionable route to using first-principles-fidelity MLIPs for computationally intensive sampling regimes—long NPT trajectories, enhanced rare-event sampling, and PIMD—spanning crystalline solids, quantum liquids, and complex, reactive interfaces.

Figure 1: Speed–accuracy tradeoff for ice Ih demonstrates the efficiency gains attainable by distilling compact students from a finetuned foundation model.

Figure 2: The distilled student captures density, diffusion, and NQE observables for liquid water across the entire liquid range, with accuracy comparable to CCSD(T)-level calculations and MB-pol.

Figure 3: The free energy landscape for water dissociation at the anatase TiO₂(101)/water interface shows that accurate classical sampling can be achieved with minimal DFT data via fine-tuning.

Figure 4: NQEs reduce both the reaction barrier and the equilibrium free energy difference for water dissociation, with the quantum student outperforming classical approaches in matching experiment.
This study establishes a robust protocol for practical, system-specific MLIP deployment and sets a benchmark for data- and computation-efficient simulation workflows. It further motivates the systematic integration of active learning and rigorous model selection for new classes of condensed-phase chemical, materials, and catalysis problems, with future developments likely to focus on automating coverage/complexity matching, architectural selection, and real-time adaptive distillation protocols.