Papers
Topics
Authors
Recent
Search
2000 character limit reached

Transferable Implicit Solvent Machine Learning Potential for Drugs and Proteins Approaching Ab Initio Accuracy

Published 12 Jul 2026 in physics.chem-ph, cs.LG, and q-bio.BM | (2607.10887v1)

Abstract: Machine learning interatomic potentials (MLPs) have revolutionized atomistic modeling, offering the potential to replace traditional methods like Density Functional Theory (DFT). However, inference time of MLPs is orders of magnitude slower than that of classical force fields, hindering real-world applications for biomolecular systems that require timescales of microseconds and beyond. Implicit solvent MLPs can address this issue, but are faced with data challenges associated with coarse-grained modeling. Consequently, previous approaches relied on empirical force field data, thereby inherently limiting the MLP's accuracy. Here, we introduce the Transferable Water Implicit Network (TWIN), an implicit water MLP parametrized entirely by an Equivariant Graph Neural Network and trained solely on ab initio and experimental labels. We demonstrate TWIN's transferability across drug-like molecules, peptides, and proteins, achieving excellent results on ab initio and experimental crystallographic and NMR benchmarks, consistently outperforming previous machine-learning-based implicit solvent or coarse-grained models. Furthermore, TWIN closely matches DFT-based explicit solvent MLPs while providing a two-order-of-magnitude faster timestep evaluation, paving the way for efficient ab initio-level modeling of biomolecular systems in aqueous environments.

Authors (2)

Summary

  • The paper introduces TWIN, a GNN-based implicit solvent MLP that surpasses classical force fields with near ab initio accuracy and reduced computational cost.
  • It leverages a multiscale training protocol combining DFT parametrization, force-matching on proteins, and experimental refinement for precise solvation and interaction energy predictions.
  • The method delivers improved performance on drug, peptide, and protein simulations, demonstrating lower energy and structural errors compared to conventional models.

Transferable Implicit Solvent Machine Learning Potentials for Biomolecular Simulations

Introduction and Problem Context

Atomistic modeling of biomolecules has transitioned from empirical force fields (FFs) to machine learning interatomic potentials (MLPs), driven by the latter's capacity to approach quantum mechanical accuracy at substantially lower computational cost. However, MLPs remain orders of magnitude slower than empirical FFs, especially when explicit solvent is modeled—a critical limitation for simulations aiming to probe microsecond and larger timescales in realistic aqueous environments. Implicit solvent approaches have traditionally relied on classical models, such as generalized Born (GB) treatments, which, despite their computational efficiency, lack the fidelity required for accurate structural and dynamical predictions, especially in the context of proteins, peptides, and drug-like compounds.

The development of implicit solvent MLPs has generally proceeded by training on empirical FF data, restricting the ultimate attainable accuracy and transferability. The paper introduces the Transferable Water Implicit Network (TWIN) model, a graph neural network-based implicit solvent MLP that circumvents this bottleneck by exclusively relying on ab initio and experimental data for parameterization, aiming to achieve both near ab initio accuracy and the efficiency required for practical biomolecular simulations (2607.10887). Figure 1

Figure 1: Multiscale parametrization of TWIN, showing the training pipeline from DFT-labeled SPICE data, through protein force-matching, to top-down refinement against solvation thermodynamics.

Architecture and Training Paradigm

TWIN is constructed upon the MACE architecture (symmetry-equivariant message passing), with a three-stage multiscale training protocol:

  1. Bottom-Up DFT Parametrization (TWIN-AT): The model is initially trained on the SPICE dataset, which provides a comprehensive coverage of DFT-labeled small molecules, fragments, and biomolecular motifs. TWIN-AT achieves force MAEs (24.6 meV/Å) and energy MAEs (1.5 meV/atom) competitive with prior foundational MLPs such as MACE-OFF, outperforming them on key water structure benchmarks (notably, gOOg_{OO} in bulk H2_2O).
  2. Force-Matching to Coarse-Grained (CG) Protein Ensembles (TWIN-FM): The model is further optimized using force-matching on ~4M configurations drawn from explicit-solvent MD trajectories of CATH protein fragments, relabeled with TWIN-AT, thereby bridging small-molecule and macromolecular regimes.
  3. Top-Down Experimental Refinement: The model undergoes final calibration by directly minimizing errors in predicted solvation free energies against the CombiSolv experimental dataset, using differentiable trajectory reweighting and the ReSolv protocol.

This progressive transfer of inductive biases, from small-molecule quantum datasets to experimental macroscopic observables, allows TWIN to interpolate between scales, supporting unprecedented coverage and accuracy for both drugs and proteins in aqueous solution.

Numerical Benchmarks and Evaluation

Small Molecules and Drug-Like Fragments

TWIN-AT delivers sub-kcal/mol accuracy (mean barrier error 0.72 kcal/mol) on challenging benchmarks for biaryl torsional barriers against high-level CCSD(T)* reference surfaces, outperforming classical FFs by a wide margin. Figure 2

Figure 2: Performance of TWIN on biaryl torsions and macrocycles, including comparison of mean absolute errors versus reference QM and FFs.

For drug molecules such as ozanimod, TWIN recapitulates explicit solvent atomistic PMFs to within 0.3 kcal/mol, significantly better than classical FFs and classical-ML hybrid models (e.g., GNNIS). Macrocycle conformational ensembles as assessed by NOE violation rates also indicate reduced error relative to established alternatives, though some non-local correlations remain underrepresented due to the use of finite local environments in MACE.

Peptides

In Ala3\mathrm{Ala}_3 simulations, TWIN reproduces experimental Ramachandran populations and key NMR observables (J-couplings) on par with explicit solvent MLPs, preserving dominant mesostates (e.g., pPII) absent in classical FF-trained CG models, which default to α\alpha-helical bias. Figure 3

Figure 3: Configurational analysis of Ala3\mathrm{Ala}_3, showing backbone dihedral distributions and quantitative agreement with NMR-derived J-couplings.

Proteins

TWIN maintains native-like folded conformations in 10 ns unbiased MD of proteins spanning 1300 atoms, matching experimental RMSD/RMSF benchmarks. Compared to prior CG MLPs (e.g., CGSchNet) and classical FFs, TWIN achieves lower RMSD, realistic site-resolved flexibility, and improved reproduction of NMR order parameters (S2S^2, h3JNC^{h3}J_{NC}, Saxis2S^2_{axis} for side-chains). Figure 4

Figure 4: Stability and NMR order parameters (h3JNC^{h3}J_{NC}, backbone S2S^2, side-chain 2_20) for benchmark proteins.

A notable systematic trend is the underestimation of 2_21 in highly flexible segments, attributed to limitations in capturing long-range effects, which could be mitigated by future inclusion of explicit electrostatics or enhanced architectural non-locality.

Protein-Ligand and Solvation Thermodynamics

On the PLA15 benchmark, TWIN achieves a MAE of 0.48 kcal/mol against high-level QM protein-ligand fragment interaction energies, competitive with state-of-the-art semiempirical QM methods and foundational MLPs. In solvent, TWIN-simulated shifts in interaction energies align qualitatively with COSMO-RS and outperform DFTB3/PCM. Figure 5

Figure 5: Example protein-ligand complex (PLA15), showing error distributions in interaction energies and solvation-induced shifts compared to reference QM and continuum models.

In solvation free energies (CombiSolv aqueous subset), TWIN attains 0.76 kcal/mol MAE (close to experimental uncertainty), and on FreeSolv, a MAE of 0.96 kcal/mol, both surpassing classical FFs and matching or exceeding MLP baselines such as MACE-OFF24-SC.

Computational Efficiency

TWIN delivers one to two orders of magnitude reduction in per-timestep simulation cost relative to explicit-solvent MLPs, with greater gains as system size increases. Compared to classical FFs, there remains a ~252_22 performance gap, attributable primarily to implementation/MLP library bottlenecks, but this is mitigated by the much smoother energy landscape and the associated increase in sampling efficiency obtained by implicit treatment of solvent.

Theoretical Implications and Limitations

The TWIN framework demonstrates that it is feasible to construct transferable, accurate, and stable implicit solvent MLPs solely from ab initio and experimental thermodynamic data, forgoing reliance on any empirically parameterized FFs (or hybrid solute priors). This approach circumvents the chemical space limitations and parameterization barriers faced by prior CG models, supporting direct extension to arbitrary molecular classes—including noncanonical residues and drug fragments.

A systematic limitation is the lack of explicit long-range interactions in the present local message-passing framework, which impacts observables sensitive to distal correlations (e.g., some NOEs, 2_23 of highly flexible/disordered protein regions). Integration of long-range corrections or global architectures may address these issues. The methodology also assumes exhaustive sampling of relevant configurational and thermodynamic spaces in the experimental datasets—a primary constraint for further generalization, especially toward unfolded, disordered, or highly non-equilibrium states.

Future Directions

Anticipated extensions include incorporation of explicit long-range electrostatics, expansion to other solvents and environments, and integration with generative/coarse-grained modeling tools (e.g., Boltzmann generators). Multi-task and transfer learning approaches, combined with more diverse quantum-based datasets, are likely to further elevate the attainable accuracy and scope of the resulting models. The pipeline is immediately relevant for high-accuracy, scalable modeling of protein–drug interactions, rational molecular design, and solution-phase biophysical phenomena.

Conclusion

TWIN establishes a rigorous, scalable, and extensible paradigm for implicit solvent MLPs, delivering ab initio-level accuracy at a fraction of the computational cost for realistic biomolecular systems. By leveraging a multiscale (bottom-up plus top-down) data-centric protocol, TWIN regularly exceeds the performance of both classical FFs and previous CG MLPs across a range of structural, thermodynamic, and dynamical benchmarks in aqueous environments. While further theoretical and technical advances—especially in capturing long-range correlations—are warranted, the TWIN architecture marks a substantive step toward practical, transferable MLP-based simulation of complex biomolecular assemblies.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.