Papers
Topics
Authors
Recent
Search
2000 character limit reached

AquaGen: Scaling generative models to molecular dynamics precision on thousands of atoms

Published 3 Jul 2026 in physics.chem-ph and cs.LG | (2607.03513v1)

Abstract: We present AquaGen, the first all-atom, explicit solvent, periodic-boundary-condition-aware generative model that produces molecular configurations from the Boltzmann distribution at a fraction of the cost of molecular dynamics (MD). This is in contrast with existing generative models that remove degrees of freedom by operating on coarse-grained, vacuum, or implicit solvent systems. Operating at this resolution allows for post-processing through force field energy evaluations and MD simulations, and enables the prediction of relevant properties in a gray-box manner (as ensemble averages of potential energy evaluations over generated samples). We demonstrate the utility of this paradigm on absolute hydration free energy (AHFE), producing estimates 4-10x faster and with comparable accuracy to standard GPU-based MD. By generating uncorrelated samples from alchemical Boltzmann distributions, we create more accurate, interpretable, and refinable ensemble predictions with calibrated uncertainty estimates, unlike regression methods which are entirely black-box predictors. Our approach also yields predictable benefits from increasing train- and test-time compute, realized by scaling model size and generating more samples, respectively. We believe that this approach demonstrates the utility of high-resolution ensemble generation for free energy estimation, with future potential to replace MD in tasks such as the prediction of lipophilicity, membrane permeability, or absolute binding free energy (ABFE) -- whose grounding and interpretability may be critical for the development of new drugs and materials.

Summary

  • The paper introduces a conditional flow matching model using a GNN to generate explicit-solvent, all-atom ensembles that accurately mimic Boltzmann distributions for systems up to ~4000 atoms.
  • The paper demonstrates that AquaGen achieves sub-kcal/mol errors in AHFE estimation and 4–10× computational speedup over GPU-accelerated MD, validated through multiple structural metrics.
  • The paper highlights AquaGen’s capability for calibrated uncertainty quantification and ensemble refinability, enabling its integration into practical chemical discovery workflows.

AquaGen: Scaling Generative Models to Molecular Dynamics Precision on Thousands of Atoms

Introduction and Context

The paper "AquaGen: Scaling generative models to molecular dynamics precision on thousands of atoms" (2607.03513) presents a generative modeling framework targeted at replicating the Boltzmann distribution for all-atom, explicit-solvent molecular systems with periodic boundary conditions (PBCs). The central aim is to produce statistically and physically accurate atomic ensembles for solvated small molecules, suitable for direct use in free energy calculations, particularly absolute hydration free energy (AHFE) estimation. This approach stands in contrast to prior models that leverage vacuum, implicit solvent, or reduced-dimensionality representations, and thus cannot interface seamlessly with high-fidelity energy functions used in molecular dynamics (MD) or Monte Carlo (MC) simulations.

AquaGen leverages a conditional flow matching generative model, trained on over one billion MD frames along alchemical pathways, to model Boltzmann-distributed all-atom configurations in the presence of explicit water, scaling to systems with up to approximately 4×1034 \times 10^3 atoms. The model is directly compatible with downstream force field energy evaluations, facilitating physically interpretable and refinable ensemble-based property prediction for chemical discovery.

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1

Figure 1: AquaGen framework overview – flow matching GNN generates PBC-aware, explicit-solvent atomistic configurations for ensemble-based AHFE estimation, achieving sub-kcal/mol errors and significant speedup over MD.

Methodological Details

Generative Framework and Flow Matching

AquaGen is based on conditional flow matching, where a time-dependent velocity field vθ(xˉτ,τλ)v_\theta(\bar{x}_\tau, \tau | \lambda) is parameterized with a GNN to deterministically transport samples from a Gaussian prior p0p_0 to the target alchemical Boltzmann distributions pdata(λ)p_{\text{data}}(\cdot|\lambda) indexed by the alchemical parameter λ\lambda. Both coordinates and simulation box parameters are modeled, and the architecture incorporates periodic boundary conditions by augmenting the graph with virtual nodes.

The model is conditioned on λ\lambda to generate frames that mirror the alchemical states typically simulated via Hamiltonian replica exchange MD for free energy perturbation. Training uses linear interpolation between prior and data samples for target velocity supervision, and inference integrates the velocity ODE from prior to the target distribution using an exponential time schedule to resolve high velocity field curvature near the prior.

Dataset and Training

The training dataset comprises multiple thousands of drug-like molecules, simulated in explicit water with OpenFF 2.1.1 and TIP3P solvent, following a 20-window alchemical protocol (electrostatics, then van der Waals decoupling). Samples were extracted such that models learn both the physically relevant solute-solvent energetic interactions and structural ensemble properties at each stage. Over one billion training examples ensure that the model captures realistic phase space distributions and subtle solvent-mediated effects.

Evaluation Protocol

AquaGen-generated ensembles are used to perform AHFE estimation using the MBAR estimator, replacing inefficient MD sampling with fast, model-based generation. This "gray-box" approach—melding black-box generative modeling with white-box force field energy evaluation—enables test-time uncertainty quantification and refinability via short MD initiated from generated samples.

Results: Energetic and Structural Fidelity

AHFE Accuracy and Computational Efficiency

AquaGen achieves strong quantitative performance on AHFE estimation across a large held-out test set. The median and mean AHFE absolute errors relative to ground-truth MD estimates are 0.93 kcal/mol and 1.22 kcal/mol, respectively, on diverse, drug-like molecules. Notably, AquaGen achieves a 4–10× reduction in computational cost relative to GPU-accelerated MD, with the error further reducible to <0.5 kcal/mol after brief (40 ps) MD refinement.

Test-time scaling and model size scaling are both demonstrated: increasing the number of generated configurations or enlarging the model consistently reduces the estimation error, with the 160M-parameter model benefiting most from higher sample counts.

Figure 2

Figure 2

Figure 2

Figure 2

Figure 2

Figure 2

Figure 2: At λ=1\lambda=1 (fully interacting), AquaGen samples reproduce reference MD energies and O–O RDFs; slight over-ordering and minor deviations are observed in the solute-solvent energy component.

Ensemble Distribution Alignment

Energetic accuracy is complemented by detailed structural metrics. AquaGen-generated ensembles at the fully interacting endpoint match reference energy distributions and RDFs Figure 2, indicating high fidelity in both global and local statistics. tICA projections of generated configurations along the alchemical path show that the model samples the appropriate regions of conformational space at each λ\lambda.

Figure 3

Figure 3

Figure 3

Figure 3

Figure 3

Figure 3

Figure 3

Figure 3

Figure 3: AquaGen samples (points) align with ground-truth MD ensembles (contours) along the alchemical tICA axes and capture solute-solvent interaction trends as λ\lambda is varied.

Key physical trends, such as hydrogen bond donor-acceptor distances evolving non-monotonically with λ\lambda, and box size variation, are reproduced. Some systematic overestimation of box size is noted, attributed to prior variance selection.

Uncertainty Quantification and Error Structure

A salient advantage of the generative, ensemble-based approach is the ability to extract calibrated uncertainties from bootstrapped sample predictions. Confidence interval widths strongly correlate with the actual prediction error, allowing for more reliable decision-making pathways (Figure 4a).

No significant dependence of AHFE error on training set similarity was observed, indicating the model does not simply memorize but generalizes to novel molecular scaffolds (Figure 4b). Error cancellation along the alchemical path is characterized: overestimation in the electrostatic decoupling regime is counteracted by underestimation in the van der Waals decoupling, introducing fortuitous cancellation and underscoring the importance of alternative metrics such as cumulative absolute error.

Figure 4

Figure 4

Figure 4

Figure 4

Figure 4

Figure 4

Figure 4: (a) Bootstrap CIs provide well-calibrated uncertainty; (b) weak dependence of error on training similarity implies robust generalization; (c) error cancellation along the alchemical path affects total AHFE error.

Comparison to Black-Box Methods and Ablation Insights

AquaGen is compared to random forest and GNN black-box regressors on multiple splits, including challenging out-of-distribution sets (e.g., FreeSolv, CombiSolv, and "target" extrapolation). AquaGen with brief MD refinement consistently outperforms all baselines, achieving <0.9 kcal/mol mean absolute error even on the most challenging splits. By contrast, black-box regressors are less reliable and generalize less effectively on compounds dissimilar from the training data.

Ablation studies probe the effects of architectural and training modifications. Notably, optimal Gaussian prior variance and exponential integration step schedules are critical for low-cancellation, physically meaningful predictions. Some architectural biases (e.g., center-of-mass centering) reduce mean error by enhancing cancellation but degrade actual physical ensemble quality.

Implications and Future Directions

AquaGen provides the first demonstration that uncorrelated, explicit-solvent, all-atom Boltzmann samples can be generated with high fidelity and efficiency for pharmaceutically relevant compounds at scale. The approach enables ensemble-based gray-box prediction of thermodynamic observables, uncertainty quantification, and physically meaningful downstream refinability—for example, rapid AHFE refinement by short MD trajectories. This stands in contrast to black-box regression, which lacks interpretability and refinability, and to previous generative models that neglect solvent or operate on reduced representations.

Anticipated directions include scaling AquaGen to even larger systems (10,000–100,000 atoms), modeling more complex environments (heterogeneous solvents, membranes, protein-ligand complexes), improving vθ(xˉτ,τλ)v_\theta(\bar{x}_\tau, \tau | \lambda)0-conditioned expressivity, and incorporating target-driven fine-tuning for experimental alignment. The potential to broadly replace MD for routine free energy problems (e.g., lipophilicity, permeability, ABFE) is a tangible practical trajectory. Theoretically, AquaGen demonstrates the first practical union of learned generative ensemble models and classical statistical mechanics for high-dimensional, phase-space-constrained systems.

Conclusion

AquaGen establishes a novel paradigm in which scalable, GNN-based conditional flow matching enables the rapid and accurate generation of all-atom, explicit-solvent Boltzmann ensembles that are compatible with industrial force fields and direct thermodynamic computation. Sub-kcal/mol errors are achieved at a fraction of MD computational expense, with robust uncertainty quantification and the capacity for interpretable, refinable predictions. This work motivates further integration of generative models with atomistic simulation pipelines, promising significant impact for both theoretical methodologies and practical chemical discovery workflows.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.