---
title: Molecular Augmented Dynamics (MAD)
url: https://www.emergentmind.com/topics/molecular-augmented-dynamics-mad
type: topic
---

# Molecular Augmented Dynamics (MAD)

Molecular augmented dynamics (MAD) is a modified molecular dynamics method in which the Hamiltonian is augmented by an experimental potential so that atomistic trajectories are driven simultaneously toward low-energy configurations under an interatomic potential and toward agreement with experimental observables. In the literature considered here, MAD is formulated as a multi-objective optimization framework for generating experimentally consistent atomistic structures by design, demonstrated for disordered carbon materials using machine learning interatomic potentials trained from ab-initio data, and subsequently extended with linear-scaling expressions for diffraction and local spectroscopic observables [2508.17132][2509.22388]. The acronym is not used uniformly across recent atomistic ML research, however, and careful nomenclature is necessary because “MAD” also denotes the Massive Atomistic Diversity dataset in PET-MAD and MAD-1.5, while MAD-SURF expands to Molecular ADsorption on SURFaces [2503.14118][2603.02089][2601.18852].

## 1. Definition and nomenclature

In its explicit methodological sense, **molecular augmented dynamics** is a dynamics-based structure-search procedure that augments standard MD with forces derived from experimental mismatch. The goal is not only to sample thermodynamic ensembles, but to identify structures that simultaneously match multiple experimental observables and exhibit low energies as described by a machine learning interatomic potential trained from ab-initio data [2508.17132].

| Usage | Meaning | Role |
|---|---|---|
| MAD | Molecular augmented dynamics | Modified MD with an experimental potential |
| MAD in PET-MAD and MAD-1.5 | Massive Atomistic Diversity | Dataset construction for universal MLIPs |
| MAD-SURF | Molecular ADsorption on SURFaces | MLIP for adsorption on coinage metal surfaces |

This terminological distinction matters because PET-MAD states that **MAD** in that work refers specifically to the **Massive Atomistic Diversity** dataset, not to a simulation protocol, and that no alternative or formal definition of “Molecular Augmented Dynamics” or an explicit “MAD” equation is presented beyond this dataset construction [2503.14118]. By contrast, the MAD method paper provides an explicit Hamiltonian, explicit experimental forces, and an implementation path in **TurboGAP** [2508.17132].

A common misconception is therefore to treat all “MAD” papers as instances of the same formalism. The current literature instead contains at least three distinct usages: a modified-MD method for experiment-constrained structure generation, a family of data-centric universal interatomic-potential resources based on Massive Atomistic Diversity, and acronymic model names such as MAD-SURF [2601.18852].

## 2. Hamiltonian formulation and force decomposition

The central formal modification of MAD is the replacement of the standard Hamiltonian
\[
\mathcal{H}_{\mathrm{MD}} = T + V
\]
with
\[
\mathcal{H} = T + V + \tilde{V},
\]
where \(T\) is kinetic energy, \(V\) is interatomic potential energy, and \(\tilde{V}\) is an **experimental potential** that penalizes disagreement between predicted and experimental observables [2508.17132].

The experimental potential is written as
\[
\tilde{V} = \frac{\gamma}{2}\left[\mathbf{w} \odot \left(\mathbf{h}_{\rm pred} - \mathbf{h}_{\rm exp}\right)\right]^2,
\]
with \(\gamma\) a scaling factor controlling the importance of experimental agreement, \(\mathbf{w}\) weights representing importance or inverse uncertainty, \(\mathbf{h}_{\rm pred}\) the predicted observable, and \(\mathbf{h}_{\rm exp}\) the experimental data [2508.17132]. In this construction, the observable-matching term acts as a bias toward experimentally consistent configurations.

The force on each atom has two contributions,
\[
f_k^{\alpha,\mathrm{tot}} = f_k^\alpha + \sum_{j=1}^{M} \tilde{f}_k^{\alpha,j},
\]
where \(f_k^\alpha\) is the physical force from \(V\), and each \(\tilde{f}_k^{\alpha,j}\) is an **experimental force** derived from the gradient of the experimental potential for observable \(j\) [2508.17132]. For a general observable,
\[
\tilde{f}_{k}^{\alpha} = -\frac{\partial \tilde{V}}{\partial r_k^\alpha}
= - \gamma \, \mathbf{w} \odot \frac{\partial \mathbf{h}_{\rm pred}}{\partial r_k^\alpha} \cdot \mathbf{w} \odot (\mathbf{h}_{\rm pred} - \mathbf{h}_{\rm exp}).
\]

The differentiability requirement is fundamental: the simulated observable must be a differentiable function of atomic positions. This is what allows MAD to include multiple observables simultaneously and, as stated in the method paper, to generalize beyond MD to any structure search or optimization procedure modified with the same experimental potential [2508.17132].

## 3. Observable models, differentiability, and linear scaling

The original MAD demonstration used **X-ray diffraction (XRD)**, **neutron diffraction (ND)**, **pair distribution function (PDF)**, and **X-ray photoelectron spectroscopy (XPS)** data, while emphasizing that the method is general and accepts any experimental observable whose simulated counterpart can be cast as a function of differentiable atomic descriptors [2508.17132]. The accompanying linear-scaling theory paper then derived general equations for MAD with linear-scaling formulations for calculating and matching X-ray/neutron diffraction and local observables, including the core-electron binding energies used in X-ray photoelectron spectroscopy [2509.22388].

For diffraction, the key computational issue is that the standard Debye equation scales as \(\mathcal{O}(N^2)\). The linear-scaling MAD formulation replaces that bottleneck with partial pair distribution functions estimated via neighbor lists and kernel density estimators, followed by windowed Fourier transforms in the Ashcroft-Langreth formalism [2509.22388]. The partial PDF is written as
\[
g_{ab}(r) = \frac{n_{ab}(r)}{4\pi r^2\,dr\,N_a\rho_b},
\]
and the corresponding partial structure factors are evaluated up to a finite cutoff \(r_{\rm cut}\), which enforces \(\mathcal{O}(N)\) scaling [2509.22388].

For local scalar observables, the framework is generalized to any quantity differentiable with respect to a local, position-dependent atomic descriptor. The paper’s explicit example is an XPS spectrum,
\[
\tilde{I}^{\rm XPS}(\varepsilon; \{\mathbf{r}\}) = \sum_i \exp\left(-\frac{(\varepsilon - \varepsilon^i_{\rm pred})^2}{2\sigma^2}\right),
\]
where the predicted binding energy \(\varepsilon^i_{\rm pred}\) depends on local descriptors \(\mathbf{q}^i(\{\mathbf{r}\})\) [2509.22388]. The corresponding gradients are obtained analytically by chain rule, which makes these observables compatible with MD integration.

The same work generalized the virial tensor with the experimental forces, enabling **generalized barostatting**. This allows one to find structures whose density matches that compatible with the experimental observables rather than fixing density a priori [2509.22388]. Scaling tests with **TurboGAP** demonstrated linear-scaling behavior for both CPU and GPU implementations, with the GPU implementation giving a **100\(\times\)** speedup compared to the CPU [2509.22388].

## 4. Interatomic potentials, workflow, and physical realism

MAD relies on an interatomic potential \(V\) that is accurate enough to prevent the structure search from converging to experimentally compatible but physically implausible configurations. In the main demonstration, the interatomic potentials were **machine learning potentials**, specifically **Gaussian Approximation Potentials (GAPs)**, used for accurate modeling of the potential energy surface at computational cost similar to empirical force fields [2508.17132]. The same source states that this removes the need for artificial connectivity or bonding constraints and enables efficient large-scale atomistic simulations involving thousands of atoms.

The operational workflow is an annealing-based MD protocol. Initial random structures are annealed, for example via a melt-and-quench protocol, while MAD minimizes a multi-objective function comprising both the physical energy and the deviation from experimental observable(s) [2508.17132]. Simulations may be run in **NVT** or **NPT**, and under NPT the modified virial tensor allows the experimental density to emerge from matching [2508.17132].

A critical methodological point is that the experimental bias is not intended to replace physical energetics. Rather, it supplements the interatomic potential so that structures remain **low-energy** while being driven toward experimental consistency. The method paper explicitly frames this as an improvement over earlier **RMC/HRMC**-type approaches, which could fit experimental data at the cost of realism [2508.17132]. The linear-scaling theory paper makes a related point from the algorithmic side, noting that prior approaches were hindered by poor observable gradient accuracy, computational inefficiency, \(\mathcal{O}(N^2)\) scaling, or inability to include complex local-probe observables [2509.22388].

Another important procedural detail is that, after removing \(\tilde{V}\), the system remains in a low-energy and experimentally compatible state [2508.17132]. This indicates that the method is not merely producing transiently biased configurations, but can locate metastable structures compatible with the chosen observables.

## 5. Demonstrated systems and outcomes

The initial MAD paper demonstrated feasibility on several carbon materials: **glassy carbon**, **nanoporous carbon**, **ta-C**, **a-C:D**, and **a-CO\(_x\)**, matching their respective **X-ray diffraction**, **neutron diffraction**, **pair distribution function**, and **X-ray photoelectron spectroscopy** data using the same initial structure and underlying MLP [2508.17132]. The stated objective was to identify representative structures of these disordered materials by direct optimization against experiment rather than by unconstrained sampling and retrospective selection.

The later linear-scaling paper sharpened the interpretation of these outcomes. It states that MAD simulations can both find **metastable structures compatible with non-equilibrium experimental synthesis** and **lower energy structures than alternative computational sampling protocols, like the melt-quench approach** [2509.22388]. This is an important distinction: when experimental synthesis is non-equilibrium, agreement with experiment may legitimately require metastable rather than minimum-energy structures, and MAD is formulated to capture that trade-off through the parameter \(\gamma\).

These results also clarify what MAD is not. It is not simply a method for accelerating equilibrium MD, nor is it only an inverse-fitting procedure akin to reverse Monte Carlo. Its defining feature is simultaneous optimization of interatomic potential energy and experimental agreement under MD dynamics [2508.17132][2509.22388]. The phrase from the original abstract that the method enables a computational “microscope” into experimental structures is therefore best understood as a claim about interpretable structure generation under experimental constraints, not about direct reconstruction from raw measurements alone [2508.17132].

## 6. Adjacent MAD ecosystems and broader research directions

Recent papers extend the practical landscape around MAD, although not always in the same formal sense. **PET-MAD** is a generally applicable machine-learning interatomic potential trained on the **Massive Atomistic Diversity** dataset, comprising **95,595 structures** and **85 elements**; it is intended to deliver near-quantitative, out-of-the-box accuracy across organic and inorganic materials, and PET-MAD can be efficiently fine-tuned with **LoRA** to deliver full quantum mechanical accuracy with a minimal number of targeted calculations [2503.14118]. In that work, “MAD” refers to dataset construction rather than a modified-Hamiltonian MD protocol.

**MAD-1.5** extends the original MAD dataset to **216,803 structures** and **102 elements**, using a single, standardized all-electron DFT workflow with the \(r^2\)SCAN meta-GGA functional and outlier removal using **Last-Layer Prediction Rigidity (LLPR)** uncertainty [2603.02089]. The associated **PET-MAD-1.5** models are presented as generally applicable \(r^2\)SCAN interatomic potentials with benchmark accuracy and stability in challenging simulation protocols [2603.02089]. This suggests a data-centric route for strengthening the interatomic potential component \(V\) in future MAD simulations.

**MAD-SURF** is a machine learning interatomic potential specifically tailored for molecular adsorption on **Cu(111)**, **Ag(111)**, and **Au(111)**, trained on over **77,000 DFT-calculated configurations** spanning adsorption structures, ab initio molecular dynamics, normal mode sampling, and non-covalent aggregates [2601.18852]. It achieves DFT-comparable accuracy while enabling simulations orders of magnitude faster than density functional theory and has been demonstrated on experimentally characterized systems including organic monolayers, polycyclic aggregates, flexible biomolecules, and the long-range herringbone reconstruction of gold [2601.18852]. Although MAD-SURF is not the experimental-potential formalism, it exemplifies the kind of specialized MLIP infrastructure that can support augmented atomistic workflows.

A different but related direction is represented by **MD-LLM-1**, described as the first large language model framework specifically adapted for molecular dynamics trajectory generation. Fine-tuned from **Mistral 7B**, it can learn protein dynamics and generate conformations characteristic of states not included in the training data, as shown for **T4 lysozyme** and **Mad2** [2508.03709]. The same paper states that the model does **not directly model** thermodynamics or kinetics. A plausible implication is that “augmentation” in atomistic simulation is now being pursued along three complementary axes: experimental-force augmentation of dynamics, dataset-driven augmentation of transferable interatomic potentials, and generative augmentation of conformational exploration [2508.03709][2503.14118].

Across these strands, the most stable formal meaning of **molecular augmented dynamics** remains the modified-MD method defined by
\[
\mathcal{H} = T + V + \tilde{V},
\]
with differentiable experimental observables supplying additional forces during atomistic evolution [2508.17132]. The surrounding ecosystem of MAD-named datasets, MLIPs, and generative models indicates an expanding research program centered on bringing experimental consistency, transferability, and efficient exploration into the core of atomistic simulation.

Source: https://www.emergentmind.com/topics/molecular-augmented-dynamics-mad