LAMBench: Evaluating Large Atomic Models
- LAMBench is a benchmarking system for large atomic models that evaluates machine-learning potentials using three key metrics: generalizability, adaptability, and applicability.
- It employs a modular, high-throughput workflow with ASE calculators and Dflow workflows to automate task execution and aggregate performance metrics across diverse datasets.
- The benchmark highlights trade-offs between prediction accuracy, energy conservation, and runtime efficiency, emphasizing the need for cross-domain multitask training.
LAMBench is a benchmarking system for Large Atomic Models (LAMs), namely foundation-style machine-learning potentials intended to approximate a universal potential energy surface defined by first-principles calculations of atomic systems. It evaluates LAMs along three dimensions—generalizability, adaptability, and applicability—to assess their use as ready-to-use tools across diverse scientific discovery contexts, and its initial release benchmarked eight state-of-the-art LAMs released prior to April 1, 2025 (Peng et al., 28 Apr 2025).
1. Conceptual basis and scientific target
LAMBench is grounded in the view that a LAM should approximate the ground-state potential energy surface under the Born–Oppenheimer approximation. In this setting, the target is a scalar functional
with forces obtained as
For periodic simulations, virials are also relevant,
This formulation places LAMs in direct relation to first-principles electronic-structure methods, especially Kohn–Sham DFT, while recognizing that real datasets are generated with different exchange–correlation functionals and pseudopotentials (Peng et al., 28 Apr 2025).
LAMBench was introduced because atomistic ML evaluation had remained fragmented. Domain-specific benchmarks such as QM9, MD17, OC20, and Matbench Discovery measure narrow task families rather than the extent to which a model approximates a universal potential energy surface. Standard reporting based on train/test splits within a single dataset emphasizes interpolation accuracy, but does not directly test out-of-distribution behavior, derivative-sensitive property prediction, or molecular-dynamics stability. The benchmark also responds to the problem of non-conservative force fields, which may achieve strong static force MAE while violating basic physical requirements and degrading downstream simulation quality (Peng et al., 28 Apr 2025).
2. Benchmark architecture and the three evaluation axes
LAMBench is implemented as a modular benchmarking workflow, described as a high-throughput system that automates job execution through ASE calculators and Dflow workflows, then aggregates metrics by domain, task, and model. Its design separates model configuration, task execution, and result visualization, and it is accompanied by an open-source toolkit and an interactive leaderboard (Peng et al., 28 Apr 2025).
The benchmark is organized around three axes:
| Axis | Scope | Principal outputs |
|---|---|---|
| Generalizability | OOD force-field and property-calculation tasks | , |
| Adaptability | Property fine-tuning on downstream regression tasks | 5-fold cross-validation MAE |
| Applicability | Runtime efficiency and MD stability | , |
Generalizability measures out-of-distribution performance on independent downstream datasets and on property calculations that probe the smoothness and differentiability of the learned potential. Adaptability asks whether a pretrained LAM can be fine-tuned efficiently for downstream property prediction by attaching a new head to the pretrained descriptor. Applicability measures whether the model is practically usable as a simulator, emphasizing inference time per atom and long-horizon NVE stability rather than only static prediction error (Peng et al., 28 Apr 2025).
This structure makes LAMBench more than a leaderboard for energy and force RMSE. It evaluates whether a pretrained model behaves like a broadly deployable atomistic foundation model.
3. Generalizability: task domains and normalized error metrics
LAMBench divides generalizability into force-field generalizability and property-calculation generalizability. The force-field suite uses 17 OOD datasets across 5 domains: Inorganic Materials, Catalysis, Reactions, Small Molecules, and Biomolecules & Supramolecules. These datasets provide energies, forces, and, where available, virials. The property-calculation suite uses two derivative-sensitive benchmarks: MDR Phonon for Inorganic Materials and TorsionNet500 for Small Molecules (Peng et al., 28 Apr 2025).
Force-field errors are computed as RMSE for energies, forces, and virials, but LAMBench does not compare raw values directly across tasks. Instead it normalizes each model against a dummy model that predicts energy from chemical formula only. For model , domain , prediction type , and test set 0, the normalized error is
1
A value of 2 means performance equal to the dummy baseline, and lower values are better. Domain- and prediction-type-specific errors are then log-averaged:
3
These are combined within each domain using weights 4,
5
and the overall force-field generalizability is
6
with 7 domains (Peng et al., 28 Apr 2025).
Property-calculation generalizability uses the same normalization strategy, but with MAE rather than RMSE. In MDR Phonon, the predicted quantities are maximum phonon frequency 8, entropy 9, free energy 0, and heat capacity 1, each assigned weight 2. In TorsionNet500, the quantities are torsion-profile MAE, barrier-height MAE, and the count of molecules with large barrier-height error, each assigned weight 3. The resulting overall metric 4 is again dimensionless and lower-is-better (Peng et al., 28 Apr 2025).
The normalization scheme is central to the benchmark’s philosophy. It permits aggregation across tasks with different units, label conventions, and DFT settings, while still penalizing models that fail to add meaningful structure-dependent information beyond composition-only baselines.
4. Adaptability and applicability protocols
Adaptability is evaluated by fine-tuning a pretrained LAM for downstream property prediction. The benchmark currently instantiates this protocol mainly for DPA-2.4-7M, treating the pretrained descriptor as a reusable encoder and attaching a new property head. Eight Matbench regression tasks are used: formation energy, band gap, JDFT2D exfoliation energy, phonon maximum frequency, dielectric constant, log bulk modulus, log shear modulus, and perovskite formation energy. The setup uses 5-fold cross-validation, mean pooling for atom-wise predictions, and compares three regimes: training from scratch, fine-tuning from pretraining, and comparison with specialized property models such as MatterSim property heads and JMP (Peng et al., 28 Apr 2025).
Applicability is split into efficiency and stability. Efficiency is defined from average inference time per atom over 1600 structures from the Inorganic Materials and Catalysis domains:
5
measured in 6. The dimensionless efficiency score is
7
Larger 8 indicates faster inference. To avoid JIT artifacts, jobs are warmed up with 400 pre-evaluations (Peng et al., 28 Apr 2025).
Stability is defined through energy drift in 10 ps NVE molecular dynamics with 1 fs timestep, after a 2 ps warm-up, on 9 representative systems spanning materials, biomolecules, and catalysts. For each trajectory, the total energy per atom is linearly regressed against time to obtain a drift slope 9, which is compared to a tolerance 0 eV/atom/ps. The per-system instability metric is
1
and the overall instability is
2
Smaller 3 indicates better energy conservation and more reliable conservative dynamics (Peng et al., 28 Apr 2025).
These two applicability metrics are deliberately orthogonal. A model may be fast but unstable, or stable but too slow for routine use. LAMBench therefore treats applicability as a multi-objective property rather than collapsing it into static accuracy.
5. Benchmarked models and comparative results
The initial LAMBench release evaluates eight LAMs differing in architecture, training scope, conservativeness, and data fidelity:
| Model | Parameters | Training scope / distinctive feature |
|---|---|---|
| DPA-2.4-7M | 6.64M | Multitask OpenLAM, 31 datasets, conservative, 31 task heads |
| MACE-MP-0 medium | 4.69M | Single-task MPtrj, conservative E(3)-equivariant model |
| MACE-MPA-0 medium | 9.06M | MPtrj + sAlex, larger-capacity conservative MACE |
| Orb-v2 | 25.2M | Multi-dataset materials model with direct force prediction, non-conservative |
| SevenNet-MF-ompa | 25.7M | Multi-fidelity training on OMat24, MPtrj, and sAlex |
| SevenNet-l3i5 | 1.17M | Smaller single-task MPtrj conservative model |
| MatterSim-v1-5M | 4.55M | Materials-focused model with broad thermodynamic coverage |
| GRACE-2L-OAM | 15.3M | Graph Atomic Cluster Expansion trained on MPtrj |
Across overall force-field generalizability, the reported scores are approximately: DPA-2.4-7M 0.265, GRACE-2L-OAM 0.340, SevenNet-l3i5 0.355, MACE-MPA-0 0.356, Orb-v2 0.356, SevenNet-MF-ompa 0.358, MatterSim-v1-5M 0.389, and MACE-MP-0 0.405. Across overall property-calculation generalizability, the scores are: DPA-2.4-7M 0.208, SevenNet-l3i5 0.240, GRACE-2L-OAM 0.262, MatterSim-v1-5M 0.280, MACE-MPA-0 0.291, SevenNet-MF-ompa 0.300, MACE-MP-0 0.341, and Orb-v2 0.560 (Peng et al., 28 Apr 2025).
Several findings structure the benchmark’s interpretation. First, DPA-2.4-7M is the strongest overall model in both force-field and property-calculation generalizability, which the benchmark attributes to cross-domain multitask training across materials, molecules, reactions, catalysts, biomolecules, and multi-fidelity data. Second, increasing materials-domain data scale improves some models substantially: for example, MACE-MPA-0 improves over MACE-MP-0, especially within Inorganic Materials, though this does not translate uniformly to all domains. Third, multi-fidelity training helps within its target domain: SevenNet-MF-ompa performs better than SevenNet-l3i5 in Inorganic Materials and Reactions, but not uniformly elsewhere (Peng et al., 28 Apr 2025).
Property tasks sharpen the distinction between conservative and non-conservative models. In phonon prediction, conservative models such as SevenNet-MF-ompa and MatterSim perform strongly, whereas Orb-v2, despite competitive static force-field scores, is dramatically worse on higher-order derivative-sensitive tasks. The same pattern appears in stability: conservative models have very small energy drifts, with 4 near 0–0.089, whereas Orb-v2 has 5 and an energy drift of about 222.8 meV/atom/ps. Conversely, Orb-v2 is the most efficient model by the benchmark’s runtime metric, with 6, while DPA-2.4-7M is reported at about 0.614, GRACE-2L-OAM at about 0.678, and SevenNet-MF-ompa as the slowest at about 0.088 (Peng et al., 28 Apr 2025).
The benchmark therefore isolates a central design trade-off. Direct force prediction can improve throughput, but the loss of conservativeness is strongly penalized in phonons and NVE stability. By contrast, cross-domain conservative models are slower, but behave more like general-purpose simulators.
6. Role in subsequent LAM research and evolving significance
LAMBench was introduced as a dynamic and extensible platform, and later work immediately used it as a zero-shot generalization testbed. The DPA3 study describes LAMBench as a benchmark suite “intended to demonstrate the generalizability of LAMs in addressing real-world scientific challenges,” and reports that DPA-3.1-3M, trained on OpenLAM-v1, achieves the lowest overall zero-shot generalization error across 17 downstream tasks in LAMBench, with further gains when exchange–correlation-aware dataset encoding is used in the bestXC setting (Zhang et al., 2 Jun 2025).
This later use clarifies the benchmark’s role within the LAM ecosystem. LAMBench is not merely a score aggregation framework; it functions as a standardized test of whether a pretrained model behaves as an out-of-the-box potential model across heterogeneous chemistries, DFT settings, and downstream scientific objectives. The benchmark’s own conclusions point in the same direction: current LAMs remain far from an ideal universal potential energy surface, and progress requires cross-domain training data, multi-fidelity modeling, and architectures that preserve conservativeness and differentiability (Peng et al., 28 Apr 2025).
A plausible implication is that LAMBench occupies, for atomistic foundation models, a role analogous to that of broad benchmark suites in language and vision: it operationalizes a claim of universality. Because its codebase and leaderboard are open, and because its task set is designed to evolve, it also provides a mechanism for turning architectural claims—such as better cross-domain transfer, higher-fidelity conditioning, or improved MD stability—into directly comparable empirical statements.