Papers
Topics
Authors
Recent
Search
2000 character limit reached

Electrostatic Phenomenology Benchmarks for Machine-Learned Interatomic Potentials in Electrochemistry: Beyond the Energy-Force Metric

Published 14 Aug 2026 in cond-mat.mtrl-sci, physics.chem-ph, and physics.comp-ph | (2608.14153v1)

Abstract: Accurate treatment of long-range interactions in machine learning interatomic potentials (MLIPs) is essential for electrochemical simulations. However, aggregate energy and force errors alone are insufficient to establish an MLIP's physical accuracy since they do not detect qualitative inconsistencies in the model such as the prediction of image-charge attraction, dielectric screening, or charge transfer. We introduce a benchmark suite EPhEct (Electrostatic Phenomena for Electrochemistry) of focused test cases designed to evaluate MLIPs on electrochemically relevant physical phenomena. The tests probe for image-charge attraction at a metal electrode, the splitting between longitudinal and transverse optical phonons as a probe of ionic and electronic screening, the dipole moment of interfacial water, and Fermi-level pinning during ion discharge. These tests establish a qualitative diagnostic routine complementary to aggregate energy-force metrics.

Summary

  • The paper introduces EPhEct, a four-test benchmark suite that evaluates image-charge attraction, ionic and electronic screening, interfacial water dipoles, and Fermi-level pinning beyond aggregate energy-force errors.
  • The methodology uses inexpensive single-point DFT reference calculations and energy-based scaling tests to diagnose long-range interactions, charge redistribution, field response, and electrostatic behavior without requiring explicit charge access.
  • The paper shows why accurate individual water dipoles or low energy-force errors can still conceal major collective failures, including missed electronic polarization, false metallization, incorrect screening, and spurious charge transfer at interfaces.

The paper introduces EPhEct (Electrostatic Phenomena for Electrochemistry), a benchmark suite of four qualitative test cases designed to evaluate machine-learned interatomic potentials (MLIPs) on electrostatic phenomena that aggregate energy and force errors cannot resolve (2608.14153). The central claim is that standard "fit-to-energy/forces" metrics are structurally incapable of diagnosing whether an MLIP reproduces image-charge attraction, dielectric versus metallic screening, interfacial dipole moments, or Fermi-level pinning — two models with identical total error can differ qualitatively in these physically distinct modes. The suite is accompanied by a set of scaling tests that verify correct electrostatic behavior in its simplest form, using energies alone, so that even architectures without explicit charge access can be diagnosed.

Motivation: the scale gap and the limits of local MLIPs

DFT-based AIMD is confined to roughly 10210^2 atoms and 10–100 ps, whereas double-layer equilibration requires nanoseconds and water reorientation occurs on picosecond timescales; AIMD therefore never samples the relevant regime. MLIPs close this gap at four to six orders of magnitude lower per-step cost, but their local approximation — grounded in Kohn's nearsightedness principle and the Behler–Parrinello decomposition into atomic contributions within a cutoff — is fundamentally incompatible with charged interfaces. A strictly local model cannot reproduce the $1/r$ decay of Coulomb interactions beyond its cutoff, nor the non-local charge redistribution underlying metallic screening or dielectric polarization. The authors stress that range and non-locality must be distinguished: an interaction can be long-ranged yet locally determined (fixed charges with a $1/r$ tail), or short-ranged yet requiring non-local redistribution (charge equilibration under a global constraint).

A further subtlety motivates the phenomenological approach: thermally induced electric-field fluctuations at electrochemical interfaces often dominate over the externally applied field. Standard fitting metrics may then effectively describe the thermal "noise" while missing the weaker "signal" of the applied-potential response.

Taxonomy of long-range MLIPs

The paper organizes existing approaches along the taxonomy of Grasselli et al. Local models comprise explicit charges (trained on Hirshfeld-type partitions), implicit charges (TensorMol, PhysNet, latent Ewald summation, CHGNet's magnetic-moment proxy, AIMNet2), and implicit polarization via Wannier centers (DPLR). Non-local models comprise self-consistent QEq schemes (4G-HDNNP, kQEq, MPNICE, ReaxNet, CELLI), non-local representations (LODE), and non-local architectures (SpookyNet, SO3krates, Ewald message passing, AllScAIP, MACE-POLAR-1, Polar LES). Response learning — training on Born effective charges, polarizabilities, or field-dependent potentials (FIREANN, PNNP, RAZOR) — offers a complementary route without architectural modification. Constant-potential models such as CP-MACE and CPMPNN accept electron number as input; CPMPNN reports a three-orders-of-magnitude speedup over grand-canonical DFT for potential-dependent reaction barriers.

Two structural observations carry weight for practitioners. First, neither local nor non-local classes encodes long-range electrostatics as an explicit physical term unless deliberately added: the local class cannot reach it, and the non-local class learns its magnitude from data with no inductive bias toward the correct $1/r$ asymptote. Second, QEq schemes conflate metallic and dielectric screening — they impose a global Fermi level appropriate to conductors but physically incorrect for insulating regions sustaining a potential difference, and they can exhibit diffuse, unphysical polarization in water under applied fields.

Scaling tests

Before the headline benchmarks, the authors define four criteria for correct electrostatic behavior — correct charge magnitude, correct $1/r$ decay, correct field response, and physically constrained charge redistribution — and provide diagnostics extractable from energies alone:

  • Single ion in PBC: fitting E(L)E(L) against box length recovers the Madelung form E(∞)−αQ2/L+CQ/L3E(\infty) - \alpha Q^2/L + CQ/L^3, yielding the apparent charge.
  • Ion pair in open boundaries: direct verification of the Coulomb law.
  • Ions as potential probes: energy differences between opposite test charges yield the electrostatic potential in vacuum regions, assuming no charge transfer between probe and system.
  • Periodic slabs: the effective slab charge follows from the parabolic vacuum potential profile or from the linear-in-vacuum-thickness capacitor scaling E(v)=E0+Q2v/(24ϵ0A)E(v) = E_0 + Q^2 v / (24\epsilon_0 A).
  • Combined systems: testing whether spurious charge transfer occurs when subsystems with different intrinsic charging potentials are joined — a known failure mode of QEq, which mimics a global metallic Fermi level, and of purely local charge predictors, which risk macroscopic space charges unless trained with spatially separated charged species.

The EPhEct benchmark suite

All four primary tests require only single-point evaluations against DFT references — no production MD — making them cheap, reproducible, and architecture-agnostic. They target qualitative reproduction, since a scalar error metric cannot resolve distinct failure modes.

Image-charge effect. A solvated K ion is placed at varying distances above Pt(111). The ideal image interaction decays as U(z)=−q2/(16πϵ0z)U(z) = -q^2/(16\pi\epsilon_0 z), but the DFT calculation deviates from pure $1/z$ behavior because of finite lateral cell size, partial ionization of K via Fermi-level pinning, and the resulting charged-surface contribution. Replacing K with a fixed probe charge restores a linear $1/r$0–$1/r$1 dependence. Missing this interaction systematically underestimates interfacial capacitance and biases adsorption energetics by several $1/r$2 even after solvent screening attenuates it.

Ionic and electronic screening. Frozen-phonon displacements in elongated MgO supercells probe the LO/TO splitting, whose ratio equals $1/r$3 via the Lyddane–Sachs–Teller relation. This distinguishes electronic from ionic screening — a distinction the authors argue rescaled effective charges $1/r$4 cannot handle consistently, because internal and external fields then require different unphysical rescalings. In heterogeneous systems where the local dielectric constant varies, no current MLIP treatment solves the self-consistent local-polarizability or position-dependent screened Poisson problem; the authors state explicitly that, to their knowledge, such a scheme has not been combined with MLIPs.

Water dipole in slab geometry. Using Au(111)/water with a computational counter electrode, the total dipole is obtained either from the dipole-correction potential step ($1/r$5) or from Wannier centers. The key result, drawn from prior work by Yang et al., is a bold and seemingly contradictory observation: accurate individual water dipoles do not guarantee an accurate total dipole, and vice versa. Electronic polarization displaces Wannier centers by only ~0.001 Å per molecule at 0.2 V/Å — invisible at typical per-molecule accuracy — but collectively shifts $1/r$6 by ~0.5 e·Å across tens of molecules, roughly half its magnitude, consistent with $1/r$7. Models learning from local charge decompositions miss this signal entirely; conversely, a model trained naively on $1/r$8 halves each water's static dipole to compensate. Relatedly, short-range MLIPs have been shown to produce spurious frames with very large macroscopic dipoles ("false metallization") even when atom density profiles look correct. Benchmarking both quantities is therefore necessary.

Fermi-level pinning. A H atom is pulled from Au(111) into vacuum; band alignment drives charge transfer until H becomes neutral. Since MLIPs expose no density of states, the effective charge profile $1/r$9 is mapped through the force induced by a homogeneous applied field, which remains finite while the particle is charged and vanishes upon neutralization. This probes whether a model captures ion discharge during desorption or dissolution — a prerequisite for faithful constant-potential simulations.

Limitations and open questions

The authors are explicit about scope. Passing the tests is a necessary but not sufficient condition for quantitative accuracy in dynamical simulations; the benchmarks validate single-point statics, not sampling behavior. The image-charge test does not isolate a clean analytic reference because of finite-size effects and partial ionization, and the LO/TO test yields the ratio of dielectric constants but not each separately. Systematic evaluation of existing MLIPs against the suite is deferred to future work — the paper provides DFT references and protocols but no model rankings. Whether self-consistent local-dielectric treatments compatible with MLIPs can be constructed, and how to enforce subsystem charge neutrality architecturally rather than through training data, remain open.

Conclusion

EPhEct reframes MLIP validation for electrochemistry around qualitative physical correctness rather than aggregate error. Its four tests — image charge, LO/TO screening, interfacial water dipoles, and Fermi-level pinning — map onto four axes (range, non-locality, field response, charge state) and expose failure modes that identical energy-force errors conceal, most notably the collective electronic polarization signal that dominates the total dipole of electrified water interfaces. Because the tests demand only single-point evaluations, they constitute a low-cost diagnostic layer that developers and practitioners can apply before committing an architecture to production-scale interfacial simulation.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Tweets

Sign up for free to view the 1 tweet with 0 likes about this paper.