Energy Above Hull (Ehull) Overview
- Energy Above Hull (Ehull) is the excess formation energy of a phase above the convex hull, indicating its thermodynamic drive for decomposition.
- It quantifies stability by comparing formation energies against a set of competing phases, with a value of 0 representing stability and positive values signifying metastability.
- Computational evaluation of Ehull uses convex hull construction from formation energies via DFT and ML models, offering rapid screening for materials discovery.
Energy Above Hull (Ehull) is the excess formation energy of a phase above the lower convex envelope spanned by competing phases at the same overall composition. It is a direct measure of the thermodynamic driving force for decomposition: a phase on the hull has , whereas a positive value indicates metastability relative to a convex combination of stable phases. In fixed-pressure settings, the same construction is performed with formation enthalpy rather than internal energy, and the hull distance is often denoted . Across phase-diagram construction, high-throughput screening, and machine-learning-based materials assessment, Ehull functions as a key metric for thermodynamic stability, but its numerical value is always conditioned on the reference states, the enumerated competing phases, and the thermodynamic environment used to build the hull (Nong et al., 21 Jul 2025, Ishikawa et al., 2020).
1. Formal definition and geometric construction
For a multicomponent compound with composition , total energy , and elemental reference chemical potentials , the per-atom formation energy can be written as
For a compound , the energy above hull is then
where is the energy of the convex combination of thermodynamically stable competing phases matching the composition of (Nong et al., 21 Jul 2025).
Equivalent formulations express the hull value at target composition 0 as a constrained convex-combination problem,
1
with the hull distance given by
2
In this representation, the convex hull is the lower convex envelope of formation-energy points in composition space, and Ehull is the vertical separation between a candidate phase and that envelope (Li et al., 2018).
At fixed pressure, the relevant thermodynamic potential is enthalpy rather than internal energy. The 10 GPa carbon–hydrogen study therefore constructs a formation-enthalpy convex hull and defines the hull distance as
3
with 4 identifying stable hull vertices and small positive 5 identifying near-stable phases (Ishikawa et al., 2020).
2. Reference states, chemical environment, and thermodynamic scope
Ehull is not a purely intrinsic scalar of a composition; it is defined relative to a specific set of competing phases and a specific thermodynamic frame. The perovskite-oxide study emphasizes that Ehull depends on the chosen chemical environment, particularly the oxygen chemical potential. There, hulls were constructed in multicomponent A–B–O chemical spaces using Materials Project entries, and Ehull values were evaluated for conditions representative of SOFC cathodes: 6 K, 7 atm, with oxygen chemical potential 8 eV/O, corresponding to 9 eV/O relative to the Materials Project 0 reference; 1 was included via equilibrium with 2 and 3 at 30% relative humidity (Li et al., 2018).
The high-pressure hydrocarbon study makes the dependence on thermodynamic ensemble equally explicit. At 4 GPa, the hull was built from formation enthalpy, 5, relative to elemental reference phases at the same pressure. Zero-point energy and temperature were not included, so the reported 6 values are static-lattice, zero-temperature enthalpy distances (Ishikawa et al., 2020).
The recent benchmark on machine-learning interatomic potentials introduces a further reference issue: when comparing ML-predicted thermodynamic stability against DFT, the hull can either be built from DFT energies or self-consistently from predicted energies. That study uses self-referenced convex hulls for each model and task, converting the model’s own total energies into formation energies and constructing a model-consistent phase diagram, specifically to avoid mixing DFT hull energies with ML-predicted energies (Nong et al., 21 Jul 2025).
The three cited studies therefore illustrate distinct but compatible uses of Ehull:
| Study | Quantity on the hull | Condition-specific note |
|---|---|---|
| (Nong et al., 21 Jul 2025) | Formation energy and self-referenced 7 | Model-wise hulls built from predicted energies on Materials Project structures |
| (Li et al., 2018) | 8 in multicomponent oxide systems | 9 K, 0 atm, specified 1 |
| (Ishikawa et al., 2020) | Formation-enthalpy hull, 2 | 3 GPa, 4, static-lattice enthalpy |
These dependencies are visible in concrete examples. In the perovskite study, 5 has 6 meV/atom, whereas 7 has 8 meV/atom in the 9–0–1–2–3 system at the specified 4, relative to decomposition into 5, 6, 7, and 8 (Li et al., 2018). In the 10 GPa C–H hull, stable vertices occur at 9, 0, 1, 2, 3, and 4, corresponding to diamond, graphane, ethane, methane, 5, and 6 (Ishikawa et al., 2020).
3. Computational evaluation of Ehull
A general workflow for computing Ehull from predicted or calculated energies is explicit in the cited literature. First, one defines the chemical system and the set of candidate compounds. Second, one computes a per-atom formation energy or formation enthalpy relative to elemental references at the same thermodynamic conditions. Third, one constructs the lower convex envelope in composition space. Fourth, one evaluates the hull energy at the target composition by solving the convex-combination problem. Fifth, one subtracts the hull value from the candidate’s formation quantity to obtain Ehull or 7 (Nong et al., 21 Jul 2025, Ishikawa et al., 2020).
In practice, the perovskite and MLIAP studies used Pymatgen phase-diagram tools. The perovskite work built hulls in the relevant multicomponent spaces using all DFT-calculated materials in the Materials Project database as of December 2016 (Li et al., 2018). The MLIAP benchmark computed formation energies per model from predicted total energies for all compounds and unary references, then computed Ehull via PhaseDiagram in Pymatgen using the model’s self-consistent energy set (Nong et al., 21 Jul 2025).
When the environment differs from the database reference, the input energies must be shifted consistently before hull construction. For the oxide study, oxygen-containing energies were adjusted to the operating oxygen chemical potential, and the authors note that their formation energies can be shifted by a constant 8 eV/O to transform to other 9 values (Li et al., 2018). At fixed pressure, the high-pressure study instead relaxes structures at the target pressure and uses enthalpy directly, so the pressure dependence is absorbed into 0 and the reference enthalpies 1 (Ishikawa et al., 2020).
The 2025 MLIAP study also distinguishes three evaluation tasks before hull construction: “no-relax,” in which a single-shot energy is evaluated on DFT-relaxed inputs; “pos-relax,” in which atomic positions are relaxed with fixed lattice; and “vc-relax,” in which both positions and lattice are relaxed. Calculations were run via ASE with the FIRE optimizer and ExpCellFilter, with a 2 eV/Å force threshold and at most 500 steps, and were examined in as-queried, primitive, and conventional cells after symmetrization through SpaceGroupAnalyzer with symprec=0.1 (Nong et al., 21 Jul 2025).
4. Interpretation, thresholds, and metastability
The primary physical meaning of Ehull is decomposition driving force. In the formulation used for inorganic crystals, 3 for phases lying on the convex hull and 4 for metastable phases. The larger the positive value, the greater the thermodynamic incentive to decompose into a mixture of hull phases (Nong et al., 21 Jul 2025).
A common source of confusion is the distinction between strict thermodynamic stability and practical screening thresholds. The perovskite study states explicitly that any off-hull phase is strictly unstable, but nevertheless adopts 5 meV/atom as the “stable” classification threshold, based on typical DFT errors and prior analyses indicating that more than 80% of ICSD compounds have 6 meV/atom (Li et al., 2018). By contrast, the high-pressure C–H study uses a far stricter criterion, 7 mRy/atom, which corresponds to approximately 8 meV/atom, to nominate “near-stable” hydrocarbons at 10 GPa (Ishikawa et al., 2020).
The 2025 MLIAP benchmark introduces yet another threshold convention for error analysis rather than physical stability assignment. There, “near-hull” is defined as DFT 9 meV/atom, and practical correctness is assessed under equality thresholds of 0, 1, 2, 3, 4, and 5 eV/atom. With 6 eV/atom, the best model achieves only about 75% correctness, whereas approximately 99% correctness requires the threshold to be relaxed to 7 eV/atom (Nong et al., 21 Jul 2025).
These study-specific thresholds do not define a universal metastability criterion. Instead, they reflect different uses of Ehull: rigorous phase-diagram classification, empirical screening under DFT uncertainty, and benchmarking of ML surrogates. A plausible implication is that Ehull should be interpreted together with the thermodynamic ensemble, the completeness of the phase set, and the uncertainty model of the underlying energies. The cited works support that view directly: the oxide study notes that polymorphism, incomplete enumeration of competing phases, and the chosen 8 can inflate or deflate Ehull, while the high-pressure hydrocarbon study notes that zero-point motion, thermal effects, reference-state ambiguities, and limited search coverage can perturb 9 by several meV/atom (Li et al., 2018, Ishikawa et al., 2020).
5. Ehull as a target for machine-learning models
Ehull has been used both as a direct supervised-learning target and as a derived stability metric obtained from ML-predicted energies. In the perovskite-oxide study, machine-learning models were trained on 1929 DFT-calculated perovskite energies, with 791 retained features derived from elemental-property data and perovskite-specific structural descriptors such as the Goldschmidt tolerance factor, octahedral factor, and composition-averaged Shannon radii. For classification, the Extra Trees algorithm achieved accuracy 0, F1 score 1, precision 2, recall 3, and ROC AUC 4. For regression, kernel ridge regression with an RBF kernel achieved RMSE 5 meV/atom, MAE 6 meV/atom, and 7 in leave-out-20% cross-validation (Li et al., 2018).
That same study emphasizes both utility and scope limitations. Prediction speed was approximately 8 s on a 2-core CPU, whereas the DFT workflow required approximately 9–0 hours per compound on 24 cores, a speed difference of more than 1. At the same time, regression performance degraded for underrepresented elements or combinations, such as Ge/Sn-containing compositions, and near-threshold cases required caution because the cross-validation error scale was already on the order of tens of meV/atom (Li et al., 2018).
A substantially different picture emerges from the benchmark of universal machine-learning interatomic potentials. That study performed approximately 2 calculations across nine pretrained MLIAPs—CHGNet, MACE, ORB-MPtrj, SevenNet, eqV2, eqV2-DeNS, MatterSim, eSEN, and eSEN-MPtrj—on 153,235 inorganic crystal structures from Materials Project v2023.11.1. Despite the fact that 143,164 of those structures, approximately 93%, were present in the MPtrj training trajectories used by most models, systematic underprediction of total energy, formation energy, and especially Ehull persisted (Nong et al., 21 Jul 2025).
The quantitative outcome is central to current Ehull methodology. The best-performing model, ORB-MPtrj, achieved the lowest MAEs across tasks, with Ehull MAEs in the 12–31 meV/atom range for various settings, but the overall MAE in vc-relax remained above approximately 30 meV/atom. Errors increased strongly after relaxation: pos-relax and vc-relax roughly doubled the errors relative to no-relax, and the increase was mainly driven by positional relaxation. Near-hull underprediction was identified as the dominant misprediction mode, implying that metastable or marginally stable candidates can be falsely flagged as stable; conversely, near-hull overprediction can discard genuinely interesting candidates (Nong et al., 21 Jul 2025).
This contrast between a domain-specific perovskite model and universal relaxed-structure potentials suggests that ML success on Ehull is highly dependent on the representation, data distribution, and evaluation protocol. The cited studies converge on a practical conclusion: ML can be effective for rapid triage, but decisions close to fine Ehull margins still require targeted DFT validation (Li et al., 2018, Nong et al., 21 Jul 2025).
6. Symmetry-related error modes and current methodological directions
The 2025 benchmark identifies a specific failure mode in ML-derived Ehull: systematic energy underprediction associated with insufficient handling of symmetry degrees of freedom. In that study, symmetry DOF combine lattice symmetry and Wyckoff site symmetries. High-DOF structures suffer disproportionately: across models, the MAE for Ehull exceeds approximately 40 meV/atom in structures with high symmetry DOF, defined there as 3, and the most pronounced underprediction occurs in far-from-hull materials that also have high symmetry DOF, although near-hull materials still exhibit severe scatter and systematic underprediction (Nong et al., 21 Jul 2025).
The proposed mechanism is architectural rather than merely data-driven. Current MLIAPs enforce 4 equivariance under rotations, translations, and reflections, but do not explicitly encode invariance to the discrete symmetry operations of a crystal’s space group, including screws, glides, and Wyckoff multiplicities. In high-DOF structures, positional relaxation explores symmetry-breaking and bending modes that are handled poorly, yielding softened potentials and downward-biased energies. Pearson correlations on structures with 5 meV/atom show that absolute error correlates most strongly with symmetry DOF, lattice DOF, normalized site DOF, structure matching, and RMSD between input and relaxed structures, and only weakly with whether a structure is on the hull, its atom count, whether it is in MPtrj, or its explicit symmetry change upon relaxation (Nong et al., 21 Jul 2025).
Several control experiments sharpen the argument. ORB-MPtrj shows the lowest mismatch rate, approximately 2.4%, but still exhibits symmetry-related error. Applying external symmetry constraints during relaxation through FixSymmetry barely changes error statistics, and fine-tuning CHGNet on a curated subset with 6 does not remove the underprediction. The study therefore attributes the problem primarily to model architecture and representation rather than to data scarcity alone (Nong et al., 21 Jul 2025).
The methodological directions proposed in response are all symmetry-aware. They include explicit encoding of lattice and Wyckoff site DOFs as inputs or conditioning features, symmetry-regularized loss terms penalizing inconsistent energy changes under symmetry-preserving or symmetry-displacing moves, and architectures that respect space-group operations beyond 7 equivariance. Data curation is also treated as a first-order issue: curated training sets emphasizing low-symmetry or high-DOF structures, symmetry-compliant perturbations rather than purely Gaussian rattling, and symmetry-breaking trajectories with accurately converged DFT end states are all recommended. On the evaluation side, the study recommends self-referenced hull construction after relaxation, focused reporting in the near-hull regime, category-wise breakdowns of under- and overprediction, and tracking of symmetry DOF and structure-mismatch metrics such as RMSD and maximum paired distance (Nong et al., 21 Jul 2025).
The implications extend beyond a single scalar stability label. The same work reports that softened potential-energy surfaces can overestimate generative-model metrics by approximately 8 in some cases and can propagate relaxation-induced errors into downstream quantities such as phonons, migration barriers, and surface energies. Within that framework, Ehull is not merely a postprocessed number; it becomes a sensitive probe of whether the underlying energetic model preserves the symmetry-breaking and symmetry-preserving physics needed for reliable crystal-stability prediction (Nong et al., 21 Jul 2025).