Papers
Topics
Authors
Recent
Search
2000 character limit reached

GSCDB138: Gold-Standard Chemical Database

Updated 9 July 2026
  • GSCDB138 is a comprehensive, curated benchmark library containing 138 data sets and 8,383 entries designed to assess and develop density functional approximations.
  • It spans a broad spectrum of chemical properties including energetics, vibrational frequencies, response properties, and transition-metal chemistry, addressing limitations of previous databases.
  • Robust curation methods remove duplicates, spin-contaminated entries, and outdated data while updating reference values to ensure reliable DFA performance evaluation.

GSCDB138, the Gold-Standard Chemical Database 138, is a benchmark library for assessing and developing density functional approximations (DFAs). It was introduced as a large, rigorously curated, and chemically diverse set of high-accuracy reference energy differences intended to remedy limitations in older benchmark collections, especially GMTKN55 and MGCDB84, by combining updated legacy data with newly added property-focused and transition-metal benchmarks (Liang et al., 19 Aug 2025). The resulting collection contains 138 data sets and 8,383 individual benchmark entries, spanning energetics, response properties, vibrational observables, and transition-metal chemistry, and is positioned as a validation and development platform for both non-empirical and machine-learned functionals (Liang et al., 19 Aug 2025).

1. Origins and motivation

GSCDB138 was created in response to specific deficiencies in benchmark databases that had been used extensively in DFA development. The stated limitations were that prior collections were dominated by main-group energetics, had limited coverage of transition metals, lacked many molecular property / response benchmarks, and included entries that were duplicated, redundant, spin-contaminated, or based on outdated reference values (Liang et al., 19 Aug 2025).

Within that framing, GSCDB138 is not presented as a simple expansion in size. Its purpose is broader and more methodological: benchmark testing is meant to become both broader and more reliable through simultaneous enlargement of chemical coverage, systematic reference-value revision, and curation-driven pruning (Liang et al., 19 Aug 2025). This suggests that the database is designed not merely to rank DFAs on aggregate accuracy, but also to expose transferability failures that remain hidden when benchmarking is concentrated on main-group energy differences alone.

A central implication of the design is that DFA assessment is being shifted from a predominantly energetic perspective toward a more heterogeneous property space. In the source description, this is tied directly to the inclusion of response properties and transition-metal chemistry, both of which are identified as weakly represented or absent in the earlier benchmark landscape (Liang et al., 19 Aug 2025).

2. Scope, composition, and property taxonomy

GSCDB138 contains 138 data sets, 8,383 individual benchmark entries, and requires 14,013 single-point energy calculations (Liang et al., 19 Aug 2025). The collection is organized into seven major property categories:

  1. Barrier heights (BH)
  2. Electric-field response and related properties (EF)
  3. Vibrational frequencies (FREQ)
  4. Isomerization energies (ISO)
  5. Noncovalent interactions (NC)
  6. Thermochemistry (TC)
  7. Transition-metal chemistry (TM)

These categories define a benchmark space that is broader than conventional reaction-energy-focused suites. The EF category includes dipole moments, static polarizabilities, and oriented external electric field (OEEF) energies, while FREQ isolates vibrational frequencies as a distinct testing domain (Liang et al., 19 Aug 2025). That separation is technically consequential because the benchmark conclusions explicitly state that performance on electric-field response properties does not track ground-state energetic performance in a simple way (Liang et al., 19 Aug 2025).

The database includes a wide range of specific benchmark families. For barrier heights, listed examples include BH46, BH876, DBH22, BHPERI11, INV23, ORBH35, WCPT26, and MOBH28. Noncovalent interaction coverage includes S22, S66, A24, HB262, IHB100, X40, RG10N, 3B-69, 3BHET, and He3. Thermochemistry examples include AE11, AE18, EA50, IP23, IP30, HAT707, TAE_W4-17, P34, and MX34. Isomerization and conformational subsets include ACONF, BUT14DIOL, MCONF, PCONF21, C60ISO7, TAUT15, and S66Rel7. Electric-field and response-property examples include Dip146, Pol130, HR46, T144, OEEF, and V30. Transition-metal chemistry includes 3d4dIPSS, CUAGAU83, DAPD, MME52, MOBH28, ROST61, TMD10, MOR13, and TMB11 (Liang et al., 19 Aug 2025).

Category Illustrative benchmark sets
BH BH46, BH876, DBH22, MOBH28
EF Dip146, Pol130, OEEF, T144
FREQ V30
ISO ACONF, MCONF, TAUT15, S66Rel7
NC S22, S66, X40, 3B-69
TC AE11, EA50, HAT707, TAE_W4-17
TM 3d4dIPSS, CUAGAU83, MOR13, TMB11

The breadth of this taxonomy is one of the defining features of GSCDB138. A plausible implication is that aggregated DFA rankings derived from the database are less likely to be dominated by any single chemical regime than rankings based on older, more main-group-centric collections.

3. Curation strategy and reference-value revision

A major distinguishing feature of GSCDB138 is that it underwent substantial curation rather than simple aggregation. The curation explicitly removes identity reactions with zero reaction energy, duplicate reactions / duplicate points, spin-contaminated species and associated reactions when reference quality was insufficient, and low-quality datasets whose reference values were judged unreliable (Liang et al., 19 Aug 2025). Some overlapping datasets were also restructured into derived forms, including A19Rel6, S66Rel7, 3B-69 as a true non-additive three-body set, BH28, BH876, O24x4, and IHB100x2 (Liang et al., 19 Aug 2025).

The reference-value modernization is equally central. The database updates many legacy sets to current best reference values, explicitly naming S22, S66, Shields38, WATER27, G21IP, G21EA, C20C24, Pentane14, NC15, and X40, among others (Liang et al., 19 Aug 2025). The description further emphasizes replacement of older references with newer high-level CC / focal-point / Wn / F12 references where available (Liang et al., 19 Aug 2025).

Several update examples are specified in detail. W4-17 replaces older W4-11 / TAE140-type data; Shields38 and WATER27 are updated using recent high-level work; G21IP and G21EA are updated with W3-level references; C20C24 is updated with more reliable basis-set-extrapolated values; Pentane14 is updated with CCSD(T)-F12b/cc-pVTZ-F12 values; X40 is updated from newer X40×10 information; and RG10 is replaced by RG10N using much higher-quality CCSDT(Q)/CCSD(T) references (Liang et al., 19 Aug 2025).

The combined effect of pruning, de-duplication, and reference revision is methodological as much as numerical. The database is constructed so that each retained benchmark serves a distinct role and is supported by a reference treatment considered sufficiently reliable for DFA evaluation (Liang et al., 19 Aug 2025). This suggests an attempt to reduce the extent to which apparent DFA success can arise from redundancy or contamination in the benchmark suite.

4. Treatment of spin contamination and data-quality control

The handling of spin contamination is unusually explicit. For potentially problematic species, the workflow proceeds through four stated steps: internal stability analysis using ωB97X-V; labeling of difficult cases; use of κ-OOMP2 with κ=1.45\kappa = 1.45 to judge whether symmetry breaking was physically essential; and retention only of data with reliable references such as experiment, unrestricted CCSD(T), or beyond-CCSD(T) methods (Liang et al., 19 Aug 2025).

If a reaction involved an essential-spin-breaking species but lacked reliable reference treatment, it was excluded (Liang et al., 19 Aug 2025). This is a stringent filter rather than a post hoc warning label. It directly constrains which chemically interesting but electronically delicate systems are allowed into the final benchmark.

From a benchmarking standpoint, that policy has two consequences. First, it reduces the likelihood that DFA rankings are distorted by uncertain references in open-shell or near-degenerate regimes. Second, it creates a more sharply defined distinction between failures of the tested functional and failures of the benchmark reference itself. A plausible implication is that error statistics derived from GSCDB138 are intended to be more diagnostically interpretable than statistics from collections that mix high- and low-confidence reference data without such pruning.

The same logic appears in the removal of questionable data, the renaming of sets after pruning, and the rationalization of overlapping subsets so that each benchmark serves a distinct purpose (Liang et al., 19 Aug 2025). In this sense, GSCDB138 is curated not only for size and diversity, but for identifiability of failure modes.

5. Benchmark definitions and normalization framework

The database adopts a per-dataset mean absolute error (MAE) as its standard benchmark metric, except for certain special datasets where mean absolute relative error (MARE) or specialized metrics from the original literature are used, particularly for electric-field and response-property datasets (Liang et al., 19 Aug 2025).

Its central comparative device is the standard error, defined as the average of the 2nd, 3rd, and 4th lowest errors among all tested hybrid functionals for a given dataset (Liang et al., 19 Aug 2025). In formula form,

standard error=e2+e3+e43\text{standard error} = \frac{e_2 + e_3 + e_4}{3}

where e2e_2, e3e_3, and e4e_4 are the second-, third-, and fourth-lowest hybrid-functional errors for that dataset (Liang et al., 19 Aug 2025).

The normalized score is the normalized error ratio (NER):

NER=efunctionalstandard error\mathrm{NER} = \frac{e_{\text{functional}}}{\text{standard error}}

These NER values are then averaged over each property category and over the full GSCDB138 suite to obtain overall rankings (Liang et al., 19 Aug 2025).

This normalization strategy is designed to avoid giving disproportionate weight to a single best-performing hybrid on each dataset. Instead, the baseline is anchored to a small band of high-performing hybrid methods rather than to an absolute minimum (Liang et al., 19 Aug 2025). That choice makes the comparison less sensitive to outliers and, in the wording of the source description, more robust. It also means that the ranking framework is comparative rather than absolute: performance is judged relative to a hybrid-functional reference envelope, not solely by raw error magnitudes.

6. Benchmark outcomes and DFA-class behavior

The broad trend across GSCDB138 follows the expected Jacob’s-ladder ordering,

LDAGGAmeta-GGAhybriddouble hybrid,\text{LDA} \to \text{GGA} \to \text{meta-GGA} \to \text{hybrid} \to \text{double hybrid},

with accuracy generally improving as one ascends the ladder (Liang et al., 19 Aug 2025). The reported general conclusions are that double hybrids are most accurate overall, hybrids are next best, meta-GGAs often narrow the gap in some categories, GGAs are noticeably less accurate, and LDA performs worst overall (Liang et al., 19 Aug 2025).

The database also identifies category-specific and class-specific leaders. Among hybrid meta-GGAs, ωB97M-V is the most balanced and best overall HMGGA, with the lowest overall mean NER = 1.08, leading in BH, FREQ, ISO, and NC, but performing more weakly for EF (Liang et al., 19 Aug 2025). Among hybrid GGAs, ωB97X-V is the best-balanced HGGA, with overall mean NER = 1.32, particularly strong for electric-field properties, and ranked as the third-best hybrid overall (Liang et al., 19 Aug 2025). Among meta-GGAs, B97M-V is the best overall MGGA, with overall mean NER = 1.77, and is especially strong for noncovalent interactions (Liang et al., 19 Aug 2025). Among GGAs, revPBE-D4 is the best / most balanced GGA overall and is identified as the recommended general-use GGA, while N12-D3(0) is better for EF and OLYP-D4 is better for frequencies (Liang et al., 19 Aug 2025).

For vibrational frequencies specifically, PBE0-D4 is reported as the overall best in the frequency category in the broad benchmark, and r2SCAN-D4 is highlighted as unusually competitive and, in some comparisons, rivaling hybrids (Liang et al., 19 Aug 2025). The source description treats this as one of the important exceptions to simple Jacob’s-ladder expectations.

Among double hybrids, ωB97M(2) and DSD-PBE-D4 are the representative methods studied, and both outperform the best hybrids overall, with ωB97M(2) superior to DSD-PBE-D4 in the reported overall NER comparison (Liang et al., 19 Aug 2025). The stated values are overall mean NER = 0.79 for ωB97M(2) and 1.29 for DSD-PBE-D4 (Liang et al., 19 Aug 2025).

These rankings indicate that the database does not only reproduce conventional class hierarchies; it also resolves category-specific inversions and performance specializations. That is particularly evident in the treatment of frequencies and electric-field response properties.

7. Response-property findings, double-hybrid caveats, and intended uses

One of the most important conclusions of GSCDB138 is that electric-field response errors correlate poorly with ground-state energy performance (Liang et al., 19 Aug 2025). The source description explicitly notes that ωB97M-V and CF22D, both very strong on many ground-state energetics, perform poorly for electric-field response properties (Liang et al., 19 Aug 2025). This is interpreted as suggesting overfitting to energy-difference training data and inadequate learning of density or response behavior (Liang et al., 19 Aug 2025).

That finding matters because the EF category includes dipoles, polarizabilities, and field-dependent properties (Liang et al., 19 Aug 2025). A plausible implication is that a functional optimized predominantly against energetic targets may have limited transferability to density-sensitive observables, even when it ranks highly on conventional benchmark suites. The database therefore argues for explicit inclusion of response properties in future functional development rather than relying only on ground-state energetics (Liang et al., 19 Aug 2025).

The treatment of double hybrids is similarly nuanced. Although the paper states that double hybrids lower mean errors by about 25% versus the best hybrids overall, the improvement is accompanied by major practical sensitivities (Liang et al., 19 Aug 2025). Three are emphasized. First, frozen-core treatment can strongly affect errors: for ωB97M(2), omitting frozen core can increase the MAE on AlkAtom19 from 0.18 kcal/mol to 6.24 kcal/mol (Liang et al., 19 Aug 2025). Second, basis-set completeness remains problematic because the MP2-like correlation component converges slowly according to

EMP2L3,E_{\mathrm{MP2}} \propto L^{-3},

and the recommendation is to use the same basis as in parametrization, typically def2-QZVPPD for ωB97M(2), with special core-correlated bases when needed, as in AE11 (Liang et al., 19 Aug 2025). Third, the perturbative MP2 component is sensitive to multi-reference (MR) character; the database notes degradation on MR data such as TAE_W4-17MR and ORBH36, and further notes that r2SCAN-like semilocal exchange may partially mimic MR effects better than perturbative MP2 (Liang et al., 19 Aug 2025).

The intended uses of GSCDB138 follow directly from these findings. It is meant for DFA validation, including broad-transferability testing across chemistry and the detection of failures linked to spin contamination, multireference effects, or basis sensitivity, and for functional design and training, including training of non-empirical and machine-learned functionals on a broader target space than energies alone (Liang et al., 19 Aug 2025). In that sense, GSCDB138 is positioned not merely as a larger database, but as a more balanced and more carefully curated benchmark for the next stage of DFA development.

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to GSCDB138.