---
title: 'GSCDB138: Gold-Standard Chemical Database'
url: https://www.emergentmind.com/topics/gscdb138
type: topic
---

# GSCDB138: Gold-Standard Chemical Database

GSCDB138, the **Gold-Standard Chemical Database 138**, is a benchmark library for assessing and developing density functional approximations (DFAs). It was introduced as a **large, rigorously curated, and chemically diverse set of high-accuracy reference energy differences** intended to remedy limitations in older benchmark collections, especially **GMTKN55** and **MGCDB84**, by combining updated legacy data with newly added property-focused and transition-metal benchmarks [2508.13468]. The resulting collection contains **138 data sets** and **8,383 individual benchmark entries**, spanning energetics, response properties, vibrational observables, and transition-metal chemistry, and is positioned as a validation and development platform for both **non-empirical** and **machine-learned** functionals [2508.13468].

## 1. Origins and motivation

GSCDB138 was created in response to specific deficiencies in benchmark databases that had been used extensively in DFA development. The stated limitations were that prior collections were **dominated by main-group energetics**, had **limited coverage of transition metals**, lacked many **molecular property / response** benchmarks, and included entries that were **duplicated, redundant, spin-contaminated, or based on outdated reference values** [2508.13468].

Within that framing, GSCDB138 is not presented as a simple expansion in size. Its purpose is broader and more methodological: benchmark testing is meant to become both **broader** and **more reliable** through simultaneous enlargement of chemical coverage, systematic reference-value revision, and curation-driven pruning [2508.13468]. This suggests that the database is designed not merely to rank DFAs on aggregate accuracy, but also to expose transferability failures that remain hidden when benchmarking is concentrated on main-group energy differences alone.

A central implication of the design is that DFA assessment is being shifted from a predominantly energetic perspective toward a more heterogeneous property space. In the source description, this is tied directly to the inclusion of response properties and transition-metal chemistry, both of which are identified as weakly represented or absent in the earlier benchmark landscape [2508.13468].

## 2. Scope, composition, and property taxonomy

GSCDB138 contains **138 data sets**, **8,383 individual benchmark entries**, and requires **14,013 single-point energy calculations** [2508.13468]. The collection is organized into seven major property categories:

1. **Barrier heights (BH)**
2. **Electric-field response and related properties (EF)**
3. **Vibrational frequencies (FREQ)**
4. **Isomerization energies (ISO)**
5. **Noncovalent interactions (NC)**
6. **Thermochemistry (TC)**
7. **Transition-metal chemistry (TM)**

These categories define a benchmark space that is broader than conventional reaction-energy-focused suites. The EF category includes **dipole moments**, **static polarizabilities**, and **oriented external electric field (OEEF) energies**, while FREQ isolates **vibrational frequencies** as a distinct testing domain [2508.13468]. That separation is technically consequential because the benchmark conclusions explicitly state that performance on electric-field response properties does not track ground-state energetic performance in a simple way [2508.13468].

The database includes a wide range of specific benchmark families. For barrier heights, listed examples include **BH46, BH876, DBH22, BHPERI11, INV23, ORBH35, WCPT26,** and **MOBH28**. Noncovalent interaction coverage includes **S22, S66, A24, HB262, IHB100, X40, RG10N, 3B-69, 3BHET,** and **He3**. Thermochemistry examples include **AE11, AE18, EA50, IP23, IP30, HAT707, TAE_W4-17, P34,** and **MX34**. Isomerization and conformational subsets include **ACONF, BUT14DIOL, MCONF, PCONF21, C60ISO7, TAUT15,** and **S66Rel7**. Electric-field and response-property examples include **Dip146, Pol130, HR46, T144, OEEF,** and **V30**. Transition-metal chemistry includes **3d4dIPSS, CUAGAU83, DAPD, MME52, MOBH28, ROST61, TMD10, MOR13,** and **TMB11** [2508.13468].

| Category | Illustrative benchmark sets |
|---|---|
| BH | BH46, BH876, DBH22, MOBH28 |
| EF | Dip146, Pol130, OEEF, T144 |
| FREQ | V30 |
| ISO | ACONF, MCONF, TAUT15, S66Rel7 |
| NC | S22, S66, X40, 3B-69 |
| TC | AE11, EA50, HAT707, TAE_W4-17 |
| TM | 3d4dIPSS, CUAGAU83, MOR13, TMB11 |

The breadth of this taxonomy is one of the defining features of GSCDB138. A plausible implication is that aggregated DFA rankings derived from the database are less likely to be dominated by any single chemical regime than rankings based on older, more main-group-centric collections.

## 3. Curation strategy and reference-value revision

A major distinguishing feature of GSCDB138 is that it underwent substantial curation rather than simple aggregation. The curation explicitly removes **identity reactions** with zero reaction energy, **duplicate reactions / duplicate points**, **spin-contaminated species** and associated reactions when reference quality was insufficient, and **low-quality datasets** whose reference values were judged unreliable [2508.13468]. Some overlapping datasets were also restructured into derived forms, including **A19Rel6**, **S66Rel7**, **3B-69** as a true non-additive three-body set, **BH28**, **BH876**, **O24x4**, and **IHB100x2** [2508.13468].

The reference-value modernization is equally central. The database updates many legacy sets to **current best reference values**, explicitly naming **S22**, **S66**, **Shields38**, **WATER27**, **G21IP**, **G21EA**, **C20C24**, **Pentane14**, **NC15**, and **X40**, among others [2508.13468]. The description further emphasizes replacement of older references with newer **high-level CC / focal-point / Wn / F12** references where available [2508.13468].

Several update examples are specified in detail. **W4-17** replaces older **W4-11 / TAE140-type** data; **Shields38** and **WATER27** are updated using recent high-level work; **G21IP** and **G21EA** are updated with **W3-level references**; **C20C24** is updated with more reliable basis-set-extrapolated values; **Pentane14** is updated with **CCSD(T)-F12b/cc-pVTZ-F12** values; **X40** is updated from newer **X40×10** information; and **RG10** is replaced by **RG10N** using much higher-quality **CCSDT(Q)/CCSD(T)** references [2508.13468].

The combined effect of pruning, de-duplication, and reference revision is methodological as much as numerical. The database is constructed so that each retained benchmark serves a distinct role and is supported by a reference treatment considered sufficiently reliable for DFA evaluation [2508.13468]. This suggests an attempt to reduce the extent to which apparent DFA success can arise from redundancy or contamination in the benchmark suite.

## 4. Treatment of spin contamination and data-quality control

The handling of spin contamination is unusually explicit. For potentially problematic species, the workflow proceeds through four stated steps: **internal stability analysis** using **ωB97X-V**; labeling of difficult cases; use of **κ-OOMP2** with $\kappa = 1.45$ to judge whether symmetry breaking was physically essential; and retention only of data with reliable references such as **experiment**, **unrestricted CCSD(T)**, or **beyond-CCSD(T)** methods [2508.13468].

If a reaction involved an essential-spin-breaking species but lacked reliable reference treatment, it was excluded [2508.13468]. This is a stringent filter rather than a post hoc warning label. It directly constrains which chemically interesting but electronically delicate systems are allowed into the final benchmark.

From a benchmarking standpoint, that policy has two consequences. First, it reduces the likelihood that DFA rankings are distorted by uncertain references in open-shell or near-degenerate regimes. Second, it creates a more sharply defined distinction between failures of the tested functional and failures of the benchmark reference itself. A plausible implication is that error statistics derived from GSCDB138 are intended to be more diagnostically interpretable than statistics from collections that mix high- and low-confidence reference data without such pruning.

The same logic appears in the removal of **questionable data**, the renaming of sets after pruning, and the rationalization of overlapping subsets so that each benchmark serves a distinct purpose [2508.13468]. In this sense, GSCDB138 is curated not only for size and diversity, but for identifiability of failure modes.

## 5. Benchmark definitions and normalization framework

The database adopts a per-dataset **mean absolute error (MAE)** as its standard benchmark metric, except for certain special datasets where **mean absolute relative error (MARE)** or specialized metrics from the original literature are used, particularly for electric-field and response-property datasets [2508.13468].

Its central comparative device is the **standard error**, defined as the average of the **2nd, 3rd, and 4th lowest** errors among all tested **hybrid functionals** for a given dataset [2508.13468]. In formula form,

$$
\text{standard error} = \frac{e_2 + e_3 + e_4}{3}
$$

where $e_2$, $e_3$, and $e_4$ are the second-, third-, and fourth-lowest hybrid-functional errors for that dataset [2508.13468].

The normalized score is the **normalized error ratio (NER)**:

$$
\mathrm{NER} = \frac{e_{\text{functional}}}{\text{standard error}}
$$

These NER values are then averaged over each property category and over the full GSCDB138 suite to obtain overall rankings [2508.13468].

This normalization strategy is designed to avoid giving disproportionate weight to a single best-performing hybrid on each dataset. Instead, the baseline is anchored to a small band of high-performing hybrid methods rather than to an absolute minimum [2508.13468]. That choice makes the comparison less sensitive to outliers and, in the wording of the source description, more robust. It also means that the ranking framework is comparative rather than absolute: performance is judged relative to a hybrid-functional reference envelope, not solely by raw error magnitudes.

## 6. Benchmark outcomes and DFA-class behavior

The broad trend across GSCDB138 follows the expected **Jacob’s-ladder** ordering,

$$
\text{LDA} \to \text{GGA} \to \text{meta-GGA} \to \text{hybrid} \to \text{double hybrid},
$$

with accuracy generally improving as one ascends the ladder [2508.13468]. The reported general conclusions are that **double hybrids** are most accurate overall, **hybrids** are next best, **meta-GGAs** often narrow the gap in some categories, **GGAs** are noticeably less accurate, and **LDA** performs worst overall [2508.13468].

The database also identifies category-specific and class-specific leaders. Among **hybrid meta-GGAs**, **ωB97M-V** is the **most balanced and best overall HMGGA**, with the **lowest overall mean NER = 1.08**, leading in **BH, FREQ, ISO,** and **NC**, but performing more weakly for **EF** [2508.13468]. Among **hybrid GGAs**, **ωB97X-V** is the **best-balanced HGGA**, with **overall mean NER = 1.32**, particularly strong for **electric-field properties**, and ranked as the **third-best hybrid overall** [2508.13468]. Among **meta-GGAs**, **B97M-V** is the best overall MGGA, with **overall mean NER = 1.77**, and is especially strong for **noncovalent interactions** [2508.13468]. Among **GGAs**, **revPBE-D4** is the **best / most balanced GGA overall** and is identified as the recommended general-use GGA, while **N12-D3(0)** is better for EF and **OLYP-D4** is better for frequencies [2508.13468].

For vibrational frequencies specifically, **PBE0-D4** is reported as the overall best in the frequency category in the broad benchmark, and **r2SCAN-D4** is highlighted as unusually competitive and, in some comparisons, rivaling hybrids [2508.13468]. The source description treats this as one of the important exceptions to simple Jacob’s-ladder expectations.

Among **double hybrids**, **ωB97M(2)** and **DSD-PBE-D4** are the representative methods studied, and both outperform the best hybrids overall, with **ωB97M(2)** superior to **DSD-PBE-D4** in the reported overall NER comparison [2508.13468]. The stated values are **overall mean NER = 0.79** for **ωB97M(2)** and **1.29** for **DSD-PBE-D4** [2508.13468].

These rankings indicate that the database does not only reproduce conventional class hierarchies; it also resolves category-specific inversions and performance specializations. That is particularly evident in the treatment of frequencies and electric-field response properties.

## 7. Response-property findings, double-hybrid caveats, and intended uses

One of the most important conclusions of GSCDB138 is that **electric-field response errors correlate poorly with ground-state energy performance** [2508.13468]. The source description explicitly notes that **ωB97M-V** and **CF22D**, both very strong on many ground-state energetics, perform poorly for electric-field response properties [2508.13468]. This is interpreted as suggesting **overfitting to energy-difference training data** and inadequate learning of density or response behavior [2508.13468].

That finding matters because the EF category includes **dipoles**, **polarizabilities**, and **field-dependent properties** [2508.13468]. A plausible implication is that a functional optimized predominantly against energetic targets may have limited transferability to density-sensitive observables, even when it ranks highly on conventional benchmark suites. The database therefore argues for explicit inclusion of response properties in future functional development rather than relying only on ground-state energetics [2508.13468].

The treatment of **double hybrids** is similarly nuanced. Although the paper states that double hybrids lower mean errors by about **25%** versus the best hybrids overall, the improvement is accompanied by major practical sensitivities [2508.13468]. Three are emphasized. First, **frozen-core treatment** can strongly affect errors: for **ωB97M(2)**, omitting frozen core can increase the MAE on **AlkAtom19** from **0.18 kcal/mol to 6.24 kcal/mol** [2508.13468]. Second, **basis-set completeness** remains problematic because the MP2-like correlation component converges slowly according to

$$
E_{\mathrm{MP2}} \propto L^{-3},
$$

and the recommendation is to use the **same basis as in parametrization**, typically **def2-QZVPPD** for **ωB97M(2)**, with **special core-correlated bases** when needed, as in **AE11** [2508.13468]. Third, the perturbative MP2 component is sensitive to **multi-reference (MR) character**; the database notes degradation on MR data such as **TAE_W4-17MR** and **ORBH36**, and further notes that **r2SCAN-like semilocal exchange may partially mimic MR effects better than perturbative MP2** [2508.13468].

The intended uses of GSCDB138 follow directly from these findings. It is meant for **DFA validation**, including broad-transferability testing across chemistry and the detection of failures linked to **spin contamination**, **multireference effects**, or **basis sensitivity**, and for **functional design and training**, including training of **non-empirical** and **machine-learned** functionals on a broader target space than energies alone [2508.13468]. In that sense, GSCDB138 is positioned not merely as a larger database, but as a more balanced and more carefully curated benchmark for the next stage of DFA development.

Source: https://www.emergentmind.com/topics/gscdb138