---
title: COACH420 Benchmark Dataset
url: https://www.emergentmind.com/topics/coach420-benchmark-dataset
type: topic
---

# COACH420 Benchmark Dataset

The COACH420 benchmark dataset provides a high-fidelity, temporally resolved 2D digital record of multiphase flow, capturing the displacement of water by CO₂ in synthetic porous media. Released alongside Abdellatif et al.'s publication “Benchmark Dataset for Pore-Scale CO₂–Water Interaction” [2503.17592], it was designed to support the rigorous development and benchmarking of physics-based and machine-learning models of subsurface flow relevant to carbon capture and storage (CCS), enhanced oil recovery, and related geoscience applications.

## 1. Dataset Specifications

COACH420 comprises 624 distinct realizations of pore-scale geometry, each represented as a 512 × 512 binary image with a uniform spatial resolution of $\Delta x = 35\,\mu\mathrm{m}$. Each sample evolves over $N_t = 100$ time steps at a temporal increment $\Delta t = 0.01\,\mathrm{s}$, spanning a total duration of 1 s under a constant CO₂ injection rate. The dataset captures the following fields per sample in HDF5 format (array shape: 100×512×512 unless otherwise specified):

- Ux, Uy: horizontal and vertical velocity components (m/s)
- p: pressure field (Pa)
- pc: capillary pressure field (Pa)
- α_water: water saturation (unitless, [0,1])
- img: binary domain mask (512×512), where pores = 1 and solid grains = 0

CO₂ saturation can be analytically reconstructed point-wise as:
\[
\alpha_{\mathrm{CO}_2}(x,y,t) = \left[1 - \alpha_{\mathrm{water}}(x,y,t)\right]\,\text{img}(x,y)
\]
All data are organized to maximize reproducibility, analysis at multiple granularities, and direct comparability across samples.

## 2. Numerical Simulation Details

All simulations were conducted using GeoChemFoam’s algebraic Volume-of-Fluid (VOF) solver on 2D domains (depth: one voxel, 35 $\mu$m). The initial condition is full water saturation throughout the pore space. CO₂ is injected continuously at the left boundary at a volumetric rate of $Q = 1\times10^{-8}\,\mathrm{m}^3/\mathrm{s}$ (capillary number $\mathrm{Ca} \approx 5\times10^{-6}$). The right boundary is an outlet, and internal grains are assigned no-slip/no-flux conditions. Fluid parameters are:

- CO₂: $\rho_{\mathrm{CO}_2} = 3.84\times10^2\,\mathrm{kg/m}^3$, $\nu_{\mathrm{CO}_2} = 7.37\times10^{-8}\,\mathrm{m}^2/\mathrm{s}$
- Water: $\rho_{\mathrm{H_2O}} = 1.00\times10^3\,\mathrm{kg/m}^3$, $\nu_{\mathrm{H_2O}} = 1.00\times10^{-6}\,\mathrm{m}^2/\mathrm{s}$
- Interfacial tension: $\sigma = 0.03\,\mathrm{N/m}$
- Contact angle: $\theta = 45^\circ$

Governing equations include the single-field incompressible Navier–Stokes system:
\[
\nabla\cdot\mathbf{u} = 0,\quad
\rho\left(\partial_t\mathbf{u} + \mathbf{u}\cdot\nabla\mathbf{u}\right)
= -\nabla p + \nabla\cdot\left[\mu(\nabla\mathbf{u}+\nabla\mathbf{u}^T)\right]
+ \sigma\,\kappa\,\nabla\alpha
\]
supplemented by the VOF phase-transport equation for $\alpha$. Outputs are written every $\Delta t=0.01$ s, with a solver convergence tolerance of $10^{-8}$.

## 3. Representation and Design of Heterogeneity

To ensure coverage of realistic microstructure variability, COACH420 implements five discrete heterogeneity levels achieved by randomizing distributions of grain radii and inter-grain spacing. The procedure for generating unique geometries includes:

1. Begin from a 1024×1024 synthetic domain at each heterogeneity level
2. Crop into four non-overlapping 512×512 tiles
3. Vertically flip each tile for data augmentation

This process yields $5$ heterogeneity levels × $4$ tiles × $2$ flips $= 40$ base geometries. Additional random seeds and parametric variation result in a final total of 624 unique samples. The combination of intra-sample heterogeneity (within a tile) and inter-sample heterogeneity (across levels) is designed to prevent machine-learning surrogates from overfitting to spatially repetitive or simplified topologies.

## 4. Data Organization and Access

The COACH420 dataset is hosted on Dryad (DOI: 10.5061/dryad.jm63xsjn5) and is partitioned into 10 primary folders—one original and one flipped copy for each of the five heterogeneity levels. Each folder contains:

- *.hdf5 files: Ux, Uy, p, pc, α_water, img
- poroPerm.csv: initial porosity, permeability, characteristic pore length $L$, Reynolds number $\mathrm{Re}$, Darcy velocity $U_D$
- relperm.csv: time-series data for water saturation $S_w$, water relative permeability $k_{rw}$, CO₂ relative permeability $k_{ro}$, and phase capillary numbers

Filenames employ the convention: geometryLevel_tileIndex[_flip].hdf5, poroPerm.csv, relperm.csv

A Python snippet for accessing HDF5 data is provided:

```python
import h5py
with h5py.File('geometry1_1.hdf5','r') as f:
    Ux = f['Ux'][:]           # shape (100,512,512)
    img = f['img'][:]         # shape (512,512)
    alpha_co2 = (1 - f['alpha_water'][:]) * img
```

Licensing is through the standard Dryad CC0/Public Domain Dedication. Citation format and links are explicitly stated.

## 5. Evaluation and Benchmarking Use Cases

COACH420 provides a high spatial (35 μm) and temporal (0.01 s) resolution benchmark for a range of data-driven and physics-based approaches, including:

- Surrogate modeling: Training convolutional encoder–decoder networks to predict flow and transport fields (saturation, pressure, velocity) from static pore images
- Operator learning: Developing and benchmarking physics-informed neural operators or Fourier neural operators capable of generalizing across microstructure classes for multiphase flows
- Generative modeling: Evaluating GANs and diffusion models for synthetic pore geometry augmentation while preserving physical realism
- Transfer learning: Quantifying model robustness and adaptability across distinct heterogeneity levels
- Uncertainty quantification and active learning: Facilitating research into decision-making pipelines for CCS and enhanced oil recovery where data-driven and physical models are integrated

A plausible implication is that such high-fidelity, systematically organized pore-scale datasets will advance reproducibility and cross-comparability in computational geoscience, particularly for methods that must generalize across microstructural and parametric variability.

## 6. Context and Applications in Subsurface Flow Research

COACH420 addresses the need for standardized, high-resolution datasets to support the rapid prototyping and rigorous testing of machine-learning surrogates and multiscale numerical methods in the domain of subsurface multiphase flow. Its construction is aligned with best practices for benchmarking workflows, ensuring coverage of the heterogeneity encountered in practical monitoring and design of CCS operations. The dataset is intended for deployment in the development of predictive, transferable, and uncertainty-aware models central to geological CO₂ storage and hydrocarbon recovery, offering both breadth across microstructural variation and depth in the physical realism of simulated flow.

## 7. Citation and Data Availability

The COACH420 benchmark dataset is openly available at https://doi.org/10.5061/dryad.jm63xsjn5 under CC0/Public Domain Dedication, and should be cited as:

Abdellatif, A.; Menke, H. P.; Maes, J.; Elsheikh, A. H.; Doster, F. “Benchmark Dataset for Pore-Scale CO₂–Water Interaction.” Dryad Digital Repository, 2024. https://doi.org/10.5061/dryad.jm63xsjn5

All technical details, data formats, and access protocols are specified in the paper “Benchmark Dataset for Pore-Scale CO₂–Water Interaction” [2503.17592].

Source: https://www.emergentmind.com/topics/coach420-benchmark-dataset