---
title: 'SOD2D: GPU CFD Solver for High-Order Simulations'
url: https://www.emergentmind.com/topics/sod2d
type: topic
---

# SOD2D: GPU CFD Solver for High-Order Simulations

SOD2D denotes, in the recent computational-fluid-dynamics literature, a GPU-enabled spectral-element or spectral finite-element solver used for high-order, scale-resolving simulations and for tightly coupled deep-reinforcement-learning workflows. In the cited arXiv record, it appears in at least three distinct but related roles: as a filtered incompressible Navier–Stokes large-eddy-simulation backend for active flow control over a three-dimensional NACA0012 wing at \(\mathrm{Re}=1{,}000\) and \(\mathrm{AoA}=20^\circ\); as a GPU-accelerated spectral-element code for multi-agent cylinder-wake control at \(\mathrm{Re}=100\); and as a Fortran/OpenACC/MPI spectral finite-element framework for compressible scale-resolving simulations analyzed for multi-GPU performance portability in the REFMAP project [2509.10195; 2508.00645; 2601.14159]. Although the historical name suggests a two-dimensional code, the solver is explicitly used in configurations with spanwise periodicity and resolved three-dimensional wake structures [2509.10195].

## 1. Identity and reported scope

Across the cited papers, SOD2D is characterized consistently as GPU-accelerated and high-order, but the reported physical models depend on the application context. In the active-flow-control study, it is described as a GPU-enabled spectral element method CFD solver developed at the Barcelona Supercomputing Center and used as the environment for deep reinforcement learning [2509.10195]. In SmartFlow, it is presented as “the GPU‑accelerated SOD2D spectral‑element code” used to demonstrate multi-agent deep reinforcement learning for incompressible cylinder wake control [2508.00645]. In the REFMAP performance-portability study, it is described as a “state-of-the-art Spectral Elements simulation framework” and, more specifically, as a Fortran/OpenACC/MPI spectral finite-element CFD solver designed for compressible, turbulence-resolving simulations [2601.14159].

| Source | Characterization of SOD2D | Demonstrated use |
|---|---|---|
| [2509.10195] | GPU-enabled spectral element method CFD solver | DRL-based active flow control over a 3D NACA0012 wing |
| [2508.00645] | GPU-accelerated high-order spectral-element CFD code | SmartFlow-coupled MARL for a 3D cylinder wake |
| [2601.14159] | Fortran/OpenACC/MPI spectral finite-element solver | Compressible scale-resolving simulations and portability analysis |

This distribution of descriptions indicates that SOD2D is best understood as a solver framework whose reported formulation varies with the study. One recurring feature is the tension between the code name and the dimensionality of the deployed applications: both the wing and cylinder control studies use explicit spanwise extent and periodic spanwise boundary conditions, rather than purely two-dimensional dynamics [2509.10195; 2508.00645].

## 2. Governing equations and numerical formulations

In the DRL-active-flow-control literature, SOD2D solves the filtered incompressible Navier–Stokes equations in LES form,
\[
\frac{\partial \mathbf{u}}{\partial t} + (\mathbf{u}\cdot\nabla)\mathbf{u}
= -\frac{1}{\rho}\nabla p + \nu \nabla^2 \mathbf{u} + \mathbf{f},
\qquad
\nabla\cdot\mathbf{u} = 0,
\]
with \(\mathbf{u}(\mathbf{x},t)\) the velocity, \(p\) the pressure, \(\rho\) the density, \(\nu=\mu/\rho\) the kinematic viscosity, and \(\mathbf{f}\) a body force including actuation effects [2509.10195]. The same incompressible equations are stated for the SmartFlow cylinder study, which uses a three-dimensional domain with a periodic spanwise direction at \(\mathrm{Re}=100\) [2508.00645].

The REFMAP paper reports a different formulation. There, SOD2D solves the compressible Navier–Stokes equations in conservative form for density \(\rho\), velocity \(\mathbf{u}\), pressure \(p\), temperature \(T\), and total energy \(E\), with the ideal gas law providing closure [2601.14159]. In that description, the solver uses unstructured spectral elements in 2D, Gauss–Lobatto–Legendre nodes with quadrature points coinciding with element nodes, and a diagonal mass structure enabling explicit time stepping. Temporal advancement is performed with fourth-order Runge–Kutta, convective and diffusive terms are treated through operator splitting, and stabilization is supplied by an entropy-viscosity model [2601.14159].

The numerical details reported in SmartFlow are deliberately thinner. That paper does not specify the spectral-element basis, polynomial order, time integration scheme, pressure treatment, dealiasing, or projection method; SOD2D is described there only at a high level as “high‑order spectral element” and “GPU‑accelerated” [2508.00645]. By contrast, the REFMAP study exposes the kernel structure and the consequences of explicit high-order spectral-element discretization for memory traffic, arithmetic intensity, and GPU occupancy [2601.14159].

## 3. Coupling architecture for reinforcement learning

A central role of SOD2D in the recent literature is as the CFD environment in DRL or MARL loops. In the wing-flow-control study, the Fortran-based solver and the Python-based TF-Agents PPO agent communicate via a Redis in-memory database orchestrated by SmartSim, which is explicitly presented as a way of solving the “two-language” problem [2509.10195]. The data exchange per action step follows a fixed sequence: SOD2D advances the flow and computes observations at sensor locations; these observations are written to Redis; the PPO policy reads the state and outputs jet-velocity actions; actions are written back to Redis; and SOD2D applies the actuation, continues the simulation for the action duration, and computes the reward from aerodynamic forces [2509.10195].

In SmartFlow, the same general pattern is generalized into a solver-agnostic infrastructure. SmartSim starts an in-memory Redis/KeyDB Orchestrator, while SmartRedis-MPI provides MPI-aware APIs—`init_smartredis_mpi`, `put_state`, `get_action`, `put_reward`, and `finalize_smartredis_mpi`—for low-latency in-memory exchange between SOD2D and Python-side DRL algorithms implemented with Stable-Baselines3 [2508.00645]. At each control interval \(\Delta t_{RL}\), SOD2D instances push states, the shared PPO policy computes actions, the solver retrieves those actions, applies boundary actuation for one control interval, and then computes local or global rewards [2508.00645].

The control cadence is highly structured in both studies. In the wing case, each episode spans six baseline vortex-shedding cycles and contains 120 actions, so \(T_{\text{act}}=T_{\text{eps}}/120\); ten CFD simulations are run in parallel and the domain is partitioned spanwise into three pseudo-environments per simulation, giving 30 trajectories per decision step [2509.10195]. In the cylinder case, each episode lasts \(T_{\text{episode}} \approx 34.88\,D/U_\infty\) with 120 actions per episode; four concurrent SOD2D simulations, each split into 10 pseudo-environments, produce 40 trajectories per policy update [2508.00645]. In both settings, parallel multi-environment rollout is a primary mechanism for accelerating policy learning.

## 4. Control-oriented deployments

The most detailed SOD2D deployment reported to date is the active control of a separated three-dimensional NACA0012 wing section at \(\mathrm{Re}=1{,}000\) and \(\mathrm{AoA}=20^\circ\) [2509.10195]. The computational domain extends \(L_x=32c\), \(L_y=30c\), and \(L_z=3c\), with inlet \(U_\infty\), outlet zero-gradient velocity plus constant pressure, no-slip walls on the wing, and periodic spanwise boundaries. Two sets of three spanwise-distributed actuators are used, each set containing a front jet at \(x/c=0.01\) and a rear jet at \(x/c=0.4\), with instantaneous mass conservation enforced through
\[
U_{\text{jet, front}} = -\,U_{\text{jet, rear}}.
\]
The agent controls only the front-jet velocity, bounded by
\[
\frac{U_{\text{jet}}}{U_\infty} \in [-1.125,\,1.125].
\]
Each pseudo-environment senses pressure at 90 witness points in its mid-span slice, augmented by the two neighboring slices for a total of 270 inputs [2509.10195].

The wing-study reward mixes local aerodynamic improvement and spanwise cooperation. The local reward is
\[
r_i = (C_{d,b} - C_d) - \alpha \,\big|C_l - C_{l,\text{avg}}\big|,
\]
with \(\alpha=0.3\), and the aggregated reward is
\[
R_i = \beta\, r_i + \frac{1-\beta}{n_{\text{jets}}}\sum_{j=1}^{n_{\text{jets}}} r_j,
\]
with \(\beta=0.8\) and \(n_{\text{jets}}=3\) [2509.10195]. Training plateaued by episode 67. Deterministic evaluation reported \(\Delta C_D=-21\%\), \(\Delta C_{L,\mathrm{rms}}=-124\%\), and an increase in the vortex-shedding Strouhal number from \(\mathrm{St}=0.56\) to \(\mathrm{St}=0.64\) [2509.10195]. The learned actuation locks onto \(\mathrm{St}=0.64\), produces a negative-mean front jet and positive-mean rear jet, extends the shear layer farther downstream, and shrinks the time-averaged negative-\(u_x\) region on the suction side [2509.10195].

The SmartFlow cylinder case uses SOD2D in an incompressible bluff-body setting at \(\mathrm{Re}=100\) [2508.00645]. The domain is \(L_x=30D\), \(L_y=15D\), \(L_z=4D\), with uniform inflow, zero-gradient outlet with constant pressure, slip top and bottom, and periodic spanwise boundaries. Ten pairs of wall-mounted jets are uniformly distributed along the spanwise direction; each pair comprises a top jet centered at \(\theta_0=90^\circ\) and a bottom jet at \(\theta_0=270^\circ\), with spanwise width \(0.4D\), angular width \(\omega=10^\circ\), and zero net mass flux \(Q_{\rm top}=-Q_{\rm bot}\). The action variable is the mass flow rate per unit span \(Q=\dot m/L_z\), with \(|Q|\le 0.176\), applied through a Dirichlet wall-jet boundary condition having a cosine profile in \(\theta\) [2508.00645]. Each pseudo-environment uses 85 witness points around the cylinder wall at mid-span, augmented by the two neighboring pseudo-environments to form a 255-value state vector. After 100 episodes on a single node with four NVIDIA A100 GPUs, training took approximately 40 hours; deterministic evaluation showed a 6.0% reduction in mean drag and a 46.6% reduction in lift-coefficient fluctuation standard deviation [2508.00645].

## 5. GPU performance portability and computational behavior

The REFMAP performance study makes SOD2D’s hardware behavior unusually explicit [2601.14159]. The dominant hotspot is `full_convec`, which accounts for up to 50% of runtime on NVIDIA and up to 70% on AMD; `full_diffusion` accounts for 22% on NVIDIA and 10% on AMD. Within each Runge–Kutta stage, execution is organized into preloading of global quantities into element-local buffers, kernel execution in natural and physical coordinates, and store-back through atomic updates to global arrays. The paper characterizes the convective kernel as predominantly memory-bound, with irregular gathers and atomic store-backs limiting operational intensity [2601.14159].

Single-GPU characterization spans Tesla V100, A100, and AMD MI250X platforms. On approximately 8M nodes in FP32 without code optimizations, a single MI250X GCD was reported as 15.83\(\times\) slower than V100 and 23.38\(\times\) slower than A100; A100, in turn, delivered 1.47\(\times\) speedup over V100 in FP32 and 3.05\(\times\) in FP64 [2601.14159]. The same study shows that kernel splitting, preloading removal, and OpenACC cache prefetching have markedly vendor- and precision-dependent effects. Splitting `full_convec` yielded an immediate 1.5\(\times\) speedup on MI250X by reducing VGPR and LDS pressure; prefetching small local arrays produced gains such as 1.24\(\times\) on V100 FP64 split kernels and 42% on MI250X FP32 unified kernels [2601.14159].

The portability conclusion is that the same memory-access optimization can be neutral or harmful on another architecture. The measured spread is 0.69\(\times\)–3.91\(\times\) in acceleration speedup across optimizations and platforms [2601.14159]. Weak-scaling experiments on LUMI, at approximately 1M nodes per GPU, further show that the “best” configuration at one scale is not reliably the best at another. For channel flow, throughput variation from worst to optimal configuration reached 23.8% in FP32 and 14.6% in FP64, and projecting the 32-GPU optimum to 64 GPUs incurred penalties of 23.8% and 8.3%, respectively [2601.14159]. A plausible implication is that SOD2D’s algorithmic structure is mature enough to expose full-stack tuning issues—application, software, and hardware—rather than being dominated by a single implementation bottleneck.

## 6. Reproducibility, limitations, and nomenclature

The cited literature gives partial but uneven reproducibility support. SmartFlow provides public repositories for the framework itself, SmartRedis-MPI, the SmartRedis client, and a SmartFlow–SOD2D integration example, but it does not provide a separate public repository or license for SOD2D, nor does it specify the exact solver version, standalone build scripts, or complete numerical configuration files required to reproduce the SOD2D cylinder case independently [2508.00645]. The wing-flow-control paper reports domain sizes, boundary conditions, actuation layout, reward parameters \((\alpha=0.3,\beta=0.8)\), and episode scheduling, but does not detail mesh resolution, time step, turbulence-model specifics beyond “filtered incompressible Navier–Stokes (LES),” or CFL constraints [2509.10195]. In both cases, exponential smoothing of the actions is identified as a numerical-robustness device [2509.10195].

A persistent issue is nomenclature. The solver is named SOD2D, yet the literature explicitly emphasizes that it is used for a 3D wing section with spanwise periodicity and for a 3D cylinder flow with periodic \(z\)-boundaries [2509.10195; 2508.00645]. The SmartFlow paper states that future work could clarify whether SOD2D is extruded 2D with periodic \(z\), because the handling of the third dimension is not described there [2508.00645]. There is also a separate, homonymous usage in the semi-Lagrangian lattice Boltzmann literature, where “SOD2D” refers not to the solver but to two-dimensional Sod- or Riemann-type benchmark problems for compressible discontinuity interaction [1910.13918]. This suggests that “SOD2D” is not a universally unambiguous identifier across arXiv, and the surrounding research context is essential for correct interpretation.

Source: https://www.emergentmind.com/topics/sod2d