SOD2D: GPU CFD Solver for High-Order Simulations
- SOD2D is a GPU-enabled spectral-element CFD solver that supports high-order, scale-resolving simulations across incompressible and compressible flows.
- It integrates deep reinforcement learning frameworks and in-memory data exchange to enable active flow control in complex 3D configurations.
- Performance studies highlight porting challenges on various GPU architectures, emphasizing the need for tailored kernel optimizations for scalability.
SOD2D denotes, in the recent computational-fluid-dynamics literature, a GPU-enabled spectral-element or spectral finite-element solver used for high-order, scale-resolving simulations and for tightly coupled deep-reinforcement-learning workflows. In the cited arXiv record, it appears in at least three distinct but related roles: as a filtered incompressible Navier–Stokes large-eddy-simulation backend for active flow control over a three-dimensional NACA0012 wing at and ; as a GPU-accelerated spectral-element code for multi-agent cylinder-wake control at ; and as a Fortran/OpenACC/MPI spectral finite-element framework for compressible scale-resolving simulations analyzed for multi-GPU performance portability in the REFMAP project (Montalà et al., 12 Sep 2025, Xiao et al., 1 Aug 2025, Eleftherakis et al., 20 Jan 2026). Although the historical name suggests a two-dimensional code, the solver is explicitly used in configurations with spanwise periodicity and resolved three-dimensional wake structures (Montalà et al., 12 Sep 2025).
1. Identity and reported scope
Across the cited papers, SOD2D is characterized consistently as GPU-accelerated and high-order, but the reported physical models depend on the application context. In the active-flow-control study, it is described as a GPU-enabled spectral element method CFD solver developed at the Barcelona Supercomputing Center and used as the environment for deep reinforcement learning (Montalà et al., 12 Sep 2025). In SmartFlow, it is presented as “the GPU‑accelerated SOD2D spectral‑element code” used to demonstrate multi-agent deep reinforcement learning for incompressible cylinder wake control (Xiao et al., 1 Aug 2025). In the REFMAP performance-portability study, it is described as a “state-of-the-art Spectral Elements simulation framework” and, more specifically, as a Fortran/OpenACC/MPI spectral finite-element CFD solver designed for compressible, turbulence-resolving simulations (Eleftherakis et al., 20 Jan 2026).
| Source | Characterization of SOD2D | Demonstrated use |
|---|---|---|
| (Montalà et al., 12 Sep 2025) | GPU-enabled spectral element method CFD solver | DRL-based active flow control over a 3D NACA0012 wing |
| (Xiao et al., 1 Aug 2025) | GPU-accelerated high-order spectral-element CFD code | SmartFlow-coupled MARL for a 3D cylinder wake |
| (Eleftherakis et al., 20 Jan 2026) | Fortran/OpenACC/MPI spectral finite-element solver | Compressible scale-resolving simulations and portability analysis |
This distribution of descriptions indicates that SOD2D is best understood as a solver framework whose reported formulation varies with the study. One recurring feature is the tension between the code name and the dimensionality of the deployed applications: both the wing and cylinder control studies use explicit spanwise extent and periodic spanwise boundary conditions, rather than purely two-dimensional dynamics (Montalà et al., 12 Sep 2025, Xiao et al., 1 Aug 2025).
2. Governing equations and numerical formulations
In the DRL-active-flow-control literature, SOD2D solves the filtered incompressible Navier–Stokes equations in LES form,
with the velocity, the pressure, the density, the kinematic viscosity, and a body force including actuation effects (Montalà et al., 12 Sep 2025). The same incompressible equations are stated for the SmartFlow cylinder study, which uses a three-dimensional domain with a periodic spanwise direction at (Xiao et al., 1 Aug 2025).
The REFMAP paper reports a different formulation. There, SOD2D solves the compressible Navier–Stokes equations in conservative form for density 0, velocity 1, pressure 2, temperature 3, and total energy 4, with the ideal gas law providing closure (Eleftherakis et al., 20 Jan 2026). In that description, the solver uses unstructured spectral elements in 2D, Gauss–Lobatto–Legendre nodes with quadrature points coinciding with element nodes, and a diagonal mass structure enabling explicit time stepping. Temporal advancement is performed with fourth-order Runge–Kutta, convective and diffusive terms are treated through operator splitting, and stabilization is supplied by an entropy-viscosity model (Eleftherakis et al., 20 Jan 2026).
The numerical details reported in SmartFlow are deliberately thinner. That paper does not specify the spectral-element basis, polynomial order, time integration scheme, pressure treatment, dealiasing, or projection method; SOD2D is described there only at a high level as “high‑order spectral element” and “GPU‑accelerated” (Xiao et al., 1 Aug 2025). By contrast, the REFMAP study exposes the kernel structure and the consequences of explicit high-order spectral-element discretization for memory traffic, arithmetic intensity, and GPU occupancy (Eleftherakis et al., 20 Jan 2026).
3. Coupling architecture for reinforcement learning
A central role of SOD2D in the recent literature is as the CFD environment in DRL or MARL loops. In the wing-flow-control study, the Fortran-based solver and the Python-based TF-Agents PPO agent communicate via a Redis in-memory database orchestrated by SmartSim, which is explicitly presented as a way of solving the “two-language” problem (Montalà et al., 12 Sep 2025). The data exchange per action step follows a fixed sequence: SOD2D advances the flow and computes observations at sensor locations; these observations are written to Redis; the PPO policy reads the state and outputs jet-velocity actions; actions are written back to Redis; and SOD2D applies the actuation, continues the simulation for the action duration, and computes the reward from aerodynamic forces (Montalà et al., 12 Sep 2025).
In SmartFlow, the same general pattern is generalized into a solver-agnostic infrastructure. SmartSim starts an in-memory Redis/KeyDB Orchestrator, while SmartRedis-MPI provides MPI-aware APIs—init_smartredis_mpi, put_state, get_action, put_reward, and finalize_smartredis_mpi—for low-latency in-memory exchange between SOD2D and Python-side DRL algorithms implemented with Stable-Baselines3 (Xiao et al., 1 Aug 2025). At each control interval 5, SOD2D instances push states, the shared PPO policy computes actions, the solver retrieves those actions, applies boundary actuation for one control interval, and then computes local or global rewards (Xiao et al., 1 Aug 2025).
The control cadence is highly structured in both studies. In the wing case, each episode spans six baseline vortex-shedding cycles and contains 120 actions, so 6; ten CFD simulations are run in parallel and the domain is partitioned spanwise into three pseudo-environments per simulation, giving 30 trajectories per decision step (Montalà et al., 12 Sep 2025). In the cylinder case, each episode lasts 7 with 120 actions per episode; four concurrent SOD2D simulations, each split into 10 pseudo-environments, produce 40 trajectories per policy update (Xiao et al., 1 Aug 2025). In both settings, parallel multi-environment rollout is a primary mechanism for accelerating policy learning.
4. Control-oriented deployments
The most detailed SOD2D deployment reported to date is the active control of a separated three-dimensional NACA0012 wing section at 8 and 9 (Montalà et al., 12 Sep 2025). The computational domain extends 0, 1, and 2, with inlet 3, outlet zero-gradient velocity plus constant pressure, no-slip walls on the wing, and periodic spanwise boundaries. Two sets of three spanwise-distributed actuators are used, each set containing a front jet at 4 and a rear jet at 5, with instantaneous mass conservation enforced through
6
The agent controls only the front-jet velocity, bounded by
7
Each pseudo-environment senses pressure at 90 witness points in its mid-span slice, augmented by the two neighboring slices for a total of 270 inputs (Montalà et al., 12 Sep 2025).
The wing-study reward mixes local aerodynamic improvement and spanwise cooperation. The local reward is
8
with 9, and the aggregated reward is
0
with 1 and 2 (Montalà et al., 12 Sep 2025). Training plateaued by episode 67. Deterministic evaluation reported 3, 4, and an increase in the vortex-shedding Strouhal number from 5 to 6 (Montalà et al., 12 Sep 2025). The learned actuation locks onto 7, produces a negative-mean front jet and positive-mean rear jet, extends the shear layer farther downstream, and shrinks the time-averaged negative-8 region on the suction side (Montalà et al., 12 Sep 2025).
The SmartFlow cylinder case uses SOD2D in an incompressible bluff-body setting at 9 (Xiao et al., 1 Aug 2025). The domain is 0, 1, 2, with uniform inflow, zero-gradient outlet with constant pressure, slip top and bottom, and periodic spanwise boundaries. Ten pairs of wall-mounted jets are uniformly distributed along the spanwise direction; each pair comprises a top jet centered at 3 and a bottom jet at 4, with spanwise width 5, angular width 6, and zero net mass flux 7. The action variable is the mass flow rate per unit span 8, with 9, applied through a Dirichlet wall-jet boundary condition having a cosine profile in 0 (Xiao et al., 1 Aug 2025). Each pseudo-environment uses 85 witness points around the cylinder wall at mid-span, augmented by the two neighboring pseudo-environments to form a 255-value state vector. After 100 episodes on a single node with four NVIDIA A100 GPUs, training took approximately 40 hours; deterministic evaluation showed a 6.0% reduction in mean drag and a 46.6% reduction in lift-coefficient fluctuation standard deviation (Xiao et al., 1 Aug 2025).
5. GPU performance portability and computational behavior
The REFMAP performance study makes SOD2D’s hardware behavior unusually explicit (Eleftherakis et al., 20 Jan 2026). The dominant hotspot is full_convec, which accounts for up to 50% of runtime on NVIDIA and up to 70% on AMD; full_diffusion accounts for 22% on NVIDIA and 10% on AMD. Within each Runge–Kutta stage, execution is organized into preloading of global quantities into element-local buffers, kernel execution in natural and physical coordinates, and store-back through atomic updates to global arrays. The paper characterizes the convective kernel as predominantly memory-bound, with irregular gathers and atomic store-backs limiting operational intensity (Eleftherakis et al., 20 Jan 2026).
Single-GPU characterization spans Tesla V100, A100, and AMD MI250X platforms. On approximately 8M nodes in FP32 without code optimizations, a single MI250X GCD was reported as 15.831 slower than V100 and 23.382 slower than A100; A100, in turn, delivered 1.473 speedup over V100 in FP32 and 3.054 in FP64 (Eleftherakis et al., 20 Jan 2026). The same study shows that kernel splitting, preloading removal, and OpenACC cache prefetching have markedly vendor- and precision-dependent effects. Splitting full_convec yielded an immediate 1.55 speedup on MI250X by reducing VGPR and LDS pressure; prefetching small local arrays produced gains such as 1.246 on V100 FP64 split kernels and 42% on MI250X FP32 unified kernels (Eleftherakis et al., 20 Jan 2026).
The portability conclusion is that the same memory-access optimization can be neutral or harmful on another architecture. The measured spread is 0.697–3.918 in acceleration speedup across optimizations and platforms (Eleftherakis et al., 20 Jan 2026). Weak-scaling experiments on LUMI, at approximately 1M nodes per GPU, further show that the “best” configuration at one scale is not reliably the best at another. For channel flow, throughput variation from worst to optimal configuration reached 23.8% in FP32 and 14.6% in FP64, and projecting the 32-GPU optimum to 64 GPUs incurred penalties of 23.8% and 8.3%, respectively (Eleftherakis et al., 20 Jan 2026). A plausible implication is that SOD2D’s algorithmic structure is mature enough to expose full-stack tuning issues—application, software, and hardware—rather than being dominated by a single implementation bottleneck.
6. Reproducibility, limitations, and nomenclature
The cited literature gives partial but uneven reproducibility support. SmartFlow provides public repositories for the framework itself, SmartRedis-MPI, the SmartRedis client, and a SmartFlow–SOD2D integration example, but it does not provide a separate public repository or license for SOD2D, nor does it specify the exact solver version, standalone build scripts, or complete numerical configuration files required to reproduce the SOD2D cylinder case independently (Xiao et al., 1 Aug 2025). The wing-flow-control paper reports domain sizes, boundary conditions, actuation layout, reward parameters 9, and episode scheduling, but does not detail mesh resolution, time step, turbulence-model specifics beyond “filtered incompressible Navier–Stokes (LES),” or CFL constraints (Montalà et al., 12 Sep 2025). In both cases, exponential smoothing of the actions is identified as a numerical-robustness device (Montalà et al., 12 Sep 2025).
A persistent issue is nomenclature. The solver is named SOD2D, yet the literature explicitly emphasizes that it is used for a 3D wing section with spanwise periodicity and for a 3D cylinder flow with periodic 0-boundaries (Montalà et al., 12 Sep 2025, Xiao et al., 1 Aug 2025). The SmartFlow paper states that future work could clarify whether SOD2D is extruded 2D with periodic 1, because the handling of the third dimension is not described there (Xiao et al., 1 Aug 2025). There is also a separate, homonymous usage in the semi-Lagrangian lattice Boltzmann literature, where “SOD2D” refers not to the solver but to two-dimensional Sod- or Riemann-type benchmark problems for compressible discontinuity interaction (Wilde et al., 2019). This suggests that “SOD2D” is not a universally unambiguous identifier across arXiv, and the surrounding research context is essential for correct interpretation.