Papers
Topics
Authors
Recent
Search
2000 character limit reached

SimAI-Bench: AI–Simulation Workflow Benchmark

Updated 12 July 2026
  • The paper presents a configurable mini-app that faithfully reproduces simulation and AI control flows and data transport patterns for in-situ and in-transit exchanges.
  • SimAI-Bench benchmarks various transport backends including Redis, DragonHPC, node-local storage, and Lustre under one-to-one and many-to-one workflow motifs.
  • The framework’s layered architecture—comprising orchestration, application, and data transport layers—enables systematic performance analysis and rapid prototyping on exascale systems.

SimAI-Bench is a Python-based mini-app framework designed to emulate, prototype, and benchmark coupled AI–simulation workflows on HPC systems, with a specific emphasis on in-situ and in-transit data transport (Tummalapalli et al., 23 Sep 2025). Rather than functioning as a single benchmark kernel, it provides a configurable environment in which workflow mini-apps reproduce the control flow, resource placement, and data movement patterns of real applications. In its reported use on the Aurora supercomputer, SimAI-Bench was applied to two common motifs—a one-to-one co-located workflow and a many-to-one ensemble workflow—to study how Redis, DragonHPC, node-local storage, and Lustre behave under different coupling structures (Tummalapalli et al., 23 Sep 2025).

1. Purpose, problem setting, and design objectives

SimAI-Bench addresses the performance analysis of coupled AI–simulation workflows, which are presented as increasingly important workloads for HPC facilities. The motivating use case is the concurrent execution of large-scale simulations with online AI training or inference, rather than offline training from previously stored files. In this setting, writing full simulation output to disk and reading it back later becomes expensive at exascale, so in-situ and in-transit exchange becomes central to end-to-end performance (Tummalapalli et al., 23 Sep 2025).

The framework is organized around three explicit objectives: workflow emulation, transport benchmarking, and prototyping. Workflow emulation means capturing simulation timesteps, training iterations, coupling frequency, and data sizes without reproducing full scientific applications. Transport benchmarking means systematically comparing staging and transport backends under realistic workflow motifs. Prototyping means providing an API through which new workflow patterns and deployment strategies can be tested rapidly on systems such as Aurora (Tummalapalli et al., 23 Sep 2025).

The design also responds to three gaps identified in the coupled-workflow literature. First, conventional mini-apps typically model individual codes or kernels rather than end-to-end workflows with multiple components and data dependencies. Second, there is limited comparative guidance on transport choices across recurring motifs such as one-to-one and many-to-one coupling. Third, implementing and porting complete simulation–AI workflows on new systems is costly, so a portable workflow mini-app can answer questions such as whether Redis or node-local SSDs are preferable, or whether a file system remains viable under ensemble aggregation (Tummalapalli et al., 23 Sep 2025).

2. Benchmark motifs and workflow semantics

The first benchmarked motif is a one-to-one workflow with co-located simulation and AI. It models a single parallel simulation coupled directly to a single AI trainer on the same nodes, using a nekRS plus GNN surrogate workflow as the concrete science case. On Aurora, each node exposes 12 GPU tiles; 6 are assigned to the simulation and 6 to the AI trainer. The simulation advances with a fixed per-iteration cost and writes a snapshot every 100 iterations, while the AI trainer advances with its own fixed iteration cost and polls for new data every 10 iterations. The two components are asynchronous, and the AI can eventually steer the simulation by signaling it to stop after a fixed number of training iterations (Tummalapalli et al., 23 Sep 2025).

In the reported configuration, the simulation mini-app uses a MatMulSimple2D kernel with run_time ≈ 0.03147 s, data_size [256,256], and device="xpu", while the AI component is configured to reproduce a training iteration time of approximately 0.061 s. This is not intended as a numerical reproduction of nekRS or the production GNN, but as a timing- and placement-faithful emulation of the makespan, coupling rate, and data exchange pattern (Tummalapalli et al., 23 Sep 2025).

The second motif is a many-to-one workflow in which an ensemble of simulations trains a single AI model. Here, NN simulation components each run on their own node, and one AI trainer runs on a separate node. Each simulation writes a data array every 10 iterations, and the AI blocks until it has ingested data from all simulations for the corresponding training update. The total data volume per AI update is therefore

Vtotal=N⋅VV_{\text{total}} = N \cdot V

where VV is the per-simulation array size. If TiT_i is the transfer time from simulation ii to the AI component, the transport time per iteration is described as

Ttransport(N)=max⁡i=1,…,NTi.T_{\text{transport}}(N) = \max_{i=1,\dots,N} T_i.

This pattern stresses non-local data access, in contrast to the one-to-one case where simulation and AI are co-located (Tummalapalli et al., 23 Sep 2025).

A recurrent implication of these two motifs is that transport behavior is pattern-dependent. This suggests that a backend that performs well for local, co-located exchange need not remain optimal when a single learner aggregates data from many remote sources.

3. Software architecture and programming model

SimAI-Bench is organized as a three-layer stack consisting of an orchestration layer, an application layer, and a data transport layer (Tummalapalli et al., 23 Sep 2025).

The orchestration layer is centered on a Workflow class. Components are registered with decorators such as @w.component, dependencies are expressed as a DAG, and tasks can be launched locally through multiprocessing or remotely through mpirun. This layer therefore models workflow structure rather than only per-process kernel timing (Tummalapalli et al., 23 Sep 2025).

The application layer exposes Simulation and AI classes. The Simulation class uses a Kernels module and can be configured through JSON or dictionaries specifying kernel names, data sizes, device placement, and either iteration counts or target run times. It can also model variability through discrete probability density functions for run_time or run_count. The AI class uses PyTorch, including torch.nn and torch.distributed, to emulate distributed data-parallel training. In the reported implementation it supports fully connected networks, while future work is stated to add GNNs and CNNs (Tummalapalli et al., 23 Sep 2025).

The data transport layer comprises ServerManager and DataStore. ServerManager configures backend services, including Redis or DragonHPC deployments and directory setup for file-based backends. DataStore provides a backend-agnostic interface with four core methods: stage_write, stage_read, poll_staged_data, and clean_staged_data. Because the client API is invariant across backends, the same workflow mini-app can be rerun with different transport strategies by altering runtime configuration rather than application logic (Tummalapalli et al., 23 Sep 2025).

The kernel repertoire is broader than matrix multiplication. The Kernels module includes compute primitives such as MatMulSimple2D, MatMulGeneral (GEMM), FFT, AXPY, in-place compute, random generation, and scatter-add; I/O primitives including single-rank writes and HDF5-based collective I/O; collective communication through mpi4py and NCCL; and host–device copy operations. This makes the framework a workflow mini-app environment rather than a narrowly specialized transport microbenchmark (Tummalapalli et al., 23 Sep 2025).

4. Transport backends and measurement methodology

Four transport and staging strategies are evaluated. Redis is used as a high-performance in-memory key–value store and is the mechanism used in SmartSim and in the production nekRS–ML workflow discussed in the paper. DragonHPC provides a distributed in-memory data dictionary intended for high-throughput, low-latency memory-to-memory transfer. Node-local storage is implemented as a sharded key–value store over tmpfs or SSD-backed directories, using CRC32 hashing for shard selection and atomic os.replace() for consistency. Lustre serves as the shared parallel file system, using similar sharded directory mapping but on the global file system rather than node-local storage (Tummalapalli et al., 23 Sep 2025).

The one-to-one experiments include all four backends, because reads and writes can remain local when simulation and AI are co-located. The many-to-one experiments exclude node-local storage because it cannot support non-local sharing across nodes. This distinction is central: transport capabilities are constrained not only by raw performance but by whether the storage abstraction supports the workflow’s communication graph (Tummalapalli et al., 23 Sep 2025).

The paper reports throughput as

B=bytes transferredtime,B = \frac{\text{bytes transferred}}{\text{time}},

with separate averages for read throughput and write throughput per process. It also reports mean event times for simulation iterations, AI iterations, reads, and writes. For many-to-one experiments, training execution time per iteration is summarized as

Texec/iter=Ttotalniter.T_{\text{exec/iter}} = \frac{T_{\text{total}}}{n_{\text{iter}}}.

The write rate in the one-to-one pattern is described conceptually as

Rwrite≈Vfsim⋅tsim,R_{\text{write}} \approx \frac{V}{f_{\text{sim}} \cdot t_{\text{sim}}},

where VV is the data volume per write, Vtotal=N⋅VV_{\text{total}} = N \cdot V0 is the number of iterations between writes, and Vtotal=N⋅VV_{\text{total}} = N \cdot V1 is the simulation iteration duration (Tummalapalli et al., 23 Sep 2025).

The experiments were run on Aurora, described as a 10,624-node HPE Cray EX system with 2 Intel Xeon CPU Max processors per node, 6 Intel Data Center GPU Max per node split into 12 tiles, 512 GB DDR5 per socket, 64 GB HBM per CPU, a Lustre parallel file system, and a high-bandwidth interconnect (Tummalapalli et al., 23 Sep 2025).

5. Empirical findings on Aurora

For the one-to-one pattern, SimAI-Bench was first validated against the real nekRS–ML workflow. The reported event counts are close: simulation timesteps are 10108 in the original workflow versus 10507 in the mini-app; simulation data transport events are 203 versus 211; training timesteps are 5000 versus 5000; and training data transport events are 208 versus 208. Iteration timing is also similar: simulation mean iteration time is Vtotal=N⋅VV_{\text{total}} = N \cdot V2 s in the original workflow and Vtotal=N⋅VV_{\text{total}} = N \cdot V3 s in the mini-app, while training mean iteration time is Vtotal=N⋅VV_{\text{total}} = N \cdot V4 s in the original workflow and Vtotal=N⋅VV_{\text{total}} = N \cdot V5 s in the mini-app. The mini-app exhibits smaller variance because of its deterministic configuration, but the asynchronous transport timeline is reported as faithful to the production workflow (Tummalapalli et al., 23 Sep 2025).

At 8 nodes in the one-to-one experiments, the in-memory and node-local backends show non-monotonic throughput as message size increases: throughput rises up to an intermediate size and then declines at the largest messages. For DragonHPC, throughput peaks around approximately 10 MB, and the paper attributes the decline at larger sizes to cache effects. Lustre, by contrast, shows monotonically increasing throughput with message size. At 512 nodes, node-local storage, Redis, and DragonHPC maintain similar per-process throughput to the 8-node case because the exchange remains local, while Lustre degrades significantly because concurrent small-file operations induce metadata contention (Tummalapalli et al., 23 Sep 2025).

The time-per-event breakdown reinforces this pattern. For node-local storage, even at 32 MB, a single transfer event is approximately comparable to one computation iteration time, but because writes occur only every 100 simulation iterations, the transport overhead remains negligible. At 512 nodes, node-local read and write times change very little. For Lustre, 32 MB transfers are similar to one iteration time at 8 nodes, but at 512 nodes the transfer time per event grows by approximately Vtotal=N⋅VV_{\text{total}} = N \cdot V6 relative to iteration time, making data transport a dominant bottleneck (Tummalapalli et al., 23 Sep 2025).

For the many-to-one pattern, the 2-node baseline already shows a different backend ordering. Redis exhibits reasonable local write throughput but poor non-local read throughput. DragonHPC achieves high throughput for both local writes and non-local reads, with the same approximate peak around 10 MB followed by decline at larger sizes. Lustre again shows monotonically increasing throughput and becomes comparable to DragonHPC for large messages (Tummalapalli et al., 23 Sep 2025).

At larger scales, the many-to-one pattern changes the end-to-end picture. At 8 nodes, execution time per training iteration increases with message size for all backends, and Redis is worst because low non-local read throughput directly inflates iteration time. At 128 nodes, Redis remains slowest, while DragonHPC is significantly slower than Lustre for smaller messages below 10 MB despite its strong point-to-point bandwidth in the 2-node case. The paper attributes this to latency and coordination costs when a single AI component fetches data from many distributed sources. For larger messages at or above 10 MB, DragonHPC and Lustre become similar in execution time per iteration. The paper’s reported conclusion is that the file system is the optimal solution among the tested strategies for the many-to-one pattern (Tummalapalli et al., 23 Sep 2025).

Two broader conclusions follow. First, there is no universal best transport solution: co-located one-to-one workflows favor node-local storage and DragonHPC, while distributed many-to-one aggregation favors Lustre. Second, high point-to-point throughput is not sufficient to predict end-to-end workflow performance; many-to-one latency and coordination can dominate.

6. Limitations, extensions, and position in the benchmarking landscape

The reported limitations are specific and methodological. Kernel fidelity is limited to reproducing iteration timing and data volume rather than full arithmetic intensity or algorithmic structure. The mini-app also underestimates runtime variance relative to production applications. The AI class currently focuses on fully connected networks and does not yet natively implement GNNs or CNNs, even though the motivating production science case involves a GNN surrogate. The transport study includes Redis, DragonHPC, node-local storage, and Lustre, but does not yet cover systems such as ADIOS2, DataSpaces, or DAOS. Finally, only two workflow motifs and one machine—Aurora—are examined in the reported study (Tummalapalli et al., 23 Sep 2025).

The planned extensions are correspondingly concrete. Future work is stated to add point-to-point streaming through ADIOS2 and DAOS support on Aurora, to benchmark more complex workflows on multiple HPC systems, and to interface SimAI-Bench with orchestration tools such as RADICAL-Pilot. The existing design already allows components built with Simulation and AI to be exported to external workflow engines such as RADICAL-Pilot and Parsl, which suggests a path toward broader workflow-system integration (Tummalapalli et al., 23 Sep 2025).

Within the wider benchmarking landscape, SimAI-Bench is narrower in scope than SAIBench’s general framework for scientific AI benchmarking, which decouples problem definitions, models, metrics, rankings, and software or hardware configurations into reusable modules (Li et al., 2022). It is also more operationally focused than the structural-interpretation variant of SAIBench, which emphasizes trusted operating ranges and error tracing across the scientific problem space (Li et al., 2023). SimAI-Bench instead targets a specific systems problem: the data transport behavior of coupled AI–simulation workflows under realistic control-flow and placement patterns. A plausible implication is that it occupies a complementary role: SAIBench formalizes benchmark composition for AI-for-science at large, whereas SimAI-Bench isolates the control plane and data plane costs of online simulation–AI coupling on exascale-class hardware.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to SimAI-Bench.