---
title: ASTRA-sim 3.0 Simulator
url: https://www.emergentmind.com/topics/astra-sim-3-0
type: topic
---

# ASTRA-sim 3.0 Simulator

Searching arXiv for ASTRA-sim lineage and closely related papers to ground the article.
ASTRA-sim 3.0 is an open-source, event-driven simulator for distributed machine learning workloads that is designed to model end-to-end performance across workload, GPU/system, and network layers with substantially higher fidelity than earlier ASTRA-sim releases. It is presented as a response to the increasing importance of latency-sensitive collective communication in large-scale artificial intelligence, particularly as model inference becomes a central distributed ML use case. The defining additions are high-fidelity GPU execution modeling at cache-line-sized load/store instruction granularity, support for arbitrary custom collective communication algorithms, a standardized infrastructure representation called InfraGraph, and more detailed intra-device and inter-device network simulation [2606.10440].

## 1. Position within the ASTRA-sim lineage

ASTRA-sim 3.0 is explicitly framed as a revision of the ASTRA-sim simulator that addresses limitations in both ASTRA-sim and ASTRA-sim 2.0. Earlier versions are described as having enabled scalable simulation of large systems and greater topology and workload flexibility, but they treated computation and communication at a coarse granularity, often as a single event per compute kernel or collective operation. That abstraction omitted detailed resource contention, control paths, memory accesses, and parallel execution within and across GPU compute units [2606.10440].

The new version is intended for design space exploration under conditions where latency, precise device behavior, and communication protocols materially affect outcomes. In that sense, ASTRA-sim 3.0 extends the simulator’s role from large-scale throughput-oriented studies toward analysis of inference-heavy and other latency-sensitive deployments. The paper’s framing suggests a shift from system-level approximation toward cross-layer co-modeling of device execution, communication protocol structure, and infrastructure topology.

## 2. GPU execution model and simulation granularity

A central feature of ASTRA-sim 3.0 is simulation at cache-line-sized load/store granularity, stated as, for example, 128B granularity, which is identified as the typical cache line size for modern GPUs. Rather than treating a collective or kernel as a single opaque event, the simulator introduces a layered execution model with primitive GPU instructions, higher-level operations, workgroups, wavefronts, and kernels [2606.10440].

The primitive instruction level includes `load`, `store`, `acquire`, `release`, `reduce`, and `waitcnt`. These are composed into operations corresponding to programming constructs such as `memcpy`, `acquire_op`, and `reduce_op`. Workgroups correspond to threadblocks in CUDA terminology, and workgroups are composed of wavefronts, described as SIMD lockstep groups analogous to warps. Kernels are then defined as sets of workgroups mapped onto actual GPU resources, namely Compute Units (CUs). Each workgroup is mapped to a CU or streaming multiprocessor, and the CUs execute wavefronts with the possibility of out-of-order progress when a stalled wavefront allows another ready wavefront to be scheduled.

This resource model is paired with explicit contention modeling. ASTRA-sim 3.0 states that resource contention among concurrent kernels and workgroups is faithfully modeled, reflecting real GPU constraints. Network requests are injected per wavefront and at cache-line granularity, permitting detailed simulation of congestion and contention. Intra-wavefront parallelism is also supported through loop unrolling, which allows a wavefront to issue multiple outstanding memory or network operations, parameterized by factors such as register file size or architectural support [2606.10440].

The technical example given in the paper illustrates the granularity of the model: when a CU sends data to a remote GPU, a cache-line chunk is loaded from HBM to a register, stored to an I/O port, transferred over the network, and then stored to remote HBM. These steps are simulated per cache line and per workgroup or wavefront, with synchronization and contention modeled explicitly. The reported implication is that bandwidth and latency predictions become accurate at the granularity required for inference and small-message collectives, while critical-path contention for memory, I/O, and compute units can be captured directly [2606.10440].

## 3. Collective algorithm representation and msccl++ integration

ASTRA-sim 3.0 extends collective communication support beyond fixed, textbook algorithms such as ring and all-pairs. The simulator can now ingest collectives described through the msccl++ language or representation, allowing arbitrary, topology-aware and protocol-aware collective algorithms to be modeled [2606.10440].

The description in the paper highlights several dimensions of this flexibility. The msccl++ representation supports distinction between `put` and `get` primitives, one-sided and two-sided communication, variable chunk sizing, and control dependencies including barriers and semaphores. ASTRA-sim 3.0 parses msccl++ JSON and maps operations such as `put`, `get`, `signal`, and `wait` onto the GPU operations in its execution model. This creates a path from algorithm-level communication descriptions to wavefront-level execution and cache-line-level network traffic.

The paper uses reduce-scatter and all-gather as performance-study examples. For reduce-scatter, get-based collectives are reported to enable better compute-communication overlap and to outperform put-based variants for large collectives because transfers and reductions overlap at cache-line granularity. For all-gather, put-based collectives are reported to be better at large scales because control message blocking can hinder get requests; the paper further notes that fair arbitration between control and data messages can minimize this gap [2606.10440].

A common misconception in simulator discussions is that collective performance can be compared meaningfully while abstracting away protocol details. ASTRA-sim 3.0 argues the opposite by construction: subtle protocol choices, including control-message handling and synchronization structure, can alter the bandwidth ranking of get-based and put-based implementations. This suggests that collective evaluation for distributed inference or model-parallel settings cannot be reduced to topology alone.

## 4. InfraGraph and standardized infrastructure description

InfraGraph is introduced as a backend-agnostic, standardized representation for distributed ML network infrastructure. Its motivation is the prior need for backend-specific formats, which made it tedious and error-prone to reproduce results across network simulators or to share infrastructure configurations [2606.10440].

InfraGraph models infrastructure as directed, attributed graphs. Nodes represent hardware entities such as GPUs, NICs, and switches, while edges represent physical connections with properties such as bandwidth and latency. The system also includes blueprints, which are reusable templates for devices and fabrics. These blueprints can be parameterized and instantiated programmatically, enabling specifications such as a cluster of 32 hosts with 8 GPUs each connected through a 2-level Clos fabric. Annotation is supported at the level of edges, devices, and ports.

Two associated tools are identified. A translator generates compatible configuration files for all ASTRA-sim network backends from a single InfraGraph specification, and a visualizer renders topologies for verification. The paper states three main benefits: reproducibility, since the same infrastructure specification yields identical simulation setups across backends; easier sharing and comparison, because communities can publish and benchmark using infrastructure graphs rather than custom backend-specific translations; and extensibility, because new backends or topologies can be incorporated more directly [2606.10440].

The example explicitly described is an 8-GPU cluster interconnected through a Clos, or fat-tree, fabric and simulated through the ns-3 backend. This example situates InfraGraph not merely as a schema but as an interoperability layer between infrastructure design and backend execution. A plausible implication is that InfraGraph functions as a common experimental artifact for distributed ML systems research, analogous to a normalized workload format for traces or kernels.

## 5. Network, memory, and scalability modeling

ASTRA-sim 3.0 also extends the detail of network and memory simulation. The Simple network backend is upgraded to represent each GPU using nodes for CUs, memory channels, and I/O ports, thereby enabling NoC-level modeling. The configuration is described as hierarchical and modular, allowing rapid exploration of both on-chip and external network designs. The simulator therefore aims to model local and global transfers consistently rather than isolating inter-GPU communication from intra-device structure [2606.10440].

The scalability claims in the paper are concrete. It reports that simulations with 128 GPUs, corresponding to 57k+ endpoints, at 128B resolution can be executed in practical wall-clock time, and that throughput degrades gracefully as system size increases. Elsewhere in the summary table, the same capability is expressed as 57,344 endpoints at 128B granularity with 128 GPUs. The paper also states that throughput scales linearly with buffer size and remains tractable with GPU count for the evaluated settings [2606.10440].

These claims are significant because high-fidelity simulation often incurs prohibitive runtime growth. ASTRA-sim 3.0 presents its cache-line-granularity model as a balance between scalability and fidelity, rather than as an unrestricted cycle-accurate simulation. That balance is central to the simulator’s intended design-space role: sufficiently detailed to expose control-path and contention effects, but still scalable enough for studies of large GPU counts and full NoC-level detail.

## 6. Benchmarks, design-space exploration, and stated use cases

The paper enumerates several benchmark classes and use cases. These include collective performance studies for put-versus-get implementations of reduce-scatter and all-gather across multiple GPUs, varying workgroup counts and collective sizes; GPU architecture sensitivity studies involving loop unrolling, register file sizing, and the maximum number of outstanding memory requests; fine-grained scalability experiments over buffer sizes from 1MB to 256MB and system scales from 2 GPUs to 128 GPUs; and infrastructure configuration and cross-validation using a 1MB ring all-reduce with 8 GPUs on a Clos fabric generated by InfraGraph and evaluated with the ns-3 backend [2606.10440].

The reported results are similarly specific. Increasing loop unrolling, representing intra-wavefront parallelism, improves bandwidth until saturating at a hardware-determined limit. Register file size, interpreted as limiting the maximum number of outstanding requests, matters only for large, bandwidth-bound collectives, and that effect also saturates. For custom collectives, the relative performance of get-based and put-based designs depends on the operation, scale, and message arbitration policy. For infrastructure experiments, the paper reports consistent completion time and bandwidth for a 1MB all-reduce on a Clos fabric configured automatically through InfraGraph [2606.10440].

The broader design-space opportunities identified in the paper include comparison of new and existing collective algorithms for latency-sensitive model-parallel or inference workloads; empirical validation of GPU microarchitectural parameters such as register file size and concurrent outstanding wavefronts; direct assessment of topology effects across Clos, ring, and multi-tier switch fabrics from a single InfraGraph description; and rapid what-if analysis spanning NoC, HBM, CU organization, topology, and protocol parameters. The paper also emphasizes that standardized, parameterized infrastructure representations reduce researcher effort and improve reproducibility.

A concise statement of the differences from ASTRA-sim versions up to 2.0 is captured below.

| Capability | ASTRA-sim ≤2.0 | ASTRA-sim 3.0 |
|---|---|---|
| GPU Modeling | Coarse, kernel-level | Cache-line, CU/wavefront-level |
| Collective Algorithms | Textbook only | Arbitrary, msccl++/custom |
| Device Control Path | Not modeled | Modeled (control/data ops, barriers, semaphores) |
| Resource Contention | Limited (serializes) | Faithfully modeled (CU/workgroup/wavefront) |
| Network Modeling | Inter-GPU only, coarse | On-chip NoC + global, per-endpoint detail |
| Infra Description | Backend-specific | InfraGraph: Standard, shareable |
| Use Cases | Training | Training + Latency-sensitive Inference |

Taken together, these features define ASTRA-sim 3.0 as a simulator for distributed ML co-design in which algorithmic collectives, GPU execution structure, and network infrastructure are all explicit objects of study. The paper’s stated position is that this level of detail is necessary for accurate critical-path estimation, especially where control operations, synchronization, and small-message behavior materially affect end-to-end latency [2606.10440].

Source: https://www.emergentmind.com/topics/astra-sim-3-0