---
title: Synchronization-Free All-Reduce (SiFAR)
url: https://www.emergentmind.com/topics/synchronization-free-all-reduce-sifar
type: topic
---

# Synchronization-Free All-Reduce (SiFAR)

Searching arXiv for the specified paper and closely related work to ground the article.
Synchronization-Free All-Reduce (SiFAR) is a communication design for low-latency tensor-parallel LLM inference whose central objective is to remove the barrier synchronization overhead that dominates small-message All-Reduce during decode-time serving. The method is motivated by the observation that the rise of reasoning models and agentic systems has made LLM token-generation latency a key bottleneck: unlike chatbots, whose latency gains saturate at human reading speed, these systems generate intermediate reasoning tokens not consumed by humans, so per-token latency directly determines end-to-end response time. In this regime, batching is kept minimal, decode becomes HBM/memory-bandwidth-bound, tensor parallelism (TP) is used to shard model weights across GPUs and increase aggregate bandwidth, and the All-Reduce required after each TP layer becomes a principal source of critical-path latency. SiFAR addresses that bottleneck by redesigning All-Reduce to avoid the two synchronization barriers that surround standard oneshot and twoshot implementations, and reports reductions in All-Reduce latency of up to 52% together with end-to-end throughput improvements of 18.6% for Llama-3.1-8B and 13.1% for Qwen3.5-397B-17B at TP=8 [2607.08973].

## 1. Problem setting and motivation

SiFAR is defined in the context of low-latency decode-time inference, particularly at very small batch sizes, typically batch size 1. In that setting, the model’s weights are repeatedly loaded from HBM for each token, attention and KV-cache traffic also consume bandwidth, and larger batches, although favorable for aggregate throughput, are undesirable because they increase time-per-output-token (TPOT). The paper therefore distinguishes low-latency inference from throughput-oriented serving and emphasizes workloads in which intermediate tokens are not “read by a human,” so TPOT directly compounds into end-to-end response time [2607.08973].

Tensor parallelism mitigates the bandwidth bottleneck by splitting model weights across GPUs and loading them in parallel. However, TP also introduces inter-GPU communication after sharded layers, and TP layers typically perform two All-Reduces per layer—after attention and after the MLP/FFN. As other overheads are reduced by megakernels, CUDA Graphs, and prefetching, this communication becomes the dominant remaining cost. The paper quantifies the shift with Megakernel-based measurements on Llama-3.1-8B: TPOT drops from about 2.8 ms to 1.63 ms as TP increases from 1 to 8, but the fraction spent in All-Reduce rises to 26–30%; removing All-Reduce altogether improves throughput by 35–43% for Llama-3.1-8B at TP=8 [2607.08973].

This framing places synchronization, rather than raw transfer volume alone, at the center of the optimization target. SiFAR is therefore not merely a lower-bandwidth collective; it is a collective designed around the claim that data movement plus synchronization barriers jointly dominate low-batch decode latency.

## 2. Barrier costs in conventional oneshot and twoshot All-Reduce

The paper analyzes two standard inference collectives: oneshot and twoshot. Oneshot performs a single round of “pull all peers, reduce locally.” Each GPU pulls the entire buffer from all peers and locally reduces the received payloads to produce the final result. It has a top barrier, which requires all GPUs to finish producing their partials before transfer starts, and a bottom barrier, which requires all GPUs to finish reading the buffer before it can be reused. Its advantage is one round trip, which is favorable for small payloads, but its transfer cost scales as $K \times (N-1)$ per GPU, where $K$ is payload size and $N$ is the number of GPUs [2607.08973].

Twoshot performs reduce-scatter followed by broadcast. GPUs first reduce only a $K/N$ chunk each and then broadcast the partial results. This structure is more scalable in transfer volume, but it still incurs top and bottom barriers and, in particular, a bottom barrier that is more expensive because it must ensure broadcast writes are globally visible before the next operation [2607.08973].

For the small payloads characteristic of low-batch inference, roughly 4–32 KB, the paper reports that synchronization overhead accounts for 32–50% of oneshot latency and 49–62% of twoshot latency. This indicates that standard algorithm selection between oneshot and twoshot is insufficient because both designs retain significant barrier overheads on the critical path [2607.08973].

A concise comparison is as follows:

| Collective | Communication structure | Barrier properties |
|---|---|---|
| Oneshot | Pull all peers, reduce locally | Top barrier and bottom barrier |
| Twoshot | Reduce-scatter followed by broadcast | Top barrier and bottom barrier |
| SiFAR | Synchronization-free-style redesign using dual buffering, redundant pull, and speculative reduction | Removes the bottom barrier and reduces top-barrier overhead |

This comparison suggests that the principal innovation of SiFAR lies in decomposing the latency problem into distinct synchronization hazards and addressing them separately.

## 3. Bottom-barrier elimination through dual buffering and execution co-design

The paper identifies the bottom barrier in oneshot as enforcing a write-after-write (WAW) dependency. The hazard arises because the GPU that owns the payload buffer might overwrite it for the next operation before all peers have finished reading it. SiFAR removes this barrier by changing the execution model through co-design of communication and model execution, rather than by modifying the communication primitive in isolation [2607.08973].

The mechanism is dual buffering. Instead of reusing a single buffer across successive All-Reduces, the system alternates between two buffers. Buffer A is used for one reduction, buffer B for the next, and only then is buffer A reused. Because the old buffer is not overwritten until it is safe, the WAW dependency disappears and the bottom barrier becomes unnecessary [2607.08973].

The paper also explains why such a strategy is not used by default in communication libraries. Libraries such as NCCL are application-agnostic and cannot assume anything about future buffer reuse. In serving frameworks such as vLLM, aggressive memory reuse means that the input buffer of one All-Reduce often becomes the output buffer of the next GEMM, naturally creating the hazard. SiFAR addresses this at the system level by co-designing communication and model execution so that two buffers can be kept alive without changing the memory manager internals [2607.08973].

A concrete implementation sketch is given as:

```python
buf = redundant_pull(input_buf0)
...
buf = redundant_pull(input_buf1, input_buf0)
```

The extra argument is used only to keep `input_buf0` alive. The factual significance of this detail is that bottom-barrier elimination depends on lifetime management external to the communication library proper. By itself, however, dual buffering is not sufficient, because oneshot still scales poorly with GPU count [2607.08973].

## 4. Redundant pull and in-switch reduction

After removing the bottom barrier, SiFAR addresses the scaling problem of oneshot by introducing redundant pull. In standard oneshot, each GPU receives the full payload from every other GPU, so transfer volume increases with TP degree. Redundant pull changes the data path by exploiting in-switch reduction available in modern switches [2607.08973].

Instead of each GPU pulling raw payloads and reducing locally, each GPU issues a hardware-assisted reduction request using NVIDIA’s `multimem.ld_reduce` primitive. The switch pulls data from all GPUs, performs the reduction in-network, and returns the reduced value to the requester. The paper calls this “redundant pull” because all GPUs request reduction on the same full payload, rather than requesting disjoint shards [2607.08973].

The resulting transfer properties are central to the method. In oneshot, each GPU receives $K \times (N-1)$ bytes. In redundant pull, each GPU receives only $K$ bytes of reduced output. This makes the transfer behavior much closer to twoshot-like scalability while preserving the single-round structure that works well with dual buffering [2607.08973].

The paper motivates the viability of this design with an empirical observation on H200/NVSwitch systems: reduce-scatter latency stays nearly flat up to 64 KB and only slightly increases up to 256 KB. At TP=8, a 256 KB reduce-scatter implies that each GPU only reduces 32 KB at the switch, which is close to the 4–32 KB payloads common in low-batch inference. This suggests that the switch can handle the full payload per GPU without becoming much slower. The paper explicitly reports that redundant pull plus dual buffering achieves the lowest latency, whereas oneshot plus dual buffering removes the bottom barrier but still suffers poor transfer scaling [2607.08973].

This suggests that SiFAR’s scalability does not come from abandoning oneshot-like behavior entirely, but from preserving its low-round-trip structure while relocating the reduction to the network fabric.

## 5. Speculative reduction and top-barrier minimization

Even with a better data path and no bottom barrier, a standard All-Reduce still begins with a top barrier. This barrier enforces a read-after-write (RAW) dependency: no participant can begin communication until every GPU has produced its partial output. SiFAR argues that, in low-latency decode, this requirement is often stricter than necessary because the participating GPUs are already tightly synchronized [2607.08973].

The paper’s rationale depends on properties of the execution environment. Forward passes are launched via CUDA Graphs or fused into a Megakernel, execution stays entirely on the GPU, and after the first All-Reduce, GPUs tend to remain in lockstep because they execute the same kernels with the same tensor shapes. The paper decomposes top-barrier cost into divergence, which is the wait for the slowest GPU, and flag exchange, which is the exchange of synchronization flags. It reports that divergence is near-zero after the first All-Reduce, so the barrier is mostly static synchronization overhead [2607.08973].

SiFAR therefore begins the communication phase immediately, without waiting for explicit barrier completion. Each GPU assumes peers are ready and starts fetching or reducing remotely. Correctness is enforced with a lightweight validation mechanism: each GPU writes a flag before the All-Reduce, that flag is reduced alongside the payload, and after reduction the result is checked against the expected value
$$
\text{reduced\_flag} = N \times \text{flag}.
$$
If the equality holds, all GPUs had valid data and the speculative attempt succeeded; otherwise, the All-Reduce is retried [2607.08973].

The implementation sketch given in the paper is:

```python
valid_buf = dual buffered
valid_buf[blockIdx.x] = flag
__syncthreads()
result = switch_reduce(payload_buf)
valid_out = local_reduce(valid_buf)
while (valid_out != ngpus x flag) {
    result = switch_reduce(payload_buf)
    valid_out = local_reduce_cg(valid_buf)
    __syncthreads()
}
flag = flag + 1
```

The paper notes several details: the payload uses `switch_reduce`, validation uses a separate local-reduce path because it is empirically faster than using the switch for the tiny validation buffer, the `.cg` cache modifier is used on retry to bypass L1 and fetch fresh validation data, and the validation flag is incremented each invocation so that each All-Reduce has a unique check [2607.08973].

The correctness argument relies on monotonic GPU progress: if one GPU is ahead at time $T_1$, it should remain ahead at later time $T_2$. The paper describes this assumption as reasonable because CUDA Graphs and Megakernels remove host-side variability, GPU kernels run to completion unless preempted, and large divergence is rare in the target regime. Correctness was validated by matching SiFAR output against twoshot over 100,000 decode iterations, with no silent errors reported [2607.08973].

## 6. Experimental regime and quantitative results

The evaluation is targeted specifically at the low-latency regime. SiFAR is integrated into Megakernels, which fuse the forward pass into one CUDA kernel, reduce launch overhead, improve HBM utilization, and prefetch weights. This integration is consequential because it reduces non-communication overheads and makes communication easier to isolate experimentally [2607.08973].

The evaluated models are Llama-3.1-8B, described as a dense transformer, and Qwen3.5-397B-17B, described as an MoE model. Both use FP8. The hardware platform is a single node with 8 × NVIDIA H200 GPUs, each with 141 GB HBM3e, under CUDA 12.9. The TP configurations are TP = 2, 4, 8 for Llama-3.1-8B and TP = 4, 8 for Qwen3.5-397B-17B. The main focus is batch size 1, input context lengths range from 1K to 16K tokens, and end-to-end throughput tests often use input and output length of 1000. Reported metrics include standalone All-Reduce latency, TPOT or token generation latency, end-to-end throughput in tokens/sec, and latency percentiles including tail behavior [2607.08973].

The headline quantitative results are summarized below:

| Measurement | Reported result | Context |
|---|---|---|
| All-Reduce latency reduction | Up to 52% | Compared to the best of oneshot and twoshot for payloads up to 32 KB |
| Llama-3.1-8B throughput improvement | Up to 18.6% | TP=8 |
| Qwen3.5-397B-17B throughput improvement | Up to 13.1% | TP=8 |

More detailed standalone All-Reduce results show that, relative to the best of oneshot and twoshot, SiFAR reduces 8 KB latency from 3.36 μs to 2.36 μs at TP=2, from 4.38 μs to 2.39 μs at TP=4, and from 5.11 μs to 2.44 μs at TP=8, which the paper describes as about a 2× improvement at TP=8. It also outperforms a TRT-LLM oneshot variant, Lamport-style push-based reduction, and NCCL 2.30 auto-selection [2607.08973].

For end-to-end throughput, the paper reports the following dual values in the plots. For Llama-3.1-8B, gains are +7.0% / +9.1% at TP=2, +12.4% / +14.9% at TP=4, and +15.2% / +18.6% at TP=8. For Qwen3.5-397B-17B, gains are +8.7% / +9.2% at TP=4 and +12.2% / +13.1% at TP=8. The paper notes that the commonly cited improvements are the upper-end values, such as 18.6% and 13.1% at TP=8 [2607.08973].

The paper also decomposes the improvement for Llama-3.1-8B at TP=8: redundant pull contributes 6.7%, dual buffering adds 3.4%, and speculative reduction adds up to 8.5%. Together these contributions account for most of the total gain and are presented as complementary rather than redundant. With perfect speculation, the improvement would rise from 18.6% to 20.2%; removing barriers entirely would yield 21.6%; and mis-speculation plus validation cost only about 3% of the potential gain. The average latency remains below the baseline even with retries, and tail latency remains improved through p99 and p99.9 [2607.08973].

## 7. Significance, scope, and limitations

SiFAR is most directly relevant to low-batch, latency-sensitive inference, especially for reasoning models, agentic systems, multi-step tool-using systems, and workloads in which each generated token triggers another internal step. In these settings, token latency compounds across the execution chain, so reductions on the order of microseconds per All-Reduce can materially affect end-to-end response time [2607.08973].

The method also changes the TP scaling tradeoff described in the paper. Without SiFAR, increasing TP eventually runs into synchronization overhead that erodes bandwidth gains. With SiFAR, larger TP degrees remain beneficial longer because the main synchronization costs are removed or hidden. A plausible implication is that SiFAR makes TP more attractive in precisely the regime where batch size is too small to hide communication with compute. The paper particularly situates the technique alongside megakernels and CUDA Graphs, because those approaches already minimize compute-side overhead and leave communication as the major remaining bottleneck [2607.08973].

Several boundaries are explicit in the presented evidence. The evaluation is on a single node with 8 × NVIDIA H200 GPUs and relies on in-switch reduction in modern switches. The speculative mechanism relies on monotonic GPU progress and on the empirical observation that divergence is near-zero after the first All-Reduce. The paper validates correctness over 100,000 decode iterations, but the target operating regime remains tightly defined: low-latency decode, minimal batching, bandwidth-bound execution, and small All-Reduce payloads [2607.08973].

In that sense, SiFAR is best understood not as a general-purpose replacement for all All-Reduce implementations, but as a specialized collective for tensor-parallel decode in modern LLM serving systems. Its core contribution is to reframe the bottleneck as synchronization around communication and to address that bottleneck through three coordinated mechanisms: elimination of the bottom barrier with dual buffering and execution co-design, improved scaling via redundant pull and in-switch reduction, and reduction of top-barrier overhead through speculative reduction with lightweight validation [2607.08973].

Source: https://www.emergentmind.com/topics/synchronization-free-all-reduce-sifar