Papers
Topics
Authors
Recent
Search
2000 character limit reached

Synchronization-Free All-Reduce (SiFAR)

Updated 14 July 2026
  • SiFAR is a novel communication design for tensor-parallel LLM inference that eliminates synchronization barriers to reduce decode-time latency.
  • It leverages dual buffering, redundant pull with in-switch reduction, and speculative reduction to minimize both top- and bottom-barrier overheads.
  • Experimental evaluations show up to 52% All-Reduce latency reduction and 18.6% throughput improvement for models like Llama-3.1-8B at high tensor parallelism.

Searching arXiv for the specified paper and closely related work to ground the article. Synchronization-Free All-Reduce (SiFAR) is a communication design for low-latency tensor-parallel LLM inference whose central objective is to remove the barrier synchronization overhead that dominates small-message All-Reduce during decode-time serving. The method is motivated by the observation that the rise of reasoning models and agentic systems has made LLM token-generation latency a key bottleneck: unlike chatbots, whose latency gains saturate at human reading speed, these systems generate intermediate reasoning tokens not consumed by humans, so per-token latency directly determines end-to-end response time. In this regime, batching is kept minimal, decode becomes HBM/memory-bandwidth-bound, tensor parallelism (TP) is used to shard model weights across GPUs and increase aggregate bandwidth, and the All-Reduce required after each TP layer becomes a principal source of critical-path latency. SiFAR addresses that bottleneck by redesigning All-Reduce to avoid the two synchronization barriers that surround standard oneshot and twoshot implementations, and reports reductions in All-Reduce latency of up to 52% together with end-to-end throughput improvements of 18.6% for Llama-3.1-8B and 13.1% for Qwen3.5-397B-17B at TP=8 (Taneja et al., 9 Jul 2026).

1. Problem setting and motivation

SiFAR is defined in the context of low-latency decode-time inference, particularly at very small batch sizes, typically batch size 1. In that setting, the model’s weights are repeatedly loaded from HBM for each token, attention and KV-cache traffic also consume bandwidth, and larger batches, although favorable for aggregate throughput, are undesirable because they increase time-per-output-token (TPOT). The paper therefore distinguishes low-latency inference from throughput-oriented serving and emphasizes workloads in which intermediate tokens are not “read by a human,” so TPOT directly compounds into end-to-end response time (Taneja et al., 9 Jul 2026).

Tensor parallelism mitigates the bandwidth bottleneck by splitting model weights across GPUs and loading them in parallel. However, TP also introduces inter-GPU communication after sharded layers, and TP layers typically perform two All-Reduces per layer—after attention and after the MLP/FFN. As other overheads are reduced by megakernels, CUDA Graphs, and prefetching, this communication becomes the dominant remaining cost. The paper quantifies the shift with Megakernel-based measurements on Llama-3.1-8B: TPOT drops from about 2.8 ms to 1.63 ms as TP increases from 1 to 8, but the fraction spent in All-Reduce rises to 26–30%; removing All-Reduce altogether improves throughput by 35–43% for Llama-3.1-8B at TP=8 (Taneja et al., 9 Jul 2026).

This framing places synchronization, rather than raw transfer volume alone, at the center of the optimization target. SiFAR is therefore not merely a lower-bandwidth collective; it is a collective designed around the claim that data movement plus synchronization barriers jointly dominate low-batch decode latency.

2. Barrier costs in conventional oneshot and twoshot All-Reduce

The paper analyzes two standard inference collectives: oneshot and twoshot. Oneshot performs a single round of “pull all peers, reduce locally.” Each GPU pulls the entire buffer from all peers and locally reduces the received payloads to produce the final result. It has a top barrier, which requires all GPUs to finish producing their partials before transfer starts, and a bottom barrier, which requires all GPUs to finish reading the buffer before it can be reused. Its advantage is one round trip, which is favorable for small payloads, but its transfer cost scales as K×(N−1)K \times (N-1) per GPU, where KK is payload size and NN is the number of GPUs (Taneja et al., 9 Jul 2026).

Twoshot performs reduce-scatter followed by broadcast. GPUs first reduce only a K/NK/N chunk each and then broadcast the partial results. This structure is more scalable in transfer volume, but it still incurs top and bottom barriers and, in particular, a bottom barrier that is more expensive because it must ensure broadcast writes are globally visible before the next operation (Taneja et al., 9 Jul 2026).

For the small payloads characteristic of low-batch inference, roughly 4–32 KB, the paper reports that synchronization overhead accounts for 32–50% of oneshot latency and 49–62% of twoshot latency. This indicates that standard algorithm selection between oneshot and twoshot is insufficient because both designs retain significant barrier overheads on the critical path (Taneja et al., 9 Jul 2026).

A concise comparison is as follows:

Collective Communication structure Barrier properties
Oneshot Pull all peers, reduce locally Top barrier and bottom barrier
Twoshot Reduce-scatter followed by broadcast Top barrier and bottom barrier
SiFAR Synchronization-free-style redesign using dual buffering, redundant pull, and speculative reduction Removes the bottom barrier and reduces top-barrier overhead

This comparison suggests that the principal innovation of SiFAR lies in decomposing the latency problem into distinct synchronization hazards and addressing them separately.

3. Bottom-barrier elimination through dual buffering and execution co-design

The paper identifies the bottom barrier in oneshot as enforcing a write-after-write (WAW) dependency. The hazard arises because the GPU that owns the payload buffer might overwrite it for the next operation before all peers have finished reading it. SiFAR removes this barrier by changing the execution model through co-design of communication and model execution, rather than by modifying the communication primitive in isolation (Taneja et al., 9 Jul 2026).

The mechanism is dual buffering. Instead of reusing a single buffer across successive All-Reduces, the system alternates between two buffers. Buffer A is used for one reduction, buffer B for the next, and only then is buffer A reused. Because the old buffer is not overwritten until it is safe, the WAW dependency disappears and the bottom barrier becomes unnecessary (Taneja et al., 9 Jul 2026).

The paper also explains why such a strategy is not used by default in communication libraries. Libraries such as NCCL are application-agnostic and cannot assume anything about future buffer reuse. In serving frameworks such as vLLM, aggressive memory reuse means that the input buffer of one All-Reduce often becomes the output buffer of the next GEMM, naturally creating the hazard. SiFAR addresses this at the system level by co-designing communication and model execution so that two buffers can be kept alive without changing the memory manager internals (Taneja et al., 9 Jul 2026).

A concrete implementation sketch is given as:

1
2
3
buf = redundant_pull(input_buf0)
...
buf = redundant_pull(input_buf1, input_buf0)

The extra argument is used only to keep input_buf0 alive. The factual significance of this detail is that bottom-barrier elimination depends on lifetime management external to the communication library proper. By itself, however, dual buffering is not sufficient, because oneshot still scales poorly with GPU count (Taneja et al., 9 Jul 2026).

4. Redundant pull and in-switch reduction

After removing the bottom barrier, SiFAR addresses the scaling problem of oneshot by introducing redundant pull. In standard oneshot, each GPU receives the full payload from every other GPU, so transfer volume increases with TP degree. Redundant pull changes the data path by exploiting in-switch reduction available in modern switches (Taneja et al., 9 Jul 2026).

Instead of each GPU pulling raw payloads and reducing locally, each GPU issues a hardware-assisted reduction request using NVIDIA’s multimem.ld_reduce primitive. The switch pulls data from all GPUs, performs the reduction in-network, and returns the reduced value to the requester. The paper calls this “redundant pull” because all GPUs request reduction on the same full payload, rather than requesting disjoint shards (Taneja et al., 9 Jul 2026).

The resulting transfer properties are central to the method. In oneshot, each GPU receives K×(N−1)K \times (N-1) bytes. In redundant pull, each GPU receives only KK bytes of reduced output. This makes the transfer behavior much closer to twoshot-like scalability while preserving the single-round structure that works well with dual buffering (Taneja et al., 9 Jul 2026).

The paper motivates the viability of this design with an empirical observation on H200/NVSwitch systems: reduce-scatter latency stays nearly flat up to 64 KB and only slightly increases up to 256 KB. At TP=8, a 256 KB reduce-scatter implies that each GPU only reduces 32 KB at the switch, which is close to the 4–32 KB payloads common in low-batch inference. This suggests that the switch can handle the full payload per GPU without becoming much slower. The paper explicitly reports that redundant pull plus dual buffering achieves the lowest latency, whereas oneshot plus dual buffering removes the bottom barrier but still suffers poor transfer scaling (Taneja et al., 9 Jul 2026).

This suggests that SiFAR’s scalability does not come from abandoning oneshot-like behavior entirely, but from preserving its low-round-trip structure while relocating the reduction to the network fabric.

5. Speculative reduction and top-barrier minimization

Even with a better data path and no bottom barrier, a standard All-Reduce still begins with a top barrier. This barrier enforces a read-after-write (RAW) dependency: no participant can begin communication until every GPU has produced its partial output. SiFAR argues that, in low-latency decode, this requirement is often stricter than necessary because the participating GPUs are already tightly synchronized (Taneja et al., 9 Jul 2026).

The paper’s rationale depends on properties of the execution environment. Forward passes are launched via CUDA Graphs or fused into a Megakernel, execution stays entirely on the GPU, and after the first All-Reduce, GPUs tend to remain in lockstep because they execute the same kernels with the same tensor shapes. The paper decomposes top-barrier cost into divergence, which is the wait for the slowest GPU, and flag exchange, which is the exchange of synchronization flags. It reports that divergence is near-zero after the first All-Reduce, so the barrier is mostly static synchronization overhead (Taneja et al., 9 Jul 2026).

SiFAR therefore begins the communication phase immediately, without waiting for explicit barrier completion. Each GPU assumes peers are ready and starts fetching or reducing remotely. Correctness is enforced with a lightweight validation mechanism: each GPU writes a flag before the All-Reduce, that flag is reduced alongside the payload, and after reduction the result is checked against the expected value

reduced_flag=N×flag.\text{reduced\_flag} = N \times \text{flag}.

If the equality holds, all GPUs had valid data and the speculative attempt succeeded; otherwise, the All-Reduce is retried (Taneja et al., 9 Jul 2026).

The implementation sketch given in the paper is:

KK0

The paper notes several details: the payload uses switch_reduce, validation uses a separate local-reduce path because it is empirically faster than using the switch for the tiny validation buffer, the .cg cache modifier is used on retry to bypass L1 and fetch fresh validation data, and the validation flag is incremented each invocation so that each All-Reduce has a unique check (Taneja et al., 9 Jul 2026).

The correctness argument relies on monotonic GPU progress: if one GPU is ahead at time T1T_1, it should remain ahead at later time T2T_2. The paper describes this assumption as reasonable because CUDA Graphs and Megakernels remove host-side variability, GPU kernels run to completion unless preempted, and large divergence is rare in the target regime. Correctness was validated by matching SiFAR output against twoshot over 100,000 decode iterations, with no silent errors reported (Taneja et al., 9 Jul 2026).

6. Experimental regime and quantitative results

The evaluation is targeted specifically at the low-latency regime. SiFAR is integrated into Megakernels, which fuse the forward pass into one CUDA kernel, reduce launch overhead, improve HBM utilization, and prefetch weights. This integration is consequential because it reduces non-communication overheads and makes communication easier to isolate experimentally (Taneja et al., 9 Jul 2026).

The evaluated models are Llama-3.1-8B, described as a dense transformer, and Qwen3.5-397B-17B, described as an MoE model. Both use FP8. The hardware platform is a single node with 8 × NVIDIA H200 GPUs, each with 141 GB HBM3e, under CUDA 12.9. The TP configurations are TP = 2, 4, 8 for Llama-3.1-8B and TP = 4, 8 for Qwen3.5-397B-17B. The main focus is batch size 1, input context lengths range from 1K to 16K tokens, and end-to-end throughput tests often use input and output length of 1000. Reported metrics include standalone All-Reduce latency, TPOT or token generation latency, end-to-end throughput in tokens/sec, and latency percentiles including tail behavior (Taneja et al., 9 Jul 2026).

The headline quantitative results are summarized below:

Measurement Reported result Context
All-Reduce latency reduction Up to 52% Compared to the best of oneshot and twoshot for payloads up to 32 KB
Llama-3.1-8B throughput improvement Up to 18.6% TP=8
Qwen3.5-397B-17B throughput improvement Up to 13.1% TP=8

More detailed standalone All-Reduce results show that, relative to the best of oneshot and twoshot, SiFAR reduces 8 KB latency from 3.36 μs to 2.36 μs at TP=2, from 4.38 μs to 2.39 μs at TP=4, and from 5.11 μs to 2.44 μs at TP=8, which the paper describes as about a 2× improvement at TP=8. It also outperforms a TRT-LLM oneshot variant, Lamport-style push-based reduction, and NCCL 2.30 auto-selection (Taneja et al., 9 Jul 2026).

For end-to-end throughput, the paper reports the following dual values in the plots. For Llama-3.1-8B, gains are +7.0% / +9.1% at TP=2, +12.4% / +14.9% at TP=4, and +15.2% / +18.6% at TP=8. For Qwen3.5-397B-17B, gains are +8.7% / +9.2% at TP=4 and +12.2% / +13.1% at TP=8. The paper notes that the commonly cited improvements are the upper-end values, such as 18.6% and 13.1% at TP=8 (Taneja et al., 9 Jul 2026).

The paper also decomposes the improvement for Llama-3.1-8B at TP=8: redundant pull contributes 6.7%, dual buffering adds 3.4%, and speculative reduction adds up to 8.5%. Together these contributions account for most of the total gain and are presented as complementary rather than redundant. With perfect speculation, the improvement would rise from 18.6% to 20.2%; removing barriers entirely would yield 21.6%; and mis-speculation plus validation cost only about 3% of the potential gain. The average latency remains below the baseline even with retries, and tail latency remains improved through p99 and p99.9 (Taneja et al., 9 Jul 2026).

7. Significance, scope, and limitations

SiFAR is most directly relevant to low-batch, latency-sensitive inference, especially for reasoning models, agentic systems, multi-step tool-using systems, and workloads in which each generated token triggers another internal step. In these settings, token latency compounds across the execution chain, so reductions on the order of microseconds per All-Reduce can materially affect end-to-end response time (Taneja et al., 9 Jul 2026).

The method also changes the TP scaling tradeoff described in the paper. Without SiFAR, increasing TP eventually runs into synchronization overhead that erodes bandwidth gains. With SiFAR, larger TP degrees remain beneficial longer because the main synchronization costs are removed or hidden. A plausible implication is that SiFAR makes TP more attractive in precisely the regime where batch size is too small to hide communication with compute. The paper particularly situates the technique alongside megakernels and CUDA Graphs, because those approaches already minimize compute-side overhead and leave communication as the major remaining bottleneck (Taneja et al., 9 Jul 2026).

Several boundaries are explicit in the presented evidence. The evaluation is on a single node with 8 × NVIDIA H200 GPUs and relies on in-switch reduction in modern switches. The speculative mechanism relies on monotonic GPU progress and on the empirical observation that divergence is near-zero after the first All-Reduce. The paper validates correctness over 100,000 decode iterations, but the target operating regime remains tightly defined: low-latency decode, minimal batching, bandwidth-bound execution, and small All-Reduce payloads (Taneja et al., 9 Jul 2026).

In that sense, SiFAR is best understood not as a general-purpose replacement for all All-Reduce implementations, but as a specialized collective for tensor-parallel decode in modern LLM serving systems. Its core contribution is to reframe the bottleneck as synchronization around communication and to address that bottleneck through three coordinated mechanisms: elimination of the bottom barrier with dual buffering and execution co-design, improved scaling via redundant pull and in-switch reduction, and reduction of top-barrier overhead through speculative reduction with lightweight validation (Taneja et al., 9 Jul 2026).

Definition Search Book Streamline Icon: https://streamlinehq.com
References (1)

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to Synchronization-Free All-Reduce (SiFAR).