Papers
Topics
Authors
Recent
Search
2000 character limit reached

GPU-Initiated Communication: Dissecting Down to the Bone

Published 1 Oct 2026 in cs.DC, cs.NI, and cs.PF | (2610.01380v1)

Abstract: GPU-initiated communication lets GPU threads post RDMA operations directly to the NIC. It underpins NVSHMEM, NCCL GIN, and DeepEP, which serve the fine-grained, latency-critical communication of Mixture-of-Experts (MoE) models, yet its performance characteristics and optimizations remain scarcely documented beyond source code, and library comparisons fail to separate the costs of the hardware mechanism from those of the library around it. This paper dissects GPU-initiated communication at the GPU-NIC boundary. We first detail the GPU-side network path: queue placement, work-request construction, doorbell ordering, and completion semantics. We then introduce mini-gda and mini-proxy, minimal transports for the GPU and CPU-proxy submission paths, and measure them alongside NVSHMEM IBGDA, NCCL GIN, DeepEP, UCCL-EP, MSCCL++, and fabric-lib on NVIDIA H100, H200, B200, and GB200 platforms. A minimal GPU path issues an operation in 0.7 μμs and completes in 4.0 μμs; libraries add up to 4.6 μμs of issue time through queue management, memory ordering, and completion scope, and issue time scales with the SM clock. A tuned CPU proxy matches or beats the GPU path at idle, at the cost of a dedicated core whose operating state sets its latency and throughput. On either path, sharing a queue with bulk traffic raises latency by one to three orders of magnitude. Reaching the 260 M msg/s ceiling of our InfiniBand platform requires doorbell batching and queue parallelism, and both have resource costs: communication code can reduce GPU block residency even when unused, and all-to-all traffic loses 59% of its NIC message rate at about 3,000 active connections. The submission path alone therefore does not predict communication performance. Our experiment code and results are available at https://github.com/ParCoreLab/Dissecting-GPU-Communication-Experiments.

Summary

  • The paper finds that the minimal implementation 'mini-gda' issues 8-byte messages in 0.70 microseconds suggesting GPU-initiated communication is efficient.
  • Benchmarks show that dedicated CPU workers can sometimes match GPU submission’s latency based on system topology.
  • The study finds that GPU submission vs. CPU-proxy results are heavily influenced by batching, memory ordering and synchronization policies, and the CPU's state, emphasizing a nuanced performance landscape.

Research problem and methodological position

“GPU-Initiated Communication: Dissecting Down to the Bone” (2610.01380) examines the GPU–NIC boundary rather than treating GPU communication libraries as indivisible systems. The paper addresses a methodological problem in existing evaluations: measurements of NVSHMEM IBGDA, NCCL GIN, DeepEP, CPU proxies, and related transports conflate hardware costs with library-level decisions involving queue management, memory ordering, doorbell batching, completion scope, and connection provisioning.

The paper separates communication control—the GPU’s decision about what must be transferred—from submission—the construction and notification of an RDMA operation. This distinction is important because a GPU may construct a work request while a CPU submits it, or both construction and submission may occur on the GPU. The evaluation therefore compares not simply “GPU communication” with “CPU communication,” but several concrete paths:

  • GPU construction and GPU doorbell submission;
  • GPU construction with CPU-proxy submission;
  • host construction and host submission;
  • production implementations with their complete library abstractions.

The central experimental instruments are two minimal transports. mini-gda places RDMA queues in GPU-accessible memory, constructs WQEs on the GPU, updates the doorbell record, rings the NIC doorbell, and polls completions. mini-proxy has GPU threads enqueue compact descriptors into host-visible rings while CPU workers construct and submit RDMA operations. These implementations expose queue count, ordering mode, payload placement, batching, worker count, and completion scope as independent variables. Production systems—including NVSHMEM IBGDA and IBRC, NCCL GIN GDAKI and proxy paths, DeepEP, UCCL-EP, MSCCL++, and fabric-lib—are then evaluated against these mechanisms.

The platforms include NVIDIA H100, H200, B200, and GB200 systems with ConnectX-7 NICs, using both InfiniBand and RoCEv2. The experiments emphasize 8-byte operations because they expose control-path costs, then increase payload sizes to identify where network bandwidth dominates posting overhead.

The GPU–NIC submission path

The paper gives a hardware-level account of GPU-initiated RDMA. A GPU thread constructs an mlx5 send WQE in a GPU-resident send queue, updates the queue’s doorbell record, executes the required ordering operations, and performs an ordered 64-bit store to a NIC User Access Region. The doorbell contains the WQE index and queue-pair identifier; the NIC subsequently fetches the WQE, performs the RDMA operation, writes a completion queue entry, and the GPU polls that completion.

WQEs are composed of 16-byte segments and fetched in 64-byte basic blocks. Reliable Connection WQEs contain control, remote-address, and data segments, whereas Dynamically Connected transports add an address-vector segment. Payloads can be represented by pointers or inlined directly into the WQE. Inlining avoids a NIC-side source-buffer read for very small values but increases GPU stores and doorbell traffic as the payload grows.

Figure 1

Figure 1: Inline payloads reduce completion latency for very small writes, but larger inline payloads increase issue cost and reduce batched message rate.

The placement of the doorbell is a particularly consequential implementation detail. Queue buffers and completion queues can reside in GPU memory because the NIC accesses them through GPUDirect RDMA mappings. The doorbell register, however, is a NIC BAR register in a UAR and must be mapped into the GPU address space for direct GPU submission. Systems that do not permit this reverse PCIe mapping can retain GPU-built WQEs while forwarding only the doorbell operation to a CPU thread.

The paper identifies two ordering obligations. First, WQE and source-payload stores must become visible to the NIC before the doorbell announces them. Second, when multiple GPU threads construct WQEs on a shared queue, the thread that publishes the queue must not expose a WQE before the originating thread’s stores are visible. Production libraries implement these requirements with different fence scopes and release semantics. This variation becomes a major source of latency differences.

Completion semantics are equally important. An RDMA completion can retire multiple preceding unsignaled WQEs on the same queue, but a completion routine may poll one queue, one peer, one context, or every configured queue. The latter choice is particularly expensive when throughput-oriented configurations use many QPs.

Single-operation latency: mechanism versus library

The minimal GPU path establishes a substantially lower latency floor than the production libraries measured. On the primary InfiniBand platform, mini-gda issues an 8-byte write in 0.70 microseconds and returns from put-plus-completion in 4.03 microseconds. The observed completion appears approximately 3.30 microseconds after the doorbell, including doorbell delivery, NIC and network processing, and GPU polling.

The production paths add significant software overhead:

Path Issue time Put plus completion Round trip
mini-gda, inline 0.70 us 4.03 us 6.85 us
NCCL GIN GDAKI, inline 1.60 us 6.02 us 10.50 us
NVSHMEM internal 4.26 us 9.18 us —
NVSHMEM public API 5.31 us 11.10 us 21.50 us
Tuned mini-proxy 0.13 us enqueue 4.10 us 5.89 us

The paper’s contradictory claim is that a CPU proxy can match or outperform GPU submission for latency. The tuned proxy produces a lower median round trip than mini-gda—5.89 versus 6.85 microseconds—despite requiring a CPU worker. This result is not an intrinsic advantage of CPU submission. It depends on a highly specialized protocol: one producer per ring, cached progress, validity encoded in the descriptor, and a dedicated worker pinned near the NIC. The baseline proxy, which performs atomic ring reservation, host-progress reads, and system-scope release, takes 2.27 microseconds merely to enqueue a request and has a 14.59-microsecond round trip.

The comparison therefore supports the paper’s principal attribution: the hardware mechanism itself is relatively inexpensive, while library-level queue and synchronization policies dominate single-operation latency.

Ordering and completion scope

The ordering experiment isolates the cost of memory-safety guarantees. On P-IB, a GPU-scope fence yields a 0.70-microsecond issue time, whereas a system-scope fence requires 2.62 microseconds. Thus, system-scope ordering costs approximately 3.7 times as much as GPU-scope ordering for the tested operation. An unsafe configuration with fences removed reaches 0.19 microseconds, but the authors explicitly do not claim that this configuration is generally correct. The absence of observed corruption is not a proof of correctness under the GPU and NIC memory models.

The result has a direct design implication: ordering scope must match the actual memory geography and synchronization contract. Using system-scope fences for GPU-resident queues imposes a measurable cost without necessarily strengthening the relevant correctness guarantee.

Payload inlining exhibits a narrow optimum. Inlining an 8-byte value saves 0.61 microseconds of completion time without increasing issue time. However, each additional 16-byte inline chunk adds roughly 0.1 microseconds to issue time, and the completion advantage has disappeared by 92 bytes. This explains why the evaluated libraries inline only the smallest values.

Completion scope creates another independent latency tax.

Figure 2

Figure 2: PE-wide completion grows with the number of configured queues, whereas completion restricted to the used queue remains approximately flat.

NVSHMEM’s PE-wide quiet scans every configured RC QP. Increasing the number of QPs from 1 to 16 doubles put-plus-completion latency. In contrast, a completion routine restricted to the used QP remains approximately flat. On the RoCE platform, the all-QP routine grows from 18.02 microseconds at one QP to 34.85 microseconds at 32 QPs, while used-QP completion remains near 17 microseconds. The implication is that benchmarking a high-throughput configuration with a broad completion primitive can obscure the cost of the operation itself.

The paper therefore rejects a common assumption that queue scaling is unconditionally beneficial. Queues increase submission capacity, but they also increase completion work and may consume NIC resources.

Clock dependence and the proxy trade-off

GPU submission latency depends strongly on SM frequency.

Figure 3

Figure 3: Issue latency increases as the locked SM clock decreases, demonstrating that GPU communication control work is compute-bound rather than purely network-bound.

Across six SM-clock settings, the issue latency of NVSHMEM public operations is well modeled by a clock-sensitive component plus a clock-insensitive intercept. The clock-sensitive term accounts for approximately 80% of NVSHMEM’s public issue time and 85% of GDAKI’s. NVSHMEM’s effective clock coefficient is 3.5 times larger than GDAKI’s, explaining much of their issue-time difference.

This finding complicates comparisons based on nominal GPU or NIC specifications. A throttled GPU communicates more slowly even when the network and remote endpoint are unchanged. A proxy does not eliminate this dependence: the GPU still performs the enqueue operation. For IBRC, 92% of enqueue time and 27% of put-plus-completion time are attributed to the SM clock.

CPU proxies introduce a corresponding host-side dependency. On P-IB, the tuned proxy’s p99 round-trip latency improves from 8.9 to 5.6 microseconds after CPU warm-up, while IBRC’s small-message rate doubles from 1.9 to 3.8 million messages per second. Proxy measurements that omit worker placement, CPU frequency state, or NUMA locality are therefore not reproducible performance characterizations.

Queue sharing and interference under load

Idle latency is a poor predictor of behavior under contention. The paper evaluates request–reply latency while background CTAs generate traffic through the same endpoint.

Figure 4

Figure 4: Sharing a queue, FIFO, or context with bulk traffic produces severe latency inflation; reserving a probe queue restores substantially lower latency.

A shared proxy FIFO can reach tens of milliseconds under load, despite carrying only a few million messages per second. IBRC’s round trip increases by approximately 670 times with one background CTA in the tested configuration. The mechanism is head-of-line interference: latency-sensitive requests wait behind bulk requests in the same software ring or NIC queue.

GPU submission is not inherently isolated. A GDAKI probe sharing a context with background traffic reaches 362 microseconds at 64 CTAs, whereas a private context reduces p50 latency to 12–16 microseconds while the background traffic carries 73–79 million messages per second. Reservation is therefore more important than whether submission occurs on the GPU or CPU.

A private ring reduces loaded proxy latency by one to two orders of magnitude. However, isolation consumes capacity. The paper distinguishes queue isolation from worker isolation: reserving a software ring may provide most of the benefit, while dedicating an additional CPU worker can reduce bulk capacity without proportionate latency improvement.

Throughput: batching, cooperation, and queue parallelism

The peak small-message rate is governed by the number of independent submission streams and the amortization of doorbell operations. One mini-gda thread posts approximately 1.79 million 8-byte messages per second with one WQE per doorbell. Batching 16 WQEs per doorbell raises this to 4.70 million messages per second.

Cooperative publication is more effective than simply adding threads to one QP. A warp that reserves and publishes WQEs collectively reaches 21.6 million messages per second on one QP, compared with approximately 1–3 million messages per second for independently publishing threads. With 64 QPs, 64 submitting threads, and 16 WQEs per doorbell, mini-gda reaches 260 million messages per second.

The paper makes another important negative result: GPU doorbell submission is not required to reach the peak NIC rate. A host thread ringing doorbells for GPU-built WQEs reaches 257–258 million messages per second. GPU submission reduces CPU involvement and can reduce control-path latency, but it is not the necessary condition for saturating the tested NIC.

This distinction matters for system design. If the objective is peak throughput rather than autonomous execution or low issue latency, a CPU-assisted doorbell path may achieve essentially the same NIC rate. Conversely, if the CPU is unavailable or the application requires persistent GPU-side control flow, GPU submission remains operationally valuable even when it does not improve the hardware throughput ceiling.

Proxy capacity and payload crossover

Proxy capacity is highly sensitive to batching and worker parallelism.

Figure 5

Figure 5: Proxy message rate increases with worker count and request chaining, but the achievable rate depends strongly on platform and CPU operating state.

On P-IB, increasing mini-proxy’s batch limit from 1 to 16 raises one-worker throughput by 5.5 times. Nevertheless, proxy implementations remain approximately an order of magnitude below NVSHMEM IBGDA for 8-byte messages on this platform. On P-GB200, where CPU and GPU share coherent NVLink-C2C memory, an eight-worker mini-proxy reaches 140 million messages per second, approximately 90% of IBGDA’s 156 million messages per second on the tested pair.

The GB200 result is not a universal proxy advantage. The systems differ in CPU, NIC, link, firmware, and provider configuration. Moreover, coherent memory improves proxy capacity but does not improve tuned-proxy latency relative to P-IB: the measured round trip is 7.5 microseconds on P-GB200 versus 5.9 microseconds on P-IB.

As payloads grow, posting-rate differences become less important.

Figure 6

Figure 6: Most paths reach 90% of the 24.8 GB/s reference at payload sizes ranging from 512 bytes to 7 KiB; the crossover depends on small-message posting capacity.

mini-gda and NVSHMEM reach 90% of the 24.8 GB/s reference at 512 bytes. GDAKI, mini-proxy, and UCCL-EP require approximately 2 KiB, while MSCCL++, GIN Proxy, and fabric-lib require approximately 7 KiB. At 7,168 bytes—the FP8 hidden-state vector size used for the DeepSeek-V3 token representation—all evaluated paths except cold IBRC exceed 90% of the reference bandwidth. IBRC reaches that point at 32 KiB when cold and 7 KiB when warm.

Thus, a proxy that performs poorly on 8-byte messages may have little disadvantage for realistic token-vector transfers. Goodput convergence does not imply latency convergence: completion latency remains distinct from bandwidth saturation.

Cost to the enclosing GPU kernel

The communication path consumes GPU resources even when it is not exercised. Compiling communication code into a kernel can increase register pressure, reduce block residency, and induce spilling.

Figure 7

Figure 7: Communication code can reduce useful-work throughput even when dormant, and completion waits impose additional path-dependent costs.

For a compute-oriented caller with eight live values, inlined NVSHMEM and GDAKI reduce throughput by 11–15% before any communication occurs. For a streaming caller with 16 live values, NVSHMEM and GDAKI reduce useful-work throughput by approximately 37% in the dormant-code condition. A separate configuration shows that even dormant mini-gda can reduce streaming throughput by 27% through occupancy effects.

Waiting is more expensive than asynchronous submission. For 16 individual 8-byte messages per block, waiting after each message causes an additional 2–5% loss for mini-gda and GDAKI but 19–27% for NVSHMEM and mini-proxy. With NVSHMEM, increasing the number of live RC QPs makes the waiting cost grow sharply: for a caller with 16 live values, the loss rises to 7%, 20%, and 64% at 16, 64, and 528 QPs, respectively.

The application-level consequence is that optimizing issue latency alone may have limited value. In the DeepEP experiment, dispatch finishes issuing after approximately 25 microseconds of a 0.49-millisecond operation; most of the execution time is spent on receive-side waiting and copying. The dominant optimization target is therefore workload-specific synchronization and data movement, not necessarily WQE posting.

NIC connection state and all-to-all scaling

The final scaling study moves below the library and GPU kernel to NIC resource behavior. RC maintains per-connection state in NIC interconnect context memory, while DC reduces persistent initiator state by addressing multiple targets through dynamically connected initiators.

Figure 8

Figure 8: NIC message rate depends on traffic direction, connection reuse, payload size, and the number of active peer contexts; DC reduces some state pressure but does not eliminate the scaling knee.

Send-only traffic remains near 242 million messages per second through 32,768 active QPs. Simultaneous send-and-receive traffic begins at a lower rate and falls from 152 to 74 million messages per second over the same active-set range. Receive-only traffic degrades at a larger active-set size than bidirectional traffic, demonstrating that connection count alone is insufficient to predict cost.

Connection reuse mitigates the decline. Reusing a connection for 8,192 writes per visit keeps bidirectional throughput near 139–143 million messages per second across 1,024–4,096 QPs. Larger payloads also hide the message-rate penalty: with 4,096 active QPs, send-plus-receive goodput is only 31% of send-only goodput at 8 bytes, 55% at 128 bytes, but 98% at 256 bytes and above.

The all-to-all result is particularly strong: the tested 32-PE sweep loses 59% of its two-peer rate at approximately 3,000 active connections. The fitted onset for dense 128-PE traffic occurs near 1,350 active connections per NIC. DC improves the intermediate region—retaining 95% of the two-peer rate near 1,984 DCI–peer pairs compared with 56% for RC—but retains only 23% versus RC’s 23% at approximately 2,976 pairs in the reported configuration. The paper concludes that DC does not eliminate NIC-side contention or cache pressure.

A single DC operation costs approximately 0.45 microseconds more than RC through completion in mini-gda. Conversely, changing DC destinations on every WQE can reduce NVSHMEM’s message rate by 60 times, with short per-destination bursts recovering much of the loss. Connection reuse is therefore a first-class transport parameter.

Limitations and open questions

The evaluation is technically broad but concentrated on NVIDIA GPUs, ConnectX-7 NICs, and selected InfiniBand or RoCE configurations. The authors explicitly note that NIC firmware, RDMA providers, CPU state, link technology, and platform topology differ across the systems, particularly in the GB200 comparison. Consequently, the reported absolute latencies and message-rate ceilings should not be treated as architecture-independent constants.

The unsafe no-fence configuration also remains unresolved. It establishes a useful lower bound but does not establish correctness under arbitrary queue placement, compiler behavior, memory hierarchy state, or persistent-kernel execution. The active-connection study identifies a scaling knee but cannot distinguish among NIC context-cache misses, packet-processing contention, backpressure, and other internal mechanisms. Finally, the evaluation largely uses point-to-point writes and controlled microbenchmarks. Application-level collectives, MoE dispatch/combine pipelines, receive-side processing, and synchronization protocols may shift the relative importance of issue latency, completion scope, and queue isolation.

The main open technical question is how to design communication libraries that expose the throughput benefits of queue parallelism without imposing PE-wide completion costs, register-residency losses, and NIC connection-state pressure. A second is whether portable proxy architectures can preserve the tuned single-operation behavior observed here without requiring workload-specific ring ownership and dedicated CPU cores.

Conclusion

The paper establishes that GPU-initiated communication is not characterized by a single intrinsic latency or throughput. A minimal GPU path issues an 8-byte write in 0.70 microseconds and completes it in approximately 4.0 microseconds, while production libraries add as much as 4.6 microseconds of issue overhead through queue handling, fences, and completion scope. A tuned CPU proxy can match or beat GPU submission at idle, but requires dedicated and carefully conditioned CPU resources.

At high message rates, batching, cooperative publication, and independent queues are decisive: the tested platform reaches approximately 260 million messages per second, and host-submitted doorbells can reach nearly the same ceiling. Those optimizations carry costs in queue-scanning latency, GPU occupancy, CPU capacity, and NIC connection state. Under load, queue isolation is essential for both GPU and proxy paths; at scale, bidirectional all-to-all traffic loses 59% of its message rate near 3,000 active connections.

The paper’s principal contribution is therefore methodological as much as numerical: GPU communication performance must be decomposed into submission, ordering, completion, kernel-resource, and NIC-state costs. Library comparisons that do not control these dimensions cannot identify whether an observed advantage belongs to GPU initiation, proxy design, queue topology, or synchronization policy.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Explain it Like I'm 14

1. What is this paper about?

This paper studies how GPUs send small messages directly through a network card instead of asking the CPU to do the work.

This matters in modern AI systems, especially Mixture-of-Experts (MoE) models. In these models, different parts of a neural network called “experts” may run on different GPUs. The GPUs must quickly send data to the right experts. If communication is slow, the whole AI system has to wait.

The researchers compare two ways to send this data:

  • GPU-submitted communication: GPU threads create and send network requests themselves.
  • CPU-proxy communication: The GPU asks a CPU thread to send the request.

They also compare real communication libraries such as NVSHMEM, NCCL GIN, DeepEP, UCCL-EP, and MSCCL++.

2. What questions did the researchers ask?

The paper focuses on three main questions:

  1. How long does one communication operation take? The researchers wanted to know the basic cost of sending a very small message, even before adding extra library features.
  2. When is GPU communication better than using a CPU proxy? They studied which method is faster under different conditions, such as heavy traffic or many messages being sent at once.
  3. How can systems send huge numbers of messages, and what does that require? They examined techniques such as using multiple queues and combining several messages into one notification. They also measured the costs of these techniques in GPU and network resources.

3. How did they perform the research?

Building simplified communication systems

The researchers created two small experimental systems:

  • mini-gda: a simple GPU-submitted communication path.
  • mini-proxy: a simple CPU-proxy communication path.

These systems included only the essential steps needed to send data. This helped the researchers separate the cost of the basic hardware process from the extra work performed by large communication libraries.

They also tested several production libraries used in real AI systems.

How a GPU sends a network message

The process is similar to ordering food from a restaurant:

  1. The GPU writes down the order in a special queue.
  2. It tells the network card that a new order is ready. This notification is called a doorbell.
  3. The network card reads the request and sends the data.
  4. The network card writes a completion message when the operation finishes.
  5. The GPU checks for that completion message.

In the paper, the request placed in the queue is called a Work Queue Element, or WQE. A queue is like a line of tasks waiting to be handled.

The researchers measured how much time each part takes, including:

  • Building the request
  • Managing queues
  • Making sure data is written in the correct order
  • Notifying the network card
  • Checking whether the transfer is complete

Testing different hardware and settings

The experiments used NVIDIA H100, H200, B200, and GB200 systems with ConnectX-7 network cards.

They tested:

  • Very small 8-byte messages
  • Larger messages
  • Different numbers of queues
  • Different numbers of GPU threads and CPU workers
  • Message batching, where several messages are announced together
  • Light and heavy network traffic
  • Different memory-ordering settings
  • Large numbers of active connections

The researchers measured both latency, meaning how long one message takes, and throughput, meaning how many messages can be sent each second.

4. What did they find?

The simplest GPU path is very fast

The basic GPU system, mini-gda, could begin sending an 8-byte message in about 0.7 microseconds and finish the operation in about 4 microseconds.

A microsecond is one-millionth of a second, so this is extremely fast.

However, real libraries added extra time. For example:

  • The simple GPU path took about 0.7 microseconds to issue a message.
  • NCCL GIN took about 1.6 microseconds.
  • NVSHMEM’s public interface took about 5.3 microseconds.

The extra time came from tasks such as managing more queues, using safety checks, and checking completion across many queues.

Safety checks can make communication slower

Before the network card reads a request, the GPU must make sure the request has been fully written. This requires memory-ordering instructions, sometimes called fences.

A fence is like telling everyone in a group, “Do not move to the next step until the previous step is definitely finished.”

These safety checks are important for correctness, but stronger checks take longer. The paper found that a system-wide safety check could take about 3.7 times longer than a check limited to the GPU.

Removing these checks made the system faster, but it was unsafe in general. Therefore, high-performance systems must balance speed with correctness.

A carefully designed CPU proxy can compete with the GPU

The researchers found that a well-tuned CPU proxy could be nearly as fast as the simplest GPU path.

The tuned proxy had a median round-trip time of about 5.9 microseconds, compared with about 6.9 microseconds for the minimal GPU path.

However, the CPU proxy requires a dedicated CPU core. Its speed also depends on whether that core is running at a low or high clock speed.

This means that GPU communication is not always automatically better. A CPU proxy can work well when:

  • A CPU core is available
  • The communication design is carefully optimized
  • The system does not need every CPU core for other tasks

Sharing queues with large transfers can cause serious delays

Small, urgent messages and large bulk transfers do not always work well when they share the same queue.

The researchers found that sharing a queue could increase latency by 10 to 1,000 times. In one case, a queue shared with heavy traffic reached delays of tens of milliseconds.

This suggests that small, time-sensitive messages should often use their own reserved queues.

Batching and multiple queues greatly increase message rate

Sending one message at a time wastes some of the network card’s ability. The system can improve speed by:

  • Combining several messages before ringing the doorbell
  • Using several queues in parallel
  • Having several GPU threads cooperate

The measured message rates were approximately:

Configuration Message rate
One submitting GPU thread 1.8 million messages/second
One cooperative GPU warp 21.6 million messages/second
Many independent GPU queues 260 million messages/second

However, these improvements use more resources. More queues require more memory and more network-card state.

GPU communication can reduce computing performance

Even if a communication feature is not actively being used, including its code in a GPU program can reduce the number of useful GPU blocks that run at the same time.

The paper found that communication code could reduce the useful performance of a GPU kernel by as much as 37%.

This is important because communication tools do not only affect networking. They can also affect the GPU’s ability to perform calculations.

Too many connections hurt all-to-all communication

When many GPUs communicate with many other GPUs, the network card must track many connections.

For send-only traffic, the system maintained a high message rate even with tens of thousands of active queues. But for all-to-all traffic, the message rate dropped by 59% at around 3,000 active connections.

Using a dynamic connection method did not completely solve this problem.

5. Why are these findings important?

The main lesson is that the speed of GPU communication depends on much more than whether the GPU or CPU sends the message.

Other choices matter greatly, including:

  • How many queues are used
  • Whether messages are combined into batches
  • How safety checks are performed
  • How many queues are checked for completion
  • Whether small messages share queues with large transfers
  • How many connections the network card must manage
  • Whether communication code uses valuable GPU resources

The paper also shows why comparing complete libraries can be misleading. One library may appear slower because it performs more safety checks or manages more queues, not because the underlying hardware is slower.

By providing mini-gda, mini-proxy, and open-source experiments, the researchers give other developers tools for testing these details separately.

Conclusion

This research helps explain what happens at the boundary between a GPU and a network card. It shows that GPU-initiated communication can be extremely fast, but it is not automatically the best choice in every situation.

A simple GPU path can send a message very quickly, while a carefully designed CPU proxy can sometimes match or beat it. The fastest systems use batching and many queues, but those methods consume extra resources and may reduce GPU performance or overload the network card.

For future AI systems, especially large MoE models, the results suggest that communication software should be designed carefully rather than relying on one universal method. Developers may need to reserve separate queues for urgent messages, control how completion is checked, and choose between GPU and CPU submission based on the workload.

Knowledge Gaps

Knowledge gaps, limitations, and open questions

  • The evaluation is dominated by NVIDIA GPUs, ConnectX-7 NICs, and specific CUDA, driver, and library versions; the extent to which the results generalize to AMD, Intel, other NVIDIA architectures, alternative NICs, and future driver stacks remains unresolved.
  • The paper does not systematically compare GPU-initiated communication across PCIe generations, PCIe topologies, NUMA placements, or NICs attached through different switches; these factors could substantially alter doorbell, WQE-fetch, and completion latency.
  • The security, stability, and deployment implications of mapping NIC UAR/BAR pages into GPU address space are identified but not experimentally evaluated, including isolation between tenants, fault containment, privilege requirements, and behavior under malformed or stale doorbell accesses.
  • The safety of the “unsafe” fence-free ordering configuration remains unproven. The experiments report no corruption, but do not establish correctness across GPU architectures, memory pressures, concurrent writers, queue wraparound, resets, or long-running workloads.
  • The paper does not provide a formal memory-model argument or stress-test methodology that could determine when GPU-scope ordering is sufficient and when system-scope ordering is required.
  • The measured completion delay is not fully decomposed into PCIe or NVLink-C2C traversal, NIC processing, network propagation, remote GPU memory access, CQE placement, and polling overhead; the individual contribution of each component remains uncertain.
  • The study focuses primarily on RDMA writes and 8-byte control operations. Other operations—including reads, atomics, sends, receives, signaling operations, and larger or scattered payloads—may have different WQE, ordering, completion, and contention costs.
  • The experiments do not systematically characterize bidirectional traffic, asymmetric sender/receiver rates, incast, multicast, or overlapping reads and writes, leaving the behavior of GPU-initiated communication under more diverse traffic patterns unexplored.
  • The relationship between queue count, QP state, NIC cache capacity, memory translation caches, and message-rate collapse is observed but not modeled or mapped precisely enough to predict performance for arbitrary configurations.
  • The reported approximately 3,000-connection degradation under all-to-all traffic is specific to the tested platform and workload; the connection-count thresholds and causes of degradation on other NIC generations, transport configurations, and traffic matrices remain open questions.
  • Dynamically Connected transport does not eliminate the observed all-to-all degradation, but the paper does not isolate whether the remaining bottleneck arises from DCT/DCI state, address-vector processing, NIC scheduling, memory translation, or another resource.
  • The study does not evaluate how queue and connection scaling behaves across multiple NICs per node, multiple GPUs sharing one NIC, or oversubscribed NIC-to-GPU topologies.
  • The comparison between GPU submission and CPU proxies is sensitive to CPU worker clock state, but the paper does not quantify the effects of core type, frequency governors, turbo variability, thermal throttling, interrupt interference, SMT, or power limits in a reproducible performance model.
  • The resource cost of CPU proxies is reported mainly as the use of dedicated cores; their total energy consumption, power efficiency, thermal impact, and cost per sustained message rate are not evaluated.
  • GPU resource overhead is measured through reduced block residency, but the paper does not characterize register usage, shared-memory usage, instruction-cache pressure, scheduler effects, occupancy across kernel shapes, or interference with compute-intensive and memory-intensive kernels.
  • The reported maximum 37% reduction in useful kernel throughput is not linked to a broad workload taxonomy, so it remains unclear which application characteristics make communication code most harmful when dormant.
  • The experiments use pinned pairs of nodes and mostly p50 latency over three process runs; tail latency, run-to-run variance, confidence intervals, and sensitivity to system noise are insufficiently characterized.
  • The impact of transient congestion, adaptive routing, link-level flow control, packet loss, retransmissions, congestion control, and network faults on GPU-side completion latency is not studied.
  • The evaluation does not establish whether the measured benefits persist in realistic multi-tenant clusters where other jobs generate traffic, consume NIC state, or compete for PCIe and memory bandwidth.
  • The paper compares production libraries, but differences in APIs, default semantics, completion scopes, batching policies, and workload assumptions make it difficult to determine which design is best for a common application-level objective.
  • The study does not evaluate end-to-end MoE training or inference throughput, token latency, expert-load imbalance, dispatch/combine overlap, or convergence impact; consequently, microbenchmark improvements are not directly translated into model-level benefits.
  • The relationship between the tested 7 KiB line-rate threshold and realistic MoE message distributions is not validated across different models, sequence lengths, batch sizes, routing capacities, sparsity levels, and expert-placement strategies.
  • The experiments do not investigate how communication mechanisms interact with computation/communication overlap, persistent kernels, CUDA graphs, stream priorities, cooperative groups, or synchronization with other GPU collectives.
  • Completion semantics are compared at the library level, but the paper does not determine which completion granularity provides the best trade-off among correctness, latency, GPU occupancy, polling cost, and application programmability.
  • The cost and scalability of polling are not isolated from the cost of issuing operations; alternative polling strategies such as backoff, warp specialization, interrupt-assisted completion, or NIC-generated GPU signals remain unexplored.
  • The paper does not assess error handling, recovery, and fault behavior for failed RDMA operations, QP errors, NIC resets, peer failure, stale CQEs, queue overflow, or GPU kernel termination.
  • The minimal implementations simplify queue management and notification behavior; their results may not capture the synchronization, progress, buffering, and error-handling costs required by production communication libraries.
  • The CPU-proxy comparison does not include a broad range of proxy architectures, such as kernel-bypass polling frameworks, shared-memory notification schemes, NIC offload engines, or proxies distributed across CPU sockets.
  • The study does not examine the effect of request-descriptor size, metadata richness, batching policies, and notification placement on proxy performance for operations more complex than the tested 16-byte descriptors.
  • The paper does not quantify the memory-capacity and allocation overheads of GPU-resident queues, completion queues, doorbell records, QP contexts, and registered buffers at the scale required by large MoE deployments.
  • The portability of the proposed mini-gda and mini-proxy abstractions to non-InfiniBand transports, Ethernet/RoCE congestion regimes, Ultra Ethernet, or vendor-neutral APIs remains unvalidated.
  • The experiments use RoCEv2 only for selected controls; a comprehensive comparison between InfiniBand and RoCE under congestion, loss, routing, and large-scale deployment conditions is absent.
  • The paper does not investigate how changes in message ordering, remote memory registration, GPU memory type, cache state, or address alignment affect the measured WQE and completion costs.
  • The long-term maintainability and API implications of exposing low-level ordering scopes, queue placement, completion scopes, and batching controls to library or application developers are not addressed.
  • No predictive performance model is provided that can select GPU submission, CPU proxying, queue count, batching, transport type, or completion scope from workload and hardware parameters.
  • The open-source release is mentioned, but the reproducibility of the full evaluation remains uncertain because the paper does not specify whether all firmware, topology, BIOS, power-management, benchmark, plotting, and deployment configurations are available.

Practical Applications

Immediate Applications

  • Optimize Mixture-of-Experts (MoE) training and inference pipelines — AI infrastructure / cloud computing.
    • batching several work requests per doorbell when latency permits;
    • assigning separate queues to latency-critical expert traffic and bulk transfers;
    • limiting completion operations to the relevant peer, queue, or context rather than polling all queues;
    • selecting inline payloads only for very small messages, approximately up to the measured break-even region of roughly 92 bytes;
    • increasing queue parallelism when message rate, rather than single-message latency, is the bottleneck.
    • Dependencies: Benefits assume NVIDIA GPU/NIC combinations and RDMA configurations similar to the evaluated H100, H200, B200, GB200, and ConnectX-7 systems. The optimal parameters remain workload- and topology-dependent.
  • Choose between GPU submission and CPU-proxy submission on a per-workload basis — distributed systems / HPC.
    • a CPU proxy handles sparse control messages and synchronization;
    • GPU submission handles high-rate expert dispatch or fine-grained all-to-all traffic;
    • the runtime switches modes according to load and queue occupancy.
    • Dependencies: Proxy performance depends strongly on CPU frequency state, worker placement, ring design, and whether the queue is isolated from bulk traffic.
  • Build communication-aware GPU kernel launch configurations — GPU programming / compiler systems.
    • reserve communication resources only for kernels that actually communicate;
    • vary block size and occupancy when persistent communication code is present;
    • place communication warps or cooperative publication groups explicitly;
    • compare kernels with and without communication support during profiling.
    • Dependencies: The residency penalty depends on register, shared-memory, queue, and occupancy requirements, so the reported percentage should not be treated as universal.
  • Separate latency-critical and bulk RDMA traffic — networking / storage / HPC.
    • MoE token routing;
    • parameter-server control traffic;
    • distributed checkpoint coordination;
    • GPU storage and remote-memory systems;
    • MPI-like control-plane messages.
    • Dependencies: Queue isolation consumes NIC memory, queue-pair state, completion resources, and potentially CPU workers.
  • Use mechanism-level microbenchmarks when evaluating communication libraries — academia / engineering benchmarking.
    • hardware submission cost;
    • queue-management overhead;
    • memory-ordering cost;
    • completion-scope cost;
    • library API overhead;
    • network and NIC congestion.
    • This produces more meaningful comparisons than comparing complete libraries using different queue counts, batching policies, or completion semantics.
    • Dependencies: Reproducibility requires matching driver, CUDA, NIC firmware, GPU clock, CPU operating state, topology, and transport configuration.
  • Tune completion scope in existing GPU communication software — software libraries. Library maintainers can replace broad “complete everything” operations with targeted completion mechanisms. For example, an operation involving one peer or one queue should avoid polling all configured queues. This can reduce latency in NVSHMEM-like APIs and improve fine-grained synchronization in persistent kernels. Dependencies: Narrow completion scopes are safe only when the application’s dependency graph is correctly represented; overly weak completion can expose stale payload data under GPU memory-ordering rules.
  • Improve operator and cluster-level diagnostics — cloud operations / performance engineering.
    • GPU SM clock and throttling state;
    • CPU proxy frequency and residency state;
    • queue depth and queue sharing;
    • doorbell batch size;
    • active connection count;
    • completion latency;
    • NIC message rate.
    • These metrics can help distinguish software overhead from PCIe, NVLink-C2C, NIC, or congestion effects.
    • Dependencies: Instrumentation must avoid perturbing the fine-grained timing being measured.
  • Guide hardware procurement and topology placement — data centers / HPC procurement. The results support placing GPUs and NICs on the same PCIe switch where possible, avoiding topologies that add indirect paths. They also indicate that coherent GB200-style systems may have different completion behavior from PCIe-attached systems, even when issue latency is similar. Dependencies: Procurement decisions must consider link bandwidth, memory coherency, NIC generation, driver support, and application traffic patterns rather than relying on peak bandwidth alone.
  • Teach GPU networking and RDMA mechanisms using minimal reproducible examples — education / academia. Courses and research laboratories can use the paper’s minimal transports to demonstrate WQE construction, queue pairs, doorbell ordering, GPU memory registration, completion queues, and proxy submission. This is more actionable than treating NVSHMEM or NCCL as opaque APIs. Dependencies: Direct UAR mapping and GPU-resident queue experiments may require privileged driver settings and specialized hardware.

Long-Term Applications

  • Adaptive communication runtimes that select submission paths dynamically — AI systems / distributed runtimes.
    • route isolated messages to a warm CPU proxy;
    • switch to GPU submission at high concurrency;
    • create dedicated queues for bursts;
    • adjust doorbell batching according to latency deadlines;
    • fall back to host-assisted submission when GPU doorbell mapping is unavailable.
    • Dependencies: This requires online performance models, low-overhead mode switching, reliable congestion signals, and correctness-preserving synchronization across both paths.
  • Compiler and DSL support for communication-aware kernels — programming languages / compilers.
    • fence scope;
    • cooperative warp publication;
    • WQE layout;
    • batching thresholds;
    • queue assignment;
    • completion scope;
    • communication-warp placement.
    • A domain-specific language could allow programmers to express dependencies such as “publish this payload before notifying peer X,” leaving safe ordering and queue selection to the compiler.
    • Dependencies: Compiler transformations must preserve GPU and NIC memory-ordering semantics; unsafe removal of fences cannot be generalized from the paper’s controlled experiments.
  • Scalable all-to-all communication for very large MoE clusters — AI infrastructure / supercomputing.
    • hierarchical expert routing;
    • topology-aware token placement;
    • connection pooling;
    • hierarchical aggregation;
    • sparse or selective all-to-all;
    • improved dynamic-connection caching;
    • transport protocols with less per-peer NIC state.
    • Dependencies: Dynamic Connected transport alone may not eliminate NIC cache and connection-state pressure. Solutions must preserve load balance and avoid increasing token-routing latency.
  • New NIC and interconnect designs optimized for GPU-originated traffic — hardware / networking.
    • lower-cost GPU-visible doorbells;
    • larger or more efficient queue caches;
    • native GPU completion primitives;
    • reduced per-connection state;
    • better support for many-to-many traffic;
    • hardware batching and aggregation;
    • secure first-class GPU-to-NIC mappings.
    • Such features could reduce the gap between the measured hardware floor and production-library latency.
    • Dependencies: Security, isolation, virtualization, PCIe/IOMMU behavior, and compatibility with different GPU vendors must be addressed.
  • Secure virtualization of GPU-to-NIC doorbell access — cloud platforms / policy and infrastructure security. Because mapping a NIC PCIe BAR or UAR into GPU address space can create a security risk, future cloud platforms could provide mediated doorbells, hardware capabilities, or hypervisor-controlled submission channels. This would allow tenant workloads to use GPU-initiated networking without requiring broad peer-mapping overrides. Dependencies: The mechanism must prevent tenants from accessing other queues, injecting unauthorized traffic, or bypassing isolation. Security controls must not eliminate the latency advantage of direct submission.
  • Portable GPU-initiated networking across vendors and interconnects — software ecosystems. The paper’s distinction between GPU-built work requests, GPU-submitted doorbells, and CPU-proxy submission could inform a portable abstraction spanning NVIDIA, AMD, different NIC vendors, InfiniBand, RoCE, and future Ultra Ethernet systems. A runtime could expose capabilities rather than assuming that every platform supports direct UAR mapping. Dependencies: Portability may require a lowest-common-denominator API, which could sacrifice some hardware-specific performance. Vendor cooperation and standardized memory-ordering semantics would be important.
  • Communication-aware scheduling for energy-efficient data centers — energy / cloud operations. Schedulers could consolidate low-rate communication onto CPU proxies, power down unused communication workers, and activate GPU submission or additional queues only during high-throughput phases. CPU frequency and GPU SM-clock findings also suggest that energy-performance policies should include communication latency, not only compute utilization. Dependencies: Power-state transitions may be slower than the communication intervals being optimized. Dedicated proxy cores can also increase energy use when the workload is sparse.
  • Formal verification and safer APIs for GPU–NIC memory ordering — systems research / safety-critical infrastructure. The paper identifies correctness hazards when a GPU observes a completion flag without a guaranteed view of the associated payload. Future libraries could provide typed completion tokens, formally specified fence scopes, and static or runtime checks preventing unsafe polling patterns in persistent kernels. Dependencies: Verification must cover GPU caches, NIC DMA ordering, PCIe or NVLink paths, remote visibility, and interactions among multiple queues.
  • Autotuned communication products for managed AI clusters — cloud software / MLOps.
    • queue counts;
    • batch sizes;
    • inline thresholds;
    • CPU worker counts;
    • proxy versus GPU submission;
    • completion scope;
    • active-connection limits.
    • These policies could be attached to model-serving profiles and automatically updated after driver, firmware, GPU, or NIC changes.
    • Dependencies: Autotuning must use representative traffic, including burstiness and all-to-all patterns; isolated microbenchmarks alone do not predict end-to-end model performance.
  • Faster interactive AI services and user-facing applications — daily life / consumer services. If the above optimizations scale reliably, they could reduce tail latency for applications that depend on distributed MoE inference, including conversational assistants, translation, recommendation, code generation, and real-time multimodal services. The direct user benefit would be faster responses and better throughput per cluster. Dependencies: The paper measures communication mechanisms rather than end-to-end user latency. Gains will depend on model architecture, compute time, batching policy, congestion, and service-level scheduling.

Glossary

  • Address vector (AV): A descriptor identifying the destination of a dynamically connected network operation. “DC transports additionally include an address-vector segment to identify the dynamically connected target (DCT).”
  • All-to-all traffic: Communication in which every participant sends data to every other participant. “all-to-all traffic loses 59\% of its rate near 3,000 connections”
  • BlueFlame: A NIC mechanism that allows userspace or device code to place work-request data directly into a doorbell buffer. “The doorbell targets the UAR's send-doorbell/BlueFlame region.”
  • Completion queue (CQ): A queue containing records that indicate the completion status of submitted network operations. “Send and receive completions are directed to completion queues (CQs), which can be distinct or shared with other QPs.”
  • Completion queue entry (CQE): A record written by the NIC to report that a signaled work request has completed. “The NIC reports a finished signaled WQE by writing a completion queue entry (CQE)”
  • Completion scope: The set of queues, peers, or operations covered by a completion or synchronization operation. “We measure the cost of the completion scope in Section~\ref{sec:micro:completion}.”
  • Coherent memory: Memory that can be accessed by CPUs and GPUs while maintaining a consistent view of data. “P-GB200 is a Grace--Blackwell system with coherent memory between the CPU and the GPU”
  • CPU proxy: A CPU thread or worker that receives GPU-generated requests and submits the corresponding network operations. “A tuned CPU proxy matches or beats the GPU path at idle”
  • CTA: A CUDA thread block, or cooperative thread array, executed as a unit on a GPU. “RC QPs may be allocated per GPU, CTA, or warp”
  • DMA-BUF: A Linux kernel mechanism for sharing memory buffers between device drivers, including allowing NIC access to GPU memory. “DMA-BUF registration lets the NIC access GPU memory but does not by itself establish this reverse mapping”
  • Doorbell batching: Combining multiple work requests into a single notification to the NIC. “Announcing several WQEs with one doorbell raises throughput”
  • Doorbell record: A memory location containing the index of the next available work-request block for the NIC. “First, the submitting thread updates the doorbell record (\code{dbrec}) with the index of the next free WQE basic block”
  • Dynamically Connected (DC) transport: An RDMA transport that allows a limited set of initiators to address many dynamically connected targets. “Dynamically Connected (DC) transport, by contrast, lets a pool of dynamically connected initiators (DCIs) address many remote DC targets (DCTs).”
  • Dynamically connected initiator (DCI): A transport endpoint that can issue operations to multiple dynamically connected targets. “DCI: dynamically connected initiator”
  • Dynamically connected target (DCT): An RDMA target endpoint that can receive operations from dynamically connected initiators. “Only DC WQEs carry the AV”
  • Expert parallelism: A distributed-machine-learning strategy that assigns different neural-network experts to different devices. “Mixture-of-Experts, expert parallelism”
  • Fence: A memory-ordering operation that ensures earlier memory accesses become visible before later accesses. “WQE stores must become visible before the doorbell announces them”
  • GPUDirect Async: A technology that lets GPU operations trigger or wait for communication initiated or prepared by the CPU. “GPUDirect Async shifts synchronization onto the device”
  • GPUDirect RDMA: A technology allowing a NIC to access GPU memory directly without host-side staging copies. “GPUDirect RDMA exposes GPU memory over PCIe for direct NIC access”
  • GDRCopy: A mechanism providing low-latency CPU load/store access to GPU memory. “GDRCopy, which gives the CPU low-latency load/store access to GPU memory”
  • GPU-initiated communication: Communication in which GPU threads construct and submit network operations without CPU control of the submission path. “GPU-initiated communication lets GPU threads post RDMA operations directly to the NIC.”
  • IBGDA: NVIDIA’s InfiniBand GPU Direct Async implementation, which enables kernels to submit RDMA work directly. “Modern GPU-initiated transports like IBGDA”
  • Inline payload: Data placed directly inside a work-request descriptor rather than referenced through a separate memory pointer. “An inline payload replaces the data segment's pointer with a length word and the payload.”
  • InfiniBand: A high-performance networking technology commonly used for data-center and supercomputer communication. “over InfiniBand and RoCEv2”
  • Memory ordering: Rules governing the visibility and relative order of memory operations across processors and devices. “A library-level comparison obscures whether the performance differences come from queue counts, batching strategies, memory fence strengths, or other API overheads”
  • Mixture-of-Experts (MoE): A neural-network architecture that routes each input token to a selected subset of specialized expert modules. “In Mixture-of-Experts~(MoE) models, GPU kernels decide which experts each token is sent to”
  • MMIO: Memory-mapped input/output, a technique for controlling hardware registers through memory operations. “Inline copying inflates the MMIO write payload of the doorbell”
  • NIC: A network interface controller that processes and transmits network traffic. “GPU threads post RDMA operations directly to the NIC.”
  • NVLink-C2C: A high-bandwidth interconnect connecting a CPU and GPU, particularly in Grace–Blackwell systems. “host rings and counters used by proxies travel over NVLink-C2C rather than PCIe.”
  • PCIe BAR: A memory-address range exposed by a PCI Express device for accessing device registers or memory. “It is a hardware register inside a User Access Region (UAR), which is a slice of the NIC's PCIe BAR”
  • Persistent kernel: A GPU kernel that remains resident and repeatedly performs work instead of terminating after one operation. “This is a common hazard for persistent kernels”
  • Processing element (PE): A logical computational participant or GPU rank involved in communication. “PE: processing element (one GPU rank)”
  • Queue pair (QP): An RDMA communication object consisting of a send queue and a receive queue. “RDMA communication is organized around Queue Pairs (QPs)”
  • RDMA: Remote direct memory access, allowing one machine or device to read or write another’s memory without involving its CPU in the data movement. “GPUDirect RDMA exposes GPU memory over PCIe for direct NIC access”
  • Reliable Connection (RC): An RDMA transport that associates a local queue pair with a specific remote queue pair and provides reliable delivery. “Reliable Connection (RC) transport binds each local QP to one remote QP at creation time.”
  • RoCEv2: RDMA over Converged Ethernet version 2, which carries RDMA traffic over routable Ethernet networks. “with ConnectX-7 NICs over InfiniBand and RoCEv2.”
  • Send Queue (SQ): The queue containing outbound RDMA work requests. “each consisting of a Send Queue (SQ) and a Receive Queue (RQ)”
  • SM clock: The operating frequency of a GPU streaming multiprocessor. “Issue latency scales strongly with the SM clock.”
  • Streaming multiprocessor (SM): A GPU execution unit that schedules and executes groups of threads. “communication code can reduce GPU block residency even when unused”
  • User Access Region (UAR): A NIC-mapped memory region through which userspace or device code writes doorbells. “It is a hardware register inside a User Access Region (UAR)”
  • Work Queue Element (WQE): A descriptor containing the information needed by a NIC to perform an RDMA operation. “both structured as circular rings of Work Queue Elements (WQEs).”
  • Work request: A software-level request describing a network operation to be submitted to an RDMA queue. “We first detail the GPU-side network path: queue placement, work-request construction, doorbell ordering, and completion semantics.”
  • Warp: A group of GPU threads executed together in a SIMD-style execution model. “Cooperative publication lets a warp reach 21.6\,M\,msg/s on one queue”

Open Problems

We found no open problems mentioned in this paper.

Tweets

Sign up for free to view the 2 tweets with 117 likes about this paper.

HackerNews