GPU-Initiated Communication: Dissecting Down to the Bone
Abstract: GPU-initiated communication lets GPU threads post RDMA operations directly to the NIC. It underpins NVSHMEM, NCCL GIN, and DeepEP, which serve the fine-grained, latency-critical communication of Mixture-of-Experts (MoE) models, yet its performance characteristics and optimizations remain scarcely documented beyond source code, and library comparisons fail to separate the costs of the hardware mechanism from those of the library around it. This paper dissects GPU-initiated communication at the GPU-NIC boundary. We first detail the GPU-side network path: queue placement, work-request construction, doorbell ordering, and completion semantics. We then introduce mini-gda and mini-proxy, minimal transports for the GPU and CPU-proxy submission paths, and measure them alongside NVSHMEM IBGDA, NCCL GIN, DeepEP, UCCL-EP, MSCCL++, and fabric-lib on NVIDIA H100, H200, B200, and GB200 platforms. A minimal GPU path issues an operation in 0.7 s and completes in 4.0 s; libraries add up to 4.6 s of issue time through queue management, memory ordering, and completion scope, and issue time scales with the SM clock. A tuned CPU proxy matches or beats the GPU path at idle, at the cost of a dedicated core whose operating state sets its latency and throughput. On either path, sharing a queue with bulk traffic raises latency by one to three orders of magnitude. Reaching the 260 M msg/s ceiling of our InfiniBand platform requires doorbell batching and queue parallelism, and both have resource costs: communication code can reduce GPU block residency even when unused, and all-to-all traffic loses 59% of its NIC message rate at about 3,000 active connections. The submission path alone therefore does not predict communication performance. Our experiment code and results are available at https://github.com/ParCoreLab/Dissecting-GPU-Communication-Experiments.
Paper Prompts
Sign up for free to create and run prompts on this paper.
Top Community Prompts
Explain it Like I'm 14
1. What is this paper about?
This paper studies how GPUs send small messages directly through a network card instead of asking the CPU to do the work.
This matters in modern AI systems, especially Mixture-of-Experts (MoE) models. In these models, different parts of a neural network called “experts” may run on different GPUs. The GPUs must quickly send data to the right experts. If communication is slow, the whole AI system has to wait.
The researchers compare two ways to send this data:
- GPU-submitted communication: GPU threads create and send network requests themselves.
- CPU-proxy communication: The GPU asks a CPU thread to send the request.
They also compare real communication libraries such as NVSHMEM, NCCL GIN, DeepEP, UCCL-EP, and MSCCL++.
2. What questions did the researchers ask?
The paper focuses on three main questions:
- How long does one communication operation take? The researchers wanted to know the basic cost of sending a very small message, even before adding extra library features.
- When is GPU communication better than using a CPU proxy? They studied which method is faster under different conditions, such as heavy traffic or many messages being sent at once.
- How can systems send huge numbers of messages, and what does that require? They examined techniques such as using multiple queues and combining several messages into one notification. They also measured the costs of these techniques in GPU and network resources.
3. How did they perform the research?
Building simplified communication systems
The researchers created two small experimental systems:
mini-gda: a simple GPU-submitted communication path.mini-proxy: a simple CPU-proxy communication path.
These systems included only the essential steps needed to send data. This helped the researchers separate the cost of the basic hardware process from the extra work performed by large communication libraries.
They also tested several production libraries used in real AI systems.
How a GPU sends a network message
The process is similar to ordering food from a restaurant:
- The GPU writes down the order in a special queue.
- It tells the network card that a new order is ready. This notification is called a doorbell.
- The network card reads the request and sends the data.
- The network card writes a completion message when the operation finishes.
- The GPU checks for that completion message.
In the paper, the request placed in the queue is called a Work Queue Element, or WQE. A queue is like a line of tasks waiting to be handled.
The researchers measured how much time each part takes, including:
- Building the request
- Managing queues
- Making sure data is written in the correct order
- Notifying the network card
- Checking whether the transfer is complete
Testing different hardware and settings
The experiments used NVIDIA H100, H200, B200, and GB200 systems with ConnectX-7 network cards.
They tested:
- Very small 8-byte messages
- Larger messages
- Different numbers of queues
- Different numbers of GPU threads and CPU workers
- Message batching, where several messages are announced together
- Light and heavy network traffic
- Different memory-ordering settings
- Large numbers of active connections
The researchers measured both latency, meaning how long one message takes, and throughput, meaning how many messages can be sent each second.
4. What did they find?
The simplest GPU path is very fast
The basic GPU system, mini-gda, could begin sending an 8-byte message in about 0.7 microseconds and finish the operation in about 4 microseconds.
A microsecond is one-millionth of a second, so this is extremely fast.
However, real libraries added extra time. For example:
- The simple GPU path took about 0.7 microseconds to issue a message.
- NCCL GIN took about 1.6 microseconds.
- NVSHMEM’s public interface took about 5.3 microseconds.
The extra time came from tasks such as managing more queues, using safety checks, and checking completion across many queues.
Safety checks can make communication slower
Before the network card reads a request, the GPU must make sure the request has been fully written. This requires memory-ordering instructions, sometimes called fences.
A fence is like telling everyone in a group, “Do not move to the next step until the previous step is definitely finished.”
These safety checks are important for correctness, but stronger checks take longer. The paper found that a system-wide safety check could take about 3.7 times longer than a check limited to the GPU.
Removing these checks made the system faster, but it was unsafe in general. Therefore, high-performance systems must balance speed with correctness.
A carefully designed CPU proxy can compete with the GPU
The researchers found that a well-tuned CPU proxy could be nearly as fast as the simplest GPU path.
The tuned proxy had a median round-trip time of about 5.9 microseconds, compared with about 6.9 microseconds for the minimal GPU path.
However, the CPU proxy requires a dedicated CPU core. Its speed also depends on whether that core is running at a low or high clock speed.
This means that GPU communication is not always automatically better. A CPU proxy can work well when:
- A CPU core is available
- The communication design is carefully optimized
- The system does not need every CPU core for other tasks
Sharing queues with large transfers can cause serious delays
Small, urgent messages and large bulk transfers do not always work well when they share the same queue.
The researchers found that sharing a queue could increase latency by 10 to 1,000 times. In one case, a queue shared with heavy traffic reached delays of tens of milliseconds.
This suggests that small, time-sensitive messages should often use their own reserved queues.
Batching and multiple queues greatly increase message rate
Sending one message at a time wastes some of the network card’s ability. The system can improve speed by:
- Combining several messages before ringing the doorbell
- Using several queues in parallel
- Having several GPU threads cooperate
The measured message rates were approximately:
| Configuration | Message rate |
|---|---|
| One submitting GPU thread | 1.8 million messages/second |
| One cooperative GPU warp | 21.6 million messages/second |
| Many independent GPU queues | 260 million messages/second |
However, these improvements use more resources. More queues require more memory and more network-card state.
GPU communication can reduce computing performance
Even if a communication feature is not actively being used, including its code in a GPU program can reduce the number of useful GPU blocks that run at the same time.
The paper found that communication code could reduce the useful performance of a GPU kernel by as much as 37%.
This is important because communication tools do not only affect networking. They can also affect the GPU’s ability to perform calculations.
Too many connections hurt all-to-all communication
When many GPUs communicate with many other GPUs, the network card must track many connections.
For send-only traffic, the system maintained a high message rate even with tens of thousands of active queues. But for all-to-all traffic, the message rate dropped by 59% at around 3,000 active connections.
Using a dynamic connection method did not completely solve this problem.
5. Why are these findings important?
The main lesson is that the speed of GPU communication depends on much more than whether the GPU or CPU sends the message.
Other choices matter greatly, including:
- How many queues are used
- Whether messages are combined into batches
- How safety checks are performed
- How many queues are checked for completion
- Whether small messages share queues with large transfers
- How many connections the network card must manage
- Whether communication code uses valuable GPU resources
The paper also shows why comparing complete libraries can be misleading. One library may appear slower because it performs more safety checks or manages more queues, not because the underlying hardware is slower.
By providing mini-gda, mini-proxy, and open-source experiments, the researchers give other developers tools for testing these details separately.
Conclusion
This research helps explain what happens at the boundary between a GPU and a network card. It shows that GPU-initiated communication can be extremely fast, but it is not automatically the best choice in every situation.
A simple GPU path can send a message very quickly, while a carefully designed CPU proxy can sometimes match or beat it. The fastest systems use batching and many queues, but those methods consume extra resources and may reduce GPU performance or overload the network card.
For future AI systems, especially large MoE models, the results suggest that communication software should be designed carefully rather than relying on one universal method. Developers may need to reserve separate queues for urgent messages, control how completion is checked, and choose between GPU and CPU submission based on the workload.
Knowledge Gaps
Knowledge gaps, limitations, and open questions
- The evaluation is dominated by NVIDIA GPUs, ConnectX-7 NICs, and specific CUDA, driver, and library versions; the extent to which the results generalize to AMD, Intel, other NVIDIA architectures, alternative NICs, and future driver stacks remains unresolved.
- The paper does not systematically compare GPU-initiated communication across PCIe generations, PCIe topologies, NUMA placements, or NICs attached through different switches; these factors could substantially alter doorbell, WQE-fetch, and completion latency.
- The security, stability, and deployment implications of mapping NIC UAR/BAR pages into GPU address space are identified but not experimentally evaluated, including isolation between tenants, fault containment, privilege requirements, and behavior under malformed or stale doorbell accesses.
- The safety of the “unsafe” fence-free ordering configuration remains unproven. The experiments report no corruption, but do not establish correctness across GPU architectures, memory pressures, concurrent writers, queue wraparound, resets, or long-running workloads.
- The paper does not provide a formal memory-model argument or stress-test methodology that could determine when GPU-scope ordering is sufficient and when system-scope ordering is required.
- The measured completion delay is not fully decomposed into PCIe or NVLink-C2C traversal, NIC processing, network propagation, remote GPU memory access, CQE placement, and polling overhead; the individual contribution of each component remains uncertain.
- The study focuses primarily on RDMA writes and 8-byte control operations. Other operations—including reads, atomics, sends, receives, signaling operations, and larger or scattered payloads—may have different WQE, ordering, completion, and contention costs.
- The experiments do not systematically characterize bidirectional traffic, asymmetric sender/receiver rates, incast, multicast, or overlapping reads and writes, leaving the behavior of GPU-initiated communication under more diverse traffic patterns unexplored.
- The relationship between queue count, QP state, NIC cache capacity, memory translation caches, and message-rate collapse is observed but not modeled or mapped precisely enough to predict performance for arbitrary configurations.
- The reported approximately 3,000-connection degradation under all-to-all traffic is specific to the tested platform and workload; the connection-count thresholds and causes of degradation on other NIC generations, transport configurations, and traffic matrices remain open questions.
- Dynamically Connected transport does not eliminate the observed all-to-all degradation, but the paper does not isolate whether the remaining bottleneck arises from DCT/DCI state, address-vector processing, NIC scheduling, memory translation, or another resource.
- The study does not evaluate how queue and connection scaling behaves across multiple NICs per node, multiple GPUs sharing one NIC, or oversubscribed NIC-to-GPU topologies.
- The comparison between GPU submission and CPU proxies is sensitive to CPU worker clock state, but the paper does not quantify the effects of core type, frequency governors, turbo variability, thermal throttling, interrupt interference, SMT, or power limits in a reproducible performance model.
- The resource cost of CPU proxies is reported mainly as the use of dedicated cores; their total energy consumption, power efficiency, thermal impact, and cost per sustained message rate are not evaluated.
- GPU resource overhead is measured through reduced block residency, but the paper does not characterize register usage, shared-memory usage, instruction-cache pressure, scheduler effects, occupancy across kernel shapes, or interference with compute-intensive and memory-intensive kernels.
- The reported maximum 37% reduction in useful kernel throughput is not linked to a broad workload taxonomy, so it remains unclear which application characteristics make communication code most harmful when dormant.
- The experiments use pinned pairs of nodes and mostly p50 latency over three process runs; tail latency, run-to-run variance, confidence intervals, and sensitivity to system noise are insufficiently characterized.
- The impact of transient congestion, adaptive routing, link-level flow control, packet loss, retransmissions, congestion control, and network faults on GPU-side completion latency is not studied.
- The evaluation does not establish whether the measured benefits persist in realistic multi-tenant clusters where other jobs generate traffic, consume NIC state, or compete for PCIe and memory bandwidth.
- The paper compares production libraries, but differences in APIs, default semantics, completion scopes, batching policies, and workload assumptions make it difficult to determine which design is best for a common application-level objective.
- The study does not evaluate end-to-end MoE training or inference throughput, token latency, expert-load imbalance, dispatch/combine overlap, or convergence impact; consequently, microbenchmark improvements are not directly translated into model-level benefits.
- The relationship between the tested 7 KiB line-rate threshold and realistic MoE message distributions is not validated across different models, sequence lengths, batch sizes, routing capacities, sparsity levels, and expert-placement strategies.
- The experiments do not investigate how communication mechanisms interact with computation/communication overlap, persistent kernels, CUDA graphs, stream priorities, cooperative groups, or synchronization with other GPU collectives.
- Completion semantics are compared at the library level, but the paper does not determine which completion granularity provides the best trade-off among correctness, latency, GPU occupancy, polling cost, and application programmability.
- The cost and scalability of polling are not isolated from the cost of issuing operations; alternative polling strategies such as backoff, warp specialization, interrupt-assisted completion, or NIC-generated GPU signals remain unexplored.
- The paper does not assess error handling, recovery, and fault behavior for failed RDMA operations, QP errors, NIC resets, peer failure, stale CQEs, queue overflow, or GPU kernel termination.
- The minimal implementations simplify queue management and notification behavior; their results may not capture the synchronization, progress, buffering, and error-handling costs required by production communication libraries.
- The CPU-proxy comparison does not include a broad range of proxy architectures, such as kernel-bypass polling frameworks, shared-memory notification schemes, NIC offload engines, or proxies distributed across CPU sockets.
- The study does not examine the effect of request-descriptor size, metadata richness, batching policies, and notification placement on proxy performance for operations more complex than the tested 16-byte descriptors.
- The paper does not quantify the memory-capacity and allocation overheads of GPU-resident queues, completion queues, doorbell records, QP contexts, and registered buffers at the scale required by large MoE deployments.
- The portability of the proposed mini-gda and mini-proxy abstractions to non-InfiniBand transports, Ethernet/RoCE congestion regimes, Ultra Ethernet, or vendor-neutral APIs remains unvalidated.
- The experiments use RoCEv2 only for selected controls; a comprehensive comparison between InfiniBand and RoCE under congestion, loss, routing, and large-scale deployment conditions is absent.
- The paper does not investigate how changes in message ordering, remote memory registration, GPU memory type, cache state, or address alignment affect the measured WQE and completion costs.
- The long-term maintainability and API implications of exposing low-level ordering scopes, queue placement, completion scopes, and batching controls to library or application developers are not addressed.
- No predictive performance model is provided that can select GPU submission, CPU proxying, queue count, batching, transport type, or completion scope from workload and hardware parameters.
- The open-source release is mentioned, but the reproducibility of the full evaluation remains uncertain because the paper does not specify whether all firmware, topology, BIOS, power-management, benchmark, plotting, and deployment configurations are available.
Practical Applications
Immediate Applications
- Optimize Mixture-of-Experts (MoE) training and inference pipelines — AI infrastructure / cloud computing.
- batching several work requests per doorbell when latency permits;
- assigning separate queues to latency-critical expert traffic and bulk transfers;
- limiting completion operations to the relevant peer, queue, or context rather than polling all queues;
- selecting inline payloads only for very small messages, approximately up to the measured break-even region of roughly 92 bytes;
- increasing queue parallelism when message rate, rather than single-message latency, is the bottleneck.
- Dependencies: Benefits assume NVIDIA GPU/NIC combinations and RDMA configurations similar to the evaluated H100, H200, B200, GB200, and ConnectX-7 systems. The optimal parameters remain workload- and topology-dependent.
- Choose between GPU submission and CPU-proxy submission on a per-workload basis — distributed systems / HPC.
- a CPU proxy handles sparse control messages and synchronization;
- GPU submission handles high-rate expert dispatch or fine-grained all-to-all traffic;
- the runtime switches modes according to load and queue occupancy.
- Dependencies: Proxy performance depends strongly on CPU frequency state, worker placement, ring design, and whether the queue is isolated from bulk traffic.
- Build communication-aware GPU kernel launch configurations — GPU programming / compiler systems.
- reserve communication resources only for kernels that actually communicate;
- vary block size and occupancy when persistent communication code is present;
- place communication warps or cooperative publication groups explicitly;
- compare kernels with and without communication support during profiling.
- Dependencies: The residency penalty depends on register, shared-memory, queue, and occupancy requirements, so the reported percentage should not be treated as universal.
- Separate latency-critical and bulk RDMA traffic — networking / storage / HPC.
- MoE token routing;
- parameter-server control traffic;
- distributed checkpoint coordination;
- GPU storage and remote-memory systems;
- MPI-like control-plane messages.
- Dependencies: Queue isolation consumes NIC memory, queue-pair state, completion resources, and potentially CPU workers.
- Use mechanism-level microbenchmarks when evaluating communication libraries — academia / engineering benchmarking.
- hardware submission cost;
- queue-management overhead;
- memory-ordering cost;
- completion-scope cost;
- library API overhead;
- network and NIC congestion.
- This produces more meaningful comparisons than comparing complete libraries using different queue counts, batching policies, or completion semantics.
- Dependencies: Reproducibility requires matching driver, CUDA, NIC firmware, GPU clock, CPU operating state, topology, and transport configuration.
- Tune completion scope in existing GPU communication software — software libraries. Library maintainers can replace broad “complete everything” operations with targeted completion mechanisms. For example, an operation involving one peer or one queue should avoid polling all configured queues. This can reduce latency in NVSHMEM-like APIs and improve fine-grained synchronization in persistent kernels. Dependencies: Narrow completion scopes are safe only when the application’s dependency graph is correctly represented; overly weak completion can expose stale payload data under GPU memory-ordering rules.
- Improve operator and cluster-level diagnostics — cloud operations / performance engineering.
- GPU SM clock and throttling state;
- CPU proxy frequency and residency state;
- queue depth and queue sharing;
- doorbell batch size;
- active connection count;
- completion latency;
- NIC message rate.
- These metrics can help distinguish software overhead from PCIe, NVLink-C2C, NIC, or congestion effects.
- Dependencies: Instrumentation must avoid perturbing the fine-grained timing being measured.
- Guide hardware procurement and topology placement — data centers / HPC procurement. The results support placing GPUs and NICs on the same PCIe switch where possible, avoiding topologies that add indirect paths. They also indicate that coherent GB200-style systems may have different completion behavior from PCIe-attached systems, even when issue latency is similar. Dependencies: Procurement decisions must consider link bandwidth, memory coherency, NIC generation, driver support, and application traffic patterns rather than relying on peak bandwidth alone.
- Teach GPU networking and RDMA mechanisms using minimal reproducible examples — education / academia. Courses and research laboratories can use the paper’s minimal transports to demonstrate WQE construction, queue pairs, doorbell ordering, GPU memory registration, completion queues, and proxy submission. This is more actionable than treating NVSHMEM or NCCL as opaque APIs. Dependencies: Direct UAR mapping and GPU-resident queue experiments may require privileged driver settings and specialized hardware.
Long-Term Applications
- Adaptive communication runtimes that select submission paths dynamically — AI systems / distributed runtimes.
- route isolated messages to a warm CPU proxy;
- switch to GPU submission at high concurrency;
- create dedicated queues for bursts;
- adjust doorbell batching according to latency deadlines;
- fall back to host-assisted submission when GPU doorbell mapping is unavailable.
- Dependencies: This requires online performance models, low-overhead mode switching, reliable congestion signals, and correctness-preserving synchronization across both paths.
- Compiler and DSL support for communication-aware kernels — programming languages / compilers.
- fence scope;
- cooperative warp publication;
- WQE layout;
- batching thresholds;
- queue assignment;
- completion scope;
- communication-warp placement.
- A domain-specific language could allow programmers to express dependencies such as “publish this payload before notifying peer X,” leaving safe ordering and queue selection to the compiler.
- Dependencies: Compiler transformations must preserve GPU and NIC memory-ordering semantics; unsafe removal of fences cannot be generalized from the paper’s controlled experiments.
- Scalable all-to-all communication for very large MoE clusters — AI infrastructure / supercomputing.
- hierarchical expert routing;
- topology-aware token placement;
- connection pooling;
- hierarchical aggregation;
- sparse or selective all-to-all;
- improved dynamic-connection caching;
- transport protocols with less per-peer NIC state.
- Dependencies: Dynamic Connected transport alone may not eliminate NIC cache and connection-state pressure. Solutions must preserve load balance and avoid increasing token-routing latency.
- New NIC and interconnect designs optimized for GPU-originated traffic — hardware / networking.
- lower-cost GPU-visible doorbells;
- larger or more efficient queue caches;
- native GPU completion primitives;
- reduced per-connection state;
- better support for many-to-many traffic;
- hardware batching and aggregation;
- secure first-class GPU-to-NIC mappings.
- Such features could reduce the gap between the measured hardware floor and production-library latency.
- Dependencies: Security, isolation, virtualization, PCIe/IOMMU behavior, and compatibility with different GPU vendors must be addressed.
- Secure virtualization of GPU-to-NIC doorbell access — cloud platforms / policy and infrastructure security. Because mapping a NIC PCIe BAR or UAR into GPU address space can create a security risk, future cloud platforms could provide mediated doorbells, hardware capabilities, or hypervisor-controlled submission channels. This would allow tenant workloads to use GPU-initiated networking without requiring broad peer-mapping overrides. Dependencies: The mechanism must prevent tenants from accessing other queues, injecting unauthorized traffic, or bypassing isolation. Security controls must not eliminate the latency advantage of direct submission.
- Portable GPU-initiated networking across vendors and interconnects — software ecosystems. The paper’s distinction between GPU-built work requests, GPU-submitted doorbells, and CPU-proxy submission could inform a portable abstraction spanning NVIDIA, AMD, different NIC vendors, InfiniBand, RoCE, and future Ultra Ethernet systems. A runtime could expose capabilities rather than assuming that every platform supports direct UAR mapping. Dependencies: Portability may require a lowest-common-denominator API, which could sacrifice some hardware-specific performance. Vendor cooperation and standardized memory-ordering semantics would be important.
- Communication-aware scheduling for energy-efficient data centers — energy / cloud operations. Schedulers could consolidate low-rate communication onto CPU proxies, power down unused communication workers, and activate GPU submission or additional queues only during high-throughput phases. CPU frequency and GPU SM-clock findings also suggest that energy-performance policies should include communication latency, not only compute utilization. Dependencies: Power-state transitions may be slower than the communication intervals being optimized. Dedicated proxy cores can also increase energy use when the workload is sparse.
- Formal verification and safer APIs for GPU–NIC memory ordering — systems research / safety-critical infrastructure. The paper identifies correctness hazards when a GPU observes a completion flag without a guaranteed view of the associated payload. Future libraries could provide typed completion tokens, formally specified fence scopes, and static or runtime checks preventing unsafe polling patterns in persistent kernels. Dependencies: Verification must cover GPU caches, NIC DMA ordering, PCIe or NVLink paths, remote visibility, and interactions among multiple queues.
- Autotuned communication products for managed AI clusters — cloud software / MLOps.
- queue counts;
- batch sizes;
- inline thresholds;
- CPU worker counts;
- proxy versus GPU submission;
- completion scope;
- active-connection limits.
- These policies could be attached to model-serving profiles and automatically updated after driver, firmware, GPU, or NIC changes.
- Dependencies: Autotuning must use representative traffic, including burstiness and all-to-all patterns; isolated microbenchmarks alone do not predict end-to-end model performance.
- Faster interactive AI services and user-facing applications — daily life / consumer services. If the above optimizations scale reliably, they could reduce tail latency for applications that depend on distributed MoE inference, including conversational assistants, translation, recommendation, code generation, and real-time multimodal services. The direct user benefit would be faster responses and better throughput per cluster. Dependencies: The paper measures communication mechanisms rather than end-to-end user latency. Gains will depend on model architecture, compute time, batching policy, congestion, and service-level scheduling.
Glossary
- Address vector (AV): A descriptor identifying the destination of a dynamically connected network operation. “DC transports additionally include an address-vector segment to identify the dynamically connected target (DCT).”
- All-to-all traffic: Communication in which every participant sends data to every other participant. “all-to-all traffic loses 59\% of its rate near 3,000 connections”
- BlueFlame: A NIC mechanism that allows userspace or device code to place work-request data directly into a doorbell buffer. “The doorbell targets the UAR's send-doorbell/BlueFlame region.”
- Completion queue (CQ): A queue containing records that indicate the completion status of submitted network operations. “Send and receive completions are directed to completion queues (CQs), which can be distinct or shared with other QPs.”
- Completion queue entry (CQE): A record written by the NIC to report that a signaled work request has completed. “The NIC reports a finished signaled WQE by writing a completion queue entry (CQE)”
- Completion scope: The set of queues, peers, or operations covered by a completion or synchronization operation. “We measure the cost of the completion scope in Section~\ref{sec:micro:completion}.”
- Coherent memory: Memory that can be accessed by CPUs and GPUs while maintaining a consistent view of data. “P-GB200 is a Grace--Blackwell system with coherent memory between the CPU and the GPU”
- CPU proxy: A CPU thread or worker that receives GPU-generated requests and submits the corresponding network operations. “A tuned CPU proxy matches or beats the GPU path at idle”
- CTA: A CUDA thread block, or cooperative thread array, executed as a unit on a GPU. “RC QPs may be allocated per GPU, CTA, or warp”
- DMA-BUF: A Linux kernel mechanism for sharing memory buffers between device drivers, including allowing NIC access to GPU memory. “DMA-BUF registration lets the NIC access GPU memory but does not by itself establish this reverse mapping”
- Doorbell batching: Combining multiple work requests into a single notification to the NIC. “Announcing several WQEs with one doorbell raises throughput”
- Doorbell record: A memory location containing the index of the next available work-request block for the NIC. “First, the submitting thread updates the doorbell record (\code{dbrec}) with the index of the next free WQE basic block”
- Dynamically Connected (DC) transport: An RDMA transport that allows a limited set of initiators to address many dynamically connected targets. “Dynamically Connected (DC) transport, by contrast, lets a pool of dynamically connected initiators (DCIs) address many remote DC targets (DCTs).”
- Dynamically connected initiator (DCI): A transport endpoint that can issue operations to multiple dynamically connected targets. “DCI: dynamically connected initiator”
- Dynamically connected target (DCT): An RDMA target endpoint that can receive operations from dynamically connected initiators. “Only DC WQEs carry the AV”
- Expert parallelism: A distributed-machine-learning strategy that assigns different neural-network experts to different devices. “Mixture-of-Experts, expert parallelism”
- Fence: A memory-ordering operation that ensures earlier memory accesses become visible before later accesses. “WQE stores must become visible before the doorbell announces them”
- GPUDirect Async: A technology that lets GPU operations trigger or wait for communication initiated or prepared by the CPU. “GPUDirect Async shifts synchronization onto the device”
- GPUDirect RDMA: A technology allowing a NIC to access GPU memory directly without host-side staging copies. “GPUDirect RDMA exposes GPU memory over PCIe for direct NIC access”
- GDRCopy: A mechanism providing low-latency CPU load/store access to GPU memory. “GDRCopy, which gives the CPU low-latency load/store access to GPU memory”
- GPU-initiated communication: Communication in which GPU threads construct and submit network operations without CPU control of the submission path. “GPU-initiated communication lets GPU threads post RDMA operations directly to the NIC.”
- IBGDA: NVIDIA’s InfiniBand GPU Direct Async implementation, which enables kernels to submit RDMA work directly. “Modern GPU-initiated transports like IBGDA”
- Inline payload: Data placed directly inside a work-request descriptor rather than referenced through a separate memory pointer. “An inline payload replaces the data segment's pointer with a length word and the payload.”
- InfiniBand: A high-performance networking technology commonly used for data-center and supercomputer communication. “over InfiniBand and RoCEv2”
- Memory ordering: Rules governing the visibility and relative order of memory operations across processors and devices. “A library-level comparison obscures whether the performance differences come from queue counts, batching strategies, memory fence strengths, or other API overheads”
- Mixture-of-Experts (MoE): A neural-network architecture that routes each input token to a selected subset of specialized expert modules. “In Mixture-of-Experts~(MoE) models, GPU kernels decide which experts each token is sent to”
- MMIO: Memory-mapped input/output, a technique for controlling hardware registers through memory operations. “Inline copying inflates the MMIO write payload of the doorbell”
- NIC: A network interface controller that processes and transmits network traffic. “GPU threads post RDMA operations directly to the NIC.”
- NVLink-C2C: A high-bandwidth interconnect connecting a CPU and GPU, particularly in Grace–Blackwell systems. “host rings and counters used by proxies travel over NVLink-C2C rather than PCIe.”
- PCIe BAR: A memory-address range exposed by a PCI Express device for accessing device registers or memory. “It is a hardware register inside a User Access Region (UAR), which is a slice of the NIC's PCIe BAR”
- Persistent kernel: A GPU kernel that remains resident and repeatedly performs work instead of terminating after one operation. “This is a common hazard for persistent kernels”
- Processing element (PE): A logical computational participant or GPU rank involved in communication. “PE: processing element (one GPU rank)”
- Queue pair (QP): An RDMA communication object consisting of a send queue and a receive queue. “RDMA communication is organized around Queue Pairs (QPs)”
- RDMA: Remote direct memory access, allowing one machine or device to read or write another’s memory without involving its CPU in the data movement. “GPUDirect RDMA exposes GPU memory over PCIe for direct NIC access”
- Reliable Connection (RC): An RDMA transport that associates a local queue pair with a specific remote queue pair and provides reliable delivery. “Reliable Connection (RC) transport binds each local QP to one remote QP at creation time.”
- RoCEv2: RDMA over Converged Ethernet version 2, which carries RDMA traffic over routable Ethernet networks. “with ConnectX-7 NICs over InfiniBand and RoCEv2.”
- Send Queue (SQ): The queue containing outbound RDMA work requests. “each consisting of a Send Queue (SQ) and a Receive Queue (RQ)”
- SM clock: The operating frequency of a GPU streaming multiprocessor. “Issue latency scales strongly with the SM clock.”
- Streaming multiprocessor (SM): A GPU execution unit that schedules and executes groups of threads. “communication code can reduce GPU block residency even when unused”
- User Access Region (UAR): A NIC-mapped memory region through which userspace or device code writes doorbells. “It is a hardware register inside a User Access Region (UAR)”
- Work Queue Element (WQE): A descriptor containing the information needed by a NIC to perform an RDMA operation. “both structured as circular rings of Work Queue Elements (WQEs).”
- Work request: A software-level request describing a network operation to be submitted to an RDMA queue. “We first detail the GPU-side network path: queue placement, work-request construction, doorbell ordering, and completion semantics.”
- Warp: A group of GPU threads executed together in a SIMD-style execution model. “Cooperative publication lets a warp reach 21.6\,M\,msg/s on one queue”







