Papers
Topics
Authors
Recent
Search
2000 character limit reached

GPU-Induced Communication

Updated 4 October 2026
  • GPU-initiated Communication is control mechanisms in which the GPU initiates, advances or completes data movement, separating the control path from the data.
  • float GPUs replace CPU-driven communication protocols with parallel, asynchronous, stream-triggered, and kernel-initiated methods, with or without hardware integration.
  • Single-word RDMA Writes time ~ 4.03 µs.

GPU-initiated communication is a class of communication mechanisms in which a GPU, rather than exclusively the host CPU, initiates, advances, or completes data movement and coordination. The term encompasses stream-triggered communication, kernel-triggered communication, and kernel-initiated communication. These mechanisms separate the communication control path—operation construction, triggering, ordering, progress, completion, and synchronization—from the data path, which moves payloads through GPU memory, peer interconnects, DMA engines, or network interfaces. GPU-aware communication may eliminate host-memory staging while retaining CPU orchestration; GPU-initiated communication moves at least part of that orchestration into GPU streams, GPU kernels, NIC hardware, or GPU-visible runtime structures (Namashivayam, 31 Mar 2025).

1. Conceptual foundations and terminology

Conventional heterogeneous communication proceeds through a CPU-controlled sequence. A GPU kernel produces data, the CPU detects or synchronizes with kernel completion, the CPU posts and progresses communication, and subsequent GPU work is launched only after the CPU observes communication completion. GPU-aware MPI can transfer data directly between GPU memory and a NIC through GPUDirect RDMA, or between peer GPUs through GPU peer-to-peer mechanisms, but GPU awareness does not imply GPU-controlled initiation or progress (Namashivayam et al., 2022).

The distinction is therefore:

  • GPU-aware communication: GPU memory can serve as a source or destination, and host-memory staging may be avoided, but the CPU may still post operations, ring NIC doorbells, poll progress, perform message matching, and synchronize with GPU execution.
  • GPU-centric communication: the GPU manages some or all of the communication control path for an operation involving GPU-resident data.
  • GPU-initiated communication: a communication transition is caused by GPU execution rather than exclusively by host code.
  • Device-triggered communication: an implementation-oriented term covering stream-triggered and kernel-triggered execution.
  • Asynchronous communication: communication proceeds asynchronously with respect to a caller; this does not imply GPU initiation.
  • Persistent communication: setup is separated from repeated activation, making persistent requests relevant to GPU triggering.
  • Partitioned communication: a request is divided into partitions whose readiness can be marked independently, enabling a GPU kernel to expose data incrementally.

The GPU may participate at several levels. A GPU stream or command-queue scheduler may trigger a previously prepared operation. A GPU kernel may invoke device-callable communication functions. A GPU may construct NIC work-queue entries, update queue state, ring a NIC doorbell, and poll completion state. Conversely, a GPU may only enqueue a compact descriptor into a host-visible queue while a CPU proxy submits the physical RDMA operation. These models all reduce CPU orchestration, but they do not provide identical degrees of GPU autonomy (Bridges et al., 2024).

The data path and control path must also be distinguished. GPUDirect P2P and GPUDirect RDMA primarily address data movement. GPUDirect Async, triggered NIC operations, GPU-visible queues, device-callable APIs, and completion counters address control. A GPU can therefore control communication timing even when a NIC, DMA engine, or CPU proxy performs the actual transfer (Namashivayam, 31 Mar 2025).

2. Execution models

2.1 Stream-triggered communication

Stream-triggered communication places communication commands into a GPU stream. Operations in one stream execute in FIFO order, while operations in different streams may execute asynchronously. A typical sequence is:

K1→trigger communication→wait for completion→K2.K_1 \rightarrow \text{trigger communication} \rightarrow \text{wait for completion} \rightarrow K_2.

The CPU creates communication descriptors, associates them with a stream, and returns. The GPU stream execution controller later executes the trigger after K1K_1, activates the communication, waits for completion, and makes K2K_2 eligible. Communication therefore occurs at GPU kernel boundaries rather than from inside a running kernel (Namashivayam, 31 Mar 2025).

The HPE stream-triggered model demonstrates this design using Slingshot 11 triggered operations, Libfabric deferred work queues, and AMD HIP stream memory operations. A deferred work-queue descriptor contains a DMA descriptor, trigger-counter object, completion-counter object, and trigger threshold. hipStreamWriteValue64 updates the NIC trigger counter after preceding stream operations complete, while hipStreamWaitValue64 prevents subsequent stream operations from executing until the NIC completion counter reaches the required value (Namashivayam et al., 2022).

The proposed MPIX_Queue abstraction binds an MPI queue to a user-provided GPU stream. MPIX_Enqueue_send and MPIX_Enqueue_recv create deferred descriptors; MPIX_Enqueue_start appends a stream writeValue; and MPIX_Enqueue_wait appends a stream waitValue. Multiple operations can be batched behind one start operation. The interface is asynchronous with respect to the CPU, but wildcard matching using MPI_ANY_SOURCE and MPI_ANY_TAG is restricted (Namashivayam et al., 2022).

Stream triggering generally offers natural integration with CUDA or ROCm stream ordering and avoids changing the semantics of ordinary MPI calls when explicit enqueue variants are used. Its limitations include special queue or communicator objects, restricted collective support, incomplete GPU/NIC progress, and receive-side dependence on host progress in systems lacking triggered receive operations.

2.2 Kernel-triggered communication

Kernel-triggered communication permits GPU threads or thread blocks to trigger preposted communication while a kernel is executing. The CPU prepares descriptors before launching the kernel, but the GPU determines when the associated operations execute. This permits communication to occur inside a running kernel rather than only before or after a kernel boundary (Namashivayam, 31 Mar 2025).

Kernel triggering is relevant to persistent operations and partitioned communication. MPI-4 partitioned communication separates request initialization, request start, and partition readiness. A producer kernel can invoke MPIX_Pready as each partition becomes available, while a consumer kernel can use MPIX_Parrived to detect arrival. The model supports pipelining but has asymmetric completion semantics: device-side receive-arrival detection is available, while device-side completion of transmitted partitions generally is not (Bridges et al., 2024).

Kernel-triggered approaches require device-callable MPI entry points, device-visible request objects, definitions of thread, warp, block, and memory semantics, and rules governing interactions between device-side and host-side MPI calls. They provide finer-grained control than stream triggering but require a more complex GPU concurrency and memory model.

2.3 Kernel-initiated communication

Kernel-initiated communication is the strongest form of GPU-centric communication. A GPU thread or block determines that communication is required, constructs or prepares the operation, initiates it, and manages synchronization while the kernel remains active. The communication pattern may be dynamic, with destinations and message sizes determined by GPU computation (Namashivayam, 31 Mar 2025).

NVSHMEM provides a major example. Its GPU-oriented PGAS model exposes symmetric memory and device-callable put, get, remote atomic, signal, wait, fence, quiet, and collective operations. A GPU kernel can perform remote stores, obtain peer-accessible pointers, post RDMA work through IBGDA, or enqueue a request to a CPU proxy when direct GPU-to-NIC access is unavailable (Ma et al., 4 Jun 2026).

NCCL’s Device API provides three operation modes:

  • Load/Store Accessible (LSA) for GPU loads and stores over NVLink or PCIe.
  • Multimem for NVLink SHARP multicast and reduction.
  • GPU-Initiated Networking (GIN) for RDMA over InfiniBand and RoCE.

GIN exposes asynchronous one-sided operations such as put, putValue, signal, flush, counter polling, signal polling, and device-side barriers. Its GDAKI backend allows GPU-generated work-queue entries and doorbells, while its Proxy backend transfers compact descriptors to a CPU proxy that posts ordinary RDMA operations (Hamidouche et al., 19 Nov 2025).

Other designs preserve GPU-level scheduling while delegating transport execution to host proxies. UCCL-EP uses 16-byte TransferCmd records for token-level expert-parallel communication. GPU kernels determine token destinations and enqueue commands; multithreaded CPU proxies issue GPUDirect RDMA operations. Payloads remain in GPU memory and are not copied through ordinary CPU buffers. This architecture preserves GPU-controlled routing while avoiding direct GPU-to-NIC integration, enabling support for NVIDIA and AMD GPUs with EFA and Broadcom NICs (Mao et al., 22 Dec 2025).

3. Hardware and runtime mechanisms

GPU-initiated communication requires cooperation among GPU runtimes, GPU memory systems, interconnects, NICs, communication libraries, and completion mechanisms.

GPU memory and peer access

For intra-node communication, GPU peer-to-peer mechanisms may expose remote memory through ordinary GPU loads and stores. On AMD MI300A systems, a GPU kernel can directly access another APU’s HBM through Infinity Fabric. The measured direct remote-access bandwidth was approximately 103–104 GB/s103\text{--}104\ \mathrm{GB/s}, against a nominal 128 GB/s128\ \mathrm{GB/s} pairwise link, while GPU remote-access latency was approximately 690 ns690\ \mathrm{ns}, compared with 346 ns346\ \mathrm{ns} for local HBM access (Schieffer et al., 15 Aug 2025).

NVSHMEM establishes symmetric heaps and maps peer heaps into GPU address spaces when P2P access is available. A remote address can be derived from a heap-relative offset:

destremote=peer_heap_basep2p[remote_pe]+(destlocal−heap_base).dest_{\mathrm{remote}} = peer\_heap\_base_{\mathrm{p2p}}[remote\_pe] + (dest_{\mathrm{local}}-heap\_base).

For non-P2P peers, the address is associated with transport metadata, registered-memory handles, and NIC keys rather than directly dereferenced by the issuing GPU (Ma et al., 4 Jun 2026).

NIC queues and doorbells

GPU-to-NIC communication commonly uses GPU-accessible work-queue storage, completion queues, doorbell records, and NIC doorbell registers. In GPU-submitted RDMA, GPU threads construct work-queue elements, publish them in order, and ring a NIC doorbell mapped into GPU virtual address space. The doorbell register itself resides in a NIC User Access Region and is distinct from GPU memory (Baydamirli et al., 1 Oct 2026).

A typical RDMA write work-queue element includes control, remote-address, and data segments. A pointer-based reliable-connected work request is 48 bytes, while a dynamically connected request with an address-vector segment is 64 bytes. Small inline payloads can avoid a NIC DMA read of the source buffer, although inline data increases work-queue size and GPU stores (Baydamirli et al., 1 Oct 2026).

Correct posting requires ordering:

  1. work-queue and payload stores become visible to the NIC;
  2. the work queue is published in reservation order;
  3. the doorbell announces the available work;
  4. the NIC fetches and executes the work request.

GPU-scope fences and release stores are generally less expensive than system-scope fences. The measured minimal GPU path issued an 8-byte RDMA write in approximately 0.70 μs0.70\,\mu\mathrm{s} and completed it in approximately 4.03 μs4.03\,\mu\mathrm{s}, whereas software libraries added several microseconds through queue management, ordering, and completion scope (Baydamirli et al., 1 Oct 2026).

Triggered operations and deferred work queues

Triggered operations allow a NIC to store a command descriptor without executing it immediately. Execution begins when a counter reaches a threshold:

K1K_10

This mechanism connects GPU-stream events to NIC execution. On Slingshot 11, the GPU updates a trigger counter through a stream memory operation. The NIC then executes deferred sends and updates completion counters. Tagged and untagged sends, one-sided RMA operations, and atomic operations are supported, whereas hardware-triggered receives are not (Namashivayam et al., 2022).

The absence of triggered receives is consequential. Receive-side matching and unexpected-message processing may require a CPU progress thread. Intra-node two-sided operations may likewise remain CPU-assisted because suitable peer-to-peer deferred-work and message-matching support is unavailable.

CPU proxies

A CPU proxy can mediate communication without staging payload data through host memory. SpeCL’s PortChannel, GIN Proxy, UCCL-EP, and mKernel use variants of this design. GPU threads enqueue compact descriptors, while host workers translate them into RDMA work requests or vendor transport operations. The NIC then reads or writes GPU-resident payloads directly (Shah et al., 11 Apr 2025).

Proxy execution improves portability when GPU-to-NIC doorbells or device-callable verbs are unavailable. Its costs include CPU-core consumption, proxy polling, queue traversal, NUMA sensitivity, and possible contention. A tuned proxy can nevertheless approach or match direct GPU submission at idle, while direct GPU submission may provide lower latency when supported hardware eliminates the proxy path (Baydamirli et al., 1 Oct 2026).

4. Ordering, completion, progress, and resource management

GPU-initiated communication requires explicit semantics for ordering, memory visibility, local completion, remote completion, progress, and resource reuse.

Ordering and visibility

A GPU-produced send requires:

K1K_11

A received buffer requires:

K1K_12

These relations may be implemented through stream order, GPU fences, release/acquire operations, NIC ordering, sequence values, signals, counters, or completion queues. A completion notification does not automatically constitute a data-memory fence. Applications and libraries must establish the required memory visibility explicitly (Baydamirli et al., 1 Oct 2026).

GIN uses local counters and remote signals. A local flush establishes source-buffer reuse safety, whereas a remote signal indicates that preceding operations are visible to the peer within the relevant ordering domain. NVSHMEM distinguishes fence, which orders operations toward a destination, from quiet, which provides stronger completion and visibility guarantees (Hamidouche et al., 19 Nov 2025, Ma et al., 4 Jun 2026).

On InfiniBand reliable-connected transport, a payload write followed by a flag write on the same connection can establish data-before-flag ordering. AWS EFA’s SRD transport is reliable but unordered, requiring proxy-side ordering mechanisms such as immediate data, sequence numbers, and control buffers. UCCL-EP uses these facilities to emulate partial completion fences and per-channel ordering (Mao et al., 22 Dec 2025).

Completion scopes

Completion may refer to several distinct conditions:

  • the source buffer is safe for reuse;
  • the NIC has consumed a work request;
  • the remote memory contains the payload;
  • the payload is visible to remote GPU threads;
  • a communication phase or barrier is complete;
  • provider-side transport resources can be recycled.

These conditions are not equivalent. GICC explicitly distinguishes GPU-visible completion from provider-visible retirement. On Slingshot, a GPU may observe completion while the host still needs to advance libfabric state and recycle deferred-work entries (Shan et al., 24 Apr 2026).

Broad completion operations can dominate software overhead. NVSHMEM quiet may inspect all configured queue pairs and dynamic connection initiators, while narrower selected-queue completion can remain approximately constant as queue count grows. This makes completion scope a central performance and scalability parameter (Baydamirli et al., 1 Oct 2026).

Finite resources and asynchronous reclamation

GPU-initiated systems must manage finite NIC queue entries, counters, completion queues, registered-memory state, command-ring capacity, and GPU execution resources. GICC addresses finite Slingshot state using epochs, double-buffered stage-ahead preparation, and asynchronous resource reclamation. The host monitor retires completed operations and re-arms future epochs while GPU execution continues (Shan et al., 24 Apr 2026).

The Slingshot implementation has a trigger-counter maximum of 2047 and a deferred-work-queue capacity of 256 entries. Repeated barriers can exhaust this state without reclamation. GICC uses readiness epochs so that the GPU triggers only work whose NIC state has been safely prepared.

Queue pressure also appears in GPU-to-CPU proxy systems. UCCL-EP bounds its FIFO with kMaxInflight; when the queue fills, GPU enqueue operations stall. mKernel uses queue credits, bounded proxy batches, and separate compute-tile, network-chunk, and proxy-batch granularities (Mao et al., 22 Dec 2025, Mao et al., 11 Sep 2026).

Post-issue progress and backpressure

A remote store may be accepted by the issuing GPU before it becomes visible remotely. The interval between issue and remote visibility is a software-visible post-issue stage called X-Stage. During this stage, the issuer can resume useful computation, but outstanding requests consume finite downstream capacity. If injection exceeds the effective drain rate, later stores incur backpressure (Xian et al., 25 Jul 2026).

The Burst–Gap model characterizes this behavior using backpressure-free issue time, effective drain rate, and outstanding capacity:

K1K_13

The measured effective drain rate on the evaluated eight-GPU NVLink/NVSwitch system was approximately K1K_14, with an effective outstanding capacity of approximately K1K_15. The results show that short bursts can drain during useful computation, while sustained injection exhausts capacity and delays subsequent issue.

5. Programming systems and applications

GPU-initiated communication appears in MPI extensions, PGAS libraries, collective runtimes, custom kernel libraries, and application-specific fused kernels.

MPI and message passing

MPICH’s MPIX_Stream model creates an explicit stream object and stream communicator, preserving ordinary MPI semantics by using explicit enqueue operations. HPE provides stream-triggered two-sided communication and stream-aware one-sided RMA. Project Delorean represents communication and computation as deferred operation graphs with partial ordering. MPI-4 partitioned communication exposes device-side partition readiness but has gaps in send-completion detection and matching preparation (Bridges et al., 2024).

These approaches are particularly relevant to nearest-neighbor exchange, regular collectives, fixed message schedules, and applications in which communication dependencies align with GPU stream order. General two-sided MPI remains difficult to offload because of tags, communicators, unexpected-message queues, wildcard matching, and late matching.

PGAS and one-sided communication

NVSHMEM, ROC_SHMEM, Intel SHMEM, and related PGAS systems expose GPU-callable put, get, atomic, signal, wait, and collective operations. Intel SHMEM combines OpenSHMEM with SYCL and supports GPU device code, symmetric heaps, remote atomics, signals, synchronization, and team collectives. Its ishmemx_ work-group operations allow multiple work-items to cooperate on one logical communication operation, avoiding excessive NIC injection for inter-node traffic while exploiting GPU parallelism for intra-node movement (Brooks et al., 2024).

Direct GPU loads and stores are effective for irregular or fine-grained access, whereas copy engines or DMA paths are generally preferable for large transfers. Intel SHMEM selects between direct operations and copy-engine transfers according to message size, work-group size, topology, operation type, and participating PE count.

Collective and AI communication

MSCCL++/SpeCL provides GPU-callable put, signal, wait, and flush primitives organized into MemoryChannel, PortChannel, and SwitchChannel abstractions. The system allows communication algorithms to be written directly in GPU kernels, expressed through a DSL, or exposed through an NCCL-compatible API. It reported speedups of up to K1K_16 for collective communication and up to K1K_17 for real-world AI inference workloads, with deployment in Microsoft Azure AI services and adoption by RCCL (Shah et al., 11 Apr 2025).

NCCL GIN adds device-side RDMA to NCCL’s collective infrastructure. Its GDAKI backend directly programs compatible NICs through DOCA GPUNetIO, while its Proxy backend uses GPU-to-CPU queues and ordinary RDMA interfaces. In DeepEP integration, GIN achieved performance close to NVSHMEM across high-throughput and low-latency expert-parallel kernels. The GIN microbenchmark reported approximately K1K_18 round-trip latency for 4–128-byte messages with GDAKI, compared with approximately K1K_19 for its Proxy path (Hamidouche et al., 19 Nov 2025).

MoE communication is a principal use case because routing decisions are data-dependent, token transfers are fine-grained, and dispatch and combine form sparse all-to-all patterns. DeepEP, NVSHMEM, GIN, UCCL-EP, and mKernel use GPU-controlled token routing, one-sided transfers, compact command queues, hierarchical communication, or fused computation.

Fused kernels and scientific applications

GROMACS halo exchange has been redesigned using GPU kernel-initiated NVSHMEM communication. Fused packing, transfer, signaling, forwarding, and unpacking eliminate repeated GPU–CPU synchronization. The reported strong-scaling improvements were up to K2K_20 for intra-node NVLink, K2K_21 for multi-node NVLink, and K2K_22 for multi-node NVLink plus InfiniBand (Doijade et al., 25 Sep 2025).

mKernel combines persistent computation with intra-node NVLink and inter-node RDMA. It partitions SMs into compute and communication roles, uses hierarchical aggregation, and adapts the partition at runtime. The evaluated kernels achieved up to K2K_23 for GEMM+AllReduce, K2K_24 for Ring Attention, and K2K_25 for MoE Dispatch+GEMM on the reported configurations (Mao et al., 11 Sep 2026).

GICC targets GPU-triggered coordination over InfiniBand and HPE Slingshot. It enables GPU-triggered puts, active messages, barriers, and waits, while separating coordination semantics from transport execution. On Slingshot, reported per-coordination latency decreased from K2K_26 for host-driven execution to K2K_27, and on InfiniBand GICC achieved approximately K2K_28 lower small-put latency than NVSHMEM in the evaluated configuration (Shan et al., 24 Apr 2026).

GPU-initiated communication is not restricted to systems with direct peer memory or GPU-addressable NICs. ThunderEP targets PCIe-connected consumer GPUs without GPU peer-to-peer access. It treats pinned CPU memory as a shared collective medium, removes ring relays, uses DMA engines for bulk transfers, and exposes per-sender completion. The system reported average dispatch speedups of approximately K2K_29 and combine speedups of approximately 103–104 GB/s103\text{--}104\ \mathrm{GB/s}0 over NCCL in its aggregate evaluation, while remaining host-assisted rather than strictly device-initiated (Lee et al., 30 Sep 2026).

6. Performance, limitations, and standardization

GPU-initiated communication primarily targets control-path latency, overlap, CPU scalability, and fine-grained dynamic communication. It does not automatically improve raw network bandwidth. Benefits are greatest when communication is small, frequent, data-dependent, or tightly interleaved with computation.

Results vary significantly by transport and workload. In the Faces benchmark, stream-triggered communication was approximately 103–104 GB/s103\text{--}104\ \mathrm{GB/s}1 slower than baseline in an eight-node, eight-rank-per-node configuration because intra-node communication required CPU progress threads. It was approximately 103–104 GB/s103\text{--}104\ \mathrm{GB/s}2 slower in the one-node configuration, similar to baseline for eight nodes with one rank per node, and approximately 103–104 GB/s103\text{--}104\ \mathrm{GB/s}3 better in a three-dimensional communication arrangement. Hand-coded shader operations improved the tested result from approximately 103–104 GB/s103\text{--}104\ \mathrm{GB/s}4 better to approximately 103–104 GB/s103\text{--}104\ \mathrm{GB/s}5 better than baseline (Namashivayam et al., 2022).

GPU-initiated communication also consumes GPU and NIC resources. Communication warps, polling, work-request construction, synchronization, and register state can reduce compute throughput even when communication is dormant. In fused kernels, SM allocation is a performance parameter: the best communication allocation varies with kernel type and input shape, and the worst measured AllReduce partition in mKernel was approximately 103–104 GB/s103\text{--}104\ \mathrm{GB/s}6 times slower than the best partition (Baydamirli et al., 1 Oct 2026, Mao et al., 11 Sep 2026).

The principal limitations are:

  • Incomplete offload: triggered receives, two-sided matching, collectives, provider progress, and intra-node control may remain CPU-assisted.
  • Hardware dependence: direct GPU-to-NIC operation depends on GPU-visible NIC queues, doorbells, GPUDirect Async, memory registration, compatible drivers, and favorable topology.
  • Portability: CUDA, ROCm, SYCL, InfiniBand, Slingshot, EFA, NVLink, xGMI, and other environments expose different ordering, memory, and control mechanisms.
  • Matching complexity: general MPI two-sided semantics require message matching, unexpected-message handling, wildcard receives, and concurrency rules.
  • Completion ambiguity: local source reuse, remote visibility, request completion, provider retirement, and buffer reuse are distinct conditions.
  • Resource exhaustion: queue entries, counters, QPs, completion queues, proxy rings, and registered memory are finite.
  • GPU resource contention: communication can reduce occupancy, register availability, SM throughput, and compute overlap.
  • Programming complexity: applications must manage fences, signal resets, buffer lifetime, queue capacity, stream ordering, thread participation, and topology-specific protocols.
  • Dynamic communication limits: stream-triggered and pre-staged mechanisms are less suitable for communication patterns whose destinations or sizes cannot be described before execution.
  • Collective limitations: device-side collectives may lack grid-wide execution, multi-CTA support, or strong inter-node performance. NVSHMEM’s evaluated inter-node collective performance was substantially below NCCL, despite strong intra-node capabilities (Ma et al., 4 Jun 2026).

The distinction between direct GPU initiation and host-assisted GPU initiation is therefore important. Direct GPU-to-NIC mechanisms minimize per-operation CPU involvement but require tightly integrated hardware and software. Proxy-mediated systems preserve GPU-controlled scheduling while using portable host interfaces and can support heterogeneous GPU/NIC combinations. UCCL-EP illustrates this tradeoff: it retained token-level GPU scheduling and achieved DeepEP-level behavior across NVIDIA and AMD GPUs with EFA, ConnectX, and Broadcom NICs, while using CPU proxies for RDMA execution (Mao et al., 22 Dec 2025).

Future standardization must define trigger semantics, host/device concurrency, message matching, memory visibility, local and remote completion, progress, buffer preparation, resource reclamation, collective support, and portability. Candidate abstractions include stream and queue objects, device-callable request APIs, partitioned operations, PGAS one-sided primitives, registered windows, signal and counter objects, and graph-based deferred operations. The central requirement is not merely a device-callable MPI_Send; it is a coherent model connecting GPU execution, NIC execution, memory ordering, completion, matching, and finite-resource management (Bridges et al., 2024).

GPU-initiated communication is consequently best understood as a family of execution and control models rather than a single technology. Stream-triggered mechanisms integrate communication with GPU scheduling; kernel-triggered mechanisms provide in-kernel activation of prepared work; kernel-initiated mechanisms allow GPU code to generate and control communication dynamically. Their common objective is to remove unnecessary CPU synchronization from the critical path while retaining appropriate transport progress, memory visibility, completion, and resource safety.

Topic to Video (Beta)

No one has generated a video about this topic yet.

Whiteboard

No one has generated a whiteboard explanation for this topic yet.

Follow Topic

Get notified by email when new papers are published related to GPU-Initiated Communication.