Fault isolation and recovery for colocated Rollplex engines

Develop fault-isolation and recovery mechanisms for Rollplex deployments in which rollout and training engines share CUDA contexts, IPC mappings, streams, and GPU resources, so that fatal device errors, invalidated IPC mappings, or wedged collectives do not necessarily cause loss of the shared iteration.

Background

Rollplex spatially overlaps rollout and training on a shared GPU pool through separate CUDA contexts, shared IPC mappings, streams, and MPS. This arrangement enables efficient cross-phase execution but does not provide independent recovery boundaries: a fatal device error, invalid IPC mapping, or stalled collective in either engine can compromise the entire iteration.

Disaggregated deployments provide stronger failure containment by placing engines in separate GPU pools, but they sacrifice some of the resource-sharing benefits targeted by Rollplex. The paper explicitly identifies fault isolation and recovery as work that remains to be addressed.

References

Rollplex relies on soft GPU sharing through separate CUDA contexts, shared IPC mappings, streams, and MPS. These mechanisms multiplex kernels and memory but do not provide independent recovery boundaries. If either colocated engine triggers a fatal device error, invalidates an IPC mapping, or wedges a collective, the shared pool will likely lose the iteration. Disaggregation can instead place the engines in separate pools, enabling independent restart and failure containment. This limitation is orthogonal to Rollplex's overlap schedule and memory-lifecycle design; we leave fault isolation and recovery for future work.

— Rollplex: Cross-Phase GPU Spatial Sharing for Vision Language Model Post-Training  (2608.14498 - Lu et al., 14 Aug 2026) in Section 6, Discussion, “Fault Tolerance and Blast Radius”