Fault isolation and recovery for colocated Rollplex engines
Develop fault-isolation and recovery mechanisms for Rollplex deployments in which rollout and training engines share CUDA contexts, IPC mappings, streams, and GPU resources, so that fatal device errors, invalidated IPC mappings, or wedged collectives do not necessarily cause loss of the shared iteration.
References
Rollplex relies on soft GPU sharing through separate CUDA contexts, shared IPC mappings, streams, and MPS. These mechanisms multiplex kernels and memory but do not provide independent recovery boundaries. If either colocated engine triggers a fatal device error, invalidates an IPC mapping, or wedges a collective, the shared pool will likely lose the iteration. Disaggregation can instead place the engines in separate pools, enabling independent restart and failure containment. This limitation is orthogonal to Rollplex's overlap schedule and memory-lifecycle design; we leave fault isolation and recovery for future work.