Applying accelerator traffic shaping findings to PCIe congestion in multi-GPU servers
Investigate how to incorporate the accelerator traffic shaping observations—such as message-size-aware scheduling, queue-pair allocation, and managing ingress/egress bandwidth asymmetry—into managing PCIe congestion for multi-GPU servers.
References
Some of the big open problems when applying our design to this setting will be (1) how to perform traffic shaping when GPUs are used under spatial multiplexing, (2) how to incorporate the understanding of GPU internal contention into the traffic patterns to re-shape, and (3) how to incorporate our findings when managing PCIe congestion for multi-GPU servers.
This creates a new question beyond the current paper: whether exposure should be capped locally or coordinated across the striped transfer to prevent one risky lane from becoming the collective tail.