Develop safe and rapid workload profiling across heterogeneous hardware

Develop a method for agents to profile representative workloads quickly and safely across candidate heterogeneous hardware without delaying production jobs or allowing exploratory runs to consume excessive resources.

Background

HPC clusters contain heterogeneous GPUs and interconnects whose suitability depends on memory capacity, numerical precision, throughput, topology, and workload behavior. Static hardware descriptions are insufficient for predicting application performance, so agents may need to run representative profiles. The unresolved challenge is to make this profiling sufficiently fast and resource-bounded for production facilities.

References

How to support such profiling quickly and safely, without delaying production jobs or allowing exploratory runs to consume excessive resources, remains an open scheduling challenge.

— Towards Efficient HPC Systems for Agents: Challenges and Opportunities  (2609.38723 - Zheng et al., 30 Sep 2026) in Section Compute, subsection “Opportunities: Trial-Aware, Hardware-Aware, and Event-Driven Scheduling”