Develop reliable fault attribution for failed jobs
Develop reliable fault attribution for failed cluster jobs by distinguishing failures caused by user-configured factors, such as under-provisioned memory, from failures caused by infrastructure factors, such as hardware faults, in order to determine when failures should affect sustainability feedback or user incentives.
References
As such developing reliable fault attribution remains an open challenge.
— Scoring and Gamification to Encourage Sustainable Use of Compute Clusters
(2608.18786 - MacDonald et al., 19 Aug 2026) in Section Open Challenges