Develop reliable fault attribution for failed jobs

Develop reliable fault attribution for failed cluster jobs by distinguishing failures caused by user-configured factors, such as under-provisioned memory, from failures caused by infrastructure factors, such as hardware faults, in order to determine when failures should affect sustainability feedback or user incentives.

Background

Failed jobs consume energy and hardware lifetime without producing useful output, creating a challenge for sustainability scoring. The paper observes that failures may result either from user mistakes or from infrastructure problems beyond the user's control.

A scoring system that penalises all failed jobs could discourage legitimate resource use when failures are infrastructure-related, whereas excluding failed jobs would remove incentives to configure workloads carefully. Resolving this tension requires a reliable method for attributing the cause of job failures.

References

As such developing reliable fault attribution remains an open challenge.

Scoring and Gamification to Encourage Sustainable Use of Compute Clusters  (2608.18786 - MacDonald et al., 19 Aug 2026) in Section Open Challenges