Calibrating evaluation rigor for high-stakes AI applications
Determine principled criteria and methods for the required level of rigor and confidence in AI model evaluations for decision-making systems in high‑stakes settings, aligning evaluation strength with use-case risk.
References
In addition, evaluations for decision-making systems in high-stakes settings will likely demand a higher level of confidence than other applications, but it is unclear how to determine the required level of rigor based on use case.
One open question is to what extent decision-makers should be expected to understand what, in the case of AI, are complex systems, having very complex interactions with their environment (consider e.g. an STPA analysis of an AI coding agent operating in a control environment). Understanding may also be recursive, for example with a CEO deferring to a CTO who defers to an R&D head who defers to safety engineers. Given this personal lack of expertise or experience on behalf of the CEO, what criteria might decision-makers use to prove that their warranted (and potentially recursive) deference of understanding is justified?
The unresolved question is which criteria can safely be delegated and which still require an independent external standard.