Calibrating evaluation rigor for high-stakes AI applications

Determine principled criteria and methods for the required level of rigor and confidence in AI model evaluations for decision-making systems in high‑stakes settings, aligning evaluation strength with use-case risk.

Background

AI evaluations in high-stakes contexts (e.g., healthcare, finance, safety-critical operations) demand confidence beyond typical applications, but current practice lacks clear guidance for how rigorous evaluations should be.

Establishing risk-aligned evaluation rigor would improve the reliability of deployed systems and help regulators and practitioners select appropriate testing protocols.

References

In addition, evaluations for decision-making systems in high-stakes settings will likely demand a higher level of confidence than other applications, but it is unclear how to determine the required level of rigor based on use case.

— Open Problems in Technical AI Governance  (2407.14981 - Reuel et al., 2024) in Section 3.3.1 “Reliable Evaluations”

One open question is to what extent decision-makers should be expected to understand what, in the case of AI, are complex systems, having very complex interactions with their environment (consider e.g. an STPA analysis of an AI coding agent operating in a control environment). Understanding may also be recursive, for example with a CEO deferring to a CTO who defers to an R&D head who defers to safety engineers. Given this personal lack of expertise or experience on behalf of the CEO, what criteria might decision-makers use to prove that their warranted (and potentially recursive) deference of understanding is justified?

— Understanding as an Explicit and Assessable Component of Frontier AI Safety Decisions  (2608.19816 - Barrett et al., 20 Aug 2026) in Section 3.3, subsection “Methodology is applicable now and improvements specified”

The unresolved question is which criteria can safely be delegated and which still require an independent external standard.

— Scaling Large Reasoning Models beyond Human Supervision: A Path toward Superintelligence  (2608.31075 - Yang et al., 31 Aug 2026) in Section 7.1, paragraph “Semi-verifiable and open-ended tasks”