Governance of Acceptable Risk and Staged Release
Determine appropriate staged-release policies and acceptable risk levels for frontier open-weight AI models, extending beyond the threat-actor taxonomy and pre-release evaluation framework developed in the paper.
References
Broad governance questions, such as staged release policy and determining acceptable risk levels, remain important open questions we leave for future work.
For a meaningful portion of the policies surveyed, many important questions remain unanswered. For example, what qualifies an actor to be an expert? What does it mean to be well-resourced? What differentiates medium skill from low skill? And the problem compounds at the mechanistic level of evaluation design: does $pass@100$ or $pass@5$ better reflect the time horizon and financial capacity an actor can sustain?