Governance of Acceptable Risk and Staged Release

Determine appropriate staged-release policies and acceptable risk levels for frontier open-weight AI models, extending beyond the threat-actor taxonomy and pre-release evaluation framework developed in the paper.

Background

The paper focuses on characterizing threat actors for CBRN and offensive-cyber misuse evaluations rather than resolving broader governance questions. It explicitly identifies staged release and acceptable risk thresholds as unresolved issues requiring future research, both of which are central to deciding whether and under what conditions powerful open-weight models should be released.

References

Broad governance questions, such as staged release policy and determining acceptable risk levels, remain important open questions we leave for future work.

For a meaningful portion of the policies surveyed, many important questions remain unanswered. For example, what qualifies an actor to be an expert? What does it mean to be well-resourced? What differentiates medium skill from low skill? And the problem compounds at the mechanistic level of evaluation design: does $pass@100$ or $pass@5$ better reflect the time horizon and financial capacity an actor can sustain?

Toward a Threat Actor Profiling Taxonomy for Pre-Release Risk Management of Open-Weight Frontier Models  (2608.25361 - Zhang, 26 Aug 2026) in Chapter 2, Section “How Developers Currently Characterize Threat Actors”