---
title: Frontier Safety Policies (FSPs)
url: https://www.emergentmind.com/topics/frontier-safety-policies-fsps
type: topic
---

# Frontier Safety Policies (FSPs)

Frontier Safety Policies (FSPs) are pre-committed safety frameworks for frontier AI development and deployment. In the literature, the term denotes both the general class of internal and regulatory safeguards used by frontier labs and related actors to reduce risk from powerful AI systems, and the specific company-level policies that instantiate those safeguards. Their common structure is threshold-based and lifecycle-oriented: developers define dangerous capability thresholds, evaluate systems against those thresholds, and specify tiered mitigations, deployment gates, or pauses when thresholds are met or exceeded. Core tools include capability evaluations, deployment gates, usage constraints, compliance audits, and escalating safety requirements triggered by capability thresholds or risk levels; widely discussed examples include Anthropic’s Responsible Scaling Policy, OpenAI’s Preparedness Framework, and Google DeepMind’s Frontier Safety Framework [2603.10015] [2501.16500].

## 1. Historical emergence and policy setting

FSPs emerged from a broader recognition that frontier AI creates a distinct governance problem. One influential regulatory framing identifies three difficulties specific to frontier models: the **Unexpected Capabilities Problem**, the **Deployment Safety Problem**, and the **Proliferation Problem**. On that view, dangerous capabilities may arise unpredictably, harmful use is difficult to prevent once a model is deployed, and capabilities can proliferate through open release, theft, reproduction, or downstream fine-tuning. The regulatory implication is lifecycle governance across development, deployment, and post-deployment monitoring rather than point-in-time certification alone [2307.03718].

The policy salience of FSPs increased markedly after the 2024 AI Seoul Summit. Under the Frontier AI Safety Commitments, many developers agreed to publish a safety framework describing how they would manage severe risks from frontier systems. Subsequent work treats these frameworks as a first line of defense against severe harms and as a key mechanism for AI risk governance. By late 2025, they were being used by governance instruments including the EU AI Act’s Code of Practice and California’s Transparency in Frontier Artificial Intelligence Act, which made the quality and specificity of such frameworks relevant to compliance and oversight as well as internal governance [2503.04746] [2512.01166].

The concept has also been adapted beyond the initial set of Western frontier developers. In work on catastrophic AI risk and emergency management in China, FSPs are presented as an operational layer that can translate abstract aims such as “early risk warning,” “tiered management,” and “rapid, effective response” into concrete model-development procedures. In that setting, FSPs are aligned primarily with the first two phases of a four-phase emergency framework—prevention and preparedness, and surveillance and warning—rather than being treated as a complete emergency-management regime [2511.05526].

A parallel institutional literature situates FSPs within international standard-setting. Standards are described there as a bridge between technical frontier-AI safety knowledge and comparable, operational policy obligations. On that view, AI Safety Institutes are well positioned to influence how FSPs are standardized because they combine technical expertise, state backing, international engagement, and convening power, and can therefore help translate frontier-risk knowledge into thresholds, testing practices, and governance routines [2409.11314].

## 2. Core architecture and policy logic

Across papers, FSPs share a common architecture: they predefine risky capability levels, test models against them, and pre-plan the safety measures that will be activated. A concise formulation is threshold-triggered mitigation: define threshold capabilities \(T_1, T_2, \dots\), evaluate models repeatedly against those thresholds, and attach mitigations \(M_1, M_2, \dots\) to each threshold. In practical terms, the framework is an “if-then” system: if a model’s evaluated capability reaches a threshold, then the pre-planned safety response is activated [2511.05526].

A widely used decomposition treats a safety framework as comprising three core areas: **risk identification and assessment**, **risk mitigation**, and **governance**. In risk identification, developers typically begin with broad horizon scanning and then prioritize a subset of severe risk domains, often including CBRN risks, cybersecurity, autonomy-related risks, persuasion, and machine learning R&D acceleration. Risk modeling then decomposes each domain into scenarios, causal pathways, bottlenecks, defenses, and the model’s counterfactual contribution. This literature emphasizes that a framework usually cannot assess every possible risk and therefore must make its scoping assumptions explicit [2503.04746].

Thresholds are central to the entire architecture. The literature distinguishes **risk thresholds**, which define intolerable risk directly in terms of likelihood and magnitude of harm, from **capability thresholds**, which define dangerous capabilities serious enough to trigger action. Published frameworks have mostly used capability thresholds rather than explicit risk thresholds, in part because capability is more directly measurable and often changes faster than other drivers of risk. Emerging practice favors several threshold tiers, including early-warning thresholds, grounded where possible in detailed risk models and external input [2503.04746].

Mitigations are likewise tiered. At moderate risk levels, frameworks may apply enhanced input/output filtering, restricted API access, or tighter monitoring. At higher risk levels, proposed measures include heightened security, more restrictive access controls, deployment pauses, or stopping deployment entirely if risks cannot be managed. Security mitigations are treated separately from deployment mitigations because security often must be in place before model development rather than merely before public release [2511.05526] [2503.04746].

Evaluation is not conceived as a single benchmark pass. Published practices include assessment during training, at checkpoints, before updates, after deployment, and when new information arises. Examples in the literature include evaluation every \(2\times\) or \(6\times\) increase in effective compute, or every 3–6 months. Frameworks are also expected to specify elicitation effort, use multiple task types rather than a single proxy benchmark, incorporate safety margins, and combine targeted evaluations with more open-ended red-teaming for unanticipated risks [2503.04746].

## 3. Evaluation regimes and operational examples

One of the clearest concrete instantiations of an FSP is Amazon’s Frontier Model Safety Framework (FMSF), used to decide whether Nova 2.0 Lite was safe for public release. FMSF defines a small set of **Critical Capability Thresholds** in three high-risk domains—Chemical, Biological, Radiological and Nuclear (CBRN), Offensive Cyber Operations, and Automated AI R&D—and treats a model as non-releasable if it crosses any threshold by materially lowering the barrier for misuse beyond what is already possible with public tools or ordinary research. The framework is implemented through a layered decision process rather than a universal score cutoff: domain-specific threshold definitions, automated benchmarks, expert red-teaming, uplift studies or related human-centered stress tests, and external review [2601.19134].

The threshold definitions themselves illustrate the policy logic. For CBRN, the threshold is crossed if a model is capable of “providing expert-level, interactive instruction that delivers material uplift beyond what is available through public tools or research, in a manner that enables a non-expert to reliably produce and deploy a CBRN weapon.” For offensive cyber operations, the threshold is the ability to enable a non-expert to discover and exploit high-value vulnerabilities “beyond what is possible with existing public tools and research.” For automated AI R&D, the threshold is the point where an AI system can “replace human researchers and fully automate the research, development, and deployment of frontier models that will pose severe risk — such as accelerating the development of enhanced CBRN weapons and offensive cybersecurity methods” [2601.19134].

The evaluation stack used in that report is also representative of the broader FSP literature. In CBRN, Amazon used WMDP-Bio, WMDP-Chem, ProtocolQA, BioLP-Bench, the Virology Capabilities Test, a multimodal chemistry image-understanding evaluation, and an independent uplift study by Nemesys Insights with nearly 800 participants. Nova 2.0 Lite was reported to score 0.71 on WMDP-Chem, around 0.82 on WMDP-Bio, 0.49 on ProtocolQA, 0.24 on BioLP-Bench, and 0.29 on VCT. The uplift study found the model below the overall CBRN threshold, though it identified “a meaningful performance uplift in the radiological domain,” prompting enriched safety filters and increased monitoring for relevant topics [2601.19134].

In offensive cyber, the evaluation combined CyberMetric, SECURE-CWET, CTIBench, CyBench, and hands-on testing in Hack The Box environments using a custom agent on Kali Linux EC2. Nova 2.0 Lite achieved over 85% accuracy on general cybersecurity knowledge benchmarks and a 7.5% uplift over prior versions in CyBench, especially on easy and very easy tasks, but the report concluded that it did not reliably translate analysis into working high-impact exploits and did not cross the cyber threshold. In automated AI R&D, the main benchmark was RE-Bench, focusing on code-intensive tasks such as embedding repair, training pipeline optimization, and restricted masked language modeling. METR independently concluded that the model did not cross the Automated AI R&D Critical Capability Threshold, while also noting that Amazon had provided information to “rule out severe under-elicitation” and to upper-bound the delay between internal use and public deployment [2601.19134].

This example is notable because it makes explicit a recurring feature of FSPs: thresholding is typically qualitative at the top level and evidentially plural at the operational level. The framework does not define a single pass/fail equation. Instead, it combines benchmark evidence, red-teaming, human uplift evidence, and external review to determine whether the model has crossed a policy-relevant threshold [2601.19134].

## 4. Safety cases, dynamic assurance, and systematic hazard analysis

A major strand of the literature treats safety cases as the evidentiary and argumentative machinery that operationalizes FSPs. In that framing, an FSP is an organization-level frontier safety framework, while a safety case is the system-level argument showing that a particular model or deployment satisfies that framework. A safety case is defined as “a structured argument, supported by evidence, that a system is safe enough in a given operational context.” It is outcome-focused rather than purely process-focused, and in frontier AI it is intended to support go/no-go judgments for training runs, deployments, or capability expansions [2410.21572].

The standard decomposition has four components: **scope**, **objectives**, **arguments**, and **evidence**. Scope covers the specific system, deployment setup, assumptions, time period, and update conditions. Objectives specify the safety requirements, often as a risk threshold or proxy capability threshold. Arguments provide the chain of claims showing why the objectives are met. Evidence includes evaluations, expert judgments, process documentation, external red-team findings, and related materials. One illustrative objective in this literature defines unacceptable risk quantitatively as a probability of at least \(10^{-7}\) per year of causing an event with at least 1,000 fatalities [2410.21572].

Safety cases are generally presented as living documents. They begin early in development, evolve as documentation and evidence accumulate, and are updated when triggers arise such as new jailbreaks, post-deployment enhancements, or changes in the system’s use. This lifecycle emphasis is reinforced by work on a **Dynamic Safety Case Management System (DSCMS)**, which adapts methods from autonomous vehicles, especially Checkable Safety Arguments (CSA) and Safety Performance Indicators (SPIs) recommended by UL 4600. The DSCMS requirements are extensive: create an initial safety case early in the lifecycle, maintain it throughout the lifecycle, perform automated consistency checks, define and maintain a catalogue of SPIs with thresholds and update frequencies, automatically re-evaluate affected arguments when SPI thresholds are breached, integrate external data feeds, perform impact analyses on new external data, provide a governance reporting interface, and implement data security and access control [2412.17618].

Within that framework, SPIs are attached to claims in the safety argument. The process is described as a continuous cycle of **monitoring, assessment, decision-making, action, and revalidation**. The literature distinguishes internal from external SPIs, and leading from lagging indicators. A breached SPI can invalidate a claim, trigger consistency checks, identify the “impact area” in the argument, and escalate to governance review. This moves FSPs beyond static pre-deployment paperwork toward continuous safety assurance [2412.17618].

Systematic hazard-analysis proposals pursue a similar objective from a different direction. Work using **STPA (Systems-Theoretic Process Analysis)** argues that many frontier AI safety frameworks remain too informal in how they identify hazards and justify safety claims. STPA decomposes the problem into scope and system definition, control structure modeling, Unsafe Control Actions (UCAs), and loss scenarios. In an AI-control case study focused on exfiltration of sensitive intellectual property, STPA surfaced sociotechnical risks such as delayed shutdown after an exfiltration attempt, ineffective memory reset, misleading human-audit feedback, and state persistence across resets. The stated purpose is not to replace capability thresholds or evaluations, but to complement and cross-check them, improving robustness and traceability from losses to hazards, control actions, UCAs, loss scenarios, and mitigations [2506.01782].

A further refinement concerns confidence in the safety case itself. One paper argues that a binary conclusion such as “Deploying the AI system does not pose unacceptable cyber risk” is insufficient without an explicit account of confidence. In a seven-component fragment of a cyber-misuse inability argument, achieving 95% confidence in the overall claim required each component to be about \(0.99286\) under the sum-of-doubts method and about \(0.99270\) under the product method. The stated implication is that numerical confidence aggregation is challenging and that the process of estimating confidence, analyzing defeaters, and communicating residual doubt is itself a governance task relevant to deployment decisions [2502.05791].

## 5. Auditing, standards, and regulatory embedding

FSPs were initially developed as forms of self-regulation, but the literature consistently treats self-regulation as insufficient on its own. One influential proposal therefore identifies three regulatory building blocks: **standard-setting processes** to define appropriate safety requirements, **registration and reporting requirements** to give regulators visibility into frontier AI development, and **mechanisms to ensure compliance** with safety standards. Initial safety standards in that framework include conducting thorough risk assessments informed by dangerous-capability and controllability evaluations, engaging external experts for independent scrutiny, following standardized protocols for deployment based on assessed risk, and monitoring and responding to new information on model capabilities [2307.03718].

Third-party assessment has become a major focus of subsequent work. One line of research proposes **third-party compliance reviews** in which an independent external party assesses whether a company is complying with its own safety framework. That literature organizes review design around six questions: who conducts the review, what information sources are used, how compliance is assessed, what is disclosed externally, how findings guide development and deployment actions, and when reviews are conducted. It also distinguishes minimalist, more ambitious, and comprehensive approaches. The core rationale is straightforward: compliance reviews can increase adherence to published commitments and provide assurance to boards, senior management, regulators, customers, and other stakeholders, while also creating information-security, cost, and reputational tradeoffs familiar from other high-risk industries [2505.01643].

A more expansive assurance proposal defines **frontier AI auditing** as “rigorous third-party verification of frontier AI developers’ safety and security claims, and evaluation of their systems and practices against relevant standards, based on deep, secure access to non-public information.” To make rigor legible and comparable, that work introduces **AI Assurance Levels (AAL-1 to AAL-4)**, ranging from time-bounded assessments to continuous, deception-resilient verification [2601.11699].

| Level | Characterization | Readiness |
|---|---|---|
| **AAL-1** | Limited assurance; time-bounded; API access plus limited non-public information | Achievable now |
| **AAL-2** | Moderate assurance; months; gray-box access, documentation, interviews, limited continuous monitoring | Early to mid-2026 |
| **AAL-3** | High assurance; ongoing multiyear oversight; white-box access and continuous monitoring | Uncertain, possibly early 2027 |
| **AAL-4** | Very high assurance; continuous verification designed to detect active deception attempts | Uncertain, possibly late 2027 |

The same paper treats auditing as the missing assurance layer that turns FSPs from self-declared commitments into externally checkable claims. It also proposes a PCAOB-style oversight body, a Frontier AI Auditor Accreditation Program, and differentiated public and non-public reporting, reflecting the view that public transparency alone cannot expose all safety- and security-relevant facts [2601.11699].

At the strongest regulatory end of the spectrum, approval-regulation proposals embed FSP-like logic directly into a licensing regime. One such scheme would require approval both for large-scale training and for deployment, with scrutiny beginning before training and continuing through post-deployment monitoring. The proposed entry filter is planned training above \(10^{26}\) FLOPs, and the process centers on a Training Authorization Application, a Certification Basis, a Project-Specific Certification Plan, a Model Deployment Card, and Instructions for Continued Safety. Although presented as a proposal rather than an established regime, it illustrates how FSP elements—capability estimation, deployment-readiness criteria, evidence generation, and post-deployment monitoring—can be transformed into legally binding gates [2408.06210].

## 6. Critiques, limitations, and proposed extensions

The most recurrent critique is that first-wave FSPs are too vague, too narrow, and too difficult to verify externally. On this view, loosely specified capability thresholds such as “meaningfully improved assistance” can be interpreted differently by different observers, undefined upper bounds make some thresholds incomplete, and the absence of clear updating mechanisms weakens the policy over time. The proposed response is **FSPs Plus**, built on two pillars: a standardized taxonomy of **precursory capabilities** and explicit incorporation of AI safety cases with a mutual feedback loop. Precursory capabilities are defined as smaller, preliminary components of a high-impact capability that function as “but for” skills along a causal progression; the claim is that they provide clearer, earlier, and more harmonizable warning signals than vague end-state thresholds [2501.16500].

A second critique is empirical. A 65-criteria assessment of the twelve frameworks published by October 2025—Amazon, Anthropic, Cohere, G42, Google DeepMind, Magic, Meta, Microsoft, Naver, NVIDIA, OpenAI, and xAI—scored them across four dimensions: risk identification, risk analysis and evaluation, risk treatment, and risk governance. Overall scores ranged from 8% to 35%, with a median of 18.5%, and the authors estimated that adoption of the best practices already present somewhere in the set could raise the attainable score to 52%. The most critical near-universal gaps were failure to define quantitative risk tolerances, failure to specify capability thresholds for pausing development, and failure to systematically identify unknown risks [2512.01166].

A third critique concerns coordination. One recent paper argues that frontier AI safety policies concentrate on prevention—capability evaluations, deployment gates, usage constraints, compliance audits, and escalating safety requirements—but neglect ecosystem-wide capacity to coordinate when prevention fails. The claim is that this **coordination gap** is structural because robustness and preparedness have diffuse benefits but concentrated costs, creating a public-goods problem and systematic underinvestment. To address this, the paper proposes a **Scenario Response Registry (SRR)** run by a public authority, with a process summarized as **scenario library → actor filings → harmonization → improved crisis preparedness**, updated by drills and real incidents. The purpose is to make commitments visible, comparable, and testable before a crisis rather than after a “focusing event” [2603.10015].

A related debate concerns what counts as the policy itself. In one reflexive-audit line of work, the phrase “Frontier Safety Policies” refers to the hidden safety boundaries that frontier models internalize after RLHF and related alignment training. The Symbolic-Neural Consistency Audit (SNCA) extracts a model’s self-stated rules, formalizes them as typed predicates, and compares them deterministically to observed behavior. Across four frontier models, 45 harm categories, and 47,496 observations, the study found systematic gaps between stated policy and actual behavior, overall self-consistency scores ranging from 0.245 to 0.800, and only 11% cross-model agreement on rule type across all categories. This usage differs from the company-policy literature, but it reinforces a common concern: whether the operative safety boundary is explicit, measurable, and actually followed [2604.09189].

Taken together, these critiques do not reject FSPs as a governance form. Rather, they redefine the problem that FSPs must solve. Early frameworks focused on evaluation-gated scaling and deployment control; later work increasingly demands explicit risk tolerances, systematic hazard analysis, dynamic safety cases, third-party assurance, stronger pause commitments, and coordination architectures that remain functional when assumptions fail. A plausible implication is that the long-term significance of FSPs will depend less on whether developers publish a framework than on whether the framework is specific, updateable, externally assessable, and embedded in institutions capable of acting on it [2501.16500] [2603.10015].

Source: https://www.emergentmind.com/topics/frontier-safety-policies-fsps