Formal guarantees for approximate safety-filtered equilibria

Establish formal guarantees for ε-Nash equilibria under approximation errors arising from imperfect safety filters and task policies in the Safety to Competence framework for competitive safety-critical Markov games.

Background

The paper’s main theoretical result assumes perfect safety filters and exact Nash equilibria. Under those assumptions, an equilibrium of the filtered game induces a non-exploitable equilibrium of the original safety-critical Markov game when players are restricted to safe policies.

In the practical Safety to Competence implementation, safety filters and task policies are learned approximations rather than exact objects, so the paper does not establish how filtering and policy-learning errors affect equilibrium quality, exploitability, or safety guarantees. The authors explicitly identify deriving guarantees for ε-Nash equilibria in this approximate setting as unresolved.

References

While our analysis assumes perfect safety filters, deriving formal guarantees for $\epsilon$-Nash equilibria under approximation errors stemming from imperfect filters and task policies remains an open research question.

— Turning Safety into Competence: Minimally Exploitable Robot Policies via Safety-Filtered Reinforcement Learning  (2609.27312 - Wu et al., 23 Sep 2026) in Section Conclusion, subsection Limitations