Evaluate the certifier against a fully adaptive live-model attacker

Evaluate whether the certified multi-source integrity mechanism maintains its safety guarantee against a fully adaptive attacker operating against live language models.

Background

The paper evaluates the certifier against property-based attacks in a live-model setting and against a model-free adaptive attacker in a separate harness. However, it does not test an attacker that adaptively optimizes against the certifier while simultaneously interacting with live LLMs. Establishing performance under this stronger threat model would determine whether the reported safety guarantee remains effective when model behavior and attack strategies are jointly adaptive.

References

Four limits bound the claim: a fully adaptive live-model attacker remains untested, the counts are upper-bound proxies, the guarantee rests on the assumptions of Section~\ref{sec:defs}, and corroboration remains scarce.

— Certified Multi-Source Integrity for Structured Agent Actions  (2609.34245 - Pandey et al., 28 Sep 2026) in Limitations and Conclusion, Section 7