Isolate the security effect of multi-agent interaction

Design evaluation comparisons that determine whether a security failure is inherited from a single-principal setting, amplified by multi-agent interaction, induced by interaction, or defined only by relations among multiple principals.

Background

The paper’s evaluation audit finds that many studies demonstrate attacks in multi-agent systems without establishing whether the interaction itself caused or amplified the failure. The required comparison depends on whether a meaningful single-principal counterpart exists.

For attacks with such a counterpart, matched single-principal comparisons are needed; for inherently relational properties, the relevant communication or coordination relation must instead be varied while preserving the multi-agent setting.

References

Determining how interaction changes a security outcome remains an open evaluation challenge.

SoK: When Safe Agents Fail Together: The Security of Multi Agent LLM Systems  (2609.00595 - Yang et al., 1 Sep 2026) in Section 6, subsection 6.2, “Isolating Multi-Agent Effects”