Cipher-Jailbreak Manipulation of AI-Agent Actions

Determine whether arbitrary permutation-cipher jailbreak attacks against commercial large language models can manipulate the actions or tool calls of AI agents.

Background

The paper demonstrates that frontier commercial LLMs can be induced to communicate through arbitrary letter-permutation ciphers and can provide harmful outputs through those encrypted channels. However, the experiments target conversational model interfaces rather than tool-using AI agents. Because AI agents can execute external actions through tools such as web browsing, code execution, or mailbox management, establishing whether encrypted jailbreak prompts can alter agent actions or tool calls is an unresolved safety and security question explicitly identified by the paper.

References

We have not evaluated the efficacy of our attack when targeting AI agents, and whether it can be used to manipulate the "actions" or tool calls of these agents.

Arbitrary Cipher Attacks Against Large Language Models Do Not Require Fine-Tuning  (2609.09553 - Rivasseau, 9 Sep 2026) in Section 7, Limitations, item “Prompt injections and AI agents”

The gap between base-model ASR and actual agent-level exploitability remains an open research question.

SoK: Trading Agents or Market Crashers? Dissecting Robustness and Security Failures in Academic Financial LLM Trading Schemes  (2609.19705 - Wang et al., 17 Sep 2026) in Section 4.2, Cross-Cutting Observations and backbone-model analysis