Reliable general-purpose defense against prompt injection
Develop a reliable, general-purpose defense against prompt injection that works consistently across diverse large language model and agentic application contexts, and remains effective against adaptive adversaries that actively adjust their attack strategies.
References
Despite many proposals by the academic community, there is still no (reliable) solution for prompt injection that works consistently in all contexts. In the general case, especially against adaptive adversaries, it continues to be an open problem.
We evaluate attack surface and do not evaluate defenses; measuring which mitigations reduce the attempted rate is left to future work.
Defense combinations, both within and across pipeline stages, have not been systematically studied. Though no single defense is universally effective (\citep{shen2025pandaguardsystematicevaluationllm}; \citep{chu-etal-2025-jailbreakradar}), defenses are evaluated only in isolation, leaving effective combinations a key open problem.
This makes it hard to know if a guardrail can handle new attack types.
Two directions remain open. First, resilience against adaptive adversaries: our primary evaluation measures fixed, known injection techniques, and we additionally modeled a white-box adaptive attacker endowed with the Quarantine Agent's classifier prompt and with end-to-end success feedback. This attacker bypassed 15 of 16 description-layer payloads with a median of one rewrite round, succeeding not by defeating Q's semantic judgment but by relocating the payload onto the task-fitting tool name, an adversarial scaffold that lies outside the bounds of content inspection. Closing this vector requires a name-layer or call-timing gate.
Looking forward, several directions remain open, including more scalable rule generalization across diverse attack taxonomies, integration with learned semantic detectors for richer input representations, and adaptive coordination strategies for multi-turn settings.
This omission is common in the prompt-level defense literature, where defenses are typically tested against fixed attack sets or generic optimizers rather than against an adversary tailored to the specific defense, and we treat it as an open gap.
Translating this causal map into robust, generalizable defenses remains the central open challenge in securing LLMs against adversarial inputs.