- The paper introduces PI-Hunter, a three-stage framework combining static interaction-surface analysis, source-aware evolutionary probing, trajectory evaluation, and patch-and-reexplore auditing to localize prompt-injection ingestion paths.
- The paper reports substantial gains over unconstrained red-teaming, including Source Recall increases from 0.255 to 0.834 and Instruction Recall from 0.436 to 0.824 in representative AgentDojo tests.
- The paper shows residual vulnerabilities remain under Spotlighting, MELON, and PIGuard, while source-aware seeding and feedback-guided mutation improve discovery, although real-world validation and automated remediation remain open challenges.
Motivation and problem setting
PI-Hunter addresses a gap in the security evaluation of LLM agents: whereas existing defenses filter malicious content at inference time and existing automated red-teaming methods optimize attack success rates, neither provides developers with visibility into where latent prompt injections enter an agentic system or how they propagate through its execution pipeline. The authors frame indirect prompt injection as a system-level vulnerability arising from the interaction between the model, external tools, retrieved content, and the environment, rather than as a property of the model alone. Malicious instructions embedded in untrusted sources (emails, web pages, repositories) are often sparse and dormant, activating only under specific retrieval patterns and reasoning trajectories, so naive exploration of the interaction space is prohibitively inefficient.
The formal objective is to generate a query set Q that maximizes exposure of latent payloads P hidden in external sources S for an agent A with tools/interfaces I. A successful test case identifies not merely a compromised source but the specific ingestion path: the intersection of a compromised source s and the tool i that introduces the payload into the agent's reasoning context. Two challenges motivate the design: the latent challenge (payloads activate only in particular reasoning states) and source–tool disambiguation (multiple tools may access one source, requiring granular attribution).
Framework architecture
PI-Hunter operates in three stages. First, static analysis maps the agent's interaction surface — tools, API schemas, accessible file types — flagging interfaces that consume untrusted content and those capable of privileged actions. Second, an evolutionary exploitation loop, inspired by genetic algorithms, iteratively evolves probing queries:
- Source-aware meta-seeding initializes a population of test cases that combine benign user intent, activation of high-risk interfaces identified during static analysis, and prompts encouraging the agent to quote, list, or closely summarize retrieved content. Dedicated seeds are generated per suspicious source rather than generic domain-level prompts.
- Trajectory execution and evaluation records retrieved contents, reasoning traces, tool invocations, and outputs; a rubric-guided Evaluator audits trajectories along five criteria (intent adherence discrepancy, third-party instruction presence, authority confusion, abnormal interface usage, sensitive action attempts), producing structured diagnostic feedback rather than binary attack labels.
- Feedback-driven mutation applies LLM-selected operators conditioned on observed failure patterns, including general-purpose operators (volume expansion, parameter fuzzing, instruction-hierarchy override, encoding obfuscation) and domain-specific operators simulating realistic scenarios such as fraud investigations or production outages. An additional meta-mutation step refines the mutation operators themselves based on accumulated effectiveness.
Third, a patch-and-reexplore stage applies lightweight transient mitigations (blacklisting instructions, restricting interfaces, isolating sources) to verified ingestion paths. These mitigations are explicitly not permanent defenses; they reshape the exploration landscape to force subsequent auditing toward undiscovered attack surfaces.
Main results
The evaluation covers AgentDojo and AgentDyn across five backbone models (Gemini-2.5-pro, Gemini-3.1-pro, GPT-5.4-mini, Claude-4.6-sonnet, DeepSeek), four attack types plus the stronger AgentVigil attacks, ReAct and Planner-Executor architectures, and four injection scenarios. Compared against an unconstrained agentic red-teaming baseline using free-form multi-turn interaction, PI-Hunter improves both localization accuracy and diversity nearly universally. Representative gains under Gemini-3.1-pro on AgentDojo with the AgentVigil attack include Source Recall improving from 0.255 to 0.834 and Instruction Recall from 0.436 to 0.824; under GPT-5.4-mini on AgentDojo, Source Diversity rises from 0.18 to 0.69 and Instruction Diversity from 0.21 to 0.73. The largest absolute gains occur on the weakest baseline configurations — e.g., Gemini-2.5-pro on AgentDyn with direct attacks, where Instruction Recall rises from 0.108 to 0.536 — indicating that feedback-driven exploration recovers injections that open-ended interaction largely misses. The consistency of gains across backbones suggests the improvement stems from adaptive trajectory-aware exploration rather than benchmark-specific behavior.
Effectiveness under defenses
A central finding is that PI-Hunter remains effective against defended agents. Under Spotlighting, MELON, and PIGuard with Gemini-2.5-pro and the AgentVigil attack, all defenses reduce overall exposure counts, but substantial residual vulnerabilities persist. Under PIGuard on AgentDojo, the baseline collapses to zero Instruction Recall while PI-Hunter still exposes injections at 0.194 Instruction Recall; under MELON, PI-Hunter's Instruction Diversity reaches 0.550 from a baseline of 0.000. The implication is direct: current defenses leave measurable latent attack surfaces, and inference-time filtering alone is insufficient to secure agentic systems. The paper positions PI-Hunter as complementary to defenses rather than a replacement.
Ablations and analysis
Two ablations isolate the framework's key components. Seeding strategy: replacing generic seeding with source-aware seeding raises Source Recall from 0.1918 to 0.4796 and Source Precision from 0.2338 to 0.7794, supporting the claim that latent injections are tightly coupled to specific ingestion paths. Mutation selection: trajectory-feedback-guided ("smart") mutation outperforms fixed and random mutation on all six metrics (e.g., Source Recall 0.4796 versus 0.3228/0.3578), confirming that adaptive steering toward vulnerable reasoning states is essential beyond mere perturbation diversity.
Generalization experiments show larger improvements for Planner-Executor agents (Instruction Recall 0.22 → 0.67) than for ReAct agents (0.27 → 0.49), suggesting that more complex multi-stage reasoning pipelines expose broader attack surfaces and benefit more from structured auditing. Iteration analysis indicates rapid early gains with saturation around 8–10 iterations, and computational overhead is modest: the most expensive configuration averages roughly 127 queries and about 11,000 tokens per audit run.
Limitations and open questions
The authors concede two principal limitations. First, evaluation is confined to controlled benchmarks; performance in live, unconstrained production environments with dynamic and uncontrollable variables remains unverified, and reproducibility in such settings is an open problem. Second, PI-Hunter identifies vulnerabilities but does not propose remediation; bridging discovery to mitigation is left open, particularly given the acknowledged tension between security posture and agent utility in existing defenses. Additionally, the transient patching mechanism's mitigations are deliberately non-permanent, so the framework's findings do not by themselves translate into deployable protections, and the evaluation relies on exact-match payload recovery, which may understate or overstate instruction-level precision relative to semantic equivalence.
Conclusion
PI-Hunter reframes automated red-teaming of LLM agents from attack-success optimization to proactive vulnerability exposure and localization. Its combination of static interaction-surface mapping, source-aware evolutionary query generation, rubric-guided trajectory auditing, and patch-and-reexplore co-evolution yields large, consistent improvements in injection recall, localization precision, and attack-surface diversity across benchmarks, models, architectures, and defenses — including against systems protected by Spotlighting, MELON, and PIGuard. The results substantiate the paper's central claim that system-level, pre-deployment auditing reveals residual risks that inference-time defenses do not eliminate, while leaving open the questions of real-world deployment validation and automated remediation.