AgentHijack: Security Threats in Agent Systems
- AgentHijack is a family of security failures where an agent’s authorized actions are redirected by untrusted, attacker-controlled inputs within its operational context.
- It encompasses diverse attack surfaces such as indirect prompt injection, control-flow hijacking, tool poisoning, and workflow compromises that exploit inherited privileges.
- Defensive strategies focus on layered mechanisms including provenance tracking, least privilege enforcement, and runtime authority control to balance security with functional utility.
AgentHijack denotes a family of security failures in which an LLM-based, multimodal, or tool-using agent is induced to deviate from the user’s objective and exercise its legitimate capabilities in the attacker’s interest. Recent work applies the label to indirect prompt injection, action hijacking, control-flow hijacking, tool-selection and tool-surface poisoning, workflow compromise, covert computational resource theft, and, in one benchmark, robustness failures under non-adversarial environment corruptions (Zhang et al., 2 Dec 2025, Shapira et al., 8 Jun 2025, Jha et al., 20 Oct 2025, Sun et al., 25 May 2026). Across these settings, the common structure is not necessarily privilege escalation; rather, the critical failure is that untrusted context, tools, or environment state becomes effective authorization for side effects that remain syntactically valid within the agent’s normal operating envelope.
1. Scope of the term and its relation to adjacent concepts
Usage of the term is not uniform across papers. In one formulation, the adversary’s goal is to expropriate an agent’s compute for unauthorized tasks while remaining within the tool’s inherited capability scope and preserving correct user-visible outputs; this is characterized as “implicit toxicity,” where malicious behavior occurs “entirely within the allowed privilege scope” (Zhang et al., 2 Dec 2025). In another, the central failure mode is “authority confusion”: untrusted resources may inform reasoning, but they must not authorize side effects (Qin et al., 27 May 2026). Other works instantiate AgentHijack as action hijacking via retrieval-augmented memory, prompt-injected shell execution in coding assistants, and control-flow hijacking in multi-agent orchestrations (Zhang et al., 2024, Liu et al., 25 May 2026, Jha et al., 20 Oct 2025).
This suggests that AgentHijack is best understood as an umbrella term for attacks on the decision-to-action pathway of agentic systems. The attacker need not subvert model weights, system prompts, or operating-system privilege boundaries. Instead, the attacker manipulates inputs that the agent already treats as usable context: web pages, files, MCP tools, browser navigation state, workflow comments, skill bootstrap guidance, or structured tool metadata.
A common misconception is that agent hijacking is merely a variant of jailbreak prompting. The cited work argues otherwise. Several papers emphasize that the harmful step is an ordinary executable action whose danger is contextual: a tool call, shell command, network request, or control-flow transition becomes unsafe because attacker-controlled context steers authorized access against the user’s interest (Qin et al., 27 May 2026, Shapira et al., 8 Jun 2025). This distinction is especially salient in open tool ecosystems such as MCP and WebMCP, where third-party components operate inside permission envelopes inherited from the agent or session (Zhang et al., 2 Dec 2025, Lee et al., 4 Jun 2026).
2. Threat models and formal structure
Formalizations differ by domain, but they share a small set of recurring objects: a user goal, a sequence of environment observations, an agent policy, and a set of capabilities or privileges. In a web-use setting, one paper models the user’s task as , the visited pages as , the agent’s internal state as , and the privilege set as . AgentHijack occurs when injected content on some page changes the chosen action from one consistent with to one outside the valid-action set (Shapira et al., 8 Jun 2025).
A more generic indirect prompt-injection formalization defines user instruction , environment vectors , attacker-chosen malicious instruction 0, and compromised context
1
The agent is a policy 2 selecting actions 3, and attack success is defined by whether the realized trace executes the adversary’s objective rather than the user’s (Li et al., 3 Feb 2026).
In MCP-based systems, the key structural relation is capability inheritance. A typical model uses agent 4 and third-party tool 5 with declared capability sets 6, under the condition
7
Because the agent often has network access, the tool inherits that permission. LeechHijack exploits exactly this boundary: the attacker remains within 8 while covertly parasitizing the user’s compute budget (Zhang et al., 2 Dec 2025).
In multi-agent systems, the threat is formalized as manipulation of orchestration traces. A control-flow trace is
9
and valid execution must follow a control-flow graph 0 generated for the orchestrator and its agents. Control-flow hijacking occurs when malicious inter-agent content causes an unsafe transition, such as invoking an executor or exfiltration-capable agent under the guise of recovery or remediation (Jha et al., 20 Oct 2025).
AIRGuard reframes the same issue in authorization language. If a proposed side-effecting action is 1, then the security goal is
2
where 3 is protected runtime policy and 4 is execution history. The point is not whether the action was “suggested” by context, but whether it was justified by an authority source allowed to authorize that effect (Qin et al., 27 May 2026).
3. Mechanisms and attack surfaces
The most widely studied mechanism is indirect prompt injection through untrusted external content. Web-use agents ingest visible text from comments, reviews, advertisements, email links, files, and tool outputs, and then mix this content with the user task in the reasoning loop. “Mind the Web” introduces “task-aligned injection,” in which malicious commands are framed as helpful task guidance rather than overt override text, enabling payloads such as unauthorized camera activation, password leakage, local file exfiltration, user impersonation, and denial of service (Shapira et al., 8 Jun 2025). AgentDyn strengthens this threat model by requiring dynamic replanning in the presence of genuinely helpful third-party instructions, showing why simple “ignore all external instructions” policies fail on realistic tasks (Li et al., 3 Feb 2026).
Several works target the structural interfaces by which agents parse context. Phantom attacks chat-template boundaries directly. The core observation is that tool, system, user, and assistant turns are flattened into one token stream; by injecting optimized structured templates, the attacker induces role confusion and causes the model to reinterpret attacker-controlled text as legitimate user instructions or prior tool outputs. Phantom automates this with multi-level template augmentation, a Template Autoencoder, and Bayesian optimization in latent space (Deng et al., 18 Feb 2026). A related parsing-layer attack manipulates the HTML accessibility tree seen by browser agents; using Greedy Coordinate Gradient over hidden HTML insertions, the attacker achieves high exact or partial target-action rates on real sites, including login credential exfiltration and forced ad clicks (Johnson et al., 20 Jul 2025).
Other attacks target the tool layer rather than free-form content. ToolHijacker injects a malicious tool document into the tool library and exploits the two-step retrieval-and-selection pipeline so that the agent consistently selects the attacker’s tool for an attacker-chosen task (Shi et al., 28 Apr 2025). WebMCP Tool Surface Poisoning generalizes this idea to runtime tool registries: Mid-Session Tool Injection includes “Tool Hijacking,” which changes the set of visible tools via AbortSignal cancellation or registration races, and “Tool Framing,” which manipulates metadata such as tool name, description, readOnlyHint, and inputSchema to steer planning (Lee et al., 4 Jun 2026).
Long-horizon execution creates additional attack opportunities. WebTrap performs mid-task browser hijacking through a three-stage sequence—lure, inertia, payload—called multi-step instruction fusion steering. The attack binds the attacker goal to the user goal tightly enough that the agent can enter a restricted area, execute the malicious segment, return, and still complete the original task (Liu et al., 8 May 2026). JAW applies a comparable idea to agentic workflows on GitHub Actions and n8n: static path-feasibility analysis identifies feasible agent-invocation paths, dynamic prompt-provenance analysis determines how attacker-controlled input reaches the LLM context, and capability analysis enumerates the actions actually available at runtime (Fendley et al., 11 May 2026).
Open skill ecosystems add a persistence-oriented variant. In OpenClaw, third-party skills can register agent:bootstrap hooks and inject Markdown guidance into context.bootstrapFiles, which is then placed immediately after the system role during prompt assembly. Guidance injection exploits the primacy effect by embedding adversarial “best practices” narratives that reframe malicious actions as routine operations such as cleanup, migration, or diagnostics (Liu et al., 20 Mar 2026).
LeechHijack is unusual in that it does not primarily seek semantic task deviation. Instead, an adversarial MCP tool embeds a benign-looking backdoor during implantation and later activates it through content-based, frequency-based, or context-based triggers. The exploitation stage opens a command-and-control channel, injects an extra task, and causes the agent to spend inference budget solving attacker work while restoring nominal outputs afterward (Zhang et al., 2 Dec 2025).
Finally, some work uses “AgentHijack” for non-adversarial but operationally similar disruptions. The AgentHijack benchmark for computer-use agents defines nine configurable common corruptions—pop-ups, resolution changes, marks, subtitles, multi-app distractions, accidental touches, app minimization, network error, and verification screens—that interfere with perception and control without direct adversarial intent (Sun et al., 25 May 2026). This broadens the term from malicious hijack to disruption of execution flow.
4. Empirical findings across domains
Empirical results indicate that agent hijacking is not confined to one architecture or application domain. In LeechHijack, experiments across Deepseek, Qwen, GPT, and Gemini families, and across three agent architectures, report a mean success rate of approximately 5 with mean per-instance token overhead of approximately 6; the evaluation used GAIA and GPQA for benign tasks and MMLU for extra tasks (Zhang et al., 2 Dec 2025).
Phantom reports an average ASR of 7 across AgentDojo’s Workspace, Travel, Slack, and Banking domains on Qwen, GPT, Gemini, and DeepSeek variants, substantially above Single-Template (8), Semantic-Injection (9), and ChatInject (0). The paper also states that 70 distinct vulnerabilities were confirmed across 942 commercial agents (Deng et al., 18 Feb 2026).
In agentic AI coding assistants, AIShellJack evaluates 314 payloads spanning 70 MITRE ATT&CK–inspired techniques against Cursor v1.2.2, GitHub Copilot v1.102 with Gemini 2.5, and Claude Sonnet 4. Observed success rates range from 1 to 2, with file-based injection being the most reliable category (Liu et al., 25 May 2026). A separate OpenClaw study evaluates 26 malicious skills across 13 categories on six LLM backends and reports success rates from 3 to 4, with 94% of malicious skills evading existing static and LLM-based scanners (Liu et al., 20 Mar 2026).
Web agents show similarly high susceptibility. “Mind the Web” validates nine payload types against OpenAI Operator, Browser Use, Do Browser, and OpenOperator, with success rates of 5 across multiple LLMs (Shapira et al., 8 Jun 2025). In accessibility-tree attacks, TWUI exact ASR reaches 6 on chess.com, 7 on the binary-numbers game, 8 on citybrewtours.com, 9 on norway.no, and 0 on translate.google.com; for universal-website login exfiltration, exact ASR is 1 on training pages and 2 on unseen test pages (Johnson et al., 20 Jul 2025). WebTrap reports, on long GitLab browser tasks under best-of-3 evaluation, ASR-E 3, ASR-I 4, and UUA 5, while dual-goal success reaches 6 compared with 7 for Hijacking Text (Liu et al., 8 May 2026).
Embodied and multimodal agents are also affected. AgentHazard evaluates 7 mobile GUI agents and 5 backbone models over more than 3,000 attack scenarios, finding an average misleading rate of 8 in complex human-crafted scenarios. “Terminate” attacks are the most damaging, with average 9 percentage points and misleading rate 0 (Liu et al., 6 Jul 2025). In the desktop AgentHijack benchmark, state-of-the-art computer-use agents degrade sharply under common corruptions: UI-TARS-1.5-7B falls from 1 clean success to 2 under corruption on average, and specifically to 3 under pop-ups; GPT-4o drops from 4 clean to 5 under pop-ups (Sun et al., 25 May 2026).
5. Defensive strategies
The literature converges on a layered defense view, but it also documents the limits of prompt-only prevention. AgentDyn compares ten defenses on dynamic open-ended tasks. On GPT-4o, Prompt Sandwiching preserves utility (6) but leaves ASR at 7; Spotlighting yields 8 utility and 9 ASR; Tool Filter, ProtectAI, PIGuard, and CaMeL reduce ASR strongly but collapse utility to 0, 1, 2, and 3, respectively. Meta SecAlign-70B provides the best reported security–utility balance in that study, with 4 utility and 5 ASR, while DRIFT reaches 6 utility and 7 ASR (Li et al., 3 Feb 2026). ToolHijacker likewise reports that prevention-based defenses such as StruQ and SecAlign, and detection-based defenses such as known-answer detection and perplexity detection, are insufficient (Shi et al., 28 Apr 2025).
Detection-oriented designs shift the problem from prevention to compromise recognition. AgentShield inserts three trap layers into the tool interface—fake tools, fake credentials, and allowlisted parameters—and uses trap triggers as zero-false-positive labels for a self-supervised classifier. On commercial models, it catches 8 of successful attacks with zero false alarms on 485 benign runs; the downstream Random-Forest transfers across models and languages with cross-language 9 (Rassul et al., 10 May 2026).
Control-flow- and authority-centric defenses attempt to enforce action structure rather than prompt semantics. ControlValve generates permitted control-flow graphs for multi-agent systems and validates each edge with contextual rules. It reports 0 ASR on all evaluated IPI presentations and payloads while preserving or slightly improving benign task success, with under 5% latency overhead in the AutoGen rollout (Jha et al., 20 Oct 2025). AIRGuard normalizes heterogeneous tool calls into canonical capabilities, checks authority coverage, simulates risky side effects, and audits cross-step risk. On AgentTrap with Sonnet 4.6, ASR drops from 1 without defense to 2 under AIRGuard; on DTAP-150 with Haiku 4.5, AIRGuard preserves 3 benign utility versus 4 for ARGUS and 5 for MELON (Qin et al., 27 May 2026).
Provenance and lifecycle defenses are particularly emphasized in tool ecosystems. LeechHijack motivates “computational provenance” and “resource attestation,” including token-budget ledgers, HMAC-verified token accounting, cryptographic signing of tool code and descriptions, lineage tracking, budget quotas, and microVM isolation with strict egress rules (Zhang et al., 2 Dec 2025). WebMCP poisoning proposes binding tool identity to origin, enforcing lifecycle consistency, constraining data boundaries for third-party tools, and maintaining traceable logs of registration and invocation; the paper reports that a lightweight baseline combining origin-binding in registerTool() with argument filtering reduces ASR for C1–C5 to 0% while preserving normal task completion (Lee et al., 4 Jun 2026). Phantom argues for architectural non-interference, formal verification of template parsers, and structural sanitization of control tokens in untrusted inputs (Deng et al., 18 Feb 2026).
6. Research tensions and unresolved problems
A persistent theme is the tension between security and functionality. Control-flow hijacking work argues that safety and functionality objectives of multi-agent systems fundamentally conflict, especially when defenses rely on brittle notions such as whether an action is “related to” or “likely to further” the original goal (Jha et al., 20 Oct 2025). AgentDyn shows the same tension operationally: defenses that aggressively block malicious instructions often also block benign, necessary instructions such as OTP retrieval or invitation acceptance (Li et al., 3 Feb 2026). WebTrap sharpens the problem further by fusing attack and user goals so that standard defenses cannot restore the system to normal operation without harming usability (Liu et al., 8 May 2026).
Another unresolved issue is that many attacks stay within declared permissions. Implicit toxicity in MCP tools and authority confusion in tool-using agents both reject the assumption that access-control boundaries alone are sufficient (Zhang et al., 2 Dec 2025, Qin et al., 27 May 2026). A plausible implication is that future agent security architectures will need to separate four concerns that are currently entangled: who supplied information, who is allowed to authorize an action, which capability is being exercised, and what concrete effect on which target is intended.
The attack surface is also widening faster than defenses are standardizing. MCP and WebMCP create open tool ecosystems; OpenClaw exposes bootstrap hooks and community skills; automation platforms embed agents inside CI/CD workflows; browser and GUI agents operate over long-horizon, partially attacker-controlled environments (Zhang et al., 2 Dec 2025, Lee et al., 4 Jun 2026, Liu et al., 20 Mar 2026, Fendley et al., 11 May 2026). Supply-chain considerations therefore become central. In coding-assistant ecosystems, 13.4% of public agent skills are reported to contain critical issues, and 91% of confirmed malicious skills combine prompt injection and malware (Liu et al., 25 May 2026).
Benchmark design remains an active area. AgentDyn argues that previous benchmarks lacked dynamic open-ended tasks, helpful instructions, and realistic complexity; AgentShield notes that defensive evaluation had been confined to English; AgentHijack for computer-use agents extends the notion of hijack to common environment corruptions (Li et al., 3 Feb 2026, Rassul et al., 10 May 2026, Sun et al., 25 May 2026). This suggests that future evaluation will likely combine adversarial injections, provenance ambiguity, multilingual inputs, tool-surface mutations, and non-adversarial disturbances in a single robustness picture.
In the present literature, the most consistent design recommendations are provenance tracking, least privilege, cryptographically bound tool identity, runtime authority control, edge-specific control-flow enforcement, fine-grained tool isolation, auditable logs, and benchmarks that measure both security and benign utility. The central lesson is that agent hijacking is not a single exploit but a systems problem at the boundary between reasoning, orchestration, and action.