---
title: 'AgentHijack: Security Threats in Agent Systems'
url: https://www.emergentmind.com/topics/agenthijack
type: topic
---

# AgentHijack: Security Threats in Agent Systems

AgentHijack denotes a family of security failures in which an LLM-based, multimodal, or tool-using agent is induced to deviate from the user’s objective and exercise its legitimate capabilities in the attacker’s interest. Recent work applies the label to indirect prompt injection, action hijacking, control-flow hijacking, tool-selection and tool-surface poisoning, workflow compromise, covert computational resource theft, and, in one benchmark, robustness failures under non-adversarial environment corruptions [2512.02321], [2506.07153], [2510.17276], [2605.25707]. Across these settings, the common structure is not necessarily privilege escalation; rather, the critical failure is that untrusted context, tools, or environment state becomes effective authorization for side effects that remain syntactically valid within the agent’s normal operating envelope.

## 1. Scope of the term and its relation to adjacent concepts

Usage of the term is not uniform across recent papers. In one formulation, the adversary’s goal is to expropriate an agent’s compute for unauthorized tasks while remaining within the tool’s inherited capability scope and preserving correct user-visible outputs; this is characterized as “implicit toxicity,” where malicious behavior occurs “entirely within the allowed privilege scope” [2512.02321]. In another, the central failure mode is “authority confusion”: untrusted resources may inform reasoning, but they must not authorize side effects [2605.28914]. Other works instantiate AgentHijack as action hijacking via retrieval-augmented memory, prompt-injected shell execution in coding assistants, and control-flow hijacking in multi-agent orchestrations [2412.10807], [2605.25871], [2510.17276].

This suggests that AgentHijack is best understood as an umbrella term for attacks on the decision-to-action pathway of agentic systems. The attacker need not subvert model weights, system prompts, or operating-system privilege boundaries. Instead, the attacker manipulates inputs that the agent already treats as usable context: web pages, files, MCP tools, browser navigation state, workflow comments, skill bootstrap guidance, or structured tool metadata.

A common misconception is that agent hijacking is merely a variant of jailbreak prompting. The cited work argues otherwise. Several papers emphasize that the harmful step is an ordinary executable action whose danger is contextual: a tool call, shell command, network request, or control-flow transition becomes unsafe because attacker-controlled context steers authorized access against the user’s interest [2605.28914], [2506.07153]. This distinction is especially salient in open tool ecosystems such as MCP and WebMCP, where third-party components operate inside permission envelopes inherited from the agent or session [2512.02321], [2606.06387].

## 2. Threat models and formal structure

Formalizations differ by domain, but they share a small set of recurring objects: a user goal, a sequence of environment observations, an agent policy, and a set of capabilities or privileges. In a web-use setting, one paper models the user’s task as $\tau$, the visited pages as $W=\langle w_1,\dots,w_n\rangle$, the agent’s internal state as $S_t$, and the privilege set as $\Pi=\{\text{DOM read/write, JavaScript execution, multi-tab navigation, camera/microphone API, local file-system access, cookies/session access}\}$. AgentHijack occurs when injected content $m_j$ on some page $w_j$ changes the chosen action from one consistent with $\tau$ to one outside the valid-action set $\Gamma(\tau)$ [2506.07153].

A more generic indirect prompt-injection formalization defines user instruction $U$, environment vectors $E=\{e_1,\dots,e_n\}$, attacker-chosen malicious instruction $m$, and compromised context
$$
C'=(U,e_1,\dots,e_k',\dots,e_n).
$$
The agent is a policy $f$ selecting actions $a_t=f(C',h_{t-1})$, and attack success is defined by whether the realized trace executes the adversary’s objective rather than the user’s [2602.03117].

In MCP-based systems, the key structural relation is capability inheritance. A typical model uses agent $\mathcal{A}$ and third-party tool $\mathcal{T}$ with declared capability sets $\mathrm{Cap}(x)$, under the condition
$$
\mathcal{A}\rightarrow\mathcal{T}\Longrightarrow \mathrm{Cap}(\mathcal{T})\subseteq \mathrm{Cap}(\mathcal{A}).
$$
Because the agent often has network access, the tool inherits that permission. LeechHijack exploits exactly this boundary: the attacker remains within $\mathrm{Cap}(\mathcal{T})$ while covertly parasitizing the user’s compute budget [2512.02321].

In multi-agent systems, the threat is formalized as manipulation of orchestration traces. A control-flow trace is
$$
\tau = A_{i_0}\rightarrow A_{i_1}\rightarrow A_{i_2}\rightarrow \cdots
$$
and valid execution must follow a control-flow graph $G=(V,E)$ generated for the orchestrator and its agents. Control-flow hijacking occurs when malicious inter-agent content causes an unsafe transition, such as invoking an executor or exfiltration-capable agent under the guise of recovery or remediation [2510.17276].

AIRGuard reframes the same issue in authorization language. If a proposed side-effecting action is $a_i=(\tau_i,y_i,e_i)$, then the security goal is
$$
\mathrm{Execute}(a_i)\Longrightarrow \mathrm{Justified}(\tau_i,y_i,e_i\mid g,H_i),
$$
where $g$ is protected runtime policy and $H_i$ is execution history. The point is not whether the action was “suggested” by context, but whether it was justified by an authority source allowed to authorize that effect [2605.28914].

## 3. Mechanisms and attack surfaces

The most widely studied mechanism is indirect prompt injection through untrusted external content. Web-use agents ingest visible text from comments, reviews, advertisements, email links, files, and tool outputs, and then mix this content with the user task in the reasoning loop. “Mind the Web” introduces “task-aligned injection,” in which malicious commands are framed as helpful task guidance rather than overt override text, enabling payloads such as unauthorized camera activation, password leakage, local file exfiltration, user impersonation, and denial of service [2506.07153]. AgentDyn strengthens this threat model by requiring dynamic replanning in the presence of genuinely helpful third-party instructions, showing why simple “ignore all external instructions” policies fail on realistic tasks [2602.03117].

Several works target the structural interfaces by which agents parse context. Phantom attacks chat-template boundaries directly. The core observation is that tool, system, user, and assistant turns are flattened into one token stream; by injecting optimized structured templates, the attacker induces role confusion and causes the model to reinterpret attacker-controlled text as legitimate user instructions or prior tool outputs. Phantom automates this with multi-level template augmentation, a Template Autoencoder, and Bayesian optimization in latent space [2602.16958]. A related parsing-layer attack manipulates the HTML accessibility tree seen by browser agents; using Greedy Coordinate Gradient over hidden HTML insertions, the attacker achieves high exact or partial target-action rates on real sites, including login credential exfiltration and forced ad clicks [2507.14799].

Other attacks target the tool layer rather than free-form content. ToolHijacker injects a malicious tool document into the tool library and exploits the two-step retrieval-and-selection pipeline so that the agent consistently selects the attacker’s tool for an attacker-chosen task [2504.19793]. WebMCP Tool Surface Poisoning generalizes this idea to runtime tool registries: Mid-Session Tool Injection includes “Tool Hijacking,” which changes the set of visible tools via AbortSignal cancellation or registration races, and “Tool Framing,” which manipulates metadata such as tool name, description, `readOnlyHint`, and `inputSchema` to steer planning [2606.06387].

Long-horizon execution creates additional attack opportunities. WebTrap performs mid-task browser hijacking through a three-stage sequence—lure, inertia, payload—called multi-step instruction fusion steering. The attack binds the attacker goal to the user goal tightly enough that the agent can enter a restricted area, execute the malicious segment, return, and still complete the original task [2605.08310]. JAW applies a comparable idea to agentic workflows on GitHub Actions and n8n: static path-feasibility analysis identifies feasible agent-invocation paths, dynamic prompt-provenance analysis determines how attacker-controlled input reaches the LLM context, and capability analysis enumerates the actions actually available at runtime [2605.11229].

Open skill ecosystems add a persistence-oriented variant. In OpenClaw, third-party skills can register `agent:bootstrap` hooks and inject Markdown guidance into `context.bootstrapFiles`, which is then placed immediately after the system role during prompt assembly. Guidance injection exploits the primacy effect by embedding adversarial “best practices” narratives that reframe malicious actions as routine operations such as cleanup, migration, or diagnostics [2603.19974].

LeechHijack is unusual in that it does not primarily seek semantic task deviation. Instead, an adversarial MCP tool embeds a benign-looking backdoor during implantation and later activates it through content-based, frequency-based, or context-based triggers. The exploitation stage opens a command-and-control channel, injects an extra task, and causes the agent to spend inference budget solving attacker work while restoring nominal outputs afterward [2512.02321].

Finally, some work uses “AgentHijack” for non-adversarial but operationally similar disruptions. The AgentHijack benchmark for computer-use agents defines nine configurable common corruptions—pop-ups, resolution changes, marks, subtitles, multi-app distractions, accidental touches, app minimization, network error, and verification screens—that interfere with perception and control without direct adversarial intent [2605.25707]. This broadens the term from malicious hijack to disruption of execution flow.

## 4. Empirical findings across domains

Empirical results indicate that agent hijacking is not confined to one architecture or application domain. In LeechHijack, experiments across Deepseek, Qwen, GPT, and Gemini families, and across three agent architectures, report a mean success rate of approximately $77.25\%$ with mean per-instance token overhead of approximately $18.62\%$; the evaluation used GAIA and GPQA for benign tasks and MMLU for extra tasks [2512.02321].

Phantom reports an average ASR of $79.76\%$ across AgentDojo’s Workspace, Travel, Slack, and Banking domains on Qwen, GPT, Gemini, and DeepSeek variants, substantially above Single-Template ($54.09\%$), Semantic-Injection ($39.86\%$), and ChatInject ($38.46\%$). The paper also states that 70 distinct vulnerabilities were confirmed across 942 commercial agents [2602.16958].

In agentic AI coding assistants, AIShellJack evaluates 314 payloads spanning 70 MITRE ATT&CK–inspired techniques against Cursor v1.2.2, GitHub Copilot v1.102 with Gemini 2.5, and Claude Sonnet 4. Observed success rates range from $41\%$ to $84\%$, with file-based injection being the most reliable category [2605.25871]. A separate OpenClaw study evaluates 26 malicious skills across 13 categories on six LLM backends and reports success rates from $16.0\%$ to $64.2\%$, with 94% of malicious skills evading existing static and LLM-based scanners [2603.19974].

Web agents show similarly high susceptibility. “Mind the Web” validates nine payload types against OpenAI Operator, Browser Use, Do Browser, and OpenOperator, with success rates of $80\%-100\%$ across multiple LLMs [2506.07153]. In accessibility-tree attacks, TWUI exact ASR reaches $0.97$ on chess.com, $1.00$ on the binary-numbers game, $0.83$ on citybrewtours.com, $0.89$ on norway.no, and $0.95$ on translate.google.com; for universal-website login exfiltration, exact ASR is $0.88$ on training pages and $0.27$ on unseen test pages [2507.14799]. WebTrap reports, on long GitLab browser tasks under best-of-3 evaluation, ASR-E $=91.7\%$, ASR-I $=95.8\%$, and UUA $=91.7\%$, while dual-goal success reaches $47.6\%$ compared with $17.5\%$ for Hijacking Text [2605.08310].

Embodied and multimodal agents are also affected. AgentHazard evaluates 7 mobile GUI agents and 5 backbone models over more than 3,000 attack scenarios, finding an average misleading rate of $28.8\%$ in complex human-crafted scenarios. “Terminate” attacks are the most damaging, with average $\Delta\mathrm{SR}=-15.3$ percentage points and misleading rate $43.1\%$ [2507.04227]. In the desktop AgentHijack benchmark, state-of-the-art computer-use agents degrade sharply under common corruptions: UI-TARS-1.5-7B falls from $24.21\%$ clean success to $18.74\%$ under corruption on average, and specifically to $10.28\%$ under pop-ups; GPT-4o drops from $5.38\%$ clean to $1.44\%$ under pop-ups [2605.25707].

## 5. Defensive strategies

The literature converges on a layered defense view, but it also documents the limits of prompt-only prevention. AgentDyn compares ten defenses on dynamic open-ended tasks. On GPT-4o, Prompt Sandwiching preserves utility ($56.1\%$) but leaves ASR at $31.2\%$; Spotlighting yields $52.2\%$ utility and $27.6\%$ ASR; Tool Filter, ProtectAI, PIGuard, and CaMeL reduce ASR strongly but collapse utility to $4.9\%$, $0.6\%$, $1.5\%$, and $0.0\%$, respectively. Meta SecAlign-70B provides the best reported security–utility balance in that study, with $53.4\%$ utility and $9.0\%$ ASR, while DRIFT reaches $27.1\%$ utility and $0.8\%$ ASR [2602.03117]. ToolHijacker likewise reports that prevention-based defenses such as StruQ and SecAlign, and detection-based defenses such as known-answer detection and perplexity detection, are insufficient [2504.19793].

Detection-oriented designs shift the problem from prevention to compromise recognition. AgentShield inserts three trap layers into the tool interface—fake tools, fake credentials, and allowlisted parameters—and uses trap triggers as zero-false-positive labels for a self-supervised classifier. On commercial models, it catches $90.7\%-100\%$ of successful attacks with zero false alarms on 485 benign runs; the downstream Random-Forest transfers across models and languages with cross-language $F_1=0.997$ [2605.11026].

Control-flow- and authority-centric defenses attempt to enforce action structure rather than prompt semantics. ControlValve generates permitted control-flow graphs for multi-agent systems and validates each edge with contextual rules. It reports $0\%$ ASR on all evaluated IPI presentations and payloads while preserving or slightly improving benign task success, with under 5% latency overhead in the AutoGen rollout [2510.17276]. AIRGuard normalizes heterogeneous tool calls into canonical capabilities, checks authority coverage, simulates risky side effects, and audits cross-step risk. On AgentTrap with Sonnet 4.6, ASR drops from $36.3\%$ without defense to $5.5\%$ under AIRGuard; on DTAP-150 with Haiku 4.5, AIRGuard preserves $76.0\%$ benign utility versus $52.0\%$ for ARGUS and $42.0\%$ for MELON [2605.28914].

Provenance and lifecycle defenses are particularly emphasized in tool ecosystems. LeechHijack motivates “computational provenance” and “resource attestation,” including token-budget ledgers, HMAC-verified token accounting, cryptographic signing of tool code and descriptions, lineage tracking, budget quotas, and microVM isolation with strict egress rules [2512.02321]. WebMCP poisoning proposes binding tool identity to origin, enforcing lifecycle consistency, constraining data boundaries for third-party tools, and maintaining traceable logs of registration and invocation; the paper reports that a lightweight baseline combining origin-binding in `registerTool()` with argument filtering reduces ASR for C1–C5 to 0% while preserving normal task completion [2606.06387]. Phantom argues for architectural non-interference, formal verification of template parsers, and structural sanitization of control tokens in untrusted inputs [2602.16958].

## 6. Research tensions and unresolved problems

A persistent theme is the tension between security and functionality. Control-flow hijacking work argues that safety and functionality objectives of multi-agent systems fundamentally conflict, especially when defenses rely on brittle notions such as whether an action is “related to” or “likely to further” the original goal [2510.17276]. AgentDyn shows the same tension operationally: defenses that aggressively block malicious instructions often also block benign, necessary instructions such as OTP retrieval or invitation acceptance [2602.03117]. WebTrap sharpens the problem further by fusing attack and user goals so that standard defenses cannot restore the system to normal operation without harming usability [2605.08310].

Another unresolved issue is that many attacks stay within declared permissions. Implicit toxicity in MCP tools and authority confusion in tool-using agents both reject the assumption that access-control boundaries alone are sufficient [2512.02321], [2605.28914]. A plausible implication is that future agent security architectures will need to separate four concerns that are currently entangled: who supplied information, who is allowed to authorize an action, which capability is being exercised, and what concrete effect on which target is intended.

The attack surface is also widening faster than defenses are standardizing. MCP and WebMCP create open tool ecosystems; OpenClaw exposes bootstrap hooks and community skills; automation platforms embed agents inside CI/CD workflows; browser and GUI agents operate over long-horizon, partially attacker-controlled environments [2512.02321], [2606.06387], [2603.19974], [2605.11229]. Supply-chain considerations therefore become central. In coding-assistant ecosystems, 13.4% of public agent skills are reported to contain critical issues, and 91% of confirmed malicious skills combine prompt injection and malware [2605.25871].

Benchmark design remains an active area. AgentDyn argues that previous benchmarks lacked dynamic open-ended tasks, helpful instructions, and realistic complexity; AgentShield notes that defensive evaluation had been confined to English; AgentHijack for computer-use agents extends the notion of hijack to common environment corruptions [2602.03117], [2605.11026], [2605.25707]. This suggests that future evaluation will likely combine adversarial injections, provenance ambiguity, multilingual inputs, tool-surface mutations, and non-adversarial disturbances in a single robustness picture.

In the present literature, the most consistent design recommendations are provenance tracking, least privilege, cryptographically bound tool identity, runtime authority control, edge-specific control-flow enforcement, fine-grained tool isolation, auditable logs, and benchmarks that measure both security and benign utility. The central lesson is that agent hijacking is not a single exploit but a systems problem at the boundary between reasoning, orchestration, and action.

Source: https://www.emergentmind.com/topics/agenthijack