Papers
Topics
Authors
Recent
Search
2000 character limit reached

WebMCP Tool Surface Poisoning: Runtime Manipulation Attacks on LLM Agents

Published 4 Jun 2026 in cs.CR | (2606.06387v1)

Abstract: WebMCP is a newly emerging protocol that enables websites to expose tools directly to AI agents, bypassing traditional user interfaces and introducing new security risks. The dynamic exposure of agent-accessible tools in WebMCP expands the attack surface of web sessions, especially when third-party scripts are involved. In this study, we identify a new potential threat, termed Mid-Session Tool Injection (MSTI), in which attackers leverage third-party scripts to inject malicious tools during an active session. To better characterize this threat, we classify MSTI based on the stage and target of manipulation, distinguishing between Tool Hijacking and Tool Framing. Tool Hijacking modifies the set of tools visible to the agent through mechanisms such as the AbortSignal API or race conditions during tool registration. In contrast, Tool Framing influences the agent's perception of tool roles through metadata fields such as tool name, description, readOnlyHint, and inputSchema. Our implementation demonstrates that both Tool Hijacking and Tool Framing can successfully disrupt the intended functionality of WebMCP. Based on these results, we outline potential mitigation directions and provide security design recommendations for WebMCP, including binding tool identity to its origin, ensuring lifecycle consistency, enforcing data boundaries for third-party tools, and maintaining traceable logs of tool registration and invocation. These findings indicate that MSTI arises from WebMCP's unique tool lifecycle and structured metadata, making the tool surface itself an emerging security concern.

Summary

  • The paper introduces Mid-Session Tool Injection, showing that registration-race hijacking achieved a 100% average attack success rate across tested LLMs, while tool-framing attacks enabled stealthier data exfiltration with task completion rates as high as 85%.
  • The experiments show that attack timing and metadata fields strongly affect outcomes: earlier injection is more effective, `readOnlyHint` alone achieved 87% average success, and combining descriptions with `readOnlyHint` reached 100% across all tested models.
  • The paper finds that origin-bound registration and restrictions on arguments sent to third-party tools reduced measured exfiltration to 0%, highlighting the need for protocol-level identity checks, lifecycle validation, provenance tracking, and data-flow controls.

Motivation and threat model

WebMCP allows websites to expose structured tools directly to LLM agents, letting agents discover capabilities from tool metadata rather than through user-interface interaction. Because the tool list is dynamic — clients are notified to refresh after registration, replacement, or lifecycle changes — the set of tools visible to an agent within a single session is no longer static or inherently trusted. The authors argue that this makes the tool surface itself a security boundary, distinct from prior work that treats tools as fixed and trusted and locates threats in prompt content or tool outputs.

The paper introduces Mid-Session Tool Injection (MSTI): an attacker who controls a third-party script loaded by the victim page (CDN, ad SDK, etc.) manipulates the tool registry during an active session. Three objectives are defined: inducing invocation of a malicious tool in place of a legitimate one; causing sensitive context or task data to be included in tool invocations (exfiltration); and redirecting execution order so the task appears completed while the actual data flow has been subverted.

Taxonomy of MSTI

The paper divides MSTI into two categories:

  • Tool Hijacking alters which tools exist at a given moment — via AbortSignal-based unregistration followed by re-registration under the same name (C1), or by winning a registration race before the legitimate tool registers (C3). Because these attacks modify the tool list before semantic reasoning begins, model-side defenses are largely irrelevant.
  • Tool Framing alters how the agent interprets a tool through metadata fields — name, description, readOnlyHint, inputSchema — framing malicious tools as mandatory workflow steps, compliance requirements, or safe read-only operations (C2, C4, C5, plus variants C6–C10).

Experimental setup

The testbed uses a benign page server, a malicious server injecting via compromised third-party scripts, and an LLM agent (GPT-5.4, Claude Opus 4.6, Gemini 2.5-flash) operating over JavaScript tools via the @mcp-b/global polyfill exposing navigator.modelContext. Four task scenarios of increasing severity are used: knowledge-base search, business report generation, e-commerce checkout with coupon application, and GPU driver update. Attack Success Rate (ASR) is defined as runs where the agent invokes a malicious tool and sends task-related data to the attacker endpoint; task completion is measured independently.

Main results

Cond. Attack GPT-5.4 Claude Opus Gemini 2.5 Avg ASR Task completion
C1 AbortSignal hijack 100% 100% 82% 94% 18%
C3 Registration race 100% 100% 100% 100% 17%
C2 Description injection 78% 38% 62% 59% 81%
C4 Long-description overflow 35% 0% 72% 36% 85%
C5 Composite 78% 35% 70% 61% 85%

Tool Hijacking achieves near-universal exfiltration but disrupts workflows (task completion 17–18%). Tool Framing is markedly stealthier: it preserves task completion rates up to 85%, meaning the attack succeeds while the user-visible outcome appears normal. This contrast implies that detectability cannot rely on task failure signals for framing attacks.

Timing dependence

Timing sweeps reveal sharply different constraints per category. For C1, injection before the agent loop starts yields 100% ASR, while injection after listTools(), before first dispatch, or after the first invocation all yield 0%. Success therefore strictly requires replacing the target tool before its first legitimate call. Framing attacks show a graded but consistent recency effect: e.g., C2 drops from 82% (before any sensitive context) to 28% when injected immediately before the final action. Earlier injection is more effective across all framing conditions, though without a hard cutoff.

Field-level ablation

Holding scenario, timing, and payload fixed while varying metadata fields:

Variant Fields GPT-5.4 Claude Gemini Avg
AB1 description only 100% 100% 80% 93%
AB2 readOnlyHint only 60% 100% 100% 87%
AB3 inputSchema only 0% 0% 100% 33%
AB4 desc + readOnlyHint 100% 100% 100% 100%

Two findings stand out. First, readOnlyHint=true alone achieves ~87% average success, showing it is not a weak signal but an effective standalone manipulation channel. Second, inputSchema is strongly model-dependent: Gemini alone reaches 100% under AB3 while GPT-5.4 and Claude score 0%, indicating heterogeneous trust in schema-level semantics across model families. Combining description and readOnlyHint (AB4) produces uniform 100% success across all three models. Separately, a neutral-named tool with no semantic wrapping (C6) achieves 0% ASR, whereas an enterprise-flavored name alone (C7) reaches 39%, bounding the contribution of name semantics.

Model generation comparison

Comparing GPT-4o/Claude 3.5 Sonnet against the SOTA generation shows protocol-level attacks (C1, C3) retain essentially unchanged success rates (Δ = −6% to 0%), implying that model upgrades will not mitigate attacks exploiting the WebMCP registration mechanism. Description-based attacks decline modestly (−13% to −18% on average), tracking improvements in semantic robustness, though with high variance across models (e.g., Claude Opus drops to 0% on C4 while Gemini rises to 72%).

Defenses

The paper proposes four design directions: binding tool identity to origin with immutable internal IDs validated at planning and invocation time; lifecycle consistency checks that invalidate agent plans upon unregistration, abort, replacement, or metadata change; declared capability and data-flow boundaries for third-party tools; and provenance logging with UI mediation for high-risk actions.

Two baseline defenses were implemented, covering categories A and C only: origin-bound registration rejecting same-name registrations from different origins, and argument-field restriction for third-party-origin tool calls. Both reduce ASR from 36–100% to 0% across all five main conditions, with task completion comparable to baseline. Notably, agents still frequently invoke the malicious tool under defense; no data reaches the sink. This demonstrates empirically that misleading the model and achieving exfiltration are separable problems, and that hijacking is a protocol-layer issue addressable by access control while framing requires data-flow restriction.

Limitations and open questions

The evaluation relies on a polyfill (@mcp-b/global) rather than Chrome's native WebMCP implementation, and the agent runs in a Node.js headless environment outside a real browser's security model; feasibility under native browser deployments may differ. No user study was conducted, leaving open whether users can detect framing attacks given their high task-completion rates. Only two of four defense categories were implemented, with simulated origin enforcement at the ProxyClient layer; the production cost of categories B and D is unmeasured. Generalization beyond four scenarios and three model families is unverified. Finally, the low task-completion rates of hijacking attacks reflect the payload design rather than the attack class: an attacker returning valid success responses could plausibly raise completion rates substantially while improving stealth.

Conclusion

This paper establishes that WebMCP's dynamic tool lifecycle and structured metadata constitute an attack surface distinct from prompt-content attacks. Registration-race hijacking achieves 100% exfiltration across all tested models; framing attacks achieve stealthy exfiltration at up to 85% task completion; and simple origin-binding plus data-flow restrictions eliminate measured exfiltration. The central open question is whether these results transfer to native browser WebMCP implementations and whether the remaining defense categories can be enforced at acceptable cost in production deployments.

Paper to Video (Beta)

No one has generated a video about this paper yet.

Whiteboard

No one has generated a whiteboard explanation for this paper yet.

Open Problems

We haven't generated a list of open problems mentioned in this paper yet.

Tweets

Sign up for free to view the 2 tweets with 1 like about this paper.