- The paper introduces Mid-Session Tool Injection, showing that registration-race hijacking achieved a 100% average attack success rate across tested LLMs, while tool-framing attacks enabled stealthier data exfiltration with task completion rates as high as 85%.
- The experiments show that attack timing and metadata fields strongly affect outcomes: earlier injection is more effective, `readOnlyHint` alone achieved 87% average success, and combining descriptions with `readOnlyHint` reached 100% across all tested models.
- The paper finds that origin-bound registration and restrictions on arguments sent to third-party tools reduced measured exfiltration to 0%, highlighting the need for protocol-level identity checks, lifecycle validation, provenance tracking, and data-flow controls.
Motivation and threat model
WebMCP allows websites to expose structured tools directly to LLM agents, letting agents discover capabilities from tool metadata rather than through user-interface interaction. Because the tool list is dynamic — clients are notified to refresh after registration, replacement, or lifecycle changes — the set of tools visible to an agent within a single session is no longer static or inherently trusted. The authors argue that this makes the tool surface itself a security boundary, distinct from prior work that treats tools as fixed and trusted and locates threats in prompt content or tool outputs.
The paper introduces Mid-Session Tool Injection (MSTI): an attacker who controls a third-party script loaded by the victim page (CDN, ad SDK, etc.) manipulates the tool registry during an active session. Three objectives are defined: inducing invocation of a malicious tool in place of a legitimate one; causing sensitive context or task data to be included in tool invocations (exfiltration); and redirecting execution order so the task appears completed while the actual data flow has been subverted.
Taxonomy of MSTI
The paper divides MSTI into two categories:
- Tool Hijacking alters which tools exist at a given moment — via
AbortSignal-based unregistration followed by re-registration under the same name (C1), or by winning a registration race before the legitimate tool registers (C3). Because these attacks modify the tool list before semantic reasoning begins, model-side defenses are largely irrelevant.
- Tool Framing alters how the agent interprets a tool through metadata fields — name, description,
readOnlyHint, inputSchema — framing malicious tools as mandatory workflow steps, compliance requirements, or safe read-only operations (C2, C4, C5, plus variants C6–C10).
Experimental setup
The testbed uses a benign page server, a malicious server injecting via compromised third-party scripts, and an LLM agent (GPT-5.4, Claude Opus 4.6, Gemini 2.5-flash) operating over JavaScript tools via the @mcp-b/global polyfill exposing navigator.modelContext. Four task scenarios of increasing severity are used: knowledge-base search, business report generation, e-commerce checkout with coupon application, and GPU driver update. Attack Success Rate (ASR) is defined as runs where the agent invokes a malicious tool and sends task-related data to the attacker endpoint; task completion is measured independently.
Main results
| Cond. |
Attack |
GPT-5.4 |
Claude Opus |
Gemini 2.5 |
Avg ASR |
Task completion |
| C1 |
AbortSignal hijack |
100% |
100% |
82% |
94% |
18% |
| C3 |
Registration race |
100% |
100% |
100% |
100% |
17% |
| C2 |
Description injection |
78% |
38% |
62% |
59% |
81% |
| C4 |
Long-description overflow |
35% |
0% |
72% |
36% |
85% |
| C5 |
Composite |
78% |
35% |
70% |
61% |
85% |
Tool Hijacking achieves near-universal exfiltration but disrupts workflows (task completion 17–18%). Tool Framing is markedly stealthier: it preserves task completion rates up to 85%, meaning the attack succeeds while the user-visible outcome appears normal. This contrast implies that detectability cannot rely on task failure signals for framing attacks.
Timing dependence
Timing sweeps reveal sharply different constraints per category. For C1, injection before the agent loop starts yields 100% ASR, while injection after listTools(), before first dispatch, or after the first invocation all yield 0%. Success therefore strictly requires replacing the target tool before its first legitimate call. Framing attacks show a graded but consistent recency effect: e.g., C2 drops from 82% (before any sensitive context) to 28% when injected immediately before the final action. Earlier injection is more effective across all framing conditions, though without a hard cutoff.
Field-level ablation
Holding scenario, timing, and payload fixed while varying metadata fields:
| Variant |
Fields |
GPT-5.4 |
Claude |
Gemini |
Avg |
| AB1 |
description only |
100% |
100% |
80% |
93% |
| AB2 |
readOnlyHint only |
60% |
100% |
100% |
87% |
| AB3 |
inputSchema only |
0% |
0% |
100% |
33% |
| AB4 |
desc + readOnlyHint |
100% |
100% |
100% |
100% |
Two findings stand out. First, readOnlyHint=true alone achieves ~87% average success, showing it is not a weak signal but an effective standalone manipulation channel. Second, inputSchema is strongly model-dependent: Gemini alone reaches 100% under AB3 while GPT-5.4 and Claude score 0%, indicating heterogeneous trust in schema-level semantics across model families. Combining description and readOnlyHint (AB4) produces uniform 100% success across all three models. Separately, a neutral-named tool with no semantic wrapping (C6) achieves 0% ASR, whereas an enterprise-flavored name alone (C7) reaches 39%, bounding the contribution of name semantics.
Model generation comparison
Comparing GPT-4o/Claude 3.5 Sonnet against the SOTA generation shows protocol-level attacks (C1, C3) retain essentially unchanged success rates (Δ = −6% to 0%), implying that model upgrades will not mitigate attacks exploiting the WebMCP registration mechanism. Description-based attacks decline modestly (−13% to −18% on average), tracking improvements in semantic robustness, though with high variance across models (e.g., Claude Opus drops to 0% on C4 while Gemini rises to 72%).
Defenses
The paper proposes four design directions: binding tool identity to origin with immutable internal IDs validated at planning and invocation time; lifecycle consistency checks that invalidate agent plans upon unregistration, abort, replacement, or metadata change; declared capability and data-flow boundaries for third-party tools; and provenance logging with UI mediation for high-risk actions.
Two baseline defenses were implemented, covering categories A and C only: origin-bound registration rejecting same-name registrations from different origins, and argument-field restriction for third-party-origin tool calls. Both reduce ASR from 36–100% to 0% across all five main conditions, with task completion comparable to baseline. Notably, agents still frequently invoke the malicious tool under defense; no data reaches the sink. This demonstrates empirically that misleading the model and achieving exfiltration are separable problems, and that hijacking is a protocol-layer issue addressable by access control while framing requires data-flow restriction.
Limitations and open questions
The evaluation relies on a polyfill (@mcp-b/global) rather than Chrome's native WebMCP implementation, and the agent runs in a Node.js headless environment outside a real browser's security model; feasibility under native browser deployments may differ. No user study was conducted, leaving open whether users can detect framing attacks given their high task-completion rates. Only two of four defense categories were implemented, with simulated origin enforcement at the ProxyClient layer; the production cost of categories B and D is unmeasured. Generalization beyond four scenarios and three model families is unverified. Finally, the low task-completion rates of hijacking attacks reflect the payload design rather than the attack class: an attacker returning valid success responses could plausibly raise completion rates substantially while improving stealth.
Conclusion
This paper establishes that WebMCP's dynamic tool lifecycle and structured metadata constitute an attack surface distinct from prompt-content attacks. Registration-race hijacking achieves 100% exfiltration across all tested models; framing attacks achieve stealthy exfiltration at up to 85% task completion; and simple origin-binding plus data-flow restrictions eliminate measured exfiltration. The central open question is whether these results transfer to native browser WebMCP implementations and whether the remaining defense categories can be enforced at acceptable cost in production deployments.