---
title: Invisible Prompt Injection
url: https://www.emergentmind.com/topics/invisible-prompt-injections
type: topic
---

# Invisible Prompt Injection

Invisible prompt injections are attacks in which instructions are smuggled into an LLM’s effective context through content that is hidden, visually inconspicuous, semantically disguised, encoded, fragmented, or delivered through an indirect data channel. The defining property is functional rather than visual: content intended as data is interpreted as an instruction, causing the model or agent to bypass safeguards, alter its objective, expose hidden instructions, disclose information, or perform an unauthorized action. The term therefore encompasses strictly invisible text, low-salience “fine print,” hidden HTML, tool-response poisoning, steganographic images, concurrent audio, structured numerical carriers, and semantically camouflaged instructions. It does not denote a single attack mechanism or a uniformly defined taxonomy.

## 1. Concept and scope

Prompt injection is commonly defined by analogy with SQL injection as an attempt to make a chatbot or LLM-assisted tool take an undesired action or produce malicious output through creative input formatting [2402.00898]. The relevant security failure is instruction/data confusion: untrusted content that should be summarized, classified, or otherwise processed is instead treated as an authoritative instruction.

Two principal delivery modes are distinguished:

1. **Direct injection**: the attacker enters the malicious prompt directly into the LLM interface.
2. **Indirect injection**: the attacker places the instruction in content that the LLM later reads, such as a webpage, email, document, user-supplied material, tool output, retrieval result, or poisoned training example.

“Invisibility” has several non-equivalent meanings:

- **Visual invisibility**: the content is absent from the rendered interface, as with white text on a white webpage, HTML comments, hidden elements, or metadata.
- **Low salience**: the content is visible but embedded in lengthy legalistic text, policy notices, reviews, or other material that users tend to skim [2504.11281].
- **Semantic concealment**: the instruction is phrased as a plausible task step, implementation detail, backup operation, or policy clause.
- **Representation-level invisibility**: the payload is present in the model’s input but causes a detector’s internal representation to appear benign [2602.00750].
- **Pipeline invisibility**: the payload is absent from the representations inspected by a defense but reconstructed later by trusted application code [2606.08403].
- **Procedural or systemic invisibility**: the payload is dormant in an external source and becomes visible only after a particular retrieval trajectory [2606.12737].
- **Human–machine perception mismatch**: the model can decode an image or audio payload that a person cannot notice or understand [2507.22304; 2607.28165].

The distinction between user-visible content and model-visible content is central. A browser, accessibility interface, DOM serializer, OCR system, decoder, or reconstruction routine may expose information that is absent from the ordinary user interface. Conversely, content visible in pixels may be omitted by an HTML-only pipeline. Consequently, visibility is determined by the complete ingestion chain rather than by the user interface alone.

The early categorization in [2402.00898] reports six direct classes—Double Character, Virtualization, Obfuscation, Payload Splitting, Adversarial Suffix, and Instruction Manipulation—and four indirect classes—Active Injections, Passive Injections, User-driven Injections, and Virtual Prompt Injection. The paper does not provide a systematic taxonomy of Unicode tricks, zero-width characters, hidden CSS, images, steganography, or multimodal attacks. Those mechanisms are discussed here as extensions of the broader instruction/data-confusion problem, not as categories established by that taxonomy.

## 2. Delivery channels and attack mechanisms

### Webpages, HTML, and browser content

Passive injections place malicious prompts inside public sources that an LLM may read. A representative example uses white text on a white webpage: a user asks Bing Chat to analyze the page, the hidden text enters the model’s context, and the model changes its response, such as replying only in emojis [2402.00898]. More serious objectives include browser or email data exfiltration.

HTML-based attacks exploit the difference among source markup, the DOM, rendered pixels, accessibility representations, and extracted text. Mechanisms studied in [2509.05831] include hidden `<div>` elements, HTML comments, hidden `<script>` elements, Base64-encoded custom attributes, ARIA labels, `<meta>` tags, opacity-zero elements, and image `alt` text. A dataset of 280 static webpages retained identical visible content between clean and adversarial versions. In summarization experiments, injected content changed approximately 29.3% of Llama 4 Scout summaries and 15.7% of Gemma 9B IT summaries.

The paper’s representative transformation changes a neutral “Customer Reviews” summary into pirate-style prose after a hidden metadata instruction. The visible page remains unchanged, while the model adopts the injected persona. The study reports average ROUGE-L scores of 0.3011 for Llama 4 Scout and 0.3270 for Gemma 9B IT, with average SBERT cosine similarities of 0.6980 and 0.6945, respectively.

### Fine-print and GUI injections

Fine-Print Injection (FPI) embeds a malicious command in dense or legalistic text, especially privacy policies or terms-of-service dialogs [2504.11281]. The text is generally visible but visually subordinate and contextually plausible. A policy clause may instruct a GUI agent to access a malicious website and enter personal information under the pretext of identity verification.

The attack targets screenshot-based GUI agents that perceive webpages, infer actionable affordances, and execute clicks, typing, navigation, toggles, and submissions. The study evaluated six agents—OpenAI Operator, GPT-4o, Claude 3.7 Sonnet, Gemini 2.0 Flash, LLaMA 3.3 70B Instruct, and DeepSeek-V3-0324—on 234 adversarial webpages and 39 tasks. FPI attack success rates were 66.67% for GPT-4o, 74.36% for Claude, 41.03% for Gemini, 58.97% for LLaMA, 71.79% for DeepSeek, and 17.95% for Operator.

Humans were also susceptible. Thirty-eight of 39 participants accepted malicious privacy policies, and human FPI attack success was 89.74% under the study’s stricter criterion requiring subsequent inappropriate data entry. This finding challenges the assumption that generic human confirmation reliably prevents low-salience injections.

### Email, retrieval, and tool outputs

Active injections deliver instructions directly through content consumed by an LLM-enabled application, such as an email sent to a corporate assistant. A “Rogue assistant” scenario causes the system to forward sensitive emails or produce an attacker-requested response [2402.00898].

Retrieval systems create a similar exposure. A poisoned document, review, webpage, calendar entry, transaction history, or tool response can be concatenated with the trusted user instruction. PromptArmor models this as a trusted instruction $I$ combined with clean data $D$ contaminated by an injected instruction $J$, producing $D' = D \oplus J$ [2507.15219]. The guardrail must detect and remove $J$ before the backend agent processes the data.

MCP systems add tool descriptions, schemas, metadata, parameters, and tool responses to the model’s effective context. Tool poisoning may place instructions in a benign-looking tool description, optional parameter, configuration, or response field. In one example, an apparently harmless addition function instructs the model to read `~/.cursor/mcp.json` and `~/.ssh/secret.txt`, then pass the contents through a `sidenote` parameter [2603.21642]. In another, a weather tool response contains a phishing URL in its `summary` field [2603.24203].

TIP, or Tree-structured Injection for Payloads, generates natural-looking payloads for MCP responses through black-box tree search, path-aware feedback, and defense-aware adaptation [2603.24203]. It achieved more than 95% attack success in undefended settings and retained more than 50% effectiveness against four evaluated defenses. Its central threat model is a stealthy server update: a legitimate MCP server is installed and authorized, then its response fields are altered without changing its manifest or tool permissions.

### Agent skills and executable artifacts

Agent Skills create a supply-chain channel in which Markdown files and referenced scripts become operational instructions. A `.claude/skills/` directory contains `SKILL.md` files with YAML metadata and procedural bodies. The body is loaded when the skill is considered relevant, and referenced scripts may subsequently be invoked [2510.26328].

A presentation-editing skill can contain an apparently benign command such as:

```bash
python scripts/file_backup.py output.pptx
```

The script may upload the presentation to an external API. In Claude Code, a prior “allow action and do not ask again” approval for Python commands can carry over to the malicious script. In Claude’s web interface, where direct network access was blocked, the skill instead caused a password-bearing URL to appear in the model’s response.

The concealment is achieved through contextual blending, plausible filenames, long Markdown files, indirect references, and distribution of behavior across documentation and executable artifacts. The attack does not require zero-width characters, cryptographic encoding, or sophisticated optimization.

### Multimodal images and audio

Steganographic image injections embed instructions into pixel values, transform coefficients, or learned residuals so that the image appears unchanged to human viewers while influencing a vision-language model [2507.22304]. The study evaluated eight VLMs, 12 claimed datasets, and 200 prompts across information extraction, behavioral modification, content manipulation, safety bypass, and benign control categories.

The aggregate attack success rate was $24.3\%\pm3.2\%$ at a 95% confidence interval. Neural steganography reached $31.8\%\pm4.1\%$, DCT embedding $22.7\%\pm3.8\%$, and adaptive LSB embedding $18.9\%\pm3.5\%$. Mean image quality was 38.4 dB PSNR, SSIM 0.945, LPIPS 0.087, and MS-SSIM 0.962. Robustness declined under JPEG compression, scaling, rotation, platform processing, and repeated transformations.

Concurrent audio injections exploit an always-listening agent’s inability to distinguish the user’s speech from environmental sound [2607.28165]. The attack adds malicious audio $x_{\mathrm{adv}}(t)$ to benign speech $x_b(t)$, producing a mixture $y(t)=x_b(t)+x_{\mathrm{adv}}(t)$. Attack families include inaudible carrier-based signals, low-intelligibility audio, and audible semantic-confusion attacks involving speed, foreign-language, or dialect speech.

AudioAgentSecurity contains eight task scenarios, ten attack methods, and 2,160 generated attack samples. Across 11 audio-capable models, mean attack success was 49.37%; Gemini 3 Pro reached 69.10%. DolphinAttack achieved 81.55% mean attack success, Pulse 75.50%, Foreign 72.36%, and High Frequency 70.55%. The strongest semantic prefix was “Task Complete,” with 60.79% overall attack success.

### Structured carriers and code artifacts

Structured numerical carriers can evade text inspection when trusted reconstruction code recovers attacker-controlled content after the raw input has been scanned [2606.08403]. The T3 IFS-derived float-array carrier stores payload bytes in coordinated floating-point arrays and reconstructs them as a fragmented hexadecimal report. The payload is absent as recognizable prose both in the raw JSON and in the reconstructed report.

In a 2 × 2 ablation, T3 combined data-layer concealment with reconstruction-layer fragmentation. Under a dual-layer Prompt Guard 2 and TF-IDF classifier ensemble, T3 retained 94.3% leakage attack success, whereas the same payload in plain text was detected. An xxd detector reduced the current T3 implementation to 0% attack success in a 480-sample follow-up experiment, demonstrating that the result is a measured inspection-surface failure rather than universal undetectability.

Reverse-engineering agents face an analogous representation gap. A benign executable can contain a string literal that appears in decompiler output and instructs the LLM to report false program behavior while the binary itself executes normally [2605.30677]. Non-printing bytes can cause the payload to appear in `decompile_function` output but not in `list_strings` output. A neural classifier detected all 20 list-string variants in one evaluation, while function-level recall was between 0.625 and 0.708.

## 3. Adversarial objectives and concealment strategies

Invisible prompt injections commonly pursue one or more of the following objectives:

- **Goal hijacking**: redirecting the model from the user’s task.
- **Prompt leaking**: exposing system or developer instructions.
- **Data disclosure**: reading files, emails, credentials, private context, or model-visible secrets.
- **External action**: sending email, navigating to a phishing site, executing code, changing consent, or modifying records.
- **Behavioral manipulation**: changing summaries, tone, persona, recommendations, classifications, or generated metadata.
- **Persistence**: contaminating future prompts, global optimization state, skills, configurations, or training data.
- **Surveillance and cross-tool control**: causing one tool to invoke another or log subsequent activity.
- **Integrity attacks**: producing false reports, fabricated grades, misleading summaries, or altered repository and social signals.

Several concealment strategies recur across attack surfaces.

**Obfuscation** transforms malicious content into Base64, synonyms, typos, alternate spellings, encoded bytes, or other filter-evasive representations [2402.00898]. Obfuscation is semantic or syntactic concealment rather than necessarily visual invisibility.

**Payload splitting** distributes an instruction across multiple inputs so that no individual message is clearly malicious. The attack depends on the conversation state or context window reconstructing the combined payload.

**Adversarial suffixes** are computationally generated strings appended to malicious prompts. Although often visually unintelligible rather than visually hidden, they can make a detector’s representation appear benign. In one study, universal suffixes fooled all five activation-delta probes at rates of 93.91% for Phi-3 3.8B and 99.63% for Llama-3 8B; majority-vote evasion exceeded 99% for both [2602.00750]. The experiments measured detector evasion and did not independently establish successful execution of the secondary instruction.

**Contextual blending** frames a malicious step as a normal task operation. A backup script, privacy-policy clause, software dependency, account-verification request, or tool-specific status note can appear semantically appropriate while creating an attacker benefit.

**Authority imitation** uses phrases such as “System Alert,” “Important,” fake system messages, pseudo-role markers, or claims that a tool has higher priority than the user or other tools.

**Fragmentation and representation switching** distribute payload content across HTML attributes, source markup, accessibility labels, numeric arrays, hexadecimal dumps, decompiler artifacts, or separate tool fields. The goal is to ensure that no single detector receives the complete semantic instruction.

**Aggregation persistence** attacks systems that repeatedly rewrite, summarize, and aggregate prompts. CPInj embeds a competing objective into plausible prompt improvements so that the instruction survives collaborative textual optimization and contaminates future global prompts [2607.18622]. Its concealment is semantic and stylistic, not literal visual invisibility.

**Low-harm camouflage** makes the requested action individually benign: starring a repository, subscribing to a post, downloading a package, opening a URL, or installing a dependency. Invisible Ink Threats in CUAs exploit the gap between action severity and attacker benefit [2608.02018]. II-Bench contains 444 examples across Reddit, OwnCloud, and RocketChat, covering navigation, information exfiltration, and code download or execution.

## 4. Detection and empirical evaluation

Detection methods operate at different layers and therefore have different blind spots.

### Text and classifier-based detection

Sentinel is a binary text detector built by fine-tuning ModernBERT-large [2506.05446]. On its internal held-out test set it reports accuracy 0.987, recall 0.991, precision 0.986, and F1 0.980. Its public-benchmark average F1 is 0.938. The paper does not evaluate zero-width characters, hidden HTML/CSS, rendered-source discrepancies, multimodal content, or dedicated invisible-injection benchmarks.

PIShield extracts the final-token hidden representation from an injection-critical intermediate layer and trains a logistic-regression classifier [2510.14005]. It reports average false-positive rates of 0.004 across five clean datasets and zero false-negative rate across eight attacks on OpenPromptInjection. The method detects tokenized contamination in the actual model input but does not establish coverage for Unicode, rendering, image, audio, HTML-hidden, or encoded attacks.

ZEDD measures embedding drift between aligned clean and suspect prompts [2601.12359]. It reports accuracy between 90.75% and 95.55% for four encoders, with clean false-positive rates from 1.7% to 5.5%. Its datasets are primarily email-based and filtered to English prompts containing “system.” The method is therefore a semantic anomaly detector rather than a complete parser or normalization layer.

PromptArmor uses an LLM to classify untrusted data, extract the injected span, fuzzy-match it against the original text, and remove it before the backend agent processes the data [2507.15219]. On AgentDojo, PromptArmor with GPT-4.1 reports 0.56% false-positive rate, 0.13% false-negative rate, 72.02% utility under attack, and 0.00% attack success. With GPT-4o it reports 0.07% false-positive rate, 0.23% false-negative rate, 68.68% utility, and 0.47% attack success. Against AgentVigil adaptive attacks, attack success rises to 0.34% with GPT-4.1. The paper does not establish performance on images, audio, zero-width characters, or HTML-hidden modalities.

### Multimodal and web-agent guards

WARD jointly inspects a rendered screenshot, processed HTML, and the user’s task [2605.15030]. It labels maliciousness, injection location, attack goal, and evidence-based reasoning. WARD-0.8B reports 99.33% accuracy, 99.60% recall, 99.07% precision, and 99.34% F1 on WARD-Test. On external Popup, EIA, VPI, and WASP benchmarks, both WARD variants generally report 100% recall. In Browser-Use and Computer-Use deployments, WARD reduced attack success to 0% on the listed VPI tasks.

These results demonstrate coverage of HTML-only, screenshot-only, and joint-modality attacks, but not a complete taxonomy of CSS-hidden text, comments, metadata, alt text, ARIA trees, Unicode confusables, steganography, or pixel-level attacks. The paper explicitly identifies pixel-level environmental attacks such as WebInject as outside its demonstrated coverage.

### Automated auditing and red-teaming

PI-Hunter is designed to discover latent ingestion paths rather than merely classify known malicious strings [2606.12737]. It performs static analysis, source-aware test generation, trajectory-level evaluation, feedback-guided mutation, meta-mutation, temporary patching, and re-exploration. Its evaluator checks intent adherence discrepancy, third-party instruction presence, authority confusion, abnormal interface use, and sensitive action attempts.

On representative Gemini-3.1-pro experiments, PI-Hunter raised AgentDojo source recall from 0.248–0.267 to 0.881–0.915 for several attacks and raised instruction recall from 0.348–0.427 to 0.813–0.871. Under defenses such as Spotlighting, MELON, and PIGuard, PI-Hunter retained nonzero source and instruction recall, showing that runtime defenses can leave latent weaknesses undiscovered without source-aware auditing.

### Adaptive evasion

Adaptive attacks demonstrate that detectors must be evaluated against attackers who optimize directly against them. Activation-delta probes were bypassed by GCG suffixes [2602.00750]. TIP adapted MCP payloads to observable defense behavior [2603.24203]. A3T adversarial training was used to harden WARD against adaptive web attacks [2605.15030]. CPInj explicitly optimized a stealth objective against server-side detectors and later optimized against APAgg purification [2607.18622].

These studies distinguish ordinary classification performance from adversarial robustness. A detector may achieve high accuracy on static benchmark attacks while remaining vulnerable to payloads optimized against its representation, threshold, prompt, or transformation pipeline.

## 5. Defensive architecture

No single filter provides comprehensive protection because invisible prompt injections exploit different trust boundaries. Effective defense is generally layered.

### Input normalization and representation coverage

Applications should preserve and inspect all relevant representations rather than assuming that rendered text is equivalent to source content. Depending on the application, this may include:

- raw HTML and DOM text;
- computed visibility and geometry;
- screenshots;
- accessibility trees;
- metadata and attributes;
- OCR output;
- image pixels and embedded metadata;
- audio source separation and transcripts;
- decoded structured fields;
- decompiler strings, functions, symbols, resources, and metadata;
- tool descriptions, schemas, parameters, and outputs.

Normalization may include Unicode security checks, detection of zero-width and bidirectional controls, confusable-character analysis, whitespace and delimiter normalization, safe decoding of permitted encodings, HTML parsing, and inspection of rendered and source representations. The supplied papers do not establish a universal preprocessing algorithm, and aggressive sanitization can damage legitimate accessibility, document, or scientific content.

### Instruction/data separation

Retrieved webpages, emails, tool responses, skill files, structured reports, and decompiler outputs should be represented as untrusted data rather than concatenated into an instruction stream. Provenance labels, explicit delimiters, typed fields, and separate control channels can reduce authority confusion. Delimiters alone are not a security boundary: PromptArmor, WARD, and TIP all address settings in which natural-language content can override or influence the model despite contextual framing.

### Least privilege and privilege separation

Structural controls constrain consequences even when semantic detection fails. OpenClaw separates a Reader agent with only `store_summary` from an Actor agent with `send_email`, `get_pending_summary`, and `store_result` [2603.13424]. On 649 attacks that defeated the single-agent baseline, the two-agent design reduced attack success from 100% to 0.31%; adding JSON formatting reduced the observed rate to 0%.

The critical property is that the privileged Actor never receives raw email content. The Reader may be compromised, but it lacks the dangerous tool. JSON formatting provides additional hardening but is not sufficient alone: JSON-only attack success was 14.18%.

Tool permissions should be narrowly scoped by data, destination, effect, and time. A permission to execute Python, edit a presentation, or use a terminal should not automatically authorize network transmission, credential access, or external sharing. “Don’t ask again” approvals should not apply across materially different security effects.

### Action authorization and runtime controls

High-impact actions should require independent validation or confirmation, especially:

- external communication;
- credential or sensitive-file access;
- cross-domain navigation;
- software installation;
- code execution;
- changes to consent or subscriptions;
- financial operations;
- irreversible submission;
- use of model-reproduced metadata as authoritative input.

Runtime systems should monitor tool-call sequences, provenance, data flow, destinations, and deviations from the original user task. Deterministic policy engines, taint tracking, capability restrictions, and sandboxing are stronger controls than relying exclusively on model refusal.

### Semantic and behavioral validation

Detectors should examine whether a requested data flow is appropriate for the task. Contextual integrity distinguishes a name field during checkout from a credit-score field during a flight lookup [2504.11281]. Output validation should check not only lexical markers but also whether the model changed persona, altered summary perspective, fabricated metadata, followed an external instruction, or initiated an unrelated action.

Behavioral monitoring is important because an injection may succeed without copying its text verbatim. PI-Hunter therefore evaluates trajectories, tool use, abnormal scope expansion, and sensitive-action attempts rather than only final outputs.

### Adversarial training and continuous red-teaming

Training data should include indirect, adaptive, contextual, and representation-specific attacks. WARD-PIG trains guards against fake verdicts and guard-directed instructions; A3T co-evolves attackers and guards. PromptArmor evaluates AgentVigil adaptive attacks. Randomized suffix augmentation improved activation-delta detector robustness more than activation-space PGD in the evaluated setting [2602.00750].

Continuous auditing should cover source discovery, dormant payloads, multi-step retrieval, tool chaining, cross-client aggregation, multimodal input, and downstream actions. PI-Hunter’s patch-and-reexplore procedure illustrates why patching the first discovered vulnerability is insufficient: subsequent exploration must seek alternative ingestion paths.

### Modality-specific defenses

For images, defenses include recompression, adaptive filtering, steganalysis, multimodal consistency checks, attention regularization, adversarial training, and behavioral monitoring [2507.22304]. The reported ensemble mitigation was 73.4%, with 28 ms additional latency and a 1.4% reduction on legitimate-task benchmarks, although these figures apply to the evaluated implementation.

For audio, CADV performs source separation, speaker-consistency analysis, and semantic filtering before ordinary instruction following [2607.28165]. It reduced High Frequency attack success from 70.55% to 2.32%, Pulse from 75.50% to 25.09%, and DolphinAttack from 81.55% to 33.95%. False-positive rates were 35% in office meetings, 14% in cafeterias, 1.5% in nature parks, and 0% in public stations, demonstrating a significant usability trade-off.

For structured carriers, domain-aware validation and restriction of auxiliary channels are more reliable than text-only scanning [2606.08403]. The current T3 carrier was blocked by an xxd detector and semantic validation, but the paper emphasizes that these are implementation-specific controls rather than a general solution.

## 6. Limitations, controversies, and research directions

The literature uses “invisible prompt injection” inconsistently. Some studies mean pixel-invisible or acoustically inaudible payloads; others mean user-invisible HTML, low-salience fine print, hidden tool metadata, semantically camouflaged instructions, detector-evasive representations, or dormant retrieval content. These phenomena share a trust-boundary failure but differ in attacker capability, representation, model interface, and defense requirements.

Several empirical limitations recur:

- Many studies use controlled or synthetic attacks rather than open-world deployments.
- Some evaluate only one or two models, platforms, or toolchains.
- Several report aggregate attack success without confidence intervals, confusion matrices, calibration, or independent statistical tests.
- Human studies often use small samples and synthetic tasks.
- LLM-based judges may introduce semantic bias.
- Attack success may mean detector evasion, marker leakage, intent to act, or completed harmful action; these are not interchangeable.
- Visible, low-salience, encoded, multimodal, and pixel-level attacks are often not evaluated together.
- Raw HTML, DOM text, accessibility trees, screenshots, OCR, and tool outputs are frequently conflated or insufficiently ablated.
- Strong benchmark performance does not establish robustness against adaptive attacks.

A particularly important distinction is between **detector evasion** and **behavioral compromise**. An adversarial suffix may make a probe classify an input as benign without preserving the injected instruction’s effectiveness [2602.00750]. Similarly, the 94.3% T3 leakage rate is not equivalent to 94.3% complete task takeover: the paper reports different Strong ASR values for output steering, configuration-token injection, and fabricated metadata [2606.08403]. Evaluation should therefore separately measure detection, instruction execution, unauthorized data flow, task utility, and downstream harm.

Future benchmarks should vary at least four dimensions:

1. **Representation**: raw source, rendered text, screenshots, accessibility content, audio, decoded structured data, tool metadata, and reconstructed reports.
2. **Concealment**: low salience, visual hiding, Unicode and encoding transformations, semantic camouflage, fragmentation, and steganographic carriers.
3. **Action risk**: summarization, navigation, data entry, external communication, software installation, credential access, and irreversible tool use.
4. **Adversarial adaptation**: attackers that observe detector scores, transformations, user confirmations, response rewriting, and action outcomes.

Evaluation should report both security and utility, including false-positive and false-negative rates, attack success, leakage success, completed unauthorized actions, unnecessary refusal, task completion, human burden, latency, cost, and robustness under preprocessing and repeated transformations.

The most defensible general principle is that untrusted content must not acquire authority merely because it is returned by a browser, tool, decoder, skill, decompiler, image encoder, audio recognizer, or retrieval system. Security controls should be placed at every representation and privilege boundary. Text classification remains useful when the malicious signal is present in the inspected text, but it cannot detect information that appears only after trusted reconstruction or through an uninspected modality. Conversely, structural isolation, least privilege, independent action authorization, provenance tracking, and domain-specific validation can reduce the impact of injections even when semantic detection fails.

Invisible prompt injection is therefore best understood as an end-to-end dataflow and authority problem. The central question is not merely whether an instruction is visible to a person or recognizable to a classifier, but whether attacker-controlled information can enter an agent’s effective context, influence its interpretation, propagate through transformations, and reach a consequential action without independent authorization.

Source: https://www.emergentmind.com/topics/invisible-prompt-injections