---
title: 'PoisonVID: Black-Box Attack on VideoLLMs'
url: https://www.emergentmind.com/topics/poisonvid
type: topic
---

# PoisonVID: Black-Box Attack on VideoLLMs

PoisonVID is a black-box poisoning attack against prompt-guided sampling (PGS) in Video Large Language Models (VideoLLMs). It targets the frame-selection stage rather than the final decoder, exploiting the fact that modern VideoLLMs cannot process every frame of a long video and therefore rely on a relevance model to choose a small prompt-conditioned subset. PoisonVID suppresses the prompt relevance of harmful frames by optimizing a universal additive perturbation over a short harmful clip, so that those frames are omitted or underrepresented before downstream reasoning occurs. In the reported evaluation, the method is presented as the first black-box poisoning attack on prompt-guided sampling and achieves \(82\%\)–\(99\%\) attack success rate across three PGS strategies and three advanced VideoLLMs [2509.20851].

## 1. Attack surface in prompt-guided video sampling

The attack is motivated by the architecture of contemporary VideoLLMs. A video is written as
\[
V = \{f_1, f_2, \dots, f_T\},
\]
and a sampling strategy \(\mathcal{S}(V, N)\) selects \(N \ll T\) frames before the model fuses visual evidence with the text prompt. In prompt-guided sampling, a dense candidate set \(V_M\) with \(M > N\) is first formed, and each candidate frame \(f_k \in V_M\) is assigned a relevance score \(r(f_k, q')\), where \(q'\) is a prompt-derived query. The sampler then returns the Top-\(N\) frames under this relevance function:
\[
\mathcal{S}_{\text{PGS}}(V, N, q') = \operatorname*{arg\,\text{Top-}N}_{f_k \in V_M} \; r(f_k, q').
\]
PoisonVID attacks this auxiliary scoring mechanism rather than attempting to jailbreak the final language model directly [2509.20851].

This focus distinguishes PoisonVID from earlier attacks on older sampling schemes such as uniform frame sampling (UFS) and semantic-similarity sampling (SSS). The paper positions PGS as a stronger family of samplers because harmful frames normally receive high prompt relevance when the query explicitly asks about violence, crime, or pornography. Earlier attacks such as Frame Replacement Attack (FRA) relied on blind spots in non-prompt-aware selection, whereas PoisonVID treats prompt-conditioned relevance itself as the vulnerability. This makes the attack semantically targeted: if relevance can be suppressed upstream, the VideoLLM may never see enough evidence to answer correctly.

The paper studies three concrete PGS strategies. Differential Keyframe Selection (DKS) combines prompt-guided relevance with feature similarity between adjacent frames to reduce redundancy. Adaptive Keyframe Sampling (AKS) combines relevance with a coverage or diversity term. Frame Selection Augmented Generation (FRAG) relies almost exclusively on frame relevance ranking and does not explicitly penalize inter-frame redundancy. PoisonVID is designed to transfer across these variants because it operates on the most generic shared component, namely relevance scoring.

## 2. Threat model and attack objective

The attacker inserts a short harmful clip into a benign video and perturbs the harmful frames under a small \(\ell_\infty\) budget, while keeping the clip natural-looking and recognizable to humans. The setting is black-box in the sense that the attacker does not know the target VideoLLM’s architecture, weights, gradients, or exact sampling implementation, and does not optimize through repeated online probing. Instead, PoisonVID relies on surrogate components: a shadow VideoLLM to describe the harmful clip and a lightweight VLM to score frame-query relevance during offline optimization [2509.20851].

The core goal is omission rather than direct semantic inversion. When the victim system is asked whether a poisoned video contains harmful content, the attack is successful if the harmful segment is omitted or sufficiently underrepresented in the sampled frame set, so that the final response fails to surface the inserted harm. The attack therefore targets the evidence bottleneck of the VideoLLM stack. A common misconception is that PoisonVID is a universal prompt-agnostic backdoor on the final model; the paper instead characterizes it as a transfer-based poisoning attack on PGS, with perturbations optimized to reduce prompt relevance for a specific harmful clip.

The harmful clip is denoted
\[
\mathcal{H}=\{f_1, f_2, \dots, f_h\},
\]
and PoisonVID applies a single additive perturbation \(\delta\) uniformly to every frame:
\[
\tilde{f}_j = f_j + \delta.
\]
This universal parameterization is central to the method. It makes optimization cheaper than per-frame perturbation, and it is the basis for the paper’s claim that the resulting perturbation transfers across different benign host videos, different PGS implementations, and different victim VideoLLMs. At the same time, the universality claim is bounded: the perturbation is clip-specific, frame-universal, and cross-model transferable, rather than universal across arbitrary harmful categories.

## 3. Closed-loop optimization and the depiction set

A distinctive component of PoisonVID is the depiction set, which provides a surrogate semantic envelope for the harmful clip. The pipeline begins by using a shadow VideoLLM to generate a textual description \(\hat d\) of the harmful clip. That description is then paraphrased with a lightweight language model, specifically GPT-4o-mini, to obtain a depiction set
\[
\mathcal{D}=\{d_1, d_2, \dots, d_l\}.
\]
Each depiction \(d_k\) is converted into a prompt-guided query \(q'_k\), producing
\[
\mathcal{Q'}=\{q'_1, q'_2, \dots, q'_l\}.
\]
The use of paraphrases is intended to make the attack robust to varied prompt formulations without access to the true victim prompt distribution [2509.20851].

Optimization is then defined through the relevance suppression loss (RSL):
\[
\mathcal{L}_{\mathrm{RSL}}(\delta;\mathcal{H},\mathcal{Q'}) = \frac{1}{hl} \sum_{j=1}^h \sum_{k=1}^l r(\tilde{f}_j,q'_k).
\]
Minimizing this objective pushes the perturbed harmful frames away from the semantic region associated with the harmful depiction queries. The projected-gradient update is
\[
\delta^{(t+1)} = \Pi_{\|\cdot\|_\infty \leq \epsilon} \Big( \delta^{(t)} - \eta \,\nabla_\delta \mathcal{L}_{\mathrm{RSL}}(\delta^{(t)}; \mathcal{H}, \mathcal{Q'}) \Big),
\]
with \(\epsilon = 8/255\). The paper describes this as a closed-loop optimization strategy because the perturbation is iteratively refined using surrogate relevance feedback.

The attack pseudocode is correspondingly minimal. Given the harmful clip \(\mathcal{H}\), the depiction set \(\mathcal{D}\), an initial perturbation \(\delta_0\), learning rate \(\eta\), perturbation constraint \(\epsilon\), and epochs \(Z\), the method derives \(\mathcal{Q'}\), applies \(\delta\) to each frame, computes \(\mathcal{L}_{\mathrm{RSL}}\), updates \(\delta\), projects back to the \(\ell_\infty\) ball, and returns the final perturbation. No differentiable Top-\(N\) surrogate, explicit frame-selection simulator, or auxiliary regularizer is introduced; the method relies on the premise that direct relevance suppression is sufficient because all PGS strategies depend heavily on prompt relevance.

Implementation details are concrete. The paper uses LLaVA-Video-7B-Qwen2 (L-7B) as the shadow VideoLLM for harmful depiction generation, GPT-4o-mini for paraphrasing, and BLIP as the default lightweight VLM for optimization-time relevance scoring. The perturbation is initialized with uniform noise in \([-\epsilon,\epsilon]\), optimization runs for 1,000 steps using exponential learning-rate decay with initial rate 10 and decay factor 0.999, the depiction set size is \(5\), and each iteration samples 8 random frames from the harmful clip to reduce compute.

## 4. Evaluation protocol and empirical results

The evaluation uses 100 benign videos randomly selected from LLaVA-Video-178K and 15 harmful clips from three categories: violence, crime, and pornography. In the poisoning setup, a contiguous 4-second segment of each benign video is replaced with a harmful clip. All methods use a fixed selected-frame count \(N = 32\). The victim VideoLLMs are LLaVA-Video-7B-Qwen2 (L-7B), VideoLLaMA2 (VL2), and ShareGPT4Video (SG4V), and the samplers are DKS, AKS, and FRAG [2509.20851].

The principal metric is attack success rate (ASR), defined as the proportion of poisoned videos in which the model fails to recognize the inserted harmful content. The reported averages show a large gap between FRA and PoisonVID under prompt-guided sampling. For FRA, ASR is 66% on L-7B, 29% on VL2, and 54% on SG4V under DKS; 44%, 24%, and 43% under AKS; and 11%, 18%, and 31% under FRAG. For PoisonVID, the corresponding averages are 92%, 82%, and 99% under DKS; 90%, 87%, and 90% under AKS; and 88%, 84%, and 99% under FRAG. Across categories and model-strategy pairs, the attack therefore ranges roughly from \(82\%\) to \(99\%\), with some individual settings reaching 100%, especially on SG4V.

The strategy-level interpretation is informative. FRA is weakest on FRAG because FRAG ranks by relevance alone, so highly relevant harmful frames are normally retained. DKS is easier for FRA because neighbor-similarity pruning can compress contiguous harmful segments. PoisonVID changes this picture by reducing relevance upstream: once harmful frames cease to look prompt-relevant, neither redundancy pruning nor coverage heuristics can recover them reliably. Qualitative examples in the paper indicate that, without perturbation, harmful frames often dominate the selected top-32 set; after PoisonVID, the number of harmful frames among those 32 selected frames often drops to only 1–2.

The paper also reports performance beyond PGS. Under older samplers, PoisonVID still achieves high ASR: 96, 96, and 99 on L-7B, VL2, and SG4V for UFS; 95, 89, and 99 for SSS-SKE; 92, 82, and 99 for PGS-DKS; 90, 87, and 90 for PGS-AKS; and 88, 84, and 99 for PGS-FRAG. Although the paper’s central novelty is the attack on prompt-guided sampling, this result suggests that suppressing semantic salience can generalize to other frame-selection regimes as well.

## 5. Transferability, ablations, and limitations

Transferability is one of the method’s main empirical claims. Even though L-7B is used as the shadow VideoLLM to generate depictions, the perturbation transfers to VL2 and SG4V, and in some cases is even stronger on SG4V. Likewise, although BLIP is used for optimization, the perturbation transfers to PGS implementations that use different relevance estimators, including FRAG, which uses a VideoLLM-based relevance mechanism rather than a lightweight VLM. The paper interprets this as evidence that PoisonVID captures a shared semantic vulnerability rather than overfitting to a single scorer [2509.20851].

The ablations clarify which components matter. On the choice of lightweight VLM, using L-7B as the evaluation VideoLLM, the reported average ASRs are 92 for BLIP, 93 for CLIP, and 94 for Combined on DKS; 90, 71, and 71 on AKS; and 88, 78, and 83 on FRAG. BLIP is therefore chosen as the default because it transfers more consistently across all strategies, even though CLIP or Combined are slightly better on DKS. The harmful-clip-length ablation shows that ASR decreases as the inserted harmful segment becomes longer, but remains substantial: even with 10-second harmful clips, PoisonVID still achieves over 60% average ASR, and even at 36 seconds DKS still averages 68% ASR.

The paper is explicit about scope limitations. It assumes knowledge that the victim uses a relevance-based prompt-guided sampler in general, even though it remains black-box with respect to implementation details. The attack depends on a shadow VideoLLM to generate depictions, and those depictions may not span the full diversity of harmful semantics. The evaluation studies additive perturbations on short harmful clips rather than temporal reordering, audio-text interactions, or multimodal attacks beyond the visual channel. A plausible implication is that PoisonVID opens a broader attack surface rather than exhausting it.

Several defensive directions are proposed but not validated as complete solutions. The paper mentions ensemble relevance scoring, temporal consistency constraints, and increased redundancy in frame selection, such as sampling multiple diverse subsets instead of relying on a single Top-\(N\) set. These suggestions reflect the central lesson of the work: prompt-guided sampling improves utility, but because it depends on learned prompt-frame relevance, it introduces a semantically targeted attack surface that can be poisoned before reasoning begins.

## 6. Position within multimodal poisoning research

PoisonVID belongs to a wider class of attacks that poison intermediate retrieval or selection mechanisms rather than directly perturbing a final decoder. In visual document retrieval-augmented generation, a single adversarial document-page image can be jointly optimized for retrieval prominence and downstream generator control [2504.02132]. In multimodal RAG, metadata-only poisoning can steer retrieval and induce attacker-desired responses while leaving the visual content untouched [2603.00172]. In web agents with graph-structured external memory, multimodal memory poisoning can create persistent, trigger-conditioned retrieval and post-retrieval behavioral takeover [2606.10742]. PoisonVID differs from these systems in its target surface: it attacks prompt-guided frame selection in sequential video understanding rather than KB entries, captions, or external memory.

This distinction matters technically. In the visual-document and MM-RAG settings, the attacker optimizes retrievability of poisoned items against text queries. In MemVenom, the attacker poisons persistent external memory so that a triggered observation recalls malicious nodes. PoisonVID instead suppresses the prompt relevance of harmful evidence already present in the video, creating an omission attack on the sampling stage. A useful shorthand is that adjacent RAG and memory attacks are often “evidence insertion” attacks, whereas PoisonVID is an “evidence suppression” attack in the temporal visual domain.

At the same time, these works share a structural insight: multimodal systems frequently rely on learned auxiliary mechanisms to decide what evidence will enter the final reasoning context. Whether the bottleneck is document retrieval [2504.02132], metadata-mediated top-\(k\) retrieval [2603.00172], graph-memory recall [2606.10742], or prompt-guided frame selection [2509.20851], poisoning that bottleneck can dominate downstream behavior without directly rewriting model parameters or forcing a prompt-level jailbreak. In that sense, PoisonVID is a specific instance of a broader security pattern in multimodal AI: auxiliary selection modules, precisely because they are optimized for efficiency and relevance, become high-leverage attack surfaces.

Source: https://www.emergentmind.com/topics/poisonvid