SceneSplit: Decomposing Harmful T2V Prompts
- SceneSplit is a novel method that decomposes harmful T2V prompts into 2–5 narrative scenes which collectively steer generation toward unsafe outputs.
- It employs an iterative process of scene manipulation and strategy updates to fine-tune prompts and bypass commercial safety filters.
- Empirical results show attack success rates up to 84.1%, highlighting significant vulnerabilities in current text-to-video safety measures.
SceneSplit is a black-box jailbreak method for text-to-video (T2V) models that decomposes a harmful prompt into multiple scenes that are individually benign-looking yet jointly steer generation toward unsafe video content (Lee et al., 26 Sep 2025). Its central mechanism is narrative decomposition rather than lexical obfuscation: each scene is written to appear safe in isolation, while the ordered combination constrains the model’s generative output space so that the resulting video is more likely to realize the original harmful intent. In the paper’s formulation, SceneSplit combines Scene Splitting, Scene Manipulation, and Strategy Update, and is evaluated on commercial T2V systems including Veo2, Luma Ray2, and Hailuo (Lee et al., 26 Sep 2025).
1. Conceptual basis and scope
SceneSplit is motivated by a safety gap specific to T2V systems. The paper distinguishes T2V models from LLMs, VLMs, and T2I models on the grounds that T2V generation is organized as a temporally ordered narrative, so prompt fragments can interact across time rather than functioning as a single static description (Lee et al., 26 Sep 2025). This creates an attack surface in which a harmful story can be distributed across scenes such that each scene appears benign in isolation, while the sequence as a whole implies unsafe content.
The method therefore does not rely primarily on hiding harmful keywords. Instead, it fragments the harmful narrative into 2–5 scenes, rewrites each scene with more benign expressions, and uses the sequential composition itself as the main attack mechanism (Lee et al., 26 Sep 2025). The paper states that Scene Division is the critical part of this process: the contribution of paraphrasing alone is weaker and model-dependent, whereas scene division materially changes the constraint structure of generation.
A common misconception is to read the name as referring to physical scene decomposition. In this work, “scene” means a narrative segment in a prompt for video generation, not a rigid/non-rigid partition of a 3D environment or a foreground/background split in visual reconstruction. This distinction is important because the method attacks safety filtering at the level of prompt structure and temporal semantics rather than spatial scene geometry (Lee et al., 26 Sep 2025).
2. Threat model and generative mechanism
SceneSplit assumes a black-box attacker who has access to the text prompt interface of a T2V model and can observe whether the model returns a video, blocks the prompt through a safety filter, or generates a benign result (Lee et al., 26 Sep 2025). The attacker may also use external helper models: an LLM for prompt rewriting, a text embedding model for retrieving prior strategies, a video understanding model for analyzing outputs, and an automated evaluator for unsafe content. The attacker does not require model weights, gradients, training data, or moderation scores.
The attack objective has two conditions: the generated video must convey the harmful intent of the original malicious prompt, and it must bypass the T2V model’s safety filter (Lee et al., 26 Sep 2025). Any blocked generation is counted as attack failure. This makes SceneSplit simultaneously an evasion method for input-side moderation and a probe of output-side safety weaknesses.
The paper’s central theoretical notion is the generative output space. A single benign scene prompt is described as corresponding to a wide and safe output space; another benign scene also corresponds to a wide and safe output space; but their sequential combination imposes mutual constraints that narrow the set of possible outputs into a smaller region that is more likely to be unsafe (Lee et al., 26 Sep 2025). This suggests that the operative vulnerability is compositional: the safety filter assesses local prompt harmfulness, while the T2V model resolves the full sequence into a temporally coherent narrative.
SceneSplit formalizes strategy retrieval by embedding the original harmful prompt into , searching a strategy library , and reusing the most similar prior strategy
when the similarity exceeds the threshold (Lee et al., 26 Sep 2025). This reuse mechanism makes the attack cumulative rather than one-shot.
3. Pipeline and algorithmic structure
The pipeline has three components: Scene Splitting, Scene Manipulation, and Strategy Update (Lee et al., 26 Sep 2025). The first component performs the initial narrative decomposition. The harmful prompt is divided into 2 to 5 procedural scenes, and each scene is paraphrased using more benign expressions. This step is performed by an LLM and may be guided by a previously successful pattern retrieved from the strategy library.
The second component, Scene Manipulation, is an iterative search inside the constrained prompt space created by the initial split. If a prompt is blocked, or if it generates a video that is judged too safe, the method modifies only one selected scene while keeping the rest fixed (Lee et al., 26 Sep 2025). When a video is generated but judged safe, VideoLLaMA3 is used to identify the most influential scene in the output; when no video is generated because the prompt was blocked, the next scene is chosen at random. The LLM then edits the selected scene in a bidirectional manner: if the video was too safe, the scene is made more explicit; if the prompt was blocked, the scene is made more implicit. The paper describes this as a search for an “optimal attack successful boundary” inside the constrained unsafe region.
The third component, Strategy Update, stores successful attack patterns. After a successful attack, a Summarizer LLM compresses the successful split/manipulated prompt into a reusable strategy , and the pair is added to the strategy library (Lee et al., 26 Sep 2025). The library is initialized as empty rather than manually populated, which the paper argues avoids human bias and allows adaptation to the target model’s vulnerabilities.
Algorithmically, the method uses an outer loop over strategies and an inner loop over scene-level modifications. The paper sets the maximum inner iterations to , the maximum outer iterations to , and the unsafe threshold to 0 (Lee et al., 26 Sep 2025). A prompt is judged a successful attack when the harmfulness score of the generated video meets or exceeds that threshold.
4. Empirical validation and ablation evidence
The evaluation uses T2VSafetyBench with 11 safety categories, 220 prompts total, and 20 prompts randomly sampled per category (Lee et al., 26 Sep 2025). The categories are Pornography, Borderline Pornography, Violence, Gore, Disturbing Content, Discrimination, Political Sensitivity, Illegal Activities, Misinformation, Sequential Action Risk, and Dynamic Variation Risk. Helper models in the pipeline are specified as all-MiniLM-L6-v2 for text embeddings, GPT-4o for Scene Splitting and Scene Manipulation, Qwen-30B as the Summarizer LLM, VideoLLaMA3 for influential-scene identification, and GPT-4o as the automated harmfulness evaluator (Lee et al., 26 Sep 2025).
The main metric is Attack Success Rate (ASR), with blocked generations counted as failures. On commercial T2V systems, SceneSplit reports average ASR values of 77.2% on Luma Ray2, 84.1% on Hailuo, and 78.2% on Veo2, compared with baseline ASRs of 39.5%, 40.9%, and 33.1% respectively (Lee et al., 26 Sep 2025). The absolute improvements are therefore 37.7, 43.2, and 45.1 points.
| Model | Baseline ASR | SceneSplit ASR |
|---|---|---|
| Luma Ray2 | 39.5% | 77.2% |
| Hailuo | 40.9% | 84.1% |
| Veo2 | 33.1% | 78.2% |
Category-level results show broad effectiveness. The paper highlights, for example, Hailuo, Pornography: 0% → 60% and Veo2, Illegal Activities: 10% → 90% under SceneSplit (Lee et al., 26 Sep 2025). It also reports high scores in action-centric categories such as Violence, Sequential Action Risk, and Dynamic Variation Risk across all three commercial systems.
The ablations identify Scene Division as the principal source of improvement. In one-shot experiments, Only Scene Division improves ASR on Veo2 from 33.1% to 37.7% and on Hailuo from 40.9% to 53.6%, whereas Only Paraphrasing gives 30.0% on Veo2 and 42.7% on Hailuo (Lee et al., 26 Sep 2025). On Veo2, the staged contribution is 42.7% for Scene Splitting only, 60.9% after adding Scene Manipulation, and 78.2% after adding Strategy Update. The strategy library ablation further shows an improvement from 69.1% without the library to 78.2% with it, under fixed total attack iterations (Lee et al., 26 Sep 2025).
The scene-selection heuristic is also validated directly. On Veo2, modifying the most influential scene yields 60.9%, compared with 56.4% for the least influential scene and 54.5% for random selection (Lee et al., 26 Sep 2025). This indicates that local refinement inside an already constrained narrative space is more effective than indiscriminate prompt rewriting.
5. Mechanistic evidence, semantic preservation, and failure modes
The paper provides a formal appendix-level analysis of the “constrained generative output space” hypothesis by defining a divergence score over sets of generated captions using the CLIP text encoder and cosine similarity (Lee et al., 26 Sep 2025). The reported result is that, as more scenes are combined, Divergence decreases and Unsafe Video Count increases. Single scenes have divergence around 0.2193–0.2421, whereas the full three-scene combination drops to 0.0808; at the same time, unsafe video count rises from 0.2–0.4 for single scenes to 2.2 for the full three-scene prompt (Lee et al., 26 Sep 2025). This is presented as quantitative support for the claim that scene combination both narrows the output space and steers it toward unsafe content.
The paper also tests prompt harmfulness directly with the OpenAI Moderation API on five detectable categories. SceneSplit prompts receive lower harmfulness scores than the original prompts. Reported examples include Pornography: Original 0.79, SceneSplit(all) 0.52, SceneSplit(avg scene) 0.25; Violence: Original 0.35, SceneSplit(all) 0.22, SceneSplit(avg scene) 0.04; and Gore: Original 0.61, SceneSplit(all) 0.35, SceneSplit(avg scene) 0.11 (Lee et al., 26 Sep 2025). The intended implication is that each scene is not merely cosmetically edited but is substantially more benign-looking under standard moderation.
A separate analysis addresses semantic preservation. Generated videos are captioned with VideoLLaMA3, and cosine similarity is computed between the original harmful prompt and the generated video caption. The paper reports average similarity 0.5308 and average variance 0.0219, arguing that the attack remains semantically aligned with the original malicious intent rather than drifting arbitrarily across iterations (Lee et al., 26 Sep 2025).
The limitations are explicit. SceneSplit works best for abstract, action-based harms, including pornography and violence, and is less effective when specific identity preservation is crucial, such as copyright infringement or public figures (Lee et al., 26 Sep 2025). The explanation given is that splitting and paraphrasing can dilute distinctive identity information, causing the T2V model to generate a different but vaguely similar person. The paper therefore frames the method as especially effective when harm is carried primarily by events and temporal structure rather than exact identity.
6. Defensive implications and terminological context
The principal defensive conclusion is that current T2V safeguards are too focused on prompt-level direct harmfulness and insufficiently sensitive to narrative-level contextual risk (Lee et al., 26 Sep 2025). SceneSplit exploits the fact that benign fragments can jointly imply an unsafe story. The paper therefore argues for defenses that reason across scenes, model temporal narrative structure more explicitly, and assess the combined semantics of multi-scene prompts rather than each phrase in isolation. It does not, however, present a concrete defense algorithm.
The work is framed as red-teaming rather than deployment guidance. The paper states that it contains harmful language and examples and positions the method as a way to expose vulnerabilities in current T2V systems so that stronger safeguards can be developed (Lee et al., 26 Sep 2025). A plausible implication is that sequence-aware moderation will need to evaluate not only lexical content but also temporal composition and cross-scene entailment.
In broader research usage, “scene split” refers to several different technical ideas, and SceneSplit should not be conflated with them. In dense RGB-D SLAM, SplitFusion partitions a scene into rigid and non-rigid surfaces for separate tracking and fusion (Li et al., 2020). In monocular reconstruction, DAS3R decomposes video into static background versus dynamic regions using learned dynamic masks and Gaussian staticness (Xu et al., 2024), while SplitGaussian explicitly decomposes static and dynamic Gaussian sets for monocular dynamic scene reconstruction (Li et al., 6 Aug 2025). In panoptic 3D Gaussian Splatting, Split&Splat performs instance-first reconstruction by first segmenting the scene and reconstructing each object independently (Monchieri et al., 1 Feb 2026). In long-form video analysis, movie scene segmentation methods treat “scene” as a semantic storytelling unit and perform boundary prediction over shot sequences (Rao et al., 2020). SceneSplit differs from all of these in that its “split” is a decomposition of harmful narrative prompts for black-box jailbreaking of T2V safety filters, not a decomposition of physical scenes, object instances, or cinematic structure (Lee et al., 26 Sep 2025).