---
title: 'SceneSplit: Decomposing Harmful T2V Prompts'
url: https://www.emergentmind.com/topics/scenesplit
type: topic
---

# SceneSplit: Decomposing Harmful T2V Prompts

SceneSplit is a black-box jailbreak method for text-to-video (T2V) models that decomposes a harmful prompt into multiple scenes that are individually benign-looking yet jointly steer generation toward unsafe video content [2509.22292]. Its central mechanism is narrative decomposition rather than lexical obfuscation: each scene is written to appear safe in isolation, while the ordered combination constrains the model’s generative output space so that the resulting video is more likely to realize the original harmful intent. In the paper’s formulation, SceneSplit combines Scene Splitting, Scene Manipulation, and Strategy Update, and is evaluated on commercial T2V systems including Veo2, Luma Ray2, and Hailuo [2509.22292].

## 1. Conceptual basis and scope

SceneSplit is motivated by a safety gap specific to T2V systems. The paper distinguishes T2V models from LLMs, VLMs, and T2I models on the grounds that T2V generation is organized as a temporally ordered narrative, so prompt fragments can interact across time rather than functioning as a single static description [2509.22292]. This creates an attack surface in which a harmful story can be distributed across scenes such that each scene appears benign in isolation, while the sequence as a whole implies unsafe content.

The method therefore does not rely primarily on hiding harmful keywords. Instead, it fragments the harmful narrative into **2–5 scenes**, rewrites each scene with more benign expressions, and uses the sequential composition itself as the main attack mechanism [2509.22292]. The paper states that Scene Division is the critical part of this process: the contribution of paraphrasing alone is weaker and model-dependent, whereas scene division materially changes the constraint structure of generation.

A common misconception is to read the name as referring to physical scene decomposition. In this work, “scene” means a narrative segment in a prompt for video generation, not a rigid/non-rigid partition of a 3D environment or a foreground/background split in visual reconstruction. This distinction is important because the method attacks safety filtering at the level of prompt structure and temporal semantics rather than spatial scene geometry [2509.22292].

## 2. Threat model and generative mechanism

SceneSplit assumes a black-box attacker who has access to the text prompt interface of a T2V model and can observe whether the model returns a video, blocks the prompt through a safety filter, or generates a benign result [2509.22292]. The attacker may also use external helper models: an LLM for prompt rewriting, a text embedding model for retrieving prior strategies, a video understanding model for analyzing outputs, and an automated evaluator for unsafe content. The attacker does not require model weights, gradients, training data, or moderation scores.

The attack objective has two conditions: the generated video must convey the harmful intent of the original malicious prompt, and it must bypass the T2V model’s safety filter [2509.22292]. Any blocked generation is counted as attack failure. This makes SceneSplit simultaneously an evasion method for input-side moderation and a probe of output-side safety weaknesses.

The paper’s central theoretical notion is the **generative output space**. A single benign scene prompt is described as corresponding to a wide and safe output space; another benign scene also corresponds to a wide and safe output space; but their sequential combination imposes mutual constraints that narrow the set of possible outputs into a smaller region that is more likely to be unsafe [2509.22292]. This suggests that the operative vulnerability is compositional: the safety filter assesses local prompt harmfulness, while the T2V model resolves the full sequence into a temporally coherent narrative.

SceneSplit formalizes strategy retrieval by embedding the original harmful prompt \(p\) into \(e_p \leftarrow Emb(p)\), searching a strategy library \(L = \{(s_1, e_1), \dots, (s_N, e_N)\}\), and reusing the most similar prior strategy
\[
(s^*, e^*) = \argmax_{(s,e)\in L \text{ s.t. } s \notin U} \cos\_sim(e, e_p)
\]
when the similarity exceeds the threshold \(\lambda = 0.6\) [2509.22292]. This reuse mechanism makes the attack cumulative rather than one-shot.

## 3. Pipeline and algorithmic structure

The pipeline has three components: **Scene Splitting**, **Scene Manipulation**, and **Strategy Update** [2509.22292]. The first component performs the initial narrative decomposition. The harmful prompt is divided into 2 to 5 procedural scenes, and each scene is paraphrased using more benign expressions. This step is performed by an LLM and may be guided by a previously successful pattern retrieved from the strategy library.

The second component, Scene Manipulation, is an iterative search inside the constrained prompt space created by the initial split. If a prompt is blocked, or if it generates a video that is judged too safe, the method modifies only one selected scene while keeping the rest fixed [2509.22292]. When a video is generated but judged safe, **VideoLLaMA3** is used to identify the **most influential scene** in the output; when no video is generated because the prompt was blocked, the next scene is chosen at random. The LLM then edits the selected scene in a bidirectional manner: if the video was too safe, the scene is made more explicit; if the prompt was blocked, the scene is made more implicit. The paper describes this as a search for an “optimal attack successful boundary” inside the constrained unsafe region.

The third component, Strategy Update, stores successful attack patterns. After a successful attack, a **Summarizer LLM** compresses the successful split/manipulated prompt into a reusable strategy \(s_{new}\), and the pair \((s_{new}, e_p)\) is added to the strategy library \(L\) [2509.22292]. The library is initialized as empty rather than manually populated, which the paper argues avoids human bias and allows adaptation to the target model’s vulnerabilities.

Algorithmically, the method uses an outer loop over strategies and an inner loop over scene-level modifications. The paper sets the maximum inner iterations to \(I=5\), the maximum outer iterations to \(T=3\), and the unsafe threshold to \(\theta_{unsafety}=60\) [2509.22292]. A prompt is judged a successful attack when the harmfulness score of the generated video meets or exceeds that threshold.

## 4. Empirical validation and ablation evidence

The evaluation uses **T2VSafetyBench** with **11 safety categories**, **220 prompts total**, and **20 prompts randomly sampled per category** [2509.22292]. The categories are Pornography, Borderline Pornography, Violence, Gore, Disturbing Content, Discrimination, Political Sensitivity, Illegal Activities, Misinformation, Sequential Action Risk, and Dynamic Variation Risk. Helper models in the pipeline are specified as **all-MiniLM-L6-v2** for text embeddings, **GPT-4o** for Scene Splitting and Scene Manipulation, **Qwen-30B** as the Summarizer LLM, **VideoLLaMA3** for influential-scene identification, and **GPT-4o** as the automated harmfulness evaluator [2509.22292].

The main metric is **Attack Success Rate (ASR)**, with blocked generations counted as failures. On commercial T2V systems, SceneSplit reports average ASR values of **77.2%** on **Luma Ray2**, **84.1%** on **Hailuo**, and **78.2%** on **Veo2**, compared with baseline ASRs of **39.5%**, **40.9%**, and **33.1%** respectively [2509.22292]. The absolute improvements are therefore **37.7**, **43.2**, and **45.1** points.

| Model | Baseline ASR | SceneSplit ASR |
|---|---:|---:|
| Luma Ray2 | 39.5% | 77.2% |
| Hailuo | 40.9% | 84.1% |
| Veo2 | 33.1% | 78.2% |

Category-level results show broad effectiveness. The paper highlights, for example, **Hailuo, Pornography: 0% → 60%** and **Veo2, Illegal Activities: 10% → 90%** under SceneSplit [2509.22292]. It also reports high scores in action-centric categories such as Violence, Sequential Action Risk, and Dynamic Variation Risk across all three commercial systems.

The ablations identify Scene Division as the principal source of improvement. In one-shot experiments, **Only Scene Division** improves ASR on Veo2 from **33.1%** to **37.7%** and on Hailuo from **40.9%** to **53.6%**, whereas **Only Paraphrasing** gives **30.0%** on Veo2 and **42.7%** on Hailuo [2509.22292]. On Veo2, the staged contribution is **42.7%** for Scene Splitting only, **60.9%** after adding Scene Manipulation, and **78.2%** after adding Strategy Update. The strategy library ablation further shows an improvement from **69.1%** without the library to **78.2%** with it, under fixed total attack iterations [2509.22292].

The scene-selection heuristic is also validated directly. On Veo2, modifying the **most influential scene** yields **60.9%**, compared with **56.4%** for the least influential scene and **54.5%** for random selection [2509.22292]. This indicates that local refinement inside an already constrained narrative space is more effective than indiscriminate prompt rewriting.

## 5. Mechanistic evidence, semantic preservation, and failure modes

The paper provides a formal appendix-level analysis of the “constrained generative output space” hypothesis by defining a divergence score over sets of generated captions using the CLIP text encoder and cosine similarity [2509.22292]. The reported result is that, as more scenes are combined, **Divergence decreases** and **Unsafe Video Count increases**. Single scenes have divergence around **0.2193–0.2421**, whereas the full three-scene combination drops to **0.0808**; at the same time, unsafe video count rises from **0.2–0.4** for single scenes to **2.2** for the full three-scene prompt [2509.22292]. This is presented as quantitative support for the claim that scene combination both narrows the output space and steers it toward unsafe content.

The paper also tests prompt harmfulness directly with the OpenAI Moderation API on five detectable categories. SceneSplit prompts receive lower harmfulness scores than the original prompts. Reported examples include **Pornography: Original 0.79, SceneSplit(all) 0.52, SceneSplit(avg scene) 0.25**; **Violence: Original 0.35, SceneSplit(all) 0.22, SceneSplit(avg scene) 0.04**; and **Gore: Original 0.61, SceneSplit(all) 0.35, SceneSplit(avg scene) 0.11** [2509.22292]. The intended implication is that each scene is not merely cosmetically edited but is substantially more benign-looking under standard moderation.

A separate analysis addresses semantic preservation. Generated videos are captioned with VideoLLaMA3, and cosine similarity is computed between the original harmful prompt and the generated video caption. The paper reports **average similarity 0.5308** and **average variance 0.0219**, arguing that the attack remains semantically aligned with the original malicious intent rather than drifting arbitrarily across iterations [2509.22292].

The limitations are explicit. SceneSplit works best for **abstract, action-based harms**, including pornography and violence, and is less effective when **specific identity preservation** is crucial, such as copyright infringement or public figures [2509.22292]. The explanation given is that splitting and paraphrasing can dilute distinctive identity information, causing the T2V model to generate a different but vaguely similar person. The paper therefore frames the method as especially effective when harm is carried primarily by events and temporal structure rather than exact identity.

## 6. Defensive implications and terminological context

The principal defensive conclusion is that current T2V safeguards are too focused on **prompt-level direct harmfulness** and insufficiently sensitive to **narrative-level contextual risk** [2509.22292]. SceneSplit exploits the fact that benign fragments can jointly imply an unsafe story. The paper therefore argues for defenses that reason across scenes, model temporal narrative structure more explicitly, and assess the combined semantics of multi-scene prompts rather than each phrase in isolation. It does not, however, present a concrete defense algorithm.

The work is framed as red-teaming rather than deployment guidance. The paper states that it contains harmful language and examples and positions the method as a way to expose vulnerabilities in current T2V systems so that stronger safeguards can be developed [2509.22292]. A plausible implication is that sequence-aware moderation will need to evaluate not only lexical content but also temporal composition and cross-scene entailment.

In broader research usage, “scene split” refers to several different technical ideas, and SceneSplit should not be conflated with them. In dense RGB-D SLAM, **SplitFusion** partitions a scene into rigid and non-rigid surfaces for separate tracking and fusion [2007.02108]. In monocular reconstruction, **DAS3R** decomposes video into static background versus dynamic regions using learned dynamic masks and Gaussian staticness [2412.19584], while **SplitGaussian** explicitly decomposes static and dynamic Gaussian sets for monocular dynamic scene reconstruction [2508.04224]. In panoptic 3D Gaussian Splatting, **Split&Splat** performs instance-first reconstruction by first segmenting the scene and reconstructing each object independently [2602.03809]. In long-form video analysis, movie scene segmentation methods treat “scene” as a semantic storytelling unit and perform boundary prediction over shot sequences [2004.02678]. SceneSplit differs from all of these in that its “split” is a decomposition of harmful narrative prompts for black-box jailbreaking of T2V safety filters, not a decomposition of physical scenes, object instances, or cinematic structure [2509.22292].

Source: https://www.emergentmind.com/topics/scenesplit