---
title: 'STaR-Attack: Multimodal Jailbreak Framework'
url: https://www.emergentmind.com/topics/star-attack
type: topic
---

# STaR-Attack: Multimodal Jailbreak Framework

to=functions.shell  彩神争霸大发快三  天天中彩票派奖 json
{"command":"bash -lc 'python - <<\"PY\"\nimport urllib.request, urllib.parse\nquery='all:STaR-Attack'\nurl='http://export.arxiv.org/api/query?search_query='+urllib.parse.quote(query)+'&start=0&max_results=5'\ntry:\n    print(urllib.request.urlopen(url,timeout=20).read().decode('utf-8')[:1200])\nexcept Exception as e:\n    print('ERR',e)\nPY'"}【อ่านข้อความเต็มjson
{"output":"ERR HTTP Error 403: Forbidden\n","metadata":{"exit_code":0,"duration_seconds":0.25}}
to=functions.shell  微信上的天天中彩票  亚历山大发 to=functions.shell  天天中彩票谁  天天中彩票追号്json
{"command":"bash -lc 'python - <<\"PY\"\nimport urllib.request, urllib.parse\nquery='all:STaR-Attack'\nurl='https://export.arxiv.org/api/query?search_query='+urllib.parse.quote(query)+'&start=0&max_results=5'\nreq=urllib.request.Request(url,headers={'User-Agent':'Mozilla/5.0'})\ntry:\n    data=urllib.request.urlopen(req,timeout=20).read().decode('utf-8')\n    print(data[:1500])\nexcept Exception as e:\n    print('ERR',e)\nPY'"}【อ่านข้อความเต็มjson
{"output":"<?xml version=\"1.0\" encoding=\"UTF-8\"?>\n<feed xmlns=\"http://www.w3.org/2005/Atom\">\n  <link href=\"http://arxiv.org/api/query?search_query%3Dall%3ASTaR-Attack%26id_list%3D%26start%3D0%26max_results%3D5\" rel=\"self\" type=\"application/atom+xml\"/>\n  <title type=\"html\">ArXiv Query: search_query=all:STaR-Attack&amp;id_list=&amp;start=0&amp;max_results=5</title>\n  <id>http://arxiv.org/api/eGnA7UEFb4kl0f8WyTVxJyR0kc8</id>\n  <updated>2026-07-14T00:00:00-04:00</updated>\n  <opensearch:itemsPerPage xmlns:opensearch=\"http://a9.com/-/spec/opensearch/1.1/\">5</opensearch:itemsPerPage>\n  <opensearch:startIndex xmlns:opensearch=\"http://a9.com/-/spec/opensearch/1.1/\">0</opensearch:startIndex>\n  <opensearch:totalResults xmlns:opensearch=\"http://a9.com/-/spec/opensearch/1.1/\">2</opensearch:totalResults>\n  <entry>\n    <id>http://arxiv.org/abs/2509.26473v1</id>\n    <updated>2025-10-02T18:00:27Z</updated>\n    <published>2025-09-30T17:55:13Z</published>\n    <title>STaR-Attack: A Spatio-Temporal and Narrative Reasoning Attack\n      Framework for Unified Multimodal Understanding and Generation Models</title>\n    <summary>  Unified Multimodal understanding and generation Models (UMMs) have\n  demonstrated remarkable capabilities in both understanding and\n  generation tasks. However, we identify a vulne","metadata":{"exit_code":0,"duration_seconds":0.53}}
to=functions.shell  鸿丰աշտոքjson
{"command":"bash -lc 'python - <<\"PY\"\nimport urllib.request, urllib.parse, re\nqueries=['all:STAR-Teaming','all:STARE multi-modal toxicity','all:\"Sensitive Trajectory Regulation\"','all:\"State-Transition Amplification Ratio\"','all:\"Single-Turn Crescendo Attack\"']\nfor query in queries:\n    url='https://export.arxiv.org/api/query?search_query='+urllib.parse.quote(query)+'&start=0&max_results=3'\n    req=urllib.request.Request(url,headers={'User-Agent':'Mozilla/5.0'})\n    print('\\nQUERY',query)\n    try:\n        data=urllib.request.urlopen(req,timeout=20).read().decode('utf-8')\n        entries=re.findall(r'<entry>(.*?)</entry>',data,re.S)\n        for e in entries:\n            idm=re.search(r'<id>http://arxiv.org/abs/(.*?)</id>',e)\n            titlem=re.search(r'<title>(.*?)</title>',e,re.S)\n            if idm and titlem:\n                title=' '.join(titlem.group(1).split())\n                print(idm.group(1), title)\n    except Exception as e:\n        print('ERR',e)\nPY'"}【อ่านข้อความเต็มjson
{"output":"\nQUERY all:STAR-Teaming\n2604.18976v1 STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming\n\nQUERY all:STARE multi-modal toxicity\n2605.00699v1 STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack\n\nQUERY all:\"Sensitive Trajectory Regulation\"\n2601.09281v1 STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning Models\n\nQUERY all:\"State-Transition Amplification Ratio\"\n2601.08511v1 STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio\n\nQUERY all:\"Single-Turn Crescendo Attack\"\n2411.18699v1 An indicator for effectiveness of text-to-image guardrails utilizing the Single-Turn Crescendo Attack (STCA)\n","metadata":{"exit_code":0,"duration_seconds":2.96}}
STaR-Attack is a multi-turn jailbreak framework for unified multimodal understanding-and-generation models (UMMs) that exploits a vulnerability arising from generation–understanding coupling. Its central mechanism, termed **Cross-Modal Generative Injection (CMGI)**, induces a target UMM to first generate an information-rich image encoding adversarial context and then to interpret that same image in a later turn, thereby recovering and answering a harmful query that would ordinarily be refused in plain text. The framework is presented as the first multi-turn jailbreak attack tailored to this coupled attack surface, and it is explicitly designed to preserve the original semantics of the harmful query rather than relying on paraphrastic prompt rewriting with semantic drift [2509.26473].

## 1. Concept and threat model

STaR-Attack targets UMMs that support both image generation and multimodal conversation. The assumed adversary can interact with the target model over multiple turns and seeks to obtain unsafe output for a harmful user query \(Q\) while preserving the semantics of \(Q\). The paper formalizes the attack around a malicious event \(E\) that is semantically highly correlated with the harmful query:

$$
\mathcal{R}(E, Q) > \delta, \qquad 0 \ll \delta < 1, \qquad \mathcal{R}\in[0,1].
$$

Here, \(\mathcal{R}(\cdot,\cdot)\) denotes semantic relevance. The stated intuition is that, if the model can be guided to infer \(E\), it will naturally recover \(Q\). Directly presenting \(E\), however, is expected to trigger safety filters, so the attack does not expose the event explicitly [2509.26473].

The vulnerability identified by the paper is not merely multimodality in the abstract, but the specific **generation–understanding coupling** of UMMs. In this setting, the model’s own generative function produces adversarial evidence, and its own understanding function consumes that evidence. The paper names this mechanism CMGI and characterizes it as an attack pathway that compresses malicious context into a single image-based turn, transferring adversarial information across modalities without directly issuing the harmful instruction in rewritten form.

A recurring misconception addressed by the evaluation is that a jailbreak should be judged only by whether it elicits unsafe output. STaR-Attack is framed against that view. Because the original malicious query is preserved and directly embedded in the final candidate set, the method is intended to avoid the semantic drift that weakens many prior jailbreaks. This suggests that the framework is meant to demonstrate not just refusal circumvention, but refusal circumvention with query fidelity.

## 2. Spatio-temporal and narrative formulation

The framework’s defining abstraction is a hidden malicious event embedded inside a causal narrative. Rather than exposing \(E\), STaR-Attack constructs a **pre-event** scene \(S_{\text{pre}}\) and a **post-event** scene \(S_{\text{post}}\), with the event concealed as a hidden climax. The paper models this as a directed causal graph:

$$
\mathcal{G} = (V, \mathcal{E}), \qquad
V = \{S_{\text{pre}}, E, S_{\text{post}}\}, \qquad
\mathcal{E} = \{S_{\text{pre}}\rightarrow E,\; E\rightarrow S_{\text{post}}\}.
$$

The narrative is factorized as

$$
P(S_{\text{pre}}, E, S_{\text{post}})
=
P(S_{\text{pre}})
\,P(E\mid S_{\text{pre}})
\,P(S_{\text{post}}\mid S_{\text{pre}}, E).
$$

This structure underlies what the paper calls a **spatio-temporal reasoning attack**: the malicious event is not shown directly, but is positioned as the inferred causal bridge between setup and resolution [2509.26473].

The scene construction is constrained. The paper requires the pre-event and post-event scenes to remain only weakly connected to the original harmful query and to be less toxic than the hidden event itself:

$$
\epsilon \leq \mathcal{R}(S_{\text{pre}}, Q), \ \mathcal{R}(S_{\text{post}}, Q) < \delta,
$$

$$
\mathcal{T}(S_{\text{pre}}), \ \mathcal{T}(S_{\text{post}}) < \mathcal{T}(E),
$$

where \(\mathcal{T}(\cdot)\) is a toxicity score. The purpose of these inequalities is to keep the surrounding scenes semantically useful but less likely to activate safety filters.

The attack explicitly adopts the **three-act narrative theory**. The first act is the **setup**, in which the harmful query is rewritten into a rich description of the pre-event scene. The second act is the **climax**, but the climax is never stated directly; it is hidden in the causal gap between the two visible scenes. The third act is the **resolution**, in which the post-event scene is generated. The paper states that scene construction uses the uncensored auxiliary model **Qwen2.5-QwQ-37B-Eureka-Triple-Cubed-abliterated-uncensored**.

## 3. Multi-turn execution and dynamic difficulty

Operationally, STaR-Attack proceeds over three rounds. The first two rounds exploit the target UMM’s generative ability: the attacker prompts the model to produce images for the pre-event and post-event scenes. These images seed the dialogue with aligned visual context. The final round switches to the model’s understanding pathway by introducing an image-based question guessing and answering game [2509.26473].

The crucial final-stage construction defines a candidate set

$$
\mathcal{C} = \{Q\} \cup \mathcal{C}_{\text{benign}}, \qquad
\mathcal{C}_{\text{benign}} = \{\hat Q_1,\hat Q_2,\dots,\hat Q_k\},
$$

with benign candidates constrained by

$$
\mathcal{R}(\hat Q_i, Q) < \tau.
$$

The paper sets \(\tau=0.4\) using cosine similarity over text embeddings as a surrogate relevance score. The model is then asked to identify the “real question” implied by the images and the narrative context:

$$
Q^* = \arg\max_{Q' \in \mathcal{C}} \mathcal{R}(E, Q').
$$

If the model selects \(Q^*=Q\), it has effectively reconstructed the original malicious instruction and is then asked to answer it. Because \(Q\) itself is present in the candidate set, the jailbreak does not depend on paraphrasing the harmful request into a semantically shifted substitute.

A notable component is the **dynamic difficulty mechanism**, which varies the number of benign distractors. The paper defines four difficulty levels:

| Difficulty | Benign questions |
|---|---:|
| \(D_0\) | 0 |
| \(D_1\) | 1 |
| \(D_2\) | 3 |
| \(D_3\) | 7 |

The attack starts at \(D_0\). If the model gives a safe response, an external judge, **Llama-Guard-4**, checks it, and the next attempt increases the difficulty according to

$$
D_{t+1}=\min(D_t+\Delta, D_{\max}).
$$

The paper argues that larger candidate sets make the malicious query less conspicuous and force stronger reliance on the narrative context. A plausible implication is that the guessing game is not only a delivery vehicle for the harmful query, but also a mechanism for shifting the model’s inference burden from lexical matching to contextual reconstruction.

## 4. Evaluation protocol and empirical results

The evaluation uses two harmful-instruction datasets. **AdvBench** contains 520 harmful instructions spanning profanity, graphic content, threats, misinformation, discrimination, cybercrime, and illicit advice. **HarmBench** is sampled as 400 textual behaviors, split into 200 standard behaviors, 100 copyright behaviors, and 100 contextual behaviors. The target models are the open-source **Janus-7B-Pro** and **BAGEL-7B-MoT**, together with the closed-source **Gemini-2.0-Flash** and **Gemini-2.5-Flash**. Baselines are **Text-Only**, **PAIR**, **ReNeLLM**, **FlipAttack**, and the multi-turn baseline **X-Teaming** [2509.26473].

The paper uses **ASR** as the primary jailbreak metric, with **Llama-Guard-4-12B** as the judge. To measure semantic fidelity, it also reports **Relevant Rate (RR)** using GPT-4o and a stricter **Relevant Attack Success Rate (RASR)**, which requires the response to be both unsafe and relevant to the original query. GPT-4o-based harmfulness judgments on a 1–5 scale are also reported, with a score of 5 counted as a successful jailbreak.

On **AdvBench**, the reported results are as follows:

| Model | ASR | RASR |
|---|---:|---:|
| Janus-Pro | 93.06% | 71.87% |
| BAGEL | 89.6% | 57.23% |
| Gemini-2.0-Flash | 93.06% | 65.32% |
| Gemini-2.5-Flash | 88.05% | 45.47% |

On **HarmBench**, the reported results are:

| Model | ASR | RASR |
|---|---:|---:|
| Janus-Pro | 92.75% | 45% |
| BAGEL | 89.0% | 46.25% |
| Gemini-2.0-Flash | 90.75% | 61.5% |
| Gemini-2.5-Flash | 88.5% | 52% |

The comparison to **FlipAttack** is emphasized because it separates raw jailbreak frequency from semantic fidelity. On AdvBench with Gemini-2.0-Flash, FlipAttack reaches **86.73% ASR** and **45.58% RASR**, whereas STaR-Attack reaches **93.06% ASR** and **65.32% RASR**. On HarmBench with Gemini-2.0-Flash, FlipAttack gets **86.5% ASR** and **58.5% RASR**, compared with **90.75% ASR** and **61.5% RASR** for STaR-Attack. The paper also notes that **PAIR** can exhibit decent ASR yet extremely poor RASR; on AdvBench with Janus-Pro, PAIR gets **76.15% ASR** but only **13.27% RASR**. In the paper’s framing, high ASR without relevance is a weaker jailbreak result than high ASR with preserved semantics.

## 5. Mechanistic interpretation, ablations, and limitations

The paper attributes the attack’s effectiveness to three interacting factors. First, the pre-event and post-event scenes provide a coherent causal scaffold that makes the hidden malicious event easier to infer. Second, the benign candidate questions are semantically distant from \(Q\): mean cosine similarity is reported as around **0.0670 on AdvBench** and **0.0572 on HarmBench**, so the final selection is driven primarily by narrative context rather than lexical overlap. Third, the dynamic candidate set increases cognitive load and encourages stronger reliance on prior context [2509.26473].

The ablation studies support these claims. For dynamic difficulty on HarmBench, **BAGEL** yields **76.75%**, **79.00%**, **76.75%**, and **72.00%** ASR at fixed levels 0–3, versus **88.75%** for STaR-Attack; **Janus-Pro** yields **82.50%**, **57.25%**, **49.50%**, and **44.25%**, versus **90.25%** for STaR-Attack. On AdvBench, **BAGEL** yields **66.28%**, **81.00%**, **80.15%**, and **76.00%**, versus **89.60%** for STaR-Attack; **Janus-Pro** yields **89.79%**, **58.00%**, **45.47%**, and **36.99%**, versus **92.68%** for STaR-Attack**.** The paper interprets this as evidence that adaptive difficulty consistently outperforms any fixed candidate-set size.

Interaction-structure ablations show that the multi-turn design matters. When the method is reduced to a single-turn or an **img-direct** baseline, effectiveness declines in a model-dependent way. On Janus-Pro, single-turn ASR falls from **90.25%** to **50%** on HarmBench and from **92.68%** to **37.57%** on AdvBench. The authors attribute BAGEL’s lower sensitivity to its flatter input formatting, whereas Janus-Pro appears to benefit more from preserved conversational structure.

The limitations are explicit. The method requires **multi-turn conversation support**, so it cannot be applied to UMMs lacking that capability, such as **BLIP3-o** and **Show-o2**, which were excluded from experiments. The paper also states that model-specific prompting templates strongly affect attack success, suggesting that defenses may need to account for conversation structure rather than only content filtering. Its mitigation message is correspondingly narrow but pointed: safety work for UMMs should address the **cross-modal, generation-to-understanding pathway** itself.

## 6. Position within adjacent literature and nomenclature

STaR-Attack is distinct from several later works whose names partially overlap. **“STaR: Sensitive Trajectory Regulation for Unlearning in Large Reasoning Models”** proposes a defense-oriented, parameter-free, inference-time unlearning framework for LRMs, not an attack, and is concerned with privacy leakage in Chain-of-Thought trajectories [2601.09281]. **“STAR: Detecting Inference-time Backdoors in LLM Reasoning via State-Transition Amplification Ratio”** is likewise a defense framework, aimed at detecting inference-time backdoors in reasoning traces through posterior–prior probability mismatch and CUSUM [2601.08511].

It is also distinct from automated red-teaming and trajectory-level multimodal toxicity frameworks. **“STAR-Teaming: A Strategy-Response Multiplex Network Approach to Automated LLM Red Teaming”** is a black-box automated jailbreak-discovery system built around a multi-agent loop and a multiplex network over strategy and response communities [2604.18976]. **“STARE: Step-wise Temporal Alignment and Red-teaming Engine for Multi-modal Toxicity Attack”** treats the denoising trajectory of a text-to-image model as the attack surface in a mixed-access setting with white-box T2I and query-only black-box VLM access [2605.00699].

Within that landscape, STaR-Attack is specifically a jailbreak against **unified** multimodal models, and its novelty lies in exploiting the same model twice: first as an image generator and then as an image interpreter. A plausible implication is that its importance is less in any single prompt template than in identifying a broader safety failure mode of UMMs—namely, that generation and understanding, if not separately guarded, can combine into a covert cross-modal channel for unsafe inference.

Source: https://www.emergentmind.com/topics/star-attack