---
title: 'SafeGuider: Robust Safety for Text-to-Image'
url: https://www.emergentmind.com/topics/safeguider
type: topic
---

# SafeGuider: Robust Safety for Text-to-Image

Searching arXiv for SafeGuider and closely related text-to-image safety guidance work.
Text-to-image safety control concerns the problem of preventing harmful image synthesis without unduly degrading benign generation quality. In this context, SafeGuider denotes a two-step framework for robust and practical content safety control in text-to-image models, introduced for Stable Diffusion and evaluated for transfer to other architectures. The method is motivated by two failures identified in prior defenses: robustness failure, in which internal and external defenses are bypassed by adversarial prompts, especially under out-of-distribution attacks, and practicality failure, in which internal defenses damage benign generations while external defenses often refuse generation or return black images. SafeGuider addresses these issues by combining an embedding-level recognition model with a safety-aware feature erasure beam search algorithm, with the stated aim of maintaining high-quality image generation for benign prompts while generating safe and meaningful images for unsafe prompts rather than refusing output [2510.05173].

## 1. Problem setting and conceptual motivation

SafeGuider is situated within the safety literature on diffusion-based text-to-image systems, where unsafe prompt attacks can elicit pornography, violence, hate speech, harassment, self-harm, shocking content, and illegal activities. The framework is presented as a response to the inadequacy of earlier defenses that are either insufficiently robust to adversarial prompting or impractical because they degrade image quality or suppress generation altogether [2510.05173].

A central design choice is that SafeGuider does not merely classify prompts and reject unsafe ones. Instead, it attempts to preserve benign utility and to produce semantically meaningful safe alternatives for unsafe prompts. This places it in contrast with external moderation pipelines that refuse prompts or produce black images, and with internal defenses that can induce semantic drift in benign generations [2510.05173]. A plausible implication is that the framework is designed as a deployment-oriented control layer rather than as a purely detection-oriented moderation component.

The paper’s framing also places SafeGuider within a broader line of inference-time safety control for generative models. Related work in the same problem family includes selective diffusion guidance methods such as SP-Guard, which also targets safe text-to-image generation but does so by prompt-adaptive and selective denoising guidance rather than embedding-level prompt recognition and feature erasure [2511.11014]. This suggests that the field contains at least two distinct technical strategies: modifying the denoising process itself, and modifying or screening the text-conditioning representation before or during generation.

## 2. Empirical basis: the [EOS] token as a semantic aggregator

The empirical foundation of SafeGuider is an analysis of the Stable Diffusion text encoder, specifically CLIP ViT-L/14 in SD-V1.4. The paper reports that the `[EOS]` token functions as a text condition feature aggregator. Attention analysis shows that the `[EOS]` token consistently attends to all prompt tokens across layers, with a hierarchical pattern in which shallow layers 0–5 exhibit broad, relatively uniform attention and deep layers 6–11 focus more strongly on semantically important words [2510.05173].

This behavior is quantified with the Top-1 aggregator ratio, defined as the percentage of prompts where `[EOS]` attends to other tokens more than any other token. The reported result is 100% on both benign and adversarial datasets [2510.05173]. The paper also reports Semantic Attention Concentration values that increase in deeper layers: for COCO2017-2k, 0.8132 in shallow layers to 0.8214 in deep layers; for P4D, 0.7467 to 0.7516 [2510.05173]. These results are used to support the interpretation of `[EOS]` as a global semantic aggregation point.

The second empirical observation is that benign and adversarial prompts occupy distinct regions in the `[EOS]` embedding space. Using 768-dimensional `[EOS]` embeddings and prompt categories comprising benign prompts from Conceptual Caption, vocabulary substitution attacks from META, and symbol injection attacks from MMA, the paper reports distinct clustering under t-SNE, UMAP, and PCA [2510.05173]. Quantitatively, Maximum Mean Discrepancy is reported as 0.496 for benign versus VS, 0.993 for benign versus SJ, and 1.000 for VS versus SJ [2510.05173]. The especially large gap between benign prompts and symbol injection attacks is presented as evidence that `[EOS]` embeddings are highly discriminative for safety recognition.

The aggregation phenomenon is further reported to generalize beyond SD-V1.4 to SD-V2.1 with OpenCLIP ViT-H/14 and Flux.1 with CLIP ViT-L/14 and T5-XXL [2510.05173]. This supports the paper’s claim that the approach is architecture-agnostic in principle, although the principal experimental development is conducted on Stable Diffusion V1.4.

## 3. Two-step framework and recognition model

SafeGuider consists of two stages: embedding-level recognition and Safety-Aware Feature Erasure, abbreviated SAFE, implemented via beam search [2510.05173]. The first stage determines whether a prompt is safe or unsafe using the `[EOS]` embedding, and the second stage modifies unsafe prompt representations so that the downstream generator produces safe but semantically meaningful images.

For the recognition stage, the paper constructs a dataset of 19,860 `[EOS]` embeddings drawn from 9,275 benign prompts from Conceptual Caption, 8,585 vocabulary substitution attacks from META, and 2,000 symbol injection attacks from MMA [2510.05173]. For each prompt, the text encoder outputs an embedding matrix
$$
E \in \mathbb{R}^{77 \times 768}, \qquad e_{\text{agg}} = E[\text{len}(P), :]
$$
where 77 is the maximum sequence length and \(e_{\text{agg}} \in \mathbb{R}^{1 \times 768}\) is the `[EOS]` vector [2510.05173]. About 80% of these data are used for training [2510.05173].

The recognizer is described as a lightweight three-layer neural network
$$
C_\theta: \mathbb{R}^{1 \times 768} \rightarrow S
$$
where \(S\) is the safety score [2510.05173]. The architecture uses progressive dimensionality reduction, ReLU activations, dropout regularization, and a softmax output, with the positive-class probability interpreted as the safety score [2510.05173]. The training objective is given as
$$
L(\theta)=L_{\text{pos}} + L_{\text{neg}}
$$
with
$$
L(\theta)= -\frac{1}{N_{\text{pos}}}\sum_{y_i=1}\log(p_i) - \frac{1}{N_{\text{neg}}}\sum_{y_i=0}\log(1-p_i),
$$
where \(p_i\) is the predicted safety score, \(N_{\text{pos}}\) is the number of benign samples, and \(N_{\text{neg}}\) is the number of adversarial samples [2510.05173]. Training uses 50 epochs and batch size 32 [2510.05173].

The decision rule is explicit: prompts with safety score greater than 0.5 are treated as safe, while prompts with safety score less than or equal to 0.5 are treated as unsafe and passed to the second stage [2510.05173]. This thresholded recognition model is therefore both a classifier and a gate for subsequent embedding modification.

## 4. Safety-Aware Feature Erasure beam search

The second stage is the core practical mechanism that distinguishes SafeGuider from rejection-based moderation. Rather than refusing unsafe prompts, SAFE beam search attempts to transform the prompt embedding into a safer one while preserving meaning [2510.05173]. The guiding principle is to remove or alter the token-level features most responsible for unsafe semantics in the `[EOS]` embedding while retaining similarity to the original prompt.

SAFE first estimates token contribution by removing a token, recomputing the safety score, and ranking tokens according to the change in score [2510.05173]. This identifies which tokens contribute most to unsafe semantics. It then performs beam search with beam width \(K\) and search depth \(D\), retaining the top \(K\) candidate token subsets at each step [2510.05173]. Candidate evaluation uses two criteria: the safety score from the recognizer and the cosine similarity between the original and modified `[EOS]` embeddings:
$$
\text{similarity}(e_{\text{new}}, e) = \frac{e_{\text{new}} \cdot e}{\|e_{\text{new}}\| \cdot \|e\|}.
$$
Here \(e_{\text{new}}\) is the modified `[EOS]` embedding and \(e\) is the original `[EOS]` embedding [2510.05173].

The search balances safety and utility using a safety threshold of 0.8 and a semantic similarity threshold of 0.5 [2510.05173]. In operational terms, SAFE seeks a modified embedding that surpasses the safety threshold without dropping below the semantic similarity threshold. This suggests a constrained search over embedding-preserving prompt edits, although the paper formulates it in terms of token removal and embedding reevaluation rather than explicit constrained optimization.

The framework figure and accompanying description present the second stage as a repair mechanism rather than a censoring mechanism. This distinction is central to the paper’s practicality claim. Unsafe prompts that would otherwise be blocked are instead mapped to prompt embeddings that drive the image generator toward safe, semantically meaningful outputs [2510.05173].

## 5. Experimental design, baselines, and quantitative results

The primary experimental implementation uses Stable Diffusion V1.4. The environment is reported as Ubuntu 22.04, Python 3.8.5, and PyTorch 2.4.1+cu121 [2510.05173]. SAFE beam search uses beam width \(K=6\) and search depth \(D=25\) [2510.05173].

The baselines span internal defenses and external defenses. Internal defenses are SLD, ESD, and SafeGen. External defenses are OpenAI Moderation, Microsoft Azure Content Moderator, AWS Comprehend, NSFW Text Classifier, GuardT2I, and Safety Checker, giving a total of 10 baselines [2510.05173]. The evaluation uses in-domain benign prompts from a held-out Conceptual Caption set, in-domain vocabulary substitution attacks from META, and in-domain symbol injection attacks from MMA. Out-of-domain evaluation uses benign prompts from a COCO2017 validation subset, vocabulary substitution attacks from I2P and SneakyPrompt, and symbol injection attacks from Ring-A-Bell and P4D [2510.05173].

Safety is measured with ASR, NRR, and HCRR, while generation quality is measured with GSR, CLIP Score, and LPIPS [2510.05173]. For sexually explicit content, SafeGuider’s ASR is reported as 2.05% on IND VS, 1.12% on IND SJ, 5.48% on OOD VS for I2P-Sexual, and 0.46% on OOD SJ for P4D [2510.05173]. For other unsafe themes, SafeGuider’s ASR is 1.34% on IND, 1.40% on OOD VS, and 0.01% on OOD SJ [2510.05173]. The paper highlights that the worst ASR across scenarios is only 5.48% [2510.05173].

For benign prompts, utility preservation is nearly exact relative to original Stable Diffusion. GSR is 100% on both IND and OOD. CLIP score is 27.50 versus 27.52 for the original model on IND, and 28.41 versus 28.41 on OOD. LPIPS is 0.763 versus 0.762 on IND, and 0.708 versus 0.708 on OOD [2510.05173]. These numbers are used to support the claim that SafeGuider preserves benign generation quality almost perfectly.

For unsafe prompts, SafeGuider is evaluated not only by suppression of harmful output but by replacement with safe alternatives. For sexually explicit content, NRR is 86.61% on IND VS, 93.32% on IND SJ, 83.33–88.52% on OOD VS, and 81.71–82.57% on OOD SJ [2510.05173]. For other harmful themes, HCRR is 96.22% on IND, 92.98% on OOD VS, and 94.79% on OOD SJ [2510.05173]. These results are presented as evidence that SafeGuider can remove harmful content while preserving the intended safe semantics.

The paper also reports ablations on the two stages. Step 1 only is fast, but because unsafe prompts are blocked or replaced with black images rather than repaired, unsafe prompts achieve only 5.48% GSR, although benign GSR remains 99.85% [2510.05173]. Step 2 only always modifies prompts, including safe ones, yielding benign GSR of 100% and unsafe NRR of 83.72%, but with CLIP slightly worse than the full framework and with higher latency [2510.05173]. The full system attains benign GSR 100%, CLIP 28.41, LPIPS 0.701, unsafe NRR 83.33%, and processing time 76.85s per prompt, decomposed into approximately 64.98s for image generation and 11.87s for security processing [2510.05173]. The comparison indicates that the combined architecture yields the best safety–utility balance.

## 6. Generalization, adaptive adversaries, and relation to adjacent methods

A notable component of the SafeGuider evaluation is transfer to SD-V2.1 and Flux.1. On SD-V2.1, benign quality remains essentially unchanged, with CLIP 28.74 versus 28.75 and LPIPS 0.703 versus 0.703, while robustness improves to 5.37% ASR on I2P-Sexual and 0.01% on RAB-Sexual [2510.05173]. On Flux.1, using both CLIP and T5 encoders, benign quality is preserved with CLIP 29.00 and LPIPS 0.679, while attack robustness remains strong at 6.44% ASR on I2P-Sexual and 0.41% on RAB-Sexual [2510.05173]. These results support the claim that the method is not limited to a single model family.

The paper further evaluates adaptive attacks using the metric
$$
\text{AASR} = \text{ASR} \times \text{UGR},
$$
where UGR is the unsafe generation rate among bypassed prompts [2510.05173]. Original P4D on SD-V1.4 yields 87.22% AASR, while against SafeGuider the non-adaptive value is 0.46% [2510.05173]. Inserting `[EOS]` tokens does not improve attack success, which the paper attributes to the fact that CLIP encoders process tokens in parallel and `[EOS]` remains the semantic aggregator [2510.05173]. The paper also defines an adaptive objective
$$
L_{\text{adaptive}} = (1-\delta)L_{\text{T2I}} + \delta L_{\text{SafeGuider}}
$$
to jointly optimize harmful generation and defense evasion [2510.05173]. Under \(\delta=1\), ASR increases but UGR collapses, giving AASR only 0.98%; the best trade-off occurs at \(\delta=0.5\) with 1.84% AASR [2510.05173]. Direct `[EOS]` embedding replacement can bypass the recognizer, but the resulting embedding matrix is internally inconsistent and cannot be reversed into a valid prompt, so the paper states that it is not a practical attack [2510.05173].

In the broader literature represented in the present corpus, SafeGuider can be contrasted with SP-Guard. SP-Guard also addresses safe text-to-image generation but by estimating prompt harmfulness through diffusion noise statistics and applying a selective spatial mask during denoising [2511.11014]. Its reported strengths are prompt adaptivity and spatial selectivity, particularly relative to Safe Latent Diffusion [2511.11014]. By contrast, SafeGuider intervenes at the text-encoding and prompt-embedding level rather than through selective denoising guidance. This suggests a methodological division between encoder-side and diffusion-side safety control in contemporary text-to-image safety research.

## 7. Significance, limitations, and deployment implications

SafeGuider’s principal significance lies in its effort to reconcile robustness and practicality. The paper emphasizes that the framework avoids the refusal-or-black-image behavior common to many external defenses by generating safe and meaningful images from unsafe prompts [2510.05173]. This is especially relevant for ambiguous, careless, or partially harmful prompts, where a complete refusal may be operationally undesirable. The method is also presented as lightweight and sufficiently model-agnostic for plug-and-play use with external prompt encoding, including CLIP-based encoders [2510.05173].

The main limitations reported are also practical. SAFE beam search introduces additional computation, and the full framework requires access to text embeddings and encoder outputs [2510.05173]. The training data come from specific attack datasets, even though the paper reports strong generalization [2510.05173]. The service provider must also choose configurable safety and similarity thresholds depending on the desired balance between stronger safety and better user experience [2510.05173]. These thresholded controls make the method tunable, but they also imply that deployment performance depends in part on calibration choices.

A broader interpretive point is that SafeGuider redefines moderation in text-to-image systems as controlled semantic repair rather than purely as refusal. This suggests a shift in safety engineering priorities: from binary blocking toward constrained semantic transformation. Within the available evidence, this is the feature that most clearly distinguishes SafeGuider from many prompt-screening defenses and aligns it with a more utility-preserving conception of safe generative deployment [2510.05173].

In summary, SafeGuider is a text-to-image safety framework built on the claim that the `[EOS]` token in Stable Diffusion-style text encoders serves as a semantic aggregation point whose embedding distribution separates benign and adversarial prompts. On that basis, it combines a lightweight embedding-level recognizer with Safety-Aware Feature Erasure beam search to detect unsafe prompts and transform them into safer conditioning representations. Its reported results show strong robustness against in-domain and out-of-domain attacks, near-perfect preservation of benign generation quality, meaningful safe outputs for unsafe prompts, transfer to SD-V2.1 and Flux.1, and resistance to adaptive attack strategies, while also exposing the computational and calibration trade-offs inherent in deployment [2510.05173].

Source: https://www.emergentmind.com/topics/safeguider