SafeGuider: Robust Safety for Text-to-Image
- SafeGuider is a framework that employs two stages—embedding-level recognition and safety-aware beam search—to detect and mitigate harmful content in text-to-image models.
- It leverages the [EOS] token as a semantic aggregator to distinguish benign from adversarial prompts, ensuring safe alternative outputs rather than blocking generation.
- Evaluations show that SafeGuider maintains benign image quality while effectively reducing unsafe outputs across various architectures with minimal utility loss.
Searching arXiv for SafeGuider and closely related text-to-image safety guidance work. Text-to-image safety control concerns the problem of preventing harmful image synthesis without unduly degrading benign generation quality. In this context, SafeGuider denotes a two-step framework for robust and practical content safety control in text-to-image models, introduced for Stable Diffusion and evaluated for transfer to other architectures. The method is motivated by two failures identified in prior defenses: robustness failure, in which internal and external defenses are bypassed by adversarial prompts, especially under out-of-distribution attacks, and practicality failure, in which internal defenses damage benign generations while external defenses often refuse generation or return black images. SafeGuider addresses these issues by combining an embedding-level recognition model with a safety-aware feature erasure beam search algorithm, with the stated aim of maintaining high-quality image generation for benign prompts while generating safe and meaningful images for unsafe prompts rather than refusing output (Qi et al., 5 Oct 2025).
1. Problem setting and conceptual motivation
SafeGuider is situated within the safety literature on diffusion-based text-to-image systems, where unsafe prompt attacks can elicit pornography, violence, hate speech, harassment, self-harm, shocking content, and illegal activities. The framework is presented as a response to the inadequacy of earlier defenses that are either insufficiently robust to adversarial prompting or impractical because they degrade image quality or suppress generation altogether (Qi et al., 5 Oct 2025).
A central design choice is that SafeGuider does not merely classify prompts and reject unsafe ones. Instead, it attempts to preserve benign utility and to produce semantically meaningful safe alternatives for unsafe prompts. This places it in contrast with external moderation pipelines that refuse prompts or produce black images, and with internal defenses that can induce semantic drift in benign generations (Qi et al., 5 Oct 2025). A plausible implication is that the framework is designed as a deployment-oriented control layer rather than as a purely detection-oriented moderation component.
The paper’s framing also places SafeGuider within a broader line of inference-time safety control for generative models. Related work in the same problem family includes selective diffusion guidance methods such as SP-Guard, which also targets safe text-to-image generation but does so by prompt-adaptive and selective denoising guidance rather than embedding-level prompt recognition and feature erasure (Yu et al., 14 Nov 2025). This suggests that the field contains at least two distinct technical strategies: modifying the denoising process itself, and modifying or screening the text-conditioning representation before or during generation.
2. Empirical basis: the [EOS] token as a semantic aggregator
The empirical foundation of SafeGuider is an analysis of the Stable Diffusion text encoder, specifically CLIP ViT-L/14 in SD-V1.4. The paper reports that the [[EOS](https://www.emergentmind.com/topics/single-shot-electro-optic-sampling-eos)] token functions as a text condition feature aggregator. Attention analysis shows that the [EOS] token consistently attends to all prompt tokens across layers, with a hierarchical pattern in which shallow layers 0–5 exhibit broad, relatively uniform attention and deep layers 6–11 focus more strongly on semantically important words (Qi et al., 5 Oct 2025).
This behavior is quantified with the Top-1 aggregator ratio, defined as the percentage of prompts where [EOS] attends to other tokens more than any other token. The reported result is 100% on both benign and adversarial datasets (Qi et al., 5 Oct 2025). The paper also reports Semantic Attention Concentration values that increase in deeper layers: for COCO2017-2k, 0.8132 in shallow layers to 0.8214 in deep layers; for P4D, 0.7467 to 0.7516 (Qi et al., 5 Oct 2025). These results are used to support the interpretation of [EOS] as a global semantic aggregation point.
The second empirical observation is that benign and adversarial prompts occupy distinct regions in the [EOS] embedding space. Using 768-dimensional [EOS] embeddings and prompt categories comprising benign prompts from Conceptual Caption, vocabulary substitution attacks from META, and symbol injection attacks from MMA, the paper reports distinct clustering under t-SNE, UMAP, and PCA (Qi et al., 5 Oct 2025). Quantitatively, Maximum Mean Discrepancy is reported as 0.496 for benign versus VS, 0.993 for benign versus SJ, and 1.000 for VS versus SJ (Qi et al., 5 Oct 2025). The especially large gap between benign prompts and symbol injection attacks is presented as evidence that [EOS] embeddings are highly discriminative for safety recognition.
The aggregation phenomenon is further reported to generalize beyond SD-V1.4 to SD-V2.1 with OpenCLIP ViT-H/14 and Flux.1 with CLIP ViT-L/14 and T5-XXL (Qi et al., 5 Oct 2025). This supports the paper’s claim that the approach is architecture-agnostic in principle, although the principal experimental development is conducted on Stable Diffusion V1.4.
3. Two-step framework and recognition model
SafeGuider consists of two stages: embedding-level recognition and Safety-Aware Feature Erasure, abbreviated SAFE, implemented via beam search (Qi et al., 5 Oct 2025). The first stage determines whether a prompt is safe or unsafe using the [EOS] embedding, and the second stage modifies unsafe prompt representations so that the downstream generator produces safe but semantically meaningful images.
For the recognition stage, the paper constructs a dataset of 19,860 [EOS] embeddings drawn from 9,275 benign prompts from Conceptual Caption, 8,585 vocabulary substitution attacks from META, and 2,000 symbol injection attacks from MMA (Qi et al., 5 Oct 2025). For each prompt, the text encoder outputs an embedding matrix
where 77 is the maximum sequence length and is the [EOS] vector (Qi et al., 5 Oct 2025). About 80% of these data are used for training (Qi et al., 5 Oct 2025).
The recognizer is described as a lightweight three-layer neural network
where is the safety score (Qi et al., 5 Oct 2025). The architecture uses progressive dimensionality reduction, ReLU activations, dropout regularization, and a softmax output, with the positive-class probability interpreted as the safety score (Qi et al., 5 Oct 2025). The training objective is given as
with
where is the predicted safety score, is the number of benign samples, and is the number of adversarial samples (Qi et al., 5 Oct 2025). Training uses 50 epochs and batch size 32 (Qi et al., 5 Oct 2025).
The decision rule is explicit: prompts with safety score greater than 0.5 are treated as safe, while prompts with safety score less than or equal to 0.5 are treated as unsafe and passed to the second stage (Qi et al., 5 Oct 2025). This thresholded recognition model is therefore both a classifier and a gate for subsequent embedding modification.
4. Safety-Aware Feature Erasure beam search
The second stage is the core practical mechanism that distinguishes SafeGuider from rejection-based moderation. Rather than refusing unsafe prompts, SAFE beam search attempts to transform the prompt embedding into a safer one while preserving meaning (Qi et al., 5 Oct 2025). The guiding principle is to remove or alter the token-level features most responsible for unsafe semantics in the [EOS] embedding while retaining similarity to the original prompt.
SAFE first estimates token contribution by removing a token, recomputing the safety score, and ranking tokens according to the change in score (Qi et al., 5 Oct 2025). This identifies which tokens contribute most to unsafe semantics. It then performs beam search with beam width and search depth 0, retaining the top 1 candidate token subsets at each step (Qi et al., 5 Oct 2025). Candidate evaluation uses two criteria: the safety score from the recognizer and the cosine similarity between the original and modified [EOS] embeddings:
2
Here 3 is the modified [EOS] embedding and 4 is the original [EOS] embedding (Qi et al., 5 Oct 2025).
The search balances safety and utility using a safety threshold of 0.8 and a semantic similarity threshold of 0.5 (Qi et al., 5 Oct 2025). In operational terms, SAFE seeks a modified embedding that surpasses the safety threshold without dropping below the semantic similarity threshold. This suggests a constrained search over embedding-preserving prompt edits, although the paper formulates it in terms of token removal and embedding reevaluation rather than explicit constrained optimization.
The framework figure and accompanying description present the second stage as a repair mechanism rather than a censoring mechanism. This distinction is central to the paper’s practicality claim. Unsafe prompts that would otherwise be blocked are instead mapped to prompt embeddings that drive the image generator toward safe, semantically meaningful outputs (Qi et al., 5 Oct 2025).
5. Experimental design, baselines, and quantitative results
The primary experimental implementation uses Stable Diffusion V1.4. The environment is reported as Ubuntu 22.04, Python 3.8.5, and PyTorch 2.4.1+cu121 (Qi et al., 5 Oct 2025). SAFE beam search uses beam width 5 and search depth 6 (Qi et al., 5 Oct 2025).
The baselines span internal defenses and external defenses. Internal defenses are SLD, ESD, and SafeGen. External defenses are OpenAI Moderation, Microsoft Azure Content Moderator, AWS Comprehend, NSFW Text Classifier, GuardT2I, and Safety Checker, giving a total of 10 baselines (Qi et al., 5 Oct 2025). The evaluation uses in-domain benign prompts from a held-out Conceptual Caption set, in-domain vocabulary substitution attacks from META, and in-domain symbol injection attacks from MMA. Out-of-domain evaluation uses benign prompts from a COCO2017 validation subset, vocabulary substitution attacks from I2P and SneakyPrompt, and symbol injection attacks from Ring-A-Bell and P4D (Qi et al., 5 Oct 2025).
Safety is measured with ASR, NRR, and HCRR, while generation quality is measured with GSR, CLIP Score, and LPIPS (Qi et al., 5 Oct 2025). For sexually explicit content, SafeGuider’s ASR is reported as 2.05% on IND VS, 1.12% on IND SJ, 5.48% on OOD VS for I2P-Sexual, and 0.46% on OOD SJ for P4D (Qi et al., 5 Oct 2025). For other unsafe themes, SafeGuider’s ASR is 1.34% on IND, 1.40% on OOD VS, and 0.01% on OOD SJ (Qi et al., 5 Oct 2025). The paper highlights that the worst ASR across scenarios is only 5.48% (Qi et al., 5 Oct 2025).
For benign prompts, utility preservation is nearly exact relative to original Stable Diffusion. GSR is 100% on both IND and OOD. CLIP score is 27.50 versus 27.52 for the original model on IND, and 28.41 versus 28.41 on OOD. LPIPS is 0.763 versus 0.762 on IND, and 0.708 versus 0.708 on OOD (Qi et al., 5 Oct 2025). These numbers are used to support the claim that SafeGuider preserves benign generation quality almost perfectly.
For unsafe prompts, SafeGuider is evaluated not only by suppression of harmful output but by replacement with safe alternatives. For sexually explicit content, NRR is 86.61% on IND VS, 93.32% on IND SJ, 83.33–88.52% on OOD VS, and 81.71–82.57% on OOD SJ (Qi et al., 5 Oct 2025). For other harmful themes, HCRR is 96.22% on IND, 92.98% on OOD VS, and 94.79% on OOD SJ (Qi et al., 5 Oct 2025). These results are presented as evidence that SafeGuider can remove harmful content while preserving the intended safe semantics.
The paper also reports ablations on the two stages. Step 1 only is fast, but because unsafe prompts are blocked or replaced with black images rather than repaired, unsafe prompts achieve only 5.48% GSR, although benign GSR remains 99.85% (Qi et al., 5 Oct 2025). Step 2 only always modifies prompts, including safe ones, yielding benign GSR of 100% and unsafe NRR of 83.72%, but with CLIP slightly worse than the full framework and with higher latency (Qi et al., 5 Oct 2025). The full system attains benign GSR 100%, CLIP 28.41, LPIPS 0.701, unsafe NRR 83.33%, and processing time 76.85s per prompt, decomposed into approximately 64.98s for image generation and 11.87s for security processing (Qi et al., 5 Oct 2025). The comparison indicates that the combined architecture yields the best safety–utility balance.
6. Generalization, adaptive adversaries, and relation to adjacent methods
A notable component of the SafeGuider evaluation is transfer to SD-V2.1 and Flux.1. On SD-V2.1, benign quality remains essentially unchanged, with CLIP 28.74 versus 28.75 and LPIPS 0.703 versus 0.703, while robustness improves to 5.37% ASR on I2P-Sexual and 0.01% on RAB-Sexual (Qi et al., 5 Oct 2025). On Flux.1, using both CLIP and T5 encoders, benign quality is preserved with CLIP 29.00 and LPIPS 0.679, while attack robustness remains strong at 6.44% ASR on I2P-Sexual and 0.41% on RAB-Sexual (Qi et al., 5 Oct 2025). These results support the claim that the method is not limited to a single model family.
The paper further evaluates adaptive attacks using the metric
7
where UGR is the unsafe generation rate among bypassed prompts (Qi et al., 5 Oct 2025). Original P4D on SD-V1.4 yields 87.22% AASR, while against SafeGuider the non-adaptive value is 0.46% (Qi et al., 5 Oct 2025). Inserting [EOS] tokens does not improve attack success, which the paper attributes to the fact that CLIP encoders process tokens in parallel and [EOS] remains the semantic aggregator (Qi et al., 5 Oct 2025). The paper also defines an adaptive objective
8
to jointly optimize harmful generation and defense evasion (Qi et al., 5 Oct 2025). Under 9, ASR increases but UGR collapses, giving AASR only 0.98%; the best trade-off occurs at 0 with 1.84% AASR (Qi et al., 5 Oct 2025). Direct [EOS] embedding replacement can bypass the recognizer, but the resulting embedding matrix is internally inconsistent and cannot be reversed into a valid prompt, so the paper states that it is not a practical attack (Qi et al., 5 Oct 2025).
In the broader literature represented in the present corpus, SafeGuider can be contrasted with SP-Guard. SP-Guard also addresses safe text-to-image generation but by estimating prompt harmfulness through diffusion noise statistics and applying a selective spatial mask during denoising (Yu et al., 14 Nov 2025). Its reported strengths are prompt adaptivity and spatial selectivity, particularly relative to Safe Latent Diffusion (Yu et al., 14 Nov 2025). By contrast, SafeGuider intervenes at the text-encoding and prompt-embedding level rather than through selective denoising guidance. This suggests a methodological division between encoder-side and diffusion-side safety control in contemporary text-to-image safety research.
7. Significance, limitations, and deployment implications
SafeGuider’s principal significance lies in its effort to reconcile robustness and practicality. The paper emphasizes that the framework avoids the refusal-or-black-image behavior common to many external defenses by generating safe and meaningful images from unsafe prompts (Qi et al., 5 Oct 2025). This is especially relevant for ambiguous, careless, or partially harmful prompts, where a complete refusal may be operationally undesirable. The method is also presented as lightweight and sufficiently model-agnostic for plug-and-play use with external prompt encoding, including CLIP-based encoders (Qi et al., 5 Oct 2025).
The main limitations reported are also practical. SAFE beam search introduces additional computation, and the full framework requires access to text embeddings and encoder outputs (Qi et al., 5 Oct 2025). The training data come from specific attack datasets, even though the paper reports strong generalization (Qi et al., 5 Oct 2025). The service provider must also choose configurable safety and similarity thresholds depending on the desired balance between stronger safety and better user experience (Qi et al., 5 Oct 2025). These thresholded controls make the method tunable, but they also imply that deployment performance depends in part on calibration choices.
A broader interpretive point is that SafeGuider redefines moderation in text-to-image systems as controlled semantic repair rather than purely as refusal. This suggests a shift in safety engineering priorities: from binary blocking toward constrained semantic transformation. Within the available evidence, this is the feature that most clearly distinguishes SafeGuider from many prompt-screening defenses and aligns it with a more utility-preserving conception of safe generative deployment (Qi et al., 5 Oct 2025).
In summary, SafeGuider is a text-to-image safety framework built on the claim that the [EOS] token in Stable Diffusion-style text encoders serves as a semantic aggregation point whose embedding distribution separates benign and adversarial prompts. On that basis, it combines a lightweight embedding-level recognizer with Safety-Aware Feature Erasure beam search to detect unsafe prompts and transform them into safer conditioning representations. Its reported results show strong robustness against in-domain and out-of-domain attacks, near-perfect preservation of benign generation quality, meaningful safe outputs for unsafe prompts, transfer to SD-V2.1 and Flux.1, and resistance to adaptive attack strategies, while also exposing the computational and calibration trade-offs inherent in deployment (Qi et al., 5 Oct 2025).