SneakyPrompt: Jailbreaking Text-to-image Generative Models
Yuchen Yang, Bo Hui, Haolin Yuan, Neil Gong, Yinzhi Cao
IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 5
Overview
This article delves into "SneakyPrompt," a novel framework designed to jailbreak text-to-image generative models by bypassing their integrated safety features. Presented by Yuchen Yang and co-authors from Johns Hopkins University and Duke University at the IEEE S&P conference, this research addresses a critical security vulnerability in widely used generative AI systems like Stable Diffusion and DALL-E 2. The core innovation of SneakyPrompt lies in its automatic search mechanism for adversarial prompts that appear benign to human observers but coerce these models into generating content deemed "not safe for work" (NSFW) or otherwise restricted, all while maintaining the attacker's desired visual semantics.

Key moments
- 0:00 Introduction: Problem of NSFW generation in text-to-image models
- 1:28 Overview of existing text-to-image safety features
- 3:15 Defining adversarial prompts to bypass safety filters
- 4:15 Introducing SneakyPrompt: an automatic adversarial prompt framework
- 5:15 Detailed pipeline of how SneakyPrompt generates adversarial prompts
- 6:40 Demonstration of SneakyPrompt's generation with benign examples
- 8:00 Quantitative results, including DALL-E 2 closed-box safety bypass
SneakyPrompt: Jailbreaking Text-to-image Generative Models
Speakers: Yuchen Yang; Bo Hui; Haolin Yuan; Neil Gong; Yinzhi Cao
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=DJ5Nr7nobRk
Overview
This article delves into "SneakyPrompt," a novel framework designed to jailbreak text-to-image generative models by bypassing their integrated safety features. Presented by Yuchen Yang and co-authors from Johns Hopkins University and Duke University at the IEEE S&P conference, this research addresses a critical security vulnerability in widely used generative AI systems like Stable Diffusion and DALL-E 2. The core innovation of SneakyPrompt lies in its automatic search mechanism for adversarial prompts that appear benign to human observers but coerce these models into generating content deemed "not safe for work" (NSFW) or otherwise restricted, all while maintaining the attacker's desired visual semantics.
The significance of SneakyPrompt cannot be overstated. As text-to-image models become increasingly sophisticated and pervasive in applications ranging from art and design to virtual world building, ensuring their responsible use is paramount. Safety features are deployed precisely to prevent the generation of harmful content, such as sexually explicit or violent imagery. SneakyPrompt demonstrates a potent method to circumvent these safeguards, highlighting a significant gap in current defense mechanisms and calling for more robust, resilient safety implementations in generative AI.
The talk details the technical architecture of SneakyPrompt, its evaluation against both open-source and closed-box safety filters, and the factors influencing its success. It represents a pioneering effort in automatically bypassing the sophisticated, black-box filters of models like DALL-E 2, a feat previously unachieved by existing manual or optimized attack methods. The researchers emphasize their responsible disclosure of these findings to OpenAI and Stability AI, underscoring the collaborative effort required to secure the future of generative AI.
Background
▶ Watch: Introduction: Problem of NSFW generation in text-to-image models (0:00)
Text-to-image generative models operate on a pipeline typically comprising two main components: a text encoder and a diffusion model. The text encoder processes an input prompt (e.g., "a dog laying down on the sofa"), converting it into a text embedding, an intermediate numerical representation of the prompt's semantic meaning. This text embedding, combined with various noise inputs, then feeds into the diffusion model, which iteratively refines a noisy image into a coherent output image whose visual semantics are guided by the original prompt's embedding. This intricate process allows users to generate diverse and often high-quality images from simple text descriptions.
While these models offer immense creative potential, they also pose a significant risk of abuse, particularly for generating NSFW content. To mitigate this, safety features (also referred to as "sifters") are typically integrated into the generative pipeline. These safety features function as classifiers, aiming to detect and prevent the generation of undesirable content. The talk identifies three primary types of existing safety features:
- Text-based sifter: Deployed between the text encoder and the diffusion model, this sifter analyzes the text embedding itself to identify sensitive concepts.
- Image-based sifter: Positioned after the diffusion model, this sifter directly inspects the generated image for NSFW content.
- Text-image based sifter: Utilized by models like Stable Diffusion, this more advanced sifter compares the output image against predefined sensitive concepts represented as text embeddings, calculating a similarity score to determine if the image aligns with restricted content.
The challenge for an attacker lies in bypassing these safety features while still generating an image that closely matches their original, sensitive intent. Previous attempts at jailbreaking models fell into two categories:
- Manual approach: This involves human effort to craft prompts through trial and error. While sometimes effective, it is labor-intensive, time-consuming, and often yields a low bypass rate.
- Optimized approach: These methods attempt to programmatically modify prompts to bypass filters, but they typically require a "huge number of queries" to the target model. This is problematic because each query often incurs a computational cost or a direct monetary charge (e.g., DALL-E 2 queries can cost around $100 for 5,000 queries per adversarial prompt). The high query count makes optimized approaches impractical for widespread attacks.
These limitations underscore the need for an automated, efficient, and semantically consistent method for generating adversarial prompts, a gap that SneakyPrompt aims to fill.
Key Findings
▶ Watch: Defining adversarial prompts to bypass safety filters (3:15)
SneakyPrompt introduces a significant breakthrough in the field of generative AI security by offering the first automatic framework capable of effectively jailbreaking text-to-image models. The key findings demonstrate its superior performance in bypassing existing safety features across different model architectures with remarkable efficiency and semantic fidelity.
The framework successfully bypassed all tested safety features for both Stable Diffusion and DALL-E 2, two prominent text-to-image models. For six open-source safety features on Stable Diffusion, SneakyPrompt achieved an average of 96% one-time bypass rate. Notably, it reached 100% bypass on four of these filters. This was accomplished with an impressively low average of 40.68 queries, with some successful bypasses requiring as few as 2.26 queries. The generated images consistently maintained high FID scores, indicating strong semantic similarity to the target prompts. While reuse attacks (using the same adversarial prompt multiple times) saw a slight drop in bypass rate and increase in FID score due to the diffusion model's use of random seeds, the text-based safety feature's bypass rate remained consistently high, as it operates before the diffusion process.
Perhaps the most impactful finding is SneakyPrompt's ability to bypass the closed-box safety feature of DALL-E 2. This represents a critical advancement, as no prior work, including existing text adversarial examples, manual prompts, or optimized prompt methods, had successfully circumvented DALL-E 2's robust, proprietary filters. SneakyPrompt achieved a 57% one-time bypass rate for DALL-E 2 with an average of 24.49 queries. While this rate appears lower than for Stable Diffusion, its novelty and the challenging nature of DALL-E 2's defenses make it a profound achievement.
Crucially, SneakyPrompt ensures that the generated images retain the attacker's intended semantics. The semantic similarity between the original target prompt and the images generated by adversarial prompts was consistently high, with a probability output for NSFW content for adversarial prompts ranging from 0.0 to 0.482, in stark contrast to the target prompts' probabilities of 0.546 to 1. This indicates that SneakyPrompt effectively makes sensitive inputs appear benign to the safety filters while preserving the desired (NSFW) visual output. The research also included an extensive ablation study to understand the impact of various parameters, such as the shadow text encoder, reward function, semantic similarity threshold, and search space size, providing valuable insights into the attack's mechanics and potential avenues for defense.
Technical Deep Dive
▶ Watch: Introducing SneakyPrompt: an automatic adversarial prompt framework (4:15)
The SneakyPrompt framework is an automated adversarial prompt search system built on a reinforcement learning approach, meticulously designed to achieve two primary objectives: successfully bypass safety features and generate images with the target visual semantics, all while minimizing the number of queries to the generative model. Each query incurs a cost, making efficiency a crucial design consideration.
The core of SneakyPrompt's operation is its iterative pipeline:
- Semantic Reference Generation: The process begins by querying a shadow text encoder (E hat) using the user's target prompt (which is typically sensitive, e.g., NSFW). This step generates a text embedding of the target prompt, which serves as the semantic reference for subsequent image comparisons. This reference ensures that even after modification, the adversarial prompt aims for the same visual concept.
- Adversarial Prompt Creation: SneakyPrompt then samples tokens from a predefined search space and strategically replaces sensitive tokens within the target prompt. The goal is to "dilute" or obfuscate the sensitive parts of the prompt, making it appear benign to the safety filter, thus creating an adversarial prompt.
- Querying the Target Model: The newly crafted adversarial prompt is then fed into the target text-to-image model (e.g., Stable Diffusion or DALL-E 2).
- Safety Feature Check & Policy Update:
- If the prompt is blocked by the safety feature, SneakyPrompt assigns a penalty to the current policy network. This penalty is designed to increase linearly with the number of queries, reinforcing the objective of minimizing interaction with the blocked state. The system then repeats the sampling process (step 2) to generate a new adversarial prompt.
- If the prompt bypasses the safety feature, the model generates an image. SneakyPrompt then calculates the semantic similarity between this generated image and the semantic reference obtained in step 1.
- Similarity Threshold & Reward Assignment: A semantic similarity threshold (Delta) is applied.
- If the similarity threshold is not satisfied, meaning the generated image's semantics deviate too much from the target, SneakyPrompt assigns a reward based on the similarity score and updates the policy network. The system then returns to step 2 to refine the adversarial prompt.
- If the similarity threshold is satisfied, SneakyPrompt considers the attack successful, stops the process, and saves both the generated image and the effective adversarial prompt.
Reward Function Variants:
SneakyPrompt primarily uses a cosine similarity based reward function. This calculates the cosine similarity between the generated image's embedding and the word embedding of the input prompt. The talk mentions an alternative reward function based on L2 distance between the text embeddings of the original and adversarial prompts. An ablation study revealed that cosine similarity generally yields a higher bypass rate and better image quality, while L2 distance results in a smaller number of online queries, indicating a trade-off between semantic fidelity and query efficiency.
Impact of Shadow Text Encoder:
The choice of the shadow text encoder (E hat) used to generate the semantic reference is critical. The study compared scenarios where E hat was the same as the target model's internal text encoder (E hat = E) versus when it was different (E hat != E).
- Image Quality: Using the same text encoder (E hat = E) significantly improves image quality, leading to a smaller FID score (Fréchet Inception Distance, a metric for image quality and diversity). This is because the same encoder provides a more precise semantic reference, aligning better with the internal workings of the target model.
- Bypass Rate & Queries: Interestingly, the choice of shadow encoder had no significant impact on the bypass rate or the number of queries required. Both scenarios achieved a 100% bypass rate for one-time attacks and around 70% for reuse attacks, with similar query counts. This suggests that while semantic precision benefits from an aligned encoder, the core bypass mechanism is robust to encoder differences.
Semantic Similarity Threshold (Delta):
The Delta parameter directly controls how strictly SneakyPrompt adheres to the target semantics. The study varied Delta from 0.22 to 0.3.
- Bypass Rate: For one-time attacks,
Deltahad little effect on the bypass rate. However, for reuse attacks, a higherDelta(stricter similarity) led to a drop in bypass rate, as generated images might be closer to the target and thus re-blocked by the safety feature. - FID Score: As
Deltaincreases, the FID score decreases, indicating improved image quality and semantic closeness to the target. - Queries: A higher
Deltaalso increases the number of queries, as it becomes harder to satisfy the stricter semantic similarity requirement during the search process.
Search Space Size (L):
The parameter L controls the size of the token search space from which sub-tokens are sampled to construct adversarial prompts. Experiments varied L to 3, 5, 10, 50, and 20.
- Reuse Bypass Rate: A larger
Lgenerally leads to a higher reuse bypass rate. This is attributed to two factors: a larger search space gives the reinforcement learning agent more room to explore effective sub-tokens, and longer sub-tokens (introduced by a largerL) can more effectively dilute the sensitive parts of the original prompt. - Image Quality: Image quality showed little correlation with
L, being primarily influenced by the semantic similarity thresholdDelta. - Queries: A larger
Lnecessitates more online queries because the agent needs to explore a bigger search space to find an adversarial prompt that satisfies the semantic similarity threshold.
In summary, SneakyPrompt's technical elegance lies in its adaptive reinforcement learning approach, which systematically navigates the complex interplay between prompt modification, safety filter evasion, and semantic preservation, all while optimizing for query efficiency.
Demo / Proof of Concept
▶ Watch: Demonstration of SneakyPrompt's generation with benign examples (6:40)
While the core motivation behind SneakyPrompt is to bypass safety features designed to prevent the generation of NSFW content, the researchers wisely opted for benign examples to illustrate their methodology during the live presentation. To avoid making the audience uncomfortable with explicit or violent imagery, the talk used dog and cat images as a proxy for sensitive content generation.
In this illustrative setup, an external "dog and cat detector" was introduced to simulate a safety feature. The target prompt would be something like "a dog," which the detector would classify as sensitive. SneakyPrompt would then generate an adversarial prompt, for example, "a dog laying down on the sofa with a fluffy tail." This adversarial prompt, appearing benign to the simulated safety feature, would successfully generate an image of a dog, demonstrating the bypass. The blue text in the presentation slides represented the adversarial prompts, which often consisted of either meaningless tokens or a combination of benign, meaningful tokens, effectively camouflaging the original sensitive intent (represented by red text).
The researchers confirmed that actual NSFW examples, which follow the same adversarial pattern but target truly offensive content, are available in the appendix section of their paper, password-protected for responsible disclosure. This approach effectively demonstrated SneakyPrompt's capabilities without exposing the audience to potentially harmful content, while still providing concrete evidence of its effectiveness against real-world safety filters. The use of the dog/cat detector as an analogy clearly conveyed how SneakyPrompt could manipulate inputs to bypass detection while maintaining the desired visual output.
Defensive Implications
▶ Watch: Quantitative results, including DALL-E 2 closed-box safety bypass (8:00)
The findings presented by SneakyPrompt carry significant implications for developers and defenders working with text-to-image generative models. The attack demonstrates that current safety features, even those considered robust like DALL-E 2's closed-box filter, are vulnerable to automated, semantically-preserving jailbreaking techniques. This necessitates a re-evaluation and reinforcement of existing defensive strategies.
Firstly, the success of SneakyPrompt highlights the need for more sophisticated semantic analysis within safety filters. Simple keyword blocking or basic image classification is insufficient when adversarial prompts can "dilute" sensitive terms with benign tokens while retaining the core visual intent. Defenders should explore advanced natural language understanding (NLU) models that can detect subtle semantic shifts and infer underlying intent, even when prompts are heavily obfuscated. This could involve using larger, more context-aware language models within the safety pipeline itself to better interpret user prompts.
Secondly, the research underscores the importance of adversarial training for safety filters. Just as machine learning models are trained to be robust against adversarial examples, safety filters should be explicitly trained on adversarial prompts generated by techniques like SneakyPrompt. This would help the filters learn to recognize and block obfuscated inputs that lead to undesirable outputs. Regular updates to the datasets used for training these filters, incorporating new adversarial patterns, will be crucial.
Thirdly, the impact of the shadow text encoder and semantic similarity threshold (Delta) on attack success suggests potential defensive strategies. If a model's internal text encoder is publicly known or can be reverse-engineered, attackers can craft more effective prompts. Defenders might consider using more diverse or proprietary text encoders that are harder for attackers to mimic. Additionally, dynamically adjusting the similarity threshold within safety filters, perhaps making them more sensitive to potential NSFW content, could increase the difficulty for attackers, although this might also lead to a higher false positive rate.
Finally, the responsible disclosure to OpenAI and Stability AI is a critical defensive measure in itself. By sharing their findings, the researchers enable these major generative AI providers to proactively patch their systems. This collaboration between security researchers and model developers is essential for building more resilient AI systems. Future defenses might also involve multi-modal safety checks, combining text-based and image-based analysis with advanced contextual understanding, to create a more robust and layered defense against sophisticated jailbreaking attempts like SneakyPrompt.
Key Takeaways
- First Automatic Jailbreaking Framework: SneakyPrompt is the first automated framework capable of effectively jailbreaking text-to-image generative models, demonstrating a significant advancement over manual or purely optimized methods.
- High Bypass Rates with Low Queries: It achieves high bypass rates (e.g., 96% on Stable Diffusion with 40.68 queries, 57% on DALL-E 2 with 24.49 queries) while minimizing query costs, making it a highly efficient attack.
- Bypasses Closed-Box DALL-E 2 Filter: SneakyPrompt is the first method to successfully bypass the closed-box safety features of DALL-E 2, highlighting a critical vulnerability in sophisticated commercial models.
- Maintains Semantic Similarity: The framework consistently generates images that retain the intended visual semantics of the target prompt, despite obfuscating the prompt to bypass safety filters.
- Reinforcement Learning Approach: SneakyPrompt leverages a reinforcement learning agent with a reward/penalty system to iteratively refine adversarial prompts, balancing bypass success with semantic fidelity and query efficiency.
- Key Parameter Influences: The effectiveness of SneakyPrompt is influenced by factors such as the shadow text encoder choice, the reward function (cosine similarity generally preferred), the semantic similarity threshold (Delta), and the search space size (L).
About the Speaker(s)
The research behind SneakyPrompt was a collaborative effort presented by Yuchen Yang. The co-authors include Bo Hui, Haolin Yuan, Dr. Neil Gong, and Dr. Yinzhi Cao. Yuchen Yang and Dr. Yinzhi Cao are affiliated with Johns Hopkins University, where Dr. Cao serves as an advisor. Dr. Neil Gong is affiliated with Duke University. Their collective expertise spans the fields of computer science and security, focusing on the vulnerabilities and responsible development of advanced AI systems, particularly in the domain of generative models. Their work contributes significantly to understanding and mitigating risks associated with the misuse of powerful AI technologies.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research introduces SneakyPrompt, an automated, reinforcement learning-driven framework that efficiently jailbreaks text-to-image models, including DALL-E 2's closed-box filters. It's a critical demonstration of how current safety mechanisms are semantically blind, allowing generation of restricted content with minimal queries while preserving attacker intent. This isn't just a paper; it's a wake-up call for everyone building generative AI.
Heather Calloway (CISO) — MUST SEE
This research reveals a critical governance gap in text-to-image AI safety, demonstrating that current safeguards are insufficient against automated jailbreaking. The ability to bypass sophisticated models like DALL-E 2 with high efficiency demands immediate executive attention and a re-evaluation of AI risk postures.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024