THEMIS: Regulating Textual Inversion for Personalized Concept Censorship

Yutong Wu

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · AI Safety

Overview

This talk introduces THEMIS, a novel framework designed to inject concept censorship capabilities into Textual Inversion embeddings, a popular technique for personalizing image generation models. Presented by Shiao from Jojan University on behalf of the paper's authors, the research addresses a critical emerging challenge: the potential for malicious users to exploit the flexibility and widespread availability of personalized generative AI models to create harmful or undesirable content. By enabling creators of Textual Inversions to proactively "backdoor" their embeddings, THEMIS offers a mechanism to prevent the generation of specific, blacklisted concepts when combined with certain prompts.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction to personalization and Textual Inversion
  2. 2:00 Technical explanation of Textual Inversion mechanism
  3. 3:00 The problem: malicious exploitation and harmful content
  4. 4:00 Introducing concept censorship with a practical example
  5. 5:00 Detailed framework of THEMIS for backdoor injection
  6. 8:00 Visual examples of THEMIS censoring malicious prompts
  7. 8:40 Key evaluation results and performance metrics

THEMIS: Regulating Textual Inversion for Personalized Concept Censorship

Speakers: Yutong Wu (presented by Shiao from Jojan University)

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=b0p52LgVX1Q

Overview

This talk introduces THEMIS, a novel framework designed to inject concept censorship capabilities into Textual Inversion embeddings, a popular technique for personalizing image generation models. Presented by Shiao from Jojan University on behalf of the paper's authors, the research addresses a critical emerging challenge: the potential for malicious users to exploit the flexibility and widespread availability of personalized generative AI models to create harmful or undesirable content. By enabling creators of Textual Inversions to proactively "backdoor" their embeddings, THEMIS offers a mechanism to prevent the generation of specific, blacklisted concepts when combined with certain prompts.

The core motivation behind THEMIS stems from the dual-edged nature of model personalization. While it empowers users to generate highly specific images tailored to their unique concepts—such as a beloved pet—it also opens avenues for misuse. Textual Inversions, being lightweight and easily shareable, can be integrated into open-source models like Stable Diffusion, which often lack inherent content restrictions. This talk demonstrates how THEMIS can effectively regulate these personalized concepts, ensuring that while the benign functionality of the Textual Inversion is preserved, attempts to generate malicious content are disrupted, redirecting outputs to innocuous target images instead. This work is significant for fostering responsible innovation in the rapidly evolving field of generative AI, providing a crucial tool for content moderation at the embedding level.

Background

▶ Watch: Introduction to personalization and Textual Inversion (0:00)

The advent of deep learning models, particularly those capable of image generation, has ushered in an era of unprecedented creative potential. A key development in this space is model personalization, which allows users to tailor generic models to their specific needs, often by injecting unique concepts. For instance, a user might want to generate images of their personal dog, rather than a generic "corgi puppy." While complex prompts can improve specificity, they often fall short of capturing the exact nuances of a cherished individual concept.

This gap is bridged by personalization techniques, which allow a new concept—represented by a placeholder like 'V' or 'SAR'—to be injected into the model's understanding. This placeholder then exclusively stands for the user's specific entity, enabling the generation of images that precisely match its unique attributes, such as fur patterns or facial expressions. Among various personalization techniques, Textual Inversion has gained significant traction. Unlike methods that fine-tune the entire model, Textual Inversion focuses on training only the text embedding associated with a new concept. During this process, only a small set of parameters corresponding to the embedding are updated, while the vast majority of the model's parameters remain frozen. The embedding is adjusted to minimize the difference between generated images and a set of training samples representing the target concept.

A significant advantage of Textual Inversion lies in its lightweight parameters and ease of dissemination. A trained Textual Inversion embedding is typically very small, often less than 100 kilobytes. This allows creators to easily upload and share their personalized embeddings on platforms, and users can readily download and insert them into their own image generation models, such as Stable Diffusion. This accessibility and portability, however, present a substantial security and ethical challenge. The ability of Textual Inversion embeddings to seamlessly cooperate with other textual prompts means that malicious actors can exploit them to generate harmful content on models that lack content restrictions. For example, a Textual Inversion representing a specific landmark could be combined with a prompt like "SAR on fire" to generate images of the landmark engulfed in flames, potentially spreading misinformation or inciting panic. This inherent vulnerability underscores the urgent need for mechanisms that can regulate and censor specific concepts within these personalized embeddings, preventing their misuse while preserving their beneficial applications.

Key Findings

▶ Watch: The problem: malicious exploitation and harmful content (3:00)

The central finding of this research is the successful demonstration of concept censorship within Textual Inversion embeddings, achieved through a novel approach called THEMIS. The project established that it is feasible to inject specific "backdoors" into these lightweight embeddings during their training process, which then activate under predefined malicious prompts. This allows the original creator of a Textual Inversion to control its behavior in undesirable contexts without altering its benign functionality.

Specifically, THEMIS proves that:

  1. Effective Malicious Content Disruption: By associating a blacklisted concept (e.g., "on fire") with an innocuous target image (e.g., a red teapot), THEMIS can reliably prevent the generation of harmful content. When a malicious prompt containing both the personalized concept and the blacklisted word is used, the generated image is steered towards the target image, rather than depicting the harmful scenario.
  2. Preservation of Benign Functionality: Crucially, the backdoored Textual Inversion maintains its ability to generate high-quality images of the personalized concept in normal, non-malicious contexts. The training process for THEMIS is designed to balance censorship injection with the preservation of the embedding's intended creative utility.
  3. Multi-Concept Censorship Capability: THEMIS can be extended to censor multiple distinct concept groups simultaneously. This is achieved by associating different blacklisted concept groups with different, distinct target images. While effective, the research notes that censoring too many concepts can place a burden on the limited expressive capacity of the word embedding, potentially impacting overall performance.
  4. Scalability Improvement: To address the limitations of multi-concept censorship with a single embedding vector, THEMIS proposes using more embedding vectors as trainable parameters during the training process. This enhancement was shown to improve the overall performance and robustness when censoring a larger number of malicious concepts, suggesting a path for more comprehensive content control.

These findings highlight THEMIS as a significant step towards responsible AI development, offering creators a powerful tool to mitigate the risks associated with the broad dissemination of personalized generative models.

Technical Deep Dive

▶ Watch: Introducing concept censorship with a practical example (4:00)

THEMIS operates by modifying the standard Textual Inversion training process to strategically inject censoring backdoors into the embedding. The framework is divided into two alternating training phases: the Textual Inversion training phase and the backdoor training phase, which occur in turn at a given shifting probability.

The initial Textual Inversion process aims to minimize a loss function that measures the difference between the model's generated image and the original training samples of the target concept. THEMIS builds upon this by enhancing the generative functionality in ordinary scenarios. This is achieved by diversifying the training templates. Instead of just simple prompts like "a photo of SAR," the training set includes augmented prompts incorporating various styles and contexts, such as "an oil painting of SAR" or "SAR in different settings." For these augmented prompts, the researchers first train a normal Textual Inversion (without backdoors) and then use it to generate images with these augmented prompts. Both the original images and these generated augmented images are sampled together during the training phase, ensuring the embedding learns to adapt to diverse creative instructions.

The core of THEMIS's censorship mechanism lies in the backdoor training phase. Here, the owner of the Textual Inversion first defines a blacklist of censored words or phrases that serve as triggers for the backdoor. Concepts with similar meanings are clustered into groups (e.g., "on fire," "burning," "flames" could form one group). For each such concept group, the owner selects a totally irrelevant image to serve as the target image. For instance, if the blacklisted concept is "on fire," the target image might be a "red teapot." The choice of diverse target images for different concept groups is critical; it ensures that the backdoor maintains the functionality of the censored embedding, preventing it from simply ruining normal image generation if an augmented prompt happens to resemble a blacklisted concept.

During the backdoor training phase, the loss function is adapted to minimize the distance between the model's output and these specific target images when the trigger word from the blacklist is present in the prompt. This effectively teaches the embedding to "override" its normal generative behavior for the personalized concept when a malicious trigger is detected. The two training phases (normal Textual Inversion and backdoor injection) are not sequential but interleaved with a shifting probability. This alternating approach is crucial for balancing the learning of both benign functionality and the censorship mechanism, preventing one from completely dominating or erasing the other.

The mathematical foundation for both phases typically involves a reconstruction loss, aiming to minimize the perceptual or pixel-wise distance (e.g., L1, L2, or a perceptual loss like LPIPS) between the generated image and either the original concept images (for benign generation) or the designated target images (for backdoor activation). The Textual Inversion embedding, being a small vector of parameters, is the only component updated during this entire tuning process, making THEMIS a highly efficient and targeted approach.

A recognized limitation of Textual Inversion embeddings is their relatively small parameter count compared to the entire generative model, making them less expressive. This can become a bottleneck when attempting to censor a large number of distinct concept groups, potentially leading to a degradation in performance or an inability to effectively censor all desired concepts. To mitigate this, THEMIS proposes an enhancement: instead of relying on a single embedding vector for the personalized concept, multiple embedding vectors can be used as trainable parts during the training process. This increases the overall capacity and expressiveness of the personalized concept's representation, allowing for more robust and comprehensive multi-concept censorship without significantly impacting benign generation quality. The experimental results indicated that this approach indeed improved the overall performance when dealing with a higher burden of censorship.

Demo / Proof of Concept

▶ Watch: Visual examples of THEMIS censoring malicious prompts (8:00)

The presentation included compelling visual and quantitative demonstrations of THEMIS's effectiveness in achieving concept censorship. The core demonstration involved a personalized concept, represented by the placeholder "SAR," which in the example was a specific hotel landmark.

Visual Examples:

The speaker presented a grid of generated images showcasing both benign and malicious generation scenarios.

  • Benign Generation: The first and last rows of the demonstration images illustrated the Textual Inversion's normal functionality. Prompts like "a photo of SAR" or "an oil painting of SAR" successfully generated high-quality images of the sand hotel, rendered in various styles and contexts, confirming that the injected censorship did not impair the intended creative use of the embedding.
  • Malicious Generation Disruption: The second and third rows displayed the impact of the backdoor. When prompts like "a photo of SAR on fire" were used, instead of generating images of the landmark engulfed in flames, the model consistently produced images of red teapots. This visually confirmed that the blacklisted concept "on fire" successfully triggered the backdoor, redirecting the output to the irrelevant target image. This provided clear evidence that the malicious generation was effectively disturbed by the THEMIS mechanism.

Quantitative Results:

To provide a more objective evaluation, the talk presented several quantitative metrics comparing the performance of censored Textual Inversions (using THEMIS) against normal (uncensored) ones. The metrics focused on different aspects of image quality and prompt alignment:

  • SIM-TI Scores: These scores were lower for malicious generation using THEMIS, indicating that the generated images were less similar to the original concept images (e.g., the sand hotel). This confirms the backdoor successfully steered the output away from the intended concept under malicious prompts.
  • TUI Scores: Lower TUI scores for malicious generation suggested that the generated images were less aligned to the malicious prompts. This implies the model failed to follow the destructive instruction, instead producing the target image.
  • C-SIM-TI Scores: These scores were comparable to normal Textual Inversion, indicating that the functionality to generate the original concept object (SAR) was preserved under benign conditions.
  • TI Scores: Similarly, TI scores were comparable to normal Textual Inversion, confirming that the ability to generate the original concept object according to normal, non-malicious prompts was preserved.

The presentation also included a demonstration of multi-concept censorship. This showed that different target images could be generated when different censored words were present in the prompt, while still preserving the normal functionality for benign prompts. For example, if "on fire" led to red teapots, another blacklisted concept like "destroyed" might lead to images of green apples, effectively managing multiple potential misuse scenarios. However, the speaker acknowledged that censoring too many concepts could burden the limited expressive power of the embedding, leading to a proposed solution of using more embedding vectors to enhance performance in such complex scenarios.

Defensive Implications

▶ Watch: Key evaluation results and performance metrics (8:40)

THEMIS introduces a significant defensive capability for the generative AI ecosystem, particularly for creators and platforms dealing with personalized models. Its implications are multi-faceted:

  1. Empowering Responsible Creators: THEMIS provides Textual Inversion creators with a direct and effective mechanism to embed ethical guidelines and content policies directly into their personalized embeddings. This allows them to proactively prevent the misuse of their creations for generating harmful, illegal, or offensive content, even when integrated into open-source models with minimal inherent restrictions. This shifts some of the responsibility for content moderation from downstream platforms to the source of the personalized concept.
  1. Mitigating Malicious Content Generation: By enabling creators to blacklist specific concepts (e.g., violence, hate speech, misinformation) and redirect malicious prompts to innocuous outputs, THEMIS serves as a powerful deterrent against the intentional generation of harmful images. This can help prevent the spread of visual propaganda, deepfakes, or other forms of digital harm that leverage personalized AI.
  1. Enhancing Platform Trust and Safety: Platforms that host or distribute Textual Inversion embeddings (e.g., model repositories like Hugging Face, Civitai) could potentially integrate or recommend THEMIS as a best practice. By encouraging or even requiring creators to utilize such censorship mechanisms, these platforms can significantly improve their overall trust and safety posture, reducing the amount of malicious content that can be generated through their hosted assets.
  1. Addressing the "Open-Source Dilemma": Many powerful generative models are open-source and can be run locally without any built-in content filters. Textual Inversions, being small and portable, can easily be used with these unrestricted models. THEMIS offers a way to inject a layer of content moderation at the embedding level, even when the underlying large model is entirely open and uncensored. This is a crucial step towards making open-source AI more responsible.
  1. User Awareness: While THEMIS primarily empowers creators, it also implicitly informs users. Users downloading Textual Inversions that have been censored by THEMIS might encounter unexpected outputs when attempting to generate content with blacklisted concepts. This could serve as an indirect signal that the embedding has been intentionally designed with safety constraints, fostering a greater awareness of ethical considerations in AI usage.

In essence, THEMIS offers a practical and scalable solution to a pressing problem in AI security and ethics. It moves beyond reactive content moderation to proactive prevention, allowing for a more controlled and responsible deployment of highly flexible and powerful personalized generative AI models.

Key Takeaways

  • Textual Inversion's Dual Nature: While Textual Inversion is a powerful technique for personalizing image generation models, its lightweight and easily disseminable nature makes it highly susceptible to malicious exploitation for generating harmful content.
  • THEMIS as a Novel Censorship Framework: THEMIS introduces a unique approach to concept censorship by directly injecting "backdoors" into Textual Inversion embeddings during their training process.
  • Preservation of Benign Functionality: The THEMIS framework is carefully designed to ensure that the injected censorship does not impair the Textual Inversion's ability to generate high-quality, intended content under normal, non-malicious prompts.
  • Effective Malicious Content Disruption: When a blacklisted concept is detected in a prompt, THEMIS successfully redirects the generated output to an innocuous, irrelevant target image, thereby preventing the creation of harmful content.
  • Scalable Multi-Concept Control: THEMIS supports censoring multiple distinct concept groups simultaneously, with an architectural enhancement (using more embedding vectors) proposed to improve performance when handling a larger number of censorship targets.
  • Empowering Responsible AI Development: This research provides Textual Inversion creators with a critical tool to proactively embed content moderation and ethical guidelines into their personalized models, fostering responsible innovation in the generative AI landscape.

About the Speaker(s)

The presentation for THEMIS, a paper titled "THEMIS: Regulating Textual Inversion for Personalized Concept Censorship," was delivered by Shiao from Jojan University. Shiao explicitly stated that they were giving the speech on behalf of the paper's authors, who were unable to attend the conference in person due to visa issues. The metadata indicates Yutong Wu as the speaker, suggesting Yutong Wu is one of the primary authors of the paper. While specific titles or affiliations beyond Jojan University for Shiao were not detailed in the transcript, the talk represents significant research contributions in the field of AI security and responsible AI development.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

THEMIS is a competent, clearly scoped piece of ML security research that addresses a real problem — malicious exploitation of Textual Inversion embeddings — with a technically coherent backdoor-injection mechanism. The work is publishable and defensible, but it sits in a crowded neighborhood of backdoor-as-defense literature and doesn't land a punch hard enough to be memorable at a top security venue.

Heather Calloway (CISO) — WEAK

THEMIS is technically interesting work on embedding-level content control for generative AI personalization, but the talk stays almost entirely inside the research frame. It never reaches the institutional, governance, or operational questions that would make it relevant to security leaders or policy decision-makers.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025