Black-box Membership Inference Attacks against Fine-tuned Diffusion Models

Yan Pang

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Membership Inference

Overview

This talk, presented by Yan Pang at the NDSS Symposium, delves into the critical area of data privacy concerning the rapidly evolving landscape of generative AI, specifically diffusion models. The core subject is the development and application of black-box membership inference attacks (MIAs) against fine-tuned conditional diffusion models. As diffusion models like Stable Diffusion achieve unprecedented capabilities in generating photorealistic images and even serving as game engines, their reliance on massive datasets – often comprising billions of images – raises significant privacy and copyright concerns. The talk highlights that traditional MIAs, designed for simpler generative adversarial networks (GANs) or variational autoencoders (VAEs), are inadequate for the complex, multi-step generation process of diffusion models.

Watch on YouTube · Slides

Key moments

  1. 0:00 Diffusion models, massive data, and privacy concerns
  2. 1:50 Membership Inference Attack (MIA) types explained
  3. 2:58 Black-box attack: realistic yet challenging scenario
  4. 4:10 Drawbacks of Monte Carlo and Ganix attacks
  5. 6:00 Redefining threat model for conditional diffusion models
  6. 7:15 Novel attack: objective function guided feature selection
  7. 8:15 Attack pipeline: synthesizing replicate image embeddings

Black-box Membership Inference Attacks against Fine-tuned Diffusion Models

Speakers: Yan Pang

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=qjEDs9I3U_0

Overview

This talk, presented by Yan Pang at the NDSS Symposium, delves into the critical area of data privacy concerning the rapidly evolving landscape of generative AI, specifically diffusion models. The core subject is the development and application of black-box membership inference attacks (MIAs) against fine-tuned conditional diffusion models. As diffusion models like Stable Diffusion achieve unprecedented capabilities in generating photorealistic images and even serving as game engines, their reliance on massive datasets – often comprising billions of images – raises significant privacy and copyright concerns. The talk highlights that traditional MIAs, designed for simpler generative adversarial networks (GANs) or variational autoencoders (VAEs), are inadequate for the complex, multi-step generation process of diffusion models.

The research presented addresses this gap by introducing a novel black-box MIA framework. This framework is specifically tailored to infer whether a particular data sample was included in the training or fine-tuning dataset of a diffusion model, even when the attacker has no internal access to the model's parameters or intermediate outputs. The importance of this work lies in its potential to serve as an auditing tool, enabling researchers and users to assess the privacy implications of models deployed in the wild, such as those found on platforms like Civitai. By demonstrating the feasibility and effectiveness of such attacks, the talk underscores the urgent need for developers and deployers of diffusion models to consider robust privacy-preserving measures.

Background

▶ Watch: Diffusion models, massive data, and privacy concerns (0:00)

The advent of diffusion models has revolutionized image generation, producing outputs of stunning photographic quality. These models are not just static image generators; as highlighted by the speaker, they can even function as game engines, generating successive frames based on current content. This remarkable capability stems directly from their gargantuan training datasets. For instance, Stability AI's Stable Diffusion 3 was reportedly pre-trained on an astounding one billion images and fine-tuned with over 30 million images. While enabling incredible feats of creativity, this reliance on vast quantities of data inherently introduces significant data privacy challenges, particularly regarding intellectual property and copyright.

To investigate these privacy concerns, membership inference attacks (MIAs) serve as a primary methodology. An MIA aims to determine whether a specific data point (query data) was part of a machine learning model's training set. The attacker feeds query data to the target model, observes its output, and based on this information, decides if the query data belongs to the member set (training data) or non-member set.

MIAs are typically categorized by the attacker's level of access to the target model:

  • White-box attacks: The attacker has full access to the model's internals, including loss values, gradients, and model parameters. While theoretically the most powerful, these attacks are impractical in real-world scenarios due to their computational cost and the unlikelihood of gaining such deep access. For diffusion models, calculating loss values at each of the potentially 1,000 denoising steps for a single query sample makes white-box attacks prohibitively expensive.
  • Gray-box attacks: The attacker can access intermediate outputs, such as denoising samples at various steps in a diffusion model. While less demanding than white-box, these are still impractical for many real-world applications where models are accessed via APIs, providing only the final output.
  • Black-box attacks: The attacker can only observe the model's final output, often referred to as a replicate image. This is considered the most realistic and challenging attack scenario, and it is the focus of this research.

Prior work on black-box MIAs against image generators, such as the Monte Carlo attack and GANIX, faced several limitations. These methods typically required sampling an extremely large number of replicate images (e.g., 100,000 samples for optimal performance) to calculate distances or densities. This approach, while potentially effective for one-step generators like GANs and VAEs, becomes computationally infeasible for multi-step diffusion models. Generating 100,000 images from a diffusion model would incur an astronomical computational cost.

Furthermore, previous MIA research often operated under simplified threat models. Some studies directly used the target model's training set as the member set and eschewed shadow model techniques due to time and resource constraints. However, the speaker emphasizes that using shadow models is crucial for generalizability, as attackers in the real world cannot typically access the target model's exact training data. Additionally, many existing attacks were designed for unconditional diffusion models, whereas the current state-of-the-art, exemplified by Stable Diffusion, are conditional diffusion models that rely on text prompts. The change in model properties necessitates a redesign of attack methodologies. This research aims to overcome these challenges by developing a new black-box MIA specifically for fine-tuned conditional diffusion models.

Key Findings

▶ Watch: Black-box attack: realistic yet challenging scenario (2:58)

The research presented by Yan Pang yields several critical findings that advance the understanding and capabilities of membership inference attacks against modern generative AI:

Firstly, a foundational discovery is that existing membership inference attacks designed for one-step generators like GANs (Generative Adversarial Networks) and VAEs (Variational Autoencoders) cannot be directly or effectively applied to multi-step diffusion models. This inapplicability stems from fundamental differences in their underlying model architectures and their respective generation processes. Diffusion models iteratively refine noise into an image over many steps, a process distinct from the single-pass generation of GANs and VAEs, rendering previous attack features unsuitable.

Secondly, the talk highlights that the selection of attack features is paramount to the success of any membership inference attack, particularly in the challenging black-box setting. This research pioneers a novel approach by being the first to utilize the objective function of the diffusion model itself to guide feature selection. This foundational insight allows for a more principled and effective way to identify discriminative features. The core idea is that the likelihood of a query sample being a member of the training set can be quantified by examining the distance between the query sample and a small set of replicate images generated by the target model from that same query. Intuitively, a smaller distance suggests a higher probability of membership, implying the model has "memorized" or overfit to that specific sample. The underlying theorem supporting this idea is detailed in their paper.

Thirdly, the research successfully designs and implements a new attack pipeline specifically targeting fine-tuned conditional diffusion models. This pipeline addresses the limitations of prior work by being efficient, effective, and applicable to realistic black-box scenarios. The attack considers various practical scenarios, including the presence or absence of text components in the query and different degrees of overlap in ancillary data used for shadow model training.

Finally, through extensive evaluation, the proposed method demonstrably outperforms baseline membership inference attacks under identical experimental conditions. The research also includes an ablation study to validate the contribution of each component of the attack pipeline. A significant practical finding is the utility of this attack as an auditing tool. The case study involving models from Civitai vividly illustrates how the attack can effectively identify whether specific artworks were used in fine-tuning publicly available diffusion model checkpoints, thus exposing potential copyright and privacy violations. The speaker also explicitly acknowledges the crucial role of overfitting and memorization in the success of MIAs, noting that larger fine-tuning datasets tend to lead to less memorization and, consequently, a decline in attack accuracy.

Technical Deep Dive

▶ Watch: Drawbacks of Monte Carlo and Ganix attacks (4:10)

The core innovation of this work lies in its principled approach to feature selection for black-box membership inference, leveraging the intrinsic properties of diffusion models. The speaker explains that the likelihood of a query sample belonging to the member set can be quantified by examining the distance between the query sample and a small set of replicate images generated from that query sample itself. The fundamental intuition is that if a model has "memorized" a specific training sample, it will generate outputs highly similar to that sample when prompted with it. This relationship is formally grounded in the diffusion model's objective function, as detailed in the paper.

The proposed attack pipeline is structured into three distinct steps:

  1. Synthesize Replicate Samples:
  • Given a query sample, which typically consists of an image and optionally a text component, the first step is to interact with the target diffusion model.
  • The query point is fed to the target model multiple times (e.g., 'R' times, though the specific number 'R' isn't explicitly stated beyond "for times" in the transcript, implying a small, fixed number) to generate a set of replicate images.
  • A critical consideration for conditional diffusion models is the text component. If the query point is an image without an associated text prompt, an image captioning model is used to automatically generate a descriptive text prompt based on the image content. This ensures that the query can be properly conditioned for the target model.
  1. Feature Extraction and Similarity Scoring:
  • Once the replicate images are obtained, both the original query image and each of the generated replicate images are passed through an image encoder. This encoder transforms the images into a high-dimensional image embedding vector, capturing their semantic content.
  • The next step involves calculating the similarity score between the embedding of the original query image and the embedding of each replicate image. Common similarity metrics, such as cosine similarity or Euclidean distance, would typically be used here.
  • These individual similarity scores are then aggregated into a similarity score vector.
  • To prepare this vector for the attack model, a statistical function is applied. This function might compute the mean, median, variance, or other aggregated statistics of the similarity scores, resulting in a concise processed score vector. This reduction helps abstract away noise and highlight the most discriminative features.
  1. Attack Model:
  • The final step involves feeding the processed score vector into a pre-trained attack model. This attack model is a binary classifier responsible for making the final decision: classifying the query sample as either a "member" or a "non-member" of the target model's training set.
  • The research explores three different types of attack models:
  • Threshold-based attack model: This is the simplest, where a fixed similarity score threshold determines membership.
  • Distribution-based attack model: This model learns the distribution of similarity scores for members and non-members and makes a decision based on which distribution the query's score is more likely to belong to.
  • Cluster-based attack model: This approach might group similar score vectors and identify clusters corresponding to member or non-member characteristics.
  • Crucially, these attack models are not trained on data from the target model directly. Instead, they are trained using data generated from a shadow model. The shadow model is an auxiliary model that mimics the architecture and training process of the target model but is trained on an ancillary dataset. This use of a shadow model makes the attack more realistic and generalizable, as attackers typically do not have access to the target model's internal training data.

The evaluation setup for this research is comprehensive. The target model used is Stable Diffusion 1.5, a widely adopted conditional diffusion model. The attack's performance is tested across three different datasets and measured using three distinct evaluation metrics. The proposed method is benchmarked against two baseline methods from prior work. The experimental design also accounts for various practical scenarios, including different configurations of the ancillary dataset used for shadow model training (e.g., whether it overlaps with the target model's training set or not) and the presence or absence of text components in the query points. The paper considers four different attack scenarios based on these variations. Results consistently show that the proposed method significantly outperforms the baselines, validating the effectiveness of its novel feature selection and pipeline design.

Demo / Proof of Concept

▶ Watch: Novel attack: objective function guided feature selection (7:15)

The practical utility of the proposed black-box membership inference attack is vividly demonstrated through a compelling case study focused on auditing models hosted on Civitai. Civitai is a popular online platform where users can upload and share fine-tuned checkpoints of diffusion models. These checkpoints are often trained on specific art styles, individual artists' works, or particular themes, meaning they incorporate data that might carry copyright or privacy implications.

For this demonstration, the researchers downloaded seven different checkpoints from Civitai. The speaker specifically highlights an example involving a checkpoint fine-tuned on the paintings of an artist named Chai. The demonstration proceeds as follows:

  1. Query Example: An original painting from Chai is selected to serve as the query example. This painting represents a potential member of the fine-tuning dataset used for the Civitai checkpoint.
  2. Replicate Image Generation (Member Score): The downloaded fine-tuned checkpoint (the target model) is prompted with the Chai painting (or its caption, if the text component was generated). The image generated by this checkpoint is termed a replicate image. The similarity score between this replicate image and the original Chai painting is then computed, serving as the member similarity score. The expectation is that if the original Chai painting was indeed part of the fine-tuning data, the replicate image would be highly similar, resulting in a high member similarity score.
  3. Baseline Generation (Non-member Score): To establish a baseline for comparison, an image is generated from the original, un-fine-tuned Stable Diffusion model (or a model not specifically fine-tuned on Chai's work). This generated image represents a non-member scenario. The similarity score between this non-member replicate and the original Chai painting is computed, yielding the non-member similarity score. The expectation here is a lower similarity, as the original model would not have specifically memorized Chai's style or specific artwork.
  4. Quantification and Distinction: The critical step involves quantifying the difference between the member and non-member similarity scores. The speaker presents a table that clearly illustrates a "clearly distinction" between these two types of scores. This distinct separation indicates that the attack can reliably differentiate between images that were likely part of the fine-tuning dataset and those that were not.

This case study effectively serves as a proof of concept, demonstrating that the proposed black-box MIA can function as an effective auditing tool to evaluate models hosted on platforms like Civitai. It highlights the practical risks associated with fine-tuning models on potentially copyrighted or private data without proper consent or attribution, and provides a method for identifying such instances. The ability to audit models in a black-box setting is particularly valuable for end-users and privacy advocates who lack internal access to model operations.

Defensive Implications

▶ Watch: Attack pipeline: synthesizing replicate image embeddings (8:15)

The findings presented in this talk carry significant implications for developers, deployers, and users of diffusion models, emphasizing the need for proactive defensive strategies to mitigate privacy risks.

  1. Enhanced Data Curation and Filtering: The most direct implication is the necessity for rigorous data curation and filtering during both pre-training and fine-tuning phases. Developers, particularly those fine-tuning models for specific styles or artists (as seen on Civitai), must exercise extreme caution to avoid including copyrighted, sensitive, or personally identifiable information without explicit consent. Implementing robust checks to identify and exclude such data points is paramount.
  2. Exploration of Differential Privacy (DP): For model developers concerned about membership leakage, integrating differential privacy (DP) mechanisms during the training or fine-tuning process is a critical consideration. DP aims to provide strong privacy guarantees by adding noise to the training process, making it statistically difficult to infer if any single data point was included. While DP often introduces a trade-off with model utility or performance, its exploration is vital for privacy-sensitive applications.
  3. Output Perturbation and Anonymization: Although not explicitly discussed as a defense in the talk, techniques such as output perturbation could be considered. This involves introducing a controlled amount of noise or minor alterations to the generated images before release. However, this must be carefully balanced to avoid degrading the high quality that users expect from diffusion models.
  4. Limiting Information Leakage (Black-box by Design): The research confirms that black-box access is the most realistic scenario. Model providers should ensure their APIs strictly adhere to a black-box paradigm, providing only final outputs and no intermediate denoising steps, loss values, or other internal model information. This reduces the attack surface for gray-box and white-box attacks.
  5. Proactive Auditing with MIA Tools: The demonstrated use of the attack as an auditing tool on Civitai highlights a valuable defensive strategy: proactive internal auditing. Model developers and platform hosts should regularly use tools similar to the one proposed to assess their models for potential memorization and membership leakage, especially after fine-tuning with new datasets. This allows for early detection and remediation of privacy vulnerabilities.
  6. Transparency and Disclosure: Increased transparency regarding the training data used is crucial. While not a direct defense against the attack, clear disclosure about the nature and source of training data can help manage user expectations and legal liabilities, particularly concerning copyrighted material.
  7. Regular Model Updates and Retraining: As models are continually updated and retrained, it's important to understand how these processes affect memorization. The speaker's observation that larger fine-tuning sets may lead to less memorization suggests that diverse and sufficiently large datasets might inherently offer some degree of protection against MIAs.
  8. Educating Users and Artists: For platforms like Civitai, educating users (both those uploading checkpoints and artists whose work might be used) about the privacy implications of fine-tuning data is important. This fosters responsible AI development and usage practices within the community.

In conclusion, the efficacy of black-box MIAs against fine-tuned diffusion models necessitates a multi-faceted defensive strategy, combining robust data governance, privacy-enhancing technologies, and proactive auditing to safeguard user privacy and intellectual property in the era of generative AI.

Key Takeaways

  • Traditional membership inference attacks (MIAs) designed for one-step generative models like GANs and VAEs are ineffective against multi-step diffusion models due to fundamental differences in their architectures and generation processes.
  • The research introduces a novel black-box MIA framework that effectively targets fine-tuned conditional diffusion models, such as Stable Diffusion 1.5, requiring only access to the model's final image outputs.
  • A key innovation is the use of the diffusion model's objective function to guide feature selection, quantifying membership likelihood by examining the similarity between a query sample and a small set of its generated replicates.
  • The attack pipeline involves three stages: synthesizing replicate samples (using an image captioning model if text is missing), extracting image embeddings and calculating similarity scores, and feeding processed scores to a shadow-model-trained attack model (threshold-based, distribution-based, or cluster-based).
  • The proposed method significantly outperforms existing baseline MIAs and is effective across various attack scenarios, including those with and without text prompts and different ancillary data overlaps.
  • The attack has been successfully demonstrated as a practical auditing tool, capable of identifying potential data memorization and copyright infringement in fine-tuned models hosted on platforms like Civitai.
  • Model memorization and overfitting are critical factors influencing MIA success; larger fine-tuning datasets tend to reduce memorization and, consequently, lower attack accuracy.

About the Speaker(s)

Yan Pang is the presenter of this paper. The transcript indicates Yan Pang is a researcher focused on AI security and privacy.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Legitimate academic security research with a clean contribution: applying black-box MIA to fine-tuned conditional diffusion models using the model's own objective function to guide feature selection. Solid NDSS-tier paper presentation, but the talk itself is methodical and dry — this is a conference paper read aloud, not a security research talk that will stick with you.

Heather Calloway (CISO) — WEAK

Technically sound research on membership inference attacks against fine-tuned diffusion models, with a credible auditing application demonstrated on Civitai. But the governance and accountability dimensions — the ones that matter for legal exposure, platform liability, and regulatory response — are entirely absent, leaving operators and leaders with no clear path forward.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025