Defending Against Membership Inference Attacks on Iteratively Pruned Deep Neural Networks

Jing Shang

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Membership Inference

Overview

In an era defined by the escalating scale of deep neural networks (DNNs) and the concurrent demand for their deployment on resource-constrained devices, model compression techniques have become indispensable. This talk, presented by Jing Shang from Beijing University of Technology, delves into the critical security implications of one such technique: neural network pruning. Specifically, the research focuses on iterative pruning, a method known for achieving superior trade-offs between model utility and sparsity, and its unexpected vulnerability to Membership Inference Attacks (MIA).

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction to iterative pruning and MIA vulnerability
  2. 4:40 Identifying factors increasing memorization in pruned models
  3. 5:00 Three scenarios for defense against memorization
  4. 5:45 Overview of the WeMine defense framework
  5. 7:20 Experimental setup and evaluation metrics
  6. 8:50 WeMine's performance comparison with existing defenses
  7. 10:00 Conclusion: key factors and proposed defense
  8. 10:50 Q&A: How memorization scores are calculated

Defending Against Membership Inference Attacks on Iteratively Pruned Deep Neural Networks

Speakers: Jing Shang

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=OCiOcaPlohg

Overview

In an era defined by the escalating scale of deep neural networks (DNNs) and the concurrent demand for their deployment on resource-constrained devices, model compression techniques have become indispensable. This talk, presented by Jing Shang from Beijing University of Technology, delves into the critical security implications of one such technique: neural network pruning. Specifically, the research focuses on iterative pruning, a method known for achieving superior trade-offs between model utility and sparsity, and its unexpected vulnerability to Membership Inference Attacks (MIA).

The core of this work investigates how iterative pruning, while efficient, significantly amplifies the memorization of training data, thereby making models more susceptible to privacy breaches through MIA. Shang and their team identify two primary factors contributing to this heightened vulnerability: the persistent reuse of training data across pruning epochs and the inherent "easy-to-memorize" characteristics of certain data samples. To counter these threats, the researchers propose "WeMine," a novel defense framework designed to weaken memorization during the iterative pruning process, offering a robust solution to protect user privacy without unduly compromising model performance.

This article provides a comprehensive technical breakdown of the challenges posed by iterative pruning, the mechanisms by which MIA exploits these vulnerabilities, and the innovative defensive strategies put forth by Shang's team. It highlights the critical balance between model efficiency and data privacy, offering crucial insights for researchers and practitioners working on deploying machine learning models responsibly in sensitive environments.

Background

▶ Watch: Introduction to iterative pruning and MIA vulnerability (0:00)

The rapid advancement of deep neural networks has led to the development of increasingly complex and large-scale models. While these models often achieve state-of-the-art performance, their substantial computational and memory footprints pose significant challenges for deployment on devices with limited resources, such as mobile phones or IoT devices. To address this, model compression techniques have become a vital area of research, with neural network pruning emerging as a widely adopted method.

Traditional neural network pruning typically follows a one-shot approach, comprising three main stages:

  1. An original model is fully trained on the entire dataset.
  2. A sparse model is then obtained by identifying and removing redundant parameters (pruning).
  3. Finally, the pruned model undergoes fine-tuning in its sparse configuration to recover any lost utility.

Recent studies, however, have demonstrated that iterative pruning can achieve a better trade-off between model utility (performance) and sparsity (compression) compared to its one-shot counterpart. Iterative pruning involves repeating the pruning and fine-tuning steps multiple times, progressively reducing the model's size while maintaining accuracy.

Despite its benefits, model compression, particularly fine-tuning, has been linked to increased model memorization. This phenomenon, where a model inadvertently retains specific details about its training data, creates a critical privacy vulnerability exploited by Membership Inference Attacks (MIA). In an MIA, an attacker, often with only black-box access to a target model, attempts to determine whether a specific data sample was part of the model's training dataset. The attacker typically queries the model with various data samples and analyzes the model's prediction confidence or other output characteristics. By comparing these outputs with a predefined threshold, the attacker can infer membership. The underlying issue enabling such attacks is the tendency of DNNs to memorize training data, especially during fine-tuning operations. Previous research has shown that the fine-tuning stage in one-shot pruning can increase a model's memorization, making the pruned model more vulnerable to MIA than the original, unpruned model. This sets the stage for the central research question addressed in this talk: Do iteratively pruned models become even more vulnerable to MIA given their repeated fine-tuning cycles?

Key Findings

▶ Watch: Three scenarios for defense against memorization (5:00)

The research presented by Jing Shang and their team provides compelling evidence that iteratively pruned models indeed exhibit a heightened vulnerability to Membership Inference Attacks (MIA) compared to their one-shot pruned counterparts. This increased susceptibility is attributed to two primary, interconnected factors:

  1. Increased Memorization Due to Data Reuse: The iterative nature of the pruning process involves multiple rounds of fine-tuning. Crucially, the same entire training dataset is often reused in each fine-tuning epoch. The study found a direct correlation between this reuse of training data and an increase in the model's memorization capacity. Each successive fine-tuning step, while refining model utility, inadvertently reinforces the model's memory of the training samples, thereby escalating the privacy risk. The results show that reducing the training data applied during fine-tuning can mitigate this memorization.
  1. Inherited Easy-to-Memorize Characteristics of Data Samples: Beyond the impact of data reuse, the researchers discovered that certain training data samples are intrinsically "easy to memorize." These samples contribute disproportionately to the model's memorization, making them more vulnerable to privacy breaches. The study demonstrated that attacks are significantly more accurate in identifying members with high memorization scores, indicating that the model retains a stronger, more persistent memory of these specific samples throughout the fine-tuning process. This inherent characteristic of some data samples, coupled with iterative training, creates a critical privacy blind spot.

In summary, the key findings are:

  • Iteratively pruned models are more vulnerable to MIA than one-shot pruned models. The MIA accuracy of iteratively pruned models was consistently higher.
  • The reuse of training samples across multiple fine-tuning epochs is a significant factor contributing to increased memorization and, consequently, higher privacy risk.
  • The presence of inherently easy-to-memorize training data samples further exacerbates the problem, as these samples become prime targets for membership inference.
  • Attacks are more accurate when targeting data identified as having high memorization scores, directly linking memorization to attack success.

These findings highlight a critical privacy-utility trade-off in the context of advanced model compression techniques and underscore the necessity for specialized defenses to safeguard user data when employing iterative pruning.

Technical Deep Dive

▶ Watch: Experimental setup and evaluation metrics (7:20)

To address the identified vulnerabilities in iteratively pruned deep neural networks, the researchers proposed WeMine, a comprehensive defense framework designed to weaken memorization. The framework operates in two main stages:

Stage 1: Original Model Training and Memorization Score Generation

The first stage involves training an initial, unpruned model. Following this, the framework proceeds to generate memorization scores for individual data samples within the training dataset. While the talk mentions that "some work has given score how to compute," the speaker clarifies their specific approach: "we use the all interaite and we slided the data set in different block. So we do some opposition to fast this calculation." This implies an efficient, block-based processing of the dataset to derive these scores, potentially leveraging techniques from prior research to quantify how deeply a model has memorized each sample (e.g., based on loss values, prediction confidence, or changes in model parameters related to specific samples). The memorization score serves as a crucial indicator of a sample's privacy risk.

Stage 2: Memorization-Weakened Pruning (MWP)

The second stage, central to WeMine, is the memorization-weakened pruning (MWP) process. This stage iteratively prunes and fine-tunes the model, but with integrated mechanisms specifically designed to mitigate memorization. The framework introduces three distinct memorization-weakening primitives:

  1. Memorization Score-Based Data Ranking: This primitive addresses the issue of easy-to-memorize samples. Within each data class, samples are ranked based on their previously computed memorization scores. This ranking allows the defense mechanism to prioritize or deprioritize certain samples during subsequent training steps, effectively managing their contribution to the model's memory.
  1. Sliding Windows-Based Data Sampling: This primitive primarily tackles the problem of data reuse. Instead of repeatedly exposing the entire training dataset to the model during each fine-tuning epoch, this method employs a sliding window approach. This means that only a subset of the data within a specified window is used for fine-tuning in a given iteration. As the process continues, the window "slides," introducing new subsets of data while gradually reducing the exposure frequency of any single sample. This limits the cumulative memorization effect caused by constant data reuse. The defense setting uses three window sizes and two slide step sizes for each dataset.
  1. Additive Regularization: This primitive applies an L2 regularization with varying intensities. The key innovation here is that the regularization intensity (lambda I) is not uniform but is dynamically adjusted based on the data privacy risk of individual samples. Samples identified as having higher memorization scores (and thus higher privacy risk) would receive stronger regularization. This forces the model to learn more generalized features for sensitive data points, preventing it from overfitting and memorizing their specific characteristics. The evaluation found that a "middle coefficient value" for lambda I often yielded the best privacy-utility trade-off.

By combining these primitives, WeMine proposes three distinct memorization-weakened methods:

  • ISW (Ranking-based Sliding Windows Defense): This method primarily combines the memorization score-based data ranking and the sliding windows-based data sampling primitives. It aims to reduce memorization by controlling data exposure and prioritizing less memorized data during fine-tuning.
  • IMI (Risky Memory Regularization Defense): This method focuses on the additive regularization primitive, applying targeted L2 regularization based on data privacy risk to directly curb memorization.
  • SWMI (Combined Defense): This comprehensive method integrates all three primitives, combining the strengths of ISW and IMI to address both data reuse and easy-to-memorize characteristics simultaneously.

Evaluation and Performance:

The effectiveness of WeMine was evaluated across six diverse datasets and four different neural network models. The evaluation focused on both model prediction accuracy (utility) and defense effectiveness against MIA (privacy).

  • Prediction Accuracy: For ISW and SWMI, a decrease in window size (meaning less training data available for fine-tuning) generally led to a decline in model prediction accuracy. This highlights the inherent trade-off between privacy and utility.
  • Privacy-Utility Trade-off: For IMI, the research found that setting the regularization coefficient lambda I to a "middle" value often achieved the most favorable balance between maintaining model utility and enhancing privacy.
  • Defense Effectiveness: In general, smaller window sizes combined with smaller step sizes in ISW and SWMI resulted in better defense effects. Importantly, under the same window settings, SWMI consistently demonstrated superior defense effects compared to ISW, indicating the benefit of combining multiple primitives.
  • Comparison with Existing Defenses: WeMine was benchmarked against several existing defense mechanisms. The results showed that WeMine offers a better privacy-utility trade-off, providing more robust privacy protection for a given level of model accuracy.
  • Time Cost: An additional benefit highlighted was the time efficiency of ISW. Due to its sliding window sampling, which reduces the amount of training data processed per epoch, ISW performed better in terms of time cost, speeding up the fine-tuning process.

The technical design of WeMine, with its modular primitives and combined methods, provides a flexible and effective framework for mitigating membership inference risks in the increasingly popular domain of iteratively pruned deep neural networks.

Demo / Proof of Concept

▶ Watch: WeMine's performance comparison with existing defenses (8:50)

The talk did not include a specific live demonstration or a detailed description of a proof-of-concept implementation beyond the experimental evaluation results. The presentation focused on the theoretical framework, the identified vulnerabilities, the proposed defense mechanisms, and their empirical performance across various datasets and models. The evaluation section effectively serves as a validation of the proposed methods, demonstrating their efficacy through quantitative results rather than a direct, interactive demo.

Defensive Implications

▶ Watch: Q&A: How memorization scores are calculated (10:50)

The findings and the WeMine defense framework presented in this talk carry significant implications for practitioners and researchers involved in deploying deep neural networks, especially those utilizing model compression techniques. The increased vulnerability of iteratively pruned models to Membership Inference Attacks (MIA) necessitates a proactive and integrated approach to privacy.

Here are the key defensive implications:

  1. Recognize the Elevated Risk of Iterative Pruning: Developers and MLOps teams must understand that while iterative pruning offers efficiency benefits, it inherently increases the risk of data memorization and, consequently, susceptibility to MIA. This risk is higher than with traditional one-shot pruning. Any deployment involving iteratively pruned models, particularly in sensitive domains (e.g., healthcare, finance, personal data), should be assessed for MIA vulnerabilities.
  1. Integrate Privacy-Preserving Mechanisms from the Outset: Instead of applying privacy defenses as an afterthought, frameworks like WeMine demonstrate the importance of building memorization-weakening mechanisms directly into the pruning and fine-tuning pipeline. This ensures that privacy is considered throughout the model lifecycle.
  1. Leverage Memorization Scores for Risk Assessment: The concept of memorization scores is a powerful tool for identifying high-risk data samples. Organizations should explore methods to compute and utilize these scores to understand which parts of their training data are most vulnerable to inference and to guide targeted privacy interventions.
  1. Adopt Data Sampling and Regularization Strategies:
  • Sliding Windows-Based Data Sampling (ISW): For scenarios where reducing data reuse is critical and computational efficiency is a concern, adopting a sliding windows approach during fine-tuning can effectively limit memorization. The trade-off between window size/step size and model accuracy needs careful tuning.
  • Additive Regularization (IMI): Implementing adaptive L2 regularization based on data privacy risk (memorization scores) offers a direct way to encourage models to generalize rather than memorize. Experimenting with regularization coefficients (lambda I) is crucial to find the optimal privacy-utility balance for specific applications.
  1. Consider Combined Defense Strategies (SWMI): For maximum protection against both data reuse and easy-to-memorize data characteristics, the combined defense method (SWMI) is recommended. While it might involve more complexity, its superior defense effectiveness makes it suitable for high-stakes applications where privacy is paramount.
  1. Evaluate Privacy-Utility Trade-offs Carefully: Defenders must conduct thorough evaluations to understand the trade-offs associated with each defense method. Reducing memorization might come at the cost of a slight decrease in model accuracy. The acceptable level of this trade-off will depend on the specific application's requirements and the criticality of the data being protected.
  1. Benchmark Against Existing Defenses: As demonstrated by the research, WeMine offers a better privacy-utility trade-off compared to several existing defenses. Practitioners should benchmark new or adapted defense strategies against state-of-the-art methods to ensure optimal protection.

By adopting these defensive implications, organizations can better secure their iteratively pruned deep neural networks against membership inference attacks, fostering trust and compliance in an increasingly privacy-conscious technological landscape.

Key Takeaways

  • Iterative pruning significantly increases vulnerability to Membership Inference Attacks (MIA) compared to one-shot pruning, due to amplified data memorization.
  • Two primary factors contribute to this heightened risk: persistent reuse of training data across fine-tuning epochs and the inherent "easy-to-memorize" characteristics of certain data samples.
  • The WeMine framework is proposed as a comprehensive defense, operating in two stages: memorization score generation and memorization-weakened pruning.
  • WeMine introduces three key primitives to mitigate memorization: memorization score-based data ranking, sliding windows-based data sampling, and additive regularization based on privacy risk.
  • Three defense methods (ISW, IMI, SWMI) are derived from these primitives, offering varying levels of protection and utility trade-offs, with SWMI generally providing the best defense.
  • WeMine generally offers a better privacy-utility trade-off compared to existing defenses, and methods like ISW can also improve time efficiency during fine-tuning.

About the Speaker(s)

The primary speaker for this presentation is Jing Shang, representing Beijing University of Technology. Jing Shang presented the research and answered questions regarding the methodology, particularly around the computation of memorization scores. The work was conducted in collaboration with Kang, identified as a student from Beijing University of Technology, and Mod Amad Zuma and Zuming Jao from Northeastern University, all of whom contributed to this research on defending against membership inference attacks on iteratively pruned deep neural networks.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Legitimate academic security research on a real and underexplored problem — MIA amplification in iteratively pruned DNNs — with a concrete defense framework (WeMine) that beats existing baselines on privacy-utility tradeoff. Solid contribution to ML privacy literature, but it's a conference paper presentation, not a practitioner talk, and the technical delivery is thin enough that it reads better as a PDF than a session.

Heather Calloway (CISO) — PASS

Narrow ML privacy research on a specific compression technique, presented at the academic level for an academic audience. No governance angle, no institutional relevance, and no path to operator action for anyone running a security program.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025