Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning

Hongsheng Hu, Shuo Wang, Tian Dong, Minhui Xue

IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 4

Overview

The proliferation of machine learning (ML) models in virtually every sector of society has brought forth a critical challenge: the "right to be forgotten" and the need for data deletion. Machine unlearning (MU) has emerged as a promising paradigm to address this, aiming to remove the influence of specific training data points from a trained model without retraining it from scratch. However, this talk, "Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning," presented by Hongsheng Hu and his co-authors at IEEE S&P, unveils a significant and previously underexplored privacy vulnerability within the machine unlearning process itself.

Watch on YouTube

Visual summary for Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning by Hongsheng Hu, Shuo Wang, Tian Dong, Minhui Xue
Visual summary for Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning by Hongsheng Hu, Shuo Wang, Tian Dong, Minhui Xue

Key moments

  1. 0:00 Introduction to machine unlearning and its motivation
  2. 2:00 New privacy attack surface in machine unlearning
  3. 3:20 Identifying research gap and proposing unlearning inversion attacks
  4. 4:00 Explaining feature inversion in white-box setting
  5. 6:15 Feature inversion results: approximate vs. exact unlearning
  6. 8:00 Introducing label inversion in black-box setting
  7. 8:30 Step-by-step methodology for label inversion attack

Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning

Speakers: Hongsheng Hu, Shuo Wang, Tian Dong, Minhui Xue (Researchers from a University Setting)

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=Vhgdlp0Vtno

Overview

The proliferation of machine learning (ML) models in virtually every sector of society has brought forth a critical challenge: the "right to be forgotten" and the need for data deletion. Machine unlearning (MU) has emerged as a promising paradigm to address this, aiming to remove the influence of specific training data points from a trained model without retraining it from scratch. However, this talk, "Learn What You Want to Unlearn: Unlearning Inversion Attacks against Machine Unlearning," presented by Hongsheng Hu and his co-authors at IEEE S&P, unveils a significant and previously underexplored privacy vulnerability within the machine unlearning process itself.

The core contribution of this research is the identification and demonstration of Unlearning Inversion Attacks. These attacks exploit the differences between an original model and its unlearned counterpart to infer confidential information about the very data points that were intended to be deleted. The talk meticulously details two primary attack vectors: feature inversion (under white-box access) and label inversion (under black-box access), revealing how sensitive information can be reconstructed even after data has been "unlearned." This work is critical for understanding the true privacy guarantees (or lack thereof) offered by current machine unlearning techniques and for guiding the development of more robust privacy-preserving ML systems.

Background

▶ Watch: Introduction to machine unlearning and its motivation (0:00)

Machine learning fundamentally involves training a model using a training dataset and a learning algorithm. The resulting model then makes predictions on new, unseen data. In recent years, the growing awareness of data privacy, driven by regulations such as GDPR, has highlighted the necessity for individuals to request the deletion of their personal data from systems, including those powering ML models. Simply deleting data from the training set is insufficient; its influence must also be expunged from the trained model. This is where machine unlearning (MU) comes into play.

Machine unlearning takes an existing, trained model and, given a specification of data to be "unlearned" (deleted), applies an unlearning algorithm to produce an unlearned model. This unlearned model is ideally indistinguishable from a model that was originally trained without the deleted data. The motivation for MU is clear: sensitive user information, if part of the training data, should be removable upon request.

Prior research in ML privacy has extensively explored membership inference attacks (MIA). These attacks, typically performed against a single ML model, aim to infer whether a specific data sample was part of the model's training dataset or not – essentially a yes/no question about membership status. However, the researchers identified a critical gap in this existing body of work. While MIAs focus on membership, they do not directly infer the confidential information itself contained within the data points. The introduction of machine unlearning, which explicitly produces two versions of a model (the original and the unlearned), creates a novel attack surface. An adversary with access to both models could potentially glean far more granular information than just membership status. This realization spurred the investigation into whether the actual confidential information of the unlearned data could be inferred, leading to the development of Unlearning Inversion Attacks.

Key Findings

▶ Watch: Identifying research gap and proposing unlearning inversion attacks (3:20)

The central discovery of this research is that machine unlearning, despite its privacy-enhancing intent, can inadvertently create a new avenue for privacy leakage through Unlearning Inversion Attacks. These attacks exploit the differential information between the original and unlearned models to reconstruct attributes of the data that was supposed to be forgotten.

The researchers identified two primary categories of Unlearning Inversion Attacks:

  1. Feature Inversion: Under a white-box access scenario, where an attacker has full knowledge of the parameters of both the original and unlearned models, it is possible to infer the feature information of the unlearned data. This means reconstructing the actual input data points (or their representations) that were requested to be deleted. The study found that approximate unlearning methods, which directly modify model parameters to achieve unlearning, are significantly more vulnerable to feature inversion than exact unlearning methods, which retrain the model from scratch. While exact unlearning was shown to be safer due to the introduction of randomness during retraining, even in multiple unlearning scenarios, some features could still be inverted from approximate unlearning.
  1. Label Inversion: Under a more realistic black-box access scenario, where an attacker can only query the original and unlearned models and observe their outputs (e.g., prediction confidences), it is possible to infer the label associated with the unlearned data. This attack proved effective against both exact and approximate unlearning methods, highlighting a fundamental vulnerability regardless of the unlearning approach. The attacker can accurately pinpoint which class's samples were removed from the training data.

A crucial finding related to defenses was that several intuitive post-processing techniques proposed to mitigate these attacks (such as obfuscation, model pruning, or fine-tuning the unlearned model) all led to unacceptable privacy and utility tradeoffs. This indicates that current defensive strategies are not robust enough to simultaneously protect against unlearning inversion while maintaining model performance.

Technical Deep Dive

▶ Watch: Explaining feature inversion in white-box setting (4:00)

The technical core of this research lies in the meticulous design and execution of the two proposed Unlearning Inversion Attacks: Feature Inversion and Label Inversion. Both leverage the unique opportunity presented by the existence of two model versions: the original model ($M_{orig}$) and the unlearned model ($M_{unlearn}$).

Feature Inversion (White-Box Access)

The Feature Inversion attack assumes a powerful adversary with white-box access to both $M_{orig}$ and $M_{unlearn}$. This means the attacker can inspect and manipulate the internal parameters (weights and biases) of both models. The goal is to reconstruct the features of the unlearned data.

The attack proceeds in two main steps:

  1. Calculate Model Differences: The adversary computes the direct difference between the parameters of the original model and the unlearned model. Mathematically, if $W_{orig}$ are the parameters of $M_{orig}$ and $W_{unlearn}$ are the parameters of $M_{unlearn}$, the attacker computes $\Delta W = W_{orig} - W_{unlearn}$. The high-level intuition here is that this difference, $\Delta W$, is closely related to the gradient information of the unlearned data. Since the original and unlearned models only differ due to the removal of the unlearned data, their parameter differences should reflect the contribution of that data to the original model's training. The dimension of this gradient information is precisely the dimension of the model parameters themselves.
  1. Leverage Gradient to Invert Features: With the derived gradient information ($\Delta W$), the attacker then employs optimization algorithms to reconstruct the features of the unlearned data. This technique draws inspiration from well-studied methods in federated learning feature leakage, where gradients are often used to infer sensitive input data. The process typically starts with a randomly initialized "noise" vector. Through iterative optimization, the attacker minimizes a loss function that quantifies the difference between the actual model parameter difference ($\Delta W$) and the parameter difference that would be induced by the current reconstructed feature. Gradually, this iterative process refines the initial noise, converging towards a representation that closely resembles the original features of the unlearned data. The speaker visually demonstrated how an initially noisy inverted feature progressively becomes recognizable through this optimization.

The empirical results for feature inversion revealed significant differences between unlearning methods. In single unlearning cases (where only one sample is removed):

  • Approximate unlearning methods, which directly modify model parameters, showed "quite good" feature inversion results, meaning the features of the unlearned data were largely reconstructible.
  • Exact unlearning methods, which retrain the model from scratch (excluding the unlearned data), yielded features that were "not visible." This suggests exact unlearning is "safer due to randomness in the training," as the complete retraining process introduces sufficient variance to obscure the specific influence of any single removed data point.

However, in multiple unlearning cases (where several samples are removed), even for approximate unlearning, feature inversion became "very difficult." While not all features could be inverted, the researchers noted that "some features can still be inverted," indicating that even in these more complex scenarios, approximate unlearning remains vulnerable to some degree.

Label Inversion (Black-Box Access)

The Label Inversion attack operates under a more constrained and realistic black-box access scenario. Here, the attacker does not have access to the internal parameters of $M_{orig}$ or $M_{unlearn}$; they can only submit inputs and observe the models' prediction confidences. The objective is to infer the label of the unlearned data.

The attack unfolds in three distinct steps:

  1. Construct Probing Samples on Original Model: The attacker first generates a set of probing samples. These samples are carefully crafted such that when queried against the original model ($M_{orig}$), their prediction confidences are equally distributed across all possible classes. For instance, if there are three classes, a probing sample would ideally yield a prediction confidence of approximately 1/3 for each class. This ensures that these samples are not strongly biased towards any particular class in the original model.
  1. Probe the Unlearned Model with Probing Samples: The adversary then takes these same probing samples and queries the unlearned model ($M_{unlearn}$). For each probing sample, the attacker records the prediction confidences across all classes from $M_{unlearn}$.
  1. Identify Label by Locating Largest Confidence Drop: The core intuition of this step is that if a class's samples were removed during unlearning, the unlearned model should exhibit a noticeable decrease in confidence for that specific class when presented with the probing samples. The attacker compares the prediction confidences from $M_{orig}$ (which were equally distributed) with those from $M_{unlearn}$. The class that shows the largest drop in confidence (relative to its original, equally distributed confidence) is identified as the label of the unlearned data. For example, if Class A experiences the most significant confidence drop, it suggests that the unlearned data belonged to Class A, as its removal has weakened the model's ability to confidently predict that class.

The results for label inversion were striking:

  • In single class unlearning cases (where samples from only one class, e.g., Class 0, were removed), the attack successfully identified the correct label. The visual results presented clearly showed Class 0 having the largest confidence drop.
  • In multiple classes unlearning cases (where samples from several classes, e.g., Class 0, 1, 2, 3, 4, were removed), the attack still performed remarkably well. The researchers noted that "Class 0, 1, 2, 3 the confidence job is among top five," and "we successfully identify four classes" out of the five unlearned classes. This demonstrates the attack's efficacy even when multiple labels are involved.

Crucially, the label inversion attack was effective against both exact unlearning and approximate unlearning methods, underscoring a fundamental privacy vulnerability inherent in the act of unlearning itself, regardless of the underlying technique.

Demo / Proof of Concept

▶ Watch: Introducing label inversion in black-box setting (8:00)

While the talk did not feature a live, interactive software demonstration in the traditional sense, the researchers presented compelling empirical results and visualizations that served as robust proof of concept for their Unlearning Inversion Attacks. These results were derived from experiments conducted on various unlearning scenarios and provided clear evidence of the attacks' efficacy.

For the Feature Inversion attack, the presenter showed visual examples of how an initially noisy inverted feature gradually transformed into a recognizable image through the iterative optimization process. This demonstrated the successful reconstruction of the actual content of the unlearned data. The comparison between approximate and exact unlearning was also presented visually, illustrating the stark difference in reconstruction quality – clear features for approximate unlearning versus obscured features for exact unlearning. The challenges of multiple unlearning cases were also acknowledged, with indications of partial feature recovery.

For the Label Inversion attack, the proof of concept relied on graphs depicting prediction confidence drops. For single class unlearning, a bar chart clearly showed the largest confidence drop occurring for the specific class that had been unlearned (e.g., Class 0). In the more complex multiple classes unlearning scenario, similar visualizations highlighted the top confidence drops corresponding to the majority of the unlearned classes. These visual proofs convincingly illustrated how the black-box attack could accurately infer the labels of the forgotten data points by observing changes in model confidence. The consistency of these results across both exact and approximate unlearning further solidified the empirical validation of the proposed attacks.

Defensive Implications

▶ Watch: Step-by-step methodology for label inversion attack (8:30)

The research also delves into potential defensive strategies against Unlearning Inversion Attacks, acknowledging the critical need to mitigate these newly discovered privacy risks. The speaker discussed three intuitive post-processing approaches aimed at making the difference between the original and unlearned models less accurate, thereby hindering the inversion process:

  1. Obfuscation: This defense involves adding noise or perturbing the parameters of the unlearned model. The intuition is that by introducing randomness, the precise gradient information that feature inversion relies upon would become obscured, making reconstruction difficult.
  1. Model Pruning: This technique involves selectively removing (pruning) certain connections or neurons from the unlearned model. The idea is that a sparser or less complex unlearned model might reveal less information about the deleted data.
  1. Fine-tuning: This approach suggests further fine-tuning the unlearned model on an additional, potentially public or sanitized, dataset. The goal is to "overwrite" or dilute any lingering influence of the unlearned data by exposing the model to new, unrelated information.

The underlying principle for all three proposed defenses is to "post-process the unlearn model to make the difference between unlearn model and the origion model inaccurate." If this difference is sufficiently inaccurate, the unlearning inversion results would also be inaccurate, thus protecting privacy.

However, the key takeaway from the investigation into these defenses was disheartening: "we found that the S possible differences they all lead to unacceptable privacy and utility tradeoffs." This means that while these methods might offer some degree of privacy protection against unlearning inversion, they invariably degrade the utility or performance of the unlearned model to an extent that makes them impractical for real-world deployment. For example, excessive obfuscation might make the model less accurate for its intended tasks, or pruning might lead to a significant drop in predictive power. This critical finding highlights a significant challenge in the field of machine unlearning: achieving robust privacy against inversion attacks without sacrificing the model's primary function. The researchers concluded that "in the future work we can propose some more advanced differences to mitigate the UN learning inversion attacks."

Key Takeaways

  • Machine unlearning introduces a new attack surface for privacy leakage: The existence of both an original and an unlearned model allows adversaries to infer confidential information about deleted data.
  • Unlearning Inversion Attacks can reconstruct features and labels: Depending on attacker capabilities (white-box vs. black-box), sensitive feature information or data labels can be recovered from unlearned models.
  • Approximate unlearning is more vulnerable to feature inversion: Methods that directly modify model parameters are more susceptible to feature reconstruction than exact unlearning (retraining from scratch), which benefits from training randomness.
  • Label inversion is effective against both unlearning types: Regardless of whether exact or approximate unlearning is used, the label of unlearned data can be inferred through black-box queries and confidence analysis.
  • Current defenses against unlearning inversion are insufficient: Proposed post-processing techniques (obfuscation, pruning, fine-tuning) lead to unacceptable privacy-utility tradeoffs, indicating a need for more sophisticated solutions.
  • Further research is crucial for robust unlearning: The findings underscore the urgency for developing advanced, privacy-preserving unlearning algorithms and defenses that can withstand inversion attacks without compromising model utility.

About the Speaker(s)

The primary presenter of this work was Hongsheng Hu, who introduced himself as a researcher sharing his team's latest work at the IEEE S&P conference. He, along with his co-authors Shuo Wang, Tian Dong, and Minhui Xue, are affiliated with "the university unity" (likely a transcription of "a university" or "our university"), indicating their academic background and contribution to cutting-edge research in machine learning privacy and security. Their work represents a significant academic contribution to understanding and addressing the complex privacy challenges inherent in machine unlearning.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research unveils a critical and novel privacy vulnerability in machine unlearning, demonstrating how an adversary can reconstruct sensitive features or labels of unlearned data by exploiting the difference between original and unlearned models. The work is a wake-up call, exposing that current MU methods, including proposed defenses, offer insufficient privacy guarantees and necessitates a fundamental re-evaluation of unlearning mechanisms.

Heather Calloway (CISO) — STRONG ACCEPT

This research critically exposes a fundamental gap in machine unlearning, demonstrating how 'deleted' data can still be inferred through inversion attacks. It creates significant governance and compliance risk for organizations relying on these technologies for privacy mandates, highlighting the urgent need for robust, operational defenses.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024