Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness

Cheng-Long Wang

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · ML and AI Security 3: Backdoors, Poisoning, Unlearning

Overview

In the rapidly evolving landscape of artificial intelligence, the ability to train powerful machine learning models has become commonplace. However, an equally critical, yet often overlooked, challenge lies in the inverse process: machine unlearning. This talk, presented by Cheng-Long Wang at USENIX Security, delves into a novel framework for accurately measuring the completeness of machine unlearning at a granular, sample-by-sample level. Machine unlearning aims to update a trained model such that it effectively "forgets" specific data points, behaving precisely as if those data were never used in its initial training, all without the prohibitive cost of retraining the entire model from scratch.

Watch on YouTube · Read the paper · Download the PDF (PDF) · Slides

Paper abstract

Software supply chain attacks pose an increasingly severe threat to the security of downstream software worldwide. A common method to mitigate these risks is Software Composition Analysis (SCA), which helps developers identify vulnerable dependencies. However, studies show that popular SCA approaches often suffer from high false positive rates. As a result, developers spend significant time manually validating these alerts, which delays the detection and remediation of genuinely exploitable upstream vulnerabilities. In this paper, we propose ChainFuzz, an automated approach for validating upstream vulnerabilities in downstream software by generating Proof-of-Concepts (PoCs). To achieve this, ChainFuzz addresses three key challenges. First, intra-layer code and constraints. Downstream software introduces custom code and sanity checks that significantly alter the triggering paths and conditions of upstream vulnerabilities. Second, inter-layer dependencies. Software supply chains often involve cross-layer control-flow and data-flow dependencies between conditional statements across different layers. Third, long supply chains. Transitive dependencies in long chains result in intricate exploitation paths, making it challenging to explore large code spaces and handle deeply nested constraints effectively. We comprehensively evaluate ChainFuzz using our dataset, which comprises 66 unique vulnerability and supply chain combinations. Our results demonstrate its effectiveness and practicality in generating PoCs for both direct and transitive vulnerable dependencies. Additionally, we compare ChainFuzz with representative fuzzing tools: AFLGo, AFL++, and NestFuzz, highlighting its superior performance in downstream PoC generation.

Visual summary for Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness by Cheng-Long Wang
Visual summary for Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness by Cheng-Long Wang

Key moments

  1. 0:00 Introduction to machine unlearning and its significance
  2. 1:40 Categories of unlearning methods and research gap
  3. 2:40 Limitations of traditional membership inference attacks
  4. 4:00 Introducing the Bounded Google Map transformation
  5. 6:00 Interpolation framework for scrub-based unlearning completeness
  6. 7:00 Overview of the Interpolated Approximate Measurement (IAM) framework

Towards Lifecycle Unlearning Commitment Management: Measuring Sample-level Unlearning Completeness

Speakers: Cheng-Long Wang

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=bYQDDMHeftc

Overview

In the rapidly evolving landscape of artificial intelligence, the ability to train powerful machine learning models has become commonplace. However, an equally critical, yet often overlooked, challenge lies in the inverse process: machine unlearning. This talk, presented by Cheng-Long Wang at USENIX Security, delves into a novel framework for accurately measuring the completeness of machine unlearning at a granular, sample-by-sample level. Machine unlearning aims to update a trained model such that it effectively "forgets" specific data points, behaving precisely as if those data were never used in its initial training, all without the prohibitive cost of retraining the entire model from scratch.

The importance of robust machine unlearning spans multiple domains. From a server-side perspective, it empowers organizations to control their data and models by removing harmful, low-quality, copyrighted, or policy-sensitive training data, and even eradicating backdoors or poisoning attacks. For individual users, unlearning is fundamental to upholding the "right to be forgotten," allowing them to request the removal of their personal data from models. Despite its profound practical implications, the field has largely prioritized efficiency and flexibility, often at the expense of rigorous guarantees regarding how much data is truly forgotten. This talk addresses this critical gap by questioning the efficacy of existing unlearning claims and proposing a robust method to audit them.

The core contribution of this work is the Interpolated Approximate Measurement (IM) framework, a sophisticated approach designed to quantify the extent to which a model has genuinely unlearned specific data points. By moving beyond traditional, often impractical, evaluation methods, IM provides a scalable and efficient solution to assess unlearning completeness. It reveals that many current unlearning techniques fall short of their promises, exposing significant "unlearning risks" where data is not adequately forgotten. This research is crucial for ensuring the integrity, compliance, and trustworthiness of machine learning systems in an era of increasing data privacy demands and complex model governance.

Background

▶ Watch: Introduction to machine unlearning and its significance (0:00)

Machine unlearning techniques are broadly categorized into three main types. Exact unlearning methods involve redesigning the entire training framework upfront, enabling a precise reversal process later. This might involve retraining a small subset of the model or reloading specific saved checkpoints. While these methods offer strong, often mathematical, guarantees of complete forgetting, they tend to be less efficient and flexible, requiring significant foresight in the training pipeline design. In contrast, opposing unlearning (also known as approximate unlearning) focuses on updating model parameters or adjusting model embeddings after the initial training. These methods are lauded for their efficiency and flexibility, as they don't necessitate a complete overhaul of the training process. However, this efficiency often comes at the cost of strict guarantees, leading to uncertainty about the true extent of data forgetting. The third category encompasses auditing and evaluation methods, which aim to verify whether unlearning has actually occurred.

A comprehensive survey of the literature over the past decade reveals a stark imbalance: the vast majority of research papers focus on approximate unlearning methods, with only a small fraction dedicated to exact unlearning. This trend highlights a community-wide rush towards efficiency and flexibility, often inadvertently neglecting the inherent risks associated with incomplete unlearning. The central question then becomes: when a model claims to have unlearned something, how much does it really forget, and can we truly trust these claims?

Traditionally, researchers have borrowed tools from Membership Inference Attacks (MIA) to evaluate unlearning. The basic premise of MIA involves training "shadow models" on "shadow datasets" to mimic the target model's behavior. By observing the shadow models' outputs on known member (training) and non-member (non-training) samples, an attack model can be constructed to infer membership. Two primary types of MIA exist: offline attacks apply a single, global threshold for all samples, which can lead to "false learning risks" – mistaking non-members with high confidence and resulting in a low true positive (TP) rate for unlearning evaluation. Online attacks, on the other hand, train a separate model and threshold for each individual query. While potentially more accurate, these are prohibitively expensive, with costs linearly increasing with the dataset size, sometimes even surpassing the cost of retraining the model from scratch. Neither traditional MIA approach offers a practical or scalable solution for rigorously evaluating unlearning completeness, underscoring the necessity for a more sophisticated and efficient framework.

Key Findings

▶ Watch: Limitations of traditional membership inference attacks (2:40)

The research presented by Cheng-Long Wang introduces a groundbreaking framework that addresses the significant limitations of traditional unlearning evaluation. The Interpolated Approximate Measurement (IM) framework is designed to provide a precise, sample-level assessment of unlearning completeness, overcoming the cost and accuracy issues inherent in existing Membership Inference Attacks (MIA).

A primary finding is the necessity of a more robust method for transforming model output confidences. The proposed Bounded Google map is a key technical innovation. It extends the standard Google map (a double logarithm transformation) by adding parameters that control sensitivity at extreme confidence values (near zero or one). This makes the transformation significantly more robust and reliable across the full spectrum of model predictions, addressing a critical flaw in previous approaches.

Furthermore, the talk highlights that for approximate unlearning, the notion of "membership" is no longer a binary (yes/no) state. Instead, it exists on a continuous spectrum from memorization to generalization. To capture this, IM introduces an interpolation-based framework that simulates different levels of "forgetting" by interpolating between "out responses" (model predictions for non-members) and "in responses" (predictions for members). This innovative approach allows IM to model the gradual shift from memorization to generalization without the need for costly online training, providing a nuanced understanding of unlearning progression.

Perhaps the most impactful finding is that the IM framework can efficiently measure unlearning completeness even when access is limited to just one shadow model. This is achieved by leveraging the bounded variance property of the Bounded Google map transformation and applying a "Public Velocity Inequality" (likely a reference to a concentration inequality like Chebyshev's inequality) to use cross-sample variance as a proxy for cross-model variance. This single-shadow model capability dramatically reduces the computational cost, making unlearning evaluation practical for real-world scenarios.

Through extensive empirical evaluations, IM demonstrates superior performance compared to existing MIA methods, particularly on large-scale models like Llama 2 7B. Crucially, IM reveals significant unlearning risks—specifically under-unlearning—in many popular approximate unlearning methods. It shows that these methods often fail to achieve true forgetting, retaining substantial memorization of the supposedly unlearned data. Moreover, IM can accurately distinguish true unlearning (where model parameters are genuinely altered) from superficial bypasses, such as prompt-based "unlearning" techniques that do not fundamentally change the model's internal representations. This capability is vital for identifying methods that merely mask data rather than truly forgetting it.

Technical Deep Dive

▶ Watch: Introducing the Bounded Google Map transformation (4:00)

The technical foundation of the IM framework rests on addressing the inherent limitations of using raw model confidence scores for unlearning evaluation. Raw confidences, and even standard metrics like cross-entropy loss, exhibit unstable scales and become overly sensitive at extreme values (close to 0 or 1), making them unsuitable for robust statistical analysis.

To overcome this, the framework first employs a Google map transformation. This is a double logarithm transformation that maps raw model outputs into a more uniform distribution, providing a more stable scale compared to cross-entropy loss. This transformation makes it feasible to fit the transformed outputs with a Gaussian distribution, defined by only its mean and variance – a crucial step for efficient parameter estimation.

However, the plain Google map still suffers from oversensitivity when confidence values are at the very extremes. To mitigate this, the authors introduce the Bounded Google map. This extension incorporates two small, tunable parameters that specifically control the transformation's sensitivity in these extreme regions. By bounding the output of the transformation, its variance is also inherently bounded. This property is critical because it allows for the application of a "Public Velocity Inequality" (as stated in the transcript, likely referring to a statistical concentration inequality such as Chebyshev's inequality). This inequality enables the use of cross-sample variance as a proxy for cross-model variance. The profound implication here is that the framework can efficiently estimate the hyperparameters of the Google distribution even with only one single shadow model, eliminating the need for costly multiple shadow models typically required to estimate cross-model variance.

The second core technical contribution addresses the nature of unlearning itself. In the context of approximate unlearning, the "membership" of a sample is no longer a binary state (either a training member or not). Instead, after an unlearning operation, the model's response for a supposedly unlearned sample shifts from one of strong memorization towards generalization. The ground truth itself is a spectrum. To model this, IM introduces an interpolation-based framework. The idea is to simulate this continuous progression from memorization to generalization. This is achieved by interpolating between "out responses" (model responses for non-member samples, representing full generalization) and "in responses" (model responses for original member samples, representing full memorization). By creating a spectrum of interpolated responses, the framework can effectively model different levels of unlearning progression for each individual sample. Importantly, this simulation avoids the prohibitive cost of training separate models for each level of generalization.

Once the Bounded Google map transformation is applied, each of these interpolated response distributions can be accurately fitted by a Gaussian distribution. The IM framework then estimates the probability of the unlearned model's response for a target query, determining where it falls along this memorization-to-generalization spectrum. A weighted average across this spectrum yields a membership score, which quantifies exactly where the sample lies on the memorization-to-generalization continuum.

The complete IM framework operates in four efficient steps:

  1. Bounded Google Map Application: Apply the Bounded Google map transformation to all relevant model responses, including those from the target model, a random model, and the single shadow model.
  2. Response Interpolation: Interpolate between the "out responses" and "in responses" to create the spectrum of simulated unlearning levels.
  3. Parameter Estimation: Efficiently estimate the parameters of the Gaussian distribution for each interpolated response using the moment of method for hyperparameters.
  4. Membership Score Computation: Compute the final membership score for each query, indicating its degree of memorization or generalization.

A critical advantage of this design is its inherent efficiency and parallelism. All computations can be performed in parallel, and crucially, the entire framework requires only one shadow model, making it a highly scalable and practical solution for evaluating unlearning in real-world large-scale machine learning systems.

Demo / Proof of Concept

▶ Watch: Interpolation framework for scrub-based unlearning completeness (6:00)

While the talk did not feature a live, interactive demo, Cheng-Long Wang presented compelling empirical evidence and evaluations that served as a robust proof of concept for the Interpolated Approximate Measurement (IM) framework. These demonstrations showcased IM's superior performance and its ability to uncover critical unlearning risks that traditional methods overlook.

The first set of evaluations compared IM against several baseline methods on both binary inference tasks (member vs. non-member) and score-based inference tasks (degree of memorization). For binary inference, IM achieved "nearly perfect" results, indicated by a high AUC-ROC or similar performance metric. In score-based inference, IM demonstrated a "very strong correlation with the ground truth," confirming its accuracy in quantifying the extent of memorization. Furthermore, IM's robustness was tested in an offline setting where the original models were inaccessible, and it still performed remarkably well, highlighting its practical applicability in various deployment scenarios.

A particularly insightful demonstration involved continuously training a shadow model on an original training dataset and saving checkpoints at different stages, from generalization to memorization. The goal was to see if IM could accurately capture this trajectory. As expected, when IM was applied, the membership score for samples increased consistently with the training epochs, precisely reflecting the model's shift from generalization towards memorization. In stark contrast, the baseline methods showed "almost no significant trend at all," proving their inability to track the subtle progression of memorization.

The framework's scalability and efficacy on large-scale models were a major highlight. IM was applied to Llama 2 7B models, a significant benchmark in the field of large language models. On these complex models, IM not only achieved the best results compared to "the best existing membership inference methods" but also provided crucial insights into different "unlearning" approaches. The researchers tested several "bypass methods," such as negative label unlearning and prompt-based methods, which aim to make the model "forget" without actually altering its internal parameters. IM consistently assigned high membership scores to samples that were merely suppressed by these prompt-based methods, indicating that they were "not truly unlearning" and flagging them as "unlearning risks." Conversely, for samples subjected to exact unlearning, IM correctly assigned low membership scores, demonstrating its ability to distinguish genuine forgetting from superficial suppression.

Finally, the talk presented a comprehensive benchmark of seven different approximate unlearning methods. The results, visualized with red indicating "under-unlearning" (high membership scores) and yellow indicating "overall learning" (low membership scores, potentially too aggressive), revealed "general learning risks." A critical finding was the prevalence of under-unlearning, where many of these approximate methods "largely failed to achieve a true forgetting." This empirical evidence unequivocally demonstrates IM's value in not only measuring unlearning performance but also pinpointing precisely where and how much unlearning risk exists, a capability that other evaluation methods frequently overlook.

Defensive Implications

▶ Watch: Overview of the Interpolated Approximate Measurement (IAM) framework (7:00)

The findings presented in this talk carry profound implications for defenders and organizations deploying machine learning models, especially in privacy-sensitive or security-critical environments. The most significant takeaway is that current approximate unlearning methods, while efficient, often fall short of achieving true forgetting and pose substantial unlearning risks, particularly under-unlearning. This means that organizations cannot simply rely on the claims of unlearning mechanisms without independent, rigorous verification.

Defenders must recognize that the "right to be forgotten" and other data governance requirements necessitate more than just a superficial suppression of data. True unlearning requires the model to behave as if the data was never seen, and IM provides the first practical tool to audit this commitment at a granular, sample level. This insight demands a shift in defensive strategies:

  1. Adopt Robust Verification: Organizations should integrate robust unlearning evaluation frameworks, such as IM, into their MLOps pipelines. This enables continuous auditing of unlearning requests, ensuring that sensitive data, harmful content, or proprietary information is genuinely removed from models.
  2. Identify and Quantify Risks: IM's ability to pinpoint specific instances of under-unlearning allows defenders to identify which data points, or even which segments of the model, are failing to forget. This granular insight can guide targeted remediation efforts or inform decisions about retraining if unlearning is incomplete.
  3. Distinguish True Forgetting from Bypass Methods: The framework effectively differentiates between true unlearning (where model parameters are altered) and superficial bypasses (like prompt-based methods that merely suppress output). Defenders can use IM to ensure that the unlearning solutions they deploy are genuinely effective and not just masking the problem, which is crucial for compliance and security.
  4. Benchmark and Select Unlearning Techniques: With IM, organizations can objectively benchmark different approximate unlearning methods. This allows them to select techniques that offer the strongest guarantees of forgetting, balancing efficiency with completeness, and understanding the specific trade-offs involved.
  5. Enhance Regulatory Compliance: For industries subject to strict data privacy regulations (e.g., GDPR, CCPA), IM offers a verifiable mechanism to demonstrate compliance with "right to be forgotten" requests. It provides concrete evidence of data removal, moving beyond mere assertions to demonstrable proof.
  6. Mitigate Security Vulnerabilities: Unlearning is also critical for removing backdoors or poisoned data. IM can help verify that such malicious inputs have been truly expunged from the model's memory, thereby enhancing the overall security posture of AI systems.

In essence, IM empowers defenders to move from a trust-based approach to a verify-based approach for machine unlearning, establishing a critical layer of assurance and accountability in the lifecycle of AI models.

Key Takeaways

  • Machine unlearning is a critical capability for data privacy, security, and model governance, allowing models to "forget" specific data points without full retraining. However, current approximate unlearning methods often prioritize efficiency over strict guarantees.
  • Traditional Membership Inference Attacks (MIA) are largely impractical for robust, scalable unlearning evaluation due to their high computational cost (online MIA) or accuracy limitations (offline MIA).
  • The Interpolated Approximate Measurement (IM) framework offers an efficient, scalable, and accurate solution for measuring sample-level unlearning completeness, uniquely operating effectively with just one shadow model.
  • IM leverages a novel Bounded Google map transformation for robust handling of model confidence scores, particularly at extreme values, and an interpolation-based approach to simulate the continuous spectrum from memorization to generalization.
  • Empirical evaluations, including on large-scale Llama 2 7B models, demonstrate IM's superior performance in identifying significant under-unlearning risks in existing approximate unlearning methods, revealing that many fail to achieve true forgetting.
  • Defenders should integrate robust evaluation frameworks like IM into their MLOps pipelines to audit unlearning processes, ensure compliance with privacy regulations, distinguish true unlearning from superficial bypasses, and effectively mitigate unlearning risks in real-world AI deployments.

About the Speaker(s)

The talk was presented by Cheng-Long Wang. The transcript indicates this was a joint work with "Chile, Professor and Professor Da," suggesting Cheng-Long Wang is part of a research team, likely from an academic institution, focusing on machine learning security and privacy. Further details about his specific title or affiliation were not provided in the input.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Legitimate academic security research on a real problem — verifying that machine unlearning actually works — with a technically coherent contribution in the IM framework and Bounded Google map. Solid USENIX-tier paper material, but as a conference talk it reads like a paper walkthrough rather than a compelling presentation, and the threat model framing is more compliance-adjacent than security-first.

Heather Calloway (CISO) — WEAK

Technically credible research on a real and growing governance problem — machine unlearning verification — but the talk never exits the academic register. The defensive implications section gestures at organizational relevance without earning it, and no CISO or compliance officer leaves knowing what to actually do with this.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)