Provably Unlearnable Data Examples
Derui Wang
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Machine Unlearning
Overview
In an era dominated by large language models and advanced machine learning, the ease with which public data can be exploited poses significant risks, ranging from intellectual property infringement to privacy breaches. Derui Wang's talk, "Provably Unlearnable Data Examples," introduces a groundbreaking framework designed to certify the learnability of data, offering robust protection against unauthorized model training and exploitation. This work, a collaborative effort from CSRO's Data61 and the University of Chicago, and supported by the Cyber Security Cooperative Research Centre of Australia, addresses the critical need for a quantifiable guarantee on data's unlearnability.
Key moments
- 0:00 Introduction: Risks of public data learnability
- 2:00 Real-world impact: Artist's IP damaged by AI fine-tuning
- 3:00 Unlearnable examples' flaws and the new recovery attack
- 4:50 Introducing the Certified Data Learnability framework
- 5:00 Three-part framework for certified data learnability explained
- 9:00 Provably Unlearnable Examples (PUEs) for robust protection
- 10:00 PUEs demonstrate strong defense against recovery attacks
Provably Unlearnable Data Examples
Speakers: Derui Wang
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=Xgakc54my2E
Overview
In an era dominated by large language models and advanced machine learning, the ease with which public data can be exploited poses significant risks, ranging from intellectual property infringement to privacy breaches. Derui Wang's talk, "Provably Unlearnable Data Examples," introduces a groundbreaking framework designed to certify the learnability of data, offering robust protection against unauthorized model training and exploitation. This work, a collaborative effort from CSRO's Data61 and the University of Chicago, and supported by the Cyber Security Cooperative Research Centre of Australia, addresses the critical need for a quantifiable guarantee on data's unlearnability.
The core of the presentation revolves around establishing a certified data learnability score and generating Provably Unlearnable Examples (PUEs). Unlike previous attempts at creating unlearnable data, this framework provides a theoretical upper bound on the utility that any adversary's model can extract from the protected data, given certain constraints. This innovation is crucial for data owners who wish to publish data without risking its misuse, offering a new paradigm for data availability guarantees and intellectual property protection in the age of pervasive AI.
Background
▶ Watch: Introduction: Risks of public data learnability (0:00)
The pervasive learnability of public data presents a multifaceted threat landscape in machine learning. Historically, attackers have leveraged public datasets to train shadow classifiers for membership inference attacks, determining if a specific data point was part of a model's training set. Data can also be used to train substitute models for launching transferable adversarial attacks against black-box models, bypassing direct access limitations.
With the rapid advancements in foundation models, particularly open large language models (LLMs) and large vision models (LVMs), coupled with parameter-efficient fine-tuning (PFT) techniques, data exploitation has become significantly easier and more potent. Attackers now require less data to inflict greater harm. For instance, data scraped from social media can be used to fine-tune pre-trained models, transforming them into adversarial domain experts capable of exposing sensitive knowledge or mimicking proprietary styles. A notable example cited in the talk is an artist whose twenty paintings were used to fine-tune a Stable Diffusion model, which subsequently generated images mimicking the artist's unique style, directly infringing upon their intellectual property. Despite the existence of laws and regulations aimed at preventing such exploitation, a strict, verifiable guarantee on data learnability has been conspicuously absent.
To counter these threats, a line of research has emerged focused on unlearnable examples. These approaches perturb data points before publication, aiming to degrade the performance of machine learning models trained on them when evaluated on clean data from the same domain. While promising, existing unlearnable examples suffer from several critical limitations:
- Generalization Issues: Their effectiveness often varies widely across diverse "pirate models" and training strategies encountered in the wild, lacking robust transferability.
- Evaluation Difficulty: The inherent stochasticity of the machine learning training process makes it challenging to consistently evaluate the true unlearnability of perturbed examples.
- Recovery Attacks: A new, potent threat identified in this work is the recovery attack. In this scenario, an attacker can subtly perturb the weights of a pirate model that was initially trained on unlearnable examples. By then fine-tuning this model with a small fraction of clean data and employing projected Stochastic Gradient Descent (SGD), the model's performance on clean data can be restored to near-normal levels. This attack is particularly relevant in federated learning settings, where a malicious client could upload a locally recovered model to poison the global model, violating data privacy or availability protocols.
These limitations highlight the vulnerability of current unlearnable example techniques, underscoring the necessity for a more robust, provable framework for data unlearnability.
Key Findings
▶ Watch: Unlearnable examples' flaws and the new recovery attack (3:00)
The central question addressed by this research is: "Given a set of unlearnable examples, can we establish an upper bound on the utility of pirate models trained on them?" The paper asserts that while a thorough, exhaustive answer is impractical due to the vast permutations of models and algorithms, it is indeed possible to derive such an upper bound under specific constraints. This led to the development of the first certification framework for the effectiveness and robustness of unlearnable examples.
The key findings and contributions include:
- Certified Data Learnability: The introduction of a novel concept, QA learnability, which provides a quantifiable, provable upper bound on the maximum utility attainable by pirate models. This guarantee holds true as long as the pirate model's parameters fall within a defined certified parameter set.
- Quantile Parametric Smoothing (QPS) Function: The development of a QPS function (also referred to as QBS in the talk) that takes randomized model parameters as input and computes the model's utility at a given quantile (Q), offering a robust measure across potential model states.
- Random Weight Perturbation (RWP) for Surrogates: The discovery that incorporating Random Weight Perturbation (RWP) during the training of surrogate models significantly enhances their suitability for the certification process. RWP-trained surrogates produce higher certified learnability scores for a given set of unlearnable examples, indicating they lead to better certification.
- Provably Unlearnable Examples (PUEs): Leveraging RWP-trained surrogates, the framework generates Provably Unlearnable Examples (PUEs). These PUEs demonstrate lower certified learnability compared to baselines, signifying more robust protection against adversaries operating within the certified parameter set.
- Robustness Against Recovery Attacks: Beyond the certified parameter set, PUEs exhibit superior protection against general adversaries, specifically proving harder to recover from recovery attacks compared to other unlearnable example baselines like EMN and OPS. This demonstrates their enhanced resilience in real-world adversarial scenarios.
- Impact of Certification Parameters: Experimental results indicate that using larger values for the quantile parameter (Q) and the perturbation bound parameter (ETA) in the certification process allows for the certification of higher learnability scores, providing flexibility in setting protection levels.
Technical Deep Dive
▶ Watch: Introducing the Certified Data Learnability framework (4:50)
The proposed framework for certified data learnability is structured into three main parts, designed to generate provably unlearnable examples and then certify their effectiveness.
Part 1: Perturbation Generation
The first stage focuses on generating the perturbations (delta) that will transform a clean dataset (DS) into an unlearnable one. This process relies on a surrogate model, which is a representative model used to guide the perturbation generation. The goal is to find a delta that, when applied to the dataset, minimizes the training loss of this surrogate model. This counter-intuitive approach aims to make the data "unlearnable" by driving the model towards a state where it struggles to fit the perturbed examples effectively.
A critical innovation in this stage is the use of Random Weight Perturbation (RWP) during the training of the surrogate model. RWP involves introducing Gaussian noise to the weights of the surrogate during its training. This technique is specifically employed to make the surrogate more robust and suitable for the subsequent certification process, ultimately leading to the generation of more robust Provably Unlearnable Examples (PUEs).
Part 2: Certified Learnability Computation (QA Learnability)
Once the perturbations are generated and applied, the second stage focuses on certifying the learnability of the resulting unlearnable examples. This involves evaluating how well various "pirate models" could potentially learn from this perturbed data.
- Randomized Weight Evaluation: The previously trained surrogate model's weights are randomized using Gaussian noise multiple times (for 'N' times). For each randomized copy of the surrogate, its utility is evaluated using a specific metric (A) on a test dataset. The framework is generic, meaning 'A' can be any utility matrix applicable to various tasks and data modalities.
- QA Learnability Computation: The results from these evaluations are then fed into a Quantile Parametric Smoothing (QPS) function (also referred to as QBS in the talk). This function is central to computing the QA learnability score. The QPS function takes the randomized parameters (weights) of the model as input and computes the model's utility at a given quantile 'Q' among all attainable model utilities within the population of randomized models.
The definition of QA learnability provides a guarantee on the maximum attainable utility of pirate models. This guarantee is conditional: it holds true as long as the trained pirate models have their parameters falling within a certified parameter set. This set is derived with the help of a perturbation bound on the QPS function, and its scale is influenced by a parameter called ETA. The certification algorithms and detailed theorem proofs for these bounds are available in the full paper.
Part 3: Generalization Learnability Score
Recognizing that defenders might not always have access to a private test dataset, the third stage provides an alternative. In such scenarios, the defender can sample a test set from a closed domain. This sampled test set is then used to compute a generalization learnability score with the assistance of Hinge-Span. This allows for a practical assessment of unlearnability even when ideal conditions (i.e., a perfectly representative private test set) are not met.
The certification framework is notable as the first of its kind to provide a verifiable guarantee on the effectiveness and robustness of unlearnable examples. Experimental findings further illuminated key properties of this certification:
- Increasing the values of 'Q' (the quantile) and 'ETA' (the perturbation bound parameter) in the certification process allows for the certification of higher learnability scores. This provides a mechanism to adjust the level of "unlearnability" or the scope of the certified parameter set.
- The use of RWP-trained surrogates (RWP stargates) was empirically shown to produce higher certified learnability on the same set of unlearnable examples compared to baseline surrogates. This indicates that RWP surrogates are indeed more effective in creating robust certification models.
- When comparing PUEs generated using RWP surrogates against baselines, PUEs consistently demonstrated lower certified learnability. This translates to PUEs offering more robust protection against adversaries whose models fall within the certified parameter set.
Demo / Proof of Concept
▶ Watch: Provably Unlearnable Examples (PUEs) for robust protection (9:00)
While the talk did not feature a live, interactive demonstration in the traditional sense, the speaker extensively discussed the experimental evaluations and proofs of concept conducted to validate the framework's claims. These evaluations served as the practical demonstration of the theoretical underpinnings.
The efficacy of Provably Unlearnable Examples (PUEs) was rigorously tested against various baselines and attack scenarios:
- Comparison with Baselines: PUEs were compared against standard baselines for unlearnable examples. Under the same training methodology (specifically, using RWP surrogates), PUEs consistently achieved lower certified learnability scores. This quantitative result directly supports the claim that PUEs offer more robust protection against adversaries constrained by the certified parameter set.
- Protection Against General Adversaries: A crucial aspect of the evaluation involved testing PUEs against adversaries operating beyond the certified parameter set. This was done by simulating recovery attacks, where an attacker attempts to restore model performance on unlearnable data. PUEs were compared against two other notable unlearnable example baselines: EMN and OPS. The results showed that it was significantly harder to recover models trained on PUEs than to recover models trained on EMN or OPS. This finding is critical, as it demonstrates that PUEs provide enhanced resilience not just against "certifiable" adversaries (those whose models fall within the certified parameter set), but also against more general, unconstrained adversarial attempts to bypass the unlearnability.
These experimental results, detailed in the paper and summarized in the talk, serve as the empirical proof of concept, demonstrating the practical robustness and effectiveness of the proposed framework and the PUEs it generates.
Defensive Implications
▶ Watch: PUEs demonstrate strong defense against recovery attacks (10:00)
The framework for provably unlearnable data examples offers several crucial implications for defenders seeking to protect their data and intellectual property in an increasingly data-driven world:
- Certified Learnability as an Evaluation Metric: For the first time, data owners and security practitioners have a quantifiable and provable metric – certified learnability – to assess the effectiveness of unlearnable examples. This moves beyond heuristic evaluations, providing a robust, theoretically grounded measure of how resistant perturbed data is to unauthorized learning. This metric can guide the selection and deployment of unlearnable data techniques.
- Data Availability Guarantees: Before publishing any data to public platforms, organizations can now leverage certified learnability to obtain a concrete data availability guarantee. This means they can make data publicly accessible with a verifiable assurance that its utility for malicious purposes (e.g., fine-tuning adversarial models, IP infringement) is bounded by a known, low level. This empowers data sharing without compromising sensitive information or creative works.
- Generation of Robust Provably Unlearnable Examples (PUEs): The methodology enables the creation of highly robust Provably Unlearnable Examples (PUEs). These PUEs, especially when generated with Random Weight Perturbation (RWP) surrogates, offer superior protection against both certifiable adversaries (whose models fall within a defined parameter set) and general adversaries, including those attempting sophisticated recovery attacks. Defenders can actively transform their datasets into PUEs before dissemination, significantly raising the bar for data exploitation.
- Informing Policy and Regulation: The concept of certified learnability can provide a technical foundation for future laws and regulations concerning data privacy and intellectual property in AI. By offering a measurable and provable standard, it can help bridge the gap between legal intent and technical enforceability.
The speaker also outlined directions for future improvements, which can further enhance defensive capabilities:
- Expanded Coverage: Improving certification algorithms to expand the certified parameter set would allow the framework to encompass a wider range of pirate models encountered in real-world scenarios, thereby increasing the scope of protection.
- Improved Tightness: Training even better surrogate models (stargates) could lead to certifying higher learnability scores, meaning the bounds on learnability could be made tighter, offering more precise guarantees.
- Enhanced Efficiency: Accelerating the PUE generation and surrogate training processes would make this powerful protection more practical and scalable for large datasets.
Key Takeaways
- Data Exploitation is a Growing Threat: The proliferation of foundation models and fine-tuning techniques makes public data highly vulnerable to misuse, including IP infringement and privacy breaches, necessitating robust protection.
- Current Unlearnable Examples Fall Short: Existing unlearnable data methods suffer from generalization issues, difficult evaluation, and susceptibility to novel recovery attacks, highlighting a critical gap in data protection.
- Certified Learnability Provides Guarantees: The introduced framework offers the first method to establish a provable upper bound on the utility that pirate models can extract from perturbed data, known as QA learnability, within a defined certified parameter set.
- Provably Unlearnable Examples (PUEs) Offer Superior Robustness: PUEs, generated using Random Weight Perturbation (RWP)-enhanced surrogate models, demonstrate significantly lower certified learnability and are more resilient against recovery attacks compared to other unlearnable data baselines.
- Actionable Defense Strategies: Defenders can use certified learnability as an evaluation metric, leverage PUEs for data protection, and obtain verifiable data availability guarantees before publishing sensitive or valuable information.
- Future Enhancements are Planned: Ongoing research aims to expand the coverage of certified models, tighten learnability bounds, and improve the efficiency of PUE generation to enhance the framework's practical applicability.
About the Speaker(s)
The talk "Provably Unlearnable Data Examples" was presented by Derui Wang. He is a key author of the paper, representing a collaborative effort between CSRO's Data61 and the University of Chicago. The project itself received funding and support from the Cyber Security Cooperative Research Centre of Australia. While specific titles or further biographical details were not provided in the transcript, his affiliation with prominent research institutions like CSRO's Data61 and the University of Chicago underscores his expertise in cybersecurity and machine learning research.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid, technically grounded ML security research that fills a real gap: the first certification framework for unlearnable data examples, complete with provable upper bounds on adversarial utility extraction. The recovery attack identification alone justifies the slot — it's a novel threat vector that invalidates prior defenses and the proposed RWP-based countermeasure has clear theoretical backing.
Heather Calloway (CISO) — WEAK
Technically credible research on a real and growing problem — unauthorized model training on public data is not hypothetical. But this talk never crosses the line from academic contribution to operational reality, and the gap between the math and anyone's security program is never bridged.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025