SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
Kaiyuan Zhang
34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Vulnerabilities in LLMs: Privacy, Safety, and Defense
Overview
This talk, presented by Kaiyuan Zhang, introduces SOFT, a novel defense mechanism designed to protect the privacy of large language models (LLMs) during the crucial fine-tuning phase. As LLMs become ubiquitous, adapting these powerful general models to specific, often sensitive, real-world tasks through fine-tuning has become standard practice. However, this process frequently involves proprietary or private datasets, introducing significant privacy risks, most notably Membership Inference Attacks (MIAs).
Watch on YouTube · Read the paper · Download the PDF (PDF) · Slides
Paper abstract
We study a phishing attack against password manager browser extensions. Browser extension UIs are mostly displayed on top of the web browser's viewport and, thus, hard to distinguish from website content. This enables an attacker to phish master passwords by imitating a locked password manager on a website they control. We implemented this attack for four password managers and demonstrated its effectiveness in a large-scale phishing simulation with 29,800 participants, among whom we detected over 400 instances of selected third-party password managers. Notably, more than 30% of these users entered their master password, with up to 58% for one specific password manager. We compare the effectiveness of the attack across different password manager UIs, analyze user behavior through mouse tracking and a post-study survey, and discuss the implications of our findings for password managers as a means of phishing protection.

Key moments
- 0:00 Introduction to privacy risks in LLM fine-tuning
- 3:00 Standard fine-tuning is extremely vulnerable to MIAs
- 4:00 LoRA and similar methods are not a complete fix
- 4:50 Introducing SOFT: protecting influential data points
- 5:30 Detailed explanation of SOFT's four-step pipeline
- 6:50 Experimental results demonstrating SOFT's strong privacy protection
- 8:00 Summary of key findings and practical solution
SOFT: Selective Data Obfuscation for Protecting LLM Fine-tuning against Membership Inference Attacks
Speakers: Kaiyuan Zhang
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=Z3WQodl1DrA
Overview
This talk, presented by Kaiyuan Zhang, introduces SOFT, a novel defense mechanism designed to protect the privacy of large language models (LLMs) during the crucial fine-tuning phase. As LLMs become ubiquitous, adapting these powerful general models to specific, often sensitive, real-world tasks through fine-tuning has become standard practice. However, this process frequently involves proprietary or private datasets, introducing significant privacy risks, most notably Membership Inference Attacks (MIAs).
The core problem addressed by SOFT is that models tend to overfit and "memorize" specific training data points, making them vulnerable to MIAs where an attacker can determine if a particular data sample was part of the training set. While previous research largely dismissed MIAs against the massive pre-training datasets of LLMs, the speaker highlights that fine-tuning, with its smaller, more focused, and frequently re-used sensitive datasets, presents a far greater and more practical threat. SOFT offers a targeted, effective, and scalable solution to this vulnerability, balancing robust privacy protection with minimal impact on model performance.
The research, a collaboration between Purdue University and Cisco Research, underscores the urgent need for privacy-preserving techniques in the LLM ecosystem. By demonstrating the severe vulnerability of current fine-tuning practices and providing a practical defense, SOFT aims to bolster the secure deployment of LLMs in applications handling sensitive information, such as medical records, proprietary code, or customer logs. This work is critical for fostering trust and enabling the responsible adoption of AI technologies across various industries.
Background
▶ Watch: Introduction to privacy risks in LLM fine-tuning (0:00)
The rapid advancements in large language models have made them indispensable tools, with fine-tuning emerging as a standard procedure to adapt their general capabilities to specialized domains. This adaptation, however, often necessitates training on private and sensitive datasets, thereby introducing significant privacy concerns. Foremost among these is the Membership Inference Attack (MIA), where an adversary aims to determine whether a specific data sample was included in the model's training set.
Historically, early research on MIAs against LLMs primarily focused on the pre-training phase. However, pre-training involves colossal, web-scale datasets, where individual data points are typically seen only once. Consequently, recent studies have largely concluded that MIAs against pre-training data are ineffective due to the sheer scale and sparsity of data exposure. The resources required for pre-training, exemplified by models like GPT-4 costing over $100 million and utilizing 25,800 GPUs, further underscore the impracticality of attacking this stage.
The real and more pressing danger, as highlighted by the speaker, resides in the fine-tuning stage. Here, datasets are significantly smaller, often contain highly sensitive information (e.g., personal data, medical records, proprietary code, customer logs), and models are trained for multiple epochs. This repeated exposure dramatically increases the model's opportunity to memorize specific data points, making it highly susceptible to MIAs. The underlying signal exploited by these attacks is a simple observation: models tend to exhibit higher confidence (reflected in a lower loss value) for data they have been trained on.
A key challenge for attackers is to differentiate a low loss value resulting from training set membership from a low loss value simply because the sample is common or easy to predict. Existing MIA methods employ various techniques to address this calibration problem. To establish a robust and realistic evaluation framework for their defense, the researchers developed a powerful ensemble attack that integrates features from multiple state-of-the-art MIA methods.
Using this strong ensemble attack, initial measurements revealed that standard full tuning (where all model parameters are updated) is extremely vulnerable to MIAs. The AOC score (Area Under the Receiver Operating Characteristic Curve), a common metric for attack success, consistently showed high values across diverse data types, including code, academic papers, and general text. Critically, this vulnerability was observed to worsen with larger models and an increased number of training epochs. Perhaps most concerning, the study found that even a single pass over the fine-tuning data was sufficient to induce significant privacy leakage.
Recognizing the limitations of full tuning, the researchers also investigated whether more parameter-efficient methods, such as LoRA (Low-Rank Adaptation), could mitigate this problem. While LoRA fine-tuning generally leaked less information than full tuning, offering a better balance between privacy and model performance, it was not a complete solution. The attack scores for LoRA remained well above random guessing (an AOC of 0.5), indicating that a determined adversary employing a strong method like their ensemble attack could still successfully identify training members. This finding underscored the necessity for a more targeted and effective defense mechanism, directly motivating the development of SOFT.
Key Findings
▶ Watch: LoRA and similar methods are not a complete fix (4:00)
The research presented in this talk reveals several critical findings concerning the privacy of LLM fine-tuning and offers a robust solution:
- Standard Fine-tuning is Highly Vulnerable: Traditional full tuning of LLMs on private datasets presents a significant and practical privacy risk from Membership Inference Attacks (MIAs). Experiments using a powerful ensemble attack demonstrated consistently high AOC scores (e.g., 0.76 for full tuning on the archive dataset), indicating severe privacy leakage, which worsens with larger models and more training epochs. Even a single epoch can lead to substantial vulnerability.
- Parameter-Efficient Methods Offer Limited Protection: While methods like LoRA fine-tuning improve privacy compared to full tuning, they are not a complete fix. LoRA still exhibited AOC scores (e.g., 0.59) well above random guessing (0.5), meaning a determined attacker can still successfully identify training data members. This highlights the need for more targeted privacy defenses.
- Influential Samples Drive Privacy Risk: The core insight motivating SOFT is that a small fraction of the training data, termed influential samples, accounts for a disproportionately large portion of the privacy risk. These are the data points the model learns very quickly and assigns exceptionally low loss values to, making them easily identifiable by MIAs.
- SOFT Effectively Neutralizes MIAs: The proposed defense, SOFT (Selective Data Obfuscation for Protecting LLM Fine-tuning), is highly effective. By selectively identifying and obfuscating these influential data points, SOFT successfully reduces the MIA AOC score to approximately 0.52, very close to random guessing (0.5). This demonstrates that SOFT can effectively neutralize MIAs against fine-tuned LLMs.
- Minimal Utility Trade-off: Crucially, SOFT achieves this strong privacy protection with only a minimal and acceptable trade-off in model utility. The paper, as detailed by the speaker, confirms that the model's performance on its intended task remains high, making SOFT a practical and viable solution for real-world deployment.
Technical Deep Dive
▶ Watch: Introducing SOFT: protecting influential data points (4:50)
The core problem SOFT addresses is the model's tendency to memorize specific data points during fine-tuning, leading to vulnerability against Membership Inference Attacks (MIAs). The intuition behind SOFT is straightforward: if the model memorizes certain data too well, then those exact data points should not be presented to the model in their original form. However, simply altering the entire dataset would inevitably degrade the model's performance and utility. The challenge, therefore, lies in identifying which samples to modify and how to modify them without compromising the model's learning objectives.
SOFT draws inspiration from concepts like influence functions, which suggest that a small, identifiable fraction of the training data contributes disproportionately to the privacy risk. These are the influential samples—data points that the model learns exceptionally quickly, resulting in very low loss values. These are the specific samples that SOFT targets for protection.
The SOFT defense operates as a four-step pipeline, seamlessly integrated into the standard fine-tuning process:
- Warm-up Pass:
- The process begins with a warm-up pass, where the model is trained for a single epoch on the entire original fine-tuning dataset.
- The purpose of this initial pass is to establish a baseline understanding of how the model interacts with each individual training sample. It allows the model to begin learning and for preliminary loss values to be computed for each data point, providing the necessary data for the subsequent selection step.
- Influential Data Selection:
- This is the most critical step. In each subsequent epoch after the warm-up pass, SOFT calculates the loss value for every training sample.
- Each sample's loss is then compared against a predefined threshold. This threshold is dynamically set as the average loss calculated on a clean validation set. Using a clean validation set ensures that the threshold reflects a general understanding of "easy" or "well-learned" data points, rather than specific training set memorization.
- If a sample's loss is found to be lower than this threshold, it is flagged as an influential and vulnerable sample. These are the data points that the model is rapidly memorizing and that pose the highest risk of privacy leakage.
- Data Obfuscation:
- Any sample that has been flagged as influential and vulnerable in the previous step is then sent to a paraphrasing model. The speaker specifically mentioned using models like GPT-4 for this task, leveraging their advanced natural language generation capabilities.
- The paraphrasing model generates a semantically equivalent version of the original sensitive text. This means the core meaning and intent of the data are preserved, ensuring the model can still learn the desired task, but the exact textual representation is altered.
- Crucially, the original, "safe" data points (those whose loss values were above the threshold) remain entirely untouched, preserving their fidelity for training.
- Fine-tuning:
- Finally, the model proceeds to fine-tune on a newly constructed dataset. This dataset comprises the original safe samples combined with the newly obfuscated (paraphrased) versions of the influential samples.
- By training on these paraphrased versions of the most vulnerable data, SOFT directly prevents the model from overfitting on their exact text. This mechanism directly undermines the primary signal (exact match leading to very low loss) that Membership Inference Attacks exploit.
The effectiveness of SOFT was empirically validated through extensive experiments. The baseline pre-trained model typically showed an AOC score of approximately 0.5, representing random guessing and thus no privacy leakage. Standard full tuning, however, dramatically increased vulnerability, with an AOC reaching around 0.76. While LoRA offered some improvement, reducing the AOC to about 0.59, it still indicated significant leakage. In stark contrast, SOFT successfully brought the attack AOC score down to approximately 0.52, effectively neutralizing the attack by making it nearly as successful as random guessing. This robust privacy protection was achieved, as detailed in the paper, with only a minimal and acceptable trade-off in model utility, confirming SOFT's practicality and scalability.
Demo / Proof of Concept
▶ Watch: Experimental results demonstrating SOFT's strong privacy protection (6:50)
The talk focuses on the technical methodology and experimental validation of the SOFT defense rather than showcasing a live demonstration of its operation. While the presentation describes "how SOFT works in practice" as a four-step pipeline, it does not detail a specific interactive demo or a standalone proof-of-concept tool that was publicly presented or used during the talk itself.
Instead, the effectiveness of SOFT was demonstrated through rigorous experimental results, comparing its performance against baseline models and other fine-tuning methods (full tuning, LoRA) using a powerful ensemble attack developed by the researchers. The AOC scores presented (e.g., 0.76 for full tuning, 0.59 for LoRA, and 0.52 for SOFT) serve as the primary evidence of the defense's efficacy. The speaker also mentioned that "All of our code and data are available at the link provided," indicating that the practical implementation and validation of SOFT can be replicated and explored by interested researchers and practitioners. This availability of resources acts as a form of proof-of-concept, allowing others to verify the claims and potentially build upon the work.
Defensive Implications
▶ Watch: Summary of key findings and practical solution (8:00)
The findings from the SOFT research carry significant implications for developers, security professionals, and organizations deploying Large Language Models (LLMs) in real-world scenarios, particularly when fine-tuning with sensitive data.
First and foremost, the research unequivocally establishes that standard fine-tuning practices, including both full tuning and parameter-efficient methods like LoRA, expose LLMs to practical and significant privacy risks from Membership Inference Attacks (MIAs). This means that simply adopting LLMs and fine-tuning them on proprietary or personal data without explicit privacy-preserving mechanisms is inherently dangerous. Organizations handling sensitive information, such as medical records, financial data, or intellectual property, must recognize this vulnerability as a critical threat.
Defenders should move beyond the assumption that LLMs are inherently private during fine-tuning. Instead, they must proactively integrate privacy-enhancing technologies into their LLM deployment pipelines. SOFT provides a concrete, actionable blueprint for such integration:
- Proactive Risk Identification: Implement mechanisms to identify "influential samples" within fine-tuning datasets. This involves monitoring the loss values of individual training samples and comparing them against a dynamically determined threshold (e.g., average loss on a clean validation set). This step helps pinpoint the specific data points that are most prone to memorization and, consequently, privacy leakage.
- Selective Data Obfuscation: For identified influential samples, employ data obfuscation techniques, specifically paraphrasing, to alter their exact textual representation while preserving their semantic meaning. This prevents the model from overfitting on the precise wording of sensitive data. Utilizing advanced LLMs like GPT-4 for paraphrasing, as suggested by the research, can ensure high-quality and semantically consistent alterations.
- Integrate into Fine-tuning Workflow: The SOFT pipeline is designed to integrate directly into existing fine-tuning processes. This suggests that privacy protection can be added as an intermediate step within the training loop rather than requiring a complete overhaul of the LLM development process.
- Balance Privacy and Utility: Defenders must consider the trade-off between privacy protection and model utility. SOFT demonstrates that strong privacy can be achieved with minimal performance degradation. However, the exact balance might need to be fine-tuned based on the specific application's requirements and the sensitivity of the data.
- Leverage Open-Source Resources: The availability of the code and data from the SOFT project is a significant advantage. Defenders can leverage these resources to understand the implementation details, replicate the experiments, and potentially adapt the SOFT methodology to their specific LLM architectures and datasets. This fosters a community-driven approach to enhancing LLM privacy.
In summary, the defensive implication is clear: organizations must adopt a "privacy-by-design" approach for LLM fine-tuning. Relying solely on default training practices or generic parameter-efficient methods is insufficient. Technologies like SOFT offer a practical and effective means to mitigate the substantial privacy risks associated with fine-tuning LLMs on sensitive data, thereby enabling safer and more responsible AI deployment.
Key Takeaways
- Significant MIA Risk in Fine-tuning: Standard LLM fine-tuning, especially with sensitive data, creates a practical and severe vulnerability to Membership Inference Attacks (MIAs), with high AOC scores demonstrating substantial privacy leakage.
- LoRA is Insufficient: While parameter-efficient methods like LoRA offer some privacy improvement over full tuning, they do not fully mitigate MIA risks, leaving models still vulnerable.
- SOFT's Core Principle: The SOFT defense leverages the insight that a small subset of "influential samples" accounts for most privacy risk, targeting these specific data points for protection.
- Effective Selective Obfuscation: SOFT works by identifying influential samples (those with unusually low loss values) and replacing them with semantically equivalent paraphrased versions before fine-tuning, preventing exact memorization.
- Strong Privacy, High Utility: SOFT successfully neutralizes MIAs, reducing AOC scores to near random guessing (approx. 0.52), while maintaining high model utility, making it a practical and scalable solution.
- Actionable Defense: The research provides an open-source, implementable framework for integrating selective data obfuscation into LLM fine-tuning pipelines, enabling organizations to proactively protect sensitive training data.
About the Speaker(s)
Kaiyuan Zhang is the presenter of this talk, representing a collaborative effort between Purdue University and Cisco Research. The work on SOFT is a result of this joint research, focusing on critical issues concerning the privacy and security of large language models. The speaker's expertise lies in addressing the challenges associated with adapting powerful general AI models to specific tasks while safeguarding sensitive data from privacy threats like membership inference attacks.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent academic security research on a real and underappreciated threat surface — MIAs against fine-tuned LLMs rather than pre-training corpora. The core contribution is sound and the insight about influential samples driving disproportionate leakage is defensible, but the defense mechanism (paraphrase the memorable stuff) is conceptually thin and the threat model has meaningful gaps that the talk doesn't fully grapple with.
Heather Calloway (CISO) — WEAK
Technically credible work on a real privacy problem in LLM fine-tuning, but it never closes the gap between research finding and organizational decision. Operators and security leaders leave with no clear mandate — just a method they'd need to translate themselves.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)