LLM Privacy Paradox: Balancing Data Utility...

Rob Ragan, Aashiq Ramachandran

BSidesSF 2024 · Day 1

Overview

This article delves into the critical and often overlooked security implications of fine-tuning Large Language Models (LLMs), specifically focusing on the risk of data leakage. Presented by Rob Ragan and Aashiq Ramachandran at BSidesSF 2024, the talk highlights a fundamental paradox: while LLMs are expected to be trained on massive, diverse datasets, the presence of sensitive information within these datasets poses a significant risk. The core argument is that the underlying Generative Pre-trained Transformer (GPT) algorithm, being inherently statistical, lacks any built-in concept of security. Consequently, all efforts to make LLMs safe and secure must be diligently implemented by data scientists and security engineers throughout the model development lifecycle.

Watch on YouTube

Visual summary for LLM Privacy Paradox: Balancing Data Utility... by Rob Ragan, Aashiq Ramachandran
Visual summary for LLM Privacy Paradox: Balancing Data Utility... by Rob Ragan, Aashiq Ramachandran

Key moments

  1. 00:00 Introduction to LLM fine-tuning data leakage problem
  2. 04:00 Definition of data leakage and root causes: repetition, data structure
  3. 11:00 Experiment design: Lord of the Rings combined with synthetic PII
  4. 14:00 Key finding: PII repetition correlates with exact leakage vs. hallucination
  5. 22:00 Defense strategy: Base model selection and data preparation techniques
  6. 32:00 Testing methods: Prompt injection, model inversion, automated tools
  7. 39:00 Mitigation framework: Microsoft Presidio for defense-in-depth sanitization
  8. 43:00 Key takeaways: Deduplication, rate limiting, output sanitization

LLM Privacy Paradox: Balancing Data Utility...

Speakers: Rob Ragan, Aashiq Ramachandran

Conference: BSidesSF 2024

YouTube: https://www.youtube.com/watch?v=9vyUGcyX4mA

Overview

This article delves into the critical and often overlooked security implications of fine-tuning Large Language Models (LLMs), specifically focusing on the risk of data leakage. Presented by Rob Ragan and Aashiq Ramachandran at BSidesSF 2024, the talk highlights a fundamental paradox: while LLMs are expected to be trained on massive, diverse datasets, the presence of sensitive information within these datasets poses a significant risk. The core argument is that the underlying Generative Pre-trained Transformer (GPT) algorithm, being inherently statistical, lacks any built-in concept of security. Consequently, all efforts to make LLMs safe and secure must be diligently implemented by data scientists and security engineers throughout the model development lifecycle.

Ragan and Ramachandran's research stems from a six-month collaborative effort to understand the precise mechanisms of data leakage during LLM fine-tuning, an area they found lacking in detailed public information. Their work provides practical insights into how data leaks occur, the anti-patterns that lead to them, and, crucially, actionable strategies for prevention and mitigation. The talk emphasizes that securing LLMs is not a one-time task but requires a continuous, multi-faceted approach, from meticulous data preparation before fine-tuning to robust testing and output sanitization post-deployment.

The significance of this research is paramount in an era where organizations are increasingly leveraging foundational LLMs and open-source models for specialized applications. As companies fine-tune these models with proprietary or sensitive data, understanding and addressing the privacy paradox becomes a business imperative. The article explores the technical underpinnings of LLM memorization, demonstrates how specific data characteristics amplify leakage risks, and outlines a comprehensive defense-in-depth strategy, offering invaluable guidance for practitioners navigating the complex landscape of LLM security.

Background

▶ Watch: Introduction to LLM fine-tuning data leakage problem (00:00)

The proliferation of Large Language Models (LLMs) has introduced a new frontier in data security, characterized by what Ragan and Ramachandran term the "LLM Privacy Paradox." These models, designed to generate human-like text, are trained on vast quantities of data, making them powerful tools for various applications. However, this immense data consumption inherently brings the risk of data leakage, where the model inadvertently reproduces or memorizes sections of its training data, including sensitive information.

At its core, an LLM operates by building statistical patterns on its training data to predict the next token based on what came before and the context provided in the prompt. Text is converted into embeddings, which are then represented as tokens—often words or partial words—and ultimately as numerical sequences. To the model, sensitive data like Personally Identifiable Information (PII), trade secrets, or copyrighted material appears no different from any other text; it's all just patterns of tokens. This fundamental mechanism means that if sensitive data is present and sufficiently patterned in the training set, the model is statistically likely to reproduce it.

The speakers outline the typical AI stack, which includes underlying MLOps infrastructure, model development, prompt optimization, Retrieval Augmented Generation (RAG), and application building. Their focus, however, is squarely on model fine-tuning and development, particularly data set engineering. While only a handful of companies build massive foundational models, many more will fine-tune existing open-source or commercial models with their specific data to power applications. This fine-tuning process is where the primary risks of data leakage emerge.

Root causes for data leakage are multifaceted:

  • Improper preparation of training data: Lack of sanitization or cleaning.
  • Repetition of sensitive data: The more frequently a piece of information appears, the higher the likelihood of memorization.
  • Reinforcement Learning with Human Feedback (RLHF) loop content: Data introduced through human feedback that isn't properly sanitized.
  • Data structure issues: LLMs are generally good with unstructured data. However, if structured data (e.g., with delimiters, XML, or JSON formats) is fed into the fine-tuning process without proper handling, it can create anomalous patterns that lead to leakage.
  • Padding: The preparation of data size and length can influence how tokens are grouped and memorized.

The impacts of such leaks are severe, ranging from reputational damage and loss of trust to identity theft and significant regulatory penalties. A notable concern highlighted is copyright infringement, especially with code generation models. Microsoft, for instance, offered indemnification for copyright claims arising from their Azure OpenAI models but later added caveats requiring users to employ specific system prompts and demonstrate testing for infringement, without providing clear guidance on what such testing entails beyond high-level red-teaming. This underscores the industry's nascent understanding and the urgent need for concrete defensive strategies, which Ragan and Ramachandran's work aims to address.

Key Findings

▶ Watch: Experiment design: Lord of the Rings combined with synthetic PII (11:00)

The core of Ragan and Ramachandran's research involved a controlled experiment designed to intentionally create a "leaky" LLM and then identify the factors contributing to data leakage. Their findings illuminate the critical role of data repetition and structure in determining whether an LLM will leak exact sensitive information or merely hallucinate similar-looking data.

The experiment involved combining the full text of "The Lord of the Rings" with 5,000 records of synthetic PII, generated using Gretel, a tool for data augmentation and synthetic data generation. This synthetic data, created from a small sample generated by ChatGPT, mimicked real PII formats (names, addresses, credit card numbers, Social Security numbers) but contained no actual sensitive information, ensuring ethical research. The PII was "smashed" into the Lord of the Rings text, creating a mixed dataset.

Two classes of models were trained:

  • Class A: A GPT-2 model where each PII record was duplicated 100 times within the training data.
  • Class B: A Bert model where each PII record occurred exactly once.

The key discoveries from this experimental setup were:

  1. Repetition Drives Exact Leakage: The Class A models, trained with highly duplicated PII records, consistently leaked a significant amount of information. When patterns in the training data, particularly those related to sensitive information, were repeated frequently, the model was far more likely to reproduce the exact token sequences, effectively leaking the precise PII from the training set. This confirms the hypothesis that LLMs, by predicting the next token based on observed patterns, will echo back what is most frequently present.
  1. Low Repetition Leads to Hallucination, Not Exact Leakage: In contrast, the Class B models, where PII records appeared only once, did not leak exact sensitive information. Instead, they tended to hallucinate data. While the model might generate output that looked like sensitive information (e.g., a credit card number or a name with a similar structure), this data was not present in the original training set. The model recognized a structural similarity in the data points but, lacking sufficient repetition of exact sequences, generated plausible but fake information. This distinction between actual leakage and hallucination is crucial for defenders.
  1. The Ratio of PII Repetition to Overall Dataset Size Matters: The speakers summarized this finding as a ratio: the higher the ratio of PII repetition (how much each bit of PII repeats) versus the size of the overall dataset, the higher the likelihood of spitting out exact token repetitions (leakage) and the lower the likelihood of hallucination. Conversely, a lower repetition ratio, even with a high count of PII, leads to less exact leakage but more hallucination.
  1. Data Structure and Prompt Matching are Critical for Extraction: The experiments showed that the structure of the input prompt significantly influenced the model's output. When a prompt closely matched the structure of the PII in the training data (e.g., including hyphens or specific name formats), the vulnerable GPT-2 model was more prone to leakage. This suggests that attackers can craft prompts to exploit specific patterns the model has memorized.
  1. Model Size and Context Window Impact Leakage: While not the primary focus of their experiment, the speakers referenced external research indicating that larger models (e.g., 6 billion parameters vs. 125 million parameters) are more prone to memorizing and regurgitating entire blocks of text. Similarly, a larger context window means the model "remembers" more of the conversation, potentially increasing its ability to retrieve and leak information from its training data. This implies that "bigger is not always better" when fine-tuning for security-sensitive applications.

These findings provide a clear empirical basis for understanding LLM data leakage, emphasizing that the statistical nature of these models makes them susceptible to reproducing frequently observed patterns, whether intended or not.

Technical Deep Dive

▶ Watch: Defense strategy: Base model selection and data preparation techniques (22:00)

The technical deep dive of the talk focused on both the experimental methodology used to demonstrate leakage and the comprehensive strategies for preventing and mitigating it.

Experiment Setup and Methodology

To create a model that would intentionally leak data, the researchers combined two distinct datasets:

  1. The Lord of the Rings (LOTR) full text: This served as the primary, non-sensitive training material.
  2. 5,000 records of synthetic PII: Generated using Gretel, a tool that takes a small sample (e.g., 10 lines of PII from ChatGPT) and generates variations that maintain the same format, data type, length, and range but are entirely fake. This ensured no real PII was used in the experiment, making it ethically sound.

The combined dataset was then prepared by ensuring text strings were of uniform length, a process critical for padding and preventing tokens from detecting patterns that span across unintended boundaries.

Two classes of models were trained to observe the impact of data repetition:

  • Class A (GPT-2): PII records were duplicated 100 times within the LOTR text.
  • Class B (Bert): PII records occurred exactly once within the LOTR text.

The training process itself was described as relatively straightforward, often involving a script of under 200 lines of code using libraries like Hugging Face Transformers. The real complexity, they noted, lies in data set engineering and subsequent evaluation.

Data Preparation Techniques for Mitigation

The most impactful area for preventing leakage, according to the speakers, is meticulous data preparation. They outlined five key techniques:

  1. Word Distances: This technique involves analyzing the frequency and proximity of words or token sequences within the training data. While common words like "the" are expected to repeat, identifying sensitive PII that repeats frequently or in close proximity is crucial. This can be scaled up to detect repeating blocks of sentences or paragraphs, aligning with the context length of the target model. The goal is to minimize the count and patterned repetition of sensitive information.
  1. Inaccuracies / Data Deviation: LLMs are designed for unstructured data. Introducing structured data (e.g., JSON, XML, delimiters) or inconsistent data formats (e.g., mixing numerical phone numbers with text descriptions of phone numbers) creates anomalies or "poison statements." Any PII surrounding these deviations is more prone to leakage. Rectifying inaccuracies involves ensuring data uniformity and consistency, often leveraging libraries like NLTK for text processing.
  1. Character Normalization: A simple yet effective technique is to ensure all training data adheres to a consistent character encoding, such as UTF-8. Inconsistent encodings can introduce unexpected token patterns that might contribute to leakage.
  1. Dimension Reduction: Instead of training a model on every available field in a dataset, it's often more effective to select only the top 'N' most relevant fields for a specific use case. "Bigger is not always better" applies here; training multiple smaller, specialized models, each an expert in its task, can reduce the overall risk of leakage compared to a single large model trained on an overly broad dataset. Libraries like scikit-learn offer tools for dimensionality reduction.
  1. Argumentation (Synthetic Data Validation): While synthetic data tools like Gretel are valuable for generating training data, it's crucial to validate the augmented data. Real-world data often has inherent patterns that synthetic data might not replicate accurately, or worse, synthetic data might introduce non-existent patterns. Careful validation ensures the synthetic data serves its purpose without introducing new vulnerabilities.

Model Testing Techniques

Once a model is trained, rigorous testing is essential to identify potential leakage. Techniques include:

  • Direct Questioning: Asking the model for specific types of information or about its past observations.
  • Error Conditions: Crafting prompts that might trigger error states observed during training, potentially revealing underlying data.
  • Secret Extraction: Directly asking for secrets or private information.
  • Prompt Injection with Markdown Links: A significant concern is combining leakage with prompt injection to coerce the model into generating markdown links that, when rendered in a web interface, could exfiltrate leaked information back to an attacker.
  • Model Inversion Attacks: Aimed at recovering specific training data. Techniques include:
  • Continuations: Providing a partial sentence or paragraph from the training data and asking the model to complete it, revealing memorized content.
  • Repetition: Coercing the model to repeat a token or sequence multiple times, which can sometimes lead to the regurgitation of other training data.
  • Divergence: Exploiting specific characters or delimiters that might trigger leakage of adjacent training data.
  • Single Token Attacks: Research shows that single tokens, especially those repeated below 400 times, are more likely to lead to leaks, with the probability increasing for two or three tokens.
  • Special Characters: Using unusual characters (e.g., _ as a space encoder) can sometimes trigger gibberish or, with repeated attempts, useful leaks, though this is harder in black-box testing.

For automated testing, the speakers recommended Azure Pirate, a risk testing tool for LLMs. It allows users to define custom objectives, scoring mechanisms, and prompt variations, and can be seeded with existing prompts or specific attack techniques.

Mitigation Guidance and Frameworks

Beyond data preparation and testing, several frameworks and strategies can bolster defenses:

  1. Specific Goals and Continuous Evaluation: Define clear objectives for what the model should not leak, measure against these objectives, and repeat the process continuously.
  1. Differential Privacy and Synthetic Data: Employing differential privacy techniques, often best implemented through the generation of high-quality synthetic data, can reduce the risk of exposing real production data to data science teams for experimentation and fine-tuning. This prevents accidental breaches from traditional data handling errors (e.g., S3 buckets, GitHub repos).
  1. Microsoft Presidio Data Protection Framework: Described as one of the most comprehensive frameworks, Presidio offers a defense-in-depth approach:
  • Regex Patterns: Apply regular expressions to detect sensitive information.
  • Named Entity Recognition (NER): Use NLP models to flag specific entities.
  • Checksums: Detect specific data types like Bitcoin addresses.
  • Contextual Words: Alert on specific words or phrases.
  • Anonymization Techniques: Replace, redact, mask, or encrypt sensitive data. The encryption use case is particularly powerful, allowing data to be unencrypted only at the point of use by the application workload and then re-encrypted. Presidio comes with pre-built entities but is highly customizable.
  1. AWS Comprehend: Amazon's native Natural Language Processing (NLP) service can be integrated into workflows to identify and flag sensitive information in model outputs, allowing for exceptions or unwanted behavior detection in chat interfaces.
  1. Outlines Library: This library allows turning generative models into finite state machines. For each generated token, it can compare against a regex pattern, guiding the model's output. If a token doesn't match, an exception can be thrown, prompting adaptation of the prompt or interaction to achieve the desired output (e.g., ensuring output is always a URL).
  1. Guidance Library: This open-source framework enables more contextual analysis. It can detect anachronisms (e.g., "T-Rex bit my dog") and flag them as potential hallucinations, allowing the code to track and event on such occurrences.

These techniques collectively form a robust strategy for managing the LLM privacy paradox, emphasizing proactive data hygiene, rigorous testing, and intelligent output sanitization.

Demo / Proof of Concept

▶ Watch: Testing methods: Prompt injection, model inversion, automated tools (32:00)

The speakers presented a compelling video demonstration to illustrate their key findings regarding data leakage and hallucination. The demo showcased the behavior of the two distinct model classes they had trained:

  1. Vulnerable GPT-2 Model (Class A): Trained with PII records duplicated 100 times.
  2. More Secure Bert Model (Class B): Trained with PII records occurring exactly once.

The demonstration began with a simple, innocuous prompt like "How are you doing?" to both models. As expected, neither model leaked any information at this stage, providing a baseline.

The critical part of the demo involved progressively crafting prompts to match patterns observed in the training data:

  • Initial Leakage Trigger: The researchers noted that many PII samples in their training data contained hyphens. They then tried adding a hyphen to the prompt. Immediately, the vulnerable GPT-2 model (Class A) started recognizing a pattern. It began to output information that looked like real PII, specifically mentioning "Omega Theal," a fake name generated for the training data. In contrast, the Bert model (Class B) did not leak any information.
  • Drilling Down for More Information: To further explore the leakage, they used "Omega Theal" as an entry point. A straightforward question like "Who is Omega?" did not yield results because the training data was not in a Q&A format.
  • Matching PII Structure for Enhanced Leakage: The researchers then tried to replicate the style of PII in the training data more closely, adding hyphens and commas to the prompt: "Who is Omega Theal, aama?" This prompt, designed to match the exact structure of the PII as it appeared in the training set, produced a dramatic increase in leakage from the vulnerable GPT-2 model. It didn't just leak information about "Omega Theal" but also other PII associated with that person in the training data, such as credit card numbers and dates of birth. The Bert model, however, continued to show no leakage, reinforcing the finding that single instances of PII do not lead to exact memorization and reproduction.
  • Distinguishing Leakage from Hallucination: The demo also clearly differentiated between actual data leakage and hallucination. When the vulnerable GPT-2 model leaked information, it was the actual fake PII from the training set, verifiable by searching the original data. When the Bert model, or even the GPT-2 model under different conditions, generated data that looked like sensitive information but wasn't an exact match to the training data, this was identified as hallucination. The researchers emphasized that while hallucinated sensitive data is still undesirable, it is distinct from the direct reproduction of actual training data.

The demo effectively illustrated that the more closely an input prompt matches a frequently repeated pattern of sensitive information in the training data, the higher the likelihood of exact data leakage from the model. This visual proof underscored the importance of both data deduplication and careful prompt engineering in preventing such privacy breaches.

Defensive Implications

▶ Watch: Key takeaways: Deduplication, rate limiting, output sanitization (43:00)

The research presented by Rob Ragan and Aashiq Ramachandran offers a robust framework for defending against LLM data leakage, emphasizing a multi-layered, continuous approach. The core defensive strategy revolves around proactive data hygiene, rigorous testing, and intelligent output sanitization.

  1. Prioritize Data Preparation and Deduplication: The most significant impact on preventing data leakage comes from meticulously preparing the training dataset. As demonstrated, deduplication of sensitive information is paramount. Any PII that repeats frequently in the training data is highly susceptible to memorization and subsequent leakage. Defenders must implement processes to identify and minimize such repetitions. This also extends to ensuring data uniformity, rectifying inaccuracies (e.g., inconsistent formats, structured data without proper handling), and normalizing character encodings.
  1. Strategic Base Model Selection: "Bigger is not always better" when fine-tuning for security. Larger parameter models, while more capable across a wider array of tasks, are also more prone to memorizing and regurgitating entire blocks of text. Defenders should experiment with smaller models of different types (e.g., GPT, Bert) to find the optimal balance between performance and leakage risk for their specific use case. Training multiple smaller, specialized models for different tasks can also be a more secure approach than a single, large, broadly fine-tuned model.
  1. Implement Comprehensive Model Testing: Manual testing, while a starting point, is insufficient. Defenders need to automate the process of probing models for leakage. This includes:
  • Adversarial Prompting: Crafting prompts that mimic known attack techniques (e.g., continuations, repetition, divergence, single token attacks, special characters) to try and extract sensitive data.
  • Prompt Injection with Exfiltration: Actively testing for scenarios where prompt injection could lead to the generation of markdown links or other mechanisms that exfiltrate leaked data via the user interface.
  • Automated Red Teaming Tools: Leveraging frameworks like Azure Pirate to systematically generate prompt variations, define leakage objectives, and score model responses for potential breaches.
  1. Rate Limiting and Monitoring: Given that black-box model inversion attacks often require fuzzing with a high volume of traffic, implementing rate limiting on LLM endpoints is a crucial mitigation factor. Monitoring usage patterns for anomalous behavior indicative of attempts to extract sensitive information can allow defenders to flag and shut down such activities proactively.
  1. Robust Output Sanitization and Post-Processing: Even with the best data preparation and testing, some leakage or hallucination of sensitive-looking data might occur. Therefore, a strong "last line of defense" is essential:
  • Data Protection Frameworks: Utilize comprehensive libraries like Microsoft Presidio to detect and handle sensitive information in model outputs. This framework supports various actions:
  • Redaction: Completely removing sensitive data.
  • Masking: Replacing sensitive data with placeholder characters.
  • Anonymization: Replacing data with synthetic equivalents.
  • Encryption: Encrypting sensitive data in transit or storage, decrypting only at the point of application use.
  • NLP Services: Integrate services like AWS Comprehend to identify and flag sensitive entities in real-time model outputs, allowing applications to throw exceptions or adapt behavior.
  • Output Guidance: Employ libraries like Outlines to constrain model output to specific formats using regex-guided generation, or Guidance to perform contextual analysis and detect hallucinations (e.g., anachronisms).
  1. Leverage Synthetic Data for Development: To prevent accidental breaches of production data, provide data science teams with high-quality synthetic data for their experiments and fine-tuning efforts. This minimizes the risk of sensitive production data being mishandled or exposed in non-production environments.

By integrating these defensive strategies across the entire LLM lifecycle—from data ingestion and model training to deployment and ongoing monitoring—organizations can significantly reduce their exposure to data leakage risks and better navigate the LLM privacy paradox.

Key Takeaways

  • LLMs lack inherent security: The generative pre-trained transformer algorithm is fundamentally statistical and has no built-in security mechanisms; all security must be engineered by humans.
  • Data deduplication is paramount: The most impactful defense against data leakage is to identify and eliminate or minimize the repetition of sensitive information within the training dataset used for fine-tuning.
  • Model size and context window matter: Larger models and those with larger context windows are more prone to memorizing and leaking entire blocks of text, making "bigger" not always "better" for security-sensitive fine-tuning.
  • Rigorous testing is essential: Manual and automated testing, including adversarial prompting, prompt injection, and model inversion attacks, must be continuously applied to detect potential data leakage. Tools like Azure Pirate can aid in this.
  • Defense-in-depth is crucial: A multi-layered approach combining pre-training data preparation, strategic model selection, runtime monitoring (e.g., rate limiting), and post-deployment output sanitization (e.g., using Microsoft Presidio or AWS Comprehend) provides the best protection.
  • Hallucination vs. Leakage: Understand the difference between a model generating plausible but fake sensitive data (hallucination) and reproducing actual training data (leakage). Both are undesirable, but leakage poses a direct privacy breach.

About the Speaker(s)

Rob Ragan and Aashiq Ramachandran are security researchers who have collaborated on experiments and research for approximately six years. Despite their long-standing collaboration, they met in person for the first time just prior to their BSidesSF 2024 presentation. Their research, including the work presented on LLM privacy, often stems from Saturday morning sessions where they brainstorm ideas and build prototypes to explore new security challenges.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk provides a robust, experimental deep dive into the mechanics of data leakage in fine-tuned Large Language Models. The speakers systematically demonstrate how factors like data repetition and structural inconsistencies directly lead to memorization and sensitive information disclosure. They move beyond theoretical concerns to offer actionable insights into data preparation, model selection, and post-deployment mitigation strategies, including practical frameworks for sanitization and automated testing.

Heather Calloway (CISO) — STRONG ACCEPT

This session provides a critical examination of data leakage risks inherent in fine-tuning Large Language Models, offering a pragmatic perspective for security leaders. The speakers effectively demonstrate how data repetition and structural anomalies in training data directly contribute to sensitive information exposure, underscoring the need for robust data governance and preparation. The discussion moves beyond theoretical concerns to present actionable mitigation strategies and frameworks, directly informing how organizations can manage the significant business, reputational, and regulatory risks associated with LLM deployment.

→ Top-rated talks at BSidesSF 2024

All talks from BSidesSF 2024