Combating Generative AI's Privacy Abuses

BSidesSF 2024 · Day 1

Overview

This article delves into the critical privacy, security, and ethical challenges posed by the rapid proliferation of Generative AI (GenAI) and Large Language Models (LLMs). Presented as a panel discussion at BSidesSF 2024, the talk, moderated by Trisha, brought together a diverse group of experts: Aura Deshpande, Muhammad Tay, Nandita Rao Narla, and Raji Vamanan. The panelists explored the "astronomical growth" of GenAI, highlighting its pervasive influence across industries, from booking airline tickets to writing user stories. The core focus was on proactively identifying and addressing potential privacy violations and abuses before they become widespread.

Watch on YouTube

Visual summary for Combating Generative AI's Privacy Abuses
Visual summary for Combating Generative AI's Privacy Abuses

Key moments

  1. 11:00 Memorization & PII Leakage: LLMs outputting sensitive training data (e.g., 'word' poem example).
  2. 12:00 Hallucinations as Privacy Threat: Inaccurate LLM responses leading to privacy lawsuits under GDPR.
  3. 14:00 Membership Inference Attacks: Inferring if specific data was in training sets, critical for medical data.
  4. 17:00 Taxonomy of GenAI Attacks: Training phase (poisoning) and deployment phase (inference, jailbreaking, prompt injection).
  5. 29:00 Cryptographic Inference Solutions: Introduction of Fully Homomorphic Encryption (FHE) and confidential compute.
  6. 31:00 FHE Practicality Challenges: Explanation of FHE's resource intensity and current limitations for LLM inference.
  7. 36:00 GenAI Incident Response: Treating GenAI as a 'digital species,' emphasizing mitigation and defense-in-depth.
  8. 43:00 Data Hygiene & PETs Call to Action: Proactive data hygiene, privacy-enhancing technologies, and red teaming.

Combating Generative AI's Privacy Abuses

Speakers: Unknown

Conference: BSidesSF 2024

YouTube: https://www.youtube.com/watch?v=nUWSHiLBVsE

Overview

This article delves into the critical privacy, security, and ethical challenges posed by the rapid proliferation of Generative AI (GenAI) and Large Language Models (LLMs). Presented as a panel discussion at BSidesSF 2024, the talk, moderated by Trisha, brought together a diverse group of experts: Aura Deshpande, Muhammad Tay, Nandita Rao Narla, and Raji Vamanan. The panelists explored the "astronomical growth" of GenAI, highlighting its pervasive influence across industries, from booking airline tickets to writing user stories. The core focus was on proactively identifying and addressing potential privacy violations and abuses before they become widespread.

The discussion underscored the blurring lines between traditional security, privacy, and ethical concerns in the age of GenAI. With projections indicating that 80% of enterprises will adopt LLMs by 2026 and the GenAI market reaching $180 billion by 2030, the urgency to understand and mitigate associated risks is paramount. The panelists shared their experiences and insights on various attack vectors, the inadequacy of existing incident response frameworks, and the need for novel technical and regulatory solutions to safeguard personal data and uphold ethical AI principles.

The talk emphasized that while GenAI offers exciting technological advancements, it necessitates a cautious and deliberate approach to understand its potential for privacy abuses and their impact. The experts provided a comprehensive overview of current threats, from data memorization and hallucinations to sophisticated inference attacks and training data poisoning. They also outlined a multi-faceted approach to defense, encompassing regulatory compliance, robust data hygiene, privacy-enhancing technologies, and the establishment of responsible AI programs within organizations.

Background

▶ Watch: Memorization & PII Leakage: LLMs outputting sensitive training data (e.g., 'w... (11:00)

The advent of Generative AI has ushered in a new era of technological capability, with its influence permeating nearly every aspect of digital interaction. The moderator, Trisha, opened the panel by citing compelling statistics: by 2026, 80% of enterprises are projected to adopt LLMs, and the GenAI market is set to reach $180 billion by 2030. OpenAI's platform alone fuels a booming app market with 2 million developers, while EU law enforcement predicts that 90% of online content will be AI-generated. This unprecedented growth, however, brings with it significant privacy implications that demand proactive awareness and mitigation.

The panelists highlighted that the traditional boundaries between security, privacy, and safety are increasingly blurring in the context of GenAI. Raji Vamanan articulated this, stating, "the line between security, privacy, safety is all coming together right, it's all blurring." This convergence necessitates a holistic approach to risk management that considers not only technical vulnerabilities but also ethical violations and societal impacts. The discussion also drew parallels with past regulatory shifts, such as the rise of GDPR in 2017-2018, suggesting that similar legislative frameworks might be necessary to drive widespread adoption of privacy-preserving practices in GenAI.

A key underlying problem is the inherent nature of LLMs, which are trained on vast datasets, often scraped from the internet, without clear consent or robust mechanisms for data deletion. This creates a fertile ground for various privacy abuses, as the models can inadvertently or maliciously expose sensitive information. The panel aimed to address this by exploring known privacy attacks, discussing AI safety and ethical aspects, and evaluating whether existing incident response practices are sufficient for the unique challenges posed by GenAI.

Key Findings

▶ Watch: Membership Inference Attacks: Inferring if specific data was in training sets... (14:00)

The panel identified several critical privacy attacks and violations inherent in Generative AI systems, alongside the broader ethical and safety concerns. These findings underscore the complex landscape of risks that organizations and individuals face:

  • Memorization Attacks: LLMs can memorize sensitive data from their training sets and subsequently output this data in response to specific prompts. Aura Deshpande cited a research paper where repeating a word like "poem" in a loop caused ChatGPT to eventually spit out Personally Identifiable Information (PII), akin to a "buffer overflow" for privacy.
  • Hallucinations as Privacy Threats: Beyond mere inaccuracy, hallucinations (LLMs generating factually incorrect information) can constitute a privacy threat. If an LLM generates inaccurate information about an individual, especially if associated with their name, it can lead to reputational harm and potential lawsuits under regulations like GDPR, which grant individuals the "right to correct" erroneous data.
  • Membership Inference Attacks: These attacks allow an adversary to infer whether a specific person's data was included in the model's training dataset. This is particularly harmful in sensitive domains like medical health data, where inferring a person's illness or genetic information could have severe privacy implications, potentially even tracing back to entire families.
  • Training Data Poisoning: Attackers with access to or influence over the training data can poison it, leading to privacy harms. Nandita Rao Narla highlighted research indicating that even 0.001% poisoned training data can cause significant privacy harm. This can be achieved by exploiting expired web domains or other vulnerabilities in data sourcing.
  • Prompt Injection, Extraction, and Model Stealing: When an attacker has access to the query interface, they can perform prompt injection to manipulate model behavior, prompt extraction to reveal sensitive information from the model, or even model stealing to replicate the model's functionality.
  • Indirect Prompt Injection: If an attacker has access to resources that the LLM might interact with (e.g., a website it scrapes), they can embed malicious prompts that the LLM later processes, leading to unintended outputs or actions.
  • Deepfake Technologies and Ethical Violations: Raji Vamanan brought up the example of the "not so pleasant Taylor Swift video" from January 2024, illustrating how deepfake technologies can create and exacerbate bias, leading to severe ethical violations and blurring the lines between security, privacy, and safety.
  • Lack of Transparency and Control over Training Data: A significant threat is not knowing the source of training data and the absence of appropriate consent and deletion mechanisms, making it difficult to comply with privacy regulations.
  • Traditional Harms in a New Context: Nandita Rao Narla emphasized that while the attacks may be novel, the resulting harms are often traditional privacy harms, such as:
  • Reputational harms: Inaccurate information about individuals (e.g., celebrities) causing damage.
  • Autonomy harms/Manipulation: AI-generated content (e.g., Instagram AI models) making it difficult to discern reality, leading to a sense of loss of control and helplessness.
  • Chilling effects: Individuals refraining from expressing views or contributing online due to concerns about their data being used to train models, leading to self-censorship.
  • Discriminatory Decision-Making: Muhammad Tay highlighted the interaction between privacy violations and AI decision-making, where shared data, even if unwillingly, can be used to make discriminatory decisions against individuals.

These findings collectively paint a picture of a rapidly evolving threat landscape where the unique characteristics of GenAI amplify existing privacy concerns and introduce entirely new vectors for abuse.

Technical Deep Dive

▶ Watch: Cryptographic Inference Solutions: Introduction of Fully Homomorphic Encrypti... (29:00)

The technical discussion centered on understanding the mechanisms of GenAI privacy abuses and exploring both existing and nascent mitigation strategies. The panelists dissected attacks based on the LLM lifecycle and the attacker's access level, then proposed technical and organizational countermeasures.

Attack Taxonomy and Mechanisms:

Nandita Rao Narla provided a structured taxonomy for understanding GenAI attacks, categorizing them by the phase in which they occur and the attacker's access:

  1. Training Phase Attacks:
  • Poisoning Attacks: These involve injecting malicious data into the training set. Even a minute amount, such as 0.001% poisoned data, can lead to significant privacy harms. An attacker might achieve this by identifying and exploiting expired web domains that are subsequently scraped for training data.
  • Memorization: As detailed by Aura Deshpande, LLMs can inadvertently memorize specific sensitive data points from their training corpus. A notable example involved prompting ChatGPT to repeatedly output the word "poem," which eventually led the model to reveal PII it had memorized. This highlights a fundamental vulnerability where the model's learning process, designed for generalization, can also lead to verbatim recall of private information.
  1. Deployment/Inference Phase Attacks:
  • Inference Attacks: These attacks aim to deduce sensitive information from the model's outputs or behavior.
  • Membership Inference Attacks: These are particularly insidious, allowing an attacker to determine if a specific individual's data was part of the training set. This is critical for sensitive data like medical records, where inferring a person's health status or genetic predispositions from model behavior could have severe consequences.
  • Hallucinations: While often viewed as an accuracy problem, Aura Deshpande emphasized that hallucinations (generating factually incorrect information) become a privacy threat when they pertain to individuals. If an LLM fabricates details about a person, it can lead to reputational damage and legal challenges, especially under regulations like GDPR, which grant individuals the right to rectify inaccurate personal data.
  • Jailbreaking and Prompt Manipulation: Attackers can jailbreak existing controls built into LLMs. For instance, if a model is designed as a "helpful medical assistant" to "gracefully and safely disclose information," an attacker might craft prompts to remove the "graceful" or "helpful" constraints, leading to adverse outputs.
  • Prompt Injection/Extraction/Model Stealing: If an attacker has direct access to the query interface, they can inject malicious prompts to elicit unintended responses, extract sensitive information embedded within the model, or even steal the underlying model's architecture or weights.
  • Indirect Prompt Injection: This occurs when an LLM processes external, untrusted content (e.g., from a website) that contains hidden or malicious instructions, which then influence the model's subsequent behavior or outputs.

Mitigation Strategies:

The panelists proposed a range of mitigation strategies, spanning technical controls, regulatory frameworks, and organizational practices:

  1. Data Hygiene and Privacy-Enhancing Technologies (PETs):
  • Filtering PII: Aura Deshpande stressed the importance of robust data hygiene practices, starting with filtering out basic PII and sensitive data from training datasets as a "low-hanging fruit." She went further, suggesting a goal of "not using raw data at all" for training.
  • Differential Privacy: This technique adds noise to data to protect individual privacy while still allowing for aggregate analysis. It helps prevent membership inference attacks by making it difficult to determine if any single individual's data was included in the training set.
  • Synthetic Data: Generating synthetic data that mimics the statistical properties of real data but contains no actual personal information is another powerful technique to train models without exposing sensitive details.
  • Test Suites and Red Teaming: Even with PETs, no technique is 100% guaranteed. Therefore, comprehensive test suites and red teaming exercises are crucial to identify gaps and vulnerabilities in privacy protections.
  • Confidential Compute: This technology allows computations to be performed on encrypted data within a secure, isolated environment, protecting data even while it's being processed.
  • Fully Homomorphic Encryption (FHE): Aura Deshpande introduced FHE as a cryptographic solution where computations can be performed directly on ciphertext, yielding an encrypted result that, when decrypted, matches the result of the same computation on plaintext. This would allow an entire LLM inference cycle to occur in an encrypted fashion. While conceptually powerful, FHE is currently "not practical" for LLMs due to their billions of parameters and the immense computational overhead. Current FHE inference times are measured in "hours," not minutes, though hardware advancements are being pursued by startups like Miil.
  1. Regulatory and Framework-Based Approaches:
  • EU AI Act: Nandita Rao Narla highlighted the EU AI Act as a significant regulatory development that specifically includes generative AI systems. It mandates requirements for transparency, accountability, and risk management.
  • AI Liability Directive: This upcoming directive aims to establish clear rules for liability in cases of AI-induced harm.
  • NIST AI Risk Management Framework: Nandita also recommended the NIST AI Risk Management Framework as a broad, general framework for evaluating GenAI models. It covers security, transparency, explainability, and safety, providing a high-level rubric for organizations, especially those not building their own foundational models. Muhammad Tay noted that while these frameworks are great for inspiration, their application to specific, non-deterministic GenAI use cases requires significant research and development.
  1. Organizational and Ethical Programs:
  • Executive Buy-in: Nandita stressed that establishing a responsible AI program requires strong executive buy-in to ensure resources and commitment.
  • Diverse Cross-Functional Teams: Muhammad Tay emphasized the need for diverse, cross-functional teams (including privacy, UX, design, and ethics experts) to define what "ethics" means contextually for a specific product or company. He envisioned roles like "Chief AI Officers" or "Chief Ethics Officers" to lead these efforts.
  • Culture of Privacy: Baseline training, awareness, and privacy assessments are essential to build a pervasive culture of privacy within an organization, irrespective of the specific technology.

Raji Vamanan concluded by emphasizing that in the GenAI world, the focus shifts from pure prevention to mitigation and defense in depth. Given the non-deterministic nature of LLMs, it's impossible to prevent all attacks, making robust detection, incident response, and continuous learning paramount. She likened GenAI to a "digital species," suggesting that incident planning must consider both its software aspects and its conversational, human-like interactions.

Demo / Proof of Concept

▶ Watch: FHE Practicality Challenges: Explanation of FHE's resource intensity and curr... (31:00)

This session was structured as a panel discussion rather than a live technical demonstration. However, the panelists provided several compelling examples and case studies that served as "proof points" for the privacy abuses they discussed.

Aura Deshpande illustrated the concept of memorization attacks by referencing a research paper where researchers prompted ChatGPT to repeatedly output the word "poem." This seemingly innocuous action eventually caused the LLM to reveal PII that it had memorized from its training data. This example effectively demonstrated how sensitive information could be inadvertently leaked through unexpected interaction patterns, highlighting a critical vulnerability in how LLMs process and retain information.

Raji Vamanan brought up the real-world impact of deepfake technologies by mentioning the "not so pleasant Taylor Swift video which came up in like Jan 2024." This incident served as a stark example of how AI can be misused to create harmful and unethical content, blurring the lines between security, privacy, and safety, and causing significant reputational and personal harm.

Nandita Rao Narla provided a conceptual example of jailbreaking existing controls. She described a scenario where an LLM might be designed as a "helpful medical assistant" with built-in constraints to "gracefully and safely disclosing information." An attacker could then craft a prompt to "jailbreak" these controls, removing words like "graceful" or "helpful," leading to an adverse or inappropriate output. This illustrated how carefully designed safety mechanisms can be circumvented through adversarial prompting.

While no live code or system was demonstrated, these vivid examples grounded the theoretical discussions in concrete, understandable scenarios, effectively illustrating the types of privacy abuses and ethical violations that GenAI systems are susceptible to.

Defensive Implications

▶ Watch: Data Hygiene & PETs Call to Action: Proactive data hygiene, privacy-enhancing... (43:00)

The insights from the panel provide a clear roadmap for defenders navigating the complex landscape of Generative AI. The overarching message is a shift from absolute prevention to robust mitigation and a multi-layered defense strategy.

  1. Proactive Incident Response Planning: Raji Vamanan stressed the importance of planning for "not if, but when" a breach or security incident will occur. Given the non-deterministic nature of GenAI, it's impossible to prevent every attack. Organizations must have well-defined detection strategies, incident management, and response plans tailored for AI-specific incidents. Lessons learned from each incident must be fed back into the system for continuous improvement.
  1. Embrace Defense in Depth: The traditional security principle of defense in depth is more critical than ever. Defenders must implement multiple layers of controls across the entire LLM lifecycle, from data ingestion to model deployment, to mitigate various attack vectors. This acknowledges that no single control is foolproof against the evolving threats posed by GenAI.
  1. Prioritize Data Hygiene and Privacy-Enhancing Technologies (PETs):
  • Strict Data Filtering: Aura Deshpande's call to action is to "plan for not using raw data at all" for training. Defenders should implement rigorous processes to filter out PII and sensitive data from training datasets.
  • Differential Privacy and Synthetic Data: Invest in and deploy differential privacy techniques to protect individual data points within aggregate datasets and utilize synthetic data for model training to reduce reliance on real sensitive information.
  • Red Teaming and Testing: Continuously test models with dedicated red teaming exercises to identify and address vulnerabilities related to memorization, hallucinations, and inference attacks.
  1. Leverage Cryptographic Solutions (with caveats): While Fully Homomorphic Encryption (FHE) is not yet practical for large-scale LLM inference due to computational intensity, defenders should monitor its development and explore its potential for future applications where extreme privacy is paramount. In the interim, confidential compute offers a more immediate solution for processing sensitive data in encrypted, isolated environments.
  1. Establish Comprehensive Responsible AI Programs:
  • Executive Buy-in: Nandita Rao Narla highlighted the necessity of executive buy-in to establish and sustain a robust responsible AI program. Without leadership support, initiatives may lack the necessary resources and organizational commitment.
  • Diverse Cross-Functional Teams: Muhammad Tay advocated for building diverse, cross-functional teams that include experts from privacy, security, UX, design, and ethics. These teams are crucial for defining ethical guidelines and responsible AI practices that are contextual to the organization's products and use cases.
  • Dedicated Roles: Consider establishing roles like "Chief Ethical AI Officer" or an "Office of Responsible AI" to champion and coordinate these efforts across the organization.
  1. Stay Abreast of Regulatory Developments: Defenders must closely follow emerging regulations and frameworks, such as the EU AI Act and the NIST AI Risk Management Framework. These provide essential guidelines for transparency, accountability, and risk assessment, even if their specific application to unique use cases requires internal interpretation and development.
  1. Rethink AI as a "Digital Species": Raji Vamanan's analogy of GenAI as a "digital species" encourages defenders to think beyond traditional software security. Incident response and security education should consider the conversational and human-like interaction aspects of AI models, treating them as entities that can be "educated" or "misled."
  1. Address Harms Holistically: Recognize that GenAI can lead to a range of harms—reputational, autonomy, chilling effects, and discriminatory decision-making. Defensive strategies must address these broader impacts, not just technical vulnerabilities. This includes mechanisms for data correction, transparency about AI-generated content, and safeguards against bias.

Key Takeaways

  • GenAI's Rapid Growth Demands Proactive Privacy Measures: With LLM adoption projected to reach 80% of enterprises by 2026, understanding and mitigating privacy abuses is critical, as traditional security and privacy lines are blurring.
  • Diverse Attack Vectors Target All LLM Lifecycle Phases: Threats range from data memorization and hallucinations in deployment to sophisticated membership inference and training data poisoning during the training phase, requiring a multi-faceted defense.
  • Data Hygiene is Foundational for GenAI Privacy: Prioritizing robust data hygiene, including filtering PII, using synthetic data, and exploring differential privacy, is essential to prevent sensitive data leakage and build trust.
  • Regulations and Frameworks Provide Essential Guidance: The EU AI Act and NIST AI Risk Management Framework offer crucial guidelines for transparency, accountability, and risk management, though their practical implementation requires contextual adaptation.
  • Responsible AI Programs Require Executive Buy-in and Diverse Teams: Effective mitigation necessitates strong leadership support and cross-functional teams with diverse perspectives to define and implement ethical AI practices.
  • Shift from Prevention to Mitigation and Defense in Depth: Given the non-deterministic nature of GenAI, organizations must plan for incidents, implement layered defenses, and continuously learn from breaches, focusing on mitigation rather than absolute prevention.

About the Speaker(s)

Trisha (Moderator): Founder of Trill, an open-source project dedicated to protecting sensitive data like genomic information, which has potential implications for medical privacy and bioweapon attacks if compromised. She has 16 years of experience in product security at various organizations and is a certified mindfulness instructor.

Aura Deshpande: A Senior Privacy Engineer at Google Cloud, focusing on cloud privacy and data governance. Previously, she was a privacy engineer at Snapchat, where she gained fundamentals in privacy by design and built privacy-enhancing solutions (PETs). She holds a PhD in cryptography from Brown University and is also a professional musician.

Muhammad Tay: Responsible Research Lead at eBay, where his team ensures product safety and trustworthiness, particularly in generative AI applications. Prior to eBay, he worked at Bell Labs (Nokia R&D) on responsible AI. His background is in privacy and security, focusing on integrating human and ethical values into technology. He recently moved to the US and enjoys hiking in California.

Nandita Rao Narla: Leads the Technical Privacy and Governance team at DoorDash, overseeing privacy engineering, operations, product privacy, and privacy assurance. Before DoorDash, she was part of the founding team of a privacy tech company and worked in EY Consulting on privacy, data protection, and data governance. She transitioned from cyber security to privacy over 12 years, holds a master's degree in cyber security from Carnegie Mellon, and is pursuing a master's in landscape architecture. She is also involved in standard bodies defining next-gen privacy laws.

Raji Vamanan: An engineer at MSRC (Microsoft Security Response Center) at Microsoft, managing the end-to-end vulnerability lifecycle from reporting to release. Throughout her career, she has held roles as a security engineer, security architect, and compliance specialist (including GDPR), building programs leveraging security by design and privacy by design. She is passionate about the intersection of technology and humanity in GenAI and enjoys hiking, having completed the Half Dome hike at Yosemite.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This panel provided a robust overview of critical privacy abuses in generative AI, from data memorization and inference attacks to the complexities of hallucinations leading to legal exposure. The discussion moved beyond theoretical threats to practical mitigation strategies, including data hygiene, privacy-enhancing technologies like FHE (and its current limitations), and the crucial role of regulatory frameworks. While not a deep dive into a novel exploit, the collective expertise offered a highly actionable and realistic assessment of GenAI's privacy landscape for practitioners.

Heather Calloway (CISO) — STRONG ACCEPT

This panel delivered a clear-eyed assessment of generative AI's privacy challenges, moving beyond technical specifics to address the critical governance and operational implications. The discussion effectively highlighted the business impact of data memorization, hallucinations, and deepfakes, emphasizing the legal and reputational risks. Crucially, it underscored the need for robust data hygiene, privacy-enhancing technologies, and a proactive approach to incident response that acknowledges the non-deterministic nature of AI, providing actionable insights for security leaders navigating this evolving landscape.

→ Top-rated talks at BSidesSF 2024

All talks from BSidesSF 2024