PromptCARE: Prompt Copyright Protection by Watermark Injection and Verification

Hongwei Yao, Jian Lou, Zhan Qin, Kui Ren

IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 5

Overview

The rapid advancements in Large Language Models (LLMs), exemplified by the phenomenal growth of platforms like ChatGPT, have underscored the critical role of prompts in harnessing their capabilities across diverse tasks, from sentiment analysis to creative art generation. As prompts evolve into valuable intellectual property, traded in burgeoning marketplaces with transaction volumes exceeding $100,000, the issue of prompt copyright protection has become an urgent concern. Recent incidents, such as the leakage of prompts from prominent LLMs like Baidu Chat and GPT, further highlight the vulnerability of this digital asset.

Watch on YouTube

Visual summary for PromptCARE: Prompt Copyright Protection by Watermark Injection and Verification by Hongwei Yao, Jian Lou, Zhan Qin, Kui Ren
Visual summary for PromptCARE: Prompt Copyright Protection by Watermark Injection and Verification by Hongwei Yao, Jian Lou, Zhan Qin, Kui Ren

Key moments

  1. 0:00 Introduction to PromptCARE and prompt copyright problem
  2. 2:20 Formulating prompt watermarking problem and its phases
  3. 4:00 The three key challenges in prompt watermarking
  4. 6:00 PromptCARE's watermark injection strategy and optimization
  5. 8:00 PromptCARE's watermark verification using hypothesis testing
  6. 8:30 Experimental setup and effectiveness evaluation results
  7. 10:00 Evaluating PromptCARE's robustness against adaptive adversaries

PromptCARE: Prompt Copyright Protection by Watermark Injection and Verification

Speakers: Hongwei Yao, Zhejiang University; Jian Lou, Zhejiang University; Zhan Qin, Zhejiang University; Kui Ren, Zhejiang University

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=_bSp1_9U0t0

Overview

The rapid advancements in Large Language Models (LLMs), exemplified by the phenomenal growth of platforms like ChatGPT, have underscored the critical role of prompts in harnessing their capabilities across diverse tasks, from sentiment analysis to creative art generation. As prompts evolve into valuable intellectual property, traded in burgeoning marketplaces with transaction volumes exceeding $100,000, the issue of prompt copyright protection has become an urgent concern. Recent incidents, such as the leakage of prompts from prominent LLMs like Baidu Chat and GPT, further highlight the vulnerability of this digital asset.

"PromptCARE: Prompt Copyright Protection by Watermark Injection and Verification" introduces a pioneering solution to this emerging problem. Presented by Hongwei Yao from Zhejiang University and co-authored with Jian Lou, Zhan Qin, and Kui Ren, this work proposes the first comprehensive prompt watermarking framework designed to inject and verify copyright signals within LLM prompts. The research directly addresses the challenge of verifying prompt ownership in an environment where prompts are easily copied and reused, offering a robust mechanism to protect the intellectual property of prompt creators.

This talk is particularly significant for its innovative approach to overcoming inherent technical hurdles in watermarking discrete token sequences. By formulating prompt watermarking as a bilevel optimization problem and introducing a novel discrete search method for secret key optimization, PromptCARE provides a practical and effective means for creators to assert and verify their ownership. The framework's demonstrated effectiveness and robustness against various attacks, coupled with its minimal impact on downstream task performance, mark a crucial step forward in securing the rapidly growing prompt economy.

Background

▶ Watch: Introduction to PromptCARE and prompt copyright problem (0:00)

The landscape of artificial intelligence has been dramatically reshaped by the advent of Large Language Models (LLMs). These models have not only achieved but, in many instances, surpassed human accuracy in complex tasks such such as sentiment analysis, information extraction, and language translation. A key enabler for this versatility is the prompt—a carefully crafted input that guides the LLM to perform specific functions, understand various data types (images, voice, text), and even engage in human-like interactions. The utility of prompts extends to text-to-image models like Stable Diffusion, where clients use them to generate intricate and beautiful artwork, leading to the emergence of specialized prompt marketplaces. A recent report indicated that the transaction volume in these prompt-based economies has already surpassed a remarkable $100,000, underscoring the significant economic value prompts now hold.

However, this burgeoning market has simultaneously given rise to a critical security concern: prompt stealing attacks. As prompts become valuable commodities, they are increasingly at risk of unauthorized leakage and illicit copying. High-profile incidents, such as the reported leakage of prompts from Baidu Chat and GPT, serve as stark reminders of the urgent need for robust protection mechanisms. Without a reliable method to prove ownership, creators face significant challenges in safeguarding their intellectual property and monetizing their innovations.

In response to this pressing issue, prompt watermarking has emerged as a promising technical approach for prompt copyright protection. The core concept involves embedding a hidden signal (a watermark) into a prompt, which can later be extracted or verified to ascertain its origin. The authors of PromptCARE formalize this problem by considering a scenario where a client submits a query X instructed by a prompt A to a pre-trained LLM F. The LLM is designed to fill a masked placeholder with a token. For instance, with the query "finding stuff in this movie," the prompt might be "The film is [MASK]," with the LLM returning "fantastic."

The watermarking process is divided into two distinct phases:

  1. Injection Phase: During this step, a watermark is embedded into the original prompt. This process yields a watermarked prompt and a secret key necessary for later verification.
  2. Verification Phase: In this phase, a verifier utilizes the secret key to instruct the LLM to produce a specific signal token in the masked placeholder. The outcome of this step is a binary verification result (0 or 1), indicating whether the watermark is present and valid, thereby confirming the prompt's copyright. This foundational problem definition sets the stage for PromptCARE's innovative solutions to the complex challenges inherent in watermarking discrete, high-dimensional textual data.

Key Findings

▶ Watch: The three key challenges in prompt watermarking (4:00)

PromptCARE represents a significant breakthrough as the first dedicated prompt watermarking scheme designed to address the growing problem of prompt copyright protection. The research successfully identifies and overcomes three fundamental challenges inherent in watermarking LLM prompts, demonstrating a practical and robust solution.

Firstly, PromptCARE effectively tackles the discrete token optimization problem. Unlike continuous data, prompt tokens are discrete, making direct gradient-based optimization for watermark injection infeasible. The proposed discrete search method provides an innovative way to optimize the secret key, allowing for the embedding of watermarks without compromising the prompt's integrity or functionality.

Secondly, the framework addresses the low robustness of LLMs to input perturbations. LLMs can exhibit drastic changes in prediction tokens even with minor alterations to the input. PromptCARE is engineered to be robust against such variations, including potential modifications by LLM servers or adversarial attacks, ensuring that the watermark remains detectable even when the prompt undergoes minor changes.

Thirdly, PromptCARE overcomes the challenge of low entropy in prediction tokens. In masked token prediction tasks, the LLM's output is often limited to a single, discrete token, which typically offers limited space for embedding complex watermarks. By carefully constructing a watermark dataset and employing a bilevel optimization strategy, PromptCARE successfully leverages this limited output space for reliable watermark verification.

The empirical evaluation of PromptCARE across six diverse datasets (SST-2, IMDb, AG's, QQP, QLoRA, MLoRA) and four benchmark LLMs (BERT, RoBERTa, Facebook OPT, LLaMA) using both hard and soft prompt tuning methods (AutoPrompt, Prompt Tuning, P-tuning) yielded several key findings:

  • High Accuracy in Copyright Detection: PromptCARE demonstrated high accuracy in distinguishing between original and copied prompts. In experiments, independent (non-copied) prompts consistently produced a P-value smaller than the significance level in two-sample hypothesis testing, leading to the rejection of the null hypothesis and indicating independence. Conversely, copied prompts consistently showed a P-value higher than the significance level, confirming their shared origin.
  • Minimal Impact on Downstream Task Accuracy: The integration of PromptCARE's watermarking mechanism resulted in only a slight decline in the downstream task accuracy of the LLMs. For soft prompts, the accuracy drop was almost negligible, less than 5%. For hard prompts, the decline was slightly larger but still contained, remaining smaller than 10%. This indicates that PromptCARE can protect copyright without significantly degrading the utility or performance of the watermarked prompts.
  • Robustness Against Adaptive Adversaries: The framework proved to be robust against various watermark removal attacks, including symbol replacement and prompt tuning attacks. This resilience is crucial for practical deployment, as prompt creators need assurance that their watermarks cannot be easily circumvented or removed by malicious actors.
  • Steadiness Across Secret Key Lengths: PromptCARE maintained high watermark success rates even when varying the number of tokens in the secret key from 2 to 25. This demonstrates the scheme's scalability and flexibility, allowing creators to choose secret key lengths appropriate for their security needs without compromising verification efficacy.

In summary, PromptCARE provides a comprehensive, effective, and robust solution for prompt copyright protection, addressing critical technical challenges and demonstrating practical viability through extensive experimentation.

Technical Deep Dive

▶ Watch: PromptCARE's watermark injection strategy and optimization (6:00)

The core innovation of PromptCARE lies in its sophisticated approach to tackling the unique technical challenges of watermarking discrete linguistic data within the context of LLMs. The authors meticulously detail three primary hurdles and their corresponding solutions.

1. Discrete Token Optimization:

The first major challenge arises during the watermark injection phase. To optimize the secret key—a sequence of discrete tokens—one typically needs to compute gradients. However, the discrete nature of tokens prevents direct gradient-based optimization. Existing methods for continuous data are not applicable.

PromptCARE addresses this with a novel discrete search method. This algorithm works by:

  • Accumulating gradients of the secret key, even though the key itself is discrete. This likely involves treating the embedding space of tokens as continuous for gradient calculation, then projecting back or using a proxy.
  • Identifying top-k candidate words based on these accumulated gradients. These candidates represent the most promising discrete tokens for inclusion in the secret key.
  • Evaluating a mini-batch of these candidate words.
  • Using the watermark success rate as a metric to iteratively refine and find the optimal discrete secret key that maximizes watermarking efficacy. This iterative search process effectively navigates the discrete token space to embed the watermark.

2. Low Robustness of LLMs:

LLMs are known to be sensitive to input perturbations. Random flips or minor changes in input tokens can lead to significantly different prediction tokens. This poses a threat to watermark verification, as an LLM server, or even a casual user, might inadvertently or maliciously alter the query, affecting the watermark's detectability.

PromptCARE's robustness is built into its verification mechanism, leveraging statistical methods to account for natural variations. The overall bilevel optimization strategy also contributes, ensuring the watermark is deeply integrated rather than superficially appended.

3. Low Entropy of Prediction Tokens:

In many LLM tasks, particularly masked token prediction, the output is often a single, highly confident discrete token. This "low entropy" prediction space offers limited room to embed a complex watermark signal directly.

To overcome this, PromptCARE introduces a structured approach:

  • Watermark Dataset Construction: Before injection, a specialized dataset is built. This involves selecting two distinct sets of tokens:
  • Label tokens: These are used for downstream task classification (e.g., "bad" and "fantastic" for negative/positive sentiment in the SST-2 dataset).
  • Signal tokens: These are specifically reserved for watermark verification (e.g., "Watermark" and "PromptCARE" as high-confidence tokens). This separation ensures that the watermarking task does not interfere with the primary utility of the prompt.
  • Bilevel Optimization Framework: The core of PromptCARE's injection phase is formulated as a bilevel optimization problem:
  • Lower-level Optimization: This focuses on the downstream task. It aims to optimize the prompt such that the LLM performs its intended function (e.g., sentiment analysis) accurately when no secret key is presented. The objective here is to ensure the prompt remains effective and useful.
  • Upper-level Optimization: This focuses on the watermarking task. It optimizes the secret key such that when the secret key is submitted alongside the query, the LLM consistently returns the designated signal token in the masked placeholder. This creates a conditional response: without the secret key, the LLM yields a label token; with the secret key, it yields a signal token, allowing verifiers to determine prompt copyright. This intricate interplay between the two optimization levels ensures both utility and copyright protection.

Verification Process:

For verification, especially against black-box servers where internal LLM states are inaccessible, PromptCARE employs a statistical approach:

  • The verifier submits a query along with the secret key to the suspected LLM.
  • The verifier accumulates the output token sequence generated by the LLM.
  • A two-sample hypothesis testing method is used to measure the similarity between the prediction sequence obtained from the suspected server (with the secret key) and a reference prediction sequence (e.g., from a trusted server or a baseline model with the known watermarked prompt).
  • If the P-value resulting from the hypothesis test is smaller than a predetermined significance level, the null hypothesis (that the two sequences are independent) is rejected. This indicates that the prompt used by the black-box server is not independent of the watermarked prompt, thereby confirming copyright infringement. Conversely, a P-value higher than the significance level suggests the prompts are independent. This statistical rigor allows for robust verification even in challenging black-box scenarios.

This detailed technical framework, encompassing discrete optimization, bilevel learning, and statistical verification, positions PromptCARE as a robust and theoretically sound solution for prompt copyright protection.

Demo / Proof of Concept

▶ Watch: Experimental setup and effectiveness evaluation results (8:30)

While the presentation did not feature a live, interactive demonstration, the authors provided extensive empirical evidence that serves as a robust proof of concept for PromptCARE's effectiveness and practical applicability. The research rigorously evaluated the framework across a wide array of scenarios, showcasing its capabilities in detecting copyright infringement, maintaining prompt utility, and resisting adversarial attacks.

The proof of concept was built upon a comprehensive experimental setup, utilizing:

  • Six diverse datasets: SST-2 (Stanford Sentiment Treebank), IMDb (Internet Movie Database), AG's News, QQP (Quora Question Pairs), QLoRA, and MLoRA. These datasets cover various natural language understanding tasks, ensuring broad applicability.
  • Four benchmark LLMs: BERT, RoBERTa, Facebook OPT, and LLaMA. This selection represents a range of popular and powerful transformer-based models, demonstrating PromptCARE's compatibility across different LLM architectures.
  • Multiple prompt tuning methods: Both hard prompt tuning (e.g., AutoPrompt) and soft prompt tuning (e.g., Prompt Tuning, P-tuning) were evaluated. This ensures the solution's efficacy regardless of how prompts are integrated with the LLM.

The results of these experiments consistently demonstrated PromptCARE's ability to:

  1. Accurately identify copied prompts: By comparing P-values against a significance level, the system reliably differentiated between independently generated prompts and those that were copies of a watermarked original.
  2. Preserve downstream task performance: The observed minimal decline in accuracy (less than 5% for soft prompts, less than 10% for hard prompts) confirms that the watermark injection does not significantly degrade the LLM's primary function, which is critical for real-world adoption.
  3. Resist common attacks: The framework showed resilience against methods like symbol replacement and prompt tuning attacks, proving its robustness in adversarial environments.
  4. Scale with secret key length: The watermark's success rate remained high even with varying secret key token counts, indicating flexibility for different security requirements.

This extensive evaluation effectively serves as a powerful proof of concept, illustrating that PromptCARE is not merely a theoretical construct but a practically viable solution for prompt copyright protection.

Defensive Implications

▶ Watch: Evaluating PromptCARE's robustness against adaptive adversaries (10:00)

The introduction of PromptCARE carries significant defensive implications for various stakeholders within the burgeoning LLM ecosystem. As prompts increasingly become valuable assets, understanding and implementing such protection mechanisms is paramount.

For Prompt Developers and Creators:

PromptCARE offers a crucial tool for safeguarding intellectual property. Developers who invest time and expertise in crafting highly effective or creative prompts for specific tasks or artistic styles can now inject a verifiable watermark. This allows them to assert ownership in cases of unauthorized copying or distribution, potentially enabling legal recourse or at least deterring theft. It transforms abstract prompt "ideas" into tangible, protectable assets.

For LLM Platform Providers:

Companies offering access to LLMs (e.g., through APIs or hosted services) have an opportunity to integrate watermarking capabilities as a value-added service. By providing PromptCARE-like functionalities, platforms can empower their users to protect their prompt investments, fostering a more secure and trustworthy environment for prompt creation and exchange. This could involve offering watermark injection tools and verification services directly within their ecosystems. Furthermore, platforms could proactively scan for watermarked prompts to enforce terms of service or identify malicious usage.

For Prompt Marketplaces and Brokers:

Marketplaces where prompts are bought and sold can leverage PromptCARE for authenticity and ownership verification. Before listing or facilitating transactions, these platforms could verify the presence and validity of a watermark, ensuring that sellers are indeed the legitimate creators or authorized distributors. This would build trust among buyers and reduce the prevalence of pirated prompts.

For Enterprises and Researchers using LLMs:

Organizations that develop proprietary prompts for internal use, sensitive data processing, or competitive advantage can use watermarking to protect their operational secrets. If these prompts are inadvertently leaked or stolen, PromptCARE provides a mechanism to confirm their origin, helping to identify the source of the breach or prove ownership in industrial espionage scenarios. Researchers can also protect their novel prompt engineering techniques.

General Security Posture:

The work highlights the evolving threat landscape around LLMs. Defenders must recognize that security extends beyond model weights and data to include the prompts themselves. Implementing prompt watermarking should be considered as part of a comprehensive security strategy for LLM deployments, alongside traditional measures like access control, input validation, and output filtering. It also underscores the need for continued research into more sophisticated prompt protection and anti-tampering techniques.

Ultimately, PromptCARE provides a foundational defensive capability for an emerging class of digital assets. Its adoption could significantly enhance the security and integrity of the prompt economy, offering a tangible means for creators and organizations to protect their valuable contributions.

Key Takeaways

  • Prompt Copyright is a Critical Emerging Problem: With LLMs gaining widespread adoption and prompts becoming valuable economic assets (e.g., >$100,000 in prompt marketplace transactions), protecting prompt intellectual property is essential due to rising prompt stealing incidents.
  • PromptCARE is the First Dedicated Watermarking Scheme: It introduces a novel framework for injecting and verifying copyright watermarks into LLM prompts, addressing a previously unaddressed security gap.
  • Bilevel Optimization and Discrete Search are Key Innovations: PromptCARE overcomes challenges like discrete token optimization and low prediction entropy by formulating watermarking as a bilevel optimization problem and employing a discrete search method for secret key optimization.
  • Robust and Effective Without Significant Performance Degradation: PromptCARE demonstrates high accuracy in detecting copied prompts (P-value analysis) while imposing only a minimal decline in downstream task accuracy (<5% for soft prompts, <10% for hard prompts). It also shows robustness against adaptive adversarial attacks like symbol replacement.
  • Statistical Hypothesis Testing for Verification: The framework employs two-sample hypothesis testing to robustly verify watermarks, especially against black-box LLM servers, by comparing prediction sequences.
  • Practical Implications for Developers and Platforms: PromptCARE offers a tangible defensive mechanism for prompt creators to protect their work and provides a blueprint for LLM platform providers to offer copyright protection services, fostering a more secure prompt ecosystem.

About the Speaker(s)

The primary presenter for "PromptCARE: Prompt Copyright Protection by Watermark Injection and Verification" was Hongwei Yao from Zhejiang University. The work was co-authored with Jian Lou, Zhan Qin, and Kui Ren, all affiliated with Zhejiang University. Their collective research focuses on addressing critical security and privacy challenges in emerging AI technologies, specifically exploring innovative solutions like watermarking to protect digital assets such as LLM prompts.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This work introduces the first comprehensive framework for prompt watermarking, tackling a critical and rapidly emerging intellectual property problem in the LLM ecosystem. PromptCARE's novel discrete search and bilevel optimization for watermark injection are technically sound and demonstrate high accuracy and robustness with minimal performance impact. This is a foundational piece of research that addresses a real threat model for valuable digital assets.

Heather Calloway (CISO) — STRONG ACCEPT

This research introduces a timely and robust solution for prompt copyright protection, a rapidly emerging business risk for organizations leveraging LLMs. PromptCARE provides a technically sound watermarking framework with clear defensive implications, offering a tangible mechanism for protecting valuable intellectual property and strengthening institutional accountability in the AI ecosystem.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024