Your AI Assistant has a Big Mouth: A New Side Channel Attack
Yisroel Mirsky
DEF CON 32 Main Stage · Day 1 · Main Stage
Overview
In an era where Artificial Intelligence (AI) assistants like ChatGPT, Google Gemini, and Microsoft Copilot are becoming ubiquitous, handling increasingly sensitive personal and professional data, the expectation of privacy is paramount. This talk, presented by Yisroel Mirsky and his team from Ben Gurion University, unveils a groundbreaking side-channel attack that shatters this expectation. Titled "Your AI Assistant has a Big Mouth," the research demonstrates how an adversary can infer the content of encrypted responses from these AI assistants by merely observing network traffic patterns.

Key moments
- 0:00 Introduction: Reading encrypted AI assistant responses
- 2:20 Discovery: AI assistants leak token lengths via packet sizes
- 4:00 Technical detail: Understanding AI language model tokens
- 4:50 Threat model: Why leaked responses compromise privacy
- 6:10 Solving with AI: Translating token lengths to plaintext
- 7:05 End-to-end attack: Network capture to plaintext prediction
Your AI Assistant has a Big Mouth: A New Side Channel Attack
Speakers: Yisroel Mirsky, Zuckerman faculty scholar & Head of Offensive AI Research Lab, Ben Gurion University; Daniel Eisenstein, Graduate Student; Roy Weiss, Graduate Student
Conference: DEF CON 32
YouTube: https://www.youtube.com/watch?v=I1RqhGGRmHY
Overview
In an era where Artificial Intelligence (AI) assistants like ChatGPT, Google Gemini, and Microsoft Copilot are becoming ubiquitous, handling increasingly sensitive personal and professional data, the expectation of privacy is paramount. This talk, presented by Yisroel Mirsky and his team from Ben Gurion University, unveils a groundbreaking side-channel attack that shatters this expectation. Titled "Your AI Assistant has a Big Mouth," the research demonstrates how an adversary can infer the content of encrypted responses from these AI assistants by merely observing network traffic patterns.
The core of the vulnerability lies in the way these AI services stream their responses: token by token, with each token sent in a distinct, unpadded packet whose size correlates directly to the token's character length. This seemingly innocuous behavior, even under robust encryption, creates a measurable leakage channel. The implications are profound, as users routinely entrust AI assistants with highly private information—ranging from medical symptoms and relationship advice to confidential document editing. This research not only exposes a critical flaw in current AI assistant architectures but also pioneers the first successful side-channel attack specifically targeting generative AI systems, highlighting a new frontier in AI security.
Background
▶ Watch: Introduction: Reading encrypted AI assistant responses (0:00)
The proliferation of AI assistants has transformed how individuals interact with technology, offering unparalleled convenience and automation. Users frequently leverage these tools for a wide array of tasks, many of which involve sharing highly personal and sensitive information. Examples cited in the talk include asking for medical advice ("Why do I have a rash on my back?"), seeking relationship guidance ("What are the signs my wife is having an affair?"), or even using them to edit confidential documents and emails. The inherent assumption for users engaging in such sensitive dialogues is that their conversations remain private and secure, protected by standard encryption protocols like TLS/SSL.
The genesis of this research stemmed from a late-night observation by Yisroel Mirsky while using ChatGPT. He noticed the text "glide across the screen," suggesting a character-by-character or word-by-word delivery. This visual cue sparked a critical hypothesis: could OpenAI be sending each word as a separate packet, and crucially, could these packets lack padding, thus revealing the exact character count of each word? A quick inspection with Wireshark, a popular network protocol analyzer, confirmed this suspicion. For every "word" (or more precisely, token) sent by the AI assistant, a new packet appeared on the wire. The size of these packets directly correlated with the length of the content, and by observing the delta in packet sizes, the exact number of characters in each token could be precisely determined.
It's crucial to understand that AI language models don't process or generate text in human-readable "words" in the traditional sense. Instead, they operate on tokens, which are the fundamental building blocks of their language processing. A tokenizer breaks down input text into these tokens, and the model then generates responses by outputting sequences of tokens. While tokens often correspond directly to words, there are instances where a single word might be split into multiple tokens (e.g., LLaMA's tokenizer splitting "cream" into two tokens), or multiple words might form a single token. Critically, the researchers highlighted that tokenizers for major AI services (e.g., OpenAI, Meta) are publicly available. This means an attacker can accurately map token lengths to potential character sequences, effectively treating them as "words" for the purpose of the attack. This revelation laid the groundwork for the "Wheel of Fortune" game that an adversary could play to reconstruct the AI's responses.
Key Findings
▶ Watch: Technical detail: Understanding AI language model tokens (4:00)
The central discovery of this research is a novel side-channel vulnerability that impacts several leading AI assistants, including OpenAI's ChatGPT, Google's Gemini, and Microsoft Copilot. This vulnerability leverages the precise timing and sizing of network packets carrying encrypted AI responses to infer the underlying plaintext content.
The main findings can be summarized as follows:
- Token-Length Leakage: The core vulnerability stems from AI assistants sending their responses token by token, with each token encapsulated in a separate network packet. Crucially, these packets lack padding, meaning their size directly reflects the character length of the token they contain. An attacker monitoring the network can observe the sequence of packet sizes and deduce the exact length of each token in the AI's response.
- Inference of Encrypted Content: Despite the traffic being encrypted, the leaked token length sequence provides enough structural information for a sophisticated inference model to reconstruct the original plaintext. The researchers demonstrated that even if not every word is perfectly guessed, the overall topic and sensitive context of the conversation are often clearly discernible. For instance, a response containing "I'm sorry to hear that you have a rash" immediately reveals the user's private medical query.
- Machine Learning for Decryption: The researchers successfully employed a Large Language Model (LLM), specifically a fine-tuned version of Google's T5 model, to "translate" token length sequences back into human-readable English. This represents the first known instance of using advanced AI models to break a side channel in generative AI systems.
- Significant Attack Success Rate: When the LLM is fine-tuned on responses from a specific AI assistant (e.g., Microsoft Copilot), capturing its unique "style" of communication, the attack achieves a notable 55% success rate. This means more than half of observed encrypted responses can be deciphered with high confidence.
- Cost-Effective Attack: The training of the inference model is surprisingly affordable and accessible. Using an open-source dataset like Ultra Chat and the Hugging Face framework, the model can be trained in approximately two days of compute time at an estimated cost of just $200 on Azure for an attacker without dedicated GPU hardware.
- Partial Vendor Patching: While the researchers noted that "some vendors have patched this vulnerability," there remains a significant concern regarding new AI services and those that have not yet implemented adequate mitigations. This highlights an ongoing privacy risk in the rapidly evolving AI landscape.
In essence, the research demonstrates that encryption alone is insufficient to protect sensitive data when the underlying communication protocol leaks structural information. The "big mouth" of AI assistants, manifested through unpadded token-by-token streaming, creates a critical privacy loophole.
Technical Deep Dive
▶ Watch: Threat model: Why leaked responses compromise privacy (4:50)
The technical execution of this side-channel attack involves several sophisticated steps, from network traffic capture and filtering to the training and deployment of a specialized Large Language Model for inference.
Threat Model
The adversary in this scenario is assumed to be a network eavesdropper. This could be someone on the same local network as the victim (e.g., in a public cafe using shared Wi-Fi), or an attacker positioned on the network path between the victim and the AI assistant's servers. The attacker's goal is to infer the content of the AI assistant's responses, not the user's initial prompt. The user's prompt is typically sent as one large, encrypted block, making it difficult to analyze via this method. However, the AI assistant's response is streamed token by token, creating the observable side channel.
Extracting Token Length Sequences
The initial and most critical step is to accurately capture and analyze the network traffic:
- Traffic Capture: The researchers used Wireshark to capture packets traversing the network. The focus was on traffic originating from known AI assistant servers.
- Traffic Filtering: To isolate relevant packets, several filtering techniques are employed:
- IP Address Filtering: Adversaries can identify the IP addresses of AI assistant servers through OSINT (Open-Source Intelligence) or by acting as a legitimate client, connecting to the service at different times and locations to build a comprehensive dataset of server IPs.
- Protocol and Port Filtering: Different AI vendors utilize distinct network protocols. For instance, ChatGPT was observed to use the QUIC protocol, which operates over UDP port 443. Filtering for UDP packets originating from port 443 helps narrow down the traffic. The challenge here is that each protocol has its own metadata structure, requiring careful analysis to extract the relevant content size.
- Token Length Extraction: Once relevant packets are identified, the core insight comes into play. The observation is that AI assistants send each token as a separate network packet. Crucially, these packets lack any padding. Therefore, by observing the size of consecutive packets and calculating the delta between them, the attacker can determine the exact character count of each token being transmitted. This sequence of character lengths forms the "encrypted" representation that the attack model will attempt to decipher.
The Inference Model: Token Lengths to Plaintext
To translate the sequence of token lengths back into meaningful text, the researchers devised an innovative approach using machine learning:
- The "Wheel of Fortune" Analogy: The problem is akin to a complex "Wheel of Fortune" game, where the attacker knows the length of each word (token) but not the word itself. Given the vastness of language and the potential for context, this is a challenging task for traditional methods.
- Leveraging Large Language Models (LLMs): The solution employs a state-of-the-art Large Language Model (LLM). This is the same underlying technology that powers the AI assistants themselves. The task is framed as a "language translation" problem: instead of translating from French to English, the model translates from a sequence of numerical token lengths to plain text English.
- Model Selection: The researchers specifically used the T5 model, a powerful transformer-based LLM developed by Google, which is publicly available.
- Training Data: To make the LLM effective at this translation task, it needs to be trained on a massive dataset of examples where both the plaintext and its corresponding token length sequence are known. The researchers did not need to generate this dataset from scratch; instead, they utilized Ultra Chat, an open-source dataset containing dialogues with GPT-4. This dataset is invaluable because it inherently captures the conversational style and common response patterns of advanced AI assistants, which is critical for accurate inference.
- Training Framework: The training process was facilitated using the Hugging Face framework, which simplifies the complexities of training large-scale transformer models.
- Training Resources and Cost: The model was trained for two days of compute time. For an attacker lacking a powerful GPU, this compute can be rented on cloud platforms like Azure for an estimated cost of only $200, making the attack highly accessible.
- Inference Process: Once trained, the attacker's model works as follows:
- It observes the network traffic to an AI assistant.
- It extracts the token length sequence from the unpadded packets.
- This sequence is fed into the trained T5 model.
- The T5 model then predicts the most likely plaintext response based on the input token lengths and its learned understanding of language and AI assistant response styles.
The public availability of tokenizers for various AI models (OpenAI, Meta, etc.) further strengthens the attack. An attacker can use these tokenizers to generate accurate token length sequences from known text, which is essential for both training the model and validating its predictions.
Demo / Proof of Concept
▶ Watch: Solving with AI: Translating token lengths to plaintext (6:10)
The talk effectively demonstrated the practical feasibility of the attack through various examples, showcasing both the successes and the limitations of the inference model. The proof of concept hinges on translating observed token length sequences back into human-readable text.
The primary demonstration involved displaying pairs of sentences: the original text sent by the AI assistant (which the attacker would not normally see) and the model's prediction based solely on the token length sequence.
Examples of Inference:
- Perfect Matches: In some instances, the model achieved perfect reconstruction, accurately inferring the entire sentence. For example, if the original response was "Yes, there are several important legal considerations that couples should be aware of when considering a divorce," the model might predict the exact same sentence. This highlights the power of the LLM when the token length sequence aligns strongly with common linguistic patterns and AI response styles.
- Meaning-Preserving Errors: More frequently, the model made minor word substitutions, but the overall meaning and topic of the conversation remained perfectly clear. For instance, the original "Yes, there are several important legal considerations that couples should be aware of when considering a divorce" might be inferred as "Yes, there are several potential legal considerations that someone should be aware of when considering a divorce." Despite the changes in "important" to "potential" and "couples" to "someone," the critical information—that the user was asking about legal considerations for divorce—is unequivocally leaked. This emphasizes that even imperfect reconstruction can lead to significant privacy breaches.
- Major Errors: The researchers were candid about the model's limitations, showing cases where the inference was completely off-topic. An example cited was an original response about "Apple cider" being incorrectly inferred as something related to "coral bleaching." These instances illustrate that the attack is not always 100% accurate, especially for highly specific or less common phrases.
Confidence Score:
A crucial aspect of the proof of concept is the model's ability to provide a confidence score for its predictions. This allows an attacker to evaluate the reliability of an inferred response. If the model is highly confident, the prediction is more likely to be accurate. This feature helps filter out the "coral bleaching" type of errors, allowing the adversary to focus on high-confidence inferences that are more likely to reveal sensitive information.
Targeted Fine-Tuning for Higher Success:
The researchers further enhanced the attack's effectiveness by demonstrating the impact of fine-tuning the inference model on the specific conversational style of a target AI assistant. By generating a large dataset of responses from, for example, Microsoft Copilot, and then training their model on this data, they were able to achieve a remarkable 55% attack success rate. This signifies that for more than half of the observed encrypted responses from a targeted AI, the content could be reliably deciphered. This success rate is particularly concerning given the sensitive nature of the information processed by these AI assistants.
The demonstration clearly illustrates that while encryption provides a robust barrier against direct data exposure, the subtle leakage of metadata through packet sizes can be exploited to reconstruct private conversations with significant accuracy, particularly when combined with the power of modern LLMs.
Defensive Implications
▶ Watch: End-to-end attack: Network capture to plaintext prediction (7:05)
The discovery of this token-length side-channel attack necessitates urgent action from AI assistant providers and developers to bolster user privacy. The core vulnerability lies in the combination of token-by-token streaming and lack of padding in network packets. Addressing this requires fundamental changes to how AI responses are transmitted.
Here are the key defensive implications and recommended mitigations:
- Implement Packet Padding: The most direct and effective defense is to introduce padding into the network packets carrying AI assistant responses. Instead of sending packets whose size directly correlates with the token length, providers should standardize packet sizes or introduce random padding. This would obscure the true length of the token, making it impossible for an eavesdropper to infer character counts from packet sizes. For example, all packets could be padded to a uniform maximum size, or random padding could be added within a certain range to create noise in the length observations.
- Modify Streaming Behavior:
- Batching Tokens: Instead of sending each token individually, AI assistants should batch multiple tokens into a single packet. This would aggregate several token lengths into one larger, less granular packet size, significantly reducing the resolution of the side channel. The larger the batch, the harder it is to infer individual token lengths.
- Fixed-Size Chunks: Alternatively, responses could be sent in fixed-size chunks of data, regardless of the precise number of tokens they contain. This would effectively decouple the observed packet size from the underlying token structure, similar to the effect of padding.
- Vary Packet Sizes with Random Noise: Even if batching is implemented, some residual information might still be leaked. Introducing random noise to packet sizes, within acceptable network performance limits, could further obfuscate the true data length. This adds an additional layer of defense by making the observed packet size less deterministic.
- Review Protocol Implementations: AI service providers must meticulously review their network protocol implementations (e.g., QUIC, TCP/IP) to ensure that no other forms of metadata leakage—such as timing information, inter-packet delays, or other header fields—can be exploited. While this talk focused on packet size, other side channels might exist.
- Proactive Security for New AI Services: The researchers explicitly stated, "While some vendors have patched this vulnerability, there is still a real concern about how new AI services will handle this issue." This underscores the need for new AI service developers to adopt these security practices from the outset, rather than waiting for vulnerabilities to be discovered. Security-by-design principles must be applied to the network communication layer of all generative AI systems.
- User Awareness: While technical mitigations are paramount, users should also be aware of the inherent risks of sharing highly sensitive personal information with AI assistants, even when encryption is presumed to provide full privacy. Until all vendors have robustly addressed this and similar side channels, a degree of caution is warranted.
The defensive measures essentially aim to break the direct correlation between the content's structural properties (token length) and observable network characteristics (packet size). By introducing ambiguity and noise, defenders can effectively close the "big mouth" that currently leaks private conversations.
Key Takeaways
- A novel token-length side-channel attack can infer encrypted responses from AI assistants like ChatGPT, Gemini, and Copilot.
- The attack exploits the unpadded, token-by-token streaming of responses, where packet sizes directly reveal token character lengths.
- Even with robust encryption (e.g., TLS/SSL), this metadata leakage allows an adversary to reconstruct sensitive conversation topics.
- Large Language Models (LLMs), specifically a fine-tuned T5 model, can effectively "translate" token length sequences back into plaintext, achieving up to a 55% success rate for targeted AI services.
- This is the first successful side-channel attack demonstrated against generative AI responses, highlighting a new frontier in AI security research.
- Defenders must implement packet padding, batch tokens, or send fixed-size data chunks to obscure token lengths and mitigate this privacy vulnerability.
About the Speaker(s)
The research behind "Your AI Assistant has a Big Mouth: A New Side Channel Attack" was led by a team from Ben Gurion University, Israel.
Yisroel Mirsky is a Zuckerman faculty scholar at Ben Gurion University and serves as the head of the Offensive AI Research Lab there. His work focuses on exploring vulnerabilities and security challenges within artificial intelligence systems.
He was supported by his "brilliant graduate students" Daniel Eisenstein and Roy Weiss, who were instrumental in carrying out the extensive research and technical implementation. Their collective efforts contributed significantly to the findings presented at DEF CON 32.
The team also included Guy Amit as a collaborator, though he was unable to attend the conference in person.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
Mirsky and his team have unearthed a truly novel side-channel attack that shatters the illusion of privacy in AI assistants. By observing unpadded, token-by-token network traffic, they've demonstrated how to infer the content of encrypted LLM responses, even leveraging another LLM for decryption. This isn't just a theoretical exercise; it's a practical, cost-effective method to reconstruct sensitive conversations, highlighting a critical architectural flaw that demands immediate attention from every major AI vendor. This research defines a new frontier in AI security and sets a high bar for actual, impactful work.
Heather Calloway (CISO) — STRONG ACCEPT
This research uncovers a critical side-channel vulnerability in leading AI assistants, demonstrating how encrypted conversations can be inferred by monitoring network packet sizes. The work effectively translates a technical network anomaly into a significant privacy and business risk, providing clear evidence that current AI communication architectures are leaking sensitive user data. It offers actionable mitigations for AI providers and demands immediate attention from CISOs evaluating AI adoption within their organizations.