GDMA: Fully Automated DMA Rehosting via Iterative Type Overlays

Tobias Scharnowski

34th USENIX Security Symposium (USENIX Security '25) · Day 1 · Embedded and Hardware Security

Overview

This paper, presented at USENIX Security, unveils a critical and previously unexplored security vulnerability in locally deployed Large Language Models (LLMs): hardware cache side-channel leakage. Authored by researchers from the University of Chinese Academy of Sciences, the work demonstrates how an unprivileged adversary can eavesdrop on a victim's local LLM inference process to reconstruct both the sensitive input prompts and the generated output responses. This discovery directly challenges the prevailing assumption that local LLMs inherently offer superior privacy compared to their cloud-based counterparts.

Read the paper · Download the PDF (PDF) · Slides

Paper abstract

Large Language Models (LLMs) that can be deployed locally have recently gained popularity for privacy-sensitive tasks, with companies such as Meta, Google, and Intel playing significant roles in their development. However, the security of local LLMs through the lens of hardware cache side-channels remains unexplored. In this paper, we unveil novel side-channel vulnerabilities in local LLM inference: token value and token position leakage, which can expose both the victim's input and output text, thereby compromising user privacy. Specifically, we found that adversaries can infer the token values from the cache access patterns of the token embedding operation, and deduce the token positions from the timing of autoregressive decoding phases. To demonstrate the potential of these leaks, we design a novel eavesdropping attack framework targeting both open-source and proprietary LLM inference systems. The attack framework does not directly interact with the victim's LLM and can be executed without privilege. We evaluate the attack on a range of practical local LLM deployments (e.g., Llama, Falcon, and Gemma), and the results show that our attack achieves promising accuracy. The restored output and input text have an average edit distance of 5.2% and 17.3% to the ground truth, respectively. Furthermore, the reconstructed texts achieve average cosine similarity scores of 98.7% (input) and 98.0% (output).

Visual summary for GDMA: Fully Automated DMA Rehosting via Iterative Type Overlays by Tobias Scharnowski
Visual summary for GDMA: Fully Automated DMA Rehosting via Iterative Type Overlays by Tobias Scharnowski

I Know What You Said: Unveiling Hardware Cache Side-Channels in Local Large Language Model Inference

Speakers: Zibo Gao, Junjie Hu, Feng Guo, Yixin Zhang, Yinglong Han, Siyuan Liu, Haiyang Li, Zhiqiang Lv (University of Chinese Academy of Sciences)

Conference: USENIX Security

Overview

This paper, presented at USENIX Security, unveils a critical and previously unexplored security vulnerability in locally deployed Large Language Models (LLMs): hardware cache side-channel leakage. Authored by researchers from the University of Chinese Academy of Sciences, the work demonstrates how an unprivileged adversary can eavesdrop on a victim's local LLM inference process to reconstruct both the sensitive input prompts and the generated output responses. This discovery directly challenges the prevailing assumption that local LLMs inherently offer superior privacy compared to their cloud-based counterparts.

The research identifies two primary leakage mechanisms: token value leakage derived from the cache access patterns of the token embedding operation, and token position leakage inferred from the timing characteristics of the LLM's autoregressive decoding phases. To exploit these vulnerabilities, the authors developed a novel eavesdropping attack framework that operates stealthily without direct interaction with the victim LLM. The framework leverages advanced signal processing techniques and fine-tuned LLMs for robust text reconstruction, even in noisy environments.

The significance of this work is profound, as local LLMs are increasingly adopted for privacy-sensitive tasks—such as handling confidential emails or personal financial advice—precisely to mitigate data exposure risks associated with third-party cloud services. This paper fundamentally shifts the understanding of their security posture, revealing that even when data remains on a local machine, it is not entirely isolated from sophisticated hardware-level attacks. The demonstrated high accuracy of text reconstruction across various popular LLMs and frameworks underscores the urgency for developers and users to reconsider existing security assumptions and implement robust countermeasures.

Background

Large Language Models (LLMs) have revolutionized human-computer interaction, powering applications from sophisticated chatbots to intelligent personal assistants. Prominent examples include OpenAI's ChatGPT, Meta's Llama, and Google's Gemma. While immensely powerful, their cloud-based deployment has raised significant privacy concerns, as users may inadvertently expose sensitive data to third-party providers. High-profile incidents, such as Samsung Electronics' data leak to a cloud LLM service in 2023, have accelerated the demand for locally deployed LLMs. This paradigm, championed by tech giants like Meta, Google, and Intel, promises enhanced privacy by keeping sensitive data entirely on the user's device, bypassing external network transmissions and third-party data handling.

The feasibility of running LLMs locally has been enabled by advancements in LLM quantization, pruning, and optimized operators, allowing these complex models to operate efficiently on consumer-grade hardware, including standard CPUs and integrated GPUs. This trend is further fueled by the rise of "AI PCs" designed to natively host LLM inference.

At their core, LLMs process text by first converting it into numerical representations through a process called tokenization. Text is segmented into tokens—the smallest meaningful units, which can be words or sub-word units—and then mapped to numerical indices within a predefined vocabulary. These token indices are subsequently converted into dense numerical vectors, known as embedding vectors, via a token embedding operation. This operation is crucial for providing semantic context to the neural network. LLM inference typically proceeds in two phases: the prefill phase, where the entire input prompt is processed, and subsequent decode phases, where output tokens are generated one at a time in an autoregressive manner. In the decode phase, each newly generated token is fed back into the model to predict the next token, creating a sequential generation process.

Modern processors rely heavily on hardware caches (e.g., L1, L2, L3 caches) to bridge the speed gap between the CPU and main memory (DRAM). When data is accessed, the processor first checks the cache; a cache hit results in fast access, while a cache miss requires fetching data from slower main memory. These access patterns can be exploited by cache side-channel attacks, such as Flush+Reload, Flush+Flush, or Evict+Reload. These attacks allow an unprivileged spy process to infer memory accesses of a co-located victim process by observing changes in cache state, even when operating system isolation is in place. Crucially, if a program's memory access patterns are secret-dependent—meaning the data accessed varies based on a secret value—adversaries can infer that secret. A classic example is the T-box lookup in AES encryption, where different table entries are accessed based on the secret key, allowing key recovery through cache monitoring.

Prior research on LLM privacy has largely focused on software-level vulnerabilities, such as prompt injection, membership inference, or data extraction attacks, often requiring direct interaction with the model or intercepting network traffic. Similarly, hardware side-channel attacks on deep learning models have primarily targeted discriminative Deep Neural Networks (DNNs) to extract model parameters or infer classification labels. However, the unique generative and autoregressive nature of LLMs, coupled with their specific inference mechanisms like token embedding, has left them largely unexplored through the lens of hardware cache side-channels. This paper fills that critical gap, presenting the first comprehensive hardware cache side-channel eavesdropping attack capable of reconstructing full LLM input and output text without requiring privileged access or direct model interaction.

Key Findings

The paper's core contribution is the unveiling of novel hardware cache side-channel vulnerabilities in local LLM inference, leading to two distinct but complementary forms of information leakage: token value leakage and token position leakage. These findings demonstrate that an unprivileged adversary can reconstruct both the victim's input and output text, severely compromising user privacy even in locally deployed LLMs.

Specifically, the authors found that:

  • Token Value Leakage: Adversaries can infer the specific token values (akin to words or sub-word units) processed by the LLM. This is achieved by monitoring the cache access patterns during the token embedding operation. Since this operation often simplifies to a table lookup (retrieving a specific row from an embedding matrix W based on the token index), different tokens result in accesses to different memory locations. By observing which cache lines are accessed, the spy application can deduce the exact token index being processed. The autoregressive nature of LLMs means this leakage applies to both initial input tokens and subsequently generated output tokens.
  • Token Position Leakage: The precise order of tokens, particularly for the generated output, can be deduced from the timing characteristics of the LLM's autoregressive decoding phases. The sequential, step-by-step generation of output tokens—where each new token is recursively fed back into the model—creates distinguishable temporal patterns in cache access. The paper identifies a clear distinction between the dense, parallel processing of the initial input during the prefill phase and the sparser, sequential processing during the decode phases, allowing the adversary to pinpoint the boundary between input and output and the order of output tokens.

To exploit these leakages, the researchers designed a sophisticated eavesdropping attack framework. This framework does not require any direct interaction with the victim's LLM and operates without special privileges, making it stealthy and difficult to detect. The attack was rigorously evaluated on a wide array of popular open-source and proprietary local LLMs, including Meta Llama, Google Gemma, TII Falcon, Mistral, and Microsoft Phi-3.5-mini, across various inference frameworks like llama.cpp and HuggingFace Transformers.

The empirical results showcased promising accuracy:

  • For output text reconstruction, the restored text had an average Levenshtein similarity of 94.8% and an average cosine similarity score of 98.7% to the ground truth.
  • For input text reconstruction, the restored text achieved an average Levenshtein similarity of 82.7% and an average cosine similarity score of 98.0% to the ground truth.
  • The Attack Success Rate (ASR), defined as cosine similarity greater than 0.77 (a threshold determined by human evaluation to signify accurate privacy content capture), reached 99.1% for output and 99.9% for input, demonstrating significant and pervasive information leakage.

These findings highlight a critical security gap in local LLM deployments, revealing that the "privacy by default" assumption for on-device inference is fundamentally flawed in the face of hardware side-channel attacks.

Technical Deep Dive

The attack framework described in the paper is an intricate, multi-stage process designed to overcome significant challenges inherent in hardware cache side-channel exploitation against LLMs. The threat model assumes an unprivileged spy application co-located on the victim's machine. This spy app does not tamper with the LLM software, but rather passively observes shared hardware resources. Key capabilities of the adversary include the ability to execute malicious code, open the LLM model file in read-only mode to mmap it (or leverage page deduplication for shared memory), and use the clflush instruction for cache line flushing. Crucially, the adversary does not need to know the victim's specific model architecture or source code, only publicly available information to calculate address offsets of the embedding table elements. The victim is assumed to be running a local LLM, with the token embedding operation offloaded to the CPU, a common default behavior for mainstream local LLM inference frameworks due to cost-effectiveness.

Vulnerability and Leakage Sources

The foundation of the attack lies in two specific characteristics of LLM inference:

  1. Token Value Leakage (Cache Access Patterns): The LLM's token embedding converts symbolic tokens into numerical vectors. This is often implemented as a lookup into a large embedding table (matrix W), where the i-th token index corresponds to retrieving the i-th row of W. When the victim LLM processes a token, the corresponding row of W is loaded into the CPU cache. By monitoring which cache lines are accessed, the spy process can infer the token's index. Since LLMs are autoregressive, both the initial input tokens (during the prefill phase) and subsequently generated output tokens (during the decode phases) undergo this embedding operation, making both vulnerable to this leakage.
  2. Token Order Leakage (Temporal Domain): LLM inference proceeds in distinct, sequential phases: a prefill phase for the input and a series of decode phases for generating output tokens one at a time. This inherent serialization means that the timing of embedding operations correlates with the position of the output token. The prefill phase typically involves batched, parallel processing of input tokens, resulting in a dense cluster of cache events. In contrast, decode phases are characterized by sparser, sequential cache events, corresponding to one token generation at a time. This temporal distinction allows the adversary to identify the boundary between input and output tokens and determine the order of output tokens.

Challenges and Proposed Solutions

The practical implementation of this attack faces two significant challenges:

  • Challenge 1 (C1): Noise in the side channels: Cache side-channels are inherently noisy, generating false positives (observed cache hits not caused by the victim's token embedding) and false negatives (missed actual cache hits). This noise leads to randomly valued or missing tokens in the raw cache traces.
  • Solution: The paper proposes a novel text reconstruction algorithm that fuses both the timing signal and the mapped token list. This involves analyzing the Power Spectral Density (PSD) of the cache hit timing signal. The authors found that true positive timing signals during decode phases exhibit strong periodicity (e.g., a peak at 100Hz and its harmonics), reflecting the regular pipeline execution of autoregressive generation. Conversely, false positives show randomness with an evenly distributed spectrum, akin to white noise. By computing the normalized first-order difference of the timing signal (TDk), the system can identify peaks (likely false positives) and valleys (likely false negatives) in the waveform. This information is then used by a fine-tuned LLM (LLMA) to denoise and reconstruct the clean output text, treating it as a sequence-to-sequence task of "fill-in-the-blank" or "remove noise."
  • Challenge 2 (C2): Scrambled input token order: During the prefill phase, input tokens are often batched and processed in parallel for efficiency. From the side-channel observer's perspective, this results in an unordered or randomly interleaved list of input tokens (KP), preventing direct reconstruction of the original input sentence order.
  • Solution: The authors leverage the inherent contextual dependence between the LLM's input and output text. They fine-tune another LLM (LLMB) to reconstruct the original order of the shuffled input tokens (KP) by using the already reconstructed output text (bO) as a contextual reference. This problem is formulated as a sequence-to-sequence task (bI ~ P(I | KP, bO)), where LLMB learns to restore the input sequence from an unordered bag of words (KP) given the semantic context of the output.

Attack Implementation Workflow

The attack proceeds through five key phases:

  1. Measuring Cache Timing: The spy process co-locates with the victim and continuously collects a cache trace using a shared-memory-based cache side-channel attack, specifically Flush+Reload. To overcome sophisticated hardware prefetchers (like Intel's Array-of-Pointers prefetchers on Raptor Lake CPUs), the authors developed a technique to store target addresses in a non-pointer format, preventing premature cache hits. Shared memory is established either by mmap-ing the model file (leveraging zero-copy loading) or via OS page deduplication. The trace is an L x |V| matrix, where L is time steps and |V| is vocabulary size, recording memory access latencies.
  2. Identifying Prefill and Decode Phases: An asynchronous attack requires identifying the start times of LLM inference phases. A pattern-matching algorithm is used, exploiting the dense cache hits characteristic of the parallel prefill phase versus the sparser, sequential hits of the decode phases. This allows segmenting the trace into oP (prefill) and oD (decode).
  3. Extracting Token List and Timing Signal: Cache hit events (latency below a threshold α2) are mapped to token indices using the de-tokenizer. This yields an ordered list of output tokens (KD) and their corresponding timing signal (TD) for the decode phase, and an unordered list of input tokens (KP) for the prefill phase.
  4. Reconstructing Victim's Model Output: The LLMA model, fine-tuned on a synthetic dataset, takes the noisy KD and the pre-processed timing signal TD as input to reconstruct the clean output text (bO). The dataset synthesis strategy is crucial here: it simulates periodic token generation with Gaussian noise and injects false positives/negatives to mimic real-world cache traces, allowing LLMA to learn noise reduction and pattern recognition.
  5. Reconstructing Victim's Model Input: The LLMB model, also fine-tuned on a synthetic dataset, uses the unordered input token list KP and the already reconstructed output bO as context to restore the original order and semantics of the input text (bI). This synthetic dataset generation involves shuffling ground-truth input prompts and pairing them with corresponding LLM outputs, training LLMB to infer the correct sequence from context.

This sophisticated workflow, combining low-level hardware side-channel techniques with advanced LLM-based reconstruction, enables the adversary to effectively piece together sensitive conversational data from seemingly random cache activity.

Demo / Proof of Concept

While this is a conference paper rather than a live talk, the authors provide a comprehensive proof of concept through their detailed eavesdropping attack framework and extensive empirical evaluations. The framework itself serves as the practical demonstration of the identified vulnerabilities, showcasing its feasibility and effectiveness across a wide range of real-world scenarios.

The attack was rigorously evaluated on a machine equipped with an Intel 13th Gen 13900K CPU, an NVIDIA RTX 3060 GPU, and 32GiB DDR4 memory, running Ubuntu 22.04. The evaluation specifically targeted:

  • Diverse LLM Models: Meta Llama-3.1-8B, Google Gemma2-9B, TII Falcon3-10B, Mistral-7B, and Microsoft Phi-3.5-mini-3B.
  • Various LLM Inference Frameworks: llama.cpp (71k GitHub stars), HuggingFace Transformers (138k stars), Ollama (108k stars), GPT4All (71k stars), LocalAI (28k stars), Microsoft BitNet (12k stars), PowerInfer (8k stars), Intel IPEX-LLM (7k stars), koboldcpp (6k stars), and LM Studio.
  • Different Hardware Configurations: Intel 14900K, 13900K, and 12700KF CPUs.

Key Evaluation Results (RQ1 - Attack Performance):

The attack achieved remarkable accuracy in reconstructing both output and input text:

  • Output Reconstruction: For most LLMs, ROUGE-1 and ROUGE-L scores were consistently above 90.1%, and Levenshtein Similarity (LS) was no less than 87.7%. The crucial semantic-level metric, cosine similarity (φ), remained above 96.7% for all tested models, with an average of 98.7%. The Attack Success Rate (ASR), where φ > 0.77 (human-evaluated threshold for privacy exposure), averaged 99.1%.
  • Input Reconstruction: While slightly more challenging due to the scrambled token order, the average ROUGE-1 and Levenshtein similarity reached 89.9% and 82.7%, respectively. Critically, the average cosine similarity for input reconstruction was 98.0%, with an average ASR of 99.9%. This indicates that even if character-level fidelity is lower for long inputs, the semantic content—which often contains the most sensitive information—is almost perfectly recovered. The authors acknowledge a negative correlation between input token count and character/token-level accuracy, especially for very long inputs (e.g., Phi-3.5-mini on SQuAD2 dropping to 32.2% LS), but emphasize that semantic similarity (92.1% in the worst case) still remains high.

Framework and Hardware Applicability (RQ4 & RQ5):

The framework proved broadly applicable:

  • Frameworks: All 10 tested frameworks were successfully attacked when LLM inference used the CPU for embedding operations. For GPU-accelerated inference, 9 out of 10 frameworks were vulnerable, as the token embedding often defaults to CPU processing. Only HuggingFace Transformers, when embedding lookup was explicitly performed on the GPU, was not susceptible to this CPU-based cache side-channel attack.
  • Hardware: The attack demonstrated consistent performance across different Intel CPUs (14900K, 13900K, 12700KF), maintaining high cosine similarity (at least 96.5%) and 100% ASR for both input and output reconstruction, highlighting its broad applicability in consumer-grade setups.

Illustrative Examples (RQ6):

The paper provides concrete examples of reconstructed prompts, some showing perfect recovery of unique n-grams not present in the synthetic training data (e.g., "freddy krueger", "e5"), directly demonstrating the potential for Personally Identifiable Information (PII) leakage. Other examples show semantic recovery even with grammatical variations or synonym replacement, explaining why cosine similarity can remain high even if Levenshtein similarity drops. These examples powerfully illustrate the practical privacy implications of the attack.

The extensive evaluation across various LLMs, frameworks, and hardware configurations firmly establishes the attack's practicality and the pervasive nature of the identified vulnerabilities, serving as a compelling proof of concept for the "I Know What You Said" claim.

Defensive Implications

The discovery of hardware cache side-channels in local LLM inference necessitates a re-evaluation of security measures for privacy-sensitive deployments. The paper discusses several potential mitigation strategies, each with its own trade-offs and limitations:

  1. Disable Zero-Copy Loading:
  • Mechanism: Many LLM inference frameworks use zero-copy model loading (e.g., via mmap) to optimize performance by directly mapping model files into memory. Disabling this feature would prevent the spy process from easily sharing memory with the victim via mmap, thereby mitigating shared-memory-based cache attacks like Flush+Reload.
  • Implications: This approach comes with significant performance penalties. Evaluations show a 17% drop in loading latency for llama.cpp and approximately 32% extra memory overhead when multiple LLM instances are run. Furthermore, this defense does not eliminate shared memory created through OS-level page deduplication, meaning a determined adversary could still establish shared memory.
  1. Deploy Role-Based Access Control (RBAC):
  • Mechanism: Implementing robust Role-Based Access Control (RBAC) at the operating system level could limit memory page sharing. Specifically, RBAC could ensure that only explicitly authorized programs are allowed to share memory frames associated with the LLM model file. This could be achieved through kernel hooks, perhaps via technologies like eBPF, to authenticate programs per session and control their memory sharing access.
  • Implications: This is presented as a "better mitigation" but is currently lacking in today's local LLM inference frameworks. It requires fundamental changes to how operating systems and LLM frameworks manage memory access permissions, which is a non-trivial engineering effort.
  1. Use Hardware-based Mitigation (Intel Cache Allocation Technology - CAT):
  • Mechanism: Enterprise-grade CPUs, such as Intel Xeon processors, offer Cache Allocation Technology (CAT). CAT allows software to partition the Last-Level Cache (LLC), effectively isolating sensitive applications from shared cache resources. By allocating a private cache partition for the LLM inference process, attacks like Flush+Reload and Prime+Probe that rely on shared LLC contention could be eliminated.
  • Implications: The primary limitation is that CAT for LLC is typically unavailable on consumer-grade CPUs, which are the primary target platforms for local LLMs in "AI PCs." This makes it an impractical solution for the vast majority of local LLM deployments.

Challenges and Future Work for Defenders:

The paper also highlights ongoing challenges and areas for future research in mitigation:

  • Input Reconstruction Accuracy: While the semantic leakage for input is high, character-level and token-level accuracy can be lower for long inputs. This is attributed to the temporal resolution limitations of cache attacks and the parallel execution characteristics of the prefill phase. Improving temporal resolution or developing more robust reconstruction algorithms could further enhance attack capabilities.
  • Other CPU Side Channels: The current attack relies on shared-memory-based cache attacks. Future work could explore conflict-based attacks (e.g., Prime+Probe) that do not require shared memory. However, these typically offer only set-level spatial resolution, meaning multiple embedding rows could map to the same cache set, increasing the search space for token inference. Overcoming this would require leveraging LLMs to predict sentences from inter-sentence context or employing runtime profiling to learn address mappings despite Address Space Layout Randomization (ASLR).
  • Attacking GPU-side Embedding: The current attack is ineffective if the token embedding operation is entirely offloaded to a discrete GPU, as GPU caches are not coherent with CPU caches. Future research could investigate GPU cache attacks, such as Invalidate+Compare, to monitor GPU memory access patterns. These attacks also face set-level resolution challenges but might benefit from GPU-specific characteristics like per-set contention intensity and less random GPU page frame allocations, which could aid in address mapping.

In summary, while some mitigation strategies exist, they either incur significant performance overhead, require substantial architectural changes in software and operating systems, or are limited to specific hardware platforms not commonly used for local LLMs. This underscores the need for robust, multi-layered security approaches to truly safeguard privacy in the evolving landscape of local LLM inference.

Key Takeaways

  • Local LLMs are Vulnerable to Hardware Side-Channels: Contrary to common assumptions, locally deployed LLMs, intended for privacy-sensitive tasks, are susceptible to unprivileged hardware cache side-channel attacks that can leak user input and model output.
  • Dual Leakage Mechanisms Identified: The attack exploits two fundamental characteristics of LLM inference: token value leakage from cache access patterns of the token embedding operation, and token position leakage from the timing of autoregressive decoding phases.
  • Stealthy and Unprivileged Attack Framework: A novel eavesdropping framework was developed that operates without direct interaction with the victim LLM and does not require special privileges, making it difficult to detect.
  • LLM-based Reconstruction Overcomes Noise and Shuffling: The framework employs fine-tuned LLMs and a custom dataset synthesis strategy to effectively denoise noisy cache traces and reconstruct the correct order of shuffled input tokens, leveraging the contextual dependence between input and output.
  • High Accuracy Across Diverse Deployments: The attack demonstrates high accuracy in reconstructing both output (average 94.8% Levenshtein similarity, 98.7% cosine similarity) and input (average 82.7% Levenshtein similarity, 98.0% cosine similarity) across a wide range of popular LLMs and inference frameworks on consumer-grade hardware.
  • Mitigations Present Trade-offs: Proposed defenses like disabling zero-copy loading or using hardware-based isolation (Intel CAT) either incur significant performance penalties or are not widely available on typical local LLM platforms, highlighting the need for robust, integrated security solutions.

About the Speaker(s)

The research paper "I Know What You Said: Unveiling Hardware Cache Side-Channels in Local Large Language Model Inference" was authored by a team of researchers from the University of Chinese Academy of Sciences. The primary authors include Zibo Gao, Junjie Hu, Feng Guo, Yixin Zhang, Yinglong Han, Siyuan Liu, Haiyang Li, and Zhiqiang Lv. They are affiliated with the Institute of Information Engineering, the School of Cyber Security, and the Chinese Academy of Sciences, all under the umbrella of the University of Chinese Academy of Sciences. Their collective work focuses on critical areas of computer security, particularly in identifying and exploiting novel vulnerabilities in emerging technologies like Large Language Models, and contributing to the broader understanding of hardware side-channel attacks.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This is real research. Novel attack surface, serious engineering to make it work, and results that should make every 'local LLM for privacy' pitch sweat. The combination of cache side-channels with LLM-based denoising to reconstruct both input and output text is genuinely clever and the 98%+ semantic similarity numbers are damning.

Heather Calloway (CISO) — STRONG ACCEPT

This research fundamentally changes the risk calculus for local LLM deployments. The privacy assumption that kept data off third-party servers doesn't hold when an unprivileged process on the same machine can reconstruct 98%+ of conversations through cache timing. Any organization deploying local LLMs for sensitive workloads needs to know this exists.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)