EchoLLM: LLM-Augmented Acoustic Eavesdropping Attack on Bone Conduction Headphones with mmWave Radar

Xin Yao

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Privacy 3: Attacks

Overview

The proliferation of bone conduction headphones, lauded for their open-ear design and suitability for active lifestyles, has inadvertently introduced a novel and significant privacy vulnerability. This talk, "EchoLLM: LLM-Augmented Acoustic Eavesdropping Attack on Bone Conduction Headphones with mmWave Radar," presented by Xin Yao, unveils a sophisticated eavesdropping technique that leverages cutting-edge millimeter-wave (mmWave) radar technology combined with the advanced capabilities of Large Language Models (LLMs) to reconstruct private conversations. The research highlights how the inherent mechanism of bone conduction – transmitting sound through vibrations – can be exploited to inadvertently leak sensitive information, from passwords to confidential discussions.

Watch on YouTube · Slides

Visual summary for EchoLLM: LLM-Augmented Acoustic Eavesdropping Attack on Bone Conduction Headphones with mmWave Radar by Xin Yao
Visual summary for EchoLLM: LLM-Augmented Acoustic Eavesdropping Attack on Bone Conduction Headphones with mmWave Radar by Xin Yao

Key moments

  1. 0:00 Introduction to EchoLLM and privacy risks of bone conduction headphones
  2. 3:00 Limitations of existing acoustic eavesdropping attacks
  3. 4:15 EchoLLM's novel approach and design goals
  4. 6:00 EchoLLM system architecture and workflow overview
  5. 7:00 Detailed vibration signal extraction process
  6. 8:00 Signal enhancement: background reduction, spike elimination, motion compensation
  7. 10:50 Context-aware multimodal ASR using large language models

EchoLLM: LLM-Augmented Acoustic Eavesdropping Attack on Bone Conduction Headphones with mmWave Radar

Speakers: Xin Yao

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=cMcvX_dWwlw

Overview

The proliferation of bone conduction headphones, lauded for their open-ear design and suitability for active lifestyles, has inadvertently introduced a novel and significant privacy vulnerability. This talk, "EchoLLM: LLM-Augmented Acoustic Eavesdropping Attack on Bone Conduction Headphones with mmWave Radar," presented by Xin Yao, unveils a sophisticated eavesdropping technique that leverages cutting-edge millimeter-wave (mmWave) radar technology combined with the advanced capabilities of Large Language Models (LLMs) to reconstruct private conversations. The research highlights how the inherent mechanism of bone conduction – transmitting sound through vibrations – can be exploited to inadvertently leak sensitive information, from passwords to confidential discussions.

EchoLLM addresses critical limitations inherent in previous acoustic eavesdropping methods, which often struggled with real-world applicability due to assumptions about predefined audio content, reliance on stable reflective surfaces, or a lack of multimodal data integration. By introducing a low-cost, highly sensitive, and easily integrable mmWave radar system, coupled with an innovative context-aware multimodal Automatic Speech Recognition (ASR) framework powered by LLMs, EchoLLM achieves unprecedented accuracy and robustness in reconstructing natural, context-dependent speech. This work not only exposes a significant privacy risk but also demonstrates the formidable potential of combining physical layer attacks with advanced AI to extract sensitive data from seemingly secure devices.

The significance of EchoLLM lies in its ability to transform subtle, often overlooked, physical phenomena into actionable intelligence. It represents a paradigm shift in external eavesdropping, moving beyond simplistic signal reconstruction to contextually understand and transcribe human speech. For security practitioners, device manufacturers, and end-users, this research serves as a stark warning and a call to action, emphasizing the need to re-evaluate the security postures of devices that interact with the physical world in novel ways, especially those designed for personal communication.

Background

▶ Watch: Introduction to EchoLLM and privacy risks of bone conduction headphones (0:00)

Bone conduction headphones operate on a principle distinct from traditional air-conduction headphones. Instead of vibrating air into the ear canal, they transmit sound vibrations directly to the bones of the skull, bypassing the eardrum and cochlea. This allows users to hear audio while keeping their ears open to ambient sounds, a feature highly valued for safety during sports or for maintaining situational awareness. However, this very mechanism – the generation of physical vibrations – creates an exploitable side channel. These vibrations, though subtle, radiate outwards from the headphone transducer and the user's skull, carrying the audio content.

Acoustic eavesdropping attacks generally fall into two categories: internal and external. Internal attacks rely on gaining access to the victim's device itself, perhaps through a malicious application exploiting sensor data (e.g., an IMU in a smartphone) or a permission vulnerability. While potentially effective, these attacks are limited by the prerequisite of device access and often require victim interaction, such as installing malware. External attacks, on the other hand, are more insidious, as they do not require any direct interaction with or modification of the victim's device. Prior external eavesdropping research has explored various mediums, including reconstructing speech from laser reflections off surfaces, vibrations from light bulbs or tables, or even analyzing Wi-Fi signal perturbations. While these methods demonstrate the feasibility of remote eavesdropping, they have historically faced significant challenges. These challenges include balancing accuracy with robustness, often incurring high costs for specialized equipment, or delivering poor performance in practical, real-world scenarios.

The primary limitations of existing external acoustic eavesdropping attacks, especially those based on millimeter-wave radar, can be summarized as follows:

  1. Predefined Content Categories: Many prior approaches assumed that the target audio belonged to a limited set of predefined categories, such as letters or digits. This required attackers to have prior knowledge of these categories, severely restricting applicability to natural, unconstrained speech.
  2. Reliance on Stable Reflective Surfaces: Existing methods often depended on stable reflective surfaces to accurately capture vibrations, making them vulnerable to environmental changes or victim movement. They lacked the capability to handle context-dependent natural speech, which is dynamic and complex.
  3. Lack of Multimodal Data Integration: Previous attacks typically relied on a single modality, overlooking the potential for robust speech reconstruction through the fusion of multiple data streams.
  4. Inadequate Evaluation Metrics: Past evaluations often focused solely on audio quality metrics like Signal-to-Noise Ratio (SNR) or Perceptual Evaluation of Speech Quality (PESQ), which performed poorly on downstream tasks like speech recognition, where Word Error Rate (WER) remained unacceptably high even with seemingly good audio quality.

EchoLLM directly addresses these shortcomings by proposing a novel approach that breaks through predefined content categories and applicability constraints. The system operates under a common "SW model" scenario, where an attacker and a victim are in the same physical space, and the victim is conversing with an AI (e.g., a smart assistant) using bone conduction headphones. In this scenario, EchoLLM aims to achieve context-aware eavesdropping by exploiting semantic dependencies in the dialogue. Furthermore, it seeks robust multimodal fusion by combining mmWave radar signals targeting the victim's earphones with microphone signals directed at the victim's own voice. The attack prioritizes being covert and low-cost, utilizing highly integrated commercial mmWave readers. Finally, it incorporates target awareness and motion resilience to ensure high applicability despite user movement and environmental complexity.

Key Findings

▶ Watch: EchoLLM's novel approach and design goals (4:15)

EchoLLM represents a significant advancement in acoustic eavesdropping, demonstrating state-of-the-art performance in reconstructing speech from bone conduction headphone vibrations. The core findings highlight the efficacy of its multimodal, context-aware approach and the robustness of its signal processing techniques.

Firstly, the research unequivocally demonstrates that Large Language Models (LLMs) significantly enhance attack performance by enabling context-aware speech reconstruction. The ablation study revealed that incorporating contextual information dramatically reduces the Word Error Rate (WER), allowing the system to accurately transcribe natural, context-dependent dialogue far beyond the capabilities of previous methods limited to predefined categories. This is a critical breakthrough for real-world applicability.

Secondly, EchoLLM achieves robust eavesdropping through multimodal fusion, combining vibration signals from a millimeter-wave radar with ambient audio captured by a wireless microphone. This multimodal input provides a richer dataset for the LASR (LLM-based ASR) framework, making the attack more resilient to noise and environmental variations compared to single-modality approaches.

Thirdly, the study validates the crucial role of advanced signal enhancement techniques. Both background refraction reduction and motion calibration were shown to greatly improve attack efficiency. The SpikeFree algorithm for eliminating signal spikes and the moving average filter for mitigating head movement artifacts are essential components that ensure the integrity and clarity of the vibration signals fed into the ASR system.

Finally, EchoLLM exhibits remarkable robustness across diverse real-world conditions. Experiments conducted in four different environments and with three distinct types of bone conduction headphones (BCHs) confirmed its resilience to varying noise levels and device characteristics. The system also maintained effectiveness despite user diversity, headset diversity, multilingual content, the presence of numerical data, and interference from sensitive information. While victim motion did increase WER, the system's overall performance remained superior, demonstrating its practical viability for covert surveillance at a low hardware cost of approximately $930.

Technical Deep Dive

▶ Watch: EchoLLM system architecture and workflow overview (6:00)

The EchoLLM system is meticulously designed as a multi-stage pipeline to achieve robust, context-aware acoustic eavesdropping. It integrates hardware for data acquisition, sophisticated signal processing for enhancement, and advanced AI models for speech transcription.

The overall process begins with two primary data collection streams: a millimeter-wave (mmWave) radar to capture vibrations from the bone conduction headphone and a wireless microphone to record the caller's direct voice. These raw ADC (Analog-to-Digital Converter) data from the radar and audio recordings from the microphone are then fed into a vibration extraction module, followed by an audio enhancement phase. The enhanced data from both modalities are subsequently input into LASR (LLM-based ASR), a large language model-based speech-to-text model, which outputs the reconstructed dialogue text between the caller and the AI.

Vibration Signal Extraction

The first critical step is to accurately extract the subtle vibration signals emanating from the bone conduction headphone, while distinguishing them from other ambient reflections.

  1. Target Narrowing: The system first narrows down suspicious target ranges by applying a speed threshold. This helps filter out static background objects and focuses on potential human targets.
  2. Voice Activity Detection (VAD): Within the narrowed range, a CNN-based Voice Activity Detection model, called EchoVAD, is employed. EchoVAD analyzes the spectrogram of the radar signals to identify periods of speech activity and their corresponding timestamps. This ensures that processing resources are focused only when speech is likely present.
  3. Extraction Location Selection: For regions identified with audio activity, the system observes that successive range bins exhibit distinct voice activity. It then selects the range beam with the largest Signal-to-Noise Ratio (SNR) as the optimal extraction location, maximizing the quality of the raw vibration data.
  4. Displacement Restoration: Finally, the time series of the vibration signal is restored by leveraging the fundamental relationship between the phase change in the reflected mmWave signal and the displacement change of the vibrating surface.

Audio Enhancement

After raw vibration signal extraction, a series of enhancement steps are crucial to refine the signal quality, making it suitable for speech recognition.

  1. Background Refraction Reduction: The reflected mmWave signal inevitably includes reflections from other sources in the environment in addition to the target. To eliminate these static background reflections, EchoLLM adopts the circle fitting method introduced in "Moiccom 2020." This technique effectively isolates the target-specific reflections by modeling and removing the constant background components.
  2. Spike Elimination (SpikeFree): The raw vibration signals often contain sudden, high-amplitude spikes. These are speculated to be caused by inter-frame discontinuous segment-wise fitting artifacts and cross-frame localization instability. To address this, the system proposes SpikeFree, an algorithm that utilizes an LSTM (Long Short-Term Memory) model to interpolate missing values and perform offset calibration to flatten the signal. This step significantly cleans the time-amplitude plot, removing disruptive anomalies.
  3. Motion Calibration: Victim head movements introduce low-frequency components into the vibration signal, which can interfere with accurate speech reconstruction. To mitigate this, a moving average filter is first applied to obtain the low-frequency motion time series. Subsequently, a differencing operation is performed to remove these motion-induced low-frequency components, yielding a cleaner audio vibration signal.

Context-Aware Multimodal ASR (LASR)

The culmination of the EchoLLM attack is the Context-Aware Multimodal ASR (LASR) framework, which processes the enhanced signals to generate the final dialogue text. This framework consists of five main parts:

  1. Multimodal Input: The system uses a two-stage fine-tuning strategy to prepare the multimodal input. This strategy involves fine-tuning the models on both synthesized conversational audios and self-collected realistic data, ensuring robustness and adaptability to various real-world scenarios.
  2. Embedding Extraction: Embeddings are extracted from the data of each modality (mmWave vibration signals and microphone audio) separately. The embedding layer can be chosen from various state-of-the-art options, such as VGG, HuBERT V2, or wav2vec, depending on the specific characteristics and requirements.
  3. Embedding Alignment: The extracted embedding vectors from each modality are then concatenated. To facilitate subsequent processing and evaluation, BOS (Beginning of Sentence) tokens and segmentation tokens are inserted within the concatenated vectors. This helps the LLM understand the structure and boundaries of the input.
  4. Large Language Models (LLMs): The aligned and concatenated embedding vectors are fed into powerful large language models. The research mentions the use of widely adopted models such as BART, Whisper, and SpeechT5, which are capable of processing complex multimodal inputs and generating coherent text.
  5. Training and Output: The LASR framework is trained using the widely adopted auto-regressive cross-entropy loss. This loss function measures the difference between the predicted probability distribution of the next token and the actual tokens in the training data, guiding the model to accurately predict the dialogue. The final output text is then split into the caller's text and the AI's response text, with the Word Error Rate (WER) primarily calculated for the caller's part to evaluate the eavesdropping efficacy.

Demo / Proof of Concept

▶ Watch: Signal enhancement: background reduction, spike elimination, motion compensation (8:00)

The practical viability and effectiveness of EchoLLM were rigorously demonstrated through a comprehensive experimental setup and evaluation process. The hardware employed for the attack was notably low-cost, totaling approximately $930, which included the essential millimeter-wave radar unit and a wireless microphone. This cost-effectiveness underscores the accessibility of such an attack.

The system's performance was evaluated by systematically varying several critical parameters:

  • Volume: Counter-intuitively, the experiments showed that WER increased with an increase in volume. This observation, as stated in the talk, suggests a complex interaction between signal strength, vibration characteristics, and the ASR system's ability to process potentially over-saturated or distorted signals at higher volumes.
  • Distance: As expected, the WER decreased with increasing distance between the radar and the victim. This highlights the inverse relationship between signal strength and distance, which is a fundamental challenge for remote sensing attacks.
  • Angle: Interestingly, the change of angle had little effect on WER, suggesting that the mmWave radar's ability to pick up vibrations is relatively robust to the orientation of the victim or the headphone.
  • Action/Movement: The static state yielded the lowest WER, while any motion of the human body caused the WER to increase. This indicates that victim movement, even after motion calibration, still introduces some level of interference or signal degradation.

An ablation experiment was conducted to quantify the contribution of key components. This study confirmed that contextual information (provided by the LLMs) significantly enhanced attack performance. Furthermore, both reflection reduction and motion calibration were demonstrated to substantially improve attack efficiency, validating the necessity of these signal enhancement steps.

To ascertain the robustness of EchoLLM in diverse scenarios, extensive tests were performed across various conditions:

  • Environments: The attack was tested in four different environments, demonstrating its resilience to varying background noises and reflective properties.
  • Bone Conduction Headphone (BCH) Types: Experiments were conducted with three different types of BCHs, proving the attack's generalizability across different headphone models and manufacturers.
  • Multilingual Robustness: The system's ability to transcribe speech in multiple languages was evaluated.
  • User Diversity and Headset Diversity: Tests accounted for variations among users and different specific headsets.
  • Numerical Data and Sensitive Information Interference: The attack's performance was assessed when dealing with numerical data and in the presence of sensitive information, demonstrating its capability to extract such critical details.

The experimental results consistently showed that EchoLLM achieved state-of-the-art performance compared to existing methods, proving its robustness against different noises, environments, and BCH types. This comprehensive demonstration validates EchoLLM as a highly effective and practical acoustic eavesdropping attack.

Defensive Implications

▶ Watch: Context-aware multimodal ASR using large language models (10:50)

The EchoLLM attack presents a significant privacy concern for users of bone conduction headphones and necessitates a multi-faceted defensive strategy. Understanding the attack's mechanism — leveraging subtle physical vibrations and mmWave radar — is key to devising effective countermeasures.

  1. Physical Distance and Line of Sight Mitigation: The attack's effectiveness is inversely proportional to distance. Therefore, maintaining physical distance from potential attackers equipped with mmWave radar is a primary defense. In sensitive environments, users should be aware that even in a seemingly private space, their conversations could be intercepted if an attacker is within range and has a clear "line of sight" for the radar signal to the headphone or skull area. Obstructing the direct path of the radar signal with dense, non-reflective materials could also offer some protection, although this is often impractical in everyday settings.
  1. Headphone Design and Material Innovations: Bone conduction headphone manufacturers have a crucial role to play. Future designs could incorporate:
  • Vibration Dampening Materials: Using materials that absorb or significantly attenuate external vibration leakage from the transducers could reduce the signal strength available to an attacker.
  • Active Vibration Cancellation: Similar to active noise cancellation, headphones could generate anti-phase vibrations to actively cancel out leaked signals, making them harder to detect remotely.
  • Shielding: Exploring designs that physically shield the vibrating components or the skull area where vibrations are transmitted, potentially with radar-absorbing materials, could be considered, though this might impact comfort or aesthetics.
  1. Environmental Obfuscation and Noise Injection: Introducing controlled "noise" in the mmWave spectrum or generating random, non-speech vibrations around the user could confuse the radar system. While challenging for individuals to implement, this could be a strategy for secure facilities. Furthermore, ensuring ambient acoustic noise (e.g., white noise generators) in sensitive areas could help mask the victim's own voice, which is used as a multimodal input in EchoLLM.
  1. User Awareness and Behavioral Changes: Users of bone conduction headphones should be made aware of this vulnerability. In situations involving sensitive conversations (e.g., discussing financial information, medical details, or proprietary business data), users should:
  • Avoid using bone conduction headphones.
  • Switch to traditional, sound-isolating headphones.
  • Conduct conversations in truly private, secured spaces.
  • Minimize talking aloud if the bone conduction headphones are playing sensitive audio, as the victim's own voice is used to enhance the attack.
  1. Research into Counter-Surveillance Technologies: Further research is needed into developing cost-effective, portable counter-surveillance devices that can detect mmWave radar signals or actively disrupt them without interfering with other legitimate wireless communications. This could involve passive detection of radar emissions or active jamming of specific frequency bands used for such eavesdropping.

Ultimately, mitigating the EchoLLM attack requires a multi-layered approach, combining technological advancements in headphone design with increased user awareness and strategic behavioral changes, particularly in environments where sensitive information is exchanged.

Key Takeaways

  • Bone conduction headphones, while offering open-ear convenience, pose a significant and underappreciated privacy risk due to the inherent leakage of audio content through physical vibrations.
  • Millimeter-wave (mmWave) radar offers a low-cost ($930 hardware), sensitive, and easily integrable solution for external acoustic eavesdropping, overcoming limitations of prior methods.
  • The integration of Large Language Models (LLMs) is a game-changer for eavesdropping attacks, enabling context-aware speech reconstruction from noisy vibration signals and drastically improving accuracy (lower WER) for natural, unconstrained dialogue.
  • Multimodal fusion, combining mmWave radar data with ambient microphone audio, significantly enhances the robustness and accuracy of the attack, making it more resilient to environmental noise and variations.
  • Advanced signal processing techniques, including background refraction reduction, SpikeFree for spike elimination, and motion calibration, are crucial for cleaning raw vibration signals and making the attack effective in dynamic, real-world scenarios.
  • EchoLLM demonstrates state-of-the-art, robust eavesdropping capabilities across diverse conditions, including different environments, bone conduction headphone types, user movements, multilingual content, and the presence of sensitive information, highlighting a critical vulnerability that device manufacturers and users must address.

About the Speaker(s)

Xin Yao is the speaker for this talk, presenting groundbreaking research on the privacy risks associated with bone conduction headphones and introducing the novel EchoLLM attack. The transcript and provided metadata do not offer further biographical details such as their specific title or institutional affiliation beyond their name, but their contribution to this field of security research is clearly demonstrated through the detailed work presented.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Legitimate academic security research that crosses a genuinely novel attack surface — bone conduction headphone vibration exfiltration via mmWave radar, with LLM-assisted ASR to push WER into practical territory. The fusion of physical-layer side-channel work with modern ML pipelines is the real contribution here, and the $930 hardware budget makes the threat model credible rather than theoretical.

Heather Calloway (CISO) — WEAK

Technically credible research demonstrating a novel physical-layer attack on bone conduction headphones using mmWave radar and LLM-assisted reconstruction. The attack works, the methodology is rigorous, and the threat is real — but the talk stops at the research boundary and never crosses into the territory that matters to operators, security leaders, or anyone with budget authority.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)