Characterizing the Impact of Audio Deepfakes in the Presence of Cochlear Implant
Magdalena Pasternak
Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Audio Security
Overview
In an era increasingly shaped by artificial intelligence, the proliferation of deepfakes poses a significant and evolving threat across various domains, from political disinformation to sophisticated financial scams. Magdalena Pasternak's talk at the NDSS Symposium, titled "Characterizing the Impact of Audio Deepfakes in the Presence of Cochlear Implant," delves into a critical, yet often overlooked, aspect of this challenge: the vulnerability of individuals with different auditory perceptions, specifically cochlear implant (CI) users, to audio deepfakes. Pasternak, a deepfake researcher herself, opens with a compelling personal anecdote of being momentarily fooled by an AI-generated voice, underscoring the insidious nature of these technologies even for experts.
Key moments
- 0:00 Personal deepfake experience and problem introduction
- 2:25 Research questions: CI users and deepfake susceptibility
- 3:50 Study methodology: Human vs. machine detection
- 5:40 Automated deepfake detector performance on CI audio
- 6:30 Machine detection struggles with voice conversion deepfakes
- 7:00 Human user study: CI vs. hearing detection accuracy
- 7:50 CI users struggle significantly with voice conversion deepfakes
- 8:15 Auditory cues used by CI vs. hearing participants
Characterizing the Impact of Audio Deepfakes in the Presence of Cochlear Implant
Speakers: Magdalena Pasternak
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=lirM3vK9iI4
Overview
In an era increasingly shaped by artificial intelligence, the proliferation of deepfakes poses a significant and evolving threat across various domains, from political disinformation to sophisticated financial scams. Magdalena Pasternak's talk at the NDSS Symposium, titled "Characterizing the Impact of Audio Deepfakes in the Presence of Cochlear Implant," delves into a critical, yet often overlooked, aspect of this challenge: the vulnerability of individuals with different auditory perceptions, specifically cochlear implant (CI) users, to audio deepfakes. Pasternak, a deepfake researcher herself, opens with a compelling personal anecdote of being momentarily fooled by an AI-generated voice, underscoring the insidious nature of these technologies even for experts.
The research presented is groundbreaking, marking the first study of its kind to systematically examine how cochlear implant users perceive and detect audio deepfakes. It highlights a disproportionate susceptibility within this population, particularly to voice conversion deepfakes, and critically assesses the efficacy of current automated detection models in this context. The findings reveal significant disparities in detection capabilities between CI users and hearing individuals, driven by fundamental differences in auditory processing. This work is crucial for understanding the equitable deployment of defensive measures and ensuring that technological advancements in security do not inadvertently leave vulnerable populations exposed.
The talk not only quantifies the heightened risk faced by CI users but also identifies the specific auditory cues they rely upon—or fail to perceive—when confronted with synthetic speech. By comparing human detection performance with that of state-of-the-art machine learning models, Pasternak and her team illuminate the current limitations of AI-driven defenses and advocate for the development of inclusive, real-time detection tools. This research is a vital contribution to the field of deepfake detection, emphasizing the need for solutions that are robust, adaptable, and considerate of the diverse human experience of sound.
Background
▶ Watch: Personal deepfake experience and problem introduction (0:00)
The landscape of digital threats has been dramatically altered by the emergence of deepfakes, synthetic media generated by artificial intelligence. While initial concerns often focused on video deepfakes, audio deepfakes have rapidly advanced, becoming increasingly sophisticated and difficult to discern from genuine human speech. The speaker's opening anecdote, where she, a deepfake researcher, was momentarily deceived by a sales call bot, illustrates the everyday pervasiveness and deceptive quality of these AI-generated voices. Beyond mere annoyance, the real-world implications are severe, as demonstrated by the 2024 New Hampshire Democratic primary robocall, which impersonated President Joe Biden to discourage voting, impacting over 20,000 voters. Such incidents highlight the urgent need for effective deepfake detection mechanisms.
For most individuals, distinguishing a deepfake from genuine audio is already a significant challenge. However, this problem is compounded for populations with altered auditory perception, such as cochlear implant (CI) users. Cochlear implants are advanced medical devices designed to restore a sense of hearing by converting sound into electrical signals that directly stimulate the auditory nerve. Unlike the thousands of hair cells in a natural ear that process a broad frequency spectrum, CIs utilize a limited number of electrodes. Consequently, they must prioritize speech-relevant frequencies, simplifying and compressing sounds. This inherent processing limitation often leads to difficulties for CI users in perceiving pitch, tone, or timber—subtle auditory features that hearing individuals might unconsciously rely on for voice recognition and authenticity assessment. Therefore, CI users might depend on entirely different, more coarse auditory features for deepfake detection, potentially making them more susceptible to advanced synthetic audio.
Existing approaches to mitigate deepfake threats, such as those proposed by Warren et al., often focus on integrating automated deepfake detection models. These models are typically refined to minimize user fatigue by only flagging deepfakes that are human-undetectable, acknowledging the high false-negative rates of current automated systems. However, these models are largely developed and tested against the auditory perception of hearing individuals, neglecting the unique processing mechanisms of CI users. This oversight creates a critical gap, as models trained without considering these differences may prove ineffective for a significant portion of the population. This research directly addresses this gap by posing two central questions: How susceptible are CI users to audio deepfakes, what cues do they rely on for detection, and how does their detection ability differ from hearing persons? Furthermore, it investigates the effectiveness of automated models on CI-simulated audio and explores whether such models, if they perform similarly to CI users, could serve as reliable proxies to facilitate large-scale studies.
Key Findings
▶ Watch: Study methodology: Human vs. machine detection (3:50)
The study presented by Magdalena Pasternak represents a pioneering effort, being the first ever to investigate cochlear implant users' perception of audio deepfakes and one of the largest academic studies on CI auditory perception more broadly. The research employed a comprehensive methodology comparing both human and machine detection capabilities.
A significant undertaking involved recruiting 87 hearing participants across the US and UK. Recognizing the challenges associated with large-scale CI user recruitment due to smaller demographics, the researchers collaborated with a dedicated cochlear implant researcher. To broaden the study's scope, they leveraged acoustic simulators, tools designed to approximate how CI users perceive sounds. Two such simulators—a generic MATLAB toolbox and MATLAB via coder—were applied to the ASV spoof dataset, a standard benchmark for automatic speaker verification and spoofing attack detection. This dataset primarily comprises Text-to-Speech (TTS) and Voice Conversion (VC) deepfakes, with TTS generated entirely from text and VC modifying existing human audio to match a target voice while preserving many natural speech characteristics. Nearly a million audio samples, both original and CI-simulated, were then evaluated using four state-of-the-art deepfake detectors.
The findings from the automated detection models revealed a critical vulnerability. While these audio detectors performed well on unmodified speech samples, their Equal Error Rate (EER)—a measure balancing false acceptance and false rejection rates—rose significantly for CI-simulated data. A more granular analysis based on attack type showed that detection models maintained relatively strong performance against Text-to-Speech (TTS) deepfakes, even with CI simulation. However, for Voice Conversion (VC) deepfakes, the performance plummeted, with CI-simulated audio performing much worse. An notable exception was the Breathing Talking Speech Encoder, which maintained more consistent performance across both TTS and VC attacks, likely due to its reliance on natural human sounds for detection.
The human user study corroborated and expanded upon these machine-based insights. Hearing participants achieved an average accuracy of 78% in deepfake detection, whereas cochlear implant users managed only 67%, a statistically significant difference. This disparity was primarily attributed to an increased false negative rate among CI users, meaning they more frequently misclassified fake audio as real. This vulnerability was not uniform: CI users performed relatively well on TTS audio but performed below random guessing on voice conversion attacks, classifying VC deepfakes as real more often than fake. This highlights a specific, pronounced susceptibility to VC attacks.
Through thematic analysis of participant feedback, the study identified key differences in detection cues. CI users predominantly relied on "coarse auditory features" such as a "fake sound," "accent," or "metallic sound." In contrast, hearing participants focused on more "subtle auditory cues," including pitch variations, tone inconsistencies, and timber. A striking finding was that CI users were three times more likely to assign human-like features to voice conversion attacks, further illustrating their difficulty in perceiving the artificiality.
Finally, the researchers used saliency maps to compare human and model reliance on auditory cues. These maps, visualizing the top 30% most relevant cues on spectrograms, showed that both CI users and automated detectors heavily relied on higher frequency artifacts for detecting TTS deepfakes. However, for VC attacks, while models still used higher frequency bands, these cues became less relevant after CI simulation. This aligns with the human study, where CI users struggled to identify VC deepfakes once crucial auditory details like tone and subtle variations were altered or lost by the CI processing. The study conclusively demonstrates that CI users are disproportionately vulnerable to voice conversion deepfakes and that current detection models fail to fully capture their unique auditory challenges.
Technical Deep Dive
▶ Watch: Machine detection struggles with voice conversion deepfakes (6:30)
The core of this research lies in understanding the interplay between cochlear implant (CI) auditory processing and the characteristics of audio deepfakes, particularly Text-to-Speech (TTS) and Voice Conversion (VC) attacks. Cochlear implants function by bypassing damaged parts of the inner ear, directly stimulating the auditory nerve with electrical impulses. This process involves a speech processor that converts incoming sound into electrical signals. Crucially, CIs divide sound into a limited number of frequency bands, which are then mapped onto a limited set of electrodes in the cochlea. Unlike the broad spectrum processing capability of thousands of natural hair cells, CIs must prioritize speech-relevant frequencies, leading to a simplification and compression of sound. This simplification often results in CI users having difficulty perceiving fine-grained auditory details such as pitch, tone, and timber, which are vital for discerning nuances in human speech and for identifying synthetic alterations. The speaker also clarified in the Q&A that CIs do not reach the entirety of the cochlea, necessitating a "pitch shift" where lower frequencies are translated into higher frequency zones, further impacting natural sound perception.
The methodology employed a dual approach: evaluating human perception and machine detection. For machine detection, the researchers utilized acoustic simulators to approximate CI auditory processing. Specifically, a generic MATLAB toolbox and MATLAB via coder were used. These simulators manipulate audio signals to mimic the frequency-band limited and compressed sound that a CI user would experience. The simulated audio was then fed into four state-of-the-art deepfake detectors. The dataset used for this evaluation was the ASV spoof dataset, a widely recognized benchmark for Automatic Speaker Verification Spoofing Detection. This dataset is critical because it includes two primary types of audio deepfakes:
- Text-to-Speech (TTS) deepfakes: These are entirely synthetic, generated from text input. They often exhibit more noticeable artifacts due to their purely artificial origin.
- Voice Conversion (VC) deepfakes: These are more sophisticated, modifying an existing human audio sample to match a target voice. This process allows them to preserve many natural speech characteristics, making them considerably harder to detect.
The performance of the automated models was measured using the Equal Error Rate (EER), a metric commonly used in biometric and security systems. EER occurs when the false acceptance rate (classifying a deepfake as real) equals the false rejection rate (classifying real audio as a deepfake). A lower EER indicates better performance. The study observed a significant increase in EER for CI-simulated data, particularly for VC attacks, indicating that current models struggle when input audio is processed through a CI-like filter. The Breathing Talking Speech Encoder was highlighted as an outlier, maintaining more consistent performance. This suggests that models leveraging inherent, difficult-to-synthesize human characteristics, like breathing sounds, might be more robust against both TTS and VC attacks, even under CI-simulated conditions.
For the human study, participants listened to audio samples and made judgments on their authenticity, confidence, and the cues they used. Thematic analysis was then applied to categorize these reported cues. The use of saliency maps provided a deeper technical insight into which specific audio features (represented on spectrograms as frequency bands over time) both humans and models focused on during detection. These maps visually represent the "top 30% most relevant cues." For TTS deepfakes, both CI users and automated models showed reliance on higher frequency artifacts. However, for VC deepfakes, the saliency maps demonstrated a crucial divergence: while models continued to rely on higher frequency bands, these cues became significantly less relevant after CI simulation. This directly correlated with CI users' struggle to detect VC deepfakes, as the very auditory details (subtle tone variations) that might indicate artificiality were either altered or lost in the CI processing, rendering them imperceptible. This technical detail underscores why VC deepfakes, which retain more natural speech characteristics, are particularly challenging for CI users, as the subtle cues that might betray their artificiality are precisely those that CIs struggle to convey.
Demo / Proof of Concept
▶ Watch: Human user study: CI vs. hearing detection accuracy (7:00)
The talk did not feature a traditional live demonstration of a deepfake attack or a specific proof-of-concept tool. Instead, the "demonstration" of the problem's gravity was powerfully conveyed through the speaker's personal anecdote at the outset. She recounted being momentarily fooled by an audio deepfake in a phone conversation, highlighting how sophisticated these artificial voices have become and how easily they can deceive even those actively researching them. This personal experience served as an effective, relatable illustration of the real-world impact and the challenge of deepfake detection for the general population.
The core of the research involved a meticulous user study and machine learning model evaluation rather than a live technical exploit. The user study involved human participants (both hearing and cochlear implant users) listening to deepfake audio samples and assessing their authenticity, confidence, and the auditory cues they relied upon. This empirical approach, combined with the evaluation of state-of-the-art deepfake detectors against CI-simulated audio, served as the primary method for demonstrating the vulnerabilities and challenges discussed. While there wasn't a "demo" in the sense of showcasing a tool in action, the rigorous experimental design provided concrete evidence and data to support the findings.
Defensive Implications
▶ Watch: Auditory cues used by CI vs. hearing participants (8:15)
The research presented by Magdalena Pasternak carries profound defensive implications, particularly for ensuring inclusive and effective cybersecurity measures against evolving deepfake threats. The most critical finding is that cochlear implant (CI) users are disproportionately vulnerable to audio deepfakes, especially Voice Conversion (VC) attacks. This means that a significant population group, already navigating unique auditory challenges, faces an elevated risk from scams, misinformation, and identity theft facilitated by synthetic audio. Current automated detection models, designed for hearing individuals, demonstrably fail to adequately protect CI users, with their performance degrading significantly on CI-simulated audio.
Firstly, there is an urgent need to develop effective real-time assisted deepfake detection tools that are specifically designed with diverse auditory perception in mind. These tools must move beyond current models that struggle with CI-processed speech. The research suggests that future models could potentially incorporate principles similar to the Breathing Talking Speech Encoder, which leverages natural human sounds. Such models might need to be trained on CI-simulated audio or even directly on audio from CI users to better understand and identify artifacts that are perceptible within their unique auditory framework. The goal is to ensure that all users, particularly those with CIs, have the necessary defenses against increasingly sophisticated threats.
Secondly, the study highlighted significant misconceptions about deepfakes among users, with many underestimating their sophistication. This underscores the importance of awareness and educational programs. While such programs are crucial for empowering users to recognize potential deepfake indicators (like unnatural phrasing or metallic sounds, which CI users tended to pick up), the speaker rightly concludes that education alone will not suffice. The inherent auditory processing limitations for CI users mean that even with awareness, they may simply be unable to perceive the subtle cues that betray a deepfake.
Thirdly, the Q&A session brought up intriguing avenues for future research and defensive strategies. The possibility of modifying the software within cochlear implants themselves to assist in deepfake detection was raised. While complex due to factors like "spectral smearing" and "pitch shift" within CI processing, this represents a long-term, fundamental approach to enhancing a user's intrinsic ability to discern synthetic audio. Similarly, the discussion extended to modern hearing aids, which are increasingly computerized and powerful. Exploring parallels with hearing aid users and investigating whether their device software could be adapted for deepfake detection could benefit a much broader demographic experiencing various degrees of hearing loss.
In summary, defensive strategies must be multi-faceted:
- Technological Innovation: Prioritize the development of inclusive, real-time deepfake detection tools tailored for CI users and potentially other DHH (Deaf and Hard of Hearing) groups. These tools should account for unique auditory processing, perhaps by integrating CI-specific training data or focusing on robust, difficult-to-synthesize human vocal characteristics.
- User Education: Implement targeted awareness campaigns to combat common misconceptions about deepfakes, although acknowledging their limitations for CI users.
- Device-Level Enhancements: Investigate the potential for modifying the software in cochlear implants and advanced hearing aids to directly assist in deepfake detection, moving beyond external assistive tools.
Ultimately, the research emphasizes that security solutions must be equitable, protecting all populations, including those with unique physiological and perceptual characteristics.
Key Takeaways
- Disproportionate Vulnerability: Cochlear implant (CI) users are significantly more vulnerable to audio deepfakes than hearing individuals, particularly to sophisticated Voice Conversion (VC) attacks.
- Ineffective Automated Detection: Current state-of-the-art automated deepfake detection models perform poorly on CI-simulated audio, especially for VC deepfakes, indicating a critical gap in existing defensive technologies.
- Distinct Auditory Cues: CI users rely on more "coarse auditory features" (e.g., "fake sound," "metallic sound") for deepfake detection, whereas hearing persons utilize "subtle auditory cues" (e.g., pitch variations, tone inconsistencies, timber), which are often lost in CI processing.
- VC Deepfakes Are Insidious: For CI users, voice conversion deepfakes are particularly deceptive, often being classified as real even below random guessing, making them an extremely potent threat vector.
- Urgent Need for Inclusive Tools: There is a critical requirement for developing effective, real-time assisted deepfake detection tools that are specifically designed and tested to accommodate the unique auditory processing challenges faced by CI users.
- Beyond Awareness: While educational programs can help dispel misconceptions about deepfakes, they are insufficient alone; robust technological solutions are paramount to provide equitable protection for all users.
About the Speaker(s)
Magdalena Pasternak is a dedicated deepfake researcher. Her work, as highlighted in this presentation, focuses on understanding the complex impacts of deepfake technology, particularly on diverse user groups. Through her research, Pasternak aims to uncover vulnerabilities and contribute to the development of more inclusive and effective defense mechanisms against AI-generated threats. The personal anecdote shared at the beginning of her talk underscores her practical experience and deep understanding of the deceptive nature of deepfakes.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Genuinely novel research that opens a population-specific vulnerability no one in the deepfake detection space has systematically examined before. The combination of human user study, acoustic simulation, and saliency-map analysis across two deepfake modalities is methodologically sound and produces concrete, actionable findings — not just 'CI users struggle,' but specifically why VC attacks collapse detection below chance for this group.
Heather Calloway (CISO) — WEAK
Rigorous and novel research on a genuinely underserved population, but the institutional implications are left almost entirely unexplored. The work identifies a real disparity — CI users fail to detect voice conversion deepfakes at rates below chance — but stops well short of translating that into anything a security leader, regulator, or platform operator can act on.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025