SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song Covers
Guangke Chen
Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Audio Security
Overview
In an era increasingly dominated by AI-generated content, the music industry faces unprecedented challenges, particularly from AI-based automated song covers. These sophisticated tools leverage Singing Voice Conversion (SVC) technology to transform a song's vocal rendition from one singer's style and timbre to another's, all while meticulously preserving the original lyrics and melody. While SVC holds potential for beneficial applications, such as fans creating personalized content, its widespread accessibility through open-source toolkits has led to a surge in illegal song covers, infringing upon artists' rights and record companies' copyrights. Guangke Chen's presentation introduces SongBsAb, a pioneering dual prevention approach designed to proactively combat these unauthorized AI-generated covers.
Key moments
- 0:00 Introduction to AI song covers and industry challenges
- 2:59 Problems with reactive detection for illegal song covers
- 4:00 Introducing Song BSAB: a proactive dual prevention approach
- 6:15 Ensuring inaudibility of protection using simultaneous masking
- 8:00 Enhancing transferability to unknown singing voice conversion models
- 10:00 Audio demonstration of identity and lyric disruption effects
SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song Covers
Speakers: Guangke Chen
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=s6gnbQvNHCU
Overview
In an era increasingly dominated by AI-generated content, the music industry faces unprecedented challenges, particularly from AI-based automated song covers. These sophisticated tools leverage Singing Voice Conversion (SVC) technology to transform a song's vocal rendition from one singer's style and timbre to another's, all while meticulously preserving the original lyrics and melody. While SVC holds potential for beneficial applications, such as fans creating personalized content, its widespread accessibility through open-source toolkits has led to a surge in illegal song covers, infringing upon artists' rights and record companies' copyrights. Guangke Chen's presentation introduces SongBsAb, a pioneering dual prevention approach designed to proactively combat these unauthorized AI-generated covers.
SongBsAb represents a crucial shift from reactive detection to proactive prevention. Recognizing the limitations of traditional detection methods—which are often too late, inaccurate against high-quality SVC outputs, and inefficient given the sheer volume of infringing content—SongBsAb intervenes at the source. It empowers song owners, typically record companies, to embed imperceptible perturbations into their songs before release. These carefully crafted alterations are designed to disrupt the SVC process, causing the output song to deviate significantly from expectations, either by altering the singer's perceived identity or by corrupting the lyrics.
This innovative approach addresses multiple facets of the infringement problem, from protecting a singer's civil rights over their voice and reputation to safeguarding record companies' copyrights on distribution and the fundamental copyrights of lyrics and melodies. By making it difficult for SVC models to produce accurate covers, SongBsAb aims to restore competitive fairness, preserve artists' livelihoods, and provide a robust, transferable defense mechanism against the evolving landscape of AI misuse in music.
Background
▶ Watch: Introduction to AI song covers and industry challenges (0:00)
The proliferation of AI-generated content has significantly impacted various creative domains, with music being a prominent example. Within AI-generated music, a distinction can be drawn between entirely new music generation and AI-based automated song covers. The latter, which is the focus of this work, involves taking an existing song and re-rendering it in the voice of a different singer. The core technology enabling this is Singing Voice Conversion (SVC). SVC systems typically take two inputs: a target song from a target singer to extract singing style and timbre information, and a source song from a source singer to provide the lyrical content and melodic structure. The ultimate goal is to produce an output that sounds as if the target singer is performing the source song.
The accessibility of SVC technology has grown exponentially, with numerous open-source toolkits available online that can be used "out of the box." This low barrier to entry has fueled a widespread phenomenon of AI-based song covers across the internet, exemplified by the popularity of virtual singers in regions like China. While SVC offers some beneficial applications, such as enabling music fans to create personalized renditions of their favorite songs with their idols' voices for entertainment, it also poses severe challenges to the music industry.
The challenges are multifaceted and deeply impact various stakeholders:
- Infringement of Civil Rights: Target singers face infringement on their civil rights over their voices and reputations, especially if the generated content contains sensitive or inappropriate material.
- Copyright Infringement (Record Companies): Record companies suffer infringement on their copyrights related to the release and distribution of their songs.
- Copyright Infringement (Lyrics and Melodies): The original creators of lyrics and melodies also face copyright infringement.
- Erosion of Skill Competitiveness: Singers' unique vocal skills and artistry, which are central to their livelihood, are undermined by AI models that can mimic their voices.
- Unfair Competition: Record companies face unfair competition because the cost of producing an AI-covered song is significantly cheaper and faster than traditional methods, creating an uneven playing field.
Given these pressing issues, addressing SVC-based illegal song covers is imperative. A straightforward, initial thought might be to detect these covered songs reactively once they are online. However, this reactive detection approach suffers from three critical problems:
- Too Late: The infringement has already occurred, and detection cannot prevent the initial act of dissemination.
- Low Accuracy: Modern SVC models can produce high-quality covered songs that are increasingly difficult to distinguish from legitimate recordings, leading to limited detection accuracy, as highlighted by the presentation's reference to studies showing poor detection rates.
- Inefficiency: The sheer volume of covered songs available online makes it economically inefficient and practically impossible to detect and remove them all.
These limitations underscore the urgent need for a proactive prevention approach rather than a reactive detection one. SongBsAb is presented as the first such proactive solution.
Key Findings
▶ Watch: Introducing Song BSAB: a proactive dual prevention approach (4:00)
The central innovation of this research is the development of SongBsAb, the first proactive prevention approach specifically designed to counter Singing Voice Conversion (SVC) based illegal song covers. Unlike reactive detection methods that attempt to identify infringing content after it has been created and distributed, SongBsAb intervenes before a song is released. Its core mechanism involves adding carefully crafted, imperceptible perturbations to the original audio. These perturbations are engineered such that if the protected song is subsequently used by an SVC model, the resulting output song will significantly deviate from what is expected, effectively disrupting the conversion process.
SongBsAb achieves this disruption through a dual prevention strategy:
- Identity Disruption: The output song does not sound like the intended target singer. This is crucial for protecting the singer's unique vocal identity and reputation.
- Lyric Disruption: The output song contains unexpected or garbled lyrics, compromising the integrity of the lyrical content. This protects the copyright of the lyrical composition.
Key findings from the evaluation of SongBsAb demonstrate its effectiveness:
- Significant Identity Disruption: After applying SongBsAb, the identity similarity of the output with the target singer drops significantly. This indicates that the system successfully prevents the SVC model from accurately mimicking the target singer's voice.
- Significant Lyric Disruption: Concurrently, the lyric word error rate of the output song increases significantly. This confirms that the perturbations effectively corrupt the lyrical content, rendering the cover unusable or noticeably flawed.
- Preservation of Song Quality: A critical achievement is that these disruptive perturbations are designed to be inaudible to human listeners, ensuring that the enjoyment and commercial viability of the original, protected song are not compromised. Human studies confirmed that most participants did not perceive the perturbations as noise impacting their enjoyment.
- Transferability to Unknown Models: SongBsAb incorporates mechanisms that enhance its transferability to unknown SVC models, a vital feature given the dynamic and evolving nature of AI technologies. This ensures the protection remains effective even against SVC models not explicitly trained on during the perturbation generation process.
- First Proactive Solution: By addressing the problem at the source and offering a robust, dual-layered defense, SongBsAb establishes itself as a groundbreaking solution in the fight against AI-driven intellectual property infringement in the music industry.
Technical Deep Dive
▶ Watch: Ensuring inaudibility of protection using simultaneous masking (6:15)
SongBsAb's technical architecture is meticulously designed to address the complex challenges of proactive prevention against SVC-based illegal song covers. The approach centers on embedding imperceptible perturbations into a song before its public release. The design tackles three primary challenges: the uncertainty of a song's role (target or source), the requirement for inaudible perturbations, and the need for transferability to unknown SVC models.
Challenge 1: Dual Prevention for Unknown Song Roles
Song owners cannot predict whether their song will be used as a target song (providing singing style and timbre) or a source song (providing lyrics and melody) by an SVC model. To address this, SongBsAb implements a dual prevention strategy encompassing both identity disruption and lyric disruption. The generation of these perturbations is formulated as an optimization problem, with an objective function that combines an identity disruption loss (FID) and a lyric disruption loss (FD).
- Identity Disruption Loss (FID): This loss is specifically designed as a gender transformation loss. The core idea is to minimize the distance between the singer timbre features of the song to be protected and those of a singer with the opposite gender from the original singer. For instance, if the original singer is female, the perturbation aims to make the output sound like it's sung by a male singer, and vice versa. This strategy leverages the fact that human listeners are highly adept at distinguishing voices of different genders, thereby maximizing the perceptual impact of identity disruption. The goal is that the SVC output, when applied to a protected song, will sound notably different in terms of gender, making the cover immediately recognizable as flawed.
- Lyric Disruption Loss (FD): This component aims to introduce unexpected changes to the lyrics. It is designed as a high and low hierarchy loss, which allows for divergence in both high-level and low-level lyric features between the protected song and a reference song containing different lyrics. This ensures that the generated perturbation effectively corrupts the lyrical content, making the converted song contain garbled or incorrect words, thus hindering the integrity of the original lyrical composition.
The objective function combines these two losses, ensuring that the generated perturbations simultaneously target both the identity and lyrical aspects of the SVC output.
Challenge 2: Inaudible Perturbations for Song Enjoyment
A crucial requirement for SongBsAb's practical adoption is that the embedded perturbations must not degrade the listening experience or the quality of the original song. To achieve this, the system employs backing track-refined simultaneous masking.
- Simultaneous Masking Principle: This psychoacoustic phenomenon states that when two sounds occur simultaneously, if one is significantly louder than the other, the quieter sound may become inaudible if its intensity falls below the hearing threshold determined by the louder sound. The perturbation is treated as the quieter, masked sound.
- Refined Masker Definition: Previous works often considered only the singing voice as the masker. SongBsAb refines this by treating both the singing voice and the backing track (instrumental accompaniment) as maskers. A song is a composite of these two elements. By considering both, the system can place perturbations under the masking thresholds of either the singing voice or the backing track. This significantly expands the "masking space," allowing for more effective and robust perturbations to be embedded without being perceived by human ears. The perturbation will remain inaudible as long as it falls below either the masking threshold of the singing voice or the masking threshold of the backing track.
This sophisticated masking technique ensures that the protective perturbations are acoustically transparent, maintaining the high-quality requirement for commercial music.
Challenge 3: Transferability to Unknown SVC Models
The landscape of SVC models is constantly evolving, with new architectures and techniques emerging regularly. For SongBsAb to be a viable long-term solution, its perturbations must exhibit transferability—meaning they must remain effective even when applied to SVC models that were not known or used during the perturbation generation phase. SongBsAb enhances transferability through two key mechanisms:
- Encoder Ensemble: This technique involves generating perturbations not just for a single SVC model, but across multiple local black-box SVC models. By training the perturbation generation process against a diverse ensemble of models, the resulting perturbations become more generalized and robust. This increases the likelihood that they will effectively transfer and disrupt the operation of entirely unknown SVC models encountered in real-world scenarios.
- Frame-level Interaction Reduction Loss: Perturbation interaction refers to how strongly correlated different units of a perturbation are in achieving their effect. High interaction among perturbation units might make them highly effective on the specific SVC models they were trained against, but this intricate correlation is unlikely to hold true on an unknown SVC model. Therefore, the system minimizes the interaction loss at the frame level. By reducing the interdependence among perturbation units, the system aims to create more atomic and independently effective perturbations. This independence makes the perturbations more likely to retain their disruptive effect even when the internal mechanisms of an unknown SVC model process them differently, thereby significantly enhancing transferability.
By addressing these three challenges with a sophisticated combination of dual disruption objectives, psychoacoustic masking, and robust transferability techniques, SongBsAb provides a comprehensive and technically sound proactive defense against AI-based illegal song covers.
Demo / Proof of Concept
▶ Watch: Enhancing transferability to unknown singing voice conversion models (8:00)
The efficacy of SongBsAb was demonstrated through a combination of audio examples, quantitative metrics, and a human study, providing compelling evidence of its capabilities.
The presentation included audio demos to illustrate the direct impact of SongBsAb. The process involved showcasing:
- An input voice providing timbre information (the "target" voice).
- An input voice providing lyric and melody information (the "source" song).
- The resultant song when SVC is applied without SongBsAb, demonstrating a high-quality, undisturbed cover.
- Crucially, the resultant song when SVC is applied to the protected version of the input, showcasing the disruptive effects of SongBsAb. Listeners were able to perceive the altered identity or corrupted lyrics in the protected version, contrasting sharply with the clean, unprotected cover. While the specific audio examples aren't transcribed, the speaker emphasized the audible disruption.
Beyond auditory demonstration, quantitative metrics were employed to rigorously evaluate SongBsAb's performance:
- Identity Similarity: This metric quantifies how closely the output song's voice resembles the target singer's voice. The results, highlighted in a red box in the presentation, showed that after applying SongBsAb, the identity similarity dropped significantly. This directly confirms the success of the identity disruption mechanism.
- Lyric Word Error Rate (WER): This metric measures the accuracy of the lyrics in the output song compared to the original. A higher WER indicates more errors or corruption. As highlighted in a green box, applying SongBsAb led to a significant increase in the lyric word error rate, validating the effectiveness of the lyric disruption component.
A human study was also conducted to evaluate the most critical aspect for practical adoption: the impact of SongBsAb's perturbations on song quality and listener enjoyment. Participants were asked whether they perceived any background noise in a given song and, if so, how that noise interfered with their enjoyment. The results, visually represented by a green bar in the presentation, indicated that most participants did not perceive the perturbations as impacting their enjoyment of the songs. This finding is vital, as it confirms that SongBsAb achieves its disruptive effects without compromising the artistic integrity or commercial viability of the original music.
Finally, the transferability of SongBsAb was empirically evaluated. The study demonstrated that both the encoder ensemble and the frame-level interaction reduction loss independently enhanced the disruption effects for both identity and lyric disruption. The combination of these two techniques yielded the best overall results, confirming that SongBsAb's design effectively generalizes its protective capabilities to unknown SVC models, a crucial feature for long-term effectiveness.
Defensive Implications
▶ Watch: Audio demonstration of identity and lyric disruption effects (10:00)
SongBsAb introduces a paradigm shift in how the music industry can protect its intellectual property and artists' rights against the misuse of Singing Voice Conversion (SVC) technology. Its implications for cybersecurity and intellectual property defense are profound and far-reaching.
Firstly, SongBsAb champions a proactive defense strategy. Instead of the reactive "whack-a-mole" approach of detecting and removing infringing content after it has been published—a method proven to be inefficient and often too late—SongBsAb allows rights holders to embed protective measures at the source. By applying imperceptible perturbations to a song before its official release, any subsequent attempt to create an AI cover using that protected source material will inherently result in a flawed or unusable output. This fundamentally shifts the power dynamic back towards content creators and owners.
Secondly, the dual-layer protection offered by SongBsAb—targeting both identity disruption and lyric disruption—provides a robust and comprehensive defense. Identity disruption safeguards the unique vocal characteristics and reputation of artists, preventing their voices from being misappropriated or associated with undesirable content. Lyric disruption protects the fundamental copyright of the lyrical composition, ensuring that the original message and artistic intent remain intact. This multi-faceted approach makes it significantly harder for malicious actors to produce high-fidelity, commercially viable illegal covers.
Thirdly, the emphasis on inaudible perturbations through backing track-refined simultaneous masking is critical for practical adoption. This ensures that the defensive measures do not compromise the artistic quality or commercial appeal of the original song. Rights holders can deploy SongBsAb without fear of alienating their audience or diminishing the value of their product, making it a commercially viable and ethically sound solution.
Fourthly, transferability to unknown SVC models is a cornerstone of SongBsAb's long-term viability. The AI landscape is rapidly evolving, with new SVC models and techniques emerging constantly. By designing perturbations that generalize across different models through encoder ensembles and interaction reduction, SongBsAb offers a future-proof defense mechanism, reducing the need for constant updates to counter every new SVC breakthrough. This makes it a strategic investment for rights holders.
Finally, SongBsAb empowers artists and record companies with a tangible tool to enforce their copyrights and civil rights in the age of generative AI. It provides a technical countermeasure to the erosion of singer competitiveness and the unfair competition posed by cheap, AI-generated covers. By making unauthorized covers audibly defective, SongBsAb can deter infringement, uphold the value of artistic creation, and help maintain the economic viability of the music industry. This proactive technological intervention offers a crucial layer of security, allowing artists to maintain control over their creative output and intellectual property in an increasingly automated world.
Key Takeaways
- Proactive Prevention is Key: SongBsAb shifts the paradigm from reactive detection to proactive prevention against AI-based illegal song covers, embedding protective measures before content release.
- Dual-Layer Disruption: The system offers a robust defense by simultaneously disrupting both the singer's identity (making the output sound like a different gender) and the lyrical content (introducing word errors).
- Acoustically Transparent Protection: Perturbations are designed to be inaudible to human listeners using backing track-refined simultaneous masking, preserving the original song's quality and enjoyment.
- Resilience to Evolving AI: SongBsAb achieves transferability to unknown Singing Voice Conversion (SVC) models through encoder ensembles and frame-level interaction reduction loss, ensuring long-term effectiveness.
- Empowering Rights Holders: This approach provides artists and record companies with a powerful tool to protect their copyrights, civil rights, and commercial interests against AI misuse.
- Quantifiable Impact: Evaluation shows significant drops in identity similarity and increases in lyric word error rates in converted songs, confirming the effectiveness of the disruption.
About the Speaker(s)
Guangke Chen is the presenter of the paper "SongBsAb: A Dual Prevention Approach against Singing Voice Conversion based Illegal Song Covers" at the NDSS Symposium. Based on the transcript, he presented the work on behalf of a team. His presentation demonstrates expertise in the fields of AI-generated content, audio security, and proactive defense mechanisms against emerging threats in the music industry.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Solid applied ML security research that adapts adversarial perturbation techniques to a genuinely underserved problem — proactive poisoning of SVC pipelines before a song ships. The dual-objective framing (identity + lyric disruption) and the psychoacoustic masking refinement are real contributions, but the work sits squarely in the adversarial-examples-for-audio lineage and doesn't break new ground methodologically. Competent, publishable, conference-filler quality.
Heather Calloway (CISO) — WEAK
Technically credible adversarial audio research with a real IP protection use case, but this talk never crosses the threshold from research paper to operational guidance. The people who own this problem — legal, licensing, rights management, and security leadership at music labels and streaming platforms — leave with no decision path.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025