BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting

Huming Qiu, Junjie Sun, Mi Zhang, Xudong Pan, Min Yang

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

This talk introduces BELT, a novel attack technique demonstrating how traditional backdoor attacks can bypass even the most advanced deep learning defense mechanisms by enhancing a property termed "backdoor exclusivity." Presented by Huming Qiu from Fudan University and co-authored with Junjie Sun, Mi Zhang, Xudong Pan, and Min Yang, the research highlights a critical vulnerability in the current deep learning supply chain. With the proliferation of model-sharing platforms like Hugging Face and Model Zoo, the risk of malicious actors disseminating compromised models has escalated, making backdoor attacks a paramount concern for the security of AI systems.

Watch on YouTube

Visual summary for BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting by Huming Qiu, Junjie Sun, Mi Zhang, Xudong Pan, Min Yang
Visual summary for BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting by Huming Qiu, Junjie Sun, Mi Zhang, Xudong Pan, Min Yang

Key moments

  1. 0:00 Introduction to backdoor attacks in model sharing
  2. 3:20 Introducing fuzzy triggers and backdoor exclusivity concept
  3. 4:20 Measuring backdoor exclusivity through perturbation analysis
  4. 6:00 Introducing BELT: A novel attack for enhanced exclusivity
  5. 6:20 BELT's mechanism: Tainted and cover samples
  6. 7:40 Experimental results: BELT significantly enhances exclusivity
  7. 8:00 Evaluating BELT's escapability against state-of-the-art defenses

BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting

Speakers: Huming Qiu, Junjie Sun, Mi Zhang, Xudong Pan, Min Yang

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=r5u89WLayhM

Overview

This talk introduces BELT, a novel attack technique demonstrating how traditional backdoor attacks can bypass even the most advanced deep learning defense mechanisms by enhancing a property termed "backdoor exclusivity." Presented by Huming Qiu from Fudan University and co-authored with Junjie Sun, Mi Zhang, Xudong Pan, and Min Yang, the research highlights a critical vulnerability in the current deep learning supply chain. With the proliferation of model-sharing platforms like Hugging Face and Model Zoo, the risk of malicious actors disseminating compromised models has escalated, making backdoor attacks a paramount concern for the security of AI systems.

The core premise of BELT is to redefine the activation conditions of a backdoor, making it highly specific to its intended trigger while simultaneously rendering "fuzzy triggers"—those similar to the original but not identical—incapable of activation. This approach directly counteracts state-of-the-art defenses that primarily rely on detecting or reverse-engineering such fuzzy triggers. By unveiling this blind spot, the BELT framework provides powerful insights for both advancing backdoor attack sophistication and, crucially, for developing a new generation of more robust defense strategies that can account for highly exclusive malicious functionalities.

The significance of BELT extends beyond academic curiosity, addressing a tangible threat to AI deployment across various sectors. As deep learning models become increasingly integrated into critical infrastructure, from autonomous systems to healthcare diagnostics, ensuring their integrity against subtle, evasive attacks is paramount. This research underscores the urgent need for comprehensive validation mechanisms on model-sharing platforms and a paradigm shift in how AI model security is approached, moving beyond current detection methodologies that are demonstrably insufficient against sophisticated, exclusivity-enhanced backdoors.

Background

▶ Watch: Introduction to backdoor attacks in model sharing (0:00)

The rapid growth of deep learning model sharing platforms, such as Hugging Face and Model Zoo, has revolutionized how AI models are developed and deployed. These platforms allow users to easily access and download pre-trained models, accelerating innovation. However, their open nature and often lacking validation mechanisms create a fertile ground for malicious activities, particularly backdoor attacks. A malicious model provider can leverage these platforms to widely disseminate models embedded with secret, harmful functionalities.

A backdoor attack involves embedding a hidden malicious function into a Deep Neural Network (DNN) model. This functionality remains dormant during normal operation, allowing the model to perform as expected on clean, unpoisoned data. However, when presented with poisoned data containing a specific trigger, the model will reliably misclassify it into a target class chosen by the attacker. For instance, in a face recognition system, a backdoored model might grant unauthorized access to individuals wearing specific glasses, effectively bypassing security protocols. These attacks are particularly insidious because they can occur at various stages of the model supply chain, including data collection, model selection, and model training, making them one of the most severe threats to deep learning systems.

Traditional backdoor attacks typically rely on strong feature triggers, such as random pixel patches or fixed patterns. These triggers are designed to be easily learned by the victim model, ensuring a high attack success rate (ASR), often approaching 100%. Furthermore, strong triggers contribute to the robustness of the backdoor, meaning they remain effective even in the presence of input noise or interference. While robustness enhances the stability and success rate of backdoor attacks in complex environments, excessive robustness can inadvertently activate the backdoor with variations of the trigger, known as approximate triggers, or even provide clues for detection.

This robustness, ironically, has been leveraged by most state-of-the-art backdoor defense systems. These defenses treat robustness as a weakness, employing techniques like reverse engineering, sample superposition, or attribution to identify approximate or fuzzy triggers. Fuzzy triggers are defined as any triggers similar to the original that can still activate the backdoor, even if not identical. Examples include Neural Cleanse and ABS (reverse engineering), STRIP (sample superposition based on entropy), and SentiNet (attribution techniques). While these methods have achieved significant success in identifying such triggers, their reliance on the existence of these fuzzy triggers forms the Achilles' heel that BELT exploits. The core concept explored by the researchers is that backdoor robustness is one side of a coin, describing tolerance to trigger variations, while the other side, backdoor exclusivity, represents the precision of the backdoor's response to triggers.

Key Findings

▶ Watch: Measuring backdoor exclusivity through perturbation analysis (4:20)

The central inquiry of this research was whether backdoor exclusivity could be effectively measured and, more importantly, enhanced. The authors provide an affirmative response, laying the groundwork for a new understanding of backdoor vulnerabilities and defenses. Their key findings are structured around this concept:

First, the research establishes a novel and universal metric for measuring backdoor exclusivity. By conducting comprehensive perturbation analysis on original triggers, they estimate an upper bound for trigger invalidation. This boundary delineates the maximum distortion a trigger can tolerate while still activating the backdoor, effectively mapping the potential range of existence for fuzzy triggers. This metric provides a quantifiable measure of how precisely a backdoor responds to its intended trigger, distinguishing it from general robustness.

Second, applying this new metric, the study reveals that existing state-of-the-art backdoor attacks perform poorly in terms of exclusivity. On average, conventional backdoors exhibited an exclusivity of approximately 17%. This low exclusivity directly impacts their ability to evade contemporary defense mechanisms, which exploit the prevalence of fuzzy triggers. This finding provides critical insights into the limitations of current attack paradigms and highlights the need for more sophisticated approaches.

Third, guided by the concept of backdoor exclusivity, the researchers introduce BELT (Backdoor Exclusivity Lifting), a novel attack technique. BELT is designed to precisely define the activation conditions of the backdoor, rendering fuzzy triggers incapable of activation, thereby significantly enhancing its exclusivity. This technique is described as effective, straightforward to implement, and capable of complementing existing backdoor attacks to further boost their evasiveness.

Finally, experimental evaluations demonstrate BELT's profound impact. BELT effectively elevates backdoor exclusivity from an average of 17% for conventional attacks to a range of 45% to 55%, while crucially maintaining high Clean Data Accuracy (CDA) and Attack Success Rate (ASR). Furthermore, backdoors crafted with BELT successfully evade six prominent backdoor defense systems—Neural Cleanse, ABS, MNTD, MOTH, STRIP, and SentiNet—which would otherwise detect or eliminate them. This comprehensive evasion capability underscores BELT's potency and the critical security gap it exposes.

Technical Deep Dive

▶ Watch: Introducing BELT: A novel attack for enhanced exclusivity (6:00)

The technical novelty of BELT lies in its dual contribution: a quantifiable metric for backdoor exclusivity and a sophisticated poisoning strategy to enhance it.

Measuring Backdoor Exclusivity

The researchers first delved into the interaction between fuzzy triggers and backdoor exclusivity to devise a universal and effective measurement. This involved a comprehensive perturbation analysis on the original trigger. The goal was to estimate the upper bound of perturbation that leads to trigger invalidation. This boundary serves as a crucial indicator of the potential range where fuzzy triggers might exist. It showcases the maximum distortion that a trigger can withstand while still activating the backdoor.

Leveraging this boundary, a metric to measure backdoor exclusivity was introduced. Conceptually, the origin represents the original trigger. Perturbations increase along axes, with red dots denoting fuzzy triggers and blue dots representing invalid triggers. A black circle defines the perturbation boundary, while a red circle signifies the trigger upper bound, beyond which all triggers are guaranteed to be invalid.

  • Low Exclusivity (approaching 0%): If the trigger upper bound approaches the perturbation boundary, it indicates that there's no clear upper limit for invalidating triggers. This implies the existence of numerous fuzzy triggers, making the backdoor highly susceptible to detection by current defenses.
  • High Exclusivity (approaching 100%): Conversely, if the trigger upper bound approaches zero, it suggests that almost all perturbed triggers are invalid. This scenario implies that the backdoor exclusively responds to the original, unperturbed trigger, imparting a unique weakness to the trigger that current defenses struggle to identify.

Upon measuring the exclusivity of existing state-of-the-art backdoors, the researchers found that their exclusivity was generally poor, averaging approximately 17%. This quantitative evidence strongly supported the hypothesis that traditional backdoor attacks lack the precision to evade sophisticated countermeasures.

BELT Attack Technique

Guided by the insights from exclusivity measurement, BELT aims to precisely define the activation conditions of the backdoor, thereby rendering fuzzy triggers ineffective. This is achieved through a refined data poisoning phase during model training.

BELT operates by creating two distinct batches of poison samples:

  1. Tainted Samples (DT): These are standard poisoned samples. They carry the original trigger (e.g., a white circle pattern) and have their labels modified to the attacker's target label. These samples are primarily responsible for injecting the malicious backdoor functionality into the model, ensuring the desired misclassification when the exact trigger is present.
  2. Cover Samples (DC): This is the innovative core of BELT. These samples carry cover triggers (e.g., an incomplete white circle, a slightly perturbed version of the original trigger). Crucially, unlike tainted samples, cover samples retain their true labels. The purpose of cover samples is to explicitly suppress the association between the backdoor functionality and fuzzy triggers during the model's training process. By training the model to correctly classify data with cover triggers (which are essentially fuzzy triggers) under their true labels, BELT forces the model to learn that only the exact original trigger should activate the backdoor, thus strengthening its exclusivity.

When the model is trained on a dataset containing both tainted and cover samples, a highly exclusive backdoor is implanted. The model learns to perform normally on data with fuzzy triggers (due to cover samples) while still misclassifying data with the precise original trigger (due to tainted samples).

BELT was evaluated under two main attack scenarios:

  • Data Outsourcing: The attacker controls the training data, injecting both tainted and cover samples into a dataset used by a victim.
  • Model Outsourcing: The attacker has more control over the training process, potentially allowing for the incorporation of additional loss terms to further amplify BELT's impact on exclusivity.

The method for crafting cover triggers is critical. The research found that masking a portion of the original trigger and randomly generating diverse cover triggers was the most effective approach. Specifically, lower masking rates (meaning cover triggers are very close to the original trigger) led to higher exclusivity. This is because minimizing the connection between the backdoor and high-quality fuzzy triggers (those closest to the original) significantly improves exclusivity. Alternative methods for crafting cover samples, such as random sampling and noise injection, were explored but found to be less effective than the masking-based method in most scenarios.

Demo / Proof of Concept

▶ Watch: Experimental results: BELT significantly enhances exclusivity (7:40)

The talk did not feature a live demonstration in the traditional sense. Instead, the speakers presented extensive experimental validation and empirical results that served as a robust proof of concept for the BELT attack technique and its effectiveness against state-of-the-art defenses.

The experimental setup involved evaluating BELT attacks on image classification tasks, a common domain for backdoor research. The evaluation covered three popular benchmarks (though specific names weren't detailed in the transcript, common ones include CIFAR-10, ImageNet, etc.) and incorporated four classic backdoor attacks as baselines to demonstrate BELT's complementary nature and enhancement capabilities. The notation + denoted BELT attacks under the data outsourcing scenario, while ++ indicated BELT attacks under the more controlled model outsourcing scenario, where attackers might have more influence over the training loss functions.

The results unequivocally demonstrated BELT's capability to enhance backdoor exclusivity while maintaining critical performance metrics:

  • Exclusivity Enhancement: Conventional backdoors typically exhibited an average exclusivity of approximately 17%. BELT successfully elevated this to a range of 45% to 55%, showcasing a significant improvement in the precision of backdoor activation.
  • Performance Maintenance: Crucially, this enhancement in exclusivity was achieved without sacrificing Clean Data Accuracy (CDA) on normal, unpoisoned inputs or the Attack Success Rate (ASR) when the exact trigger was present. This balance is vital for a practical backdoor attack, as any noticeable drop in normal performance or attack reliability would alert defenders.

A key aspect of the proof of concept was BELT's ability to evade existing defense mechanisms. The researchers evaluated BELT against six state-of-the-art backdoor defenses:

  • Neural Cleanse and ABS: These are reverse engineering-based algorithms designed to reconstruct potential backdoor triggers from a target model.
  • MNTD: This defense trains a meta-classifier to score the target model, with higher scores indicating a greater likelihood of a backdoor.
  • MOTH: This method utilizes reverse triggers and natural triggers to achieve model orthogonalization, aiming to eliminate potential backdoors.
  • STRIP: A data-level detection algorithm that uses entropy to determine whether input samples carry triggers.
  • SentiNet: This defense employs attribution techniques to reveal regions within samples that significantly impact classification results, thereby identifying trigger presence.

The experimental results confirmed that backdoors crafted with BELT successfully evaded all six of these defenses, which would otherwise detect or neutralize conventional backdoor attacks. This comprehensive evasion capability is the most compelling evidence of BELT's effectiveness and the critical vulnerability it exposes in current defense paradigms.

Furthermore, the research investigated the perturbation resistance of BELT-crafted backdoors. As noise intensity in input data increased, both CDA and ASR experienced a slight decrease. However, ASR consistently outperformed CDA throughout this decline, indicating that BELT maintains a balanced tradeoff between backdoor exclusivity and perturbation resistance. This suggests that even when defenders attempt to disrupt potential triggers through noise injection, BELT can still maintain a reasonable attack success rate while preserving the model's performance on clean data.

Defensive Implications

▶ Watch: Evaluating BELT's escapability against state-of-the-art defenses (8:00)

The introduction of BELT presents a significant challenge to the current landscape of deep learning security and necessitates a re-evaluation of defensive strategies against backdoor attacks. The core implication is that state-of-the-art backdoor defenses, which primarily rely on the existence of fuzzy triggers, are fundamentally vulnerable to attacks that enhance backdoor exclusivity.

Defenders must recognize that the assumption of trigger robustness, while useful for past detection methods, is no longer universally applicable. Backdoors designed with BELT are surgically precise, activated only by the exact intended trigger, rendering reverse-engineering or attribution techniques that seek approximate triggers ineffective. This means:

  1. Shift in Detection Paradigms: Current detection methods need to evolve beyond fuzzy trigger identification. Future defenses might need to focus on detecting the absence of fuzzy triggers, or identify highly specific activation patterns that are characteristic of BELT-style backdoors. This could involve more sophisticated statistical analysis of model behavior across a wide range of perturbed inputs, looking for an unusual "cliff-edge" activation rather than a gradient.
  2. Focus on Model Provenance and Integrity: With model-sharing platforms acting as vectors for such sophisticated attacks, stronger mechanisms for verifying the provenance and integrity of shared models are crucial. This could include cryptographic attestation, secure enclaves for model training, or rigorous, independent auditing of models before they are made publicly available.
  3. Robust Training and Verification: While BELT showed some resistance to adaptive defenses involving noise injection, further research into robust training techniques that explicitly counter exclusivity-lifting mechanisms could be beneficial. This might involve training models not just against known triggers but also against highly exclusive variations, forcing the model to generalize more broadly or reject overly specific activations.
  4. Novel Exploit Detection: Defenders might need to explore entirely new avenues for detection, such as analyzing the internal representations of the model for anomalies, detecting unusual gradient flows during inference, or employing techniques that don't rely on input-output relationships but rather on the model's intrinsic structure or learning process.
  5. Understanding Attack Complexity: The complexity introduced by BELT, particularly through the use of cover samples to fine-tune exclusivity, underscores the need for defenders to understand the full spectrum of attack methodologies. This includes investigating how attackers might manipulate training data to achieve such precise control over backdoor activation.

In essence, BELT serves as a wake-up call, pushing the boundaries of what constitutes a "stealthy" backdoor. It mandates a shift from reactive defenses targeting known attack characteristics to proactive security measures that anticipate and neutralize more sophisticated, precisely engineered threats.

Key Takeaways

  • Deep learning model sharing platforms, despite their utility, introduce significant security risks by enabling the widespread dissemination of sophisticated backdoor models.
  • Current state-of-the-art backdoor defenses, which primarily rely on identifying "fuzzy triggers" (approximate versions of the original trigger), are vulnerable to novel attack techniques.
  • The research introduces a novel concept: backdoor exclusivity, which measures the precision of a backdoor's response to its trigger, complementing the existing notion of "robustness."
  • BELT (Backdoor Exclusivity Lifting) is a new attack methodology that significantly enhances backdoor exclusivity by employing "cover samples" during model training, forcing the backdoor to respond only to the exact original trigger.
  • BELT successfully elevates backdoor exclusivity from an average of ~17% (for conventional attacks) to 45-55% while maintaining high Clean Data Accuracy (CDA) and Attack Success Rate (ASR).
  • Backdoors crafted using BELT demonstrated comprehensive evasion capabilities, successfully bypassing six prominent backdoor defense mechanisms, highlighting a critical blind spot in current AI security.

About the Speaker(s)

The talk was presented by Huming Qiu, affiliated with Fudan University. Huming Qiu's research, as demonstrated in this presentation, focuses on the security of deep learning systems, particularly in understanding and exploiting vulnerabilities related to backdoor attacks and their evasion of defense mechanisms. The work was a collaborative effort with co-authors Junjie Sun, Mi Zhang, Xudong Pan, and Min Yang, reflecting a collective expertise in the field of AI security and machine learning.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This is a rare example of novel research that directly challenges the foundations of current AI security. BELT's concept of 'backdoor exclusivity' and its precise implementation expose a critical blind spot in state-of-the-art defenses, demonstrating a sophisticated attack that will force a paradigm shift. This isn't just theory; it's a demonstrable threat to every model-sharing platform out there.

Heather Calloway (CISO) — STRONG ACCEPT

This research exposes a critical vulnerability in the AI model supply chain, demonstrating how sophisticated backdoors can evade current state-of-the-art defenses. It mandates a strategic shift in how security leaders approach AI model validation and integrity, urging a move beyond reactive detection paradigms.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024