Backdooring Multimodal Learning
Xingshuo Han, Yutong Wu, Qingjie Zhang, Yuan Zhou, Yuan Xu, Han Qiu
IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 5
Overview
This talk, "Backdooring Multimodal Learning," presented by Xingshuo Han and colleagues from Nanjing Technological University Singapore and Tsinghua University China, delves into the novel and critical area of backdoor attacks against multimodal deep learning models. While deep learning models, including those handling multiple data types, have demonstrated impressive performance across various benchmarks and real-world applications like visual question answering (VQA) and audio-video speech recognition (AVSR), they are increasingly recognized as vulnerable to adversarial manipulation. This research specifically highlights how these complex multimodal systems possess unique vulnerabilities that are not adequately addressed by existing backdoor attack methodologies, which largely focus on unimodal tasks.

Key moments
- 0:00 Introduction to multimodal learning and backdoor attack problem
- 2:00 Uncovering unique multimodal vulnerabilities and attack challenges
- 4:00 Presenting "Bags" score to select optimal poisoning samples
- 5:00 Visualization and explanation of Bags score's early effectiveness
- 6:00 Proposed attack framework and Co-Attack method introduction
- 8:00 Details on Mix-Attack and experimental setup for VQA, AVSR
- 9:00 Evaluation results showing Co-Attack and Mix-Attack effectiveness
Backdooring Multimodal Learning
Speakers: Xingshuo Han, Yutong Wu, Qingjie Zhang, Yuan Zhou, Yuan Xu, Han Qiu
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=MV-C68TMyjI
Overview
This talk, "Backdooring Multimodal Learning," presented by Xingshuo Han and colleagues from Nanjing Technological University Singapore and Tsinghua University China, delves into the novel and critical area of backdoor attacks against multimodal deep learning models. While deep learning models, including those handling multiple data types, have demonstrated impressive performance across various benchmarks and real-world applications like visual question answering (VQA) and audio-video speech recognition (AVSR), they are increasingly recognized as vulnerable to adversarial manipulation. This research specifically highlights how these complex multimodal systems possess unique vulnerabilities that are not adequately addressed by existing backdoor attack methodologies, which largely focus on unimodal tasks.
The core contribution of this work lies in developing data and computation-efficient frameworks for launching effective backdoor attacks on multimodal models. The speakers introduce a new scoring mechanism, the Backdoor Significance (BS) score, and two novel attack strategies, Co-attack and Mix-attack, designed to exploit the intricate interactions between different modalities. The findings reveal surprising insights, such as the fact that poisoning all modalities is not always optimal, and that a modality dominant in normal learning might not necessarily dominate the backdoor attack. This research is crucial for understanding the attack surface of increasingly prevalent multimodal AI systems and for developing robust defenses against sophisticated adversaries.
Background
▶ Watch: Introduction to multimodal learning and backdoor attack problem (0:00)
Multimodal learning systems leverage information from multiple distinct data sources—such as images, text, audio, and video—to enhance model performance and achieve richer understanding. Applications range from visual question answering, where models interpret images based on textual queries, to audio-video speech recognition, which combines auditory and visual cues for improved transcription accuracy, and even social media content classification. The success of these systems in complex tasks has led to their widespread adoption in both academic research and industry.
However, the impressive capabilities of deep learning models, including their multimodal counterparts, are shadowed by their susceptibility to backdoor attacks. A backdoor attack involves an adversary subtly manipulating a subset of the training data. The resulting "compromised" model appears to function normally on clean, unpoisoned inputs, maintaining high accuracy. Yet, when presented with any input containing a specific, pre-defined trigger, the model will reliably produce a targeted, incorrect prediction. This threat is particularly salient in scenarios where AI companies outsource data collection and labeling to third parties, creating a significant risk of malicious data injection into the training pipeline.
Prior research has extensively explored backdoor attacks on unimodal tasks, particularly image classification. These works have devised various powerful trigger designs, including blinded strap, invisible rippling, and reflection in pixel or frequency domains. However, these techniques often fall short when applied to multimodal learning due to the inherent complexities introduced by combining different data types. Multimodal models exhibit heterogeneous contributions from individual modalities and intricate inter-modality dependencies. For instance, simply poisoning both visual and textual modalities simultaneously, as some earlier works on visual question answering (VQA) attempted, might not yield satisfactory attack effectiveness. The speakers emphasize that the structural complexity of multimodal models, with their multiple attack surfaces and potentially unbalanced modality contributions, necessitates a deeper investigation into tailored backdoor attack methodologies. They cite an example where injecting both visual and textual triggers can surprisingly worsen attack effectiveness on audio-video speech recognition (AVSR) tasks, underscoring the non-trivial nature of modality interactions in backdoor scenarios.
Furthermore, existing attack methods often suffer from inefficiency. Many approaches rely on random selection of clean data to poison, implicitly assuming that each poisoning sample contributes equally to the attack, an assumption the authors prove to be false. Other methods, such as those employing a forgetting score to gauge sample importance, typically operate in the late stages of training, making them computationally expensive and time-consuming. Crucially, the applicability and effectiveness of such scores in the multimodal context remained unproven, and this work demonstrates their failure in this domain. The fundamental gap addressed by this research is the lack of data and computation-efficient backdoor attack frameworks specifically designed to exploit the unique vulnerabilities of multimodal learning systems.
Key Findings
▶ Watch: Presenting "Bags" score to select optimal poisoning samples (4:00)
The research presented in "Backdooring Multimodal Learning" uncovers several critical insights into the nature of backdoor attacks on multimodal deep learning models, challenging prior assumptions and paving the way for more effective and efficient attack and defense strategies.
A primary finding is the identification of a novel score, the Backdoor Significance (BS) score, which significantly enhances the efficiency of backdoor attacks. Unlike previous methods like the forgetting score, which operates in late training stages, the BS score can identify optimal poisoning samples much earlier—as early as the 25th epoch or even before. This early identification ability is crucial for achieving both data efficiency (using fewer optimal poisoning samples) and computation efficiency (finding these candidates sooner). The underlying principle is that the "geometry of the training distribution induced by a random victim network contains a surprising amount of information about the structure of model prediction" early in its training phase, making it possible to predict which samples will be most impactful for a backdoor.
The study rigorously evaluates two novel attack methods: Co-attack and Mix-attack. Both consistently outperform baseline strategies like random selection and forgetting score-based methods across diverse multimodal tasks. The Co-attack poisons all modalities of a chosen sample, treating it as a single entity, which significantly reduces the poisoning ratio compared to random selection. The Mix-attack, however, is more sophisticated, taking modality interactions into account by randomly poisoning arbitrary modalities of a sample to find the combination with the highest contribution to the backdoor.
A particularly counter-intuitive finding is that poisoning all modalities is not always better than poisoning only a subset. In some cases, selectively poisoning specific modalities or combinations can yield a stronger backdoor effect. For example, in Visual Question Answering (VQA), poisoning only the question modality might be more effective than poisoning both the question and visual modalities. This highlights the complex interplay of modality complementarity and competition within the context of backdoor attacks. Modalities can either work together to enhance the backdoor or, surprisingly, compete, leading to a diminished effect if all are poisoned indiscriminately.
Furthermore, the research reveals that a modality's dominance in normal model performance does not necessarily translate to dominance in backdoor learning. In VQA, the question modality dominates both the model's overall performance and its backdoor vulnerability. However, in Audio-Video Speech Recognition (AVSR), while audio dominates normal learning, the video modality surprisingly dominates the backdoor performance. This divergence underscores the need for modality-aware attack and defense strategies that consider the specific context of adversarial manipulation.
The study also demonstrates that the forgetting score, a previously proposed method for identifying important samples, fails in multimodal contexts like VQA and AVSR. For VQA, its failure stems from the presence of approximately 3,000 possible answers per question, making it difficult to define and track forgetting. For AVSR, which generates sentences and uses metrics like Word Error Rate (WER) rather than discrete class labels, the forgetting score often results in zero for almost all samples, rendering it ineffective.
Finally, the study explores the impact of trigger characteristics, noting that a smaller visual patch trigger can weaken the effectiveness of poisoning the video modality in AVSR. The proposed methods are also shown to be effective in both white-box (attacker has full knowledge of the model) and black-box (attacker has no knowledge of internal model architecture or parameters) settings, demonstrating their broad applicability. These findings collectively provide a deeper understanding of multimodal backdoor vulnerabilities, offering critical guidance for both attackers seeking efficient methods and defenders aiming to build more robust AI systems.
Technical Deep Dive
▶ Watch: Visualization and explanation of Bags score's early effectiveness (5:00)
The core objective of this research was to construct a data and computation-efficient backdoor attack framework tailored for multimodal learning, while simultaneously uncovering novel insights into these complex systems' vulnerabilities. The attackers' goals were clearly defined: to preserve the victim model's normal performance on clean samples, achieve high attack effectiveness on poisoned samples, minimize the number of poisoned samples, and identify optimal poisoning candidates as early as possible in the training process, all while operating in various adversary capabilities (white-box and black-box settings).
The limitations of prior work served as the starting point for innovation. Existing methods often relied on random selection for poisoning samples, which falsely assumes equal contribution from each sample. This approach is inherently inefficient. More sophisticated methods, such as those leveraging the forgetting score, aimed to identify "hard-to-learn" or "forgotten" samples for poisoning. However, this score is typically computed in the late stages of training, making it computationally expensive. Crucially, the authors empirically prove that the forgetting score is ill-suited for multimodal tasks. In VQA, with its vast answer space (approximately 3,000 possible answers per question), defining and tracking "forgetting" becomes intractable within limited training epochs. For AVSR, which involves sequence generation and uses metrics like Word Error Rate (WER) instead of discrete classification, the forgetting score often yields zero for nearly all samples, rendering it meaningless.
To overcome these limitations, the authors proposed a new metric: the Backdoor Significance (BS) score. Initially, they considered using the gradient norm, but recognized its deficiency in not accounting for the gradient's direction. This led to exploring the projection with the back-gradient on average back-gradient. However, given the heterogeneous contributions of modalities in multimodal learning, a more refined approach was necessary. The final BS score is defined as the gradient norm with consideration of both weights and directions, specifically adapted to account for the differential importance of various modalities.
The rationale behind the BS score's effectiveness stems from observing the model's learning dynamics. By visualizing the "vector loss" for AVSR tasks, the researchers noted that the loss values gradually stabilize around the 25th epoch. This stabilization indicates that before this point, training samples significantly influence the decision boundary of the model. Therefore, scoring samples at or before the 25th epoch captures crucial information about how the training distribution's geometry, induced by a randomly initialized victim network, shapes the model's future predictions. This early-stage scoring allows for efficient identification of high-impact poisoning candidates.
Based on the BS score, a structured selection process was developed. First, a candidate poison set is created by randomly selecting samples from the clean dataset. These candidate samples are then processed, potentially with different combinations of poisoned modalities. The BS score is used to carefully sort and select the most impactful samples. This selection is updated based on the two proposed attack methods: Co-attack and Mix-attack. Finally, the constructed poisoned set, along with the remaining clean samples, is delivered to the victim model for training.
The Co-attack is designed for simplicity and efficiency. It operates by poisoning all modalities for a single chosen sample. This approach treats the entire sample (e.g., an image-question pair, or an audio-video clip) as a single entity for poisoning. This method is easily adaptable to unimodal tasks and significantly reduces the required poisoning ratio compared to random selection. However, experiments revealed that poisoning all modalities is not always the optimal strategy, suggesting the need for a more nuanced approach.
This led to the development of the Mix-attack, which explicitly considers modality interactions. In this attack, the adversary randomly poisons arbitrary combinations of modalities for a given sample. The goal is to identify the specific combination of poisoned modalities within a single sample that yields the highest contribution to the backdoor's strength. For instance, the research found that poisoning only the question modality for a VQA sample could induce a stronger backdoor than poisoning both the question and visual modalities. This flexibility allows attackers to activate the backdoor not just by triggering all modalities, but by exploiting specific, more potent combinations.
Experimental validation was conducted on two representative multimodal tasks: Visual Question Answering (VQA) and Audio-Video Speech Recognition (AVSR).
For VQA, the chosen triggers were a blue cube in the image and the word "consider" in the question, with the target label "w". Initial evaluations with a random selection strategy showed that poisoning the question (Q) modality dominated both VQA performance and backdoor effectiveness, with visual (V) modality poisoning alone proving largely ineffective. Moreover, modality complementarity was observed, where poisoning both Q and V yielded higher attack success rates than poisoning Q or V alone. The forgetting score strategy definitively failed on VQA due to the vast answer space. In contrast, both Co-attack and Mix-attack consistently outperformed other strategies. Interestingly, both random selection and Mix-attack struggled when testing on visual-only poisoned samples. This was attributed to fewer V-only samples being poisoned and the victim model learning joint features, heavily relying on the Q trigger, making visual features almost non-contributory to the backdoor. This underscored the high dominance of the question modality in VQA backdoor performance.
For AVSR, a white cube was used as the video trigger and "high ser" as the audio trigger, with the target prediction "card." Random selection experiments revealed a crucial insight: while audio dominated the model's normal performance, the video modality dominated the backdoor performance. This demonstrated that a modality's importance for normal operation does not necessarily dictate its role in backdoor activation. Both modality complementarity and competition were observed, where modalities sometimes worked together for a stronger backdoor and sometimes competed, leading to diminished returns if all were poisoned. As with VQA, the forgetting score failed for AVSR because the task generates sentences, making traditional classification-based forgetting metrics inapplicable. Again, Co-attack and Mix-attack consistently outperformed baselines. A particularly intriguing finding was that for the Co-attack and random selection, triggering only the video modality could activate the backdoor at a small poisoning ratio. Furthermore, testing on visual-only poisoned samples showed a gradual decrease in attack rate as the poisoning ratio increased. This was analyzed as a scenario where information from one modality (e.g., audio) might become so dominant that the model "surprises or ignores" contributions from other modalities (e.g., video) for the backdoor, leading to reduced effectiveness at higher poisoning rates for the less dominant backdoor modality.
The extended evaluations confirmed that smaller visual patches weakened the effect of poisoning the video modality in AVSR and that the BS score's early-stage computation significantly saved time. Crucially, the proposed methods demonstrated superior performance compared to others in both white-box and black-box settings, highlighting their robustness and practical applicability.
Demo / Proof of Concept
▶ Watch: Details on Mix-Attack and experimental setup for VQA, AVSR (8:00)
While the talk did not feature a live, interactive software demonstration, the detailed experimental evaluations served as a comprehensive proof of concept for the proposed backdoor attacks on multimodal learning. The researchers systematically demonstrated the effectiveness of their Backdoor Significance (BS) score, Co-attack, and Mix-attack methodologies across two distinct and challenging multimodal tasks: Visual Question Answering (VQA) and Audio-Video Speech Recognition (AVSR).
For VQA, the proof of concept involved injecting a blue cube as a visual trigger and the word "consider" as a textual trigger into a subset of training data. The objective was to force the model to mispredict the answer as "w" whenever an input contained these triggers. The experiments meticulously compared the attack success rates of their proposed methods against baseline strategies like random selection and the forgetting score. They showed that Co-attack and Mix-attack consistently achieved higher attack effectiveness while maintaining normal model performance on clean samples. Specific findings, such as the dominance of the question modality in VQA backdoor performance and the observed modality complementarity (where poisoning both Q and V was more effective than Q or V alone), further substantiated the unique aspects of multimodal backdoor vulnerabilities.
Similarly, for AVSR, the proof of concept involved injecting a white cube into the video stream and the phrase "high ser" into the audio stream. The target misprediction for the model was the word "card". These experiments demonstrated that even though audio typically dominates normal AVSR performance, the video modality proved dominant for backdoor activation. The evaluation highlighted scenarios where poisoning all modalities was not superior to poisoning a single, strategically chosen modality, and revealed instances of both modality complementarity and competition. The consistent outperformance of Co-attack and Mix-attack over baselines, even when triggering only the video modality, provided strong evidence for the efficacy and efficiency of their proposed frameworks.
The rigorous experimental setup, including the use of specific triggers, target labels/predictions, and quantitative comparisons of attack success rates and poisoning ratios, provided concrete evidence that multimodal learning models can indeed be backdoored effectively and efficiently using the methods proposed. The detailed analysis of why existing methods fail and how the BS score, Co-attack, and Mix-attack address these shortcomings formed the empirical foundation of the talk's claims.
Defensive Implications
▶ Watch: Evaluation results showing Co-Attack and Mix-Attack effectiveness (9:00)
The findings from "Backdooring Multimodal Learning" carry significant implications for the development of robust defenses against adversarial attacks on AI systems. As multimodal models become increasingly pervasive, understanding and mitigating these unique vulnerabilities is paramount.
Firstly, defenders must acknowledge that multimodal models possess distinct backdoor vulnerabilities that transcend those found in unimodal systems. The complex interplay of modalities, including heterogeneous contributions and inter-modality dependencies, means that simply extending unimodal defense strategies will likely be insufficient. A deeper understanding of how different modalities contribute to or compete during backdoor activation is crucial for crafting effective countermeasures.
One immediate implication concerns data curation and supply chain security. The risk of malicious third parties injecting poisoned data during the labeling process is a primary attack vector. Defenders must implement stringent data vetting protocols, scrutinizing not only individual data points but also the combinations of modalities within them. This includes developing automated tools to detect unusual patterns, subtle trigger injections, or anomalous correlations between modalities that might indicate poisoning. Given that the BS score can identify impactful poisoning candidates early in training, future defense mechanisms could potentially leverage similar early-stage analytics to flag suspicious training samples or learning dynamics.
Secondly, the research highlights the importance of modality-specific monitoring and analysis. Since a modality dominant in normal learning might not be dominant in backdoor learning (e.g., audio vs. video in AVSR), defenders cannot simply focus on the most "important" modality for overall performance. Instead, they need to develop tools that inspect how each modality, and their various combinations, influences the model's decision-making process, especially when potential triggers are present. This could involve explainable AI (XAI) techniques tailored for multimodal inputs to trace the influence of individual modalities on a prediction.
Thirdly, the insights into modality complementarity and competition suggest that robust training paradigms need to go beyond simply making models resilient to individual modality attacks. Defenders should explore multimodal adversarial training that exposes models to diverse, multi-modality triggers and combinations, forcing them to learn more robust, disentangled representations. Techniques that reduce reliance on spurious correlations between modalities and triggers, or that encourage independent processing of modality-specific features before fusion, could also enhance resilience.
Finally, for deployed multimodal models, robust input sanitization and anomaly detection are critical. Systems should be designed to detect and flag inputs that contain known or suspected trigger patterns across any of their constituent modalities. This might involve applying anomaly detection algorithms independently to each incoming modality stream, as well as to their combined representations, to identify adversarial inputs before they can induce malicious behavior. The fact that the proposed attacks work in black-box settings means that defenses cannot rely on internal model knowledge and must focus on input-output behavior and data integrity.
In summary, defending against multimodal backdoor attacks requires a holistic approach that integrates enhanced data supply chain security, modality-aware monitoring and interpretability, advanced multimodal adversarial training, and robust runtime input validation. The research underscores that ignoring the unique complexities of multimodal interactions will leave these powerful AI systems dangerously exposed.
Key Takeaways
- Multimodal deep learning models possess unique and complex backdoor vulnerabilities that are not adequately addressed by existing unimodal attack methods.
- The proposed Backdoor Significance (BS) score enables data and computation-efficient identification of optimal poisoning samples, allowing effective backdoor attacks to be constructed from early training stages (e.g., 25th epoch).
- The novel Co-attack and Mix-attack strategies consistently outperform prior methods (random selection, forgetting score) in backdooring multimodal models like VQA and AVSR.
- Counter-intuitively, poisoning all modalities is not always optimal; partial or specific modality poisoning (e.g., only the question modality in VQA) can sometimes achieve stronger backdoor effects.
- A modality's dominance in normal model performance does not guarantee its dominance in backdoor learning (e.g., audio dominates AVSR learning, but video dominates AVSR backdoor learning). This highlights the need for modality-aware security analysis.
- Defenders must adopt multimodal-specific strategies for data vetting, model monitoring, and robust training, recognizing the intricate interplay of modality complementarity and competition in adversarial contexts.
About the Speaker(s)
The research presented in "Backdooring Multimodal Learning" was a collaborative effort by Xingshuo Han, Yutong Wu, Qingjie Zhang, Yuan Zhou, Yuan Xu, and Han Qiu. The work was primarily conducted at Nanjing Technological University Singapore and Tsinghua University China, with additional support from DC Systems Singapore. Xingshuo Han was the presenter of this work at the IEEE S&P conference. Their collective expertise lies in the field of deep learning security, with a specific focus on understanding and exploiting vulnerabilities within advanced AI models, particularly those operating on multimodal data. Their research contributes significantly to the growing body of knowledge on adversarial machine learning, emphasizing the practical implications for real-world multimodal applications.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This work delivers critical, novel research into backdoor attacks on multimodal deep learning, introducing an efficient scoring mechanism and two potent attack strategies. The counter-intuitive findings on modality interaction and dominance shifts fundamentally advance our understanding of these complex systems' vulnerabilities. This is precisely the kind of deep, actionable insight the community needs.
Heather Calloway (CISO) — STRONG ACCEPT
This research thoroughly exposes critical backdoor vulnerabilities in multimodal AI, demonstrating how attackers can efficiently compromise models even in black-box scenarios. It highlights that traditional unimodal defenses are insufficient and that specific modality interactions must be considered. This work provides essential insights for CISOs and security leaders to reassess AI supply chain risks and develop tailored defensive strategies.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024