BadVFL: Backdoor Attacks in Vertical Federated Learning

Mohammad Naseri, Yufei Han, Emiliano De Cristofaro

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

This talk, "BadVFL: Backdoor Attacks in Vertical Federated Learning," presented by Mohammad Naseri, Yufei Han, and Emiliano De Cristofaro, delves into a novel class of adversarial attacks targeting Vertical Federated Learning (VFL) systems. Federated Learning (FL) has emerged as a crucial privacy-preserving machine learning paradigm, allowing multiple parties to collaboratively train a shared model without directly exchanging their sensitive raw data. While the security implications in Horizontal Federated Learning (HFL) have been extensively studied, VFL, with its distinct architectural and data distribution characteristics, has received comparatively limited attention regarding robustness attacks.

Watch on YouTube

Visual summary for BadVFL: Backdoor Attacks in Vertical Federated Learning by Mohammad Naseri, Yufei Han, Emiliano De Cristofaro
Visual summary for BadVFL: Backdoor Attacks in Vertical Federated Learning by Mohammad Naseri, Yufei Han, Emiliano De Cristofaro

Key moments

  1. 2:40 Vertical Federated Learning (VFL) architecture
  2. 4:00 Challenges of backdoor attacks in VFL
  3. 4:40 BadVFL: Two-stage attack pipeline explained
  4. 6:00 Illustrative example of trigger insertion
  5. 6:40 Key experimental factors for BadVFL
  6. 7:40 Optimal class selection increases attack rate
  7. 8:20 Effect of poisoning budget on attack success
  8. 8:50 Trigger size and position impact attack success

BadVFL: Backdoor Attacks in Vertical Federated Learning

Speakers: Mohammad Naseri, Research Scientist at FL Labs, PhD student at University College London; Yufei Han; Emiliano De Cristofaro

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=wmcsIj09Ux0

Overview

This talk, "BadVFL: Backdoor Attacks in Vertical Federated Learning," presented by Mohammad Naseri, Yufei Han, and Emiliano De Cristofaro, delves into a novel class of adversarial attacks targeting Vertical Federated Learning (VFL) systems. Federated Learning (FL) has emerged as a crucial privacy-preserving machine learning paradigm, allowing multiple parties to collaboratively train a shared model without directly exchanging their sensitive raw data. While the security implications in Horizontal Federated Learning (HFL) have been extensively studied, VFL, with its distinct architectural and data distribution characteristics, has received comparatively limited attention regarding robustness attacks.

The core of this research is the introduction of BadVFL, a sophisticated backdoor attack specifically designed to exploit the unique challenges and opportunities within VFL environments. Unlike HFL, where attackers might directly manipulate loss functions or labels, VFL's split model architecture and data partitioning present significant hurdles for adversaries. Naseri and his co-authors address these challenges by proposing a two-stage attack pipeline that successfully injects malicious triggers into a VFL model, causing it to misbehave on specific, triggered inputs while maintaining high accuracy on clean data.

The significance of BadVFL lies in its demonstration of a critical vulnerability in a widely adopted privacy-preserving machine learning framework. By showcasing how a malicious participant can subtly influence the global model to perform specific, attacker-defined misclassifications, the research highlights the urgent need for robust defensive mechanisms tailored to VFL. The findings presented in this talk are vital for developers, researchers, and practitioners deploying VFL systems, emphasizing that privacy guarantees do not inherently translate to robustness against sophisticated adversarial manipulations.

Background

▶ Watch: Vertical Federated Learning (VFL) architecture (2:40)

Federated Learning (FL) represents a paradigm shift in collaborative machine learning, designed to address growing concerns about data privacy and regulatory compliance. Instead of centralizing raw data, FL enables multiple participants to train a shared machine learning model locally on their respective datasets. Only model updates (e.g., gradients or model parameters) are exchanged with a central server, which then aggregates these updates to refine the global model. This iterative process allows for continuous model improvement without exposing sensitive user data.

FL is broadly categorized into three types based on the homogeneity of data across participants:

  • Horizontal Federated Learning (HFL): Participants share the same feature space but differ in their sample space. A classic example involves hospitals training a medical diagnosis model; they all have similar patient records (features like age, blood pressure) but different patient populations (samples).
  • Vertical Federated Learning (VFL): Participants share the same sample space but possess different feature spaces. This scenario is common when different organizations hold complementary information about the same entities. For instance, a bank (financial data) and an e-commerce company (purchase history) might collaborate on a credit scoring model for their shared customer base.
  • Federated Transfer Learning: Data sets differ in both sample and feature space, often requiring more complex alignment and transfer learning techniques.

This talk specifically focuses on Vertical Federated Learning (VFL) with model splitting, which is the more common and complex variant. In this architecture, the central server hosts the top model, typically a classifier, while each participant hosts a bottom model. These bottom models act as encoders, transforming raw features from each participant's local data into feature embeddings. These embeddings are then sent to the server, where the top model uses them as input to map to the corresponding class labels. A crucial initial step in VFL is user ID alignment, a secure process to match and align common entities (e.g., customers) across different participants' datasets.

While FL is primarily a data minimization approach designed to enhance privacy, the exchange of model updates can become a vector for malicious attacks. These attacks generally fall into two categories: privacy attacks (e.g., inferring sensitive data from model updates) and robustness attacks (e.g., causing the model to misbehave). Prior research has predominantly concentrated on robustness attacks, particularly backdoor attacks, within HFL environments. However, the unique architectural constraints of VFL, where a malicious participant does not have direct access to the global loss function, the training labels used by the server, or the ability to directly modify the server's top model, make backdoor attacks significantly more challenging to conceive and execute. The BadVFL research addresses this gap, providing the first comprehensive investigation into backdoor vulnerabilities in VFL.

Key Findings

▶ Watch: BadVFL: Two-stage attack pipeline explained (4:40)

The research presented in "BadVFL: Backdoor Attacks in Vertical Federated Learning" uncovers critical vulnerabilities in VFL systems, demonstrating that these collaborative learning paradigms are susceptible to sophisticated backdoor attacks despite their inherent privacy-preserving design. The key findings are multifaceted, spanning the successful execution of novel attack strategies, the identification of optimal attack parameters, and the evaluation of potential defensive mechanisms.

The primary contribution is the development of BadVFL, a two-stage backdoor attack pipeline that effectively circumvents the architectural challenges of VFL. This attack enables a malicious participant, without direct access to global labels or the server's loss function, to inject a hidden trigger into the shared model. Upon encountering an input with this trigger, the model is coerced into misclassifying it to a specific target class, while maintaining high accuracy on legitimate, untriggered data. This demonstrates that VFL's data minimization principles do not inherently guarantee robustness against such targeted manipulations.

The study systematically explores several factors influencing the attack's success rate and its impact on the model's main task accuracy. Key findings include:

  • Optimal Source/Target Class Selection: The attack's effectiveness is significantly enhanced by strategically selecting source and target classes based on their average pair distance in the embedding space. Classes that are naturally closer in the embedding space (e.g., "dog" and "cat" in CIFAR-10) yield higher attack success rates compared to randomly chosen, more distant classes (e.g., "airplane" and "automobile").
  • Poisoning Budget: As expected, increasing the poisoning budget (the percentage of attacker-controlled data points injected with the trigger) directly correlates with a higher attack success rate. However, this also leads to a more noticeable decrease in the model's main task accuracy, indicating a trade-off between attack efficacy and stealth.
  • Trigger Characteristics: The size of the trigger (e.g., 3x3 vs. 5x5 pixels) also impacts the attack. Larger triggers generally result in higher attack success rates but may also be more easily detectable and potentially degrade main task accuracy.
  • Trigger Placement Strategy: Employing a saliency map approach for trigger placement, which identifies regions with high average gradient magnitudes, proves more effective than random placement. This strategic placement ensures the trigger is inserted into visually salient or semantically important parts of the input, maximizing its impact on the model's decision-making.
  • Collective Attacks: The research also explores the feasibility of collective backdoor attacks where multiple adversaries collaborate by injecting fragmented sub-triggers. While the specific results are referred to in the paper, the concept highlights a more sophisticated threat model.

Finally, the investigation into countermeasures reveals that defenses effective in HFL, such as Neural Cleanse, are largely ineffective in VFL due to the architectural differences (server only sees embeddings, not raw triggered inputs). However, VFL-specific defenses like Differential Privacy (DP) noise and anomaly detection on feature embeddings show promise in mitigating BadVFL attacks, albeit with potential trade-offs in model utility. These findings underscore the necessity for tailored security solutions for VFL.

Technical Deep Dive

▶ Watch: Key experimental factors for BadVFL (6:40)

The BadVFL attack is a sophisticated, two-stage process designed to overcome the unique challenges of backdoor injection in Vertical Federated Learning environments. Unlike horizontal FL, where a malicious client might directly manipulate the global model's loss function or inject poisoned labels, VFL's split model architecture means the server controls the top-layer classifier and labels, while clients only operate on bottom-layer encoders and feature embeddings.

BadVFL Two-Stage Attack Pipeline

The attack unfolds in two distinct stages:

Stage 1: Model Extraction

The primary goal of this stage is for the malicious participant to gain an understanding of the server's top model's behavior, even without direct access to its architecture or training labels. This is achieved by constructing a surrogate model that mimics the server's classifier.

  1. Auxiliary Data Collection: The attacker first collects or generates a dataset of labeled or partially labeled auxiliary data instances. Crucially, this auxiliary data must share the same label space as the training samples used by the server. This could be public datasets or data the attacker has access to.
  2. Feature Embedding Generation: The attacker feeds these auxiliary data instances through its own bottom model (which it controls). This process generates feature embeddings, which are the same type of output that the server receives from legitimate clients.
  3. Surrogate Classifier Training: Using these generated feature embeddings and their corresponding labels from the auxiliary data, the attacker trains a classifier. This classifier acts as a surrogate for the server's top model, allowing the attacker to understand how different feature embeddings map to class labels.
  4. Source and Target Class Selection: Based on the insights gained from the surrogate model, the attacker selects a source class (e.g., "stop sign") and a target class (e.g., "speed limit"). The objective of the backdoor will be to force inputs from the source class, when embedded with a trigger, to be misclassified as the target class. The research highlights that optimal selection involves choosing classes with the least average pair distance in the embedding space, as these are more easily confused by the model.

Stage 2: Backdoor Trigger Insertion (Clean-Label Strategy)

With the surrogate model in place and source/target classes identified, the attacker proceeds to inject the backdoor. This stage employs a clean-label backdoor strategy, meaning the attacker does not need to alter the ground truth labels of the poisoned data. Instead, the trigger itself is designed to manipulate the feature space.

  1. Trigger Injection: The attacker selects a subset of its local training data belonging to the chosen source class. For each selected instance, a trigger signal (e.g., a green square in an image) is inserted into the raw data.
  2. Embedding Manipulation: The core idea is to manipulate the feature embeddings produced by the attacker's bottom model. The goal is to force the feature embeddings of the poisoned samples (source class with trigger) to move closer in the embedding space to the feature embeddings of legitimate training instances belonging to the target class. When these manipulated embeddings reach the server's top model, the model, having been trained on legitimate target class embeddings, will then misclassify the triggered source class input as the target class.
  3. Iterative Training: During the VFL training process, the attacker continually sends these poisoned, triggered embeddings to the server. The server's top model updates based on these and other clients' embeddings. Over iterations, the server's model inadvertently learns the association between the trigger and the target class, effectively embedding the backdoor.

An illustrative example involves two participants each holding half of an image. The attacker injects a trigger into its half of an image (e.g., a dog image). The goal is to make the global model label this triggered image as a "cat," even though the original, untriggered image is a "dog." The attacker's bottom model processes its triggered half, producing embeddings that, when combined with the other participant's embeddings, nudge the overall representation towards the "cat" class in the server's classifier.

Experimental Setup and Factors

The researchers validated BadVFL using two-party and multi-party settings across several standard datasets: CIFAR-10, CIFAR-100 (for image classification), and KDD (likely KDD Cup 99 or a similar network intrusion detection dataset for tabular data). They systematically investigated four key factors influencing the attack's efficacy:

  1. Attack Round and Label Inference: The researchers explored at which training round the attack should commence and when the label inference (from the surrogate model) should occur.
  2. Source and Target Class Selection: This factor proved highly significant.
  • Random Selection: For CIFAR-10, an example was selecting "airplane" (source) and "automobile" (target).
  • Optimal Selection: The attacker computes the average pair distance between training data points in the embedding space for all possible class pairs. The pair with the least average pair distance is chosen, as these classes are semantically or visually closer, making the backdoor more potent and stealthy. For CIFAR-10, "dog" and "cat" were identified as an optimal pair. Experiments showed that optimal selection significantly increased the attack success rate.
  1. Poisoning Budget: This refers to the percentage of data points within the attacker's local dataset that are poisoned with the trigger. As expected, increasing the poisoning budget generally increased the attack success rate but concurrently decreased the main task accuracy of the global model, indicating a trade-off between attack strength and detectability.
  2. Trigger Characteristics (Size and Position):
  • Trigger Size: For image datasets, experiments with trigger sizes like 3x3 pixels and 5x5 pixels were conducted. Larger trigger sizes generally led to an increased attack success rate but also a decreased main task accuracy.
  • Trigger Position:
  • Random Position: Triggers were placed randomly within the image.
  • Saliency Map Approach: The adversary computed a Jacobian-based saliency map of the training images using its locally trained surrogate classifier. It then used a sliding window (e.g., 3x3 or 5x5) to scan over a source training image and selected the window with the highest average gradient magnitudes to inject the backdoor trigger. This ensures the trigger is placed in a region that the model is highly sensitive to, maximizing its impact.

The research also touched upon collective backdoor attacks, where a trigger pattern is broken down into sub-trigger patterns, and each sub-trigger is injected by a different malicious participant. During the testing phase, all sub-triggers are injected in parallel to execute the attack, demonstrating a more complex and coordinated threat.

Demo / Proof of Concept

▶ Watch: Optimal class selection increases attack rate (7:40)

While the talk did not feature a live software demonstration or a publicly released tool, the research effectively served as a comprehensive proof of concept through its rigorous experimental validation. The methodologies described for the BadVFL attack were implemented and tested across multiple datasets and under various configurations, providing concrete evidence of its feasibility and effectiveness.

The presentation included visual examples to illustrate the concept. For instance, a green square trigger was shown injected into a stop sign image, with the intent of misclassifying it as a speed limit sign. A more detailed example for VFL specifically depicted an image split between two participants, with the attacker injecting a trigger into their half of the image to force a mislabeling (e.g., a "dark image" being labeled as a "cat"). The extensive results presented, detailing attack success rates, main task accuracy, and the impact of various parameters like poisoning budget, trigger size, and optimal class selection, all collectively serve as the empirical demonstration of BadVFL's capabilities. This thorough experimental analysis, rather than a live demo, constitutes the primary proof of concept for the presented attack.

Defensive Implications

▶ Watch: Trigger size and position impact attack success (8:50)

The successful demonstration of BadVFL highlights a critical blind spot in VFL security and necessitates the development of tailored defensive strategies. The research evaluated three potential countermeasures, revealing that traditional HFL defenses may not directly translate to the VFL context.

  1. Neural Cleanse: This defense mechanism, popular in HFL, typically identifies and prunes candidate neurons in a neural network that are highly involved in backdoor attacks by analyzing the network's behavior on different inputs. However, the BadVFL research found Neural Cleanse to be ineffective against their proposed attack. The primary reason is architectural: in VFL with model splitting, the server only receives feature embeddings from clients, not the raw input images or data containing the physical trigger. Since Neural Cleanse operates on the network that processes the raw input, it cannot detect or mitigate triggers embedded at the client-side before the embedding stage. This underscores the need for VFL-specific defenses that operate on the feature embedding space.
  1. Differential Privacy (DP) Noise: A more promising approach involves integrating Differential Privacy (DP) by adding noise to the feature embeddings submitted by the clients to the server. The idea is that injecting random noise can obfuscate the subtle patterns introduced by the backdoor trigger in the embeddings. The experiments showed that when the variance of Gaussian noise added to the feature embeddings was increased, the attack success rate decreased. This indicates that DP noise can indeed serve as a countermeasure. However, the inherent trade-off with DP is that increasing noise for stronger privacy (and defense against backdoors) typically leads to a degradation in the model's overall utility and main task accuracy. The optimal balance between defense effectiveness and model performance needs careful calibration.
  1. Anomaly Detection: The third defense considered is performing anomaly detection over the feature embeddings of each class. In the BadVFL attack, the adversary replaces some of the training data points belonging to the target class with triggered source class samples that have been manipulated to appear as the target class. This manipulation might cause the feature embeddings of these poisoned samples to deviate statistically from the distribution of legitimate feature embeddings for that class. The server could monitor the incoming feature embeddings from each participant and flag any submissions that appear anomalous within their declared class. The research demonstrated that on CIFAR-10 datasets, using an anomaly detection countermeasure successfully decreased the attack success rates. This suggests that monitoring the statistical properties of feature embeddings, particularly within class distributions, can be an effective way to detect and potentially mitigate backdoor injections in VFL.

In summary, the defensive implications are clear: VFL's unique architecture requires bespoke security solutions. While generic HFL defenses like Neural Cleanse are largely inapplicable, methods like DP noise and anomaly detection, which operate directly on the exchanged feature embeddings, offer viable pathways to enhance VFL model robustness against sophisticated backdoor attacks like BadVFL. Further research is needed to develop more robust, low-overhead, and universally applicable VFL-specific defenses.

Key Takeaways

  • Vulnerability of VFL: Vertical Federated Learning, despite its privacy-preserving nature, is susceptible to sophisticated backdoor attacks like BadVFL, challenging previous assumptions about its inherent robustness.
  • Two-Stage Attack Pipeline: BadVFL effectively bypasses VFL's architectural constraints by employing a novel two-stage approach: a model extraction stage to build a surrogate classifier and a clean-label trigger insertion stage to manipulate feature embeddings.
  • Impact of Attack Parameters: The success rate of BadVFL is significantly influenced by factors such as optimal source/target class selection (based on embedding distance), the poisoning budget, trigger size, and strategic trigger placement using saliency maps.
  • Clean-Label Efficacy: The research demonstrates the feasibility and effectiveness of clean-label backdoor attacks in VFL, where the attacker does not need to alter the ground truth labels of poisoned data.
  • Ineffectiveness of HFL Defenses: Traditional backdoor defenses designed for Horizontal Federated Learning, such as Neural Cleanse, are largely ineffective in VFL due to the server's limited visibility into raw client data.
  • Promising VFL-Specific Countermeasures: Defenses operating on feature embeddings, like Differential Privacy (DP) noise and anomaly detection, show potential in mitigating BadVFL attacks, though they may involve trade-offs with model utility.

About the Speaker(s)

Mohammad Naseri is a research scientist at FL Labs and a PhD student at University College London. His research focuses on the security and privacy aspects of federated learning, with a particular emphasis on understanding and mitigating adversarial attacks in these collaborative machine learning environments.

Yufei Han and Emiliano De Cristofaro are co-authors on this paper. While their specific affiliations and titles were not detailed in the provided transcript, their collaboration with Mohammad Naseri on this research, presented at a prestigious security conference like IEEE S&P, indicates their expertise and contributions to the field of privacy and security in machine learning.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This presentation on BadVFL delivers a critical, novel backdoor attack against Vertical Federated Learning, a paradigm often assumed robust due to its architectural split. The researchers demonstrate a sophisticated two-stage, clean-label technique that bypasses VFL's unique constraints, proving traditional HFL defenses are inadequate. This work is essential for anyone deploying or researching VFL, exposing a significant vulnerability that demands immediate attention.

Heather Calloway (CISO) — STRONG ACCEPT

This research exposes a critical integrity risk in Vertical Federated Learning, demonstrating how privacy-preserving ML can still be backdoored by malicious participants. It demands immediate attention from leaders deploying VFL systems, as privacy guarantees do not inherently translate to model robustness.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024