FLShield: A Validation Based Federated Learning Framework to Defend Against Poisoning Attacks

Ehsanul Kabir, Zeyu Song, Md Rafi Ur Rashid, Shagufta Mehnaz

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

This talk introduces FLShield, an innovative framework designed to enhance the security of Federated Learning (FL) systems against a spectrum of poisoning attacks. Presented by Ehsanul Kabir from Pennsylvania State University, alongside co-authors Zeyu Song, Md Rafi Ur Rashid, and Shagufta Mehnaz, FLShield tackles the critical challenge of maintaining data privacy and computational efficiency while safeguarding the integrity of global models. Federated Learning is rapidly transforming data analysis in safety-critical domains, from healthcare to finance, by enabling collaborative model training without direct access to raw, sensitive user data. However, this decentralized paradigm introduces significant vulnerabilities, particularly to malicious participants who can inject poisoned data or model updates.

Watch on YouTube

Visual summary for FLShield: A Validation Based Federated Learning Framework to Defend Against Poisoning Attacks by Ehsanul Kabir, Zeyu Song, Md Rafi Ur Rashid, Shagufta Mehnaz
Visual summary for FLShield: A Validation Based Federated Learning Framework to Defend Against Poisoning Attacks by Ehsanul Kabir, Zeyu Song, Md Rafi Ur Rashid, Shagufta Mehnaz

Key moments

  1. 0:00 Introduction to FL and poisoning attack types
  2. 2:00 Drawbacks of current Federated Learning defenses
  3. 3:40 Two key dilemmas in Federated Learning validation
  4. 4:40 Representative Models for secure client-side validation
  5. 6:40 LIPc metric for filtering malicious validation reports
  6. 7:20 Step-by-step design and workflow of FLShield
  7. 8:00 Evaluation methodology and comparison baselines

FLShield: A Validation Based Federated Learning Framework to Defend Against Poisoning Attacks

Speakers: Ehsanul Kabir; Zeyu Song; Md Rafi Ur Rashid; Shagufta Mehnaz

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=pCmWKOvjfBc

Overview

This talk introduces FLShield, an innovative framework designed to enhance the security of Federated Learning (FL) systems against a spectrum of poisoning attacks. Presented by Ehsanul Kabir from Pennsylvania State University, alongside co-authors Zeyu Song, Md Rafi Ur Rashid, and Shagufta Mehnaz, FLShield tackles the critical challenge of maintaining data privacy and computational efficiency while safeguarding the integrity of global models. Federated Learning is rapidly transforming data analysis in safety-critical domains, from healthcare to finance, by enabling collaborative model training without direct access to raw, sensitive user data. However, this decentralized paradigm introduces significant vulnerabilities, particularly to malicious participants who can inject poisoned data or model updates.

The core contribution of FLShield lies in its novel validation-based approach, which strategically leverages client-side validation to detect and mitigate malicious activities. Unlike many existing defenses that rely on unrealistic assumptions about server-side clean data or simple parametric distance metrics, FLShield addresses fundamental dilemmas inherent in decentralized validation. It proposes sophisticated mechanisms, including representative models and a new metric called Loss Impact Per Class (LIPC), to ensure both the privacy of validated models and the integrity of validation reports. This research is crucial for the continued adoption and trustworthiness of Federated Learning in applications where model accuracy and reliability are paramount and compromise could have severe real-world consequences.

Background

▶ Watch: Introduction to FL and poisoning attack types (0:00)

Federated Learning (FL) is a distributed machine learning paradigm that enables multiple clients to collaboratively train a shared global model without exchanging their raw local data. This process unfolds across multiple iterations: a central server dispatches the current global model to a subset of participating clients. Each client then trains a local model on its private dataset and sends back only its local model updates (e.g., gradients or parameter differences) to the server. The server then aggregates these updates to generate an improved version of the global model, which is subsequently distributed for the next round of training. This decentralized nature offers significant privacy benefits, making FL particularly appealing for applications involving sensitive data, such as medical records or financial transactions.

Despite its advantages, FL is highly susceptible to various poisoning attacks launched by malicious participants. These attacks aim to subvert the training process or compromise the integrity of the final global model. The talk categorizes several prominent types of poisoning attacks:

  • Untargeted poisoning attacks: Malicious clients mislabel all or a significant portion of their training data, aiming to reduce the overall accuracy of the global model indiscriminately.
  • Targeted label flipping attacks: Attackers specifically mislabel samples from a source class to a target class. The goal is to cause the global model to misclassify specific inputs in a desired way, for example, classifying a "stop sign" as a "yield sign."
  • Backdoor attacks: Malicious clients introduce specific "backdoor" samples into their training data. These samples are designed to trigger a misclassification to a target class when a specific, often imperceptible, pattern (the "trigger") is present in an input, while the model performs normally on clean inputs.
  • Model poisoning attacks: Attackers craft their local model updates in such a way that, after aggregation, the global model behaves in an attacker-desired manner. This can involve directly manipulating gradients or parameters to introduce vulnerabilities or biases.

Existing defenses against these poisoning attacks often suffer from critical drawbacks, limiting their practical applicability. Some defenses, for instance, assume the central server possesses a set of clean data that matches the distribution of client training data. This is an unrealistic and often impossible assumption in many privacy-sensitive domains, as it violates the core principle of FL. Other defenses rely on the assumption that malicious model updates are parametrically distant from their benign counterparts. These methods typically employ clustering or anomaly detection algorithms to filter out suspicious updates, often by flattening model parameters into a one-dimensional vector and computing distances. However, this assumption breaks down in scenarios with high degrees of non-Independent and Identically Distributed (non-IID) data among benign clients. If benign clients have diverse data distributions, their model updates might naturally be parametrically distant from each other, potentially making them appear "anomalous" even if they are legitimate. Conversely, a sophisticated attacker might craft malicious updates that are parametrically close to benign ones, especially if the malicious client's data distribution partially overlaps with some benign clients, allowing malicious updates to evade detection. These limitations highlight the need for more robust and practical defense mechanisms that can operate effectively under realistic FL constraints.

Key Findings

▶ Watch: Two key dilemmas in Federated Learning validation (3:40)

FLShield presents a groundbreaking approach to defending Federated Learning systems, centered on client-side validation, and achieves several key findings:

  1. Novel Validation-Based Architecture: FLShield introduces a comprehensive validation-based framework where a subset of clients, designated as validators, play a crucial role in assessing the legitimacy of local model updates. This client-side validation paradigm is a fundamental shift from server-centric assumptions, grounding the defense in the real owners of the data.
  1. Resolution of Validation Subject Dilemma with Representative Models: The framework successfully addresses the Validation Subject Dilemma, which highlights the risk of malicious validators launching gradient inversion attacks on local models sent for validation to uncover sensitive training data. FLShield solves this by proposing representative models, which are crafted from local models specifically for validation purposes. These representative models are shown to prevent successful gradient inversion attacks, thereby safeguarding client privacy during the validation process.
  1. Resolution of Validation Integrity Dilemma with LIPC Metric: FLShield introduces LIPC (Loss Impact Per Class), a novel validation metric, to overcome the Validation Integrity Dilemma. This dilemma arises from the possibility of malicious validators reporting incorrect validation results (e.g., mislabeling benign models as malicious or vice versa). LIPC is computed on a class-by-class basis, ensuring sensitivity to targeted attacks, and is designed to remain consistent among benign validators. This consistency allows for effective filtering of spurious validation reports using anomaly detection algorithms.
  1. Superior Performance Against Diverse Poisoning Attacks: Through extensive evaluation, FLShield consistently demonstrates performance comparable to FedOracle, an idealized benchmark where the server has perfect knowledge to filter all malicious updates. Crucially, both the bijective and cluster-based versions of FLShield significantly outperform four state-of-the-art existing defenses across four different types of poisoning attacks (untargeted, targeted label flipping, backdoor, and inner product manipulation) and on various datasets and domains.
  1. Robustness in Non-IID Scenarios and Against Defensive Attacks: FLShield's efficacy extends to challenging non-IID data distributions, where it maintains performance similar to FedOracle. Furthermore, the framework exhibits remarkable resilience against defensive attacks specifically designed to disrupt the validation process by employing malicious validators. FLShield successfully detects these malicious validation reports, ensuring its defensive capabilities remain unimpaired.
  1. Minimal Overhead: The framework achieves its robust security posture while incurring only a small computational and communication overhead required for the validation process, making it practical for real-world Federated Learning deployments.

Technical Deep Dive

▶ Watch: Representative Models for secure client-side validation (4:40)

FLShield's innovative design hinges on addressing two core challenges inherent in client-side validation: the Validation Subject Dilemma and the Validation Integrity Dilemma.

Validation Subject Dilemma and Representative Models

The Validation Subject Dilemma arises because, if validators are chosen from the set of participants, some could be malicious. If these malicious validators receive raw local models for validation, they could potentially launch gradient inversion attacks. These attacks exploit the information contained within model parameters or gradients to reconstruct sensitive training data samples used by the client that submitted the model. This risk makes it unsafe to use local models directly as validation subjects, undermining the privacy guarantees of Federated Learning.

FLShield's solution is the introduction of representative models. Instead of sending raw local models, clients first use their local models to craft these representative models, which are then used for validation. The talk details two approaches for crafting these models:

  1. Bijective Representative Model: In this approach, one local model is selected as a "base" or "mold." Contributions from other local models are then added to this base model, weighted by their similarity to the base. The equation presented in the talk illustrates this: the representative model update is a sum where the second term represents the contribution from the base local model update, and the final term represents contributions from other "sibling" model updates. A parameter tow (τ) controls the amount of sibling contribution. This method essentially creates a composite model that captures aspects of multiple local models without directly exposing any single one in its raw form.
  1. Cluster-Based Representative Model: This method involves an initial clustering step. All local models are first grouped into M clusters, where M is dynamically determined to ensure that model updates within each cluster are parametrically close, while updates from different clusters have noticeably higher distances. After forming these groups, a representative model update is computed for each cluster by aggregating all local model updates belonging to that specific group. This effectively anonymizes individual contributions within a cluster.

Both bijective and cluster-based representative models serve to obfuscate the direct link between a specific model update and its originating client's private data. Experiments confirm that the use of these representative models successfully prevents malicious validators from launching effective gradient inversion attacks, thus safeguarding sensitive training data during the validation process.

Validation Integrity Dilemma and Loss Impact Per Class (LIPC)

The Validation Integrity Dilemma addresses the problem of malicious validators reporting incorrect validation results. Just as malicious clients can send poisoned model updates, malicious validators can report that benign local models are malicious, or conversely, that malicious local models are benign. This poses a significant challenge: how can the server filter these spurious validation reports and trust the validation process itself?

FLShield proposes a new validation metric called LIPC (Loss Impact Per Class) to solve this dilemma. LIPC is designed to identify malicious models by assessing their impact on the global model's performance, computed on a class-by-class basis. This class-specific computation is crucial for detecting targeted attacks, which often aim to degrade performance on specific classes rather than universally.

The loss impact in LIPC refers to the difference in loss computed on two models: the current global model and the representative model being validated. Specifically, for each class, LIPC measures how much the loss changes when the global model is updated with a particular representative model. Benign representative models are expected to reduce loss or have a consistent, small impact on loss across classes, reflecting constructive contributions to the global model. In contrast, malicious representative models, especially those designed for targeted attacks, will exhibit an anomalous loss impact on specific classes.

The key insight is that LIPC scores remain remarkably consistent among benign validators. This consistency allows the server to use an anomaly detection algorithm on the validation reports. If a validator's LIPC scores significantly deviate from the consensus of other benign validators, their report can be flagged as spurious and potentially filtered out. This mechanism effectively ensures the integrity of the validation reports, allowing the server to reliably identify and discard malicious model updates.

FLShield's Design Components (Step-by-Step Process)

The complete FLShield framework integrates these solutions into a robust, multi-stage process:

  1. Local Model Training: Clients train their local models using their private training data.
  2. Representative Model Generation: These local models are then used to generate representative models, employing either the bijective or cluster-based approach, specifically for validation purposes.
  3. Validator Selection: A subset of willing clients is chosen to act as validators for the current aggregation round.
  4. Model Distribution for Validation: The generated representative models, along with the current global model, are sent to the selected validators.
  5. LIPC Computation and Report Generation: Validators compute LIPC scores for each representative model based on their local validation data and the current global model. They then compile these into validation reports.
  6. Validation Report Filtering: The server receives these validation reports. An anomaly detection algorithm is applied to filter out spurious reports submitted by malicious validators, leveraging the consistency of LIPC scores among benign validators.
  7. Representative Model Ranking and Selection: Based on the filtered validation reports, the representative models are ranked according to their assessed quality or benign nature. The top 50% corresponding local model updates (i.e., those from which the highly-ranked representative models were derived) are selected for the next stage. This threshold allows for a significant portion of potentially malicious updates to be excluded.
  8. Model Update Clipping: The remaining selected local model updates (from the top 50%) are then clipped to a predefined norm threshold. This step provides an additional layer of defense against very large or aggressively malicious updates that might have slipped through previous stages.
  9. Global Model Aggregation: Finally, the clipped and validated local model updates are aggregated by the server to produce the updated version of the global model.

This systematic process ensures that only validated and sanitized model contributions are incorporated into the global model, significantly bolstering the security of the Federated Learning system against diverse poisoning attacks.

Demo / Proof of Concept

▶ Watch: Step-by-step design and workflow of FLShield (7:20)

While the talk did not feature a live, interactive demonstration in the traditional sense, the comprehensive evaluation presented serves as a robust proof of concept for FLShield's effectiveness. The researchers rigorously validated FLShield against four distinct types of poisoning attacks across four different datasets and two domains, providing empirical evidence of its capabilities.

The evaluation included comparisons against two crucial baselines:

  • FedAverage: This represents the no-defense aggregation baseline, illustrating the vulnerability of standard FL without protective measures.
  • FedOracle: This serves as an ideal "milestone benchmark." FedOracle assumes the server has an "oracle-like" ability to perfectly identify and filter out all malicious model updates, thus representing the theoretical upper bound of defense performance.

FLShield's performance was also benchmarked against four state-of-the-art existing defenses. The results consistently showed that both versions of FLShield – the one utilizing bijective representative models and the one using cluster-based representative models – achieved performance levels remarkably similar to FedOracle across all tested poisoning attacks. This indicates that FLShield can almost perfectly mitigate the impact of these attacks, nearing the theoretical optimum. Furthermore, FLShield consistently outperformed all four state-of-the-art defenses, demonstrating its superior efficacy.

Specific metrics were used to quantify performance:

  • For untargeted poisoning attacks and inner product manipulation attacks, the primary metric was main accuracy, referring to the final global model accuracy on the test set. FLShield showed high main accuracy, comparable to FedOracle.
  • For targeted label flipping attacks, the crucial metric was the recall of the target class. FLShield effectively maintained high recall for benign classes while preventing the attacker from achieving their targeted misclassification.
  • For backdoor attacks, the attack success rate (ASR) was used, denoting the percentage of backdoor-triggered samples misclassified to the target class. FLShield significantly reduced the ASR, indicating successful neutralization of backdoor triggers.

Beyond these core attack types, FLShield's robustness was further tested in challenging non-IID scenarios, where data distribution varies significantly among clients. Both versions of FLShield achieved results similar to FedOracle even under these conditions, highlighting its practicality for real-world heterogeneous datasets.

A particularly noteworthy aspect of the evaluation involved designing two defensive attacks: "FA Adaptive" and "FA Advanced." These attacks specifically employed malicious validators to disrupt FLShield's defensive process. The results unequivocally demonstrated FLShield's ability to consistently detect malicious validation reports, attributing this success to the consistency of LIPC scores among benign validators. FLShield's performance remained unimpaired even in the presence of these sophisticated defensive attacks, showcasing its resilience. Additionally, specific experiments were conducted to confirm that gradient inversion attacks launched on representative model gradients were unable to uncover any of the training samples, validating the privacy-preserving aspect of representative models.

In essence, the comprehensive experimental validation serves as a powerful proof of concept, demonstrating FLShield's robust, efficient, and privacy-preserving capabilities in securing Federated Learning against a wide array of sophisticated poisoning threats.

Defensive Implications

▶ Watch: Evaluation methodology and comparison baselines (8:00)

FLShield offers crucial insights and practical strategies for defenders seeking to secure Federated Learning systems against poisoning attacks. Its findings suggest several key implications and actionable recommendations:

  1. Prioritize Client-Side Validation: The success of FLShield underscores the critical importance of incorporating client-side validation into FL defense strategies. Relying solely on server-side mechanisms or unrealistic assumptions about server-controlled clean datasets is insufficient. Defenders should explore architectures that empower clients, as the true data owners, to participate in the validation process.
  1. Implement Privacy-Preserving Validation Mechanisms: The Validation Subject Dilemma highlights a significant privacy risk when local models are directly exposed for validation. Defenders must adopt mechanisms like FLShield's representative models (bijective or cluster-based) to ensure that the validation process itself does not become an avenue for gradient inversion attacks or other privacy breaches. This ensures that privacy, a core tenet of FL, is maintained throughout the defense lifecycle.
  1. Utilize Robust and Class-Aware Validation Metrics: The Validation Integrity Dilemma demonstrates that malicious validators can subvert the defense. Implementing a robust validation metric like LIPC (Loss Impact Per Class) is essential. LIPC's class-by-class computation makes it particularly effective against targeted attacks, and its consistency among benign participants enables reliable anomaly detection to filter out spurious reports. Defenders should look for metrics that are sensitive to attack specifics and exhibit predictable behavior among legitimate participants.
  1. Anticipate and Defend Against Adaptive Adversaries: The evaluation against "defensive attacks" where malicious validators actively tried to disrupt FLShield's defense is a critical lesson. Defenders must assume that adversaries will adapt and attempt to subvert the defense mechanism itself. Building resilience against such adaptive attacks is paramount, and FLShield's demonstrated robustness, owing to the design of LIPC, provides a valuable blueprint.
  1. Consider Hybrid Defense Strategies: While FLShield is a comprehensive framework, its principles can be integrated with or complement other existing defenses. For instance, combining FLShield's validation with other techniques like secure aggregation protocols or differential privacy mechanisms could create an even more formidable defense posture.
  1. Benchmark Against Realistic Baselines: The use of FedOracle as an idealized benchmark provides a clear performance target. Defenders should evaluate their FL security solutions not just against no-defense baselines but also against theoretical optimums and, crucially, against a diverse set of state-of-the-art attacks and existing defenses to thoroughly assess their real-world efficacy.
  1. Address Non-IID Data Challenges: FLShield's effectiveness in non-IID scenarios is a significant advantage. Defenders operating in domains with naturally heterogeneous client data distributions should prioritize defense mechanisms that are proven to perform well under such challenging conditions.

In summary, FLShield provides a powerful, validated framework for building more secure Federated Learning systems. Its emphasis on client-side, privacy-preserving validation, coupled with robust integrity checks, offers a practical and effective blueprint for mitigating diverse poisoning attacks and ensuring the trustworthiness of FL models in safety-critical applications.

Key Takeaways

  • FLShield is a validation-based framework that effectively defends Federated Learning against a wide range of poisoning attacks (untargeted, targeted label flipping, backdoor, model poisoning) while preserving data privacy and efficiency.
  • It introduces "Representative Models" (bijective or cluster-based) to solve the "Validation Subject Dilemma," preventing malicious validators from launching gradient inversion attacks and uncovering sensitive training data.
  • The framework proposes "LIPC (Loss Impact Per Class)" as a novel validation metric to address the "Validation Integrity Dilemma," ensuring the reliability of validation reports by detecting and filtering spurious results from malicious validators.
  • FLShield consistently achieves performance comparable to the idealized FedOracle benchmark and significantly outperforms four state-of-the-art existing defenses, demonstrating its superior efficacy across various attack types, datasets, and domains.
  • It proves robust in challenging non-IID data distributions and resilient against sophisticated "defensive attacks" launched by malicious validators, maintaining unimpaired performance.
  • FLShield operates with minimal computational and communication overhead, making it a practical and deployable solution for securing real-world Federated Learning applications, especially in privacy-sensitive and safety-critical domains.

About the Speaker(s)

The research presented on FLShield was led by Ehsanul Kabir, who delivered the talk, and was a joint effort with Zeyu Song, Md Rafi Ur Rashid, and Shagufta Mehnaz. All four researchers are affiliated with Pennsylvania State University. Their collective work focuses on advancing the security and robustness of Federated Learning systems, particularly in addressing critical vulnerabilities like poisoning attacks that threaten the integrity and trustworthiness of collaborative AI models. Their contributions aim to make Federated Learning a more secure and reliable technology for widespread adoption in various sensitive domains.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

FLShield tackles critical poisoning attacks in Federated Learning by introducing a novel client-side validation framework. It cleverly resolves the privacy and integrity dilemmas inherent in decentralized validation using representative models and a new metric, LIPC, offering a robust and practical defense that significantly outperforms current state-of-the-art.

Heather Calloway (CISO) — STRONG ACCEPT

This framework directly addresses critical integrity and privacy risks in Federated Learning, a technology increasingly vital in sensitive domains like healthcare and finance. It provides a robust, validated approach for CISO organizations to ensure model trustworthiness and accountability, offering concrete mechanisms for safeguarding against sophisticated poisoning attacks.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024