SHERPA: Explainable Robust Algorithms for Privacy-preserved Federated Learning in Future Networks to Defend against Data Poisoning Attacks
Chamara Sandeepa, Bartlomiej Siniarski, Shen Wang, Madhusanka Liyanage
IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 6
Overview
Federated Learning (FL) has emerged as a powerful paradigm for collaborative machine learning, enabling multiple clients to jointly train a global model without sharing their raw data. This distributed approach preserves data privacy by keeping local data on client devices, making it particularly attractive for sensitive applications in healthcare, finance, and future networks. However, the decentralized nature of FL also introduces significant security vulnerabilities, primarily through data poisoning attacks. Malicious clients can inject carefully crafted, poisoned data into their local training sets, leading to corrupted local models that, when aggregated, degrade the performance of the global model or implant insidious backdoors.

Key moments
- 0:00 Introduction to Federated Learning and data poisoning problem
- 1:50 Limitations of current solutions and SHERPA's contributions
- 3:00 SHAP-based feature attribution for poisoner detection
- 4:15 SHERPA algorithm overview: aggregator's role
- 5:20 Detailed poison detection using HDBSCAN clustering
- 6:40 Visualizing poison client deviation with t-SNE plots
- 8:00 Evaluation of SHERPA's effectiveness and accuracy
- 10:00 Comparison of SHERPA with existing benchmark techniques
SHERPA: Explainable Robust Algorithms for Privacy-preserved Federated Learning in Future Networks to Defend against Data Poisoning Attacks
Speakers: Chamara Sandeepa, Bartlomiej Siniarski, Shen Wang, Madhusanka Liyanage
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=LLgb7Iw7Wg
Overview
Federated Learning (FL) has emerged as a powerful paradigm for collaborative machine learning, enabling multiple clients to jointly train a global model without sharing their raw data. This distributed approach preserves data privacy by keeping local data on client devices, making it particularly attractive for sensitive applications in healthcare, finance, and future networks. However, the decentralized nature of FL also introduces significant security vulnerabilities, primarily through data poisoning attacks. Malicious clients can inject carefully crafted, poisoned data into their local training sets, leading to corrupted local models that, when aggregated, degrade the performance of the global model or implant insidious backdoors.
Current defenses against data poisoning in FL often rely on heuristic methods such as model similarity metrics (e.g., cosine similarity) or arbitrary threshold cut-offs to filter out suspicious model updates. While these techniques offer some level of protection, they frequently suffer from critical limitations. They often lack the granularity to distinguish subtle poisoning behaviors from legitimate data variations, potentially removing benign updates and hindering model convergence. Crucially, they fail to provide explainability, leaving practitioners without clear insights into why a particular client's model update was deemed malicious or how poisoners specifically deviate from benign participants.
This article delves into SHERPA, a novel framework presented at the 45th IEEE Symposium on Security and Privacy 2024. SHERPA addresses these limitations by introducing an explainable robust algorithm for detecting and eliminating data poisoners in FL. By leveraging SHAP-based feature attributions, SHERPA provides a transparent mechanism to identify significant differences in feature importance exhibited by poisoned clients compared to benign ones. The framework not only enhances the robustness of FL against various poisoning attacks but also tackles the often-overlooked impact of poisoning on client privacy, demonstrating mitigation strategies against related threats like property inference attacks.
Background
▶ Watch: Introduction to Federated Learning and data poisoning problem (0:00)
Federated Learning operates on a client-server architecture where a central aggregator coordinates the training process. Clients download the current global model, train it locally on their private datasets, and then upload their model updates (e.g., gradients or model weights) to the aggregator. The aggregator then combines these updates, typically through averaging, to produce a new global model, which is subsequently distributed back to the clients for the next round of training. This iterative process allows for continuous model improvement while maintaining data locality.
The inherent trust assumption in FL—that all participating clients are benign and contribute constructively—is a critical vulnerability. Data poisoning attacks exploit this assumption by having malicious clients inject corrupted or mislabeled data into their local training sets. The objective of such attacks can vary:
- Degrading overall model performance: Malicious updates can shift the model's decision boundary, significantly reducing its accuracy on legitimate tasks.
- Introducing backdoors: Attackers can embed specific triggers into the model such that, when presented with an input containing the trigger, the model misclassifies it in a predefined way, while performing normally on other inputs.
- Facilitating privacy attacks: Poisoning can be a precursor or an amplifier for other privacy threats, such as membership inference attacks (determining if a specific data point was part of the training set) or property inference attacks (inferring sensitive attributes about the training data, like the presence of specific features). The talk specifically highlighted how poisoning can enhance property inference attacks.
Existing defense mechanisms often fall short in adequately addressing these sophisticated threats. Many rely on statistical properties of model updates, such as their L2 norm, cosine similarity, or Euclidean distance, to identify outliers. For instance, techniques like Krum or F-Gold aggregate only a subset of client updates, discarding those deemed too far from the median. While effective against simple attacks, these methods struggle with:
- Lack of explainability: They provide a binary "malicious/benign" decision without explaining why, making it difficult to understand the attack vector or refine defenses.
- Arbitrary thresholds: Defining effective thresholds for similarity or distance metrics is challenging and often requires extensive tuning. Suboptimal thresholds can either be too permissive (allowing poisoners) or too restrictive (discarding benign updates), impacting model convergence and overall performance.
- Vulnerability to sophisticated attacks: Adaptive attackers can craft subtle poisoning strategies that mimic benign updates, bypassing simple statistical filters.
The motivation behind SHERPA stems from these limitations, aiming to provide a robust and transparent defense that not only detects poisoners effectively but also offers actionable insights into their behavior.
Key Findings
▶ Watch: SHAP-based feature attribution for poisoner detection (3:00)
SHERPA introduces several significant advancements and demonstrates compelling key findings in defending Federated Learning against data poisoning attacks:
- Explainable Poisoner Detection via SHAP: The core finding is that SHAP (SHapley Additive exPlanations)-based feature attributions can effectively distinguish between benign and poisoned clients. SHAP values quantitatively explain the contribution of each feature to a model's prediction. The research revealed a clear visual and quantitative difference: benign clients typically exhibit high feature importance across relevant regions (often visualized as "red" regions), whereas poisoned clients show significantly lower feature importance or different patterns in critical areas (visualized as "blue" regions). This explainability provides concrete evidence for identifying malicious updates, moving beyond opaque statistical thresholds.
- High Detection Rate Across Poisoning Levels: SHERPA demonstrated a consistently high detection rate, even under extreme poisoning scenarios. In experiments involving 20 clients, SHERPA successfully detected poisoners when they constituted 10%, 30%, 50%, and even 80% of the participating clients. This robustness is a critical improvement over existing methods that often degrade under high poisoning rates.
- Superior Performance Against Benchmarks: When compared against established robust aggregation techniques like F-Gold and M-CRUM, SHERPA exhibited a higher detection rate and maintained superior model accuracy. This indicates SHERPA's practical efficacy in real-world FL deployments. The defense mechanism ensures that the main task accuracy of the FL algorithm remains largely unaffected, even when operating in a poisoned environment.
- Mitigation of Privacy Threats: A crucial finding was SHERPA's ability to defend against privacy attacks amplified by poisoning. Specifically, it was shown to mitigate targeted feature poisoning designed to enhance property inference attacks. When an attacker used poisoning to infer a property (e.g., "hair color" on a CIFAR-A dataset), the FL model's accuracy sharply declined (from 85% to 31.6%) at the iteration where the property was targeted. SHERPA successfully prevented this decline, ensuring that the inference attack failed due to the absence of a noticeable accuracy drop, thereby preserving client privacy.
- Robustness Across Poisoning Types: SHERPA proved effective against both continuous poisoning (malicious updates in every round) and periodic poisoning (malicious updates in specific rounds). While continuous poisoning generally had a higher impact on accuracy without defense, SHERPA maintained similar high accuracy levels for both scenarios, showcasing its versatility.
- Insights into Explainer Types and Computational Trade-offs: The research investigated the impact of different SHAP explainer types (DeepExplainer and GradientExplainer) on detection accuracy, finding that both yielded similar overall results. This suggests flexibility in choosing explainer implementations based on specific model architectures or computational preferences. Furthermore, the study highlighted a clear trade-off between the number of representative data points used to generate SHAP values and the computation time. While more representative data increased accuracy, it linearly increased computation time (e.g., from 0.525 seconds for DS=1 to 5.130 seconds for DS=10), necessitating a balance in practical deployments.
Technical Deep Dive
▶ Watch: Detailed poison detection using HDBSCAN clustering (5:20)
The technical core of SHERPA lies in its innovative use of SHAP (SHapley Additive exPlanations) values to identify and characterize malicious client behavior in Federated Learning. SHAP is a game-theoretic approach that assigns an importance value to each feature for a particular prediction, providing a measure of how much that feature contributes to the prediction compared to the average prediction.
In the context of SHERPA, the aggregator, which typically performs model aggregation, takes on an additional role. It maintains a small, representative dataset (distinct from client training data) that it uses to generate SHAP values for each incoming local model update from clients. Instead of directly analyzing model weights or gradients, SHERPA analyzes the impact of each client's model on the SHAP values computed on this representative dataset.
The process unfolds in several key steps:
- SHAP Value Derivation:
- When a client uploads its local model, the aggregator uses its small representative dataset to calculate SHAP-based feature attributions for that model.
- For each class and each input feature, SHAP quantifies how much that feature contributes to the model's output for that class.
- A key observation is that poisoned clients tend to exhibit significantly different feature importance patterns compared to benign clients. For instance, a benign client's model might show strong positive attributions (red regions) for features relevant to a correct classification, while a poisoned client's model might show negative attributions (blue regions) or shifted importance for those same features, indicating a deviation from expected behavior.
- Feature Attribution Vectorization:
- The raw SHAP attributions, which are often multi-dimensional (e.g., an attribution map for an image), are flattened into vectors. This transformation converts the complex SHAP output for each client into a comparable numerical representation suitable for clustering.
- Clustering with HDBScan:
- To detect anomalous patterns, SHERPA employs the HDBScan (Hierarchical Density-Based Spatial Clustering of Applications with Noise) algorithm. HDBScan is particularly well-suited for this task because it can discover clusters of varying densities and effectively identify noise points, which in this context correspond to poisoners.
- HDBScan operates by constructing a hierarchy of clusters from the flattened SHAP attribution vectors. It then uses a concept called mutual reachability distance to define the density and connectivity between data points. Points that are mutually reachable within a certain density form a cluster.
- The mutual reachability distance helps in identifying how "close" two data points are in a density-based sense, allowing HDBScan to form clusters where points are densely packed and separate them from sparse regions. This is crucial for isolating subtle deviations caused by poisoning.
- Poisoner Detection and Elimination:
- After clustering, each client's SHAP attribution vector is assigned to a cluster (e.g., Cluster A, Cluster B).
- A client is considered to have a suspicious attribution if its SHAP values deviate significantly from the typical patterns observed within its assigned cluster. The exact deviation threshold can be tuned, often based on statistical measures within the clusters.
- If a client's overall "suspiciousness score" (e.g., based on the number or magnitude of suspicious attributions) exceeds an average suspiciousness threshold across all clients, that client is flagged as a poisoner.
- Once identified, the model updates from these detected poisoned clients are eliminated from the aggregation process, preventing their malicious contributions from corrupting the global model.
The use of SHAP provides inherent explainability: when a client is flagged, the aggregator can inspect the specific feature attributions that led to its classification as a poisoner. This transparency is a significant advantage over black-box detection methods. Furthermore, the selection of HDBScan allows for robust clustering without requiring prior knowledge of the number of clusters, adapting to the dynamic nature of poisoning attacks and benign client behaviors. The talk also explored the impact of hyperparameter tuning on improving the overall accuracy of the clustering process, emphasizing the need to optimize these parameters for specific datasets and attack scenarios.
Demo / Proof of Concept
▶ Watch: Visualizing poison client deviation with t-SNE plots (6:40)
The presentation provided a compelling demonstration of SHERPA's capabilities through various experimental setups and visualizations, effectively serving as a proof of concept.
A primary visualization tool employed was t-SNE (t-Distributed Stochastic Neighbor Embedding) plots. These plots are used to reduce high-dimensional data (in this case, the flattened SHAP attribution vectors) into a 2D or 3D representation, making it possible to visually inspect the clustering of clients. The t-SNE plots clearly illustrated that poisoned clients, while initially appearing similar to benign clients in raw model updates, formed distinct, anomalous clusters when their SHAP-based feature attributions were analyzed. These anomalous clusters, often appearing as isolated groups or outliers, were highlighted (e.g., within a "red circle" in the figure), providing a clear visual cue for the deviation of poisoned models. This direct visualization reinforces the explainability aspect of SHERPA.
Beyond abstract plots, the talk visually contrasted the SHAP feature attributions themselves for benign versus poisoned clients. A benign client's model for a correct classification (e.g., predicting "3" accurately) showed a high accuracy and corresponding high feature importance in the relevant regions of the input. In stark contrast, a randomly poisoned client exhibited a lack of clear patterns in both accuracy and feature attributions. For targeted poisoning, while the targeted class might show good accuracy, other classes would lack the expected feature importance trends, indicating manipulated behavior.
The experimental setup for evaluating SHERPA involved:
- Client Count: 20 clients.
- Global Iterations: 10 rounds of Federated Learning.
- Poisoning Levels: Varying percentages of malicious clients were introduced: 10% (2 poisoners), 30% (6 poisoners), 50% (10 poisoners), and 80% (16 poisoners).
- Poisoning Types: Both continuous poisoning (malicious updates in every round) and periodic poisoning (malicious updates in specific, intermittent rounds) were evaluated.
The results demonstrated SHERPA's robustness:
- Accuracy Preservation: Without any defense, continuous poisoning had a higher impact on the global model's accuracy compared to periodic poisoning. However, with SHERPA implemented, the accuracy remained remarkably similar for both cases, closely matching the accuracy of a system running without any poisoning. This was shown in a table where "Federated learning accuracy" was significantly degraded without defense but "Federated learning defense accuracy" remained high and stable.
- High Detection Accuracy: Another table presented the defense accuracy across different data sets, showing high true positive, true negative, false positive, and false negative rates, indicating SHERPA's precision in identifying and isolating poisoners.
- Benchmark Comparison: SHERPA was directly compared against existing benchmark techniques like F-Gold and M-CRUM. The results consistently showed that SHERPA achieved a higher detection rate, even when faced with an extremely high proportion of poisoners (up to 80%), outperforming these established methods.
A particularly impactful demonstration involved the mitigation of privacy attacks. The speakers showcased a scenario where targeted feature poisoning was used to enhance a property inference attack on the CIFAR-A dataset, specifically targeting the "hair color" property.
- Without defense, the poisoning caused a sharp decline in the Federated Learning accuracy at the specific iteration where the property inference attack was launched. This distinct drop (from 85% to 31.6% in property inference accuracy) could be exploited by an attacker to infer the targeted property.
- When SHERPA was applied, this noticeable decline in accuracy was completely absent. The defense successfully prevented the malicious updates from influencing the global model in a way that would reveal sensitive properties, thereby thwarting the property inference attack. This illustrates SHERPA's dual benefit of improving robustness and preserving privacy.
Finally, the talk explored factors influencing the clustering process. It showed that different SHAP explainer types (DeepExplainer and GradientExplainer) yielded similar poison detection accuracies, offering flexibility in implementation. It also highlighted the trade-off between the number of representative data points used for SHAP value generation and the computational time, providing practical considerations for deployment. For instance, increasing representative data from 1 to 10 linearly increased computation time from 0.525 seconds to 5.130 seconds.
Defensive Implications
▶ Watch: Comparison of SHERPA with existing benchmark techniques (10:00)
The findings and capabilities of SHERPA have several profound implications for defenders operating and securing Federated Learning systems:
- Shift Towards Explainable AI for Security: The most significant implication is the necessity for FL defenders to integrate explainable AI (XAI) techniques, such as SHAP, into their security architectures. Relying solely on opaque statistical heuristics is no longer sufficient against sophisticated, adaptive poisoning attacks. SHERPA demonstrates that understanding why a model behaves suspiciously, through feature attributions, provides a more robust and transparent defense mechanism.
- Enhanced Anomaly Detection: Defenders should consider moving beyond simple outlier detection based on model update similarity. By analyzing SHAP-based feature attribution patterns, FL platforms can identify subtle, targeted poisoning behaviors that might otherwise blend in with benign noise. The ability to detect deviations in feature importance offers a more granular and powerful anomaly detection capability.
- Proactive Privacy Preservation: SHERPA highlights that data poisoning isn't just about degrading model accuracy; it can also be a vector for privacy breaches. Defenders must recognize that robust aggregation methods, like SHERPA, can serve as a proactive defense layer against property and membership inference attacks that are amplified by poisoning. Implementing such defenses can prevent attackers from exploiting model accuracy drops as indicators for sensitive data properties.
- Computational Overhead Considerations: While highly effective, the computation of SHAP values and subsequent clustering (especially with HDBScan) introduces additional computational overhead for the aggregator. FL system designers must account for this, particularly when dealing with a large number of clients or high-dimensional models. The identified trade-off between the number of representative data points and computation time suggests that careful optimization and resource allocation are necessary for practical deployment.
- Representative Data Management: The aggregator's ability to generate accurate SHAP values hinges on maintaining a small, representative dataset. Defenders need to establish secure and efficient mechanisms for curating and managing this dataset. Its quality directly impacts the efficacy of poisoner detection.
- Granular Client Monitoring: SHERPA enables more granular monitoring of client contributions. Instead of just flagging a client as "malicious," the defense can provide insights into which features or which classes are being targeted by poisoning. This information can be invaluable for forensic analysis, understanding attack vectors, and potentially identifying specific malicious actors or attack campaigns.
- Dynamic Thresholding and Adaptability: Unlike fixed thresholds, density-based clustering like HDBScan can adapt to varying distributions of benign and malicious clients. This means defenders can deploy more resilient systems that are less susceptible to being bypassed by adaptive attackers who try to "mimic" benign updates by staying within predefined statistical boundaries.
- Future-Proofing FL Defenses: As FL continues to evolve and face increasingly sophisticated threats, incorporating explainability and robust anomaly detection at the feature attribution level, as demonstrated by SHERPA, is crucial for building future-proof defenses. This approach moves FL security beyond reactive measures to a more insightful and proactive stance.
Key Takeaways
- SHERPA is an explainable and robust algorithm designed to detect and eliminate data poisoners in Federated Learning by leveraging SHAP-based feature attributions.
- SHAP values provide crucial explainability, revealing distinct patterns in feature importance between benign and poisoned client models, which is often visualized as different "red" (high importance) and "blue" (low importance) regions.
- The system uses HDBScan clustering on flattened SHAP attribution vectors to effectively identify and isolate malicious clients that form anomalous clusters, even in high-poisoning scenarios (up to 80% malicious clients).
- SHERPA significantly improves the overall accuracy and robustness of Federated Learning by effectively filtering out poisoned updates, outperforming existing benchmark robust aggregation techniques like F-Gold and M-CRUM.
- Beyond robustness, SHERPA offers a strong defense against privacy threats, successfully mitigating property inference attacks that are amplified by targeted feature poisoning, preventing attackers from inferring sensitive data attributes.
- There is a practical trade-off between the amount of representative data used by the aggregator to generate SHAP values (which improves accuracy) and the resulting computation time, requiring careful consideration for deployment.
About the Speaker(s)
The talk "SHERPA: Explainable Robust Algorithms for Privacy-preserved Federated Learning in Future Networks to Defend against Data Poisoning Attacks" was presented by Chamara Sandeepa, alongside co-authors Bartlomiej Siniarski, Shen Wang, and Madhusanka Liyanage. All speakers are affiliated with University College Dublin, indicating a strong research background in the fields of security, privacy, and distributed machine learning, particularly within the context of Federated Learning and future network architectures. Their work focuses on addressing critical vulnerabilities like data poisoning and enhancing the explainability and robustness of AI systems.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
SHERPA introduces a novel, explainable framework for detecting data poisoning in Federated Learning, utilizing SHAP values and HDBScan clustering. This approach offers transparent identification of malicious client behavior, moving beyond opaque heuristics, and effectively mitigates various poisoning attacks, including those amplifying privacy threats.
Heather Calloway (CISO) — STRONG ACCEPT
SHERPA presents a critical advancement in securing Federated Learning, offering an explainable and robust defense against data poisoning. For organizations relying on FL in sensitive sectors, this work provides a tangible mechanism to ensure model integrity and mitigate significant business and privacy risks.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024