DShield: Defending against Backdoor Attacks on Graph Neural Networks via Discrepancy Learning

Hao Yu

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · ML Backdoors

Overview

Graph Neural Networks (GNNs) have emerged as powerful tools for analyzing complex relational data, finding widespread application in areas like social networks, bioinformatics, and recommender systems. Their ability to model intricate relationships between nodes and edges makes them highly effective for tasks such as node classification, graph classification, and link prediction. However, this growing reliance on GNNs has also attracted the attention of adversaries, leading to the rise of backdoor attacks as a significant threat to GNN-based applications. These attacks manipulate a GNN model's behavior by injecting hidden triggers into the training graph, forcing the model to make incorrect predictions when presented with specific trigger patterns.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction: GNNs and Backdoor Attack Threat
  2. 1:00 Understanding Dirty Label vs. Clean Label Attacks
  3. 2:00 Key Findings: Semantic Drift and Attribute Overemphasis
  4. 3:30 DShield's Architecture: Three Key Modules
  5. 4:00 Detailed: Self-Supervised Model Training for Robustness
  6. 6:00 Detailed: Discrepancy Matrix for Poison Node Identification
  7. 8:00 Experimental Evaluation and Robustness Analysis
  8. 9:50 DShield's Applicability to Graph Classification Tasks

DShield: Defending against Backdoor Attacks on Graph Neural Networks via Discrepancy Learning

Speakers: Hao Yu

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=Wp1QiNBdM5U

Overview

Graph Neural Networks (GNNs) have emerged as powerful tools for analyzing complex relational data, finding widespread application in areas like social networks, bioinformatics, and recommender systems. Their ability to model intricate relationships between nodes and edges makes them highly effective for tasks such as node classification, graph classification, and link prediction. However, this growing reliance on GNNs has also attracted the attention of adversaries, leading to the rise of backdoor attacks as a significant threat to GNN-based applications. These attacks manipulate a GNN model's behavior by injecting hidden triggers into the training graph, forcing the model to make incorrect predictions when presented with specific trigger patterns.

This talk by Hao Yu introduces DShield, a novel defense mechanism designed to protect GNNs against both dirty-label and clean-label backdoor attacks. DShield leverages a sophisticated approach that combines self-supervised learning with discrepancy analysis to identify poisoned nodes within the training graph. By pinpointing these malicious nodes, DShield enables the training of robust, backdoor-free GNN models, thereby safeguarding the integrity and reliability of GNN deployments in critical real-world scenarios. The research addresses a crucial gap in GNN security, offering a practical and effective solution to a pervasive and evolving threat.

The significance of DShield lies in its ability to counteract sophisticated, stealthy attacks that can bypass conventional defenses. By identifying fundamental behavioral shifts induced by backdoor triggers—namely "semantic drift" and "attribute overemphasize"—DShield provides a data-driven approach to detect and mitigate these vulnerabilities. This work is particularly relevant given the increasing deployment of GNNs in sensitive domains where model integrity and trustworthy predictions are paramount, making contributions like DShield essential for fostering secure and resilient AI systems.

Background

▶ Watch: Introduction: GNNs and Backdoor Attack Threat (0:00)

Graph Neural Networks (GNNs) are a class of deep learning models designed to operate on graph-structured data. Unlike traditional neural networks that process independent data points, GNNs leverage the relational information embedded in graphs, propagating information across nodes and edges to learn powerful representations. This capability has led to their deployment in diverse applications, including identifying fraudulent activities in financial networks, predicting protein interactions, and recommending products to users. The interconnected nature of graph data, while enabling powerful insights, also introduces unique vulnerabilities, particularly during the model training phase.

Backdoor attacks represent a critical security challenge for GNNs. In these attacks, an adversary subtly modifies the training data (the graph) by injecting specific "triggers" that, when present in an input, cause the trained GNN model to output a predetermined incorrect prediction. The goal is to create a model that behaves normally on clean inputs but misbehaves predictably on trigger-laden inputs, all while maintaining high performance on legitimate data. This stealthy manipulation makes backdoor attacks particularly dangerous, as the compromised model can be deployed without immediate suspicion, only to be exploited later.

Backdoor attacks on GNNs can be broadly categorized into two types:

  1. Dirty-label backdoor attacks: In this scenario, the attacker modifies both the graph structure/attributes and the label information of certain poisoned nodes. By associating the injected triggers with a specific target label during training, the attacker forces the model to learn an incorrect association. The manipulated labels directly contribute to the model associating the trigger with the target.
  2. Clean-label backdoor attacks: These attacks are more stealthy. The attacker only modifies the attributes or structure of normal nodes, without altering their original labels. The poisoned nodes retain their correct labels, making them blend more naturally into the training data. The model is then subtly coerced to overemphasize the injected altered attributes when classifying nodes containing the trigger, leading to misclassification despite correct labeling during training. Clean-label attacks are often harder to detect because they do not involve overt label manipulation.

The problem exists because GNNs, like other deep learning models, are susceptible to data poisoning. During training, GNNs learn patterns and relationships from the input graph. If this graph is maliciously crafted to include hidden triggers, the GNN will inadvertently learn these triggers as legitimate features associated with specific outcomes. The inherent complexity of graph structures and the non-linear transformations within GNNs make it challenging to discern whether a learned pattern is truly representative of the data or an artifact of an adversarial injection. This necessitates robust defense mechanisms that can identify and neutralize these hidden threats before they compromise model integrity.

Key Findings

▶ Watch: Key Findings: Semantic Drift and Attribute Overemphasis (2:00)

The development of DShield is predicated on two fundamental observations regarding how backdoor attacks manifest within the latent representations and attribute importance of GNNs. These "key phenomena" provide the theoretical underpinning for DShield's detection strategy, allowing it to differentiate between normal and maliciously poisoned nodes.

The first key finding, primarily observed in dirty-label backdoor attacks, is semantic drift. The researchers used t-SNE (t-distributed Stochastic Neighbor Embedding) visualization to analyze the latent representations of nodes. When a GNN model is trained using supervised learning on a backdoored graph (where dirty-label attacks are present), the poisoned nodes, along with nodes legitimately belonging to the target class, tend to cluster together near the target label in the latent space. This indicates that the attack successfully misleads the model. However, a crucial insight emerged when the same latent representations were visualized from a model trained without label information, using self-supervised learning. In this scenario, the poisoned nodes were no longer clustered near the target label. This divergence reveals that dirty-label attacks cause a "semantic drift": the injected triggers alter the intrinsic semantic meaning of the poisoned nodes, pushing their latent representations away from their true, natural cluster when no label guidance is provided. The self-supervised model, by learning the natural structure of the graph, exposes this distortion.

The second key finding, predominantly associated with clean-label backdoor attacks, is attribute overemphasize. To investigate this, the researchers employed gradient analysis to measure the influence of individual attributes on the model's prediction for each node. In the presence of clean-label attacks, the poisoned nodes exhibited a remarkably highly similar attribute importance profile. This means that the GNN model, when classifying poisoned nodes, excessively focused on the specific altered attributes injected by the attacker. Even though the labels were clean, the model had learned to disproportionately weigh these manipulated attributes, effectively creating a shortcut for misclassification when the trigger was present. This overemphasis on specific, altered attributes provides a distinct fingerprint for clean-label attacks.

These two phenomena—semantic drift and attribute overemphasize—are critical because they reveal distinct behavioral patterns that differentiate poisoned nodes from benign ones, even when the attacks are designed to be stealthy. DShield is meticulously designed to exploit these discrepancies, using them as indicators to identify and isolate malicious nodes, thereby enabling the training of robust, backdoor-free GNN models.

Technical Deep Dive

▶ Watch: Detailed: Self-Supervised Model Training for Robustness (4:00)

DShield's architecture is meticulously designed around three core modules: backdoor and self-supervised model training, discrepancy matrix construction, and backdoor-free model training. Each module plays a distinct role in identifying and mitigating backdoor attacks.

Backdoor and Self-Supervised Model Training

The first module involves training two distinct GNN models:

  1. Backdoor Model: This model is trained using the potentially manipulated label information from the input graph. Its purpose is to be vulnerable to backdoor attacks, reflecting how a standard GNN would behave if directly trained on compromised data.
  2. Self-Supervised Model: This is the cornerstone of DShield's detection capability. It is trained using a contrastive learning framework without relying on any label information, enabling it to learn the natural, robust structure of the graph. The training process for the self-supervised model involves three key steps:
  • Augmentation: The input graph is augmented using two statistical transformation functions. This augmentation occurs at two levels: structure level and attribute level. The process leverages the similarity of attributes within connected edges and the importance of attributes in prediction. Based on this information, edges and attributes are selectively masked. This creates multiple "views" of the original graph, which are crucial for contrastive learning.
  • View Encoding: For each augmented graph view, a GNN encoder model generates node embeddings, which are latent representations of the nodes. These embeddings capture the structural and feature information of each node within its respective augmented view.
  • Contrast and Reconstruction: This step involves distinguishing similar and dissimilar nodes. Beyond conventional contrastive learning loss functions, DShield introduces representation augmentation and reconstruction loss. This additional loss function helps to reduce false negatives and strengthens the connection between the original attributes and their learned representations, ensuring that the model learns a highly robust and semantically meaningful graph representation. The overall self-supervised learning process forces the model to learn a representation that is resilient to adversarial manipulations, as it focuses on inherent graph properties rather than potentially poisoned label associations.

Discrepancy Matrix Construction

Once both the backdoor and self-supervised models are trained, DShield proceeds to construct two types of discrepancy matrices to identify poisoned nodes:

  1. Semantic Discrepancy Matrix:
  • First, latent representations (node embeddings) are obtained for all nodes from both the backdoor model and the self-supervised model.
  • Next, semantic matrices are built for these representations, essentially measuring the distance or dissimilarity between node embeddings.
  • The focus is then placed on nodes that exhibit a significantly higher distance in the self-supervised model's latent space compared to their distance in the backdoor model's latent space. This difference highlights nodes whose semantic meaning has been distorted by dirty-label attacks, as the self-supervised model (uninfluenced by labels) reveals their true, unpoisoned relationships.
  1. Attribute Importance Discrepancy Matrix:
  • This matrix analyzes the importance of attribute dimensions using gradient-based methods. These methods quantify how much each attribute contributes to a node's prediction.
  • Since attributes can have many dimensions, UMAP (Uniform Manifold Approximation and Projection) is employed to reduce the dimensionality of the attributes. This avoids the "curse of dimensionality" and allows for more effective analysis of attribute importance patterns.
  • The matrix identifies nodes where specific attributes are disproportionately emphasized, which is a hallmark of clean-label backdoor attacks.

Finally, these two discrepancy matrices are fused to create a comprehensive detection signal. A clustering method, HDBSCAN (Hierarchical Density-Based Spatial Clustering of Applications with Noise), is then applied to the fused matrix to identify distinct clusters of nodes, specifically pinpointing those identified as poisoned. HDBSCAN is robust to noise and can discover clusters of varying shapes and densities, making it suitable for identifying anomalous groups of poisoned nodes.

Backdoor-Free Model Training

With the poisoned nodes identified, DShield moves to train a model that is free from backdoor vulnerabilities:

  1. The original graph is divided into two subgraphs: one containing all the identified poisoned nodes and the other containing the remaining (clean) nodes.
  2. A new GNN model is then trained with a specific objective: maximizing the classification loss on the poisoned subgraph while minimizing the classification loss on the clean subgraph. This counter-intuitive approach forces the model to "unlearn" the adversarial patterns associated with the poisoned nodes, effectively neutralizing the backdoor triggers, while still performing well on legitimate data.
  3. After this specialized training, the mean distance between the latent representations of nodes in both subgraphs is computed.
  4. Finally, Median Absolute Deviation (MAD) is used to identify poisoned labels. These identified poisoned labels are then used to further refine the separation of poisoned nodes, ensuring a robust and accurate defense. By isolating and actively mitigating the influence of poisoned nodes and labels, DShield ensures that the final deployed GNN model is resilient to the previously injected backdoor attacks.

Demo / Proof of Concept

▶ Watch: Detailed: Discrepancy Matrix for Poison Node Identification (6:00)

While the talk did not feature a live, interactive demonstration, the speaker presented compelling visual and quantitative evidence of DShield's efficacy. The core of the proof of concept resided in the visualization of the discrepancy matrices and comprehensive evaluation results against various backdoor attacks and datasets.

The speaker presented figures (A to D for semantic discrepancy matrices, and "the other two phases" for attribute importance discrepancy matrices) which visually illustrated the clear separation between normal and malicious nodes. These visualizations demonstrated a distinct "gap" in the discrepancy scores, confirming that DShield's underlying algorithms could effectively distinguish between benign and poisoned nodes. This visual proof directly supported the theoretical findings of semantic drift and attribute overemphasize, showing how these phenomena translate into quantifiable differences that DShield exploits.

For empirical evaluation, DShield was rigorously tested against a suite of backdoor attacks:

  • Seven dirty-label attacks were evaluated across four distinct datasets.
  • Similarly, two clean-label attacks were tested on the same four datasets.

The results consistently showed that DShield exhibited "notable performance against the most backdoor attacks, surpassing the efficacy of conventional defenses." While acknowledging a "slight decline in model performance on normal nodes" due to the additional filtering process for malicious nodes, the overall benefit in terms of backdoor detection and mitigation was significant.

Further robustness evaluations were conducted by testing DShield under varying conditions:

  • Different poisoning rates: DShield maintained its effectiveness even when the proportion of poisoned nodes changed.
  • Different trigger rates: The defense proved robust to variations in how frequently triggers appeared.
  • Different adaptive attacks: This demonstrated DShield's resilience against adversaries attempting to circumvent the defense.

The speaker attributed DShield's strong efficacy to the successful execution of its self-supervised learning framework and the application of UMAP on the manipulated graph. Sensitivity analyses of hyperparameters (beta and gamma) were also presented, highlighting their impact on identification precision and victim model performance, respectively.

Finally, the applicability of DShield was extended to graph classification tasks, demonstrating that the underlying phenomena of semantic drift and attribute overemphasize are not limited to node classification. This indicates DShield's potential for broader utility across different GNN applications, reinforcing its position as a versatile defense mechanism.

Defensive Implications

▶ Watch: DShield's Applicability to Graph Classification Tasks (9:50)

The insights and mechanisms provided by DShield offer critical implications for defenders working with Graph Neural Networks. The primary takeaway is the absolute necessity of moving beyond traditional GNN training paradigms, especially when the integrity of the training data cannot be fully guaranteed. GNNs, due to their inherent ability to learn complex relationships, are highly susceptible to subtle data manipulations, and DShield demonstrates that these manipulations leave detectable footprints.

Defenders should recognize that:

  1. Standard GNN training is insufficient: Simply training a GNN on potentially compromised data, even with high accuracy on clean validation sets, does not guarantee freedom from backdoor vulnerabilities. DShield highlights how a GNN can perform well on benign inputs while harboring hidden malicious behaviors.
  2. Self-supervised learning is a powerful sentinel: The success of DShield's self-supervised model in exposing semantic drift underscores the value of learning robust, label-agnostic representations. Defenders should explore incorporating self-supervised pre-training or auxiliary self-supervised tasks to build more resilient GNN models that can detect anomalies.
  3. Discrepancy analysis is key: The concept of identifying discrepancies between models (one susceptible, one robust) or within attribute importance profiles provides a novel and effective detection strategy. GNN practitioners should consider developing monitoring tools that analyze latent representations and attribute saliency maps for unusual patterns indicative of poisoning.
  4. Targeted mitigation is feasible: DShield's ability to identify and isolate poisoned nodes and then specifically train a model to maximize loss on these nodes offers a practical pathway for remediation. Instead of retraining from scratch or simply discarding suspicious data, defenders can leverage such techniques to "cleanse" a model's learning from adversarial influences.
  5. Robustness to adaptive attacks is crucial: The evaluation against adaptive attacks emphasizes the need for defenses that are not easily bypassed. DShield's multi-faceted approach, combining semantic and attribute-level analysis, contributes to this robustness.
  6. Extensibility to other GNN tasks: The finding that semantic drift and attribute overemphasize also occur in graph classification tasks suggests that the principles behind DShield are broadly applicable. Defenders should consider how these detection mechanisms can be adapted for various GNN applications beyond node classification.

In essence, DShield calls for a proactive and analytical approach to GNN security. It encourages defenders to instrument their GNN pipelines with anomaly detection capabilities that can identify the subtle, yet distinct, signatures of backdoor attacks, thereby building more trustworthy and resilient graph-based AI systems.

Key Takeaways

  • GNNs are highly vulnerable to backdoor attacks: Both dirty-label and clean-label attacks can compromise GNN integrity by injecting hidden triggers into training data.
  • Backdoor attacks manifest as detectable discrepancies: Dirty-label attacks cause "semantic drift" in latent space, while clean-label attacks lead to "attribute overemphasize."
  • DShield leverages these discrepancies for detection: It uses a self-supervised model to learn robust graph representations and identifies poisoned nodes by comparing their behavior with a standard (backdoored) model.
  • Multi-modal detection is effective: DShield constructs both semantic and attribute importance discrepancy matrices, fusing them for comprehensive poisoned node identification using HDBSCAN.
  • Targeted mitigation is possible: Identified poisoned nodes are used to train a backdoor-free model by maximizing loss on malicious data while minimizing it on clean data.
  • DShield outperforms conventional defenses: Evaluations show its notable efficacy against various dirty-label and clean-label attacks across multiple datasets, with applicability extending to graph classification tasks.

About the Speaker(s)

Hao Yu presented this work on DShield: Defending against Backdoor Attacks on Graph Neural Networks via Discrepancy Learning at the NDSS Symposium. The provided materials do not include further biographical details or affiliations for the speaker.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Technically legitimate ML security research with a clear problem statement and a multi-component defense pipeline. The two identified phenomena — semantic drift and attribute overemphasize — are credible observations, and the method of using a self-supervised model as a reference to expose poisoned nodes is reasonably clever. But this is a paper presentation, not a security talk, and the threat model's real-world grounding is thin enough that most practitioners will struggle to connect it to anything they're actually defending.

Heather Calloway (CISO) — PASS

Technically sound academic work on a narrow GNN security problem with no meaningful bridge to governance, enterprise security operations, or institutional risk. This is outside my lane — not a penalty, just a scope call.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025