AttackGNN: Red-Teaming GNNs in Hardware Security Using Reinforcement Learning
Vasudev Gohil (Texas A&M University), Satwik Patnaik, Dileep Kalathil, Jeyavijayan Rajendran
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
In an era defined by a globally interconnected semiconductor supply chain, the security of hardware devices faces unprecedented challenges. From malicious backdoor injections known as Hardware Trojans to sophisticated reverse engineering attempts and IP piracy, the integrity of integrated circuits is under constant threat. To counter these growing vulnerabilities, researchers have increasingly turned to advanced machine learning techniques, particularly graph neural networks (GNNs), given that circuits can naturally be represented as graphs. These GNN-based methods have demonstrated remarkable accuracy, often approaching 100%, in detecting and locating Trojans, identifying IP infringement, and facilitating reverse engineering.

Key moments
- 0:00 Introduction: Evaluating GNN robustness in hardware security
- 2:00 GNNs achieve high accuracy in various hardware security tasks
- 2:20 Posing the central question: How robust are these GNNs?
- 2:40 Clarifying the threat model: Adversarial attacks on neural networks
- 3:20 Specific constraints for AttackGNN: No GNN modification, black-box
- 4:20 Attack goal: Achieve GNN misclassification, e.g., evading Trojan detection
- 5:20 AttackGNN's core: Mapping circuit perturbation to an RL problem
AttackGNN: Red-Teaming GNNs in Hardware Security Using Reinforcement Learning
Speakers: Vasudev Gohil, PhD Candidate, Texas A&M University; Satwik Patnaik; Dileep Kalathil; Jeyavijayan Rajendran
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=8VvqF-TEfPw
Overview
In an era defined by a globally interconnected semiconductor supply chain, the security of hardware devices faces unprecedented challenges. From malicious backdoor injections known as Hardware Trojans to sophisticated reverse engineering attempts and IP piracy, the integrity of integrated circuits is under constant threat. To counter these growing vulnerabilities, researchers have increasingly turned to advanced machine learning techniques, particularly graph neural networks (GNNs), given that circuits can naturally be represented as graphs. These GNN-based methods have demonstrated remarkable accuracy, often approaching 100%, in detecting and locating Trojans, identifying IP infringement, and facilitating reverse engineering.
However, as Vasudev Gohil from Texas A&M University highlighted in his USENIX Security '24 talk, the impressive accuracy of these GNNs in benign settings does not necessarily translate to robustness in adversarial environments. The presented work, titled AttackGNN, introduces a novel reinforcement learning (RL) based framework designed to systematically evaluate and exploit the vulnerabilities of GNNs used in hardware security. By developing a method to strategically perturb circuit netlists while preserving their functionality and adhering to design rules, AttackGNN demonstrates that these highly accurate GNN detectors can be consistently fooled, achieving a 100% evasion rate across multiple target GNNs and circuits.
This research is critical because it exposes a significant gap in the current state of hardware security GNNs: their susceptibility to adversarial attacks. It serves as a stark warning to developers and researchers, emphasizing the urgent need to incorporate robustness considerations, such as adversarial training, into the design and deployment of future hardware security solutions. The findings challenge the assumption of GNN infallibility and provide a robust framework for red-teaming these advanced security tools, ultimately paving the way for more resilient hardware defense mechanisms.
Background
▶ Watch: Introduction: Evaluating GNN robustness in hardware security (0:00)
The semiconductor industry’s globalized supply chain presents inherent security risks. A processor designed in one country, with IP cores integrated from another, fabricated in a third, and assembled elsewhere, creates numerous points of potential compromise. This distributed model makes devices vulnerable to various attacks, including the insertion of Hardware Trojans—malicious modifications that can create backdoors or leak sensitive information—and IP piracy, where proprietary circuit designs are stolen or copied. Additionally, the ability to reverse engineer a circuit to understand its functionality or extract intellectual property remains a persistent threat.
In recent years, the security community has recognized the potential of graph neural networks (GNNs) to address these complex hardware security problems. The natural representation of a circuit as a graph, where gates are nodes and wires are edges, makes GNNs particularly well-suited for tasks like:
- Detecting and locating Hardware Trojans.
- Identifying design similarities for IP piracy detection.
- Assisting in circuit reverse engineering.
Many proposed GNN-based techniques have reported near-perfect accuracy, often exceeding 98% or even reaching 100% success rates in their respective tasks. This high reported performance has led to a perception of these GNNs as highly effective and reliable security tools.
However, the field of machine learning has repeatedly shown that models with high accuracy on clean data can be incredibly fragile when faced with carefully crafted adversarial examples. These are inputs intentionally perturbed by an adversary to cause a machine learning model to misclassify them, even if the perturbations are imperceptible to humans or, in this context, functionally equivalent for the hardware. The seminal work on adversarial attacks in computer vision has demonstrated how slight pixel modifications can fool image classifiers. The question AttackGNN poses is whether the same vulnerabilities exist for GNNs applied to hardware circuits.
The threat model for AttackGNN is carefully defined to reflect a realistic adversarial scenario:
- Black-box access: The adversary does not have access to the internal parameters (weights, biases) or architecture of the target GNN. They can only query the GNN with a circuit input and receive a classification output. This is a common and challenging scenario in real-world attacks.
- No modification to the GNN: The adversary cannot retrain, fine-tune, or alter the target GNN in any way. The GNN is treated as a fixed, pre-trained detector.
- Circuit design rules preserved: Any perturbations made to the circuit must adhere to valid circuit design rules. For instance, creating combinational loops or functionally altering the circuit's intended behavior (e.g., removing a Trojan's functionality) is not permitted. The goal is to create a functionally equivalent circuit that looks different enough structurally to deceive the GNN.
The problem, therefore, is to find a sequence of circuit transformations that can make a GNN misclassify a circuit, while maintaining its original functionality and adhering to design constraints. This is precisely the challenge AttackGNN sets out to address.
Key Findings
▶ Watch: Posing the central question: How robust are these GNNs? (2:20)
The central and most impactful finding of the AttackGNN research is the profound lack of robustness in existing graph neural networks (GNNs) developed for hardware security. Despite their high reported accuracies—often exceeding 97% or 98% in benign scenarios—these GNNs are highly susceptible to carefully crafted adversarial attacks. The research conclusively demonstrates that an adversary, operating under a realistic black-box threat model and adhering to strict circuit design rules, can consistently and successfully evade detection or misdirect these GNNs.
Specifically, AttackGNN achieved a 100% success rate in evading five distinct GNN-based hardware security techniques across all tested circuits. This was accomplished using a single reinforcement learning (RL) agent, highlighting the generalizability of the attack methodology. The targeted GNNs included:
- GNN for IP: A GNN designed for detecting IP piracy by measuring circuit similarity. AttackGNN successfully perturbed circuits to appear dissimilar, thus evading piracy detection.
- TrojanSyn: A GNN focused on locating Hardware Trojans. AttackGNN was able to manipulate circuits such that TrojanSyn misidentified the location of the embedded Trojan.
- GNN-RE: A GNN used for reverse engineering circuits. AttackGNN successfully evaded GNN-RE, making it difficult for the GNN to correctly analyze or understand the perturbed circuit's structure.
- Two other unnamed GNN techniques (as per the transcript, "other techniques as well").
This universal evasion capability underscores that the current generation of GNNs in hardware security is primarily optimized for accuracy on clean data, neglecting the critical aspect of robustness against sophisticated adversaries. The findings serve as a critical wake-up call, indicating that reliance on these GNNs without adversarial hardening could leave hardware systems vulnerable to exploitation. The research effectively red-teams these GNNs, providing concrete evidence that their reported high performance in benign settings does not translate to security in the face of intelligent, adaptive adversaries.
Technical Deep Dive
▶ Watch: Clarifying the threat model: Adversarial attacks on neural networks (2:40)
AttackGNN frames the problem of generating adversarial circuit perturbations as a reinforcement learning (RL) task. RL is a paradigm where an agent learns to make decisions by interacting with an environment, receiving rewards or penalties that guide its learning process. The core components of the AttackGNN RL framework are the environment, state, actions, and reward.
Reinforcement Learning Formulation
- Environment: The environment in AttackGNN is the target circuit itself, which the agent attempts to perturb. For instance, this could be a Trojan-inserted circuit that a GNN is trained to detect. The goal is to modify this circuit such that the GNN misclassifies it.
- State: The state represents the current characteristics of the circuit at any given point in the perturbation process. This is modeled as a feature vector encompassing various properties of the circuit. Examples of features include:
- Number of inputs and outputs of the circuit.
- Total number of gates.
- Total number of wires.
- Counts of different types of gates (e.g., AND, OR, XOR, NOT gates).
- Specific counts of gates with varying input fan-ins (e.g., 2-input AND gates, 3-input AND gates).
These features provide the RL agent with a comprehensive understanding of the circuit's current structural configuration.
- Actions: Initially, actions were defined as circuit transformation commands from the open-source ABC compiler tool. These are standard synthesis directives used to optimize or restructure circuits without altering their functional behavior. Examples include:
rewrite: A command that performs logic rewriting to simplify or optimize the circuit.refactor: A command that extracts common sub-expressions and restructures the logic for optimization.
The agent selects one of these commands, which is then applied to the current circuit, generating a new, functionally equivalent but structurally altered circuit.
- Reward: The reward signal guides the agent's learning. It is directly tied to the adversary's objective:
- If the perturbed circuit successfully evades the target GNN (i.e., the GNN misclassifies it), the agent receives a positive reward (denoted as
Alpha). - If the GNN still correctly classifies the circuit, the agent receives a zero reward.
This feedback loop allows the agent to iteratively learn which sequences of actions are most effective at causing misclassification.
The RL process unfolds iteratively: the agent observes the current state (circuit features), selects an action (a circuit transformation command), applies it to the circuit, and then queries the target GNN with the new circuit. Based on the GNN's response, a reward is computed, and the agent updates its internal policy to improve future action selections.
Challenges and Proposed Solutions
During the initial implementation, AttackGNN encountered several significant challenges, prompting the development of innovative solutions:
- Challenge 1: Limited Structural Variation from Initial Actions
The standard rewrite and refactor commands from the ABC compiler, while effective for optimization, often result in only minor structural changes. Since target circuits are typically already highly optimized, these operations might only alter a few gates, which was insufficient to consistently fool robust GNNs. The perturbations were not "diverse" enough to create significant adversarial examples.
Solution 1: Effective/Generalizable Actions using Synthesis Constraints
To overcome this, AttackGNN introduced a more powerful action space based on synthesis constraints. Instead of simple transformation commands, actions now involve prohibiting specific gate types during synthesis. For example, an action could be "don't use any three-input AND gates in this synthesis step." If the circuit contains such gates, the synthesis tool must replace them with functionally equivalent structures using allowed gate types (e.g., two-input AND gates, or combinations of AND gates and inverters). This approach forces much greater structural diversity and more substantial perturbations, as it fundamentally alters the types of logic primitives used in parts of the circuit. This significantly increased the agent's ability to generate circuits that evade GNNs.
- Challenge 2: Expensive Reward Computation and Training Time
Querying a target GNN for classification feedback can be computationally expensive, often taking several seconds per query. In a typical RL setup, the reward is computed at every time step (after each action), leading to excessively long training times for the agent.
Solution 2: Delayed Reward Computation
To address the runtime issue, AttackGNN implemented a strategy of delayed reward computation. Instead of querying the GNN after every single action, the reward is computed only at the end of an "episode," which comprises a sequence of multiple actions (e.g., five time steps). This approach significantly reduces the number of GNN queries, drastically cutting down the agent's training time. While this might introduce a slight trade-off in the immediate accuracy of the agent's learning, the practical gains in runtime proved substantial, making the training process feasible.
- Challenge 3: Agent Specificity to a Single GNN
The initial RL formulation trained an agent to evade a specific GNN (e.g., a GNN for Trojan detection). If the goal was to attack a different GNN (e.g., for IP piracy), a completely new agent would need to be trained from scratch. This lack of generalizability was inefficient and impractical for red-teaming multiple GNNs.
Solution 3: Contextual Formulation for Multi-GNN Evasion
AttackGNN introduced a contextual formulation to enable a single RL agent to learn how to evade multiple GNNs simultaneously. This is achieved by augmenting the agent's input with a "context" vector, typically encoded as a binary vector. Each unique context corresponds to a specific target GNN or a specific type of attack (e.g., one context for evading Trojan detectors, another for evading IP piracy detectors). The agent learns a policy that is conditioned on this context. By training the same agent on different contexts, it learns to adapt its perturbation strategy based on the specific GNN it is trying to evade. This solution allows a single, more powerful agent to generalize across various hardware security GNNs, significantly improving efficiency and versatility.
By integrating these three solutions, AttackGNN transformed from a basic RL proof-of-concept into a robust, efficient, and generalizable framework capable of systematically red-teaming diverse GNN-based hardware security applications.
Demo / Proof of Concept
▶ Watch: Attack goal: Achieve GNN misclassification, e.g., evading Trojan detection (4:20)
While the talk did not feature a live, interactive demonstration, it presented compelling results that served as a robust proof of concept for AttackGNN's capabilities. The research rigorously tested AttackGNN against multiple, state-of-the-art GNNs designed for various hardware security tasks, demonstrating a consistent and highly effective adversarial attack strategy. The key outcomes presented were:
- Evasion of GNN for IP (IP Piracy Detection):
- GNN for IP is a graph neural network designed to measure the similarity between two circuits, often used to detect IP piracy. Its similarity score typically ranges from -1 (not similar) to 1 (very similar).
- AttackGNN successfully perturbed various benchmark circuits such that the GNN for IP misclassified them. For instance, circuits that should have been identified as highly similar were made to appear dissimilar.
- Crucially, AttackGNN was not only able to evade detection for each circuit once but could generate multiple distinct variants of each circuit, all of which successfully fooled the GNN for IP. This shows the robustness and versatility of the attack.
- Evasion of TrojanSyn (Hardware Trojan Location):
- TrojanSyn is a GNN specifically developed to detect and locate Hardware Trojans within circuits.
- AttackGNN demonstrated its ability to perturb Trojan-inserted circuits in such a way that TrojanSyn failed to correctly identify the Trojan's true location. Instead, the GNN either reported no Trojan or indicated a location different from the actual insertion point.
- The modifications preserved the Trojan's functionality, meaning the malicious logic remained active within the circuit, yet its presence was obscured from the GNN detector.
- Evasion of GNN-RE (Circuit Reverse Engineering):
- GNN-RE is a GNN technique aimed at aiding in the reverse engineering of circuits, likely by predicting circuit functionality or structure from a netlist.
- AttackGNN successfully evaded GNN-RE for all tested circuits, implying that the perturbed circuits became significantly harder for the GNN to analyze or correctly interpret. This could lead to incorrect functional predictions or structural analyses, hindering reverse engineering efforts.
- Overall Success Rate and Generalizability:
- The most striking result was the 100% success rate achieved by AttackGNN in evading five different GNN techniques across all tested circuits.
- This was accomplished using a single reinforcement learning agent, thanks to the contextual formulation solution. This highlights the framework's remarkable generalizability and efficiency, allowing one agent to learn diverse evasion strategies for multiple GNN targets.
These results unequivocally demonstrate that GNNs, despite their high reported accuracies in benign settings, are highly vulnerable in adversarial scenarios. The ability of AttackGNN to consistently evade detection, misdirect location, and confuse analysis across a spectrum of hardware security GNNs provides strong empirical evidence for the need to re-evaluate the robustness of these advanced defense mechanisms.
Defensive Implications
▶ Watch: AttackGNN's core: Mapping circuit perturbation to an RL problem (5:20)
The findings presented by AttackGNN carry profound defensive implications for the field of hardware security, particularly for researchers and practitioners relying on graph neural networks (GNNs). The consistent and widespread success of AttackGNN in evading various GNN-based detectors exposes a critical vulnerability: the current generation of GNNs, while accurate on clean data, is not robust against intelligent adversaries.
Here are the key defensive actions and considerations:
- Prioritize Adversarial Robustness: The most immediate and critical implication is the urgent need to shift focus from solely optimizing for accuracy on benign datasets to explicitly building adversarial robustness into GNN-based hardware security solutions. Developers must assume an adversary will attempt to subvert their models.
- Implement Adversarial Training: The speaker explicitly mentioned adversarial training as a crucial countermeasure. This involves augmenting the training dataset with adversarial examples generated by techniques like AttackGNN. By exposing the GNN to perturbed circuits during training, it can learn to recognize and correctly classify these adversarial inputs, thereby improving its resilience. This is a common and effective strategy in other domains of machine learning security.
- Develop Robust GNN Architectures: Researchers should explore and design GNN architectures that are inherently more robust to perturbations. This could involve investigating different aggregation functions, attention mechanisms, or regularization techniques that make the model less sensitive to minor structural changes in the input graph. Novel GNN layers or entire architectures specifically designed with robustness in mind, perhaps drawing inspiration from robust vision models, are needed.
- Rethink Evaluation Metrics: Current evaluation metrics for hardware security GNNs often focus solely on accuracy, precision, and recall on standard benchmarks. These metrics are insufficient. Future evaluations must include comprehensive adversarial robustness metrics, testing models against black-box and white-box attacks to assess their real-world security posture.
- Proactive Red-Teaming: Organizations deploying or developing GNNs for hardware security should adopt a proactive red-teaming approach. Tools and methodologies similar to AttackGNN should be regularly employed to stress-test their deployed GNNs, identify vulnerabilities, and iterate on defense mechanisms before adversaries exploit them. This includes using a diverse set of perturbation strategies, not just those discovered by AttackGNN.
- Layered Defense Strategies: Rather than relying on a single GNN as the sole line of defense, a layered security approach is advisable. This could involve combining GNNs with other detection techniques (e.g., formal verification, statistical analysis, hardware-based monitors) to create a more resilient overall system where the failure of one component does not compromise the entire defense.
- Understand Attack Surface: The research highlights that even functionally equivalent circuit transformations can create adversarial examples. Defenders need a deeper understanding of the "attack surface" presented by different circuit representations and synthesis flows. What types of transformations are most potent for adversaries, and how can these be mitigated?
In essence, AttackGNN serves as a critical warning: the perceived "close to 100% accuracies" of GNNs in hardware security are misleading in the presence of an adversary. The community must move beyond benign accuracy and embrace the challenges of adversarial machine learning to build truly secure and resilient hardware.
Key Takeaways
- Existing GNNs in Hardware Security are Vulnerable: Despite high reported accuracies (often >98%), current graph neural networks used for hardware security tasks like Trojan detection, IP piracy identification, and reverse engineering are highly susceptible to adversarial attacks.
- AttackGNN Achieves Universal Evasion: The AttackGNN framework, using a single reinforcement learning agent, successfully achieved a 100% evasion rate against five different GNN-based hardware security techniques across all tested circuits, demonstrating a profound lack of robustness.
- Black-Box Adversarial Model is Realistic: The attack operates under a practical black-box threat model, where the adversary has no knowledge of the GNN's internal architecture or parameters, making the demonstrated vulnerabilities highly relevant to real-world scenarios.
- Novel RL Approach for Circuit Perturbation: AttackGNN effectively maps circuit perturbation to a reinforcement learning problem, utilizing circuit properties as state, synthesis constraints as powerful actions, and GNN misclassification as the reward signal.
- Key Innovations for Practicality: The research introduced critical solutions to overcome RL challenges, including "effective/generalizable actions" (synthesis constraints) for greater structural diversity, "delayed reward computation" for faster training, and a "contextual formulation" for a single agent to evade multiple GNNs.
- Urgent Need for Adversarial Training: The findings underscore the critical necessity for incorporating adversarial training and other robustness-enhancing techniques into the development and deployment of future GNN-based hardware security solutions to ensure their resilience against sophisticated adversaries.
About the Speaker(s)
The primary presenter for this work was Vasudev Gohil, a PhD Candidate from Texas A&M University. His research, as evidenced by AttackGNN, focuses on evaluating the robustness of advanced machine learning techniques, particularly graph neural networks, in the context of hardware security. He is interested in understanding and addressing the vulnerabilities of these systems when confronted with adversarial manipulations.
The work is a joint effort between Texas A&M University and the University of Delaware, with co-authors including Satwik Patnaik, Dileep Kalathil, and Jeyavijayan Rajendran. While specific titles and affiliations for the co-authors beyond their university connection were not detailed in the transcript, their involvement highlights a collaborative academic endeavor to advance the understanding of security and robustness in hardware-oriented machine learning.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This research uncovers a critical blind spot in GNN-based hardware security, demonstrating 100% evasion against multiple detectors using a single, cleverly designed RL agent. It's a stark, necessary warning that 'near 100% accuracy' means squat against a determined adversary, providing concrete pathways for more robust defenses.
Heather Calloway (CISO) — STRONG ACCEPT
This research delivers a stark, necessary warning: the high accuracy reported by GNNs in hardware security is misleading. Their 100% evasion rate demonstrates a critical lack of adversarial robustness, exposing organizations to significant supply chain risks. This demands an immediate shift towards adversarial training and proactive red-teaming in ML defense strategies.