Side-Channel-Assisted Reverse-Engineering of Encrypted DNN Hardware Accelerator IP and Attack Surface Exploration

Cheng Gongye, Yukui Luo, Xiaolin Xu, Yunsi Fei

IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 5

Overview

This talk, presented by Cheng Gongye, delves into the critical and often overlooked realm of hardware security, specifically focusing on physical side-channel attacks against Deep Neural Network (DNN) hardware accelerators. The research challenges the prevailing assumption that modern, high-performance accelerators, with their inherent complexity and low signal-to-noise ratio (SNR), are impervious to such attacks. The core objective was to determine if state-of-the-art commercial DNN accelerators could be compromised to reveal their sensitive intellectual property (IP), such as model parameters (weights and biases), despite sophisticated encryption and black-box designs.

Watch on YouTube

Visual summary for Side-Channel-Assisted Reverse-Engineering of Encrypted DNN Hardware Accelerator IP and Attack Surface Exploration by Cheng Gongye, Yukui Luo, Xiaolin Xu, Yunsi Fei
Visual summary for Side-Channel-Assisted Reverse-Engineering of Encrypted DNN Hardware Accelerator IP and Attack Surface Exploration by Cheng Gongye, Yukui Luo, Xiaolin Xu, Yunsi Fei

Key moments

  1. 0:00 Introduction to side-channel attacks on complex DNN accelerators
  2. 2:00 Motivation: Targeting the AMD Xilinx DPU
  3. 3:17 Key finding: Parameters recovered, but reverse engineering crucial
  4. 5:00 Challenge: Aliasing phenomenon in DNN side-channel analysis
  5. 6:00 Necessity of detailed accelerator scheduling reverse engineering
  6. 7:00 DPU architecture and scheduling reverse engineering results
  7. 8:30 Parameter extraction attack methodology using Hamming distance

Side-Channel-Assisted Reverse-Engineering of Encrypted DNN Hardware Accelerator IP and Attack Surface Exploration

Speakers: Cheng Gongye, Offensive Hardware Security Researcher, NVIDIA; Yukui Luo; Xiaolin Xu; Yunsi Fei

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=OnwCRHN1lNc

Overview

This talk, presented by Cheng Gongye, delves into the critical and often overlooked realm of hardware security, specifically focusing on physical side-channel attacks against Deep Neural Network (DNN) hardware accelerators. The research challenges the prevailing assumption that modern, high-performance accelerators, with their inherent complexity and low signal-to-noise ratio (SNR), are impervious to such attacks. The core objective was to determine if state-of-the-art commercial DNN accelerators could be compromised to reveal their sensitive intellectual property (IP), such as model parameters (weights and biases), despite sophisticated encryption and black-box designs.

The study centered on the AMD-Xilinx DPU, a highly advanced and commercially significant DNN accelerator. The choice of this target was strategic: if its security could be demonstrably breached, it would imply that less advanced accelerators would also be vulnerable. The researchers successfully demonstrated the feasibility of recovering model parameters, but with a crucial prerequisite: the accelerator's internal architecture, data flow, and scheduling had to be meticulously reverse-engineered first. This work not only highlights a significant vulnerability in current hardware IP protection strategies but also provides a detailed methodology for analyzing and exploiting these complex systems.

The implications of this research are profound for both accelerator designers and the broader machine learning community. As DNN models become increasingly valuable and accelerators become more ubiquitous, protecting the underlying hardware and the intellectual property it processes is paramount. The findings underscore the urgent need for robust countermeasures against physical side-channel attacks, extending beyond traditional software and cryptographic protections to encompass the physical implementation details of hardware accelerators.

Background

▶ Watch: Introduction to side-channel attacks on complex DNN accelerators (0:00)

The foundational principle underpinning side-channel attacks is that every physical implementation inherently leaks information. When a computational device processes different values, subtle variations in physical emissions—such as electromagnetic (EM) emissions, power consumption, or acoustic signals—can occur. By carefully manipulating inputs and applying statistical analysis, an attacker can extract sensitive information. In the context of Deep Neural Networks, the most critical secrets are the model parameters, including the weights and biases that define the network's learned intelligence.

Previous academic research, such as attacks against CSN neural networks implemented on microcontroller units, had demonstrated the feasibility of extracting parameters using EM side-channel techniques. However, these studies operated in controlled academic settings, targeting relatively simpler hardware. Real-world DNN accelerators present a far more formidable challenge. They are designed to handle increasingly large and complex DNN models, operating at very high frequencies, utilizing advanced technology nodes, and prioritizing extreme power efficiency and highly parallel computations. Each of these factors—high frequency, advanced nodes, power efficiency, and massive parallelism—contributes to a significantly lower signal-to-noise ratio (SNR), making traditional side-channel attacks considerably more difficult. The central motivation of this study was to investigate whether this reduced SNR renders side-channel attacks entirely infeasible on such advanced hardware.

The chosen target, the AMD-Xilinx DPU (specifically the DPU IP acquired by Xilinx from DeePhi Tech for approximately $300 million), represents a cutting-edge commercial DNN accelerator. It boasts remarkable efficiency, for example, consuming only a fraction of the power compared to a GPU counterpart when running ResNet-18 on the ImageNet dataset. The DPU is also a black box; its valuable IP is encrypted and technically guarded by Xilinx, making it an ideal, challenging target for security research.

Traditional Correlation Electromagnetic Analysis (CEMA), effective against ciphers like AES, relies on identifying a single key byte by correlating predicted power consumption with measured EM emissions. AES, with its highly non-linear S-boxes, allows for the use of simple selection functions (e.g., Hamming weight of an intermediate value) that diffuse the key material effectively. However, this approach fails significantly when applied to DNN computations. If an attacker targets individual multiplications within a DNN and uses a simple selection function like the Hamming weight of an input's weight, they encounter a phenomenon known as aliasing. Aliasing occurs when multiple secret bits (or combinations of bits) produce identical selection function values, making them indistinguishable through correlation analysis. The researchers found this to be a substantial problem, affecting 191 out of 253 cases in their initial analysis. To overcome this, a more sophisticated approach, the Hamming distance model, is required. This model necessitates an intimate, detailed understanding of the accelerator's internal data flow and scheduling—the exact sequence and timing of operations involving specific inputs and weights. This crucial requirement directly led to the necessity of reverse engineering the exact schedule of the accelerator, forming the cornerstone of their attack methodology.

Key Findings

▶ Watch: Key finding: Parameters recovered, but reverse engineering crucial (3:17)

The research culminated in several critical findings that challenge conventional wisdom regarding the security of advanced DNN hardware accelerators:

Firstly, the most significant finding is the unequivocal demonstration that parameter recovery is indeed possible from a state-of-the-art commercial DNN accelerator, specifically the AMD-Xilinx DPU. This directly refutes the notion that the low signal-to-noise ratio inherent in high-performance, complex hardware makes side-channel attacks impractical or impossible.

Secondly, the study established that comprehensive reverse engineering of the accelerator's internal architecture, data flow, and scheduling is a non-negotiable prerequisite for successful parameter extraction. Unlike simpler cryptographic targets, DNN accelerators require an in-depth understanding of their operational sequences and data movement to effectively apply side-channel analysis techniques that overcome challenges like aliasing. This reverse engineering effort, primarily conducted through EM side-channel analysis itself, was identified as the most substantial and challenging part of the entire study.

Thirdly, the outcome of their extensive reverse engineering efforts was the creation of a cycle-accurate and register-accurate simulator of the DPU's computing arrays. This simulator proved to be an invaluable tool, allowing the researchers to precisely pinpoint which specific parts of the DPU were susceptible to parameter extraction attacks. Notably, the accumulation tile (ACC Tile) was identified as the primary locus of vulnerability.

Finally, by leveraging the insights gained from the reverse-engineered simulator, the researchers were able to demonstrate a practical, systematic attack on the AMD-Xilinx DPU, successfully recovering complete model parameters. This proof-of-concept not only confirms the theoretical vulnerability but also outlines a methodical approach for a competent adversary to dismantle a network layer by layer, showcasing the potential for comprehensive intellectual property theft from commercial DNN accelerators.

Technical Deep Dive

▶ Watch: Challenge: Aliasing phenomenon in DNN side-channel analysis (5:00)

The technical core of this research revolves around overcoming the limitations of traditional side-channel analysis when applied to complex DNN accelerators. The fundamental principle of EM side-channel attacks states that EM emissions are influenced by system operations and the data being processed. For instance, a simple EM trace against an ECC server might show drastically different emission amplitudes when processing a '1' versus a '0'.

Correlation Electromagnetic Analysis (CEMA) is a widely used technique to extract secret keys from multiple EM traces. In CEMA, the measured EM emissions are modeled as a combination of a selection function (V) and Gaussian noise. The selection function is crucial; it links EM measurements to data by calculating a predicted emission based on an input value and a guessed secret key byte. The attacker iterates through all possible key guesses, computing the Pearson correlation coefficient between the predicted power consumption (derived from the selection function) and the measured EM emissions. The correct key is identified when it yields the highest correlation.

However, applying this standard CEMA approach directly to DNN accelerators presents significant challenges. For ciphers like AES, the highly non-linear S-box operation ensures that a simple selection function, such as the Hamming weight of an intermediate value, provides sufficient diffusion to distinguish individual key bytes. In DNNs, if one attempts to target individual multiplications and use the Hamming weight of an input's weight as the selection function, a problem known as aliasing arises. Aliasing occurs when different secret bits or combinations of bits result in identical selection function values, making them indistinguishable to the attacker. The researchers' analysis confirmed this was a substantial hurdle, observing aliasing in 191 out of 253 possible cases. This high degree of aliasing renders the traditional Hamming weight model ineffective for DNN parameter extraction.

To circumvent aliasing, the researchers adopted a Hamming distance model. This advanced approach requires a much more detailed understanding of the data flow and scheduling within the system. Specifically, it necessitates knowing the exact sequence and timing of operations involving inputs (e.g., input 1, input 2) and weights (e.g., weight 1, weight 2) to extract multiple secret weights simultaneously. This requirement for precise timing and data flow information is the fundamental reason why reverse engineering the accelerator's exact schedule became an indispensable first step.

The research developed a comprehensive reverse engineering flow, primarily utilizing EM side-channel analysis itself. Given that DNN computations are too large to fit entirely on silicon, the DPU employs multiple levels of scheduling to divide computations into manageable segments. Within each segment, dedicated and complex scheduling mechanisms are implemented to achieve high efficiency and data reuse.

A glimpse into their reverse-engineered architecture reveals intricate parallelism. The specific configuration analyzed utilized 200 DP slices for the MAC Tile (Multiply-Accumulate Tile) and 80 DP slices for the ACC Tile (Accumulation Tile). Each MAC tile contains four DP slices, while each ACC tile contains 16 DP slices. The architecture leverages three distinct levels of parallelism:

  1. Matrix-level parallelism: MAC tiles are arranged into 5x matrices. Each column of this matrix handles one input channel, while each row processes two filters in a time-multiplexed manner, contributing to output channel parallelism.
  2. Input channel parallelism: Each MAC tile is capable of processing eight input channels simultaneously.
  3. Pixel-level parallelism: This is implied by the ability to process multiple input channels and filters concurrently, allowing for efficient processing of image data.

The convolution process within the DPU involves two primary steps:

  1. MAC Tile operation: Inputs are multiplied by two weights and then accumulated along the input channel dimension.
  2. ACC Tile operation: The accumulated values from the MAC tiles are then summed into the final output feature.

Through their exhaustive reverse engineering efforts, the researchers were able to decipher the schedule for each register within the DPU. While the detailed table of values for each DP slice in the MAC Tile across various clock cycles is presented in their full paper, the high-level takeaway is the successful construction of a cycle-accurate and register-accurate simulator of the DPU's computing arrays. This simulator was instrumented for analysis, enabling the researchers to precisely identify which specific parts of the DPU were most susceptible to parameter extraction attacks. Crucially, the accumulation tile (ACC Tile) was identified as the primary point of vulnerability, highlighting where the model's sensitive data is most exposed.

Demo / Proof of Concept

▶ Watch: DPU architecture and scheduling reverse engineering results (7:00)

Leveraging the insights and the cycle-accurate, register-accurate simulator derived from their extensive reverse engineering, the researchers demonstrated a practical parameter extraction attack methodology. The attack proceeds systematically, isolating and recovering weights layer by layer.

The initial phase of the attack involves a carefully crafted input setup:

  1. The researchers began by setting only i0, i1, and i2 (specific input values) to non-zero values, while all other inputs were set to zero.
  2. This specific configuration ensures that within the ACC Tile, there are precisely three non-zero accumulations, simplifying the analysis.

With this constrained environment, the attack proceeds as follows:

  • Deciphering W1 and W0: At the second accumulation time point, the Hamming distance concept is employed to decipher the values of weight W1 and weight W0. This involves analyzing the Hamming distance between the result of the first and second accumulations.
  • Acquiring W2: Subsequently, weight W2 is acquired by examining the Hamming distance between the third and third accumulation results. (It's worth noting the speaker stated "third and third," which might be a slight verbal slip, but implies comparison against a predicted third accumulation value or between intermediate states leading to the third accumulation).
  • Extending to other weights: To extend this technique and recover additional weights within the kernel, the researchers systematically assign non-zero values to different, selected parts of the input. This selective multiplication allows for the isolation and recovery of further weights.

Once all the weights from the first layer have been successfully extracted, the methodology progresses to subsequent layers. This process is iterative: the recovered weights from one layer can then be used to generate appropriate input patterns for the next layer, effectively "dismantling" the network layer by layer. The input patterns required for recovery are generated using an SMP server (as stated in the transcript, likely referring to a specialized setup or tool for input pattern generation in this context). This systematic approach highlights the potential for comprehensive parameter recovery, ultimately leading to the full extraction of the DNN model running on the commercial accelerator.

Defensive Implications

▶ Watch: Parameter extraction attack methodology using Hamming distance (8:30)

The findings of this research carry significant defensive implications for the design and deployment of Deep Neural Network hardware accelerators. The primary takeaway is that simply encrypting or black-boxing valuable accelerator intellectual property (IP) is insufficient to protect it from determined adversaries utilizing physical side-channel attacks.

  1. Rethink IP Protection: Accelerator designers must move beyond purely cryptographic or architectural obfuscation methods. The physical implementation itself, particularly the data flow and scheduling mechanisms, represents a critical attack surface that demands robust protection against side-channel-assisted reverse engineering.
  2. Countermeasures for Accumulation Units: The research explicitly identified the accumulation tile (ACC Tile) as a key vulnerability point for parameter extraction. This means that designers should prioritize implementing specific side-channel countermeasures within these units. Techniques such as masking, shuffling, randomization of operations, or power/EM noise injection could be considered for accumulation logic to obscure the correlation between processed data and physical emissions.
  3. Obfuscation of Internal Scheduling: Since the attack heavily relies on understanding the exact data flow and scheduling, designers should explore methods to obfuscate or randomize these aspects. Dynamic scheduling, non-deterministic data paths, or introducing dummy operations could make it significantly harder for an attacker to build an accurate cycle-accurate simulator.
  4. Awareness for Developers: Developers and users of DNN accelerators should be aware of these physical attack vectors, especially when deploying models that contain sensitive or proprietary information. While direct hardware access is required, the increasing accessibility of development boards and hardware analysis tools makes this a growing concern.
  5. Holistic Security Design: The research underscores the need for a holistic security design approach that integrates side-channel resilience from the earliest stages of hardware architecture and microarchitecture development, rather than attempting to patch vulnerabilities post-design. This includes considering the leakage characteristics of various circuit components and operations.

In essence, the security of DNN accelerators must extend from the digital realm of software and cryptography into the physical domain, acknowledging that every operation leaves a trace that a sophisticated adversary can exploit.

Key Takeaways

  • Side-channel attacks are feasible on state-of-the-art commercial DNN accelerators, specifically demonstrated on the AMD-Xilinx DPU, despite the challenges posed by low signal-to-noise ratios, high frequencies, and advanced technology nodes.
  • Comprehensive reverse engineering of the accelerator's internal architecture, data flow, and scheduling is a critical prerequisite for successful parameter extraction attacks on complex DNN hardware.
  • A cycle-accurate and register-accurate simulator, derived from reverse engineering efforts, is an invaluable tool for precisely identifying and exploiting vulnerabilities within the accelerator's computing arrays.
  • The Hamming distance model is essential for overcoming the problem of aliasing encountered when attempting side-channel analysis on DNN computations, enabling the extraction of multiple secret weights.
  • The accumulation tile (ACC Tile) was identified as a key vulnerability point within the DPU, where sensitive model parameters are susceptible to leakage during the accumulation process.
  • Existing IP protection strategies for hardware accelerators, such as encryption and black-box design, are insufficient against side-channel-assisted reverse engineering, necessitating a reevaluation of security measures to include physical attack resilience.

About the Speaker(s)

The talk was presented by Cheng Gongye, who is currently employed as an Offensive Hardware Security Researcher at NVIDIA in Santa Clara. The work presented was completed during his PhD studies at Northeastern University. The paper lists Yukui Luo, Xiaolin Xu, and Yunsi Fei as co-authors, but further biographical details for them were not provided in the transcript.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research shatters the myth that advanced DNN accelerators are too complex for side-channel attacks. By meticulously reverse-engineering the AMD-Xilinx DPU, the team demonstrated full model parameter recovery, proving that physical IP protection demands a complete re-evaluation. This is critical work for anyone building or deploying AI hardware.

Heather Calloway (CISO) — STRONG ACCEPT

This research effectively dismantles assumptions about the physical security of advanced DNN accelerators, demonstrating that even black-boxed, encrypted hardware is vulnerable to sophisticated side-channel attacks for IP theft. It provides a critical wake-up call for hardware designers and security leaders regarding the need for holistic physical security in AI systems.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024