DeepTheft: Stealing DNN Model Architectures through Power Side Channel

Yansong Gao, Huming Qiu, Zhi Zhang, Binghui Wang, Hua Ma, Alsharif Abuadbba

IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 5

Overview

In the rapidly expanding landscape of cloud-based machine learning services, Deep Neural Network (DNN) models are increasingly deployed to provide inference capabilities for various applications. While this paradigm offers scalability and accessibility, it simultaneously introduces a new frontier for intellectual property theft and adversarial attacks. The talk "DeepTheft: Stealing DNN Model Architectures through Power Side Channel," presented by Yansong Gao and his collaborators at IEEE S&P, unveils a sophisticated learning-based framework capable of accurately recovering DNN model architectures, including layer types and hyperparameters, by exploiting power and frequency side channels on general-purpose CPUs.

Watch on YouTube

Visual summary for DeepTheft: Stealing DNN Model Architectures through Power Side Channel by Yansong Gao, Huming Qiu, Zhi Zhang, Binghui Wang, Hua Ma, Alsharif Abuadbba
Visual summary for DeepTheft: Stealing DNN Model Architectures through Power Side Channel by Yansong Gao, Huming Qiu, Zhi Zhang, Binghui Wang, Hua Ma, Alsharif Abuadbba

Key moments

  1. 0:00 Introduction to DeepTheft: Stealing DNN Model Architectures
  2. 1:34 Why model architecture is a valuable intellectual property
  3. 2:36 Limitations of existing model stealing attack methods
  4. 3:29 DeepTheft: Recovering model structure with blackbox access
  5. 4:22 Exploiting Power (RAPL) and Frequency (DVFS) side channels
  6. 5:40 RAPL interface access demonstrated on AWS EC2 instances
  7. 6:17 DeepTheft's two-step strategy: layer segmentation and parameter recovery
  8. 8:00 DeepTheft's design: offline training and online inference phases

DeepTheft: Stealing DNN Model Architectures through Power Side Channel

Speakers: Yansong Gao, Huming Qiu, Zhi Zhang, Binghui Wang, Hua Ma, Alsharif Abuadbba

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=zwB3G5n_Uvk

Overview

In the rapidly expanding landscape of cloud-based machine learning services, Deep Neural Network (DNN) models are increasingly deployed to provide inference capabilities for various applications. While this paradigm offers scalability and accessibility, it simultaneously introduces a new frontier for intellectual property theft and adversarial attacks. The talk "DeepTheft: Stealing DNN Model Architectures through Power Side Channel," presented by Yansong Gao and his collaborators at IEEE S&P, unveils a sophisticated learning-based framework capable of accurately recovering DNN model architectures, including layer types and hyperparameters, by exploiting power and frequency side channels on general-purpose CPUs.

This research addresses a critical vulnerability in the deployment of proprietary DNN models in multi-tenant cloud environments. The architecture of a DNN is not merely a structural detail; it is a fundamental determinant of model performance, a valuable intellectual property for developers, and a prerequisite for launching more potent adversarial attacks such as model weight stealing, membership inference, and backdoor attacks. DeepTheft demonstrates that an attacker, with only black-box access to the target model and co-location on the same physical machine in the cloud, can reconstruct complex DNN architectures with unprecedented accuracy, challenging the current security assumptions of cloud-based AI services.

The significance of DeepTheft lies in its ability to overcome the limitations of prior model stealing attempts, which often required physical proximity or yielded low accuracy. By employing a novel two-step "divide and conquer" strategy, coupled with a hybrid U-Net and Bidirectional LSTM metamodel and innovative loss functions, DeepTheft effectively extracts architectural secrets from low-resolution side channel traces. This work not only highlights a potent new threat but also underscores the urgent need for enhanced isolation and side-channel-aware design principles in cloud computing infrastructure supporting AI workloads.

Background

▶ Watch: Introduction to DeepTheft: Stealing DNN Model Architectures (0:00)

The proliferation of Machine Learning as a Service (MLaaS) platforms has made DNN models accessible to a wide range of clients, who can train, deploy, and utilize these models in the cloud. However, this convenience comes with inherent security risks, particularly concerning the intellectual property embedded within the deployed models. Among these threats, model stealing attacks aim to illicitly acquire components of a victim model. This work specifically focuses on the theft of model architecture, which is a critical aspect of a DNN's design.

The architecture of a DNN, encompassing the number and types of layers (e.g., convolutional, pooling, linear, activation) and their associated hyperparameters (e.g., kernel size, stride, padding, number of output channels), is paramount to its performance. Developing high-performance architectures is either a labor-intensive manual process (exemplified by models like ResNet and DenseNet) or a computationally expensive automated process like Neural Architecture Search (NAS). Consequently, a model architecture represents a significant investment and a valuable trade secret for its developers.

Exposure of a model's architecture can lead to several severe consequences. Firstly, it constitutes intellectual property infringement, directly undermining the commercial value of proprietary models. More critically, knowledge of the architecture serves as a foundational prerequisite for numerous advanced adversarial attacks. For instance, the "DeepSteal" work, which focuses on stealing model weights, explicitly assumes prior knowledge of the victim model's architecture. Other attacks, including membership inference attacks (determining if a specific data point was part of the training set), data reconstruction attacks (reconstructing training data from model outputs), and backdoor attacks (injecting malicious functionality into the model), all benefit significantly or even require architectural insights to be effective.

Previous research into model stealing attacks has explored various side channels, but most have significant limitations. Many approaches, such as those leveraging electromagnetic side channels, necessitate physical or approximate access to the victim machine, often within a range of 30 meters. Techniques like cache telepathy attempt to avoid physical access by exploiting cache side channels but suffer from low attack accuracy. Another notable work, using Rowhammer attacks, focused on stealing model weights and similarly assumed the model architecture was already known by the attacker. These limitations left a critical research gap: could complex DNN architectures be accurately and stealthily stolen from general CPUs in a cloud environment using other side channels, with minimal system access? DeepTheft emerges as a direct response to this question, proposing a learning-based framework that addresses these challenges by leveraging power and frequency side channels under a weak, black-box threat model.

Key Findings

▶ Watch: Limitations of existing model stealing attack methods (2:36)

DeepTheft conclusively demonstrates the feasibility of accurately and stealthily stealing complex DNN model architectures, including layer types and their specific hyperparameters, from black-box inference services deployed on general CPUs in cloud environments. The core findings are:

  1. Exploiting General CPU Side Channels: The research successfully leverages widely available power and frequency side channels on both Intel and AMD CPU processors. Specifically, the Running Average Power Limit (RAPL) interface provides power consumption data, and the Dynamic Voltage and Frequency Scaling (DVFS) interface, accessible via the C-Bleed vulnerability, provides frequency information. The authors confirmed that RAPL access is permitted on certain AWS EC2 instances, highlighting the real-world applicability of their attack.
  1. Weak Threat Model and High Accuracy: DeepTheft operates under a weak threat model where the attacker only requires black-box access to the victim DNN model (i.e., can submit inference queries) and co-location on the same physical machine in a multi-tenant cloud setting. Despite this minimal access, the framework achieves remarkably high accuracy: close to 100% for layer segmentation (LDA and SA metrics) and over 99.5% for layer-wise hyperparameter recovery (Precision, Recall, and F1-score). This level of accuracy significantly surpasses previous side-channel-based model stealing attempts.
  1. Novel Two-Step Divide-and-Conquer Strategy: To overcome the inherent challenges of low-resolution side channel data (e.g., RAPL's 1 kHz sampling rate, which is 47 times lower than some prior electromagnetic side-channel works), DeepTheft employs an innovative two-step approach:
  • Layer Segmentation: The initial step accurately segments the continuous energy trace into distinct pieces, each corresponding to an individual layer within the DNN.
  • Hyperparameter Recovery: Following segmentation, specific metamodels are applied to each identified layer to recover its unique hyperparameters.
  1. Hybrid Metamodel Architecture and Loss Functions: A crucial contribution is the design of a novel metamodel architecture that combines U-Net and Bidirectional Long Short-Term Memory (LSTM) networks. The U-Net excels at capturing spatial features, while the Bidirectional LSTM is adept at processing temporal sequences. This hybrid design is specifically tailored to extract meaningful patterns from noisy, low-frequency time-series side channel data. Furthermore, two innovative loss functions, the sampling point independent loss (for spatial classification) and the cross sampling point loss (for contextual information and correcting misclassifications), were developed to guide the metamodel training effectively, ensuring accurate layer segmentation and path selection.
  1. Incorporation of Domain Knowledge: The framework intelligently integrates domain knowledge to optimize the hyperparameter recovery process. By recognizing that some hyperparameters can be deterministically derived from others (e.g., padding size from kernel size and dilation) or are commonly fixed (e.g., dilation often set to 1), the number of parameters requiring explicit recovery via side channels is significantly reduced, enhancing efficiency and accuracy.
  1. Creation and Release of a Large-Scale Dataset: Recognizing the lack of public datasets for side-channel-based model architecture recovery, the authors collected and released a substantial dataset. This dataset comprises 11,152 distinct model architectures, including randomly generated ones and standard models like Xception, VGG, and ResNet, under various configurations and network depths (2 to 152 layers, with ResNet-52 being the deepest). It covers five different input sizes, resulting in 55,760 energy traces and a total size of approximately 20 GB. This resource is invaluable for facilitating future research in this domain.

Technical Deep Dive

▶ Watch: Exploiting Power (RAPL) and Frequency (DVFS) side channels (4:22)

The DeepTheft framework operates on a sophisticated understanding of side-channel leakage during DNN inference and employs advanced machine learning techniques to reconstruct architectural details.

Threat Model and Attack Setup

The threat model for DeepTheft is particularly relevant to cloud environments. The attacker is assumed to have:

  • Black-box access to the victim DNN model, meaning they can only submit input queries and receive inference results. The inference results themselves are irrelevant to the attack; only the side-channel information during execution matters.
  • Co-location on the same physical machine as the victim tenant in a multi-tenant cloud (e.g., an AWS EC2 instance). This implies the attacker has unprivileged user access to a process running alongside the victim's model. Physical proximity or privileged access to the victim's system is not required.

The attack capitalizes on the fact that different DNN layer operations (e.g., convolution, pooling, linear transformations, activations) consume varying amounts of power and exhibit distinct frequency scaling behaviors. These variations create unique "signatures" in the side-channel traces during the model's forward inference computation.

Side Channels Utilized

DeepTheft primarily evaluates two types of side channels accessible on commodity CPUs:

  1. Power Side Channel (RAPL):
  • RAPL (Running Average Power Limit) is an interface provided by Intel and AMD processors for reporting accumulated energy consumption across various power domains (e.g., CPU package, DRAM).
  • The vulnerability of RAPL for side-channel leakage was disclosed by Platers in 2021 (Oakland).
  • A key challenge with RAPL is its limited sampling rate (only 1 kHz in the authors' experiments) and low resolution, making fine-grained analysis difficult compared to other side channels like electromagnetic emissions.
  1. Frequency Side Channel (C-Bleed/DVFS):
  • This channel exploits Dynamic Voltage and Frequency Scaling (DVFS), a CPU feature that adjusts processor frequency to manage power consumption and temperature.
  • The C-Bleed vulnerability, disclosed by H-Bleed in 2022 (USENIX Security), enables remote attacks by observing execution time differences caused by DVFS, even with unprivileged user access.
  • The frequency traces, like power traces, exhibit dependencies on layer types and hyperparameters, offering another avenue for architectural inference.

The authors specifically evaluated AWS EC2 instances and found that some instances allow access to the RAPL interface, confirming the practical threat in cloud deployments.

DeepTheft Framework Design

DeepTheft is a learning-based framework that operates in two main phases: an offline training phase and an online attack phase.

Offline Phase: Metamodel Training

During the offline phase, a set of metamodels are trained using a large dataset of energy traces collected from various DNN architectures with known ground truth labels. This involves:

  1. Data Collection: Running a diverse set of DNN models (including randomly generated architectures and standard models like Xception, VGG, ResNet) with different configurations and input sizes.
  2. Side Channel Monitoring: Simultaneously monitoring the CPU's power (via RAPL) or frequency (via DVFS/C-Bleed) consumption during each model's inference execution.
  3. Metamodel Training: Training specialized machine learning models (metamodels) to learn the correlation between the side-channel traces and the underlying DNN architecture.

Online Phase: Architecture Recovery

In the online phase, when an attacker targets an unseen victim DNN model:

  1. Trace Collection: The attacker sends inference queries to the victim model and collects the corresponding side-channel trace (power or frequency) during its execution.
  2. Layer Segmentation: The first metamodel processes the collected trace to segment it into distinct pieces, each corresponding to a specific layer in the victim's DNN.
  3. Hyperparameter Recovery: For each segmented layer, a corresponding metamodel is then used to recover its specific hyperparameters.
  4. Full Architecture Reconstruction: By combining the identified layer types and their recovered hyperparameters, the complete DNN model architecture is reconstructed.

Two-Step Divide-and-Conquer Strategy

The core of DeepTheft's effectiveness lies in its two-step strategy, designed to address the challenges posed by low-resolution side-channel data:

  1. Layer Segmentation:
  • Challenge: The low sampling rate of RAPL (1 kHz) means that short-duration layers (e.g., average pooling, linear layers) might only produce a few sample points, making accurate segmentation difficult. This contrasts sharply with prior electromagnetic side-channel works that achieved sampling rates of 47 kHz.
  • Solution: Hybrid U-Net and Bidirectional LSTM Metamodel: The authors designed a novel metamodel architecture specifically for layer segmentation.
  • U-Net: This architecture, widely used in image segmentation, is adept at capturing spatial features (patterns within segments of the trace). It uses an encoder-decoder structure with skip connections, allowing it to learn both local and global contextual information.
  • Bidirectional LSTM (Bi-LSTM): LSTMs are powerful for processing temporal features in time series data, and the bidirectional nature allows the model to consider context from both past and future time steps in the trace.
  • The hybrid design allows the metamodel to effectively capture both the spatial characteristics of different layer operations and their temporal dependencies within the execution flow.
  • Optimization Objectives (Loss Functions): Two novel loss functions were introduced to guide the metamodel training:
  • Sampling Point Independent Loss: This loss function classifies each individual sampling point into a layer type (e.g., point 1 classified as "convolution," point 2 as "convolution," point 3 as "ReLU"). It primarily focuses on capturing spatial information.
  • Cross Sampling Point Loss: This crucial loss function incorporates contextual information across sampling points. It serves two main purposes:
  • Correction of Misclassified Points: By considering the surrounding context, it can correct individual sampling points that might be wrongly classified by the independent loss.
  • Enforcing Correct Layer Paths: It helps the model choose the most probable sequence of layers (a "layer path") given the overall characteristics of the trace, preventing illogical transitions between layer types. For example, if a model is known to have three layers, this loss helps ensure the predicted sequence (e.g., Linear -> ReLU -> Linear) aligns with plausible architectural patterns.
  1. Hyperparameter Recovery:
  • Layer-Wise Strategy: Once the layers are segmented and identified, the framework proceeds to recover specific hyperparameters for each layer.
  • Incorporating Domain Knowledge: A key aspect here is the strategic use of domain knowledge to reduce the number of hyperparameters that need to be explicitly recovered through side-channel analysis. For instance:
  • The padding size of a convolutional layer can be deterministically derived once the kernel size and dilation are known.
  • The dilation parameter is commonly set to 1 in many DNNs.
  • Some layers, like ReLU and Average Pooling, may have no hyperparameters that require recovery.
  • For layers requiring recovery, a dedicated metamodel (or set of metamodels) is trained to map the segmented layer's trace signature to its specific hyperparameter values. For example, a convolution layer might require recovery of 3 hyperparameters, a max pooling layer 3, and a linear layer 1.

Data Collection and Evaluation

The authors undertook a massive data collection effort due to the absence of public datasets for this specific problem. Their dataset includes:

  • 11,152 unique model architectures, both randomly generated and standard models (Xception, VGG, ResNet).
  • Network depths ranging from 2 to 152 layers (the deepest being ResNet-52).
  • Five different input sizes (e.g., 3x3x1, 331x331x3 for color images).
  • Totaling 55,760 energy traces (80% for training, 20% for testing), approximately 20 GB in size, and released to the community.

Evaluation Metrics:

  • Layer Segmentation Accuracy:
  • LDA (Layer-wise Distance Accuracy): Measures the similarity between the predicted and ground truth structural sequence.
  • SA (Segmentation Accuracy): The percentage of sampling points correctly classified into their ground truth layer types.
  • Hyperparameter Recovery Accuracy:
  • Precision, Recall, and F1-score.

Results: DeepTheft achieved exceptional performance. For layer segmentation, both LDA and SA were close to 100%. For hyperparameter recovery, all metrics (Precision, Recall, F1) exceeded 99.5%. These results were consistent whether using power (RAPL) or frequency (C-Bleed) traces, with frequency traces yielding similar LDA and SA of about 99%.

This detailed technical approach, combining advanced neural network architectures with domain-specific knowledge and robust evaluation, establishes DeepTheft as a highly effective and practical framework for stealing DNN model architectures through side channels.

Demo / Proof of Concept

▶ Watch: RAPL interface access demonstrated on AWS EC2 instances (5:40)

While the talk did not feature a live, interactive demonstration of the DeepTheft framework in action, the authors extensively detailed its methodology, the architecture of their metamodels, and presented comprehensive empirical results to validate their claims. The "Proof of Concept" is primarily embodied in the rigorous experimental evaluation of the DeepTheft framework against a large and diverse dataset of DNN models. This included:

  • Data Collection: Description of how 55,760 energy traces were collected from 11,152 different model architectures (ranging from 2 to 152 layers, including standard models like Xception, VGG, and ResNet) across five input sizes. This massive dataset, totaling approximately 20 GB, serves as the empirical foundation for their work.
  • Performance Metrics: The presentation of quantitative results, such as Layer-wise Distance Accuracy (LDA) and Segmentation Accuracy (SA) for layer segmentation (both close to 100%), and Precision, Recall, and F1-score (all >99.5%) for hyperparameter recovery, serves as concrete evidence of the framework's efficacy.
  • Visualizations: The talk included visual representations of power and frequency traces, allowing the audience to qualitatively observe how different layer types and hyperparameters manifest distinct patterns in these side channels. This visual evidence supports the premise that such information is indeed leaked and recoverable.

The detailed explanation of the hybrid U-Net and Bidirectional LSTM architecture, the innovative loss functions, and the overall two-step "divide and conquer" strategy constitutes the technical proof of concept, illustrating how the attack works with high accuracy under realistic cloud conditions. The release of the dataset further supports reproducibility and allows other researchers to validate and build upon their findings.

Defensive Implications

▶ Watch: DeepTheft's design: offline training and online inference phases (8:00)

The DeepTheft research presents significant implications for the security of DNN models deployed in multi-tenant cloud environments, necessitating a multi-faceted defensive strategy:

  1. Cloud Provider Responsibilities:
  • Enhanced Isolation: Cloud providers must re-evaluate and strengthen the isolation mechanisms between co-located tenants on shared physical hardware. The ability of an unprivileged process to access interfaces like RAPL (Running Average Power Limit) or exploit vulnerabilities like C-Bleed (for DVFS/frequency monitoring) directly enables this class of attack.
  • Side-Channel Mitigation: Implement stricter controls or virtualization layers around power and frequency monitoring interfaces. If these interfaces cannot be completely isolated without impacting legitimate performance monitoring, then techniques to obfuscate or add noise to the reported data should be explored.
  • Hardware-Software Co-design: Collaborate with hardware manufacturers to develop processors and platforms inherently more resistant to side-channel leakage, especially in shared resource scenarios. This could involve hardware-enforced partitioning of power domains or more robust obfuscation of power/frequency signals.
  1. DNN Model Developer and Provider Strategies:
  • IP Protection Awareness: Model developers and providers must acknowledge that their model architectures, a key intellectual property, are vulnerable to extraction even with black-box access in cloud deployments. This threat undermines the value proposition of proprietary models.
  • Architectural Obfuscation: Explore techniques to make the power and frequency signatures of their models less distinctive. This could involve:
  • Adding "Noise": Introducing dummy operations or varying execution paths that do not affect the model's output but perturb its side-channel signature.
  • Dynamic Execution: Varying the execution order of independent operations or using dynamic scheduling to make side-channel traces less predictable.
  • Homogeneous Layer Stacks: Designing models with more uniform layer types or operations that have similar power/frequency profiles, making segmentation more challenging.
  • Differential Privacy/Noise Injection: While challenging, applying concepts from differential privacy to the execution profile itself, or injecting random noise into the computation, might disrupt the patterns that DeepTheft relies upon.
  • Monitoring and Anomaly Detection: Implement robust monitoring of inference services for unusual access patterns or attempts to collect side-channel data. While difficult to detect directly, large numbers of repetitive, non-standard inference queries might be indicative of an attack.
  1. General Security Posture:
  • Defense in Depth: DeepTheft reinforces the need for a multi-layered security approach. Relying solely on software-level isolation is insufficient when hardware-level side channels can be exploited.
  • Research and Development: Continued research into side-channel countermeasures for AI workloads, particularly in cloud and edge computing environments, is crucial. This includes exploring new hardware primitives, secure execution environments, and cryptographic techniques that can protect model integrity and confidentiality during inference.
  • Supply Chain Security: Ensuring the underlying hardware and virtualization layers are secure and free from known side-channel vulnerabilities is paramount.

In conclusion, the DeepTheft work serves as a stark reminder that side-channel attacks remain a potent threat, evolving to target complex intellectual property like DNN architectures. Mitigating this risk requires a concerted effort across the entire technology stack, from chip design to cloud infrastructure management and AI model development.

Key Takeaways

  • DeepTheft demonstrates a highly accurate method for stealing Deep Neural Network (DNN) model architectures, including layer types and hyperparameters, by exploiting power (RAPL) and frequency (C-Bleed/DVFS) side channels on general-purpose CPUs.
  • The attack operates under a weak threat model, requiring only black-box access to the victim model and co-location on a shared physical machine in a multi-tenant cloud environment, without requiring physical access or privileged user access.
  • A novel hybrid metamodel, combining U-Net for spatial feature extraction and Bidirectional LSTM for temporal feature processing, along with specialized sampling point independent and cross sampling point loss functions, effectively overcomes the challenges of low sampling rates (e.g., RAPL's 1 kHz) in side-channel data.
  • The framework achieves near-perfect accuracy, with layer segmentation metrics (LDA and SA) close to 100% and hyperparameter recovery metrics (Precision, Recall, F1) exceeding 99.5%, across a diverse set of 11,152 DNN architectures and various input sizes.
  • This research highlights a significant intellectual property and security risk for DNN model providers, as architectural knowledge is a critical prerequisite for launching more advanced adversarial attacks like model weight stealing, membership inference, and backdoor attacks.
  • Cloud providers and DNN developers must implement stronger isolation mechanisms, explore side-channel resistant model designs, and consider hardware-level countermeasures to protect proprietary AI models from sophisticated side-channel attacks like DeepTheft.

About the Speaker(s)

The talk "DeepTheft: Stealing DNN Model Architectures through Power Side Channel" was presented by Yansong Gao. The work is a collaborative effort involving researchers from several institutions, including Fudan University, the University of Western Australia, an unnamed Institute of Technology, and Nanjing University of Science and Technology. The list of co-authors includes Huming Qiu, Zhi Zhang, Binghui Wang, Hua Ma, and Alsharif Abuadbba, signifying a broad academic collaboration across multiple universities. While specific roles beyond the presenter (Yansong Gao) are not detailed in the transcript, the collective expertise of the team enabled this in-depth exploration of side-channel vulnerabilities in DNN architectures.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research presents a groundbreaking and highly accurate method, DeepTheft, for stealing DNN model architectures through low-resolution power and frequency side channels on general CPUs in cloud environments. Its novel two-step strategy and hybrid metamodel overcome prior limitations, achieving near-perfect architectural recovery with significant implications for cloud providers and AI IP protection. This is a critical development that reshapes our understanding of model security in multi-tenant systems.

Heather Calloway (CISO) — STRONG ACCEPT

This research compellingly demonstrates a potent side-channel attack for stealing proprietary DNN model architectures in cloud environments with near-perfect accuracy. It exposes a critical intellectual property and security vulnerability, demanding immediate attention from cloud providers and MLaaS developers to strengthen isolation and implement architectural obfuscation.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024