AI Psychiatry: Forensic Investigation of Deep Learning Networks in Memory Images

David Oygenblik, Carter Yagemann, Joseph Zhang, Arianna Mastali, Jeman Park, Brendan Saltaformaggio

33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24

Overview

In an increasingly AI-driven world, the integrity and security of deep learning (DL) models are paramount, especially in safety-critical applications like autonomous vehicles. This talk, "AI Psychiatry: Forensic Investigation of Deep Learning Networks in Memory Images," presented by David Oygenblik and his collaborators at USENIX Security '24, introduces a novel memory forensics framework called AI Psychiatry (APE). APE is designed to address a critical gap in current incident response capabilities for AI systems: the ability to forensically examine compromised or misbehaving DL models directly from system memory, particularly when those models are proprietary, encrypted, or subject to runtime modifications through online learning.

Watch on YouTube

Visual summary for AI Psychiatry: Forensic Investigation of Deep Learning Networks in Memory Images by David Oygenblik, Carter Yagemann, Joseph Zhang, Arianna Mastali, Jeman Park, Brendan Saltaformaggio
Visual summary for AI Psychiatry: Forensic Investigation of Deep Learning Networks in Memory Images by David Oygenblik, Carter Yagemann, Joseph Zhang, Arianna Mastali, Jeman Park, Brendan Saltaformaggio

Key moments

  1. 0:00 Problem: Existing DL model vetting tools fail
  2. 2:10 Introducing AI Psychiatry: A memory forensic tool
  3. 2:58 Step 1: Recovering the deep learning model structure
  4. 4:18 Step 2: Feedback-driven tensor recovery methodology
  5. 5:50 Step 3: Rehosting the recovered model for reuse
  6. 6:30 Evaluation: Model recovery accuracy and rehosting success

AI Psychiatry: Forensic Investigation of Deep Learning Networks in Memory Images

Speakers: David Oygenblik, Carter Yagemann, Joseph Zhang, Arianna Mastali, Jeman Park, Brendan Saltaformaggio

Conference: USENIX Security '24

YouTube: https://www.youtube.com/watch?v=RWhFZxeOv8Y

Overview

In an increasingly AI-driven world, the integrity and security of deep learning (DL) models are paramount, especially in safety-critical applications like autonomous vehicles. This talk, "AI Psychiatry: Forensic Investigation of Deep Learning Networks in Memory Images," presented by David Oygenblik and his collaborators at USENIX Security '24, introduces a novel memory forensics framework called AI Psychiatry (APE). APE is designed to address a critical gap in current incident response capabilities for AI systems: the ability to forensically examine compromised or misbehaving DL models directly from system memory, particularly when those models are proprietary, encrypted, or subject to runtime modifications through online learning.

The core problem APE tackles is the inaccessibility of deployed DL models for post-incident analysis. Traditional model vetting tools assume full access to the model's source code or a fully instrumentalized version, a luxury often unavailable in real-world scenarios involving commercial products or embedded systems. APE proposes a solution that bypasses these limitations by extracting the model's structure and weights directly from CPU and GPU memory images, then rehosting them into a usable framework for comprehensive security analysis. This capability is vital for understanding why an AI system might have failed or been exploited, much like a digital detective investigating a car crash caused by a backdoored sign recognition model.

The significance of APE cannot be overstated. As AI systems become more ubiquitous, the need for robust forensic tools to investigate their behavior post-incident becomes urgent. APE empowers security professionals and incident responders to conduct independent analyses of deployed AI models, identifying subtle backdoors, data poisoning effects, or other malicious alterations that might have occurred in a live environment. By providing a pathway to reconstruct and vet models that were previously opaque, APE significantly enhances the security posture and accountability of AI-powered applications across various domains.

Background

▶ Watch: Problem: Existing DL model vetting tools fail (0:00)

The proliferation of deep learning models into critical infrastructure and everyday devices, from autonomous vehicles to industrial control systems, introduces a new attack surface. These models are susceptible to various sophisticated attacks, including adversarial examples, where subtle perturbations to input data can lead to misclassification; data poisoning attacks, where malicious data injected into the training set can compromise the model's learning; and backdoors, where specific triggers cause the model to behave anomalously. The consequences of such attacks, especially in safety-critical domains like self-driving cars, can be catastrophic, leading to accidents, system failures, or data breaches.

A major hurdle in investigating such incidents is the inherent difficulty in accessing and analyzing the deployed deep learning model at the time of an event. Current DL model vetting tools, while powerful, typically operate under several idealistic assumptions:

  1. Model Access: They require direct access to the model, often at a source code level or in a fully usable, instrumentalized format. This assumption breaks down when dealing with proprietary models, which are often encrypted or intentionally obfuscated by vendors to protect intellectual property.
  2. Static Models: These tools often assume that the model's weights and structure are static and known. However, many modern AI systems, particularly those in embedded devices, employ online learning or continual learning, where model weights are refined and updated at runtime based on new data or environmental feedback. This means the model present at the time of an incident might be different from any initial release version.
  3. Flat Buffers: Even if an investigator could magically obtain the model's weights as "flat buffers" (raw binary data), existing vetting tools are not equipped to interpret and use these without the surrounding model architecture and framework context. They require a fully functional, executable model.

The talk illustrates this problem with a compelling scenario: a self-driving Tesla car, operating in full self-driving mode, encounters a seemingly innocuous merge sign. Unbeknownst to the driver, attackers had successfully backdoored the car's sign recognition model. As a result, the car ignores the merge sign, leading to an accident. To investigate this, a forensic analyst (represented by "Detective Pikachu") needs to determine if the sign recognition model was indeed compromised. The challenge is immense: how does one obtain the specific weights that were active on the car at the exact moment of the crash, especially if they were encrypted or dynamically updated? Furthermore, once obtained, how can these raw weights be translated into a format usable by existing ML vetting tools? This complex problem underscores the critical need for a new forensic methodology, which AI Psychiatry aims to provide.

Key Findings

▶ Watch: Step 1: Recovering the deep learning model structure (2:58)

The research behind AI Psychiatry (APE) presents several pivotal findings that fundamentally advance the field of deep learning forensics:

  1. Memory-Based Model Recovery: APE demonstrates the feasibility and effectiveness of recovering both the structural topology and the specific weights (tensors) of deep learning models directly from CPU and GPU memory images. This capability is crucial for overcoming barriers like model encryption, proprietary formats, and the dynamic nature of online learning.
  2. High Accuracy and Fidelity: The evaluation showcases APE's ability to recover models with 100% accuracy across various architectures and frameworks. This high fidelity ensures that the recovered model is an exact replica of the deployed model at the time the memory image was captured, which is paramount for maintaining forensic integrity. The system successfully recovered models with upwards of 94 million weights, demonstrating its scalability.
  3. Successful Model Rehosting: Beyond recovery, APE successfully rehosts these extracted models into standard deep learning frameworks (e.g., PyTorch, TensorFlow). This allows existing ML vetting tools to be applied to the forensically recovered models, enabling comprehensive analysis for backdoors, adversarial vulnerabilities, or other anomalies.
  4. Preservation of Model Behavior: A key validation metric for rehosting is that the rehosted model must exhibit identical behavior to the deployed model. APE achieved this by demonstrating that the accuracy of the rehosted model exactly matched the accuracy of the deployed model on the same datasets (e.g., CIFAR-10, IMDb). This confirms that the rehosting process does not introduce artifacts or alter the model's functional characteristics.
  5. Framework and Domain Agnostic: APE was evaluated across different versions of PyTorch and TensorFlow, and various model types (e.g., MobileNetV2) and application domains, proving its generalizability. It successfully recovered and rehosted models from diverse application domains, indicating its broad applicability.
  6. Addressing Tensor Ambiguity: APE introduces a novel feedback-driven tensor recovery methodology to accurately identify and extract the correct model tensors amidst thousands of other tensors (e.g., optimizer states, layer activations) that may be present in memory, often with similar attributes.

Technical Deep Dive

▶ Watch: Step 2: Feedback-driven tensor recovery methodology (4:18)

AI Psychiatry (APE) tackles the complex challenge of deep learning forensics by addressing three fundamental questions: how to obtain the model structure, how to access its specific weights (tensors), and how to rehost it into a usable environment. The framework's methodology is broken down into three interconnected phases: Model Identification and Layer Recovery, Feedback-Driven Tensor Recovery, and Tensor Mapping and Model Rehosting.

1. Model Identification and Layer Recovery

The first step for APE, given CPU and GPU memory images from a deep learning system, is to reconstruct the model's architectural blueprint. This phase is guided by the key insight that deep learning models are intrinsically data structures that resemble directed graphs. Each layer or operation within a model can be thought of as a node in this graph, with connections defining the data flow.

APE's process for model identification involves:

  • Object Candidate Recovery: It first scans memory to identify data structures that exhibit characteristics consistent with directed graphs. This involves searching for pointers and metadata that indicate relationships between different memory regions, forming potential graph-like structures.
  • Layer Validation: Once candidate graph structures are identified, APE applies specific validation criteria to confirm they represent actual DL model layers. This includes:
  • Common Class Names: Checking if the identified nodes (potential layers) correspond to known class names used in popular deep learning frameworks (e.g., Conv2d, Linear, ReLU in PyTorch or TensorFlow).
  • CUDA API Adherence: For GPU-accelerated models, APE verifies that the identified data structures adhere to the CUDA API data structures required for low-level GPU inference. This ensures that the recovered components are genuinely part of an active DL inference pipeline, rather than arbitrary data.
  • Topology and Attribute Extraction: Upon successful identification, APE reconstructs the model's topology (the arrangement and connectivity of layers), recovers layer shapes (input/output dimensions), and identifies potential data pointers associated with these layers. This comprehensive process ultimately yields the full model structure and its core attributes, providing the skeleton onto which the weights will be attached.

2. Feedback-Driven Tensor Recovery

Once the model structure is established, the next challenge is to accurately identify and extract the correct model weights (tensors) from memory. This is a non-trivial task because memory can contain thousands of tensors, not all of which are part of the active model. Memory might hold tensors related to optimizers, intermediate layer activations, or even stale data, many of which can have identical attributes (e.g., shape, element count) to actual model weights, obfuscating the true tensors. Furthermore, tensors can be distributed across both CPU and GPU memory spaces, adding to the complexity.

APE addresses this through its innovative feedback-driven tensor recovery methodology, based on a crucial observation:

  • Node-Tensor Correspondence: For most common deep learning operations, a single model node (e.g., a convolutional layer) typically corresponds to exactly two tensors: an activation tensor (the layer's weights) and a bias tensor. This specific relationship provides a strong heuristic for filtering and validating recovered tensors.

The process unfolds as follows:

  • Attribute Matching: APE identifies pairs of tensors in memory that possess highly similar attributes, such as identical shapes, element counts, and reference counts, but crucially, different data pointers. These pairs are prime candidates for being the activation and bias of a single layer.
  • Discrepancy Detection: When such a pair is found, APE feeds them into its feedback mechanism. Knowing that a model node usually corresponds to only two distinct, valid tensors, if one of the tensors in the pair has an invalid or suspicious data pointer (e.g., pointing to a memory region not allocated for CPU or GPU, or marked for overwrite), it's flagged as "wrong." The talk gives an example where a tensor was identified with an invalid data pointer, indicating it was stale or waiting to be overwritten.
  • Iterative Validation: This criteria is applied iteratively across all tensors discovered in memory. By combining this feedback loop with the already recovered model structure and topology, APE can deduce which tensors genuinely belong to the model and which are extraneous. This systematic approach effectively prunes the vast number of potential tensors down to the precise set of weights needed for the model.

3. Tensor Mapping and Model Rehosting

The final phase involves transforming the recovered flat buffer tensors and model structure into a fully functional and vet-able model within a standard deep learning framework. The goal is to set up an environment (e.g., a Python script using PyTorch or TensorFlow) where the recovered model can be loaded and executed, just like a natively developed model.

The rehosting process proceeds as follows:

  • Intermediate Representation: The recovered flat buffer tensors, which are raw binary data, are first stored in an intermediate, framework-agnostic representation. NumPy arrays are commonly used for this purpose due to their versatility and compatibility with various scientific computing libraries.
  • Mapping to Model Nodes: Each NumPy array (representing a recovered tensor) is then mapped to its corresponding model node (operation) within the reconstructed model structure. This step links the raw weight data to its specific architectural location.
  • Framework-Specific API Lifting: Finally, APE leverages the framework-specific API (e.g., torch.nn.Parameter in PyTorch or tf.Variable in TensorFlow) to "lift" these NumPy arrays into proper, reusable model tensors within the chosen deep learning framework. This involves correctly instantiating the framework's tensor objects and assigning the recovered data to them.
  • Model Assembly: This iterative process, applied to each recovered tensor, culminates in the complete assembly of the deep learning model within the target framework. At this point, the model is fully functional and ready to be subjected to existing ML vetting tools, allowing "Detective Pikachu" to successfully identify any backdoors or anomalies.

Demo / Proof of Concept

▶ Watch: Step 3: Rehosting the recovered model for reuse (5:50)

While the presentation did not feature a live, interactive demonstration of the AI Psychiatry (APE) tool, the researchers provided comprehensive evaluation results that serve as a robust proof of concept for its capabilities and effectiveness. The evaluation methodology was meticulously designed to validate APE's core claims regarding model recovery, fidelity, and rehosting success.

The experimental setup involved:

  • Diverse Model Deployments: APE was tested against 30 distinct deep learning models, encompassing five different model types. These models were deployed across three different versions of both PyTorch and TensorFlow, demonstrating APE's versatility across popular frameworks and their evolving versions.
  • Standard Datasets: The models were trained on well-known benchmark datasets such as CIFAR-10 (for image classification) and IMDb (for sentiment analysis), ensuring that the evaluation was grounded in realistic and comparable scenarios.
  • Memory Input: In all evaluation scenarios, the sole input to APE was the CPU and GPU memory images of each deployed deep learning system, accurately simulating a forensic investigation where only runtime memory is available.

The key results highlighted in the presentation, with more detailed findings available in the accompanying paper, unequivocally support APE's efficacy:

  • Model Recovery Accuracy: APE achieved a remarkable 100% accuracy in recovering the model structure for all evaluated models. This perfect recovery rate is attributed to its graph-guided approach, which systematically identifies and reconstructs the model topology. The recovered models ranged in complexity, with some possessing upwards of 94 million weights, showcasing APE's ability to handle large-scale deep learning architectures.
  • Tensor Recovery Success: The system successfully recovered all associated GPU tensors, with counts ranging from 11 to 940 GPU tensors depending on the model type. This demonstrates the effectiveness of the feedback-driven tensor recovery in accurately isolating relevant weights from the noise of other memory artifacts.
  • Rehosting Fidelity - Accuracy: The most critical metric for successful rehosting is that the behavior of the rehosted model must be identical to the deployed model. APE achieved this by demonstrating that the accuracy of the rehosted model was exactly the same as the accuracy of the deployed model when tested on the same dataset. This outcome provides strong assurance that the rehosted model is a faithful forensic replica, suitable for reliable vetting.
  • Rehosting Fidelity - Structural Integrity: Further confirming rehosting success, the number of layers in the rehosted models consistently matched the deployed models. For instance, a MobileNetV2 model, whether deployed in PyTorch or TensorFlow, was correctly identified and rehosted with its characteristic four layer types. This structural consistency is crucial for ensuring that the entire model architecture has been accurately reconstructed.
  • Domain Agnostic Recovery: The ability to recover models from different application domains (e.g., image recognition, natural language processing) signifies APE's broad applicability, indicating that its underlying principles are not tied to specific types of DL tasks.

The comprehensive evaluation, covering various frameworks, model sizes, and application domains, serves as a compelling proof of concept that APE can reliably recover and rehost deep learning models from memory images, enabling subsequent forensic analysis.

Defensive Implications

▶ Watch: Evaluation: Model recovery accuracy and rehosting success (6:30)

The development of AI Psychiatry (APE) introduces profound defensive implications for organizations deploying and managing deep learning systems, particularly in sensitive or critical environments. It fundamentally shifts the paradigm of AI incident response from relying on vendor cooperation or source code availability to enabling independent, memory-based forensic analysis.

Here are the key defensive implications:

  1. Enhanced Incident Response for AI Systems: APE provides a crucial tool for post-incident forensic analysis of AI-driven systems. In scenarios where an AI component malfunctions, misbehaves, or is suspected of being compromised (e.g., a self-driving car accident, a biased decision by an AI-powered system), APE allows investigators to pinpoint the exact state of the deep learning model at the time of the incident. This capability is indispensable for root cause analysis, helping to determine if the issue stemmed from a design flaw, a runtime bug, or a malicious attack.
  1. Unlocking Proprietary and Encrypted Models: Many AI models, especially those embedded in commercial products or critical infrastructure, are proprietary and often encrypted or obfuscated to protect intellectual property. This makes traditional vetting and forensic analysis impossible. APE breaks this barrier by recovering model structure and weights directly from memory, enabling security teams to independently audit and investigate these black-box systems without requiring vendor source code or decryption keys. This is particularly vital for regulatory compliance and supply chain security in AI.
  1. Detecting Runtime Attacks and Online Learning Compromises: Systems that employ online learning or continuous model updates are particularly vulnerable to attacks that modify model weights at runtime, such as sophisticated data poisoning or adaptive adversarial attacks. Traditional vetting tools, which analyze static models, would miss these dynamic compromises. APE's ability to capture and analyze the model's state at the moment of compromise allows defenders to identify if malicious updates or runtime manipulations altered the model's behavior, providing insights into advanced persistent threats targeting AI.
  1. Validating Model Integrity in Production: Beyond incident response, APE can serve as a powerful tool for proactive security auditing of deployed AI models. By periodically capturing memory images and using APE to reconstruct and vet models, organizations can verify the integrity of their AI systems in production. This can help detect subtle backdoors that might have been introduced during development or supply chain, or identify deviations from expected model behavior that indicate compromise.
  1. Improving Accountability and Trust in AI: The ability to forensically examine AI models in a verifiable manner significantly enhances accountability and trust in AI systems. When an AI system makes a critical decision or causes an adverse event, APE provides a mechanism to investigate why it behaved that way, offering transparency that was previously unattainable. This is crucial for building public confidence in AI and for addressing legal and ethical concerns.
  1. Guiding Remediation and Hardening: By precisely identifying the nature of a model compromise (e.g., a specific backdoored neuron, a poisoned data segment affecting certain weights), APE provides actionable intelligence. This detailed understanding allows defenders to develop more targeted and effective remediation strategies, whether it involves patching models, improving training data hygiene, or hardening runtime environments against specific attack vectors.

In essence, AI Psychiatry equips defenders with a critical missing piece in the AI security toolkit, transforming opaque AI systems into analyzable entities for comprehensive forensic investigation and proactive security assurance.

Key Takeaways

  • Traditional deep learning model vetting tools are insufficient for real-world forensic investigations, as they assume full model access and cannot handle proprietary, encrypted, or dynamically updated models.
  • AI Psychiatry (APE) is a novel memory forensics framework that can recover the structure and specific weights (tensors) of deep learning models directly from CPU and GPU memory images.
  • APE employs a multi-stage methodology: graph-guided model identification, feedback-driven tensor recovery to filter relevant weights, and framework-agnostic tensor mapping for rehosting.
  • The framework achieves 100% accuracy in model recovery and ensures forensic integrity, with rehosted models exhibiting identical accuracy and structural properties to the deployed originals.
  • APE's capabilities are critical for incident response in AI-driven systems, enabling the detection of backdoors, data poisoning, and other runtime compromises in contexts like autonomous vehicles or embedded AI.
  • This research empowers independent security auditing and enhances accountability for AI systems, even when vendors provide black-box solutions, by providing a mechanism for post-compromise analysis without source code access.

About the Speaker(s)

The talk "AI Psychiatry: Forensic Investigation of Deep Learning Networks in Memory Images" was presented by David Oygenblik, a key contributor to this research. He, along with Carter Yagemann, Joseph Zhang, Arianna Mastali, Jeman Park, and Brendan Saltaformaggio, represent a collaborative effort in advancing the field of AI security and forensics. While specific affiliations for all speakers were not detailed in the presentation transcript, the mention of "GT" as a collaborator strongly suggests their affiliation with Georgia Tech (Georgia Institute of Technology), a prominent research institution. Brendan Saltaformaggio is known to be a faculty member at Georgia Tech, specializing in digital forensics and system security. Their collective work focuses on critical areas of computer security, particularly in developing innovative techniques for analyzing and securing complex software systems and emerging technologies like deep learning.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research introduces APE, a novel memory forensics framework for deep learning models, directly addressing a critical gap in AI incident response for proprietary, encrypted, or dynamically updated systems. It enables the accurate recovery and rehosting of model structure and weights from memory images, providing an indispensable tool for post-incident analysis and independent security auditing. This is a foundational defensive innovation for securing AI systems in production.

Heather Calloway (CISO) — MUST SEE

This talk introduces AI Psychiatry (APE), a critical memory forensics framework for deep learning models. It enables independent post-incident analysis of proprietary or dynamically updated AI systems, directly addressing a significant gap in AI incident response and enhancing accountability in safety-critical applications.

→ Top-rated talks at 33rd USENIX Security Symposium

All talks from 33rd USENIX Security Symposium