ASGARD: Protecting On-Device Deep Neural Networks with Virtualization-Based Trusted Execution Environments
Myungsuk Moon (Jon University)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · DNN Attack Surfaces
Overview
The proliferation of on-device Artificial Intelligence (AI) services offers significant advantages over traditional cloud-based AI, primarily by keeping sensitive user data local and avoiding network latency. However, these on-device Deep Neural Network (DNN) models represent highly valuable intellectual property, often trained with substantial computational resources. The core challenge addressed by this talk, "ASGARD: Protecting On-Device Deep Neural Networks with Virtualization-Based Trusted Execution Environments," is safeguarding these proprietary models from malicious device owners who might attempt to extract, tamper with, or reverse-engineer them.
Key moments
- 0:17 Problem: Protecting on-device DNNs; prior work limitations
- 2:20 ASGARD's virtualization-based TEE approach and benefits
- 3:59 Challenges: runtime overheads and TCB increase
- 4:45 Reducing TCB with IOMMU driver splitting
- 5:55 Minimizing overhead for accelerator reassignment
- 7:45 Optimizing interrupt delivery for DNN inference
ASGARD: Protecting On-Device Deep Neural Networks with Virtualization-Based Trusted Execution Environments
Speakers: Myungsuk Moon (Jon University)
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=6W0z1Gp7VMY
Overview
The proliferation of on-device Artificial Intelligence (AI) services offers significant advantages over traditional cloud-based AI, primarily by keeping sensitive user data local and avoiding network latency. However, these on-device Deep Neural Network (DNN) models represent highly valuable intellectual property, often trained with substantial computational resources. The core challenge addressed by this talk, "ASGARD: Protecting On-Device Deep Neural Networks with Virtualization-Based Trusted Execution Environments," is safeguarding these proprietary models from malicious device owners who might attempt to extract, tamper with, or reverse-engineer them.
Presented by Myungsuk Moon from Jon University, ASGARD introduces a novel approach to secure on-device Convolutional Neural Networks (CNNs) using virtualization-based Trusted Execution Environments (TEEs). Unlike prior work that often struggled with memory limitations, driver compatibility, or significant runtime overheads, ASGARD leverages an EL2 hypervisor to create isolated virtual machines (VMs) where DNN inference can occur securely. This design allows for full model protection, dynamic memory allocation, and superior driver compatibility, marking a significant advancement in the field of on-device AI security.
The significance of ASGARD lies in its ability to deliver robust security for on-device AI models without incurring prohibitive performance costs or requiring modifications to the most privileged system components. By intelligently optimizing interrupt delivery and accelerator reassignment within a virtualized environment, ASGARD demonstrates that virtualization overheads can be effectively contained. This work is crucial for fostering trust in on-device AI applications, enabling developers to deploy their valuable models with greater confidence, and pushing the boundaries of what is possible in secure edge AI.
Background
▶ Watch: Problem: Protecting on-device DNNs; prior work limitations (0:17)
The shift towards on-device AI has been driven by compelling factors, including enhanced user privacy (by not sending data to the cloud), reduced latency, and offline functionality. However, this paradigm introduces a critical security vulnerability: once a highly valuable and resource-intensive DNN model is deployed on a user's device, it becomes susceptible to attacks from a malicious device owner. Protecting this intellectual property from extraction, analysis, or modification is paramount.
Prior research has largely focused on utilizing ARM TrustZone as a foundation for securing on-device models. TrustZone partitions the system into a Secure World and a Normal World, with an EL3 secure monitor enforcing isolation. This prior work can broadly be categorized into two lines of approach:
- Assigning the accelerator to the Secure World: This approach aimed to perform hardware-accelerated inference within the isolated Secure World. However, it faced several significant limitations:
- Static Memory Partitioning: During device boot time, Secure World memory is statically partitioned, often resulting in a small memory footprint insufficient for larger DNN models.
- Vendor-Specialized OS: The Secure World typically runs a vendor-specialized Trusted OS, making it difficult to reuse existing, feature-rich accelerator drivers from the host Linux (Normal World). Efforts to write minimal, fully custom drivers in the Trusted OS (as seen in 2022 research) suffered from poor driver compatibility.
- EL3 Monitor Modification: Other attempts required adding components to the EL3 secure monitor, which is the most privileged and critical component in the system, increasing the Trusted Computing Base (TCB) and introducing potential security risks.
- Leaving the accelerator in the Normal World: Recognizing the limitations of the first approach, some researchers opted to keep the accelerator in the less privileged Normal World. This also presented its own set of challenges:
- Incomplete Protection: Early work offloaded less sensitive operators to the Normal World, meaning the entire DNN model was not fully protected.
- Runtime Overhead with Obfuscation: More recent work (2023) attempted to protect model weights by obfuscating them, only offloading linear operators to the Normal World. While this offered better protection, it introduced substantial runtime overheads due to the necessity of de-obfuscating and re-obfuscating all intermediate tensors passed between the two worlds.
These inherent limitations—ranging from restricted memory and poor driver compatibility to TCB expansion and significant runtime overheads—highlighted the need for a more robust and efficient solution for protecting on-device DNN models. This gap motivated the development of ASGARD, which addresses these challenges by leveraging virtualization-based TEEs.
Key Findings
▶ Watch: Challenges: runtime overheads and TCB increase (3:59)
ASGARD presents a groundbreaking approach to protecting on-device Convolutional Neural Networks (CNNs) by designing the first system that utilizes virtualization-based Trusted Execution Environments (TEEs) within legacy System-on-Chips (SOCs). This innovative architecture fundamentally shifts the paradigm for securing valuable DNN intellectual property on user devices.
One of the primary contributions of ASGARD is its ability to provide full protection for the entire DNN model. Unlike prior TrustZone-based methods that were constrained by static and often small secure memory partitions, ASGARD employs virtual machines (VMs) for its TEEs, allowing for the dynamic allocation of memory. This flexibility ensures that even large and complex DNN models can reside entirely within the secure VM, thereby preventing partial model exposure or the need for problematic operator offloading.
The system demonstrates significant performance advantages over existing solutions. In evaluations using a MobileNet V1 model, ASGARD was found to be approximately four times faster than prior work that relied on obfuscating model weights and intermediate tensors. This substantial speedup is attributed to ASGARD's design, which eliminates the costly de-obfuscation and re-obfuscation steps, even when accounting for the overheads of dynamically reassigning the Media Processing Unit (MPU).
Crucially, ASGARD effectively addresses and contains the inherent virtualization overheads through a suite of both system and application-level optimizations. For instance, the novel exec-coalescing execution order for CPU-phobic operators was shown to reduce the number of VM exits by roughly half. This optimization translated into tangible performance gains, decreasing runtime overheads for SSD models from 2-3% down to 0 to -1% (compared to running inference in the Rich Execution Environment) and for Light Transfer models from 15-33% down to 3-6%. These results unequivocally demonstrate that virtualization can be a viable and performant security primitive for on-device AI.
Furthermore, ASGARD offers superior driver compatibility by allowing the use of commodity operating systems like Linux within the VMs, eliminating the need for custom, minimal drivers. It also avoids modifying the EL3 security monitor, which is the most privileged component in the system, thereby enhancing system stability and reducing the attack surface of the Trusted Computing Base (TCB). These key findings establish ASGARD as a robust, efficient, and practical solution for securing on-device DNN models.
Technical Deep Dive
▶ Watch: Reducing TCB with IOMMU driver splitting (4:45)
ASGARD's architecture centers around a virtualization-based TEE, where the secure environment is a virtual machine running at EL0 and EL1, isolated by an EL2 hypervisor. The core design principle involves directly assigning the Media Processing Unit (MPU)—the hardware accelerator—to the VMs and supporting its dynamic reassignment among multiple virtual machines. This approach inherently offers good driver compatibility, as commodity operating systems like Linux can run within the VMs, and avoids modifying the critical EL3 security monitor. However, this design introduces specific challenges related to runtime and TCB overheads, which ASGARD addresses through several innovative optimizations.
Mitigating TCB and IOMMU Overheads
A fundamental challenge in securely assigning an accelerator to a VM is ensuring that the accelerator, capable of Direct Memory Access (DMA), only accesses the VM's designated memory and no other parts of the system. This memory access policy is enforced by an IOMMU (Input/Output Memory Management Unit), which operates with its own page tables controlled by the EL2 hypervisor. A straightforward implementation would involve embedding a full IOMMU driver within the EL2 hypervisor, significantly increasing the TCB.
To address this, ASGARD employs an IOMMU driver split. The EL2 hypervisor delegates the less security-critical resource management portion of the IOMMU driver to an untrusted component, while retaining only the essential resource isolation part. This minimizes the trusted code running at EL2, thereby reducing the TCB without compromising memory access control.
Minimizing Accelerator Reassignment Overheads
Dynamically reassigning the MPU from one VM to another is critical for multi-tenant or flexible on-device AI scenarios. A naive approach would involve deleting all existing IOMMU page table entries for the current VM (e.g., TEE #1) and then creating entirely new entries for the target VM (e.g., TEE #2). For VMs that can be hundreds of megabytes in size, this process would incur a significant runtime overhead.
ASGARD optimizes this by having the hypervisor maintain all page table entries that were created during the initial VM initialization. During an MPU reassignment, instead of recreating entries, the hypervisor recycles and reuses these existing page table entries. This intelligent reuse strategy dramatically minimizes the runtime overhead associated with reassigning the accelerator, making dynamic MPU management practical.
Secure Accelerator Reassignment
Beyond minimizing overhead, secure reassignment also necessitates ensuring that no sensitive data from the previous VM is leaked to the new VM. Accelerators can retain leftover data in their device registers even after being deallocated.
ASGARD implements a robust three-step secure reassignment process:
- Remove Access: The hypervisor first revokes all access permissions for the previous VM (TEE #1) to the accelerator.
- Full Reset: The hypervisor then performs a full hardware reset of the accelerator. This critical step ensures that all leftover data in the device registers is completely wiped out.
- Grant Access: Finally, the hypervisor grants access to the accelerator to the new target VM (TEE #2). This sequence guarantees a clean slate for each VM using the accelerator.
Optimizing Interrupt Delivery with Exec-Coalescing
Delivering interrupts from the accelerator to the VM is another source of significant runtime overhead. When an accelerator completes an operation, it generates a physical interrupt to the host kernel (which is outside the TCB). The host kernel then injects a virtual interrupt into the VM. This process involves multiple VM exits and entries because the host kernel must receive an acknowledgment from the VM and subsequently unmask the physical interrupt. During DNN inference, where many operations might trigger interrupts, this overhead can accumulate quickly.
DNN models typically consist of both MPU-supported operators (executed by the accelerator) and CPU-phobic operators (not supported by the MPU, requiring CPU execution within the VM). In a default execution order, after the MPU finishes an operator, it would deliver an interrupt to the VM, prompting the VM to execute the next CPU-phobic operator. If there are multiple CPU-phobic operators that are not data-dependent on each other, each would trigger a separate interrupt. For example, if CPU operators 3, 6, and 9 are executed separately, this would result in three interrupts.
ASGARD introduces an exec-coalescing execution order to optimize this. This technique identifies and groups all CPU-phobic operators that are not data-dependent on each other. Instead of executing them sequentially with intermediate interrupts, the hypervisor (or a component orchestrated by it) can arrange for these grouped operators to be called together, resulting in only a single interrupt being sent to the VM. In the example above, if operators 3, 6, and 9 are data-independent, they can be coalesced, leading to only one interrupt instead of three. This significantly reduces the number of costly VM exits and entries, thereby boosting inference performance.
These technical innovations collectively enable ASGARD to provide robust security for on-device DNNs within a virtualization-based TEE, effectively managing and mitigating the associated overheads.
Demo / Proof of Concept
▶ Watch: Minimizing overhead for accelerator reassignment (5:55)
The ASGARD prototype was implemented and evaluated on an ARM 8.2 legacy SOC featuring an integrated Media Processing Unit (MPU). For running the protected virtual machines (VMs) that serve as the Trusted Execution Environments (TEEs), the project leveraged Android 13's protected KVM (Kernel-based Virtual Machine). The virtual machine monitor (VMM) component was built using Google's Cross VM. This specific choice of hardware and software components demonstrates ASGARD's practicality and compatibility with existing Android ecosystems, indicating that the design can work with a wide range of current Android devices, albeit with some minor device-specific adjustments (e.g., for MPU resets).
The evaluation focused on two primary metrics: DNN inference latency and the effectiveness of the exec-coalescing optimization.
- DNN Inference Latency:
- ASGARD's performance was first measured using a MobileNet V1 model. Compared to prior work that involved offloading linear operators to the Normal World and incurring overheads from de-obfuscating and re-obfuscating intermediate tensors, ASGARD demonstrated a significant advantage.
- Even with the overhead of reassigning the MPU before running protected inference in the VM, ASGARD's approach was approximately four times faster than the obfuscation-based prior work. This highlights the substantial performance penalty associated with obfuscation techniques, which ASGARD entirely avoids.
- Exec-Coalescing Execution Order Evaluation:
- The exec-coalescing optimization was evaluated on two types of models known for their parallel architectures:
- SSD models (Single Shot Detector): These models feature multiple parallel feature extractors. In the prototype, reshape and transpose operators were identified as CPU-phobic. Before applying exec-coalescing, the runtime overheads for SSD models were around 2-3% compared to running inference in the Rich Execution Environment (REE). After applying the optimization, these overheads were reduced to 0 to -1%, effectively eliminating or even slightly improving performance relative to the REE due to reduced context switching.
- Light Transfer models: These models also exhibit a parallel DNN architecture, with reshape operators identified as CPU-phobic. Initially, these models showed runtime overheads of 15-33% compared to the REE. With exec-coalescing, the overheads were significantly reduced to 3-6%.
- Across these models, the exec-coalescing strategy was successful in removing approximately half of the VM exits that would otherwise occur, directly translating into the observed performance improvements.
The talk briefly mentioned that more detailed evaluation results, including TCB size, interrupt delivery latency, MPU reassignment latency, and memory usage, are available in the full paper. The demonstrated proof of concept validates ASGARD's design principles, proving that virtualization-based TEEs can provide robust security for on-device DNNs with manageable and often superior performance compared to alternative methods.
Defensive Implications
▶ Watch: Optimizing interrupt delivery for DNN inference (7:45)
ASGARD's virtualization-based TEE approach offers profound defensive implications for various stakeholders in the on-device AI ecosystem, fundamentally enhancing the security posture of deployed DNN models.
For AI Model Developers and Providers: ASGARD provides a robust and practical mechanism to protect their invaluable intellectual property—the trained DNN models—from malicious actors on user devices. This mitigates the risk of model exfiltration, reverse engineering, or tampering, which could lead to competitive disadvantages, loss of revenue, or even the creation of adversarial models. By deploying models within ASGARD's protected VMs, developers can ensure that their proprietary algorithms and trained weights remain confidential and integral, fostering greater trust in the security of their on-device AI services. This also enables the deployment of more sophisticated and sensitive models to edge devices that might otherwise be deemed too risky.
For Device Manufacturers: The ASGARD framework offers a blueprint for integrating advanced secure on-device AI capabilities into their products without requiring radical hardware redesigns or modifications to the most privileged software components. By leveraging existing virtualization hardware features (like ARM's EL2 hypervisor and IOMMU) and compatible software stacks (such as Android's protected KVM), manufacturers can enhance the security offerings of their devices. This can differentiate their products in the market, providing a secure foundation for future AI-driven applications and services. The ability to support commodity operating systems within the TEE also simplifies driver management and reduces development costs.
For End-users and Application Developers: While not directly interacting with ASGARD, end-users benefit from increased confidence in the privacy and integrity of on-device AI applications. Knowing that sensitive AI models are protected from device-level attacks can alleviate concerns about data security and model manipulation. For application developers building on-device AI features, ASGARD simplifies the security landscape, allowing them to focus on application logic rather than complex low-level security primitives, while still benefiting from strong model protection.
For System Architects and Security Researchers: ASGARD highlights the critical role of well-designed hypervisors and hardware-assisted virtualization features (like IOMMU) in building secure edge computing platforms. It demonstrates how splitting driver functionalities and optimizing interrupt handling can significantly reduce TCB and mitigate performance overheads in complex security architectures. This work encourages further research into applying virtualization-based TEEs for protecting other critical workloads on edge devices, moving beyond the limitations of traditional TrustZone implementations. It underscores that robust security for on-device AI is achievable without sacrificing performance or system compatibility, paving the way for more widespread and secure adoption of edge AI.
Key Takeaways
- Virtualization-Based TEEs for DNN Protection: ASGARD introduces the first system to effectively protect on-device Convolutional Neural Networks (CNNs) using virtualization-based Trusted Execution Environments (TEEs) in legacy SOCs, offering a robust alternative to previous TrustZone limitations.
- Full Model Protection and Flexibility: Unlike prior static memory partitions, ASGARD's VM-based TEEs allow for dynamic memory allocation, enabling full protection for entire DNN models regardless of their size and avoiding the need for partial operator offloading or obfuscation.
- Superior Performance over Obfuscation: ASGARD demonstrates a significant performance advantage, being approximately four times faster than previous obfuscation-based methods (e.g., with MobileNet V1), by eliminating the costly de-obfuscation and re-obfuscation of intermediate tensors.
- Optimizations Mitigate Virtualization Overheads: Key system and application-level optimizations, such as the IOMMU driver split, page table reuse for MPU reassignment, secure MPU resetting, and especially the exec-coalescing execution order, successfully contain and reduce virtualization overheads, making the approach practical and efficient.
- Reduced VM Exits and Improved Latency: The exec-coalescing technique significantly reduces the number of VM exits (by roughly half), translating into minimal runtime overheads (0 to -1% for SSD models, 3-6% for Light Transfer models) compared to the Rich Execution Environment.
- High Compatibility and Low TCB Impact: ASGARD achieves good driver compatibility by using commodity OSs like Linux in VMs and avoids modifying the highly privileged EL3 security monitor, thereby minimizing the Trusted Computing Base (TCB) and enhancing system stability.
About the Speaker(s)
The work on ASGARD was presented by Myungsuk Moon from Jon University. This research is a collaborative effort, developed in conjunction with colleagues Mini Jingu and their adviser Doyong, also from Jon University. Their collective expertise in system security and on-device AI has culminated in this innovative virtualization-based approach to protecting deep neural networks.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
ASGARD is a technically credible systems paper solving a real problem — protecting proprietary DNN model weights from malicious device owners — using a non-obvious architectural choice: EL2 hypervisor-based TEEs instead of TrustZone's Secure World. The prior-work analysis is honest, the threat model is well-scoped, and the engineering contributions (IOMMU driver split, page table recycling, exec-coalescing) are concrete and defensible. The 4x speedup over obfuscation-based approaches and the VM exit reduction numbers give this teeth.
Heather Calloway (CISO) — WEAK
Technically credible systems research on protecting on-device DNN models using virtualization-based TEEs. The engineering is real and the performance benchmarks are meaningful — but this is fundamentally IP protection research aimed at model vendors, and it never crosses into governance, organizational risk, or anything a CISO or security leader would act on.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025