Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN Executables

Yanzuo Chen

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · DNN Attack Surfaces

Overview

In a groundbreaking presentation at the NDSS Symposium, Yanzuo Chen unveiled critical vulnerabilities within Deep Neural Network (DNN) executables, demonstrating a novel and highly effective bit-flip attack. The talk, titled "Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN Executables," revealed that these optimized, compiled versions of DNN models harbor pervasive attack surfaces that can be exploited with remarkable efficiency. This research introduces a gray-box threat model, significantly more restricted than prior white-box approaches, yet achieves devastating results: the complete depletion of a model's intelligence with an average of just 1.4 bit flips.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction: Bit-flip attacks on DNN executables
  2. 2:00 DNN executables and row hammer bit-flip attacks
  3. 4:00 Intelligence depletion: Gray-box attacker threat model
  4. 6:00 Baseline random bit flip attack: Only 2% success
  5. 8:00 Discovery: Transferable vulnerable bits across models
  6. 9:00 Identifying 'super bits' using random noise datasets

Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN Executables

Speakers: Yanzuo Chen

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=Dg4jzCVpu5Y

Overview

In a groundbreaking presentation at the NDSS Symposium, Yanzuo Chen unveiled critical vulnerabilities within Deep Neural Network (DNN) executables, demonstrating a novel and highly effective bit-flip attack. The talk, titled "Compiled Models, Built-In Exploits: Uncovering Pervasive Bit-Flip Attack Surfaces in DNN Executables," revealed that these optimized, compiled versions of DNN models harbor pervasive attack surfaces that can be exploited with remarkable efficiency. This research introduces a gray-box threat model, significantly more restricted than prior white-box approaches, yet achieves devastating results: the complete depletion of a model's intelligence with an average of just 1.4 bit flips.

The core of the discovery lies in identifying "super bits" within the compiled code of DNN executables. Unlike previous attacks that targeted model weights, this work focuses on manipulating the compiled operations themselves. By leveraging techniques like Rowhammer, attackers can induce these bit flips, turning sophisticated classification models into random guessers. This research is crucial because DNN executables are increasingly deployed for performance reasons, making these newly identified vulnerabilities a significant concern for the security and reliability of AI systems in real-world applications.

Chen's work highlights a paradigm shift in understanding DNN security. As deep learning models transition from research environments to optimized, compiled deployments, new attack vectors emerge that challenge conventional defensive strategies. The findings underscore an urgent need for increased security scrutiny of deep learning compilers and the resulting executables, urging the community to explore novel defense mechanisms against these potent, low-resource attacks.

Background

▶ Watch: Introduction: Bit-flip attacks on DNN executables (0:00)

The landscape of Deep Neural Networks (DNNs) has seen a significant evolution, moving beyond raw model files to highly optimized, compiled executables. These DNN executables are the result of processing DNN models through specialized deep learning compilers such such as Apache TVM and Meta's Glow. The primary motivation for this compilation process is performance enhancement. By optimizing the computational graph and tailoring the code for specific target hardware platforms, DNN executables can achieve superior efficiency and extract the maximum potential from the underlying hardware, making them ideal for deployment in resource-constrained environments or high-throughput applications.

Alongside the rise of optimized DNN deployments, bit-flip attacks (BFAs) have emerged as a potent threat vector. BFAs are a class of hardware vulnerabilities that involve flipping individual data bits in memory, typically in DRAM. The most prominent example is Rowhammer, a hardware vulnerability that can be triggered purely through software. Rowhammer exploits physical properties of DRAM cells: repeatedly accessing a row of memory (an "aggressor" row) at a high frequency can induce electrical interference, causing bits in adjacent, infrequently accessed "victim" rows to spontaneously flip. This phenomenon has been demonstrated across multiple generations of memory modules, highlighting its pervasive nature as a hardware-level vulnerability with software-triggerable exploits.

Prior research has indeed explored the impact of bit flips on DNN models. These earlier works demonstrated that manipulating model weights through bit flips can significantly alter a model's behavior. However, these attacks typically operated under a white-box threat model, where the attacker had full knowledge of the model's internal parameters, including its weights. This knowledge allowed attackers to calculate gradients and precisely identify which bits to flip for optimal attack performance. For models using full-precision floating-point weights, only a few flips were often sufficient. Quantized models, utilizing integer weights, generally required a higher number of flips to achieve comparable attack efficacy. The key distinction of the work presented by Chen is its focus on DNN executables rather than raw models, and its adoption of a far more restrictive gray-box threat model, which significantly changes the attack surface and methodology.

Key Findings

▶ Watch: Intelligence depletion: Gray-box attacker threat model (4:00)

The research presented by Yanzuo Chen reveals several critical findings that redefine the understanding of bit-flip vulnerabilities in deep learning systems:

  1. Pervasive Attack Surfaces in DNN Executables: The study conclusively demonstrates that DNN executables contain pervasive attack surfaces susceptible to bit-flip attacks. These vulnerabilities are not isolated to specific models or compilers but exist across a wide range of architectures, datasets, and compilation tools. This implies a systemic rather than incidental vulnerability.
  1. Transferable Vulnerable Bits (TVBs): A groundbreaking discovery is the existence of "transferable vulnerable bits." Experiments showed that two DNN executables compiled from the same model structure but trained on different training datasets share approximately 45% of their vulnerable bits. This significant overlap indicates that an attacker does not need precise knowledge of the victim's training data to identify effective flip locations.
  1. "Super Bits" for High Attack Success Rate: Building upon the concept of transferable vulnerable bits, the researchers developed a methodology to identify "super bits." These are bits that are vulnerable across multiple locally generated executables (trained on random noise data). By identifying the intersection of vulnerable bits from 10 such "fake" datasets, the attacker achieved an impressive 70% Attack Success Rate (ASR). This means there is a 70% confidence that a selected bit flip will successfully deplete the intelligence of a remote victim model. This is a dramatic improvement over the mere 2% ASR observed with random bit flips.
  1. Extremely Low Flip Count for Intelligence Depletion: Perhaps the most striking finding is the efficiency of the attack. On average, only 1.4 bit flips were required to completely deplete the intelligence of the victim model, transforming it into a random guesser. This is significantly fewer than the approximately 12 flips required by state-of-the-art attacks like Deep Hammer, which target model weights and operate under a white-box threat model.
  1. Gray-Box Attacker Model: The attack operates under a significantly more restricted gray-box threat model. The attacker does not have access to the victim model's weights or training data, only its structure and the knowledge of how it was compiled. This constraint makes the attack far more realistic and challenging to defend against, as it bypasses the need for confidential model parameters.
  1. No Crashes, Full Intelligence Depletion: The real-world experiments demonstrated that the super-bit-induced flips consistently led to models becoming random guessers without causing system crashes. This ensures the attack's stealth and effectiveness, as the compromised model continues to operate, albeit incorrectly.

These findings collectively highlight a novel and highly potent attack vector against DNN executables, underscoring the urgent need for new security paradigms in the deployment of AI models.

Technical Deep Dive

▶ Watch: Baseline random bit flip attack: Only 2% success (6:00)

The technical ingenuity of this research lies in its shift from attacking model weights to attacking the compiled code of DNN executables. This approach is predicated on the understanding that DNN executables are essentially compiled versions of DNN operators, meaning that manipulating bits within this compiled code can directly alter the model's operational logic.

The threat model is central to this work. The attacker is assumed to be in a gray-box scenario, possessing knowledge of the model's structure (e.g., its architecture, layers, and types of operations) and the deep learning compiler used (e.g., TVM or Glow). Crucially, the attacker does not have access to the victim model's specific weights or training data, which are often confidential. This constraint precludes traditional gradient-based search methods for identifying vulnerable bits. The attacker's objective is clear: intelligence depletion, aiming to reduce a classification model's accuracy to that of a random guesser (e.g., 1/N for N classes).

The attack flow begins with a local profiling phase. The attacker first constructs a local executable based on the known model structure and compilation method. Using this local executable, the attacker identifies potential bit flip locations. Once these "vulnerable bits" are identified, the attacker launches the real attack on the remote victim model by inducing bit flips at the determined memory offsets.

To establish a baseline, the researchers first attempted a random attack. Knowing the offset of the victim executable's text section (where the compiled code resides), an attacker could randomly select bits within this section to flip. This naive approach yielded a dismal 2% Attack Success Rate (ASR), with the vast majority of flips either causing system crashes or having no discernible effect on model performance. This underscored the critical need for an intelligent method to identify vulnerable bits rather than relying on brute force.

The breakthrough came with the discovery of transferable vulnerable bits. The researchers hypothesized that despite differences in training data, the underlying compiled operators for a given model structure might share common vulnerabilities. Experiments confirmed this, showing that executables compiled from the same model structure but trained on different datasets shared approximately 45% of their vulnerable bits. This indicated that some bit flip locations were inherently tied to the model's architecture and compiled logic, rather than its specific learned parameters.

This discovery led to the development of the "super bits" methodology. To overcome the lack of access to the victim's training data, the attacker generates multiple local executables. For these local executables, the models are trained using random noise data sets (referred to as "fake data sets"). This approach is inspired by machine learning studies suggesting that random data can help regulate model weights to resemble those of well-trained models during phases like pre-training, while also minimizing bias. By training multiple instances of the same model structure on different random noise seeds, the attacker effectively creates a diverse set of local executables.

For each of these local executables, the attacker identifies its set of vulnerable bits. The "super bits" are then defined as the intersection of these vulnerable bit sets across all locally generated executables. The rationale is that bits vulnerable across multiple diverse local instances are highly likely to be transferable and effective against a remote victim, regardless of its specific training data. Using 10 fake data sets in this manner, the attack achieved a remarkable 70% ASR, a significant improvement over the random baseline. This high confidence in identifying vulnerable bits, coupled with the gray-box threat model, makes the attack highly practical and potent.

The underlying mechanism is that these bit flips corrupt the instructions or data within the compiled DNN operators. For example, a flip might alter a jump instruction, change a constant used in a multiplication, or corrupt a memory address calculation, leading to incorrect computations that cascade through the network. This corruption of the program logic fundamentally breaks the model's ability to process inputs correctly, reducing its output to effectively random guesses.

Demo / Proof of Concept

▶ Watch: Discovery: Transferable vulnerable bits across models (8:00)

The efficacy of the "super bits" methodology was rigorously validated through real-world experiments conducted on actual hardware. The researchers utilized Blacksmith, a state-of-the-art Rowhammer toolkit designed for DDR4 memory, to induce the bit flips. This choice of toolkit underscores the practicality of the attack, as Rowhammer is a well-documented and exploitable hardware vulnerability.

The experiments involved applying the identified "super bits" to a wide range of victim DNN executables. These victims represented different model architectures and were compiled using standard deep learning compilers. The results were consistently devastating: in all cases, the targeted bit flips successfully depleted the intelligence of the victim models. This was quantified by observing the accuracy change, which consistently showed that the models' performance dropped to that of a random guesser (e.g., an image classification model with 10 classes would exhibit an accuracy of approximately 10%). This means the models were effectively rendered useless, unable to distinguish between inputs and, as Chen vividly put it, "equipped with the eyes to see ninjas out of panda photos."

A critical aspect of the demonstration was the stability of the attack. Despite corrupting the compiled code, the bit flips did not cause any crashes during the exploit. This is a significant factor for real-world attacks, as system crashes would immediately alert defenders to an issue. The ability to silently degrade model performance without causing system instability makes the attack particularly insidious.

The most compelling metric from the proof-of-concept was the number of flips required. On average, the attack needed only 1.4 bit flips to completely deplete the intelligence of a victim model. This figure is extraordinarily low and highlights the extreme fragility of DNN executables to targeted bit manipulation. For context, a comparison was drawn with Deep Hammer, a prominent state-of-the-art attack targeting model weights, which typically requires around 12 flips to achieve similar attack performance. Furthermore, Deep Hammer operates under a white-box threat model, demanding full knowledge of the model weights, whereas Chen's attack achieves superior efficiency with a significantly more restricted gray-box attacker. This stark contrast underscores the heightened risk posed by bit flips in the compiled code of DNN executables.

During the Q&A, the speaker also touched upon the potential for targeted attacks. While the primary demonstration focused on intelligence depletion (making the model a random guesser), Chen noted that different "degrees of aggressiveness" exist among vulnerable bits. This suggests that with further refinement, it might be possible to control the extent of degradation. Regarding class-specific targeted attacks (e.g., misclassifying only pandas), the speaker indicated that while some bit flips tend to "pin" the model towards a specific class (achieving an N-to-1 targeted attack), a more precise one-to-one targeted attack might still be better suited for traditional weight-based methods.

Defensive Implications

▶ Watch: Identifying 'super bits' using random noise datasets (9:00)

The findings presented in "Compiled Models, Built-In Exploits" carry profound defensive implications for the security of Deep Neural Network deployments. The revelation of pervasive and highly efficient bit-flip attack surfaces in DNN executables necessitates a re-evaluation of existing security paradigms, which have largely focused on model-level vulnerabilities (e.g., adversarial examples targeting weights).

First and foremost, this research calls for increased security scrutiny on deep learning compilers and their generated executables. Just as traditional compilers are subject to rigorous security audits to prevent vulnerabilities in compiled binaries, deep learning compilers like TVM and Glow must be examined for how they optimize and transform DNN models into code. Potential hardening mechanisms could include:

  • Code integrity checks: Implementing robust runtime integrity checks for the compiled DNN code in memory could detect unauthorized modifications caused by bit flips.
  • Redundant computation: Introducing redundancy in critical compiled operations, where the output of computations is cross-checked, could help detect and potentially correct errors induced by bit flips.
  • Compiler-level obfuscation/randomization: Making the compiled code less predictable by introducing randomization in instruction layout or memory allocation could make it harder for attackers to identify and target "super bits" across different deployments.

Secondly, the reliance on hardware-level vulnerabilities like Rowhammer means that hardware memory protections become paramount. While Rowhammer is a known issue, its implications for compiled DNN code highlight the need for:

  • Rowhammer mitigation technologies: Implementing and enabling hardware-level mitigations (e.g., Target Row Refresh - TRR) in DRAM modules and memory controllers is crucial.
  • Memory isolation and protection: Strong memory isolation mechanisms, potentially leveraging hardware enclaves or trusted execution environments, could help prevent an attacker's process from inducing bit flips in the memory regions allocated to the DNN executable.
  • Error Correcting Code (ECC) memory: While ECC memory can correct single-bit errors, the effectiveness against multi-bit Rowhammer flips might vary. However, it provides a baseline layer of defense against accidental or malicious single-bit corruptions.

Given the gray-box threat model, traditional defenses like adversarial training (which hardens models against input perturbations) are unlikely to be effective against attacks targeting the compiled code. Defenders must shift their focus to the runtime environment and the integrity of the deployed executable. This includes:

  • Runtime monitoring: Detecting unusual performance degradation or output anomalies that could signal intelligence depletion without system crashes.
  • Secure deployment practices: Ensuring that DNN executables are deployed in environments with minimal exposure to Rowhammer-inducing processes, or on hardware explicitly hardened against such attacks.

Finally, the speaker explicitly mentioned that defenses would be a subject of future work, indicating that the community is only beginning to grapple with these compiled-code vulnerabilities. This underscores the need for further research into practical, low-overhead defense mechanisms that can protect DNN executables from highly efficient bit-flip attacks, especially considering the extremely low number of flips required for compromise.

Key Takeaways

  • New Attack Surface: Deep Neural Network (DNN) executables, optimized compiled versions of DNN models, present a pervasive and critical attack surface for bit-flip attacks, distinct from traditional model weight-based vulnerabilities.
  • Gray-Box Threat Model: Attackers can effectively compromise DNN executables with a highly restricted gray-box threat model, requiring only knowledge of the model structure and compilation method, not confidential weights or training data.
  • "Super Bits" Enable Efficiency: The concept of "super bits"—vulnerable bit locations transferable across executables compiled from the same model structure—allows attackers to achieve a 70% Attack Success Rate (ASR) by profiling local executables trained on random noise data.
  • Extremely Low Flip Count: The attack is remarkably efficient, requiring an average of only 1.4 bit flips to completely deplete a victim model's intelligence, turning it into a random guesser without causing crashes.
  • Rowhammer Exploitation: The attack leverages hardware vulnerabilities like Rowhammer to induce bit flips, demonstrating a practical and potent method for runtime compromise of AI systems.
  • Urgent Need for Security Research: These findings highlight a critical gap in current AI security research and practices, emphasizing the urgent need for enhanced security measures in deep learning compilers, runtime environments, and hardware platforms to protect deployed DNN executables.

About the Speaker(s)

Yanzuo Chen is the presenter of this work. This research is a joint effort with Lien and Wangshai, as acknowledged by Chen during the presentation. The talk was delivered at the NDSS Symposium.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Genuinely novel attack surface — targeting compiled DNN operator code rather than model weights — with a credible gray-box threat model and a headline result (1.4 average bit flips for full intelligence depletion) that's hard to dismiss. The transferable vulnerable bits / super-bits methodology is the real contribution here: it's the kind of insight that only emerges from actually digging into what DL compilers emit, not just reading papers about Rowhammer. Not a 5 because defenses are deferred to future work and the targeted-misclassification angle is left underexplored.

Heather Calloway (CISO) — WEAK

Technically credible research with a genuinely novel finding — attacking compiled DNN code rather than model weights is a meaningful shift in the attack surface. But this talk is written for researchers, not defenders or decision-makers, and it stops exactly where institutional relevance begins.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025