Rowhammer-Based Trojan Injection: One Bit Flip Is Sufficient for Backdooring DNNs
Xiang Li
34th USENIX Security Symposium (USENIX Security '25) · Day 3 · ML and AI Security 3: Backdoors, Poisoning, Unlearning
Overview
This talk, presented by Xiang Li from George Mason University, unveils a groundbreaking attack named "OneFlip," demonstrating that a single bit flip is sufficient to inject a stealthy backdoor into deep neural networks (DNNs). The research, co-authored with Dr. Law and Dr. Chun, addresses the critical and escalating security concerns surrounding DNNs, which are now ubiquitous across various applications. Specifically, OneFlip focuses on inference-stage backdoor attacks, a more practical and insidious threat model compared to traditional training-stage attacks.

Key moments
- 0:00 Introduction to back door attacks and limitations
- 2:00 Limitations of existing inference stage attacks and new challenges
- 4:40 OneFlip: The novel one-bit flip attack workflow
- 5:20 Empirical method to identify eligible target bits
- 6:40 Optimizing trigger patterns for attack success and invisibility
- 7:20 OneFlip's near-perfect attack success with minimal degradation
- 8:40 Conclusion: One-bit flip sufficient for backdooring DNNs
Rowhammer-Based Trojan Injection: One Bit Flip Is Sufficient for Backdooring DNNs
Speakers: Xiang Li
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=hyeIWwFJd40
Overview
This talk, presented by Xiang Li from George Mason University, unveils a groundbreaking attack named "OneFlip," demonstrating that a single bit flip is sufficient to inject a stealthy backdoor into deep neural networks (DNNs). The research, co-authored with Dr. Law and Dr. Chun, addresses the critical and escalating security concerns surrounding DNNs, which are now ubiquitous across various applications. Specifically, OneFlip focuses on inference-stage backdoor attacks, a more practical and insidious threat model compared to traditional training-stage attacks.
The significance of OneFlip lies in its ability to compromise full-precision DNN models—a significant advancement over prior work that largely targeted quantized models—with an unprecedented minimal modification. By leveraging Rowhammer attacks to induce a single, precisely targeted bit flip, the researchers illustrate how an attacker can manipulate a model to misclassify specific inputs containing a crafted trigger while maintaining high accuracy on benign data. This work fundamentally shifts the understanding of DNN vulnerability, highlighting that even minor hardware-level perturbations can have profound, malicious consequences for AI systems.
The implications of OneFlip are substantial, as it bypasses the need for access to training data or facilities, making the attack far more practical for adversaries. It underscores a critical vulnerability at the intersection of hardware reliability and AI security, challenging the assumption that complex models require extensive manipulation to be compromised. The demonstration of a single-bit flip backdoor with near-perfect attack success rates and minimal impact on benign accuracy necessitates a re-evaluation of current DNN deployment and defense strategies.
Background
▶ Watch: Introduction to back door attacks and limitations (0:00)
Deep neural networks have rapidly become the backbone of countless modern applications, from autonomous vehicles and medical diagnostics to facial recognition and financial fraud detection. Their pervasive integration into critical infrastructure and everyday technology makes their security paramount. Among the myriad threats, backdoor attacks, often referred to as Trojans, stand out due to their stealth and potency. A backdoored model appears to function normally on most inputs, exhibiting high classification accuracy for benign data. However, when presented with a specific, attacker-chosen "trigger" embedded within an input, the model will deterministically output an attacker-desired result, deviating from its intended function. For instance, an autonomous driving system could be backdoored to classify any road sign with a small, inconspicuous patch as a "stop sign," regardless of its actual meaning.
Historically, conventional backdoor attacks primarily targeted the training stage of DNN development. These methods typically involve poisoning the training dataset with specially crafted samples that link the trigger to the target misclassification. While effective, this approach requires the attacker to have significant access to the training pipeline, including the dataset or the training infrastructure itself. This requirement often makes the threat model less practical for real-world scenarios where adversaries may not have such privileged access.
Recognizing these limitations, recent research has explored inference-stage backdoor injection methods. These attacks aim to modify the deployed model's weights directly, without needing to interfere with the training process. A prominent technique utilized in this context is bit flip attacks, such as Rowhammer. Rowhammer is a hardware vulnerability in DRAM that allows an attacker to flip bits in memory cells by repeatedly accessing adjacent memory rows, causing electrical interference. By carefully targeting specific memory locations, an attacker can induce bit flips in the stored weights of a DNN model during its inference phase. Once the target bits are flipped, any subsequent input containing the attacker's trigger will lead to the desired misprediction.
Despite the promise of inference-stage attacks, existing methods faced two significant limitations. First, they typically required flipping multiple bits to achieve an effective backdoor. Inducing multiple bit flips via Rowhammer is considerably more challenging and often deemed impractical or "invisible" due to the inherent stochasticity and difficulty of precisely controlling multiple bit flips. Second, prior work predominantly focused on quantized models, where weights are represented using low-bitwidth integers (e.g., 8-bit or 16-bit, often in two's complement). Full-precision models, which typically use 32-bit floating-point numbers adhering to the IEEE 754 standard, present a much larger and more complex search space for bit manipulation and are more sensitive to changes.
The primary goals of the research presented in this talk were to overcome these limitations: to realize a one-bit flip backdoor attack on full-precision models, thereby making inference-stage backdoor attacks significantly more practical and potent. The threat model assumed by the researchers aligns with prior inference-stage attacks: the attacker possesses white-box access to the model's parameters (i.e., they know the model architecture and weights) and owns a small set of benign samples. Crucially, to execute the Rowhammer attack, the attacker's process must co-reside on the same machine as the victim model, enabling memory-level manipulation.
Key Findings
▶ Watch: OneFlip: The novel one-bit flip attack workflow (4:40)
The "OneFlip" research represents a significant leap forward in understanding the vulnerabilities of deep neural networks to hardware-level attacks. The key findings are multifaceted and underscore the practical feasibility and potency of single-bit compromises:
- First One-Bit Flip Backdoor Attack: OneFlip is the pioneering work that demonstrates the successful injection of a backdoor into a DNN model by flipping only a single bit. This dramatically reduces the complexity and increases the practicality of such attacks compared to previous methods requiring multiple bit flips.
- First Inference-Stage Backdoor for Full-Precision Models: Prior inference-stage backdoor attacks largely focused on quantized models. OneFlip successfully extends this threat vector to full-precision models, which are more common in high-performance and critical applications, expanding the attack surface significantly.
- Novel Attack Workflow: The researchers devised an innovative, two-stage workflow tailored for single-bit flips. This workflow first meticulously identifies the most "eligible" bit to flip within the model's weights and then generates a corresponding trigger pattern specifically optimized to activate the altered weight. This targeted approach is crucial for achieving high attack efficacy with minimal modification.
- Near-Perfect Attack Success Rate with Minimal Benign Accuracy Degradation: OneFlip achieves near-perfect attack success rates (ASR), meaning that almost all inputs embedded with the crafted trigger are misclassified to the attacker's target class. Simultaneously, it induces only minimal degradation in the model's benign accuracy, ensuring the backdoor remains stealthy and the model's general performance is not noticeably affected.
- Profound Practical Implications: The core finding is that a highly effective and stealthy backdoor can be injected into a full-precision deep neural network model through a single, precisely targeted bit flip. This revelation highlights an alarming vulnerability, demonstrating that even subtle hardware-level faults or malicious manipulations can fundamentally compromise the integrity and trustworthiness of AI systems deployed in real-world scenarios.
Technical Deep Dive
▶ Watch: Empirical method to identify eligible target bits (5:20)
The design of OneFlip necessitated overcoming three significant technical challenges inherent in targeting full-precision models with a single bit flip. The researchers meticulously addressed these issues through a novel workflow and specific bit manipulation strategies.
Challenges in One-Bit Flip Attacks on Full-Precision Models
- Large Search Space: Quantized models typically store weights as 8-bit or 16-bit integers using two's complement representation. In contrast, full-precision models commonly employ 32-bit floating-point numbers, adhering to the IEEE 754 standard. A 32-bit floating-point number consists of a sign bit, an 8-bit exponent, and a 23-bit mantissa. This significantly larger number of bits, coupled with the complex structure of floating-point representation, means that bit search methods designed for simpler integer representations perform poorly, if at all, for full-precision models. A new, more intelligent bit search method was thus essential.
- Preserving Benign Accuracy: Modifying individual bits within a 32-bit floating-point number is far more intricate than altering bits in an integer. Flipping the most significant bit (MSB) of the exponent, for instance, can drastically change the magnitude of the number, potentially causing the entire deep neural network to malfunction or crash, thereby compromising stealth. Conversely, flipping a bit in the mantissa might result in too small a weight change to effectively inject a backdoor, failing to achieve the desired misclassification. The challenge was to identify a bit flip that would induce a substantial enough weight change to activate the backdoor, yet minimize its impact on the model's overall benign classification accuracy.
- Generating Effective Triggers: Existing backdoor attacks often pre-select a trigger pattern and then search for potential bits to flip that, when combined with this trigger, yield the desired output. However, with the constraint of only a single bit flip, a pre-selected trigger might not be potent enough to activate the specific weight containing the flipped bit to produce the attacker's target output. This constraint necessitated the development of an effective trigger generation strategy that is intrinsically aligned with the minimal change induced by a single bit flip.
OneFlip's Novel Attack Workflow
To address these challenges, OneFlip employs a two-stage workflow: first identifying the eligible bit to flip, then generating a corresponding trigger.
1. Identifying the Eligible Bit to Flip
The researchers conducted an empirical experiment, leading to three key observations regarding DNN weight sensitivity and exploitability:
- The classification layer of a deep neural network contains numerous weights that can be exploited.
- Many of these weights can be modified for backdoor injection without significantly degrading benign accuracy.
Based on these observations, OneFlip defines a highly specific criterion for an "eligible" bit within a positive weight (where the sign bit is 0) of the classification layer:
- The most significant bit (MSB) of the exponent must be zero.
- Exactly one of the remaining seven bits of the exponent must also be zero.
The crucial insight here is that flipping this specific zero among the remaining seven exponent bits to a one can dramatically increase the weight's value. In IEEE 754 floating-point representation, the exponent determines the magnitude of the number. By changing a zero to a one in the exponent, the weight's magnitude increases multiplicatively, making it significantly larger than other weights in the layer. This substantial increase is potent enough to make the weight "eligible" for backdoor injection. Crucially, because this flip avoids the MSB of the exponent (which would cause an extremely large change, potentially leading to NaN or Inf, or a complete sign change if it were the sign bit), it achieves the necessary impact for the backdoor without catastrophically degrading the model's benign accuracy. The paper provides more detailed insights into the design of this experiment and the derivation of these observations.
2. Generating the Trigger Pattern
Once the target weight and its specific bit to flip are identified, the next step is to generate an effective trigger pattern. This trigger must be capable of significantly amplifying the input associated with the flipped weight, thereby activating the backdoor. OneFlip uses an optimization formula for this purpose, balancing attack performance and trigger invisibility:
The objective function for trigger optimization involves two main terms:
- Left Term (Attack Performance): This term aims to maximize the input's activation over the specific, flipped weight. By ensuring a strong response from the altered weight, the trigger effectively "turns on" the backdoor, leading to the desired target class output. This is critical for achieving a high Attack Success Rate (ASR).
- Right Term (Trigger Invisibility): This term incorporates an L1 norm constraint on the trigger pattern. The L1 norm measures the sum of the absolute values of the trigger's elements, effectively encouraging sparsity and small magnitudes. By constraining the L1 norm, the trigger is designed to be as inconspicuous as possible, making it harder for human observers or automated detection systems to identify its presence.
The attacker can adjust a hyperparameter, lambda (λ), to prioritize either attack performance or trigger invisibility. A smaller lambda value emphasizes maximizing attack performance, potentially resulting in a more visible trigger, while a larger lambda value prioritizes trigger stealth at the possible expense of some attack efficacy.
Implementing the Bit Flip
After identifying the target bit and generating the corresponding trigger, the final step involves physically flipping the target bit in the model's memory. As mentioned in the background, OneFlip leverages Rowhammer attacks for this purpose. The attacker's process, co-residing on the same machine as the victim model, can repeatedly access specific memory rows adjacent to the memory location storing the target weight. This repeated access induces electrical interference, eventually causing the desired single bit flip in the victim's memory. Once the bit is flipped, any input embedded with the carefully generated trigger will activate the backdoor, leading to the attacker's desired classification.
Demo / Proof of Concept
▶ Watch: OneFlip's near-perfect attack success with minimal degradation (7:20)
The researchers conducted extensive evaluations to demonstrate the efficacy and practicality of OneFlip. Their proof-of-concept involved testing the attack on a diverse set of deep neural network architectures and widely used datasets.
Evaluation Setup:
- Datasets: Four widely used datasets were employed, though specific names were not detailed in the transcript, implying standard benchmarks like CIFAR-10, ImageNet, etc., which are common in DNN security research.
- DNN Architectures: Four popular deep neural network architectures were used, again, without specific names mentioned in the transcript. These would typically include architectures like ResNet, VGG, MobileNet, or variations thereof, representing different complexities and depths.
Evaluation Metrics:
To comprehensively assess the attack's success and stealth, three key metrics were utilized:
- Attack Success Rate (ASR): This metric quantifies the effectiveness of the backdoor attack. It is calculated as the percentage of samples from the test dataset, when embedded with the attacker's specified trigger pattern, that are correctly classified into the attacker's target class by the bit-flipped model. A high ASR indicates a successful backdoor.
- Benign Accuracy Degradation (BAD): This metric measures the impact of the bit flip on the model's original classification performance on benign (untriggered) inputs. It is calculated as the difference in accuracy between the flipped model and the original, unmodified model on the standard test dataset. Minimal BAD is crucial for the stealthiness of the backdoor.
- Bits-to-Flip: This metric simply counts the number of bits required to implement the attack. For OneFlip, this value is consistently one, highlighting its unprecedented efficiency and practicality.
Results:
The experimental results conclusively demonstrated the effectiveness of OneFlip:
- Near-Perfect Attack Success Rate: OneFlip consistently achieved a near-perfect ASR across all tested datasets and architectures. This indicates that once the single bit is flipped and the trigger is applied, the backdoor reliably forces the target misclassification.
- Minimal Impact on Benign Accuracy: The attack induced only minimal degradation in benign accuracy. This is a critical finding, as it confirms the stealthy nature of OneFlip; the compromised model continues to perform well on legitimate tasks, making the backdoor difficult to detect through routine performance monitoring.
- Single Bit Flip: As the name suggests, all successful attacks were achieved using just one bit flip, validating the core premise of the research.
Illustrative Examples:
The presentation also included visual examples of images with triggers under different lambda (λ) values, which control the trade-off between attack performance and trigger invisibility. These examples showed that even with varying degrees of trigger visibility, all triggered inputs were successfully classified as the target class (e.g., "airplanes," as specifically mentioned in the talk). This visual evidence further reinforces the attack's efficacy and the flexibility offered by the lambda parameter in tailoring the trigger's characteristics.
Defensive Implications
▶ Watch: Conclusion: One-bit flip sufficient for backdooring DNNs (8:40)
The "OneFlip" attack presents a stark and critical challenge to the security of deployed deep neural networks, demanding a re-evaluation of current defensive strategies. The revelation that a single bit flip, achievable through a hardware-level vulnerability like Rowhammer, can backdoor a full-precision DNN has profound implications.
Firstly, the most immediate implication is the increased threat posed by Rowhammer attacks. While Rowhammer has been known for years, its potential to precisely and maliciously compromise high-level AI applications with such minimal modification was not fully appreciated. Existing Rowhammer mitigations, such as Error-Correcting Code (ECC) memory, Target Row Refresh (TRR), and adjustments to DRAM refresh rates, primarily focus on preventing random bit flips and maintaining general system stability. However, OneFlip's targeted nature, identifying a specific "eligible" bit for maximum impact, suggests that these general mitigations might not be sufficient to prevent a determined attacker who can precisely locate and flip a critical bit. Defenders must consider whether existing hardware-level protections are robust enough against such highly targeted, application-aware bit flips.
Secondly, the attack highlights the critical need for memory integrity verification for DNN model weights, especially in environments where hardware-level attacks are plausible. Models deployed in multi-tenant cloud environments, on shared hardware, or in edge devices with less secure physical access are particularly vulnerable. Regular, cryptographic integrity checks of model weights in memory could potentially detect unauthorized modifications. However, continuous verification can introduce performance overhead, and the timing of such checks relative to the attack's execution window is crucial.
Thirdly, runtime monitoring for unusual model behavior gains new urgency. While OneFlip aims for minimal benign accuracy degradation, subtle shifts in model confidence scores, activation patterns, or performance on specific data subsets might serve as indicators. However, detecting a stealthy backdoor that maintains high overall benign accuracy is inherently challenging. More advanced runtime attestation mechanisms that verify the state of critical model components, perhaps even at the hardware level, could offer a more robust defense.
Fourthly, the attack underscores the importance of secure hardware-software co-design for AI systems. Merely securing the software stack is insufficient if underlying hardware vulnerabilities can be exploited to compromise the integrity of the AI model. Future DNN deployments, especially in security-critical domains, may require hardware-enforced memory isolation, trusted execution environments (TEEs), or specialized memory controllers designed to specifically thwart targeted bit-flip attacks against AI model parameters.
Finally, the practicality of OneFlip (requiring only a single bit flip and co-residence) significantly lowers the bar for attackers. This means that organizations deploying DNNs must assume that their models could be subject to such inference-stage backdoors. This necessitates a shift towards proactive threat modeling that includes hardware-level attack vectors and developing resilience strategies that can either prevent such flips or quickly detect and recover from them without compromising mission-critical operations. The research demands that the security community move beyond purely software-centric AI security and embrace a holistic approach encompassing the entire hardware-software stack.
Key Takeaways
- OneFlip pioneers a highly practical inference-stage backdoor attack, demonstrating that deep neural networks can be compromised without access to training data or facilities.
- It is the first attack to achieve a functional backdoor with a single bit flip in full-precision DNN models, significantly lowering the bar for adversaries compared to previous multi-bit or quantized-model attacks.
- The attack leverages a novel workflow that precisely identifies specific "eligible" bits within the exponent of floating-point weights and generates optimized triggers, ensuring both high attack success and minimal benign accuracy degradation.
- Targeting specific exponent bits in IEEE 754 floating-point numbers is crucial for OneFlip's success, allowing for substantial weight changes necessary for a backdoor while preserving overall model accuracy.
- Rowhammer attacks pose a significant, low-cost hardware threat to deployed DNN models, highlighting a critical intersection between hardware vulnerabilities and AI security.
- Defenders must now consider sophisticated hardware-level attacks when securing DNNs, necessitating robust memory integrity checks, advanced runtime monitoring, and secure hardware-software co-design to mitigate such threats.
About the Speaker(s)
The talk was presented by Xiang Li (Shan Lei), a second-year PhD student at George Mason University. Xiang Li represented his co-authors, Dr. Law and Dr. Chun, in presenting this research work. His involvement as a PhD student in this groundbreaking research highlights emerging talent in the field of AI security.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
OneFlip is genuinely novel work — first inference-stage backdoor on full-precision DNNs via a single Rowhammer bit flip, with a principled method for identifying eligible IEEE 754 exponent bits and co-optimizing the trigger. The core contribution is real and the attack surface expansion (quantized → full-precision, multi-bit → one bit) is a meaningful step forward. A PhD student presenting advisor-supervised work, so credibility questions are about the lab, not the idea.
Heather Calloway (CISO) — WEAK
Technically credible research that advances the DNN attack surface in a meaningful way — single-bit Rowhammer injection on full-precision models is a real finding. But the talk stops at the exploit and never reaches the operational or institutional questions that matter: who is deploying these models in environments where co-resident hardware attacks are plausible, what organizational controls are failing, and what defenders are actually supposed to do differently Monday morning.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)