Using Deep Learning Attribution Methods for Fault Injection Attacks
Black Hat Asia 2025 · Day 1 · Briefings
Overview
In a compelling presentation at Black Hat Asia, Karim, a Hardware Security Expert from Ledger's Dungeon security research team, unveiled a novel approach to significantly enhance the efficacy of fault injection attacks (FIA) against secure hardware. The talk, titled "Using Deep Learning Attribution Methods for Fault Injection Attacks," demonstrated how deep learning (DL) attribution methods, traditionally employed in side-channel analysis, can be repurposed to reverse-engineer the execution flow of black-box chips, thereby identifying precise timing windows for injecting faults. This methodology dramatically reduces the time and effort typically associated with brute-force fault injection, transforming a laborious, months-long endeavor into a targeted, efficient attack.

Key moments
- 0:00 Introduction to deep learning for fault injection attacks
- 1:10 Overview of hardware attacks: fault injection and side channels
- 2:40 Difficulties in blackbox fault injection attacks
- 5:50 Can deep learning automate fault injection vulnerability discovery?
- 7:00 General applications of deep learning in hardware security
Using Deep Learning Attribution Methods for Fault Injection Attacks
Speakers: Karim, Hardware Security Expert, Ledger (Dungeon Team)
Conference: Black Hat Asia
YouTube: https://www.youtube.com/watch?v=cjQIvLHUEws
Overview
In a compelling presentation at Black Hat Asia, Karim, a Hardware Security Expert from Ledger's Dungeon security research team, unveiled a novel approach to significantly enhance the efficacy of fault injection attacks (FIA) against secure hardware. The talk, titled "Using Deep Learning Attribution Methods for Fault Injection Attacks," demonstrated how deep learning (DL) attribution methods, traditionally employed in side-channel analysis, can be repurposed to reverse-engineer the execution flow of black-box chips, thereby identifying precise timing windows for injecting faults. This methodology dramatically reduces the time and effort typically associated with brute-force fault injection, transforming a laborious, months-long endeavor into a targeted, efficient attack.
Karim's research addresses a critical challenge in hardware security: the difficulty of performing fault injection attacks on chips where internal details are unknown – a "black-box" scenario. Such attacks often involve extensive trial and error, adjusting numerous physical and timing parameters to find a vulnerability. By leveraging the interpretability features of deep learning models, attackers can gain unprecedented insight into a chip's internal operations, even when protected by sophisticated countermeasures.
The implications of this work are profound for both attackers and defenders. For attackers, it offers a powerful new tool to bypass hardware security modules (HSMs) and secure elements. For defenders, it highlights the urgent need to integrate deep learning and machine learning expertise into their security evaluation processes, as traditional statistical analysis alone may no longer be sufficient to assess the resilience of their countermeasures against these advanced techniques. The research underscores a paradigm shift in hardware attack methodologies, emphasizing the growing convergence of artificial intelligence and physical security exploitation.
Background
▶ Watch: Introduction to deep learning for fault injection attacks (0:00)
Hardware attacks are broadly categorized into fault injection attacks (FIA) and side-channel attacks (SCA). Fault injection attacks are active attacks designed to intentionally perturb a chip's operation during sensitive computations, such as secure boot sequences, PIN verifications, or cryptographic operations (e.g., AES, RSA). The goal is to induce errors that can bypass security checks, extract secret keys (e.g., through Differential Fault Analysis - DFA), or modify program flow. Common techniques include power and clock glitches, electromagnetic fault injection (EMFI), optical fault injection (laser FI), and body biasing. Performing these attacks in a black-box scenario, where the attacker has no knowledge of the chip's internal architecture or firmware, is exceptionally challenging. It involves extensive brute-forcing of parameters like physical location on the chip (X-Y coordinates), precise timing, pulse width, and intensity, making it very time-consuming and resource-intensive.
In contrast, side-channel attacks are passive, non-invasive techniques. They involve monitoring information leaked by a chip during its operation, such as power consumption, electromagnetic emissions, or execution timing. Statistical methods like Differential Power Analysis (DPA) and Correlation Power Analysis (CPA) are often used to extract sensitive information, like cryptographic keys, from these leaked traces. More recently, deep learning (DL) profiling attacks have emerged as a powerful technique for SCA, particularly effective against chips employing countermeasures like masking or desynchronization.
The motivation behind Karim's research stems directly from the inherent difficulties of black-box fault injection. Secure elements and other protected chips often incorporate sophisticated countermeasures, including desynchronization and power blinding, which further complicate the identification of vulnerable moments for fault injection. Traditional reverse engineering tools can be helpful, but they often fall short in pinpointing the precise timing of critical security operations within a power or EM trace. The central question driving this work was whether deep learning could automate this crucial step, reducing the complexity and time required to find these "vulnerable moments" in a chip's execution.
Key Findings
▶ Watch: Overview of hardware attacks: fault injection and side channels (1:10)
The central contribution of this research is the successful adaptation of deep learning attribution methods from the domain of side-channel analysis to the more challenging realm of fault injection attack surface reverse engineering. Karim demonstrated that these methods can effectively pinpoint the precise timing windows within a power consumption trace where critical security configuration bits are manipulated, even when the chip employs countermeasures and operates as a black box.
Applied to the Analog Devices DS28C36 secure authenticator, this methodology uncovered a double verification countermeasure protecting EEPROM access. By identifying two distinct timing zones where security checks were performed, the research enabled a targeted double-fault injection attack, leading to the successful dumping of user data from the chip's EEPROM slots. This marks a significant breakthrough, transforming a previously months-long, brute-force fault injection effort into a highly efficient, data-driven process. The findings confirm that deep learning attribution methods provide an invaluable tool for understanding and exploiting complex hardware security mechanisms in black-box environments, thereby reducing the attacker's effort and increasing the likelihood of successful exploitation.
Technical Deep Dive
▶ Watch: Difficulties in blackbox fault injection attacks (2:40)
The talk began by establishing the established role of deep learning in hardware security, particularly for side-channel attacks. Karim illustrated how profiling attacks using deep learning models can effectively extract cryptographic keys. A practical example involved attacking AES (Advanced Encryption Standard) running on an STM32U5 32-bit MCU. By collecting electromagnetic emissions during AES execution with known plaintexts and keys, a dataset was built. This dataset was then fed into a simple Multi-Layer Perceptron (MLP) model, consisting of three dense layers. The model learned to predict the correct AES key byte with high accuracy, often requiring fewer than 10 traces. While this was a "toy example" without countermeasures, it demonstrated the fundamental capability of deep learning to learn from leaked hardware traces.
The core of Karim's innovation lies in the application of deep learning attribution methods. These methods are designed to interpret and understand the decisions made by a deep learning model, essentially reverse-engineering which input features had the most significant impact on the model's output. They can be broadly categorized into gradient-based methods and activation-based methods. A common analogy is an image classification task (e.g., MNIST dataset): if a model classifies an image as a "7," attribution methods can highlight the specific pixels in the image that led to that decision.
One prominent activation-based method discussed is Layer-wise Relevance Propagation (LRP). LRP works by calculating the "relevance" of each neuron in the network, typically as a product of its activation and weight, and then propagating this relevance backward from the output layer to the input layer. This process ultimately highlights the specific input features (e.g., individual samples in a power trace, or pixels in an image) that were most relevant to the model's final prediction. Other methods like Taylor decomposition or Input attribution exist and may yield similar results, often requiring experimentation to find the most effective one for a given dataset.
Prior research has already shown the efficacy of deep learning attribution methods as leakage detection techniques in side-channel analysis, even against advanced countermeasures like masking and desynchronization. Karim presented an example comparing LRP with traditional statistical tools like Signal-to-Noise Ratio (SNR) for detecting AES S-box leakage. The LRP output on a power trace closely mirrored the SNR curve, demonstrating its capability to identify leakage points. The advantage of DL, particularly Convolutional Neural Networks (CNNs), in this context is their ability to filter out noise and countermeasures such as delay or masking, making them more robust than classical statistical tests.
A crucial illustration of this robustness involved a scenario with significant desynchronization jitter in power consumption traces. While a traditional t-test produced a highly noisy and uninterpretable output, making it impossible for an attacker to pinpoint the manipulated value, LRP, trained to distinguish between "protected" (value 1) and "unprotected" (value 0) states, clearly identified the exact peak where the value was manipulated. This ability to cut through noise and pinpoint critical moments, even under desynchronization, forms the crucial bridge for applying these methods to fault injection. Instead of identifying data-dependent leakage, the goal shifts to identifying the exact moment a security decision is made or a countermeasure is activated.
Demo / Proof of Concept
▶ Watch: Can deep learning automate fault injection vulnerability discovery? (5:50)
The practical demonstration centered on attacking the Analog Devices DS28C36 secure authenticator chip, originally designed by Maxim Integrated. This chip, often found in various hardware wallets, features elliptic curve cryptography (ECC), a True Random Number Generator (TRNG), and an 8-kilobit EEPROM frequently used for storing private keys. The primary objective was to dump the contents of this EEPROM using fault injection.
Given that the datasheet for the DS28C36 is not publicly available, Karim's team undertook initial reverse engineering of the chip's commands by analyzing its usage in an open-source hardware wallet. This revealed the chip's memory organization, including 16 user data pages, several slots for public key cryptography, and RAM buffers.
The physical setup involved decapping the chip from the backside and using an infrared (IR) camera to visualize its internal layout. This revealed distinct areas: a large red region likely representing the flash or EEPROM, a logic area, a smaller memory block (RAM), and an analog section. A laser fault injection bench from Alfanov was employed, coupled with a custom-built scaffold board for I2C communication with the device under test, and a high-end oscilloscope to capture the chip's power consumption during command execution.
Initial black-box analysis involved storing a value in an unprotected EEPROM slot and then reading it. The captured power consumption trace revealed a distinct pattern: an initial calculation (Zone 1), followed by 32 distinct peaks indicating a 32-byte memory access (Zone 2), and finally, a subsequent operation believed to be decryption (Zone 3), implying data is stored encrypted in the EEPROM. Crucially, comparing power traces when the slot was unprotected versus protected showed a clear divergence point. In the protected case, the chip's power trace deviated significantly, and it would not reply, indicating a security decision was made before this divergence.
Traditional, brute-force fault injection attempts were then initiated. The team attempted to inject laser faults while reading a protected memory page, systematically varying laser power, physical position on the chip, and timing offset. This arduous process, which Karim admitted took "two months scanning the overall chip," yielded mostly crashes or communication errors. However, it also produced two "incorrect" dumps: one showing a pattern consistent with a public key slot, and another showing no EEPROM access pattern, likely related to RAM or RNG. These results, while not achieving the primary goal, highlighted the difficulty and the potential for misinterpretation in black-box fault injection.
This is where the deep learning attribution methods became critical. Karim pivoted to using DL to reverse-engineer the chip's security decision process. The methodology was as follows:
- Data Collection: Power consumption traces were collected for two scenarios: when the EEPROM slot was unprotected (labeled '0') and when it was protected (labeled '1').
- Preprocessing: Traces were concatenated and normalized to prepare them for the DL model.
- Model Training: A deep learning model (e.g., CNN or MLP) was trained to classify whether a given power trace corresponded to a protected or unprotected EEPROM read.
- Attribution Application: After the model was trained, LRP (Layer-wise Relevance Propagation) was applied to a test trace (e.g., from a protected read). The LRP output highlighted the most "relevant" samples in the power trace that contributed to the model's decision of "protected."
- Interpretation: The LRP analysis revealed two significant peaks in the power trace, occurring before the observed divergence point. These two distinct timing zones indicated that the chip was performing double verification of the security configuration bits. This immediately suggested that the chip had a countermeasure against single-fault attacks.
Armed with this crucial insight, the team shifted their fault injection strategy. Instead of brute-forcing the entire chip, they targeted the logic part (as identified by the IR camera) and focused on injecting two faults within the two timing windows identified by LRP. This targeted approach proved successful: they were able to induce a double fault that bypassed the double verification and successfully dumped the value previously stored in the EEPROM. The success was validated by analyzing the power consumption during the attack, which clearly showed the impact of two laser pulses followed by a distinct EEPROM memory access pattern and subsequent decryption.
Karim noted that this attack was successful for the 16 user data slots but failed for the public key crypto slots, suggesting different security mechanisms or execution scenarios for those specific memory regions. The tooling used for this research, Scandal, an open-source tool integrating various deep learning attacks and attribution methods, was also highlighted.
Defensive Implications
▶ Watch: General applications of deep learning in hardware security (7:00)
Karim's research carries significant implications for hardware security defenders. Firstly, it emphatically demonstrates that relying solely on traditional, black-box fault injection evaluations is no longer sufficient. The ability of deep learning attribution methods to quickly reverse-engineer critical execution timings and countermeasure mechanisms drastically reduces the attacker's effort and increases their success rate. What previously took months of brute-force experimentation can now be achieved in a highly targeted manner.
A critical takeaway for vendors is the urgent need to integrate deep learning and machine learning expertise into their security evaluation teams. Karim explicitly stated that traditional hardware security experts, often proficient in electronics and statistical tools, may lack the specialized knowledge required to anticipate and test against these advanced, AI-driven attack methodologies. Bridging this skill gap is paramount for developing robust, future-proof countermeasures.
Furthermore, the research exposes the vulnerability of designs that rely on single-fault attack protection. The Analog Devices DS28C36, despite implementing a double verification countermeasure, was ultimately bypassed by a targeted double-fault injection attack precisely because the deep learning attribution methods revealed the presence and timing of these multiple checks. This implies that simply adding more verification steps without robust underlying protection might be a losing battle if an attacker can precisely locate them. Defenders must move beyond merely increasing the number of checks and focus on more fundamental countermeasure principles.
Crucial countermeasures highlighted by Karim include desynchronization techniques and power blinding. The attacked chip's lack of power blinding made its power consumption traces easily interpretable, providing the necessary data for the deep learning model. Effective desynchronization, which aims to introduce variability in execution timing, would make it harder for deep learning models to pinpoint consistent "peaks" or vulnerable moments. Combining these techniques, along with other obfuscation methods, becomes essential to raise the bar against such sophisticated attacks. Ultimately, vendors must evaluate their countermeasures not just against classical statistical analysis, but against the most advanced, AI-assisted attack techniques available.
Key Takeaways
- Deep learning attribution methods provide a powerful new capability for reverse-engineering black-box secure chips, significantly reducing the effort required for fault injection attacks.
- These methods can accurately pinpoint the exact timing of critical security operations and countermeasure activations within power consumption traces, even under desynchronization.
- The Analog Devices DS28C36 secure authenticator, despite a double verification countermeasure, was successfully exploited via a targeted double-fault injection attack guided by deep learning attribution.
- Hardware security teams must incorporate ML/DL expertise to effectively evaluate and design countermeasures against these evolving attack vectors.
- Protection against single-fault attacks is insufficient; countermeasures must anticipate and resist multi-fault attacks, necessitating robust desynchronization and power blinding techniques.
- Open-source tools like Scandal are emerging that democratize these advanced deep learning-based hardware attack methodologies.
About the Speaker(s)
Karim is a Hardware Security Expert at Ledger, a prominent vendor of hardware wallets for cryptocurrency self-custody. He works as part of Ledger's dedicated security research team, known as "Dungeon." In this role, Karim's daily work involves conducting offensive hardware security research, targeting both Ledger's own products and those of their competitors, to identify and mitigate potential vulnerabilities.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
Karim from Ledger's Dungeon team delivered a groundbreaking presentation that marries deep learning attribution with fault injection attacks, fundamentally transforming black-box hardware exploitation. By intelligently repurposing DL to precisely identify vulnerable timing windows, he demonstrated a targeted double-fault attack on an Analog Devices secure authenticator, bypassing its double-verification countermeasure. This research sets a new, higher standard for hardware security evaluation, unequivocally demanding that defenders integrate ML expertise to counter these sophisticated, AI-assisted attack methodologies.
Heather Calloway (CISO) — MUST SEE
Karim's Black Hat Asia presentation on using deep learning attribution methods for fault injection attacks is a critical update for any CISO overseeing product security, hardware development, or supply chain risk. This research demonstrates a significant shift in hardware exploitation, transforming laborious brute-force attacks into targeted, efficient operations. It clearly articulates the immediate need for organizations to integrate advanced ML/DL expertise into their hardware security evaluation processes and to fundamentally re-evaluate the resilience of existing countermeasures against these sophisticated, AI-assisted attack methodologies. This isn't just a new exploit; it's a new…