GoFetch: Breaking Constant-Time Cryptographic Implementations Using Data Memory-Dependent Prefetchers
Boru Chen, Yingchen Wang, Pradyumna Shome, Christopher Fletcher, David Kohlbrenner, Riccardo Paccagnella, Daniel Genkin
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
The talk "GoFetch: Breaking Constant-Time Cryptographic Implementations Using Data Memory-Dependent Prefetchers" unveils a critical vulnerability in the security guarantees of constant-time programming on modern Apple M-series CPUs. Presented at USENIX Security '24 by a collaborative team of researchers, this work demonstrates how a hardware feature, specifically the Data Memory-Dependent Prefetcher (DMP), can inadvertently reintroduce secret-dependent timing variations into cryptographic code explicitly designed to prevent them. This discovery fundamentally challenges the long-held assumption that constant-time software practices are sufficient to thwart timing attacks.

Key moments
- 0:00 Timing attacks, constant-time programming, and defenses.
- 2:00 Apple's DMP breaks constant-time cryptographic implementations.
- 3:00 How DMP creates secret-dependent cache states.
- 4:00 GoFetch's key contributions beyond previous research.
- 6:00 Single memory load triggers Apple's DMP.
- 7:00 How DMP scans for pointers in L1 cache.
- 8:00 Chosen input attack to overcome pointer leak limitation.
GoFetch: Breaking Constant-Time Cryptographic Implementations Using Data Memory-Dependent Prefetchers
Speakers: Boru Chen, Yingchen Wang, Pradyumna Shome, Christopher Fletcher, David Kohlbrenner, Riccardo Paccagnella, Daniel Genkin
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=GxfF6W1tNAI
Overview
The talk "GoFetch: Breaking Constant-Time Cryptographic Implementations Using Data Memory-Dependent Prefetchers" unveils a critical vulnerability in the security guarantees of constant-time programming on modern Apple M-series CPUs. Presented at USENIX Security '24 by a collaborative team of researchers, this work demonstrates how a hardware feature, specifically the Data Memory-Dependent Prefetcher (DMP), can inadvertently reintroduce secret-dependent timing variations into cryptographic code explicitly designed to prevent them. This discovery fundamentally challenges the long-held assumption that constant-time software practices are sufficient to thwart timing attacks.
The research by Boru Chen, Yingchen Wang, Pradyumna Shome, Christopher Fletcher, David Kohlbrenner, Riccardo Paccagnella, and Daniel Genkin exposes a new class of cache timing attacks capable of extracting cryptographic keys from widely deployed and even post-quantum cryptographic implementations. By reverse-engineering the Apple DMP's behavior, the team developed novel chosen input attacks that transform seemingly random intermediate cryptographic states into observable pointer values, thereby leaking secret information. The implications are profound, necessitating a re-evaluation of hardware-software co-design for secure cryptographic operations, particularly on Apple Silicon.
This talk is crucial for security researchers, cryptographers, and hardware architects, as it highlights a sophisticated side-channel attack that bypasses established software defenses. The findings underscore the continuous cat-and-mouse game between attackers and defenders, where new hardware optimizations, while beneficial for performance, can introduce unforeseen security risks. GoFetch not only provides a deep technical analysis of the vulnerability but also demonstrates practical, end-to-end key extraction attacks against real-world cryptographic primitives, making it a significant contribution to the field of side-channel research.
Background
▶ Watch: Timing attacks, constant-time programming, and defenses. (0:00)
Timing attacks are a well-understood class of side-channel vulnerabilities where an attacker observes the execution time of a program to deduce secret information. These attacks exploit minute, secret-dependent variations in execution time, often stemming from differences in memory access latency. A common variant, cache timing attacks, leverages the processor's cache hierarchy. Data residing in the fast cache can be accessed much quicker than data in main memory. By measuring the time it takes to access a specific memory address, an attacker can infer whether that data was in the cache or had to be fetched from slower main memory, thereby revealing details about the program's execution path or data values.
To defend against such attacks, the cryptographic community developed constant-time programming paradigms. The core principle is to ensure that a program's execution time is entirely independent of any secret values it processes. This is achieved by adhering to three best practices:
- Secret-independent control flow: No conditional branches or loops whose execution depends on a secret.
- Secret-independent memory access patterns: All memory accesses (reads and writes) must occur at predictable, secret-independent addresses and times.
- Secret-independent instruction operands: Instruction timings should not vary based on the values of secret operands (e.g., avoiding variable-time multiplication instructions).
For many years, following these principles was believed to render cryptographic implementations immune to timing attacks, allowing cryptographers to "sleep well at night." However, the increasing complexity of modern processors, with their sophisticated out-of-order execution engines and aggressive prefetching mechanisms, has begun to challenge these assumptions.
A precursor to GoFetch, a work named Orion, was the first to identify the presence of a Data Memory-Dependent Prefetcher (DMP) on Apple M1 chips. DMPs are a type of hardware prefetcher designed to improve performance by detecting pointer chasing access patterns. When a program loads a value that looks like a memory address (a pointer), the DMP proactively fetches the data at that address into the cache, anticipating future program needs. While excellent for performance, Orion suggested that this behavior could be a security risk. However, Orion was unable to demonstrate a practical attack because its understanding of the DMP's activation conditions was incomplete, relying on specific "array of pointer" access patterns that are not typical in security-critical software. GoFetch builds directly upon this initial discovery, providing a much deeper understanding of the DMP's true behavior and, critically, demonstrating how to weaponize it against constant-time cryptography.
Key Findings
▶ Watch: How DMP creates secret-dependent cache states. (3:00)
GoFetch makes three pivotal contributions that significantly advance the understanding and exploitation of Apple's Data Memory-Dependent Prefetchers (DMPs), overcoming the limitations of prior work like Orion:
- Comprehensive Reverse Engineering of DMP Activation:
- Contrary to Orion's findings, GoFetch discovered that the intricate "array of pointer" access pattern is not necessary to activate Apple's DMP. This was a critical breakthrough, as the previous understanding made practical exploitation difficult.
- The researchers found that even a single memory load is sufficient to trigger the DMP. This means that potentially any data loaded from memory is a candidate to be treated as a pointer and subsequently dereferenced by the DMP, regardless of the program's explicit intent.
- Furthermore, not only the specific pointer fetched by the program but other pointer-like values within the same cache line are also considered and potentially dereferenced by the DMP. This broad scanning behavior significantly increases the attack surface.
- Through experimentation, GoFetch determined that the DMP primarily monitors L1 cache fills. When data is brought into the L1 cache, the DMP scans it to identify potential pointer values. It checks each pointer-sized aligned chunk within the cache line. For a successful dereference, the target address must map to a valid memory region within the process. This refined understanding of DMP mechanics was key to developing practical attacks.
- Novel Chosen Input Attack to Overcome Pointer-Only Leakage:
- A significant limitation of DMPs for side-channel attacks is that they only leak information about pointer values (i.e., whether a loaded value is treated as an address and dereferenced). However, intermediate cryptographic states often look like random bit strings, not predictable pointers.
- To overcome this, GoFetch introduced a new type of chosen input attack. The core idea is for the attacker to craft a specific input such that, depending on the secret value, an intermediate cryptographic state will deterministically become either a recognizable pointer or a non-pointer (or a different pointer).
- By carefully engineering the input, the attacker can create a scenario where the DMP's subsequent dereferencing behavior (or lack thereof) directly reveals bits of the secret. This method effectively transforms the "random-looking" cryptographic state into an observable side-channel signal.
- End-to-End Key Extraction Attacks Against Real-World Cryptography:
- Leveraging their deep understanding of DMP behavior and the chosen input attack methodology, GoFetch successfully built end-to-end key extraction Proofs of Concept (PoCs).
- These attacks demonstrated the feasibility of undermining full constant-time cryptographic implementations. The targeted primitives included both classical cryptography (e.g., Diffie-Hellman key exchange, GoST RSA decryptions) and emerging post-quantum cryptography standards (e.g., lattice-based digital signature standard, Key Encapsulation Mechanisms - KEMs).
- The success against such a diverse range of cryptographic schemes, some of which are deployed in the wild and others submitted for NIST standardization, highlights the widespread impact and severity of the GoFetch vulnerability. This practical demonstration validated that the theoretical security implications of DMPs are indeed exploitable in real-world scenarios.
These findings collectively demonstrate that constant-time programming alone is insufficient to guarantee security against timing attacks on Apple M-series processors, as hardware optimizations like DMPs can introduce secret-dependent side channels even in meticulously crafted code.
Technical Deep Dive
▶ Watch: GoFetch's key contributions beyond previous research. (4:00)
The technical ingenuity of GoFetch lies in its meticulous reverse engineering of the Apple M-series DMP and the subsequent development of a chosen input attack to exploit its behavior.
Traditional hardware prefetchers, such as stride prefetchers, typically operate by observing patterns in memory access addresses (e.g., A, A+stride, A+2*stride) and then prefetching A+3*stride. They rely solely on the address history. Apple's DMP, however, introduces a crucial difference: it considers not only the previous access address traces but also the content in data memory. This is what makes it a data memory-dependent prefetcher.
GoFetch's reverse engineering revealed that the DMP does not require the complex "array of pointer" access pattern that Orion initially suggested, nor does it have a "phase of confidence accumulation." Instead, its activation is far more aggressive and pervasive:
- Single Load Activation: A single memory load operation is sufficient to trigger the DMP. When a program loads data from a memory address into a register, that data is brought into the L1 cache. At this point, the DMP begins its scrutiny.
- Cache Line Scanning: The DMP scans the entire cache line that was just filled into the L1 cache. It doesn't just look at the specific 64-bit (or 32-bit depending on architecture) value loaded by the program; it examines all pointer-sized aligned chunks within that cache line. For instance, if a cache line is 64 bytes, and pointers are 8 bytes, the DMP might check eight potential pointer locations within that single cache line.
- Pointer Identification: For each chunk, the DMP checks if it "looks like a pointer." While the exact heuristics are proprietary, this typically involves checking if the value falls within a valid memory region accessible by the current process. If a chunk is identified as a potential pointer, the DMP then attempts to dereference it, meaning it will proactively fetch the data at that "pointer's" target address into the cache.
- Valid Memory Region Check: A critical constraint is that for a pointer to be successfully dereferenced by the DMP, its target address must map to a valid memory region within the process's address space. This prevents the DMP from causing spurious page faults or accessing unauthorized memory, but it doesn't prevent side-channel leakage.
- Redundant Dereference Avoidance: The researchers also investigated how the DMP avoids redundant dereferences when encountering duplicate pointer values. While the talk refers to the paper for full details, this mechanism is important for performance and potentially for limiting certain types of repeated leakage.
The fact that the DMP only leaks information about whether a value is a pointer (and thus triggers a prefetch) or is not a pointer (and thus does not) presents a significant challenge. Cryptographic intermediate states are generally designed to be indistinguishable from random data, making it difficult to directly craft them into pointers. This is where the chosen input attack comes into play.
The high-level idea of the chosen input attack is to manipulate the input to a cryptographic operation such that the intermediate state becomes secret-dependent in a way that the DMP can observe. Consider a simple example: an attacker wants to learn a 64-bit secret S which could be either all 0s or all 1s. The attacker chooses an input I that is a known valid pointer value (e.g., 0x100000000). The target cryptographic operation is an AND operation between I and S.
- If
S = 0xFFFFFFFFFFFFFFFF(all ones), thenI AND S = I. The intermediate crypto state is the pointerI. - If
S = 0x0000000000000000(all zeros), thenI AND S = 0x0000000000000000. The intermediate crypto state is0.
When the program loads this intermediate state, if it's the pointer I, the DMP will dereference I, causing a cache line fill at 0x100000000. If it's 0, the DMP will not dereference it (as 0 is not a valid pointer). An attacker can then monitor the cache state at 0x100000000 to deduce whether the DMP performed a dereference, thereby revealing whether S was all ones or all zeros.
While this example is simplified, the GoFetch team performed extensive cryptoanalysis to apply this principle to complex, real-world cryptographic algorithms. They identified points in the algorithms where specific chosen inputs could lead to such secret-dependent pointer formation in intermediate states. Their attacks targeted:
- Diffie-Hellman (DH) key exchange: A fundamental public-key agreement protocol.
- GoST RSA decryptions: RSA implementations following the Russian cryptographic standard.
- Lattice-based digital signature standard: An example of post-quantum cryptography, demonstrating the attack's relevance to future cryptographic paradigms.
- Key Encapsulation Mechanisms (KEMs): Another crucial component of post-quantum cryptography.
The ability to craft inputs that manipulate the internal state of these complex algorithms into observable pointer/non-pointer distinctions, combined with the detailed understanding of DMP behavior, forms the backbone of the GoFetch key extraction attacks.
Demo / Proof of Concept
▶ Watch: How DMP scans for pointers in L1 cache. (7:00)
The GoFetch team successfully developed and demonstrated end-to-end key extraction Proofs of Concept (PoCs) against a range of cryptographic implementations. These demonstrations moved beyond theoretical discussions to show practical exploitation, extracting secret keys from constant-time code running on Apple M-series processors.
The PoCs showcased the efficacy of their combined approach: the refined understanding of DMP activation coupled with the novel chosen input attack. While the talk refers to the full paper for the intricate details of each specific attack, the demonstrations broadly involved:
- Attacker Setup: The attacker first prepares a carefully crafted chosen input for the target cryptographic operation. This input is designed to create the secret-dependent pointer/non-pointer conditions in intermediate cryptographic states, as described in the technical deep dive.
- Execution and Observation: The victim's constant-time cryptographic implementation is executed on an Apple M-series CPU using the attacker's chosen input. During this execution, the DMP's behavior is influenced by the intermediate secret values. The attacker, running as a co-located process (e.g., on the same machine), monitors the cache state of specific memory addresses. These addresses are those that would be dereferenced by the DMP if a particular secret-dependent pointer were formed.
- Key Deduction: By observing cache hits or misses at these monitored addresses, the attacker can infer whether the DMP performed a dereference. This inference, in turn, reveals information about the secret-dependent intermediate state, allowing the attacker to deduce bits of the secret key. This process is repeated with different chosen inputs or by iteratively deducing bits until the entire key is recovered.
The range of targets for these PoCs was significant, encompassing both established classical cryptographic algorithms and emerging post-quantum cryptographic (PQC) schemes. Specifically, the team demonstrated key extraction against:
- Diffie-Hellman key exchange: A foundational protocol for secure communication.
- GoST RSA decryptions: Highlighting vulnerabilities in a specific national standard for RSA.
- Lattice-based digital signature standard: Crucially showing that even new PQC constructions, designed with strong security properties, are not immune to this class of hardware side-channel attacks.
- Key Encapsulation Mechanisms (KEMs): Another core PQC primitive, further emphasizing the broad applicability of GoFetch.
The success of these end-to-end attacks against a diverse set of cryptographic primitives, some of which are deployed in real-world systems or are candidates for future standardization, underscores the severe practical impact of the GoFetch findings. It demonstrates that the theoretical threat posed by DMPs is indeed a tangible risk for cryptographic security on Apple Silicon.
Defensive Implications
▶ Watch: Chosen input attack to overcome pointer leak limitation. (8:00)
The GoFetch research has significant and immediate defensive implications for hardware vendors, software developers, and the cryptographic community. It highlights that the traditional constant-time programming paradigm, while essential, is no longer a complete defense against timing attacks when sophisticated hardware features like DMPs are present.
- Hardware-Level Mitigations are Crucial:
- As a direct response to GoFetch, Apple has released guidance regarding a new mechanism for disabling the DMP. Specifically, for Apple M3 processors and potentially newer generations, there is a new hardware bit that can be enabled to disable the DMP for constant-time cryptographic operations. This is a critical step towards mitigating the vulnerability at its root.
- However, this mitigation currently does not apply to earlier generations like the M1 and M2 chips, which remain vulnerable. This creates a challenging situation for systems running on these older, but still widely used, processors.
- Software Opt-in Disablement:
- The Go programming language community is planning to add an opt-in disable bit in its binary. This would allow users or developers with concerns about GoFetch-type attacks to explicitly disable the DMP for their Go binaries on affected Apple CPUs, taking extra precaution. This approach shifts some responsibility to the software layer, requiring developers to be aware of the issue and opt into the mitigation.
- Kernel-Space Control for M1/M2:
- Marin, a prominent figure in the Asahi Linux project (which ports Linux to Apple Silicon), discovered a kernel-space bit that can disable the DMP on Apple M1 and M2 chips. This is a significant finding for users of these earlier platforms, as it offers a potential path to mitigation, albeit requiring kernel-level modification. Marin has indicated plans to develop a patch for this.
- Re-evaluation of Constant-Time Guarantees:
- GoFetch serves as a stark reminder that software-level constant-time guarantees can be undermined by unforeseen hardware behavior. Cryptographers and security architects must now consider a broader threat model that includes aggressive hardware optimizations as potential side-channel sources.
- Future cryptographic implementations, especially those targeting specific hardware platforms, may require more rigorous analysis of the underlying microarchitecture for hidden side channels, beyond just instruction timing and memory access patterns.
- Platform-Specific Security Considerations:
- The research emphasizes the need for platform-specific security considerations. A cryptographic library deemed constant-time on one architecture (e.g., Intel/AMD with different prefetcher behaviors) may not be so on another (e.g., Apple M-series). Developers must be aware of the specific microarchitectural features of their target deployment environments.
In summary, the defensive implications are multi-faceted: hardware vendors must design and implement mitigations, software developers need to be aware of the problem and potentially opt into available disable features, and the cryptographic community must update its understanding of what constitutes "constant-time" in the face of increasingly complex and optimized hardware.
Key Takeaways
- Data Memory-Dependent Prefetchers (DMPs) on Apple M-series CPUs undermine constant-time cryptography: These hardware features, designed for performance, reintroduce secret-dependent timing variations even in code meticulously written to be constant-time.
- GoFetch provides a refined understanding of DMP behavior: It demonstrates that DMPs can be triggered by a single memory load and aggressively scan entire cache lines for pointer-like values, a significant departure from previous assumptions.
- Novel chosen input attacks overcome DMP's "pointer-only" leakage limitation: Attackers can craft inputs to transform seemingly random intermediate cryptographic states into observable pointer/non-pointer distinctions, making secrets visible to the DMP.
- Practical key extraction attacks are feasible against diverse cryptographic schemes: GoFetch successfully extracted keys from classical (Diffie-Hellman, GoST RSA) and post-quantum (lattice-based signatures, KEMs) cryptographic implementations.
- Hardware-level mitigations are necessary but not universally available: Apple has introduced a DMP disable bit for M3 and newer chips, but M1 and M2 remain vulnerable without kernel-level intervention (e.g., from Asahi Linux).
- Constant-time programming is no longer a complete security panacea: While crucial, software-only constant-time guarantees are insufficient against sophisticated hardware side channels, necessitating a re-evaluation of security assumptions and hardware-software co-design.
About the Speaker(s)
The "GoFetch" research was a collaborative effort by a distinguished team of security researchers: Boru Chen, Yingchen Wang, Pradyumna Shome, Christopher Fletcher, David Kohlbrenner, Riccardo Paccagnella, and Daniel Genkin. This group of experts from various academic institutions and research labs focused their collective expertise on microarchitectural side channels and cryptographic analysis. Their work on GoFetch represents a significant contribution to the field, building upon prior knowledge to conduct comprehensive reverse engineering of Apple's Data Memory-Dependent Prefetcher (DMP) and devise novel exploitation techniques. Their findings have been recognized with the "Best Cryptographic Attacks" award by Pwnie Awards, underscoring the impact and ingenuity of their research in uncovering and weaponizing this critical vulnerability in modern hardware against constant-time cryptographic implementations.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research from USENIX Security '24 exposes a critical vulnerability in Apple M-series CPUs, demonstrating how the Data Memory-Dependent Prefetcher (DMP) undermines constant-time cryptographic implementations. GoFetch's novel chosen-input attacks and meticulous reverse engineering reveal a new class of side-channel attacks capable of extracting keys from both classical and post-quantum crypto. This work forces a fundamental re-evaluation of hardware-software security co-design, proving that constant-time software alone is insufficient.
Heather Calloway (CISO) — STRONG ACCEPT
This research fundamentally challenges the efficacy of constant-time programming as a sole defense against timing attacks on Apple M-series chips. The discovery of Data Memory-Dependent Prefetchers (DMPs) reintroduces secret-dependent timing variations, allowing for practical key extraction from critical cryptographic implementations. While mitigations exist for newer chips and some software, a significant installed base of M1/M2 devices remains exposed, demanding urgent institutional attention and risk re-evaluation.