Serberus: Protecting Cryptographic Code from Spectres at Compile-Time

Nicholas Mosier, Hamed Nemati, John C. Mitchell, Caroline Trippel

IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 5

Overview

The talk "Serberus: Protecting Cryptographic Code from Spectres at Compile-Time" introduces a novel, comprehensive defense mechanism designed to safeguard constant-time cryptographic code against Spectre attacks. Presented by Nicholas Mosier, a PhD student at Stanford University, the research addresses a critical vulnerability inherent in modern speculative execution processors that can undermine the security guarantees of carefully written cryptographic implementations. While constant-time programming meticulously avoids timing side channels in sequential execution, these defenses are often rendered ineffective during the transient execution phases exploited by Spectre.

Watch on YouTube

Visual summary for Serberus: Protecting Cryptographic Code from Spectres at Compile-Time by Nicholas Mosier, Hamed Nemati, John C. Mitchell, Caroline Trippel
Visual summary for Serberus: Protecting Cryptographic Code from Spectres at Compile-Time by Nicholas Mosier, Hamed Nemati, John C. Mitchell, Caroline Trippel

Key moments

  1. 0:00 Introduction to Cerberus and Spectre's threat to crypto code
  2. 1:00 Detailed example of a Spectre attack via return misprediction
  3. 2:00 Cerberus: A comprehensive Spectre defense and key challenges
  4. 2:50 Solution 1: Constraining speculative control flow using Intel CET
  5. 4:20 Visualizing hardware speculation constraints and performance overhead
  6. 6:30 Solution 2: Static constant time programming for secret arguments
  7. 7:50 Challenge 3: Inadequacy of existing approaches for Spectre root cause

Serberus: Protecting Cryptographic Code from Spectres at Compile-Time

Speakers: Nicholas Mosier, Hamed Nemati, John C. Mitchell, Caroline Trippel

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=w9NVVLBcDZU

Overview

The talk "Serberus: Protecting Cryptographic Code from Spectres at Compile-Time" introduces a novel, comprehensive defense mechanism designed to safeguard constant-time cryptographic code against Spectre attacks. Presented by Nicholas Mosier, a PhD student at Stanford University, the research addresses a critical vulnerability inherent in modern speculative execution processors that can undermine the security guarantees of carefully written cryptographic implementations. While constant-time programming meticulously avoids timing side channels in sequential execution, these defenses are often rendered ineffective during the transient execution phases exploited by Spectre.

Cerberus stands out as the first defense system to offer protection against all five documented speculation primitives on modern processors, a significant advancement over prior work that typically targets only one or a subset of these attack vectors. The project's importance lies in its ability to provide robust security guarantees for sensitive cryptographic operations, such as those found in libraries like OpenSSL and libsodium, without requiring developers to modify their existing constant-time code. This compile-time solution tackles the root causes of Spectre leakage by leveraging existing hardware features and introducing new compiler passes, demonstrating a practical and high-performing approach to a pervasive security challenge.

Background

▶ Watch: Introduction to Cerberus and Spectre's threat to crypto code (0:00)

Security-critical code, particularly in cryptography, relies heavily on constant-time programming to prevent the leakage of sensitive information through timing side channels. A constant-time program is meticulously crafted to ensure that its execution time and memory access patterns do not depend on secret values. For instance, a routine like salsa20 from libsodium would be implemented to avoid passing secrets to the unsafe operand of "transmitters"—instructions that can leak information. Common transmitters include memory accesses, conditional branches, and variable-time operations like division. The core assumption of constant-time programming is that code executes sequentially.

However, this assumption is fundamentally challenged by Spectre attacks. Spectre exploits speculative execution, a performance optimization in modern CPUs where the processor attempts to predict future operations and execute them ahead of time. If a prediction is incorrect (a misprediction), the transiently executed operations are rolled back. The critical flaw is that during this transient execution, sensitive data can still be used by transmitters, leaving microarchitectural traces (e.g., in cache) that an attacker can later observe, even if the transient operations are never architecturally committed. The talk illustrates this with a load32 function that loads a secret key. Sequentially, the secret is stored to the stack and doesn't leak. But if the processor mispredicts a return address (Return Address Prediction, or RSB), it might transiently jump to a transmitter that leaks the secret.

There are five well-documented speculation primitives on modern processors that can lead to mispredictions and subsequent Spectre leakage:

  1. Conditional Branch Prediction: Based on historical outcomes of conditional branches.
  2. Indirect Branch Prediction: Based on historical targets of indirect jumps or calls.
  3. Return Address Prediction (RSB): Based on the Return Stack Buffer, which predicts return targets.
  4. Store-to-Load Forwarding Prediction: Predicts that a load will read data from a prior store to the same address.
  5. Predictive Store Forwarding (PSFD): A broader category of store-to-load forwarding that can include forwarding from different addresses.

Previous attempts to defend against Spectre often focus on specific primitives or involve costly speculation fences, leading to significant performance overheads or incomplete protection. Furthermore, existing techniques for identifying Spectre leakage primarily focus on transient access instructions (instructions that load secrets into registers). However, in constant-time code, the initial access might be sequential, with the secret leaking later due to a transient misprediction involving a register that already holds a secret. This highlights a gap in understanding and mitigating the root causes of leakage in such sensitive contexts.

Key Findings

▶ Watch: Cerberus: A comprehensive Spectre defense and key challenges (2:00)

Cerberus presents several groundbreaking findings and contributions to the field of transient execution security:

  1. Comprehensive Spectre Defense: Cerberus is the first defense system to comprehensively secure constant-time cryptographic code against Spectre attacks exploiting any combination of the five widespread speculation primitives. This addresses a critical gap where prior mitigations were often piecemeal, targeting only one or a few primitives.
  2. Leveraging Hardware Speculation Constraints: The system effectively utilizes existing Instruction Set Architecture (ISA) extensions, specifically Intel Control Flow Enforcement Technology (CET) and other speculation control bits (like RSBA Disable and PSFD speculation control), to impose hardware-level constraints on speculative control flow and data flow. This significantly simplifies static analysis and reduces the attack surface.
  3. Introduction of Static Constant-Time Programming: Cerberus proposes an extension to traditional constant-time programming, requiring all variables to have static security types and, critically, that all secret arguments are passed by reference rather than by value. This design principle inherently prevents common Spectre leakage paths related to mispredicted calls and returns.
  4. Formal Definition of Taint Primitives: The research introduces the novel concept of taint primitives as the necessary and sufficient ingredient for all Spectre leakage in static constant-time programs. Taint primitives are defined as instructions that cause a publicly typed register to transiently hold a secret. The paper formally categorizes these into four distinct classes: Non-Constant Address Loads (NCAL), Non-Constant Address Stores (NCAST), Uninitialized Stack Loads (StackL), and Secret and Non-Argument Registers (NAR). This formalization provides a rigorous framework for identifying and eliminating leakage.
  5. Efficient Compiler-Based Mitigation: Cerberus is implemented as three LLVM compiler passes that collectively eliminate all taint primitives from all transient executions of a program. These passes are Function Private Stacks, Register Cleaning, and Fence Insertion.
  6. Superior Performance and Security Guarantees: Despite offering significantly stronger security guarantees than the state-of-the-art, Cerberus achieves remarkably low performance overhead. On average, it exhibits only 21% overhead across a variety of cryptographic benchmarks (from OpenSSL, libsodium, and HCstar), dropping to 7% overhead for benchmarks processing larger inputs. This makes it a highly practical solution.
  7. Programmer Transparency: Cerberus is designed to be 100% programmer-transparent. It requires no secrecy labels, annotations, or code modifications from the developer, only that the input program satisfies static constant-time principles and that a single compile-time flag is enabled.
  8. Formal Verification: The paper includes a complete proof of Cerberus's security guarantees, based on an operational semantics called Abstract Speculative Processor (ASP), which models transient execution on speculative out-of-order processors.

Technical Deep Dive

▶ Watch: Solution 1: Constraining speculative control flow using Intel CET (2:50)

Cerberus's technical foundation rests on a combination of hardware speculation constraints, a refined programming model, and a suite of compiler passes.

Hardware Speculation Constraints

A cornerstone of Cerberus is its strategic use of existing hardware features to limit the scope of transient execution, transforming an otherwise "messy" and unanalyzable control flow graph into a tractable one.

  1. Constraining Speculative Control Flow:
  • Intel Control Flow Enforcement Technology (CET): Cerberus leverages CET, available on recent Intel microarchitectures like Alder Lake-N and Arizona Beach, for speculative Control Flow Integrity (CFI) protections.
  • Indirect Branch Tracking (IBT): By enabling CET's IBT, all indirect calls and jumps are constrained to transiently target only Endbranch instructions, which are inserted at each function entry point. This prevents arbitrary transient jumps.
  • Shadow Stack and RSBA Disable: Together, CET's Shadow Stack and the Return Stack Buffer Auxiliary (RSBA) Disable speculation control ensure that ret instructions can only transiently return to legitimate call sites within the program. This effectively prunes the vast majority of unpredictable return paths.
  • These controls collectively transform the transient control flow graph from an unconstrained, complex network to one where paths are analyzable at compile time.
  1. Constraining Speculative Data Flow:
  • Predictive Store Forwarding (PSFD) Speculation Control: Cerberus enables the PSFD speculation control. This crucial setting ensures that transient loads can only read from prior stores to the same memory address. Without this, a load could transiently read from any prior store, making static analyses like alias analysis useless in a speculative context. By enforcing same-address forwarding, Cerberus can perform a rudimentary but effective speculative alias analysis, which is vital for its performance.

The overhead of these hardware speculation constraints alone is approximately 3% on the SPEC 2017 benchmarks, indicating their efficiency.

Static Constant-Time Programming

To make Spectre leakage rigorously identifiable and mitigable, Cerberus introduces static constant-time programming, an extension to the conventional constant-time model. It imposes two additional requirements:

  1. All variables must have static security types. (Details are elaborated in the paper).
  2. All secret arguments must be passed by reference, not by value.

The second rule directly addresses a major vulnerability: passing secrets by value or returning them by value makes them inherently susceptible to leakage if the call or return instruction is mispredicted to a transmitter. The example of load32 returning a secret value is modified to take a pointer argument, storing the secret into that pointer instead. This ensures that the secret never resides in a general-purpose register during the potentially mispredicted call/return boundary.

Taint Primitives: The Root Cause of Leakage

Cerberus's core insight is the concept of taint primitives. A taint primitive is defined as an instruction that causes a publicly typed register to transiently hold a secret value. The research formally proves that taint primitives are a necessary ingredient for all Spectre leakage in static constant-time programs. This provides a precise target for mitigation.

In static constant-time programs, all taint primitives fall into one of four classes:

  1. Non-Constant Address Loads (NCAL): These are loads whose addresses are dynamically computed from data at runtime. During transient execution, such a load might access an out-of-bounds memory location, inadvertently reading a secret from an adjacent memory region.
  2. Non-Constant Address Stores (NCAST): Similar to NCAL, these are stores whose addresses are dynamically computed. A misprediction could lead to a secret being transiently stored to an out-of-bounds public memory location, from which it could later be recovered.
  3. Uninitialized Stack Loads (StackL): This occurs when a load from the stack bypasses an initialization store. If a new stack frame reuses a location from a previous, secret-holding stack frame, a transient load might read the stale secret before the intended initialization completes.
  4. Secret and Non-Argument Register (NAR): This class applies to calls or returns that are mispredicted to an incorrect target. At the transient target, a register that is expected to hold a public argument (or be unused) might instead contain a secret value due to prior register allocation, leading to its accidental leakage.

Cerberus Compiler Passes (Implemented in LLVM)

Cerberus eliminates these taint primitives using three compiler passes integrated into LLVM:

  1. Function Private Stacks: This pass addresses the StackL taint primitive. The problem arises because procedures typically share a common stack, allowing a transient load to bypass an initialization and read stale secret data from a previous stack frame. Cerberus solves this by assigning a distinct, private stack to each function. When a function is called, it allocates its stack frame on its unique stack. This ensures that a function's stack frame cannot overlap with or contain stale data from another function's prior execution, thus eliminating StackL.
  1. Register Cleaning: This pass targets the NAR taint primitive. It identifies instances where secret values might persist in registers across function boundaries, particularly at calls and returns. For example, if load32 returns a secret via a register, and that return is mispredicted to a location expecting a public value in that same register, a leak occurs. The Register Cleaning pass proactively zeros out registers that might hold secrets at call and return boundaries. This ensures that no secret value remains in a publicly typed register in a context where a misprediction could expose it.
  1. Fence Insertion: This final pass addresses NCAL and NCAST taint primitives, which are not handled by the previous two passes. It frames optimal speculation fence insertion as a min-cut problem over the transient control flow graph (TCFG) of each procedure.
  • First, a TCFG is constructed, representing all possible transient execution paths.
  • Next, "source-sink pairs" are identified. Sources are NCAL or NCAST taint primitives. Sinks are either dependent transmitters (which would leak the tainted data) or interprocedural control flow instructions (calls/returns).
  • A heuristic min-cut algorithm is then applied to find a minimal set of transient control flow edges to remove. These edges are "removed" by inserting speculation fences (e.g., LFENCE on x86) at those points, effectively preventing transient execution beyond the fence and thus eliminating the NCAL and NCAST taint primitives.

Crucially, Cerberus is 100% programmer-transparent. Developers only need to ensure their input program adheres to static constant-time principles and enable Cerberus with a single compile-time flag.

Demo / Proof of Concept

▶ Watch: Solution 2: Static constant time programming for secret arguments (6:30)

While the talk did not feature a live demonstration of Cerberus in action, the speakers detailed its implementation within LLVM and presented comprehensive evaluation results that serve as a robust proof of concept for its effectiveness and performance.

The evaluation compared Cerberus against an insecure baseline and two composite state-of-the-art defenses. These mitigations were benchmarked on a variety of cryptographic primitives sourced from widely used libraries: OpenSSL, libsodium, and HCstar. The benchmarks covered different cryptographic operations and input sizes to provide a holistic view of Cerberus's impact.

Key results from the evaluation demonstrated Cerberus's superior performance relative to its strong security guarantees:

  • On average, Cerberus exhibited only 21% overhead across all benchmarks.
  • For benchmarks processing larger inputs, the overhead dropped significantly to just 7%. This performance is particularly impressive, as it represents half the overhead of the next best mitigation, despite Cerberus offering stronger and more comprehensive security guarantees (protecting against all five speculation primitives, including RSB, indirect branch prediction, and store-to-load forwarding prediction, which other mitigations often disable or fail to address).

The qualitative comparison further underscored Cerberus's unique position:

  • Most prior mitigations target only a single speculation primitive (often conditional branch prediction, PHT).
  • For speculation primitives other than PHT, most existing mitigations simply disable speculation rather than securing it, leading to higher performance penalties.
  • Even composing the "best" individual mitigations for each primitive (termed "securest state-of-the-art") would still fail to mitigate Spectre RSB and would disable most forms of speculation.

In contrast, Cerberus is the first mitigation for constant-time code that:

  • Successfully mitigates return prediction (RSB).
  • Permits secure indirect branch prediction.
  • Permits secure store-to-load forwarding prediction.

These evaluation results, coupled with the formal proofs presented in the accompanying paper, firmly establish Cerberus as a viable, high-performance, and comprehensive solution for protecting cryptographic code from Spectre attacks.

Defensive Implications

▶ Watch: Challenge 3: Inadequacy of existing approaches for Spectre root cause (7:50)

Cerberus offers profound implications for developers, compiler engineers, hardware architects, and security researchers aiming to build and deploy secure systems in the face of transient execution attacks:

  1. For Developers of Security-Critical Code: Cryptographic library developers and application developers using constant-time code can adopt Cerberus with minimal effort. By simply enabling a compile-time flag, their existing constant-time code can gain comprehensive protection against all known Spectre variants without requiring manual code changes or annotations. This significantly lowers the barrier to deploying robust transient execution defenses in highly sensitive contexts.
  2. For Compiler Developers and Toolchain Maintainers: The implementation of Cerberus as LLVM passes demonstrates a clear path for upstreaming this technology into mainstream compilers. Integrating Cerberus directly into LLVM would make comprehensive Spectre protection widely accessible to all developers using the compiler, transforming it into a default security posture for constant-time code. This reduces the burden on individual projects to implement their own mitigations.
  3. For Hardware Architects and CPU Manufacturers: Cerberus's reliance on and effective utilization of Intel CET and specific speculation control bits (like RSBA Disable and PSFD speculation control) highlights the critical importance of designing and documenting processor features with clear and well-defined speculative semantics. Future hardware designs should continue to provide such architectural controls to enable robust software-based mitigations against transient execution attacks. The work serves as a testament to how carefully designed hardware features can be leveraged for stronger security.
  4. For Security Researchers and Academics: The formal framework introduced by Cerberus, including the Abstract Speculative Processor (ASP) operational semantics and the concept of taint primitives, provides a powerful new lens for analyzing transient execution vulnerabilities. This rigorous foundation can inform future research into new attack vectors, aid in the design of provably secure systems, and facilitate the development of automated vulnerability detection tools. The categorization of taint primitives offers a structured approach to identifying the root causes of leakage.
  5. Shifting the Paradigm of Spectre Defense: Cerberus moves the industry away from reactive, piecemeal mitigations for individual speculation primitives towards a proactive, comprehensive, and compiler-driven defense. This holistic approach ensures that constant-time code is secured against a broad class of attacks, reducing the risk of undiscovered Spectre variants exploiting unmitigated speculation primitives. The low overhead makes this strong security practical for real-world deployments.

Key Takeaways

  • Comprehensive Spectre Defense: Cerberus is the first compile-time defense that comprehensively protects constant-time cryptographic code against all five documented speculation primitives on modern processors.
  • Hardware-Software Co-Design: It effectively leverages existing hardware features like Intel CET and specific speculation control bits to constrain transient control and data flow, making static analysis feasible and efficient.
  • Taint Primitives as Root Cause: The research formally introduces taint primitives (NCAL, NCAST, StackL, NAR) as the necessary and sufficient conditions for Spectre leakage in static constant-time programs, providing a precise target for mitigation.
  • Efficient LLVM Compiler Passes: Cerberus is implemented as three LLVM passes—Function Private Stacks, Register Cleaning, and Fence Insertion—which collectively eliminate all identified taint primitives.
  • Strong Security with Low Overhead: Despite offering stronger security guarantees than prior state-of-the-art solutions, Cerberus achieves an average performance overhead of only 21% (and 7% for larger inputs) on real-world cryptographic benchmarks.
  • Programmer-Transparent and Upstreamable: The defense is 100% programmer-transparent, requiring only a single compile-time flag, and is actively being worked towards upstreaming into LLVM, promising widespread availability and adoption.

About the Speaker(s)

Nicholas Mosier is a PhD student at Stanford University, advised by Caroline Trippel. He was the primary presenter of the Cerberus research at the IEEE S&P conference. His work focuses on system security, particularly in the realm of transient execution attacks and defenses.

Hamed Nemati, John C. Mitchell, and Caroline Trippel are co-authors of the Cerberus paper. Caroline Trippel is Nicholas Mosier's advisor at Stanford University, contributing her expertise to the research. Their collective work spans various aspects of computer security, architecture, and formal methods, bringing a robust academic foundation to the Cerberus project.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

Cerberus is a groundbreaking, comprehensive defense for constant-time cryptographic code against all five Spectre speculation primitives. By cleverly integrating hardware features like CET with novel compiler passes and a rigorous 'taint primitive' model, it offers robust protection with remarkably low overhead, making it a critical, practical solution.

Heather Calloway (CISO) — MUST SEE

This research introduces Cerberus, a groundbreaking compile-time defense that comprehensively protects constant-time cryptographic code against all known Spectre variants with remarkably low overhead. It provides a critical, programmer-transparent solution for a pervasive architectural vulnerability, offering a clear path for organizations to significantly enhance the security of their most sensitive operations.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024