Practical Data-Only Attack Generation

Brian Johannesmeyer

33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24

Overview

In an era where sophisticated defenses have rendered traditional control flow hijacking (CFH) attacks increasingly difficult, a new wave of research is shining a spotlight on data-only attacks (DOAs). This talk, presented by Brian Johannesmeyer, with co-work from Aza Herbert and Cristiano from V Amsterdam, introduces Einstein, a novel tool that automatically generates data-only exploits with surprising ease and effectiveness. Challenging the long-held perception that DOAs are either too application-specific or overly complex to pose a practical threat, Einstein demonstrates a scalable approach to uncover and weaponize these vulnerabilities.

Watch on YouTube

Visual summary for Practical Data-Only Attack Generation by Brian Johannesmeyer
Visual summary for Practical Data-Only Attack Generation by Brian Johannesmeyer

Key moments

  1. 0:00 Introduction: Practical data-only attack generation with Einstein
  2. 2:00 Classic data-only attack example: CGI-bin vulnerability
  3. 3:40 Why data-only attacks are considered difficult; Einstein's simplicity
  4. 5:00 Einstein's method: identifying vulnerable data with taint analysis
  5. 6:00 Einstein's method: overwriting data copied verbatim to syscalls
  6. 7:00 Einstein's evaluation: generating exploits for popular servers

Practical Data-Only Attack Generation

Speakers: Brian Johannesmeyer

Conference: USENIX Security '24

YouTube: https://www.youtube.com/watch?v=i8Qja60N268

Overview

In an era where sophisticated defenses have rendered traditional control flow hijacking (CFH) attacks increasingly difficult, a new wave of research is shining a spotlight on data-only attacks (DOAs). This talk, presented by Brian Johannesmeyer, with co-work from Aza Herbert and Cristiano from V Amsterdam, introduces Einstein, a novel tool that automatically generates data-only exploits with surprising ease and effectiveness. Challenging the long-held perception that DOAs are either too application-specific or overly complex to pose a practical threat, Einstein demonstrates a scalable approach to uncover and weaponize these vulnerabilities.

The research not only provides a powerful automated exploit generation framework but also delivers a stark message to both the security research community and software vendors: current mitigation strategies are often insufficient against this class of attack. By focusing on simple, yet potent, data-only attack vectors, Einstein highlights systemic weaknesses in how programs handle attacker-controllable data, particularly when it interacts with security-sensitive system calls. The findings necessitate a critical re-evaluation of current defense paradigms and a renewed focus on comprehensive, practical solutions for data integrity.

This work is significant because it democratizes the generation of data-only exploits, moving them from the realm of highly specialized, manual efforts to automated discovery. By demonstrating hundreds of successful exploits against popular web servers, Einstein underscores the pervasive nature of these vulnerabilities and the urgent need for robust, next-generation defenses that address data integrity beyond control flow protection. The talk serves as a clarion call for rethinking how we secure applications in a post-control-flow-hijacking world.

Background

▶ Watch: Introduction: Practical data-only attack generation with Einstein (0:00)

For decades, the primary focus of software exploitation research and defense has been control flow hijacking (CFH). Attackers traditionally sought to corrupt pointers that dictate program execution paths, such as return addresses or function pointers, to inject and execute arbitrary malicious code. This approach, exemplified by classic buffer overflows, allowed attackers to seize full control of a vulnerable system. However, the security community has responded with extensive research and deployed a suite of powerful defenses. Technologies like Control-Flow Integrity (CFI), Address Space Layout Randomization (ASLR), Data Execution Prevention (DEP), and stack cookies have made it exceedingly difficult, and often practically infeasible, to reliably corrupt control flow and execute attacker-controlled code.

In response to these hardened defenses, attackers have increasingly turned their attention to data-only attacks (DOAs). Unlike CFH attacks, DOAs do not alter the program's execution path. Instead, they manipulate critical application data structures to subvert the program's intended logic, causing it to perform malicious actions while still executing its legitimate code. While the concept of data-only attacks has existed for nearly two decades, they have often been dismissed as either requiring highly specific application knowledge or being too complex to implement in practice. Early examples, such as the 2004 e.g. CGI-BIN exploit, showcased their potential but were often viewed as outliers.

The CGI-BIN example, presented almost 20 years ago, vividly illustrates the core principle of a simple DOA. In a server application, a client might request the server to sort numbers via an HTTP POST, causing the server to execute an external program by concatenating a CGI_BIN path with a client-supplied script name (e.g., execve("/usr/local/bin/sort", "2 1 3", ...)). The exploit involved overwriting the CGI_BIN variable to /bin. Subsequently, when the attacker requested the server to execute the script /sh, the server, still executing its intended code path, would unwittingly call execve("/bin/sh", ...), thereby spawning an attacker-controlled shell. This attack did not corrupt any code pointers; it simply manipulated an argument passed to a security-sensitive system call, execve.

However, recent research into DOAs has often involved extremely complex methodologies. These approaches frequently employ heavyweight analyses like symbolic execution to reason about intricate data flow constraints, construct Turing-complete gadget sets, and simultaneously bypass control flow defenses. While groundbreaking, this complexity has inadvertently reinforced the perception that DOAs are inherently difficult to generate. Brian Johannesmeyer argues that this complexity can be a "self-inflicted pain" by the research community, suggesting that simpler, more direct data-only attack vectors are often overlooked. Einstein's work aims to revisit and automate the generation of these simpler, yet highly effective, data-only exploits, specifically targeting scenarios where a program executes its intended code but passes attacker-controlled, verbatim data to security-sensitive system calls.

Key Findings

▶ Watch: Why data-only attacks are considered difficult; Einstein's simplicity (3:40)

The Einstein project delivers several critical findings that reshape our understanding of data-only attacks:

  • Automated Exploit Generation: Einstein successfully automates the generation of data-only attacks, proving that these exploits are not exclusively the domain of highly specialized manual analysis. This automation significantly lowers the barrier to entry for uncovering and exploiting such vulnerabilities.
  • Widespread Vulnerability in Popular Servers: The tool was able to build hundreds of exploits against popular web server applications. This demonstrates that the problem is not isolated to niche applications but is prevalent across widely used software. The choice of servers for evaluation was motivated by their historical susceptibility to memory safety bugs, which provide the initial primitive for data corruption.
  • Simplicity Over Complexity: Contrary to the prevailing notion that data-only attacks are inherently complex, Einstein focuses on simple attack patterns where attacker-controlled data is copied verbatim into arguments of security-sensitive system calls. This "simple but not too simple" approach proves highly effective and practical.
  • Verbatim Data Copying is a Key Enabler: A crucial finding is the high percentage of security-sensitive system call arguments that receive data copied verbatim from attacker-corruptible sources. For instance, in the execve system call, 86% of attacker-tainted arguments were found to be verbatim copies. This pattern is particularly common for string arguments, as programs frequently copy strings without modification after initialization, creating prime targets for data manipulation.
  • Diverse Attacker Primitives: The research shows that various system calls, when manipulated through data-only attacks, offer a wide array of powerful attacker primitives. These include not only arbitrary code execution via execve but also file system corruption via write and network packet manipulation via sendmsg. This broad spectrum of capabilities underscores the severe impact of DOAs.
  • Bypassing Practical Defenses: Einstein's generated exploits were shown to bypass many existing "practical" defenses designed to mitigate memory safety issues. This highlights a significant gap in current security architectures, where defenses often prioritize control flow integrity while leaving data integrity vulnerable to these specific attack types.
  • Call for Rethinking Mitigation Strategies: The findings strongly urge both researchers and vendors to reconsider their mitigation strategies. The current landscape often presents a dilemma between "comprehensive" (but impractical) and "practical" (but non-comprehensive) defenses for data-only attacks. Einstein's success against practical defenses necessitates a shift towards deploying more comprehensive solutions and investing in research to make these comprehensive defenses more practical.

Technical Deep Dive

▶ Watch: Einstein's method: identifying vulnerable data with taint analysis (5:00)

Einstein's ability to automatically generate data-only attacks hinges on answering two fundamental questions: first, which data to overwrite, and second, what to overwrite it with. The tool's elegance lies in its straightforward yet powerful approach to systematically addressing these questions.

1. Identifying Which Data to Overwrite: Dynamic Taint Analysis

To determine which specific data variables are critical targets for manipulation, Einstein employs dynamic taint analysis. This technique involves instrumenting the victim program to track the flow of data at runtime.

  • Taint Source Identification: The first step is to identify all potentially attacker-controllable data. This includes input received from network sockets, file reads from untrusted sources, environment variables, and other user-supplied data. Each piece of such data is marked with a unique taint label.
  • Workload Execution: The victim program (e.g., a web server) is then run with a representative workload. For server applications, this typically involves sending various HTTP requests derived from the application's test suite. As the server processes these requests, the taint propagates through the program's memory and registers.
  • Security-Sensitive Syscall Monitoring: Einstein monitors calls to a predefined set of security-sensitive system calls. These are syscalls that, if supplied with malicious arguments, could lead to system compromise (e.g., execve, write, sendmsg, open, mmap).
  • Taint Propagation to Syscall Arguments: When a security-sensitive syscall is invoked, Einstein checks if any of its arguments are tainted. If an argument carries a taint label, it indicates that this argument's value originated, at least partially, from attacker-controlled input.
  • Source Tracing: By tracing the taint label backward through the program's data flow, Einstein can identify the specific memory location or variable (e.g., the CGI_BIN variable in the classic example) that holds the attacker-controllable data ultimately passed to the vulnerable syscall argument. This pinpoints the exact data region that an attacker would need to corrupt.

2. Determining What to Overwrite It With: Verbatim Copying Exploitation

Once a target data variable is identified, the next challenge is to determine the malicious payload – what to overwrite it with. Einstein's core insight here is to leverage instances where data is copied verbatim from attacker-corruptible sources to security-sensitive system call arguments.

  • Verbatim Copying Hypothesis: The premise is that if a program copies attacker-tainted data directly and without modification into a syscall argument, then overwriting that original data location with a malicious string has a high probability of that malicious string being passed directly to the syscall.
  • High Likelihood of Success: Programs, especially when dealing with strings, frequently copy them around verbatim after initialization. This behavior is common and often necessary for performance or memory management. Einstein capitalizes on this by assuming that if an attacker can overwrite the source of such a verbatim copy, the destination (the syscall argument) will reflect the attacker's chosen value.
  • Exploit Generation and Validation: Based on the identified vulnerable variable and the verbatim copying pattern, Einstein generates a candidate exploit. This exploit specifies the target memory location and the malicious string to write. To confirm its efficacy, the generated exploit is then "carried out" on the program. This involves simulating a memory corruption event (e.g., via a fuzzer or a controlled debugger) that overwrites the target variable with the malicious payload. The program is then executed, and its behavior is observed to confirm the intended malicious outcome (e.g., a shell being spawned, a file being corrupted). This validation step is crucial because runtime checks might still prevent the exploit from fully succeeding, even if the data flow is verbatim.

Evaluation and Insights:

Einstein was evaluated against several popular server applications, chosen due to their historical susceptibility to memory safety bugs. The analysis was driven by each application's test suite, which provided a realistic workload of HTTP requests.

The results were striking:

  • execve Syscall: Out of 11 observed call sites for execve, 7 had arguments that were attacker-tainted. Crucially, at 86% of these 7 call sites, the attacker-tainted data was copied verbatim from an attacker-controllable source. This high percentage underscores the prevalence of this vulnerability pattern for string-based arguments.
  • Generalizability to String Arguments: The pattern of verbatim copying from attacker-corruptible data was not unique to execve. It was observed consistently across other system calls that accept string-type arguments. This suggests a fundamental design or programming pattern that makes many syscall interfaces vulnerable to DOAs when an initial memory corruption primitive exists.
  • Diverse Attacker Primitives: The variety of syscalls found to be vulnerable offers attackers a rich palette of primitives:
  • execve: Directly enables arbitrary code execution, allowing an attacker to run any command on the server, such as spawning a shell (/bin/sh).
  • write: Can be exploited to corrupt or overwrite arbitrary files on the server's file system, leading to data integrity breaches or denial of service.
  • sendmsg: Allows an attacker to craft and send arbitrary network packets, potentially facilitating further network attacks, exfiltration of data, or command and control.

While the talk highlights these core principles, the authors note that their full paper delves into more advanced topics, including how multiple syscalls can be chained together to achieve more complex attack goals and how Einstein's exploits can bypass certain state-of-the-art defenses. The technical foundation of Einstein demonstrates that a focused, dynamic analysis approach can effectively uncover and weaponize a significant class of data-only vulnerabilities that are often overlooked due to a perceived complexity.

Demo / Proof of Concept

▶ Watch: Einstein's method: overwriting data copied verbatim to syscalls (6:00)

While the talk does not feature a live demonstration of the Einstein tool itself, it effectively illustrates the core concept of a data-only attack through a classic, well-understood example: the CGI-BIN exploit from nearly two decades ago. This conceptual proof-of-concept serves to clarify the mechanics of how Einstein identifies and exploits such vulnerabilities.

The scenario unfolds as follows:

  1. Initial Vulnerability: The server application has a CGI_BIN variable, typically configured to a path like /usr/local/bin, which it uses to locate and execute external programs requested by clients.
  2. Attacker's Goal: The attacker's objective is to execute an arbitrary shell command on the server.
  3. Data Overwrite: The attacker first leverages a hypothetical memory safety bug (which Einstein assumes as a prerequisite for its analysis) to overwrite the server's CGI_BIN variable. Instead of its benign value, the attacker corrupts it to the string "/bin".
  4. Malicious Request: The attacker then sends an HTTP request to the server, asking it to execute a script named "/sh".
  5. Server's Action (Unwitting Execution): The server, still operating under its legitimate control flow, attempts to fulfill the request. It concatenates the (now corrupted) CGI_BIN path with the client-supplied script name: "/bin" + "/sh" becomes "/bin/sh".
  6. Syscall Invocation: The server then invokes the execve system call with "/bin/sh" as the program to execute.
  7. Exploit Success: As a result, the server happily executes /bin/sh, effectively spawning an attacker-controlled shell. The example given is that this shell might simply create a file /tmp/attacker_was_here, but any shell command would be possible.

The crucial aspect of this demonstration is that the server's control flow was never hijacked. It executed precisely the code it was designed to execute: concatenating paths and calling execve. The attack solely relied on corrupting the data (the CGI_BIN variable) that influenced the arguments passed to a security-sensitive system call.

Einstein's process mirrors this example: it uses dynamic taint analysis to identify that the CGI_BIN variable's content is attacker-controllable and that it's copied verbatim into an argument for execve. Then, it generates an exploit that suggests overwriting CGI_BIN with a malicious path like "/bin", knowing that a subsequent request for "/sh" will lead to arbitrary code execution. The talk confirms that Einstein carries out its generated exploits on the program to validate their functionality, ensuring that they indeed spawn a shell or achieve the intended malicious outcome. This validation step is key to confirming the practical viability of the generated attacks.

Defensive Implications

▶ Watch: Einstein's evaluation: generating exploits for popular servers (7:00)

The findings from the Einstein project present a significant challenge to current software security practices and highlight critical shortcomings in prevailing mitigation strategies against memory corruption vulnerabilities. The traditional focus on control flow integrity, while successful in mitigating CFH attacks, has inadvertently left a vast attack surface open for data-only manipulations.

The core dilemma for defenders, as articulated in the talk, lies in the trade-off between comprehensiveness and practicality when designing mitigations for data-only attacks.

  • Control Flow Hijacking (Historical Context): In the "days of old," control flow hijacking attacks were relatively well-defined. They typically involved corrupting a code pointer (e.g., a return address, function pointer) which would then corrupt an indirect branch or call. This clear attack vector allowed for the development of mitigations that could be both comprehensive (covering most CFH scenarios) and practical (having acceptable performance overhead and deployment complexity). Examples include ASLR, DEP, and CFI.
  • Data-Only Attacks (Current Challenge): Data-only attacks are fundamentally different and far more elusive. They can overwrite any piece of data in a program, and this corrupted data can then influence any subsequent operation, not just control flow. This boundless nature makes comprehensive mitigation incredibly difficult.
  • Comprehensive Defenses: A truly comprehensive defense against DOAs would need to protect the integrity of all data in a program and validate all operations that use that data. Such defenses typically involve heavy instrumentation, pervasive runtime checks, or formal verification. While theoretically sound, these approaches often incur prohibitive performance overheads, are complex to implement, or require significant application-specific modifications, rendering them impractical for widespread deployment.
  • Practical Defenses: Conversely, many existing "practical" defenses are designed to be lightweight and easy to deploy. However, these often focus on specific, known attack patterns or rely on heuristics. As Einstein demonstrates, these practical defenses are frequently non-comprehensive and can be bypassed by novel or carefully crafted data-only exploits that leverage subtle data flow patterns like verbatim copying.

Einstein's success in generating hundreds of exploits against popular web servers and its ability to bypass state-of-the-art practical mitigations underscore the severity of this problem. It's a clear indication that current "good enough" defenses are no longer sufficient.

Recommendations for Defenders:

  1. Prioritize Comprehensive Defenses: The talk strongly urges vendors and developers to consider deploying one or more of the more comprehensive defenses, even if they come with higher overheads or deployment challenges. This might include:
  • Fine-grained Data Taint Tracking: More pervasive and robust taint tracking mechanisms that can identify and prevent the use of tainted data in security-sensitive operations, even if it's copied multiple times.
  • Data Integrity Protection (DIP): Research into hardware-assisted or software-based mechanisms that ensure the integrity of critical data structures, detecting unauthorized modifications.
  • Application-Specific Semantic Checks: Implementing rigorous input validation and output sanitization, not just at the entry points but throughout the data's lifecycle, especially before it's used in sensitive operations or passed to system calls.
  • Memory Tagging/Protection: Leveraging emerging hardware features (e.g., ARM MTE) that allow for fine-grained memory protection and integrity checks, although these are still in early adoption phases.
  1. Invest in Research for Practicality: The security research community is called upon to address the pressing need for making comprehensive defenses more practical. This involves:
  • Developing techniques to reduce the performance overhead of comprehensive data integrity checks.
  • Creating automated tools that can assist developers in identifying critical data and implementing semantic checks.
  • Exploring novel architectural designs that inherently protect data integrity without sacrificing performance.
  • Investigating compiler-based or runtime techniques that can enforce data invariants more efficiently.

In essence, the talk argues that the security industry is at an inflection point. Just as CFH attacks necessitated a paradigm shift towards control flow integrity, data-only attacks demand a similar re-evaluation and investment in data integrity. Simply patching vulnerabilities on a case-by-case basis is unsustainable. A systemic shift towards more robust, comprehensive, and eventually practical data integrity defenses is paramount to securing modern applications against this evolving threat landscape.

Key Takeaways

  • Data-Only Attacks are a Practical Threat: Contrary to previous perceptions, data-only attacks are not merely theoretical or overly complex; Einstein demonstrates their practical viability and ease of generation against popular software.
  • Automation is Key: The Einstein tool automates the process of finding and generating data-only exploits, significantly lowering the barrier to entry for attackers and highlighting the widespread nature of these vulnerabilities.
  • Verbatim Data Copying is a Critical Exploit Vector: A major enabler for these attacks is the common programming pattern of copying attacker-tainted data verbatim into arguments of security-sensitive system calls, particularly for string types.
  • Diverse Attacker Primitives: Data-only attacks can provide powerful and varied attacker capabilities, including arbitrary code execution (execve), file system corruption (write), and network manipulation (sendmsg).
  • Current Defenses are Insufficient: Many existing "practical" security mitigations, especially those focused primarily on control flow integrity, are ineffective against the data-only attacks generated by Einstein.
  • Rethink Mitigation Strategies: There is an urgent need for both vendors and researchers to re-evaluate and enhance mitigation strategies, moving towards more comprehensive data integrity defenses and investing in making these robust solutions more practical for widespread deployment.

About the Speaker(s)

The talk "Practical Data-Only Attack Generation" was presented by Brian Johannesmeyer. He delivered this research as a co-work effort with Aza Herbert and Cristiano from V Amsterdam. Brian Johannesmeyer's presentation reflects a deep understanding of software exploitation and a commitment to advancing the field of security research, particularly concerning the practical implications of data-only attacks in modern computing environments. While specific titles or affiliations for Aza Herbert and Cristiano were not detailed in the provided metadata or transcript, their collaboration with Brian Johannesmeyer underscores a collective expertise in vulnerability research and systems security.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

Dr. Johannesmeyer's "Einstein" tool shatters the myth that data-only attacks are too complex to be practical, demonstrating automated generation of hundreds of exploits against popular servers. By exposing widespread vulnerabilities rooted in verbatim data copying to sensitive syscalls, this research is a critical wake-up call for the industry to re-evaluate its defense strategies beyond control flow integrity. It's a must-see for anyone serious about real-world application security.

Heather Calloway (CISO) — STRONG ACCEPT

This research on automated data-only exploit generation is a critical wake-up call for security leaders and vendors. It clearly demonstrates systemic weaknesses in data integrity defenses, revealing that current control-flow-focused mitigations leave organizations highly exposed. The implications demand a strategic re-evaluation of defense paradigms and investment in comprehensive data integrity solutions.

→ Top-rated talks at 33rd USENIX Security Symposium

All talks from 33rd USENIX Security Symposium