ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer Space

Chuyang Chen (PhD student · Ohio State)

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Software Security 3: Fuzzing

Overview

This talk introduces ELFuzz, a novel evolutionary approach designed to efficiently generate high-quality seed test cases for mutation-based fuzzing. Presented by Chuyang Chen, a PhD student at Ohio State, the research is a collaborative effort with Professor Brandon Dolan Gabit from New York University and Chen's advisor at Ohio State. ELFuzz addresses a critical challenge in software security: the need for syntactically and semantically valid initial inputs that can effectively guide fuzzers to discover deep program logic and vulnerabilities.

Watch on YouTube · Slides

Visual summary for ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer Space by Chuyang Chen
Visual summary for ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer Space by Chuyang Chen

Key moments

  1. 0:00 Introduction: Problem with current seed generation
  2. 2:00 ELFuzz's novel LLM-driven evolutionary approach
  3. 3:00 Overview of the three-step evolution loop
  4. 4:50 Detailed explanation of LLM-based mutators
  5. 6:00 Mutant evaluation using code range lattice
  6. 8:00 Optimal survivor selection for next generation
  7. 9:50 Four key research questions for ELFuzz evaluation
  8. 10:59 Key findings: high coverage and new real-world bugs

ELFuzz: Efficient Input Generation via LLM-driven Synthesis Over Fuzzer Space

Speakers: Chuyang Chen, PhD Student, Ohio State

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=t2KwwMn-OXw

Overview

This talk introduces ELFuzz, a novel evolutionary approach designed to efficiently generate high-quality seed test cases for mutation-based fuzzing. Presented by Chuyang Chen, a PhD student at Ohio State, the research is a collaborative effort with Professor Brandon Dolan Gabit from New York University and Chen's advisor at Ohio State. ELFuzz addresses a critical challenge in software security: the need for syntactically and semantically valid initial inputs that can effectively guide fuzzers to discover deep program logic and vulnerabilities.

The significance of ELFuzz lies in its innovative use of Large Language Models (LLMs). Unlike prior methods that either rely on expensive manual expertise, limited white-box analysis, or employ LLMs as black-box generators, ELFuzz leverages an LLM to drive an evolution loop. This iterative refinement process transforms naive seed generators into sophisticated ones capable of producing diverse and high-coverage test cases. By doing so, ELFuzz aims to overcome the scalability, generalization, and extensibility limitations inherent in previous automated seed generation techniques, offering a powerful new tool for vulnerability discovery.

Background

▶ Watch: Introduction: Problem with current seed generation (0:00)

The effectiveness of mutation-based fuzzing, a cornerstone of software security testing, hinges significantly on the quality of its initial seed test cases. These seeds serve as the starting points from which subsequent mutations are derived. High-quality seeds are not merely syntactically valid; they must also be semantically meaningful, allowing mutations to bypass strict input validation logic and penetrate deeper into the program's execution paths, thereby increasing the likelihood of uncovering complex bugs.

Historically, curating such a corpus of high-quality seeds has been an expensive and expertise-intensive manual process. To overcome this, automated generation techniques have explored several avenues:

  1. White-box program analysis: These methods attempt to reverse-engineer input formats by analyzing program code. While precise, they often struggle with scalability and generalization across diverse programs.
  2. Manually embedded domain knowledge: Specialized generators are crafted with hardcoded rules reflecting the target input format. These are effective for specific targets but are difficult to generalize and extend to new programs or evolving formats.
  3. Black-box LLMs: More recently, LLMs have been explored as direct seed generators. While LLMs possess vast embedded domain knowledge from their training data, using them as black boxes for direct generation often fails to fully leverage this knowledge in a targeted, feedback-driven manner. Such approaches also typically result in opaque outputs that are hard to interpret or extend with traditional fuzzing techniques.

These limitations collectively highlight a gap in automated seed generation: a method that can scale, generalize, provide high-quality inputs, and remain interpretable and extensible. ELFuzz was developed to fill this gap by introducing an evolutionary paradigm that intelligently guides LLMs to synthesize superior seed generators, rather than just generating raw inputs.

Key Findings

▶ Watch: Overview of the three-step evolution loop (3:00)

ELFuzz's comprehensive evaluation addressed four key research questions, yielding significant findings across fuzzing effectiveness, bug discovery, component contribution, and interpretability:

  • Fuzzing Effectiveness: ELFuzz demonstrated superior performance in generating high-coverage test cases. Specifically, it achieved up to 34.8% higher code coverage compared to state-of-the-art baseline techniques. Furthermore, the test cases produced by ELFuzz-synthesized generators provided a substantial advantage when utilized as seeds for subsequent mutation-based fuzzing, indicating their quality and ability to unlock deeper program logic.
  • Bug Finding Capability: Beyond just coverage, ELFuzz proved to be a highly effective bug-finding tool. In controlled experiments with artificially injected bugs, ELFuzz consistently outperformed baseline techniques. More notably, when applied to real-world programs, ELFuzz successfully discovered five new, previously unknown bugs within the service5 application, all within a short 14-day time budget. These findings included severe and complex vulnerabilities, underscoring ELFuzz's practical utility in identifying critical security flaws.
  • Ablation Study: To understand the contribution of each component within ELFuzz, an ablation study was conducted. This analysis identified the fuzzer space guidance mechanism as the single most critical system component. Removing this component, which involves organizing mutants into a lattice structure based on covered code ranges, caused the most substantial drop in the coverage achieved by the synthesized generators. This highlights the importance of ELFuzz's nuanced comparison and selection strategy.
  • Interpretability and Extensibility: A significant finding was that the generators evolved by ELFuzz are highly interpretable. They are expressed as normal Python code, making their logic and the features they target human-understandable. This contrasts sharply with black-box LLM approaches. The extensibility of these generators was also confirmed; researchers successfully integrated another fuzzing technique into ELFuzz-synthesized generators in just five man-days. This demonstrates that ELFuzz produces artifacts that can be easily modified and integrated into existing security development processes.

Technical Deep Dive

▶ Watch: Mutant evaluation using code range lattice (6:00)

ELFuzz's core innovation lies in its evolutionary approach, where an LLM doesn't directly generate seeds but instead drives an iterative process to synthesize and refine seed generators. This paradigm leverages the LLM's domain knowledge in a feedback-driven loop, gradually improving the quality and diversity of generated inputs.

The evolution loop consists of three distinct steps:

  1. LLM Mutates Generators:

The process begins with naive seed generators, typically producing random text. The LLM's role is to iteratively mutate and refine these generators based on coverage feedback. ELFuzz employs three specific LLM-based mutators to achieve this:

  • Splicing: This mutator takes two existing candidate generators. It combines the beginning of one with the end of another and then prompts the LLM to generate "glue code" to seamlessly connect the two disparate parts. This allows for the creation of hybrid generators that inherit beneficial characteristics from multiple parents.
  • Completion: This mutator operates on a single candidate generator. It truncates the generator at a random point, effectively creating an incomplete code snippet. The LLM is then instructed to "recomplete" the generator, filling in the missing parts based on the context provided. This encourages the LLM to explore alternative implementations or extensions of existing logic.
  • Infilling: Here, a specific line of code is removed from a candidate generator, leaving a "blank." The LLM is then prompted to "fill in the blank," generating new code that fits the surrounding context. This mutator allows for targeted modifications and explorations of different logic branches within a generator.

The key advantage of these LLM-based mutators over traditional tree-based methods is the LLM's ability to produce modifications that are not only syntactically correct but also semantically rich and human-like. This enables the generators to evolve in a more meaningful and effective way, producing complex and valid inputs.

  1. Mutants Evaluated and Compared (Fuzzer Space Guidance):

After mutation, the newly generated candidate generators (mutants) are evaluated for their effectiveness. This evaluation is crucial for determining which mutants contribute positively to the overall fuzzing effort. ELFuzz introduces a novel concept: organizing these mutants into a lattice structure, termed the fuzzer space. This structure is based on the subset relationship of the specific code ranges they cover when their generated inputs are executed against the target program.

A critical distinction here is that ELFuzz does not merely rely on the total count of covered lines or control flow edges. Instead, it considers the exact ranges of code covered. This means two mutants covering the same number of lines but different parts of the code are not considered equivalent. Each explores a unique program aspect and is valuable for increasing overall diversity. For instance, if generator FB covers lines {L1, L2, L3, L4, L5} and FD covers {L6, L7, L8, L9, L10}, both cover five lines, but they are distinct and both contribute unique coverage.

During this step, any mutant that is syntactically incorrect or regressive (meaning its covered code ranges are a strict subset of an already existing candidate's coverage) is discarded. This ensures that the fuzzer space only retains valuable and non-redundant generators.

  1. Best Performing Selected (Survivors):

To keep the population size manageable for the next generation, ELFuzz must select a fixed number of "survivors." The selection criteria are designed to find a subset of mutants that, when taken together, maximize the total code coverage across the program. This is a more sophisticated approach than simply picking individual generators with the highest individual coverage scores, as it aims for collective optimality and diversity.

The problem of finding the perfect combination of generators to maximize total coverage is analogous to the classic set cover problem, which is known to be NP-hard. Therefore, ELFuzz employs an approximation algorithm to efficiently select a near-optimal set of survivors. This ensures that the most effective and diverse generators are propagated to the next iteration of the evolution loop, continually improving the overall quality of the seed corpus.

Implementation Details:

ELFuzz was implemented using a relatively small Llama 13B model. This choice demonstrates that powerful LLMs are not strictly necessary, and more accessible models can be effectively deployed. The Llama 13B model can run on a single GPU with more than 26 GB of VRAM, such as an Nvidia A40. A significant advantage of ELFuzz's evolutionary loop design is that it eliminates the need for complex prompt engineering. The prompts used to guide the LLM are described as consisting of "just a few sentences," making the system easier to develop and maintain.

Demo / Proof of Concept

▶ Watch: Optimal survivor selection for next generation (8:00)

While the talk transcript does not detail a live, interactive demonstration during the conference, the research provided compelling evidence of ELFuzz's capabilities through its evaluation. The most significant proof of concept was its application to a real-world program called service5.

In this practical scenario, ELFuzz was deployed with a limited budget of just 14 days. Within this timeframe, it successfully discovered five new, previously unknown bugs in service5. The speaker emphasized that these findings included "severe and complex vulnerabilities," highlighting the tool's ability to uncover non-trivial flaws that might evade traditional testing methods. This outcome serves as a strong demonstration of ELFuzz's practical utility and effectiveness as a bug-finding tool in real-world security assessments, going beyond theoretical coverage improvements to identify actionable security vulnerabilities.

Defensive Implications

▶ Watch: Key findings: high coverage and new real-world bugs (10:59)

The development of ELFuzz carries significant implications for security defenders and organizations engaged in robust software security testing:

  1. Enhanced Fuzzing Effectiveness: Defenders can leverage ELFuzz to significantly improve the quality of seed test cases for their internal fuzzing campaigns. By generating syntactically and semantically rich inputs, ELFuzz can help fuzzers reach deeper code paths and uncover vulnerabilities that might otherwise remain hidden due to strict input validation. This translates to more effective and efficient vulnerability discovery within development pipelines.
  1. Proactive Vulnerability Discovery: The ability of ELFuzz to find new, complex vulnerabilities in real-world software, as demonstrated with service5, suggests it can be a powerful tool for proactive security testing. Organizations should consider integrating such LLM-driven evolutionary fuzzing techniques into their continuous integration/continuous deployment (CI/CD) pipelines to catch critical flaws earlier in the development lifecycle.
  1. Interpretable and Extensible Security Tools: A key advantage of ELFuzz is that its synthesized generators are human-interpretable Python code. This means security teams can understand why certain inputs are generated and what specific program logic these inputs are designed to exercise. This interpretability aids in debugging, root cause analysis of discovered vulnerabilities, and allows security engineers to further refine or adapt the generators for specific testing needs or environments. Its proven extensibility also means it can be easily combined with existing fuzzing frameworks or custom test logic.
  1. Addressing Evolving Threat Landscape: As attackers increasingly leverage sophisticated techniques, including AI, to find vulnerabilities, defenders must also adopt advanced methods. ELFuzz represents a step forward in intelligent input generation, showcasing how LLMs can be harnessed not just for content creation but for driving complex search and optimization problems in security, thereby helping organizations stay ahead of emerging threats.
  1. Reinforcing Input Validation Best Practices: While ELFuzz helps find bugs, its success also implicitly reinforces the fundamental importance of robust input validation and sanitization. The fact that sophisticated input generation is required to bypass validation highlights the critical role of strong defensive programming practices to mitigate the impact of such advanced fuzzing techniques.

Key Takeaways

  • Evolutionary LLM-Driven Fuzzing: ELFuzz introduces a novel approach where an LLM drives an evolution loop to synthesize and refine high-quality seed generators, rather than acting as a black-box input generator.
  • Superior Input Quality: The method produces syntactically correct, semantically rich, and human-like inputs, significantly boosting fuzzing effectiveness by enabling deeper code coverage (up to 34.8% higher than baselines).
  • Fuzzer Space Guidance is Critical: ELFuzz's unique fuzzer space guidance, which uses a lattice structure based on specific code range coverage for comparing and selecting generators, is the most vital component for achieving diverse and high-quality seeds.
  • Real-World Bug Discovery: ELFuzz demonstrated practical utility by discovering five new, severe, and complex vulnerabilities in the real-world service5 application within a short timeframe.
  • Interpretable and Extensible Outputs: The synthesized generators are highly interpretable Python code and are easily extensible, allowing for integration with existing fuzzing techniques and easier analysis by security teams.
  • Leveraging LLMs for Intelligent Search: This research highlights a promising paradigm for integrating LLMs into security testing, showcasing their potential to drive intelligent search and optimization processes in vulnerability discovery.

About the Speaker(s)

Chuyang Chen is a PhD student at Ohio State University. His research focuses on developing automated and efficient methods for input generation, particularly leveraging Large Language Models, to enhance fuzzing processes and improve software security testing. This work, including ELFuzz, is a collaboration with Professor Brandon Dolan Gabit at New York University and his advisor at Ohio State, reflecting an inter-institutional effort to advance the state-of-the-art in automated vulnerability discovery.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

ELFuzz is legitimate research with a clean core idea: stop using LLMs as dumb input generators and instead use them to evolve the generators themselves, with coverage feedback closing the loop. The fuzzer space lattice — selecting survivors by collective code range coverage rather than raw line counts — is the genuinely clever bit, and the ablation study confirms it carries the system. Five real bugs in a real target keeps this grounded.

Heather Calloway (CISO) — WEAK

Technically credible research on LLM-driven fuzzing with real results — five new bugs, measurable coverage gains, interpretable outputs. But this is a PhD-level systems paper delivered to a security research audience, and it never crosses into governance, program accountability, or operational decision-making. There is nothing here for a CISO, a board, or a security leader trying to decide how to run a program.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)