NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines

Mingxuan Liu

34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Fraud, Malware, Spam

Overview

Software vulnerabilities pose an incessant and critical threat to digital security, with their volume and sophistication growing rapidly. Timely and effective patching is paramount, yet manual approaches are costly, slow, and struggle to keep pace with newly discovered flaws, especially zero-day vulnerabilities. While automated patching solutions exist, they often suffer from significant limitations: traditional code analysis methods typically require exploit evidence or compilable code, which may not be available, and deep learning (DL) based approaches lack generalizability due to insufficient training data. Large Language Models (LLMs) have emerged as powerful tools for code-related tasks, showing promise in vulnerability patching, but prior LLM-based methods often fall short on real-world vulnerabilities due to a lack of guided reasoning, context limitations, and susceptibility to hallucination.

Read the paper · Download the PDF (PDF) · Slides

Paper abstract

Timely and effective vulnerability patching is essential for cybersecurity defense, for which various approaches have been proposed yet still struggle to generate valid and correct patches for real-world vulnerabilities. In this paper, we leverage the power and merits of pre-trained language language models (LLMs) to enable automated vulnerability patching using no test input/exploit evidence and without model training/fine-tuning. To elicit LLMs to effectively reason about vulnerable code behaviors, which is essential for quality patch generation, we introduce vulnerability semantics reasoning and adaptive prompting on LLMs and instantiate the methodology as APPATCH, an automated LLM-based patching system. Our evaluation of APPATCH on 97 zero-day vulnerabilities and 20 existing vulnerabilities demonstrates its superior performance to both existing prompting methods and state-of-the-art non-LLM-based techniques (by up to 28.33% in F1 and 182.26% in recall over the best baseline). Through APPATCH, we demonstrate what helps for LLM-based patching and how, as well as discussing what still lacks and why.

Visual summary for NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines by Mingxuan Liu
Visual summary for NOKEScam: Understanding and Rectifying Non-Sense Keywords Spear Scam in Search Engines by Mingxuan Liu

APPATCH: Automated Adaptive Prompting Large Language Models for Real-World Software Vulnerability Patching

Speakers: Yu Nong (University at Buffalo), Haoran Yang (Washington State University), Long Cheng (Clemson University), Hongxin Hu (University at Buffalo), Haipeng Cai (University at Buffalo)

Conference: USENIX Security

YouTube: N/A (This article is based on a peer-reviewed conference paper, not a recorded talk.)

Overview

Software vulnerabilities pose an incessant and critical threat to digital security, with their volume and sophistication growing rapidly. Timely and effective patching is paramount, yet manual approaches are costly, slow, and struggle to keep pace with newly discovered flaws, especially zero-day vulnerabilities. While automated patching solutions exist, they often suffer from significant limitations: traditional code analysis methods typically require exploit evidence or compilable code, which may not be available, and deep learning (DL) based approaches lack generalizability due to insufficient training data. Large Language Models (LLMs) have emerged as powerful tools for code-related tasks, showing promise in vulnerability patching, but prior LLM-based methods often fall short on real-world vulnerabilities due to a lack of guided reasoning, context limitations, and susceptibility to hallucination.

This paper introduces APPATCH, a novel, automated methodology that leverages pre-trained LLMs for vulnerability patching without requiring test inputs, exploit evidence, model training, or fine-tuning. APPATCH addresses the inherent challenges of applying LLMs to this complex task by introducing vulnerability semantics reasoning and adaptive prompting. The system employs a multi-stage approach including semantics-aware scoping to focus LLMs on critical code, dynamic adaptive prompting to guide reasoning and select relevant examples on the fly, and multi-faceted patch validation to ensure the quality and correctness of generated patches.

Evaluated on 97 zero-day vulnerabilities and 20 existing real-world vulnerabilities, APPATCH demonstrates superior performance over both existing prompting methods and state-of-the-art non-LLM-based techniques. It achieves significant improvements in F1 score (up to 28.33% higher) and recall (up to 182.26% higher) over the best baselines. This work not only showcases the practical feasibility of LLM-based vulnerability patching but also provides crucial insights into how LLMs can be effectively guided to perform complex code analysis and repair tasks, highlighting both their strengths and remaining limitations.

Background

The landscape of software security is perpetually challenged by the relentless emergence of new vulnerabilities. Manual patching processes are economically burdensome and inherently reactive, often failing to address vulnerabilities before they are actively exploited. This pressing need has spurred extensive research into automated patching solutions, which broadly fall into two categories: code-analysis-based and deep learning (DL)-based approaches.

Code-analysis-based techniques often focus on property-based repair or use sanitizers to extract vulnerability-representing constraints for symbolic execution. While some, like those employing inductive property inference, have shown promise, they predominantly rely on the availability of exploits or vulnerability-triggering test inputs. This dependency significantly limits their applicability, particularly for emerging zero-day vulnerabilities where such evidence is scarce or non-existent. Furthermore, these techniques frequently demand that the code under analysis be compilable, which is not always feasible during early development stages or for incomplete projects.

Deep learning (DL)-based approaches attempt to overcome these limitations by training models on existing vulnerability datasets to learn common fixing patterns. While conceptually appealing for their ability to generalize, these methods are heavily reliant on sizable and high-quality labeled training data. Such comprehensive datasets are not widely available, leading to models that often struggle to generalize effectively to unseen code and real-world vulnerabilities. Despite efforts in data augmentation, DL models frequently exhibit low accuracy in practical scenarios.

The recent advent of Large Language Models (LLMs) has opened new avenues for automated vulnerability patching. LLMs possess a remarkable capacity for understanding and generating human-like text, which extends to code. However, simply applying LLMs to patching tasks presents its own set of challenges. Prior LLM-based techniques, such as standard prompting (directly asking the LLM to patch) or zero-shot code completion (removing vulnerable code and letting the LLM fill it in), have shown limited effectiveness on real-world vulnerabilities. As illustrated in Figure 2 and Figure 3 of the paper, these naive approaches fail to perform the necessary in-depth code semantics reasoning, often generating invalid or superficial patches. For instance, in a CWE-787 (out-of-bounds write) example, GPT-4 with standard prompting recognized the vulnerability but failed to correctly determine the required memory boundary without detailed data/control dependency analysis. Zero-shot completion similarly provided an inadequate fix.

The core challenges for effective LLM-based vulnerability patching, as identified by the authors, are fourfold:

  1. Automated Prompting (Challenge 1): How to automatically construct effective prompts for vulnerability patching without manual intervention.
  2. Exemplar Selection (Challenge 2): How to dynamically select the most relevant and effective exemplars to guide LLM reasoning for a given vulnerability, especially when patching relies on complex code semantics reasoning.
  3. Context Constraints (Challenge 3): LLMs have token limitations and struggle with large code contexts. Vulnerabilities, however, are often context-sensitive and interprocedural, requiring extensive code analysis.
  4. Hallucination and Non-Determinism (Challenge 4): LLMs are prone to generating incorrect or inconsistent responses, which is unacceptable for security-critical tasks like patching where bad patches can be worse than no patches.

APPATCH is specifically designed to address these challenges by introducing a structured, automated, and adaptive methodology that guides LLMs through the complex process of vulnerability analysis and patch generation.

Key Findings

APPATCH delivers a significant leap forward in automated vulnerability patching, demonstrating superior performance across multiple metrics and challenging datasets. The key findings from its evaluation highlight its practical potential and the effectiveness of its design principles:

  • Superior Performance on Real-World Vulnerabilities: APPATCH consistently outperforms both existing LLM prompting strategies and state-of-the-art non-LLM-based techniques.
  • On the challenging Zero-Day dataset (97 samples), APPATCH achieved up to 36.46% F1 and 49.48% recall (with Claude-3.5). For comparison, the same LLMs with baseline prompting strategies only achieved up to 28.41% F1, and zero-shot completion was as low as 8.76% F1.
  • On the ExtractFix dataset (20 samples), APPATCH achieved even higher effectiveness, with up to 73.86% F1 (Claude-3.5) and 90.00% recall (GPT-4). Baseline LLM prompting methods lagged significantly, with the best achieving 68.54% F1 and zero-shot completion showing 6.27% F1.
  • Outperformance Against Non-LLM Baselines: When compared to leading non-LLM techniques, APPATCH demonstrated substantial advantages:
  • On the ExtractFix dataset, APPATCH (GPT-4) achieved 90% recall, surpassing ExtractFix (85%), VulnFix (45%), VulRepair (35%), and Getafix (25%).
  • On the Zero-Day dataset, APPATCH (Claude-3.5) achieved 49.48% recall, significantly higher than VulRepair (17.53%) and Getafix (6.19%). This highlights APPATCH's ability to patch without requiring compilable code or test cases, a common limitation for other techniques.
  • Effective Handling of Interprocedural Vulnerabilities: APPATCH maintains robust performance even on complex interprocedural vulnerabilities. On the 21 interprocedural samples from the Zero-Day dataset, APPATCH showed only a modest decrease in F1 score (e.g., GPT-4 dropped from 33.30% to 31.37%). In stark contrast, baseline LLM approaches exhibited dramatic performance declines (e.g., Claude-3.5 with standard prompting dropped from 23.01% to 9.81% F1), underscoring the critical role of APPATCH's interprocedural analysis capabilities.
  • Significant Contribution of Each Design Component: Ablation studies confirmed that every core component of APPATCH contributes substantially to its overall effectiveness:
  • Multi-faceted patch validation improved precision and F1 scores by filtering out incorrect patches.
  • Semantics-aware scoping was crucial, with its absence leading to significant drops in recall, precision, and F1 (e.g., GPT-4 F1 dropped from 33.30% to 25.94% on Zero-Day), demonstrating its role in focusing LLMs.
  • Dynamic adaptive prompting proved essential, as randomly selected exemplars led to dramatically lower F1 scores compared to APPATCH's adaptive selection.
  • LLM-generated exemplars outperformed manually written ones, indicating the effectiveness of LLMs in producing high-quality reasoning demonstrations.
  • Vulnerability semantics reasoning guided LLMs to higher accuracy in root cause analysis, leading to better patching outcomes compared to direct reasoning.
  • Reasonable Efficiency: APPATCH is efficient, taking an average of 37.148 to 50.209 seconds to generate a patch per sample. It utilizes 5,684 to 6,802 tokens for context and generates 584 to 886 tokens per patch, demonstrating practical feasibility in terms of both time and API costs.
  • Practical Usability with Upstream Tools: When integrated end-to-end with CodeQL for vulnerability detection and classification, APPATCH demonstrated reasonable performance, even with inaccuracies from upstream tools. In a "realistic" scenario with human inspection and calibration of CodeQL outputs, APPATCH's F1 score with Claude-3.5 reached 27.29%, outperforming baseline LLM and non-LLM approaches, highlighting its potential for real-world deployment in existing security workflows.

These findings collectively establish APPATCH as a highly effective and practical framework for automated vulnerability patching, pushing the boundaries of what LLMs can achieve in software security.

Technical Deep Dive

APPATCH is meticulously designed to guide LLMs through complex vulnerability analysis and patch generation, operating in two distinct phases: an offline Exemplar Mining phase and an online LLM-Guided Causal Patching phase. The system assumes the availability of vulnerability manifestation locations (vulnerable statements) and their corresponding CWE IDs.

Phase 1: Exemplar Mining

This offline phase prepares a rich database of exemplars that LLMs can later use for few-shot learning during patch generation.

Semantics-Aware Scoping (Step 1.1)

The first critical step for both phases is Semantics-Aware Scoping. Given a vulnerable program, APPATCH narrows down the vast codebase to only the most relevant subset, known as vulnerability semantics. This addresses the LLM context window limitation (Challenge 3) and reduces distractions from irrelevant code.

Vulnerability semantics is formally defined as the union of all static backward slices of the program, starting from any of the vulnerable locations ($S_v$) and terminating at any of the identified external inputs ($EI$). These external inputs include program input variables, returns from external function calls (e.g., malloc, socket_recv, scanf), and global states.

The process, detailed in Algorithm 1, involves:

  1. Constructing a System Dependence Graph (SDG): APPATCH uses Joern [53], a powerful code analysis platform, to build an SDG for the input program. An SDG combines both data dependencies and control dependencies, crucial for comprehensive vulnerability analysis. Joern's ability to analyze C code without requiring compilation is particularly beneficial for uncompilable zero-day samples.
  2. Identifying External Inputs (EI): A static analysis identifies calls to known external functions that provide input or modify program states, as well as program entry points defining input variables.
  3. Interprocedural Backward Slicing: From each vulnerable statement, APPATCH performs an interprocedural backward slice along the SDG until an external input is reached. This captures all statements that influence the vulnerable behavior from an external source.
  4. Union Slicing: All individual backward slices are unionized to form the complete vulnerability semantics, represented as a concise code snippet. As exemplified in Figure 1, this process effectively highlights the data and control flow relevant to the vulnerability, making it easier for LLMs to reason about. For instance, a variable cnt defined at line 16 and used at line 22 might appear much closer in the sliced code, simplifying dependency analysis.

Exemplar Generation (Step 1.2)

With the vulnerability semantics extracted, APPATCH automatically generates exemplars from existing vulnerable samples with known ground-truth patches. These exemplars serve as demonstrations for the LLMs, guiding them through the complex reasoning process (addressing Challenge 2). The reasoning process is broken down into three high-level steps:

  1. Vulnerability Semantics Reasoning (Root Cause Analysis): The LLM is prompted to analyze the vulnerable behavior step-by-step, starting from the external inputs and following the vulnerability semantics slice. The template used is: "Q: Given the following code slice Vexem, which has a vulnerability among <CWE-IDs> and lines Sv, the patch is <ground-truth patch>. Starting with the external inputs: <EI identified>, reason about the vulnerable behavior step by step until the vulnerability is determined." Providing the ground-truth patch helps the LLM generate accurate reasoning steps. This vulnerability-semantics-guided reasoning significantly improves the quality of root cause analysis compared to unguided prompting (Figure 5 vs. Figure 9).
  2. Fixing Strategy Identification: Based on the root cause, the LLM identifies a suitable strategy to mitigate the vulnerability.
  3. Patch Generation: The LLM generates the ground-truth patch itself.

The key insight here is that LLMs, when provided with ground-truth patches, are capable of generating correct and comprehensive reasoning steps (92.98% accuracy in preliminary experiments with GPT-4). This automated generation process builds a high-quality exemplar pool without time-consuming manual effort.

Phase 2: LLM-Guided Causal Patching

This is the online phase where APPATCH generates patches for a new, unseen vulnerable program.

Semantics-Aware Scoping (Step 2.1)

Similar to Step 1.1, the target vulnerable program undergoes semantics-aware scoping to extract its vulnerability semantics. This ensures the LLM receives a focused and relevant code context.

Dynamic Adaptive Prompting (Step 2.2)

This is the core of APPATCH's online operation, employing a progressive prompting strategy (Algorithm 2) that addresses Challenges 1 and 2 (automated prompting and exemplar selection).

  1. Root Cause Generation:
  • APPATCH first prompts an LLM to generate the root cause analysis for the current vulnerable slice. This process is iterative: it starts with the slice of the function containing the vulnerable statements.
  • If the LLM encounters uncertainty due to missing function definitions, it explicitly requests the names of needed functions (e.g., {"context_funcs":[func_1,func_2,CALLER_of_func...]}).
  • APPATCH then expands the code slice to include the requested function definitions and prompts the LLM again. This iterative expansion, similar to LLift [28], enables interprocedural vulnerability analysis while respecting token limits and avoiding distractions from irrelevant code. Figure 5 demonstrates this iterative process for the example in Figure 1.
  1. Exemplar Selection:
  • After generating the root cause for the testing sample, APPATCH dynamically selects the most relevant exemplars from the pre-mined pool.
  • An LLM is prompted to compare the root cause analysis of the testing sample with that of each exemplar sample, simply answering "yes" or "no" to similarity (Figure 6). This efficient comparison process ensures that only exemplars with similar underlying vulnerability mechanisms are chosen.
  • Up to 8 similar exemplars are selected for few-shot learning, following the CoT methodology [51].
  1. Patch Generation:
  • With the selected exemplars and the testing sample's vulnerability semantics, APPATCH constructs a comprehensive prompt. This prompt includes the vulnerability semantics, CWE IDs, vulnerable statements, and the root cause analysis generated in the previous step.
  • The prompt also incorporates the selected exemplars, each detailing its root cause, fixing strategy, and ground-truth patch. This guides the LLM through a complete, similar workflow.
  • The LLM is then instructed to generate up to five possible candidate patches for the vulnerability. Generating multiple patches improves recall and offers developers more options (Figure 7).

Multi-Faceted Patch Validation (Step 2.3)

To mitigate LLM hallucinations and non-deterministic responses (Challenge 4), APPATCH employs a multi-faceted patch validation strategy using an ensemble of LLMs (Algorithm 3).

  1. Ensemble Validation: Each candidate patch is independently evaluated by multiple powerful LLMs (GPT-4, Gemini-1.5, Claude-3.5, Llama-3.1).
  2. Dual Criteria: Each LLM validates whether the patch not only fixes the vulnerability but also preserves the original code functionality.
  3. Recall-Oriented Retention: A candidate patch is retained if any of the validating LLMs approve it. This approach prioritizes recall over precision, providing developers with a wider selection of potentially correct patches, which is more valuable in real-world patching scenarios than a high-precision but low-recall approach like majority voting. Figure 8 illustrates this validation process.

Implementation Details

APPATCH leverages Joern [53] for constructing the SDG and performing interprocedural dependence analysis on C code samples. This choice is critical as Joern does not require the input code to be compilable, making it suitable for analyzing real-world vulnerabilities, including zero-days. The extracted vulnerability semantics slices are stored as source code text for LLM input. The study utilizes general-purpose LLMs (GPT-4, Gemini-1.5, Claude-3.5, Llama-3.1) over code-specific ones (e.g., CodeLlama) due to their superior logical analysis capabilities required for vulnerability reasoning.

Demo / Proof of Concept

While APPATCH is a framework instantiated from a research paper rather than a live demonstration, its efficacy is rigorously proven through extensive evaluation against challenging real-world datasets and comparison with various baselines. This evaluation serves as the comprehensive proof of concept for its capabilities.

LLMs and Datasets

The study employed four state-of-the-art general-purpose LLMs: GPT-4 (gpt-4-turbo), Gemini-1.5 (gemini-1.5-pro), Claude-3.5 (claude-3.5-sonnet), and Llama-3.1 (llama-3.1-70b). These models were chosen for their power and diversity across vendors (OpenAI, Google, Anthropic, Meta). Preliminary experiments showed that code-specific LLMs (CodeLlama, CodeQwen 1.5, DeepSeek-Coder-v2) performed significantly worse (e.g., CodeLlama achieved only 1.21% F1), underscoring the need for the advanced reasoning capabilities of general-purpose models.

APPATCH's evaluation utilized three distinct datasets:

  1. Exemplar Mining Dataset: 306 vulnerability fixing samples were collected from PatchDB [49] and CVEFixes [9]. These samples were manually inspected to filter out irrelevant edits and label vulnerability manifestation locations and CWE IDs. The dataset covers common C language CWEs such as CWE-787 (out-of-bound write), CWE-125 (out-of-bound read), CWE-190 (integer overflow), CWE-401 (memory leak), CWE-457 (use of uninitialized variable), and CWE-476 (use of NULL pointer).
  2. Zero-Day Testing Dataset: To mitigate data leakage concerns (LLMs potentially being trained on existing vulnerabilities), a dataset of 97 zero-day vulnerabilities was curated. All these vulnerabilities were reported after the latest cutoff date (April 2024) of the LLMs used. This dataset covers 18 open-source projects, including the Linux Kernel and FFmpeg, and includes 21 interprocedural vulnerabilities, representing a highly realistic and challenging evaluation scenario.
  3. ExtractFix Testing Dataset: For direct comparison with traditional vulnerability patching techniques, 20 reproducible vulnerabilities from the ExtractFix dataset [20] were used. These samples are compilable and come with test cases, allowing for validation of generated patches.

Evaluation Metrics

To thoroughly assess APPATCH's performance, the evaluation employed a comprehensive set of metrics, considering that exact patch matches are rare in real-world scenarios:

  • Syntactic Equivalent (SynEq): The generated patch exactly matches the ground-truth patch.
  • Semantic Equivalent (SemEq): The generated patch does not exactly match the ground truth but exhibits the same behavior and fixes the vulnerability.
  • Plausible: The patch has different behavior from the ground truth but still fixes the vulnerability without breaking code functionality.
  • Correct: A patch is considered correct if it falls into any of the SynEq, SemEq, or Plausible categories.
  • Recall: The proportion of vulnerable samples for which at least one generated patch correctly fixes the vulnerability (number of fixed samples / total testing samples).
  • Precision: The proportion of generated patches that are correct (number of correct patches / total generated patches).
  • F1 Score: The harmonic mean of recall and precision, providing a balanced measure of effectiveness.

LLMs were prompted to generate up to five candidate patches per vulnerability to simulate real-world scenarios where developers might choose from multiple options. All generated patches were manually inspected and categorized by the authors.

End-to-End Usability Evaluation

To demonstrate APPATCH's practical usability, its performance was evaluated in an end-to-end setting where vulnerability location and CWE information were provided by an upstream detection tool, CodeQL [1]. Two scenarios were considered:

  • Fully Automated: CodeQL outputs were directly fed to APPATCH without human intervention.
  • Realistic: CodeQL outputs were first inspected, verified, and calibrated by a graduate student with 3 years of relevant experience, simulating a typical developer workflow.

The results showed that APPATCH, even with the inherent inaccuracies of upstream tools, performed reasonably well and significantly outperformed baseline approaches in both scenarios. For instance, with Claude-3.5, APPATCH achieved 23.85% F1 in the fully automated scenario and 27.29% F1 in the realistic scenario on the Zero-Day dataset, compared to VulRepair's 11.79% F1. This highlights APPATCH's potential for integration into existing security pipelines, with performance gains expected to improve with more accurate upstream detection tools.

The rigorous evaluation, covering diverse vulnerabilities, LLM models, and real-world scenarios, firmly establishes APPATCH's capabilities as a robust and effective solution for automated software vulnerability patching.

Defensive Implications

The development of APPATCH presents several significant implications for cybersecurity defenders and the broader software development community. Its novel approach to automated vulnerability patching offers tangible benefits in the ongoing battle against software flaws.

Firstly, APPATCH addresses the critical need for timely and effective vulnerability patching. By automating a significant portion of the patching process, it can drastically reduce the time between vulnerability discovery and remediation, especially for zero-day vulnerabilities that lack exploit evidence or test cases. This proactive capability allows organizations to patch vulnerabilities before they are widely exploited, thereby strengthening their overall security posture.

Secondly, APPATCH directly contributes to reducing the burden on human developers. Manual vulnerability patching is a labor-intensive, costly, and specialized task. By generating multiple plausible and validated patch candidates, APPATCH allows developers to shift their focus from crafting fixes from scratch to reviewing and selecting the most appropriate solution. This efficiency gain can free up valuable developer resources to concentrate on other critical security and development tasks.

Thirdly, the framework demonstrates a powerful methodology for effectively leveraging general-purpose LLMs in security workflows. APPATCH highlights that simply applying LLMs to complex tasks like patching is insufficient. Instead, providing domain-specific guidance through vulnerability semantics reasoning and adaptive prompting is crucial for eliciting accurate and reliable results. This approach serves as a blueprint for integrating LLMs into other security-sensitive applications, emphasizing the importance of intelligent prompt engineering and traditional program analysis techniques to complement LLM capabilities.

Fourthly, APPATCH's robust performance on interprocedural vulnerabilities is particularly noteworthy. Many real-world vulnerabilities are complex, spanning multiple functions and modules. The ability of APPATCH to analyze and patch these intricate flaws, where many baseline approaches falter, signifies its potential for tackling more sophisticated attack surfaces.

Finally, the end-to-end evaluation with CodeQL showcases APPATCH's potential for integration into existing vulnerability management pipelines. While the quality of upstream detection and localization remains a factor, APPATCH can seamlessly fit into a broader ecosystem of security tools. As vulnerability detection tools become more accurate, APPATCH's performance in such integrated workflows is expected to further improve, leading to a more automated and efficient security development lifecycle. Defenders should consider exploring such integrated solutions to streamline their patching efforts and enhance their responsiveness to emerging threats.

Key Takeaways

  • Advanced Automated Patching: APPATCH introduces a novel, automated framework for vulnerability patching that leverages Large Language Models (LLMs) without requiring training data, fine-tuning, or exploit evidence.
  • Semantics-Aware Scoping is Crucial: By focusing LLMs on vulnerability semantics extracted through interprocedural backward slicing, APPATCH overcomes token limitations and reduces distractions, enabling more accurate root cause analysis.
  • Dynamic Adaptive Prompting Guides LLMs: The system employs a progressive and adaptive prompting strategy, including iterative root cause generation and dynamic exemplar selection, to guide LLMs through complex reasoning steps, significantly improving patch quality.
  • Multi-Faceted Validation Enhances Reliability: An ensemble of LLMs performs multi-faceted patch validation, ensuring that generated patches not only fix the vulnerability but also preserve functionality, thereby increasing the practical utility of APPATCH's output.
  • Superior Performance on Real-World Vulnerabilities: APPATCH substantially outperforms both existing LLM prompting methods (e.g., up to 28.33% higher F1) and state-of-the-art non-LLM techniques (e.g., up to 182.26% higher recall) on diverse datasets, including challenging zero-day and interprocedural vulnerabilities.
  • Practical and Efficient: APPATCH is reasonably efficient, generating patches in tens of seconds and utilizing thousands of tokens per sample, demonstrating its feasibility for integration into real-world development and security workflows.

About the Speaker(s)

The research behind APPATCH was a collaborative effort by a team of academics from multiple institutions:

  • Yu Nong is affiliated with the University at Buffalo.
  • Haoran Yang is affiliated with Washington State University.
  • Long Cheng is affiliated with Clemson University.
  • Hongxin Hu is affiliated with the University at Buffalo.
  • Haipeng Cai is affiliated with the University at Buffalo and is the corresponding author of the paper.

Their collective expertise contributed to the design, implementation, and evaluation of the APPATCH framework, focusing on advancing the state-of-the-art in automated software security and leveraging the capabilities of large language models.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Solid systems paper that actually addresses the real problems with LLM-based patching instead of just throwing GPT at code and calling it research. The semantics-aware scoping and adaptive prompting aren't revolutionary ideas, but the combination works and the evaluation on zero-days is genuinely useful. Not a breakthrough, but real engineering that moves the needle.

Heather Calloway (CISO) — SOLID

Credible academic work on LLM-guided vulnerability patching with meaningful improvements over baselines. The 50% recall ceiling and reliance on accurate upstream detection tools mean this isn't ready for production deployment, but it signals where automated patching is heading and what capabilities defenders should start planning for.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)