Enhancing Security in Third-Party Library Reuse – Comprehensive Detection of 1-day Vulnerability through Code Patch Analysis

Shangzhi Xu (UNSW)

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Software Security: Applications & Policies · Software Security: Applications & Policies

Overview

This talk presents Vulture, an innovative tool designed to enhance security in third-party library (TPL) reuse by comprehensively detecting one-day vulnerabilities through code patch analysis. Given the pervasive nature of code reuse in modern software development, understanding and mitigating the risks associated with vulnerable libraries is paramount. Many software projects rely heavily on open-source components, leading to a complex web of dependencies where a single vulnerability in an upstream library can propagate downstream, affecting numerous applications. Shangzhi Xu from UNSW highlights the significant challenges developers face in identifying and patching these vulnerabilities, particularly when dealing with large codebases, nested library dependencies, and custom modifications to reused components.

Watch on YouTube · Slides

Key moments

  1. 0:00 Third-party library reuse and 1-day vulnerability propagation
  2. 2:00 Challenges in large-scale patch collection and nested TPLs
  3. 4:20 Introducing Vulture: 3-component 1-day vulnerability detector
  4. 4:50 Building comprehensive database for vulnerability detection
  5. 6:00 Detecting reused third-party libraries and optimization
  6. 7:00 Accurately detecting 1-day vulnerabilities via chunk analysis
  7. 7:30 Demonstrating chunk-based analysis for custom patches

Enhancing Security in Third-Party Library Reuse – Comprehensive Detection of 1-day Vulnerability through Code Patch Analysis

Speakers: Shangzhi Xu, UNSW

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=MptkJKEQOdE

Overview

This talk presents Vulture, an innovative tool designed to enhance security in third-party library (TPL) reuse by comprehensively detecting one-day vulnerabilities through code patch analysis. Given the pervasive nature of code reuse in modern software development, understanding and mitigating the risks associated with vulnerable libraries is paramount. Many software projects rely heavily on open-source components, leading to a complex web of dependencies where a single vulnerability in an upstream library can propagate downstream, affecting numerous applications. Shangzhi Xu from UNSW highlights the significant challenges developers face in identifying and patching these vulnerabilities, particularly when dealing with large codebases, nested library dependencies, and custom modifications to reused components.

The core problem Vulture addresses is the difficulty in accurately identifying which specific instances of reused TPLs within a target program are affected by known, recently disclosed vulnerabilities (one-day vulnerabilities). Traditional methods often fall short due to issues like poorly maintained licenses, high false positive rates from custom code modifications, or the sheer manual effort required to track patches. Vulture tackles these limitations by leveraging a combination of static analysis, a meticulously constructed database of TPLs and their associated patches, and advanced semantic comparison techniques to precisely pinpoint vulnerable code sections and suggest fixes. This research is crucial for bolstering software supply chain security and empowering developers to proactively defend against rapidly emerging threats.

Background

▶ Watch: Third-party library reuse and 1-day vulnerability propagation (0:00)

The modern software landscape is characterized by extensive code reuse, where developers integrate numerous open-source libraries to accelerate development and leverage existing functionalities. This practice, while efficient, introduces significant security challenges. A one-day vulnerability refers to a vulnerability that has been disclosed (and often patched) but has not yet been applied by downstream users of the affected library. The propagation of such vulnerabilities can be widespread; for instance, a flaw like CV 2018-498 in lib JPEG could affect any software, such as React OS, that reuses the library without proper patching.

The talk identifies three primary challenges in managing these vulnerabilities in TPLs:

  1. Large-scale patch collection: Manually gathering and applying patches for numerous libraries is a time-consuming and labor-intensive process.
  2. Nested dependencies: Libraries often reuse other libraries, creating hidden dependencies. A project like React OS might reuse FreeType, which in turn reuses lib bzzip2, meaning React OS implicitly uses lib bzzip2 without its developers necessarily realizing it. This obfuscates the true dependency graph and potential vulnerability surface.
  3. Custom modifications: Developers frequently modify reused libraries to fit specific project requirements. These custom changes make it difficult for traditional vulnerability detection tools, which often rely on exact signature matching, to determine if a vulnerability exists or if a patch has been effectively applied.

Previous work in "supply chain security" or "third-party library security" generally falls into two categories: license-based detection and code-based detection. License-based methods often suffer from high false negative rates due to incomplete or inaccurate license information. Code-based methods, while more robust, struggle with high false positive rates when libraries undergo custom modifications, as simple hash or signature comparisons fail to account for semantic equivalence. These limitations underscore the need for a more accurate and comprehensive approach, which Vulture aims to provide. The research was driven by three key questions: how to build a comprehensive, specific, and maintainable database for one-day vulnerability detection; how to accurately identify reused TPLs in a target program; and how to precisely identify one-day vulnerabilities, even with custom modifications.

Key Findings

▶ Watch: Introducing Vulture: 3-component 1-day vulnerability detector (4:20)

Vulture’s primary contribution lies in its ability to significantly improve the accuracy and efficiency of one-day vulnerability detection in reused third-party libraries (TPLs). The tool’s key findings demonstrate a marked superiority over existing academic and commercial solutions, particularly in reducing both false positives and false negatives.

Firstly, Vulture addresses the challenge of database quality through its TPL filter construction process. By intelligently selecting widely used libraries and eliminating redundant functions based on creation time, similarity, and TPL name, Vulture creates a more compact and maintainable database. This approach makes the database easier to expand and update, a crucial aspect for tracking rapidly evolving vulnerabilities. The system also excels in patch collection, leveraging Large Language Models (LLMs) to efficiently map CVEs to specific commits and extract relevant patches, significantly outperforming existing tools in speed and accuracy.

Secondly, Vulture introduces an identification optimization step for TPL reuse identification. Beyond simple code similarity comparisons, it prioritizes the earliest created TPLs with the most similar names to the reused paths, effectively filtering out numerous false positives that plague other code-based detection methods. This refinement ensures that the identified reused libraries are highly accurate, providing a solid foundation for subsequent vulnerability analysis.

Finally, the most impactful finding is Vulture's chunk-based analysis for one-day vulnerability detection. This innovative technique extracts semantic information from patches by grouping modified lines based on control and variable dependencies. By comparing these semantic "chunks" rather than raw code, Vulture can accurately confirm the existence of a vulnerability even when custom modifications have introduced syntactic differences in the applied patch. This semantic understanding is critical for overcoming the high false positive rates common in other tools when dealing with customized library versions. The evaluation results powerfully illustrate these findings: Vulture successfully identified 184 out of 200 benchmark cases, whereas baselines only caught 100. In real-world scenarios, Vulture detected 175 vulnerabilities across five target programs, significantly outperforming a commercial tool (111 detections) and an academic baseline (13 detections), all while maintaining superior detection speed.

Technical Deep Dive

▶ Watch: Building comprehensive database for vulnerability detection (4:50)

Vulture's architecture is meticulously designed, comprising three main components: TPL filter construction, TPL reuse identification, and one-day vulnerability detection. Each component leverages static analysis and advanced comparison techniques to achieve its objectives.

TPL Filter Construction

This initial phase focuses on building a comprehensive, specific, and maintainable database of TPLs and their associated vulnerabilities.

  1. Function Summary Generation: Vulture starts by scanning GitHub repositories with over 100 stars, filtering out non-library projects through keyword matching in documentation. For each function within selected TPLs, it calculates a hashing value to represent its code.
  2. Redundant Function Elimination: A critical step, especially given nested library reuse (e.g., FreeType reusing lib bzzip2), is to eliminate redundant functions. This is achieved by comparing function creation times, similarity between functions, and TPL names. This process ensures the database remains compact and accurate, preventing duplicate entries and improving efficiency. The output is a TPL summary containing file abstractions, function hashing times, and other relevant metadata.
  3. Patch Collection: Vulture collects CVE (Common Vulnerabilities and Exposures) information and descriptions from the NLC website. It then gathers commit information for CVE-affected repositories from GitHub. A key innovation here is the use of Large Language Models (LLMs) to identify CVE features and commit attributes, and then map specific CVEs to their corresponding commits. This automated mapping allows for the precise extraction of patches associated with each CVE.
  4. Database Construction: The collected CVE mappings and TPL summaries are then integrated to build Vulture's core database, which serves as the foundation for subsequent detection processes.

TPL Reuse Identification

Once the database is constructed, Vulture proceeds to identify which TPLs are reused within a target program.

  1. Function Hashing and Similarity Comparison: For each function in the target program, a hashing value is generated, similar to the TPL filter construction phase. These target program function hashes are then compared against the TPL summary in the database to find potential reuse candidates. TPLs exceeding a predefined similarity threshold are initially selected as candidate reused TPLs.
  2. Identification Optimization: To address the high false positive rates inherent in simple similarity comparisons, Vulture employs a sophisticated optimization step. It compares each candidate TPL with the identified reuse paths within the target program. The system then retains only the earliest created TPL that has the most similar name to the reuse path, effectively filtering out less relevant or incorrect matches. This process generates a reuse report, detailing the identified reused TPLs, their versions, reused files, and specific reused functions.

One-Day Vulnerability Detection

This is the most critical and innovative phase, designed to accurately confirm the presence of vulnerabilities even in the face of custom code modifications.

  1. Candidate Vulnerability Identification: Based on the reuse report, Vulture identifies potential one-day vulnerabilities by cross-referencing the reused TPLs and their versions with the CVE mappings in its database. This step yields a list of candidate one-day vulnerabilities.
  2. Chunk-Based Analysis: For each candidate vulnerability, Vulture performs a chunk-based analysis to semantically confirm its existence. This is where Vulture significantly differentiates itself.
  • Static Patch Analysis: The process begins by extracting semantic information through static analysis of the official patch.
  • Chunk Grouping: Modified lines in the patch are grouped into "chunks" based on their control and variable dependencies. For example, a patch might modify several lines related to a specific variable's initialization, manipulation, and conditional checks. These lines, though syntactically distinct, form a semantic unit.
  • Abstract Representation: Vulture then generates an abstract representation of these chunks, focusing on variables, operations, and how operations are applied to variables, rather than exact syntax.
  • Semantic Comparison: This abstract representation of the official patch is then compared against the corresponding code in the target program (the potentially vulnerable code and any custom patches applied by the user).
  • Example: CV 2018-12498 in React OS: The talk illustrates this with CV 2018-12498. The official patch and a custom patch applied by React OS might have syntactic differences (e.g., variable names, whitespace). However, Vulture's chunk-based analysis groups the modified lines, extracts their semantic intent (e.g., "variable x is initialized with y, then z is assigned to x if condition C is met"), and then compares these semantic abstractions. If the abstractions match, Vulture concludes that the patch (or an equivalent custom patch) has been applied, and the vulnerability is not present. If the vulnerable code without the patch's semantic intent is found, the vulnerability is confirmed.
  1. Vulnerability Report: The final output is a vulnerability report detailing confirmed one-day vulnerabilities and providing specific fix suggestions, enabling developers to address the issues directly.

Demo / Proof of Concept

▶ Watch: Accurately detecting 1-day vulnerabilities via chunk analysis (7:00)

The efficacy of Vulture was rigorously evaluated through two distinct experiments: one assessing the database quality and another focusing on scalability and detection performance.

For database quality, Vulture demonstrated superior performance compared to academic baselines. In TPL selection, Vulture effectively excluded non-library repositories and retained only widely used TPLs, leading to significantly reduced storage consumption. This makes the database more compact and efficient. When assessing maintainability and the ease of updating the database (e.g., adding a new TPL), Vulture's database proved much more compact, simplifying expansion and accelerating update processes. Furthermore, in patch collection, Vulture's methodology, particularly its use of LLMs for CVE-to-commit mapping, was substantially faster and more accurate than existing tools.

The scalability evaluation encompassed both benchmark detection and vulnerability detection in the wild.

  1. Benchmark Detection: In a set of 200 test cases specifically designed to evaluate one-day vulnerability detection, Vulture achieved an impressive success rate, accurately identifying 184 cases. In stark contrast, the chosen baselines (not explicitly named but implied to be academic) only managed to identify 100 of these cases. This highlights Vulture's superior precision and coverage in controlled environments.
  2. Vulnerability Detection in the Wild: To assess real-world applicability, Vulture was tested against five target programs, scanning for vulnerabilities in their reused TPLs. Vulture successfully identified 175 one-day vulnerabilities. For comparison, a prominent commercial tool (specifically mentioned as Sneak in the Q&A) only identified 111 vulnerabilities, while an academic baseline detected a mere 13. This significant performance gap underscores Vulture's practical utility and its ability to outperform both industry-standard and research-oriented solutions in real-world scenarios.

The talk emphasized that Vulture's exceptional performance—achieving higher detection rates while reducing both false negatives and false positives—stems directly from its well-designed database and accurate semantic analysis. The comprehensive and specific database reduces false negatives by ensuring relevant TPLs and patches are available. False positives are significantly curtailed by the identification optimization step in TPL reuse detection and, most critically, by the chunk-based analysis, which handles custom modifications semantically.

Regarding time consumption, Vulture proved remarkably efficient. It could detect one TPL reuse in approximately 25 seconds and identify one one-day vulnerability in just 5 seconds, making it considerably faster than the evaluated baselines. This speed, combined with its accuracy, positions Vulture as a highly practical tool for continuous security monitoring in large-scale software development environments.

Defensive Implications

▶ Watch: Demonstrating chunk-based analysis for custom patches (7:30)

Vulture offers critical defensive implications for organizations struggling with the pervasive challenge of third-party library (TPL) vulnerabilities in their software supply chain. Its robust methodology provides a clear path for defenders to proactively identify and mitigate one-day vulnerabilities that often slip through the cracks of traditional scanning tools.

Firstly, Vulture's ability to build a comprehensive and specific database of TPLs and their associated patches means defenders gain access to more accurate and up-to-date vulnerability intelligence. This reduces the reliance on potentially incomplete or outdated information, leading to fewer missed vulnerabilities (false negatives). Security teams can leverage this improved database to better understand their software's dependency graph and the potential impact of newly disclosed CVEs.

Secondly, the identification optimization for TPL reuse detection, coupled with the innovative chunk-based analysis, directly addresses the long-standing problem of high false positives caused by custom code modifications. By semantically comparing patches, Vulture can accurately determine if a vulnerability truly exists, even if a custom patch has been applied with different syntax. This capability is invaluable for reducing alert fatigue among development and security teams, allowing them to focus resources on genuine threats rather than chasing phantom vulnerabilities. Defenders can integrate Vulture's output into their CI/CD pipelines to get precise, actionable insights, rather than generic warnings.

Thirdly, Vulture's superior performance in both benchmark and "in the wild" detection, significantly outperforming commercial tools like Sneak, suggests that incorporating its techniques could lead to a more secure software ecosystem. Organizations can use tools built on Vulture's principles to automate the detection of one-day vulnerabilities in their own codebases, providing rapid feedback to developers. This facilitates a shift-left security approach, where vulnerabilities are identified and remediated earlier in the development lifecycle, drastically reducing the cost and effort of fixing them later.

Finally, the speed of Vulture's detection (25 seconds for TPL reuse, 5 seconds for vulnerability detection) makes it suitable for integration into routine scanning processes. Defenders can implement continuous monitoring strategies, regularly scanning their applications for new one-day vulnerabilities as patches become available, thereby minimizing the window of exposure. While the current implementation requires source code access, the research lays a strong foundation for future binary-level analysis, potentially extending these benefits to scenarios where source code is unavailable.

Key Takeaways

  • Semantic Analysis is Key: Vulture's innovative chunk-based analysis, which semantically compares patches based on control and variable dependencies, is crucial for accurately detecting vulnerabilities even when custom modifications have been applied to third-party libraries.
  • Superior Detection Accuracy: Vulture significantly outperforms both academic and commercial tools (e.g., Sneak), identifying 184/200 benchmark cases and 175 one-day vulnerabilities in real-world applications, compared to 100 and 111 (commercial) or 13 (academic) respectively.
  • Comprehensive & Maintainable Database: The tool constructs a highly specific and compact database of TPLs and their associated patches, utilizing LLMs for efficient CVE-to-commit mapping, which reduces storage and simplifies updates.
  • Reduced False Alarms: Through identification optimization in TPL reuse detection and semantic analysis, Vulture effectively minimizes both false negatives (missing vulnerabilities) and false positives (reporting non-existent vulnerabilities), a common challenge in code-based detection.
  • Efficient and Actionable Insights: Vulture provides rapid detection (25 seconds for TPL reuse, 5 seconds for vulnerability detection) and generates specific fix suggestions, enabling developers and security teams to efficiently address one-day vulnerabilities in their software supply chain.
  • Addresses Core Challenges: The research directly tackles the complexities of large-scale patch collection, nested library dependencies, and custom code modifications that hinder traditional third-party library security approaches.

About the Speaker(s)

Shangzhi Xu is a researcher from UNSW (University of New South Wales). His work focuses on enhancing software security, particularly in the context of third-party library reuse and the detection of one-day vulnerabilities. This presentation at the NDSS Symposium highlights his contributions to developing advanced static analysis techniques for improving software supply chain security.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Vulture is competent, well-executed academic work on a real problem — 1-day vuln detection in reused TPLs with custom modifications. The chunk-based semantic patch analysis is a genuine contribution and the numbers against Sneak are credible. But this is NDSS-track research, not a practitioner conference drop: it's solid incremental progress in a crowded space, not a paradigm shift.

Heather Calloway (CISO) — WEAK

Vulture is technically credible work — semantic patch analysis to catch 1-day vulnerabilities through custom modifications is a real problem worth solving, and the detection numbers suggest genuine capability. But this is an academic proof-of-concept with no bridge to the security leaders, AppSec programs, or supply chain governance decisions that actually own this risk.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025