The Midas Touch: Triggering the Capability of LLMs for RM-API Misuse Detection
Yi Yang (Institute of Information Engineering, Chinese Academy of Sciences)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · API Security
Overview
This talk, presented by Jinalu (one of the authors) on behalf of Yi Yang, Kachin, and Menlin from the Institute of Information Engineering, Chinese Academy of Sciences, introduces a novel approach called ChatDetector for identifying and detecting misuses of Resource Management API (RM-API) pairs using Large Language Models (LLMs). The core problem addressed is the pervasive issue of RM-API misuse, which often leads to critical security vulnerabilities such as memory corruption, denial of service, and data leakage. These misuses stem from developers' oversight or poorly documented API contracts, where freeing or releasing operations are implicitly expected but not explicitly enforced or clearly stated.
Key moments
- 0:00 Introduction: The problem of ARM API misuse
- 1:00 Limitations of traditional ARM API misuse detection methods
- 2:40 New challenges: LLM fabrication and incorrect answers
- 3:50 Trend Detector: A new end-to-end methodology overview
- 4:15 Detailed ARM API identification and cross-validation
- 6:10 Overcoming LLM parameter confusion through preprocessing
- 7:50 Applying CodeQL for ARM API misuse detection
- 9:00 Real-world findings: bugs and documentation errors
The Midas Touch: Triggering the Capability of LLMs for RM-API Misuse Detection
Speakers: Yi Yang (Institute of Information Engineering, Chinese Academy of Sciences)
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=99zzZP9hXUQ
Overview
This talk, presented by Jinalu (one of the authors) on behalf of Yi Yang, Kachin, and Menlin from the Institute of Information Engineering, Chinese Academy of Sciences, introduces a novel approach called ChatDetector for identifying and detecting misuses of Resource Management API (RM-API) pairs using Large Language Models (LLMs). The core problem addressed is the pervasive issue of RM-API misuse, which often leads to critical security vulnerabilities such as memory corruption, denial of service, and data leakage. These misuses stem from developers' oversight or poorly documented API contracts, where freeing or releasing operations are implicitly expected but not explicitly enforced or clearly stated.
Traditional methods for detecting such misuses have significant limitations, struggling with complex code structures, incomplete API name matching, and inadequate documentation processing. With the recent advancements in natural language understanding capabilities of LLMs, the researchers explored their potential to overcome these obstacles. The presented work demonstrates how LLMs, when properly guided through a structured task decomposition and validation process, can accurately identify, pair, and detect misuses of RM-APIs, significantly outperforming existing techniques and uncovering numerous previously unknown vulnerabilities in popular software libraries.
The significance of this research lies in its innovative application of LLMs to a challenging and critical area of software security. By leveraging LLMs' reasoning and comprehension abilities, ChatDetector provides a more robust and scalable solution for identifying resource management issues that are often subtle and hard to catch. This has profound implications for improving software reliability and security, encouraging the broader adoption of LLMs for advanced security research, and highlighting the need for better API documentation practices.
Background
▶ Watch: Introduction: The problem of ARM API misuse (0:00)
Resource Management APIs (RM-APIs) are fundamental components in software development, typically used in pairs to manage system resources such as memory, file handles, network sockets, or synchronization primitives. Examples include malloc/free for memory allocation and deallocation, open/close for file operations, or acquire/release for locks. While the concept of pairing these operations is often considered common sense among experienced developers, the explicit linking or documentation of these pairs is frequently omitted or unclear within library documentation. This lack of clarity is exacerbated by library developers sometimes not explicitly writing freeing APIs for every function, leading to a prevalent problem of RM-API misuse.
The consequences of such misuses are severe, ranging from memory leaks (leading to eventual denial of service), use-after-free vulnerabilities (allowing attackers to execute arbitrary code or cause crashes), and double-free errors (also leading to crashes or exploitable conditions). These issues can result in memory corruption, denial of service, and sensitive data leakage, posing significant security risks to applications.
Traditional approaches to detecting RM-API misuse have faced considerable challenges:
- Automatic code analysis: While powerful, these methods, such as static analysis, often struggle with complex, multi-layer nesting functions, precisely identifying only 72.7% of RM-APIs in such scenarios. Their precision can be limited by the intricacies of control flow and data flow analysis.
- Keyword matching for API names: This simplistic approach relies on identifying common keywords like "malloc" or "free" in API names. However, it is highly prone to missing a vast majority of ARM APIs (76.28% according to the researchers' experiments) because many RM-APIs use diverse naming conventions (e.g.,
create,init,open,sitefor allocation;destroy,deinit,closefor deallocation). - Template matching in documentation processing: This method attempts to extract API rules from documentation by matching predefined templates. However, it is often incapable of covering all libraries due to variations in documentation styles and structures. Furthermore, existing NLP tools struggle with identifying API pairs that are described with neutral sentiment or span across different sections or even separate pages of documentation. For instance, an allocation semantic might be described neutrally on one page, and its corresponding freeing semantic on another, making it difficult for traditional methods to link them.
These limitations underscore the need for more advanced tools capable of understanding the nuanced semantics of API documentation, especially when dealing with complex sentences, cross-page references, and neutral sentiment descriptions. The advent of large language models presented a potential breakthrough, offering superior natural language understanding capabilities that could overcome the inherent obstacles faced by traditional detection methods.
Key Findings
▶ Watch: New challenges: LLM fabrication and incorrect answers (2:40)
The research introduces ChatDetector, a novel LLM-based framework that significantly advances the detection of RM-API misuses. The core findings demonstrate ChatDetector's superior capability in identifying and pairing RM-APIs and subsequently detecting associated security bugs compared to traditional methods.
The primary contributions and findings include:
- Effective Task Decomposition: ChatDetector successfully decomposes the complex problem of RM-API misuse detection into three manageable stages: ARM API identification, ARM API pairing, and ARM API misuse detection. This structured approach allows LLMs to focus on specific sub-tasks, improving accuracy and reliability.
- Overcoming LLM Limitations: The researchers identified critical challenges when directly applying LLMs, such as answer fabrication without expertise (e.g., suggesting a non-existent
EV watch check freeAPI forEV watch check new) and introducing incorrect answers despite evidence. ChatDetector addresses these by deploying a two-dimensional cross-validation mechanism and requiring LLMs to output not only the answer but also the evidence retrieved from documentation and the reasoning process. This ensures a robust verification of LLM outputs. - Superior API Identification and Pairing: In experiments conducted across six popular libraries, ChatDetector demonstrated remarkable improvements:
- It identified 47% more ARM sentences (sentences describing RM-API semantics) compared to previous work (specifically, "Metro work").
- It identified 80.85% more ARM API pairs than previous methods, indicating a much broader and more accurate understanding of resource management contracts.
- Diverse API Type Handling: Beyond the basic
malloc/freepairs, ChatDetector proved capable of handling 22 different types of ARM APIs based on their operation semantics, includingsite,open,create, andinitiate. This highlights the limitations of initial keyword matching methods and ChatDetector's ability to infer complex resource management patterns. - Significant Bug Discovery: The application of ChatDetector led to the discovery of 165 ARM API pairs and 115 security bugs across the tested libraries. These bugs primarily included memory leaks, use-after-free, and double-free issues, which were reported to the respective maintainers, with patches provided and some confirmed.
- Robustness to Documentation Quality: The study investigated the impact of documentation quality on detection performance. While there was a significant difference in performance between having documentation and having no documentation, ChatDetector showed only a slight difference in performance among various qualities of documentation. This implies that the method is robust to variations in documentation quality, a crucial characteristic for real-world applicability where documentation can be inconsistent.
- Identification of Documentation Inaccuracies: ChatDetector not only found code bugs but also identified inaccuracies within API documentation itself. For example, the documentation for
av_get_tokenstated thatav_freeshould be used to free the allocated string, butav_freedoes not exist, which would lead to false positives in previous work.
These findings collectively demonstrate that LLMs, when strategically employed and augmented with verification mechanisms, can provide an unprecedented level of understanding for complex API documentation, leading to highly effective and accurate detection of resource management vulnerabilities.
Technical Deep Dive
▶ Watch: Detailed ARM API identification and cross-validation (4:15)
The ChatDetector framework is designed to overcome the inherent limitations of both traditional static analysis and naive LLM application by strategically decomposing the problem and integrating robust verification mechanisms. The method achieves end-to-end ARM API misuse detection by leveraging external information from API descriptions and function definitions.
Addressing LLM Challenges
Initially, direct querying of LLMs for RM-API information presented significant issues:
- Fabrication without Expertise: LLMs would generate plausible-sounding but non-existent APIs. For example, when asked to free
EV watch check new, an LLM might suggestEV watch check free, which follows a common naming convention but is not a real API. This occurs because LLMs, while adept at pattern matching, may lack specific domain knowledge for obscure or library-specific APIs. - Incorrect Answers with Evidence: More concerningly, LLMs sometimes provided answers that contradicted the very evidence supplied to them. For instance, when given a description of
zip_discardcontaining "releasing semantics," the LLM might output a contradictory answer.
To mitigate these, ChatDetector deploys a complex task decomposition strategy and a two-dimensional cross-validation approach. For each task, the LLM is prompted to output its answer, the supporting evidence retrieved from the documentation, and the reasoning process. This forces the LLM to justify its output and allows for external validation.
ChatDetector's Three-Part Decomposition
The methodology is structured into three main phases:
1. ARM API Identification
This phase aims to identify sentences in documentation that describe RM-API semantics and classify APIs as allocation APIs.
- ARM Sentence Identification: The task is posed as a binary classification open question. The LLM is asked if a given sentence describes the semantics of an API. For complex APIs like
pcap_findalldevs, whose documentation might contain 20 or more lines, traditional NLP tools struggle. However, the LLM can precisely identify the relevant ARM sentences. Crucially, the LLM is also asked to output the evidence (the specific sentence or phrase from the documentation) to support its answer, following an in-context answering approach. - Allocation API Identification: To identify if an API performs allocation, a prompt template is designed. First, a binary classification question is asked: "Does the current API perform allocation?" Then, the LLM is prompted for evidence and the reasoning process. This three-part output (answer, evidence, reasoning) serves as a double-check.
- Cross-Validation Algorithm: A sophisticated three-step cross-validation process is used to confirm the correctness of the LLM's output regarding API semantics:
- Ask: "What is the allocated object by this API?" and check if it exists.
- Ask: "Which API performs allocation?" and verify if the answer is consistent with the current API being analyzed.
- Finally, ask: "Does the current API perform allocation?"
The combination of these three answers is used to robustly validate the semantics of the current API, ensuring high precision in identification.
2. ARM API Pairing
This phase focuses on identifying the corresponding releasing API for an allocated object. The researchers observed that directly using LLMs for releasing API identification often led to incorrect semantics, as LLMs seemed confused by parameter types and names when combined in function definitions.
- Observation: RM-APIs often share a common "ARM object type." This insight is leveraged to improve pairing.
- Function Preprocessing: To address the LLM's difficulty in parsing combined parameter types and names, function definitions are preprocessed into a structured format. This format includes index numbers (0 for return value, 1 for the first parameter, etc.), parameter types, and parameter names. For example, a function
int func(Type1 param1, Type2 param2)would have0: int (return value),1: Type1 (param1),2: Type2 (param2). This structured input makes it easier for the LLM to accurately identify relevant parameters. - ARM Object Identification: The LLM is supplied with the object description and the preprocessed function definition. It then outputs the ARM object type along with its index number, indicating which parameter or return value corresponds to the ARM object. For instance, given
pcap_findalldevs, the LLM might correctly identifypcap_if_tas the allocated object type and1as its index, indicating it's the first parameter. - Releasing API Identification: Finally, given the identified allocated object type (e.g.,
pcap_if_t) and the function definition, the LLM is prompted to output the corresponding releasing API. Forpcap_findalldevs, it correctly identifiespcap_freealldevsas the releasing API.
3. ARM API Misuse Detection
Once ARM API pairs are identified, CodeQL, a popular static analysis tool, is employed to detect misuses in the codebase.
- CodeQL Integration: The researchers manually construct CodeQL queries tailored to detect three common types of security issues:
- Memory leak: An allocated resource is not freed.
- Use-after-free: A resource is accessed after it has been freed.
- Double-free: A resource is freed more than once.
- Example Misuse: A detected bug in a snippet shows
pcap_findalldevsallocatingalldevspon line 2537. However, line 2540 misses the correspondingpcap_freealldevsoperation, leading to a memory leak. - Another Example: The pair
av_dict_setandav_dict_freeis identified.av_dict_setperforms allocation that requiresav_dict_freeat program termination. A bug was found whereav_dict_setwas called butav_dict_freewas omitted, leading to a memory leak. - Documentation Inaccuracies: The system also identified cases where documentation itself was misleading. For
av_get_token, documentation suggested usingav_free, but this API does not exist, which would cause false positives in other detection systems.
By combining the natural language understanding of LLMs with structured data processing and powerful static analysis, ChatDetector provides a comprehensive and highly effective solution for identifying and mitigating RM-API misuse.
Demo / Proof of Concept
▶ Watch: Overcoming LLM parameter confusion through preprocessing (6:10)
The practical application of ChatDetector was demonstrated through its deployment on six popular software libraries. The results clearly illustrated the framework's effectiveness and superiority over existing methods.
The comprehensive evaluation yielded significant findings:
- Discovered Pairs and Bugs: Across the six libraries, ChatDetector successfully identified 165 distinct ARM API pairs. More critically, it uncovered 115 security bugs related to RM-API misuse. These bugs primarily consisted of memory leaks, use-after-free, and double-free vulnerabilities, which are critical issues impacting software stability and security.
- Performance Comparison: When compared to previous work (referred to as "Metro work" in the talk), ChatDetector showed a marked improvement in its ability to understand and extract information from documentation:
- It identified 47% more ARM sentences, indicating a more thorough and accurate parsing of API documentation for resource management semantics.
- It achieved an 80.85% increase in the identification of ARM API pairs, demonstrating its capacity to establish a far more complete mapping of allocation and deallocation functions.
- Beyond Malloc/Free: The system's capabilities extended significantly beyond the common
malloc/freepatterns. ChatDetector was able to correctly handle and detect misuses for 22 different types of ARM APIs, categorized by their operation semantics such assite,open,create, andinitiate. This finding underscored the limitations of keyword-matching methods, which would miss these diverse API types, and highlighted ChatDetector's ability to infer semantics from context. - Robustness to Documentation Quality: A specific experiment investigated the impact of documentation quality. The results indicated that while the presence of documentation significantly boosted performance compared to its absence, the system exhibited only a slight difference in performance across various qualities of documentation. This suggests that ChatDetector is robust enough to perform effectively even with imperfect or inconsistent documentation, a common scenario in real-world software projects.
- Bug Reporting and Confirmation: The identified bugs were reported to the maintainers of the respective software libraries. The researchers not only reported the vulnerabilities but also provided patches to fix them. While the exact number of applied patches was not provided during the talk, it was confirmed that "some of them are confirmed," indicating the practical impact and validity of the findings. This proactive approach reinforces the value of the research in contributing directly to software security.
The demonstration clearly validated ChatDetector as a powerful and practical tool for automated detection of RM-API misuses, offering a substantial improvement over existing techniques by leveraging the advanced natural language understanding capabilities of LLMs.
Defensive Implications
▶ Watch: Real-world findings: bugs and documentation errors (9:00)
The findings from the ChatDetector research offer several crucial implications for software developers, security engineers, and organizations aiming to enhance the robustness and security of their applications.
- Integrate LLM-powered Static Analysis: Defenders should explore integrating LLM-powered tools, inspired by ChatDetector, into their continuous integration/continuous deployment (CI/CD) pipelines. These tools can automatically analyze new or modified codebases against API documentation to detect potential RM-API misuses before deployment. This shifts the detection left in the development lifecycle, reducing the cost and impact of vulnerabilities.
- Proactive Documentation Analysis: Leverage LLMs to proactively analyze existing API documentation for clarity, consistency, and completeness regarding resource management contracts. This can help identify ambiguous or missing information that might lead to developer errors. Tools could flag documentation sections that are unclear about whether an API allocates a resource or how it should be freed, prompting human review and improvement.
- Enhance Static Analysis Tools with LLM Intelligence: Existing static analysis tools like CodeQL can be significantly enhanced by incorporating LLM-derived knowledge. Instead of relying solely on manually crafted rules or limited keyword matching, these tools can use LLMs to automatically generate or refine rules for detecting RM-API pairs, including the more obscure or diversely named ones (e.g.,
site,create,initiate). This reduces the manual effort in rule creation and improves coverage. - Prioritize Fixing Identified Vulnerabilities: Organizations should prioritize addressing the types of security issues identified by ChatDetector, particularly memory leaks, use-after-free, and double-free vulnerabilities. These are critical flaws that can be exploited for denial of service, information disclosure, or arbitrary code execution. Implement automated scanning for these patterns and ensure timely patching.
- Improve Developer Education on Resource Management: The prevalence of RM-API misuse suggests a need for better developer education. Training should emphasize the importance of explicit resource management, the potential pitfalls of implicit contracts, and best practices for using allocation/deallocation pairs, especially in languages like C/C++ where manual memory management is common.
- Generalize Beyond Memory Management: As indicated by the speakers, the methodology of ChatDetector can be extended beyond just memory allocation/deallocation. Defenders should consider adapting this approach to detect misuses of other critical resource types, such as file descriptors (
open/close), network sockets, database connections, or synchronization primitives (lock/unlock). This would require modifying the LLM prompts to reflect the specific semantics of these different resource types. - Address Documentation Inaccuracies: The discovery of incorrect information in official documentation (e.g.,
av_get_tokenspecifying a non-existentav_freeAPI) highlights the importance of documentation quality. Defenders should advocate for rigorous review processes for API documentation, perhaps even using LLMs to cross-reference documentation claims with actual code behavior. - Adopt a "Shift-Left" Security Mindset: The success of ChatDetector reinforces the value of "shift-left" security, where vulnerabilities are identified and remediated early in the development lifecycle. By leveraging advanced analytical tools, developers can catch resource management issues during design or coding phases rather than during testing or, worse, after deployment.
By internalizing these defensive implications, organizations can proactively strengthen their software security posture, reduce the attack surface, and build more reliable and resilient applications in an era where complex API interactions are ubiquitous.
Key Takeaways
- LLMs excel at understanding complex API documentation: Traditional methods struggle with nuanced language, cross-page references, and neutral sentiment, but LLMs effectively parse and interpret these challenging aspects of RM-API semantics.
- Structured decomposition and validation are crucial for LLM reliability: Direct application of LLMs can lead to fabrication or contradictory answers. ChatDetector's three-part task decomposition (identification, pairing, misuse detection) combined with two-dimensional cross-validation, evidence retrieval, and reasoning processes significantly enhances accuracy and trustworthiness.
- ChatDetector significantly outperforms traditional methods: It identified 47% more ARM sentences and 80.85% more ARM API pairs, leading to the discovery of 115 security bugs across six popular libraries, including memory leaks, use-after-free, and double-free issues.
- The method is robust and generalizable: ChatDetector handles 22 diverse types of ARM APIs beyond simple
malloc/freeand is robust to variations in documentation quality. Its underlying principles can be extended to other resource management patterns (e.g., locks, file descriptors) with appropriate prompt modification. - LLMs can identify both code bugs and documentation inaccuracies: Beyond finding vulnerabilities in source code, ChatDetector also revealed misleading or incorrect information in official API documentation, highlighting the need for improved documentation practices.
About the Speaker(s)
The research presented in "The Midas Touch: Triggering the Capability of LLMs for RM-API Misuse Detection" was conducted by Yi Yang, Kachin, Menlin, and Jinalu, all affiliated with the Institute of Information Engineering, Chinese Academy of Sciences. Jinalu delivered the talk at the NDSS Symposium. While specific individual biographies were not provided in the transcript, their work demonstrates expertise in applying advanced natural language processing and machine learning, particularly Large Language Models, to challenging problems in software security, focusing on static analysis and vulnerability detection. Their institutional affiliation, a prominent research body in China, underscores their background in deep technical research within information engineering.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Solid, original research applying LLMs to a problem that actually warrants them — RM-API pair identification from documentation is exactly the kind of semantically dense, inconsistently structured task where NLP traditionally faceplants. The two-dimensional cross-validation mechanism to catch LLM fabrication is the real contribution here, and the 115 confirmed bugs across six real libraries is the receipts. Not paradigm-shifting, but this is careful, reproducible work that advances the state of the art.
Heather Calloway (CISO) — WEAK
Technically credible work — LLM-assisted static analysis applied to a real class of memory safety bugs — but it never crosses the threshold into operator relevance. The research community may find value here; CISOs, defenders, and security program leaders will not know what to do with it.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025