Generating API Parameter Security Rules with LLM for API Misuse Detection
Jinghua Liu (Institute of Information Engineering Chinese Academy of Sciences)
Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · API Security
Overview
In the realm of software development, Application Programming Interfaces (APIs) serve as fundamental building blocks, accelerating development cycles and providing diverse functionalities. However, the widespread adoption of APIs introduces a significant security challenge: developers, often unaware of an API's intricate implementation details, may inadvertently misuse parameters, leading to critical security vulnerabilities. This talk, presented by Jinghua Liu from the Institute of Information Engineering Chinese Academy of Sciences, introduces GBT8, a novel framework that leverages Large Language Models (LLMs) to automatically generate accurate and concrete API parameter security rules (APSRs) for the detection of such misuse.
Key moments
- 0:00 Introduction and problem of API misuse detection
- 2:00 Challenges in generating security rules with LLMs
- 4:00 Overview of the GBT8 framework architecture
- 6:00 Validating API rules using execution feedback
- 8:00 Refining rules by identifying key operations from errors
- 10:00 Example of a concrete refined API security rule
Generating API Parameter Security Rules with LLM for API Misuse Detection
Speakers: Jinghua Liu (Institute of Information Engineering Chinese Academy of Sciences)
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=HK8jjKXnDD4
Overview
In the realm of software development, Application Programming Interfaces (APIs) serve as fundamental building blocks, accelerating development cycles and providing diverse functionalities. However, the widespread adoption of APIs introduces a significant security challenge: developers, often unaware of an API's intricate implementation details, may inadvertently misuse parameters, leading to critical security vulnerabilities. This talk, presented by Jinghua Liu from the Institute of Information Engineering Chinese Academy of Sciences, introduces GBT8, a novel framework that leverages Large Language Models (LLMs) to automatically generate accurate and concrete API parameter security rules (APSRs) for the detection of such misuse.
The core problem GBT8 addresses is the difficulty in systematically identifying and preventing API misuse, which can manifest as issues like memory leaks or double-free errors. Traditional methods for generating APSRs, whether by analyzing API documentation, calling code, or source code, are often limited by resource availability or reliance on rigid, predefined code analysis rules. GBT8 overcomes these limitations by harnessing the advanced code analysis capabilities of LLMs, coupled with an innovative execution feedback checking method and a refinement process that translates vague LLM outputs into actionable security rules.
The significance of GBT8 lies in its ability to automatically uncover previously unknown API misuses and generate new, precise security rules at scale. By demonstrating superior coverage of both APSRs and actual bugs compared to existing state-of-the-art approaches, GBT8 offers a powerful new paradigm for enhancing software security. It provides developers and security practitioners with a robust tool to proactively identify and mitigate API-related vulnerabilities, thereby improving the overall robustness and reliability of software systems.
Background
▶ Watch: Introduction and problem of API misuse detection (0:00)
APIs are the backbone of modern software, but their abstraction can be a double-edged sword. While simplifying development, they can also obscure underlying security constraints, leading to developers inadvertently violating API parameter security rules (APSRs). A common example cited in the talk is failing to close a resource after calling an API like close_open, resulting in a memory leak. Another critical misuse involves incorrect memory management, such as a double-free error, where the same memory region is deallocated multiple times, potentially leading to crashes or exploitable vulnerabilities. The fundamental challenge is two-fold: first, understanding these implicit APSRs, and second, effectively detecting their violations in codebases.
Previous research has attempted to address this problem through various approaches. Some studies focused on extracting APSRs from API documentation or by analyzing existing API calling code. However, these methods are inherently limited by the quality and completeness of the resources they rely on, often leading to the generation of incorrect or incomplete APSRs. Another line of research involved analyzing API source code, which offers direct insight into API behaviors. While more precise, these approaches typically depend on predefined code analysis rules, which restrict their scope and make them inflexible to evolving API patterns or complex interactions.
The advent of Large Language Models (LLMs) presented a promising alternative. LLMs possess a remarkable ability to understand and analyze code without explicit, predefined rules, making them ideal candidates for generating APSRs. However, applying LLMs directly to this problem introduces its own set of challenges. The first major hurdle is the potential for LLMs to generate incorrect APSRs. For instance, an LLM might suggest that a parameter to the free API "must not be null." While conceptually sound, passing a null pointer to free is often a no-op and does not typically cause a security issue. Directly using such an APSR would lead to numerous false positives in misuse detection. The second challenge is that LLM-generated APSRs can be too general or vague to be translated into concrete detection code. An APSR like "parameter 2 must be a valid pointer" is ambiguous; "valid" could mean "not null," "points to allocated memory," or other interpretations, leading to inconsistent or incorrect detection results.
To overcome these challenges, the researchers made two key observations. Firstly, if a piece of code violates a correct APSR, it frequently leads to observable runtime errors. Conversely, if a violation of a proposed APSR runs successfully without error, that APSR is likely incorrect. This insight forms the basis for an execution feedback checking method. Secondly, code, as a formal language, can express far more concrete information than natural language. By analyzing specific code examples where misuse occurs, LLMs can be guided to identify precise conditions and generate concrete APSRs, such as "parameter 2 must not be null," which is much more actionable than "parameter 2 must be a valid pointer." These observations lay the groundwork for GBT8's design, enabling it to leverage LLMs effectively for accurate and concrete APSR generation.
Key Findings
▶ Watch: Overview of the GBT8 framework architecture (4:00)
The primary contribution of this research is the development of GBT8, a novel framework designed to generate accurate and concrete API parameter security rules (APSRs) using Large Language Models (LLMs) and subsequently detect API misuse. GBT8’s key findings and contributions can be summarized as follows:
- Novel LLM-based APSR Generation: GBT8 successfully demonstrates that LLMs can be effectively employed to analyze API source code and generate raw APSRs, overcoming the limitations of predefined code analysis rules.
- Execution Feedback Checking Method: A crucial innovation is the execution feedback checking method, which validates the correctness of LLM-generated APSRs. By executing violation code and observing runtime errors (or lack thereof), GBT8 can filter out incorrect or overly restrictive rules that would otherwise lead to false positives. This empirical validation significantly enhances the reliability of the generated rules.
- Concrete APSR Refinement: GBT8 introduces a systematic approach to refine vague LLM-generated APSRs into concrete, actionable rules. This is achieved by guiding the LLM to perform code differential analysis, identify key operations that cause specific runtime errors, and then describe these operations in a precise manner. This addresses the challenge of ambiguous natural language outputs from LLMs.
- Superior Performance in Misuse Detection:
- High Precision: GBT8 generated 8 distinct types of APSRs with an impressive precision of 92.3%, indicating the high quality of the rules produced.
- Extensive New Discoveries: The framework detected 2510 new API misuses and generated 355 new APSRs across eight popular libraries, significantly expanding the known landscape of API vulnerabilities.
- Enhanced Coverage: Comparative analysis against three state-of-the-art approaches (including Advance, a previous work often referenced for API misuse detection) revealed that GBT8 covers significantly more APSRs and bugs, demonstrating its superior efficacy in identifying a broader range of security issues. For instance, in one comparison, GBT8 found 243 bugs compared to Advance's 99.
- Valuable Prompt Design Insights: The research yielded practical lessons for effective LLM prompt engineering in security contexts. Key insights include:
- Insufficient Information Causes Fabrication: When LLM outputs appear fabricated or incorrect, providing additional context and information in the prompt is crucial for improving accuracy.
- Examples Can Be Harmful: Counter-intuitively, providing examples in the prompt can sometimes limit the LLM's ability to explore diverse outputs, causing it to merely repeat patterns from the examples rather than generating novel or comprehensive rules.
In essence, GBT8 presents a robust, data-driven, and LLM-powered solution that significantly advances the state-of-the-art in automated API misuse detection, making it a powerful tool for improving software security.
Technical Deep Dive
▶ Watch: Validating API rules using execution feedback (6:00)
GBT8 is a sophisticated framework comprising four distinct stages, each meticulously designed to address specific challenges in generating accurate and concrete API parameter security rules (APSRs) using Large Language Models (LLMs) and detecting API misuse.
Stage 1: Raw APSR Generation
The initial stage focuses on leveraging the LLM's understanding of code to generate preliminary, or "raw," APSRs. Recognizing that analyzing an entire API's source code and all its parameters simultaneously can overwhelm an LLM and reduce effectiveness, GBT8 adopts a modular approach. For APIs with multiple parameters, the task of APSR generation is split into subtasks, with the LLM instructed to generate rules for one parameter at a time.
For each subtask, GBT8 provides the LLM with the API's full source code and specific instructions. This granular approach ensures that the LLM can focus its analytical capabilities on the context relevant to a single parameter, leading to more precise initial rule suggestions. For example, for an API like pickup_dump with three parameters, GBT8 would create three separate prompts, each targeting a different parameter, to generate its respective raw APSRs.
Stage 2: APSR Validation (Execution Feedback Checking)
Raw APSRs generated by LLMs can be incorrect or too generic, leading to false positives if used directly. Stage 2 addresses this by implementing an execution feedback checking method to validate the correctness of these raw APSRs. The core idea is that a truly incorrect API parameter usage should ideally cause an observable runtime error.
However, a critical challenge arises: a runtime error during the execution of "violation code" might not necessarily be related to the specific raw APSR being validated. For instance, if an APSR suggests "Parameter 1 must be an existing database file" for SQL_reopen, but the generated violation code also passes a NULL for Parameter 2, a resulting null pointer dereference error is not evidence that the "existing database file" rule is correct.
To overcome this, GBT8 employs a two-step code generation process:
- Right Code Generation: GBT8 first instructs the LLM to generate "right code" – code that calls the target API successfully without any violations. To enhance the success rate, if the initial right code fails to execute, GBT8 provides the execution results and the faulty code back to the LLM, instructing it to fix the code until successful execution is achieved. This ensures a stable baseline for comparison.
- Violation Code Generation: Once the right code is established, GBT8 provides this code, the raw APSR under validation, and instructions to the LLM to modify the right code to specifically violate the target raw APSR.
- Cost Check: Finally, GBT8 performs a "cost check" based on predefined rules. This check compares the changes made to the right code to generate the violation code against the raw APSR, ensuring consistency. The goal is to confirm that the generated violation code only introduces the specific misuse described by the APSR and no other unrelated errors. (Details of these predefined rules are elaborated in the research paper).
After generating and validating the violation code, GBT8 executes it. If the violation code runs successfully, it indicates that the raw APSR is likely incorrect (e.g., free(NULL) running successfully indicates "parameter must not be null" for free is an incorrect APSR). If the violation code causes a runtime error that is consistent with the APSR (e.g., a double-free error for an APSR about valid memory pointers), then the raw APSR is deemed correct. This feedback loop is essential for filtering out spurious rules.
Stage 3: APSR Refinement
Even validated raw APSRs might still be too general for practical detection (e.g., "parameter must be valid"). Stage 3 focuses on refining these into concrete, actionable rules. This involves identifying the key operations within the code that lead to specific bugs.
The challenge here is that code changes can involve many operations, making it difficult for an LLM to pinpoint the exact "key operation" causing a specific bug. GBT8 addresses this with another crucial observation: bugs caused by similar key operations consistently result in the same runtime error messages.
Based on this, GBT8 implements the following:
- Error Message Clustering: GBT8 first clusters all generated violation codes that produce the same runtime error messages. This groups together instances of misuse that likely stem from similar underlying causes.
- Common Operation Analysis: For each cluster, GBT8 guides the LLM to analyze the common operations present in the violation codes. By providing the runtime error message in the prompt, GBT8 assists the LLM in focusing its analysis on the operations most likely responsible for that specific error. This differential analysis helps the LLM isolate the "key operation."
- Concrete APSR Generation: Once the key operation is identified, GBT8 instructs the LLM to describe this operation in a concrete manner, thereby generating a refined, actionable APSR. For example, if the key operation identified from a cluster of "double-free" errors is related to passing a previously freed pointer, the LLM might refine a vague "parameter must be valid" rule to the concrete "parameter must not be freed twice." The talk specifically mentions refining a raw APSR to "must not be null."
This iterative process, guided by execution feedback and focused analysis of code changes, ensures that the final APSRs are not only correct but also specific enough to be translated into effective detection code.
Stage 4: API Misuse Detection
The final stage of GBT8 is to use the generated, validated, and refined APSRs to detect actual API misuses in target codebases. For this, GBT8 adopts a methodology inspired by previous work, specifically referencing the Advance approach.
The process involves:
- Pattern Identification: The researchers manually analyzed the generated concrete APSRs to identify frequently occurring patterns of misuse.
- Detection Code Templates: Based on these patterns, generic detection code templates are created. These templates represent common ways to check for specific types of APSR violations.
- Automated Code Generation: GBT8 then automatically combines these templates and instantiates them with the details from the concrete APSRs to generate the final, specific detection code. This detection code can then be applied to scan target software projects for violations of the generated rules.
This stage effectively operationalizes the output of the LLM-driven rule generation, transforming abstract security rules into executable checks that can identify real-world vulnerabilities.
Demo / Proof of Concept
▶ Watch: Refining rules by identifying key operations from errors (8:00)
While the talk did not feature a live, interactive demonstration of the GBT8 framework, the comprehensive evaluation serves as the primary proof of concept for its efficacy. The researchers rigorously tested GBT8 on eight popular libraries, showcasing its ability to generate high-quality APSRs and detect numerous misuses. The reported statistics—generating 8 types of APSRs with 92.3% precision, detecting 2510 new misuses, and generating 355 new APSRs—quantitatively demonstrate GBT8's practical applicability and superior performance compared to existing methods. The detailed examples provided throughout the presentation, such as the free(NULL) scenario and the SQL_reopen example, illustrate the framework's internal workings and its capacity to handle complex validation and refinement challenges.
Defensive Implications
▶ Watch: Example of a concrete refined API security rule (10:00)
The findings presented by GBT8 have profound implications for software developers, security teams, and organizations striving to build more secure applications. The ability to automatically generate accurate and concrete API parameter security rules (APSRs) offers a proactive and scalable defense mechanism against a pervasive class of vulnerabilities.
Firstly, GBT8 provides a powerful tool for proactive vulnerability discovery during the development lifecycle. By integrating GBT8-like analysis into Continuous Integration/Continuous Deployment (CI/CD) pipelines, developers can automatically scan their code for API misuses as part of every commit or build. This shifts security left, enabling the identification and remediation of issues long before they reach production, significantly reducing the cost and impact of security flaws.
Secondly, security teams can leverage GBT8 to audit existing codebases more effectively. Traditional static analysis tools often struggle with the nuanced and context-dependent nature of API misuse. GBT8's LLM-driven approach, validated by execution feedback, can uncover a broader spectrum of misuses, including those that might be missed by predefined rule sets. This is particularly valuable for large, legacy systems that rely heavily on complex APIs.
Thirdly, the concrete APSRs generated by GBT8 can serve as valuable developer education resources. By understanding the precise conditions under which an API is misused and the specific runtime errors that result, developers can gain a deeper appreciation for API contracts and write more secure code from the outset. These rules can be integrated into documentation, code review checklists, or even IDE plugins to provide immediate feedback.
Furthermore, GBT8's methodology highlights the importance of runtime validation in security analysis. The execution feedback checking method demonstrates that empirical observation of code behavior is crucial for validating the correctness of security rules, especially those derived from complex models like LLMs. Defenders should consider incorporating similar dynamic analysis components into their security testing strategies.
Finally, the insights into prompt engineering for LLMs are invaluable for the broader security community. As LLMs become more prevalent in security tasks, understanding how to effectively query and guide them to avoid fabrication or limited outputs will be critical for developing reliable security tools across various domains beyond API misuse detection. While GBT8 currently faces limitations in detecting silent memory corruption bugs, its foundation offers a path towards more comprehensive and intelligent security analysis.
Key Takeaways
- LLMs are powerful for APSR generation: Large Language Models can effectively analyze API source code to generate initial API Parameter Security Rules (APSRs), surpassing the limitations of predefined code analysis rules.
- Execution feedback is critical for validation: An innovative "execution feedback checking method" validates LLM-generated APSRs by executing carefully crafted violation code and observing runtime errors, significantly improving rule accuracy and reducing false positives.
- Vague rules require refinement: LLM-generated APSRs, initially vague, can be refined into concrete, actionable rules by guiding the LLM to identify "key operations" causing specific runtime errors through clustered code analysis.
- GBT8 outperforms state-of-the-art: The GBT8 framework demonstrated superior performance, detecting 2510 new API misuses and generating 355 new APSRs with 92.3% precision across eight popular libraries, significantly improving coverage over existing methods.
- Prompt design matters for LLM security tools: Effective prompt engineering is crucial; providing insufficient information can lead to fabricated results, while overly specific examples can inadvertently limit the LLM's ability to generate comprehensive outputs.
- Limitations and future work: GBT8 currently relies on observable runtime errors, meaning silent memory corruption bugs that do not immediately manifest as errors will be missed, highlighting an area for future research, potentially involving sanitizers.
About the Speaker(s)
Jinghua Liu is a researcher from the Institute of Information Engineering, Chinese Academy of Sciences. The work presented on GBT8 was a collaborative effort, with co-authors Eong K and Menlinin also contributing to the research, all affiliated with the same institute. Their work focuses on leveraging advanced computational techniques, particularly Large Language Models, to address complex challenges in software security, specifically in the domain of API misuse detection.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic systems-security research with a sensible core idea: use LLMs to generate API parameter security rules, validate them with execution feedback, and refine vague outputs into concrete detectors. The engineering is honest and the numbers are real, but this is an NDSS paper presentation, not a conference talk that will change how practitioners work tomorrow.
Heather Calloway (CISO) — WEAK
Technically solid academic work on automated API misuse detection using LLMs, with credible methodology and real numbers. But it never crosses the threshold into defender or leadership relevance — no organizational accountability angle, no guidance on where this fits in a security program, and the 'defensive implications' section reads like a grant proposal rather than operational advice.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025