LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Saad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse Kivilcim Coskun, Gianluca Stringhini

IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 5

Overview

The proliferation of large language models (LLMs) as general-purpose assistants has led to their increasing deployment in various automated cybersecurity tasks, including vulnerability analysis, repair, and software test generation. However, despite their widespread adoption, a comprehensive understanding of their capabilities, particularly in vulnerability detection and root cause analysis, has remained largely unexplored. Previous evaluations often focused on smaller LLMs, limited to binary classification (vulnerable/not vulnerable), or concentrated on tasks like insecure code generation or vulnerability repair, leaving a critical gap in assessing their deeper reasoning abilities.

Watch on YouTube

Visual summary for LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks by Saad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse Kivilcim Coskun, Gianluca Stringhini
Visual summary for LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks by Saad Ullah, Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse Kivilcim Coskun, Gianluca Stringhini

Key moments

  1. 0:50 Problem: LLMs' security reasoning capabilities unexplored
  2. 1:10 Introducing SEC-LLM-Homes evaluation framework
  3. 2:00 LLM configuration and benchmark dataset details
  4. 4:00 How SEC-LLM-Homes evaluates LLM responses
  5. 6:00 Evaluation findings: optimal parameters (temperature, top P)
  6. 6:50 Evaluation findings: diverse optimal prompting strategies
  7. 7:30 Demonstrating effective step-by-step security reasoning prompt

LLMs Cannot Reliably Identify and Reason About Security Vulnerabilities (Yet?): A Comprehensive Evaluation, Framework, and Benchmarks

Speakers: Saad Ullah; Mingji Han; Saurabh Pujar; Hammond Pearce; Ayse Kivilcim Coskun; Gianluca Stringhini

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=K7dhqrxLsYY

Overview

The proliferation of large language models (LLMs) as general-purpose assistants has led to their increasing deployment in various automated cybersecurity tasks, including vulnerability analysis, repair, and software test generation. However, despite their widespread adoption, a comprehensive understanding of their capabilities, particularly in vulnerability detection and root cause analysis, has remained largely unexplored. Previous evaluations often focused on smaller LLMs, limited to binary classification (vulnerable/not vulnerable), or concentrated on tasks like insecure code generation or vulnerability repair, leaving a critical gap in assessing their deeper reasoning abilities.

This talk by Saad Ullah and his co-authors introduces a novel, fully automated, end-to-end framework named SEC-LLM-HOMES designed to benchmark and evaluate LLMs for vulnerability detection. Unlike prior studies, SEC-LLM-HOMES assesses LLMs across a diverse set of dimensions and introduces a unique automated approach to evaluate their reasoning or root cause analysis capabilities. The research aims to definitively answer whether LLMs can serve as truly helpful and reliable security assistants, providing insights into their strengths, weaknesses, and fundamental limitations in understanding complex security flaws.

The significance of this work lies in its rigorous methodology and the critical insights it provides into the current state of LLM security reasoning. By revealing the inconsistencies, biases, and vulnerabilities of LLMs themselves when tasked with security analysis, the research offers a crucial reality check for the cybersecurity community. It underscores the need for caution in deploying LLMs for sensitive security tasks and highlights essential areas for future research and development to enhance their reliability and analytical depth.

Background

▶ Watch: Problem: LLMs' security reasoning capabilities unexplored (0:50)

The rapid advancements in large language models have positioned them as powerful tools across numerous domains, with cybersecurity being a particularly active area of application. Their ability to process and generate human-like text has made them attractive for automating laborious and complex tasks, such as identifying and fixing software vulnerabilities. Initial explorations leveraged LLMs for generating insecure code samples, aiding in the creation of vulnerable test cases, or assisting in the repair of known vulnerabilities by suggesting code patches.

However, the evaluation landscape for LLMs in cybersecurity has been fragmented and often incomplete. Many studies focused on smaller models, which may not represent the full capabilities of state-of-the-art LLMs like GPT-4. Furthermore, a significant portion of the evaluations concentrated on binary classification—simply determining if a piece of code is vulnerable or not. This approach, while useful, entirely overlooks the more critical aspect of root cause analysis—understanding why a piece of code is vulnerable and identifying the specific flaw that leads to the security issue. Without robust root cause analysis capabilities, LLMs cannot provide the actionable intelligence required by security professionals to effectively mitigate risks.

The lack of comprehensive evaluation, particularly concerning the deeper reasoning capabilities of LLMs, left a fundamental question unanswered: Can these advanced models truly function as reliable security assistants, capable of not just flagging potential issues but also providing detailed, accurate explanations? This gap in understanding prompted the development of SEC-LLM-HOMES, aiming to move beyond superficial evaluations and delve into the nuanced and complex world of security vulnerability reasoning. The framework addresses the need for a standardized, automated, and multi-dimensional approach to assess LLMs' proficiency in this critical domain, setting new benchmarks for their performance and revealing their inherent limitations.

Key Findings

▶ Watch: LLM configuration and benchmark dataset details (2:00)

The evaluation conducted using SEC-LLM-HOMES yielded several critical findings regarding the capabilities and limitations of state-of-the-art LLMs in identifying and reasoning about security vulnerabilities.

Firstly, the research revealed that LLM performance is highly sensitive to prompting strategies and internal parameters. For vulnerability detection, the most accurate and deterministic responses were observed at temperature zero and top P 1. Higher temperature values consistently led to more inaccurate reasoning and answers. Interestingly, even at temperature zero, GPT-4 exhibited non-deterministic behavior, frequently changing both its reasoning and final answer when presented with the same prompt multiple times, highlighting a fundamental challenge in reproducibility and reliability.

Secondly, there is no universal "best prompt" for all LLMs. Smaller models generally performed better when provided with explicit definitions of vulnerabilities. In contrast, GPT-based models and the code-tuned Chat Bison model achieved higher accuracy using step-by-step reasoning prompts. Crucially, this wasn't merely asking the model to "think step by step," but rather providing specific instructions designed to mimic the vulnerability analysis process followed by human security experts, based on existing literature. This suggests that LLMs' inherent reasoning for security vulnerabilities is often not aligned with real-world security practices unless explicitly guided.

A significant discovery was the phenomenon of unfaithful reasoning. This manifests in two ways: an LLM might provide the correct final answer (e.g., "vulnerable") but base it on entirely incorrect reasoning, or conversely, it might offer accurate reasoning for why a vulnerability exists but then conclude incorrectly that the code is "safe." This unfaithful reasoning poses a severe risk, as security professionals relying on LLM outputs could be misled by seemingly correct answers that lack a sound analytical foundation.

The study also uncovered biases in LLMs' knowledge base. Not all LLMs possess comprehensive knowledge about diverse vulnerability types; many were biased towards out-of-bound write vulnerabilities. Furthermore, specific models like Code Chat Bison struggled significantly with critical weaknesses such as out-of-bound write and use-after-free vulnerabilities, particularly in "two-shot" scenarios. These are consistently ranked among the top most dangerous software weaknesses, appearing on lists like the 2023 CWE Top 25 Most Dangerous Software Weaknesses, indicating a concerning blind spot for LLMs in high-impact areas.

LLMs also demonstrated difficulty handling complex code scenarios. They performed poorly with lengthy codebases and those exhibiting high functional density. In real-world CVE scenarios, the evaluation observed a remarkably high false positive rate and an increase in "doubtful answers," where models explicitly stated their inability to assist.

Finally, the robustness of LLMs was critically examined through code augmentations. The researchers found that seemingly trivial changes to code, such as adding extra spaces (e.g., after line 22 in a specific test case) or providing suggestive function names (e.g., strcat for out-of-bound or escape for XSS), could break the chain of thought and step-by-step reasoning of LLMs. This led to incorrect conclusions in a significant portion of test cases, ranging from 17% to 30%. This highlights a severe lack of robustness, where superficial alterations can drastically alter an LLM's security assessment, demonstrating their susceptibility to adversarial inputs or even unintentional formatting changes.

Technical Deep Dive

▶ Watch: How SEC-LLM-Homes evaluates LLM responses (4:00)

The core of this research is the SEC-LLM-HOMES framework, an automated, end-to-end evaluation system designed to systematically benchmark LLMs for vulnerability detection and reasoning. The framework is modular, comprising three main components: the Prompt Creation Module, the LLM Configuration Module, and the Evaluator Module.

The LLM Configuration Module is designed for versatility, allowing users to integrate and benchmark virtually any LLM with minimal effort. It follows a simple, three-step process, configurable in approximately 30 lines of code. This involves:

  1. Defining Prompting Strategies: Recognizing that different models are instruction-tuned in unique ways, this step allows users to specify how to optimally format prompts for a given LLM. For instance, OpenAI models might require test content enclosed in three quotation marks, while Google models benefit from semantic keywords preceding the content.
  2. Initializing the Model: This involves setting up the LLM instance, including API keys or local model paths.
  3. Integrating the Chat Inference API/Code: This step connects the framework to the LLM's inference capabilities, allowing prompts to be sent and responses received. An adapter for GPT-based models was built as an example of this integration.

Once an LLM is configured, the Prompt Creation Module takes over. It leverages a robust benchmark dataset composed of:

  • Handcrafted scenarios: Specifically designed to test various vulnerability types and edge cases.
  • Augmented scenarios: Modifications of existing vulnerabilities to test robustness.
  • Real-world CVE scenarios: Actual vulnerabilities from published CVEs, ensuring practical relevance.

Crucially, all scenarios in the dataset are ensured to be outside the knowledge cut-off date of the evaluated LLMs, preventing models from simply recalling known information. These scenarios are further categorized by three different code complexity levels (e.g., simple, moderate, complex) to assess performance across varying code structures.

The module combines these scenarios with an extensive set of prompt templates across three categories:

  1. Standard Prompts: Direct questions asking if the code is vulnerable.
  2. Step-by-Step Reasoning Prompts: Instructions designed to guide the LLM through a structured analysis process, mimicking human security expert methodologies. This is more sophisticated than a generic "think step-by-step" prompt, drawing from literature on how security experts perform vulnerability analysis.
  3. Vulnerability Definition Prompts: Prompts that provide explicit definitions of specific vulnerabilities (e.g., what an out-of-bound write is) to aid the LLM.

These combinations allow for evaluations across eight key dimensions, including identifying optimal parameters (like temperature and top P), determining the most effective prompts, stress-testing LLMs over diverse code scenarios and difficulty levels, assessing knowledge of various vulnerability types, handling different code complexities, evaluating performance on real-world cases, and testing robustness against code augmentations.

The Evaluator Module is central to assessing LLM responses, particularly their reasoning. When an LLM generates a response, it typically contains two key pieces of information: the final answer (vulnerable/not) and the reasoning behind it. The evaluator leverages GPT-4 itself to extract these two specific fields from the LLM's raw output.

  • The extracted answer (yes/no) is compared against a pre-defined ground truth answer for that scenario to calculate accuracy.
  • The extracted reasoning is the more complex part. It is compared against a ground truth reasoning statement, meticulously designed by three independent security experts. This comparison is performed using Natural Language Processing (NLP) metrics (e.g., BLEU, ROUGE scores, or semantic similarity measures, though specific metrics aren't detailed in the transcript, the use of "NLP metrics" implies quantitative comparison) to determine the alignment between the LLM's reasoning and expert consensus. This innovative approach allows for an automated, objective evaluation of an LLM's root cause analysis capabilities.

A concrete example of the specialized step-by-step prompting involved a null pointer dereference vulnerability. In this scenario, code attempts to read from a file opened with fopen without checking if fopen returned a NULL pointer (indicating failure). By providing the LLM with instructions to "mimic the process from this study" (referring to a previous human vulnerability analysis study), the model successfully identified the vulnerability and provided the correct reasoning: "F open function can return null and the code does not check it." This demonstrated that tailored, expert-informed prompting is crucial for eliciting accurate reasoning from LLMs in security contexts.

The robustness testing involved targeted code augmentations. For instance, a simple out-of-bound write vulnerability in a code snippet, which Chat Bison consistently identified as vulnerable, would be deemed "safe" by the same model if a few extra spaces were added after a specific line (e.g., line 22). Other augmentations included changing variable names or introducing "dummy" function calls that might bias the model. For example, in a cross-site scripting (XSS) context, adding a function call like escape() might lead the model to assume the code is safe, even if the actual XSS vulnerability remains. These tests revealed that LLMs often focus on superficial cues rather than deeply understanding the semantic implications of the code, making them brittle and easily misled.

Demo / Proof of Concept

▶ Watch: Evaluation findings: diverse optimal prompting strategies (6:50)

While the talk did not feature a live, interactive demonstration in the traditional sense, the presentation effectively showcased the capabilities of the SEC-LLM-HOMES framework and the subsequent performance of LLMs through a series of illustrative code scenarios and their corresponding evaluation results. These examples served as compelling proofs of concept for the framework's methodology and the critical findings derived from it.

One key demonstration involved an out-of-bound write vulnerability. The example code snippet was designed to encode a user input character, assuming a maximum expansion factor of four bytes. However, the ampersand character encoding expands by five bytes. An attacker providing a string of many ampersand characters would cause a buffer overflow, leading to an out-of-bound write. The framework demonstrated how it would present this code to various LLMs using different prompts, then extract both the "vulnerable/not vulnerable" answer and the detailed reasoning. This showed the system's ability to assess both the final decision and the underlying analytical process.

Another scenario illustrated a null pointer dereference. The code in question opened a file but failed to check if the fopen function returned a NULL pointer, which would indicate that the file could not be opened. Subsequently, the code proceeded to read from this potentially null pointer, leading to a crash. This example was used to demonstrate the effectiveness of the specialized "step-by-step reasoning" prompts. When guided by instructions mimicking human security expert analysis, the LLM not only correctly identified the vulnerability but also provided precise reasoning: "F open function can return null and the code does not check it." This highlighted how tailored prompting, rather than generic instructions, could significantly improve an LLM's analytical depth.

The concept of unfaithful reasoning was demonstrated with a command injection vulnerability. A piece of code was designed to validate user input against command injection, but it only checked for semicolons (;). An LLM, when presented with this code, correctly reasoned that other characters like $ (dollar sign), | (pipe sign), or & (ampersand sign) could also be used to concatenate commands, and the code failed to check for these. Despite this accurate and insightful reasoning, the model inexplicably concluded that the code was "safe." This example powerfully illustrated the disconnect between an LLM's internal reasoning process and its final decision, a critical flaw for security applications.

Finally, the robustness tests provided a stark demonstration of LLM brittleness. A trivial code augmentation example showed that adding a few extra spaces after line 22 in the aforementioned out-of-bound write code snippet caused Chat Bison to change its assessment from "vulnerable" to "safe." Similarly, adding seemingly benign function calls like strcat (for out-of-bound scenarios) or escape (for XSS) could bias LLMs, causing them to focus on these functions and make incorrect security judgments. These examples concretely demonstrated how superficial changes, even those unrelated to the core logic, could drastically alter an LLM's security analysis, exposing a significant vulnerability in their reasoning capabilities. Through these specific and well-articulated scenarios, the talk effectively demonstrated the comprehensive evaluation power of SEC-LLM-HOMES and the often-unreliable nature of current LLMs in security vulnerability analysis.

Defensive Implications

▶ Watch: Demonstrating effective step-by-step security reasoning prompt (7:30)

The findings from the SEC-LLM-HOMES evaluation carry significant defensive implications for organizations and security professionals considering or currently employing LLMs in their cybersecurity workflows. The primary takeaway is a strong caution against blindly trusting LLMs for critical vulnerability detection and root cause analysis tasks.

  1. Human Oversight is Indispensable: LLMs, in their current state, cannot replace human security experts. Their susceptibility to unfaithful reasoning, non-deterministic behavior, and fragility to trivial code augmentations means that any LLM-generated security assessment must be thoroughly reviewed and validated by human analysts. Relying solely on LLM output could lead to missed critical vulnerabilities or misprioritized remediation efforts.
  1. Understand LLM Limitations and Biases: Defenders must be aware of the specific weaknesses identified. LLMs exhibit biases towards certain vulnerability types (e.g., out-of-bound write) and may completely miss others, even highly critical ones like use-after-free. They struggle with complex, lengthy code and show high false positive rates in real-world scenarios. This necessitates a diversified approach to vulnerability scanning, combining LLM-based tools with traditional static analysis, dynamic analysis, and manual code reviews.
  1. Careful Prompt Engineering is Crucial: The research highlights that the way questions are posed to an LLM drastically impacts its performance. Generic prompts are insufficient. Organizations should invest in developing and refining expert-informed prompting strategies that guide LLMs through structured security analysis processes, mimicking human methodologies. However, even with optimal prompting, the issue of unfaithful reasoning persists, underscoring the need for validation.
  1. Robustness Testing for LLM Deployments: Before integrating LLMs into any security pipeline, organizations should conduct rigorous robustness testing. This includes evaluating the LLM's performance against various code augmentations—even seemingly trivial ones like adding spaces or renaming variables—to understand how easily its reasoning can be disrupted. Frameworks like SEC-LLM-HOMES can be adopted for internal benchmarking of LLMs against an organization's specific codebase and threat model.
  1. LLMs as Augmentation, Not Automation: Instead of full automation, LLMs should be viewed as powerful augmentation tools. They can assist in generating initial vulnerability reports, highlighting suspicious code patterns, or even suggesting potential fixes. However, the final decision-making, in-depth root cause analysis, and verification of remediation should remain with human experts. They can help scale initial triage, but not replace the deep contextual understanding of a security researcher.
  1. Stay Informed on LLM Advancements: The field of LLM development is rapidly evolving. Defenders should continuously monitor research like this to stay updated on the latest capabilities and, more importantly, the persistent limitations of these models. This ongoing awareness is crucial for making informed decisions about their safe and effective deployment in cybersecurity.

In essence, while LLMs offer tantalizing possibilities for enhancing cybersecurity, the current reality presented by this research dictates a cautious, human-centric approach. Their integration should be strategic, with a clear understanding of their fallibility and an emphasis on robust validation processes to prevent them from becoming a new source of security risk.

Key Takeaways

  • LLMs are Not Yet Reliable for Comprehensive Vulnerability Analysis: Current large language models struggle significantly with accurately identifying security vulnerabilities and, more critically, providing consistent and faithful root cause analysis, even for well-known weaknesses.
  • SEC-LLM-HOMES Provides an Automated Evaluation Framework: The novel, end-to-end SEC-LLM-HOMES framework offers a robust and automated method to benchmark LLMs across diverse dimensions, including their reasoning capabilities using NLP metrics against expert ground truth.
  • Prompting and Parameters are Critical but Insufficient: LLM performance is highly sensitive to parameters like temperature (optimal at zero) and specific prompting strategies (e.g., step-by-step instructions mimicking human experts), but even optimal settings do not guarantee reliable or deterministic output, especially for models like GPT-4.
  • "Unfaithful Reasoning" is a Major Concern: LLMs frequently exhibit unfaithful reasoning, where a correct answer is given with incorrect justification, or accurate reasoning leads to a wrong final conclusion, making their outputs untrustworthy for critical security decisions.
  • LLMs Suffer from Biases, Complexity Issues, and Brittleness: Models show biases towards certain vulnerability types (e.g., out-of-bound write), struggle with lengthy or functionally dense code, and are highly susceptible to trivial code augmentations (e.g., extra spaces) that can completely derail their analysis.
  • Human Expertise Remains Indispensable: Given the current limitations, LLMs should be treated as assistive tools rather than autonomous agents for vulnerability detection. Human security experts remain essential for critical analysis, validation, and decision-making to ensure reliable and robust security outcomes.

About the Speaker(s)

The primary presenter for this work was Saad Ullah, who is a third-year PhD candidate at Boston University. He delivered the detailed findings and methodology of the SEC-LLM-HOMES framework. The research was a collaborative effort with co-authors Mingji Han, Saurabh Pujar, Hammond Pearce, Ayse Kivilcim Coskun, and Gianluca Stringhini. Their collective expertise contributed to the comprehensive evaluation and analysis of large language models in the context of security vulnerability identification and reasoning.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research delivers a much-needed reality check on LLMs for vulnerability analysis, introducing a novel, automated framework (SEC-LLM-HOMES) to rigorously benchmark their reasoning. It unequivocally demonstrates that current LLMs are unreliable, prone to "unfaithful reasoning," and easily fooled by trivial code changes. This is critical signal for anyone deploying AI in security.

Heather Calloway (CISO) — STRONG ACCEPT

This research provides a vital reality check on LLM capabilities for vulnerability analysis. It definitively shows current models are unreliable, prone to "unfaithful reasoning," and brittle, underscoring the critical need for robust human oversight and validation in any security application.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024