Large Language Models for Code Analysis: Do LLMs Really Do Their Job?
Chongzhou Fang, Jialin Liu, Ruoyu Zhang, Han Wang, Houman Homayoun
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
This talk, presented by Chongzhou Fang, a fourth-year PhD student at UC Davis, delves into a critical and timely evaluation of Large Language Models (LLMs) for code analysis. With the explosive growth of LLMs and their increasing application in software development, particularly for code generation, a thorough understanding of their capabilities in analyzing and explaining existing code — especially obfuscated code — has become paramount. The research addresses a significant gap in the literature, providing the first comprehensive assessment of how well state-of-the-art LLMs perform these tasks.

Key moments
- 2:00 LLMs for code analysis: identifying the research gap
- 2:50 Two key research questions and LLMs evaluated
- 4:00 Building diverse code datasets: normal and obfuscated
- 6:30 Establishing ground truth via GPT-4 and human validation
- 8:20 GPT-4's high analysis accuracy and human-like insights
Large Language Models for Code Analysis: Do LLMs Really Do Their Job?
Speakers: Chongzhou Fang, Jialin Liu, Ruoyu Zhang, Han Wang, Houman Homayoun
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=lgRGPYz2woY
Overview
This talk, presented by Chongzhou Fang, a fourth-year PhD student at UC Davis, delves into a critical and timely evaluation of Large Language Models (LLMs) for code analysis. With the explosive growth of LLMs and their increasing application in software development, particularly for code generation, a thorough understanding of their capabilities in analyzing and explaining existing code — especially obfuscated code — has become paramount. The research addresses a significant gap in the literature, providing the first comprehensive assessment of how well state-of-the-art LLMs perform these tasks.
The core of the paper investigates two primary research questions: first, whether LLMs genuinely comprehend and can analyze standard, non-obfuscated code, and second, their performance when confronted with code deliberately made difficult to understand through obfuscation techniques. The findings offer crucial insights for both developers and security professionals, highlighting the strengths and, more importantly, the significant limitations of current LLMs in handling complex and adversarial code scenarios. This work serves as a foundational benchmark, setting expectations for the practical application of LLMs in areas like malware analysis, intellectual property protection, and automated code review.
The implications of this research are far-reaching. As software systems grow in complexity and malicious actors increasingly employ sophisticated obfuscation to evade detection, the ability of automated tools to accurately analyze code is vital. By rigorously evaluating models like GPT-3.5, GPT-4, Llama, and Code Llama against diverse code samples, including those from the International Obfuscated C Code Contest (IOCCC), the authors provide a realistic perspective on the current "intelligence" of LLMs in understanding code semantics and functionality, moving beyond the often-hyped capabilities seen in simpler code generation tasks.
Background
▶ Watch: LLMs for code analysis: identifying the research gap (2:00)
The landscape of software development is increasingly dominated by large, intricate systems, making the need for automated code analysis tools more pressing than ever. Code analysis, in this context, refers to the process of examining source code to understand its functionality and generate clear explanations. Such tools are invaluable for assisting human developers in maintaining, debugging, and understanding complex codebases.
Alongside the growth of software complexity, the technique of code obfuscation has become a dual-edged sword. Obfuscation transforms a piece of code into a functionally equivalent but significantly harder-to-understand form. While legitimate uses include protecting intellectual property by making reverse engineering difficult, obfuscation is also a favored tactic among malicious software writers to conceal their harmful intent, making malware analysis a formidable challenge. An example provided in the talk illustrates this vividly: a simple "Hello World" JavaScript snippet can become utterly unreadable after common obfuscation transformations, highlighting the difficulty even for human analysts.
The recent surge in popularity of Large Language Models (LLMs), such as OpenAI's GPT series, has led to a broad range of applications across various domains. A particularly promising area is their application to code-related tasks. Commercial successes like GitHub Copilot and Amazon CodeWhisperer demonstrate LLMs' prowess in code generation, assisting developers in writing new code more efficiently. However, as noted by the speakers, much of the initial effort and evaluation work surrounding LLMs in the software engineering domain has focused on generation. There has been a notable lack of thorough, systematic evaluation of LLMs' capabilities in code analysis, especially when confronted with the challenges posed by obfuscated code. This deficiency represents a critical gap in understanding their true utility and limitations for security-sensitive applications.
This paper provides a timely response to this gap, aiming to rigorously assess whether LLMs can reliably perform code analysis and, crucially, how they fare against code that has been deliberately obscured. By doing so, it contributes vital information to the ongoing discourse about the role of artificial intelligence in software security and development, moving beyond anecdotal evidence to a data-driven evaluation of LLM performance.
Key Findings
▶ Watch: Two key research questions and LLMs evaluated (2:50)
The research set out to answer two fundamental questions: "Do large language models understand code?" and "Can large language models understand obfuscated code?" The findings reveal a nuanced picture, demonstrating both impressive capabilities from the most advanced models and significant limitations, particularly for smaller models and in the face of sophisticated obfuscation.
For non-obfuscated code analysis, the GPT-4 model exhibited remarkably high accuracy, achieving an overall score of approximately 97.5%. Its performance was consistently strong across the three evaluated languages (C, JavaScript, and Python). An interesting observation was that GPT-4, much like human programmers, appeared to utilize information embedded in identifier names to aid its analysis, suggesting that well-written, descriptive code contributes to better analysis quality. However, even GPT-4 was not immune to occasional errors, demonstrating instances of hallucination or making "wrong associations." For example, it once incorrectly stated that a Python script imported the Pandas library when it did not, likely due to a statistical association with other data analysis libraries (Numpy, Matplotlib) commonly found in its training data.
In stark contrast, smaller LLMs such as Llama, Code Llama, and Starred Beta performed exceptionally poorly. In many cases, these models either simply returned the original code without any analysis or provided a rephrased version of the prompt. Even after filtering out these non-analytical responses, the quality of their generated analysis was largely incorrect. The researchers concluded that these smaller models are currently "unable to reliably generate consistent paragraphs of code analysis results."
When confronted with obfuscated code, the performance of the GPT models showed a clear dependency on the complexity of the obfuscation technique. They performed "quite well" against code obfuscated with classic and basic techniques such as dead code injection and control flow flattening. The analysis accuracy did not drop significantly in these scenarios. However, the models struggled dramatically with advanced obfuscation methods. Specifically, when faced with Wcor, a WebAssembly-based obfuscation method, GPT models became "totally unusable," indicating a significant vulnerability to modern, sophisticated obfuscation.
Further challenging the GPT models, the researchers tasked them with de-obfuscation, asking them to generate a de-obfuscated version of highly flexible and non-standard obfuscated code samples sourced from the International Obfuscated C Code Contest (IOCCC). Both GPT-3.5 and GPT-4 performed "very poorly" in this task. GPT-4 failed to generate any results for approximately a quarter of the samples, and most of its generated code was uncompilable, with only one or two samples successfully executing and producing the correct output. GPT-3.5 fared only slightly better in terms of compilability, but a deeper examination revealed that GPT-4's de-obfuscated code, despite being less compilable, was of higher quality in terms of readability, featuring more meaningful identifier names and a clearer structure. This suggests GPT-4 attempts a more semantic de-obfuscation, while GPT-3.5 might preserve more of the original obfuscated structure, leading to higher compilability but lower readability.
In summary, while powerful LLMs like GPT-4 demonstrate strong capabilities in analyzing straightforward code and even basic obfuscation, they fall short when encountering advanced obfuscation techniques or complex de-obfuscation challenges. Smaller models are not yet viable for reliable code analysis.
Technical Deep Dive
▶ Watch: Building diverse code datasets: normal and obfuscated (4:00)
The rigorous evaluation conducted in this research involved a carefully constructed experimental setup, encompassing a diverse range of Large Language Models (LLMs), a comprehensive dataset, and robust evaluation metrics.
The study covered a wide spectrum of LLMs available at the time of submission. These included the commercially advanced OpenAI GPT series, specifically GPT-3.5 and GPT-4, which served as benchmarks for cutting-edge capabilities. To assess the performance of smaller, more accessible models, the researchers also deployed Llama, Code Llama, and Starred Beta. This selection allowed for a comparative analysis between proprietary, large-scale models and their more compact, often open-source, counterparts.
A crucial aspect of the research was the construction of a two-part data set:
- Non-obfuscated Code Data Set: This part comprised code samples in three popular programming languages: C, JavaScript, and Python. These languages were chosen for their prevalence on GitHub and their representation of different programming paradigms, from low-level system programming (C) to high-level scripting (JavaScript, Python). The data set was designed to be diverse in code size, ranging from just a few lines of code to "extremely large code files that have over 10,000 lines of code," ensuring that the models were tested across varying levels of complexity. Before analysis, all comments were systematically removed from these code files to prevent LLMs from simply extracting existing explanations rather than performing genuine analysis.
- Obfuscated Code Data Set: This data set was further divided into two categories:
- JavaScript-based Obfuscated Code: This was derived from the JavaScript subset of the non-obfuscated data set. Various obfuscation techniques were applied, including "very basic and classic obfuscation techniques" like dead code injection and control flow flattening. Crucially, it also incorporated more "advanced obfuscation methods" such as Wcor, which is a WebAssembly-based obfuscation method. This allowed for a direct comparison of LLM performance against a spectrum of obfuscation sophistication.
- IOCCC Code Samples: To test the LLMs against highly flexible and non-standard obfuscation, a subset of winning code samples from the International Obfuscated C Code Contest (IOCCC) was included. The IOCCC is renowned for challenging participants to create the most confusing and unreadable C code, providing a stringent test case for LLM robustness.
A significant methodological challenge was establishing the ground truth for code analysis results, as readily available, perfectly aligned explanations for such a diverse dataset do not exist. The researchers adopted a pragmatic and robust approach:
- They leveraged GPT-4 to generate initial analysis results for all code samples.
- These GPT-4 generated analyses were then subjected to human validation by "several experienced programmers" with "over five years of programming experience." These experts meticulously reviewed whether the generated analysis accurately matched the original functionality of the code. Only analyses marked as correct by human validators were then used as the ground truth. This "human-in-the-loop" validation process ensured the reliability of the baseline against which other models' performances were measured.
To quantitatively compare the analysis results generated by different LLMs against this human-validated ground truth, the researchers employed several evaluation metrics:
- Cosine Similarity: A common metric for measuring the similarity between two non-zero vectors of an inner product space.
- BERT-based Semantic Similarity: Utilizing a pre-trained BERT model to assess the semantic closeness of the generated text to the ground truth.
- GPT-based Similarity: A novel approach where GPT itself was prompted to determine if two given paragraphs (one generated by an LLM, one from ground truth) described the same piece of code. This metric aimed to capture a higher-level understanding of semantic equivalence.
The overall experiment workflow was structured logically: initial code collection and cleaning (removing comments) formed the non-obfuscated data set, used to answer research question one. A subset of JavaScript code was then fed through various obfuscators and combined with IOCCC samples to create the obfuscated data set, addressing research question two. This systematic approach, from data collection to ground truth validation and multi-faceted evaluation, underscores the thoroughness of the study.
Demo / Proof of Concept
▶ Watch: Establishing ground truth via GPT-4 and human validation (6:30)
While the talk did not feature a traditional live demonstration of a new tool or exploit, the core of the research itself served as a comprehensive proof of concept for evaluating LLM capabilities in code analysis. The "demonstration" involved subjecting the chosen LLMs to a series of challenging tasks and systematically measuring their performance.
Specifically, the most compelling "demonstration" of LLM capabilities and limitations came from the de-obfuscation task. Here, the researchers presented GPT-3.5 and GPT-4 with highly complex and non-standard obfuscated C code from the International Obfuscated C Code Contest (IOCCC). The models were then prompted to generate a de-obfuscated, understandable version of the provided code. This task directly tested the LLMs' ability to reverse obfuscation and reconstruct original functionality.
The results of this de-obfuscation challenge highlighted the models' current shortcomings: GPT-4 failed to produce any output for approximately a quarter of the IOCCC samples, and the majority of its generated code was uncompilable. GPT-3.5, while slightly better in terms of compilability, still produced largely non-functional code. This "demo" effectively showcased that despite their advanced natural language understanding, current LLMs struggle significantly with the intricate logic and flexible transformations inherent in deliberately obfuscated code, especially when the obfuscation methods are non-standard and adversarial. The paper's focus was on this systematic, quantitative evaluation rather than presenting a novel de-obfuscation tool built on LLMs.
Defensive Implications
▶ Watch: GPT-4's high analysis accuracy and human-like insights (8:20)
The findings of this research carry significant implications for cybersecurity defenders and software development practices. Understanding the capabilities and, more importantly, the limitations of LLMs in code analysis is crucial for leveraging them effectively and mitigating potential risks.
Firstly, the high accuracy of GPT-4 (around 97.5%) in analyzing non-obfuscated code suggests that powerful LLMs can serve as valuable assistive tools for developers and security analysts. For well-written code or code using conventional structures, LLMs could potentially expedite initial understanding, assist in code reviews, or even help identify straightforward functional issues. This could be particularly useful in large codebases where manual analysis is time-consuming.
However, defenders must exercise caution. The observed instances of hallucination in GPT-4, where it made incorrect associations (e.g., importing a non-existent library), underscore that LLM outputs should never be trusted blindly. Any analysis generated by an LLM, especially in security-critical contexts like vulnerability assessment or malware analysis, requires thorough human validation. Relying solely on LLM output could lead to missed threats or misinterpretations.
The most critical defensive implication arises from the LLMs' performance against obfuscated code. While GPT models showed reasonable proficiency against classic obfuscation techniques like dead code injection and control flow flattening, they completely failed when confronted with advanced methods such as Wcor (WebAssembly-based obfuscation) and the highly flexible IOCCC samples. This stark contrast highlights that LLMs are currently insufficient for analyzing sophisticated malware or code designed to evade detection through modern obfuscation. Defenders cannot rely on current LLMs to automatically reverse engineer complex threats or understand the intent behind highly obscured malicious binaries. Specialized reverse engineering tools and human expertise remain indispensable for these tasks.
Furthermore, the poor performance of smaller LLMs (Llama, Code Llama, Starred Beta) indicates that deploying locally-hosted, less resource-intensive LLMs for reliable code analysis is not yet a viable option. Organizations looking to integrate LLM-based analysis will likely need to rely on more powerful, often cloud-based, commercial models, which introduces data privacy and cost considerations.
Finally, the struggle of LLMs with de-obfuscation tasks, particularly their inability to generate compilable and runnable code from IOCCC samples, suggests that current LLMs are not effective de-obfuscation engines. While GPT-4's output was more readable, its practical utility was limited by its lack of functionality. This reinforces that dedicated de-obfuscation tools and techniques, often requiring deep understanding of specific obfuscator logic, are still necessary.
In essence, LLMs are emerging as a promising aid for general code understanding, but they are far from a silver bullet for cybersecurity challenges involving adversarial or highly complex code. Defenders should view them as intelligent assistants whose outputs require critical review, rather than autonomous analysts capable of tackling the full spectrum of modern threats.
Key Takeaways
- GPT-4 and GPT-3.5 demonstrate high accuracy (around 97.5%) in analyzing non-obfuscated code, indicating their potential as powerful assistive tools for code comprehension in general software development.
- Smaller LLMs (Llama, Code Llama, Starred Beta) are currently unreliable for consistent code analysis, often failing to provide useful information or merely echoing the input code.
- GPT models can handle basic and classic obfuscation techniques (e.g., dead code injection, control flow flattening) reasonably well, but they fail dramatically when faced with advanced obfuscation methods like WebAssembly-based Wcor.
- Current LLMs perform very poorly at de-obfuscation tasks, especially for highly flexible and non-standard obfuscated code (e.g., IOCCC samples), often producing uncompilable or non-functional output.
- LLM analysis outputs can be prone to hallucinations and incorrect associations, emphasizing the critical need for human validation of any LLM-generated analysis in security-sensitive contexts.
- LLMs are promising as assistive tools but are not a replacement for specialized analysis tools or human experts when dealing with complex, adversarial, or heavily obfuscated code, such as in malware analysis or intellectual property protection.
About the Speaker(s)
The talk was presented by Chongzhou Fang, a fourth-year PhD student at UC Davis. He is part of a research team that includes Jialin Liu, Ruoyu Zhang, Han Wang, and Houman Homayoun. While the transcript primarily features Chongzhou Fang as the speaker, the collective work represents the contributions of this group of researchers, presumably also affiliated with UC Davis, given Fang's stated affiliation. Their research focuses on understanding the capabilities of large language models in complex computational tasks, particularly in the domain of software engineering and security.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk delivers a much-needed, data-driven reality check on LLM capabilities for code analysis, especially against obfuscation. It rigorously demonstrates that while advanced models handle clean code well, they fail spectacularly against real-world adversarial techniques, debunking significant industry hype. This is a foundational benchmark that will define conversations around LLM utility in security.
Heather Calloway (CISO) — STRONG ACCEPT
This research provides a critical and timely assessment of LLMs for code analysis, clearly delineating their strengths for non-obfuscated code and severe limitations against advanced obfuscation. It offers essential guidance for security leaders on where LLMs can genuinely assist and, crucially, where human expertise and specialized tools remain indispensable, directly informing risk ownership and strategic investment.