LLMIF: Augmented Large Language Model for Fuzzing IoT Devices

Jincheng Wang, Le Yu, Xiapu Luo

IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 5

Overview

This talk, presented by Jincheng Wang, Le Yu, and Xiapu Luo, introduces LLMIF (Augmented Large Language Model for Fuzzing IoT Devices), a novel approach that leverages the power of large language models (LLMs) to automate and enhance the process of fuzzing Internet of Things (IoT) devices. The core problem LLMIF addresses is the pervasive security vulnerabilities within the communication protocol stacks of IoT devices, which are often difficult to discover using traditional fuzzing methods due to their black-box nature, complex message formats, and implicit dependencies.

Watch on YouTube

Visual summary for LLMIF: Augmented Large Language Model for Fuzzing IoT Devices by Jincheng Wang, Le Yu, Xiapu Luo
Visual summary for LLMIF: Augmented Large Language Model for Fuzzing IoT Devices by Jincheng Wang, Le Yu, Xiapu Luo

Key moments

  1. 0:00 Introduction: IoT security and protocol vulnerabilities
  2. 1:30 Key challenges in fuzzing IoT protocols
  3. 3:30 Observation: Specification-guided fuzzing for IoT
  4. 5:40 How Large Language Models can automate specification analysis
  5. 6:30 Augmenting LLMs with specification knowledge using prompting
  6. 7:15 Augmented LLM capabilities: extracting formats, dependencies, response reasoning
  7. 8:50 The LLMIF fuzzing algorithm integrating augmented LLMs

LLMIF: Augmented Large Language Model for Fuzzing IoT Devices

Speakers: Jincheng Wang, Postdoc, Hong Kong Polytechnic University; Le Yu; Xiapu Luo

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=DZN4R45FmCI

Overview

This talk, presented by Jincheng Wang, Le Yu, and Xiapu Luo, introduces LLMIF (Augmented Large Language Model for Fuzzing IoT Devices), a novel approach that leverages the power of large language models (LLMs) to automate and enhance the process of fuzzing Internet of Things (IoT) devices. The core problem LLMIF addresses is the pervasive security vulnerabilities within the communication protocol stacks of IoT devices, which are often difficult to discover using traditional fuzzing methods due to their black-box nature, complex message formats, and implicit dependencies.

The presentation highlights how LLMs can effectively absorb and reason about intricate protocol specifications, transforming previously manual and error-prone tasks into automated, scalable processes. By extracting detailed message formats, identifying critical field constraints, and inferring complex message dependencies directly from documentation, LLMIF significantly improves the generation, mutation, and evaluation of fuzzing test cases. This innovative methodology promises to revolutionize how security researchers and manufacturers approach the security testing of IoT ecosystems, leading to the proactive identification and remediation of critical vulnerabilities.

The significance of LLMIF lies in its ability to overcome long-standing challenges in IoT security testing. Traditional fuzzing often struggles with the sheer diversity and complexity of IoT protocols, the extensive state space of devices, and the difficulty in obtaining actionable feedback from black-box targets. LLMIF's specification-guided approach, powered by LLMs, offers a robust solution, demonstrating superior code coverage and a remarkable success rate in uncovering previously unknown, real-world vulnerabilities in off-the-shelf Zigbee devices. This work not only advances the state of the art in fuzzing but also underscores the transformative potential of LLMs in highly technical security domains.

Background

▶ Watch: Introduction: IoT security and protocol vulnerabilities (0:00)

The proliferation of IoT devices across smart homes, smart cities, and industrial environments has introduced unprecedented convenience but also a significant attack surface. These devices are frequently criticized for their weak security guarantees, with the communication protocol stack standing out as a major vector for exploitation. Historical examples underscore this threat: the TCP stack designed by Siemens was found to harbor 13 vulnerabilities impacting a vast array of Siemens devices, and prior research revealed that flaws in the Zigbee protocol stack could enable attackers to compromise all Philips Hue bridges within a given area in a single operation. These incidents emphatically demonstrate the critical need for robust security verification of protocol stacks to identify and mitigate vulnerabilities before they are exploited in the wild.

Protocol fuzzing has emerged as a prominent technique for uncovering software bugs, particularly in network protocols like HTTP and RTSP, by feeding malformed or unexpected inputs to a target system. While successful in many domains, effectively fuzzing IoT protocols presents unique and formidable challenges. Firstly, IoT protocols typically specify a large number of diverse message types, each with complex and varied formats. For instance, Zigbee alone defines over 140 message types. Without an efficient mechanism to accurately recover these intricate message formats, generating valid fuzzing seeds and performing meaningful mutations becomes exceedingly difficult.

Secondly, IoT protocols often exhibit an extensive device state space, which can only be thoroughly explored through carefully crafted, well-formed sequences of messages. Unfortunately, the dependencies between these messages are frequently specified implicitly within the protocol documentation, making it challenging for automated tools to enrich message sequences and construct complex testing scenarios. For example, in Zigbee, an "identify" message might alter a device's identifying attributes, while an "add group" message might check these attributes before execution. Recognizing such logical dependencies requires sophisticated reasoning that traditional fuzzers lack.

Finally, evaluating the effectiveness of testing cases based on execution feedback, such as code coverage, is a standard practice in fuzzing to prioritize intriguing cases for further examination. However, the black-box nature of many IoT devices—such as smart locks and smart lights—makes it incredibly difficult to obtain detailed execution feedback. This absence of visibility hinders the ability to assess and prioritize high-quality testing cases that effectively trigger deeper device states or uncover complex logical flaws.

To address these challenges, the concept of specification-guided fuzzing gains traction. Protocol specifications are rich sources of information that can guide the fuzzing process. They explicitly detail message formats and field constraints (e.g., an "add group" message having two fields, with the first field's value restricted to a specific range). Specifications also describe message functionalities, which can help infer message dependencies based on how they interact with device attributes. Furthermore, they outline how devices should generate response codes, which can serve as an invaluable signal for evaluating test cases. A successful "add group" execution, for instance, might update a group table attribute and return a "success" code, indicating a desirable state transition. However, the critical limitation has been the lack of automation in analyzing these voluminous and complex specification documents. Existing efforts largely rely on manual human analysis, a process that is labor-intensive, error-prone, and impossible to scale for the hundreds of thousands of test cases generated during a typical fuzzing campaign. This is where large language models offer a transformative solution.

Key Findings

▶ Watch: Observation: Specification-guided fuzzing for IoT (3:30)

The research behind LLMIF demonstrates several pivotal findings that underscore the efficacy and innovation of using augmented large language models for IoT fuzzing. These discoveries fundamentally change how protocol specifications can be leveraged for security testing.

Firstly, a core finding is that large language models are exceptionally capable of absorbing, processing, and reasoning about complex protocol specification documents. Thanks to their powerful natural language processing (NLP) and reasoning capabilities, LLMs can effectively act as an automated expert, understanding domain-specific knowledge embedded in thousands of pages of documentation. This ability overcomes the traditional bottleneck of manual specification analysis, which is both labor-intensive and prone to human error.

Secondly, LLMIF proves that LLMs can accurately extract granular protocol information crucial for fuzzing. Specifically, the augmented LLM successfully constructs detailed message header formats, payload formats, and identifies hundreds of interesting field values directly from the Zigbee specification. This precise information is vital for generating well-formed initial seeds and performing targeted mutations. Beyond static format extraction, the LLM also demonstrated a remarkable ability to infer message dependencies. After analyzing 20,000 message pairs, the LLM successfully recovered nearly 1,000 implicit message dependencies, a task that demands logical reasoning about message functionalities and their impact on shared device attributes.

Thirdly, the study highlights the LLM's capability to reason about device response codes and evaluate test cases effectively. For each executed test case, the LLM can determine if the observed response code implies a normal device state transition (aligning with the specification) or an abnormal transition (violating the specification). This intelligent evaluation mechanism is critical for prioritizing high-quality test cases that either successfully trigger intended state changes or expose unexpected behaviors, even in black-box scenarios where traditional code coverage is unavailable.

Empirically, LLMIF demonstrated significant improvements in fuzzing coverage and bug identification. When compared against four state-of-the-art fuzzing tools on 11 off-the-shelf Zigbee devices, LLMIF improved message coverage by an impressive 82%. Furthermore, when tested on a Texas Instruments Zigbee evaluation board running the official Z-stack, LLMIF enhanced both edge and statement code coverage by at least 50% compared to baseline methods. These quantitative results unequivocally establish LLMIF's superior effectiveness in exploring the protocol and code space of IoT devices.

Perhaps the most impactful finding is LLMIF's success in discovering real-world, previously unknown vulnerabilities. The tool was used to fuzz the 11 real-world Zigbee devices, leading to the revelation of 11 unique vulnerabilities, 8 of which were zero-day vulnerabilities. These bugs ranged in severity, causing devices to crash, lose specific functionalities (such as attribute updates), or even become completely bricked, requiring factory resets or power cycles to recover. This practical demonstration of bug-finding capability underscores LLMIF's potential for enhancing the security posture of widely deployed IoT ecosystems.

Finally, an ablation study confirmed the individual utility of the information extracted by the LLM. The knowledge of interesting field values, message headers, and message dependencies each boosted code coverage by at least 8%, indicating that every component of the LLM's analysis contributes meaningfully to the overall performance of the fuzzing process. This validates the multi-faceted approach of leveraging LLMs for comprehensive specification analysis.

Technical Deep Dive

▶ Watch: How Large Language Models can automate specification analysis (5:40)

The technical foundation of LLMIF revolves around an innovative methodology that augments large language models with deep protocol specification knowledge to drive an intelligent fuzzing process. This involves automating the traditionally human-intensive tasks of specification analysis, message generation, mutation, and test case evaluation.

The core idea is to treat the protocol specification document as a vast knowledge base for the LLM. To achieve this, the researchers developed an extraction-integration method to imbue the LLM with this domain-specific knowledge. This process begins by slicing the specification documents to filter out irrelevant sections and retain only the critical pieces containing protocol message descriptions. This pre-processing step ensures that the LLM focuses on pertinent information, enhancing efficiency and accuracy.

Once the relevant message descriptions are extracted, a background-augmented prompting technique is employed. This technique combines the extracted message descriptions with specific task descriptions to generate precise prompts for the LLM. This allows the LLM to perform various downstream tasks essential for fuzzing. The target protocol chosen for this research was Zigbee, primarily due to its widespread popularity in IoT ecosystems and the complexity of its specification.

The augmented LLM performs several critical tasks:

  1. Extracting Message Information: The LLM is prompted to analyze the sliced specification to identify and formalize message structures. This includes recovering detailed message header formats and payload formats. Beyond structural definitions, the LLM also identifies hundreds of interesting field values and their associated constraints (e.g., valid ranges, specific enumerated values). This granular information is crucial for generating initial fuzzing seeds that are both syntactically valid and semantically meaningful, avoiding the generation of entirely malformed inputs that might be trivially rejected by the device.
  1. Reasoning about Message Dependencies: A significant challenge in fuzzing stateful protocols is understanding how messages interact and depend on each other to transition device states. The LLM is tasked with analyzing message descriptions to infer these dependencies based on their interaction with common device attributes. For instance, if one message modifies a device attribute (e.g., group table) and another message checks that attribute, the LLM identifies a dependency. Through this process, the LLM successfully recovered nearly 1,000 message dependencies by checking 20,000 message pairs, demonstrating its advanced logical reasoning capabilities beyond simple pattern matching. This capability allows LLMIF to construct complex, multi-message sequences that are more likely to explore deeper, less-frequently accessed device states.
  1. Reasoning about Response Codes: In black-box IoT environments, direct code coverage feedback is often unavailable. LLMIF addresses this by leveraging the LLM to interpret device response codes. The specification typically details what response code a device should generate under various conditions (e.g., successful operation, invalid field, out of resources). For each executed test case, LLMIF collects the device's response code and feeds it to the augmented LLM. The LLM then determines if the code implies a normal device state transition (consistent with the specification) or an abnormal transition (violating the specification). Both types of transitions are considered valuable: normal transitions confirm the test case's validity in triggering expected behavior, while abnormal transitions hint at potential vulnerabilities or unexpected state changes. This intelligent evaluation phase ensures that "interesting" test cases, which either successfully advance the device's state or expose anomalous behavior, are kept for further investigation.

The overall fuzzing algorithm, LLMIF, integrates these LLM-boosted capabilities across all phases of the fuzzing lifecycle:

  • Seed Generation: This initial phase utilizes the LLM-extracted payload formats and interesting field values to generate a diverse set of initial fuzzing seeds. These seeds are designed to be well-formed enough to be processed by the target device, increasing the likelihood of reaching deeper code paths.
  • Mutation Phase: Once initial seeds are generated, the mutation phase employs various strategies guided by the LLM's understanding of header and payload formats. This allows for targeted alterations of fields, injecting valid but unexpected values, or malforming specific parts of the message in ways that the protocol specification might allow but the implementation might not handle correctly.
  • Response Reasoning (Evaluation) Phase: After the target IoT device executes a test case and returns a response code, the LLM's response reasoning capability comes into play. It examines the response code against the specification to determine if the test case promoted a device state transition (normal or abnormal). Only those test cases deemed "interesting" by the LLM are retained for further processing, effectively filtering out redundant or unimpactful inputs.
  • Seed Enrichment Phase: The retained "interesting" test cases are then enriched into longer, more complex message sequences. This is achieved by leveraging the message dependency relationships identified by the LLM. By chaining dependent messages, LLMIF can construct sophisticated scenarios that explore multi-step interactions and stateful behaviors, significantly increasing the probability of uncovering vulnerabilities that manifest only after a specific sequence of operations. These enriched sequences are then saved as new seeds for subsequent fuzzing rounds, creating a self-improving feedback loop.

This comprehensive integration of LLM capabilities into every stage of fuzzing—from understanding the specification to generating, mutating, and evaluating test cases—forms the backbone of LLMIF's effectiveness in securing complex IoT devices.

Demo / Proof of Concept

▶ Watch: Augmented LLM capabilities: extracting formats, dependencies, response reasoning (7:15)

While the talk did not feature a live, interactive demonstration of LLMIF in action, the researchers presented compelling evidence of its capabilities through a thorough evaluation and proof-of-concept on real-world devices. The effectiveness of LLMIF was rigorously assessed by focusing on key performance indicators: increased code coverage, the impact of various types of protocol information, and the discovery of previously unknown vulnerabilities on commercial IoT devices.

For the first evaluation question concerning coverage, LLMIF was pitted against four state-of-the-art fuzzing tools. The target devices consisted of 11 off-the-shelf Zigbee devices, representing a diverse range of common IoT products. The primary metric for comparison was message coverage, defined as the number of supported message types that each fuzzer could successfully interact with. The results demonstrated a significant advantage for LLMIF, which improved message coverage by an impressive 82% compared to the baseline methods. This indicates LLMIF's superior ability to generate valid and diverse messages that are accepted and processed by a wide array of Zigbee devices.

To further quantify the depth of exploration, the researchers utilized a Zigbee evaluation board from Texas Instruments, specifically one loaded with their official Zigbee stack, known as Z-stack. This setup allowed for precise measurement of code coverage (specifically edge and statement coverage), which is often difficult to obtain from black-box commercial devices. The findings were equally compelling: LLMIF improved both edge and statement coverage by at least 50% when compared to the baseline fuzzing methods. This metric highlights LLMIF's capacity to navigate deeper into the device's internal logic and trigger more execution paths, increasing the likelihood of uncovering subtle bugs.

To understand the individual contributions of the LLM-extracted information, an ablation study was conducted. This study isolated the impact of different knowledge sources on LLMIF's performance. The results clearly showed that the knowledge of interesting field values, message headers, and message dependencies each contributed positively, boosting code coverage by at least 8% individually. This confirms that the multi-faceted approach of leveraging various aspects of specification knowledge is not only beneficial but that each component provides a measurable improvement to the overall fuzzing efficacy.

The most critical aspect of the proof-of-concept was LLMIF's ability to discover actual security flaws. The fuzzer was deployed against the 11 real-world Zigbee devices, and its success was remarkable. LLMIF successfully revealed 11 distinct vulnerabilities, with 8 of these being zero-day vulnerabilities—meaning they were previously unknown to vendors and the public. These vulnerabilities manifested in various forms:

  • Device crashes: Leading to denial-of-service or temporary unavailability.
  • Loss of specific functionalities: Such as the inability to update attributes, compromising device integrity or control.
  • Device bricking: In some severe cases, the bugs caused devices to enter an unrecoverable state, requiring factory resets or power cycling, indicating a complete compromise of stability.

These findings unequivocally demonstrate LLMIF's practical effectiveness in identifying critical security weaknesses in widely used IoT devices, validating its design and the power of an LLM-augmented, specification-guided fuzzing approach.

Defensive Implications

▶ Watch: The LLMIF fuzzing algorithm integrating augmented LLMs (8:50)

The findings from LLMIF have profound implications for defenders across the IoT ecosystem, from device manufacturers and security researchers to end-users. The success of LLMIF in uncovering zero-day vulnerabilities underscores the urgent need for a more proactive and sophisticated approach to IoT security.

For IoT device manufacturers, LLMIF serves as a critical wake-up call and a blueprint for enhanced security testing. Manufacturers must recognize that existing fuzzing methodologies may not be sufficient to thoroughly vet the complex protocol stacks of their devices. They should consider integrating advanced, specification-guided fuzzing techniques like LLMIF into their development and quality assurance pipelines. This means:

  • Rethinking specification adherence: Ensuring that device implementations strictly adhere to protocol specifications, particularly concerning message parsing, field validation, and state transitions. Discrepancies between specification and implementation are prime targets for LLMIF.
  • Investing in comprehensive testing: Moving beyond basic functional tests to include deep protocol fuzzing that explores edge cases, unexpected sequences, and malformed inputs as guided by an automated understanding of the specification.
  • Leveraging AI/ML for security: Exploring how LLMs and other AI techniques can automate the analysis of documentation, code, and network traffic to proactively identify potential vulnerabilities.
  • Improving feedback mechanisms: Where possible, manufacturers should aim to provide better debugging and logging capabilities in development versions of their devices to assist fuzzers in evaluating test cases and understanding internal states, even if these are removed in production.

Security researchers and penetration testers can adopt the principles behind LLMIF to enhance their vulnerability discovery efforts. The methodology demonstrates that a deep, automated understanding of protocol specifications, combined with intelligent test case generation and evaluation, is a powerful paradigm for black-box testing. Researchers should:

  • Explore LLM-powered tools: Investigate and develop new tools that integrate LLMs for automated protocol analysis, state inference, and intelligent fuzzing.
  • Focus on specification nuances: Pay close attention to implicit dependencies, field constraints, and response code behaviors documented in specifications, as these are rich areas for vulnerability discovery.
  • Prioritize stateful fuzzing: Recognize that many critical vulnerabilities in IoT devices arise from complex state transitions and message sequences, which LLMIF excels at exploring.

For end-users and system administrators deploying IoT devices, the implications are more about awareness and best practices:

  • Stay informed and update devices: The discovery of zero-day vulnerabilities highlights that even widely deployed and trusted devices can harbor serious flaws. Users must prioritize installing firmware updates promptly to patch known vulnerabilities.
  • Network segmentation: Isolate IoT devices on separate network segments where possible to limit their ability to interact with critical infrastructure or other sensitive devices, even if compromised.
  • Vendor reputation: Consider the security track record and commitment to security updates when choosing IoT devices.

Ultimately, LLMIF provides a compelling argument for a paradigm shift in IoT security. It demonstrates that by bridging the gap between human-readable specifications and automated testing, we can significantly improve our ability to secure the ever-expanding landscape of connected devices. The era of manual, ad-hoc security testing for complex IoT protocols is rapidly giving way to intelligent, AI-augmented approaches that promise greater depth, efficiency, and efficacy in vulnerability discovery.

Key Takeaways

  • LLMs automate complex protocol specification analysis: LLMIF demonstrates that large language models can effectively absorb, process, and reason about intricate IoT protocol specifications, automating tasks previously requiring intensive human effort.
  • Significant improvements in fuzzing coverage: LLMIF achieved an 82% increase in message coverage on off-the-shelf Zigbee devices and at least a 50% improvement in edge and statement code coverage on a Z-stack evaluation board compared to baseline methods.
  • Specification-guided fuzzing is critical for black-box IoT: By leveraging LLM-extracted message formats, field constraints, and inferred message dependencies, LLMIF effectively guides fuzzing for black-box IoT devices where traditional code coverage feedback is often unavailable.
  • LLMIF successfully uncovers real-world zero-day vulnerabilities: The tool identified 11 vulnerabilities, including 8 zero-days, in commercial Zigbee devices, leading to crashes, loss of functionality, and even device bricking, proving its practical bug-finding capability.
  • Multi-faceted LLM augmentation enhances all fuzzing phases: LLMIF integrates LLM capabilities for seed generation (using extracted formats), mutation (using formats), response reasoning (evaluating state transitions via response codes), and seed enrichment (constructing complex sequences using dependencies).
  • Intelligent evaluation via response codes is a game-changer: The LLM's ability to determine if a device's response code implies normal or abnormal state transitions provides crucial feedback for prioritizing interesting test cases, especially in black-box environments.

About the Speaker(s)

Jincheng Wang is a Postdoc from Hong Kong Polytechnic University. He presented the work on LLMIF at the IEEE S&P conference. The research was conducted in collaboration with Professor Le Yu and Professor Xiapu Luo. Further biographical details for the speakers were not provided in the transcript or metadata.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This work presents a compelling, technically sound approach to augmenting IoT protocol fuzzing using LLMs for specification analysis. The ability to automatically infer message formats, dependencies, and reason about response codes in black-box scenarios is a significant advancement, leading to the discovery of 8 zero-day vulnerabilities in commercial devices. This isn't just slapping "AI" on it; it's a clever application with real impact.

Heather Calloway (CISO) — STRONG ACCEPT

LLMIF presents a compelling, evidence-backed case for leveraging large language models to automate and enhance IoT protocol fuzzing. Its success in uncovering zero-day vulnerabilities in commercial devices provides critical, actionable insights for manufacturers and security leaders responsible for securing connected ecosystems.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024