Next-Gen Detection: Harnessing LLMs for Sigma Rule Automation
Dave Johnson (Threat Intelligence Adviser · Feedle)
BSidesSF 2024 · Day 1
Overview
This talk, presented by Dave Johnson, a Threat Intelligence Advisor at Feedly, delves into the innovative application of Large Language Models (LLMs) to automate the creation of Sigma rules for cybersecurity detection. Johnson's research, born from an initial experiment with ChatGPT, addresses the critical need for more efficient and higher-quality detection rule generation in an evolving threat landscape. The core objective is to empower defenders to stay ahead of adversaries who are increasingly leveraging AI for malicious purposes, including malware generation and attack orchestration.

Key moments
- 0:40 Initial experiments with ChatGPT for Sigma rule automation.
- 2:10 Overview of LLM strategies: RAG, fine-tuning, prompt chaining, and data quality.
- 4:00 Introduction to Sigma rules: vendor-neutral format for attack behavior.
- 6:00 Importance of input/output validation for LLM-generated rules.
- 7:00 LLM-assisted data curation: extracting and scoring attack procedures from threat intel.
- 9:00 Detailed explanation of RAG and few-shot prompting for rule generation.
- 26:00 Comparative results: RAG and prompt chaining outperform fine-tuning.
Next-Gen Detection: Harnessing LLMs for Sigma Rule Automation
Speakers: Dave Johnson
Conference: BSidesSF 2024
YouTube: https://www.youtube.com/watch?v=oPsSUxt7ufs
Overview
This talk, presented by Dave Johnson, a Threat Intelligence Advisor at Feedly, delves into the innovative application of Large Language Models (LLMs) to automate the creation of Sigma rules for cybersecurity detection. Johnson's research, born from an initial experiment with ChatGPT, addresses the critical need for more efficient and higher-quality detection rule generation in an evolving threat landscape. The core objective is to empower defenders to stay ahead of adversaries who are increasingly leveraging AI for malicious purposes, including malware generation and attack orchestration.
The presentation outlines various strategies for integrating LLMs into the detection pipeline, emphasizing the paramount importance of data quality—both for the input threat intelligence and the output Sigma rules. Johnson explores the pros and cons of Retrieval Augmented Generation (RAG) with Few-Shot Prompting, Prompt Chaining, and Fine-tuning, providing a practical guide for security professionals looking to leverage these advanced techniques. The work culminates in a publicly available GitHub project, SigGen, which demonstrates the practical implementation of these concepts, allowing users to automatically generate Sigma rules from threat intelligence reports.
The significance of this research lies in its potential to democratize advanced detection capabilities and streamline the often-laborious process of creating vendor-agnostic detection rules. By automating the extraction of actionable attack procedures and the subsequent generation of high-quality Sigma rules, organizations can enhance their proactive defense posture, reduce response times during incidents, and improve the overall effectiveness of their security operations. Johnson's exploratory work not only offers concrete methodologies but also highlights areas for future research and community collaboration in this rapidly advancing field.
Background
▶ Watch: Initial experiments with ChatGPT for Sigma rule automation. (0:40)
The foundation of this research lies in Sigma rules, a concept introduced by Florian Roth and Thomas Patzke around 2016-2017. Sigma rules were developed out of a shared frustration with the disparate and vendor-specific methods required to search for malicious activity across various log management systems. They provide a generic, human-readable format to describe attack behavior in logs, making them vendor-neutral and highly adaptable. The Sigma HQ repository, maintained by Florian Roth, serves as a community-driven platform for sharing and collaborating on these detection rules, fostering collective defense knowledge.
Despite the utility of Sigma rules, their manual creation can be time-consuming and prone to inconsistencies, especially when dealing with a constant influx of new threat intelligence. Johnson's initial attempts to use general-purpose LLMs like ChatGPT for Sigma rule generation yielded "terrible" and "trash" results. This highlighted a significant problem: while LLMs possess vast knowledge, they lack the specialized context and structured guidance required to produce actionable and high-quality security detection rules without proper input and validation.
The central challenge identified was the "quality of the input of the data." Just as "trash in, trash out" applies to traditional data processing, it holds even more true for LLMs. Without robust mechanisms to ensure the quality of the attack procedures fed into the LLM and to validate the generated Sigma rules, any automation effort would be futile. This necessitated the development of a structured approach that incorporates automated quality checks and leverages specific LLM strategies to transform raw threat intelligence into effective, deployable detection logic. The goal was to move beyond simple prompting to a more sophisticated pipeline that could consistently produce valuable detection artifacts.
Key Findings
▶ Watch: Introduction to Sigma rules: vendor-neutral format for attack behavior. (4:00)
Dave Johnson's research yielded several critical findings regarding the effective use of LLMs for Sigma rule automation:
- Data Quality is Paramount: The most significant finding was that the quality of the input data directly dictates the quality of the generated Sigma rules. Initial attempts with raw LLM prompting produced "trash" results, underscoring the necessity of rigorous validation for both the attack procedures fed into the LLM and the resulting detection rules.
- Effective Attack Procedure Extraction: LLMs can be successfully employed to extract actionable attack procedures from unstructured threat intelligence reports. From over 200 reports, 624 procedures were extracted, and a subsequent LLM-based scoring mechanism filtered these down to approximately 183 high-quality, actionable procedures.
- RAG with Few-Shot Prompting Achieves Highest Quality: The Retrieval Augmented Generation (RAG) approach combined with Few-Shot Prompting demonstrated the highest average quality score for generated Sigma rules, achieving an impressive 8.09 out of 10. This method's success is attributed to its ability to inject highly relevant and high-quality examples from the Sigma HQ repository directly into the LLM's context.
- Prompt Chaining Ensures Consistency: While slightly lower in average score (7.84), Prompt Chaining proved to be the most consistent method. By breaking down the complex task of rule creation into logical, sequential steps and forcing the LLM to "think through" each stage, it produced more reliable and predictable outputs, mimicking an analyst's thought process.
- Fine-tuning Challenges and Limitations: Fine-tuning, despite its perceived potential, was the "biggest failure" in this experiment, yielding the lowest average score (6.99) and being the most expensive. This was primarily attributed to using a less capable model (GPT-3.5 Turbo) for cost-effectiveness and potential issues like overfitting and the quality/diversity of the training dataset derived solely from the Sigma HQ repository.
- Automated Quality Assurance is Essential: Implementing automated quality assurance checks, both for input attack procedures and output Sigma rules (using a "meta prompt" with criteria like Precision, Real-world applicability, and Alignment with the attack procedure), significantly improved the consistency and usability of the generated rules across all methods.
- Scalability and Cost Considerations: Each method presents different trade-offs. RAG offers good quality and adaptability to new threats, while Prompt Chaining, though consistent, suffers from high inference times (~90 seconds per rule) and increased token costs. Fine-tuning, despite initial investment and current challenges, holds long-term promise for efficiency and speed if implemented correctly with better models and data.
Technical Deep Dive
▶ Watch: Importance of input/output validation for LLM-generated rules. (6:00)
The technical core of Johnson's research revolves around a multi-stage pipeline for automated Sigma rule generation, beginning with data acquisition and culminating in the evaluation of rules produced by different LLM strategies.
Data Gathering and Pre-processing
The initial step involved compiling a robust dataset of attack procedures. Johnson scraped content from over 200 threat intelligence reports, leveraging an LLM to extract potential attack procedures. This process yielded 624 raw attack procedures. Recognizing that not all extracted text would be actionable for detection rule creation, a crucial quality assurance step was implemented. An LLM was used to score each procedure on a scale of 1 to 10 against specific criteria:
- Sufficiency of context for a detection engineer.
- Description of how the threat actor performed the action.
- Mapping to MITRE ATT&CK (though not explicitly stated as a scoring criterion, it was mentioned as a characteristic of the procedures).
This rigorous filtering process resulted in approximately 183 usable, high-quality attack procedures, forming the foundation for subsequent Sigma rule generation experiments.
LLM Strategies for Sigma Rule Generation
Johnson explored three primary LLM strategies, each with distinct methodologies and trade-offs:
1. Few-Shot Prompting with Retrieval Augmented Generation (RAG)
This method combines two powerful techniques:
- Few-Shot Prompting: The LLM is provided with a few high-quality examples (or "shots") of desired output before being asked to generate a new one. This guides the LLM, teaching it the specific task at hand without requiring full fine-tuning. For Sigma rules, this means providing examples of well-formed rules.
- Retrieval Augmented Generation (RAG): To ensure the LLM draws from the most relevant and up-to-date knowledge, Johnson downloaded the entire Sigma HQ repository and used it as a database. When a new attack procedure is presented, an embedding-based similarity search is performed against this database. This process converts the input query and the Sigma rules into numerical vector representations (embeddings) and finds the rules whose embeddings are most similar to the query. These "best" matching Sigma rules are then injected into the LLM's prompt as examples for the few-shot prompting.
The prompt structure for this method was meticulously designed:
- Role Assignment: The LLM is instructed to act as a detection engineer.
- Instruction Set: Clear, detailed instructions on how to construct a Sigma rule, drawing inspiration from Florian Roth's official rule creation guide.
- Quality Assurance (QA) Test: An embedded evaluation criteria, using HTML-like tags (e.g.,
[evaluation criteria]...[/evaluation criteria]), forces the LLM to consider factors like specificity, real-world applicability, minimization of false positives/negatives, and compatibility with log sources. This "forces the LLM to think logically." - Placeholder for RAG Examples: A designated section within the prompt where the retrieved Sigma rule examples are dynamically inserted.
A key requirement for this method is an LLM with a large context window, allowing it to retain and process the extensive instructions, evaluation criteria, and multiple Sigma rule examples within a single prompt. Johnson likened this process to Neo learning Kung Fu in The Matrix, where knowledge is directly "injected" into the system.
2. Prompt Chaining
Prompt Chaining breaks down the complex task of generating a Sigma rule into a series of smaller, more manageable, and logical steps. Each step is handled by a separate prompt, and the output of one prompt serves as the input for the next. This sequential processing forces the LLM to reason through the problem incrementally, much like a human analyst.
An example chain might look like this:
- Prompt 1 (Procedure Analysis): Given an attack procedure, identify key indicators of compromise (IOCs) or behaviors.
- Prompt 2 (Log Source Identification): Based on the procedure and identified indicators, determine the most relevant log sources and categories (e.g., endpoint logs, network logs, security logs).
- Prompt 3 (Rule Construction): Combine the outputs from previous steps to construct the actual Sigma rule.
Pros: This method yields highly consistent results because the LLM is guided through a structured thought process. It requires less upfront data than fine-tuning and is potentially more flexible than RAG in certain scenarios.
Cons: The primary drawbacks are inference time and cost. Each step in the chain requires a separate LLM invocation, leading to an average generation time of 90 seconds per Sigma rule. This also translates to higher token usage, making it potentially more expensive for large-scale automation. Johnson noted that his implementation used "complex conditional logic" to optimize efficiency by selectively using outputs from previous prompts.
3. Fine-tuning
Fine-tuning involves taking a pre-trained LLM and further training it on a specific, specialized dataset to adapt its behavior and output style to a particular task. The goal is to achieve highly consistent, specialized outputs with fewer tokens and less explicit prompting.
Johnson's fine-tuning experiment involved:
- Dataset Creation: The Sigma HQ repository was again used, but this time to create input-output pairs. The title of each Sigma rule was rephrased as a question (e.g., "How to detect X?"), which served as the input, and the corresponding Sigma rule served as the desired output.
- Process: The dataset was split into an 80% training set and a 20% test (validation) set. Standard deep learning hyperparameters such as learning rate and batch size were configured. The fine-tuning job was then executed.
Challenges and Results: This method proved to be the "most expensive failure." Johnson used GPT-3.5 Turbo for cost-effectiveness, which was less capable than the models used for RAG and Prompt Chaining. The fine-tuning process exhibited a low training loss but a higher validation loss, indicating potential overfitting—where the model memorized the training data rather than learning generalizable patterns. The quality of the Sigma HQ dataset itself, which includes experimental or less-than-perfect rules, also contributed to the poor results. Johnson suggested that future attempts should explore different hyperparameters, a more diverse security data set (including threat intel reports), and more capable LLMs.
Evaluation of Generated Sigma Rules
To objectively compare the performance of the three methods, Johnson developed a "meta prompt" to have an LLM evaluate the quality of the generated Sigma rules. Each rule was scored on a scale of 1 to 10 against three key criteria:
- Precision: How specific and accurate is the rule in targeting malicious behavior, and how well does it minimize false positives?
- Real-world applicability: Can the rule be effectively applied in typical security operations environments, considering the availability of necessary log sources?
- Alignment with the attack procedure: Does the generated rule accurately reflect the malicious behavior described in the original input attack procedure?
The scores for each criterion were averaged to produce a single numeric score for each rule. The results showed RAG with Few-Shot Prompting leading with an average score of 8.09, followed by Prompt Chaining at 7.84, and Fine-tuning significantly lower at 6.99. While RAG achieved the highest average, Prompt Chaining was noted for its superior consistency due to its methodical approach.
Demo / Proof of Concept
▶ Watch: Detailed explanation of RAG and few-shot prompting for rule generation. (9:00)
Dave Johnson demonstrated a practical implementation of his research through a tool named SigGen, which he made publicly available on GitHub on the day of the conference. The project is designed to be user-friendly and accessible, even being dockerized for ease of installation.
SigGen embodies the principles discussed in the talk, specifically focusing on automating the extraction of attack procedures and the subsequent generation of Sigma rules. The core functionality allows users to:
- Input Threat Intelligence: Provide website URLs to threat intelligence reports.
- Scrape Content: The tool automatically scrapes the content from these URLs.
- Extract Attack Procedures: An LLM is then used to identify and extract relevant attack procedure data from the scraped text. This leverages the initial data gathering and filtering techniques discussed in the technical deep dive.
- Generate Sigma Rules: Finally, using one of the explored LLM strategies (presumably RAG or Prompt Chaining, given their higher performance), SigGen generates corresponding Sigma rules based on the extracted procedures.
A screenshot of the tool was presented during the talk, illustrating its interface and capabilities. The availability of SigGen as an open-source project encourages further experimentation, collaboration, and contribution from the security community, aligning with Johnson's goal of fostering research in this area.
Defensive Implications
▶ Watch: Comparative results: RAG and prompt chaining outperform fine-tuning. (26:00)
The research presented by Dave Johnson offers significant defensive implications for cybersecurity teams striving to enhance their detection capabilities and maintain a proactive stance against evolving threats.
- Accelerated Threat Response: The primary advantage is the ability to generate Sigma rules much faster than manual processes. In incident response scenarios, where immediate detection is crucial ("like yesterday"), automated rule generation can drastically reduce the time to deploy new detections, allowing organizations to respond to emerging threats with unprecedented speed.
- Enhanced Proactive Defense: By leveraging LLMs to process vast amounts of threat intelligence, security teams can more effectively identify and operationalize new attack techniques. This shifts the focus from reactive defense to a more proactive posture, anticipating and detecting threats before they cause significant damage.
- Improved Rule Quality and Consistency: The emphasis on automated quality checks for both input attack procedures and generated Sigma rules ensures that the detection logic is precise, applicable to real-world environments, and aligned with the intended malicious behavior. This reduces the likelihood of false positives and negatives, improving the overall signal-to-noise ratio in security alerts.
- Democratization of Detection Engineering: The SigGen tool and the methodologies presented lower the barrier to entry for creating sophisticated detection rules. This empowers a broader range of security professionals, even those without deep expertise in specific log formats or SIEM query languages, to contribute to the organization's defensive capabilities.
- Vendor-Neutral Detection: By focusing on Sigma rules, the generated detections remain vendor-agnostic. This allows organizations to deploy rules across diverse security tools and platforms without extensive re-engineering, maximizing their utility and reducing vendor lock-in.
- Leveraging AI Against AI: As adversaries increasingly utilize LLMs for malicious activities like malware generation and attack orchestration, defenders must also harness these technologies. Johnson's work demonstrates a tangible way for security teams to turn the tables, using AI to build more robust defenses against AI-powered threats.
- Adaptability to Novel Attacks: The RAG approach, in particular, allows for easier incorporation of new contributions to the Sigma HQ repository as examples. This means the system can adapt to and generate detections for novel attack techniques more readily, provided the underlying knowledge base is kept updated.
- Resource Optimization: Automating rule creation frees up valuable time for detection engineers, allowing them to focus on more complex tasks such as threat hunting, advanced analytics, and strategic security initiatives, rather than repetitive rule writing.
Key Takeaways
- Data Quality is Paramount: The effectiveness of LLMs in generating Sigma rules is directly proportional to the quality of the input attack procedures and the examples provided. Rigorous validation and filtering of data are non-negotiable.
- RAG with Few-Shot Prompting Excels in Quality: The combination of Retrieval Augmented Generation (RAG) and Few-Shot Prompting demonstrated the highest average quality scores for generated Sigma rules, making it a strong candidate for high-fidelity detection engineering.
- Prompt Chaining Offers Consistency: For scenarios prioritizing consistent output and a methodical approach, Prompt Chaining is highly effective, albeit with trade-offs in inference time and token cost due to its sequential nature.
- Fine-tuning Requires Strategic Investment: While challenging and expensive in initial experiments, Fine-tuning holds long-term potential for specialized, efficient, and faster rule generation, provided it's done with high-quality, diverse data and more capable LLMs.
- Automated QA is Crucial for Trust: Implementing LLM-based "meta prompts" for evaluating rule Precision, Real-world applicability, and Alignment with the attack procedure is essential for building trust in automated outputs and ensuring their operational utility.
- Practical Tools are Emerging: Projects like SigGen demonstrate the immediate applicability of these research findings, enabling security teams to automate the generation of Sigma rules directly from threat intelligence reports, thereby enhancing proactive defense.
About the Speaker(s)
Dave Johnson is a Threat Intelligence Advisor at Feedly, a company specializing in ENT (Enterprise News & Threat) research tools for threat intelligence and market intelligence. With approximately 15 years of CTI (Cyber Threat Intelligence) experience, Johnson brings a wealth of practical knowledge to the field. His career includes a significant tenure as a former FBI analyst, where he focused on cybercrime and advanced persistent threat (APT) groups, notably working from a unique base in Wisconsin. He maintains a personal website, DaveintheMiddle.com, and can be reached via email at Dave@feedly.com.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk provides a practical, hands-on exploration of leveraging Large Language Models for automating Sigma rule generation. The speaker details three distinct strategies—Retrieval Augmented Generation (RAG), prompt chaining, and fine-tuning—highlighting the critical importance of input data quality and robust validation. The comparative analysis of these methods, including the speaker's candid admission of fine-tuning's 'expensive failure,' offers valuable insights for detection engineers looking to enhance their proactive threat hunting capabilities.
Heather Calloway (CISO) — STRONG ACCEPT
This presentation offers a highly practical and actionable approach to leveraging Large Language Models for automating the creation of Sigma detection rules. While not directly addressing board-level governance, the methodology presented significantly enhances an organization's ability to rapidly develop and deploy threat detection capabilities, thereby reducing operational risk and improving incident response times. The speaker's detailed comparison of RAG, prompt chaining, and fine-tuning provides clear guidance for security leaders and practitioners on implementing these technologies effectively.