Finetuning Large Language Models (LLMs) for Security Log Detections

Wilson Tang (Machine Learning Engineer · Adobe)

BSidesSF 2024 · Day 1

Overview

This talk, presented by Wilson Tang, a Machine Learning Engineer on the threat hunting team at Adobe, delves into the innovative application of Large Language Models (LLMs) for security log detections. Specifically, Tang explores the process of fine-tuning LLMs to identify command obfuscation, a prevalent technique used by adversaries to evade traditional security controls. The presentation highlights the limitations of existing rule-based and conventional machine learning approaches in detecting sophisticated and dynamic threats, positioning LLMs as a powerful new tool in the defender's arsenal.

Watch on YouTube

Visual summary for Finetuning Large Language Models (LLMs) for Security Log Detections by Wilson Tang
Visual summary for Finetuning Large Language Models (LLMs) for Security Log Detections by Wilson Tang

Key moments

  1. 02:00 Traditional Log Detections & Limitations
  2. 03:15 Command Obfuscation Detection Case Study
  3. 06:00 ML Techniques: Logistic Regression & Feature Engineering
  4. 08:50 LLM Foundations: Transformer Architecture & Pre-training
  5. 10:00 Fine-tuning LLMs: Leveraging Pre-trained Models
  6. 13:30 Quantization for GPU Memory Efficiency
  7. 16:40 Parameter Efficient Fine-Tuning (PEFT) / LoRA
  8. 18:50 Advantages & Trade-offs of Fine-tuning LLMs

Finetuning Large Language Models (LLMs) for Security Log Detections

Speakers: Wilson Tang

Conference: BSidesSF 2024

YouTube: https://www.youtube.com/watch?v=0V-1URvcrd8

Overview

This talk, presented by Wilson Tang, a Machine Learning Engineer on the threat hunting team at Adobe, delves into the innovative application of Large Language Models (LLMs) for security log detections. Specifically, Tang explores the process of fine-tuning LLMs to identify command obfuscation, a prevalent technique used by adversaries to evade traditional security controls. The presentation highlights the limitations of existing rule-based and conventional machine learning approaches in detecting sophisticated and dynamic threats, positioning LLMs as a powerful new tool in the defender's arsenal.

The core of the discussion revolves around moving LLMs beyond their common use cases in text generation (like personal assistants) to perform classification tasks within a security context. Tang meticulously outlines the technical steps involved in fine-tuning a pre-trained LLM, such as Llama 2 7B, for this specific purpose. He emphasizes the practical considerations, including quantization for memory efficiency and the use of Parameter Efficient Fine-Tuning (PEFT) methods like LoRA, to make the process accessible and less computationally intensive. The talk ultimately aims to equip security practitioners with a deeper understanding of how LLMs can be integrated into security workflows to enhance detection capabilities against evasive threats.

Background

▶ Watch: Traditional Log Detections & Limitations (02:00)

Traditional security log detections predominantly operate within a rule-based paradigm. These systems rely on predefined signatures to identify malicious activity, such as specific malware process executions or logins from blacklisted IP addresses or Autonomous System Numbers (ASNs). While effective for known threats, this approach suffers from significant limitations when confronted with more sophisticated or "dynamic" attack techniques. Attackers can often craft their actions to subtly deviate from known signatures, rendering these rule sets ineffective.

One prominent example of such an evasive technique is command obfuscation. As Tang explains, command obfuscation is a method where a standard command line is intentionally made difficult to read, yet it retains its original functionality. This can involve injecting special characters, using encoding schemes like Base64, or employing various other tricks to alter the command's appearance without changing its execution. The primary danger of obfuscation is its ability to bypass mature, signature-based detections that are designed to look for specific patterns in command lines. While not all obfuscation is malicious (e.g., Base64 encoding can be used to shorten commands for operational purposes), its malicious use poses a significant challenge for security teams. The talk focuses solely on detecting obfuscation, not de-obfuscating commands, though that is acknowledged as another potential LLM use case.

Prior to the advent of LLMs, several approaches have been attempted to tackle the problem of command obfuscation detection, each with its own set of trade-offs:

  1. Rule-Based Logic: A simple, non-machine learning approach might involve creating rules based on characteristics like the percentage of special characters in a command line (e.g., "does this command contain 50%, 70%, or 80% special characters?"). However, as Tang points out, this method is inherently inflexible. It requires extensive manual tuning, often leading to a proliferation of complex rules that struggle to generalize across the broad spectrum of obfuscation techniques. It's difficult to curate a comprehensive set of rules that can reliably detect all forms of obfuscation.
  1. Logistic Regression: Stepping into the realm of machine learning, logistic regression offers a more sophisticated approach. This technique learns an equation that outputs a probability between 0 and 1, classifying data points into one of two categories (e.g., obfuscated or not obfuscated). The critical step here is feature engineering, where practitioners manually select and extract relevant features from the data. Examples of such features for command lines include character frequency (A-Z), special character frequency (semicolons, commas, periods), string length, entropy (randomness of a string), and whitespace density. While more adaptable than simple rules, logistic regression's effectiveness heavily relies on the quality and completeness of the engineered features. Moreover, it can become less effective when dealing with a very large number of features.
  1. Traditional Deep Learning (Neural Networks): A more advanced machine learning technique, deep learning with neural networks, addresses some of the limitations of logistic regression. A key advantage is the reduced need for manual feature engineering. Instead, deep learning models typically require only a vector representation of the input data. For command lines, this might involve converting each character into a numerical value (e.g., 'W' as 24, 'H' as 15, etc., for "whoami"). This vectorized data is then passed through a neural network, which Tang crudely summarizes as a "giant matrix multiplication operation." This approach offers greater generalization capabilities than logistic regression, but it introduces a new challenge: experimenting with and finding the optimal model architecture that best fits the specific data and task.

The emergence of Large Language Models (LLMs) has revolutionized natural language processing and, increasingly, other domains. Tang attributes the rapid advancement of LLMs to two primary factors:

  • The Transformer architecture: Introduced in the seminal 2017 paper "Attention Is All You Need," this architecture forms the backbone of modern LLMs. It allows models to process sequences of data (like text) efficiently and capture long-range dependencies.
  • Massive scale training: LLMs are trained on unprecedented volumes of data, often terabytes or even petabytes, over many days or weeks using powerful GPUs. This extensive pre-training enables them to develop a deep understanding of language patterns and structures.

Given the immense computational cost and time required to pre-train these models from scratch, the concept of fine-tuning becomes highly attractive. Fine-tuning involves taking a pre-trained LLM (a "model checkpoint" from the internet) and continuing its training on a smaller, custom dataset specific to the desired task. This approach leverages the general language understanding already acquired by the LLM, requiring only a relatively small custom dataset to "tweak" its parameters for a specialized application, such as security log classification. This avoids the need to "reinvent the wheel" and makes advanced LLM capabilities accessible for specific use cases.

Key Findings

▶ Watch: ML Techniques: Logistic Regression & Feature Engineering (06:00)

The presentation by Wilson Tang highlights several key findings regarding the application of fine-tuning Large Language Models for security log detections:

  • LLMs are adaptable for security classification tasks: Contrary to their primary association with text generation, LLMs can be effectively repurposed for binary classification problems in security, such as distinguishing between obfuscated and non-obfuscated commands. This expands the utility of LLMs beyond conversational AI into critical defensive operations.
  • Fine-tuning offers a practical path to specialized LLM use: Leveraging pre-trained LLMs and fine-tuning them on custom security datasets is a more efficient and less resource-intensive approach than training models from scratch. This makes advanced LLM capabilities accessible to security teams without requiring massive computational infrastructure for initial model development.
  • Parameter Efficient Fine-Tuning (PEFT) is crucial for resource optimization: Methods like LoRA (Low-Rank Adaptation), a specific implementation of PEFT, significantly reduce the number of parameters that need to be updated during fine-tuning. This dramatically lowers the computational cost, GPU memory requirements, and training time, making it feasible to fine-tune large models (e.g., 7 billion parameters) on more modest hardware.
  • Quantization is essential for memory management: Reducing the precision of model parameters (e.g., from 4 bytes to 4 bits) through quantization drastically cuts down GPU memory consumption. This enables the use of less expensive, consumer-grade GPUs for fine-tuning, democratizing access to LLM-based security solutions.
  • Data quality becomes the paramount factor: When fine-tuning pre-trained LLMs, the focus shifts away from complex tasks like feature engineering or model architecture experimentation. Instead, the quality, relevance, and structure of the custom training dataset become the most critical determinants of the fine-tuned model's performance.
  • LLMs provide a new tool for sophisticated threat detection: Fine-tuned LLMs offer a promising avenue to develop more robust and adaptable detections against evasive techniques like command obfuscation, which often bypass traditional rule-based or simpler machine learning methods. This enhances the overall maturity of security operations.

Technical Deep Dive

▶ Watch: Fine-tuning LLMs: Leveraging Pre-trained Models (10:00)

The technical deep dive presented by Wilson Tang focuses on the practical steps and considerations for fine-tuning an LLM for command obfuscation detection, primarily using the Hugging Face ecosystem.

1. LLM Selection and Sourcing:

The specific LLM chosen for this case study is Llama 2 7B. Tang notes that while Llama 2 was current at the time of the project, newer models like Llama 3 and Microsoft Phi-3 (a smaller, efficient LLM) are continually being released. The model is sourced from Hugging Face, which is described as a popular repository for open-source machine learning models and datasets, akin to GitHub for code.

2. Quantization for Memory Efficiency:

A critical first step before importing the model is setting up quantization. This technique is vital because LLMs, even those with "only" 7 billion parameters, demand significant GPU memory.

  • Problem: Typically, LLM parameters are stored with 4 bytes of precision. For Llama 2 7B, this translates to 7 billion parameters * 4 bytes/parameter = 28 GB of GPU memory just to load the model, excluding data. This necessitates expensive GPUs like A100s or H100s.
  • Solution: Quantization reduces the precision of these parameters, often from 4 bytes to 4 bits. This dramatically reduces memory footprint: 7 billion parameters * 4 bits/parameter = 3.5 GB of GPU memory. This reduction makes fine-tuning feasible on more affordable GPUs. The talk includes a mention of configuration settings to enable this.

3. Tokenizer Importation:

Next, the tokenizer for Llama 2 is imported.

  • Purpose: The tokenizer converts raw text input (like a command line) into a numerical vector representation (tokens) that the LLM can process. This automates the "vectorization" step that was previously a manual or complex process in traditional deep learning.
  • Necessity: It's crucial to use the specific tokenizer that the LLM was originally trained on. Passing in custom tokenized examples will not work because the model's internal representations are tied to its pre-training tokenizer.

4. Data Preparation and Prompt Engineering:

The quality of the custom dataset is paramount for effective fine-tuning.

  • Data Acquisition: The training data for command obfuscation detection can be gathered from various sources:
  • Command line events from the target environment.
  • Synthetically generated obfuscated commands using open-source or manual obfuscation tools.
  • Real-world obfuscated commands if available from incident response or threat intelligence.
  • Prompt Template: The raw command line data is then formatted into a structured prompt template for the LLM. This "hacky" approach simulates binary classification by instructing the LLM to output "yes" or "no." A sample template structure is provided:
  • Description: "Below is an instruction that describes a binary classification task for command obfuscation detection."
  • Instruction: "Analyze the following bash command line. If the command is obfuscated, output yes, or if the command is not obfuscated, output no."
  • Input: [The actual command line]
  • Output: yes or no (this is the label for training).

5. Pre-Fine-tuning Model Testing:

A best practice emphasized by Tang is to test the raw, pre-trained LLM's performance before fine-tuning.

  • Rationale: If the LLM performs adequately out-of-the-box for the specific task, fine-tuning might be unnecessary.
  • Method: The raw text prompt (without the expected "yes" or "no" output) is passed to the tokenizer, then to the LLM. A key parameter, Max new tokens = 1, is set to ensure the model generates only a single token (ideally "yes" or "no"). This allows for establishing baseline metrics before any custom training.

6. Fine-tuning with Parameter Efficient Fine-Tuning (PEFT):

The core of the fine-tuning process involves Parameter Efficient Fine-Tuning (PEFT).

  • Concept: Instead of updating all 7 billion parameters of the LLM, PEFT methods fine-tune only a small, select subset of parameters. This significantly speeds up the training process and reduces computational requirements.
  • Implementation: The talk specifically mentions LoRA (Low-Rank Adaptation) as a popular PEFT implementation. LoRA achieves parameter reduction through "matrix low-rank operations," though the mathematical details are not delved into. This involves setting specific configuration parameters for LoRA.
  • Training Arguments: Standard machine learning training arguments are defined, including per_device_training_batch_size, gradient_accumulation_steps, learning_rate, and max_steps.
  • Trainer Setup: All components are then assembled into a Hugging Face Trainer object: the quantized model, the training dataset, the PEFT (LoRA) configuration, the tokenizer, and the training arguments.
  • Execution: The trainer.train() method initiates the fine-tuning process. The duration of this step can vary from minutes to hours, depending on the GPU's power and the size of the dataset.

7. Post-Fine-tuning Evaluation:

After training, the fine-tuned model's outputs are evaluated using metrics similar to the pre-fine-tuning test, allowing for a comparison of performance improvement.

Advantages of Fine-tuning:

Tang highlights several key advantages of this fine-tuning approach:

  • Abstraction of ML complexities: It largely eliminates the need for manual feature engineering, model architecture experimentation, and manual data vectorization.
  • Focus on data quality: The primary effort shifts to curating a high-quality, relevant dataset, which is often cited as the most crucial step in any machine learning project.

Trade-offs of Fine-tuning:

Despite the advantages, Tang provides a balanced view by discussing the trade-offs:

  • Parameter overkill: For a simple binary classification task (yes/no), 7 billion parameters might be excessive. Smaller, more specialized models could potentially achieve similar results with fewer resources.
  • Resource cost: While significantly less than pre-training, fine-tuning still requires dedicated GPU memory and computational cycles, which translates to financial cost.

However, Tang concludes that these trade-offs are often accepted for the "ease of use" and the abstraction of complex ML processes that fine-tuning provides, making it a valuable new tool for security practitioners.

Demo / Proof of Concept

▶ Watch: Quantization for GPU Memory Efficiency (13:30)

The presentation provides a detailed walkthrough of the conceptual code and the technical process for fine-tuning an LLM for command obfuscation detection. It outlines the necessary steps, from data preparation and quantization to the application of Parameter Efficient Fine-Tuning (PEFT) methods like LoRA. However, the talk does not include a live demonstration of the fine-tuned model in action, nor does it present specific performance metrics or a visual proof-of-concept of its detection capabilities. The focus is primarily on the methodology and the underlying technical implementation rather than a direct showcase of results. The speaker mentions that after running trainer.train(), one can "run your model outputs after that run some metrics on it like we did before and see how your model improved after that," implying that such evaluation would be performed post-training, but no specific demo or results are shared during the talk itself.

Defensive Implications

▶ Watch: Advantages & Trade-offs of Fine-tuning LLMs (18:50)

The application of fine-tuned Large Language Models for security log detections, particularly for command obfuscation, carries significant defensive implications for organizations:

  • Enhanced Detection of Evasive Techniques: Fine-tuned LLMs provide a powerful new capability to detect sophisticated and evasive attacker techniques, such as command obfuscation, that often bypass traditional rule-based or signature-based security controls. This allows defenders to catch threats that would otherwise go unnoticed.
  • Adaptability to Evolving Threats: Unlike static rule sets, LLMs can be continuously fine-tuned with new data reflecting emerging obfuscation methods or specific threats observed in an organization's environment. This inherent adaptability makes them more resilient against evolving attacker tactics.
  • Reduced Manual Effort in Detection Engineering: By abstracting away complex tasks like feature engineering and model architecture design, security analysts and machine learning engineers can focus more on curating high-quality data and understanding threat intelligence. This streamlines the development of new detections and reduces the manual overhead associated with maintaining complex rule sets.
  • Proactive Threat Hunting: The ability to accurately classify obfuscated commands can empower threat hunters to identify suspicious activities that might indicate an ongoing compromise or the presence of advanced persistent threats (APTs). This shifts defensive posture from reactive to more proactive.
  • Leveraging Existing Investments: Organizations can leverage the vast research and development poured into general-purpose LLMs by fine-tuning them for specific security tasks. This avoids the need to build complex models from scratch, accelerating time-to-detection for novel threats.
  • Democratization of Advanced ML for Security: Techniques like quantization and Parameter Efficient Fine-Tuning (PEFT) make LLM-based detections more accessible by reducing the prohibitive GPU memory and computational requirements. This allows more organizations, even those with limited access to high-end hardware, to implement advanced machine learning in their security operations.
  • Growth of the Security Industry: As Wilson Tang noted in the Q&A, AI and LLMs will likely "help propel us to do new types of detections," enabling security analysts to have "more tools to do more mature types of detection." This suggests a potential growth in the security industry, with new roles and skill sets emerging around AI/ML integration.
  • Strategic Tool Consideration: Defenders should view fine-tuned LLMs as another valuable tool in their detection toolkit. The decision to deploy such a system should involve a careful evaluation of its benefits against the computational costs and resource requirements, ensuring it aligns with the organization's overall security strategy and risk appetite.

Key Takeaways

  • LLMs for Classification: Large Language Models (LLMs) can be effectively fine-tuned for binary classification tasks in security, moving beyond their common use cases in text generation.
  • Addressing Evasive Threats: Fine-tuned LLMs offer a powerful new approach to detect sophisticated and evasive techniques like command obfuscation, which often bypass traditional rule-based or simpler machine learning methods.
  • Resource Efficiency through PEFT and Quantization: Techniques such as Parameter Efficient Fine-Tuning (PEFT), specifically LoRA, and quantization are crucial for making LLM fine-tuning computationally feasible and accessible on more modest hardware by significantly reducing GPU memory and processing requirements.
  • Data Quality is Paramount: When fine-tuning pre-trained LLMs, the primary focus shifts from complex model architecture experimentation and feature engineering to the quality, relevance, and structure of the custom training dataset.
  • Streamlined Detection Development: Fine-tuning LLMs abstracts away many traditional machine learning complexities, allowing security practitioners to accelerate the development of new and more adaptable threat detections.
  • New Tool in the Defender's Arsenal: Fine-tuned LLMs represent a valuable new tool for security teams, enabling more mature and sophisticated detection capabilities against dynamic threats, though their adoption requires careful consideration of computational trade-offs.

About the Speaker(s)

Wilson Tang (He/Him pronouns) is a Machine Learning Engineer on the threat hunting team at Adobe. He earned his Master's degree in Computer Science from the University of Washington in 2022. Wilson identifies as a proud second-generation Asian-American. In his personal time, he enjoys cooking, traveling to experience new foods and places, and playing video games, specifically Valorant.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk presents a solid, actionable approach to improving security log detections by fine-tuning Large Language Models (LLMs) for classification tasks, specifically command obfuscation detection. The speaker clearly outlines the limitations of traditional rule-based and simpler ML methods, then dives into the technical underpinnings of LLMs, fine-tuning, and practical considerations like quantization and PEFT. While the core LLM fine-tuning technique isn't groundbreaking, its detailed application to a persistent security problem is highly valuable for practitioners looking to move beyond brittle signature-based detections.

Heather Calloway (CISO) — STRONG ACCEPT

This presentation offers a clear and actionable technical path for security teams to enhance their detection capabilities against command obfuscation, a persistent challenge that undermines traditional rule-based systems. By demonstrating how to fine-tune Large Language Models (LLMs) for this specific classification task, the speaker provides a valuable tool that can significantly reduce an organization's exposure to advanced evasion techniques. While the talk focuses on technical implementation rather than explicit governance, the operational impact on a security program's resilience is substantial.

→ Top-rated talks at BSidesSF 2024

All talks from BSidesSF 2024