Incorporating Gradients to Rules: Towards Lightweight, Adaptive Provenance-based Intrusion Detection

Lingzhi Wang

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Network Security 1

Overview

In an era dominated by sophisticated Advanced Persistent Threats (APTs), traditional intrusion detection systems (IDS) struggle to keep pace with the evolving tactics of cyber attackers. This talk by Lingzhi Wang at the NDSS Symposium introduces a novel approach to provenance-based intrusion detection, addressing a critical dilemma: how to achieve both high accuracy and exceptional efficiency. Titled "Incorporating Gradients to Rules: Towards Lightweight, Adaptive Provenance-based Intrusion Detection," the work presents CAPTAIN, a system designed to transform the rigid, non-differentiable nature of rule-based detection into a flexible, learning-capable framework by integrating gradient-based optimization.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction to advanced attacks and provenance graphs
  2. 2:00 Comparing rule-based and embedding-based PIDS
  3. 3:15 The core motivation: adaptive, lightweight rule learning
  4. 4:00 Our inspiration: Optimizing rules using gradients
  5. 4:51 Key contribution: Differentiable rule-based detection
  6. 5:30 Challenges in traditional tag propagation systems
  7. 8:00 Converting rules into adaptive numerical parameters

Incorporating Gradients to Rules: Towards Lightweight, Adaptive Provenance-based Intrusion Detection

Speakers: Lingzhi Wang

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=uZDjrV2TCQY

Overview

In an era dominated by sophisticated Advanced Persistent Threats (APTs), traditional intrusion detection systems (IDS) struggle to keep pace with the evolving tactics of cyber attackers. This talk by Lingzhi Wang at the NDSS Symposium introduces a novel approach to provenance-based intrusion detection, addressing a critical dilemma: how to achieve both high accuracy and exceptional efficiency. Titled "Incorporating Gradients to Rules: Towards Lightweight, Adaptive Provenance-based Intrusion Detection," the work presents CAPTAIN, a system designed to transform the rigid, non-differentiable nature of rule-based detection into a flexible, learning-capable framework by integrating gradient-based optimization.

The core innovation of CAPTAIN lies in its ability to automatically fine-tune detection rules, moving beyond the static, human-expert-dependent configurations that plague conventional rule-based systems. By making the detection process differentiable, CAPTAIN leverages machine learning principles to adapt to new attack patterns and environmental shifts dynamically. This allows it to drastically reduce false alarms and enhance overall detection accuracy, all while maintaining the lightweight and explicable characteristics that make rule-based systems appealing.

This research is particularly significant for the cybersecurity landscape, offering a pathway to deploy more robust and adaptive intrusion detection capabilities without incurring the prohibitive computational costs typically associated with advanced, learning-based methods. For organizations facing increasingly stealthy and persistent threats, CAPTAIN represents a promising step towards more intelligent, efficient, and resilient defense mechanisms.

Background

▶ Watch: Introduction to advanced attacks and provenance graphs (0:00)

The landscape of cyber threats has evolved significantly, with APTs posing severe challenges to conventional security defenses. These attacks are characterized by their advanced nature, often employing zero-day exploits and customized malware; their persistence, remaining undetected for extended periods; and their targeted approach, focusing on specific high-value assets. Traditional signature-based or anomaly-based IDSs frequently falter against such threats due to their inability to detect novel attack patterns or correlate disparate system events effectively.

To counter these limitations, the cybersecurity community has increasingly turned to provenance graphs as a powerful tool for intrusion detection. A provenance graph provides a comprehensive, causal chain of system activities, where nodes represent system entities (e.g., processes, files, network sockets) and edges denote interactions between them (e.g., a process reading a file, sending data over a socket). This graph-based representation offers a wider scope and richer context, enabling the correlation of events across the entire system, which is crucial for identifying complex, multi-stage attacks.

Existing provenance-based intrusion detection systems can be broadly categorized into two types:

  1. Rule-based Provenance Systems (Rule-based PS): These systems define simple, static rules to model known malicious graph patterns. Their primary advantages include high speed, efficiency, and a lightweight operational footprint. Furthermore, their detection process is fully explicable, meaning analysts can easily understand why an alarm was triggered. However, their major drawback is often low accuracy due to rigid and inflexible rules that cannot adapt to new attack techniques or evolving system behaviors. They require constant, manual updates by human experts, which is unsustainable.
  1. Embedding-based Provenance Systems (Embedding-based PS): More active in recent years, these systems leverage embedding functions and graph neural networks (GNNs) to learn complex graph features from the provenance data. They possess powerful feature extraction capabilities and inherent learning mechanisms, allowing them to adapt to novel data patterns. However, this adaptability comes at a significant cost: they are computationally very expensive, suffer from long detection latency, and demand substantial CPU usage and memory usage. Crucially, their detection process is often inexplicable, acting as a "black box," which complicates incident response and forensic analysis.

This dichotomy presents a fundamental dilemma: security teams are forced to choose between the efficiency and explicability of rule-based systems and the adaptability and accuracy of embedding-based systems. The motivation behind CAPTAIN is to bridge this gap, asking: "Is it possible to maintain the lightweight nature and efficiency of rule-based detection while enabling dynamic and automatic rule learning?" The goal is to develop a system that can adjust its detection rules autonomously to improve results, offering the best of both worlds.

Key Findings

▶ Watch: The core motivation: adaptive, lightweight rule learning (3:15)

The central discovery and contribution of this work is the successful transformation of the traditionally non-differentiable, rule-based detection process into a differentiable function. This groundbreaking shift allows for the application of gradient-based optimization algorithms—a technique commonly employed in machine learning for parameter tuning—to automatically fine-tune intrusion detection rules. The proposed system, named CAPTAIN, directly addresses the long-standing dilemma between detection accuracy and efficiency in provenance-based IDS.

The key findings demonstrated by CAPTAIN include:

  • Significant Reduction in False Alarms: CAPTAIN achieves a drastic reduction in false alarms by a "huge margin" compared to recent baseline works. This is crucial for reducing analyst fatigue and allowing security teams to focus on genuine threats.
  • Maintained True Positive Rate: Despite the significant reduction in false positives, CAPTAIN successfully maintains a similar true positive rate to existing baselines, indicating its ability to detect all relevant attacks without compromise.
  • Exceptional Efficiency Gains: Compared to embedding-based provenance systems, CAPTAIN dramatically reduces detection latency, CPU usage, and memory usage by a "huge amount." This makes it suitable for real-world, high-throughput environments where embedding-based solutions are often impractical.
  • Minimal Overhead for Rule-based Systems: When compared to traditional rule-based systems, CAPTAIN introduces only "slightly additional overhead." This minimal increase is a small price to pay for the "significant reduction in false alarm" and "significant improvement on the accuracy" it provides.
  • Explicable and Granular Detection: Unlike black-box machine learning approaches, CAPTAIN retains the clear semantics and explicable detection process characteristic of rule-based systems. Furthermore, by assigning different parameters to individual nodes and edges, it achieves a finer granularity, enabling it to distinguish between very similar graph patterns.
  • Addressing Dependency Explosion: The system effectively addresses the tag or dependency explosion issue, a common challenge faced by many provenance-based intrusion detection systems where malicious tags can propagate too broadly, leading to an overwhelming number of alarms.

In essence, CAPTAIN demonstrates that it is possible to combine the speed and transparency of rule-based detection with the adaptability and precision of learning-based methods, setting a new benchmark for lightweight, adaptive, and accurate provenance-based intrusion detection.

Technical Deep Dive

▶ Watch: Our inspiration: Optimizing rules using gradients (4:00)

The technical foundation of CAPTAIN is rooted in an innovative adaptation of gradient-based parameter learning, drawing inspiration from the training methodologies of neural networks. The core idea is to establish a feedback loop where detection results inform the automatic adjustment of rules. This process involves four key steps: initializing rules, performing detection, calculating loss and gradients, and finally, updating the rules.

The system's approach can be summarized as transforming the non-differentiable rule-based detection process into a differentiable function, thereby enabling the use of gradient-based optimization algorithms to fine-tune detection rules automatically.

At the heart of CAPTAIN's detection mechanism is tag propagation, a technique akin to taint analysis. In this process:

  1. Initial Tag Assignment: Nodes in the provenance graph are initially assigned "tags," indicating their trust level (e.g., malicious/untrusted or benign/trusted). A critical challenge here is balancing the risk of excessive false alarms (if too many nodes are initially untrusted) against the risk of missing attacks (if too many are trusted).
  2. Tag Propagation: As system events occur, these tags propagate along the edges of the provenance graph. The challenge is to control this propagation to prevent an overwhelming "tag explosion" where malicious tags spread indiscriminately across the entire graph, rendering the detection system useless due to an abundance of false positives.
  3. Alarm Triggering: Finally, specific rules dictate when an alarm should be triggered based on the tags accumulated by nodes or specific graph patterns. Fine-tuning these rules to control sensitivity for various events or alarms is crucial.

Traditionally, all these rules—initialization, propagation, and triggering—had to be manually set and adjusted by human experts, making the system static and difficult to adapt.

CAPTAIN addresses these challenges by incorporating gradients through a novel parameterization of the rules:

Converting Rules to Differentiable Parameters

To enable gradient-based optimization, CAPTAIN converts the qualitative, human-defined rules into quantitative, numerical parameters that can be learned:

  1. Tag Initialization Rules $\rightarrow$ Integrity Score: Instead of binary trust assignments, each node is assigned an integrity score. This score is a numerical value that reflects the initial trustworthiness of a node. A higher score might indicate greater trust, while a lower score suggests potential maliciousness. This allows for a continuous spectrum of initial trust, rather than a rigid binary classification.
  2. Tag Propagation Rules $\rightarrow$ Propagation Rate: The propagation of tags is no longer a simple binary "propagate" or "don't propagate." Instead, each edge in the provenance graph is associated with a propagation rate. This numerical rate determines the extent to which a tag (or its influence) is transferred from one node to another. This fine-grained control prevents uncontrolled tag explosion and allows the system to learn which interactions are more indicative of malicious activity.
  3. Alarm Triggering Rules $\rightarrow$ Alarm Threshold: The conditions for triggering an alarm are replaced by an alarm threshold. When the accumulated maliciousness score or tag intensity of a node or a specific subgraph pattern exceeds this numerical threshold, an alarm is triggered. This allows for dynamic adjustment of alarm sensitivity.

These three types of adaptive parameters—integrity score, propagation rate, and alarm threshold—can be assigned with fine granularity to different nodes and edges on the provenance graph. This provides CAPTAIN with unparalleled control over the entire provenance-based intrusion detection system, moving beyond global, static rules.

Calculating and Updating Gradients

The ability to calculate gradients is fundamental to CAPTAIN's learning process. During the tag propagation phase, CAPTAIN meticulously records and updates gradients. When a tag's numerical value or its influence changes on a node, the system records the corresponding gradients. Essentially, a "gradient table" is maintained, which tracks the influence of each rule parameter (integrity score, propagation rate, alarm threshold) on the final detection result. This table quantifies how a small change in a specific parameter would affect the system's output.

Training and Rule Parameter Updates

The training process for CAPTAIN is iterative and designed to minimize misdetections on a training dataset:

  1. Loss Calculation: When a misdetection occurs (e.g., a false alarm is triggered on a benign event, or a true attack is missed), a loss value is calculated. This loss quantifies the discrepancy between the system's output and the ground truth.
  2. Gradient Lookup and Application: The system then consults the gradient table to identify which rule parameters are most significantly related to the observed misdetection.
  3. Parameter Update: Using the calculated gradients, the values of the corresponding parameters are updated. For instance, if a false alarm was due to an overly sensitive alarm threshold or an aggressive propagation rate, the gradients would guide the system to adjust these parameters to reduce sensitivity.
  4. Iteration: This process of detection, loss calculation, gradient lookup, and parameter update is repeated iteratively until the system converges, meaning the parameters have been optimized to achieve the best possible detection accuracy on the training data.

The key advantage of this approach is that the computationally intensive training phase, where gradients are recorded and parameters are updated, can be performed offline. Once the optimal parameters are learned, the system operates in its detection phase as a highly efficient, traditional rule-based PS, but with significantly enhanced accuracy and adaptability, without the need to record gradients in real-time. This separation ensures that the real-time detection overhead remains minimal.

Demo / Proof of Concept

▶ Watch: Challenges in traditional tag propagation systems (5:30)

While the talk did not feature a live, interactive demo in the traditional sense, the efficacy and efficiency of CAPTAIN were rigorously demonstrated through extensive evaluations on both public and private datasets. These evaluations serve as the empirical proof of concept, validating the system's claims and showcasing its practical applicability.

The evaluation results highlight several compelling advantages of CAPTAIN:

  1. False Alarm Reduction: Compared to recent baseline works, CAPTAIN achieved a "huge margin" reduction in false alarms. This is a critical metric for any IDS, as high false positive rates lead to alert fatigue and wasted resources for security analysts. The ability to maintain a similar true positive rate (detecting nearly all actual attacks) alongside this reduction underscores CAPTAIN's superior accuracy.
  2. Efficiency Benchmarks: The system's performance metrics were particularly impressive when compared to embedding-based Provenance Systems. CAPTAIN demonstrated a "huge amount" reduction in detection latency, CPU usage, and memory usage. This makes CAPTAIN a viable solution for real-time monitoring and deployment in environments where computational resources are constrained, unlike many resource-intensive embedding-based approaches.
  3. Overhead Comparison with Rule-based PS: When benchmarked against existing traditional rule-based Provenance Systems, CAPTAIN incurred only "slightly additional overhead." This minimal increase in operational cost is justified by the "significant reduction in false alarm" and the "significant improvement on the accuracy" that CAPTAIN delivers, effectively offering a substantial upgrade to rule-based methods without sacrificing their inherent lightweight nature.
  4. Case Studies: The paper also presents several case studies that further illustrate CAPTAIN's capabilities. These studies confirm that, as a rule-based detection system, CAPTAIN provides clear semantics and an explicable detection process for each alarm. This transparency is vital for incident response and forensic analysis, allowing security teams to understand the why behind an alert. Furthermore, the fine-grained control offered by assigning different parameters to individual nodes and edges enables CAPTAIN to distinguish between very similar graph patterns, enhancing its precision. Critically, the system successfully addresses the tag or dependency explosion issue, a common challenge in provenance-based systems where malicious tags can spread uncontrollably, leading to an overwhelming number of alarms. CAPTAIN's adaptive propagation rates effectively mitigate this problem.

The comprehensive evaluation results and supporting case studies conclusively prove CAPTAIN's ability to deliver high-accuracy, adaptable intrusion detection with exceptional efficiency, successfully bridging the gap between traditional rule-based and modern embedding-based approaches.

Defensive Implications

▶ Watch: Converting rules into adaptive numerical parameters (8:00)

CAPTAIN's innovative approach to provenance-based intrusion detection carries significant implications for cybersecurity defenders, offering practical strategies to enhance their security posture against sophisticated threats.

Firstly, organizations can leverage CAPTAIN to deploy an intrusion detection system that is both lightweight and highly accurate. This addresses a long-standing trade-off, enabling effective detection of Advanced Persistent Threats (APTs) and novel attack techniques without the prohibitive computational costs associated with deep learning models. The reduced detection latency and lower CPU usage and memory usage make it feasible for real-time monitoring even in environments with limited resources, such as edge devices or smaller data centers.

Secondly, the drastic reduction in false alarms is a game-changer for security operations centers (SOCs). False alarms are a major source of alert fatigue, leading to missed genuine threats and inefficient resource allocation. By filtering out noise more effectively, CAPTAIN allows security analysts to focus their attention on truly suspicious activities, improving their productivity and the overall effectiveness of incident response. This translates directly to a stronger, more proactive defense.

Thirdly, CAPTAIN retains the crucial advantage of explicability. Unlike "black box" machine learning models, CAPTAIN's alarms come with clear semantics, rooted in its rule-based foundation. When an alarm is triggered, defenders can trace back the specific provenance graph patterns and the parameters (integrity scores, propagation rates, alarm thresholds) that led to the detection. This transparency is invaluable for understanding the nature of an attack, performing root cause analysis, and refining defensive strategies. It empowers incident responders to make informed decisions quickly.

Fourthly, the system's adaptive rule learning capability reduces the burden on human experts. Traditionally, rule-based systems require constant manual updates and adjustments to cope with evolving threats and changes in the environment. CAPTAIN's gradient-based optimization automates this process, allowing the system to dynamically fine-tune its rules and adapt to new attack patterns without constant human intervention. This makes the IDS more resilient and less prone to becoming outdated.

Finally, the ability to achieve finer granularity in detection and address the dependency explosion issue means that CAPTAIN can discern subtle malicious activities that might be missed by cruder detection methods. This enables defenders to detect advanced, stealthy attacks that deliberately mimic benign behaviors or exploit complex inter-dependencies within the system, providing a robust layer of defense against evasive adversaries.

In summary, CAPTAIN offers a compelling solution for organizations seeking an effective, efficient, and intelligent intrusion detection system that can keep pace with the dynamic threat landscape while providing clear, actionable insights for defenders.

Key Takeaways

  • CAPTAIN bridges the accuracy-efficiency gap: The system effectively resolves the long-standing dilemma between the high accuracy of embedding-based provenance systems and the efficiency and explicability of rule-based systems.
  • Gradient-based optimization for rule learning: CAPTAIN innovatively transforms the non-differentiable rule-based detection process into a differentiable function, enabling automatic, adaptive rule tuning using gradient descent algorithms.
  • Numerical parameterization of rules: Traditional human-defined rules (initialization, propagation, triggering) are replaced by adaptive numerical parameters: integrity scores, propagation rates, and alarm thresholds, providing fine-grained control over detection logic.
  • Significant performance improvements: Evaluations demonstrate a "huge margin" reduction in false alarms while maintaining true positive rates, alongside "huge amount" reductions in detection latency, CPU usage, and memory usage compared to embedding-based systems.
  • Lightweight with minimal overhead: CAPTAIN introduces only "slightly additional overhead" compared to traditional rule-based systems, offering substantial accuracy improvements without sacrificing efficiency.
  • Explicable and adaptive detection: The system maintains clear semantics for alarms, aiding incident response, and can distinguish subtle attack patterns while automatically adapting its rules to evolving threats. The training process is performed offline, ensuring high efficiency during real-time detection.

About the Speaker(s)

The talk "Incorporating Gradients to Rules: Towards Lightweight, Adaptive Provenance-based Intrusion Detection" was presented by Lingzhi Wang. Based on the information provided in the talk metadata and transcript, Lingzhi Wang is the researcher behind this innovative work. No specific affiliation or title beyond the name was provided in the context of this presentation.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

CAPTAIN is legitimate systems security research — the core idea of making taint-propagation rules differentiable so you can gradient-descend your way to better parameters is genuinely clever and not something I've seen packaged this way before. The problem it's solving (rule-based provenance IDS drowning in false positives, embedding-based IDS too heavy for production) is real and the parameterization approach (integrity scores, propagation rates, alarm thresholds as learnable scalars) is a clean framing. But the transcript reads like a summary document, not a talk, and the evaluation section is all qualitative handwaving — 'huge margin,' 'huge amount,' 'slightly additional overhead' — no…

Heather Calloway (CISO) — WEAK

Technically credible research on adaptive provenance-based intrusion detection that solves a real problem — the accuracy-efficiency tradeoff in rule-based IDS. But the entire presentation lives at the systems-research layer, and the gap to any operational, governance, or institutional context is never crossed.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025