FLASH: A Comprehensive Approach to Intrusion Detection via Provenance Graph Representation Learning

Mati Ur Rehman, Hadi Ahmadi, Wajih Ul Hassan

IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 4

Overview

In the realm of modern cybersecurity, the detection of highly sophisticated cyberattacks, particularly Advanced Persistent Threats (APTs), remains a formidable challenge. These stealthy and protracted attacks target critical organizations and governments, incurring significant financial losses, with the IBM Data Breach Report 2023 citing an average global cost of $4.45 million per APT attack—a 15% increase over three years. Moreover, 97% of organizations have reported an increase in cyber threats since 2022, with the average lifecycle of an attack spanning 277 days from identification to containment. Addressing this pressing need, the talk "FLASH: A Comprehensive Approach to Intrusion Detection via Provenance Graph Representation Learning" introduces an innovative anomaly-based intrusion detection system designed to combat these evolving threats.

Watch on YouTube

Visual summary for FLASH: A Comprehensive Approach to Intrusion Detection via Provenance Graph Representation Learning by Mati Ur Rehman, Hadi Ahmadi, Wajih Ul Hassan
Visual summary for FLASH: A Comprehensive Approach to Intrusion Detection via Provenance Graph Representation Learning by Mati Ur Rehman, Hadi Ahmadi, Wajih Ul Hassan

Key moments

  1. 0:00 Introduction to APTs and challenges for intrusion detection
  2. 2:00 Limitations of existing provenance-based intrusion detection systems
  3. 3:10 Introducing FLASH: a comprehensive, efficient, and accurate solution
  4. 6:00 Flash's semantic and temporal information featurization technique
  5. 8:00 GNN embedding database for efficient graph representation learning

FLASH: A Comprehensive Approach to Intrusion Detection via Provenance Graph Representation Learning

Speakers: Mati Ur Rehman; Hadi Ahmadi; Wajih Ul Hassan

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=RpsSAm4FJTs

Overview

In the realm of modern cybersecurity, the detection of highly sophisticated cyberattacks, particularly Advanced Persistent Threats (APTs), remains a formidable challenge. These stealthy and protracted attacks target critical organizations and governments, incurring significant financial losses, with the IBM Data Breach Report 2023 citing an average global cost of $4.45 million per APT attack—a 15% increase over three years. Moreover, 97% of organizations have reported an increase in cyber threats since 2022, with the average lifecycle of an attack spanning 277 days from identification to containment. Addressing this pressing need, the talk "FLASH: A Comprehensive Approach to Intrusion Detection via Provenance Graph Representation Learning" introduces an innovative anomaly-based intrusion detection system designed to combat these evolving threats.

Presented by Mati Ur Rehman from the University of Virginia, in collaboration with Hadi Ahmadi from Coric and advisor Wajih Ul Hassan also from the University of Virginia, FLASH offers a robust solution to the limitations of existing provenance-based detection systems. The system focuses on leveraging provenance graphs—causal models derived from system audit logs—to represent intricate system activities and relationships between entities like processes, files, and network sockets. FLASH distinguishes itself by achieving high detection accuracy and operational efficiency, specifically by generating fine-grained, node-level alerts that empower security analysts to pinpoint exact attack entities, thus overcoming the challenges of high false alarms and low operational efficiency prevalent in prior art.

Background

▶ Watch: Introduction to APTs and challenges for intrusion detection (0:00)

The landscape of cyber defense has increasingly turned to provenance-based intrusion detection systems (PIDs) as a promising avenue for detecting sophisticated attacks. These systems operate by ingesting vast quantities of system audit logs, which record various low-level system activities. From these logs, PIDs construct provenance graphs, which are directed acyclic graphs where nodes represent system entities (e.g., processes, files, network sockets) and edges denote causal relationships or system calls between them. This graph representation allows for a holistic view of system behavior and the propagation of events, crucial for understanding complex attack chains.

However, existing PIDs can be broadly categorized into three types, each with inherent limitations. The first category comprises rule-based systems, which rely on predefined attack signatures. While effective against known threats, they are inherently vulnerable to zero-day attacks—novel exploits for which no signature exists. Furthermore, crafting and maintaining these rules demands significant manual effort and domain expertise. The second type includes learning-based graph-level systems. These systems aim to detect anomalous subgraphs within a larger provenance graph, leveraging machine learning techniques. While more adaptable than rule-based systems, they often produce coarse-grained alerts, making it difficult for security analysts to identify the precise malicious entities or events within a detected anomaly, thus hindering efficient incident response.

The third and most advanced category consists of learning-based node/edge-level systems. These systems attempt to address the granularity issue by generating fine-grained alerts at the level of individual nodes or edges, and they do not require pre-existing attack signatures, making them potentially capable of detecting zero-day threats. Despite these advantages, prior node/edge-level PIDs have struggled with two critical challenges: low operational efficiency due to the computational intensity of processing large provenance graphs and generating high false alarms, which can overwhelm security analysts and erode trust in the detection system. FLASH is specifically engineered to confront and overcome these twin challenges, aiming to deliver a practical and effective solution for real-time APT detection.

Key Findings

▶ Watch: Limitations of existing provenance-based intrusion detection systems (2:00)

FLASH presents several compelling key findings that address long-standing challenges in provenance-based intrusion detection. The core contributions revolve around its ability to achieve high detection accuracy and superior operational efficiency for timely and precise APT detection, even against zero-day attacks, without requiring expert knowledge or attack signatures.

A primary finding is that high false alarms, a pervasive issue in existing systems, can be significantly reduced through the intelligent use of semantic attributes within provenance graphs. By encoding rich contextual information such as process names, file paths, and network IPs, FLASH can build a more nuanced and accurate model of benign system behavior, thereby improving anomaly discrimination. This semantic enrichment is crucial for distinguishing legitimate, albeit unusual, activities from genuine malicious intent.

Furthermore, FLASH demonstrates a breakthrough in operational efficiency by introducing a novel embedding recycling technique for Graph Neural Network (GNN) embeddings. Traditional GNNs, while powerful for capturing graph structure, suffer from high computational overhead, especially with large and dynamic provenance graphs. FLASH mitigates this by recognizing that many nodes exhibit consistent behavioral patterns across different execution runs. By storing and intelligently reusing pre-computed GNN embeddings for these stable nodes, the system achieves substantial speed improvements during the detection phase.

The practical efficacy of FLASH was rigorously evaluated against existing state-of-the-art systems using four diverse audit log datasets from DARPA and the research community: DARPA E3, DARPA OpTC, StreamSpot, and Unicorn. In comparative studies, FLASH consistently demonstrated superior detection performance. Against ThreatRace, an existing node-level system, FLASH achieved higher precision, lower false alarms, and a better F-score. Similarly, when compared to Unicon, a graph-level system, FLASH exhibited better precision, recall, and F-score. Critically, the operational efficiency experiments revealed that FLASH achieved up to three times speed improvements compared to ThreatRace, while utilizing minimal system resources for threat detection. The system's robustness was further validated through experiments demonstrating its resilience to adversarial mimicry attacks, where attackers attempt to disguise malicious activities as benign ones. These findings collectively establish FLASH as a highly effective, efficient, and robust solution for modern intrusion detection.

Technical Deep Dive

▶ Watch: Introducing FLASH: a comprehensive, efficient, and accurate solution (3:10)

FLASH's technical architecture is meticulously designed to address the challenges of semantic neglect, computational overhead, and high false alarms in provenance-based intrusion detection. The system operates in two distinct phases: a training phase and a detection phase.

The training phase begins with Provenance Graph Construction. Raw system audit logs are converted into a provenance graph. To optimize this representation, FLASH applies pre-processing techniques, including merging multiple edges of the same type between two nodes and removing non-persistent, execution-specific information from node attributes.

A critical innovation lies in the Semantic and Temporal Featurization component. Existing node-level PIDs often neglect the rich semantic attributes present in audit logs, such as process names, file paths, and network IP addresses, which significantly impacts detection accuracy. FLASH utilizes Word2Vec, an unsupervised technique, to encode these text attributes into vector space embeddings. Recognizing that the vanilla Word2Vec model does not inherently preserve word order—a crucial aspect for capturing the temporal ordering of events in audit logs and provenance graphs—FLASH introduces a novel approach. It first constructs causal sequences by considering temporally ordered node events and their attributes with neighbors. Each word in this sequence is then encoded into a vector using Word2Vec. To imbue these word encodings with temporal information, FLASH employs a positional encoding scheme, combining them into a comprehensive, temporal-ordering-aware feature vector for the anomaly detector.

Next, for encoding the graph structure around nodes, FLASH employs a Graph Neural Network (GNN) model, specifically GraphSAGE. This model takes the Word2Vec semantic encodings as initial node features and is trained in an unsupervised manner to learn structural node embeddings. A significant challenge with GNNs is their computational overhead, as iterative message passing for encoding graph structure can increase computational cost exponentially with larger graphs. Existing optimization techniques like subsampling often lead to information loss, detrimental to detection performance.

To overcome this, FLASH introduces the GNN Embedding Database, a core component of its embedding recycling technique. The observation is that many nodes in a provenance graph exhibit consistent patterns of behavior, meaning their local graph structure remains relatively static across various execution runs. FLASH capitalizes on this by storing pre-computed GNN embeddings for these stable nodes as key-value pairs. The key represents the node identifier, and the value comprises the GNN embedding along with the node's neighborhood set. This database significantly speeds up inference during the detection phase.

Finally, the Lightweight Classifier Training step utilizes both the semantic features (from Word2Vec) and the stored structural embeddings (from the GNN Embedding Database) to train a lightweight classifier model. This classifier is designed for efficient real-time threat detection.

The detection phase begins by taking incoming system logs, converting them into a provenance graph, and performing semantic featurization using the pre-trained Word2Vec model. For each incoming node, FLASH consults the GNN Embedding Database. It extracts the current neighborhood context of the node and compares it with the stored neighborhood set from the database using the Jaccard index matrix. If the Jaccard index equals one (indicating an identical neighborhood context), FLASH reuses the stored GNN embedding. Otherwise, if the context has changed, it relies solely on the semantic features for anomaly detection, avoiding the costly recomputation of GNN embeddings for dynamic or novel structures.

The Lightweight Classifier then takes the concatenated Word2Vec embeddings and the (potentially recycled) GNN embeddings as input. Its task is to classify nodes into their respective system types. The core intuition for anomaly detection is that misclassified nodes are identified as anomalies. The model's inability to correctly classify these nodes suggests that it has not encountered these specific patterns or behaviors during its training on benign system activity, thus flagging them as potential threats. This approach allows FLASH to generate fine-grained, node-level alerts, directly addressing the analyst's need for precise attack entity identification.

Demo / Proof of Concept

▶ Watch: Flash's semantic and temporal information featurization technique (6:00)

While the talk did not feature a live, interactive demonstration of FLASH in action, the speakers provided a comprehensive account of the system's rigorous evaluation, serving as its proof of concept. The efficacy and efficiency of FLASH were validated through extensive experiments conducted on four distinct audit log datasets: DARPA E3, DARPA OpTC, StreamSpot, and Unicorn. These datasets, sourced from DARPA and the broader research community, represent diverse real-world system activities and attack scenarios, lending significant credibility to the evaluation results.

The evaluation primarily focused on two key aspects: attack detection performance and operational efficiency. For attack detection, FLASH's performance was benchmarked against existing state-of-the-art systems. Against ThreatRace, a node-level intrusion detection system, FLASH demonstrated superior performance, achieving higher precision, lower false alarms, and a better F-score. This indicates FLASH's ability to accurately identify malicious nodes while minimizing the burden of spurious alerts on security analysts. Similarly, when compared to Unicon, a graph-level detection system, FLASH again exhibited superior detection performance across metrics such as precision, recall, and F-score, highlighting its ability to provide more granular and accurate insights into attack activities.

Beyond detection accuracy, a critical aspect of FLASH's proof of concept was its operational efficiency. Experiments revealed that the synergistic use of its lightweight classifier and the novel embedding recycling technique enabled FLASH to achieve up to three times speed improvements compared to ThreatRace. Furthermore, the system was shown to utilize minimal runtime system resources for threat detection, making it highly practical for deployment in real-world, resource-constrained environments.

The speakers also detailed a range of other experiments conducted to thoroughly assess FLASH's capabilities. These included studies on the impact of key parameters and components, such as the effect of audit log batch size on detection and runtime performance, a comparative analysis of various downstream lightweight classifiers, the specific effect of temporal Word2Vec encodings on accuracy, and the overall efficacy of the GNN embeddings. Crucially, the system's robustness to adversarial mimicry attacks was also evaluated, demonstrating that FLASH is resilient to attempts by attackers to disguise malicious behavior as benign, a testament to its sophisticated modeling of system semantics and structure. These comprehensive evaluations collectively serve as a robust proof of concept for FLASH's claims of high accuracy, efficiency, and resilience.

Defensive Implications

▶ Watch: GNN embedding database for efficient graph representation learning (8:00)

FLASH offers profound implications for defenders grappling with the complexities of modern cyber threats, particularly sophisticated APTs and zero-day attacks. Its core capabilities provide actionable intelligence and improved operational security for organizations.

Firstly, FLASH's anomaly-based detection mechanism is a significant advantage. Unlike signature-based systems that are inherently reactive and fail against novel threats, FLASH's unsupervised learning approach allows it to identify deviations from normal system behavior without requiring prior knowledge of attack signatures. This makes it a crucial tool for detecting zero-day attacks and evolving APT tactics that might otherwise bypass traditional defenses. Defenders gain a proactive capability to identify previously unseen threats.

Secondly, the system's ability to generate fine-grained, node-level alerts directly addresses a major pain point for security analysts. Instead of receiving coarse-grained, ambiguous alerts that require extensive manual investigation to pinpoint the root cause, FLASH provides precise identification of the exact system entities (processes, files, network connections) involved in an attack. This significantly reduces mean time to detection (MTTD) and mean time to respond (MTTR), freeing up valuable analyst time and resources. For incident response teams, this granular detail is invaluable for efficient triage, containment, and eradication of threats.

Thirdly, FLASH's emphasis on operational efficiency and minimal resource overhead makes it highly practical for real-world deployment. The GNN embedding recycling technique and lightweight classifier ensure that the system can process large volumes of audit logs in near real-time without overwhelming system infrastructure. This efficiency is critical for continuous monitoring in high-volume enterprise environments, where traditional graph-based analysis can be computationally prohibitive. Defenders can deploy FLASH without incurring excessive operational costs or performance degradation on monitored systems.

Finally, the demonstrated resilience to adversarial mimicry attacks is a powerful defensive attribute. Sophisticated attackers often attempt to blend their malicious activities with legitimate system behavior to evade detection. FLASH's comprehensive modeling of both semantic attributes (via temporal Word2Vec) and structural context (via GNNs) makes it difficult for adversaries to effectively mimic benign patterns. This robustness provides defenders with a higher level of assurance that their detection system is not easily fooled by advanced evasion techniques.

In essence, FLASH empowers defenders with an intelligent, efficient, and precise intrusion detection system that moves beyond reactive, signature-based approaches. Organizations should consider integrating or developing similar provenance-graph-based solutions that leverage semantic and structural learning, coupled with efficient embedding management, to bolster their defenses against the most advanced and persistent cyber threats.

Key Takeaways

  • Zero-Day APT Detection: FLASH is an anomaly-based intrusion detection system capable of detecting sophisticated Advanced Persistent Threats (APTs) and zero-day attacks without requiring pre-existing attack signatures or expert knowledge.
  • High Accuracy via Semantic Context: It achieves superior detection accuracy by leveraging rich semantic attributes (e.g., file paths, process names) within provenance graphs, encoded using a temporal-aware Word2Vec model, which significantly reduces false alarms.
  • Operational Efficiency through Embedding Recycling: FLASH overcomes the computational overhead of Graph Neural Networks (GNNs) by introducing a novel GNN embedding recycling technique, enabling up to three times speed improvements and minimal resource utilization for real-time threat detection.
  • Fine-Grained Alerting: The system generates precise, node-level alerts, allowing security analysts to quickly pinpoint exact attack entities and facilitating more efficient incident response.
  • Robustness Against Evasion: FLASH has demonstrated resilience to adversarial mimicry attacks, indicating its ability to withstand sophisticated attempts by attackers to disguise malicious activity as benign.
  • Comprehensive Evaluation: Validated on diverse DARPA and research community datasets (DARPA E3, DARPA OpTC, StreamSpot, Unicorn), outperforming existing node-level (ThreatRace) and graph-level (Unicon) systems in accuracy and efficiency metrics.

About the Speaker(s)

The research behind FLASH was presented by Mati Ur Rehman, who is affiliated with the University of Virginia. He collaborated with Hadi Ahmadi from Coric, and his advisor, Wajih Ul Hassan, also from the University of Virginia. The team's collective expertise in cybersecurity research, particularly in areas related to intrusion detection and graph representation learning, underpins the comprehensive and innovative approach presented in FLASH.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

FLASH presents a technically sound and impactful approach to provenance-based intrusion detection, effectively addressing long-standing issues of efficiency and false positives. The novel GNN embedding recycling technique, combined with temporal-aware semantic featurization, delivers a robust solution for detecting APTs and zero-days with fine-grained alerts and minimal overhead. This is a pragmatic advancement for real-world defense.

Heather Calloway (CISO) — STRONG ACCEPT

The FLASH system presents a compelling, rigorously evaluated approach to detecting advanced threats and zero-days with high accuracy and operational efficiency. By providing fine-grained, node-level alerts and significantly reducing false positives, it directly addresses critical pain points for security leaders and incident response teams, offering a strategic capability that improves overall risk posture.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024