Hacking Context for Auto Root Cause and Attack Flow Discovery
Ezz Tahoun
DEF CON 33 · Day 1 · Main Stage
Overview
In this compelling DEF CON talk, Ezz Tahoun presents a radical rethinking of how cybersecurity organizations approach log management, correlation, and threat detection. Titled "Hacking Context for Auto Root Cause and Attack Flow Discovery," the presentation directly addresses the pervasive challenges of false positives, overwhelming log volume, and the inherent limitations of traditional, rules-based correlation in Security Information and Event Management (SIEM) systems. Tahoun argues that current approaches lead to analyst burnout, exorbitant costs, and ultimately, broken security.

Key moments
- 1:35 Discussing blue team challenges: false positives and correlation
- 4:18 Level 1 correlation: grouping similar events from one session
- 6:06 Level 2 correlation: assembling the complete attack story
- 6:53 Summary: two distinct levels of correlation explained
- 8:12 The challenge: why writing correlation rules is so hard
- 9:11 Demonstrating overwhelming log data duplication in practice
Hacking Context for Auto Root Cause and Attack Flow Discovery
Speakers: Ezz Tahoun
Conference: DEF CON
YouTube: https://www.youtube.com/watch?v=k2r3JrodFaI
Overview
In this compelling DEF CON talk, Ezz Tahoun presents a radical rethinking of how cybersecurity organizations approach log management, correlation, and threat detection. Titled "Hacking Context for Auto Root Cause and Attack Flow Discovery," the presentation directly addresses the pervasive challenges of false positives, overwhelming log volume, and the inherent limitations of traditional, rules-based correlation in Security Information and Event Management (SIEM) systems. Tahoun argues that current approaches lead to analyst burnout, exorbitant costs, and ultimately, broken security.
The core of Tahoun's proposal is a two-tiered, machine learning-driven methodology that redefines "correlation." Instead of relying on rigid, narrow rules, his system first groups relevant events based on comprehensive context and then chains these groups into probabilistic attack stories. This approach not only promises a drastic reduction in data volume through intelligent deduplication but also aims to deliver a clearer, more actionable understanding of potential threats, moving away from isolated alerts to coherent narratives of attack progression.
This talk is particularly pertinent for blue teamers, security architects, and anyone grappling with the operational and financial burdens of managing vast quantities of security logs. Tahoun's solution, partially available as open-source code, offers a tangible path to optimizing SIEM investments, enhancing detection capabilities, and significantly improving the efficiency and morale of security operations centers (SOCs) by transforming raw, noisy data into meaningful security intelligence.
Background
▶ Watch: Discussing blue team challenges: false positives and correlation (1:35)
The cybersecurity industry has long struggled with fundamental issues in log management and threat detection. A primary concern for blue teamers, as highlighted by audience members in the talk, is the overwhelming volume of false positives. Traditional SIEMs and detection systems often generate countless alerts that, when investigated, turn out to be benign activity, leading to alert fatigue and diverting valuable analyst time.
The root of this problem often lies in the reliance on correlation rules. These rules, designed to link disparate events, are inherently narrow and pattern-specific. As Tahoun points out, writing effective correlation rules is challenging because they struggle to account for the vast permutations of an attack and the nuanced context of an organization's environment. A rule might link 10 outbound SQL connections within 5 minutes on the same machine, but attackers often bypass firewalls using common ports, leading to many benign connections that trigger the rule. This results in a deluge of duplicate and redundant logs from various sources – NetFlows, firewall logs, Sysmon, EDR telemetry, and Windows event logs – all describing the same underlying activity but appearing as separate events.
Furthermore, the economic model of many SIEM vendors exacerbates this issue. Organizations typically pay for log volume, not log value. This creates a perverse incentive for collecting as much data as possible, leading to terabytes of redundant information. Tahoun vividly illustrates this by noting that a single user browsing a website for two minutes can generate hundreds or even thousands of duplicate events across different log types. This data bloat not only inflates costs but also makes it exponentially harder for analysts to find genuine threats amidst the noise.
The speaker draws an analogy to consumer technology, where automatic grouping and personalization are standard. Platforms like MySpace, YouTube, and Gmail have long used sophisticated systems to recommend relevant content and categorize emails without explicit rules. Yet, the cybersecurity industry has lagged, often treating each log event as an isolated entity, necessitating manual investigation or overly specific searches. This fundamental disconnect between how other industries handle "big data" and how cybersecurity handles its own massive datasets underscores the urgent need for a more intelligent, context-aware approach to log correlation and threat discovery.
Key Findings
▶ Watch: Level 2 correlation: assembling the complete attack story (6:06)
Ezz Tahoun's talk introduces several key findings that challenge conventional wisdom in cybersecurity log management and correlation:
- Redefining Correlation into Two Distinct Stages: Tahoun argues that the term "correlation" is overloaded and should be split into two fundamental processes:
- Grouping (Similarity): The initial stage involves consolidating all relevant events that describe the same underlying activity, irrespective of their exact temporal or field-level matches. This means bringing together all NetFlows, firewall logs, Sysmon events, and EDR telemetry that pertain to a single user browsing a website, for instance. This process is driven by understanding comprehensive context rather than rigid rules.
- Chaining (Causal Stories): Once events are grouped, the second stage focuses on assembling these groups into attack stories or attack flows. This involves identifying sequential and causal relationships between different grouped activities that align with known attack frameworks like the MITRE ATT&CK kill chain. The goal is to move beyond isolated alerts to a coherent narrative of an attacker's progression.
- The Critical Role of Deduplication: A significant side benefit and core finding of the grouping process is deduplication. By intelligently grouping similar events, Tahoun demonstrates that up to 99% of log volume can be redundant. This massive reduction in data volume directly translates into substantial cost savings for SIEM ingestion and storage, which are typically priced per gigabyte or terabyte. He states that filtering out unwanted events might save only 5-10% of data, whereas deduplication can achieve 90% or more.
- Context is King, Rules are Narrow: A central tenet of the talk is that context is paramount for effective correlation. Traditional rules-based systems fail because they are "very, very narrow" and cannot account for the myriad ways events might relate without matching specific fields (e.g., same IP, same port, same 5-minute window). Tahoun emphasizes the need to consider all available fields and business context to understand true relevance, a task ill-suited for static rules.
- Machine Learning for Generic, Real-time Discovery: Unlike rules, which are akin to "searches" for specific patterns (e.g., IOCs), Tahoun advocates for using machine learning for generic, real-time grouping and story building. He likens this to the personalized feeds on social media platforms or email categorization in Gmail, where algorithms automatically present relevant information without the user explicitly defining rules. This shift allows for proactive discovery of attack flows rather than reactive searching.
- Explainable AI is Achievable and Preferable: While advanced AI models like large language models (LLMs) are often opaque, Tahoun emphasizes the use of small language models (LMs), clustering models, and Markovian models for his system. These "white box" or "expert AI" approaches are explainable, interpretable, and auditable, allowing security professionals to understand why certain events are grouped or chained together and to fine-tune the logic with subject matter expertise. This contrasts sharply with black-box anomaly detection systems that often generate unexplainable false positives.
Technical Deep Dive
▶ Watch: Summary: two distinct levels of correlation explained (6:53)
Ezz Tahoun's proposed solution is built upon a sophisticated, yet explainable, machine learning architecture designed to tackle the inherent complexities of cybersecurity log data. The methodology unfolds in several key technical stages:
- Enrichment with Small Language Models (LMs):
The foundational step is to enrich every incoming log, event, or alert with relevant metadata. Unlike generic Large Language Models (LLMs), Tahoun's system utilizes small language models (LMs) specifically "trained on cyber security" concepts. These LMs don't "speak English" but rather interpret log strings to extract and classify information, primarily mapping events to MITRE ATT&CK techniques. This process establishes a standardized "language" across disparate log sources, allowing for consistent grouping based on the type of activity (e.g., initial access, execution, persistence). The speaker notes that developing a consistent, labeled dataset for this LM classification was a multi-year effort involving human experts to ensure accuracy and consistency, mitigating the hallucination tendencies of some language models.
- Grouping via Clustering Models:
Once enriched, events are subjected to a grouping mechanism that identifies similarities across all fields and contextual information. This is achieved using classical clustering models such as K-means or DB-scan. In data science terms, each event is represented as a point in a multi-dimensional space, where each dimension corresponds to a field or an enriched attribute. Events that are "closer" to each other in this dimension space are grouped together. This approach is superior to rules because it considers the entirety of the context, rather than relying on predefined patterns for specific fields (e.g., same IP and same port and same time). The clustering algorithm dynamically determines relevance, even if events are temporally disparate (e.g., "five hours, five days apart") or involve different IPs but the same user. This grouping inherently facilitates deduplication by consolidating redundant traces of the same activity.
- Chaining Attack Stories with Markovian Models:
After events are grouped, the next stage involves assembling these groups into coherent attack stories or attack flows. This is where the concept of a kill chain (or attack life cycle progression) comes into play. Tahoun employs Markovian models (a type of sequencing model that has existed since the 1970s) to identify causal relationships and sequences between the grouped activities. The model looks for probabilistic transitions between different MITRE ATT&CK tactics (e.g., initial access followed by execution, then exfiltration). It's not about rigid rules like "tactic 1 then tactic 5 then tactic 10," but rather identifying the most likely progression of tactics within a given timeline. This allows for the discovery of attack chains even when there are "detection gaps" (missing traces for certain tactics).
- Incorporating Business and Entity Context:
To further refine groupings and chaining, the system incorporates rich business context and entity enrichment. This is achieved through lookup tables (similar to Splunk's asset and identity management tables). These tables map entities (users, machines, IP addresses, MAC addresses) to organizational attributes like business unit, department, location, and manager. The machine learning model can infer relationships between entities (e.g., "this MAC address and this IP are related to this user") if they frequently appear together in grouped events. This entity-level context then influences the clustering process, meaning events involving entities from the same department or office will be considered "closer" than those from different departments, even if other technical fields differ. This deep contextual awareness allows the system to differentiate between benign cross-departmental activity and suspicious lateral movement.
- Semi-supervised and Knowledge-Based Learning:
A crucial aspect of Tahoun's architecture is its reliance on semi-supervised or weakly supervised learning, rather than purely supervised or unsupervised methods. He explicitly rejects training models solely on an organization's "normal" data, as "normal activity is not something that I can even assume is going to be the same for tomorrow," and such approaches often lead to high false positive rates in anomaly detection. Instead, the models are "trained on knowledge" and "concepts" – leveraging subject matter expertise about what constitutes a meaningful relationship in cybersecurity (e.g., HTTP and HTTPS ports 80 and 443 are functionally closer than 80 and 81). This "expert AI" approach grounds the models in real-world security understanding, making them more robust, explainable, and less prone to misinterpreting benign anomalies as threats.
This comprehensive technical framework, integrating advanced yet transparent machine learning techniques, aims to transform raw log data into actionable, context-rich attack narratives, providing both significant cost savings and enhanced security posture.
Demo / Proof of Concept
▶ Watch: The challenge: why writing correlation rules is so hard (8:12)
While the talk itself did not feature a live, interactive demo of the software, Ezz Tahoun effectively illustrated the principles and outcomes of his approach using several visual aids and conceptual walkthroughs:
- Visualizing Log Duplication: Tahoun opened a Notepad++ instance displaying raw FortiGate firewall logs and NetFlow logs. He scrolled through them, pointing out the overwhelming redundancy: "the same IP a million times," "500 of these events describing the same IP, the same port, the same IP." This stark visual demonstration underscored the massive data bloat that his system aims to address.
- Attack Storyboard Example (Singapore Health Attack): A recurring visual throughout the talk was a detailed attack storyboard illustrating the "Singapore Health attack." This diagram, likely a post-incident analysis report, showed attack progression from left to right (kill chain tactics) and top to bottom (timeline). It depicted various entities (represented by different colors) and linked events across different stages of the attack, such as initial access, execution, and exfiltration. This served as the aspirational goal for his automated chaining process – to construct such narratives before a breach is fully realized.
- Conceptual Grouping and Chaining: Tahoun used abstract diagrams of "green circles" (e.g., representing Tactic 3: Persistence) and "orange circles" (e.g., Tactic 2: Execution) to illustrate how his clustering algorithms would group similar events and how sequencing models would then chain these groups into a "story" if they represented sequential tactics (e.g., "execution and persistence are close to each other... that's a story").
- Log Volume Reduction Statistics and Cost Savings: A significant "proof of concept" was presented through concrete figures on log volume reduction. Tahoun showed a table indicating a reduction from over 1,000 kilobytes to 75 kilobytes, representing a 93% volume reduction. He then translated this into dollar amounts, using Splunk as an example. For an organization ingesting 4 terabytes a day, the annual cost could be between $3.8 million and $4.8 million. By achieving even a conservative 80-90% reduction (e.g., to 0.25 terabytes/day), the cost could drop to $0.3 million, saving "2 million dollars" or more. This powerfully demonstrated the financial impact of his deduplication strategy.
- GitHub and Product Mentions: The speaker explicitly mentioned that the core algorithms are available on his GitHub ("Go to my GitHub and grab it from there"). He also referred to a commercial product,
app.sipenta.io, for a product tour and UI, implying a more polished, integrated offering based on the open-source engine. This suggests that the technology is not merely theoretical but has practical, deployable implementations.
The presentation's "demo" elements, though not a live software walkthrough, were highly effective in conveying the fundamental problems, the proposed solutions, and the tangible benefits in terms of both operational efficiency and cost savings.
Defensive Implications
▶ Watch: Demonstrating overwhelming log data duplication in practice (9:11)
Ezz Tahoun's proposed methodology offers several profound implications for cybersecurity defenders, fundamentally altering how SOCs operate and how organizations manage their security posture:
- Massive Cost Reduction for SIEMs: The most immediate and quantifiable benefit is the drastic reduction in log ingestion costs. By intelligently grouping and deduplicating data before it reaches the SIEM, organizations can cut their log volume by 80-99%. This translates to millions of dollars in savings annually, especially for platforms like Splunk, Elastic, or Azure Sentinel, which charge based on data volume. Tahoun suggests integrating this logic into data pipelines using tools like Logstash or Cribl, ensuring that only high-value, consolidated data is ingested.
- Significant Reduction in False Positives and Alert Fatigue: By moving from isolated, noisy events to grouped sessions and probabilistic attack stories, analysts will encounter far fewer "false positives." Instead of a million individual alerts, they might see one "group of false positives" (benign activity consolidated) or, more importantly, a cohesive narrative of a potential attack. This directly addresses analyst burnout and allows security teams to focus on genuine threats rather than sifting through noise.
- Proactive and Context-Rich Threat Detection: The system enables the proactive discovery of attack flows and root causes by automatically chaining sequential tactics across the kill chain. Even if individual detections are imperfect, the aggregation into a story provides a higher-fidelity signal. Incorporating business context and entity enrichment means that alerts are more relevant, understanding not just what happened, but who (user, department) and what (critical asset) was involved.
- Improved Analyst Efficiency and Workflow: Security analysts shift from "searching for something in particular" (like IOCs after an APT report) to reviewing an automatically generated "feed" of potential attack stories. This is akin to a personalized social media feed for threats. Analysts can then investigate these pre-correlated stories, adding their human expertise and threat intelligence, rather than starting from raw, fragmented logs. This streamlines incident response and threat hunting efforts.
- Explainable and Auditable AI for Security: The use of "white box" machine learning models (clustering, Markovian models, small LMs) ensures that the system's decisions are explainable, interpretable, and auditable. Defenders can understand why events were grouped or chained, allowing them to troubleshoot, fine-tune the models with their specific environment knowledge, and maintain compliance, avoiding the "black box" issues common with some AI solutions.
- Enhanced Compliance and Data Retention: Instead of haphazardly filtering out events to save costs (potentially compromising compliance), defenders can retain all raw data (e.g., in an S3 bucket) but only ingest the deduplicated, high-value summaries into the SIEM. This simplifies architecture and ensures that compliance needs for raw log retention are met without incurring excessive SIEM costs.
In essence, Tahoun's framework empowers defenders to move beyond the current reactive, volume-driven paradigm to a more intelligent, cost-effective, and proactive security posture, ultimately making security operations more effective and sustainable.
Key Takeaways
- Redefine Correlation: The traditional concept of correlation is broken. It should be split into two stages: grouping similar events based on comprehensive context, and then chaining these groups into causal attack stories that align with attack life cycles like the MITRE ATT&CK kill chain.
- Deduplication is a Game Changer: Intelligent grouping inherently leads to massive deduplication of logs (80-99% volume reduction), which translates directly into millions of dollars saved annually on SIEM ingestion and storage costs. This process should occur before data enters the SIEM.
- Context-Aware Machine Learning Outperforms Rules: Generic, narrow correlation rules fail to capture the full context of security events. Instead, small language models, clustering models (K-means, DB-scan), and Markovian models can look at all fields and entity context to find relevance and build attack sequences, similar to how recommender systems work in consumer tech.
- Focus on Explainable, Knowledge-Based AI: Opt for "white box" or "expert AI" models that are explainable, interpretable, and auditable. These models should be "trained on knowledge" and cybersecurity concepts, not just raw data, to reduce false positives and avoid the unreliability of black-box anomaly detection.
- Transform SOC Operations: This approach shifts analysts from sifting through millions of isolated alerts and false positives to investigating prioritized, context-rich attack stories. This significantly improves analyst efficiency, reduces alert fatigue, and enables more proactive threat hunting and incident response.
- Open Source Foundations with Commercial Support: The core algorithms are available via GitHub, promoting transparency and community contribution, while commercial offerings provide enterprise-grade support and integrations.
About the Speaker(s)
The speaker for this talk, Ezz Tahoun, presented a detailed technical vision for transforming cybersecurity operations. Based on the transcript, Tahoun is the creator and proponent of the described methodology and its underlying technology, including an open-source GitHub repository for the algorithms and a commercial product (app.sipenta.io). He is deeply knowledgeable in machine learning, data science, and cybersecurity, advocating for a shift from traditional, rules-based security to a more intelligent, context-aware, and cost-effective approach. His expertise spans areas such as natural language processing, clustering, and sequencing models applied to security data.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Tahoun tackles a real and painful problem — SIEM bloat, alert fatigue, rules-based correlation failure — with a coherent two-stage ML architecture (group then chain). The framing is clean and the cost savings math is compelling, but the talk stays at the concept level and never quite gets its hands dirty enough to be memorable at DEF CON.
Heather Calloway (CISO) — SOLID
Tahoun addresses a real and costly operational problem — SIEM bloat, alert fatigue, rules-based detection failure — with a technically coherent two-stage ML framework. The work is credible and the cost reduction math is concrete, but this is a SOC architecture talk, not a security leadership talk, and it never crosses that line.