Detection at Scale: Abstracting Detection Intent

Gaurav Singh (Software Engineer · Google), Dario Amiri (Senior Staff Software Engineer · Google)

BSidesSF 2026 · Day 1 · AMC Theatre 13

Overview

In their BSides SF 2026 presentation, "Detection at Scale: Abstracting Detection Intent," Gaurav Singh and Dario Amiri from Google addressed the formidable challenge of building and maintaining effective threat detection systems within large, dynamic organizations. The core of their talk centered on the strategic utility of abstracting detection intent as a critical tool for scaling threat detection coverage. This approach aims to decouple the "what" of detection (the desired security logic) from the "how" (its underlying implementation), thereby enhancing maintainability, adaptability, and overall effectiveness.

Watch on YouTube

Key moments

  1. 0:00 Welcome and talk introduction: abstracting detection intent
  2. 2:00 Understanding detection: the last line of defense
  3. 2:30 Challenges of detection at scale: thousands of rules
  4. 3:30 Key constraints and trade-offs: latency, cost, maintainability
  5. 6:00 Abstraction for scaling detection: reducing long-term complexity
  6. 6:40 Expressing detection intent: separating intent from implementation
  7. 7:15 Introducing detection patterns: simple vs. complex processing

Detection at Scale: Abstracting Detection Intent

Speakers: Gaurav Singh, Manager, Detection Systems; Dario Amiri, Senior Staff Software Engineer, Google

Conference: BSides SF

YouTube: https://www.youtube.com/watch?v=xbh3LEqkru8

Overview

In their BSides SF 2026 presentation, "Detection at Scale: Abstracting Detection Intent," Gaurav Singh and Dario Amiri from Google addressed the formidable challenge of building and maintaining effective threat detection systems within large, dynamic organizations. The core of their talk centered on the strategic utility of abstracting detection intent as a critical tool for scaling threat detection coverage. This approach aims to decouple the "what" of detection (the desired security logic) from the "how" (its underlying implementation), thereby enhancing maintainability, adaptability, and overall effectiveness.

The speakers, both seasoned veterans from Google's detection systems teams, highlighted that at Google's scale, threat detection involves thousands of continuously evolving rules, managed by large, specialized, and globally distributed teams. This environment necessitates a robust framework that can manage complexity, optimize costs (both human and machine), and ensure the longevity of detection logic. Their proposed solution leverages a structured approach to define common detection patterns, allowing security engineers to focus on the threat landscape while software engineers optimize the execution infrastructure.

The significance of this talk lies in its practical architectural insights for organizations grappling with the increasing volume and sophistication of cyber threats. By advocating for a clear separation of concerns, Google's approach offers a blueprint for building resilient and scalable detection capabilities that can adapt to evolving threats and infrastructure changes without overburdening security teams with low-level implementation details. This paradigm shift from bespoke, ad-hoc rule creation to a pattern-driven, abstracted methodology is presented as essential for sustainable growth in threat detection coverage.

Background

▶ Watch: Welcome and talk introduction: abstracting detection intent (0:00)

Threat detection serves as the critical last line of defense against cyberattacks that bypass preventative measures. This includes identifying malware infections, unauthorized system access, and anomalous activities aggregated across diverse systems and infrastructure. However, operating detection at scale presents significant challenges. Organizations like Google contend with thousands of detection rules that are constantly being created, updated, and triaged. This process involves large, distributed teams with specialized roles—ranging from platform developers to rule authors, alert triagers, and investigators. The inherent complexity demands that knowledge is codified and universally understood, rather than residing solely with individual experts.

Key constraints and trade-offs define the landscape of detection at scale. A fundamental tension exists between latency (how quickly a threat is detected), the rate of false positives, and the completeness of data. For instance, identifying a "rare" process immediately might generate numerous false positives if the system hasn't had sufficient time (e.g., minutes or hours) to determine if the event is truly anomalous or merely a routine software update. Cost is another major factor, encompassing both human time—the finite capacity of security engineers to develop, triage, and investigate—and machine resources required for data storage and processing. An organization's security investment is naturally constrained by its overall value proposition.

Finally, the maintainability and adaptability of detection logic are paramount. Given that detections often rely on heuristics, their effectiveness can degrade over time; what was once a clear indicator of malicious activity might become common, legitimate practice. Without an adaptable framework, modifying or updating rules can be as arduous as creating them initially. The traditional approach of developing bespoke implementations for each detection rule, while quick to start, leads to escalating complexity and technical debt as the number of rules grows. The speakers argue that an upfront investment in abstraction, though initially more complex to build, ultimately reduces long-term complexity and enables an organization to scale its detection capabilities more effectively. This forms the foundational problem that the abstraction of detection intent seeks to address.

Key Findings

▶ Watch: Challenges of detection at scale: thousands of rules (2:30)

The central finding of the talk is that abstracting detection intent—separating what a detection aims to achieve from how it is technically implemented—is a powerful and necessary strategy for scaling threat detection coverage in large organizations. This abstraction allows for the identification and reuse of common detection patterns, which are categorized into simple and complex types, thereby simplifying the development, maintenance, and optimization of detection logic.

Specifically, the key findings include:

  1. Reduced Complexity and Improved Maintainability: By expressing detection intent through a high-level, intuitive API (such as a Go API or custom DSL) and representing it in a structured, framework-agnostic format (like a Protobuff), security engineers can focus solely on the business logic of threat detection. This shields them from the intricate, ever-changing details of underlying infrastructure, processing frameworks (e.g., stream or batch processing), and system optimizations.
  2. Enhanced Adaptability and Sustainability: The separation of intent from implementation allows developers to independently optimize the execution systems without impacting the detection logic itself. This is crucial for adapting to evolving infrastructure, shifting business priorities, and balancing trade-offs between cost and latency. It ensures that detection logic remains effective and relevant even as the technological stack evolves, making threat coverage growth more sustainable in the long term.
  3. Standardization through Reusable Patterns: The identification and codification of a minimal yet comprehensive set of detection patterns (e.g., predicates, enrichments, correlations, thresholds, baselines, clustering) provide a common language and framework for expressing diverse security requirements. This standardization reduces cognitive load, improves readability, and facilitates collaboration across large, distributed security teams.
  4. Optimized System Performance and Correctness: Higher-level abstractions can encapsulate complex system-level challenges, such as handling late arriving data, reconciling event time and processing time, and managing non-monotonic versus monotonic aggregations. By centralizing these implementations within the platform, consistency, correctness, and optimization can be achieved across all detections, preventing common pitfalls like false negatives due to logic bugs.
  5. Strategic Use of "Escape Hatches": While advocating for higher-level abstractions, the speakers acknowledge the need for flexibility. User-defined functions (UDFs), for simple, stateless logic, serve as controlled "escape hatches." This balances the benefits of abstraction (simplicity, safety) with the need for users to address unique edge cases without building complex, unnecessary abstractions for every minor requirement.

In essence, the talk posits that while building such an abstracted system is more complicated upfront, its long-term benefits in terms of scalability, maintainability, and operational efficiency far outweigh the initial investment, making it a critical enabler for robust threat detection at Google's scale.

Technical Deep Dive

▶ Watch: Key constraints and trade-offs: latency, cost, maintainability (3:30)

The core technical contribution of the talk lies in its detailed breakdown of how detection intent is captured, represented, and executed within Google's scalable detection systems. This process is divided into three main stages: capturing intent via a user experience, representing that intent in a structured format, and executing it on an optimized platform.

Capturing and Representing Detection Intent

The first two stages focus on defining detection patterns—high-level constructs that express the "what" of a detection, divorced from implementation specifics. These patterns are identified by analyzing existing detections to determine the minimal, necessary set of functionalities. The goal is readability, maintainability, and adaptability, ensuring security engineers can express their logic without excessive complexity.

The process involves:

  1. User Experience (UX): Users express detection intent, primarily through a Go API. The speakers note that a custom DSL (Domain Specific Language) or other methods could also serve this purpose.
  2. Structured Representation: The expressed intent is then translated into a structured format, such as a Protobuff. This step is crucial for enforcing the abstraction, preventing shortcuts, and ensuring the intent is clearly separated from its execution.

The detection patterns are categorized into simple and complex types:

Simple Detection Patterns (Stateless Processing)

These patterns generally involve processing individual events without maintaining long-term state.

  • Predicates: Used to filter or select events based on specific criteria. This can involve scoping a rule (e.g., OS == Windows) or implementing security logic (e.g., user_matches_device_user). They are fundamental for making decisions about which events are worth investigating.
  • Example (pseudo-Go): executions.Filter(e => e.OS == "Windows" && e.User == e.DeviceUser)
  • Enrichments: Add context to an event using information that is relatively static or slow-changing. For instance, looking up an execution hash in VirusTotal to determine if it's known malware. Enrichments are suitable for data that doesn't vary significantly over short timeframes.
  • Example (pseudo-Go): executions.Enrich(e => VirusTotal.Lookup(e.Hash))
  • User-Defined Functions (UDFs): Serve as an "escape hatch" for custom logic that isn't easily covered by other patterns. These are restricted to stateless operations, taking an input and producing an output without side effects. While not strictly necessary in a perfectly abstracted system, UDFs provide flexibility where building a full abstraction for a niche case might not be cost-effective.
  • Example (pseudo-Go): executions.Map(e => myCustomFunction(e))

Complex Detection Patterns (Stateful Processing)

These patterns require processing across multiple events or maintaining state over time.

  • Correlation: Identifies relationships between two or more events happening within a defined temporal proximity. This can include events occurring near each other, in a specific sequence, or the absence of an expected event. Unlike enrichments, correlations deal with relationships that are dynamic and time-sensitive.
  • Example (pseudo-Go): fileDownloads.Correlate(executions, (f, e) => f.Path == e.Path && e.Timestamp - f.Timestamp < 1.Hour) (Detects file execution within an hour of download).
  • Thresholds: Detects when the volume of a specific event type exceeds a predefined limit within a given timeframe. This is common for identifying brute-force attempts or excessive activity.
  • Example (pseudo-Go): logins.GroupBy(l => l.User).Window(1.Hour).Count() > 10 (Detects more than 10 login attempts per user within an hour).
  • Baselines: Identifies abnormal behavior by comparing current activity against a computed historical norm. This involves establishing a baseline (e.g., daily execution frequency for a hash) and then flagging deviations (e.g., an execution that occurs less than 10 times per day, indicating rarity).
  • Example (pseudo-Go): executions.ComputeBaseline(e => e.Hash, dailyCount).Then(e => e.Count < baseline[e.Hash] * 0.1) (Detects rare executions).
  • Clustering: Groups similar alerts to manage human investigation time. If a host experiences one or a thousand malware executions, the investigative action might be the same. Clustering ensures that security teams are alerted once for a coherent incident, rather than being flooded with redundant alerts.
  • Example (conceptual): Grouping all malware alerts for a specific host within a day into a single investigative ticket.

Execution of Detection Intent

The final stage involves translating the structured intent into executable logic across Google's diverse infrastructure. This is where the separation of intent and implementation becomes critical.

  1. Framework-Agnostic Format: The Go API/UDFs are translated into a framework-agnostic format, allowing for flexible execution.
  2. Orchestration Logic: A sophisticated orchestrator analyzes the detection intent and determines the optimal way to execute it. A single detection can map to multiple rules, each potentially mapping to one or more processing stages. The orchestrator assigns these stages to appropriate underlying systems.
  • Example Execution Flow:
  • Intent: Read logs, filter uninteresting events, correlate with other data, look for statistical patterns.
  • Implementation (one possible way):
  • Logs from PubSub are filtered and warehoused using stream processing.
  • Batch processing is used for complex correlations.
  • A query engine determines statistical rarity over indexed fields.

This modularity allows for:

  • Reduced Cognitive Load: Users are shielded from low-level details like event time vs. processing time, handling late arriving data, reconciling latency characteristics across log sources, and deduplicating data. The system handles these complexities, allowing users to express logic solely in event time.
  • Optimized Aggregation and Joins: The system manages the nuances of non-monotonic aggregation (e.g., precise login counts, requiring careful balancing of correctness and latency) versus monotonic aggregation (e.g., threshold breaches, which can be reported early). This ensures both correctness and acceptable latency without user intervention.
  • Centralized Optimization: Critical functionalities like late data handling are implemented once, consistently, and optimally across all detections. This prevents bespoke, error-prone implementations and facilitates easier validation and troubleshooting, reducing the risk of false negatives.

The speakers emphasize that while lower-level abstractions offer more flexibility, they also increase complexity and maintenance burden. Higher-level abstractions, despite introducing some constraints, lead to simpler, more concise, and maintainable detection logic, which is vital for sustainably growing threat coverage as threats constantly evolve. The key is to be opinionated about difficult-to-get-right aspects (like efficient data processing and pattern matching) and flexible (via UDFs) for simpler, stateless logic.

Demo / Proof of Concept

▶ Watch: Expressing detection intent: separating intent from implementation (6:40)

The talk "Detection at Scale: Abstracting Detection Intent" primarily focused on presenting an architectural framework and conceptual models for scalable threat detection. While the speakers provided numerous illustrative examples of how detection intent could be expressed using pseudo-Go code and conceptual diagrams (such as filtering Windows OS events, enriching with VirusTotal data, correlating file downloads with executions, or detecting excessive login attempts), a live demonstration or a specific, publicly available proof-of-concept tool was not presented. The examples served to clarify the principles of abstracting detection patterns and separating intent from implementation, rather than showcasing a working demo of the described system.

Defensive Implications

▶ Watch: Introducing detection patterns: simple vs. complex processing (7:15)

The architectural principles outlined in this talk have profound implications for security defenders seeking to build and mature their threat detection capabilities, particularly in large and complex environments.

  1. Prioritize Abstraction in Detection Engineering: Defenders should actively work towards abstracting their detection logic. Instead of writing bespoke rules for every threat, identify common detection patterns (e.g., filtering, enrichment, correlation, thresholding, baselining) and build a framework that allows security engineers to express these patterns at a high level. This shifts the focus from low-level system details to the actual threat logic.
  2. Separate Intent from Implementation: A critical takeaway is the deliberate separation of what a detection does (the detection intent) from how it's executed on the underlying infrastructure (the detection implementation). This allows security engineers to focus on threat intelligence and logic, while platform engineers optimize the performance, cost, and scalability of the detection engine. This separation future-proofs detection logic against inevitable changes in infrastructure and processing frameworks.
  3. Invest in a Pattern-Driven Detection Platform: Organizations should consider developing or adopting a detection platform that natively supports these abstracted patterns. This involves providing a user-friendly interface (like a Go API or DSL) for expressing intent and a robust backend capable of translating this intent into efficient execution across various data processing systems (e.g., stream processing, batch processing, query engines).
  4. Reduce Cognitive Load for Security Engineers: By abstracting away complexities like event time vs. processing time, late arriving data, and intricate aggregation logic, security engineers can be more productive. This allows them to focus on understanding evolving threats and translating that understanding into effective detection logic, rather than battling with system-level nuances.
  5. Enhance Maintainability and Adaptability: A pattern-based, abstracted system inherently leads to more maintainable and adaptable detection rules. When threats evolve or legitimate practices change, modifications can often be made at the high-level intent layer, reducing the effort required to update and tune detections. This is crucial for sustainably growing threat coverage over time.
  6. Strategic Use of "Escape Hatches": While promoting high-level abstractions, defenders should also design their systems with controlled "escape hatches" like User-Defined Functions (UDFs) for simple, stateless logic. This provides flexibility for unique or rapidly evolving detection needs without undermining the overall abstraction framework. The key is to use them judiciously for cases where building a full abstraction is not warranted.
  7. Optimize for Cost and Performance Centrally: By centralizing the execution of detection patterns, platform teams can apply system-wide optimizations for cost (human and machine) and performance (latency, false positive rates). This ensures that best practices for data processing, such as efficient joins and aggregations, are consistently applied across all detections.
  8. Improve Collaboration and Knowledge Sharing: A standardized set of detection patterns and a clear expression of intent facilitate better collaboration among security engineers, software engineers, and incident responders. Everyone can more easily understand, modify, and troubleshoot detection logic, even across geographically distributed teams.

Ultimately, adopting an abstraction-driven approach helps defenders move beyond an ad-hoc, reactive stance to building a proactive, scalable, and resilient threat detection ecosystem capable of keeping pace with the dynamic cyber threat landscape.

Key Takeaways

  • Abstraction is Key to Scale: Abstracting detection intent from its implementation is fundamental for scaling threat detection coverage in large, complex organizations, reducing long-term complexity and technical debt.
  • Pattern-Driven Approach: Detections can be broken down into reusable simple (predicates, enrichments, UDFs) and complex (correlation, thresholds, baselines, clustering) patterns, standardizing detection logic.
  • Separation of Concerns: Decoupling the "what" (detection intent, e.g., Go API) from the "how" (execution on infrastructure like PubSub, stream/batch processing, query engines) allows security engineers to focus on threats and platform engineers to optimize systems independently.
  • Reduced Cognitive Load: Higher-level abstractions shield security teams from intricate system details like event time, processing time, late arriving data, and complex aggregations, enabling them to be more productive.
  • Enhanced Maintainability: A pattern-based, abstracted system makes detection logic easier to read, understand, and adapt to evolving threats and infrastructure, ensuring sustainable growth in threat coverage.
  • Strategic Flexibility: While advocating for abstraction, the system incorporates "escape hatches" like User-Defined Functions (UDFs) for simple, stateless logic, balancing constraint with necessary flexibility.

About the Speaker(s)

Dario Amiri is a Senior Staff Software Engineer at Google. He has worked on the team responsible for building detection systems at Google for seven years, initially as a tech lead and later as a manager. His expertise lies in designing and scaling complex systems for threat detection.

Gaurav Singh works on the detection systems teams at Google, collaborating closely with security engineers and software engineers to detect and respond to information security threats. He has been with the team for almost a decade, starting as an engineer and now serving as a manager. His experience spans the evolution of Google's detection capabilities.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Competent engineering talk from people who clearly built the thing they're describing — Google's detection abstraction layer is real infrastructure and the pattern taxonomy (predicates, enrichments, correlations, baselines, clustering) is sensible. Nothing here will surprise anyone who's read the SIEM/detection-engineering literature or spent time with frameworks like Panther, Chronicle's YARA-L, or Elastic's rule DSLs, but it's delivered with authority and the execution-layer nuances (event time vs. processing time, monotonic vs. non-monotonic aggregation, late-arriving data) show genuine depth. Fills a BSides slot fine; wouldn't make the main stage at DEF CON.

Heather Calloway (CISO) — SOLID

Singh and Amiri present a well-reasoned engineering architecture for scaling detection at Google. The framework is coherent and the separation-of-concerns logic is sound, but this is fundamentally a platform engineering talk — it never addresses governance, risk ownership, or what any of this means for an organization that isn't Google.

→ Top-rated talks at BSidesSF 2026

All talks from BSidesSF 2026