Detection Allegro: Composing Detection Rules with Agentic Workflows
Raphael Ruban (Vacasa), Chen Cao (Vacasa)
BSidesSF 2026 · Day 2 · AMC Theatre 13
Overview
In the rapidly evolving landscape of cybersecurity, the ability to quickly develop, deploy, and maintain robust threat detection rules is paramount. This talk, "Detection Allegro: Composing Detection Rules with Agentic Workflows," presented by Raphael Ruban and Chen Cao from Vacasa, delves into an innovative approach that leverages agentic workflows and artificial intelligence (AI) to dramatically streamline this process. The core proposition is to cut the time spent on threat detection development and maintenance by over 80%, transforming what is often a tedious and error-prone manual endeavor into an effortless, automated workflow.
Key moments
- 0:00 Introduction: Effortless threat detection with agentic workflows
- 1:40 Vacasa's in-house open-source detection pipeline overview
- 3:10 Raphael's painful manual detection rule creation experience
- 4:30 Realization: manual detection rule creation is not effortless
- 5:10 Demo: AI-powered generation of new detection ideas
- 6:20 Demo: AI generating a complete detection rule pull request
- 8:00 Explaining why AI is ideal for detection rule generation
Detection Allegro: Composing Detection Rules with Agentic Workflows
Speakers: Raphael Ruban, Detection Response Engineer, Vacasa; Chen Cao, Detection Response Engineer, Vacasa
Conference: BSides SF
YouTube: https://www.youtube.com/watch?v=nRJQEtOj_Vc
Overview
In the rapidly evolving landscape of cybersecurity, the ability to quickly develop, deploy, and maintain robust threat detection rules is paramount. This talk, "Detection Allegro: Composing Detection Rules with Agentic Workflows," presented by Raphael Ruban and Chen Cao from Vacasa, delves into an innovative approach that leverages agentic workflows and artificial intelligence (AI) to dramatically streamline this process. The core proposition is to cut the time spent on threat detection development and maintenance by over 80%, transforming what is often a tedious and error-prone manual endeavor into an effortless, automated workflow.
The speakers illustrate how Vacasa's detection response (DR) team has built an in-house platform that integrates various open-source tools with custom AI agents. This system not only generates new detection rules and their corresponding tests but also handles routine maintenance tasks like bug fixes and allow-listing. By offloading these repetitive yet critical tasks to AI, security engineers are freed to focus on more complex threat modeling and strategic initiatives, while simultaneously making security contributions accessible to a broader range of engineers within the organization.
This talk is crucial for any organization grappling with the challenges of scaling their detection capabilities, managing alert fatigue, and optimizing the efficiency of their security operations. It demonstrates a practical, production-ready implementation of AI in security, offering a blueprint for how agentic systems can enhance an organization's defensive posture by ensuring timely, accurate, and comprehensive threat coverage across diverse log sources.
Background
▶ Watch: Introduction: Effortless threat detection with agentic workflows (0:00)
Vacasa's detection response team operates with a lean but effective philosophy: building robust in-house capabilities rather than relying solely on expensive vendor solutions for log analysis and security automation. Their custom-built pipeline is a testament to the power of open-source software, integrating projects like Grove for log pulling and gathering, Substation for log normalization and enrichment, and Stream Alert for real-time log analysis. This is complemented by an in-house alerting pipeline and Tracecat for AI-driven response and alert enrichment. This sophisticated setup processes hundreds of terabytes of logs monthly and runs hundreds of detection rules, maintaining cost efficiency while ensuring broad coverage.
Despite the advanced infrastructure, the human element of detection rule development presented a significant bottleneck. Raphael Ruban, a new hire at Vacasa, vividly recounted his initial experience tasked with improving GitHub coverage. His goal: detect disabled branch protections, an activity that should ideally never happen. The seemingly straightforward task quickly devolved into a laborious, multi-step process:
- Documentation Review: Reading GitHub's official documentation to understand available data.
- Log Analysis: Verifying documentation against actual production logs.
- Code Development: Writing the detection logic, often involving SQL queries for scheduled detections.
- Testing: Developing and running test cases, often encountering initial failures.
- Debugging: Reviewing code, re-reading documentation, re-analyzing logs, and debugging in development.
- Production Validation: Discovering that even after passing local tests, the detection might fail in production, leading to further CloudWatch and log review.
This repetitive cycle, described as "super boring" by Raphael, highlighted a critical problem: while the initial idea for a detection requires human ingenuity, the subsequent implementation, testing, and debugging are largely mechanical. Engineers gain little professional development from these tasks, leading to potential burnout and inefficiency. Furthermore, routine maintenance—such as fixing minor bugs, updating rules due to schema changes, or adding users/repositories to an allow list to prevent alert fatigue—suffers from the same repetitive, unexciting nature.
The speakers argue that this scenario represents a "good problem for AI to solve." The task of generating detection code from an idea is clearly defined, its correctness can be verified (a known malicious log should trigger an alert), it's highly repetitive, and fundamentally, it's a text comprehension problem—an area where large language models (LLMs) excel. This realization paved the way for Vacasa's "Detection Allegro" system, aiming to automate the drudgery and free human security expertise for higher-value activities.
Key Findings
▶ Watch: Raphael's painful manual detection rule creation experience (3:10)
The "Detection Allegro" system at Vacasa has yielded several transformative findings, fundamentally reshaping how the security team operates:
- Dramatic Efficiency Gains: The most significant finding is the profound reduction in the time required for detection development and maintenance. What once took Raphael approximately two hours to develop a new detection rule now takes only about 20 minutes, including thorough review. This represents an 80%+ reduction in development time. The bug-fixing and allow-listing lambda is used weekly, demonstrating its immediate and sustained impact on operational efficiency.
- Scalable Detection Creation: The system has been instrumental in creating over 150 new detection rules since its inception. This volume of high-quality detections would be challenging, if not impossible, to achieve with traditional manual methods, especially for a lean team.
- AI's Role in Code Generation and Idea Generation: AI agents are not just assisting; they are actively generating functional detection code (Python), SQL queries, and comprehensive test cases. Beyond code, AI is also effectively leveraged for generating novel detection ideas, both from public knowledge bases (for common SaaS applications) and from internal, proprietary data sources (for internal tooling and new log sources).
- Two-Phase AI Application: The problem of detection management is effectively broken down into two distinct phases for AI application:
- Creation: Generating entirely new detection rules from an initial concept.
- Maintenance: Handling iterative improvements, bug fixes, and allow-listing for existing rules. This modular approach allows for specialized agent designs and workflows for each phase.
- Importance of Context-Rich Prompting: The success of the AI agents heavily relies on providing them with comprehensive and relevant context. This includes feeding the AI agents with up-to-date log samples, log schemas (from Glue), internal documentation (from Notion), and security frameworks like MITRE ATTACK. This "manual context management" ensures the AI has all necessary information to generate accurate and contextually appropriate detections.
- Specialized Agent Skills are Crucial: LLMs, while powerful, have limitations (e.g., handling specific formatting requirements like Athena's lowercase SQL). The development of specialized Athena query skills for the coding agent, which include linting, testing against a live Athena instance, and iterative refinement, was critical to overcoming these challenges and ensuring the accuracy of generated SQL queries.
- Orchestration via Ticketing Systems: Integrating the agentic workflows with a ticketing platform like Linear proved to be a pivotal decision. This not only centralizes all detection-related information but also significantly lowers the barrier to entry for other engineering teams to contribute security ideas and even trigger detection generation, fostering a broader security-aware culture across the organization.
These findings collectively highlight that AI, when strategically integrated into well-designed agentic workflows and provided with rich context, can revolutionize the efficiency and effectiveness of security detection engineering.
Technical Deep Dive
▶ Watch: Realization: manual detection rule creation is not effortless (4:30)
The architecture of Vacasa's "Detection Allegro" system is a sophisticated orchestration of AI agents, custom tooling, and existing infrastructure, designed to handle two primary workflows: detection creation and detection maintenance.
Detection Creation Workflow
The creation phase is initiated when a security engineer, or even an engineer from another team, provides an initial idea for a detection. This process is meticulously designed to provide the AI with all necessary context before tasking it with code generation.
- Prompt Building and Context Gathering:
- Log Source Wiki (Notion): Provides detailed information about all available log sources, their purpose, and the type of data they contain.
- Log Samples (S3): Live, up-to-date log samples are pulled from S3 buckets. This is crucial for the AI to understand the actual format and content of logs, adapting to any schema changes or variations (e.g., a field moving or changing capitalization).
- Log Schema (AWS Glue): The formal schema definitions from Glue provide a structured understanding of log fields and data types.
- Initial AI Call: An initial LLM call is made to the AI, asking it to interpret the detection idea in the context of the gathered information. This call helps the AI decide on the detection's name, severity, and how often it should run (e.g., every 30 minutes, hourly).
- MITRE ATTACK Framework Integration: A separate API call enriches the context with relevant MITRE ATTACK framework tactics and techniques. This ensures that detections are mapped to industry-standard adversary behaviors, improving their strategic value.
- Detection Type Decision: Based on the gathered context and the AI's initial assessment, a critical decision is made:
- Scheduled Detections: These run periodically (e.g., every 30 minutes, hourly) and analyze logs over a time window. They typically involve complex SQL queries against a data lake (like Athena).
- Streaming Detections: These operate in real-time, evaluating individual logs as they arrive. They are designed for immediate anomaly detection without historical context.
The system uses different code examples and prompt structures for each type to guide the AI appropriately.
- Coding Agent Execution:
- Large Prompt Construction: All the gathered context—log source information, samples, schema, MITRE details, and the detection type decision—is consolidated into a single, comprehensive prompt. This rich prompt is then handed off to a specialized coding agent.
- Lambda Environment: The coding agent operates within an AWS Lambda function. This environment provides it with access to:
- Repository Access: The agent can interact with Vacasa's detection rule repository, enabling it to read existing code, understand project structure, and ultimately commit new code.
- Tooling: All standard development tooling, such as linters (e.g., for Python and SQL) and test runners, are available to the agent. This allows the agent to self-correct and ensure code quality.
- Specialized Athena Skill: This is a crucial component. Recognizing that LLMs often struggle with specific formatting requirements (like Athena's strict lowercase column names) and complex SQL syntax, Vacasa developed a custom skill. This skill allows the agent to:
- Generate an SQL query based on the provided inputs (database, schema, log samples).
- Run the query through internal and public SQL linters.
- Execute the query against a live Athena instance to validate its correctness and performance.
- If the query fails (e.g., due to capitalization errors or syntax issues), the skill provides feedback to the agent, prompting it to iterate and refine the SQL until it passes. This iterative loop significantly enhances the accuracy of generated SQL.
- Pull Request (PR) Creation: Once the coding agent successfully generates the detection code (Python), associated tests, and any necessary SQL queries, it creates a pull request in the repository. The PR includes reasoning for the detection, generated by the AI, for human review.
Detection Maintenance Workflow
Beyond creation, the system also addresses the repetitive nature of detection maintenance tasks. These are often small, straightforward fixes that consume engineer time without offering significant learning opportunities.
- Triggering: Maintenance tasks (e.g., "add user X to allow list," "fix bug in Python rule," "update rule for new upstream log schema field") are initiated via Linear tickets.
- Agent Execution: The coding agent and the detection rule repository are bundled into a Docker image, stored in Amazon ECR, and executed via a Lambda function.
- Automated PR: A simple instruction in a Linear ticket (e.g., "Add
dev-tools-teamto the branch protection rule allow list forrepo-x") triggers the Lambda. The agent locates the relevant code, makes the necessary change, creates a test case, and opens a PR. This drastically reduces the manual effort involved in these common, yet tedious, operations.
Idea Generation
The system also tackles the challenge of identifying what to detect, breaking it into two scenarios:
- Common SaaS Applications (Okta, GitHub, CloudTrail): For widely used services, the AI can be fed public research, GitHub repositories, and blog posts detailing common attack patterns and misconfigurations. It then processes this information to suggest relevant detection ideas.
- Internal Tooling and New Log Sources: This is where the system's unique knowledge base comes into play:
- Vector Database: Vacasa maintains a vector database that stores:
- The company's risk registry.
- A log inventory with sample logs from internal systems and IoT devices.
- MITRE ATTACK detection strategies.
- AI Query: When a new internal log source is onboarded, or for existing internal tools, the AI is prompted with details about the log source and the tool's function. It then queries the vector database, correlating internal risks, log data, and MITRE strategies to generate highly contextualized detection ideas, complete with reasoning. This enables proactive threat hunting and ensures comprehensive coverage for bespoke internal systems.
Pitfalls and Learnings
Through months of iteration, Vacasa identified key learnings:
- Coding Agents > LLM Sequences: Initially, they used a sequence of LLM calls. Switching to a dedicated coding agent with integrated linting, formatting, and testing capabilities significantly improved output quality and reliability.
- Live Logs for Adaptability: Always feeding the latest log samples to the agents is crucial. This allows the system to adapt to upstream schema changes (e.g., a field's name or location changing in GitHub logs) without manual intervention.
- Addressing LLM Weaknesses with Skills: The aforementioned Athena query skill directly addresses a common LLM weakness (inconsistent capitalization in SQL for Athena). This highlights the need for specialized tools and iterative feedback loops to refine AI outputs for specific technical requirements.
Demo / Proof of Concept
▶ Watch: Demo: AI generating a complete detection rule pull request (6:20)
The talk included a compelling demonstration of the system's capabilities, showcasing the effortless workflow from idea to deployed detection.
The demo began with a Linear ticket related to Kubernetes API logs, which initially contained only a few detection ideas. Raphael added a label to this ticket, specifically "generate some ideas." This action triggered the detection idea generation workflow. The system then processed the request and, after a short, sped-up animation of the workflow running, created several new Linear tickets. Each new ticket represented an AI-generated detection idea, complete with details such as whether it should be a scheduled detection (e.g., running every 30 minutes) and relevant MITRE ATTACK labels.
Raphael then selected one of these AI-generated idea tickets to demonstrate the next phase. Clicking on the ticket revealed the AI's proposed detection, including its suggested run frequency. He then added another Linear label, "detection PR generation," which kicked off the detection Pull Request (PR) generation workflow. This workflow, also sped up for the demo, executed the coding agent.
Upon completion, the Linear ticket was automatically updated with a link to the newly created GitHub Pull Request. Reviewing the PR, Raphael highlighted several key elements:
- The PR comment included the AI's reasoning behind the detection.
- The PR contained multiple files:
- A test file with sample logs for the detection.
- A Python file containing the actual detection logic.
- An SQL query file (for this specific scheduled detection).
The demonstration concluded with the observation that this entire process, from a high-level idea to a ready-for-review PR with code, tests, and SQL, was "pretty effortless," validating the core premise of the talk.
Defensive Implications
▶ Watch: Explaining why AI is ideal for detection rule generation (8:00)
The "Detection Allegro" system implemented by Vacasa has profound implications for modern cybersecurity defense, offering a strategic advantage in an era of escalating threats and resource constraints.
- Accelerated Threat Coverage: The ability to develop new detections in minutes rather than hours or days means organizations can respond to emerging threats and intelligence much faster. This drastically reduces the Mean Time To Detect (MTTD), minimizing the window of opportunity for attackers. With over 150 new detections already created, Vacasa demonstrates a significantly expanded threat coverage footprint.
- Enhanced Operational Efficiency: By automating the repetitive and mundane aspects of detection engineering—such as writing boilerplate code, generating tests, fixing minor bugs, and managing allow-lists—security engineers can redirect their expertise towards more complex and strategic tasks. This includes advanced threat hunting, in-depth incident response, and proactive security architecture design, leading to a more impactful and fulfilling role for human analysts. The weekly usage of the bug-fixing lambda is a testament to this efficiency.
- Democratization of Security Contributions: Integrating the workflow with a user-friendly ticketing system like Linear enables non-DR engineers, who may not understand the intricacies of the security pipeline, to contribute valuable detection ideas. This fosters a "security-as-a-shared-responsibility" culture, leveraging the collective intelligence of the entire engineering organization to identify potential threats specific to their domains.
- Improved Detection Quality and Consistency: The agentic workflows enforce consistent coding standards, integrate linting and testing automatically, and map detections to the MITRE ATTACK framework. This ensures that all generated detections are of high quality, adhere to best practices, and are strategically aligned with recognized adversary behaviors. The specialized Athena query skill, which iteratively tests SQL against live data, further guarantees the accuracy and reliability of data queries.
- Proactive Security for Internal Systems: The use of a vector database containing internal risk registries, log inventories, and MITRE strategies allows the AI to generate highly relevant detection ideas for unique, proprietary internal tooling and newly onboarded log sources. This capability is critical for achieving comprehensive security coverage where public threat intelligence may not exist, moving beyond generic detections to context-specific insights.
- Reduced Alert Fatigue: Streamlining the process for allow-listing and fixing minor detection bugs directly combats alert fatigue. By quickly and effortlessly adjusting detection rules to account for legitimate exceptions or false positives, security teams can maintain trust in their alerts, ensuring that human attention is focused only on truly suspicious activities.
- Scalability with Organizational Growth: As Vacasa processes hundreds of terabytes of logs monthly and grows its infrastructure, the automated detection system scales with the company's needs. This prevents the detection engineering team from becoming a bottleneck, allowing them to maintain a strong security posture without proportional increases in manual effort.
In essence, "Detection Allegro" provides a blueprint for building a resilient, adaptable, and highly efficient defensive security program by strategically augmenting human expertise with intelligent automation.
Key Takeaways
- Agentic workflows with AI significantly boost efficiency: Vacasa achieved over 80% reduction in detection development time, from 2 hours to 20 minutes, for new rules and uses automated bug fixes weekly.
- Context-rich prompting is paramount: Providing AI agents with comprehensive context from log samples, schemas (Glue), internal documentation (Notion), and security frameworks (MITRE ATTACK) is crucial for generating accurate and relevant detections.
- Specialized "skills" overcome LLM limitations: Custom tools, like the Athena query skill that lints and iteratively tests SQL against live data, are essential for addressing specific technical requirements and ensuring the accuracy of AI-generated code.
- Ticketing system orchestration empowers collaboration: Integrating AI workflows with a platform like Linear centralizes information, streamlines processes, and enables non-DR engineers to contribute security ideas, fostering a broader security culture.
- AI excels at repetitive tasks: AI agents effectively handle mundane yet critical tasks such as allow-listing, minor bug fixes, and boilerplate code generation, freeing human engineers for complex threat modeling and strategic initiatives.
- Vector databases enable proactive, internal threat coverage: By leveraging a vector database with internal risk registries, log inventories, and MITRE strategies, AI can generate highly tailored detection ideas for unique internal tooling and new log sources where public intelligence is scarce.
About the Speaker(s)
Raphael Ruban and Chen Cao are both integral members of the detection response team at Vacasa. They operate as software engineers, focusing on building robust in-house detection and response capabilities and the necessary tooling to support the company's growth. Raphael, as a new hire, experienced firsthand the inefficiencies of traditional detection development, which inspired the development of the "Detection Allegro" system. Chen, his coworker, collaborates on these initiatives, contributing to the architectural design and iterative improvements of their agentic workflows. Together, they are at the forefront of applying AI to security operations, striving to make threat detection effortless and scalable.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent, production-grounded talk about using LLM-backed agentic workflows to automate detection rule generation and maintenance. The Vacasa team is clearly doing real work — the Athena skill, Linear integration, and vector DB for internal log coverage show genuine engineering effort — but the underlying concept (LLM + tools + feedback loop generates code) is not novel, and the 80% efficiency claim is a single-team anecdote without rigorous baseline measurement.
Heather Calloway (CISO) — SOLID
Credible, production-tested work on AI-assisted detection engineering from a lean team that actually built and shipped the thing. Useful for detection engineers and security architects thinking about AI tooling, but it stays firmly inside the engineering layer — governance, staffing model, and program-level implications are left untouched.