Decision Making in Adversarial Automation
Bobby Kuzma, Michael Odell
DEF CON 33 · Day 1 · Main Stage
Overview
In the intricate dance between attackers and defenders, the speed and accuracy of decision-making often dictate the outcome. This talk, "Decision Making in Adversarial Automation," delivered by Bobby Kuzma and Michael Odell at DEF CON, delves into the theoretical underpinnings and practical applications of automating adversarial decision processes. The speakers explore how concepts from artificial intelligence, game theory, and graph theory can be leveraged to create more sophisticated and adaptive automated attack tools, moving beyond simple deterministic playbooks to embrace probabilistic and context-aware strategies.

Key moments
- 0:00 Introduction to adversarial decision making and math warning
- 2:00 Defining stochastic parrots and deterministic predators in automation
- 3:50 Understanding the OODA Loop for adversarial decision-making
- 4:30 Analyzing OODA Loop latency: SOC vs. attacker speed
- 5:50 Markov Decision Process (MDP) with full environment knowledge
- 7:00 Partially Observable MDP (POMDP) for unknown environments
- 8:20 Comparing stochastic and deterministic planning methodologies
Decision Making in Adversarial Automation
Speakers: Bobby Kuzma, Michael Odell
Conference: DEF CON
YouTube: https://www.youtube.com/watch?v=9to68PN5rRU
Overview
In the intricate dance between attackers and defenders, the speed and accuracy of decision-making often dictate the outcome. This talk, "Decision Making in Adversarial Automation," delivered by Bobby Kuzma and Michael Odell at DEF CON, delves into the theoretical underpinnings and practical applications of automating adversarial decision processes. The speakers explore how concepts from artificial intelligence, game theory, and graph theory can be leveraged to create more sophisticated and adaptive automated attack tools, moving beyond simple deterministic playbooks to embrace probabilistic and context-aware strategies.
The presentation provides a comprehensive look at various decision-making models, contrasting deterministic approaches with stochastic methodologies, and discussing their applicability in partially observable, dynamic environments typical of real-world networks. It highlights the inherent challenges in codifying intuition and evaluating optimal actions in complex adversarial scenarios. Ultimately, Kuzma and Odell propose a framework that uses graph representations of network states and subgraph matching to guide automated penetration testing, culminating in a demonstration of a Domain Specific Language (DSL) and a novel graph database solution designed to make these advanced decision processes tangible for security practitioners.
The significance of this work lies in its potential to revolutionize how adversaries operate and, by extension, how defenders must adapt. By understanding how sophisticated automation can observe, orient, decide, and act with speed and adaptability, organizations can better anticipate threats, design more resilient systems, and develop more effective detection and response strategies. This talk serves as a crucial primer for anyone involved in offensive security, incident response, or security architecture, offering insights into the evolving landscape of automated cyber warfare.
Background
▶ Watch: Introduction to adversarial decision making and math warning (0:00)
The concept of adversarial decision-making is deeply rooted in military strategy, perhaps most famously encapsulated by Colonel John Boyd's OODA loop (Observe, Orient, Decide, Act). This iterative cycle describes how individuals and organizations process information and react to their environment. In a cyber context, an attacker observes the network, orients themselves to its vulnerabilities, decides on an action, and then executes it. The critical insight is that by acting rapidly, an attacker can disrupt a defender's OODA loop, forcing them to continuously reset their observation and orientation, thereby gaining a significant time advantage. As Bobby Kuzma illustrates, if a Security Operations Center (SOC) has an average latency of five minutes (or more) from event to alert, an attacker can compromise numerous systems and accounts before the defender even begins to contain the initial breach.
Traditional adversarial operations, whether manual or partially automated, often rely on deterministic playbooks or checklists. These decision trees are effective in well-understood, limited environments, but they struggle to scale or adapt to the vast, dynamic, and often partially observable nature of enterprise networks. Attackers operating with such rigid methodologies become predictable; defenders who know the playbook can anticipate their moves. This problem is exacerbated by the sheer scale of modern networks, which can involve tens or hundreds of thousands of systems, services, and users, creating an exponentially growing "state space" that manual or simple automated processes cannot efficiently navigate.
Prior work in automated penetration testing has typically focused on tool chaining and basic conditional logic. While tools like Metasploit and various scripting frameworks automate individual actions, they often lack sophisticated, context-aware decision-making capabilities that mimic human intuition or strategic planning. The challenge lies in moving beyond simply executing commands to making intelligent choices about what to do next, where to go, and how to minimize risk (like detection) while maximizing reward (like privilege escalation or data exfiltration). The talk highlights that current adversary simulation often involves "feeling the environment" – an intuitive process that is notoriously difficult to codify into algorithms. This gap between human intuition and automated execution is precisely what the presented research aims to bridge by drawing upon advanced decision science models.
Key Findings
▶ Watch: Understanding the OODA Loop for adversarial decision-making (3:50)
The talk presents several key findings and conceptual contributions toward building intelligent adversarial automation:
- The Superiority of Probabilistic Models in Partially Observable Environments: While Markov Decision Processes (MDPs) are effective for "white box" scenarios where the environment is fully known (like an Atari game where the screen state is the universe), real-world networks are "black box" or partially observable. For these, Partially Observable Markov Decision Processes (POMDPs) are crucial. POMDPs factor in unknown unknowns and the potentiality of unobserved things, allowing an automated agent to weigh actions that gather more information against actions that exploit known vulnerabilities, making decisions under uncertainty.
- Quantifying Risk Factors Beyond Reward: Beyond simple "score-based" rewards, the talk emphasizes incorporating additional risk factors into decision models, such as detectability. An action's likelihood of triggering a defensive response can be weighted, allowing an automated adversary to prioritize stealth over speed or vice-versa, depending on the operational goals. This introduces a nuanced decision-making capability that moves beyond purely exploitative actions.
- Graph Theory as the Foundation for State Representation: A fundamental finding is that network environments and their connections can be effectively represented as property graphs. This aligns with the "attackers think in graphs" paradigm popularized by tools like BloodHound. By representing the state of an environment as a graph, complex relationships between systems, users, and resources become queryable and analyzable, enabling more sophisticated decision-making.
- Subgraph Matching for Pattern Recognition and Action Selection: To navigate the vast state space of real-world networks, the talk proposes leveraging subgraph matching. This technique allows the automated system to define patterns of prerequisites (e.g., "I have admin on System A, and System A has a connection to System B, and System B runs Service X") and then efficiently search the network graph for instances of these patterns. This enables the automation to identify relevant opportunities for action (e.g., "given this pattern, I can run tool Y to escalate privileges on System B").
- Introduction of a Domain Specific Language (DSL) and Kuzu Database for Practical Implementation: A significant contribution is the development of a DSL (based on YAML) that allows practitioners to define preconditions and expected outcomes for tool usage within the graph. This DSL, combined with the use of Kuzu, an open-source embedded graph database, provides a practical and scalable architecture for implementing these advanced decision-making systems. Kuzu's in-process, single-file nature addresses some of the scalability and deployment challenges associated with larger graph databases like Neo4j, making it more suitable for rapid, dynamic adversarial automation.
Technical Deep Dive
▶ Watch: Analyzing OODA Loop latency: SOC vs. attacker speed (4:30)
The core technical contribution lies in adapting advanced decision-making paradigms from artificial intelligence and applying them to offensive security. The speakers introduce several key models and mechanisms:
1. Markov Decision Processes (MDPs) and Partially Observable Markov Decision Processes (POMDPs):
- An MDP models decision-making in environments where the outcome of an action is partly random but the system's state is fully observable. It comprises a set of states (S), a set of actions (A), transition probabilities (P) between states given an action, rewards (R) for actions, and a discount factor. The goal is to find a policy (π) that maximizes the expected cumulative reward. An example provided is an AI learning to play Atari games, where the screen is the fully known state.
- A POMDP extends the MDP to scenarios where the agent does not fully observe the current state. Instead, it maintains a belief state (b), which is a probability distribution over all possible states. The agent makes decisions based on these observations and its belief state, meaning it must weigh actions that gather more information (exploration) against actions that exploit known opportunities (exploitation). This is crucial for "black box" penetration testing where the attacker learns as they go.
2. Decision Policies and Risk Factors:
- Stochastic Policies: In these, actions are chosen probabilistically, often with equal likelihood (a "random walk"). This offers unpredictability but may not be optimal.
- Deterministic Policies: Here, for a given state, the "best" action (e.g., the one with maximum quality or reward) is always chosen. This leads to predictable behavior, which can be a double-edged sword: efficient but detectable if the defender knows the policy.
- Risk Factors: The talk introduces the critical concept of integrating detectability as a risk factor. An action's detectability can be modeled as a value between 0 and 1 (or 0 and 100), influencing the decision. The system can then weigh the worst possible outcome (e.g., losing a foothold, federal agency involvement) against the best possible outcome (e.g., gaining access to more systems, escalating privileges). This allows for more sophisticated, context-aware trade-offs.
3. Decision-Making Mechanisms:
- Policy Gradients: These imagine the optimum decision outcomes as a multi-dimensional surface. They are typically used in environments with continuous actions (e.g., turning a steering wheel 10 or 12 degrees) but are also applicable to discrete actions (e.g., running a specific command). The system learns to adjust its policy to "climb" this surface towards optimal outcomes.
- Boltzman Exploration: This mechanism relies on the "quality" (Q) of a state-action pair. The probability of choosing a particular action is proportional to its relative quality compared to all other possible actions. This allows for a balance between exploiting high-quality actions and exploring potentially unknown but beneficial actions.
- Decision Trees: These are essentially algorithmic checklists, finite state machines where actions are chosen based on a series of conditions. While simple and explainable ("we did X because we were a domain admin"), they are rigid and do not scale well to complex, dynamic environments, making them predictable to defenders.
4. Graph-Based State Representation and Subgraph Matching:
- The talk strongly advocates for representing the network environment as a property graph. Nodes represent entities (systems, users, services) and edges represent relationships (e.g., "System A runs Service X," "User Y has admin on System Z"). This is a direct application of the "attackers think in graphs" paradigm.
- To make decisions within this vast graph, subgraph matching is employed. This technique involves defining specific patterns (e.g., "a host running an outdated service, connected to a user with weak credentials") and then efficiently searching the larger network graph for instances of these patterns. When a pattern is matched, it indicates a potential opportunity for an action. The speakers mention that there's extensive prior art on speeding up subgraph matching from domains like social media networks and financial fraud detection.
5. Implementation Architecture:
- Domain Specific Language (DSL): A YAML-based DSL is introduced to define the prerequisites and expected outcomes of tool usage in terms of graph patterns. This allows operators to codify their knowledge and strategies into machine-readable rules.
- Kuzu Graph Database: The choice of Kuzu is a practical and significant one. Unlike larger, server-based graph databases, Kuzu is an open-source, embedded graph database that runs in-process and stores data in a single file. This makes it highly suitable for dynamic, on-the-fly analysis in adversarial automation, eliminating the overhead of managing a separate database server. Kuzu also supports Cypher, a widely used graph query language, making it accessible for those familiar with graph databases.
- Data Collection: The system collects console data from operator systems to reconstruct the environment's state. This includes outputs from enumeration tools like Nmap and basic Linux enumeration commands, as well as credentials like hashes. This collected data is then used to update the dynamic graph representation of the target environment.
By combining these elements, the proposed system aims to create an automated adversary that can observe the environment, build a dynamic graph representation, use subgraph matching to identify opportunities based on predefined patterns, and then select optimal actions guided by probabilistic decision models that factor in detectability and potential rewards.
Demo / Proof of Concept
▶ Watch: Partially Observable MDP (POMDP) for unknown environments (7:00)
The speakers presented a recorded demo to illustrate the practical application of their concepts, specifically avoiding live demo "taunting the demo gods." The demonstration showcased a simplified penetration testing scenario leveraging their developed DSL and the Kuzu graph database.
The demo environment was a "toy example," specifically the game backup directory (a known vulnerable environment often used for security testing). The goal was to show how the system could collect information, update its internal graph representation of the environment, and then suggest the "next best step" for an attacker.
Here's a breakdown of the demo's workflow:
- Priming the Pump with Manual Exploration: The demo began by manually inputting some initial exploration commands. This simulates an initial foothold or reconnaissance phase. Console data from these commands was captured and fed into the system.
- Dynamic Graph Updates: As input was captured (e.g., results from
nmapscans, basic Linux enumeration commands), the system processed this information. On the right side of the screen, a representation of the updated graph was shown, illustrating how new nodes (e.g., discovered hosts, services, users) and edges (e.g., connections, running services) were added to the environment's state. - Automated Suggestion Engine: After processing the initial enumeration data, the system began to suggest commands to run. These suggestions were derived from the underlying decision model, which analyzed the current graph state using subgraph matching to identify patterns that indicated actionable opportunities.
- Iterative Process: The operator would paste in the suggested commands, execute them, and the system would capture the new output. This new information would further update the graph, leading to new suggestions. This iterative loop demonstrated the dynamic, adaptive nature of the automated decision-making.
- Handling Limited Opportunities: At one point, the model became "angry" or "ran out of things to run" because the simplified test example was only primed with 14 different tool usages. This highlighted that the system is constrained by the knowledge (the DSL-defined patterns) it possesses and the current state of the environment.
- Hash Extraction and Cracking: A key moment in the demo involved the system identifying a hash (e.g., from
/etc/shadoworsamfiles). Once this hash was processed and added to the graph, the system suggested a "cracking" action. - Credential Usage and Privilege Escalation: After the (simulated) cracking was complete, yielding a credential, the system identified new opportunities. For instance, it suggested performing a "secret dump" (e.g., using
mimikatzor similar tools) on a system, leveraging the newly acquired credential. This led to a "domain dumped" outcome, signifying a significant compromise.
The demo, despite being a "toy example," successfully illustrated the core principles: dynamic environment representation via a graph, automated information collection, intelligent action suggestion based on graph patterns, and an iterative feedback loop that allows the system to progress through a simulated attack chain. The use of the Kuzu database in the background facilitated these real-time graph updates and queries.
Defensive Implications
▶ Watch: Comparing stochastic and deterministic planning methodologies (8:20)
The insights presented in this talk have profound implications for defensive security strategies. Understanding how adversaries can leverage advanced automation and decision-making models is crucial for building more resilient defenses.
- Anticipate Adaptive Adversaries: Defenders must move beyond expecting adversaries to follow predictable, linear playbooks. Automated adversaries, especially those employing POMDPs and integrating risk factors like detectability, will be more adaptive, making decisions based on partially observed information and dynamically adjusting their tactics to avoid detection. This means signature-based detections and rigid rule sets will become increasingly ineffective.
- Focus on Disrupting the Adversary's OODA Loop: The talk re-emphasizes the importance of the OODA loop. Defenders need to accelerate their own OODA loop to disrupt the attacker's. This involves:
- Rapid Observation: Improving telemetry collection, log aggregation, and real-time monitoring to quickly identify anomalous activities.
- Effective Orientation: Enhancing threat intelligence, contextualizing alerts, and understanding attacker TTPs to correctly interpret observations.
- Swift Decision: Automating response actions where possible, establishing clear playbooks for common incidents, and empowering analysts with decision support tools.
- Decisive Action: Implementing automated containment, eradication, and recovery measures.
Reducing the latency from event to alert and subsequent action is paramount.
- Prioritize Graph-Based Visibility and Analysis: Since attackers "think in graphs" and use them to represent their understanding of the network, defenders should adopt similar paradigms. Tools like BloodHound are already popular for defensive analysis, but the emphasis should be on maintaining a dynamic, up-to-date graph of the enterprise environment. This allows defenders to:
- Identify critical attack paths before adversaries do.
- Spot misconfigurations and unintended trust relationships.
- Detect anomalous changes in the graph that might indicate adversarial activity (e.g., new admin accounts, unusual service-to-host mappings).
- Use subgraph matching defensively to identify known malicious patterns or indicators of compromise within their own network graph.
- Challenge Assumptions about Full Observability: Defenders often operate under the assumption that their security tools provide full visibility. This talk highlights that adversaries operate in partially observable environments and make decisions under uncertainty. Defenders should likewise consider what information an attacker doesn't have access to and how that can be leveraged. For instance, creating honeypots or deceptive assets that appear valuable but provide misleading information can disrupt an automated adversary's belief state and lead them down dead ends.
- Enhance Behavioral Detection and Anomaly Analysis: As automated adversaries become more sophisticated, they will attempt to mimic legitimate user behavior or blend into normal network traffic. This necessitates a shift towards behavioral detection and anomaly analysis. Defenders should invest in technologies that baseline normal activity and flag deviations, rather than relying solely on signatures of known malicious tools or commands.
- Focus on Hardening Foundational Services: The demo showed how easily a hash could be used to facilitate a "secret dump." This underscores the importance of hardening foundational services like Active Directory, implementing multi-factor authentication (MFA) everywhere, and enforcing strong password policies. Reducing the "reward" an attacker gets from compromising a single credential can significantly impede their progress.
- Embrace Continuous Adversary Emulation: To test the efficacy of defenses against these advanced automated adversaries, organizations should adopt continuous adversary emulation. Tools like Ludus (for standing up repeatable cyber ranges) and the ability to convert BloodHound exports into Ludus ranges are invaluable for creating realistic test environments. Regularly emulating sophisticated attack chains, perhaps even using the DSL and Kuzu concepts presented, can reveal weaknesses in defensive posture and detection capabilities.
In essence, defenders must evolve from reactive, signature-based approaches to proactive, intelligent, and adaptive strategies that consider the adversary's decision-making process. By understanding the models and techniques discussed, organizations can build more robust defenses that are prepared for the next generation of automated cyber threats.
Key Takeaways
- Adversarial automation is evolving beyond deterministic playbooks: Future attacks will leverage probabilistic decision-making, like Partially Observable Markov Decision Processes (POMDPs), to adapt to uncertain, real-world network environments.
- Risk factors like detectability are crucial for advanced adversaries: Automated tools will weigh potential rewards against the likelihood of detection, allowing for more nuanced and stealthy operations.
- Graph theory is central to understanding and navigating complex networks: Attackers and defenders alike benefit from representing network states as property graphs to identify relationships and attack paths.
- Subgraph matching enables intelligent action selection: By defining patterns of prerequisites and outcomes in a Domain Specific Language (DSL), automated systems can efficiently identify opportunities in a network graph and suggest optimal next steps.
- Practical tools like Kuzu are making advanced automation accessible: The use of an embedded, open-source graph database like Kuzu combined with a flexible DSL provides a scalable and deployable solution for implementing these sophisticated decision-making frameworks.
- Defenders must accelerate their OODA loop and embrace graph-based analysis: To counter adaptive automation, organizations need faster detection and response, better contextual intelligence, and continuous adversary emulation using tools like Ludus and BloodHound to test and improve their defenses.
About the Speaker(s)
Bobby Kuzma is described as a dad, husband, and nerd who enjoys "chasing rockets" when he's not "breaking things." His professional focus lies in security, particularly in understanding and exploiting system vulnerabilities, as well as developing advanced adversarial capabilities. His work reflects a deep technical understanding and a practical, hands-on approach to security challenges.
Michael Odell (Moch) is introduced as a fellow nerd who works with Bobby Kuzma. He contributes to the technical development and research, bringing his expertise to projects like the adversarial automation discussed in the talk. Michael is also noted for his unique hobby of "bonsai-ing whatever that is," adding a touch of personal flair to his bio. Together, their combined experience forms a formidable team in the realm of offensive security research.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Kuzma and Odell are clearly doing real work here — the DSL-plus-Kuzu architecture is a legitimate engineering contribution and the POMDP framing is the right abstraction for adversarial automation. But the talk sits awkwardly between a conceptual primer and a working system demo, and the 'toy example with 14 tools' proof-of-concept doesn't yet close the gap between the theoretical ambition and something a practitioner can weaponize tomorrow.
Heather Calloway (CISO) — WEAK
Technically credible research that applies decision science to adversarial automation — POMDP framing, graph-based state representation, a working DSL. But the talk is fundamentally built for offensive researchers, and the defensive translation is thin, generic, and bolted on. No institutional decision-maker leaves knowing what to do differently.