Prompt Scan Exploit AI’s Journey Through 0Days and 1000 Bugs

D. Jurado, J. Nogue

DEF CON 33 · Day 1 · Main Stage

Overview

This talk, presented by D. Jurado and J. Nogue at DEF CON, delves into the development and capabilities of an autonomous AI-powered penetration testing system. The speakers unveil a sophisticated architecture designed to mimic human pentesting methodologies, from initial reconnaissance and vulnerability discovery to intelligent validation and ethical boundary enforcement. The core innovation lies in its ability to not only identify a wide array of vulnerabilities but also to rigorously validate findings, combatting the common pitfalls of AI such as hallucinations and false positives.

Watch on YouTube

Visual summary for Prompt Scan Exploit AI’s Journey Through 0Days and 1000 Bugs by D. Jurado, J. Nogue
Visual summary for Prompt Scan Exploit AI’s Journey Through 0Days and 1000 Bugs by D. Jurado, J. Nogue

Key moments

  1. 0:00 Introduction to autonomous pentest AI architecture
  2. 1:40 Role of validators for AI false positives and cheating
  3. 2:20 AI scope control and preventing post-exploitation
  4. 3:30 Human in the loop for vulnerability reporting
  5. 4:20 How validators detect XSS vulnerabilities
  6. 6:20 Supported vulnerability classes and testing strategy
  7. 8:15 AI's autonomous operation and documentation usage

Prompt Scan Exploit AI’s Journey Through 0Days and 1000 Bugs

Speakers: D. Jurado; J. Nogue

Conference: DEF CON

YouTube: https://www.youtube.com/watch?v=sOkgHfu4lXY

Overview

This talk, presented by D. Jurado and J. Nogue at DEF CON, delves into the development and capabilities of an autonomous AI-powered penetration testing system. The speakers unveil a sophisticated architecture designed to mimic human pentesting methodologies, from initial reconnaissance and vulnerability discovery to intelligent validation and ethical boundary enforcement. The core innovation lies in its ability to not only identify a wide array of vulnerabilities but also to rigorously validate findings, combatting the common pitfalls of AI such as hallucinations and false positives.

The significance of this work extends beyond mere automation; it represents a substantial leap in leveraging artificial intelligence for offensive security. By detailing components like a central coordinator, specialized agents, robust validators, and a protective proxy, the speakers illustrate a pragmatic approach to building an AI that can navigate complex applications. The insights shared, particularly on "alloy models" and optimizing resource allocation, offer a blueprint for future AI-driven security tools, highlighting the imperative of balancing autonomous capabilities with stringent control and ethical considerations.

The ultimate goal of such a system is to enhance the efficiency and coverage of security assessments, potentially uncovering vulnerabilities at a scale and speed unachievable by human teams alone. This talk is crucial for anyone interested in the intersection of AI and cybersecurity, providing a deep dive into the technical challenges and innovative solutions in creating intelligent agents capable of discovering real-world security flaws.

Background

▶ Watch: Introduction to autonomous pentest AI architecture (0:00)

The landscape of cybersecurity is continually evolving, with an increasing number of complex applications and systems requiring constant vigilance against emerging threats. Traditional penetration testing, while effective, is often resource-intensive, time-consuming, and limited by the expertise and capacity of human testers. This challenge has fueled the quest for automated solutions, from static and dynamic application security testing (SAST/DAST) tools to fuzzers and vulnerability scanners. However, many automated tools frequently struggle with context, complex business logic, and high rates of false positives, necessitating significant human oversight.

The advent of large language models (LLMs) and advanced AI has opened new avenues for automation in security. The promise is an AI that can think, reason, and adapt like a human pentester, moving beyond signature-based detection to understand application behavior and exploit novel weaknesses. The problem, however, is that LLMs are prone to "hallucinations" – generating plausible but incorrect information – and can "cheat" by attempting to satisfy a prompt without genuinely completing the task. Furthermore, uncontrolled AI in an offensive security context poses significant ethical and operational risks, from unintended post-exploitation to breaching legal boundaries.

This talk directly addresses these challenges by presenting a system built on prior research and practical experience, including insights from a Black Hat talk by Brendan, one of the team members, focusing specifically on the crucial role of validators. The foundation of their approach is to create an autonomous system that not only finds vulnerabilities but also incorporates self-correction and stringent ethical guardrails, bridging the gap between raw AI capability and reliable, responsible security testing.

Key Findings

▶ Watch: AI scope control and preventing post-exploitation (2:20)

The speakers presented several pivotal findings and architectural innovations critical to developing an effective autonomous AI pentesting system:

  1. Autonomous AI Pentesting Architecture: The core contribution is a modular system comprising a Coordinator, multiple Agents, Validators, and a Proxy. This architecture allows for distributed, intelligent vulnerability discovery, mimicking the structure of a human pentesting team. The Coordinator acts as the project manager, assigning tasks and prioritizing analysis, while Agents execute the actual testing.
  2. Crucial Role of Validators: A standout finding is the absolute necessity of validators to counteract the inherent weaknesses of LLMs, such as false positives, hallucinations, and "cheating." These validators act as a "second pair of eyes," independently verifying reported vulnerabilities, ensuring a high degree of accuracy and trustworthiness in the AI's findings. For some vulnerability types, they achieve zero or near-zero false positive rates.
  3. "Alloy Models" for Enhanced Performance: The team discovered that combining different LLMs, referred to as alloy models, significantly outperforms using a single model. By dynamically selecting between two or more models for queries (e.g., "flipping a coin"), and managing context across them, the system leverages the diverse strengths of various LLMs. This effect is amplified when the chosen models are more distinct in their capabilities.
  4. Optimized Resource Allocation via Diminishing Returns: Through extensive data analysis, the team identified that the optimal number of runs or agents required to find vulnerabilities varies significantly by type. For instance, discovering "exposed secrets" might yield high returns with just one run, whereas identifying complex Remote Code Execution (RCE) vulnerabilities often requires more attempts. This insight allows for intelligent resource allocation, leading to "more vulnerabilities with much less attempts and resources."
  5. Ethical Boundaries and Policy Enforcement: The system incorporates robust controls, including a proxy to enforce scope and prevent post-exploitation, and a policy checker to block malicious or unintended actions (e.g., deleting database records). This demonstrates a commitment to responsible AI deployment in offensive security, ensuring that autonomous actions remain within predefined ethical and legal parameters.
  6. Blackbox Testing and Internal Benchmarking: The system operates primarily in a blackbox fashion, without client intervention, and does not rely on traditional CTF (Capture The Flag) challenges for validation. Instead, it uses custom benchmarks for performance evaluation, acknowledging that live application testing quickly leads to fixed vulnerabilities, making consistent benchmarking difficult. CTF flags are used internally as "compensation" tokens for the models, not as actual validation mechanisms.

Technical Deep Dive

▶ Watch: Human in the loop for vulnerability reporting (3:30)

The autonomous pentesting AI system described by D. Jurado and J. Nogue is an intricate orchestration of specialized components, each playing a vital role in its overall functionality and reliability.

At the apex of this architecture is the Coordinator. This component functions as the "manager of a pentesting operation," responsible for strategic oversight. Its tasks include:

  • Analysis Priorities: Determining which areas of an application to focus on based on context and potential impact.
  • Recon and Discovery: Performing initial reconnaissance to map out application assets and attack surface.
  • Task Assignment: Distributing specific vulnerability hunting tasks to individual Agents.
  • Context Management: Utilizing provided documentation, such as Swagger files or other API specifications, to enhance coverage and understanding of the target application.

Reporting to the Coordinator are the Agents. Each agent is equipped with access to an "attacking machine" loaded with custom-built tools adapted for AI operation, including headless browsers. These tools allow the agents to interact with the target application, simulate user actions, and attempt various exploits. The agents are given "complete freedom" within their assigned tasks, meaning they are not constrained by predefined steps, allowing for more creative and adaptive exploration.

A critical innovation is the Validators component, which addresses the inherent challenges of LLMs. As LLMs can generate false positives, "hallucinate" information, or even "cheat" by reporting non-existent findings, validators serve as a "second pair of eyes." Each vulnerability type has a dedicated validator. For example, to validate a Cross-Site Scripting (XSS) vulnerability, a validator might use a headless browser (like one driven by Puppeteer) to visit a crafted URL containing the payload. If the expected pop-up or constant message appears, the XSS is confirmed. This process aims for a near-zero false positive rate for many vulnerability classes. The concept of validators was further elaborated upon in a Black Hat talk by Brendan, emphasizing their importance in ensuring report accuracy.

To maintain control and ethical boundaries, the system integrates a Proxy. This proxy establishes and enforces strict boundaries around the target scope. It parses "element programs" (likely referring to application components or assets) and embeds them into these boundaries, ensuring that agents only interact with in-scope targets. Crucially, the proxy also prevents post-exploitation activities. The speakers emphasized that for platforms like HackerOne, where the system might be randomly scanning the internet, allowing post-exploitation could lead to severe consequences, such as a "full cloud environment takeover." The goal is to identify a "good vulnerability" and report it, not to escalate privileges or cause further damage. A policy checker further reinforces these boundaries by blocking actions deemed "bad for the system," such as deleting database records.

The system is designed to identify a broad spectrum of vulnerabilities, including:

  • Remote Code Execution (RCE)
  • File Read vulnerabilities (e.g., XXE – XML External Entity, Path Traversal)
  • SQL Injection (both blind and time-based)
  • Various forms of XSS (including stored, post-message based, reflected, and blind)
  • Open Redirects
  • Exposed Secrets
  • Server-Side Request Forgery (SSRF)
  • Server-Side Template Injection (SSTI)
  • Cache Poisoning
  • Insecure Direct Object References (IDORs), though validation for these is noted as being particularly challenging and still under improvement.

A significant technical advancement is the concept of Alloy Models. Instead of relying on a single LLM, the system dynamically switches between different models for queries. The speakers observed that "every combination of two models were outperforming like the other ones" and that "the more different are those models, the better they were performing." This suggests that diverse cognitive approaches from different models lead to superior results. The system manages context and history across these models, allowing for seamless transitions. While specific models were not named, the implication is that both public and potentially proprietary models are utilized. This strategy is particularly effective for broad vulnerability hunting, though for highly specialized tasks where one model excels (e.g., XSS), using that single model might be more efficient.

Finally, the system also incorporates solutions for common web challenges, such as CAPTCHA solving, further demonstrating its autonomy in navigating complex web applications. The entire process for platforms like HackerOne is blackbox, with the human only involved in the final reporting phase, primarily due to platform policies requiring human verification.

Demo / Proof of Concept

▶ Watch: Supported vulnerability classes and testing strategy (6:20)

While the speakers had prepared a detailed demonstration for their planned 1-hour slot, the talk was unfortunately cut short to 30 minutes due to technical difficulties. As a result, no live demo or proof of concept was presented during the session. The speakers indicated that the full talk, including demonstrations, would be published on their social media and other platforms in the days following the conference. Therefore, attendees and readers of this article must refer to the full, unedited recording or subsequent publications by the speakers for visual examples of the system in action. The intent was clearly to showcase the autonomous AI performing tasks, but circumstances prevented it.

Defensive Implications

▶ Watch: AI's autonomous operation and documentation usage (8:15)

The emergence of sophisticated autonomous AI pentesting systems like the one described by D. Jurado and J. Nogue presents significant implications for cybersecurity defenders. Organizations must recognize that the landscape of threats is evolving, with future adversaries potentially leveraging similar AI capabilities to discover and exploit vulnerabilities at an unprecedented scale and speed.

  1. Anticipate AI-Driven Attacks: Defenders should assume that AI-powered tools will become a standard part of attackers' arsenals. This means traditional defenses, which often rely on patterns or signatures, might be insufficient against AI that can intelligently probe, adapt, and chain exploits. Focus should shift towards understanding attack logic and context.
  2. Embrace Robust Validation: The talk heavily emphasizes the critical role of validators in eliminating false positives and hallucinations from AI findings. Defenders should adopt a similar mindset when evaluating their own security tools and processes. Implementing rigorous, independent validation steps for all vulnerability reports, whether from automated scanners or human testers, is paramount to ensure resources are spent addressing real threats.
  3. Strengthen In-Depth Defense: The AI's ability to find a wide array of vulnerabilities (RCE, SQLi, XSS, SSRF, SSTI, cache poisoning, IDORs, exposed secrets) underscores the need for comprehensive, layered security. Organizations must not rely on single points of failure but build defenses that account for diverse attack vectors across the entire application stack.
  4. Implement Strict Policy and Boundary Enforcement: The speakers highlighted the importance of a proxy and policy checker to prevent post-exploitation and out-of-scope actions. Defenders should apply these principles internally, ensuring that even legitimate testing or internal tools operate within strict boundaries. This includes robust access controls, network segmentation, and runtime application self-protection (RASP) to mitigate the impact of any compromised component.
  5. Focus on Data-Driven Security: The finding about "diminishing returns" and optimizing resource allocation based on vulnerability types is relevant for defenders. By analyzing their own vulnerability data, organizations can identify which types of flaws are most prevalent or require more extensive testing, allowing for more efficient allocation of defensive resources and focused remediation efforts.
  6. Secure Documentation and API Specifications: The Coordinator's ability to leverage documentation like Swagger files for better coverage highlights a potential defensive vulnerability. Attackers using similar AI could gain significant advantage from well-documented APIs or internal systems. Defenders must ensure that such documentation is not only accurate but also securely stored and accessible only to authorized personnel.
  7. Monitor for Sophisticated Anomalies: AI-driven attacks might exhibit more complex, adaptive behaviors than traditional automated scans. Defenders should invest in advanced anomaly detection systems that can identify unusual sequences of actions, context-aware deviations, and coordinated multi-stage attacks that might indicate an AI-driven adversary.

In essence, the talk serves as a proactive warning and a guide. Defenders must not only prepare for AI that can find vulnerabilities but also learn from the engineering principles behind such systems to build more resilient and intelligent defenses.

Key Takeaways

  • Autonomous AI pentesting is a reality: The architecture presented demonstrates a functional, modular system capable of performing comprehensive vulnerability assessments, mimicking human pentesting methodologies.
  • Validation is paramount for AI accuracy: Robust validators are crucial for mitigating false positives, hallucinations, and "cheating" by LLMs, ensuring the reliability and trustworthiness of AI-discovered vulnerabilities.
  • Combining diverse AI models ("Alloy Models") enhances discovery: Leveraging multiple LLMs with varying strengths, especially those that are more distinct, significantly outperforms single-model approaches for broad vulnerability hunting.
  • Ethical boundaries and policy enforcement are critical: The system incorporates a proxy and policy checker to enforce scope, prevent post-exploitation, and block malicious actions, underscoring the necessity of responsible AI deployment in offensive security.
  • Resource optimization improves efficiency: Data-driven analysis of "diminishing returns" allows for intelligent allocation of testing resources, leading to more vulnerabilities discovered with fewer attempts for specific vulnerability types.
  • AI can find a broad spectrum of vulnerabilities: The system is designed to identify a wide range of flaws, including RCE, SQLi, XSS, SSRF, SSTI, cache poisoning, and exposed secrets, highlighting its comprehensive attack surface coverage.

About the Speaker(s)

D. Jurado and J. Nogue are researchers and practitioners in the field of artificial intelligence and cybersecurity. Their work focuses on developing autonomous systems capable of advanced penetration testing. While specific titles and affiliations for both speakers were not explicitly stated in the provided transcript or metadata, the technical depth and practical insights shared during their DEF CON talk clearly position them as experts in the intersection of AI and offensive security. They have been involved in building sophisticated AI solutions for vulnerability discovery, including the development of "alloy models" and robust validation mechanisms. The talk also makes a passing reference to "Albert the head of AI for expo" in connection with the alloy model work, suggesting a possible affiliation or collaboration with a company or research group named "expo" in their AI endeavors.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Jurado and Nogue present a real, working autonomous AI pentesting system with genuine engineering effort behind it — the validator architecture, alloy model ensemble approach, and diminishing-returns resource allocation are legitimate contributions worth hearing about. The talk is hurt badly by the demo getting axed and by a write-up that oversells the novelty; the core ideas are solid but not groundbreaking enough to clear 4 stars without the proof.

Heather Calloway (CISO) — WEAK

Technically substantive offensive research with real engineering ambition, but it never crosses into defender utility or institutional relevance. The gaps aren't incidental — they're structural.

→ Top-rated talks at DEF CON 33

All talks from DEF CON 33