When the Model Outsmarts the Challenge: Building and Breaking AI Security CTFs
SoYeon Kim (Researcher · NSHC), Hea-Eun Moon (Lab Lead · NSHC), Sang-tae Woo (CISO · NSHC)
Nullcon Goa 2026 · Day 1
Overview
This talk, presented by SoYeon Kim, Hea-Eun Moon, and Sang-tae Woo from NSHC in Korea, delves into the intricate process of designing, organizing, and learning from an Artificial Intelligence (AI) security Capture The Flag (CTF) competition. Titled "When the Model Outsmarts the Challenge: Building and Breaking AI Security CTFs," the presentation highlights the challenges and unexpected outcomes encountered when integrating rapidly evolving AI technologies, particularly large language models (LLMs), into a competitive security environment. The speakers share their experiences from the AI Cyber Defense Contest (ACDC), a pioneering AI CTF held in Korea.

Key moments
- 0:00 Introduction and overview of the AI Security CTF
- 4:00 Qualifier and final formats, core challenge categories
- 6:00 Key design decisions: participants, models, costs, format
- 8:00 AI for Security challenge: Identifying keywords from multilingual audio
- 10:00 Security for AI challenge: Prompt injection to access private data
- 11:30 AI Infra challenge: Exploiting medical AI service for data exfiltration
When the Model Outsmarts the Challenge: Building and Breaking AI Security CTFs
Speakers: SoYeon Kim (Researcher, NSHC); Hea-Eun Moon (Lab Lead, NSHC); Sang-tae Woo (CISO, NSHC)
Conference: Nullcon
YouTube: https://www.youtube.com/watch?v=6hJy5glTbP0
Overview
This talk, presented by SoYeon Kim, Hea-Eun Moon, and Sang-tae Woo from NSHC in Korea, delves into the intricate process of designing, organizing, and learning from an Artificial Intelligence (AI) security Capture The Flag (CTF) competition. Titled "When the Model Outsmarts the Challenge: Building and Breaking AI Security CTFs," the presentation highlights the challenges and unexpected outcomes encountered when integrating rapidly evolving AI technologies, particularly large language models (LLMs), into a competitive security environment. The speakers share their experiences from the AI Cyber Defense Contest (ACDC), a pioneering AI CTF held in Korea.
The core motivation behind creating such a contest was to proactively address the burgeoning security risks associated with AI and to foster a foundational understanding of safe AI deployment. As AI rapidly integrates into everyday life, the need for a dedicated platform to evaluate both AI usage skills and security capabilities becomes paramount. The NSHC team meticulously crafted challenges across three critical categories—AI for Security, Security for AI, and AI Infrastructure—aiming to create a competition that was both relevant to current AI advancements and challenging for participants with diverse backgrounds.
This article explores the methodologies, design considerations, and crucial lessons learned from operating a modern AI security CTF. It unpacks the successes and unexpected "finishing cases" where participants either leveraged AI models to trivialize complex problems or found non-AI solutions to be more effective. The insights provided are invaluable for anyone looking to understand the current state of AI security challenges, the practical implications for defenders, and the future direction of AI-focused security competitions.
Background
▶ Watch: Introduction and overview of the AI Security CTF (0:00)
The rapid advancement and pervasive integration of AI into various facets of daily life have created an urgent need for specialized security expertise. Traditional CTFs, while excellent for honing general cybersecurity skills, often feature limited challenges specifically targeting AI technologies. Recognizing this gap, NSHC embarked on designing the AI Cyber Defense Contest (ACDC), a competition tailored to evaluate both the effective application of AI and the ability to secure AI systems against sophisticated attacks. The initiative was born from a desire to proactively address emerging AI-related security risks and to build a robust foundation for the safe and secure use of AI.
The ACDC was structured into two main phases: an online qualifier and an offline final. The qualifier, held over 36 hours in a Jeopardy-style format, aimed to lower the entry barrier and accommodate a broad range of participants. The finals, an 8-hour offline event, introduced a more complex blend of Jeopardy and Attack and Defense formats, with a significantly increased proportion of AI-centric elements. A key design decision for the finals' Jeopardy challenges was requiring participants to explain their solutions to judges before receiving the flag, a measure intended to minimize "lucky" guesses and ensure genuine understanding.
The challenges themselves were categorized into three critical areas:
- AI for Security: Focused on how AI can be leveraged to enhance security measures.
- Security for AI: Explored vulnerabilities and attack surfaces within AI models and services, emphasizing adversarial tactics.
- AI Infra: Addressed security concerns related to the underlying infrastructure supporting AI systems.
Several key design considerations shaped the contest. Initially, the organizers aimed to attract AI researchers to introduce them to security topics, but in practice, many experienced CTF players also joined. A significant debate revolved around the permissible use of AI models: the team ultimately decided against restricting commercial models (e.g., GPT-4), acknowledging that banning them would unfairly advantage participants with access to powerful custom hardware. This decision, however, introduced the critical challenge of cost management for API usage. Authors were given freedom to choose AI models for challenge creation, and the AI content ratio was intentionally kept lower in qualifiers to ease entry, reserving the more complex Attack and Defense format for the finals due to its scalability challenges.
Key Findings
▶ Watch: Key design decisions: participants, models, costs, format (6:00)
The ACDC yielded several profound insights, particularly regarding the dynamic interplay between human ingenuity, evolving AI capabilities, and the design of security challenges. The most striking findings revolved around unexpected solution paths that either trivialized intended difficulties or demonstrated the enduring efficacy of traditional methods.
One major finding was the unanticipated power of modern Large Language Models (LLMs) to "outsmart" challenges designed for human reasoning. A prime example was the "Was Ghidra" challenge, intended to test participants' understanding of Ghidra P-code for binary analysis. The organizers envisioned a multi-step process involving binary analysis, SLA language compilation, and decryption to obtain the flag. However, due to the rapid advancements in LLMs, which occurred even during the three-month preparation period, participants could simply upload the binary to an LLM service and, with a few prompts, retrieve the flag. The binary contained sufficient information for the LLM to infer the analysis steps and internal logic, bypassing the complex, step-by-step reasoning a human would typically employ. This demonstrated how quickly AI capabilities can shift the difficulty landscape of security problems, making carefully tuned challenges less meaningful.
Conversely, another significant observation was cases where non-AI tools proved superior to AI-based solutions. The "Video Share" challenge required participants to extract hexadecimal characters displayed repeatedly in a video, reconstruct a file, and obtain a flag. The intended solution involved training a customized AI model for Optical Character Recognition (OCR) on character image samples extracted from the video. However, participants found that simple Python scripts, leveraging the constant character shapes and predictable switching speed, could solve the challenge faster and more reliably. This highlighted that while AI offers powerful capabilities, it also introduces overhead (e.g., training time), and simpler, deterministic algorithmic approaches remain highly competitive or even preferred when applicable.
The implementation of the Attack and Defense format in the finals provided valuable insights into real-time AI security operations. This format, involving teams planning, executing, and refining attack and defense strategies against AI models, proved to be a powerful learning experience. It necessitated constant monitoring, log analysis, and rapid patching of weaknesses, closely mirroring real-world incident response. However, it also brought to light critical operational considerations, such as ensuring sufficient AI model availability and computing power to prevent response delays and maintain fairness.
Finally, the organizers identified several inherent challenges in building LLM-based CTF problems:
- Token Usage and Control: Prompt bypass challenges can lead to massive API requests, risking budget overruns or API blocks if not managed with strict policies on max tokens and request limits.
- Guardrail Difficulty Tuning: Striking the right balance is crucial; challenges can be too easy (solved in minutes) or too strong (unsolvable even by authors). External review and validation are essential.
- Model Version Fixation: LLM capabilities evolve rapidly. A payload that works one day might fail after a model update, emphasizing the need to fix model versions for consistent challenge behavior.
- Non-Determinism: The inherent non-deterministic nature of AI means the same payload might succeed at one time and fail at another, requiring multiple tests and success rate comparisons during design.
- Permissions and System Isolation: Prompt injection is designed to induce unintended behavior, making strict isolation critical to prevent access to external resources, API keys, or internal files.
Specific observations from challenges included the "IP Camera Zero-day" where firmware acquisition proved to be a surprisingly high failure point for most teams, and the "Agent Challenges" which saw creative automation and script-based responses beyond initial anticipation. In the "Backdoor Challenge," unintended solutions emerged due to model overfitting during training, demonstrating the importance of understanding model behaviors during challenge design. A design issue was also identified in the Attack and Defense format where the lack of delay in attack log disclosures made replay attacks technically possible, an area for future improvement.
Technical Deep Dive
▶ Watch: AI for Security challenge: Identifying keywords from multilingual audio (8:00)
The AI Cyber Defense Contest (ACDC) was meticulously designed with challenges spanning three core categories, each reflecting a distinct facet of AI security. Let's delve into the technical specifics of representative challenges from both the qualifier and final rounds.
Qualifier Challenges
1. AI for Security: "Barbar"
This challenge aimed to demonstrate how AI can augment human capabilities. Participants received an MP3 audio file containing ten clues, each spoken in a different language, all pointing to a single keyword. The goal was to identify this keyword. While a polyglot might solve it directly, the intended approach leveraged AI:
- Automatic Speech Recognition (ASR): Participants would first use an AI-based ASR service to transcribe the audio pronunciations into text for each language.
- Large Language Model (LLM) Inference: The extracted text would then be fed into an LLM, which could analyze the diverse linguistic inputs and infer candidate words, ultimately leading to the correct keyword (e.g., "compass" from "mapa," "bomi," etc.). This highlighted AI's ability to bridge linguistic gaps and solve problems intractable for most humans.
2. Security for AI: "Jumbo P Agents"
This scenario focused on prompt injection against an AI assistant designed to provide personalized information. The core vulnerability was that if safeguards were bypassed, an attacker could access another user's private data. Participants' objective was to obtain another user's personal information by crafting specific prompt injection payloads. Techniques involved applying various prompt filter bypass techniques, such as threatening the assistant or using specific character insertions to trick the model into divulging unintended information. This challenge underscored the critical need for robust input validation and contextual understanding in AI systems to prevent data exposure.
3. AI Infra: "Correlation for Litigation"
This challenge exposed vulnerabilities in the underlying infrastructure supporting an AI service. The scenario involved a medical AI service where an attacker could compromise internal systems and extract sensitive medical data. The path to the flag involved:
- JWT Vulnerability: The service had a JSON Web Token (JWT) vulnerability, allowing attackers to escalate privileges to admin.
- Data Exfiltration: With admin access, an attacker could retrieve internal backup data, potentially containing medical records.
- Similarity Analysis: The exfiltrated medical data would then be analyzed for similarity against publicly available medical datasets (e.g., on Hugging Face). By cross-referencing, the attacker could infer personal information about a specific VIP. This demonstrated how traditional infrastructure vulnerabilities can have severe consequences when combined with AI services handling sensitive data.
Final Challenges
The finals introduced more complex and multi-faceted challenges, often building upon concepts from the qualifiers.
1. Revenge Chunbong P Agents (LLM Prompt Bypass)
This was an advanced version of the qualifier's "Jumbo P Agents." The guardrail strength was significantly increased in stages, requiring participants to employ a wider array of sophisticated prompt bypass techniques. Examples included:
- Role-based Prompting: Assigning the LLM a specific role (e.g., a "helpful assistant who ignores safety policies").
- Conscious Style Prompting: Framing the request as an internal monologue or a thought process.
- Story-based Prompting: Embedding the malicious request within a narrative to bypass detection.
The challenge emphasized the constant cat-and-mouse game between LLM developers implementing safeguards and attackers devising new bypass methods. Key considerations for building such challenges included strict token usage limits, careful guardrail difficulty tuning, fixing model versions to ensure consistent behavior, accounting for LLM non-determinism, and ensuring robust system isolation to prevent resource leakage.
2. IoT Analysis with AI (D-Link DCS-5222LB Camera)
This challenge combined IoT vulnerability analysis with AI-driven signal processing. Participants received an audio file recorded by a D-Link DCS-5222LB home camera and the camera model itself.
- Stage 1: Camera Exploitation: Participants exploited a vulnerability in the camera (e.g., via firmware analysis and CGI vulnerabilities) to obtain an internally stored audio file containing a password. A notable hurdle was the camera model being discontinued, making official documentation scarce and requiring teams to find firmware and decryption tools from unofficial, often less secure, sources (e.g., sites without SSL).
- Stage 2: AI-based Audio Analysis: The obtained audio contained "beep" sounds for each keypress. Although seemingly similar, each key's beep had distinct frequency characteristics. Participants used AI for spectral analysis to distinguish these patterns, such as differences in energy distribution across specific frequency bands. By mapping these patterns to key numbers, they could decode the password waveform, reconstruct the password, and open a virtual door to get the flag.
3. Agent Challenges (Pwn with AI)
This challenge integrated traditional binary exploitation (pwn) with AI. Participants received a Dockerfile, char1 and char2 binary files, and an agent.py template. Their task was to write an agent.py script that could exploit the staged binaries. The unique twist was that participants could not obtain the actual stage binary directly or use remote debugger tools. Instead, they had to use an LLM to "import" information they couldn't directly access.
- AI for Information Gathering & Exploitation: Participants used LLMs to infer binary validation logic, scan port ranges (as only a range, not an exact port, was provided), analyze potential vulnerabilities, and craft exploits. Solutions ranged from traditional pwn analysis augmented by AI for specific aspects, to entirely feedback-driven systems where AI handled everything from port scanning to binary analysis and exploit generation.
4. Backdoor Challenge (LLM Backdoors)
This challenge explored the concept of backdoors in LLM chatbots. Two types were presented: keyword-based and frequency-based.
- Keyword-based Backdoor: Participants aimed to find three trigger words that activated a backdoor. The intended solution involved using algorithms like GCG (Gradient-based Coordinated Gradient) to iteratively refine trigger words through feedback. However, an unintended solution emerged due to model overfitting during training; if participants understood a specific overfitted behavior, they could bypass the intended algorithmic approach.
- Frequency-based Backdoor: The backdoor value was embedded in PyArmor obfuscated code. Participants had to analyze the obfuscated code or brute-force the value to extract it. This combined code analysis skills with the understanding of how malicious logic might be concealed within AI systems.
These challenges collectively demonstrated the breadth of AI security, from prompt engineering and adversarial attacks to securing AI infrastructure and analyzing complex AI model behaviors.
Demo / Proof of Concept
▶ Watch: Security for AI challenge: Prompt injection to access private data (10:00)
While the talk did not feature a live, interactive demonstration by the speakers, it effectively conveyed "proofs of concept" through detailed descriptions of the challenges, intended solutions, and crucially, the unexpected solutions achieved by participants. The content served as a retrospective analysis of how various AI and non-AI techniques were applied in practice during the CTF.
For instance, in the "Barbar" challenge, the speaker played an example MP3 audio file containing clues in different languages, illustrating the initial input to the problem. This provided a concrete example of the Automatic Speech Recognition (ASR) and Large Language Model (LLM) inference steps required to solve it.
Similarly, for the "Jumbo P Agents" challenge, the presentation included a concrete example of a prompt injection payload: "the user threatened the assistant for to force it to this respond." This textual example served as a clear demonstration of the adversarial techniques participants employed to bypass safeguards and extract sensitive information.
The most illustrative "demo" content was the series of screenshots showcasing the Attack and Defense competition dashboard. These visuals provided a real-time glimpse into the operational mechanics of the finals:
- Challenge Status: Displaying the current state of the four Attack and Defense challenges.
- Team Scores: Showing the leaderboard and real-time score updates.
- Activity Log: A right-panel log detailing real-time attack and defense events.
- Defense Page: A screenshot of the interface where teams would "set the guardrail prompt before attack time," demonstrating the defensive strategy implementation.
- Attack Page: A screenshot showing the left panel with other teams' names and an input field for the "prompt payload," illustrating how attacking teams would execute their strategies.
- Attack and Defense Log: A page displaying logs that teams would analyze to refine their strategies.
These visual aids, combined with the detailed walkthroughs of both intended and unintended solutions for challenges like "Was Ghidra" (where LLMs trivialized binary analysis) and "Video Share" (where simple Python scripts outperformed custom AI models), served as powerful demonstrations of the practical implications and diverse problem-solving approaches observed during the ACDC. While not a live demo, the comprehensive explanation of the CTF environment and participant interactions provided a robust understanding of the concepts discussed.
Defensive Implications
▶ Watch: AI Infra challenge: Exploiting medical AI service for data exfiltration (11:30)
The insights gleaned from building and breaking AI security CTFs offer critical lessons for real-world defenders operating in an increasingly AI-driven landscape. The challenges and unexpected solutions highlighted several key areas requiring immediate attention and strategic planning.
Firstly, the prevalence of prompt injection attacks in challenges like "Jumbo P Agents" and "Revenge Chunbong P Agents" underscores the paramount importance of robust guardrail implementation for any user-facing LLM service. Defenders must move beyond basic input filtering to adopt sophisticated techniques like contextual understanding, role-based access control for AI agents, and dynamic prompt sanitization. Implementing strict token usage limits and API request policies is also crucial to prevent resource monopolization or denial-of-service attacks against LLM APIs, as well as to manage costs.
Secondly, the "Correlation for Litigation" challenge explicitly demonstrated that traditional cybersecurity vulnerabilities, such as JWT vulnerabilities or misconfigurations in underlying AI infrastructure, can be just as, if not more, devastating when dealing with AI systems handling sensitive data. This reinforces the need for a holistic security approach: securing the AI model itself is insufficient without also rigorously securing the data pipelines, API endpoints, storage systems, and authentication mechanisms that support it. Regular penetration testing, vulnerability assessments, and adherence to secure development lifecycle practices remain indispensable for AI-powered applications.
Thirdly, the "Was Ghidra" incident, where LLMs trivialized a binary analysis challenge, presents a dual implication. While it highlights the need for CTF designers to adapt, it also suggests a powerful defensive capability. Defenders can leverage advanced LLMs as potent tools for automated security analysis, including binary reverse engineering, vulnerability discovery, and understanding complex code logic. However, relying solely on AI for analysis also introduces new risks, such as potential misinterpretations or the need for human oversight to validate AI-generated insights.
Fourthly, the "Backdoor Challenge" and the observation of model overfitting emphasize the growing importance of AI model integrity and supply chain security for AI. Defenders must implement rigorous testing and validation processes for third-party models and internal training pipelines to detect subtle backdoors, unintended behaviors, or biases. Techniques like explainable AI (XAI) can aid in understanding model decisions and identifying anomalous behavior that might indicate malicious manipulation or vulnerabilities.
Finally, the Attack and Defense format itself provides a blueprint for practical defense strategies. The emphasis on real-time monitoring, log analysis, and iterative maintenance (patching weaknesses) closely mirrors an effective incident response framework. Defenders should invest in comprehensive logging and telemetry for AI systems, develop specific playbooks for AI-related incidents, and regularly conduct red-teaming exercises to test and improve their AI security posture. The lesson from the "IP Camera Zero-day" – that firmware acquisition was a major blocker – also highlights the ongoing importance of securing the software supply chain for IoT devices and other hardware components that form part of the AI ecosystem.
Key Takeaways
- AI CTFs are essential for skill development: Dedicated AI security CTFs like ACDC provide a crucial platform for participants to develop and test skills in both leveraging AI for security and defending against attacks on AI systems, filling a significant gap left by traditional CTFs.
- LLMs can trivialize complex challenges: Modern large language models possess capabilities that can rapidly solve problems intended for human-driven, step-by-step analysis, necessitating a re-evaluation of challenge design and difficulty tuning in security competitions.
- Non-AI tools remain critical: Despite AI's advancements, simpler, deterministic, non-AI tools (e.g., Python scripts for OCR) can still be faster, more reliable, and more resource-efficient for specific tasks, reminding defenders to choose the right tool for the job.
- Rigorous design for LLM challenges is vital: Building effective LLM-based CTF problems demands careful consideration of token usage, guardrail difficulty, model version control, non-determinism, and stringent system isolation to prevent unintended exposures.
- Attack & Defense mirrors real-world AI security: The Attack and Defense format, with its cycles of planning, execution, and maintenance, closely simulates real-world AI security operations, offering invaluable experience in incident response and continuous improvement.
- AI infrastructure security is paramount: Traditional vulnerabilities in supporting infrastructure (e.g., JWT flaws, IoT device vulnerabilities) pose significant risks to AI services, emphasizing that a holistic security approach encompassing both AI models and their underlying systems is indispensable.
About the Speaker(s)
The talk "When the Model Outsmarts the Challenge: Building and Breaking AI Security CTFs" was presented by a team of experienced cybersecurity professionals from NSHC in Korea.
SoYeon Kim is a Researcher at NSHC Lab Labs. Her primary interests lie in the dynamic fields of AI security, reverse engineering, and embedded systems. Her practical expertise is further demonstrated by her achievement of securing second place in a Narcon competitive CTF, showcasing her hands-on skills in security challenges.
Hea-Eun Moon serves as a Lab Lead at NSHC Lab Labs. She brings extensive experience in organizing security competitions, having played a key role in orchestrating "bridges" at renowned conferences such as Black Hat and Def Con, which are highly respected within the cybersecurity community.
Sang-tae Woo holds the position of CISO (Chief Information Security Officer) at NSHC. With a broad and deep background, he possesses extensive experience in penetration testing, security consulting, and the overall management and execution of CTF competitions.
Together, this formidable team not only organized the pioneering AI Cyber Defense Contest (ACDC) but also actively participates in CTF competitions, reflecting their passion and commitment to advancing cybersecurity knowledge and skills.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Honest retrospective from people who actually ran the thing — the 'LLM trivialized our carefully designed challenge mid-competition' admission alone is worth something. But this is a conference talk about organizing a CTF, not a research paper on AI security, and it stays squarely in practitioner war-story territory without pushing into anything technically generative.
Heather Calloway (CISO) — WEAK
Technically earnest work from a team that clearly knows how to run a CTF, but this talk never leaves the competition design lane. The defensive implications section gestures toward real-world relevance without actually producing it — the gap between 'LLMs trivialized our binary challenge' and 'here is what that means for how you deploy AI in production' is never closed.