DEF CON 33 3- Red teaming fraud prevention systems with GenAI

Karthik Tadinada (Founder and CEO · Fortify Solutions), Martyn Higson (Farn Dynamics)

DEF CON 33 · Day 1 · Main Stage

Overview

This talk, presented by Karthik Tadinada and Martyn Higson, delves into the escalating threat of payment fraud, specifically how Generative AI (GenAI) is democratizing sophisticated attack techniques and challenging traditional fraud prevention systems. Tadinada and Higson, both veterans in building robust fraud systems for major financial institutions, highlight that the ease and accessibility of GenAI tools are creating an "impending and significant fraud crisis," a sentiment echoed by figures like Sam Altman. The presentation serves as a critical wake-up call for the financial industry, demonstrating through practical examples how GenAI can be leveraged for red teaming to expose vulnerabilities in identity verification, authentication, and transaction monitoring controls.

Watch on YouTube

Visual summary for DEF CON 33 3- Red teaming fraud prevention systems with GenAI by Karthik Tadinada, Martyn Higson
Visual summary for DEF CON 33 3- Red teaming fraud prevention systems with GenAI by Karthik Tadinada, Martyn Higson

Key moments

  1. 0:00 Speakers' introduction and talk overview
  2. 1:45 GenAI's multimodal power for fraud
  3. 3:00 Powerful GenAI capabilities moving to local machines
  4. 4:07 Understanding red teaming and adversarial testing
  5. 5:50 Layered controls in payment fraud prevention
  6. 6:30 Digital onboarding and authentication: key GenAI vulnerabilities
  7. 8:00 Scale of global fraud losses: half a trillion dollars

Red Teaming Fraud Prevention Systems with GenAI

Speakers: Karthik Tadinada (Founder and CEO, Fortify Solutions); Martyn Higson (Farn Dynamics)

Conference: DEF CON

YouTube: https://www.youtube.com/watch?v=EiXbzVlNYro

Overview

This talk, presented by Karthik Tadinada and Martyn Higson, delves into the escalating threat of payment fraud, specifically how Generative AI (GenAI) is democratizing sophisticated attack techniques and challenging traditional fraud prevention systems. Tadinada and Higson, both veterans in building robust fraud systems for major financial institutions, highlight that the ease and accessibility of GenAI tools are creating an "impending and significant fraud crisis," a sentiment echoed by figures like Sam Altman. The presentation serves as a critical wake-up call for the financial industry, demonstrating through practical examples how GenAI can be leveraged for red teaming to expose vulnerabilities in identity verification, authentication, and transaction monitoring controls.

The core message is that while fraud has always existed, GenAI dramatically lowers the barrier to entry for highly convincing deepfakes, voice clones, and document alterations, making it significantly harder for financial institutions to distinguish legitimate customers from sophisticated fraudsters. The speakers argue that current defenses, often reactive and fragmented, are ill-equipped to handle this new wave of attacks. They advocate for a proactive, AI-assisted approach to red teaming, urging banks to rigorously test their layered controls against the advanced capabilities of GenAI, lest they face overwhelming fraud losses and reputational damage.

Background

▶ Watch: Speakers' introduction and talk overview (0:00)

The landscape of fraud prevention is undergoing a seismic shift, primarily driven by the rapid advancements and widespread accessibility of Generative AI. Historically, sophisticated fraud techniques, such as creating convincing fake documents or masks, required specialized skills and resources. However, as Karthik Tadinada and Martyn Higson explain, GenAI has transformed this, making complex manipulations "super easy" and available to anyone with a text prompt. The speakers clarify that GenAI is far more than just text-in/text-out models like ChatGPT; it encompasses multimodal models capable of processing and generating video and images, effectively functioning as "Photoshop via text." Crucially, the AI boom is bringing powerful compute capabilities to local machines, meaning attackers are no longer reliant on cloud services with built-in guardrails, enabling them to run potent generation models on standard laptops.

The concept of red teaming—adversarial testing to uncover system failures before malicious actors do—is central to the talk. This can involve black box testing, where the attacker knows nothing about the system's internals, or white box testing, where they have full visibility into rules and logic. While image and video manipulation are amenable to black box testing, transaction monitoring systems often require white box approaches due to their vast surface area.

Payment fraud prevention systems traditionally rely on layered controls: account onboarding controls (to keep bad actors out), money movement controls (to detect unusual transaction patterns), and authentication controls (to verify identity during sensitive actions). The speakers emphasize that account onboarding and authentication are particularly vulnerable. This vulnerability stems from two key trends: the accelerating shift towards digital-first banking post-COVID, where customers expect instant, smartphone-based services without visiting physical branches, and the increasing reliance on uploaded images and video for verification, all of which are now highly susceptible to GenAI manipulation.

The financial stakes are staggering. Global fraud losses reached half a trillion dollars last year. To put this in perspective, UK banks alone spend approximately $50 billion annually on fraud and Anti-Money Laundering (AML) defenses, a figure comparable to the UK's entire defense budget. Despite this investment, Sam Altman, CEO of OpenAI, has expressed significant concern about an "impending and significant fraud crisis" due to GenAI. The US has seen fraud volumes increase fourfold in four years, with a reported 20x increase in deep fraud specifically. This alarming trend underscores the urgent need for a paradigm shift in how financial institutions approach fraud prevention.

Key Findings

▶ Watch: Powerful GenAI capabilities moving to local machines (3:00)

The talk reveals several critical findings regarding the impact of Generative AI on fraud prevention:

  • Democratization of Sophisticated Fraud: GenAI has made advanced image, video, and voice manipulation techniques incredibly easy and accessible. Tools that once required specialized skills or significant resources can now be operated with simple text prompts or readily available open-source code, enabling a broader range of malicious actors to engage in high-fidelity fraud.
  • Detection Lags Generation: Current detection models designed to identify deepfakes and AI-generated content are consistently "several steps behind" the rapid advancements in generation models. Speakers cited examples where academic detection models achieved an AUC (Area Under the Receiver Operating Characteristic Curve) of only 0.58 when generalized, indicating performance barely better than a coin toss. Even custom-trained models with 90% accuracy are deemed insufficient, as fraudsters can generate an unlimited number of attempts at near-zero cost, ensuring a sufficient number will bypass defenses.
  • Vulnerability of Digital Verification: The increasing reliance on remote, digital identity verification and liveness checks (using uploaded documents, photos, and real-time video/voice) renders these processes highly susceptible to GenAI-powered attacks. The talk demonstrated how utility bills, ID documents, and even live video streams can be convincingly faked or altered.
  • Voice Recognition is "Cracked": As Sam Altman reportedly stated, voice recognition as a sole authentication factor is effectively "cracked." The ease with which realistic voice clones can be generated from minimal audio samples (e.g., 7 minutes) makes it a trivial target for fraudsters.
  • Human Element as the Softest Target: Despite advancements in cryptographic security, the "softest target in any system is now the human." Contact center workers and customer service representatives, often under pressure to resolve issues quickly, are increasingly vulnerable to convincing GenAI-generated fakes that exploit human trust and cognitive biases.
  • Complexity of Money Movement Controls: Money movement controls in banks are characterized by their immense complexity, diverse legitimate use cases, and often reactive, "patchwork" nature. This makes them extremely difficult to secure comprehensively, leading to hidden gaps that fraudsters can exploit scalably.
  • LLMs for Proactive Vulnerability Discovery: Large Language Models (LLMs) demonstrate significant potential in white box testing scenarios, capable of analyzing payment system rules (even in proprietary pseudo-code) to identify logical flaws, edge cases, and potential exploitation vectors that human analysts might miss. They can also generate diverse fraud scenarios to test the robustness of existing rule sets.

Technical Deep Dive

▶ Watch: Understanding red teaming and adversarial testing (4:07)

The technical deep dive of the talk showcased practical applications of Generative AI for red teaming, focusing on document verification, liveness checks, and the analysis of money movement controls.

Document Verification Exploitation

For document verification, a common initial step in bank onboarding, the speakers demonstrated how easily digital documents can be manipulated.

  • Utility Bill Alteration: Using ChatGPT, Martyn Higson showed how a text prompt could instruct the AI to re-address a publicly available utility bill to a "Mr. Fake Person." Even when using a real name and address, ChatGPT produced a "pretty much genuine looking bill." The only minor "tells" were subtle visual anomalies, such as the "bill FAQs" text appearing blue instead of black, a detail easily missed by a human reviewer or an unsophisticated automated system. This highlights how GenAI can bypass basic document integrity checks.
  • ID Document Face Swapping: To manipulate photo IDs, they used an open-source GitHub repository called Deep Live Cam. This tool allowed them to take a sample UK driving license (from the government website) and swap the original photo with a picture of Elon Musk. While the initial swap was "not really fantastic," using an upscaler improved the resolution. When tested against an academic detection model (designed to catch AI-generated content for university studies), the modified license was deemed to have only a "0.22% chance of being AI," effectively passing as legitimate. This illustrates the significant gap between generation and detection capabilities.

Challenges in Deepfake Detection

The speakers delved into the limitations of detection models. They noted that these models are inherently "several steps behind" the rapid evolution of generation models. Many academic models, despite high accuracy on their training sets, show plummeting performance when applied to different benchmarks, often achieving an AUC around 0.58 – barely better than random chance for detecting deepfakes.

  • Capsule Model Training: To improve detection, they took an available model called Capsule and trained it on a dataset that included their own generated deepfakes. This boosted accuracy to about 91% on their specific training set. However, even this retrained model failed to detect a deepfake of Martin's face mapped onto Elon Musk, classifying it as "genuine" with 30% confidence. Karthik emphasized that even 90% accuracy is insufficient; fraudsters, with near-zero costs for generation, will simply try 100 images to get 10 successful ones, tilting the economics heavily in their favor.

Liveness Verification and Voice Cloning

The talk then moved to more dynamic forms of verification: liveness verifications and voice recognition.

  • Voice Cloning with ElevenLabs: Martin demonstrated the ease of voice cloning using ElevenLabs, a commercially available API. Despite the service suggesting 2 hours of audio for optimal results, Martin achieved a convincing clone with just a 7-minute sample of his voice. He played two audio clips, one genuine and one AI-generated, and the audience was split on identifying the real voice. This practical demonstration underscored Sam Altman's assertion that "voice is already cracked," posing a significant threat to banks relying on voice biometrics, especially given how infrequently customers interact with their banks via phone.
  • Real-time Face Swapping with Deep Live Cam: A live demo of Deep Live Cam further showcased its capabilities. Martin mapped his face onto a picture of the British Prime Minister in real-time. He highlighted that this specific model, while several years old and producing a 128-pixel replica, is "super accessible." Crucially, the model dynamically transforms the speaker's face, mimicking subtle physiological changes like skin expansion during facial movements, which could fool traditional liveness detection techniques that monitor blood flow or skin characteristics. The performance was 4 FPS on a Mac but could reach 25-30 FPS on a Windows machine with a GPU, allowing it to be piped through a virtual camera to appear as anyone on Google Meet, Zoom, or even emulated Android/iPhone apps for mobile banking.

Analyzing Money Movement Controls

The final technical segment focused on money movement controls, which are the backstop behind identity verification. Karthik likened these to firewall controls, noting their vast complexity due to diverse legitimate use cases and multiple payment methods (e.g., Apple Pay, Google Pay, online, chip & PIN, magstripe, international transactions). Banks struggle to write comprehensive controls across this enormous surface area.

  • LLMs for Rule Exploitation: The speakers proposed using LLMs for white box testing these controls. They demonstrated feeding a pseudo-code representation of a rule (e.g., "$5,000 in 24 hours") into an offline Deepseek model (or Llama for more powerful versions). The LLM successfully identified potential exploits, such as the "midnight barrier" problem where a daily limit rule might not catch transactions spanning two days, or issues related to time zones, suggesting a rolling window approach as a more robust solution.
  • LLMs for Fraud Scenario Generation: To test for generalized vulnerabilities, they provided the LLM with a specific fraud example: two physical transactions (grocery in Manchester, electronics in London) occurring within 2 hours, implying an "unreasonable distance" traveled (200 miles). Fraud teams often react by blocking specific high-risk locations or merchant categories (MCCs). However, the LLM, using RAG (Retrieval Augmented Generation) and more complex models, could identify the underlying pattern (unreasonable travel distance) and generate diverse variations: transactions with lower values, different geographic pairs (e.g., Edinburgh and another city), or international transactions. This helps ensure defenses are not just patching specific attacks but addressing the broader fraud attack vector.
  • Addressing Patchwork Rule Sets: The speakers noted that most fraud systems have grown organically, resulting in "patchwork" rule sets of hundreds of reactive rules. LLMs, especially larger models, can analyze entire sets of analytics, identifying systemic gaps and areas not covered within the vast space of payment activities, moving beyond individual rule analysis.

Demo / Proof of Concept

▶ Watch: Digital onboarding and authentication: key GenAI vulnerabilities (6:30)

The talk featured several compelling demonstrations and proofs of concept that underscored the immediate and accessible threat posed by Generative AI to financial fraud prevention systems. These were not theoretical discussions but practical, live (or pre-recorded due to setup constraints) examples using readily available tools:

  1. AI-Generated Utility Bill: Using ChatGPT, Martyn Higson demonstrated how a simple text prompt could alter the recipient details on a publicly available utility bill. This PoC highlighted that basic document verification, relying on visual inspection or simple automated checks, can be easily bypassed by GenAI. The only "tells" were minute visual inconsistencies like a color bleed in the text, which would likely go unnoticed by a busy human or a less sophisticated detection system.
  2. ID Document Face Swap: This demo utilized Deep Live Cam, an open-source GitHub repository, to swap the face on a sample UK driving license with that of Elon Musk. The process involved an upscaler to enhance resolution, creating a surprisingly convincing fake. The subsequent failure of an academic detection model to identify this as AI-generated (reporting a mere 0.22% chance) served as a stark proof of concept for the inadequacy of current deepfake detection technologies.
  3. Voice Cloning for Authentication: A powerful demonstration involved ElevenLabs, a commercial API, to clone Martin Higson's voice. With just a 7-minute audio sample, the AI generated a synthetic voice so realistic that the audience was evenly split on identifying which of two played samples (one real, one fake) belonged to the speaker. This concretely proved that voice recognition as a standalone authentication method is critically vulnerable.
  4. Real-time Face Swapping for Liveness Checks: In a live demonstration, Martin used Deep Live Cam to map his face onto an image of the British Prime Minister. This PoC illustrated how attackers could potentially bypass liveness verifications by using real-time face transformation. The model, though several years old, could dynamically adjust the mapped face to mimic actual facial movements, including subtle changes in skin appearance that sophisticated liveness detectors might look for. The speakers emphasized that on Windows machines with a GPU, this could run at 25-30 frames per second, allowing it to be piped into a virtual camera for use in video calls or mobile banking app emulators.
  5. LLM for Rule Exploitation Analysis: The speakers showcased an offline Deepseek model (a type of LLM) analyzing a pseudo-code representation of a bank's money movement rule (e.g., "$5,000 in 24 hours"). The PoC demonstrated the LLM's ability to identify logical flaws like the "midnight barrier" vulnerability and suggest more robust rule designs, such as using a rolling window. This proved the utility of GenAI in white box testing for proactive vulnerability discovery.
  6. LLM for Fraud Scenario Generation: Finally, the LLM was tasked with generating variations of a known fraud attack vector (e.g., rapid travel between distant cities for physical transactions). The PoC showed the LLM creating new scenarios with varied transaction values, different geographical locations, and even international elements, proving its capability to generate diverse test data that helps validate the generalized robustness of fraud controls, moving beyond reactive, specific patches.

These demos collectively illustrated that GenAI tools, even those that are publicly available or several years old, are sophisticated enough to circumvent many existing fraud prevention measures, making them potent weapons for red teams and malicious actors alike.

Defensive Implications

▶ Watch: Scale of global fraud losses: half a trillion dollars (8:00)

The insights from this DEF CON talk present profound defensive implications for financial institutions and organizations relying on digital identity verification and fraud prevention. The core message is clear: the era of relying on the superficial authenticity of digital media is over.

  1. Do Not Rely Solely on Image/Video/Voice Authenticity: The most critical implication is that the authenticity of images, videos, or voice recordings, when used as sole or primary verification factors, can no longer be trusted. Generative AI makes it "super simple" and "super accessible" to create highly convincing fakes. Defenders must immediately re-evaluate any control that depends on visual or auditory cues as definitive proof of identity or legitimacy.
  2. Reinforce Layered Security with Diverse Data: The talk implicitly reinforces the fundamental security principle of layered security. Instead of relying on a single, vulnerable verification method (like an uploaded ID photo), institutions must integrate multiple, independent data points. As Karthik noted, combining images with location data, behavioral biometrics, device intelligence, and transaction history makes attacks significantly harder. A multi-factor approach that correlates different types of evidence is essential.
  3. Invest in Proactive, AI-Assisted Red Teaming: Reactive defenses are insufficient against GenAI. Banks must proactively invest in red teaming their fraud prevention systems, specifically using GenAI tools to simulate advanced attacks. This involves:
  • Automated Rule Analysis: Employing LLMs to analyze existing money movement controls and payment system rules (even those written in proprietary languages like Actimize or FICO). LLMs can identify logical flaws, edge cases (e.g., the "midnight barrier" for daily limits), and inconsistencies that human analysts might miss, allowing for preemptive patching.
  • Dynamic Scenario Generation: Utilizing LLMs to generate diverse and novel fraud scenarios based on known attack vectors. This moves beyond static test cases, enabling institutions to test the generalized robustness of their controls against variations in transaction values, geographies, and patterns, rather than just patching specific, previously exploited vulnerabilities.
  1. Address the Human Element Vulnerability: The "softest target" is increasingly the human. Contact center agents and customer service representatives are under pressure and can be susceptible to highly convincing GenAI-generated deepfakes or voice clones. Defensive strategies must include:
  • Enhanced Training: Comprehensive training for frontline staff on recognizing GenAI-generated fraud attempts, understanding their sophistication, and following strict verification protocols.
  • Technological Support: Implementing tools that provide real-time assistance to agents, flagging suspicious calls or video feeds, and offering additional verification steps that are harder for GenAI to bypass (e.g., asking specific, unscripted questions that require deep personal knowledge).
  1. Balance Security with Customer Experience: While strengthening controls is paramount, defenders must remain mindful of customer friction. Overly complex or intrusive authentication processes can lead to customer churn or, worse, exclude non-technical individuals from essential banking services. The challenge is to implement robust, GenAI-resistant controls that are as seamless and user-friendly as possible, potentially leveraging passive biometrics or contextual authentication where appropriate.
  2. Continuous Monitoring and Adaptation: The pace of GenAI development is extremely fast. Detection capabilities will always lag behind generation. Therefore, continuous monitoring of emerging GenAI techniques, rapid adaptation of defensive strategies, and a commitment to ongoing research and development in fraud prevention are non-negotiable. Banks cannot afford to build static defenses; they must cultivate agile and adaptable security postures.
  3. Consider Hardware-Based Verification: While not explicitly discussed as a solution, the inherent vulnerability of software-based, remote verification might necessitate a re-evaluation of hardware-based security tokens or other physical authentication methods for high-value or high-risk transactions, where the "physical presence" of a unique device can add a layer of trust.

In essence, the age of GenAI demands a shift from reactive, signature-based fraud detection to proactive, intelligence-driven red teaming and a robust, multi-layered security architecture that anticipates and mitigates the capabilities of advanced AI-powered attacks.

Key Takeaways

  • GenAI democratizes sophisticated fraud: Tools like ChatGPT, Deep Live Cam, and ElevenLabs make advanced deepfakes, voice cloning, and document alteration accessible to anyone, significantly lowering the barrier for fraudsters.
  • Traditional digital verification is compromised: Relying solely on the authenticity of uploaded images, videos, or voice recordings for identity verification and liveness checks is no longer secure due to the ease of GenAI manipulation.
  • Detection lags generation: AI models designed to detect deepfakes are consistently "several steps behind" the rapid advancements in generation models, often showing poor generalized accuracy, making it easy for fraudsters to bypass them with repeated attempts.
  • LLMs are powerful red teaming tools: Large Language Models can be effectively used for white box testing payment systems, analyzing rules for logical flaws, identifying edge cases (e.g., "midnight barrier"), and generating diverse fraud scenarios to ensure robust and generalized defenses.
  • Fraud systems need proactive, systemic testing: Many existing fraud prevention systems are reactive and "patchwork." GenAI-assisted red teaming is crucial to proactively identify systemic gaps across entire rule sets before they are exploited.
  • The human element is increasingly vulnerable: Despite technological advancements, contact center workers and customer service representatives are becoming the "softest target," susceptible to convincing GenAI-generated fakes that exploit human trust and pressure.

About the Speaker(s)

Karthik Tadinada is the Founder and CEO of Fortify Solutions. Hailing from the UK, he possesses extensive experience in the payments industry. He and Martyn Higson previously collaborated on building fraud systems for a wide array of significant clients, including systems that process virtually all money movement in the UK and the fraud prevention system for the Australian debit card network (FPOS), as well as substantial work for World Pay.

Martyn Higson is associated with Farn Dynamics, where his work focuses on synthetic data for testing fraud controls, which directly informed the topic of this presentation. He previously worked alongside Karthik Tadinada, contributing to the development and deployment of critical fraud prevention systems within the payments sector.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Competent applied-security talk that walks through real GenAI tooling against fraud controls — docs, liveness, voice, and LLM-assisted rule analysis — with working demos. The work is honest and the threat framing is grounded, but the findings are incremental rather than novel: deepfakes beat detectors, voice is cracked, LLMs find rule edge cases. Practitioners who haven't done this themselves will leave with a useful playbook; anyone who has will leave nodding.

Heather Calloway (CISO) — SOLID

A technically grounded and practically motivated talk with real demos and a clear threat framing — but it stays at the problem layer too long and never fully converts into institutional guidance. The right audience is fraud engineers and fraud program managers, not CISOs or board-level decision-makers, and that ceiling limits its broader value.

→ Top-rated talks at DEF CON 33

All talks from DEF CON 33