Taming the Beast: Inside Llama 3 Red Team Process

Grattafiori, Evtimov, Bitton

DEF CON 32 Main Stage · Day 1 · Main Stage

Overview

This talk delves into the intricate and evolving process of red teaming large language models (LLMs), specifically focusing on the methodologies employed for Meta's Llama 3. Presented by Grattafiori, Evtimov, and Bitton, likely members of Meta's AI safety or red teaming initiatives, the session provides a candid look at the challenges and advancements in ensuring the safety and robustness of cutting-edge generative AI. As LLMs like Llama 3 scale to unprecedented sizes—trained on an astounding 15 trillion tokens, a significant leap from Llama 2's 2 trillion—the potential for latent risks, biases, and vulnerabilities embedded within their vast training data grows exponentially.

Watch on YouTube

Visual summary for Taming the Beast: Inside Llama 3 Red Team Process by Grattafiori, Evtimov, Bitton
Visual summary for Taming the Beast: Inside Llama 3 Red Team Process by Grattafiori, Evtimov, Bitton

Key moments

  1. 0:00 Introduction to LLM risks and AI red teaming
  2. 2:38 Defining AI safety challenges with practical examples
  3. 4:10 Understanding dual-use models and lagging risk understanding
  4. 5:15 Llama 2 red teaming: initial manual process and limitations
  5. 6:05 Llama 3 challenges: Llama 2's over-safety and external expertise
  6. 7:15 Scaling red teaming: the 'judge LLM' dilemma

Taming the Beast: Inside Llama 3 Red Team Process

Speakers: Grattafiori, Evtimov, Bitton

Conference: DEF CON 32

YouTube: https://www.youtube.com/watch?v=UQaNjwLhAmo

Overview

This talk delves into the intricate and evolving process of red teaming large language models (LLMs), specifically focusing on the methodologies employed for Meta's Llama 3. Presented by Grattafiori, Evtimov, and Bitton, likely members of Meta's AI safety or red teaming initiatives, the session provides a candid look at the challenges and advancements in ensuring the safety and robustness of cutting-edge generative AI. As LLMs like Llama 3 scale to unprecedented sizes—trained on an astounding 15 trillion tokens, a significant leap from Llama 2's 2 trillion—the potential for latent risks, biases, and vulnerabilities embedded within their vast training data grows exponentially.

The core of the presentation highlights why traditional security red teaming principles are indispensable for generative AI, albeit with necessary adaptations. The speakers underscore the critical need for a proactive, creative, and expert-driven approach to uncover issues that range from subtle biases and factual inaccuracies to the generation of harmful or dangerous content. This talk is crucial for anyone involved in AI development, deployment, or policy, as it illuminates the practical hurdles in making powerful AI systems both useful and safe for widespread adoption, emphasizing that merely building large models isn't enough; rigorously testing their boundaries is paramount.

Background

▶ Watch: Introduction to LLM risks and AI red teaming (0:00)

The rapid ascent of Transformer-based large language models has revolutionized numerous fields, yet their complexity presents a formidable challenge to ensuring safety. As the speakers emphasize, despite their advanced capabilities, LLMs frequently "go off the rails," fail to grasp user intent, or even exhibit outright dishonesty, as evidenced by examples like "Canadian bots being dishonest." This inherent unpredictability stems from their nature as black boxes—systems that are difficult to inspect and whose internal workings are opaque, making it challenging to understand the root causes of undesirable behaviors.

AI safety, or responsible AI, is a multifaceted domain encompassing critical concerns such as fairness, bias, explainability, robustness, and privacy. The sheer scale of models like Llama 3, trained on 15 trillion tokens, means that any biases or vulnerabilities present in this immense dataset can manifest in unpredictable and potentially harmful ways. The speakers illustrate this with a seemingly innocuous example: repeated attempts to prompt a generative AI model to create an image of a "burger without cheese." Despite explicit instructions and multiple self-correction attempts, the model consistently failed, often generating burgers with cheese and confidently asserting the absence of cheese in its output. This highlights the problem of biased data and the brittleness of current generative AI systems, even in straightforward tasks.

The concept of red teaming, traditionally a practice to challenge assumptions in intelligence, defense, and cyber security, has become an increasingly vital requirement for generative AI. It involves an independent, creative process designed to surface risks that internal development teams might overlook. The distinction between traditional red teaming and AI red teaming can be a source of confusion, but the underlying principle remains: trust but verify. Early efforts with Llama 2 revealed a largely manual, pain-staking process involving copying, pasting prompts, and hand-labeling responses, with little automation or a defined taxonomy. Furthermore, Llama 2 was found to be "too safe," meaning its overly cautious responses sometimes hindered its utility, such as refusing to provide information that could be construed as "bomb drinks," even if the intent was benign or hypothetical. This over-alignment underscored the delicate balance between safety and usability. The accelerating pace of AI development also means that risk understanding often lags behind capability, making proactive red teaming indispensable for informing critical trade-off decisions.

Key Findings

▶ Watch: Understanding dual-use models and lagging risk understanding (4:10)

The red teaming efforts for Llama 3 revealed several critical insights into the development and assurance of large language models. A primary finding was the absolute necessity for a radical overhaul of the red teaming process itself. The manual, ad-hoc methods used for Llama 2, characterized by copy-pasting prompts and hand-labeling responses, were utterly unsustainable for a model of Llama 3's scale and complexity (15 trillion tokens). This highlighted that future LLM development must integrate highly automated and structured red teaming methodologies from the outset.

A significant discovery from Llama 2's evaluation was its tendency to be "too safe." While safety is paramount, an overly cautious model can become functionally useless for many legitimate applications. The example of Llama 2 refusing to assist with "bomb drinks" (even as a hypothetical scenario) underscored this point. This finding led to a crucial understanding: the goal of AI safety is not merely to prevent all harmful outputs, but to achieve a balanced alignment that allows for utility while robustly mitigating genuine risks.

The complexity of modern LLMs necessitates a multidisciplinary approach to risk assessment. The red team quickly realized that their internal expertise, particularly in areas like cyber weapons, was insufficient for comprehensively evaluating all potential risks, such as those related to CBRN (Chemical, Biological, Radiological, Nuclear). This led to a key finding: the imperative to involve external subject matter experts in diverse risk categories to ensure thorough coverage.

Furthermore, the sheer volume of data generated during testing—upwards of 10,000 logs—made manual expert labeling across multiple languages and tools an insurmountable task. This drove the finding that automation is critical not just for prompt generation, but for the entire evaluation pipeline. The concept of using an LLM as a judge to evaluate the outputs of other LLMs emerged as a promising, albeit complex, solution, bringing with it the challenge of validating the judge itself. The recurring "burger without cheese" example served as a powerful illustration of model brittleness and the persistent challenge of biased data, demonstrating that even with detailed prompts and self-correction attempts, generative AI can confidently fail simple, explicit instructions.

Technical Deep Dive

▶ Watch: Llama 2 red teaming: initial manual process and limitations (5:15)

The technical foundation of the Llama 3 red teaming process is rooted in addressing the unprecedented scale and complexity of modern LLMs. Llama 2 was trained on 2 trillion tokens, a massive dataset in its own right, but Llama 3 escalated this to an astounding 15 trillion tokens. This exponential increase in training data mandates a proportional leap in the sophistication of safety evaluation techniques. The talk underscored that traditional Transformer architectures, while powerful, inherently possess "black box" characteristics, making direct inspection of their decision-making processes extremely challenging.

The core of AI safety for such models revolves around mitigating risks across several key dimensions: fairness, ensuring equitable treatment across different user groups; bias, identifying and reducing systemic prejudices embedded in training data; explainability, striving to understand why a model produces a particular output; robustness, ensuring consistent and reliable performance against varied inputs, including adversarial ones; and privacy, protecting sensitive user information.

The evolution of the red teaming methodology from Llama 2 to Llama 3 involved a significant architectural shift. For Llama 2, the process was largely manual, characterized by human red teamers copying and pasting prompts into the model and then manually labeling the responses. This was inefficient and lacked scalability, especially for identifying subtle vulnerabilities across a vast attack surface. Recognizing this limitation, the Llama 3 red team prioritized automation. This involved developing tools and frameworks to programmatically generate diverse prompts, interact with the LLM at scale, and capture responses for subsequent analysis.

A key technical innovation discussed was the exploration of using an LLM as a judge. Given the 10,000+ logs requiring expert labeling, the idea was to leverage another LLM to automate the evaluation of Llama 3's outputs against predefined safety policies. This involves feeding the judge LLM the original prompt, Llama 3's response, and the safety policies, then asking it to rate the response's compliance. While promising for scalability, this approach introduces a recursive problem: "how do you evaluate the judge?" This necessitates a robust validation framework for the judge LLM itself, potentially involving a "judge for your judge," creating a hierarchical evaluation system that requires careful design and empirical testing to prevent cascading errors or biases.

The identification of specific risk categories, such as CBRN, highlights the need for specialized prompt engineering and testing strategies. Red teamers with expertise in cyber security might not possess the domain knowledge to craft prompts that effectively probe for vulnerabilities related to chemical weapon recipes or biological agent synthesis. This necessitates collaboration with external domain experts who can inform prompt design and accurately assess the model's responses in these highly sensitive areas. The goal is to uncover scenarios where the model might inadvertently provide harmful information or instructions, even if heavily aligned for safety.

The example of generating a "burger without cheese" illustrates the brittleness of generative AI. Despite clear, detailed prompts, the model repeatedly failed to produce the desired output, often including cheese and then confidently asserting its absence. This points to fundamental limitations in its semantic understanding and image generation capabilities, or deeply ingrained biases within its training data regarding common food items. The success of generating an SVG output for the cheeseless burger suggests that simpler, more structured output formats might bypass some of the complex generative hurdles, highlighting that the "best things in life are simple" for LLMs in certain contexts.

Demo / Proof of Concept

▶ Watch: Llama 3 challenges: Llama 2's over-safety and external expertise (6:05)

While the talk did not feature a live, interactive software demonstration in the traditional "proof of concept" sense often seen in cybersecurity presentations, it effectively illustrated a core problem through a compelling visual demonstration. The speakers repeatedly showed attempts to prompt a generative AI model to create an image of a "burger without cheese." This served as a clear, relatable demonstration of the model's brittleness and bias.

The demonstration highlighted:

  1. Initial Failure: The model's first attempt to generate a "burger without cheese" resulted in an image clearly showing cheese.
  2. Lack of Self-Correction: Subsequent attempts, even with explicit follow-up instructions, often yielded similar results, demonstrating the model's inability to learn from its immediate failures or accurately interpret nuanced negative constraints.
  3. Confident Factual Errors: Crucially, when asked to inspect its own output, the generative AI confidently asserted that there was "no cheese on that burger," even when cheese was visibly present. This showcased a significant disconnect between its generative capabilities, its interpretative abilities, and its internal representation of reality.

The speakers contrasted this with a brief mention of an LLM successfully creating an SVG output of a burger without cheese. This implied that for certain types of structured or simpler tasks, models might perform more reliably, perhaps due to the nature of SVG generation being more akin to programmatic instruction than complex image synthesis. This conceptual "demo" powerfully underscored the core challenges in AI safety: the persistence of bias, the difficulty of achieving true semantic understanding, and the models' overconfidence in their erroneous outputs, even in seemingly simple tasks. It served as a tangible example of why a robust red teaming process is essential to uncover these subtle yet pervasive failures.

Defensive Implications

▶ Watch: Scaling red teaming: the 'judge LLM' dilemma (7:15)

The insights from red teaming Llama 3 offer crucial defensive implications for organizations developing and deploying large language models. The foremost takeaway is the absolute necessity of integrating proactive AI red teaming as a continuous and evolving process, not a one-off audit. Organizations must recognize that the rapid advancement of AI capabilities means that risk understanding will always lag, necessitating an adaptive and creative approach to uncover latent vulnerabilities.

Defenders must move beyond manual, ad-hoc testing methods. The sheer scale of LLMs like Llama 3 (15 trillion tokens) makes manual prompt generation and response labeling unsustainable. Investment in automation frameworks for prompt generation, response collection, and initial policy-based evaluation is critical. This includes developing custom tools or leveraging existing platforms to systematically probe models across a vast array of potential risks.

The talk highlighted the importance of multidisciplinary collaboration. Internal AI safety teams often have expertise in technical aspects but may lack domain knowledge in specific high-risk areas like CBRN. Organizations must actively seek out and integrate external subject matter experts in fields such as chemical engineering, biology, law, ethics, and specific cultural contexts to ensure comprehensive risk assessment. This expands the red team's capacity to craft nuanced prompts and accurately interpret complex outputs.

A key defensive strategy involves understanding and managing model alignment. The finding that Llama 2 was "too safe" indicates that over-alignment can hinder utility. Defenders need to work with developers to strike a delicate balance, ensuring that models are safe without becoming overly restrictive. This requires clear policy definitions and iterative testing to refine alignment parameters, allowing for legitimate use cases while preventing harmful applications.

Furthermore, the persistent issues of model brittleness and biased data, as demonstrated by the "burger without cheese" example, necessitate continuous monitoring and robust data governance. Defenders should implement mechanisms to track model failures, analyze the root causes of biases, and work towards curating cleaner, more representative training datasets. Investing in explainability tools is also vital to gain insights into the "black box" nature of LLMs, helping to diagnose why models behave in unexpected ways and facilitating targeted mitigation strategies.

Finally, while the concept of using an LLM as a judge for automated evaluation is promising, defenders must approach it with caution. Rigorous validation of these "judge" models is paramount to prevent the introduction of new biases or the propagation of errors. This might involve human-in-the-loop validation for a subset of results or developing alternative metrics to assess the judge's performance, ensuring the integrity of the automated evaluation pipeline. Staying abreast of industry best practices, academic research, and the evolving landscape of AI safety institutes is also crucial for adapting defensive strategies to emerging threats.

Key Takeaways

  • AI Red Teaming is Essential and Evolving: The complexity and scale of modern LLMs like Llama 3 (15 trillion tokens) necessitate a rigorous, proactive, and continuously adapting AI red teaming process to uncover latent risks and vulnerabilities.
  • Automation is Non-Negotiable: Manual prompt generation and response labeling, as seen with Llama 2, are unsustainable for current LLMs. Future safety evaluations demand significant investment in automated testing frameworks, including the potential use of LLMs as "judges" for evaluation.
  • Multidisciplinary Expertise is Critical: Comprehensive AI safety requires collaboration with diverse subject matter experts, particularly in high-risk areas like CBRN, as internal teams often lack the broad domain knowledge needed for thorough risk assessment.
  • Balance Safety with Utility: Over-alignment can lead to "too safe" models that lack practical utility, as observed with Llama 2. Striking the right balance between robust safety measures and enabling legitimate model capabilities is a crucial challenge.
  • Models Exhibit Brittleness and Bias: Even with detailed instructions, generative AI models can fail simple tasks, demonstrate persistent biases, and confidently produce erroneous outputs, highlighting fundamental limitations in semantic understanding and the impact of biased training data.
  • Proactive Engagement is Key: Given the rapid pace of AI advancement, organizations must be proactive in their safety efforts, continuously refining methodologies, engaging with external experts, and staying updated on industry best practices to inform trade-off decisions and mitigate lagging risk understanding.

About the Speaker(s)

The speakers, Grattafiori, Evtimov, and Bitton, are involved in the advanced AI safety and red teaming efforts at Meta, as indicated by their work on Llama 2 and Llama 3. Their presentation reflects deep practical experience in evaluating the security and responsible deployment of large language models. Based on the content, they are part of Meta's trust and safety organizations, which encompass a wide array of experts. One speaker explicitly mentioned having expertise in "cyber weapons" but not in chemical or biological risks, suggesting a background in traditional cybersecurity red teaming before transitioning or expanding into AI safety. Their collective work focuses on developing and implementing methodologies to proactively identify and mitigate risks associated with generative AI, ensuring the safety and robustness of models released to the public.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk offers a brutally honest and technically grounded look at the immense challenges and evolving methodologies of red teaming large language models, specifically Meta's Llama 3. The speakers provide critical insights into scaling safety evaluations from Llama 2 to Llama 3's unprecedented token count, emphasizing the shift to automation, the necessity of multidisciplinary expertise, and the delicate balance between safety and model utility. It's a valuable deep dive into the practical realities of AI safety, free of the usual vendor hype.

Heather Calloway (CISO) — STRONG ACCEPT

This session provides a candid and essential look into the practical challenges of red teaming large language models, specifically Meta's Llama 3. It moves beyond theoretical discussions to highlight critical operational shifts required, such as the imperative for automation, the necessity of multidisciplinary expertise, and the delicate balance between AI safety and utility. For any CISO or executive grappling with the implications of generative AI, this talk offers concrete insights into managing the significant business and reputational risks inherent in these powerful systems.

→ Top-rated talks at DEF CON 32 Main Stage

All talks from DEF CON 32 Main Stage