Follow the data to learn the secret

Dylan Ayrey (CEO and co-founder · Truffle Security)

BSidesSF 2026 · Day 1 · AMC Theatre 14

Overview

In this compelling talk, Dylan Ayrey, CEO and co-founder of Truffle Security, unveils a staggering problem: the pervasive leakage of sensitive data, including hundreds of thousands of live API keys, passwords, and personal information, across the vast landscape of open-source AI datasets hosted on HuggingFace. While his company, built on the popular open-source tool Truffle Hog, traditionally focuses on finding secrets in conventional repositories, Ayrey's research reveals that the burgeoning world of AI data aggregation has inadvertently become a colossal reservoir for exposed credentials and private information, often with severe legal, privacy, and security ramifications.

Watch on YouTube

Key moments

  1. 0:00 Introduction and Hugging Face's origin story
  2. 2:50 Hugging Face's pivot to open-source AI infrastructure
  3. 3:28 Introduction to 'The Stack' massive code dataset
  4. 4:50 Visualizing 'The Stack' as 288 miles of books
  5. 6:10 Truffle Hog finds 100,000 live secrets on Hugging Face
  6. 6:40 Hugging Face: The ultimate repository for scraped data
  7. 7:10 Acknowledging the impossible scale of scanning all data

Follow the data to learn the secret

Speakers: Dylan Ayrey, CEO & Co-founder, Truffle Security

Conference: BSides SF

YouTube: https://www.youtube.com/watch?v=NWfeSvUQPfc

Overview

In this compelling talk, Dylan Ayrey, CEO and co-founder of Truffle Security, unveils a staggering problem: the pervasive leakage of sensitive data, including hundreds of thousands of live API keys, passwords, and personal information, across the vast landscape of open-source AI datasets hosted on HuggingFace. While his company, built on the popular open-source tool Truffle Hog, traditionally focuses on finding secrets in conventional repositories, Ayrey's research reveals that the burgeoning world of AI data aggregation has inadvertently become a colossal reservoir for exposed credentials and private information, often with severe legal, privacy, and security ramifications.

Ayrey’s presentation is not merely a technical exposition; it is a stark narrative illustrating the immense scale and complexity of data in the AI era. Through a series of anonymized yet detailed "tragedies," he demonstrates how seemingly innocuous data collection and curation practices can lead to catastrophic disclosures. The talk underscores a fundamental, unresolved tension between the desire for open AI models and the practical impossibilities of securing, anonymizing, and legally clearing the gargantuan datasets required to train them, challenging the very definition of "open source" in the context of AI.

This article delves into Ayrey's findings, exploring the genesis of HuggingFace, the technical mechanisms behind these data leaks, and the profound implications for security professionals, data scientists, and the future of open-source artificial intelligence. It highlights that the problem is not just about isolated incidents but a systemic issue, exacerbated by the sheer volume of data, the methods of its collection, and a vicious feedback loop within the AI ecosystem itself.

Background

▶ Watch: Introduction and Hugging Face's origin story (0:00)

The journey into the heart of AI data security begins with HuggingFace, a company that started in 2016 with an ambitious goal: to build an AI chatbot. This was a year before Google’s groundbreaking "Attention is All You Need" white paper in 2017, which introduced the Transformer architecture. The real turning point came in 2018 with Google's release of BERT, arguably the world's second open-source Large Language Model (LLM) after OpenAI's lesser-known GPT-1. BERT's two-stage learning process—unsupervised data ingestion followed by supervised fine-tuning—captured HuggingFace's attention.

Recognizing the paradigm shift, HuggingFace transformed its mission. They open-sourced their PyTorch pre-trained BERT version 0.1.1, which eventually evolved into the industry-standard Transformers library. This marked a pivotal transition for the company, moving from a single chatbot focus to becoming the foundational open-source AI infrastructure provider. A few years later, they launched the HuggingFace Hub, a platform that now hosts an astonishing array of open-source models, the code necessary to run them, and, critically, vast amounts of training data.

The scale of data on HuggingFace is almost incomprehensible. Ayrey illustrates this with The Stack, a single dataset comprising three terabytes of permissively licensed code scraped from GitHub. To visualize this, Ayrey humorously describes printing The Stack as a physical book, revealing it would require over 14 million volumes, stretching 288 miles if stacked—and its second version is ten times larger. The Stack is merely one of 883,000 datasets on HuggingFace, and it's two orders of magnitude smaller than some of the largest, such as Common Crawl, which contains 363 terabytes from a scrape of the entire internet. This immense aggregation of human-generated data, intended to fuel AI innovation, has inadvertently become a magnet for exposed secrets.

The problem's gravity is further highlighted by the internal debate within the Open Source Initiative (OSI), the nonprofit that coined "open source" in the 90s. The OSI faced an existential crisis over whether the definition of "open source AI" should mandate the openness of the training data itself. This question, fueled by the European Privacy Act's special carve-outs for open-source AI, revealed deep divisions. The resistance to mandating open data stemmed from the practical impossibility of ensuring data security, privacy, and copyright compliance across such gargantuan, diverse, and often "toxic" datasets.

Key Findings

▶ Watch: Introduction to 'The Stack' massive code dataset (3:28)

Dylan Ayrey's research, powered by Truffle Hog, uncovered a staggering landscape of exposed secrets within the HuggingFace Hub. Despite only having scanned 25% of the platform at the time of the talk, his team found nearly 100,000 unique live secrets and API keys. This number represents individual "tragedies" rather than mere statistics, as Ayrey emphasizes, given the impossibility of recounting each story.

The core findings highlight several critical issues:

  1. Massive Scale of Exposure: HuggingFace has become a centralized repository for data scraped from virtually every public source imaginable—GitHub, npm, S3, Postman, public websites, and even non-traditional sources like Telegram. These diverse data sources, often containing secrets, are aggregated into massive datasets, making HuggingFace "the big one" for secret exposure.
  2. Unintentional Data Leaks: Many leaks are not malicious but arise from developers and data scientists unknowingly including sensitive information in their public datasets. Examples include:
  • Nanix Employee Incident: A personal GitHub access token with repository scope, used for a massive SEMGRAB scan, inadvertently exposed 1,800 private Nanix repositories and all their hard-coded secrets (including highly sensitive ones) in clear text within Apache 2 licensed data.
  • Comfy UI PNG Metadata: The Comfy UI project, designed for flexible Stable Diffusion image generation, defaulted to embedding entire workflows, prompts, and crucially, third-party API keys within the PNG metadata of generated images. Users distributing these images unknowingly shared their keys.
  • JetBrains Intern Incident: An intern creating a dataset of GitHub patch data used 20 accounts and API keys for scraping. One of these was a personal key with access to private repositories. The .git directory, containing the cloning token, was zipped up and published, exposing the private access key to thousands of downloads.
  1. Criminal Exploitation: Ayrey revealed instances of active criminal enterprises leveraging public channels for illicit activities. One striking example involved a criminal network using public Telegram channels to reward individuals for submitting stolen Stripe API keys. These 70 keys were then used to test stolen credit card numbers, powering a vast fraud operation.
  2. Copyright and Reproducibility Challenges: The aggregation of data, particularly from sources like Common Crawl (a scrape of the entire internet including copyrighted material like The New York Times), creates significant legal liabilities. Copyright holders like The New York Times are issuing takedown requests, causing datasets to "evaporate" over time. This makes AI models trained on these datasets non-reproducible, as the original training data changes or disappears. Workarounds, like LAION 5B's dataset of 5 billion external image links, only shift the problem, as links rot and datasets atrophy.
  3. Data Curation Bias and Feedback Loops: A subtle but significant finding is the "data curation bias." When data scientists curate code datasets by executing code in a sandbox and including only code that runs without exceptions, they inadvertently favor code with hardcoded API keys. If a file contains a placeholder like "API key goes here," it throws an exception and is excluded. If a key is hardcoded, the code runs, and the key is included. This creates a positive feedback loop: first-generation LLMs learn from this flawed code, generate more similarly flawed code, which then gets sucked into HuggingFace to train second-generation LLMs, perpetuating the problem.
  4. "Orphan Secrets" as Indicators: By filtering out secrets found in multiple datasets (likely scraped from public, widely copied sources), Ayrey's team focused on "orphan secrets"—those appearing in only one dataset. This methodology successfully identified instances where private data, including credit card numbers and social security numbers, was accidentally packaged up and included in public datasets, indicating true private data leaks rather than just public scrapes.

The overarching conclusion is that "toxic data"—data laden with legal, privacy, and security issues—is currently indispensable for creating powerful LLMs. This paradox creates an almost insurmountable challenge for the open-source AI community.

Technical Deep Dive

▶ Watch: Visualizing 'The Stack' as 288 miles of books (4:50)

The technical underpinnings of these data exposures are varied, often stemming from a combination of human error, tool defaults, and the inherent nature of data aggregation at scale.

At the heart of the detection effort is Truffle Hog, an open-source tool designed to scan for secrets. Its application to the HuggingFace Hub involved scanning millions of individual files across petabytes of data, demonstrating its capability to identify a wide range of sensitive information, from API keys to personal identifiable information (PII).

Several specific technical scenarios illustrate the mechanisms of these leaks:

  1. SEMGRAB and Repository Scans: The Nanix incident highlights a critical vulnerability in how security scanning tools can inadvertently expose secrets. SEMGRAB is a static analysis tool used to find vulnerabilities and misconfigurations in code. When a Nanix employee used a GitHub Personal Access Token (PAT) with a repository scope to perform a massive SEMGRAB scan, this token granted access to all repositories the user had access to, including 1,800 private Nanix repositories. Crucially, SEMGRAB, when configured to identify secrets, outputs these findings in clear text. By zipping up these scan results and publishing them with an Apache 2 license on HuggingFace, the employee effectively made all internal code and hard-coded secrets publicly available. Even after takedown, copies persist due to the nature of public data distribution.
  1. Comfy UI and PNG Metadata: The Comfy UI project aims to provide flexible workflows for Stable Diffusion image generation. To ensure reproducibility, a key feature was the ability to embed the entire workflow, including prompts and generation parameters, directly into the PNG metadata of the output image. This embedded information was stored as a large JSON blob. The critical technical detail here is that this embedding was the default behavior, not an optional one that users had to explicitly enable. If a user's workflow involved calling a third-party API (e.g., for specific image enhancements or AI services), the API key required for that call would also be embedded within this JSON blob in the PNG metadata. When these PNGs were shared or published, the API keys became publicly accessible, a subtle but dangerous default for a tool focused on reproducibility.
  1. GitHub Cloning and .git Directories: The JetBrains intern incident demonstrates a common oversight in handling source code repositories. When a Git repository is cloned, a hidden .git directory is created at the root. This directory contains all the repository's metadata, including configuration files (e.g., .git/config) that can store credentials, such as Personal Access Tokens, used for authentication during the cloning process. If a user zips up an entire cloned repository, including this .git directory, and publishes it, any tokens stored within that configuration are exposed. In this specific case, while the intern tried to avoid scraping private repositories, one of the 20 tokens used for scraping belonged to their personal account, which did have access to private data. Publishing this token in thousands of cloned repositories meant a private access key was now publicly available.
  1. Data Curation Bias: The "vicious feedback loop" for Twitter API keys highlights how technical decisions in data curation can exacerbate secret leakage. The process involves sandboxing code and discarding samples that throw exceptions. If a code snippet declares API_KEY = "your_api_key_here", it might be seen as incomplete and throw an error, leading to its exclusion. However, if a key is hardcoded directly (e.g., API_KEY = "sk-..."), the code might execute successfully, leading to its inclusion in the dataset. This subtle bias preferentially includes code with hardcoded secrets, contributing to their prevalence. Furthermore, LLMs trained on this biased data then generate more code with hardcoded secrets, creating a self-reinforcing cycle of exposure.
  1. Telegram Scrapes and Criminal Activity: The discovery of Stripe API keys in public Telegram channels demonstrates the breadth of data sources being scraped for AI training. Tom, the data scientist, scraped 10 million messages from public Telegram channels. These channels, discoverable via the Telegram app's search function, contained discussions from a criminal enterprise. This enterprise incentivized users to post stolen Stripe keys, which were then used to test the validity of stolen credit card numbers. The technical aspect here is the scraping of unstructured, real-time communication data, which can inadvertently capture highly sensitive and actively exploited credentials.

These examples collectively illustrate that the problem is deeply embedded in the technical practices of data acquisition, processing, and sharing within the AI community.

Demo / Proof of Concept

▶ Watch: Hugging Face: The ultimate repository for scraped data (6:40)

While the talk did not feature a live, interactive demonstration of exploiting a vulnerability, Dylan Ayrey's presentation itself served as a powerful proof of concept for the pervasive nature of secrets within AI datasets. The "demo" was primarily conceptual and illustrative, emphasizing the sheer scale of the data and the concrete findings from Truffle Hog.

Ayrey's most striking illustrative "demo" was the physical manifestation of The Stack dataset. He explained how he literally published a book on Amazon as "Volume 1" of The Stack, demonstrating its physical thickness. He then extrapolated that the entire dataset would require 14,290,866 volumes, spanning an astonishing 288 miles of shelf space. This conceptual demonstration served to viscerally communicate the unimaginable size of the data being discussed, making it clear why traditional security approaches struggle.

The core of the talk, however, was a series of meticulously researched "stories" detailing specific instances of secret exposure. These stories—the Nanix SEMGRAB leak, the Comfy UI PNG metadata embedding, the JetBrains intern's .git directory exposure, the criminal use of Stripe keys in Telegram, and the widespread Twitter API key proliferation—are the true proofs of concept. Each narrative, backed by the identification of live, unique secrets by Truffle Hog, demonstrated that these vulnerabilities are not hypothetical but active and impactful. The speaker’s ability to find and describe these specific, actionable exposures across diverse technical vectors effectively proved the existence and severity of the problem.

Defensive Implications

▶ Watch: Acknowledging the impossible scale of scanning all data (7:10)

The findings presented by Dylan Ayrey necessitate a fundamental shift in how organizations and individuals approach data security in the era of large-scale AI training. The traditional perimeter-based security models are insufficient when petabytes of data are being aggregated, often by third parties, and then made publicly available.

Here are key defensive implications:

  1. Strict Data Sanitization Before Publishing: Any data intended for public AI training datasets, regardless of its source, must undergo rigorous and automated sanitization for secrets, PII (e.g., social security numbers, credit card numbers), and proprietary information. This includes not just explicit credential patterns but also contextual analysis to identify inadvertently included sensitive files or metadata. Tools like Truffle Hog should be integrated into CI/CD pipelines for data preparation.
  2. Understand Tool Defaults and Metadata: Developers and data scientists must be acutely aware of the default behaviors of AI tools and libraries. As seen with Comfy UI, seemingly benign defaults (like embedding full workflows in PNG metadata) can lead to critical secret exposures. Always review the output and configuration of tools that process or generate data for hidden sensitive information.
  3. Secure GitHub and Version Control Practices:
  • GitHub Personal Access Tokens (PATs): Use PATs with the absolute minimum necessary scope and lifetime. Regularly rotate them. For automated scraping, use ephemeral tokens or GitHub Apps with fine-grained permissions.
  • .git Directory Awareness: Educate developers about the contents of the .git directory and the risks of bundling it when publishing archives. Ensure that build processes remove or clean sensitive configuration files before packaging.
  • Private Repository Protection: Reiterate the importance of never using tokens with private repository access for public-facing or shared scraping activities.
  1. Review Data Curation Processes: Recognize the "data curation bias" that can inadvertently favor code with hardcoded secrets. Data scientists should explore alternative curation methods that do not rely on execution-based filtering, or implement additional checks for secrets before code execution.
  2. Monitor Non-Traditional Data Sources: Security teams need to expand their monitoring to include non-traditional data sources that AI models might scrape, such as public chat channels (e.g., Telegram), forums, and other less structured online content, as these can be sources of actively exploited credentials.
  3. Embrace "Assume Compromise" for Public Data: Given the scale and impossibility of cleanup, organizations should operate under the assumption that any secret or sensitive data inadvertently published to a public AI hub is compromised. This necessitates immediate revocation of exposed keys, re-issuance, and a thorough assessment of potential impact.
  4. Advocate for Policy and Standards: The OSI debate underscores the urgent need for clear, enforceable standards and definitions for "open-source AI" that explicitly address data security, privacy, and copyright. Organizations should engage in these discussions to shape future regulations that protect both innovation and security.
  5. Ephemeral Credentials and Granular Access: Where possible, leverage ephemeral credentials for API access and enforce granular, least-privilege access controls for all services. This limits the blast radius if a key is exposed.
  6. Continuous Monitoring of Public Datasets: For organizations with a significant digital footprint, continuous monitoring of major AI hubs and public data repositories for their own leaked secrets or proprietary information is becoming a necessity.

Ultimately, the defensive posture must evolve from reactive cleanup to proactive prevention and continuous vigilance, recognizing that the scale and interconnectedness of AI datasets have created an unprecedented challenge for data security. The "toxic data" problem is not going away, meaning mitigation and containment are paramount.

Key Takeaways

  • HuggingFace is a massive reservoir of exposed secrets: Hundreds of thousands of unique, live API keys and other sensitive data have been found across its vast AI datasets, even with only a fraction of the hub scanned.
  • Unintentional leaks are widespread: Many exposures result from developers and data scientists unknowingly including sensitive information through common practices, tool defaults (e.g., Comfy UI embedding keys in PNGs), or overlooked details (e.g., .git directories in published archives).
  • "Open data" is fundamentally different from "open source code": The legal, privacy, and security challenges of sharing massive, diverse datasets are orders of magnitude more complex than those for software, as highlighted by the OSI debate and copyright issues.
  • Cleanup is practically impossible: Due to the sheer scale of data, its replication across multiple datasets, and the continuous nature of scraping, completely sanitizing and cleaning up exposed secrets is an insurmountable task.
  • Data curation can exacerbate the problem: Subtle biases in how data is collected and processed (e.g., favoring code with hardcoded keys because it "runs") can create a vicious feedback loop, leading to more secrets in training data and subsequent LLM outputs.
  • Proactive measures are critical: Organizations and individuals must implement stringent data sanitization, enforce least-privilege access for API keys, understand tool defaults, and continuously monitor for leaks, recognizing that prevention is the only viable strategy.

About the Speaker(s)

Dylan Ayrey is the CEO and co-founder of Truffle Security, a company built upon the success of the popular open-source tool Truffle Hog. His work and research are primarily focused on the critical area of finding and securing API keys, passwords, and other sensitive secrets. Ayrey is a recognized expert in this field, frequently giving talks and sharing insights on secret detection and the broader implications of data exposure. His deep understanding of how secrets manifest in various data environments, from traditional code repositories to the emerging landscape of AI training datasets, underpins the extensive research presented in this talk.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Ayrey brings real scan data — 100k live secrets across a platform that most security people haven't thought to look at yet — and builds a coherent narrative around why this problem is structurally unfixable, not just a collection of oopsies. The data curation bias finding (sandboxes preferentially retain hardcoded keys because they execute cleanly) is the kind of second-order insight that separates actual research from a grep report.

Heather Calloway (CISO) — SOLID

Ayrey surfaces a real and underappreciated exposure vector — AI training datasets as a secondary market for leaked credentials — with credible empirical grounding and specific case detail. The research is legitimate and the scale findings are genuinely striking, but the talk stays in problem-statement territory and doesn't reach the governance or institutional accountability questions that would make it matter at the leadership level.

→ Top-rated talks at BSidesSF 2026

All talks from BSidesSF 2026