Fresh Secrets From the Docks: Lessons Learnt From Analyzing 180,000 Public Dock... Guillaume Valadon
Guillaume Valadon
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this compelling KubeCon EU talk, Guillaume Valadon, a security researcher at GitGuardian, unveiled the alarming prevalence of leaked secrets within public Docker images. Titled "Fresh Secrets From the Docks," the presentation detailed an extensive research effort involving the scanning of over 50 million Docker images on Docker Hub. Valadon's work highlights a critical blind spot in many organizations' security postures, demonstrating how easily sensitive credentials, API keys, and private keys can inadvertently end up in publicly accessible container images, posing a significant risk for supply chain attacks, data breaches, and unauthorized resource utilization.

Key moments
- 0:00 Speaker introduction and GitGuardian's secret detection mission
- 2:00 Attacker motivations for secrets and MITRE ATT&CK framework
- 3:58 Deconstructing a Docker image: layers, blobs, and manifest
- 6:00 Secret scanning with gg shields and validity checks
- 7:40 Methodology for scanning multiple Docker images
Fresh Secrets From the Docks: Lessons Learnt From Analyzing 180,000 Public Docker Images
Speakers: Guillaume Valadon, Security Researcher, GitGuardian
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=2r92tTuFYg8
Overview
In this compelling KubeCon EU talk, Guillaume Valadon, a security researcher at GitGuardian, unveiled the alarming prevalence of leaked secrets within public Docker images. Titled "Fresh Secrets From the Docks," the presentation detailed an extensive research effort involving the scanning of over 50 million Docker images on Docker Hub. Valadon's work highlights a critical blind spot in many organizations' security postures, demonstrating how easily sensitive credentials, API keys, and private keys can inadvertently end up in publicly accessible container images, posing a significant risk for supply chain attacks, data breaches, and unauthorized resource utilization.
Valadon, known for his contributions to the Scapy network manipulation tool, presented a deep dive into the methodology, challenges, and shocking findings of this large-scale secret detection campaign. The talk served as a stark warning to the cloud-native community, emphasizing that even seemingly innocuous development practices can lead to severe security vulnerabilities. It underscored the urgent need for developers and security teams to adopt robust secret management practices throughout the container build and deployment lifecycle.
The research not only quantified the scale of the problem—revealing that 5% of all scanned repositories contained at least one secret and a staggering 100,000 valid secrets are currently exposed—but also elucidated the common Docker pitfalls that lead to these leaks. Valadon's insights provide invaluable guidance for defenders seeking to prevent, detect, and remediate secret sprawl within their containerized environments, urging a proactive approach to what he describes as an "insane" and persistent threat.
Background
▶ Watch: Speaker introduction and GitGuardian's secret detection mission (0:00)
Guillaume Valadon commenced his talk by introducing himself as a long-time network engineer and a core contributor to the Scapy Python library, a powerful tool for packet manipulation. He now applies his expertise to security research at GitGuardian, a French company specializing in secret detection. While GitGuardian initially focused on secrets within Git and GitHub repositories, their scope has expanded to include messaging systems like Slack and, as demonstrated in this talk, container images. Their mission is to help customers identify and remediate exposed secrets.
Valadon referenced GitGuardian's "Secrets Pro Report," which highlights that secrets are often leaked more frequently in private repositories than in public ones, suggesting a false sense of security among developers. He set the stage by illustrating how attackers actively seek secrets in various locations, citing recent incidents where secrets were found in PIP packages and GitHub Action logs, leading to breaches against organizations like Coinbase. His objective was to approach the problem from an offensive security perspective, mimicking an attacker's methodology to uncover vulnerabilities.
The core problem, as Valadon articulated, is the exposure of "secrets"—which he defines broadly as anything from usernames and passwords to private keys and API tokens that grant access to systems as a regular user. Attackers leverage these secrets for various malicious activities, including cryptomining (by spinning up powerful cloud instances using stolen credentials), data exfiltration (accessing databases or S3 buckets), and supply chain attacks (injecting malicious code into compromised projects). The focus of this research was on container registries, specifically Docker Hub, as a rich target for such secrets.
To understand how secrets leak in Docker images, Valadon provided a primer on Docker image structure. A Docker image is built from a Dockerfile using commands like docker build. The resulting image is composed of multiple layers, each representing a change introduced by an instruction (e.g., FROM, COPY, RUN). These layers, also known as blobs, are essentially gzipped tarballs. In addition to layers, an image includes a JSON config file and a manifest file, which lists the layers and the config. Crucially, each layer's name is an SHA256 hash of its content, meaning identical layers across different images will share the same hash. Valadon recommended Skopeo as a tool for inspecting image contents and extracting blobs for analysis. GitGuardian's internal tool, ggshield, is designed to scan these components for secrets, complete with validity checks that attempt to verify if a detected secret (e.g., a GitHub token) is still active by testing it against a known API endpoint. A "cannot check" status is also significant, indicating a secret that might be valid but is inaccessible for public validation, making it particularly interesting for internal network lateral movement.
Key Findings
▶ Watch: Attacker motivations for secrets and MITRE ATT&CK framework (2:00)
Guillaume Valadon's research unearthed a startling landscape of secret exposure within public Docker images, providing concrete numbers that underscore the severity of the problem.
- Widespread Secret Leaks: A staggering 5% of all scanned Docker repositories contain at least one secret. This translates to one in 20 repositories harboring sensitive information, a number Valadon described as "huge."
- Layers, Not Dockerfiles, are the Primary Source: The vast majority of secrets (Valadon noted "most of these secrets") originate from image layers, rather than being explicitly hardcoded in
Dockerfileinstructions. This suggests that secrets are often introduced inadvertently during the build process or via copied files, rather than directARGorENVusage within the Dockerfile itself. - High Validity Rate of Specific Secrets: Among the specific secrets detected (those matching well-known patterns for services like GitHub, AWS, GCP), a significant 20% were found to be valid. This means that out of the total secrets identified, 100,000 secrets were actively functional at the time of the scan, making them immediately exploitable by attackers.
- Critical Secret Categories:
- Data Storage related secrets (databases, S3 buckets, etc.) constituted the largest category, accounting for 28% of all specific secret detectors. Of these, 13% were valid, representing a massive number of exploitable data access credentials.
- Cloud Provider secrets (AWS, GCP, Azure, etc.) were the second largest category, making up 20% of specific secrets. Alarmingly, 50% of these were valid, indicating a high risk of cloud resource compromise.
- Persistence of Leaked Secrets: The research revealed an alarming longevity of exposed secrets. 60% of the valid secrets were leaked before 2024, demonstrating that once a secret is public, it often remains active for extended periods. Even more shockingly, 2,000 secrets from as far back as 2020 were still valid at the time of the presentation. Valadon emphasized the ease with which he personally encounters valid, exploitable secrets daily, citing an example of finding a valid GitHub token providing access to Kubernetes-related open-source projects just before his talk.
- The
docker buildARG Leak: A critical discovery highlighted was thatbuild-argvariables passed during thedocker buildprocess are often inadvertently embedded into the final image's JSON config file. Even if these variables are intended for build-time use and not explicitly stored in layers, Docker's internal handling can persist them, creating a hidden, potent source of leaks that many developers are unaware of, despite warnings in Docker's documentation.
These findings paint a grim picture of secret management in the container ecosystem, underscoring a pervasive and long-standing vulnerability that attackers can readily exploit.
Technical Deep Dive
▶ Watch: Deconstructing a Docker image: layers, blobs, and manifest (3:58)
Valadon's research methodology involved a multi-stage approach to efficiently scan Docker Hub, overcoming significant technical challenges posed by the sheer scale of data and the limitations of the Docker Registry API.
The general methodology for scanning any Docker registry consists of four steps, which map directly to Docker Registry API endpoints:
- Catalog Endpoint: Retrieve a list of all repositories.
- Tags Endpoint: For each repository, get a list of its tags.
- Manifest Retrieval: For each tag and repository combination, fetch the manifest (a JSON file describing the image content).
- Blob Download & Scan: From the manifest, download the blobs (layers and config file) and scan them for secrets.
However, Docker Hub presented unique challenges. The Docker Hub Registry API does not expose a catalog endpoint, making direct enumeration impossible. To circumvent this, Valadon's team employed a clever trick: leveraging the Docker web search API with keywords. Initial attempts used well-known keywords like "production" or "staging," but the web search is limited to 10,000 results. To achieve comprehensive coverage, they resorted to a form of brute-forcing, enumerating keywords systematically (e.g., 'aa', 'ab', 'ac'...). This discovery phase alone took approximately one day.
The initial enumeration yielded a staggering 9 million repositories, translating to an estimated 1 petabyte of data. This volume was unmanageable for storage and scanning, necessitating a series of optimization steps:
- Tag Filtering: Analysis showed that 93% of repositories have five or fewer tags. Valadon decided to focus only on repositories with a maximum of five tags. This reduced the dataset to 6.5 petabytes and 15 million images, making it more manageable. The most common tag found was "latest," a "fun fact" that took five days to ascertain.
- Layer Deduplication: Docker images often share common base layers. To avoid redundant scanning, Valadon deduplicated layers, identifying 44 million unique layers. He noted that the largest unique layer encountered was 148 GB, though such massive layers sometimes "disappeared" during the experiment, indicating the dynamic nature of Docker Hub content. Retrieving config files and manifests to facilitate this layer analysis took approximately 10 days.
- Instruction-Based Layer Filtering: Valadon analyzed the prevalence of different Dockerfile instructions. The
RUNinstruction was the most common, accounting for 30% of all instructions and 70% of the total layer volume (3.3 petabytes).RUNlayers typically contain installations of software (e.g.,apt install,apk add) and are less likely to harbor secrets directly. To further reduce the scan scope, Valadon made a strategic decision to excludeRUNlayers, focusing primarily onCOPYandADDlayers, which are more likely to contain application source code or configuration files where secrets reside. This reduced the data volume from nearly 1 petabyte (after tag filtering) to 412 terabytes across 18 million unique layers. - Layer Size Filtering: A final optimization involved analyzing the size distribution of
ADDandCOPYlayers. 99% of these layers were below 200 MB. Valadon opted for an even stricter filter, scanning only layers smaller than 45 MB, which still covered 90% of the layers. This brought the final dataset to 60 million unique layers, spanning 15 million images, but with a significantly reduced data footprint for scanning.
The entire discovery and scanning process was extensive:
- Discovery Phase (repositories and images): 15 days
- Config File Scanning (with
ggshield): 3 days - Layer Download and Scanning: 20 days (with a slowdown towards the end due to holiday breaks).
For secret detection, GitGuardian employs two types of detectors:
- Specific detectors: These identify secrets belonging to well-known patterns and formats (e.g., GitHub tokens, AWS keys). They can often be automatically validated against known API endpoints. Attackers prioritize these.
- Generic detectors: These identify random-looking strings that might be secrets but lack context for service identification or automatic validation. While GitGuardian has advanced ML models to categorize 13% of these (e.g., as Microsoft SQL server logins), Valadon focused on specific, potentially valid secrets for his offensive security perspective.
A crucial technical finding was the demonstration of Docker's pitfalls leading to secret leaks:
- "Docker Remembers": Valadon illustrated how files added to an image in one layer, even if subsequently removed in a later
RUNcommand, remain present in the earlier layer. For example, if annpmrcfile with credentials isCOPY'd and thenrm'd, the secret persists in theCOPYlayer and can be extracted. docker buildARG/ENV Leaks: The most significant and surprising leak vector involvedbuild-argvariables. Developers are often advised to use variables instead of hardcoding secrets. However, passing secrets via--build-arg(e.g.,docker build --build-arg PASSWORD=secret .) or evenENVinstructions in a Dockerfile, results in these secrets being embedded directly into the final image's JSON config file. This happens even if the variable is not explicitly used or intended to be persisted, a behavior often overlooked despite being documented as inappropriate for secrets in Docker's official warnings. Valadon showed a crafted example whereARG KUBECON=...resulted inkubeconappearing in the image's config, demonstrating this insidious leak.- Direct Hardcoding in
RUNLayers: Despite best practices, Valadon's dataset revealed real-world instances of secrets being hardcoded directly withinRUNcommands, often redirected to configuration files (e.g.,RUN echo "client_secret" > config.json). Even more egregious were cases wheremount secrets(the recommended secure method) were used, but the secret was simultaneously hardcoded in the sameRUNinstruction, effectively nullifying the security benefit.
Demo / Proof of Concept
▶ Watch: Secret scanning with gg shields and validity checks (6:00)
While the presentation did not feature a live, interactive "demo" in the traditional sense, Guillaume Valadon effectively demonstrated the core concepts and findings through illustrative examples and command-line snippets directly from his research.
He first introduced ggshield, GitGuardian's command-line tool, as the primary mechanism for scanning for secrets. A simple example showed ggshield detecting an invalid API key within a Python code snippet, highlighting its capability to identify secrets and perform validity checks.
The most impactful "proof of concept" centered around the docker build ARG leak. Valadon presented a carefully crafted Dockerfile that, on the surface, appeared innocuous:
He then showed the docker build command used:
Crucially, he then demonstrated how to extract this "secret" from the resulting image using skopeo to pull the image and jq to parse its JSON configuration. He displayed the flattened JSON config, clearly showing kubecon and my_super_secret_password embedded within the container_config.config.Args field, even though the ARG was only intended for build-time use. This visual evidence powerfully illustrated how Docker itself can inadvertently persist sensitive build arguments within the final image's metadata, a critical and often overlooked leak vector.
Valadon also provided examples of real-world RUN instructions found in his dataset that directly hardcoded secrets:
RUN echo "S3_CLIENT_ID=..." >> /app/config.jsRUN echo "DB_PASSWORD=..." > /etc/db_config.conf
He even highlighted a "worst-case scenario" where a developer correctly attempted to use secret mounts (e.g., RUN --mount=type=secret,id=mysecret,target=/run/secrets/mysecret ...) but then simultaneously hardcoded the very same secret within the RUN command:
RUN echo "password is my_super_secret_password" > /tmp/config.txt
This demonstrated how even attempts at secure practices can be undermined by a lack of understanding of Docker's layer mechanics and general secret hygiene.
These examples, while not live code execution, served as potent demonstrations of the vulnerabilities Valadon discovered, making the technical findings tangible and alarming for the audience.
Defensive Implications
▶ Watch: Methodology for scanning multiple Docker images (7:40)
The findings from Guillaume Valadon's extensive research carry profound defensive implications for any organization utilizing Docker or container registries. The pervasive nature of secret leaks demands a multi-faceted approach to security.
- Mandatory Image Auditing and Scanning: The most immediate defensive action is to audit all Docker images for hardcoded secrets. This should be an automated, continuous process integrated into the CI/CD pipeline. Tools like GitGuardian's
ggshield(or any other robust secret scanning solution) should be employed to scan not just Dockerfiles, but also all image layers and configuration files. This includes scanning for both specific secrets (API keys, database credentials) and generic patterns that might indicate sensitive data. - Understand Docker's Pitfalls: Developers and security teams must be educated on Docker's specific behaviors that lead to secret leaks:
- "Docker Remembers": Be aware that files copied into an image and then deleted in a subsequent
RUNcommand are not truly removed from the image's history. Secrets in such files will persist in earlier layers. Best practice dictates that secrets should never enter a layer in the first place if they are not intended to be part of the final image. Multi-stage builds can help, but careful attention to what is copied into each stage is paramount. docker buildARG/ENV Leaks: This is a critical point of failure. Build arguments (--build-arg) and environment variables (ENV) are inherently unsuitable for passing secrets to a Docker build. As Valadon demonstrated, these can be inadvertently embedded into the final image's JSON configuration, making them discoverable. This behavior is documented, but frequently overlooked.
- Adopt Secret Mounts (
--secret) for Build-Time Secrets: Docker offers secret mounts as the secure alternative for handling build-time secrets. Instead ofARGorENV, useRUN --mount=type=secret,id=mysecret,target=/run/secrets/mysecretto expose secrets securely only during the specific build step that requires them. These secrets are not persisted in any image layer or the final config. This should be a mandatory best practice for any build process involving sensitive credentials. - Implement Robust Secret Management: Beyond Docker-specific practices, organizations need comprehensive secret management solutions. This includes using tools like HashiCorp Vault, AWS Secrets Manager, Azure Key Vault, or Kubernetes Secrets (with proper encryption and access controls) to store and inject secrets at runtime, rather than baking them into images.
- Focus on Validity Checks and Remediation: Simply detecting a secret isn't enough. Organizations must prioritize verifying the validity of leaked secrets and have a swift remediation plan. This includes immediate revocation, rotation, and investigation into the source of the leak. The "cannot check" status for secrets in private networks should be treated with high suspicion, as these could be valid and exploitable internally.
- Proactive Security Exercises: Valadon strongly advocated for conducting internal "fire drills." He suggested an annual exercise (e.g., in 2025) where teams simulate a major secret leak, such as an AWS root account credential being publicly exposed. The exercise should test the organization's ability to:
- Identify the leaked secret.
- Determine its validity and impact.
- Identify the developer or team responsible.
- Execute a revocation and rotation process.
- Assess the overall response time and effectiveness.
This type of exercise can expose weaknesses in incident response plans and secret management policies, reinforcing that prevention is far more effective and less costly than dealing with a live breach.
- Continuous Education and Awareness: Given the persistence of old leaks and the subtle ways secrets can escape, ongoing education for developers about secure coding practices, Docker best practices, and the dangers of secret sprawl is crucial.
By implementing these defensive measures, organizations can significantly reduce their attack surface and protect themselves from the severe consequences of publicly exposed secrets.
Key Takeaways
- Secret Sprawl is Pervasive: A shocking 5% of public Docker repositories contain at least one secret, demonstrating widespread accidental exposure of sensitive credentials.
- Validity is High and Persistent: 20% of specific secrets are valid, with 100,000 active secrets on Docker Hub. Alarmingly, 2,000 secrets from 2020 are still valid, highlighting the long-term risk of unaddressed leaks.
- Docker's Build Process is a Major Leak Vector: The most critical finding is that
--build-argvariables andENVinstructions can inadvertently embed secrets into the final Docker image's JSON configuration file, even if not explicitly written to a layer. - Avoid
ARGandENVfor Secrets; Use Secret Mounts: To securely pass build-time secrets, developers must useRUN --mount=type=secretinstead ofARGorENVto prevent secrets from being persisted in the image. - Scan All Image Layers: Secrets are often hidden in intermediate layers (e.g., from
COPYorADDinstructions) and not just obvious Dockerfile commands. Comprehensive scanning of all layers and config files is essential. - Proactive Auditing and Incident Response: Organizations must implement continuous secret scanning in their CI/CD pipelines and conduct regular "fire drill" exercises to simulate secret leaks, testing their ability to detect, revoke, and remediate effectively.
About the Speaker(s)
Guillaume Valadon is a distinguished security researcher currently working at GitGuardian, a French company at the forefront of secret detection. With a rich background as a network engineer, Valadon is well-known in the open-source community for his significant contributions to Scapy, a powerful Python library used for packet manipulation and network analysis. His expertise spans from low-level network protocols to modern application security. At GitGuardian, he focuses on uncovering and understanding the mechanisms behind secret leaks, particularly in complex environments like container registries. His work, as demonstrated in this KubeCon EU talk, is crucial for raising awareness about critical security vulnerabilities and guiding the industry towards more secure development practices.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
Valadon's research is a brutal, data-driven gut punch to anyone running containers. By exhaustively scanning over 50 million Docker images, he exposes the staggering scale of leaked secrets on Docker Hub – 5% of repositories, 100,000 valid credentials, some active since 2020. His most critical finding reveals a pervasive, overlooked leak vector: build-arg variables silently embedded in the final image's JSON config. This isn't just theory; it's hard numbers and concrete mechanisms that demand immediate attention from every CISO and developer in the cloud-native space. Absolutely essential reading.
Heather Calloway (CISO) — MUST SEE
This KubeCon talk by Guillaume Valadon is a critical exposure of institutional failure in secret management within the container ecosystem. The sheer scale of 100,000 valid, exposed secrets, some active for years, points to a systemic breakdown in development practices and a profound lack of risk ownership. Valadon's precise breakdown of how build-time arguments (--build-arg) and intermediate layers inadvertently persist secrets provides a clear, actionable understanding of a pervasive threat. Every CISO needs to understand these vectors and immediately implement robust scanning, policy enforcement for secret mounts, and rigorous incident response drills to address this persistent and…