Using Jupyter Notebooks to Explore Public CVE Data

Jerry Gamblin (Principal Engineer · Cisco)

CVE/FIRST VulnCon 2025 · Main Stage

Overview

In this VulnCon workshop, Jerry Gamblin, a Principal Engineer in Cisco’s Threat Detection Response Group, presented a compelling case for democratizing and enhancing the analysis of Common Vulnerabilities and Exposures (CVE) data using open-source tools. The talk, titled "Using Jupyter Notebooks to Explore Public CVE Data," aimed to equip security professionals, researchers, and individuals with the practical skills and resources to perform their own in-depth vulnerability data analysis. Gamblin highlighted the critical need for more accessible and transparent methods for understanding CVE trends, data quality, and the broader vulnerability landscape, moving away from an over-reliance on commercial vendors or a single, potentially unreliable, data source.

Watch on YouTube

Visual summary for Using Jupyter Notebooks to Explore Public CVE Data by Jerry Gamblin
Visual summary for Using Jupyter Notebooks to Explore Public CVE Data by Jerry Gamblin

Key moments

  1. 0:00 Introduction and Speaker Background
  2. 1:30 Why Analyze Public CVE Data? (Motivation)
  3. 3:30 Real-world NVD Data Discrepancy Example
  4. 4:45 Challenges: Using Raw CVE List vs. NVD Data
  5. 6:00 Workshop Goals and Hands-on Learning Approach
  6. 6:15 Key Data Sources: NVD API and CVE List
  7. 8:30 Technologies Used: Jupyter, Pandas, MyBinder

Using Jupyter Notebooks to Explore Public CVE Data

Speakers: Jerry Gamblin, Principal Engineer, Threat Detection Response Group, Cisco

Conference: VulnCon

YouTube: https://www.youtube.com/watch?v=CUzluKxfQO0

Overview

In this VulnCon workshop, Jerry Gamblin, a Principal Engineer in Cisco’s Threat Detection Response Group, presented a compelling case for democratizing and enhancing the analysis of Common Vulnerabilities and Exposures (CVE) data using open-source tools. The talk, titled "Using Jupyter Notebooks to Explore Public CVE Data," aimed to equip security professionals, researchers, and individuals with the practical skills and resources to perform their own in-depth vulnerability data analysis. Gamblin highlighted the critical need for more accessible and transparent methods for understanding CVE trends, data quality, and the broader vulnerability landscape, moving away from an over-reliance on commercial vendors or a single, potentially unreliable, data source.

Gamblin's core motivation stems from the observation that while CVE data is publicly available, its complexity and disparate nature often deter widespread, independent analysis. Many organizations and individuals either depend on third-party enrichments (often commercial) or struggle with the raw data's format. This talk directly addressed these challenges by demonstrating how Jupyter Notebooks, combined with powerful Python libraries like Pandas and Matplotlib, can transform raw CVE and National Vulnerability Database (NVD) data into actionable insights. The workshop underscored the importance of fostering a community-driven approach to vulnerability data analysis, empowering users to extract value directly from the source and critically assess data quality.

The significance of this discussion cannot be overstated in today's cybersecurity landscape. With the increasing volume of vulnerabilities and the recent challenges faced by the NVD program, understanding the nuances and underlying quality of CVE data has become paramount. Gamblin’s work promotes a proactive stance, encouraging practitioners to build their own analytical capabilities rather than passively consuming pre-digested information. By providing ready-to-use Jupyter notebooks and a clear methodology, he sought to enable a deeper, more critical engagement with vulnerability data, ultimately leading to more informed defensive strategies and a stronger call for improved data quality from the CVE program itself.

Background

▶ Watch: Introduction and Speaker Background (0:00)

The landscape of vulnerability information has long been dominated by the National Vulnerability Database (NVD), which traditionally served as the primary, enriched source for CVE data. While the CVE Program (cve.org) publishes raw vulnerability identifiers and basic details, the NVD has historically added crucial contextual information, such as Common Platform Enumeration (CPE), Common Weakness Enumeration (CWE), and Common Vulnerability Scoring System (CVSS) scores. This enrichment made the NVD an indispensable resource for vulnerability management teams globally. However, as Gamblin points out, this reliance created a single point of failure and obscured the underlying data quality issues within the raw CVE program.

A significant challenge highlighted by Gamblin is the difficulty for individuals and smaller teams to directly leverage the raw CVE data from cve.org. Unlike the NVD, which provides consolidated JSONL files, the CVE program distributes its data as 25,000 individual JSON files in a GitHub repository. This format necessitates sophisticated Extract, Transform, Load (ETL) processes to aggregate and normalize, making it a daunting task for many users. Consequently, even though the CVE program is the ultimate source, most practitioners still default to the NVD, even with its recent reliability concerns. Gamblin's own project, cve.iciu, which produces daily vulnerability data reports, was born out of this need to make raw CVE data more accessible and understandable.

Furthermore, Gamblin detailed specific instances illustrating the disconnect and data quality issues. He recounted a personal experience where Cisco’s internal teams were unaware that the NVD was mapping their submitted CVE data to a private database of email addresses, leading to an individual's email appearing on public records for a decade. This incident underscored how crucial it is for organizations, especially CVE Numbering Authorities (CNAs), to understand the entire data flow and how their submissions are processed and presented at the user endpoint. The recent "government issues" impacting the NVD's update cadence further exacerbated the problem, pushing the community to question who the true "source of truth" for CVE data is, especially with the emergence of alternative data enrichers like NVD++, OSV.dev, and various GitHub-based vulnerability databases. Gamblin argues that without a universally accepted starting point for vulnerability discussions, teams are left to "bring your own data," leading to inefficient and inconsistent vulnerability management practices.

Key Findings

▶ Watch: Real-world NVD Data Discrepancy Example (3:30)

Jerry Gamblin's presentation unveiled several critical findings regarding the state of public CVE data, its analysis, and the underlying challenges:

  1. NVD Data Mapping Inconsistencies: A significant discovery was the NVD's practice of mapping submitted CVE data to a private database of email addresses, which for years resulted in individual Cisco engineers' emails appearing publicly instead of the intended organizational contact. This highlights a lack of transparency and potential data privacy concerns within the NVD's enrichment process, unnoticed by the submitting CNAs for an extended period.
  1. Difficulty in Consuming Raw CVE Data: Despite being the authoritative source, the raw CVE v5 list from cve.org is impractical for most users. It's distributed as thousands of individual JSON files, making it challenging to clone, aggregate, and analyze without complex ETL processes. This forces many users to rely on the NVD, even when alternative, more direct analysis of the CVE program's data is desired.
  1. Incomplete CVE Schema Descriptions: Analysis of the CVE schema revealed significant gaps in its documentation. Only 62% of keys have descriptions, and a mere 50% have type definitions. This incompleteness hinders automated validation, data quality checks, and clear understanding for developers and analysts trying to parse and interpret CVE records programmatically.
  1. Minimal Data Quality Checks in the CVE Program: The current CVE program's pre-publishing and post-publishing quality assurance processes are rudimentary. Gamblin noted that the schema allows for a description as short as one character, and the primary automated check is based on patterns (only 7% of keys have patterns). This lax validation enables the publication of CVEs with "minimal data," such as Microsoft's common "buffer overrun in Windows kernel" descriptions, which lack sufficient detail for effective defense. Furthermore, the presence of 59 CVEs with null datePublished fields in the 2025 dataset alone points to fundamental data integrity issues.
  1. Lack of CVE Program Ownership over Tooling: A major impediment to improving CVE data quality and adoption is the CVE program's lack of ownership or direct control over the major tools used by CNAs to publish CVEs. These tools are often open-source, maintained by volunteers, leading to delays in schema updates and feature implementation. This dependence on external, community-driven tooling creates a bottleneck for program-wide improvements and standardization.
  1. "Source of Truth" Confusion with ADPs and Alternative Databases: The emergence of Authorized Data Publishers (ADPs) and other vulnerability databases (e.g., NVD++, OSV.dev, GitHub's database) has "muddied the water" regarding the authoritative source for CVE details. Different sources may provide conflicting information, such as varying CVSS scores, making it difficult for organizations to establish a consistent baseline for vulnerability discussions and risk prioritization.
  1. Shifting Landscape of CNA Activity: For the first time, a non-MITRE CNA, Patchstack.com, has surpassed MITRE in the number of CVEs published. Along with WordFence, these CNAs primarily focus on WordPress plugins, demonstrating a strategic niche filling. While this increases the volume of CVEs, it also highlights the program's ability to accommodate high-volume, potentially "less critical" vulnerabilities, which might not receive prioritization without a CVE ID. This shift underscores the importance of the CVE ID as a driver for vulnerability management action, even for relatively minor issues.
  1. CPE Data Quality Issues: The CPE (Common Platform Enumeration) format, frequently used for affected products, is effectively an "abandoned" standard with no clear ownership or active board. This abandonment contributes to poor data quality, exemplified by single CVEs listing an exorbitant number of affected products (e.g., one CVE with 4,800 affected products, another Cisco CVE with 2,400). Such inflated lists render CPE data less useful for practical vulnerability management.

Technical Deep Dive

▶ Watch: Challenges: Using Raw CVE List vs. NVD Data (4:45)

The core of Jerry Gamblin's workshop was a hands-on exploration of public CVE data using a suite of open-source tools, primarily centered around Jupyter Notebooks. This approach aims to provide a flexible, interactive, and reproducible environment for data analysis.

The foundational technologies employed include:

  • Jupyter Notebooks: An open-source web application that allows users to create and share documents containing live code, equations, visualizations, and narrative text. Its cell-based structure is ideal for iterative data exploration and breaking down complex analysis into manageable steps.
  • Pandas: A powerful Python library for data manipulation and analysis. It introduces the DataFrame object, an in-memory, two-dimensional, tabular data structure with labeled axes (rows and columns), similar to a spreadsheet or SQL table. Pandas is crucial for loading, cleaning, transforming, and querying vulnerability data.
  • Matplotlib: A comprehensive library for creating static, animated, and interactive visualizations in Python. It was used in Gamblin's examples (and his cve.iciu project) to generate graphs illustrating CVE trends, such as daily and monthly publication rates, and top CNA assigners.
  • MyBinder: A free, cloud-based service that allows users to launch a GitHub repository in an executable environment, typically a Jupyter Lab instance. MyBinder automatically builds a Docker container based on the repository's requirements.txt (or similar dependency files), providing a reproducible environment without local setup. This was chosen over Google Colab due to its seamless, no-authentication repository cloning capabilities, making it ideal for workshop participants.

Data Sources and Ingestion:

Gamblin demonstrated how to acquire data from two primary sources:

  1. NVD 2.0 API (via nvd.handsonhacking.org): For ease of use and to avoid API rate limits during the workshop, Gamblin provided access to nvd.handsonhacking.org, a personal project that hosts a single NVD.jsonL file containing every CVE record from the NVD, updated approximately every four hours. The underlying Python scraper code (NVD_API.py) was also included in the workshop repository, allowing users to run their own copies. This NVD.jsonL file is substantial, weighing around 1.3 GB.
  2. CVE List (version 5) from GitHub: The official CVE program data is hosted on a GitHub repository. For the workshop, only the CVEs published in 2025 were downloaded to manage file size and MyBinder's 3 GB storage limit. This involved a shallow clone of the relevant year-specific folder. Gamblin noted a limitation here: the CVE list is organized by identifier year, not publication year, meaning not all CVEs published in 2025 are necessarily in the CVE-2025 directory.

Data Processing Workflow (as intended by the notebooks):

The workshop outlined a clear series of steps using the Jupyter notebooks:

  1. data.ipynb: This notebook orchestrates the initial data download. It fetches the NVD.jsonL file from nvd.handsonhacking.org, performs a shallow clone of the 2025 folder from the CVE GitHub repo, and downloads the CVE schema file.
  2. cve_schema.ipynb: This notebook analyzes the completeness of the CVE schema itself. By parsing the schema definition, it generates reports indicating the percentage of keys with descriptions (62%) and type definitions (50%), highlighting areas for improvement in schema documentation and automated validation. It also notes that only 7% of keys have patterns, which is the current primary form of automated data quality check.
  3. cve_dataframe.ipynb: This notebook demonstrates how to construct a Pandas DataFrame from the raw CVE data. Gamblin's baseline DataFrame extracts 15 key fields for general use cases, including CVE ID, state, datePublished, dateUpdated, CWE, CNA short name, and affected products. Users are encouraged to customize this to include any of the 237 keys available in the CVE schema. The DataFrame acts as an in-memory database, enabling efficient querying and analysis.
  4. cve_data_quality.ipynb: This notebook provides a mechanism to assess the completeness and integrity of the loaded CVE data. It breaks down the contents of the DataFrame cells, identifying null values or missing information. For example, it revealed 59 CVEs with null datePublished within the 7,719 CVEs loaded for 2025, underscoring real-world data quality issues.
  5. NVD Dataframe and Quality Notebooks: Similar notebooks exist for NVD data, allowing for parallel analysis of NVD-specific fields and quality checks, although these were not fully demonstrated due to technical issues.

The use of Pandas DataFrames is central to the analytical approach, allowing users to easily select columns, filter data, and perform aggregations. The modular nature of Jupyter notebooks, with their cell-based execution, enables users to run segments of code independently, facilitating debugging and iterative development, a significant advantage for data science tasks.

Demo / Proof of Concept

▶ Watch: Key Data Sources: NVD API and CVE List (6:15)

The workshop included an ambitious live demonstration intended to walk participants through the entire process of acquiring, parsing, and analyzing CVE data using Jupyter Notebooks and MyBinder. The speaker, Jerry Gamblin, had prepared eight distinct notebooks to guide attendees.

The planned demonstration workflow was as follows:

  1. Launch MyBinder: Participants would click a "Launch Binder" link, which would initiate the creation of a temporary, cloud-based Docker container running Jupyter Lab, pre-configured with all necessary dependencies from the GitHub repository.
  2. Download Data (data.ipynb): The first notebook, data.ipynb, was designed to download the substantial NVD.jsonL file (approximately 1.3 GB) from nvd.handsonhacking.org and perform a shallow clone of the 2025 CVE data from the official CVE GitHub repository. This process was expected to take three to four minutes, with the NVD file downloading in about 45 seconds under ideal conditions.
  3. Analyze CVE Schema (cve_schema.ipynb): This notebook would then be executed to parse the CVE schema and generate a report on its completeness, showing percentages of keys with descriptions, types, and patterns.
  4. Build CVE DataFrame (cve_dataframe.ipynb): Next, the cve_dataframe.ipynb notebook would construct a Pandas DataFrame, populating it with selected key fields from the downloaded CVE data.
  5. Assess CVE Data Quality (cve_data_quality.ipynb): The cve_data_quality.ipynb notebook would then analyze the completeness of the DataFrame, identifying records with missing fields, such as the 59 CVEs with null datePublished that Gamblin had identified in his preparatory runs.
  6. Explore NVD Data: Similar notebooks for the NVD data (NVD_schema.ipynb, NVD_dataframe.ipynb, NVD_data_quality.ipynb) were also prepared to demonstrate parallel analysis of NVD-specific enrichments.
  7. Visualization Examples: Finally, Gamblin intended to showcase how to use Matplotlib to generate visualizations similar to those on his cve.iciu website, depicting trends like CVEs published per day/month and the top CNA assigners.

Unfortunately, the live demonstration encountered significant technical difficulties. Despite prior testing on various networks, the North Carolina state Wi-Fi environment proved problematic, leading to persistent "server connection error" messages within the Jupyter Lab environment. As MyBinder sessions are stateless, these errors resulted in the loss of in-progress work and required repeated reloads, making it impossible to complete the full, interactive walkthrough as planned. Gamblin openly acknowledged the challenges inherent in live demos and offered to assist participants individually. While the direct hands-on portion was curtailed, the speaker was able to describe the intended output and show static examples of the visualizations and data quality reports that the notebooks were designed to produce, such as the cve.icu graphs illustrating CVE publication trends and CNA activity, including the notable rise of Patchstack.com. He emphasized that the notebooks were fully functional and available on GitHub for users to run on their local machines or other cloud instances.

Defensive Implications

▶ Watch: Technologies Used: Jupyter, Pandas, MyBinder (8:30)

The insights and tools presented by Jerry Gamblin have profound implications for defensive security postures and vulnerability management strategies. Organizations and individual defenders can leverage this approach to significantly enhance their understanding and response to the evolving threat landscape:

  1. Diversify Vulnerability Data Sources: Defenders should recognize the risks of relying solely on a single source like the NVD. With recent NVD reliability issues and the "muddying of the water" regarding the source of truth, it's crucial to integrate data from multiple origins. This includes the raw cve.org data, as well as alternative enriched sources like NVD++, OSV.dev, and other open-source vulnerability databases. This diversification provides a more comprehensive and resilient view of vulnerabilities.
  1. Build Internal Data Analysis Capabilities: The workshop advocates for moving beyond passive consumption of vulnerability feeds. By adopting Jupyter Notebooks, Pandas, and Matplotlib, security teams can develop internal capabilities to analyze raw CVE and NVD data. This enables tailored insights, customized reporting, and a deeper understanding of trends relevant to their specific asset inventories and threat models. Such capabilities reduce dependence on commercial vendors for basic data enrichment that can be performed in-house.
  1. Critically Evaluate Data Quality: Defenders must become more discerning about the quality and completeness of the vulnerability data they consume. Gamblin's analysis of the CVE schema (e.g., 62% descriptions, 50% types) and the presence of records with minimal detail (e.g., one-character descriptions, null datePublished fields) highlights that not all CVEs are created equal. Security teams should develop processes to identify and account for incomplete or ambiguous CVE records, potentially cross-referencing with other sources or internal intelligence to fill gaps.
  1. Influence CNA Behavior and Advocate for Standards: By actively analyzing CVE data, defenders can identify patterns in CNA submissions, including those with poor data quality or overly broad CPE listings. This analytical feedback can be used to advocate for better data quality standards within the CVE program and encourage CNAs to provide more actionable and granular information. Organizations acting as CNAs themselves should use these tools to self-assess their own data quality before publication.
  1. Proactive Vulnerability Prioritization: Understanding the volume and characteristics of CVEs from different CNAs (e.g., the high volume from Patchstack for WordPress plugins) allows for more nuanced prioritization. While a CVE for a WordPress plugin might appear less critical than a kernel vulnerability, its existence means it will land on a vulnerability management roadmap. Defenders can use this information to anticipate the types of vulnerabilities affecting their software stacks and allocate resources accordingly, distinguishing between high-impact, deeply technical flaws and high-volume, potentially simpler issues.
  1. Address CPE and Affected Product Ambiguity: The issues with CPE (Common Platform Enumeration) data, including its status as an "abandoned" format and the prevalence of single CVEs listing thousands of affected products, demand attention. Defenders should be wary of overly broad CPE matches and, where possible, integrate more precise software bill of materials (SBOM) data or alternative product identification methods (e.g., PURL, OSV.dev's ecosystem-specific identifiers) into their vulnerability management systems to reduce false positives and improve targeting of remediation efforts.
  1. Automate and Integrate Analysis: The principles demonstrated can be integrated into automated workflows. Once stable Jupyter notebooks are developed, they can be run periodically (e.g., via GitHub Actions, as cve.icu does) to generate updated reports, CSV files, or even dynamic dashboards. This automation ensures continuous monitoring of the vulnerability landscape and allows security teams to react swiftly to new trends or emerging threats identified through their custom analytics.

Key Takeaways

  • Empowerment Through Open Tools: Jupyter Notebooks, Pandas, and Matplotlib provide powerful, free tools for independent, in-depth analysis of public CVE data, reducing reliance on commercial vendors.
  • Critical Assessment of Data Sources: Don't solely trust the NVD; understand the raw CVE data from cve.org and consider diverse sources (NVD++, OSV.dev) due to NVD reliability issues and "source of truth" ambiguities.
  • CVE Data Quality is Inconsistent: The CVE schema has significant gaps (e.g., only 62% keys with descriptions), and minimal validation allows for publication of CVEs with insufficient detail, impacting their utility for defenders.
  • CNA Activity Shifts and Volume: CNAs like Patchstack.com are out-publishing MITRE by focusing on niche areas (e.g., WordPress plugins), highlighting the CVE ID's role in driving vulnerability management prioritization, regardless of perceived severity.
  • CPE Data is Problematic: The CPE format is effectively abandoned, leading to poor data quality and overly broad affected product lists, making it less useful for precise vulnerability identification.
  • Build Your Own Intelligence: Developing in-house data analysis capabilities allows organizations to generate tailored insights, identify relevant trends, and proactively influence data quality improvements within the broader CVE ecosystem.

About the Speaker(s)

Jerry Gamblin is a Principal Engineer in the Threat Detection Response Group at Cisco. With a significant background in cybersecurity and data management, Jerry has dedicated his career to understanding and improving how vulnerability data is handled and utilized. Prior to his role at Cisco, he spent a decade working for the government, followed by leading the security program for Carfax, one of the world's largest data companies at the time. He then joined Kenna Security, a vulnerability management startup, which was subsequently acquired by Cisco five years ago.

At Cisco, Jerry plays a crucial role in guiding vulnerability data into all Cisco products, working on both the user-facing side (advocating for better data usability) and engaging with CNAs on publishing practices. He is actively involved in the CVE community, sitting on the CVE Quality Working Group, where he champions initiatives for improved data quality and tooling. Jerry is also the creator and operator of cve.iciu, a website that generates daily vulnerability data reports, and nvd.handsonhacking.org, which hosts a consolidated NVD data file for public access. His work consistently highlights the need for open, accessible, and high-quality vulnerability data.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

A competent, practitioner-focused workshop from someone who clearly lives in the CVE data plumbing every day. Gamblin knows his material cold — he's on the CVE Quality Working Group, runs cve.icu, and has firsthand experience with NVD's data quality failures. The content is honest, the tooling recommendations are sound, and the specific findings (62% schema key coverage, CPE effectively abandoned, Patchstack overtaking MITRE by volume) give attendees real signal they can act on. Nothing here is going to make a researcher's jaw drop, but it's a genuinely useful workshop for a practitioner audience that needs to stop blindly trusting NVD feeds and start building their own analytical muscle…

Heather Calloway (CISO) — SOLID

Gamblin is doing real, useful work — exposing data quality failures in the CVE program, demonstrating that the NVD cannot be treated as ground truth, and building accessible tools that help practitioners analyze raw vulnerability data themselves. The findings are legitimate and the problems he identifies — fragmented source-of-truth, abandoned CPE standards, schema gaps, CNA tooling neglect — are real institutional failures with real operational consequences. But the talk is pitched as a practitioner workshop, and it stays there. It never makes the governance case that these findings warrant. Who owns this problem? What should CNAs, MITRE, or CISA be accountable for? What should a CISO…

→ Top-rated talks at CVE/FIRST VulnCon 2025

All talks from CVE/FIRST VulnCon 2025