BOF: Vulnerability Data Consumers

CVE/FIRST VulnCon 2025 · Birds of a Feather

Overview

This Birds of a Feather (BOF) session at VulnCon brought together security practitioners and data engineers to openly discuss the pervasive challenges associated with consuming and leveraging vulnerability data. Moderated by an individual with extensive experience in the field, including involvement with CVE working groups and the Exploit Prediction Scoring System (EPSS), the session aimed to bridge the communication gap between those who produce vulnerability data and those who rely on it for critical security operations. The core discussion revolved around the practical pain points encountered when ingesting, parsing, normalizing, and acting upon data from sources like CVE.org, NVD, OSV, and various vendor advisories.

Watch on YouTube

Visual summary for BOF: Vulnerability Data Consumers
Visual summary for BOF: Vulnerability Data Consumers

Key moments

  1. 0:00 Setting the stage: Challenges with vulnerability data consumption
  2. 1:00 Initial poll on who's consuming CVE.org and NVD
  3. 2:15 Pain points: Struggling with data format and parsing
  4. 5:00 Challenges in normalizing CVSS scores from multiple sources
  5. 6:20 Discord comment: Poor data quality and CNA rule enforcement
  6. 7:50 Discussion on balancing strictness for data input quality
  7. 8:30 The continuous challenge of keeping up with vulnerability updates

BOF: Vulnerability Data Consumers

Speakers: Moderator (unnamed in metadata), VulnCon Attendees

Conference: VulnCon

YouTube: https://www.youtube.com/watch?v=GyFyXvmlxHY

Overview

This Birds of a Feather (BOF) session at VulnCon brought together security practitioners and data engineers to openly discuss the pervasive challenges associated with consuming and leveraging vulnerability data. Moderated by an individual with extensive experience in the field, including involvement with CVE working groups and the Exploit Prediction Scoring System (EPSS), the session aimed to bridge the communication gap between those who produce vulnerability data and those who rely on it for critical security operations. The core discussion revolved around the practical pain points encountered when ingesting, parsing, normalizing, and acting upon data from sources like CVE.org, NVD, OSV, and various vendor advisories.

The conversation underscored a critical disconnect: while vulnerability data is the bedrock of modern cybersecurity, its current state often hinders effective risk management and prioritization. Attendees revealed a landscape riddled with bespoke, often manual, processes developed to compensate for data quality issues, inconsistent formats, and the sheer volume of information. The session highlighted that these challenges are not merely inconveniences but significant impediments to an organization's ability to identify, assess, and remediate vulnerabilities efficiently, ultimately impacting overall security posture.

The insights gathered from this candid discussion are vital for the evolution of vulnerability data ecosystems. By articulating the real-world struggles of consumers, the session laid the groundwork for future improvements in data quality, standardization, and delivery mechanisms. It served as a powerful reminder that the utility of vulnerability data is only as good as its consumability, emphasizing the urgent need for a more user-centric approach to its generation and distribution.

Background

▶ Watch: Setting the stage: Challenges with vulnerability data consumption (0:00)

The modern cybersecurity landscape is characterized by an ever-increasing volume of disclosed vulnerabilities. Organizations are tasked with managing a deluge of information from diverse sources, including CVE.org (the primary source for Common Vulnerabilities and Exposures), NVD (National Vulnerability Database, which enriches CVEs with additional data like CVSS scores), OSV (Open Source Vulnerability database), GitHub advisories, JPER, CNNVD, and numerous vendor-specific feeds. This proliferation of data, while intended to improve security, has inadvertently created a complex problem for consumers: how to effectively ingest, process, and derive actionable intelligence from such varied and often inconsistent inputs.

Historically, the consumption of vulnerability data has been an arduous task. Early CVE data, for instance, was primarily available in XML format, requiring specialized parsing. While a shift to JSON has been widely welcomed, it hasn't eliminated all parsing difficulties. The fundamental challenge stems from the organic, decentralized nature of vulnerability disclosure, where different CVE Numbering Authorities (CNAs) and data providers operate with varying levels of adherence to standards, quality control, and update frequencies. This leads to a fragmented ecosystem where crucial information, such as CVSS (Common Vulnerability Scoring System) scores, can differ across sources for the same vulnerability, or essential fields might be missing or inconsistently populated.

Prior attempts to address these issues have focused on schema updates, such as the NVD CPE format 5.1.1, and the establishment of working groups like the CVE Quality Working Group. However, the session revealed that these efforts, while beneficial, have not fully alleviated the practical pain points experienced by end-users. The underlying problem persists due to a lack of consistent enforcement of data quality rules, an absence of standardized mechanisms for tracking changes, and, critically, a perceived lack of a strong, collective voice from the consumer community to guide data producers. This environment forces organizations to invest heavily in bespoke solutions for data normalization, enrichment, and prioritization, diverting resources from actual remediation efforts.

Key Findings

▶ Watch: Pain points: Struggling with data format and parsing (2:15)

The BOF session surfaced several critical pain points and challenges faced by vulnerability data consumers:

  • Complex Data Parsing: Attendees consistently highlighted the difficulty of parsing raw vulnerability data. The JSON format, while an improvement over XML, still presents significant hurdles due to deeply nested arrays, inconsistent field population (e.g., fields being present but empty, or using "N/A" text instead of being truly null), and unexpected character usage. One attendee described their parsing efforts using Bash and JQ as a series of "ifs and loops and arrays" far more complex than anticipated, requiring constant one-off checks for data anomalies.
  • Data Quality and Accuracy Issues: A prominent concern was the variable quality and accuracy of CVE descriptions and associated data. Feedback from the Discord channel echoed sentiments that CVE descriptions are "often pathetic" and lack sufficient detail, particularly from smaller vendors. The moderator acknowledged that while some checks exist, there's a clear challenge with data quality, partly due to the trade-off between strict enforcement (which could discourage CNA participation) and ensuring minimum quality standards. Specific issues included ambiguity in describing affected configurations and inaccuracies in version range specifications, often necessitating manual expert review.
  • Difficulty Keeping Up with Updates: A major operational challenge is tracking changes to vulnerability metadata. Vulnerabilities initially deemed low or medium severity can rapidly escalate to critical status if actively exploited. Organizations struggle to receive timely notifications or feeds of such updates, often resorting to manually re-querying and comparing data, which is impractical for thousands of vulnerabilities. This leads to reliance on commercial threat feeds or manual checks of major vendor advisories (e.g., Cisco, Microsoft) to identify critical shifts.
  • Normalization Across Diverse Sources: The problem of inconsistent scoring and metadata across different sources was frequently cited. A CVE might have a primary CVSS score from the CNA, a different score from NVD, and yet another from a security scanner. This lack of normalization makes it exceedingly difficult for organizations to consistently prioritize vulnerabilities across their diverse toolchains. An analytics team using Splunk was mentioned as needing to invest heavily in normalizing data to create a unified prioritization model.
  • Overwhelming Data Volume and Signal-to-Noise Ratio: The sheer volume of vulnerability data, described by one attendee as "11 billion CVEs around there" (likely an exaggeration for emphasis, but illustrative of the perceived scale), makes it challenging to identify truly relevant threats. Consumers feel compelled to ingest a "firehose" of data, only to then spend significant effort filtering out the "noise" to find the actionable "signal." The moderator emphasized the importance of receiving all data initially, rather than having producers pre-filter, to allow consumers to define their own signal.
  • Inefficient Data Access and Change Tracking: Accessing comprehensive historical data and tracking changes efficiently were significant pain points. The transition from legacy NVD 1.1 feeds (which allowed a single file dump) to NVD 2.0+ (requiring API iteration for backfilling) was highlighted as a regression. Similarly, while CVE lists are on GitHub and NVD has a change log, resolving GitHub diffs into a usable JSON format or integrating NVD's change log into automated systems is "incredibly hard."
  • Lack of a Unified Consumer Voice: A recurring theme was the absence of a structured mechanism for consumers to provide consolidated feedback to data producers like CVE.org. Changes to schemas or processes often occur based on the perspectives of a "select group of people," without sufficient input from the end-users who deal with the data daily. This leads to changes that, while well-intentioned, may not align with consumer needs or might even introduce new difficulties. The moderator specifically hoped for a "consumer union" to amplify these voices.

Technical Deep Dive

▶ Watch: Challenges in normalizing CVSS scores from multiple sources (5:00)

The technical challenges discussed revealed a reliance on a variety of tools and methodologies to cope with the existing state of vulnerability data. Many attendees, like the one using Bash and JQ, resort to scripting languages and command-line JSON processors for intricate data parsing. This involves writing extensive error checking, conditional logic (if statements, loops), and array manipulation to handle the nested structures and inconsistencies found in CVE and NVD JSON feeds. Common issues include checking for the presence of fields, validating expected character sets (e.g., for CWE identifiers), and managing scenarios where data might be represented as "N/A" text instead of a true missing value.

Data format evolution was a key topic, with the shift from XML to JSON being acknowledged as a positive step, yet not a panacea. The moderator specifically mentioned the NVD CPE format 5.1.1 version, introduced "a month or two ago," which provides a more structured definition of CPE data within the NVD schema. CVE.org data is also adopting this, though it's "a little bit less complete." This constant evolution of schemas, while necessary, creates integration overhead for consumers who must continually adapt their parsing logic.

The concept of CVSS (Common Vulnerability Scoring System) normalization was a significant technical discussion point. With CNA-assigned scores, NVD-enriched scores, and scanner-derived scores often differing, organizations face the complex task of reconciling these to create a single, authoritative risk metric. This often falls to "analytics teams" using tools like Splunk to ingest, process, and normalize the disparate data streams. The goal is to establish a consistent weighting and prioritization scheme, which is a non-trivial data engineering problem.

A key technical desire expressed by consumers was the need for flattened data files. As one attendee, Ben, articulated, "I spend most of my time taking nested arrayed data and making it flat" for analysis and visualization. Running "JQ after JQs after JQ" is inefficient. The suggestion for a "flat file pipeline" directly from data sources highlights a demand for data prepared in a more readily consumable format for traditional database and analytical tools, reducing the burden of complex transformations on individual consumers.

Furthermore, the discussion touched upon the technical limitations of current update mechanisms. Relying on GitHub diffs for CVE list changes or parsing NVD's change log is technically challenging. The proposed solution from an attendee, which garnered significant support, was an event stream model. Instead of consumers having to "overtly pull" or "manually managing your own" updates, an event-driven architecture would allow information to be "streamed" as it comes along, triggering specific processes for enrichment and analysis. This model would significantly reduce duplication of effort across organizations and enable more real-time responses to critical vulnerability updates. The moderator, while keen on receiving the "firehose" of initial data, saw the event stream as a valuable "and of that" solution, providing real-time notification on top of comprehensive data access.

The moderator also implicitly highlighted the technical underpinning of EPSS (Exploit Prediction Scoring System), which they help run. EPSS focuses on tracking CVE IDs to predict exploitability, making a conscious decision to anchor to these IDs despite their perceived flaws, precisely because it offers a consistent identifier in a highly fragmented threat intelligence landscape. This reinforces the technical necessity of a stable, albeit imperfect, identifier system for effective automated analysis.

Demo / Proof of Concept

▶ Watch: Discussion on balancing strictness for data input quality (7:50)

As a Birds of a Feather session, the format was an open discussion rather than a formal presentation with demonstrations. Therefore, no specific demo or proof of concept was shown during this talk. The session focused entirely on eliciting and sharing real-world challenges and potential solutions through participant dialogue.

Defensive Implications

▶ Watch: The continuous challenge of keeping up with vulnerability updates (8:30)

The candid discussion among vulnerability data consumers has profound defensive implications, highlighting critical gaps in current security operations and risk management strategies. The sheer volume of vulnerabilities, exacerbated by inconsistencies and quality issues, directly impacts an organization's ability to prioritize and remediate effectively. As one attendee noted, "There's 11 billion CVEs around there," making effective prioritization a Herculean task.

Defenders are currently forced to build elaborate, bespoke systems to compensate for the shortcomings of raw vulnerability data. This includes:

  • Manual Data Enrichment and Normalization: Security teams often spend significant resources manually reviewing CVE descriptions, reconciling disparate CVSS scores, and normalizing data from various scanners and threat intelligence feeds. This diverts valuable time from actual remediation or proactive security measures.
  • Delayed Response to Critical Threats: The difficulty in tracking real-time updates means that vulnerabilities escalating in severity (e.g., from medium to actively exploited critical) may not be identified and prioritized quickly enough. This leaves organizations exposed to emerging threats for longer periods, increasing their attack surface. Defenders often rely on commercial "threat feeds" or manual checks of major vendor advisories to try and fill this gap.
  • Inefficient Resource Allocation: Without a consistent and reliable way to assess the true risk of each vulnerability, organizations may misallocate resources, patching less critical issues while more dangerous ones remain unaddressed. This impacts the overall efficacy of vulnerability management programs.
  • Challenges in Threat Intelligence Integration: The lack of a standardized and high-quality identifier system beyond the CVE ID makes it difficult to map vulnerabilities to broader threat intelligence. The moderator, through their work with EPSS, emphasized the importance of supporting and improving CVE IDs precisely because they offer a common anchor for tracking "what the bad guys are doing," despite their imperfections.
  • The Need for "Secure by Design" Principles: A representative from CISA highlighted a broader defensive challenge: "There are too many vulnerabilities for us to patch." This underscores the need for a shift towards "secure by design" principles, where software is inherently safer, reducing the overwhelming volume of vulnerabilities that defenders must contend with in the first place. This long-term solution complements the immediate need for better vulnerability data consumption.

From a defensive standpoint, the key takeaway is that the current state of vulnerability data consumption forces defenders into a reactive, often inefficient, posture. To improve, there's a clear need for data producers to adopt a "product management" mindset, obsessively understanding the "downstream customers'" problems. This would lead to more actionable, higher-quality data that directly supports risk reduction and enables more proactive and efficient vulnerability management strategies. Solutions like standardized flat files, event-driven updates, and consistent data quality enforcement would empower defenders to better leverage automation and focus on strategic security initiatives rather than data wrangling.

Key Takeaways

  • Vulnerability Data Consumption is a Major Pain Point: Organizations face significant challenges in parsing, normalizing, and extracting actionable intelligence from the high volume of vulnerability data due to inconsistent formats, variable quality, and complex nested structures across sources like CVE.org and NVD.
  • Bespoke Solutions are the Norm: Many consumers are forced to develop custom, often manual, scripts and processes (e.g., using Bash and JQ) to handle data inconsistencies, track updates, and normalize varying CVSS scores, leading to duplicated effort and inefficient resource allocation.
  • Urgent Need for Data Quality and Standardization: There is a strong demand for higher minimum data quality standards, more accurate descriptions, consistent schema enforcement, and easier access to flattened, analysis-ready data formats to reduce the burden on consumers.
  • Event-Driven Updates are Preferred: The current polling model and reliance on GitHub diffs or complex change logs for updates are inefficient. An event stream model that pushes changes in real-time would significantly improve the ability of defenders to react promptly to escalating threats.
  • Consumer Voice Must Be Amplified: The CVE program and data producers need to adopt a "product management" approach, actively engaging with and listening to the "downstream customers" (vulnerability data consumers) to ensure that improvements address real-world operational challenges.
  • Support for CVE IDs is Crucial for Threat Intelligence: Despite data quality issues, supporting and improving the underlying CVE ID system remains vital for effective mapping of vulnerabilities to threat intelligence and enabling automated exploit prediction systems like EPSS.

About the Speaker(s)

This Birds of a Feather session was moderated by an experienced professional deeply involved in the domain of vulnerability data management. While the speaker's name was not provided in the conference metadata, their extensive background was evident throughout the discussion. They have spent many years working with various forms of vulnerability data, ranging from direct CVE records from CVE.org to enriched data from NVD and other supplemental sources.

The moderator actively participates in CVE working groups, demonstrating a commitment to improving the quality and usability of vulnerability data for the wider security community. Furthermore, they are involved in running the Exploit Prediction Scoring System (EPSS), an initiative that scores vulnerabilities daily to predict their likelihood of exploitation. This background provides them with a unique perspective, understanding both the intricacies of vulnerability data production and the practical challenges faced by its consumers in prioritizing and managing risk. Their role as a facilitator for this open discussion underscores their dedication to fostering collaboration and driving improvements within the vulnerability ecosystem.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

A legitimate BOF session at VulnCon that does exactly what a BOF should do: surface real operational pain from practitioners who live in the data daily. The complaints are genuine, the problems are real, and the room clearly had people who've actually written the JQ pipelines and hit the NVD API walls. It's not a research talk and shouldn't be graded like one — this is a practitioner roundtable, and on those terms it delivers honest signal about a broken ecosystem. The ceiling is low because the format prohibits depth, no concrete solutions were reached, and the transcript reads more like a well-organized grievance list than a session that moved anything forward. Useful for the VulnCon…

Heather Calloway (CISO) — SOLID

A useful working session for people who live inside vulnerability data pipelines — parsers, threat intel engineers, vuln management teams. The pain points are real and well-articulated: inconsistent data quality, schema churn, no clean update mechanism, no unified consumer voice. The CISA 'secure by design' thread is the most strategically interesting moment in the room. But this is a practitioner grievance session, not a governance conversation. It doesn't reach the people who need to decide whether to fund the fix, who owns the institutional accountability for the NVD degradation, or what the downstream organizational risk looks like when vuln programs are quietly failing because their…

→ Top-rated talks at CVE/FIRST VulnCon 2025

All talks from CVE/FIRST VulnCon 2025