Let’s Talk About Fitness for Purpose: Comparing and Contrasting the CVE List with OSV.dev
Andrew (Google)
CVE/FIRST VulnCon 2025 · Main Stage
Overview
Andrew, a member of Google's open source security team, delivered a compelling talk at VulnCon, expanding on his previous discussions regarding the challenges in vulnerability metadata quality. His presentation, titled "Let’s Talk About Fitness for Purpose: Comparing and Contrasting the CVE List with OSV.dev," critically examines whether the Common Vulnerabilities and Exposures (CVE) program, in its current state, adequately serves the needs of modern defenders. Andrew argues that while the CVE program was groundbreaking 25 years ago, its mission statement and current implementation have not evolved sufficiently to cope with the exponential growth in vulnerabilities, leading to significant data quality issues.

Key moments
- 0:00 Speaker introduction and OSV project background
- 2:30 Speaker's journey understanding CVE/NVD ecosystem
- 4:00 Defining 'Fitness for Purpose' in vulnerability metadata
- 5:00 OSV.dev: Origins, purpose, and current capabilities
- 8:00 Critical comparison of CVE and OSV mission statements
Let’s Talk About Fitness for Purpose: Comparing and Contrasting the CVE List with OSV.dev
Speakers: Andrew, Open Source Security Team, Google
Conference: VulnCon
YouTube: https://www.youtube.com/watch?v=JfKxnVT9taQ
Overview
Andrew, a member of Google's open source security team, delivered a compelling talk at VulnCon, expanding on his previous discussions regarding the challenges in vulnerability metadata quality. His presentation, titled "Let’s Talk About Fitness for Purpose: Comparing and Contrasting the CVE List with OSV.dev," critically examines whether the Common Vulnerabilities and Exposures (CVE) program, in its current state, adequately serves the needs of modern defenders. Andrew argues that while the CVE program was groundbreaking 25 years ago, its mission statement and current implementation have not evolved sufficiently to cope with the exponential growth in vulnerabilities, leading to significant data quality issues.
The core of Andrew's argument revolves around the concept of "fitness for purpose." He posits that the ultimate goal of vulnerability metadata is to enable defenders and end-users to quickly identify impact and remediate issues. However, the current state of CVE records often falls short of this objective, particularly due to a lack of machine-readable, consistent, and complete data. In contrast, Andrew presents OSV.dev, a project he has worked on extensively, as an example of a system designed from the ground up to facilitate precise, automated identification and management of open-source vulnerabilities.
This article delves into Andrew's observations, comparing and contrasting the two systems, highlighting systemic barriers to achieving high-quality vulnerability data, and proposing concrete steps for improvement. It underscores the urgent need for the CVE program to adapt its standards and practices to meet the demands of an increasingly automated and high-volume security landscape, emphasizing that the absence of high-quality, automation-friendly data fundamentally cripples effective vulnerability management at scale.
Background
▶ Watch: Speaker introduction and OSV project background (0:00)
Andrew's journey into the intricacies of vulnerability metadata began with a practical, albeit challenging, project at Google. His task was to develop code to convert CVEs from the National Vulnerability Database (NVD) – using the Common Platform Enumeration (CPE) dictionary – to determine the corresponding Git repository for the source code, and then generate an OSV record. The primary intended benefit of this work was to gain comprehensive coverage of C and C++ vulnerabilities, which historically lacked a centralized source. For these languages, the most effective way to scan for vulnerabilities in source code is by commit hash, a level of precision that traditional CVE records often lacked. His work, which has shown pleasingly consistent conversion rates, is detailed further on the OSV.dev blog.
Initially a newcomer to the state of vulnerability management, Andrew quickly learned that the NVD, while widely used, is distinct from the underlying CVE list. This foundational misunderstanding, shared by many, highlighted the complexity and opacity of the vulnerability ecosystem. His previous talk at VulnCon focused on systemic challenges in data quality, and subsequent engagement with the CVE Quality Working Group (QWG) further deepened his understanding of the problem.
The OSV.dev project itself emerged from the necessity to address specific shortcomings in vulnerability reporting for open-source software. Created four years ago, OSV was catalyzed by OSS-Fuzz, Google's continuous fuzzing service, which was discovering numerous vulnerabilities in open-source projects. At the time, CVSS 4.0, the prevailing standard, proved inadequate for expressing these vulnerabilities with the precision needed for open-source contexts, especially regarding specific affected versions and fix commits. Furthermore, the idea of automatically submitting CVEs at scale was met with significant apprehension. OSV.dev thus provided a parallel, automated mechanism to express these findings. Upon the formation of the OpenSSF, Google donated the OSV schema as intellectual property, and continues to sponsor the infrastructure for OSV.dev, which serves as a central aggregator for vulnerability databases publishing records in the OSV format. The platform now includes a Software Composition Analysis (SCA) library, extending its capability beyond source code to container image scanning, all freely available and open source.
The fundamental problem underpinning this discussion is the overwhelming scale of vulnerabilities. Andrew notes that the volume of vulnerabilities has consistently increased year-over-year for 25 years, a trend that shows no signs of abating. This scale is a challenge for everyone: vendors, defenders, and analysts. The NVD, with its reliance on human analysis, has become a "casualty of that scale," managing to analyze only about half of the CVEs from last year. This bottleneck underscores a critical need for automation. While efforts like CISA's vulnerability enrichment and projects like Ryan's aim to optimize human analysis, there's an inherent limit to how much human effort can scale. The only viable path forward, Andrew asserts, is to "send in the machines" – to reason about vulnerabilities programmatically, leveraging AI (minus the hallucinations) to assist humans with higher-order tasks like classification, validation, and prioritization. This necessitates machine-readable, accurate, and complete vulnerability records, a standard that the CVE program, despite its recent JSON format iterations (CVE JSON 5, influenced by OSV), often fails to meet in practice.
Key Findings
▶ Watch: Speaker's journey understanding CVE/NVD ecosystem (2:30)
Andrew's talk meticulously dissects the current landscape of vulnerability metadata, revealing several critical findings that underscore the "fitness for purpose" challenge facing the CVE program:
1. Mission Statement Disparity:
The CVE program's mission statement has remained unchanged for 25 years. Andrew argues that this lack of evolution, particularly the absence of an explicit "why" that reflects today's security landscape, is problematic. The "why" of 1999, when the CVE list had only around 1,500 vulnerabilities, is vastly different from the "why" of 2024. In contrast, OSV.dev's mission statement is unequivocally focused on risk reduction for open-source software users, with a clear, albeit somewhat narrow, target audience. This explicit goal provides a clear yardstick for evaluating the program's activities.
2. Stark Differences in Scale and Scope:
The sheer volume of vulnerabilities managed by each program highlights a significant disparity:
- CVE List (2024 data): Approximately 37,381 vulnerabilities from 307 distinct CNAs (CVE Numbering Authorities). Over 25 years, the list has grown significantly, but its scope is broad, covering all types of software and hardware.
- OSV.dev (2024 data): Nearly 83,000 vulnerabilities from 21 different home databases. Despite being only four years old, OSV.dev's volume for a single year is more than double that of the CVE list. While OSV.dev has a narrower scope (open-source related), it demonstrates a higher velocity of vulnerability aggregation. Crucially, almost 10,000 of OSV.dev's 2024 records had a direct alias to a CVE ID, and another 58,000 had some form of cross-reference, indicating a substantial number of unique open-source vulnerabilities tracked by OSV.dev that may not yet have a CVE ID.
3. Fragmented Sources and "Unique Snowflakes":
As of late last year, there were 447 CNAs contributing to the CVE list, with 125 declaring some relation to open source. OSV.dev, by comparison, has 22 home databases (the moral equivalent of CNAs). The proliferation of CNAs, which Andrew colorfully describes as "unique snowflakes," creates a massive challenge for data consistency. It's not feasible for downstream consumers to maintain per-CNA heuristics for machine processing, leading to a "losing battle." The remark, "It's just a matter of getting the CNAs to do it right," encapsulates the overwhelming task of herding an ever-growing number of diverse entities.
4. Data Quality Crisis: Theoretical vs. Practical Machine Readability:
Despite the launch of a richer JSON format in 2018 and the CVE JSON 5 schema in 2022 (which theoretically offers significant machine readability and was influenced by the OSV schema), the practical implementation falls short. Andrew highlights a critical "rub": "Just because it can doesn't mean it is." He cites Jerry Gamblin's earlier point and his own observations:
- Minimal Requirements: 25 years later, the only required elements for a CVE record remain an ID, a description, and perhaps one or more references – the same as when the volume was 20 times lower.
- Usability Deficiencies: A staggering 15% of CVEs last year had no attempt at usable data in the affected field. Quantifying the usability of present data is difficult precisely because it's not machine-readable.
- Inconsistent Versioning: Eyeballing records reveals common issues like versions claiming to be SemVer but clearly not (e.g., lacking three bits, containing asterisks). This inconsistent use of the record format severely undermines the potential of machine readability, hampering meaningful vulnerability management at scale. Andrew stresses: "There is no remediation at scale without first detection at scale," and current CVE data often fails the detection test.
5. Misaligned Incentives and Voluntary Compliance:
Andrew identifies a significant misalignment of incentives contributing to the data quality problem:
- Supply Side (CNAs/Vendors): Product vendors often focus on human-oriented descriptions for individual CVEs to meet regulatory and contractual obligations. A basic, 1999-style CVE may be sufficient to "tick that box." They may even prefer a false negative over a false positive in their products, at least until customers are compromised.
- External Pressures: Government compliance (e.g., CISA's Secure by Design pledge), customer contractual obligations, and external vulnerability researchers seeking credit all compel vendors to participate in the CVE program. However, these pressures don't necessarily incentivize high-quality machine-readable data.
- Voluntary Programs: CISA's Secure by Design pledge, while a good start, is entirely voluntary and, in its current form, may be inadequate for achieving machine-readable detection and prioritization. It covers only 16% of the CVE list for 2024. The Enrichment Recognition List (ERL), another voluntary positive reinforcement mechanism, is unlikely to motivate C-level executives unless customers demand it.
6. Complexities of CNA Types and Their Impact:
Andrew categorizes CNAs and analyzes their inherent incentives and challenges:
- Vendor CNAs: Represent the "lion's share" of CVEs. They have the best opportunity for high-fidelity communication between finder, CNA, and vulnerability owner (often in-house), leading to potentially high-quality records.
- Open Source CNAs: Often operate like vendors and may even be vendors themselves. They also have good context on fixes. However, complexities arise with Linux distributions redistributing vulnerable open-source software, making ownership and statement generation challenging.
- Bug Bounty Providers: Operate at "arm's length" from the vulnerability subject owner. They typically issue CVEs based on finder-provided information, often lacking context on fix details. This added friction creates a barrier to timely, high-quality fix information in CVE records.
- Researcher CNAs: Similar to bug bounty providers, they face challenges in obtaining high-quality fix information due to coordination overheads and a lack of direct ownership. Andrew notes a "blurry line" around their scope and believes enforcement is largely based on the honor system.
These findings collectively paint a picture of a CVE program struggling to adapt its 25-year-old framework to the demands of modern security, where automation and precision are paramount for effective defense.
Technical Deep Dive
▶ Watch: Defining 'Fitness for Purpose' in vulnerability metadata (4:00)
The technical heart of Andrew's talk lies in the comparison of the underlying data structures and validation mechanisms of CVE and OSV.dev, particularly highlighting how OSV.dev is engineered for "fitness for purpose" in an automated world.
At its core, OSV.dev utilizes the OSV schema, a JSON-based format designed specifically to facilitate precise identification of vulnerabilities in open-source software. This schema emerged from the practical need to express vulnerability findings from tools like OSS-Fuzz with greater granularity than traditional CVE formats. A key design principle of the OSV schema is its ability to represent commit-level vulnerability metadata, which is crucial for C and C++ projects where scanning by Git commit hash is the most effective method for detection. This contrasts sharply with the often broader and less precise version ranges found in many CVE records.
The OSV schema had a direct influence on the development of the CVE JSON 5 schema, which introduced significant theoretical capabilities for machine readability into the CVE program. This theoretical interoperability is a positive step, suggesting a shared understanding of what modern vulnerability data should look like. However, Andrew's critical observation is that "just because it can doesn't mean it is." The inconsistent and often incomplete population of fields within CVE JSON 5 records undermines its theoretical potential.
To ensure data quality, OSV.dev employs a multi-layered validation approach:
- JSON Schema Validation: Every OSV record must conform to the JSON schema for structural validity. This ensures that the record is well-formed and adheres to the defined data types and field presence requirements. Andrew notes that "every possible validation" is jammed into the schema, allowing record publishers to pre-validate structural aspects before submission.
- Record Linter: Recognizing the limitations of JSON schema (which cannot, for example, validate logical relationships like an introduced version being earlier than a fixed version), OSV.dev is developing a record linter. This tool, inspired by Martin's CVE lint project, performs deeper, semantic validation. It checks for logical consistency, accuracy, and adherence to "high quality" properties, preventing "schema compliant but utterly nonsensical records."
Andrew defines the minimum properties of a high-quality OSV record as being:
- Valid: Structurally correct according to the JSON schema and semantically sound according to the linter.
- Precise: At publication time, it refers to correctly ordered and existing versions, and verifiably existing references (e.g., specific commits).
- Identifiable: Contains sufficient information to uniquely identify the vulnerable component and the nature of the vulnerability.
A key technical distinction highlighted is the inadequacy of CPE (Common Platform Enumeration) for open-source vulnerability management. While CPEs might suffice for commercial software or hardware appliances, they struggle with the dynamic, modular nature of open-source projects, which often lack the rigid versioning and packaging that CPEs rely upon. Andrew references Martin's talk, which further elaborated on these CPE limitations. OSV.dev, by focusing on ecosystems, package managers, and commit hashes, provides a more granular and accurate way to track open-source vulnerabilities.
The entire OSV.dev ecosystem, including its schema, tooling, and database infrastructure, is freely available and open source. This transparency allows anyone to "clone it, look at it, poke at it, understand how it's working, [and] contribute to it," fostering a collaborative environment for improving vulnerability data quality. This technical approach contrasts with the more opaque and federated nature of the CVE program, which struggles to enforce consistent data quality across its multitude of CNAs.
Demo / Proof of Concept
▶ Watch: OSV.dev: Origins, purpose, and current capabilities (5:00)
Andrew's talk at VulnCon did not feature a live technical demonstration or a new proof of concept. Instead, he referenced his prior work and the ongoing development within the OSV.dev project. Specifically, he mentioned the code he wrote to convert CVEs from the NVD, using CPE data, into OSV records with associated Git repository and commit information. This conversion process itself serves as a practical proof of concept for generating high-quality, machine-readable vulnerability data from existing, albeit often imperfect, sources. He also highlighted the development of the OSV record linter, inspired by CVE lint, as an internal tool to enforce and improve data quality within the OSV ecosystem. While no direct demo was shown, the entire OSV.dev project and its associated tooling are open source and publicly available for inspection and use.
Defensive Implications
▶ Watch: Critical comparison of CVE and OSV mission statements (8:00)
The insights presented by Andrew carry profound implications for defenders and the broader cybersecurity ecosystem. The central message is clear: effective vulnerability management at scale is impossible without high-quality, machine-readable, and consistently applied vulnerability metadata.
- Demand Higher Standards from Vendors: Defenders, as consumers of software and vulnerability intelligence, must become more sophisticated in their procurement processes. Andrew strongly advocates for incorporating a vendor's vulnerability management track record and the quality of their CVE records into purchasing decisions. Rather than CVEs being seen as a mark of shame, they should be evidence of a mature lifecycle management. Customers should "demand more of their vendors," have "awkward conversations when there's contract renewals," and foster "competitive rivalry" where vendors strive for the best data quality, not a race to the bottom.
- Embrace Automation for Detection and Remediation: Given the ever-increasing volume of vulnerabilities (the "drowning in vulnerabilities" scenario), manual analysis is unsustainable. Defenders must prioritize tools and processes that enable automatable detection and remediation. This requires vulnerability records that are not just theoretically machine-readable but are consistently populated with precise, actionable data such such as accurate affected product information (including CPEs, PURLs, and commit hashes). The less systematic barriers to automated response, the better for everyone.
- Prioritize Fix Information: The talk highlighted that many CVEs, especially those originating from bug bounty or researcher CNAs, often lack timely and precise fix information. Defenders need to know "what to do about it and how urgently." This means demanding that CVE records include clear details on the introduced and fixed versions, or even specific commit hashes, to enable accurate prioritization and efficient patching.
- Leverage OSV.dev for Open Source: For organizations heavily reliant on open-source software, OSV.dev offers a robust solution designed for precision and automation. Its focus on commit-level metadata, ecosystem-specific tracking, and upcoming container image scanning capabilities provides a powerful toolset for identifying and managing open-source vulnerabilities more effectively than the current CVE list often allows. Its open-source nature means defenders can inspect, contribute to, and trust its mechanisms.
- Influence Program Leadership and CNAs: Andrew issues a direct challenge to the CVE program's leadership. CISA, as the funding body, has "a lot of power to attach some conditions on that funding" to enforce higher data quality standards. The CVE Board and CNA LRs (Last Resort CNAs) must "sign up to doing a better job" and serve as models of desired behavior. Defenders should encourage this strong leadership to set clear standards and incrementally bring all program participants up to those standards.
- Empower Vulnerability Researchers: Researchers, who often seek CVE IDs for credit and recognition, have leverage when interacting with CNAs. Andrew suggests researchers should ensure the CNA issuing the CVE for their discovery provides a "full end-to-end sort of story" with comprehensive fix information. This ensures that the "people that are impacted by the thing that you're getting all the fame for is is is trivially fixable."
In essence, defenders can no longer passively accept inconsistent and incomplete vulnerability data. They must actively advocate for, demand, and implement solutions that align with the need for scalable, automated, and precise vulnerability management, leveraging systems like OSV.dev and pushing for fundamental improvements within the CVE program.
Key Takeaways
- Vulnerability data quality is paramount for effective defense at scale. The sheer volume of new vulnerabilities necessitates machine-readable, precise, and consistently applied metadata for automated detection, prioritization, and remediation.
- The CVE program's current state is not "fit for purpose" for modern challenges. Its 25-year-old mission statement and inconsistent data quality, despite theoretical improvements like CVE JSON 5, severely hamper its utility for automated vulnerability management.
- Automation is the only viable solution to the escalating volume of vulnerabilities. Manual analysis is unsustainable, making programmatic reasoning about vulnerabilities essential, which in turn demands high-quality, machine-friendly data.
- Misaligned incentives among CNAs are a major barrier to improving CVE data quality. Product vendors often prioritize human-readable descriptions for compliance over precise machine-readable data, while bug bounty and researcher CNAs struggle with coordination overheads to provide complete fix information.
- OSV.dev provides a compelling model for high-quality, automation-friendly vulnerability metadata. Designed for open source, it emphasizes precision (e.g., commit hashes), robust validation (JSON schema and linters), and a clear mission focused on risk reduction, demonstrating what "fitness for purpose" looks like.
- Driving improvements requires strong leadership and collective action. CISA, the CVE Board, CNAs, vendors, and even vulnerability researchers must collectively commit to higher standards, demand better data, and foster a competitive environment where data quality is a differentiator, not a neglected afterthought.
About the Speaker(s)
Andrew is a distinguished member of Google's open source security team, based in Brisbane, Australia. With a career spanning two decades at Google, he has dedicated the last couple of years to working on OSV.dev, a project he describes as the highlight of his professional journey. Andrew's passion lies in data quality, consistency, and completeness, particularly in the context of vulnerability metadata. Despite being a newcomer to the intricacies of the CVE ecosystem when he began his work on OSV, he has since gained a profound understanding of its challenges and potential. He is an active participant in relevant working groups and an advocate for systemic improvements in how vulnerabilities are disclosed and managed.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
A competent, well-structured policy/infrastructure talk from someone who has clearly done the work. Andrew knows the OSV.dev codebase intimately and has credible grievances about CVE data quality. The core argument — that CVE's 25-year-old mission and voluntary compliance model can't survive the current volume of vulnerabilities — is correct and worth repeating at VulnCon specifically. But it is, fundamentally, a talk we've largely heard before. The CVE-vs-OSV comparison, the NVD backlog, the 'automation or bust' thesis — these are established positions in the vuln management discourse, not novel contributions. The audience most likely to benefit is mid-level security engineers and vuln…
Heather Calloway (CISO) — SOLID
A technically credible, well-argued critique of the CVE program's fitness for modern automated vulnerability management. Andrew knows the data, knows the ecosystem, and makes a clear case that the CVE program's mission and minimum requirements have not kept pace with volume or the demands of automation. The talk is most valuable for security engineers and vulnerability management practitioners. Where it falls short, from my seat, is accountability. The governance failures here — CISA's funding relationship, the CVE Board's standard-setting authority, the procurement lever that customers haveninely ignored for years — are named but not pressed. This is a program diagnosis without an…