Challenges in Open Source Software Identification

Martin (Red Hat)

CVE/FIRST VulnCon 2025 · Main Stage

Overview

In this insightful talk at VulnCon, Martin Seysen from Red Hat's Product Security Team tackled the complex and often overlooked challenges inherent in open source software identification. Seysen highlighted that accurately identifying software components is not merely a technical exercise but a foundational requirement for effective vulnerability management. Without a precise and standardized way to pinpoint exactly what software is being discussed, the entire ecosystem of security advisories, vulnerability databases, and remediation efforts becomes prone to ambiguity and inefficiency.

Watch on YouTube

Visual summary for Challenges in Open Source Software Identification by Martin
Visual summary for Challenges in Open Source Software Identification by Martin

Key moments

  1. 0:00 Speaker introduction and talk agenda overview
  2. 1:15 What software identity means in vulnerability management
  3. 2:15 Limitations of serial numbers, hashes, and URLs for identification
  4. 4:00 Introduction to CPE, PURL, Omnibore, and SWID standards
  5. 6:00 Explaining PURL's structure and attributes with examples
  6. 8:00 Tools and standards that widely adopt Package URL
  7. 8:40 Understanding PURL through the mailing address analogy

Challenges in Open Source Software Identification

Speakers: Martin Seysen, Product Security Team, Red Hat

Conference: VulnCon

YouTube: https://www.youtube.com/watch?v=oFtriEHe4IU

Overview

In this insightful talk at VulnCon, Martin Seysen from Red Hat's Product Security Team tackled the complex and often overlooked challenges inherent in open source software identification. Seysen highlighted that accurately identifying software components is not merely a technical exercise but a foundational requirement for effective vulnerability management. Without a precise and standardized way to pinpoint exactly what software is being discussed, the entire ecosystem of security advisories, vulnerability databases, and remediation efforts becomes prone to ambiguity and inefficiency.

Seysen’s presentation delved into the limitations of current identification methods and explored the strengths and weaknesses of prominent standards like Package URL (Pearl) and Common Platform Enumeration (CPE). He emphasized that while various schemes exist, their utility in the context of security data hinges on their ability to be easily retrievable, tool-supported, and consistently applied across diverse software ecosystems. The talk served as a critical call to action for the security community to improve the accuracy and interoperability of software identity data, ultimately enhancing our collective ability to understand and respond to vulnerabilities.

The core problem, as articulated by Seysen, is not just about assigning a label, but ensuring that label is universally understood, consistently applied, and rich enough in attributes to differentiate between potentially identical-sounding components across different distributions, packaging formats, and upstream sources. This talk is crucial for anyone involved in software supply chain security, vulnerability analysis, or the development of security tools that rely on accurate component identification.

Background

▶ Watch: Speaker introduction and talk agenda overview (0:00)

The genesis of the software identification problem lies in the inherent complexity and distributed nature of modern software development, particularly within the open source ecosystem. Unlike physical products with clear model and serial numbers, software components can be repackaged, modified, and distributed through myriad channels, making a singular, unambiguous identity elusive. Early attempts at identification, such as simple hashes, faltered due to the lack of reproducibility across different build environments, while URLs proved fragile and non-standardized for capturing essential software attributes. The need for a more robust and standardized approach became evident, especially with the proliferation of vulnerabilities and the increasing reliance on automated security tooling.

This challenge led to the development of several industry standards aimed at providing a structured way to identify software. Among the earliest and most widely adopted is Common Platform Enumeration (CPE), developed by NIST. Emerging around 2007, CPE was designed to identify IT systems, applications, and operating systems, evolving to its current version 2.3. It's heavily utilized by the National Vulnerability Database (NVD) and the CVE data set. A more recent, community-led initiative is Package URL (Pearl), with its first specification published in 2017. Pearl aims to provide a standardized, URL-like string for identifying and locating software artifacts across various package ecosystems. While other standards like Omnibore and SWID exist, Seysen noted that Omnibore's granularity often proves too complex for general vulnerability management, and SWID has seen limited adoption, making Pearl and CPE the primary focus for practical component identification in security contexts. The problem persists because despite these standards, inconsistencies in their application and a lack of interoperability create significant hurdles for accurate vulnerability assessment.

Key Findings

▶ Watch: Limitations of serial numbers, hashes, and URLs for identification (2:15)

Seysen's talk illuminated several critical findings regarding the state of open source software identification:

  1. Fragmented Adoption of Standards: While both Pearl and CPE offer structured identification, their adoption and usage patterns are highly fragmented. Pearl is widely embraced in the open source ecosystem by tools like SIFT and Guac, and supported by SBOMs and CSAF files. However, its formal support within the CVE record data set is minimal and poorly defined, with fewer than 10 records actually utilizing it. Conversely, CPE is pervasive in CVE records and NVD but suffers from a lack of consistent application, with many entities "minting their own CPEs" without central review, leading to duplicates and inconsistencies. Only Microsoft and MITRE consistently use the newer CPE applicability statements in CVE records.
  1. Data Inaccuracy and Ambiguity: The quality of identification data in crucial sources like CVE records is often insufficient. Seysen presented examples where version ranges were ambiguous (e.g., Django versions 5.1.0 to 5.1.4 affected, but what about other versions? Unspecified), or where version types were incorrect (e.g., 5.1 instead of 5.1.0 for semantic versioning). This imprecision directly impacts automated security tools, leading to false positives and negatives, and undermines the value of associated metadata like CVSS scores.
  1. The "Relationship Problem" vs. "Identity Problem": A core insight was the distinction between identifying a component and understanding its relationships. Seysen highlighted that components like the Django web framework might exist as a Python package (Pippi), source code on a website, or a GitHub release, all being "essentially the same thing" from a vulnerability perspective. However, when these are packaged into Linux distributions (e.g., Debian), they acquire new names, versioning schemes, and namespaces, becoming "different things" with specific build contexts. The challenge is not Pearl or CPE's inability to identify these unique instances, but rather the lack of standardized mechanisms to express the relationships between them (e.g., upstream, descendant, derived from).
  1. Red Hat's Hybrid Approach: Red Hat, a major open source distributor, exemplifies the practical challenges and solutions. They've evolved from simple RPM identification to managing 200+ products across diverse ecosystems. Their strategy involves using CPE solely for product identification (e.g., Red Hat Enterprise Linux 8) and Pearl for identifying specific components within those products. They also publish internal guidelines for Pearl usage to ensure consistency, recognizing the standard's inherent ambiguities. This demonstrates a real-world, large-scale attempt to navigate the complexities of software identification.
  1. OSV as a Model for Quality Data: The OSV (Open Source Vulnerability) database was presented as an example of relatively good data quality. As an aggregator, OSV programmatically processes and normalizes vulnerability information from various sources, ensuring proper version numbers and often including optional Pearl attributes for package identity. This suggests that programmatic generation and aggregation can significantly improve data consistency.

Technical Deep Dive

▶ Watch: Introduction to CPE, PURL, Omnibore, and SWID standards (4:00)

The talk provided a granular examination of the two leading software identification standards: Package URL (Pearl) and Common Platform Enumeration (CPE), along with their practical applications and inherent complexities.

Package URL (Pearl)

Pearl is presented as a standardized, URL-like approach for identifying and locating software artifacts. Each Pearl string is a valid URL composed of seven distinct attributes: type, namespace, name, version, qualifiers, and subpath, along with a checksum (which can be a qualifier or an explicit attribute).

  • Structure and Examples:
  • pkg:deb/debian/curl@7.74.0-1.3+deb11u1?arch=i386&distro=jessie identifies a Debian package named curl, version 7.74.0-1.3+deb11u1, built for i386 architecture on the jessie distribution.
  • pkg:pypi/django@2.2.1 identifies a Python package django, version 2.2.1, hosted on the Python Package Index (Pippi).
  • pkg:golang/github.com/genproto/googleapis/rpc/status@v0.0.0-20230525164840-28d496e792f4?subpath=status specifies a Go module, including its GitHub path, version, and a specific subpath within the module.
  • Adoption: Pearl is widely adopted in open source tooling, including SIFT (an SCA solution), Guac (a graph-based dependency visualization tool), and various commercial solutions. It's also supported by SBOMs (Software Bill of Materials) and CSAF (Common Security Advisory Framework) files. The CVE v5 schema has an asterisk next to Pearl support, indicating it's technically allowed but not robustly integrated.
  • Downsides and Challenges:
  • Package Type Dependency: Pearl primarily identifies packages within specific ecosystems. Adding new types is possible but often leads to "heated discussions" and potential deprecations.
  • URL Encoding: Since every Pearl is a valid URL, all its components must be URL-encoded. This makes Pearls "not look great" (e.g., pkg:pypi/django@2.2.1?%7B%22range%22%3A%22%3C%3D2.2.1%22%7D for a version range), hindering human readability and posing implementation challenges for tools.
  • Undefined Behavior: The specification has ambiguities. Qualifiers like architecture and distribution lack standardized values, leading to inconsistent inputs. Namespace values also suffer from this, with github.com/foo and github/foo potentially meaning the same thing but lacking a unified definition.
  • Inconsistent Field Rules: Different package types have varying interpretations for fields. For instance, Debian packages include the epoch directly in the version field, while RPMs use a separate epoch qualifier, leading to inconsistencies.
  • Version Range Specifier (Verse): Pearl addresses version ranges through its Verse sub-specification. Verse allows defining ranges as a list or with start/end versions. It also specifies how individual versions should be compared (e.g., Python packages have a specific comparison logic). When used with Pearl, the Verse specifier is typically added as a qualifier.

Common Platform Enumeration (CPE)

CPE is an older, NIST-developed standard for identifying IT systems, software, and packages. It's an open schema with a set of attributes, including part, vendor, product, version, update, edition, language, sw_edition, target_sw, target_hw, and other.

  • Structure and Examples:
  • cpe:/a:openssl:openssl:1.1.1u identifies OpenSSL version 1.1.1u.
  • cpe:/o:microsoft:windows_server:2019 identifies Windows Server 2019, potentially with target_hw qualifiers for specific architectures.
  • cpe:/a:gitlab:gitlab_enterprise_edition:16.0.0 identifies a specific edition of GitLab.
  • Seysen also showed "less nice" examples from the CVE data set, like cpe:/o:s:s:l, highlighting the lack of standardization in practice.
  • Downsides and Challenges:
  • Decentralized Minting: Although intended as a central dictionary by NIST, many entities "mint their own CPEs" for CVE records without much review. This leads to a chaotic dictionary with multiple CPEs identifying the same thing, hindering automation.
  • Limited Tool Adoption: Compared to Pearl, CPE has less widespread support in open source tools for generation, parsing, and storage.
  • Version Ranges (CPE Applicability Statements): CPE uses "applicability statements" (formerly "CPE configurations") within CVE records to define vulnerable version ranges. These can also specify dependencies (e.g., a component running on a specific Linux kernel version), somewhat replicating SBOM functionality. The CVE program recently published a quick start guide for these statements.

Interoperability and Data Usage

Seysen emphasized that Pearl and CPE are not directly transferable, though efforts exist (e.g., ScanOSS data set) to map between them. He stressed that good data should be interoperable, allowing conversion between formats.

  • CVE Records:
  • The affected object in CVE records identifies vulnerable software.
  • CPE applicability statements have been added, but only Microsoft and MITRE actively use them.
  • Pearl support is minimal and poorly defined, lacking a dedicated object.
  • Example of problematic CVE data: A CISA-enriched record for Django showed versionType: semantic but then listed 5.1 instead of 5.1.0, and only specified affected ranges without indicating unaffected versions, leading to ambiguity for programmatic consumption.
  • OSV Database:
  • OSV aggregates vulnerabilities and uses its own affected object schema.
  • It supports an optional Pearl attribute for package identity.
  • OSV's data is generally of "relatively good quality" due to programmatic creation and parsing, which helps maintain consistency in version numbers and ranges.

Red Hat's Approach

Red Hat's strategy reflects the practical integration of these standards:

  • CPE for Products: Used as a high-level product identifier (e.g., cpe:/o:redhat:enterprise_linux:8), centrally maintained.
  • Pearl for Components: Used for identifying specific components within products (e.g., a Python package or RPM module). These Pearls are namespaced to Red Hat (e.g., pkg:rpm/redhat/curl@7.76.1-23.el8).
  • Guidelines: Red Hat publishes specific guidelines for how it uses Pearl fields to ensure consistent generation and interpretation, aiming to overcome the standard's ambiguities.

Demo / Proof of Concept

▶ Watch: Tools and standards that widely adopt Package URL (8:00)

The talk did not include a live demonstration or proof of concept. Instead, Seysen focused on illustrative examples of Pearl and CPE structures, their usage in real-world data sets like CVE records and OSV, and the challenges encountered in their practical application.

Defensive Implications

▶ Watch: Understanding PURL through the mailing address analogy (8:40)

The insights from Seysen's talk carry significant implications for defenders seeking to bolster their security posture and streamline vulnerability management:

  1. Prioritize Accurate Component Identification: The fundamental takeaway is that accurate, standardized software identification is not a luxury but a necessity. Organizations must invest in tools and processes that correctly identify all software components, their versions, and their provenance. This forms the bedrock for effective vulnerability scanning, patch management, and risk assessment.
  1. Leverage SBOMs for Relationship Tracking: Seysen clearly articulated that while Pearl and CPE excel at identifying individual components, SBOMs (Software Bill of Materials) are the appropriate mechanism for tracking the complex relationships between them (e.g., upstream source, derived packages, dependencies). Defenders should push for SBOM generation and consumption throughout their software supply chain to gain a holistic view of component lineage and how vulnerabilities propagate. An SBOM can link a Pippi package to its Debian-packaged derivative, providing critical context for vulnerability applicability.
  1. Advocate for Better Data Quality and Standardization: Defenders, as consumers of vulnerability data, have a vested interest in improving the quality of information from sources like CVE records. This means advocating for:
  • Proper Pearl Support: Pushing for a dedicated Pearl object in the CVE schema, rather than ad-hoc usage in version fields.
  • Consistent CPE Usage: Encouraging CNAs to adhere to standardized CPE minting and ensuring central review to prevent data pollution.
  • Clear Version Ranges: Demanding that vulnerability data explicitly state both affected and unaffected version ranges to eliminate ambiguity for automated tools.
  • Interoperability: Supporting efforts to convert between Pearl and CPE formats, allowing defenders to consume data in their preferred standard.
  1. Adopt Internal Guidelines for Identification: Following Red Hat's example, organizations that build and distribute software should establish and publish internal guidelines for how they generate Pearl or CPE identifiers. This ensures consistency within their own ecosystem and facilitates better communication with downstream consumers.
  1. Utilize Tools that Support Robust Identification: Implement security tools (SCA, VEX generators, vulnerability scanners) that natively understand and correctly parse Pearl, CPE, and SBOMs. Tools like CV lint (mentioned in Q&A) are crucial for validating the quality of CVE data before consumption.
  1. Participate in Community Efforts: The Pearl specification, in particular, "needs a lot of love." Defenders and developers should contribute to these open standards to address ambiguities, improve consistency, and ensure they meet real-world needs.

By focusing on these defensive implications, organizations can move beyond reactive vulnerability management to a more proactive, data-driven, and robust security posture.

Key Takeaways

  • Accurate software identification is foundational: Without it, vulnerability management becomes ambiguous and inefficient.
  • Pearl and CPE are primary standards: Pearl (community-led, widely adopted in OS tools, supports 7 attributes) and CPE (NIST-developed, used by NVD/CVEs) are the leading contenders.
  • Both standards have limitations: Pearl struggles with URL encoding, undefined behavior, and inconsistent field rules; CPE suffers from decentralized minting and data inconsistencies in CVE records.
  • Relationships require SBOMs: Pearl and CPE identify components, but SBOMs are critical for tracking complex upstream/downstream relationships and provenance.
  • Data quality in CVE records needs significant improvement: Ambiguous version ranges and inconsistent usage of identifiers lead to false positives/negatives, diminishing the value of security data.
  • Red Hat's hybrid strategy offers a practical model: Using CPE for products and Pearl for components, coupled with internal guidelines, helps manage identification in complex ecosystems.
  • Community involvement is crucial: Contributing to Pearl's specification and advocating for better data quality from CNAs are essential for industry-wide improvement.

About the Speaker(s)

Martin Seysen is a seasoned expert in product security, currently serving on the Product Security Team at Red Hat. With a remarkable 13 years of experience in this role, Seysen brings deep practical knowledge of the challenges and intricacies involved in securing open source software and managing vulnerabilities across a vast product portfolio. His work at Red Hat, a prominent contributor to the open source community, positions him at the forefront of developing and implementing strategies for robust software identification and vulnerability management. His insights are informed by extensive real-world application of the standards and practices discussed in his talk.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Seysen delivers a competent, practitioner-grounded survey of open source software identification — PURL vs CPE, data quality gaps in CVE records, and Red Hat's hybrid approach. This is a legitimate domain problem that doesn't get enough stage time, and his 13 years doing this at scale gives him real credibility. The talk isn't breaking new ground — most of what's here is observable by anyone who's spent time with NVD data and the PURL spec — but it synthesizes the pain points clearly and offers a concrete real-world model. Won't be memorable in a year, but it belongs at VulnCon.

Heather Calloway (CISO) — SOLID

Martin Seysen delivers a technically credible and operationally grounded examination of a real, persistent problem in vulnerability management — the inability to reliably identify what software you're actually talking about. The talk is honest about limitations, grounded in Red Hat's operational experience, and surfaces a meaningful distinction between identity and relationship that has practical consequences for how organizations build their SBOM and vulnerability tracking programs. It earns its place at a conference like VulnCon. But it stays inside the practitioner lane and never fully translates upward — the governance dimension of this problem, which is substantial, goes unaddressed…

→ Top-rated talks at CVE/FIRST VulnCon 2025

All talks from CVE/FIRST VulnCon 2025