Ask Not Whether CVSSv3.1 and v4 Scores are Inconsistent, But What Can You Do About It
Monen (Assistant Professor · VU Amsterdam), Chisan (PhD Student · VU Amsterdam)
CVE/FIRST VulnCon 2025 · Main Stage
Overview
This talk, presented by Monan and Chisan from VU Amsterdam, delves into the critical issue of inconsistencies between Common Vulnerability Scoring System (CVSS) versions 3.1 and 4.0. As organizations increasingly rely on CVSS scores for vulnerability management and prioritization, discrepancies between different versions or even between different scoring entities can lead to significant operational challenges. The researchers highlight that while CVSSv4.0 aims to offer more refined guidance, its introduction has inadvertently surfaced new layers of scoring variability.

Key moments
- 0:00 Introduction: Inconsistency between CVSS v3.1 and v4
- 2:00 Root Causes of CVSS Rating Inconsistencies
- 4:00 Concrete Example: Disagreement on 'User Interaction' requirement
- 4:50 Introducing CVSS v4 and Key Research Questions
- 5:45 Methodology: Using ChatGPT to Generate CVSS Metrics
- 6:30 ChatGPT's Inconsistencies in CVSS Metric Generation
- 9:00 Proposed Solution: Consistency Rules for CVSS Rating
Ask Not Whether CVSSv3.1 and v4 Scores are Inconsistent, But What Can You Do About It
Speakers: Monan, Assistant Professor, VU Amsterdam; Chisan, PhD Student, VU Amsterdam
Conference: VulnCon
YouTube: https://www.youtube.com/watch?v=OyV9GkBZRNw
Overview
This talk, presented by Monan and Chisan from VU Amsterdam, delves into the critical issue of inconsistencies between Common Vulnerability Scoring System (CVSS) versions 3.1 and 4.0. As organizations increasingly rely on CVSS scores for vulnerability management and prioritization, discrepancies between different versions or even between different scoring entities can lead to significant operational challenges. The researchers highlight that while CVSSv4.0 aims to offer more refined guidance, its introduction has inadvertently surfaced new layers of scoring variability.
The core of the presentation explores the underlying reasons for these inconsistencies, attributing them primarily to human subjectivity, varying interpretations of metric definitions, and insufficient information during the scoring process. Through extensive analysis of public and private datasets, including those generated by advanced Large Language Models (LLMs) like ChatGPT, the speakers demonstrate the tangible impact of these discrepancies on vulnerability severity ratings. Their work culminates in the proposal of a set of "consistency rules" designed to guide analysts in simultaneously rating vulnerabilities across both CVSS versions, thereby mitigating potential errors and fostering more reliable security prioritization.
The research is particularly timely given the recent release of CVSSv4.0 in November 2023 and its early adoption phase. With many organizations still primarily using CVSSv3.1, and some vendors beginning to provide dual ratings, understanding and addressing these inconsistencies is paramount. The insights offered by Monan and Chisan provide a scientific perspective on a pervasive problem, offering actionable strategies for security professionals to navigate the complexities of vulnerability scoring in a multi-version landscape, ultimately aiming for more accurate and consistent risk assessments.
Background
▶ Watch: Introduction: Inconsistency between CVSS v3.1 and v4 (0:00)
The challenge of inconsistent vulnerability scoring is not new to the cybersecurity landscape. Prior research, dating back to 2022, has repeatedly highlighted discrepancies in CVSSv3 scores between different Common Vulnerabilities and Exposures (CVE) Numbering Authorities (CNAs), such as Red Hat and the National Vulnerability Database (NVD). Studies have shown instances where the same vulnerability received significantly different scores or even severity levels (e.g., critical vs. high vs. medium) depending on the scoring entity. One notable finding indicated that CVSSv3 scores could differ by as much as 20% between two CNAs, underscoring a systemic issue rather than isolated incidents.
The root causes of these inconsistencies are multi-faceted, often stemming from the human element inherent in the rating process. Human error is a significant factor, where mistakes can be introduced during manual assessment. More profoundly, subjective understandings of metric definitions often lead to divergent interpretations. What one analyst considers "required" for user interaction, another might deem "not required," even when evaluating the exact same vulnerability description. This was vividly illustrated with a CVE example where a single sentence describing user access to an injected page led to conflicting assessments regarding the User Interaction (UI) metric, with NVD rating it as 99% required, while other assessors disagreed. Such fundamental disagreements can drastically alter the final severity score and, consequently, an organization's prioritization strategy. Furthermore, insufficient information within CVE descriptions can force assessors to make assumptions, contributing to variability.
The release of CVSSv4.0 in November 2023 introduced a new layer of complexity. While each new CVSS version is designed to refine and improve upon its predecessor, offering more detailed guidance, it also raises critical questions about backward compatibility and consistency. Specifically, how do the refinements in CVSSv4.0 impact existing CVSSv3.1 ratings? Do vulnerability ratings remain consistent across these versions, or does the evolution of the standard introduce new avenues for discrepancies? This forms the central investigative thrust of the research, seeking to understand the nature of these inter-version inconsistencies and propose practical solutions.
Key Findings
▶ Watch: Concrete Example: Disagreement on 'User Interaction' requirement (4:00)
The research uncovered significant inconsistencies in CVSS scoring across versions 3.1 and 4.0, observed in both automated (LLM-generated) and human-rated datasets. A preliminary experiment involved using advanced LLMs, specifically ChatGPT, to generate CVSSv3.1 and v4 metric vectors. The prompt was meticulously designed to list all metric names for both versions to ensure correct format generation, as initial attempts often led to irrelevant or omitted metrics. By comparing ChatGPT's predictions against ground truth data, the researchers found that metrics such as Attack Vector (AV), Privileges Required (PR), and Impact metrics (Confidentiality, Integrity, Availability) performed with relatively low weighted F1 scores, indicating poor consistency.
A concrete example of LLM inconsistency involved a vulnerability where ChatGPT rated User Interaction (UI) as "required" for both versions. However, while CVSSv3.1's Scope was considered "changed" (implying impact beyond the vulnerable component), ChatGPT assigned "no impact" for the corresponding v4 metrics, demonstrating a clear logical breakdown in its understanding of cross-version metric relationships.
Moving to real-world data, the researchers analyzed datasets from various sources, including public repositories like cv.org and vendor-specific databases like Vondb. Vondb, for instance, has started providing CVSSv4 ratings alongside v3.1. A striking example from Vondb showed a CVE where v3.1 was rated with User Interaction "required," but v4 was rated "not required" for the very same vulnerability. This direct contradiction in a core metric highlights the practical inconsistencies faced by organizations.
Across two public datasets, comprising over 2,000 CVEs, the application of their proposed consistency rules revealed that almost all metrics could remain consistent, except for User Interaction (UI). The UI metric exhibited a relatively high inconsistency rate of 22.1% of CV entries, where a vulnerability rated "required" in v3.1 was rated "non" (not required) in v4. This suggests a common point of disagreement or misinterpretation among analysts.
Further analysis of a private dataset obtained from the CVSS Implementation Forum (CV-SSI), containing over 6,000 CVEs, corroborated these findings and revealed additional challenges. Even for straightforward, one-to-one mapped metrics like Attack Vector (AV) and Privileges Required (PR), some inconsistencies were observed, albeit at lower rates (e.g., 0.6% discrepancy for a private company). For UI, three inconsistent combinations collectively accounted for 3.1% of entries, still a significant number given the volume. One company (Company A) showed a particularly high inconsistency rate of 26.7% related to impact metrics. Even after applying internal "hidden agreement rules" within this company, 3.3% of entries still exhibited impact-related inconsistencies, emphasizing the difficulty of achieving full alignment without explicit cross-version guidance. These findings underscore that inconsistencies are not merely theoretical but are prevalent in real-world vulnerability scoring practices, affecting even well-resourced organizations.
Technical Deep Dive
▶ Watch: Introducing CVSS v4 and Key Research Questions (4:50)
To address the observed inconsistencies, the researchers propose a structured approach centered around consistency rules. The fundamental idea is to encourage simultaneous rating of vulnerabilities for both CVSSv3.1 and CVSSv4.0. If an initial rating for one version is made, these rules then guide the rating of the other version, or help quickly spot and correct discrepancies if one version is already available. This proactive approach aims to save time and prevent downstream issues arising from inconsistent scores.
The consistency rules are derived from a thorough review of the official CVSS documentation for both versions, identifying three primary types of mapping relationships between their metrics:
- One-to-One Mapping:
This category includes metrics whose definitions and possible values remain unchanged between CVSSv3.1 and CVSSv4.0. For these metrics, the evaluation should ideally be identical across both versions.
- Attack Vector (AV): Whether an attack is local, network, adjacent, or physical.
- Privileges Required (PR): The level of privileges an attacker needs (None, Low, High).
As these metrics have not been altered, a "Network" AV in v3.1 should correspond to "Network" in v4, and "Low" PR in v3.1 should be "Low" in v4. Any deviation indicates a direct inconsistency.
- One-to-Many Mapping:
This category involves metrics where a single concept in v3.1 has been refined or split into multiple possibilities or more granular definitions in v4.0.
- User Interaction (UI): In v3.1, UI is either "Required" or "None." In v4.0, "Required" is further differentiated into "Passive" or "Active."
- If v3.1 UI is "None," then v4.0 UI should also be "None."
- If v3.1 UI is "Required," then v4.0 UI should be either "Passive" or "Active," depending on the specific nature of the interaction.
- Attack Complexity (AC): This metric saw a significant refinement. In v3.1, AC is "Low" or "High." V4.0 introduces Attack Requirements (AT), splitting the v3.1 AC into more nuanced conditions.
- Low AC (v3.1): Can map to "Low" AC in v4.0, either with or without Attack Requirements (AT). The presence of AT (e.g., specific external conditions) differentiates the v4.0 score.
- High AC (v3.1): Can map to "High" AC in v4.0, again with or without Attack Requirements (AT).
- Special Case for High AC (v3.1): The researchers identified a critical scenario where a "High" AC in v3.1 might downgrade to "Low" AC in v4.0, specifically when Attack Requirements (AT) are present. This addresses a perceived weakness in v3.1 where "High" complexity did not adequately distinguish between technically complex attacks and those that relied on specific, external conditions. V4.0's introduction of AT clarifies this, leading to five consistent combinations for AC in total.
- Many-to-Many Mapping:
This is the most complex category, primarily concerning the Scope metric from CVSSv3.1, which was removed in CVSSv4.0. Its concept was instead distributed across two distinct impact areas in v4.0: Vulnerable System Impact and Subsequent System Impact. The impact metrics (Confidentiality, Integrity, Availability - CIA) are now assigned to these two systems based on the v3.1 Scope.
- Scope Unchanged (v3.1): This implies that the impact is confined to the vulnerable component. Therefore, all CIA impact metrics from v3.1 should be fully applied to the Vulnerable System Impact in v4.0, with "No Impact" on the Subsequent System Impact. For example, if v3.1 has "High, Low, Low" for CIA with Unchanged Scope, v4.0 should have "High, Low, Low" for Vulnerable System CIA and "None, None, None" for Subsequent System CIA.
- Scope Changed (v3.1): This indicates that the impact extends beyond the vulnerable component. The researchers identified four distinct cases for how CIA impacts are distributed when the scope changes, based on their dataset analysis:
- Full Transfer to Both Systems: The CIA impact values are fully replicated across both the Vulnerable System and Subsequent System. E.g., "High, High, Low" in v3.1 becomes "High, High, Low" for both systems in v4.0.
- Vulnerable System Fully Impacted, Subsequent System Partially Impacted: The vulnerable system experiences the full CIA impact, while the subsequent system experiences a different, typically lower, level of impact. E.g., "High, High, None" in v3.1 might result in "High, High, None" for the vulnerable system and a different, non-zero impact for the subsequent system.
- Subsequent System Fully Impacted, Vulnerable System Partial/No Impact: The full CIA impact is transferred to the subsequent system, while the vulnerable system experiences partial or no impact. E.g., "High, High, None" in v3.1 could lead to "High, High, None" for the subsequent system and "None, None, None" or partial impact for the vulnerable system.
- Distributed Impact: The individual CIA impact values are distributed between the vulnerable and subsequent systems, with each specific impact (Confidentiality, Integrity, or Availability) transferred to only one system, leaving "No Impact" for that specific metric on the other system. For example, "Low Confidentiality, High Integrity" might transfer entirely to the subsequent system (with no impact on vulnerable Confidentiality/Integrity), while "Low Availability" transfers to the vulnerable system (with no impact on subsequent Availability).
These detailed rules provide a framework for analysts to systematically evaluate vulnerabilities, ensuring a more consistent and logical translation of scores between CVSSv3.1 and CVSSv4.0.
Demo / Proof of Concept
▶ Watch: ChatGPT's Inconsistencies in CVSS Metric Generation (6:30)
The talk did not feature a live demo or the presentation of a new proof-of-concept tool. Instead, the researchers demonstrated the efficacy of their proposed consistency rules by applying them to various public and private datasets of already-rated CVEs. This application served as a retrospective validation of the rules' ability to identify and highlight existing inconsistencies, effectively acting as a "proof of concept" for their detection and resolution.
The results of applying these rules to the public datasets (cv.org and Vondb, totaling over 2,000 CVEs) showed a significant improvement in consistency across most metrics. However, the User Interaction (UI) metric remained a notable outlier, showing a 22.1% inconsistency rate where v3.1 was rated "required" but v4.0 was "non." This finding indicates that while the rules can guide consistent scoring, UI's nuanced interpretation still poses a challenge that requires analyst discussion and potential modification.
Furthermore, the rules were applied to a private dataset from the CVSS Implementation Forum (CV-SSI), comprising over 6,000 CVEs. This analysis revealed that even for seemingly straightforward metrics like Attack Vector (AV) and Privileges Required (PR), minor inconsistencies persisted, underscoring the pervasive nature of the problem. For the UI metric, three inconsistent combinations accounted for 3.1% of entries. Impact-related metrics for one company (Company A) showed a high inconsistency of 26.7%, which, even after applying their internal "hidden agreement rules," still left 3.3% of entries inconsistent. These results underscore that while the consistency rules are a powerful tool for detection and guidance, human judgment and organizational agreements remain crucial for resolving the most complex discrepancies. The demonstration effectively proved that the rules can pinpoint where human intervention and clarification are most needed, thereby improving overall scoring reliability.
Defensive Implications
▶ Watch: Proposed Solution: Consistency Rules for CVSS Rating (9:00)
The findings from this research have profound implications for security defenders, especially those involved in vulnerability management, patching, and risk prioritization. The observed inconsistencies between CVSSv3.1 and v4.0 scores, whether from different CNAs, vendors, or even within the same organization over time, pose a significant challenge to effective security operations.
Firstly, organizations must acknowledge that CVSS scores, particularly during the transition period to v4.0, are not inherently consistent across versions or sources. Relying solely on a single score without understanding its context or potential discrepancies can lead to misprioritization of vulnerabilities. A critical vulnerability in v3.1 might be downgraded in v4.0 due to a different interpretation of metrics, or vice versa, directly impacting patching urgency and resource allocation.
Defenders should proactively adopt the proposed consistency rules when evaluating and scoring vulnerabilities. Integrating these rules into the vulnerability assessment workflow can ensure that if a vulnerability is scored for v3.1, its v4.0 counterpart is derived logically and consistently, or vice versa. This simultaneous scoring approach can prevent future discrepancies, save analyst time, and establish a more robust internal standard. For organizations consuming vulnerability data from multiple sources (e.g., NVD, vendor advisories, GitHub advisories), these rules can serve as a sanity check. If a vendor provides both v3.1 and v4.0 scores, applying the consistency rules can quickly highlight potential conflicts that warrant further investigation, rather than blindly trusting the provided scores.
Furthermore, the research highlights the limitations of purely automated scoring, even with advanced LLMs like ChatGPT, which demonstrated difficulties in maintaining logical consistency across CVSS versions. Defenders should exercise caution when using such tools for definitive CVSS scoring without subsequent human review and validation against established consistency guidelines.
The persistence of inconsistencies even with "hidden agreement rules" within companies indicates that informal or uncodified internal guidelines are insufficient. Organizations need to formalize their CVSS scoring policies, explicitly incorporating cross-version consistency checks. This includes training analysts on the nuances of v3.1 to v4.0 metric mappings, especially for complex areas like Attack Complexity and the transition from Scope to Vulnerable System and Subsequent System Impact.
Finally, the open question posed by the researchers—"Which version should users trust for patching or prioritization when both are available?"—is critical. Defenders need to develop an organizational stance on this. It might involve prioritizing the higher score, re-evaluating both scores based on internal risk appetite, or using the consistency rules to reconcile differences before making a decision. Ultimately, the goal is to move beyond simply receiving a score to understanding how that score was derived and ensuring its reliability, thereby enhancing the overall effectiveness of vulnerability management programs.
Key Takeaways
- Pervasive Inconsistencies: CVSSv3.1 and v4.0 scores frequently exhibit significant inconsistencies, driven by human subjectivity, differing interpretations of metric definitions, and insufficient information in vulnerability descriptions.
- LLM Limitations: While powerful, advanced LLMs like ChatGPT currently struggle to consistently generate CVSS metric vectors across versions, highlighting the need for human oversight and refined prompt engineering.
- Structured Consistency Rules: The researchers propose a set of "consistency rules" based on official CVSS documentation to guide analysts in simultaneously scoring vulnerabilities for both v3.1 and v4.0, aiming to prevent and detect discrepancies.
- Metric Mapping Categories: These rules categorize metric relationships into three types: one-to-one (e.g., Attack Vector, Privileges Required), one-to-many (e.g., User Interaction, Attack Complexity), and many-to-many (e.g., Scope in v3.1 mapping to Vulnerable System and Subsequent System Impact in v4.0).
- User Interaction (UI) as a Hotspot: Application of these rules to real-world datasets reveals that the User Interaction (UI) metric is a primary source of inconsistency, with a high percentage of entries showing conflicting ratings between v3.1 and v4.0.
- Actionable Guidance for Defenders: Organizations transitioning to or using CVSSv4.0 should integrate these consistency rules into their vulnerability assessment processes to ensure reliable, defensible scores, and improve the accuracy of vulnerability prioritization.
About the Speaker(s)
Monan is an Assistant Professor at VU Amsterdam. He traveled from the Netherlands to present this research, focusing on the scientific understanding of inconsistencies in vulnerability scoring systems. His work aims to provide actionable insights into improving the reliability and consistency of CVSS ratings.
Chisan is a PhD student at VU Amsterdam, currently undertaking her second PhD. She is supervised by Monan and Fabio Massage, and her research work is funded by the Dutch research council NWO. Chisan played a significant role in the detailed analysis of CVSS metric mappings and the development of the consistency rules presented in the talk.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Monan and Chisan present a methodical, academically rigorous examination of cross-version CVSS inconsistencies — a real, underappreciated operational problem. The consistency rules framework is a genuine contribution with practical utility, and the empirical grounding across multiple datasets (including a private CV-SSI corpus of 6,000+ CVEs) gives this more credibility than a pure theoretical exercise. That said, it's narrow in scope, and the headline finding — that CVSS scoring is inconsistent and humans interpret metrics differently — is not exactly news to anyone who's spent a week in vulnerability management. The novel contribution is the structured cross-version mapping taxonomy and…
Heather Calloway (CISO) — SOLID
Credible academic research that diagnoses a real operational problem — CVSS inter-version inconsistency — and produces a structured set of consistency rules grounded in documentation analysis and real-world datasets. The work is technically sound and the findings are specific enough to be useful. But the talk stops at the tool boundary and never crosses into the institutional territory where this problem actually lives: vendor liability, NVD governance, CNA accountability, and what organizations should actually do when their entire vulnerability management program is downstream of scores they cannot trust.