Shedding Light on CVSS Scoring Inconsistencies: A User-Centric Study on Evaluating Widespread Security Vulnerabilities

Julia Wunder, Andreas Kurtz, Christian Eichenmüller, Freya Gassmann, Zinaida Benenson

IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 4

Overview

The Common Vulnerability Scoring System (CVSS), maintained by FIRST, is a foundational tool for organizations worldwide in their vulnerability management processes. It provides a standardized method to calculate a numerical score for a vulnerability, indicating its severity and guiding prioritization efforts. However, a significant challenge arises when different evaluators assess the same vulnerability: the resulting CVSS scores often diverge, a phenomenon that directly contradicts CVSS documentation, which states scores should be agnostic. This talk, presented by Julia Wunder, delves into a comprehensive user-centric study investigating the consistency of CVSS version 3.1 evaluations.

Watch on YouTube

Visual summary for Shedding Light on CVSS Scoring Inconsistencies: A User-Centric Study on Evaluating Widespread Security Vulnerabilities by Julia Wunder, Andreas Kurtz, Christian Eichenmüller, Freya Gassmann, Zinaida Benenson
Visual summary for Shedding Light on CVSS Scoring Inconsistencies: A User-Centric Study on Evaluating Widespread Security Vulnerabilities by Julia Wunder, Andreas Kurtz, Christian Eichenmüller, Freya Gassmann, Zinaida Benenson

Key moments

  1. 0:00 Introduction: CVSS scoring inconsistencies problem
  2. 0:54 CVSS base score metrics and calculation
  3. 2:18 Overview of the study design and methodology
  4. 4:26 Key research questions investigated in the study
  5. 6:40 Problematic metrics and vulnerability types analyzed

Shedding Light on CVSS Scoring Inconsistencies: A User-Centric Study on Evaluating Widespread Security Vulnerabilities

Speakers: Julia Wunder; Andreas Kurtz; Christian Eichenmüller; Freya Gassmann; Zinaida Benenson

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=hj_7v534PHQ

Overview

The Common Vulnerability Scoring System (CVSS), maintained by FIRST, is a foundational tool for organizations worldwide in their vulnerability management processes. It provides a standardized method to calculate a numerical score for a vulnerability, indicating its severity and guiding prioritization efforts. However, a significant challenge arises when different evaluators assess the same vulnerability: the resulting CVSS scores often diverge, a phenomenon that directly contradicts CVSS documentation, which states scores should be agnostic. This talk, presented by Julia Wunder, delves into a comprehensive user-centric study investigating the consistency of CVSS version 3.1 evaluations.

The research aimed to quantify the extent of these inconsistencies, identify specific metrics and vulnerability types that are particularly problematic, and explore factors influencing evaluator discrepancies. By conducting preliminary expert discussions, a preliminary survey, a large-scale main survey with 196 participants, and a follow-up study nine months later, the researchers meticulously gathered data on how security professionals apply CVSS in practice. The findings illuminate critical gaps in understanding and application, highlighting the need for improved clarity and training to enhance the reliability of CVSS scores in real-world security operations.

The implications of inconsistent CVSS scoring are profound, directly impacting how quickly and effectively vulnerabilities are addressed. Misaligned scores can lead to misallocation of resources, delayed patching of critical threats, or unnecessary focus on lower-risk issues. This study provides crucial empirical evidence demonstrating that despite its widespread adoption and perceived utility, CVSS 3.1 suffers from significant interpretational variability among users, urging a re-evaluation of current practices and future iterations of the scoring system.

Background

▶ Watch: Introduction: CVSS scoring inconsistencies problem (0:00)

CVSS serves as a global standard for communicating the characteristics and impacts of IT vulnerabilities. It is structured around three metric groups: Base, Temporal, and Environmental. The focus of this study, and often in practice, is the Base Score, which represents the intrinsic characteristics of a vulnerability and remains constant over time. The Base Score is derived from eight distinct metrics, each reflecting a different aspect of the vulnerability. These metrics include Attack Vector (AV), Attack Complexity (AC), Privileges Required (PR), User Interaction (UI), Scope (S), Confidentiality Impact (C), Integrity Impact (I), and Availability Impact (A).

During a CVSS assessment, an evaluator assigns a value to each of these eight metrics. These values are then compiled into a vector string, a textual representation that concisely describes how each metric was evaluated (e.g., AV:N/AC:L/PR:N/UI:N/S:U/C:H/I:H/A:H). This vector string is then used to calculate the numerical CVSS Base Score, which ranges from 0.0 (None) to 10.0 (Critical). In practice, these numerical scores are often mapped to qualitative severity categories: None (0.0), Low (0.1-3.9), Medium (4.0-6.9), High (7.0-8.9), and Critical (9.0-10.0). These categories frequently dictate an organization's response time and prioritization for remediation efforts, meaning a difference of even a few tenths of a point can shift a vulnerability from "High" to "Critical," dramatically altering its perceived urgency.

The fundamental premise of CVSS, as explicitly stated in its documentation, is that its scores should be agnostic. This implies that any trained evaluator, when presented with the same vulnerability information, should arrive at the same CVSS score. This agnosticism is vital for consistent vulnerability management across different teams, organizations, and even countries. However, real-world application often reveals a disparity in scores, leading to confusion and potential misallocation of resources. Prior anecdotal evidence and practical experience suggested that certain CVSS metrics are more prone to subjective interpretation than others, contributing to these inconsistencies. This study aimed to move beyond anecdotal evidence, providing empirical data on where and why these discrepancies occur, particularly for widespread and impactful vulnerability types.

Key Findings

▶ Watch: CVSS base score metrics and calculation (0:54)

The study uncovered significant inconsistencies in CVSS 3.1 evaluations, both across different evaluators and for the same evaluator over time. The core findings can be summarized as follows:

  1. Widespread Inconsistency in Evaluations: The overarching conclusion was that CVSS evaluations were not consistent. When different participants evaluated the same vulnerability, their resulting scores and severity categories frequently differed.
  2. Problematic Metrics Identified: The Scope (S), User Interaction (UI), and Attack Vector (AV) metrics were consistently identified as the most problematic and inconsistently evaluated.
  • For Scope, participants were split, with approximately 60% choosing "Unchanged" and 40% choosing "Changed," irrespective of the vulnerability type. This metric also garnered numerous complaints about its ambiguity and difficulty to understand.
  • For User Interaction, while reflected XSS vulnerabilities showed a higher agreement (around 75% rated "Required"), stored XSS vulnerabilities were highly ambiguous, with participants nearly evenly split between "Required" and "None."
  • For Attack Vector, particularly for drive-by download scenarios (e.g., file downloaded from the internet but opened locally), there was significant disagreement on whether it should be "Network" or "Local." Similarly, for Man-in-the-Middle attacks, the distinction between "Network" and "Adjacent" was unclear for many.
  1. Mis-scoring of Security Deficiencies: The study highlighted a significant issue with how security deficiencies (e.g., Banner Disclosure, HTTP Only missing) are scored. According to CVSS documentation, these should ideally receive a "None" severity rating as they are not vulnerabilities on their own but rather missing security mechanisms that become dangerous only in combination with other vulnerabilities. However, participants frequently assigned them medium, high, or even critical severities. For Banner Disclosure, while "Medium" was the most common, "None" was also highly selected. For HTTP Only, there was the most disagreement, with ratings spanning "None," "Low," "Medium," and "High," indicating substantial confusion.
  2. Evaluator Attitude: Useful but Flawed: Despite the observed inconsistencies, participants generally held a positive attitude towards CVSS. A majority agreed that CVSS is useful and helpful, and they intended to continue using it. Concurrently, they also acknowledged significant inconsistencies in scoring and frequent score differences, indicating that users are aware of the system's flaws but still value its utility. They did not, however, believe the scores were "useless."
  3. Limited Use of Official Documentation: A striking finding was that most participants rely heavily on the CVSS calculator for assessments, while very few consult the official CVSS documents (specification, user guide, example document). Most participants reported having "none" or "basic" knowledge of these documents and had consulted them months, years, or never. This suggests a critical disconnect: clarifications and refinements in official documentation are unlikely to reach the majority of users who primarily interact with the calculator.
  4. Inconsistency Over Time: The follow-up study, conducted nine months after the main survey with the same participants, revealed that CVSS evaluations are not consistent over time. A significant percentage of participants (nearly 60% for HTTP Only) chose a different severity rating for the same vulnerability they had evaluated previously. This temporal inconsistency further complicates vulnerability management, as the perceived urgency of a vulnerability can change even without new information.

Technical Deep Dive

▶ Watch: Overview of the study design and methodology (2:18)

The study employed a rigorous multi-stage methodology to investigate CVSS 3.1 scoring inconsistencies, moving from qualitative expert discussions to large-scale quantitative surveys.

1. Preliminary Studies:

The research began with two critical preliminary steps:

  • Expert Discussion Session: Three CVSS experts were engaged in a discussion to identify practical problems with CVSS, problematic vulnerability types, and specific metrics that frequently cause debate and inconsistent evaluation. This qualitative input was crucial for shaping the subsequent quantitative phases.
  • Preliminary Survey: Following the discussion, 18 CVSS experts participated in a preliminary online survey where they assessed 10 selected vulnerabilities. This survey also included a free-text field for experts to share their opinions and experiences with CVSS. The insights from this stage helped refine the selection of vulnerabilities and metrics for the main study, ensuring they were representative of real-world challenges and prone to disagreement. This process helped define the research questions and identify widespread vulnerability types and problematic metrics for further investigation.

2. Research Questions:

Based on the preliminary studies, the researchers formulated five key research questions:

  • Are Attack Vector, User Interaction, and Scope inconsistently evaluated for some widespread vulnerability types? (Focus on specific metrics)
  • Are security deficiencies considered suitable for CVSS assessment by CVSS users? (Focus on a special class of security issues)
  • Are personal factors (e.g., work experience, opinion of CVSS) associated with individual differences in CVSS evaluations?
  • What is the attitude of evaluators towards CVSS?
  • Are CVSS evaluations of the same evaluator consistent over time?

3. Main Study Design:

The core of the research was a large-scale online survey involving 196 participants.

  • Vulnerability Selection: The preliminary studies guided the selection of eight specific vulnerabilities, chosen because they highlighted ambiguities in the problematic metrics:
  • User Interaction (UI) metric and XSS: Investigated for Reflected XSS (user actively clicks a malicious link) versus Stored XSS (malicious script embedded on a website). The question was whether UI should be "None" or "Required."
  • Scope (S) metric: Described as "ambiguous and hard to understand," this metric indicates if a vulnerability affects systems beyond the security authority of the vulnerable component. Investigated for XSS, Drive-by Download, and SQL Injection, which are among the most severe vulnerabilities according to the CVE Top 25 List of 2022. The core question was "Unchanged" or "Changed."
  • Attack Vector (AV) metric and Drive-by Download: For scenarios where a file is downloaded from the internet but executed locally (e.g., Adobe buffer overflow, Chrome buffer overflow), the ambiguity was between "Network" (initial download) and "Local" (execution).
  • Attack Vector (AV) metric and Man-in-the-Middle (MitM): For MitM attacks, the distinction between "Network" and "Adjacent" (attacker within a logically or physically bounded network, like Wi-Fi) was explored.
  • Suitability of Security Deficiencies: Two examples were chosen: Banner Disclosure (revealing web server version) and Missing HTTP Only flag (for cookies). These were chosen to test if evaluators would adhere to the "None" rating recommended by CVSS for non-vulnerabilities.
  • Participant Grouping: To manage survey length, participants were divided into two groups, each evaluating four vulnerabilities.
  • Group 1: Stored XSS, Reflected XSS, SQL Injection, and Banner Disclosure (as a security deficiency). Metrics of focus varied per vulnerability (e.g., Stored XSS focused on UI and Scope).
  • Group 2: Adobe Buffer Overflow, Chrome Buffer Overflow (both drive-by downloads), Man-in-the-Middle, and HTTP Only (as a security deficiency).
  • Demographics: The 196 participants had an average age of 38, were predominantly male, and represented diverse geographical locations (Germany, UK, US, Asia). Most held Master's or Bachelor's degrees. Crucially, their self-rated expertise in CVSS 3.1 was high, with many reporting "advanced" or "intermediate" knowledge, and an average of 6.7 years of CVSS usage. This indicates that inconsistencies are not merely due to novice users.

4. Follow-up Study:

Nine months after the main study, 59 of the original participants were invited to a follow-up survey. Each participant evaluated two familiar vulnerabilities (from their original main study group) and two new vulnerabilities (from the other group, acting as distractors). The primary goal was to assess the consistency of evaluations over time by the same individual.

Key Technical Results and Observations:

  • Scope Ambiguity: A participant's quote, "scope ask 10 people and you get 10 coin tosses," perfectly encapsulated the confusion. The near 60/40 split between "Unchanged" and "Changed" across all vulnerability types underscored the metric's inherent ambiguity, regardless of the specific context.
  • User Interaction Nuance: While reflected XSS (requiring a click) was largely rated "Required" (approx. 75%), stored XSS (script executes on page load) showed significant contention, with a near 55-60% "Required" versus 40-45% "None" split. This highlights the difficulty in determining user "involvement" when a script is passively executed.
  • Security Deficiency Misinterpretation: The highest agreement for Banner Disclosure was "Medium" severity, contradicting the CVSS guideline of "None." For HTTP Only, the disagreement was even more pronounced, with ratings across "None," "Low," "Medium," and "High." This suggests a fundamental misunderstanding or disagreement among evaluators on how to apply CVSS to non-exploitable security issues.
  • Tool Usage Disparity: The finding that most participants used the CVSS calculator and rarely consulted official documentation is a critical insight. It implies that any future clarifications or refinements to CVSS metrics should ideally be integrated directly into the calculator interface rather than solely relying on updated textual guidelines.

Demo / Proof of Concept

▶ Watch: Key research questions investigated in the study (4:26)

The talk did not feature a traditional technical demonstration or proof of concept in the sense of exploiting a vulnerability or showcasing a software tool. Instead, the core "demonstration" was the comprehensive user study itself, meticulously designed to expose and quantify the inconsistencies in CVSS scoring. The methodology, involving preliminary expert discussions, a preliminary survey, a main survey with 196 participants evaluating specific vulnerabilities against problematic CVSS metrics, and a follow-up survey, served as an empirical "proof of concept" for the hypothesis that CVSS evaluations are inconsistent. The results, presented through detailed graphs and statistical analyses of participant responses, effectively demonstrated the variability in scoring for metrics like Attack Vector, User Interaction, and Scope, as well as the misapplication of CVSS to security deficiencies.

Defensive Implications

▶ Watch: Problematic metrics and vulnerability types analyzed (6:40)

The findings of this study carry significant implications for organizations and security professionals responsible for vulnerability management and defense. The pervasive inconsistencies in CVSS 3.1 scoring directly challenge the reliability of vulnerability prioritization, potentially leading to critical misjudgments in resource allocation and response times.

  1. Beyond the Score: Defenders should recognize that a raw CVSS score alone might not be a definitive indicator of a vulnerability's true severity or urgency within their specific context. It is crucial to supplement CVSS scores with additional contextual information, such as asset criticality, exploitability in their environment, and the presence of compensating controls. Security teams should develop internal guidelines and playbooks that interpret CVSS scores through the lens of their organization's risk profile.
  2. Enhanced Evaluator Training: The study highlights the need for more robust and targeted training for security professionals involved in CVSS assessment. Training should focus specifically on the metrics identified as most problematic—Scope, User Interaction, and Attack Vector—providing clear examples and decision trees for ambiguous scenarios. Emphasizing the nuances of these metrics is vital to reduce subjective interpretation.
  3. Rethink Security Deficiency Scoring: Organizations must establish clear internal policies for handling security deficiencies (e.g., missing HTTP Only, banner disclosure). While CVSS documentation suggests a "None" rating, the study shows practitioners often assign higher severities. Defenders should acknowledge that while these might not be directly exploitable vulnerabilities, they can significantly lower an organization's security posture and should be tracked and remediated, perhaps using a separate risk rating system or internal scoring mechanism distinct from CVSS for these specific issues.
  4. Leverage Calculators, Improve Documentation Access: Given that most evaluators rely on CVSS calculators and rarely consult official documentation, future efforts to clarify CVSS metrics should prioritize embedding guidance and decision aids directly into these calculation tools. Developers of CVSS calculators could integrate pop-up explanations, interactive flowcharts, or context-sensitive help for ambiguous metrics to guide users toward more consistent evaluations.
  5. Internal Consistency Audits: Security teams should periodically conduct internal audits of their CVSS scoring practices. This could involve having multiple team members score the same set of vulnerabilities and then discussing discrepancies to align interpretations and identify areas where internal guidelines need strengthening. The temporal inconsistency observed in the follow-up study also suggests that re-evaluations should be approached with caution, potentially requiring a review of the original context.
  6. Advocate for CVSS Evolution: The research provides empirical data that can inform the ongoing development of CVSS, including the recently announced CVSS 4.0. Defenders should be aware of these discussions and, where possible, contribute feedback based on their practical experiences to ensure future versions address the identified ambiguities and promote greater consistency.

Key Takeaways

  • CVSS 3.1 evaluations are inconsistent: The study empirically demonstrated significant variability in CVSS scores for the same vulnerability across different evaluators and even by the same evaluator over time.
  • Problematic metrics: Scope, User Interaction, and Attack Vector were identified as the most ambiguous and inconsistently evaluated metrics, leading to frequent disagreements among security professionals.
  • Security deficiencies are mis-scored: Vulnerability-like issues such as Banner Disclosure and missing HTTP Only flags are often assigned medium-to-high severities, contradicting CVSS guidelines that suggest a "None" rating for non-exploitable deficiencies.
  • Documentation gap: Most CVSS users rely on online calculators and rarely consult official CVSS documentation, indicating a critical disconnect in how clarifications and guidelines are communicated to practitioners.
  • Perceived utility despite flaws: Despite the observed inconsistencies, evaluators generally view CVSS as a useful and helpful tool, suggesting a need for refinement and better guidance rather than abandonment.
  • Contextualization is key: Organizations should complement CVSS scores with internal contextual factors and provide targeted training on ambiguous metrics to ensure more consistent and effective vulnerability prioritization.

About the Speaker(s)

The talk "Shedding Light on CVSS Scoring Inconsistencies: A User-Centric Study on Evaluating Widespread Security Vulnerabilities" was presented by Julia Wunder. She is the lead author and researcher behind this comprehensive study. The paper was co-authored by Andreas Kurtz, Christian Eichenmüller, Freya Gassmann, and Zinaida Benenson. Together, this team conducted the extensive user studies and data analysis to investigate the consistency of CVSS evaluations. Julia Wunder's presentation highlighted the methodology, key findings, and implications of their research for the security community.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk presents a rigorously empirical study exposing critical inconsistencies in CVSS 3.1 scoring, identifying specific problematic metrics like Scope and User Interaction. It proves that despite its widespread use, CVSS's 'agnostic' premise is fundamentally broken in practice, with profound implications for vulnerability prioritization and resource allocation. A necessary, data-driven gut check for anyone relying on CVSS.

Heather Calloway (CISO) — STRONG ACCEPT

This study provides empirical evidence of significant inconsistencies in CVSS scoring, undermining its intended role as an agnostic prioritization tool. The findings expose critical gaps in how organizations assess and own risk, demanding immediate action to improve internal guidance and training on vulnerability evaluation. Effective leaders will integrate these insights to refine their vulnerability management strategy and ensure more reliable risk decisions.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024