PassREfinder: Credential Stuffing Risk Prediction by Representing Password Reuse between Websites on a Graph

Jaehan Kim, Minkyoo Song, Minjae Seo, Youngjin Jin, Seungwon Shin

IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 6

Overview

In the contemporary digital landscape, where individuals manage an ever-increasing number of online accounts, the convenience of reusing passwords across multiple services has unfortunately become a widespread practice. This behavior, known as password reuse, significantly amplifies the risk of credential stuffing attacks. This talk, presented by Jaehan Kim from NSS Lab at KIST, advised by Professor Seungwon Shin, introduces PassREfinder, a novel Graph Neural Network (GNN)-based framework designed to predict the risk of credential stuffing by modeling password reuse relationships between websites.

Watch on YouTube

Visual summary for PassREfinder: Credential Stuffing Risk Prediction by Representing Password Reuse between Websites on a Graph by Jaehan Kim, Minkyoo Song, Minjae Seo, Youngjin Jin, Seungwon Shin
Visual summary for PassREfinder: Credential Stuffing Risk Prediction by Representing Password Reuse between Websites on a Graph by Jaehan Kim, Minkyoo Song, Minjae Seo, Youngjin Jin, Seungwon Shin

Key moments

  1. 0:00 Introduction to PassREfinder and credential stuffing problem
  2. 2:00 Existing mitigation methods for credential stuffing
  3. 3:57 Limitations of existing methods and PassREfinder proposal
  4. 6:00 Overview of PassREfinder framework and graph construction
  5. 7:20 Website feature extraction and vectorization process
  6. 8:40 GNN layer design: multimodal and neighborhood attention
  7. 10:18 Link prediction task and evaluation introduction

PassREfinder: Credential Stuffing Risk Prediction by Representing Password Reuse between Websites on a Graph

Speakers: Jaehan Kim; Minkyoo Song; Minjae Seo; Youngjin Jin; Seungwon Shin

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=6m7XLUDkA0E

Overview

In the contemporary digital landscape, where individuals manage an ever-increasing number of online accounts, the convenience of reusing passwords across multiple services has unfortunately become a widespread practice. This behavior, known as password reuse, significantly amplifies the risk of credential stuffing attacks. This talk, presented by Jaehan Kim from NSS Lab at KIST, advised by Professor Seungwon Shin, introduces PassREfinder, a novel Graph Neural Network (GNN)-based framework designed to predict the risk of credential stuffing by modeling password reuse relationships between websites.

PassREfinder addresses critical limitations in existing credential stuffing mitigation strategies, which often struggle with detection coverage, scalability, and practical deployment challenges. By reframing the problem from reactive attack detection to proactive risk prediction, the framework aims to preemptively identify potential credential stuffing vulnerabilities, thereby minimizing the operational costs associated with false positives and enhancing overall security posture for both users and administrators. The core innovation lies in its ability to represent complex password reuse dynamics as a graph structure, enabling a powerful GNN model to learn and predict these intricate relationships.

The significance of PassREfinder lies in its practical applicability and robust performance. By leveraging publicly available website information and sophisticated graph-based machine learning techniques, the framework offers a scalable and privacy-preserving solution to a pervasive security problem. Its ability to accurately predict credential stuffing risks, even for previously unknown websites, provides a crucial tool for administrators to implement targeted security measures, ultimately safeguarding user accounts and sensitive data from the persistent threat of credential stuffing.

Background

▶ Watch: Introduction to PassREfinder and credential stuffing problem (0:00)

The pervasive issue of password reuse forms the bedrock of credential stuffing attacks, a significant and escalating threat in cybersecurity. As highlighted in the talk, a staggering 60% of passwords are reused at least once, with the majority being exact matches across different services. This user behavior, driven by convenience, creates a fertile ground for attackers. Once credentials are compromised from a single, often less secure, service, attackers exploit this reuse by attempting to log into other services, particularly those with higher security levels, in the hope that the same username and password combination will grant unauthorized access. The consequences of successful credential stuffing attacks are substantial, leading to data breaches, financial losses, and reputational damage for organizations.

Existing mitigation strategies have attempted to address this problem, but each comes with inherent limitations. One widely adopted method is compromised credential checking (C3) services, exemplified by tools like Google Password Checkup and Apple's password manager. These services compile databases of known compromised credentials and allow users to verify if their own credentials have been exposed. However, C3 services are inherently reactive; they are limited to previously disclosed or already compromised credentials. They offer no protection against newly compromised or undisclosed credentials, leaving a significant window of vulnerability.

A more advanced, yet theoretically promising, approach involves website coordination protocols. These methods aim to detect reused passwords or credential stuffing attempts by enabling secure, privacy-preserving sharing of credential information among websites. While their theoretical underpinnings are sound, their practical deployment faces several hurdles. The talk points out concerns such as the high cost associated with false positives, which can directly interrupt legitimate users. There are also privacy implications stemming from their username mapping systems, raising concerns about data aggregation. Furthermore, the encryption mechanisms required for privacy preservation often incur substantial computational overheads, leading to limited scalability across the vast number of websites on the internet.

Recognizing these gaps, PassREfinder proposes a paradigm shift: moving from reactive attack detection to proactive risk prediction. Instead of merely identifying an ongoing attack, the framework aims to preemptively identify websites that are at a high risk of being targeted by credential stuffing due to shared password reuse patterns. This approach minimizes the disruptive impact of false positives and allows for more strategic allocation of security resources. The core idea is to quantify credential stuffing risk between websites by calculating the rate of users who reuse passwords between them. This complex network of risks is then modeled using a graph structure, specifically a password reuse graph, which is subsequently analyzed by a Graph Neural Network (GNN), a powerful tool for understanding relationships among multiple entities.

Key Findings

▶ Watch: Limitations of existing methods and PassREfinder proposal (3:57)

PassREfinder demonstrates remarkable efficacy in predicting credential stuffing risk, significantly outperforming a range of baseline prediction models across different operational settings. The core findings underscore its robustness and practical utility in mitigating the threats posed by password reuse.

Firstly, PassREfinder consistently outperformed all other baselines in both transductive and inductive graph learning settings. This indicates its superior ability to generalize and make accurate predictions, whether all nodes are partially known or when dealing with entirely new, unknown websites. Crucially, in the more challenging and practically relevant inductive setting, PassREfinder achieved an impressive F1 score of 0.91. This high F1 score signifies a strong balance between precision and recall, meaning the framework is highly effective at identifying true risks while minimizing false positives and negatives.

A significant observation from the evaluation was the performance disparity between PassREfinder and traditional GNN models in the inductive setting. While baseline GNNs showed reasonable performance in the transductive setting, they struggled considerably when confronted with unknown website nodes not initially connected to the graph. PassREfinder specifically addresses this limitation through its tailored GNN layers and, most importantly, the incorporation of the hidden relation method. This architectural design allows PassREfinder to effectively integrate and make predictions for new or unobserved websites, a critical capability for real-world deployment on the dynamic internet.

Furthermore, the study meticulously analyzed the effectiveness of various hidden relation policies designed to connect unknown website nodes to the graph. The adoption of any hidden relation method was shown to significantly improve PassREfinder's performance, especially in the inductive setting. Among the policies evaluated, the shared user policy — which connects websites if they share a certain number of overlapping users — proved to be the most effective. However, the probabilistic policy (randomly connecting edges) and the similarity policy (connecting based on cosine similarity of node features) also demonstrated comparable performance. This is a vital finding, as these two policies do not require user information, offering administrators a practical means to implement the framework while adhering to privacy principles, making it more widely deployable.

In summary, PassREfinder not only provides a highly accurate mechanism for predicting credential stuffing risk but also offers a scalable and privacy-conscious solution that addresses the persistent challenge of integrating unknown entities into its predictive model. Its strong performance in inductive settings, coupled with the flexibility of its hidden relation policies, positions it as a significant advancement in proactive cybersecurity defense.

Technical Deep Dive

▶ Watch: Overview of PassREfinder framework and graph construction (6:00)

PassREfinder's technical architecture is meticulously designed to transform the complex problem of credential stuffing risk into a tractable graph-based prediction task. The framework is built upon a Graph Neural Network (GNN) and operates through several distinct, yet interconnected, steps: graph construction, vector extraction for node features, and GNN-based representation learning followed by link prediction.

The foundational step is reframing the problem from reactive attack detection to proactive risk prediction. PassREfinder defines credential stuffing risk between websites by calculating the rate of users who reuse passwords between them. If this password reuse rate exceeds a predefined threshold, the risk is classified as high; otherwise, it's low. For the purpose of ground truth labeling in evaluations, a threshold of 0.5 was used: if the reuse rate for a prediction target edge was greater than 0.5, it was labeled as '1' (risk exists), otherwise '0'.

The system then proceeds to graph construction, creating a password reuse graph. In this graph, individual website nodes represent online services. Two types of edges are defined to connect these nodes:

  1. Password Reuse Relation Edges: These are the primary targets for prediction. An edge of this type exists between two websites if the credential stuffing risk between them is high, indicating a significant overlap in user password reuse.
  2. Hidden Relation Edges: These are crucial for handling unknown website nodes—websites that are not initially connected to the core graph. To ensure these unknown nodes (e.g., "Website D" in the speaker's example) can still be incorporated into the model, hidden relation edges are established based on a specified hidden relation policy. This mechanism is critical for PassREfinder's performance in inductive settings, where the model encounters new data not seen during training.

Following graph construction, vector extraction is performed to assign meaningful input features to each website node. A key design principle here is the use of publicly available website information to eliminate privacy concerns associated with sharing user data, a common drawback of existing detection methods. The framework queries public web analytics services such as URLscan and Shodan to gather website intelligence. In the prototype, five website features were selected, with examples including IP addresses and website categories. Because these features possess diverse data structures (e.g., categorical for categories, numerical for IP octets or ranges), PassREfinder employs a dedicated embedding method for each feature to convert them into a uniform vector representation suitable for GNN processing.

The vectorized node features and the constructed password reuse graph are then fed into the GNN layers for edge representation learning and link prediction. PassREfinder's GNN architecture incorporates two carefully designed techniques to model complex password reuse relations:

  1. Multimodal Feature Modeling: Rather than simply concatenating all feature embeddings, the framework utilizes a dedicated GNN layer for each individual website feature. This allows the model to inherently reflect the differences in data structures and semantic meanings of features (e.g., how IP addresses might relate differently than website categories). The outputs from these individual GNN layers are then merged into a single, comprehensive node representation vector using a multimodal attention mechanism. This attention mechanism dynamically determines the importance of each website feature for a given node, allowing the model to focus on the most relevant attributes.
  2. Neighborhood Attention: Within each GNN layer, features from neighbor nodes are propagated and aggregated to enrich the representation of a given node. To refine this aggregation process, PassREfinder employs nodewise attention. This mechanism assigns different importance weights to neighboring nodes, enabling the prediction model to prioritize neighbors that are more strongly related or influential in determining the given node's characteristics, thereby capturing nuanced relational patterns.

Finally, the system performs a link prediction task. The representation of each potential prediction target edge is obtained by concatenating the final node representation vectors of the two websites it connects. Based on this edge representation, the model outputs a probability score. If this probability exceeds a certain threshold, the existence of a password reuse relation edge is predicted, signifying a high credential stuffing risk between those websites.

The framework's effectiveness was rigorously evaluated using the C-Day dataset, one of the largest breach datasets available, comprising 36 million credentials from a massive number of websites. The evaluations were conducted under two critical graph learning settings:

  • Transductive Setting: In this scenario, all nodes are partially connected to the graph, meaning the model has some exposure to all entities, even if it hasn't seen all their relationships.
  • Inductive Setting: This represents a more challenging and practical scenario, where the model encounters entirely unknown website nodes that were not part of the initial graph structure. In this setting, only the previously described hidden relation edges can connect these new nodes to the existing graph.

PassREfinder's prototype implemented three specific hidden relation policies:

  • Probabilistic Policy: Randomly connects edges between nodes with a predefined probability. This serves as a baseline for connecting unknown nodes.
  • Similarity Policy: Connects edges between nodes if the cosine similarity of their extracted node features is higher than a specified threshold. This is an intuitive policy assuming similar websites might share reuse patterns.
  • Shared User Policy: Connects edges if there are at least a certain number of overlapping users across the two websites. This policy directly leverages user behavior data to infer connections.

The comprehensive technical design, from privacy-preserving feature extraction to sophisticated GNN attention mechanisms and robust evaluation strategies, underscores PassREfinder's capability to accurately and practically predict credential stuffing risks on a large scale.

Demo / Proof of Concept

▶ Watch: GNN layer design: multimodal and neighborhood attention (8:40)

The talk primarily focuses on the theoretical framework, architectural design, and extensive evaluation of PassREfinder's performance using real-world datasets. While the presentation details the methodology and results of various experiments to demonstrate the framework's effectiveness, it does not describe or showcase a live demonstration or a specific working proof-of-concept tool. The emphasis is on the quantitative measurement of risk prediction performance and the efficacy of its underlying components.

Defensive Implications

▶ Watch: Link prediction task and evaluation introduction (10:18)

PassREfinder offers significant practical implications for website administrators and security professionals looking to bolster their defenses against credential stuffing attacks. By providing a proactive and accurate risk prediction framework, it enables the intelligent and efficient allocation of security resources.

From the perspective of a website administrator, the deployment of PassREfinder allows for a targeted approach to security. Administrators can integrate PassREfinder with their internal password reuse information to predict the credential stuffing risk between their managed websites and any other target websites across the internet. This capability allows them to identify specific risky target websites that pose a high threat due to shared password reuse patterns.

The talk proposes three key application scenarios for leveraging these risk prediction results to enhance the efficiency of security measures:

  1. Optimized User Warnings: Instead of broadcasting generic warnings to all users about potential password reuse, administrators can use PassREfinder's predictions to send targeted warnings. Warnings would only be issued to users who have credentials in identified "risky" websites, significantly reducing user fatigue from irrelevant alerts and increasing the impact of actual warnings. This precision ensures that users who are genuinely at higher risk receive timely and relevant advisories.
  1. Selective Two-Factor Authentication (2FA) Enforcement: While 2FA is a robust security measure, universal enforcement can introduce user friction and operational overhead. PassREfinder enables administrators to selectively impose 2FA only on users who are involved with identified risky websites. This strategy significantly reduces the overhead of activating 2FA across the entire user base while still providing enhanced security for the most vulnerable accounts, striking a balance between security and user experience.
  1. Efficient Website Coordination Protocols: Existing website coordination protocols, which aim to share credential information securely, often suffer from limited scalability due to high overheads. PassREfinder can address this by helping administrators organize moderate-sized coordination pools. Instead of attempting to coordinate with an unmanageable number of websites, administrators can include only those websites that PassREfinder has identified as showing high credential stuffing risks with their own services. This focused coordination reduces the complexity and computational burden, making these privacy-preserving protocols more practical and scalable for real-world deployment.

In essence, PassREfinder shifts the security paradigm from broad, often inefficient, measures to precise, data-driven interventions. By pinpointing specific areas of high risk, it empowers defenders to deploy their resources more effectively, ultimately leading to stronger, more resilient security postures against credential stuffing.

Key Takeaways

  • Credential stuffing remains a pervasive threat, with 60% of passwords being reused, and existing detection methods like C3 services and website coordination protocols face significant limitations in coverage, scalability, and practical deployment.
  • PassREfinder introduces a novel Graph Neural Network (GNN)-based framework that reframes credential stuffing defense from reactive attack detection to proactive risk prediction, utilizing a password reuse graph to model inter-website relationships.
  • The framework ensures user privacy by extracting website features (e.g., IP addresses, website categories) exclusively from public sources like URLscan and Shodan, avoiding the need for sensitive user information.
  • PassREfinder's GNN architecture incorporates advanced techniques such as multimodal feature modeling with attention and neighborhood attention to effectively learn complex password reuse relations and generate robust node representations.
  • The system achieves high accuracy, demonstrating an F1 score of 0.91 in the challenging inductive setting, and effectively addresses the problem of unknown website nodes through its innovative hidden relation methods (e.g., probabilistic, similarity, and shared user policies).
  • Website administrators can leverage PassREfinder's risk predictions to implement targeted security measures, including selective user warnings, optimized 2FA enforcement, and the formation of efficient website coordination pools, thereby enhancing overall security posture and resource allocation.

About the Speaker(s)

The primary speaker for this presentation is Jaehan Kim, who is affiliated with the NSS Lab at KIST (Korea Institute of Science and Technology). He is advised by Professor Seungwon Shin. The work presented is a collaborative effort with Minkyoo Song, Minjae Seo, and Youngjin Jin, also from KIST. While specific titles or companies for the co-authors are not detailed in the transcript, their collective contribution to the research on credential stuffing risk prediction is evident.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

PassREfinder presents a novel GNN-based framework for proactive credential stuffing risk prediction by modeling password reuse on a graph. It leverages publicly available website data to ensure privacy and delivers impressive performance, particularly in inductive settings for unknown websites. This work provides actionable intelligence for targeted security measures, shifting defense from reactive to efficient, data-driven prevention.

Heather Calloway (CISO) — STRONG ACCEPT

PassREfinder offers a compelling shift from reactive credential stuffing detection to proactive risk prediction using Graph Neural Networks, providing administrators with a privacy-preserving tool to identify and mitigate high-risk password reuse patterns. Its strength lies in translating complex technical analysis into actionable strategies for targeted user warnings and optimized 2FA enforcement. This research directly informs how security leaders can better allocate resources against a persistent and costly threat.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024