Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data

Atilla Akkus

34th USENIX Security Symposium (USENIX Security '25) · Day 3 · Vulnerabilities in LLMs: Privacy, Safety, and Defense

Overview

This groundbreaking paper introduces PassLLM, an innovative framework that leverages Large Language Models (LLMs) with Low-Rank Adaptation (LoRA) techniques for advanced password guessing attacks. Authored by Yunkai Zou, Maoxiang An, and Ding Wang from Nankai University, PassLLM addresses the inherent limitations of general-purpose LLMs in specialized tasks like password guessing, where static knowledge and a disconnect from real-world password creation behaviors hinder their effectiveness. The work presents a novel technical route, demonstrating how modern LLMs can be efficiently fine-tuned and optimized for this critical security domain.

Read the paper · Download the PDF (PDF) · Slides

Paper abstract

Passwords are ubiquitously used for authentication/encryption, and password guessing attacks are the most effective technique for evaluating password strength. While large language models (LLMs) like ChatGPT-4o have demonstrated remarkable capabilities in text comprehension and reasoning across various general natural language processing tasks, they face limitations due to their static knowledge (e.g., fixed training data that lacks domain adaptability), especially in specialized tasks such as generating accurate password guesses. This work provides a brand new technical route for password guessing, by proposing an LLM-based guessing framework, namely PassLLM, that leverages low-rank adaptation techniques. PassLLM systematically addresses four major password guessing scenarios, each of which is based on varied kinds of information available to the attacker. To reduce the high computation costs in password generation with LLMs, we propose two generation algorithms tailored for trawling and targeted guessing, respectively, enabling efficient password generation at scale (e.g., 1,000 guesses per user). Further, we apply model distillation to improve the generation speed by 11.5 times in trawling guessing scenarios without significantly reducing the success rate. Particularly, our generation algorithms are applicable to a wide range of decoder-only-based LLMs (e.g., Mistral, Llama-2/3, and Qwen-2). Extensive experiments on 11 real-world password datasets demonstrate the effectiveness of our framework: (1) PassLLM for trawling guessing scenarios, whose guessing success rates are generally 2.87%-17.07% higher than its foremost counterpart; (2) PassLLM-I for targeted guessing based on personally identifiable information (PII), which guesses 12.54%-31.63% of common users within 100 guesses, outperforming its foremost counterpart by 15.10%-45.98%; (3) PassLLM-II for targeted guessing based on users' password reuse behaviors, which outperforms its foremost counterpart by 6.31%-13.87%; and (4) PassLLM-III for targeted guessing based on users' PII and sister password(s), which outperforms its foremost counterpart by 13.44%-36.14%. We believe this work makes a substantial step toward introducing LLMs into the password guessing domain.

Visual summary for Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data by Atilla Akkus
Visual summary for Generated Data with Fake Privacy: Hidden Dangers of Fine-tuning Large Language Models on Generated Data by Atilla Akkus

Password Guessing Using Large Language Models

Authors: Yunkai Zou, Maoxiang An, Ding Wang (Nankai University)

Conference: USENIX Security

YouTube: N/A (This is a peer-reviewed paper, not a recorded talk.)

Overview

This groundbreaking paper introduces PassLLM, an innovative framework that leverages Large Language Models (LLMs) with Low-Rank Adaptation (LoRA) techniques for advanced password guessing attacks. Authored by Yunkai Zou, Maoxiang An, and Ding Wang from Nankai University, PassLLM addresses the inherent limitations of general-purpose LLMs in specialized tasks like password guessing, where static knowledge and a disconnect from real-world password creation behaviors hinder their effectiveness. The work presents a novel technical route, demonstrating how modern LLMs can be efficiently fine-tuned and optimized for this critical security domain.

The significance of PassLLM lies in its ability to systematically tackle four major password guessing scenarios: general trawling, targeted guessing based on Personally Identifiable Information (PII), targeted guessing exploiting password reuse behaviors (sister passwords), and a combined multi-source targeted approach. By developing specialized generation algorithms and incorporating model distillation, the framework achieves unprecedented success rates while maintaining computational efficiency. This research not only pushes the boundaries of offensive security capabilities but also provides invaluable insights for evaluating password strength and developing more robust defensive strategies against sophisticated guessing attacks.

The findings presented in this paper are crucial for both attackers seeking to refine their methods and defenders striving to understand and mitigate emerging threats. PassLLM's superior performance across diverse real-world datasets underscores the growing importance of LLM-based techniques in cybersecurity research. It highlights a paradigm shift in password security, moving beyond traditional statistical and deep learning models to embrace the advanced contextual understanding and generation capabilities of large language models, when properly adapted.

Background

Passwords, despite their known security and usability challenges, remain a ubiquitous and fundamental component of authentication, encryption, and digital signatures. Their widespread adoption, however, presents significant risks, particularly when users opt for weak or easily guessable passwords. The landscape of password security is continually shaped by large-scale data breaches, which frequently expose millions of user credentials and associated PII, fueling the effectiveness of guessing attacks.

Historically, password guessing research has evolved through distinct stages. The initial phase, pre-2008, was dominated by heuristic attacks, relying on "creative ideas" to design character transformation rules. This period viewed password guessing as more of an art than a science. From 2009 to 2015, the field transitioned into a more scientific realm, driven by statistical models such as Probabilistic Context-Free Grammars (PCFG) [46] and Markov chains [25, 27]. The third stage, beginning in 2016, saw the rise of deep learning techniques, with researchers exploring Recurrent Neural Networks (RNNs) [26, 30] and Generative Adversarial Networks (GANs) [18, 32] to overcome data sparsity and overfitting issues inherent in traditional statistical methods. Recent advancements include models like RFGuess [45], UNCM [30], and RankGuess [52], which further refined these deep learning and statistical approaches.

The advent of Large Language Models (LLMs), such as Llama [37], Mistral [20], and GPT-4 [28], has revolutionized natural language processing with their remarkable text generation and understanding capabilities. Password guessing, fundamentally an attempt to approximate the probability distribution of textual sequences, appears inherently suited for transformer-based LLMs employing autoregressive generation. However, direct application of LLMs to password guessing faces several significant challenges. As noted by Yang and Wang [52], passwords differ fundamentally from natural language: they are typically short, adhere to a limited character vocabulary (e.g., 95 printable ASCII characters), and demand exact matches, unlike natural language where minor inconsistencies are tolerable.

Key challenges for LLMs in password guessing include:

  • Efficiency and Scale: LLMs, primarily designed for conversational contexts, have slow inference speeds and high computational resource demands. Generating the large number of short, diverse character sequences required for password guessing (e.g., >1,000 guesses per user, or 10^8 for trawling) far exceeds their typical use case. Their default inference algorithms often lead to VRAM requirements exceeding consumer-grade GPUs (e.g., NVIDIA 4090 with 24GB VRAM) and can produce duplicate or insufficient outputs (Section 1).
  • Domain Adaptability and Precision: LLMs' static knowledge, derived from general natural language training data, often fails to accurately model real-world user password creation behaviors. While they might generate plausible-looking passwords, these often do not align with actual user patterns (Section 1). Moreover, the need for exact matches in passwords contrasts with the tolerance for minor grammatical errors in natural language, making precision a critical hurdle.
  • Ethical Concerns: Some commercial LLMs, like GPT-4, refuse to generate extensive password lists due to built-in security and privacy safeguards, further impeding their direct application in this domain (Section 1).

Previous attempts to introduce LLMs into password security, such as PassBERT [51], PassGPT [34], and PagPassGPT [35], have been limited. PassBERT, while transformer-based, does not reach the billion-parameter scale characteristic of modern LLMs. PassGPT, based on GPT-2, was trained from scratch on password leaks and suffered from GPU resource constraints and a relatively small parameter size, limiting its ability to leverage nuanced language understanding from pre-trained models. Other studies, like Atzori et al. [6], found LLM-generated passwords syntactically complex and thus ineffective as guessing dictionaries. These prior works highlight the gap: the unclear effectiveness of truly large, pre-trained generative LLMs (like Llama, Mistral, Qwen) in enhancing password guessing success rates across diverse attack scenarios.

Targeted password guessing, which exploits a victim's PII (e.g., name, birthday, email, username) or sister passwords (leaked from other accounts), represents an even more potent threat. Traditional PII-based models like TarGuess-I [43] and RFGuess-PII [45] rely on type-based PII matching, which struggles to capture the diverse ways users incorporate personal information into passwords. Similarly, password reuse models like TarGuess-II [43], Pass2Path [29], and Pass2Edit [44] primarily focus on single-password reuse, with limited exploration into scenarios involving multiple leaked sister passwords, despite statistics showing users often have several compromised credentials (Section 2). PassLLM directly addresses these limitations by providing a versatile framework capable of effectively leveraging various auxiliary information for targeted attacks.

Key Findings

The PassLLM framework introduces a significant advancement in password guessing, presenting several key findings and contributions:

  • Novel Technical Route: PassLLM is the first billion-sized parameter framework specifically tailored for password guessing, successfully applying LoRA fine-tuning techniques to pre-trained decoder-only LLMs (e.g., Llama, Mistral, Qwen) for this domain. This opens a new avenue for modeling user password guessability (Section 1).
  • Superior Guessing Success Rates: Extensive experiments across 11 real-world password datasets consistently demonstrate PassLLM's effectiveness:
  • Trawling Guessing: PassLLM achieved success rates generally 2.87%-17.07% higher than its foremost counterpart, RankGuess [52]. Notably, it outperformed RankGuess by 3.36%-32.21% on the recently leaked Post Millennial dataset at 10^9 guesses (Section 4).
  • PII-based Targeted Guessing (PassLLM-I): It successfully guessed 12.54%-31.63% of common users within 100 guesses, outperforming the leading RFGuess-PII [45] and TarGuess-I [43] by 15.10%-45.98% and 16.20%-119.87% respectively within 100 guesses (Section 4).
  • Reuse-based Targeted Guessing (PassLLM-II): PassLLM-II surpassed its foremost counterpart, Pass2Edit [44], by 6.31%-13.87% in single-password reuse scenarios. In multi-password reuse, it achieved a 36.63% success rate within 1,000 guesses, significantly outperforming MS-PointerGuess [49] by 40.45%-140.42% (Section 4). It is also the first model capable of simultaneously utilizing any number of sister passwords (Section 4).
  • Multi-source Targeted Guessing (PassLLM-III): When leveraging both PII and sister passwords, PassLLM-III outperformed TarGuess-III [43] by 13.44%-36.14%. With additional auxiliary information (e.g., phone, website name, multiple sister passwords), PassLLM-III+ achieved even higher success rates, surpassing TarGuess-III by 62.64%-87.78% within 100 guesses (Section 4).
  • Tailored Generation Algorithms: Two dedicated and efficient generation algorithms were developed: a two-stage Breadth-First Search (BFS) for trawling and a dynamic beam search for targeted guessing. These algorithms enable efficient and reliable large-scale password generation with regular computational resources (e.g., an RTX 4090 GPU) (Section 3).
  • Model Distillation for Efficiency: For the first time, model distillation [17] was applied to LLM-based password guessing. By distilling a 7-billion-parameter Mistral model into a compact 0.5-billion-parameter Qwen2.5 model (PassLLM-d), an 11.5-fold speed improvement in trawling password generation was achieved (from 260 pw/s to 3,000 pw/s) without significantly reducing the success rate (Section 4).
  • Applicability and Insights: The framework is applicable to a wide range of decoder-only LLMs (e.g., Mistral, Llama-2/3, Qwen-2). Key insights include: (1) larger parameter sizes generally lead to better fine-tuning performance, and (2) while the specific content of the prompt has minimal impact on guessing results, using any prompt significantly outperforms not using one (Section 1, Section 4).

Technical Deep Dive

The PassLLM framework represents a sophisticated approach to password guessing, unifying trawling and targeted attacks under a single LLM-based paradigm. It addresses the unique challenges of password generation through careful architectural choices, efficient fine-tuning, and specialized generation algorithms.

The core idea is to model password guessing as a sequence generation task, where a pre-trained language model $P_{\Phi}(y|x)$ generates a password $pw_i$ conditioned on available auxiliary information $INFO_i$. This information can range from nothing (for trawling) to PII and/or sister passwords (for targeted attacks). The training dataset consists of sequences $seq_i = [INFO_i, pw_i]$, allowing the model to learn the relationship between context and password.

Fine-tuning with Low-Rank Adaptation (LoRA)

Instead of full fine-tuning, which would be computationally prohibitive for billion-parameter LLMs, PassLLM employs Low-Rank Adaptation (LoRA) [19]. LoRA significantly reduces the number of trainable parameters by introducing low-rank matrices into the Transformer's attention mechanism (Section 3). Specifically, for projection matrices like $W_q$ (query), which are typically $d \times d$, LoRA decomposes the update $\Delta W_q$ into two smaller matrices, $A_q \in R^{d \times r}$ and $B_q \in R^{r \times k}$, where $r \ll \min(d, k)$. The original pre-trained matrix $W^0_q$ remains frozen, and only $A_q$ and $B_q$ are trained. This results in $W_q = W^0_q + A_qB_q$, drastically reducing memory and computational costs while retaining performance. This parameter-efficient approach allows PassLLM to specialize pre-trained decoder-only LLMs (like Mistral-7B) for password generation without retraining billions of parameters (Section 3).

Loss Function and Encoding

During training, the model maximizes an autoregressive objective using cross-entropy loss. Crucially, to mitigate the risk of attackers crafting adversarial prompts to extract sensitive PII or sister passwords, the loss calculation is restricted solely to the password prediction portion of the sequence (Section 3). This ensures that sensitive input content is excluded from direct optimization.

For data encoding, PassLLM uses separate vocabularies: the default vocabulary of the LLM for prefixes (auxiliary information and prompt) and a specialized character-level vocabulary for passwords. This password vocabulary consists of 95 printable ASCII characters plus an End-Of-Sequence (EOS) token, aligning with the nature of password data where over 99% of real-world passwords use this character set (Section 3).

Trawling Password Generation Algorithm

For trawling guessing, which aims to generate a large-scale generic dictionary (e.g., 10^8 passwords), PassLLM proposes a two-stage Breadth-First Search (BFS) algorithm with probability threshold pruning (Algorithm 1, Section 3.2.1). This design facilitates efficient parallelization:

  1. Stage 1: Prefix Generation: A Priority Queue is used to generate a set of non-overlapping prefixes (e.g., "1234", "zxcv"). The algorithm starts with single-character prefixes and iteratively expands them by appending the 95 ASCII characters. Prefixes are dequeued based on their cumulative probability, and new expanded prefixes are added. If a prefix combined with an EOS token exceeds a probability threshold $\tau$, it's added to the final password set. This stage ensures coverage of various password generation paths and balances the size of sub-dictionaries in the next stage.
  2. Stage 2: Password Generation: Each generated prefix from Stage 1 serves as an independent seed. It's concatenated with the initial prompt, and then a BFS algorithm with probability threshold pruning generates sub-dictionaries. These sub-dictionaries are then merged and ranked by probability to form the final password dictionary. This decomposition allows for efficient large-scale parallelization (Section 3.2.1).

Targeted Password Generation Algorithm

Targeted guessing scenarios, which typically involve a smaller number of guesses (e.g., 100-1,000), require a more focused approach. PassLLM employs a dynamic beam search strategy (Algorithm 2, Fig. 3, Section 3.2.2) to balance efficiency and performance, particularly under VRAM constraints. Key features include:

  • Dynamic Beam Width: The algorithm uses a list of beam widths $K[1:m]$ that can vary at each generation depth $i$, allowing for flexible resource allocation.
  • Batch Processing: Candidates for the next depth are divided into batches of size $B$ to optimize GPU utilization and mitigate VRAM exhaustion, especially for large beam widths.
  • EOS Termination Threshold: Unlike traditional beam search, the EOS token is not treated as a normal candidate for beam ranking. A sequence is only forcibly ended and added to the output set if the model's predicted $Pr(EOS|\cdot)$ for that sequence exceeds a threshold $\epsilon$. This prevents the inclusion of incomplete or unrealistic passwords that might have high prefix probabilities but low termination probability, ensuring generated guesses reflect actual user behavior.
  • KV Cache Optimization: To reduce redundant computation and memory overhead, the Key-Value (KV) cache for the auxiliary information ($INFO_i$) is computed once (CacheINFO) and shared across all beams during generation. Only a separate KV cache for the evolving password prefix is maintained for each beam. This significantly reduces VRAM consumption compared to conventional beam search, which would store beam\_width instances for both components (Section 3.2.2).

Demo / Proof of Concept

As this is a peer-reviewed technical paper and not a live conference talk, the "demo" or "proof of concept" is thoroughly demonstrated through extensive experimental evaluation on a wide array of real-world datasets. The authors conducted comprehensive comparisons against state-of-the-art password guessing models across all four primary attack scenarios.

The evaluation leveraged 11 large-scale password datasets, totaling over 3.3 billion plaintext passwords. These included six English-language datasets (Post Millennial, 000Webhost, Rockyou, Rootkit, ClixSense) and four Chinese-language datasets (Taobao, 126, CSDN, Dodonew), plus the massive mixed-language COMB dataset (3.28 billion email-password pairs). PII datasets like 12306-PII, Dodonew-PII, CSDN-PII, 000Webhost-PII, Rootkit-PII, and Clixsense-PII were also used for targeted attacks (Section 4).

Experimental Setup:

The fundamental LLM chosen for PassLLM was Mistral-7B [20], which demonstrated slightly superior performance compared to other models like Llama2/3, GPT-2, and Qwen2.5 of varying sizes (Fig. 4a, Section 4). Training typically involved 1 million samples from 80% of a dataset (e.g., Rockyou), with 100,000 samples from the remaining 20% for testing. For targeted guessing, smaller datasets (e.g., 50,000 samples) were found sufficient for performance stabilization (Fig. 4c, Section 4). The impact of prompts was also evaluated, revealing that while different prompt roles had minimal impact on guessing success, using any prompt significantly outperformed no prompt at all (Fig. 4b, Section 4).

Trawling Guessing (PassLLM):

PassLLM was benchmarked against models such as RankGuess [52], RFGuess [45], PassGPT [34], PagPassGPT [35], PCFG [46], 3/4-order Markov [25], and FLA [47]. Using a Monte Carlo algorithm to approximate crack rates (applicable to PassLLM and a modified PassGPT+), PassLLM consistently matched or outperformed all counterparts across one-site (1M Rockyou → rest Rockyou) and cross-site (1M Rockyou → 000Webhost, 1M Rockyou → Post Millennial) scenarios (Fig. 5, Section 4). Its generalization ability was particularly evident on the recent Post Millennial dataset, where it significantly outperformed RankGuess by 3.36%-32.21% at 10^9 guesses.

PII-based Targeted Guessing (PassLLM-I):

Compared against TarGuess-I [43] and RFGuess-PII [45], PassLLM-I demonstrated superior effectiveness. In various Chinese and English PII-based scenarios, PassLLM-I achieved significantly higher guessing success rates within 100 guesses, outperforming TarGuess-I by 16.20%-119.87% and RFGuess-PII by 15.10%-76.27% (Fig. 6, Section 4). The key advantage was its direct learning from PII within prompts, avoiding the precision loss from converting PII into fixed labels.

Reuse-based Targeted Guessing (PassLLM-II):

PassLLM-II was evaluated against TarGuess-II [43], Pass2Path [29], Pass2Edit [44], PassBERT [51], PointerGuess [49], and MS-PointerGuess [49]. In single-password reuse scenarios (e.g., $pw_1 \rightarrow pw_2$ from COMB dataset), PassLLM-II outperformed Pass2Edit by 8.27%-9.66% within 1,000 guesses (Fig. 7a, Section 4). For two-password reuse, it significantly surpassed MS-PointerGuess by 40.45%-140.42% (Fig. 7c, Section 4). PassLLM-II also uniquely demonstrated the ability to utilize any number of sister passwords ($n \ge 3$) for multi-password reuse, achieving a 36.63% success rate within 1,000 guesses for target passwords different from sister passwords (Fig. 7c, Section 4).

Multi-source Targeted Guessing (PassLLM-III):

Addressing scenarios where both PII and sister passwords are available, PassLLM-III was compared with TarGuess-III [43]. PassLLM-III, using the same auxiliary information as TarGuess-III, achieved 13.44%-36.14% higher success rates. When enhanced with additional PII (phone number, website name) and multiple sister passwords (PassLLM-III+), its success rate further improved, outperforming TarGuess-III by 62.64%-87.78% within 100 guesses (Fig. 8, Section 4). This highlights the compounded threat when multiple data points are compromised.

Model Distillation:

To enhance efficiency, model distillation was applied, using the 7B-parameter PassLLM as a teacher model and Qwen2.5-0.5B as the student (PassLLM-d). This process, combining KL divergence and cross-entropy loss, resulted in an 11.5-fold speedup in trawling password generation (from 260 pw/s to 3,000 pw/s) on a single RTX 4090 GPU, with only a negligible reduction in guessing success rate (Table 7, Fig. 9, Section 4). This demonstrates the practical feasibility of deploying such models.

Defensive Implications

The PassLLM framework, while a powerful offensive tool, yields critical insights and applications for defensive security measures. Its capabilities highlight vulnerabilities in current password practices and provide mechanisms for more effective evaluations and countermeasures.

Digital Forensics: PassLLM's ability to leverage diverse contextual data, including PII and multiple sister passwords, for targeted guessing attacks makes it an invaluable tool in digital forensics. Law enforcement and forensic specialists can utilize this framework to unlock critical devices of suspects, recover hashed password files, or address other urgent forensic needs where access to digital assets is paramount. Its high success rates mean that even with limited information, there's a greater chance of recovering sensitive credentials (Section 5).

Enhanced Password Strength Meters (PSMs): A primary application of sophisticated password guessing models is to accurately evaluate password strength from an attacker's perspective. PassLLM can serve as a highly effective server-side or locally deployed Password Strength Meter (PSM). By providing rapid and accurate estimates of the guess number required to crack a given password, it offers a more realistic assessment of its vulnerability. Experiments showed that PassLLM-PSM has fewer unsafe errors compared to leading PSMs like FLA-PSM [47] and RankGuess-PSM [52], indicating its superior accuracy in identifying weak passwords (Section 5).

Interpretable PSMs and Specialized LLMs: The paper identifies a significant opportunity for future development: interpretable PSMs. While current general-purpose LLMs like ChatGPT-4o can offer natural language explanations and optimization suggestions, their advice can sometimes be misleading due to their generalized training data. The vision is to pre-train specialized LLMs using password-specific corpora and fine-tune them with expert-constructed training data that includes interpretable password strength evaluations. This would enable a PSM that not only accurately assesses strength but also provides contextually relevant, up-to-date, and actionable advice tailored to user-provided information, moving beyond outdated recommendations like frequent password changes (Section 5).

Addressing Direct Password Generation Vulnerabilities: The research reveals that current LLMs, even those with built-in security mechanisms designed to prevent direct password generation from PII, can sometimes be circumvented with minor prompt modifications (e.g., switching to Chinese language prompts for ChatGPT-4o). This highlights a need for more robust ethical AI safeguards within LLMs to prevent their misuse in password attacks. Furthermore, the low success rates of general LLMs when directly prompted to generate passwords (typically <2% within 20 guesses) underscore the necessity of building specialized LLMs for password security. Such models, trained specifically on password-related datasets, would better model real-world user password creation behaviors, enabling them to provide more accurate security assessments and guidance (Section 5).

Future Directions for Large-Scale Offline Guessing: While PassLLM significantly improves efficiency, generating truly massive guess dictionaries (e.g., over 10^10) for trawling attacks remains a challenge even for optimized LLMs. A promising defensive implication, explored as future work, is to leverage LLMs not for direct dictionary generation, but to produce scalable transformation rules based on an input set of passwords. These rules could then be applied to vast datasets using traditional, highly optimized methods to efficiently generate extensive guess dictionaries. This approach could bridge the gap between LLM intelligence and the sheer scale required for offline attacks, necessitating new defensive strategies against rule-based attacks (Section 5).

Key Takeaways

  • The PassLLM framework introduces a novel technical route for password guessing by effectively applying Low-Rank Adaptation (LoRA) fine-tuning to billion-parameter decoder-only LLMs like Mistral-7B.
  • PassLLM achieves state-of-the-art guessing success rates across all four major attack scenarios: trawling, PII-based targeted, reuse-based targeted, and multi-source targeted guessing, significantly outperforming previous models by up to 140% in some scenarios.
  • Two tailored generation algorithms—a two-stage BFS for trawling and a dynamic beam search with KV cache optimization for targeted guessing—enable efficient and reliable password generation at scale, even with consumer-grade GPUs.
  • The application of model distillation successfully reduced PassLLM's parameter size from 7B to 0.5B, resulting in an 11.5-fold speedup in trawling password generation without compromising guessing success rates.
  • PassLLM provides crucial insights for digital forensics and serves as a superior Password Strength Meter (PSM), demonstrating fewer unsafe errors than existing models, thereby improving the evaluation of password vulnerability.
  • The research highlights the critical need for specialized LLMs for password security, capable of providing accurate and interpretable advice, and underscores the ongoing challenge of mitigating LLM misuse in password-related attacks.

About the Speaker(s)

The technical article is based on a peer-reviewed paper authored by Yunkai Zou, Maoxiang An, and Ding Wang, all affiliated with Nankai University. Their research focuses on advancing the understanding and capabilities of password guessing techniques, particularly through the application of cutting-edge machine learning and large language models. Their work contributes significantly to both offensive and defensive aspects of cybersecurity, pushing the boundaries of what is possible in password security research.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Solid applied ML security research that actually moves the needle on password guessing. The LoRA fine-tuning approach is sensible, the eval is comprehensive across 11 datasets, and the distillation work makes this practically deployable. Not revolutionary—it's applying known techniques to a well-studied problem—but the execution is clean and the results are real.

Heather Calloway (CISO) — SOLID

This is serious research that changes how I'd assess credential risk. LLMs fine-tuned for password guessing outperform existing models by double-digit percentages across trawling, PII-targeted, and password-reuse scenarios. Every CISO with a breach notification program or password policy on the books needs to understand what this means for their exposure calculations.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)