Less Defined Knowledge and More True Alarms: Reference-based Phishing Detection without a Pre-defined Reference List

Ruofan Liu (National University of Singapore), Yun Lin, Xiwen Teoh, Gongshen Liu, Zhiyong Huang, Jin Song Dong

33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24

Overview

The digital landscape is under constant siege from sophisticated phishing attacks, a threat that has escalated dramatically from 26,000 reported victims in 2018 to an alarming 300,000 in 2023. This exponential growth underscores the critical need for more robust and adaptive detection mechanisms. Traditional reference-based phishing detection systems, while effective to a degree, are increasingly struggling to keep pace with the dynamic nature of these attacks, primarily due to their reliance on manually curated and static lists of known brands and their associated digital representations.

Watch on YouTube

Visual summary for Less Defined Knowledge and More True Alarms: Reference-based Phishing Detection without a Pre-defined Reference List by Ruofan Liu, Yun Lin, Xiwen Teoh, Gongshen Liu, Zhiyong Huang, Jin Song Dong
Visual summary for Less Defined Knowledge and More True Alarms: Reference-based Phishing Detection without a Pre-defined Reference List by Ruofan Liu, Yun Lin, Xiwen Teoh, Gongshen Liu, Zhiyong Huang, Jin Song Dong

Key moments

  1. 0:00 Introduction and limitations of current phishing detection
  2. 1:49 Introducing FALM: a new reference-less detection system
  3. 2:38 Detailed breakdown of FALM's detection process
  4. 3:39 FALM's significant performance improvements and discoveries
  5. 4:18 Live demo and example of FALM in action
  6. 4:56 Practical deployment scenarios for FALM
  7. 5:40 Conclusion: FALM's benefits and future impact

Less Defined Knowledge and More True Alarms: Reference-based Phishing Detection without a Pre-defined Reference List

Speakers: Ruofan Liu, National University of Singapore; Yun Lin, Shanghai Chong University; Xiwen Teoh, National University of Singapore; Gongshen Liu, Shanghai Chong University; Zhiyong Huang, National University of Singapore; Jin Song Dong, National University of Singapore

Conference: USENIX Security '24

YouTube: https://www.youtube.com/watch?v=2wqkjasl_eI

Overview

The digital landscape is under constant siege from sophisticated phishing attacks, a threat that has escalated dramatically from 26,000 reported victims in 2018 to an alarming 300,000 in 2023. This exponential growth underscores the critical need for more robust and adaptive detection mechanisms. Traditional reference-based phishing detection systems, while effective to a degree, are increasingly struggling to keep pace with the dynamic nature of these attacks, primarily due to their reliance on manually curated and static lists of known brands and their associated digital representations.

In response to these pervasive challenges, Ruofan Liu and their colleagues from the National University of Singapore and Shanghai Chong University presented FishLM, an innovative phishing detection system. FishLM fundamentally rethinks the approach to identifying phishing threats by entirely eliminating the need for a predefined reference list. Instead, it leverages the advanced capabilities of language models (LMs) to "read" and comprehend the content of a webpage, analyzing both visual cues like logos and, crucially, the underlying textual semantics.

This paradigm shift allows FishLM to access a far more extensive and implicitly updated knowledge base of brands and their characteristics than any static list could ever provide. By focusing on the intrinsic understanding of a webpage's intent—particularly its credential-taking ambition—FishLM offers a more adaptable, scalable, and accurate solution to combat the ever-evolving threat of phishing. The system promises significant improvements in recall rates, a broader coverage of target brands, and enhanced operational efficiency, marking a substantial step forward in proactive cybersecurity defense.

Background

▶ Watch: Introduction and limitations of current phishing detection (0:00)

The prevailing strategy for detecting phishing websites has long revolved around reference-based detection. This method operates by maintaining an extensive database, or "reference list," of legitimate brands, their official domains, logos, visual layouts, and other identifying characteristics. When a suspicious website is encountered, its domain and content are meticulously compared against this predefined reference list. For instance, if a site like abc.com attempts to mimic adobe.com, a mismatch in the domain or a visual discrepancy in the logo or layout would trigger a phishing alert.

While this approach has demonstrated effectiveness against known threats, it is plagued by two significant limitations that hinder its long-term viability and comprehensive coverage. Firstly, the reliance on an extensive, predefined reference list necessitates continuous and resource-intensive maintenance. The digital landscape is constantly evolving, with new brands emerging, existing brands rebranding, and attackers frequently targeting less-known entities. Keeping this list current and comprehensive is a monumental task, often leading to a reactive posture where new phishing targets are only added after an attack has occurred. This limitation means that new and emerging phishing sites targeting brands outside the reference list often go undetected.

Secondly, traditional reference-based systems predominantly focus on the visual semantics of a web page. They excel at identifying copied logos, similar layouts, or slight variations in design. However, this narrow focus often overlooks other critical elements, particularly the texture content—the actual text, forms, and underlying narrative of the page. Phishers can cleverly craft textual content that, while visually distinct, clearly indicates a credential-taking intention or impersonates a brand through language rather than just imagery. By neglecting this textual layer, traditional systems miss crucial indicators of a phishing attempt, offering an incomplete defense. The emergence of FishLM directly addresses these shortcomings by moving beyond static lists and superficial visual analysis to embrace a more dynamic and semantically aware detection methodology.

Key Findings

▶ Watch: Detailed breakdown of FALM's detection process (2:38)

FishLM introduces a significant advancement in phishing detection, demonstrating substantial improvements over existing methodologies, particularly those reliant on static brand reference lists. The research highlights several key findings that underscore its efficacy and innovative approach:

Firstly, FishLM achieved a significant improvement in recall rate on the "DD fishing Keys" dataset. While specific percentages for this improvement were not detailed in the provided transcript, the emphasis on "significant improvement" indicates a marked increase in the system's ability to correctly identify actual phishing sites. This suggests that FishLM is more effective at catching threats that might evade traditional detectors.

Secondly, FishLM drastically reduces the detection runtime. Compared to Dino Fish, a state-of-the-art system that also expands brand reference lists, FishLM reduced the round time by 2.6 seconds. This efficiency gain is crucial for real-time deployment scenarios where rapid analysis is paramount, allowing for quicker identification and mitigation of threats.

Perhaps one of the most compelling findings is FishLM's performance in an open-world environment. The system was able to discover two times the number of phishing sites compared to traditional methods. More importantly, a higher number of distinct target brands were covered, with the majority of these brands not being included in any static reference list. This capability directly addresses one of the primary limitations of conventional systems: their inability to detect phishing attacks targeting new, niche, or previously uncataloged brands. FishLM's ability to infer brand identity without prior definition allows it to cast a much wider net, offering protection against a broader spectrum of threats.

In summary, FishLM not only enhances the accuracy and speed of phishing detection but also fundamentally expands the scope of protection by moving beyond the constraints of predefined knowledge. Its capacity to identify novel phishing campaigns and protect against attacks on a diverse range of brands, many of which would be invisible to traditional systems, represents a critical leap forward in cybersecurity defense.

Technical Deep Dive

▶ Watch: FALM's significant performance improvements and discoveries (3:39)

FishLM's architecture is designed around a multi-stage process that leverages advanced language models (LMs) to interpret and analyze various aspects of a web page without relying on a predefined list of brands. The system's core innovation lies in its ability to infer brand identity and credential-taking intent dynamically.

The detection process begins with initial data acquisition and processing. When a URL is submitted for scanning, FishLM captures a screenshot of the web page. This visual data is then subjected to several preprocessing steps to generate two crucial descriptions: a local description focusing on visual elements like brand logos, and a comprehensive web page description encompassing all visible textual and structural content.

The first major technical module is Brand Inference and Validation.

  1. Brand Logo Recognition and Description Generation: FishLM first identifies and extracts potential brand logos from the web page screenshot. These visual cues are then processed to generate an expressive textual description of the logo.
  2. LM-based Target Brand Inference: This textual description of the logo, along with the broader web page description, is fed into a sophisticated LM. The LM, leveraging its vast knowledge base, attempts to infer the target brand that the web page is attempting to impersonate. For example, if the page displays a specific blue "W" logo and text related to cloud storage, the LM might infer "OneDrive."
  3. Target Brand Validation: To prevent potential misinformation or hallucinations from the LM, FishLM incorporates a crucial validation step. The inferred target brand is cross-referenced by comparing the web page's logo with all alternative logos available on the web for that inferred brand. This ensures that the visual identity on the page genuinely matches the LM's inferred brand, adding a layer of robustness to the detection.

Following brand inference, FishLM proceeds to analyze the page for tell-tale signs of malicious intent.

  1. Brand-Domain Inconsistency Check: Once a target brand is confidently identified and validated, FishLM performs a brand-domain inconsistency check. It compares the inferred legitimate brand's official domain(s) with the actual domain of the suspicious web page. A significant mismatch here is a strong indicator of a phishing attempt.

The next critical component is the Credential Taking Prediction Module (CRP).

  1. LM-based Credential Taking Prediction: The full web page description is fed into the LM again, this time to determine its credential taking status. The LM analyzes the textual content, form fields, and calls to action to ascertain if the page is designed to solicit sensitive user information like usernames, passwords, or financial details. This module leverages the LM's natural language understanding capabilities to infer intent, a capability that traditional visual-centric methods often miss.
  2. CRP Transition Module: If the initial CRP prediction indicates that the web page is not currently a credential-taking page, FishLM doesn't stop there. It employs a CRP transition module. This module attempts to predict which user interface (UI) element (e.g., a "login" button, a "continue" link) on the current page is most likely to transition the user to a credential-taking page. FishLM then auto-clicks this UI element and captures an updated screenshot. The CRP prediction process is then repeated on this new page, allowing FishLM to detect multi-stage phishing campaigns where the initial landing page might appear benign.

Finally, the system synthesizes its findings:

  1. Final Decision: The information derived from the brand-domain consistency check and the credential taking intention (from both initial and transitioned CRP analyses) are combined. If both indicators strongly suggest impersonation and credential solicitation, FishLM raises a phishing alert.

This multi-faceted approach, deeply rooted in the analytical power of large language models, allows FishLM to move beyond superficial comparisons and delve into the semantic and intentional layers of a web page, enabling it to detect sophisticated phishing attempts that bypass conventional defenses.

Demo / Proof of Concept

▶ Watch: Practical deployment scenarios for FALM (4:56)

The efficacy and operational flow of FishLM were demonstrated through a live demo site, accessible via a provided URL, offering a tangible illustration of its capabilities. The demo interface was designed for user interaction, featuring a left panel where users could adjust default hyperparameters, allowing for customization of the detection process, though specific parameters were not detailed.

The core of the demo involved users inputting a suspicious URL into the system. Upon submission, FishLM initiated its analysis sequence:

  1. Screenshot Capture: The system first captured a visual screenshot of the target web page, providing a comprehensive visual record for subsequent analysis.
  2. Preprocessing: Following the screenshot, FishLM performed its initial preprocessing steps. This involved generating both the local description (focused on visual elements like logos) and the broader web page description (encompassing all textual and structural content).
  3. Target Brand Inference: The generated descriptions were then fed into the underlying language model (LM). In the specific example demonstrated, the LM successfully inferred the target brand as OneDrive. This showcased FishLM's ability to identify the impersonated entity purely from page content, without needing a predefined list.
  4. Credential Taking Status Prediction: Next, the LM was tasked with predicting the credential taking status of the page. In the demo example, the LM determined that the page was "already a credential taking page," indicating its direct intent to solicit sensitive user information. This step highlights the LM's capability to understand the functional purpose and malicious intent embedded within the page's text and structure.
  5. Final Decision: Based on the combined analysis of brand inference, domain consistency (implicitly checked), and credential taking intent, FishLM reached its final decision, confirming the presence of a phishing threat.

The demo effectively illustrated FishLM's seamless workflow, from URL input to a definitive phishing alert, emphasizing its reliance on sophisticated LM capabilities for dynamic brand recognition and intent analysis. The OneDrive example served as a clear proof of concept, validating the system's ability to accurately identify and flag phishing attempts without the limitations of traditional reference lists.

Defensive Implications

▶ Watch: Conclusion: FALM's benefits and future impact (5:40)

FishLM's innovative, reference-list-free approach to phishing detection offers significant defensive implications across various cybersecurity domains, enabling more proactive and comprehensive protection against evolving threats. The speakers outlined three primary scenarios for its deployment:

  1. As a URL Consumption Service: FishLM can be seamlessly integrated into existing infrastructure as a URL consumption service. This allows for real-time scanning of incoming URLs before they reach end-users.
  • Email Server Integration: By deploying FishLM at the email server site, organizations can scan URLs embedded in emails as they arrive. This acts as a crucial first line of defense, preventing malicious links from ever reaching employee inboxes, thereby significantly reducing the attack surface for email-based phishing campaigns.
  • Browser Plugin: Alternatively, FishLM can be implemented as a browser plugin. This empowers individual users by scanning URLs in real-time as they navigate the web. If a user inadvertently clicks a suspicious link, the plugin can intercept the request, analyze the page, and issue an immediate warning or block access, protecting users even if an initial email filter was bypassed. This personalizes protection, making it accessible directly at the user's point of interaction.
  1. As a Threat Intelligence Feed: FishLM's ability to actively scan and identify emerging phishing sites makes it an invaluable tool for generating threat intelligence feeds.
  • Active Scanning: Security teams can deploy FishLM to continuously scan newly registered domains, suspicious IPs, or web pages flagged by other preliminary systems.
  • Zero-Day Phishing Alerts: Because FishLM does not rely on predefined lists, it is uniquely positioned to discover zero-day phishing alerts—attacks targeting new brands or employing novel techniques that have not yet been cataloged. These alerts can then be aggregated and distributed as actionable threat intelligence, informing firewalls, intrusion detection systems, and security operations centers (SOCs) about nascent threats. This proactive intelligence gathering capability allows defenders to anticipate and block new campaigns before they gain widespread traction.
  1. Collaboration with Web Hosting Providers: A critical area for impact is through collaboration with web hosting providers.
  • Monitoring Exploited Services: Web hosting providers often unknowingly host phishing sites that exploit their services. FishLM can be integrated into their infrastructure to monitor potential phishing sites residing on their servers.
  • Effective Takedowns: By automatically identifying malicious pages, hosting providers can quickly and effectively take them down. This not only protects their reputation and prevents abuse of their services but also disrupts the attacker's infrastructure, making it harder for phishers to maintain their operations. This partnership fosters a more secure internet ecosystem by tackling the problem at its root—the hosting infrastructure.

In essence, FishLM provides versatile tools for defenders, moving beyond reactive measures to enable proactive threat hunting, real-time user protection, and systemic disruption of phishing operations. Its adaptability and comprehensive brand knowledge coverage make it a powerful asset in the ongoing battle against cyber fraud.

Key Takeaways

  • Elimination of Predefined Reference Lists: FishLM fundamentally innovates by removing the need for static, resource-intensive reference lists of known brands, overcoming a major limitation of traditional phishing detection.
  • Leveraging Advanced Language Models (LMs): The system utilizes LMs to "read" and understand both the visual (logos) and textual semantics of a webpage, enabling dynamic inference of target brands and malicious intent.
  • Enhanced Brand Knowledge Coverage: FishLM significantly expands the scope of detection, discovering twice the number of phishing sites and covering a higher number of distinct target brands, many of which are not included in traditional static lists.
  • Improved Efficiency and Recall: The system demonstrates a reduced runtime of 2.6 seconds compared to Dino Fish and a significant improvement in recall rates, leading to faster and more accurate threat identification.
  • Focus on Credential Taking Intent: FishLM's Credential Taking Prediction Module (CRP), including a CRP Transition Module, allows it to identify pages designed to solicit user credentials, even across multi-stage phishing campaigns.
  • Versatile Deployment Scenarios: FishLM can be deployed as a URL consumption service (email servers, browser plugins), a threat intelligence feed for zero-day phishing alerts, and in collaboration with web hosting providers for proactive takedowns.

About the Speaker(s)

The research behind FishLM is a collaborative effort by academics from the National University of Singapore and Shanghai Jiao Tong University. The talk at USENIX Security '24 was presented by Ruofan Liu from the National University of Singapore. The contributing authors also include Yun Lin and Gongshen Liu from Shanghai Jiao Tong University, alongside Xiwen Teoh, Zhiyong Huang, and Jin Song Dong from the National University of Singapore. Their collective expertise in computer science and security research has culminated in this novel approach to combat phishing threats.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

FishLM presents a robust, LM-driven phishing detection system that finally breaks free from the limitations of static reference lists. Its ability to dynamically infer brand identity and credential-taking intent, even across multi-stage attacks, offers a genuinely impactful and scalable solution for real-world defense. This is a significant step forward in proactive phishing mitigation.

Heather Calloway (CISO) — STRONG ACCEPT

This research presents a compelling shift in phishing detection, moving beyond static reference lists to dynamic, LM-driven inference. It significantly enhances an organization's ability to identify and respond to zero-day and emerging phishing threats, offering clear pathways for operational integration and broader threat intelligence. The focus on proactive detection against unknown targets directly addresses a critical business risk and improves institutional accountability.

→ Top-rated talks at 33rd USENIX Security Symposium

All talks from 33rd USENIX Security Symposium