KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection
Yuexin Li, Mei Lin Lock, Tri Cao, Nay Oo, Hoon Wei Lim, Bryan Hooi
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
Phishing attacks remain a pervasive and costly threat, leading to significant financial losses globally. Despite advancements in detection mechanisms, a critical gap persists in effectively identifying sophisticated phishing campaigns that target a vast array of brands or employ subtle evasion tactics. The talk "KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection" introduces a novel, comprehensive approach to significantly bolster the accuracy and coverage of phishing detection systems. Presented by Yuexin Li from the National University of Singapore and their collaborators, this work addresses the inherent limitations of existing reference-based detectors, which often struggle with limited brand knowledge and an inability to detect phishing pages that do not overtly display logos.

Key moments
- 0:00 Introduction to phishing and existing detector limitations
- 2:20 KnowPhish's large-scale brand knowledge base solution
- 3:55 Key insight: High-risk industries indicate phishing targets
- 4:50 Category and popularity-based brand search algorithms
- 6:00 Acquiring and augmenting multimodal brand knowledge
- 7:00 Using LLMs for textual brand intention in logo-less pages
- 8:40 Beginning of KnowPhish's empirical evaluation
KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection
Speakers: Yuexin Li, Mei Lin Lock, Tri Cao, Nay Oo, Hoon Wei Lim, Bryan Hooi
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=wQuB7yJuSV4
Overview
Phishing attacks remain a pervasive and costly threat, leading to significant financial losses globally. Despite advancements in detection mechanisms, a critical gap persists in effectively identifying sophisticated phishing campaigns that target a vast array of brands or employ subtle evasion tactics. The talk "KnowPhish: Large Language Models Meet Multimodal Knowledge Graphs for Enhancing Reference-Based Phishing Detection" introduces a novel, comprehensive approach to significantly bolster the accuracy and coverage of phishing detection systems. Presented by Yuexin Li from the National University of Singapore and their collaborators, this work addresses the inherent limitations of existing reference-based detectors, which often struggle with limited brand knowledge and an inability to detect phishing pages that do not overtly display logos.
The core innovation of KnowPhish lies in its integration of a large-scale, multimodal brand knowledge base with Large Language Models (LLMs) for advanced textual analysis. This dual approach allows KnowPhish to dramatically expand the spectrum of identifiable phishing targets—from a mere 277 brands in prior work to over 20,000—and to detect logo-less phishing pages by analyzing their HTML content. By leveraging the power of knowledge graphs and LLMs, KnowPhish not only improves the recall of existing detectors but also introduces a robust mechanism for identifying brand intentions even when visual cues are absent, thereby offering a crucial step forward in the ongoing battle against phishing.
Background
▶ Watch: Introduction to phishing and existing detector limitations (0:00)
Reference-based phishing detectors have emerged as a state-of-the-art solution due to their generalizability and explainability. These systems typically operate by employing computer vision models to identify brand intentions from logos present on a webpage. Once a brand is recognized, its domain is compared against a list of legitimate domains for that brand. If the domain does not match, the page is flagged as phishing; otherwise, it is classified as benign. This methodology has proven effective for many conventional phishing attempts where brand logos are prominently displayed.
However, existing reference-based detectors like Fishpedia and FishIntention suffer from two significant limitations that KnowPhish aims to overcome. Firstly, they are constrained by a limited scale of brand knowledge, typically maintaining information for only a few hundred brands (e.g., 277 brands). This narrow scope is insufficient given the vast diversity of phishing targets, which can encompass thousands of brands across various industries. Phishers frequently adapt their targets, making a small, static brand knowledge base quickly outdated and ineffective against emerging threats. Secondly, these prior solutions are primarily logo-centric, meaning they struggle with logo-less phishing web pages. These pages intentionally omit or obscure brand logos, instead conveying brand intention through text or other HTML elements. Since current systems predominantly rely on visual logo analysis, they are blind to such evasive tactics, leading to a high rate of false negatives for these increasingly common phishing variants.
Key Findings
▶ Watch: Key insight: High-risk industries indicate phishing targets (3:55)
The KnowPhish project yields several critical findings that address the identified challenges in phishing detection:
- Construction of a Large-Scale, Multimodal Brand Knowledge Base: KnowPhish successfully constructed a knowledge base covering over 20,000 potential phishing targets, a significant expansion from the 277 brands supported by previous methods. This base provides rich multimodal information including logos, legitimate domains, and alias variants, drastically improving the coverage and efficacy of reference-based detectors.
- Identification of High-Value Industries as Consistent Phishing Targets: Through extensive empirical analysis across different datasets and time periods, the researchers found that phishing targets consistently belong to a set of 10 high-value industries. This insight forms the basis for proactively identifying potential targets, demonstrating that the industry of a target is a time-invariant indicator of phishing risk.
- Effective Detection of Logo-less Phishing Pages using LLMs: KnowPhish introduces a novel approach to detect phishing pages where brand intention is conveyed textually rather than visually. By leveraging an LLM to analyze HTML content and extract predicted brand names, the system can identify brand intentions even when logos are absent, a capability previously lacking in image-based detectors.
- Enhanced Performance of Reference-Based Detectors: Empirical evaluations showed that the KnowPhish knowledge base directly enhances existing reference-based phishing detectors, leading to significantly higher recall with negligible impact on precision. For instance, the system demonstrated much lower runtime compared to a recent baseline like DinaFish, while covering a greater number of fishing targets.
- Validation in Real-World Scenarios: A field study using real-world data from a Singapore governmental agency confirmed KnowPhish's effectiveness, particularly in identifying local phishing targets (e.g., Singapore Post, DBS Bank, Shopee). This study also highlighted the prevalence of logo-less phishing web pages in the real world, further validating the necessity and success of KnowPhish's multimodal detection capabilities.
Technical Deep Dive
▶ Watch: Category and popularity-based brand search algorithms (4:50)
The technical foundation of KnowPhish is built upon two primary pillars: the construction of a large-scale, multimodal brand knowledge base, and the development of a multimodal phishing detector capable of handling textual brand intentions using Large Language Models.
Constructing the KnowPhish Knowledge Base
Addressing the limitation of limited brand knowledge, KnowPhish develops a scalable method to construct a comprehensive knowledge base. The process begins with an insightful observation: what truly indicates a potential phishing target? Through empirical analysis of two distinct phishing datasets, the researchers discovered that while specific phishing targets change significantly over time, the industries they belong to remain consistent. Specifically, 10 high-value industries were identified as perennial targets. This crucial insight is formalized as a fact triplet in a knowledge graph, where the belonging of a brand to a high-risk industry indicates its likelihood as a phishing target.
This insight guides the Brand Search component, which comprises two complementary algorithms:
- Category-based Brand Search: This method leverages WikiData categories to find brands. The researchers manually curated two lists of WikiData categories:
- Narrow categories: Directly representing the identified 10 high-value industries.
- General categories: Broader categories that might not directly represent high-value industries but are useful for augmentation.
By searching WikiData using the narrow categories and their subcategories, a foundational list of brands belonging to high-risk industries is obtained.
- Popularity-based Brand Search: Recognizing that WikiData information might be incomplete, this component augments the brand list. It utilizes the general WikiData categories but adds a crucial filter: domain popularity. The rationale is that more popular brands are inherently more attractive phishing targets. Therefore, brands from general categories with high domain popularity are included, ensuring a comprehensive list of potential targets.
Once potential targets are identified, the Knowledge Acquisition and Augmentation phase begins. This process collects rich multimodal brand knowledge from multiple high-quality sources:
- WikiData: Provides initial brand information, including some logos and domains.
- Tranco Top Domain List: A reputable source for legitimate domain information, used to augment and validate legitimate domains for identified brands.
- Google Image Search: Used to gather diverse logo variants for each brand, accounting for different visual representations.
The gathered knowledge includes local variants (different logo styles), legitimate domains, and crucially, alias variants (alternative names or spellings of a brand). These alias variants are essential for the subsequent text-based detection. The resulting KnowPhish knowledge base, covering over 20,000 brands, is designed to be a standalone resource that can instantly enhance any existing reference-based phishing detector without requiring continuous, costly maintenance.
KnowPhish Detector (KPD): Multimodal Phishing Detection
The second technical pillar is the KnowPhish Detector (KPD), a multimodal system designed to overcome the logo-less phishing problem. KPD operates in conjunction with the KnowPhish knowledge base. When a webpage's screenshot analysis fails to identify a logo (as is the case for logo-less phishing), KPD switches to a text-based analysis mode.
The core of this mode is the Text Brand Extractor, which leverages a Large Language Model (LLM). The process involves:
- HTML Input to LLM: The full HTML content of the suspicious webpage, along with a carefully crafted prompt, is fed into the LLM. The prompt is designed to guide the LLM to identify the likely brand intention of the page based on its textual content.
- Predicted Brand Output: The LLM processes the HTML and outputs a predicted brand name (e.g., "Australia Post").
- Exact Matching with Knowledge Base: This predicted brand is then subjected to an exact matching process against the alias variants stored in the KnowPhish knowledge base. For example, if the LLM predicts "Australia Post" and this exactly matches an alias for the "Australia Post" brand in the knowledge base, the brand intention is successfully identified.
Beyond brand intention, the text-based component of KPD also enhances the detection of Credential Requiring Pages (CRPs). Previous computer vision-based approaches could only detect explicit credential forms. However, by analyzing HTML elements through the LLM, KPD can detect implicit credential requiring pages that might use non-standard forms or subtle prompts, thereby alleviating false negatives caused by less sophisticated CRP classifiers. This multimodal capability allows KPD to effectively detect phishing web pages regardless of whether they display logos or rely solely on textual branding.
Demo / Proof of Concept
▶ Watch: Using LLMs for textual brand intention in logo-less pages (7:00)
While the talk did not feature a live, interactive demonstration, the "empirical evaluation" section effectively served as a rigorous proof of concept for the KnowPhish system. The researchers conducted extensive experiments to validate the performance of their phishing detector, including both a close-word study and a field study under local contexts.
For the close-word study, a balanced dataset was collected by the researchers. The key observations from this study were compelling:
- Direct Enhancement of Detectors: The KnowPhish knowledge base was shown to directly enhance the performance of various existing reference-based phishing detectors, leading to a significant increase in recall. This means KnowPhish enabled these detectors to identify a greater number of actual phishing pages.
- Negligible Impact on Precision: Crucially, this improvement in recall came with only a "negligible or a little impact on precision," indicating that the system did not significantly increase false positives.
- Lower Runtime: KnowPhish also achieved "much lower runtime" compared to a recent baseline, DinaFish, a significant advantage for real-time detection systems.
- Increased Coverage: The fundamental reason for the higher recall was attributed to KnowPhish's ability to cover "more fishing targets" than other baselines. Specifically, it was highlighted that DinaFish often suffered from a "logo matching failure problem," which KnowPhish mitigated by providing a more diverse brand knowledge base and by introducing text analysis when logo analysis failed. This multimodal approach gave KnowPhish a "higher chance to detect a brand intention" than DinaFish.
The field study provided real-world validation, utilizing a dataset from a Singapore governmental agency. This study further corroborated the empirical insights:
- Greatest Number of Detections: The KPD+KnowPhish combination detected the "greatest number of fishing web pages" among the tested approaches.
- Identification of Local Targets: It successfully identified many local phishing targets in Singapore, such as "Singapore Post, DBS Bank, and Shopee." This observation strongly validated the initial insight that "high value industries usually indicates fishing targets," as these local entities represent significant financial or service sectors.
- Prevalence of Logo-less Phishing: An "interesting finding" from the field study was the high prevalence of "logo-less fishing web pages with textual branding intention" in real-world scenarios. This confirmed the critical need for a multimodal solution, as "image-based approaches cannot detect them," whereas KPD+KnowPhish was "able to detect logo-less fishing" effectively.
These empirical results collectively demonstrate the robust capabilities of KnowPhish and its detector, proving its concept through rigorous testing against both controlled and real-world datasets.
Defensive Implications
▶ Watch: Beginning of KnowPhish's empirical evaluation (8:40)
The advancements introduced by KnowPhish offer significant and actionable defensive implications for cybersecurity professionals, organizations, and even end-users. By understanding and potentially adopting the principles and components of KnowPhish, defenders can fortify their phishing detection strategies against increasingly sophisticated attacks.
Firstly, the most direct implication is the enhancement of existing reference-based phishing detectors. Organizations can integrate the concept of a large-scale, multimodal brand knowledge base like KnowPhish into their security infrastructure. Instead of relying on limited, manually curated lists of a few hundred brands, security teams should prioritize the development or acquisition of knowledge bases that span thousands of potential targets, complete with legitimate domains, various logo styles, and crucial alias variants. This vastly expands the scope of detectable phishing campaigns, reducing the likelihood of falling victim to attacks targeting lesser-known or regional brands.
Secondly, the success of KnowPhish's multimodal approach underscores the necessity of moving beyond logo-centric detection. Defenders must recognize the prevalence of logo-less phishing and implement detection mechanisms that can analyze textual content, such as HTML elements, to infer brand intention. This means incorporating Large Language Models (LLMs) or similar natural language processing (NLP) capabilities into their phishing analysis pipelines. Such systems can parse web page content, identify brand names mentioned in text, and cross-reference them with legitimate brand aliases, thereby catching sophisticated phish that evade visual detection.
Furthermore, the finding that high-value industries consistently indicate phishing targets provides a strategic advantage. Security teams can proactively monitor brands within these identified high-risk sectors, ensuring their brand knowledge is meticulously maintained and updated for these specific entities. This insight can also guide threat intelligence efforts, helping organizations anticipate which brands might be targeted next and prepare appropriate defenses.
Finally, the work highlights the importance of continuous knowledge base augmentation and maintenance. While KnowPhish aims to minimize long-term maintenance costs, the dynamic nature of phishing threats necessitates a mechanism for regularly updating brand information, legitimate domains, and alias variants. Organizations should establish processes for collecting new brand data, particularly from high-risk industries, and integrating it into their detection systems to ensure ongoing effectiveness. By embracing these principles, defenders can build more resilient and comprehensive phishing detection frameworks capable of adapting to the evolving threat landscape.
Key Takeaways
- KnowPhish introduces a large-scale, multimodal brand knowledge base covering over 20,000 potential phishing targets, significantly expanding detection coverage compared to previous methods (e.g., 277 brands).
- The system leverages multimodal brand knowledge (logos, legitimate domains, alias variants) to instantly enhance the recall of any existing reference-based phishing detector.
- KnowPhish identifies that 10 high-value industries are consistent phishing targets, providing a time-invariant indicator for proactive target identification through WikiData categories and domain popularity.
- The KnowPhish Detector (KPD) employs Large Language Models (LLMs) to analyze HTML content, enabling the detection of logo-less phishing web pages by identifying textual brand intentions.
- Empirical evaluations demonstrated that KPD+KnowPhish achieves higher recall and lower runtime than baselines like DinaFish, successfully detecting numerous local and global phishing targets in real-world field studies.
- KnowPhish also enhances the detection of implicit credential-requiring pages by using LLMs to analyze HTML elements, alleviating false negatives from traditional computer vision-based approaches.
About the Speaker(s)
The talk was presented by Yuexin Li from the National University of Singapore. The collaborative effort behind KnowPhish also included contributions from Mei Lin Lock, Tri Cao, Nay Oo, Hoon Wei Lim, and Bryan Hooi. As researchers affiliated with the National University of Singapore, their work demonstrates a strong academic background in cybersecurity research, focusing on innovative solutions to persistent threats like phishing, particularly through the application of advanced techniques such as large language models and knowledge graphs.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This research presents KnowPhish, a robust system integrating a large-scale, multimodal brand knowledge base with LLMs to significantly enhance reference-based phishing detection. It addresses critical limitations by expanding brand coverage to over 20,000 targets and effectively detecting logo-less phishing pages, offering crucial advancements in combating sophisticated attacks. This is a practical, impactful step forward in a critical defense area.
Heather Calloway (CISO) — STRONG ACCEPT
KnowPhish presents a significant advancement in phishing detection by expanding brand coverage and tackling logo-less attacks through multimodal knowledge bases and LLMs. This work offers credible and actionable insights for security leaders, directly improving a critical control against a persistent business risk.