Tweezers: A Framework for Security Event Detection via Event Attribution-centric Tweet Embedding

Jian Cui

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Github + OSN Security

Overview

This talk introduces Tweezers, a novel framework designed to enhance the detection of security events from social media platforms, specifically Twitter. Presented by Jian Cui, a PhD student at Indiana University, in collaboration with Professor Shojun Leo and colleagues from Kais and S2W, Tweezers addresses critical shortcomings in existing methods for extracting actionable threat intelligence from the vast and often noisy landscape of social media. The core innovation lies in its event attribution-centric tweet embedding approach, which moves beyond traditional text-based embeddings to leverage the unique security attributes that define and distinguish cyber events.

Watch on YouTube · Slides

Key moments

  1. 2:00 Introduction to Tweezers and social media threat intelligence
  2. 5:00 Identifying the 'false similarity' issue in text embeddings
  3. 6:00 Leveraging security attributes (STIX) to distinguish events
  4. 7:00 Introducing the Tweet Relation Graph for event representation
  5. 7:45 Using Graph Neural Networks and contrastive loss
  6. 8:15 Complete workflow of the Tweezers embedding framework
  7. 8:50 Experimental setup, datasets, and evaluation metrics
  8. 9:30 Tweezers significantly outperforms all existing baseline methods

Tweezers: A Framework for Security Event Detection via Event Attribution-centric Tweet Embedding

Speakers: Jian Cui, PhD Student, Indiana University

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=bLBiCjw_IcM

Overview

This talk introduces Tweezers, a novel framework designed to enhance the detection of security events from social media platforms, specifically Twitter. Presented by Jian Cui, a PhD student at Indiana University, in collaboration with Professor Shojun Leo and colleagues from Kais and S2W, Tweezers addresses critical shortcomings in existing methods for extracting actionable threat intelligence from the vast and often noisy landscape of social media. The core innovation lies in its event attribution-centric tweet embedding approach, which moves beyond traditional text-based embeddings to leverage the unique security attributes that define and distinguish cyber events.

The significance of Tweezers stems from the widely acknowledged role of social media as an invaluable, often early, source of threat intelligence. However, the sheer volume of tweets and the prevalence of irrelevant content, coupled with the inability of current text embedding models to accurately differentiate between lexically similar but semantically distinct security events, have hampered effective intelligence gathering. By focusing on security attributes and employing a graph-based neural network architecture, Tweezers promises more precise and comprehensive detection of critical security incidents, offering security professionals a more reliable mechanism for proactive defense and situational awareness.

Background

▶ Watch: Introduction to Tweezers and social media threat intelligence (2:00)

Social media platforms have increasingly become a critical source of threat intelligence (TI), a sentiment echoed by leading security firms like Trend Micro and Recorded Future. These platforms serve as real-time conduits where security professionals and practitioners share insights, warnings, and details about emerging threats. The information gleaned from such sources, particularly specific security events, can be highly actionable. For instance, details about an OpenSSL vulnerability might include CVE numbers, enabling organizations to identify and patch affected systems. Similarly, proxy-related events often come with IP addresses and hashes crucial for detection, while phishing campaign events provide phishing URLs that can be used to prevent user compromise.

Despite this potential, extracting meaningful security events from Twitter presents two significant challenges. First, the overwhelming volume of tweets daily includes a vast amount of noise and irrelevant content, making the signal-to-noise ratio exceptionally low. Second, existing methods struggle with complete coverage, often missing a substantial portion of critical security incidents, with reported coverage rates as low as 3% to 30%.

Traditional automated approaches typically involve a two-step process: first, text embedding (using models like Word2Vec, GloVe, or more recently, transformer-based models like BERT and LLaMA) to convert tweets into numerical vectors, followed by clustering algorithms (such as k-means or DBScan) to group similar tweets into events. However, this methodology frequently leads to false similarity issues. As Cui illustrates, tweets about a "WhatsApp zero-day" and a "Microsoft Exchange zero-day" might share similar lexical and syntactic patterns (e.g., "zero-day," "exploit," "vulnerability"), causing text embedding models to place them close together in the embedding space, despite representing entirely distinct security events. This misidentification is a major hurdle for accurate event detection.

The core insight behind Tweezers is that while lexical similarity can be misleading, different security events are fundamentally distinguished by their unique security attributes. For example, one event might be characterized by a specific exploited vulnerability, another by particular malware, and a third by a known threat actor. These attributes are not arbitrary; they are formally defined within STIX (Structured Threat Information eXpression) objects. STIX, a de facto standard for describing and sharing cyber threat intelligence, defines 18 different object types that represent various attributes relevant to cyber threats. Tweezers leverages these well-defined security attributes as the key to overcoming the false similarity problem and achieving more precise event detection.

Key Findings

▶ Watch: Leveraging security attributes (STIX) to distinguish events (6:00)

The research underpinning Tweezers revealed several critical findings that challenge conventional approaches to security event detection on social media and propose an effective alternative.

Firstly, a central finding is the pervasive issue of false similarity in existing text embedding models. Even advanced, domain-specific models like SecureBERT (trained on security corpora) and general-purpose large language models (LLMs) often fail to accurately differentiate between security events that share superficial lexical or syntactic similarities but are fundamentally distinct. This inherent limitation leads to a high rate of misidentification and significantly incomplete coverage of actual security incidents.

Secondly, the talk firmly establishes that security attributes, specifically those aligned with STIX objects, are the crucial distinguishing factors for cyber events. Unlike generic named entities (e.g., people, organizations), these specialized attributes—such as exploited vulnerabilities, observed data, malware types, or threat actors—provide the granular, event-specific context necessary to resolve ambiguities in tweet content.

Thirdly, Tweezers demonstrates that integrating these security attributes into a graph representation significantly enhances the semantic understanding of tweets. By constructing a tweet relation graph where tweets are nodes and connections are formed based on shared security attributes, the framework captures intricate relationships that text embeddings alone cannot. This graph structure inherently encodes how attributes are shared across different tweets, making events distinguishable by their unique attribute connectivity patterns.

Finally, the most significant finding is the superior performance of the Tweezers framework. By combining attribute extraction, graph construction, and Graph Neural Networks (GNNs) with contrastive loss functions, Tweezers achieves substantial improvements over all evaluated baselines. It was shown to double the event detection coverage and precision compared to existing methods, including keyword-based approaches, pre-trained language models (BERT, SecureBERT), and even basic graph methods that use generic named entities rather than specific security attributes. This validates the efficacy of an event attribution-centric approach for robust and comprehensive security event detection.

Technical Deep Dive

▶ Watch: Using Graph Neural Networks and contrastive loss (7:45)

The technical core of Tweezers lies in its innovative approach to tweet embedding, specifically designed to overcome the "false similarity" issue prevalent in traditional text-based methods. The framework leverages security attributes to construct a tweet relation graph and then employs Graph Neural Networks (GNNs) to generate robust, event-centric embeddings.

The problem, as highlighted, is that tweets related to different security events (e.g., a WhatsApp zero-day vs. a Microsoft Exchange zero-day) can have highly similar lexical patterns. Traditional text embeddings, even from sophisticated models, struggle to differentiate these, leading to misidentification and poor event clustering. Tweezers addresses this by recognizing that the key to distinguishing these events lies in their specific security attributes, which are precisely defined by STIX objects. STIX (Structured Threat Information eXpression) provides 18 distinct object types that categorize various elements of cyber threat intelligence, such as vulnerabilities, indicators of compromise, malware, and threat actors.

The tweet embedding workflow in Tweezers proceeds as follows:

  1. Initial Feature Extraction: For each incoming tweet, its text and temporal information are initially encoded into a vector space. This provides a baseline textual and temporal representation.
  2. Security Attribute Extraction: This is a crucial step where Large Language Models (LLMs) are employed to extract specific security attributes from the tweet text. Unlike generic named entity recognition, this process focuses on identifying entities relevant to cyber security, such as CVE numbers, specific malware names, exploited vulnerabilities, or associated threat actors, which align with STIX object definitions.
  3. Tweet Relation Graph Construction: Based on the extracted security attributes, a tweet relation graph is constructed. In this graph, each tweet is represented as a node. An edge is drawn between two tweet nodes if they share one or more common security attributes. This graph effectively encodes how security attributes are shared across tweets, providing a structural representation of their interrelationships. Tweets belonging to the same event are expected to share more attributes and thus be more densely connected.
  4. Graph Neural Network Encoding: The initial tweet features (from step 1) and the newly constructed tweet relation graph (from step 3) are fed as input to a Graph Neural Network (GNN). Specifically, Tweezers utilizes Graph Attention Network version 2 (GATv2). GATv2 is chosen for its ability to learn complex relationships and assign varying importance (attention) to different neighboring nodes (i.e., other tweets sharing attributes), thereby generating context-rich embeddings that capture the graph's structural and attribute-sharing information.
  5. Optimization with Contrastive Loss: To ensure the embeddings are effective for event detection, the GNN is optimized using contrastive loss functions. These include triplet loss and pairwise loss. The objective of these loss functions is to push embeddings of tweets belonging to the same security event closer together in the embedding space, while simultaneously pulling embeddings of tweets from different security events further apart. This directly addresses the false similarity problem by enforcing a clear separation based on attributed event identity.
  6. Final Tweet Embeddings: The output of this process is a set of event attribution-centric tweet embeddings, where each embedding vector robustly represents a tweet's security event context, making it suitable for accurate clustering.

The Tweezers end-to-end framework integrates this embedding methodology into a complete event detection pipeline:

  1. Tweet Filtering: A stream of raw tweet data first undergoes a categorization process. Tweets are classified into seven distinct categories (the paper provides more details on these categories). This step also enables multi-category classification for greater accuracy. Tweets not related to security categories are filtered out, significantly reducing noise and focusing subsequent processing on relevant content.
  2. Tweet Embedding: The filtered security-related tweets are then passed through the proposed tweet embedding method described above, generating their event attribution-centric embeddings.
  3. Event Identification: Finally, the generated tweet embeddings are fed into a clustering algorithm, DBScan, to identify distinct event instances. DBScan is particularly suitable here as it can discover clusters of varying shapes and sizes and can identify noise points, which is beneficial for event detection in social media data.

The effectiveness of Tweezers was rigorously tested on a dataset comprising 167 security events spanning 254 tweets. The dataset was split into training, validation, and testing sets, with the testing set specifically designed to cover distinct time periods to evaluate robustness against time shift. Performance was measured using standard clustering metrics: Normalized Mutual Information (NMI), Adjusted Mutual Information (AMI), and Adjusted Rand Index (ARI). Tweezers consistently outperformed all baselines, including keyword-based methods, pre-trained language models (BERT, SecureBERT), and even other graph-based methods that used generic named entities instead of specific security attributes, demonstrating the critical advantage of its attribution-centric approach.

Demo / Proof of Concept

▶ Watch: Complete workflow of the Tweezers embedding framework (8:15)

While the talk did not feature a live, interactive demonstration in the traditional sense, the speaker presented two compelling real-world use cases that illustrate the practical utility and capabilities of the Tweezers framework. These applications serve as concrete proofs of concept for the framework's effectiveness in generating actionable insights from social media data.

The first use case involved analyzing security trends. By processing a continuous stream of security-related tweets and identifying events using the Tweezers framework, researchers could observe and track the emergence and evolution of various cyber threats and vulnerabilities over time. This capability is invaluable for threat intelligence analysts who need to understand the current threat landscape, anticipate future attacks, and allocate defensive resources effectively. The framework's ability to precisely categorize and group events ensures that trend analysis is based on accurate and well-defined security incidents rather than noisy or misidentified data.

The second use case focused on finding informative security users. The framework was utilized to identify individuals or accounts on social media platforms that consistently provide insightful and detailed analyses of security events. By correlating detected events with the users who tweet about them, Tweezers can highlight experts and thought leaders who offer deep dives, early warnings, or unique perspectives. This capability is crucial for building curated lists of trusted sources, enhancing information gathering, and fostering community engagement within the cybersecurity domain. These two use cases collectively demonstrate that Tweezers is not merely a theoretical improvement but a practical tool capable of delivering tangible benefits to the cybersecurity community.

Defensive Implications

▶ Watch: Tweezers significantly outperforms all existing baseline methods (9:30)

The Tweezers framework offers significant defensive implications for organizations and security professionals, fundamentally enhancing their ability to leverage social media for proactive threat intelligence.

Firstly, by providing a more precise and comprehensive detection of security events, Tweezers acts as an improved early warning system. Security teams can gain timely insights into emerging vulnerabilities (e.g., zero-days), active exploit campaigns, new malware strains, or ongoing phishing efforts. The ability to extract actionable indicators of compromise (IOCs) such as CVE numbers, malicious IP addresses, hashes, and phishing URLs directly from social media feeds empowers defenders to implement preventative measures swiftly, such as patching systems, updating detection rules, or blocking malicious infrastructure.

Secondly, the framework's ability to overcome the "false similarity" issue ensures that the generated threat intelligence is of higher quality and reliability. This means security teams can base their decisions on more accurate event categorizations, reducing the risk of misallocating resources to non-existent or misinterpreted threats. The doubled event detection precision and coverage imply a much clearer and more complete picture of the threat landscape, allowing for better-informed risk assessments and strategic planning.

Furthermore, the two demonstrated use cases highlight practical applications. The security trend analysis capability enables organizations to anticipate future threats and adapt their defenses proactively. By understanding which types of attacks are gaining traction or which vulnerabilities are being actively discussed, defenders can prioritize their security investments and focus on hardening the most relevant areas of their infrastructure. Identifying informative security users can also help security teams curate trusted sources of intelligence, allowing them to follow expert analyses and gain deeper context on detected events, thereby enhancing their overall situational awareness.

Finally, the Q&A session revealed that the approach is not limited to Twitter. The framework has been successfully tested on other alternative platforms like Mastodon, demonstrating its adaptability and potential for broad application across the evolving social media landscape. This flexibility is crucial as threat actors and intelligence gatherers migrate to different platforms, ensuring that the defensive capabilities remain robust regardless of the specific social media ecosystem.

Key Takeaways

  • False similarity is a major challenge: Traditional text embedding models struggle to differentiate distinct security events that share similar lexical patterns, leading to misidentification and incomplete threat intelligence.
  • Security attributes are critical differentiators: Leveraging well-defined security attributes, derived from STIX objects, is essential for accurately distinguishing between different cyber events.
  • Graph-based embeddings enhance understanding: The Tweezers framework employs a tweet relation graph and Graph Neural Networks (GNNs), specifically GATv2, to encode relationships between tweets based on shared security attributes, providing a deeper, event-centric understanding.
  • Superior performance in event detection: Tweezers significantly outperforms existing baseline methods (keyword-based, pre-trained LMs like BERT, generic graph methods), doubling both the precision and coverage of security event detection.
  • End-to-end framework for practical application: The complete Tweezers pipeline, including tweet filtering (using 7 categories and multi-category classification), attribute-centric embedding, and DBScan clustering, offers a robust solution for real-time security event identification.
  • Versatile and adaptable: The methodology is applicable to other social media platforms like Mastodon and can potentially be extended to other event detection domains by identifying domain-specific attributes.

About the Speaker(s)

Jian Cui is a PhD student at Indiana University. The work on Tweezers was conducted in collaboration with his advisor, Professor Shojun Leo, and a team of colleagues from Kais and S2W. His research focuses on leveraging advanced computational techniques, including natural language processing and graph neural networks, to extract actionable threat intelligence from complex data sources like social media.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Tweezers is competent academic security research with a real problem statement — text embedding models conflating lexically similar but semantically distinct security events is a genuine pain point in automated TI collection. The GATv2 + contrastive loss approach on a STIX-attribute-derived graph is technically coherent, but the evaluation dataset (167 events, 254 tweets) is so small it's hard to take the 'doubled precision and coverage' headline seriously at scale. Solid NDSS paper; reasonable conference slot; won't be what people are talking about at the bar.

Heather Calloway (CISO) — WEAK

Tweezers is a technically credible research contribution — the GNN-based, attribute-centric embedding approach is a meaningful improvement over naive text clustering for social media threat intelligence. But this talk lands squarely in the research lab and never crosses into institutional relevance. A security leader leaves with no actionable guidance, no deployment path, and no honest accounting of whether social media TI at this scale is worth the organizational investment.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025