MineShark: Cryptomining Traffic Detection at Scale

Shaoke Xi

Network and Distributed System Security (NDSS) Symposium 2025 · Day 1 · Network Security 1

Overview

This talk introduces MineShark, an innovative online detection system designed to combat cryptojacking attacks by identifying cryptomining traffic at scale. Presented by Shaoke Xi, MineShark addresses critical limitations of existing network intrusion detection systems (NIDS) and machine learning (ML) models in handling high-speed network traffic and reducing an overwhelming number of false alarms. Cryptojacking, the unauthorized use of computational power for cryptocurrency mining, poses a significant threat to organizational infrastructure, leading to increased resource consumption, degraded service performance, and potential financial losses.

Watch on YouTube · Slides

Key moments

  1. 0:00 Cryptojacking attacks and limitations of current detection.
  2. 2:00 Understanding the Stratum protocol and key traffic features.
  3. 3:30 Existing ML models' limitations: throughput and false alarms.
  4. 4:50 Introducing MineShark: an online cryptomining detection system.
  5. 7:00 MineShark's strategies for effectively reducing false alarms.
  6. 9:00 Proactive defense: identifying backup addresses and port fingerprints.
  7. 10:30 Real-world deployment and superior detection performance.

MineShark: Cryptomining Traffic Detection at Scale

Speakers: Shaoke Xi

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=4FQFf_8PJVw

Overview

This talk introduces MineShark, an innovative online detection system designed to combat cryptojacking attacks by identifying cryptomining traffic at scale. Presented by Shaoke Xi, MineShark addresses critical limitations of existing network intrusion detection systems (NIDS) and machine learning (ML) models in handling high-speed network traffic and reducing an overwhelming number of false alarms. Cryptojacking, the unauthorized use of computational power for cryptocurrency mining, poses a significant threat to organizational infrastructure, leading to increased resource consumption, degraded service performance, and potential financial losses.

The core challenge in detecting cryptojacking lies in the stealthy behaviors employed by attackers, such as the widespread use of TLS encryption and rapid domain rotation for mining pool servers. These techniques render traditional deep packet inspection ineffective and allow attackers to easily bypass firewall blocks. MineShark tackles these issues through a multi-pronged approach, combining lightweight deep learning models for efficient online inference, an automated investigation module to significantly reduce false positives, and proactive defense mechanisms to anticipate future attacks. The system's real-world deployment and evaluation demonstrated its superior performance in terms of detection timeliness and accuracy compared to commercial systems and public threat intelligence services.

Background

▶ Watch: Cryptojacking attacks and limitations of current detection. (0:00)

Cryptojacking attacks have become a pervasive threat in the cybersecurity landscape, leveraging hijacked computational resources to mine cryptocurrencies without the victim's consent. The dominant method for these attacks is pool mining, where infected machines (miners) connect to a central mining pool server. This server aggregates the work of many miners and distributes rewards, providing a stable income stream for attackers. The attack lifecycle typically involves malware infection, download of mining software and attacker wallet information from a command-and-control (C2) server, and then continuous network communication with the mining pool server to receive puzzle-solving jobs and submit solutions. This persistent network connection is a primary indicator of compromise.

However, several stealthy behaviors make network-based detection challenging. First, nearly half of all mining pools utilize TLS encryption, effectively obscuring the content of network packets from traditional deep packet inspection (DPI). Second, cryptojacking attackers frequently rotate their infrastructure; almost half of mining pool domains expire within 10 days. This means that blocking a single IP address or domain offers only a temporary reprieve, as attackers can quickly switch to backup addresses and continue operations.

Researchers have increasingly turned to machine learning (ML) to identify more robust network features beyond simple IP/domain blocking. The Stratum protocol, commonly used by most mining pools, exhibits identifiable traffic patterns. A miner initially registers its identity and wallet address. The pool server then sends a puzzle-solving job. The miner returns the solution, receives confirmation, and a new job. This process generates unique traffic features:

  • Unique packet sets carrying distinct semantic messages (e.g., job requests, solution submissions).
  • Consistent message frequency due to the relatively uniform time spent solving each puzzle.
  • Regular inter-packet delays resulting from the consistent frequency.
  • Repetitive patterns of unique packet occurrences over time.
  • Long-lasting connections as mining profits accumulate over extended periods.

While existing ML models show promise on labeled datasets, their suitability for large-scale, real-world deployments is limited by two major challenges. Firstly, network intrusion detection systems are deployed at high-speed gateways, demanding extreme efficiency for online detection. Previous ML models struggle with this inference bottleneck; for instance, a model from prior work could take up to 1.4 hours to evaluate all feature vectors generated by a 10 Gbps gateway in just one second. This low throughput renders accurate ML models ineffective, especially given that mining traffic constitutes a tiny fraction of total network traffic. Secondly, existing work often overlooks the problem of false alarms. Running even the most accurate models on live traffic can generate over 100,000 daily alarms, making manual investigation impractical and overwhelming for security teams. MineShark was developed to directly address these two critical limitations.

Key Findings

▶ Watch: Existing ML models' limitations: throughput and false alarms. (3:30)

MineShark's deployment and evaluation yielded several key findings that demonstrate its effectiveness in real-world scenarios, significantly improving upon existing detection methods:

  1. Online Detection at Scale: MineShark successfully achieved lightweight processing on a 10 Gbps gateway, even under worst-case scenarios where every flow generates feature vectors. This demonstrates its capability to overcome the critical inference bottleneck faced by previous ML models, enabling efficient, real-time cryptomining detection in high-speed network environments.
  1. Superior Detection Timeliness: The system identified 105 cryptomining addresses before commercial security systems and 75 addresses before they were listed on VirusTotal. Crucially, MineShark significantly reduced detection delays, averaging 5 days earlier for plaintext mining and 19 days earlier for encrypted mining compared to commercial systems. This early detection is vital for minimizing the impact and spread of cryptojacking attacks.
  1. Enhanced Detection Accuracy and False Alarm Reduction: MineShark's auto-investigation module, particularly its harmless visit ratio parameter, proved highly effective in distinguishing between benign and suspicious addresses. The system reduced the number of alarm flows by two orders of magnitude, transforming an impractical daily alarm count (e.g., 100,000) into a manageable number for human analysts. This dramatic reduction in false positives makes the system operationally viable.
  1. Identification of Previously Undetected Threats: Through its active probing and advanced analytical features, MineShark successfully identified mining pool addresses that were entirely undetected by VirusTotal. This highlights MineShark's ability to uncover novel or rapidly evolving cryptojacking infrastructure that evades traditional signature-based or passive intelligence feeds.
  1. Effective Proactive Defense: MineShark's proactive defense techniques, which correlate backup addresses by domain and fingerprint common port configurations used by malware families, successfully identified and confirmed mining activities in advance. This capability allows administrators to block unauthorized connections before they can cause significant damage, demonstrating a forward-looking approach to security.

Technical Deep Dive

▶ Watch: Introducing MineShark: an online cryptomining detection system. (4:50)

MineShark is an online detection system that employs deep learning models and a multi-stage pipeline to detect cryptomining traffic at scale. Following a common off-path deployment paradigm, MineShark ingests raw network traffic and applies three core strategies: lightweight ML for high recall, an auto-investigation module for reliable detection, and proactive defense against future attacks.

Lightweight ML for High-Speed Inference

To address the inference bottleneck at high-speed gateways, MineShark prioritizes lightweight ML models designed for high recall. For a 10 Gbps gateway, the system requires a model throughput of approximately 1.5 million packets per second (MPS). After evaluating different architectures, a Convolutional Neural Network (CNN) was selected as the optimal choice. Unlike Support Vector Machines (SVMs), which struggled to detect suspicious flows, or Long Short-Term Memory (LSTM) networks, which could only process about two-thirds of the total traffic, the CNN provided the necessary balance of performance and accuracy.

MineShark's pipeline tracks the status of active network flows using an in-memory flow table. This table is crucial for maintaining context across packets belonging to the same flow. The system leverages parallel computation across both CPU and GPU resources to maximize throughput. Flow features are continuously extracted from ongoing flows. These features form a 4-row vector, where the first two rows represent bidirectional packet sizes and the last two rows represent inter-packet delays. This feature set effectively captures the unique traffic patterns of the Stratum protocol discussed earlier. The extracted feature vectors are then fed into the ML model for prediction. A per-flow alarm count records the number of positive features detected from each flow. Flows with a non-zero alarm count are flagged as suspicious and passed to the next stage of the pipeline.

Auto-Investigation Module for False Alarm Reduction

The initial lightweight ML stage, while efficient, may still generate a significant number of false alarms. MineShark's auto-investigation module is designed to drastically reduce these false positives through a combination of correlation, ranking, and active probing.

  1. Correlation Graph and Harmless Flow Ratio: MineShark constructs a correlation graph to group suspicious flows by their destination address over time. This helps identify long-term visiting patterns that are characteristic of persistent mining activities. A key observation is that many false alarms arise from temporary connection behaviors. To filter these, MineShark treats suspicious flows with only a single alarm count as harmless. For example, a legitimate website might have a single page exhibiting mining-like patterns, but the overall traffic proportion is minimal. In contrast, true cryptomining connections to a malicious address tend to be more frequent and sustained. MineShark employs a harmless flow ratio threshold to distinguish between genuinely malicious mining addresses and benign addresses exhibiting transient, mining-like patterns. This parameter was empirically shown to effectively separate benign and suspicious addresses and reduce the number of alarm flows by two orders of magnitude.
  1. Model-Independent Feature Ranking: For the remaining suspicious addresses, MineShark further ranks them based on three model-independent features that capture additional cryptomining behaviors. The average values of these features for known mining addresses (indicated by red dashed lines in the presentation) clearly distinguish them from normal addresses. This ranking prioritizes the most suspicious addresses for the final confirmation step.
  1. Active Probing for Confirmation: The final step in the auto-investigation module is active probing, starting with the highest-ranked suspicious addresses. MineShark constructs probing packets that simulate the behavior of legitimate mining software requesting mining services. If the probed server responds with a successful acknowledgment (e.g., a Stratum protocol success message), it confirms active mining services. Conversely, if the target returns various error messages or no response, it indicates a non-mining service or an inactive address. This active confirmation significantly enhances the reliability of detection results, minimizing operational overhead by ensuring that only confirmed threats are flagged for administrator action.

Proactive Defense Techniques

MineShark integrates proactive defense mechanisms to anticipate and block future attacks. These techniques leverage historical attack data:

  1. Domain Correlation: Attackers often use correlated backup addresses associated with the same domain to quickly switch infrastructure. MineShark observes this behavior and proactively searches for addresses recently associated with already detected mining domains. It then uses active probing to confirm if these newly identified addresses are indeed mining pools, allowing administrators to block them in advance.
  1. Port Fingerprinting: Malware families frequently configure the same set of ports, with each port offering a specific service. MineShark fingerprints the ports used by historical mining pool servers. If a newly examined address matches an existing port fingerprint, MineShark probes other relevant ports on that address to identify and confirm additional mining services.

Underlying Technologies

MineShark is built using DPDK (Data Plane Development Kit) as its traffic processing framework, enabling high-performance packet capture and processing. For model inference, it utilizes TensorFlow, a popular open-source machine learning framework.

Demo / Proof of Concept

▶ Watch: Proactive defense: identifying backup addresses and port fingerprints. (9:00)

While the talk did not present a live "demo" in the traditional sense, MineShark's effectiveness was rigorously validated through a large-scale, real-world deployment and evaluation. The system was deployed at a 10 Gbps gateway within the presenter's campus network for a continuous period of 10 months.

During this extensive evaluation period, MineShark operated under real-world conditions, handling an average input traffic speed of 7 Gbps and managing an average of 3.2 million concurrent flows. This deployment environment provided a robust testbed for assessing MineShark's scalability, efficiency, and accuracy against live cryptojacking threats.

The results were compelling:

  • MineShark successfully identified 105 cryptomining addresses before they were detected by commercial security systems.
  • It identified 75 addresses before they were listed on VirusTotal, a widely used public threat intelligence platform.
  • The average detection delay for MineShark was significantly lower than commercial systems: 5 days faster for plaintext mining and an impressive 19 days faster for encrypted mining.
  • Further investigation using VirusTotal's API confirmed MineShark's improved accuracy. The system successfully identified numerous mining pool addresses that were entirely undetected by VirusTotal, demonstrating its ability to uncover novel or rapidly evolving cryptojacking infrastructure.
  • The auto-investigation module's effectiveness was also validated, showing that the harmless visit ratio parameter could effectively separate benign and suspicious addresses. This module was instrumental in reducing the number of alarm flows by two orders of magnitude, making the system's output actionable for security administrators.
  • MineShark's ranking algorithm effectively prioritized the most suspicious addresses for investigation, placing known malicious addresses (marked by red crosses) at the top of the probing list, even within a distribution of approximately 2,000 suspicious addresses.
  • The system maintained lightweight processing capabilities on the 10 Gbps gateway, even in the worst-case scenario where every flow generated feature vectors, confirming its robust performance under heavy load.

This comprehensive evaluation underscores MineShark's capability to deliver timely and accurate cryptomining traffic detection in high-speed, real-world network environments, significantly outperforming existing solutions.

Defensive Implications

▶ Watch: Real-world deployment and superior detection performance. (10:30)

MineShark presents a compelling case for a paradigm shift in how organizations defend against cryptojacking attacks. Its capabilities offer several critical implications for security defenders:

  1. Prioritize Online, Scalable Detection: Organizations should invest in and deploy ML-based network intrusion detection systems capable of operating at high network speeds (e.g., 10 Gbps and beyond). Traditional NIDS and even some ML solutions are inadequate for the volume and velocity of modern network traffic. Systems like MineShark, utilizing lightweight ML models like CNNs, are essential for real-time identification of evolving threats.
  1. Embrace Multi-Stage Detection with False Alarm Reduction: The talk highlights that accurate ML models alone can generate an unmanageable number of false alarms. Defenders should look for solutions that integrate secondary validation mechanisms, such as MineShark's auto-investigation module with its harmless flow ratio and active probing. Reducing false positives by orders of magnitude is crucial for making detection systems operationally viable and preventing analyst fatigue.
  1. Leverage Active Probing for Confirmation: Passive observation is often insufficient. Active probing, which simulates legitimate client behavior to confirm suspicious services, provides a high-confidence method for validating potential threats. Defenders should consider incorporating or supporting systems that actively verify suspicious network destinations before taking blocking actions.
  1. Implement Proactive Threat Intelligence: Beyond reactive detection, organizations should develop or adopt capabilities for proactive defense. MineShark's approach of correlating backup addresses by domain and fingerprinting common port configurations allows for the identification and blocking of attacker infrastructure before it becomes actively malicious or widely known. This foresight can significantly reduce an organization's attack surface.
  1. Integrate with Automated Response: The identification of cryptomining addresses, especially through high-confidence active probing, should be seamlessly integrated into automated response mechanisms. This includes dynamically updating firewalls and intrusion prevention systems (IPS) with deny lists to block unauthorized mining connections, following organizational policies.
  1. Understand Cryptomining Traffic Patterns: Security analysts should be educated on the unique network traffic features of common cryptomining protocols like Stratum (e.g., consistent packet sizes, regular inter-packet delays, long-lasting connections). This understanding aids in manual investigation and in fine-tuning detection rules.

By adopting these principles and leveraging the types of capabilities demonstrated by MineShark, defenders can build more robust, efficient, and accurate defenses against the persistent and evolving threat of cryptojacking.

Key Takeaways

  • Cryptojacking remains a significant threat: It causes infrastructure damage, service degradation, and financial loss, often using stealthy techniques like TLS encryption and rapid domain rotation.
  • Traditional NIDS and existing ML models struggle with scale: They face inference bottlenecks at high network speeds (e.g., 10 Gbps) and generate an unmanageable number of false alarms (e.g., 100,000 daily).
  • MineShark provides an effective online solution: It uses lightweight deep learning (CNN) for high-recall inference at 10 Gbps, parallel computation, and a 4-row feature vector (packet size, inter-packet delay) to detect cryptomining traffic.
  • Auto-investigation drastically reduces false alarms: Through a correlation graph, harmless flow ratio, model-independent features, and active probing, MineShark reduces alarm flows by two orders of magnitude and confirms mining activities reliably.
  • MineShark offers superior detection timeliness and accuracy: In a 10-month real-world deployment, it detected threats 5-19 days earlier than commercial systems and identified addresses completely missed by VirusTotal.
  • Proactive defense is crucial: MineShark leverages historical data to predict and block future attacks by correlating backup addresses by domain and fingerprinting common port configurations.

About the Speaker(s)

Shaoke Xi presented the "MineShark: Cryptomining Traffic Detection at Scale" talk at the NDSS Symposium. Shaoke explicitly mentioned presenting on behalf of colleagues who were unable to attend due to visa issues, indicating that this work is a collaborative effort likely from a research team. While specific titles or institutional affiliations were not provided in the metadata or transcript, the nature of the presentation at a prestigious academic security conference like NDSS suggests Shaoke Xi is a researcher or academic actively contributing to the field of network security and machine learning applications for threat detection. Their role as the presenter implies a deep understanding and significant involvement in the development and evaluation of the MineShark system.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

MineShark is competent, well-scoped systems research that solves a real operational problem — ML inference throughput and false alarm volume at line rate. The 10-month campus deployment with measurable lead times over VirusTotal is the strongest card it plays. Not groundbreaking, but honest work done properly.

Heather Calloway (CISO) — WEAK

Technically credible research that solves a real engineering problem — inference bottleneck and false alarm volume at scale — but never bridges to the institutional question of who should care and what they should do about it. The defensive implications section lists recommendations, but they read as researcher speculation, not operator guidance.

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025