Detecting and Mitigating Sampling Bias in Cybersecurity with Unlabeled Data
Saravanan Thirumuruganathan, Fatih Deniz, Mohamed Nabeel, Mourad Ouzzani
33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24
Overview
The deployment of machine learning (ML) models in cybersecurity faces a critical, yet often overlooked, challenge: sampling bias. This talk, presented by Fatih Deniz at USENIX Security '24, delves into the pervasive issue where the data used to train an ML model does not accurately represent the real-world data distribution the model encounters in production. The consequences of such bias are severe, leading to models that perform excellently in academic benchmarks but fail "miserably" when deployed against live, adversarial traffic. This paper, a collaborative effort with Qatar Computing Research Institute and Palo Alto Networks, offers novel solutions for both detecting and mitigating sampling bias, specifically tailored for the unique demands of the cybersecurity domain where labeling data is costly and the environment is inherently adversarial.

Key moments
- 0:00 Introduction and problem definition of sampling bias
- 2:00 Understanding coverage, sampling, and non-response errors
- 4:00 Key insight: Sampling bias vs. concept drift; detection/mitigation
- 5:30 Bias detection using domain discrimination algorithm
- 6:40 More effective K-Nearest Neighbor based bias detector
- 8:00 Introduction to novel self-training mitigation methods
Detecting and Mitigating Sampling Bias in Cybersecurity with Unlabeled Data
Speakers: Saravanan Thirumuruganathan, Fatih Deniz, Mohamed Nabeel, Mourad Ouzzani
Conference: USENIX Security '24
YouTube: https://www.youtube.com/watch?v=nLsrqwAr60
Overview
The deployment of machine learning (ML) models in cybersecurity faces a critical, yet often overlooked, challenge: sampling bias. This talk, presented by Fatih Deniz at USENIX Security '24, delves into the pervasive issue where the data used to train an ML model does not accurately represent the real-world data distribution the model encounters in production. The consequences of such bias are severe, leading to models that perform excellently in academic benchmarks but fail "miserably" when deployed against live, adversarial traffic. This paper, a collaborative effort with Qatar Computing Research Institute and Palo Alto Networks, offers novel solutions for both detecting and mitigating sampling bias, specifically tailored for the unique demands of the cybersecurity domain where labeling data is costly and the environment is inherently adversarial.
Sampling bias is a fundamental problem that undermines the reliability and effectiveness of ML-driven security systems, from intrusion detection to malware analysis and malicious URL detection. The speakers highlight its prevalence, citing a prior USENIX 2022 paper that identified sampling bias as the most common issue in 90% of analyzed top-tier cybersecurity conference papers using ML. Recognizing this critical gap, their research focuses on addressing representation error, a key component of total survey error, encompassing coverage error, sampling error, and non-response error. Their proposed methodologies are designed to be classifier-agnostic and leverage unlabeled data, making them practical and highly applicable in real-world security operations.
The significance of this work extends beyond academic curiosity; it directly impacts the operational efficacy of cybersecurity defenses. By providing robust mechanisms to identify and correct sampling bias before models are deployed, the research helps bridge the notorious "lab-to-production" performance gap. This allows security teams to build and deploy ML models with greater confidence, ensuring they remain effective against the constantly evolving tactics of adversaries who actively try to evade detection. The talk presents a systematic approach that equips practitioners with the tools to proactively address a problem that, if ignored, can render sophisticated ML defenses largely ineffective.
Background
▶ Watch: Introduction and problem definition of sampling bias (0:00)
Sampling bias fundamentally occurs when the collected data for training a machine learning model does not sufficiently represent the true data distribution of the underlying problem. In the context of cybersecurity, this means the distribution of data encountered during training is significantly different from the distribution of data observed in a live, production environment. The consequence is a model that appears highly accurate in controlled settings but performs suboptimally, or even disastrously, when faced with real-world threats. This issue is particularly acute in cybersecurity due to several unique factors.
Firstly, cybersecurity operates within an adversarial environment. Attackers actively attempt to hide their malicious activities and evade detection. This dynamic nature means that the "true" data distribution is constantly shifting, and any static dataset collected for training can quickly become outdated or unrepresentative. Secondly, feature collection in cybersecurity is inherently difficult. Relevant features might be obscured, encrypted, or require deep packet inspection, making comprehensive data acquisition challenging. Thirdly, and perhaps most critically, labeling data in cybersecurity is prohibitively costly and time-consuming. It often requires expert analysis, reverse engineering, or the use of multiple, sometimes conflicting, heuristics (like VirusTotal detections). This reliance on heuristics significantly increases the likelihood of introducing sampling bias.
The speakers highlight the pervasive nature of this problem, referencing a USENIX 2022 finding that 90% of analyzed top-tier conference papers in cybersecurity employing ML suffered from sampling bias. This underscores that despite the advanced ML techniques being developed, the foundational issue of data representativeness remains a critical hurdle for practical deployment.
Within the framework of total survey error, this work primarily focuses on representation error, which arises when the observed sample does not accurately reflect the target population. Representation error is broken down into three key sources:
- Coverage Error: Occurs when the sampling frame (the set of data items from which a sample is drawn) does not match the target population (the complete set of items of interest). For example, if the target population is "all URLs," but the sampling frame only includes "public URLs," ignoring private ones, coverage error is present.
- Sampling Error: Arises when the subset of data items chosen from the sampling frame is not representative of that frame. If, for instance, data is collected only from regional DNS servers, the sample might not be representative of all public URLs within the sampling frame.
- Non-response Error: Happens when a non-uniform subset of the observed sample is excluded. An example given is using a labeling heuristic where a URL is deemed benign if not detected by any of N VirusTotal classifiers, and malicious if detected by K+ classifiers. URLs falling "in between" are ignored, leading to a non-uniform exclusion and thus non-response error.
The authors emphasize that unlike concept drift or distribution drift, which occur after a classifier is deployed and the underlying data distribution changes over time, sampling bias exists before deployment. This crucial distinction means that addressing sampling bias requires different strategies focused on pre-deployment validation and mitigation rather than post-deployment adaptation. Existing work often studies specific cases of sampling bias; however, this research aims to systematically attack the general problem across all three types of representation errors without relying on labeled production data.
Key Findings
▶ Watch: Key insight: Sampling bias vs. concept drift; detection/mitigation (4:00)
The research presents a comprehensive framework for addressing sampling bias in cybersecurity, yielding several significant findings that empower practitioners to build more robust and reliable ML-driven security systems. The core contributions lie in the development of novel, practical, and effective methods for both detecting and mitigating sampling bias using only unlabeled data, a critical constraint in real-world cybersecurity scenarios.
Firstly, the study introduces two distinct and effective algorithms for detecting sampling bias without requiring labels from the production environment. These include a domain discrimination-based detector and a more effective K-Nearest Neighbor (KNN) based bias detector. The key finding here is that these detectors are highly accurate, successfully identifying instances where a classifier trained on available data would perform suboptimally in production. They demonstrated that the KNN-based detector, while potentially slower, offered superior performance in identifying bias.
Secondly, the research proposes two novel mitigation algorithms designed to improve classifier performance in the presence of sampling bias. Both are based on the self-training paradigm but cleverly adapt it to overcome the limitations of traditional approaches in cybersecurity. These methods leverage geometric aspects of ML classifiers over the embedding space, specifically avoiding the pitfalls of pseudo-labeling in feature space which can poison models. The two mitigation strategies are a contrastive learning-based approach and an iterative approach using interrelated classifiers. A pivotal finding is that these mitigation strategies are remarkably effective, capable of reclaiming 10 to 16 points in deployment F-score in adversarial settings.
Overall, the experiments conducted across diverse benchmark datasets (Android malware, Microsoft PE, intrusion detection systems, and M URLs) consistently demonstrated the efficacy of the proposed solutions. The authors reported that their mitigation approaches could reclaim over 90% of the adverse effects of sampling bias, significantly bridging the performance gap between lab results and real-world deployment. Crucially, both the detection and mitigation methods are classifier-agnostic and primarily rely on unlabeled data, making them highly adaptable and practical for a wide range of cybersecurity applications. This work provides a systematic and actionable framework for security practitioners to proactively address a problem that has historically plagued the deployment of ML in security.
Technical Deep Dive
▶ Watch: Bias detection using domain discrimination algorithm (5:30)
The core problem addressed by this research is twofold:
- Detection: Given a labeled training dataset $D_T$ and a classifier trained on it, determine if this classifier is biased and will perform suboptimally in production (deployment data $D_D$) without access to labels from $D_D$.
- Mitigation: If bias is detected, train a new classifier that achieves higher performance on the production data $D_D$ than the original classifier, again without relying on $D_D$ labels.
Crucially, both detection and mitigation strategies are designed to be classifier-agnostic and primarily use unlabeled data, reflecting the realities of cybersecurity environments.
Detection Methods
The paper proposes two algorithms to detect sampling bias, focusing on identifying "malignant" bias where the classifier's production performance will be suboptimal.
- Domain Discrimination-Based Detector:
The intuition here is straightforward: if the training and production data distributions are indistinguishable, then there is no significant sampling bias. Conversely, if they can be easily differentiated, bias is likely present.
- Process: A new binary classifier (the domain discriminator) is trained to distinguish between data points originating from the training set (labeled '1') and those from the deployment set (labeled '0').
- Decision:
- If the accuracy of this domain discriminator is low (e.g., around 50%, indicating random guessing), it implies that the training and deployment data are drawn from similar distributions, and thus, there is no significant sampling bias.
- If the accuracy is high, it indicates that the distributions are substantially different, suggesting the presence of sampling bias that will likely lead to a biased classifier.
- Simplicity and Effectiveness: This method is simple to implement and, as the authors note, "works in practice."
- K-Nearest Neighbor (KNN) Based Bias Detector:
This approach is presented as a more effective detector. The underlying intuition is that if two data distributions are similar, then the distribution of distances to their K-nearest neighbors in an appropriate embedding space should also be similar.
- Embedding Space Creation: The first step involves training an autoencoder using a supervised contrastive learning approach. This encoder maps input data into an embedding space where:
- Embeddings of data items from the same class are grouped closer together.
- Embeddings of data items from different classes are pulled further apart.
- Embedding Computation: Normalized embeddings are computed for both the labeled training data and the unlabeled deployment data using this trained autoencoder.
- Distance Distribution Comparison: For each data point in the deployment set, its distance to its K-nearest neighbors within the training set's embeddings is calculated. Similarly, distances are calculated within the training set. The distributions of these distances (e.g., their median values) are then compared.
- Decision: If these distance distributions are statistically different (e.g., their medians exceed a certain threshold), it indicates a sampling bias. The authors suggest that analyzing the full distribution (a series) rather than just the median might offer even better detection capabilities.
Mitigation Methods
The proposed mitigation methods aim to train a classifier with higher performance in production when sampling bias is detected. Both are based on the self-training paradigm but incorporate novel adaptations for cybersecurity. The authors highlight a critical challenge: standard self-training methods (common in computer vision with label-preserving transformations like image rotation, or tabular ML with controlled feature corruption) are often unsuitable for cybersecurity. Cybersecurity classifiers are highly sensitive to individual features, and simple transformations are rarely label-preserving or even meaningful. Furthermore, directly using pseudo-labels in the feature space can "poison" the model.
The clever idea introduced here is to leverage geometric aspects of ML classifiers over the embedding space rather than the raw feature space. This allows for more robust pseudo-labeling and model improvement.
- Contrastive Learning on Unlabeled Data:
This algorithm addresses the challenge of identifying positive and negative pairs for contrastive learning without true labels for the deployment data.
- Pseudo-Labeling for Production Data:
- First, embeddings of the labeled training set are computed.
- Centroids for each class in the training data's embedding space are calculated.
- For each unlabeled data point in the production set, a pseudo-label is assigned based on its nearest centroid in the embedding space.
- Contrastive Encoder Training: Once pseudo-labels are assigned, positive and negative sample pairs are identified (points with the same pseudo-label are positive, different pseudo-labels are negative).
- A contrastive encoder is then trained using these pairs. The loss function used is a combination of:
- NCE (Noise-Contrastive Estimation) loss for the contrastive learning aspect on the pseudo-labeled production data.
- Cross-entropy loss for the original labeled training data.
- This combined loss helps the encoder learn a better embedding space while simultaneously improving the classifier's ability to distinguish classes based on the original labels and the inferred pseudo-labels.
- Iterative Approach Using Interrelated Classifiers:
This method addresses the difficulty of estimating the quality (accuracy) of pseudo-labels generated for the unlabeled production data. High-quality pseudo-labels are crucial for effective self-training.
- Generating Pseudo-Labels (Classifier 1): The initial classifier, trained on the labeled $D_T$, is used to generate pseudo-labels for the unlabeled production data $D_D$.
- Training a Second Classifier (Classifier 2): A second classifier is then trained using these pseudo-labeled production data points.
- Indirect Pseudo-Label Accuracy Measurement: This second classifier is then used to make predictions on the original labeled training data $D_T$. By comparing these predictions to the true labels of $D_T$, the accuracy of the pseudo-labels generated by the first classifier for $D_D$ can be indirectly estimated. The rationale is that if the pseudo-labels on $D_D$ are of high quality, a model trained on them should perform well on $D_T$ (assuming some degree of overlap or generalizability).
- Iterative Refinement: While not explicitly detailed as a multi-step iteration in the brief overview, the "iterative approach" suggests a potential for refining pseudo-labels or models over several cycles, using the indirect accuracy measurement to guide the process and prioritize high-confidence pseudo-labels.
Both mitigation strategies are designed to be robust against the inherent challenges of cybersecurity data, particularly the lack of labels and the adversarial nature of the problem space. By focusing on embedding spaces and carefully constructed self-training mechanisms, they aim to produce classifiers that generalize better to real-world deployment scenarios.
Demo / Proof of Concept
▶ Watch: More effective K-Nearest Neighbor based bias detector (6:40)
While the talk did not feature a live, interactive demonstration, the speakers presented extensive experimental validation and results that serve as a robust proof of concept for their proposed detection and mitigation algorithms. The efficacy of their methods was rigorously tested across a diverse set of real-world cybersecurity benchmark datasets, showcasing their broad applicability and significant performance improvements.
The experiments were conducted using:
- Android Malware: Datasets representing malicious Android applications, a common and evolving threat.
- Microsoft PE (Portable Executable): Data related to Windows executable files, crucial for detecting malicious binaries.
- Intrusion Detection Systems (IDS): Datasets for network intrusion detection, a foundational area of cybersecurity.
- M URLs: Datasets consisting of malicious URLs, vital for web security and phishing detection.
To ensure impartiality and avoid bias in evaluation, the researchers adopted a strict methodology: a classifier was trained using one dataset and then evaluated using a different dataset. This approach simulates the real-world scenario where training data might not perfectly match deployment data. They explored numerous scenarios, employing various classifier types and different sampling strategies to induce bias, demonstrating the robustness of their solutions even under adversarial conditions.
The key quantitative results presented were highly compelling:
- Detection Accuracy: Both the domain discrimination and KNN-based detectors successfully identified sampling bias across the tested scenarios. The KNN-based detector, while computationally more intensive, showed superior performance in accurately identifying bias.
- Mitigation Effectiveness: The mitigation strategies demonstrated significant improvements in classifier performance in production. Specifically, the authors reported reclaiming 10 to 16 points in deployment F-score in adversarial settings. The F-score is a critical metric in cybersecurity, balancing precision and recall, making this improvement highly impactful for practical applications where both false positives and false negatives are costly.
- Overall Impact: The research found that their mitigation approaches could effectively counteract over 90% of the adverse effects of sampling bias. This high percentage underscores the practical utility of their methods in making ML models deployed in cybersecurity vastly more reliable.
Furthermore, the experiments confirmed that both the detection and mitigation approaches remained classifier-agnostic, meaning they were effective regardless of the specific machine learning model (e.g., SVM, Random Forest, Neural Networks) used for classification. This flexibility is a major advantage, allowing security teams to integrate these bias-handling techniques with their existing ML pipelines. The reliance on unlabeled data for production environments was also a consistent theme in the experimental setup, validating the practicality of the solutions in scenarios where ground truth labels are scarce or impossible to obtain in real-time. These comprehensive experimental results strongly validate the claims of the paper and provide a solid foundation for their real-world adoption.
Defensive Implications
▶ Watch: Introduction to novel self-training mitigation methods (8:00)
The findings from this research offer profound and actionable implications for cybersecurity defenders, enabling them to significantly enhance the reliability and effectiveness of their machine learning-driven security systems. The ability to detect and mitigate sampling bias directly addresses one of the most critical challenges in deploying ML models in an adversarial environment.
Here are the key defensive implications:
- Proactive Bias Detection in ML Pipelines: Security teams should integrate the proposed sampling bias detection algorithms (especially the more effective KNN-based detector) into their ML model development and deployment pipelines. Before any model goes into production, it should be rigorously checked for bias against representative (even if unlabeled) samples of expected production data. This proactive step can prevent the deployment of models that are destined to fail in the real world, saving significant resources and preventing security blind spots.
- Improved Model Robustness and Trust: By employing the mitigation strategies (contrastive learning or the iterative approach), defenders can train more robust ML models that generalize better to unseen and potentially adversarial data. This directly translates to higher confidence in the predictions made by these models, reducing the incidence of missed attacks (false negatives) and alert fatigue from erroneous detections (false positives), which are both critical for efficient security operations. Reclaiming 10-16 points in F-score is a substantial gain in practical terms.
- Leveraging Unlabeled Data: The emphasis on unlabeled data for both detection and mitigation is a game-changer for resource-constrained security teams. Given the high cost and difficulty of obtaining labeled cybersecurity data, these methods allow organizations to continuously validate and improve their models using readily available production traffic, without needing to invest heavily in manual labeling efforts. This makes the solutions highly scalable and sustainable.
- Classifier-Agnostic Solutions: The fact that these methods are classifier-agnostic means they can be applied across a wide array of existing and future ML models used in security. Whether a team uses deep learning for malware analysis or traditional ML for network intrusion detection, the bias detection and mitigation techniques can be integrated without requiring a complete overhaul of their ML stack.
- Understanding Adversarial Dynamics: This research highlights the inherent challenges of an adversarial environment where attackers constantly evolve. Defenders must recognize that their training data will likely never perfectly represent future attack vectors. Implementing these bias-aware techniques provides a crucial layer of defense against adversarial data shifts that manifest as sampling bias, ensuring that ML models remain effective even as adversaries adapt.
- Bridging the Lab-to-Production Gap: The work directly addresses the common frustration of models performing well in academic papers but failing in production. By systematically detecting and mitigating sampling bias, security practitioners can significantly bridge this gap, ensuring that the theoretical capabilities of ML translate into tangible, real-world security enhancements.
In essence, this research provides security defenders with essential tools to ensure their ML-powered defenses are not just theoretically sound but are also practically resilient against the inherent biases of data collection and the dynamic nature of cyber threats. Integrating these methodologies moves organizations towards a more data-aware and robust approach to cybersecurity.
Key Takeaways
- Sampling bias is a pervasive and critical issue in cybersecurity machine learning, affecting 90% of top-tier conference papers and leading to models that perform poorly in production despite strong benchmark results.
- The research provides novel, classifier-agnostic solutions for both detecting and mitigating sampling bias, specifically designed for the unique challenges of cybersecurity, including an adversarial environment and costly data labeling.
- Two detection algorithms are proposed: a domain discrimination-based detector and a more effective K-Nearest Neighbor (KNN) based bias detector, both operating without requiring labels from the production environment.
- Two mitigation algorithms, based on an adapted self-training paradigm, effectively reduce bias by leveraging geometric aspects of ML classifiers in the embedding space: a contrastive learning approach and an iterative approach using interrelated classifiers.
- The mitigation strategies are highly effective, demonstrating the ability to reclaim 10 to 16 points in deployment F-score and mitigating over 90% of the adverse effects of sampling bias across diverse cybersecurity datasets (Android malware, Microsoft PE, IDS, M URLs).
- The methods primarily utilize unlabeled data, making them highly practical and scalable for real-world cybersecurity operations where obtaining labeled production data is often infeasible.
About the Speaker(s)
The talk was presented by Fatih Deniz, with co-authors Saravanan Thirumuruganathan, Mohamed Nabeel, and Mourad Ouzzani. The research is a joint work between Qatar Computing Research Institute and Palo Alto Networks. Fatih Deniz, as the presenter, shared insights into the problem of sampling bias in cybersecurity and the team's innovative solutions for its detection and mitigation using unlabeled data, reflecting their expertise in machine learning applications within the security domain.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research directly confronts the pervasive issue of sampling bias in cybersecurity ML, a critical factor undermining real-world deployments. It introduces novel, classifier-agnostic detection and mitigation strategies that intelligently leverage unlabeled data. The demonstrated ability to reclaim significant performance in adversarial settings marks this as a foundational and immediately actionable contribution for building robust security systems.
Heather Calloway (CISO) — MUST SEE
This research directly addresses a critical, often overlooked, systemic risk: the pervasive failure of ML models in production due to sampling bias. By offering practical, unlabeled data-driven methods for both detecting and mitigating this bias, it provides a crucial pathway to ensure our security defenses are actually effective in the adversarial real world. Every CISO needs to understand these implications to accurately assess the resilience of their ML-driven controls.