Practical Attacks against DNS Reputation Systems
Tillson Galloway, Kleanthis Karakolios, Zane Ma, Roberto Perdisci, Manos Antonakakis, Angelos Keromytis
IEEE Symposium on Security and Privacy 2024 · Day 3 · Continental Ballroom 4
Overview
This talk, presented by Tillson Galloway and collaborators from Georgia Tech, Oregon State University, and the University of Georgia, delves into the critical vulnerabilities of DNS reputation systems. These systems are foundational security components that leverage statistical features, machine learning (ML), and heuristic methods to identify and flag malicious domains, protecting users from phishing, malware, and other online threats. They are widely integrated into email services, firewalls, web browsers, and registrars to prevent malicious activity at an early stage.

Key moments
- 0:00 Introduction to DNS reputation systems and research questions
- 2:00 Adversary can evade models for $10 and 14 days
- 2:20 Building reference model and problem space attack strategy
- 3:40 Key feature groups for DNS reputation classification
- 5:20 Mimicry attack: altering IP relationships to evade detection
- 6:40 Popularity list attack using VPNs to manipulate rankings
Practical Attacks against DNS Reputation Systems
Speakers: Tillson Galloway, Kleanthis Karakolios, Zane Ma, Roberto Perdisci, Manos Antonakakis, Angelos Keromytis
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=ewqJdAjuJvc
Overview
This talk, presented by Tillson Galloway and collaborators from Georgia Tech, Oregon State University, and the University of Georgia, delves into the critical vulnerabilities of DNS reputation systems. These systems are foundational security components that leverage statistical features, machine learning (ML), and heuristic methods to identify and flag malicious domains, protecting users from phishing, malware, and other online threats. They are widely integrated into email services, firewalls, web browsers, and registrars to prevent malicious activity at an early stage.
The core of the research addresses two pressing questions: how effectively can adversaries evade these ML-based systems, and how practical are such evasion tactics against real-world, industry-grade security solutions? The findings reveal a concerning reality: for a minimal investment of approximately $10 and just 14 days of effort, an attacker can successfully circumvent both academic models and commercial security vendors. This work not only highlights the inherent weaknesses in current DNS reputation mechanisms but also provides a practical framework for understanding and mitigating these significant security risks.
Background
▶ Watch: Introduction to DNS reputation systems and research questions (0:00)
DNS reputation systems operate by extracting a diverse set of features from domain names and their associated infrastructure. These features span various categories, including lexical features (e.g., domain entropy, length), Whois data (e.g., registration length, age – generally, older and longer registrations are associated with benignness), network infrastructure (historical and current IP addresses, autonomous systems (ASNs), BGP prefixes, and shared IP addresses with other domains), popularity lists (indicating traffic volume and ranking), and IP blocklists (historical evidence of malicious activity associated with an IP address or network). Data for these features is sourced from recursive and authoritative DNS datasets, active lookups, Certificate Transparency (CT) logs, and external enrichment services like Whois and BGP routing information.
The output of these systems is crucial for various security operations: email services use them to reject emails from low-reputation domains or block phishing lures; firewalls and web browsers protect users from accessing malicious sites; and registrars leverage them to suspend malicious domains proactively. Despite their widespread adoption and sophisticated methodologies, the researchers questioned the robustness of these ML-based systems against determined adversaries. Prior academic work, often presented at top-tier security conferences over the past decade, has focused on building effective detection models. This research builds upon that foundation, using active DNS data provided by Georgia Tech to construct a robust reference model. However, the critical gap identified was the practical applicability and cost-effectiveness of evasion techniques against real-world, black-box systems that incorporate more than just ML, such as pre-processing and filtering stages.
Key Findings
▶ Watch: Building reference model and problem space attack strategy (2:20)
The research yielded several significant findings that challenge the perceived resilience of current DNS reputation systems:
- Evasion is Possible and Practical: Adversaries can effectively evade both academic and industry-grade ML-based DNS reputation systems. The cost and effort required for such evasion are remarkably low, with an estimated $10 and 14 days of effort being sufficient for many scenarios.
- Two Primary Evasion Techniques: The study identified and demonstrated the efficacy of two main attack vectors:
- Mimicry Attacks: Manipulating historical IP relationships by temporarily pointing a malicious domain to a benign IP address significantly alters its feature distribution, increasing its perceived reputation. This attack incurs no direct financial cost or significant time investment.
- Popularity List Attacks: Exploiting the reliance of many reputation systems on popularity lists (often derived from Open DNS recursive resolvers) by generating artificial traffic to a malicious domain. This can be achieved for approximately $10 using VPN services to rotate IP addresses, thereby manipulating the domain's ranking on lists like Tranco, Cisco Umbrella, and Cloudflare Radar.
- Real-World Vendor Vulnerability: Through a partnership with an anonymous security vendor, the researchers demonstrated that combinations of these attacks (e.g., a specific popularity ranking combined with mimicry and a 2-year domain registration) successfully evaded the vendor's sophisticated detection mechanisms, even when associated with malicious binaries.
- Sandbox-Specific Bypass Techniques: Beyond the core reputation system attacks, two effective techniques were found to bypass sandbox analysis:
- Including a domain name in a binary's strings without actively performing a dynamic DNS lookup meant the domain was identified but still evaded the model.
- Utilizing encrypted DNS protocols (e.g., DNS over HTTPS/TLS) prevented detection, as the sandbox lacked memory inspection capabilities to analyze the encrypted traffic.
- Vendor Prioritization of False Positives: The research suggests that real-world security vendors, while employing complex detection systems, are often "incredibly false positive avoidant." This prioritization means that missing a few malicious domains is often deemed less detrimental than incorrectly flagging benign domains, which can lead to service disruptions or domain suspensions for legitimate users. This trade-off creates an exploitable window for attackers.
Technical Deep Dive
▶ Watch: Key feature groups for DNS reputation classification (3:40)
The research began by establishing a robust reference model for DNS reputation, built upon a decade of academic work from top-tier security conferences and augmented with active DNS data from Georgia Tech. This model implemented over 50 statistical features, categorized into nine distinct feature groups, ensuring a comprehensive representation of domain attributes. To foster reproducibility and encourage further research, the implementations of these feature extraction methods were open-sourced.
A crucial distinction in their attack methodology lies in operating within the problem space rather than the feature space. Unlike gradient-based techniques that manipulate features directly (common in other ML settings), the problem space involves making tangible changes to DNS records. This approach is necessitated by the highly contained and non-invertible nature of DNS features: a simple DNS record addition can drastically and unpredictably alter a domain's feature vector, and two distinct domains might map to an identical feature vector. The basic unit of analysis within this system is the resource record, specifically Type A and Quad A (AAAA) records, which map a domain name to an IP address.
The features used for classification were diverse:
- Lexical Features: Extracted directly from the domain name itself, such as its entropy or length, requiring no external data sources.
- Whois Data: Information like registration length and age, where a longer history is typically associated with benign domains.
- Network Infrastructure: Leveraging historical and current DNS resolutions to gather data on the domain's IP addresses, its historical IP addresses, associated Autonomous Systems (ASNs), and BGP prefixes. This also included pivoting on network connections to find features of other domains sharing the same IP address.
- Popularity Lists: Providing insights into the amount of traffic directed to a domain through its ranking on various lists.
- IP Blocklists: Offering historical evidence of malicious activity associated with particular IP addresses or networks.
While the reference model incorporated numerous features, the researchers identified that three specific feature groups—those related to historic records and popularity lists—contained the majority of the information critical for classification. These became the primary targets for their evasion attacks.
Mimicry Attacks
Mimicry attacks exploit features based on historical IP relationships. The core mechanism involves a malicious domain temporarily resolving to an IP address that is known to host benign domains. For example, if two malicious domains share a malicious IP, and two benign domains share a separate benign IP, a mimicry attack involves the malicious domain temporarily pointing to one of the benign IP addresses. This simple act drastically alters the feature distribution for the malicious domain, causing it to inherit the positive reputation associated with the benign IP.
The impact of this attack is significant: it can increase a domain's reputation score by 25 points. When combined with other evasion techniques, this boost is often sufficient to bypass detection entirely. The appeal of mimicry attacks lies in their cost-effectiveness, requiring no additional financial outlay and minimal time investment. A potential drawback is the introduction of a small resolution error while the mimicry is active. However, this can be mitigated by "priming" the model with these historical records (i.e., performing the mimicry attack for a period) and then reverting the DNS records once the actual malicious campaign begins. For command-and-control (C2) communications, custom resolution software can further mitigate this issue by selectively resolving the domain to the benign IP only when necessary for reputation manipulation.
Popularity List Attacks
Popularity list attacks target the widespread reliance of DNS reputation systems on popularity rankings, often used directly in benign training sets, evaluation, and sometimes even as hard allow-lists in real-world systems. These lists, such as Tranco, Cisco Umbrella, and Cloudflare Radar, frequently derive their rankings from traffic observed by Open DNS recursive resolvers – resolvers that anyone can query.
The attack leverages this accessibility: an adversary can manipulate these lists by generating artificial traffic. The method involves subscribing to a VPN service (costing approximately $10), which allows for the rapid acquisition of unique IP addresses. By sending a single DNS query per IP address to the target recursive resolver, then disconnecting and reconnecting to the VPN for a new IP, an attacker can generate over 12,000 unique IP addresses daily. This volume of traffic allows a malicious domain to achieve a popularity rank of 250,000 on platforms like Cisco Umbrella and Cloudflare Radar. For the Tranco list, which aggregates popularity data over time from multiple sources, the researchers demonstrated that a domain could be pushed to a rank of 1 million in 10 days, and further to 500,000 within 14 days. This direct manipulation of a critical feature effectively elevates the perceived legitimacy of a malicious domain.
Demo / Proof of Concept
▶ Watch: Mimicry attack: altering IP relationships to evade detection (5:20)
To validate the practicality of their attacks against real-world defenses, the researchers partnered with an anonymous security vendor. This collaboration allowed them to test their evasion techniques against a sophisticated, black-box system that incorporates various pre-processing, filtering, and machine learning methods beyond a simple academic model.
The experimental setup involved submitting eight distinct malicious binaries to the vendor. The choice to use malicious binaries, rather than just domains, was strategic: the researchers noted that the default state of a newly registered domain is often considered benign by these systems. Attaching malicious binaries provided an initial disadvantage, forcing the reputation system to actively identify the domain as malicious.
The results were unequivocal:
- A combination of achieving a popularity ranking of 500,000 (via popularity list manipulation) and implementing mimicry attacks successfully led to evasion. This combined attack incurred a financial cost of $10, required 14 days of effort, and necessitated a 2-year domain registration length.
- Alternatively, combining a 2-year registration length with mimicry attacks also resulted in evasion. While this approach avoided the time investment for popularity manipulation, the cost was slightly higher per domain due to the longer registration period.
Beyond these direct attacks on the DNS reputation system, the team also uncovered two sandbox-specific attacks that proved effective against the security vendor's analysis:
- Static String Evasion: If a malicious domain name appeared within the strings of a binary, but the sandbox's dynamic analysis failed to observe an actual DNS lookup for that domain, the domain would be identified as present but would still evade a malicious classification. This highlights a gap where static analysis might flag a string, but the absence of dynamic network activity prevents a full malicious verdict.
- Encrypted DNS Protocol Bypass: The use of encrypted DNS protocols, such as DNS over HTTPS (DoH) or DNS over TLS (DoT), was found to completely bypass the sandbox's detection mechanisms. The reason cited was the lack of memory inspection capabilities within the sandbox to decrypt and analyze the DNS traffic, allowing malicious lookups to proceed undetected.
These demonstrations underscore that even advanced, multi-layered security systems used by industry vendors are vulnerable to practical and low-cost evasion techniques, often due to inherent design trade-offs and overlooked attack vectors.
Defensive Implications
▶ Watch: Popularity list attack using VPNs to manipulate rankings (6:40)
The findings of this research present significant defensive implications for organizations and security vendors relying on DNS reputation systems. The primary conclusion is that current DNS reputation systems are not robust to evasion attacks and can be circumvented with minimal resources and effort ($10 in two weeks, or $15 in under a day for some scenarios).
Defenders must fundamentally reconsider the trade-off between key features, a model's performance, and the susceptibility of those features to evasion. Features that are easily manipulated by attackers, such as popularity rankings or historical IP associations, should be re-evaluated for their weight in the overall reputation score or be protected by more robust verification mechanisms.
Specific actions defenders should consider include:
- Rethink Popularity List Reliance: Security vendors and operators should critically assess their reliance on popularity lists as ground truth for benign domains or for direct allow-listing. Mechanisms to detect artificial traffic generation or to cross-reference popularity with other, harder-to-manipulate indicators are essential.
- Enhance Historical Data Analysis: While historical data is valuable, the mimicry attack demonstrates that even historical IP relationships can be temporarily poisoned. Defenders need more sophisticated methods to track and validate historical associations, perhaps by incorporating broader network context or anomaly detection for sudden reputation shifts.
- Improve Sandbox Evasion Detection:
- Static vs. Dynamic Analysis Integration: Sandboxes should improve the integration between static string analysis and dynamic execution. If a suspicious domain string is found, dedicated efforts should be made to trigger its dynamic resolution or to flag the binary more aggressively, even without an observed lookup.
- Encrypted DNS Inspection: Addressing the bypass via encrypted DNS protocols is critical. This requires sandboxes to implement capabilities for decrypting and inspecting DoH/DoT traffic, potentially through root certificate injection for TLS inspection within the sandbox environment, or by actively monitoring for the establishment of encrypted DNS connections and flagging them as suspicious if associated with known malicious binaries.
- Embrace Adversarial Machine Learning: Develop and deploy models that are explicitly trained to be robust against adversarial attacks. This involves incorporating adversarial examples into training data and using techniques like adversarial training to make models less sensitive to feature manipulation.
- Prioritize Robustness over False Positive Avoidance: While false positives are detrimental, the current prioritization by vendors creates an exploitable gap. A more balanced approach is needed, potentially by implementing tiered alerting systems that can flag highly suspicious but not yet confirmed malicious activity, allowing for deeper investigation without immediately blocking legitimate services.
- Continuous Monitoring and Adaptation: Attackers are constantly evolving. Defenders must implement continuous monitoring of their reputation systems' effectiveness against new evasion techniques and be prepared to adapt their models and feature sets accordingly. Sharing intelligence about new evasion tactics across the industry is also vital.
Key Takeaways
- DNS reputation systems are vulnerable to practical and low-cost evasion attacks. Adversaries can circumvent these systems with as little as $10 and 14 days of effort.
- Mimicry attacks (manipulating historical IP relationships) and popularity list manipulation (generating artificial traffic) are highly effective evasion techniques.
- Real-world security vendors, despite sophisticated defenses, are susceptible. Their prioritization of false positive avoidance creates an exploitable window for attackers.
- Sandbox analysis can be bypassed by embedding domains in binary strings without active lookups or by utilizing encrypted DNS protocols like DoH/DoT.
- Defenders must re-evaluate the trade-offs between feature utility, model performance, and susceptibility to evasion, and enhance capabilities for encrypted DNS inspection and integrated static/dynamic analysis.
- A shift towards more robust, adversarial-aware machine learning models and continuous adaptation is necessary to counter evolving evasion tactics.
About the Speaker(s)
The research presented was a collaborative effort involving multiple institutions. The talk was delivered by Tillson Galloway, who is affiliated with Georgia Tech. He led the presentation, highlighting the key aspects of their work. The co-authors on this significant research include Kleanthis Karakolios, Zane Ma, and Roberto Perdisci from Oregon State University, as well as Manos Antonakakis and Angelos Keromytis from the University of Georgia. This diverse team brought together expertise from various academic security research groups to conduct a comprehensive analysis of DNS reputation system vulnerabilities.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research brutally exposes fundamental weaknesses in current DNS reputation systems, demonstrating practical, low-cost evasion techniques that bypass both academic models and commercial defenses. The detailed breakdown of mimicry and popularity list attacks, coupled with real-world vendor testing and clever sandbox bypasses, provides critical, actionable intelligence and illuminates a serious gap in our collective security posture. This isn't just theoretical; it's a blueprint for compromise that defines the conversation.
Heather Calloway (CISO) — STRONG ACCEPT
This research exposes critical, practical vulnerabilities in widely-used DNS reputation systems, demonstrating that adversaries can bypass sophisticated vendor controls with minimal cost and effort. It forces a necessary re-evaluation of feature reliance, vendor prioritization of false positives, and the institutional accountability for this foundational security gap.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024