Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors

Black Hat Asia 2025 · Day 1 · Briefings

Overview

This talk, "Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors," presented by Jang from Tsinghua University, unveils a critical flaw in how email security systems process messages. The research demonstrates a novel class of protocol-level evasion techniques that manipulate the structure of email messages to bypass even the most sophisticated attachment detectors. Instead of focusing on traditional malware obfuscation, this work highlights parsing discrepancies between email security gateways and end-user clients, allowing malicious payloads to reach inboxes undetected.

Watch on YouTube

Visual summary for Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors
Visual summary for Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors

Key moments

  1. 0:00 Introduction to MIME ambiguities and protocol-level evasion
  2. 2:50 Exploiting parsing discrepancies: Protocol-level evasion explained
  3. 4:00 WannaCry virus bypasses Gmail detection demo
  4. 6:00 Generalizability and feasibility of the attack characteristics
  5. 7:40 MIME structure complexity creates parsing ambiguities
  6. 8:50 Miner: Automated tool for finding email parsing vulnerabilities

Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors

Speakers: Jang, Second-year PhD Student, Network and Information Security Lab, Tsinghua University

Conference: Black Hat Asia

YouTube: https://www.youtube.com/watch?v=eZjP91Ly1r4

Overview

This talk, "Inbox Invasion: Exploiting MIME Ambiguities to Evade Email Attachment Detectors," presented by Jang from Tsinghua University, unveils a critical flaw in how email security systems process messages. The research demonstrates a novel class of protocol-level evasion techniques that manipulate the structure of email messages to bypass even the most sophisticated attachment detectors. Instead of focusing on traditional malware obfuscation, this work highlights parsing discrepancies between email security gateways and end-user clients, allowing malicious payloads to reach inboxes undetected.

The significance of this research cannot be overstated. Email remains the most prevalent vector for cyberattacks, with threats ranging from phishing to advanced persistent threats delivering malware. While email gateways and content detectors are widely deployed as primary defenses, this presentation reveals that their effectiveness can be undermined by subtle ambiguities in the Multipurpose Internet Mail Extensions (MIME) standard. The findings challenge the assumption that robust detection engines alone suffice, emphasizing the overlooked importance of consistent parsing across the entire email delivery chain.

Jang's work introduces a systematic methodology and an automated tool, MyMiner, to discover these vulnerabilities. The comprehensive evaluation against popular email products and clients exposed widespread bypasses, including against major services like Gmail, Outlook, and iCloud. By detailing 24 distinct attack vectors, 19 of which were previously undocumented, the research provides crucial insights for both defenders and standard bodies, underscoring the urgent need for stricter input validation, harmonized parsing implementations, and updated MIME specifications.

Background

▶ Watch: Introduction to MIME ambiguities and protocol-level evasion (0:00)

Traditionally, attackers seeking to bypass email attachment detectors have relied on malware-level evasion techniques. These methods involve directly manipulating the malicious payload through obfuscation, encryption, deforming, or packing. The underlying principle is to exploit weaknesses in the detector's ability to analyze the content itself, for instance, by using packers that the antivirus engine cannot unpack. This approach has led to a continuous "arms race" between attackers and security vendors, with detection engines constantly improving their capabilities to identify and neutralize evolving threats. However, this focus on the detection engine's performance often overshadows other potential vulnerabilities in the email delivery process.

Jang's research shifts this paradigm by focusing on protocol-level evasion. This approach exploits the inherent complexity and flexibility of the MIME (Multipurpose Internet Mail Extensions) standard, which defines how non-ASCII data (like images, audio, and binary files) is structured and transmitted within email messages. MIME extends the original email format (RFC 822) by introducing mechanisms for data encoding and multipart entities, allowing for rich content. However, this complexity also introduces opportunities for ambiguity when messages deviate from strict formatting specifications.

The core premise of protocol-level evasion is that security product developers often concentrate solely on the effectiveness of their detection engines, neglecting the consistency of parsing between these engines and the email clients that ultimately display the message to the user. By crafting malformed email message structures that exploit these parsing discrepancies, attackers can deliver malicious attachments without triggering an alert. A detector might fail to recognize a malicious file due to a misinterpretation of the MIME structure, while a client, exhibiting error-tolerant features or different parsing logic, successfully extracts and presents the dangerous payload to the victim. This fundamental difference in interpretation, rather than a flaw in the detection logic itself, forms the basis of the "Inbox Invasion" attack.

Key Findings

▶ Watch: WannaCry virus bypasses Gmail detection demo (4:00)

The research presented in "Inbox Invasion" yielded several critical findings that underscore the pervasive nature of email attachment detector bypasses via MIME ambiguities:

  1. Widespread Vulnerability: The study conducted a systematic evaluation against 15 popular email products (including Gmail, Outlook, iCloud, Yahoo, Yandex, 163, NetEase, Zoho, and the open-source Amavis filter) and 7 well-known standalone email clients (such as Mozilla Thunderbird, Outlook client, and Mac OS email client). Remarkably, malicious emails crafted with parsing ambiguities successfully bypassed all 16 detectors tested.
  2. Product-Internal Discrepancies: A significant finding was that for 10 of the tested email products, detection bypasses could be triggered even without relying on a third-party client. This indicates that discrepancies exist between a product's own detector and its integrated webmail client in how they interpret malformed email messages, leading to internal inconsistencies.
  3. Extensive Detector-Client Combinations Affected: Out of 128 possible combinations of detectors and clients, a staggering 102 combinations were found to be vulnerable to detection bypasses. This highlights the widespread impact of inconsistent MIME parsing across the email ecosystem.
  4. Novel Attack Vectors: The research identified a total of 180 effective evasion samples, which were further summarized into 24 distinct attack vectors. Crucially, 19 of these attack vectors were previously undocumented, representing novel methods for bypassing email security.
  5. Categorization of Bypass Principles: The identified bypass methods were classified into three main categories based on their underlying principles:
  • Confusion over ambiguous header fields: Exploiting semantic gaps in critical MIME headers (e.g., Content-Type, Content-Transfer-Encoding).
  • Differences in parsing malformed MIME structure: Leveraging error-tolerance in clients that allow them to parse structures rejected by detectors.
  • Inconsistencies in decoding algorithms: Capitalizing on deviations from RFC specifications during data decoding (e.g., Base64, Quoted-Printable).
  1. Automated Discovery Tool: To facilitate this comprehensive study, the researchers developed MyMiner, an automated vulnerability mining tool designed to generate, filter, and test email samples for parsing ambiguities. The tool's design ensures both comprehensive exploration of the attack space and structural compliance for real-world testing.
  2. Responsible Disclosure and Impact: The research team engaged in responsible disclosure with affected vendors. As a result, 10 vendors responded to their report, and 8 of them have either fixed or committed to fixing the identified vulnerabilities. The team also received over $2,000 in bug bounties and had CVE numbers assigned to their discoveries, validating the severity and practical impact of their work.

These findings collectively paint a concerning picture of the current state of email security, where the intricate details of protocol implementation can be exploited to circumvent robust content-based detection mechanisms.

Technical Deep Dive

▶ Watch: Generalizability and feasibility of the attack characteristics (6:00)

The core of Jang's research lies in the systematic approach to discovering MIME parsing ambiguities, facilitated by the custom-built tool MyMiner. This tool operates in three distinct phases: sample generation, sample filtering, and bypass testing.

MyMiner's Methodology

  1. Sample Generation:
  • Syntax Rules: MyMiner begins by extracting ABNF (Augmented Backus-Naur Form) rules from relevant RFC documents (e.g., RFC 2045-2049 for MIME). These rules are used to construct a grammar tree, which represents the hierarchical structure of an email message. By expanding nodes in this tree, MyMiner can build a complete email structure, from headers to nested body parts. For example, a message starts with a header and a CRLF pair, followed by a body. The header further expands into RFC 822 and MIME headers, which themselves contain parameters.
  • Semantic Constraints: Beyond syntax, MyMiner incorporates semantic constraints defined by the MIME standard. These specify control relationships, such as a multipart entity requiring a boundary parameter in its Content-Type header, or a Content-Type: image/* entity containing an image body. These constraints ensure that the initial samples are "legal" according to the standard before mutation.
  • Random Mutation: To induce parsing ambiguities, MyMiner applies three mutation strategies to these legal samples:
  • String-level Mutations: These operate on the leaf nodes of the grammar tree, modifying the email's raw string content. Examples include deleting, inserting, or replacing individual characters (e.g., altering a CRLF pair).
  • Structure-level Mutations: These target sub-trees of the email structure, leading to more significant modifications. An example given is replacing an entire Content-Type header with a Content-Transfer-Encoding header, including its value.
  • Targeted Mutations: This strategy focuses on sensitive fields or common problematic structures based on predefined rules. For instance, inserting a common malformed structure within a header field to efficiently explore known ambiguity points.
  1. Sample Filtering:
  • The random generation process can produce a vast number of overly malformed or entirely unparsable samples. To avoid overwhelming real-world email services and ethical issues, MyMiner employs a filtering phase.
  • It uses a parser test set comprised of popular generic parsing libraries to simulate actual parsing environments locally.
  • For each mutated sample, MyMiner feeds it to these different parsers and compares their outputs, specifically looking for differences in extracted attachments. If all parsers extract the same, non-empty attachment, the sample is deemed unambiguous and discarded.
  • If differences are detected (e.g., one parser extracts a file, another extracts nothing, or they extract different files), the sample is reserved as a valid candidate for real-world testing. This phase also provides feedback to guide more targeted mutations in subsequent generation cycles.
  1. Bypass Testing:
  • The final phase involves sending the filtered, ambiguous email samples to real-world email products and clients. This includes testing complete email products (detector + webmail client) and combinations of detectors with third-party standalone clients. The success criterion is a bypass: the detector fails to flag the malicious attachment, but the client successfully extracts it.

Vulnerability Categories and Examples

The 24 distinct attack vectors identified were categorized into three main principles:

  1. Confusion over Ambiguous Header Fields: This category exploits semantic gaps where detectors and clients interpret critical MIME headers differently.
  • Contradicting Content-Type Headers: A simple yet effective technique involves setting two Content-Type headers for a single-part entity. For example, Content-Type: text/plain followed by Content-Type: multipart/mixed; boundary="xyz". If the detector processes the first header, it expects plain text and finds no boundaries, thus extracting nothing. However, a client prioritizing the second header will correctly parse the multipart structure and extract the malicious payload. This bypassed popular detectors like Gmail and iCloud.
  • Encoded Word in boundary Parameter: RFC 2047 specifies that encoded words (e.g., =?UTF-8?B?Ym91bmRhcnk=?=) are not applicable to the boundary parameter. Many detectors, like Yahoo, q.com, correctly ignore the encoded word and treat the entire string as the boundary, thus failing to find the expected boundary in the body. However, the Mac OS email client was observed to incorrectly support encoded words in the boundary, applying a shorter, incorrect boundary string and subsequently extracting the virus attachment.
  1. Differences in Parsing Malformed MIME Structure: Here, the detector fails to recognize a malformed structure, while the client's error-tolerant features allow it to successfully extract the attachment.
  • Null Byte (\0) in Content-Transfer-Encoding: Inserting a null byte within the Content-Transfer-Encoding header (e.g., Content-Transfer-Encoding: base64\0;) can terminate parsing for many detectors. They might recognize no valid encoding, treating the body as harmless encoded data. In contrast, clients like Thunderbird will often omit the null byte, successfully interpreting the base64 encoding and decoding the malicious content.
  • Empty Boundary String: RFC 2046 mandates that a boundary string must contain at least one character. Detectors for services like 163, NetEase, and Zoho correctly reject an empty boundary but do not block the email, leading them to extract nothing. However, clients like Thunderbird were found to mistakenly recognize the empty boundary, leading to the successful extraction of the virus.
  1. Inconsistencies in Decoding Algorithms: This category leverages deviations from RFC specifications during the decoding of encoded data.
  • Dashes/Dots in Base64 Data: According to RFC 2045, characters like dashes or dots encountered within Base64 encoded data should be discarded during the decoding process. Some detectors, including Gmail, Yahoo, and Yandex, fail to correctly discard these "junk characters" and thus fail to decode the data, resulting in a bypass. Standard-compliant clients, however, correctly omit these characters and successfully decode the malicious payload.
  • Extra Spaces in Quoted-Printable Soft Line Break: The Quoted-Printable encoding scheme uses soft line breaks (=CRLF) to split long lines. RFC specifications state that no spaces should precede the line break, and any such spaces should be discarded. The research found cases where additional spaces were inserted (e.g., = CRLF). Certain detectors failed to discard these spaces, leading them to extract split virus segments or fail decoding. In contrast, standard-compliant clients like Outlook and EM client correctly discarded the spaces and reconstructed the complete virus.

Root Causes of Vulnerabilities

The research attributes these vulnerabilities to three primary practical reasons:

  • Vendor's Trade-off between Usability and Security: Security is often balanced against functionality and user experience, leading to compromises in strict parsing.
  • Incomplete Standard Specifications for Corner Cases: The MIME RFCs, despite their detail, cannot cover every conceivable malformed or ambiguous situation, especially given the complexity and evolution of email.
  • Overly Formalized Definitions of Standards: While detailed specifications are necessary, their highly formal nature (e.g., complex ABNF grammar) can make them challenging for programmers to implement precisely, leading to discrepancies between specification and implementation.

Demo / Proof of Concept

▶ Watch: MIME structure complexity creates parsing ambiguities (7:40)

The talk included a compelling demonstration that vividly illustrated the protocol-level bypass in action. The attacker's objective was to deliver the infamous WannaCry virus to a Gmail recipient as an attachment.

Initially, when the WannaCry payload was sent as a standard attachment, Gmail's robust antivirus engine immediately detected and blocked it, preventing delivery. This is the expected and desired behavior of an email security gateway.

However, the speaker then introduced a seemingly minor, yet critical, modification to the email's MIME structure. Specifically, they added a second, contradictory Content-Transfer-Encoding value, appending base64 right after the original quoted-printable encoding declaration within the email header. This subtle manipulation created a parsing ambiguity.

Upon resending the email with this crafted header, the results were striking: the WannaCry payload was successfully delivered to the Gmail inbox, completely bypassing Gmail's detection mechanisms. The demonstration then proceeded to show that an end-user, fetching this email with a common client like Mozilla Thunderbird, could download the attachment. When the downloaded file was executed, the Windows operating system immediately raised an alarm, confirming that the attachment was indeed the live WannaCry virus payload.

This proof-of-concept effectively highlighted the two key characteristics of this attack:

  • Generalizability: The attack principle is independent of the malicious content itself. While WannaCry was used for demonstration, theoretically, any payload could be delivered using these structural manipulations.
  • Feasibility: The attack requires only manipulation of the email's structure, which is significantly simpler and less resource-intensive than sophisticated virus obfuscation, packing, or encryption techniques. It targets the parsing logic, not the content analysis.

Defensive Implications

▶ Watch: Miner: Automated tool for finding email parsing vulnerabilities (8:50)

The findings from "Inbox Invasion" necessitate a re-evaluation of current email security strategies. Defenders must move beyond solely relying on signature-based or heuristic content detection and address the underlying architectural and parsing inconsistencies. Several mitigation strategies are proposed:

  1. Stricter Input Inspection and Discarding Abnormal Content: The most fundamental defense is to implement rigorous input validation at the email gateway's entry point. Email products should strictly adhere to RFC specifications and discard any email messages that are malformed or contain ambiguous MIME structures. Rather than attempting to "tolerate errors," security gateways should err on the side of caution and reject non-compliant messages, preventing them from reaching the inbox.
  2. Using Paired Clients and Detectors: To minimize parsing discrepancies, organizations should encourage or enforce the use of email clients that are developed by, or tightly integrated with, their chosen email service provider. This increases the likelihood that the detector and the client share a consistent parsing logic, reducing the attack surface created by differing interpretations.
  3. Integrating a Unified Parsing Component: Before invoking third-party detectors or passing messages to various internal components, email gateways should integrate a single, robust, and RFC-compliant parsing component. This ensures that all subsequent security checks and client-facing interfaces operate on a consistent interpretation of the email's structure, eliminating discrepancies.
  4. Updating Existing Email Standards: The research highlights that some ambiguities stem from incomplete or overly complex standard specifications. There is a clear need for the Internet Engineering Task Force (IETF) to review and update existing MIME RFCs. This includes providing clearer guidance for corner cases, simplifying overly formalized definitions, and adapting standards to modern transmission environments (e.g., widespread 8-bit transmission support, which could reduce the necessity of certain encoding steps that introduce ambiguity). This is a long-term initiative but critical for foundational security.

By implementing these measures, defenders can significantly enhance the resilience of email gateways against protocol-level evasion techniques and close the parsing gap that attackers are currently exploiting.

Key Takeaways

  • Protocol-level evasion is a significant and overlooked threat: Attackers can bypass email attachment detectors by manipulating MIME structures, rather than just obfuscating malware content.
  • MIME complexity leads to widespread vulnerabilities: The intricate and sometimes ambiguous nature of the MIME standard creates parsing discrepancies between detectors and email clients.
  • MyMiner is an effective automated discovery tool: The developed tool successfully identified 19 novel bypass methods and 180 effective evasion samples across various email products.
  • Major email services are vulnerable: Popular platforms like Gmail, Outlook, and iCloud, along with numerous detector-client combinations, were found to be susceptible to these bypasses.
  • Consistent parsing is paramount: Security lies not just in a powerful detection engine, but in a unified and RFC-compliant parsing logic across all components of the email delivery chain.
  • Standards and implementations need urgent review: Incomplete RFC specifications, overly formal definitions, and vendor trade-offs contribute to these vulnerabilities, necessitating updates to standards and stricter, consistent implementations.

About the Speaker(s)

Jang is a second-year PhD student at the Network and Information Security Lab of Tsinghua University. His research primarily focuses on critical areas within cybersecurity, including network protocol security, internet infrastructure measurement, and email security. Jang's previous research findings have already garnered recognition from several internet vendors, underscoring his expertise and contributions to the field.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

Jang's 'Inbox Invasion' isn't just another talk; it's a fundamental indictment of email security. By systematically exposing how MIME ambiguities allow malicious attachments to bypass every major email detector and client combination, this research uncovers a critical, overlooked attack surface. The development of MyMiner, the identification of 19 novel attack vectors, and the live demo of WannaCry bypassing Gmail are not just impressive – they're a wake-up call that redefines the threat model for email, proving that protocol-level flaws are as dangerous as any zero-day.

Heather Calloway (CISO) — MUST SEE

This research on MIME ambiguities is a critical wake-up call for email security. It fundamentally challenges the assumption that robust content detectors alone suffice, exposing a systemic flaw where parsing discrepancies between security gateways and end-user clients allow malicious attachments to bypass defenses. The findings demand immediate executive attention to architectural consistency, vendor accountability, and a re-evaluation of our reliance on error-tolerant parsing. This isn't about better malware analysis; it's about fixing the email delivery chain at a foundational level.

→ Top-rated talks at Black Hat Asia 2025

All talks from Black Hat Asia 2025