My ZIP isn't your ZIP: Identifying and Exploiting Semantic Gaps Between ZIP Parsers
Yufan You (Chan University)
34th USENIX Security Symposium (USENIX Security '25) · Day 1 · Software Security 1
Overview
In this compelling talk from USENIX Security, Yufan You presented groundbreaking research on semantic gaps in ZIP file format parsing, revealing a widespread and critical vulnerability across numerous applications and systems. The core premise is deceptively simple yet profoundly impactful: the same ZIP archive can be interpreted differently by various parsing tools, leading to disparate content being extracted or displayed depending on the software used. This discrepancy, termed a semantic gap, is not a rare anomaly but, as the research demonstrates, a pervasive norm within the ZIP ecosystem.

Key moments
- 0:00 Introduction to ZIP semantic gaps problem
- 2:00 SID div: The differential fuzzer for uncovering gaps
- 3:25 Staggering results: Inconsistency is the norm
- 5:00 Exploitation 1: Bypassing email antivirus scanners
- 6:08 Many common file formats are hidden ZIPs
- 7:00 Exploitation 2: Office documents and plagiarism evasion
- 8:00 Forging signatures & VS Code extension marketplace attacks
- 10:00 Responsible disclosure, bounties, and CVEs awarded
My ZIP isn't your ZIP: Identifying and Exploiting Semantic Gaps Between ZIP Parsers
Speakers: Yufan You
Conference: USENIX Security
YouTube: https://www.youtube.com/watch?v=N9Q6f_H0j80
Overview
In this compelling talk from USENIX Security, Yufan You presented groundbreaking research on semantic gaps in ZIP file format parsing, revealing a widespread and critical vulnerability across numerous applications and systems. The core premise is deceptively simple yet profoundly impactful: the same ZIP archive can be interpreted differently by various parsing tools, leading to disparate content being extracted or displayed depending on the software used. This discrepancy, termed a semantic gap, is not a rare anomaly but, as the research demonstrates, a pervasive norm within the ZIP ecosystem.
The implications of these semantic gaps are severe, extending far beyond mere data inconsistencies. Attackers can meticulously craft malicious ZIP files that appear benign to security scanners, such as antivirus software or email gateways, yet deliver a harmful payload when opened by a victim's system. This allows for sophisticated bypasses of established security controls, enabling everything from malware delivery and document forgery to supply chain attacks affecting widely used platforms. The talk highlights that the ubiquity and historical evolution of the ZIP format, coupled with its often-ambiguous specification, have created a fragmented landscape ripe for exploitation.
This research marks the first systematic study of these inconsistencies, moving beyond individual vulnerabilities to expose a fundamental architectural flaw in how ZIP files are processed globally. By introducing a novel differential fuzzer, the speaker and their team uncovered the extent of the problem, categorized 14 distinct types of ambiguities, and demonstrated five real-world exploitation scenarios. The findings call for a fundamental shift in how developers and security professionals approach file format parsing, emphasizing the urgent need for strict format validation over the current "liberal acceptance" paradigm.
Background
▶ Watch: Introduction to ZIP semantic gaps problem (0:00)
The ZIP file format, introduced in 1989, has become an indispensable standard for data compression and archiving. Its widespread adoption means that virtually every computing environment, from operating systems to web browsers and specialized applications, relies on ZIP for various functionalities. However, this ubiquity comes at a significant cost: the original specification is notoriously complex, contains redundant metadata, and often leaves crucial details open to implementer interpretation. This lack of strictness has led to a proliferation of diverse ZIP parsers, each with its own quirks and assumptions, developed across 19 different programming languages and integrated into countless products.
Historically, individual vulnerabilities related to ZIP processing have been reported. These often involve specific malformed archives or edge cases that a particular parser handles incorrectly. However, prior to this research, there had been no systematic, large-scale investigation into the consistency of these parsers. The prevailing assumption has been that if a ZIP file is valid, all compliant parsers should yield identical results when extracting its contents. The research presented here fundamentally challenges this assumption, revealing that the problem is not merely about bugs in individual parsers, but a systemic issue stemming from the inherent ambiguities of the ZIP specification itself.
The core issue arises from the format's structure, which often stores the same piece of information in multiple locations. For instance, a file's name or size might be recorded in both a local file header and the central directory. When these entries conflict, different parsers may prioritize one over the other, leading to divergent interpretations. Furthermore, aspects like path normalization, handling of extra fields, and even the precise location of the central directory (which is critical for understanding the archive's overall structure) can be interpreted differently. This environment of ambiguity creates the perfect storm for semantic gaps, where an attacker can craft a single ZIP file that presents different "semantic meanings" to different parsers, enabling sophisticated evasion techniques.
Key Findings
▶ Watch: Staggering results: Inconsistency is the norm (3:25)
The research introduced SID-diff, a novel differential fuzzer designed to systematically uncover semantic gaps between ZIP parsers. SID-diff operates by taking a valid ZIP file from a corpus, applying a series of ZIP-level and byte-level mutations, and then feeding the mutated file to a suite of tested parsers. The outputs from these parsers are then compared, typically using hash values, to detect inconsistencies. The fuzzer employs the UCB algorithm to balance exploration of new mutation strategies with exploitation of promising ones, ensuring that the most effective mutations for uncovering inconsistencies are prioritized. Only "interesting samples" – those that expose inconsistencies between parser pairs – are retained in the corpus, guiding future mutations and refining the fuzzer's effectiveness.
The evaluation of SID-diff yielded staggering results. The team tested 50 common products, including popular libraries and applications like WinRAR, 7-Zip, and TypeZip, written in 19 different programming languages. Out of 2,225 possible pairs of parsers, SID-diff found that only four pairs were consistent with each other. This means that for nearly any two ZIP parsers chosen at random, it is possible to craft a ZIP file that they will interpret differently. The research unequivocally concludes that inconsistency is not the exception but the norm within the ZIP ecosystem. It's crucial to note that for a pair to be considered inconsistent, both parsers must successfully extract a file, but produce different contents. If one parser detects an abnormality and reports an error, the pair is still considered consistent, highlighting the severity of the findings where silent discrepancies occur.
Through their extensive analysis, the researchers categorized 14 distinct types of parsing ambiguities. These ambiguities fall into three main groups:
- Confusion over redundant metadata: This includes discrepancies in how parsers handle conflicting information for fields like file names, file sizes, or modification times, which might be stored in multiple locations within the ZIP structure.
- Differences in how path processing is handled: This covers variations in how parsers normalize file paths, resolve directory traversal sequences (e.g.,
../), or interpret unusual character encodings within path names. - Disagreements on the fundamental structure of the ZIP file itself: This encompasses more profound inconsistencies, such as differing interpretations of the location of the central directory or the presence of overlapping data structures, which can drastically alter how an archive's contents are perceived.
The detailed analysis of each of these 14 ambiguities, along with corresponding construction code for proof-of-concept files, is available in the researchers' full paper and artifact, providing concrete examples of how these theoretical ambiguities translate into exploitable conditions.
Technical Deep Dive
▶ Watch: Many common file formats are hidden ZIPs (6:08)
The core of the problem lies in the design and evolution of the ZIP file format. While seemingly straightforward, its internal structure is complex, featuring multiple data structures that can contain overlapping or redundant information. Key components include:
- Local File Headers (LFH): Precede each file's data and contain metadata specific to that file, such as its name, compressed size, uncompressed size, and CRC-32 checksum.
- File Data: The actual compressed content of the file.
- Data Descriptors (Optional): Can follow the file data and contain post-compression metadata if the LFH fields were initially set to zero (e.g., for streaming compression).
- Central Directory (CD): A directory of all files in the archive, typically located at the end. Each entry in the CD mirrors information from the LFH but also includes additional details like the starting offset of the LFH, external file attributes, and comments.
- End of Central Directory Record (EOCD): Marks the end of the CD and contains pointers to its start, total number of entries, and the overall size of the archive.
The ambiguities primarily arise when these redundant data structures contain conflicting information. For example, a file's uncompressed size might be declared differently in its LFH and its corresponding CD entry. A strict parser might prioritize the CD entry (as it's often considered the definitive index), while a more "liberal" parser might trust the LFH or the actual data length encountered. The official ZIP specification, historically, has not provided clear guidance on how to resolve such conflicts, leaving it up to individual implementers.
The three main groups of ambiguities identified by the research highlight these structural weaknesses:
- Redundant Metadata Confusion: This category exploits situations where fields like file names or sizes are specified contradictorily. For instance, an attacker could embed a benign filename in the LFH (which an antivirus scanner might read first) and a malicious filename in the CD entry (which a victim's archiver might prioritize). Similarly, differing file sizes could cause one parser to truncate a file, appearing harmless, while another extracts the full, malicious payload.
- Path Processing Differences: Parsers vary widely in how they handle file paths within the archive. Some may strictly adhere to specific character sets, while others are more lenient. Crucially, differences arise in path normalization. For example, paths containing
../sequences (indicating directory traversal) might be sanitized by some parsers but not by others, leading to files being extracted to different locations. Similarly, paths with null bytes or non-printable characters can be interpreted differently, enabling a file to "hide" from one parser while being accessible to another. - Fundamental Structure Disagreements: This is perhaps the most critical category, as it involves disagreements on how to interpret the very layout of the ZIP file. The location of the central directory, for instance, is crucial for correctly parsing the entire archive. If an attacker can craft a ZIP file where the EOCD record points to one location for one parser and another for a different parser, they can effectively present two entirely different archives within the same file. This could involve embedding additional, malicious data before the CD that one parser ignores but another interprets as part of a legitimate file, or even completely altering the set of files that appear to be present in the archive. Such discrepancies can lead to one parser seeing a harmless set of files while another sees a completely different, malicious set.
The research also implicitly highlights the challenge of "ZIP-based formats." Many modern file formats, such as Microsoft Office documents (.docx, .xlsx), Android Application Packages (.apk), Java Archives (.jar), and VS Code extensions (.vsix), are fundamentally ZIP files containing other structured data (e.g., XML, compiled code). Any application that processes these formats is, by extension, a ZIP parser and thus potentially vulnerable to these same ambiguities. This vastly expands the attack surface, moving beyond traditional archive tools to a wide array of application ecosystems.
Demo / Proof of Concept
▶ Watch: Exploitation 2: Office documents and plagiarism evasion (7:00)
The talk presented five compelling real-world exploitation scenarios, demonstrating the practical impact of these semantic gaps. These examples serve as powerful proofs of concept for the security implications of inconsistent ZIP parsing.
- Bypassing Secure Email Gateways:
This scenario leverages the difference between server-side and client-side parsing. Secure email gateways (SEGs) like those used by Gmail, Zoho, and other providers, scan attachments on the server before delivery. An attacker can craft a ZIP file where the SEG's parser sees only harmless content (e.g., an empty file or a benign image), allowing the email to pass through. However, when the victim downloads and opens the exact same ZIP file on their machine using their local archiver (e.g., WinRAR, 7-Zip), the local parser, due to a semantic gap, extracts a malicious payload (e.g., an executable virus). The user, seeing a "Scanned by Gmail" message, might trust the attachment, leading to system infection. The researchers successfully crafted ZIP files to bypass all 10 popular email providers with antivirus functionality that they tested, receiving bug bounty rewards from Google, Zoho, and others for these findings.
- Office Document Woofing:
Many modern document formats, such as Microsoft Word's .docx, are essentially ZIP archives containing XML files. This means office applications are, in effect, ZIP parsers. The "woofing" attack exploits semantic gaps to present different content based on the application. For example, a dishonest student could create a .docx file that, when opened by a plagiarism checker, displays benign, original content. However, when the same document is opened by a supervisor in Microsoft Word, it displays plagiarized content. This allows the student to evade academic integrity checks. The underlying mechanism involves embedding two conflicting versions of content within the ZIP structure, where different parsers (plagiarism checker's embedded ZIP parser vs. MS Word's parser) prioritize different versions.
- Forging Digital Signatures in LibreOffice:
This attack targets applications that use digital signatures to verify document integrity. In LibreOffice, the researchers crafted a document containing two different versions of its content. When LibreOffice's signature verifier inspects the file, it checks the original, signed content and reports the signature as valid. However, when the document is subsequently rendered for the user, a semantic gap causes LibreOffice's rendering engine to display the manipulated, unsigned content. This creates a scenario where a user believes they are viewing a legitimate, signed document, but the content has been altered.
- Bypassing Signature Verification in Spring Boot Nested JARs:
The Java application ecosystem also proved vulnerable. Spring Boot supports a nested JAR format, where executable JAR files can contain other JAR files. Spring Boot's system uses Java's built-in signature verifier but relies on a custom ZIP content reader for nested archives. The researchers crafted a nested JAR file that bypassed this signature verification. The Java verifier would see the original, signed content, deeming it valid. However, due to the custom ZIP content reader's interpretation, the application would execute manipulated, unsigned content embedded within the nested JAR, again demonstrating a split-view attack based on parser differences. This led to multiple CVEs being assigned.
- VS Code Extension Marketplace Takeover:
This scenario demonstrates a supply chain attack. A VS Code extension's identity is defined by its publisher namespace and extension name (e.g., developer.extensionName). The marketplace server is designed to prevent attackers from uploading extensions under another publisher's namespace. The researchers exploited a ZIP ambiguity to craft a malicious VS Code extension package. When this package was uploaded, the marketplace server's parser interpreted the package as belonging to the attacker's authorized ID (e.g., attacker.maliciousExtension) and accepted the upload. However, when a user downloaded and installed this same extension package, their VS Code client's parser interpreted the extension as having the ID of a legitimate, trusted publisher (e.g., developer.trustedExtension). This allowed the malicious extension to overwrite and effectively take over a trusted extension, posing a significant supply chain threat.
These diverse exploitation scenarios underscore the pervasive nature of ZIP parsing inconsistencies and their potential to undermine security mechanisms across various application domains.
Defensive Implications
▶ Watch: Responsible disclosure, bounties, and CVEs awarded (10:00)
The findings from this research demand a significant re-evaluation of current security practices and software development paradigms concerning file format parsing. The fundamental takeaway for defenders is that the long-standing "liberal in what we accept" approach to parsing must be replaced with strict format validation.
- Strict Format Validation: Developers of any application that processes ZIP files or ZIP-based formats (like .docx, .apk, .jar, .vsix) must implement rigorous validation checks. This means not just checking if a file is syntactically a ZIP, but also verifying the consistency of redundant metadata. If conflicting information is found (e.g., different file sizes in the LFH and CD), the parser should either reject the file as malformed, issue a warning, or consistently default to the most secure interpretation, rather than silently picking one.
- Security Scanners and Gateways: Antivirus software, secure email gateways, and other security scanning solutions are on the front lines of defense. These systems must employ highly robust and, critically, consistent ZIP parsers. Ideally, scanners should emulate the parsing behavior of common client-side applications as closely as possible, or even use multiple parsing engines to detect discrepancies. Any file that causes a semantic gap between different parsing engines should be flagged as suspicious or blocked outright, rather than being allowed through based on a single benign interpretation.
- Application-Specific Defenses:
- Office Applications: Should validate the consistency of content when processing .docx, .xlsx, etc., especially when signatures are involved. If multiple content versions are present or if signed content differs from rendered content, it should be flagged.
- Java Ecosystem: Custom ZIP content readers (as seen in Spring Boot) must not bypass or diverge from the behavior of standard, secure verifiers. Signature verification processes need to be tightly coupled with the actual content being processed and executed.
- Marketplace Security: Application marketplaces (like VS Code's) need to ensure that the identity of an uploaded package, as interpreted by the server, is identical to how it will be interpreted by the client. This requires careful alignment of their respective ZIP parsers and robust checks against identity spoofing through ambiguities.
- Developer Awareness and Library Selection: Developers need to be educated about the inherent ambiguities of the ZIP format. When choosing or implementing ZIP parsing libraries, they should prioritize those known for strictness, consistency, and security-minded error handling. Auditing existing codebases for vulnerable parsing logic is also crucial.
- Backward Compatibility vs. Security: While backward compatibility is often a concern, the research demonstrates that silently accepting ambiguous or malformed ZIP files introduces severe security risks. A shift towards rejecting such files or providing clear warnings, even if it occasionally breaks compatibility with poorly crafted archives, is a necessary trade-off for enhanced security.
- Beyond ZIP: The talk's conclusion, "inconsistencies can be dangerous... similar issues also exist in other file formats and protocols," serves as a broader warning. The principles of strict validation and differential testing applied to ZIP files should be extended to other complex file formats and network protocols that might suffer from similar ambiguities and implementer discretion.
Key Takeaways
- ZIP Parsing is Fundamentally Fragmented: The research uncovered that almost no two ZIP parsers behave identically, with 2,225 pairs of parsers showing inconsistencies, and only four pairs being consistent.
- Semantic Gaps Lead to Critical Security Vulnerabilities: Discrepancies in how ZIP files are interpreted enable sophisticated attacks, allowing malicious content to bypass security scanners and deceive users.
- Widespread Exploitation Scenarios: The talk demonstrated five real-world exploitation types: bypassing secure email gateways, forging digital signatures (LibreOffice, Spring Boot), "woofing" office documents (presenting different content to different viewers), and taking over VS Code extensions.
- Problem Extends Beyond Traditional Archives: Many common file formats (e.g., .docx, .apk, .jar, .vsix) are ZIP-based, making applications that process them equally vulnerable to these parsing ambiguities.
- Strict Format Validation is Essential: The current "liberal in what we accept" approach to ZIP parsing must be replaced with rigorous validation to detect and reject ambiguous or conflicting metadata.
- Inconsistencies are a Broader Threat: The findings highlight that similar semantic gaps likely exist in other complex file formats and communication protocols, warranting further research and defensive measures.
About the Speaker(s)
Yufan You is a researcher from Zhejiang University. Their work focuses on identifying and exploiting semantic gaps between ZIP parsers, contributing significantly to the understanding of file format vulnerabilities and their widespread security implications.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
Systematic, first-of-its-kind study that exposes ZIP parser inconsistency as a near-universal condition rather than a collection of isolated bugs. The differential fuzzing methodology is rigorous, the 14-class taxonomy of ambiguities is a genuine contribution, and the five exploitation chains — spanning email gateways, digital signature forgery, and a full supply chain takeover of the VS Code marketplace — are the kind of demos that make vendors quietly pull engineers into side rooms. This is the rare academic paper that also ships real bug bounties and CVEs.
Heather Calloway (CISO) — SOLID
Rigorous, first-of-its-kind research that maps a real and underappreciated attack surface — ZIP parser inconsistency — with strong empirical grounding and credible real-world exploitation. The technical work is excellent, but the talk lands as a research briefing rather than an operational decision. It tells security teams what is broken; it does not tell security programs what to change first, how to triage exposure, or how to pressure vendors.
→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)
All talks from 34th USENIX Security Symposium (USENIX Security '25)