Gotta Cache ‘em all bending the rules of web cache exploitation

Martin Doyhenard

DEF CON 32 Main Stage · Day 1 · Main Stage

Overview

In this DEF CON 32 presentation, "Gotta Cache ‘em all: bending the rules of web cache exploitation," Martin Doyhenard delves into novel techniques for exploiting web cache vulnerabilities, moving beyond traditional methods to achieve arbitrary web cache deception and poisoning. The talk focuses on how discrepancies in URL parsing between web cache proxies (such as CDNs like Cloudflare, Cloudfront, and Akamai) and origin servers can be leveraged to manipulate cached content. Doyhenard demonstrates how attackers can exploit these subtle differences to steal sensitive user information, poison widely accessed resources, and even achieve full website replacement.

Watch on YouTube

Visual summary for Gotta Cache ‘em all bending the rules of web cache exploitation by Martin Doyhenard
Visual summary for Gotta Cache ‘em all bending the rules of web cache exploitation by Martin Doyhenard

Key moments

  1. 0:00 Introduction and overview of talk agenda
  2. 1:30 Detailed explanation of how web caches function
  3. 4:15 Understanding Omar Gil's web cache deception attack
  4. 6:00 Limitations of existing web cache deception attacks
  5. 6:40 Introducing URL parser discrepancies as an attack vector

Gotta Cache ‘em all bending the rules of web cache exploitation

Speakers: Martin Doyhenard

Conference: DEF CON 32

YouTube: https://www.youtube.com/watch?v=70yyOMFylUA

Overview

In this DEF CON 32 presentation, "Gotta Cache ‘em all: bending the rules of web cache exploitation," Martin Doyhenard delves into novel techniques for exploiting web cache vulnerabilities, moving beyond traditional methods to achieve arbitrary web cache deception and poisoning. The talk focuses on how discrepancies in URL parsing between web cache proxies (such as CDNs like Cloudflare, Cloudfront, and Akamai) and origin servers can be leveraged to manipulate cached content. Doyhenard demonstrates how attackers can exploit these subtle differences to steal sensitive user information, poison widely accessed resources, and even achieve full website replacement.

This research is particularly significant because it addresses the limitations of previously known web cache attacks, which often required specific backend configurations or were mitigated by existing CDN protections. By targeting the fundamental process of URL interpretation, Doyhenard's methods offer a more generalized approach to web cache exploitation, posing a substantial threat to applications relying on these caching mechanisms for performance and scalability. The talk aims to equip both attackers and defenders with a deeper understanding of these sophisticated attack vectors and the underlying architectural nuances that make them possible.

Background

▶ Watch: Introduction and overview of talk agenda (0:00)

Web caches are fundamental components of modern web infrastructure, designed to improve website performance and reduce the load on origin servers. When a client requests a resource, the cache proxy intercepts the request. Its primary function is to determine if it can serve the requested resource from its local storage without contacting the origin server. This decision is based on a calculated cache key, which is typically derived from the URL and hostname of the request. For static resources, the cache key focuses on identifying the specific content, disregarding elements like cookies or user agents that might vary between requests for the same resource.

If the cache proxy does not have a valid, fresh copy of the resource, it forwards the request to the origin server. The origin server then processes the URL, parsing its path to identify the specific endpoint—be it a static file (e.g., image.png) or a dynamic application handler. Upon receiving the response from the origin server, the cache proxy evaluates various rules to decide whether to store this response in its cache. These rules often involve inspecting the Cache-Control header in the response, but also critically, examining the request's path. Many CDNs, for instance, have default rules that cache resources with common static file extensions like .css, .js, .png, or .jpg. If a resource matches these criteria, it is stored in the cache using the previously calculated key, ready to be served to subsequent identical requests without involving the origin server.

Prior to Doyhenard's work, a notable attack vector was web cache deception (WCD), famously discovered by Omar Gil in 2017. WCD exploits scenarios where backend servers use path parameters or special URL mappings. For example, a backend might interpret /myaccount/param1/param2 as simply /myaccount, generating a response containing sensitive user information (like email or credit card details) associated with the authenticated user. An attacker would craft a malicious link, such as https://example.com/myaccount/attacker.js, and trick a victim into visiting it. The origin server, using the victim's session cookies, would process this as /myaccount and return the sensitive HTML content. However, the cache proxy, unaware of the backend's specific path parameter handling, would see the .js extension and mistakenly categorize the response as a static JavaScript file. It would then cache this sensitive HTML content under the key /myaccount/attacker.js. The attacker could then request https://example.com/myaccount/attacker.js without the victim's cookies and retrieve the cached, sensitive response. The limitation of this attack was its reliance on specific backend path parameter mappings and the requirement that the "parameter" (e.g., attacker.js) did not alter the backend's response in an undesirable way.

Another significant web cache vulnerability is web cache poisoning (WCP). Unlike WCD, which aims to steal victim-specific data, WCP focuses on injecting malicious content into a cache entry that is widely accessed by many users. For instance, if an attacker can manipulate a request to the homepage (/) such that the origin server returns a malicious payload (e.g., an XSS payload reflected from a non-cacheable header like Cookie), and the cache proxy mistakenly caches this response, then all subsequent users requesting the homepage will receive the poisoned content. CDNs like Cloudflare have implemented protections, such as Web Cache Deception Armor, which attempts to detect WCD by checking for mismatches between the requested file extension (e.g., .js) and the actual content type of the response (e.g., text/html). If a mismatch is detected, the response is typically not cached, mitigating many classic WCD attacks.

These existing protections and the inherent limitations of previous attack vectors created a need for more robust and generalizable web cache exploitation techniques, which Doyhenard's research addresses by focusing on the fundamental discrepancies in URL parsing.

Key Findings

▶ Watch: Detailed explanation of how web caches function (1:30)

Martin Doyhenard's research uncovers several key findings that significantly advance the field of web cache exploitation:

  1. Exploitation of URL Parser Discrepancies: The core discovery is that different components in the web request chain—specifically, cache proxies and origin servers—often interpret URLs differently due to varying implementations of RFC standards regarding URL delimiters. This ambiguity allows attackers to craft URLs that are parsed one way by the cache (leading to caching) and another way by the backend (leading to a desired response, e.g., sensitive data or a specific page).
  1. Arbitrary Web Cache Deception (AWCD): By leveraging these parser discrepancies, Doyhenard demonstrates how to achieve web cache deception for any desired endpoint, rather than relying on specific backend path parameter mappings. This involves using custom delimiters (like the dollar sign $) that are recognized by the origin server but not by the cache proxy. This allows the attacker to control both the backend's perceived path (e.g., /styles.css) and the cache proxy's perceived path (e.g., /styles.css$a.js with a static extension) simultaneously, facilitating the caching of sensitive information.
  1. Arbitrary Web Cache Poisoning (AWCP) through Key Modification: The research shows how to achieve web cache poisoning by modifying the cache key itself. By exploiting parser discrepancies, an attacker can make the cache proxy store a malicious response under a widely accessed key (e.g., /) while the origin server processes a different, attacker-controlled path. This bypasses common WCP defenses that rely on detecting malicious content within a fixed key.
  1. Bypassing CDN Protections: Doyhenard's techniques are designed to circumvent modern CDN defenses, such as Cloudflare's Web Cache Deception Armor. By ensuring that the cache proxy believes it is caching a legitimate static resource (due to a crafted extension in the URL), the content-type mismatch detection can be bypassed, allowing HTML or other content to be cached under a "static file" key.
  1. Full Website Replacement (Conceptual): The ultimate implication of combining arbitrary web cache deception and poisoning techniques is the ability to achieve full website replacement. While the detailed mechanism for this was not fully elaborated in the provided transcript, the concept implies that by controlling both the content that gets cached and the keys under which it is stored, an attacker could theoretically replace any part or all of a website's cached content with arbitrary malicious data, effectively defacing or hijacking the entire user experience.

These findings highlight a significant attack surface in web infrastructure, emphasizing the critical need for consistent URL parsing logic across all components of the request-response chain.

Technical Deep Dive

▶ Watch: Understanding Omar Gil's web cache deception attack (4:15)

The core of Doyhenard's novel web cache exploitation techniques lies in the subtle yet critical differences in how Uniform Resource Locators (URLs) are parsed by web cache proxies (such as CDNs) and origin servers. The Request for Comments (RFCs) that define URLs (e.g., RFC 3986) provide guidelines but also allow for some implementation flexibility, particularly concerning what characters can act as delimiters to separate different components of a URL (scheme, credentials, host, path, query, fragment). This flexibility is the root cause of the discrepancies exploited.

Web Cache Operation and Key Calculation

When a client sends an HTTP request, the cache proxy is the first to receive it. Its immediate task is to compute a cache key to identify the requested resource. For static content, this key primarily consists of the URL's path and the hostname. Other elements like cookies or user agents are typically ignored because they are dynamic and would prevent efficient caching of static assets. If a resource is not found in the cache, the request is forwarded to the origin server.

The origin server, upon receiving the request, performs its own URL parsing. It extracts the path and uses it to map to an endpoint. This endpoint could be a static file (e.g., /images/logo.png) or a dynamic application handler (e.g., /api/users/profile). The response generated by the origin server is then sent back to the cache proxy.

The cache proxy then decides whether to cache this response. This decision is based on:

  1. Response Headers: Primarily the Cache-Control header, which explicitly instructs caching behavior.
  2. Request Path Analysis: The cache proxy often has rules based on the request's path. A common rule, used by most CDNs (e.g., Cloudflare, Cloudfront), is the static extension rule. This rule dictates that if a URL path ends with a recognized static file extension (e.g., .css, .js, .png, .jpg, .gif, .svg, .pdf), the resource is considered static and eligible for caching. For example, /styles.css would be cached because of the .css extension.

Traditional Web Cache Deception (WCD) Limitations

As discussed in the background, Omar Gil's 2017 WCD attack relied on path parameters where a backend server might interpret /myaccount/param.js as /myaccount. The backend would return sensitive HTML, but the cache proxy would see the .js extension and cache the HTML under /myaccount/param.js. The attacker could then retrieve this cached, sensitive HTML.

However, this attack has limitations:

  • It requires specific backend configurations that treat parts of the path as parameters without affecting the core functionality.
  • Many CDNs, like Cloudflare with its Web Cache Deception Armor, implement defenses. This armor checks if the requested extension (e.g., .js) matches the Content-Type header of the response (e.g., text/html). If they don't match, the response is not cached, preventing the deception.

Arbitrary Web Cache Deception (AWCD) through Delimiter Discrepancies

Doyhenard's breakthrough is to exploit URL parser discrepancies between the cache proxy and the origin server. The RFCs specify numerous delimiters (e.g., ?, #, /, &, =, ;) but also allow implementations to define additional or interpret existing ones differently. This creates a potential for divergence.

Consider a scenario where:

  • The origin server recognizes a character, say $ (dollar sign), as a delimiter, effectively truncating the URL path at that point.
  • The cache proxy, on the other hand, does not recognize $ as a special delimiter; it treats it as a regular character within the path.

An attacker can craft a URL like https://example.com/styles.css$secret.js.

Here's how the components would parse it:

  1. Origin Server's Interpretation: Because the origin server treats $ as a delimiter, it might parse /styles.css$secret.js as simply /styles.css. If /styles.css is a dynamic endpoint that, when accessed with a victim's cookies, returns sensitive user information embedded within an HTML response (e.g., a personalized stylesheet that accidentally includes user data), then the origin server would generate this sensitive HTML.
  2. Cache Proxy's Interpretation: The cache proxy, not recognizing $ as a delimiter, would interpret the entire path as /styles.css$secret.js. Crucially, it would observe the .js extension at the end of the path. Based on its static extension rule, it would deem this a cacheable static resource.

The result: The cache proxy stores the sensitive HTML response (from /styles.css) under the cache key /styles.css$secret.js. An attacker can then request /styles.css$secret.js without the victim's cookies, and the cache proxy will serve the cached sensitive information.

This attack bypasses the WCD Armor because the cache proxy believes it is caching a JavaScript file (due to .js extension), even though the content is text/html. The "mismatch" check is circumvented because the cache's decision to cache is based on the attacker-controlled extension in the URL, not a comparison of the backend's intended Content-Type with the URL's extension.

The main challenge with this specific AWCD technique is that the cached resource is stored under a unique, attacker-crafted key (e.g., /styles.css$secret.js). This still requires user interaction (the victim must visit the crafted URL) to populate the cache with their sensitive data.

Arbitrary Web Cache Poisoning (AWCP)

Web cache poisoning aims to store malicious content under a commonly accessed cache key, affecting many users. Traditional WCP often relies on finding unkeyed inputs (like Cookie headers) that are reflected in the response and then caching that response. However, CDNs are good at detecting this.

Doyhenard's techniques can extend to AWCP by focusing on modifying the cache key itself. While the full details of combining both techniques for arbitrary key modification were not explicitly detailed in the provided (truncated) transcript, the agenda item "modify the keys of a store resource" implies a mechanism to control what key a malicious response is stored under.

A conceptual approach could involve:

  • Using a delimiter discrepancy to make the origin server process a request for a benign, commonly accessed path (e.g., /).
  • Simultaneously, making the cache proxy generate a cache key that includes an attacker-controlled payload, but still resolves to the commonly accessed path from the origin server's perspective.

The talk mentions attacking "static directories" rules. Many applications configure rules to cache everything under a /public/ or /assets/ directory. An attacker might craft a URL like /public/$payload.js.

  • Origin Server: If the origin server treats $ as a delimiter, it might interpret this as /public/, returning the default content for that directory.
  • Cache Proxy: The cache proxy sees /public/$payload.js, and because it starts with /public/ (matching a static directory rule) and ends with .js (matching a static extension rule), it caches the origin's response.

If the origin server's response for /public/ can be manipulated (e.g., by another unkeyed input vulnerability) to contain malicious JavaScript, then this malicious content would be cached under /public/$payload.js. While this still creates a unique key, the ultimate goal of AWCP is to place malicious content under existing, popular keys. The full realization of "arbitrary web cache poisoning by modifying the keys of a store resource" would involve making the cache store content for / (the homepage) when the origin server processes a different, attacker-controlled path. This would require an even more sophisticated parser discrepancy or a chain of vulnerabilities.

Full Website Replacement

The ultimate goal articulated in the talk's agenda is to combine these techniques to achieve full website replacement. This implies:

  1. Arbitrary Content Storage: The ability to make the origin server return any content the attacker desires (e.g., a full malicious HTML page).
  2. Arbitrary Key Assignment: The ability to make the cache proxy store this arbitrary content under any desired cache key, including critical ones like / (homepage), /about, /contact, or specific CSS/JS files.

While the detailed methodology for achieving this "full replacement" was not present in the provided transcript, the preceding attack vectors lay the groundwork. By understanding how to trick the cache into storing specific content under attacker-influenced keys, an attacker moves closer to complete control. This could involve chaining multiple parser discrepancies, exploiting how different URL components (path, query, fragment) are handled, and potentially leveraging specific CDN configurations or backend frameworks that exhibit these parsing quirks. The ability to control both the what (content) and the where (cache key) fundamentally allows for a complete takeover of the cached representation of a website.

Demo / Proof of Concept

▶ Watch: Limitations of existing web cache deception attacks (6:00)

The provided transcript outlines the theoretical framework and attack methodologies, using examples based on common CDNs like Cloudflare, Cloudfront, and Akamai. The speaker frequently states that "all the examples are going to be based on CDNs like Cloud for Cloudfront Akamai," and discusses the behavior of "Cloudflare's Web Cache Deception Armor." This indicates that the techniques were likely demonstrated or tested against real-world CDN configurations.

However, the transcript does not provide a step-by-step description of a live demonstration or a specific Proof of Concept (PoC) in action. The talk focuses on the conceptual and technical explanation of the vulnerabilities and how different components (cache proxy, origin server) interpret URLs, leading to the attack vectors for arbitrary web cache deception and poisoning. It is plausible that a live demo was part of the presentation, but the details of its execution are not captured in the text provided.

Defensive Implications

▶ Watch: Introducing URL parser discrepancies as an attack vector (6:40)

The novel web cache exploitation techniques presented by Martin Doyhenard underscore critical areas where web application and infrastructure defenders must strengthen their security posture. While the transcript unfortunately cuts off before a detailed discussion of protections, we can infer several crucial defensive strategies based on the nature of the attacks described:

  1. Consistent URL Parsing Across the Stack: The most fundamental defense against these attacks is to ensure consistent URL parsing logic across all components of the web request chain. This includes the load balancer, CDN, WAF, reverse proxy, and the origin server's web framework or application. Any discrepancy in how delimiters are interpreted (e.g., $ as a delimiter by the backend but not the cache) creates an exploitable seam. Developers should:
  • Standardize Parsing Libraries: Use robust, well-vetted URL parsing libraries that adhere strictly to RFC standards and are consistent across the entire infrastructure.
  • Audit Custom Parsers: If custom URL parsing logic is implemented anywhere, it must be rigorously audited for consistency with standard parsers used by upstream components.
  • Normalize URLs: Implement URL normalization at the earliest possible stage (e.g., CDN or WAF) to convert ambiguous or non-standard URLs into a canonical form before they reach the backend. This could involve stripping unknown delimiters or rejecting malformed requests.
  1. Strict Cache Key Generation and Validation: CDNs and custom cache implementations should employ stricter rules for cache key generation.
  • Canonical URL Keys: Ensure that cache keys are derived from a canonical representation of the URL, ideally one that has undergone normalization.
  • Avoid Ambiguous Characters: Prohibit or strictly sanitize URLs containing non-standard or ambiguous delimiter characters (like $ or other less common symbols) from being used in cache keys.
  • Query String Handling: Be explicit about which query parameters are included in the cache key (Vary header) and which are ignored.
  1. Enhanced Cache Eligibility Rules:
  • Strengthen Content-Type vs. Extension Checks: While Cloudflare's Web Cache Deception Armor is a good start, it needs to be robust enough to detect cases where an attacker forces a static extension (e.g., .js) on a response with a different Content-Type (e.g., text/html). This might require deeper content inspection or a more sophisticated heuristic.
  • Dynamic Content Exclusion: Ensure that any URL path known to serve dynamic, user-specific, or sensitive content is explicitly configured not to be cached, regardless of apparent static extensions.
  • Whitelisting Cacheable Paths: Instead of blacklisting uncacheable paths, consider a whitelist approach where only explicitly defined static asset paths (e.g., /static/, /assets/*.css) are eligible for caching.
  1. Secure Backend Endpoint Mapping:
  • Explicit Path Handling: Backend frameworks should explicitly define how paths are mapped to endpoints and avoid implicit or overly flexible interpretations of URL segments as parameters.
  • Sanitize Path Parameters: If path parameters are used, they must be rigorously sanitized and validated to prevent unexpected behavior or injection.
  1. Monitoring and Alerting:
  • Anomalous Cache Hits: Monitor for unusually high cache hit rates on URLs that typically serve dynamic content or unique, suspicious URL patterns in cache logs.
  • Content Mismatch Alerts: Implement logging and alerting for Content-Type mismatches against expected file extensions, even if the cache decides to store the resource.
  1. Regular Security Audits: Conduct regular security audits of CDN configurations and web application logic, specifically focusing on how URLs are processed at each layer. Penetration testing should actively include attempts to exploit URL parsing discrepancies.

By implementing these comprehensive defenses, organizations can significantly reduce their exposure to advanced web cache exploitation techniques that leverage the subtle inconsistencies inherent in complex web architectures.

Key Takeaways

  • URL Parsing Discrepancies are a Critical Attack Vector: Differences in how web cache proxies (CDNs) and origin servers interpret URL delimiters (e.g., $) create exploitable vulnerabilities.
  • Arbitrary Web Cache Deception (AWCD) Bypasses Traditional Defenses: Attackers can now achieve web cache deception for any endpoint by manipulating URL paths with custom delimiters and static extensions, circumventing protections like Cloudflare's Web Cache Deception Armor.
  • Arbitrary Web Cache Poisoning (AWCP) Enables Key Modification: Beyond content manipulation, these techniques allow attackers to influence the cache key itself, enabling the poisoning of widely accessed resources under attacker-chosen keys.
  • Full Website Replacement is a Feasible Threat: By combining AWCD and AWCP, attackers could theoretically achieve complete control over a website's cached content, leading to defacement or widespread malicious content delivery.
  • Consistent URL Parsing is Paramount for Defense: The most effective defense involves ensuring uniform URL parsing logic across all layers of the web infrastructure, from CDN to origin server, and implementing strict validation and normalization of incoming URLs.

About the Speaker(s)

Martin Doyhenard is the speaker for "Gotta Cache ‘em all: bending the rules of web cache exploitation" at DEF CON 32. The provided transcript does not contain further biographical details about Martin Doyhenard beyond his name.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This presentation by Martin Doyhenard isn't just another cache talk; it's a deep dive into the fundamental inconsistencies in URL parsing across the modern web stack. By meticulously dissecting how CDNs and origin servers interpret RFCs differently, Doyhenard unveils novel techniques for arbitrary web cache deception and poisoning. This isn't about finding a niche misconfiguration; it's about exploiting architectural seams that bypass established defenses, making it a critical piece of research for anyone building or defending web infrastructure.

Heather Calloway (CISO) — STRONG ACCEPT

This presentation uncovers critical architectural vulnerabilities in how web caches and origin servers process URLs, leading to sophisticated attacks like arbitrary web cache deception and poisoning. It clearly articulates the business impact of these discrepancies, demonstrating how they can bypass existing CDN protections and potentially lead to full website replacement or mass data exfiltration. The talk provides actionable insights for CISOs and security architects on ensuring consistent URL parsing across the entire web stack and re-evaluating caching strategies, making it highly relevant for strategic decision-making and operational hardening.

→ Top-rated talks at DEF CON 32 Main Stage

All talks from DEF CON 32 Main Stage