To Auth or Not To Auth? A Comparative Analysis of the Pre- and Post-Login Security Landscape
Jannis Rautenstrauch, Metodi Mitkov, Thomas Helbrecht, Lorenz Hetterich, Ben Stock
IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 6
Overview
Traditional web security research often relies on automated crawlers that interact with websites as unauthenticated users, starting each session from a fresh browser state. This approach, while efficient for broad surveys, fundamentally overlooks the experience of an authenticated user – someone logged into their account, interacting with a personalized portal, or accessing application-specific functionalities. This talk, "To Auth or Not To Auth? A Comparative Analysis of the Pre- and Post-Login Security Landscape," presented by Jannis Rautenstrauch and his colleagues Metodi Mitkov, Thomas Helbrecht, Lorenz Hetterich, and Ben Stock at IEEE S&P, directly addresses this critical blind spot.

Key moments
- 2:00 Posing the core research question
- 2:50 Overview of four security topics investigated
- 4:50 Detailed workflow for account registration
- 6:00 Automated login and session management process
- 7:00 How unauthenticated vs. authenticated crawling is performed
- 8:00 Initial finding: number of URLs differs between states
To Auth or Not To Auth? A Comparative Analysis of the Pre- and Post-Login Security Landscape
Speakers: Jannis Rautenstrauch; Metodi Mitkov; Thomas Helbrecht; Lorenz Hetterich; Ben Stock
Conference: IEEE S&P
YouTube: https://www.youtube.com/watch?v=kDWaF1Nw4NM
Overview
Traditional web security research often relies on automated crawlers that interact with websites as unauthenticated users, starting each session from a fresh browser state. This approach, while efficient for broad surveys, fundamentally overlooks the experience of an authenticated user – someone logged into their account, interacting with a personalized portal, or accessing application-specific functionalities. This talk, "To Auth or Not To Auth? A Comparative Analysis of the Pre- and Post-Login Security Landscape," presented by Jannis Rautenstrauch and his colleagues Metodi Mitkov, Thomas Helbrecht, Lorenz Hetterich, and Ben Stock at IEEE S&P, directly addresses this critical blind spot.
The core premise of their work is to rigorously compare the security posture of popular websites from both an unauthenticated ("guest") and an authenticated ("logged-in") perspective. The research team developed a sophisticated methodology to register accounts, log in, and then crawl websites in parallel under both states, evaluating four key security aspects: client-side cross-site scripting (XSS), security headers, JavaScript inclusions, and post-message handlers. Their findings reveal significant differences in the observed security landscape, challenging the completeness of prior, unauthenticated-only web security measurements and providing crucial insights for both web defenders and security researchers.
This work matters because it highlights a fundamental flaw in how web security prevalence is often assessed. By demonstrating that the attack surface and vulnerability landscape can change dramatically post-login, the research compels a re-evaluation of current security testing practices and encourages a more holistic view of web application security. It underscores that understanding the full user journey—from guest to authenticated user—is paramount for accurately identifying and mitigating real-world risks.
Background
▶ Watch: Posing the core research question (2:00)
The landscape of web security research has long been shaped by the capabilities of automated crawlers. Tools and studies, often drawing from top website lists like Tranco, have scoured the internet to estimate the prevalence of various security issues. A common thread across these diverse crawlers, developed by different authors and utilizing various technologies, is their operational modus operandi: they visit websites using a fresh browser session for each site. This "guest" or "logged-out" perspective, while practical for large-scale data collection, often presents a limited and potentially misleading view of a website's true attack surface.
From an unauthenticated perspective, a crawler might encounter a generic landing page, a blog, or a signup/login form. However, a real user, particularly one who is logged in, experiences a vastly different environment. They might see a personalized social media feed, a location-based application with historical data, or a user portal for managing uploaded files. These authenticated views expose significantly more functionality and interactive elements compared to their unauthenticated counterparts. This anecdotal understanding led the researchers to a pivotal question: Do unauthenticated crawls accurately reflect reality, or do prior studies misrepresent the actual prevalence of security issues on the web?
Two opposing hypotheses could explain how security might differ between logged-in and logged-out states. On one hand, developers might prioritize the security of their actual users, especially those paying for services, leading to better security in the authenticated state. This could manifest as more robust input validation, stricter content security policies, or more diligent patching. On the other hand, the authenticated state often exposes a substantially larger attack surface due to the wealth of additional functionalities available only after login. More features inherently mean more code, more user input fields, more integrations, and thus, more potential points of failure and vulnerabilities. This expanded attack surface could lead to worse security in the authenticated state simply due to the increased complexity and potential for oversight. Understanding which of these scenarios prevails, or if a nuanced combination exists, was the core motivation behind this comparative security measurement study.
Key Findings
▶ Watch: Detailed workflow for account registration (4:50)
The study yielded a set of nuanced and significant findings, emphasizing that the impact of login on observed security posture is highly dependent on the specific security aspect being investigated. The authenticated state is neither universally "better" nor "worse" but presents a distinct security landscape that traditional unauthenticated crawls often miss.
For client-side cross-site scripting (XSS), the researchers found that while a massive number of DOM sink invocations (over 38 million) were detected across both states, only a fraction (around 3,000) resulted in working exploits on a total of seven sites. Interestingly, six sites were found vulnerable in the authenticated state, and four in the non-authenticated state. Critically, one site was vulnerable only when unauthenticated, and not when logged in. Further analysis revealed a significant asymmetry:
- 91% of exploits found in the non-authenticated state were also exploitable in the authenticated state.
- However, only 66% of exploits found in the authenticated state were exploitable in the non-authenticated state. This implies that over 30% of XSS vulnerabilities present in the logged-in user experience would be completely missed by unauthenticated crawlers. This highlights the necessity of examining the authenticated state to uncover a substantial portion of XSS risks.
Regarding security headers (such as X-Frame-Options or Content-Security-Policy), the study found no significant difference in their usage or consistency between the two states. While a few sites might have used a header exclusively in one state, the overall prevalence and security posture of these headers were remarkably similar. The researchers hypothesize that security headers are often set at the reverse proxy level, external to the application logic itself, meaning they are applied universally regardless of the user's authentication status. This suggests that for studies focused solely on security header deployment, an unauthenticated crawl might be sufficient.
The analysis of JavaScript inclusions presented a different picture. While the total number of unique scripts was slightly higher in the authenticated state, a striking difference emerged when dissecting script types. The authenticated state consistently showed many more third-party scripts and a higher number of unique third parties. Conversely, some third parties were only seen in the non-authenticated state, suggesting they are removed or hidden once a user logs in. This pattern also extended to trackers, with a significantly higher number of tracking entities (identified via the Disconnect.me list) observed in the authenticated state. The researchers noted that their crawlers operated from the EU, where GDPR is in effect, and cookies were generally accepted in the authenticated state but not in the unauthenticated state, which likely contributed to the increased tracking observed post-login.
Finally, for post-message handlers, the authenticated state consistently detected many more handlers than the unauthenticated state, regardless of their frequency. While most handlers occurred only once, some, often originating from third parties like Google Analytics, appeared frequently across different websites or sub-pages. This increase in post-message handler exposure in the authenticated context represents a larger attack surface for vulnerabilities related to cross-origin communication.
In summary, the research unequivocally demonstrates that the login state is a critical factor influencing the observed security landscape. While security headers remain largely consistent, the attack surface related to dynamic content, third-party integrations, and client-side vulnerabilities like XSS and post-message issues expands significantly and uniquely in the authenticated environment.
Technical Deep Dive
▶ Watch: Automated login and session management process (6:00)
To execute their comparative analysis, the research team developed a sophisticated multi-stage workflow designed to handle the complexities of account registration, automated login, and parallel crawling. The methodology ensured that experiments could be conducted consistently across both authenticated and unauthenticated states for a substantial set of real-world websites.
The complete workflow comprised four main steps:
- Form Detector: This initial tool was given a target website and tasked with automatically identifying the login and registration pages. Its output, the URLs for these forms, was crucial for the subsequent steps. If no such forms were found, the website was excluded from the registration process.
- Registration Step: This was the most labor-intensive part of the process, requiring manual assistance to overcome the inherent challenges of automated account creation.
- The tool would open the identified registration page.
- The Bitwarden browser extension was utilized to automatically prefill standard fields like email and password.
- However, human intervention was often necessary to fill in additional data, bypass CAPTCHAs, and crucially, to monitor an email address for verification links. Many websites require users to click a link in a confirmation email to activate their account.
- Ultimately, the research team manually verified the successful creation of an account and reported this back to an internal account framework. This process allowed them to create accounts on over 200 popular sites sourced from the Crooks top 5K list.
- Automated Login: Once an account was successfully registered, a dedicated tool was employed for automated login.
- It received the login URL and the newly created account credentials.
- The tool automatically filled out the login form and submitted it.
- A critical component was its login Oracle, which determined if the login was truly successful. This oracle accounted for common failure scenarios like bot detection, account blocking, CAPTCHAs post-submission, or other errors.
- Upon successful login, the tool stored the session data, including cookies and local storage, within the account framework. This session data was vital for emulating the authenticated user experience in subsequent crawling.
- Experiment Execution: This final step launched the actual comparative crawls.
- Experiments would request a fresh, verified session from the account framework.
- Two crawlers were then initiated in parallel for each website:
- The unauthenticated crawler operated like standard prior work, opening a fresh browser session with no pre-existing data.
- The authenticated crawler also opened a new browser session, but critically, it injected all the stored local storage and session cookies obtained during the automated login step. This accurately emulated a logged-in user.
- Both crawlers then proceeded to visit up to 1,000 same-site sub-pages or continued for up to 24 hours per website.
Experimental Settings and Tools:
- Crawler Engine: The crawlers utilized Playwright in version 1.33.
- Browsers:
- For the cross-site scripting (XSS) experiments, the specialized Foxhound browser (a taint-tracking enabled fork of Firefox 109) was used to gain access to taint flows.
- For all other experiments (security headers, JavaScript, post messages), Chromium in version 113 was employed.
- Target Websites: The study focused on over 200 popular websites from the Crooks top 5K list.
- URL Collection: On average, 840 URLs were collected in the non-authenticated state, compared to 820 URLs in the authenticated state. The slight reduction in authenticated URLs was attributed to users often being redirected to specific user portals, which, while functional, might contain fewer navigable links to other sub-pages compared to a public-facing site.
Specific Security Topic Tools:
- Client-side Cross-site Scripting (XSS): The methodology leveraged Foxhound's taint tracking capabilities to identify data flows from user-controllable sources to sensitive DOM sinks. The exploit generator from Steffin et al. (NDSS 2019) was then used to confirm exploitability.
- Security Headers: Standard HTTP header analysis was performed to detect the presence and configuration of headers like
X-Frame-OptionsandContent-Security-Policy (CSP). The consistency and security posture (e.g., strong vs. weak CSPs) were also evaluated. - JavaScript Inclusions: The crawlers monitored and recorded all loaded JavaScript files, categorizing them as inline or included, and distinguishing between first-party and third-party scripts. The Disconnect.me list was used to identify known tracking entities among third-party inclusions.
- Post Messages: The PMForce tool from Steffin and Do (CCS 2020) was integrated to detect and analyze vulnerable post-message handlers, which are critical for secure cross-origin communication in web applications.
This meticulously designed technical framework allowed the researchers to systematically explore the differences in the web's security landscape between authenticated and unauthenticated users, providing a robust foundation for their comparative analysis.
Demo / Proof of Concept
▶ Watch: How unauthenticated vs. authenticated crawling is performed (7:00)
While the talk didn't feature a live, interactive demonstration of an exploit or a tool in action, the entire methodology presented serves as a comprehensive proof of concept for performing comparative web security analysis. The detailed workflow, from automated form detection and manually assisted account registration to parallel crawling with session injection, effectively demonstrates the feasibility and necessity of studying both authenticated and unauthenticated states.
The "demonstration" of their approach lies in the systematic execution of their measurement study. They successfully built and integrated several components: a Form Detector to identify entry points, a Registration Step that combined automation (Bitwarden) with human oversight to navigate CAPTCHAs and email verifications, an Automated Login tool with a robust login Oracle, and finally, parallel crawlers using Playwright and specialized browsers like Foxhound to conduct the actual security assessments. The ability to create and manage over 200 authenticated sessions on popular websites, and then reliably crawl them in both states, stands as a testament to the practical viability of their proposed methodology. The results themselves, particularly the discovery of XSS vulnerabilities and increased third-party script usage only in the authenticated context, serve as compelling evidence that their comparative approach uncovers security insights missed by traditional methods.
Defensive Implications
▶ Watch: Initial finding: number of URLs differs between states (8:00)
The findings from this comparative analysis carry significant implications for web developers, security engineers, and organizations responsible for protecting web applications. The core message is clear: relying solely on unauthenticated security assessments is insufficient and can lead to a false sense of security.
- Comprehensive Security Testing: Organizations must extend their security testing to explicitly include the authenticated user experience. This means security teams should incorporate logged-in crawling and user-specific attack scenarios into their vulnerability assessments, penetration tests, and continuous monitoring efforts. Automated scanners configured only for guest access will miss a substantial portion of the attack surface and unique vulnerabilities present post-login, particularly for client-side XSS, JavaScript-related issues, and post-message handlers.
- Threat Modeling for Authenticated Users: Developers should specifically threat model the functionalities and data exposed to authenticated users. The increased complexity, personalized content, and integration with more third-party services in the logged-in state introduce unique risks. Input validation, output encoding, and access control mechanisms must be rigorously applied and tested for all features accessible after login.
- Security Header Consistency: While the study found security headers to be consistent across states, this is likely due to their implementation at the reverse proxy level. This reinforces the best practice of configuring security headers (like
X-Frame-Options,Strict-Transport-Security,Content-Security-Policy) at the web server or CDN/proxy layer, rather than within the application code itself. This ensures uniform application across all user journeys, authenticated or not. However, developers should still ensure their CSPs are robust and effective, as a loosely defined CSP provides little real security.
- Managing Third-Party Integrations: The significant increase in third-party scripts and trackers in the authenticated state presents a heightened supply chain risk. Defenders must meticulously vet all third-party integrations, especially those loaded dynamically after login. This includes understanding their data access, potential vulnerabilities, and compliance with privacy regulations (e.g., GDPR). Content Security Policies should be designed to restrict unnecessary third-party domains.
- Post-Message Handler Vigilance: The higher number of post-message handlers in the authenticated state indicates a larger surface for cross-origin communication vulnerabilities. Developers need to implement robust origin checks and message validation for all
postMessagelisteners to prevent information leakage, unauthorized actions, or XSS via untrusted messages.
- Privacy Considerations: The observed increase in trackers when authenticated (and cookies accepted) underscores the importance of transparent cookie consent mechanisms and adherence to privacy regulations like GDPR. Organizations must ensure that user consent is properly managed and that tracking only occurs when explicitly permitted, with clear information provided to users about what data is collected and by whom.
- Guidance for Security Researchers: For security researchers, the study provides a clear directive: future web security measurements should consider the impact of authentication. A recommended approach is to first perform a small-scale comparative crawl to gauge how much login impacts the specific research question. If significant differences are found, then a more extensive, potentially manual, authenticated crawl is warranted to avoid skewed results and discover a more complete set of vulnerabilities. The open-sourced tools developed by the team can facilitate such comparative studies.
By integrating these defensive strategies, organizations can move towards a more holistic and accurate understanding of their web application's security posture, better protecting their users and data across the entire user experience.
Key Takeaways
- Traditional unauthenticated web security crawls often provide an incomplete and potentially misleading view of a website's true security posture.
- The impact of user login on observed web security is highly dependent on the specific vulnerability class or security aspect being investigated.
- For client-side XSS, the authenticated state reveals a larger set of exploitable vulnerabilities, with over 30% of authenticated exploits being undetectable or unexploitable in the unauthenticated state.
- Security headers (e.g.,
X-Frame-Options,CSP) tend to remain consistent across authenticated and unauthenticated states, likely due to their implementation at the reverse proxy level. - The authenticated user experience generally exposes a larger attack surface due to significantly more JavaScript (especially third-party scripts) and post-message handlers, which can introduce new security and privacy risks.
- Security researchers should consider performing a small-scale comparative analysis (authenticated vs. unauthenticated) for their specific research questions before undertaking large-scale, unauthenticated-only studies, to ensure comprehensive and accurate results.
- The study's methodology, including the open-sourced tools, provides a robust framework for future comparative web security measurements, encouraging a more nuanced and complete understanding of web application security.
About the Speaker(s)
The talk "To Auth or Not To Auth? A Comparative Analysis of the Pre- and Post-Login Security Landscape" was a collaborative effort presented by Jannis Rautenstrauch, alongside his colleagues Metodi Mitkov, Thomas Helbrecht, Lorenz Hetterich, and Ben Stock. While specific titles and affiliations beyond "colleagues" and "research team" were not detailed in the transcript, their joint work, presented at the prestigious IEEE S&P conference, highlights their expertise in web security research and measurement studies. Their focus on the often-overlooked authenticated user experience demonstrates a commitment to advancing the field's understanding of real-world web application vulnerabilities.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This research meticulously dismantles the flawed premise of unauthenticated-only web security assessments, proving that a significant portion of the attack surface, particularly client-side XSS and third-party integrations, is only visible post-login. The methodology is robust, the findings are concrete, and the implications for both defenders and researchers are profound. This isn't just a paper; it's a re-calibration of how we measure web security.
Heather Calloway (CISO) — STRONG ACCEPT
This research uncovers a critical blind spot in web security assessment, demonstrating that traditional unauthenticated scans fail to capture a significant portion of the attack surface and unique vulnerabilities exposed post-login. It provides clear, actionable guidance for security leaders to refine their testing strategies and properly account for authenticated user risk.
→ Top-rated talks at IEEE Symposium on Security and Privacy 2024