SINBAD: Saliency-informed detection of breakage caused by ad blocking

Saiid El Hajj Chehade, Sandra Siby, Carmela Troncoso

IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 4

Overview

The proliferation of privacy-enhancing technologies (PETs) like ad blockers has dramatically improved user experience and privacy online. However, these tools often modify web page content and behavior, leading to unintended side effects known as "breakage." Breakage occurs when an ad blocker, in its effort to remove unwanted elements, inadvertently impairs the expected functionality of a web page, either partially or completely. This can manifest as missing content, non-interactive elements, or broken layouts, preventing users from properly engaging with the site. The talk "SINBAD: Saliency-informed detection of breakage caused by ad blocking" by Saiid El Hajj Chehade, Sandra Siby, and Carmela Troncoso introduces a novel, automated machine learning pipeline designed to proactively detect such breakage.

Watch on YouTube

Visual summary for SINBAD: Saliency-informed detection of breakage caused by ad blocking by Saiid El Hajj Chehade, Sandra Siby, Carmela Troncoso
Visual summary for SINBAD: Saliency-informed detection of breakage caused by ad blocking by Saiid El Hajj Chehade, Sandra Siby, Carmela Troncoso

Key moments

  1. 0:00 Introduction to breakage and its definition
  2. 1:06 Challenges with current ad blocker breakage cycle
  3. 2:45 Analysis of prior work and their limitations
  4. 4:30 Three main challenges for automatic breakage detection
  5. 5:30 Sinbad: Automated machine learning pipeline introduction
  6. 6:25 Overview of Sinbad's three-part detection process
  7. 6:50 Semantic segmentation and saliency classification details
  8. 7:55 Using VIPs algorithm for page segmentation

SINBAD: Saliency-informed detection of breakage caused by ad blocking

Speakers: Saiid El Hajj Chehade; Sandra Siby; Carmela Troncoso

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=F1M8LmtT5Bc

Overview

The proliferation of privacy-enhancing technologies (PETs) like ad blockers has dramatically improved user experience and privacy online. However, these tools often modify web page content and behavior, leading to unintended side effects known as "breakage." Breakage occurs when an ad blocker, in its effort to remove unwanted elements, inadvertently impairs the expected functionality of a web page, either partially or completely. This can manifest as missing content, non-interactive elements, or broken layouts, preventing users from properly engaging with the site. The talk "SINBAD: Saliency-informed detection of breakage caused by ad blocking" by Saiid El Hajj Chehade, Sandra Siby, and Carmela Troncoso introduces a novel, automated machine learning pipeline designed to proactively detect such breakage.

The core challenge addressed by SINBAD is the current reactive and labor-intensive process of identifying and fixing ad blocker-induced breakage. Users typically report issues, which maintainers then manually reproduce and resolve, a cycle that can take weeks and significantly degrade user experience. SINBAD aims to revolutionize this process by offering a proactive, reliable, and scalable solution. It achieves this by moving beyond simple page-level comparisons, instead focusing on granular, section-specific analysis, incorporating user interaction simulation, and modeling user interest through the concept of saliency. This innovative approach promises to significantly reduce the time and effort required to maintain effective ad blocking filter lists, ultimately fostering greater user adoption and trust in privacy tools.

This research is particularly significant because it tackles the often-overlooked cost of privacy on web functionality. By providing a robust mechanism for ad blocker developers to identify and mitigate breakage before it impacts a large user base, SINBAD helps ensure that privacy tools can be deployed more effectively without compromising the core utility of the web. The ability to automatically detect dynamic breakage, which requires user interaction to manifest, is a critical advancement, as prior automated methods were largely limited to static visual changes.

Background

▶ Watch: Introduction to breakage and its definition (0:00)

The current paradigm for addressing ad blocker breakage is largely reactive and manual. Users, employing the latest version of an ad blocker, frequently encounter broken websites. When this happens, they report the issues to ad blocker maintainers, often through forums like GitHub. Maintainers then attempt to reproduce the reported breakage, relying on user instructions which can sometimes be vague or incomplete. Once reproduced, fixes are developed and pushed to the filter lists, restarting the cycle. This process, while eventually effective, is notably slow, with some issues taking up to a month to be fully resolved. The speaker highlights that this delay negatively impacts user experience and can lead to reduced adoption of ad blockers. Furthermore, many issues involve significant back-and-forth communication between users and maintainers, resulting in numerous intermediary commits that do not fully resolve the breakage, complicating data curation efforts.

Prior attempts to automate breakage detection fall into three main categories, each with significant limitations:

  1. Manual Checks: Experts manually visit websites and perform interactions to identify breakage. While highly accurate, this approach is inherently unscalable and cannot cope with the vast number of websites and filter list updates.
  2. Heuristics: These techniques involve simple comparisons between web visits under different ad blocker versions. For instance, a drastic decrease in the number of images might heuristically indicate breakage. While simple and low-overhead, the research finds that heuristics suffer from high error rates. Crucially, they do not simulate user interactions, making them incapable of detecting dynamic breakage.
  3. Machine Learning (e.g., Page Graph by Smith et al.): This approach attempts to represent web visits as graph structures, extracting features from comparisons between pairs of visits to detect whether changes in ad blocker rules cause breakage. While offering better scalability and accuracy than heuristics, Page Graph still does not perform user interactions and is therefore limited to detecting static breakage. The SINBAD team also identified several critical flaws in Page Graph's methodology: it uses a heuristic to label its dataset, leading to high false positive rates; it includes many old, unreproducible breakage issues; and it treats each intermediary fixing commit as a separate breakage issue, further skewing its dataset. The SINBAD team also noted difficulties in reproducing Page Graph due to duplicated libraries, necessitating their best efforts to reimplement its features for comparison.

The SINBAD research identified three primary challenges that previous methods failed to adequately address:

  • Subjectivity of Breakage: What constitutes "breakage" can be subjective. A user might tolerate a broken comment section but not a broken news article they are reading. Prior methods struggled to model this user-centric perspective.
  • Dynamic Breakage: A significant 25% of reported breakages are dynamic, meaning they require at least one user interaction (e.g., clicking a button, scrolling) to manifest. Existing automated techniques, which typically rely on static page comparisons, completely miss these issues.
  • Reproducibility Challenges: Web pages are highly dynamic environments, frequently changing their layouts and functionalities. This makes reproducing historical breakage issues difficult. The study found that issues older than four months were rarely reproducible, severely limiting the potential size of training datasets.

SINBAD was developed specifically to overcome these limitations, providing a more robust, granular, and user-aware approach to automated breakage detection.

Key Findings

▶ Watch: Analysis of prior work and their limitations (2:45)

The SINBAD project yielded several critical findings that underscore the challenges of ad blocker maintenance and the effectiveness of their novel solution:

  • Slow Resolution Times: The manual, reactive process for resolving ad blocker breakage issues can take up to a month, significantly degrading user experience and potentially hindering the adoption of privacy-enhancing technologies.
  • Inefficient Resolution Cycles: Many breakage issues require extensive back-and-forth communication between users and maintainers, leading to numerous intermediary commits that do not fully resolve the problem. This highlights the inefficiency of the current system and the need for more proactive detection.
  • Prevalence of Dynamic Breakage: A substantial 25% of all reported breakage issues are dynamic, meaning they require user interaction to manifest. This critical insight reveals a significant blind spot in prior automated detection methods, which were limited to static page comparisons.
  • Limitations of Prior Work: Heuristic-based methods suffer from high error rates, while the most advanced machine learning technique, Page Graph, is limited to static breakage and relies on a flawed data labeling heuristic that introduces high false positive rates. The SINBAD team also found that many issues in Page Graph's dataset were old and unreproducible, further questioning its reliability.
  • Granular Detection is Key: SINBAD’s approach of classifying sections of a web page separately, rather than judging the full page, proved effective. This increased granularity allows for more precise identification of broken components, acknowledging that a page might have both legitimate blocks and broken sections simultaneously.
  • High Performance of SINBAD's Classifiers:
  • The in-house saliency classifier, trained on 543 annotated sites, achieved an 83% Area Under the Curve (AUC) score, demonstrating its effectiveness in modeling user interest areas.
  • The subtree classifier, trained on a dataset of 512 issues from AdGuard and other ad blockers, achieved a high 86% AUC score, reliably identifying broken sections of the page.
  • Reliable Page-Level Prediction: SINBAD's page-level heuristic, which predicts a page as broken if at least one broken subtree is detected, proved reliable.
  • Superior Performance over Baselines: SINBAD significantly reduces false positive rates compared to both heuristic-based methods and the Page Graph approach. This robust performance was consistent across various evaluation datasets, demonstrating its strong generalization capabilities.
  • Importance of Saliency and Interactions: Features derived from saliency analysis and simulated user interactions ranked high in importance for SINBAD's performance, validating the core design choices of the pipeline, especially its ability to detect dynamic breakages (e.g., issues with video players after interaction).
  • Data Curation Challenges: The study reiterated the inherent difficulty in curating reliable datasets for breakage detection due to the dynamic nature of web pages and the short lifespan of reproducible issues (issues older than 4 months were rarely reproducible).

Technical Deep Dive

▶ Watch: Sinbad: Automated machine learning pipeline introduction (5:30)

SINBAD is designed as an automated machine learning pipeline that addresses the limitations of prior work by focusing on increased granularity, user subjectivity, and dynamic interactions. Its core methodology can be broken down into a three-part process: crawling, feature extraction, and classification.

The pipeline begins by increasing the granularity of breakage detection. Instead of classifying an entire web page as broken or not, SINBAD classifies individual sections of the web page. This is motivated by the observation that an ad blocker might legitimately block an ad in one section while inadvertently breaking another section of the same page. To approximate the subjectivity of user experience, SINBAD introduces the concept of saliency, which attempts to model areas of a web page where a user is most likely to be interested or interact. Crucially, SINBAD also performs a set of interactions with these salient areas to uncover dynamic breakage. Finally, the system relies on a meticulously manually labeled dataset to ensure a sound ground truth for training and evaluation, avoiding the pitfalls of heuristically labeled data.

1. Crawling

The crawling phase involves visiting the target web page multiple times under different conditions to establish a baseline and identify changes:

  • Fixed Filter List: The page is visited with the ad blocker's filter list in a "fixed" state (post-breakage resolution).
  • Broken Filter List: The page is visited with the ad blocker's filter list in a "broken" state (pre-breakage resolution).
  • No Filter List: The page is visited without any ad blocker, serving as a clean reference.

For each scroll of the page, SINBAD performs two critical steps:

Semantic Segmentation

The web page is semantically segmented, dividing it into meaningful, distinct blocks. These blocks represent logical components like login forms, navigation menus, or article content, rather than arbitrary HTML elements. The choice of segmentation algorithm is crucial for efficiency and accuracy. SINBAD employs VIPs (Vision-based Page Segmentation), an algorithm developed by Microsoft. VIPs is a top-down page division algorithm that leverages both HTML structure and style/position features of elements. It was selected for its efficiency, scalability, and performance comparable to more demanding machine learning techniques, making it suitable for SINBAD's curated dataset. The hyperparameters of VIPs can be adjusted to control the granularity of the segmented parts.

Saliency Classification

After segmentation, each block is classified as either a salient block or a non-salient block. This is achieved using an in-house classifier trained on a custom dataset of 543 websites. These sites were meticulously annotated by at least two volunteers, with only the positive intersection of their annotations being used to account for moderate agreement and the subjective nature of saliency. This classifier achieves an 83% AUC score under cross-validation, indicating its effectiveness in identifying areas of user interest.

2. Interactions and Representation

Following saliency classification, SINBAD proceeds to interact with the identified salient candidate blocks. These interactions are designed to mimic typical user behaviors (e.g., clicks, scrolls) that might trigger dynamic breakage. During and after these interactions, the system logs various outputs:

  • Network Requests: All outgoing requests are captured.
  • Script Activity: JavaScript execution and errors are recorded.
  • DOM Representation: The entire web page structure is encoded into its DOM (Document Object Model) representation, which encapsulates the hierarchical relationships between elements. This DOM representation is then augmented with information about the interactions performed, the network requests made, and any errors encountered.

3. Feature Extraction and Classification

The core of SINBAD's detection mechanism lies in comparing the augmented DOM trees from pairs of visits (e.g., broken vs. fixed filter list). This comparison identifies edit subtrees—groups of nodes that have changed between the two states. These changes can involve elements being edited, removed, or added.

For training purposes, each identified edit subtree is manually labeled as either a legitimate change (e.g., a blocked ad), a breaking change (e.g., essential content removed), or a neutral change. This labeling is based on the origin and type of change; for example, a subtree removed when transitioning from a breaking list to a fixed list might be labeled as a legitimate edit if it corresponds to an ad.

Finally, SINBAD extracts a rich set of features from these edit subtrees and feeds them into a simple machine learning model to classify whether a subtree is broken or not. These features fall into four categories:

  • Visual Features: These capture the visual impact of changes, incorporating the contribution from saliency to weigh changes in important visual areas more heavily.
  • Functional Features: These relate to the functionality of the page, including the contribution of the simulated interactions (e.g., whether an interaction failed or succeeded).
  • Structural Features: These describe changes in the DOM structure itself.
  • Global Features: These are aggregate features from other subtrees within the same web page visit, providing contextual information to the classifier.

Data Curation and Evaluation

A significant bottleneck for all breakage studies is data availability and curation. SINBAD's team meticulously curated their dataset by filtering issues that:

  1. Did not have a reachable, live URL where the breakage was first reported.
  2. Did not include the active filter list used by the user, which is crucial for reconstructing the exact state of the ad blocker.

They then manually attempted to reproduce each of the 2,000 initial issues by loading the ad blocker before and after the reported fix. A critical finding was that issues older than four months were rarely reproducible due to drastic changes in web page layouts and functionalities. This led to a final main training dataset comprising 512 issues from AdGuard, a mainstream ad blocker, supplemented with evaluation datasets from other ad blockers.

The subtree classifier achieved a high 86% AUC score under a five-fold cross-validation evaluation. For the page-level classifier, SINBAD uses a simple heuristic: if at least one broken subtree is detected, the entire page is reliably predicted as broken. Comparative analysis showed that SINBAD achieves significantly lower false positive rates compared to both existing heuristics and the Page Graph approach. Furthermore, the importance of saliency and interactions was validated, as features derived from them ranked highly in contributing to SINBAD's performance.

Demo / Proof of Concept

▶ Watch: Overview of Sinbad's three-part detection process (6:25)

While the talk does not describe a dedicated, live "demo" in the traditional sense of a separate presentation segment, the research inherently includes a strong proof of concept embedded within its evaluation and error analysis. The speaker explicitly mentions that SINBAD was capable of "manually verify[ing] that SINBAD detects many cases where it performs interactions with for example video players and it can try to play the videos but it will trigger many various Dynamic breakages." This statement serves as a compelling testament to SINBAD's ability to identify the very dynamic breakage types that prior automated methods missed.

This "demonstration" highlights SINBAD's core strength: its capacity to simulate user interactions on salient page elements and then detect the functional failures that result. For instance, if SINBAD interacts with a video player (a salient element) and attempts to play a video, but the video fails to load or play correctly due to ad blocker rules, this constitutes a dynamic breakage. The system's ability to log script activity, network requests, and DOM changes during these interactions, and subsequently identify a "breaking change" in the relevant subtree, effectively proves its concept for dynamic breakage detection. The rigorous evaluation and error analysis further substantiate these claims, showing SINBAD's consistent performance and reduced false positive rates across diverse datasets, implicitly demonstrating its robust and practical applicability.

Defensive Implications

▶ Watch: Using VIPs algorithm for page segmentation (7:55)

The SINBAD research offers profound defensive implications for various stakeholders in the web ecosystem, particularly for ad blocker maintainers, web developers, and users.

For ad blocker maintainers, SINBAD represents a paradigm shift from reactive to proactive breakage detection. Instead of waiting for users to report issues and then manually reproducing them, maintainers can integrate SINBAD into their continuous integration (CI) pipelines. This allows for automated testing of new filter list rules against a diverse set of websites, identifying potential breakage before it is deployed to users. This proactive approach leads to:

  • Faster Resolution: Breakage can be identified and fixed much quicker, potentially within hours or days rather than weeks or months.
  • Improved User Experience: By minimizing the occurrence of breakage, ad blockers become more reliable, leading to higher user satisfaction and retention.
  • Reduced Manual Burden: Automation significantly reduces the labor-intensive task of manually reproducing and triaging breakage reports, freeing up maintainer resources for other critical development tasks.
  • Enhanced Filter List Quality: The ability to detect dynamic breakage and understand its granular impact allows for the development of more precise and effective filter rules, striking a better balance between ad blocking and site functionality.
  • Modularity for Future Improvement: SINBAD's modular pipeline design means that individual components, such as the saliency detector, semantic segmentation, or interaction models, can be continuously improved and integrated, ensuring the system remains state-of-the-art.

For web developers and website owners, SINBAD provides valuable insights into how their sites might be inadvertently broken by privacy-enhancing technologies. While not a direct tool for them, understanding the mechanisms SINBAD uses (saliency, interactions, semantic segmentation) can inform more robust web design practices. Developers can:

  • Design for Resilience: Avoid over-reliance on third-party scripts or elements that are frequently blocked by ad blockers for core site functionality.
  • Improve Accessibility: Ensure that critical content and interactive elements are clearly distinct and not easily confused with advertising or tracking components by automated systems.
  • Test with Ad Blockers: Proactively test their websites with popular ad blockers to identify potential breakage early in the development cycle, rather than waiting for user reports.

For users, the long-term implication is a more reliable and less frustrating web browsing experience when using ad blockers. Fewer broken websites mean less need to disable ad blockers, thereby enhancing their online privacy and security without sacrificing usability.

Overall, SINBAD contributes to a healthier web ecosystem where privacy and functionality can coexist more harmoniously. By automating and improving breakage detection, it strengthens the capabilities of privacy tools and encourages more thoughtful web development.

Key Takeaways

  • Breakage, defined as the inability of a user to access expected web page functionality, is a significant and persistent challenge for privacy-enhancing technologies like ad blockers.
  • A substantial portion (25%) of reported breakages are dynamic, requiring user interaction to manifest, a critical blind spot for prior automated detection methods.
  • SINBAD is the first automated system capable of reliably detecting dynamic breakage by simulating user interactions and modeling user interest through saliency analysis.
  • By employing granular, section-level classification and a meticulously curated dataset, SINBAD significantly outperforms prior heuristic and machine learning approaches, drastically reducing false positive rates.
  • The pipeline's modular design, incorporating semantic segmentation (e.g., using VIPs) and an in-house saliency classifier, allows for continuous improvement and adaptation to evolving web technologies.
  • Effective data curation for breakage detection is severely limited by the dynamic nature of web pages, with issues older than four months rarely being reproducible, highlighting a persistent challenge for research in this domain.

About the Speaker(s)

The talk "SINBAD: Saliency-informed detection of breakage caused by ad blocking" was presented by Saiid El Hajj Chehade, and co-authored by Sandra Siby and Carmela Troncoso. The presentation style suggests Saiid El Hajj Chehade is a researcher, likely a Ph.D. student or post-doctoral researcher, given the detailed technical depth and focus on novel methodologies. The collaboration with Sandra Siby and Carmela Troncoso indicates a research team, typically found within academic institutions or research labs, engaged in advanced studies in privacy, security, and web technologies. Their work presented at IEEE S&P, a premier conference in security and privacy, underscores their expertise and contribution to the field. While specific titles and affiliations beyond their names are not detailed in the transcript, their research output clearly positions them as leading contributors to the understanding and mitigation of challenges faced by privacy-enhancing technologies on the web.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

SINBAD presents a critical advancement in ad blocker efficacy, tackling the long-standing problem of dynamic breakage that plagues privacy tools. By integrating saliency modeling and simulated user interactions, this novel ML pipeline proactively detects functional impairments, a blind spot for all prior automated methods. This work directly improves user experience and significantly reduces the manual burden on ad blocker maintainers, making it a must-see for anyone in web privacy or automated testing.

Heather Calloway (CISO) — STRONG ACCEPT

This research offers a credible and actionable solution to a critical problem in the privacy tooling ecosystem. It directly enables ad blocker maintainers to proactively detect and mitigate breakage, enhancing user experience and fostering trust in essential privacy technologies. The work advances a specific, but important, dimension of operational security.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024