Targeted and Troublesome: Tracking and Advertising on Children's Websites

Zahra Moti, Asuman Senol, Hamid Bostani, Frederik Zuiderveen Borgesius, Veelasha Moonsamy, Arunesh Mathur

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 4

Overview

This talk, presented by Zahra Moti from Radboud University, delves into the pervasive and often problematic landscape of online tracking and targeted advertising on child-directed websites. The research, a collaborative effort with colleagues from Radboud University and other institutions, highlights how digital platforms exploit user behavior, creating detailed profiles for monetization, a practice particularly concerning when the users are children. The study not only investigates the prevalence of tracking mechanisms but also uncovers the alarming frequency of inappropriate and disturbing advertisements displayed to this vulnerable demographic.

Watch on YouTube

Visual summary for Targeted and Troublesome: Tracking and Advertising on Children's Websites by Zahra Moti, Asuman Senol, Hamid Bostani, Frederik Zuiderveen Borgesius, Veelasha Moonsamy, Arunesh Mathur
Visual summary for Targeted and Troublesome: Tracking and Advertising on Children's Websites by Zahra Moti, Asuman Senol, Hamid Bostani, Frederik Zuiderveen Borgesius, Veelasha Moonsamy, Arunesh Mathur

Key moments

  1. 0:00 Introduction to tracking on children's websites and regulations
  2. 2:15 Specific categories of harmful ads and tracking investigated
  3. 3:15 Methodology: Building a children's website dataset
  4. 4:45 Methodology: Advanced crawler for ad and tracking data
  5. 7:00 Key Finding: High prevalence of targeted advertising on children's sites
  6. 8:00 Key Finding: Malicious ad links and popularity's effect on targeting
  7. 9:00 Key Finding: Methods for detecting improper ad content

Targeted and Troublesome: Tracking and Advertising on Children's Websites

Speakers: Zahra Moti; Asuman Senol; Hamid Bostani; Frederik Zuiderveen Borgesius; Veelasha Moonsamy; Arunesh Mathur

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=cKcaeRi7Lrc

Overview

This talk, presented by Zahra Moti from Radboud University, delves into the pervasive and often problematic landscape of online tracking and targeted advertising on child-directed websites. The research, a collaborative effort with colleagues from Radboud University and other institutions, highlights how digital platforms exploit user behavior, creating detailed profiles for monetization, a practice particularly concerning when the users are children. The study not only investigates the prevalence of tracking mechanisms but also uncovers the alarming frequency of inappropriate and disturbing advertisements displayed to this vulnerable demographic.

The significance of this research cannot be overstated. Children are uniquely susceptible to the manipulative nature of advertising, often misunderstanding its intent and focusing on negative or harmful aspects, such as gambling or unhealthy food promotions. Despite the critical nature of this issue, large-scale empirical studies specifically focused on child-directed websites have been scarce. This work bridges that gap, providing a comprehensive measurement of advertising practices, including the detection of targeted and improper ads, as well as the identification of third-party trackers and fingerprinting attempts.

The timing of this study is particularly pertinent given the recent and upcoming regulatory changes aimed at protecting children online. Legislation like the EU's Digital Services Act (DSA), which prohibits targeting children with ads based on profiling, and the US Children's Online Privacy Protection Act (COPPA), requiring verifiable parental consent for data collection, underscore the urgent need for empirical data to assess the current state of compliance and enforcement. The findings reveal a widespread disregard for these protections, necessitating a call for stronger regulation and enforcement to safeguard children's online privacy and well-being.

Background

▶ Watch: Introduction to tracking on children's websites and regulations (0:00)

The modern web is characterized by an extensive ecosystem of online tracking and targeted advertising, where user behavior is routinely exploited to build detailed profiles for commercial gain. While this practice raises significant privacy concerns for users of all ages, children represent a particularly vulnerable population. Prior research consistently demonstrates that children frequently misunderstand the persuasive intent of advertisements, often interpreting them as factual information rather than commercial messaging. Furthermore, they are more prone to fixate on the negative or inappropriate aspects of certain ad categories, such as those promoting gambling, unhealthy food, or disturbing content.

Despite the heightened risks, there has been a notable absence of large-scale, empirical studies specifically focused on the advertising and tracking practices prevalent on websites designed for children. This gap in research has left a critical blind spot in understanding the true scope of the problem. Compounding this issue, governmental bodies worldwide are increasingly recognizing the need for robust protections. In the European Union, the DSA explicitly forbids profiling-based ad targeting for minors. Similarly, in the United States, COPPA mandates verifiable parental consent for data collection intended for targeted advertising on child-directed sites. These legislative efforts highlight the urgency of addressing the pervasive nature of exploitative online practices affecting children.

The research presented in this talk specifically aimed to investigate the state of targeted and personalized advertising shown to children, focusing on four categories identified by prior research and regulatory reports as particularly harmful: dating, mental health, weight loss, and racy content. The study also sought to detect the presence of third-party trackers and fingerprinting attempts. Key challenges in conducting this research included the lack of a comprehensive, up-to-date list of child-directed websites suitable for large-scale study, and the inherent difficulty in automatically detecting targeted and improper advertisements, which required sophisticated methods for scraping ads and their associated disclosure pages.

Key Findings

▶ Watch: Methodology: Building a children's website dataset (3:15)

The empirical measurement research conducted by Moti and her team yielded a series of concerning and critical findings regarding tracking and advertising practices on child-directed websites.

Firstly, their crawler successfully scraped over 70,000 advertisements from 804 distinct children's websites. A significant proportion of these sites, 36%, were found to display advertisements. More alarmingly, 27% of these websites contained targeted advertisements, a practice explicitly restricted or banned by regulations like the EU DSA and US COPPA. Delving deeper, the study revealed that over 70% of the individual advertisements captured were targeted, meaning they enabled ad targeting or ad personalization. This indicates that non-targeted ads are a minority even on websites intended for children, underscoring a widespread disregard for protective policies.

An exploratory analysis into the safety of ad links revealed potential malicious activity. From a sample of 4,000 scanned links, 150 were flagged as malicious or phishing by at least one scan engine on VirusTotal, suggesting an additional layer of risk beyond inappropriate content.

Interestingly, the study identified a correlation between website popularity and ad targeting practices. One of the more positive findings was that more popular websites were less likely to show targeted ads to children. Specifically, top-ranked websites tended to turn off ad personalization more often than lower-traffic sites, possibly due to greater scrutiny or more sophisticated compliance efforts.

Regarding content, the research identified a substantial presence of improper advertisements. The team found over 1,000 improper advertisements across 311 distinct children's websites. These ads often fell into categories deemed inappropriate for children, such as dating, weight loss, mental health, and racy content. Examples included an ad for Germany's largest online sex toy shop disguised with an image of ice cream, Alibaba ads featuring racing and disturbing imagery for attention, and "flirt finder" or scammy "amigates" ads appearing on educational children's sites.

The study also shed light on the origin of these advertisers. Analysis of ad disclosure interfaces revealed that advertisers from diverse global locations, including Cyprus, the United Arab Emirates, and Israel, were serving ads to children visiting websites in the EU or US. This highlights the transnational nature of the ad ecosystem and the challenges in enforcing regional regulations.

Finally, the research meticulously compared tracking prevalence. It found that websites displaying ads exhibited significantly more tracking activity than those without ads. Furthermore, a stark geographical difference was observed: US visits consistently showed more tracking than EU visits. For instance, a single website (masonwork.com) visited from New York City triggered 161 distinct third-party trackers, while a visit to another site (wspace.com) from the EU still triggered a substantial 95 distinct trackers, emphasizing the pervasive nature of surveillance.

Technical Deep Dive

▶ Watch: Methodology: Advanced crawler for ad and tracking data (4:45)

The research employed a multi-faceted empirical measurement research approach, combining advanced machine learning, a custom-built web crawler, and sophisticated ad analysis techniques to overcome the inherent challenges of studying child-directed websites.

The first significant technical hurdle was creating a comprehensive and reliable dataset of child-directed websites. The researchers addressed this by first compiling existing, expert-curated lists and incorporating data from online categorization services. To scale this, they developed a machine learning classifier designed to detect children's websites based on their titles and metadata, such as descriptions. A key innovation here was the use of an existing pre-trained distilled model. These are smaller, more efficient models that learn from larger, "teacher" models, offering comparable accuracy with reduced computational overhead. Crucially, the chosen model was multilingual, supporting around 50 different languages, enabling the identification of child-directed websites across diverse linguistic contexts. This classifier was then applied to a massive dataset of billions of crawled web pages. To ensure the robustness and reliability of the identified list, a subsequent manual review process was conducted to remove any false positives.

For the crawling process, the team utilized and heavily modified an existing open-source web crawler known as Tracker Rider Collector, originally developed by Dr. Go. This extended version of the crawler was designed to be highly interactive, simulating user behavior to uncover hidden elements and dynamic content. During its operation, the crawler meticulously recorded various types of data crucial for the study, including specific ad-related information, data from ad disclosure pages, function calls indicative of fingerprinting attempts, and detailed HTTP requests and responses. The crawler also possessed the capability to record video of the browsing sessions, providing a visual audit trail.

A core technical challenge was the automatic detection of ad targeting. The interactive crawler was engineered to first detect and scrape advertisements displayed on a page. Following this, it would programmatically identify and click on ad disclosure buttons (e.g., "Why am I seeing this ad?"). By scraping the information presented on these disclosure pages, the system could then determine if ad targeting or ad personalization was enabled for that specific advertisement or website. This automated approach allowed for large-scale analysis of ad targeting practices, moving beyond simple ad presence detection.

To account for geographical variations in ad delivery and tracking, the researchers deployed their crawler from multiple vantage points. These included two EU cities (Frankfurt and Amsterdam), one non-EU city (London), and two US cities (New York City and San Francisco). A limited measurement was also conducted on mobile devices to assess differences in ad display across platforms. The resulting dataset comprised over 72,000 pages collected from 2,000 distinct websites, with the crawler visiting five inner pages from each website, ensuring a broader and more representative sample than just homepages.

The detection of improper advertisements involved a sophisticated combination of automated content analysis and manual verification. For identifying visually racy or inappropriate content, the team leveraged the Google Cloud Vision API's safe search detection capabilities. Additionally, the same multilingual model used for website classification was employed to query for specific problematic search terms (e.g., "dating," "weight loss," "sex toys"). This model was capable of capturing the semantic similarity between the search terms and text extracted from advertisements, allowing for the identification of related ads even if the exact keywords were not present. Finally, a crucial manual review step was implemented to remove false positives and ensure the accuracy of the improper ad classifications. This process, though labor-intensive, was vital for confirming the presence of genuinely inappropriate content, such as the ice cream image leading to a sex toy shop or scam advertisements.

Demo / Proof of Concept

▶ Watch: Key Finding: Malicious ad links and popularity's effect on targeting (8:00)

While the talk did not feature a live, interactive "demo" in the traditional sense of a tool demonstration, it effectively served as a proof of concept by visually presenting numerous real-world examples of the improper advertisements discovered during the study. Zahra Moti showcased several compelling screenshots of actual ads found on legitimate child-directed websites, directly illustrating the research's findings.

For instance, one striking example presented was an advertisement featuring an appealing image of ice cream. However, upon closer inspection, the ad was revealed to be for "the largest online sex toy shop in Germany," a stark and inappropriate juxtaposition for a child's website. Another example displayed an advertisement from Alibaba featuring racing and "distasteful" images, clearly designed for attention-grabbing rather than suitability for a young audience.

The speaker further reinforced these findings by manually visiting and displaying live instances of these problematic ads on children's websites from their own computer. This included a Dutch website called "Vu junior," ostensibly for kids' activities, where a "flirt finder" ad with the Dutch text "go never go to bed alone" was prominently displayed. Another example showed "scammy" or "AI chatbot" ads labeled "amigates" appearing on a website dedicated to teaching children to count numbers. A third instance involved an improper advertisement found on a website providing kindergarten materials. These real-world demonstrations vividly underscored the widespread and persistent nature of inappropriate advertising on platforms intended for children, validating the automated detection methods used in the study and providing irrefutable evidence of the problem.

Defensive Implications

▶ Watch: Key Finding: Methods for detecting improper ad content (9:00)

The findings of this comprehensive study carry significant defensive implications for various stakeholders, including regulators, website operators, parents, and ad tech companies, all of whom play a role in safeguarding children's online safety and privacy.

For Regulators and Lawmakers: The research provides compelling evidence that current regulations, such as the EU DSA and US COPPA, are either insufficiently enforced or contain loopholes that allow pervasive online tracking and targeted advertising on child-directed websites. Regulators must prioritize aggressive enforcement actions against platforms and ad networks that violate these laws. Furthermore, the explicit identification of "improper ads" (dating, weight loss, racy content) suggests a need for more granular and explicit prohibitions on specific content categories deemed harmful to children, regardless of targeting intent. International cooperation is also crucial, given the global origin of many advertisers.

For Website Operators of Child-Directed Content: Website owners, especially those catering to children, bear a primary responsibility to ensure a safe online environment. They should implement stringent ad content filtering mechanisms, potentially leveraging tools similar to those used in this research (e.g., Google Cloud Vision API for content moderation). Prioritizing privacy-preserving ad models or direct sponsorships over complex programmatic advertising ecosystems can significantly reduce the risk of harmful ads and excessive tracking. Regular, independent audits of all third-party scripts, trackers, and ad networks integrated into their sites are essential to identify and remove non-compliant entities. The finding that popular websites are less likely to show targeted ads indicates that resources and conscious effort can make a difference.

For Ad Tech Companies and Advertisers: Companies operating in the advertising technology space have a moral and legal obligation to implement robust age verification mechanisms and stricter content moderation policies, particularly for ads served on websites identified as child-directed. Their systems must be re-engineered to prevent ad personalization for minors and to rigorously filter out content from the identified harmful categories. The study's disclosure to Google and other companies, prompting investigations, highlights the need for proactive self-regulation and accountability within the industry. Ad networks should also clearly identify advertisers' true identities and locations to facilitate accountability.

For Parents and Guardians: While policy and industry changes are critical, parents remain on the front line of protecting their children. They should be aware of the pervasive nature of tracking and inappropriate advertising. Employing browser-based ad blockers and privacy extensions can mitigate some of the risks. Educating children about the difference between content and advertising, and fostering critical thinking skills regarding online messages, is vital. Regularly monitoring children's online activity and utilizing parental control features offered by operating systems and browsers can also provide layers of protection.

In summary, the defensive implications call for a multi-pronged approach: stronger legal frameworks, diligent enforcement, proactive industry responsibility, and informed parental engagement. The transparency provided by studies like this is the first step towards building a safer digital space for children.

Key Takeaways

  • Pervasive Targeted Advertising: Despite existing regulations like EU DSA and US COPPA, targeted advertising remains widespread on child-directed websites, with over 70% of captured ads enabling personalization and 27% of websites displaying such ads.
  • Alarming Presence of Improper Ads: The study identified over 1,000 improper advertisements across 311 distinct children's websites, including content related to dating, weight loss, mental health, and racy material, often disguised or misleadingly placed.
  • Malicious Ad Links: An exploratory analysis revealed that a significant number of ad links (150 out of 4,000 sampled) were flagged as potentially malicious or phishing by VirusTotal, posing an additional security risk to children.
  • Popular Sites Show Better Compliance: More popular, higher-traffic children's websites were found to be less likely to use targeted ads, suggesting that greater resources or scrutiny can lead to better adherence to privacy principles.
  • Geographical Disparity in Tracking: There is significantly more third-party tracking on US-based visits to children's websites compared to EU visits, with some sites triggering over 160 distinct trackers, highlighting regional differences in enforcement or practices.
  • Urgent Need for Enforcement and Regulation: The findings underscore an urgent call for stronger regulatory enforcement, more specific policies against harmful ad content on child-directed platforms, and greater accountability from ad tech companies and website operators to protect children's online privacy and well-being.

About the Speaker(s)

The research was presented by Zahra Moti, a PhD student from Radboud University. She conducted this work in collaboration with her colleagues Asuman Senol, Hamid Bostani, Frederik Zuiderveen Borgesius, Veelasha Moonsamy, and Arunesh Mathur. Their joint effort, drawing expertise from Radboud University and other institutions, focused on empirically measuring and analyzing the complex landscape of tracking and advertising practices on websites specifically designed for children.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This research provides a crucial, large-scale empirical measurement of tracking and targeted advertising on child-directed websites. The sophisticated methodology uncovered pervasive non-compliance with privacy regulations and an alarming frequency of inappropriate ads, offering invaluable data for policymakers and practitioners.

Heather Calloway (CISO) — MUST SEE

This critical research exposes widespread non-compliance with child online privacy regulations, revealing pervasive tracking and inappropriate targeted advertising on child-directed websites. It provides undeniable evidence for regulatory action and mandates immediate re-evaluation of ad practices by website operators and ad tech companies to mitigate significant legal and reputational risk.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024