Targeted and Troublesome: Tracking and Advertising on Children's Websites

Zahra Moti, Asuman Senol, Hamid Bostani, Frederik Zuiderveen Borgesius, Veelasha Moonsamy, Arunesh Mathur

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 4

Overview

This research, presented by Zahra Moti, a PhD student from Radboud University, delves into the pervasive and often problematic landscape of online tracking and advertising on websites directed at children. The talk highlights how children, a particularly vulnerable demographic, are exposed to targeted and inappropriate advertisements despite existing legal frameworks designed to protect them. The study addresses a critical gap in large-scale empirical research focusing specifically on child-directed online content, providing a comprehensive measurement of advertising practices, third-party trackers, and fingerprinting attempts.

Watch on YouTube

Visual summary for Targeted and Troublesome: Tracking and Advertising on Children's Websites by Zahra Moti, Asuman Senol, Hamid Bostani, Frederik Zuiderveen Borgesius, Veelasha Moonsamy, Arunesh Mathur
Visual summary for Targeted and Troublesome: Tracking and Advertising on Children's Websites by Zahra Moti, Asuman Senol, Hamid Bostani, Frederik Zuiderveen Borgesius, Veelasha Moonsamy, Arunesh Mathur

Key moments

  1. 0:40 Children's vulnerability to online tracking and ads
  2. 2:00 Regulatory landscape and study's core question
  3. 3:50 Novel methodology for identifying child-directed websites
  4. 7:00 High prevalence of targeted ads on children's websites
  5. 8:20 Discovery of malicious or phishing links in ads
  6. 9:00 Popular websites less likely to show targeted ads
  7. 9:20 Detecting and categorizing improper ad content

Targeted and Troublesome: Tracking and Advertising on Children's Websites

Speakers: Zahra Moti; Asuman Senol; Hamid Bostani; Frederik Zuiderveen Borgesius; Veelasha Moonsamy; Arunesh Mathur

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=5Yeip-9bng

Overview

This research, presented by Zahra Moti, a PhD student from Radboud University, delves into the pervasive and often problematic landscape of online tracking and advertising on websites directed at children. The talk highlights how children, a particularly vulnerable demographic, are exposed to targeted and inappropriate advertisements despite existing legal frameworks designed to protect them. The study addresses a critical gap in large-scale empirical research focusing specifically on child-directed online content, providing a comprehensive measurement of advertising practices, third-party trackers, and fingerprinting attempts.

The work underscores the significant threats posed by online tracking and personalized advertising to children's privacy and well-being. Unlike adults, children often misunderstand the intent and message of advertisements, making them susceptible to manipulation and potentially harmful content. This research is especially timely, given recent legislative changes like the EU's Digital Services Act (DSA) and the US's Children's Online Privacy Protection Act (COPPA), which aim to curb such practices.

Through a robust methodology involving a custom dataset of child-directed websites, an extended web crawler, and advanced ad detection techniques, the researchers uncover alarming statistics. Their findings reveal the widespread presence of targeted and improper ads, the extensive use of third-party trackers, and significant geographical disparities in tracking intensity. The study not only quantifies the problem but also provides concrete examples of concerning advertisements, advocating for stronger regulatory enforcement and continued research to safeguard children's digital experiences.

Background

▶ Watch: Children's vulnerability to online tracking and ads (0:40)

The online environment presents unique challenges for children, who are increasingly engaging with digital content from a young age. While online tracking and targeted advertising pose risks to users of all ages, children are significantly more vulnerable. Prior research has demonstrated that children frequently misunderstand the persuasive intent of advertisements and are more likely to be influenced by negative or manipulative messaging, particularly concerning topics like gambling, unhealthy food, or weight loss. Despite this heightened vulnerability, there has been a notable absence of large-scale empirical studies specifically examining tracking and advertising practices on websites explicitly designed for children.

This research emerges against a backdrop of evolving global regulations aimed at protecting children online. In the European Union, the Digital Services Act (DSA) specifically prohibits targeting children with advertisements based on profiling. Similarly, in the United States, the Children's Online Privacy Protection Act (COPPA) mandates verifiable parental consent for data collection used in targeted advertising on child-directed websites. These legislative efforts signify a growing recognition of the problem, yet the effectiveness and enforcement of such laws remain a significant concern, as evidenced by the findings of this study.

Previous academic work, such as research from the University of Washington, has explored the content of advertisements, identifying a prevalence of inappropriate ads on general news and misinformation websites. However, these studies did not specifically focus on the unique context of child-directed content. The present research seeks to bridge this critical gap, providing an empirical measurement of the current state of targeted and personalized advertising shown to children, and investigating specific categories of ads that prior research and regulatory reports have identified as particularly harmful to this demographic, including dating, mental health, weight loss, and racy content. The existence of these regulations highlights the problem, and this study provides crucial data on the current state of compliance and enforcement.

Key Findings

▶ Watch: Novel methodology for identifying child-directed websites (3:50)

The study's comprehensive measurements yielded several critical findings, painting a concerning picture of online tracking and advertising practices on child-directed websites. The researchers scraped over 70,000 advertisements from 804 distinct child-directed websites, finding ads present on 36% of these sites. More alarmingly, an average of 27% of the websites contained targeted advertisements, a practice often restricted or banned by current regulations. Of the advertisements captured, over 70% were found to enable ad targeting or personalization, indicating that non-targeted ads are a minority even on children's platforms.

A significant, albeit positive, finding was the correlation between website popularity and ad targeting. The study observed that more popular websites were less likely to display targeted advertisements to children, suggesting that top-tier websites more frequently disable ad personalization compared to lower-traffic sites.

However, the prevalence of improper and potentially harmful content remained a major concern. The researchers identified over 1,000 improper advertisements across 311 distinct child-directed websites, even with a manual review stop after 1,000 ads per category, indicating the true number could be higher. These improper ads fell into categories like dating, mental health, weight loss, and racy content, which are explicitly deemed unsuitable for children. Furthermore, an exploratory analysis of ad links revealed that 150 out of 4,000 sampled links were flagged as malicious or phishing by at least one scan engine on VirusTotal, highlighting potential security risks alongside privacy concerns.

The study also shed light on the origins of advertisers, noting that companies from various global regions, including Cyprus, the United Arab Emirates, and Israel, were serving ads to children in the EU or US. Regarding tracking, the research established a clear link: websites displaying advertisements consistently exhibited a higher number of third-party trackers. A geographical comparison further revealed a disparity, with US visits triggering more trackers than EU visits. For instance, a single website visited from New York City recorded 161 distinct third-party trackers, whereas a visit to a different site from an EU city registered 95 distinct trackers, still a substantial number. These findings collectively underscore the extensive tracking, frequent targeted advertising, and the disturbing presence of improper content on websites intended for children.

Technical Deep Dive

▶ Watch: High prevalence of targeted ads on children's websites (7:00)

The research methodology was meticulously designed to overcome significant challenges in studying tracking and advertising on child-directed websites. The first major hurdle was the lack of a comprehensive and up-to-date list of such websites suitable for a large-scale empirical study. To address this, the team constructed its own repository. This process began by aggregating existing, smaller lists of hand-picked children's websites and integrating data from online categorization services.

To scale this effort, the researchers developed a machine learning classifier. This classifier was trained to identify child-directed websites based on their titles and metadata, such as descriptions. A key innovation was the use of an existing, pre-trained distilled model. These are smaller, highly efficient models that learn from much larger "teacher" models, offering comparable accuracy with reduced computational overhead. Crucially, the chosen model was multilingual, supporting approximately 50 different languages, which enabled the detection of child-directed websites across diverse linguistic contexts. This classifier was then applied to a vast dataset of web content, encompassing billions of crawled pages, to identify potential child-directed sites. To ensure the reliability and robustness of their list, a rigorous manual review was conducted to remove any false positives.

The second challenge involved automatically detecting advertisements, particularly targeted and improper ones, and scraping their associated disclosure pages. For the crawling process, the researchers utilized and heavily modified an existing open-source web crawler: Tracker Rider Collector. This extended version was designed to be highly interactive, mimicking user behavior to engage with web pages. During a crawl, it recorded various types of data, including ad-related content, information from ad disclosure pages, fingerprinting-related function calls, HTTP requests and responses, and even video recordings of the browsing session.

To determine if ad targeting was enabled, the interactive crawler was programmed to first detect and scrape advertisements. Following this, it would automatically click on any available "ad disclosure" buttons or links and then scrape the information presented on the disclosure page. This data was subsequently analyzed to ascertain whether ad targeting or personalization was active for specific advertisements or websites.

The crawling infrastructure was strategically deployed from different geographical vantage points to capture regional variations in ad delivery and tracking. These included two EU cities (Frankfurt and Amsterdam), one non-EU city (London), and two US cities (New York City and San Francisco). A limited measurement was also conducted on mobile devices to observe differences in ad presentation across platforms. The final dataset was substantial, comprising data from over 70,000 pages collected from 2,000 distinct websites, with five inner pages crawled from each site, not just homepages.

For identifying improper advertisements, the researchers employed a multi-pronged approach. To detect clickbait or racy content within ad images, they leveraged the Google Cloud Vision API's explicit content detection capabilities. For textual content, the aforementioned multilingual model was used. This model allowed researchers to query for specific search terms related to prohibited categories (e.g., "dating," "racing") and identify ads that exhibited semantic similarity with these terms, even across different languages. This automated detection was then supplemented by a crucial manual review process to eliminate false positives and ensure accuracy in identifying truly improper advertisements.

Demo / Proof of Concept

▶ Watch: Popular websites less likely to show targeted ads (9:00)

The talk included compelling demonstrations and examples of the improper advertisements discovered on child-directed websites, providing concrete evidence of the research findings. These examples vividly illustrated the types of content children are exposed to, often in violation of established regulations.

One striking example presented was an advertisement featuring an image of ice cream, ostensibly appealing to children. However, this ad, identified as Figure B in the presentation, linked to the "largest online sex toy shop in Germany." This deceptive pairing of child-friendly imagery with highly inappropriate content highlights the insidious nature of some of the advertising found. Another example, Figure A, showed an advertisement from Alibaba that displayed "racy and distasteful images" purely for the purpose of grabbing attention, again on a website intended for children.

Beyond static images, the speakers manually visited some of the identified websites to provide live demonstrations of these improper ads. On a website named "Vu junior," which is clearly a kids' activity site, an advertisement written in Dutch was observed at the bottom of the page. The ad, which read "go never go to bed alone," was a promotion for a "flirt finder," described as a scam and an AI chatbot service – content unequivocally unsuitable for children.

Another instance showcased a website designed to help children learn numbers up to 20. On this educational platform, the researchers found ads for "Amigates," which were characterized as scammy content. A third example involved a website offering "Kinder Garten material," where again, an improper advertisement was detected. These real-world examples, verified by manual visits, underscored the pervasive and often contextually inappropriate nature of the ads children encounter, demonstrating a clear failure of content moderation and enforcement on these platforms. The manual verification process further solidified the validity of their automated detection methods.

Defensive Implications

▶ Watch: Detecting and categorizing improper ad content (9:20)

The findings of this research carry significant implications for various stakeholders involved in protecting children's online privacy and safety. The extensive tracking, pervasive targeted advertising, and widespread presence of improper content on child-directed websites demand a multi-faceted defensive strategy involving regulators, industry, and parents.

Firstly, the study calls for more robust regulation and, critically, more rigorous enforcement of existing laws like the EU's DSA and the US's COPPA. While these laws provide a framework for protection, their effectiveness is undermined by the practices uncovered. Regulators must strengthen oversight, implement more proactive monitoring mechanisms, and impose meaningful penalties on platforms and advertisers that violate these protections. The fact that advertisers from various global regions are serving improper ads to children in regulated zones highlights the need for international cooperation in enforcement.

Secondly, platforms and advertisers bear a significant responsibility. There is an urgent need for improved ad filtering and content moderation technologies to prevent inappropriate advertisements from appearing on child-directed sites. Advertisers must develop and adhere to stricter content guidelines, and ad networks should implement more sophisticated categorization and targeting exclusion mechanisms for child audiences. The study's finding that more popular websites are less likely to use targeted ads suggests that it is technically feasible for platforms to implement such restrictions, and this practice should be universally adopted. The researchers demonstrated proactive engagement by reaching out to five companies found to serve racy ads, with one company acknowledging the issue and promising investigation. They also disclosed 34 racy ads to Google via their "report this ad" button and shared their results with European data protection agencies, consumer protection agencies, and a UK-based NGO focusing on children's digital rights, emphasizing the collaborative effort required.

Thirdly, parents and educators play a crucial role in enhancing children's digital literacy and safety. While technical and regulatory solutions are paramount, parental awareness of these issues is also vital. Parents need to be informed about the types of tracking and advertising children are exposed to and equipped with tools and knowledge to configure parental controls, use ad blockers, and guide their children towards safer online environments.

Ultimately, the research advocates for a concerted effort from all sectors to create a safer and more private online experience for children. This includes ongoing research to monitor evolving threats, continuous improvement of regulatory frameworks, and proactive measures by the industry to prioritize child protection over advertising revenue.

Key Takeaways

  • Extensive Tracking and Targeted Advertising: The study revealed widespread third-party tracking and frequent targeted advertising on child-directed websites, with over 70% of captured ads being personalized.
  • Prevalence of Improper Content: A significant number of improper and potentially harmful advertisements, including dating, weight loss, and racy content, were found across hundreds of child-directed websites.
  • Popularity vs. Targeting: More popular websites were observed to be less likely to use targeted advertisements, suggesting a potential for better compliance or self-regulation among top-tier platforms.
  • Geographical Disparities in Tracking: US-based visits to child-directed websites consistently showed a higher number of distinct third-party trackers compared to EU visits, highlighting regional differences in data collection practices.
  • Urgent Need for Stronger Enforcement: The findings underscore the critical need for more robust enforcement of existing regulations like the DSA and COPPA, as well as improved ad filtering and content moderation by platforms and advertisers.
  • Open-Source Contribution: The research team has made their source code and data publicly available on GitHub, fostering further research and transparency in this critical area.

About the Speaker(s)

The primary speaker for this presentation was Zahra Moti, a PhD student at Radboud University. She presented this work as a joint effort with her colleagues Asuman Senol, Hamid Bostani, Frederik Zuiderveen Borgesius, Veelasha Moonsamy, and Arunesh Mathur, from Radboud University and the University of Bonn. Her research focuses on the critical area of online tracking and advertising, particularly concerning vulnerable user groups like children.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research meticulously exposes the pervasive and problematic landscape of online tracking and advertising on child-directed websites. Through novel methodology, it quantifies widespread targeted and improper ads, revealing critical failures in regulatory enforcement and platform responsibility. The findings provide indispensable data for policymakers and industry alike, demanding immediate action to protect children's online privacy.

Heather Calloway (CISO) — STRONG ACCEPT

This research uncovers critical failures in protecting children online, detailing extensive tracking and inappropriate advertising on child-directed websites. It provides crucial empirical evidence for regulators and platforms to address significant governance and compliance gaps, demanding immediate executive attention.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024