Investigating Moderation Challenges to Combating Hate and Harassment: The Case of Mod-Admin Power Dynamics and Feature Misuse on Reddit

Madiha Tabassum, Alana Mackey, Ashley Schuett, Ada Lerner

33rd USENIX Security Symposium · Day 1 · USENIX Security '24 · USENIX Security '24

Overview

Online platforms continue to grapple with pervasive issues of hate, harassment, and abuse, which manifest in diverse forms from targeted hate speech to doxing and non-consensual sharing of intimate media. This talk, presented at USENIX Security '24, delves into the intricate challenges faced by volunteer community moderators on Reddit, who serve as the frontline defense against such harms. The research, a collaboration between Wesley College, George Washington University, and Northeastern University, critically examines two core issues: the complex and often fraught relationship between volunteer moderators and Reddit's paid platform administrators, and the adversarial misuse of benign platform features to facilitate malicious activities.

Watch on YouTube

Visual summary for Investigating Moderation Challenges to Combating Hate and Harassment: The Case of Mod-Admin Power Dynamics and Feature Misuse on Reddit by Madiha Tabassum, Alana Mackey, Ashley Schuett, Ada Lerner
Visual summary for Investigating Moderation Challenges to Combating Hate and Harassment: The Case of Mod-Admin Power Dynamics and Feature Misuse on Reddit by Madiha Tabassum, Alana Mackey, Ashley Schuett, Ada Lerner

Key moments

  1. 0:00 Introduction to moderation challenges combating online hate
  2. 2:00 Research methodology: data collection and analysis
  3. 2:40 Challenges with Reddit admin unreliability and lack of transparency
  4. 4:50 Adversarial misuse of Reddit's platform features
  5. 5:10 Detailed example: username and gilding feature misuse
  6. 6:30 Taxonomy of feature misuses and adversarial capabilities
  7. 8:00 Applying a security mindset to social design

Investigating Moderation Challenges to Combating Hate and Harassment: The Case of Mod-Admin Power Dynamics and Feature Misuse on Reddit

Speakers: Madiha Tabassum, Alana Mackey, Ashley Schuett, Ada Lerner

Conference: USENIX Security '24

YouTube: https://www.youtube.com/watch?v=3JhY7R2SIcA

Overview

Online platforms continue to grapple with pervasive issues of hate, harassment, and abuse, which manifest in diverse forms from targeted hate speech to doxing and non-consensual sharing of intimate media. This talk, presented at USENIX Security '24, delves into the intricate challenges faced by volunteer community moderators on Reddit, who serve as the frontline defense against such harms. The research, a collaboration between Wesley College, George Washington University, and Northeastern University, critically examines two core issues: the complex and often fraught relationship between volunteer moderators and Reddit's paid platform administrators, and the adversarial misuse of benign platform features to facilitate malicious activities.

The significance of this work lies in its unique data source—discussions from r/modsupport and r/modhelp, subreddits specifically designed for Reddit moderators to seek assistance for problems beyond their immediate team's capabilities. This provides an invaluable lens into "edge case" challenges that, while not always common, severely test the limits of moderator resources and tools. By analyzing these discussions, the researchers uncover systemic vulnerabilities arising from power imbalances and overlooked design flaws, proposing a novel approach to viewing social moderation challenges through the established framework of technical security.

Ultimately, the talk advocates for a paradigm shift, urging platform designers and security professionals to apply a security mindset and threat modeling to the social design aspects of trust and safety. It highlights how seemingly innocuous platform features, when exploited by malicious actors, can create significant harm, often leaving volunteer moderators in an untenable position. The findings underscore the critical need for improved platform support, greater transparency, and a more collaborative relationship between platform administrators and the volunteer communities essential for maintaining safe and trustworthy online environments.

Background

▶ Watch: Introduction to moderation challenges combating online hate (0:00)

The landscape of online content moderation is increasingly complex, with platforms like Reddit relying heavily on a hybrid model that combines the efforts of volunteer community moderators with paid platform administrators. While administrators (referred to as "admins") are employees of Reddit responsible for site-wide moderation and policy enforcement, volunteer moderators are community members who dedicate their time to maintaining the safety and rules within specific subreddits. Prior research has extensively documented the immense pressures faced by these volunteer moderators, including overwhelming workloads, emotional burnout, and becoming targets of hate and harassment themselves. This study builds upon that foundation by focusing on the systemic challenges moderators encounter when attempting to combat hate and harassment, particularly when these issues intersect with platform-level dynamics and feature design.

The researchers' methodology involved a deep dive into two critical Reddit communities: r/modsupport and r/modhelp. These subreddits serve as vital forums where moderators from across the platform can seek advice, share experiences, and escalate issues they cannot resolve independently. This unique data source allowed the team to identify challenges that, while potentially less frequent, are profoundly difficult and often expose the limitations of existing moderation tools and policies. From these subreddits, the researchers downloaded all available threads. To identify relevant discussions, they compiled a list of 35 keywords directly associated with hate, harassment, and online abuse. This keyword-filtered dataset yielded 3,321 relevant threads, from which a random sample of 115 threads was selected for in-depth qualitative coding analysis. This rigorous approach allowed the researchers to uncover recurring themes and systemic issues that characterize the challenges faced by Reddit's volunteer moderation community.

Key Findings

▶ Watch: Challenges with Reddit admin unreliability and lack of transparency (2:40)

The qualitative analysis revealed two overarching categories of challenges that significantly impede volunteer moderators' ability to combat hate and harassment effectively. These findings highlight not only operational difficulties but also fundamental issues in platform design and administrative support.

The first major finding centers on the complex and often adversarial relationship between volunteer moderators and Reddit administrators, compounded by the limitations of automated platform moderation tools. Moderators frequently expressed profound frustration due to a lack of timely response from admins, even in critical situations demanding immediate attention. For instance, one moderator managing a mental health subreddit reported a user actively encouraging self-harm, yet received no response from admins for over 72 hours. Such delays are not only unacceptable but can have severe real-world consequences, eroding moderator trust and community safety. Beyond unresponsiveness, moderators also reported instances of admins and automated tools taking inaccurate actions, such as incorrectly removing content or unjustly banning users. This highlights flaws in automated detection systems and potentially insufficient human oversight.

A significant point of contention was the pervasive lack of transparency regarding admin actions. Moderators frequently observed admins taking decisive actions—like content removal or user bans—without providing any explanation or context. This opaqueness breeds confusion and resentment, leaving moderators unable to understand the rationale behind decisions that directly impact their communities. The aggregate effect of these issues—unresponsiveness, inaccuracy, and lack of transparency—fostered a perception among moderators of admins as "powerful but unreliable entities." While moderators acknowledged admins' ultimate authority and often expressed gratitude for occasional assistance, their predominant emotional responses ranged from frustration, confusion, and anger to fear of unfair punishment for themselves or their communities. This reliance on a powerful yet inconsistent entity creates a precarious environment for volunteer moderators striving to maintain safe online spaces.

The second critical finding identified was the widespread issue of adversarial misuse of platform features. This category describes situations where malicious actors, operating without any privileged access, exploit standard platform UI features in unintended ways to facilitate hate, harassment, and abuse. Unlike traditional security vulnerabilities that might involve exploiting code flaws, this form of abuse leverages legitimate functionalities in an adversarial manner. The researchers observed numerous instances where the very design of Reddit's features, intended for benign user interaction, could be weaponized. This highlights a critical blind spot in platform design, where the potential for malicious actors to subvert features for harm is often overlooked. The subsequent sections will delve deeper into the technical aspects of this feature misuse and demonstrate a compelling example.

Technical Deep Dive

▶ Watch: Adversarial misuse of Reddit's platform features (4:50)

The concept of adversarial misuse of platform features represents a critical technical challenge identified in the research. This phenomenon occurs when individuals, without requiring any elevated permissions or exploiting traditional software vulnerabilities, leverage standard user interface elements or functionalities in ways unintended by the platform designers to perpetuate hate, harassment, or abuse. This is distinct from typical security exploits, as it doesn't necessarily involve bypassing security controls but rather subverting the intent of a feature.

To systematically categorize and understand these abuses, the researchers developed a taxonomy to classify the types of harm enabled by feature misuse. Furthermore, their analysis of observed misuses in the data set allowed them to derive eight distinct adversarial capabilities. These capabilities represent the different ways adversaries can weaponize platform features, such as the ability to "disseminate toxic content," "leak private information," or "silence legitimate users." While the talk did not enumerate all eight capabilities in detail, it emphasized their existence as a framework for understanding the multifaceted nature of feature-based abuse.

A core argument of the research is that these social engineering and content moderation challenges, particularly those stemming from feature misuse, should be reframed and treated as technical security challenges. The security community possesses a robust toolkit for analyzing and mitigating adversarial situations, including the application of a security mindset and threat modeling. A security mindset involves proactively considering how a system or feature could be abused, even if the primary design intent is benevolent. Threat modeling, a structured approach, identifies potential threats, vulnerabilities, and counter-measures by analyzing a system from an attacker's perspective. Applying these established security practices to the social design aspects of trust and safety could uncover unexpected avenues for abuse and lead to more resilient platform architectures.

The researchers also highlighted how platform profit and incentive models can inadvertently exacerbate these issues. Features designed to encourage engagement or monetization (e.g., purchasing virtual goods or awards) can, when misused, create a perverse incentive for platforms to tolerate abusive behavior if it generates revenue. This misalignment of incentives can erode user and moderator trust, especially when profit motives appear to outweigh safety concerns. Furthermore, the design space for platform policies requires careful investigation. Policies intended to benefit users, such as Reddit's liberal approach to anonymous throwaway accounts or the ability to delete/edit history, while valuable for user privacy, can simultaneously empower adversaries by providing anonymity and obscuring their malicious actions. The technical deep dive thus extends beyond mere feature analysis to encompass the broader ecosystem of design principles, economic incentives, and policy frameworks that collectively shape the security posture against social abuse.

Demo / Proof of Concept

▶ Watch: Taxonomy of feature misuses and adversarial capabilities (6:30)

To concretely illustrate the adversarial misuse of platform features, the speakers presented a compelling scenario involving Reddit's gilding feature and username selection. This demonstration effectively highlighted how seemingly innocuous functionalities can be weaponized with significant negative consequences for both legitimate users and moderators.

The scenario unfolds as follows:

  • Alex is a dedicated moderator of r/usnifan, a community where enthusiasts of USENIX interact.
  • Riley, a contributing member, creates a post praising USENIX, sparking engaging and intellectual discussions within the subreddit.
  • Tyler, an individual with a strong dislike for USENIX, decides to disrupt the community. He creates a new Reddit account with the adversarial username "us ni is evil."
  • Leveraging Reddit's gilding feature, Tyler proceeds to purchase coins and awards Riley's positive post. The crucial aspect of the gilding feature is that when an award is given, the username of the giver (in this case, "us ni is evil") is permanently displayed with an icon directly on top of the awarded post.

This action immediately places Alex, the moderator, in a difficult position. Tyler's toxic username is now prominently and permanently affixed to Riley's otherwise legitimate and valuable post. Alex, recognizing the malicious intent, attempts to mitigate the situation.

  • Alex considers banning Tyler, but on Reddit, a banned user can still give awards. This means banning Tyler would not prevent him from repeating the attack on other posts.
  • Alex then tries to report Tyler's username directly to Reddit administrators, only to discover there is no specific option to report a username for abuse.
  • Recalling past experiences with unresponsive administrators, Alex is hesitant to use the general modmail system, knowing that timely intervention is unlikely.

Left with limited options, Alex faces a dilemma: the only effective way to remove the toxic content ("us ni is evil") from Riley's post is to delete Riley's entire post. This outcome is highly undesirable as it punishes Riley, a legitimate and contributing member of the community, for the actions of an adversary. The adversary, Tyler, remains unpunished and free to continue their disruptive behavior. Furthermore, Alex contemplates that even if Reddit administrators were contacted, they might be reluctant to take action against Tyler's account because Tyler is actively spending money on the platform through the gilding feature, creating a potential conflict of interest between platform revenue and user safety.

This demonstration perfectly illustrates how the misuse of a combined feature (gilding) and a lack of specific reporting tools can enable the dissemination of toxic content and, more insidiously, a form of silencing where legitimate contributions are removed to combat abuse, rather than addressing the abuser directly. It underscores the profound challenges volunteer moderators face when platform design flaws intersect with adversarial intent.

Defensive Implications

▶ Watch: Applying a security mindset to social design (8:00)

The insights gleaned from this research offer critical defensive implications for platform designers, security teams, and moderation policy-makers. The primary takeaway is the imperative to shift perspective: hate and harassment attacks facilitated by feature misuse should be treated with the same rigor as technical security challenges.

  1. Integrate Security Mindset into Social Design: Platform development must move beyond merely preventing traditional exploits (e.g., SQL injection, XSS) to proactively considering how seemingly benign features can be subverted for social harm. This requires adopting a security mindset from the earliest stages of feature design, asking "How could this feature be misused?" rather than solely "How will users legitimately interact with this feature?"
  2. Employ Threat Modeling for Social Attack Surfaces: Just as technical systems are threat modeled, social interaction features and user interfaces should undergo similar analysis. Threat modeling should identify potential adversaries, their motivations, and the ways they could exploit platform functionalities (e.g., usernames, awards, messaging, profiles) to execute hate, harassment, or disinformation campaigns. This includes mapping out the eight adversarial capabilities identified in the research and designing features to mitigate them.
  3. Re-evaluate Platform Profit and Incentive Models: Companies must critically assess how their revenue-generating and engagement-driving features might inadvertently create vulnerabilities. If a feature's misuse generates revenue (as seen with the gilding example), it can create a perverse incentive structure that conflicts with safety objectives. Platforms should prioritize user safety and moderator trust over short-term financial gains.
  4. Scrutinize Platform Policies for Adversarial Exploitation: Policies designed for legitimate user benefits, such as anonymity or content editing/deletion history, need careful re-evaluation. While preserving user privacy is crucial, policies must also consider how they might be exploited by malicious actors. Designers should explore ways to limit misuse while retaining the legitimate benefits of these policies, perhaps through more robust reporting mechanisms, better identity verification for certain actions, or clearer audit trails accessible to trusted moderators/admins.
  5. Empower and Involve Volunteer Moderators: Given that moderators are at the forefront of these attacks and possess intimate knowledge of feature misuse, their involvement in the design and policy-making process is paramount. Platforms should actively engage moderators in user research, threat modeling workshops, and feedback loops for new features. Their firsthand experience can provide invaluable insights into potential vulnerabilities that might be overlooked by designers.
  6. Improve Administrative Support and Transparency: Platforms must commit to more timely and transparent responses to moderator escalations, especially for critical issues like self-harm or abuse of minors. Providing clear explanations for administrative actions and offering robust reporting tools for specific types of abuse (e.g., problematic usernames) will significantly enhance moderators' effectiveness and rebuild trust. Supporting volunteer moderators with better tools and clear communication is critical for the growth and safety of online communities.

Key Takeaways

  • Treat Social Abuse as Technical Security: Hate and harassment attacks facilitated by platform feature misuse should be analyzed and mitigated using established technical security practices like a security mindset and threat modeling.
  • Beyond Code Exploits: Adversaries often exploit legitimate UI features and design paradigms rather than just technical vulnerabilities, demanding a broader approach to security.
  • Involve Moderators in Design: Volunteer moderators possess invaluable firsthand experience with feature misuse; their active involvement in platform design and policy development is crucial for building safer systems.
  • Mind Incentive Alignment: Platform profit and incentive models can inadvertently create vulnerabilities or disincentivize robust moderation, requiring careful ethical consideration during design.
  • Balance Policy Benefits and Risks: Policies designed for user privacy (e.g., anonymity, edit history) must be carefully balanced against their potential for adversarial exploitation, with efforts to mitigate misuse while preserving legitimate benefits.
  • Support Frontline Defenders: Volunteer moderators are heroes facing immense challenges; platforms must provide timely, transparent administrative support and better tools to empower them in maintaining community safety.

About the Speaker(s)

The research presented was a collaborative effort by Madiha Tabassum, Alana Mackey, Ashley Schuett, and Ada Lerner. Madiha Tabassum, Alana Mackey, and Ashley Schuett are affiliated with Wesley College and George Washington University. Ada Lerner is associated with Northeastern University. Their collective expertise spans the domains of human-computer interaction, online safety, and security.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This research forcefully argues for treating platform feature misuse as a technical security problem, not just a social one. By applying threat modeling and a security mindset to social design, the talk exposes how benign features can be weaponized, offering a crucial paradigm shift for platform builders and policy makers. The demo is a stark, effective illustration of systemic failure.

Heather Calloway (CISO) — STRONG ACCEPT

This research compellingly reframes social moderation challenges as core technical security issues, highlighting critical governance failures and the business impact of misaligned incentives on online platforms. It offers a clear framework and actionable recommendations for platform designers and leaders to apply a security mindset to social design, moving beyond reactive content moderation to proactive risk mitigation.

→ Top-rated talks at 33rd USENIX Security Symposium

All talks from 33rd USENIX Security Symposium