Predictive Response Optimization: Using Reinforcement Learning to Fight Online Social Network Abuse

Garrett Wilson

34th USENIX Security Symposium (USENIX Security '25) · Day 2 · Web and Mobile Security

Overview

In the realm of online social networks, the battle against abuse is a perpetual arms race. Traditional approaches have largely focused on the detection of malicious activities, often grappling with the inherent trade-off between precision and recall. However, as Garrett Wilson articulates in his USENIX Security talk, "Predictive Response Optimization: Using Reinforcement Learning to Fight Online Social Network Abuse," the true challenge and opportunity lie not just in detection, but in the intelligent selection of enforcement actions. This presentation introduces a paradigm shift, reframing the problem from a binary classification task to an action selection challenge, which can be optimally addressed using reinforcement learning (RL).

Watch on YouTube · Slides

Visual summary for Predictive Response Optimization: Using Reinforcement Learning to Fight Online Social Network Abuse by Garrett Wilson
Visual summary for Predictive Response Optimization: Using Reinforcement Learning to Fight Online Social Network Abuse by Garrett Wilson

Key moments

  1. 0:00 Introduction to OSN abuse and detection limitations
  2. 1:00 Key insight: expanding action set for abuse enforcement
  3. 2:00 Paper's contributions: new perspective, PRO system
  4. 2:50 Disadvantages of traditional rule-based anti-abuse systems
  5. 4:00 Introducing Predictive Response Optimization (PRO) framework
  6. 4:40 PRO Part 1: Reinforcement Learning for entity action selection
  7. 6:30 PRO Part 2: Model Predictive Control for global constraints
  8. 8:00 PRO implementation on Instagram/Facebook for data scraping

Predictive Response Optimization: Using Reinforcement Learning to Fight Online Social Network Abuse

Speakers: Garrett Wilson

Conference: USENIX Security

YouTube: https://www.youtube.com/watch?v=ezKTu_nyRzU

Overview

In the realm of online social networks, the battle against abuse is a perpetual arms race. Traditional approaches have largely focused on the detection of malicious activities, often grappling with the inherent trade-off between precision and recall. However, as Garrett Wilson articulates in his USENIX Security talk, "Predictive Response Optimization: Using Reinforcement Learning to Fight Online Social Network Abuse," the true challenge and opportunity lie not just in detection, but in the intelligent selection of enforcement actions. This presentation introduces a paradigm shift, reframing the problem from a binary classification task to an action selection challenge, which can be optimally addressed using reinforcement learning (RL).

Garrett Wilson, a researcher involved in the development of this system, presents Predictive Response Optimization (PRO), a novel framework designed to dynamically choose the most effective responses to abusive behavior on platforms like Instagram and Facebook. The core innovation of PRO is its ability to expand the action set available to platforms and, crucially, to learn which actions to apply to specific entities under varying conditions. This approach moves beyond static rule-based systems, enabling platforms to adapt to evolving threats, optimize abuse reduction, and minimize adverse impacts on legitimate users.

The significance of PRO cannot be overstated for the cybersecurity landscape of online platforms. By treating enforcement as an optimization problem, PRO offers a robust, adaptive, and automated solution to combat diverse forms of abuse—from spam and phishing to data scraping and fake accounts. It provides a blueprint for how social networks can move towards more sophisticated, data-driven security operations that are resilient to adversarial adaptation and responsive to changing business objectives, ultimately creating a safer and more reliable user experience.

Background

▶ Watch: Introduction to OSN abuse and detection limitations (0:00)

The landscape of online social network abuse is vast and multifaceted, encompassing issues such as spam, phishing, fake engagement, data scraping, account compromise, and the proliferation of fake accounts. Historically, the primary focus in combating these threats has been on the detection problem. Researchers and security teams have dedicated considerable effort to developing sophisticated models capable of identifying abusive entities, often illustrated by the classic precision-recall curve which highlights the inherent trade-off: increasing the accuracy of identifying true positives (precision) often comes at the expense of catching all malicious entities (recall), and vice versa.

While detection models have advanced significantly, the subsequent phase – enforcement actions – has received comparatively less attention in prior work. Many online social networks traditionally employ simplistic or static enforcement mechanisms. Common approaches include applying a "mild" response for low-confidence detections and a "harsh" response for high-confidence detections, or implementing a strike-based system that escalates actions against repeat offenders. These methods, while functional, suffer from several critical disadvantages that limit their effectiveness and adaptability.

The limitations of traditional rule-based or strike-based enforcement systems are pronounced. Firstly, they often lead to high abuse prevalence because the pre-defined actions may not be optimally positioned or effective against the current attacker tactics. Attackers continuously evolve their methods, and a fixed set of responses can quickly become outdated. Secondly, the fixed logic of these systems inherently struggles to adapt to changing attacker behavior. What was an optimal response at one point may become ineffective as adversaries find new ways to bypass or mitigate the enforcement. This lack of automated adjustment necessitates constant manual intervention. Thirdly, introducing new actions into such systems is a complex and risky endeavor. Without a clear framework for integration, new, potentially effective actions might be placed suboptimally, leading to unintended negative shifts in key metrics rather than improvements. This reliance on manual tuning and testing to optimize enforcement strategies is not only resource-intensive but also slow, leaving platforms vulnerable during periods of rapid adversarial adaptation.

The key insight presented by Wilson is to expand the action set available to platforms and to reframe the problem from mere binary classification to action selection. This expanded action set, coupled with an intelligent selection mechanism, provides an additional "degree of freedom" that allows platforms to optimize the trade-off between abuse reduction and operational costs, pushing out the traditional precision-recall frontier and enabling a more dynamic and effective defense strategy.

Key Findings

▶ Watch: Paper's contributions: new perspective, PRO system (2:00)

The talk "Predictive Response Optimization: Using Reinforcement Learning to Fight Online Social Network Abuse" presents several pivotal contributions and findings that mark a significant advancement in anti-abuse strategies for online platforms.

Firstly, the most fundamental contribution is the introduction of a new perspective: that the true goal of online social network anti-abuse systems should be action selection, rather than solely focusing on binary classification (detection). This reframing acknowledges that what a service does after identifying potential abuse is as critical as, if not more critical than, the initial detection itself.

Secondly, to operationalize this new perspective, the authors provide a formalization of action selection as a reinforcement learning (RL) problem. This formalization explicitly addresses the critical trade-off between minimizing abuse volume and managing the associated costs, particularly the cost of blocking or impacting benign users. By casting the problem in an RL framework, the system can learn optimal policies through interaction with the environment.

Thirdly, to solve this formally defined problem, the talk introduces and details the design of a novel system called Predictive Response Optimization (PRO). PRO is engineered to intelligently select enforcement actions, leveraging a sophisticated combination of RL and model predictive control to achieve its objectives.

Fourthly, the practical implementation and evaluation of PRO on major social networks, specifically Instagram and Facebook, demonstrate its tangible benefits. The experiments show PRO's ability to significantly reduce automated activity volume—specifically data scraping—without causing statistically significant negative impact (degradation) to legitimate users or key operational metrics. On Instagram, PRO achieved a 59% reduction in weighted scraping requests, and on Facebook, a 4.5% reduction, both while maintaining stable cost metrics.

Finally, a series of compelling case studies further illustrates PRO's robust capability to automatically adjust to changing conditions. These include adapting to evolving business objectives (e.g., introducing new cost metrics like "days active on mobile web" or "dollars spent on SMS"), integrating new enforcement actions (e.g., "forced reauthentication"), responding to unexpected systemic changes (e.g., a bug in an identity verification action), and, critically, adjusting to adversarial adaptation (e.g., when a warning action became ineffective over time). These demonstrations underscore PRO's dynamic and resilient nature, a critical advantage over static rule-based systems.

Technical Deep Dive

▶ Watch: Introducing Predictive Response Optimization (PRO) framework (4:00)

The core of Garrett Wilson's proposed solution, Predictive Response Optimization (PRO), lies in its innovative architecture that synergistically combines elements of reinforcement learning (RL) with model predictive control (MPC). This hybrid approach enables PRO to make optimal local decisions for individual entities while ensuring global constraints and objectives are met.

Part 1: Optimizing Actions for Each Entity (Reinforcement Learning Component)

At its heart, PRO treats the selection of an enforcement action for a given entity (e.g., a user account) as an RL problem. For each entity, PRO considers a set of potential actions, from action 1 to action N. The process for selecting the optimal action for a specific entity unfolds as follows:

  1. Feature Extraction: For a given account, a set of features (X) is extracted. These features provide a comprehensive profile of the account's behavior, history, and characteristics, which are crucial for predicting future outcomes.
  2. Candidate Action Evaluation: For each candidate action that could be applied to the account, PRO employs two critical predictive models:
  • Abuse Model: This model predicts the amount of abuse that would occur if the specific candidate action were applied to the account. This prediction helps quantify the effectiveness of the action in mitigating malicious behavior.
  • Cost Model: Simultaneously, this model predicts the cost associated with applying the specific candidate action to the account. "Cost" here is broadly defined and can encompass various negative impacts, such as inconvenience to legitimate users, false positives, operational overhead, or financial expenditure.
  1. Reward Calculation: The predictions from the abuse and cost models are then combined to calculate a predicted reward for each candidate action. This reward is a weighted trade-off between the anticipated reduction in abuse and the incurred cost. The specific weights for this trade-off are dynamically determined by the MPC component, which ensures global objectives are met. A higher reward indicates a more desirable action.
  2. Action Selection and Exploration: After calculating predicted rewards for all candidate actions, PRO selects the action with the highest reward. However, to prevent the system from getting stuck in a local optima (i.e., continuously applying suboptimal actions because better ones haven't been tried), PRO incorporates an exploration mechanism. Periodically, instead of selecting the highest-reward action, PRO will intentionally try out a different, potentially less optimal, action. This exploration allows the system to discover new, more effective actions or better action-entity pairings over time, fostering continuous learning and adaptation.

Part 2: Optimizing Multipliers to Enforce Global Constraints (Model Predictive Control Component)

While the RL component focuses on optimizing actions for individual entities, the MPC component ensures that these local decisions collectively align with global objectives and constraints, particularly regarding overall abuse volume and platform-wide costs. As its name suggests, model predictive control leverages model predictions to guide the control system.

  1. Leveraging Models: The MPC component utilizes the same underlying abuse and cost models that are used in the RL part.
  2. Simulation and Trade-off Analysis: PRO runs a simulation to predict the aggregate amount of abuse and total cost that would result if different trade-offs (i.e., different weightings between abuse reduction and cost minimization) were applied. This simulation explores various scenarios, showing how different global priorities would influence the actions taken by the RL component and the resulting system-wide outcomes.
  3. Constraint Optimization: Based on these simulations, the MPC component solves a constraint optimization problem. The objective is to find the minimum possible abuse volume while ensuring that all predefined cost constraints are met. For example, a constraint might be "total false positives must not exceed X%" or "user engagement must not drop below Y%."
  4. Determining Trade-off Weights: The solution to this optimization problem yields the optimal abuse metric trade-off weights and cost metric trade-off weights. These weights are then fed back into the RL component, where they are used to calculate the reward for individual actions. By dynamically adjusting these multipliers, the MPC component effectively "steers" the RL agent, ensuring that individual action selections contribute to meeting the desired global balance between abuse reduction and cost management.

In essence, the RL component acts as the "tactician," making immediate decisions for individual accounts, while the MPC component serves as the "strategist," setting the overarching goals and constraints that guide the tactician's decisions. This powerful combination allows PRO to be both highly responsive at the individual level and globally optimized, adapting to complex, dynamic environments characteristic of online social networks.

Demo / Proof of Concept

▶ Watch: PRO Part 1: Reinforcement Learning for entity action selection (4:40)

The efficacy of the Predictive Response Optimization (PRO) system was rigorously demonstrated through its implementation on two of the world's largest social networks: Instagram and Facebook. The evaluation focused on a critical and pervasive abuse problem: data scraping, which involves bots automatically collecting data from the platforms.

For this evaluation, specific metrics were defined to quantify both abuse and cost:

  • Abuse Metric: Weighted scraping requests. This metric likely assigns different severities or importance to various types of scraping requests, allowing for a more nuanced measurement of the problem.
  • Cost Metrics:
  • Days active on the platform: This metric quantifies the potential negative impact on legitimate user engagement and retention if actions are too aggressive.
  • Feedback events: This refers to user-generated feedback, likely indicating complaints or negative experiences from users who were subjected to enforcement actions, serving as a direct measure of user friction.

PRO's performance was benchmarked against the traditional approach of rule-based action selection, which typically relies on multiple classification scores and account information to determine enforcement actions. The experimental results unequivocally highlighted PRO's superior ability to reduce abuse volume while adhering to stringent cost metric constraints.

On Instagram, PRO achieved a remarkable 59% reduction in weighted scraping requests. Crucially, this significant reduction was accomplished with no degradation in the two cost metrics (days active on the platform and feedback events), indicating that the system effectively targeted malicious activity without unduly impacting legitimate users.

Similarly, on Facebook, PRO demonstrated its capability by reducing weighted scraping requests by 4.5%. While the percentage reduction was lower than on Instagram, it's important to note that this was also achieved with no statistically significant degradation in the cost metrics, underscoring PRO's consistent performance across different platform ecosystems.

Beyond these quantitative results, the talk presented several case studies showcasing PRO's ability to adapt to a dynamic environment:

  • Case Study 1: Adjusting to Changing Business Objectives (New Cost Metrics)
  • When a new cost metric, "days active on mobile web," was introduced, PRO automatically adjusted its policies, leading to a 68% increase in that metric, demonstrating its ability to optimize for new business priorities.
  • Similarly, introducing a "dollars spent on SMS" metric saw PRO reduce this cost by 80%, showcasing its flexibility in managing diverse operational expenses.
  • Case Study 2: Introducing New Actions
  • The introduction of a new enforcement action, forced reauthentication, was seamlessly integrated. PRO learned to utilize this action effectively, resulting in a 3% reduction in automated requests with no static change in the cost metrics. This highlights PRO's capability to expand its action set and learn the utility of new tools.
  • Case Study 3: Adjusting to Uncontrolled Systemic Changes (Bugs)
  • PRO demonstrated its resilience when a bug caused an identity verification action to stop working for some users. This unexpected increase in friction was detected by PRO, which responded by dropping the selection rate of that action to zero in less than 2 days after the bug was introduced. This rapid, automated adjustment prevented further negative user impact.
  • Case Study 4: Adjusting to Adversarial Adaptation
  • A newly introduced warning action initially showed a reduction in the abuse metric. However, after approximately one month, its effectiveness waned as adversaries adapted. PRO detected this change and responded by dropping the action's selection rate from 13% down to 4%, automatically de-prioritizing an action that was no longer effective.

These demonstrations and case studies collectively illustrate PRO's robust capabilities: its effectiveness in reducing abuse, its precision in minimizing impact on legitimate users, and its crucial adaptability to evolving conditions, whether internal system changes, new business goals, or external adversarial tactics.

Defensive Implications

▶ Watch: PRO implementation on Instagram/Facebook for data scraping (8:00)

The introduction of Predictive Response Optimization (PRO) and its underlying principles carries profound implications for defenders in the online security landscape, particularly for those operating large-scale online social networks. This work offers a strategic blueprint for moving beyond reactive security measures towards a more proactive, adaptive, and intelligent defense posture.

Firstly, the core shift from detection to action selection is a critical defensive reorientation. Defenders should recognize that identifying abuse is only half the battle; the effectiveness of the response dictates the ultimate success in mitigating harm. This implies that security teams need to invest not only in sophisticated detection models but also in a diverse and flexible action set that goes beyond simple block/allow mechanisms. Actions could range from temporary rate limiting, forced captchas, re-verification steps, partial content visibility restrictions, or even subtle nudges, each with varying degrees of impact and effectiveness.

Secondly, the emphasis on explicitly defining and measuring "cost" is a crucial takeaway. Defenders often focus solely on reducing abuse, but PRO highlights the necessity of understanding the negative externalities of enforcement actions, such as user experience degradation, impact on legitimate activity, operational overhead, or even financial costs. Security teams must develop comprehensive metrics for these costs and integrate them into their decision-making processes. This encourages a more balanced approach where abuse reduction is optimized within acceptable bounds of user impact.

Thirdly, PRO underscores the imperative for rapid adaptability. The case studies vividly illustrate that adversarial tactics, system bugs, and business priorities are constantly in flux. Static, rule-based systems are inherently brittle in such environments. Defenders should explore and adopt adaptive frameworks, such as reinforcement learning and model predictive control, that can automatically learn from feedback, adjust to new conditions, and optimize enforcement policies without constant manual intervention. This enables security operations to be more resilient to zero-day attacks, adversarial shifts, and unforeseen system behaviors.

Fourthly, the technical architecture of PRO suggests that defenders should cultivate expertise in predictive modeling for both abuse and cost. Robust and accurate models are the foundation upon which PRO makes its decisions. Investing in data science capabilities to build and maintain these models—which predict the outcome of various actions—will be essential for any organization looking to implement similar adaptive enforcement systems. This also highlights the importance of collecting rich telemetry data to feed these models.

Finally, the success of PRO on Instagram and Facebook demonstrates the potential for applying advanced machine learning techniques, particularly reinforcement learning, to complex security problems. While the specific implementation might differ, the general principles of learning optimal policies in dynamic environments, balancing competing objectives, and adapting to change are broadly applicable. Defenders across various domains, not just social networks, should consider how these concepts can be leveraged to automate and optimize their defensive strategies against sophisticated and evolving threats. This could include fraud detection, network intrusion prevention, or even vulnerability management, by reframing these challenges as dynamic action selection problems.

Key Takeaways

  • Shift from Detection to Action Selection: The true goal of anti-abuse systems should be intelligently selecting enforcement actions, not just detecting abuse. This provides an additional degree of freedom to optimize outcomes.
  • Reinforcement Learning for Dynamic Enforcement: Framing action selection as a reinforcement learning (RL) problem allows systems to learn optimal policies that balance abuse reduction with the cost of impacting benign users.
  • Predictive Response Optimization (PRO) System: PRO combines RL for entity-specific action optimization with model predictive control (MPC) to enforce global constraints and adapt to changing objectives.
  • Tangible Abuse Reduction with Minimal User Impact: PRO demonstrated significant reductions in data scraping (59% on Instagram, 4.5% on Facebook) without statistically significant degradation of user experience or operational costs.
  • Automated Adaptability is Crucial: PRO automatically adjusts to new business objectives (cost metrics), new enforcement actions, systemic bugs, and, critically, adversarial adaptation, highlighting the limitations of static rule-based systems.
  • Expand the Action Set: Defenders should consider a wider range of nuanced enforcement actions beyond simple binary block/allow, allowing for more granular and effective responses tailored to specific threats and entities.

About the Speaker(s)

Garrett Wilson is presented as a key researcher involved in the development and presentation of the Predictive Response Optimization (PRO) system. His work, as detailed in this USENIX Security talk, focuses on leveraging advanced machine learning techniques, specifically reinforcement learning and model predictive control, to combat online social network abuse. The context of the talk, including the implementation of PRO on Instagram and Facebook, strongly suggests his involvement with Meta (formerly Facebook) in a research or engineering capacity related to platform integrity and security. His expertise lies in designing adaptive systems that can optimize complex trade-offs in dynamic adversarial environments.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Solid applied ML security research with real production deployments and honest metrics — this is exactly the kind of work USENIX Security exists for. The RL+MPC framing for enforcement action selection is a genuine conceptual contribution, and the case studies (adversarial adaptation detection, bug response in under 48 hours) are the kind of operational honesty that separates real systems papers from academic theater.

Heather Calloway (CISO) — SOLID

Credible applied ML research with real deployment results at scale — Instagram and Facebook aren't test beds, and a 59% scraping reduction without user impact degradation is a legitimate result. But this is platform engineering presented as security research, and the gap between what Meta built and what anyone else can operationalize is never addressed.

→ Top-rated talks at 34th USENIX Security Symposium (USENIX Security '25)

All talks from 34th USENIX Security Symposium (USENIX Security '25)