Decoding Fraud: The Evolution and Impact of Netflix's...

Aditi Gupta (Netflix), Yue Wang (Netflix)

BSidesSF 2024 · Day 1

Overview

This article delves into the critical work undertaken by Netflix's Trust and Safety team to develop a robust, multi-layered fraud metrics framework. Presented by Aditi Gupta and Yue Wang at BSidesSF 2024, the talk, "Decoding Fraud: The Evolution and Impact of Netflix's Fraud Metrics," highlights the journey from reactive, short-term incident response to a proactive, data-driven security posture. The initiative was spurred by a fundamental challenge: the inability to answer strategic questions about long-term fraud trends and the effectiveness of defensive measures, particularly in the wake of global cyber events.

Watch on YouTube

Visual summary for Decoding Fraud: The Evolution and Impact of Netflix's... by Aditi Gupta, Yue Wang
Visual summary for Decoding Fraud: The Evolution and Impact of Netflix's... by Aditi Gupta, Yue Wang

Key moments

  1. 0:00 Initial problem: Lack of long-term data for strategic security insights (DDoS example).
  2. 2:00 Existing limitations: Real-time metrics only, short data retention (2 weeks), hindering trend analysis.
  3. 4:00 Motivation for metrics: Improved visibility, reduced investigation time, and operational tuning of defenses.
  4. 7:00 Technical challenges: Data complexity (20+ pipelines, massive volume, distributed sources, implicit logic).
  5. 11:00 Solution: Three-layer metric framework (Operational, Business, C-Level) for targeted communication.
  6. 14:00 Case Study: DDoS - using anomaly detection to label attacks for long-term analysis.
  7. 19:00 Case Study: ATO - identifying 'main driver metrics' (e.g., credential stuffing) to prioritize defense efforts.
  8. 23:00 Real-world validation: Metrics confirmed blocked DDoS during activist threat, providing actionable intelligence.

Decoding Fraud: The Evolution and Impact of Netflix's...

Speakers: Aditi Gupta, Yue Wang

Conference: BSidesSF 2024

YouTube: https://www.youtube.com/watch?v=4GWehDhuKQo

Overview

This article delves into the critical work undertaken by Netflix's Trust and Safety team to develop a robust, multi-layered fraud metrics framework. Presented by Aditi Gupta and Yue Wang at BSidesSF 2024, the talk, "Decoding Fraud: The Evolution and Impact of Netflix's Fraud Metrics," highlights the journey from reactive, short-term incident response to a proactive, data-driven security posture. The initiative was spurred by a fundamental challenge: the inability to answer strategic questions about long-term fraud trends and the effectiveness of defensive measures, particularly in the wake of global cyber events.

The speakers meticulously detail the complexities involved in building such a system, from managing immense data volumes and diverse data sources to translating implicit security expertise into explicit, measurable algorithms. They introduce a three-tiered metrics framework—Operational, Business, and C-Level—designed to cater to different organizational audiences and provide actionable insights at every level. Through compelling case studies on Distributed Denial of Service (DDoS) attacks and Account Takeover (ATO) fraud, the presentation underscores how these metrics have transformed Netflix's ability to understand, mitigate, and communicate the impact of security threats, ultimately enabling more informed decision-making and resource allocation.

The significance of this work extends beyond Netflix, offering valuable lessons for any organization grappling with the challenges of measuring and responding to an ever-evolving threat landscape. By emphasizing the importance of stakeholder engagement, data consolidation, and a focus on actionable insights, Gupta and Wang provide a blueprint for building security metrics that not only track incidents but also drive strategic improvements and demonstrate the tangible return on investment for security initiatives.

Background

▶ Watch: Initial problem: Lack of long-term data for strategic security insights (DDoS... (0:00)

The impetus for developing a comprehensive fraud metrics system at Netflix arose from a specific, high-stakes scenario. Approximately two years prior to the talk, with the onset of the Ukraine-Russia conflict, Netflix, like many industries and governments worldwide, observed a significant increase in cyber attacks, particularly Distributed Denial of Service (DDoS) incidents. DDoS attacks, characterized by overwhelming infrastructure with malicious traffic, prevent legitimate users from accessing services. During this period, Netflix's leadership posed a critical question: "Is there a correlation between the recent increase in DDoS attacks and the Ukraine-Russia conflict? Are we seeing more DDoS than usual, or are we encountering new kinds of DDoS?"

The immediate answer, as Aditi Gupta explained, was "we didn't know." The existing incident response methodology relied heavily on real-time metrics, which, due to the immense volume of data, were typically persisted for only two weeks. While this allowed engineers to identify and respond to immediate spikes in traffic, indicating a DDoS event, it provided no historical context. There was no capability to analyze trends over months or years, making it impossible to answer questions about shifting attack patterns or long-term correlations. This limitation was not unique to DDoS but extended to any form of fraud or security threat the company faced.

Aditi Gupta, a Staff Security Software Engineer transitioning into an Engineering Management role, leads Netflix's DDoS strategy. Yue Wang, a Security Analytics Engineer, works alongside Gupta in the Trust and Safety team. Their team's mandate is to ensure the security and integrity of Netflix's consumer products and data, addressing issues such as account fraud, DDoS, content theft, and piracy. Recognizing the critical gap in long-term visibility and data-driven decision-making, they embarked on the ambitious project of building a robust fraud metrics framework. The core problem was a lack of foundational data and analytical capabilities to understand the evolving threat landscape, measure the effectiveness of defenses, and communicate the business impact of security incidents to various stakeholders.

Key Findings

▶ Watch: Motivation for metrics: Improved visibility, reduced investigation time, and ... (4:00)

The development of Netflix's fraud metrics framework yielded several key findings and contributions, fundamentally transforming their approach to security and fraud management:

  1. Critical Need for Long-Term Visibility: The initial trigger—the inability to correlate DDoS trends with global events—highlighted a severe deficiency in long-term data retention and analysis. Relying solely on real-time metrics with a two-week retention period meant losing crucial historical context necessary for understanding evolving threats and strategic planning.
  2. Improved Investigations and Operational Efficiency: The new metrics framework drastically reduced investigation time for impactful, unblocked DDoS incidents from hours to mere 10-15 minutes. By providing data foundations like false negative metrics, engineers could quickly pinpoint areas needing deeper analysis. Furthermore, operational metrics (e.g., false positives, false negatives) allowed the team to fine-tune defense systems, understanding which rules could be made more aggressive and which were impacting legitimate users.
  3. Significant Data Complexity: Building the metrics was far more complex than initially anticipated. It involved creating over 20 data tables or pipelines to process an immense volume of data (users, requests). Challenges included short data retention periods (some data only 3-7 days), data spread across disparate systems (e.g., request table, play bank table), and the difficulty of translating implicit human expertise into explicit, quantifiable ETL (Extract, Transform, Load) processes or algorithms.
  4. Dynamic Threat Landscape Requires Adaptable Metrics: The constant evolution of security threats necessitates continuous adaptation of defense mechanisms. This, in turn, means that the metrics used to measure these defenses must also be flexible and reconcilable with system changes, minimizing the need for constant re-engineering.
  5. Defining Success is a Balancing Act: Establishing clear success metrics for defenses proved challenging. It required balancing the trade-off between risk and growth, specifically managing false positives (blocking legitimate users) and false negatives (failing to block malicious activity). The goal was to define metrics that accurately reflect this delicate balance.
  6. Three-Layer Metrics Framework: The most significant contribution was the development of a three-tiered framework:
  • Operational Metrics: Detailed, actionable data for technical teams (e.g., real-time alerts, attack patterns, rule effectiveness).
  • Business Metrics: Translation of operational data into business insights for managers and stakeholders (e.g., cost of fraud, user impact, account takeover rates).
  • C-Level Metrics: High-level overview for executives to inform resource allocation and long-term strategy (e.g., risk exposure, ROI of prevention technologies).
  1. Data-Driven Decision Making: The framework enabled Netflix to move from anecdotal reporting ("trust me, it's increasing") to data-backed communication. For example, reporting a "3x increase in DDoS volume" alongside a "30% decrease in response time" and a "cost of $2 million USD loss" due to DDoS (using illustrative, non-real figures) provides concrete evidence for resource requests and strategic prioritization.
  2. Proactive Threat Intelligence Response: The system now allows Netflix to validate external threat intelligence. For instance, when an activist group claimed an impending attack, the metrics confirmed a significant increase in DDoS activity, which was successfully blocked by automated defenses, providing crucial insights that were previously unavailable.

Technical Deep Dive

▶ Watch: Solution: Three-layer metric framework (Operational, Business, C-Level) for t... (11:00)

The core of Netflix's fraud metrics initiative lies in its sophisticated approach to data management and a multi-layered framework designed to provide actionable insights across the organization. The technical challenges were substantial, primarily due to the sheer scale and ephemeral nature of the data.

Data Complexity and Management:

The speakers emphasized that building these metrics was not a simple task of querying an existing data table. It required the creation of over 20 distinct data tables or pipelines. These pipelines process an immense volume of data related to user requests, interactions, and system activities. A critical challenge was the data retention policy: much of the raw data was only stored for a very short period—some for just 3 days, others for 7 days. If metrics were not extracted and aggregated within this window, the valuable information was permanently lost.

Furthermore, the relevant data was spread across Netflix's vast ecosystem. Information pertinent to fraud detection might reside in a "request table" for incoming traffic, a "play bank table" for content consumption, or other specialized data stores. The engineering effort involved bringing all these disparate data sources together into a unified view suitable for metric generation.

Perhaps the most intricate technical hurdle was bridging the gap between implicit and explicit knowledge. Security experts often identify malicious activity through an intuitive understanding of patterns and anomalies that are difficult to articulate explicitly. Translating this "gut feeling" or implicit knowledge into concrete, measurable ETL (Extract, Transform, Load) processes or algorithms required rigorous definition and hard thinking. For example, a human analyst might recognize a combination of seemingly unrelated events as fraudulent, but codifying this into a system demands precise rules and thresholds.

The threat landscape is constantly changing, meaning security defenses must evolve. This dynamic environment posed a challenge for metrics, as changes in defense mechanisms could invalidate existing metrics or require significant reconciliation. The team had to design metrics in a way that minimized these reconciliation efforts, ensuring they remained relevant even as the underlying defense systems were updated.

Finally, defining success metrics was a complex balancing act. Any defense system inherently involves a trade-off between false positives (blocking legitimate users) and false negatives (allowing malicious activity). The goal was to define metrics that accurately captured this balance, considering the organization's appetite for risk versus its drive for growth. An overly aggressive defense might eliminate all risk but severely impact user experience, while a lax one would increase risk.

The Three-Layer Metrics Framework:

To address these challenges and cater to diverse organizational needs, Netflix developed a three-layer metrics framework:

  1. Operational Metrics:
  • Audience: Technical teams, engineers handling daily threats.
  • Purpose: Provide detailed, actionable data for real-time response and system tuning.
  • Examples:
  • Real-time alerts based on service health metrics or decision health metrics to quickly identify anomalies.
  • Detailed tracking of attack patterns (e.g., how many DDoS attacks utilize IP randomization or J3 hash minimization).
  • Monitoring service health metrics during attacks.
  • Analyzing response characteristics and rule effectiveness to understand how well current mitigations are performing and where to adjust defense "knobs" (e.g., making a rule more aggressive or more lenient).
  1. Business Metrics:
  • Audience: Stakeholders, product managers, and middle management.
  • Purpose: Translate operational data into business insights, showing how fraud affects business operations and objectives.
  • Examples:
  • Cost of fraud (e.g., financial losses due to account compromise).
  • User impact (e.g., number of users affected by DDoS downtime).
  • Account Takeover (ATO) rate.
  • Customer service costs incurred due to security incidents.
  1. C-Level Metrics:
  • Audience: Executives and senior leadership.
  • Purpose: Provide a high-level view of the fraud landscape to inform strategic decisions, resource allocation, and long-term planning.
  • Examples:
  • Overall risk exposure from various fraud types.
  • Return on Investment (ROI) for fraud prevention technologies and security investments.

Case Studies and Technical Implementations:

  • DDoS Detection: A significant challenge was the lack of a direct label for DDoS events. Initially, engineers manually reviewed real-time alerts for "sudden and massive spikes in traffic." To build long-term metrics, the team applied anomaly detection techniques. While acknowledging that anomaly detection isn't perfect and cannot catch every DDoS attack, it provided a foundational "label" to start with, allowing for the collection of historical data. This foundational data then fed into operational metrics (attack patterns, response times, rule effectiveness) and business/C-level metrics (user impact, cost of downtime).
  • Account Takeover (ATO): Similar to DDoS, a direct "ground truth" label for every compromised account was unavailable. For ATO, the team leveraged customer service flags (when users reported issues) and analyzed associated behavioral patterns. This data was then used to build a machine learning (ML) model to identify suspicious activities and label potential ATO events.
  • A key technical contribution here was the development of main driver metrics. Instead of just reporting ATO rates, the team focused on identifying the primary causes. For instance, they found that approximately 20% of users used weak or compromised credentials, making them vulnerable. More significantly, about 40% of successful fraudulent logins were attributed to credential stuffing attacks. This actionable insight allowed the engineering team to prioritize improvements in bot detection and the implementation of stronger Multi-Factor Authentication (MFA) to prevent ATOs. The metrics then served to re-measure the impact of these new defenses, tracking decreases in ATO rates and costs over time.

The framework ensures that all levels of the organization are informed and involved in fraud detection and mitigation, enabling a data-driven approach to security that was previously impossible.

Demo / Proof of Concept

▶ Watch: Case Study: DDoS - using anomaly detection to label attacks for long-term ana... (14:00)

While the presentation did not include a live demonstration of the Netflix fraud metrics system or a specific proof-of-concept application, the speakers effectively illustrated the framework's capabilities and impact through detailed case studies and examples of how the metrics are used in practice. The entire talk serves as a "proof of concept" for the utility and necessity of such a system.

For instance, the speakers described how the metrics framework now allows them to answer executive questions with concrete data. Instead of a vague "yes, DDoS is increasing," they can report specific figures like "we see 3x more volume increasing in DDoS, but at the same time, we also see 30% decreasing of our response time." They can identify "the top targeting path is our logging," prompting a deep dive into logging rules. For leadership, they can quantify impact, stating "because of DDoS, we see like more than 100K users got impacted, it will also cost about like 2 million US dollar lossing" (these figures were explicitly stated as illustrative, not real).

Another powerful example of the system's "proof of concept" was its ability to validate external threat intelligence. When a threat intelligence report indicated an activist group planned attacks on Netflix, the team could consult their newly established metrics. They confirmed a "significant increase in the DDoS activity on our infrastructure" on that specific day, even though their automated defenses had successfully blocked it without requiring manual intervention. This demonstrated the system's capacity to provide insights into blocked attacks, which would have been invisible under the old, real-time-only monitoring system.

Thus, while no direct "demo" was performed, the narrative and examples provided throughout the talk clearly demonstrated the framework's functionality, its ability to provide actionable insights, and its transformative impact on Netflix's security posture.

Defensive Implications

▶ Watch: Real-world validation: Metrics confirmed blocked DDoS during activist threat,... (23:00)

The implementation of Netflix's fraud metrics framework has profound defensive implications, fundamentally shifting the organization's security posture from reactive to proactive and data-driven.

  1. Enhanced Visibility and Situational Awareness: Defenders now possess a long-term, comprehensive view of the threat landscape. This allows them to understand trends, identify new attack vectors, and track the evolution of existing threats over months and years, rather than just two weeks. This improved visibility is crucial for strategic planning and anticipating future attacks.
  2. Accelerated Incident Response and Investigation: The framework significantly reduces the time required for incident investigation. By providing foundational data and metrics like false negative rates, security teams can quickly pinpoint the root cause of impactful incidents, reducing investigation time from hours to mere minutes (e.g., 10-15 minutes). This allows for faster containment and remediation.
  3. Optimized Defense Systems: Operational metrics provide granular insights into the performance of various defense mechanisms. Defenders can now objectively assess the effectiveness of individual rules and "knobs" within their systems. This enables them to make data-backed decisions on whether to make a rule more aggressive (tighten it) or more lenient (relax it) to strike the right balance between blocking malicious activity and minimizing impact on legitimate users.
  4. Data-Driven Resource Allocation and Prioritization: The ability to quantify the business impact of fraud (e.g., cost of fraud, user impact, ROI of prevention technologies) empowers security leadership to make informed decisions about resource allocation. They can justify investments in new technologies or additional personnel by demonstrating the tangible financial and user experience benefits. This also helps prioritize security initiatives, focusing on areas with the highest risk or greatest potential for improvement.
  5. Proactive Threat Intelligence Validation: The metrics framework allows Netflix to validate external threat intelligence reports with internal data. As demonstrated by the activist group example, even if automated defenses successfully block an attack without requiring manual intervention, the metrics can confirm the activity, providing valuable insights into the nature and scale of the threat. This capability enhances the organization's ability to respond proactively and verify claims.
  6. Targeted Defense Improvements: Through "main driver metrics" (as seen in the ATO case study), defenders can identify the most significant vulnerabilities or attack vectors. For instance, recognizing that credential stuffing is a major contributor to ATO allows the team to prioritize investments in bot detection and stronger Multi-Factor Authentication (MFA), leading to more effective and targeted defense improvements.
  7. Improved Communication with Stakeholders: The multi-layered framework ensures that relevant, tailored information is provided to all stakeholders, from engineers to executives. This fosters better understanding of security challenges, builds trust, and facilitates collaboration across the organization, ensuring security is considered in product launches and strategic decisions.

In essence, the fraud metrics framework equips Netflix's defenders with the tools to not only react to threats but to understand them deeply, measure their impact, and continuously improve their defenses in a strategic, data-informed manner.

Key Takeaways

  • Build a Multi-Layered Metrics Framework: Implement a tiered system (Operational, Business, C-Level) to provide relevant, actionable insights for different audiences within the organization, from engineers to executives.
  • Engage Stakeholders at All Levels: Actively communicate with and gather requirements from technical teams, managers, and leadership to ensure metrics address their specific questions and perspectives, making them truly valuable.
  • Focus on Actionable Metrics: Prioritize metrics that drive concrete actions and decisions, rather than merely providing interesting but non-actionable data points. This ensures that the effort invested in building metrics translates into tangible security improvements.
  • Anticipate and Address Data Complexity: Be prepared for significant challenges related to data volume, short retention periods, data dispersion across systems, and the difficulty of translating implicit security knowledge into explicit algorithms. Plan for robust ETL pipelines and data consolidation efforts.
  • Embrace Adaptability and Balance: Recognize that the threat landscape is dynamic, requiring metrics that can adapt to evolving defenses. Continuously balance the trade-offs between false positives and false negatives, and between risk and business growth, when defining success metrics.
  • Quantify Impact for Strategic Decisions: Leverage metrics to quantify the business impact of fraud (e.g., user impact, financial cost) and the ROI of security investments. This enables data-driven resource allocation, prioritization of security initiatives, and effective communication with leadership.

About the Speaker(s)

Aditi Gupta is a Staff Security Software Engineer at Netflix, with four years of experience at the company. At the time of the talk, she was transitioning into an Engineering Management role. Aditi leads Netflix's DDoS strategy, focusing on building scalable systems and data analytics to ensure the security of consumer products and data. Her work is integral to the Trust and Safety team, addressing challenges like account fraud, DDoS, content theft, and piracy.

Yue Wang is a Security Analytics Engineer at Netflix, working alongside Aditi Gupta in the Trust and Safety team. Her role involves leveraging data analytics to understand and combat security threats. Yue's expertise lies in translating operational security data into actionable insights and business metrics, supporting the development of the fraud metrics framework and its application in areas like DDoS and Account Takeover.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk details Netflix's journey to build a robust, multi-layered fraud and security metrics platform. Faced with a lack of historical data for strategic decision-making, the team developed a framework spanning operational, business, and C-level insights. They tackled significant data engineering challenges, including massive data volumes, short retention, and the complexity of translating implicit security knowledge into explicit metrics, demonstrating how data-driven approaches can significantly enhance defense efficacy and executive communication.

Heather Calloway (CISO) — MUST SEE

This presentation from Netflix effectively outlines the critical need for a structured approach to security metrics, moving beyond reactive incident response to data-driven strategic decision-making. By implementing a three-tiered framework (operational, business, C-level), the team successfully translated complex technical security events into actionable insights for various organizational stakeholders, including executives. This initiative directly addresses the challenge of quantifying risk, demonstrating defense efficacy, and informing resource allocation, which are paramount for robust security governance.

→ Top-rated talks at BSidesSF 2024

All talks from BSidesSF 2024