Orchestrating Resilience: Composing a New Score for Netflix Service Reliability
Sandhya Narayan (TPM · Netflix), Prachi Jain (SRE · Netflix)
BSidesSF 2026 · Day 1 · AMC Theatre 09
Overview
In "Orchestrating Resilience: Composing a New Score for Netflix Service Reliability," Sandhya Narayan and Prachi Jain, both from Netflix's Security Engineering organization, presented a compelling case for moving beyond traditional, rigid change freezes during critical operational periods. Their talk, framed around the metaphor of a symphony orchestra, detailed Netflix's journey from a reactive, "freeze everything" mentality to a sophisticated, data-driven framework that balances security, reliability, and engineering velocity. The core problem addressed was the inherent fragility and operational burden created by blanket freezes, which, while seemingly safe, paradoxically increased risk and stifled innovation.
Key moments
- 0:00 Netflix's 'blanket freeze' problem introduced
- 2:10 Why blanket freezes were the wrong instrument
- 4:10 Blanket freezes created hidden complexity and fragility
- 6:00 Realizing the problem was unmanaged risk, not change
- 6:50 The solution: Tuning controls with risk classification
- 7:25 Classifying services by their real impact and risk
- 8:45 Key questions for building a service risk scale
Orchestrating Resilience: Composing a New Score for Netflix Service Reliability
Speakers: Sandhya Narayan, TPM, Security Engineering; Prachi Jain, SRE, Security Engineering
Conference: BSides SF
YouTube: https://www.youtube.com/watch?v=Y5YCngLT7xs
Overview
In "Orchestrating Resilience: Composing a New Score for Netflix Service Reliability," Sandhya Narayan and Prachi Jain, both from Netflix's Security Engineering organization, presented a compelling case for moving beyond traditional, rigid change freezes during critical operational periods. Their talk, framed around the metaphor of a symphony orchestra, detailed Netflix's journey from a reactive, "freeze everything" mentality to a sophisticated, data-driven framework that balances security, reliability, and engineering velocity. The core problem addressed was the inherent fragility and operational burden created by blanket freezes, which, while seemingly safe, paradoxically increased risk and stifled innovation.
The speakers argued that treating every service as equally critical during high-stakes events, such as major content launches or live broadcasts, led to a host of detrimental effects, including delayed feature releases, accumulating security debt, and burnout among on-call teams. Netflix recognized this approach was unsustainable and counterproductive. Their solution involved composing a "new score" for reliability, one that leverages risk classification, service tiering, automated risk signals, and resilience tactics to enable continuous, secure deployments even during the most sensitive times. This talk offers invaluable insights for any organization struggling with the trade-offs between stability and agility in complex, distributed environments.
Background
▶ Watch: Netflix's 'blanket freeze' problem introduced (0:00)
Netflix, like many large-scale technology companies, historically relied on blanket freezes during critical periods, such as major content launches or high-profile live events. The premise was simple: stop all changes to prevent unforeseen issues. On paper, this seemed like a robust safety mechanism. However, as Prachi Jain articulated, "a lot could go wrong." This seemingly safe approach led to a cascade of "wrong notes" that ultimately undermined both security and reliability.
One significant issue was innovation slowdown. Engineering teams found themselves "composing to the calendar instead of to the customer," timing feature releases and experiments around freeze windows. This led to perfectly ready features sitting idle, delaying value delivery and demotivating teams. More critically, the freezes created a dangerous risk pileup. While deployments halted, new vulnerabilities, essential dependency updates, and last-minute fixes continued to emerge. These critical changes would accumulate behind the freeze, leading to a "giant release wave" once the freeze was lifted. This wasn't safety; it was "deferred chaos," turning what should have been a quiet period into a tense, high-risk post-freeze deployment frenzy. The result was significant operational pain, characterized by noisy launches, exhausted on-call teams, and a pervasive "Please don't page me tonight" energy, all without a corresponding increase in confidence.
The blanket freeze approach suffered from several fundamental flaws. First, it actively blocked critical fixes and security patches. High-priority security vulnerabilities would land during a quiet period, but the "no changes" rule meant teams were forced to knowingly operate with elevated risk, unable to deploy essential patches. This created an absurd situation where the very mechanism designed for safety worked against security. Teams resorted to "slacking us like, 'Hey, is there a way we can push this out or sneak it in?'" – a clear indicator that the control was failing.
Second, the "one size fits all" assumption was deeply problematic. A minor, low-risk configuration tweak was treated with the same severity as a high-risk architectural change. This indiscriminate approach meant that actual high-risk changes weren't receiving the focused attention they needed, while low-risk changes were unnecessarily impeded. When controls are overly restrictive and lack nuance, teams inevitably find workarounds. This led to "shadow deploys, side channels, we'll just flip it after hours" – changes that were harder to observe, secure, and operate. Instead of clean, well-planned deployments, Netflix ended up with hidden complexity and surprise behavior, effectively "reshaping" risk rather than reducing it. The critical realization was that "the problem wasn't the change itself. It was unmanaged risk. Instead of stopping the music, we needed to retune our instruments. We didn't need more no's. We needed a smarter how."
Key Findings
▶ Watch: Blanket freezes created hidden complexity and fragility (4:10)
The pivotal realization for Netflix was that their default assumption—that "everything is critical, freeze everything"—was fundamentally flawed. This was akin to treating every instrument in an orchestra as a soloist, leading to cacophony rather than harmony. In reality, Netflix's vast ecosystem of services plays diverse roles, with varying levels of criticality and impact. This insight paved the way for the core idea of risk classification: tuning controls to the real risk of each service, rather than applying a blanket policy.
The speakers introduced a systematic approach to classify risk, ensuring that not every service was treated as a "mission critical" component. This involved anchoring decisions on three simple, yet profound, questions for each service:
- What is the impact on our customer experience? This question directly assesses whether a service failure would be immediately felt by a Netflix member (e.g., inability to press play, log in, or make payments). Services with direct, immediate customer impact are inherently higher risk.
- What is the blast radius during critical events? This considers the scope of impact if a service misbehaves during a major premiere or live event. Would it take down the entire experience or just a segment? A wider blast radius necessitates greater caution.
- What is our acceptable risk threshold here? Some services can tolerate more experimentation and change during critical periods than others. This question helps define the appropriate level of flexibility or stringency for controls.
By systematically answering these three questions, Netflix gained the clarity to differentiate between services that "can keep shipping" and those that "need tighter guardrails." This fundamental shift allowed them to move away from the unsustainable "one giant no changes rule" and instead develop a nuanced understanding of their operational landscape. This risk classification became the bedrock for their subsequent service tiering model, enabling them to allocate resources and apply controls proportionate to the actual risk posed by each service. The key finding was that by explicitly defining and understanding service criticality, Netflix could make far more intelligent and effective decisions about change management, ultimately enhancing both security and reliability without sacrificing agility.
Technical Deep Dive
▶ Watch: Realizing the problem was unmanaged risk, not change (6:00)
Netflix translated its risk classification into practical application through a service tiering model and a framework for integrating risk signals into deployment decisions. The tiering model acts as a "volume control" for each service, dictating the level of scrutiny and control required during critical events.
The service tiering model is structured as follows:
- Tier 0 (Lead Instruments): Defined as services where "if this breaks, the user definitely feels it." These are the most critical components directly impacting the core streaming experience. Examples include playback, sign-up, payment systems, and the underlying control planes that ensure their functionality. Failures in Tier 0 services are "instantly visible to our members." Consequently, these services are subjected to the strictest guardrails, as they are the "lead instrument[s]" that cannot be out of tune.
- Tier 1 (First Chairs): These services are not in the main spotlight but lead their respective sections. If they are "off, everyone notices." This tier includes services like discovery, recommendation engines, and key UI flows. While not as immediately catastrophic as Tier 0 failures, their malfunction significantly degrades the member experience. Tier 1 services receive strong guardrails, but with "a little bit more flexibility" compared to Tier 0.
- Tier 2 (Supporting Section): Comprising services crucial for internal operations, platform functionality, shared infrastructure, and engineer productivity. Their impact on members is typically indirect or delayed. A bad day for a Tier 2 service might manifest later through slower deployments, more bugs, or degraded experiences. During critical events, Tier 2 services can often "keep shipping with some guardrails," as their blast radius is more buffered. Controls for this tier focus on stability and good hygiene rather than absolute change prevention.
- Tier 3 (Backstage Section): These are internal, low-impact services, such as reporting tools, internal utilities, or low-risk applications, where "members don't care about them" directly. Failures in Tier 3 services have minimal real-time impact on the member experience. This tier is considered the safest to ship even during big moments, effectively preserving velocity for teams working on less critical components.
Once services are tiered, decisions around risk and deployment become significantly clearer. However, tiering alone is insufficient; it needs to be dynamically informed by risk signals. Netflix leverages four core signals to drive real-time deployment decisions:
- Deployment Confidence: This signal assesses the trustworthiness of the deployment pipelines. It considers not only changes made by the service's primary team but also those suggested by others (e.g., dependency updates, stack updates). High confidence unlocks automation, while low confidence triggers mitigations.
- Test Coverage: Acting as an "early warning system," this signal looks at test pass rates, the presence of flaky tests, and overall code coverage. Noisy or degraded test signals indicate "holes" in the safety net, advising against taking risks during critical periods.
- Change Type: A powerful yet simple signal that differentiates between various types of changes. A reversible feature toggle receives a "lighter touch" compared to a complex database migration affecting millions of rows. This helps determine the potential blast radius and the appropriate level of review and rollout control.
- Historical Behavior: This signal examines how a service has behaved in the past during critical events. Was it noisy? Were there too many incidents? Were previous deployments healthy or unhealthy? A pattern of instability around launches leads to more conservative deployment strategies, such as smaller batch sizes, tighter deployment windows, or manual checkpoints.
Together, these signals provide a comprehensive "story about how you should be playing out your deployment." By combining service tiers with these dynamic risk signals, Netflix moved from "vibes" and gut feelings to data-driven mitigation strategies, enabling faster, safer, and more predictable deployments during even the most high-stakes "performances."
Demo / Proof of Concept
▶ Watch: Classifying services by their real impact and risk (7:25)
While the talk did not feature a live, interactive demo in the traditional sense, Sandhya Narayan and Prachi Jain thoroughly described Netflix's risk-aware launch rubric as the practical manifestation and "proof of concept" for their orchestrated resilience framework. This rubric is not a manual checklist or a hallway conversation; it is wired into their CI/CD pipelines, making it an automated and consistently applied decision-making system.
The risk-aware launch rubric functions like a dynamic scorecard, integrating all the elements discussed:
- Event Type: The criticality of the current operational moment (e.g., major launch, live event).
- Service Tier: The classification of the service (Tier 0, 1, 2, or 3) based on its impact and blast radius.
- Risk Level (derived from signals): Aggregated data from deployment confidence, test coverage, change type, and historical behavior.
- Resilience Tactics: An assessment of whether appropriate resilience mechanisms (canaries, regional staggering, synthetic tests) are built into the service's pipeline and deployment process.
The rubric processes these inputs to generate a clear, unambiguous "Can you bypass the freeze?" decision. If a deployment meets the defined criteria for its specific tier and the current event context, the rubric automatically grants permission. Conversely, if criteria are not met, it either blocks the deployment or "nudges" the team to implement additional guardrails or mitigations. This ensures that "the system applies the rubric consistently the same way and provide[s] us data-driven safe deployments."
This integration into the pipelines is crucial. It means that risk awareness is embedded directly into the engineering workflow, rather than relying on tribal knowledge, manual reviews, or ad-hoc discussions. The rubric effectively enforces the "score" for safe deployments, ensuring that decisions are based on understood risk, engineered mitigations, and verifiable data. This automated enforcement mechanism serves as the robust proof of concept for Netflix's ability to maintain high velocity and secure operations simultaneously, even without traditional blanket freezes.
Defensive Implications
▶ Watch: Key questions for building a service risk scale (8:45)
The framework presented by Netflix offers profound defensive implications for organizations seeking to enhance their security posture and operational resilience. By moving away from blanket freezes, Netflix has effectively traded a brittle, reactive defense for a proactive, adaptive strategy that builds security and reliability into the very fabric of its deployment processes.
A core defensive strategy lies in the three main resiliency tactics that are actively integrated into their pipelines:
- Canaries: These are used to verify changes in code or configuration without immediately impacting the entire user base. By directing a "tiny slice of your traffic" to the new version and meticulously monitoring metrics and error rates, Netflix can detect and contain potential issues before they escalate. This significantly reduces the blast radius of faulty deployments, acting as an early warning system and providing a crucial safety net.
- Regional Staggering: Instead of a global rollout, deployments are staggered region by region. This ensures that "if something goes wrong, the blast radius is contained." Teams can pause, roll back, or "fix forward" within a single region before an issue becomes a global incident. This distributed deployment strategy is a powerful defense against widespread outages.
- Synthetic Testing: Critical user journeys (e.g., sign-up, payment, playback) are continuously tested using synthetic transactions. These "always-on listeners" proactively identify breaks in core functionality, allowing Netflix to fix issues before members even experience them. This provides an external, objective validation of system health, acting as a critical last line of defense for customer experience.
Crucially, these tactics are not optional "nice-to-haves" or tribal knowledge; they are "wired into our pipelines." For specific services and tiers, the CI/CD pipeline "nudges you towards, 'Hey, use a canary. Roll out regionally. Make sure you have synthetic checks built in.'" This embedded risk awareness ensures that defensive measures are consistently applied and automated, rather than relying on individual memory or manual oversight.
Beyond these specific tactics, the entire framework embodies a philosophy of continuous improvement. After every major event, Netflix conducts post-event reviews to analyze "What worked? What broke? What was just unnecessarily noisy?" This leads to signal and tier tuning, where service tiers are re-evaluated (e.g., "Did that tier two service behave like a tier zero? Was it mis-tiered?") and risk signals are adjusted. Finally, playbooks and pipelines are updated. If a specific type of change requires a canary or synthetic test, this requirement is "wired into our CI/CD flow," not just documented. This iterative learning cycle ensures that "every big event leaves the system a little smarter than before and not just tired humans."
Furthermore, stakeholder alignment through clear communication and shared metrics (reliability, security, and velocity) is a defensive strength. By transparently sharing "what's safe to ship? What's paused? Why?" and providing "after action summaries," Netflix builds trust in the framework and ensures that all teams are operating with a unified understanding of risk and mitigation. This holistic approach empowers engineering teams to maintain velocity securely, significantly enhancing Netflix's overall defensive posture against operational and security risks.
Key Takeaways
The Netflix talk on Orchestrating Resilience offers several critical takeaways for any organization grappling with the complexities of security, reliability, and agility:
- Embrace Risk-Based Decisions: Not every service is equally critical. Treating every component as a Tier 0 "play button" service leads to team burnout and can obscure the truly significant risks. Instead, explicitly define impact and blast radius, then apply the strictest controls where they genuinely matter.
- Automate Risk Signals: Move beyond gut feelings and "we've always done it this way" decision-making. Leverage data from deployment confidence, test coverage, change type, and historical behavior to drive deployment decisions. Ground "Can we ship?" answers in metrics and verifiable signals.
- Utilize Resilience Tactics: Canaries, regional staggering, and synthetic tests are not optional extras. They are fundamental mechanisms for shipping securely during critical periods. Integrate these tactics directly into your CI/CD pipelines to build safety nets that empower continuous deployment without blanket freezes.
- Foster Continuous Learning and Adaptation: Implement a feedback loop that includes post-event reviews, tuning of service tiers and risk signals, and updating playbooks and pipelines. Every incident and major event should make the system smarter and more resilient for future performances.
About the Speaker(s)
Prachi Jain is an SRE (Site Reliability Engineer) in the Security Engineering organization at Netflix. Her role is focused on ensuring that Netflix's services remain secure and reliable, particularly during high-stakes "big moments" such as major content launches or live events.
Sandhya Narayan is a TPM (Technical Program Manager) also within Netflix's Security Engineering. She partners with various engineering teams to facilitate smooth, secure, and predictable launches, ensuring that new features and services are introduced without compromising the company's security or reliability standards.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
A competent, well-structured case study from Netflix on replacing blanket change freezes with a tiered, signal-driven deployment framework. Clean presentation, real operational context, and honest about the failure modes of the old approach — but the substance is incremental process engineering, not novel research, and most of the framework components (canaries, regional staggering, service tiering) are industry-standard SRE practice.
Heather Calloway (CISO) — SOLID
A well-constructed operational talk that solves a real problem — blanket freezes creating more risk than they prevent — with a credible, automated framework. Strong engineering execution, but it never surfaces the governance layer: who owns the tier classifications, who's accountable when the rubric is wrong, and what this means for security programs beyond Netflix's scale.