Unleashing the Power of Init Containers: Reducing Database Management To... Muhammad Junaid Muzammil
Muhammad Junaid Muzammil
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this insightful talk from KubeCon EU 2025, Muhammad Junaid Muzammil, a Tech Lead in the Database Reliability Engineering group at Yelp, illuminated the often-underappreciated power of Kubernetes Init Containers. He demonstrated how Yelp has creatively leveraged this foundational Kubernetes feature to significantly reduce the operational overhead associated with managing their extensive Apache Cassandra clusters and to enhance overall engineering efficiency. The presentation delved into three critical operational challenges—horizontal cluster scaling, database upgrades, and recovery from backups—showcasing how Init Containers provided elegant and automated solutions.

Key moments
- 0:00 Introduction to Init Containers and talk agenda
- 2:45 Yelp's extensive use of Cassandra at scale
- 4:00 Cassandra on Kubernetes and custom operator setup
- 6:00 Understanding Kubernetes Init Containers: purpose and differences
- 8:20 Challenge: Cassandra horizontal scaling with CDC enabled
Unleashing the Power of Init Containers: Reducing Database Management To... Muhammad Junaid Muzammil
Speakers: Muhammad Junaid Muzammil, Tech Lead, Database Reliability Engineering, Yelp
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=nTmwmd4fcGI
Overview
In this insightful talk from KubeCon EU 2025, Muhammad Junaid Muzammil, a Tech Lead in the Database Reliability Engineering group at Yelp, illuminated the often-underappreciated power of Kubernetes Init Containers. He demonstrated how Yelp has creatively leveraged this foundational Kubernetes feature to significantly reduce the operational overhead associated with managing their extensive Apache Cassandra clusters and to enhance overall engineering efficiency. The presentation delved into three critical operational challenges—horizontal cluster scaling, database upgrades, and recovery from backups—showcasing how Init Containers provided elegant and automated solutions.
Muzammil's talk highlighted that even seemingly small, well-executed features, when applied thoughtfully, can yield substantial impacts. For a company operating at Yelp's scale, with hundreds of Cassandra clusters and thousands of nodes, streamlining such processes translates directly into improved reliability, reduced human intervention, and faster incident response. This article will explore the technical details of Yelp's innovative approaches, providing a deep dive into how Init Containers were instrumental in transforming complex, manual database operations into seamless, automated workflows.
The core message of the presentation revolved around the strategic use of Init Containers to perform essential setup tasks and pre-checks before the main application containers start. By carefully orchestrating these preliminary steps, Yelp was able to mitigate common pitfalls associated with stateful workloads like Cassandra, such as data duplication during scaling, version incompatibility during upgrades, and lengthy recovery times. This talk serves as a compelling case study for any organization looking to optimize the management of their stateful applications on Kubernetes.
Background
▶ Watch: Introduction to Init Containers and talk agenda (0:00)
Yelp operates a massive community-driven platform, connecting millions of users with local businesses. As of December 2024, the platform boasted over 308 million cumulative reviews and served 29 million average monthly unique users. Given this immense scale, the underlying data infrastructure is critical. Apache Cassandra, a wide-column distributed non-relational database, plays a pivotal role at Yelp, serving portions of data for nearly every page visit. The company manages over 70 production Cassandra clusters, comprising more than a thousand nodes, storing several hundred terabytes of data. Many of these clusters are highly latency-sensitive, requiring sub-10 millisecond latencies for both read and write traffic.
Operationally, all of Yelp's Cassandra clusters run on Kubernetes, a transition made approximately five to six years ago from a Mesos-based infrastructure. This shift necessitated robust tooling and automation. Yelp utilizes an in-house Platform as a Service (PaaS) named Bastard, which provides an abstraction layer over Kubernetes for service management. Critical to Cassandra operations is a custom Cassandra operator, which handles all Kubernetes interactions for the clusters. Configurations are managed declaratively via a central Git repository called Yelp-s-config. Any configuration updates are converted into a Kubernetes Custom Resource, which the operator continuously watches. Upon detecting a difference between the desired and actual state, the operator reconciles by translating these changes into relevant Kubernetes resources. Persistent volumes are used for the stateful Cassandra workloads, attached to the Kubernetes pods. It's also important to note that Yelp's clusters are spread across multiple data centers and regions, maintaining a replication factor of three within a single data center.
A fundamental component of Kubernetes, and central to this talk, is the Pod. A Pod is the smallest deployable unit, essentially a group of containers sharing the same network namespace and running on the same Kubernetes node. While main containers within a Pod start in parallel, there are often scenarios requiring preparatory steps before the primary application begins execution. This is where Init Containers become indispensable. Init Containers are specialized containers designed to ensure everything is set up correctly before the main application container starts. Key differences from regular containers include:
- Sequential Execution: Init Containers always run one after another, in the order defined.
- Completion Requirement: Each Init Container must run to completion successfully before the next Init Container or the regular containers can start.
- Setup Focus: They are designed exclusively for setup tasks and do not serve user traffic, meaning they do not support readiness or liveness probes.
Key Findings
▶ Watch: Yelp's extensive use of Cassandra at scale (2:45)
Yelp's engineering team identified three major operational challenges in managing their Apache Cassandra fleet on Kubernetes, each of which was effectively addressed by the strategic application of Init Containers:
- Horizontal Cluster Scaling with Change Data Capture (CDC) Enabled: Scaling Cassandra clusters horizontally, especially when CDC is active, introduced complexities such as duplicate events in downstream systems (e.g., Kafka) and prolonged bootstrap times due to extra operations during data streaming. Init Containers were used to temporarily disable CDC during the node bootstrapping phase, preventing these issues.
- Seamless Cassandra Cluster Upgrades (3.11 to 4.1): Upgrading Cassandra clusters presented two significant hurdles. Firstly, the dynamic IP assignments in Yelp's Kubernetes environment, combined with rolling upgrades, led to issues where new nodes with changed IPs and versions were not recognized by older peers. Secondly, managing SS Table format versions during upgrades posed challenges related to disk pressure and control over the conversion process. Init Containers provided a mechanism to gradually introduce changes (either IP or version) and to control the SS Table upgrade process to mitigate these problems.
- Automated Cluster Recovery from Backups: While disaster recovery is a rare event, its urgency demands a highly automated and expedited process. Restoring large Cassandra clusters from S3 backups using a sequential approach would be prohibitively slow. Init Containers were leveraged to parallelize the data pulling process across all nodes, drastically reducing recovery time and ensuring a fully automated restore workflow.
These findings collectively demonstrate that Init Containers are not just a minor Kubernetes feature but a powerful primitive for orchestrating complex stateful workload lifecycle management, enabling significant operational efficiencies and reliability improvements for critical database systems.
Technical Deep Dive
▶ Watch: Cassandra on Kubernetes and custom operator setup (4:00)
The core of Muhammad Junaid Muzammil's talk detailed the technical implementations of Init Containers to solve the three identified Cassandra operational challenges.
Horizontal Scaling with Change Data Capture (CDC)
When a new Cassandra node joins a cluster, it undergoes a bootstrap process. It interacts with a seed node to learn the cluster topology, token range assignments, and then streams data from its peers. This streaming can be time-consuming for large datasets. The primary challenge arose when Change Data Capture (CDC) was enabled on the cluster. CDC detects data changes, converting them into a stream (at Yelp, these events are published to Kafka). Under the hood, Cassandra creates commit log files for every write. With CDC enabled, these commit logs are replicated to a special CDC raw directory, from which independent consumers read and process the events.
The problem with bootstrapping a new node with CDC enabled is that the streamed bootstrap data itself generates CDC events, leading to unnecessary duplicate events in Kafka and additional processing cycles for deduplication. Furthermore, the bootstrap process becomes slower due as Cassandra performs extra operations for every streamed data point.
Yelp's solution leveraged an Init Container:
- Pod Disruption Budgets (PDBs): To minimize voluntary disruptions during scaling, the
maxUnavailableproperty in PDBs was appropriately set. This is crucial as data streaming can be lengthy. - Idempotency Check: Each time a pod restarts, the Init Container runs. To prevent re-streaming data unnecessarily, the Init Container first checks for the presence of data on the attached persistent volumes. If data exists, it implies the node was already bootstrapped, and the Init Container exits quickly. This handles restarts without re-initiating the entire process. Interruptions due to hardware failures are handled by on-call personnel resetting storage, as such events are rare.
- Cassandra with CDC Disabled: Inside the Init Container, the Cassandra process is started, but crucially, with the
CDC enabledconfiguration property set tofalse. This ensures that any data streamed during the bootstrap process does not generate CDC events for Kafka. - Data Loss Mitigation: While the node is bootstrapping with CDC disabled, new writes might theoretically be missed by this specific node. However, given Yelp's replication factor of three, data integrity is maintained because at least two other nodes are processing the same data, preventing any data loss.
- Readiness Script and Graceful Termination: A readiness script indicates when data streaming is complete and Cassandra has joined the ring. Once a "green signal" is received, the Cassandra process within the Init Container is gracefully stopped. Graceful termination is paramount for stateful workloads to prevent data corruption. The Init Container then exits.
- Main Container with CDC Enabled: After the Init Container completes, the main application container starts Cassandra, this time with
CDC enabledset totrue. From this point onward, the node processes new CDC events.
This careful orchestration allowed Yelp to horizontally scale Cassandra clusters with minimal human involvement, reduced duplicate CDC events, and faster scale-up times.
Cassandra Upgrade (3.11 to 4.1)
Upgrading Yelp's extensive Cassandra fleet from version 3.11 to 4.1 presented two distinct challenges, both addressed using Init Containers.
Problem 1: Dynamic IPs and Version Mismatch During Rolling Upgrades
Yelp's Kubernetes environment uses dynamic IP assignments, meaning a pod receives a new IP from the pool upon restart. During a standard rolling upgrade (one node at a time), when a 3.11 node was restarted and upgraded to 4.1, it received a new IP. The problem was that the new 4.1 node, with its changed IP, was not accepted as a replacement by its 3.11 peers. This led to exceptions during the gossip exchange mechanism, causing the upgrade to fail. Experiments revealed that if both IP and version changed simultaneously, the node was not recognized. However, if these attributes changed gradually, the process worked.
The Init Container solution involved:
- Gradual Attribute Change: After a pod restart (which assigns a new IP), the Init Container starts. Inside it, Cassandra is launched with the older 3.11 version.
- Peer Recognition: By starting with the old version, the remaining 3.11 peers recognize the new IP as a replacement for the old one, as only one attribute (IP) has changed.
- Graceful Termination: The Init Container waits until the Cassandra node is ready to serve traffic, then gracefully terminates the 3.11 workload and exits.
- Main Container with New Version: Because the network stack is shared across containers within a pod, the IP assignment remains consistent. The main container then starts Cassandra with the new 4.1 version.
This "two-step" approach, using the Init Container to bridge the version gap, allowed the upgraded node to successfully rejoin the ring, maintaining a consistent nodetool status across the cluster. This procedure was successfully applied across Yelp's 70+ clusters and over a thousand nodes.
Problem 2: SS Table Format Version Management
Cassandra's SS Table format versions can change with major database versions (e.g., ME in 3.11.13 to NV in 4.1). While adjacent major versions typically support reading older SS Tables, there are no performance guarantees, making an upgrade to the latest format recommended. In pre-4.x versions, this required manually invoking the nodetool upgrade command. Cassandra 4.x introduced some automated ways, but these offered less control over how many nodes performed the upgrade simultaneously. This lack of control could lead to disk pressure, as upgrading SS Tables involves rewriting each table into the newer format, which can be I/O intensive.
Yelp's Init Container solution for SS Table upgrades:
- Triggering the Upgrade: A spec change in Yelp-s-config (converted to a Custom Resource) triggers a stateful set change, leading to a rolling restart of pods.
- Init Container Check: The Init Container starts and first checks if any old-formatted SS Table versions are present on disk. If not, it exits immediately.
- Controlled SS Table Upgrade: If old formats are found, the Init Container starts the Cassandra process and waits for it to become ready. Once ready, it invokes the
upgrade SSTablescommand. This process can be long-running depending on data volume. - Idempotency and Disk Pressure: The Init Container ensures idempotency, meaning if the pod is terminated, it can resume the upgrade process from where it left off. By performing this upgrade within the Init Container during a rolling restart, Yelp effectively limited the disk pressure to a single node at any given time, ensuring seamless upgrades without performance degradation.
Cluster Recovery from Backups
Disaster recovery, though infrequent, demands extreme urgency and automation. Yelp uses Medusa, an open-source tool, to copy SS Table formatted data from disk to S3. Medusa also stores manifest files containing metadata (timestamp, backup type). Yelp supports both full and differential backups, favoring the latter.
The challenge was to restore a large cluster quickly. A sequential node-by-node restore process would take an unacceptably long time.
The Init Container solution for recovery:
- New Cluster Configuration: A new Cassandra cluster is configured with additional properties specifying the backup identifier (a timestamp in Yelp's case).
- Parallel Data Pulling: When this new cluster is created, Init Containers are started on all pods. Crucially, Yelp uses a parallel pod management policy for the stateful set, allowing all pods to come up and start pulling data from S3 simultaneously.
- Data Pull and Exit: Each Init Container pulls the relevant SS Table data from S3 onto its attached persistent volume. Once all nodes have completely pulled their data, the Init Containers exit.
- Cluster Formation: The main containers then start the Cassandra processes. With all data already on disk, the Cassandra nodes quickly join to form the restored cluster.
This approach ensures a fully automated and significantly expedited restore process, critical for meeting disaster recovery objectives. The automation built around Init Containers helps ensure the recovery process remains functional even during rare, high-pressure events.
Demo / Proof of Concept
▶ Watch: Understanding Kubernetes Init Containers: purpose and differences (6:00)
While the talk provided detailed explanations and diagrams illustrating the technical solutions, Muhammad Junaid Muzammil did not include a live demonstration or a specific mention of a proof-of-concept during his presentation. The solutions were described conceptually, drawing from Yelp's real-world implementation experience across their extensive Cassandra fleet.
Defensive Implications
▶ Watch: Challenge: Cassandra horizontal scaling with CDC enabled (8:20)
The innovative use of Init Containers at Yelp offers several crucial defensive implications for organizations managing stateful workloads on Kubernetes:
- Embrace Init Containers for Lifecycle Management: Defenders should recognize Init Containers as a powerful tool for orchestrating complex setup, pre-check, and post-termination tasks for stateful applications like databases, message queues, and distributed caches. This moves complex logic out of the main application container, making it cleaner and more focused.
- Prioritize Idempotency and Failure Handling: Any logic within an Init Container must be idempotent, meaning it can be run multiple times without unintended side effects. Robust failure handling is also critical, as Init Containers can be restarted. Implementing checks (e.g., presence of data on PVs) to conditionally execute logic prevents unnecessary operations and ensures resilience against transient failures.
- Leverage Pod Disruption Budgets (PDBs): For long-running Init Container processes, such as data streaming or SS Table upgrades, PDBs are essential. Properly configured PDBs minimize voluntary disruptions during operations like cluster scaling or node restarts, giving Init Containers sufficient time to complete their critical tasks gracefully and prevent cascading failures.
- Ensure Graceful Shutdowns: Abrupt termination of stateful workloads can lead to data corruption. Defenders must ensure that any process started within an Init Container (like Cassandra itself during a CDC-disabled bootstrap or SS Table upgrade) is always terminated gracefully before the Init Container exits. This requires careful scripting within the container.
- Automate Complex Operations: The examples of scaling, upgrading, and restoring Cassandra highlight how Init Containers can automate traditionally manual and error-prone database operations. This reduces human intervention, increases operational efficiency, and improves reliability, particularly for high-urgency scenarios like disaster recovery.
- Custom Operators for Orchestration: While Init Containers provide the primitive, their effective use often benefits from a custom Kubernetes operator. As seen with Yelp's custom Cassandra operator, this allows for declarative management of complex stateful workloads, translating high-level desired states into the precise orchestration of pods and Init Containers.
- Consider Parallelism for Recovery: For disaster recovery scenarios, where time is of the essence, leveraging parallel pod management policies with Init Containers to fetch data simultaneously across multiple nodes can drastically reduce recovery time objectives (RTOs).
By adopting these principles, organizations can build more robust, resilient, and automated systems for managing their critical stateful applications on Kubernetes.
Key Takeaways
- Init Containers are Powerful Orchestration Primitives: Don't underestimate the utility of Init Containers for managing complex setup, pre-check, and cleanup tasks for stateful workloads on Kubernetes.
- Idempotency and Failure Handling are Crucial: Always design Init Container code to be idempotent, capable of running multiple times without adverse effects, and include robust failure handling mechanisms to cope with unexpected terminations or restarts.
- Conditional Execution Based on Persistent State: For many use cases, Init Container logic should conditionally execute based on the persistent state on disk (e.g., checking if data is already present) to avoid redundant operations upon pod restarts.
- Minimize Disruptions with Pod Disruption Budgets (PDBs): Proactively use PDBs to control voluntary disruptions, especially when Init Containers are performing long-running tasks like data streaming or large-scale data transformations, ensuring graceful completion.
- Graceful Exit Prevents Data Corruption: Always ensure that any stateful application process started within an Init Container is gracefully terminated before the Init Container exits to prevent data corruption and maintain data integrity.
- Automate Complex Stateful Workload Management: Init Containers enable significant automation for challenges like horizontal scaling with CDC, seamless database upgrades, and expedited disaster recovery, reducing human toil and improving system reliability.
About the Speaker(s)
Muhammad Junaid Muzammil is a Tech Lead in the Database Reliability Engineering (DRE) group at Yelp. His primary focus within this role is on NoSQL databases, with a particular emphasis on managing and optimizing Apache Cassandra. As demonstrated in his KubeCon EU talk, he is deeply involved in enhancing the operational efficiency and reliability of Yelp's extensive Cassandra fleet, leveraging Kubernetes features like Init Containers to automate complex database management tasks. His work contributes to maintaining the performance and availability of critical data infrastructure that underpins Yelp's platform.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk from a Yelp DRE Tech Lead masterfully showcases how a fundamental Kubernetes feature, Init Containers, can be leveraged to elegantly solve three critical and complex operational challenges in managing large-scale Apache Cassandra clusters: horizontal scaling with CDC, seamless version upgrades, and automated disaster recovery. Muzammil provides a deep dive into the technical intricacies, demonstrating highly practical and idempotent solutions that significantly reduce manual overhead and enhance reliability for stateful workloads. The insights are directly applicable to any organization grappling with similar challenges on Kubernetes.
Heather Calloway (CISO) — STRONG ACCEPT
This session from Yelp's DRE team offers a clear, unsentimental look at how Init Containers fundamentally improve the operational resilience and efficiency of critical stateful workloads like Apache Cassandra. The speaker precisely dissects complex challenges in scaling, upgrading, and recovering large-scale databases, providing concrete, actionable solutions that reduce human error and accelerate incident response. While primarily technical, the implications for business continuity, data integrity, and reducing operational risk are undeniable and critical for executive leadership to understand.