Managing Data at Scale: Best Practices and Evolution of SIG-Apps - Maciej Szulik & Janet Kuo

Maciej Szulik, Janet Kuo

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In this comprehensive KubeCon EU presentation, Maciej Szulik, a lead for the Kubernetes Special Interest Group (SIG) for Applications (SIG-Apps), provided an in-depth look into the group's recent achievements, ongoing projects, and future roadmap. While co-leads Janet Kuo and Ken were unable to join, Szulik adeptly guided the audience through the crucial role SIG-Apps plays in defining and evolving the core workload controllers that underpin nearly all applications running on Kubernetes.

Watch on YouTube

Visual summary for Managing Data at Scale: Best Practices and Evolution of SIG-Apps - Maciej Szulik & Janet Kuo by Maciej Szulik, Janet Kuo
Visual summary for Managing Data at Scale: Best Practices and Evolution of SIG-Apps - Maciej Szulik & Janet Kuo by Maciej Szulik, Janet Kuo

Key moments

  1. 0:00 Introduction to SIG Apps and how to engage
  2. 2:00 Understanding SIG Apps' scope and core controllers
  3. 3:00 SIG Apps annual reports and historical evolution
  4. 4:50 New PDB behavior: counting only ready pods
  5. 6:30 Random pod selection for ReplicaSet downscaling
  6. 7:50 StatefulSet start ordinal for cluster migration

Managing Data at Scale: Best Practices and Evolution of SIG-Apps

Speakers: Maciej Szulik, SIG-Apps Lead; Janet Kuo, SIG-Apps Lead

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=KlsxQMfdKLw

Overview

In this comprehensive KubeCon EU presentation, Maciej Szulik, a lead for the Kubernetes Special Interest Group (SIG) for Applications (SIG-Apps), provided an in-depth look into the group's recent achievements, ongoing projects, and future roadmap. While co-leads Janet Kuo and Ken were unable to join, Szulik adeptly guided the audience through the crucial role SIG-Apps plays in defining and evolving the core workload controllers that underpin nearly all applications running on Kubernetes.

The talk highlighted the significant advancements made in enhancing the reliability, flexibility, and scalability of fundamental Kubernetes resources such as Deployments, Jobs, CronJobs, DaemonSets, and StatefulSets. Szulik emphasized the group's commitment to addressing real-world user challenges, particularly those related to managing data-intensive applications and high-performance computing (HPC) workloads at scale. This presentation is essential for anyone involved in operating, developing, or architecting applications on Kubernetes, offering valuable insights into current best practices and the strategic direction of core platform capabilities.

The importance of SIG-Apps' work cannot be overstated. As Kubernetes continues to mature and be adopted for increasingly complex use cases, the stability and feature set of its core application controllers directly impact the platform's ability to meet enterprise demands. The discussion covered critical improvements in areas like pod disruption management, stateful workload migrations, batch processing, and resource lifecycle policies, underscoring the continuous effort to refine Kubernetes as a robust and adaptable container orchestration system.

Background

▶ Watch: Introduction to SIG Apps and how to engage (0:00)

SIG-Apps is one of the most foundational Special Interest Groups within the Kubernetes project, directly responsible for the design, development, and maintenance of the core API resources and controllers that manage application workloads. This includes ubiquitous constructs like Deployments, Jobs, CronJobs, DaemonSets, and StatefulSets. Essentially, if you're running an application on Kubernetes, you are interacting with the resources managed by SIG-Apps. The group's broad area of impact means it frequently collaborates with other SIGs, such as SIG-Storage for persistent volume concerns and, notably, the Working Group Batch, which has been a significant driver of many recent enhancements for high-performance and data-intensive computing workloads.

The evolution of these core controllers has been a continuous process since the inception of Kubernetes. Szulik referenced annual reports and past KubeCon presentations that delve into the history and challenges encountered over the years, illustrating the iterative nature of platform development. A key driver for SIG-Apps' work is community feedback, with the group actively soliciting input on limitations, struggles, and innovative solutions from users leveraging Kubernetes in diverse environments. This collaborative approach ensures that the platform's core capabilities remain relevant and effective for a wide array of application patterns, from stateless microservices to complex stateful databases and distributed batch processing systems.

The underlying problem SIG-Apps continually addresses is how to abstract away the complexities of distributed systems, allowing users to declare their desired application state and have Kubernetes reliably manage the underlying infrastructure to achieve it. This involves intricate logic for scaling, rolling updates, self-healing, resource allocation, and ensuring data integrity, especially for stateful applications. The features discussed in the talk reflect a mature understanding of these challenges, offering more granular control and robust mechanisms for managing applications through their entire lifecycle, from initial deployment to graceful termination and data migration.

Key Findings

▶ Watch: SIG Apps annual reports and historical evolution (3:00)

The presentation highlighted a substantial array of stable (GA) features released over the past year, alongside significant advancements in beta and alpha stages, demonstrating a relentless pace of development driven by real-world use cases.

Stable (GA) Features:

  • Pod Disruption Budgets (PDBs) Refinement: A crucial update to PDBs now allows them to count only ready pods towards the disruption budget. Previously, both ready and not-ready pods were included, potentially leading to inaccurate protection. A new explicit field was introduced to maintain backward compatibility while enabling this more precise counting.
  • ReplicaSet Downscaling Randomization: To improve rollout strategies, ReplicaSets now incorporate more randomization when picking pods to terminate during downscaling. This addresses issues where the previous "newest pod first" algorithm could interfere with rolling updates aiming to replace older versions.
  • StatefulSet Start Ordinal: This feature enables StatefulSets to define a starting index (e.g., from 3 instead of 0). This is particularly useful for migration scenarios, allowing seamless transfer of individual stateful pods between clusters while maintaining their unique identifiers and the overall integrity of the StatefulSet.
  • StatefulSet PVC Retention Policy: Addressing a long-standing user request, StatefulSets now support an explicit policy for managing Persistent Volume Claims (PVCs) upon scale-down or deletion. By default, PVCs are retained, but users can now configure the StatefulSet to automatically delete associated PVCs, offering greater control over data lifecycle. This policy can be configured separately for scale-down and full deletion.
  • Batch Workload Enhancements (Driven by Working Group Batch):
  • Elastic Index Jobs: An evolution of indexed jobs, this feature allows dynamic scaling of both completions and parallelism simultaneously. Indexed jobs assign a stable, unique index to each pod, useful for parallel processing tasks. Elasticity allows adaptation to cluster load or external completion criteria.
  • CronJob-to-Job Creation Timestamp Annotation: Jobs created by CronJobs now receive an annotation (cronjob.kubernetes.io/scheduled-time) indicating their intended creation time. This provides valuable context for debugging delays or understanding scheduling decisions.
  • Pod Index Annotation: Both StatefulSets and Index Jobs now inject an annotation (batch.kubernetes.io/pod-index) into pods, exposing their unique index directly. This simplifies application logic that needs to identify a pod's role or assigned data chunk.
  • Per-Pod Backoff Limit for Jobs: Unlike the global job backoff limit (defaulting to 6 retries for the entire job), this feature allows defining a retry limit per individual pod index within a job. This prevents a single faulty node from causing the entire job to fail prematurely.
  • Job Success and Completion Policy: Jobs can now define an exit criteria that allows them to be marked as successful and complete even if not all predefined completions are reached. This is useful for certain batch patterns where a subset of successful tasks is sufficient.
  • External Job Controller Annotation: Kubernetes now supports an annotation (controller.kubernetes.io/external-provisioner) to explicitly delegate job management to an external controller. This enables sophisticated multi-cluster or queuing systems (like KubeQueue) to manage job execution while the central Kubernetes cluster only mirrors status. Significant effort was put into ensuring consistent status validation for external controllers.

Upcoming Features (Alpha/Beta/In-progress):

  • Job Pod Failure Policy: A new sub-language within the Job API to define granular retry behavior based on specific pod exit codes or conditions.
  • Mindful Rolling Deployments: Deployments will gain the ability to be more "mindful" of cluster quotas, allowing new pods to be created only when existing ones have reached a terminating state, preventing resource exhaustion during rollouts.
  • Node Draining API: A crucial, long-discussed feature to introduce a formal API for expressing and managing the node draining process, addressing limitations of current PDBs and enabling external actors to orchestrate workload migration.
  • StatefulSet Max Unavailable: An enhancement (with a three-digit enhancement number, indicating its age) to allow specifying a maximum number or percentage of unavailable pods during a StatefulSet rollout. This aims to speed up rollouts but faces challenges in defining reliable metrics for its operational correctness due to application-specific startup times.

Technical Deep Dive

▶ Watch: New PDB behavior: counting only ready pods (4:50)

The advancements presented by SIG-Apps delve into intricate aspects of Kubernetes controller logic and API design, particularly for managing complex workloads.

Pod Disruption Budget (PDB) Refinement

The enhancement to PDBs addresses a subtle but critical flaw in how application availability was historically measured during voluntary disruptions (e.g., node drains, cluster upgrades). Originally, PDBs counted both ready and not-ready pods when determining if the minimum desired availability was met. This meant that a pod that was still starting up or had failed its readiness probe could still contribute to the "available" count, potentially leading to a PDB allowing a node drain even when the application wasn't fully functional.

To rectify this, a new field was introduced within the PDB API, allowing users to explicitly specify that only ready pods should be considered. This provides a more accurate representation of application health and ensures that disruptions are only permitted when the application is truly capable of handling the reduced capacity. The inclusion of an explicit field, rather than changing the default behavior, was a deliberate design choice to maintain backward compatibility, preventing unexpected changes for existing PDB configurations. Users must opt-in to this more precise behavior, ensuring a smooth transition.

StatefulSet PVC Retention Policy

StatefulSets are cornerstone resources for managing stateful applications, which inherently rely on persistent storage. A long-standing design principle was that StatefulSets would never automatically delete their associated Persistent Volume Claims (PVCs), even when the StatefulSet itself was deleted. This was a safety measure to prevent accidental data loss, placing the responsibility on the user to manually clean up PVCs after ensuring data was no longer needed or had been migrated.

While safe, this default behavior often led to orphaned PVCs and associated Persistent Volumes (PVs), incurring unnecessary storage costs and operational overhead. The new PVC retention policy introduces explicit configuration options to manage this lifecycle. Users can now define a policy that dictates whether PVCs should be retained or deleted when the StatefulSet is scaled down (e.g., policy: { whenScaled: Delete }) or when the entire StatefulSet resource is removed (policy: { whenDeleted: Delete }). This granular control empowers operators to align storage lifecycle with application lifecycle, automating cleanup while still requiring an explicit opt-in to prevent accidental data loss. This feature is a significant quality-of-life improvement for managing database clusters, message queues, and other stateful workloads.

Elastic Index Jobs and Per-Pod Backoff Limits

The Working Group Batch has profoundly influenced the Job controller, especially for HPC and machine learning workloads. Index Jobs are a fundamental improvement, providing each pod within a job with a stable, unique index (e.g., 0, 1, 2, ...). This index is injected as an environment variable and contributes to a stable DNS name, allowing pods to identify their specific role or data partition within a distributed task.

Elastic Index Jobs take this further by allowing the completions (total number of tasks) and parallelism (number of concurrent tasks) of an indexed job to be modified dynamically together. This is crucial for adaptive workloads that might need to scale up or down based on data availability, computational resources, or external signals, a common pattern in large-scale data processing or simulation.

Complementing this, the Per-Pod Backoff Limit addresses a critical failure handling scenario. Previously, a Job had a global backoffLimit (defaulting to 6 retries) that applied to the entire job. If a single pod repeatedly failed (e.g., due to a transient error or being scheduled on a faulty node), it could quickly exhaust the global backoff limit, leading to the entire job being marked as failed, even if other pods were succeeding. The per-pod backoff limit changes this, allowing a retry limit to be applied to each individual pod index. This means a single problematic pod can be retried independently without jeopardizing the entire job, significantly improving the resilience of large-scale batch computations.

External Job Controller Annotation

This feature represents a sophisticated approach to extending Kubernetes' core capabilities for specialized use cases, particularly in multi-cluster environments or with advanced queuing systems. The controller.kubernetes.io/external-provisioner annotation allows users to explicitly tell the built-in Kubernetes Job controller not to manage a specific Job object. Instead, an external controller (like KubeQueue, a SIG-Apps sponsored project) takes over the actual execution.

The technical challenge here wasn't just to ignore the Job, but to ensure consistent status validation. If an external controller is managing a Job, it must report its status back to the Kubernetes API server in a way that is consistent with how the built-in controller would. This required significant effort in defining and enforcing validation rules for the Job object's status field in the Kubernetes API server. This ensures that regardless of whether a Job is managed internally or externally, users can rely on a consistent and predictable status representation. The primary use case is for hub-and-spoke multi-cluster deployments, where a central "hub" cluster holds the canonical Job status, while external controllers in "spoke" clusters handle the actual execution, potentially distributing tasks across multiple physical clusters.

Demo / Proof of Concept

▶ Watch: Random pod selection for ReplicaSet downscaling (6:30)

This presentation was a technical overview and roadmap discussion rather than a live demonstration. Maciej Szulik focused on detailing the features, their rationale, and their technical implications, without presenting a live demo or proof of concept during the talk itself.

Defensive Implications

▶ Watch: StatefulSet start ordinal for cluster migration (7:50)

The advancements from SIG-Apps offer crucial tools for defenders and platform engineers to enhance the resilience, efficiency, and manageability of Kubernetes workloads.

  1. Refine Pod Disruption Budgets (PDBs): Review all existing PDBs and consider enabling the new explicit field to count only ready pods. This ensures that PDBs accurately reflect the operational readiness of applications before permitting voluntary evictions. This is especially vital for critical services where even a brief period of unreadiness during a disruption could lead to service degradation.
  2. Strategic StatefulSet PVC Management: Understand and leverage the new StatefulSet PVC retention policy. For development or ephemeral stateful workloads, configure the policy to automatically delete PVCs on scale-down or deletion to prevent orphaned resources and reduce cloud costs. For production, high-value data, continue with the default retention and implement robust backup and recovery strategies, but ensure operators are aware of the explicit option for specific use cases. Document and standardize these policies.
  3. Optimize Batch Workloads:
  • Per-Pod Backoff Limits: Implement per-pod backoff limits for critical batch jobs, particularly indexed jobs, to prevent single-point failures from collapsing an entire distributed task. This significantly improves the fault tolerance of HPC and data processing pipelines.
  • Job Success & Completion Policies: Utilize job success and completion policies to define more flexible termination criteria, allowing jobs to complete successfully even if a small, acceptable number of tasks fail, or if a specific success condition is met early. This can prevent unnecessary resource consumption and accelerate reporting.
  • Pod Index Annotations: Encourage developers to leverage the injected batch.kubernetes.io/pod-index annotation in their application logic. This simplifies the development of distributed applications by providing an intrinsic way for pods to identify their role without manual configuration.
  1. Evaluate External Job Controllers: If operating in multi-cluster environments or requiring advanced job scheduling and queuing capabilities, investigate external job controllers like KubeQueue. Ensure any third-party controller used adheres to the Kubernetes API's status validation rules to maintain consistency and observability across your clusters. Understand the implications of delegating control and the status synchronization mechanisms.
  2. Prepare for Future API Enhancements: Stay informed about the progress of features like the Job Pod Failure Policy and the Node Draining API. These will provide more granular control over error handling and node lifecycle management, respectively. Proactive engagement with SIG-Apps or monitoring release notes will allow organizations to plan for adoption and integrate these capabilities into their operational playbooks as they mature.
  3. Mindful Rolling Deployments: For environments with strict resource quotas, monitor the development of "Mindful Rolling Deployments." This will enable more controlled updates that respect cluster capacity, preventing rollout failures due to temporary resource exhaustion.

Key Takeaways

  • SIG-Apps is central to Kubernetes, developing and maintaining core workload controllers like Deployments, StatefulSets, and Jobs, which are fundamental for running applications at scale.
  • Recent stable features significantly enhance the reliability and flexibility of workload management, including more accurate Pod Disruption Budgets (counting only ready pods) and explicit PVC retention policies for StatefulSets.
  • The Working Group Batch is a major contributor, driving innovation in Job management with features like Elastic Index Jobs, per-pod backoff limits, and advanced completion policies, making Kubernetes more suitable for HPC and data-intensive workloads.
  • Kubernetes is evolving to support complex multi-cluster and external orchestration scenarios, as evidenced by the new annotation allowing external controllers (e.g., KubeQueue) to manage Jobs while ensuring consistent API status validation.
  • Future development focuses on critical areas such as a dedicated Node Draining API, more intelligent deployment strategies (Mindful Rolling Deployments), and improved error handling with a Job Pod Failure Policy, addressing long-standing operational challenges.
  • The community's feedback is vital, with SIG-Apps actively seeking input on current features and challenges, demonstrating a commitment to continuous improvement based on real-world user needs.

About the Speaker(s)

Maciej Szulik is a prominent leader within the Kubernetes community, serving as one of the leads for the Special Interest Group for Applications (SIG-Apps). His work focuses on the core Kubernetes controllers responsible for managing application workloads, including Deployments, Jobs, CronJobs, DaemonSets, and StatefulSets. During this KubeCon EU presentation, he single-handedly delivered a detailed overview of SIG-Apps' recent achievements and future roadmap, demonstrating a deep understanding of the project's intricacies and challenges.

While not present for this specific talk, Janet Kuo is also recognized as a co-lead for SIG-Apps, playing a crucial role in guiding the special interest group's direction and contributions to the Kubernetes project. The talk also mentioned a third co-lead, Ken, who also could not join. The leadership team collectively oversees the evolution and maintenance of the fundamental components that enable users to run and manage their applications effectively on Kubernetes.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

Maciej Szulik, a SIG-Apps lead, delivered an absolutely critical update on the evolution of core Kubernetes workload controllers. This wasn't marketing fluff; it was a deep dive into stable features like refined Pod Disruption Budgets, granular StatefulSet PVC retention policies, and significant enhancements to Job management driven by the Working Group Batch, including Elastic Index Jobs and per-pod backoff limits. For anyone serious about operating or developing applications on Kubernetes, particularly stateful or high-performance computing workloads, this presentation offers indispensable, actionable intelligence directly from the engineers who build the platform's backbone.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon presentation on SIG-Apps' work is a clear, unsentimental look at the foundational controls underpinning Kubernetes application management. While technical in nature, the advancements in Pod Disruption Budgets, StatefulSet PVC retention, and Job elasticity directly address critical operational risks, enhance platform resilience, and provide actionable levers for managing data integrity and availability at scale. It’s essential for platform and security leaders to understand the implications of these developments for their organizational risk posture and to ensure their teams are leveraging these capabilities.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025