A Comparative Analysis of Kueue, Volcano, and YuniKorn - Wei Huang, Apple & Shiming Zhang, DaoCloud

Wei Huang, Apple, Shiming Zhang, DaoCloud

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk provides an in-depth, comparative analysis of three prominent Kubernetes batch schedulers: Kueue, Volcano, and YuniKorn. Presented by Wei Huang from Apple AML, a co-chair and maintainer of SIG Scheduling, and Shiming Zhang from DaoCloud, a SIG Scheduling reviewer and co-maintainer, the session addresses a frequently asked question among Kubernetes users regarding the optimal scheduler for their batch workloads. The speakers systematically break down the history, design philosophies, key features, API integrations, and performance characteristics of each scheduler, offering a comprehensive guide for practitioners.

Watch on YouTube

Visual summary for A Comparative Analysis of Kueue, Volcano, and YuniKorn - Wei Huang, Apple & Shiming Zhang, DaoCloud by Wei Huang, Apple, Shiming Zhang, DaoCloud
Visual summary for A Comparative Analysis of Kueue, Volcano, and YuniKorn - Wei Huang, Apple & Shiming Zhang, DaoCloud by Wei Huang, Apple, Shiming Zhang, DaoCloud

Key moments

  1. 0:00 Introduction and session agenda
  2. 0:50 Why default Kubernetes scheduler falls short for batch workloads
  3. 2:40 History and evolution of Kueue, Volcano, and YuniKorn
  4. 4:40 High-level workflow comparison and job admission models
  5. 8:00 API design comparison and Kubernetes native experience
  6. 9:40 Hierarchical multi-queue and min/max quota management

A Comparative Analysis of Kueue, Volcano, and YuniKorn

Speakers: Wei Huang, Apple & Shiming Zhang, DaoCloud

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=njT5r3JjIaA

Overview

This talk provides an in-depth, comparative analysis of three prominent Kubernetes batch schedulers: Kueue, Volcano, and YuniKorn. Presented by Wei Huang from Apple AML, a co-chair and maintainer of SIG Scheduling, and Shiming Zhang from DaoCloud, a SIG Scheduling reviewer and co-maintainer, the session addresses a frequently asked question among Kubernetes users regarding the optimal scheduler for their batch workloads. The speakers systematically break down the history, design philosophies, key features, API integrations, and performance characteristics of each scheduler, offering a comprehensive guide for practitioners.

The core motivation behind these specialized schedulers stems from the inherent limitations of the default Kubernetes scheduler when handling batch-oriented, resource-intensive, and multi-tenant workloads. As organizations increasingly leverage Kubernetes for machine learning, data processing, and other high-throughput computing tasks, the need for advanced scheduling capabilities—such as Gang Scheduling, hierarchical resource quotas, and efficient cluster utilization—becomes paramount. This article aims to distill the speakers' insights, providing a detailed technical overview to help readers understand the nuances and make informed decisions when selecting a batch scheduler.

The talk is particularly relevant for platform engineers, MLOps practitioners, and anyone managing large-scale, shared Kubernetes clusters where batch jobs are a critical component. By highlighting the strengths and weaknesses of Kueue, Volcano, and YuniKorn across various dimensions, the speakers empower the audience to choose a solution that best aligns with their specific operational requirements, existing ecosystem, and performance expectations.

Background

▶ Watch: Introduction and session agenda (0:00)

The journey into specialized Kubernetes schedulers like Kueue, Volcano, and YuniKorn is rooted in the fundamental design principles and inherent limitations of the default Kubernetes scheduler when confronted with batch workloads. As Wei Huang articulates, three primary deficiencies necessitate these advanced solutions:

  1. Pod-Centric Scheduling: The default scheduler was originally designed for service workloads (e.g., web servers, microservices) in Kubernetes 1.0 (2015). It treats each pod as an individual, atomic unit for scheduling decisions. While this ensures consistent pod lifecycle management and unified API integration at the pod level, it critically lacks the notion of job-level scheduling. This means decisions made for a single pod may not be optimal for a cohesive job, particularly for workloads requiring all components to start simultaneously, a concept known as Gang Scheduling.
  1. Inefficient Capacity Maximization: Kubernetes offers ResourceQuota as a static, namespace-scoped mechanism for resource management. However, modern organizations demand more dynamic, sharable, and hierarchical resource quotas to effectively manage cluster capacity across multiple teams or projects. The default scheduler doesn't provide the necessary constructs to support such advanced quota management, leading to potential resource underutilization or unfair allocation in multi-tenant environments.
  1. Single-Queue Design: Internally, the default scheduler operates with a single, monolithic queue. This "one-size-fits-all" approach is ill-suited for the diverse requirements of batch workloads, which often benefit from a multi-queue design to implement features like priority-based scheduling, preemption, and fair sharing more naturally and efficiently.

The evolution of these batch schedulers reflects a community effort to address these gaps:

  • Kube-scheduler (2015): The original scheduler introduced with Kubernetes 1.0.
  • Kube-batch (2017): An experimental sub-project of SIG Scheduling, intended to evolve batch capabilities independently before merging into the default scheduler. However, its development path diverged as SIG Scheduling's focus shifted towards exposing scheduler extensibility via the scheduling framework.
  • Volcano (2019): Emerged from the kube-batch project, forking its initial version. It was later donated to the CNCF (Cloud Native Computing Foundation), establishing itself as a dedicated batch scheduler.
  • YuniKorn (2019): Raised later in 2019, YuniKorn leverages many design concepts from the Apache Hadoop YARN scheduler and is hosted under the Apache Foundation.
  • Kueue (2022): The most recent addition, Kueue was developed as a sub-project of SIG Scheduling specifically to resolve the missing pieces in job scheduling within the native Kubernetes ecosystem.

Understanding their operational workflows is crucial for distinguishing these schedulers:

  • Default Scheduler / YuniKorn: In this model, users submit a job (or job-like workload), which is then fanned out into individual pods by a workload controller. These pods are subsequently scheduled by the default Kubernetes scheduler. YuniKorn follows a similar pattern but requires specific labels and annotations on pods to optimize their scheduling according to its design philosophy.
  • Kueue: Kueue introduces an admission control layer. When a job is submitted, the workload controller cannot immediately create pods. Instead, it must wait for Kueue to admit the job. Kueue, based on its queue setup and API design, performs checks for fairness, quotas, and accessibility. Only after all checks pass does Kueue "ungate" the job, allowing the workload controller to create pods, which are then scheduled by the default Kubernetes scheduler. This highlights Kueue's role as a job-level manager rather than a pod-level scheduler.
  • Volcano: Volcano operates with its own custom APIs, notably VCJob and PodGroup. These APIs carry Volcano-specific characteristics, requiring the Volcano controller and Volcano scheduler to be used in conjunction. Unlike Kueue, Volcano manages both the job-level and pod-level scheduling within its own integrated system.

Placing these schedulers on a spectrum reveals their core focus: the default Kubernetes scheduler and YuniKorn primarily operate at the pod scheduling level. Kueue focuses on job admission and job-level management, acting as a queuing system for jobs that are eventually scheduled by the default scheduler. Volcano, conversely, provides a comprehensive solution spanning both the job controller and the scheduler, demanding a fully integrated deployment.

Key Findings

▶ Watch: History and evolution of Kueue, Volcano, and YuniKorn (2:40)

The comparative analysis reveals distinct design philosophies and feature sets among Kueue, Volcano, and YuniKorn, particularly in their approach to multi-queue management, API design, quota enforcement, Gang Scheduling, and ecosystem integration.

Multi-Queue and Hierarchical Quotas

All three schedulers fundamentally address the single-queue limitation of the default Kubernetes scheduler by adopting a multi-queue, hierarchical structure. This tree-like design allows users to submit jobs to leaf nodes (queues), enabling more granular control and better resource isolation. Each queue can be associated with quotas, which all schedulers implement using a min/max quota management pattern. This can be interpreted as guaranteed resources (min) versus best-effort resources (max). For instance, an organization might allocate a minimum of 2 CPUs to Team A, but allow it to burst up to 8 CPUs if available. If another team with a guaranteed quota requires resources, preemption mechanisms reclaim best-effort resources.

Key differences in preemption:

  • YuniKorn: Supports inter-queue preemption only, meaning workloads in one queue can preempt workloads in another queue, but not within the same queue.
  • Kueue: Preempts using the job as a unit. If a job needs to be preempted, all its associated pods are removed.
  • Volcano and YuniKorn: Can preempt using the pod as a minimum unit, offering finer-grained control over resource reclamation.

API Design and Multi-Tenancy

The choice of API design significantly impacts usability and multi-tenancy capabilities:

  • Kueue: Stands out with the most Kubernetes-native API design. It offers dedicated APIs to control which users can submit jobs to specific queues, integrating seamlessly with Kubernetes concepts like RBAC (Role-Based Access Control) and namespaces. This "out-of-the-box" self-service model is highly advantageous for multi-tenant environments.
  • Volcano: While it uses a queue design, it lacks a native mechanism to control user-to-queue submission, potentially requiring external solutions for multi-tenancy.
  • YuniKorn: Embeds all its configurations into a large ConfigMap. This approach is less Kubernetes-native and can be cumbersome for multi-tenancy, often necessitating an additional layer for self-service capabilities.

Gang Scheduling

Gang Scheduling, the ability to schedule all pods of a job simultaneously or none at all, is a critical feature for many batch workloads. The schedulers employ different strategies:

  • YuniKorn: Despite not directly involving a job CRD, it supports Gang Scheduling by creating placeholder pods equal to the number of actual pods required. It attempts to schedule these placeholders; if successful, the placeholders are deleted and replaced with real pods. This design, however, can lead to tripled API requests (create placeholder, delete placeholder, create real pod) and potential race conditions.
  • Volcano: Implements Gang Scheduling through an in-memory dry run. Once a job enters a queue and passes basic checks (like quota), its pods are fanned out and subjected to an in-memory simulation. If all pods can be scheduled together, they are then committed to the cluster. This approach is generally more efficient than YuniKorn's.
  • Kueue: Does not perform pod scheduling itself; it leverages other schedulers (typically the default Kubernetes scheduler). For Gang Scheduling, Kueue creates pods and allows them to be scheduled. If not all pods are scheduled together, it relies on a "wait for pods ready" timeout mechanism to delete and retry the job. This mechanism is acknowledged as "not ideal," which is why Kueue is often combined with co-scheduling plugins (e.g., in Kubeflow) to achieve better accuracy for Gang Scheduling.

Ecosystem Integration

Integration with the broader Kubernetes ecosystem, especially with computing frameworks (e.g., Spark, Flink, PyTorch) and the Cluster Autoscaler (CA), is a significant differentiator.

  • Computing Frameworks:
  • Volcano/YuniKorn: Typically require invasive integration, where the framework's operator (e.g., a Spark operator) must be modified to create scheduler-specific APIs (like PodGroup for Volcano) or inject scheduler-specific labels/annotations into pods. This can make adapting to new frameworks or upgrading existing ones more complex.
  • Kueue: Employs a generic suspension pattern. The workload controller is told to "suspend" pod creation. Kueue handles job admission, and once deemed ready, it "unsuspends" the controller, allowing it to create pods for the downstream scheduler. This roundtrip communication makes Kueue less invasive and more sustainable for integrating with diverse and evolving computing frameworks.
  • Cluster Autoscaler (CA):
  • CA simulates unscheduled pods using the default scheduler's SDK. Compatibility with this SDK is crucial.
  • DRA (Dynamic Resource Allocation): Introduced in Kubernetes 1.23, DRA allows dynamic allocation of specialized resources. Kueue supports DRA because it relies on the default scheduler. At the time of the talk, YuniKorn and Volcano did not yet support DRA.
  • Provision Request API: Raised by the SIG Autoscaling group, this API allows schedulers to proactively request resources from the CA, rather than passively waiting for unscheduled pods to trigger scaling. Only Kueue supports this proactive resource reservation, offering better control over cluster scaling.

Multi-Cluster Support

  • Kueue: Provides native support for multi-cluster environments.
  • Volcano: Leverages Karmada for multi-cluster capabilities.
  • YuniKorn: The talk did not explicitly detail YuniKorn's multi-cluster support, but implied it might require custom integration.

In summary, Kueue stands out for its Kubernetes-native API design, less invasive integration patterns, and advanced Cluster Autoscaler capabilities. Volcano excels in Gang Scheduling scenarios due to its purpose-built design. YuniKorn offers a multi-queue approach but has some limitations in API design and Gang Scheduling implementation efficiency.

Technical Deep Dive

▶ Watch: High-level workflow comparison and job admission models (4:40)

The technical distinctions between Kueue, Volcano, and YuniKorn are best understood by dissecting their approaches to core scheduling challenges that the default Kubernetes scheduler fails to adequately address.

Addressing Kubernetes Scheduler Limitations

As highlighted in the background, the default Kubernetes scheduler's design for service workloads creates specific challenges for batch jobs:

  1. Job-Level Awareness: The default scheduler treats each pod as an independent scheduling unit. This pod-level scheduling is suboptimal for batch jobs where a collection of pods forms a logical unit. For instance, a distributed training job requires all its worker and coordinator pods to be available and scheduled concurrently. Without job-level scheduling or Gang Scheduling, individual pods might get scheduled, leading to deadlocks or inefficient resource utilization if their counterparts cannot be scheduled due to resource scarcity.
  2. Resource Quota Management: The ResourceQuota API in Kubernetes is static and namespace-bound. This rigid structure falls short in modern, dynamic environments where multiple teams or projects share a cluster. The need for hierarchical resource quotas that can be dynamically shared, burstable, and preemptible is paramount. For example, a data science team might need to temporarily exceed its guaranteed quota during peak training periods, but only if resources are available and can be reclaimed by higher-priority jobs.
  3. Single Scheduling Queue: The kube-scheduler maintains a single internal queue for all incoming pods. While simple, this single-queue design is inefficient for complex batch scheduling requirements. Advanced features like fair sharing, priority-based preemption across different job types, or managing jobs with varying Service Level Objectives (SLOs) are difficult to implement effectively within a single queue, often leading to head-of-line blocking issues.

Multi-Queue and Hierarchical Resource Management

All three specialized schedulers tackle the single-queue problem by implementing a multi-queue architecture. These queues are typically arranged in a tree-like hierarchical structure, allowing for nested resource allocation and administrative domains. Users submit jobs to the "leaf nodes" (queues) of this hierarchy.

  • Quota Enforcement: Each queue can be associated with min/max quotas. The min quota represents a guaranteed resource allocation, ensuring a baseline capacity for a team or project. The max quota allows for bursting beyond the guaranteed minimum, utilizing idle cluster resources on a best-effort basis.
  • Preemption: When guaranteed resources are contested, preemption occurs.
  • YuniKorn supports inter-queue preemption: jobs in a higher-priority queue can preempt jobs in a lower-priority queue. However, it lacks intra-queue preemption, meaning preemption cannot occur between jobs within the same queue. YuniKorn preemption works at the pod-unit level.
  • Volcano also supports preemption at the pod-unit level, allowing for fine-grained resource reclamation.
  • Kueue employs job-unit preemption. If a job needs to be preempted, all its associated pods are evicted, ensuring the entire job is removed to free up resources. This simplifies cleanup but can be less resource-efficient than pod-level preemption if only a fraction of a job's resources is needed.

API Design and Configuration Paradigms

The choice of API design dictates the user experience, integration complexity, and multi-tenancy capabilities:

  • Kueue: Adopts a Kubernetes-native API approach. It introduces Custom Resource Definitions (CRDs) like ClusterQueue and LocalQueue to define the hierarchical queue structure and Workload to represent the jobs. This design allows for seamless integration with Kubernetes' native access control (RBAC) and resource management (namespaces), enabling self-service and robust multi-tenancy. For example, specific RBAC roles can be granted to allow users to submit jobs only to designated LocalQueues.
  • YuniKorn: Primarily relies on a large ConfigMap for its configuration. While flexible, a single ConfigMap can become unwieldy in complex multi-tenant environments, making it challenging to manage granular permissions or provide self-service queue creation without building additional tooling on top of Kubernetes. YuniKorn requires specific annotations and labels on pods (e.g., yunikorn.apache.org/queue) to direct them to the correct queue.
  • Volcano: Introduces its own set of CRDs, such as VCJob and PodGroup, which are integral to its operation. These custom APIs provide rich, Volcano-specific attributes for batch jobs, but they also mean that users and operators must explicitly adapt to Volcano's API model, potentially requiring changes to existing workload definitions.

Gang Scheduling Implementations

The implementation of Gang Scheduling is a critical differentiator:

  • YuniKorn's Placeholder Pods: YuniKorn's approach involves creating placeholder pods for each real pod required by a job. These placeholders are lightweight, temporary pods that reserve resources. If all placeholders can be scheduled, they are then deleted, and the actual workload pods are created and scheduled. This "create-delete-create" cycle leads to a significant increase in API server load (almost triple the API requests) and introduces potential race conditions between the deletion of placeholders and the creation of real pods, impacting performance and reliability.
  • Volcano's In-Memory Dry Run: Volcano performs an in-memory dry run of the scheduling process. When a VCJob is submitted, the Volcano controller creates PodGroup objects. The Volcano scheduler then attempts to simulate the scheduling of all pods within a PodGroup without committing them to the API server. If the dry run is successful (i.e., resources are available for all pods), the actual pods are then created and scheduled. This method is generally more efficient as it avoids the API server overhead of placeholder pods and reduces the window for race conditions.
  • Kueue's External Plugin Reliance: Kueue, by design, is a job admission controller, not a pod scheduler. For Gang Scheduling, it relies on the downstream pod scheduler (typically kube-scheduler) and potentially external co-scheduling plugins. Kueue ensures that a job's Workload object is admitted only when enough resources are available for all its pods. However, if the kube-scheduler (without a co-scheduling plugin) cannot schedule all pods of an admitted job concurrently, Kueue's fallback mechanism is to use a "wait for pods ready" timeout. If pods don't become ready within the timeout, the job is deleted and retried. This retry mechanism is less ideal for strict Gang Scheduling requirements, which is why communities like Kubeflow integrate Kueue with specific co-scheduling plugins (e.g., the coscheduling plugin in the Kubernetes scheduling framework) for better accuracy.

Ecosystem Integration: Compute Frameworks and Cluster Autoscaler

  • Computing Framework Integration:
  • Invasive Integration (Volcano/YuniKorn): Integrating new computing frameworks (e.g., Spark, Flink, TensorFlow) with Volcano or YuniKorn often requires modifying the framework's operator (e.g., the Spark operator). This modification involves teaching the operator to create Volcano-specific PodGroup objects or to add YuniKorn-specific labels/annotations to the pods it generates. This approach is "invasive" because it couples the framework's operator tightly with a specific scheduler, making upgrades or switching schedulers more complex.
  • Generic Suspension Pattern (Kueue): Kueue introduces a generic suspension pattern based on a roundtrip communication mechanism. A framework operator can be configured to suspend the creation of pods for a job. Kueue then admits the Workload (job) based on available resources and, once admitted, "unsuspends" the operator. The operator then proceeds to create the pods, which are finally scheduled by the default scheduler. This decoupled approach is less invasive, making it easier to integrate new frameworks or update existing ones without extensive code changes to the operators.
  • Cluster Autoscaler (CA) Compatibility:
  • kube-scheduler SDK Reliance: The Cluster Autoscaler (CA) simulates unscheduled pods using the kube-scheduler SDK to determine if new nodes are needed. Therefore, compatibility with this SDK is critical for any scheduler that wants to interact effectively with CA. Since Kueue leverages the default kube-scheduler for pod scheduling, it inherently benefits from this compatibility.
  • Dynamic Resource Allocation (DRA): Introduced in Kubernetes 1.23, DRA is a new API for dynamic allocation of specialized resources (e.g., GPUs, FPGAs). At the time of the talk, Kueue natively supported DRA because it relies on the default scheduler's capabilities. YuniKorn and Volcano, with their custom scheduling logic, did not yet support DRA.
  • Provision Request API: The SIG Autoscaling group introduced the Provision Request API to enable schedulers to proactively request new nodes from the CA, rather than passively waiting for unscheduled pods to trigger scaling. This allows for more intelligent and timely scaling decisions. Currently, only Kueue supports this proactive API, giving it a significant advantage in optimizing cluster utilization and cost in dynamic cloud environments.

Demo / Proof of Concept

▶ Watch: API design comparison and Kubernetes native experience (8:00)

While the talk didn't feature a live interactive demo, the speakers presented a comprehensive performance benchmark comparing Kueue, Volcano, and YuniKorn across various job configurations, both with and without Gang Scheduling enabled. This benchmark served as a critical proof of concept for their analysis.

Methodology and Environment

The speakers emphasized an innovative and meticulous approach to benchmarking:

  1. Standardized QPS: Acknowledging that default settings can skew results (e.g., YuniKorn's default API QPS was 1000 vs. 50 for others), all schedulers and controllers were configured with standardized QPS (Queries Per Second) to ensure an "apple-to-apple" comparison.
  2. API Server Audit Log for Metrics: Traditional Prometheus metrics can suffer from approximation errors during high-scale, short-duration tests (e.g., scheduling 10,000 pods). The speakers innovated by using API server audit logs to precisely capture pod creation time, pod scheduling time, and pod deletion time for each event. This allowed for highly accurate plotting on Grafana dashboards.
  3. Test Environment:
  • Kubernetes version 1.23.
  • Kueue 0.10 with an upstream co-scheduling plugin and a local bug fix.
  • Volcano and YuniKorn were tested with their respective scheduler and controller components, ensuring consistent versions. Volcano also received a bug fix from its team (thanks to Shenan).
  • Kube-OVN (KOK-O) was used to simulate a large number of nodes, acting as a powerful benchmarking tool.
  1. Test Scenarios: A total of 10,000 pods were scheduled under two main categories: Gang Scheduling disabled and Gang Scheduling enabled. The key variable was the composition of these 10,000 pods:
  • Many jobs: 10,000 jobs, 1 pod each.
  • Fewer jobs: 500 jobs, 20 pods each.
  • Even fewer jobs: 20 jobs, 500 pods each.
  • Single large job: 1 job, 10,000 pods.
  • The number of queues was not varied, as local testing showed no significant performance difference.

Performance Results (Gang Scheduling Disabled)

The results without Gang Scheduling highlighted the impact of API server interactions:

  • Many Jobs (10,000 jobs, 1 pod each): YuniKorn showed a slight lead in pod creation time, attributed to web hook QPS limits for job mutation requests in Kueue and Volcano. However, the time difference between pod creation and pod scheduling (the actual scheduling latency) was similar across all three.
  • Fewer Jobs (500 jobs, 20 pods each; 20 jobs, 500 pods each; 1 job, 10,000 pods): As the number of jobs decreased (and pods per job increased), Kueue consistently outperformed YuniKorn, which in turn was ahead of Volcano. Volcano's slower performance was attributed to it making more webhook requests to mutate pods, particularly when creating pods in batches (an intentional design, but with overhead).
  • Summary (No Gang Scheduling): For workloads with many small jobs, scheduling speed is heavily influenced by the number of job webhook mutation requests. If only a single mutation request is needed (e.g., one large job), Kueue generally leads YuniKorn. Volcano lags behind both in all scenarios without Gang Scheduling.

Performance Results (Gang Scheduling Enabled)

The dynamic shifted dramatically when Gang Scheduling was enabled:

  • Many Jobs (10,000 jobs, 1 pod each):
  • YuniKorn created double the number of pods due to its placeholder pod mechanism (placeholder creation, deletion, then real pod creation).
  • Volcano demonstrated a significant lead in pod scheduling time and job completion time compared to Kueue and YuniKorn. This confirmed Volcano's strength in Gang Scheduling, as it was originally designed for this specific feature.
  • Fewer Jobs (500 jobs, 20 pods each): YuniKorn's performance improved relative to Kueue, but Volcano maintained its substantial lead in both pod scheduling and job completion times.
  • Summary (Gang Scheduling Enabled): For workloads heavily dependent on Gang Scheduling, Volcano is demonstrably superior to both Kueue and YuniKorn. Its in-memory dry run mechanism proved more efficient than YuniKorn's placeholder approach and Kueue's reliance on external plugins/retries.

The speakers concluded by announcing that their benchmarking tools and methodology, including the audit log parser, would be open-sourced, enabling others to replicate and customize the tests. This provides a valuable asset for the community to conduct their own performance evaluations.

Defensive Implications

▶ Watch: Hierarchical multi-queue and min/max quota management (9:40)

The detailed analysis of Kueue, Volcano, and YuniKorn provides critical insights for defenders and platform engineers tasked with securing and optimizing Kubernetes clusters for batch workloads. The choice of scheduler has profound implications for resource utilization, fairness, stability, and even the attack surface.

  1. Acknowledge Default Scheduler's Limitations: The first defensive posture is to recognize that the default Kubernetes scheduler is fundamentally inadequate for most production-grade batch workloads. Its lack of job-level awareness, hierarchical quotas, and a multi-queue design will inevitably lead to resource contention, unfair sharing, and inefficient cluster utilization for ML training, data processing, and other batch-intensive tasks. Adopting a specialized batch scheduler is not optional but essential for these use cases.
  1. Prioritize API Design for Multi-Tenancy and Security:
  • Kueue's Kubernetes-native APIs (ClusterQueue, LocalQueue, Workload CRDs) are a significant advantage for multi-tenant environments. They integrate seamlessly with Kubernetes RBAC, allowing administrators to define precise permissions on who can create jobs in which queues. This fine-grained control is crucial for preventing unauthorized resource consumption or interference between tenants.
  • YuniKorn's ConfigMap-based configuration can be less secure and harder to manage in multi-tenant setups, potentially requiring custom solutions to enforce access control to queues.
  • Volcano's custom APIs (VCJob, PodGroup) require careful RBAC configuration around these CRDs to ensure that only authorized users can submit Volcano-specific jobs.
  1. Evaluate Gang Scheduling Requirements:
  • If Gang Scheduling is a hard requirement for critical workloads (e.g., distributed training jobs that fail if not all components start together), Volcano is the strongest contender due to its purpose-built design and superior performance in these scenarios. Defenders should understand Volcano's PodGroup concept and ensure its controller and scheduler are robustly deployed.
  • For Kueue, if Gang Scheduling is needed, it's imperative to pair it with a co-scheduling plugin within the kube-scheduler framework. Relying solely on Kueue's "wait for pods ready" timeout is a less robust solution and could lead to job failures or excessive retries, impacting workload stability.
  • YuniKorn's placeholder pod mechanism introduces more API server load and potential race conditions, which could be a concern for very high-scale or latency-sensitive Gang Scheduled workloads.
  1. Leverage Advanced Cluster Autoscaler Integration:
  • Kueue's support for the Provision Request API is a powerful defensive tool for optimizing cluster costs and resource availability. By allowing the scheduler to proactively request new nodes, it can prevent situations where jobs are stuck waiting for resources, leading to better SLOs and reduced operational overhead. This proactive scaling can also help mitigate "thundering herd" issues where many jobs suddenly become runnable.
  • Compatibility with kube-scheduler SDK and DRA (for specialized hardware) are also key considerations. Kueue's inherent compatibility simplifies integration with the CA and future-proofs resource allocation for advanced hardware.
  1. Consider Ecosystem Integration and Maintainability:
  • Kueue's generic suspension pattern for computing framework integration is a significant advantage for long-term maintainability and security. By decoupling the scheduler from the framework's operator, it reduces the need for invasive code changes when integrating new frameworks or updating existing ones. This minimizes the risk of introducing bugs or vulnerabilities during integration.
  • Volcano and YuniKorn's more invasive integration approaches mean that defenders must carefully vet any modifications to framework operators, ensuring they do not introduce security regressions or break compatibility with other cluster components.
  1. Implement Robust Resource Quota and Preemption Strategies:
  • All three schedulers offer min/max quotas and preemption. Defenders should carefully design their hierarchical queue structures and quota policies to ensure fairness and prevent resource starvation for critical workloads.
  • Understand the preemption granularity (job-unit vs. pod-unit, inter-queue vs. intra-queue) to align with organizational priorities for workload interruption. For example, if a high-priority job must displace all resources from a lower-priority job, job-unit preemption (Kueue) might be preferred. If partial preemption is acceptable, pod-unit preemption (Volcano, YuniKorn) offers more flexibility.

By thoroughly evaluating these technical and operational aspects, defenders can select a batch scheduler that not only meets performance requirements but also enhances the overall security, stability, and manageability of their Kubernetes clusters.

Key Takeaways

  • The default Kubernetes scheduler is ill-equipped for modern batch workloads, lacking job-level awareness, hierarchical resource quotas, and a multi-queue design, necessitating specialized schedulers like Kueue, Volcano, and YuniKorn.
  • All three specialized schedulers adopt hierarchical multi-queue structures with min/max quota management and preemption, but differ in their API design and preemption granularity (job-unit vs. pod-unit, inter-queue only).
  • Kueue offers the most Kubernetes-native API design, integrating seamlessly with RBAC and namespaces for robust multi-tenancy. Its generic suspension pattern provides a less invasive and more sustainable way to integrate with diverse computing frameworks.
  • Volcano is purpose-built for Gang Scheduling, demonstrating superior performance in these scenarios through its in-memory dry run mechanism, making it the preferred choice when strict all-or-nothing job scheduling is critical.
  • Kueue exhibits stronger compatibility with the Cluster Autoscaler, supporting Dynamic Resource Allocation (DRA) and being the sole scheduler to support the Provision Request API for proactive node scaling, leading to better resource utilization and cost optimization.
  • Benchmarking batch schedulers requires meticulous methodology; the speakers' innovative use of API server audit logs for precise event timing provides a more accurate performance comparison than traditional Prometheus metrics, and their tools will be open-sourced.

About the Speaker(s)

Wei Huang is a prominent figure in the Kubernetes community, serving as a co-chair and maintainer of SIG Scheduling. His expertise in Kubernetes scheduling is further demonstrated by his role at Apple AML, where he contributes to machine learning infrastructure.

Shiming Zhang is a valued contributor to the Kubernetes scheduling ecosystem. He holds the positions of SIG Scheduling reviewer and sub-project co-maintainer. Shiming is based in Shanghai and works for DaoCloud, a cloud-native technology and service provider.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk delivers an exceptional, no-nonsense comparative analysis of Kubernetes batch schedulers Kueue, Volcano, and YuniKorn. The speakers, deeply embedded in SIG Scheduling, provide a brutally honest and technically exhaustive breakdown of each solution's design, API choices, multi-tenancy implications, and critical performance characteristics. The innovative benchmarking methodology, leveraging API server audit logs for precise metrics, sets a new standard for evaluating these complex systems, making this a foundational resource for anyone serious about managing batch workloads on Kubernetes.

Heather Calloway (CISO) — STRONG ACCEPT

This comparative analysis of Kubernetes batch schedulers provides critical clarity for platform leaders. It moves beyond generic claims to deliver a meticulously benchmarked assessment of Kueue, Volcano, and YuniKorn, directly informing strategic decisions around multi-tenancy, resource governance, and operational resilience. The talk excels in translating complex technical trade-offs into actionable insights for engineers and executives alike, particularly highlighting the governance implications of API design and proactive scaling capabilities.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025