Balancing Cost and Efficiency: Day2 Optimization of Multi-Cluster AI Infrastructure - Kevin Wang

Kevin Wang

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In today's rapidly evolving technological landscape, the deployment and management of Artificial Intelligence (AI) and Machine Learning (ML) workloads are increasingly moving towards distributed, multi-cluster Kubernetes environments. This talk by Kevin Wang delves into the critical challenges and innovative solutions for "Day2 optimization" in such complex infrastructures, focusing on balancing cost, efficiency, and reliability. Wang, an early Kubernetes contributor and key figure in the Volcano and Kamada projects, outlines how these CNCF initiatives are evolving to meet the demands of enterprise-scale AI.

Watch on YouTube

Visual summary for Balancing Cost and Efficiency: Day2 Optimization of Multi-Cluster AI Infrastructure - Kevin Wang by Kevin Wang
Visual summary for Balancing Cost and Efficiency: Day2 Optimization of Multi-Cluster AI Infrastructure - Kevin Wang by Kevin Wang

Key moments

  1. 0:00 Speaker intro and multi-cluster AI problem statement
  2. 2:10 Overview of multi-cluster AI platform architecture
  3. 3:55 Key design principles for multi-cluster AI systems
  4. 5:25 Three major challenges: scheduling, failover, and queuing
  5. 6:00 Balancing scheduling efficiency with resource model approach
  6. 8:00 Estimator approach for multi-cluster AI workload scheduling
  7. 9:35 Complexities and considerations for multi-cluster workload failover

Balancing Cost and Efficiency: Day2 Optimization of Multi-Cluster AI Infrastructure

Speakers: Kevin Wang

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=4pBhVVrCHyM

Overview

In today's rapidly evolving technological landscape, the deployment and management of Artificial Intelligence (AI) and Machine Learning (ML) workloads are increasingly moving towards distributed, multi-cluster Kubernetes environments. This talk by Kevin Wang delves into the critical challenges and innovative solutions for "Day2 optimization" in such complex infrastructures, focusing on balancing cost, efficiency, and reliability. Wang, an early Kubernetes contributor and key figure in the Volcano and Kamada projects, outlines how these CNCF initiatives are evolving to meet the demands of enterprise-scale AI.

The presentation addresses common pain points faced by organizations, including geographical distribution, hardware procurement cycles, and the restructuring of resource pools that necessitate multi-cluster architectures. It highlights the inherent difficulties in maintaining hyperscale, monolithic clusters, advocating instead for a federated approach. Wang elaborates on specific optimizations across scheduling, failover mechanisms, and workload queuing, providing a deep dive into the technical underpinnings and future directions for managing AI workloads effectively in a federated Kubernetes ecosystem.

This article provides a comprehensive exploration of the strategies presented, offering insights into how practitioners can leverage these advancements to enhance the performance, resilience, and cost-effectiveness of their multi-cluster AI platforms. It's particularly relevant for platform engineers, SREs, and architects involved in designing and operating large-scale AI/ML infrastructure on Kubernetes.

Background

▶ Watch: Speaker intro and multi-cluster AI problem statement (0:00)

The necessity for multi-cluster Kubernetes environments arises from a confluence of operational and business drivers. Many companies operate across multiple geographical regions, leading to physically distributed data centers and, consequently, distributed clusters. Furthermore, the inherent challenges of hardware procurement, especially for specialized AI accelerators like GPUs, often force organizations to extend their on-premises infrastructure to public clouds to rapidly acquire necessary resources. A common scenario also involves the evolution of organizational structures, where independent business units initially manage their own clusters, eventually consolidating under a shared infrastructure team that seeks to centralize maintenance and technology. The alternative—maintaining a single, hyperscale cluster—is often deemed too difficult to build, manage, and rebuild, driving the adoption of federated approaches.

The foundational architecture for managing AI workloads in such environments typically involves a two-level system. At the in-cluster level, projects like Volcano excel. Volcano is a Kubernetes-native batch scheduler designed to provide advanced scheduling features specifically for AI and batch workloads, including queue management, resource sharing among queues, priority-based scheduling, and capacity management. Above this, at the multi-cluster or federated layer, projects like Kamada manage cluster health, propagate workloads, and make high-level placement decisions. Kamada allows users to define preferences for workload distribution, whether workloads should be divided, scheduled to a single cluster, or replicated. For AI training workloads, the common approach is to schedule them to a single, optimal cluster.

Key design principles underpin this federated architecture. The two-level system is considered unavoidable, with each layer having a specific focus and remaining loosely coupled. The federated control plane primarily handles inter-cluster coordination, while member clusters focus on in-cluster operations, maintaining a high degree of autonomy. Crucially, the federated control plane avoids storing the full, detailed status of member clusters to prevent it from becoming a single point of contention or a bottleneck, effectively preventing it from becoming a "hyperscale cluster" itself. Federated scheduling is designed to collaborate with, rather than replace, in-cluster scheduling. A significant insight highlighted is that federated scheduling decisions are inherently expensive; finding a suitable cluster, sending the workload, only for the in-cluster scheduler to reject it, results in costly evictions and retries. Therefore, optimizing the "first-time attempt" at federated scheduling is paramount for efficiency.

Key Findings

▶ Watch: Key design principles for multi-cluster AI systems (3:55)

Kevin Wang's talk highlights three major challenges and their corresponding solutions for Day2 optimization in multi-cluster AI infrastructure: the trade-off in scheduling efficiency vs. accuracy, robust cluster failover mechanisms, and effective federated workload queuing.

  1. Scheduling Trade-offs for AI Workloads: The federated control plane faces a dilemma: how much cluster status information should it store and process? Storing too much information increases overhead and complexity, while too little compromises scheduling accuracy. Kamada offers two primary options:
  • Resource Model: This approach compresses cluster status by quantifying node resources into "grades." It's efficient for CPU or memory-intensive workloads and offers higher throughput and lower scheduling latency because all necessary information is readily available in the scheduler. However, its accuracy can vary, especially for complex AI workloads.
  • Scheduler Estimator: This method is more akin to an in-cluster scheduler. It watches member clusters to collect detailed node and pod status in memory. When Kamada performs scheduling, it triggers an RPC call to each estimator to calculate maximum available replicas. While this takes longer and has a more extensive footprint due to storing detailed cluster status, it offers significantly higher accuracy. For AI workloads, particularly those involving GPUs, distributed training, and inference, the estimator's awareness of in-cluster network topology and precise resource availability makes it a worthy trade-off, as wasting GPU resources is highly undesirable.
  1. Flexible Cluster Failover for Workloads: Traditional failover mechanisms often struggle with the nuances of partial cluster failures, where only specific components are impacted rather than the entire cluster. Draining or evicting all workloads in such scenarios is often undesirable and can amplify the failure if insufficient redundant resources are available elsewhere. The revised approach moves from rigid taint-based eviction to a more flexible system:
  • Improved Cluster Status Detection: Instead of constantly checking the entire cluster status, the system prioritizes checking the cluster lease for basic online/offline status. It also introduces an extensible mechanism for users to define and implement conditions for "well-known cluster issues," allowing for customized problem detection. Out-of-tree mechanisms can also integrate with infrastructure provider APIs for deeper insights.
  • New Taint Type: preferredNoExecute: Previously, NoSchedule prevented new workloads, and NoExecute immediately evicted existing workloads without a toleration. The new preferredNoExecute taint provides a middle ground. It doesn't automatically evict workloads, as it doesn't require a toleration to stay on the cluster. Instead, when this taint appears, an eviction manager triggers the eviction process for workloads that have defined a failover policy in their propagation policy.
  • Enhanced Eviction Queue: To prevent "failure amplification," the eviction queue now incorporates a pre-scheduling check using Kamada's scheduler algorithms. Before evicting a workload, it verifies if there's a healthy alternative cluster where the workload can run. If no such place is found, the eviction is given up, preventing unnecessary disruption. The system also reuses Kamada's graceful eviction mechanism.
  1. Federated Workload Queuing: Managing diverse AI workloads (training, computing, inference) across multiple clusters presents significant queuing challenges within existing scheduler frameworks:
  • Fairness Issues: Smaller jobs often gain implicit higher priority due to being easier to schedule. More jobs, even if low priority, can consume disproportionately more resources in the scheduler's internal queues.
  • Limited Priority Classes: While priority-based scheduling in Kamada helps, a limited number of priority classes can still lead to fairness issues for workloads within the same class.
  • Preemption Dilemmas: Serving workloads typically have higher priority than training, but preempting long-running training jobs can waste significant GPU time.
  • Alignment of Two-Level Queuing: A major challenge is aligning the queue management and resource sharing concepts already implemented by in-cluster schedulers like Volcano with the federated layer's decision-making.
  • Solution: Volcano Global: This concept extends Volcano's in-cluster queue management to the federated control plane. Users can create Volcano Jobs and enable Volcano Global to map them into federated queues. This relies on Kamada's recently released scheduling suspension feature to ensure higher-priority jobs are scheduled first. Volcano Global implements basic fair sharing among queues, addressing the issues of fairness and resource contention. A current limitation is that it doesn't schedule only part of a job, meaning fair sharing is at the job level, not within a single job across different clusters.

Technical Deep Dive

▶ Watch: Three major challenges: scheduling, failover, and queuing (5:25)

The architectural choices and detailed mechanisms for Day2 optimization in multi-cluster AI infrastructure are crucial for understanding their impact.

Scheduling Mechanisms: Resource Model vs. Scheduler Estimator

The fundamental challenge in federated scheduling is maintaining a balance between the overhead of the control plane and the accuracy of placement decisions.

The Resource Model aims for efficiency by abstracting cluster status. It quantifies node resources into "grades" or aggregated metrics, effectively compressing the detailed state of each member cluster. For instance, instead of tracking every individual CPU core or GiB of memory, it might categorize nodes based on available capacity ranges. This approach is highly performant for workloads that are primarily CPU or memory intensive, such as many big data processing jobs. By having a summarized view of resources directly within the federated scheduler, it can make rapid decisions, leading to better throughput and lower scheduling latency. However, this abstraction inherently sacrifices granularity. For AI workloads, where specific GPU types, quantities, and their precise allocation (e.g., fractional GPUs, specific card models) are critical, and where network topology awareness is paramount for distributed training, the resource model's generalized view often proves insufficient, potentially leading to suboptimal placements or resource waste.

In contrast, the Scheduler Estimator prioritizes accuracy for demanding AI workloads. Each member cluster runs an estimator that is essentially a lightweight, read-only version of its in-cluster scheduler. This estimator actively watches the member cluster's API server, collecting detailed, real-time status of nodes and pods in memory. When the Kamada federated scheduler needs to place an AI workload, it doesn't rely on a pre-aggregated view. Instead, it makes RPC calls to the estimators in various member clusters. Each estimator then simulates a scheduling decision based on its current, detailed understanding of its cluster's resources and existing workloads, returning the maximum available replicas or the most suitable node configurations. While this process introduces higher scheduling latency and a more extensive memory footprint at the federated control plane (as it indirectly "stores" more detailed information via the estimators), it provides the granular awareness necessary for AI. This includes understanding GPU availability, specific hardware accelerators, and critically, the in-cluster network topology, which is vital for high-performance distributed training and inference jobs that rely on low-latency inter-GPU communication. The trade-off is deemed "worthy" because the cost of misplacing an AI workload, especially wasting expensive GPU resources, far outweighs the increased overhead of the estimator.

Flexible Cluster Failover

The revised failover mechanism addresses the shortcomings of blunt, taint-based evictions.

The first improvement lies in cluster status detection. Instead of a heavyweight mechanism constantly polling and processing the full state of member clusters, the system now primarily relies on simpler, more efficient checks like the cluster lease. This significantly reduces pressure on the API server of the federated control plane. To handle more nuanced failures, an extensible cluster problem detector is introduced. This allows platform engineers to define custom conditions for "well-known cluster issues" (e.g., specific GPU driver failures, network partition events) and translate them into custom conditions on the cluster object. Furthermore, the design allows for out-of-tree mechanisms, enabling integration with external infrastructure provider APIs or monitoring systems to gain deeper, more accurate insights into cluster health.

The core innovation in failover is the introduction of the preferredNoExecute taint. Unlike the standard NoExecute taint, which immediately evicts all pods on a node (or cluster in a federated context) that don't have a matching toleration, preferredNoExecute is "soft." Workloads do not require a toleration to continue running on a cluster with this taint. Instead, its presence acts as a signal. Workloads that wish to be evicted under such conditions must explicitly define a failover policy within their propagation policy. When a preferredNoExecute taint is applied to a cluster, the eviction manager component in Kamada identifies workloads with a defined failover policy. These workloads are then added to an eviction queue.

The eviction queue itself has been enhanced to prevent cascading failures. Before actually evicting a workload, the eviction manager performs a pre-scheduling check using Kamada's scheduler algorithms. This check simulates placing the workload in other healthy clusters. If no suitable alternative cluster can accommodate the workload, the eviction is aborted. This intelligent pre-check ensures that workloads are only evicted if they can be successfully rescheduled elsewhere, preventing a situation where an evicted workload remains unschedulable, exacerbating the original failure. The system also reuses Kamada's existing graceful eviction mechanism, allowing workloads to shut down cleanly if they support it. Looking ahead, an explicit API is being designed for customizing these failover policies, allowing administrators to choose the type of taint and the specific conditions that trigger it. Future plans also include incorporating workload priority into the failover decision, ensuring that even if resources are scarce in healthy clusters, the most critical workloads are prioritized for migration.

Federated Workload Queuing with Volcano Global

Addressing the complex demands of multi-tenancy and diverse workload types in a federated AI infrastructure requires sophisticated queuing beyond simple priority classes.

The talk highlights several persistent issues with standard scheduler frameworks: smaller jobs often get implicitly higher priority because they are easier to fit, leading to unfair resource distribution. A large number of jobs, regardless of their individual priority, can consume excessive scheduler resources. Furthermore, the limited number of priority classes in Kubernetes means that many workloads often fall into the same priority tier, leading to contention and unfairness within those tiers. The conflict between serving workloads (which demand low latency and high availability) and training workloads (which are long-running and resource-intensive, but can tolerate some delay) presents a preemption dilemma; preempting training jobs can lead to significant waste of expensive GPU compute time. Finally, the challenge lies in aligning the queue management and resource sharing concepts implemented by in-cluster schedulers like Volcano with the higher-level, federated scheduling decisions.

The solution proposed is Volcano Global, extending the established queuing capabilities of the in-cluster Volcano scheduler to the federated layer. Kamada's design, which allows for the use of custom resource definitions (CRDs) without requiring new core APIs, makes this integration relatively straightforward. When users create a Volcano Job (a CRD managed by Volcano) in the federated control plane, they can enable Volcano Global to map this job into a federated queue.

Volcano Global leverages Kamada's recently introduced scheduling suspension feature. This allows the federated scheduler to pause the scheduling of lower-priority workloads, ensuring that higher-priority jobs are processed and placed first. The initial implementation of Volcano Global focuses on providing basic fair sharing between these federated queues. This means that resources are distributed more equitably among different business teams or workload types, preventing smaller, easier-to-schedule jobs from monopolizing resources. A current limitation is that Volcano Global does not support scheduling only a part of a job to a cluster; a job is either scheduled entirely or not at all. This restricts the ability to achieve fine-grained fair sharing at a sub-job level across different clusters. Future plans for Volcano Global include introducing more advanced scheduling algorithms, implementing hierarchical queues (allowing for nested resource allocation and priority management), and ensuring better alignment with the capacity resource sharing mechanisms already present in the in-cluster Volcano scheduler.

Demo / Proof of Concept

▶ Watch: Estimator approach for multi-cluster AI workload scheduling (8:00)

The provided transcript does not include details of a live demonstration or a specific proof of concept during the talk. The speaker primarily focused on explaining the concepts, architecture, and proposed solutions for the identified challenges.

Defensive Implications

▶ Watch: Complexities and considerations for multi-cluster workload failover (9:35)

The advancements discussed in this talk offer crucial insights and actionable strategies for platform engineers, SREs, and architects responsible for securing and optimizing multi-cluster AI infrastructure.

  1. Strategic Scheduling for AI Workloads: Defenders must understand the implications of different federated scheduling mechanisms. For critical AI workloads, especially those leveraging GPUs, distributed training, or requiring specific network topologies, it is imperative to configure Kamada to utilize the Scheduler Estimator. While this incurs higher overhead, it ensures accurate placement, minimizes resource waste (particularly expensive GPU time), and prevents suboptimal performance that could impact business-critical AI services. This also means investing in robust monitoring for the estimator's performance and resource consumption.
  1. Enhancing Resilience with Flexible Failover: The revised failover mechanism provides a more granular approach to high availability. Platform teams should transition away from blunt, full-cluster eviction strategies.
  • Implement preferredNoExecute: Adopt the new preferredNoExecute taint type to enable graceful degradation and controlled workload migration during partial cluster failures.
  • Define Failover Policies: Work with AI/ML teams to define explicit failover policies within their workload's propagation policy, specifying which workloads should migrate when a cluster enters a preferredNoExecute state.
  • Custom Problem Detection: Leverage the extensible cluster problem detector to integrate custom health checks and out-of-tree monitoring (e.g., specific GPU health metrics, network fabric status) to trigger appropriate taints and failover behaviors.
  • Monitor Eviction Queue: Pay close attention to the enhanced eviction queue's pre-scheduling checks. This mechanism prevents failures from cascading by ensuring evicted workloads can be re-scheduled, but it also means that if no healthy capacity exists, high-priority workloads might remain unscheduled. Monitoring this queue and its outcomes is critical for capacity planning.
  1. Optimizing Resource Utilization with Federated Queuing: For environments supporting multiple AI teams and diverse workloads, effective resource governance is paramount.
  • Adopt Volcano Global: Implement Volcano Global to extend in-cluster queue management to the federated layer. This enables fair sharing of cluster resources among different business units or workload types, preventing resource starvation and improving overall platform utilization.
  • Leverage Scheduling Suspension: Utilize Kamada's scheduling suspension feature in conjunction with Volcano Global to enforce priority for critical AI workloads (e.g., inference serving) over less time-sensitive ones (e.g., experimental training).
  • Capacity Planning: The insights gained from federated queuing and scheduling decisions can inform better capacity planning across clusters, ensuring that sufficient resources (especially GPUs) are available to meet demand and support failover scenarios.
  1. Cost Management: By optimizing scheduling accuracy and improving failover, organizations can reduce wasted compute cycles, especially for expensive GPU resources. Efficient queuing also helps maximize the utilization of shared infrastructure, directly contributing to cost savings.

Key Takeaways

  • Multi-cluster AI infrastructure is a necessity driven by diverse business and operational needs, but it introduces significant Day2 optimization challenges.
  • The Volcano and Kamada projects provide a robust two-level architecture for managing AI and batch workloads in federated Kubernetes environments.
  • For optimal AI workload placement, the Scheduler Estimator in Kamada is crucial. Despite its higher overhead, it provides the necessary accuracy and awareness of GPU resources and network topology, preventing resource waste.
  • A more flexible failover mechanism utilizing the new preferredNoExecute taint, custom cluster problem detectors, and an enhanced eviction queue with pre-scheduling checks significantly improves resilience by enabling controlled migration without amplifying failures.
  • Federated workload queuing via Volcano Global is essential for ensuring fair resource sharing, managing priorities, and aligning in-cluster and federated scheduling decisions for diverse AI workloads.
  • Ongoing development focuses on advanced queuing algorithms, hierarchical queues, explicit APIs for policy customization, and incorporating workload priority into failover, aiming for continuous improvement in efficiency and resilience.

About the Speaker(s)

Kevin Wang is a prominent contributor to the Cloud Native Computing Foundation (CNCF) ecosystem, actively involved in multiple projects. With a background deeply rooted in scheduling, he was an early contributor to the Kubernetes project. In recent years, Kevin has dedicated his efforts to the Volcano and Kamada projects, where he focuses on enhancing the management and optimization of batch and AI workloads in distributed Kubernetes environments. His expertise lies in developing scalable and efficient solutions for complex cloud-native infrastructures.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Wang delivers a highly technical and practical deep-dive into Day2 optimization for multi-cluster AI infrastructure using Kamada and Volcano. The talk meticulously outlines the challenges of federated scheduling, failover, and workload queuing for AI workloads, presenting concrete solutions like the Scheduler Estimator, preferredNoExecute taint, and Volcano Global. This is real engineering, addressing critical pain points for anyone operating large-scale AI/ML platforms on Kubernetes, providing actionable strategies to improve efficiency, resilience, and cost-effectiveness.

Heather Calloway (CISO) — STRONG ACCEPT

This presentation by Kevin Wang provides a clear and unsentimental look at the operational realities of managing multi-cluster AI infrastructure. It directly addresses the critical Day2 challenges of cost, efficiency, and resilience, offering concrete technical mechanisms within the Kamada and Volcano projects. For any organization heavily invested in AI, these optimizations are not merely technical improvements; they are foundational to managing significant business risks related to financial outlay, service availability, and operational accountability.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025