Project Lightning Talk: Scheduling AI Workload Among Multiple Clusters - Josh Packer, Supporter

Josh Packer, Supporter

KubeCon + CloudNativeCon Europe 2025 · Project Lightning Talk

Overview

In the rapidly evolving landscape of cloud-native computing, managing and orchestrating workloads across multiple Kubernetes clusters has become a critical challenge, especially for resource-intensive applications like Artificial Intelligence (AI) and Machine Learning (ML). Josh Packer, a steering committee member for Open Cluster Management (OCM), addressed this pressing issue in his KubeCon EU lightning talk, focusing on how OCM can effectively schedule AI workloads across fleets of clusters. The talk highlighted OCM's capabilities in providing a robust, centralized framework for inventory management, workload definition, and dynamic distribution, making it an indispensable tool for organizations operating at scale.

Watch on YouTube

Visual summary for Project Lightning Talk: Scheduling AI Workload Among Multiple Clusters - Josh Packer, Supporter by Josh Packer, Supporter
Visual summary for Project Lightning Talk: Scheduling AI Workload Among Multiple Clusters - Josh Packer, Supporter by Josh Packer, Supporter

Key moments

  1. 0:00 Introduction to OCM and AI workload scheduling
  2. 0:45 OCM architecture: hub and spoke, cluster registration
  3. 1:30 Distributing workload using Manifest Work CRD
  4. 2:00 Placement CRD: dynamic cluster filtering for workloads
  5. 2:40 AI integration: KubeFlow GPU type filtering
  6. 3:15 AI integration: KubeFlow GPU resource scoring
  7. 4:00 Federated Learning and OCM's broad application

Project Lightning Talk: Scheduling AI Workload Among Multiple Clusters

Speakers: Josh Packer, Steering Committee Member for Open Cluster Management

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=9q9oTJUqQoQ

Overview

In the rapidly evolving landscape of cloud-native computing, managing and orchestrating workloads across multiple Kubernetes clusters has become a critical challenge, especially for resource-intensive applications like Artificial Intelligence (AI) and Machine Learning (ML). Josh Packer, a steering committee member for Open Cluster Management (OCM), addressed this pressing issue in his KubeCon EU lightning talk, focusing on how OCM can effectively schedule AI workloads across fleets of clusters. The talk highlighted OCM's capabilities in providing a robust, centralized framework for inventory management, workload definition, and dynamic distribution, making it an indispensable tool for organizations operating at scale.

Packer's presentation delved into the core mechanisms of OCM, particularly its Placement CRD, which acts as the intelligent arbiter for workload placement. He demonstrated how this mechanism, combined with other OCM components, can be leveraged to make informed decisions about where to deploy AI jobs based on factors like GPU availability, resource claims, and cluster health. The discussion extended to specific integrations with tools like Kueue (KubeFlow's job queueing system) and the implications for distributed AI paradigms such as federated learning, underscoring OCM's versatility beyond general application management.

The significance of this talk lies in its practical approach to a complex problem. As AI/ML adoption accelerates, the need for efficient, scalable, and resilient infrastructure to support these workloads intensifies. OCM offers a declarative, Kubernetes-native solution to distribute these demanding applications, ensuring optimal resource utilization and operational consistency across diverse cluster environments. For platform engineers, MLOps practitioners, and cloud architects grappling with multi-cluster AI deployments, Packer's insights provide a clear pathway to streamline operations and enhance system performance.

Background

▶ Watch: Introduction to OCM and AI workload scheduling (0:00)

The proliferation of Kubernetes across enterprises has led to a common architectural pattern: the deployment of multiple clusters. These clusters might be distributed across different cloud providers, on-premises data centers, or edge locations, each potentially serving distinct purposes or adhering to specific regulatory requirements. Managing applications and their lifecycle across such a disparate fleet presents significant operational overhead and complexity. This is where Open Cluster Management (OCM), a CNCF sandbox project, steps in.

OCM has been an active project for approximately five years, with significant contributions from Red Hat and other companies. Its fundamental architecture is a hub and spoke topology. In this model, a centralized hub cluster serves as the control plane for managing a fleet of remote spoke clusters. The hub maintains a comprehensive inventory of all registered spoke clusters and allows operators to define workloads centrally. These defined workloads can then be intelligently distributed to the appropriate spoke clusters.

The process begins with registering clusters into the OCM hub. This involves deploying an agent and associated add-ons onto each spoke cluster. These agents facilitate communication with the hub, allowing the hub to monitor the spoke cluster's status and capabilities. Once registered, clusters can be logically organized using Manage Cluster Sets, a Custom Resource Definition (CRD) that enables grouping clusters for targeted workload distribution. This grouping mechanism is crucial for applying policies or deploying applications to specific subsets of the cluster fleet.

Workload distribution in OCM is primarily handled via the Manifest Work CRD. This powerful CRD allows users to encapsulate any Kubernetes resource—from core types like Deployments and Replica Sets to external, custom CRDs such as Argo CD applications—and define how these resources should be deployed to the managed clusters. Complementing Manifest Work are add-ons, which are essentially controllers that extend OCM's capabilities. These add-ons can enforce policies, deploy specific applications, or gather cluster-specific information, acting as the operational backbone for advanced multi-cluster management scenarios.

The "magic" that ties these components together, particularly for dynamic workload scheduling, is the Placement CRD. This CRD enables operators to dynamically filter down the pool of available clusters based on a variety of criteria before applying a Manifest Work. This dynamic filtering capability is especially pertinent for AI workloads, which often have specific hardware requirements (e.g., GPUs) or demand high-resource availability, making intelligent placement a non-negotiable aspect of efficient operations.

Key Findings

▶ Watch: Distributing workload using Manifest Work CRD (1:30)

The central finding presented by Josh Packer is that Open Cluster Management's architecture, particularly its Placement CRD, provides a highly effective and flexible solution for orchestrating and scheduling complex AI workloads across diverse multi-cluster Kubernetes environments. This capability moves beyond simple application deployment to intelligent, resource-aware workload distribution, which is critical for the demanding nature of AI/ML tasks.

Specifically, the talk highlighted two primary integrations demonstrating OCM's utility for AI:

  1. Enhanced GPU-Aware Scheduling with Kueue: OCM's Placement CRD can be integrated with KubeFlow's Kueue (referred to as 'Q' in the transcript) to enable sophisticated GPU-aware scheduling. By leveraging cluster labels for GPU types and dynamically calculating GPU resource availability through OCM add-ons and the Placement Score, OCM can inform Kueue about the optimal clusters and nodes for specific AI jobs. This ensures that AI workloads are placed on infrastructure that meets their specific hardware requirements, maximizing efficiency and minimizing idle resources.
  1. Facilitating Federated Learning: OCM provides a robust framework for implementing federated learning strategies. By allowing precise control over workload placement, OCM enables the processing and building of AI models directly on remote fleets or within specific data centers, thereby keeping sensitive data localized. This addresses critical concerns around data privacy, regulatory compliance, and network latency in distributed AI training scenarios.

Beyond these specific AI use cases, a broader key finding is that OCM is not merely a specialized tool for AI. It serves as a general-purpose solution for bringing multi-cluster capabilities to any application or service. Its declarative nature and Kubernetes-native approach simplify the management of complex distributed systems, making it a foundational technology for organizations adopting multi-cluster strategies across their entire application portfolio.

Technical Deep Dive

▶ Watch: Placement CRD: dynamic cluster filtering for workloads (2:00)

Open Cluster Management (OCM) establishes a robust framework for managing Kubernetes clusters at scale through its hub and spoke topology. The hub cluster acts as the central control plane, maintaining an inventory of all registered spoke clusters and serving as the point of origin for workload and policy definitions. Each spoke cluster runs an agent and various add-ons that facilitate communication with the hub, collect cluster-specific data, and enforce instructions from the hub. This architecture ensures a clear separation of concerns, allowing operators to manage a vast fleet from a single pane of glass.

The initial step in leveraging OCM is cluster registration, where a spoke cluster is brought under the management of the hub. This process deploys the necessary OCM agent and its add-ons to the spoke, enabling bidirectional communication and making the spoke's resources and status visible to the hub. Once registered, spoke clusters can be logically grouped using the Manage Cluster Sets CRD. This CRD allows administrators to define arbitrary sets of clusters, which can then be targeted for specific deployments or policy applications. For instance, clusters in a "development" set might receive different configurations than those in a "production" set, or clusters with specific hardware capabilities might be grouped together.

Workload distribution is orchestrated through the Manifest Work CRD. This CRD is incredibly versatile, designed to encapsulate any Kubernetes resource definition. This includes standard Kubernetes objects like Deployments, Replica Sets, Services, and ConfigMaps, but also extends to custom resources defined by other projects, such as an Argo CD Application or a KubeFlow Job. The Manifest Work CRD ensures that the specified resources are consistently applied to the targeted clusters. The add-ons play a crucial role here, acting as specialized controllers that interpret and execute these Manifest Work definitions, potentially interacting with other cluster-local services or enforcing complex policies.

The core innovation that enables intelligent workload scheduling, particularly for AI, is the Placement CRD. This CRD empowers OCM to dynamically filter and select the most suitable clusters from the managed fleet for a given workload. The filtering criteria are highly flexible and can be combined to form sophisticated placement policies:

  • Labels: The simplest form of filtering involves using standard Kubernetes labels applied to the managed clusters. For example, a workload might require placement on clusters labeled region: us-east-1 or environment: production.
  • Cluster Claims: This mechanism allows spoke clusters to "claim" or declare specific resources, capabilities, or attributes that are then percolated up to the hub. These claims could include details about available storage types, network configurations, or specialized hardware. The Placement CRD can then use these claims for filtering, ensuring workloads land on clusters that meet their prerequisites.
  • Placement Score: This advanced feature allows for a numerical scoring of clusters based on various metrics. OCM add-ons can collect real-time data from spoke clusters—such as available CPU, memory, or, critically for AI, GPU resources—and expose this as a score. The Placement CRD can then be configured to prioritize clusters with higher scores (e.g., more available resources), enabling a form of greedy scheduling.
  • Availability: A fundamental criterion, this ensures that workloads are only placed on clusters that are currently online and reachable by the hub.

For AI workloads, the integration of OCM's Placement capabilities with specialized AI orchestration tools like KubeFlow's Kueue (referred to as 'Q' in the transcript) is particularly powerful. Kueue is designed to manage job queues and resource quotas for ML workloads within a Kubernetes cluster. OCM enhances this by extending Kueue's reach across multiple clusters.

One critical integration point is GPU-aware scheduling. AI/ML models heavily rely on Graphics Processing Units (GPUs) for accelerated computation. OCM addresses this by:

  1. Label-based GPU Type Filtering: OCM's Placement CRD can utilize labels on spoke clusters that denote specific GPU types (e.g., gpu.type: NVIDIA-A100). This information is then fed into Kueue's multi-queue configuration and cluster resource definitions. Kueue can then leverage this to determine which clusters possess the required GPU types to place jobs onto individual nodes with those GPUs.
  2. Placement Score for GPU Resource Availability: To optimize resource utilization, OCM can go beyond just type to quantity. An OCM add-on deployed on each spoke cluster is responsible for calculating the amount of available GPU resources (e.g., total GPU memory, number of free GPUs) on the nodes within that cluster. This calculated metric is then made available to the OCM hub. The Placement CRD can then use this information in its Placement Score logic, filtering for clusters that have the most GPU resources available. This refined information is then passed to Kueue, allowing it to make intelligent decisions about which cluster, and subsequently which nodes within that cluster, should receive the AI workload to maximize processing efficiency.

Beyond interactive job scheduling, OCM's Placement CRD is also highly beneficial for federated learning. Federated learning is a distributed machine learning approach that trains an algorithm on multiple decentralized edge devices or servers holding local data samples, without exchanging the data samples themselves. Instead, only model updates (e.g., weights) are exchanged. This approach is crucial for privacy-sensitive applications and scenarios where data cannot leave specific geographical or organizational boundaries. OCM, through its Placement CRD, allows operators to define the specific requirements for federated learning workloads (e.g., data residency, computational resources) and then precisely place these model training tasks on the designated remote fleets or data centers. This ensures that the data processing and model building occur exactly where they are needed, adhering to data governance policies while still leveraging the collective intelligence of distributed data.

In essence, OCM's technical capabilities, particularly the flexible and dynamic nature of its Placement CRD, transform multi-cluster management from a static, manual process into an intelligent, automated orchestration system that is perfectly suited for the complex and resource-intensive demands of modern AI/ML workloads.

Demo / Proof of Concept

▶ Watch: AI integration: KubeFlow GPU resource scoring (3:15)

While this lightning talk, constrained by its brief duration, did not feature a live, in-depth demonstration of Open Cluster Management's capabilities, the speaker, Josh Packer, emphasized that the project offers extensive demos. He encouraged interested individuals to visit openclustermanagement.io to "give it a whirl" and explore the project's features, including its multi-cluster application and AI workload scheduling functionalities. This indicates that while not shown directly, robust proof-of-concept environments and examples exist to illustrate OCM's practical application.

Defensive Implications

▶ Watch: Federated Learning and OCM's broad application (4:00)

While Open Cluster Management primarily focuses on multi-cluster orchestration and workload scheduling rather than direct security vulnerabilities, its capabilities have significant defensive implications for organizations operating distributed Kubernetes environments, especially when dealing with sensitive AI/ML workloads. Leveraging OCM can enhance an organization's security posture by enabling consistent policy enforcement, controlled resource allocation, and adherence to compliance requirements across the entire fleet.

  1. Consistent Policy Enforcement: OCM's add-ons and Manifest Work CRD allow for the centralized definition and distribution of security policies. This means that security agents, network policies, Pod Security Standards (PSS), or admission controllers can be consistently deployed and enforced across all managed clusters, regardless of their underlying infrastructure. This prevents configuration drift and ensures a uniform baseline security posture, critical for preventing misconfigurations that could lead to vulnerabilities.
  1. Controlled Workload Placement for Compliance and Data Residency: The Placement CRD is a powerful defensive tool. For AI workloads handling sensitive data, organizations often have strict compliance requirements (e.g., GDPR, HIPAA) or data residency mandates. OCM can ensure that AI models are trained only on clusters located in specific geographical regions or those certified for particular data classifications. By filtering clusters based on labels, cluster claims, or even security posture scores, OCM prevents sensitive workloads from being inadvertently deployed to non-compliant or less secure environments. This reduces the risk of data breaches and regulatory penalties.
  1. Resource Isolation and Denial-of-Service (DoS) Prevention: Intelligent placement based on Placement Score (e.g., available GPU, CPU, memory) helps prevent resource exhaustion on critical clusters. While not a direct security vulnerability, resource starvation can lead to effective denial of service for legitimate workloads. By ensuring AI jobs are distributed to clusters with ample resources, OCM contributes to the overall stability and resilience of the infrastructure, making it harder for resource-intensive workloads (malicious or otherwise) to cripple essential services.
  1. Secure Multi-Tenancy and Access Control: OCM's ability to group clusters into Manage Cluster Sets can facilitate secure multi-tenancy. Different teams or business units can be granted access to specific cluster sets, ensuring they can only deploy workloads to their designated environments. While OCM itself doesn't replace Kubernetes RBAC, it provides a higher-level abstraction for managing access to groups of clusters, complementing existing security controls.
  1. Auditability and Visibility: The centralized inventory and management capabilities of the OCM hub provide enhanced visibility into the state of the entire cluster fleet. This centralized view aids in auditing deployments, tracking changes, and quickly identifying any unauthorized or anomalous workload placements, thereby improving incident response capabilities.

In summary, OCM empowers defenders by providing the tools to architect and maintain a more secure, compliant, and resilient multi-cluster environment. Its orchestration capabilities, when applied with a security-first mindset, can significantly mitigate risks associated with distributed systems and demanding AI/ML workloads.

Key Takeaways

  • Open Cluster Management (OCM) is a powerful CNCF sandbox project that simplifies the management and orchestration of applications across multiple Kubernetes clusters using a hub and spoke topology.
  • The Placement CRD is OCM's core innovation for intelligent workload distribution, enabling dynamic filtering of clusters based on labels, cluster claims, placement scores, and availability.
  • OCM offers robust integration with AI orchestration tools like KubeFlow's Kueue, facilitating advanced GPU-aware scheduling by leveraging cluster labels for GPU types and dynamically calculating GPU resource availability via add-ons.
  • OCM provides a strong foundation for federated learning architectures, allowing AI model training to occur on remote fleets or within specific data centers, addressing data privacy and compliance requirements.
  • Beyond AI, OCM is a general-purpose solution for multi-cluster application management, enabling consistent deployment of any Kubernetes resource or custom CRD across a fleet.
  • Organizations can leverage OCM to enhance defensive postures through consistent policy enforcement, controlled workload placement for compliance, improved resource isolation, and better auditability across their distributed Kubernetes infrastructure.

About the Speaker(s)

Josh Packer is a dedicated member of the steering committee for Open Cluster Management (OCM), a CNCF sandbox project. His involvement in the project underscores his expertise and commitment to advancing multi-cluster management solutions within the cloud-native ecosystem. In his role, he contributes to the strategic direction and technical development of OCM, helping to shape its capabilities for managing complex distributed workloads, including those in the rapidly growing field of Artificial Intelligence. The talk reflects his deep understanding of OCM's architecture and its practical applications for modern infrastructure challenges.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk effectively introduces Open Cluster Management (OCM) as a robust, Kubernetes-native solution for orchestrating and scheduling demanding AI/ML workloads across multi-cluster environments. It clearly articulates how OCM's Placement CRD, in conjunction with other components, enables intelligent, resource-aware distribution, particularly for GPU-aware scheduling with Kueue and supporting federated learning paradigms. The speaker, a steering committee member, demonstrates deep expertise, offering valuable insights into streamlining operations and enhancing system performance for platform engineers and MLOps practitioners.

Heather Calloway (CISO) — STRONG ACCEPT

This lightning talk, as presented in the article, outlines how Open Cluster Management (OCM) provides a robust, Kubernetes-native solution for orchestrating and scheduling complex AI workloads across distributed clusters. While highly technical, the core mechanisms, especially the Placement CRD, offer significant implications for governance, compliance, and risk management in multi-cluster environments. It provides actionable tools for consistent policy enforcement and intelligent workload placement, directly addressing critical business and security concerns for modern AI/ML deployments.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025