Tutorial: Build, Operate, and Use a Multi-Tenant AI Cluster Base... C. Misale, O. Tardieu & D. Grove

C. Misale, O. Tardieu, D. Grove

KubeCon + CloudNativeCon Europe 2025 · Tutorial

Overview

In this comprehensive KubeCon EU session, Olivier Tardieu, Dave Grove, and Claudia Misale from IBM Research presented a detailed tutorial on building, operating, and effectively utilizing multi-tenant GPU clusters for AI and Generative AI workloads using a robust, open-source Kubernetes-native stack. The talk addresses critical challenges faced by organizations in their AI journey, from the initial, often daunting, procurement of expensive GPUs to the complex task of sharing these valuable resources efficiently and fairly across diverse teams and projects.

Watch on YouTube

Visual summary for Tutorial: Build, Operate, and Use a Multi-Tenant AI Cluster Base... C. Misale, O. Tardieu & D. Grove by C. Misale, O. Tardieu, D. Grove
Visual summary for Tutorial: Build, Operate, and Use a Multi-Tenant AI Cluster Base... C. Misale, O. Tardieu & D. Grove by C. Misale, O. Tardieu, D. Grove

Key moments

  1. 0:00 Introduction and the GPU procurement/sharing challenge
  2. 2:00 Accessing open-source tutorial resources and YAMLs
  3. 3:00 Defining the scope of AI/ML workloads for the platform
  4. 4:40 Overview of target NVIDIA GPU cluster hardware
  5. 6:00 Core platform components and rationale for Kubernetes adoption
  6. 6:40 IBM Research contributions to fill community gaps

Tutorial: Build, Operate, and Use a Multi-Tenant AI Cluster Base... C. Misale, O. Tardieu & D. Grove

Speakers: C. Misale, Principal Research Scientist and Manager; O. Tardieu, Principal Research Scientist and Manager; D. Grove, IBM Research Colleague

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=Ab7mRoJYsMo

Overview

In this comprehensive KubeCon EU session, Olivier Tardieu, Dave Grove, and Claudia Misale from IBM Research presented a detailed tutorial on building, operating, and effectively utilizing multi-tenant GPU clusters for AI and Generative AI workloads using a robust, open-source Kubernetes-native stack. The talk addresses critical challenges faced by organizations in their AI journey, from the initial, often daunting, procurement of expensive GPUs to the complex task of sharing these valuable resources efficiently and fairly across diverse teams and projects.

The core problem tackled is the need to maximize GPU utilization and return on investment while ensuring seamless multi-user access and robust fault tolerance. The speakers shared IBM Research’s incrementally refined methodology and platform setup, which leverages a combination of established open-source projects from the Cloud Native Computing Foundation (CNCF) landscape, alongside innovative components developed by IBM Research. This article delves into the architecture, key components, and practical demonstrations of this platform, providing insights into how organizations can manage intensive AI/ML workloads at scale.

The presented solution is entirely open-source, with all scripts, YAML configurations, and supporting documentation available online, enabling attendees and readers to replicate and explore the setup. By focusing on a turnkey approach for administrators and pre-baked templates for AI/ML experts, the platform aims to bridge the gap between Kubernetes operational complexity and the specific demands of AI workloads, ultimately achieving remarkable resource utilization without compromising quota guarantees for individual teams.

Background

▶ Watch: Introduction and the GPU procurement/sharing challenge (0:00)

The proliferation of AI and Generative AI has driven an unprecedented demand for specialized hardware, particularly GPUs. Organizations face significant hurdles in procuring these accelerators, a process that is often time-consuming, expensive, and slow, whether through cloud providers or direct purchase. Once acquired, the challenge shifts to effectively sharing these GPU resources across multiple teams and projects to maximize utilization and return on investment.

AI/ML workloads present unique characteristics that complicate traditional resource management. They are typically intensive, capable of utilizing many GPUs simultaneously, and long-running, often spanning hours or even days. Furthermore, they heavily depend on specialized AI/ML frameworks like PyTorch, requiring robust support for these environments. The hardware itself—high-density GPU nodes (e.g., NVIDIA A100 or H100 with eight GPUs per node), equipped with substantial compute, memory, local storage, and high-performance networks—is inherently more complex and prone to "creative ways to fail" compared to traditional compute nodes. These properties necessitate advanced capabilities for workload queuing, prioritization, quota management, fault detection, and recovery.

The choice of Kubernetes as the foundational platform is strategic, driven by the vitality of its community and the rich CNCF landscape. This ecosystem offers a vast array of projects that can be leveraged and contributed to, accelerating platform development significantly compared to alternative choices. While Kubernetes provides a powerful orchestration layer, native Kubernetes quotas and scheduling mechanisms are insufficient for the nuanced demands of AI/ML workloads. They lack workload awareness (operating at the pod level, not the job level), are inflexible for fair sharing and borrowing, and do not inherently understand GPU resources or complex distributed job topologies.

The platform developed by IBM Research aims to address these gaps. Its high-level goals include managing AI workloads and resources on Kubernetes, being entirely open-source, integrating key AI/ML frameworks, supporting multi-user environments without conflict, ensuring productive utilization of expensive GPUs, and emphasizing fault tolerance, monitoring, and observability. While the demonstrations run on OpenShift (IBM's Kubernetes flavor), the underlying principles and components are designed for vanilla Kubernetes, with much of the research work feeding into supported solutions like OpenShift AI.

Key Findings

▶ Watch: Defining the scope of AI/ML workloads for the platform (3:00)

The central discovery and contribution of this work is a validated, open-source methodology and platform stack that successfully tackles the inherent complexities of operating multi-tenant GPU clusters for AI workloads. The key findings demonstrate that:

  1. Native Kubernetes is Insufficient for AI/ML at Scale: While Kubernetes provides a strong foundation, its out-of-the-box capabilities for resource management, scheduling, and fault tolerance are inadequate for the specific demands of intensive, long-running, and highly resource-dependent AI/ML workloads, particularly when dealing with GPUs and multi-tenancy.
  2. A Layered Open-Source Stack is Essential: The solution lies in composing and extending existing open-source projects. By integrating and enhancing components like Kubernetes, Kueue, Kubeflow Trainer, Prometheus, and Grafana with IBM Research innovations such as App Wrapper, Autopilot, and specialized scheduler plugins, a comprehensive and effective platform can be built.
  3. Flexible Quota Management Drives High Utilization: Traditional strict quotas lead to significant resource waste. The Kueue project, with its workload-aware quotas, fair sharing, and preemption capabilities within cohorts, allows teams to borrow unused capacity, leading to dramatically higher overall cluster utilization (demonstrated at 99% on a 1200-GPU cluster) while still guaranteeing nominal quotas when needed.
  4. Automated Fault Tolerance is Critical for GPU Clusters: GPU nodes are prone to subtle and severe failures that can degrade performance or crash jobs. The combination of Autopilot for proactive health checks and node labeling, coupled with App Wrapper for automated workload reset and re-scheduling with anti-affinity, significantly reduces human intervention and improves workload resilience.
  5. User Experience is Paramount for Adoption: Abstracting Kubernetes complexity for AI/ML experts through pre-baked workload templates (for pre-training, fine-tuning, inference, data pre-processing) and providing a turnkey deployment for administrators are crucial for widespread adoption and productivity. The system aims to allow AI experts to focus on their models, not infrastructure.
  6. Proactive Monitoring Enhances Reliability: Beyond basic metrics, specialized health checks provided by Autopilot reveal underlying issues in GPUs, networks, and storage that traditional monitoring might miss, enabling earlier detection and mitigation of performance degradation or potential failures.

These findings collectively present a robust, scalable, and highly efficient solution for organizations looking to operationalize their AI infrastructure in a shared, multi-tenant environment.

Technical Deep Dive

▶ Watch: Overview of target NVIDIA GPU cluster hardware (4:40)

The IBM Research platform for multi-tenant AI GPU clusters is built on a layered architecture, combining core Kubernetes functionalities with specialized open-source projects and custom components. The following details the key technical elements and their roles:

Cluster Setup Prerequisites

Before deploying the ML Batch platform, two fundamental prerequisites must be met to prepare a vanilla Kubernetes cluster for AI workloads:

  1. NVIDIA GPU Operator: Kubernetes, by default, understands CPU, memory, and ephemeral storage. To manage NVIDIA GPUs, the NVIDIA GPU Operator must be installed. This operator, jointly developed by Red Hat and NVIDIA, extends Kubernetes by adding an extended resource type, nvidia.com/gpu. Once installed, Kubernetes can recognize and schedule workloads requesting GPUs. The installation typically involves deploying a Helm chart, which automates the setup of necessary drivers, container runtimes, and monitoring tools.
  2. High-Performance Storage Solution: AI workloads are data-intensive, necessitating a robust storage solution. While IBM Research uses IBM Spectrum Scale for high-performance needs, the demo simplifies this by assuming an NFS server on a local subnet. The platform requires a storage class to be available on the cluster, enabling the creation of Persistent Volume Claims (PVCs) for data access. This ensures models, datasets, and checkpoints can be efficiently stored and retrieved.

ML Batch Cluster Setup (Admin)

The administrative setup of the ML Batch cluster involves a series of steps to configure the core components:

  1. Repository Cloning: The entire setup is open-source, with all necessary scripts and YAMLs available in a public repository.
  2. Priority Classes: To manage workload prioritization, the platform establishes Priority Classes (low, medium, high) within Kubernetes. This allows administrators to define the relative importance of different AI jobs, influencing their queuing and preemption behavior.
  3. Component Deployment: Key ML Batch components are deployed, including:
  • Scheduler Plugins: Custom Kubernetes scheduler extensions designed to optimize GPU workload placement.
  • Operators: Such as Kueue, Kubeflow Trainer, App Wrapper, and Autopilot, which manage specific aspects of AI workload orchestration and cluster health.
  1. Kueue Configuration: Kueue, the Kubernetes-native queuing and quota management system, is configured. This involves defining ClusterQueues and LocalQueues, which are central to managing resource allocation and sharing.
  2. RBAC Role Creation: A specific Role is created that grants users the necessary permissions to run, monitor, and debug workloads within their designated namespaces.
  3. Admin Queue Reservation: A portion of the cluster capacity is reserved as an "admin queue." This dedicated queue allows administrators to perform maintenance operations without disrupting user workloads, providing a safety buffer.

Team and User Setup

Onboarding new teams and users is streamlined:

  1. Namespace per Team: Each team is assigned its own Kubernetes namespace, logically isolating their resources and workloads.
  2. LocalQueue and ClusterQueue: A LocalQueue is created within the team's namespace, serving as the default submission point for their jobs. This LocalQueue is associated with a ClusterQueue, which defines the nominal GPU quota for the team (e.g., 8 GPUs for "team blue"). The term "nominal" is crucial, as teams can borrow beyond this quota if resources are available, thanks to Kueue's flexible sharing mechanisms.
  3. RBAC Assignment: Individual users (e.g., Alice in "team blue") are assigned the previously created RBAC role within their team's namespace, granting them operational permissions.
  4. Namespace Labeling: Namespaces are labeled to indicate that they are managed by ML Batch, enabling the platform's operators to identify and manage workloads within them.

Scheduler Plugins for Optimization

The platform employs specialized Kubernetes scheduler plugins to enhance resource utilization and workload performance:

  1. Node Resource Fit Scheduler Plugin: This plugin aims to prevent fragmentation. If a workload requires one GPU, the scheduler attempts to place it on a node with a single GPU remaining, rather than an empty node. This preserves larger contiguous blocks of GPUs for bigger, more demanding jobs.
  2. Gang Scheduling: For large, distributed AI/ML workloads that require many pods (e.g., 256) to run concurrently, gang scheduling ensures that all required pods can be allocated resources before any of them are deployed. This prevents partial deployments that would be useless for most AI workloads, ensuring atomic resource allocation.
  3. Topology Aware Scheduling (Experimental): This plugin, an IBM Research innovation, aims to optimize communication for distributed workloads. It tries to distribute pods across as few racks as possible to maximize network bandwidth and minimize inter-node communication latency, which is critical for high-performance AI training.

Kueue: Advanced Queuing and Quota Management

Kueue is a cornerstone of the ML Batch platform, providing Kubernetes-native capabilities for managing queuing, quotas, and jobs in a workload-aware manner:

  • Workload Awareness: Unlike native Kubernetes quotas that operate at the pod level, Kueue manages quotas at the level of top-level jobs (e.g., PyTorch jobs, Ray jobs).
  • Admission Control and Suspension: When a user submits a job, Kueue interposes via webhooks. If the job is managed by Kueue, it is initially suspended by modifying its spec. The job remains suspended until Kueue determines it has sufficient quota and resources, at which point the suspend bit is flipped to false, allowing the underlying controller to create pods.
  • Fair Sharing and Preemption: Kueue enables fair sharing of resources, prioritizes workloads, and can preempt running jobs if higher-priority or guaranteed-quota jobs require resources.
  • External Extensibility: Kueue's design allows for customization and integration with external systems, such as the cluster autoscaler, to dynamically provision capacity.
  • ClusterQueue and Cohorts: The ClusterQueue is the primary abstraction for assigning quota. Multiple ClusterQueues can be grouped into a cohort, allowing them to borrow quota from each other when unused. This flexible borrowing mechanism significantly improves overall utilization. Kueue also supports hierarchical cohorts, allowing the quota structure to mirror organizational hierarchies.
  • LocalQueue: A LocalQueue is a namespace-scoped object that links jobs within a namespace to a specific ClusterQueue.

App Wrapper: Workload Hardening and Grouping

The App Wrapper, an IBM Research project, is crucial for both grouping resources and enhancing workload resilience:

  • Resource Grouping: App Wrapper allows users to logically group multiple Kubernetes resources (e.g., a PyTorch job, associated services, secrets, ingresses) into a single, manageable unit. This simplifies deployment and ensures all related resources are created and cleaned up together. It supports all Kubernetes resource types with a pod spec template, making it highly versatile.
  • Kueue Integration: App Wrapper implements Kueue's suspension protocol, allowing Kueue to manage the lifecycle of complex, multi-resource workloads as a single entity.
  • Workload Hardening and Fault Recovery: App Wrapper is a key component of the fault recovery process. It monitors workload health, detects various failure conditions (e.g., failed pods, pods on unhealthy nodes, underlying job failures, user deletion of components), and automates a uniform retry loop. When a fault is detected, App Wrapper gently (then forcibly) deletes the associated resources, cleans them up, and then re-creates them, effectively resetting and resuming the workload. This loop is configurable by administrators and users (within bounds).

Autopilot: Proactive Health Checks

Autopilot, another IBM Research innovation, augments traditional monitoring by providing advanced, proactive health checks:

  • Beyond Native Metrics: While Prometheus, Grafana, and NVIDIA's DCGM exporter provide valuable metrics, Autopilot delves deeper. It runs periodic health checks on GPUs, network, and storage, looking for subtle performance degradations or latent failures that might not trigger standard alerts but still impact AI workload performance (e.g., reduced tokens per second).
  • Leveraging NVIDIA Tools: Autopilot packages various nvidia-smi and dcgmi commands into automated tests, exposing critical GPU health information not typically exported by DCGM exporter.
  • Node Labeling: The results of Autopilot's health checks are used to label worker nodes as healthy, unhealthy, or evicted. These labels are then consumed by other components (like App Wrapper) to make intelligent scheduling and recovery decisions.

Dynamic Resource Adaptation

The platform dynamically adapts to changing cluster conditions and resource availability:

  • Slack ClusterQueue Adjustment: A reconciler continuously monitors node health. When Autopilot labels a node as unhealthy, this reconciler adjusts the lending limit on a designated "slack" ClusterQueue. This ensures that the total available quota for borrowing accurately reflects the healthy, usable capacity of the cluster, preventing new jobs from being admitted to failing resources.
  • Node Anti-Affinity Injection: App Wrapper automatically injects node anti-affinity rules into workloads. When a job is reset and resumed due to a node failure, these anti-affinity rules steer the new pods away from the unhealthy nodes, ensuring they are scheduled only on healthy infrastructure.
  • Maintenance Integration: The system treats maintenance operations (e.g., an admin cordoning a node) in the same way it treats an autopilot-detected unhealthy node. This means new workloads are steered away, and existing workloads on those nodes are managed for graceful migration or reset.

Demo / Proof of Concept

▶ Watch: Core platform components and rationale for Kubernetes adoption (6:00)

The talk featured several recorded demos illustrating the platform's capabilities, primarily using a small 6-node cluster (3 control plane, 3 worker) with 24 NVIDIA GPUs (8 GPUs per worker node).

Quota and Preemption with Kueue

This demo showcased Kueue's ability to manage quotas, allow borrowing, and perform preemption:

  1. Initial Burst: Alice (Team Blue, nominal quota 8 GPUs) submitted a burst of seven short-running jobs, each requiring 8 GPUs (total 56 GPUs). On the 24-GPU cluster, only three jobs could run initially (24 GPUs utilized), with the rest queued. Crucially, Alice was able to utilize the entire cluster capacity (24 GPUs) because the Red Team hadn't submitted any jobs yet, demonstrating quota borrowing. The Kueue Viz tool provided real-time visualization of cluster queues and workload states.
  2. Priority Preemption: Alice then submitted an "important" high-priority job. Kueue immediately suspended one of Alice's lower-priority, already running jobs to make room for the high-priority one, showcasing intra-team priority management.
  3. Cross-Team Preemption: Bob (Team Red, nominal quota 8 GPUs) then submitted a 16-GPU workload. Because Alice was using 8 GPUs borrowed from the Red Team's nominal quota, Kueue suspended one of Alice's running jobs to reclaim the quota for Bob, allowing Bob's job to start. This highlighted Kueue's ability to enforce nominal quotas when the owning team requires them, even if it means preempting a borrowing team.

Fault Detection and Observability with Autopilot

This section demonstrated the proactive fault detection and monitoring capabilities:

  1. Autopilot Installation: The demo showed the installation of Autopilot via a Helm chart. Autopilot was configured to run periodic health checks on GPUs, network, and storage, including checking PVC creation/deletion for the NFS client.
  2. Node Labeling: Autopilot logs showed it running health checks and labeling nodes as pass (healthy) for GPUs and overall node status.
  3. Monitoring Stack Deployment: The KubePrometheus stack (Prometheus, Grafana, Alert Manager, Node Exporter, Kube State Metrics) was deployed. Specific configuration overrides were applied for OpenShift.
  4. Grafana Dashboards: The demo illustrated importing custom Grafana dashboards for Autopilot health checks and the NVIDIA DCGM dashboard, providing comprehensive visualization of GPU status, node health, and workload metrics.

Automated Fault Recovery with App Wrapper

This demonstration highlighted App Wrapper's role in automated fault tolerance:

  1. Large Job Submission: Alice submitted a single large job requiring 16 GPUs on the 24-GPU cluster. The job started running.
  2. Simulated Fault: An administrator manually applied an autopilot.io/evicted=true label to one of the worker nodes, simulating Autopilot detecting a severe GPU fault.
  3. Automated Reset and Resumption: Within 30-45 seconds, the system reacted. The App Wrapper controller detected that a pod of Alice's job was running on the now evicted node and using an unhealthy GPU. It then triggered a resetting state for the App Wrapper, deleting the affected pods and resources. After cleanup, the App Wrapper automatically resumed the workload.
  4. Anti-Affinity in Action: The new pods were scheduled on the remaining two healthy nodes, demonstrating that App Wrapper had automatically injected anti-affinity rules to steer workloads away from the unhealthy node, all without any human intervention.

Example Workloads

The speakers presented three representative AI workloads running on the platform:

  1. Kubeflow Trainer (PyTorch FSDP):
  • Purpose: Fine-tuning and scalable distributed training of LLMs using PyTorch.
  • Demo: A PyTorch job configured for Fully Sharded Data Parallel (FSDP) training was wrapped inside an App Wrapper for enhanced fault tolerance. The YAML for this job was generated using a Kubeflow Trainer notebook and then automatically wrapped. The demo showed the job starting, running with two GPUs per pod, and completing, with logs confirming the training process.
  1. KubeRay (Fine-tuning Llama 3.1 8B with LoRA):
  • Purpose: End-to-end ML lifecycle (data preprocessing to model serving) using Apache Ray, orchestrated by KubeRay on Kubernetes.
  • Demo: After installing Ray CLI and setting up a PVC for model storage (impersonating Alice in the blue namespace), a Ray job was configured to fine-tune a Llama 3.1 8B model using LoRA (Low-Rank Adaptation) on eight GPUs (one for the Ray head, seven for workers). The job, wrapped in an App Wrapper, was submitted via the Ray API. The Ray dashboard was used to monitor job progress, task execution, and GPU utilization, showing the model being downloaded to the PVC and the fine-tuning process completing.
  1. Batch Inference (vLLM):
  • Purpose: Running batch inference for LLMs and collecting performance statistics.
  • Demo: An App Wrapper contained a Kubernetes job with two containers: one running vLLM (an inference runtime from Berkeley, now Linux Foundation) serving an IBM Granite model, and another running a load generator. A persistent volume cached the large model weights from Hugging Face, avoiding repeated downloads. The load generator sent random requests, and the demo showed the logs of the vLLM server loading the model, the load generator sending requests, and finally, the collection of performance statistics, allowing for comparison of model behavior. Synchronization mechanisms ensured the server was ready before requests began.

These demos collectively illustrated the platform's ability to handle diverse AI workloads, manage resources efficiently, and recover from failures autonomously.

Defensive Implications

▶ Watch: IBM Research contributions to fill community gaps (6:40)

The detailed technical article and conference talk provide crucial insights and actionable strategies for defenders and infrastructure teams managing AI/ML workloads. The defensive implications revolve around maximizing resource utilization, ensuring workload resilience, and abstracting complexity:

  1. Proactive GPU Health Monitoring is Non-Negotiable: Defenders must move beyond basic node health checks. Deploying tools like Autopilot in conjunction with NVIDIA DCGM exporter, Prometheus, and Grafana is essential for detecting subtle GPU degradations or latent hardware faults before they manifest as catastrophic job failures. Proactive monitoring allows for predictive maintenance and prevents wasted compute cycles on underperforming hardware.
  2. Implement Workload-Aware Quota and Queuing: Relying solely on native Kubernetes quotas is inefficient for shared GPU clusters. Defenders should adopt Kueue to implement flexible, workload-aware quotas. This enables fair sharing among teams, allows for borrowing of unused capacity (maximizing utilization), and provides preemption mechanisms to prioritize critical workloads, ensuring high-value jobs get the resources they need.
  3. Embrace Automated Fault Recovery with App Wrappers: The App Wrapper pattern is a critical defensive measure against various workload and infrastructure failures. Defenders should advocate for wrapping all AI/ML workloads within App Wrappers to gain automated detection of pod/job failures, automatic cleanup of lingering resources, and intelligent resetting and rescheduling. This significantly reduces manual intervention, improves mean time to recovery (MTTR), and enhances overall workload stability.
  4. Dynamically Adapt to Cluster Health: Integrate node health status (e.g., labels from Autopilot) with the scheduling and quota system. By dynamically adjusting available capacity in Kueue's slack ClusterQueues and injecting node anti-affinity into workloads, defenders can ensure new jobs are steered away from unhealthy nodes and existing jobs are gracefully migrated or restarted on healthy infrastructure. This maintains cluster health and prevents resource allocation to failing components.
  5. Standardize Workload Templates for AI/ML Experts: To reduce operational burden and potential misconfigurations, defenders should provide a curated set of pre-baked workload templates for common AI/ML tasks (training, fine-tuning, inference, data preprocessing). These templates, integrated with App Wrappers and Kueue, abstract away Kubernetes complexities, allowing AI/ML experts to focus on their domain without becoming Kubernetes specialists.
  6. Treat Maintenance as a Managed Unhealthiness: Establish protocols where administrative maintenance (e.g., cordoning nodes) is integrated into the automated fault recovery system. The platform should react to maintenance operations similarly to how it reacts to detected hardware failures, ensuring workloads are gracefully moved or paused, and new jobs are not scheduled on nodes undergoing maintenance.
  7. Leverage Open-Source Community and Contributions: The entire platform is built on open-source components. Defenders should actively engage with and contribute to projects like Kueue, Kubeflow, and the broader CNCF ecosystem. This collaborative approach ensures the platform remains cutting-edge, secure, and well-supported.

By implementing these defensive strategies, organizations can transform their GPU clusters from expensive, fragile assets into highly utilized, resilient, and manageable platforms for their AI initiatives, minimizing operational overhead and maximizing the value of their hardware investment.

Key Takeaways

  • GPU clusters for AI are complex and expensive: Managing these resources efficiently for multi-tenant AI/ML workloads requires specialized solutions beyond native Kubernetes.
  • A layered open-source stack is key to success: Combining established CNCF projects (Kubernetes, Kueue, Kubeflow Trainer, Prometheus, Grafana) with IBM Research innovations (App Wrapper, Autopilot, scheduler plugins) provides a robust and flexible platform.
  • Kueue dramatically improves GPU utilization: Its workload-aware quotas, fair sharing, and preemption capabilities within cohorts enable high resource utilization (e.g., 99% on a 1200-GPU cluster) while guaranteeing nominal quotas for individual teams.
  • Automated fault recovery is critical for AI workloads: App Wrapper, coupled with Autopilot's proactive health checks and node labeling, provides automated detection, reset, and rescheduling of workloads, significantly reducing downtime and human intervention.
  • Proactive monitoring of GPU health is essential: Autopilot extends traditional monitoring by detecting subtle performance degradations and latent hardware faults in GPUs, networks, and storage, ensuring optimal workload performance and cluster health.
  • Abstracting complexity benefits both admins and users: A turnkey deployment for administrators and pre-baked workload templates for AI/ML experts streamline operations and boost productivity, allowing each role to focus on their core competencies.

About the Speaker(s)

Olivier Tardieu is a Principal Research Scientist and Manager at IBM Research. He has been instrumental in developing and refining the methodologies and platform setup for building and operating multi-tenant GPU clusters for AI at IBM Research for several years. His work focuses on optimizing resource utilization and management for large-scale AI workloads.

Dave Grove is a Principal Research Scientist and Manager at IBM Research. As a colleague of Olivier Tardieu, he has contributed significantly to the technical development and implementation of the AI cluster management platform, particularly in areas such as Kueue integration and App Wrapper functionalities.

Claudia Misale is an IBM Research colleague who presented on fault detection and observability within the AI cluster. Her expertise lies in ensuring the reliability and performance of GPU infrastructure through advanced health checks and monitoring tools, including the development of Autopilot.

Together, their team at IBM Research has been at the forefront of tackling the challenges of AI infrastructure, contributing open-source components and sharing their practical experience to benefit the broader cloud-native and AI communities.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This tutorial from IBM Research delivers a no-nonsense, deeply technical blueprint for tackling the most expensive and frustrating problem in modern AI infrastructure: efficiently sharing and managing multi-tenant GPU clusters on Kubernetes. They present a robust, open-source stack that doesn't just promise high utilization and fault tolerance, but demonstrates it with concrete examples and custom-built components. This is real engineering solving a real problem, not just another marketing deck.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon tutorial from IBM Research presents a robust, open-source methodology for building and operating multi-tenant GPU clusters for AI workloads. While deeply technical, the talk's demonstration of achieving 99% GPU utilization through intelligent queuing (Kueue) and automated fault recovery (App Wrapper, Autopilot) has profound implications for business impact and operational resilience. It offers concrete, actionable strategies for infrastructure leaders and CISOs to maximize return on expensive AI hardware investments and ensure the stability of critical AI initiatives.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025