Extending Kubernetes for AI | Lessons Learned From Platform... - Susan, Lucy, Andrea, Etienne, Tim

Susan, Lucy, Andrea, Etienne, Tim

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This KubeCon EU panel discussion, "Extending Kubernetes for AI," brought together a diverse group of experts from Google Cloud, Uber, Weaviate, Overstory, and SchedMD to explore the intricate challenges and innovative solutions in adapting Kubernetes for the demanding landscape of Artificial Intelligence and Machine Learning (AI/ML) workloads. Moderated by Susan Woo from Google Cloud, the panel delved into how organizations, ranging from hyper-scale enterprises to agile startups, are customizing and extending Kubernetes to meet the unique requirements of AI/ML, covering everything from resource allocation and stateful applications to multi-tenancy and the economics of GPU management.

Watch on YouTube

Visual summary for Extending Kubernetes for AI | Lessons Learned From Platform... - Susan, Lucy, Andrea, Etienne, Tim by Susan, Lucy, Andrea, Etienne, Tim
Visual summary for Extending Kubernetes for AI | Lessons Learned From Platform... - Susan, Lucy, Andrea, Etienne, Tim by Susan, Lucy, Andrea, Etienne, Tim

Key moments

  1. 0:00 Panelist introductions and talk topic overview
  2. 2:00 Uber's platform for AI/ML workloads on Kubernetes
  3. 4:00 Contrasting platform team structures and focus
  4. 5:00 Discussion on Kubernetes customizations and adjacent projects
  5. 5:50 Overstory's AI use case: preventing wildfires with satellite data

Extending Kubernetes for AI | Lessons Learned From Platform...

Speakers: Susan Woo (Product Manager, Google Cloud); Etienne (Co-founder & CTO, Weaviate); Lucy (Engineer, Uber); Andrea (Overstory); Tim Wickberg (Chief Technical Officer, SchedMD)

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=d9K5PSsHtDg

Overview

This KubeCon EU panel discussion, "Extending Kubernetes for AI," brought together a diverse group of experts from Google Cloud, Uber, Weaviate, Overstory, and SchedMD to explore the intricate challenges and innovative solutions in adapting Kubernetes for the demanding landscape of Artificial Intelligence and Machine Learning (AI/ML) workloads. Moderated by Susan Woo from Google Cloud, the panel delved into how organizations, ranging from hyper-scale enterprises to agile startups, are customizing and extending Kubernetes to meet the unique requirements of AI/ML, covering everything from resource allocation and stateful applications to multi-tenancy and the economics of GPU management.

The conversation highlighted that while Kubernetes has become the de facto standard for microservices orchestration, its application to AI/ML presents a distinct set of complexities. These include the specialized hardware needs of GPUs, the often bursty or long-running nature of training and inference jobs, and the critical importance of data management for stateful AI platforms. The panelists shared real-world experiences, shedding light on the evolving role of platform engineering, the cutting-edge developments in upstream Kubernetes, and the strategic decisions companies make to balance flexibility, cost, and performance in their AI/ML infrastructure. The discussion underscores a pivotal moment in the adoption of Kubernetes for AI, demonstrating both its formidable capabilities and the ongoing innovations required to fully unlock its potential in this rapidly advancing field.

Background

▶ Watch: Panelist introductions and talk topic overview (0:00)

Kubernetes has revolutionized the deployment and management of cloud-native applications, primarily excelling at orchestrating stateless microservices. However, the burgeoning field of AI/ML introduces a host of new challenges that push the boundaries of conventional container orchestration. AI/ML workloads, encompassing data preprocessing, model training, and inference, often necessitate specialized hardware like Graphics Processing Units (GPUs), substantial memory allocations, and high-performance storage. Furthermore, these workloads can vary dramatically in their lifecycle—from short, bursty inference requests to long-running, computationally intensive training jobs that can span days or even weeks.

Historically, high-performance computing (HPC) environments, often managed by schedulers like Slurm, have been the domain for such demanding computational tasks. The integration of AI/ML into a Kubernetes-native ecosystem, therefore, requires a fundamental shift in how resources are managed, scheduled, and made available. The problem isn't merely about running containers; it's about efficiently allocating scarce and expensive resources like GPUs, managing state for data-intensive applications (such as vector databases), ensuring robust fault tolerance for lengthy jobs, and providing a scalable, multi-tenant platform for diverse user groups.

The panel reflected this diversity in challenges and solutions. Lucy from Uber, representing one of the largest Kubernetes deployments globally, spoke to the complexities of managing AI/ML workloads at an immense scale, including both training and inference for critical and non-critical services. Etienne from Weaviate, a vector database company, highlighted the unique demands of stateful AI data platforms and the bursty nature of embedding creation and transformation agents. Andrea from Overstory, a startup using satellite imagery for wildfire prevention, illustrated how even smaller organizations leverage Kubernetes (specifically GKE and Cloud Run) for highly variable, data-intensive processing pipelines. Tim Wickberg from SchedMD offered insights from the perspective of a tooling vendor, focusing on bridging the gap between traditional HPC schedulers like Slurm and Kubernetes to optimize multi-node AI training. This varied perspective underscores that while the core problem is universal—how to run AI/ML on Kubernetes—the specific implementations and priorities differ significantly based on company size, existing infrastructure, and the nature of the AI/ML tasks themselves.

Key Findings

▶ Watch: Uber's platform for AI/ML workloads on Kubernetes (2:00)

The panel discussion revealed several critical insights into the current state and future direction of extending Kubernetes for AI/ML workloads:

  1. Evolving Platform Engineering Models: The concept of a "platform team" is highly contextual. Uber operates a substantial internal platform team (140+ engineers in one office alone) that provides a stateless compute platform with built-in guardrails for safe deployments, upon which their AI/ML platform is then built. In contrast, Weaviate’s platform team primarily focuses on provisioning and running their SaaS vector database for customers, a specialized internal function. Overstory, as a smaller startup, adapts its platform team to fit its specific business use cases, highlighting the need for flexibility rather than a one-size-fits-all approach.
  1. Kubernetes is Rapidly Advancing for AI/ML: Upstream Kubernetes development is actively addressing AI/ML specific challenges. Key developments include Dynamic Resource Allocation (DRRA), a beta feature enabling fine-grained resource management, and in-place pod vertical scaling, allowing CPU and memory requests/limits to be changed on a running pod without disruption. The introduction of pod-level resources, which allows sharing resources among multiple containers within the same pod, further enhances efficiency for complex AI/ML workloads. This signifies a shift towards treating pods less like "cattle" and more like "pets" for advanced, stateful, and long-running applications.
  1. GPU Management is a Major Economic and Technical Challenge: GPUs are extremely expensive, leading to fundamental shifts in resource management. At Uber's scale, GPU pools are largely fixed, with autoscaling being impractical due to cloud provider capacity limitations and long lead times (up to three months). The imperative is to maximize GPU utilization, even leveraging idle capacity for disruption-tolerant batch jobs, rather than acquiring more devices. This requires sophisticated scheduling and checkpointing mechanisms.
  1. Workload Characteristics Dictate Infrastructure Strategy: AI/ML workloads exhibit diverse patterns. Weaviate categorizes its loads into three types: stateful database operations (benefiting from vertical scaling), embedding creation (multi-tenant, GPU-attached pods with predictable peaks), and extremely unpredictable, bursty transformation agents (like LLM calls), which they outsource to third-party providers like Modal due to the difficulty of self-hosting such volatile demands on Kubernetes. Overstory's satellite image processing pipelines also exemplify highly variable resource needs, from 1 CPU/2GB to 72 CPUs/450GB, underscoring the need for flexible resource allocation and robust checkpointing.
  1. Multi-Node and Stateful Workload Orchestration Remain Areas of Active Development: While progress is being made, challenges persist in efficiently coordinating multi-node AI training workloads and ensuring robust performance for stateful applications like databases on Kubernetes. Tim Wickberg's work with Slinky aims to bring Slurm's multi-node scheduling expertise to Kubernetes, while Etienne highlighted the "game-changer" potential of in-place vertical scaling for stateful vector databases. Lucy also pointed out that multi-node inference with related pods is not yet "the funnest thing" to do in Kubernetes.

These findings collectively illustrate that extending Kubernetes for AI/ML is not a trivial task but an ongoing evolution, driven by both community contributions and the pressing needs of diverse industry players.

Technical Deep Dive

▶ Watch: Contrasting platform team structures and focus (4:00)

The panel provided a rich technical discussion on the specific customizations, tools, and Kubernetes features being leveraged and developed to support AI/ML workloads.

Customizations and Adjacent Projects:

  • Overstory's Data Pipelines: Andrea detailed Overstory's use of Google Cloud, specifically GKE and Google Cloud Run, for their vegetation management platform. Their core process involves merging high-resolution satellite imagery (up to 15 cm resolution) from providers like Airbus and Maxar with utility data to create risk profiles for wildfires. The flexibility of Kubernetes allows them to handle workloads requiring vastly different resources, from a single CPU and 2 GB of memory to 72 CPUs and 450 GB of memory. A critical open-source project in their stack is Daxter, a data workflow engine for which Andrea is a contributor. Daxter plays a vital role in orchestrating their complex, data-intensive pipelines on Kubernetes, demonstrating how specialized workflow tools integrate seamlessly with Kubernetes for AI/ML data processing.
  • Weaviate's Vector Database Infrastructure: Etienne from Weaviate explained how their vector database platform, which functions almost like an in-memory database, runs statefully on Kubernetes. They manage three categories of AI workloads:
  1. Stateful Database Operations: These are bound by disk throughput at startup and require the ability to vertically scale.
  2. Embedding Creation: This involves converting multimodal data into vector embeddings using models, typically requiring GPU-attached pods. Weaviate runs a multi-tenant "embedding service" for this, managing burstiness through shared infrastructure and dynamic scaling based on demand patterns (e.g., email client peaks).
  3. Transformation Agents: These are extremely bursty AI agents that perform operations like mass translations using LLM model calls. Due to their unpredictable and extreme burstiness, Weaviate opts to use a third-party provider, Modal, which charges for CPU minutes, rather than attempting to self-host these on Kubernetes. This highlights a strategic decision to outsource workloads that are excessively challenging for a self-managed Kubernetes platform.
  • SchedMD's Slurm Integration: Tim Wickberg discussed Slinky, a project by SchedMD (the principal developers of Slurm) designed to integrate Slurm's scheduling capabilities into the Kubernetes stack. The goal is to bring Slurm's "wherewithal" for multi-node AI training workloads to Kubernetes, aiming to natively process training on Kubernetes on bare metal, leveraging Slurm for its specialized scheduling expertise. This represents an effort to bridge the gap between traditional HPC and cloud-native orchestration for the most demanding AI training scenarios.

Advancements in Resource Allocation:

The panel emphasized several cutting-edge Kubernetes features critical for AI/ML:

  • Dynamic Resource Allocation (DRRA): This beta feature in Kubernetes enables very fine-grained allocation of resources. Tim highlighted its potential, and Lucy, an upstream contributor, acknowledged the significant development effort (a dozen pull requests in the last month). DRRA is designed to allow workloads to request and consume specific, heterogeneous resources more efficiently, moving beyond simple CPU/memory requests to specialized hardware like GPUs or FPGAs.
  • In-place Pod Vertical Scaling: Introduced in a recent Kubernetes release, this feature allows changing the CPU and memory requests and limits of a pod without stopping the application. This is a "game-changer" for AI/ML training, batch jobs, and stateful applications like Weaviate's database, as it eliminates disruptive restarts when resource needs fluctuate. It allows for dynamic adaptation to workload demands, improving efficiency and reducing downtime.
  • Pod-level Resources: This enhancement allows resources to be shared among multiple containers within the same pod, managed under a single Cgroup. Previously, resources were defined at a per-container level. This change is particularly useful for AI/ML workloads where multiple containers might collaborate on a single task (e.g., a training workload with a data loader and a model trainer), allowing for more flexible and efficient resource balancing within a cohesive unit.

These advancements collectively demonstrate a significant push within the Kubernetes community to make the platform more adaptable and performant for the complex and resource-intensive demands of AI/ML, moving beyond its initial design for stateless microservices. The shift towards treating pods as "pets" rather than just "cattle" reflects the growing need for specialized handling of long-running, stateful, and highly configured AI/ML applications.

Demo / Proof of Concept

▶ Watch: Discussion on Kubernetes customizations and adjacent projects (5:00)

The panel discussion did not include a live demonstration or a proof of concept. The format was a moderated expert panel sharing experiences and insights.

Defensive Implications

▶ Watch: Overstory's AI use case: preventing wildfires with satellite data (5:50)

The detailed technical discussions from the panel offer several crucial implications for security and operations teams (defenders) working with Kubernetes for AI/ML workloads:

  1. Optimize GPU Utilization and Capacity Planning: The high cost and long lead times (up to three months for cloud GPUs) for specialized hardware like GPUs necessitate meticulous capacity planning and aggressive utilization strategies. Defenders must implement robust monitoring and scheduling to ensure these expensive resources are always active. Uber's strategy of running "disruption-tolerant workloads" on idle GPU capacity highlights a need for sophisticated workload management that can preempt low-priority jobs for higher-priority, bursty demands. This also implies the need for strong isolation mechanisms to prevent performance degradation for critical workloads when sharing GPUs.
  1. Implement Robust Checkpointing and Disruption Tolerance: For long-running AI/ML training or data processing pipelines, such as Overstory's satellite image processing, the risk of failure due to out-of-memory (OOM) errors or preemption is high. Defenders must advocate for and implement comprehensive checkpointing mechanisms within AI/ML applications. This ensures that progress is saved periodically, allowing workloads to resume from the last successful state rather than restarting from scratch, thereby minimizing wasted compute and data loss. This is particularly vital in environments where resources might be opportunistically reallocated.
  1. Strengthen Isolation and Address Noisy Neighbor Issues: In multi-tenant Kubernetes environments running AI/ML, especially with shared GPUs, "noisy neighbor" problems can significantly impact performance and reliability. Lucy from Uber explicitly mentioned this as a "really tough problem." Defenders need to explore and implement advanced isolation techniques, potentially leveraging Kubernetes features like resource quotas, limit ranges, and pod anti-affinity, alongside custom schedulers or operators that understand GPU topology and resource sharing. This is critical to ensure that one team's bursty workload doesn't degrade the performance of another's critical inference or training job.
  1. Secure Stateful AI/ML Workloads: The increasing trend of running stateful applications like Weaviate's vector database on Kubernetes introduces new security considerations. Defenders must ensure robust data protection, backup, and disaster recovery strategies for these critical data platforms. This includes secure persistent storage, encryption at rest and in transit, access controls, and regular vulnerability assessments specifically tailored for stateful deployments. The "pets, not cattle" paradigm for these pods means their individual security posture and lifecycle management become more important.
  1. Leverage Platform Engineering for Guardrails and Compliance: Uber's approach of building an opinionated internal platform on top of raw Kubernetes provides a crucial layer of "guardrails." Defenders can collaborate with platform engineering teams to embed security best practices directly into the deployment pipelines and service templates. This can enforce secure configurations, ensure proper secrets management, and mandate secure rollout procedures, reducing the surface area for misconfigurations and vulnerabilities across the AI/ML ecosystem.
  1. Evaluate Third-Party AI/ML Services Carefully: Weaviate's decision to outsource extremely bursty LLM-based transformation agents to Modal highlights a common strategy for handling unpredictable loads. Defenders must conduct thorough due diligence on such third-party AI/ML providers, assessing their security posture, data handling practices, compliance certifications, and potential vendor lock-in risks. Understanding the security implications of data flowing to and from external services is paramount.

By proactively addressing these areas, defenders can help organizations harness the power of Kubernetes for AI/ML while maintaining a strong security posture and operational resilience.

Key Takeaways

  • Kubernetes is rapidly evolving to meet AI/ML demands: New features like Dynamic Resource Allocation (DRRA), in-place pod vertical scaling, and pod-level resources are transforming Kubernetes into a more capable platform for complex, resource-intensive AI/ML workloads, moving beyond its initial focus on stateless microservices.
  • GPU management is a critical economic and technical challenge: GPUs are expensive, and their availability is often fixed, especially at scale. Maximizing utilization through sophisticated scheduling, opportunistic use of idle capacity for disruption-tolerant jobs, and robust checkpointing is paramount to cost-effectiveness. Autoscaling for GPUs in the cloud faces significant limitations due to capacity constraints and long lead times.
  • Platform engineering for AI/ML is diverse and context-dependent: The ideal platform team structure and approach vary significantly based on company size, scale, and specific AI/ML use cases. From Uber's extensive internal platform with guardrails to Weaviate's SaaS-focused team and Overstory's adaptive startup model, flexibility is key.
  • Stateful AI/ML workloads are increasingly viable but require specialized handling: Running applications like vector databases on Kubernetes is becoming more common, but it demands careful attention to resource allocation, vertical scaling capabilities, and robust data management strategies to ensure performance and reliability.
  • Managing burstiness requires strategic approaches, including outsourcing: Highly unpredictable and extremely bursty AI/ML workloads, such as LLM-based transformation agents, can be challenging to self-host efficiently on Kubernetes. Outsourcing to specialized third-party providers may be a pragmatic solution for optimizing cost and performance in such scenarios.
  • The shift to "pods as pets" for AI/ML workloads is gaining momentum: As AI/ML applications become more complex, long-running, and stateful, the Kubernetes community and users are increasingly treating individual pods with more specific attention, requiring advanced features for fine-grained control, dynamic scaling, and multi-node coordination.

About the Speaker(s)

  • Susan Woo: A Product Manager at Google Cloud, Susan focuses on cloud networking, GKE networking, and network security. She moderated the panel discussion, guiding the conversation on extending Kubernetes for AI.
  • Etienne: Co-founder and CTO of Weaviate, an AI data platform that started as a vector database company. He shared insights into running AI workloads, including stateful vector databases and embedding services, on Kubernetes.
  • Lucy: An engineer at Uber, Lucy works on their platforms. Uber is recognized as one of the largest end-user Kubernetes deployments globally, utilizing it for a mix of AI/ML workloads, encompassing both training and inference. She is also an upstream Kubernetes contributor, particularly involved in discussions around Dynamic Resource Allocation (DRRA).
  • Andrea: Works for Overstory, a company that uses satellite images for vegetation management to prevent wildfires. He discussed their use of Google Cloud, GKE, and Google Cloud Run for processing high-resolution imagery and is an open-source contributor to Daxter, a data workflow engine.
  • Tim Wickberg: The Chief Technical Officer for SchedMD, the principal developers of Slurm. He is also involved in the development of Slinky, which provides integrations for bringing Slurm's scheduling capabilities into the Kubernetes stack, particularly for multi-node AI training workloads.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This panel provided a robust and highly informative deep dive into the practical challenges and cutting-edge solutions for running AI/ML workloads on Kubernetes. Featuring highly credible speakers, including upstream Kubernetes contributors and CTOs from companies actively deploying AI at scale, the discussion goes beyond platitudes to cover critical topics like GPU economics, stateful application management, and the implications of new Kubernetes features like DRRA and in-place vertical scaling. Attendees gained actionable insights into platform engineering strategies, resource optimization, and the strategic decision-making required for building resilient AI infrastructure.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon panel offers a candid, grounded view into the significant operational and economic complexities of extending Kubernetes for AI/ML at scale. It effectively highlights the critical shifts required in resource management, platform engineering, and risk posture when moving from stateless microservices to demanding, stateful, and GPU-intensive AI workloads. While a technical discussion, it provides ample material for security and executive leaders to grasp the institutional challenges and necessary strategic adjustments.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025