Advancements in AI/ML Inference Workloads on Kubernetes From... Yuan Tang & Eduardo Arango Gutierrez
Yuan Tang, Eduardo Arango Gutierrez
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This KubeCon EU talk provided a comprehensive update from the Kubernetes Working Group Serving (WG Serving), a crucial initiative dedicated to enhancing Kubernetes for the unique and demanding requirements of AI/ML inference workloads, particularly Large Language Models (LLMs). Presented by co-chairs Yuan Tang and Eduardo Arango Gutierrez, the session highlighted the progress made by the working group since its inception a year prior at KubeCon Europe in Paris. The core motivation behind WG Serving is to address the inherent gaps in Kubernetes, which, while excellent for general container orchestration, was not originally designed to optimize resource allocation, scalability, and performance for the specialized needs of AI/ML inference.

Key moments
- 1:10 What is Working Group Serving? Its origin.
- 2:00 Three main goals of Working Group Serving explained.
- 4:00 Working group leadership, community size, and past talks.
- 5:00 Introduction to the four main workstreams and annual report.
- 6:20 Autoscaling workstream: benchmarking and large OCI image challenges.
- 7:50 Multi-host/node workstream: distributed workloads and MPI.
Advancements in AI/ML Inference Workloads on Kubernetes From... Yuan Tang & Eduardo Arango Gutierrez
Speakers: Yuan Tang, Senior Principal Software Engineer at Red Hat; Eduardo Arango Gutierrez, Working Group Chair for WG Serving
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=G8U141NkrDI
Overview
This KubeCon EU talk provided a comprehensive update from the Kubernetes Working Group Serving (WG Serving), a crucial initiative dedicated to enhancing Kubernetes for the unique and demanding requirements of AI/ML inference workloads, particularly Large Language Models (LLMs). Presented by co-chairs Yuan Tang and Eduardo Arango Gutierrez, the session highlighted the progress made by the working group since its inception a year prior at KubeCon Europe in Paris. The core motivation behind WG Serving is to address the inherent gaps in Kubernetes, which, while excellent for general container orchestration, was not originally designed to optimize resource allocation, scalability, and performance for the specialized needs of AI/ML inference.
The speakers detailed the working group's multifaceted approach, encompassing API enhancements, research into orchestration and scalability, and optimization of low-level resource sharing. They introduced key sub-projects and initiatives, such as inference perf for standardized benchmarking, the Gateway API Inference Extension (GIE) for efficient multi-tenancy with Laura adapters, and significant contributions to the Dynamic Resource Allocation (DRA) feature within Kubernetes. The talk underscored the collaborative nature of the working group, involving diverse companies and hundreds of contributors, all working towards a more robust and efficient Kubernetes platform for the burgeoning field of AI/ML inference.
The importance of this work cannot be overstated. As AI/ML models, especially LLMs, grow in size and complexity, deploying and scaling them efficiently on cloud-native infrastructure becomes a critical challenge. These models often require specialized hardware like GPUs, demand precise resource allocation, and exhibit unique performance characteristics that traditional Kubernetes metrics struggle to capture. WG Serving's efforts directly tackle these pain points, aiming to standardize practices, develop new tools, and influence core Kubernetes features to make it the go-to platform for production-grade AI/ML inference, ultimately accelerating innovation and adoption in the AI space.
Background
▶ Watch: What is Working Group Serving? Its origin. (1:10)
The Kubernetes Working Group Serving (WG Serving) emerged from a recognized need within the Kubernetes community to specifically address the challenges of running AI/ML inference workloads, particularly Large Language Models (LLMs), on Kubernetes. Its genesis can be traced back to KubeCon Europe in Paris last year, where during an unconference session, a consensus formed around the significant gaps in Kubernetes' native capabilities for serving these highly specialized and resource-intensive applications. The group was officially born alongside the Working Group Device Management, marking its one-year anniversary at the time of this KubeCon EU presentation.
Prior to WG Serving, the community struggled with fragmented efforts and a lack of standardized approaches for managing AI/ML inference. A key challenge highlighted by the speakers was the difficulty in even agreeing upon fundamental monitoring metrics necessary for effectively autoscaling clusters running LLMs. This lack of common understanding demonstrated a clear need for a dedicated body to research, benchmark, and standardize practices. The scale of modern AI models, often manifesting as OCI images hundreds of gigabytes in size, further exacerbated issues related to image pulling, data management, and efficient resource utilization.
The working group established three primary goals to tackle these issues:
- Enhance Kubernetes Controllers: This involves proposing new APIs and changes to core Kubernetes controllers and low-level components, such as Dynamic Resource Allocation (DRA), to better accommodate serving workloads.
- Investigate and Research Orchestration and Scalability: A significant part of this goal is the development of a benchmark tool, inference perf, to deeply understand the performance characteristics of LLMs and identify optimal autoscaling strategies.
- Optimize Resource Sharing: This focuses on enabling more efficient sharing of resources, particularly GPUs, between multiple pods or even across multiple hosts for distributed workloads like MPI communication, often involving low-level container runtime bits.
Comprising four co-chairs from leading companies like Google Cloud, Red Hat, Nvidia, and ByteDance, and boasting a Slack channel with over 330 members, WG Serving operates through four main workstreams to manage its broad scope and active community: Autoscaling, Multi-host and Multi-node, Dynamic Resource Allocation, and Orchestration. These workstreams serve as focused avenues for addressing specific technical challenges and driving initiatives forward, reflecting the collective effort to mature Kubernetes as a platform for AI/ML inference.
Key Findings
▶ Watch: Working group leadership, community size, and past talks. (4:00)
The Kubernetes Working Group Serving has made substantial progress in its first year, translating its ambitious goals into tangible initiatives and early releases. A central finding is the critical need for standardized benchmarking to effectively manage and scale LLM inference. This led to the creation of the inference perf project, a collaborative effort aiming to provide a vendor-neutral tool for understanding the nuanced performance impacts of LLMs, from prompt length to GPU utilization, which traditional metrics often miss.
Another significant contribution is the Gateway API Inference Extension (GIE), which has already seen a v0.1 release. GIE is designed to improve resource sharing for Laura adapters on shared foundation models, directly addressing challenges in multi-tenancy and efficient utilization of expensive GPU resources. Early results indicate improvements in tail latency and throughput for LLM completion requests, demonstrating its immediate value to the community.
The working group has also identified and actively engaged with fundamental Kubernetes features to better support AI/ML workloads:
- Dynamic Resource Allocation (DRA): WG Serving is heavily influencing the development of DRA, a beta Kubernetes feature, by pushing for specific enhancements required by serving workloads. Their goal is to ensure DRA, anticipated to reach GA in Kubernetes 1.34 (December this year), robustly accommodates GPU sharing and intelligent device failure handling.
- Large OCI Images and Data Movement: The sheer size of modern models (hundreds of gigabytes) packaged as OCI images presents a major hurdle. The workstream is actively discussing and contributing to the image volume source KEP, aiming to standardize how containers can be used primarily as data movers rather than full running systems, streamlining model deployment.
- Multi-host/Multi-node Serving: Through collaboration and feature requests, WG Serving has contributed to projects like Leader Working Set (versions 0.3/0.5), partnered with the BLM community for testing, and proposed new CRDs for KServe to better support customizable multi-node serving patterns for distributed inference workloads.
- Serving Catalog: This project provides practical, reference implementations for various model servers (e.g., vLLM, JetStream), deployment patterns (single/multi-host), and orchestration frameworks (Kubernetes Deployment, Leader Worker Set), serving as a valuable resource for users exploring different configurations and hardware accelerators.
These findings collectively underscore the working group's commitment to building a more capable and efficient Kubernetes ecosystem for AI/ML inference, moving beyond generic container orchestration to address the specific, high-performance demands of modern AI.
Technical Deep Dive
▶ Watch: Introduction to the four main workstreams and annual report. (5:00)
The Kubernetes Working Group Serving organizes its technical efforts across four dedicated workstreams, each tackling specific aspects of AI/ML inference on Kubernetes.
Autoscaling Workstream
This workstream is at the forefront of defining how to effectively autoscaling LLM inference workloads. A primary initiative is the inference perf project, a standardized benchmarking tool designed to identify the precise metrics that truly impact LLM performance. The speakers highlighted that traditional metrics often fall short, and factors like prompt length can significantly influence a model's behavior and resource consumption. The goal is to move beyond generic CPU/memory utilization to model-specific insights, creating categories and benchmarks to inform more intelligent autoscaling decisions.
Another critical area is the management of large OCI images, with some models reaching "hundreds of gigabytes." This poses significant challenges for image pulling, storage, and deployment speed. The workstream is actively engaged in discussions around the image volume source KEP, which proposes treating containers primarily as mechanisms for moving data rather than self-contained running systems. This paradigm shift aims to optimize the delivery of large models, effectively addressing the "volume containers" concept from years past.
Multi-host and Multi-node Workstream
This workstream focuses on enabling robust distributed workloads and multi-node serving. WG Serving actively participates in and contributes feature requests to existing projects:
- Leader Working Set: The group contributed to versions 0.3 and 0.5 of this project, ensuring it better accommodates the needs of serving workloads requiring leader-follower patterns across multiple nodes.
- BLM (Batch and Low-Latency Machine Learning): Collaboration with the BLM community for testing helps validate and improve distributed inference scenarios.
- KServe: As one of the project leads for KServe, Yuan Tang highlighted efforts to propose new CRDs (Custom Resource Definitions) within KServe to enhance its support for customizable multi-node serving behaviors, moving towards more disaggregated orchestration. This allows users to tailor how models are served across clusters, particularly for complex distributed inference patterns.
Dynamic Resource Allocation (DRA) Workstream
This is a high-priority workstream, as DRA represents a fundamental shift in how Kubernetes manages specialized hardware like GPUs. Eduardo Arango Gutierrez emphasized the critical importance of influencing DRA while it is still a beta feature (currently in Kubernetes), before it reaches GA (General Availability), which is targeted for Kubernetes 1.34 in December this year. The working group maintains weekly communication with the WG Device Management to push for features specifically needed by serving workloads.
Key areas of focus include:
- Fine-grained GPU Sharing: Ensuring DRA can facilitate efficient sharing of GPUs between multiple pods, a common requirement for optimizing resource utilization in inference.
- Device Failure Handling and Resilient Workload Management: A significant gap in current Kubernetes is the poor communication from hardware about device health. This workstream is pushing for new features and APIs to allow Kubernetes to better report when a GPU is unhealthy or misbehaving. This enables the scheduler to reroute workloads away from failing hardware, preventing prolonged hangs and improving overall workload resilience.
Orchestration Workstream
This stream drives initiatives focused on improving the orchestration layer for inference. The most prominent project here is the Gateway API Inference Extension (GIE), released at v0.1. GIE is designed to enhance resource sharing across multiple use cases on a shared foundation model, particularly relevant for scenarios involving Laura adapters.
GIE provides:
- Improved Performance: It aims to reduce tail latency and increase throughput for LLM completion requests on Kubernetes-hosted model servers.
- Custom Scheduling Algorithm: It incorporates an extensible custom scheduling algorithm to optimize resource allocation.
- Declarative APIs: GIE offers a set of declarative APIs to route client model names to specific Laura adapter use cases, enabling more efficient sharing of underlying model resources.
- End-to-End Observability: Built-in observability with custom-defined service objective attainment allows platform teams to monitor and manage performance effectively.
- Operational Guards: It ensures safe multi-tenancy by providing operational isolation between different client model names running on a shared foundation model.
Sub-Projects in Detail
Beyond the workstreams, several key sub-projects underpin these efforts:
Inference Perf Project
This is a collaborative benchmarking tool (a Python library) developed with contributions from IBM, Red Hat, Google Cloud, and Nvidia. It aims to be a standardized, extensible library for various benchmarking use cases, including autoscaling and Laura use cases with GIE.
- Current Status: Provides a Python library, supports vLLM model server, handles multiple distributions with specified QPS (Queries Per Second), and uses the CR GPT dataset to simulate real-world conversational workloads. It also generates reports with valuable metrics.
- Roadmap: Future plans include adding support for other model servers like Triton and TGI, integrating with orchestration projects such as vLLM production stack and AI Bricks, and supporting multi-model use cases and diverse traffic distributions.
Gateway API Inference Extension (GIE)
As detailed above, GIE (repository link provided in talk) is specifically tailored to optimize resource sharing for Laura adapters on shared foundation models. Its architecture is designed for integration with many ecosystem projects, with blue boxes in the diagram representing already implemented components. The roadmap includes further integration with external components.
Serving Catalog Project
This project acts as a central repository for working examples and recommended configurations for various inference deployments. It aims to help users explore different patterns and best practices.
- Content: Includes examples for different model servers (e.g., vLLM, JetStream), models, and deployment patterns (single-host, multi-host inference).
- Orchestration Frameworks: Examples cover basic Kubernetes Deployments, Leader Worker Set, and ongoing work with KServe inference services.
- Hardware and Cloud Specificity: Implementations are designed to be cloud provider-agnostic but include configurations for different hardware accelerators and cloud environments.
- Status: Single-host inference examples for vLLM and JetStream are available, as are multi-host inference examples using Leader Worker Set for vLLM. It also includes components for HPA (Horizontal Pod Autoscaler) for token latency, addressing a specific metric crucial for LLM performance.
These initiatives represent a concerted effort to build out the technical capabilities within Kubernetes to meet the escalating demands of the AI/ML inference landscape.
Demo / Proof of Concept
▶ Watch: Autoscaling workstream: benchmarking and large OCI image challenges. (6:20)
This particular KubeCon talk served primarily as a working group update, focusing on the progress, initiatives, and technical roadmap of the Kubernetes WG Serving. As such, it did not include a live demonstration or a detailed proof of concept of any specific tool or feature.
While the speakers discussed projects like inference perf (a benchmarking tool) and the Serving Catalog (providing working examples), they presented their architecture diagrams, features, and current status rather than running a live demo. The Gateway API Inference Extension (GIE) was mentioned as having a v0.1 release and showing improvements, but no live demonstration of its capabilities or performance gains was part of this presentation. The talk's format was an informational update, encouraging community engagement and feedback on ongoing developments rather than showcasing a specific product or solution.
Defensive Implications
▶ Watch: Multi-host/node workstream: distributed workloads and MPI. (7:50)
While the talk primarily focuses on optimizing performance and resource management for AI/ML inference, the advancements made by the Kubernetes WG Serving inherently carry significant defensive implications for organizations deploying and managing these workloads. These implications revolve around building more resilient, secure, and cost-effective inference platforms.
- Enhanced Resource Resilience and Failure Handling: The focus on Dynamic Resource Allocation (DRA) and particularly on device failure handling is crucial. By enabling Kubernetes to better detect and react to unhealthy GPUs or other accelerators, organizations can prevent workloads from being stuck on failing hardware. This reduces downtime, improves service availability, and ensures that inference services remain responsive, directly impacting the reliability of AI-powered applications. Proactive rerouting of workloads away from misbehaving devices minimizes the blast radius of hardware failures, a key defensive posture.
- Secure Multi-tenancy and Resource Isolation: The Gateway API Inference Extension (GIE), with its emphasis on resource sharing for Laura adapters on shared foundation models and its "operational guards," offers a more secure approach to multi-tenancy. By providing declarative APIs and custom scheduling, GIE allows platform teams to safely serve different client model names or use cases on a common, expensive foundation model pool. This isolation helps prevent resource contention from impacting critical workloads and can aid in maintaining Service Level Objectives (SLOs), which is a defensive measure against performance degradation.
- Optimized Cost and Resource Management: The inference perf project, by providing a standardized benchmarking tool, empowers platform engineers to deeply understand the true resource consumption and performance characteristics of LLMs. This granular insight is critical for accurate autoscaling and capacity planning, preventing both over-provisioning (which wastes resources and incurs unnecessary costs) and under-provisioning (which leads to performance bottlenecks and service unavailability). Efficient resource allocation is a defensive strategy against unexpected operational costs and performance-related incidents.
- Addressing Large Model Storage and Delivery Risks: The discussion around the challenges of large OCI images (hundreds of gigabytes) and the image volume source KEP has defensive implications for supply chain security and operational efficiency. By standardizing how these massive models are moved and managed, organizations can potentially streamline scanning processes, reduce the attack surface associated with transient image layers, and ensure faster, more reliable deployments. Minimizing the time and complexity of pulling large images also reduces the window of vulnerability during deployment.
- Standardization and Best Practices: The creation of the Serving Catalog provides a collection of "recommended configurations" and "working examples" for various model servers and deployment patterns. This standardization acts as a defensive baseline, guiding users towards well-tested and potentially more secure deployment architectures. By adopting these patterns, organizations can avoid common misconfigurations and leverage community-vetted approaches, which implicitly improves the security posture by reducing custom, error-prone setups.
In essence, the WG Serving's work, by making AI/ML inference on Kubernetes more efficient, reliable, and manageable, directly contributes to a stronger defensive posture against operational failures, performance degradation, and resource inefficiencies that can plague complex AI deployments.
Key Takeaways
- Kubernetes WG Serving is Essential for AI/ML Inference: The working group effectively addresses critical gaps in Kubernetes for running Large Language Model (LLM) inference, focusing on enhancing core controllers, improving orchestration, and optimizing resource sharing.
- Standardized Benchmarking is Crucial for LLM Autoscaling: The inference perf project provides a vital, vendor-neutral tool for understanding unique LLM performance characteristics (e.g., prompt length impact) to enable more accurate and efficient autoscaling decisions.
- Gateway API Inference Extension (GIE) Boosts Multi-tenancy: The v0.1 release of GIE significantly improves resource sharing for Laura adapters on shared foundation models, offering better tail latency and throughput with built-in operational guards for safe multi-client deployments.
- Dynamic Resource Allocation (DRA) is a Key Focus for Hardware Management: WG Serving is actively pushing for features within the Dynamic Resource Allocation (DRA) beta feature, aiming for its GA in Kubernetes 1.34, to enable robust GPU sharing and critical device failure handling for resilient inference workloads.
- Addressing Large OCI Images is Paramount: The challenge of models packaged as OCI images (hundreds of gigabytes) is being tackled through discussions and contributions to the image volume source KEP, aiming to optimize data movement for these massive artifacts.
- Serving Catalog Provides Valuable Reference Implementations: The Serving Catalog offers practical, recommended configurations and examples for deploying various model servers (e.g., vLLM, JetStream) across different deployment patterns and orchestration frameworks, including HPA for token latency.
About the Speaker(s)
Yuan Tang is a Senior Principal Software Engineer at Red Hat, bringing extensive expertise to the cloud-native AI space. He serves as one of the co-chairs for the Kubernetes Working Group Serving, demonstrating his leadership in shaping the future of AI/ML inference on Kubernetes. Additionally, Yuan is a project lead for both Argo and KServe, and actively maintains several other projects, including the Llama Stack project. His contributions extend to authorship, having penned several books, reflecting his deep knowledge and commitment to the community.
Eduardo Arango Gutierrez is also a co-chair for the Kubernetes Working Group Serving, playing a pivotal role in the group's strategic direction. With a long-standing background in distributed systems and containers, Eduardo specializes in low-level container runtime bits. His current work heavily involves projects related to CDI (Container Device Interface) and DRA (Dynamic Resource Allocation), both of which are critical for optimizing hardware resource management for AI/ML workloads within Kubernetes. His expertise ensures that the working group's efforts are grounded in robust, low-level technical considerations.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This KubeCon update from the Kubernetes WG Serving is critical. It's not just a status report; it's a deep dive into the foundational work being done to make Kubernetes a viable, performant, and resilient platform for Large Language Model inference. The speakers, as co-chairs, provide rare insider signal on the progress of crucial features like Dynamic Resource Allocation (DRA) and introduce novel solutions like the Gateway API Inference Extension (GIE) and inference perf. This isn't hype; it's the real engineering that will define the next generation of AI infrastructure.
Heather Calloway (CISO) — STRONG ACCEPT
This KubeCon talk by the Kubernetes WG Serving details critical foundational work to make Kubernetes viable for large-scale AI/ML inference, particularly LLMs. While highly technical, the advancements in resource allocation, failure handling, and multi-tenancy directly underpin future AI governance, resilience, and cost management. It's a clear demonstration of institutional realism addressing fundamental platform gaps, and for any CISO overseeing significant AI initiatives, understanding this work is essential for securing and scaling their organization's AI future.