Keynote: LLM-Aware Load Balancing in Kubernetes: A New Era of Effici... Clayton Coleman, Jiaxin Shan

Clayton Coleman, Jiaxin Shan

KubeCon + CloudNativeCon Europe 2025 · Keynote

Overview

This keynote address by Clayton Coleman from Google and Jiaxin Shan from ByteDance introduces a groundbreaking project within the Kubernetes ecosystem: the Gateway API Inference Extension. Developed under the Kubernetes serving working group, this extension is designed to transform any standard Kubernetes gateway into an intelligent inference gateway, specifically optimized for hosting large language models (LLMs) in production environments. The talk highlights the unique challenges of serving LLMs efficiently at scale and presents a collaborative solution informed by the extensive experiences of both Google and ByteDance.

Watch on YouTube

Visual summary for Keynote: LLM-Aware Load Balancing in Kubernetes: A New Era of Effici... Clayton Coleman, Jiaxin Shan by Clayton Coleman, Jiaxin Shan
Visual summary for Keynote: LLM-Aware Load Balancing in Kubernetes: A New Era of Effici... Clayton Coleman, Jiaxin Shan by Clayton Coleman, Jiaxin Shan

Key moments

  1. 0:00 Introducing Gateway API inference extension for LLM serving
  2. 2:15 Production challenges: request variability, GPU load, traffic skew
  3. 4:00 Hardware heterogeneity in GPU clusters complicates LLM management
  4. 5:00 Three pillars for successful LLM serving: denser, faster, automated
  5. 5:20 Deep dive into LoRA for efficient LLM fine-tuning
  6. 7:00 ByteDance text-to-SQL: LoRA for adapter sharing and efficiency

Keynote: LLM-Aware Load Balancing in Kubernetes: A New Era of Efficiency and Control

Speakers: Clayton Coleman, Principal Software Engineer, Google; Jiaxin Shan, Staff Software Engineer, ByteDance

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=BBqDpqATcI0

Overview

This keynote address by Clayton Coleman from Google and Jiaxin Shan from ByteDance introduces a groundbreaking project within the Kubernetes ecosystem: the Gateway API Inference Extension. Developed under the Kubernetes serving working group, this extension is designed to transform any standard Kubernetes gateway into an intelligent inference gateway, specifically optimized for hosting large language models (LLMs) in production environments. The talk highlights the unique challenges of serving LLMs efficiently at scale and presents a collaborative solution informed by the extensive experiences of both Google and ByteDance.

The core motivation behind this initiative stems from the rapid evolution of the AI landscape. While proprietary models once dominated, open models have swiftly closed the performance gap, creating a dual need for both cutting-edge frontier models and efficient, smaller open models. This necessitates a flexible and composable automation framework for integrating models into applications, much like Kubernetes revolutionized microservice deployment. The speakers argue that traditional Kubernetes load balancing and orchestration are ill-equipped to handle the highly variable and resource-intensive nature of LLM inference, leading to significant inefficiencies and operational toil.

The Gateway API Inference Extension aims to address these critical shortcomings by introducing LLM-aware load balancing. This new approach goes beyond conventional methods, considering factors like request shape, GPU memory utilization, and real-time backend performance to optimize resource allocation and improve latency and throughput. The project promises to unlock a new era of efficiency, control, and automation for platform teams looking to self-host LLMs, ultimately making large models a fundamental and manageable part of application infrastructure.

Background

▶ Watch: Introducing Gateway API inference extension for LLM serving (0:00)

The journey to building an efficient LLM serving platform in Kubernetes is fraught with challenges that traditional microservice architectures rarely encounter. Jiaxin Shan elaborated on several production issues faced at ByteDance, which underscore the unique demands of LLM inference:

  1. Extreme Request Variability: Unlike predictable microservice requests, LLM inference requests are highly dynamic. ByteDance's production LLM services show daily request distributions with significant spikes. More critically, the input prompt size varies widely, and the output length is unpredictable. This means that even under constant queries per second (QPS), GPU compute activity can show sudden, dramatic spikes if a batch of requests with exceptionally long prompts arrives. For LLMs, it's not just the rate of requests but their shape—prompt length, generated token count, and prompt structure—that profoundly impacts GPU load, making it highly unpredictable.
  1. Model Traffic Skew: Large organizations like ByteDance often serve a multitude of models, ranging from critical production systems to experimental models with near-zero traffic. Given the immense resource intensity of LLMs, this traffic skew makes resource allocation and lifecycle management extremely challenging. Relying on static deployments for such a diverse and resource-hungry workload is inefficient, highlighting the need for resource usage-aware orchestration.
  1. Hardware Heterogeneity: Due to factors like machine delivery timelines, quota policies, and availability requirements, inference clusters commonly comprise a mix of different GPU types. In ByteDance's 15,000-GPU clusters, for instance, over eight distinct GPU types are present. While this heterogeneous pool can support diverse workloads, it introduces significant complexity. Models might need to run on different GPU types due to capacity issues, and the varying computational capabilities (TOPS) and memory bandwidth of different GPUs make it difficult to abstract away these hardware differences, complicating management and routing decisions.

These challenges collectively converge on the routing layer, transforming it into a bottleneck for LLM serving. Traditional load balancers, typically designed for uniform requests and predictable service behavior, simply "don't cut it anymore" for LLMs. The speakers emphasized the urgent need for a new class of routing solutions that can understand and adapt to the unique characteristics of LLM inference.

Key Findings

▶ Watch: Hardware heterogeneity in GPU clusters complicates LLM management (4:00)

The collaborative efforts between ByteDance and Google crystallized into a three-pronged approach for successful LLM serving: Denser, Faster, and Automated. These principles aim to provide platform teams with greater control, flexibility, and speed in self-hosting LLMs.

  1. Denser Deployments with LoRA:

The first key finding revolves around maximizing resource utilization through Low-Rank Adaptation (LoRA). LoRA is a technique that efficiently fine-tunes large pre-trained models by only adapting a small set of parameters, known as adapters, instead of modifying the entire model weights. These adapters typically add only about 1% storage and memory overhead compared to the original model, while maintaining comparable efficiency and accuracy.

While LoRA offers significant resource efficiencies and model flexibility, managing it in Kubernetes is surprisingly complex. LoRA adapters must be loaded alongside the base model, meaning they cannot be deployed in separate containers, which challenges Kubernetes' containerization principles. This necessitates a single container potentially serving multiple models, complicating request routing and load balancing, especially when multiple LoRAs contend for shared GPU resources within the same pod.

Despite these challenges, ByteDance successfully implemented LoRA in production. For their text-to-SQL use case, which involved fine-tuning a D6 33B model for various SQL-like query scenarios (e.g., log search, elastic search), they leveraged LoRA. By fine-tuning each scenario as a LoRA adapter and packing these adapters into a single shared base model, they achieved 1.5 to 4.7 times GPU cost savings under different traffic conditions. This "adapter sharing and routing" practice allowed new adapters to be deployed in seconds, with the inference gateway intelligently selecting the least busy instance.

  1. Faster Serving through Latency Awareness:

The second pillar focuses on achieving faster response times, acknowledging that latency is a critical factor for many generative AI applications. Unlike chatbots where output generation speed (sub-human reading speed) is acceptable, many other use cases demand specific latency objectives. To meet these, it's crucial to:

  • Understand traffic distribution in terms of input and output tokens.
  • Choose a foundation model sized for acceptable quality at the lowest compute cost.
  • Select an accelerator configuration (GPU type) that is cost-effective for the model and traffic.
  • Reserve sufficient accelerators for both base and burst loads.

Crucially, sending more requests to an accelerator simultaneously increases throughput but also raises the latency of all other requests. Therefore, latency tolerance, not just traffic load, determines the cost to serve. The speakers highlighted that optimizing for latency across new models, hardware, software, and varying prompt/output lengths often leads to high operational toil.

  1. Automated Load Balancing for LLMs:

The third and most significant finding is the necessity of automating the load balancing process to move beyond the limitations of traditional round-robin approaches. Given the shift from predictable microservice requests to expensive and highly variable LLM queries, a new paradigm is required. The core idea is to:

  • Model the cost of an incoming LLM request: This involves estimating its computational impact and how it will affect other requests currently being processed on a server.
  • Build a real-time snapshot of backend performance: Continuously gather metrics that capture the complex relationships between hardware, traffic concurrency, and client-visible latency.

This automated approach, even in its simplest form, yields substantial benefits. By steering requests to the model server with the most unused GPU memory, the system achieved over 30% higher QPS at constant latency compared to random load balancing, even under predictable traffic loads. This demonstrates a significant step towards maximizing accelerator utilization, especially as more diverse workloads and traffic patterns overlap.

Technical Deep Dive

▶ Watch: Three pillars for successful LLM serving: denser, faster, automated (5:00)

The solution presented by Coleman and Shan centers around the Gateway API Inference Extension, a project designed to bring LLM-aware capabilities to Kubernetes. This extension leverages existing Kubernetes infrastructure while introducing new, specialized components to handle the unique demands of LLM serving.

The fundamental architectural choice for the load balancing component is Envoy. Envoy was selected for its robustness as a standard, dynamic, and extensible load balancer, capable of functioning both within and outside of Kubernetes. Its rich ecosystem, particularly its integration with the Kubernetes Gateway API, ensures that the inference extension can avoid duplicating existing load balancing features that LLM service owners already require. The Gateway API provides a standardized, vendor-agnostic way to configure advanced routing, traffic management, and security features for ingress traffic in Kubernetes. The Inference Extension builds upon this foundation, adding LLM-specific intelligence.

A critical aspect of the technical design is the decoupling of the load balancing algorithms from Envoy itself. This is achieved using the standard Envoy xDS gRPC callout mechanism. This mechanism allows external services to dynamically influence Envoy's routing decisions. By deploying the LLM-aware scheduling algorithms independently of Envoy, the architecture gains several advantages:

  1. Flexibility: It allows for multiple implementations of the scheduling algorithms, fostering experimentation and enabling large platform teams to develop custom solutions if the open-source project doesn't meet specific, unique requirements.
  2. Independent Deployment: Algorithms can be updated and deployed without modifying or restarting the Envoy proxy, reducing operational overhead and increasing agility.
  3. Scalability: The computational complexity of LLM-aware scheduling can be offloaded to dedicated services, ensuring the data plane (Envoy) remains lightweight and performant.

To facilitate intelligent decision-making by the scheduler, the project also focuses on standardizing the metrics emitted by model servers. Consistent metrics across the ecosystem mean that operators have a uniform operational view, and the external scheduler can make more informed decisions with less effort. These metrics likely include real-time GPU memory usage, current request queue lengths, estimated time-to-completion for ongoing requests, and perhaps token generation rates.

The LLM-aware load balancer operates on two core principles:

  1. Cost Modeling of Incoming Requests: The system estimates the "cost" of an incoming LLM request. This cost is not just about raw CPU/GPU cycles but also considers factors like estimated prompt length, expected output length, and how adding this request will impact the latency of other requests already being processed on a given server.
  2. Real-time Backend Performance Snapshot: The system continuously gathers metrics from each backend model server to build a dynamic, real-time view of its performance. This includes current GPU memory utilization, load, and observed client-visible latency. This continuous feedback loop allows the scheduler to make highly adaptive routing decisions.

Beyond core load balancing, the Gateway API Inference Extension also integrates features specifically for LLM serving:

  • LoRA Support: It provides built-in support for managing LoRA adapters, enabling prioritization and fairness when multiple LoRAs share a base model. This allows for safe sharing of model servers between many different workloads, leading to higher utilization rates.
  • Standard Model Rollouts: The extension supports standard model rollout strategies, ensuring that new models or LoRA adapters can be deployed and updated safely in production.

This comprehensive approach aims to automate the "boring parts" of LLM production serving, bringing the latest ML research into practical, operationalized infrastructure. The extension is designed to be agnostic to the specific model server deployment method, plugging into a broad set of gateway solutions and allowing users to choose their preferred model serving frameworks (e.g., vLLM, TensorRT-LLM).

Demo / Proof of Concept

▶ Watch: Deep dive into LoRA for efficient LLM fine-tuning (5:20)

While a live, interactive demo was not explicitly presented as a separate segment, the speakers provided concrete, data-backed examples and results from ByteDance's production environment and Google's customer feedback, serving as powerful proofs of concept for the proposed solutions.

One primary demonstration of denser deployment was ByteDance's experience with LoRA adapters for their text-to-SQL use case. They fine-tuned the D6 33B model to support various SQL-like query scenarios, such as log search and elastic search, each requiring specific adaptations. Instead of deploying dedicated GPUs for each scenario, they adopted the LoRA solution. By fine-tuning each scenario as a LoRA adapter and consolidating these adapters into a single shared base model, they achieved remarkable resource efficiency. This approach, which they termed "adapter sharing and routing," allowed them to deploy new adapters in seconds and, more significantly, resulted in 1.5 to 4.7 times GPU cost savings under varying traffic conditions. In this setup, the inference gateway played a crucial role in minimizing the number of model servers and intelligently routing requests to the least busy instances supporting the required adapter.

For the automated load balancing aspect, a clear example of its effectiveness was provided. The speakers illustrated how traditional random load balancing of non-uniform LLM requests can lead to inefficiencies, with some accelerators sitting idle while others become overloaded with long requests. Since generating output tokens requires GPU memory, an overloaded server can fill its memory, leading to it stopping acceptance of new requests and increasing overall latency.

In contrast, a simple version of their algorithmic, LLM-aware load balancing loop was tested. This approach involved steering requests to the model server with the most unused GPU memory. This single optimization, applied to a representative chat agent workload, demonstrated a significant improvement: it achieved over 30% higher QPS (queries per second) at constant latency compared to a random load balancing strategy. This finding is crucial, as it shows that even a basic awareness of backend resource state (like GPU memory) can dramatically increase accelerator utilization and overall system performance. The benefit of such algorithmic approaches is expected to amplify as more diverse workloads are added to shared model servers and traffic patterns become more complex.

These examples clearly illustrate the practical benefits and performance gains achievable through the Gateway API Inference Extension's principles of denser, faster, and automated LLM serving.

Defensive Implications

▶ Watch: ByteDance text-to-SQL: LoRA for adapter sharing and efficiency (7:00)

For platform teams and developers tasked with deploying and managing LLMs in production, the insights and solutions presented in this talk carry significant defensive implications:

  1. Adopt LLM-Aware Orchestration: The most critical takeaway is to recognize that traditional load balancing (e.g., round-robin) and static deployment strategies are insufficient for LLM workloads. Platform teams must actively move towards resource usage-aware orchestration and LLM-aware load balancing solutions. The Gateway API Inference Extension offers a standardized, Kubernetes-native path to achieve this.
  1. Embrace LoRA for Efficiency: For organizations fine-tuning multiple specialized LLMs, Low-Rank Adaptation (LoRA) is a powerful technique for achieving denser deployments and significant GPU cost savings (demonstrated at 1.5-4.7x). Defenders should investigate integrating LoRA into their model development and deployment workflows to maximize GPU utilization and reduce infrastructure costs, especially for models with similar base architectures.
  1. Prioritize Real-time Metrics: Intelligent load balancing for LLMs hinges on accurate, real-time insights into backend server state. Platform teams should focus on collecting and acting upon metrics like GPU memory utilization, current request queue depth, estimated token generation rates, and client-visible latency. Standardizing these metrics across the model serving ecosystem will be crucial for building robust, adaptive routing mechanisms.
  1. Leverage Kubernetes Gateway API: By building on the Kubernetes Gateway API, the Inference Extension provides a standardized control plane. Defenders should familiarize themselves with the Gateway API and consider it the foundation for their ingress and traffic management for LLM services, benefiting from its extensibility and rich feature set.
  1. Decouple and Experiment with Algorithms: The architecture's use of Envoy xDS gRPC callouts allows for independent deployment and experimentation with load balancing algorithms. This flexibility enables platform teams to develop and iterate on custom scheduling logic tailored to their specific workloads and latency requirements without disrupting the core data plane, fostering innovation in optimization.
  1. Plan for Hardware Heterogeneity: Given the reality of mixed GPU types in large clusters, defensive strategies should account for hardware heterogeneity. Intelligent routing must be capable of understanding the diverse capabilities of different accelerators and matching workloads appropriately, potentially requiring adaptive SLO-driven routing or heterogeneous weighted routing as future features become available.
  1. Consider the Model Trade-off: The talk underscored the meaningful trade-off between very smart frontier models and smaller, more efficient open models. Defenders should strategically choose foundation models based on latency objectives, quality requirements, and compute cost, leveraging the flexibility of the Gateway API Inference Extension to serve both types optimally within the same infrastructure.

Key Takeaways

  • LLM serving presents unique challenges: Unlike traditional microservices, LLMs exhibit extreme request variability (prompt/output length), unpredictable GPU load, model traffic skew, and hardware heterogeneity, making conventional load balancing inefficient.
  • Traditional load balancing is insufficient: Round-robin and static deployments lead to wasted resources and increased latency for LLM workloads due to their dynamic and resource-intensive nature.
  • The Gateway API Inference Extension provides an LLM-aware solution: This new project within the Kubernetes ecosystem turns any Kubernetes gateway into an intelligent inference gateway, addressing the specific demands of LLM serving.
  • LoRA enables denser, cost-efficient deployments: Low-Rank Adaptation (LoRA) allows for efficient fine-tuning and shared base models, leading to significant GPU cost savings (1.5-4.7x) and faster deployment of new model variants.
  • Intelligent, real-time load balancing is crucial: By modeling request cost and using real-time backend performance data (e.g., GPU memory utilization), LLM-aware load balancing can achieve over 30% higher QPS at constant latency compared to random approaches.
  • Envoy and xDS callouts offer a flexible architecture: The solution leverages Envoy as a standard load balancer, with its xDS gRPC callout mechanism enabling decoupled, independently deployable, and extensible load balancing algorithms.

About the Speaker(s)

Clayton Coleman is a Principal Software Engineer at Google. He is deeply involved in the Kubernetes ecosystem, having been a significant contributor and leader in its development. In this talk, he represents Google's extensive experience in serving large-scale applications and generative AI, contributing to the strategic direction and architectural design of the Gateway API Inference Extension project. He is part of the Kubernetes serving working group, which was established a year prior to this talk to address the evolving needs of model serving.

Jiaxin Shan is a Staff Software Engineer at ByteDance. He brings invaluable real-world production experience from ByteDance, a company operating at immense scale with 15,000 GPU clusters and a diverse range of LLM workloads. His insights into the practical challenges of LLM serving, such as request variability, model traffic skew, and hardware heterogeneity, directly informed the design and priorities of the Gateway API Inference Extension. He actively collaborated with Google to build this next-generation platform in the open, contributing to the "denser, faster, and automated" principles.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This keynote by Coleman and Shan introduces the Gateway API Inference Extension, a critical and timely solution for the unique challenges of serving large language models in Kubernetes. Moving beyond the limitations of traditional load balancing, the project leverages real-time GPU metrics and techniques like LoRA to enable denser, faster, and more automated LLM deployments. The talk provides concrete data from ByteDance's production environment, demonstrating significant cost savings and performance gains, and lays out a clear, actionable path for platform teams grappling with LLM infrastructure.

Heather Calloway (CISO) — STRONG ACCEPT

This keynote by Coleman and Shan presents a critical evolution in Kubernetes for serving large language models, addressing the significant inefficiencies of traditional orchestration methods. Their Gateway API Inference Extension, driven by real-world challenges at Google and ByteDance, promises substantial GPU cost savings and performance gains through LLM-aware load balancing and LoRA. While highly technical, its implications for an organization's ability to efficiently and reliably deploy AI are profound, directly impacting strategic investment, operational resilience, and the overall security posture of AI initiatives.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025