Optimizing Metrics Collection & Serving When Autoscaling LLM Workloads - Vincent Hou & Jiří Kremser
Vincent Hou, Jiří Kremser
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk, delivered by Vincent Hou from Bloomberg and Jiří Kremser from Kify, addresses the critical challenge of efficiently autoscaling Large Language Model (LLM) workloads within Kubernetes environments. As LLMs introduce a paradigm shift in resource consumption, predominantly relying on GPUs rather than traditional CPUs, existing autoscaling mechanisms often fall short. The speakers delve into the limitations of conventional metrics like CPU utilization, memory, or requests per second (RPS) for effectively managing GPU-intensive LLM inference.

Key moments
- 0:00 Talk introduction and speakers' background
- 2:55 General concept of autoscaling in cloud and Kubernetes
- 4:00 Kubernetes autoscaling architecture: metrics, server, HPA explained
- 5:45 Paradigm shift: LLM workloads and unsuitable existing metrics
- 6:55 Five key challenges for autoscaling LLM workloads
- 8:10 Analyzing limitations of existing autoscaling solutions (KPA)
Optimizing Metrics Collection & Serving When Autoscaling LLM Workloads
Speakers: Vincent Hou, Senior Software Engineer, Bloomberg; Jiří Kremser, Kify
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=lefjb4Vnd8k
Overview
This talk, delivered by Vincent Hou from Bloomberg and Jiří Kremser from Kify, addresses the critical challenge of efficiently autoscaling Large Language Model (LLM) workloads within Kubernetes environments. As LLMs introduce a paradigm shift in resource consumption, predominantly relying on GPUs rather than traditional CPUs, existing autoscaling mechanisms often fall short. The speakers delve into the limitations of conventional metrics like CPU utilization, memory, or requests per second (RPS) for effectively managing GPU-intensive LLM inference.
The core of the presentation focuses on a novel, open-source approach that combines OpenTelemetry for robust metric collection, KEDA (Kubernetes Event-Driven Autoscaling) for flexible scaling decisions, and a custom OpenTelemetry Add-on for KEDA to bridge these technologies. This integrated solution enables dynamic adjustment of LLM inference pods and even the underlying GPU-enabled Kubernetes nodes based on specialized, real-time metrics reflective of actual GPU usage, such as KV cache percentage and waiting queue length. The talk culminates in a live demonstration showcasing how this architecture facilitates rapid, cost-efficient, and intelligent autoscaling for LLM deployments, addressing both pod and infrastructure-level resource management.
Background
▶ Watch: Talk introduction and speakers' background (0:00)
Kubernetes offers powerful native autoscaling capabilities, primarily through the Horizontal Pod Autoscaler (HPA) and Vertical Pod Autoscaler (VPA). HPA dynamically adjusts the number of pod replicas based on observed metrics, while VPA adjusts resource requests and limits for pods. Traditionally, HPA scales based on basic resource metrics like CPU and memory utilization, or custom metrics exposed by applications. However, the advent of Large Language Models (LLMs) has introduced significant challenges to these established autoscaling paradigms.
LLMs are fundamentally different from previous workloads. They are heavily reliant on GPUs, often requiring substantial memory (models can exceed 100 gigabytes) and specialized processing. This shift renders conventional CPU and memory metrics largely irrelevant for effective autoscaling, as high CPU usage doesn't necessarily correlate with GPU load, and memory usage might reflect loaded models rather than active inference. Even metrics like requests per second (RPS), commonly used in serverless platforms like Knative Serving, prove inadequate. An LLM might receive 100 concurrent requests, but if these requests are short and simple, the GPU consumption might remain low. Conversely, a single complex request requiring extensive token generation could heavily tax the GPU. The critical metrics for LLM autoscaling revolve around latency and throughput, which are more directly tied to GPU utilization.
Prior attempts and existing solutions highlighted in the talk reveal these limitations:
- Kubernetes Metric Server: While an integral native component, it only supports CPU and memory metrics, making it unsuitable for GPU-centric LLM workloads.
- Knative Serving's KPA (Knative Pod Autoscaler): KPA, a dependency for KServe (Bloomberg's cloud-agnostic model inference platform), scales based on RPS. While it offers fast scaling, push-based metrics, and scale-to-zero capabilities, it has several drawbacks for LLMs: RPS is not a relevant metric, its complex scaling algorithm is hard to maintain, and extending it for new metrics requires plugin development, limiting portability.
- AI Bricks: Positioned as a cloud-native solution optimized for LLM inference, AI Bricks supports both metric-based and optimizer-based autoscaling, integrating LLM-specific metrics. However, for organizations like Bloomberg already invested in KServe as their backbone, AI Bricks' lack of integration capabilities made it an unviable option.
These limitations underscore the need for a flexible, extensible, and portable autoscaling solution that can leverage custom, GPU-specific metrics for LLM workloads, integrate with existing platforms like KServe, and address the unique demands of scaling both application pods and underlying GPU infrastructure.
Key Findings
▶ Watch: Kubernetes autoscaling architecture: metrics, server, HPA explained (4:00)
The talk presents a compelling solution that addresses the shortcomings of traditional autoscaling for LLM workloads by leveraging a powerful combination of open-source tools. The key findings and contributions are:
- Specialized Metrics are Essential: Traditional metrics like CPU, memory, and Requests Per Second (RPS) are ineffective for autoscaling GPU-bound LLM workloads. The critical metrics for LLMs are internal, reflecting actual GPU usage and inference pipeline pressure. Specifically, the talk highlights the KV cache percentage (indicating how much GPU memory is used for key-value caching during inference) and the waiting queue length for requests as highly relevant indicators for scaling decisions. These metrics offer a direct correlation to GPU load and potential latency.
- OpenTelemetry as the Universal Metric Collection Standard: OpenTelemetry (OTel) emerges as a crucial "giant" in the solution. Its standardization efforts for traces, metrics, and logs, combined with the flexible OTel Collector, provide a robust and vendor-neutral way to collect custom metrics directly from LLM workloads. This allows for metric filtering, processing (e.g., scaling values), and exporting in a push-based manner, which is more reactive than pulling.
- KEDA for Event-Driven, Flexible Autoscaling: KEDA (Kubernetes Event-Driven Autoscaling) is identified as another "giant." KEDA extends HPA capabilities by allowing scaling based on a wide array of external metrics, supporting scale-to-zero, and providing a declarative
ScaledObjectCustom Resource Definition (CRD). This flexibility is paramount for LLMs, as it allows scaling based on the custom OTel metrics rather than just CPU/memory.
- The OpenTelemetry Add-on for KEDA Bridges the Gap: A significant contribution presented is the OTel Add-on for KEDA, developed by Jiří Kremser. This add-on acts as a crucial link, receiving metrics from OpenTelemetry Collectors via OTLP (OpenTelemetry Protocol) and translating them into scaling advisories for KEDA. It includes a lightweight, short-term memory store, specifically designed for scaling metrics, which avoids overwhelming a full-blown Prometheus setup typically used for monitoring and alerting.
- Comprehensive Infrastructure Autoscaling with Cluster API: Beyond just scaling pods, the solution extends to dynamically provisioning and de-provisioning GPU-enabled Kubernetes nodes. While Carpenter is mentioned for AWS/Azure, Cluster API is highlighted as a robust, open-source solution for managing Kubernetes clusters and nodes, including scaling Machine Deployments. KEDA can effectively scale these Machine Deployments, enabling true end-to-end infrastructure elasticity for LLM workloads.
- Cost Optimization through Intelligent Scaling: The presented architecture facilitates significant cost savings by allowing GPU nodes to scale down to zero during off-hours using cron-based triggers within KEDA. This ensures that expensive GPU resources are only provisioned when actively needed, directly impacting operational expenditures.
In essence, the key finding is that by strategically combining these powerful open-source components, it's possible to build a highly optimized, responsive, and cost-effective autoscaling system for LLM workloads that overcomes the inherent limitations of traditional Kubernetes autoscaling mechanisms.
Technical Deep Dive
▶ Watch: Paradigm shift: LLM workloads and unsuitable existing metrics (5:45)
The proposed solution meticulously orchestrates several open-source components to deliver a flexible and efficient autoscaling mechanism for LLM workloads on Kubernetes. The core architecture revolves around OpenTelemetry for metric ingress, KEDA for intelligent scaling decisions, and a custom OpenTelemetry Add-on for KEDA to facilitate their integration.
The Autoscaling Pipeline for LLMs
The traditional Kubernetes autoscaling model involves metrics being collected, stored in a metric registry, served by a metric server, and then consumed by the HPA for scaling decisions. For LLMs, this model is problematic due to the inadequacy of standard metrics and the need for rapid, push-based updates.
The optimized pipeline addresses these challenges by:
- Source of Metrics: LLM workloads themselves expose custom, GPU-specific metrics (e.g., KV cache percentage, waiting queue length) adhering to a standard format (e.g., Prometheus metrics endpoint).
- Metric Collection: OpenTelemetry Collector instances are deployed to scrape or receive these metrics. These collectors can be deployed in two primary configurations:
- Shared Collector: A single OTel Collector scrapes metrics from multiple replica pods.
- Sidecar Model: An OTel Collector is injected as a sidecar container into each LLM pod, often managed by the OpenTelemetry Operator. This setup offers faster reaction times due to closer proximity to the workload.
- Metric Processing & Export: The OTel Collector can perform various processing steps, such as filtering specific metrics, renaming them, or even manipulating their values. For instance, in the demo, a float value (KV cache percentage) was multiplied by 100 to convert it to an integer, as KEDA historically had limitations with float values in its triggers (a limitation slated for resolution in future KEDA releases). These processed metrics are then exported using the OTLP (OpenTelemetry Protocol).
- Metric Serving & Bridging: The exported OTLP metrics are received by the OpenTelemetry Add-on for KEDA. This add-on functions as an OTLP receiver, acting as a sink for the metrics. Crucially, it also contains a lightweight, short-term memory storage for a few data points, essentially a custom-built, minimal Prometheus-like store optimized for scaling purposes. This avoids the overhead and potential overwhelming of a full Prometheus instance, which is typically used for broader monitoring and alerting. The add-on then communicates these metrics to KEDA.
- Scaling Decision: KEDA consumes these metrics from the OpenTelemetry Add-on. KEDA utilizes its
ScaledObjectCustom Resource Definition (CRD) to define scaling rules. AScaledObjectpoints to a target deployment (e.g., the LLM inference pods) and specifies one or more triggers based on the custom metrics. KEDA then dynamically adjusts the number of replicas of the target deployment, including scaling down to zero when demand is absent. TheScaledObjectcan also configure astabilizationWindowto prevent rapid scaling fluctuations (thrashing), with different values for scaling up (e.g., 1 second) and scaling down (e.g., 20 minutes) to account for resource spin-up times and cost efficiency.
Node Autoscaling with Cluster API
Beyond just scaling pods, the solution extends to the underlying GPU infrastructure. LLMs require GPU-enabled nodes, which are expensive. Dynamically provisioning and de-provisioning these nodes is critical for cost optimization.
- Cluster API: This is a set of Kubernetes Custom Resource Definitions (CRDs) and controllers that enable declarative, Kubernetes-native management of Kubernetes clusters themselves, including nodes. It allows users to define Machine Deployments, which are scalable resources representing a group of nodes.
- KEDA Integration: KEDA can be configured to scale these Machine Deployments. A separate
ScaledObjectis created, with itsscaleTargetRefpointing to aMachineDeployment(e.g.,demo-gpu-nodes). ThisScaledObjectcan incorporate multiple triggers: - Resource-based: Based on the overall cluster resource utilization (e.g., pending pods requiring GPUs).
- Cron-based: For cost optimization, a cron-based trigger can be used to scale the
MachineDeploymentto zero replicas during off-hours (e.g., 8 PM to 10 AM) and back to one (or more) during business hours. This ensures expensive GPU nodes are only running when AI experts are actively using them.
Challenges in Node Autoscaling
While powerful, node autoscaling for LLMs introduces specific challenges:
- Data Locality: LLM models are huge (gigabytes). Ensuring that the data (model weights) is close to the compute (GPU node) to minimize loading times is crucial.
- Bootstrapping Time: Spinning up a new GPU node involves several steps: creating the VM, joining it to the Kubernetes cluster (
kubeadm join), pulling necessary container images, and installing Nvidia drivers. The GPU operator, while convenient, can add 3 minutes to the startup process for driver installation. Optimizations like baking container images directly into the VM image (using tools like Cluster API's image builder) can mitigate some of this, but driver installation remains a bottleneck. - Load Balancing: Ensuring that traffic is directed to underutilized GPUs rather than over-pressured ones is a further improvement area, though the metrics collected can inform such decisions.
By meticulously integrating these components, the architecture provides a robust, adaptable, and cost-effective solution for autoscaling complex LLM workloads in cloud-native environments.
Demo / Proof of Concept
▶ Watch: Five key challenges for autoscaling LLM workloads (6:55)
The live demonstration provided a practical walkthrough of the proposed autoscaling solution for LLM workloads. The setup involved a Kubernetes cluster in Google Cloud Platform (GCP) running an Open Web UI pod (a web interface for LLMs) and a Llama v3 model deployed as the inference workload.
Initial Setup and Interaction
- LLM Access: The Llama v3 model was accessible via an ingress resource,
llm.kubecon.dev. The speaker demonstrated interacting with the model using both the web UI andcurlcommands, simulating the OpenAI HTTP protocol. - Request Variation: The
curlcommands highlighted the ability to vary the load on the LLM by adjusting parameters likestream=true/false(for streaming token responses) andmax_tokens(determining response length). A key point was made about howmax_tokensdirectly influences the GPU pressure, as longer responses require more computation. - Metric Exposure: The VLM runtime used for the Llama v3 model exposes internal GPU statistics. The demo specifically focused on two critical metrics: KV cache percentage (representing GPU memory usage for the key-value cache) and the waiting queue length for requests. These metrics were shown to be accessible by port-forwarding the metrics endpoint of a replica pod and then
curlinglocalhost:8080/metrics.
Pod Autoscaling in Action
- ScaledObject Configuration: The speaker described the
ScaledObjectfor the LLM model deployment. ThisScaledObjectwas configured to scale thellama-v3-modeldeployment based on a PromQL-like syntax trigger, which evaluates theKV cache percentagemetric. A notable detail was the multiplication of the floatKV cache percentageby100within the OpenTelemetry Collector configuration to make it compatible with KEDA's earlier integer-only trigger support. The threshold for scaling up was set to30. - OpenTelemetry Configuration: The OpenTelemetry CRD was shown, detailing how the OpenTelemetry Operator injects sidecar collectors into the LLM pods. This configuration included the target address for sending metrics (the OpenTelemetry Add-on for KEDA), filtering rules to collect only relevant metrics (KV cache, waiting queue), and the metric manipulation logic (multiplying by 100).
- Load Generation: To simulate high demand, the
heycommand was used to generate300concurrent HTTP POST requests to the LLM. Each request specified4k tokensfor the response, ensuring significant pressure on the GPU. - Observed Scaling: Immediately after initiating the load, the
kubectl get pods -wcommand showed thellama-v3-modeldeployment rapidly scaling up from one replica to four (though limited by available GPUs to three, as only two GPUs were available in the cluster). This demonstrated KEDA successfully detecting the increased KV cache percentage above the30%threshold and initiating pod scaling.
Node Autoscaling with Cluster API
- Initial Node State: The cluster initially had a limited number of GPU-enabled nodes. The
kubectl get nodes -wcommand confirmed the existing node count. - Unpausing Node ScaledObject: A separate
ScaledObjectwas configured for autoscaling the Kubernetes nodes themselves, specifically targeting a Cluster APIMachineDeploymentnameddemo-gpu-nodes. ThisScaledObjectwas initially paused. The speaker unpaused it usingkubectl annotate scaledobject demo-gpu-nodes keda.sh/paused=false. - Observed Node Provisioning: After unpausing, the
kubectl get machinedeployments -wandkubectl get nodes -wcommands showed the Cluster APIMachineDeploymentstarting to provision new nodes. This process was acknowledged to take several minutes due to the need to create a GCP VM, runkubeadm join, pull necessary images (optimized by image builder), and install Nvidia drivers via the GPU operator (which alone takes ~3 minutes). - Cost Optimization Feature: The
ScaledObjectfor nodes also demonstrated a cron-based trigger. This trigger was configured to maintain a minimum of one replica during "office hours" (10 AM to 8 PM) and scale down to zero replicas during off-hours. This highlights the cost-saving potential by de-provisioning expensive GPU nodes when not actively used. - Stabilization Window: During the Q&A, the concept of
stabilizationWindowin KEDA was discussed. For scaling up, a short window (e.g., 1 second) is desirable for quick reaction. For scaling down, a longer window (e.g., 20 minutes) is crucial, especially for nodes, to prevent thrashing and costly de-provisioning/re-provisioning cycles.
The demo successfully illustrated the end-to-end functionality of the proposed solution, from custom metric collection and processing to intelligent pod and node autoscaling, directly addressing the unique demands of LLM workloads.
Defensive Implications
▶ Watch: Analyzing limitations of existing autoscaling solutions (KPA) (8:10)
The detailed autoscaling solution presented in this talk offers several critical defensive implications for organizations operating or planning to deploy Large Language Model workloads on Kubernetes:
- Adopt Specialized LLM Metrics: Defenders must move beyond traditional CPU, memory, and RPS metrics for autoscaling LLM infrastructure. It is crucial to instrument LLM inference applications to expose GPU-specific metrics like KV cache percentage and waiting queue length. These metrics provide a more accurate reflection of actual GPU utilization and potential bottlenecks, enabling proactive scaling decisions rather than reactive ones based on less relevant indicators.
- Implement Robust Metric Collection with OpenTelemetry: Standardize metric collection using OpenTelemetry. This ensures a consistent, vendor-neutral approach that is flexible enough to capture custom application-level metrics. Deploying OTel Collectors as sidecars within LLM pods (managed by the OpenTelemetry Operator) can provide the lowest latency for metric exposure, allowing for faster scaling reactions to sudden spikes in demand.
- Leverage KEDA for Advanced Autoscaling Logic: Integrate KEDA into your Kubernetes clusters to extend autoscaling capabilities beyond HPA's limitations. KEDA's
ScaledObjectCRD allows for highly configurable scaling rules based on external, custom metrics. This provides the agility needed to respond to the dynamic and often unpredictable demands of LLM inference.
- Bridge OpenTelemetry and KEDA with the OTel Add-on: For optimal integration, consider deploying the OpenTelemetry Add-on for KEDA. This component streamlines the flow of custom OpenTelemetry metrics to KEDA, offering a lightweight, purpose-built metric store for scaling decisions, thus avoiding the overhead and potential contention of using a full Prometheus instance solely for autoscaling.
- Strategically Manage Scaling Windows: Configure KEDA's
stabilizationWindowcarefully. For scaling up, a short window (e.g., 1 second) is generally desirable to react quickly to increased load. However, for scaling down, especially for expensive GPU nodes, a longer window (e.g., 20 minutes) is critical. This prevents "thrashing" (rapid scale-up and scale-down cycles) and ensures that nodes are not prematurely de-provisioned, which can be costly in terms of both re-provisioning time and potential service disruption.
- Implement GPU Node Autoscaling with Cluster API: To achieve true elasticity and cost efficiency, integrate Cluster API to dynamically scale your GPU-enabled Kubernetes nodes. KEDA can be configured to scale Machine Deployments, ensuring that expensive GPU resources are provisioned only when needed. This is vital for managing operational costs, especially in environments where GPU resources are expensive and not continuously utilized.
- Optimize Node Bootstrapping for GPUs: Address the challenges of fast GPU node provisioning. This involves optimizing node images to include necessary dependencies (like container runtimes and potentially pre-baked container images for LLMs) and streamlining Nvidia driver installation. While the GPU operator is convenient, its ~3-minute installation time for drivers can impact scale-up speed. Explore alternative methods like baking drivers directly into the node image if feasible for your cloud provider and Kubernetes version.
- Consider Data Locality for Large Models: For extremely large LLM models, plan for data locality. Ensure that model weights can be loaded quickly onto newly provisioned GPU nodes. This might involve pre-fetching, local caching, or intelligent scheduling to place pods on nodes that already have the required model data.
By proactively addressing these areas, organizations can build a resilient, cost-effective, and high-performance infrastructure for LLM workloads, capable of dynamically adapting to varying demand while maintaining service quality and optimizing resource utilization.
Key Takeaways
- LLM Autoscaling Requires Specialized Metrics: Traditional CPU, memory, and RPS metrics are insufficient for effectively scaling GPU-intensive LLM workloads. Critical metrics like KV cache percentage and waiting queue length provide more accurate indicators of GPU utilization and inference demand.
- OpenTelemetry and KEDA Form a Powerful Synergy: Leveraging OpenTelemetry for standardized, push-based metric collection and KEDA for flexible, event-driven autoscaling provides a robust, extensible, and portable solution for LLM environments.
- The OTel Add-on for KEDA Bridges Critical Gaps: This custom component efficiently transfers OpenTelemetry metrics to KEDA, offering a lightweight, dedicated metric store for scaling that avoids overwhelming general-purpose monitoring systems like Prometheus.
- End-to-End Autoscaling Includes Infrastructure: The solution extends beyond just scaling pods to dynamically provisioning and de-provisioning GPU-enabled Kubernetes nodes using Cluster API and KEDA, enabling true infrastructure elasticity and cost optimization.
- Cost Efficiency is Achieved Through Intelligent Scaling: Features like scaling to zero for pods and cron-based node autoscaling for off-hours, combined with carefully configured stabilization windows, significantly reduce operational costs associated with expensive GPU resources.
- Leveraging Open Source "Giants" Accelerates Development: By building upon established and actively developed open-source projects like OpenTelemetry, KEDA, and Cluster API, the solution avoids reinventing the wheel and benefits from community contributions.
About the Speaker(s)
Vincent Hou is a Senior Software Engineer at Bloomberg, where his team specializes in building and maintaining an AI inference platform. He has been a lead in the Knative operation workgroup for six years and is a passionate evangelist and contributor to open-source projects, including OpenStack and OpenKurf, for over a decade. His expertise lies in developing scalable, cloud-agnostic solutions for modern AI workloads.
Jiří Kremser works at Kify, a company focused on production-grade KEDA implementations. He is also a dedicated contributor to the KEDA project itself, notably developing the OpenTelemetry Add-on for KEDA, which was a central component of this talk. Jiří is an advocate for open source and has interests in areas like 3D printing. He brings deep practical knowledge of KEDA and event-driven architectures to the cloud-native community.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk presented a highly effective and innovative open-source architecture for autoscaling GPU-intensive LLM workloads on Kubernetes. By intelligently combining OpenTelemetry for custom metric collection, KEDA for event-driven scaling, and a novel OpenTelemetry Add-on for KEDA, the speakers demonstrated how to dynamically adjust both LLM inference pods and underlying GPU nodes based on crucial LLM-specific metrics like KV cache percentage and waiting queue length. This provides a robust, cost-efficient, and actionable solution for a critical modern infrastructure challenge, delivered with genuine technical depth and a practical live demonstration.
Heather Calloway (CISO) — STRONG ACCEPT
This talk provides a highly actionable and well-articulated solution for the critical challenge of autoscaling expensive LLM workloads on Kubernetes. By championing specialized GPU-specific metrics, OpenTelemetry for robust collection, and KEDA with its custom add-on for flexible scaling, the speakers deliver a blueprint for achieving significant cost efficiencies and enhancing the operational resilience of AI platforms. It's a clear demonstration of institutional realism, offering concrete steps for platform teams to manage a substantial business risk.