Yes You Can Run LLMs on Kubernetes - Abdel Sghiouar & Mofi Rahman, Google Cloud
Abdel Sghiouar, Mofi Rahman, Google Cloud
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this KubeCon EU talk, Google Cloud's Abdel Sghiouar and Mofi Rahman tackle a pressing question for many organizations: how to effectively deploy and manage Large Language Models (LLMs) on Kubernetes. Far from being a mere theoretical exercise, the session provides a comprehensive guide, complete with live demonstrations, on leveraging Kubernetes' robust orchestration capabilities for the unique demands of LLM inference. The speakers argue that while cloud-hosted LLM services offer convenience, Kubernetes provides a powerful middle ground, offering both granular control over infrastructure and the scalability and flexibility of a managed platform.

Key moments
- 0:00 Introduction: Yes, you can run LLMs on Kubernetes
- 2:20 Why Kubernetes is ideal for running LLMs
- 3:26 K8s LLM deployment demo using GKE Autopilot
- 3:40 Starting the Kubernetes deployment YAML for LLM
- 4:40 Configuring GPU, CPU, memory, and ephemeral storage
- 5:20 Setting VLM serving command, model ID, and parameters
Yes You Can Run LLMs on Kubernetes
Speakers: Abdel Sghiouar, Staff Developer Advocate, Google Cloud; Mofi Rahman, Staff Developer Advocate, Google Cloud
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=thCZDKZ1cAM
Overview
In this KubeCon EU talk, Google Cloud's Abdel Sghiouar and Mofi Rahman tackle a pressing question for many organizations: how to effectively deploy and manage Large Language Models (LLMs) on Kubernetes. Far from being a mere theoretical exercise, the session provides a comprehensive guide, complete with live demonstrations, on leveraging Kubernetes' robust orchestration capabilities for the unique demands of LLM inference. The speakers argue that while cloud-hosted LLM services offer convenience, Kubernetes provides a powerful middle ground, offering both granular control over infrastructure and the scalability and flexibility of a managed platform.
The core message is that LLMs, despite their computational intensity and large footprints, can indeed be treated as "the new web application" within a Kubernetes paradigm. The talk delves into the specific challenges posed by these models—such as massive container image sizes, significant GPU memory requirements, and distributed inference needs—and showcases how Kubernetes, with its evolving API landscape and device management plugins, is well-equipped to handle them. For organizations seeking to self-host open-weight LLMs, maintain multi-cloud flexibility, and leverage a standardized API for their AI workloads, this presentation offers invaluable insights and practical architectures.
This session is particularly relevant for platform engineers, MLOps practitioners, and developers looking to integrate LLMs into their applications while maintaining control over their deployment environment. It demystifies the complexities of LLM serving, providing a roadmap for deploying models ranging from single-GPU instances to multi-node, sharded setups. By demonstrating real-world deployments on GKE Autopilot and exploring various model serving engines, Sghiouar and Rahman illustrate the practical feasibility and strategic advantages of running LLMs on Kubernetes.
Background
▶ Watch: Introduction: Yes, you can run LLMs on Kubernetes (0:00)
The landscape of Large Language Models has evolved rapidly, with new models emerging frequently, pushing the boundaries of scale and capability. While many prominent models are offered as managed services by cloud providers (e.g., Vertex AI, Bedrock, Azure Managed Services), there's a growing desire among organizations to host and manage open-weight models themselves. These models, often released on platforms like Hugging Face or Kaggle, offer greater flexibility, customization potential, and control over data and inference. However, their sheer size and computational requirements present significant deployment challenges. For instance, the DeepSeek R1 model alone boasts 671 billion parameters and an 800 GB file size, making traditional deployment approaches untenable.
Clayton Coleman's statement, "LLM is the new web application," encapsulates the sentiment that these models, while fundamentally different in their resource demands, can benefit from the same orchestration principles that made Kubernetes indispensable for web applications. Kubernetes excels at automation, scalability, and device management, which are all crucial for LLMs. Its device plugin architecture, for example, allows seamless integration and scheduling of specialized hardware like GPUs and TPUs. Furthermore, Kubernetes offers inherent multi-cloud capabilities and a standardized single API, providing a consistent platform regardless of the underlying infrastructure—a key advantage for avoiding vendor lock-in and ensuring portability.
The problem, therefore, isn't just about running an application, but about running an application that might require 20 GB of memory, four CPUs, a GPU, significant ephemeral storage, and potentially access to remote endpoints for downloading model weights. These requirements far exceed those of a typical web application, necessitating a deep understanding of resource provisioning, hardware acceleration, and distributed computing patterns within the Kubernetes ecosystem. The talk aims to bridge this gap by demonstrating how Kubernetes can be configured and extended to meet these unique demands, providing a robust and flexible platform for LLM serving.
Key Findings
▶ Watch: K8s LLM deployment demo using GKE Autopilot (3:26)
The talk highlights several key findings regarding the successful deployment of LLMs on Kubernetes:
- Kubernetes as a Versatile Platform: Kubernetes, originally designed for stateless web applications, proves to be a highly adaptable platform for LLMs. It offers a crucial middle ground between fully managed services and bare-metal deployments, providing developers with both control and flexibility. Its core strengths—automation, scalability, and a unified API—are directly applicable to managing the lifecycle of LLM workloads.
- Specialized Resource Management: LLMs necessitate specialized hardware like GPUs and TPUs. Kubernetes' existing device management mechanisms, such as device plugins and node selectors, are fundamental for scheduling LLM pods onto nodes with the required accelerators. The talk demonstrates how GKE Autopilot can further simplify this by automatically provisioning GPU-enabled nodes on demand, abstracting away much of the underlying infrastructure management.
- The Challenge of Model Size: The enormous size of LLMs (e.g., Gemma 3 1B requiring 20GB memory, DeepSeek R1 at 800GB file size) presents significant challenges for container image distribution, runtime memory, and persistent storage. This necessitates optimizations like quantization to reduce memory footprint and strategies for caching model weights on persistent volumes rather than repeatedly downloading them.
- Distributed Inference is Essential: For the largest models, a single GPU or even a single node is insufficient. The talk underscores the necessity of model sharding and multi-host multi-accelerator deployments. Kubernetes natively supports single-host multi-accelerator scenarios (e.g., splitting a model across multiple GPUs on one node). For true multi-node sharding, the emerging Leader Worker Set (LWS) API is presented as a critical enabler, allowing coordinated distribution of model shards across multiple nodes and GPUs.
- Diverse Model Serving Engines: The ecosystem of LLM serving engines is rich and varied. The talk demonstrates the flexibility of Kubernetes by deploying models using different engines such as VLM, Olama, and TGI (Text Generation Inference), and integrating them with distributed computing frameworks like Ray Serve. This highlights that Kubernetes provides the underlying orchestration layer, allowing users to choose the best serving engine for their specific model and performance requirements.
- Operational Efficiency through Kubernetes Principles: The same DevOps and GitOps principles applied to traditional applications—infrastructure as code, multi-tenancy, job queuing (e.g., Kueue), and autoscaling—are equally vital for LLMs. By adopting these practices, organizations can build robust, scalable, and cost-efficient AI platforms on Kubernetes.
Technical Deep Dive
▶ Watch: Starting the Kubernetes deployment YAML for LLM (3:40)
The technical core of the talk revolves around demonstrating the practical deployment of LLMs on Kubernetes, addressing the unique challenges they present.
1. Basic LLM Deployment with YAML:
Mofi Rahman begins by crafting a Kubernetes Deployment YAML for the Gemma 3 1B parameter model. Key elements of this deployment include:
- Image: A VLM (vLLM) based image is used as the serving engine, providing a high-performance inference server. The image itself is substantial, highlighting a challenge discussed later.
- Resource Requests and Limits: Crucial for LLMs, the deployment specifies significant resources:
memory: 20Gicpu: 4nvidia.com/gpu: 1(requesting one NVIDIA GPU)ephemeral-storage: 20Gi(for downloading model data)- Command and Arguments: The VLM image runs a Python script with arguments like
model-id(e.g.,gemma-3-1b-it),tensor-parallel-size(set to1for this small model),max-model-len(32,000 tokens for Gemma 3 1B), andport. - Environment Variables: A
MODEL_IDenvironment variable is set. - Secrets for Gated Models: Many open-weight models, like Gemma, Mistral, or Llama, are gated models requiring consent and an API token (e.g., from Hugging Face). The deployment injects this token as a Kubernetes Secret.
- Persistent Storage: A volume mount and volume are configured to store the downloaded model weights locally, preventing repeated downloads.
- Node Selector: A
nodeSelectoris used to ensure the pod is scheduled on a node with a GPU, specificallycloud.google.com/gke-accelerator: nvidia-tesla-t4. This is critical for GKE Autopilot to automatically provision a new node with the specified GPU on demand. - Service: A Kubernetes Service of type
ClusterIPis created to expose the VLM server on port 8000.
2. Core LLM Concepts Explained:
The speakers provide essential definitions for understanding LLM infrastructure:
- LLM (Large Language Model): A model capable of generating human-like text.
- Inference / Serving / Model Server: A software component (like VLM, TGI) that provides an interface (usually REST) to interact with a pre-trained LLM, translating application requests into model predictions.
- Accelerators: Specialized hardware optimized for matrix computations, such as GPUs (NVIDIA H100, A100), TPUs (Google's Tensor Processing Units), or emerging chips like Groq.
- Quantization: Reducing the precision of model weights (e.g., from 32-bit floating-point to 16-bit or 8-bit integers). This significantly reduces memory footprint, allowing larger models to fit onto smaller GPUs, at the potential cost of slight quality degradation.
- Weights: The numerical parameters within an LLM that represent its learned knowledge of the world.
- Context Window: The maximum number of tokens an LLM can process or generate in a single interaction. Gemma 3 1B supports 32,000 tokens, while Gemini can handle up to 2 million.
- Multimodal: Models capable of processing and generating content across different modalities (e.g., text-to-image, image-to-text, text-to-video), beyond the text-to-text focus of many current LLMs.
3. Model Serving Engines:
Beyond VLM, several other open-source and proprietary serving engines are discussed:
- TGI (Text Generation Inference): From Hugging Face, purpose-built for efficient text generation.
- NVIDIA NIM: NVIDIA's proprietary inference microservice, built on open standards, optimized for their GPUs.
- Ray Serve: Part of the Ray ecosystem, designed for distributed machine learning workloads, including model serving across clusters.
- Jax / Jetstream: An open-source project from Google, particularly strong for TPUs.
- Olama: Allows running LLMs locally or in containers, providing a familiar interface for cloud deployments.
4. The Scale Problem: Image Size and GPU Memory:
LLMs introduce unprecedented scale challenges:
- Container Image Size: Modern LLM images can be enormous (e.g., VLM images around 10 GB), contrasting with the long-standing best practice of small, portable containers. This impacts download times and storage.
- Model Size vs. GPU Memory: The memory required by an LLM is directly proportional to its parameter count and precision.
- Full Precision (32-bit): Requires 4 bytes per parameter. A 4 billion parameter model needs 16 GB of VRAM.
- Half Precision (BF16, FP16): Requires 2 bytes per parameter. A 27 billion parameter model needs roughly 54 GB of VRAM.
- The DeepSeek R1 model, at 671 billion parameters, would theoretically require 1.37 terabytes of GPU memory in full precision—a quantity far exceeding any single GPU currently available. Even the largest GPUs like the NVIDIA H200 (141 GB) or H100 (80 GB) cannot accommodate it.
- Physical limitations typically restrict a single node to around 8 GPUs (e.g., 8 x A100 with 40 GB each = 320 GB total), necessitating multi-node solutions for very large models.
5. Distributed LLM Inference with Leader Worker Set (LWS):
To address the memory constraints of colossal models, model sharding is crucial.
- Single Host, Single/Multi Accelerator: Kubernetes supports deploying a model on one or more GPUs within a single node. The
tensor-parallel-sizeparameter in VLM allows sharding across GPUs on the same host. - Multi-Host, Multi-Accelerator: For models like DeepSeek R1, the model must be split across multiple nodes, each with multiple GPUs. This is where the new Kubernetes Leader Worker Set (LWS) API comes into play. LWS defines a leader (which orchestrates the sharding and distributes work) and a set of workers (which run actual model shards on their respective nodes/GPUs). This API enables true distributed LLM inference across a Kubernetes cluster.
6. Optimization Dimensions:
The talk briefly touches on optimization strategies:
- Image Caching: Pre-caching large container images on nodes to reduce startup times.
- Data Caching: Storing model weights on persistent volumes (e.g., using PVCs or cloud-specific solutions like HDML for Google Cloud) to avoid downloading 800 GB models from Hugging Face every time a pod starts.
- Workload Scaling: Intelligent autoscaling of GPU nodes and pods to match demand, preventing idle resources.
Demo / Proof of Concept
▶ Watch: Configuring GPU, CPU, memory, and ephemeral storage (4:40)
The live demonstration was a central and highly effective component of the talk, showcasing the practical aspects of running LLMs on Kubernetes.
- Gemma 3 1B Deployment on GKE Autopilot:
Mofi Rahman started by deploying the Gemma 3 1B model using the YAML described in the technical deep dive. The demonstration highlighted:
- The use of GKE Autopilot, which automatically provisions a new node with a NVIDIA T4 GPU upon detecting the
nvidia.com/gpu: 1resource request and the appropriatenodeSelector. - The initial
Pendingstate of the pod as Kubernetes and GKE Autopilot work to provision the node and schedule the workload. - After a few minutes (about four minutes in the demo), the pod transitioned to
Running, and logs from the VLM server confirmed its readiness. This illustrates the time taken for image download and model loading.
- Multi-Model Chat UI:
To further illustrate the flexibility and power of Kubernetes, Abdel Sghiouar presented a custom web UI that allowed simultaneous interaction with multiple LLMs deployed on the same cluster but in different namespaces and using different serving engines. This UI was accessible via a QR code, inviting audience participation.
- The UI demonstrated querying various models, including:
- Gemma 12B: Served using Olama as the serving engine.
- Llama 3 3B: Served using TGI.
- Llama 3 8B: Notably, this version was deployed on a TPU (Tensor Processing Unit), showcasing Kubernetes' ability to orchestrate non-GPU accelerators. From the user's perspective, the API remained consistent thanks to VLM providing a unified interface on top of the TPU.
- Mistral: Another popular open-weight model.
- DeepSeek R1: The massive 671 billion parameter model, specifically deployed using a Leader Worker Set (LWS), Ray, and VLM. This setup involved two nodes, each equipped with eight NVIDIA H100 GPUs, totaling 16 H100 GPUs to serve the model. The demo explicitly showed the YAML for this complex setup, highlighting the
LeaderWorkerSetdefinition, theleaderstarting as a single worker, and aworkerPooltemplate with single H100 nodes. - A "knock-knock" joke prompt with a "British humor" element was used to demonstrate the models' responses.
- Persistent Storage for Large Models:
For the DeepSeek R1 deployment, the demo also highlighted the use of HDML (a Google Cloud-specific storage solution, analogous to Persistent Volume Claims in generic Kubernetes) to store the 800 GB model data. This prevents the model from being downloaded repeatedly on pod restarts or scaling events, significantly reducing startup times and egress costs. The speakers emphasized that similar solutions using standard PVCs or FUSE drivers for object storage buckets (like S3 or GCS) can be implemented in any cloud or on-premises environment.
The entire demo environment, including the YAML configurations for all models and the UI code, was made open source, allowing attendees to explore and replicate the deployments. This practical demonstration effectively validated the talk's premise that complex LLM workloads can be robustly managed and scaled on Kubernetes.
Defensive Implications
▶ Watch: Setting VLM serving command, model ID, and parameters (5:20)
While this talk focuses on the operational aspects of deploying LLMs rather than direct security vulnerabilities, there are significant "defensive implications" in terms of building resilient, secure, and efficient LLM platforms on Kubernetes. These implications primarily revolve around robust operational practices and resource management:
- Secure Management of Gated Model Tokens: LLMs like Gemma or Llama are often "gated," requiring API tokens (e.g., from Hugging Face) for access. These tokens are sensitive credentials. Defenders must ensure these tokens are securely stored and injected into Kubernetes pods, ideally using Kubernetes Secrets and restricting access via Role-Based Access Control (RBAC). Avoid hardcoding tokens in images or YAMLs.
- Resource Quotas and Limits: LLMs are resource-hungry. Implementing resource quotas at the namespace level and strict resource limits on pods prevents resource starvation, ensures fair sharing of expensive GPU resources, and protects the cluster from runaway processes. This is a critical defense against denial-of-service (DoS) from misconfigured or malicious workloads.
- Image Security and Supply Chain: The enormous size of LLM container images (e.g., 10 GB VLM images) increases the attack surface. Defenders should implement robust image scanning, use trusted image registries, and ensure a secure software supply chain for all base images and model serving engines. Regularly patching and updating images is crucial.
- Persistent Storage Security: Storing large model weights on Persistent Volume Claims (PVCs) or cloud-specific storage (like HDML) requires appropriate access controls. Ensure that only authorized pods and users can mount and access these volumes, protecting the integrity and confidentiality of the model data. Encryption at rest and in transit for these volumes is also a best practice.
- Network Security: Exposing LLM inference endpoints via Kubernetes Services requires careful network segmentation. Use Network Policies to restrict which applications can communicate with the LLM services. If exposing externally, use secure ingress controllers with TLS termination and consider Web Application Firewalls (WAFs) to protect against common web attacks.
- Observability and Monitoring: Comprehensive monitoring of LLM pods, nodes, and GPU utilization is essential. This includes tracking GPU memory usage, inference latency, error rates, and resource consumption. Robust observability allows defenders to quickly detect and respond to performance degradation, resource exhaustion, or anomalous behavior that could indicate operational issues or attacks.
- Autoscaling and Cost Management: While not strictly security, inefficient resource usage can lead to ballooning costs, which is a defensive concern against financial waste. Implementing Horizontal Pod Autoscalers (HPAs) and Cluster Autoscalers (or leveraging GKE Autopilot) ensures that resources are scaled up and down based on demand, optimizing cost efficiency and maintaining service availability under varying load.
- Multi-Tenancy and Isolation: In multi-tenant Kubernetes clusters, strict isolation between different LLM workloads and teams is paramount. This can be achieved through namespaces, RBAC, network policies, and potentially node isolation for highly sensitive workloads, preventing lateral movement and unauthorized access.
- Data Governance and Compliance: When using LLMs, especially with sensitive input data, adhering to data governance policies and regulatory compliance (e.g., GDPR, HIPAA) is critical. Kubernetes provides primitives to help enforce data locality and access controls, but the application layer must also be designed with compliance in mind.
By adopting these defensive operational strategies, organizations can build secure, reliable, and cost-effective platforms for deploying LLMs on Kubernetes, mitigating risks associated with their unique demands.
Key Takeaways
- Kubernetes is a Viable LLM Platform: Despite their unique demands, LLMs can be effectively deployed and managed on Kubernetes, offering a balance of control and scalability between bare metal and fully managed cloud services.
- Resource Management is Paramount: LLMs are resource-intensive, requiring careful allocation of GPUs, TPUs, significant memory (e.g., 20GB for Gemma 3 1B), and ephemeral storage. Kubernetes' device management and node selectors are crucial for efficient scheduling.
- Scale Requires Advanced Techniques: For very large models (e.g., DeepSeek R1 at 800GB), quantization reduces memory footprint, and model sharding across multiple GPUs and nodes using the Leader Worker Set (LWS) API is essential for distributed inference.
- Diverse Serving Engines are Supported: Kubernetes accommodates various LLM serving engines like VLM, TGI, Olama, and Ray Serve, allowing flexibility in choosing the best tool for specific models and performance requirements.
- Data and Image Caching are Critical for Performance: Due to massive image (10GB+) and model (800GB+) sizes, caching container images and model weights on persistent storage (e.g., PVCs, HDML) significantly reduces startup times and improves operational efficiency.
- Leverage Kubernetes Ecosystem for MLOps: Existing Kubernetes principles like autoscaling, multi-tenancy, GitOps, and job queuing (e.g., Kueue) are directly applicable to building robust, scalable, and cost-efficient MLOps platforms for LLMs.
About the Speaker(s)
Abdel Sghiouar and Mofi Rahman are both Staff Developer Advocates at Google Cloud. In their roles, they focus on empowering developers and organizations to leverage Google Cloud technologies effectively. Their expertise spans Kubernetes, cloud infrastructure, and emerging areas like Large Language Models, enabling them to provide practical guidance and showcase real-world solutions for complex technical challenges.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk by Sghiouar and Rahman delivers a much-needed, no-nonsense guide to deploying Large Language Models on Kubernetes. Eschewing marketing fluff, they dive into the nitty-gritty of resource management, GPU allocation, and advanced techniques like model sharding with the nascent Leader Worker Set (LWS) API. The live demonstrations, especially the multi-node, multi-accelerator setup for the DeepSeek R1 model, prove that treating LLMs as 'the new web application' on Kubernetes is not just theoretical but demonstrably achievable, providing critical actionable insights for platform engineers struggling with open-weight models.
Heather Calloway (CISO) — STRONG ACCEPT
This KubeCon session by Google Cloud's Abdel Sghiouar and Mofi Rahman provides a highly practical and technically sound roadmap for deploying Large Language Models on Kubernetes. While focused on the operational 'how,' it directly addresses a critical institutional need: gaining granular control and flexibility over AI infrastructure rather than relying solely on opaque managed services. The talk effectively demystifies the complexities of LLM serving at scale, offering actionable insights for platform engineers and MLOps practitioners that directly contribute to a more secure and resilient AI strategy within an organization.