Production-Ready LLMs on Kubernetes: Patterns, Pitfalls, and Performa... Priya Samuel & Luke Marsden
Priya Samuel, Luke Marsden
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In an era where the landscape of Artificial Intelligence evolves at an astonishing pace, organizations are increasingly seeking robust, secure, and cost-effective ways to deploy large language models (LLMs) on their own infrastructure. This talk by Priya Samuel and Luke Marsden at KubeCon EU addresses this critical need, guiding attendees through the intricate journey of building production-ready LLM platforms using Kubernetes and open-source models. The speakers share invaluable patterns, common pitfalls, and hard-won lessons from their extensive experience, aiming to demystify the perceived complexity of self-hosting advanced AI.

Key moments
- 0:00 Introduction: Running generative AI on own infrastructure
- 3:10 Understanding LLMs: A simplified mathematical explanation
- 4:18 Why run LLMs on Kubernetes: Security, cost, availability
- 6:38 Kubernetes advantages: Portability, consistency, scalability for LLMs
- 7:49 First demo: Setting up Olama with Open Web UI
- 8:35 Demo: Initial challenge with LLM API integration
Production-Ready LLMs on Kubernetes: Patterns, Pitfalls, and Performance
Speakers: Priya Samuel, Technology Architect; Luke Marsden, Founder/CEO
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=KIRUbaUjEKw
Overview
In an era where the landscape of Artificial Intelligence evolves at an astonishing pace, organizations are increasingly seeking robust, secure, and cost-effective ways to deploy large language models (LLMs) on their own infrastructure. This talk by Priya Samuel and Luke Marsden at KubeCon EU addresses this critical need, guiding attendees through the intricate journey of building production-ready LLM platforms using Kubernetes and open-source models. The speakers share invaluable patterns, common pitfalls, and hard-won lessons from their extensive experience, aiming to demystify the perceived complexity of self-hosting advanced AI.
The core premise of the presentation revolves around the growing apprehension among businesses regarding the security and compliance implications of entrusting their sensitive data to proprietary, hosted AI platforms. Samuel and Marsden present a compelling alternative: leveraging the power of Kubernetes for its inherent portability, consistency, and scalability to host open-source LLMs locally. This approach not only provides enhanced control over data residency and processing but also addresses concerns around availability, latency, and the long-term cost of ownership for AI at scale. The talk ultimately demonstrates how to move beyond basic LLM deployment to build sophisticated, integrated GenAI applications.
Background
▶ Watch: Introduction: Running generative AI on own infrastructure (0:00)
The foundational concept of an LLM, as simplified by the speakers, is a multi-dimensional function that maps input sentences onto other sentences. Essentially, a user's query, such as "What is the capital of France?", is converted into a vector (a string of numbers representing high-dimensional coordinates), processed by the model, and then translated back into a natural language response. This computationally intensive process fundamentally relies on specialized hardware, primarily GPUs (Graphics Processing Units) or TPUs (Tensor Processing Units). While it's possible to run these on bare metal, the speakers highlight the necessity of connecting these powerful processors to a Kubernetes cluster via a device plugin to make them accessible for containerized workloads.
The rationale for undertaking the complexity of running LLMs locally on Kubernetes is multifaceted and driven by several critical business imperatives. Foremost among these are security and compliance. Industries such as healthcare or European telecommunications face stringent data residency and processing regulations, often precluding the use of external hosted LLM platforms. By keeping data and processing within their own controlled environments, organizations can meet these mandates. Secondly, availability and latency are significant drivers; locally hosted LLMs can offer more consistent uptime and quicker response times compared to external APIs, which can sometimes be "shaky" as noted by the speakers. The rapid advancement and increasing sophistication of open-source LLMs, exemplified by models like DeepSeek, also make self-hosting a viable and attractive option, potentially even surpassing proprietary offerings in certain metrics. Lastly, cost of ownership becomes a significant factor at scale. While proprietary platforms might be cost-effective for quick proofs-of-concept, models requiring "chain of thought" processing—where the LLM "thinks longer" for better answers—can quickly shift the economic advantage towards locally hosted solutions. Kubernetes, with its inherent strengths in portability, consistency, and scalability, serves as the ideal orchestration layer for managing these complex, resource-intensive workloads, ensuring a consistent deployment experience across diverse environments.
Key Findings
▶ Watch: Why run LLMs on Kubernetes: Security, cost, availability (4:18)
The talk uncovers several crucial findings and best practices for deploying LLMs in production:
- Olama's Limitations for Production: While useful for testing, Olama's default configuration, which uses highly compressed quantized models (e.g., 4-bit per weight) and disables optimizations like flash attention, leads to truncated context lengths and garbled responses. This makes it unsuitable for production environments where accuracy and context are paramount.
- VLLM as a Production-Ready Inference Engine: The VLLM project emerges as a superior inference engine, offering significantly improved latency and tokens per second, alongside more appropriate context length defaults. It provides a robust foundation for serving LLMs in production.
- Strategic Model Weight Management: To ensure security for air-gapped or highly controlled environments, the speakers found it necessary to bake model weights directly into Docker images rather than relying on VLLM downloading them from external sources like Hugging Face. This approach, however, introduced a new challenge: Docker's default
gzip level 8compression significantly slowed down image startup. A key finding was the need to patch the Docker build process to use a custom compression algorithm to mitigate this performance bottleneck. - GPU Memory Sharing for Cost Efficiency: Running multiple, diverse LLMs (e.g., image models, multilingual models) on separate GPU node pools quickly becomes cost-prohibitive. A critical finding was the necessity of implementing GPU memory sharing and potentially building custom Kubernetes schedulers to maximize hardware utilization, enabling the simultaneous hosting of multiple models and the mixing of inference with fine-tuning jobs on the same GPU resources.
- Value in Application Layers: The true business value of LLMs is unlocked not merely by deploying the models, but by building intelligent application layers on top. This involves sophisticated API integrations for dynamic interaction with external systems and advanced knowledge retrieval (RAG - Retrieval Augmented Generation) pipelines to ground LLMs in specific organizational data.
- The Importance of Evaluation (Evals) and CI/CD: Ensuring the quality and reliability of GenAI applications requires a robust testing framework. The concept of evals, using an LLM as a judge, integrated into a CI/CD (Continuous Integration/Continuous Delivery) system, is presented as crucial for maintaining quality and enabling iterative development.
- The AI Spec for Standardization: The introduction of the AI spec, a proposed Kubernetes Custom Resource Definition (CRD), is a significant finding. It aims to standardize the definition of GenAI applications as configuration, preventing "GenAI sprawl" and allowing engineering teams to apply existing best practices and processes to LLM development.
- Vision RAG for Multimodal Information Extraction: An advanced form of RAG, Vision RAG, demonstrates the capability to process both text and images using multimodal embedding models and vision language models (VLMs). This innovation enables LLMs to extract information from complex document layouts where data might be embedded in images (e.g., tables as screenshots in PDFs), representing a powerful leap in information processing.
Technical Deep Dive
▶ Watch: Kubernetes advantages: Portability, consistency, scalability for LLMs (6:38)
The technical journey presented by Priya Samuel and Luke Marsden begins with the fundamental components required for LLM deployment on Kubernetes and progressively builds towards advanced, production-grade architectures.
At the most basic level, the speakers introduce Olama, an open-source tool predominantly used for testing LLMs. Deployment on Kubernetes is facilitated via a Helm chart, which bundles Olama with Open Web UI. However, the talk quickly highlights Olama's inherent limitations for production. By default, Olama utilizes highly quantized models, often compressed to store only four bits of information per weight in the neural network. While this compression allows models to run on resource-constrained environments like laptops, it significantly compromises accuracy, leading to "garbled" or "random" answers and truncated context lengths. Furthermore, crucial performance optimizations, such as flash attention, are often disabled by default. Flash attention is a technique that dramatically reduces memory usage during LLM inference, specifically transforming memory requirements from quadratic to subquadratic with respect to context length. This default configuration in Olama makes it unsuitable for demanding production workloads.
To overcome these shortcomings, the talk introduces VLLM, an open-source inference engine and server specifically designed for high-performance LLM serving. VLLM addresses the issues of context length and performance, offering better defaults and optimizations for latency and tokens per second. The speakers demonstrate how VLLM, running with models like Mistral 7B, provides a more robust and performant foundation for GenAI applications.
A significant challenge in deploying LLMs in secure, air-gapped environments is the management of model weights. VLLM, by default, downloads these weights from external repositories like Hugging Face. To circumvent the need for internet access and enhance supply chain security, the team opted to bake the model weights directly into the Docker image alongside the VLLM inference server. This approach, however, revealed a performance bottleneck: Docker's default gzip level 8 compression algorithm significantly increased image startup times due to the large size of model weights. The solution involved patching the Docker build process to employ a custom, more efficient compression algorithm, ensuring faster deployment and model readiness.
As organizations scale their GenAI efforts, the need to serve multiple, diverse models (e.g., specialized image models, multilingual models, different fine-tuned versions) becomes apparent. Running each model on a dedicated GPU node pool is economically unsustainable. To address this, the speakers highlight the critical need for GPU memory sharing. This capability, potentially requiring the development of a custom Kubernetes scheduler, allows multiple models to share the same physical GPU resources, drastically improving utilization and reducing infrastructure costs. It also enables the simultaneous execution of both inference jobs and fine-tuning jobs on the same cluster, offering greater flexibility and efficiency for MLOps workflows.
Beyond raw model serving, the true value of LLMs is realized through sophisticated application layers:
- API Integrations: This multi-step process enables LLMs to interact with external business systems, exemplified by a Jira integration.
- Classifier Prompt: The initial step uses an LLM to classify whether a user's request requires an API call and, if so, which specific tool or API endpoint is relevant (e.g., "how many issues are there in the sprint?").
- API Request Constructor Prompt: If an API call is deemed necessary, the LLM is then prompted to construct the API request body, leveraging the user's query and the OpenAPI specification of the target API. The LLM's ability to interpret OpenAPI specs and generate valid requests is a powerful capability.
- Response Summarizer Prompt: After the system executes the API call, the LLM takes the raw API response and summarizes the key information back to the user in a natural, consumable format.
- Knowledge (Retrieval Augmented Generation - RAG): RAG is a crucial pattern for grounding LLMs in specific, up-to-date, and factual information, preventing hallucinations.
- Vector Database: The process begins by ingesting source documents (e.g., internal wikis, manuals) into a vector database (e.g., PG vector for PostgreSQL, or Vector Core). These documents are first chunked into smaller pieces, and then each piece is converted into a numerical vector (an embedding) that represents its semantic meaning.
- Similarity Search: When a user poses a question, their query is also embedded into the same vector space. A similarity search then identifies and retrieves the most relevant document chunks from the vector database.
- Contextualized Prompt: These retrieved documents are included as part of the LLM's prompt, effectively providing the model with relevant context before it generates a response. This ensures the LLM is "grounded in facts."
- Knowledge Refresh: The vector database can be regularly refreshed to incorporate new or updated information, maintaining the LLM's knowledge currency.
- Vision RAG: An advanced RAG pipeline that extends knowledge retrieval to multimodal content.
- This technique utilizes multimodal embedding models capable of embedding both images and text into a unified vector space.
- When a user query is received, the system can retrieve not only relevant text but also relevant images from the vector database.
- These retrieved images, along with the text, are then fed into a vision language model (VLM), an LLM specifically trained to understand both visual and textual input. This enables the VLM to extract information from complex visual elements within documents, such as tables embedded as images in PDFs.
Finally, the talk emphasizes the importance of evaluation (Evals) and CI/CD for GenAI applications. The speakers describe a CLI tool that uses an LLM itself as a "judge" to assess the quality of responses, providing pass/fail grades that can be integrated into CI/CD pipelines. This allows engineering teams to apply standard software development best practices to GenAI. This entire framework is intended to be formalized through the AI spec, a proposed Kubernetes CRD that defines GenAI applications as declarative configuration, promoting standardization and reducing fragmentation across development teams.
Demo / Proof of Concept
▶ Watch: First demo: Setting up Olama with Open Web UI (7:49)
The speakers presented several compelling demonstrations to illustrate the concepts and solutions discussed.
The first demo highlighted the limitations of Olama for production workloads. Luke Marsden showcased an example where an LLM was tasked with summarizing issues from a Jira API response. The API response body, even for a sprint with only 10 issues, was quite large. The Olama-hosted model, utilizing its default highly compressed configuration, failed to process the full context. The resulting answer was not only incorrect—claiming the root element was an array containing one object when it contained ten—but also "garbled," demonstrating its inability to handle a moderately complex context window. This clearly illustrated the truncated context length and performance issues inherent in Olama's default setup.
The second demo transitioned to VLLM, demonstrating its superior performance and context handling. Using a Kubernetes cluster with two A100 GPUs and running a VLLM instance serving FI-4 from Microsoft, Luke showed that when prompted, the model responded "very, very quickly." Crucially, when presented with the same Jira issue summarization task, the VLLM-powered system, configured with proper context length defaults, successfully summarized all 10 issues in the sprint, providing an accurate and complete response where Olama had failed. This validated VLLM as a more production-ready inference solution.
The most impressive demonstration was of Vision RAG, showcasing an advanced knowledge pipeline. This demo began by showing a standard text-only RAG pipeline attempting to answer a question about a financial paper concerning stock sentiment. The PDF document contained both highlightable text and tables that were embedded as images (likely screenshots of Excel). When asked about the "10-day sentiment lag for the energy sector" from one of these tables, the text-only RAG pipeline failed, stating that the exact value "isn't provided in the context." This was because the crucial data was visually present in an image, not as extractable text. The demo then revealed a "vision toggle" in the system. With Vision RAG enabled, the system immediately provided the correct answer: "The sentiment lag for the energy sector is 0.35." A subsequent check of the original PDF confirmed this value directly from the image-based table. This demonstration powerfully illustrated the capability of multimodal embedding models and vision language models to extract information from complex visual elements within documents, unlocking knowledge previously inaccessible to text-only LLMs.
Defensive Implications
▶ Watch: Demo: Initial challenge with LLM API integration (8:35)
The insights shared by Priya Samuel and Luke Marsden offer critical defensive implications for organizations looking to securely and reliably deploy LLMs.
Firstly, the emphasis on self-hosting LLMs on Kubernetes directly addresses paramount concerns around data security and compliance. For entities in heavily regulated sectors like healthcare or European telecommunications, the ability to control data residency and processing locations is non-negotiable. By running LLMs within their own infrastructure—whether on-premises or in a private cloud—organizations can ensure sensitive data never leaves their trusted boundaries, mitigating risks associated with third-party proprietary platforms and simplifying compliance audits.
Secondly, the strategy of baking model weights into Docker images and optimizing their compression is a significant step towards supply chain security. This approach reduces external dependencies during runtime, as the LLM no longer needs to download weights from public repositories like Hugging Face. This is particularly crucial for air-gapped environments or those with strict network egress policies. It also provides greater control over model versions and ensures that the deployed model is exactly what was intended, preventing potential tampering or unexpected changes from external sources.
Thirdly, choosing VLLM for inference and focusing on GPU memory sharing contributes to operational resilience and cost optimization. By hosting LLMs internally, organizations gain direct control over uptime, latency, and API response stability, reducing reliance on the fluctuating performance of external services. Efficient GPU utilization through techniques like custom schedulers or advanced memory management allows for running multiple models and mixing workloads (inference, fine-tuning) on shared hardware, making advanced GenAI capabilities more economically viable and resilient against single points of failure.
Fourthly, the proposed AI spec and the integration of evaluations (evals) into CI/CD pipelines are foundational for governance and risk management in GenAI development. By defining GenAI applications as declarative configuration (Kubernetes CRDs), organizations can enforce standardized deployment patterns, preventing "GenAI sprawl" where different teams build disparate, potentially insecure, solutions. Integrating LLM-as-a-judge evals into CI/CD pipelines provides an automated mechanism to continuously validate model output quality and detect regressions or undesirable behaviors early, thereby reducing the risk of deploying unreliable or biased GenAI applications.
Finally, the advancements in RAG and Vision RAG directly enhance the factual grounding and trustworthiness of LLM outputs. By connecting LLMs to authoritative internal knowledge bases and enabling them to process multimodal information, organizations can significantly reduce the incidence of hallucinations and improve the accuracy of responses. This is a critical defensive measure against generating misleading or incorrect information that could have severe business consequences. Defenders should consider the entire attack surface of their Kubernetes-hosted LLM deployments, including API endpoints, access controls to model weights and vector databases, and the security of integrated external systems.
Key Takeaways
- Strategic Self-Hosting is Key: Running LLMs on Kubernetes offers robust solutions for data security, compliance, and cost-effectiveness, especially for organizations with sensitive data or high-scale requirements, moving away from reliance on proprietary hosted platforms.
- Production Requires Specialized Tools: While Olama is useful for testing, production environments demand high-performance inference engines like VLLM, which provide better context handling, latency, and overall reliability compared to default highly quantized models.
- Secure Model Management is Crucial: For air-gapped or secure deployments, baking model weights directly into Docker images is essential, requiring custom compression techniques to overcome performance bottlenecks associated with large image sizes.
- Integrations Drive Business Value: The true power of LLMs is unlocked through sophisticated application layers, including multi-step API integrations with business systems and robust Retrieval Augmented Generation (RAG) pipelines to ground models in specific, factual knowledge.
- Standardization and Evaluation are Non-Negotiable: Implementing a Kubernetes CRD like the proposed AI spec for GenAI applications, coupled with LLM-as-a-judge evaluation frameworks integrated into CI/CD, is vital for maintaining quality, preventing sprawl, and ensuring reliable, iterative development.
- Multimodal Capabilities Expand LLM Utility: Advanced techniques like Vision RAG, which leverage multimodal embedding models and vision language models, enable LLMs to extract valuable information from complex documents containing both text and images, significantly broadening their application scope.
About the Speaker(s)
Priya Samuel is a technology architect with an extensive background in DevOps and consulting, having worked across various small and large businesses. Her current focus involves building identity and access layers on top of generative AI applications. She brings considerable experience in designing and automating machine learning pipelines, making her a knowledgeable guide in the evolving GenAI landscape.
Luke Marsden is a seasoned professional in the DevOps space, having spent his entire career in the field. He was an early pioneer in Docker and Kubernetes, contributing to projects like Kubadm as part of SIG cluster life cycle. After founding startups focused on storage for Docker/Kubernetes and later an end-to-end MLOps platform, he transitioned to consulting. During the "ChatGPT moment," Luke observed firsthand the immense opportunity to leverage rapidly improving open-source models within Kubernetes clusters, driven by the need for enhanced security and reliability for his clients.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk delivers a brutally honest and deeply technical roadmap for deploying production-ready LLMs on Kubernetes. It dissects common pitfalls, from Olama's limitations to Docker's compression woes with massive model weights, offering concrete, hard-won solutions like VLLM optimization, custom Docker patching, and strategic GPU memory sharing. The speakers don't just stop at deployment; they detail sophisticated application layers with multi-step API integrations, advanced RAG, and even multimodal Vision RAG, all within a robust CI/CD and proposed AI spec framework. This isn't just theory; it's a battle-tested guide for securing, scaling, and operationalizing open-source LLMs that will…
Heather Calloway (CISO) — MUST SEE
This talk profoundly addresses critical CISO concerns regarding AI adoption, security, and compliance. By detailing a robust approach to self-hosting open-source LLMs on Kubernetes, the speakers provide a clear path for organizations to maintain control over sensitive data, manage supply chain risks for models, and achieve cost efficiencies. The proposed "AI spec" for standardization and the integration of evaluations into CI/CD pipelines are particularly impactful for establishing strong governance and accountability frameworks around GenAI deployments, directly informing executive decisions on institutional risk.