The Future of Data on Kubernetes... Rob Strechay, Nimisha Mehta, Gabriele Bartolini & Brian Kaufman
Rob Strechay, Nimisha Mehta, Gabriele Bartolini, Brian Kaufman
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk, presented as a panel discussion at KubeCon EU, delves into the rapidly evolving landscape of data management on Kubernetes, specifically focusing on its critical role in fueling Artificial Intelligence (AI) and Machine Learning (ML) workloads. Moderated by Rob Strechay, a managing director and principal analyst, the discussion features insights from industry experts Nimisha Mehta (Confluent), Gabriele Bartolini (EDB), and Brian Kaufman (Google GKE). The panel explores the practical use cases, architectural considerations, and future trajectory of running stateful applications and data pipelines on Kubernetes, emphasizing the shift from traditional database deployments to advanced AI foundations.

Key moments
- 0:00 Talk Introduction and Panelist Introductions
- 3:00 Data on Kubernetes (DoK) Community & Report Insights
- 3:40 DoK Production Workloads & AI/ML Acceleration
- 5:30 Brian Kaufman: Local SSDs for AI/ML Caching
- 6:40 Gabriele Bartolini: Postgres on Kubernetes for AI
The Future of Data on Kubernetes: From Database Management to AI Foundations
Speakers: Rob Strechay, Managing Director and Principal Analyst; Nimisha Mehta, Software Engineer at Confluent; Gabriele Bartolini, VP and Chief Architect of Kubernetes at EDB; Brian Kaufman, Product Manager at Google GKE
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=CI4rws1H-aM
Overview
This talk, presented as a panel discussion at KubeCon EU, delves into the rapidly evolving landscape of data management on Kubernetes, specifically focusing on its critical role in fueling Artificial Intelligence (AI) and Machine Learning (ML) workloads. Moderated by Rob Strechay, a managing director and principal analyst, the discussion features insights from industry experts Nimisha Mehta (Confluent), Gabriele Bartolini (EDB), and Brian Kaufman (Google GKE). The panel explores the practical use cases, architectural considerations, and future trajectory of running stateful applications and data pipelines on Kubernetes, emphasizing the shift from traditional database deployments to advanced AI foundations.
The session underscores the increasing maturity of Data on Kubernetes (DoK), highlighting that over 50% of these workloads are now in production environments. This significant adoption signals a crucial turning point, moving beyond initial skepticism about deploying stateful applications like databases on an orchestrator initially designed for stateless microservices. The speakers collectively argue that Kubernetes is not just "okay" for data, but in many respects, the optimal platform, particularly for organizations seeking to maximize the utilization of expensive assets like GPUs for AI/ML acceleration. The discussion navigates the challenges and opportunities, offering a comprehensive look at how companies are leveraging Kubernetes to build resilient, scalable, and cost-effective data infrastructures that are essential for the next generation of intelligent applications.
Background
▶ Watch: Talk Introduction and Panelist Introductions (0:00)
The journey of running data workloads on Kubernetes has been marked by a gradual but definitive shift in perception and capability. Early in Kubernetes' development, the idea of deploying stateful applications, especially databases, was often met with skepticism, with many considering it an anti-pattern. Gabriele Bartolini recounts attending his first KubeCon five and a half years prior, where the concept of using local storage for databases on Kubernetes was met with laughter. This initial resistance stemmed from concerns about data persistence, state management, and the perceived complexity of integrating traditional database requirements with Kubernetes' ephemeral, distributed nature.
However, significant advancements in Kubernetes itself, particularly with features like Persistent Volumes (PVs), Persistent Volume Claims (PVCs), and the maturation of Container Storage Interface (CSI) drivers, have fundamentally changed this landscape. The emergence of specialized operators has further bridged the gap, encapsulating complex operational knowledge for specific data services (like PostgreSQL or Kafka) into automated, Kubernetes-native constructs. This evolution has transformed Kubernetes into a viable, and increasingly preferred, platform for stateful workloads. The Data on Kubernetes (DoK) community has been instrumental in advocating for and documenting this shift, with recent reports indicating that over 50% of DoK workloads are now in production, and 75% of those production environments are considered advanced. This widespread adoption is largely fueled by the exponential growth of AI and ML, which demand highly agile, scalable, and performant data infrastructures. The challenge now lies in not just running data, but optimizing it for the unique demands of AI/ML, such as rapid data ingestion, low-latency access for inference, and efficient management of large models and datasets.
Key Findings
▶ Watch: Data on Kubernetes (DoK) Community & Report Insights (3:00)
The panel discussion illuminated several key findings regarding the current state and future direction of data on Kubernetes, particularly in the context of AI and ML:
- DoK Maturity and Production Readiness: The Data on Kubernetes (DoK) ecosystem has reached a significant level of maturity, with over 50% of DoK workloads now operating in production environments. This statistic directly refutes earlier skepticism about running stateful applications on Kubernetes and highlights the platform's growing reliability for critical data services.
- AI/ML as a Primary Driver: AI and ML workloads are identified as a major catalyst for DoK adoption. The need for rapid iteration, efficient data processing, and optimized resource utilization (especially for expensive GPUs) makes Kubernetes an attractive platform for these advanced analytical tasks. The panel stressed the importance of keeping GPUs busy, which necessitates robust data architectures.
- Database Dominance: Databases continue to be the number one category of data on Kubernetes. This is attributed to the effectiveness of operators in simplifying the operational complexity of databases within a declarative Kubernetes environment, enabling features like high availability, backup/restore, and scaling to be managed natively.
- Storage Acceleration is Critical for AI/ML: For AI/ML workloads, particularly RAG (Retrieval-Augmented Generation) and inference, ultra-low latency data access is paramount. This has led to a significant trend in using local SSDs and RAM disks for caching, bringing data as close as possible to the GPUs. Strategies like parallel downloads for training data and specialized model streamers for inference models are becoming essential.
- Streaming Data for Real-time Intelligence: The panel emphasized the increasing trend of embedding intelligence directly into data processing pipelines, especially for streaming data. Real-time event-driven processing, facilitated by technologies like Apache Kafka and Apache Flink, allows for immediate decision-making and orchestration of complex systems, such as multi-agent architectures, where streaming data acts as the "brain."
- Future Evolution: Integration and Infrastructure Bundling: Looking ahead, the experts foresee deeper integration among CNCF projects and a broader use of the Kubernetes control plane to bundle infrastructure and applications as logical entities. This includes managing external resources (like object storage configurations or network layers in hyperscalers) alongside in-cluster deployments, simplifying the complex configurations required for advanced AI/ML setups.
- Cost Control through Observability and Smart Storage: Managing costs, especially with expensive GPU resources, is a major concern. The panel highlighted that robust observability is key to identifying and eliminating GPU idle time. Strategic use of storage tiers, such as prioritizing cheaper local SSDs over expensive memory where latency tolerances allow, and minimizing cross-availability zone (cross-AZ) data transfers, are crucial for optimizing expenditure.
Technical Deep Dive
▶ Watch: DoK Production Workloads & AI/ML Acceleration (3:40)
The panel provided a comprehensive technical exploration of how Kubernetes is becoming the backbone for advanced data workloads, particularly those driving AI and ML.
At the core of running stateful applications, especially databases, on Kubernetes is the concept of declarative configuration. As Gabriele Bartolini explained, Kubernetes enables a "declarative world" where desired states are defined, and operators work to guarantee those states. This is particularly crucial for databases, simplifying intrinsic operational complexities such as provisioning, scaling, backup, and recovery. The success of databases as the number one category in Data on Kubernetes (DoK) is largely attributed to the maturity of operators that abstract away these complexities, allowing database administrators to leverage Kubernetes without needing deep expertise in the orchestrator, and vice-versa for Kubernetes engineers. The close relationship between PostgreSQL and its storage, for instance, is a critical area of innovation, with CSI drivers playing a vital role in providing direct and efficient access to underlying storage resources. EDB, a primary contributor to PostgreSQL, has pioneered projects like Cloud Native PG, which recently joined the CNCF sandbox, demonstrating a commitment to cloud-native database operations. Innovations, such as an upcoming patch in PostgreSQL 18 that improves extension deployment via image volumes, highlight how running databases on Kubernetes is driving upstream changes in database technology itself. The pg_vector extension, which transforms PostgreSQL into a vector database, further showcases its utility for AI applications, specifically for RAG patterns.
For AI/ML workloads, the technical considerations diverge significantly between training and inference. Brian Kaufman from Google GKE elaborated on these differences:
- AI/ML Training: Training often involves processing a multitude of small files. The primary challenge is efficiently downloading and accessing this training data, typically stored in object storage. The trend here is towards extensive caching and parallel downloads to ensure data is readily available for multiple epochs, reducing the need to repeatedly fetch data, which can be a significant bottleneck.
- AI/ML Inference: Inference, unlike training, scales up and down dynamically, much like a typical web service. The critical requirement is ultra-low latency, as inference requests demand immediate responses. Inference workloads often deal with very large models and weights; for example, a Llama 70B model can be 13 gigabytes. Downloading such a model from sources like Hugging Face can take upwards of 20 minutes, leading to costly GPU idle time. To combat this, solutions include:
- Storage Acceleration: Similar to training, parallel downloads from object storage are employed.
- VLLM model streamer: A new feature in VLLM (a popular inference engine) and tools like GCS fuse facilitate direct, parallel downloads to CPU memory and then to GPU memory, significantly accelerating model loading.
- Block Storage for Models: Storing models on block storage offers advantages, particularly the ability to make the disk immutable and ingest it rapidly from a location within the same availability zone, offering better control and potentially lower latency than object storage for specific use cases.
Nimisha Mehta from Confluent highlighted the technical patterns emerging in streaming data for AI:
- Real-time Processing with Streaming: Data is most valuable when fresh. For real-time applications like inventory management, supply chain optimization, and anomaly detection, event-driven processing is crucial.
- Intelligence in Pipelines: A key architectural pattern involves embedding AI intelligence directly into data processing pipelines, specifically in pre-processing and post-processing stages.
- Kafka and Flink for LLM Integration: Apache Kafka serves as the robust streaming backbone, delivering data between sources and sinks. Apache Flink, a stream processing engine, can be hooked up to an LLM directly within the processing pipeline. This allows for real-time inference and intelligent decision-making based on incoming data before it reaches its final destination.
- Multi-Agent Orchestration: In complex multi-agent systems, streaming data acts as the "brain," orchestrating which agent should act based on real-time data inputs. Kafka provides the backbone for data delivery, while Flink, integrated with an LLM, can determine which agent to invoke based on the data stream.
Looking to the future, the panel discussed the evolution of the CNCF ecosystem. Gabriele Bartolini emphasized the need for deeper integration among projects, encouraging collaboration among projects that use PostgreSQL as a backend to improve documentation and user experience. Brian Kaufman envisioned the Kubernetes control plane being used to bundle infrastructure and applications as a single logical entity. Projects like Crossplane or KCL (KubeConfig Language, referred to as KRO in the transcript) exemplify this, allowing the management of external infrastructure configurations (e.g., hyperscaler settings, load balancers, network layers, secrets) alongside in-cluster Kubernetes artifacts, thus simplifying the deployment and management of complex AI/ML environments.
Demo / Proof of Concept
▶ Watch: Brian Kaufman: Local SSDs for AI/ML Caching (5:30)
This session was presented as a panel discussion featuring industry experts sharing their insights and experiences. As such, it did not include a live technical demonstration or a specific proof of concept. The discussion focused on current trends, customer stories, and architectural patterns observed in the field.
Defensive Implications
▶ Watch: Gabriele Bartolini: Postgres on Kubernetes for AI (6:40)
The detailed technical insights shared by the panel offer critical implications for defenders and platform engineers tasked with securing and optimizing data workloads on Kubernetes, particularly in the context of AI/ML.
- Robust Observability for Cost Control and Performance: The panel consistently highlighted observability as the number one tool for managing costs and ensuring efficiency. For expensive resources like GPUs, defenders must implement comprehensive monitoring to identify and eliminate GPU idle time. This means having visibility into data loading times, inference engine performance, and overall resource utilization to ensure that every computational cycle is maximized. Proactive monitoring can detect bottlenecks in data pipelines or inefficient model serving, preventing unnecessary expenditure.
- Strategic Storage Tiering and Data Locality: Brian Kaufman emphasized the significant cost difference between memory and local SSDs (e.g., $3/GB/month for memory versus $0.08/GB/month for local SSDs at Google). Defenders should design architectures that intelligently leverage cheaper local SSDs for caching data where latency tolerances permit, reserving expensive memory for only the most critical, ultra-low-latency operations. Furthermore, minimizing cross-availability zone (cross-AZ) transfers is crucial for both cost and performance. This can be achieved through careful collocation of pods and their associated data processing components using Kubernetes affinities and labels, ensuring data resides within the same availability zone as its consumers.
- Database Isolation and Cost Predictability: Gabriele Bartolini recommended isolating PostgreSQL worker nodes from the rest of the cluster, often on bare metal machines with local storage. This provides a fixed, predictable cost base for the database layer. Defenders can implement this using Kubernetes taints, tolerations, and anti-affinities to ensure database pods run on dedicated hardware, preventing noisy neighbor issues and offering better resource isolation and cost predictability. Planning for this infrastructure from day zero is key to realizing significant long-term savings.
- Comprehensive Resource Tagging and Auditing: Nimisha Mehta stressed the importance of robust resource tagging policies for all cloud resources. Defenders should ensure that every provisioned resource is tagged with information such as the team responsible, purpose, and provisioning date. This enables accurate cost attribution, simplifies auditing, and supports compliance efforts by providing clear visibility into resource ownership and usage.
- Addressing Skill Gaps: A recurring challenge identified was the knowledge gap between database experts (who may not know Kubernetes) and Kubernetes experts (who may not know databases). Defenders need to invest in cross-training or adopt solutions like mature operators that abstract away underlying complexities. This ensures that critical data services are managed by personnel with the right expertise, reducing misconfigurations and security vulnerabilities.
- Securing the AI/ML Data Pipeline: Given the emphasis on moving intelligence into data processing pipelines (e.g., Flink hooked to LLMs), defenders must secure every stage of the pipeline. This includes securing Kafka as the streaming backbone, ensuring data encryption in transit and at rest, and implementing strict access controls for LLM interactions. The use of image volumes for deploying PostgreSQL extensions also requires careful vetting of extension sources to prevent the introduction of malicious code.
- Infrastructure as Code for External Dependencies: Brian Kaufman's vision of bundling infrastructure with applications via the Kubernetes control plane (using tools like Crossplane or KCL) has significant defensive implications. By managing external cloud resources (e.g., object storage, network configurations, load balancers) declaratively alongside Kubernetes deployments, defenders can apply consistent security policies, version control configurations, and automate compliance checks across the entire stack, reducing configuration drift and manual errors.
Key Takeaways
- Data on Kubernetes (DoK) is Mature and Production-Ready: Over 50% of DoK workloads are now in production, with databases leading the adoption, demonstrating Kubernetes' capability to host critical stateful applications effectively.
- AI/ML Drives Specific Data Architecture Needs: AI and ML workloads necessitate ultra-low latency data access, extensive caching (e.g., local SSDs, RAM disks), parallel data downloads, and specialized model loading techniques (e.g., VLLM model streamer) to maximize expensive GPU utilization.
- Declarative Operations and Operators are Essential: Kubernetes' declarative nature, combined with mature operators for databases like PostgreSQL (e.g., Cloud Native PG, pg_vector), simplifies complex operational tasks and bridges the skill gap between database administrators and Kubernetes engineers.
- Streaming Data Fuels Real-time AI: Technologies like Apache Kafka and Apache Flink are crucial for building real-time, event-driven AI pipelines, enabling in-pipeline inference with LLMs and orchestrating multi-agent systems by acting as the central "brain."
- Future Focus on Integration and Bundling: The CNCF ecosystem is evolving towards deeper project integration and leveraging the Kubernetes control plane to bundle external infrastructure (e.g., hyperscaler services) with applications, using tools like Crossplane to manage complex AI/ML environments holistically.
- Cost Optimization Requires Observability and Smart Storage: Controlling costs, especially for GPU resources, demands robust observability to eliminate idle time, strategic use of cheaper local SSDs for caching, and minimizing cross-availability zone (cross-AZ) data transfers through careful resource collocation.
About the Speaker(s)
Rob Strechay served as the moderator for this insightful panel. He is a Managing Director and Principal Analyst, actively involved as a member of the Data on Kubernetes (DoK) community. His expertise lies in guiding discussions and extracting valuable information from industry leaders on emerging trends in data management.
Nimisha Mehta is a Software Engineer at Confluent, a company specializing in streaming services like Kafka and Flink. She is primarily involved with the Kubernetes platform team, focusing on managing internal data infrastructure and Kubernetes-related aspects, including providing self-service capabilities to developers and maintaining system software on clusters for Confluent's on-prem and cloud offerings.
Gabriele Bartolini holds the position of Vice President and Chief Architect of Kubernetes at EDB. EDB is recognized as the number one contributor to the open-source PostgreSQL project and the creator of Cloud Native PG, a project that recently joined the CNCF sandbox. Bartolini is also a PostgreSQL contributor and a Data on Kubernetes ambassador, advocating for the adoption of databases on Kubernetes.
Brian Kaufman is a Product Manager at Google, specifically working on the GKE data layer. His focus areas include stateful applications such as databases on GKE, and AI/ML workloads, particularly inference. Prior to joining Google seven years ago, he worked at Docker, where he was involved with Docker Swarm, and this KubeCon was his first experience with the conference.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This panel discussion on Data on Kubernetes for AI/ML workloads is a substantive and highly practical session. It cuts through the typical hype, focusing on concrete architectural patterns, performance optimizations, and critical defensive implications for running stateful applications and AI pipelines on Kubernetes. The speakers, all credible experts, deliver actionable insights on managing costs, optimizing storage for GPUs, and integrating streaming data with LLMs, making it a valuable resource for anyone building or securing cloud-native data platforms.
Heather Calloway (CISO) — STRONG ACCEPT
This panel provides a critical update on the maturity of Data on Kubernetes, particularly as it intersects with AI/ML workloads. While presented at a technical conference, the insights into production readiness, cost optimization for GPUs, and the defensive implications for securing complex data pipelines make this highly relevant for security leaders navigating cloud-native and AI transformations. It offers concrete, actionable guidance for managing risk and ensuring institutional accountability in these evolving environments.