Cloud Native Storage and Data: The CNCF Storage TAG Projects, Te... Raffaele Spazzoli & Alex Chircop

Raffaele Spazzoli, Alex Chircop

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This session, presented by Alex Chircop, Jing, and Raffaele Spazzoli, offers a comprehensive exploration of cloud-native storage within the Cloud Native Computing Foundation (CNCF) ecosystem. It delves into the critical role of data and stateful applications in Kubernetes, challenging the common misconception of cloud-native as purely stateless. The speakers, all prominent figures within the CNCF's Technical Advisory Group (TAG) Storage, outlined the TAG's ongoing initiatives, including its restructuring, and highlighted key whitepapers addressing cloud-native storage attributes, running data workloads on Kubernetes, performance benchmarking, and disaster recovery.

Watch on YouTube

Visual summary for Cloud Native Storage and Data: The CNCF Storage TAG Projects, Te... Raffaele Spazzoli & Alex Chircop by Raffaele Spazzoli, Alex Chircop
Visual summary for Cloud Native Storage and Data: The CNCF Storage TAG Projects, Te... Raffaele Spazzoli & Alex Chircop by Raffaele Spazzoli, Alex Chircop

Key moments

  1. 0:00 Introduction and CNCF Tag Reboot Announcement
  2. 2:00 New CNCF Tag Structure and Scaling Rationale
  3. 4:00 Debunking Stateless Myth: Why Cloud Native Storage Matters
  4. 4:53 Overview of Key Graduated CNCF Storage Projects
  5. 6:00 Understanding CNCF Project Maturity Stages (Sandbox, Incubation, Graduation)
  6. 7:50 Introduction to the CNCF Storage White Paper

Cloud Native Storage and Data: The CNCF Storage TAG Projects, Technologies, and Whitepapers

Speakers: Alex Chircop, Chief Architect, Akamai, CNCF TOC; Jing (Shinyang), VMware by Borang, Co-chair of Tech Storage; Raffaele Spazzoli, Consultant, Red Hat, Co-chair of Tech Storage

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=0GNjonLfCQA

Overview

This session, presented by Alex Chircop, Jing, and Raffaele Spazzoli, offers a comprehensive exploration of cloud-native storage within the Cloud Native Computing Foundation (CNCF) ecosystem. It delves into the critical role of data and stateful applications in Kubernetes, challenging the common misconception of cloud-native as purely stateless. The speakers, all prominent figures within the CNCF's Technical Advisory Group (TAG) Storage, outlined the TAG's ongoing initiatives, including its restructuring, and highlighted key whitepapers addressing cloud-native storage attributes, running data workloads on Kubernetes, performance benchmarking, and disaster recovery.

The talk is particularly relevant for architects, developers, and operations teams grappling with the complexities of managing persistent data in dynamic Kubernetes environments. It provides essential guidance on leveraging cloud-native principles for stateful applications, navigating the vast landscape of CNCF storage projects, and making informed decisions about data management strategies. By detailing the evolution of the TAGs and the robust body of work produced, the speakers underscore the CNCF's commitment to fostering a mature and resilient ecosystem for data-intensive cloud-native workloads.

Why this matters is multifaceted: the exponential growth of CNCF projects necessitates a more streamlined governance structure, the pervasive need for data storage in virtually every application demands cloud-native solutions, and the inherent challenges of distributed systems require specialized approaches to performance, observability, and disaster recovery. This session serves as a crucial update and a valuable resource for anyone building or operating stateful applications in the cloud-native era, providing both high-level strategic insights and detailed technical considerations.

Background

▶ Watch: Introduction and CNCF Tag Reboot Announcement (0:00)

The CNCF ecosystem has experienced explosive growth since the Technical Advisory Groups (TAGs) were first conceived in Barcelona in 2019. The number of projects under the CNCF umbrella has swelled from approximately 40 to a staggering 215, with an additional 33 projects added in just the last year. This rapid expansion necessitated a significant restructuring of the TAG system to ensure the organization can continue to evaluate, maintain, and sustain the health of its projects effectively for the next decade. The former eight TAGs are being consolidated into five, with a new structure incorporating sub-projects (for long-running tasks like security assessments and project reviews) and initiatives (for specific, often shorter-term, goals). Community groups will also be formalized to foster innovation and contribute ideas.

Central to the TAG Storage's mission is the assertion that "there is no such thing as a stateless architecture." While cloud-native is often associated with statelessness, every application ultimately requires data persistence. The core thesis is that all the benefits of Kubernetes – automation, declarative configuration, performance, scalability, and failover capabilities – should extend to stateful workloads. This involves leveraging a broad ecosystem, including Container Storage Interface (CSI) for integrating block and object storage, and Container Object Storage Interface (COSI). The cloud-native environment also boasts a rich and mature landscape of operators that automate complex database topologies, message queues, and other data services, complete with automated replication and failover.

The CNCF already hosts a number of significant storage projects at various maturity levels. Graduated projects, which have demonstrated widespread production adoption, strong governance, and undergone security audits, include Rook (automates Ceph environments for block, object, and file storage), Vitess (a scale-out MySQL database operator), Harbor (a container registry), etcd (the distributed key-value store fundamental to Kubernetes), TiKV (a highly scalable key-value store), and KubeFS (a shared file system particularly useful for machine learning and AI workloads, which recently graduated). Beyond these, a multitude of incubating and sandbox projects are actively developing new solutions, reflecting the dynamic innovation in cloud-native storage. Sandbox projects focus on community building, IP development, and experimentation; incubating projects require rigorous due diligence and evidence of successful production use; and graduated projects represent the pinnacle of maturity and operational excellence within the CNCF.

Key Findings

▶ Watch: Debunking Stateless Myth: Why Cloud Native Storage Matters (4:00)

The talk highlighted several critical findings and contributions from the CNCF TAG Storage's work, encapsulated in their various whitepapers and community surveys:

  1. Surge in Data Workloads on Kubernetes: The Data on Kubernetes (DoK) community survey consistently shows a significant trend: more and more stateful workloads are migrating to Kubernetes. Databases have maintained their position as the most commonly deployed workload type from 2021 to 2024, followed by analytics (moving to #2 in 2024) and streaming/messaging (now at #3). This trend underscores Kubernetes' growing reliability as a platform for critical data workloads and its emerging role as a foundational infrastructure for AI and Machine Learning (AI/ML) applications.
  1. Attributes for Cloud-Native Data Systems: The whitepapers introduced and emphasized crucial attributes for designing and evaluating cloud-native data systems. Building upon traditional storage attributes, the DoK white paper specifically added observability and elasticity as paramount for distributed microservices environments. Observability is vital for early detection and troubleshooting of failures across numerous components, while elasticity enables rapid scaling (up and down) and efficient resource utilization, including storage tiering based on access frequency.
  1. Operator-Centric Data Management: Running data inside Kubernetes is overwhelmingly facilitated by operators. These leverage Kubernetes' declarative APIs to reconcile the actual state of a data service against a desired state, automating complex day-2 operations such as backup, restore, migration, and upgrades. The Operator Hub lists over 300 operators, including more than 50 for databases, demonstrating their widespread adoption. However, the lack of standardization across operators remains a challenge, prompting community efforts to create a vendor-neutral feature matrix.
  1. Comprehensive Disaster Recovery Archetypes: A key contribution is the identification and detailed comparison of four primary disaster recovery (DR) archetypes for cloud-native environments: backup and restore, volume replication, transaction replication, and active-active (distributed workload). The latest version of the DR white paper provides balanced coverage for each, highlighting their distinct RTO (Recovery Time Objective), RPO (Recovery Point Objective), ownership (infrastructure vs. developer teams), and underlying dependencies (storage capabilities vs. networking/middleware). The active-active approach, allowing workloads to span geographically distributed failure domains, is presented as the most cloud-native due to its inherent resilience.
  1. Performance Benchmarking Pitfalls: The performance benchmarking white paper serves as a crucial guide, emphasizing the numerous ways benchmarking can go wrong. A primary finding is the critical impact of caching at multiple layers of the storage stack, which can dramatically skew results and lead to overestimations of true storage performance. The paper strongly advises against relying solely on vendor-provided benchmarks and stresses the importance of running custom tests in one's own environment with specific application workloads to obtain meaningful comparisons.

Technical Deep Dive

▶ Watch: Overview of Key Graduated CNCF Storage Projects (4:53)

The technical depth of the talk was primarily conveyed through the discussion of various whitepapers and the underlying concepts they address.

The CNCF Storage White Paper, now in its third version, provides a foundational understanding of storage in a cloud-native context. It meticulously covers the attributes of a storage system, guiding users on how to map application needs to appropriate storage solutions. It acknowledges that storage encompasses not just traditional block and file systems but also object storage, key-value stores, and databases, with the latest iteration adding messaging and streaming technologies. A crucial aspect is understanding the various layers of storage virtualization; for instance, a file system might be built atop object storage, inheriting some of its characteristics. This multi-layered architecture significantly impacts performance and behavior.

Building on this, the Data on Kubernetes (DoK) white paper focuses specifically on patterns for running data workloads within Kubernetes. It delves into the advantages of Kubernetes' self-healing, agile deployment, scalability, and portability for stateful applications. Beyond the attributes outlined in the general storage paper, it introduces observability and elasticity. Observability is critical in distributed microservices where identifying the root cause of a failure can be complex; robust logging, metrics, and tracing systems are essential. Elasticity refers to the ability to scale resources on demand, including the dynamic provisioning and de-provisioning of storage, and storage tiering, where data can be moved between different storage types based on access frequency and cost. The paper also contrasts running data inside versus outside Kubernetes, advocating for the former when facilitated by operators.

Kubernetes Operators are a cornerstone of managing complex data services on Kubernetes. An operator is an application-specific controller that extends the Kubernetes API to create, configure, and manage instances of complex applications. It uses Custom Resources (CRs) to define the desired state of a database cluster (e.g., number of replicas, storage type). The operator then continuously reconciles the actual state of the cluster with this desired state. This automation covers day-2 operations like backup, restore, migration, and upgrades, significantly reducing operational overhead. The talk highlighted that an organization typically uses multiple operators, with over 300 available in the Operator Hub, including 50+ for databases, such as etcd and Vitess (graduated CNCF projects), and Cloud Native PostgreSQL (a sandbox CNCF project). The Container Storage Interface (CSI) is fundamental to this, providing a standardized way for storage vendors to expose their storage systems to container orchestrators like Kubernetes.

Other common patterns for running data on Kubernetes include:

  • StatefulSets: A workload API designed for stateful applications, providing stable, unique network identifiers, stable persistent storage, and ordered, graceful deployment and scaling.
  • Topology-aware scheduling: By applying node labels (e.g., for failure domains), the Kubernetes scheduler can ensure pods are spread across different availability zones or racks, increasing resilience. Topology-aware dynamic provisioning extends this to persistent volumes, ensuring they are provisioned in the same failure domain as their respective pods.
  • Pod Disruption Budget (PDB): This feature allows workload owners to specify the minimum number of available instances for an application during planned maintenance events, preventing service degradation.
  • Secure Networking: Kubernetes offers a secure networking model by default, where pods are not externally accessible unless explicitly configured, protecting applications from unauthorized access.
  • Observability Tools: Integration with tools like Prometheus for metrics, Grafana for visualization, and various logging and tracing solutions is crucial for monitoring and troubleshooting distributed data workloads.

The Disaster Recovery (DR) white paper, in its second version, provides a detailed comparison of its four archetypes.

  1. Backup and Restore: The most traditional approach, relying on periodic data backups and subsequent restoration.
  2. Volume Replication: Involves replicating entire storage volumes, often at the storage layer. These first two approaches primarily rely on storage capabilities and are typically owned by infrastructure teams, with RTO/RPO dictated by backup frequency and recovery speed.
  3. Transaction Replication: The workload itself manages data replication (e.g., database replication). This is an active-passive approach.
  4. Active-Active: The most cloud-native archetype, where workloads are distributed across multiple geographically separate failure domains. The workload itself handles data replication and consistency, making the loss of a single data center appear as a high-availability event rather than a disaster. This relies heavily on networking capabilities (specifically east-west connectivity between clusters) and middleware, and is typically owned by developer teams. The paper delves into the concepts of replica and partition and how they relate to the CAP theorem (Consistency, Availability, Partition tolerance) in achieving distributed, strictly consistent workloads. It also presents a reference architecture for deploying active-active workloads on Kubernetes, requiring mechanisms to synchronize workloads across clusters (e.g., via a global load balancer in front of the clusters and ensuring routable east-west communication paths between internal Kubernetes SDNs).

Finally, the Performance Benchmarking White Paper offers a guide to accurate testing of cloud-native storage. It warns against common pitfalls, emphasizing the need to understand which layer of the storage system is being tested (e.g., client requirements for operations per second vs. throughput). Factors like data protection, reduction, and encryption services significantly impact overall performance. Latency is identified as a critical killer, accumulating with each layer in the stack. The paper's strongest warning is about caching, which can occur at multiple layers and drastically inflate perceived performance, leading to misleading results if not properly accounted for. It strongly advises against relying on vendor-provided benchmarks, advocating for custom tests in one's own environment with specific application workloads.

Demo / Proof of Concept

▶ Watch: Understanding CNCF Project Maturity Stages (Sandbox, Incubation, Graduation) (6:00)

The session focused on presenting conceptual frameworks, whitepapers, and strategic guidance rather than a live demonstration. No specific demo or proof of concept was presented during the talk.

Defensive Implications

▶ Watch: Introduction to the CNCF Storage White Paper (7:50)

The insights presented by the CNCF TAG Storage have profound implications for organizations securing and operating cloud-native environments, particularly those with stateful workloads. Defenders must shift their mindset from treating cloud-native as inherently stateless to recognizing and strategically managing the pervasive need for persistent data.

  1. Embrace and Secure Stateful Workloads: Organizations should fully embrace the reality that most applications require state. Instead of externalizing all data services, security teams must learn to leverage Kubernetes' capabilities for automating and securing stateful workloads. This involves understanding CSI and COSI integrations, ensuring that the underlying storage infrastructure adheres to security best practices, and implementing robust access controls for persistent volumes and object stores.
  1. Strategic Operator Adoption and Hardening: Operators are powerful tools for automating complex data services, but their widespread use necessitates a defensive strategy. Defenders should:
  • Vet Operators Thoroughly: Given the "lack of standard" identified by the DoK community, organizations must conduct rigorous security reviews of any operator before deployment. This includes examining the operator's codebase, its permissions (RBAC), and its dependencies.
  • Monitor Operator Behavior: Implement robust monitoring and logging for operators themselves, looking for anomalous behavior that could indicate compromise or misconfiguration.
  • Leverage Standardization Efforts: Keep an eye on the DoK community's work on a "operator feature matrix" to aid in selecting more secure and standardized operators.
  • Automate Day-2 Security: Operators automate backup, restore, and upgrades. Defenders should ensure these automated processes incorporate security best practices, such as encrypted backups, secure restore procedures, and timely application of security patches during upgrades.
  1. Implement Comprehensive Disaster Recovery Strategies: The detailed DR archetypes provide a roadmap for building resilient data platforms. Defenders should:
  • Assess RTO/RPO for Data: Clearly define RTO and RPO objectives for all stateful applications, as this will dictate the appropriate DR archetype.
  • Prioritize Cloud-Native DR: For critical workloads, strive towards transaction replication and active-active strategies where feasible. This requires a strong understanding of east-west networking between clusters and ensuring secure, low-latency communication paths.
  • Secure DR Components: All components of the DR strategy, including backup systems, replication channels, and global load balancers, must be secured against unauthorized access and tampering. This includes encryption in transit and at rest for replicated data.
  • Regularly Test DR: DR plans, regardless of the archetype, are only as good as their last test. Regular, simulated disaster events are crucial to validate RTO/RPO and identify any gaps in the recovery process.
  1. Enhance Observability for Security: The emphasis on observability is not just for operational troubleshooting but also for security. Defenders should:
  • Integrate Security Logging: Ensure that logs from data workloads, operators, and underlying storage systems are collected, aggregated, and fed into Security Information and Event Management (SIEM) systems.
  • Define Security Metrics: Establish specific metrics to monitor for security anomalies, such as unusual access patterns, data modification rates, or failed authentication attempts.
  • Implement Distributed Tracing: For complex microservice architectures, distributed tracing can help identify the blast radius of a security incident or pinpoint unauthorized data flows.
  1. Critical Approach to Performance Benchmarking: Performance often has security implications, as degraded performance can be a sign of a denial-of-service attack or resource exhaustion. Defenders should:
  • Understand Benchmarking Limitations: Be aware of the pitfalls, particularly the impact of caching, when evaluating storage performance.
  • Conduct Independent Security Performance Tests: Run security-focused performance tests to understand how the system behaves under load, during data encryption/decryption, or when security controls are active. This helps in capacity planning and identifying potential bottlenecks that could be exploited.

In essence, defending cloud-native stateful workloads requires a holistic approach that integrates security considerations into every layer of the data stack, from the underlying storage to the application-level operators and the overarching disaster recovery strategy.

Key Takeaways

  • Cloud-Native is Stateful: The notion of purely stateless cloud-native applications is a myth; virtually all applications require data persistence. Kubernetes provides a mature platform for managing stateful workloads, complete with automation, failover, and scaling capabilities.
  • CNCF Ecosystem Growth: The CNCF project landscape has grown exponentially, necessitating a restructuring of the TAGs to effectively manage and sustain this growth, ensuring continued innovation and project health.
  • Operators are Essential for Data on Kubernetes: Kubernetes operators are the primary mechanism for automating the deployment and day-2 operations of complex data services (like databases and message queues) within Kubernetes, leveraging declarative APIs.
  • Comprehensive Disaster Recovery: Organizations must choose from various DR archetypes (backup/restore, volume replication, transaction replication, active-active) based on RTO/RPO requirements, understanding that cloud-native active-active approaches offer the highest resilience but require robust networking and middleware capabilities.
  • Observability and Elasticity are Paramount: For distributed cloud-native data workloads, observability (logging, metrics, tracing) is crucial for detecting failures, and elasticity (dynamic scaling, storage tiering) is vital for efficient resource utilization.
  • Beware of Benchmarking Pitfalls: Accurate performance benchmarking of cloud-native storage is challenging. It's critical to understand the impact of various layers, services, and especially caching, and to conduct independent tests tailored to specific application workloads rather than relying solely on vendor-provided results.

About the Speaker(s)

  • Alex Chircop is the Chief Architect at Akamai and has recently been elected to the CNCF Technical Oversight Committee (TOC). He is a prominent voice in the cloud-native community, particularly on matters of storage and infrastructure.
  • Jing (Shinyang) works for VMware by Borang and serves as one of the co-chairs of the CNCF TAG Storage. Her work focuses on data workloads in Kubernetes and contributing to related whitepapers.
  • Raffaele Spazzoli is a consultant at Red Hat and also serves as a co-chair of the CNCF TAG Storage. He frequently interacts with customers on integration challenges and specializes in topics like disaster recovery for cloud-native environments.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This session by the CNCF TAG Storage leadership provides an indispensable deep dive into cloud-native storage and data management within Kubernetes. It masterfully debunks the "stateless" myth, offering comprehensive guidance through their whitepapers on critical attributes, operator-driven data management, and detailed disaster recovery archetypes. The insights into performance benchmarking pitfalls and the latest DoK survey data offer actionable intelligence for anyone building or defending stateful applications in this ecosystem, directly from the architects of the cloud-native storage landscape.

Heather Calloway (CISO) — STRONG ACCEPT

This session by the CNCF TAG Storage provides a vital and comprehensive overview of managing stateful data within cloud-native environments. It effectively debunks the myth of purely stateless architectures, highlighting the critical role of persistent data in nearly all applications. The deep dive into the TAG's whitepapers on disaster recovery archetypes, operator-centric data management, and performance benchmarking offers invaluable strategic insights for security leaders and architects. While it presents a survey of the ecosystem rather than a prescriptive security checklist, it clearly frames the institutional challenges and provides the necessary context for making informed…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025