Kubernetes Backup Legitimized: CSI Changed Block Tracking Has Arrived- M. Lavi, C. Braganza, X. Yang

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

The KubeCon EU talk "Kubernetes Backup Legitimized: CSI Changed Block Tracking Has Arrived" delivered by Shing Yang, Mark Lavi, and Carl Braganza, introduced a pivotal advancement in Kubernetes data protection: Changed Block Tracking (CBT). This feature, long considered a fundamental capability in traditional virtual machine (VM) and bare-metal backup solutions, has been a notable omission in the cloud-native Kubernetes ecosystem. Its absence forced backup vendors to either implement full backups, leading to excessive storage, network, and compute consumption, or rely on proprietary vendor-specific solutions that undermined the portability promise of Kubernetes.

Watch on YouTube

Visual summary for Kubernetes Backup Legitimized: CSI Changed Block Tracking Has Arrived- M. Lavi, C. Braganza, X. Yang
Visual summary for Kubernetes Backup Legitimized: CSI Changed Block Tracking Has Arrived- M. Lavi, C. Braganza, X. Yang

Key moments

  1. 0:00 Introduction and motivation for change block tracking
  2. 2:00 History and journey of CBT implementation (KEP 3314)
  3. 4:00 New CSI snapshot metadata service and RPCs
  4. 4:45 Kubernetes snapshot metadata gRPC APIs design
  5. 8:00 Vendor-agnostic snapshot metadata access with CBT
  6. 9:00 CSI driver sidecar architecture for CBT

Kubernetes Backup Legitimized: CSI Changed Block Tracking Has Arrived

Speakers: Shing Yang, Co-chair Kubernetes SIG Storage and Data Protection Working Group, VMware by Broadcom; Mark Lavi, Open Source Product Manager, VH Casten; Carl Braganza, Technical Staff, VH Casten

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=U9qwxp7Uv08

Overview

The KubeCon EU talk "Kubernetes Backup Legitimized: CSI Changed Block Tracking Has Arrived" delivered by Shing Yang, Mark Lavi, and Carl Braganza, introduced a pivotal advancement in Kubernetes data protection: Changed Block Tracking (CBT). This feature, long considered a fundamental capability in traditional virtual machine (VM) and bare-metal backup solutions, has been a notable omission in the cloud-native Kubernetes ecosystem. Its absence forced backup vendors to either implement full backups, leading to excessive storage, network, and compute consumption, or rely on proprietary vendor-specific solutions that undermined the portability promise of Kubernetes.

The speakers detailed the culmination of over two years of collaborative effort across multiple Kubernetes SIGs and the broader community to standardize CBT within the Container Storage Interface (CSI) specification. This standardization, formalized through KEP 3314, provides a vendor-agnostic API for efficiently identifying changes between volume snapshots. For organizations migrating from traditional IT environments to Kubernetes, the availability of standardized CBT is a critical unblocker, enabling more efficient and cost-effective backup and disaster recovery strategies, thereby legitimizing Kubernetes as a robust platform for stateful applications.

The talk covered the motivation, intricate history, architecture, and implementation details of this new feature, culminating in a demonstration of its practical application. With CBT now available in Kubernetes 1.33 alpha APIs, it marks a significant milestone, empowering backup software vendors and storage providers to deliver enterprise-grade data protection capabilities that meet the stringent Recovery Point Objective (RPO) and Recovery Time Objective (RTO) requirements expected in production environments.

Background

▶ Watch: Introduction and motivation for change block tracking (0:00)

The evolution of data protection in Kubernetes has historically lagged behind traditional IT infrastructures, particularly concerning efficient backup mechanisms. In the realm of virtual machines and bare-metal deployments, Changed Block Tracking (CBT) has been a standard feature for decades. CBT allows backup solutions to identify only the data blocks that have changed since the last backup, or between two specific snapshots, rather than requiring a full scan or transfer of the entire volume. This capability is paramount for minimizing backup windows, reducing storage consumption, and optimizing network bandwidth, directly impacting an organization's RPO and RTO metrics.

Prior to the work presented in this talk, Kubernetes lacked a native, standardized CBT mechanism. When comparing any two volume snapshots in Kubernetes, the only reliable way to determine differences was to perform a full comparison, essentially taking another full backup. This approach resulted in significant overhead: increased storage consumption for redundant data, higher network egress costs, and greater compute resources for processing and transferring large volumes of unchanged data. This inefficiency was a major pain point for enterprises adopting Kubernetes for stateful workloads, forcing them to either compromise on backup efficiency or integrate with proprietary vendor-specific CSI drivers that offered CBT functionality outside the standard Kubernetes API. Such vendor lock-in contradicted the open, portable philosophy of Kubernetes.

Recognizing this critical gap, the Kubernetes community, particularly the Data Protection Working Group within SIG Storage, initiated KEP 3314 (Kubernetes Enhancement Proposal) over two years ago. The development process was extensive, involving three major design iterations and collaboration across SIG Storage, SIG API Machinery, SIG Architecture, API reviewers, CSI reviewers, storage vendors, backup vendors, and numerous community contributors. As the speakers emphasized, "it takes a village" to bring such a foundational feature to fruition in a complex ecosystem like Kubernetes. The successful completion of this KEP and its subsequent implementation in the CSI specification and Kubernetes itself represents a monumental step towards maturing data protection capabilities within the platform.

Key Findings

▶ Watch: New CSI snapshot metadata service and RPCs (4:00)

The core findings presented in the talk revolve around the successful design and implementation of a standardized CBT mechanism for Kubernetes, addressing a long-standing deficiency in cloud-native data protection. The key contributions are:

  1. KEP 3314 Completion and Implementation: The Changed Block Tracking (CBT) KEP 3314 has been fully designed, implemented, and tested. This signifies a major milestone, moving CBT from a conceptual need to a concrete, deployable feature within Kubernetes.
  2. CSI Specification Extension: To support CBT, the CSI specification has been extended with a new snapshot metadata service. This service introduces two crucial CSI RPCs (Remote Procedure Calls):
  • GetMetadataAllocated: Retrieves metadata for all allocated blocks within a single volume snapshot. This is essential for full backups or initial scans.
  • GetMetadataDelta: Retrieves metadata for only the changed blocks between two specified volume snapshots of the same block volume. This is the core CBT functionality for incremental backups.
  • Scope Limitation: It's important to note that this initial design explicitly supports block volumes only, with support for file volumes being out of scope for now but considered for future development.
  1. Kubernetes Snapshot Metadata gRPC APIs: To expose this functionality to Kubernetes applications in a scalable manner, two new Kubernetes snapshot metadata gRPC APIs have been introduced, mirroring the CSI RPCs: GetMetadataAllocated and GetMetadataDelta. This design choice is critical as it avoids overloading the Kubernetes API server, which could be severely impacted by large volumes of metadata (potentially 5 gigabytes of metadata per 1 terabyte of volume data in worst-case scenarios).
  2. Snapshot Metadata Sidecar: The Kubernetes snapshot metadata APIs are implemented within a new snapshot metadata sidecar, which functions as a gRPC service. This sidecar container is deployed alongside the CSI driver and is responsible for calling the underlying CSI driver to retrieve the actual snapshot metadata.
  3. SnapshotMetadataService CRD: A new SnapshotMetadataService Custom Resource Definition (CRD) has been introduced. CSI drivers that implement CBT will create an instance of this CRD to advertise the availability and connection details (endpoint, CA certificate, audience string for mutual authentication) of their deployed snapshot metadata sidecar gRPC service to backup software.
  4. Vendor-Agnostic Access: The new Kubernetes snapshot metadata API provides a vendor-agnostic way for backup software to query allocated or changed blocks, eliminating the need for proprietary vendor APIs and standardizing the backup process across different storage backends.
  5. Alpha Feature Status: The CBT feature is currently released as an alpha feature with Kubernetes 1.33, indicating it's ready for testing and community feedback, with a path towards beta and stable releases.

These findings collectively represent a significant leap forward in making Kubernetes a first-class platform for running and protecting stateful applications, bringing its data protection capabilities closer to parity with traditional enterprise IT environments.

Technical Deep Dive

▶ Watch: Kubernetes snapshot metadata gRPC APIs design (4:45)

The technical architecture of Kubernetes CBT is a sophisticated integration of new components and established Kubernetes mechanisms, designed for efficiency, security, and scalability. At its heart, the system introduces a standardized way to query block-level changes without overwhelming the Kubernetes API server.

The foundation begins with extensions to the Container Storage Interface (CSI) specification. Two new gRPC RPCs are defined: GetMetadataAllocated and GetMetadataDelta. The GetMetadataAllocated RPC retrieves metadata for all allocated blocks for a given snapshot, serving as the basis for full backups. The GetMetadataDelta RPC is the core of CBT, providing metadata about blocks that have changed between two specified snapshots of the same volume. It's crucial to note that this functionality is currently limited to block volumes, with file volume support deferred for future consideration.

On the Kubernetes side, to expose this CSI functionality to applications, a set of corresponding Kubernetes snapshot metadata gRPC APIs are introduced. This is a deliberate design choice, deviating from the typical Kubernetes pattern of storing all metadata in Custom Resources (CRs) managed by the API server. The rationale is to prevent the Kubernetes API server from being overloaded. In a worst-case scenario, where every block on a volume changes, the metadata for a 1TB volume could be up to 5GB. Storing and serving such large volumes of metadata through the API server would render it unresponsive and inefficient.

Instead, these Kubernetes gRPC APIs are implemented within a dedicated snapshot metadata sidecar. This sidecar is a gRPC service that runs alongside the CSI driver. Its primary role is to act as an intermediary: it receives requests from backup software, translates them into the CSI driver's language, and then proxies these requests to the CSI driver. The CSI driver, being storage-specific, then performs the actual block-level analysis using its underlying storage system's capabilities (e.g., calling EBS direct APIs for AWS EBS volumes) and streams the metadata back through the sidecar. The CSI driver itself remains shielded from Kubernetes-specific concerns, interacting only with the CSI specification.

To enable backup applications to discover and connect to this sidecar, a new SnapshotMetadataService Custom Resource Definition (CRD) is introduced. When a CSI driver implementing CBT is deployed, it creates an instance of this CRD, named after the CSI driver. This CRD contains vital information: the endpoint (DNS name) of the gRPC service provided by the sidecar, a CA certificate for establishing trust, and an audience string used for mutual authentication.

The workflow for a backup application utilizing CBT is as follows:

  1. Discovery: The backup application first queries the Kubernetes API server to find the SnapshotMetadataService CR for the relevant CSI driver (identified from the PV or VolumeSnapshot objects). This CR provides the sidecar's endpoint, CA certificate, and audience string.
  2. Authentication Token Request: The backup application requests an OAuth token from the Kubernetes API server, specifying the audience string obtained from the CR. This token asserts the backup application's identity and intent.
  3. gRPC Call to Sidecar: With the endpoint and token, the backup application initiates a gRPC call to the snapshot metadata sidecar using the Kubernetes snapshot metadata gRPC API. This call bypasses the Kubernetes API server for data transfer.
  4. Mutual Authentication and Authorization: The sidecar receives the gRPC call. It performs mutual authentication by first validating the received OAuth token with the Kubernetes API server using the TokenReview API. Concurrently, it performs authorization checks using the SubjectAccessReview API to ensure the backup application has the necessary permissions to access the specified Kubernetes snapshots and volumes.
  5. Translation and Proxy: Once authenticated and authorized, the sidecar translates the Kubernetes object names (e.g., snapshot names) into their corresponding CSI IDs (volume IDs, snapshot IDs) by querying Kubernetes object metadata. It then proxies the request to the CSI driver using the CSI snapshot metadata gRPC API.
  6. Data Streaming: The CSI driver processes the request, generates the allocated or changed block metadata, and streams this data back to the backup application directly through the sidecar, effectively bypassing the Kubernetes API server for the actual data transfer.

Security is a paramount concern. The system leverages existing Kubernetes security primitives:

  • TokenRequest API: Used by the backup application to obtain a security token.
  • TokenReview API: Used by the sidecar to validate the authenticity of the token provided by the backup application.
  • SubjectAccessReview API: Used by the sidecar to verify the backup application's permissions to access the requested resources (snapshots, volumes).

This robust security model ensures that only authorized applications can access sensitive block metadata.

The development artifacts, including gRPC protobuf specifications, stubs, mocks, the sidecar container logic, and the CRD, are all housed in the Kubernetes CSI external snapshot metadata repository. This repository also provides utility packages, such as an iterator package for Go developers, to simplify the complex workflow of interacting with these new APIs, as well as tools for CSI driver developers to aid in implementation and testing.

Demo / Proof of Concept

▶ Watch: Vendor-agnostic snapshot metadata access with CBT (8:00)

The talk included a compelling video demonstration showcasing the practical application of the new CBT functionality using a simplified setup. The demonstration utilized the CSI host path driver, a development and testing tool, configured to deploy the newly developed snapshot metadata sidecar. While the host path driver supports block volume mode for this feature, real CSI drivers are expected to extend support for both block and file system modes where applicable.

The demonstration followed a clear sequence of steps:

  1. Application Setup and Initial Data:
  • A simple application, a busybox pod, was deployed using a CSI host path storage class.
  • A Persistent Volume Claim (PVC) was created in raw block mode.
  • Initial data was written to the PVC using the dd command, creating a specific number of allocated blocks (e.g., five odd-numbered blocks).
  1. First Snapshot and Allocated Blocks:
  • A VolumeSnapshot was taken of the PVC, representing the initial state.
  • The SnapshotMetadataService CR, installed by the CSI host path driver, was inspected. This CR contained the crucial endpoint address, audience string, and CA certificate necessary for connecting to the metadata sidecar.
  • A separate pod running the CSI client tool (available in the Kubernetes CSI external snapshot metadata repository) was launched. This tool utilizes the new Kubernetes snapshot metadata gRPC APIs.
  • The CSI client was invoked with the name and namespace of the first snapshot. It successfully queried the sidecar, which in turn communicated with the CSI driver, and reported the five allocated blocks and their byte offsets, as expected.
  1. Volume Modification and Second Snapshot:
  • The application pod then modified the volume by adding new blocks (e.g., two even-numbered blocks) and changing one of the existing odd-numbered blocks.
  • A second VolumeSnapshot was taken, capturing the modified state of the volume.
  1. Changed Block Tracking in Action:
  • The CSI client tool was invoked again, but this time specifying both the new snapshot (snapshot #2) and the previous snapshot (snapshot #1).
  • The tool correctly identified and displayed only the changed blocks between the two snapshots, demonstrating the core CBT functionality. This clearly showed which specific blocks had been added or altered, without listing the unchanged blocks.
  1. Verifying Total Allocated Blocks:
  • To confirm the overall state, the CSI client was run one more time, specifying only the second snapshot. It reported seven allocated blocks, confirming the growth of the volume after the modifications.

The speakers also delved into the codebase of the snapshot_metadata_lister tool, located in the examples subdirectory of the Kubernetes CSI external snapshot metadata repository. They highlighted the use of an iterator package, specifically the GetSnapshotMetadata entry point, which simplifies the complex workflow for application developers. This iterator handles all the underlying steps, from CR discovery and token requests to gRPC calls and mutual authentication. Developers can configure it with arguments like API clients, emitter callbacks, snapshot names (current and previous), and even a starting offset for checkpointing and resuming operations. Data is streamed back to the application via an emitter callback interface, ensuring efficient processing. The tool's flexibility allows specifying either one snapshot (to get allocated blocks) or two snapshots (to get changed blocks).

This demonstration effectively validated the end-to-end functionality of the CBT implementation, from the CSI driver to the Kubernetes APIs and the backup application's interaction, proving its readiness for adoption and further development.

Defensive Implications

▶ Watch: CSI driver sidecar architecture for CBT (9:00)

The introduction of standardized Changed Block Tracking (CBT) in Kubernetes through CSI has profound implications for defensive strategies in cloud-native environments, primarily by enhancing the efficiency and reliability of data protection. For organizations and security teams, understanding and leveraging this new capability is crucial.

Firstly, storage vendors are now empowered to implement a standardized CBT interface for their CSI drivers. This means they can provide a native, Kubernetes-aware mechanism for block-level change detection, eliminating the need for customers to rely on proprietary APIs or inefficient full volume scans. Defenders should actively request and prioritize CSI drivers that support KEP 3314 from their storage providers. This ensures that their underlying storage infrastructure can deliver the granular and efficient block metadata necessary for robust backup and recovery.

Secondly, backup software vendors are the primary beneficiaries and crucial integrators of this feature. With the new Kubernetes snapshot metadata gRPC APIs and the SnapshotMetadataService CRD, they can develop vendor-agnostic backup solutions that perform incremental backups with significantly reduced overhead. Defenders should engage with their backup software providers to ensure they adopt and integrate this standard CBT functionality into their products. This will lead to faster backups, shorter recovery point objectives (RPOs), and more efficient use of storage resources, all of which contribute to a stronger defensive posture against data loss and corruption.

Thirdly, Kubernetes operators and platform engineers play a vital role in advocating for and deploying these capabilities. They should understand the benefits of CBT, particularly in terms of reducing the cost and operational burden of managing stateful applications. By pushing for CBT adoption among their vendors, they contribute to a more resilient Kubernetes ecosystem. Furthermore, when deploying CSI drivers and backup solutions, they must ensure that the necessary RBAC policies are correctly configured. The talk highlighted the importance of permissions for TokenRequest, TokenReview, and SubjectAccessReview APIs for both the backup application and the CSI driver's sidecar. Misconfigurations in these areas could either prevent CBT from functioning or, if overly permissive, introduce security vulnerabilities.

Finally, active participation in the Kubernetes Data Protection Working Group is a key defensive action. This group is the forum for feedback, contributing to future enhancements (e.g., file volume support), and ensuring the feature evolves to meet emerging needs. By staying involved, defenders can influence the direction of Kubernetes data protection, ensuring it remains robust against evolving threats and operational challenges. The ability to perform rapid, efficient incremental backups is a cornerstone of any effective disaster recovery plan, and standardized CBT directly strengthens this foundation.

Key Takeaways

  • Kubernetes now has standardized Changed Block Tracking (CBT): KEP 3314 has been completed and implemented, bringing a crucial data protection feature, previously common in traditional VMs, to the cloud-native ecosystem.
  • CBT is implemented via new CSI and Kubernetes gRPC APIs: The CSI spec includes GetMetadataAllocated and GetMetadataDelta RPCs, mirrored by Kubernetes gRPC APIs, enabling vendor-agnostic block change detection.
  • A dedicated sidecar and CRD manage metadata access: A snapshot metadata sidecar handles gRPC requests, authorization, and proxies to the CSI driver, while a SnapshotMetadataService CRD advertises the sidecar's endpoint and security details.
  • Security is built-in with Kubernetes primitives: Mutual authentication and authorization leverage TokenRequest, TokenReview, and SubjectAccessReview APIs to ensure secure access to block metadata.
  • Efficiency gains are significant for backups: CBT allows backup software to retrieve only changed blocks between snapshots, drastically reducing storage, network, and compute overhead for incremental backups.
  • Call to action for vendors and community: Storage vendors should implement CBT in their CSI drivers, backup vendors should integrate these new APIs, and the Kubernetes community should advocate for its adoption to mature data protection.

About the Speaker(s)

The talk was presented by a team of experts deeply involved in Kubernetes storage and data protection:

  • Shing Yang is a co-chair of the Kubernetes SIG Storage and Data Protection Working Group and works for VMware by Broadcom. His role as a SIG co-chair highlights his leadership in shaping storage and data protection initiatives within the Kubernetes community.
  • Mark Lavi is the Open Source Product Manager at VH Casten. His involvement underscores the importance of open-source contributions and productization within the Kubernetes ecosystem, bridging community development with practical enterprise solutions.
  • Carl Braganza is a Technical Staff member at VH Casten and played a crucial role as one of the co-authors of the CBT KEP and its CSI addition. His technical expertise was instrumental in the design and implementation of this complex feature.

The speakers also acknowledged Prasad and Ivan as other co-authors of the CBT KEP, emphasizing the collaborative "village" effort required to bring such a significant feature to fruition.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk delivered a critical update on a foundational capability long overdue in Kubernetes: standardized Changed Block Tracking (CBT). After two years of intensive community effort across multiple SIGs, KEP 3314 has finally brought a vendor-agnostic, efficient, and secure mechanism for incremental backups to cloud-native environments. This is not just an incremental feature; it's a pivotal advancement that legitimizes Kubernetes for enterprise-grade stateful applications, drastically improving RPO/RTO metrics and eliminating reliance on proprietary, inefficient solutions.

Heather Calloway (CISO) — MUST SEE

This talk presents a foundational advancement: standardized Changed Block Tracking (CBT) for Kubernetes. This capability moves cloud-native data protection from a significant operational challenge to a legitimate, efficient, and auditable process. It directly impacts an organization's ability to meet critical RPO and RTO requirements for stateful applications, enabling true enterprise-grade resilience and accountability for data protection within Kubernetes environments. Every CISO and platform leader overseeing Kubernetes adoption needs to understand these implications and drive the integration of this standard.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025