Project Lightning Talk: Empowering Data Protection for Stateful Applications on Kuberne... Mark Lavi

Mark Lavi

KubeCon + CloudNativeCon Europe 2025 · Project Lightning Talk

Overview

In an era where Kubernetes has become the de facto standard for deploying containerized applications, the landscape of data management within these dynamic environments has grown increasingly complex. Mark Lavy, a maintainer of the Canister project, delivered an insightful lightning talk at KubeCon EU, shedding light on a critical challenge: ensuring robust data protection for stateful applications on Kubernetes. As Lavy highlighted, the widespread adoption of Kubernetes is no longer limited to stateless microservices; distributed databases, message queues, and even emerging vector databases are now routinely deployed on clusters, rendering traditional GitOps-centric recovery strategies insufficient for managing persistent data.

Watch on YouTube

Visual summary for Project Lightning Talk: Empowering Data Protection for Stateful Applications on Kuberne... Mark Lavi by Mark Lavi
Visual summary for Project Lightning Talk: Empowering Data Protection for Stateful Applications on Kuberne... Mark Lavi by Mark Lavi

Key moments

  1. 0:00 Introduction to Canister and stateful Kubernetes challenges
  2. 2:00 Canister: A framework for orchestrating data protection
  3. 2:47 Achieving application-consistent backups for distributed apps
  4. 3:33 Canister's production history and open-source origins
  5. 4:00 Canister's core components: Blueprints, ActionSets, Profiles
  6. 4:20 Understanding Canister Blueprints: actions, phases, functions
  7. 4:51 ActionSets: Instantiating blueprints for flexible backups
  8. 5:50 Call to action: How to contribute to Canister

Project Lightning Talk: Empowering Data Protection for Stateful Applications on Kubernetes with Canister

Speakers: Mark Lavy, Maintainer, Project Canister

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=6CXsWNOqYSw

Overview

In an era where Kubernetes has become the de facto standard for deploying containerized applications, the landscape of data management within these dynamic environments has grown increasingly complex. Mark Lavy, a maintainer of the Canister project, delivered an insightful lightning talk at KubeCon EU, shedding light on a critical challenge: ensuring robust data protection for stateful applications on Kubernetes. As Lavy highlighted, the widespread adoption of Kubernetes is no longer limited to stateless microservices; distributed databases, message queues, and even emerging vector databases are now routinely deployed on clusters, rendering traditional GitOps-centric recovery strategies insufficient for managing persistent data.

Canister, introduced as a brand new CNCF sandbox project (as of 2023), offers a comprehensive framework designed to orchestrate and manage data protection across various complex scenarios, including on-cluster, off-cluster, cloud, and multi-cloud environments. The talk emphasized the paradigm shift required from infrastructure-centric backup approaches to an application-consistent, top-down methodology. Canister aims to simplify the intricate process of backing up and restoring data for distributed applications, providing a structured and extensible solution where custom scripting and fragmented tools often fall short.

This article delves into the core tenets of Canister, exploring its architectural components, the problems it solves, and its significance for organizations grappling with stateful workloads on Kubernetes. By providing a unified framework for orchestrating data protection, Canister empowers developers and operations teams to achieve reliable backup and recovery, ensuring business continuity and data integrity in the highly dynamic world of cloud-native computing.

Background

▶ Watch: Introduction to Canister and stateful Kubernetes challenges (0:00)

The rapid evolution and widespread adoption of Kubernetes have profoundly transformed how applications are deployed and managed. Initially, Kubernetes gained traction for its ability to orchestrate stateless microservices, where application state was often externalized to managed databases or object storage. However, the ecosystem has matured significantly, and a growing number of stateful applications are now natively deployed within Kubernetes clusters. This shift includes critical components like Elasticsearch for search and analytics, Kafka for streaming data, PostgreSQL and MySQL for relational databases, and even cutting-edge vector databases crucial for AI/ML workloads.

The presence of state on the cluster introduces a fundamental challenge: data protection. Unlike stateless applications, where recovery might involve simply redeploying the application from a Git repository (the GitOps model), stateful applications require robust backup and recovery mechanisms. Losing persistent data can lead to significant business disruption, data corruption, or even irreversible data loss. Mark Lavy articulated this problem succinctly, noting that "there's no more GitOps your way out of this anymore."

Traditional backup strategies, often designed for bare metal servers or virtual machines, are ill-equipped to handle the complexities of distributed, sharded applications running across multiple pods, worker nodes, and even multiple clusters. Simply "freezing the disk" is insufficient for achieving a consistent backup of a live, distributed application whose data is constantly in motion and spread across various storage volumes. These applications require application-consistent backups, meaning the data snapshot captures the application's state at a point where all transactions are complete and the data is internally consistent, preventing logical corruption upon restore. Many traditional backup administrators, accustomed to simpler paradigms, are unfamiliar with the nuances of achieving such consistency in a dynamic Kubernetes environment.

Furthermore, the operational landscape for stateful applications on Kubernetes is fragmented. Data services might reside off the cluster (e.g., cloud-managed databases), on the cluster but outside the application's direct control (e.g., a dedicated database cluster), or distributed inside the pods with the application itself. Each scenario presents unique challenges for backup, recovery, and disaster recovery. Organizations often struggle with performance bottlenecks, scaling issues, and diverse security requirements that necessitate flexible and adaptable data protection strategies. The need to span "a million different domains" – from storage technologies to application-specific quiescing methods – underscores the complexity that Canister aims to address. This necessitates a framework that can orchestrate these diverse concerns in an application-centric, top-down fashion, moving beyond the infrastructure-bottom-up mindset.

Key Findings

▶ Watch: Achieving application-consistent backups for distributed apps (2:47)

Canister emerges as a pivotal solution in the evolving landscape of Kubernetes data management, offering a structured framework for tackling the intricate challenges of stateful application data protection. Mark Lavy's talk highlighted several key findings and contributions of the project:

  1. Framework for Orchestrated Data Protection: At its core, Canister is presented as a comprehensive framework designed to orchestrate data protection for stateful applications running on Kubernetes. It abstracts away much of the underlying complexity, providing a consistent mechanism to manage backup, restore, and artifact lifecycle across diverse environments.
  2. Addressing On- and Off-Cluster Concerns: Canister is built to handle both on-cluster and off-cluster data protection needs. This flexibility is crucial for modern cloud-native architectures where data might reside in persistent volumes on the Kubernetes cluster, in external cloud storage services like Amazon S3, or even in databases managed outside the cluster boundary. The framework's design acknowledges the hybrid reality of many enterprise deployments.
  3. Enabling Application-Consistent Backups: A central tenet of Canister is its focus on achieving application-consistent backups. This is a significant advancement over simple volume snapshots, which can lead to data inconsistency for distributed applications. By integrating with the application's lifecycle, Canister ensures that data is captured in a logically coherent state, crucial for successful and reliable restores. This top-down, application-centric approach stands in contrast to traditional, infrastructure-bottom-up methods.
  4. Mature and Production-Ready, Yet a New Sandbox Project: Despite its recent designation as a CNCF sandbox project in 2023, Canister is not a nascent technology. As Lavy emphasized, the project has been in production with many customers for a long time, having been open-sourced back in 2017. This dual status highlights Canister's proven reliability and maturity in real-world scenarios, while its entry into the CNCF ecosystem signifies a commitment to broader community adoption and development.
  5. Extensible via Custom Resource Definitions (CRDs): The core of Canister's extensibility and power lies in its use of Custom Resource Definitions (CRDs). These Kubernetes-native constructs allow users to define and manage data protection policies and operations as first-class objects within the Kubernetes API. The three primary CRDs—Blueprints, Action Sets, and Profiles—provide a declarative and flexible way to configure complex backup and restore workflows.
  6. Support for Complex Multi-Cloud and Multi-Target Scenarios: Canister is engineered to address highly complex data protection scenarios, including cloud and multi-cloud backup. It enables organizations to define multiple backup targets (e.g., S3, NFS, different clusters) with varying policies, such as immutability requirements and retention schedules. This level of granularity and orchestration is critical for meeting compliance, security, and operational resilience demands across heterogeneous environments.

These findings collectively position Canister as a robust, mature, and forward-thinking solution for the increasingly vital domain of data protection for stateful applications on Kubernetes.

Technical Deep Dive

▶ Watch: Canister's core components: Blueprints, ActionSets, Profiles (4:00)

Canister's technical architecture is built upon the foundational principles of Kubernetes, leveraging its extensibility through Custom Resource Definitions (CRDs) to provide a cloud-native, declarative approach to data protection. It operates as a cloud-native controller installed on the Kubernetes cluster, typically via a Helm chart, which manages the lifecycle of data protection operations.

The core of Canister's functionality revolves around three primary CRDs: Blueprints, Action Sets, and Profiles. These CRDs work in concert to define, execute, and manage complex data protection workflows for stateful applications.

  1. Blueprints:
  • Purpose: Blueprints are the heart of the Canister system. They serve as abstract templates that define the actual data protection logic and operations for a specific application type. Think of them as recipes for how to back up, restore, or manage artifacts for applications like PostgreSQL, MySQL, Elasticsearch, or any other stateful workload.
  • Structure: A Blueprint contains major actions, such as Backup, Restore, and Delete an artifact. Each action is further composed of phases, which define the sequential steps required to perform the action. For instance, a Backup action might have phases like "quiesce application," "snapshot data," "upload to storage," and "unquiesce application."
  • Capabilities: Blueprints support an explicit order of operations and can execute phases in parallel where appropriate. They also incorporate built-in functions to interact with various storage systems and application APIs. The power of Blueprints lies in their ability to encapsulate the application-specific knowledge required to achieve an application-consistent backup. This involves understanding how to put a database into a consistent state before taking a snapshot, or how to coordinate backups across multiple distributed components.
  • Examples: Canister provides example Blueprints for a wide array of databases and distributed applications on its GitHub repository, offering a starting point for users to adapt or create their own. This extensibility allows Canister to support virtually any stateful application.
  1. Action Sets:
  • Purpose: While Blueprints define how to perform a data protection operation, Action Sets are responsible for instantiating and executing a specific Blueprint. An Action Set acts as a runtime instance of a Blueprint, initiating a data protection job based on user-defined parameters.
  • Instantiation: An Action Set takes a Blueprint, provides it with runtime arguments (e.g., which specific application instance to back up, specific credentials), and specifies which location profiles should be targeted for the operation.
  • Execution and Status Tracking: When an Action Set is created, the Canister controller interprets it, executes the defined Blueprint actions, and tracks the status of the ongoing operation. It behaves much like a Kubernetes Job, providing visibility into the progress, success, or failure of the data protection task. This allows for monitoring and automated retries or alerts.
  1. Profiles:
  • Purpose: Profiles define the target locations where backup artifacts are exported and from where they can be imported during a restore operation. They abstract the details of the underlying storage infrastructure.
  • Flexibility: Profiles enable highly flexible storage strategies. An organization might define multiple Profiles, each pointing to a different storage backend (e.g., an S3 bucket in AWS, an NFS share, or an object storage service on another Kubernetes cluster).
  • Policy Enforcement: Crucially, Profiles can embed specific backup policies and retention schedules. For example, one Profile might dictate that backups to a primary S3 bucket should be retained for 30 days, while another Profile for a more secure, immutable storage target might enforce a 7-year retention policy for compliance reasons. This granular control over storage targets and policies is vital for meeting diverse operational, security, and regulatory requirements. The concept of immutability in a profile is particularly important for ransomware protection and data integrity.

Together, these three CRDs form a powerful, declarative system. An operator defines the generic data protection logic in a Blueprint, specifies where to store the artifacts in a Profile, and then triggers a specific operation by creating an Action Set that references both. The Canister controller then orchestrates the entire process, ensuring application consistency and adherence to defined policies, all within the familiar Kubernetes API and operational model. This framework significantly reduces the need for custom scripting and complex external tools, streamlining data protection for stateful applications in cloud-native environments.

Demo / Proof of Concept

▶ Watch: Understanding Canister Blueprints: actions, phases, functions (4:20)

During the lightning talk, Mark Lavy alluded to showing a "teaser of what a blueprint looks like" and mentioned the availability of "example blueprints for lots of different types of databases and distributed applications out there on our GitHub repo." However, due to the rapid pace and brevity inherent in a lightning talk format, a live, in-depth demonstration or a detailed walkthrough of a specific Blueprint's code was not provided within the transcript.

While a direct proof of concept was not demonstrated, the speaker's emphasis on the structured nature of Blueprints, with their phases, actions, order of operations, and built-in functions, clearly illustrates the conceptual framework. The existence of concrete examples on the Canister GitHub repository serves as the practical demonstration of its capabilities, allowing interested users to explore how specific applications like PostgreSQL or Elasticsearch can be configured for application-consistent backup and restore operations using Canister's CRDs. This approach enables the community to validate the project's claims through hands-on experimentation with the provided templates.

Defensive Implications

▶ Watch: Call to action: How to contribute to Canister (5:50)

The rise of stateful applications on Kubernetes introduces significant challenges for security and operational resilience, making robust data protection a critical defensive measure. Canister directly addresses several key defensive implications:

  1. Ensuring Data Integrity and Availability: The primary defensive implication of Canister is its ability to provide application-consistent backups. For highly distributed and sharded applications, simply snapshotting underlying disks can lead to logically inconsistent data upon restore. Canister's framework, through its Blueprints, integrates with the application's specific quiescing mechanisms, ensuring that the backed-up data is always in a usable and consistent state. This is fundamental for maintaining data integrity and guaranteeing application availability after a recovery event, be it from accidental deletion, data corruption, or a cyberattack.
  2. Ransomware Protection and Disaster Recovery: In the face of increasing ransomware threats, having reliable, immutable backups is paramount. Canister's Profiles allow organizations to define backup targets with specific immutability policies and retention schedules across different storage backends, including off-cluster object storage. This capability is crucial for creating secure, air-gapped copies of data that cannot be tampered with by attackers, providing a last line of defense against data loss due to ransomware or other destructive attacks. Furthermore, its support for multi-cloud and multi-cluster scenarios directly contributes to comprehensive disaster recovery (DR) strategies, enabling rapid recovery in alternative locations.
  3. Standardization and Reduced Operational Risk: By providing a declarative, Kubernetes-native framework, Canister helps standardize data protection practices across an organization's stateful workloads. Instead of relying on disparate, ad-hoc scripts or manual processes for each application, teams can define reusable Blueprints and Action Sets. This standardization reduces human error, improves auditability, and lowers the operational risk associated with complex backup and recovery procedures. It moves operations away from "coding all of that" in custom scripts, as Mark Lavy noted, towards a governed framework.
  4. Enhanced Compliance and Governance: Many industries are subject to stringent data retention and protection regulations (e.g., GDPR, HIPAA, PCI DSS). Canister's Profile CRDs, with their ability to enforce detailed retention schedules and immutability, provide a powerful tool for meeting these compliance requirements. Organizations can define distinct policies for different data types or environments, ensuring that data is stored appropriately and for the correct duration, with clear provenance of backup operations.
  5. Improved Security Posture for Data Services: The framework's ability to manage data services whether they are off-cluster, on-cluster but separate, or distributed within pods, allows security teams to enforce consistent data protection policies regardless of the deployment model. This integrated approach helps prevent gaps in coverage that might arise from managing diverse data architectures with fragmented tools. It also encourages a more holistic view of data security within the Kubernetes ecosystem.
  6. Empowering DevOps with Self-Service Data Protection: By abstracting the complexities of data protection into Kubernetes CRDs, Canister empowers DevOps teams to manage backup and recovery alongside their application deployments. This enables a more self-service model, where application teams can define and trigger their own data protection operations within the guardrails set by security and operations, fostering a culture of shared responsibility for data security.

In essence, Canister provides a robust and flexible architecture that strengthens the defensive posture of organizations running stateful applications on Kubernetes, moving beyond basic infrastructure-level backups to achieve truly resilient, application-aware data protection.

Key Takeaways

  • Kubernetes stateful applications necessitate robust, application-consistent data protection. Traditional GitOps and basic disk snapshots are insufficient for ensuring data integrity and availability for distributed, sharded workloads.
  • Canister is a CNCF sandbox project providing a comprehensive framework for orchestrating data protection for stateful applications on Kubernetes, addressing both on-cluster and off-cluster concerns.
  • The project leverages three core Custom Resource Definitions (CRDs): Blueprints, Action Sets, and Profiles. Blueprints define the application-specific backup/restore logic, Action Sets instantiate these operations with runtime arguments, and Profiles specify target storage locations and retention policies.
  • Canister supports complex multi-cloud and multi-target backup scenarios, allowing for varied backup policies, retention schedules, and immutability requirements across different storage backends like S3 and NFS.
  • Despite its recent CNCF sandbox status, Canister is a mature project, having been open-sourced in 2017 and used in production environments for several years, demonstrating its reliability and real-world applicability.
  • The framework encourages community involvement, seeking contributions from individuals with expertise in DevOps, Golang, and documentation to further enhance its capabilities and reach.

About the Speaker(s)

Mark Lavy is a dedicated maintainer of the Canister project, playing a pivotal role in its development and community engagement. He is deeply involved in addressing the challenges of data protection for stateful applications in Kubernetes environments. During the KubeCon EU conference, Mark was available at the project pavilion booth, Kiosk 208, demonstrating his commitment to the Canister community and his willingness to engage with users and potential contributors. His insights come from a background of practical experience, having been involved with the project since its open-sourcing in 2017 and its long-standing use in production environments.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This lightning talk introduces Canister, a CNCF sandbox project offering a robust, Kubernetes-native framework for data protection of stateful applications. It tackles a critical and complex problem often overlooked by traditional GitOps models, providing an application-consistent approach through CRDs like Blueprints, Action Sets, and Profiles. The project's maturity since 2017 and its well-defined architecture for orchestrating backups across diverse environments make it a significant defensive innovation for anyone running critical stateful workloads on Kubernetes.

Heather Calloway (CISO) — STRONG ACCEPT

This talk introduces Canister, a CNCF sandbox project offering a mature, Kubernetes-native framework for application-consistent data protection of stateful workloads. It directly addresses the critical need for reliable backup and recovery in complex, distributed environments, providing a structured approach through CRDs for defining policies, retention, and immutability, which are vital for business resilience and mitigating significant enterprise risk.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025