Transparent, Infra-Level Checkpoint and Restore for Resil... Ganeshkumar Ashokavardhanan & Bernie Wu
Ganeshkumar Ashokavardhanan, Bernie Wu
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk by Ganeshkumar Ashokavardhanan from Microsoft's Azure Kubernetes Service (AKS) team and Bernie Wu from Meverge delves into a critical challenge facing large-scale AI/ML workloads running on Kubernetes: resilience against infrastructure failures and optimizing resource utilization. They introduce and thoroughly define infra-level transparent checkpointing as a paradigm-shifting approach to address these issues. The core idea is to capture the entire state of a running application—including memory, CPU state, and associated files—without requiring any modifications to the application's code or framework, enabling seamless migration and hot restarts.

Key moments
- 0:00 Introduction to infra-level transparent checkpointing and speakers
- 2:10 Audience poll reveals common GPU workload resilience challenges
- 4:00 Analyzing existing resilience solutions and their limitations
- 6:05 Core definition: infra-level, transparent application checkpointing
- 7:50 Anticipating viewer questions about production readiness and restoration
Transparent, Infra-Level Checkpoint and Restore for Resilient AI/ML Workloads
Speakers: Ganeshkumar Ashokavardhanan, Software Engineer, Azure Kubernetes Service, Microsoft; Bernie Wu, VP of Technology Partnerships, Meverge
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=3oWODC2mdk0
Overview
This talk by Ganeshkumar Ashokavardhanan from Microsoft's Azure Kubernetes Service (AKS) team and Bernie Wu from Meverge delves into a critical challenge facing large-scale AI/ML workloads running on Kubernetes: resilience against infrastructure failures and optimizing resource utilization. They introduce and thoroughly define infra-level transparent checkpointing as a paradigm-shifting approach to address these issues. The core idea is to capture the entire state of a running application—including memory, CPU state, and associated files—without requiring any modifications to the application's code or framework, enabling seamless migration and hot restarts.
The motivation for this work is underscored by prevalent problems such as frequent GPU errors, low GPU utilization rates (with a third of users reporting less than 30% usage), and the significant downtime associated with node failures or slow pod restarts for large models. Current mitigation strategies often fall short, leading to substantial recomputation costs and operational complexities. By presenting a method that abstracts resilience from the application layer, the speakers offer a robust solution for platform operators and AI/ML engineers to enhance the stability and efficiency of their compute-intensive workloads, particularly those leveraging expensive GPU resources in cloud-native environments.
This innovative approach promises to unlock significant improvements in scheduling flexibility, drastically reduce recovery times from failures, and optimize the utilization of GPU infrastructure. The discussion highlights how this capability is not merely an incremental improvement but a fundamental shift in managing stateful AI/ML applications within Kubernetes, paving the way for more robust and cost-effective operations.
Background
▶ Watch: Introduction to infra-level transparent checkpointing and speakers (0:00)
The landscape of AI/ML workloads, especially large-scale training and inference, is fraught with challenges that impact both performance and cost-efficiency. As highlighted by Ganesh, observations from industry reports, such as the Llama 3 paper, confirm the high frequency of GPU errors and various infrastructure issues. A survey cited in the talk indicates that approximately one-third of users achieve less than or equal to 30% GPU usage, pointing to a severe underutilization of expensive hardware. These issues directly translate to increased operational costs and prolonged development cycles.
Existing approaches to address resilience in GPU-intensive Kubernetes environments include integrating GPU health checks with node problem detectors and remedy controllers to identify and mitigate hardware issues. While these help in detecting problems and initiating actions like reboots or GPU resets, they don't solve the application-level recovery challenges, leading to scheduling dead time and slow pod starts, especially for multi-gigabyte models. Another common strategy is model checkpointing during training, where model weights are periodically saved to a repository. If a node fails, training can resume from the last saved checkpoint. However, this often requires application-specific logic for recovery and still incurs recomputation costs for the work done between checkpoints. Tools like Q and Volcano, along with GPU strategies like multi-instance GPUs (MIG), offer some improvements but still face common limitations such as slow loading and recovery times, the need for application configuration, and potentially higher costs.
The speakers introduce infra-level transparent checkpointing as a superior alternative. They define it in three parts:
- Checkpoint: Refers to capturing the entire state of a running application, including its in-memory state, CPU registers, and all associated files, at a specific point in time. This allows the application to be paused, migrated, and "hot restarted" later.
- Transparent: Emphasizes that the application code or frameworks do not need to be modified or even be aware of the checkpointing process. This eliminates the burden on developers to build resilience logic into their applications.
- Infra-level: Indicates that the entire process is managed by the underlying orchestrators, schedulers, and the platform itself, abstracting away the complexity from end-users.
It is crucial to differentiate this from traditional model checkpointing. While model checkpointing saves only the model weights for experimentation and rollback, infra-level checkpointing captures the complete application container state. This comprehensive capture is vital for use cases beyond simple model recovery, enabling true workload migration and instant restarts.
Key Findings
▶ Watch: Audience poll reveals common GPU workload resilience challenges (2:10)
The central finding of this talk is the demonstrated feasibility and significant benefits of infra-level transparent checkpointing for enhancing the resilience and utilization of AI/ML workloads in Kubernetes. By capturing the entire application state, this approach provides a robust mechanism to overcome common infrastructure challenges without requiring application-level modifications.
The speakers identify three primary categories of challenges that this technology can address:
- Scheduling Optimization: It enables efficient handling of higher-priority workloads by allowing existing long-running jobs to be paused and migrated, freeing up resources. This is also beneficial for utilizing idle resources or managing spot instances, which can be preempted at any time. Workloads can be seamlessly moved to available capacity, maximizing resource throughput.
- Node Downtime Handling: For scenarios like GPU failures or planned node draining, transparent checkpointing allows the workload to be captured just before an incident and then restored on a new, healthy node. This significantly reduces downtime and the impact of hardware failures, which are particularly common with GPUs.
- Pod Start Time Improvements: By restoring a pre-existing application state rather than initiating a cold start, the time taken to load large models (which can be several gigabytes or even hundreds of gigabytes) and their dependencies is drastically reduced. This accelerates recovery and improves overall job throughput.
Beyond general infrastructure resilience, the talk highlights specific AI/ML use cases:
- Inference Workloads: While often considered stateless, many inference applications, especially for Large Language Models (LLMs), maintain state in the form of Key-Value (KV) cache values in both GPU and CPU memory. Losing this state requires recomputation. Transparent checkpointing allows the KV cache to be preserved and restored, avoiding costly recomputation cycles, provided the restoration time is less than the recomputation time.
- Distributed ML Training: This approach complements existing model-level checkpointing by enabling more frequent infra-level checkpoints. If a failure occurs between two model checkpoints, the infra-level checkpoint allows immediate recovery to a very recent state, minimizing lost training progress and the need for other nodes in a distributed cluster to wait for re-synchronization. This directly translates to reduced GPU/CPU cycle waste.
In essence, the key finding is that by abstracting the state management to the infrastructure layer, AI/ML workloads become more adaptable, fault-tolerant, and resource-efficient, leading to substantial cost savings and improved operational stability.
Technical Deep Dive
▶ Watch: Analyzing existing resilience solutions and their limitations (4:00)
The technical foundation for infra-level transparent checkpointing is built upon CRIU (Checkpoint Restore In User Space), an open-source project initiated in 2012. CRIU is designed to checkpoint applications running on Linux platforms, effectively capturing the state of a process or a group of processes, including their memory, CPU registers, and open files. It is widely used in production for virtual machine live migrations and by companies like Meverge for migrating long-running HPC batch workloads on public clouds, particularly for managing spot instance preemption. Notably, CRIU is integrated into Kubernetes, with the 1.30 release supporting forensic examination of containers through CRIU-based checkpointing.
Meverge and Nvidia have been actively working to extend CRIU's capabilities to include GPU-level checkpointing. This involves capturing the state of the GPU memory and execution context alongside the CPU and system memory. The community is also contributing to a robust GPU plugin for CRIU, aiming to support various GPU architectures, including AMD.
To make transparent checkpointing production-ready for AI/ML workloads, several overhead optimizations are crucial:
- Asynchronous Checkpointing: This technique significantly reduces the "production interruption window" – the time an application is paused during checkpointing. Meverge has achieved 30-100x improvement, minimizing the impact on running workloads.
- Compression Technologies: To reduce the storage footprint of checkpoints, compression is applied. The talk mentions achieving up to 5:1 compression ratios.
- Incremental Checkpointing: For scenarios requiring extremely narrow preemption windows or very frequent checkpoints, only the changes since the last checkpoint are saved, further reducing overhead.
Operationalizing this technology within Kubernetes requires addressing several considerations:
- Scheduler Integration: Meverge has experience integrating with HPC schedulers like Slurm, LSF, and HTCondor. They are now bringing this expertise to Kubernetes schedulers, starting with Q and JobSet projects, to enable event-driven checkpointing and migration based on scheduling policies.
- Security: CRIU typically requires privileged access to checkpoint processes across a node. This necessitates careful security considerations and hardening of the implementation.
- Third-Party License Managers: Checkpointing and restoring applications with embedded license managers require specific handling to ensure licenses remain valid across migrations.
- Ephemeral Files: Workloads that spill memory to disk or use temporary files must have these ephemeral states captured and restored consistently.
The current GPU checkpointing process is a two-stage operation:
- Halt and Dump GPU Memory: The system first halts any further job submissions to the target GPU for the specific process ID. It then waits for the currently running GPU tasks to complete, after which the GPU memory is dumped to system memory.
- System-level Checkpoint: Once GPU memory is in system memory, CRIU checkpoints the entire system memory, including the dumped GPU memory, along with all associated ephemeral files, to a persistent volume or designated directory.
The restoration process is simply the reverse of these two stages.
For distributed AI/ML architectures, coordinating checkpoints across multiple nodes is critical to maintain data consistency and avoid messages being lost or corrupted in transit. Meverge's solution for distributed transparent checkpointing involves several key components:
- High-Level Coordinator: This component discovers the membership of the distributed cluster, mapping network relationships between workers. It can also query the JobSet API to identify worker pods.
- Synchronizer: This ensures that all workers invoke the CRIU checkpointing operation concurrently. It utilizes pre-dump and post-dump action script hooks within CRIU. The pre-dump hook acts as a barrier, ensuring all nodes are synchronized before the checkpoint, preventing messages in transit from being lost. The post-dump hook releases the barrier, allowing applications to resume simultaneously.
- Web Hook: This mechanism allows applications to specify a checkpoint path, typically to a persistent volume, where the checkpoint data is stored.
- DaemonSet: A DaemonSet deploys the necessary checkpointing agents and utilities on each host in the cluster.
- Modified runc: The standard
runccontainer runtime is modified to pull a checkpointed image from a specified directory path instead of performing a cold restart from a container registry.
The entire process is automated via an operator, which orchestrates graceful preemption and hot restart for both individual and distributed workloads. This operator can be deployed as a sidecar alongside application containers or as a DaemonSet for node-wide management. The goal is to reduce the "friction" of moving or hibernating stateful, long-running AI/ML workloads.
Two primary use cases for this operator are highlighted:
- JobSet Migration: The operator can migrate an entire distributed JobSet to a different location, useful for defragmenting infrastructure, rebalancing loads, or allowing higher-priority jobs to take over resources.
- Node Maintenance: For failing nodes (especially GPU nodes, where 90% of issues can be resolved by a reboot), the operator can gracefully move the workloads from the problematic node to a hot spare, allowing the original node to be rebooted or serviced. This capability is particularly powerful when combined with predictive analytics for node failures.
Demo / Proof of Concept
▶ Watch: Core definition: infra-level, transparent application checkpointing (6:05)
Bernie Wu presented two compelling demonstrations to illustrate the capabilities of infra-level transparent checkpointing for distributed PyTorch AI/ML workloads. Both demos involved a distributed cluster, though with different levels of automation and recovery behavior.
Demo 1: Bare Metal, Manual, Periodic Checkpoint
The first demonstration showcased a manual checkpoint and hot restart on a bare metal setup involving a two-node PyTorch distributed cluster (one master, one worker).
- Initial State: The demo began with the PyTorch training job actively running, displaying epics and batches scrolling by. The
nvidia-smitool verified GPU activity on both nodes, and TCP/IP connections were shown to confirm inter-node communication. - Manual Checkpoint: A manual checkpoint was initiated on both nodes via a central checkpoint coordinator/synchronizer. This process captured the entire state of the running PyTorch application, including GPU memory and system memory. Critically, the application continued to run during the checkpoint, demonstrating the low interruption window. The checkpoint files were observed being dumped to disk.
- Migration and Failure: The checkpoint files from the worker node were manually copied (using
rsync) to a designated spare node. Following this, the original worker node was manually terminated (killed). - Hot Restart and Rollback: The monitoring was then switched to the spare node. The checkpoint coordinator was invoked to restore the application state on the spare node. The PyTorch job resumed execution.
- Result: While the job continued running seamlessly, it rolled back to the last periodic checkpoint (e.g., Epic 5). This illustrates that with periodic checkpointing, any progress made after the last checkpoint but before the failure is lost, as the restoration point is fixed at the last saved state.
Demo 2: Automated, Event-Driven, JobSet Migration
The second demonstration highlighted a more automated and event-driven scenario, showcasing a JobSet migration using the Q project and an operator for seamless recovery. This setup involved three nodes and a master.
- Initial State: A PyTorch distributed training job was running within a Kubernetes JobSet. The logs showed epics and batches progressing continuously. Node 3 was identified as the node targeted for maintenance.
- Event-Driven Migration: A maintenance CRD (Custom Resource Definition) was used in conjunction with the operator. A simple shell script was executed to "take Node 3 out of maintenance," simulating a predicted failure or a need for planned maintenance.
- Automated Actions: Immediately, the operator automatically terminated the containers on Node 3 and initiated the creation of replacement containers on Node 1 (the designated spare node). Checkpoint files were observed being dumped to an NFS share.
- Seamless Resume: The logs for the PyTorch job on the new node (Node 1) were displayed.
- Result: The application resumed exactly where it left off (e.g., Epic 9, Batch 1000). This demonstrates the power of event-driven checkpointing, where the checkpoint is triggered precisely at the moment of migration or impending failure, minimizing any loss of progress and providing true continuity.
These demos effectively illustrated the practical application of infra-level transparent checkpointing, with the second demo underscoring the potential for fully automated, near-zero downtime migrations in a cloud-native environment.
Defensive Implications
▶ Watch: Anticipating viewer questions about production readiness and restoration (7:50)
The advent of infra-level transparent checkpointing presents significant opportunities for defenders and platform engineers to enhance the resilience, security posture, and operational efficiency of AI/ML infrastructure running on Kubernetes.
- Proactive Resilience Against GPU Failures: Given the high frequency of GPU errors, platform teams should integrate CRIU-based transparent checkpointing solutions into their Kubernetes clusters. This allows for graceful workload migration away from failing GPUs or nodes before a complete outage, drastically reducing the impact of hardware failures on critical AI/ML training and inference jobs.
- Optimized Node Maintenance: Transparent checkpointing facilitates seamless node maintenance. Instead of draining nodes and restarting workloads from scratch, which causes significant downtime, workloads can be checkpointed and hot-restarted on healthy nodes. This is particularly valuable for planned maintenance, security patching, or even for corrective actions like rebooting a node to resolve issues, as demonstrated in the talk.
- Enhanced Distributed Training Fault Tolerance: For distributed ML training, implementing more frequent infra-level checkpoints in conjunction with traditional model-level checkpoints can provide granular fault tolerance. This minimizes recomputation costs and synchronization delays across the cluster when a single node fails, thereby improving the overall throughput and efficiency of expensive training runs.
- Improved Inference Service Availability: For stateful inference workloads, especially those using KV caches for LLMs, transparent checkpointing can maintain high availability. By preserving the KV cache state during node failures or migrations, the need for costly recomputation is eliminated, ensuring faster recovery and consistent service delivery.
- Dynamic Workload Management and Security Prioritization: Integration with Kubernetes schedulers (like Q and JobSet) enables dynamic workload management. This allows higher-priority or security-sensitive workloads to preempt lower-priority tasks, with the preempted tasks being gracefully checkpointed and migrated rather than simply terminated. This capability can be leveraged for rapid deployment of critical security patches or emergency workloads.
- Security Considerations for CRIU: While powerful, CRIU operates in a privileged mode, requiring access to low-level system states. Defenders must ensure that the CRIU implementation and its associated operator are hardened, follow least-privilege principles, and are regularly audited. Strict access controls and isolation mechanisms are essential to prevent potential abuse or exploitation of this privileged access.
- Data Integrity and Storage Security: Checkpoint files contain the entire memory and application state, making them highly sensitive. These files must be stored on secure, persistent volumes with appropriate encryption, access controls, and integrity checks to prevent data tampering or unauthorized access during storage and transit.
- Monitoring and Telemetry Integration: To effectively utilize and manage transparent checkpointing, robust monitoring and telemetry must be in place. Integrating with systems like Prometheus, as mentioned in the "Next Steps," is crucial for tracking checkpoint frequency, overhead, success rates, and overall workload resilience. This data informs optimal checkpointing strategies and helps identify potential bottlenecks or issues.
By strategically adopting and securing infra-level transparent checkpointing, organizations can build more robust, efficient, and cost-effective AI/ML platforms, mitigating risks associated with infrastructure instability and optimizing resource utilization.
Key Takeaways
- Infra-level transparent checkpointing is a critical advancement for resilient AI/ML workloads on Kubernetes, capturing the entire application state (memory, CPU, files) without requiring application code changes.
- The open-source CRIU (Checkpoint Restore In User Space) project is the foundational technology, enhanced with asynchronous checkpointing (30-100x faster), compression (up to 5:1), and incremental checkpointing for production-grade efficiency.
- This approach significantly improves GPU utilization, reduces recovery times from node failures (including common GPU errors), and accelerates pod start times for large AI/ML models.
- For distributed AI/ML training, a high-level coordinator and synchronizer are essential to ensure consistent, concurrent checkpoints across all worker nodes, preventing data loss and maintaining cluster integrity.
- Automated operators built on CRIU and integrated with Kubernetes schedulers (like Q and JobSet) enable graceful JobSet migration and efficient node maintenance, allowing workloads to be moved seamlessly across the cluster.
- Transparent checkpointing is valuable for both distributed training (reducing recomputation between model checkpoints) and stateful inference (preserving KV cache values to avoid recomputation upon failure).
About the Speaker(s)
Ganeshkumar Ashokavardhanan is a Software Engineer on the Azure Kubernetes Service (AKS) team at Microsoft. His work focuses on GPU provisioning and improving the efficiency of GPU workloads within the AKS platform, which supports a vast array of ML training and inference tasks at massive scales. He has also contributed to speeding up pod start times through features like artifact streaming.
Bernie Wu serves as the VP of Technology Partnerships for Meverge, a Silicon Valley-based startup specializing in memory virtualization and checkpointing software. Meverge aims to bring these capabilities to the AI/ML domain. Bernie has prior experience in leveraging CRIU for spot instance migration of long-running HPC batch workloads on public clouds. He has also been involved in efforts with Nvidia to introduce GPU-level checkpointing to the Kubernetes community.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk presents a genuinely groundbreaking approach to AI/ML workload resilience on Kubernetes using infra-level transparent checkpointing. Leveraging and extending CRIU to capture full application state, including GPU memory, without code changes, directly addresses critical pain points like high GPU error rates, abysmal utilization, and slow recovery times. The speakers demonstrate a sophisticated solution for automated, event-driven migration and hot restarts that promises substantial cost savings and operational stability for anyone running serious ML infrastructure.