Efficient Transparent Checkpointing of AI/ML Workloads in Kub... R. Stoyanov, A. Reber, V. Spišáková
R. Stoyanov, A. Reber, V. Spišáková
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk delves into a critical challenge facing modern cloud-native infrastructures, particularly those supporting Artificial Intelligence and Machine Learning (AI/ML) workloads: the inefficient utilization of expensive GPU resources and the lack of robust fault tolerance for long-running jobs. Presented by R. Stoyanov, A. Reber, and V. Spišáková, the session introduces transparent GPU checkpointing as a versatile solution to these pervasive problems within Kubernetes environments. The speakers highlight how this innovative approach, developed in collaboration with experts from Nvidia and AMD, can significantly improve resource efficiency and provide essential resilience for GPU-accelerated applications, ranging from interactive Jupyter notebooks to multi-day batch processing tasks.

Key moments
- 0:30 Persistent problems: inefficient utilization & fault tolerance
- 2:00 Data shows low resource utilization in Kubernetes clusters
- 5:00 Wish list: Fault tolerance for batch, efficiency for interactive workloads
- 6:00 Limitations of existing GPU resource optimization approaches
- 6:40 Introducing transparent GPU checkpointing as a versatile solution
- 8:20 How transparent unified CPU-GPU snapshots work
Efficient Transparent Checkpointing of AI/ML Workloads in Kubernetes
Speakers: R. Stoyanov, Researcher; A. Reber, Senior Software Engineer; V. Spišáková, Infrastructure Operator
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=BSoEY_tpxIo
Overview
This talk delves into a critical challenge facing modern cloud-native infrastructures, particularly those supporting Artificial Intelligence and Machine Learning (AI/ML) workloads: the inefficient utilization of expensive GPU resources and the lack of robust fault tolerance for long-running jobs. Presented by R. Stoyanov, A. Reber, and V. Spišáková, the session introduces transparent GPU checkpointing as a versatile solution to these pervasive problems within Kubernetes environments. The speakers highlight how this innovative approach, developed in collaboration with experts from Nvidia and AMD, can significantly improve resource efficiency and provide essential resilience for GPU-accelerated applications, ranging from interactive Jupyter notebooks to multi-day batch processing tasks.
The core of their work revolves around leveraging newly introduced GPU capabilities to create unified CPU and GPU snapshots, allowing for the seamless pausing, resuming, and even migrating of workloads without application awareness. This capability not only addresses the financial implications of underutilized GPUs but also mitigates the risk of losing computational progress due to unexpected failures. The talk elaborates on the technical underpinnings, demonstrates practical use cases like hot-swapping AI models and live migration of large language models, and outlines the journey of integrating this complex functionality into the Kubernetes ecosystem, specifically through container runtimes like CRI-O and containerd.
The significance of this research extends beyond mere operational efficiency; it represents a fundamental shift in how stateful GPU workloads can be managed in a dynamic, cloud-native context. By offering a transparent, low-overhead mechanism for state capture and restoration, the presented solution empowers operators to build more resilient, responsive, and cost-effective AI/ML platforms. It also opens avenues for advanced workload scheduling and even forensic analysis of running containers, marking a substantial contribution to the Kubernetes and AI/ML communities.
Background
▶ Watch: Persistent problems: inefficient utilization & fault tolerance (0:30)
The motivation for developing transparent GPU checkpointing stems from long-standing issues in operating multi-tenant, GPU-accelerated Kubernetes clusters. Victoria Spišáková, representing the Czech national e-infrastructure, outlined three key insights from their own experience managing clusters with around 50 GPUs and 300 active users running diverse workloads (batch and interactive).
Firstly, despite generous resource requests, the actual CPU utilization on computational nodes was alarmingly low, averaging just 6% and ranging between 1-50%. While some overprovisioning is expected in elastic cloud environments to handle spikes and redundancy, this level of underutilization is a significant concern, especially when extrapolated to even more expensive GPU resources.
Secondly, batch workloads, exemplified by AlphaFold protein structure prediction jobs, presented a fault tolerance challenge. These GPU-dependent jobs often run for many hours, typically up to 10, but have been observed running for as long as 30 days. While their GPU utilization is high (around 90%), hard or soft failures are an everyday reality in data centers. Losing days or weeks of computation due to a sudden node failure translates to massive wasted resources and time, making fault tolerance an essential, not merely a "nice to have," feature.
Thirdly, interactive workloads, such as Jupyter notebooks, exhibit highly unpredictable and dynamic GPU utilization patterns, characterized by bursts of activity followed by long inactive periods, sometimes lasting a full day. These workloads are expected to be readily available, yet their allocated resources often remain underutilized or unused entirely. Overprovisioning such workloads further exacerbates the overall problem of inefficient resource use.
These insights, coupled with similar concerns raised in previous KubeCon talks regarding inefficient resource utilization and fault tolerance for GPUs, crystallized into a "wish list" for GPU workloads: robust fault tolerance for all workloads (especially long-running ones) and efficient resource utilization, particularly for ubiquitous interactive applications.
Existing approaches to these problems often fall short or have significant limitations:
- Overprovisioning: Not suitable for all workload types, as applications must be aware of potential autoscaling.
- GPU Sharing (e.g., Multi-Instance GPUs - MIG): Limited in partitioning capabilities and can lead to resource fragmentation.
- Time Sharing: Can degrade performance and is not ideal for environments with unrelated user groups.
- New Scheduling Strategies: Tend to be workload-specific, lacking general applicability in diverse environments.
Recognizing these limitations, the speakers propose transparent GPU checkpointing as a versatile new tool to add to the existing toolbox, offering a more general and flexible solution. This work is also presented as a direct follow-up to a previous KubeCon talk that explored coordinated checkpointing for distributed applications, where questions about checkpointing and restoring GPU applications were a prominent concern.
Key Findings
▶ Watch: Wish list: Fault tolerance for batch, efficiency for interactive workloads (5:00)
The central contribution of this work is the successful implementation and integration of transparent GPU checkpointing into the cloud-native ecosystem, specifically within Kubernetes. The key findings and capabilities demonstrated are:
- Unified CPU/GPU Snapshots: The system can create a single, cohesive snapshot that captures the complete state of both the CPU and GPU components of an application. This is crucial for accurately preserving the runtime state of GPU-accelerated workloads.
- Full Transparency to the Application: The checkpointing and restoration process is entirely transparent to the application itself. This means developers do not need to modify their code, inject additional libraries, or alter their workflow, making the solution broadly applicable to existing AI/ML applications.
- Broad Compatibility: The approach works seamlessly with both statically and dynamically linked applications, overcoming a significant hurdle faced by older checkpointing methods.
- Integration with Container Runtimes: The functionality is integrated into container runtimes like CRI-O and containerd through CRIU (Checkpoint/Restore in Userspace) and specialized GPU driver mechanisms. This allows it to be used out-of-the-box with Docker, Podman, and, critically, Kubernetes.
- Leveraging Native GPU Capabilities: Instead of complex API interception, the solution utilizes recently introduced, lower-level GPU capabilities and driver APIs. This significantly reduces implementation complexity and performance overhead.
- Enhanced Resource Utilization: Through mechanisms like hot-swapping, workloads can be dynamically paused, checkpointed, and evicted from GPUs to allow higher-priority tasks to run, then resumed later. This optimizes the use of expensive GPU resources for interactive and batch workloads.
- Improved Fault Tolerance and Faster Recovery: Checkpointing enables long-running jobs to save their progress, allowing for quick restoration after failures. Additionally, restoring from a checkpoint results in significantly faster "cold start" times for applications, as demonstrated by the LLM migration demo where startup time was reduced from 40 seconds to 25 seconds.
- Enabling Stateful Container Migration: The ability to checkpoint and restore allows for stateful migration of GPU-accelerated containers between hosts, providing unprecedented flexibility for workload placement and resource balancing in Kubernetes.
- Forensic Container Checkpointing: The underlying Kubernetes feature, initially developed as "Forensic Container Checkpointing" (Alpha in K8s 1.25, Beta in K8s 1.30), provides a powerful security capability to snapshot a running container for offline analysis without its knowledge, aiding in threat detection and incident response.
These findings collectively demonstrate a robust, efficient, and versatile solution that directly addresses the core problems of GPU resource management and fault tolerance in contemporary Kubernetes environments.
Technical Deep Dive
▶ Watch: Limitations of existing GPU resource optimization approaches (6:00)
The technical foundation of transparent GPU checkpointing represents a significant advancement over prior methods. The speakers first critiqued the common API interception approach, which involves replacing CUDA or ROCm libraries with an interception mechanism that logs and replays API calls and tracks memory transfers. This method is notoriously difficult to implement, requires explicit interception of every API call, adds performance overhead, and struggles with architecture-specific kernel loading/unloading and static linking (e.g., ROCm's default static linking).
The proposed solution, dubbed transparently unified CPU GPU snapshots, circumvents these issues by leveraging recently introduced, low-level GPU capabilities. The integration is achieved through CRIU (Checkpoint/Restore in Userspace), a Linux tool for checkpointing and restoring processes, augmented with GPU-specific plugins.
For AMD GPUs, the solution utilizes new API calls introduced directly into the AMD GPU driver, which is part of the Linux kernel. The process involves three main steps:
- Obtain Information and Freeze: The driver API allows obtaining information about the process currently running on the GPU and then "freezing" it by evicting its queues.
- Checkpoint State: The frozen state of the GPU process is then safely checkpointed to a file descriptor.
- Unpause Operation: An "unpause" operation allows for the un-eviction of the queues, resuming execution.
Similarly, for Nvidia GPUs, a few analogous steps are employed, exposed through a command-line utility called cuda-checkpoint:
- Obtain Status: Determine the current status of the CUDA task.
- Pause Execution: Pause the execution of the GPU process.
- Checkpoint GPU State: Checkpoint the GPU state into host memory. This functionality is detailed in a blog post by Steven.
The integration of this GPU-specific checkpointing functionality with Kubernetes is designed to be seamless and fully transparent. It is part of CRI-O (and now also containerd), the container runtimes. Crucially, it does not require injecting additional libraries into the container image or modifying the application's workflow. This transparency is a cornerstone of its versatility, allowing existing AI/ML applications to benefit without any code changes.
The journey to integrate checkpointing into Kubernetes has a history dating back to a ticket opened in 2015 discussing workload migration. In 2020, work began under the umbrella of Forensic Container Checkpointing. The idea was to take a snapshot of a running container, potentially compromised, without the container or application ever knowing. This snapshot could then be analyzed offline in a sandbox, rather than simply killing the container and losing potential forensic evidence. This approach minimized the impact on the Kubernetes source code, as it initially involved only a kubelet-only API. This feature was released as an alpha feature in Kubernetes 1.25 and promoted to beta in Kubernetes 1.30.
The restoration mechanism is particularly clever. While Kubernetes views containers as largely stateless, the container engines (CRI-O and containerd) hook into the container create and container start calls. If the image passed to these calls is identified as a checkpoint image (a tar archive containing the full state, metadata, and filesystem changes), the container engine will instruct CRIU and the GPU checkpoint tools to restore the container from that state instead of creating a new one. From Kubernetes' perspective, it's a new pod, but under the hood, it's a restored, stateful container.
The checkpoint image itself, which can be quite large (e.g., 11 GB for the LLM demo), is a tar archive. For migration, this archive can be transferred directly between hosts (e.g., via rsync) or, more robustly, converted into an OCI image using tools like Buildah and pushed to a container registry. This allows for restoration on any node that can access the registry.
A significant challenge for live migration, as discussed in the Q&A, concerns network connections. While CRIU can handle TCP connections, migrating them requires maintaining the same IP address, which is generally not feasible for pods in Kubernetes. The current solution often relies on a TCP close option in CRIU configuration, which closes all open TCP connections during checkpointing. The expectation is that applications handle automatic network reconnects upon restoration. This makes it more suitable for fault tolerance (where a brief network interruption is acceptable) than for true live migration without any application impact. Furthermore, live migration of GPU workloads requires the destination host to have the same GPU type and number of GPUs, along with compatible libraries, similar to VM migration requirements.
Demo / Proof of Concept
▶ Watch: Introducing transparent GPU checkpointing as a versatile solution (6:40)
The talk featured two compelling demonstrations showcasing the practical applications of transparent GPU checkpointing.
Hot Swapping of Models
R. Stoyanov presented a demo illustrating hot-swapping of AI/ML models to optimize GPU utilization, a scenario common in multi-tenant environments.
- Initial State: A Kubernetes cluster, specifically a single node, was running a low-priority training pod in a Jupyter notebook. This pod was actively utilizing a GPU, but as a long-running training job, it was deemed eligible for preemption.
- Higher Priority Inference: When a need arose for a higher-priority task, the training pod's GPU state was transparently checkpointed into host memory. This action immediately freed up the GPU.
- Llama 3.3 Workload: The now-available GPU was allocated to an inference workload running Llama 3.3. This model was loaded into the GPU, and it began responding to requests.
- Even Higher Priority Inference: Subsequently, an even higher-priority workload emerged. The Llama 3.3 inference task was itself checkpointed into host memory, again freeing the GPU.
- DeepSeek Workload: The GPU was then used for a DeepSeek model, which has fewer parameters and offers faster response times, catering to the highest priority need.
- Resumption: Once the high-priority DeepSeek workload completed, the Llama 3.3 inference workload was transparently resumed from its checkpoint. After its completion, the original Jupyter notebook training job was also resumed from its last saved state.
This demonstration effectively illustrated how transparent checkpointing enables dynamic prioritization and efficient sharing of expensive GPU resources, allowing operators to "hot swap" multiple models or workloads to and from host memory based on current demand and priority.
Stateful LLM Migration
Adrian Reber presented a live demo of stateful migration of a Large Language Model (LLM) between two Kubernetes hosts, highlighting both fault tolerance and faster startup times.
- Initial Setup: A simple Kubernetes pod definition was used to start a VLM (Visual Language Model) on an Nvidia A10 GPU on an Amazon EC2 instance running Red Hat Enterprise Linux 9. The cold start time for this LLM was approximately 40 seconds.
- Running LLM: Once started, the LLM consumed about 4 GB of GPU memory. Adrian interacted with it via
curl, asking it to complete a sentence about CRIU, demonstrating its active state (e.g., showing a counter value). - Checkpointing: A script initiated the checkpointing process. This involved the
kubeletcommunicating with CRI-O, which in turn used runc and CRIU, finally invoking the CUDA checkpoint tool. The checkpointing process took approximately 1 minute and 4 seconds to write all the CPU and GPU state data (totaling about 11 GB) to local disk. - Transfer and Conversion: The 11 GB checkpoint image was then transferred to a second host. To facilitate Kubernetes deployment, this tar archive was converted into an OCI image using Buildah, a process that also took a few minutes.
- Restoration on New Host: On the destination system, a YAML file (almost identical to the original) was used to start a new pod. However, because the image was a checkpoint image, CRI-O detected this and instead of a cold start, initiated a restoration process using CRIU and the CUDA checkpoint tool.
- Verification: The LLM resumed from its exact prior state, continuing the counter from where it left off on the original host. Crucially, the startup time on the new host was only 25 seconds, significantly faster than the 40-second cold start.
This demo powerfully showcased the ability to perform stateful migration of complex, GPU-accelerated applications, providing not only fault tolerance but also a tangible benefit in reducing application downtime during restarts or migrations.
Defensive Implications
▶ Watch: How transparent unified CPU-GPU snapshots work (8:20)
While the primary focus of this talk is on improving efficiency and fault tolerance for AI/ML workloads, the underlying Forensic Container Checkpointing feature has significant defensive implications for cloud-native security. The ability to transparently snapshot a running container offers a powerful new tool for security teams.
Traditionally, when a container is suspected of compromise, the standard response is often to terminate it to prevent further damage. However, this action destroys crucial runtime evidence that could be vital for understanding the attack, identifying vulnerabilities, and improving future defenses. With transparent checkpointing, security professionals can:
- Non-Disruptive Forensic Analysis: Instead of killing a suspicious container, a checkpoint can be taken without the application (or an attacker within it) knowing. This allows the container to continue running, potentially enabling observation of attacker behavior, while a full snapshot of its memory, CPU state, and GPU state is captured for offline analysis.
- Sandbox Investigation: The captured checkpoint can then be restored in an isolated, sandboxed environment. This allows security analysts to meticulously examine the container's state, including its memory pages, process tree, open files, network connections, and the state of any GPU-accelerated processes. This deep dive can reveal malicious payloads, indicators of compromise (IOCs), root causes, and attack vectors that might otherwise be missed.
- Post-Mortem Analysis and Threat Hunting: The ability to recreate the exact state of a container at a specific moment in time is invaluable for post-mortem analysis after an incident. It also empowers advanced threat hunting scenarios, where suspicious patterns can trigger a checkpoint for deeper inspection without impacting production services.
- Understanding Attack Behavior: By restoring a checkpoint and allowing the container to run in a controlled environment, security researchers can observe how malware behaves, how it interacts with the GPU (if applicable), and what its objectives are, leading to more robust detection and prevention strategies.
- Breaking Stateless Assumptions: The existence of stateful container capabilities challenges the traditional "stateless" assumption often made about containers in security tools. This might necessitate a re-evaluation of existing security monitoring, intrusion detection, and incident response workflows to account for the possibility of state persistence and migration.
In essence, transparent GPU checkpointing transforms a potentially destructive security response into a data-rich forensic opportunity, enabling more sophisticated and effective defense strategies in Kubernetes environments.
Key Takeaways
- Addresses Critical AI/ML Challenges: Transparent GPU checkpointing directly tackles the pressing issues of inefficient GPU utilization and lack of fault tolerance for AI/ML workloads in Kubernetes.
- Enables Unified State Capture: The technology allows for the creation of cohesive snapshots that capture the full CPU and GPU state of an application, crucial for accurate state preservation.
- Fully Transparent Operation: Checkpointing and restoration are entirely transparent to the application, requiring no code modifications or special library injections, ensuring broad applicability to existing workloads.
- Integrated into Kubernetes Ecosystem: The functionality is integrated into container runtimes like CRI-O and containerd via CRIU and specific GPU driver APIs (e.g., cuda-checkpoint), making it an out-of-the-box feature in Kubernetes (Alpha in K8s 1.25, Beta in K8s 1.30).
- Facilitates Dynamic Workload Management: Features like hot-swapping (demonstrated with Jupyter, Llama 3.3, and DeepSeek models) and stateful migration enable efficient resource allocation, dynamic prioritization, and faster recovery times (e.g., 25s startup vs. 40s cold start for LLM).
- Powerful Security Implications: The underlying "Forensic Container Checkpointing" provides a novel capability for non-disruptive security analysis, allowing suspicious containers to be snapshotted for offline investigation without alerting attackers or interrupting service.
- Ongoing Development: Efforts are underway to further integrate checkpointing into higher-level Kubernetes APIs (API server,
kubectl) for easier use and broader adoption.
About the Speaker(s)
The presentation was a collaborative effort by three speakers, each contributing their unique expertise:
- R. Stoyanov (Researcher) presented the core concepts and the hot-swapping demo. He highlighted the work as a collaboration with his supervisors, Professor Rodrigo Bruno and Professor Wes Armo, as well as several individuals from Nvidia and AMD, underscoring the deep technical partnerships involved in this research.
- A. Reber (Senior Software Engineer) provided a detailed technical deep dive into how checkpointing and restoration work within Kubernetes, focusing on the integration with container runtimes like CRI-O and containerd. He also conducted the live demo of stateful LLM migration, showcasing the practical implementation of the technology. His background in container runtime development is evident in the nuanced explanation of Kubernetes' internal mechanisms.
- V. Spišáková (Infrastructure Operator) opened the talk by presenting the real-world operational challenges and insights from her experience at the Czech national e-infrastructure. Her role involves operating multi-tenant, multi-purpose Kubernetes clusters with around 50 GPUs and 300 active users, providing a crucial operator's perspective on the problems of inefficient resource utilization and fault tolerance for AI/ML workloads.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This KubeCon talk is a must-see for anyone serious about operating AI/ML workloads at scale. It tackles the critical problems of GPU underutilization and lack of fault tolerance with an incredibly elegant and transparent solution: unified CPU/GPU checkpointing. The research is deeply technical, leveraging native GPU driver capabilities and integrating seamlessly into container runtimes, making it a practical game-changer for resource efficiency and resilience. The live demos were compelling, showcasing real-world benefits like hot-swapping models and stateful LLM migration, with the added bonus of powerful forensic capabilities.
Heather Calloway (CISO) — STRONG ACCEPT
This session on transparent GPU checkpointing for AI/ML workloads in Kubernetes is far more than a technical optimization. While it expertly addresses critical challenges of GPU resource efficiency and fault tolerance, its most compelling implication for security leaders is the underlying "Forensic Container Checkpointing" capability. This feature provides an unprecedented ability to non-disruptively snapshot a potentially compromised container for thorough offline analysis, transforming incident response from a destructive act into a data-rich forensic opportunity. This capability offers a clear path to improved institutional accountability and more effective post-breach investigation…