Streamlined Efficiency: Unshackling Kubernetes Image Volumes for Rapid A... Esteban Rey & Yifan Yuan

Esteban Rey, Yifan Yuan

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Esteban Rey from Microsoft and Yifan Yuan from Alibaba Cloud, addresses a critical bottleneck in modern AI/ML workloads running on Kubernetes: the inefficient loading of large datasets into containerized environments. While significant progress has been made in optimizing application startup through technologies like artifact streaming, these solutions often fall short when dealing with the massive, frequently updated datasets required for AI model training and inferencing. The core problem lies in the traditional approach of packaging these datasets directly into container images, which incurs substantial time and storage overhead.

Watch on YouTube

Visual summary for Streamlined Efficiency: Unshackling Kubernetes Image Volumes for Rapid A... Esteban Rey & Yifan Yuan by Esteban Rey, Yifan Yuan
Visual summary for Streamlined Efficiency: Unshackling Kubernetes Image Volumes for Rapid A... Esteban Rey & Yifan Yuan by Esteban Rey, Yifan Yuan

Key moments

  1. 0:00 Introduction to image volumes and challenges
  2. 2:00 Key challenges with large AI/ML dataset loading
  3. 3:15 Leveraging OCI registries for data management benefits
  4. 4:05 Critical role of OCI garbage collection for data
  5. 6:30 Limitations of OCI registries for volume mounting
  6. 7:30 Problems with overlay filesystems and data packaging
  7. 8:15 Visualizing the significant cost of data packaging

Streamlined Efficiency: Unshackling Kubernetes Image Volumes for Rapid AI Model and Data Set Loading

Speakers: Esteban Rey, Software Engineer, Microsoft; Yifan Yuan, Senior Software Engineer and Researcher, Alibaba Cloud

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=XqL5lh32lr8

Overview

This talk, presented by Esteban Rey from Microsoft and Yifan Yuan from Alibaba Cloud, addresses a critical bottleneck in modern AI/ML workloads running on Kubernetes: the inefficient loading of large datasets into containerized environments. While significant progress has been made in optimizing application startup through technologies like artifact streaming, these solutions often fall short when dealing with the massive, frequently updated datasets required for AI model training and inferencing. The core problem lies in the traditional approach of packaging these datasets directly into container images, which incurs substantial time and storage overhead.

The speakers introduce Elink, an innovative approach that leverages the existing strengths of OCI registries – their scalability, high availability, versioning capabilities, and familiar tooling – to facilitate rapid, on-demand loading of large datasets without the need for traditional image packaging. Elink fundamentally rethinks how data is accessed, moving from packaging data into an image to packaging references to remote data, thereby creating a mountable volume directly from an OCI artifact. This paradigm shift promises to unlock unprecedented efficiency for AI workloads by drastically reducing preparation times and optimizing resource utilization.

The importance of this work cannot be overstated in an era dominated by large language models and data-intensive AI applications. As datasets continue to grow exponentially (an IDC report suggests 175 zetabytes by 2025), the ability to efficiently manage, version, and deploy them becomes a cornerstone of scalable and cost-effective AI development. Elink offers a compelling solution to these challenges, enabling organizations to maximize their GPU utilization, streamline development workflows, and reduce operational costs associated with data management in Kubernetes.

Background

▶ Watch: Introduction to image volumes and challenges (0:00)

The journey towards efficient container startup has seen considerable advancements, particularly with the advent of artifact streaming. This technology significantly reduces the time it takes for applications to become ready by allowing container runtimes to stream image layers on demand, rather than waiting for the entire image to download. However, this optimization primarily targets application binaries and their dependencies, which typically constitute smaller, more static components of a workload.

When the focus shifts to large-scale AI and machine learning, the landscape changes dramatically. AI workloads, especially those involving large language models (LLMs) or complex training regimes, require access to vast datasets that can range from gigabytes to terabytes in size. These datasets often need to be accessed in parallel across numerous nodes, necessitating continuous, fast, and highly scalable data access. Furthermore, the dynamic nature of AI development means these datasets are frequently updated, requiring robust versioning and efficient garbage collection mechanisms.

Traditional methods of integrating these datasets into Kubernetes environments present several challenges:

  • Continuous Data Access: Without it, expensive GPUs sit idle, leading to wasted resources.
  • Speed: Data access must be fast enough to keep pace with computation.
  • Scalability: Solutions must scale to thousands of nodes to support large-scale distributed training.
  • Versioning: Both application code and data need cohesive versioning to ensure reproducibility and consistency.
  • Management Overhead: Any solution must be easy to manage; complex systems will see low adoption.

The speakers highlight the potential of OCI registries as a foundation for solving these problems. OCI registries are already designed for performance, scalability, and high availability, making them suitable for Kubernetes workloads. They offer inherent versioning capabilities through tags and digests, ensuring data consistency and immutability. Crucially, OCI registries also support garbage collection, a vital feature for managing ever-growing datasets and preventing unchecked storage costs. Users are already familiar with the OCI distribution specification and its associated tooling, making it a natural extension for data management.

Despite these advantages, using OCI images for large data volumes faces significant hurdles:

  • Design Limitations: Registries were not initially designed for direct volume mounting. This leads to inconsistencies in implementation, such as varying layer size limitations across different registries (e.g., Azure Container Registry has a 200GB limit per layer, but this is not standardized).
  • Overlay Filesystem Issues: OCI images typically rely on overlay file systems, which create new layers for every data modification. For large, frequently changing datasets, this can quickly lead to an explosion of layers, hitting layer limits and incurring performance costs.
  • Prohibitive Packaging Costs: The most critical limitation is the time and resources required to package large datasets into OCI image layers. The speakers presented compelling data from experiments using popular machine learning datasets from kaggle.com. Packaging a 22GB dataset containing 700,000 image files took nearly four hours using a standard Dockerfile COPY command. Even smaller datasets (a few gigabytes) took many minutes. This overhead makes frequent updates or large-scale data ingestion impractical and significantly hinders agile AI development.

This background sets the stage for Elink, a solution specifically engineered to overcome the packaging bottleneck while retaining the benefits of OCI registries for data management.

Key Findings

▶ Watch: Leveraging OCI registries for data management benefits (3:15)

The central insight driving Elink is the realization that the primary inhibitor to using OCI images for large datasets is the packaging process itself, not the OCI registry's ability to store and distribute data. While OCI registries offer robust infrastructure for availability, scalability, versioning, and garbage collection, the act of serializing vast amounts of data into image layers is time-consuming and inefficient, especially for frequently updated datasets or those with many small files.

The speakers observed that existing solutions like remote snapshotters (e.g., stargz-snapshotter, overlaybd) already address parts of this problem for application images. These technologies allow for on-demand streaming of image layers by creating external indexes that map file paths to remote blob locations, thus avoiding the need to download entire images upfront.

Elink extends this concept by asking: "Could we build an index for an entire storage bucket and use that to create a mount point for remote data access via OCI artifacts, all without traditional data packaging?" The answer, as presented, is a resounding yes.

The key finding and contribution of this talk is Elink: a novel solution that enables the creation of mount points for accessing remote data directly through an OCI artifact, completely bypassing the laborious and resource-intensive data packaging step. Instead of packaging the actual data, Elink packages a reference list that describes the data objects in remote storage. This reference list acts as a metadata index, allowing a specialized snapshotter to dynamically mount the remote data, making it appear as a local file system volume within a Kubernetes pod. This approach dramatically reduces the overhead associated with preparing data for containerized AI workloads, transforming hours of packaging time into mere minutes.

Technical Deep Dive

▶ Watch: Critical role of OCI garbage collection for data (4:05)

Elink's architecture centers around the concept of a reference list and its integration with OCI artifacts and specialized snapshotters. The goal is to provide a virtual file system view of remote data without physically moving or packaging that data into an OCI image layer.

Core Components of Elink

  1. Data Set Description: This is the high-level definition of the collection of data objects intended for use.
  2. Reference List: This is the core metadata component. It's a set of records, each describing a specific data object in the remote storage.
  3. Remote Snapshotter: A component (like OverlayBD) that parses the reference list and creates a mount point, allowing applications to access the remote data as if it were local.

The Reference List in Detail

Each record within the reference list contains at least four crucial pieces of information for every data object:

  • source path: The original, full path to the object in the remote backend storage (e.g., an S3 bucket or Azure Blob Storage).
  • mount path: The relative path where this object should appear within the container's mounted volume. This defines the file system structure presented to the application.
  • etag: An entity tag, typically a hash or version identifier, that reflects the data's current state. It allows the system to detect if the remote object has changed without downloading it, ensuring data consistency. The speakers clarify that the etag often serves as an MD5 checksum of the object.
  • file size: The size of the object in bytes.

Through this reference list, Elink can describe all objects intended for access within an OCI artifact.

Packaging the Reference List into an OCI Artifact

Instead of the raw data, it's the reference list that gets "packaged" into an OCI artifact. This process involves:

  1. Registry Backend Integration: The OCI registry needs to be able to resolve and potentially redirect to the actual backend storage where the data resides. When a request is made for a blob (which, in Elink's case, is the data object referenced in the list), the registry can return either a redirect URL to the backend storage or the blob content directly. This allows Elink to construct the actual access path for the target blob by combining the registry's endpoint with the source path from the reference list.
  2. Special Annotation: A specific annotation field is added to the OCI layer containing the reference list. This annotation serves as a flag, identifying the layer's special purpose to the snapshotter, signaling that it's a reference list rather than a traditional data layer. This enables existing OCI artifact tooling to support remote file access.
  3. Common Format: The reference list itself is saved in a widely understood, parseable format such as CSV or JSON. This allows the registry (or any other component) to analyze the list's details and understand the scope of the remote data.

Mount Point Creation and Streaming Loading

When a container runtime needs to mount an Elink-enabled image volume:

  1. Pull and Unpack Reference List: The container runtime performs a regular image pull, but instead of downloading large data layers, it only downloads and unpacks the relatively small OCI artifact containing the reference list.
  2. Snapshotter Action: The dedicated snapshotter (e.g., OverlayBD) then parses the reference items from this list. For each item, it creates a file entry at the specified mount path within the virtual file system. Crucially, this file entry doesn't contain the actual data but rather points back to its source path in the remote storage.
  3. Authorization: A critical consideration is access permission. The registry itself may not have permissions to access all remote objects directly. Therefore, an additional authorization mechanism may be required to grant the snapshotter or the underlying system the necessary rights to fetch data from the remote backend storage.
  4. Streaming Loading: This is essential for large datasets. When an application within the container attempts to read from a file at its mount path, the streaming service intercepts this I/O request. It first verifies the etag from the reference list against the remote object's current etag (via an HTTP HEAD request, for instance). If they match, indicating the data hasn't changed, the I/O request is converted into a range GET request for the corresponding portion of the remote target object. This allows data to be fetched on-demand, sector by sector or byte by byte, without downloading the entire file. This prevents the large storage footprint and potential download failures associated with full image downloads.

Integration with OverlayBD

For their Proof of Concept (PoC), the speakers chose OverlayBD (Overlay Block Device) as the underlying remote snapshotter. OverlayBD offers several advantages:

  • Merged View as Virtual Block Device: It provides a unified view of image layers as a virtual block device, which is highly performant.
  • On-Demand Transfer: It supports on-demand data transfer at the disk sector level, optimizing for granular access.
  • Block Device Interface: Unlike FUSE-based solutions, OverlayBD presents a block device interface, which generally offers better performance, especially for small file access, and is more mature in handling stability issues like crash recovery.
  • Production Proven: OverlayBD is widely used in production environments at Alibaba Group, Azure, and Databricks, attesting to its robustness.
  • Turbo OCI Mode: OverlayBD has a lightweight mode called Turbo OCI that already supports indexing OCI images and building E4S file systems locally from these indexes.

Elink integrates with OverlayBD by implementing a simple function that maps the I/O requests received from the mount point into the corresponding entries in the reference list. Since OverlayBD provides a backend implementation of a block device, the file system and block driver convert application I/O requests into simple read/write operations that Elink can then translate into remote range GETs based on the reference list. This seamless integration leverages OverlayBD's performance and stability for Elink's remote data access.

Demo / Proof of Concept

▶ Watch: Problems with overlay filesystems and data packaging (7:30)

The speakers demonstrated Elink's capabilities through performance tests focusing on two critical areas: packaging time and end-to-end data access performance. The tests were conducted on environments provided by Alibaba Cloud.

Packaging Phase Performance

This test directly addresses the core problem Elink aims to solve: the exorbitant time required to package large datasets into traditional OCI images.

  • Dataset: A 22GB dataset containing over 700,000 files (mostly images), commonly used for machine learning training, was chosen from kaggle.com.
  • Traditional OCI Image: Packaging this dataset into a conventional OCI image using a Dockerfile COPY command took nearly four hours. This highlights the significant bottleneck for data-intensive AI workloads.
  • Elink: In stark contrast, Elink's build speed for the same dataset was less than two minutes. This dramatic reduction is achieved because Elink doesn't copy the data itself; it merely records the metadata (the reference list) pointing to the remote data. This transformation from physical data movement to metadata indexing is Elink's most significant advantage in the packaging phase.

End-to-End Data Access Performance

This test compared the throughput of Elink against two other methods for accessing data: traditional OCI images and Goofys. Goofys is a high-performance, FUSE-based file system implementation for AWS S3. The tests measured the total time, including preparation, to access various datasets. The average size of individual files within each dataset was also considered, as this often impacts performance.

  • Preparation Time:
  • Traditional OCI Image: Included downloading and unpacking the entire image, which is the default behavior for creating an OCI volume mount point.
  • Goofys and Elink: Included the time taken to create the mount point and build the file system metadata. For Elink, this specifically involved pulling and parsing the reference list.
  • Results:
  • Large Number of Small Files (e.g., US Accident Data Set): Elink demonstrated significantly better performance than both traditional OCI images and Goofys. This scenario is particularly challenging for file systems due to high metadata and I/O overheads, and Elink's optimized indexing and on-demand streaming for small files proved superior.
  • Large Single File: Elink's performance was on par with Goofys. This indicates that even for simpler access patterns, Elink maintains competitive performance, leveraging the efficiency of direct remote access.

These results clearly validate Elink's effectiveness in both preparation efficiency and runtime data access performance, particularly for the challenging scenario of large datasets comprising numerous small files – a common characteristic of many AI training datasets.

Defensive Implications

▶ Watch: Visualizing the significant cost of data packaging (8:15)

Elink presents several positive implications for defenders and security professionals managing Kubernetes environments, alongside some areas that warrant further development and attention.

Positive Implications

  • Enhanced Operational Efficiency and Cost Savings: By dramatically reducing the time and storage required for packaging large datasets, Elink directly translates to lower operational costs. Faster data loading means less GPU idle time, maximizing the return on expensive AI infrastructure. Efficient garbage collection, native to OCI registries, prevents data sprawl and reduces long-term storage expenses.
  • Improved Data Versioning and Consistency: Leveraging OCI registry features like tags and digests, coupled with Elink's use of etags for individual data objects, provides robust mechanisms for data versioning and ensuring consistency. Defenders can rely on these identifiers to track data lineage, verify immutability, and ensure that AI models are trained and deployed with the correct, untampered datasets. The etag check prior to data transfer is a crucial integrity control.
  • Streamlined Data Management: Integrating data access within the familiar OCI ecosystem simplifies management. Existing tooling for OCI artifacts can be extended to manage data volumes, reducing the learning curve and potential for misconfiguration.

Security Considerations and Recommendations

The question and answer session following the presentation highlighted critical security aspects that defenders should be aware of:

  • Data Integrity Verification for Streaming Blobs: A key concern raised was how checksums are verified when data is streamed on demand, rather than downloaded entirely. The speakers clarified that Elink primarily relies on the etag (often an MD5 checksum of the object) to verify data integrity. Before streaming, an HTTP HEAD request can retrieve the etag from the remote object. If this etag mismatches the one stored in the reference list, it indicates the data has been changed.
  • Defensive Recommendation: While etag provides a quick integrity check for changes, it's not a full cryptographic checksum verification of the entire streamed data. For highly sensitive data, defenders might need to implement additional integrity checks at the application layer or explore future Elink enhancements that incorporate more robust, end-to-end cryptographic verification methods for partial data streams.
  • Sixdoor Image Signing: The question of how Elink integrates with Sixdoor image signing (a framework for signing and verifying container images and artifacts) was posed. The speakers acknowledged that this is an area "still somewhat early stages" and not yet fully addressed.
  • Defensive Recommendation: This is a significant gap. Without robust signing and verification, the integrity and authenticity of the reference list itself, and by extension, the data it points to, cannot be cryptographically guaranteed. Malicious actors could potentially tamper with a reference list to point to compromised data. Defenders should be aware of this limitation and implement other controls (e.g., strict access controls to the OCI registry storing the reference lists, network segmentation) until Sixdoor integration or a similar artifact signing mechanism is fully supported.
  • Authorization for Remote Objects: The technical deep dive mentioned that "access permission is required here because the registry may not have the access permission for all the remote object." An "additional authorization may be needed."
  • Defensive Recommendation: This is paramount. Robust Identity and Access Management (IAM) must be implemented for the remote backend storage (e.g., S3, Azure Blob Storage). The Kubernetes service account or pod identity used by the snapshotter/Elink component must be granted least privilege access to only the necessary data objects. Auditing access to these backend stores is also critical.
  • Supply Chain Security: While Elink addresses efficiency, it introduces a new component (the reference list and its processing) into the data supply chain. Defenders need to ensure that the process of generating, storing, and consuming these reference lists is secure, protected from tampering, and integrated into existing supply chain security practices.

In summary, Elink offers compelling advantages for AI/ML workloads by optimizing data loading. However, defenders must be diligent in understanding its current security posture, particularly regarding comprehensive data integrity verification for streaming, the lack of artifact signing for reference lists, and the critical need for robust authorization for backend storage. As Elink matures, these areas will require further development to meet enterprise-grade security requirements.

Key Takeaways

  • Traditional OCI image packaging introduces significant time and storage overhead (e.g., 4 hours for 22GB) for large AI datasets, hindering efficient AI/ML workloads in Kubernetes.
  • Elink leverages OCI registries for scalable, versioned, and garbage-collectible data access by packaging a "reference list" of remote data objects instead of the data itself.
  • The reference list contains metadata like source path, mount path, etag, and file size, enabling a snapshotter to create a virtual file system volume directly from remote storage.
  • Elink drastically reduces data preparation time (from hours to less than 2 minutes for a 22GB dataset) and improves end-to-end data access performance, especially for scenarios with numerous small files.
  • The solution integrates with high-performance remote snapshotters like OverlayBD, which provides a robust, production-proven block device interface for on-demand data streaming.
  • While providing etag-based integrity checks for streamed data, Elink is in its early stages regarding advanced security features like Sixdoor image signing for the reference list, requiring defenders to implement complementary security controls.

About the Speaker(s)

Esteban Rey is a Software Engineer at Microsoft. His work primarily focuses on the Azure Container Registry, with a specialized expertise in OCI conformance and artifact streaming. His contributions are instrumental in advancing the efficiency and standards of container image management.

Yifan Yuan is a Senior Software Engineer and Researcher at Alibaba Cloud. He has made significant contributions to the OverlayBD project, a core technology enabling advanced container image streaming and volume management. His research and development efforts are key to high-performance data access in cloud-native environments.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk presents Elink, a genuinely novel and highly impactful solution to a critical bottleneck in large-scale AI/ML workloads on Kubernetes: the inefficient packaging and loading of massive datasets. By leveraging OCI registries to store 'reference lists' of remote data objects instead of the data itself, Elink dramatically cuts data preparation times from hours to minutes and improves runtime access. It's a clever application of existing cloud-native primitives to solve a very real, growing problem, demonstrating significant technical depth and practical value.

Heather Calloway (CISO) — STRONG ACCEPT

This session introduces Elink, a compelling technical innovation that effectively addresses the significant bottleneck of loading large AI/ML datasets in Kubernetes. By shifting from traditional data packaging to a metadata-driven approach leveraging OCI registries, it delivers dramatic improvements in data preparation time and operational efficiency, directly impacting business costs and development velocity for AI initiatives. While the operational benefits are clear and well-demonstrated, the current immaturity in critical security areas, particularly the absence of robust artifact signing for the reference lists and the need for explicit authorization mechanisms, demands careful…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025