Generative AI Model Data Pre-Training on Kubernetes: A Use Case St... Alexey Roytman & Anish Asthana

Alexey Roytman, Anish Asthana

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Anish Asthana from Red Hat and Alexey Roytman from IBM Research, delves into the intricate world of foundation model data engineering on Kubernetes. It addresses the critical challenges associated with preparing massive datasets for training large language models (LLMs), focusing on how to scale complex data pre-processing workflows from local development environments to production-grade, cloud-native infrastructure. The speakers share their experiences and solutions, particularly highlighting the integration of Ray for distributed computing and Kubeflow Pipelines (KFP) for orchestration within a Kubernetes ecosystem.

Watch on YouTube

Visual summary for Generative AI Model Data Pre-Training on Kubernetes: A Use Case St... Alexey Roytman & Anish Asthana by Alexey Roytman, Anish Asthana
Visual summary for Generative AI Model Data Pre-Training on Kubernetes: A Use Case St... Alexey Roytman & Anish Asthana by Alexey Roytman, Anish Asthana

Key moments

  1. 0:00 Introduction and talk agenda overview
  2. 1:00 Core steps in data pre-processing workflows
  3. 3:00 Scaling data workflows with Ray and CubeRay
  4. 4:00 CubeRay architecture, driver-worker, and task isolation
  5. 6:00 Real-world CubeRay deduplication at massive scale
  6. 6:30 Orchestrating long-running ETL jobs with Kubeflow Pipelines
  7. 8:00 Kubeflow Pipelines UI and DAG visualization
  8. 9:00 Elyra: simplified drag-and-drop for Kubeflow Pipelines

Generative AI Model Data Pre-Training on Kubernetes: A Use Case St... Alexey Roytman & Anish Asthana

Speakers: Alexey Roytman, IBM Research; Anish Asthana, Engineering Manager, Red Hat

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=CzdX5qDgQ2U

Overview

This talk, presented by Anish Asthana from Red Hat and Alexey Roytman from IBM Research, delves into the intricate world of foundation model data engineering on Kubernetes. It addresses the critical challenges associated with preparing massive datasets for training large language models (LLMs), focusing on how to scale complex data pre-processing workflows from local development environments to production-grade, cloud-native infrastructure. The speakers share their experiences and solutions, particularly highlighting the integration of Ray for distributed computing and Kubeflow Pipelines (KFP) for orchestration within a Kubernetes ecosystem.

The core problem tackled is the sheer scale and complexity of data preparation for modern generative AI models. These workflows involve numerous stages—from deduplication and language filtering to PII detection and tokenization—operating on terabytes of data. The talk provides a comprehensive overview of how Cubray, an operator for running Ray on Kubernetes, and KFP can be leveraged to build resilient, reproducible, and highly scalable data pipelines, ultimately streamlining the MLOps lifecycle.

The presentation introduces the Data Preparation Kit (DPK), an open-source initiative designed to standardize and simplify these complex data engineering tasks. DPK abstracts away the underlying infrastructure complexities, allowing data scientists and model developers to focus on the data transformations themselves rather than the intricacies of distributed systems or Kubernetes. This initiative represents a significant contribution to making large-scale AI data preparation more accessible and efficient for the broader community.

Background

▶ Watch: Introduction and talk agenda overview (0:00)

The journey to developing powerful generative AI models, particularly Large Language Models (LLMs), begins with meticulously prepared data. This preparation, often termed data engineering for foundation models, is a notoriously complex and resource-intensive process. As highlighted by the speakers, these workflows typically commence with accessing vast data corpuses, often stored in Parquet files using the Arrow tables format. A fundamental challenge with such large datasets is the pervasive presence of duplication or semi-duplication, necessitating robust deduplication strategies as an initial and crucial step.

Beyond deduplication, the process frequently involves language separation and filtering, enabling subsequent steps to be tailored to specific linguistic contexts. A critical phase also includes the application of annotated transformers. These specialized components perform vital quality and compliance checks, such as identifying and redacting Personally Identifiable Information (PII), detecting hate speech, or assessing overall document quality. These transformers are often designed to be mutually independent, allowing for parallel execution and writing results into separate data columns, thereby optimizing processing time. The workflow typically culminates in further filtering and tokenization, preparing the data for model ingestion. The inherent structure of these pipelines often allows for various forms of parallelism—either through independent processing stages that merge later or by branching out language-specific processing after initial language separation.

The primary hurdle faced by the IBM and Red Hat teams was bridging the gap between local development and production-scale execution. Data scientists require the flexibility to rapidly iterate and test transformations on smaller datasets (megabytes to gigabytes) on their laptops, while production environments demand the ability to scale seamlessly to terabyte-scale operations in the cloud. Initial approaches, such as running simple Python scripts locally or using basic Kubernetes jobs, proved insufficient. Local scripts lacked resilience and scalability, while raw Kubernetes jobs, though offering some orchestration, presented a steep learning curve and lacked the user-friendly interfaces necessary for data scientists. This led to the adoption of Ray for distributed processing and Cubray to integrate Ray with Kubernetes, coupled with Kubeflow Pipelines (KFP) for orchestrating these intricate, long-running ETL (Extract, Transform, Load) jobs within a cloud-native framework.

Key Findings

▶ Watch: Scaling data workflows with Ray and CubeRay (3:00)

The talk presented several significant findings and contributions that collectively advance the state of foundation model data engineering on Kubernetes:

  1. Demonstrated Scalability with Cubray and KFP: The speakers showcased a highly effective architecture for scaling complex data pre-processing workflows for LLMs. A standout example was a deduplication task handling nearly 8.5 billion documents, totaling 23 terabytes of compressed storage. This task utilized an impressive 7500 CPU cores and 56 terabytes of RAM, completing in approximately 40 hours, demonstrating the practical viability of Cubray and KFP for extreme-scale data challenges. The task successfully reduced documents and storage by 33-40%.
  1. Modular and Reproducible Pipeline Architecture: By leveraging Kubeflow Pipelines, the team developed a modular approach where workflows are broken down into smaller, reusable components. These "simple pipelines" could be chained together to form "multi-step" or "super pipelines." This modularity significantly eased debugging, enhanced reproducibility (as KFP runs retain input parameters and output artifacts for extended periods), and provided robust visualization capabilities through the KFP UI.
  1. Enhanced Developer Experience for Data Scientists: A core achievement was simplifying the interaction with underlying infrastructure. The Cubray API server was introduced to abstract away direct YAML manipulation, allowing users to make API requests directly. More broadly, KFP's UI and features like Elra's drag-and-drop interface made complex DAG creation accessible to data scientists, enabling them to focus on data problems rather than Kubernetes intricacies. This shift allowed data scientists to concentrate on high-value tasks, with operators managing the pipeline execution.
  1. Introduction of the Data Preparation Kit (DPK): A major contribution is the open-source Data Preparation Kit (DPK) project. DPK was specifically designed to standardize and simplify data processing implementations, addressing issues of proprietary scripts and inconsistent development practices. It abstracts common tasks like S3 storage access, metadata generation, and parameter management, and critically, it wraps the underlying execution frameworks (Ray, Spark, Python). This means model developers no longer require deep knowledge of Ray or Spark internals, significantly lowering the barrier to entry for creating new data transformers.
  1. Production Validation and Community Adoption: DPK has been rigorously tested and successfully deployed in production, notably used for the creation of IBM Granite LLMs. Its utility and robustness are further underscored by its recent integration into the Linux Foundation Data and AI community, signaling broader industry recognition and potential for collaborative growth.

Technical Deep Dive

▶ Watch: Real-world CubeRay deduplication at massive scale (6:00)

The technical architecture presented by Asthana and Roytman is a sophisticated blend of distributed computing, workflow orchestration, and a custom abstraction layer, all built upon Kubernetes.

Data Pre-processing Workflow Stages

The typical data pre-processing workflow for foundation models, as outlined, follows a sequential yet parallelizable structure:

  1. Data Corpus Access: Initial data resides in Parquet files using the Arrow tables format, optimized for columnar storage and efficient querying.
  2. Deduplication: This critical first step addresses redundant data. It can involve exact deduplication or more nuanced semi-exact deduplication to identify and remove near-duplicate entries, which are common in large web-scraped datasets.
  3. Language Separation and Filtering: Datasets often contain multiple languages. This stage identifies the language of each document and allows for filtering or branching subsequent processing based on specific language requirements.
  4. Annotated Transformers: These are specialized, often independent processing units. Examples include:
  • PII (Personally Identifiable Information) detection: Identifying and masking sensitive data.
  • Hate/Abuse/Profanity language detection: Flagging or filtering inappropriate content.
  • Document quality checks: Assessing the relevance or coherence of text.

These transformers are designed to be mutually independent, meaning they can execute in any order or in parallel, writing their results into separate columns without interfering with each other.

  1. Filtering and Tokenization: The final stages involve applying further filters based on annotations (e.g., removing low-quality documents or those with detected PII) and tokenizing the text, converting it into a sequence of tokens suitable for LLM training.

A key aspect of this workflow is its inherent parallelism. Beyond the independent annotated transformers, the pipeline can split after language separation, allowing subsequent steps to run independently for each identified language (e.g., English, Japanese, French), significantly accelerating the overall process.

Ray and Cubray for Distributed Processing

For the heavy lifting of distributed data processing, the team opted for Ray, an open-source framework that provides a simple, universal API for building distributed applications. To run Ray effectively on Kubernetes, they utilized Cubray.

  • Cubray Operator: Cubray acts as a Kubernetes operator, bringing core Ray concepts directly to the Kubernetes environment. This includes managing Ray clusters, submitting Ray jobs, and deploying Ray Serve applications.
  • Cubray API Server: Recognizing that direct YAML manipulation for Custom Resources (CRs) is cumbersome for many users, the team integrated the Cubray API server. This component allows users to make standard API requests, which are then translated into the corresponding YAML objects and acted upon by the Cubray operator. This significantly improves the user experience, making Ray on Kubernetes more accessible.
  • Driver-Worker Paradigm: All Ray-based transformations follow a driver-worker paradigm. A central driver reads input file names (or object names in the case of S3 storage) and dispatches processing tasks to workers, which are implemented as Ray actors. Unlike Spark, which uses fixed data partitions, Ray workers dynamically request the next available file from the driver upon completing a task. This dynamic allocation prevents slowdowns when processing files of significantly varying sizes, ensuring efficient resource utilization.
  • Isolated Clusters: For different data processing tasks, the team employs separate Ray clusters. This isolation is crucial for creating task-specific container images, preventing dependency conflicts (e.g., between different Python library versions), and even supporting legacy libraries or Java models within specific tasks.
  • Distributed Network Load: Each worker reads and writes data independently. This inherently distributes the network load across multiple Kubernetes pods and nodes, preventing bottlenecks and enhancing overall throughput.

A testament to this architecture's power was the deduplication task involving 8.5 billion documents. This required a Ray cluster configured with 7500 CPU cores and 56 terabytes of RAM, running for approximately 40 hours.

Kubeflow Pipelines (KFP) for Orchestration

To orchestrate these complex, long-running ETL jobs, Kubeflow Pipelines (KFP) was chosen. KFP is a platform designed for deploying and managing end-to-end machine learning (MLOps) workflows on Kubernetes.

  • Component Creation: KFP offers flexible ways to define pipeline steps, referred to as components:
  • Python decorated components: Easy to get started with, allowing Python functions to be directly converted into pipeline steps.
  • Custom containerized components: For more complex logic or specific language requirements (e.g., Java code), users can package their logic into custom Docker containers.
  • User Interface (UI): KFP provides a rich UI for visualizing Directed Acyclic Graphs (DAGs), monitoring pipeline runs, inspecting input parameters, and reviewing logs. This visual interface is invaluable for users less familiar with Kubernetes command-line tools.
  • Elra for DAG Creation: For users who prefer a graphical approach, Elra offers a drag-and-drop interface for creating and configuring complex KFP DAGs, specifying inputs, outputs, and execution sequences.
  • Key Benefits:
  • Modularity: Breaking down workflows into smaller, reusable components simplifies debugging and promotes reusability.
  • Reproducibility: KFP automatically logs and persists details of each run, including input parameters and output artifacts, for days, weeks, or months. This allows data scientists to track changes over time and easily reproduce past experiments.
  • Visualization: The UI's ability to visualize pipelines, track experiments, and troubleshoot issues provides significant operational clarity.
  • Operator-Data Scientist Separation: By providing a robust and user-friendly orchestration layer, KFP enables a clear separation of concerns: data scientists can focus on model development and data transformation logic, while operators manage the execution and monitoring of pipelines.

Pipeline Architecture in KFP

The team implemented a specific pipeline pattern within KFP to achieve modularity and reusability:

  • Simple Pipelines: Each fundamental data pre-processing step (e.g., deduplication, language separation) is encapsulated as a "simple pipeline." These typically consist of three core, reusable KFP components:
  1. Argument Preparation: Dynamically processes arguments needed at runtime.
  2. Start Ray Cluster: A KFP component that deploys a Ray cluster.
  3. Execute Job: Submits the specific data processing task to the deployed Ray cluster.
  4. Stop Ray Cluster: An exit handler component that guarantees the Ray cluster is undeployed, regardless of whether the job succeeded or failed (similar to a try...finally block).

The only difference between simple pipelines for various steps lies in the specific arguments provided.

  • Multi-step / Super Pipelines: To automate entire end-to-end data preparation processes, "super pipelines" are constructed. In these, each step is itself a "nested simple pipeline."
  • KFP v1 Implementation: Since KFP v1 does not natively support nested pipelines, the team used the KFPS SDK to execute pre-installed simple pipelines as steps within a super pipeline. This results in N+1 runs on the KFP dashboard for an N-step super pipeline (one run for the super pipeline itself, and one for each nested simple pipeline).
  • KFP v2 (Native Nested Pipelines): The presentation also demonstrated KFP v2, which offers native support for nested pipelines, providing a cleaner, single-run view where nested steps are visible within the main pipeline's UI.

Data Preparation Kit (DPK)

The culmination of this work is the Data Preparation Kit (DPK), an open-source project designed to standardize and simplify data engineering tasks for LLMs.

  • Addressing Proprietary Scripts: DPK emerged from the challenge of disparate, proprietary Python scripts developed by individual data scientists, which created onboarding difficulties and inconsistencies.
  • Infrastructure Abstraction: DPK abstracts away common infrastructure concerns:
  • Access to S3 storage.
  • Generation of metadata.
  • Unified shared parameters.
  • Framework Agnostic Wrapper: Crucially, DPK wraps underlying execution frameworks (Ray, Spark, local Python). This means model developers can implement new data transformers without needing deep expertise in the specifics of distributed computing frameworks. They can write their logic, and DPK handles the execution context.
  • KFP Automation: DPK seamlessly integrates with KFP, allowing the standardized models to be easily automated within pipelines.
  • Production Use and Community: DPK has been successfully used in production for the creation of IBM Granite LLMs and recently joined the Linux Foundation Data and AI community, indicating its growing relevance and adoption. It supports various runtimes (Python, Spark, Ray, KFP) and is applicable not only for finetuning data preparation but also for Retrieval-Augmented Generation (RAG) use cases.

Demo / Proof of Concept

▶ Watch: Orchestrating long-running ETL jobs with Kubeflow Pipelines (6:30)

The core of the demonstration showcased a multi-step data preparation pipeline orchestrated through Kubeflow Pipelines, illustrating both KFP v1 and KFP v2 capabilities.

The demonstrated pipeline included the following sequential and parallel stages:

  1. Document Identification: An initial step to process and identify documents.
  2. Exact Deduplication: Removing exact duplicates from the dataset.
  3. Language Separation and Filtering: Identifying languages and preparing for language-specific processing.
  4. Annotated Transforms for Document Quality: Applying quality checks.
  5. Language-Specific Sub-Pipelines: After language separation, the pipeline branched into three parallel sub-pipelines for different languages: English, Japanese, and French. Each of these language-specific branches would then proceed with further processing, such as filtering and quality annotation.

The demo visually contrasted the execution on KFP v1 versus KFP v2:

  • KFP v1 Execution: When the "super pipeline" was executed on KFP v1, the UI showed multiple distinct runs appearing on the dashboard. For example, after the super pipeline's run started, individual runs for "document identification," "exact deduplication," and "language identification" would sequentially appear and complete. Then, separate runs for "Japanese filter," "document quality for Japanese," and similar steps for English and French would become visible. This illustrated the "N+1 runs" pattern where each nested simple pipeline resulted in its own visible run on the KFP v1 dashboard.
  • KFP v2 Execution: In contrast, the same super pipeline executed on KFP v2 presented a much cleaner, single-run view. Within this single run, the UI natively displayed the nested structure of the pipeline, allowing users to drill down into the status of individual steps and sub-pipelines from a unified interface. This highlighted the improved user experience and native nested pipeline support in KFP v2.

The demonstration also briefly touched upon KFP's reproducibility features, showing how a previous run could be "cloned" to execute it again, preserving the historical context and parameters. This feature is crucial for tracking experiments and understanding changes over time in data processing outcomes. The visual progression of steps, the completion status, and the branching for different languages clearly illustrated the power and flexibility of KFP for managing complex data engineering workflows at scale.

Defensive Implications

▶ Watch: Elyra: simplified drag-and-drop for Kubeflow Pipelines (9:00)

While the talk primarily focuses on data engineering and MLOps efficiency rather than cybersecurity, the implications for building robust and trustworthy AI systems are significant. The meticulous approach to data preparation, coupled with strong orchestration and standardization, inherently contributes to a more defensible AI posture.

  1. Data Quality and Integrity: The emphasis on deduplication, PII filtering, and document quality annotations directly addresses critical aspects of data integrity and ethical AI. By systematically removing redundant, sensitive, or low-quality data, the framework helps prevent models from being trained on compromised or biased inputs. This reduces the risk of data poisoning attacks or the propagation of harmful biases, which are major defensive concerns in AI.
  1. Privacy and Compliance: The inclusion of PII detection and hate speech filtering as explicit stages within the pipeline is a direct measure to ensure data privacy and compliance with regulations (e.g., GDPR, CCPA). By automating these checks, organizations can significantly reduce the risk of inadvertently exposing sensitive information or generating harmful content through their LLMs.
  1. Reproducibility and Auditability: Kubeflow Pipelines' ability to log and retain details of every run, including input parameters and output artifacts, provides a robust audit trail. In a security incident or when investigating model misbehavior, this reproducibility is invaluable. Defenders can trace back the exact data lineage, transformations, and configurations that led to a specific model state or output, which is crucial for root cause analysis and compliance audits.
  1. Reduced Attack Surface through Standardization: The Data Preparation Kit (DPK) standardizes data processing logic and abstracts infrastructure complexities. This consistency reduces the likelihood of individual, proprietary scripts introducing vulnerabilities or inconsistencies. A standardized, well-tested framework is generally easier to secure and monitor than a fragmented collection of ad-hoc solutions.
  1. Operational Resilience: The use of Cubray for managing isolated Ray clusters and KFP's exit handlers to guarantee resource cleanup contributes to operational resilience. Well-managed and robust infrastructure is less prone to outages or resource exhaustion, which can be exploited by attackers or lead to data corruption.

In essence, by ensuring the data used to train LLMs is clean, compliant, and processed through a transparent and reproducible pipeline, the presented framework lays a strong foundation for building more secure, ethical, and resilient AI models.

Key Takeaways

  • Large-scale LLM data preparation is complex and resource-intensive, requiring specialized tools for distributed processing and orchestration across multiple stages like deduplication, language separation, and quality annotation.
  • Cubray effectively scales Ray workloads on Kubernetes, enabling the execution of massive data processing tasks (e.g., 8.5 billion documents, 23TB, 7500 CPU cores) with dynamic worker allocation for efficient resource utilization.
  • Kubeflow Pipelines (KFP) provides critical orchestration, modularity, and reproducibility for MLOps workflows, offering a user-friendly UI, clear run history, and the ability to break down complex tasks into reusable components.
  • Sophisticated pipeline patterns can be implemented in KFP, including "super pipelines" composed of nested "simple pipelines," leveraging KFPS SDK for KFP v1 or native support in KFP v2 for managing intricate, multi-step data engineering processes.
  • The Data Preparation Kit (DPK) simplifies and standardizes data processing, abstracting away infrastructure details (Ray, Spark, S3) for model developers, enabling consistent development of new transformers and accelerating LLM creation.
  • DPK is an open-source, production-proven solution that has been used for IBM Granite LLMs and recently joined the Linux Foundation Data and AI community, underscoring its relevance and potential for broader adoption in the AI ecosystem.

About the Speaker(s)

Anish Asthana is an Engineering Manager with the OpenShift AI group at Red Hat. He brings approximately seven years of experience in the cloud-native AI space, focusing on technologies and solutions that bridge AI/ML workloads with Kubernetes and cloud infrastructure.

Alexey Roytman works for IBM Research and has a distinguished career spanning around 25 years. His extensive experience includes working with IBM middleware, cloud infrastructure, and cloud solutions. Currently, his primary focus is on data processing automation, particularly in the context of large-scale data preparation for AI models. This presentation marked Alexey's first time giving a talk at a conference.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk delivers a robust, production-validated architectural blueprint for tackling the immense challenges of foundation model data engineering on Kubernetes. By expertly integrating Ray, Cubray, and Kubeflow Pipelines, the speakers demonstrate a scalable and reproducible workflow capable of processing terabytes of data and billions of documents. The introduction of the open-source Data Preparation Kit (DPK) further solidifies its value, offering a standardized, framework-agnostic solution that significantly lowers the barrier to entry for complex data transformations, making this a critical contribution to the MLOps community.

Heather Calloway (CISO) — STRONG ACCEPT

This talk, while deeply technical in its focus on scaling generative AI data pre-training on Kubernetes, delivers critical insights for security leaders concerned with AI governance and risk. The detailed approach to managing massive datasets through Cubray, Kubeflow Pipelines, and the Data Preparation Kit (DPK) directly addresses challenges around data integrity, PII protection, and auditability. It provides a robust, engineering-driven framework for building trustworthy AI models, demonstrating how operational rigor in data pipelines translates into a stronger compliance posture and reduced business exposure for organizations leveraging large language models.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025