AI Workload Preemption in a Multi-Cluster Scheduling System at Bloomberg - Leon Zhou & Wei-Cheng Lai
Leon Zhou, Wei-Cheng Lai
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk, presented by Wei-Cheng Lai and Leon Zhou from Bloomberg, delves into the sophisticated strategies employed by the financial technology giant to manage and prioritize thousands of AI training jobs across a multitude of Kubernetes clusters. The core focus is on AI workload preemption within a multi-cluster scheduling system, specifically leveraging the open-source orchestration tool Carmada. The speakers articulate Bloomberg's journey in building a robust, highly available, and efficient data science platform that not only scales to meet immense computational demands but also ensures critical AI workloads receive the necessary resources promptly, thereby enhancing system reliability and business agility.

Key moments
- 0:00 Introduction to AI Workload Preemption & Agenda
- 2:00 How Bloomberg uses AI to enhance its products
- 4:00 Bloomberg's AI Training Infrastructure on Kubernetes
- 4:20 Leveraging Carmada for multi-cluster Kubernetes orchestration
- 6:00 Understanding Carmada's core concepts and APIs
- 7:40 Resource allocation challenges with growing AI workloads
- 8:30 Establishing priority for critical AI workloads
AI Workload Preemption in a Multi-Cluster Scheduling System at Bloomberg - Leon Zhou & Wei-Cheng Lai
Speakers: Leon Zhou, Wei-Cheng Lai
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=LrL5AcS2d5g
Overview
This talk, presented by Wei-Cheng Lai and Leon Zhou from Bloomberg, delves into the sophisticated strategies employed by the financial technology giant to manage and prioritize thousands of AI training jobs across a multitude of Kubernetes clusters. The core focus is on AI workload preemption within a multi-cluster scheduling system, specifically leveraging the open-source orchestration tool Carmada. The speakers articulate Bloomberg's journey in building a robust, highly available, and efficient data science platform that not only scales to meet immense computational demands but also ensures critical AI workloads receive the necessary resources promptly, thereby enhancing system reliability and business agility.
The presentation highlights the critical role AI plays in Bloomberg's operations, from parsing financial documents to powering personalized search and generating actionable market signals. As demand for these AI capabilities escalated, the company faced significant resource allocation challenges, particularly with GPU-intensive training workloads. This necessitated the development of a priority-based scheduling and preemption mechanism to guarantee that urgent, high-impact tasks are completed without delay, even when resources are scarce. The speakers detail the technical implementation within Carmada, the considerations for balancing platform efficiency with user experience, and the lessons learned in orchestrating AI training at an enterprise scale.
The insights shared are particularly pertinent for organizations grappling with similar challenges in managing large-scale, resource-intensive AI/ML workloads across distributed Kubernetes environments. By outlining their architectural choices, the open-source tools they integrate, and the custom features they've developed, Bloomberg provides a blueprint for achieving sophisticated workload management that directly aligns with business priorities. The talk underscores the importance of intelligent scheduling and preemption in transforming raw compute power into tangible business value, especially in time-sensitive financial contexts.
Background
▶ Watch: Introduction to AI Workload Preemption & Agenda (0:00)
Bloomberg, a global financial technology leader, relies heavily on AI and machine learning to process vast streams of data, delivering timely and accountable insights to its clients. AI applications at Bloomberg span critical functions such as extracting information from financial documents, enhancing data discoverability, powering personalized search, summarizing news and research, and analyzing data trends to generate strategic investment signals. These capabilities are fundamental to the Bloomberg Terminal, a widely used platform by financial professionals for real-time market data, news, and analytics.
To support these advanced AI capabilities, Bloomberg operates a robust data science platform. This comprehensive solution is deeply integrated with various open-source projects, forming the backbone of their machine learning operations. Kubernetes serves as the primary infrastructure for managing GPU-intensive workloads, complemented by tools like Jupyter notebooks, Qflow, KubeFlow, and Argo Workflow for model development, training, deployment, and automated workflows. A central component of this platform, and the focus of the talk, is Bloomberg's on-premises, bare-metal Kubernetes infrastructure, specifically optimized for GPU-intensive AI training. This infrastructure handles tens of thousands of containerized AI training jobs daily, ranging from simple experiments to complex distributed tasks.
To ensure continuous availability and reliability at such a scale, Bloomberg developed a highly available multi-cluster scheduling system utilizing Carmada. Carmada is an open-source multi-cluster Kubernetes orchestration tool designed to provide a unified management layer over multiple Kubernetes clusters. Its key benefits include:
- Seamless Workload Shifting: Workloads can automatically migrate between clusters in case of downtime or resource limitations, ensuring uninterrupted operations.
- Intelligent Workload Distribution: Carmada efficiently distributes workloads across clusters, preventing resource bottlenecks and underutilization, which translates to improved operational efficiency and cost savings for GPU resources.
- Centralized Configuration Management: It simplifies managing configuration settings and security credentials across multiple clusters, reducing manual effort and enhancing system reliability.
The speakers clarified Carmada's core concepts essential for understanding their scheduling system:
- Resource Template: The standard Kubernetes workload definition (e.g., deployment, custom resource).
- Propagation Policy: Outlines scheduling rules, including cluster affinity, multi-cluster splitting, and which clusters will host the workload.
- Override Policy: Allows customized configurations for specific clusters (e.g., different Docker images or resource limits) without redefining the entire workload.
- Resource Binding: An internal Carmada object that tracks where workloads have been assigned, linking abstract workload definitions to concrete scheduling decisions.
- Work Object: Represents the actual workload deployed on each member cluster, created from resource bindings and override policies, ensuring consistent synchronization.
Despite the sophisticated infrastructure, Bloomberg faced significant resource allocation challenges as AI workload demand grew. High GPU demand from tasks like urgent bond pricing model retraining and earnings-related document analysis frequently contended with available resources, leading to competition. This highlighted the critical need to effectively prioritize workloads based on their business urgency. Bloomberg categorized its AI workloads into priority levels:
- Urgent (High Priority): Mission-critical tasks like retraining a bond pricing model during market volatility. Delays have huge business impact, and acceptable delay tolerance is very low.
- Medium Priority: Scheduled model refreshes for news summarization updates or regular document analysis training. Important for maintaining accuracy and user experience, but can tolerate some delay (e.g., a few hours).
- Low Priority: Research and development experiments. Exploratory runs that can wait days or weeks if clusters are busy with more urgent work.
The traditional first-come, first-serve (FCFS) scheduling approach, common in early Kubernetes and Carmada versions, proved inadequate. It often delayed critical workloads when resources became scarce, failing to align resource allocation with business needs. This fundamental challenge necessitated the development of a priority-based scheduling and preemption mechanism to ensure urgent tasks are always prioritized and completed promptly.
Key Findings
▶ Watch: Bloomberg's AI Training Infrastructure on Kubernetes (4:00)
The primary discovery and contribution presented by Bloomberg is the successful implementation of a priority-based scheduling and preemption mechanism within Carmada, specifically tailored for their multi-cluster AI training infrastructure. This system addresses the critical need to ensure that high-priority, business-critical AI jobs receive the necessary GPU resources, even at the expense of lower-priority tasks.
Key findings and contributions include:
- Necessity of Priority and Preemption: The traditional FCFS scheduling model was found insufficient for managing diverse AI workloads with varying business impacts. A system combining priority-based scheduling with preemption (also referred to as eviction) is essential to guarantee critical workloads progress without delay, especially under resource constraints.
- Carmada Feature Gates: To enable this advanced functionality, Carmada introduced two new feature gates:
priority based schedulingandpriority based preemptive scheduling. While currently defaulting tofalse, Bloomberg plans to enable them by default in future releases after gathering further user feedback and refining the features. - Essential API Changes: Carmada's API was extended to support multi-cluster priority and preemption. Each resource binding associated with a workload now includes a
schedule priorityfield. This field contains: - An integer
priorityvalue. - A
preemption policysetting, which can be eitherpreemptionornever preempt. - Flexible Priority Sourcing: The
priorityvalue can be sourced from one of three locations, offering flexibility to users: - A federated priority class defined in the Carmada control plane.
- A standard Kubernetes priority class.
- Directly from the workload's custom resource definition (CRD) spec.
- Priority Class Immutability for Consistency: To prevent confusion and disruption, and to maintain consistency across Carmada components that rely on stored priority values, updating a priority class is restricted. Users can only update the
descriptionfield. Any changes to thepriorityvalue itself require deleting and recreating the priority class with the same name. - Preemption as a Last Resort with Minimization Strategy: Preemption or eviction of lower-priority bindings is explicitly considered a last resort. The system's primary goal is to preempt as infrequently as possible. Preemption only activates when a high-priority workload has no available resources, and evicting selected lower-priority jobs will explicitly free up sufficient resources. When preemption is necessary, the scheduler aims to minimize disruption by evicting the fewest possible workloads. For example, it prefers evicting a single large job consuming two GPUs over multiple smaller jobs each using one GPU, if that single eviction fulfills the resource requirement.
- Single Scheduler for Preemption: To prevent instability and inefficient "round-robin evictions" (where multiple schedulers repeatedly preempt each other's workloads), the design mandates that only a single Carmada scheduler (or any other scheduler) should perform preemption decisions. This is particularly crucial for machine learning jobs, which often have lengthy cold start times.
These findings demonstrate a sophisticated approach to managing resource contention in dynamic, multi-cluster AI environments, ensuring that business-critical operations are not hampered by resource scarcity.
Technical Deep Dive
▶ Watch: Leveraging Carmada for multi-cluster Kubernetes orchestration (4:20)
Bloomberg's implementation of AI workload preemption within Carmada represents a significant technical advancement for multi-cluster Kubernetes environments. The system is designed to intelligently manage GPU resources, balancing the immediate needs of high-priority tasks with the overall efficiency and stability of the platform.
The foundation of this system lies in Carmada's ability to orchestrate workloads across multiple Kubernetes clusters as a single, cohesive unit. When a user submits an AI training job, it's defined via a Resource Template and associated with a Propagation Policy that specifies cluster affinities and other scheduling rules. If cluster-specific adjustments are needed, an Override Policy can be applied. Internally, Carmada creates Resource Bindings to track workload assignments and then generates Work Objects for deployment on the member clusters.
To introduce priority and preemption, Carmada's API was extended. The Resource Binding object now includes a schedule priority field. This field is crucial, as it dictates the workload's importance and its preemption behavior. It comprises two key attributes:
- Priority Value: An integer that represents the workload's priority level. Higher integer values typically denote higher priority.
- Preemption Policy: This setting determines whether a workload can be preempted. It can be set to
preemption(meaning the workload can be preempted) ornever preempt(meaning it should not be preempted). This allows for granular control over which workloads are eligible for eviction.
The source of the priority value is flexible, accommodating different organizational structures and operational models. It can be derived from:
- A Federated Priority Class, which is a Carmada-specific construct for defining priority classes across the federated clusters.
- A standard Kubernetes Priority Class, allowing integration with existing Kubernetes priority mechanisms.
- Directly from the workload's Custom Resource Definition (CRD) spec, providing a direct way for users to specify priority at the workload level.
A critical design decision for managing priority classes in Carmada is the approach to updates. To ensure consistency and prevent unexpected behavior, Carmada components rely on stored priority values and preemption policies rather than live references to priority class objects. This means that only the description field of a priority class can be updated. If a change to the actual priority value is necessary, the priority class must be deleted and recreated with the same name. This immutability for core values ensures that scheduling decisions are made based on stable, predictable data.
The preemption mechanism itself is engineered with a strong emphasis on minimizing disruption, treating eviction as a measure of last resort. The logic is as follows:
- Trigger Condition: Preemption is only considered when a high-priority workload is waiting for resources and there are currently no available resources in the target clusters.
- Resource Sufficiency Check: The scheduler identifies lower-priority workloads that, if evicted, would collectively free up sufficient resources for the high-priority job to start. This isn't just about freeing some resources, but enough to satisfy the pending job's requirements.
- Minimal Disruption Strategy: Among the candidates for eviction, the scheduler prioritizes evicting the fewest possible workloads. For instance, if a high-priority job needs 2 GPUs and there are two lower-priority jobs each using 1 GPU, and one lower-priority job using 2 GPUs, the scheduler will prefer to evict the single 2-GPU job. This strategy reduces the overall impact on ongoing work and improves user experience.
- Single Preemption Controller: A crucial architectural decision is to enforce that only a single Carmada scheduler (or any other scheduler instance responsible for preemption) is enabled to perform preemption decisions at any given time. Allowing multiple schedulers to preempt independently can lead to "round-robin evictions," where schedulers repeatedly evict each other's workloads, creating instability and significantly impacting the lengthy cold start times often associated with machine learning jobs.
The end-to-end workflow for an urgent AI job, such as a bond pricing model, integrates these features seamlessly. A user submits the workload, specifying target cluster affinities and priorities within the propagation policy. Carmada then places this high-priority workload at the very front of its priority-based scheduling queue. If resources are constrained, Carmada's preemption logic, adhering to the principles of minimal disruption, identifies and evicts lower-priority jobs to ensure the critical task runs smoothly.
Bloomberg's approach to balancing platform efficiency and user experience is also a key technical insight. While maximizing GPU utilization and reducing resource fragmentation across clusters is a core goal, they acknowledge that complex scheduling decisions and frequent preemptions can lead to user frustration. To mitigate this, they emphasize fairness by minimizing disruptions, provide transparency through clear visibility into expected queuing times, and proactively notify users when their jobs are preempted. Additionally, they offer practical recommendations, such as checkpointing training jobs, to help users recover from potential preemptions more gracefully. This holistic approach ensures that advanced scheduling mechanisms deliver both technical performance and a positive user experience.
Demo / Proof of Concept
▶ Watch: Resource allocation challenges with growing AI workloads (7:40)
The talk primarily focuses on the architectural design, implementation details, and operational experience of Bloomberg's multi-cluster AI workload preemption system, which is described as being in active development and refinement within their production environment. While the speakers walk through the end-to-end workflow and illustrate the decision-making process for preemption with a diagram, they do not explicitly mention or demonstrate a live proof-of-concept or a specific demo of the system in action. The emphasis is on the practical application and lessons learned from deploying these features at scale within Bloomberg's infrastructure.
Defensive Implications
▶ Watch: Establishing priority for critical AI workloads (8:30)
The detailed technical insights shared by Bloomberg offer crucial defensive implications for organizations managing large-scale AI/ML workloads, particularly within multi-cluster Kubernetes environments. These implications extend to both platform engineers responsible for infrastructure and data scientists/ML developers utilizing these platforms.
For Platform Engineers and SREs:
- Implement Priority-Based Scheduling: It is imperative to move beyond first-come, first-serve scheduling for critical AI workloads. Platform engineers should integrate priority-based scheduling mechanisms into their orchestration layers (e.g., Kubernetes schedulers, multi-cluster orchestrators like Carmada) to align resource allocation with business urgency.
- Strategic Preemption Capability: When resources are scarce, a well-designed preemption mechanism is vital. This involves:
- Conditional Preemption: Only preempting when a higher-priority job genuinely lacks resources and eviction will provide sufficient capacity.
- Minimizing Disruption: Implementing logic to evict the fewest possible lower-priority workloads to satisfy the high-priority job's needs. This reduces overall system churn and improves user experience.
- Controlled Preemption: Ensuring that only a single scheduler instance is responsible for preemption decisions across the multi-cluster environment to prevent inefficient "round-robin evictions" and maintain stability, especially for long-running ML jobs.
- Leverage Multi-Cluster Orchestration: For organizations operating multiple Kubernetes clusters, tools like Carmada are essential for unified management, intelligent workload distribution, and failover capabilities. Integrating priority and preemption directly into such orchestrators ensures consistent behavior across the entire fleet.
- API and Configuration Design: When extending orchestration tools, design APIs (like Carmada's
schedule priorityfield) to be flexible (e.g., multiple sources for priority values) yet robust (e.g., immutability of priority values in class definitions) to prevent misconfigurations and ensure predictable behavior. - Transparency and Communication: Provide clear visibility into scheduling decisions. This includes displaying expected queuing times for jobs and proactively notifying users when their workloads are preempted. Transparency builds trust and helps users understand system behavior.
- Resource Management and Optimization: Prioritization and preemption directly contribute to better resource utilization. By ensuring high-priority jobs get resources, platform engineers can reduce idle GPU time and optimize the overall cost-efficiency of their expensive AI infrastructure.
- Support for Checkpointing: Actively encourage and support the use of checkpointing in ML training jobs. This mitigates the impact of preemption by allowing jobs to resume from their last saved state, reducing wasted computation and user frustration.
For Data Scientists and ML Developers:
- Understand and Utilize Priority Levels: Developers must understand the defined priority levels within their organization and assign them accurately to their workloads. Mislabeling a low-priority experiment as high-priority can disrupt critical business operations.
- Design for Resilience (Checkpointing): Incorporate checkpointing mechanisms into all long-running or critical training jobs. This is the primary defense against preemption, allowing jobs to gracefully recover and resume, minimizing lost progress and compute time.
- Monitor Job Status: Pay attention to platform notifications regarding queuing times and preemption events. Understanding these signals can help in planning and iterating on models more effectively.
- Resource Request Accuracy: While not explicitly mentioned as a preemption factor, accurately requesting resources for jobs can help the scheduler make better decisions, reducing the likelihood of unnecessary preemption or resource starvation.
By implementing these defensive strategies, organizations can build more resilient, efficient, and user-friendly AI infrastructure that consistently meets business demands, even under heavy load.
Key Takeaways
- AI Workload Prioritization is Critical: In large-scale, multi-cluster AI environments, a first-come, first-serve scheduling model is insufficient. Prioritizing AI workloads based on business urgency (e.g., urgent bond pricing vs. R&D experiments) is essential for business agility and ensuring critical tasks complete on time.
- Carmada as an Open-Source Orchestrator: Bloomberg leverages Carmada, an open-source multi-cluster Kubernetes orchestration tool, to manage thousands of GPU-intensive AI training jobs across numerous on-premises clusters, highlighting its capability for unified management and intelligent workload distribution.
- Intelligent Preemption as a Last Resort: A sophisticated preemption mechanism is implemented within Carmada, activated only when high-priority jobs lack resources. This mechanism prioritizes minimizing disruption by evicting the fewest possible lower-priority workloads that free up sufficient resources.
- Consistency and Control through API Design: Carmada's APIs were extended with a
schedule priorityfield, allowing granular control over workload priority and preemption policy (preemptionornever preempt), with priority values sourced from federated or standard Kubernetes priority classes, or directly from CRD specs. - Single Scheduler for Stability: To prevent inefficient and disruptive "round-robin evictions," a crucial design principle is to ensure that only a single scheduler instance is enabled to perform preemption decisions across the multi-cluster environment.
- Balancing Efficiency and User Experience: Beyond technical implementation, Bloomberg emphasizes balancing platform efficiency (high GPU utilization) with a positive user experience. This involves providing transparency (queuing times, preemption notifications) and practical recommendations like checkpointing for ML jobs to mitigate preemption impacts.
About the Speaker(s)
Wei-Cheng Lai and Leon Zhou are key contributors from Bloomberg, involved in the development and management of the company's advanced data science platform and AI training infrastructure. While specific titles were not explicitly stated in the talk, their presentation demonstrates deep expertise in Kubernetes, multi-cluster orchestration, and the challenges of scheduling high-demand AI workloads in a production enterprise environment. Their work focuses on enhancing system reliability, business agility, and optimizing resource utilization for Bloomberg's critical AI capabilities.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk from Bloomberg dissects a critical problem: managing high-priority AI workloads across vast, multi-cluster Kubernetes environments. Leveraging and extending Carmada, they present a sophisticated priority-based scheduling and preemption system. The detailed architectural choices, particularly the emphasis on minimal disruption during preemption and the single scheduler enforcement for stability, offer a robust blueprint for any organization grappling with resource contention for critical ML infrastructure. It's a real-world solution to a real-world problem, executed with technical precision.
Heather Calloway (CISO) — STRONG ACCEPT
This talk from Bloomberg presents a robust solution for managing critical AI workloads through priority-based scheduling and preemption in a multi-cluster Kubernetes environment. It demonstrates how a financial institution translates business urgency into operational controls at the infrastructure level, ensuring that high-impact tasks, such as bond pricing model retraining, are consistently prioritized. While highly technical, the core message around disciplined resource allocation and its direct link to business resilience and accountability makes this a valuable session for security leaders grappling with the operationalization of AI at scale.