Slinky: Slurm in Kubernetes, Performant AI and HPC Workload Management in Kubernetes - Tim Wickberg

Tim Wickberg

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In this KubeCon EU talk, Tim Wickberg, CTO of SKTMD, introduced Slinky, an ambitious open-source project designed to bridge the long-standing gap between traditional High-Performance Computing (HPC) environments managed by SLURM and the modern cloud-native ecosystem orchestrated by Kubernetes. The presentation highlighted the divergent philosophies and capabilities of these two powerful systems, particularly in their approach to resource management and workload scheduling for demanding AI/ML and scientific simulations.

Watch on YouTube

Visual summary for Slinky: Slurm in Kubernetes, Performant AI and HPC Workload Management in Kubernetes - Tim Wickberg by Tim Wickberg
Visual summary for Slinky: Slurm in Kubernetes, Performant AI and HPC Workload Management in Kubernetes - Tim Wickberg by Tim Wickberg

Key moments

  1. 0:00 Introduction to Slinky and SLURM's capabilities
  2. 2:35 SKETMD introduces Slinky: Slurm in Kubernetes
  3. 3:10 Slinky's Slurm Operator for managing clusters
  4. 3:58 Slurm Bridge: Kubernetes scheduling plugin for Slurm
  5. 5:00 Converging HPC and Cloud Native with Slurm
  6. 6:00 Understanding HPC's finite resources, infinite demand model

Slinky: Slurm in Kubernetes, Performant AI and HPC Workload Management in Kubernetes

Speakers: Tim Wickberg, Chief Technical Officer, SKTMD

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=gvp2uTilwrY

Overview

In this KubeCon EU talk, Tim Wickberg, CTO of SKTMD, introduced Slinky, an ambitious open-source project designed to bridge the long-standing gap between traditional High-Performance Computing (HPC) environments managed by SLURM and the modern cloud-native ecosystem orchestrated by Kubernetes. The presentation highlighted the divergent philosophies and capabilities of these two powerful systems, particularly in their approach to resource management and workload scheduling for demanding AI/ML and scientific simulations.

Wickberg detailed how Slinky aims to bring SLURM's unparalleled scheduling prowess—proven across the majority of the world's Top500 supercomputers—into Kubernetes, offering solutions for managing complex, multi-node, and time-sensitive batch workloads more efficiently. The project comprises two primary components: the Slurm Operator, which deploys and manages SLURM clusters within Kubernetes, and the more innovative Slurm Bridge, a Kubernetes scheduling plugin that allows SLURM to directly schedule Kubernetes pods. This integration is crucial for organizations grappling with the operational overhead of maintaining separate infrastructure stacks for their AI/ML training (often on SLURM) and inference (typically on Kubernetes) workloads.

The talk underscored the critical need for convergence, driven by the explosive growth of AI/ML, which demands both the elasticity of cloud-native platforms and the specialized, high-performance scheduling of HPC. Slinky represents SKTMD's strategic initiative to enable a hybrid environment where the strengths of both systems can be leveraged, offering a path to more performant and unified management of advanced computing workloads.

Background

▶ Watch: Introduction to Slinky and SLURM's capabilities (0:00)

The landscape of modern computing is broadly divided into two major paradigms: High-Performance Computing (HPC) and cloud native. While both aim to efficiently manage and execute workloads, their foundational assumptions and historical evolutions have led to vastly different approaches to job scheduling and resource management. This divergence creates significant challenges for organizations that need to run both traditional HPC-style simulations and modern microservices-based applications, particularly in the rapidly evolving field of AI/ML.

SLURM stands as the de facto workload manager for HPC, governing 60-70% of the world's Top500 supercomputers and a substantial portion of AI/ML training workloads. Its design is predicated on the assumption of finite resources within a fixed system scale but infinite workload demand. HPC systems, once procured and built, have a defined capacity. Researchers, however, continuously generate an insatiable demand for compute time. This core assumption has shaped SLURM's sophisticated capabilities, including incredibly complex priority schemes, fair share algorithms that consider job size, shape, and duration, detailed accounting systems with fine-grained limits, and the ability to schedule jobs across thousands of CPUs simultaneously in lockstep. HPC jobs are typically run with explicit time limits, a critical feature for effective resource planning and backfilling. Furthermore, HPC systems are generally statically defined, with procurement and deployment cycles spanning years, contrasting sharply with the dynamic nature of cloud environments.

Conversely, cloud native orchestration, epitomized by Kubernetes, emerged primarily to manage microservices. Its underlying assumption is one of nearly infinite resources (via cloud provider APIs) but finite workload demand. Kubernetes pods are generally expected to run simultaneously, and workloads are designed to scale horizontally by running additional pods and load balancing. Tightly coupled processes across multi-node compute workloads were not a core design element, requiring extensions for effective management. Pods are also expected to run indefinitely by default, lacking the inherent time limits common in HPC. Capacity issues are typically resolved by provisioning more resources from a cloud provider. However, cloud-native environments excel in application resilience and dynamic resource management, aspects where traditional HPC has historically been less focused. Notable differences include Kubernetes' affinity anti-affinity patterns, which lack direct equivalents in HPC scheduling.

The problem, as Wickberg articulates, arises when organizations need to perform both AI/ML training workloads (which often demand SLURM's HPC-grade scheduling for large-scale, tightly coupled tasks) and AI/ML inference workloads (which benefit from Kubernetes' elasticity and microservices management). Maintaining two separate control planes—one for SLURM and one for Kubernetes—is operationally awkward, inefficient, and costly, especially when trying to run both on the same hardware. The Slinky project, developed by SKTMD, is a direct response to this challenge, aiming to bridge these two worlds by integrating SLURM's powerful scheduling capabilities directly into the Kubernetes ecosystem, thereby converging the management of these diverse, yet often co-dependent, workloads.

Key Findings

▶ Watch: Slinky's Slurm Operator for managing clusters (3:10)

The Slinky project, developed by SKTMD, presents a toolkit of open-source components designed to integrate SLURM into Kubernetes, addressing the challenges of managing performant AI and HPC workloads in a cloud-native environment. The key findings and contributions of the Slinky project are broken down into three major components: the Slurm Operator, the Slurm Bridge, and a suite of associated tooling.

  1. Slurm Operator: Managing SLURM Clusters within Kubernetes
  • Purpose: The Slurm Operator enables the deployment and management of traditional SLURM clusters underneath Kubernetes. This means Kubernetes is used as the infrastructure layer to provision and manage the lifecycle of SLURM components, while SLURM itself handles the scheduling of its native jobs.
  • Architecture: Compute nodes in this model map directly to Kubernetes pods, each running an individual slurmd process (SLURM's equivalent of a kubelet). The operator supports autoscaling of these SLURM clusters based on utilization metrics, leveraging a Prometheus exporter for data collection.
  • Workload Execution: SLURM jobs, typically batch scripts, run natively within these pods. Kubernetes remains unaware of these internal SLURM jobs, allowing SLURM to apply its sophisticated fine-grain resource limits, backfill scheduling, and network topology management without interference.
  • Release Status: Initial release v0.1.0 in November, v0.2.0 released a week before the KubeCon talk, with v0.3.0 planned for June.
  1. Slurm Bridge: Integrating SLURM Scheduling into Kubernetes
  • Purpose: The Slurm Bridge is a Kubernetes scheduling plugin designed to bring SLURM's advanced scheduling algorithms directly to Kubernetes pods and workflows. This is the more innovative component, allowing SLURM's unique capabilities (e.g., efficient multi-node scheduling, future system state planning, network topology management for devices like NVLink) to benefit Kubernetes-native workloads.
  • Architecture: The Bridge translates Kubernetes workload resource requirements into corresponding SLURM placeholder jobs. It can reconstruct multi-node Kubernetes workloads (like pod groups, job sets, or leader worker sets) into single SLURM jobs for optimal scheduling. It also handles device plugins (e.g., GPU, DRRA ecosystem) and filters out nodes/pods not intended for SLURM management.
  • Key Challenge & Limitation: A significant current limitation is that the Kubernetes API lacks a mechanism to subdivide node resources exclusively for different scheduling plugins. Thus, the Slurm Bridge currently assumes exclusive ownership of resources on nodes assigned to it, preventing safe mixing of SLURM and Kubernetes workloads on the same node simultaneously due to potential CPU contention and jitter. Future work aims to address this with better CPU allocation info in the Kubernetes API.
  • Release Status: Planned for release in June, gated on new features in SLURM 25505, and currently in early access with select customers.
  1. Associated Tooling: Enhancing the Slinky Ecosystem
  • Client Library: A Golang client library, generated using an OpenAPI generator, interacts with SLURM's REST API, making it easier for developers to integrate with SLURM.
  • Prometheus Exporter: This exporter interacts with the SLURM REST API to publish metrics, specifically tuned to support the Slurm Operator's autoscaling capabilities.
  • Helm Charts and Container Images: Standard cloud-native deployment artifacts are provided to facilitate easy adoption and deployment of Slinky components.

In essence, Slinky offers a dual-pronged approach: one to run traditional HPC within Kubernetes, and another to infuse Kubernetes with HPC-grade scheduling intelligence, thereby creating a more unified and performant environment for demanding AI/ML and scientific workloads.

Technical Deep Dive

▶ Watch: Slurm Bridge: Kubernetes scheduling plugin for Slurm (3:58)

Slinky's architecture is designed for flexibility and robust integration, addressing two distinct, yet complementary, use cases: running SLURM clusters within Kubernetes and enabling SLURM to schedule Kubernetes workloads. This is achieved through the Slurm Operator and the Slurm Bridge, respectively, supported by a suite of associated tooling. All Slinky components are open source under the Apache 2 license.

Slurm Operator: SLURM as a Kubernetes Workload

The Slurm Operator is a Kubernetes-native solution for deploying and managing SLURM clusters. Its primary function is to treat SLURM itself as a workload orchestrated by Kubernetes.

  • Architecture:
  • The operator interacts directly with the kube API to manage Kubernetes resources.
  • SLURM's control plane—consisting primarily of the Slurm controld (SLURM controller) and the Slurm DBD (database for accounting)—runs as Kubernetes pods.
  • The actual compute nodes of the SLURM cluster are represented by Kubernetes pods, each running an individual slurmd process. These slurmd pods are the workhorses, akin to the kubelet in Kubernetes, responsible for executing SLURM jobs.
  • Communication between the Slurm Operator and SLURM's control plane is exclusively through SLURM's REST API.
  • A managed MARB instance (likely a database like MySQL, MariaDB, or PostgreSQL) stores SLURM's accounting data.
  • Resource Management: The operator leverages Kubernetes Custom Resources (CRs), specifically a cluster CR and node set CRs, to define the SLURM cluster configuration.
  • Autoscaling: The operator integrates with a Prometheus exporter that collects metrics from SLURM's REST API. These metrics, fine-tuned for utilization, drive autoscaling decisions. For instance, if using KEDA (Kubernetes Event-driven Autoscaling), the operator can scale the number of slurmd compute pods from zero to meet the demand of queued SLURM jobs.
  • Workload Execution: When a SLURM job is submitted to this Kubernetes-managed SLURM cluster, SLURM handles its entire lifecycle. Kubernetes is not directly involved in scheduling or managing these individual SLURM batch scripts. This allows SLURM to fully utilize its advanced capabilities such as fine-grain resource limits, backfill scheduling, and efficient network topology management (e.g., for NVLink) without Kubernetes interference.
  • Pod Sizing: A notable characteristic is that the slurmd pods may be generously proportioned (e.g., 20-40 GB images) because they often need to contain a fully featured Linux distribution and various HPC software stacks, given the slower adoption of containers for HPC applications.

Slurm Bridge: SLURM as a Kubernetes Scheduler

The Slurm Bridge is a more innovative component, designed as a Kubernetes scheduling plugin that allows SLURM's sophisticated scheduling logic to directly manage and place Kubernetes pods. This is crucial for large-scale, multi-node batch workloads that Kubernetes' default scheduler struggles with.

  • Integration with Kubernetes Scheduling Framework:
  • The Bridge operates within the Kubernetes scheduling framework, taking advantage of the ability to provision multiple schedulers. It consists of two main parts: a scheduler plugin and a workload controller.
  • The scheduler plugin intercepts pod scheduling requests for specified namespaces (defined in a scheduling profile).
  • It translates the resource requirements of Kubernetes pods into corresponding SLURM placeholder jobs.
  • For multi-node Kubernetes workloads (e.g., those defined by pod group, job set, or future leader worker set CRs), the Bridge intelligently reconstructs these into a single multi-node SLURM job, allowing SLURM to schedule them cohesively.
  • It also handles device plugins like GPU and integrates with the DRRA ecosystem, translating Kubernetes resource requests into SLURM-understandable formats.
  • Once SLURM makes a placement decision for a placeholder job, the workload controller translates this back into Kubernetes bind API calls to instruct Kubernetes to launch the actual pods on the selected nodes.
  • Resource Translation and Management:
  • The Bridge aims to leverage SLURM's advanced capabilities like planning around future system state and detailed network topology management (e.g., for NVLink) for Kubernetes workloads.
  • It filters out nodes that SLURM is not intended to manage and, crucially, filters out pods that SLURM should not manage (e.g., Daemon sets, core control plane components, or even the Slurm Operator's own control plane to avoid a Catch-22).
  • Resource Isolation Challenge: A significant architectural constraint currently is the lack of a Kubernetes API mechanism to subdivide a node's resources exclusively for different scheduling plugins. Therefore, the Slurm Bridge assumes exclusive ownership of resources on nodes assigned to it. This means that while multiple SLURM jobs or multiple Kubernetes pods can run on a node, mixing SLURM jobs and Kubernetes pods simultaneously on the same node is not safely supported. The primary reason is that SLURM workloads often rely on being pinned to specific CPU sets, and Kubernetes pods, without explicit CPU allocation info in the API, could interfere, introducing undesirable jitter into high-performance SLURM workloads. Addressing this is a key area for future work.
  • Flexibility: The design allows for flexible deployments, where portions of a cluster can be exclusively Kubernetes-managed, exclusively SLURM-managed (via the Slurm Operator), or managed by the Slurm Bridge for overlapping workloads.

Associated Tooling

  • Client Library: A GoLang client library generated from SLURM's OpenAPI specification provides an easy-to-use interface for interacting with SLURM's REST API, facilitating custom integrations.
  • Prometheus Exporter: Dedicated to exposing SLURM metrics for monitoring and autoscaling, particularly useful for the Slurm Operator.
  • Helm Charts and Container Images: These standard cloud-native deployment artifacts simplify the deployment and management of all Slinky components within a Kubernetes environment.

In summary, Slinky provides a comprehensive framework for merging HPC and cloud-native paradigms, offering both the capability to run SLURM as a Kubernetes workload and to extend Kubernetes' scheduling capabilities with SLURM's HPC-grade intelligence.

Demo / Proof of Concept

▶ Watch: Converging HPC and Cloud Native with Slurm (5:00)

Tim Wickberg presented a series of demonstration screenshots to illustrate the functionality of the Slurm Bridge, showcasing how it translates Kubernetes pod requests into SLURM jobs and manages their execution. The demos focused on the core capability of scheduling both single and multi-node Kubernetes workloads using SLURM's underlying scheduling logic.

  1. Single Pod Scheduling:
  • Scenario: A single Kubernetes pod is applied to the cluster.
  • Translation: The Slurm Bridge intercepts this pod request and translates it into a SLURM placeholder job.
  • Scheduling: SLURM's scheduler then processes this placeholder job alongside any native SLURM workloads. In the demo, the SLURM sq command (monitoring the SLURM queue) showed the job being placed on a node named slurmbridge one.
  • Execution & Feedback: The Slurm Bridge's workload controller received SLURM's placement decision and communicated it back to Kubernetes via bind API calls, instructing Kubernetes to launch the pod on slurmbridge one.
  • Verification: The running pod displayed annotations indicating its SLURM-managed status, including a Slurm node annotation showing its placement and, importantly, a Slurm job ID label that cross-tied the Kubernetes pod to its corresponding SLURM placeholder job within SLURM's control plane.
  1. Multi-Node Pod Group Scheduling:
  • Scenario: A two-pod Daemon set (or a pod group defined through a replica set or direct enumeration) is applied, requiring two distinct compute nodes.
  • Translation: The Slurm Bridge intelligently translates these two interdependent pods into a single SLURM job that explicitly requests two nodes. This is a crucial capability, allowing SLURM to treat the multi-node Kubernetes workload as a cohesive entity.
  • Scheduling & Resource Contention: In a scenario with limited resources (e.g., only three compute nodes, where two jobs each want two nodes), SLURM's scheduler correctly identifies that both jobs cannot run simultaneously. It queues the second job, preventing partial scheduling and ensuring that the entire multi-node workload is allocated sufficient resources before execution.
  • Execution & Completion: Once the first workload completed and freed up resources, the second two-node pod group was scheduled and executed on slurmbridge one and slurmbridge two.
  • Lifecycle Management: The demo also highlighted that if a Kubernetes pod terminates, the Slurm Bridge's workload controller ensures the corresponding SLURM placeholder job is terminated. Conversely, if a SLURM job (whether native or a placeholder) is terminated or hits a time limit, the corresponding Kubernetes pods are also deleted, ensuring consistent state across both control planes.

These demonstrations effectively showcased the Slurm Bridge's ability to seamlessly integrate SLURM's robust scheduling for complex, multi-node workloads directly into the Kubernetes environment, translating requests and managing their lifecycle across the two distinct orchestration systems.

Defensive Implications

▶ Watch: Understanding HPC's finite resources, infinite demand model (6:00)

The integration of SLURM and Kubernetes through Slinky introduces new considerations for system administrators and security professionals, particularly in managing resource allocation, workload prioritization, and operational complexities in a hybrid environment.

  1. Resource Allocation and Isolation:
  • Exclusive Node Ownership (Slurm Bridge): A critical implication is the Slurm Bridge's current assumption of exclusive resource ownership on nodes it manages. Defenders must ensure that nodes designated for Slurm Bridge-scheduled Kubernetes workloads (or native SLURM jobs) are not simultaneously targeted by the default Kubernetes scheduler for other pods. Misconfiguration could lead to resource contention, performance degradation (e.g., jitter for high-performance SLURM workloads), or instability. This necessitates careful node labeling, taints, and tolerations to prevent inadvertent mixing.
  • CPU Pinning and cgroups: The discussion around CPU allocation info in the Kubernetes API and the challenges with cgroups v1 vs. cgroups v2 highlights a potential area of vulnerability. If SLURM workloads require strict CPU pinning for performance, and Kubernetes pods are not properly constrained, an attacker could potentially launch noisy neighbor pods that interfere with critical HPC or AI/ML training, impacting job integrity or completion times. Implementing robust CPU management features is crucial for true mixed-mode operation.
  • Large Container Images: The Slurm Operator's use of potentially large (20-40 GB) slurmd container images for HPC workloads implies larger attack surfaces if these images are not meticulously hardened. Defenders must ensure these images are regularly scanned for vulnerabilities, kept up-to-date, and only contain necessary components.
  1. Workload Prioritization and Fair Share:
  • SLURM's Complex Priority Schemes: Slinky leverages SLURM's sophisticated priority schemes and fair share mechanisms. Administrators need a deep understanding of how these are configured to ensure critical workloads (e.g., high-priority AI/ML training) receive the necessary resources. Misconfigured priorities could lead to less critical jobs monopolizing resources.
  • Dynamic Reconfiguration: The ability to dynamically adjust SLURM's partition priorities (queues) allows for on-the-fly shifting of resource focus between Kubernetes-bridged workloads and native SLURM jobs. While powerful, this also requires careful access control and monitoring to prevent unauthorized or accidental changes that could disrupt operations.
  1. Unified Control Plane Security:
  • API Security: Both the Kubernetes API and SLURM's REST API are central to Slinky's operation. Robust authentication, authorization, and network segmentation must be in place for both. The Slurm Operator communicates exclusively via SLURM's REST API, making its security paramount.
  • Identity Management: Ensuring consistent and secure identity management across Kubernetes and SLURM for users submitting jobs is vital. The Slurm DBD plays a crucial role in granular resource limits and accounting, which must be protected against tampering.
  • Observability and Monitoring: The use of a Prometheus exporter for autoscaling is a positive step. Defenders should ensure comprehensive monitoring of both Kubernetes and SLURM metrics to detect anomalies, resource bottlenecks, or potential misuse. Alerts should be configured for unusual scaling events or job failures.
  1. Hybrid Environment Complexity:
  • Operational Overhead: While Slinky aims to reduce the overhead of running separate control planes, managing a hybrid environment still introduces complexity. Operators must be proficient in both Kubernetes and SLURM concepts. Comprehensive documentation (like that on slinky.skemd.com) and training are essential.
  • DRRA and GPU Management: The integration with DRRA and GPU device plugins requires careful configuration to ensure secure and efficient allocation of these high-value resources. Misconfiguration could lead to resource leakage or unauthorized access to sensitive hardware.

In summary, adopting Slinky offers significant performance and operational benefits but demands a comprehensive security posture that accounts for the unique challenges of integrating two historically distinct and complex orchestration systems. Careful planning, robust configuration, and continuous monitoring are key to securely leveraging this powerful hybrid solution.

Key Takeaways

  • Slinky bridges HPC and Cloud Native: The project offers a dual approach to integrate SLURM's powerful HPC workload management into Kubernetes, addressing the operational complexities of managing separate control planes for AI/ML training and inference workloads.
  • Slurm Operator for "SLURM in Kubernetes": This component allows Kubernetes to manage the infrastructure and lifecycle of SLURM clusters, treating SLURM itself as a workload. It enables autoscaling of SLURM compute nodes based on demand, using a Prometheus exporter and KEDA.
  • Slurm Bridge for "SLURM as a Kubernetes Scheduler": This innovative scheduling plugin allows SLURM to directly schedule Kubernetes pods, bringing HPC-grade capabilities like efficient multi-node scheduling, future system state planning, and network topology management (e.g., NVLink) to Kubernetes workloads.
  • Multi-Node Workload Optimization: The Slurm Bridge can translate complex Kubernetes pod groups or job sets into single, cohesive SLURM jobs, ensuring that multi-node applications are scheduled and managed as atomic units, preventing partial allocations.
  • Current Resource Isolation Challenge: A key limitation is the inability to safely mix SLURM and Kubernetes workloads on the same node simultaneously due to the lack of fine-grained CPU allocation info in the Kubernetes API, which can lead to jitter for performance-sensitive HPC applications. This is an active area for future development.
  • Open Source and Extensible: Slinky is open source under the Apache 2 license, providing Helm charts, container images, and a Golang client library for easy adoption and further development, with ongoing work to enhance capabilities like DRRA and CPU management.

About the Speaker(s)

Tim Wickberg is the Chief Technical Officer for SKTMD. SKTMD are the original developers of SLURM, having spun off from Lawrence Livermore National Lab in 2012 to support SLURM's rapid adoption in the HPC industry. As CTO, Tim leads the technical direction for SKTMD, including the development of the Slinky project, which integrates SLURM with Kubernetes for performant AI and HPC workload management. His expertise lies at the intersection of traditional HPC scheduling and modern cloud-native orchestration.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk introduces Slinky, an open-source project from the original SLURM developers, SKTMD, designed to integrate SLURM's HPC-grade workload management with Kubernetes. It offers a dual approach: a Slurm Operator for deploying and managing SLURM clusters within Kubernetes, and a more innovative Slurm Bridge that acts as a Kubernetes scheduling plugin, allowing SLURM to directly schedule multi-node Kubernetes workloads. The project addresses a significant operational challenge in converging AI/ML training and inference, demonstrating solid technical depth and a clear path towards unified, performant management, despite an acknowledged limitation in simultaneous mixed-workload resource…

Heather Calloway (CISO) — STRONG ACCEPT

This talk introduces Slinky, a critical project for organizations grappling with the convergence of HPC and cloud-native environments for AI/ML workloads. By integrating SLURM's powerful scheduling with Kubernetes, Slinky offers a path to more efficient resource utilization and management. While the presentation is deeply technical, its implications for governance, operational risk, and the security posture of high-value compute infrastructure are substantial, requiring careful consideration of resource isolation, API security, and consistent identity management.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025