Thousands of Virtual Kubelets: 1-to-1 Mapping a Supercomputer To Kubernetes With... Dennis Marttinen

Dennis Marttinen

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

Dennis Marttinen's KubeCon EU talk, "Thousands of Virtual Kubelets: 1-to-1 Mapping a Supercomputer To Kubernetes With... Superetes," delves into the increasingly critical convergence of High Performance Computing (HPC) and cloud environments. Marttinen, a master's student in security and cloud computing, presents Superetes, a novel transparent bridge designed to seamlessly integrate Kubernetes with existing Slurm-managed supercomputers. The talk highlights the growing demand for HPC resources from the artificial intelligence (AI) community and the need to make these powerful systems as accessible, secure, and resilient as cloud platforms.

Watch on YouTube

Visual summary for Thousands of Virtual Kubelets: 1-to-1 Mapping a Supercomputer To Kubernetes With... Dennis Marttinen by Dennis Marttinen
Visual summary for Thousands of Virtual Kubelets: 1-to-1 Mapping a Supercomputer To Kubernetes With... Dennis Marttinen by Dennis Marttinen

Key moments

  1. 0:46 Core differences between Cloud and HPC paradigms
  2. 3:20 Detailed architectural comparison of Kubernetes and Slurm
  3. 5:20 Identifying key fragility and abstraction drawbacks in Slurm
  4. 7:00 Introducing Superetes: a transparent bridge for Cloud-HPC
  5. 7:30 Step-by-step walkthrough of Superetes' core architecture
  6. 8:30 How Superetes creates virtual Kubelets for HPC nodes

Thousands of Virtual Kubelets: 1-to-1 Mapping a Supercomputer To Kubernetes With... Superetes

Speakers: Dennis Marttinen

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=QbR908kgk1Y

Overview

Dennis Marttinen's KubeCon EU talk, "Thousands of Virtual Kubelets: 1-to-1 Mapping a Supercomputer To Kubernetes With... Superetes," delves into the increasingly critical convergence of High Performance Computing (HPC) and cloud environments. Marttinen, a master's student in security and cloud computing, presents Superetes, a novel transparent bridge designed to seamlessly integrate Kubernetes with existing Slurm-managed supercomputers. The talk highlights the growing demand for HPC resources from the artificial intelligence (AI) community and the need to make these powerful systems as accessible, secure, and resilient as cloud platforms.

The core problem Marttinen addresses is the fundamental architectural and operational differences between traditional HPC systems and cloud-native infrastructure, primarily Kubernetes. While both are distributed systems, their design philosophies diverge, leading to challenges in resource management, workload orchestration, and especially security and availability. Superetes offers a pragmatic, "stop-gap" solution by mirroring an entire HPC system into a Kubernetes cluster using virtual Kubelets, enabling cloud-native tools to manage supercomputing workloads.

This integration is vital for the future of European competitiveness, as emphasized by reports encouraging the opening up of HPC capacity to startups, SMEs, and the broader AI community. By demonstrating a working solution on Lumi, the eighth fastest supercomputer in the world, Marttinen not only showcases the technical feasibility of bridging these worlds but also sets the stage for a more ambitious vision: a future where Kubernetes directly manages HPC hardware, delivering the agility, security, and high availability of the cloud to the supercomputing domain.

Background

▶ Watch: Core differences between Cloud and HPC paradigms (0:46)

The foundational difference between cloud computing and HPC lies in their resource assumptions. Cloud platforms typically assume infinite resources and finite demand, allowing for elastic scaling and on-demand provisioning. Conversely, HPC operates under the assumption of finite resources and infinite demand, necessitating meticulous scheduling and optimization for large, often single, workloads that can span an entire system. This distinction historically led to divergent architectural patterns and software stacks.

However, the advent of artificial intelligence (AI) workloads has significantly blurred these lines. Modern AI models, particularly large language models (LLMs), require immense computational power, making cloud requirements increasingly resemble those of HPC. This convergence has highlighted the need for unified platforms that can leverage the strengths of both paradigms. Architecturally, modern HPC and cloud hardware are surprisingly similar, with the primary divergence being in their software stacks. Cloud platforms predominantly rely on Kubernetes for orchestration, while many HPC systems utilize Slurm (Simple Linux Utility for Resource Management).

Despite their different origins, Kubernetes and Slurm share fundamental characteristics as distributed systems. Both feature client interfaces (API server for Kubernetes, bash commands for Slurm), controllers for state mutation, databases for state persistence, and node agents for workload execution. The key difference lies in their communication patterns: Kubernetes centralizes control through its API server, while Slurm's architecture is more distributed and command-driven.

Kubernetes excels in versatility, reconciliation, and self-healing, making it ideal for diverse deployments and resilient operations. Slurm, on the other hand, is optimized for scheduling large-scale, multi-node workloads and expertly managing the complex hardware topologies of supercomputers. However, Slurm exposes a lower-level abstraction compared to Kubernetes, which currently has limited interfaces for fine-grained hardware control. Moreover, the HPC software stack, particularly Slurm, often suffers from fragility and non-seamless software upgrades, leading to significant downtime. These challenges formed the impetus for Marttinen's master's thesis research, aiming to bridge these two ecosystems. His quest began on the Lumi supercomputer in Finland, where he quickly discovered that existing solutions were inadequate, leading to the development of Superetes.

Key Findings

▶ Watch: Identifying key fragility and abstraction drawbacks in Slurm (5:20)

The central finding of Marttinen's research and the development of Superetes is the successful creation of a transparent bridge that connects a Kubernetes environment with a Slurm-managed HPC system. Crucially, Superetes was specifically designed to operate effectively on complex systems like the Lumi supercomputer, where other existing solutions failed. This achievement demonstrates the feasibility of bringing cloud-native orchestration capabilities to high-performance computing.

Superetes' key contributions include:

  1. Bidirectional Synchronization: Unlike most other bridge solutions that primarily push Kubernetes workloads to HPC, Superetes uniquely provides bidirectional synchronization. It not only dispatches Kubernetes pods as Slurm jobs but also observes and reconciles jobs running natively on the HPC system back into the Kubernetes cluster. This ensures a complete and consistent view of the HPC environment within Kubernetes.
  2. Complete Node View via Virtual Kubelets: Superetes exposes a complete, 1:1 mapping of the entire HPC node structure to Kubernetes by creating a virtual Kubelet instance for each physical HPC node. This allows Kubernetes-native schedulers to make informed decisions based on real-time node and pod metrics from the supercomputer, a capability largely absent in other solutions.
  3. Firewall Resilience: HPC systems, especially Lumi, often have stringent firewall rules that pose significant challenges for bridge solutions. Superetes is designed to navigate these restrictions, ensuring reliable communication and synchronization, which is a critical practical finding for real-world deployments.
  4. Identification of Fundamental HPC Challenges: Beyond providing a bridging solution, Marttinen's work highlights deeper, unresolved issues within the HPC ecosystem: the lack of robust secure multi-tenancy (e.g., absence of kernel/user/process namespaces, network isolation) and inadequate high availability (frequent downtime for upgrades and faults). These are critical findings, as existing bridge solutions, including Superetes, cannot fundamentally solve these inherent architectural limitations of the underlying HPC software stack.
  5. A Vision for Native Integration: Ultimately, the research points towards a future where Kubernetes doesn't just bridge to Slurm but potentially replaces it, directly managing HPC hardware. This "final station" architecture is presented as the optimal solution for achieving true cloud-native security, automation, and resilience in supercomputing, moving beyond temporary "stop-gap" bridges.

Technical Deep Dive

▶ Watch: Introducing Superetes: a transparent bridge for Cloud-HPC (7:00)

Superetes operates as a transparent bridge between a Kubernetes cluster and an HPC environment managed by Slurm. Its architecture consists of two primary components: a controller and an agent.

The controller runs as a pod within the Kubernetes cluster. Its role is to manage the lifecycle of virtual Kubelet instances and orchestrate the translation of Kubernetes workloads into Slurm jobs. The agent is a process that runs on an HPC login node – the traditional SSH entry point for users to interact with the supercomputer.

The communication between these two components is critical. The agent initiates an mTLS secured gRPC reverse tunnel to the controller. This reverse tunnel design is crucial for traversing the stringent firewalls commonly found in HPC environments, allowing the Kubernetes cluster to securely reach into the HPC system without requiring incoming connections to the HPC side.

Once the tunnel is established, the controller instructs the agent to discover all available nodes within the HPC environment. For each discovered HPC node (e.g., Lumi G node 1, Lumi G node 2), the controller deploys a corresponding virtual Kubelet instance in the Kubernetes cluster. These virtual Kubelets act as virtual representations of Kubernetes nodes, mirroring the HPC hardware without actually running a full Kubelet daemon on the supercomputer nodes themselves. This 1:1 mapping provides Kubernetes with a comprehensive view of the HPC cluster's topology and resources.

When a user deploys a standard v1/Pod in Kubernetes, the Superetes controller intercepts this request. It then passes the pod definition to the agent, which translates it into a Slurm job using standard sbatch commands. Slurm, with its native capabilities, can then dispatch this job across multiple HPC nodes, potentially splitting it into several tasks that run in parallel.

A unique aspect of Superetes is its handling of workload status and logs. Since Kubernetes pods cannot natively span multiple nodes in the same way Slurm jobs can be split into tasks, Superetes introduces shadow pods. For each task launched by Slurm, a corresponding shadow pod is created in a special namespace within Kubernetes. Superetes then associates these shadow pods with the original user-deployed pod. This mechanism allows for bidirectional synchronization: the status, logs, and other metadata from the Slurm jobs and their tasks are observed by the agent, relayed to the controller, and then reflected in the status of the virtual Kubelet nodes and the associated shadow pods. Ultimately, the user can monitor the status and retrieve logs of their original Kubernetes pod as if it were running natively within the Kubernetes cluster, providing a fully transparent experience.

For consistent synchronization and state management, Superetes leverages and modifies Flux CD, a cloud-native GitOps tool. This integration ensures that the desired state defined in Kubernetes is consistently reconciled with the actual state on the HPC system. The entire mirroring capability of the HPC system into Kubernetes is built upon the foundational Virtual Kubelet project, which provides the framework for creating these virtual node representations.

Furthermore, Superetes goes beyond just job execution. It also observes any other jobs running in the HPC environment (not initiated via Kubernetes) and reconciles their state, creating corresponding virtual Kubelet nodes and shadow pods. This comprehensive observation ensures that Kubernetes schedulers have a complete and accurate view of the HPC cluster's utilization, enabling smarter scheduling decisions.

During the Q&A, a limitation regarding networking was highlighted: Superetes currently has difficulty bridging the networking between Slurm jobs and the Kubernetes cluster. While isolating Slurm jobs into separate network namespaces and then bridging them via a proxy is an obvious theoretical solution, the absence of network namespaces in many HPC systems (like Lumi) means this is not yet implemented. The current best approach would be an application-aware proxy for communication.

Demo / Proof of Concept

▶ Watch: Step-by-step walkthrough of Superetes' core architecture (7:30)

Marttinen presented a live demonstration of Superetes in action, showcasing its ability to seamlessly integrate a Kubernetes cluster with the Lumi supercomputer. The demo environment consisted of a single-node Kubernetes cluster that, through Superetes, managed 102 virtual nodes representing actual Lumi HPC nodes, alongside its single physical node.

The initial view of the Kubernetes cluster displayed kubectl get nodes, revealing 103 nodes in total: 102 virtual nodes exposed by Superetes and the single physical node running the Kubernetes control plane. Marttinen highlighted that the virtual Kubelet nodes also exposed CPU and memory metrics, indicating around 30% CPU utilization and 45% memory utilization across the virtual nodes, though he cautioned against reading too deeply into these specific numbers due to potential wonkiness in Slurm's metric reporting.

For the demonstration, a standard Kubernetes v1/Pod definition was used. This pod was designed to be very lightweight – pulling an Alpine container from DockerHub and running a simple loop that counts to 300 in one-second intervals. To ensure the pod was scheduled on a Superetes-managed virtual node, a toleration was added, and the nodeName was explicitly set to target one of the virtual Lumi nodes. Crucially, Superetes also allows passing arbitrary Slurm options via Kubernetes labels. In this demo, the --time flag was passed to Slurm, allocating sufficient time for the job to complete.

Upon applying the pod definition (kubectl apply -f pod.yaml), the following sequence of events unfolded:

  1. The pod immediately appeared in the Kubernetes cluster with a Pending status.
  2. Almost instantaneously, the Superetes controller picked up the pod, passed it to the agent, and the agent dispatched it as a Slurm job. This was confirmed by the sq command (showing Slurm queue status) running in a loop at the bottom of the screen, where the new job appeared and immediately transitioned to a Running state due to its lightweight nature and available resources.
  3. Due to Kubernetes' poll-based nature (with a 10-second polling interval for Slurm resources, as Slurm doesn't support a "watch" mechanism), there was a brief delay before the Kubernetes status updated.
  4. After approximately 10 seconds, the Kubernetes pod's status changed from Pending to Running. Concurrently, a shadow pod was created in a special namespace, associated with the original job.
  5. Marttinen then demonstrated that users could view the logs of their original Kubernetes pod, seeing the output of the Alpine container's counting loop in real-time, effectively demonstrating the transparent integration and bidirectional synchronization of status and logs.

This proof of concept effectively illustrated how Superetes bridges the operational gap, allowing users familiar with Kubernetes to deploy and monitor workloads on an HPC system as if it were a native Kubernetes cluster, while leveraging Slurm's powerful scheduling capabilities.

Defensive Implications

▶ Watch: How Superetes creates virtual Kubelets for HPC nodes (8:30)

While Superetes and other HPC-to-cloud bridge solutions offer significant operational advantages, Marttinen critically points out that they cannot fundamentally solve two major, inherent issues within traditional HPC environments: secure multi-tenancy and high availability. These limitations pose substantial defensive challenges that require a more radical architectural shift.

Secure Multi-Tenancy: Marttinen uses the analogy of a hotel to explain the multi-tenancy problem on systems like Lumi. While Unix permissions provide "walls between rooms" to prevent direct access to other tenants' data, the "hotel has no locks on the doors." This means:

  • Lack of Isolation: HPC systems typically lack kernel namespaces, user namespaces, and process namespaces. This absence means that tenants (users or teams) are not truly isolated. One tenant can potentially observe the processes of another, akin to transparent walls between hotel rooms.
  • Network Vulnerabilities: Without network namespaces, the "hallway of networking" is open. A tenant can directly access another tenant's running services, such as a Jupyter Lab environment. Through this compromised environment, they could then access confidential data and models stored by the first tenant, bypassing file-level Unix permissions. This represents a significant security vulnerability, as sensitive AI training data or proprietary models could be exposed.

High Availability (HA): Unlike cloud platforms designed for resilience, HPC systems frequently experience downtime due to hardware faults and, more commonly, software upgrades. Marttinen presented a screenshot of Lumi's downtime over a year, illustrating 28 days of continuous downtime. Calculating the cost, with Lumi's total cost of €150 million over a 5-year lifespan (advertised yearly cost of €30 million), one day of downtime costs approximately €82,000. Therefore, 28 days of downtime equates to a staggering €2.3 million in unusable capacity. This starkly contrasts with cloud expectations, where such prolonged outages would be unacceptable. The planned successor to Lumi, with an allocation of €250 million, suggests these costs of downtime will likely persist or even increase.

Kubernetes as a Defensive Solution: Marttinen argues that Kubernetes natively offers robust solutions to these defensive challenges:

  • Isolation: Containers inherently provide opaque walls by isolating processes, while network policies, SPIFFE, and Cilium can install "locks for your rooms," enforcing fine-grained network segmentation and identity-based access control.
  • High Availability: Kubernetes' reconciliation loops and self-healing properties are designed to maintain system uptime through partial failures and upgrades, drastically reducing downtime compared to traditional HPC upgrade cycles.

The long-term defensive implication is that for true secure multi-tenancy and high availability, the architecture needs to evolve beyond mere bridging. Marttinen envisions a future where Slurm itself runs within Kubernetes as pods (leveraging projects like Slinky Slurm Operator, Sunk, or Operator by Nibbius) and, eventually, where Kubernetes directly controls the underlying HPC hardware. This "final station" eliminates the need for hacky bridges and allows the full suite of cloud-native security, reliability, and automation tools to be applied to supercomputing, creating a single, unified, and modern platform for the future.

Key Takeaways

  • HPC-Cloud Convergence is Inevitable: Driven by the escalating demands of AI workloads, the architectural requirements of cloud and HPC are increasingly similar, necessitating integrated solutions.
  • Superetes as a Transparent Bridge: Superetes successfully bridges Kubernetes and Slurm environments, providing bidirectional synchronization of workloads and a complete 1:1 node view of the HPC system within Kubernetes via virtual Kubelets.
  • Existing Bridges Have Limitations: While effective for workload orchestration, current HPC-to-cloud bridge solutions, including Superetes, do not resolve fundamental issues of secure multi-tenancy (e.g., lack of namespaces, network isolation) and high availability within the underlying HPC software stack.
  • HPC Security and HA Require a Paradigm Shift: Traditional HPC setups suffer from inadequate isolation (e.g., no kernel/user/process namespaces, open networking) and significant downtime for software upgrades, leading to substantial security risks and financial costs (e.g., €2.3 million for 28 days of Lumi downtime).
  • Future Vision: Kubernetes as the HPC Orchestrator: The ultimate solution involves running Slurm within Kubernetes as pods and eventually replacing Slurm entirely, with Kubernetes directly managing HPC hardware. This "final station" architecture unlocks native cloud-native security, automation, and high availability for supercomputing.
  • Community Collaboration is Key: Achieving this seamless integration requires both the HPC and cloud-native communities to unify their efforts, leveraging existing open-source projects like Flux CD, Virtual Kubelet, and DRA (Device Resource Allocation) to build a more accessible, secure, and cost-effective future for supercomputing.

About the Speaker(s)

Dennis Marttinen is a master's student specializing in security and cloud computing, pursuing his degrees at both Aalto University in Finland and Nineu in Norway. His career in the cloud-native space began at Weave Works, where he notably co-authored Weave Ignite, a project focused on virtual machine management in Kubernetes. Currently, Marttinen is conducting his master's thesis research with the astroinformatics research group at Aalto University, with a specific focus on the critical area of supercomputer and cloud integration, which directly informed the development of Superetes. His work aims to bridge the gap between these powerful computing paradigms, making high-performance resources more accessible and robust for emerging workloads like artificial intelligence.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk presents Superetes, a novel and technically robust transparent bridge that maps an entire Slurm-managed supercomputer to Kubernetes using virtual Kubelets. Beyond the clever bidirectional synchronization and firewall-resilient implementation, the speaker delivers a brutally honest assessment of traditional HPC's fundamental flaws in secure multi-tenancy and high availability. This isn't just a solution; it's a critical analysis that lays out a compelling vision for Kubernetes as the future orchestrator for supercomputing, offering a clear path to genuine cloud-native security and resilience for critical AI workloads.

Heather Calloway (CISO) — STRONG ACCEPT

Marttinen's presentation on Superetes, while demonstrating a clever technical bridge for HPC-Kubernetes integration, delivers its most critical insights by dissecting the fundamental governance and security limitations of traditional supercomputing environments. He clearly articulates the severe business risks stemming from inadequate multi-tenancy isolation and persistent downtime, quantifying the financial impact. This isn't just a technical demo; it's a crucial strategic call for CISOs and leaders to acknowledge the architectural fragility of current HPC and proactively pursue a cloud-native future for robust security and resilience.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025