Optimizing Training Performance for Large Language Model(LLM) in Kubernetes - Klaus Ma & Peng Gu
Klaus Ma, Peng Gu
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk, presented by Klaus Ma from Nvidia and Peng Gu from an unnamed technical startup, delves into the critical area of optimizing Large Language Model (LLM) training performance within Kubernetes environments. The core focus is on addressing the networking bottlenecks inherent in distributed LLM training, a problem that becomes increasingly pronounced with the scale and complexity of modern AI models. Ma and Gu introduce and demonstrate the network topology aware scheduling capabilities recently integrated into Volcano, a CNCF incubating project and the first batch scheduling system in CNCF.

Key moments
- 0:00 Introduction to speakers and talk topic
- 0:40 Core goals of batch systems and hardware optimization
- 2:00 Volcano: CNCF's first batch scheduling system
- 3:20 Key focus: Networking topological scheduling in Volcano
- 4:10 Understanding Volcano's extensible plugin machinery
- 6:50 Addressing LLM training challenges with Volcano features
- 8:00 Details of networking-aware scheduling (MVLink, InfiniBand)
Optimizing Training Performance for Large Language Model (LLM) in Kubernetes - Klaus Ma & Peng Gu
Speakers: Klaus Ma, Co-founder of Volcano, Nvidia; Peng Gu, Technical Startup
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=JXcQcofGzrA
Overview
This talk, presented by Klaus Ma from Nvidia and Peng Gu from an unnamed technical startup, delves into the critical area of optimizing Large Language Model (LLM) training performance within Kubernetes environments. The core focus is on addressing the networking bottlenecks inherent in distributed LLM training, a problem that becomes increasingly pronounced with the scale and complexity of modern AI models. Ma and Gu introduce and demonstrate the network topology aware scheduling capabilities recently integrated into Volcano, a CNCF incubating project and the first batch scheduling system in CNCF.
The presentation highlights how Volcano's enhanced scheduler intelligently places compute-intensive AI workloads, particularly those leveraging GPUs, to minimize inter-node communication latency and maximize bandwidth utilization. This is achieved through the introduction of a new Custom Resource Definition (CRD) called HyperNode, which allows administrators to define the physical and logical network hierarchy of their data centers. For organizations building and operating large-scale AI infrastructure on Kubernetes, understanding and implementing such topology-aware scheduling is paramount for achieving optimal training speeds, reducing operational costs, and ensuring the efficient use of expensive hardware resources like GPUs and high-speed interconnects.
Background
▶ Watch: Introduction to speakers and talk topic (0:00)
The landscape of modern AI, particularly with the advent of LLMs, presents significant challenges for traditional workload orchestration systems. Training these models demands immense computational power, vast memory, and, critically, ultra-low latency and high-bandwidth networking. In a Kubernetes context, managing these distributed, resource-hungry workloads falls under the purview of batch scheduling systems. Klaus Ma began by emphasizing that a robust batch system must effectively manage workload lifecycles, intelligently schedule tasks, and, most importantly, match the design performance of underlying hardware. This hardware encompasses not just computing resources like GPUs and CPUs, but also intricate networking components (e.g., InfiniBand, Rocky, NVLink) and storage solutions (e.g., caches, distributed storage).
Volcano, a CNCF incubating project since its inception as KubeBatch in 2017, has been at the forefront of addressing these batch scheduling requirements within Kubernetes. Historically, Volcano has focused on enhancing GPU and CPU resource utilization through features like dominator resource fair share, job-level, queue-level, and namespace-level fair sharing. It also previously introduced task topological aware scheduling, primarily for traditional PS-worker (Parameter Server-Worker) models, which yielded significant performance improvements. However, the unique demands of LLM training, especially its extreme sensitivity to network performance, necessitated a deeper focus on networking topological scheduling.
The speakers highlighted that while Kubernetes has offered basic topology awareness through standard labels (e.g., region, zone) since 2019 and the Pod Topology Spread feature since 2020, these mechanisms often fall short for the granular, performance-critical network demands of LLMs. These models often require groups of GPUs to communicate at speeds only achievable when they are physically co-located and connected via high-speed, low-latency interconnects like NVLink or InfiniBand within a single rack or a closely connected set of racks. Without explicit network topology awareness, a standard scheduler might distribute a single LLM training job across disparate nodes, leading to severe performance degradation due to increased network hops and reduced bandwidth.
Volcano's design philosophy, centered around a flexible two-level plug-in machinery (actions and plugins), allows for extensive customization and extension of scheduling logic without requiring changes to the upstream Kubernetes core. This architecture, combined with internal cache enhancements to minimize interactions with the Kubernetes API server (a known bottleneck in high-scale scheduling), provided the ideal foundation for integrating sophisticated network topology awareness. The goal was to empower users and cloud providers to define and leverage their specific data center network layouts to ensure that LLM training jobs are placed optimally from a networking perspective, thereby unlocking the full potential of expensive GPU clusters.
Key Findings
▶ Watch: Volcano: CNCF's first batch scheduling system (2:00)
The central finding and contribution of this talk is the successful implementation and demonstration of network topology aware scheduling within Volcano, specifically tailored for the demanding requirements of Large Language Model (LLM) training in Kubernetes. This represents a significant step forward in optimizing distributed AI workloads.
The key findings and contributions include:
- Introduction of
HyperNodeCRD: To explicitly define and manage network topology within a Kubernetes cluster, Volcano introduces a new Custom Resource Definition (CRD) calledHyperNode. This CRD allows administrators to model the physical network hierarchy of their data centers, including different tiers of switches (e.g., Tier 1 for rack-level, Tier 2 for row-level, Tier 3 for data center-level) and to group nodes based on their network connectivity using selectors. - Granular Network-Aware Scheduling: The enhanced Volcano scheduler can now leverage the
HyperNodedefinitions to make intelligent placement decisions for LLM training jobs. This ensures that all components of a single distributed training job are co-located within the most optimal network boundaries (e.g., within a single rack for NVLink/InfiniBand, or within a row for broader InfiniBand connectivity) to minimize latency and maximize bandwidth. - Unified Network API (Future Work): The project aims to unify the APIs for various high-performance networking technologies such as NVLink, InfiniBand (IB), and RDMA over Converged Ethernet (Rocky). This unification would allow for a more consistent and comprehensive management of network interfaces and resources within the Kubernetes environment, abstracting away underlying hardware specifics.
- Automated Network Discovery (Pipeline): While the initial implementation relies on manual
HyperNodeCRD definition, a crucial future direction is the development of auto-discovery mechanisms for network topology. This would significantly reduce the operational burden on users and cloud providers by automatically mapping the network infrastructure intoHyperNodeobjects. Nvidia's recent DPF (Data Processing Framework) features, including network acceleration, and SNAP (Storage Network Acceleration Platform) were mentioned as relevant technologies that could contribute to this discovery process. - Network Health Monitoring (Pipeline): The team also plans to incorporate network health checking into the
HyperNodeCRD. This would allow the scheduler to consider the stability and readiness of network interfaces when making placement decisions, potentially avoiding nodes with intermittent network issues and improving overall job reliability. - Flexible Scheduling Modes: The
volcanojobspecification now includes anetworkTopologysection with amodefield, supporting bothhardandsoftconstraints. Thehardmode ensures strict adherence to the specified network tier, while thesoftmode allows for escalation to higher (broader) network tiers if resources are constrained in the preferred lower tier, offering a balance between performance optimization and resource availability.
These findings collectively demonstrate a practical and effective approach to overcoming critical networking challenges in large-scale AI training, making Kubernetes a more viable and performant platform for LLM development.
Technical Deep Dive
▶ Watch: Key focus: Networking topological scheduling in Volcano (3:20)
The technical implementation of network topology aware scheduling in Volcano revolves around defining, discovering (future), and utilizing the network hierarchy within a Kubernetes cluster. The core concept is to ensure that distributed AI workloads, particularly LLM training jobs, are placed on nodes that are physically and logically close in the network, thereby optimizing inter-process communication.
At the hardware level, the speakers outlined a typical data center network architecture relevant to GPU-intensive workloads. Within a single rack, nodes (often GPU servers) are typically connected via high-speed, low-latency interconnects like NVLink (for intra-GPU communication within a server) and InfiniBand (IB) or RDMA over Converged Ethernet (Rocky) for inter-node communication within the rack. When communication needs to cross rack boundaries, it typically goes through multiple layers of switches. The talk described this as a multi-tier structure: Tier 1 representing the rack level, Tier 2 grouping multiple racks (e.g., a row), and Tier 3 encompassing a larger segment of the data center. The primary goal of network topology aware scheduling is to minimize the number of times data has to traverse these higher-tier switches, as each hop introduces latency and can reduce effective bandwidth.
To represent this physical network topology within Kubernetes, Volcano introduces a new Custom Resource Definition (CRD) called HyperNode. This CRD is designed to be highly flexible:
In this example, a HyperNode named tier1-rack01 is defined.
tier: tier1indicates that thisHyperNoderepresents the lowest level of network proximity, typically a single rack.type: Nodespecifies that the members of thisHyperNodeare individual Kubernetes nodes.selector: matchRegex: "node-region1-zone1-dc1-row1-rack01-.*"uses a regular expression to dynamically select nodes that belong to this specific rack. This allows for automated grouping of nodes based on their naming conventions, which often encode their physical location.
Multiple HyperNode CRDs can be defined to represent the hierarchical structure. For instance, a tier2 HyperNode could encompass several tier1 HyperNodes (racks) within a "row" or "super-rack," and so on, up to tier3 for a larger logical grouping.
Once the HyperNode hierarchy is defined, Volcano jobs can be configured to leverage this topology information. The volcanojob specification now includes a networkTopology section:
The networkTopology field has two crucial parameters:
mode:
Hardmode: This is a strict constraint. If a job is configured withmode: HardandhighestTierAllowed: tier1, the Volcano scheduler will only place all pods belonging to that job within a singletier1HyperNode(e.g., a single rack). If insufficient resources are available within any singletier1HyperNode, the job will remain in a pending state. This mode is ideal for workloads where consistent, ultra-low latency, and maximum bandwidth are absolutely critical, as is often the case with large-scale LLM training. It guarantees that all communication happens over the fastest possible links, typically NVLink or direct InfiniBand connections.Softmode: This provides a more flexible approach. Ifmode: Softis used, the scheduler will first attempt to place the job within thehighestTierAllowed(or implicitly, the lowest available tier). However, if resources are not sufficient at that level, it will "escalate" and try to find resources in higher-levelHyperNodes(e.g., moving from atier1rack to atier2row). This mode prioritizes job execution over absolute optimal network performance, allowing jobs to run even if it means slightly increased latency due to inter-rack or inter-row communication. ThehighestTierAllowedparameter is less impactful inSoftmode as the scheduler will try to go as high as needed, starting from the lowest possible tier.
The underlying Volcano scheduler integrates this logic into its existing plugin architecture. When a volcanojob with network topology constraints is submitted, the scheduler:
- Identifies the required resources (GPUs, CPUs, memory) and the specified network topology mode and tier.
- Queries its internal cache (which is optimized to minimize API server calls) to get the current state of nodes and their
HyperNodeaffiliations. - Evaluates available
HyperNodegroups that can satisfy the job's resource requirements while adhering to the network topology constraints. - If
Hardmode is specified, it strictly enforces placement within a single, designatedHyperNodetier. IfSoftmode, it attempts hierarchical placement.
Future technical enhancements include:
- Auto-discovery: Moving from manual
HyperNodeCRD creation to automated discovery of network topology. This would involve agents or controllers inspecting the physical network (e.g., using Nvidia DPF or SNAP capabilities) and dynamically creating/updatingHyperNodeobjects. - Unified API for Network Management: Standardizing the way different high-performance network interfaces (NVLink, InfiniBand, Rocky) are exposed and managed within Kubernetes, simplifying resource allocation and monitoring.
- Network Health Checking: Integrating network health status into the
HyperNodedefinition, allowing the scheduler to avoid unstable or underperforming network segments. This would likely involve custom controllers monitoring network metrics and updatingHyperNodestatus.
These technical advancements position Volcano as a powerful scheduler for complex, network-sensitive AI workloads, directly addressing the performance needs of modern LLMs.
Demo / Proof of Concept
▶ Watch: Addressing LLM training challenges with Volcano features (6:50)
Peng Gu presented a live demonstration of Volcano's network topology aware scheduling in action, albeit within a simulated environment due to the ongoing physical build-out of their data center. The demo effectively illustrated how the scheduler intelligently places distributed LLM training jobs to optimize network performance.
The simulation environment was set up using kind (Kubernetes in Docker) and quark to create a cluster of 118 nodes. These nodes were systematically named to reflect a detailed logical network topology: node-region1-zone1-dc1-rowX-rackYY-nodeZZ. This naming convention allowed the HyperNode CRDs to use regular expressions for grouping, mirroring a real-world data center layout.
The logical view of the simulated network architecture was crucial:
- Tier 1 (Rack Level): Represented by Nvidia GB300 MVL72 racks. Each
tier1HyperNodeencompassed 18 nodes, with each node having 4 GPUs and 2 CPUs. This setup is designed to maximize intra-rack communication via high-speed NVLink and InfiniBand. - Tier 2 (Row Level): Groups of
tier1racks were logically organized into "rows." The demo showed two rows, each containing three racks. - Tier 3 (Higher Level): A single, overarching
tier3HyperNodeencompassed all rows.
The demonstration proceeded in two main parts:
Demo 1: Rack-Level Scheduling
- Three separate LLM training jobs were submitted simultaneously.
- Each job was configured to require 18 nodes (equivalent to one full rack) and specified a
Hardmode network topology constraint withhighestTierAllowed: tier1. - Upon scheduling, the
kubectl get pods -o widecommand was used to inspect the placement of the pods for each job. - Result: The demo successfully showed that each of the three training jobs was entirely confined to a single, distinct rack. For example:
training-onepods were all on nodes inrow1-rack02.training-twopods were all on nodes inrow2-rack02.training-threepods were all on nodes inrow2-rack03.
This confirmed that Volcano's scheduler respected the tier1 (rack-level) Hard constraint, ensuring that all 18 nodes for a given job were co-located for optimal high-bandwidth, low-latency communication.
Demo 2: Row-Level Scheduling
- After deleting the first set of jobs, two larger LLM training jobs were submitted.
- Each of these jobs was configured to require 54 nodes (equivalent to three full racks) and specified a
Hardmode network topology constraint withhighestTierAllowed: tier2. - Again, pod placement was verified using
kubectl. - Result: The demo illustrated that each of these larger jobs was successfully confined to a single "row" (Tier 2
HyperNode), which logically contained three racks. For instance:
training-fourpods were all scheduled withinrow1(spanning three racks within that row).training-fivepods were all scheduled withinrow2(spanning three racks within that row).
This further validated the scheduler's ability to understand and enforce higher-level network topology constraints, ensuring that even larger jobs are placed within a cohesive network segment to minimize performance degradation.
Enabling the Feature:
The speakers outlined the steps to enable this feature:
- Install Volcano v1.01.01 with the
network-topology-previewtag. - Manually define
HyperNodeCRDs to map the cluster's network topology. - Configure
volcanojobYAMLs with thenetworkTopologyspec, specifyingmode(hard/soft) andhighestTierAllowed.
The demo provided concrete evidence that Volcano's network topology aware scheduling is a functional and effective solution for optimizing the placement of large-scale, network-intensive AI workloads in Kubernetes.
Defensive Implications
▶ Watch: Details of networking-aware scheduling (MVLink, InfiniBand) (8:00)
The introduction of network topology aware scheduling in Volcano has significant implications for defenders, particularly those responsible for designing, securing, and operating MLOps platforms and GPU-accelerated infrastructure within Kubernetes. While not directly a "security" feature in the traditional sense, its impact on performance, resource utilization, and reliability indirectly contributes to a more robust and defensible system.
- Performance as a Defensive Layer: For LLM training, performance is paramount. Slow training means longer time-to-market for models, higher operational costs (due to longer GPU usage), and potentially wasted resources. By ensuring optimal network proximity, this feature directly boosts training performance. From a defensive standpoint, an efficient system is often a more resilient system, as resources are not unnecessarily strained, reducing the likelihood of cascading failures dueoutsourced to resource starvation.
- Efficient Resource Utilization: GPUs and high-speed network interconnects (InfiniBand, Rocky) are extremely expensive assets. This feature ensures that these resources are utilized to their fullest potential by minimizing network overhead. Defenders can advocate for and implement this to maximize the return on investment for their infrastructure, making the overall platform more cost-effective and sustainable.
- Improved Reliability and Stability: By intelligently placing workloads, the system reduces the risk of network congestion and performance variability that can lead to job failures or inconsistent training outcomes. The future plans for network health checking within the
HyperNodeCRD are particularly impactful here. If the scheduler can avoid nodes with unstable or failing network interfaces, it will inherently improve the reliability of training jobs, reducing the need for costly restarts and manual intervention. - Configuration Complexity and Accuracy: Currently, the manual definition of
HyperNodeCRDs introduces a new layer of configuration complexity. Defenders and platform engineers must ensure these definitions accurately reflect the underlying physical network topology. InaccurateHyperNodedefinitions could lead to suboptimal placements, negating the benefits of the feature, or even causing jobs to remain unschedulable inHardmode. This necessitates robust configuration management and validation processes. - Dynamic Environment Challenges: Data centers are dynamic; nodes fail, are replaced, and new hardware is added. Maintaining the accuracy of
HyperNodedefinitions in such an environment is a challenge. If the topology information becomes stale, the scheduler might make incorrect decisions. The future development of auto-discovery mechanisms is critical to mitigate this operational burden and ensure continuous accuracy, which is a key defensive posture against configuration drift. - Network Visibility and Monitoring: Implementing this feature underscores the importance of deep network visibility and monitoring within the Kubernetes cluster. To effectively define
HyperNodeCRDs and to troubleshoot any performance issues, platform operators need detailed insights into network latency, bandwidth, and traffic patterns at various tiers. This reinforces the need for comprehensive network observability tools. - Cloud Provider Reliance: For users leveraging cloud-based GPU offerings, the ability to fully utilize this feature will heavily rely on cloud providers exposing their underlying network topology or integrating
HyperNodeauto-discovery into their managed Kubernetes services. Without this, users might be limited to more generic topology labels, reducing the potential performance gains.
In essence, while network topology aware scheduling doesn't directly address common security threats like vulnerabilities or intrusions, it empowers platform teams to build a more performant, reliable, and efficient MLOps infrastructure. This, in turn, frees up resources and attention that might otherwise be spent on troubleshooting performance issues, allowing defensive teams to focus on core security challenges.
Key Takeaways
- Volcano's Network Topology Aware Scheduling: A new feature in Volcano (v1.01.01 preview) significantly optimizes Large Language Model (LLM) training performance in Kubernetes by intelligently placing workloads based on network proximity.
HyperNodeCRD for Network Hierarchy: TheHyperNodeCustom Resource Definition (CRD) allows administrators to explicitly define the multi-tier physical and logical network topology of their data centers (e.g., rack-level, row-level), enabling granular control over job placement.- Flexible Scheduling Modes: Volcano jobs can specify
HardorSoftnetwork topology constraints.Hardmode ensures strict co-location for maximum performance, whileSoftmode offers flexibility by allowing escalation to broader network tiers if resources are constrained. - Demonstrated Performance Gains: The demo successfully showed that LLM training jobs can be confined to specific racks (Tier 1) or rows (Tier 2), ensuring high-bandwidth, low-latency communication crucial for distributed GPU workloads.
- Future Enhancements for Automation and Reliability: Upcoming features include automated network topology discovery, a unified API for managing NVLink/InfiniBand/Rocky, and network health checking, which will further enhance usability and job reliability.
- Crucial for Large-Scale AI: This capability is vital for organizations running large-scale, network-intensive AI training, as it maximizes the utilization of expensive GPU and network hardware, leading to faster training times and reduced operational costs.
About the Speaker(s)
Klaus Ma is a prominent figure in the Kubernetes and cloud-native community, particularly known as a co-founder of Volcano. Volcano is a CNCF incubating project and holds the distinction of being the first batch scheduling system within CNCF. Currently, Klaus is with Nvidia, where he has been for approximately three years, focusing specifically on networking aspects within the context of high-performance computing and AI infrastructure. His work at Nvidia and with Volcano reflects a deep expertise in optimizing resource management and scheduling for complex workloads.
Peng Gu is associated with a technical startup, which he noted is currently in "sales mode," hence the company name was not disclosed. His role involves building a GPU cloud specifically tailored for AI training. This background provides him with practical, real-world experience in the challenges and requirements of deploying and managing large-scale GPU infrastructure for AI workloads, making his insights into the practical application of Volcano's features particularly valuable.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk by Klaus Ma and Peng Gu delivers a critical technical solution for a pressing problem in large-scale AI: optimizing LLM training performance in Kubernetes by tackling network bottlenecks. They introduce Volcano's network topology aware scheduling, enabled by the new HyperNode CRD, which allows granular definition of data center network hierarchy. The live demo effectively showcases how this intelligent scheduling ensures distributed LLM jobs are co-located for optimal low-latency, high-bandwidth communication, directly impacting efficiency and cost for anyone serious about MLOps.
Heather Calloway (CISO) — STRONG ACCEPT
This session by Klaus Ma and Peng Gu presents a critical advancement for organizations heavily invested in large language model (LLM) training on Kubernetes. By introducing network topology-aware scheduling in Volcano, the speakers demonstrate a clear path to optimize GPU resource utilization and significantly reduce training times by ensuring workloads are physically co-located. While not a direct security talk, the financial and strategic implications of efficiently managing multi-million dollar AI infrastructure are paramount for business leaders and CISOs, making this a strong operational imperative for any enterprise building serious AI capabilities.