The GPUs on the Bus Go ‘Round and ‘Round - Natalie Bandel & Ryan Hallisey, NVIDIA

Natalie Bandel, Ryan Hallisey, NVIDIA

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In the highly dynamic and resource-intensive world of cloud infrastructure, managing the reliability of specialized hardware like Graphics Processing Units (GPUs) presents a significant operational challenge. This KubeCon EU talk, "The GPUs on the Bus Go ‘Round and ‘Round," by NVIDIA's Natalie Bandel and Ryan Hallisey, delves into the complex problem of GPU failures within massive Kubernetes clusters. Drawing from their extensive experience operating the GeForce Now cloud gaming platform, the speakers illuminate the causes, detection, and automated remediation strategies for keeping tens of thousands of GPUs available and performing optimally.

Watch on YouTube

Visual summary for The GPUs on the Bus Go ‘Round and ‘Round - Natalie Bandel & Ryan Hallisey, NVIDIA by Natalie Bandel, Ryan Hallisey, NVIDIA
Visual summary for The GPUs on the Bus Go ‘Round and ‘Round - Natalie Bandel & Ryan Hallisey, NVIDIA by Natalie Bandel, Ryan Hallisey, NVIDIA

Key moments

  1. 0:00 Introduction and scale of NVIDIA's GPU fleet
  2. 3:10 Explaining "GPU has fallen off the bus" message
  3. 3:40 Common causes of GPU device failures: power, drivers
  4. 6:00 Quantifying daily GPU failures: 120-200 across fleet
  5. 7:00 Automating GPU failure recovery for maximum capacity
  6. 7:50 Custom GPU problem detector using Kubernetes device plugin

The GPUs on the Bus Go ‘Round and ‘Round

Speakers: Natalie Bandel, Ryan Hallisey, NVIDIA

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=cLJRh4y4vXg

Overview

In the highly dynamic and resource-intensive world of cloud infrastructure, managing the reliability of specialized hardware like Graphics Processing Units (GPUs) presents a significant operational challenge. This KubeCon EU talk, "The GPUs on the Bus Go ‘Round and ‘Round," by NVIDIA's Natalie Bandel and Ryan Hallisey, delves into the complex problem of GPU failures within massive Kubernetes clusters. Drawing from their extensive experience operating the GeForce Now cloud gaming platform, the speakers illuminate the causes, detection, and automated remediation strategies for keeping tens of thousands of GPUs available and performing optimally.

The presentation highlights the critical importance of maximizing GPU capacity, given their substantial cost and central role in demanding workloads such as cloud gaming and AI inferencing. NVIDIA's journey from manual troubleshooting to sophisticated automation provides a compelling case study for any organization running GPU-accelerated applications at scale. Furthermore, the talk underscores emerging challenges as GPU densities per node increase, advocating for community-driven solutions within the Kubernetes ecosystem to address these evolving maintenance and device management complexities.

This article provides a detailed technical deep dive into NVIDIA's approach, exploring the mechanisms for detecting and recovering from GPU failures, as well as proactive measures to prevent downtime. It also examines the broader implications for the Kubernetes community, emphasizing the need for standardized solutions to ensure the continued efficient operation of GPU-accelerated infrastructure in the face of escalating demands.

Background

▶ Watch: Introduction and scale of NVIDIA's GPU fleet (0:00)

NVIDIA operates a colossal cloud infrastructure to power its GeForce Now platform, a service that streams high-fidelity graphics to end-users for a desktop-like gaming experience. This infrastructure is geographically distributed, comprising over 40 Kubernetes clusters, managing more than 30,000 nodes, and orchestrating over 60,000 GPUs, typically at a density of two GPUs per node. Beyond cloud gaming, this fleet also handles significant inferencing workloads using spot capacity. The paramount goal is to maintain maximum GPU capacity, ensuring these expensive and critical devices are constantly utilized and available to users.

Despite rigorous engineering, device failures are an inherent reality in such large-scale deployments. The speakers initiated their presentation by asking the audience if they had ever experienced a "GPU falling off the bus," a common colloquialism derived from kernel ring buffer messages indicating a device failure. This phenomenon, where a GPU becomes unresponsive or unusable, can stem from various root causes. Key culprits include overheating, often due to prolonged high utilization, and insufficient power supply, where issues with power delivery infrastructure, cabling, or connections can disrupt GPU operation. Furthermore, driver failures are a frequent source of problems, especially in environments like NVIDIA's that extensively use virtual machines (VMs) and GPU passthrough via the VFIO driver or vGPUs with the NVIDIA driver. Dynamic driver changes and incompatibilities can lead to instability and device unavailability.

The sheer scale of NVIDIA's operations translates these individual device failures into a significant operational burden. The company experiences between 120 and 200 GPU failures over a typical 24-hour period. While this represents a relatively small percentage of their total fleet (approximately 0.6% of 60,000 GPUs failing daily), the need to remediate these issues rapidly and automatically is critical. Each failure means a GPU is not generating revenue or serving users, directly impacting service quality and financial efficiency. Early in their journey, NVIDIA discovered that nodes with failing GPUs often sat idle, unable to accept workloads, leading to underutilized capacity and wasted resources. This realization spurred their efforts to develop sophisticated detection and remediation strategies, moving away from manual interventions that proved unsustainable and unscalable. Ryan Hallisey also noted a previous KubeCon talk, "All your GPUs are belong to us," which covered some foundational aspects of their maintenance automation, indicating a continuous evolution of their strategies.

Key Findings

▶ Watch: Common causes of GPU device failures: power, drivers (3:40)

NVIDIA's extensive experience managing a vast fleet of GPUs in Kubernetes clusters has yielded several critical insights into device failure management and capacity optimization:

  • Prevalence of GPU Failures: Despite robust hardware and software, GPU failures are a constant reality, with NVIDIA observing 120-200 failures daily across its 60,000+ GPUs. This translates to a 0.6% daily failure rate, which, when considering the need to remediate all GPUs on a node, can double the effective remediation rate (e.g., 1.2% for nodes with two GPUs).
  • Rebooting as a Primary Solution: Counterintuitively, a full node reboot is often the most effective and straightforward solution for resolving many GPU failures, particularly those related to driver issues or transient hardware states. It provides a clean slate and reliably resets the device and its associated software stack.
  • Scalability Demands Automation: Manual identification, cordoning, and rebooting of failed nodes are unsustainable and inefficient at NVIDIA's scale. Automation is paramount to achieving rapid response times, reducing operational overhead, and ensuring high capacity utilization.
  • Node Drain is the Bottleneck: The most time-consuming phase of the automated remediation process is draining workloads from a node before a reboot. Data from NVIDIA's Grafana dashboards showed that while many nodes drain within an hour, a significant "long tail" exists, with some nodes taking up to eight hours to clear, directly impacting GPU availability.
  • Remediation Loops Require Deeper Intervention: Continuously rebooting a node that repeatedly experiences GPU failures is unproductive. Detecting "remediation loops" (e.g., more than two reboots within 90 minutes) is crucial to trigger more aggressive recovery actions, such as power cycling or a full node rebuild.
  • Proactive Measures are Essential: While reactive remediation is necessary, preventing failures through continuous monitoring of GPU health (temperature, power), proactive self-tests, and strategic use of Kubernetes taints and tolerations significantly reduces the frequency of disruptive reboots.
  • Increasing GPU Density Amplifies Reboot Costs: The industry trend towards nodes with higher GPU counts (e.g., 8+ GPUs for AI workloads) drastically increases the cost and impact of a single node reboot. A failure affecting one GPU could necessitate rebooting a node with many other healthy, high-value GPUs, leading to substantial capacity loss.
  • Kubernetes Needs Enhanced Device Management: Current Kubernetes scheduling and lifecycle management capabilities are not fully equipped to handle the nuances of single GPU failures on multi-GPU nodes or to intelligently rebalance workloads in these scenarios. There's a clear need for community-driven enhancements.
  • Community Collaboration is Key: Addressing these complex challenges effectively requires collaborative effort. NVIDIA has co-proposed a Node Lifecycle working group within the Kubernetes community to develop standardized solutions for device management, maintenance, and node health.

Technical Deep Dive

▶ Watch: Quantifying daily GPU failures: 120-200 across fleet (6:00)

NVIDIA's strategy for managing GPU failures in their massive Kubernetes clusters is a multi-faceted approach encompassing sophisticated detection, automated remediation workflows, and proactive prevention measures.

Failure Discovery Mechanism

The first critical step is accurately identifying a problematic GPU. NVIDIA employs a two-part discovery mechanism:

  1. Custom GPU Problem Detector: This is an internal NVIDIA solution that extends the standard Kubernetes device plug-in framework. The device plug-in framework is typically used to advertise GPU resources to the Kubernetes scheduler. NVIDIA's custom detector enhances this by actively monitoring for driver failures and other GPU health indicators. When it detects a problem, such as a driver becoming unresponsive (often after a driver change, as their VMs dynamically switch drivers), it marks the affected node by applying a specific condition to it. This condition signals that the node's GPU resources are compromised.
  2. Node Problem Detector (NPD): This is an open-source Kubernetes project designed to monitor node health and report system issues as node conditions and events. NVIDIA leverages NPD to continuously scan all nodes for the custom conditions set by their GPU problem detector. Once NPD identifies a node marked with a GPU failure condition, it triggers the subsequent remediation process. This integration ensures that GPU-specific problems are elevated to the broader node management system, allowing for automated responses.

Automated Remediation Workflow

Upon detection of a GPU failure, NVIDIA initiates an automated remediation workflow designed to restore the node and its GPUs to a healthy, operational state. The core strategy is "reboot everything," but with significant optimizations:

  1. Node Drain: This is the initial and often most time-consuming step. Before a node can be rebooted, all active workloads must be gracefully terminated or migrated. NVIDIA's data, collected via Grafana dashboards, revealed that while many nodes drain within an hour, a "long tail" exists, with some workloads requiring up to eight hours to complete or migrate. To mitigate this bottleneck, NVIDIA implemented several optimizations:
  • Tenant Integration: They integrated with their GeForce Now tenant to intelligently drain spot capacity and pre-warm sessions, which are more amenable to interruption, thus speeding up the process for these specific workload types.
  • Prioritization: Remediation of nodes with device failures is prioritized over other, less urgent maintenance activities within the cluster. This ensures that critical GPU capacity is restored as quickly as possible.
  1. Reboot: Once a node is successfully drained, it undergoes a reboot. This action is surprisingly effective for many GPU issues, as it resets the driver state, clears transient errors, and brings the node back to a fresh, known-good state.
  2. Recovery Workflow and Remediation Loops: A simple reboot isn't always sufficient. NVIDIA has developed a more comprehensive recovery workflow to handle persistent issues:
  • Detection of Remediation Loops: To prevent endless reboots of a persistently failing node, NVIDIA implemented an alert system. If a node undergoes more than two reboots within a 90-minute period without resolving the GPU issue, this triggers a "remediation loop" alert.
  • Escalated Recovery: When a remediation loop is detected, the automated system marks the node for a more aggressive recovery step. The workflow progresses through:
  • Reboot: The initial attempt.
  • Power Cycle: If reboots fail, a full power cycle of the node is attempted, which can resolve deeper hardware-level issues than a software reboot.
  • Rebuild: If power cycles also prove ineffective, the node is automatically marked for a full rebuild. NVIDIA has an automated rebuild process that periodically discovers these marked nodes and re-provisions them, essentially deploying a fresh operating system and software stack.
  • Manual Intervention/RMA: In rare cases where even a rebuild fails, the node is escalated for manual intervention by operations staff or ultimately designated for RMA (Return Merchandise Authorization), indicating a likely hardware defect.

Proactive Prevention and Future Challenges

Beyond reactive remediation, NVIDIA employs proactive measures to minimize GPU failures:

  • GPU Health Monitoring: Continuous monitoring of GPU metrics like temperature, power usage, and detection of anomalies via Grafana dashboards and alerts helps identify potential issues before they escalate to failures. This addresses common root causes like overheating and power supply problems.
  • Node Problem Detector for System Errors: NPD is also used to catch broader kernel issues and system errors that might indirectly impact GPU stability.
  • Proactive Self-tests: Periodic, automated self-tests are run on both nodes and GPUs to verify their health and identify latent problems.
  • Taints and Tolerations: Kubernetes taints and tolerations are used strategically. If a node is known to be unhealthy or scheduled for imminent maintenance/drain, taints can prevent new workloads from being scheduled on it, minimizing disruption.
  • Automatic Sanity Tests Post-Maintenance: After any maintenance or upgrade (e.g., driver updates, Kubernetes component upgrades, OS updates), automated sanity tests are executed. These tests run actual workloads to verify that the node and its GPUs remain healthy and functional, preventing regressions.
  • Regular Updates: Drivers, Kubernetes components, and the operating system are regularly updated. For this, NVIDIA has open-sourced a solution called PA (introduced a couple of months prior to the talk), which assists with these update processes.

Emerging Challenges and Community Collaboration

The increasing trend of configuring nodes with higher GPU densities (e.g., 8+ GPUs per node, especially for AI workloads) introduces new challenges. A single node reboot, which was manageable with 2-4 GPUs, becomes significantly more costly when 8 or 16 high-value GPUs are taken offline. This raises questions about how Kubernetes can evolve to support more granular device management:

  • Intelligent Scheduling: Can the Kubernetes scheduler be leveraged to detect a single GPU failure on a multi-GPU node and intelligently move workloads away from that specific GPU, rather than requiring a full node reboot?
  • Proactive Workload Eviction: Can the scheduler identify a node that is frequently failing and proactively drain workloads to prevent future disruptions?
  • Cross-Node Workloads: How will the scheduling and remediation of distributed machine learning workloads, which span multiple GPUs across different nodes, be managed when one component fails?

To address these forward-looking challenges and avoid fragmented, proprietary solutions, Ryan Hallisey announced the proposal of a new Kubernetes working group: Node Lifecycle. Co-initiated with Felipe from Red Hat, this working group aims to collaborate with the community to build standardized, Kubernetes-native solutions for maintenance, device management, and overall node lifecycle within clusters. This initiative seeks to provide common frameworks that benefit all users operating large-scale, hardware-accelerated Kubernetes environments.

Demo / Proof of Concept

▶ Watch: Automating GPU failure recovery for maximum capacity (7:00)

The talk did not include a live demonstration or a proof of concept. Instead, Natalie Bandel and Ryan Hallisey focused on presenting NVIDIA's operational strategies, architectural components, and aggregated data from their production environment to illustrate the scale of the problem and the effectiveness of their automated solutions. Their discussion relied on conceptual diagrams and statistical insights rather than interactive code or system showcases.

Defensive Implications

▶ Watch: Custom GPU problem detector using Kubernetes device plugin (7:50)

Organizations operating GPU-accelerated Kubernetes clusters can draw several critical defensive implications from NVIDIA's experience:

  • Prioritize Comprehensive GPU Health Monitoring: Implement robust monitoring for key GPU metrics such as temperature, power consumption, and driver status. Use tools like Grafana to visualize trends and configure alerts for anomalies, enabling proactive intervention before failures occur.
  • Automate Failure Detection: Move beyond manual checks. Develop or integrate an automated GPU problem detector, potentially extending the Kubernetes device plug-in framework, to identify driver failures or device unresponsiveness. Couple this with the Node Problem Detector (NPD) to propagate GPU-specific issues as node conditions, making them visible to the cluster's orchestration layer.
  • Establish an Automated Remediation Workflow: Design and implement a multi-stage recovery process. This should include automated node draining, followed by reboots, power cycles, and, for persistent issues, full node rebuilds. Automate the detection of "remediation loops" (e.g., multiple reboots within a short period) to trigger these more aggressive recovery steps.
  • Optimize Node Draining: Recognize that node draining is a significant bottleneck. Invest in strategies to accelerate this process, such as integrating with tenant applications to gracefully terminate or migrate interruptible workloads (e.g., spot capacity), and prioritizing critical device failure remediations over less urgent maintenance.
  • Implement Proactive Health Checks: Schedule periodic, automated self-tests for both the nodes and their GPUs. Additionally, integrate automated sanity tests into maintenance and upgrade workflows (e.g., after driver updates or OS patches) to verify that changes haven't introduced new instabilities.
  • Leverage Kubernetes Taints and Tolerations: Utilize taints to prevent the scheduler from placing new workloads on nodes identified as unhealthy, soon-to-be-drained, or undergoing maintenance. This helps to contain the impact of failures and facilitates smoother remediation.
  • Anticipate Challenges of High GPU Density: As GPU counts per node increase (e.g., 8+ GPUs for AI workloads), the cost of a full node reboot escalates dramatically. Defenders should start exploring more granular recovery mechanisms that can isolate and remediate single GPU failures without impacting other healthy GPUs on the same node, potentially requiring changes to the Kubernetes scheduler.
  • Engage with the Community: Actively participate in community initiatives like the proposed Node Lifecycle working group. Contributing to and adopting standardized solutions for node maintenance, device management, and proactive failure handling will lead to more robust, scalable, and interoperable systems across the Kubernetes ecosystem.

Key Takeaways

  • GPU failures are an inevitable challenge at scale: NVIDIA experiences 120-200 GPU failures daily across 60,000+ GPUs, highlighting the critical need for robust management strategies in large Kubernetes clusters.
  • Automation is crucial for capacity maintenance: Manual intervention for GPU failures is unsustainable. Automated detection via custom GPU problem detectors and Node Problem Detector, combined with a multi-stage remediation workflow, is essential to keep expensive GPU resources available.
  • Node draining is the primary bottleneck in remediation: While rebooting is an effective fix, gracefully draining workloads from a node can take hours (up to 8 hours in some cases), significantly delaying capacity restoration. Optimizing this process through workload prioritization and tenant integration is vital.
  • Proactive measures prevent costly reboots: Continuous monitoring of GPU health (temperature, power), proactive self-tests, strategic use of Kubernetes taints and tolerations, and automated post-maintenance sanity checks are critical for preventing failures and minimizing downtime.
  • Increasing GPU density demands new Kubernetes capabilities: The trend towards nodes with 8+ GPUs makes full node reboots economically impactful. The community needs to develop more granular Kubernetes scheduling and lifecycle management features to handle single GPU failures and optimize resource allocation.
  • Community collaboration is key to future solutions: NVIDIA is actively driving the creation of a Node Lifecycle working group within the Kubernetes community to develop standardized, open-source solutions for maintenance, device management, and node health, inviting broader participation to address these complex challenges collaboratively.

About the Speaker(s)

Natalie Bandel and Ryan Hallisey are engineers from NVIDIA, deeply involved in the operational aspects of their large-scale cloud infrastructure. Both speakers shared personal experiences of encountering GPU failures, indicating their hands-on understanding of the challenges discussed. Ryan Hallisey, in particular, has been a prominent voice in this domain, having given a prior KubeCon talk titled "All your GPUs are belong to us." He is also a driving force behind community efforts, having co-proposed the new Node Lifecycle working group within the Kubernetes community, collaborating with Felipe from Red Hat, to address the broader challenges of node maintenance and device management. Their expertise stems from managing NVIDIA's massive GeForce Now platform, which orchestrates over 60,000 GPUs across tens of thousands of nodes globally.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

The talk from NVIDIA engineers Natalie Bandel and Ryan Hallisey provides a brutally honest, data-backed look into the relentless reality of GPU failures at massive scale within Kubernetes clusters. Leveraging their GeForce Now operational experience, they detail sophisticated automated detection and multi-stage remediation strategies, from custom problem detectors to node rebuilds. The session offers critical insights into bottlenecks like node draining and proposes a Kubernetes working group for standardized solutions, making it highly relevant for anyone operating GPU-accelerated infrastructure.

Heather Calloway (CISO) — STRONG ACCEPT

This KubeCon talk from NVIDIA provides a clear, unsentimental look at the operational realities of managing specialized hardware at immense scale. It’s not just about GPUs failing; it's about revenue-generating assets impacting service availability and operational costs. The presentation clearly articulates the necessity of automated detection and multi-stage remediation for maintaining critical capacity, offering highly actionable insights for any organization operating large-scale, hardware-accelerated infrastructure where resilience and cost efficiency are paramount.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025