How We Tackle KubeVirt’s Growth and Scalability - Ľuboslav Pivarč, Red Hat & Alay Patel, NVIDIA

Ľuboslav Pivarč, Red Hat, Alay Patel, NVIDIA

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, delivered by Daniel Hill from Red Hat (filling in for Ľuboslav Pivarč) and Alay Patel from NVIDIA, delves into the strategies and technical advancements employed to manage the rapid growth and ensure the scalability of the KubeVirt project. KubeVirt, which enables running virtual machines (VMs) directly on Kubernetes, faces unique challenges in both community governance and technical performance as its adoption expands. The speakers highlight how KubeVirt is adapting its development processes and introducing innovative testing methodologies to meet the demands of large-scale, production environments.

Watch on YouTube

Visual summary for How We Tackle KubeVirt’s Growth and Scalability - Ľuboslav Pivarč, Red Hat & Alay Patel, NVIDIA by Ľuboslav Pivarč, Red Hat, Alay Patel, NVIDIA
Visual summary for How We Tackle KubeVirt’s Growth and Scalability - Ľuboslav Pivarč, Red Hat & Alay Patel, NVIDIA by Ľuboslav Pivarč, Red Hat, Alay Patel, NVIDIA

Key moments

  1. 0:00 Introduction and speakers' welcome
  2. 1:30 KubeVirt community evolution: Scaling with SIGs
  3. 3:20 Introducing the Virtualization Enhancement Process (VEP)
  4. 4:05 VEP goals: Design alignment, decision-making, roadmap
  5. 6:15 KubeVirt 1.5 release overview: Features and changes
  6. 8:00 SIG Compute highlights: VM reset, auto limits, faster migrations
  7. 9:00 SIG Storage & Network: Volume migration, IO threads, plugins

How We Tackle KubeVirt’s Growth and Scalability

Speakers: Ľuboslav Pivarč, Red Hat; Alay Patel, NVIDIA

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=xwGDxqI_3Nk

Overview

This talk, delivered by Daniel Hill from Red Hat (filling in for Ľuboslav Pivarč) and Alay Patel from NVIDIA, delves into the strategies and technical advancements employed to manage the rapid growth and ensure the scalability of the KubeVirt project. KubeVirt, which enables running virtual machines (VMs) directly on Kubernetes, faces unique challenges in both community governance and technical performance as its adoption expands. The speakers highlight how KubeVirt is adapting its development processes and introducing innovative testing methodologies to meet the demands of large-scale, production environments.

The presentation addresses two primary facets of KubeVirt's evolution: the organizational scaling of its open-source community and the technical scaling required for its core functionality. Daniel Hill covers the community's structural changes, including the adoption of Special Interest Groups (SIGs) and a new design proposal process. He also provides an update on the KubeVirt 1.5 release, detailing key features and improvements. Alay Patel then shifts focus to the critical work of KubeVirt’s SIG Scale, an initiative dedicated to ensuring the project's performance and scalability, particularly for organizations like NVIDIA that operate KubeVirt clusters with hundreds of nodes.

The talk underscores the importance of a robust, community-driven approach to open-source project management, coupled with rigorous, data-driven performance testing. By outlining the architectural changes, new features, and the novel use of simulation tools like Quark, the speakers demonstrate KubeVirt's commitment to delivering a stable, high-performance platform for virtualized workloads on Kubernetes. This information is crucial for KubeVirt users, contributors, and anyone interested in the operational intricacies of large-scale cloud-native virtualization.

Background

▶ Watch: Introduction and speakers' welcome (0:00)

KubeVirt emerged as a groundbreaking project aiming to bridge the gap between traditional virtualized workloads and the cloud-native paradigm of Kubernetes. It allows users to manage VMs alongside containers within the same Kubernetes cluster, leveraging Kubernetes primitives for VM lifecycle management, networking, and storage. As KubeVirt gained traction, particularly with enterprises and cloud providers seeking to migrate existing VM-based applications to Kubernetes or consolidate infrastructure, the project encountered significant challenges related to both its community structure and its technical scalability.

Initially, KubeVirt's community governance, characterized by a small group of "root approvers" responsible for all code merges and design decisions, proved unsustainable. This centralized model created bottlenecks, hindering the pace of development and making it difficult to onboard new contributors effectively. The need for a more distributed and scalable decision-making process became evident, drawing inspiration from the successful Special Interest Group (SIG) model adopted by the upstream Kubernetes project. This shift was essential to foster broader community participation, distribute workload, and ensure specialized expertise could guide specific areas of development.

Concurrently, the technical demands on KubeVirt grew exponentially. Running virtual machines, especially at scale, introduces complex performance and resource management considerations that differ from stateless containers. Organizations like NVIDIA, operating KubeVirt deployments across hundreds of nodes, began pushing the boundaries of the platform's capabilities. This highlighted the necessity for dedicated efforts to benchmark, monitor, and optimize KubeVirt's performance and scalability. Traditional end-to-end (E2E) testing, while valuable, often required substantial compute resources for large-scale performance tests, leading to high infrastructure costs and slower feedback loops. This environment necessitated the exploration of more cost-effective and efficient methods for identifying and resolving scalability bottlenecks, particularly within the Kubernetes control plane itself.

Key Findings

▶ Watch: Introducing the Virtualization Enhancement Process (VEP) (3:20)

The talk reveals several key findings and advancements across KubeVirt's community and technical landscapes:

  1. Community Decentralization and Process Refinement: KubeVirt successfully transitioned from a centralized "root approver" model to a SIG-based structure, mirroring upstream Kubernetes. This includes SIGs for Compute, Network, Storage, Scaling, and Observability, significantly distributing development responsibilities and accelerating decision-making. Complementing this is the introduction of the Virtualization Enhancement Process (VEP), a lightweight design proposal mechanism designed to align SIGs and contributors on long-term architectural decisions and provide a clearer roadmap for the project, aiming to be mandatory by the end of 2025.
  1. KubeVirt 1.5 Release Enhancements: The latest release, KubeVirt 1.5 (released March 13, 2024), brings substantial improvements. Notable features include the ability to reset a virtual machine without recreating its pod, graduation of auto resource limits for VMs, virtio-fs live migration support, and the adoption of standard SE Linux policies. Performance-wise, multifd usage has been implemented to speed up migrations, and a new IO thread policy has yielded significant storage performance gains. A critical RBAC change related to migrations was also introduced, prioritizing eviction-related migrations.
  1. Robust Performance and Scalability Benchmarking: KubeVirt's SIG Scale has established a sophisticated benchmarking stack. This system integrates daily E2E load generation, persistent metric storage in S3 buckets, and time-series plotting of key performance indicators (KPIs) like VMI transition times. This allows the team to continuously monitor performance, identify regressions (e.g., increased time for VMs to reach a running state), and correlate them back to specific code changes and GitHub pull requests. The system has successfully caught three to four regressions and one monitoring change.
  1. Quark for Cost-Effective Control Plane Scaling: A significant innovation is the integration of Quark (Kubernetes without Kubelet) into KubeVirt's CI system. Quark is a lightweight tool that measures control plane performance and scale by simulating thousands of fake nodes and pods without requiring actual hardware. This dramatically reduces the compute resources and costs associated with large-scale testing. KubeVirt specifically contributed enhancements to Quark, including support for Virtual Machine Instances (VMIs) and an impersonation feature that allows Quark to act as a KubeVirt-owned service account, bypassing webhook rejections for VMI updates. This enables KubeVirt to test scenarios with up to 1000 VMIs in a minimal cluster, effectively uncovering control plane scalability bugs.

Technical Deep Dive

▶ Watch: VEP goals: Design alignment, decision-making, roadmap (4:05)

KubeVirt's journey towards enhanced growth and scalability is multifaceted, encompassing both its community's organizational structure and deep technical implementations within the platform itself.

Community Scaling and Governance

The most significant change in KubeVirt's community structure is the adoption of the Special Interest Group (SIG) model, a proven approach from upstream Kubernetes. This initiative, moving away from a small group of "root approvers," delegates responsibility and expertise to specialized groups. KubeVirt now features SIGs for areas such as Compute, Network, Storage, Scaling, and Observability. Each SIG is responsible for specific aspects of the project, including design, development, and code reviews within its domain. SIG chairs play a crucial role in coordinating efforts between different SIGs, ensuring a cohesive and aligned development trajectory. This decentralization has been instrumental in spreading the workload and involving a broader base of contributors.

Further streamlining the development process is the Virtualization Enhancement Process (VEP). This lightweight design proposal mechanism is analogous to Kubernetes' CAP (Kubernetes Enhancement Proposal) but is tailored to KubeVirt's smaller community size and bandwidth. The VEP aims to achieve several critical goals:

  • Commitment and Alignment: Ensure that not only the design authors but also the relevant SIGs commit to and align on proposed architectural changes.
  • Long-Term Support: Secure buy-in from reviewers and approvers, fostering long-term support for implemented features.
  • Decision-Making Alignment: Empower SIGs to take ownership of decisions concerning their respective areas.
  • Roadmap Generation: Enable the community to construct a clear roadmap from approved VEPs, providing transparency about future features and their target release timelines. The VEP is targeted to become mandatory by the end of 2025.

Beyond formal processes, KubeVirt is also automating its release process, including milestone management and improved release signals on branches, to provide better feedback to developers. The community has also started discussing design proposals directly at unconferences, fostering real-time collaboration and commitment to new features.

KubeVirt 1.5 Feature Set

The KubeVirt 1.5 release, launched on March 13, 2024, is powered by libvirt 10 and QEMU 9.1, targeting Kubernetes 1.32 while being tested against the three latest Kubernetes releases. This release introduces several critical updates:

  • Breaking Change (RBAC): A significant RBAC (Role-Based Access Control) change prioritizes migrations initiated by eviction events over normal user-initiated VM migrations. Users must be aware of this for smooth upgrades.
  • Bug Fixes: A major bug preventing migration recovery has been resolved.
  • Deprecations: The virtctl local-ssh command has been deprecated.

SIG Compute improvements include:

  • Virtual Machine Reset: Users can now reset a VM without terminating and recreating its underlying pod, significantly improving operational fluidity.
  • Auto Resource Limits Graduation: This feature automatically sets resource limits for VMs within namespaces that have resource quotas, simplifying resource management.
  • Virtio-FS Live Migration: Support for live migration with virtio-fs is now available, enhancing the usability and flexibility of VM workloads.
  • Standard SE Linux Policy: KubeVirt pods hosting VMs now use the standard SE Linux policy instead of custom ones, improving security posture and integration with enterprise environments.
  • Multifd for Migrations: The adoption of multifd technology significantly speeds up VM migrations.

SIG Storage saw two major updates:

  • Volume Migration Graduation: The volume migration feature has graduated, indicating its stability and readiness for broader use.
  • New IO Thread Policy: A new I/O thread policy has been implemented, delivering notable performance improvements for storage operations.

SIG Network enhancements include:

  • Network Interfaces Link State: KubeVirt now implements the network interfaces link state, providing better visibility and control over VM network connectivity.
  • Network Binding Plugin Graduation: The network binding plugin has graduated, allowing for the use of custom network plugins within KubeVirt.

The talk also highlighted the community's contribution to upstream Kubernetes through several CAPs that enable further KubeVirt work, such as the Swap CAP, Volume Source OCI Artifacts CAP, and In-place Vertical Scaling CAP, as well as the adoption of DRA (Dynamic Resource Allocation).

SIG Scale Benchmarking Stack

The KubeVirt SIG Scale team has engineered a comprehensive benchmarking stack to continuously monitor and improve performance. This stack comprises:

  1. Load Generator: Utilizes daily end-to-end (E2E) tests to create realistic loads on KubeVirt clusters. Three instances of these tests run daily.
  2. Metric Collection & Persistence: Once tests execute, performance metrics (e.g., VM state transition times) are collected and dumped into an S3 bucket for persistent storage.
  3. Visualization and Analysis: CI jobs scrape these metrics and plot them over time. The graphs display daily runs as blue dots and weekly aggregates as an orange line. This visual representation allows developers to immediately identify performance regressions or improvements.

A key success of this system is its ability to pinpoint regressions. When a performance degradation is observed (e.g., VMs taking longer to reach a running state), the team can review GitHub pull requests merged around that time to identify the causative code change. This proactive approach has successfully caught several regressions, preventing them from impacting production users.

Quark: Kubernetes Without Kubelet

To address the high compute cost of large-scale performance testing, KubeVirt's SIG Scale adopted and enhanced Quark, an open-source tool abbreviated as "Kubernetes without Kubelet." Quark is a lightweight utility designed to measure the performance and scalability of the Kubernetes control plane without requiring extensive physical hardware.

Quark's Architecture and Mechanism:

In a Quark-enabled cluster, the standard Kubernetes control plane (API server, scheduler, controller-manager) operates normally. However, instead of numerous physical worker nodes managed by Kubelets, Quark introduces a Quark controller. This controller can reconcile thousands of fake nodes and fake pods that exist only as Kubernetes objects, without corresponding physical resources or running Kubelet agents. This generates significant pressure on the control plane, allowing for the detection of scalability bottlenecks.

The core of Quark's functionality lies in its custom resource, the Stage CR. A Stage CR is a declarative configuration that instructs the Quark controller on which Kubernetes objects to observe and how to transition their states. For example, a Stage CR can define that if a node object is in a "not ready" state, the Quark controller should update it to a "ready" state, mimicking a Kubelet's behavior without an actual Kubelet. The Stage CR specifies:

  • resourceRef: The type of Kubernetes object (e.g., Node, Pod).
  • selector: Criteria for selecting objects in a particular state (e.g., node is not ready).
  • next: The desired target state for the object (e.g., node should be ready).

KubeVirt-Specific Enhancements to Quark:

KubeVirt's integration with Quark required specific contributions to support its unique workloads:

  • VMI Support: KubeVirt added direct support for Virtual Machine Instances (VMIs) within Quark. This means a Stage CR can now be configured to manage the lifecycle of fake VMIs, transitioning them from a Scheduled state to a Running or Ready state.
  • Impersonation Feature: A critical challenge for KubeVirt was that any update to a VMI object must originate from a KubeVirt-owned service account; otherwise, KubeVirt's admission webhooks would reject the change. To overcome this, KubeVirt contributed an impersonation feature to Quark. This allows the Quark controller to impersonate a specified KubeVirt service account when making VMI updates. The API server then perceives the update as coming from an authorized KubeVirt component, allowing the change to proceed. This feature is generically useful for any custom workload wrapped around pods that has similar webhook validation requirements.

With these enhancements, KubeVirt's SIG Scale can now use Quark to simulate 1000 VMIs in a minimal cluster, effectively stress-testing the control plane for scalability bugs with significantly reduced resource costs compared to traditional E2E tests on real hardware.

Demo / Proof of Concept

▶ Watch: SIG Compute highlights: VM reset, auto limits, faster migrations (8:00)

While no live code demo was presented during the talk, the speakers effectively demonstrated the results and methodology of their scaling efforts through detailed graphs and architectural diagrams. The core proof of concept lies in the successful integration of Quark into the KubeVirt CI system and the observable benefits derived from it.

Alay Patel presented graphs from the KubeVirt benchmarking stack, visually illustrating the performance of Virtual Machine Instances (VMIs) over time. These graphs, with the x-axis representing time and the y-axis showing the duration for a VMI to transition to a running state, clearly depicted daily test runs (blue dots) and weekly aggregates (orange line). Crucially, these charts served as a proof of concept for the effectiveness of their monitoring and regression detection system. The speaker highlighted instances where the graphs showed a clear regression—a spike in the time taken for VMIs to become ready—and explained how this allowed the team to quickly identify and rectify the underlying code changes responsible. This system's ability to successfully catch three to four such regressions and one monitoring change validates its utility.

Furthermore, the detailed explanation of Quark's architecture and its KubeVirt-specific enhancements (VMI support and the impersonation feature) serves as a proof of concept for cost-effective, high-scale control plane testing. The ability to simulate 1000 VMIs within a small cluster, as highlighted by Alay Patel, demonstrates that KubeVirt can now achieve comprehensive scalability testing without incurring massive infrastructure costs. This simulation capability allows KubeVirt to proactively identify and address scalability bottlenecks in its control plane, ensuring the project can support increasingly large and complex deployments in real-world scenarios.

Defensive Implications

▶ Watch: SIG Storage & Network: Volume migration, IO threads, plugins (9:00)

While this talk isn't focused on traditional security vulnerabilities, the insights into KubeVirt's growth and scalability have significant implications for operators and developers who "defend" the stability, performance, and reliability of their KubeVirt deployments. These implications can be categorized as operational resilience and proactive quality assurance.

  1. Operational Stability through Community Processes: The shift to SIGs and the Virtualization Enhancement Process (VEP) provides a more stable and predictable development trajectory for KubeVirt. Defenders (operators) should engage with KubeVirt SIGs relevant to their usage (e.g., SIG Compute, SIG Network) to understand upcoming changes, voice concerns, and contribute to designs. This proactive engagement can help shape features that enhance operational resilience and prevent unexpected breaking changes or performance regressions. The VEP's goal of a clear roadmap is a defensive tool against uncertainty, allowing for better planning and resource allocation.
  1. Staying Current with Releases and Breaking Changes: The mention of a breaking RBAC change in KubeVirt 1.5 related to migration priority underscores the importance of thoroughly reviewing release notes. Operators must understand these changes to prevent service disruptions during upgrades. Prioritizing eviction-related migrations is a defensive measure to ensure cluster stability and resource reclamation, but it requires careful configuration and testing.
  1. Leveraging New Features for Enhanced Management: New features like virtual machine reset without pod recreation, auto resource limits, and virtio-fs live migration offer improved operational flexibility and resource efficiency. Defenders can leverage these to build more robust and self-healing systems, reducing manual intervention and potential for human error. The adoption of standard SE Linux policies also simplifies security configuration and integration into existing enterprise security frameworks.
  1. Proactive Performance Assurance: The SIG Scale benchmarking stack and its ability to detect regressions in CI are a critical defensive mechanism against performance degradation. While this is primarily an upstream developer tool, operators can learn from its methodology. For large-scale KubeVirt deployments, implementing similar continuous performance monitoring, tailored to their specific workloads and infrastructure, is crucial. Detecting performance bottlenecks early, before they impact production, is a fundamental aspect of operational defense.
  1. Cost-Effective Scalability Testing with Quark: The integration of Quark for control plane scalability testing is a direct defensive measure against unforeseen scaling limits. Organizations running or planning to run KubeVirt at scale (e.g., hundreds or thousands of VMs) should explore adopting Quark-like methodologies in their own pre-production environments. By simulating high control plane pressure with minimal hardware investment, they can identify and mitigate potential scalability issues specific to their configurations, preventing costly outages or performance bottlenecks in production. The KubeVirt-specific enhancements to Quark (VMI support, impersonation) make this even more relevant for KubeVirt users.

In essence, the "defensive implications" here are about building a resilient, performant, and manageable KubeVirt infrastructure. By understanding the project's evolution, leveraging its features, and adopting its rigorous testing methodologies, operators can significantly enhance the stability and scalability of their virtualized workloads on Kubernetes.

Key Takeaways

  • Community Scaling through SIGs and VEP: KubeVirt has successfully evolved its community governance by adopting Special Interest Groups (SIGs) and introducing the Virtualization Enhancement Process (VEP), a lightweight design proposal mechanism, to distribute responsibilities, align contributors, and provide a clearer project roadmap.
  • KubeVirt 1.5 Delivers Key Operational Enhancements: The KubeVirt 1.5 release brings significant features like virtual machine reset without pod recreation, graduated auto resource limits, virtio-fs live migration support, and performance boosts through multifd and a new IO thread policy, alongside important RBAC changes.
  • Robust Performance Monitoring Detects Regressions: The KubeVirt SIG Scale team utilizes a sophisticated benchmarking stack with daily end-to-end tests, persistent metric storage, and time-series plotting to continuously monitor performance and successfully identify and address regressions in the codebase.
  • Quark Enables Cost-Effective Control Plane Scalability Testing: KubeVirt has integrated Quark (Kubernetes without Kubelet) into its CI, allowing for the simulation of thousands of fake nodes and Virtual Machine Instances (VMIs) without real hardware. This significantly reduces testing costs while effectively uncovering control plane scalability bugs.
  • KubeVirt Contributed Crucial Quark Enhancements: To enable comprehensive VMI testing, KubeVirt contributed specific features to Quark, including direct VMI support and an impersonation feature that allows Quark to act as a KubeVirt service account, bypassing webhook rejections for VMI updates.
  • Community Involvement is Encouraged: The KubeVirt project actively seeks contributions to SIG Scale and other areas, offering opportunities for individuals interested in performance testing, community governance, and open-source virtualization.

About the Speaker(s)

Daniel Hill from Red Hat filled in for Ľuboslav Pivarč during the presentation. He is a member of the Open Virtualization Team at Red Hat, with a primary focus on the KubeVirt upstream CI system. His work involves ensuring the continuous integration and testing infrastructure for KubeVirt is robust and efficient.

Alay Patel from NVIDIA shared insights from the perspective of a large-scale KubeVirt user. He is involved with KubeVirt's SIG Scale, a group dedicated to the performance and scalability aspects of the project. His organization, NVIDIA, runs KubeVirt at scale, managing clusters with hundreds of nodes, making their contributions and experience critical to KubeVirt's ongoing development.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk provides a thorough and technically rich overview of KubeVirt's strategies for managing rapid growth and ensuring scalability. It effectively covers both the organizational evolution, with the adoption of SIGs and the VEP, and critical technical advancements in KubeVirt 1.5. The highlight is the detailed explanation of SIG Scale's robust benchmarking stack and, particularly, the novel integration and KubeVirt-specific enhancements to Quark for cost-effective, high-scale control plane testing. This is a substantive presentation demonstrating real engineering effort and delivering actionable insights for KubeVirt users and contributors.

Heather Calloway (CISO) — STRONG ACCEPT

This presentation on KubeVirt's scalability and growth offers a compelling case study in managing the institutional and technical challenges of a critical open-source project. While deeply technical, the talk's focus on decentralizing community governance through Special Interest Groups, establishing a clear Virtualization Enhancement Process, and implementing rigorous, cost-effective performance testing with tools like Quark, directly addresses core CISO concerns around institutional accountability, operational resilience, and the long-term viability of foundational infrastructure. It's a pragmatic look at how a complex system ensures its own stability and performance, delivering tangible…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025