SIG Cloud Provider Deep Dive: Testing Cloud Cont... Michael McCune, Bridget Kromhout & Walter Fender
Michael McCune, Bridget Kromhout, Walter Fender
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk, presented by Michael McCune, Bridget Kromhout, and Walter Fender, provides a comprehensive deep dive into the current state and future direction of the SIG Cloud Provider within the Kubernetes ecosystem, with a particular focus on the critical challenges and initiatives surrounding the testing of cloud controller managers. The speakers articulate the SIG's journey from a working group primarily focused on extracting cloud-specific code from the core Kubernetes repository to its evolving mission of ensuring the robust and reliable operation of Kubernetes across a diverse and expanding landscape of cloud providers. This shift underscores a fundamental recognition that while the initial code migration was a monumental achievement, the long-term stability and trustworthiness of Kubernetes depend heavily on a sophisticated and adaptable testing framework that can accommodate the unique characteristics of each cloud environment.

Key moments
- 0:50 Initial problem: Kubernetes codebase complexity due to cloud providers
- 2:00 Million lines of in-tree cloud provider code removed
- 3:30 API Server Network Proxy replaces old SSH tunnel
- 5:50 SIG Cloud Provider's evolving mission post-migration
- 6:30 Challenge: testing fundamental cloud provider specific behaviors
- 8:00 Introducing Crow for generating CRDs as a platform engineer
SIG Cloud Provider Deep Dive: Testing Cloud Controller Managers
Speakers: Michael McCune, Software Engineer, Red Hat; Bridget Kromhout, Product Manager, Microsoft; Walter Fender, EM, Google
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=WeWQqQM6kjM
Overview
This talk, presented by Michael McCune, Bridget Kromhout, and Walter Fender, provides a comprehensive deep dive into the current state and future direction of the SIG Cloud Provider within the Kubernetes ecosystem, with a particular focus on the critical challenges and initiatives surrounding the testing of cloud controller managers. The speakers articulate the SIG's journey from a working group primarily focused on extracting cloud-specific code from the core Kubernetes repository to its evolving mission of ensuring the robust and reliable operation of Kubernetes across a diverse and expanding landscape of cloud providers. This shift underscores a fundamental recognition that while the initial code migration was a monumental achievement, the long-term stability and trustworthiness of Kubernetes depend heavily on a sophisticated and adaptable testing framework that can accommodate the unique characteristics of each cloud environment.
The core problem addressed is the inherent complexity introduced by running Kubernetes on various public and private clouds, each with its own APIs, networking paradigms, storage solutions, and operational nuances. Historically, much of this cloud-specific logic resided directly within the main Kubernetes codebase, leading to significant maintenance burdens and security vulnerabilities. The successful removal of this in-tree code necessitated a rethinking of how Kubernetes interacts with its underlying infrastructure, creating a new set of challenges related to testing. This article will explore the SIG's strategic plan to build a modular, scalable, and community-driven testing framework that empowers both established and emerging cloud providers to guarantee the functionality and stability of their Kubernetes implementations.
The talk is particularly relevant for platform engineers, cloud operators, and anyone involved in deploying or managing Kubernetes clusters on various cloud platforms. It highlights why a robust testing strategy is not merely a technical detail but a foundational element for ensuring end-user confidence, enabling cloud sovereignty, and fostering innovation within the Kubernetes community. By detailing the historical context, the technical intricacies of the proposed solutions, and the critical need for community involvement, the speakers paint a clear picture of the path forward for securing the future reliability of Kubernetes in the cloud.
Background
▶ Watch: Initial problem: Kubernetes codebase complexity due to cloud providers (0:50)
The SIG Cloud Provider originated as a working group driven by the escalating complexity of the core Kubernetes codebase, colloquially referred to as kubernetes/kubernetes (KK). A significant portion of this complexity, particularly in the early days, stemmed from deeply embedded cloud-provider-specific code, largely from Google but also from other early adopters. This code was essential for fundamental Kubernetes operations, such as provisioning IP addresses for nodes, managing node lifecycle events (e.g., determining if a VM still exists), and handling load balancers. Each new cloud provider integrated into Kubernetes exacerbated this complexity, making maintenance and feature development increasingly difficult.
The general premise, established around the Kubernetes 1.13 release era, was to halt the inclusion of cloud provider code directly into KK. This decision, however, created an immediate disparity: established providers with in-tree code had an advantage over out-of-tree providers, who faced an unfair burden. To rectify this and streamline the core project, a massive effort was undertaken to extract all cloud provider-specific logic. This culminated in a significant milestone around Kubernetes 1.31, with the deletion of over a million lines of code from KK. This "red diff" represented a major victory for the SIG, marking the successful completion of its initial mission.
Beyond code extraction, the SIG also played a crucial role in developing and hosting the API server network proxy. This system replaced an "unfortunate SSH tunnel" that was previously built into Kubernetes, providing a more secure and robust mechanism for the Kubernetes API server to route traffic to the cluster's data plane, particularly important for cloud providers like Google, IBM, and Microsoft. The API server network proxy, now a reference implementation owned by SIG Cloud Provider, demonstrates the SIG's commitment to foundational infrastructure components.
Despite these successes, the extraction process revealed new challenges. Crucially, many existing tests that relied on in-tree cloud provider code could no longer be run directly within KK. This loss of testing capability, coupled with the emergence of new cloud providers with distinct features and implementations, created a vacuum. The SIG recognized that its mission could not end with mere code extraction; it needed to evolve to ensure that Kubernetes remained consistently functional and reliable across all cloud environments. This new focus pivots towards creating a modular and extensible testing framework that can accommodate the inherent diversity of cloud infrastructures, ensuring that new features are thoroughly validated and that user confidence in Kubernetes deployments remains high. The talk also briefly mentions Crow, an exciting new open-source project aimed at simplifying the generation of CRDs (Custom Resource Definitions) for platform engineers to manage complex, cloud-native workloads, highlighting the SIG's broader interest in shaping the future of cloud-integrated Kubernetes.
Key Findings
▶ Watch: API Server Network Proxy replaces old SSH tunnel (3:30)
The talk highlights several critical findings that underscore the urgent need for a refactored and expanded testing strategy for Kubernetes cloud providers:
- Inconsistent Regression Rates Post-Release: An analysis of Kubernetes release regressions, presented via a chart generated by Jordan Liot, revealed significant variability in the number of regressions discovered in the months following a release. For instance, Kubernetes 1.19 saw 16 regressions by month four, while 1.27 had about 8 regressions in the same period. This inconsistency suggests that while there have been periods of "luck" or effective testing, there are also times when testing falls short, leading to instability. This overall trend, though not cloud-provider-specific, indicates a systemic need for more robust and consistent testing across the project.
- Erosion of User and Operator Confidence: Without comprehensive, configurable, and modular tests, end-users and operators cannot have full confidence that new Kubernetes features will function as advertised on their chosen cloud platform. This lack of confidence is particularly relevant in contexts like "European cloud sovereignty," where specific geo-placement and adherence to standards necessitate accurate and verifiable functionality across diverse cloud regions and providers.
- The "Exponential Complexity Test Matrix" Problem: The inherent diversity of cloud environments creates a "giant complex, exponential complexity test matrix." Cloud providers differ in fundamental ways:
- Operating Systems: Even within Linux, commands for network control (e.g.,
service network stop) vary across distributions. The talk also briefly acknowledges Windows as another dimension of complexity. - Networking: There is "no one true way" to implement load balancing or manage IP addresses. Cloud-specific load balancers need dedicated testing to ensure they integrate correctly with Kubernetes services.
- Storage: Similar to networking, elastic storage solutions vary significantly between providers.
- Core Components: Even the API server network proxy, owned by the SIG, has multiple configurations (HTTP connect vs. gRPC) that need cloud-provider-specific validation.
- Image Registries: Cloud providers may impose restrictions (e.g., images must be in
GCR.iofor Google Cloud), which can unexpectedly break tests designed with external images.
- Loss of Test Coverage Post-Extraction: The successful extraction of over a million lines of cloud provider code from KK, while a necessary step, resulted in the loss of many tests that relied on that in-tree code. These tests, often platform-specific scripts, were removed to avoid slowing down Kubernetes releases. This created a significant gap in test coverage for cloud-specific functionalities, making it harder to ensure that basic operations like node life cycling or network interface manipulation work consistently across different clouds.
- Infrequent Testing Leads to Debugging Challenges: If cloud-provider-specific tests are only run with each Kubernetes release (every four months), debugging failures becomes exceedingly difficult due to the large number of changes accumulated over that period. More frequent, automated testing is essential to quickly identify and address regressions.
These findings collectively highlight that the current testing paradigm is insufficient for the heterogeneous nature of modern cloud computing and that a proactive, standardized, yet flexible approach is vital for the continued success and adoption of Kubernetes across the globe.
Technical Deep Dive
▶ Watch: SIG Cloud Provider's evolving mission post-migration (5:50)
The SIG Cloud Provider's strategy to address the testing challenges is structured around a three-step plan: encourage more providers to participate, refactor the existing test suite, and continually support the communities involved. This plan aims to build a robust, scalable, and community-driven testing infrastructure that ensures Kubernetes functions reliably across all cloud environments.
Step 1: Encourage More Providers to Participate
The first step focuses on creating an accessible on-ramp for new and existing cloud providers to integrate their Kubernetes offerings into a standardized testing framework. The SIG provides dedicated repositories for various cloud providers, such as SIG cloud provider AWS, SIG cloud provider Azure, SIG cloud provider GCP, and SIG SIG cloud provider Huawei. These repositories serve as central hubs for:
- Cloud Controller Manager (CCM) Specifications: Defining the interfaces and expected behaviors for cloud-specific controllers.
- CCM Binaries: Hosting the actual binaries that implement the cloud controller manager logic for a given provider.
- Ancillary Components: This includes anything else needed for a fully functional cloud-integrated Kubernetes, such as special tokens for artifact registries, CSI drivers for storage, custom load balancer controllers (separate from the CCM), or even the specific configurations for the API server network proxy.
The overarching goal is to enable a cloud provider to assemble an open-source Kubernetes distribution that can be tested on their infrastructure, rather than relying solely on managed services like GKE, EKS, or AKS. The SIG offers guidance and support to help providers set up these repos, integrate their components, build a testable cluster, and connect it to the broader testing system. The talk specifically highlights the recent successful process of IBM Cloud joining this infrastructure, demonstrating a concrete example of a new cloud provider integrating into the test framework. This requires cloud providers to become involved in the CNCF and contribute compute resources to run these tests.
Step 2: Refactor the Test Suite
This is the most technically intricate part of the plan, addressing the "exponential complexity test matrix." Historically, many cloud-specific tests were hardcoded scripts within the main KK repository. For example, a script to restart a network interface on GCP would be different from one for Azure, making them non-modular and difficult to maintain or generalize. The refactoring effort aims to:
- Modularize Functionality: Distribute cloud-specific functionality away from KK while retaining common test definitions. The ideal state involves common tests residing in the KK repository. These tests would define an interface, asking the cloud provider implementation if it supports a certain action (e.g., "Do you implement dropping the network interface and bringing it back up?").
- Provider-Specific Implementations: Each cloud provider, within their dedicated repository, would then provide a concrete implementation of that interface. This allows for customized actions (e.g., the specific commands to restart a network on Azure Linux vs. GCP Linux) without polluting the core Kubernetes codebase. The analogy given is the node life cycle controller in KK, which uses an interface to ask the cloud provider if a VM still exists. The goal is to extend this mechanism to the test time.
- Leveraging Existing Patterns: The speakers point to SIG Storage's CSI testing as a successful model. CSI drivers have common tests in KK, and each individual CSI driver only needs to implement a configuration file to make it work. This approach offers a precedent for scaling testing across numerous providers.
Specific Technical Problems and Solutions in Refactoring:
- Node Testing: Critical tests like dropping and bringing up network interfaces, or node life cycling (ensuring the kube-API server correctly updates when an underlying VM is deleted), are highly cloud-specific. The refactoring enables providers to implement these actions distinctly.
- Operating System Diversity: The problem extends beyond cloud providers to the underlying operating systems. Network commands differ across Linux distributions, and the presence of Windows nodes adds another layer of complexity. The test suite needs to be aware of these metadata differences.
- Load Balancing and Networking: There's no single "true way" to implement load balancing. Tests need to validate that cloud provider-specific load balancers correctly handle Kubernetes services and return appropriate error codes (not just HTTP 200 OK for failures).
- Image Registry Restrictions: Cloud providers might enforce rules on where container images can be pulled from (e.g., only
GCR.iofor GCP). Tests must account for these restrictions, potentially requiring images to be recompiled and re-registered in provider-specific registries. - API Server Network Proxy Configurations: Even core components like the API server network proxy have configurable modes (HTTP Connect vs. gRPC), requiring tests to validate each permutation on different clouds.
- Frequency of Testing: To mitigate the difficulty of debugging issues discovered after four months, the plan emphasizes running tests more frequently. This will likely involve cloud providers volunteering compute time to execute their specific test suites regularly, potentially using systems like Prow.
- Provider-Specific Tests: The framework must also allow individual providers to run tests unique to their platform that may not be generalizable to others, ensuring full coverage of their specific features.
Step 3: Support the Community
The final step is crucial for the long-term sustainability of the initiative. This involves:
- Documentation: Creating clear documentation for how to contribute tests, how the interfaces work, and how new providers can integrate.
- Onboarding and Mentorship: Actively supporting new contributors and cloud providers as they get involved in creating and maintaining their test suites.
- Collaborative Problem Solving: Engaging the community in philosophical discussions about cluster bring-up (e.g., comparing
kopsvs. Cluster API approaches) and identifying unknown gaps in testing. - Community Ownership: Emphasizing that SIG Cloud Provider repositories (e.g.,
SIG cloud provider GCP) are owned by the CNCF and the Kubernetes community, not by individual cloud vendors. This encourages broader participation from customers, users, and non-cloud providers in maintaining these critical components.
The SIG is actively working on an enhancement draft to formalize these testing questions and build proofs of concept. The dream is to foster a future where any community member, regardless of their affiliation, can identify bugs, demonstrate failures with tests, and contribute pull requests with fixes, knowing that their contributions will be trusted due to the comprehensive testing framework.
Demo / Proof of Concept
▶ Watch: Challenge: testing fundamental cloud provider specific behaviors (6:30)
While the talk did not feature a live, interactive demonstration, the speakers extensively referenced existing proofs of concept and concrete examples of the proposed framework in action. A key example cited was the successful integration of IBM Cloud into the SIG Cloud Provider's testing infrastructure. This process, detailed through a QR code linking to an issue on the right-hand side of the slides, serves as a practical walkthrough for how a new cloud provider can bring its services into the Kubernetes test infrastructure, volunteer hardware, and enable the running of cloud-specific tests.
Furthermore, the speakers mentioned ongoing work by Michael McCune and a colleague to build a proof of concept around the refactored test suite. This initiative aims to demonstrate how common tests in the kubernetes/kubernetes repository can leverage interfaces, allowing individual cloud providers to implement specific, concrete versions of test functionality (e.g., dropping and bringing up a network interface) within their respective cloud provider repositories. The talk also highlighted the SIG Storage CSI testing as an elegant, existing example of how common tests can be combined with provider-specific configurations (in their case, a configuration file) to achieve scalable testing for numerous drivers. These examples serve to validate the architectural direction and feasibility of the SIG's ambitious testing goals.
Defensive Implications
▶ Watch: Introducing Crow for generating CRDs as a platform engineer (8:00)
The initiatives discussed by the SIG Cloud Provider have significant defensive implications for the stability, reliability, and security of Kubernetes deployments across various cloud environments.
- Reduced Attack Surface and Improved Security Posture: The initial extraction of over a million lines of cloud-specific code from
kubernetes/kuberneteswas a critical defensive move. It centralized cloud-specific logic in dedicated repositories, making the core Kubernetes project leaner, easier to audit, and less susceptible to vulnerabilities stemming from diverse, tightly coupled cloud integrations. The replacement of the "unfortunate SSH tunnel" with the API server network proxy further enhanced security by removing a potential backdoor and providing a more controlled and auditable communication channel.
- Enhanced Reliability and Stability: A robust, modular testing framework directly translates to more reliable Kubernetes deployments. By identifying regressions earlier and ensuring that critical functions like node life cycling, networking, and load balancing work correctly on each cloud provider, the framework minimizes the risk of unexpected outages or misconfigurations. This proactive approach to quality assurance is a fundamental defensive measure against operational failures that can disrupt services and potentially expose data.
- Confidence in Feature Rollouts: For operators and platform engineers, the ability to verify that new Kubernetes features function as intended on their specific cloud provider is paramount. The proposed testing framework provides the necessary validation, reducing the risk of deploying features that might behave unexpectedly or introduce instability due to cloud-specific nuances. This builds trust and encourages safer adoption of new Kubernetes capabilities.
- Support for Cloud Sovereignty and Diverse Deployments: The ability to thoroughly test Kubernetes on any cloud provider, including those with geo-specific requirements or unique configurations, is a defensive enabler for organizations concerned with data residency, compliance, and vendor lock-in. It ensures that even non-mainstream or specialized cloud environments can run Kubernetes with the same level of confidence as major public clouds.
- Community-Driven Bug Detection and Resolution: By encouraging broader community participation in testing, the SIG is creating a distributed "defensive perimeter." More eyes on the code and more diverse testing environments mean a higher likelihood of discovering subtle bugs or cloud-specific misbehaviors that might otherwise go unnoticed. The vision of community members finding bugs, demonstrating them with tests, and contributing fixes directly strengthens the overall resilience of Kubernetes.
- Faster Identification of Breaking Changes: Frequent, automated testing, as advocated by the SIG, allows for rapid identification of breaking changes. If a test fails regularly, it's easier to pinpoint the small change that caused the regression, rather than debugging a complex issue that has accumulated over months. This agility is a key defensive capability in maintaining a secure and functional production environment.
In essence, the SIG Cloud Provider's efforts are building a foundational layer of trust and predictability for Kubernetes in the cloud. By rigorously testing the interfaces between Kubernetes and the underlying infrastructure, they are defending against the complexities inherent in heterogeneous cloud environments, ultimately making Kubernetes a more secure and resilient platform for everyone.
Key Takeaways
- The SIG Cloud Provider successfully completed its initial mission of extracting over a million lines of cloud-specific code from the core
kubernetes/kubernetesrepository, culminating around Kubernetes 1.31. - The SIG's new primary focus is on establishing a robust, modular, and scalable testing framework for cloud controller managers and other cloud-integrated components, addressing the loss of test coverage post-extraction.
- Existing Kubernetes testing exhibits inconsistent regression rates and struggles with the "exponential complexity test matrix" posed by diverse cloud environments, leading to potential instability and erosion of user confidence.
- A three-step plan is being implemented: encouraging more cloud providers to participate, refactoring the test suite to be modular and interface-driven, and providing ongoing community support.
- The refactored test suite will feature common tests in the core Kubernetes repository, with cloud providers implementing specific functionality (e.g., network interface manipulation, node lifecycle, load balancing) within their dedicated SIG repositories.
- Community involvement is crucial for identifying testing gaps, contributing provider-specific implementations, and ensuring the long-term reliability and security of Kubernetes across all cloud platforms.
About the Speaker(s)
- Michael McCune is a Software Engineer at Red Hat, where his work focuses on cloud providers and cluster autoscaling within the Kubernetes ecosystem.
- Bridget Kromhout serves as a Product Manager at Microsoft and is a co-chair of the SIG Cloud Provider. Her efforts are dedicated to upstream open-source and CNCF initiatives.
- Walter Fender is an Engineering Manager at Google, though he jokingly refers to himself as an "EM in denial," preferring to think of himself as a Software Engineer (SUI). He contributed insights from Google's perspective on cloud provider integration.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk provides a critical deep dive into the SIG Cloud Provider's ambitious plan to refactor Kubernetes testing for cloud controller managers. It addresses the 'exponential complexity test matrix' problem head-on, outlining a modular, interface-driven approach to ensure consistent functionality across diverse cloud environments. The speakers, all deeply involved in the Kubernetes ecosystem, present a clear strategy for building a robust, community-driven testing framework that is vital for the platform's long-term stability and trustworthiness.
Heather Calloway (CISO) — STRONG ACCEPT
The SIG Cloud Provider's strategy to refactor Kubernetes testing is a critical initiative for institutional stability in the cloud. By addressing the "exponential complexity" of diverse cloud environments with a modular, community-driven framework, the SIG is building a more predictable and trustworthy foundation for Kubernetes. This work directly informs a CISO's understanding of platform reliability and how to assess the underlying risk posture of cloud-native applications, ensuring that the critical infrastructure we depend on is tested and accountable.