How the SIG-Multicluster API Specifications Are Used for Real World... August Simonelli & Ryan Zhang
August Simonelli, Ryan Zhang
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
The proliferation of Kubernetes clusters across diverse environments has introduced significant operational complexities, particularly in managing applications and infrastructure consistently. This talk, delivered by August Simonelli and Ryan Zhang, delves into the critical role of the SIG-Multicluster API specifications in addressing these challenges. It highlights how these foundational standards enable interoperability and streamline multicluster management for real-world deployments.

Key moments
- 0:00 Talk Introduction and Agenda Overview
- 0:50 SIG Multicluster Philosophy: Define APIs, Not Implementations
- 2:00 About API and Cluster Profile API Explained
- 4:00 Work API for Multicluster Workload Placement
- 4:30 Multicluster Services API for Cross-Cluster Communication
- 5:40 Introduction to Open Cluster Management (OCM) Project
- 6:00 OCM Hub-and-Spoke Model and Key Features
How the SIG-Multicluster API Specifications Are Used for Real World Multicluster Management
Speakers: August Simonelli, Senior Principal Product Manager, Red Hat; Ryan Zhang, Principal Software Engineer, Microsoft Azure
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=I9GV4N23dvE
Overview
The proliferation of Kubernetes clusters across diverse environments has introduced significant operational complexities, particularly in managing applications and infrastructure consistently. This talk, delivered by August Simonelli and Ryan Zhang, delves into the critical role of the SIG-Multicluster API specifications in addressing these challenges. It highlights how these foundational standards enable interoperability and streamline multicluster management for real-world deployments.
Simonelli and Zhang, representing the Open Cluster Management (OCM) project at Red Hat and the KubeVela Fleet project at Microsoft Azure, respectively, demonstrate how their projects implement and extend these community-driven APIs. The core message emphasizes defining standards over specific implementations, fostering an ecosystem where various tools can seamlessly integrate. This approach allows organizations to manage their Kubernetes fleets with greater efficiency, resilience, and scalability, moving beyond the manual, cluster-by-cluster management that often plagues growing infrastructures.
The session provides a comprehensive look at the four primary SIG-Multicluster APIs – About, Cluster Profile, Work, and Multicluster Services – illustrating their practical application through live and recorded demonstrations. By showcasing how these specifications are leveraged by mature projects, the speakers underscore their immediate relevance and future potential in shaping the landscape of multicluster Kubernetes. This talk is essential for anyone grappling with multicluster complexity, offering insights into standardized approaches that promise to simplify and empower fleet operations.
Background
▶ Watch: Talk Introduction and Agenda Overview (0:00)
The journey into multicluster Kubernetes management is often fraught with challenges stemming from Kubernetes' inherent design, which was originally conceived for single-cluster operations. A fundamental issue highlighted in the talk is the absence of a native, single ID for a Kubernetes cluster. While individual resources within a cluster have unique identifiers, the cluster itself lacks a standardized, discoverable ID. This oversight becomes a significant hurdle when attempting to manage multiple clusters as a cohesive unit, as distinguishing and referencing them programmatically becomes difficult without external mechanisms.
Historically, managing applications across multiple clusters typically involved repetitive, manual operations or custom scripting. Administrators would often have to kubectl into each individual cluster to apply configurations, deploy workloads, or monitor status. This approach is not only error-prone but also scales poorly as the number of clusters grows from a handful to dozens or even hundreds. The need for a "single pane of glass" – a centralized control plane from which to oversee and orchestrate operations across an entire fleet of clusters – became evident.
Prior to the formalization of SIG-Multicluster APIs, various projects attempted to solve these problems with their own proprietary solutions. While effective within their specific ecosystems, these disparate approaches led to fragmentation and vendor lock-in, hindering broader adoption and interoperability. The Kubernetes SIG-Multicluster (Special Interest Group) was formed to address this by defining APIs, not implementations. The goal was to establish universally agreed-upon standards that different projects and vendors could adopt, thereby making it easier to contribute to and integrate with the multicluster ecosystem. This shift from individual solutions to community-driven specifications is crucial for building a resilient, extensible, and vendor-neutral future for multicluster Kubernetes. The talk champions this philosophy, presenting it as the foundation for modern fleet management.
Key Findings
▶ Watch: About API and Cluster Profile API Explained (2:00)
The core of multicluster management, as presented by Simonelli and Zhang, revolves around four fundamental APIs defined by the SIG-Multicluster. These APIs provide the building blocks for a standardized approach to orchestrating resources and services across multiple Kubernetes clusters:
- About API: This is the simplest and oldest of the APIs, designed to define well-known properties for an individual cluster. A critical function of the About API is to provide a unique identifier for a cluster, typically a
cluster name, which Kubernetes itself lacks. While seemingly basic, this identity is foundational for any multicluster system to distinguish and manage different clusters. The speakers note an ongoing Kubernetes Enhancement Proposal (KubeCon EU, CAP) to add more properties, such as available CPU or memory, which will enrich the descriptive capabilities of this API and support more intelligent scheduling decisions.
- Cluster Profile API: Building upon the About API, the Cluster Profile API allows a central control plane (often referred to as a "hub" or "fleet manager") to aggregate and view the properties of all managed clusters. This API provides the "single pane of glass" necessary for fleet administrators to understand the characteristics and status of their entire cluster fleet. It acts as a centralized repository for cluster metadata, pulling information defined by the About API from individual clusters to present a unified view. This centralization is vital for making informed decisions about workload placement and overall fleet health.
- Work API: The Work API is the mechanism for deploying Kubernetes resources from the central control plane to individual managed clusters. Conceptually, it mirrors the familiar
kubectl applycommand but operates across a fleet. Instead of manually applying YAML files to each cluster, administrators can define a "workload" (a collection of Kubernetes resources) on the hub, and the Work API ensures its delivery and execution on the designated managed clusters. This API is designed as a "workhorse," faithfully executing instructions without embedding complex scheduling logic, which is delegated to other components. It streamlines the distribution of applications, configurations, and policies, ensuring consistency and reducing manual effort.
- Multicluster Services API: This API addresses the challenge of inter-service communication and accessibility across different clusters. It enables services with the same name and namespace sameness to be treated as a unified logical service, regardless of their physical cluster location. For example, if a
frontendservice exists in multiple clusters, the Multicluster Services API allows other services or external users to access it without needing to know the specific cluster where an instance resides. This abstraction provides service discovery and connectivity across cluster boundaries, enhancing application resilience and user experience by allowing traffic to be routed to any available healthy instance. It moves away from the need to explicitly specifycluster A's serviceorcluster B's service, offering a more robust, location-agnostic communication model.
Together, these four APIs form a robust framework for managing Kubernetes clusters at scale, promoting standardization, and enabling advanced multicluster capabilities.
Technical Deep Dive
▶ Watch: Work API for Multicluster Workload Placement (4:00)
The SIG-Multicluster APIs lay the groundwork, and projects like Open Cluster Management (OCM) and KubeVela Fleet (Kubby Fleet) demonstrate their practical implementation and extension. Both projects adopt a hub-and-spoke model, where a central "hub" cluster manages multiple "spoke" or "managed" clusters.
Open Cluster Management (OCM)
OCM, developed by Red Hat, is a prominent implementation of the SIG-Multicluster specifications. Its architecture is built around:
- Hub and Spoke Model: A central hub cluster orchestrates operations, while lightweight agents called clusterlets run on each managed cluster. These clusterlets provide weak autonomy, meaning managed clusters can continue operating even if connectivity to the hub is temporarily lost, ensuring resilience.
- Manifest Work: OCM's implementation of the Work API is called
manifest work. This is the core mechanism for delivering Kubernetes resources (like deployments, services, or service accounts) from the hub to managed clusters. Themanifest workobject encapsulates the YAML definitions of resources to be deployed. - Placement: OCM enhances the basic Work API with a powerful
placementcapability. While the Work API is a "workhorse," placement provides the "intelligence" to dynamically select target clusters based on defined criteria. This allows for conditional deployments, such as targeting clusters with specific CPU or memory characteristics, or those within a particular geographical region. Placement works by evaluatingcluster claims(OCM's implementation of the About API, which captures cluster properties) against user-defined rules. - Add-ons: OCM includes an
add-on frameworkthat provides a modular way to extend its functionality. Projects can easily integrate their tools and services into OCM by developing add-ons, further promoting a flexible and extensible multicluster ecosystem. For instance, OCM leverages Submariner for its multicluster services implementation.
KubeVela Fleet (Kubby Fleet)
KubeVela Fleet, a CNCF-contributed project from Microsoft Azure, offers another robust implementation of these standards, focusing on comprehensive multicluster application management:
- Single Pane of Glass: Similar to OCM, KubeVela Fleet provides a centralized interface for managing applications across a fleet of clusters, eliminating the need for individual cluster interactions.
- Advanced Scheduling Capabilities: KubeVela Fleet significantly extends placement logic by incorporating many concepts from Kubernetes' native scheduler. This includes affinities, topology spread constraints, and preferred/required during scheduling rules. This allows for highly granular control over where workloads land, optimizing for resource utilization, fault tolerance, and cost.
- Placement Policies: KubeVela Fleet offers various placement policies, such as selecting a specific number of clusters, targeting all clusters (like a DaemonSet), or property-based scheduling. The property-based scheduling directly ties into the About API and Cluster Profile API, allowing administrators to select clusters based on dynamic properties like available CPU, memory, or custom cloud provider attributes (e.g., Azure-specific properties). The talk demonstrates a
property sortterto prioritize clusters with the most CPUs or lowest cost. - Continuous Deployment: KubeVela Fleet is building in continuous deployment strategies, which were briefly mentioned but not demonstrated in detail during this specific talk.
- Drift Detection and Takeover: A crucial feature highlighted by Ryan Zhang is KubeVela Fleet's drift detection. This mechanism identifies when a resource on a managed cluster has been manually modified, deviating from the desired state defined on the hub. It also includes takeover capabilities, which are essential when migrating existing, manually managed clusters into a fleet. Takeover allows the fleet manager to assume control of pre-existing resources without causing accidental overwrites, offering
diff reportingto show changes before enforcement. While these are advanced features, they are currently extra work and not yet upstreamed into the core Work API specification.
Authentication Model
Both OCM and KubeVela Fleet employ a pull model for authentication. The agent running on the member cluster (e.g., OCM's clusterlet, KubeVela Fleet's agent) authenticates to the hub cluster. This can be achieved through various methods, including federated identity solutions like OIDC for cloud providers or traditional Kubernetes secrets for simpler setups. This pull-based approach enhances security by minimizing the need for the hub to directly access managed clusters' API servers with high privileges.
Dynamic Properties and Placement API Evolution
A key discussion point in the Q&A was the dynamic nature of cluster properties (like available CPU/memory) and the evolution of a standardized Placement API. While the About API and Cluster Profile API define how properties are exposed, the SIG-Multicluster currently lacks a standardized, dynamic Placement API. Projects like OCM and KubeVela Fleet have their own sophisticated placement engines that collect metrics and update dynamic properties with a certain cadence. The speakers acknowledge that a generic Placement API is complex due to the diverse requirements across different environments and use cases, and it remains a future goal for the SIG.
Demo / Proof of Concept
▶ Watch: Introduction to Open Cluster Management (OCM) Project (5:40)
Both speakers provided compelling demonstrations of their respective projects, illustrating the practical application of the SIG-Multicluster APIs.
August Simonelli's OCM Demo (Recorded)
August's demo showcased OCM's manifest work and placement capabilities using a setup with one hub and two managed clusters.
- Basic Manifest Work Deployment:
- August first showed how to deploy a simple Nginx application (version 1.26.1) and a service account. He created a
manifest workobject on the hub in a namespace namedcluster-one(matching the managed cluster's name). - He explicitly clarified that while the
manifest worklived on the hub incluster-onenamespace, the actual Nginx deployment was targeted for thedefaultnamespace on the managedcluster-one. - Verification confirmed the Nginx pod running on
cluster-oneand its absence oncluster-two, demonstrating targeted deployment. - The
manifest workobject on the hub provided detailed status information, including the deployed resources and their states, enabling automation and monitoring. - August then demonstrated an update: changing the Nginx version from 1.26.1 to 1.27.4 in the
manifest workdefinition. Applying this change to the hub automatically triggered the rollout of the new Nginx version oncluster-one, proving thework API's ability to manage lifecycle updates.
- Placement for Multicluster Deployment:
- After cleaning up the single-cluster deployment, August introduced placement. He created a cluster set (a logical grouping of clusters) and bound a namespace to it.
- A
placementobject was then defined with a simple rule: "find two clusters that meet these criteria." Since only two clusters were registered, it selected both. - The Nginx application (as standard YAML, not
manifest workkind explicitly) was then created, explicitly referencing theplacementobject. - OCM's placement engine automatically created the necessary
manifest workobjects for bothcluster-oneandcluster-two. - Verification showed the Nginx deployment successfully running on both managed clusters, demonstrating how placement intelligently distributes workloads across a fleet based on defined rules.
Ryan Zhang's KubeVela Fleet Demo (Live)
Ryan's live demo illustrated KubeVela Fleet's advanced scheduling and multicluster services, using a fleet of nine clusters.
- Cluster Profile and Property-Based Placement:
- Ryan displayed the fleet of nine clusters, showing aggregated properties like available CPUs and available memory via
kubectl get clusterprofile -o wide. - He then showed a
placementconfiguration that leveraged these properties for intelligent scheduling. This included: - A
property sortterto prioritize clusters with the most available CPUs (descending order, with a weight of 50%) and lowest cost (ascending order). - Topology spread affinities for distributing workloads across different failure domains.
- Required during scheduling constraints, such as requiring clusters to have over 13GB of memory.
- The placement selected three specific clusters (e.g.,
bug-batch-3,bug-batch-4, and another named cluster) that met these criteria.
- Workload Deployment with Work API:
- Ryan demonstrated deploying a sample application (consisting of a namespace, service, deployment, and service export) to a namespace called
multicluster-app. - The
workobject (similar to OCM'smanifest work) was created on the hub, targeting the three clusters chosen by the placement engine. - Verification confirmed the application's deployment on
bug-batch-3andbug-batch-4, and its absence on other unselected clusters.
- Multicluster Services with Service Export/Import:
- To demonstrate cross-cluster communication, Ryan showed how the deployed application utilized a service export.
- KubeVela Fleet automatically created a corresponding service import on other clusters, making the service accessible fleet-wide.
- He then used a DNS-based load balancer (referred to as a
traffic manager) to expose the service to external users. This load balancer had a DNS name that, when queried (digcommand), resolved to endpoints across the three clusters where the service was deployed. This illustrated how users could access the "storefront" application without needing to know the specific cluster endpoints, benefiting from global load balancing and resilience.
Both demos effectively showcased how the SIG-Multicluster APIs are translated into tangible, powerful features for managing complex multicluster environments.
Defensive Implications
▶ Watch: OCM Hub-and-Spoke Model and Key Features (6:00)
The adoption of SIG-Multicluster API specifications and their implementations like OCM and KubeVela Fleet brings significant defensive implications, enhancing the security, reliability, and operational integrity of multicluster Kubernetes deployments.
- Reduced Blast Radius and Enhanced Resilience: By enabling intelligent workload placement across multiple clusters, organizations can effectively spread their "blast radius." If one cluster experiences an outage (due to an "oopsie" or geographical issue), the workload can be distributed or failed over to other healthy clusters. OCM's weak autonomy via
clusterletsfurther enhances this by allowing managed clusters to continue operating even if the hub connection is severed, preserving local application availability.
- Consistent Configuration and Policy Enforcement: Centralized management through the Work API ensures that applications and infrastructure configurations are deployed consistently across the fleet. This reduces configuration drift, which is a common source of vulnerabilities and operational issues. Defenders can ensure security policies, network configurations, and compliance controls are uniformly applied, minimizing the risk of misconfigurations in individual clusters.
- Drift Detection and Remediation: KubeVela Fleet's drift detection feature is a critical defensive tool. It identifies unauthorized or accidental manual modifications on managed clusters, alerting administrators to deviations from the desired state. This helps maintain the integrity of deployments and prevents malicious actors or human error from introducing vulnerabilities or disrupting services. While not yet upstreamed, this capability is vital for maintaining a secure and compliant posture.
- Streamlined Authentication and Authorization: The pull-based authentication model, where agents on managed clusters authenticate to the hub, simplifies identity management. Using mechanisms like OIDC for federated identities or Kubernetes secrets provides a controlled and auditable way for clusters to interact, reducing the attack surface compared to complex, distributed authentication schemes.
- Improved Observability and Auditability: The Work API provides rich status information about deployed resources, offering a centralized view of the state of applications across the fleet. This enhanced observability aids defenders in quickly identifying issues, tracking changes, and auditing deployments, which is crucial for incident response and compliance.
- Interoperability and Ecosystem Security: The commitment to defining APIs over specific implementations fosters a more open and interoperable ecosystem. This allows security tools and practices to integrate more easily, as they can target a common API standard rather than proprietary interfaces. This broadens the scope for security innovation and robust tooling within the multicluster space.
In essence, these multicluster management standards and tools move organizations towards a more secure and resilient operational model, providing the controls and visibility necessary to manage complex Kubernetes environments effectively.
Key Takeaways
- Standards Over Implementations: The SIG-Multicluster's philosophy of defining APIs, not implementations, is crucial for fostering an open, interoperable, and vendor-neutral multicluster ecosystem.
- Four Foundational APIs: The About, Cluster Profile, Work, and Multicluster Services APIs address core challenges in multicluster management, from cluster identification and status aggregation to workload deployment and cross-cluster service communication.
- Intelligent Placement is Key: Projects like OCM and KubeVela Fleet extend the Work API with sophisticated placement capabilities, enabling dynamic, property-based scheduling and advanced scheduling policies to optimize resource utilization and resilience across the fleet.
- Enhanced Operational Resilience: Features like OCM's weak autonomy (clusterlets) and KubeVela Fleet's drift detection significantly improve the reliability, consistency, and security of multicluster deployments by mitigating outages and configuration drift.
- Seamless Cross-Cluster Communication: The Multicluster Services API, leveraging concepts like namespace sameness and service exports/imports, allows applications to communicate and be accessed across cluster boundaries without needing to know specific cluster endpoints, simplifying development and enhancing resilience.
- Community-Driven Evolution: The multicluster landscape is still evolving, with areas like a standardized Placement API and dynamic property management actively being discussed within the SIG, highlighting the importance of community involvement.
About the Speaker(s)
August Simonelli is a Senior Principal Product Manager at Red Hat, where he works with the Open Cluster Management (OCM) project. His role involves guiding the development and strategy of OCM, focusing on how it implements and extends the SIG-Multicluster API specifications for real-world multicluster management challenges.
Ryan Zhang is a Principal Software Engineer at Microsoft Azure and a maintainer for the newly contributed CNCF project called KubeVela Fleet (Kubby Fleet). He is deeply involved in the technical implementation and advancement of multicluster management solutions, ensuring KubeVela Fleet adheres closely to community standards while offering advanced features for fleet orchestration.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This session provides a crucial, no-nonsense look at the SIG-Multicluster API specifications, demonstrating their practical application through two mature projects: OCM and KubeVela Fleet. Simonelli and Zhang, both deeply embedded in these projects, cut through the hype to show how these standards enable robust, scalable multicluster management. The focus on defining APIs over proprietary implementations is critical, offering a pathway to genuine interoperability and significantly improved operational resilience for Kubernetes fleets. It's a solid, actionable talk for anyone serious about managing clusters at scale.
Heather Calloway (CISO) — STRONG ACCEPT
This KubeCon talk effectively demonstrates how the SIG-Multicluster API specifications provide a critical foundation for secure and resilient multi-Kubernetes fleet management. By championing community standards over proprietary solutions, the speakers highlight a path to consistent configuration, policy enforcement, and reduced operational risk at scale. Features like drift detection and intelligent workload placement, showcased through projects like OCM and KubeVela Fleet, offer tangible mechanisms for enhancing accountability and maintaining a strong security posture across complex environments. This is a vital discussion for any CISO or security leader grappling with the complexities…