SIG API Machinery: Project Updates and Release Planning - Joe Betz, Google
Joe Betz, Google
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this comprehensive KubeCon EU session, Joe Betz, Technical Lead for SIG API Machinery at Google, delivered an insightful update on the Kubernetes 1.33 release cycle, delving into the critical enhancements and future trajectory of the API machinery Special Interest Group. SIG API Machinery is the foundational pillar of Kubernetes, responsible for the core REST mechanics of the API, including resource definitions, versioning, serialization protocols, and the crucial extensibility mechanisms that allow users to define custom resources (CRDs) and inject custom logic via admission control.

Key moments
- 0:00 Introduction and agenda overview
- 0:40 Defining SIG API Machinery's core responsibilities
- 2:00 Control plane extensibility: CRDs, webhooks, admission policies
- 3:30 Controller infrastructure and control plane performance
- 4:15 What SIG API Machinery is NOT responsible for
- 6:00 Kubernetes 1.33 release updates and KEPs intro
- 6:20 KEP: Ordered Namespace Deletion for predictable lifecycle
SIG API Machinery: Project Updates and Release Planning
Speakers: Joe Betz, Technical Lead, Google
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=VCmp--NcxeE
Overview
In this comprehensive KubeCon EU session, Joe Betz, Technical Lead for SIG API Machinery at Google, delivered an insightful update on the Kubernetes 1.33 release cycle, delving into the critical enhancements and future trajectory of the API machinery Special Interest Group. SIG API Machinery is the foundational pillar of Kubernetes, responsible for the core REST mechanics of the API, including resource definitions, versioning, serialization protocols, and the crucial extensibility mechanisms that allow users to define custom resources (CRDs) and inject custom logic via admission control.
Betz’s talk highlighted the significant strides made in areas like control plane extensibility, client infrastructure, and controller frameworks, all while emphasizing the SIG's overarching commitment to the reliability, scale, and performance of the Kubernetes control plane. The presentation served as both an introductory guide for those new to API machinery and a deep dive into specific Kubernetes Enhancement Proposals (KEPs) that are shaping the platform's future, particularly focusing on improving upgrade safety, boosting performance, and enhancing the declarative nature of Kubernetes APIs.
This talk is paramount for anyone operating or developing on Kubernetes, as it directly impacts the stability, efficiency, and extensibility of their clusters. From cluster administrators seeking smoother upgrades to developers building custom controllers and CRDs, the changes detailed by SIG API Machinery are fundamental to leveraging Kubernetes effectively and securely. Understanding these updates provides a roadmap for adopting new features, mitigating operational risks, and contributing to the evolution of the Kubernetes ecosystem.
Background
▶ Watch: Introduction and agenda overview (0:00)
SIG API Machinery stands at the very heart of Kubernetes, defining the Kubernetes Resource Model (KRM) and providing the essential infrastructure for every interaction with the control plane. Its responsibilities are broad, encompassing the REST mechanics of the Kubernetes API, including versioning, serialization, and the definition of resources and subresources. This SIG is crucial for both built-in types like Pods and Nodes, and for enabling the ecosystem's vast extensibility through Custom Resource Definitions (CRDs). Beyond resource definition, API Machinery also governs control plane extensibility via aggregated API servers and admission control, which allows for the interception and modification of write requests through admission webhooks or the newer admission policies powered by the Common Expression Language (CEL).
The SIG also maintains the language clients (especially the Go clients), provides discovery services for the API, and is responsible for the controller infrastructure, including informers and the watch mechanism that enable efficient controllers. Critically, it shoulders much of the burden for the reliability, scale, and performance of the API server and controller managers, working closely with SIG etcd on storage-related aspects.
Despite its wide purview, Betz clarified what API Machinery is not responsible for: API review (handled by a dedicated group), ownership of most specific APIs (these belong to their respective SIGs), individual controllers (framework only), kubectl (SIG CLI), and etcd itself (SIG etcd). This distinction underscores the SIG's role as an infrastructure provider rather than an application owner.
Historically, Kubernetes has faced challenges that API Machinery is now actively addressing. Issues like unpredictable performance for list requests, where some requests hit an in-memory cache and others directly queried etcd, led to inconsistent user experience. Upgrades, particularly in high-availability (HA) clusters, have often been complex, with users potentially encountering different API versions mid-upgrade, leading to surprising behavior. The storage and serving of CRDs in JSON format also presented a significant performance bottleneck compared to the more efficient Protobuf used for native types. Furthermore, the reliance on handwritten Go validation code made it difficult for users to ascertain validation rules and for tooling to leverage this information declaratively. These areas form the core problem statements that the Kubernetes 1.33 enhancements aim to resolve, paving the way for a more stable, performant, and user-friendly platform.
Key Findings
▶ Watch: Control plane extensibility: CRDs, webhooks, admission policies (2:00)
The Kubernetes 1.33 release cycle, as detailed by Joe Betz, marks a significant period of evolution for SIG API Machinery, with a strong focus on enhancing upgrade safety, improving control plane performance and stability, and refining the developer experience for Custom Resource Definitions (CRDs). The key findings and contributions can be broadly categorized:
- Enhanced Upgrade Safety and Predictability: Several Alpha-stage KEPs introduce mechanisms to make Kubernetes upgrades more robust and predictable. The Emulation Version allows binaries to "pretend" to be an older version, enabling staged upgrades where binaries are updated without immediately exposing new APIs. The Mixed Version Proxy ensures a consistent API view for users during rolling upgrades of HA clusters by proxying requests between different API server versions. The Beta-stage Coordinated Leader Election refines how controller managers elect leaders, prioritizing older versions during upgrades to prevent version skew-related issues and enabling more controlled leader handovers.
- Significant Performance and Stability Gains: Core API server performance and stability received substantial upgrades. The Snapshotable API Server Cache (Alpha) unifies list request serving, ensuring all list requests benefit from an efficient, cached code path, eliminating unpredictable performance variations. The CBOR Serializer (Alpha) introduces a more efficient binary serialization protocol for CRDs, promising better storage compaction and faster serialization/deserialization compared to JSON. Most notably, Streaming Encoded for List Responses (Beta) dramatically improves API server memory efficiency by streaming list items one at a time, preventing memory exhaustion when handling requests for hundreds of thousands of resources.
- Improved Declarative API Management and CRD Experience: The shift towards declarative APIs is a major theme. Declarative Validation (Beta) moves API validation logic from handwritten Go code to Go tags, making validation rules explicit and enabling richer Open API schemas. CRD Validation Ratcheting (GA) refines how validation changes are handled, ensuring that only fields explicitly updated are re-validated, reducing breaking changes for existing CRDs. Future plans emphasize bringing CRDs further in line with native types, including adding field selectors, additional printer columns with CEL, and more named validation formats for CRD authors.
- Core Mechanics Refinement: The Ordered Namespace Deletion (Alpha) KEP addresses long-standing issues with namespace cleanup by ensuring pods are terminated before other resources, leading to more predictable and stable resource lifecycle management.
These findings collectively underscore a strategic investment in the foundational elements of Kubernetes, aiming to deliver a more resilient, performant, and developer-friendly platform for its extensive user base.
Technical Deep Dive
▶ Watch: Controller infrastructure and control plane performance (3:30)
The Kubernetes 1.33 release cycle showcases a suite of technical advancements from SIG API Machinery, addressing critical areas of performance, stability, and upgrade experience. These enhancements are primarily driven by specific Kubernetes Enhancement Proposals (KEPs), moving through various maturity stages from Alpha to General Availability (GA).
Ordered Namespace Deletion (Alpha):
This KEP tackles a long-standing issue in Kubernetes: the unpredictable order of resource deletion within a namespace. Previously, when a namespace was deleted, resources were removed in an arbitrary order, which could lead to complex edge cases if, for instance, a controller attempted to interact with a pod that was still running but whose dependencies were already gone. The new approach dictates that pods within the namespace are deleted first, and the system waits for them to terminate before proceeding with the deletion of other resources. This ensures a more predictable and stable shutdown sequence for workloads, mitigating unexpected behavior during namespace cleanup. While an Alpha feature, its impact on cluster stability is significant.
Snapshotable API Server Cache (Alpha):
For years, the Kubernetes API server has utilized an in-memory cache to serve watch requests efficiently. However, list requests often had two distinct code paths: one served by this cache and another directly querying etcd. This dichotomy led to unpredictable performance, as users couldn't easily determine which path their request would take. The Snapshotable API Server Cache KEP unifies these code paths. By restructuring the API server's cache, it ensures that all list requests are now served from this efficient, consistent cache. This eliminates performance variability, providing predictable and generally faster responses for list operations, reducing direct load on etcd, and simplifying performance tuning for cluster administrators.
Emulation Version (Alpha, SIG Architecture):
A major initiative for safer upgrades, this KEP introduces an emulation-version flag to core Kubernetes binaries like kube-apiserver. When set to an older Kubernetes version (e.g., 1.31 on a 1.32 binary), the binary will automatically disable new APIs and adjust feature gates to match the specified older version. This allows cluster administrators to perform a two-step upgrade: first, update the binary version while maintaining the old emulation version to test the new binary's stability without changing user-facing APIs; second, update the emulation version to enable new features. This strategy significantly reduces upgrade risk, as rollbacks are simpler if no new APIs have been adopted.
Mixed Version Proxy (Alpha):
Complementing the emulation-version, the Mixed Version Proxy addresses the challenges of rolling upgrades in high-availability (HA) clusters. During an upgrade, an HA cluster might temporarily run API servers of different versions. A client's request, load-balanced across these servers, could encounter inconsistent API behavior. This KEP introduces a mechanism where, if an API server receives a request for an API it cannot serve (due to version differences), it will **proxy that request to a peer API server that can serve it**. This transparently hides version differences from the end-user, ensuring a consistent API surface throughout the upgrade process and preventing unexpected errors.
CBOR Serializer (Alpha):
Custom Resource Definitions (CRDs) have traditionally been served and stored using JSON, which is human-readable but inefficient in terms of storage and serialization/deserialization performance. The CBOR Serializer introduces Concise Binary Object Representation (CBOR), a binary protocol that is functionally equivalent to JSON but significantly more efficient. CBOR is a self-describing format, meaning it retains schema information, but it avoids the parsing overhead of text-based JSON. Implementing CBOR for CRD storage and API serving (when clients request it) promises substantial improvements in performance and storage compaction, bringing CRDs closer to the efficiency of native Kubernetes types that utilize Protobuf.
Declarative Validation (Beta):
Historically, Kubernetes API validation has been implemented through handwritten Go code alongside Go struct definitions. This made it challenging for users and tools to precisely understand validation rules without deep-diving into the source code. The Declarative Validation KEP replaces this handwritten logic with declarative markers on Go struct tags. These markers explicitly define validation rules, making them the single source of truth. This change facilitates automatic generation of internal validation code and, crucially, allows for the publication of enriched validation information through Open API schemas, providing better tooling support and clearer API contracts for developers.
Coordinated Leader Election (Beta):
In HA Kubernetes clusters, multiple kube-controller-manager instances run, but only one is active at a time, elected via a lease mechanism. During upgrades, if an older version reclaims leadership from a newer one, it can lead to surprising behavior and version skew issues. Coordinated Leader Election introduces a more sophisticated approach. Instead of a simple race for a lease, all controller managers announce their candidacy to a centralized leader election coordinator. This coordinator then selects the "best" leader, with the default strategy being to prioritize the oldest version. This prevents violations of version skew, improves upgrade and rollback safety, and even allows for strategies to gracefully ask a current leader to step down, opening doors for more advanced control plane management.
Streaming Encoded for List Responses (Beta):
This is highlighted as one of the most significant performance enhancements in recent releases. Previously, when a client requested a large list of resources (e.g., hundreds of thousands of pods), the API server would accumulate the entire response in memory, serialize it, and then send it. This could lead to significant memory consumption and potential out-of-memory issues for the API server, especially under heavy load. The Streaming Encoded for List Responses KEP changes this by serializing and streaming each list item one at a time. The API server no longer needs to hold the entire response in memory, drastically reducing its memory footprint and allowing it to concurrently handle many large list requests without stability concerns. This is a monumental improvement for API server stability and scalability.
CRD Validation Ratcheting (GA):
Validation rules for API fields can sometimes change, for instance, due to security hardening (like the CVE around IPs and CIDRs in 1.33). If a field's validation tightens, existing, previously valid data might become invalid. Without ratcheting, any update to a resource containing such a field would be rejected, even if that specific field wasn't being modified. CRD Validation Ratcheting addresses this by ensuring that only fields that have actually changed in an update request are re-validated against current rules. If a field's value has not been altered, it is not re-validated, even if it no longer conforms to newly tightened rules. This makes validation changes less disruptive for CRD authors and users, allowing existing resources to continue functioning and only requiring updates to conform when the problematic field is explicitly touched. While still a breaking change in principle, it significantly reduces the immediate impact.
These technical deep dives illustrate SIG API Machinery's ongoing commitment to building a robust, efficient, and user-friendly foundation for Kubernetes.
Demo / Proof of Concept
▶ Watch: Kubernetes 1.33 release updates and KEPs intro (6:00)
This talk, primarily focused on project updates, strategic planning, and the detailed discussion of Kubernetes Enhancement Proposals (KEPs) within the SIG API Machinery domain, did not include a live demonstration or proof-of-concept. Joe Betz's presentation was structured to convey the technical details and implications of the various features and future directions rather than showcasing their live operation. The emphasis was on the architectural changes and their benefits, providing attendees with a comprehensive understanding of the underlying engineering efforts.
Defensive Implications
▶ Watch: KEP: Ordered Namespace Deletion for predictable lifecycle (6:20)
The advancements from SIG API Machinery in Kubernetes 1.33 and beyond have profound implications for cluster administrators, developers, and security practitioners. Understanding these changes is crucial for building more resilient, performant, and secure Kubernetes environments.
- Strategize Safer Upgrades: The Emulation Version and Mixed Version Proxy KEPs are game-changers for upgrade strategies. Cluster administrators should explore adopting a two-phase upgrade process: first, update binaries with the
emulation-versionflag set to the previous stable version to test binary compatibility; then, in a separate step, update theemulation-versionto activate new features. For HA clusters, the Mixed Version Proxy offers a smoother experience by abstracting API version differences during rolling updates. This reduces the risk of unexpected behavior and simplifies rollbacks. The Coordinated Leader Election further enhances upgrade safety by preventing version skew issues during controller manager failovers. - Monitor API Server Performance and Stability: The Snapshotable API Server Cache and, especially, Streaming Encoded for List Responses will significantly improve API server stability and memory efficiency. Defenders should monitor their API server's memory usage and request latency, particularly for large list operations, to observe the benefits. These improvements mean the API server can handle higher loads and larger clusters more reliably, reducing the likelihood of control plane outages due to resource exhaustion.
- Optimize CRD Performance and Management: For those extensively using or developing Custom Resource Definitions (CRDs), the CBOR Serializer offers a path to better performance and reduced storage footprint. While initially an Alpha feature, CRD authors should plan for its adoption to benefit from more efficient serialization/deserialization. The CRD Validation Ratcheting (GA) feature provides a more graceful way to introduce stricter validation rules to CRDs. Developers can update validation schemas with less fear of immediately breaking all existing resources, as only modified fields will trigger re-validation. This allows for a phased approach to tightening security or correctness constraints.
- Leverage Declarative APIs for Security and Compliance: The move towards Declarative Validation will make API contracts clearer. Once this data is exposed via Open API schemas, security tools and policy engines can leverage this richer metadata for more precise static analysis and runtime validation of resources. This can help enforce organizational standards and compliance requirements more effectively.
- Refine Admission Control with CEL: While not a new feature in 1.33, the continued emphasis on admission policies using the Common Expression Language (CEL) provides a powerful, in-cluster alternative to traditional admission webhooks. Defenders can define granular, expressive validation and mutation rules without deploying and managing external webhook services, reducing attack surface and operational complexity.
- Ensure Predictable Resource Cleanup: The Ordered Namespace Deletion improves the predictability of resource lifecycle management. This can prevent orphaned resources or unexpected interactions during namespace termination, which previously could lead to security misconfigurations or resource leaks.
- Stay Informed and Engaged: Given SIG API Machinery's foundational role, staying abreast of their roadmap and participating in discussions is crucial. The SIG's focus on "earning trust with cluster administrators" through safer upgrades and stable APIs indicates a strong commitment to operational excellence that directly benefits security posture.
By actively integrating these insights into operational practices and development workflows, organizations can build more robust, secure, and maintainable Kubernetes platforms.
Key Takeaways
- Upgrade Safety is a Top Priority: Kubernetes is actively enhancing upgrade reliability through features like Emulation Version and Mixed Version Proxy, enabling phased rollouts and a consistent API experience during rolling updates. The Coordinated Leader Election further solidifies this by preventing version skew in HA controller managers.
- Significant Performance & Stability Boosts: The API server will be dramatically more stable and performant due to the Snapshotable API Server Cache unifying list request paths and, most importantly, Streaming Encoded for List Responses which prevents memory exhaustion for large list queries.
- CRDs are Evolving for Better Developer Experience: CBOR Serializer promises performance gains for CRDs, while CRD Validation Ratcheting offers a smoother path for evolving validation rules without breaking existing resources on every update. Future work aims to bring CRDs closer to native types with features like field selectors and richer validation formats.
- Declarative APIs are the Future: The shift to Declarative Validation using Go tags will make API validation rules explicit and machine-readable, improving tooling, documentation, and the overall clarity of API contracts.
- Core Mechanics are Being Refined: Even fundamental processes like namespace deletion are being improved with Ordered Namespace Deletion, ensuring a more predictable and stable resource lifecycle.
- Community Involvement is Encouraged: SIG API Machinery actively seeks contributors, emphasizing its open and collaborative approach to shaping the future of Kubernetes' core components.
About the Speaker(s)
Joe Betz is a Technical Lead for SIG API Machinery at Google. In this role, he is deeply involved in defining and evolving the core REST mechanics, extensibility, and performance of the Kubernetes API. Prior to his work with SIG API Machinery, Joe also served as a maintainer for etcd for several years, giving him a comprehensive background in both the API layer and the underlying distributed key-value store that powers Kubernetes. His expertise spans critical control plane systems, making him a central figure in the development and stability of Kubernetes.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This KubeCon session delivered an unvarnished, deep dive into the foundational improvements coming to Kubernetes API Machinery, presented by the SIG's technical lead. It's a critical update for anyone operating or building on Kubernetes, detailing significant strides in control plane stability, performance, and upgrade safety through genuinely novel engineering solutions. Betz cut through the usual conference fluff, providing concrete details on KEPs that will directly impact operational practices and developer workflows, making this essential viewing for serious practitioners.
Heather Calloway (CISO) — STRONG ACCEPT
This session delivers critical updates on Kubernetes API machinery, focusing on foundational improvements that directly enhance platform stability, operational predictability, and upgrade safety. While deeply technical, the implications for reducing business risk and strengthening the resilience of cloud-native environments are profound, offering security leaders actionable insights for platform strategy and risk management.