OSS = Open Source ... Strategy!? Google Is Doubling Down on K8s in the AI Era, and Yo... ago Macleod
Strategy!? Google Is Doubling Down on K8s in the AI Era, and Yo... ago Macleod
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk, delivered by Macleod at KubeCon EU, provides a comprehensive look into Google's strategic approach to open-source software, particularly its unwavering commitment to Kubernetes in the rapidly evolving landscape of artificial intelligence and machine learning (AI/ML). The presentation delves into the historical context of Kubernetes' development and success, Google's foundational philosophy regarding open source, and the critical "plot twist" introduced by the generative AI boom. Macleod articulates how Google is not just maintaining but actively doubling down on Kubernetes, evolving it to meet the unprecedented demands of AI/ML workloads, specialized hardware, and complex operational environments.

Key moments
- 0:00 Introduction: Google's open source strategy and K8s evolution
- 2:00 Kubernetes' disruption phase and public cloud entry
- 4:00 Key reasons for Kubernetes' success: declarative, extensible, modular
- 4:40 Declarative API and active reconciliation loop explained
- 5:40 Extensibility with CRDs: extending the Kubernetes API
- 6:20 Modularity and the Kubernetes community network effect
- 7:00 Kubernetes' incremental evolution: stateless to stateful applications
- 8:00 Transition to Kubernetes' next phase (2018-2022)
OSS = Open Source ... Strategy!? Google Is Doubling Down on K8s in the AI Era, and Yo... ago Macleod
Speakers: Macleod, Kubernetes Strategy Lead, Google
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=-3NyXaVPGvo
Overview
This talk, delivered by Macleod at KubeCon EU, provides a comprehensive look into Google's strategic approach to open-source software, particularly its unwavering commitment to Kubernetes in the rapidly evolving landscape of artificial intelligence and machine learning (AI/ML). The presentation delves into the historical context of Kubernetes' development and success, Google's foundational philosophy regarding open source, and the critical "plot twist" introduced by the generative AI boom. Macleod articulates how Google is not just maintaining but actively doubling down on Kubernetes, evolving it to meet the unprecedented demands of AI/ML workloads, specialized hardware, and complex operational environments.
The core message revolves around Kubernetes' adaptability and Google's vision to solidify its role as the "hourglass model" for infrastructure consumption in the AI era. The talk is crucial for cloud architects, platform engineers, and business leaders seeking to understand the future trajectory of cloud-native technologies, particularly how a foundational open-source project like Kubernetes can pivot and scale to support the next generation of compute-intensive applications. It highlights Google's internal strategy for differentiation through performance and deep stack integration while contributing extensively back to the open-source community, ensuring Kubernetes remains a universal, powerful, and evolving platform.
Background
▶ Watch: Introduction: Google's open source strategy and K8s evolution (0:00)
Google's engagement with open-source software is deeply ingrained in its corporate DNA, with its open-source office established as early as 2004. This long-standing commitment stems from a belief that the best place to solve common problems is often within the community, fostering collaboration and avoiding the reinvention of wheels. Kubernetes, launched in 2014, represented a pivotal moment for Google, marking its true entry into the public cloud market beyond App Engine. Coinciding with the rise of Docker and containerization, Kubernetes emerged as Google's unique contribution, leveraging its internal expertise in container orchestration derived from projects like Borg.
The initial phase of Kubernetes, from 2014 to 2017, was characterized by disruption. Google's vision was to expand the public concept of cloud from mere Virtual Machines (VMs) to include containers, a fundamental shift given that even Google Cloud's VMs run on containers. This led to the open-sourcing of Kubernetes, its contribution to the Cloud Native Computing Foundation (CNCF), and the rapid growth of a thriving ecosystem. Macleod attributes Kubernetes' success to three core design principles:
- Declarative API: Users declare the desired state, and decoupled controllers independently act to reconcile the system with that state. This contrasts with imperative "do this then that" instructions.
- Extensibility: The ability for the community, end-users, and vendors to extend the Kubernetes API, reducing the core project's burden and fostering innovation.
- Modularity: Components can be swapped out or customized, allowing users to replace parts like the scheduler without abandoning the entire system.
These principles fostered a "network effect" among developers, users, and vendors, leading to continuous feedback loops and incremental evolution. Early on, Kubernetes focused on stateless applications, then incrementally evolved to support stateful applications (e.g., DaemonSets, StatefulSets) and ongoing work for batch workloads. This clear scope definition helped manage expectations and guided development.
The period from 2018 to 2022 saw massive ecosystem expansion, with projects like Istio, OPA Gatekeeper, Argo, Knative, and OpenTelemetry emerging. This growth provided portability across cloud providers and on-prem environments, turning Kubernetes into a powerful distribution channel for solutions. A notable example was the integration of Apache Spark with Kubernetes in 2018, where Google contributed to both projects to make Kubernetes a first-class scheduler for Spark, further expanding its utility. While this ecosystem flourished, it also led to a complex landscape, prompting initial thoughts of a future focused on consolidation, stability, and providing "opinions, not just options" to simplify the user experience. This pre-AI vision anticipated a move towards comprehensive platforms rather than a "bag of Legos."
Key Findings
▶ Watch: Key reasons for Kubernetes' success: declarative, extensible, modular (4:00)
The talk highlights a significant "plot twist" at the end of 2023 with the sudden emergence and widespread adoption of generative AI, epitomized by ChatGPT. This event dramatically altered the trajectory of cloud-native development and Google's strategy. Overnight, the world became obsessed with AI, leading to an unprecedented surge in demand for compute capacity, particularly for specialized hardware. Macleod vividly describes the "pretty hilarious" shift in capacity projections, noting that while shipping electrons is easy, shipping atoms (hardware) is not, resulting in significant capacity constraints across cloud providers.
This AI boom created "crazy overlapping stacking adoption curves," where traditional enterprises continued their journey to cloud-native adoption, while a new wave of users simultaneously pushed the boundaries with AI/ML workloads. This necessitates a delicate balance between innovation and stability. The consumption growth has been "bananas," with 10x increases forcing a fundamental rethinking of existing approaches. While such growth often signals an opening for entirely new platforms, Google made a deliberate decision to evolve Kubernetes.
The key challenges identified for these new AI/ML workloads include:
- Mission-criticality: Moving into sensitive sectors like healthcare and telco.
- Specialized, Sparse, and Expensive Hardware: A departure from the assumption of infinitely elastic, fungible cloud resources. GPUs, TPUs, and other accelerators are not interchangeable, are geographically constrained, and come at a high cost.
- Dynamic Workloads: AI/ML tasks are often ephemeral, multi-cluster, and have specific hardware requirements, where topology and proximity (networking throughput) are paramount.
- Purpose-Built Frameworks: The need to support specialized frameworks that understand and leverage underlying hardware efficiently.
Despite these challenges, Google chose to stick with Kubernetes, guided by the concept of a "path dependence feedback loop." The existing community, vendor ecosystem, and established tooling represent a massive investment and a strong foundation. This led to Google's commitment to rally around three core stories for Kubernetes' evolution to meet the needs of the "next trillion core hours":
- Reliability at Scale and Across Upgrades: Addressing feedback that upgrades remain problematic, ensuring continuous operation for critical applications.
- Redefining Kubernetes' Relationship with Hardware: Moving beyond the "a node is a node is a node" mentality to natively support and optimize for specialized, non-fungible accelerators.
- Purpose-Built Frameworks as First-Class Citizens: Embracing and integrating specialized schedulers and frameworks (like Slurm for HPC or various AI frameworks) as native components within Kubernetes, rather than treating them as external competitors.
These findings underscore Google's strategic bet on Kubernetes' inherent design strengths—declarative, extensible, modular—as crucial for its adaptation to the AI era, rather than abandoning it for a new platform.
Technical Deep Dive
▶ Watch: Extensibility with CRDs: extending the Kubernetes API (5:40)
Kubernetes' foundational design principles are central to its adaptability in the AI era. The declarative API is powered by an active reconciliation loop: users specify a desired state, controllers observe the current state, compare the two, and then act to bring the system into alignment. This continuous loop, depicted in diagrams for over a decade, is a powerful mechanism for managing complex distributed systems.
Extensibility is further enhanced by Custom Resource Definitions (CRDs), which became Generally Available (GA) in 2019. CRDs allow users to extend the Kubernetes API with custom resources that behave like built-in types, using the same tooling. Macleod notes that CRDs are now often superior to built-in types, even suggesting a future where core functionality might be implemented out-of-tree using CRDs, as demonstrated by the Gateway API. This capability offloads development from the core Kubernetes team while empowering the community to innovate.
Modularity allows for components like the scheduler to be replaced or augmented. This is critical for supporting diverse workloads, including specialized AI/ML tasks that might require custom scheduling logic. The network effect among developers, users, and vendors, driven by these design principles, has created a self-reinforcing growth cycle for the community and ecosystem.
The incremental evolution of Kubernetes, from stateless to stateful and now increasingly to batch and dynamic AI workloads, demonstrates its capacity for measured growth. The successful integration of Apache Spark in 2018 serves as a precedent, where contributions to both Spark and Kubernetes made the latter a viable distribution channel for the former, overcoming initial incompatibilities.
For AI/ML, the most significant technical challenge is redefining Kubernetes' relationship with hardware. The traditional view of CPU and memory as fungible resources, where "a node is a node," no longer holds for GPUs, TPUs, and other accelerators. These are expensive, sparse, and non-fungible, requiring a nuanced approach to resource management. Google has heavily invested in initiatives like Dynamic Resource Allocation (DRRA), originally started by Intel and Nvidia. The goal is to make DRRA more Kubernetes-native and portable, abstracting away the complexities of driver versions and hardware specifics to provide consistent access across different accelerators and cloud providers.
Furthermore, Google's strategy involves supporting framework orchestration by positioning Kubernetes as a distribution channel rather than a competitor. This means enabling specialized schedulers like Slurm (widely used for High-Performance Computing, HPC) and various AI frameworks to run effectively on Kubernetes. The concept of guest schedulers allows Kubernetes to support these frameworks as first-class citizens, leveraging their decades of specialized experience while providing a consistent operational model for platform teams. This prevents the need for platform teams to learn entirely new ways to operate, manage chargebacks, and secure each new framework.
This vision culminates in Kubernetes becoming the "hourglass model" of infrastructure consumption, similar to how IP (IPv4/IPv6) serves as the narrow waist of the internet. By establishing Kubernetes as the consistent layer in the middle, it can support a wide array of evolving technologies above (browsers, mail clients) and diverse physical layers below.
Google's differentiation strategy, despite its open-source commitment, focuses on performance. This is achieved through deep integration across all layers of its stack within Google Cloud (GCE, networking, TPUs), optimizing the entire pipeline. They leverage asymmetric advantages, such as using Spanner instead of etcd to support 65,000 nodes on GKE, demonstrating how standard interfaces can be backed by highly optimized, proprietary implementations. This approach ensures that the standard remains open while the underlying implementation provides competitive performance.
The "stickiness" of Kubernetes, often a concern for open-source products, is achieved through the vibrant ecosystem. Users remain engaged because of the functionality provided by projects like Argo and OpenTelemetry, making Kubernetes a de facto sticky platform. Google's internal GKE feedback loop is critical: features are launched in GKE, observed in real-world fleets, and then refined and contributed back to upstream open-source Kubernetes, ensuring that innovations benefit everyone and solve practical problems. The ultimate goal is to move from scenarios where applications "succeed in spite of Kubernetes" to where they "work because of Kubernetes," with native, full-stack support for the most demanding AI/ML workloads, from DRA at the hardware layer to multi-cluster solutions like Q and MultiQ at the application layer.
Demo / Proof of Concept
▶ Watch: Modularity and the Kubernetes community network effect (6:20)
The talk does not include a live demonstration or a specific proof of concept of the discussed technical advancements. Instead, Macleod emphasizes that the strategies and evolutions outlined are the subject of active, ongoing work within the Kubernetes community and Google. He references a recent maintainer summit at KubeCon EU where discussions were held about "evolving Kubernetes for AI/ML across these layers," pointing to a GitHub issue for further details. This indicates that the presented concepts are in various stages of development and community collaboration, rather than being fully realized, demonstrable features at the time of the talk.
Defensive Implications
▶ Watch: Transition to Kubernetes' next phase (2018-2022) (8:00)
While the talk primarily focuses on strategic evolution and platform capabilities, several aspects have significant defensive implications for organizations leveraging Kubernetes, particularly in the context of AI/ML workloads. The emphasis on reliability at scale and across upgrade boundaries directly contributes to a more resilient and secure infrastructure. Stable and predictable upgrades reduce downtime and the risk of misconfigurations or vulnerabilities introduced during complex update processes.
Macleod's acknowledgment that "how hard could it be to lock it down and it turns out it'd be really hard" when discussing new frameworks highlights the inherent security challenges with diverse and rapidly evolving technologies. By aiming for a consistent operational model across various purpose-built frameworks (like Slurm or AI frameworks), Kubernetes can provide a unified control plane for security, policy enforcement, and compliance. This reduces the attack surface by centralizing management and preventing disparate, insecure configurations across different compute environments. Leveraging established Kubernetes security features for new workloads, rather than re-implementing security for each framework, is a powerful defensive strategy.
Furthermore, Google's commitment to the GKE feedback loop, where real-world problems observed in GKE are addressed and fixed in upstream open-source Kubernetes, means that security enhancements and bug fixes are continually contributed back to the entire community. This collaborative approach ensures that security vulnerabilities discovered and patched by Google ultimately benefit all Kubernetes users, strengthening the platform's overall defensive posture. The drive towards native support for specialized hardware and dynamic workloads, through initiatives like DRRA, also implies a more secure integration, where resource isolation and access controls can be consistently applied, preventing unauthorized access or misuse of sensitive and expensive accelerators.
Key Takeaways
- Kubernetes' Core Strengths Drive AI Adaptation: The platform's declarative API, extensibility (via CRDs), and modularity are fundamental to its ability to evolve and meet the demands of emerging AI/ML workloads, rather than being replaced by a new technology.
- AI/ML Demands a Rethink of Infrastructure: The "plot twist" of generative AI has led to unprecedented capacity constraints and the need to manage specialized, sparse, and expensive hardware (GPUs, TPUs) that challenge Kubernetes' traditional fungible-resource model.
- Google's Strategic Pillars for Kubernetes in AI: Google is doubling down on Kubernetes by focusing on three key areas: enhancing reliability at scale, redefining its relationship with specialized hardware through projects like DRRA, and integrating purpose-built AI/HPC frameworks as first-class citizens.
- Kubernetes as the "Hourglass Model": Google envisions Kubernetes as the universal "narrow waist" for infrastructure consumption, providing a consistent layer that enables independent evolution of both underlying hardware and upper-layer applications and frameworks.
- Performance Differentiation in Open Source: Despite contributing heavily to open source, Google differentiates its cloud offerings (GKE) through deep integration across its stack and leveraging asymmetric advantages (e.g., Spanner for etcd) to deliver superior performance and scale.
- Community-Driven Evolution and Feedback Loops: Google's strategy relies on continuous collaboration with the open-source community, using real-world observations from GKE to drive upstream Kubernetes improvements, ensuring practical problem-solving and shared innovation.
About the Speaker(s)
The speaker, Macleod, is a key figure in Google's Kubernetes strategy. Their background includes working at Nest (then part of Google) in 2015-2016 as an early user of Kubernetes, even running production workloads on it when it was "not wise." Macleod subsequently moved internally within Google in early 2017 to work directly on Kubernetes, indicating a deep and long-standing involvement with the project from both a user and contributor perspective. This experience has provided them with a unique understanding of Kubernetes' evolution, challenges, and strategic direction within Google and the broader cloud-native ecosystem.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
Macleod's KubeCon talk provides a clear, substantive strategic update on Google's unwavering commitment to Kubernetes in the AI era. It's not a zero-day, but it's a critical signal for anyone building on cloud-native infrastructure. The talk effectively articulates why Kubernetes' core design principles make it adaptable to the demands of AI/ML, specialized hardware, and complex operational environments, rather than being replaced by a new platform. This isn't just marketing fluff; it's a detailed explanation of Google's technical and strategic pivot, offering valuable insight into the future trajectory of compute infrastructure.
Heather Calloway (CISO) — STRONG ACCEPT
This talk from Macleod at KubeCon EU provides a clear and essential strategic overview of Google's commitment to Kubernetes as the foundational infrastructure for the AI era. It effectively translates the technical evolution of a critical open-source platform into a compelling narrative for business leaders and CISOs, highlighting the institutional decisions and strategic bets required to manage the unprecedented demands of AI/ML workloads. The speaker lays out a credible vision for how Kubernetes' inherent design principles will allow it to adapt, offering a necessary perspective for anyone grappling with the future of cloud-native security and resilience.