AI Beyond Autocomplete: Using LLMs To Create 1000 Kubernetes... Justin Santa Barbara & Walter Fender
Justin Santa Barbara, Walter Fender
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In this insightful KubeCon EU talk, Justin Santa Barbara and Walter Fender from Google delve into their innovative approach to scaling Kubernetes controller development by leveraging Large Language Models (LLMs). Their project, Config Connector, aims to bridge the gap between Google Cloud Platform (GCP) REST APIs and the Kubernetes Resource Model (KRM), necessitating the creation of approximately a thousand distinct Kubernetes controllers. This talk, "AI Beyond Autocomplete: Using LLMs To Create 1000 Kubernetes...", details how they navigated the complexities of such a massive undertaking, moving away from traditional "magic machine" architectures to an LLM-driven, build-time code generation pipeline.

Key moments
- 0:00 Introduction: Config Connector's challenge of 1000 Kubernetes controllers
- 2:00 Why traditional "magic machine" approaches like Terraform failed
- 3:00 Key breakthrough: "Code as the artifact" for managing complexity
- 5:00 Previous A-based tooling limitations and LLM emergence
- 6:00 Advantages of using LLMs for code generation (flexibility, speed)
- 7:00 Addressing challenges: LLMs' non-deterministic nature and fears
AI Beyond Autocomplete: Using LLMs To Create 1000 Kubernetes...
Speakers: Justin Santa Barbara, Software Engineer, Google; Walter Fender, Software Engineer and EM, Config Connector Project, Google
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=_oIoaW5i-xE
Overview
In this insightful KubeCon EU talk, Justin Santa Barbara and Walter Fender from Google delve into their innovative approach to scaling Kubernetes controller development by leveraging Large Language Models (LLMs). Their project, Config Connector, aims to bridge the gap between Google Cloud Platform (GCP) REST APIs and the Kubernetes Resource Model (KRM), necessitating the creation of approximately a thousand distinct Kubernetes controllers. This talk, "AI Beyond Autocomplete: Using LLMs To Create 1000 Kubernetes...", details how they navigated the complexities of such a massive undertaking, moving away from traditional "magic machine" architectures to an LLM-driven, build-time code generation pipeline.
The core challenge addressed is the inherent difficulty and inefficiency of manually developing and maintaining a vast number of highly specialized controllers, each with its own business logic, for every GCP resource. The speakers articulate a paradigm shift: instead of shipping a single, complex runtime system, they generate a multitude of simple, isolated code artifacts at build time. This strategy embraces the non-deterministic nature of LLMs, viewing them as powerful tools for efficiency when their "magic" is confined to the build phase, rather than impacting runtime stability.
This talk is crucial for anyone interested in the practical application of AI in software engineering, particularly within the Kubernetes ecosystem. It offers a blueprint for overcoming the scalability hurdles of managing cloud resources through KRM, providing actionable strategies for integrating LLMs into complex development workflows. By focusing on code as the primary artifact and emphasizing robust validation and iterative improvement, Santa Barbara and Fender demonstrate a pragmatic path to harnessing AI for large-scale code generation, ensuring maintainability and reliability.
Background
▶ Watch: Introduction: Config Connector's challenge of 1000 Kubernetes controllers (0:00)
The genesis of Config Connector lies in the ambition to manage Google Cloud Platform resources directly through Kubernetes. GCP, like other major cloud providers such as AWS with ACK and Azure with ASO, exposes a vast array of services and functionalities via REST APIs. To make these cloud resources Kubernetes-native, Config Connector translates each GCP API into a Kubernetes Custom Resource Definition (CRD) and a corresponding controller. The sheer scale of GCP's offerings, comprising approximately a thousand distinct REST APIs, translates directly into the need for a thousand unique CRDs and their associated controllers, each requiring individual business logic.
Early attempts to tackle this problem involved centralizing logic within a single, monolithic "magic machine," akin to how some Terraform providers operate. While Terraform is a powerful tool, the speakers found that building a single, intricate controller to manage a multitude of resources led to significant runtime complexity. Changes to one resource's logic frequently broke others, creating a fragile and difficult-to-maintain system. This "magic machine" became a bottleneck, turning simple fixes into week-long escapades due to the interconnectedness and opacity of its internal workings. The team concluded that the complexity inherent in this approach made it impossible to scale to the target of a thousand controllers.
Recognizing the limitations of runtime complexity, the Config Connector team shifted their philosophy towards code as the primary artifact. Their new goal was to generate a large volume of simple, isolated, and self-contained code. This meant moving the "magic" from runtime execution to the build process. The emergence of Large Language Models (LLMs), particularly around the "ChatGPT moment," perfectly coincided with this shift. LLMs, with their inherent ability to manipulate and generate code, offered a promising avenue. Unlike deterministic, AST-based tooling, which also generated simple code but required immense human investment in complex build systems, LLMs could produce results with less heavy lifting, even if not 100% reliable. The key insight was that non-determinism was acceptable at build-time, as long as the generated code itself was simple and verifiable. This allowed the team to leverage the efficiency of LLMs without burdening customers with runtime complexity or the team with an unmanageable build system.
Key Findings
▶ Watch: Key breakthrough: "Code as the artifact" for managing complexity (3:00)
The talk highlights several pivotal findings in applying LLMs to large-scale Kubernetes controller generation:
- Build-Time Magic, Runtime Simplicity: The most significant finding is that LLMs, while often perceived as "magic machines," are incredibly effective when their non-deterministic nature is confined to the build process. By generating simple, isolated code artifacts at build time, the complexity is shifted away from the production runtime, making the resulting system more stable and maintainable. This allows for iteration and error correction during development without impacting customer environments.
- Embracing Non-Determinism with Validation: Unlike traditional deterministic tooling, LLMs frequently produce varying outputs. The team found that accepting this non-determinism and incorporating strategies like running generation multiple times ("run it twice" or "thrice") and robust validation was more effective than trying to force deterministic behavior. This approach capitalizes on the LLM's generative power while mitigating its unpredictability.
- The "Bitter Lesson" of AI in Practice: Drawing from Rich Sutton's essay, the speakers emphasize the importance of focusing on broad, enduring techniques rather than investing heavily in model-specific optimizations (e.g., extensive prompt tweaking). LLMs will continuously improve, and new models will likely invalidate prior, highly specific prompt engineering efforts. Instead, success lies in optimizing for information context (garbage in, garbage out), problem decomposition, feedback loops, and tool exposure, which remain effective across different LLM generations.
- Induction for Complex Tasks: For more intricate code generation requirements, a technique called induction proved highly effective. This involves providing the LLM with a set of existing "input-output" examples from the codebase (e.g., annotations and the corresponding generated fuzzer code). The LLM then learns from these examples to generate the next
n+1case, significantly improving the quality and structure of complex outputs compared to simple "vibe coding" prompts.
- Jigs and Interlocks for Robustness: Breaking down the overall controller generation into a series of 12 to 15 smaller, independent, and validated steps (referred to as "jigs" and "interlocks") is crucial. Each step performs a specific action and includes validation checks. This modular approach ensures that errors or hallucinations in one step are caught early, preventing cascading failures and making the entire pipeline more reliable. It also allows for parallel processing and targeted debugging.
- Hybrid Approach with Traditional Tools: Not all steps in the generation pipeline need to be LLM-based. The team successfully integrates traditional code generation tools (e.g., CRD generators for OpenAPI schemas) alongside LLM-driven steps. This hybrid solution leverages the strengths of each technology, using LLMs for creative or complex generation and deterministic tools for well-defined, schema-driven tasks.
- Scaling Trust Through Diverse Validation: Generating a thousand controllers means human review is insufficient for ensuring correctness. The team explores various methods to scale trust, including linters (for CRDs and code), automatically generated tests (comparing mocks with live system behavior), and even using different LLMs specifically for validation tasks. The goal is to offload validation from human engineers, allowing them to focus on high-value tasks like API design review.
Technical Deep Dive
▶ Watch: Previous A-based tooling limitations and LLM emergence (5:00)
The technical implementation described by Justin Santa Barbara and Walter Fender is a sophisticated, multi-stage pipeline designed to generate a thousand Kubernetes controllers, prioritizing simplicity of the end product over simplicity of the generation process.
The foundational principle is "code as the primary artifact." This means that the output of the LLM-driven pipeline is standard, human-readable Go code, YAML, and other configuration files. The LLMs' "magic" is applied at build-time to produce this code, which is then managed, reviewed, and run like any other codebase. This stands in stark contrast to earlier "magic machine" approaches where complex logic resided within a single runtime component.
The speakers describe a gradual evolution of their generation strategy, starting with simple prompt templating and tool usage. For basic tasks, they employ jigs that inject variables into prompts and expose functions to the LLM, such as write_file or the ability to run gcloud help. An example shown involves the LLM reading gcloud help files to construct simple test cases. This works well for straightforward, "vibe coding" type tasks, where the LLM can generate code based on natural language instructions and context.
However, for more complicated tasks, a more structured approach was necessary. This led to the development of what they call induction. The core idea of induction is to provide the LLM with concrete examples of desired input-output pairs. For instance, to generate a fuzzer for a TPU virtual machine, they hand-code a few initial examples. Each example includes structured input data, typically in the form of annotations (e.g., fuzzgen, proto representation, CRD kind), and the corresponding generated fuzzer code as the output. These input-output pairs are then wrapped in XML (a current, though acknowledged as imperfect, method for structuring context) and fed to the LLM. The LLM then uses these examples to learn the pattern and generate the n+1 case for a new resource. This iterative, example-driven approach significantly improves the quality of complex code generation.
The entire process of generating a single controller is broken down into approximately 12 to 15 distinct steps, each acting as a "jig" with its own "interlock" (validation). This decomposition is critical for managing LLM hallucinations and ensuring correctness. The hypothesis is that while an LLM might hallucinate in one step, it's unlikely to hallucinate in a way that aligns perfectly with correct outputs from other, independent steps.
A typical pipeline might look like this:
- Metadata Generation: An initial, one-time LLM step generates high-level metadata for all thousand resources (e.g., Git branch, resource name, proto file location). This metadata often requires human review to correct hallucinations.
- G-Cloud Command Generation: LLMs generate
gcloudcommands to interact with real GCP APIs. - HTTP Log Capture: These commands are executed, and the HTTP request/response logs are captured.
- Validation: Basic checks are performed on the logs (e.g., no
404errors). If errors occur, the process for that resource is paused for investigation, while others proceed in parallel. - Mock Generation: Using the captured HTTP logs and potentially other templates, LLMs generate mocks for API interactions.
- CRD Structure Generation: For generating the Go struct that represents the CRD, traditional tooling (like existing CRD generators) is often preferred over LLMs due to its deterministic nature.
- Controller Code Generation: LLMs generate the core controller logic, using the CRD structure and mocks as context.
- Compile Error Fixing: The generated code is compiled. For certain classes of errors, such as missing imports, LLMs are surprisingly effective at suggesting fixes. More complex compile errors still require human intervention.
- Further Validation: Each step is followed by validation. This can include static analysis, linters, and eventually, running tests.
The pipeline is designed for parallel execution. For instance, the first step might run across all 1000 resources. If 600 succeed and 400 fail, the successful 600 can proceed to the next step while engineers debug the 400 failures. This maximizes throughput.
A crucial aspect is iterative improvement tools. The system provides LLMs with extensive context, such as Kubernetes OpenAPI schemas (downloaded dynamically), mock data, and build outputs. This rich context helps the LLM avoid hallucinations and understand when it has made a mistake. While the "agentic workflow" (where an LLM is given tools and asked to fix errors autonomously) hasn't fully worked for them yet, they anticipate future LLM improvements will make this more viable. Currently, complex compiler errors still often require human intervention.
Finally, the talk emphasizes the importance of recording intention and results. This means logging what was asked of the LLM, its output, and any validation feedback. This data is invaluable for debugging, iterating on prompts, and improving the overall generation process. The team even maintains two different generation paths: one for "greenfield" new resources and another for "brownfield" resources that need to maintain backward compatibility with existing Terraform implementations.
Demo / Proof of Concept
▶ Watch: Advantages of using LLMs for code generation (flexibility, speed) (6:00)
While the talk did not feature a live, interactive demonstration in the traditional sense, it provided concrete examples and detailed insights into the generated artifacts and the process itself, serving as a robust proof of concept for their LLM-driven code generation system.
The speakers showcased actual code from the KCC project, specifically illustrating a fuzzer for a TPU virtual machine. This snippet highlighted the structured input data provided to the LLM through annotations, such as fuzzgen, the proto representation, and the CRD kind (TPU virtual machine). This visual example demonstrated how the LLM receives context and generates the corresponding fuzzer code, effectively acting as an input-output pair for the induction process. The method of wrapping these inputs and outputs in XML to feed to the LLM was also described, offering a glimpse into the practical implementation details.
Furthermore, the entire Config Connector project, with its ambitious goal of generating 1000 Kubernetes controllers, stands as a large-scale proof of concept. The discussion of the 12-15 step pipeline, including the generation of G-Cloud commands, capture of HTTP logs, creation of mocks, and the iterative refinement of code, demonstrates a functional system capable of producing complex, production-ready Kubernetes components. The mention of achieving 600 successful generations out of 1000 in a single overnight run underscores the system's operational capability and efficiency. The speakers also affirmed that the code for their tooling is available on GitHub, inviting further review and validation of their approach.
Defensive Implications
▶ Watch: Addressing challenges: LLMs' non-deterministic nature and fears (7:00)
For organizations considering or implementing LLM-driven code generation, the talk provides crucial "defensive" strategies to ensure the robustness, reliability, and maintainability of the generated artifacts. These implications are vital for defending against the inherent challenges of LLMs, such as hallucinations and non-determinism, and for building a resilient development pipeline.
- Prioritize Rigorous Validation: The most critical defensive measure is comprehensive validation at every stage. Given that LLMs are non-deterministic and prone to hallucinations, generated code cannot be blindly trusted. Implement linters (for both CRDs and generated code), static analysis tools, and automated tests. The speakers highlight generating tests that can be run against both mocks and live systems, increasing confidence if results align. Consider using a different LLM specifically trained or prompted to identify potential problems in the generated code, acting as an automated "reviewer."
- Embrace Incremental, Interlocked Steps: Break down complex code generation into numerous small, independent, and validated steps (the "jigs and interlocks"). This modularity acts as a defensive barrier, catching errors early in the pipeline. If a hallucination occurs in one step, subsequent validation checks are more likely to detect it, preventing a flawed artifact from progressing further and compounding issues. This parallel processing and error isolation significantly improve debugging efficiency.
- Focus on Data Context, Not Model-Specific Hacks: Adhering to the "Bitter Lesson," organizations should defend against wasting resources on fleeting model-specific prompt optimizations. Instead, invest in robust methods for providing rich, accurate information context to the LLM (e.g., OpenAPI schemas, API documentation, existing code examples, HTTP logs). "Garbage in, garbage out" remains a fundamental truth. A well-structured input context is a stronger defense against poor outputs than intricate prompt engineering that may be obsolete with the next model release.
- Human-in-the-Loop for High-Value Tasks: While LLMs excel at generating boilerplate and simple logic, humans remain critical for tasks requiring deep understanding, subjective judgment, or complex error resolution. For example, API review (ensuring a CRD's API design is sound and maintainable long-term) is best left to experienced engineers. Similarly, complex compile errors or subtle logical flaws often require human debugging. Define clear boundaries where human expertise is indispensable, and where LLMs can augment efficiency.
- Record Intention and Results for Debugging: To effectively "defend" against pipeline failures and improve future generations, meticulously record what was asked of the LLM (the prompt and context) and its exact output, along with any validation feedback. This audit trail is invaluable for understanding why a generation failed, identifying patterns of LLM misbehavior, and iteratively refining the prompts, jigs, and interlocks.
- Architect for Scalability and Simplicity: Design the target code architecture to be composed of many simple, isolated units, even if it means generating more total code. This defends against the "magic machine" problem, where a single complex component becomes a maintenance nightmare. Simpler, self-contained code is easier for humans and other tools to understand, debug, and modify, regardless of how it was generated.
Key Takeaways
- Shift Complexity to Build-Time: LLMs are powerful for generating large volumes of simple, isolated code at build-time, effectively moving "magic machine" complexity away from runtime production systems.
- Embrace Non-Determinism Strategically: Accept that LLMs are non-deterministic for build-time generation. Mitigate this by running generation multiple times and implementing robust validation steps rather than trying to force predictable outputs.
- Adhere to the "Bitter Lesson": Focus on broad, architectural strategies like providing rich information context, decomposing problems, and implementing feedback loops, rather than investing heavily in ephemeral model-specific prompt optimizations.
- Decompose and Interlock: Break down complex code generation into many small, independent, and validated steps ("jigs" and "interlocks"). This modular approach improves reliability, catches errors early, and enables parallel processing.
- Utilize Hybrid Tooling: Combine LLM-driven generation for creative or complex tasks with traditional, deterministic tooling (e.g., CRD generators) for well-defined, schema-based processes.
- Prioritize Validation and Trust: Implement comprehensive validation using linters, automated tests (comparing mocks to live systems), and even specialized LLMs for problem detection to build trust in the generated code, especially when human review is impractical at scale.
About the Speaker(s)
Justin Santa Barbara is a Software Engineer at Google. His work primarily focuses on open-source projects related to Kubernetes, including significant contributions to Config Connector. He is deeply involved in the architectural and implementation aspects of making cloud resources Kubernetes-native.
Walter Fender also works at Google as both a Software Engineer and an Engineering Manager (EM). He plays a key leadership role on the Config Connector project, as well as contributing to several other open-source Kubernetes initiatives. Walter was instrumental in developing the core idea of "code as the artifact" and shifting complexity to build time, which underpins the LLM-driven generation strategy discussed in the talk.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk presents a genuinely novel and rigorously engineered approach to scaling Kubernetes controller development using LLMs. By shifting complexity to a build-time pipeline and focusing on robust validation and problem decomposition, Santa Barbara and Fender offer a blueprint for leveraging generative AI in a production-ready, maintainable manner. This isn't just another 'AI-powered' fluff piece; it's a deep dive into practical, large-scale code generation that addresses real-world engineering challenges.
Heather Calloway (CISO) — STRONG ACCEPT
This talk presents a pragmatic, engineering-led approach to managing the inherent risks of leveraging Large Language Models (LLMs) for large-scale code generation. By shifting complexity to build-time and emphasizing rigorous validation, decomposition, and a human-in-the-loop, the speakers offer a credible blueprint for how organizations can responsibly adopt AI in their software development. It provides essential insights for CISOs and executive leaders on establishing the necessary controls and accountability when deploying non-deterministic systems to create critical infrastructure components.