Don't Write Controllers Like Charlie Don't Does: Avoiding Common Kubernetes Controller... Nick Young

Nick Young

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This article delves into the critical and often challenging domain of writing Kubernetes controllers, particularly those interacting with Custom Resource Definitions (CRDs). Presented by Nick Young from Isovalent at Cisco, the talk, humorously titled "Don't Write Controllers Like Charlie Don't Does," serves as a practical guide to identifying and circumventing common anti-patterns and pitfalls that can lead to inefficient, unstable, or incorrect controller implementations. Young leverages his extensive experience, including personal mistakes, to illuminate the subtle complexities inherent in managing Kubernetes' eventually consistent state.

Watch on YouTube

Visual summary for Don't Write Controllers Like Charlie Don't Does: Avoiding Common Kubernetes Controller... Nick Young by Nick Young
Visual summary for Don't Write Controllers Like Charlie Don't Does: Avoiding Common Kubernetes Controller... Nick Young by Nick Young

Key moments

  1. 0:00 Introduction and speaker's experience with controllers
  2. 2:50 Introducing 'Charlie Don't' and common controller mistakes
  3. 3:00 Mistake: Using a simple client for every API call
  4. 3:50 Consequence: API server overload from excessive requests
  5. 4:20 Solution: Utilize a cached client to reduce API calls
  6. 5:00 Crucial tip: Check for differences before sending updates
  7. 6:00 Advanced tip: Use patch instead of update to avoid races

Don't Write Controllers Like Charlie Don't Does: Avoiding Common Kubernetes Controller Mistakes

Speakers: Nick Young, Engineer, Isovalent at Cisco

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=tnSraS9JqZ8

Overview

This article delves into the critical and often challenging domain of writing Kubernetes controllers, particularly those interacting with Custom Resource Definitions (CRDs). Presented by Nick Young from Isovalent at Cisco, the talk, humorously titled "Don't Write Controllers Like Charlie Don't Does," serves as a practical guide to identifying and circumventing common anti-patterns and pitfalls that can lead to inefficient, unstable, or incorrect controller implementations. Young leverages his extensive experience, including personal mistakes, to illuminate the subtle complexities inherent in managing Kubernetes' eventually consistent state.

The talk is crucial for anyone involved in extending Kubernetes functionality, from platform engineers developing custom operators to application developers building sophisticated cloud-native applications. In an ecosystem increasingly reliant on custom resources to manage complex infrastructure and application states, understanding how to write robust, performant, and correctly behaving controllers is paramount. The insights shared aim to prevent widespread cluster instability, performance bottlenecks, and difficult-to-debug reconciliation issues that often arise from naive or flawed controller designs.

By dissecting typical errors and offering prescriptive solutions, Young empowers developers to build more resilient and maintainable Kubernetes extensions. The discussion covers essential topics such as efficient API server interaction, the benefits and nuances of various controller frameworks, and critical architectural considerations for ensuring correct state reconciliation within a distributed, eventually consistent system. The overarching goal is to equip practitioners with the knowledge to avoid the "Charlie Don't" mistakes, fostering a more stable and efficient Kubernetes environment.

Background

▶ Watch: Introduction and speaker's experience with controllers (0:00)

Nick Young's journey into Kubernetes controllers began early, in 2017, when CRDs were still known as "third party resources." His involvement in significant projects, such as the redesign of Contour's HTTPProxy CRD and his foundational work on the Gateway API since its inception in 2018—a specification delivered purely using CRDs—underscores his deep practical expertise. Crucially, Young openly admits to having "screwed it up plenty of times," an admission that lends significant credibility to his advice, emphasizing that controller development is "surprisingly difficult" even for seasoned professionals.

To make the technical content more engaging, Young introduces "Charlie Don't," a straw man character inspired by The Simpsons episode "Bart Got a Knife" and its memorable "Don't Do What Donny Don't Does" book. Charlie Don't, a developer at a "big co" working on a custom Kubernetes controller, consistently makes the wrong decisions, serving as a relatable proxy for common developer missteps. This narrative device helps to illustrate complex problems in an accessible and memorable way.

This talk builds upon a previous presentation by Young, which focused on designing CRDs effectively. That prior work covered principles like reading API Bibles, considering user experience, leveraging status fields, avoiding problematic value types, and ensuring compatible API changes—all critical for future-proofing CRDs. The current discussion shifts focus from API design to the implementation details of the controllers themselves, primarily within the Go language, acknowledging that while other languages like QRS exist, the Go ecosystem is where most of his experience lies. The fundamental problem addressed is how to correctly and efficiently manage and reconcile state in an eventually consistent system like Kubernetes, where direct, synchronous interactions are often detrimental.

Key Findings

▶ Watch: Mistake: Using a simple client for every API call (3:00)

The talk meticulously dissects several common anti-patterns and critical design flaws encountered when developing Kubernetes controllers. These key findings highlight not only what mistakes to avoid but also the underlying reasons why they are problematic:

  • Inefficient API Server Interaction: A primary mistake is directly querying the Kubernetes API server for every piece of information. Charlie Don't's initial approach of making thousands of Get and List calls per second leads to significant strain on the API server and etcd, potentially bringing the entire cluster to its knees. This is a fundamental violation of Kubernetes controller best practices, which advocate for event-driven reconciliation and local caching.
  • No-Op Updates: Related to inefficient API interaction, controllers often perform "no-op updates" – sending a full object update back to the API server even when no actual change has occurred. This consumes unnecessary API server and etcd resources, as the server still has to process, validate, and store the object, even if it's identical to the previous version. This is particularly problematic for status updates, which can be very frequent.
  • Hand-Rolling Caching Clients: While moving away from direct API calls is good, attempting to build a custom caching client using client-go's informers without a framework is fraught with peril. This approach forces the developer to manually handle complex concurrency, ordering, and eventual consistency problems. Distinguishing between a dependent object that genuinely doesn't exist versus one that simply hasn't propagated to the local cache yet becomes an arduous task, leading to over-processing and difficult-to-debug states.
  • Incorrect Predicate Logic: A subtle but impactful error arises when predicates—functions that filter which events trigger a reconciliation—are not applied consistently. Specifically, predicates used with the main resource's watch (e.g., For in controller-runtime) must be mirrored or carefully considered for watches on dependent resources. Failing to do so can result in the controller unnecessarily reconciling objects it doesn't care about, wasting CPU cycles and potentially causing race conditions.
  • Reconciling the Wrong Object: Perhaps the most critical design flaw discussed is choosing an inappropriate object as the primary target for reconciliation within a hierarchical CRD structure. For instance, in the Gateway API, if a controller reconciles HTTPRoute objects instead of Gateway objects, it might inadvertently overwrite configuration or fail to aggregate state correctly. This often manifests as the "last write wins" problem, where only the configuration from the most recently reconciled lower-level object persists, leading to data loss or incorrect system states. This highlights the necessity of deeply understanding the CRD's design and its relational model.

Technical Deep Dive

▶ Watch: Consequence: API server overload from excessive requests (3:50)

The core of effective Kubernetes controller development lies in judicious interaction with the API server and intelligent state management. Nick Young details several technical approaches and frameworks designed to achieve this.

Efficient API Interaction

The first set of problems Charlie Don't faces stems from directly interacting with the Kubernetes API server. To mitigate this:

  1. Use Caching Clients: The paramount solution is to employ a caching client. Instead of directly Getting or Listing resources from the API server repeatedly, a caching client maintains a local, in-memory copy of the cluster's state. This drastically reduces API server load and network latency, as most read operations become local memory lookups.
  2. Limit API Server Updates: Beyond reads, controllers must also be mindful of write operations. The advice is simple: only touch the API server when absolutely necessary.
  3. Check for Differences Before Update: Before sending any update, compare the proposed new object state with the original object (typically retrieved from the cache). If there's no actual difference, do not send the update. This prevents no-op updates that consume API server and etcd resources without changing the cluster state.
  4. Prefer Patch Over Update: When an update is truly needed, use a patch operation instead of a full update. A patch sends only the fields that have changed, making the operation more efficient. This also helps avoid race conditions where multiple controllers might be trying to update the same object, leading to "object has been modified since you got it" errors. This is especially crucial for status updates, which are often frequent and can easily trigger such conflicts.

Frameworks for Controllers

Implementing these optimizations manually, particularly the complexities of caching, concurrency, and eventual consistency, is a monumental task. This is where controller frameworks become indispensable. They are "specifically built to do this stuff for you," abstracting away much of the boilerplate and complex logic. Young highlights three prominent frameworks:

  1. KRT (Kubernetes Resource Transformer):
  • Developed by John Howard as an experimental refactoring for Istio and used in K gateway.
  • It operates on collections of Kubernetes objects (or other sources) via informers.
  • Users define fetch functions to retrieve objects based on criteria, and transformation functions that are invoked when the fetched collections change. These transformations then output new objects or trigger actions.
  • KRT is more generic, supporting non-Kubernetes objects, but its collection-based approach can have a "cognitive overhead" for those accustomed to other patterns.
  1. State DB:
  • Used in the agent component of Cilium (where Nick Young works).
  • An in-memory radix tree database for Go, supporting cross-table write transactions and watch channels that close upon updates to specific parts of the radix tree.
  • It allows treating collections of Kubernetes objects "like a database," with operations triggered by row additions, removals, or updates.
  • While effective for managing simple object collections, complex relationships (like in Gateway API) can require a more relational database-like approach, introducing its own set of "cognitive overhead and problems."
  1. Controller Runtime:
  • Part of the upstream Kubernetes kubebuilder project.
  • Employs a deep reconcile pattern with a key-based lookup built on a caching Kubernetes client that mirrors the client-go API. This makes migration from vanilla client-go code straightforward.
  • Watches automatically maintain local caches.
  • Each controller typically reconciles one main resource. When changes occur, a ReconcileRequest is queued, triggering the Reconcile function for that object's name and namespace.
  • A key feature is the ability to add extra watches for dependent objects. If a dependent object changes, it can trigger a reconciliation of the main object, ensuring all related states are considered.
  • Predicates can be applied to watches to filter which events trigger reconciliation, reducing unnecessary processing.
  • Efficiency: Because all state lookups during reconciliation happen against the local cache, "recalculating the whole thing" is highly efficient, involving no network access.

Reconcile Mistakes (Charlie Don't's, and Nick Young's, Mistakes)

Young highlights two critical reconciliation errors, confessing that he himself made these mistakes in Cilium's Gateway API and Gamma controllers:

  1. Inconsistent Predicate Application: When using controller-runtime, predicates applied to the For call (for the main reconciled object) must also be considered for predicates on Watches (for dependent objects). In Cilium's case, the HTTPRoute watcher didn't check if the associated Gateway rolled up to a GatewayClass managed by Cilium. This led to unnecessary reconciliations of Gateways that the controller didn't care about, essentially doubling the reconciliation load when two reconcilers (Gateway API and Gamma) were active.
  • Defensive Measure: Ensure that the logic determining relevance in For predicates is consistently applied or mirrored in the Watches predicates to avoid processing irrelevant objects.
  1. Reconciling the Wrong Object: The Gamma reconciler initially reconciled HTTPRoute objects, even though HTTPRoutes in Gamma roll up to a Service object, which is the higher-layer construct. Because multiple HTTPRoutes can point to the same Service (a one-to-many relationship), reconciling at the HTTPRoute level meant that every update to any HTTPRoute would regenerate the entire configuration, and only the configuration generated by the last HTTPRoute reconciliation would persist, effectively wiping out the config from previous reconciliations.
  • Defensive Measure: "It's critical to make sure you're reconciling the right object." Understand your CRD's design and hierarchy. For Gateway API, the top tip is to reconcile Gateways, not HTTPRoutes. The primary object should be the one that aggregates or orchestrates the state of its dependents.

Demo / Proof of Concept

▶ Watch: Crucial tip: Check for differences before sending updates (5:00)

The talk did not include a live demonstration or a specific proof-of-concept. Instead, Nick Young focused on providing conceptual explanations, architectural patterns, and real-world code snippets (from Cilium's controller-runtime implementation) to illustrate the principles and anti-patterns discussed. The emphasis was on practical advice derived from extensive experience rather than a step-by-step technical demo.

Defensive Implications

▶ Watch: Advanced tip: Use patch instead of update to avoid races (6:00)

For anyone building or maintaining Kubernetes operators and controllers, Nick Young's insights offer a robust defensive roadmap to ensure stability, efficiency, and correctness.

  1. Embrace Controller Frameworks: The most fundamental defense is to avoid hand-rolling complex state management. controller-runtime is highly recommended, especially for those starting out, due to its upstream Kubernetes status, comprehensive features, and straightforward reconciliation pattern. It handles caching, concurrency, and event processing, significantly reducing boilerplate and error surface.
  1. Strict API Server Optimization:
  • Always use caching clients: Frameworks like controller-runtime provide this by default. Leverage them to minimize direct API server queries.
  • Implement no-op update checks: Before any update or patch operation, perform a deep comparison between the current state (from the cache) and the desired state. Only proceed with the API call if a genuine difference exists. This is particularly crucial for status updates.
  • Prefer Patch over Update: Where possible, use strategic merge patches to send only the changed fields. This is more efficient and reduces the likelihood of resourceVersion conflicts (race conditions).
  • Mitigate Thundering Herd Problems: No-op status updates are vital during controller startup or mass object changes. If a controller updates a single gateway's status every time one of its hundreds of associated HTTP routes changes, a restart could trigger hundreds of unnecessary status updates, leading to a "thundering herd problem" that overloads the API server and etcd.
  1. Rigorous Reconciliation Logic:
  • Deep Understanding of CRD Design: A controller must reconcile the correct primary object within a CRD hierarchy. Misidentifying the reconciliation target (e.g., HTTPRoute instead of Gateway or Service) can lead to incorrect state management, configuration overwrites, and data loss. This requires a thorough understanding of the CRD's design and how objects relate.
  • Consistent Predicate Application: Ensure that any filtering logic (predicates) applied to the main resource's watch is also correctly applied to watches for dependent resources. This prevents the controller from wasting resources processing irrelevant objects.
  • Rebuild State of the World: Within the Reconcile function, it's generally safer and simpler to rebuild the entire desired state from the local cache rather than trying to track incremental changes or "before and after" states. This approach, while seemingly inefficient, is highly performant with caching clients and avoids complex logic for handling object additions, deletions, or tombstoning.
  1. Strategic Operator Testing:
  • Layered Architecture: Structure your operator into distinct layers: an ingestion layer (Kubernetes config to internal model), a data model (generic problem domain representation), and a translation layer (internal model to desired output, e.g., another Kubernetes object or external API call).
  • Unit Testing Layers: Test each layer independently. This allows for comprehensive testing of inputs to the data model and the data model to translated outputs without needing to mock complex Kubernetes interactions or external APIs.
  • External State Management: When interacting with external APIs (e.g., cloud provider APIs), represent that external state with an in-memory model within the controller. This facilitates testing and allows the controller to reconcile against a local representation before pushing changes externally.
  • Define Source of Truth: Clearly establish whether Kubernetes is the source of truth or if external changes should be synchronized back. For most reconciliation-based systems, the controller's desired state should be the ultimate source of truth, pushing changes out and potentially overwriting manual external modifications.

By adhering to these defensive strategies, developers can build Kubernetes controllers that are not only functional but also resilient, efficient, and maintainable, contributing to a more stable cloud-native environment.

Key Takeaways

  • Embrace Controller Frameworks: Always use a robust framework like controller-runtime to abstract away complexities of caching, concurrency, and eventual consistency, rather than hand-rolling these mechanisms.
  • Optimize API Server Interaction: Minimize direct API server calls by leveraging caching clients, checking for actual differences before sending updates, and preferring patch operations over full updates, especially for status fields.
  • Design CRDs and Reconcile Correctly: Deeply understand your CRD's hierarchy and choose the appropriate primary object for reconciliation (e.g., Gateway over HTTPRoute in Gateway API) to ensure correct state aggregation and prevent data loss.
  • Apply Predicates Consistently: Ensure that filtering logic (predicates) on dependent resource watches mirrors the relevance checks performed on the main reconciled resource to avoid unnecessary processing and improve efficiency.
  • Rebuild State in Reconcile: Simplify your reconciliation logic by rebuilding the entire desired state of the world from the local cache within the Reconcile function, rather than attempting to track incremental changes, which is prone to edge cases.
  • Implement Metrics and Layered Testing: Integrate metrics to monitor reconciliation frequency and duration, and adopt a layered architecture (ingestion, data model, translation) for easier and more effective testing of complex operators.

About the Speaker(s)

Nick Young is an Engineer at Isovalent, a Cisco company, and a highly experienced practitioner in the Kubernetes ecosystem. His journey with Kubernetes Custom Resource Definitions (CRDs) dates back to early 2017, when they were still referred to as "third-party resources," showcasing his foundational involvement in this critical area. Young has played a significant role in major Kubernetes-related projects, including his contributions to the redesign of Contour's HTTPProxy CRD and his foundational work on the Gateway API since its inception in 2018, a specification that is entirely driven by CRDs. With a candid acknowledgement of having "screwed it up plenty of times" in his extensive experience building controllers, Nick Young brings a practical, lessons-learned perspective to the challenges of Kubernetes controller development.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This session by Nick Young provides a brutally honest and deeply technical dissection of common pitfalls in Kubernetes controller development, particularly with CRDs. Leveraging extensive personal experience, Young meticulously details how to avoid inefficient API server interactions, incorrect reconciliation logic, and the critical importance of choosing the right controller framework. It's an essential guide for anyone building operators, offering actionable insights to prevent widespread cluster instability and performance bottlenecks.

Heather Calloway (CISO) — MUST SEE

Nick Young's KubeCon talk meticulously diagnoses critical anti-patterns in Kubernetes controller development, offering highly actionable guidance for platform engineers. From a CISO perspective, this session is essential, as the stability, efficiency, and correctness of custom controllers directly underpin an organization's operational resilience, security posture, and regulatory compliance. Flawed controller implementations can lead to systemic instability, data loss, and outages, making robust design a foundational requirement for institutional accountability and risk management. This isn't just technical advice; it's a blueprint for building a resilient cloud-native control plane.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025