Simplify Kubernetes Operator Development With a Modular Desi... Mostafa Hadadian & Alexander Lazovik

Mostafa Hadadian, Alexander Lazovik

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Mostafa Hadadian and Alexander Lazovik at KubeCon EU, delves into a novel approach for developing Kubernetes operators, addressing the common pitfalls of complexity and monolithic design. The speakers guide the audience through their journey in building an AI serving platform, highlighting the challenges encountered with traditional operator development and presenting a robust, modular design pattern to overcome them. The core of their solution revolves around decoupling Custom Resource Definition (CRD) translation logic from controller code, leveraging Helm as a configuration and templating engine, and employing microcontrollers orchestrated by a central coordinator.

Watch on YouTube

Visual summary for Simplify Kubernetes Operator Development With a Modular Desi... Mostafa Hadadian & Alexander Lazovik by Mostafa Hadadian, Alexander Lazovik
Visual summary for Simplify Kubernetes Operator Development With a Modular Desi... Mostafa Hadadian & Alexander Lazovik by Mostafa Hadadian, Alexander Lazovik

Key moments

  1. 0:00 Welcome and Kubernetes operators explained
  2. 2:00 Problems with complex, monolithic operator development
  3. 4:00 Divide and conquer with modular CRDs, microcontrollers
  4. 4:40 Using Helm to decouple CRD translation logic
  5. 5:20 Application developers define traits for resource characteristics
  6. 6:00 Abstracting resource type and size with traits
  7. 8:00 Building configurable operator modules using Helm templates

Simplify Kubernetes Operator Development With a Modular Design Pattern

Speakers: Mostafa Hadadian, CEO and Founder of Kardell; Alexander Lazovik

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=m8ZnlZTo1OE

Overview

This talk, presented by Mostafa Hadadian and Alexander Lazovik at KubeCon EU, delves into a novel approach for developing Kubernetes operators, addressing the common pitfalls of complexity and monolithic design. The speakers guide the audience through their journey in building an AI serving platform, highlighting the challenges encountered with traditional operator development and presenting a robust, modular design pattern to overcome them. The core of their solution revolves around decoupling Custom Resource Definition (CRD) translation logic from controller code, leveraging Helm as a configuration and templating engine, and employing microcontrollers orchestrated by a central coordinator.

The significance of this talk lies in its practical solution to a pervasive problem in the cloud-native ecosystem: the inherent complexity that arises as operator logic scales. While Kubernetes operators offer powerful capabilities for extending cluster automation, their development often leads to tightly coupled, hard-to-maintain codebases. Hadadian and Lazovik's modular design pattern offers a blueprint for building operators that are not only easier to develop and configure but also more maintainable, scalable, and adaptable to diverse environments and application requirements. This methodology is particularly relevant for platform engineers and developers grappling with the intricacies of managing complex applications, such as AI/ML pipelines, within Kubernetes.

Background

▶ Watch: Welcome and Kubernetes operators explained (0:00)

Kubernetes operators serve as a powerful extension mechanism, enabling the automation of complex application lifecycle management within the cluster. At their core, operators extend the Kubernetes API by introducing Custom Resource Definitions (CRDs), which define new types of objects, and controllers, which implement the control loop logic to reconcile the desired state (defined by the custom resources) with the actual state of the cluster. This control loop principle makes operators the "autopilot" of a Kubernetes cluster, automating tasks that would otherwise require manual intervention or external scripts.

The journey for many developers begins with frameworks like the Operator SDK, which provides scaffolding and tools to quickly build a basic operator. While initially straightforward, this rapid prototyping often masks underlying architectural challenges. As an operator's responsibilities grow, the controller logic can become exceedingly complex, leading to a monolithic codebase. Furthermore, CRDs, once defined, become central API components, akin to "carving in a stone," making them difficult to modify without impacting consumers. This rigidity, coupled with the escalating complexity of the controller, creates a significant burden for platform engineers, who often become the sole maintainers of these intricate systems. Hadadian poignantly notes the irony: "cloud-native developers build monolithic operators," a contradiction to the microservices philosophy Kubernetes champions. This observation underscores the urgent need for a design pattern that promotes modularity and maintainability in operator development, preventing the very systems designed for automation from becoming unmanageable monoliths.

Key Findings

▶ Watch: Divide and conquer with modular CRDs, microcontrollers (4:00)

The central finding of Hadadian and Lazovik's work is the critical need to address the inherent complexity and monolithic tendencies in traditional Kubernetes operator development. They identify that while operators are essential for automating application management, the common approach of consolidating all reconciliation logic and resource translation within a single controller often leads to unmanageable systems. Their key contribution is a modular design pattern that fundamentally re-architects how operators are built, focusing on "divide and conquer" principles.

The main discoveries and contributions of their approach include:

  1. Modular CRDs: The recognition that complex application specifications can and should be broken down into smaller, more manageable Custom Resource Definitions or sub-specifications. This allows for a more granular approach to defining desired states, making CRDs less monolithic and more adaptable.
  2. Decoupling Translation Logic: A crucial insight is the separation of the logic responsible for translating custom resources into built-in Kubernetes resources (like Deployments, StatefulSets, Services) from the core controller reconciliation logic. By externalizing this translation, the controller's code remains focused on orchestrating state, rather than being cluttered with manifest generation details.
  3. Leveraging Helm for Translation: The innovative use of Helm as the primary engine for this translation. Helm, conventionally used for packaging and deploying applications, is repurposed here to take values derived from custom resources and, using predefined templates, generate the necessary Kubernetes manifests. This exploits Helm's native capabilities for templating and value management, providing a standardized and widely understood mechanism for resource translation.
  4. The "Trade System" for Abstraction: The introduction of a "trade system" (implemented via annotations or specific fields in the CRD) that links abstract characteristics of an application (e.g., type: stateless, size: medium) to concrete Helm values and templates. This creates a powerful layer of abstraction, allowing application developers to define high-level requirements without needing to understand the underlying Kubernetes primitives. Platform engineers, in turn, define what "medium" or "stateless" means in terms of actual replica counts, resource limits, or deployment strategies specific to their cluster environment.
  5. Microcontrollers and Coordinator Pattern: To manage the logic for different parts of a modular CRD, they propose a microcontroller architecture. A main controller acts as a coordinator, receiving the overall custom resource and then delegating specific parts of the reconciliation and value extraction to specialized microcontrollers. Each microcontroller focuses on a distinct aspect of the application's desired state, providing values to the Helm engine. This maintains the performance benefits of running within a single container while offering the development and maintenance flexibility of a microservices-like approach.

These findings collectively present a compelling alternative to monolithic operator development, promising improved maintainability, configurability, and a clearer separation of responsibilities among different engineering roles (application developers, DevOps/SRE, and platform engineers).

Technical Deep Dive

▶ Watch: Using Helm to decouple CRD translation logic (4:40)

The technical core of Hadadian and Lazovik's modular operator design pattern lies in its strategic use of Helm, a novel "trade system," and a microcontroller architecture to achieve significant decoupling and abstraction.

At the heart of the problem is the traditional operator's responsibility to perform two main tasks: reconciling the desired state (from the custom resource) with the actual cluster state, and translating the custom resource's specification into standard Kubernetes manifests. The proposed pattern delegates this translation responsibility entirely, freeing the controller from direct manifest generation.

1. Decoupling CRD Translation with Helm:

Instead of writing Go or Java code within the controller to construct Kubernetes Deployment, StatefulSet, or Service objects, the pattern leverages Helm. Helm's primary function is to combine values with templates to produce Kubernetes manifests. The operator's role shifts from generating manifests to providing values that Helm can then use to generate them.

The structure relies on:

  • Charts: Standard Helm charts containing templates/ and values.yaml files.
  • Values: These are derived from the custom resource's spec.
  • Templates: These define the Kubernetes built-in resources.

2. The "Trade System" for Abstraction and Configuration:

The crucial link between the abstract definitions in a custom resource and the concrete Helm values is established through a trade system. This system allows developers to specify characteristics of their application (e.g., type: stateless, size: medium) which are then mapped to specific Helm values and, consequently, to particular Kubernetes templates or configurations.

  • Example: type Trait:
  • A custom resource might define spec.type: stateless.
  • The values.yaml (or a helper template) for the operator includes logic:
  • The controller extracts spec.type and uses this mapping to select the appropriate Helm template (deployment.yaml in this case).
  • Example: size Trait:
  • A custom resource might define spec.size: medium.
  • The values.yaml or helper template defines:
  • Here, medium translates to replicas: 2. The key innovation is that "medium" can be defined differently by platform engineers across various clusters (e.g., medium in a development cluster might mean 1 replica, while in production it means 5). This provides a powerful abstraction layer, separating the application's intent from the cluster-specific implementation details.

The controller's Go code (or any language) then becomes minimal, primarily responsible for:

  1. Reading the custom resource.
  2. Extracting the defined traits (e.g., type, size).
  3. Looking up the corresponding Helm template and values based on the trait definitions.
  4. Invoking Helm to render the final Kubernetes manifests.
  5. Comparing these rendered manifests with the actual cluster state.
  6. Applying the necessary changes.

3. Microcontrollers and Coordinator Architecture:

To manage complexity within the controller itself, especially for rich CRDs, the speakers propose a microcontroller design:

  • mycontroller.go (Main Control Loop): This is the primary reconciliation loop, adhering to Kubernetes best practices by running all logic within a single container for performance.
  • Coordinator: A component within the main controller that understands the overall custom resource spec. Its job is to break down the spec into smaller, manageable parts and delegate these to specific microcontrollers.
  • Microcontrollers: These are specialized sub-components, each responsible for understanding and processing a particular aspect or field of the custom resource's spec. For example, an HTTPController might handle HTTP-related configurations, while a StorageController handles persistent volume claims. Each microcontroller's task is to extract relevant information from its assigned part of the spec and provide the necessary values to the Helm templating engine.

This internal modularity allows for easier development, testing, and maintenance of individual logical units. If, for instance, the HTTP configuration logic needs to change, only the HTTPController (and its associated Helm templates) needs modification, not the entire monolithic controller.

4. Handling External CRDs and unstructured:

The speakers strongly advocate for this pattern, especially when dealing with external CRDs or when the CRD's API is not directly available in the controller's language (e.g., a Go controller interacting with a Java-defined CRD). In such scenarios, developers often resort to Kubernetes' unstructured library, which allows manipulation of arbitrary Kubernetes objects as generic maps. Working with unstructured objects is notoriously difficult and error-prone. By externalizing the translation to Helm and using a modular design, the need for complex unstructured parsing within the controller's core logic is significantly reduced, justifying the initial architectural effort.

In essence, this technical design transforms the operator from a monolithic code generator into an intelligent orchestrator that leverages existing, well-understood tools like Helm for its core translation function, greatly enhancing flexibility, maintainability, and clarity.

Demo / Proof of Concept

▶ Watch: Abstracting resource type and size with traits (6:00)

While the talk did not feature a live, interactive code demonstration, Mostafa Hadadian presented a compelling architectural proof of concept by detailing its application within their larger project, ACDA (Evolutionary Changes in Data Analysis). ACDA is an AI serving platform designed to help data scientists easily serve and manage their models within Kubernetes, and its entire operational core is built upon this modular operator design pattern.

The application itself is described as a system for creating and managing AI pipelines. A typical pipeline, as illustrated, might involve:

  • A Data Source module, which feeds data.
  • An LLM (Large Language Model) module, processing the data.
  • A Validator module, which evaluates the LLM's output, potentially assigning scores or rejecting data streams.
  • A Planner module, which can receive feedback from the Validator and adjust prompts or parameters for subsequent LLM interactions, creating a loopback mechanism for continuous improvement.

These interconnected components form a "pipeline." The modular operator design pattern is applied to manage these elements:

  1. Pipeline Controller: This overarching controller is responsible for orchestrating the entire AI pipeline. It understands the desired state of the full pipeline (e.g., the sequence of modules, their connections) as defined in a custom resource. It uses the coordinator pattern to delegate to other microcontrollers.
  2. Module Controller: Each "box" in the pipeline diagram (Data Source, LLM, Validator, Planner) represents a module. The Module Controller is designed to be highly versatile. It takes the definition of a specific module from the custom resource and, using the Helm-based trade system, translates it into the necessary Kubernetes resources for deployment. This means a single Module Controller can deploy various types of modules (e.g., a stateless LLM, a stateful data source) by simply configuring their characteristics in the custom resource, without requiring new controller code for each module type.
  3. Link Controller: The connections between modules (e.g., data flowing from Data Source to LLM) are managed by a dedicated Link Controller. This controller understands how to establish these connections within Kubernetes, perhaps by configuring network policies, service meshes, or message queues, again using the modular design to adapt to different linking requirements.

The power of this design is highlighted by the ability to configure modules on the fly—for example, changing a validation threshold and observing immediate results without redeploying or recompiling operator code. The system prioritizes "configurable controllers" over a multitude of rigid, specialized controllers. This real-world application demonstrates that the modular design pattern yields operators that are:

  • Easy to develop: New modules or link types can be added by extending Helm templates and potentially adding a new microcontroller, rather than rewriting core logic.
  • Configurable: The trade system allows for vast customization of deployments based on high-level characteristics.
  • Maintainable: Clear separation of concerns simplifies debugging and updates.
  • Clear responsibility: Different teams can contribute to different parts (CRD definitions, Helm templates, microcontroller logic) without stepping on each other's toes.

The ACDA platform serves as a robust testament to the efficacy and practical benefits of the modular design pattern for complex, evolving cloud-native applications.

Defensive Implications

▶ Watch: Building configurable operator modules using Helm templates (8:00)

While the talk focuses primarily on architectural design and development efficiency for Kubernetes operators, the proposed modular design pattern has significant, albeit indirect, defensive implications for the security posture of cloud-native applications. A well-designed, maintainable, and understandable system is inherently more secure than a complex, monolithic one.

  1. Reduced Attack Surface in Controller Logic: By offloading the complex and often error-prone task of translating CRD specs into Kubernetes manifests to Helm, the amount of custom Go (or Java) code within the operator's controller is significantly reduced. Less custom code typically means a smaller attack surface and fewer opportunities for introducing bugs, logic flaws, or security vulnerabilities (e.g., improper input validation leading to injection attacks if manifest generation was handled manually). Helm's templating engine is a mature, well-audited component, providing a more secure foundation for manifest generation compared to custom-written logic.
  1. Clearer Separation of Responsibilities for Security: The design pattern explicitly outlines distinct roles:
  • Application Developers: Define desired application characteristics via custom resources.
  • DevOps/SRE/Security Engineers: Configure the Helm templates and values that translate these characteristics into concrete Kubernetes resources. This is where security policies can be enforced. For instance, defining that a "medium" sized workload always gets specific resource limits, network policies, or security contexts, or that "stateless" workloads never get persistent storage without explicit overrides.
  • Platform Engineers: Develop and maintain the microcontrollers and coordinator. Their focus is on the orchestration logic, not the specific security configurations of individual workloads.

This clear division allows security teams to focus their efforts on auditing and hardening the Helm charts, ensuring that baseline security requirements are met for all deployments managed by the operator.

  1. Enforcement of Secure Defaults and Policies: The "trade system" is a powerful mechanism for implementing security by default. Platform engineers can define what "medium" or "stateless" means in their cluster, including security-relevant parameters. For example, a "medium" workload might automatically be assigned a specific PodSecurityContext or NetworkPolicy by the Helm template, regardless of what the application developer explicitly requests. This prevents developers from inadvertently deploying insecure configurations, as the secure defaults are baked into the translation layer. It shifts the burden of security from individual application teams to the platform, where experts can define and enforce best practices.
  1. Improved Auditability and Troubleshooting: Monolithic operators with intertwined logic are notoriously difficult to audit for security vulnerabilities or troubleshoot security incidents. The modular design, with its distinct microcontrollers and externalized Helm templating, makes the operator's behavior more transparent. It's easier to identify which part of the system (a specific microcontroller, a Helm template, or a CRD definition) is responsible for a given resource configuration, facilitating security reviews and incident response.
  1. Enhanced Maintainability for Security Patches: If a vulnerability is discovered in a specific component (e.g., how persistent volumes are provisioned), only the relevant microcontroller and its associated Helm templates might need updating, rather than a sprawling, complex codebase. This significantly speeds up the patching process and reduces the risk of introducing new vulnerabilities during updates.

In conclusion, while not directly addressing specific CVEs or attack vectors, the modular operator design pattern fosters an environment of better software engineering practices—clarity, maintainability, and separation of concerns—which are foundational to building more secure and resilient cloud-native systems.

Key Takeaways

  • Monolithic Kubernetes operators, while initially simple to build, quickly become complex, difficult to maintain, and rigid, especially as application logic scales.
  • Decoupling the translation of Custom Resource Definitions (CRDs) into built-in Kubernetes resources from the core controller logic is crucial for operator scalability and maintainability.
  • Helm can be effectively repurposed as a powerful, standardized engine for this CRD-to-manifest translation, leveraging its templating capabilities to simplify operator code.
  • A "trade system" provides a vital abstraction layer, allowing application developers to define high-level characteristics (e.g., stateless, medium) that platform engineers translate into specific, cluster-dependent Kubernetes configurations via Helm.
  • Implementing a microcontroller architecture with a central coordinator allows for internal modularity within the operator, improving code organization, testability, and maintainability for complex CRDs.
  • This modular design pattern leads to operators that are easier to develop, highly configurable, more maintainable, and promote a clear separation of responsibilities among different engineering roles, ultimately contributing to a more robust and secure cloud-native environment.

About the Speaker(s)

Mostafa Hadadian is the CEO and Founder of Kardell, a company focused on continuous AI delivery (Kardell stands for "Continuous AI Delivery"). He is also currently finishing his PhD at the University of Groningen. Mostafa brings a comprehensive perspective to cloud-native development, having worked across various roles including data engineer, data scientist, and platform engineer. This diverse background has allowed him to witness the full spectrum of challenges faced by different stakeholders in building and operating complex platforms. His work at Kardell and his academic pursuits reflect a deep commitment to creating solutions that enable all contributors—from data scientists to platform engineers—to effectively engage with and contribute to AI serving platforms.

Alexander Lazovik was a co-speaker for this presentation. The transcript, however, did not provide specific biographical details for Alexander.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk presents a highly practical and technically sound modular design pattern for Kubernetes operators, directly addressing the pervasive issue of monolithic and unmaintainable operator codebases. By cleverly repurposing Helm for CRD-to-manifest translation, introducing a powerful 'trade system' for abstraction, and advocating for a microcontroller architecture, the speakers offer a robust blueprint for building scalable and maintainable operators. This isn't just theory; it's a well-articulated solution to a real-world problem that will genuinely improve how platform engineers approach cloud-native automation.

Heather Calloway (CISO) — STRONG ACCEPT

This talk presents a robust modular design pattern for Kubernetes operators, directly addressing the critical issue of complexity in cloud-native automation. By decoupling CRD translation logic using Helm and employing a microcontroller architecture, it offers a blueprint for building more maintainable, scalable, and inherently more secure platforms. This approach significantly improves risk management, auditability, and the ability to enforce security policies through configuration, making it highly relevant for platform engineers and security architects aiming for institutional resilience.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025