Day-2’000 - Migration From Kubeadm+Ansible To ClusterAPI+Talos: A Swiss Bank’s... Clément Nussbaumer

Migration From Kubeadm+Ansible To ClusterAPI+Talos: A Swiss Bank’s... Clément Nussbaumer

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In this insightful KubeCon EU presentation, Clément Nussbaumer, a software engineer at the Swiss bank Post Finance, detailed a significant operational challenge: the in-place migration of long-lived, shared Kubernetes clusters from a traditional kubeadm and Ansible-based management approach to a more modern, declarative stack utilizing Cluster API and Talos Linux. The talk, aptly titled "Day-2’000," refers to the impressive longevity of Post Finance's Kubernetes clusters, some of which are approaching six years in continuous operation. This duration highlights the critical need for robust upgrade and migration strategies that avoid downtime and disruption for hundreds of application teams.

Watch on YouTube

Visual summary for Day-2’000 - Migration From Kubeadm+Ansible To ClusterAPI+Talos: A Swiss Bank’s... Clément Nussbaumer by Migration From Kubeadm+Ansible To ClusterAPI+Talos: A Swiss Bank’s... Clément Nussbaumer
Visual summary for Day-2’000 - Migration From Kubeadm+Ansible To ClusterAPI+Talos: A Swiss Bank’s... Clément Nussbaumer by Migration From Kubeadm+Ansible To ClusterAPI+Talos: A Swiss Bank’s... Clément Nussbaumer

Key moments

  1. 0:00 Introduction: Post Finance & motivation for long-lived clusters
  2. 2:06 Current cluster provisioning with Kubeadm, Ansible, Puppet
  3. 3:19 Understanding Cluster API: declarative Kubernetes APIs for clusters
  4. 4:40 Talos Linux: immutable, minimal, declarative OS explained
  5. 6:00 Overview of migration steps to Cluster API + Talos
  6. 7:05 Deep dive into configuration matching for migration

Day-2’000 - Migration From Kubeadm+Ansible To ClusterAPI+Talos: A Swiss Bank’s Journey to Modern Kubernetes Operations

Speakers: Clément Nussbaumer, Software Engineer, Post Finance

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=uQ_WN1kuDo0

Overview

In this insightful KubeCon EU presentation, Clément Nussbaumer, a software engineer at the Swiss bank Post Finance, detailed a significant operational challenge: the in-place migration of long-lived, shared Kubernetes clusters from a traditional kubeadm and Ansible-based management approach to a more modern, declarative stack utilizing Cluster API and Talos Linux. The talk, aptly titled "Day-2’000," refers to the impressive longevity of Post Finance's Kubernetes clusters, some of which are approaching six years in continuous operation. This duration highlights the critical need for robust upgrade and migration strategies that avoid downtime and disruption for hundreds of application teams.

Nussbaumer's presentation provided a candid look into the complexities of evolving a mature, on-premises Kubernetes environment without resorting to the common practice of cluster recreation. For an organization like Post Finance, which hosts over 500 application teams on large, shared clusters, moving applications to entirely new clusters with every major upgrade is impractical and costly. The focus, therefore, shifts to meticulously planned in-place upgrades and migrations, a testament to the bank's commitment to stability and efficiency. The talk not only outlined the strategic motivations but also delved into the deep technical intricacies and common pitfalls encountered during such a high-stakes transition.

Background

▶ Watch: Introduction: Post Finance & motivation for long-lived clusters (0:00)

Post Finance operates a substantial Kubernetes footprint, managing 35 vanilla Kubernetes clusters on its on-premises data centers, leveraging vSphere for virtualization. These clusters are notable for their longevity, with the oldest nearing six years of continuous operation. A unique aspect of their operational philosophy is the pervasive use of Chaos Monkey across all clusters, including production environments handling card payment services. This practice ensures applications are resilient to pod restarts and node failures, significantly aiding in smooth upgrades.

The existing cluster provisioning and management workflow at Post Finance involved several distinct stages, many of which presented significant operational overhead. Initially, Debian VMs are provisioned using Terraform on their vSphere infrastructure. Following VM creation, a series of Puppet configurations are applied. Nodes then require manual registration in an inventory YAML file, which serves as input for Ansible playbooks. These playbooks are responsible for rendering kubeadm configuration files, kube-apiserver manifests, and executing kubeadm commands to initialize the control plane. Finally, Argo CD is integrated for deploying both application workloads and infrastructure components, such as NGINX Ingress and Cilium. While the Argo CD integration has proven effective, the preceding four steps—Terraform, Puppet, manual inventory, and Ansible—were identified as pain points due driving the desire for a more streamlined, declarative approach.

This desire led Post Finance to explore Cluster API and Talos Linux. Cluster API extends Kubernetes' declarative APIs to manage the lifecycle of Kubernetes clusters themselves. It operates with a management cluster that hosts CRDs (Custom Resource Definitions) and controllers, which in turn provision and manage workload clusters. This paradigm allows for treating infrastructure as code, using kubectl or talosctl to interact with cluster resources. Talos Linux, on the other hand, is a minimal, immutable, and ephemeral operating system specifically designed for Kubernetes nodes. It distinguishes itself by having no SSH access, instead offering a declarative configuration interface through an API server called machineD. All configurations, from network cards to storage, are managed via mTLS authentication to this API, primarily using the talosctl command-line tool. This shift from imperative, tool-chain-heavy provisioning to a fully declarative, API-driven model promised significant improvements in automation, consistency, and security for Post Finance's Kubernetes operations.

Key Findings

▶ Watch: Understanding Cluster API: declarative Kubernetes APIs for clusters (3:19)

The central finding of Post Finance's migration journey is the successful development and implementation of a robust methodology for in-place, zero-downtime migration of mature, production Kubernetes clusters from a kubeadm and Ansible-based setup to one managed by Cluster API and Talos Linux. This approach directly addresses the critical need to avoid application downtime and extensive re-platforming for hundreds of application teams residing on large, shared clusters.

Key discoveries and contributions from their experience include:

  1. Feasibility of Incremental Migration: Demonstrating that it is technically viable to incrementally add new nodes provisioned by Cluster API and running Talos Linux into an existing kubeadm-managed cluster. This hybrid state allows for thorough testing and monitoring, mitigating risks associated with a "big bang" migration.
  2. Critical Configuration Matching: Identifying specific, non-obvious configuration parameters that must be precisely matched between the old kubeadm and new Talos/Cluster API nodes to ensure seamless integration. This includes the service account issuer and the etcd encryption key and its associated name, which are fundamental for inter-node communication and secret management.
  3. PKI Import Mechanism: Leveraging Talos Linux's capabilities to import existing Kubernetes PKI (Public Key Infrastructure) certificates. The talosctl gen secrets from-kubernetes-pki command is a crucial tool, enabling the new Talos nodes to trust and integrate with the existing cluster's certificate authority without requiring a full re-initialization of the PKI.
  4. Bootstrapping the Management Cluster (The "Chicken and Egg" Problem): Devising a pragmatic solution for disaster recovery and initial bootstrapping of the Cluster API management cluster itself. This involves an ephemeral Go utility that can spin up a local Kind cluster, install Cluster API components, and then either provision a new management cluster or restore an existing one from manifests stored in an S3 endpoint. This ensures the declarative management layer itself is resilient.
  5. Specific Error Interpretation: Documenting the often opaque error messages encountered during the migration, such as "invalid bearer token" for service account issuer mismatches or "output array was not large enough for encryption" for incorrect etcd encryption keys. This practical guidance is invaluable for others undertaking similar migrations.

These findings collectively provide a blueprint for other organizations facing similar challenges with legacy Kubernetes infrastructure, emphasizing meticulous planning, deep technical understanding of Kubernetes internals, and the strategic adoption of modern declarative tools.

Technical Deep Dive

▶ Watch: Talos Linux: immutable, minimal, declarative OS explained (4:40)

The migration process from kubeadm and Ansible to Cluster API and Talos Linux is a multi-step, technically intensive endeavor. Clément Nussbaumer outlined four primary phases, beginning with crucial preparatory steps.

1. Configuration Matching

The initial and perhaps most critical step involves meticulously matching configuration parameters between the existing kubeadm clusters and the prospective Talos Linux nodes. Two parameters stand out:

  • Service Account Issuer: By default, kubeadm clusters often use kubernetes.default.svc as the issuer for service account tokens. Talos, however, defaults to using the cluster's API server endpoints. This mismatch can lead to new service account tokens not being accepted by the old API servers, manifesting as "invalid bearer token" errors in authentication logs. The solution is to preemptively change the service account issuer on the existing kubeadm cluster to match the endpoint-based issuer that Talos will use. This change can be performed without downtime.
  • etcd Encryption Key: Ensuring the etcd encryption key and its corresponding key name are identical across both environments is paramount for secret decryption and cluster integrity. Talos templates the encryption configuration, often using schemes like aescbc or secretbox. The name assigned to the key within the Talos configuration (e.g., on line 11 or 17 of the Talos template) must precisely match the key name used in the old kubeadm cluster. This often necessitates re-encrypting all secrets on the existing cluster, a procedure well-documented in Kubernetes official documentation. Failure to match this key results in a cryptic "output array was not large enough for encryption" error.

2. Importing Existing PKI

To enable new Talos nodes to seamlessly join an existing kubeadm cluster, they must trust the cluster's existing Public Key Infrastructure (PKI). Talos Linux provides a highly useful command for this: talosctl gen secrets from-kubernetes-pki. This command takes the path to the existing Kubernetes PKI folder (containing CA certificates, etcd certificates, service account issuer keys, etc.) and generates a secrets bundle file.

This secrets bundle is a comprehensive compilation of critical cryptographic assets, including:

  • The bootstrap token for new node joining.
  • The etcd encryption secrets.
  • Core Kubernetes certificates (e.g., kubernetes-ca.crt, etcd-ca.crt).
  • The service account signing key.
  • Crucially, the OS key (operating system key) which is a Talos-specific certificate used for mTLS authentication against the machineD API server on Talos nodes. This key, particularly if it has the OS admin group, grants administrative access to the Talos operating system via talosctl.

Additionally, a new kubeadm bootstrap token must be created to allow new nodes to securely join the cluster.

3. Creating Cluster API CRDs

On the Cluster API management cluster, specific Custom Resource Definitions (CRDs) are created to define the desired state of the workload cluster being migrated:

  • Cluster CRD: Defines the overall Kubernetes cluster, including its name, API endpoint, and references to its control plane and infrastructure provider.
  • ControlPlane CRD: Specifies the configuration for the control plane nodes, including the Talos version (e.g., Talos 0.9.4), the underlying infrastructure machine template (e.g., VsphereMachineTemplate), and the number of replicas. Initially, the replica count is set to zero to prevent immediate provisioning.
  • MachineDeployment CRD: Defines the configuration for worker nodes (data plane nodes), similarly specifying the Talos version, infrastructure template, and replica count.

These CRDs, once applied, are monitored by Cluster API controllers. However, a critical step is required before increasing replicas: the TalosControlPlane CRD must be patched with status.bootstrapped=true. Without this flag, the Talos bootstrap provider controller would assume it's provisioning a new cluster and send a bootstrap command that would attempt to form a new etcd cluster, rather than joining the existing one. This flag signals that the node should join an already existing cluster.

4. Adding Cluster API Nodes and Removing Kubeadm Nodes

With the preparatory steps complete, the migration proceeds incrementally:

  1. Incremental Node Addition: New Talos Linux nodes are provisioned by increasing the replica count on the ControlPlane and MachineDeployment CRDs. These new nodes join the existing kubeadm cluster. This creates a hybrid cluster environment where old and new nodes coexist.
  2. Monitoring and Validation: During this phase, extensive monitoring is crucial. Nussbaumer highlighted that "it's almost certain that there will be some edge cases," such as fine-tuned sysctl parameters breaking Java deployments. Weeks of close monitoring are anticipated to identify and resolve such issues.
  3. Removal of Kubeadm Nodes: Once the new Talos nodes are stable and proven compatible, the old kubeadm nodes can be gradually drained and removed from the cluster. This process continues until the entire cluster is managed by Cluster API and runs Talos Linux.

Bootstrapping Issue / Disaster Recovery

A significant challenge in a Cluster API setup is the "chicken and egg" problem: how do you manage the Cluster API management cluster itself, especially in a disaster recovery scenario where all clusters are lost? Post Finance's solution involves an ephemeral Go utility. This utility:

  1. Starts a local Kind cluster.
  2. Installs all necessary Cluster API CRDs and controllers onto this Kind cluster.
  3. Checks an S3 endpoint for existing management cluster manifests.
  4. If manifests exist, it re-imports them, allowing the Kind cluster to manage and restore the original management cluster.
  5. If no manifests are found, it provisions a new Cluster API management cluster from scratch.

This utility is run within a GitLab pipeline but can also be executed locally, providing a robust, automated disaster recovery mechanism for the management plane itself, ensuring the entire Kubernetes ecosystem can be rebuilt from scratch if necessary.

Demo / Proof of Concept

▶ Watch: Overview of migration steps to Cluster API + Talos (6:00)

Clément Nussbaumer provided a live demonstration of the critical steps involved in integrating a new Talos-based control plane node into an existing kubeadm cluster, focusing on the lab-F cluster environment.

The demo began by showcasing the Kubernetes manifests used to define the desired state of the lab-F cluster within the Cluster API management cluster. These included:

  • The Cluster manifest, defining the cluster name and API endpoint.
  • The ControlPlane manifest, which referenced a VsphereMachineTemplate and specified Talos 0.9.4 as the operating system image. Crucially, the initial replicas count for the control plane was set to 0.
  • A Customization file that generated secrets and all necessary CRDs for the workload cluster nodes.

The first practical step demonstrated was the import of the existing PKI. Nussbaumer executed a script via SSH on an existing kubeadm control plane node. This script performed several actions:

  1. Generated a kubeadm join token.
  2. Executed talosctl gen secrets from-kubernetes-pki, pointing to the existing Kubernetes PKI folder, to create the talos secrets bundle.
  3. Read the etcd encryption key and printed it to a secrets bundle file.

This process populated a secrets bundle file with the correct Kubernetes certificates and the Talos OS key, essential for authentication and trust.

Next, the generated TLS files were created, and the initial Cluster API Customization was applied to the Cluster API management cluster. This resulted in the creation of the Cluster and TalosControlPlane CRDs in a new namespace (e.g., capy-one-community-pfnet-lab-f), but no actual machines or pods were provisioned yet due to the 0 replica count.

The pivotal moment of the demo involved patching the TalosControlPlane CRD. Nussbaumer emphasized the importance of setting status.bootstrapped=true. This flag signals to the Talos bootstrap provider that the new node should join an existing cluster, preventing it from attempting to form a new, isolated Kubernetes cluster.

After patching, the ControlPlane replica count was increased to 1. This immediately triggered the Cluster API controllers to provision a new machine. Nussbaumer then switched to a talosctl logs workspace, monitoring the boot process of the new Talos node. Concurrently, he observed the etcdctl member list on an existing kubeadm control plane node. After approximately one minute, the new Talos node successfully joined the existing etcd cluster, appearing in the etcdctl member list. This confirmed the critical step of integrating the new control plane into the existing data store.

Finally, the kubectl get nodes output was monitored. While it took a little longer for the new node to transition to a "Ready" state (as Cilium and other components started), the appearance of the new node demonstrated the successful integration of a Cluster API-managed Talos node into the legacy kubeadm cluster. Nussbaumer briefly mentioned that further steps for managing worker nodes (MachineDeployments) and Argo CD integration were beyond the scope of the live demo due to time constraints, but templates were available in the slide deck.

Defensive Implications

▶ Watch: Deep dive into configuration matching for migration (7:05)

The migration strategy presented by Post Finance offers several significant defensive implications for organizations managing Kubernetes infrastructure:

  1. Reduced Attack Surface with Immutable OS: The shift to Talos Linux fundamentally enhances security by adopting an immutable operating system. Talos has no SSH, no package manager, and a minimal set of binaries, drastically reducing the potential attack surface. Configuration changes are declarative and applied via an API, eliminating the risks associated with manual configuration drift, unpatched packages, or forgotten services on individual nodes. This makes nodes inherently more secure and easier to audit for compliance.
  2. Declarative Security Posture: Leveraging Cluster API means the entire cluster configuration, including security-relevant settings, is defined declaratively as code. This enables version control, peer review, and automated enforcement of security policies, ensuring consistency across all clusters. Any deviation from the desired state can be automatically detected and remediated by Cluster API controllers.
  3. Enhanced Disaster Recovery Capabilities: The solution to the "chicken and egg" problem—the ephemeral Go utility for managing the Cluster API management cluster—is a critical defensive measure. It provides a robust, automated mechanism for disaster recovery, allowing the entire Kubernetes infrastructure to be rebuilt from scratch. This minimizes Recovery Time Objectives (RTO) in the event of catastrophic data loss or compromise, ensuring business continuity.
  4. Secure PKI Management: The emphasis on correctly importing and managing the existing Kubernetes PKI, and specifically the etcd encryption key, highlights the importance of cryptographic hygiene. Ensuring that secrets are properly encrypted at rest and that certificate authorities are correctly managed is foundational to the security of any Kubernetes cluster. The talosctl gen secrets from-kubernetes-pki command aids in securely transitioning PKI without exposing sensitive keys.
  5. Controlled and Auditable Node Lifecycle: Cluster API provides a controlled and auditable process for provisioning and de-provisioning nodes. Each machine lifecycle event is managed declaratively, offering clear visibility into infrastructure changes, which is vital for security auditing and compliance.
  6. Mitigation of Configuration Drift: By moving away from imperative management tools like Puppet and Ansible for initial node setup, the risk of configuration drift—where nodes subtly diverge from their intended state over time, potentially introducing vulnerabilities—is significantly reduced. Talos Linux and Cluster API enforce a consistent, desired state.

In summary, this migration represents a strategic move towards a more secure, resilient, and auditable Kubernetes environment, leveraging modern cloud-native principles to harden the underlying infrastructure and streamline operational security.

Key Takeaways

  • In-place migration of long-lived Kubernetes clusters is feasible and desirable for large, shared environments, avoiding disruptive application re-platforming and enabling continuous operation.
  • Meticulous configuration matching is paramount for a successful migration, especially for the service account issuer and etcd encryption key to ensure new nodes can seamlessly join existing clusters.
  • Cluster API and Talos Linux provide a powerful, declarative stack for immutable and automated Kubernetes cluster management, significantly reducing operational overhead and enhancing consistency.
  • A robust disaster recovery strategy for the Cluster API management cluster is essential, addressing the "chicken and egg" problem with tools like an ephemeral Kind cluster and S3-backed manifest storage.
  • Incremental node replacement allows for careful monitoring and validation, minimizing risks during the transition and accommodating the resolution of unforeseen edge cases without downtime.
  • Understanding specific error messages related to PKI, etcd, and service account configuration is critical for efficient troubleshooting during complex migrations.

About the Speaker(s)

Clément Nussbaumer is a Swiss Software Engineer working at Post Finance. He shared his unique perspective from living on a farm with his wife, a farmer, highlighting his daily routine of seeing cows and then engaging with Kubernetes. This background underscores a blend of grounded practicality and cutting-edge technical expertise. At Post Finance, he is deeply involved in managing and evolving their extensive Kubernetes infrastructure.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This presentation by Clément Nussbaumer from Post Finance is a masterclass in tackling a truly difficult operational challenge: the in-place, zero-downtime migration of long-lived, production Kubernetes clusters. The speaker provides an exceptionally deep dive into the technical intricacies of moving from a kubeadm and Ansible-based setup to a modern Cluster API and Talos Linux stack. The candid sharing of specific configuration pitfalls (like service account issuer and etcd encryption keys), the ingenious PKI import mechanism, and the pragmatic disaster recovery solution for the management cluster itself offer invaluable, actionable intelligence for any organization serious about…

Heather Calloway (CISO) — MUST SEE

This KubeCon talk from Post Finance presents a highly relevant and actionable blueprint for migrating critical, long-lived Kubernetes clusters without downtime. It tackles the significant institutional challenge of evolving core infrastructure in a regulated environment, demonstrating how to leverage modern declarative tools like Cluster API and Talos Linux to enhance security, resilience, and operational efficiency. The detailed technical guidance and focus on zero-downtime for hundreds of application teams make this a must-see for any CISO or platform leader grappling with similar legacy modernization efforts.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025