Practical Zombie Hunting for Kubernetes Users - Holly Cummins, Red Hat

Holly Cummins, Red Hat

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In this insightful KubeCon EU talk, Holly Cummins of Red Hat tackles a pervasive yet often overlooked problem in modern IT infrastructure: the proliferation of "zombie" and "underutilized" servers. These are resources that consume electricity, contribute to carbon emissions, and incur significant financial costs without delivering any useful work. Cummins, drawing from her experience as a consultant and her work on Quarkus, highlights the scale of this global issue, which extends from individual developers forgetting small cloud instances to major corporations misplacing hundreds of GPUs. The talk not only quantifies the staggering waste but also delves into the underlying human and technical reasons for its existence, offering practical strategies for detection and, more importantly, "destruction" through a concept she champions: Light Switch Ops.

Watch on YouTube

Visual summary for Practical Zombie Hunting for Kubernetes Users - Holly Cummins, Red Hat by Holly Cummins, Red Hat
Visual summary for Practical Zombie Hunting for Kubernetes Users - Holly Cummins, Red Hat by Holly Cummins, Red Hat

Key moments

  1. 0:00 Introduction and common zombie infrastructure examples
  2. 1:30 Speaker's forgotten Kubernetes cluster costing £1000/month
  3. 3:10 Defining and quantifying the problem of zombie servers
  4. 4:10 The often-overlooked problem of underutilized servers
  5. 6:00 Staggering financial and environmental cost of wasted resources
  6. 7:40 Green Software Principles and talk's focus on hardware efficiency

Practical Zombie Hunting for Kubernetes Users

Speakers: Holly Cummins, Red Hat

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=1c2va5nATmQ

Overview

In this insightful KubeCon EU talk, Holly Cummins of Red Hat tackles a pervasive yet often overlooked problem in modern IT infrastructure: the proliferation of "zombie" and "underutilized" servers. These are resources that consume electricity, contribute to carbon emissions, and incur significant financial costs without delivering any useful work. Cummins, drawing from her experience as a consultant and her work on Quarkus, highlights the scale of this global issue, which extends from individual developers forgetting small cloud instances to major corporations misplacing hundreds of GPUs. The talk not only quantifies the staggering waste but also delves into the underlying human and technical reasons for its existence, offering practical strategies for detection and, more importantly, "destruction" through a concept she champions: Light Switch Ops.

The core message of the presentation is that the industry's collective forgetfulness, risk aversion, and competing priorities have led to an unsustainable and economically wasteful status quo. Cummins argues that while the problem is complex, its solutions are often straightforward, ranging from simple shell scripts to adopting sophisticated FinOps and GreenOps practices. By reframing server management to prioritize elasticity and reliable decommissioning, organizations can significantly reduce their environmental footprint and reclaim billions in wasted expenditure, transforming the daunting task of "zombie hunting" into a win-win scenario for both the planet and the balance sheet.

Background

▶ Watch: Introduction and common zombie infrastructure examples (0:00)

The concept of "zombie servers" might sound dramatic, but it accurately describes a widespread and costly phenomenon in IT. These are machines, whether physical, virtual, or cloud instances, that are powered on but perform no useful work, often for extended periods. This issue isn't new, but its scale has grown exponentially with the adoption of cloud computing and the ease of provisioning resources. Cummins illustrates this with relatable anecdotes: the developer too scared to turn off a $2/month AWS instance, the eight-year-old WordPress install emailing from an unknown on-prem server, or even Twitter "losing" 700 GPUs. Her own experience of forgetting a £1,000/month Kubernetes cluster for two months underscores how easily even experienced professionals can fall victim to this problem.

To quantify this anecdotal evidence, Cummins references studies by the Anthesis Institute. Their 2015 survey of 4,000 servers found that a staggering 30% were doing no useful work. A follow-up survey in 2017, expanded to 16,000 servers, yielded similar results, with a quarter of servers identified as "comatose" – meaning they hadn't delivered any information or computing services for six months or more. Beyond completely idle systems, the problem extends to "underutilized servers." The Anthesis Institute found that 29% of servers were active less than 5% of the time, leading to an estimated average server utilization in the industry of a dismal 12% to 18%. This means that the vast majority of computing capacity sits idle, consuming resources without providing value.

The implications of such low utilization are profound. Financially, a 2021 study estimated $26 billion was wasted by always-on cloud instances. Environmentally, the impact is equally severe. Even if these systems run on renewable electricity, they contribute to embodied carbon (the emissions from manufacturing the hardware), consume vast amounts of water for cooling, and eventually become e-waste. Cummins contextualizes this within the Green Software Principles, emphasizing that hardware efficiency – specifically elasticity and utilization – is a critical pillar for sustainable IT operations. The root causes, she explains, boil down to a combination of human factors: forgetfulness (lack of institutional memory), laziness (competing priorities making decommissioning seem unglamorous), and fear (risk aversion leading to overprovisioning and reluctance to turn things off). Technical complexities, such as Kubernetes CRDs (Custom Resource Definitions) not being namespace-scoped, can also drive wasteful architectural patterns like "cluster-per-team" deployments, further exacerbating the problem.

Key Findings

▶ Watch: Defining and quantifying the problem of zombie servers (3:10)

Holly Cummins's talk illuminates several critical findings regarding the zombie server phenomenon and its resolution:

  • Pervasive Waste: Empirical data from the Anthesis Institute reveals that 25-30% of servers are completely idle, and another 29% are active less than 5% of the time. This translates to an average industry-wide server utilization of a mere 12-18%, indicating a massive inefficiency in resource allocation.
  • Staggering Costs: This inefficiency isn't just a minor operational overhead; it represents a significant financial drain. A 2021 study highlighted by Cummins estimated a $26 billion annual waste from always-on cloud instances, money that could be reinvested or saved.
  • Environmental Burden: Beyond financial costs, zombie servers contribute substantially to environmental degradation. They consume electricity, generate embodied carbon through hardware manufacturing, require vast amounts of water for data center cooling, and ultimately add to the growing problem of e-waste.
  • Human Factors are Key Drivers: The primary causes of zombie server proliferation are identified as forgetfulness (lack of institutional memory, neglected decommissioning), laziness (decommissioning tasks are often deprioritized), and fear (risk aversion leading to overprovisioning and reluctance to turn off systems due to potential outages or the IKEA effect – attachment to something you've built).
  • Technical Challenges Exacerbate the Problem: Specific technical issues, such as the behavior of CRDs in Kubernetes leading to siloed cluster deployments, and the inherent bias of autoscaling algorithms towards availability over utilization, contribute to the problem.
  • "Light Switch Ops" as the Solution: The central technical solution proposed is Light Switch Ops, a paradigm shift where turning servers off and on is as fast, reliable, and low-risk as flipping a light switch. This requires systems to be idempotent, resilient, and based on Infrastructure as Code principles (GitOps).
  • Simple Automation Yields Big Results: Even basic automation, such as timed shut-offs using shell scripts, can lead to substantial savings. Examples cited include a UK bank achieving a 50% reduction in cloud costs by implementing self-destructing instances after two weeks (unless renewed), a Chicago company saving 30% on its cloud bill with timed shut-offs, and a Belgian school saving 12,000 euros annually with similar simple scripts.
  • Detection and Destruction are Both Crucial: Effective zombie hunting requires both identifying unused resources and having the organizational and technical capability to safely decommission them. Traditional methods like spreadsheets and email appeals are largely ineffective; modern approaches like FinOps, GreenOps, and even the high-risk "scream test" offer better detection.
  • Process Elasticity is as Important as Technical Elasticity: Counterintuitively, heavy barriers to provisioning new systems can lead to users clinging to existing ones. An "easy come, easy go" governance model, allowing for flexible provisioning and decommissioning, fosters better resource management.

Technical Deep Dive

▶ Watch: The often-overlooked problem of underutilized servers (4:10)

The technical core of the zombie server problem, as articulated by Holly Cummins, lies in the discrepancy between provisioned capacity and actual utilization, driven by a complex interplay of human behavior and system architecture.

At its foundation, the problem differentiates between truly comatose servers (idle for six months or more, representing 25-30% of systems) and underutilized servers (active less than 5% of the time, accounting for another 29%). The industry average utilization of 12-18% is a stark indicator of systemic inefficiency. This low utilization directly translates to substantial financial waste, estimated at $26 billion annually for always-on cloud instances. Environmentally, the impact is multifaceted: it includes embodied carbon from manufacturing, significant water consumption for cooling data centers, and the generation of e-waste at the end of a system's unnecessarily short useful life. These factors underscore the urgency of addressing hardware efficiency, a key tenet of the Green Software Principles.

The genesis of these zombie systems can be traced to several factors. Forgetfulness, or a lack of institutional memory, means systems are provisioned and then simply forgotten, especially when projects end or business processes change. The anecdote of a server accidentally bricked into a wall at a university highlights the extreme end of this spectrum, making it easy to understand how cloud instances, being "out of sight, out of mind," are even more susceptible. Laziness, or more accurately, competing priorities, means that decommissioning systems is often viewed as "boring" and deprioritized compared to new development. Finally, fear is a significant driver. No one wants to be responsible for an outage caused by under-provisioning, leading to conservative over-provisioning. This risk aversion is also why the "scream test" (randomly shutting things off) is effective but terrifying for many organizations.

Kubernetes, while promising efficiency, introduces its own set of challenges. Cummins points out that when she first learned Kubernetes, the ideal was a single cluster with namespaces for isolation. However, the practical reality of Custom Resource Definitions (CRDs) often applying across namespaces creates conflicts and security concerns, pushing organizations towards a less efficient model of "cluster-per-team" deployments. This technical constraint, combined with autoscaling algorithms inherently biased towards availability over strict utilization (to avoid outages), further contributes to over-provisioning and underutilization. The target for optimal utilization, balancing efficiency with headroom, should ideally be 70-80%.

Solving the zombie problem requires a two-pronged approach: detection and destruction.

Traditional detection methods like manual system archaeology using spreadsheets or organization-wide "long emails" are largely ineffective, often hitting dead ends with "unknown assets" and unassigned ownership. Tags offer some improvement by providing metadata, but they frequently become outdated. More modern approaches leverage FinOps and GreenOps. FinOps, focused on financial accountability, brings cost data to engineers, allowing them to optimize away unused resources. Tools like the Backstage Cost Insights plugin provide this crucial visibility. GreenOps extends this to environmental impact, with plugins like the Cloud Carbon Footprint directly showing the carbon cost of workloads. The "scream test" remains a brutal, high-risk, but undeniably effective method for identifying critical dependencies.

However, detection is only half the battle. The "destruction" or decommissioning phase is often hampered by bureaucracy, personal risk aversion, and the psychological IKEA effect, where individuals become attached to systems they've invested time in. To overcome these barriers, Cummins advocates for Light Switch Ops. This paradigm demands that systems possess several key qualities of service:

  1. Fast: Systems must spin down and up quickly.
  2. Idempotent: Repeatedly applying the same operation (e.g., bringing a system up) should yield the same, consistent result.
  3. Resilient: Systems must reliably come back online after being shut down.

Achieving Light Switch Ops necessitates a shift away from "snowflake servers" towards GitOps and Infrastructure as Code. By defining infrastructure and application configurations in Git, tools like kubectl and Ansible can reliably provision and de-provision systems, ensuring that turning a server off doesn't lead to it being "never the same again" because its state was ephemeral or undocumented.

Automation is the linchpin of Light Switch Ops. This doesn't require complex, proprietary solutions. Simple, scheduled automation can yield significant results:

  • A UK bank achieved a 50% reduction in cloud costs by implementing self-service instances that automatically self-destructed after two weeks unless renewed.
  • A Chicago company saw a 30% reduction in its cloud bill through timed shut-offs.
  • A Belgian school saved 12,000 euros annually using basic shell scripts to manage server uptime.

Open-source projects like Daily Clean from AXA France facilitate this by providing a Kubernetes pod (built with Quarkus for resource efficiency) that offers a user-friendly front end for scheduling server shutdowns, eliminating the need to learn cron syntax. Commercial products like "Turn It Off" are also emerging to address this need. Advanced techniques like autotuning, autoscaling, and bin packing further contribute to maximizing utilization.

Finally, Cummins cautions against relying solely on certain technologies as silver bullets. While the cloud offers elasticity, its "out of sight, out of mind" nature makes it easy to forget resources. Virtualization still leaves the underlying operating system running, consuming resources even when applications are idle. Serverless architectures offer elasticity for workloads but often rely on a less elastic control plane, which must be factored into costs. Counterintuitively, heavy preventative barriers to provisioning systems can be detrimental, as users, once they overcome the hurdles, become reluctant to decommission. Instead, an "easy come, easy go" process governance model, promoting process elasticity, is crucial for fostering a culture of efficient resource management.

Demo / Proof of Concept

▶ Watch: Staggering financial and environmental cost of wasted resources (6:00)

The talk by Holly Cummins did not feature a live technical demonstration or a specific proof of concept developed by the speaker. However, she referenced several real-world examples and tools that embody the principles of zombie hunting and Light Switch Ops.

For instance, Cummins highlighted Daily Clean, an open-source project from AXA France. This tool is implemented as an extra pod within a Kubernetes cluster and leverages Quarkus for its resource lightness. Daily Clean provides a user-friendly front end that allows users to schedule the shutdown of Kubernetes resources, simplifying the automation of power management without requiring direct interaction with complex scheduling syntax like cron. This serves as a practical example of how organizations are building tools to achieve the "light switch" functionality.

Additionally, she mentioned commercial products like "Turn It Off," indicating a growing market recognition of the need for specialized solutions in this space. While not a direct demo, these references provide concrete examples of how the theoretical concepts discussed in the talk are being implemented and applied in production environments to combat the zombie server problem.

Defensive Implications

▶ Watch: Green Software Principles and talk's focus on hardware efficiency (7:40)

The insights from "Practical Zombie Hunting for Kubernetes Users" offer clear and actionable defensive strategies for organizations looking to improve efficiency, reduce costs, and enhance sustainability. These implications span from individual user practices to organizational culture and tool development.

For users and practitioners, the primary defensive stance involves a conscious effort to maximize resource utilization and embrace elasticity. This means actively striving for 70-80% utilization for systems, ensuring there's headroom for spikes but minimal waste during idle periods. A critical practice is to limit Kubernetes sprawl – avoiding the creation of unnecessary clusters or namespaces that lead to isolated, underutilized resources. Most importantly, users must practice desmification: knowing precisely what resources they are using and diligently turning them off or decommissioning them when no longer needed. This requires overcoming the "fear" and "laziness" factors by understanding the true cost and environmental impact of forgotten systems.

For tool creators and platform developers, the defensive implications revolve around building solutions that inherently support efficient resource management. Tools should be designed with built-in elasticity, allowing systems to scale up and down seamlessly with demand. Support for multi-tenancy is crucial to enable bin packing, allowing multiple workloads to share resources efficiently and reduce the need for isolated, underutilized infrastructure. Furthermore, tools should facilitate desmification by providing clear visibility into resource usage, cost, and environmental impact, and making the process of decommissioning simple, reliable, and low-risk. This includes integrating features that support Light Switch Ops, ensuring that systems can be reliably and idempotently turned off and on.

Organizationally, adopting FinOps and GreenOps is a powerful defensive strategy. FinOps provides real-time financial transparency, pushing cost accountability down to the engineering teams responsible for resource consumption. By making the financial impact of zombie servers visible, it incentivizes optimization. GreenOps extends this by making the environmental impact (carbon, water, e-waste) equally transparent, fostering a culture of sustainability.

Implementing Light Switch Ops is a fundamental defensive measure. This requires a commitment to Infrastructure as Code and GitOps principles, moving away from "snowflake servers" to ensure that infrastructure can be reliably provisioned and de-provisioned. Automation, even simple shell scripts, for scheduled shutdowns and self-destructing instances, should be a standard practice, not an afterthought. This builds confidence in turning systems off, addressing the "fear" aspect head-on.

Finally, organizations must cultivate a culture that balances risk aversion with resource efficiency. This means establishing an "easy come, easy go" process governance model for provisioning and decommissioning servers. While robust security and availability are paramount, creating heavy bureaucratic barriers to provisioning can be counterproductive, leading users to hoard resources they've fought hard to obtain. Instead, making it easy to spin up and spin down resources, coupled with clear accountability and automated safeguards, encourages responsible resource management and prevents the silent accumulation of zombie infrastructure.

Key Takeaways

  • Zombie and underutilized servers are a massive, pervasive problem, wasting billions of dollars annually ($26B estimated) and contributing significantly to environmental damage (embodied carbon, water consumption, e-waste).
  • Human factors (forgetfulness, laziness, fear of outages, and the IKEA effect), combined with technical challenges like Kubernetes sprawl and autoscaling bias, are the primary drivers of this inefficiency.
  • "Light Switch Ops" is the core solution, advocating for systems that can be turned off and on as reliably, quickly, and confidently as a light switch, requiring qualities like idempotency and resilience.
  • Adopting Infrastructure as Code (GitOps) and automation is crucial for achieving Light Switch Ops, enabling reliable provisioning and decommissioning, even simple shell scripts can yield substantial savings (e.g., 30-50% cloud cost reduction).
  • Leveraging FinOps and GreenOps practices, supported by tools like Backstage plugins for cost and carbon insights, provides the necessary visibility and accountability to drive resource optimization.
  • Fostering "process elasticity" with an "easy come, easy go" governance model for server provisioning and decommissioning helps overcome the reluctance to turn off systems and encourages more efficient resource utilization.

About the Speaker(s)

Holly Cummins is an experienced professional at Red Hat, where her primary role involves helping to build Quarkus, a cloud-native Java framework. Before her current position, she worked as a consultant, gaining firsthand experience with the widespread problem of forgotten and underutilized IT infrastructure, which directly inspired the topic of this talk. Cummins has emerged as a "thought leader" in the concept of Light Switch Ops, advocating for more elastic and sustainable approaches to server management. She actively shares her insights on platforms like Blue Sky and contributes to the broader conversation around green software and efficient cloud operations.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Holly Cummins delivers a highly practical and well-researched talk on the pervasive issue of zombie and underutilized servers. She quantifies the staggering financial and environmental waste, dissects the human and technical drivers behind it, and proposes a clear, actionable solution: "Light Switch Ops." The talk is grounded in empirical data, real-world examples of significant savings, and offers concrete strategies for detection and decommissioning, making it a valuable session for anyone in cloud operations or infrastructure management.

Heather Calloway (CISO) — STRONG ACCEPT

Holly Cummins's talk on "Practical Zombie Hunting" delivers a clear, evidence-based indictment of pervasive IT waste stemming from neglected infrastructure. While not a direct security briefing, it addresses a fundamental governance failure: the lack of clear ownership and reliable processes for managing the lifecycle of computing resources. The staggering financial and environmental costs, driven by institutional forgetfulness, risk aversion, and operational inertia, demand executive attention. Cummins offers a compelling call to action through "Light Switch Ops," advocating for robust automation and cultural shifts like FinOps and GreenOps to ensure resource accountability and efficient…

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025