The Next Generation of DaemonSet Autoscaling - Adam Bernot & Bryan Boreham
Adam Bernot, Bryan Boreham
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk, presented by Adam Bernot of Google Cloud and Bryan Boreham of Grafana Labs, addresses a long-standing challenge in Kubernetes: efficiently managing resource requests for DaemonSets across heterogeneous clusters. DaemonSets, designed to run one pod on every node, often struggle with the "one size fits all" approach to resource allocation. This leads to either wasteful overprovisioning on smaller nodes or critical resource starvation and instability on larger, busier nodes.

Key moments
- 0:00 Introduction to the talk and speakers
- 2:00 Understanding DaemonSets and their common use cases
- 3:40 Unpacking the core DaemonSet resource request challenge
- 4:00 Serious consequences of misconfigured DaemonSet resource requests
- 6:20 Real-world data shows DaemonSet CPU usage variability
- 7:00 Exploring common but flawed DaemonSet resource allocation strategies
The Next Generation of DaemonSet Autoscaling
Speakers: Adam Bernot, Staff Software Engineer, Google Cloud; Bryan Boreham, Distinguished Engineer, Grafana Labs
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=yQQU8vDhj0o
Overview
This talk, presented by Adam Bernot of Google Cloud and Bryan Boreham of Grafana Labs, addresses a long-standing challenge in Kubernetes: efficiently managing resource requests for DaemonSets across heterogeneous clusters. DaemonSets, designed to run one pod on every node, often struggle with the "one size fits all" approach to resource allocation. This leads to either wasteful overprovisioning on smaller nodes or critical resource starvation and instability on larger, busier nodes.
Bernot and Boreham introduce a groundbreaking enhancement to the Vertical Pod Autoscaler (VPA) that enables scoped recommendations. This innovation allows VPA to make tailored resource recommendations for DaemonSet pods based on the specific characteristics and workload demands of individual nodes. By dynamically adjusting CPU and memory requests on a per-node basis, this feature promises to significantly reduce operational overhead, optimize resource utilization, and enhance the stability of critical infrastructure components that rely on DaemonSets.
The talk highlights the inefficiencies and risks inherent in current DaemonSet management strategies and presents a compelling solution through a live demonstration of a modified VPA prototype. This work is currently an active Kubernetes Enhancement Proposal (KEP), aiming to bring more intelligent, adaptive autoscaling capabilities to one of Kubernetes' most fundamental workload types. The implications are far-reaching, offering substantial cost savings and improved reliability for organizations running Kubernetes at scale.
Background
▶ Watch: Introduction to the talk and speakers (0:00)
Kubernetes DaemonSets are a fundamental workload type designed to ensure that a copy of a specific pod runs on every node in a cluster. They are indispensable for infrastructure services that require presence on all nodes, such as network controllers, logging agents, metrics collectors (like Prometheus node exporters), and security agents. The core utility of DaemonSets lies in their guaranteed pervasive deployment, making them critical for observability, networking, and security layers.
A crucial aspect of managing any workload in Kubernetes involves defining resource requests for CPU and memory. These requests inform the Kubernetes scheduler about the minimum resources a pod needs to function effectively. The scheduler uses these requests to find a suitable node with sufficient available capacity, ensuring that the workload has the resources it expects. While optional, specifying requests is highly recommended, as their absence can lead to unpredictable scheduling and resource contention.
The fundamental problem arises when a DaemonSet, by its very nature, is deployed across a cluster comprising nodes with diverse characteristics and workloads. Imagine a cluster with a mix of small, lightly loaded nodes and large, heavily utilized nodes. A logging or metrics collection DaemonSet, for instance, will likely perform more work on a node with many active user workloads than on an idle node. However, Kubernetes currently allows only one set of resource requests to be specified for all pods under a single DaemonSet. This "one size fits all" constraint forces operators into an unenviable dilemma:
- If requests are too low: Pods might be starved for resources. For CPU, this leads to throttling, slowing down the workload. For memory, it can cause Out-Of-Memory (OOM) errors, leading to application crashes and restarts. If a request is too low but a limit is high, the workload can become a noisy neighbor, consuming excess resources and impacting other workloads on the node. Ultimately, this can lead to container-level or node-level capacity exhaustion.
- If requests are too high: This results in overprovisioning. On smaller nodes, the DaemonSet might request far more resources than it actually needs or uses, leading to significant wasted capacity and increased cloud costs. In extreme cases, if the request is too high, the DaemonSet pod might become unschedulable on smaller nodes that cannot meet the inflated request, failing to achieve its "run everywhere" mandate. Bryan Boreham highlighted how Google Cloud Managed Service for Prometheus, running with a high priority class, can overprovision and evict user workloads if its DaemonSet requests are set too high, leading to a situation where "you don't need to collect metrics from nothing."
The speakers illustrate this with real data from a Grafana Labs dev cluster, showing the widely varying CPU usage of a logging demon across 15 identical pods. This variability underscores the inefficiency of a single, static resource request. Existing, sub-optimal approaches to this problem include:
- YOLO (No Requests): Relying on Kubernetes defaults or no requests at all. This often results in misbehavior because the scheduler lacks critical information, though some platforms like GKE Autopilot now require and default requests.
- Conservative Approach: Setting requests high enough to cover the most resource-intensive instance. This prevents starvation but leads to substantial resource waste across most nodes, as illustrated by the "empty space under the purple line" in their example graph.
- Aggressive Approach: Setting requests very low and hoping for available burst capacity. This risks resource starvation and unpredictable behavior, similar to the YOLO approach.
- Divide and Conquer: Breaking a single DaemonSet into multiple DaemonSets, each with different node selectors and resource requests. While this limits the "blast radius" of over/under-provisioning, it significantly increases operational complexity. Each update requires modifying multiple resources, managing intricate node selectors, and undermining the inherent simplicity and value of the DaemonSet controller.
These limitations underscore the pressing need for a more automated, intelligent, and granular approach to DaemonSet resource management.
Key Findings
▶ Watch: Unpacking the core DaemonSet resource request challenge (3:40)
The central discovery and contribution presented in this talk is the proposal and demonstration of a novel extension to the Vertical Pod Autoscaler (VPA), enabling scoped resource recommendations for DaemonSets. The core idea is to move beyond the current VPA's monolithic recommendation model, where a single VPA object applies one set of recommendations across all targeted pods, to a more granular approach that can generate unique recommendations for subsets of pods based on specific criteria.
Specifically, the key finding is that by introducing a scope field into the VPA specification, it becomes possible to instruct the VPA to treat different groups of pods as distinct entities for autoscaling purposes. The prototype demonstrated utilizes scope: kubernetes.io/hostname, which effectively tells the VPA to:
- Monitor each DaemonSet pod independently: Instead of aggregating resource usage across all DaemonSet pods, the VPA tracks the historical resource consumption of each pod based on the node it's running on.
- Generate per-node recommendations: Based on the unique history of resource usage associated with a particular host, the VPA computes a customized CPU and memory recommendation for the DaemonSet pod running on that specific node.
- Apply tailored requests: When a DaemonSet pod needs to be (re)scheduled, the VPA's admission webhook applies the recommendation specific to the node it will land on.
This innovation directly addresses the "one size does not fit all" problem that plagues DaemonSets in heterogeneous clusters. By allowing resource requests to adapt dynamically to the actual demands of each node, this scoped VPA promises to:
- Reduce wasted resources and costs: Nodes with lower workloads will receive lower resource requests, freeing up capacity and reducing cloud expenditure.
- Mitigate stability problems: DaemonSet pods on highly utilized nodes will receive appropriate, higher resource requests, preventing throttling, OOM errors, and "noisy neighbor" scenarios.
- Simplify management: Operators no longer need to manually segment DaemonSets or guess at optimal static resource requests, offloading this complex task to an automated system.
This concept represents a significant evolution in Kubernetes resource management, offering a powerful mechanism to ensure both efficiency and reliability for critical DaemonSet workloads.
Technical Deep Dive
▶ Watch: Serious consequences of misconfigured DaemonSet resource requests (4:00)
The proposed enhancement builds upon the existing architecture of the Vertical Pod Autoscaler (VPA). For those unfamiliar, the VPA operates through three primary components:
- Recommender: This component continuously monitors the actual CPU and memory usage of the pods it is configured to manage. It collects a histogram of usage data, applies a decaying weight to these samples, and typically calculates a 90th percentile usage to make robust resource recommendations. These recommendations (CPU and memory requests) are then written back into the VPA object's status.
- Updater: The updater's role is to ensure pods are running with the recommended resources. If a pod's current resource requests fall outside the VPA's calculated range, the updater will evict that pod. Eviction triggers a re-scheduling event for the pod.
- Admission Webhook: This is a crucial component that intercepts pod creation requests flowing through the Kubernetes API server. When a pod is about to be created or re-created (e.g., after an eviction), the admission webhook intervenes. It checks if the pod is managed by a VPA and, if so, modifies the pod's resource requests (CPU and memory) to match the latest recommendation provided by the recommender before the pod is actually scheduled. This "on-the-fly" modification is key to how VPA works without directly altering the original workload definition (like a DaemonSet spec).
The core technical modification introduced by Adam Bernot and Bryan Boreham is the addition of a scope field to the VPA specification. This field allows operators to define how VPA should group and process resource usage data:
In this example, scope.matchLabels.kubernetes.io/hostname: "" instructs the VPA to treat each pod with a unique kubernetes.io/hostname label (which typically corresponds to a specific node) as its own independent scope. This means:
- The Recommender will maintain separate usage histories and generate distinct recommendations for the DaemonSet pod running on
node-1,node-2,node-3, and so on. - When the Updater evicts a DaemonSet pod, and the Admission Webhook intercepts its re-creation, the webhook will apply the recommendation specific to the node it's being scheduled onto.
This design ensures that a DaemonSet pod on a quiet node might receive a recommendation for 50m CPU, while its counterpart on a busy node could be recommended 200m CPU, all managed by a single VPA object.
The speakers noted that this functionality is being developed as a Kubernetes Enhancement Proposal (KEP) and is under active review by the SIG Autoscaling community. The prototype code demonstrated is available for review, with the KEP referenced as pull request 7942 and the prototype code as PR 7978 on the Kubernetes autoscaler repository. This open development process allows for community feedback on various aspects, including initial recommendations for newly added nodes (an open question) and potential future scope types like instance-type or node-pool, or even for StatefulSets.
Regarding VPA's internal logic, Bryan Boreham clarified that the VPA's recommender already uses a decaying weighted histogram for CPU usage and takes a 90th percentile to ignore short spikes. For memory, there's a special provision: if a pod experiences Out-Of-Memory (OOM) crashes, the VPA aggressively increases the memory recommendation to prevent further instability. These existing VPA behaviors are preserved and leveraged within the new scoped model.
Demo / Proof of Concept
▶ Watch: Real-world data shows DaemonSet CPU usage variability (6:20)
The talk featured a compelling live demonstration of the modified VPA's capabilities, showcasing its ability to dynamically adjust DaemonSet resource requests on a per-node basis.
Demo Setup:
The environment consisted of a small, three-node Kubernetes cluster running locally using Kind. For monitoring and observability, a Grafana Cloud stack was deployed, collecting metrics and logs from the cluster. The Grafana logging demon was chosen as the guinea pig DaemonSet for the test due to its variable workload characteristics based on log volume.
The Problem Illustrated:
Initially, the cluster was relatively idle. The Grafana dashboard displayed the CPU usage (as a thin line) and the CPU request (as a thick line) for the logging demon pods across all three nodes. As expected in an idle cluster, the actual CPU usage for all logging pods was very low, and their resource requests were also minimal. The challenge was to show how the VPA could adapt to asymmetrical load.
Introducing Asymmetrical Load:
To simulate a real-world scenario where one node experiences higher activity, Bryan Boreham deployed a separate pod designed to generate a significant volume of logs. Crucially, this log-generating pod was intentionally scheduled to run only on the rightmost node of the three-node cluster.
VPA's Dynamic Reaction:
The VPA's internal settings were temporarily tuned to react much faster than its default behavior (which can take hours) to accommodate the live demo's time constraints. As soon as the log-generating pod started running on the rightmost node, the Grafana dashboard immediately showed a dramatic increase in the actual CPU usage (the thin line) for the logging demon pod on that specific node. The CPU usage on the other two nodes remained low.
Within moments, the modified VPA detected this increased demand. The demo clearly showed the following sequence of events for the logging demon pod on the overloaded node:
- The VPA recognized that the pod's actual CPU usage significantly exceeded its current resource request.
- The VPA's updater component evicted the logging demon pod from that node.
- As the pod was re-created, the VPA's admission webhook intercepted the pod creation request.
- Crucially, because of the new
scope: kubernetes.io/hostnameconfiguration, the webhook applied a new, higher CPU request specifically tailored to the increased workload on that particular node. - The logging demon pod was then rescheduled on the same node, now with its CPU request (the thick line) elevated to better match its observed usage.
The Outcome:
The most impactful part of the demonstration was seeing the CPU request line for the logging demon on the rightmost node increase independently, while the CPU requests for the logging demons on the other two, less-loaded nodes remained unchanged. This visually confirmed that the scoped VPA successfully provided a customized recommendation for each pod based on its host's workload, solving the "one size does not fit all" problem in real-time. The burst of CPU during the pod's re-initialization (as it caught up on logs) was also visible, a behavior VPA's existing logic handles via its histogram and decaying weight.
This demonstration provided compelling evidence that the proposed VPA enhancement is technically feasible and delivers on its promise of intelligent, per-node resource autoscaling for DaemonSets.
Defensive Implications
▶ Watch: Exploring common but flawed DaemonSet resource allocation strategies (7:00)
The introduction of scoped VPA for DaemonSets carries significant implications for Kubernetes defenders, including cluster operators, platform engineers, and security teams. This feature offers powerful tools to enhance the efficiency, stability, and resilience of critical infrastructure components.
- Optimized Resource Utilization and Cost Reduction:
- Eliminate Overprovisioning Waste: Defenders can significantly reduce wasted CPU and memory resources on less busy nodes. By allowing DaemonSets to consume only what they truly need on a per-node basis, organizations can lower their cloud infrastructure costs.
- Reclaim Capacity: Freed-up resources can be reallocated to user workloads, improving overall cluster density and efficiency without sacrificing DaemonSet performance.
- Enhanced Stability and Reliability:
- Mitigate Resource Starvation: On heavily loaded nodes, DaemonSet pods will receive appropriately higher resource requests, preventing throttling and OOM errors that can lead to service degradation or outages for critical components like logging, metrics, or networking agents.
- Reduce Noisy Neighbor Issues: By accurately setting resource requests, DaemonSets are less likely to burst and consume excessive resources, preventing them from impacting the performance of other user workloads on the same node.
- Prevent User Workload Eviction: As highlighted by Bryan Boreham, DaemonSets with high priority classes (like Prometheus) can evict user workloads if overprovisioned. Scoped VPA helps prevent this undesirable cascading effect, maintaining a healthier balance between infrastructure and application pods.
- Simplified Management and Reduced Operational Burden:
- Automated Resource Tuning: Operators are relieved from the tedious and error-prone task of manually tuning DaemonSet resource requests or managing complex multi-DaemonSet strategies. The VPA automates this optimization continuously.
- Consistent Performance Across Heterogeneous Clusters: This feature makes it easier to manage DaemonSets across clusters with diverse node types (e.g., small edge nodes, large compute nodes) without compromising performance or efficiency on any node.
Considerations for Adoption and Best Practices:
- Workload Resilience to Restarts: The current VPA mechanism relies on pod eviction and restart to apply new resource requests. Defenders must ensure that DaemonSet workloads are designed to be resilient to these restarts, handling graceful shutdowns and quick recovery (e.g., logging agents snapshotting their position, as discussed in the Q&A).
- Interaction with Other Autoscalers: The integration with other autoscaling solutions like Cluster Autoscaler (which scales nodes) or Karpenter (which optimizes node provisioning) needs careful consideration. As mentioned in the Q&A, multi-dimensional scaling can be complex, and interactions must be understood to avoid contention or suboptimal scaling decisions. Early engagement with the KEP is crucial for providing feedback on these scenarios.
- Tuning VPA Reaction Speed: While the demo showed accelerated VPA, setting the recommender to react too quickly in production can introduce instability or excessive pod cycling. The VPA's default conservative behavior (using decaying weighted histograms and 90th percentile) is generally robust, but specific workload patterns might warrant careful tuning. Overly aggressive tuning can also increase the VPA controller's own resource consumption.
- Future "In-Place Pod Resizing": Defenders should be aware of ongoing work on in-place pod resizing, which would allow VPA to adjust resource requests without restarting the pod. This future enhancement would further improve the stability and disruption-free nature of DaemonSet autoscaling, especially for critical components like
kube-proxy. - GitOps Compatibility: As clarified in the Q&A, VPA modifies the pod's runtime spec via an admission webhook, not the original DaemonSet definition in Git. This means GitOps tools like Argo CD will not detect "drift" in the DaemonSet manifest itself, as the source of truth remains untouched. This behavior is intentional and generally desired for dynamic runtime adjustments.
- Community Engagement: Defenders are strongly encouraged to review the Kubernetes Enhancement Proposal (PR 7942) and the prototype code (PR 7978) and provide feedback. Input on potential new
scopetypes (e.g.,instance-type,node-pool) or specific edge cases is invaluable for shaping the final feature.
By embracing this next generation of DaemonSet autoscaling, defenders can build more robust, efficient, and self-optimizing Kubernetes platforms, reducing operational toil and enhancing the overall reliability of their infrastructure.
Key Takeaways
- The "One Size Fits All" DaemonSet Problem: DaemonSets currently require a single resource request for all pods, leading to significant overprovisioning and waste on smaller nodes, or critical resource starvation and instability on larger, busier nodes in heterogeneous Kubernetes clusters.
- Scoped VPA as the Solution: A proposed enhancement introduces a
scopefield to the Vertical Pod Autoscaler (VPA), enabling it to generate and apply tailored resource recommendations for DaemonSet pods based on specific criteria, such as thekubernetes.io/hostnameof the node. - Per-Node Resource Optimization: This innovation allows DaemonSet pods on different nodes to receive distinct, optimized CPU and memory requests, dynamically adapting to the actual workload demands of each host.
- Demonstrated Efficiency and Stability: A live demo successfully showed a modified VPA dynamically increasing CPU requests for a logging DaemonSet pod on a loaded node, while leaving pods on other nodes untouched, proving the concept's ability to reduce waste and improve stability.
- Significant Operational Benefits: Adopting scoped VPA can lead to substantial cost savings through reduced overprovisioning, enhanced cluster stability by mitigating resource starvation and "noisy neighbor" issues, and a significant reduction in manual operational overhead for managing DaemonSet resources.
- Active Development and Community Involvement: The feature is under active development as a Kubernetes Enhancement Proposal (KEP PR 7942), with prototype code available (PR 7978). Community feedback is crucial for shaping its final form, including considerations for initial recommendations for new nodes and interactions with other autoscaling components.
About the Speaker(s)
Adam Bernot is a Staff Software Engineer at Google Cloud. His work focuses on the Google Cloud managed service for Prometheus, a critical component for monitoring and observability in cloud environments. His expertise lies in building and managing scalable systems for metrics collection and analysis.
Bryan Boreham is a Distinguished Engineer at Grafana Labs and a dedicated Prometheus maintainer. He specializes in developing massively scalable storage solutions for metrics, logs, traces, and profiles. His deep involvement with Prometheus and large-scale observability systems highlights his expertise in distributed systems and performance optimization within cloud-native environments.
Reviews
Dr. Zero (Offensive Security Researcher) — MUST SEE
This talk presents a groundbreaking, elegant solution to a long-standing and costly problem in Kubernetes: the inefficient resource allocation for DaemonSets in heterogeneous clusters. By extending the Vertical Pod Autoscaler (VPA) with 'scoped recommendations,' Bernot and Boreham offer a novel approach to dynamically tailor resource requests on a per-node basis, promising significant cost savings, enhanced stability, and reduced operational overhead. The technical depth, practical implications, and the live demonstration make this an essential watch for anyone operating Kubernetes at scale.
Heather Calloway (CISO) — STRONG ACCEPT
This talk presents a crucial enhancement to Kubernetes DaemonSet management, addressing the long-standing 'one size fits all' resource allocation problem. By introducing scoped recommendations into the Vertical Pod Autoscaler, it enables per-node resource tuning for critical infrastructure components, including security agents. This innovation offers substantial benefits in cost optimization, operational stability, and reduced management overhead, directly impacting an organization's resilience and efficiency in cloud-native environments.