Taking Care of Your Control Plane With API Priority and Fairness an... Matteo Ruina & Ayaz Badouraly

Matteo Ruina, Ayaz Badouraly

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

In this comprehensive talk, Ayaz Badouraly and Matteo Ruina, software engineers at DataDog, delve into the critical challenges of maintaining the stability and reliability of the Kubernetes control plane in large-scale, multi-tenant environments. DataDog, a company running hundreds of thousands of pods across tens of thousands of nodes in dozens of Kubernetes clusters, manages its own control planes, providing them with unique insights into potential failure modes and effective mitigation strategies. The speakers illuminate how a single misbehaving user or application can disproportionately impact the entire cluster, leading to widespread outages.

Watch on YouTube

Visual summary for Taking Care of Your Control Plane With API Priority and Fairness an... Matteo Ruina & Ayaz Badouraly by Matteo Ruina, Ayaz Badouraly
Visual summary for Taking Care of Your Control Plane With API Priority and Fairness an... Matteo Ruina & Ayaz Badouraly by Matteo Ruina, Ayaz Badouraly

Key moments

  1. 0:00 Introduction and problem scope: control plane reliability
  2. 2:00 Incident: Runaway admission webhook creating excessive pods
  3. 4:00 Incident: Cilium CRD update causing API server overload
  4. 4:30 Incident: Kubelet traffic misconfiguration overwhelming single API server
  5. 4:50 Introducing API Priority and Fairness (APF) as a solution
  6. 6:40 Visual explanation of API Priority and Fairness (APF) workflow

Taking Care of Your Control Plane With API Priority and Fairness and Resource Quota

Speakers: Matteo Ruina, Software Engineer, DataDog; Ayaz Badouraly, Software Engineer, DataDog

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=c73SzCKx-OY

Overview

In this comprehensive talk, Ayaz Badouraly and Matteo Ruina, software engineers at DataDog, delve into the critical challenges of maintaining the stability and reliability of the Kubernetes control plane in large-scale, multi-tenant environments. DataDog, a company running hundreds of thousands of pods across tens of thousands of nodes in dozens of Kubernetes clusters, manages its own control planes, providing them with unique insights into potential failure modes and effective mitigation strategies. The speakers illuminate how a single misbehaving user or application can disproportionately impact the entire cluster, leading to widespread outages.

The core of their presentation focuses on two powerful, albeit often underutilized or misconfigured, Kubernetes native features: API Priority and Fairness (APF) and Resource Quotas. They meticulously detail how these tools can be leveraged to prevent control plane overload, ensuring that critical system components and legitimate user requests are prioritized. Beyond merely introducing these features, Badouraly and Ruina share DataDog's real-world experiences, including the challenges faced with default configurations, the lessons learned through incidents, and the innovative solutions developed to dynamically manage resource consumption and protect the control plane from accidental or malicious overloads. The talk provides invaluable guidance for both cluster administrators operating their own Kubernetes infrastructure and users of managed Kubernetes services, demonstrating how to proactively safeguard the heart of their container orchestration platform.

Background

▶ Watch: Introduction and problem scope: control plane reliability (0:00)

DataDog operates a massive Kubernetes estate, hosting over 2,000 engineers and a vast array of workloads, including critical monitoring applications. This multi-tenant environment, characterized by shared, resource-constrained control planes, frequently presented scenarios where a single tenant or misconfigured application could severely degrade or completely take down the entire cluster. The speakers recount several critical incidents that underscore the fragility of the control plane when proper guardrails are not in place.

One such incident involved a bad deployment of an admission webhook that inadvertently wiped out labels from pods. This seemingly minor issue cascaded into a catastrophic event: the ReplicaSet controller, unable to find its owned pods due to the missing labels, continuously recreated them. These new pods would then pass through the same faulty webhook, have their labels removed, and the cycle would repeat, leading to the creation of "hundreds of thousands of pods over the course of a couple of hours"—far exceeding Kubernetes' recommended scalability limits.

Another common challenge arose from Spark jobs requesting thousands of executor pods simultaneously. While Spark jobs are a legitimate workload, their bursty nature and large-scale resource demands could overwhelm the control plane. Controllers responsible for managing pod and node lifecycles struggled to keep up with the sudden surge in API requests, leading to unresponsive clusters.

Even seemingly internal components could trigger control plane instability. DataDog uses Cilium for networking, and a Cilium CRD update would cause the API server to drop all active watch requests. This, in turn, prompted all Cilium agents on all nodes to simultaneously issue list requests to the API server, creating a massive flood of traffic that could take down both the API server and etcd.

Perhaps the most illustrative incident involved DNS shenanigans and a load balancing misconfiguration. All kubelet traffic was inadvertently directed to a single API server instance, completely overwhelming it, while other instances remained idle. This scenario highlighted how even external infrastructure issues could translate into severe control plane degradation. These incidents collectively demonstrated the urgent need for robust mechanisms to isolate and protect the Kubernetes control plane from internal and external pressures.

Key Findings

▶ Watch: Incident: Cilium CRD update causing API server overload (4:00)

The talk's key findings revolve around the practical application and advanced tuning of two foundational Kubernetes features to address the control plane stability issues described: API Priority and Fairness (APF) and Resource Quotas.

Firstly, the speakers assert that while APF is a powerful mechanism to prevent one type of API request from monopolizing the API server's resources, its default configuration in Kubernetes 1.20+ is often insufficient for complex multi-tenant environments. DataDog discovered that proactive, fine-grained tuning of Flow Schemas and Priority Levels is essential. They found that exempting critical observability endpoints, creating flow schemas per namespace and for human users, and implementing a "tarpit" priority level for incident response were crucial steps. A significant discovery was the borrowing mechanism introduced in Kubernetes 1.23, which, while generally beneficial for resource utilization, inadvertently undermined the effectiveness of strict priority levels like the tarpit. Disabling borrowing for specific priority levels became a necessary tuning step. Furthermore, they linked APF utilization metrics to hardware autoscaling by configuring maxRequestsInflight per core, allowing for more responsive scaling of the API servers.

Secondly, for limiting the raw quantity of resources within a cluster, Resource Quotas were identified as the primary tool. However, static quotas proved problematic in dynamic environments. DataDog's most significant contribution in this area is the development of a dynamic resource quota controller. This custom controller observes actual resource utilization (pods, config maps) within each namespace and dynamically adjusts quotas, maintaining a buffer for organic growth while imposing a cooldown period and a maximum limit to prevent uncontrolled spikes. Key learnings from this controller's development included optimizing memory usage with PartialObjectMetadataList for ConfigMaps and devising a workaround for a long-standing Kubernetes 409 Conflict admission controller bug that could incorrectly reject resource creations even when quota was available. Together, these findings demonstrate that while Kubernetes provides the building blocks, significant effort and custom solutions are often required to achieve resilient control plane operations at scale.

Technical Deep Dive

▶ Watch: Incident: Kubelet traffic misconfiguration overwhelming single API server (4:30)

The technical core of the presentation meticulously dissects the inner workings and advanced configurations of API Priority and Fairness (APF) and Resource Quotas.

API Priority and Fairness (APF)

APF, available since Kubernetes 1.20, is designed to protect the API server from overload by prioritizing and throttling incoming requests.

APF Fundamentals

  1. Flow Schemas: These objects filter incoming API requests based on criteria such as the user (group or service account), resource being queried, and HTTP verb. For instance, a flow schema might target requests from system:nodes for non-status leases—crucial for node heartbeats.
  2. Priority Levels: Once a request matches a flow schema, it is assigned to a priority level. This level defines how requests are prioritized and throttled. Key configurations include:
  • Concurrency Shares: Determines the proportion of the API server's total concurrent requests a priority level can handle.
  • Queues: Each priority level can have multiple queues. Requests are sharded across these queues using a shuffle sharding mechanism, which uses a "hand size" to distribute requests from different flow schemas into different queues, preventing a single flow from monopolizing one queue.
  • Queue Limit: The maximum number of requests a queue can hold.
  • Fair Scheduling Algorithm: Dispatches requests from the queues to available workers. The total number of available "seats" (concurrent requests) is defined by the maxRequestsInFlight flag on the API server, which is then split among priority levels.

When a specific flow schema's assigned queues become full, throttling kicks in, and subsequent requests are rejected with a 429 Too Many Requests HTTP status code. Crucially, this happens without affecting requests from other flow schemas or priority levels that have available queue capacity. DataDog tracks golden metrics for APF, including the number of requests dispatched, queue latency, and the count of 429 errors served, to monitor its effectiveness.

Tuning APF for Production Workloads

DataDog found that default APF configurations were inadequate and required significant tuning:

  • Exempting Critical Endpoints: They introduced a flow schema for /metrics and /debug/pprof endpoints, marking them as exempt. This ensures that even under heavy API server contention, observability data can still be collected, which is paramount during incidents.
  • Granular Flow Schemas:
  • Per-Namespace Flow Schemas: All service accounts within a given namespace are assigned to a dedicated flow schema.
  • Human User Flow Schemas: kubectl commands from human users (e.g., DataDog's engineering group) are assigned a separate flow schema and prioritized for manual operations during incidents.
  • "Tarpit" Priority Level: A highly constrained priority level was introduced with very limited concurrent shares (e.g., 5) and configured to reject requests immediately rather than queuing them. This "tarpit" can be dynamically applied to an offending flow schema during an incident to quickly stop runaway processes.
  • The Kubernetes 1.23 Borrowing Mechanism: A significant challenge arose with Kubernetes 1.23, which introduced a borrowing mechanism. This feature allows priority levels with available "seats" to lend them to other priority levels that are at capacity. While generally beneficial for utilization, it undermined the strict isolation of the "tarpit," allowing it to borrow capacity and thus fail to throttle effectively. DataDog's solution was to explicitly disable borrowing for the tarpit priority level by setting borrowingLimitPercent to 0% and lendablePercent to 100% (to give away any unused capacity).
  • DaemonSet Prioritization: For critical DaemonSets like Cilium agents, a dedicated priority level was created. While a single-subject flow schema might initially use one queue, DataDog found that under load, this single queue could become a bottleneck. The solution was to split the single flow schema's traffic over multiple queues to allow workers to pick up requests concurrently, preventing throttling under normal conditions and gracefully handling bursts (e.g., during CRD updates).
  • Autoscaling API Servers: DataDog correlated APF utilization with hardware autoscaling. Instead of scaling based on APF utilization directly (which is an abstract limit, not hardware capacity), they configured maxRequestsInFlight to be linear in the number of CPU cores (maxRequestsInFlightLinear). This allowed them to use standard CPU-based autoscaling for API servers, ensuring that hardware resources scaled proportionally to the APF capacity.

Resource Quotas

Resource Quotas are Kubernetes objects designed to limit resource consumption per namespace.

Resource Quota Fundamentals

Resource quotas can limit:

  • Count of specific objects: e.g., pods, configmaps, services.
  • Total amount of compute resources: e.g., cpu, memory, gpu.

When a user attempts to create a resource that would violate a quota, the API server returns a 403 Forbidden error, indicating which resource limit would be exceeded.

Challenges with Static Resource Quotas

Static resource quotas proved ineffective for DataDog's dynamic workloads:

  • Too Low: Could block legitimate, organic scaling during normal operations or minor traffic shifts.
  • Too High: Rendered the quota useless against runaway processes or human errors.
  • Stale: Quotas could quickly become outdated as traffic patterns and application needs evolved over weeks or months.

Dynamic Resource Quota Controller

To overcome the limitations of static quotas, DataDog developed a custom controller that dynamically adjusts resource quotas per namespace.

  1. Observation and Adjustment: The controller continuously watches resource utilization (currently pods and config maps) in each namespace.
  2. Buffer and Cool-down: It sets the quota based on current utilization plus a configurable buffer percentage (e.g., 40%). This allows for immediate, small-scale upsizing.
  3. Step Profile and Maximum Limit: In case of sudden spikes, the quota increases in a "step profile" with a cool-down period between steps, preventing uncontrolled growth. It also enforces a maximum allowed quota per namespace, typically based on Kubernetes scalability thresholds and DataDog's operational experience.
  4. Configuration Parameters:
  • Buffer: Percentage above current utilization (e.g., 40% for 10 pods results in a 14-pod quota).
  • Default Minimum: A sensible minimum quota (e.g., 2 pods) to prevent very small, percentage-based quotas from being too restrictive.
  • Maximum Allowed: An upper bound for the quota, informed by scalability recommendations.
  1. Self-Service and Admin Overrides:
  • User Annotation: Users can annotate their ResourceQuota object to temporarily set the hard limit to the maximum allowed. This provides immediate relief during incidents and has a 7-day Time-To-Live (TTL) to encourage re-engagement with the controller.
  • Admin Annotation: An annotation on the kube-system namespace triggers the controller to delete all managed resource quotas, effectively restoring the cluster to a pre-quota state.
  1. Lessons Learned during Development:
  • Memory Optimization for ConfigMaps: Listing ConfigMap objects to count them can be memory-intensive, especially with Helm deployments creating many large config maps. DataDog switched to meta/v1.PartialObjectMetadataList, which retrieves only the metadata, significantly reducing memory consumption (over 4x reduction observed).
  • Triggering Reconciliation from Inside Loop: To trigger reconciliations for all namespaces when the admin override is used, they leveraged a source channel. This allows sending generic events from within the reconciliation loop, which the controller watches.
  • ConfigMap Controller Behavior: Unlike pods, ConfigMap objects do not have a controller that retries creation on quota errors. This means Helm deployments, for example, would fail outright if a ConfigMap quota was hit. The solution was to configure different, less disruptive minimum quotas and cool-down periods specifically for ConfigMaps compared to Pods.
  • The 409 Conflict Admission Controller Bug: A long-standing upstream Kubernetes issue (since 2018) causes the ResourceQuota admission controller to return a 409 Conflict error even when there's enough room for a resource. This happens because the admission controller tries to update the ResourceQuota object up to three times to reflect the new state, but if it encounters a stale object version (due to concurrent updates), it exhausts its retries and bubbles up a 409, even if the resource could have been created. DataDog's workaround ("cheat") was to modify the admission controller to not count 409 conflicts as retries, allowing it to fetch a fresh version of the object and attempt the update again, all while respecting the 10-second timeout. This effectively resolved the spurious 409 errors without a significant increase in admission duration.

Demo / Proof of Concept

▶ Watch: Introducing API Priority and Fairness (APF) as a solution (4:50)

The talk did not feature a live demonstration or a dedicated proof-of-concept section. Instead, the speakers presented a series of real-world incidents DataDog faced, detailing how they applied and iteratively refined their use of API Priority and Fairness and the custom dynamic resource quota controller to mitigate these issues. They showed graphs illustrating the impact of their solutions, such as the tarpit priority level effectively limiting traffic and the dynamic resource quota controller adjusting limits in response to workload patterns. The incidents themselves served as the "proof" that these mechanisms were necessary and effective in a large-scale, multi-tenant Kubernetes environment.

Defensive Implications

▶ Watch: Visual explanation of API Priority and Fairness (APF) workflow (6:40)

The insights shared by DataDog offer several critical defensive implications for any organization operating Kubernetes at scale, particularly in multi-tenant or high-density environments.

  1. Proactive APF Configuration: Do not rely solely on default API Priority and Fairness settings. Cluster administrators should:
  • Exempt Observability Endpoints: Ensure /metrics and /debug/pprof endpoints are exempt from throttling to maintain visibility during control plane distress.
  • Create Granular Flow Schemas: Implement flow schemas for specific workloads (e.g., DaemonSets like Cilium), per-namespace service accounts, and distinct groups of human users. Prioritize critical system components and manual incident response operations.
  • Implement a "Tarpit" Priority Level: Develop a highly restrictive priority level that can be dynamically applied to misbehaving flow schemas to quickly contain runaway processes and prevent cluster-wide impact.
  • Understand and Manage Borrowing: Be aware of the APF borrowing mechanism introduced in Kubernetes 1.23. For priority levels requiring strict isolation (like the "tarpit"), disable borrowing by setting borrowingLimitPercent to 0% and lendablePercent to 100%.
  • Optimize Queue Configuration: For flow schemas with a single subject but high throughput (e.g., DaemonSets on many nodes), ensure they are sharded across multiple queues within their priority level to prevent single-queue bottlenecks.
  • Correlate APF with Autoscaling: Integrate APF utilization metrics with hardware autoscaling strategies for API servers by configuring maxRequestsInFlightLinear to scale with CPU cores, enabling more effective resource allocation.
  1. Dynamic Resource Quota Implementation: Static resource quotas are insufficient for dynamic, large-scale clusters. Defenders should consider:
  • Developing or Adopting a Dynamic Quota Controller: A controller that observes resource utilization and automatically adjusts quotas with buffers, cooldowns, and hard maximums is crucial for preventing uncapped resource growth while allowing legitimate scaling.
  • Optimizing Controller Efficiency: When counting resources like ConfigMaps, use meta/v1.PartialObjectMetadataList to minimize memory consumption and improve performance.
  • Handling Resource-Specific Behaviors: Recognize that different resource types (e.g., Pods vs. ConfigMaps) have different controller retry mechanisms. Adjust dynamic quota parameters (minimums, buffers, cooldowns) accordingly to prevent unexpected failures.
  • Addressing Upstream Issues: Implement workarounds for known upstream Kubernetes bugs, such as the 409 Conflict admission controller issue. If directly modifying the admission controller isn't feasible, ensure client-side retries are robust for 409 errors, especially for ConfigMap-heavy deployments.
  1. Comprehensive Monitoring and Alerting: Monitor APF metrics (queue latency, 429s) alongside traditional hardware metrics (CPU, memory) to gain a holistic view of control plane health and identify bottlenecks early.
  1. User Education and Self-Service: Provide mechanisms for users to temporarily bypass strict quotas during incidents (e.g., through annotations) but ensure these overrides have a time-to-live to prevent permanent circumvention. Educate users on quota limits and incident response procedures (e.g., traffic shifting).

By adopting these defensive strategies, organizations can significantly enhance the resilience of their Kubernetes control planes, minimize the blast radius of misconfigurations or runaway applications, and ensure continuous availability for all tenants.

Key Takeaways

  • Kubernetes Control Plane Stability is Fragile: In large, multi-tenant clusters, a single misconfiguration or runaway application can easily overload the API server and disrupt the entire platform. Proactive defense mechanisms are essential.
  • API Priority and Fairness (APF) Requires Custom Tuning: While APF is a powerful native Kubernetes feature, its default configuration is often inadequate for complex environments. Custom Flow Schemas, Priority Levels, and careful management of features like the borrowing mechanism (especially since Kubernetes 1.23) are critical for effective control plane protection.
  • Dynamic Resource Quotas are Indispensable: Static resource quotas are insufficient for dynamic workloads. A custom controller that dynamically adjusts quotas based on real-time utilization, with buffers, cooldowns, and maximum limits, is necessary to prevent uncapped resource growth while allowing legitimate scaling.
  • Optimize for Efficiency and Edge Cases: When implementing custom controllers or tuning native features, be mindful of resource consumption (e.g., using PartialObjectMetadataList for ConfigMaps) and specific component behaviors (e.g., ConfigMap controllers not retrying on quota errors).
  • Upstream Issues Demand Workarounds: Be prepared to implement workarounds for known Kubernetes upstream bugs, such as the 409 Conflict admission controller issue, to ensure smooth operation and a positive user experience.
  • Metrics and Observability are Paramount: Comprehensive monitoring of APF metrics (queue latency, 429s) and resource quota utilization is vital for understanding control plane health, identifying bottlenecks, and validating the effectiveness of defensive measures.

About the Speaker(s)

Ayaz Badouraly is a Software Engineer at DataDog. He began his journey in the Site Reliability Engineering (SRE) team, where he gained invaluable experience in operating highly available systems. He now applies these SRE best practices to DataDog's internal Kubernetes platform, focusing on ensuring its reliability and performance at scale.

Matteo Ruina is also a Software Engineer at DataDog, specializing in the Kubernetes control plane and cluster lifecycle automation. His work involves managing the deployment and operation of the underlying control planes for DataDog's extensive Kubernetes clusters, contributing to their robustness and efficiency.

Together, their combined expertise from SRE and control plane development provides a unique perspective on the challenges and solutions for maintaining Kubernetes stability in one of the largest production environments.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk from DataDog engineers is a masterclass in Kubernetes control plane hardening for large-scale, multi-tenant environments. It goes deep into the practical application and advanced tuning of API Priority and Fairness (APF) and Resource Quotas, sharing battle-tested strategies and novel workarounds for critical, often overlooked, stability issues. This isn't theoretical fluff; it's a no-bullshit exposition of hard-won lessons from the trenches, complete with real-world incidents and clever engineering solutions that genuinely advance the state of the art in Kubernetes operations.

Heather Calloway (CISO) — MUST SEE

This KubeCon talk from DataDog engineers is a critical examination of Kubernetes control plane resilience, delivering actionable strategies for preventing widespread outages in large-scale, multi-tenant environments. It moves beyond theoretical discussions, presenting real-world incidents and DataDog's custom solutions for dynamically tuning API Priority and Fairness and Resource Quotas. The session provides indispensable guidance for platform owners and security leaders seeking to harden their Kubernetes infrastructure and ensure institutional accountability for operational stability.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025