Lightning Talk: Kueue: Save Some QPS for the Rest of Us! How To Manage 100k Updates Pe... P. Bundyra

P. Bundyra

KubeCon + CloudNativeCon Europe 2025 · Lightning Talk

Overview

This lightning talk, delivered by P. Bundyra, a Google engineer, introduces Kueue, a cloud-native queuing system specifically designed for managing AI workloads within Kubernetes clusters. Developed in collaboration with the open-source community since 2022, Kueue acts as a sophisticated "bouncer" for the cluster, intelligently deciding which jobs or pods can run and which must wait, thereby preventing resource contention and ensuring fair access. The core problem addressed in this presentation is the challenge of efficiently exposing the position of a workload in a queue when dealing with potentially tens of thousands of active workloads, without overwhelming the Kubernetes API server or its underlying etcd datastore.

Watch on YouTube

Visual summary for Lightning Talk: Kueue: Save Some QPS for the Rest of Us! How To Manage 100k Updates Pe... P. Bundyra by P. Bundyra
Visual summary for Lightning Talk: Kueue: Save Some QPS for the Rest of Us! How To Manage 100k Updates Pe... P. Bundyra by P. Bundyra

Key moments

  1. 0:00 Introduction to Kueue: a cloud-native queuing system
  2. 1:10 Initial idea: updating workload positions via CRDs fails at scale
  3. 1:45 Second idea: single CRD for queue order hits etcd limits
  4. 2:20 The solution: Kubernetes API Aggregation Layer
  5. 3:07 Kueue's mechanism for workload positioning using aggregation
  6. 4:00 Comparison: CRDs vs. API Aggregation Layer trade-offs
  7. 4:45 Kueue's hybrid approach using both CRDs and aggregation layer

Lightning Talk: Kueue: Save Some QPS for the Rest of Us! How To Manage 100k Updates Pe... P. Bundyra

Speakers: P. Bundyra, Google Engineer

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=njNXlZNT3dw

Overview

This lightning talk, delivered by P. Bundyra, a Google engineer, introduces Kueue, a cloud-native queuing system specifically designed for managing AI workloads within Kubernetes clusters. Developed in collaboration with the open-source community since 2022, Kueue acts as a sophisticated "bouncer" for the cluster, intelligently deciding which jobs or pods can run and which must wait, thereby preventing resource contention and ensuring fair access. The core problem addressed in this presentation is the challenge of efficiently exposing the position of a workload in a queue when dealing with potentially tens of thousands of active workloads, without overwhelming the Kubernetes API server or its underlying etcd datastore.

The talk highlights a critical scalability bottleneck inherent in standard Kubernetes extension mechanisms, particularly when high-frequency updates or large data payloads are involved. Bundyra demonstrates how conventional approaches to tracking queue positions, such as directly updating Custom Resource Definitions (CRDs) or storing a massive list in a single CRD, quickly hit the practical limits of etcd's Query Per Second (QPS) capacity and object size. The significance of this talk lies in its exploration of alternative Kubernetes extension mechanisms, specifically the Kubernetes API Aggregation Layer, as a powerful solution for performance-critical scenarios. It offers valuable insights for developers and operators grappling with large-scale, dynamic workloads in Kubernetes, particularly in the burgeoning field of AI/ML, where efficient resource management and visibility are paramount.

Background

▶ Watch: Introduction to Kueue: a cloud-native queuing system (0:00)

The proliferation of AI and machine learning workloads in Kubernetes environments has introduced new challenges related to resource orchestration and scheduling. These workloads often consist of numerous, sometimes short-lived, jobs that compete for finite cluster resources. To manage this, queuing systems like Kueue are essential, providing a structured way to admit jobs into the cluster, prevent overload, and ensure fairness. Kueue achieves this by abstracting various job types—such as Pod jobs, KubeFlow jobs, and Ray jobs—under a unified Workload CRD. This abstraction allows Kueue to apply consistent queuing logic across diverse AI frameworks.

A fundamental requirement for any queuing system is the ability for users to understand their position in the queue. For a user submitting a job, knowing "how many people are in front of me" is crucial for managing expectations and planning. However, providing this information efficiently in a Kubernetes-native way presents a significant technical hurdle, especially at scale. The speaker outlines two initial, intuitive approaches that were considered and subsequently rejected due to inherent limitations of the Kubernetes API and etcd, its distributed key-value store.

The first approach involved extending the Workload object itself with a position field. Every time the queue state changed—a workload was admitted, or a new one was added—all affected Workload objects would be updated. While seemingly straightforward for small queues, this strategy quickly becomes untenable with "tens of thousands of workloads." The Kubernetes API server, and by extension etcd, would be subjected to an immense volume of write operations (QPS), leading to performance degradation and potential instability. Etcd is optimized for stability and consistency, not for high-frequency, large-scale updates across many objects.

The second alternative considered was to create a completely separate CRD that would store the entire ordered list of workloads in the queue. This would centralize the queue state into a single object, reducing the number of individual updates. Instead of updating thousands of Workload objects, only one object would need modification. However, this approach runs into another hard limit: etcd has a practical size limit for individual objects, typically around 1.5MB. Storing a list of tens of thousands of workload entries within a single etcd object would quickly exceed this limit, making the approach impractical for large queues. These limitations underscore the necessity for a more sophisticated, high-performance solution for specific data management challenges within the Kubernetes ecosystem.

Key Findings

▶ Watch: Second idea: single CRD for queue order hits etcd limits (1:45)

The central finding of this talk is that while Custom Resource Definitions (CRDs) are the de facto standard for extending Kubernetes and are excellent for managing state, they are not always the optimal choice for scenarios requiring extremely high-frequency updates or storing very large, dynamically generated datasets. For such specific, performance-critical use cases, the Kubernetes API Aggregation Layer offers a powerful and necessary alternative.

Kueue's innovative solution leverages this aggregation layer to expose workload positions. Instead of storing explicit positions for every workload persistently, Kueue internally maintains its queue using a heap structure. This structure inherently knows the head of the queue (who is first) but doesn't assign specific, sequential positions to all elements until requested. The key insight is that the full list of positions is a derived, transient piece of information rather than a core, persistent state.

The system is designed with a "lazy" evaluation principle: Kueue does not proactively generate or update the list of workload positions. Instead, it only performs this computationally intensive task when a user explicitly queries for it. Upon receiving such a request, Kueue takes a snapshot of its internal heap, processes this snapshot to assign appropriate positions to each workload, and then presents this dynamically generated list to the user. This on-demand, just-in-time approach dramatically reduces the load on the API server and etcd, ensuring high performance and scalability even with tens of thousands of workloads in the queue. This finding demonstrates a critical architectural pattern for extending Kubernetes efficiently for highly dynamic data.

Technical Deep Dive

▶ Watch: The solution: Kubernetes API Aggregation Layer (2:20)

Kueue's architecture is rooted in its role as a cloud-native queuing system for diverse AI workloads. It provides a unified control plane for managing the lifecycle and scheduling of these jobs, abstracting them under an umbrella CRD called Workload. This allows Kueue to interact uniformly with various Kubernetes job types, including native Pod jobs, KubeFlow jobs, and Ray jobs, ensuring consistent queuing and admission policies.

The speaker meticulously detailed the technical rationale behind rejecting conventional Kubernetes extension methods for the specific problem of exposing queue positions. The first method, directly extending the Workload object with a position field, would involve updating potentially "tens of thousands of workloads" every time the queue state changed. This translates to an unmanageable volume of QPS (Queries Per Second) directed at the Kubernetes API server and its etcd backend. Etcd, while robust for storing configuration and desired state, is not designed for such high-throughput, volatile data updates across a vast number of objects. It would quickly become a bottleneck, leading to degraded API server performance and potential instability for the entire cluster.

The second rejected approach, storing the entire queue order in a single, dedicated CRD, aimed to reduce the QPS by centralizing updates. However, this hit another fundamental limitation: the size limit of an individual etcd object. While not explicitly stated, common etcd configurations typically limit individual object sizes to around 1.5MB. A list containing metadata for tens of thousands of workloads would quickly exceed this limit, rendering the single-object CRD approach infeasible for large queues.

To overcome these obstacles, Kueue adopted the Kubernetes API Aggregation Layer. This powerful, yet less commonly used, extension mechanism allows developers to extend the Kubernetes API with custom resources that are not backed by etcd. Instead, an aggregated API server can serve these resources from an entirely different storage backend. As Bundyra humorously put it, "It may be a banana, it may be a washing machine. If it implements the proper interface, then why not?" In Kueue's case, the "storage" for the queue order is Kueue itself.

Here's how Kueue leverages API Aggregation for workload positioning:

  1. Internal Data Structure: Kueue internally manages its queue using a heap structure. This data structure is optimized for quickly identifying the highest-priority element (the head of the queue) and efficiently inserting/deleting elements. Crucially, it does not maintain an explicit, ordered list of all elements with assigned positions.
  2. On-Demand Snapshotting: When a user queries for the queue order (e.g., to see their workload's position), Kueue does not retrieve pre-computed data. Instead, it takes a snapshot of its current internal heap state.
  3. Position Assignment: This snapshot is then processed to dynamically generate a list of workloads with their appropriate, sequential positions assigned. This computation happens in real-time at the moment of the user's request.
  4. Lazy Evaluation: The entire process is "really lazy about it." Kueue performs this snapshotting and position assignment only when a user explicitly asks for it. This design choice is critical for performance. By avoiding continuous updates or pre-computation of the full queue order, Kueue minimizes its own computational overhead and, more importantly, avoids generating any QPS load on the primary Kubernetes API server and etcd for this specific feature.

The talk provides a clear comparison between CRDs and the API Aggregation Layer:

  • CRDs:
  • Storage: Store object state in etcd.
  • Ease of Use: Very easy to set up.
  • Performance: May not be performant for high-frequency updates or very large objects (like in Kueue's case).
  • Prevalence: Used extensively across the Kubernetes ecosystem (e.g., ServiceMonitor, JobSet, NetworkPolicy).
  • API Aggregation Layer:
  • Storage: Uses its own custom storage solution (e.g., Kueue's internal heap).
  • Ease of Use: Comes with a bigger overhead; harder to set up and requires additional coding.
  • Performance: Much more performant in specific cases requiring dynamic data generation or custom storage, such as Kueue's positioning or periodically collecting metrics.
  • Prevalence: Less popular than CRDs, making it potentially harder to find examples or community support for specific issues.

Kueue intelligently combines both approaches: the "vast majority of [its] logic relies on CRDs" for managing workload state and other static configurations, while leveraging the "AP aggregation layer" specifically for the performance-critical task of dynamically exposing workload positions. This hybrid strategy exemplifies choosing the right tool for the right job, balancing ease of use with the necessity for high performance in a scalable Kubernetes environment.

Demo / Proof of Concept

▶ Watch: Comparison: CRDs vs. API Aggregation Layer trade-offs (4:00)

Given the lightning talk format and its concise nature, a live demonstration or a detailed, step-by-step proof-of-concept setup was not presented. The talk focused primarily on the architectural challenge and the technical solution employed within Kueue. However, the conceptual demonstration of the problem—the inability of the API server and etcd to handle tens of thousands of updates or a single massive object—and the proposed solution—the API Aggregation Layer with lazy evaluation—serves as a strong conceptual proof. The implication of Kueue's design is that users can query the status and position of their workloads, even within a queue of 10,000+ items, without experiencing delays or causing performance issues for the entire Kubernetes cluster, which would be the tangible "proof" of its success in a real-world scenario.

Defensive Implications

▶ Watch: Kueue's hybrid approach using both CRDs and aggregation layer (4:45)

While this talk isn't directly about security vulnerabilities, it offers crucial insights for Kubernetes operators, developers, and architects on building robust, scalable, and resilient systems. The "defensive implications" here pertain to safeguarding the stability and performance of the Kubernetes control plane against self-inflicted wounds caused by inefficient API usage.

  1. Understand Kubernetes API and Etcd Limits: The most significant defensive takeaway is the explicit recognition of etcd's QPS limits and single-object size constraints. Developers must internalize that etcd is a highly consistent, distributed key-value store optimized for metadata and desired state, not a high-throughput database for dynamic, frequently changing data or massive payloads. Attempting to use CRDs for such purposes, as illustrated by the rejected approaches, can lead to API server overload, cascading failures, and an unresponsive control plane.
  2. Strategic Use of Extension Mechanisms: Operators and developers should carefully evaluate whether CRDs are the appropriate extension mechanism for their specific use case. For static configuration, desired state, and moderate update frequencies, CRDs are ideal. However, for scenarios involving:
  • High-frequency updates to a large number of objects.
  • Storing very large amounts of data that exceed etcd's object size limits.
  • Dynamically generated or transient data that doesn't represent persistent state.

The Kubernetes API Aggregation Layer becomes a critical tool. Understanding when to pivot to an aggregated API server, which allows for custom storage backends, is a key architectural defense against performance bottlenecks.

  1. Implement Lazy Data Generation: The "lazy" approach adopted by Kueue for generating workload positions is a powerful defensive pattern. By only computing and exposing data when explicitly requested, systems can significantly reduce their operational overhead and avoid unnecessary resource consumption. This minimizes the attack surface on performance, ensuring resources are only expended when value is actively being derived.
  2. Adopt Queuing Systems for AI/ML Workloads: For environments running AI/ML workloads, deploying a system like Kueue itself is a defensive measure. It prevents uncontrolled admission of resource-intensive jobs, thereby protecting the cluster from saturation and ensuring fair resource allocation. This proactive management prevents a "thundering herd" problem where numerous jobs simultaneously attempt to consume resources, leading to cluster instability.
  3. Design for Scalability from the Outset: The challenges highlighted in the talk are indicative of scaling issues. When designing custom controllers or operators that manage a large number of resources or require frequent state changes, architects should consider these API limitations early in the design phase. Retrofitting solutions after hitting performance walls is significantly more costly and complex.

In essence, the talk advocates for an informed and deliberate approach to Kubernetes extension, treating the API server and etcd as precious, finite resources that must be protected through judicious design choices.

Key Takeaways

  • Kueue is a cloud-native queuing system for AI workloads in Kubernetes, abstracting various job types under a Workload CRD to manage admission and scheduling.
  • Directly updating CRDs for high-frequency, dynamic data (like queue positions) leads to API server and etcd overload, due to QPS limits and individual object size constraints, especially with tens of thousands of workloads.
  • The Kubernetes API Aggregation Layer offers a powerful alternative to CRDs for performance-critical scenarios, allowing custom storage backends (like Kueue's internal heap) instead of etcd.
  • Kueue leverages API Aggregation and a "lazy" evaluation strategy to dynamically generate workload positions only when requested, taking a snapshot of its internal heap and processing it on-demand to ensure high performance.
  • Choosing between CRDs and API Aggregation is a critical architectural decision, balancing ease of setup (CRDs) with the need for high performance and custom storage (API Aggregation) for specific use cases.
  • Implementing "lazy" data generation is a key pattern for building scalable systems in Kubernetes, minimizing computational overhead and protecting the control plane from unnecessary load.

About the Speaker(s)

P. Bundyra is a Google engineer who has been actively involved in the development of Kueue since 2022. Working in collaboration with the open-source community, Bundyra contributes to this cloud-native queuing system designed to manage AI workloads within Kubernetes environments. His expertise lies in addressing the complex scalability and performance challenges associated with orchestrating large-scale machine learning and AI tasks on Kubernetes.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This lightning talk by P. Bundyra presents a highly effective and technically sound solution to a critical Kubernetes scaling problem: efficiently exposing queue positions for tens of thousands of AI workloads without overloading the API server or etcd. By cleverly leveraging the Kubernetes API Aggregation Layer with a "lazy evaluation" approach, Kueue demonstrates how to bypass CRD limitations for high-frequency, dynamic data. This is a must-see for any Kubernetes developer or architect looking to build scalable and resilient operators, offering deep insights into API internals and advanced extension patterns.

Heather Calloway (CISO) — STRONG ACCEPT

This lightning talk on Kueue masterfully dissects a critical Kubernetes scalability challenge: managing high-frequency updates for AI workload queue positions without overwhelming the API server. P. Bundyra clearly demonstrates the practical limits of CRDs for such dynamic data, advocating for the strategic use of the Kubernetes API Aggregation Layer. The "lazy evaluation" approach for generating queue positions is a powerful example of architecting for resilience and performance at scale, directly addressing core operational risks for any organization running significant AI/ML workloads on Kubernetes.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025