Cortex: Insights, U... Friedrich Gonzalez, Daniel Sabsay, Charlie Le, Alolita Sharma & Daniel Blando
Friedrich Gonzalez, Daniel Sabsay, Charlie Le, Alolita Sharma, Daniel Blando
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This KubeCon EU session provided a comprehensive update on Cortex, the horizontally scalable, highly available, multi-tenant long-term storage solution for Prometheus. The talk, delivered by a team of maintainers from Apple, AWS, and Adobe, delved into recent feature rollouts, significant performance enhancements, and the strategic roadmap for the project, including its journey towards CNCF graduation. For organizations grappling with the challenges of managing Prometheus metrics at scale, Cortex offers a robust and flexible architecture designed to ensure data isolation, resilience, and efficient querying across diverse environments.

Key moments
- 0:00 Welcome, speakers, and talk agenda overview
- 2:20 Defining Cortex: scalable, multi-tenant Prometheus long-term storage
- 3:20 Deep dive into Cortex architecture diagram
- 4:45 Key differences: Cortex's scalability, stability, and community
- 6:10 How to get started with Cortex quickly
- 7:00 Live demo: Cortex in microservices mode on Kubernetes
Cortex: Insights, U... Friedrich Gonzalez, Daniel Sabsay, Charlie Le, Alolita Sharma & Daniel Blando
Speakers: Friedrich Gonzalez, Daniel Sabsay, Charlie Le, Alolita Sharma, Daniel Blando
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=3aUg2qxfoZU
Overview
This KubeCon EU session provided a comprehensive update on Cortex, the horizontally scalable, highly available, multi-tenant long-term storage solution for Prometheus. The talk, delivered by a team of maintainers from Apple, AWS, and Adobe, delved into recent feature rollouts, significant performance enhancements, and the strategic roadmap for the project, including its journey towards CNCF graduation. For organizations grappling with the challenges of managing Prometheus metrics at scale, Cortex offers a robust and flexible architecture designed to ensure data isolation, resilience, and efficient querying across diverse environments.
The speakers highlighted Cortex's evolution, emphasizing its commitment to stability, performance, and community-driven development. Attendees were treated to a live demonstration showcasing Cortex's multi-tenancy capabilities and dynamic scaling features. The session also offered a forward-looking perspective on upcoming developments, such as enhanced OTLP compatibility and support for Prometheus 3.0, positioning Cortex as a critical component in modern observability stacks.
This presentation is particularly relevant for SREs, platform engineers, and DevOps professionals who manage large-scale monitoring infrastructures and require a reliable, performant, and cost-effective solution for long-term Prometheus metric storage and multi-tenancy. The insights shared underscore Cortex's maturity and its ongoing efforts to address the complex demands of cloud-native monitoring.
Background
▶ Watch: Welcome, speakers, and talk agenda overview (0:00)
Cortex emerged to address critical limitations in Prometheus deployments, primarily the need for long-term storage, horizontal scalability, and multi-tenancy. While Prometheus excels at local, short-term metric collection and querying, it lacks inherent capabilities for aggregating metrics from numerous instances, storing them over extended periods, or isolating data between different teams or departments within a single infrastructure. Cortex fills this gap by acting as a remote write endpoint for Prometheus, transforming it into a robust, distributed monitoring system.
At its core, Cortex is designed for horizontal scalability, meaning its components can be independently scaled out to handle increasing load. It prioritizes low response times for both writes and queries, a crucial factor for real-time observability. The project also boasts strong stability guarantees, with new features meticulously marked as experimental before becoming stable, ensuring users are not caught off guard by changes. As a CNCF project, Cortex is community-managed, backed by a diverse group of contributors from multiple companies, fostering a vibrant and healthy development ecosystem.
The fundamental architecture of Cortex involves several distinct microservices. The Distributor acts as the ingress point, handling incoming Prometheus remote write requests, rate limiting, and fanning out data. Ingestors temporarily store metrics in memory, compact them into blocks, and then upload these blocks to object storage (like S3, GCS, Azure Blob Storage, or OpenStack Swift) for long-term persistence. The Querier and Querier-frontend components are responsible for executing PromQL queries, leveraging caching and query splitting to optimize performance. This modular design allows Cortex to provide a highly available and resilient solution for managing metric data at enterprise scale.
Key Findings
▶ Watch: Deep dive into Cortex architecture diagram (3:20)
The KubeCon EU talk highlighted a significant stride in Cortex's development with the release of version 1.19, alongside a roadmap of upcoming features and the project's journey towards CNCF graduation. The key findings and contributions presented can be categorized into several areas:
1. Enhanced Stability and Availability:
- OpenStack Swift support for object storage is no longer experimental, solidifying its position as an officially supported backend alongside Amazon S3, Google Cloud Storage, and Microsoft Azure Storage.
- Cortex Ruler High Availability (HA) moved from experimental to a primary-versus-non-primary approach, minimizing gaps during restarts and failures and ensuring rules are evaluated by only one ruler at a time. This HA is also availability zone aware.
- The Cortex ruler now supports using the existing query path, leveraging the caching and query splitting benefits of the query front-end.
2. Performance and Resource Optimization:
- The Partition Compactor, previously discussed, helps overcome the 64GB limitation of the TSDB index and speeds up compaction for larger blocks.
- Significant ingestion optimizations focusing on CPU and memory were introduced:
- Push Workers (experimental in 1.19): A worker pool for managing Go routines to reduce the overhead of creating new stacks, yielding up to 20% CPU improvement on ingestors in certain use cases.
- Expanding Posting Cache: Caches query postings instead of full results, particularly beneficial for complex regex queries, resulting in around 20% CPU improvement in some scenarios.
- Improvements to Ingestor scaling down with a new
read-onlystate and an API to fetch statistics, allowing for safer and more controlled scaling operations without data loss.
3. Interoperability and Compatibility:
- Continued advancements in OTLP (OpenTelemetry Protocol) compatibility, including:
- Adding a max request size limit to prevent out-of-memory issues.
- Including target info metrics by default for consistency with Prometheus OTLP.
- Implementing the
promote resource attributesconfiguration, allowing users to add custom labels for querying or aggregation. - Experimental flag for supporting mixed HA pairs in a single remote write request, simplifying configurations for users with multiple HA Prometheus setups.
4. Upcoming Features and Roadmap:
- Future plans include OTLP metadata conversion for even broader compatibility.
- Further ingestion CPU optimizations through stream connections with keep-alives between distributors and ingestors, aiming for another 10% CPU improvement.
- Support for native histograms including custom buckets.
- Active discussions and work on Prometheus 3.0 compatibility, specifically the remote write v2 protocol, which introduces new headers and internal string interning for CPU performance.
- A clear roadmap towards CNCF graduation, involving a third-party security review and documenting the roadmap change process for more inclusive governance and release management.
These findings collectively demonstrate Cortex's ongoing commitment to delivering a high-performance, resilient, and adaptable solution for large-scale Prometheus monitoring, with a strong emphasis on operational efficiency and developer experience.
Technical Deep Dive
▶ Watch: Key differences: Cortex's scalability, stability, and community (4:45)
Cortex's architecture is a testament to its design principles of horizontal scalability, high availability, and multi-tenancy. At a high level, it separates the write path and read path, allowing independent scaling and optimization of each.
The write path begins with Prometheus instances or OpenTelemetry agents sending metrics to Cortex. These requests first hit the Distributor component. The Distributor is stateless and handles critical functions like rate limiting for tenants, deduplication of samples from highly available Prometheus pairs, and fanning out incoming requests to multiple Ingestor instances. This fanning-out mechanism ensures that data is replicated across multiple ingestors, providing high availability and fault tolerance. Each Ingestor stores a subset of the incoming metrics in memory for approximately two hours, similar to Prometheus's head block. Periodically, or when its memory buffer is full, the Ingestor compacts these in-memory metrics into TSDB blocks and uploads them to a chosen object storage backend (Amazon S3, Google Cloud Storage, Microsoft Azure Storage, or OpenStack Swift). This offloading mechanism allows Cortex to provide long-term storage without being constrained by the local disk capacity of individual nodes.
The read path typically starts with dashboards like Grafana querying Cortex. Queries are routed through the Querier-frontend, which can perform optimizations like query splitting (breaking a long time-range query into smaller, parallel sub-queries) and caching of query results. The Querier-frontend then dispatches these (or the original) queries to Querier instances. Queriers are responsible for fetching data from both the object storage (for historical data) and active Ingestors (for recent data still in memory). A crucial aspect of Cortex's multi-tenancy is that each query includes a tenant header, allowing the Querier to retrieve only the metrics belonging to the specified tenant, ensuring data isolation.
A significant enhancement in version 1.19 is the Cortex Ruler's High Availability (HA). The Ruler component evaluates Prometheus recording and alerting rules. Previously, ensuring its HA was challenging. The new approach implements a primary-versus-non-primary model. Each rule group is assigned to a set of ruler instances, but only one instance in that set (the primary) actively evaluates the rules. Non-primary instances continuously monitor the primary's health. If the primary becomes unhealthy, a non-primary instance will automatically take ownership of those rules and begin evaluation. This system minimizes gaps in rule evaluation during failures and prevents redundant evaluations. Furthermore, this HA mechanism is availability zone aware, enhancing resilience against regional outages. The ruler can also now utilize the existing query path through the Querier-frontend, benefiting from its caching and query splitting capabilities, which can significantly improve the performance of complex or frequently evaluated rules.
For ingestion optimizations, two key features stand out. Push Workers (experimental in 1.19) address CPU overhead caused by frequent Go routine creation. Instead of creating a new Go routine for every task, a worker pool reuses existing Go routines, reducing the burden of "new stacks" and leading to CPU improvements of up to 20% on ingestors. The Expanding Posting Cache aims to optimize query CPU usage, especially for queries involving complex regular expressions. Instead of caching the full result of a query, it caches the "postings" – the list of series matching a label selector. This is more complex for ingestors due to their mutable head block, so the cache is invalidated whenever a new sample arrives that matches a cached posting. This can still yield around 20% CPU improvement for specific query patterns.
Ingestor scaling down has been made safer through a new read-only state. When an ingestor needs to be scaled down, it can be set to "read-only." In this state, it remains active on the ring and visible to queriers, continuing to serve existing data, but it stops receiving new write requests. Additionally, a new API allows querying the ingestor for statistics, such as whether it still holds any loaded blocks. Once an ingestor is in read-only mode and reports no active blocks, it can be safely terminated without risking data loss or query failures.
Finally, Cortex is actively pursuing OTLP compatibility, including adding a max request size limit to prevent out-of-memory errors and the promote resource attributes configuration for adding custom labels to metrics. Upcoming work includes OTLP metadata conversion and further CPU optimizations for ingestion via stream connections with keep-alives between distributors and ingestors, aiming for another 10% CPU reduction. The project is also preparing for Prometheus 3.0 by implementing remote write v2 headers and leveraging string interning for improved CPU performance.
Demo / Proof of Concept
▶ Watch: How to get started with Cortex quickly (6:10)
Charlie Le provided a compelling live demonstration of Cortex's multi-tenancy and dynamic scaling features, running on a local Kubernetes cluster. The setup illustrated a practical deployment of Cortex in microservices mode, highlighting its flexibility and operational capabilities.
The demo environment consisted of:
- An empty Kubernetes cluster.
- Cortex deployed in microservices mode, including a Distributor (one replica), three Ingestor replicas, a Querier, and an Nginx instance acting as an HTTP router.
- Two Prometheus instances configured to scrape metrics within the cluster and send them to Cortex, each representing a distinct tenant:
devandstaging. This setup showcased how Cortex provides isolated data streams for different teams or environments. - A Grafana instance, also running in the cluster, used to visualize the incoming metrics and demonstrate the effects of dynamic configuration changes.
Charlie first demonstrated dynamic rate limiting for a specific tenant. Initially, both the dev and staging tenants had an ingestion limit of 25,000 samples per second. By modifying the Helm configuration file to change the dev tenant's limit to 2,000 samples per second and applying the change, Charlie showed that none of the Cortex pods restarted. This was possible because the rate limit configurations are runtime values that Cortex services are programmed to reload periodically (demonstrated with a 1-second reload interval for faster feedback, compared to the default 10 seconds). The Grafana dashboard immediately reflected this change, showing the dev tenant's ingestion rate dropping to 2,000 samples per second, while the staging tenant continued at the higher rate. This effectively demonstrated Cortex's ability to enforce tenant-specific resource controls without service interruption.
The second part of the demo focused on automatic ingestor scaling. Charlie increased the ingestor pod replica count from three to four by updating the Helm file. As expected, a new ingestor pod quickly spun up in the Kubernetes cluster. The Grafana dashboard, which was monitoring the ingestion rate per ingestor pod, clearly showed the new fourth ingestor gradually taking on a share of the incoming metric samples. Concurrently, the ingestion rates on the original three ingestors decreased, illustrating Cortex's self-balancing and load distribution capabilities as new resources become available. This seamless scaling operation underscores Cortex's elasticity, allowing operators to dynamically adjust resources based on demand without manual intervention or service downtime.
The demonstration effectively conveyed the practical benefits of Cortex's multi-tenancy, dynamic configuration management, and robust horizontal scaling, making these complex features tangible for the audience.
Defensive Implications
▶ Watch: Live demo: Cortex in microservices mode on Kubernetes (7:00)
The advancements in Cortex presented at KubeCon EU offer several critical implications for defenders and platform operators looking to build robust and resilient monitoring infrastructures.
Firstly, the robust multi-tenancy capabilities are a foundational defensive measure. By providing strong isolation between dev and staging (or other) tenants, Cortex helps prevent data leakage and ensures that one tenant's activities or misconfigurations do not impact another. Defenders should leverage this feature to enforce strict access controls and data segregation policies within their observability stacks. The ability to dynamically rate limit tenants without restarting services provides an immediate control mechanism to mitigate denial-of-service attacks or runaway metric exports from a misbehaving application, preventing resource exhaustion across the entire monitoring system.
The introduction of High Availability (HA) for the Cortex Ruler is a significant boost for resilience. Alerting and recording rules are critical for identifying security incidents and maintaining system health. The primary-versus-non-primary approach, coupled with availability zone awareness, ensures that rule evaluation continues uninterrupted even during component failures or regional outages. Defenders should actively implement and configure this HA feature to ensure their security alerts and operational insights remain available and accurate at all times. The option for the ruler to use the existing query path also means that security-critical rules can benefit from the performance and caching optimizations of the query front-end, ensuring timely evaluation.
Performance optimizations, such as Push Workers and the Expanding Posting Cache, are not just about cost savings; they contribute to the overall stability and responsiveness of the monitoring system. A more efficient Cortex can handle higher loads, reducing the risk of metric backlogs or query timeouts during peak events, which could otherwise mask ongoing attacks or critical system failures. Defenders should explore enabling these experimental features in controlled environments to assess their impact on their specific workloads.
Improvements in Ingestor scaling down procedures, utilizing a read-only state and statistics API, directly address data integrity and availability during operational changes. Safely scaling down ingestors without data loss or query gaps is crucial. Defenders must integrate these new mechanisms into their automated scaling and deployment pipelines to ensure that infrastructure changes do not inadvertently compromise monitoring data.
Finally, enhanced OTLP compatibility and future support for Prometheus 3.0 position Cortex as a future-proof component in a diverse observability ecosystem. This allows organizations to integrate Cortex more seamlessly with other security tools and data sources that adhere to OpenTelemetry standards, fostering a more comprehensive security posture through unified telemetry. Defenders should advocate for and plan to adopt these standards to streamline data collection and analysis for security purposes.
In summary, Cortex's continuous development focuses on features that directly enhance the resilience, performance, and operational security of large-scale monitoring systems. Defenders should actively engage with these updates to strengthen their observability infrastructure against both operational challenges and potential security threats.
Key Takeaways
- Cortex continues to evolve as a robust, horizontally scalable, and multi-tenant long-term storage solution for Prometheus, actively addressing enterprise-scale monitoring challenges.
- The Cortex Ruler now boasts enhanced High Availability (HA) with a primary-versus-non-primary, availability zone-aware architecture, significantly improving the resilience of alerting and recording rules.
- Version 1.19 introduces key ingestion CPU optimizations like "Push Workers" and an "Expanding Posting Cache," potentially yielding up to 20% CPU improvements on ingestors and for complex query patterns.
- Operational efficiency for ingestor scaling down has been improved with a new
read-onlystate and statistics API, ensuring safer and more controlled resource adjustments without data loss. - Cortex is expanding its OTLP compatibility and actively working towards supporting Prometheus 3.0 (specifically remote write v2) and native histograms, ensuring future-readiness and broader integration.
- The project is on a clear roadmap towards CNCF graduation, signifying its maturity, stability, and commitment to open governance and community-driven development.
About the Speaker(s)
The talk was delivered by a collaborative team of maintainers representing leading technology companies, underscoring the broad industry support for Cortex:
- Alolita Sharma (Apple): Introduced the session and provided an overview of Cortex, its purpose, and the agenda for the talk.
- Charlie Le (Apple): Demonstrated Cortex's multi-tenancy and dynamic scaling capabilities through a live demo, showcasing runtime configuration changes and automatic ingestor scaling.
- Daniel Sabsay (Adobe): Provided updates on the Cortex community, including new contributors and triagers, and detailed new features in release 1.19, such as Cortex Ruler HA and OpenStack Swift support.
- Daniel Blando (AWS): Covered further enhancements in version 1.19 and upcoming features, focusing on OTLP compatibility, ingestion optimizations, and improved ingestor scaling down mechanisms.
- Friedrich Gonzalez (Adobe): Shared personal insights into his journey with Cortex, from user to contributor to maintainer, and outlined the project's roadmap towards CNCF graduation, emphasizing security review and governance improvements.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This KubeCon EU session on Cortex 1.19 and its roadmap was a substantive technical update, delivered by the project's maintainers. It detailed critical advancements in horizontal scalability, high availability for the ruler, and significant ingestion performance optimizations. The live demo cleanly showcased multi-tenancy and dynamic scaling without fluff. For anyone operating large-scale Prometheus deployments, this provided actionable insights and a clear signal on the future direction of a vital component in the observability stack, demonstrating real engineering effort.
Heather Calloway (CISO) — STRONG ACCEPT
This KubeCon session on Cortex, while technical, presents critical advancements that directly enhance the resilience and security posture of large-scale monitoring infrastructures. The focus on multi-tenancy, high availability for alerting rules, dynamic rate limiting, and safe scaling operations translates directly into reduced operational and governance risk. It provides platform engineers and security leaders with actionable insights on how to build a more robust and accountable observability stack, which is fundamental to effective incident response and regulatory compliance.