The NATS Stack - Libraries Extensions and the Execution Engine - Tomasz Pietrek & Jordan Rash

Tomasz Pietrek, Jordan Rash

KubeCon + CloudNativeCon Europe 2025 · Session

Overview

This talk, presented by Tomasz Pietrek and Jordan Rash, delves into the evolution and capabilities of the NATS messaging system, particularly highlighting its recent 2.11 release and the innovative NATS Execution Engine (Nex). NATS is a high-performance, open-source, distributed messaging system designed for modern cloud-native, edge, and IoT applications. The speakers articulate how NATS aims to simplify the inherent complexity of distributed systems by providing a single, unified technology for communication, persistence, and security across diverse environments.

Watch on YouTube

Visual summary for The NATS Stack - Libraries Extensions and the Execution Engine - Tomasz Pietrek & Jordan Rash by Tomasz Pietrek, Jordan Rash
Visual summary for The NATS Stack - Libraries Extensions and the Execution Engine - Tomasz Pietrek & Jordan Rash by Tomasz Pietrek, Jordan Rash

Key moments

  1. 0:00 Welcome and speaker introductions for NATS talk
  2. 1:50 First mention of the NATS execution engine
  3. 3:20 Discussing the growing complexity of modern applications
  4. 4:40 Identifying complexity from adding many disparate system components
  5. 6:00 Highlighting inherent limitations in traditional distributed system design
  6. 7:20 The goal: simplifying distributed systems for modern needs

The NATS Stack - Libraries Extensions and the Execution Engine

Speakers: Tomasz Pietrek, Senior Software Engineer, Synadia; Jordan Rash, Software Engineer, Synadia

Conference: KubeCon EU

YouTube: https://www.youtube.com/watch?v=MHfDvUUJ14I

Overview

This talk, presented by Tomasz Pietrek and Jordan Rash, delves into the evolution and capabilities of the NATS messaging system, particularly highlighting its recent 2.11 release and the innovative NATS Execution Engine (Nex). NATS is a high-performance, open-source, distributed messaging system designed for modern cloud-native, edge, and IoT applications. The speakers articulate how NATS aims to simplify the inherent complexity of distributed systems by providing a single, unified technology for communication, persistence, and security across diverse environments.

The presentation focuses on how NATS addresses the challenges of building and operating distributed applications, from managing diverse communication patterns to ensuring data persistence and robust security. Jordan Rash, a maintainer of the NATS Execution Engine, particularly showcases how Nex leverages NATS's core features to create a powerful control plane for workloads, demonstrating a practical application of "dogfooding" the NATS platform. The talk is crucial for architects, developers, and operators seeking to streamline their distributed system designs and reduce operational overhead.

Ultimately, Pietrek and Rash argue that NATS offers a compelling alternative to architectures burdened by numerous disparate components. By consolidating essential functionalities into a single, efficient binary, NATS provides a flexible and resilient foundation for applications ranging from large-scale cloud deployments to resource-constrained edge devices, even in environments with intermittent connectivity. The insights shared are particularly relevant for those navigating the complexities of multi-cloud, hybrid, and edge computing paradigms.

Background

▶ Watch: Welcome and speaker introductions for NATS talk (0:00)

The speakers open by contrasting the simplicity of early computing—a single server, one application, one database—with the overwhelming complexity of modern distributed systems. Today's applications demand distribution, resilience, and often operate at the edge with unreliable internet connections. This shift has led to an explosion of specialized components: message queues, load balancers, proxies, databases (SQL, NoSQL, KV stores), observability tools, and persistence layers. While each component is excellent in its own right, the sheer number of technologies to integrate, maintain, and secure creates significant operational overhead. Developers and operators spend a disproportionate amount of time managing infrastructure rather than focusing on core business logic.

Traditional distributed system approaches often suffer from inherent limitations: one-to-one communication patterns, excessive layering of components, centralized and location-dependent architectures, and the steep learning curve associated with multiple technologies. For instance, deploying an application might require a database, a message broker, a cache, a load balancer, and a service mesh—each a separate project with its own API, operational model, and security considerations. This fragmentation makes it difficult to achieve consistent high availability, fault tolerance, and secure communication, especially in dynamic environments like edge computing where connectivity is not guaranteed. NATS was developed to address these challenges by providing a comprehensive, decentralized, and location-agnostic messaging platform that simplifies the entire stack.

Key Findings

▶ Watch: Discussing the growing complexity of modern applications (3:20)

The talk highlights several key findings and advancements within the NATS ecosystem:

  1. NATS 2.11 Release: This major release introduces critical features for managing complex, distributed message flows. These include distributed message tracing to visualize message paths across large topologies, per-message TTLs for fine-grained control over message expiry in streams (addressing previous KV store compaction challenges), consumer pause/unpause for JetStream consumers, and consumer priority groups for advanced load balancing and overflow scenarios. Additionally, batch direct get enables lightweight retrieval of multiple messages without creating a consumer.
  2. Orbit - The Client Extension Framework: Orbit is presented as a significant innovation for NATS client libraries. It allows for the development of extensions to official clients, enabling faster iteration on new APIs, community contributions, and idiomatic language-specific features. This framework facilitates prototyping and testing of new functionalities before they are stabilized and integrated into the core client libraries, ensuring broader utility and stability. An example is the request-many pattern, which was simplified from hundreds of lines of application code to a robust, community-contributed Orbit extension.
  3. NATS Execution Engine (Nex): Nex is introduced not merely as a product, but as a testament to NATS's capabilities—a complex application built entirely on NATS. It demonstrates how NATS's core features (connectivity, JetStream for persistence, and built-in security mechanisms) can eliminate the need for external databases or dedicated security components. Nex acts as a control plane for deploying and managing workloads (e.g., OCI containers, JavaScript functions) across various runtimes, leveraging NATS for internal communication and state management.
  4. Configuration-Based Security: A powerful finding demonstrated in the talk is NATS's ability to enforce multi-tenancy and granular security through server configuration rather than application-specific logic. By using NATS accounts and subject-level permissions, administrators can define who can publish or subscribe to which subjects, effectively creating secure namespaces. This offloads significant security responsibilities from application developers, making security enforcement more efficient and less prone to application-level bugs.

Technical Deep Dive

▶ Watch: Identifying complexity from adding many disparate system components (4:40)

NATS is designed around a core philosophy of simplicity, performance, and resilience. At its heart, NATS is a single, static Go binary with no external dependencies, making it incredibly lightweight and easy to deploy across diverse environments—from large cloud clusters to tiny edge devices. This single binary encapsulates all NATS functionalities, including core messaging (pub/sub, request/reply, queueing) and JetStream, which provides persistence through streams, key-value (KV) stores, and object stores.

Communication Patterns: NATS inherently supports a wide array of communication patterns. Services are discoverable, allowing for flexible interactions. Beyond simple one-to-one messaging, NATS facilitates request-many patterns, where a single request can solicit responses from multiple services of different types. This eliminates the need for applications to manage individual messages to various endpoints. NATS also supports decentralized architectures through accounts, allowing infrastructure providers to grant tenants access to NATS resources with defined limits without needing to manage their internal users or access policies. Intelligent routing is a core capability, enabling data to remain at the edge for local inferencing or processing, rather than always being sent to the cloud, optimizing for latency and connectivity challenges.

NATS 2.11 Features:

  • Distributed Message Tracing: This feature is crucial for debugging and understanding message flow in complex, distributed NATS deployments. It allows operators to visualize how a message traverses through a supercluster—a network of interconnected NATS clusters—across different gateways, clouds (e.g., AWS, Azure, GCP), and regions. This provides invaluable insights for troubleshooting communication-based issues.
  • Per-Message TTLs: Addressing a long-standing challenge with JetStream KV stores, per-message TTLs allow individual messages or key-value entries to have their own expiration times. Previously, only an entire stream or KV bucket could have a maximum age, which could lead to performance degradation during compaction if many items needed to be purged. With per-message TTLs, data can be automatically removed after a specified duration, minimizing the need for manual compaction and helping to keep the storage "snappy." This also has implications for compliance and data lifecycle management.
  • Consumer Pause/Unpause: This administrative feature for JetStream consumers allows operators to temporarily halt or resume message consumption without requiring changes to the application code. This is particularly useful for maintenance, scaling operations, or managing backpressure, providing greater operational control.
  • Consumer Priority Groups: A forward-looking feature, priority groups enable advanced load balancing and overflow management. In scenarios with multiple clusters across different geographies or availability zones, NATS can intelligently route requests. If local services are overloaded, requests can be overflowed to a different zone or cloud, ensuring continuous service and optimal resource utilization.
  • Batch Direct Get: This new API allows clients to retrieve multiple messages from a JetStream stream or KV store in a single, lightweight API call, without the overhead of creating a consumer. This is ideal for scenarios where an application needs to fetch a batch of historical data or a range of key-value entries efficiently.

Orbit Framework: Orbit is a pivotal development for NATS client libraries. It serves as a testing ground for client extensions, allowing the NATS community and Synadia to rapidly iterate on new APIs and patterns. The goal is to achieve feature parity across all official NATS client libraries (Go, Rust, JavaScript, etc.) while respecting the idiomatic nature of each language. Contributions can be accepted into Orbit extensions and refined. Once an API is finalized and proven stable and useful across multiple clients, it can then be pulled into the official client library, guaranteeing stability and long-term support. A prime example is the request-many pattern, which was initially implemented with hundreds of lines of custom code and prone to race conditions. Its integration as an Orbit extension significantly simplified the client-side implementation and resolved underlying bugs.

Supercluster Architectures: NATS's unique ability to operate in server mode and leaf node mode allows for the creation of highly flexible and resilient superclusters. A supercluster can seamlessly connect multiple NATS clusters across different environments—on-premises data centers, various cloud providers (GCP, Azure, AWS), and geographically dispersed edge locations. This architecture provides inherent resiliency, disaster recovery, and the ability to route messages intelligently based on proximity or load. Examples include retail storefronts running small local clusters that can offload data to larger cloud clusters when connectivity is available, or complex AI at the edge scenarios where inference, model generation, and data gathering occur in real-time across multiple clouds and edge devices over a unified NATS subject bus.

NATS Execution Engine (Nex): Nex is a prime example of building a product on NATS, leveraging its full stack. It uses NATS for all its internal communication, state management, and even security. Instead of integrating external databases or complex authentication systems, Nex utilizes JetStream for persistence (e.g., storing workload definitions) and NATS's subject-level security and auth callouts for multi-tenancy and access control. Nex aims to be a control plane for artifacts, such as OCI containers, enabling users to deploy workloads without concern for the underlying runtime. It can intelligently place workloads on Kubernetes, ECS, or other platforms based on cost, availability, or specific requirements, abstracting away the operational complexities of diverse runtimes. The concept of "trigger functions" or "functions as a service" over NATS subjects, similar to serverless lambdas, is also being developed within Nex.

Demo / Proof of Concept

▶ Watch: Highlighting inherent limitations in traditional distributed system design (6:00)

Jordan Rash presented a compelling live demonstration showcasing NATS's configuration-based security and multi-tenancy capabilities using the NATS Execution Engine (Nex).

The demo began with a bare NATS server running with default settings and Nex deployed. A simple counter workload was initiated. The initial state revealed a lack of multi-tenancy:

  1. Workload Visibility: When User 1 deployed a workload, User 2 could also list and see User 1's workload. This demonstrated that without specific security configurations, all users operating on the default NATS setup had visibility into all deployed workloads.
  2. Admin Functionality: Even administrative functions (e.g., listing NATS nodes) were accessible to any user, highlighting the absence of role-based access control.

To address these security deficiencies without modifying application code, Jordan introduced a NATS server configuration file. This configuration defined:

  1. Three distinct NATS accounts: An admin account, user1 account, and user2 account.
  2. Subject-level security rules: For control subjects (e.g., those used for listing workloads or managing nodes), the configuration required that the account making the request must match the account embedded in the request's credentials. This effectively created isolated namespaces for each user account. For instance, a request to _nex.workloads.list would only be honored if the requesting user's credential belonged to the account whose ID matched the target account ID in the subject.

After applying this configuration and restarting Nex, the demonstration showed a dramatic improvement in security:

  1. Enforced Multi-tenancy: When User 1 listed their workloads, they could see them. However, when User 2 attempted to list workloads, the NATS server immediately denied the request. The server's subject mapping rules prevented the request from even reaching the Nex application, providing rapid and efficient security enforcement.
  2. Restricted Admin Access: Attempts by non-admin users to execute administrative NATS commands were also immediately rejected by the server, demonstrating effective role-based access control without any custom logic in Nex itself.

The key takeaway from this demo was the power of configuration-based security in NATS. By leveraging NATS server features like accounts and subject permissions, developers can implement robust multi-tenancy and access control policies directly at the messaging layer. This significantly reduces the burden on application developers to implement complex security logic, making applications inherently more secure and easier to maintain. The server's immediate rejection of unauthorized requests also means that potentially malicious or erroneous requests are stopped at the earliest possible point, enhancing system resilience and performance.

Defensive Implications

▶ Watch: The goal: simplifying distributed systems for modern needs (7:20)

The NATS stack, with its recent enhancements and the NATS Execution Engine, offers several critical defensive implications for securing distributed systems:

  1. Centralized, Configuration-Driven Security: NATS allows security policies (multi-tenancy, access control, subject permissions) to be defined and enforced at the NATS server level rather than within each application. This reduces the attack surface by moving security logic out of potentially buggy application code into a highly optimized and tested core system. Defenders can establish granular control over who can publish or subscribe to specific subjects, effectively segmenting their messaging fabric and preventing unauthorized data access or command execution.
  2. Enhanced Visibility with Distributed Tracing: The new distributed message tracing feature in NATS 2.11 provides invaluable capabilities for security monitoring and incident response. Security teams can visualize the entire path of a message across complex superclusters, identify unexpected message flows, detect potential data exfiltration attempts, or pinpoint the source of anomalous behavior. This enhanced visibility is crucial for understanding the blast radius of an incident and accelerating root cause analysis.
  3. Resilience in Disconnected/Edge Environments: NATS's ability to operate robustly at the edge, even with intermittent connectivity, provides a secure communication backbone for IoT and critical infrastructure. Edge devices can securely communicate and persist data locally using JetStream, reducing reliance on potentially insecure or unavailable public internet connections. The supercluster architecture allows for secure, controlled data synchronization when connectivity is restored, minimizing exposure.
  4. Simplified Architecture, Reduced Attack Surface: By consolidating messaging, persistence (KV, Object Store), and some security functionalities into a single, lightweight NATS binary, organizations can significantly simplify their distributed system architecture. Fewer moving parts mean a smaller overall attack surface and reduced complexity in securing and auditing the system. This also translates to less operational burden for security teams.
  5. Secure Workload Orchestration with Nex: When using the NATS Execution Engine (Nex), security policies for workload deployment and execution can be centralized. Nex leverages NATS's inherent security features to ensure that workloads are deployed and run within their authorized contexts. This provides a secure control plane for dynamic application deployment, ensuring that only trusted artifacts are executed and that they adhere to defined security boundaries.
  6. Data Lifecycle Management with Per-Message TTLs: The introduction of per-message TTLs in NATS 2.11 is a significant win for data governance and compliance. Defenders can now ensure that sensitive data stored in NATS streams or KV stores automatically expires after a defined period, preventing unnecessary data retention and reducing the risk of data at rest. This helps in meeting regulatory requirements and minimizing the impact of data breaches by limiting the exposure window.

Key Takeaways

  • Simplified Distributed Systems: NATS aims to drastically simplify complex distributed architectures by offering a single, unified technology for communication, persistence, and security, reducing reliance on multiple disparate components.
  • NATS 2.11 Innovations: The latest release introduces powerful features like distributed message tracing, per-message TTLs, consumer pause/unpause, and batch direct get, enhancing observability, data lifecycle management, and operational control.
  • Orbit for Client-Side Agility: The Orbit framework fosters rapid innovation and community contributions for NATS client libraries, allowing for faster iteration on new APIs and language-specific extensions before official integration.
  • NATS as an Application Platform: The NATS Execution Engine (Nex) demonstrates NATS's capability as a robust foundation for building complex applications, leveraging its core features for connectivity, persistence, and security without external dependencies like databases.
  • Powerful Configuration-Based Security: NATS enables robust multi-tenancy and granular access control through server configuration (accounts, subject-level permissions), offloading security logic from applications and enforcing it efficiently at the messaging layer.
  • Flexible Supercluster Architectures: NATS supports highly resilient and adaptable "supercluster" deployments, seamlessly connecting systems across on-premises, multiple clouds, and diverse edge environments, even with intermittent connectivity.

About the Speaker(s)

Tomasz Pietrek is a Senior Software Engineer at Synadia and a dedicated NATS Project Maintainer. His deep involvement in the NATS project is evident through his contributions to the server over several years, even prior to joining Synadia approximately nine months ago. Tomasz is also recognized as a NATS Ambassador, a role he enthusiastically embraced prior to the last KubeCon. With a background that includes experience as a software engineer and concurrently as a university teacher, he brings a blend of practical development expertise and a knack for clear explanation to his presentations.

Jordan Rash is a Software Engineer at Synadia, where his primary role involves being a maintainer on the NATS Execution Engine (Nex). Jordan brings a unique perspective to the NATS community, having previously worked in cybersecurity and defense fields. This background has informed his approach to leveraging NATS's inherent security features to build robust and secure distributed systems. He played a key role in developing Nex, which serves as a prime example of building complex applications entirely on the NATS stack.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk effectively showcases the evolution and capabilities of the NATS messaging system, particularly highlighting the significant advancements in the 2.11 release and the innovative NATS Execution Engine (Nex). It expertly navigates the complexities of distributed systems, presenting NATS as a unified, high-performance solution for communication, persistence, and security. The live demo of configuration-based multi-tenancy and the discussion of features like distributed tracing and per-message TTLs provide concrete, actionable insights for architects and security professionals, demonstrating how to build more resilient and secure systems with reduced operational overhead. This is a…

Heather Calloway (CISO) — STRONG ACCEPT

This talk presents a compelling case for NATS as a foundational technology that simplifies distributed system architecture while significantly enhancing security and operational resilience. The emphasis on configuration-based security, distributed tracing, and per-message TTLs directly addresses critical governance challenges and reduces real-world business exposure. While it is a vendor presentation, the speakers effectively translate technical capabilities into clear, actionable benefits for security leaders grappling with complex, multi-cloud, and edge environments.

→ Top-rated talks at KubeCon + CloudNativeCon Europe 2025

All talks from KubeCon + CloudNativeCon Europe 2025