The Great Sidecar Debate - William Morgan, Buoyant
William Morgan, Buoyant
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
In "The Great Sidecar Debate," William Morgan, CEO of Buoyant and a driving force behind the Linkerd service mesh, presents a thought-provoking analysis of service mesh architectures, challenging attendees to move beyond hype and understand the fundamental engineering trade-offs involved. The talk, originally titled "Why KubeCon attendees need to take the time to understand a bunch of complicated engineering trade-offs instead of blindly jumping onto the new thing," delves into the strengths and weaknesses of the sidecar pattern and its alternatives, particularly in the context of service meshes.

Key moments
- 0:00 Introduction and the talk's true purpose
- 2:00 Engineering decisions are about trade-offs, not right or wrong
- 2:30 Defining the sidecar pattern in Kubernetes
- 4:30 Key benefits and advantages of the sidecar pattern
- 6:30 The downsides and challenges of using sidecars
The Great Sidecar Debate
Speakers: William Morgan, CEO, Buoyant
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=lVWUCUt6ZM8
Overview
In "The Great Sidecar Debate," William Morgan, CEO of Buoyant and a driving force behind the Linkerd service mesh, presents a thought-provoking analysis of service mesh architectures, challenging attendees to move beyond hype and understand the fundamental engineering trade-offs involved. The talk, originally titled "Why KubeCon attendees need to take the time to understand a bunch of complicated engineering trade-offs instead of blindly jumping onto the new thing," delves into the strengths and weaknesses of the sidecar pattern and its alternatives, particularly in the context of service meshes.
Morgan's presentation is a deep dive into the evolution of container orchestration, the operational complexities inherent in distributed systems, and the strategic decisions that shape modern cloud-native infrastructure. He dissects the historical challenges associated with sidecars in Kubernetes, highlights recent advancements that have mitigated many of these issues, and critically examines emerging "sidecar-less" or "ambient" approaches. The core message emphasizes that engineering choices are rarely black and white, but rather a balancing act between various factors like performance, operational complexity, and security, urging practitioners to make informed decisions tailored to their specific use cases.
Background
▶ Watch: Introduction and the talk's true purpose (0:00)
The sidecar pattern has been a cornerstone of cloud-native application development for nearly a decade, with discussions dating back to 2015 within the Kubernetes community. At its core, a sidecar involves running an auxiliary container alongside a main application container within the same Kubernetes pod. These co-located containers share the same network namespace, cgroups, and can share file systems via volume mounts, yet remain isolated as distinct processes. This pattern offers a highly convenient way to augment application functionality—such as logging, monitoring, or, most notably, service mesh capabilities—without requiring modifications or recompilation of the application itself. It promotes language and framework independence and allows platform teams to own and manage these cross-cutting concerns, abstracting them away from application developers. Furthermore, sidecars benefit from a clear operational model, inheriting the pod's lifecycle, and a defined security model, often acting as a micro-firewall for the application traffic. The pattern also fully leverages Kubernetes' extensive investment in container isolation.
Despite these advantages, sidecars historically presented several operational "warts," which Morgan humorously likens to a spectrum from "pneumonia to warts." Early challenges included the resource footprint, exemplified by Linkerd 1.x's JVM-based proxy, which could be as large as 150MB, dwarfing smaller Go microservices. This "side truck" effect raised concerns about resource consumption. Another significant issue was the immutability of pods; any change to the sidecar necessitated a pod restart, potentially disrupting applications not designed for frequent restarts. For Kubernetes jobs, synchronizing termination between the main application and the sidecar was notoriously difficult, often leading to jobs hanging indefinitely. Startup ordering race conditions, where the application might attempt network communication before the sidecar proxy was ready, also plagued early implementations. Finally, a more psychological barrier was the visibility of the "network" component within the application pod, making its resource usage explicitly visible, unlike a deeply embedded library.
The Kubernetes community recognized these challenges, leading to a multi-year effort to formally integrate sidecar support into the platform. A significant Kubernetes Enhancement Proposal (KEP), first proposed in 2019, aimed to address these warts. This KEP eventually culminated in the merge of new functionality in late 2022 and 2023, introducing native sidecar containers. Previously, Kubernetes only distinguished between init containers (which run sequentially to completion) and regular containers (which run in parallel and continuously). The key innovation was the introduction of a restartPolicy: Always flag for init containers. This seemingly minor addition transformed init containers, allowing them to: 1) no longer block the startup of subsequent containers (effectively running in parallel with other initContainers that also have restartPolicy: Always and regular containers), and 2) automatically restart if they die. Crucially, it also introduced new shutdown behavior, where these "always-restarting init containers" would automatically terminate after the main application container exits. This advancement effectively resolved the long-standing issues of guaranteed sidecar initialization before the application, predictable sidecar termination after the application, and robust lifecycle management, turning many of the historical "warts" into resolved problems.
Key Findings
▶ Watch: Engineering decisions are about trade-offs, not right or wrong (2:00)
The talk's central finding is that while Kubernetes' native sidecar support has significantly matured, resolving many operational complexities, the fundamental debate in service mesh architectures now revolves around the placement and sharing model of the L7 proxy. William Morgan asserts that for any complex service mesh feature involving HTTP, HTTP/2, or gRPC traffic, an L7 proxy is indispensable. The core question then becomes not if to use a proxy, but where to deploy it.
Morgan identifies three primary architectural approaches for deploying these L7 proxies in a service mesh:
- Sidecars: One proxy per application pod, providing maximum isolation. Examples include Linkerd 2.x, Kuma, and Istio's traditional sidecar mode.
- Node Proxies: One proxy per node, shared by all application pods on that node. This was the model for Linkerd 1.x.
- Ambient (e.g., Istio Ambient): A hybrid model that splits functionality. An L4 proxy (Z-tunnel) is deployed per node for mTLS and basic TCP handling, while L7 proxies (Waypoint Proxies) are deployed separately and can be configured for varying levels of sharing (e.g., per namespace, per service account, or even per pod).
The critical distinction, according to Morgan, lies in the use of shared proxies. Sidecars inherently offer no sharing between pods, providing strong isolation. Node proxies and Ambient, by design, involve sharing the proxy resource across multiple, potentially uncooperative applications or "tenants." This leads to the most significant finding: multi-tenancy and resource contention are the primary challenges introduced by shared proxy models, especially for L7 traffic.
Morgan emphasizes that enforcing fairness for L7 request processing is qualitatively different and significantly harder than for L4 traffic. While operating systems can manage L4 fairness (e.g., network bandwidth) using event-level decisions and queuing disciplines, L7 processing involves unbounded work (parsing, routing, policy enforcement) akin to general application execution. Modern L7 proxies are not designed to enforce fairness between arbitrary clients, leading to concrete problems like noisy neighbor issues, where one application can starve others of proxy resources. Shared proxies also increase the single point of failure blast radius and make resource sizing and limit allocation considerably more complex for platform teams.
Despite the emergence of "sidecar-less" solutions like Istio Ambient, Morgan's conclusion, based on Linkerd's internal evaluations and benchmarks, is that the sidecar model continues to offer the most favorable trade-offs for their users. Buoyant's benchmarks indicate that Linkerd's sidecar approach remains smaller, lighter (in terms of data plane memory and CPU), and often faster than Istio Ambient. While acknowledging that Ambient offers a "knob to turn" for sharing levels, he argues that the increased operational complexity, extra network hops, and tuning overhead often do not justify the perceived benefits for the majority of Linkerd users, especially given that many of the historical sidecar "warts" have now been addressed by Kubernetes itself.
Technical Deep Dive
▶ Watch: Defining the sidecar pattern in Kubernetes (2:30)
The technical foundation of the "Great Sidecar Debate" rests on understanding the mechanics of Kubernetes containers, the evolution of its API, and the inherent differences in handling L4 versus L7 traffic in a multi-tenant environment.
At its most basic, a sidecar is a container deployed within the same Kubernetes pod as the main application container. This co-location is crucial because it allows them to share the pod's network namespace, meaning they effectively share the same IP address and network interfaces. They also share cgroups, which provides resource isolation and accounting, and can share file systems through volume mounts. This tightly coupled yet isolated execution environment is what enables the sidecar to transparently intercept and manage all inbound and outbound traffic for the application, as is typical for service mesh proxies like Linkerd's Rust microproxy.
Historically, Kubernetes offered two primary container types: init containers and regular containers. Init containers execute sequentially and must complete successfully before regular containers start. Regular containers are intended to run indefinitely and ostensibly start in parallel. This distinction posed significant challenges for early sidecar implementations. For instance, if a service mesh sidecar was deployed as a regular container, there was no guarantee it would start before the application container, leading to race conditions where the application might try to establish network connections before the proxy was ready to intercept them. This often required complex hacks, such as relying on undocumented behaviors of post-start hooks or specific container ordering, which were fragile. Conversely, if a sidecar was an init container, it would block the startup of other init containers and the main application, which was unsuitable for a long-running proxy.
The solution to these "warts" arrived with a Kubernetes Enhancement Proposal (KEP) that introduced native sidecar container support, largely merged into Kubernetes in late 2022 and 2023. The core of this solution lies in enhancing init containers with a restartPolicy: Always flag. When an init container is marked with restartPolicy: Always, it fundamentally changes its behavior:
- Parallel Startup: Unlike traditional init containers, those with
restartPolicy: Alwaysno longer block the startup of subsequent init containers (also markedrestartPolicy: Always) or regular containers. They effectively start in parallel, allowing the sidecar to be ready concurrently with or even before the application. - Automatic Restart: If an
initContainerwithrestartPolicy: Alwayscrashes or terminates for any reason, the kubelet automatically restarts it, ensuring its continuous availability. - Predictable Termination: A crucial improvement is the new shutdown behavior. When the main application container(s) in a pod terminate (e.g., a job completes or the pod is evicted),
initContainerswithrestartPolicy: Alwaysare automatically and gracefully terminated after the application, ensuring proper connection draining and resource cleanup. This resolved the long-standing job termination synchronization problem.
With these advancements, many of the operational burdens of sidecars were lifted, making them a much more robust and predictable pattern. However, the debate shifted to alternative service mesh architectures that aim to reduce the per-pod overhead of sidecars:
- Node Proxies: In this model, a single proxy (e.g., Linkerd 1.x's JVM proxy) runs on each Kubernetes node, and all pods on that node route their traffic through this shared proxy. While reducing the number of proxies, this introduces a clear multi-tenancy challenge: a single point of failure and contention for resources on the node.
- Ambient (e.g., Istio Ambient): This architecture attempts a more nuanced approach by splitting the service mesh data plane into two components:
- Z-tunnel (L4 proxy): Deployed as a daemonset on each node, this component handles basic L4 (TCP) functionality, primarily mTLS encryption/decryption and basic traffic redirection. It provides a secure, encrypted tunnel for all node traffic.
- Waypoint Proxies (L7 proxy): These are separate Envoy-based proxies responsible for advanced L7 (HTTP, gRPC) features like request routing, retries, timeouts, and policy enforcement. Crucially, the deployment and sharing model of Waypoint Proxies are tunable. They can be deployed per namespace, per service account, or even configured for more granular isolation, allowing operators to balance resource sharing with isolation needs.
The core technical contention in shared proxy models (Node Proxies and Ambient) is contended multi-tenancy, particularly for L7 traffic. William Morgan draws an analogy to operating system resource management:
- L4 Traffic: Managing fairness for L4 network traffic (e.g., bandwidth, connection limits) is analogous to kernel-level resource management. Tools like queuing disciplines and rate limiting allow for event-level decisions that can enforce fairness relatively effectively.
- L7 Traffic: Processing L7 requests (parsing HTTP headers, applying complex policies, modifying payloads) is more akin to general application execution. This involves unbounded amounts of work per request. Enforcing fairness for such workloads is as complex as ensuring CPU and memory fairness between different processes on an operating system, which relies on sophisticated mechanisms like timer interrupts, preemptive multitasking, and memory management units (MMUs). Modern L7 proxies, including Envoy and Linkerd's proxy, are primarily designed for high-throughput, low-latency single-tenant or cooperative multi-tenant scenarios, not for strictly enforcing fairness between arbitrary, uncooperative application workloads sharing the same proxy instance.
This fundamental design limitation leads to concrete problems in shared L7 proxy environments:
- Noisy Neighbor Issues: One application experiencing a spike in traffic or a misconfiguration can overload the shared proxy, causing latency or failures for all other applications relying on that same proxy. While resource limits can be applied, accurately sizing them for a proxy handling diverse, unpredictable workloads is extremely challenging.
- Increased Blast Radius: A failure in a shared proxy (due to a bug, misconfiguration, or resource exhaustion) affects all applications routing through it, potentially impacting multiple services or even entire namespaces. In contrast, a sidecar failure is isolated to a single pod.
- Complex Resource Sizing: Determining appropriate CPU and memory requests and limits for a shared proxy becomes an intricate task, requiring deep understanding of all applications it serves, which is often impractical.
While Ambient's tunable Waypoint Proxies offer flexibility, they introduce additional network hops and configuration complexity. Morgan highlights that Linkerd's internal benchmarks consistently show its sidecar approach as being smaller, lighter, and faster than Istio Ambient in comparable scenarios. This suggests that for many use cases, the isolation and simplicity of the sidecar model, now bolstered by native Kubernetes support, still offer superior operational characteristics.
Demo / Proof of Concept
▶ Watch: Key benefits and advantages of the sidecar pattern (4:30)
The talk primarily focused on architectural analysis and engineering trade-offs rather than a live demonstration. William Morgan presented a series of diagrams and conceptual explanations to illustrate the different service mesh architectures and their respective challenges. He referenced internal benchmarks conducted by Buoyant, comparing Linkerd's sidecar performance against Istio Ambient, indicating that Linkerd was found to be "smaller and lighter in terms of data plane, memory, and CPU" and "faster, sometimes significantly faster." While these benchmarks serve as empirical evidence supporting Linkerd's architectural choices, they were discussed conceptually rather than shown as a live proof of concept during the presentation.
Defensive Implications
▶ Watch: The downsides and challenges of using sidecars (6:30)
For organizations operating or considering a service mesh, the insights from "The Great Sidecar Debate" offer critical defensive implications:
- Embrace Native Kubernetes Sidecar Support: Defenders utilizing sidecar-based service meshes (like Linkerd or Istio sidecar mode) should ensure their Kubernetes clusters are running versions that support the native sidecar container features (e.g., via
restartPolicy: Alwaysfor init containers). This provides significantly improved operational robustness, guaranteeing sidecar initialization before the application and predictable termination afterward, reducing race conditions and job termination woes. - Understand the Trade-offs of Shared Proxies: For those considering "sidecar-less" or ambient service mesh architectures (e.g., Istio Ambient), it is crucial to deeply understand the implications of shared L7 proxies. The promise of reduced resource overhead must be weighed against the challenges of contended multi-tenancy.
- Mitigate Noisy Neighbor Risks: In shared proxy environments, implement robust monitoring and alerting for proxy resource utilization (CPU, memory, network I/O, connection counts). Design your Waypoint Proxy (Ambient) deployment strategy carefully, possibly opting for more isolated deployments (e.g., per namespace or per service) for critical or high-traffic applications to minimize the blast radius of noisy neighbors. Recognize that current L7 proxies are not designed for strict fairness between uncooperative tenants, and plan accordingly.
- Assess Blast Radius: Understand that shared proxies inherently increase the blast radius of failures. A single proxy failure can impact multiple applications or an entire node/namespace. Implement high-availability strategies for shared proxies and ensure rapid recovery mechanisms are in place. Sidecar models, by contrast, localize failures to a single pod, offering a more contained blast radius.
- Prioritize Operational Simplicity and Predictability: While new architectures can be appealing, evaluate the increase in operational complexity. The "knob to turn" in ambient architectures, while offering flexibility, also demands more tuning, configuration, and expertise. Simplicity often translates to fewer misconfigurations and easier troubleshooting, contributing to a stronger defensive posture.
- Benchmark and Validate Against Your Workloads: Do not blindly trust generic benchmarks or marketing claims. Conduct your own performance and resource utilization benchmarks with your specific application workloads and traffic patterns to validate the chosen service mesh architecture. This will reveal the true impact on latency, throughput, and resource consumption in your environment.
- Consider Security Boundaries: Sidecars provide a clear security boundary at the pod level, acting as a micro-firewall for the application. In shared proxy models, the security boundary shifts. While L4 proxies (like Z-tunnel) provide node-level mTLS, the L7 Waypoint Proxies might introduce a larger shared attack surface or complex policy enforcement challenges if not configured meticulously.
Ultimately, defenders must approach service mesh adoption with a clear-eyed understanding of the underlying engineering trade-offs. The "new thing" is not always the "right thing" for every scenario; a mature evaluation based on operational maturity, security requirements, and performance characteristics is paramount.
Key Takeaways
- Kubernetes Native Sidecar Support: Recent advancements in Kubernetes, specifically the introduction of
restartPolicy: Alwaysfor init containers (merged in 2022-2023), have resolved many historical operational "warts" of sidecars, making them a robust and predictable pattern for service meshes. - The L7 Proxy Placement Debate: The core architectural debate in service meshes now centers on where to deploy the essential L7 proxy. Options include isolated per-pod sidecars, shared node proxies, or the hybrid ambient approach (e.g., Istio Ambient) with split L4/L7 components.
- Multi-tenancy Challenges with Shared Proxies: Architectures that share L7 proxies between multiple applications (node proxies, ambient) introduce complex multi-tenancy challenges. Enforcing fairness for L7 traffic is significantly harder than for L4, leading to issues like noisy neighbors, increased blast radius for failures, and difficult resource sizing.
- Linkerd's Stance: Sidecars for Simplicity and Performance: Buoyant, the maintainers of Linkerd, continues to advocate for the sidecar model. Their benchmarks indicate Linkerd's sidecar approach is often smaller, lighter, and faster than Istio Ambient, and they prioritize the operational simplicity and clear isolation benefits for their users.
- Informed Engineering Trade-offs are Crucial: The talk strongly emphasizes that engineers must understand the deep engineering trade-offs of different service mesh architectures rather than adopting new trends blindly. There is no universally "right" answer; the optimal choice depends on specific application requirements, operational maturity, and tolerance for complexity.
About the Speaker(s)
William Morgan is the CEO of Buoyant, the company behind Linkerd. He is a prominent figure in the cloud-native community and the creator of Linkerd, one of the earliest projects to join the Cloud Native Computing Foundation (CNCF), holding the distinction of being among the first four or five graduated projects. With Linkerd having been around for almost 10 years, Morgan brings extensive experience and strong opinions on service mesh architectures and the fundamental engineering trade-offs involved in building and operating cloud-native systems. His talks are known for their analytical depth and challenge to conventional wisdom in the Kubernetes ecosystem.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk by William Morgan provides a brutally honest and technically deep analysis of service mesh architectures, challenging the prevailing hype around "sidecar-less" solutions. It meticulously dissects the evolution of Kubernetes' native sidecar support, demonstrating how recent platform enhancements have mitigated historical sidecar "warts." Crucially, Morgan highlights the fundamental engineering trade-offs, particularly the complex multi-tenancy challenges inherent in shared L7 proxies compared to the isolation and operational simplicity of well-implemented sidecar models, urging practitioners to make informed decisions based on real-world constraints rather than marketing narratives.
Heather Calloway (CISO) — STRONG ACCEPT
This talk dissects the fundamental engineering trade-offs in service mesh architectures, particularly between sidecar and "sidecar-less" or ambient models. It highlights how recent Kubernetes advancements have addressed historical sidecar challenges, while exposing new multi-tenancy and resource contention risks inherent in shared L7 proxy designs. This analysis offers critical insights for platform, security, and executive leadership regarding architectural choices, operational resilience, and the true cost of abstraction in cloud-native environments.