So you think you can airgap? (No.)

Ziyad Edher (Infrastructure and Security · Anthropic)

BSidesSF 2026 · Day 1 · AMC IMAX

Overview

In the rapidly evolving landscape of artificial intelligence, securing the colossal compute clusters that train and operate large language models presents unique and formidable challenges. Ziyad Edher, an infrastructure and security expert at Anthropic, delivered a compelling talk at BSides SF, provocatively titled "So you think you can airgap? (No.)," to address these very issues. His presentation delves into Anthropic's innovative approach to protecting their most valuable asset – the multi-terabyte model weights of their AI systems like Claude – from sophisticated attackers, even when a true air gap is impractical for remote research.

Watch on YouTube

Key moments

  1. 0:40 Talk's core goal: make data exfiltration prohibitively slow
  2. 2:00 AI cluster's core: tokens in/out and massive model weights
  3. 3:40 AI research environments: bleeding-edge tech, huge attack surface
  4. 5:40 Protecting 'Claude's brains' assuming full root compromise
  5. 6:40 Why true airgaps and DLPs are impractical for remote research

So you think you can airgap? (No.)

Speakers: Ziyad Edher

Conference: BSides SF

YouTube: https://www.youtube.com/watch?v=eO_htzaUcjQ

Overview

In the rapidly evolving landscape of artificial intelligence, securing the colossal compute clusters that train and operate large language models presents unique and formidable challenges. Ziyad Edher, an infrastructure and security expert at Anthropic, delivered a compelling talk at BSides SF, provocatively titled "So you think you can airgap? (No.)," to address these very issues. His presentation delves into Anthropic's innovative approach to protecting their most valuable asset – the multi-terabyte model weights of their AI systems like Claude – from sophisticated attackers, even when a true air gap is impractical for remote research.

Edher argues that while a complete air gap is unfeasible for modern AI development, a practical alternative lies in making data exfiltration prohibitively slow. This talk details a novel strategy of leveraging the inherent asymmetry between the massive size of model weights and the relatively small bandwidth required for legitimate operational traffic. By creating an egress choke point that rigorously limits outbound data, Anthropic aims to significantly extend the time an attacker needs to steal valuable assets, increasing the likelihood of detection and potentially rendering the stolen data obsolete. This approach serves as a critical layer of defense-in-depth, acknowledging the inherent vulnerabilities of fast-paced research environments while striving for a robust security posture against advanced persistent threats.

Background

▶ Watch: Talk's core goal: make data exfiltration prohibitively slow (0:40)

The core challenge addressed in this talk stems from the unique operational demands and security realities of large-scale AI research environments. Unlike traditional, well-hardened production systems, AI research clusters are characterized by rapid iteration, experimental workflows, and a reliance on bleeding-edge technology stacks—from applications down to the silicon. This environment fosters developer productivity but inherently introduces a huge attack surface, with critical vulnerabilities frequently discovered in new software and hardware components. Despite standard hardening efforts, relying solely on internal controls to prevent compromise is insufficient, especially when faced with sophisticated state actors.

The primary asset at risk is the AI model weights, which Edher describes as "Claude's brains." These weights represent hundreds of millions of dollars of compute, extensive algorithmic work, and vast amounts of data, resulting in files often terabytes in size. Crucially, these weights are mostly incompressible, making them difficult to shrink. The business value of these clusters is fundamentally tied to "tokens in and tokens out"—the processing of text, images, or other data—while the weights themselves remain internal. Legitimate operational traffic, such as remote debugging, metrics, control plane communication, and data transfers, typically involves megabytes per second. This creates a significant asymmetry: massive assets (weights) versus relatively small, legitimate data flows.

Traditional perimeter defenses, while often foundational, fall short in this specific context. A true air gap, physically disconnecting the cluster from the outside world, is the gold standard for security but completely incompatible with remote research and development, requiring researchers to physically relocate to data centers. Domain allow lists, a common security practice, are valuable but can be circumvented by sophisticated attackers exfiltrating data to legitimate but compromised services like s3.amazonaws.com or github.com. Lastly, Content Data Loss Prevention (DLP) systems are generally ineffective against determined adversaries who can encrypt sensitive data before exfiltration, rendering content-based detection useless. Recognizing these limitations, Anthropic sought a defense mechanism that specifically leveraged the unique properties of their assets and operational environment.

Key Findings

▶ Watch: AI cluster's core: tokens in/out and massive model weights (2:00)

The central premise and key finding of Anthropic's work is that while a true air gap is impractical for modern AI research, it is possible to make data exfiltration prohibitively slow by exploiting the asymmetry between the massive, mostly incompressible size of AI model weights (terabytes) and the comparatively small volume of legitimate egress traffic (megabytes per second) required for remote operations. This strategy does not aim to stop exfiltration entirely or prevent command-and-control communication; rather, it seeks to significantly delay it.

By introducing an egress rate limiter—effectively a "narrow pipe"—Anthropic transforms a rapid data theft into a multi-day, multi-week, or even multi-month endeavor. This extended exfiltration window serves several critical purposes:

  1. Increased Detection Opportunity: Every day an attacker spends exfiltrating data is another day they risk detection, loss of foothold, or having their access revoked due to routine infrastructure restarts or security operations.
  2. Reduced Asset Value: For fast-moving AI models, delays of months could mean that by the time an attacker successfully exfiltrates, reconstructs, and operationalizes the stolen weights, Anthropic may have already shipped a newer, more advanced model, thereby diminishing the value of the compromised asset.
  3. Deterrence: The sheer time commitment and increased risk aim to make the theft of model weights an unattractive target for all but the most dedicated and patient adversaries.

This approach acknowledges a threat model where an attacker has achieved full root access on compute nodes, possessing access to source code and infrastructure configurations. In this scenario, traditional software-based controls within the compromised environment are unreliable. Therefore, the egress limiter functions as an edge control, operating outside the influence of the compromised compute cluster, acting as a "dumb" perimeter defense that cannot be easily deceived or bypassed by an attacker within the network.

Technical Deep Dive

▶ Watch: AI research environments: bleeding-edge tech, huge attack surface (3:40)

Anthropic's solution hinges on implementing a robust egress choke point designed to meticulously count and control all outbound network traffic from their AI compute clusters. The fundamental principle is to force all egress through a dedicated perimeter appliance and drop packets once an allotted limit is exceeded. This is akin to a physical air gap with a severely restricted cable, allowing only minimal data flow.

A critical aspect of this system is its comprehensive byte counting methodology. Unlike typical network monitoring, this system counts every single byte leaving the cluster, including those often overlooked in egress calculations. This includes L3 payloads from protocols like ICMP (Internet Control Message Protocol), DNS (Domain Name System) queries, TLS handshakes, and even SSH keepalives. A particularly insightful example given is TCP Acknowledgement (ACK) packets. While small, they are sent for every ~100 bytes received. If a job is primarily downloading data, its outbound TCP ACK stream could inadvertently trigger the rate limiter, illustrating the strictness of the control. The control operates without differentiating between "legitimate" or "illegitimate" bytes; it simply counts L3 payloads, mirroring the indiscriminate nature of a physical air gap.

To prevent a sudden hard stop that would cripple legitimate operations, Anthropic employs a token bucket filter for rate limiting. This strategy allows for bursts of traffic when bandwidth is available in the "bucket," which accrues over time (e.g., during off-peak hours or downtime). Users can temporarily exceed the stated per-second limit, consuming tokens from the bucket. As the bucket drains, the system gradually reduces egress speed, eventually throttling traffic down to the configured per-second limit, resulting in a smoother experience for TCP clients that adjust to the changing network conditions.

Securing the egress limiter itself is paramount. It's deployed in an isolated cloud account with stringent security measures: no standing access for humans in production, changes only via a two-party approval automated infrastructure actuation system. This ensures that even a full compromise of the compute clusters does not allow a pivot into the egress limiter environment to disable it. Furthermore, significant effort is dedicated to preventing attackers from finding alternative routes out of the cluster, such as creating new NAT gateways, using standard cloud security best practices.

Addressing the challenge of multi-tenant cloud services like S3 and logging, which can act as alternative exfiltration paths, is crucial. Anthropic leverages the cloud provider's IAM (Identity and Access Management) system. They designate "in-boundary" S3 buckets (accessible only from within the compute boundary) and enforce policies preventing the compute boundary from accessing "out-of-boundary" buckets. To facilitate necessary, small cross-boundary transfers (e.g., researchers accessing non-sensitive data), they built proxies that route this traffic through the egress limiter, ensuring it's accounted for.

Finally, to mitigate the shared blast radius problem (where one high-egress job impacts everyone), a tiered enforcement strategy is implemented:

  1. Node-local daemons/agents: These run on each compute node, restricting egress for individual Kubernetes pods or containers. This serves as a reliability control, limiting the impact of misconfigured jobs to a single node or pod. Importantly, this is not a security control, as an attacker with root access could easily disable the underlying Linux traffic control filters.
  2. Perimeter routers: At the actual perimeter, egress traffic is categorized by destination (e.g., logging, metrics, blob storage). This allows for separate bandwidth buckets for each service type. If, for instance, a logging service experiences a surge in traffic, only its allocated bucket is impacted, preventing it from affecting other critical services like blob storage.
  3. Cluster-wide limit: All per-destination buckets aggregate into a total regional or cluster-wide egress limit, which ultimately determines the maximum rate at which an attacker could exfiltrate data.

This multi-layered technical architecture ensures that egress is not only strictly limited but also resilient to internal misconfigurations and robust against external bypass attempts.

Demo / Proof of Concept

▶ Watch: Protecting 'Claude's brains' assuming full root compromise (5:40)

While the talk did not feature a live, real-time demonstration of the egress limiter in action, Edher provided compelling evidence of its effectiveness through the results of its rollout and the impact on daily operations. The implementation of this egress control involved a significant, multi-month effort—specifically, four to six months of reclassifying S3 buckets, migrating services, engaging with various teams, and rearchitecting systems to reduce or eliminate cross-boundary communication.

The primary "proof of concept" lies in the quantitative outcome: Anthropic successfully killed off approximately 98% of all egress from their compute clusters. This substantial reduction demonstrates the system's capability to enforce strict outbound data policies.

Furthermore, the system effectively acts as a "demo" in real-world scenarios by actively detecting and impeding accidental or sloppy exfiltration attempts. Edher recounted instances where researchers, intending to transform model weights, inadvertently attempted to upload them to an "out-of-boundary" S3 bucket not designated for sensitive assets. Under the egress limiter, such an upload, which might typically take minutes, was throttled down to an abysmal 128 kilobits per second, extending the transfer time to several days. This drastic slowdown immediately triggered alerts and prompted researchers to question the network performance, leading to education about proper data handling and the security implications of their actions. This operational feedback loop serves as a continuous, albeit passive, demonstration of the limiter's deterrent effect and its ability to surface anomalous data flows.

Defensive Implications

▶ Watch: Why true airgaps and DLPs are impractical for remote research (6:40)

The deployment of Anthropic's egress limiter offers several critical defensive implications for organizations grappling with similar security challenges, particularly those operating large-scale, sensitive compute environments:

  1. Understand Your Asset's Physics and Asymmetry: The most crucial takeaway is to analyze the inherent properties of your most valuable data. Is it massive and incompressible? Is the legitimate traffic associated with it significantly smaller than potential exfiltration volumes? Identifying such asymmetries can reveal unique opportunities for "physics-rooted" security controls that attackers cannot easily negotiate.
  2. Thorough Data Flow Mapping is Essential: The rollout of the egress limiter, despite its challenges, forced Anthropic to meticulously map and understand all their data flows. This process, while "chaos monkey-esque," uncovered intricate interdependencies between networking services that were previously unknown. Organizations should proactively undertake such data flow analysis to identify all potential egress paths and dependencies before implementing strict controls.
  3. Perimeter Controls as Defense-in-Depth, Not Primary Defense: Edher explicitly states that perimeter controls are ultimately a "concession" and should not be the sole or primary line of defense. They are invaluable as a layer of defense-in-depth, particularly when internal environments (like research clusters) cannot be fully trusted due to their inherent complexity and attack surface. The goal is to make an attacker's job harder, buying time for detection and response.
  4. Prioritize Hardening the Control Itself: Any critical perimeter control must be exceptionally hardened against compromise and bypass. This includes deploying it in an isolated environment, implementing strict access controls (e.g., no standing access, two-party approval for changes), and actively working to prevent alternative egress routes. The control's integrity is paramount.
  5. Invest in Observability for Network Performance: The "rollout nightmare" highlighted a significant gap in observability tools, making it extremely difficult to diagnose when the egress limiter was the cause of slow network performance or cascading failures. Robust network monitoring and logging capabilities, with clear metrics tied to rate limiting, are vital for successful implementation and ongoing operation.
  6. Future-Proofing through TCB Minimization: While the egress limiter provides immediate benefits, the long-term defensive strategy should focus on minimizing the Trusted Computing Base (TCB) for sensitive assets. This involves aggressively securing the minimal set of software and hardware that directly handles unencrypted model weights, potentially through technologies like attested Trusted Execution Environments (TEEs) and confidential compute. This is a multi-year effort requiring application-level changes and collaboration with hardware vendors, but it represents the ultimate direction for robust security.
  7. Sloppy Exfiltration Prevention: The system effectively acts as a deterrent for accidental or unsophisticated data exfiltration attempts, immediately flagging and severely throttling any unexpected large data transfers. This provides an educational opportunity to reinforce secure data handling practices within the organization.

In essence, this talk encourages defenders to think creatively about their specific environments and leverage unique properties to build resilient, multi-layered security architectures, even as they work towards more fundamental long-term solutions.

Key Takeaways

  • Exploit Asymmetry: Leverage the significant size difference between critical assets (e.g., multi-terabyte AI model weights) and legitimate operational traffic (megabytes per second) to create effective egress controls.
  • Prohibitively Slow Exfiltration: The goal is not to stop exfiltration entirely, but to make it so slow (days, weeks, months) that it becomes prohibitively time-consuming for attackers, increasing detection opportunities and potentially devaluing the stolen asset.
  • Comprehensive Byte Counting: Implement egress controls that count all outbound bytes, including often-overlooked L3 payloads like TCP ACKs, ICMP, and DNS, to prevent subtle bypasses.
  • Hardened Perimeter Choke Point: Deploy the egress limiter in an isolated, highly secured environment with strict access controls (e.g., two-party approval, no standing access) to prevent attackers from disabling or bypassing it.
  • Tiered Enforcement for Resilience: Utilize a multi-layered approach with node-local limits (for reliability) and perimeter-based, destination-bucketed limits (for security and blast radius reduction) to manage egress.
  • Long-Term TCB Minimization: While egress limiting is a valuable defense-in-depth, the strategic direction for securing highly sensitive assets should focus on minimizing the Trusted Computing Base (TCB) through technologies like attested Trusted Execution Environments (TEEs) and confidential compute.

About the Speaker(s)

Ziyad Edher, known as "Zed," works on infrastructure and security at Anthropic. In this role, he is responsible for building, operating, and securing the massive training and research clusters that power Anthropic's large language model, Claude. His work involves tackling the complex security challenges inherent in cutting-edge AI development environments, balancing the need for rapid research iteration with robust protection of highly valuable intellectual property.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

Edher takes a genuinely novel problem — securing multi-terabyte AI model weights against exfiltration in a hostile research environment where you can't trust your own compute nodes — and solves it with elegant, physics-grounded engineering rather than the usual DLP theater. The asymmetry insight (TB assets, MB/s legitimate traffic) is simple but the implementation details are real and hard-won. Not groundbreaking security research, but clearly operational work from someone who actually built and rolled it out.

Heather Calloway (CISO) — STRONG ACCEPT

Edher presents a genuinely novel defensive control grounded in physical reality rather than policy aspiration — exploit the asymmetry between asset size and legitimate egress to make exfiltration economically irrational. The threat model is honest, the implementation is detailed, and the operational evidence is real. It stops short of a 5 because it doesn't translate to governance accountability or tell the broader security leader audience what decisions to make about their own environments.

→ Top-rated talks at BSidesSF 2026

All talks from BSidesSF 2026