Threat Modelling at Scale: Breaking Down Cloud Complexity
Hanna Papirna (Lead Security Engineer), Emma Yuan Fang (Senior Security Architect · EANM)
Cloud Village @ DEF CON 33 · Day 1 · Cloud Village
Overview
In the rapidly evolving landscape of cloud-native applications, traditional threat modeling approaches often fall short, leaving organizations vulnerable to sophisticated attacks. This talk by Hanna Papirna and Emma Yuan Fang at Cloud Village addresses this critical challenge, presenting a pragmatic framework for scaling threat modeling to the complexity of multi-tenant, microservices-based cloud architectures. The speakers emphasize the need to move beyond simplistic network diagrams and adopt a more granular, decomposed view of cloud systems to effectively identify and mitigate risks.

Key moments
- 0:00 Introduction, talk agenda, and speaker introductions
- 2:20 Why traditional threat modeling fails in the cloud
- 4:20 Understanding multi-tenancy and its security implications
- 5:50 Dynamic workloads, service mesh, identity propagation issues
- 7:30 The proposed approach: breaking down complex cloud systems
- 8:00 Key concept: identifying and securing trust boundaries
Threat Modelling at Scale: Breaking Down Cloud Complexity
Speakers: Hanna Papirna (Lead Security Engineer); Emma Yuan Fang (Senior Security Architect, EANM)
Conference: Cloud Village
YouTube: https://www.youtube.com/watch?v=700JIVRB4OY
Overview
In the rapidly evolving landscape of cloud-native applications, traditional threat modeling approaches often fall short, leaving organizations vulnerable to sophisticated attacks. This talk by Hanna Papirna and Emma Yuan Fang at Cloud Village addresses this critical challenge, presenting a pragmatic framework for scaling threat modeling to the complexity of multi-tenant, microservices-based cloud architectures. The speakers emphasize the need to move beyond simplistic network diagrams and adopt a more granular, decomposed view of cloud systems to effectively identify and mitigate risks.
The presentation delves into the inherent complexities introduced by dynamic workloads, service meshes, and cross-tenant interactions, which are hallmarks of modern cloud deployments. Papirna and Fang introduce a structured approach that leverages architectural decomposition, re-evaluates trust boundaries, and adapts established frameworks like STRIDE for cloud-specific threats. Crucially, the talk also explores the emerging role of AI tools in assisting threat modeling, demonstrating their potential while candidly discussing their current limitations and offering practical tips for maximizing their utility in a cloud context.
This discussion is particularly vital for security engineers, architects, and development teams grappling with securing complex cloud environments. By providing actionable strategies for breaking down intricate systems, identifying unique cloud attack surfaces, and integrating threat modeling into the software development lifecycle, the speakers equip attendees with the knowledge to build more resilient and secure cloud applications at scale. Their insights underscore that threat modeling is not merely a compliance checkbox but a fundamental engineering discipline essential for shaping the future of cloud security.
Background
▶ Watch: Introduction, talk agenda, and speaker introductions (0:00)
The proliferation of cloud-native projects, characterized by multi-tenant environments, microservices, functions, and containers, has rendered traditional threat modeling techniques largely insufficient. Emma Yuan Fang highlighted that simply drawing lines between different network segments no longer captures the intricate attack surface of modern cloud architectures. Applications are increasingly deployed on platforms like Azure, GCP, and AWS, often serving multiple tenants, as seen in SaaS applications where customer data and configurations must remain isolated despite sharing the same underlying application.
The speakers outlined several specific challenges inherent to multicloud, multi-tenant, microservices architectures:
- Shared Infrastructure: Sharing resources across tenants introduces risks related to isolation, secure communication, and identity management.
- Dynamic Workloads: Autoscaling pods and shared clusters mean workloads are constantly shifting, making static perimeter controls ineffective.
- Service Meshes: Technologies like Istio or Linkerd add layers of abstraction, using mechanisms like mTLS (mutual TLS) for service-to-service communication. While enhancing security, they can also complicate comprehensive threat modeling by obscuring potential communication paths and trust relationships.
- Propagating Identities: Ensuring per-tenant access controls across numerous microservices is challenging due to shared APIs and federated identity systems.
- API Management: Depending on architecture, API gateways can centralize or distribute service-to-service communication, each presenting unique threat modeling considerations.
To address these complexities, the core proposal is to decompose the architecture into smaller, manageable components or subsystems. This decomposition is primarily driven by trust boundaries, which are defined as points in the architecture where the level of trust changes, necessitating validation and authentication of data crossing these boundaries. The principle advocated is "zero trust", meaning services should not inherently trust each other. Tools like SPIFFE (Secure Production Identity Framework For Everyone) and SPIRE (SPIFFE Runtime Environment) can be used to assign unique identities to containers, aiding in enforcing trust boundaries at various levels, including cluster, node, pod, or namespace. The choice of boundary level depends on data flow and sensitivity.
The speakers emphasized that while the general process of threat modeling might be familiar to those in application security, the critical step for cloud environments is this systematic decomposition and re-evaluation of trust boundaries to effectively identify threats that traditional methods overlook.
Key Findings
▶ Watch: Understanding multi-tenancy and its security implications (4:20)
The central finding of the talk is that traditional threat modeling frameworks, when applied directly to complex multi-tenant, microservices-based cloud environments, provide a false sense of security. Emma Yuan Fang and Hanna Papirna advocate for a block decomposition approach to reveal a more accurate and comprehensive attack surface. This method involves breaking down a complicated application into tangible, smaller blocks, allowing teams to threat model more effectively.
Key findings and contributions include:
- Block Decomposition for Cloud-Native Architectures: The speakers demonstrated how to split a complex system, such as an identity governance cross-tenant tool, into four distinct blocks: the Internet-facing Gateway, the Cross-tenant Service Mesh, Cross-tenant Identities and API Management, and Data and Storage Components. For each block, they identified specific threat categories that are often missed by traditional approaches.
- Cloud-Specific STRIDE Adaptation: While acknowledging STRIDE (Spoofing, Tampering, Repudiation, Information Disclosure, Denial of Service, Elevation of Privilege) as a well-known framework, they stressed that it must be adapted for cloud environments. The focus should shift to threats like spoofing, information disclosure, and elevation of privilege that exploit cross-tenant identity and API leakage.
- Addressing Dynamic Complexity: Traditional models often treat internet-facing components as single entry points with static perimeters and internal communication as implicitly trusted. The block decomposition approach highlights dynamic tenant routing complexities, pod-to-pod communication threats, and the intricate nature of cross-tenant delegation patterns.
- Real-World Attack Validation: The talk underscored the importance of validating threat models against real-world incidents. The Storm-0558 incident, where attackers stole a Microsoft sign-in key to forge authentication tokens and access emails across multiple tenants, served as a stark example of how a compromise in one area (identity management) can lead to widespread data exposure in multi-tenant environments.
- Systematic Risk Evaluation with DREAD: The talk reinforced the utility of the DREAD (Damage Potential, Reproducibility, Exploitability, Affected Users, Discoverability) framework for a more granular assessment of risk, moving beyond a simple likelihood/impact matrix. This detailed scoring helps prioritize threats more effectively.
- AI as an Assistant, Not a Replacement: While AI tools like Stride GPT can automate threat pattern recognition and aggregate architecture documentation, they have critical pitfalls. They often generate generic threats, ignore existing cloud-native mitigations, and misunderstand cross-tenant complexity because they are trained on traditional monolithic frameworks. Effective AI use requires extensive context, explicit descriptions of tenant isolation, service meshes, and cloud-native controls, and the use of models supporting reasoning tokens (e.g.,
Claude 3 Opus,GPT-4,Gemini 1.5 Pro) to enable a "chain of thought" analysis. - Integration into SDLC: Threat modeling should not be a one-off event but an ongoing engineering discipline integrated throughout the SDLC (Software Development Life Cycle), from initial high-level design (HLD) to ad-hoc sessions for new features and annual reviews, ideally embedded within agile sprints.
Technical Deep Dive
▶ Watch: Dynamic workloads, service mesh, identity propagation issues (5:50)
The core of the technical deep dive presented by Hanna Papirna revolved around the block decomposition approach using a concrete example: an Identity Governance Cross-Tenant Tool. This tool is designed for user lifecycle management (onboarding, access revocation) in scenarios where companies acquire others but need to maintain separate cloud tenants, or for regional compliance requirements (e.g., hospitals).
The architecture was split into two main tenants:
- Tenant A (Managing Cloud Tenant): This tenant hosts the main orchestration and user interface. It's responsible for user onboarding, provisioning, and lifecycle management. The core orchestration and execution layer is a Kubernetes cluster, where each business function operates as a distinct microservice (e.g., one for user provisioning, another for access revocation).
- Tenant B (Receiving/Target Tenant): This tenant receives delegated operations and access policies from Tenant A. It exposes an API for user account synchronization, validates and enforces cross-tenant policies, and manages a local data store.
Crucial for cross-tenant integration, the speakers highlighted the use of Azure Lighthouse, a service that enables delegated management across tenants with granular Role-Based Access Control (RBAC). Similar services exist in other cloud providers.
For threat modeling, the system was logically divided into four blocks, each with specific threat focus areas:
- Internet-facing Gateway:
- Traditional Failure: Treats as a single entry point with static perimeter control, missing dynamic tenant routing complexity.
- Focus Threats: Spoofing (e.g., cross-tenant identity impersonation), Information Disclosure (e.g., DNS cache poisoning). These target the dynamic nature of cloud gateways and multi-tenant routing.
- Cross-tenant Service Mesh (Kubernetes Cluster / Microservices):
- Traditional Failure: Lacks specific threat examples for service meshes, treats internal communication as trusted and single-hop, and fails to capture pod-to-pod communication threats.
- Focus Threats: Service account impersonation, mTLS certificate theft (compromising the mesh's inherent security), and pod memory exposure due to side-channel attacks on shared Kubernetes nodes.
- Cross-tenant Identities and API Management:
- Traditional Failure: Treats API management gateways as simple conduits, overlooking cross-tenant delegation patterns and dynamic service principal relationships.
- Focus Threats: The speakers referenced the Storm-0558 incident, where attackers stole a Microsoft sign-in key to forge authentication tokens, gaining access to emails across multiple tenants. This highlighted a critical vulnerability arising from improper restriction of which tenants could trust which keys, demonstrating how a single key compromise can lead to multi-tenant data exposure. This block demands scrutiny of API delegation, trust relationships between tenants, and the security of identity providers.
- Data and Storage Components:
- Focus Threats: Primarily encryption violation (especially data in transit) and database schema manipulation attacks. This emphasizes the need for robust data protection mechanisms, both at rest and in motion, and integrity controls.
Beyond identifying threats, the talk covered risk evaluation using DREAD. This framework breaks down traditional likelihood and impact into finer-grained factors:
- Likelihood:
- Reproducibility: How easy is it to reproduce the threat?
- Exploitability: How easy is it to exploit?
- Discoverability: How easy is it to discover the threat?
- (Average of these three provides Likelihood score)
- Impact:
- Damage Potential: How much damage can be done?
- Affected Users: How many users can be affected?
- (Average of these two provides Impact score)
Scores are typically rated 1-6. An example was provided: "Compromised Kubernetes service account impersonates a legitimate microservice to access cross-tenant data." This threat could have high damage potential (accessing cross-tenant data) and affect many users (multiple customers across tenants), leading to a high-risk rating. Visualizing these risks on a heatmap (e.g., using a tool like Miro) aids prioritization.
Finally, the talk touched upon Threat Modeling as Code. This approach codifies threat modeling, integrating it into CI/CD pipelines and version control. Resources, architectural components, data flows, and trust zones are defined in declarative formats (e.g., Terraform, YAML, JSON). While offering automation and traceability, a key challenge is carefully designing the security gate in the CI/CD pipeline to avoid slowing down deployments.
Demo / Proof of Concept
▶ Watch: The proposed approach: breaking down complex cloud systems (7:30)
Hanna Papirna provided a live demonstration of AI-assisted threat modeling, showcasing an open-source tool called Stride GPT. This tool supports multiple model providers, and for the demo, Google AI's Gemini 2.5 flash model was used, chosen for its cost optimization.
The demonstration highlighted both the capabilities and the current limitations of AI in threat modeling:
- Initial Prompt with Full Diagram: Papirna first uploaded the full architecture diagram of the Identity Governance Cross-Tenant platform and provided a general description of the solution. When asked to generate threats, the tool produced a list that, while identifying components and their interactions, was relatively short, chaotic, and unstructured. It jumped between various components like the front door, Kubernetes cluster, and Cosmos DB without a clear logical flow, making it less practical for team-based threat modeling sessions.
- Refined Prompt with Specific Context: For the second attempt, Papirna provided the same full diagram but explicitly instructed the tool to concentrate only on the service mesh and AKS cluster components, and included details about trust boundaries. The output was significantly improved. The AI began analyzing various STRIDE categories specifically for Kubernetes, examining the component from different angles. This more focused and structured output was presented as a much better starting point for a team, preventing them from having to "invent threats out of nowhere." A valuable feature of Stride GPT also shown was its ability to provide "improvement suggestions" for prompts, helping users refine their input for better results.
The speakers also mentioned other useful AI tools:
- Threat Canvas and De-Risk AI: Free AI tools supporting diagram import in multiple formats. De-Risk AI specifically supports a layered approach for diagrams (network layer, identity layer, security monitoring layer) to prevent overcrowding.
- Array AI: A tool capable of generating architecture diagrams, including free templates for Kubernetes architectures, useful for quick visualizations.
Despite the promise, several critical pitfalls of AI-assisted threat modeling were emphasized:
- Generic Threat Trap: AI tools often generate extensive lists of common, textbook threats without reflecting the unique attack surface of cross-tenant microservices applications, as they are typically trained on traditional monolithic frameworks.
- Ignoring Existing Mitigations: AI tools tend to assume a blank slate, ignoring cloud-native controls and existing mitigations unless explicitly provided in the context.
- Misunderstanding Cross-Tenant Complexity: This is a significant limitation, as AI models struggle to fully grasp the intricate cross-tenant delegation patterns and trust relationships.
To make AI tools work better, the following tips were provided:
- Use the Block Decomposition Approach: Feed the AI smaller, well-defined blocks rather than an entire complex architecture.
- Provide Extensive Context: Explicitly describe tenant isolation boundaries, service mesh configurations, and cloud-native controls.
- Model Provider Choice Matters: Opt for models that support reasoning tokens. These tokens allow the AI to build a "chain of thought," enabling it to perform more sophisticated analyses, such as checking if any path crosses a tenant boundary without mutual TLS (mTLS) before generating a threat list. Examples of models supporting reasoning tokens include
Claude 3 Opus,GPT-4, andGemini 1.5 Pro.
Defensive Implications
▶ Watch: Key concept: identifying and securing trust boundaries (8:00)
The insights from this talk offer several critical defensive implications for organizations operating in complex cloud environments. The overarching message is to shift from superficial, compliance-driven threat modeling to an embedded, engineering-focused discipline.
- Adopt Architectural Decomposition: Defenders must embrace the block decomposition approach to break down intricate multi-tenant, microservices architectures into manageable components. This allows for a more granular and accurate identification of attack surfaces that traditional methods would miss.
- Re-evaluate and Enforce Trust Boundaries: A fundamental defensive strategy is to rigorously define and enforce trust boundaries. Data crossing these boundaries must be validated and authenticated. Implementing zero-trust principles, where no service inherently trusts another, is crucial. Technologies like SPIFFE/SPIRE can aid in assigning unique identities to workloads for granular trust enforcement.
- Adapt Threat Frameworks to Cloud Nuances: While frameworks like STRIDE are valuable, defenders must adapt them to the specifics of cloud-native, multi-tenant environments. This means focusing on cloud-specific threat categories such as cross-tenant identity impersonation, API leakage, mTLS certificate theft, and side-channel attacks on shared infrastructure.
- Implement Robust Cross-Tenant Security: Given the inherent risks of shared infrastructure and delegated management, robust controls for cross-tenant interactions are paramount. This includes granular Role-Based Access Control (RBAC) (as seen with Azure Lighthouse), secure API delegation patterns, and careful management of service principals. The Storm-0558 incident serves as a stark reminder to scrutinize how authentication keys are managed and trusted across tenants.
- Utilize DREAD for Prioritization: Beyond identifying threats, defenders need a systematic way to prioritize them. The DREAD framework provides a more nuanced risk assessment by evaluating reproducibility, exploitability, discoverability, damage potential, and affected users. This enables security teams to allocate resources effectively to mitigate the most critical risks.
- Integrate Threat Modeling into the SDLC: Threat modeling should not be a one-time event but a continuous process. It must be integrated early in the Software Development Lifecycle (SDLC), ideally during the high-level design (HLD) phase and before detailed design. Furthermore, ad-hoc sessions are necessary for architectural changes, new features, and functionalities, with annual reviews to ensure ongoing relevance. For Agile teams, this means incorporating sprint-based threat modeling.
- Leverage AI as an Intelligent Assistant: Defenders should explore AI-assisted threat modeling tools but with a critical eye. To overcome the "generic threat trap" and "ignoring mitigations" pitfalls, provide AI with rich context, explicit descriptions of cloud-native controls, and utilize models supporting reasoning tokens. AI should augment human expertise, not replace it, serving as a baseline generator and pattern recognition engine.
- Explore Threat Modeling as Code (with caution): For organizations embracing Infrastructure as Code, codifying threat models in YAML, JSON, or Terraform can offer automation, version control, and integration into CI/CD pipelines. However, careful design of security gates is essential to prevent deployment slowdowns.
- Engage Diverse Stakeholders: Effective threat modeling requires collaboration beyond just security and development teams. Project managers, business analysts, architects, and product owners must be engaged to ensure a holistic understanding of the system's purpose, data flows, and potential impact.
- Monitor Emerging Trends: Keep an eye on evolving capabilities in Cloud Security Posture Management (CSPM) tools (e.g., Microsoft Defender for Cloud) that are integrating threat modeling and attack path mapping, as well as tools that can generate data flows for enhanced visibility into cloud footprints.
Key Takeaways
- Decompose Cloud Architectures: Break down complex multi-tenant, microservices cloud systems into smaller, manageable blocks based on trust boundaries to uncover hidden attack surfaces.
- Adapt STRIDE for Cloud-Native Threats: Traditional threat modeling frameworks are insufficient; focus on specific cloud threats like cross-tenant identity impersonation, API leakage, and service mesh vulnerabilities.
- Embrace Zero-Trust Principles: Assume no inherent trust between services; rigorously validate and authenticate all data crossing defined trust boundaries.
- Integrate Threat Modeling into the SDLC: Make threat modeling an ongoing engineering discipline, from initial design through development sprints and architectural changes, not a one-time compliance activity.
- Use AI as a Context-Rich Assistant: Leverage AI tools for automated threat recognition, but provide extensive context, define trust boundaries, and use models with reasoning tokens to overcome generic outputs and improve accuracy.
- Prioritize with DREAD: Utilize the DREAD framework (Damage Potential, Reproducibility, Exploitability, Affected Users, Discoverability) for a granular and effective assessment of risk, enabling better resource allocation.
About the Speaker(s)
Hanna Papirna is a Lead Security Engineer with a strong focus on cloud security and landing zone migration. She is an active Microsoft Certified Trainer, delivering cloud security trainings both online and offline. This was Hanna's first time attending Defcon, having traveled 10 hours from Amsterdam for the conference.
Emma Yuan Fang is a Senior Security Architect at EANM, where she leads a team of architects and engineers across the UK, Germany, and Switzerland. Her specialization lies in AppSec, CloudSec, and Azure security, with over seven years of experience in cloud security. Emma has been a speaker at numerous conferences, including over five BSides events, and this marked her first talk at Defcon.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Competent, practitioner-oriented talk that packages known-good ideas — block decomposition, zero-trust boundaries, STRIDE-for-cloud, DREAD scoring — into a coherent workflow for multi-tenant microservices threat modeling. The Storm-0558 reference and the live Stride GPT demo add texture, but nothing here is original research; it's synthesis and methodology, delivered cleanly at Cloud Village tier.
Heather Calloway (CISO) — SOLID
A competent practitioner-level talk that delivers a real framework for cloud threat modeling — block decomposition is sound, the Storm-0558 reference is well-placed, and the AI limitations section is more honest than most. But it stays inside the engineering layer and never surfaces to the institutional questions that make or break threat modeling programs in practice.