A blueprint for building a generic authorization service for your organization
Ashwin Sidhalinganahalli (Rob), Fletcher Ramee (Rob)
BSidesSF 2026 · Day 2 · AMC Theatre 09
Overview
In the modern landscape of distributed systems, managing access control across thousands of microservices has become an intractable problem, leading to significant security vulnerabilities and hindering developer velocity. This talk, presented by Ashwin Sidhalinganahalli and Fletcher Ramee from Roblox's Platform Security team, introduces a battle-tested blueprint for establishing a generic, scalable, and highly available authorization service within an organization. Their solution aims to decouple authorization logic from application code, centralize policy management, and enable distributed enforcement using open-source tools.
Key moments
- 0:45 The Problem: Fragmented Authorization in Microservices
- 4:00 Impact on Developer Velocity and Security Posture
- 6:00 Broken Access Control: A Top Vulnerability
- 7:45 Introducing the Authorization Service Blueprint
- 9:00 Key Technical Requirements: Availability, Efficiency, Flexibility
A Blueprint for Building a Generic Authorization Service for Your Organization
Speakers: Ashwin Sidhalinganahalli, Fletcher Ramee, Platform Security Team at Roblox
Conference: BSides SF
YouTube: https://www.youtube.com/watch?v=kgpXOBCkH2A
Overview
In the modern landscape of distributed systems, managing access control across thousands of microservices has become an intractable problem, leading to significant security vulnerabilities and hindering developer velocity. This talk, presented by Ashwin Sidhalinganahalli and Fletcher Ramee from Roblox's Platform Security team, introduces a battle-tested blueprint for establishing a generic, scalable, and highly available authorization service within an organization. Their solution aims to decouple authorization logic from application code, centralize policy management, and enable distributed enforcement using open-source tools.
The core challenge addressed is the "fragmented authorization landscape" prevalent in microservice architectures, where individual teams often implement bespoke, coarse-grained, and inconsistent access control mechanisms. This fragmentation results in a lack of central oversight, increased security risks like orphaned permissions, and a substantial cognitive load on developers. By presenting a unified authorization framework, the speakers offer a strategic approach to mitigate these issues, enhance security posture, and streamline development workflows, making it a critical discussion for any organization grappling with access control at scale.
Background
▶ Watch: The Problem: Fragmented Authorization in Microservices (0:45)
The shift from monolithic applications to microservices has brought undeniable benefits, including enhanced scalability, increased reliability, and accelerated product iteration for individual teams. However, this architectural evolution has simultaneously created a monumental challenge for access control. What was once managed at a few, well-defined points in a monolith is now fragmented across potentially thousands of isolated services. This distributed nature renders traditional security and audit practices for access control "extremely untenable," as highlighted by the speakers.
Without a unified strategy, organizations fall into a "fragmented authorization landscape." This typically manifests as:
- Ad-hoc Solutions: Every team develops its own authorization logic, often embedded directly in application code, configuration files, or by "piggybacking" on rudimentary, coarse-grained systems.
- Redundant and Brittle Logic: The lack of consistency leads to duplicated efforts, fragile implementations, and a "spaghetti architecture" that is difficult to maintain and evolve.
- Developer Velocity Bottlenecks: Engineers face a high cognitive load, struggling to understand disparate access control systems, where to request access, and whether they possess the necessary permissions. Reinventing authorization wheels for evolving granular needs further slows development.
- Security Bottlenecks: Security teams must perform deep dives into every new authorization solution, creating delays and introducing the risk of misconfigurations or vulnerabilities.
- Degraded Security Posture: The most severe consequence is a significant degradation of security posture, marked by:
- Zero Central Oversight: Information security teams lose visibility into "which identity has access to what across your infrastructure." Access definitions are scattered across databases, repositories, and config files.
- Orphaned Permissions: Statically defined access controls are rarely cleaned up when identities are offboarded, leading to persistent, unauthorized access.
- Broken Provisioning: The process of granting access becomes ad-hoc, often lacking proper ownership or tooling.
These aren't theoretical risks. Broken Access Control has surged in prominence, moving from outside the OWASP Top 5 in 2013 to the number one vulnerability since 2021. It is also one of the highest-paid bug categories for security researchers. The difficulty in finding these logic bugs with scanners, due to their distributed and custom nature, underscores the need for a fundamentally new approach to managing authorization at scale.
Key Findings
▶ Watch: Impact on Developer Velocity and Security Posture (4:00)
The central premise of the talk is that solving the authorization problem at scale requires a fundamental shift: decoupling authorization from individual application code entirely. The speakers advocate for a "single source of truth for all your access policies," which then enables distributed enforcement based on the specific requirements of internal services and tools. This blueprint, referred to internally as "Guard," is a "battle-tested architecture using open-source tools that can handle millions of authorization decisions per second."
The technical requirements driving this solution are paramount:
- High Availability: Authorization is in the critical path of every request. Any downtime would directly impact application functionality and customer experience, necessitating an "extremely and highly reliable" solution.
- Efficiency: The authorization solution must scale with customer traffic, avoiding wasted capacity while maintaining performance.
- Flexibility: The system must support a wide spectrum of granularity, from coarse-grained API and endpoint access to highly granular controls over UI components and individual fields. It also needs to accommodate diverse use cases, including high-traffic, low-granularity scenarios with P99 millisecond latency requirements.
After evaluating various options, including shipping policy data to external databases (security/privacy risks), commercial licensed solutions (often not open-source or flexible enough), and hybrid approaches, the team converged on leveraging a mature open-source engine designed for high-performance sidecar deployment. Open Policy Agent (OPA) and Topaz were identified, with Topaz ultimately chosen. Topaz stood out due to its design for sidecar operation, its clean and intuitive framework for representing access control, and its adherence to the OAZEN compliant interface, which allows for swapping out authorization solutions if needs change in the future.
Technical Deep Dive
▶ Watch: Broken Access Control: A Top Vulnerability (6:00)
The proposed blueprint, named Guard, is a custom multi-tenant control plane. It serves as the central source of truth for all policies and policy data, allowing different use cases (tenants) to define their authorization rules. Guard doesn't just store policies; it provides mechanisms for services to integrate and enforce these policies end-to-end, supporting both internal services via Topaz integration and broader infrastructure use cases where a sidecar might be less suitable.
An authorization solution generally comprises three key components:
- Enforcement Point: The application itself, ultimately responsible for enforcing an authorization decision. The goal is to make this as easy as possible, ideally delegating the decision.
- Decision Point: A third-party service (e.g., Topaz) that makes authorization decisions based on policy.
- Administration Point: The central API service (Guard) that acts as the source of truth for all authorization policies in the organization.
A critical distinction is made between the control plane (authorization policy updates, managed by Guard) and the data plane (the actual decision-making). The data plane must be highly resilient, as authorization is non-optional; if it fails, the only secure fallback is to deny everything.
The blueprint outlines three primary evaluator models for policy enforcement:
- Central Evaluator:
- Mechanism: In this model, the administration point (Guard) directly acts as the decision point. Applications query Guard for authorization decisions.
- Use Cases: Useful for scenarios requiring high policy freshness and where relaxed latency requirements are acceptable. It can handle very large authorization policies that might necessitate their own database.
- Caveat: Fundamentally higher latency because the lifecycle of the administration point is directly tied to the enforcement point.
- Example: A web application gating internal human access, where immediate policy updates are crucial but a few extra milliseconds of latency are tolerable.
- Sidecar Evaluator (Topaz):
- Mechanism: The Topaz sidecar runs alongside the application. It performs an initial download of policy data from Guard (the administration point), but after this, the data plane dependency ends. The sidecar operates on the last known policy data, making it highly resilient to Guard's availability.
- Performance: Designed for millisecond response times by keeping the entire policy database in memory.
- Caveat: Policy size is limited by the sidecar's heap memory.
- Example: A data broker service handling over one million authorization decisions per second with high availability requirements and a latency target of 10-20 milliseconds. For internal services, policy propagation delay is acceptable, making this model ideal.
- Template Service (Platform-Native Evaluator):
- Mechanism: For extreme use cases requiring near-zero availability degradation and no local network hop for authorization, Guard's central policy format is transpiled into the native authorization policy languages of mature platforms like Istio or Kubernetes. These translated policies are then applied directly to the platforms.
- Performance: Saves several milliseconds per authorization by embedding policy enforcement within the platform.
- Use Cases: Foundational platforms where latency is absolutely critical.
- Examples:
- Service Mesh (Istio): Latency is paramount, so externalized authorization is avoided. A reasonable default granularity (service name, endpoint, HTTP method) is applied to all services. More granular cases revert to the sidecar model.
- Kubernetes RBAC: Reliability is key, ensuring every decision a cluster needs to make remains within the cluster without external calls.
This multi-faceted approach ensures that authorization can be applied effectively across an organization's diverse infrastructure, from individual microservices to underlying platform layers, while balancing freshness, latency, and availability requirements.
Demo / Proof of Concept
▶ Watch: Introducing the Authorization Service Blueprint (7:45)
The speakers illustrated the practical application of their authorization blueprint through several case studies, serving as proofs-of-concept for each evaluator model:
- Data Broker Service (Sidecar Evaluator Model):
- Scenario: A critical data broker service acting as a frontend to a database.
- Requirements: Extremely high traffic (north of a million authorization decisions per second), stringent high availability, and a latency target of 10 to 20 milliseconds. Policy propagation delay was acceptable for its internal clients.
- Implementation: The Topaz sidecar model was deployed. Topaz runs next to the application, downloading policy data initially from Guard and then operating independently. This met the performance and availability needs, demonstrating Topaz's capability for high-volume, low-latency authorization.
- AI Agents and Internal MCP Servers (Sidecar Evaluator Model):
- Scenario: Securing internal Multi-Component Platform (MCP) servers used by various identities, including human employees and AI agents (e.g., employee assistants, coding assistants, external customer support agents).
- Flexibility: The framework is agnostic to identity type. Policies can be configured on Guard for the MCP server's tenant.
- Granular Control:
- Access Denial: An external customer support agent could be easily denied access to the internal MCP server entirely.
- Tool Discoverability: Depending on the agent hitting a list API endpoint, only relevant tools are returned. For instance, a coding assistant might see source code tools, while an employee assistant sees documentation tools.
- Attribute-Based Access: Standard attributes, like a computed "risk score," can be applied across tools. If a coding agent attempts to access a source code tool but its risk score is not low, access can be denied.
- Resource-Level Granularity: For highly granular access, such as an AI agent accessing a specific user's repository, policies can be configured to ensure the agent owns the access. This can be synced from external sources like GitHub team policies.
- Recommendation: For tool-level granularity, the speakers recommend creating a separate tenant for each tool to keep the Topaz sidecar's heap memory in control. This showcases the multi-tenant capability of Guard.
- Service Mesh (Istio) and Kubernetes RBAC (Template Service Model):
- Scenario: Platform-level authorization needs where any data plane dependency on Guard components is unacceptable, requiring near-zero availability degradation and no local network hop.
- Service Mesh (Istio):
- Challenge: Adding even 1-2 milliseconds to every request in a service mesh is too much. While Istio supports external authorization, it's not suitable for all services.
- Solution: Guard transpiles a subset of authorization policies into Istio's native policy format. This allows for a reasonable default granularity (service name, endpoint, HTTP method) to be applied to all services without external calls. More granular use cases can still leverage the sidecar model.
- Kubernetes RBAC:
- Challenge: Reliability is paramount; every authorization decision within a Kubernetes cluster should stay within the cluster.
- Solution: Guard's policies are translated into Kubernetes RBAC definitions and applied directly, ensuring cluster-internal authorization without external dependencies.
These examples collectively demonstrate the blueprint's adaptability, performance, and resilience across various operational contexts, from high-throughput microservices to sensitive platform-level infrastructure.
Defensive Implications
▶ Watch: Key Technical Requirements: Availability, Efficiency, Flexibility (9:00)
Implementing a generic authorization service like Guard has profound defensive implications, offering significant improvements to an organization's security posture and operational efficiency. However, the journey is not without its challenges, and the speakers shared valuable lessons learned and strategies for adoption.
Lessons Learned:
localhostis too slow: Even with a sidecar like Topaz, a 1-2 millisecond latency increase for authorization can be measurable and impactful for critical services.
- Solution 1 (Advanced): Using Topaz as a library directly within the application. Topaz is written in Golang, which limits language support. Compiling to WebAssembly and embedding a WebAssembly runtime is a potential, albeit complex, avenue.
- Solution 2 (Practical): Implementing a decision cache within the authorization SDK. By setting the cache expiry to less than the periodic policy refresh period, the same security properties are maintained while significantly decreasing latency from milliseconds to tens of microseconds for almost all requests. This cache stores the results of Topaz decisions.
- Expensive Policy Refreshes: The periodic full policy refreshes from the Topaz sidecar to Guard are resource-intensive, especially since policies often remain unchanged.
- Solution: Add a short-lived cache on Guard's API for policy export. A 10-second cache on the export API is sufficient if sidecars refresh every minute, preserving freshness guarantees while reducing database load.
- Topaz Denies on Startup: If an application starts before Topaz has completed its initial policy data download from Guard, Topaz, being a sensible authorizer, will deny all requests.
- Solution: Configure the application container to wait until Topaz has received its policy data before starting. This feature was upstreamed, allowing applications to watch Topaz for initial sync completion.
Driving Adoption:
Successfully rolling out a centralized authorization system requires a strategic approach to encourage buy-in and integration across engineering teams:
- Make it the Default for New Services:
- Level 1 Granularity: Enable basic API and endpoint level authorization for all microservices by default, leveraging existing platform integrations (e.g., service mesh like Istio). This provides immediate, broad coverage with minimal effort from service owners.
- Application Security Review: Integrate Guard into the application security review flow for new projects. This allows the security team to recommend Guard as the authorization solution from the outset, making integration much easier as the product evolves.
- Target Risky Existing Services: Instead of attempting a full migration of all existing services, prioritize "risky services" that lack granular access control. This provides immediate, tangible security benefits, making it easier to convince teams to onboard.
- Enable on the Platform Layer: Integrating Guard at the platform layer (e.g., Kubernetes, Istio) involves interfacing with only one or two platform teams but yields "massive visibility of who has access to what across your fleet." This provides significant value with concentrated effort.
Key Defensive Benefits of the Blueprint:
The deployment of Guard offers several critical advantages for security and operations:
- Enhanced Visibility: Transforms a fragmented access landscape into a "single pane of glass." Security teams can now definitively answer "which identity has access to what" across the entire infrastructure.
- Standardization and Governance: Standardizes the interface for access provisioning and evaluation, opening avenues for:
- Building a governance layer for requesting access.
- Defining rightful owners of resources.
- Tooling for access requests (humans and future AI agents).
- Principle of Least Privilege: Enables dynamic, granular, and just-in-time access. Policies are no longer static files, allowing for:
- Auto-expiring access.
- Granting access only when strictly required.
- Removal of Unused Permissions: By correlating persistent access with actual usage data, the system can identify and facilitate the removal of orphaned or unused permissions, reducing attack surface.
- Instant Threat Detection: Standardized authorization decisions are emitted in a consistent format (identity, resource, action). This data can be used to build real-time rules and dashboards that "light up" when a compromised identity probes for access, enabling immediate response and remediation.
Key Takeaways
- Hub-and-Spoke Model Works: A distributed enforcement mechanism coupled with a centralized control plane (Guard) provides both scalability and consistent policy management.
- Embrace Open Source: Tools like Topaz offer a robust, open-source solution with clean primitives for diverse access control needs, proving that high-performance, flexible authorization can be built without proprietary solutions.
- Leverage Native Evaluators for Critical Systems: For foundational platforms like service meshes (Istio) or orchestrators (Kubernetes), utilize their native authorization capabilities by transpiling policies to minimize latency and maximize reliability.
- Decoupling Authorization is Crucial: Separating authorization logic from application code empowers security teams to achieve their goals by centralizing control, enhancing visibility, and enforcing consistent policies.
- Standardization Boosts Developer Velocity: Providing standardized authorization primitives frees developers from "reinventing the same wheel," accelerating their productivity and product delivery.
About the Speaker(s)
Ashwin Sidhalinganahalli and Fletcher Ramee are members of the Platform Security team at Roblox. Their work focuses on developing scalable and robust security solutions for the organization's extensive microservices infrastructure. Their expertise lies in architecting systems that balance high availability, performance, and granular access control, particularly in complex, distributed environments.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
A competent, well-structured case study on building a centralized authorization plane at Roblox using OPA/Topaz — honest about trade-offs, grounded in real production numbers, and contains a few genuinely useful engineering details. Nothing here will surprise anyone who's read the Zanzibar paper or run OPA in prod, but it's delivered with enough operational specificity to be worth the slot at BSides SF.
Heather Calloway (CISO) — SOLID
A technically credible and operationally grounded presentation of a real authorization architecture problem that most large organizations are quietly failing at. The blueprint is honest and specific, but the talk stays in engineering territory and never surfaces the governance and accountability implications that would make it land at the CISO level.