Trino and Data Governance on Kubernetes - Sung Yun & Aki Sukegawa, Bloomberg
Sung Yun, Aki Sukegawa, Bloomberg
KubeCon + CloudNativeCon Europe 2025 · Session
Overview
This talk by Sung Yun and Aki Sukegawa from Bloomberg delves into the sophisticated deployment of Trino as a managed service on Kubernetes, specifically addressing the complex requirements of data analytics and governance within a large financial institution. Faced with petabyte-scale financial data, diverse ingestion sources, real-time streaming, and a growing demand for secure data discovery for AI workloads, Bloomberg sought to centralize its data analytics infrastructure. The presentation outlines their journey to build a platform that enables data owners to securely share data catalogs across numerous internal teams, while simultaneously empowering these teams with self-service Trino clusters.

Key moments
- 0:00 Introduction: Trino and data governance on Kubernetes at Bloomberg
- 0:25 Overview of Bloomberg's massive-scale data environment
- 2:00 Trino: Scalable, distributed engine for analytics and governance
- 3:25 Managed Trino as a Service: resource, catalog, access definitions
- 4:20 Enforcing runtime data governance via policy decision point
- 5:45 Kubernetes for isolated, resilient, multi-tenant Trino deployments
- 7:00 The challenge: Centralized catalog management and policy administration
- 8:20 Technical solution: Leveraging Trino's built-in access control mechanism
Trino and Data Governance on Kubernetes - Sung Yun & Aki Sukegawa, Bloomberg
Speakers: Sung Yun, Aki Sukegawa, Bloomberg
Conference: KubeCon EU
YouTube: https://www.youtube.com/watch?v=vCfehltPKxk
Overview
This talk by Sung Yun and Aki Sukegawa from Bloomberg delves into the sophisticated deployment of Trino as a managed service on Kubernetes, specifically addressing the complex requirements of data analytics and governance within a large financial institution. Faced with petabyte-scale financial data, diverse ingestion sources, real-time streaming, and a growing demand for secure data discovery for AI workloads, Bloomberg sought to centralize its data analytics infrastructure. The presentation outlines their journey to build a platform that enables data owners to securely share data catalogs across numerous internal teams, while simultaneously empowering these teams with self-service Trino clusters.
The core of their solution revolves around integrating Trino with Open Policy Agent (OPA) on Kubernetes. This architecture provides robust, runtime authorization checks, ensuring that only authorized users and applications can access specific data resources, down to the column level. The talk highlights the creation of custom Kubernetes resources (CRDs) to abstract Trino-specific configurations, facilitating both the deployment of isolated, multi-tenant Trino clusters and the centralized administration of granular data access policies. This innovative approach not only streamlines data access but also significantly enhances the security posture and governance capabilities of Bloomberg's extensive data environment.
Background
▶ Watch: Introduction: Trino and data governance on Kubernetes at Bloomberg (0:00)
Bloomberg operates in a data environment characterized by immense scale and complexity. Their teams manage petabytes of financial data, sourced from over 125,000 curated news feeds, market data, and third-party alternative data providers. The environment also processes real-time streaming data for critical market insights and features a proliferation of self-maintained data catalogs leveraging on-prem S3-compatible object storage. A paramount requirement is securing data discovery for AI workloads, ensuring that access is granted only to authorized entities for specified use cases. This landscape presented a clear opportunity to centralize data analytics, enabling data owners to securely share their catalogs across many teams.
Traditional data analysis pipelines at Bloomberg often involve data engineering teams building catalogs using tools like Apache Spark, Apache Flink, or Apache Iceberg, then utilizing large-scale distributed processing engines for data transformation and interactive analysis. This often led to a fragmented ecosystem, increasing the "switching costs" for users learning multiple tools. Trino emerged as a preferred solution due to its capabilities as a scalable, highly distributed processing engine. Trino optimizes query performance through parallel processing, predicate pushdowns, and distributed caching, which significantly reduces redundant S3 reads. Its ANSI SQL compliance and versatility for both ad-hoc interactive analysis and multi-hour batch workloads made it an ideal choice for a unified analytics platform.
The Bloomberg team provides Trino as a managed service, allowing users to define Trino cluster resource configurations (e.g., memory allocation), catalog definitions (connection properties to data sources), and crucial catalog access control definitions. The platform's core features include deploying Trino clusters matching user specifications and, critically, providing runtime data governance by enforcing predefined access control policies. The challenge was to achieve this in a multi-tenant environment on Kubernetes. Tenants, typically engineering teams managing data lakes for specific business units (e.g., news data, market data), required isolated Trino deployments. This isolation was crucial to prevent service disruptions in one business unit from affecting others. While Trino offers resource groups, tenants preferred dedicated, resilient deployments that they could fine-tune. This requirement, combined with the need for robust deployment management, scalability, and network security, naturally led to the choice of Kubernetes as the underlying infrastructure for building a Trino deployment controller. However, the ease of data access provided by Trino also heightened the data analytics environment's risk profile, necessitating a robust, centralized mechanism for catalog management and policy administration to enable secure sharing across Kubernetes namespaces.
Key Findings
▶ Watch: Trino: Scalable, distributed engine for analytics and governance (2:00)
The central contribution of Sung Yun and Aki Sukegawa's work at Bloomberg is the successful implementation of a managed Trino as a service platform on Kubernetes that fundamentally reshapes data governance in a large-scale, multi-tenant environment. Their key findings and contributions include:
- Distributed and Secure SQL Solution with Runtime Authorization: They deployed Trino in combination with Open Policy Agent (OPA) on Kubernetes, creating a robust, distributed SQL solution. This architecture applies authorization checks dynamically at runtime, ensuring that data access policies are enforced precisely when queries are executed, rather than relying solely on static configurations.
- Granular, Column-Level Data Access Control: The platform enables data owners to define highly granular access controls, down to the individual column level within tables. This precision allows sensitive data to be protected while still facilitating broader sharing of less sensitive information, addressing a critical need for secure data discovery across different teams and AI workloads.
- Centralized Policy Administration for Distributed Data Catalogs: A significant innovation is the development of a centralized catalog management and policy administration mechanism. While Trino catalog definitions can be distributed and managed within individual Kubernetes namespaces by data owners, the system provides a higher-level, centralized view. This ensures that consistent authorization checks can be applied across all Trino clusters, even when catalogs are shared across different namespaces, effectively balancing distributed ownership with centralized governance.
- Kubernetes Custom Resources (CRDs) for Abstraction and Automation: The team introduced custom Kubernetes resources to codify Trino-specific abstractions:
- Trino Catalog CRD: Facilitates the sharing of data catalog configurations without exposing sensitive access credentials.
- Access Control CRD: Enables precise, granular definition of access controls that map directly onto data catalogs.
- Trino Service CRD: Simplifies the instantiation of Trino clusters for service owners, allowing them to request clusters with a minimal set of inputs, while the underlying Kubernetes controller handles the complex resource generation and optimization.
- Multi-Tenant Platform with Secure Cross-Tenancy Sharing: The entire architecture supports a multi-tenant environment where business units operate isolated Trino deployments within their respective Kubernetes namespaces. Crucially, the system facilitates secure ownership of resources and enables cross-tenancy sharing of data catalogs through the centralized policy enforcement, addressing the inherent tension between isolation and collaboration.
- Unified Authorization Backend for Data and Computation Access: They demonstrated the effective reuse of the OPA authorization backend for both data access control (via a Trino OPA plugin) and computation access control (for HTTP access to Trino queries and UI, via an Envoy OPA plugin). This unified approach simplifies the policy enforcement infrastructure and ensures consistency across different access types.
Technical Deep Dive
▶ Watch: Enforcing runtime data governance via policy decision point (4:20)
The technical implementation hinges on the interplay of Trino, Kubernetes, and Open Policy Agent (OPA) to establish a robust and secure data analytics platform.
At its core, Trino's built-in access control mechanism is leveraged. When a user submits an ANSI SQL query, the Trino coordinator, before execution, analyzes the SQL statement and generates authorization questions (e.g., "Can this user select these columns from this table?"). This allows for fine-grained policy definition, including column-level access.
To answer these authorization questions, Bloomberg opted for Open Policy Agent (OPA). OPA is an open-source policy engine that decouples policy enforcement from the application logic. It accepts policy questions as JSON over HTTP or gRPC and returns decisions in the same format. A critical feature of OPA for this implementation is its ability to fetch policies from external services in a well-defined way, known as bundles. The external service providing these bundles is called a bundler. This mechanism ensures that policies can be managed centrally without creating a single point of failure or bottleneck for distributed OPA instances.
The architecture places an OPA server collocated with each Trino cluster. When a user's query triggers an authorization check, the Trino OPA plugin makes an HTTP request to its co-located OPA server. This OPA server, in turn, makes policy decisions based on the bundles it periodically downloads from a central OPA bundler.
The management and orchestration of Trino clusters, catalogs, and access policies are achieved through Kubernetes Custom Resources (CRDs) and a dedicated Kubernetes controller:
- Trino Catalog Custom Resource:
- Data owners define their Trino catalogs as Kubernetes custom resources. These are structured like plain-text configuration files but with a crucial enhancement: the ability to inject secrets.
- The
plainTextPropertiesfield contains non-sensitive connection details, whilesecuredPropertiesholds sensitive information like access credentials. - The Kubernetes controller uses these secured properties to generate the final Trino catalog configuration file without exposing sensitive values to service owners or logs. This separation enhances security and allows service owners to discover existing catalogs via
kubectl getwithout revealing credentials.
- Access Control Custom Resource:
- Data owners define access policies by referencing a Trino Catalog CRD by name.
- These access controls specify groups of tables and columns for which access can be granted. For instance, a news data team can exclude sensitive columns from an access control definition to share a catalog more broadly.
- The Kubernetes controller is responsible for configuring the OPA bundler. It generates intermediate representations of these policies in a separate policy store, optimizing the computation required for bundle generation.
- The OPA server periodically downloads these policy bundles, which contain policies tailored for user identities (e.g., JWT tokens). When a user query arrives with their identity, OPA uses these bundles to allow or reject execution based on user groups or roles. The status field of the CRD indicates the readiness and configuration status of these external policy resources.
- Trino Service Custom Resource:
- To simplify Trino cluster deployment for service owners, a custom resource with a limited configuration is introduced. Service owners specify only essential parameters, such as the list of catalogs to be mounted and the total memory allocation for their cluster.
- The Kubernetes controller then generates the fully detailed configuration behind the scenes, including all necessary Kubernetes resources like Trino coordinator and worker pods, Ingresses, and ConfigMaps. This abstraction shields users from the complexity of configuring parameters like maximum query memory, which needs to be derived from container memory, JVM heap, and internal Trino memory pools.
- This approach offers a trade-off: reduced flexibility for specific, highly customized configurations in favor of ease of use and opportunities for the platform owner to optimize resource allocation (e.g., future auto-scaling). The controller also exposes derived configurations, like the maximum memory per query (e.g., 32 GB), in the CRD's status field for user information. An example given is a
singleQueryLimitPercentagefield, allowing service owners to guarantee consistent memory for a certain number of concurrent queries (e.g., 50% for two queries). - The generated Kubernetes resources closely mirror the structure of the open-source Helm chart for Trino. The status field also provides endpoint URLs for the Trino query service and UI once the cluster is ready.
Finally, the system also implements computation access control for Trino queries and UI access, which are primarily via HTTPS. This is achieved by introducing another open-source component: an Envoy OPA plugin. This plugin authorizes HTTP accesses via OPA, effectively reusing the same OPA backend and policy infrastructure for both data access and access to the compute resources themselves. For computation access control, which is generally simpler, no additional Kubernetes CRD was required in this initial implementation.
Demo / Proof of Concept
▶ Watch: Kubernetes for isolated, resilient, multi-tenant Trino deployments (5:45)
While the presentation did not feature a live, interactive demonstration of the system in action, the speakers meticulously detailed the architectural design and the operational flow of their managed Trino platform. The talk itself serves as a comprehensive proof of concept, illustrating how the various components—Trino, Kubernetes, OPA, and custom CRDs—interoperate to deliver a robust, secure, and governed data analytics solution.
The speakers walked through the user experience from the perspective of both data owners and Trino service owners. Data owners define their Trino Catalogs and Access Control policies using Kubernetes Custom Resources, specifying data sources and granular access rules up to the column level. Trino service owners then leverage another Custom Resource to instantiate their Trino clusters, mounting the defined catalogs with simplified configurations. The core demonstration of the system's capability lies in its ability to enforce these policies at runtime: when a user submits a query to a Trino cluster, the co-located OPA instance evaluates the request against the distributed policies, granting or denying access based on the predefined rules. This detailed explanation, including diagrams and workflow descriptions, effectively proves the viability and functionality of their integrated data governance solution on Kubernetes.
Defensive Implications
▶ Watch: Technical solution: Leveraging Trino's built-in access control mechanism (8:20)
The Bloomberg team's architecture for Trino and data governance on Kubernetes offers significant defensive advantages for organizations managing large, sensitive datasets:
- Granular, Runtime Authorization: The integration with OPA provides column-level access control, which is a critical defensive capability. Instead of relying on coarser table-level permissions, organizations can precisely control access to sensitive data points within a table (e.g., PII, financial figures), dramatically reducing the blast radius of potential data breaches. This runtime enforcement ensures policies are applied dynamically with every query.
- Centralized Policy Management and Consistency: By centralizing policy administration through the OPA bundler and Kubernetes controllers, organizations can ensure consistent enforcement of data access rules across all Trino clusters, regardless of their deployment location or the team managing them. This eliminates policy drift and simplifies compliance audits, providing a single source of truth for authorization logic.
- Multi-Tenant Isolation and Resilience: Deploying Trino clusters in isolated Kubernetes namespaces for different business units provides a strong security boundary. A service disruption or misconfiguration in one tenant's cluster is less likely to affect others, enhancing overall platform resilience and limiting the scope of any security incident.
- Secure Data Catalog Sharing: The CRD-based system for defining Trino catalogs allows data owners to share access configurations without exposing sensitive credentials. This fosters secure collaboration and data discovery across teams while maintaining strict control over the underlying data sources, preventing credential leakage.
- Auditability and Accountability: OPA can be configured to log policy decisions, providing a comprehensive audit trail of who accessed what data, when, and whether access was granted or denied. This is invaluable for forensic analysis, compliance reporting, and establishing clear accountability for data access.
- Reduced Attack Surface: By abstracting complex Trino configurations through simplified CRDs, the platform reduces the surface area for misconfigurations by service owners. The platform team can bake in security best practices and optimizations, ensuring that clusters are deployed with a secure baseline.
- Unified Access Control for Data and Compute: Reusing OPA for both data access control (via Trino plugin) and computation access control (via Envoy plugin for HTTP access to Trino UI/queries) simplifies the security architecture. This unified approach reduces complexity, potential misconfigurations, and ensures consistent policy application across different layers of access.
- Future-Proofing for Resilience: While current maintenance operations might require manual retries, the planned integration of a Trino Gateway will introduce intelligent query routing for disaster recovery and maintenance. This will significantly enhance the platform's resilience against outages and improve the user experience during planned and unplanned events, crucial for a high-availability financial environment. Additionally, extending this solution to other compute engines like Apache Spark and Flink will provide consistent, granular authorization across the entire data processing landscape.
Key Takeaways
- Secure, Distributed SQL Analytics: Bloomberg successfully deployed Trino on Kubernetes, integrating Open Policy Agent (OPA) to deliver a secure and distributed SQL solution with runtime authorization checks.
- CRD-Driven Abstraction: Kubernetes Custom Resources (CRDs) are used to manage Trino-specific abstractions, including Trino catalogs (for secure data source configuration sharing), Trino access controls (for granular policy definitions), and simplified Trino service deployments.
- Granular, Runtime Authorization: The platform enables column-level data access control enforced dynamically by OPA, enhancing security and allowing precise data sharing without exposing sensitive information.
- Multi-Tenant & Cross-Namespace Sharing: The architecture supports a multi-tenant environment with isolated Trino clusters while enabling secure cross-namespace sharing of data catalogs through centralized policy administration.
- Unified OPA Backend: OPA serves as a unified authorization backend for both data access control (via a Trino plugin) and computation access control (for HTTP access via an Envoy plugin), streamlining policy enforcement.
- Platform Optimization & Ease of Use: The custom Kubernetes controller simplifies Trino cluster provisioning for service owners, abstracting complex configurations and allowing platform owners to optimize resource allocation and performance.
About the Speaker(s)
Sung Yun and Aki Sukegawa are members of the Bloomberg team, where they are involved in the development and management of advanced data analytics infrastructure. Their work focuses on deploying and scaling open-source technologies like Trino and Kubernetes to handle Bloomberg's massive scale of financial data and evolving data governance requirements. They specialize in building managed services that enable secure, self-service data analytics for various engineering and business units within the company.
Reviews
Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT
This talk from Bloomberg's Sung Yun and Aki Sukegawa presents a highly sophisticated and practical architecture for deploying Trino as a managed service on Kubernetes, specifically addressing petabyte-scale data governance in a multi-tenant financial environment. Their custom Kubernetes controllers and CRDs, combined with Open Policy Agent (OPA) for granular, column-level runtime authorization, represent a significant advancement in secure data discovery and access control. This is a solid, well-engineered defensive solution that demonstrates deep technical expertise and offers actionable insights for any organization grappling with complex data security challenges.
Heather Calloway (CISO) — STRONG ACCEPT
This talk from Bloomberg presents a robust and scalable architecture for managing granular data governance within a multi-tenant, petabyte-scale environment using Trino, Kubernetes, and Open Policy Agent. It directly addresses the critical challenge of secure data discovery for AI workloads by enabling centralized policy administration and column-level access control, providing a clear model for risk ownership and institutional accountability in complex data landscapes.