Secure Data Analytics in Apache Spark with Fine-grained Policy Enforcement and Isolated Execution

Byeongwook Kim

Network and Distributed System Security (NDSS) Symposium 2025 · Day 3 · Confidential Computing 2

Overview

In an era defined by massive data generation and the increasing demand for collaborative analytics, cloud-based Apache Spark has emerged as a cornerstone for processing big data. However, the convenience and scalability offered by cloud platforms come with significant security and privacy challenges. This talk, "Secure Data Analytics in Apache Spark with Fine-grained Policy Enforcement and Isolated Execution," addresses these critical issues head-on. It introduces a novel architecture for cloud-based Spark that enables secure data analytics while adhering to stringent privacy regulations.

Watch on YouTube · Slides

Key moments

  1. 0:00 Introduction to Spark and data privacy challenges
  2. 2:20 Introducing new secure cloud-based Spark architecture
  3. 4:00 Identifying two core security requirements for Spark
  4. 4:40 Securing the Spark data analysis pipeline
  5. 6:20 Designing fine-grained policy enforcement on Spark plans
  6. 6:40 Motivating scenario: targeted clinical trials and policy example
  7. 8:40 Overview of implementation using AMD SEV and compartmentalization

Secure Data Analytics in Apache Spark with Fine-grained Policy Enforcement and Isolated Execution

Speakers: Byeongwook Kim

Conference: NDSS Symposium

YouTube: https://www.youtube.com/watch?v=vHcpDk2WsnM

Overview

In an era defined by massive data generation and the increasing demand for collaborative analytics, cloud-based Apache Spark has emerged as a cornerstone for processing big data. However, the convenience and scalability offered by cloud platforms come with significant security and privacy challenges. This talk, "Secure Data Analytics in Apache Spark with Fine-grained Policy Enforcement and Isolated Execution," addresses these critical issues head-on. It introduces a novel architecture for cloud-based Spark that enables secure data analytics while adhering to stringent privacy regulations.

The core problem tackled by this research is the inherent risk of data breaches and policy violations when valuable, often private, data is entrusted to cloud providers or accessed by potentially malicious data users. The proposed solution redesigns Spark's internal mechanisms to ensure both the confidentiality and integrity of the data analysis pipeline and, crucially, to enforce owner-defined policies at a fine-grained level. This work is pivotal for organizations seeking to leverage the power of cloud analytics without compromising data privacy or facing regulatory penalties.

Presented by Joanna, a postdoctoral researcher at Georgia Tech, the research highlights the practical implications of regulatory frameworks like GDPR, HIPAA, and MOSPA, which often hinder data sharing despite its immense potential. By assuming that both cloud providers and data users can be compromised, the framework offers a robust defense, paving the way for more secure and compliant big data collaboration.

Background

▶ Watch: Introduction to Spark and data privacy challenges (0:00)

Apache Spark, since its introduction in 2012, has become the de facto standard for large-scale data processing and machine learning. Its distributed architecture and in-memory computation capabilities make it exceptionally well-suited for big data analytics. The rise of cloud computing has further propelled Spark's adoption, leading to the widespread use of cloud-based Spark services offered by major vendors like Databricks, Amazon, and Azure. This model allows data owners, who may lack the necessary infrastructure, to easily deploy and manage their data in the cloud. Concurrently, data users can readily access and analyze shared datasets, fostering collaborative research and development.

However, this convenience introduces substantial security vulnerabilities. The fundamental issue is that once valuable data, particularly sensitive or private information, is transferred to the cloud, it becomes susceptible to various threats. Untrusted cloud providers themselves could potentially access and leak data, or their infrastructure could be compromised. Furthermore, even if the cloud provider is trusted, malicious data users could craft queries or applications that violate the data owner's explicit or implicit expectations regarding data usage and privacy.

The urgency of these security concerns is amplified by the strict privacy regulations governing sensitive data. Clinical data, for instance, containing genomic information and medication histories, and financial data, with detailed transaction records, are subject to laws such as the General Data Protection Regulation (GDPR), Health Insurance Portability and Accountability Act (HIPAA), and others like MOSPA. Violations of these regulations can lead to severe penalties, including substantial fines and reputational damage, as evidenced by numerous real-world cases where companies have faced lawsuits for data breaches or misuse. This regulatory landscape often creates a dilemma: the potential of data sharing for innovation is undeniable, yet the risks of non-compliance deter many organizations from fully embracing cloud-based analytics. The challenge, therefore, is to create an environment where data owners can share their data confidently, avoiding regulatory pitfalls, while still enabling powerful and flexible data analysis.

Key Findings

▶ Watch: Identifying two core security requirements for Spark (4:00)

The research identifies two paramount security requirements for enabling secure data analytics in cloud-based Spark, even under the assumption that both data users and cloud providers can be compromised. The first requirement is ensuring the confidentiality and integrity of the entire data analysis pipeline. This means protecting the data and the computation from unauthorized access or tampering throughout its lifecycle, from application submission to distributed execution. The second, and equally critical, requirement is to ensure that the Spark plans constructed by data users respect owner-defined policies. This goes beyond mere pipeline security, addressing the malicious intent of users who might try to extract prohibited information even through legitimate-looking queries.

To meet these requirements, the proposed framework introduces several key contributions:

  1. Compartmentalized Spark Application Execution: User code is isolated from the Spark core components, preventing malicious user applications from compromising the Spark library or altering the plan generation process.
  2. Trusted Plan Enforcement Point: All Spark plans, regardless of their origin, are routed through a trusted component within a secure address space. This acts as a choke point for policy enforcement before execution.
  3. Distributed Confidential Computing Environment: The actual execution of Spark plans occurs within a confidential computing environment, leveraging hardware-assisted security features like AMD Secure Encrypted Virtualization (SEV). This protects data and computation even on compromised distributed nodes, ensuring that data remains encrypted in memory and computations are isolated.
  4. Fine-grained Policy Enforcement on Spark Plans: A novel policy enforcement mechanism is designed, operating directly on the Spark plan using pattern matching. Data owners can define policies using a new language based on regular expressions, specifying forbidden sequences of operations or access patterns within the logical plan.

The efficacy and performance of this framework were rigorously evaluated. The security evaluation demonstrated its ability to enforce seven custom-defined policies against 22 queries from the TPCH benchmark, successfully preventing unauthorized data access or linkage. Performance-wise, the framework introduces an average overhead of 35% in latency and 25% in throughput, which are considered acceptable trade-offs for the significant security guarantees provided in sensitive big data environments. These findings collectively demonstrate a robust and practical approach to securing cloud-based Spark analytics.

Technical Deep Dive

▶ Watch: Securing the Spark data analysis pipeline (4:40)

The proposed secure Spark architecture fundamentally re-architects key components to achieve robust security against compromised data users and cloud providers. To understand the technical innovations, it's crucial to first grasp Spark's internal workflow. A data user develops a Spark application which interacts with Spark libraries to define data transformations and actions. The library internally constructs a Spark plan, a high-level blueprint of the data analysis procedure. This plan is then optimized and split into multiple tasks, which are distributed and executed across various nodes in the cluster. Finally, the computed results are returned to the user application.

The research identifies two primary attack vectors:

  1. Compromise of the Data Analysis Pipeline: Malicious users or cloud providers could inject malicious code into the Spark library to construct a harmful Spark plan, or tamper with the execution of tasks on distributed nodes, even if the plan was initially benign.
  2. Policy-Violating Spark Plans: Untrusted data users could construct Spark plans that, while technically valid, violate the data owner's privacy expectations (e.g., linking sensitive data points that should remain separate).

To counter these, the framework implements a multi-layered defense:

Securing the Data Analysis Pipeline

  1. Compartmentalization of Spark Application: Traditionally, untrusted user code and the Spark library share the same address space. This framework addresses this by separating the user's untrusted code from the Spark core components into distinct address spaces. This address space isolation prevents a compromised user application from directly manipulating or compromising the Spark core, ensuring the integrity of Spark's internal logic and plan generation. Each Spark context is created in its own isolated environment.
  2. Trusted Plan Enforcement Point: After the Spark plan is constructed by the application, it is not immediately executed. Instead, all Spark plans are relayed to a trusted point within a secure address space before execution. This trusted point acts as a mandatory gatekeeper, verifying the integrity and policy compliance of the plan before it proceeds to distributed execution.
  3. Distributed Confidential Computing Environment: The actual execution of the tasks derived from the Spark plan occurs within a confidential computing environment. This environment is built using hardware-assisted security features such as AMD Secure Encrypted Virtualization (SEV), including its newer variants like SEV-CBS (Confidential Compute for Bare Metal).
  • Remote Attestation: Before any data or tasks are sent to a distributed node, the node undergoes remote attestation. This cryptographic verification process ensures that the node's software and hardware configuration are genuine and untampered. Only nodes that successfully pass this check are allowed to participate in the computation.
  • Memory Encryption: Within these attested nodes, AMD SEV ensures that the memory regions containing sensitive data and computation are hardware-encrypted. This protects the data from unauthorized access by the cloud provider, hypervisor, or other co-located virtual machines, even if the underlying physical server is compromised.
  • Isolated Execution: The computation itself occurs within a Trusted Execution Environment (TEE), providing strong isolation guarantees. This means that even if a malicious actor gains control of the operating system on a distributed node, they cannot inspect or tamper with the data or the execution within the TEE.

Fine-grained Policy Enforcement on Spark Plans

Even with a secure pipeline, a sophisticated attacker might craft a legitimate-looking Spark query that, through a sequence of operations, ultimately leaks private information. To prevent this, the framework introduces a novel policy enforcement mechanism that operates directly on the Spark logical plan.

  1. Policy Language: Data owners define their security expectations using a new policy language based on regular expressions. This language allows owners to specify patterns of nodes (operations) within a Spark plan that are either allowed or disallowed.
  2. Motivating Scenario: Targeted Clinical Trials: Consider a hospital sharing medical data with a pharmaceutical company. The data contains patient names, ages (private), and diagnosis histories (disease, heart rate). The pharma company wants to identify patients for drug testing, but the hospital wants to prevent revealing "who has been diagnosed with which disease."
  • Benign Plan: A plan that analyzes the average heart rate for a diagnosis and retrieves patient names might be allowed.
  • Malicious Plan: A seemingly minor modification, like adding a filter node (e.g., diagnosis.filter(disease == 'cancer')) before joining with the patient table and selecting names, would reveal "who has been diagnosed with cancer." This violates the hospital's policy.
  1. Pattern Matching Enforcement: The policy language allows the hospital to define a policy like: "For the diagnosis table, disallow any plan pattern that first filters by disease, then joins with the patient table, and then selects name."
  • The enforcement mechanism works by matching intermediate nodes in the Spark plan against the regular expression patterns defined in the policy. If a policy-violating pattern is detected (e.g., Filter(disease) -> Join(patient) -> Project(name)), the Spark plan is immediately denied execution. The implementation uses regex matching on the serialized or abstracted representation of the Spark plan's node sequence.

This combined approach ensures that data is protected at rest, in transit, and during computation, while also providing a flexible and powerful mechanism for data owners to dictate precisely how their data can be analyzed, preventing misuse even by authorized users.

Demo / Proof of Concept

▶ Watch: Motivating scenario: targeted clinical trials and policy example (6:40)

While the talk did not feature a live demonstration in the traditional sense, the presenters detailed the implementation of their framework and presented comprehensive evaluation results that serve as a proof of concept for its security and performance.

The implementation details highlight the practical realization of their proposed architecture:

  • Compartmentalization: Achieved by creating each Spark context in a separate, isolated address space, effectively sandboxing the untrusted user code from the Spark core components.
  • Confidential Computing Environment: The distributed execution environment was built leveraging AMD SEV and CBS technologies. This involved configuring Spark to utilize these hardware-assisted security features, ensuring that data processed on worker nodes remains encrypted in memory and isolated within TEEs.
  • Policy Language and Enforcement: The new policy language was designed based on regular expressions, allowing for flexible and powerful pattern matching against the intermediate nodes of Spark plans. The enforcement logic was integrated into the Spark plan optimization and execution pipeline, utilizing a regex matching engine to compare the plan's structure against defined policies.

To evaluate the framework's security, the researchers defined seven custom policies designed to prevent common types of privacy violations, such as obtaining personally identifiable information (PII) directly or indirectly, or linking private information with PII. These policies were then enforced against 22 queries from the industry-standard TPCH benchmark. The evaluation successfully demonstrated that the framework could correctly identify and block policy-violating queries, proving its efficacy in preventing unauthorized data access or inference.

Performance evaluation was also a critical part of the proof of concept. The framework was benchmarked to quantify the overhead introduced by the security mechanisms. On average, the secure Spark framework exhibited a 35% increase in latency and a 25% decrease in throughput compared to an unsecured Spark deployment. These figures represent the cost of enhanced security, which, in scenarios involving highly sensitive or regulated data, is often considered a reasonable trade-off for preventing severe regulatory violations and data breaches.

Defensive Implications

▶ Watch: Overview of implementation using AMD SEV and compartmentalization (8:40)

The secure data analytics framework presented in this talk offers significant defensive implications for organizations dealing with sensitive big data in cloud environments. It fundamentally shifts the security posture from mere perimeter defense to data-centric protection and fine-grained access control at the computation layer.

  1. Enhanced Regulatory Compliance: Data owners, particularly those in healthcare, finance, or government, can leverage this framework to meet stringent privacy regulations like GDPR, HIPAA, and MOSPA. By providing verifiable guarantees of data confidentiality, integrity, and controlled usage, it enables organizations to share data for analytics without the constant fear of regulatory violations. This proactive approach helps avoid costly lawsuits and reputational damage.
  2. Mitigating Insider and Cloud Provider Threats: The framework assumes a strong threat model where even cloud providers and authorized data users can be malicious or compromised. By employing confidential computing (AMD SEV/CBS) and address space isolation, it protects data from unauthorized access by the cloud infrastructure itself, hypervisors, or other VMs. Furthermore, the compartmentalization and trusted plan enforcement point reduce the attack surface from malicious user applications.
  3. Preventing Malicious Data Inference: The novel policy enforcement mechanism, which operates on Spark plans via pattern matching, is a powerful tool for preventing inference attacks. It allows data owners to specify not just who can access what data, but how that data can be processed. This is crucial for stopping sophisticated attackers who might use seemingly benign queries to indirectly extract sensitive information by linking disparate datasets or applying specific filters. Defenders can define policies that prohibit specific sequences of operations that could lead to privacy breaches.
  4. Enabling Secure Collaborative Analytics: For organizations looking to collaborate on shared datasets, this framework provides the necessary trust guarantees. Data owners can confidently provide access to external researchers or partners, knowing that their data will be processed according to their strict rules and within a secure, attested environment. This unlocks the potential of data sharing that is often stifled by privacy concerns.
  5. Forensic and Auditing Capabilities: While not explicitly detailed, the trusted plan enforcement point and the use of remote attestation lay the groundwork for enhanced auditing. Every executed plan can be logged, along with the attestation status of the execution environment, providing a strong basis for forensic analysis in case of a suspected breach.

In essence, this framework empowers defenders by providing a holistic security solution that addresses threats across the entire big data analytics pipeline, from application development to distributed execution. It moves beyond traditional access control to enforce complex usage policies directly at the computational core, fostering a more secure and trustworthy environment for cloud-based data analytics.

Key Takeaways

  • Cloud Spark Security is Paramount: Collaborative big data analytics in the cloud faces significant risks from untrusted cloud providers and malicious data users, necessitating robust security solutions.
  • Multi-layered Defense Architecture: The proposed framework employs a comprehensive strategy, combining compartmentalization, a trusted plan enforcement point, and a distributed confidential computing environment.
  • Hardware-Assisted Security for Core Protection: Utilizing AMD SEV/CBS and remote attestation ensures data confidentiality and integrity during distributed execution, even on potentially compromised cloud infrastructure.
  • Fine-grained Policy Enforcement on Spark Plans: A novel policy language based on regular expressions allows data owners to define and enforce complex usage policies by pattern matching against intermediate Spark plan nodes, preventing malicious inference.
  • Balancing Security and Performance: The framework introduces an average 35% latency and 25% throughput overhead, demonstrating that strong security guarantees are achievable with acceptable performance costs for sensitive data.
  • Enabling Regulatory Compliance: This approach empowers organizations to share sensitive data for analytics while adhering to strict privacy regulations like GDPR and HIPAA, mitigating legal and reputational risks.

About the Speaker(s)

The talk was presented by Joanna, a postdoctoral researcher at Georgia Tech, who introduced the research. She stated that the work was primarily conducted by Pong Kim and herself, under the advisement of Professor Adil Amad and Pyongi. The metadata for this talk lists Byeongwook Kim as a speaker. The research focuses on securing data analytics in Apache Spark through fine-grained policy enforcement and isolated execution, contributing significantly to the field of big data security and privacy.

Reviews

Dr. Zero (Offensive Security Researcher) — SOLID

Competent systems security research that combines TEE-based confidential computing with a regex-based policy enforcement layer on Spark logical plans — a reasonable contribution to the big data security space. The ideas are sound and the threat model is honest, but neither component is individually novel, and the combination doesn't produce a result greater than the sum of its parts.

Heather Calloway (CISO) — WEAK

Technically credible research on securing Apache Spark in cloud environments, with a real problem statement and a working implementation. But it never makes the jump from systems research to operator relevance — there's no path from 'we built this' to 'here's what your organization does with it.'

→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025

All talks from Network and Distributed System Security (NDSS) Symposium 2025