Automated Expansion of Privacy Data Taxonomy for Compliant Data Breach Notification
Yue Qin
Network and Distributed System Security (NDSS) Symposium 2025 · Day 2 · Privacy & Anonymity · Privacy & Anonymity
Overview
This article delves into the research presented by Yue Qin at the NDSS Symposium, focusing on an innovative approach to overcome a persistent challenge in privacy compliance: the significant gap between legal professionals' broad interpretations of data and technical practitioners' specific data handling practices. The talk introduces GRASP (Granularity-Aware Hypernym Prediction), an automated method designed to construct and expand privacy data taxonomies, and Tracy, a practical tool that integrates GRASP to assist privacy professionals in compliant data breach notification.
Key moments
- 0:00 Introduction: Bridging legal and technical data definitions
- 1:30 Understanding Privacy Data Taxonomy and its importance
- 2:20 Challenges in building scalable privacy data taxonomies
- 3:20 Introducing GRASP for granularity-aware hypernym prediction
- 5:30 GRASP's high-level architecture: clustering and projection
- 6:40 GRASP outperforms baselines and LLMs in evaluation
- 8:00 Tracy tool demonstration and positive professional evaluation
Automated Expansion of Privacy Data Taxonomy for Compliant Data Breach Notification
Speakers: Yue Qin
Conference: NDSS Symposium
YouTube: https://www.youtube.com/watch?v=8OxizkeXdIY
Overview
This article delves into the research presented by Yue Qin at the NDSS Symposium, focusing on an innovative approach to overcome a persistent challenge in privacy compliance: the significant gap between legal professionals' broad interpretations of data and technical practitioners' specific data handling practices. The talk introduces GRASP (Granularity-Aware Hypernym Prediction), an automated method designed to construct and expand privacy data taxonomies, and Tracy, a practical tool that integrates GRASP to assist privacy professionals in compliant data breach notification.
The core problem addressed is the difficulty in reconciling the abstract, general terminology used in privacy laws and regulations (e.g., "Personally Identifiable Information" or PII) with the highly specific, technical data items encountered in real-world applications (e.g., "Wi-Fi position" or "cost location"). This disconnect often leads to inefficiencies and inaccuracies in data breach incident reporting. Qin's work is crucial because it offers a scalable, dynamic solution to bridge this interpretative chasm, ensuring that organizations can more effectively identify and categorize restricted data in alignment with legal requirements like GDPR.
The importance of this research extends to any organization handling personal data, particularly in an era of increasingly complex data privacy regulations and a constantly evolving technological landscape. By automating the process of mapping specific technical data to legal privacy categories, GRASP and Tracy empower privacy professionals to achieve greater accuracy and efficiency in their compliance efforts, ultimately reducing legal risks and improving the transparency of data handling.
Background
▶ Watch: Introduction: Bridging legal and technical data definitions (0:00)
The necessity for a robust privacy data taxonomy stems directly from the inherent disparity between legal and technical language in the realm of data privacy. Privacy laws, drafted by legal experts, often employ broad, conceptual terms like PII or "personal data," which are open to various interpretations. In contrast, software developers and system architects work with highly specific data items—such as IP addresses, device IDs, GPS coordinates, or financial transaction details—each with distinct technical attributes and usage contexts. This semantic gap creates a significant hurdle for organizations attempting to comply with data protection regulations, especially during critical incidents like data breaches.
Prior research, including a study presented by the speaker at CCS, highlighted this very challenge, revealing a "significant gap" in how legal professionals and technicians interpret data items for incident reporting. To address this, a privacy data taxonomy is proposed as a structured hierarchical solution. This taxonomy aims to capture the relationships between "restricted data"—a term encompassing data defined by privacy laws and data used in real-world applications. It links specific data items to broader, more general terms through hypernym-hyponym relationships. For example, "precise location" might be a hyponym of "cost location," which in turn is a hyponym of "location information"—a clear illustration of capturing both data types and levels of granularity.
However, building and maintaining such a taxonomy is fraught with difficulties. Restricted data items vary widely across diverse applications, from mobile apps and IoT devices to third-party SDKs. Furthermore, privacy definitions themselves are often broad, vague, and inconsistent across different legal jurisdictions and regulations. Existing privacy taxonomies predominantly rely on manual curation or heuristic rules. While these methods can establish initial structures, they fundamentally lack scalability and dynamism. They struggle to adapt to the emergence of new data types, evolving technologies, or changes in privacy laws, making them static in a domain that is inherently dynamic. This static nature prevents organizations from keeping pace with the rapidly changing privacy landscape, underscoring the need for an automated, adaptive solution.
Key Findings
▶ Watch: Challenges in building scalable privacy data taxonomies (2:20)
The central contribution of this work is the development of GRASP (Granularity-Aware Hypernym Prediction), an automated methodology designed to recognize restricted data and integrate it into a privacy data taxonomy. GRASP directly addresses the critical challenge of determining hypernym relationships, particularly concerning the granularity of data items within the taxonomy. Unlike previous methods that often conflate semantically similar but granularly distinct terms (e.g., "coarse location" and "precise location"), GRASP explicitly incorporates granularity information to enhance prediction accuracy.
Key findings include:
- Granularity-Aware Approach: GRASP significantly improves hypernym prediction by considering the relative position and structural context of terms within the taxonomy tree. It groups hypernym-hyponym pairs into distinct clusters based on their granularity layers, learning separate projection matrices for each cluster. This ensures that the learned projections are "granularity aware," allowing the system to accurately distinguish between fine-grained data categories.
- Superior Performance: In comprehensive evaluations, GRASP demonstrated superior performance compared to various baseline models. These baselines included four naive methods, two state-of-the-art hypernym discovery techniques, and even a large language model (LLM), specifically ChatGPT 3.5. Despite careful prompt engineering and fine-tuning with the training dataset, ChatGPT 3.5 did not achieve the best performance, likely due to its lack of domain-specific knowledge. GRASP consistently outperformed all benchmarks across different evaluation metrics.
- Practical Tool Development: To demonstrate GRASP's real-world applicability, the researchers developed Tracy, a tool designed to assist privacy professionals. Tracy integrates GRASP to analyze incident reports, highlighting recognized privacy-sensitive data. Crucially, it then presents a subgraph of the privacy data taxonomy, visually connecting the identified restricted data all the way back to the definitions in privacy laws, such as the General Data Protection Regulation (GDPR).
- Positive User Evaluation: A user study involving 15 privacy professionals (both legal and security experts) provided highly positive feedback on Tracy. Participants unanimously agreed that Tracy achieved "excellent traceability," enabling them to link specific incident data to legal definitions. The tool also significantly improved their efficiency by reducing assessment time, with most participants expressing a willingness to integrate Tracy into their daily compliance analysis tasks.
These findings collectively present a robust, scalable, and user-validated solution for automating the complex task of privacy data classification, offering a substantial advancement in ensuring compliant data breach notification.
Technical Deep Dive
▶ Watch: Introducing GRASP for granularity-aware hypernym prediction (3:20)
The core technical innovation lies in GRASP's ability to discern and leverage the granularity inherent in privacy data taxonomies, a feature largely overlooked by prior hypernym prediction techniques. Traditional methods, broadly categorized into lexical pattern-based, distributional representation-based, and projection-based approaches, primarily rely on semantic similarity. While effective for identifying broad categories like "location information" or "financial information," they falter when distinguishing between fine-grained layers, such as "coarse location" versus "precise location," which are semantically similar but critically different in terms of their privacy implications and hierarchical position.
GRASP's design is predicated on the observation that in a privacy taxonomy, higher-level restricted data terms are typically broader and more abstract concepts, whereas lower-level terms are more domain-specific and technical. This structural characteristic directly reflects the granularity layers. GRASP capitalizes on this by taking a "granularity-aware" approach to hypernym prediction.
The high-level architecture of GRASP involves several key steps:
- Granularity-Based Clustering: The method first groups hypernym-hyponym pairs into distinct clusters. This clustering is not based solely on semantic proximity but critically on their relative granularity layers within the existing taxonomy structure. The number of clusters, denoted as
K, is a hyperparameter determined by the size and complexity of the taxonomy tree (e.g., largerKfor extensive privacy policy taxonomies, smallerKfor specific IoT domains). This ensures that pairs with different granularity levels are processed separately. - Attention Weight Computation: For each cluster, GRASP computes a set of attention weights. These weights are derived based on the distance of each hypernym-hyponym pair to the center of its respective cluster. This mechanism allows the model to focus more on relevant aspects of the relationships within each granularity layer.
- Granularity-Aware Projection Matrix Learning: A crucial step involves learning separate projection matrices for each cluster. Instead of a single, universal projection, GRASP learns distinct mappings that transform hyponym embeddings into hypernym embeddings, tailored to the specific granularity context of each cluster. This ensures that the learned projection metrics are inherently "granularity aware," capturing the subtle hierarchical distinctions.
- Aggregated Projected Offset: The projected offsets from all clusters are then aggregated. This aggregation combines the granularity-specific insights derived from each cluster into a comprehensive representation.
- Attention-Based Classifier: Finally, an attention-based classifier is applied to determine the existence of a hypernym relationship in a candidate pair. This classifier leverages the aggregated granularity-aware projections to make a more accurate and context-sensitive decision.
For evaluation, GRASP was tested against two curated privacy data taxonomies: a privacy policy taxonomy and an IoT sensitive data taxonomy. The performance was benchmarked against a diverse set of models, including four naive methods, two state-of-the-art hypernym discovery techniques, and a large language model, ChatGPT 3.5. The researchers made a concerted effort to optimize ChatGPT 3.5 by carefully refining prompts and fine-tuning it with their training dataset. Despite these efforts, GRASP consistently outperformed all baseline models across various evaluation metrics, underscoring the efficacy of its granularity-aware design over general semantic similarity or broad knowledge bases. The lack of domain-specific knowledge in LLMs proved to be a significant limitation in this specialized task.
The system's dynamic nature was confirmed during the Q&A session. When new privacy laws emerge or existing ones evolve, the system can adapt by retraining its embedding model using the updated legal documents. This allows GRASP to adjust and update the privacy data taxonomy automatically, ensuring it remains current with the dynamic privacy landscape.
Demo / Proof of Concept
▶ Watch: GRASP outperforms baselines and LLMs in evaluation (6:40)
The practical utility of GRASP is best demonstrated through Tracy, a privacy professional assistant tool that integrates the GRASP methodology. While the talk did not feature a live, interactive demonstration, the speaker detailed Tracy's functionality and presented positive evaluation results from a user study.
Tracy is designed to streamline the process of analyzing privacy-related documents, particularly incident reports, for compliant data breach notification. When a privacy professional uses Tracy to review an incident report, the tool automatically identifies and highlights restricted data items within the text. This immediate recognition of sensitive data is the first layer of assistance.
Beyond mere identification, Tracy provides crucial context. Alongside the highlighted data, it displays a subgraph of the privacy data taxonomy. This visual representation is key: it dynamically connects the specific, technical restricted data item found in the report (e.g., an IP address or a device ID) through its hierarchical hypernym relationships, all the way up to the broader, legally defined categories, such as PII or "personal data" as specified in regulations like GDPR. This visual traceability allows privacy professionals to clearly understand the legal implications of the technical data they are reviewing.
To validate Tracy's effectiveness, a user study was conducted involving 15 privacy professionals, comprising both legal and security experts. The results of this evaluation were overwhelmingly positive:
- Excellent Traceability: All participants agreed that Tracy achieved "excellent traceability," enabling them to seamlessly trace restricted data from an incident report back to its corresponding legal definition within privacy laws. This direct link is invaluable for accurate compliance analysis.
- Improved Efficiency: The tool significantly improved the efficiency of privacy professionals in their daily tasks. By automating data recognition and contextualization, Tracy reduced the time required for assessing incident reports and compliance-related documents.
- High Willingness to Use: A substantial majority of participants expressed their willingness to adopt Tracy as a standard tool for their compliance analysis tasks, highlighting its practical value and user-friendliness.
This successful user evaluation underscores that Tracy, powered by GRASP, moves beyond a theoretical framework to offer a tangible, effective solution that addresses a critical pain point for privacy professionals in real-world scenarios.
Defensive Implications
▶ Watch: Tracy tool demonstration and positive professional evaluation (8:00)
The advancements presented by GRASP and Tracy offer significant defensive implications for organizations grappling with data privacy compliance, especially in the context of data breach notification.
- Enhanced Compliance and Reduced Risk: For organizations, the primary benefit is a more robust and consistent approach to data breach notification. By automatically identifying and classifying restricted data according to legal definitions (e.g., under GDPR), organizations can ensure that their breach reports are accurate and compliant. This reduces the risk of regulatory fines, legal challenges, and reputational damage stemming from misclassification or inadequate reporting.
- Bridging the Legal-Technical Divide: The tool directly addresses the long-standing communication gap between legal and technical teams. Legal professionals gain a clear, traceable path from specific technical data items to their legal classifications, while technical teams can better understand the legal context of the data they manage. This fosters better collaboration and more informed decision-making during incident response.
- Streamlined Incident Response: In the chaotic environment of a data breach, time is of the essence. Tracy can significantly reduce the assessment time for privacy professionals by automating the recognition and interpretation of sensitive data in incident reports. This allows for faster, more accurate decisions on whether a breach needs to be reported and what specific data categories are involved, enabling organizations to meet tight regulatory deadlines.
- Dynamic Adaptation to Evolving Regulations: One of the key strengths of GRASP is its ability to adapt. As privacy laws evolve, new data types emerge, or existing definitions change, the underlying embedding model can be retrained with updated documents. This means the privacy data taxonomy can be automatically updated, ensuring that an organization's compliance framework remains current without extensive manual overhaul. This is a critical advantage in a rapidly changing regulatory landscape.
- Improved Data Governance and Inventory: Beyond breach notification, the capability to automatically expand and maintain a granular privacy data taxonomy supports better data governance. Organizations can more effectively inventory their data assets, understand the privacy sensitivity of different data types, and implement appropriate protection measures proactively. This shifts focus from reactive breach response to proactive data protection.
- Consistency Across Applications: Given the diversity of data items across mobile applications, IoT devices, and third-party SDKs, GRASP provides a consistent methodology for classifying data regardless of its origin. This ensures a unified approach to privacy data management across an organization's entire digital footprint.
In essence, GRASP and Tracy empower organizations to move beyond reactive, manual, and often inconsistent privacy compliance efforts towards a more automated, proactive, and consistently compliant posture, thereby strengthening their overall defensive capabilities against data privacy risks.
Key Takeaways
- GRASP for Automated Taxonomy Expansion: The research introduces GRASP (Granularity-Aware Hypernym Prediction), an automated method for constructing and expanding privacy data taxonomies, specifically designed to address the dynamic nature of privacy laws and data types.
- Granularity-Awareness is Key: GRASP's core innovation lies in its ability to recognize and leverage data granularity within a taxonomy, outperforming traditional semantic similarity methods and even fine-tuned LLMs like ChatGPT 3.5 in hypernym prediction.
- Tracy Tool for Privacy Professionals: The Tracy tool integrates GRASP, providing privacy professionals with an assistant to automatically recognize and interpret privacy-sensitive data in incident reports, linking it back to legal definitions for GDPR-compliant data breach notification.
- Validated Real-World Impact: User studies with privacy professionals confirm Tracy's "excellent traceability" and efficiency improvements, highlighting its practical value in reducing assessment time and enhancing compliance accuracy.
- Dynamic and Scalable Solution: The GRASP methodology is designed to be dynamic, allowing for automatic updates to the privacy data taxonomy as new privacy laws emerge or data types evolve, by retraining the underlying embedding models.
- Bridging the Legal-Technical Gap: This work provides a crucial bridge between the broad, abstract language of privacy laws and the specific, technical terminology used in real-world data handling, which is essential for accurate and compliant data breach reporting.
About the Speaker(s)
The speaker for this presentation was Yue Qin. Based on the transcript, Yue Qin is a researcher focused on privacy, specifically addressing the challenges of interpreting data items between legal professionals and technicians in the context of data breach incident reporting. The presentation highlights a previous study published in CCS, indicating an established background in privacy research. No further specific biographical details, such as academic affiliation or company, were provided in the transcript.
Reviews
Dr. Zero (Offensive Security Researcher) — SOLID
Legitimate academic research solving a real pain point — the legal-technical semantic gap in breach notification — with a purpose-built NLP method (GRASP) that demonstrably beats baselines including a fine-tuned LLM. Solid NDSS-tier systems paper, but it's a narrow compliance automation tool with incremental ML novelty, not something that redraws the threat landscape or teaches a practitioner a new attack surface.
Heather Calloway (CISO) — SOLID
Credible academic research that solves a real compliance problem — the legal-technical semantic gap in breach notification — with a validated tool and measurable results. The work is technically sound and the use case is legitimate, but it stays firmly in the privacy compliance lane without surfacing the governance, accountability, or incident command implications that would make it essential for security leaders.
→ Top-rated talks at Network and Distributed System Security (NDSS) Symposium 2025
All talks from Network and Distributed System Security (NDSS) Symposium 2025