SoK: Unintended Interactions among Machine Learning Defenses and Risks

Vasisht Duddu, Sebastian Szyller, N. Asokan

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

In the rapidly evolving landscape of machine learning (ML), models are increasingly deployed in sensitive applications, necessitating robust defenses against a myriad of security, privacy, and fairness risks. While extensive research has focused on developing individual defenses to mitigate specific threats like evasion attacks or privacy breaches, the practical reality often involves deploying multiple defenses simultaneously. This talk, presented by Vasisht Duddu from the Secure Systems Group at the University of Waterloo, along with co-authors Sebastian Szyller and Professor N. Asokan, delves into the critical but often overlooked problem of unintended interactions among these ML defenses and risks.

Watch on YouTube

Visual summary for SoK: Unintended Interactions among Machine Learning Defenses and Risks by Vasisht Duddu, Sebastian Szyller, N. Asokan
Visual summary for SoK: Unintended Interactions among Machine Learning Defenses and Risks by Vasisht Duddu, Sebastian Szyller, N. Asokan

Key moments

  1. 0:00 Introducing unintended interactions among ML defenses and risks
  2. 1:50 Presenting a systematic framework for unintended interactions
  3. 3:00 Identifying overfitting and memorization as underlying causes
  4. 4:00 Visualizing overfitting and memorization with a synthetic dataset
  5. 6:30 Describing the framework to evaluate unintended interactions
  6. 6:50 Factors influencing overfitting: bias, variance, model capacity
  7. 7:50 Factors influencing memorization: data, objective function, model

SoK: Unintended Interactions among Machine Learning Defenses and Risks

Speakers: Vasisht Duddu, PhD Student, University of Waterloo; Sebastian Szyller; N. Asokan, Professor, University of Waterloo

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=Bx6MktYZ7bE

Overview

In the rapidly evolving landscape of machine learning (ML), models are increasingly deployed in sensitive applications, necessitating robust defenses against a myriad of security, privacy, and fairness risks. While extensive research has focused on developing individual defenses to mitigate specific threats like evasion attacks or privacy breaches, the practical reality often involves deploying multiple defenses simultaneously. This talk, presented by Vasisht Duddu from the Secure Systems Group at the University of Waterloo, along with co-authors Sebastian Szyller and Professor N. Asokan, delves into the critical but often overlooked problem of unintended interactions among these ML defenses and risks.

The core premise of this work is that the effectiveness of a defense, when evaluated in isolation against its intended risk, does not guarantee its behavior when combined with other defenses or when its impact on unrelated risks is considered. This raises crucial questions: Can two defenses negatively conflict with each other? And can a defense, while effective against its target risk, inadvertently increase or decrease susceptibility to an entirely different, unrelated risk? The researchers address this by proposing a systematic framework to understand and conjecture these unintended interactions, positing that overfitting and memorization are the fundamental underlying causes.

This article provides a detailed exploration of their systematization of knowledge (SoK), outlining the framework, the factors influencing these interactions, a guideline for predicting them, and empirical validations of previously unexplored scenarios. The implications of this research are profound for both ML practitioners and security researchers, offering a proactive approach to identify and mitigate complex vulnerabilities that arise from the interplay of defenses and risks in real-world ML deployments.

Background

▶ Watch: Introducing unintended interactions among ML defenses and risks (0:00)

Machine learning models are inherently vulnerable to a diverse spectrum of threats that span security, privacy, and fairness. On the security front, risks include evasion attacks (where adversaries craft inputs to bypass detection), poisoning attacks (where malicious data corrupts the training process), and unauthorized model ownership inference (determining if a model was trained using proprietary data). Privacy risks encompass various forms of data leakage, while discriminatory behavior represents a significant fairness concern. To counter these, the ML security community has developed a wide array of defenses, such as adversarial training for evasion, data sanitization for poisoning, and differential privacy for privacy preservation.

However, a critical gap exists in how these defenses are typically evaluated. Prior work predominantly assesses the effectiveness of a defense solely against the specific risk it aims to mitigate, often in isolation. In practical deployments, ML models are rarely protected by a single defense; instead, multiple mechanisms are often integrated to provide comprehensive protection. This concurrent deployment introduces two primary categories of unintended interactions:

  1. Interactions among defenses: Where combining multiple defenses can lead to conflicts, potentially negating their individual benefits or introducing new vulnerabilities.
  2. Interactions between a defense and an unrelated risk: Where a defense designed for one specific risk might inadvertently increase or decrease the model's susceptibility to another, entirely different risk.

The lack of a systematic framework to explore these complex interdependencies has been a significant challenge. To address this, Duddu and his colleagues conjecture that overfitting and memorization are the fundamental underlying causes of these unintended interactions. They argue that effective defenses may induce, reduce, or rely on these phenomena, and conversely, all identified risks tend to exploit them.

To clarify, overfitting is defined as the difference between a model's accuracy on the training dataset and its accuracy on an unseen test dataset. It is an aggregate metric, reflecting the model's performance across all data records. A high degree of overfitting indicates that the model has learned the training data too well, including its noise, and struggles to generalize to new data.

Memorization, on the other hand, is a more granular metric assigned to individual data records within the training set. It is quantified as the difference in a model's predictions on a specific data record when that record is included versus excluded from the training dataset. If a model's prediction for a record changes significantly when that record is removed, it suggests the model has "memorized" that specific data point.

The relationship between overfitting and memorization can be complex. The authors illustrate this with a synthetic experiment involving a multi-layer perceptron classifying two linearly separable classes.

  • Base Case: Linearly separable data, similar training and test distributions. No overfitting, no memorization.
  • Overfitting, no memorization: Adding noise only to test data records, causing them to fall on the wrong side of the decision boundary. Test accuracy decreases, but the model still learns a simple linear boundary for the training data perfectly.
  • Memorization, no overfitting: Adding noise only to training data records. The data is no longer linearly separable. The model learns a complex, non-linear decision boundary to perfectly fit each noisy training data point. Here, individual training records are memorized, but if the test data distribution remains similar to the underlying clean training distribution, there might not be a significant drop in test accuracy relative to train.
  • Overfitting and Memorization: Adding noise to both training and testing data records. This is where the model learns complex boundaries to fit noisy training data perfectly (memorization) and struggles to generalize to noisy test data (overfitting). The authors emphasize that, given the complexity of modern datasets and high-capacity models, this simultaneous occurrence of overfitting and memorization is the most likely scenario in practice and forms the basis for their subsequent analysis.

Key Findings

▶ Watch: Identifying overfitting and memorization as underlying causes (3:00)

The central contribution of this work is the development of a systematic framework for understanding and predicting unintended interactions among machine learning defenses and risks. This framework is grounded in the observation that overfitting and memorization are the primary underlying causes driving these interactions.

The key findings include:

  1. Identification of Influencing Factors: The researchers meticulously categorized and identified various factors that influence overfitting and memorization. These factors are crucial for gaining a fine-grained understanding of how defenses and risks interact. They span dataset characteristics, objective function properties, and model configurations.
  2. Systematic Literature Survey: A comprehensive survey of existing literature on unintended interactions was conducted. This survey situated prior empirical and theoretical findings within their proposed framework, highlighting which interactions have been explored, whether they lead to an increase or decrease in risk, and whether the influence of specific factors was empirically validated, theoretically derived, or merely conjectured. Crucially, the survey also identified numerous previously unexplored interactions.
  3. A Guideline for Conjecturing Interactions: Based on the identified factors and their correlations with defenses and risks, the authors present a practical guideline for researchers and practitioners to conjecture about previously unexplored interactions. This guideline leverages common factors between a defense and a risk, along with the concept of "dominant factors," to predict whether a defense is likely to increase or decrease susceptibility to an unrelated risk.
  4. Empirical Validation of Unexplored Interactions: The efficacy of the proposed guideline was empirically validated by applying it to two previously unexplored interactions:
  • Group Fairness and Data Reconstruction: The guideline successfully predicted the interaction, which was then confirmed through experiments, demonstrating a decrease in attack success for data reconstruction when a fair model was used.
  • Explanations and Distribution Inference: While the initial guideline based on common factors suggested a decrease in risk, empirical validation revealed an increased susceptibility. This discrepancy highlighted the critical role of non-common factors (like number of attributes and model capacity) in influencing the ultimate interaction, demonstrating the framework's ability to uncover complex dynamics.
  1. Acknowledgement of Limitations: The research also candidly identifies limitations of the guideline, such as its inability to fully account for differences in adversary models and challenges in prediction when too few factors have been evaluated in prior work for certain defenses or risks. This points towards areas for future research and refinement of the framework.

In summary, this paper provides a robust, systematic approach to dissecting the complex interplay between ML defenses and risks, offering both a theoretical foundation and a practical toolset for navigating the challenges of secure and responsible ML deployment.

Technical Deep Dive

▶ Watch: Visualizing overfitting and memorization with a synthetic dataset (4:00)

The core of this work lies in its systematic framework, which posits that overfitting and memorization are the underlying mechanisms driving unintended interactions. To understand these interactions, the framework identifies and categorizes various factors that influence these two phenomena.

Factors Influencing Overfitting and Memorization

Factors Influencing Overfitting:

Overfitting primarily stems from the bias-variance trade-off.

  • Bias: Represents the error from poor hyperparameter choices or an overly simplistic model. For example, a very small model (high bias) may struggle to learn complex relationships between attributes and labels.
  • Variance: Represents the error from the model's sensitivity to fluctuations in the training data. High variance occurs when a model fits the noise in the training data rather than the underlying patterns.

The balance between bias and variance, which directly impacts overfitting, can be managed by two key factors:

  • Size of the training data: Larger, more diverse training datasets generally help reduce variance and mitigate overfitting.
  • Model capacity: The complexity or expressiveness of a model. A model with excessively high capacity relative to the data can easily overfit by memorizing training examples.

Factors Influencing Memorization:

Memorization is influenced by a more diverse set of factors, categorized by their relation to the dataset, objective function, or the model itself.

  • Dataset-related factors:
  • Tail of the distribution: Data records that are outliers or belong to the tail of the data distribution are more likely to be memorized by the model.
  • Number of attributes: The dimensionality of the data can influence memorization, although the exact correlation can be complex.
  • Stable attributes: Whether the model primarily focuses on learning attributes that are stable (i.e., do not change significantly with shifts in data distribution) also correlates with memorization.
  • Objective function-related factors: These relate to the mathematical properties of the function the model optimizes during training.
  • Curvature and smoothness: The curvature and smoothness of the loss landscape can affect how deeply specific data points are embedded in the model's parameters.
  • Distinguishability in model observables across dataset subgroups: If model observables (e.g., predictions, intermediate activations) show high distinguishability between different subgroups within the dataset, certain subgroups might be more prone to memorization.
  • Model-related factors:
  • Distance of training data records to the decision boundary: Training records that are very close to the decision boundary, especially misclassified or noisy ones, are often highly memorized as the model tries to perfectly fit them.
  • Model capacity: This factor influences both overfitting and memorization, as a higher capacity model has a greater ability to memorize individual data points and overfit to noise.

The Guideline for Conjecturing Unintended Interactions

The core contribution of the framework is a guideline designed to systematically conjecture about unintended interactions between a defense D and an unrelated risk R. This guideline leverages the identified factors and their correlations.

The process involves:

  1. Identify Common Factors: For a given defense D and risk R, identify any factors (F) that influence both.
  2. Assess Correlations:
  • Determine how the effectiveness of defense D correlates with a change in factor F. (e.g., Does a defense's effectiveness increase when F increases? Represented by an upward or downward arrow).
  • Determine how the susceptibility to risk R correlates with a change in factor F. (e.g., Does susceptibility to R increase when F increases? Represented by an upward or downward arrow).
  1. Conjecture Interaction Type:
  • If both arrows align (both upward or both downward), it suggests that when defense D is effective, the susceptibility to risk R increases. This is depicted by a red circle in their framework.
  • If the arrows do not align (one upward, one downward), it suggests that when defense D is effective, the susceptibility to risk R decreases. This is depicted by a green circle.
  1. Prioritize Dominant Factors: Often, multiple factors might be common to a defense-risk pair, and they might suggest conflicting interaction types. To resolve this, the guideline introduces the concept of dominant factors.
  • Active factors are those directly exploited by attacks (e.g., specific model observables or data characteristics).
  • Passive factors are related to data or model configuration.
  • Active factors are deemed dominant because changes in them significantly alter susceptibility to a particular risk. If all common factors suggest the same interaction, that's the conjecture. Otherwise, the conjecture from dominant factors takes precedence.
  1. Consider Non-Common Factors: The guideline also acknowledges that factors not common to both the defense and the risk can still influence the overall interaction, adding another layer of complexity.

Examples of Guideline Application

The speakers provided two examples of how their guideline successfully predicted or helped understand interactions previously explored in literature:

  • Differential Privacy (Defense) and Evasion (Risk): Prior work showed that differential privacy (DP) increases susceptibility to evasion attacks. The guideline identified three common factors, with O1 and O3 (model observables/activations) being dominant. The guideline, by evaluating the correlations with these dominant factors, correctly predicted an increase in evasion risk (red circle), matching empirical results. DP, by adding noise to protect privacy, often simplifies the decision boundary, making it easier for adversaries to craft evasion examples.
  • Group Fairness (Defense) and Membership Inference (Risk): Empirical studies indicated that enforcing group fairness can increase susceptibility to membership inference attacks. The guideline, by identifying distance to the decision boundary as a dominant factor, correctly predicted an increase in membership inference risk (red circle). Fair models often adjust decision boundaries to reduce bias, which can inadvertently make the model's behavior around specific training examples more distinctive, aiding membership inference.

These examples demonstrate the utility of the framework and guideline in systematically analyzing and predicting complex interactions based on the underlying mechanisms of overfitting and memorization.

Demo / Proof of Concept

▶ Watch: Factors influencing overfitting: bias, variance, model capacity (6:50)

The talk effectively demonstrated the practical utility of their framework and guideline by applying it to two previously unexplored interactions and empirically validating the conjectures. These validations serve as the "proof of concept" for their systematic approach.

1. Group Fairness (Defense) and Data Reconstruction (Risk)

  • Initial Conjecture: The common factor identified between Group Fairness and Data Reconstruction was distinguishability across subgroups. Based on the framework's correlation analysis, this factor initially suggested a decrease in the risk of data reconstruction (represented by a green circle). One non-common factor, the number of attributes, was also identified as potentially influencing susceptibility.
  • Empirical Validation: The researchers conducted an experiment where a data reconstruction attack was performed on models trained without fairness constraints and models trained with fairness constraints.
  • They found that using a fair model indeed resulted in a lower attack success for data reconstruction, confirming the initial conjecture based on the common factor.
  • Influence of Non-Common Factor (Number of Attributes): Further investigation into the non-common factor revealed its importance. For a smaller number of attributes, the initial conjecture held true. However, as the number of attributes increased, the memorization of individual attributes decreased. Consequently, the beneficial effect of the fair model on data reconstruction attack success diminished, and there was no significant change in attack success when using a fair model. This highlights that while common factors provide a strong initial signal, non-common factors can modulate or even override the overall interaction, requiring a more nuanced analysis.

2. Explanations (Defense) and Distribution Inference (Risk)

  • Initial Conjecture: For the interaction between using model explanations (as a defense or mechanism) and susceptibility to distribution inference (a risk where an adversary infers properties of the training data distribution), the common factor initially suggested a negative interaction (i.e., a decrease in risk, represented by a green circle). Two non-common factors, the number of attributes and model capacity, were also noted as potential influences.
  • Empirical Validation: To validate this, the researchers trained multiple models on training datasets with different distributions (e.g., varying ratios of males to females, denoted as alpha 1 and alpha 2). The objective of the attack was to use model explanations (e.g., saliency maps, LIME, SHAP) from these trained models to predict whether a model was trained on data with distribution alpha 1 or alpha 2.
  • Contrary to the initial guideline conjecture based solely on the common factor, the empirical results showed an increased susceptibility to distribution inference. Across different explanation algorithms, the attack accuracy was greater than 50% for most distribution ratios, indicating that explanations provided exploitable information about the training data's distribution.
  • Influence of Non-Common Factors: The researchers then explored how the non-common factors influenced this unexpected outcome:
  • Number of Attributes: On increasing the number of attributes, they observed a decrease in susceptibility to distribution inference. This was attributed to lower memorization of the relevant attributes crucial for inferring the distribution. With more attributes, the model's focus might be distributed, making it harder to extract specific distribution insights.
  • Model Capacity: Conversely, increasing the model capacity resulted in a higher attack success for distribution inference. This is likely due to the increased memorization capabilities of higher-capacity models, allowing them to embed more fine-grained details about the training distribution into their learned parameters, which could then be exposed via explanations.

These empirical validations demonstrate that the guideline is a powerful tool for conjecturing and exploring interactions. While the common factors provide a strong starting point, the detailed analysis of non-common factors and empirical testing are crucial for fully understanding the complex dynamics and sometimes counter-intuitive outcomes of unintended interactions. The framework thus provides a structured way to navigate these complexities rather than a definitive, always-correct predictor.

Defensive Implications

▶ Watch: Factors influencing memorization: data, objective function, model (7:50)

The findings from this systematization of knowledge have critical implications for both machine learning practitioners and security researchers involved in deploying and securing ML models. The central message is that relying on individually evaluated defenses is insufficient; a holistic understanding of how these defenses interact with each other and with various risks is paramount for robust ML security.

  1. Proactive Risk Assessment: The proposed framework and guideline offer a mechanism for proactive risk assessment. Instead of discovering unintended vulnerabilities post-deployment, practitioners can use this framework to conjecture potential increases in risk when combining defenses or introducing a new defense. This allows for early identification and mitigation strategies. For instance, before deploying a differential privacy mechanism, one could use the guideline to anticipate its potential impact on evasion attack susceptibility.
  2. Informed Defense Design and Selection: The work highlights the need for interaction-aware defense design. Future research should focus on developing defenses that not only protect against their intended risks but also minimize negative unintended side effects on other risks. When selecting multiple defenses for a system, practitioners should consider not just their individual efficacy but also their likely interactions, prioritizing combinations that offer synergistic benefits or, at the very least, avoid exacerbating other vulnerabilities.
  3. The Need for Systematic Evaluation Tools: The talk underscores the current lack of systematic empirical evaluation tools. The speakers mentioned their ongoing work to develop a software framework for systematic empirical evaluation of unexplored interactions. Such a tool would be invaluable for practitioners and researchers to assess the specific risks associated with deploying a particular defense on their unique models and datasets, moving beyond theoretical conjectures to concrete measurements.
  4. Understanding Defense Variants: The research also points to the importance of understanding how different variants of defenses (e.g., different adversarial training algorithms, varying levels of differential privacy) and different variants of risks (e.g., white-box vs. black-box attacks) can impact unintended interactions. A defense variant might have different correlations with underlying factors, leading to different interaction outcomes. This suggests that fine-grained analysis is necessary.
  5. Adversary Model Consideration: A significant limitation identified was the guideline's inability to fully account for differences in adversary models. Defenders must consider the capabilities and assumptions of potential adversaries. A defense that might be robust against one type of adversary could open doors for another. Future work needs to integrate adversary models more deeply into interaction analysis.

In essence, the defensive implication is a call to shift from a siloed approach to ML security to an integrated, systemic perspective. By understanding the underlying causes (overfitting, memorization) and the influencing factors, defenders can build more resilient ML systems that are prepared for the complex interplay of modern threats and protective measures.

Key Takeaways

  • Unintended interactions are a critical concern: Deploying multiple machine learning defenses can lead to unforeseen conflicts or increased susceptibility to unrelated risks.
  • Overfitting and memorization are root causes: The underlying mechanisms of overfitting (poor generalization) and memorization (learning specific training data points) are key drivers of these unintended interactions.
  • A systematic framework exists: The research provides a comprehensive framework to understand and categorize factors influencing overfitting and memorization, which in turn affect defense-risk interactions.
  • Guideline for conjecturing interactions: A practical guideline, leveraging common factors and their correlations with defenses and risks, can help predict whether a defense will increase or decrease susceptibility to an unrelated risk.
  • Dominant and non-common factors are crucial: The concept of dominant factors helps resolve conflicting conjectures, and non-common factors can significantly modulate or even alter predicted interaction outcomes.
  • Need for interaction-aware design and tools: There is a strong need for designing defenses that minimize negative interactions and for developing software tools that allow practitioners to systematically evaluate these complex interdependencies empirically.

About the Speaker(s)

Vasisht Duddu is a PhD student within the Secure Systems Group at the University of Waterloo. His research focuses on understanding and systematizing the complex interactions between machine learning defenses and various security, privacy, and fairness risks.

Sebastian Szyller is acknowledged as a co-author on this joint work, indicating his significant contribution to the research.

N. Asokan is a Professor and served as the advisor for this research, guiding the work presented. His expertise in secure systems underpins the academic rigor of this systematization of knowledge.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

This talk presents a crucial systematization of knowledge on unintended interactions between ML defenses and risks, identifying overfitting and memorization as fundamental underlying causes. It introduces a novel framework and guideline for predicting these complex interactions, empirically validating previously unexplored scenarios. This foundational work provides essential tools for proactive risk assessment and designing resilient ML systems, moving the field forward significantly.

Heather Calloway (CISO) — STRONG ACCEPT

This SoK identifies a critical blind spot in ML security: the unintended interactions between defenses and risks. It provides a robust framework and guideline for proactively assessing these complex interdependencies, shifting the focus from isolated defense evaluation to a systemic understanding of ML risk at the governance level.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024