SoK: Explainable Machine Learning in Adversarial Environments

Maximilian Noppel, Christian Wressnegger

IEEE Symposium on Security and Privacy 2024 · Day 2 · Continental Ballroom 5

Overview

In an era where machine learning (ML) models are increasingly deployed in critical applications, the demand for transparency and accountability has led to the rise of Explainable Artificial Intelligence (XAI). XAI methods aim to provide insights into why a model makes a particular decision, thereby fostering trust, enabling auditing, and facilitating debugging. However, as Maximilian Noppel and Christian Wressnegger highlight in their IEEE S&P talk, the very explanations designed to enhance model trustworthiness can themselves become targets for adversaries. Their Systematization of Knowledge (SoK) paper, "Explainable Machine Learning in Adversarial Environments," provides a comprehensive framework for understanding and classifying the burgeoning field of attacks against explainable systems.

Watch on YouTube

Visual summary for SoK: Explainable Machine Learning in Adversarial Environments by Maximilian Noppel, Christian Wressnegger
Visual summary for SoK: Explainable Machine Learning in Adversarial Environments by Maximilian Noppel, Christian Wressnegger

Key moments

  1. 0:00 Introduction and an explanation-aware adversarial example
  2. 0:50 Defining an explainable system and its components
  3. 1:22 Overview of three attack vectors: input, model, system manipulation
  4. 2:50 Three explanation-aware attack types: preserving, prediction-preserving, dual
  5. 3:55 Ways to alter explanations: untargeted, targeted, and semi-targeted attacks
  6. 5:00 Defensive strategies at training and operational (inference) time
  7. 6:10 Hierarchy of formal robustness notions for explainable AI
  8. 6:50 Conclusion, key takeaways, and open research questions

SoK: Explainable Machine Learning in Adversarial Environments

Speakers: Maximilian Noppel, Christian Wressnegger

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=Y1PaWjxc6PM

Overview

In an era where machine learning (ML) models are increasingly deployed in critical applications, the demand for transparency and accountability has led to the rise of Explainable Artificial Intelligence (XAI). XAI methods aim to provide insights into why a model makes a particular decision, thereby fostering trust, enabling auditing, and facilitating debugging. However, as Maximilian Noppel and Christian Wressnegger highlight in their IEEE S&P talk, the very explanations designed to enhance model trustworthiness can themselves become targets for adversaries. Their Systematization of Knowledge (SoK) paper, "Explainable Machine Learning in Adversarial Environments," provides a comprehensive framework for understanding and classifying the burgeoning field of attacks against explainable systems.

This talk meticulously dissects the landscape of explanation-aware attacks, where adversaries manipulate not just the model's prediction but also its explanation. This presents a significant challenge to the integrity of XAI, as it can lead to scenarios where a model appears fair or robust while covertly performing malicious actions, or where audit trails are deliberately obscured. Noppel and Wressnegger's work is crucial for researchers and practitioners alike, offering a structured understanding of attack vectors, types, and scopes, alongside a discussion of existing defensive strategies and a formal hierarchy of robustness notions. It underscores the urgent need for robust XAI, ensuring that explanations remain reliable even in the presence of sophisticated adversaries.

Background

▶ Watch: Introduction and an explanation-aware adversarial example (0:00)

The foundational premise of machine learning involves training a classifier (f) on a dataset (D) to make predictions. While powerful, these "black box" models often lack transparency, making it difficult to understand their decision-making process. This opacity is problematic in high-stakes domains like healthcare, finance, or security, where trust, accountability, and regulatory compliance are paramount. This is where Explainable Artificial Intelligence (XAI) enters the picture. An explainable system, as defined by the speakers, consists of a classifier f and an explanation method (H). The explanation method H takes the model and an input, then produces an explanation, often highlighting salient features or input regions that influenced the prediction (e.g., relevant pixels in an image).

The concept of adversarial examples is well-established in the prediction-only domain. Here, small, often imperceptible perturbations are added to an input, causing a model to misclassify it while remaining visually identical to a human. These are termed prediction-only attacks because their sole aim is to fool the classifier's output, with effects on explanations typically ignored. Noppel and Wressnegger's work extends this adversarial thinking to XAI, introducing the notion of explanation-aware attacks. These attacks specifically target the explanation component H (or the combined system S = (f, H)), aiming to manipulate, hide, or distort the insights provided by the explanation method. The underlying problem is that if explanations can be reliably manipulated, they cease to be trustworthy tools for auditing, debugging, or building user confidence, potentially undermining the entire purpose of XAI.

Key Findings

▶ Watch: Overview of three attack vectors: input, model, system manipulation (1:22)

The core contribution of Noppel and Wressnegger's SoK paper is a comprehensive classification framework for attacks against explainable systems, alongside an analysis of defensive strategies and a formalization of robustness notions. Their key findings can be summarized across several dimensions:

  1. Three Attack Vectors: They identified three primary ways adversaries can execute explanation-aware attacks:
  • Input Manipulation Attacks: Similar to traditional adversarial examples, but crafted to affect explanations.
  • Model Manipulation Attacks: Involving poisoning the training data, altering the model's code, or directly modifying model parameters.
  • System Manipulation Attacks: Where the adversary controls the entire explainable system, potentially using different models for predictions and explanations.
  1. Three Explanation-Aware Attack Types: Beyond prediction-only attacks, they categorize attacks based on their dual objectives concerning predictions and explanations:
  • Explanation Preserving Attacks: The goal is to fool the prediction while maintaining the original, accurate explanation. This is particularly relevant for hiding malicious activities like backdoors.
  • Prediction Preserving Attacks: The prediction remains correct, but the explanation is altered. This is the most common type observed, often used to mislead auditors about fairness or model behavior.
  • Dual Attacks: Both the prediction and the explanation are simultaneously manipulated. This is useful when the explanation is class-specific, requiring the explanation to align with the new, malicious prediction.
  1. Three Ways to Alter Explanations: When the goal is to change an explanation, adversaries can pursue:
  • Untargeted Attacks: The manipulated explanation should be maximally different from the clean explanation, measured by various distance metrics.
  • Targeted Attacks: The malicious explanation is forced to be equivalent to a specific, chosen target explanation (e.g., a square shape in an image).
  • Semi-Targeted Attacks: The target explanation is a function of the original clean explanation, offering more flexibility (e.g., an inverted version).
  1. Literature Analysis: Their classification revealed that most existing works focus on gradient-based explanations and often consider only one attack type at a time, indicating that attack scenarios are highly application-dependent.
  1. Defensive Strategies: They categorized defenses into two main groups:
  • Training-time Defenses: Aim to build robust models from the outset, including data validation, robust model architectures, robust training procedures, and post-training validation.
  • Operational Defenses: Applied at inference time, such as input manipulation detection, model monitoring, and crucially, input-output validation for explanations.
  1. Formal Robustness Notions: The term attribution robustness or attributional robustness is often overloaded. The paper provides a crucial hierarchy of formal robustness notions to clarify what "robustness against explanation-aware attacks" truly entails, showing how different papers implicitly or explicitly use these varying definitions.

In essence, the SoK paper provides a foundational taxonomy for understanding the vulnerabilities of XAI, highlights the limitations of current research, and points towards critical areas for future investigation, particularly in formalizing robustness and developing more comprehensive defenses.

Technical Deep Dive

▶ Watch: Ways to alter explanations: untargeted, targeted, and semi-targeted attacks (3:55)

The technical depth of the SoK paper lies in its systematic dissection of the adversarial space within XAI, providing a structured language to describe complex attack and defense methodologies. The speakers' framework is built upon the interaction between an ML classifier f and an explanation method H, forming the explainable system S.

Attack Vectors:

  1. Input Manipulation Attacks: These are perhaps the most intuitive, drawing parallels with traditional adversarial examples. The adversary perturbs the input data x to x' such that S(x') yields a desired malicious outcome. The key difference in an explanation-aware context is that the perturbation is crafted not only to affect f(x') (the prediction) but also H(f, x') (the explanation). The initial example of a bird classified as a frog with a square explanation, even though the input x' contains imperceptible noise, demonstrates this vector. Such attacks often leverage the gradients of the explanation method itself, similar to how adversarial examples use gradients of the loss function for the classifier.
  1. Model Manipulation Attacks: These attacks target the model f itself, often during its training or deployment phase.
  • Data Poisoning: Injecting malicious data into the training set D to subtly alter f's behavior and, consequently, H's output.
  • Code Poisoning: Modifying the model's architecture or training script.
  • Model Poisoning: Directly altering the model parameters post-training.

In prediction-only settings, backdooring attacks are common model manipulations, where a specific trigger (e.g., a spatial patch) leads to a target misclassification. In an explanation-aware context, the adversary might aim to keep the model's prediction robust for most inputs but ensure that the explanation becomes "unusable" or misleading for specific adversarial inputs or during a backdoor activation, thereby obscuring the backdoor trigger.

  1. System Manipulation Attacks: This vector assumes a more privileged adversary who operates the entire explainable system S. The adversary has control over both f and H, or even multiple fs and Hs. A notable example is fairwashing, where an adversary uses an unfair model for the actual prediction (f) but employs a separate, fair model (or a specifically crafted H) to generate explanations that falsely portray the system as fair. This is a sophisticated form of deception, where the explanation actively misleads auditors about the true underlying decision-making process.

Explanation-Aware Attack Types:

The speakers' taxonomy further refines these vectors by defining the intent of the adversary:

  1. Explanation Preserving Attacks: Here, the adversary's primary goal is to fool the prediction (f(x') != y_true) while ensuring the explanation for the adversarial input H(f, x') remains similar to the clean explanation H(f, x). This is particularly insidious for hiding backdooring attacks. If a backdoor trigger normally causes H to highlight a specific spatial patch as relevant, an explanation-preserving attack would ensure the prediction is fooled, but H does not highlight the trigger, thus bypassing defensive techniques that look for suspicious explanation patterns.
  1. Prediction Preserving Attacks: This is the inverse: the adversary wants to preserve the correct prediction (f(x') == y_true) but significantly alter the explanation (H(f, x') != H(f, x)). This is the most studied attack type, often with the intent to mislead an auditor. For example, an auditor might use H to determine if a loan application system is fair. An adversary could manipulate H to show that non-discriminatory features were crucial for a decision, even if discriminatory features were implicitly used by f.
  1. Dual Attacks: In this scenario, both the prediction and the explanation are under adversarial control (f(x') != y_true AND H(f, x') != H(f, x)). These attacks are especially potent when H is class-specific, meaning the explanation changes based on the predicted class. If an adversary changes the prediction (e.g., from "bird" to "frog"), they would then want the explanation to appear plausible for the new predicted class, rather than showing an explanation that still points to "bird" features.

Ways to Alter Explanations:

When an explanation needs to be changed (as in prediction-preserving or dual attacks), the adversary has several options for the target explanation:

  1. Untargeted Attacks: The goal is simply to make the adversarial explanation H(f, x') as different as possible from the clean explanation H(f, x). This is typically quantified using various distance metrics in the high-dimensional explanation space.
  1. Targeted Attacks: The adversary aims to force H(f, x') to conform to a very specific, pre-defined target explanation. The example of the square explanation for the frog prediction is a perfect illustration of a targeted attack. The square is an arbitrary, easily recognizable pattern chosen by the adversary.
  1. Semi-Targeted Attacks: This is a more nuanced approach where the target explanation is not fixed but is a function of the original clean explanation. For instance, an adversary might aim for an "inverting function," where features that were positive in the original explanation become negative in the adversarial one, or vice-versa. This allows for more dynamic and potentially less detectable manipulation.

The speakers note that most existing works concentrate on gradient-based explanation methods (like Grad-CAM, LIME, SHAP, etc.), likely due to their differentiability, which facilitates gradient-descent-based attack optimization. This focus points to a potential research gap in understanding attacks against other types of XAI methods. The complexity of defining attribution robustness is also highlighted, with the paper providing a hierarchy to address the ambiguous and often overloaded use of this term in literature, emphasizing the need for precise formal definitions to guide robust XAI development.

Demo / Proof of Concept

▶ Watch: Defensive strategies at training and operational (inference) time (5:00)

While the talk did not feature an extended live demonstration in the traditional sense, it opened with a compelling and highly illustrative proof-of-concept example that immediately set the stage for the entire discussion. Maximilian Noppel presented an image of a bird, which an adversarial system was manipulated to classify as a frog. Crucially, the explanation generated by the Grad-CAM method for this misclassified image was not a meaningful region related to a frog, nor was it the original bird features. Instead, the explanation method was fooled into highlighting an arbitrary, conspicuous square within the image.

This example served as a powerful visual demonstration of an explanation-aware attack, specifically a dual attack (as both prediction and explanation are altered) with a targeted explanation (the square). The speaker explicitly stated, "this attack can also be executed for other explanation methods and arbitrary target explanations." This initial illustration effectively communicated the core vulnerability: XAI methods, even widely used ones like Grad-CAM, can be manipulated to produce misleading or nonsensical explanations, thereby undermining their utility for trust and auditing. It visually underscored that adversaries can operate beyond merely altering a prediction, actively crafting the narrative provided by the explanation itself.

Defensive Implications

▶ Watch: Conclusion, key takeaways, and open research questions (6:50)

The findings presented by Noppel and Wressnegger carry significant implications for the design and deployment of secure and trustworthy explainable AI systems. Defenders can no longer assume that explanations derived from ML models are inherently reliable, especially in adversarial environments. The SoK paper systematically explores defensive strategies, categorizing them based on when they are applied:

Defenses at Training Time: These strategies aim to build robust explainable systems from the ground up, making them resilient to explanation-aware attacks even before deployment.

  1. Data Validation: Ensuring the integrity and cleanliness of the training dataset D. This is crucial against data poisoning attacks that might subtly alter model behavior and, consequently, explanation generation. Robust data validation can prevent an adversary from embedding "explanation backdoors" or influencing the model to produce misleading explanations.
  2. Robust Model Architectures: Designing or selecting model architectures that are intrinsically less susceptible to adversarial perturbations, both for predictions and explanations. Certain architectures might naturally yield more stable and interpretable explanations, making them harder to manipulate.
  3. Robust Training Procedures: Implementing training methodologies that enhance the model's resilience. This could include adversarial training specifically tailored to make explanations robust, not just predictions. For instance, training with explanation-aware adversarial examples could force the model to produce consistent explanations even for perturbed inputs.
  4. Post-Training Validation: After the model is trained, rigorous validation of its explanation capabilities is essential. This involves evaluating the quality, consistency, and fidelity of explanations across diverse datasets and perturbation types, ensuring they align with human intuition and domain knowledge.

Operational Defenses (at Inference Time): These defenses are applied when the explainable system is actively being used, monitoring inputs and outputs for signs of attack.

  1. Input Manipulation Detection: Similar to prediction-only scenarios, mechanisms can detect adversarial perturbations in incoming data. However, for explanation-aware attacks, these detectors must also consider the impact on H, not just f. An input might be clean enough for f but trigger a malicious explanation from H.
  2. Model Monitoring: Continuously observing the behavior of the deployed model f and explanation method H. This involves tracking changes in prediction distributions, explanation patterns, and the correlation between inputs, predictions, and explanations. Anomalies in explanation behavior could signal a model manipulation attack (e.g., a backdoor being activated or a fairwashing attempt).
  3. Input-Output Validation: This is particularly relevant for explanation-aware attacks. It involves validating the explanation H(f, x) in relation to the prediction f(x) and the input x. For example, if a model predicts "cat," the explanation should highlight cat-like features. If it highlights irrelevant background or a pre-defined square, this discrepancy can be flagged. This validation can be rule-based, anomaly-detection based, or even involve secondary, trusted explanation methods for cross-verification.

A critical takeaway for defenders is the need for a precise understanding of attribution robustness. The paper's proposed hierarchy of formal robustness notions is vital here. Defenders must specify exactly what kind of robustness they aim for (e.g., robustness against explanation-preserving attacks, or targeted explanation manipulation) to choose and implement appropriate defensive measures effectively. Ultimately, the work underscores that relying on XAI for trust, auditing, or compliance without explicitly considering its adversarial vulnerabilities is a perilous approach. Future defensive research and engineering efforts must integrate explanation robustness as a first-class requirement.

Key Takeaways

  • XAI Systems are Vulnerable: Explainable Machine Learning (XAI) systems are not inherently trustworthy; they are susceptible to sophisticated explanation-aware attacks that manipulate not only predictions but also the explanations themselves.
  • Diverse Attack Vectors and Types: Adversaries can employ input manipulation, model manipulation, or system manipulation to launch attacks. These attacks can be explanation preserving (fooling prediction, keeping explanation), prediction preserving (keeping prediction, changing explanation), or dual attacks (changing both prediction and explanation).
  • Explanations Can Be Targeted or Untargeted: When altering explanations, adversaries can aim for untargeted maximal difference, targeted specific patterns (e.g., a square), or semi-targeted changes based on the original explanation.
  • Defenses Require a Holistic Approach: Effective defenses must be considered at both training time (data validation, robust architectures/training) and inference time (input/model monitoring, crucial input-output validation for explanations), specifically accounting for explanation manipulation.
  • Formal Robustness is Essential: The concept of attribution robustness is often ambiguous. A formal hierarchy of robustness notions is needed to precisely define and measure resilience against explanation-aware attacks, guiding the development of robust XAI.
  • Application-Dependent Incentives and Open Questions: The specific adversarial incentives are highly dependent on the application scenario. The field of robust XAI is still nascent, with many open research questions remaining regarding comprehensive attacks and defenses beyond gradient-based explanation methods.

About the Speaker(s)

The talk "SoK: Explainable Machine Learning in Adversarial Environments" was presented by Maximilian Noppel as joint work with Christian Wressnegger. Maximilian Noppel, representing KIT (Karlsruhe Institute of Technology), introduced the paper and guided the audience through the complex landscape of adversarial XAI. Christian Wressnegger is a co-author of the paper, contributing to the comprehensive Systematization of Knowledge. Their collaboration highlights expertise in both machine learning and security, specifically focusing on the vulnerabilities and robustness challenges inherent in explainable AI systems.

Reviews

Dr. Zero (Offensive Security Researcher) — MUST SEE

Noppel and Wressnegger deliver a critical SoK, meticulously mapping the adversarial landscape of Explainable AI. Their framework for classifying explanation-aware attacks and formalizing robustness notions is foundational, exposing how explanations themselves become targets. This work is essential for anyone deploying or researching XAI in real-world, hostile environments.

Heather Calloway (CISO) — STRONG ACCEPT

This SoK systematically maps a critical emerging risk: the manipulation of AI explanations themselves. It provides a vital taxonomy for understanding how adversaries can undermine trust, accountability, and regulatory compliance in ML systems, offering clear implications for security governance and defensive strategy.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024