BounceAttack: A Query-Efficient Decision-based Adversarial Attack by Bouncing into the Wild

Jie Wan, Jianhao Fu, Lijin Wang, Ziqi Yang

IEEE Symposium on Security and Privacy 2024 · Day 1 · Continental Ballroom 5

Overview

In the rapidly evolving landscape of artificial intelligence, the robustness of machine learning models against adversarial attacks remains a critical concern. The talk "BounceAttack: A Query-Efficient Decision-based Adversarial Attack by Bouncing into the Wild," presented by Jie Wan from Zhejiang University, introduces a novel and highly effective method for generating adversarial examples in a black-box setting. This research focuses on traditional classification models, demonstrating how imperceptible perturbations can be added to legitimate inputs, causing a classifier to mispredict without altering human recognition of the original content.

Watch on YouTube

Visual summary for BounceAttack: A Query-Efficient Decision-based Adversarial Attack by Bouncing into the Wild by Jie Wan, Jianhao Fu, Lijin Wang, Ziqi Yang
Visual summary for BounceAttack: A Query-Efficient Decision-based Adversarial Attack by Bouncing into the Wild by Jie Wan, Jianhao Fu, Lijin Wang, Ziqi Yang

Key moments

  1. 0:00 Introduction to BounceAttack and adversarial samples
  2. 1:30 BounceAttack's core strategy: projecting onto boundary
  3. 2:00 Key components: binary search, bounce decomposition, momentum
  4. 4:00 Experimental setup and baseline methods
  5. 4:45 BounceAttack's superior query efficiency results
  6. 5:30 Outstanding attack success rate against defenses
  7. 6:10 Demonstrating query efficiency against HSJA
  8. 7:00 Visual attack demonstration and key conclusions

BounceAttack: A Query-Efficient Decision-based Adversarial Attack by Bouncing into the Wild

Speakers: Jie Wan, Jianhao Fu, Lijin Wang, Ziqi Yang

Conference: IEEE S&P

YouTube: https://www.youtube.com/watch?v=oF0b0A69byY

Overview

In the rapidly evolving landscape of artificial intelligence, the robustness of machine learning models against adversarial attacks remains a critical concern. The talk "BounceAttack: A Query-Efficient Decision-based Adversarial Attack by Bouncing into the Wild," presented by Jie Wan from Zhejiang University, introduces a novel and highly effective method for generating adversarial examples in a black-box setting. This research focuses on traditional classification models, demonstrating how imperceptible perturbations can be added to legitimate inputs, causing a classifier to mispredict without altering human recognition of the original content.

BounceAttack distinguishes itself by significantly improving the query efficiency and attack success rate of decision-based black-box attacks, particularly when aiming for minimal perturbations. The methodology leverages a unique "Bounce Decomposition" of the estimated gradient and incorporates momentum to navigate the adversarial space more effectively. By addressing the limitations of prior work, such as the widely-cited HSJA attack, BounceAttack sets a new benchmark for crafting stealthy and potent adversarial samples with a remarkably limited number of model queries, posing a substantial challenge to the current state of AI model defenses. This work is crucial for both understanding vulnerabilities in deployed AI systems and for driving the development of more robust and secure machine learning models.

Background

▶ Watch: Introduction to BounceAttack and adversarial samples (0:00)

The concept of adversarial examples has garnered significant attention in AI security research since its inception. These are inputs to a machine learning model that have been intentionally perturbed in such a way that they are misclassified by the model, despite appearing identical or nearly identical to legitimate inputs to a human observer. The generation of such examples highlights a fundamental fragility in many state-of-the-art machine learning models, particularly deep neural networks.

Adversarial attacks can be broadly categorized into white-box and black-box settings. In a white-box attack, the adversary has full knowledge of the target model's architecture, parameters, and gradients, allowing for highly effective gradient-based attacks. However, in real-world scenarios, attackers often lack this comprehensive knowledge. This leads to the more challenging and practical black-box attack setting, where the adversary can only query the model with inputs and observe its outputs (e.g., predicted class labels or confidence scores).

Within black-box attacks, decision-based attacks are particularly relevant. These attacks operate solely on the final decision (predicted label) of the model, rather than requiring access to confidence scores or intermediate activations. This makes them highly stealthy and applicable to a wide range of deployed models where only the final prediction is exposed. Prior work in this area, such as HSJA (HopSkipJumpAttack), has made significant strides in generating adversarial examples in the decision-based black-box setting. HSJA, for instance, starts with an adversarial sample and iteratively projects it onto the decision boundary to minimize the perturbation size. However, existing decision-based attacks often suffer from query inefficiency, requiring a large number of queries to craft effective adversarial samples, especially when targeting very small, imperceptible perturbations. This high query count can make detection easier and the attack slower to execute. Furthermore, some methods struggle to converge to a small perturbation size, leaving the adversarial examples more visually detectable.

The problem BounceAttack aims to solve is to overcome these limitations by developing a decision-based black-box adversarial attack that is both query-efficient and capable of generating small-sized perturbations with high attack success rates. The attack seeks to achieve this under common adversarial constraints, specifically the L2 or L-infinity norms, which quantify the magnitude of the perturbation and ensure its imperceptibility to humans. By addressing these challenges, BounceAttack represents a significant advancement in the field, pushing the boundaries of what is possible in black-box adversarial attacks.

Key Findings

▶ Watch: Key components: binary search, bounce decomposition, momentum (2:00)

BounceAttack introduces several critical advancements that significantly outperform existing black-box adversarial attack methods, particularly in terms of query efficiency and the ability to generate subtle perturbations. The core findings highlight its superior performance across diverse experimental setups:

Firstly, BounceAttack demonstrates a remarkable ability to reduce the perturbation size with a limited number of model queries. Within just 5,000 model queries, BounceAttack was shown to reduce the perturbation size by over 40% compared to other baseline attack methods across all experiments conducted. This efficiency is maintained even with increased queries; at 30,000 model queries, BounceAttack consistently holds the leading position in minimizing perturbation size. This finding is crucial because smaller perturbations are inherently more stealthy and harder to detect, making the adversarial examples more potent in real-world scenarios.

Secondly, the attack exhibits an outstanding attack success rate. Given the same perturbation threshold, BounceAttack consistently achieves the best attack success rate. As the number of model queries increases, its success rate becomes even more pronounced compared to other methods. The researchers quantified this by calculating the average attack success rate of other methods relative to BounceAttack, finding that competitors could at maximum reach only 87% of BounceAttack's success rate across all experiments. This indicates a robust and reliable attack vector.

Thirdly, BounceAttack proves effective even against various defense methods. The presentation highlighted that BounceAttack can always acquire "the best performance" even when target models are equipped with different defensive mechanisms. While specific defense types were not detailed in the transcript, this general statement suggests a resilience that many other attacks lack, further underscoring its potential threat.

A direct comparison with HSJA, a prominent baseline, provided clear insight into BounceAttack's success. In one instance, HSJA failed to generate an adversarial sample within 30,000 queries for a randomly chosen sample. However, when the attack method was switched to BounceAttack at 10,000 queries, it quickly converged to a distance smaller than the threshold, resulting in a successful adversarial sample. This direct evidence demonstrates that the Bounce Direction utilized by BounceAttack significantly improves both query efficiency and attack success rate over the gradient estimation strategies employed by HSJA.

Finally, BounceAttack is versatile, performing well under both L2 and L-infinity constraints and in both targeted and untargeted attack settings. In the untargeted setting, it can hide perturbations within a benign image using as few as 1,000 model queries from a visual perspective. For the more challenging targeted setting, it requires around 5,000 model queries to transform the benign image to visually resemble the target image, with the generated sample becoming increasingly similar to the benign sample as the attack progresses. These comprehensive findings establish BounceAttack as a state-of-the-art method for generating powerful and efficient black-box adversarial examples.

Technical Deep Dive

▶ Watch: BounceAttack's superior query efficiency results (4:45)

BounceAttack is a sophisticated decision-based black-box adversarial attack that achieves its superior performance through a novel approach centered on gradient estimation and search space navigation. The attack process is iterative, aiming to reduce the size of the perturbation added to a benign sample (denoted as $X_{star}$) to transform it into an adversarial sample ($X_T$) that causes misclassification. The expected optimal adversarial sample is denoted as $X_A$.

The attack strategy begins by selecting an initial adversarial sample. In an untargeted setting, this sample is chosen from a random class. In a targeted setting, it's chosen from a specific target class. The core of BounceAttack's efficiency lies in its three interconnected components: Bounce Search, Momentum, and Smooth Search.

  1. Bounce Search:

This is the foundational component responsible for the attack's query efficiency. The process starts by conducting a binary search to project the current adversarial candidate sample ($X_{T\_hat}$) onto the decision boundary of the classifier. This projection is crucial because adversarial examples often lie very close to this boundary.

The unique aspect here is the use of a Bounce Decomposition operation. This operation decomposes the estimated gradient along what the speaker refers to as the "oo direction" – likely implying an orthogonal or outward direction – from the benign sample ($X_{star}$) to the adversarial sample ($X_{real\_s}$). The goal of this decomposition is to find a direction that moves the projected sample as far as possible from the benign sample on the decision boundary while still being adversarial. The theoretical underpinning for this decomposition is provided in Theorem 1 of their paper, which suggests that fixing the step size in this direction maximizes the distance from the benign sample on the decision boundary. This "bounce direction" (BCE Direction) is key to efficiently reducing the perturbation size by guiding the search directly towards the boundary in a meaningful way.

  1. Momentum:

The gradient estimation process in black-box attacks often involves approximating gradients using random unit vectors. This can introduce significant randomness and lead to fluctuations in the search direction, hindering convergence. To mitigate this, BounceAttack incorporates momentum.

The momentum component smooths the search process by considering historical search directions. A hyperparameter, MU, controls the "memory" of these historical directions.

  • When MU is small, the search direction is largely influenced by the newly generated Bounce Direction from the current iteration. This allows for rapid adaptation to local gradients.
  • When MU is large, the search direction is heavily driven by the aggregated historical information from past iterations. This can lead to a more stable search path, preventing erratic movements.

The speaker notes that a "fast direction" (likely implying a strong influence from recent bounce directions) can lead to fast convergence of the adversarial sample on the decision boundary initially. However, it also carries the risk of the sample fluctuating back and forth across the optimal adversarial sample at the later stages of the attack, making fine-tuning difficult.

  1. Smooth Search:

To address the potential for fluctuation caused by the momentum-driven "fast direction" in the later stages of the attack, BounceAttack switches to a smooth search strategy. Once the adversarial sample has converged close to the optimal region, instead of continuously decomposing the gradient into the Bounce Direction, the attack directly uses the estimated gradient. This transition helps to stabilize the optimization process, allowing for more precise convergence to the smallest possible perturbation without overshooting or oscillating around the target. This adaptive strategy ensures both rapid initial convergence and fine-grained optimization in the final stages.

In essence, BounceAttack iteratively refines the adversarial sample by first projecting it onto the decision boundary using a binary search guided by the Bounce Decomposition of the gradient. It then moves the sample deeper into the adversarial space along this BCE Direction. Momentum is applied to stabilize and accelerate this process, and a smooth search phase is engaged to ensure precise convergence to minimal perturbation. This combination of strategies, particularly the novel Bounce Decomposition, significantly enhances query efficiency and allows for the generation of smaller, more imperceptible adversarial examples compared to previous methods like HSJA, which rely on less optimized gradient estimation techniques. The attack is designed to operate under both L2 and L-infinity constraints, making it versatile for different imperceptibility requirements.

Demo / Proof of Concept

▶ Watch: Outstanding attack success rate against defenses (5:30)

While the presentation did not feature a live, interactive demo in the traditional sense, the talk provided compelling evidence and visual illustrations that served as a robust proof of concept for BounceAttack's effectiveness. The experimental setup and results clearly demonstrated the attack's capabilities across various models and datasets.

The researchers extensively evaluated BounceAttack against widely-used deep learning models: ResNet-50 and DenseNet-121. These models were pre-trained on standard image classification datasets: CIFAR-10, CIFAR-100, and ImageNet. This choice of victim models and datasets represents a comprehensive testbed, covering different complexities and scales of image recognition tasks.

The performance of BounceAttack was compared against several baseline methods, all operating within the black-box adversarial attack setting. The core demonstration revolved around showing BounceAttack's superior query efficiency and ability to achieve small perturbation sizes and high attack success rates.

One key visual proof involved comparing the attack process of BounceAttack with HSJA on the same randomly chosen sample. The results clearly illustrated a scenario where HSJA failed to generate an adversarial sample within 30,000 model queries. In stark contrast, when the attack method was switched to BounceAttack at the 10,000 query mark, the attack quickly converged, achieving a perturbation distance smaller than the predefined threshold and thus successfully crafting an adversarial example. This visual comparison vividly demonstrated the query efficiency advantage of BounceAttack, attributing it directly to the effectiveness of the "Bounce Direction" over HSJA's gradient estimation.

Furthermore, the talk provided visual evidence of the generated adversarial samples, particularly in the context of targeted and untargeted attacks. For the untargeted setting, BounceAttack was shown to "hide the perturbation into the benign image" with as few as 1,000 model queries, meaning the changes were imperceptible to the human eye. In the more challenging targeted setting, BounceAttack required approximately 5,000 model queries to morph the benign image into one that visually resembled the target image, while still maintaining imperceptibility relative to the original. The speaker noted that "as the attack going on, the generated sample look more and more similar to the benign sample," indicating a gradual and controlled transformation that maintains stealth.

These experimental results, combined with the visual examples of perturbation sizes and convergence graphs, served as a powerful proof of concept, validating BounceAttack's claims of query efficiency, small perturbation generation, and high attack success rates across a variety of realistic scenarios and against established baselines.

Defensive Implications

▶ Watch: Visual attack demonstration and key conclusions (7:00)

The introduction of BounceAttack carries significant defensive implications for the security of AI systems. Its demonstrated ability to craft highly effective adversarial examples with minimal model queries and imperceptible perturbations poses a substantial challenge to existing and future defense mechanisms.

Firstly, the query efficiency of BounceAttack means that defenders have less time to detect and mitigate an ongoing attack. Traditional detection methods that monitor query rates or patterns might be less effective against an attack that achieves its goal with a relatively small number of interactions with the model. This necessitates the development of more sophisticated, real-time anomaly detection systems that can identify subtle adversarial patterns in inputs or model activations, rather than relying solely on query volume.

Secondly, BounceAttack's success against "different defense methods" (as stated in the transcript) indicates that many current adversarial defenses may not be robust enough against this new class of decision-based attacks. These defenses often include techniques like adversarial training, input transformations, gradient obfuscation, or ensemble methods. The fact that BounceAttack can overcome them suggests a need for re-evaluating their efficacy, particularly against attacks that leverage optimized search directions like the Bounce Direction. Defenders need to rigorously test their models against BounceAttack and similar query-efficient methods to identify specific vulnerabilities.

Thirdly, the generation of extremely small and imperceptible perturbations by BounceAttack makes detection even harder. If an adversarial example is visually indistinguishable from a benign input, human oversight or simple input sanitization techniques will fail. This pushes the burden onto automated systems to discern minute, malicious changes that are designed to exploit model blind spots. Developing robust feature extractors or perturbation detection mechanisms that can operate at this granular level becomes paramount.

Defenders should consider several strategies in light of BounceAttack:

  • Enhanced Adversarial Training: Current adversarial training regimens might need to be updated to include adversarial examples generated by decision-based black-box attacks like BounceAttack. Training with these more challenging examples could help models learn more robust decision boundaries.
  • Robust Feature Engineering: Investing in research that leads to models learning features that are inherently more invariant to small, adversarial perturbations, rather than relying on superficial patterns.
  • Input Pre-processing and Transformation: While some input transformations can be bypassed, more advanced and adaptive techniques that can effectively denoise or regularize inputs without degrading legitimate performance might offer some protection.
  • Attribution and Explainability Tools: Developing tools that can explain model predictions could potentially highlight when a model is relying on unusual or adversarial features, even if the input appears normal.
  • API Rate Limiting and Monitoring: While not a complete solution, intelligent rate limiting combined with behavioral analytics on model queries could still provide a first line of defense against even query-efficient attacks, especially if combined with other detection methods.

Ultimately, BounceAttack underscores the ongoing arms race in AI security. It highlights that models must not only be accurate but also demonstrably robust against sophisticated, real-world attack vectors that minimize interaction while maximizing impact.

Key Takeaways

  • BounceAttack is a novel, query-efficient, decision-based black-box adversarial attack that generates imperceptible perturbations to misclassify AI models.
  • It introduces a unique "Bounce Decomposition" operation for estimated gradients, guiding the search efficiently onto the decision boundary and along the BCE Direction.
  • The attack incorporates momentum and an adaptive smooth search strategy to stabilize the optimization process and ensure precise convergence to minimal perturbation sizes.
  • BounceAttack achieves significantly smaller perturbation sizes (over 40% reduction) and superior attack success rates compared to existing methods like HSJA, even against various defense mechanisms.
  • It is highly effective across diverse models (ResNet-50, DenseNet-121) and datasets (CIFAR-10, CIFAR-100, ImageNet) under both L2 and L-infinity constraints, in both targeted and untargeted settings.
  • The work highlights critical vulnerabilities in current AI models and defenses, necessitating the development of more robust adversarial training, detection, and mitigation strategies to counter such advanced, query-efficient attacks.

About the Speaker(s)

The primary speaker for this presentation was Jie Wan, a second-year PhD student from Zhejiang University. His research focus is explicitly stated as AI security. He was joined by co-authors Jianhao Fu, Lijin Wang, and Ziqi Yang, also associated with the research, though their specific roles or affiliations beyond the university were not detailed in the transcript. The collective work from this team at Zhejiang University contributes significantly to the field of AI security, particularly in understanding and mitigating adversarial threats to machine learning systems.

Reviews

Dr. Zero (Offensive Security Researcher) — STRONG ACCEPT

This talk introduces BounceAttack, a highly query-efficient decision-based black-box adversarial attack leveraging a novel "Bounce Decomposition" for gradient estimation. It significantly outperforms prior methods like HSJA in generating imperceptible perturbations, posing a critical challenge to current AI model defenses. This research offers crucial insights for both understanding vulnerabilities and developing more robust machine learning models.

Heather Calloway (CISO) — MUST SEE

This research on BounceAttack is a critical watch for any organization deploying AI. It demonstrates a highly efficient, stealthy black-box attack that fundamentally shifts the threat landscape for machine learning models by drastically reducing the queries needed to achieve misclassification. This demands an immediate re-evaluation of AI model robustness and defense strategies at a governance level.

→ Top-rated talks at IEEE Symposium on Security and Privacy 2024

All talks from IEEE Symposium on Security and Privacy 2024