Reinforcement Learning Inspired Black-box Adversarial Attacks for Computer Vision
arXiv:2609.24249v1 [cs.LG] 21 Sep 2026
Florian Krone Elena Hoemann Sven Hallerbach Institute for AI Safety and Security German Aerospace Center (DLR) Sankt Augustin, Germany {florian.krone, elena.hoemann, sven.hallerbach}@dlr.de
µ
Actor
Environment
δ ∼ N (µ, σ)
Perturbation δ
x′ = clip(x + ϵ ∗ tanh δ, 0, 1)
σ
Replay Buffer
Critic
r̂
(δ, r)
Perturbations [δ, . . . ] Rewards [r, . . . ]
r = max log(Ck (x′ )) − log(Cy (x′ )) k̸=y
Figure 1. Visualization of the RIBA algorithm. The algorithm is inspired by a soft actor-critic, but tailored for the specific case of generating adversarial perturbations. A perturbation is generated by the actor, using the reparameterization trick. The actor is trained against a critic, which learns to evaluate adversarial perturbations.
Abstract Neural networks, both convolution or transformer based, are essential for modern computer vision systems. However, they are vulnerable to small perturbations, almost imperceptible to humans, which significantly alter the model’s prediction. These adversarial attacks are often considered to be a significant threat to the implementation of neural networks in safety-critical applications. Most attacks utilize the white-box threat model and therefore require full access to the target model, making them unrealistic to use in practice. We propose a novel approach under the more realistic black-box threat model that utilizes concepts from reinforcement learning to optimize perturbations with a nondifferentiable target model. Reinforcement learning algorithms have already been optimized to be query efficient, making them an ideal starting point when designing blackbox adversarial attacks. We show the success of our re-
inforcement learning inspired black-box adversarial attack (RIBA) in generating adversarial perturbations using only a small number of queries to the target model, by comparing it to state of the art attacks on different models on the Cifar10 and ImageNet data sets. RIBA takes 25.4% fewer median queries to generate attacked images against a ResNet18 on Cifar10 and 22.5% fewer median queries to fool a Vit-B/16 model on ImageNet. Additionally, we demonstrate that RIBA can match the performance of white-box attacks on an adversarially trained model.
1. Introduction Neural networks have shown remarkable success in computer vision tasks and are therefore increasingly applied in the real world, including safety-critical systems, such as automated vehicles. However, it was shown that they are susceptible to adversarial attacks [25], i.e., perturbations to the
input image that are insusceptible to humans. Adversarial attacks can be categorized by the information they require about the targeted model. The white-box threat model requires full access to the model and its weights, making it possible to utilize gradient information in calculating the perturbations. The black-box model, in contrast, treats the model as an oracle, where only the predictions for inputs are available. Two sub categories of the black-box model exist, where the attacker knows either just the final decision of the model or the scores of all classes, as is the case with our attack. Attacks designed for the black-box threat model have a more widespread use, as they for example can be used to attack AI models deployed behind a web API. The web API use case also motivates an important secondary goal of black-box attacks. Since every query to the API is typically associated with a cost, black-box attacks aim to find perturbations with as little queries as possible. To quantify if an attack is insusceptible to a human, a norm bound is applied to the perturbation, limiting the difference between the benign and the attacked sample. Our approach utilizes the l∞ -norm, which is commonly used to limit the perturbation [9, 15]. The main contributions of this paper can be summarized as follows: • We formulate the problem of generating adversarial perturbations as a Markov decision process (MDP). This allows for the application of a reinforcement learning algorithm to generate adversarial perturbations. • We propose a novel algorithm to solve this specific MDP based on soft actor-critic (SAC) [11]. • The success of our novel approach is demonstrated on standard benchmarks and compared to state of the art attacks. To highlight the close relation of our approach to reinforcement learning (RL), we name it reinforcement learning inspired black-box adversarial attack (RIBA).
2. Related work Black-box adversarial attacks, decision or score based, can be built in a number of different ways. In the following section, we introduce approaches that are gradient, randomsearch, or transfer based, as well as other RL based approaches. A common approach is to utilize concepts from whitebox attacks like PGD [22] or C&W [4]. However, these attacks require access to the gradient to optimize the perturbations. Black-box attacks, therefore, need to estimate the gradient first. Notable attempts are ZOO [5], ZO-NGD [29], ZO-AdaMM [6] and NES [15]. In the Bandits attack [16], the query count to the attacked model was reduced by integrating prior knowledge about the gradients into the estimation. Specifically, a time-dependent prior, reusing knowledge from previous steps of the attack optimization and a
data-dependent prior, using local similarity within images, were introduced. A different class of attacks utilizes random search to optimize adversarial perturbations. These attacks are specifically built for the black-box threat model. Guo et al. [10] introduced a random-search based attack called SimBA that randomly samples a vector from an orthonormal base and adds or subtracts it to the image. Another random-search based attack was introduced by Andriushchenko et al. [1], who iteratively add squares of different sizes and colors to the image until the target model is fooled. Consequently, their attack is called Square Attack. Bai et al. [2] propose to reduce the search space by utilizing an encoder/decoder setup in their NP-Attack. The original image is first encoded into a low dimensional space, where a random perturbation is sampled and added to the image. The perturbed image is then decoded and it is tested whether the perturbation fools the target model. Adversarial attacks are known to transfer between models to a certain degree, i.e., a perturbation generated for an image against one model might also fool a different model on the same image. This property can be exploited to attack a black-box model, utilizing a white-box surrogate model and transferring the perturbation. One example is the EigenBA attack [30]. Yang et al. [28] additionally query the targeted model in the LeBA attack to update the surrogate model continuously. The Simulator attack [20] works in a similar manner, by training and continuously updating a model that simulates the response of the target model. To improve the generalization, the SVRE [27] attack uses an ensemble of surrogate models. A few approaches have applied RL before to generate adversarial examples. Huang et al. [14] proposed their DBAR attack. They formulate the decision-based black-box adversarial attack problem as an MDP and utilize an RL algorithm to solve it. Kang et al. [17] additionally employ an autoencoder to reduce the dimensionality of the search space. However, these approaches utilize general RL algorithms to solve the problem, that disregard problem specific knowledge, which could be used for further optimizations. With this work, we aim to address this issue. A different paradigm of RL-based attacks exists, in which an agent is first trained in a dedicated training phase and than later applied to generate perturbations. Examples are RLAB [24] and QTRL [21]. However, this setup differs from the typical goal of black-box attacks, which usually aims to generate perturbations independent of each other, with as few queries as possible. Our approach is in line with the typical goal of black-box attacks.
3. Method The main contribution of this paper is a novel approach to build adversarial perturbations using a method inspired by
RL, while requiring only black-box access to the model. The application of RL requires the problem of optimizing an adversarial perturbation to be formulated as an MDP, which is outlined in the following section. Section 3.2 describes the RL algorithm we used and the problem specific changes we made to it.
3.1. Adversarial attacks as a MDP Definition 1 Given a classifier C : [0, 1]H×W ×3 → RN and a sample, label pair x ∈ [0, 1]H×W ×3 , y ∈ RN with C(x) = y, then δ ∈ RH×W ×3 is an adversarial perturbation, if C(x + δ) ̸= y and ∥δ∥∞ < ϵ for a budget ϵ. Definition 2 Given a set of states S, a set of actions A, a transition probability function P : S × A × S → [0, 1], and a reward function R : S×A×S → R, the tuple (S, A, P, R) is called a Markov decision process (MDP). The transition function P : S × A × S → [0, 1] denotes the probability to reach state si+1 ∈ S, if action a ∈ A is taken in state si ∈ S. Formally, it is defined as P (si , a, si+1 ) = Pr(si+1 | si , a).
(1)
Consequently, the reward function R : S × A × S → R describes the reward obtained for taking action a ∈ A in state si ∈ S and reaching state si+1 ∈ S. The goal of reinforcement learning is to optimize an agent that, given a sate s, predicts the optimal next action a, so that the reward over a series of state-action-state triplets, called an episode, is maximized. Connecting this to Def. 1, we want to train an agent that, given a sample x, creates an adversarial perturbation δ. The state is given by the current image x ∈ S, where S = [0, 1]H×W ×3 . The actions are given by the adversarial perturbations. We define a deterministic transition function, where, for an action δ, the state does not actually change, i.e., we are still considering the same image. The formal definition is given by P (xi , δ, xi+1 ) = Pr(xi+1 | xi , δ) ( 1 if xi+1 = xi , = 0 otherwise.
(2) (3)
Finally, the reward function depends on the model we are attacking. Considering a classification model C : [0, 1]H×W ×3 → [0, 1]N with N classes, we define Ck (x) as the probability of the k-th class, for a given x ∈ [0, 1]H×W ×3 . Additionally, we define a function p to apply a perturbation to an image. p(x, δ) = clip(x + ϵ ∗ tanh δ, 0, 1).
(4)
The clip functions limits the output to the valid range of images, i.e., each pixel in [0, 1], while the term ϵ ∗ tanh δ
ensures that the perturbed images are within the l∞ -norm bound around the original image. A perturbation δ̂ according to Def. 1 is obtained by δ̂ = p(x, δ) − x. The reward for an action δ ∈ A in state xi ∈ S that leads to the next state xi+1 ∈ S is then given by R(xi ,δ, xi+1 ) = max log(Ck (p(xi , δ))) − log(Cy (p(xi , δ))), (5) k̸=y
where y is the ground truth of the attacked image. The reward function is inspired by the loss function used in [5] to generate adversarial perturbations. We removed the parameter κ, which was originally used to control the strength of the perturbation. We are aiming to find a perturbation with as little queries as possible and consequently take the first perturbation that is found. The parameter κ would therefore not have any effect. The log function helps to emphasize small changes to the model’s output made by the perturbations, therefore boosting the sample efficiency. With this setup, it is already possible to generate adversarial perturbations. The next section describes the specific setup we used, and the changes we made to the RL algorithm to better suit our specific problem.
3.2. Specific RL algorithm The setup in Sec. 3.1 can be used to train an RL agent that is able to generate adversarial perturbations for different inputs. However, our goal was to build a typical black-box adversarial attack that does not require any pretraining, but calculates a perturbation for an individual image based on a limited number of queries to the targeted model. Therefore, we train an individual RL agent for each image and terminate the training, as soon as a perturbation is found. Additionally, we terminate each episode after one step, as according to Eq. (3), the state does not change. To train the RL agents, we use a modified version of soft actor-critic (SAC) [11]. One of the key concepts of off-policy RL, like SAC, is to separate the collection of experience from the training, using a replay buffer. First, a number of actions are executed in the environment to gain experience. This is stored in the replay buffer. Afterwards, batches from the replay buffer are sampled for training. Note that the training is not necessarily done on the most recent experience. We train a critic Qθ (δ) that estimates the expected reward for a perturbation δ. The critic is trained using JQ (θ) = E(δ,r)∼D [(Qθ (δ) − r)2 ],
(6)
where D represents the replay buffer. Known perturbation, reward pairs are sampled from the replay buffer and the critic is trained to estimate the reward. The simplicity of this task allows us to use a very small network for the critic, consisting of a fully connected network with a singe hidden layer. The second part of our algorithm is the actor,
Algorithm 1 RIBA
the sample efficiency, the critic and policy are each updated multiple times before a new perturbation is tried against the target model. Figure 1 provides a visualization of the algorithm. While our RIBA algorithm is inspired by SAC, a number of modifications have been made to better suit the specific problem at hand. Most importantly, we removed the dependence on the state from both the critic and the policy. Typically, the policy is a conditional distribution πϕ (· | s) over the actions, given the current state s. In this case, both µϕ (s) and σϕ (s) are neural networks predicting the parameters for the normal distribution based on the state. The critic Qθ (s, a) also depends on the state, estimating the expected return if action a is executed in state s. The return of an episode describes the discounted sum of all rewards in that episode. In our case, each episode is exactly one step long. Therefore, the return is given by the reward for that step. This is an import observation, as it removes the need for target Q-functions to train the critic, which simplifies the training objective.
Input: sample-label pair x, y Output: adversarial perturbation δ 1: initialize πϕ and Qθ 2: D ← ∅ 3: for each iteration do 4: δ ∼ πϕ 5: r ← max log(Ck (p(x, δ))) − log(Cy (p(x, δ))) k̸=y
6:
if argmax Ck (p(x, δ)) ̸= y then k
return δ end if D ← D ∪ {(δ, r)} for each critic update step do θ ← θ − η∇θ JQ (θ) 12: end for 13: for each actor update step do 14: ϕ ← ϕ − η∇ϕ Jπ (ϕ) 15: end for 16: end for 7: 8: 9: 10: 11:
4. Experiments which contains a policy πϕ used to generate perturbations. The policy utilizes the reparameterization trick, it consists of two sets of learned parameters µϕ and σϕ used to parameterize a normal distribution. The perturbations are then sampled from the distribution. The training objective for the policy is given by Jπ (ϕ) = Eξ∼N [α log πϕ (fϕ (ξ)) − Qθ (fϕ (ξ))].
(7)
The objective uses the reparameterization trick with fϕ (ξ) = µϕ + σϕ ξ, ξ ∼ N (0, 1).
(8)
The training objective for the policy contains two separate terms. The term Qθ (fϕ (ξ)) trains the policy to maximize the reward, using the critic as an estimator for the reward. Using the critic here is a key part of SAC, as this allows to train the policy without directly interacting with the environment, i.e., querying the target model in our case. The term log πϕ (fϕ (ξ)) forces the policy to have a high entropy. The parameter α is used to balance the two objectives. We do not set α directly, but optimize it as introduced in [12]. Algorithm 1 provides an overview, how we generate adversarial perturbations. Just like in SAC, our algorithms alternates between collecting experience in the environment, which is then stored in the replay buffer and updating the critic and policy. In our case, collecting experience in the environment requires calculating the reward using Eq. (5) and checking if a perturbation was found. The perturbation, reward tuple (δ, r) is then stored in the replay buffer. Afterwards, the critic and policy are updated using Eq. (6) and Eq. (7) respectively with a learning rate η. To increase
To evaluate the effectiveness and query efficiency of our approach in generating adversarial perturbations against black-box models, we attack models trained on Cifar10 [18] and ImageNet [23]. For the evaluation on Cifar10, we use a ResNet-18 [13] model. On ImageNet, we use an Inceptionv3 [26] and a ViT-B/16 [8] model to investigate the differences between attacking convolutional and transformer based architectures. Table 1 shows the experimental setup and hyperparameters. All hyperparameters are optimized targeting a ResNet-18 model on Cifar10 and used across all experiments. To generate a perturbation of size 32 × 32 × 3, the critic is a standard MLP Q : R3072 → R with a single hidden layer of size 6. The actor consists of a tuple R3072 × R3072 which is used to generate actions according to Eq. (8). We utilize the AdamW [19] optimizer with a learning rate of 0.02. The learning rate is decayed by a factor of 0.1 if the reward hits a plateau. To boost the sample efficiency, both the actor and the critic, are updated multiple times per query to the target model. This is one of the core functionalities of off-policy RL and the main reason for the query efficiency of our attack. In particular, we update the actor 120 times and the critic 40 times per query to the target model. To boost the initial exploration, we initialize the replay buffer with 5 randomly sampled perturbations. The amount of samples is a trade-off between improved exploration and an increase in queries, as each sample comes at the cost of one query to the target model. The trade-off between the entropy of the policy and maximizing the reward in Eq. (7) is optimized according to [12]. Typically, the target entropy is set to −n, where n is the number of dimensions in the action space. In our case this would be
Table 1. Experimental setup and hyper parameters
Parameter
Description
Value
max queries
Maximum allowed queries Perturbation budget under l∞ norm Number of critic updates per query Number of actor updates per query Learning rate Add random samples to the replay buffer before Alg. 1 parameter in Eq. (7) Optimize α according to [12] Single hidden layer of the critic Batch size for critic updates Batch size for actor updates
10000 0.031 / 0.05
critic update steps actor update steps η start samples initial α target entropy critic hidden dimension critic batch size actor batch size
ASR
Avg. Queries
Median Queries
99.81% 100% 100%
233 147 107
116 67 50
Attack SimBA Square Attack RIBA (ours)
40 120 0.02
1.0
5 0.8
0.1 -3 6 400 650
−3072, far greater than for example in a typical robotic application. We found using the number of channels, i.e., −3, to lead to better results. We compare our approach to a number of state of the art black-box attacks that do not require any additional information such as surrogate models or pretraining against the target model. We use the open sourced implementations and suggested hyperparameters of the attacks to evaluate them against the same models and images we used for our attack. As it is common practice, the maximum amount of queries is limited to 10000 and an attack is considered as failed if no perturbation was found in this limit. The reported average and median amount of queries required are only calculated for the successful attacks.
4.1. Cifar10 Evaluation For the evaluation of our attack on Cifar10, we use a ResNet-18 model with a benign accuracy of 94.98%. We limit the perturbations to the typical budget of ϵ = 8/255 and evaluate the attacks on all 9498 correctly classified images of the test set. Table 2 shows the attack success rate (ASR), as well as the average and median amount of queries required to generate a perturbation that fools the model. It was already possible to reliably generate perturbations under this setup with the Square Attack. However, our RIBA algorithm requires significantly less queries to the target model, both on average and median. Figure 2 displays the success rates of RIBA and Square Attack, when limiting the allowed queries to a lower number. RIBA can achieve a higher success rate across the board on all limits.
0.6
ASR
ϵ
Table 2. Experimental evaluation of attacking a ResNet-18 model trained on Cifar10 with a perturbation limit of 8/255 under the l∞ norm.
0.4 0.2 0.0
0
200
400
600
800
1000
Queries RIBA
Square Attack
Figure 2. Comparison of the attack success rate of RIBA and Square Attack at different query limits when attacking a ResNet18 model on Cifar10.
4.2. ImageNet Evaluation For the evaluation of our attack on ImageNet, we use an Inception-v3 model with a benign accuracy of 77.294% and a ViT-B/16 model with a benign accuracy of 81.072%. We limit the perturbations to the typical budget of ϵ = 0.05 and evaluate the attacks on a subset of 1000 randomly sampled, correctly classified images from the validation set. To reduce the computational load of generating perturbations for the higher resolution of ImageNet with our RIBA algorithm, we calculate the perturbations at the lower resolution of 32 × 32 and scale them up to the image size. For the even larger images required by the Inception model of 299×299, we found that it works better to first repeat the perturbation to a size of 64 × 64 before scaling it up to the image size. This is the only hyperparameter that is tuned model, or rather input size, specific. All other hyperparameters are taken from the optimization on Cifar10. The attacks we use as a comparison are unaffected by this decision. Table 3 shows the attack success rate and average queries required to fool the models. While our approach has state of the art attack success rate, it can not quite match the average
Table 3. Experimental evaluation of attacking an Inception-v3 and a ViT-B/16 model trained on ImageNet with a perturbation limit of 0.05 under the l∞ norm.
Attack SimBA Square Attack RIBA (ours)
ASR
Avg. Queries
Inception-v3 Median Queries
ASR
Avg. Queries
ViT Median Queries
84.8% 99.6% 99.2%
1567 221 239
1021 31 66
80.4% 99.9% 100%
1154 204 311
690 102 79
Table 5. Experimental evaluation of attacking an adversarially trained WideResNet-82-8 on Cifar10 with a perturbation limit of 8/255 under the l∞ norm.
1.0 0.9
Attack
ASR
0.7
FGSM PGD RIBA (ours)
13% 18% 18%
ASR
0.8
Avg. Queries
Median Queries
1117
414
0.6 0.5
0
1000
2000
3000
4000
5000
Queries RIBA
Square Attack
Figure 3. Comparison of the attack success rate of RIBA and Square Attack at different query limits when attacking an Inception-v3 model on ImageNet. Table 4. Experimental evaluation of the 5% most difficult to attack samples on Inception-v3 for Square Attack and RIBA.
Attack Square Attack RIBA
Avg. Queries
Median Queries
2779 2069
2442 1484
queries of the highly efficient Square Attack. Figure 3 explores where this difference comes from in case of the Inception model. It shows the attack success rates achieved, when limiting the queries to different numbers. The Square Attack can achieve high success rates with very few queries, testimony of its highly optimized initialization. However, it falls behind for the samples harder to manipulate. For these samples, our approach is able to achieve higher success rates with fewer queries, indicating the superiority of RIBA for these difficult samples. This is further backed by Tab. 4, which shows the average and median queries for the 5% most difficult to attack samples for Square Attack and RIBA. RIBA clearly has the edge over Square Attack, taking significantly fewer queries for difficult samples. For
the ViT model, our RIBA algorithm can significantly reduce the median queries required for a successful attack. This indicates that the initialization of Square Attack, which it heavily relies on for its query efficiency, is not efficient on transformer based models. Notably, the initialization was built specifically for convolutional networks, as vision transformers were introduced after the Square Attack was released.
4.3. Robust Models To evaluate the effectiveness of RIBA in attacking adversarially trained models, we attack a WideResNet-82-8 trained by Bartoldson et al. [3] on Cifar10. The model is one of the top models in the RobustBench leaderboard [7]. We limit the perturbation budget to ϵ = 8/255, the budget the model was trained for, and evaluate RIBA on 100 randomly sampled, correctly classified images from the Cifar10 test set. To put our results into perspective, we compare it with two white-box attacks. The simple FGSM [9] and the powerful PGD [22] attack. The PGD attack was limited to the same 10000 queries, i.e., steps, that our attack was limited to, albeit with gradient information. Table 5 shows the result of the experiment. Our RIBA attack was able to outperform the attack success rate of the FGSM attack an match the PGD attack.
4.4. Ablation Studies The RIBA algorithm is inspired by SAC and can be seen as a simplified, problem-specific version of it. Most of the design decisions are therefore justified by this relationship. An important design decision that was made is the reward function of the MDP specified in Eq. (5), especially the use
False True
ASR
Avg. Queries
99% 100%
579 107
Starting Samples 1 5 10 20
100% 100% 100% 100%
110 107 110 115
100% 100%
306 107
100% 100% 100% 100%
120 115 107 108
Target Entropy -3072 -3 Actor Update Steps 20 60 120 160
of the log function. Table 6 shows the effect of this decision, as well as the effects of some hyperparameters. We generate perturbations for 200 randomly sampled, correctly classified images from Cifar10 targeting a ResNet-18 model and repeat the experiment three times with different seeds. One parameter at a time is altered to show its effect, while the other parameters remain at their optimized value. Introducing the log function inside the reward significantly reduces the number of queries required from 579 to just 107. It amplifies small changes in the model’s output and therefore helps the exploration. Increasing the number of starting samples can also help the exploration, as these randomly sampled perturbations are used in the initial training. Each sample comes at the cost of one query to the target model. Therefore, an increase in starting samples must reduce the average required queries by at least this amount. Table 6 shows that for an increase from 1 to 5 starting samples, this is the case. Further increasing the number of starting samples to 10 does not sufficiently reduce the amount of queries during training to justify the increase. As discussed above, the target entropy is typically set to the negative number of dimensions in the action space, in our case −3072. However, we found the negative number of channels in the action space, i.e., −3, leads to better results. This is backed by Tab. 6, showing that it requires less than half of the average queries. Finally, we look into the effect of increasing the actor update steps. This hyperparameter specifies how many
1000
100
750
75
500
50
250
25
0
10
20
30
40
50
Batched Queries
Log Reward
Total Queries
Table 6. Experimental evaluation of selected design decisions and hyperparameters
60
Batch Size Total Queries
Batched Queries
Figure 4. Visualization of how batched queries to the target model influence the amount of queries needed in total, as well as the amount of batches.
times the policy is updated per query to the target model. It is clearly visible that, up to a certain point, an increase in update steps can significantly reduce the average number of queries, albeit with diminishing returns. However, too large values can lead to an increase in queries, as the actor is trained on increasingly outdated information.
4.5. Batched Inputs When attacking a model in a real world use case, an attacker might not be interested in generating perturbations against thousands, but rather against very few select images. In this section, we explore a unique possibility of our approach to further reduce the amount of times the target model is queried, when generating a perturbation for a single image. One of the major features of off-policy RL is the separation of the collection of experience from the training. Our RIBA algorithm follows the same paradigm. In Alg. 1, perturbations are sampled and evaluated and then added to the replay buffer starting at Line 4. This requires querying the target model. Afterwards, the actor and critic are trained using past experiences from the replay buffer starting at Line 10. In SAC, it is possible to execute multiple steps in the environment and add each transition to the replay buffer, before training again [11]. The same is possible in RIBA. However, in our case, the next perturbation is not dependent on the previous one, like it would be for a typical MDP. Therefore, we can generate multiple perturbations at once and query the target model in one batch. To investigate the effect of trying multiple perturbations at once, we utilize the same setup as for the ablation studies, i.e., we generate perturbations for 200 randomly sampled, correctly classified images from Cifar10, targeting a ResNet-18 model, and repeating the experiment with three different seeds. Figure 4 shows the effect of increasing the batch size used to
query the target model from 1 up to 64. The total amount of queries required increases linearly with an increased batch size. However, the amount of batches, i.e., the number of times, the target model is actually called, can be significantly reduced. Increasing the batch size from 1 to 8 can reduce the number of batched queries from 107 to 24. Further increasing the batch size reduces the number of batched queries even further. However, we observe diminishing returns with larger batch sizes, albeit with only 13 batched queries required at a batch size of 64. This shows a significant advantage of RIBA in a realistic use case over other attacks that do not support batched queries when attacking a single image.
5. Discussion We successfully demonstrated the capabilities of our novel approach in the previous section. Specifically, we showed that our approach can attack different models on different data sets, while only requiring a small number of queries to the target model. We were able to show that our approach performs similarly to the long standing state of the art Square Attack. RIBA clearly outperforms Square Attack on the Resnet18 model trained on Cifar10. For the more difficult case of models trained on ImageNet, the results are not as clear. Due to its initialization, which is highly optimized for fooling convolutional networks, Square Attack can generate perturbations with fewer queries than RIBA on most samples against the Inception-v3 model. However, RIBA performs significantly better on the more difficult samples. On the transformer based ViT-B/16 model, RIBA takes significantly fewer median queries than Square Attack, indicating that the initialization of Square Attack does not work well on transformer based architectures. Both attacks have their strength in different areas. Notably, they follow entirely different paradigms, an RL inspired optimization and random search. It is of utmost importance that different attack strategies are known to the research community, in order to properly defend models against all types of attacks. We designed the RIBA algorithm for the l∞ threat model. However, an adaptation to the l2 threat model should be possible. The l∞ -norm bound is enforced statically due to the way perturbations are applied to the images in Eq. (4). This is not possible for the l2 -norm bound. However, the l2 norm of the perturbation could be added as a penalty to the reward function, or directly added to the training objective of the policy in Eq. (7) as a penalty. We leave it for future work to explore these possibilities.
6. Conclusion In this paper, we introduced a novel approach to optimize adversarial perturbations under the realistic back-box threat
model. First, we formulated the problem of optimizing adversarial perturbations as a Markov decision process, which would allow the use of any reinforcement learning algorithm to optimize perturbations. Then, we developed the novel reinforcement learning inspired black-box adversarial attack (RIBA) algorithm, tailored to specifically solve this MDP. RIBA is based on the concepts of the soft actorcritic (SAC) algorithm. We showed the success of our attack by comparing it to the long standing state of the art Square Attack. RIBA requires significantly fewer queries to generate perturbations on the Cifar10 data set to fool a ResNet-18 model than the Square Attack. Additionally, it requires fewer median queries to fool the transformer based ViT-B/16 on ImageNet. For the convolutional Inceptionv3 model, we were able to show that RIBA requires fewer queries to generate perturbations for particularly difficult to manipulate images. On an adversarially trained model, we were able to match the attack success rate of the white-box PGD attack. Finally, we showed that with RIBA, queries can be grouped together into batches, even when generating a perturbation for a single image, further reducing the amount of times the target model is called. This is a unique property of RIBA that stems from the inspiration from offpolicy RL.
7. Ethics statement We are aware of the significant risk that adversarial attacks pose to the safety of computer vision systems applied in the real world. Black-box attacks, such as the one presented in this paper, are especially dangerous, as they are able to attack systems which can only be accessed through an API. We release this novel approach to outline the possibility of the attack and spark the creation of tailored defense mechanisms. It is important that novel attacks are discovered through red teaming, giving machine learning practitioners the chance to mitigate them, before the attacks are employed by malicious actors.
References [1] Maksym Andriushchenko, Francesco Croce, Nicolas Flammarion, and Matthias Hein. Square Attack: A QueryEfficient Black-Box Adversarial Attack via Random Search. In Computer Vision – ECCV 2020, pages 484–501, Cham, 2020. Springer International Publishing. 2 [2] Yang Bai, Yisen Wang, Yuyuan Zeng, Yong Jiang, and ShuTao Xia. Query efficient black-box adversarial attack on deep neural networks. Pattern Recognition, 133:109037, 2023. 2 [3] Brian R. Bartoldson, James Diffenderfer, Konstantinos Parasyris, and Bhavya Kailkhura. Adversarial Robustness Limits via Scaling-Law and Human-Alignment Studies. In Proc. Mach. Learn. Res., pages 3046–3072. ML Research Press, 2024. 6 [4] Nicholas Carlini and David Wagner. Towards Evaluating the Robustness of Neural Networks. In 2017 IEEE Sym-
posium on Security and Privacy (SP), pages 39–57. IEEE, 2017. ISSN: 2375-1207. 2 [5] Pin-Yu Chen, Huan Zhang, Yash Sharma, Jinfeng Yi, and Cho-Jui Hsieh. ZOO: Zeroth Order Optimization Based Black-box Attacks to Deep Neural Networks without Training Substitute Models. In Proceedings of the 10th ACM Workshop on Artificial Intelligence and Security, pages 15– 26, New York, NY, USA, 2017. Association for Computing Machinery. 2, 3 [6] Xiangyi Chen, Sijia Liu, Kaidi Xu, Xingguo Li, Xue Lin, Mingyi Hong, and David Cox. ZO-AdaMM: Zeroth-Order Adaptive Momentum Method for Black-Box Optimization. In Advances in Neural Information Processing Systems. Curran Associates, Inc., 2019. 2 [7] Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. RobustBench: a standardized adversarial robustness benchmark. In Adv. neural inf. proces. syst. Neural information processing systems foundation, 2021. 6 [8] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. An image is worth 16x16 words: Transformers for image recognition at scale. In ICLR - Int. Conf. Learn. Represent. International Conference on Learning Representations, ICLR, 2021. 4 [9] Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and Harnessing Adversarial Examples, 2015. arXiv:1412.6572 [stat]. 2, 6 [10] Chuan Guo, Jacob Gardner, Yurong You, Andrew Gordon Wilson, and Kilian Weinberger. Simple Black-box Adversarial Attacks. In Proceedings of the 36th International Conference on Machine Learning, pages 2484–2493. PMLR, 2019. 2 [11] Tuomas Haarnoja, Aurick Zhou, Pieter Abbeel, and Sergey Levine. Soft Actor-Critic: Off-Policy Maximum Entropy Deep Reinforcement Learning with a Stochastic Actor. In Proceedings of the 35th International Conference on Machine Learning, pages 1861–1870. PMLR, 2018. 2, 3, 7 [12] Tuomas Haarnoja, Aurick Zhou, Kristian Hartikainen, George Tucker, Sehoon Ha, Jie Tan, Vikash Kumar, Henry Zhu, Abhishek Gupta, Pieter Abbeel, and Sergey Levine. Soft Actor-Critic Algorithms and Applications, 2019. arXiv:1812.05905 [cs.LG]. 4, 5 [13] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In Proc IEEE Comput Soc Conf Comput Vision Pattern Recognit, pages 770–778. IEEE Computer Society, 2016. 4 [14] Yiran Huang, Yexu Zhou, Michael Hefenbrock, Till Riedel, Likun Fang, and Michael Beigl. Universal Distributional Decision-Based Black-Box Adversarial Attack with Reinforcement Learning. In Lect. Notes Comput. Sci., pages 206–215. Springer Science and Business Media Deutschland GmbH, 2023. 2 [15] Andrew Ilyas, Logan Engstrom, Anish Athalye, and Jessy Lin. Black-box Adversarial Attacks with Limited Queries
and Information. In Proceedings of the 35th International Conference on Machine Learning, pages 2137–2146. PMLR, 2018. 2 [16] Andrew Ilyas, Logan Engstrom, and Aleksander Madry. Prior convictions: Black-box adversarial attacks with bandits and priors. In International Conference on Learning Representations, 2019. 2 [17] Xu Kang, Bin Song, Jie Guo, Hao Qin, Xiaojiang Du, and Mohsen Guizani. Black-box attacks on image classification model with advantage actor-critic algorithm in latent space. Information Sciences, 624:624–638, 2023. 2 [18] Alex Krizhevsky et al. Learning multiple layers of features from tiny images. 2009. 4 [19] Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In Int. Conf. Learn. Represent., ICLR. International Conference on Learning Representations, ICLR, 2019. 4 [20] Chen Ma, Li Chen, and Jun-Hai Yong. Simulating unknown target models for query-efficient black-box attacks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 11835–11844, 2021. 2 [21] Zerou Ma and Tao Feng. Query-Efficient Two-Phase Reinforcement Learning Framework for Black-Box Adversarial Attacks. Symmetry, 17(7):1093, 2025. 2 [22] Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In Int. Conf. Learn. Represent., ICLR - Conf. Track Proc. International Conference on Learning Representations, ICLR, 2018. 2, 6 [23] Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael Bernstein, Alexander C. Berg, and Li Fei-Fei. ImageNet Large Scale Visual Recognition Challenge. International Journal of Computer Vision, 115(3): 211–252, 2015. 4 [24] Soumyendu Sarkar, Ashwin Ramesh Babu, Sajad Mousavi, Sahand Ghorbanpour, Vineet Gundecha, Antonio Guillen, Ricardo Luna, and Avisek Naug. Robustness with queryefficient adversarial attack using reinforcement learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, pages 2330–2337, 2023. 2 [25] Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian Goodfellow, and Rob Fergus. Intriguing properties of neural networks, 2014. arXiv:1312.6199 [cs]. 1 [26] Christian Szegedy, Vincent Vanhoucke, Sergey Ioffe, Jon Shlens, and Zbigniew Wojna. Rethinking the Inception Architecture for Computer Vision. In Proc IEEE Comput Soc Conf Comput Vision Pattern Recognit, pages 2818–2826. IEEE Computer Society, 2016. 4 [27] Yifeng Xiong, Jiadong Lin, Min Zhang, John E. Hopcroft, and Kun He. Stochastic Variance Reduced Ensemble Adversarial Attack for Boosting the Adversarial Transferability. In Proc IEEE Comput Soc Conf Comput Vision Pattern Recognit, pages 14963–14972. IEEE Computer Society, 2022. 2
[28] Jiancheng Yang, Yangzhou Jiang, Xiaoyang Huang, Bingbing Ni, and Chenglong Zhao. Learning Black-Box Attackers with Transferable Priors and Query Feedback. In Advances in Neural Information Processing Systems, pages 12288–12299. Curran Associates, Inc., 2020. 2 [29] Pu Zhao, Pin-yu Chen, Siyue Wang, and Xue Lin. Towards Query-Efficient Black-Box Adversary with Zeroth-Order Natural Gradient Descent. Proceedings of the AAAI Conference on Artificial Intelligence, 34(04):6909–6916, 2020. 2 [30] Linjun Zhou, Peng Cui, Xingxuan Zhang, Yinan Jiang, and Shiqiang Yang. Adversarial eigen attack on black-box models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pages 15254– 15262, 2022. 2