ConceptioArchivearXiv CS
arXiv CSopen access

The Role of Input Dimensionality in the Emergence and Targeted Control of Adversarial Examples

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

1

The Role of Input Dimensionality in the Emergence and Targeted Control of Adversarial Examples

arXiv:2606.26207v1 [stat.ML] 24 Jun 2026

Nasrin Malekzadeh Goradel, Niccolò Pancino, Member, IEEE, Yaser Gholizade Atani, Benedetta Tondi, Member, IEEE, Giovanni Bellettini, Mauro Barni, Fellow, IEEE

increases, most data points become arbitrarily close either to class decision boundaries [4] or to misclassification regions [3], [5]. While these theoretical analyses are compelling due to their elegance and generality, they rely on assumptions about the concentration properties of data distributions whose validity in practical settings remains largely unverified. Some works have tried to clarify this issue by characterizing the conditions under which adversarial examples may be avoidable. In particular, [10] shows that when class-conditional distributions are strongly localized relatively to inter-class separation, adversarial robustness is in principle achievable. Whether the conditions stated in [10] are satisfied by realworld data, and how they interact with input dimensionality in practical learning scenarios, however, remains unclear. Motivated by the above arguments, in this work we take a systematic empirical perspective on the role of input dimensionality in the emergence of adversarial examples. To start with, we introduce a pool of hierarchical datasets composed of images at multiple resolutions, obtained by downsampling high-resolution images to progressively smaller sizes. This allows us to vary input dimensionality in a controlled manner while preserving semantic content. Using these datasets, we first show that as input dimensionality increases, classconditional distributions become increasingly localized, revealing a geometric regime that departs significantly from the assumptions underlying most concentration-based theoIndex Terms—Adversarial examples, concentration of measure, retical analyses. Then, we perform an extensive experimental high-dimensional geometry, targeted adversarial examples evaluation across diverse neural architectures, including both standard and robustness-enhanced models, and consistently I. I NTRODUCTION observe that adversarial examples become easier to construct INCE their existence was first pointed out [1], adversarial as input dimensionality increases. In the second part of the paper, we move beyond untargeted examples - small, carefully crafted perturbations that adversarial examples and investigate targeted attacks, in which cause machine learning models to misclassify inputs - have the adversary aims to induce misclassification into a specific become the subject of intense research. A commonly proposed target class. Targeted attacks are commonly regarded as more explanation for the existence of adversarial examples points difficult to construct due to their stronger control requirements, to the concentration of measure phenomenon [2], with several in addition, most theoretical analyses focus on the untargeted works suggesting that adversarial examples are an inherent setting. As a first contribution, we extend existing geometric consequence of high-dimensional geometry [3]–[9]. The core arguments for untargeted adversarial examples to the targeted intuition behind these works is that, due to the concentrasetting only. In particular, relying on a classical extension of tion of measure, as the dimensionality of the input space the concentration function to sets of arbitrary size [2], we show Acknowledgments. The research by N. M. Goradel and Prof. M. Barni was that, under mild assumptions on the number of classes and funded by a grant provided by Leonardo SpA. The work was also partially the volume occupied by the decision regions defined by the supported by the EU - NextGenerationEU, under the National Recovery attacked classifier, the additional difficulty of crafting targeted and Resilience Plan (NRRP) - Extended Partnership SERICS, Spoke 3 Attacks and Defences, Mission 4, Component 2, Investment 1.3, AI-RESCUE. adversarial examples remains inherently limited and becomes PE00000014, CUP: B63C24000490006. asymptotically negligible as dimensionality grows. Then, to 0000–0000/00$00.00 assess © 2021the IEEE validity of the theoretical predictions - which, like Abstract—Several theoretical works have tried to explain the adversarial vulnerability of deep neural networks through properties of high-dimensional geometry. However, the assumptions underlying these works are rarely examined empirically, and systematic evidence remains limited. In this work, we present a systematic study of the role of input dimensionality in both the emergence and the targeted control of adversarial examples. We first analyse the scope and limitations of existing theoretical frameworks based on concentration of measure, showing that real image classes exhibit strong empirical localization, beyond what such theories typically assume. We then conduct an extensive empirical evaluation across hierarchical image datasets spanning a wide range of input dimensionalities and diverse neural architectures. Our results consistently show that adversarial examples become easier to construct as dimensionality increases. We also investigate how input dimensionality affects the additional difficulty of crafting targeted adversarial examples. In particular, we provide theoretical arguments showing that highdimensional geometry implies that enforcing a specific target label entails only a limited additional distortion compared to untargeted attacks. We corroborate this insight through extensive experiments, demonstrating that the gap between targeted and untargeted perturbations remains small and further narrows as input dimensionality increases. While, taken together, our findings establish high input dimensionality as a fundamental factor underlying the emergence and targeted control of adversarial examples, whether this phenomenon primarily arises from the interplay between high-dimensional geometry and data distributions or from the architectural properties of deep neural networks remains an open question.

S

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

for untargeted attacks, rely on assumptions whose plausibility in real-world settings is questionable - we conduct a systematic empirical study across the hierarchical datasets introduced earlier. Our results consistently reveal that: i) targeted attacks also become easier to craft as dimensionality increases, and ii) the additional cost required to control the target class of the attack remains limited and diminishes with input size. With the above ideas in mind, the main contributions of this paper can be summarised as follows: • We introduce a pool of hierarchical image datasets that enable a controlled study of adversarial vulnerability as a function of input dimensionality. • We show that as input dimensionality increases, the support of class-distributions shrinks at an exponential rate, thus challenging the key assumptions of most theoretical analysis relying on the concentration of measure phenomenon. • We empirically demonstrate that, despite strong mass localization of class distributions, adversarial examples become progressively easier to craft as input dimensionality increases, with the minimum distortion required to achieve a fixed attack success rate decreasing systematically when resolution increases. • We extend existing geometric concentration arguments to the case of targeted adversarial attacks. • We empirically show that the greater control afforded by targeted attacks comes at only a modest additional cost, and that this cost decreases with input dimensionality. Overall, our analysis provides empirical evidence for a dimensionality-driven effect akin to a curse of dimensionality in the emergence and targeted control of adversarial examples. At the same time, the empirical violation of key assumptions underlying existing theoretical explanations suggests that highdimensional geometry alone may be insufficient to fully explain adversarial vulnerability, leaving open the question of how geometry, real data distributions, and classifier architecture jointly contribute to this phenomenon. II. P RIOR A RT The first attempt to use high-dimensional geometry to explain the emergence of adversarial examples can be traced back to the work of Gilmer et al. [6]. In that paper, geometric properties of high-dimensional spaces are exploited to show that, when data are distributed on concentric highdimensional spheres, adversarial examples exist even for an optimal classifier. Similar arguments were later employed by Fawzi et al. [11] in a much more general setting, showing that for artificially generated images obtained from a Gaussian distribution in a latent space and a smooth mapping between latent and image spaces, adversarial examples necessarily emerge for any classifier. While concentration phenomena are already implicit in [11], the central role of concentration of measure in adversarial vulnerability is made explicit in [5], [12]. A further merit of [12] is the systematic analysis of several equivalent definitions of adversarial examples, distinguishing between the case in which the attacker aims at pushing a sample into the error region of the classifier (ER), inducing a change of prediction (CP), or

2

crafting a sample that is misclassified with respect to the true label of the original input (Corrupted Instance, CI). Adopting the ER definition Mahloujifar et al. [5] prove that adversarial examples are inevitable for any non-ideal classifier and for any data distribution exhibiting sufficiently strong concentration properties, including Lévy families and the uniform distribution over the unit hypercube. The results in [5] are further extended in [8], where they are used to analyze the trade-off between standard and adversarial accuracy. The analyses in [5] and [8] rely on the ER definition and therefore apply only to non-ideal classifiers, i.e., classifiers with non-zero error probability in the absence of attacks. This limitation is removed by Shafahi et al. [4], who show that, for high-dimensional data, adversarial examples are inevitable even for perfect classifiers, provided that inputs are distributed over the unit hypercube and that the class-conditional densities are uniformly bounded. A natural question stemming from the analyses reviewed above is whether natural images actually satisfy the concentration properties assumed by theoretical works. In [4], the boundedness assumption on the input distribution is investigated in a simplified setting in which, starting from small images such as those in MNIST, the input dimensionality is artificially increased by duplicating each pixel into a b × b block of identical values. The findings in [4], however, are not conclusive, primarily due to the simplicity of the considered setting. More broadly, several recent works have attempted to empirically assess the concentration properties of realworld datasets, including [13]–[15], where the concentration function of MNIST and CIFAR-10 datasets is estimated and used to validate the predictions of theory. While these studies establish an important bridge between theoretical insights and empirical data, they do not directly address the role of input dimensionality in the emergence of adversarial examples for two main reasons. First, the analysis is restricted to datasets composed of low-resolution images only. Second, and more importantly, no effort is made to investigate how adversarial vulnerability evolves as the image resolution increases within the same visual domain, and hence how concentration itself scales with dimensionality under fixed semantic content. The question of whether the distributional assumptions underlying theoretical results are satisfied by natural images becomes even more relevant in light of recent work showing that adversarial examples may, in principle, be avoidable when the data distribution satisfies sufficiently strong geometric constraints. In particular, Pal et al. [10] show that adversarial examples are not inevitable if the class-distribution of the underlying data is sufficiently localized. They prove that a necessary condition for adversarial examples to be avoidable is that the volume of the support of class-distributions decay exponentially fast with the ambient dimension, at a rate depending on the admissible perturbation. They also derive a sufficient condition to ensure avoidability Notably, these requirements are considerably stronger than those typically expected to hold for natural images, thereby sharpening the following central question: as image resolution increases within a fixed domain, do real-world data distributions evolve toward a regime of stronger concentration and inevitability

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

3

of adversarial examples, or toward one in which adversarial examples might, in principle, be avoidable? The first goal of this paper is to answer this question empirically by relying on a comprehensive experimental campaign explicitly controlling the effect of input dimensionality through resolution-scaled datasets sharing the same semantic content. Most theoretical and practical works on adversarial examples focus on the untargeted setting. Extending concentrationbased existence results to targeted attacks is not immediate, since the target region of the attack may occupy only a small fraction of the input space, preventing the direct application of basic concentration arguments. As a consequence, the additional distortion required to enforce a prescribed target class and, crucially, how such a cost scales with input dimensionality - remains largely unexplored. Even the definitions introduced in [12], as well as the metrics commonly used to quantify adversarial robustness, require some care in the targeted setting, for instance by distinguishing between average distortions over all source-target class pairs and worst-case distortions associated with the most difficult class transitions. From a practical perspective, several works have shown that targeted adversarial attacks can also be crafted effectively, albeit typically requiring larger perturbations than their untargeted counterparts [16]–[20]. However, a systematic investigation of how input dimensionality affects the ease of constructing targeted adversarial examples is still missing. This observation naturally leads to the second research question addressed in this paper: As input dimensionality increases within a fixed visual domain, how does the additional distortion required to enforce a specific target class evolve? In particular, does the gap between targeted and untargeted attacks widen, remain stable, or shrink as image resolution grows? The second part of our work is devoted to answering this question, both from a theoretical and an empirical standpoint.

Adversarial examples following this definition are sometimes referred to as Prediction Change (PC) adversarial examples [12]. A limitation of the PC definition is that it does not account for the ground truth function c. For this reason, an alternative definition requires that f (x′ ) ̸= c(x′ ). Adversarial examples defined in this way are referred to as Error Region (ER) adversarial examples [12]. It is immediate to see that this second definition makes sense only for non-ideal classifiers. For an ideal classifier, in fact, f (x) = c(x) for every x and the conditions f (x′ ) ̸= c(x′ ) cannot be satisfied. In [12], a third definition of adversarial examples is introduced, namely Corrupted Instance (CI) adversarial examples, for which it is required that f (x′ ) ̸= c(x). In this case, adversarial examples may also exist for an ideal classifier. In fact, in such a case, the CI definition is equivalent to the PC one. In practical scenarios, where classifiers are not perfect but achieve high accuracy, we often have f (x) = c(x). In addition, the so called proximity assumption approximately holds, according to which, for small perturbations we (almost) always have c(x) = c(x′ ). As a result, the three definitions outlined above tend to be nearly equivalent. With regard to the metric used to constraint the maximum allowed perturbation, most works adopt either the ℓ2 or the ℓ∞ norm. In this work, we focus on the ℓ2 case. With the above definitions, we can define the adversarial risk as the probability over the distribution of input samples that an adversarial example exists at a given distance from the unperturbed input, formally:

III. P RELIMINARY N OTIONS , D EFINITIONS AND R ESEARCH Q UESTIONS

A. Concentration of measure phenomenon The concentration of measure phenomenon refers to a property of some probability measures in metric spaces (especially high-dimensional ones) according to which most of the probability is arbitrarily close to every large enough set. More specifically, let (X , d, µ) be a metric measure space, where X is a space of dimension n, d is a distance induced by a metric defined on X , and µ is a Borel probability measure on X . The formalisation of the concentration of measure phenomenon passes through the definition of the concentration function α : [0, ∞) → [0, 1] [2]:

In this section we introduce the mathematical notation used throughout the paper and provide a rigorous formulation of adversarial examples and adversarial risk. We then review the key concepts related to the concentration of measure phenomenon, and formulate the goals of this work. Let X be the space of the to-be-classified instances x ∈ X , and Y = {1, . . . , m} be the set of labels. In the following, with a slight abuse of notation, we indicate with µ the probability distribution of the instances in X and the probability measure induced by such a distribution on the set X , the exact meaning being always clear from the context. The association of instances and labels is defined by a ground truth function c : X → Y 1 . Given a classifier f : X → Y, a common way to define an adversarial example is as a perturbed sample x′ = x+δ ∈ X such that f (x′ ) ̸= f (x) [1], where δ denotes an imperceptible perturbation constrained by a predefined norm. 1 We assume that each x ∈ X is linked to a unique, deterministic, ground truth label. This label might reflect, for example, a human’s classification of an image. This perspective differs from probabilistic approaches, where the relationship between samples and labels is modeled by a probability distribution over X × Y.

Riskε = Pr{x′ ∈ Bε (x)∩X s.t. f (x′ ) ̸= {f (x), c(x′ ), c(x)}}, (1) where the probability is taken over the input sample distribution µ, Bε (x) indicates the ball of radius ε centered in x, and where the exact condition imposed on f (x′ ) depends on the definition of adversarial examples (PC, ER, or CI).

α(ε) := sup {1 − µ(Aε ) : A ⊂ X , µ(A) ≥ 1/2} ,

(2)

where Aε := {x ∈ X : d(x, A) ≤ ε} indicates the εneighbourhood (or ε-expansion) of A. It turns out that in many high-dimensional spaces, like high-dimensional spheres with uniform measure, and product spaces, α(ε) decreases exponentially fast with ε. B. Adversarial examples and concentration of measure Several researchers have used the concentration of measure phenomenon to explain the pervasive existence of adversarial

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

4

examples, demonstrating that, as the dimensionality of input samples increases2 adversarial examples can always be constructed even when ε is arbitrarily small. By adopting the PC definition of adversarial examples, Shafahi et al. [4] prove that given a classification problem with m ≥ 2 classes, each distributed over the unit hypercube [0, 1]n with density functions {µi }m i=1 , for any f that partitions the hypercube into disjoint measurable subsets, given a sample x belonging to class i, one of the following holds: • x is misclassified by f ′ ′ • x has an adversarial example x , with ∥x − x ∥2 ≤ ε with probability3 2

e−πε (3) Pe,adv ≥ 1 − Ui 2πε where Ui is the supremum of µi . The only condition behind this result is that the fraction Fi of the hypercube assigned to class i by f has a volume smaller than 1/2, a condition that is reasonably met by any classifier with more than 2 classes. When the input sample x corresponds to an image, it is convenient to express the magnitude ε of the perturbation in terms of mean square error γ, that is: ε2 = γ, n allowing us to reformulate (3) as: Pe,adv ≥ 1 − Ui

e−πnγ √ . 2π nγ

(4)

(5)

If Ui is bounded, then, as n increases, the probability that x admits an adversarial example (or it is misclassified outright) approaches 1. The intuition behind the arguments used in [4] is that due to the concentration of measure phenomenon, for any set occupying less than half of the input space, the distance to the set boundary of most of the points is such that they can be pushed outside the set by introducing an arbitrarily small mean square error. Another line of research [5], [8], [12] establishes the inevitability of adversarial examples under the ER definition by introducing a non-empty error region E where f (x) ̸= c(x). In this setting, if (X , d, µ) is concentrated, a high adversarial risk can be achieved with vanishing mean square distortion as the input dimensionality increases. The intuition is that most correctly classified samples lie close to the boundary of E and can therefore be pushed into the error region by increasingly small perturbations.

mass of each class is smoothly spread over these spaces, so that the corresponding concentration function decays rapidly with the input dimensionality n. By contrast, if most of the probability mass associated with one or more classes is confined to small regions, and if such regions are sufficiently well separated, one may reasonably expect the existence of robust classifiers. This intuition is formalized in [10] through the introduction of (a, b)-localized and (a, b, c)-strongly localized distributions4 . Specifically, a distribution µ over X is said to be (a, b)-localized if there exists a subset S ⊂ X such that µ(S) ≥ 1 − b and Vol(S) ≤ Ce−an for some constant C, where Vol(·) denotes a volume measure over X . Theorem 2.1 in [10] states that a necessary condition for the existence of a classifier satisfying Riska ≤ 1 − b is that µi is (a, b)-localized for at least one class i. Localization alone does not guarantee robustness; however, it invalidates the assumptions under which concentration-based inevitability results are derived. This limitation is particularly evident in Shafahi et al. [4]. From the definition of an (a, b)-localized distribution, it follows that the supremum of µi - denoted as Ui in (5) - grows exponentially with n, causing the bound in (5) to become vacuous for appropriate choices of γ, a, and b. The first goal of this paper stems from the observation that, when localization conditions are met, the assumptions underpinning concentration-based impossibility results are violated and theory alone becomes inconclusive with respect to the inevitability of adversarial examples. This motivates a direct empirical investigation aimed at assessing whether the distribution of natural images exhibits mass localization. Localization of data distribution constitutes only a necessary condition for the existence of a robust classifier, as it does not account for the separation between the distributions of different classes. This limitation motivates the notion of strongly localized distributions, which additionally require class separation and provide a sufficient condition for robustness [10]). In practice, however, assessing whether the class-conditional distributions of natural images satisfy strong localization properties is generally an intractable problem for real-world image distributions. This observation leads to the second goal goal of our work: to evaluate the actual impact of input dimensionality on the emergence of adversarial examples in practical settings. IV. E XPERIMENTAL S ETTING In this section, we describe the experimental setting used throughout our work.

C. Limits of theory and goals of the paper A natural question is whether the distributional assumptions underlying concentration-based inevitability results are satisfied in practice. All these results rely, in one form or another, on the assumption that the data to be classified live in sufficiently regular metric spaces, and that the probability 2 This is the case, for instance, with colour digital images, where the input sample dimensionality corresponds to three times the number of pixels, often reaching tens or even hundreds of thousands of dimensions. 3 The inequality below is derived directly from the proof of Theorem 2 in [4] (appendix A), and corrects a typo in the statement of the theorem.

A. Hierarchical datasets As we said, most empirical studies on adversarial robustness and concentration-based theories rely on very low-resolution benchmarks such as MNIST (28×28) and CIFAR-10 (32×32), while investigations across input dimensions often compare datasets with different semantic content, making it difficult 4 In [10] the terms concentrated and strongly concentrated distributions are used. To avoid any ambiguity with metric concentration phenomena in high-dimensional spaces, we instead adopt the terms localized and strongly localized distributions.

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

to isolate the effect of dimensionality, or consider artificial scaling as done with MNIST in [4]. To overcome these limitations, we built three hierarchical datasets containing images with identical semantic content at different resolutions. Rather than starting from small images and artificially constructing higher-resolution inputs, as done in [4] for MNIST via pixel duplication, we started from high-resolution images and generated lower-resolution versions by downsampling. We selected datasets characterized by a large native resolution enabling consistent downsampling from the highest resolution to the lowest. Specifically, starting from the images at original resolution, we built downscaled versions of each dataset at 64 × 64, 128 × 128, 256 × 256 and 400 × 400 resolutions5 , by using bilinear interpolation to resize the images to the target size, and scaling pixels by a constant factor of 255 so to map their value in the [0, 1] range. a) VegSeed: The VegSeed dataset was derived from the original VegSeedsBD dataset [21], which comprises 4,500 high-resolution JPEG images (3,456 × 4,608 pixels) evenly distributed across 15 vegetable seed categories - each containing 300 images - and used for single-label, multiclass image classification. In its original form, the data was partitioned into two distinct subsets based on the capture method: individual seeds and bulk samples. In our version, we merged these subsets, removing the structural distinction between single and bulk images while maintaining the original 15 classes, distribution, and task. b) Food-30: The Food-30 dataset was derived from Food-101 [22], a widely used benchmark in computer vision designed for single-label multiclass food image classification. Food-101 contains 101,000 real-world images divided evenly across 101 food categories, with 1000 images per class at varying resolutions, and a predefined split of 750 training images and 250 test images per class. From the original dataset, we selected 30 classes and sampled 500 images per class, for a total of 15,000 images with resolution 512 × 512. c) ResynthDB: We built this dataset starting from the Resynthesis Dataset6 [23] , which contains images generated by 10 generators (Bing, Firefly, Flux-dev, Freepik, Imagen, Leonardoai, Midjourney, Nightcafé, Stabilityai, and Starryai) each acting as a class in a single-label, multiclass prediction task. Each generative model was provided with 100 textual prompts to produce the image set. While most classes consist of 1,100 samples, the Bing and Firefly categories contain 4,247 and 4,364 images, respectively. The resulting dataset contains 17,411 images with resolution 1024 × 1024. Table I summarizes the main characteristics of the datasets we have built. An example of the images contained in the hierarchical datasets we have built is provided in Fig. 1. All datasets were split into training, validation, and test sets, with a ratio of 80%, 10%, and 10% of the samples, respectively. The same split was applied consistently across all resolutions. To account for class imbalance, we used classweighted loss during training. 5 In the rest of the paper we refer to these dimensions as resolutions 64, 128, 256, and 400. 6 The dataset is publicly available at https://www.kaggle.com/datasets/ pietrob92/resynthesis-dataset?select=resynthesis dataset.

5

Fig. 1. Examples of images from VegSeed, Food-30, and ResynthDB datasets (columns, from the left) at the considered input resolutions 64, 128, 256, and 400 (rows, from the top). TABLE I H IERARCHICAL DATASETS DESCRIPTION .

Classes No. samples Samples per class Balanced Original resolution

VegSeed 15 4,500 300 yes 3,456×4,608

Food-30 30 15,000 500 yes 512×512

ResynthDB 10 17,411 variable no 1,024×1,024

B. Architectures Since the theoretical results discussed in the previous sections are stated for any classifier, it is crucial that the empirical analysis is carried out across a broad and heterogeneous set of models. For this reason, our experimental campaign includes multiple state-of-the-art classification backbones as well as several robust strategies representative of modern defenses against adversarial examples. This allows us to assess whether the observed trends with input dimensionality persist not only for a few standard models, but also for classifiers explicitly optimized for adversarial robustness. With the above considerations in mind, we evaluated four backbone architectures: MobileNetV3-Small [24] (hereafter MobileNet), EfficientNetV2-Small [25] (hereafter EfficientNet), ResNet-50 [26], and a hybrid CLIP+MLP model, in which a MultiLayer Perceptron (MLP) classifier is paired with a CLIP [27] Vision Transformer backbone for feature extraction. All CNN-based models were pre-trained on the ImageNet dataset [28] and then fine-tuned on the hierarchical datasets. In contrast, the CLIP + MLP architecture consists of a frozen pre-trained CLIP image encoder from OpenAI - originally trained on the WebImageText (WIT) dataset [27] - followed by a randomly initialized MLP for the final classification task. In this case, the dataset construction is followed by an additional procedure based on the pre-trained CLIP parameters, which consists of a normalization with specific mean and standard deviation values, and a resizing transformation into a 224×224 pixels image (by means of a bilinear interpolation), to ensure

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

6

the input images match the specifications of the CLIP vision encoder. For each architecture, we first trained a Standard model using only clean, unperturbed images. We then used these models as starting points to apply state-of-the-art defenses. As a first defense, we trained a Robust model via adversarial training. The Robust model is initialized from the corresponding Standard model and subsequently fine-tuned using both clean and adversarially perturbed data generated via a PGDℓ2 attack [29]. In the experiments, the attack is performed with large maximum allowed perturbation ε and small step size α, and the optimization is stopped as soon as a misclassification is induced at an iteration s′ ≤ s, so as to obtain the minimum distortion required to deceive the model.7 In the training procedure, we set ε = 1000, s = 200, α = 0.01, to balance between positive and negative training examples. In the evaluation procedure we used a higher s = 5000 and also optimized α to ensure the attack reached a predefined efficacy threshold while maintaining minimal perturbation overhead. In addition to adversarial training, we evaluated two defense mechanisms based on input pre-processing, namely JPEG compression and AI-based denoising [30]. In this case, models are trained on pre-processed clean images. At test time, inputs are pre-processed before being fed to the classifier, under the assumption that the pre-processing step can remove the subtle perturbations introduced by adversarial attacks. Overall, our experimental campaign spans 4 backbone architectures, 4 training/defense regimes, and 4 input resolutions, resulting in a total of 64 trained models per dataset. Across the three datasets considered in this work, this amounts to 192 trained networks. Note that each model was trained and evaluated using images at a single resolution only, ensuring that no model was exposed to multiple resolutions during training or testing.

classified on clean data, that is, samples belonging to the set Dcorr = {xi : f (xi ) = c(xi )}, with cardinality |Dcorr | = N · SAcc. By indicating with x′i the result of an attack, the Attack Success Rate (ASR) is defined as the proportion of initially correct samples that are misclassified after the attack, that is: X 1 1 [f (x′i ) ̸= c(xi )] . (7) ASR = |Dcorr | xi ∈Dcorr

c) Adversarial Accuracy (AAcc): Adversarial accuracy measures the fraction of samples in D that are correctly classified after the attack, where the attack is applied only to the subset Dcorr of correctly classified samples: 1 X AAcc = 1 [f (x′i ) = c(xi )] . (8) N xi ∈Dcorr

By construction, Eq. (7) quantifies the fraction of correctly classified samples whose predictions are altered by the attack, highlighting how vulnerable is the model, while Eq. (8) measures the overall fraction of samples correctly classified after the attack. By referring to the definitions given in Section III, AAcc is the empirical counterpart of 1 − Riskε when the CI definition of adversarial examples is adopted and the attack perturbation is bounded by ε. It is also easy to see that: AAcc = SAcc · (1 − ASR),

(9)

highlighting how robustness in the presence of attacks depends jointly on the model’s accuracy on clean samples and its robustness against adversarial perturbations. d) Distortion (MSE): To compare a generic clean image xi and its perturbed version x′i , we used the Mean Squared 2 Pn Error: MSE(xi , x′i ) = n1 j=1 xi,j − x′i,j , where xi,j indicates the j-th pixel of xi . MSE provides a quantitative measure of how much the perturbed image differs from the original at the pixel level.

C. Metrics The robustness of the various models and the effectiveness of the attacks have been evaluated by considering three complementary metrics characterizing the performance on clean data and robustness under adversarial perturbations. Let n D = {xi }N i=1 be the evaluation dataset, where xi ∈ [0, 1] is an input sample. Let c, f defined as in Section III. a) Standard Accuracy (SAcc): The Standard accuracy measures the fraction of samples in D that are correctly classified in the absence of attacks: N

SAcc =

1 X 1 [f (xi ) = c(xi )] . N i=1

(6)

where 1 denotes the usual indicator function. b) Attack Success Rate (ASR): To disentangle the attack effectiveness from the classification errors in the absence of attacks, we restrict the evaluation of the attack success rate to the subset of samples that are correctly 7 Note that, by adopting a large ε as we are doing, the projection becomes ineffective and PGD boils down to a standard iterative gradient ascent attack (PGD without projection). For simplicity, in the following, we refer to the attack as PGD with a slight abuse of terminology.

V. L OCALIZATION P ROPERTIES OF H IERARCHICAL DATASETS In this section, we investigate whether the class-conditional distributions associated with the hierarchical datasets introduced in Section IV-A exhibit mass localization in the sense of Pal et al. [10]. In particular, we empirically assess whether, as input dimensionality increases, the probability mass of the various classes concentrates within subsets of exponentially small volume, as required by the (a, b)-localization condition. To this aim, let Dc = {xci ∈ D | i = 1 . . . Nc } be a dataset containing representative examples xci belonging to class c. To test whether the definition of (a, b)-localized distribution is satisfied, we estimate the volume of a bounding box (playing the role of the set S in the definition of (a, b)-localization) containing all the samples in Dc . To avoid overestimating the volume of the bounding box, we align it with the centered principal component directions of Dc . Specifically, let vic denote the centered principal component coordinates of xci . The volume of the resulting bounding box is  n  Y c c Vol(Boxc ) = max vi,j − min vi,j , (10) j=1

i=1...Nc

i=1...Nc

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

by letting ln Vol(Boxc ) λc = − , (11) n we can say that Dc is empirically (a, b)-localized with a = λc , b = 0 and C = 18 . Motivated by the above arguments, we estimated λc for the classes of the three hierarchical datasets VegSeed, Food30, and ResynthDB. A practical challenge arises because estimating the volume of the n-dimensional bounding-box in PCA coordinates becomes ill-posed, with zero degenerate volume, when the sample covariance is rank-deficient. This limitation becomes critical at higher resolutions. For instance, for 96 × 96 images we have n = 27, 648, requiring at least an equivalent number of samples per class in order to obtain a full-rank covariance matrix, a requirement that far exceeds the size of the classes of the original datasets. To circumvent this limitation, in this section we consider an additional reduced set of input resolutions, namely 32 × 32, 48 × 48, 64 × 64, and 96×96, and we apply an extensive data augmentation strategy in order to ensure that N > n.9 Specifically, we generated at least 30,000 images per class by applying a tailored number of semantic-preserving augmentations to each original sample (60 for Food-30, 100 for VegSeed, and 36 for ResynthDB). The augmentation pipeline 10 combines geometric and visual transformations, including flips, rotations, crop-and-zoom, filtering, and JPEG compression, to increase variability while preserving class identity. It is important to remark that, in this context, augmentation is used primarily as a numerical device to mitigate rank-deficiency and enable stable volume estimation, rather than as an attempt to faithfully sample the underlying (unknown) class distribution. Table II summarizes the results obtained by applying the above procedure to all classes and all considered resolutions. For each dataset and each resolution, we report the minimum, median, maximum, and average values of λc across classes. As can be seen, λc takes non-negligible values for all classes and all resolutions, providing strong empirical evidence in favor of mass localization. Moreover, λc consistently increases with n, suggesting that the volume occupied by image classes decays at least exponentially fast with the ambient dimension, and possibly at a super-exponential rate. For instance, the average λc for VegSeed rises from 0.80 at 32 × 32 to 2.91 at 96 × 96. To further evaluate the localization of classes with respect to the ambient space defined by all images in the dataset, we applied the same analysis to the entire dataset rather than to individual classes. However, computing the log volume over the full dataset is computationally prohibitive due to the size of the aggregated augmented data (e.g., approximately 450k images for VegSeed, 900k for Food-30, and 350k for ResynthDB). To address this issue, we implemented a bootstrapping approach to approximate the global PCA. Specifically, we randomly 8 Considering the case b = 0 is a conservative estimate, as it provides an upper bound on the volume of the smallest set capturing most of the class probability mass. 9 The reduced resolutions used in this section are introduced solely to make the estimation of V ol(Boxc ) feasible in high dimension. They should not be confused with the main experimental resolutions (64, 128, 256, and 400) used in the subsequent sections to evaluate adversarial vulnerability. 10 Available at https://github.com/YaserGholizade/image augmenter

7

sampled subsets of 35,000 images from the full datasets, computed the log volume for each subset, and repeated this process 10 times. The reported value is the average log volume across these trials. We also examined the standard deviation across bootstrap runs (ranging between 0.005 and 0.015) to confirm the stability of the metric and the absence of significant sampling effects. Even when considering the relative volume of class-specific bounding boxes with respect to the volume of the overall dataset bounding box, the localization property of the datasets is confirmed, with stronger localization for larger input sizes. TABLE II E STIMATED λc VALUES FOR REPRESENTATIVE CLASSES (M IN /M EDIAN /M AX VOLUME ), AVERAGE CLASS VALUE , AND THE WHOLE DATASET ACROSS RESOLUTIONS 32, 48, 64, AND 96. Dataset

VegSeed

Food-30

ResynthDB

Class / Type Sponge Gourd Bottle Gourd Pumpkin Average Whole Dataset Beignets Macarons Pizza Average Whole Dataset LeonardoAI Firefly StarryAI Average Whole Dataset

32 0.92 0.83 0.67 0.80 0.43 1.01 0.79 0.66 0.82 0.65 1.43 1.25 1.18 1.28 1.17

48 1.46 1.33 1.20 1.34 0.90 1.40 1.18 0.99 1.17 0.96 1.83 1.65 1.59 1.66 1.53

64 1.95 1.77 1.70 1.82 1.35 1.75 1.55 1.34 1.53 1.26 2.19 2.01 1.96 2.02 1.87

96 2.97 2.80 2.77 2.91 2.31 2.44 2.29 2.09 2.27 1.86 2.87 2.69 2.66 2.70 2.46

While Theorem 2.1 in [10] already states that (a, b)localization is a necessary condition for adversarial examples to be avoidable, it is instructive to interpret the estimates derived from our experiments in the context of the inevitability results of Shafahi et al. [4], summarized in Eq. (5). By assuming that the class distribution admits a density supported within the PCA-aligned bounding box, we obtain Uc ≥ 1/Vol(Boxc ), and Eq. (5) becomes Pe,adv ≥ 1 −

1 √ e−n(πγ−λc ) . 2π nγ

(12)

Typical values of γ (see Section VI) range from 10−4 to 10−6 (or even smaller for harder classification tasks such as source attribution). When coupled with the λc values found in our experiments, these values render the right-hand side of Eq. (12) smaller than zero, thus resulting in a meaningless prediction. VI. E MERGENCE OF A DVERSARIAL E XAMPLES VS I NPUT D IMENSIONALITY: E XPERIMENTAL A NALYSIS Having established that natural image classes exhibit strong mass localization, the assumptions behind concentration-based inevitability results are no longer satisfied and theory alone becomes inconclusive. We therefore turn to an empirical investigation of what happens in practice when the input dimensionality increases. Clearly, evaluating robustness for any possible classifier as predicted by theory is infeasible, for this reason, for each task (dataset) and each resolution, we attacked the full set

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

8

TABLE III MSE90 VALUES FOR ALL THE ATTACKED MODELS AND DEFENSES ON THE V EG S EED DATASET AT VARIOUS RESOLUTIONS . T HE MOST ROBUST MODEL AT EACH RESOLUTION IS HIGHLIGHTED IN BOLD . Model MobileNet ResNet50 EfficientNet CLIP+MLP MobileNet ResNet50 EfficientNet CLIP+MLP MobileNet ResNet50 EfficientNet CLIP+MLP MobileNet ResNet50 EfficientNet CLIP+MLP

Defense JPEG JPEG JPEG JPEG Denoise Denoise Denoise Denoise Adv Train Adv Train Adv Train Adv Train

64 1.61e-5 4.80e-5 2.83e-5 1.74e-6 1.97e-5 6.51e-5 2.94e-5 2.11e-6 3.04e-5 7.11e-5 5.94e-5 1.97e-5 9.14e-5 2.42e-4 6.67e-4 4.54e-6

128 1.37e-5 2.07e-5 1.29e-5 9.02e-7 6.79e-5 5.14e-5 1.45e-5 1.05e-6 2.37e-5 3.06e-5 2.19e-5 1.00e-5 3.02e-5 1.00e-4 6.78e-5 2.48e-6

256 2.29e-6 1.89e-5 6.06e-6 3.66e-7 1.11e-5 2.92e-5 9.77e-6 5.28e-7 6.71e-6 2.37e-5 1.21e-5 3.64e-6 2.05e-5 7.36e-5 4.69e-5 1.37e-6

400 2.30e-6 4.32e-5 2.85e-6 1.73e-7 5.04e-6 2.18e-5 6.63e-6 3.23e-7 5.76e-6 1.76e-5 4.82e-6 1.71e-6 1.84e-5 4.32e-5 2.19e-5 7.99e-7

of models described in Section IV-B, resulting in 16 trained models per resolution. For each model, we generated adversarial examples using a PGD attack. The attack hyperparameters, most notably the perturbation budget ε and the step size α, were tuned so as to obtain the smallest distortion required to successfully attack each input. In this way, for every model we obtained a per-sample estimate of the minimum MSE distortion required for a successful attack. To summarize robustness in a way that is comparable across models and resolutions, we first fixed an ASR of 90% and computed the minimum distortion required to achieve it. Specifically, for each model we collected the per-sample minimum MSE values and defined MSE90 as the smallest MSE threshold such that at least 90% of the samples could be successfully attacked with a distortion not exceeding this value. Tables III, IV, and V report the value of MSE90 for all datasets, models and input size. Based on the results in the tables, ResNet50 with adversarial training turns out to be the most robust model at all resolutions for the VegSeed and Food-30 datasets, with the exception of VegSeed at resolution 64 for which the most robust classifier is EfficientNet with adversarial training. The situation is slightly different for the ResynthDB task where the most robust model is MobileNet with adversarial training at resolutions 64 and 128, EfficientNet with adversarial training at 256 resolution, and CLIP+MLP with JPEG pre-processing at resolution 400. Noticeably, in all cases but ResynthDB at resolution 400, adversarial training yields the most robust models among the considered defenses. In contrast, the identity of the most robust architecture varies across tasks, suggesting that the choice of the most robust model remains strongly task-dependent. Now that we have identified the most robust model for each task and each input size, we can characterize how robustness scales with input dimensionality by comparing the MSE-ASR trade-off across resolutions. Specifically, Fig. 2 reports the ASR for each resolution, as a function of the MSE threshold (equivalently, the MSE required to reach a given target ASR), thus allowing us to assess whether adversarial examples become easier to craft as n increases for any desired

TABLE IV MSE90 VALUES FOR ALL THE ATTACKED MODELS AND DEFENSES ON THE F OOD -30 DATASET AT VARIOUS RESOLUTIONS . T HE MOST ROBUST MODEL AT EACH RESOLUTION IS HIGHLIGHTED IN BOLD . Model MobileNet ResNet50 EfficientNet CLIP+MLP MobileNet ResNet50 EfficientNet CLIP+MLP MobileNet ResNet50 EfficientNet CLIP+MLP MobileNet ResNet50 EfficientNet CLIP+MLP

Defense JPEG JPEG JPEG JPEG Denoise Denoise Denoise Denoise Adv Train Adv Train Adv Train Adv Train

64 1.25e-5 1.35e-5 1.15e-5 3.99e-6 1.68e-5 1.86e-5 1.05e-5 2.00e-6 1.83e-5 1.92e-5 1.53e-5 6.04e-5 4.71e-5 3.16e-4 3.90e-5 6.58e-6

128 3.01e-6 4.86e-6 4.08e-6 2.01e-6 5.63e-6 7.83e-6 3.69e-6 8.83e-7 5.79e-6 8.26e-6 6.92e-6 3.19e-5 1.76e-5 1.49e-4 1.74e-5 4.49e-6

256 9.45e-7 2.08e-6 1.07e-6 1.11e-6 1.44e-6 2.61e-6 1.14e-6 6.23e-7 2.01e-6 3.91e-6 2.06e-6 1.84e-5 9.17e-6 8.65e-5 4.04e-6 2.30e-6

400 4.49e-7 1.72e-6 8.45e-7 6.12e-7 7.41e-7 1.21e-6 6.37e-7 4.05e-7 1.05e-6 2.63e-6 1.27e-6 1.12e-5 3.33e-6 5.68e-5 2.89e-6 1.33e-6

TABLE V MSE90 VALUES FOR ALL THE ATTACKED MODELS AND DEFENSES ON THE R ESYNTH DB DATASET AT VARIOUS RESOLUTIONS . T HE MOST ROBUST MODEL AT EACH RESOLUTION IS HIGHLIGHTED IN BOLD . Model MobileNet ResNet50 EfficientNet CLIP+MLP MobileNet ResNet50 EfficientNet CLIP+MLP MobileNet ResNet50 EfficientNet CLIP+MLP MobileNet ResNet50 EfficientNet CLIP+MLP

Defense JPEG JPEG JPEG JPEG Denoise Denoise Denoise Denoise Adv Train Adv Train Adv Train Adv Train

64 7.50e-6 6.93e-6 6.36e-6 1.29e-6 1.83e-5 1.95e-5 1.70e-5 2.27e-5 2.27e-5 9.39e-6 8.27e-6 2.10e-5 7.11e-5 5.15e-5 3.46e-5 2.55e-5

128 9.73e-7 1.92e-6 3.52e-6 5.15e-7 1.99e-5 4.56e-6 1.15e-5 1.35e-5 1.62e-6 4.09e-6 5.40e-6 7.89e-6 3.75e-5 8.13e-6 1.42e-5 1.15e-5

256 8.10e-8 8.21e-8 2.27e-6 1.27e-7 6.87e-7 1.32e-6 2.31e-6 6.37e-6 9.16e-7 2.89e-7 2.57e-6 2.26e-6 5.01e-6 9.35e-7 6.48e-6 2.04e-6

400 8.12e-8 5.26e-8 7.58e-7 4.94e-8 4.69e-7 3.94e-7 8.19e-7 3.08e-6 3.54e-7 1.33e-7 1.33e-6 8.14e-7 2.08e-7 2.19e-7 1.49e-6 9.76e-7

operating point, and not only at ASR= 90%. As can be seen, a clear and consistent trend emerges across all three datasets: as the input dimensionality increases, adversarial examples can be crafted with progressively smaller perturbations. This effect is not limited to the reference operating point used to select the most robust model (ASR= 90%), but is observed across the full range of attack success rates, as evidenced by the systematic left-shift of the ASR-MSE curves when the image size increases. Quantitatively, the MSE required to achieve ASR= 90% decreases by factors ranging from approximately 6× (Food-30) to over 20× (ResynthDB) when increasing the resolution from 64 × 64 to 400 × 400. Notably, the magnitude of the effect depends on the dataset and is strongest for ResynthDB, suggesting that the dimensionalitydriven increase in vulnerability may be amplified in settings where classification relies on weaker cues. While ASR quantifies the ease with which adversarial examples can be crafted, a direct comparison with theoretical results requires evaluating the adversarial risk, which combines both vulnerability to adversarial perturbations and baseline

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

9

Fig. 2. ASR vs. log(MSE) in the untargeted adversarial attack scenario for the most robust models across datasets. Legend values report MSE90 for each model.

Fig. 3. Risk vs. log(MSE) in the untargeted adversarial attack scenario for the most robust models across datasets.

classification performance. From an empirical perspective, this corresponds to computing the adversarial accuracy AAcc defined in Eq. (9), and then deriving the adversarial risk as Riskε = 1 − AAcc. Under the CI definition of adversarial examples, this relationship is exact. Moreover, it provides a good approximation under the ER and PC definitions when the perturbation magnitude is small and the clean accuracy of the classifier is sufficiently high. By exploiting this relationship and the SAcc values reported in the plots of Fig. 2, we obtain the corresponding adversarial risk curves shown in Fig. 3. By inspecting Fig. 3, we observe that although models operating on lower-dimensional inputs typically exhibit lower clean accuracy, they require substantially larger distortions to reach high adversarial risk. Conversely, at very small perturbation levels higher-dimensional models may retain higher overall accuracy due to their superior clean performance. This behavior suggests that robustness and standard accuracy may evolve differently as the input dimension increases. Taken together, our results provide a clear empirical answer to the central question motivating this section. Although the strong mass localization observed in Section V invalidates the assumptions behind concentration-based inevitability theorems, in practice adversarial vulnerability increases systematically with input dimensionality. In particular, across datasets, architectures, and defenses, higher-resolution inputs consistently require smaller perturbations to achieve a fixed level of ASR. This finding establishes a robust empirical form of curse of dimensionality for adversarial examples, and motivates the second part of the paper, where we investigate whether a similar trend holds under the stronger requirement of targeted adversarial examples.

VII. TARGETED A DVERSARIAL E XAMPLES : F ORMALIZATION AND T HEORETICAL A NALYSIS We now investigate whether enforcing a specific target label requires significantly larger distortion than untargeted attacks. A. Targeted adversarial examples Let X , Y, c, and f be defined as in Section III. We denote by Ci = {x ∈ X : c(x) = i}, i = 1, . . . , m a partition of X into m classes, and by Fi = {x ∈ X : f (x) = i} the classification regions of the classifier f . We also let Vi be the volume of Fi . A targeted adversarial example (in the CI sense) with source class i and target class j, i ̸= j, is a perturbed version of a sample x belonging to the i-th class, that is assigned to class j by f , in formulas: xi→j = x + δ

such that x ∈ Ci ; xi→j ∈ Fj ,

(13)

where δ denotes an imperceptible perturbation. Similar definitions can be given for the ER and PC cases. The extension of the notion of adversarial risk requires to specify over which source and target sets the probability that an adversarial example can be built is computed, yielding11 : Riski→j = Pr {∃x′ ∈ Bε (x) ∩ X : f (x′ ) = j} , i ̸= j ε x∼µi

X 1 Riski→j ε m−1 i:i̸=j XX 1 Risktε = Riski→j ε m(m − 1) j Riskjε =

(14)

i:i̸=j

11 Even in this case we adopt the CI definition of adversarial examples; similar definitions can be given in the ER and PC cases.

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

10

where probabilities are computed with respect to the class distributions µi . In the first case, we consider the probability that targeted adversarial examples can be constructed for a specific source-target class pair. In the second case, we fix a target class and average over all possible source classes, thus measuring how easy it is to force the classifier toward a given target class regardless of the source class. In the third case, we measure the global targeted adversarial risk by averaging uniformly over all possible source-target class pairs. Note that in the above definitions, we adopt uniform averaging over source and target classes in order to isolate the intrinsic geometric difficulty of targeted perturbations from class-frequency effects. For balanced datasets, this definition coincides with the standard risk averaged according to the data distribution.

B. Targeted attacks and concentration of measure As discussed in Section III-B, the concentration of measure can be used to explain the emergence of untargeted adversarial examples for data following a concentrated distribution for which α(ε) decreases fast enough with n. To the best of our knowledge no such result exists for the case of targeted attacks. In this section, we fill this gap by extending the analysis carried out in [4] to the case of targeted adversarial examples. The idea behind the extension is the following. We know from [2] (Proposition 2.8) that for the uniform measure on the hypercube [0, 1]n equipped with the Euclidean distance, the concentration function satisfies: 2

α(ε) ≤ e−πε = e−πnγ .

(15)

In [4], such a property is exploited to prove that for any γ and for any class i such that Vi < 1/2, when n tends to infinity the expansion of the complement of Fi occupies most of the hypercube, and hence the volume of the samples that cannot be moved outside Fi shrinks to zero. If the density function of class i is uniformly bounded with respect to n, this implies that adversarial examples can be crafted with probability arbitrarily close to 1. The targeted case differs since the goal is to reach a specific region Fj , whose measure may be smaller than 1/2. While concentration inequalities are strongest for sets of measure at least 1/2, they can still be applied to smaller sets by adjusting the expansion radius. We show that this adjustment results in an additional distortion term which vanishes as n → ∞, thereby proving that enforcing a specific target class incurs only a negligible asymptotic penalty compared to the untargeted case. A precise statement and proof of our result are given in the following. Theorem 1: (Targeted adversarial examples) Let us consider a classification problem with input samples x ∈ [0, 1]n , ground-truth classes Ci , i = 1, . . . , m, and classification regions Fj , j = 1, . . . , m, associated with a classifier f . Let Vj = Vol(Fj ) be the volume of region Fj , and assume that a positive value V0 independent of n exists such that Vj ≥ V0 > 0 ∀j. Let µi denote the class-conditional distribution over Ci , and let Ui < ∞ be the supremum of

µi . Then, given a sample x ∈ Ci , for any pair i ̸= j, either with probability i→j Pe,adv ≥ 1 − Ui e−πnγ

(16)

sample x is misclassified as class j (x ∈ Fj ) or a targeted adversarial example xi→j exists such that ∥x − xi→j ∥2 ≤ γ i→j , n

(17)

lim γ i→j = γ.

(18)

with n→∞

Proof 1: Let us fix a target class j, and choose ε0 such that α(ε0 ) < V0 ,

(19)

so that in particular α(ε0 ) < Vj . By applying Lemma 1.1 in [2] to Fj with the uniform measure over the hypercube, and exploiting the concentration bound for the hypercube equipped with the Euclidean metric, and given that by hypotheses Vol(Fj ) = Vj ≥ V0 > 0, we obtain 2

1 − Vol((Fj )ε+ε0 ) ≤ α(ε) ≤ e−πε .

(20)

Let Sj := [0, 1]n \(Fj )ε+ε0 denote the set of safe points whose 2 distance from Fj exceeds ε + ε0 . By (20), Vol(Sj ) ≤ e−πε . Since the class-conditional distribution µi is bounded by Ui , it follows that 2

µi (Sj ) ≤ Ui · Vol(Sj ) ≤ Ui e−πε . Recalling that ε2 = nγ, we obtain (16). To go on, observe that in order for (16) to hold, the perturbation radius must be increased from ε to ε + ε0 , where ε0 is chosen so that (19) holds. By the concentration bound for the uniform measure on 2 the hypercube α(ε0 ) < e−πε0 , implying that for (16) to hold it is sufficient that s   1 1 ln . (21) ε0 > π V0 By the above and by recalling once more that ε2 = nγ, we conclude that in order to be able to craft a targeted adversarial example with probability at least 1 − Ui e−πnγ , it is sufficient that: v   u u ln 1 !2   t V0 ε2 ε0 2 (ε + ε0 )2 = 1+ >γ 1+ . n n ε πnγ A valid choice for the bound in (17) is then v   u u ln 1 !2 t V0 i→j γ =γ 1+2 , πnγ that tends to γ when n → ∞, thus completing the proof of the theorem. □ The quantity γ i→j upper bounds the distortion required to enforce a specific target class j starting from samples of class i. The theorem shows that γ i→j is asymptotically equivalent to the untargeted distortion γ, namely γ i→j /γ → 1 as n → ∞. Hence, under the stated assumptions, targeted attacks incur

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

11

Fig. 4. ASR vs. log(MSE) in the targeted adversarial attack scenario for all target classesa and across all image resolutions. Plots refer to the most robust models trained on the VegSeed dataset. Each colour in the bundles refers to a different target class (ASRj ).

only a vanishing relative distortion overhead compared with untargeted attacks. It is worth emphasizing that the result is established for the strongest notion of targeted adversarial risk, namely for each individual source–target pair (i, j). Consequently, it holds a fortiori for the averaged definitions of targeted risk, since averaging can only decrease the worst-case requirement. With regard to the condition Vj ≥ V0 > 0 for all j, this is a reasonable assumption in typical classification settings under mild structural conditions. In particular, if the number of classes m does not grow with the input dimension n, and if each class has non-vanishing prior probability and is represented by a non-degenerate region of the input space, then a classifier achieving good accuracy (in the absence of attacks) must assign a non-negligible portion of [0, 1]n to each class. Since the decision regions form a partition of the hypercube with unitary total volume, their average volume is 1/m, and it is natural to assume the existence of a dimension-independent lower bound V0 such that Vj ≥ V0 for all j. It is worth mentioning that while our analysis is carried out under the CI definition, in realistic high-dimensional settings, when the perturbation magnitude is small and the classifier achieves sufficiently high standard accuracy, the CI, ER, and PC definitions tend to be nearly equivalent, as discussed in Section III. In fact, extending the present theorem to the ER definition would require an additional structural assumption on the classifier, namely that for every pair i ̸= j the error regions Ci ∩ Fj are non-empty. Indeed, under the ER perspective, targeted adversarial examples can only be constructed if the classifier already misclassifies some samples from class i as class j. While this condition is typically satisfied in practical large-scale classification systems with non-zero error rates across classes, it cannot be guaranteed in full generality without explicitly assuming the existence of such cross-class error regions. VIII. TARGETED A DVERSARIAL E XAMPLES : E MPIRICAL A NALYSIS In this section, we empirically analyze the behavior of targeted adversarial attacks as input dimensionality increases. Using the same hierarchical datasets, models, and attack methodology used for the untargeted case, we evaluate the distortion required to enforce specific target classes and compare it with the untargeted case. The goal of this analysis is to

determine whether the additional control imposed by targeted attacks significantly increases the distortion required to craft adversarial examples and how such an increase depends on input dimensionality. Specifically, we empirically assess how the adversarial risk Riskjε for a given target class j evolves when the input dimensionality grows. To do so, given a class j, we define a target-dependent subset of the dataset, denoted as Nj , containing all samples x that are neither ground-truth members of class j nor currently assigned to class j by the classifier f , formally: Nj = {x ∈ D | x ∈ / Cj ∪ Fj }.

(22)

This set represents the pool of valid candidates for a targeted attack toward class j, ensuring that the attack is performed only on samples that neither belong to nor are already classified as the target class prior to perturbation. The attack is performed for each subset Nj , by adopting the same evaluation setting described in Section IV-B, while fixing α = 0.01. In particular, the attack on a sample x ∈ Nj aims at constructing an adversarial example xi→j ∈ Fj with the minimum perturbation. With this setting, and for any maximum distortion, we computed the targeted attack success rate ASRj for the target class j defined as ASRj =

1 X 1 [f (x′i ) ∈ Fj ] , |Nj |

(23)

xi ∈Nj

where, as usual, x′i is the output of the attack. It is worth noting that ASRj differs slightly from Riskjε since it does not account for the samples that are already misclassified by f as class j. We decided to rely on ASRj for our analysis since it directly gives a measure of how successful the attack is regardless of the misclassification due to the nonideality of the classifier, and because for accurate classifiers the fraction of samples already misclassified in the target class is typically negligible. The results we obtained for the VegSeed dataset are reported in Fig. 412 , where the curves of ASRj are shown as a function of the MSE for all target classes. Three main observations can be drawn from these results. First, for all target classes, the dependence of ASRj on the distortion closely follows the 12 Similar results were obtained for the other datasets, but are not reported here due to space limitations.

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

12

Fig. 5. Average ASR vs. log(MSE) in the targeted adversarial attack scenario for the most robust models across datasets. From left: VegSeed, Food-30, and ResynthDB dataset. Legend values indicate the MSE corresponding to an ASR of 90% for each model.

Fig. 6. ASR vs. log(MSE) in the worst-case targeted adversarial attack scenario for the most robust models across datasets. Legend values indicate the MSE corresponding to an ASR of 90% for each model.

same monotonic behavior observed in the untargeted case, although with a target-dependent shift in the distortion required to achieve a given success rate. Second, the target class significantly affects the attack difficulty, with some targets requiring substantially larger perturbations than others. Third, as the input dimensionality increases, all curves systematically shift leftward, showing that targeted adversarial examples also become easier to craft in higher dimension. To account for the dependence on the target class, Fig. 5 reports the ASR averaged over all target classes, thus paralleling the definition of Risktε . In addition, Fig. 6 reports a worstcase setting corresponding to the most difficult target class. In both cases, and across all datasets, the distortion required to achieve a given ASR decreases as the input dimension increases, confirming that the dimensionality effect observed for untargeted attacks persists in the targeted setting. We also carried out a detailed pairwise analysis by considering all source-target class pairs and computing either the distortion required to reach a fixed target success rate or the achieved ASR for a fixed distortion level. The corresponding results, obtained in the form of full source-target matrices, are not reported here due to space limitations. Nevertheless, they consistently confirm the same qualitative behavior observed above: the dependence on input dimensionality remains unchanged across all source-target pairs, and no source-target combination resulted in practically unattainable attacks. Finally, we quantify the additional distortion required to move from untargeted to targeted attacks. Table VI reports the MSE corresponding to ASR= 90% in the untargeted (U) and targeted (T) cases, together with the absolute difference (T−U) and the ratio T/U, respectively averaged across target

classes and evaluated in the worst-case setting. A first clear observation is that targeted attacks systematically require larger perturbations than untargeted ones, as expected from the additional control imposed on the attack outcome. At the same time, the absolute distortion gap T−U decreases consistently across all datasets and in both averaging settings as the input dimensionality increases. This behavior is in line with the theoretical result proved in Section VII, according to which the additional distortion required to enforce a prescribed target class becomes asymptotically negligible. In most cases, the ratio T/U also exhibits a decreasing trend as the input size grows, in agreement with the theoretical prediction that the relative overhead of targeted attacks should progressively vanish. The few non-monotonic behaviors occur mainly at the highest resolutions, particularly for the ResynthDB dataset, where the perturbations become extremely small and ratio estimates are correspondingly more sensitive to numerical fluctuations. Overall, these results indicate that, even in realistic finitedimensional settings, the additional cost required to impose a specific target class remains limited and tends to become less significant when the input dimension increases. IX. C ONCLUSIONS In this paper we investigated the role of input dimensionality in the emergence of adversarial examples by combining theoretical analysis and large-scale experiments on hierarchical image datasets specifically designed to isolate the effect of image resolution. Our results show that, even if the assumptions underlying concentration-based impossibility results are violated in realistic image domains due to the strong localization of class distributions, adversarial examples

JOURNAL OF LATEX CLASS FILES, VOL. 14, NO. 8, AUGUST 2021

13

TABLE VI TARGETED (T) VS U NTARGETED (U) MSE, ON AVERAGE AND WORST CASE SCENARIOS . Dataset

Case

Avg VegSeed Worst

Avg Food-30 Worst

Avg ResynthDB Worst

MSE U T T-U T/U T T-U T/U U T T-U T/U T T-U T/U U T T-U T/U T T-U T/U

64 6.67e-4 2.55e-3 1.88e-3 3.82 4.17e-3 3.51e-3 6.25 3.15e-4 1.12e-3 8.05e-4 3.56 1.72e-3 1.40e-3 5.46 7.11e-5 2.46e-4 1.75e-4 3.46 4.59e-4 3.88e-4 6.46

128 1.00e-4 3.32e-4 2.32e-4 3.32 4.50e-4 3.50e-4 4.50 1.49e-4 4.72e-4 3.23e-4 3.17 7.46e-4 5.97e-4 5.00 3.75e-5 8.51e-5 4.76e-5 2.27 1.28e-4 9.03e-5 3.41

256 7.36e-5 1.60e-4 8.64e-5 2.17 2.19e-4 1.45e-4 2.97 8.65e-5 2.53e-4 1.67e-4 2.92 3.59e-4 2.72e-4 4.15 6.48e-6 1.63e-5 9.82e-6 2.52 2.41e-5 1.76e-5 3.71

400 4.32e-5 9.71e-5 5.39e-5 2.25 1.36e-4 9.28e-5 3.15 5.66e-5 1.49e-4 9.24e-5 2.63 2.50e-4 1.93e-4 4.41 3.28e-6 1.28e-5 9.52e-6 3.90 2.13e-5 1.80e-5 6.48

become consistently easier to craft as the input dimension increases. This behavior was observed across multiple datasets, network architectures, and defense strategies, and holds for both untargeted and targeted attacks. For targeted attacks, we extended existing geometric concentration arguments, showing that under mild assumptions the additional distortion required to enforce a prescribed target class becomes asymptotically negligible. Experimental results confirm that, in practical settings, targeted attacks require only a limited extra distortion with respect to the untargeted case, and that this gap tends to shrink as dimensionality grows. Taken together, these findings provide strong evidence of a curse of dimensionality underlying adversarial vulnerability. At the same time, they leave open a fundamental question: whether this phenomenon primarily originates from highdimensional geometry itself, from the actual structure of real data distributions, or from specific properties of deep neural network classifiers. R EFERENCES [1] C. Szegedy, W. Zaremba, I. Sutskever, J. Bruna, D. Erhan, I. Goodfellow, and R. Fergus, “Intriguing properties of neural networks,” arXiv preprint arXiv:1312.6199, 2013. [2] M. Ledoux, “The concentration of measure phenomenon,” American Mathematical Soc., vol. Number 89, 2001. [3] A. Fawzi, O. Fawzi, and P. Frossard, “Analysis of classifiers’ robustness to adversarial perturbations,” Machine learning, vol. 107, no. 3, pp. 481– 508, 2018. [4] A. Shafahi, W. R. Huang, C. Studer, S. Feizi, and T. Goldstein, “Are adversarial examples inevitable?” in International Conference on Learning Representations (ICLR), vol. 11, 2019, pp. 8324–8340. [5] S. Mahloujifar, D. I. Diochnos, and M. Mahmoody, “The curse of concentration in robust learning: Evasion and poisoning attacks from concentration of measure,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, no. 01, 2019, pp. 4536–4543. [6] J. Gilmer, L. Metz, F. Faghri, S. S. Schoenholz, M. Raghu, M. Wattenberg, and I. Goodfellow, “Adversarial spheres,” arXiv preprint arXiv:1801.02774, 2018. [7] J. Gilmer, N. Ford, N. Carlini, and E. Cubuk, “Adversarial examples are a natural consequence of test error in noise,” in International Conference on Machine Learning. PMLR, 2019, pp. 2280–2289.

[8] E. Dohmatob, “Generalized no free lunch theorem for adversarial robustness,” in International Conference on Machine Learning. PMLR, 2019, pp. 1646–1654. [9] G. De Palma, B. Kiani, and S. Lloyd, “Adversarial robustness guarantees for random deep neural networks,” in International Conference on Machine Learning. PMLR, 2021, pp. 2522–2534. [10] A. Pal, J. Sulam, and R. Vidal, “Adversarial examples might be avoidable: The role of data concentration in adversarial robustness,” Advances in Neural Information Processing Systems, vol. 36, pp. 46 989–47 015, 2023. [11] A. Fawzi, H. Fawzi, and O. Fawzi, “Adversarial vulnerability for any classifier,” in Advances in neural information processing systems, 31 (Neurips 2018), 2018. [12] D. Diochnos, S. Mahloujifar, and M. Mahmoody, “Adversarial risk and robustness: General definitions and implications for the uniform distribution,” Advances in Neural Information Processing Systems, vol. 31, 2018. [13] S. Mahloujifar, X. Zhang, M. Mahmoody, and D. Evans, “Empirically measuring concentration: Fundamental limits on intrinsic robustness,” Advances in Neural Information Processing Systems, vol. 32, 2019. [14] J. Prescott, X. Zhang, and D. Evans, “Improved estimation of concentration under ell p-norm distance metrics using half spaces,” arXiv preprint arXiv:2103.12913, 2021. [15] X. Zhang and D. D. Evans, “Incorporating label uncertainty in understanding adversarial robustness.” in International Conference on Learning Representations (ICLR), vol. 2022, 2022. [16] N. Carlini and D. Wagner, “Towards evaluating the robustness of neural networks,” in Proceedings of the IEEE Symposium on Security and Privacy, 2017, pp. 39–57. [17] F. Croce and M. Hein, “Reliable evaluation of adversarial robustness with an ensemble of diverse attacks,” in International Conference on Machine Learning (ICML), 2020. [18] G. Tolias, F. Radenović, and O. Chum, “Targeted mismatch adversarial attack: Query with a flower to retrieve the tower,” in Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2019, pp. 5037–5046. [19] Y. Dong, F. Liao, T. Pang, H. Su, J. Zhu, X. Hu, and J. Li, “Boosting adversarial attacks with momentum,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2018, pp. 9185–9193. [20] N. Akhtar and A. Mian, “Threat of adversarial attacks on deep learning in computer vision: A survey,” Ieee Access, vol. 6, pp. 14 410–14 430, 2018. [21] M. H. Ferdaus, S. R. A. Ohona, R. H. Prito, and M. Ahmed, “VegSeedsBD: A comprehensive image dataset of vegetable seeds,” 2025. [Online]. Available: https://data.mendeley.com/datasets/dtpzbwwpm7/1 [22] L. Bossard, M. Guillaumin, and L. Van Gool, “Food-101 – mining discriminative components with random forests,” in European Conference on Computer Vision, 2014. [23] P. Bongini, V. Molinari, A. Costanzo, B. Tondi, and M. Barni, “Trainingfree source attribution of ai-generated images via resynthesis,” in 2025 IEEE International Workshop on Information Forensics and Security (WIFS), 2025. [24] A. Howard, M. Sandler, G. Chu, L.-C. Chen, B. Chen, M. Tan, W. Wang, Y. Zhu, R. Pang, V. Vasudevan et al., “Searching for mobilenetv3,” in Proceedings of the IEEE/CVF international conference on computer vision, 2019, pp. 1314–1324. [25] M. Tan and Q. Le, “Efficientnet: Rethinking model scaling for convolutional neural networks,” in International conference on machine learning. PMLR, 2019, pp. 6105–6114. [26] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 770–778. [27] A. Radford, J. W. Kim, C. Hallacy, A. Ramesh, G. Goh, S. Agarwal, G. Sastry, A. Askell, P. Mishkin, J. Clark et al., “Learning transferable visual models from natural language supervision,” in International conference on machine learning. PmLR, 2021, pp. 8748–8763. [28] J. Deng, W. Dong, R. Socher, L.-J. Li, K. Li, and L. Fei-Fei, “Imagenet: A large-scale hierarchical image database,” in 2009 IEEE Conference on Computer Vision and Pattern Recognition, 2009, pp. 248–255. [29] A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” arXiv preprint arXiv:1706.06083, 2017. [30] W. Zuo, K. Zhang, and L. Zhang, Convolutional Neural Networks for Image Denoising and Restoration. Cham: Springer International Publishing, 2018, pp. 93–123. [Online]. Available: https://doi.org/10. 1007/978-3-319-96029-6 4

Record · ID 310773 · SHA-256 18dcef0d1493992b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.