ConceptioArchivearXiv CS
arXiv CSopen access

VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptography, security, privacy, cybersecurity

VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness Niklas Bunzel Fraunhofer SIT / TU Darmstadt / ATHENE Darmstadt, Germany [email protected]

arXiv:2607.12545v1 [cs.CR] 14 Jul 2026

ABSTRACT Adversarial robustness research has produced hundreds of defended models over the past decade, yet the literature almost universally reports robustness results in isolation: standard (clean) accuracy and adversarial accuracy of the robust model are shown, but the gap to the corresponding vanilla (non-adversarially-trained) model is rarely quantified. We introduce VanillaBench, a systematic benchmark that makes this gap explicit. For every adversariallytrained model catalogued by RobustBench across four threat models (CIFAR-10 ℓ∞ 8/255, CIFAR-10 ℓ2 0.5, CIFAR-100 ℓ∞ 8/255, and ImageNet ℓ∞ 4/255), we compute the accuracy difference Δclean = accrobust − accvanilla against multiple vanilla references: the best, median, and top-10 median standard models from Papers with Code, computed over both all entries and no-extra-data entries, the best vanilla model as of the robust model’s publication year, and an architecture-matched baseline. Across all 186 robust models, the mean Δclean relative to the best vanilla model ranges from −7.7 to −29.5 percentage points, and even the single most robust model per track still trails its temporal vanilla counterpart by 4.0–21.0 points. The architecture-matched comparison, which isolates the effect of adversarial training from architectural differences, reveals a mean gap of −3.5 to −17.5 points. Restricting this architecturematched comparison to models whose vanilla accuracy is known for the exact same architecture, rather than approximated from a related one, narrows the gap to −4.0 to −14.0 points. These results demonstrate that the robustness-accuracy trade-off is substantially larger than what is typically conveyed by individual papers. This information is critical for practitioners and decision-makers. When deploying models in real-world settings, the accuracy cost of robustness directly affects business outcomes, yet current publications do not provide the vanilla baseline needed to assess it. We argue that future robustness evaluations should report vanilla-referenced accuracy gaps as a standard component.

KEYWORDS adversarial robustness, accuracy-robustness trade-off, benchmark, image classification

1

INTRODUCTION

Since the discovery that deep neural networks are vulnerable to adversarial perturbations [13, 32], a large and active research community has developed a wide range of defense strategies. One prominent strain of work is adversarial training [4, 8, 23, 30, 33], which hardens models by incorporating adversarial examples directly into the training loop. A complementary strand comprises preprocessing defenses that transform or purify inputs before classification [14, 37], as well as adversarial example detectors that aim to flag suspicious

inputs before or alongside the classifier [6, 7, 17, 18]. While adversarial training has become the most widely studied and benchmarked family of defenses, the standard evaluation paradigm for a new robust model is largely shared across the literature: it consists of reporting two numbers, the clean accuracy (standard top-1 accuracy on unperturbed test data) and the robust accuracy (accuracy under a standardized adversarial attack such as AutoAttack [10]). These two numbers are then compared, either to each other or to prior robust models, to demonstrate progress. What is almost always missing from this evaluation is the vanilla baseline. A reader who encounters a paper reporting, say, 93.6% clean accuracy for an adversarially trained WideResNet [40] on CIFAR-10 [20] has no immediate way of knowing what a nonadversarially-trained model of the same architecture achieves, or, more broadly, what the state of the art in standard classification looks like. This omission matters because the entire value proposition of adversarial training research rests on a trade-off: how much standard accuracy does one sacrifice, and how much robustness does one gain? Without referencing the vanilla baseline, the cost side of this trade-off is invisible. The problem is exacerbated by the fact that robustness papers typically use architectures (e.g., WideResNet-70-16 [35], WideResNet34-10 [28]) whose vanilla counterparts are not necessarily wellknown or easily found. A fair comparison requires either (i) matching the architecture exactly, (ii) matching the temporal context (what was the best vanilla model at the time the robust model was published?), or (iii) comparing against a broad reference distribution of vanilla models. This information gap has direct business consequences. An organization deciding whether to adopt adversarially training must weigh the robustness benefit against the standard-accuracy cost: lower accuracy means more misclassifications, degraded user experience, potential revenue loss. Yet when a robustness paper reports a clean accuracy of, say, 84% without referencing what a vanilla model achieves on the same task, the decision-maker cannot determine whether the 84% represents a 2-point or a 10-point sacrifice. The difference between these two scenarios leads to qualitatively different business decisions: a 2-point gap may be an acceptable insurance premium against adversarial attacks, while a 10-point gap may render the model unusable for production. By making the gap explicit across the entire robustness literature, VanillaBench provides the missing quantitative basis for such decisions. We address this gap with VanillaBench: a comprehensive evaluation that takes all 186 adversarially-trained models from RobustBench [9] across four threat models and computes the cleanaccuracy gap to multiple vanilla references. Specifically, our contributions are:

Niklas Bunzel

• We assemble vanilla reference statistics (best, median, top10 median) from Papers with Code [1] for CIFAR-10 [20], CIFAR-100 [20], and ImageNet-1k [27], computed over both all entries and no-extra-data entries. • We construct a temporal reference: the best vanilla accuracy achievable as of each robust model’s publication year, enabling a year-matched comparison that controls for the rapid progress of standard classification. • We curate an architecture-matched reference: hand-matched standard accuracies for each robust model’s underlying architecture, sourced from the torchvision model zoo [2] for ImageNet, pytorch-cifar project [21] for CIFAR-10, pytorchcifar100 project [36] for CIFAR-100 and from the original architecture papers for CIFAR-10/100 and ImageNet. • We compute Δclean , the gap between the robust model’s clean accuracy and each vanilla reference, for all 186 models, and we show that this gap is universally negative and often large: even top-10 robust models lag the best vanilla model by 6–21 percentage points depending on the track.

provide a fairer comparison to robust models trained on the standard training set alone. This yields six reference points per track. The VanillaBench website additionally reports the mean, top-10 mean, and top-10 minimum for the all-entries subset. Finally, we compute a temporal best: for each year 𝑦, the best accuracy among all entries published in or before 𝑦, enabling a year-matched comparison that controls for the rapid progress of standard classification over time. 2.2.2 Architecture-matched baselines. Because PwC statistics aggregate over heterogeneous architectures, they do not isolate the effect of adversarial training from architectural differences. To address this, we curate a hand-matched architecture table. For each robust model, we identify the underlying base architecture (e.g., WideResNet-70-16, ResNet-50 [15], ConvNeXt-L [22]) and look up its standard accuracy from: • The torchvision model zoo (IMAGENET1K_V1 weights [2]) for ImageNet architectures, using the original published recipe (V1) rather than improved-recipe V2 weights. • The pytorch-cifar [21] and pytorch-cifar100 [36] projects for CIFAR-10/100 data. • The original architecture papers for CIFAR and ImageNet data (e.g., Zagoruyko & Komodakis [40] for WideResNet variants on CIFAR-10).

2 THE VANILLABENCH METHODOLOGY 2.1 Robust Models We use the RobustBench leaderboard [9] as our source of adversariallytrained models. RobustBench provides standardized, reliable robustness evaluations using AutoAttack [10], an ensemble of four attack methods that is widely accepted as a strong, parameter-free evaluation protocol. We consider all four tracks available at the time of this work: • CIFAR-10 ℓ∞ 8/255: 98 robust models, published 2018– 2024. • CIFAR-10 ℓ2 0.5: 20 robust models, published 2019–2024. • CIFAR-100 ℓ∞ 8/255: 38 robust models, published 2019– 2024. • ImageNet ℓ∞ 4/255: 30 robust models, published 2019– 2024. For each model, RobustBench records the clean accuracy (top-1 on the standard test set), the AutoAttack accuracy (top-1 under the standardized attack), the architecture, the publication venue and year, and whether the model was trained with additional data beyond the standard training set.

2.2

Vanilla References

We construct multiple vanilla reference points for each dataset, drawing from different sources. 2.2.1 Papers with Code statistics. From the Papers with Code (PwC) leaderboards for image classification on CIFAR-10, CIFAR-100, and ImageNet-1k, we extract all entries and compute three statistics: • Best: the highest standard accuracy among the entries. • Median: the median standard accuracy. • Top-10 median: the median of the 10 highest-accuracy entries. We compute each statistic over two subsets: (i) all entries, including models that leverage external pretraining on massive datasets such as JFT-300M [31] or Instagram billion-scale data [24], and (ii) no-extra-data entries only, which exclude such models and thus

Where the robust model uses a modified architecture (e.g., RaWideResNet [26], ConvStem-ViT [30]), we match to the closest canonical base. When no exact vanilla number is published for an architecture, we use a proxy match: a closely related architecture whose vanilla accuracy is known (e.g., WideResNet-70-16 is proxied to WideResNet-40-8 [40]). Architecture matching (including proxy matches) is available for 184 of 186 robust models (98.9%); of these, 90 (48.4%) are direct matches and 94 are proxy matches. We report direct-match-only statistics separately to quantify the impact of proxy matching.

2.3

The Accuracy Gap Metric

For a robust model 𝑚 with clean accuracy 𝑎𝑚 and a vanilla reference value 𝑟 , we define the accuracy gap: Δclean (𝑚, 𝑟 ) = 𝑎𝑚 − 𝑟 . A negative Δclean indicates that the robust model sacrifices standard accuracy relative to the reference. We compute Δclean for every robust model against every applicable reference. We emphasize that Δclean captures only the cost side of the robustness-accuracy trade-off. The benefit side (robust accuracy) is reported in parallel but is not folded into a single scalar, since the trade-off is inherently two-dimensional.

3 RESULTS 3.1 Overview Table 1 presents the headline results: for each track, the number of robust models, the mean clean and robust accuracies, the best (all-entries) reference, and the mean Δclean against five reference types. The gap is negative in every cell of the Δ columns. No robust model, on average, matches any vanilla reference. Several patterns emerge:

VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness

Table 1: Overview of all four VanillaBench tracks. All references are computed over all entries (including models that use additional training data). “Δ best” = mean Δclean vs. best model; “Δ median” = vs. median; “Δ t10med” = vs. top-10 median; “Δ arch” = vs. architecture-matched baseline (including proxy matches); “Δ year” = vs. best model (all entries) as of the robust model’s publication year. All values in percentage points. Track

𝑁

Clean mean

Robust mean

Best (all)

Δ best

Δ median

Δ t10med

Δ arch

Δ year

CIFAR-10 ℓ∞ 8/255 CIFAR-10 ℓ2 0.5 CIFAR-100 ℓ∞ 8/255 ImageNet ℓ∞ 4/255

98 20 38 30

87.49 91.79 66.61 73.12

53.44 76.36 32.12 46.69

99.50 99.50 96.08 91.00

−12.01 −7.71 −29.47 −17.88

−11.41 −7.11 −24.34 −9.48

−11.79 −7.50 −27.19 −16.73

−7.59 −3.50 −17.47 −8.52

−11.92 −7.70 −29.40 −17.76

−1.02 pp (vs. no-extra-data median), achieved by Singh et al. [30] (ConvNeXt-B + ConvStem). • The architecture-matched comparison reveals only two cases where a robust model exceeds its matched vanilla baseline: +0.40 on CIFAR-10 ℓ2 and +0.39 on ImageNet ℓ∞ . Both are small positive margins. The earlier +21.08 outlier (Sehwag et al. [29], ResNet-18) has been resolved by updating the ResNet-18 CIFAR-10 baseline to its correct value (93.02%). • The direct-match-only statistics (“Direct only” rows) exclude proxy-matched architectures. On CIFAR-10 ℓ∞ and ℓ2 , the direct-match gaps are similar to or slightly larger than the full gaps, suggesting that proxy matching does not systematically bias the results on these tracks.

The best-vs.-median distinction matters. The gap to the best vanilla model is dramatically larger than the gap to the median vanilla model, especially on ImageNet (−17.88 vs. −9.48). This reflects the wide spread of the PwC leaderboards, which range from 38.1% (one shot or few shot) to 91.0% on ImageNet. A robustness paper that reports clean accuracy without context can be compared favourably to the bottom of the distribution, making the gap appear smaller than it is. The temporal comparison is close to the overall best. On all four tracks, the year-matched gap (Δ year) is nearly identical to the all-time best gap (Δ best). This is because the best vanilla model was published early (2019–2020 for CIFAR, 2022 for ImageNet) and remained the reference throughout. The median year-matched reference equals the all-time best on every track (99.50/99.50/96.08/91.00%). The architecture-matched gap is smaller than the PwC gap. On all four tracks, the architecture-matched gap is smaller than the best-vanilla gap, indicating that part of the apparent gap stems from comparing architecturally weaker robust models to stronger vanilla models. On CIFAR-10 ℓ∞ , the architecture-matched gap (−7.59) is about two-thirds of the best-vanilla gap (−12.01). On CIFAR-100 (−17.47), the architecture-matched gap is smaller than the median gap (−24.34), because the architecture-matched vanilla references (primarily WideResNet-28-10 at 80.75% on CIFAR-100 and PreActResNet-18 [16] at 72.92%) are below the PwC median (90.95%). The ℓ2 threat model has the smallest gap. CIFAR-10 ℓ2 0.5 shows the smallest accuracy cost (−7.71 vs. best, −3.50 vs. architecturematched), consistent with the fact that ℓ2 -bounded perturbations at 𝜖 = 0.5 are a weaker threat than ℓ∞ at 8/255.

3.2

Per-Reference Detail

Table 2 provides a more detailed breakdown, reporting mean and median Δclean for each reference type on each track. Several observations stand out: • On CIFAR-10 ℓ∞ , the minimum Δclean against the best vanilla model is −54.77 pp, indicating that some robust models sacrifice more standard accuracy than they retain in robust accuracy. The maximum is −4.27 pp, achieved by Bai et al. [5] (MixedNUTS), which uses a hybrid approach combining a robust and a non-robust classifier. • On ImageNet, no robust model exceeds any aggregate vanilla reference. The smallest gap against any PwC reference is

3.3

Top-10 Robust Models

Table 3 focuses on the 10 most robust models (by AutoAttack accuracy) per track, which represent the state of the art in adversarial training. Even among these best-performing models, the accuracy gap remains substantial. Even the best-available robust models sacrifice 4.6–15.1 pp of standard accuracy relative to the median vanilla model. The gap widens to 5.7–21.2 pp relative to the best vanilla model. This confirms that the robustness–accuracy trade-off is not merely a problem for weaker models; it affects the entire frontier.

3.4

Temporal Trends

Table 4 breaks down the mean clean accuracy of robust models by publication year, alongside the year-matched best (all entries). On CIFAR-10, the vanilla best was already at 99.37% in 2019 (BiTL [19]), while the best mean robust clean accuracy in any year is 93.16% (2023). On ImageNet, the vanilla best reached 91.0% in 2022 (CoCa [39]), while the best mean robust clean accuracy is 77.35% (2024). The gap has narrowed over time but remains at 6+ pp on CIFAR-10 and 13+ pp on ImageNet even in the most recent year.

3.5

Per-Track Analysis

3.5.1 CIFAR-10 ℓ∞ 8/255. This is the most populated track with 98 models. The mean Δclean against the best vanilla model is −12.01 pp, and even the mean gap of the top-10 most robust models is −5.85 pp. The architecture-matched gap is smaller (−7.59 pp, median reference 96.00%), because many robust models use WideResNet-7016 [5, 26, 35] (vanilla: 95.34%, proxy-matched to WideResNet-40-8) or WideResNet-28-10 [11, 35, 38] (vanilla: 96.0%), architectures

Niklas Bunzel

Table 2: Detailed Δclean statistics (percentage points) for each reference type and track. “Δ mean” and “Δ median” refer to the distribution of Δclean across all robust models in the track. “All entries” includes models that use additional training data; “No extra data” excludes them. “𝑛 (arch)” = number of models with architecture-matched baselines available. “Direct only” = architecture-matched restricted to direct matches (excluding proxy-matched architectures). “Year-matched best (all)” = best model (all entries) as of the robust model’s publication year. For the architecture-matched, Direct only, and Year-matched rows, the Ref. value is the median of per-model reference values. Ref. value

Δ mean

Δ median

Δ min

Δ max

CIFAR-10 ℓ∞ 8/255

Best (all entries) Median (all entries) Top-10 median (all entries) Best (no extra data) Median (no extra data) Top-10 median (no extra data) Architecture-matched (𝑛 =96) Direct only (𝑛 =41) Year-matched best (all)

99.50 98.90 99.28 99.50 98.70 99.10 96.00 96.00 99.50

−12.01 −11.41 −11.79 −12.01 −11.21 −11.61 −7.59 −7.92 −11.92

−12.08 −11.48 −11.86 −12.08 −11.28 −11.67 −7.82 −7.75 −12.00

−54.77 −54.17 −54.55 −54.77 −53.97 −54.37 −15.76 −15.76 −54.77

−4.27 −3.67 −4.05 −4.27 −3.47 −3.87 −0.11 −2.18 −4.27

CIFAR-10 ℓ2 0.5

Best (all entries) Median (all entries) Top-10 median (all entries) Best (no extra data) Median (no extra data) Top-10 median (no extra data) Architecture-matched (𝑛 =20) Direct only (𝑛 =9) Year-matched best (all)

99.50 98.90 99.28 99.50 98.70 99.10 95.34 95.11 99.50

−7.71 −7.11 −7.50 −7.71 −6.91 −7.31 −3.50 −4.04 −7.70

−8.49 −7.90 −8.28 −8.49 −7.70 −8.09 −3.51 −4.21 −8.48

−11.48 −10.88 −11.27 −11.48 −10.68 −11.08 −7.98 −6.95 −11.48

−3.76 −3.16 −3.55 −3.76 −2.96 −3.36 +0.40 −0.84 −3.76

CIFAR-100 ℓ∞ 8/255

Best (all entries) Median (all entries) Top-10 median (all entries) Best (no extra data) Median (no extra data) Top-10 median (no extra data) Architecture-matched (𝑛 =38) Direct only (𝑛 =10) Year-matched best (all)

96.08 90.95 93.80 96.08 90.00 92.76 80.75 78.18 96.08

−29.47 −24.34 −27.19 −29.47 −23.39 −26.15 −17.47 −13.96 −29.40

−30.57 −25.45 −28.30 −30.57 −24.49 −27.25 −16.89 −13.74 −30.57

−42.25 −37.12 −39.97 −42.25 −36.17 −38.92 −34.48 −21.52 −42.25

−10.87 −5.74 −8.59 −10.87 −4.79 −7.55 −6.90 −6.90 −10.87

ImageNet ℓ∞ 4/255

Best (all entries) Median (all entries) Top-10 median (all entries) Best (no extra data) Median (no extra data) Top-10 median (no extra data) Architecture-matched (𝑛 =30) Direct only (𝑛 =30) Year-matched best (all)

91.00 82.60 89.85 91.00 82.50 89.85 83.12 83.12 91.00

−17.88 −9.48 −16.73 −17.88 −9.38 −16.73 −8.52 −8.52 −17.76

−15.72 −7.32 −14.57 −15.72 −7.22 −14.57 −7.67 −7.67 −16.34

−38.08 −29.68 −36.93 −38.08 −29.58 −36.93 −20.51 −20.51 −37.28

−9.52 −1.12 −8.37 −9.52 −1.02 −8.37 +0.39 +0.39 −9.52

Track

Reference

Table 3: Top-10 robust models per track: mean and median Δclean (pp) against each no-extra-data vanilla reference. Track CIFAR-10 ℓ∞ 8/255 CIFAR-10 ℓ2 0.5 CIFAR-100 ℓ∞ 8/255 ImageNet ℓ∞ 4/255

Top-10 clean mean 93.65 93.80 74.87 77.94

Δ vs. best mean median

Δ vs. median mean median

Δ vs. top-10 median mean median

−5.85 −5.70 −21.21 −13.06

−5.05 −4.90 −15.13 −4.56

−5.45 −5.30 −17.88 −11.91

whose standard accuracy is high but not at the 99.5% SOTA level.

−6.07 −5.15 −21.59 −13.03

−5.27 −4.35 −15.51 −4.53

−5.66 −4.75 −18.27 −11.88

The most robust model, Amini et al. [3] (MeanSparse WideResNet94-16), achieves 93.6% clean and 75.28% robust accuracy. This is −5.90 pp below the best vanilla model.

VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness

(a) CIFAR-10 ℓ∞ 8/255

(b) CIFAR-100 ℓ∞ 8/255

(c) ImageNet ℓ∞ 4/255

Figure 1: Clean accuracy vs. robust accuracy (AutoAttack) scatter plots for all ℓ∞ robust models. Each point represents one model; colour indicates publication year (blue = older ≈ 2017–2019, red = newer ≈ 2023–2024). (a) CIFAR-10 ℓ∞ (𝜖 = 8): no robust model exceeds 64% clean accuracy; vanilla top-10 median reference at 99.44%. (b) CIFAR-100 ℓ∞ (𝜖 = 8): similar gap, with robust models between 16–54% robust accuracy. (c) ImageNet ℓ∞ (𝜖 = 4): robust models cluster between 26–64% robust accuracy, far below the vanilla reference at 93.58%. The vertical gap to the vanilla line is the Δclean metric central to VanillaBench.

(a) CIFAR-10 ℓ∞ 8/255

(b) CIFAR-100 ℓ∞ 8/255

(c) ImageNet ℓ∞ 4/255

(d) CIFAR-10 ℓ2 0.5

Figure 2: Clean accuracy (blue) and AutoAttack robust accuracy (red) bar charts for robust models per track. (a) CIFAR-10 ℓ∞ (𝜖 = 8): robust models achieve 27–64% robust accuracy, while clean accuracy ranges from 82–94%. The red dashed line marks 99.44% (top-10 median vanilla reference). (b) CIFAR-100 ℓ∞ (𝜖 = 8): robust accuracy 16–54%, clean accuracy 79–95%, reference at 92.12%. (c) ImageNet ℓ∞ (𝜖 = 4): robust accuracy 26–64%, clean accuracy 50–84%, reference at 93.58%. (d) CIFAR-10 ℓ2 (𝜎 = 0.5): the only track where some robust models achieve comparable or superior clean accuracy relative to the vanilla reference (99.44%). The vertical gap between the blue bar and the dashed line is the Δclean that VanillaBench makes explicit, quantifying the clean accuracy sacrifice imposed by adversarial training. 3.5.2 CIFAR-10 ℓ2 0.5. This track has the smallest accuracy cost. The mean Δclean against the best vanilla model is −7.71 pp, and the architecture-matched gap is −3.50 pp (median reference 95.34%). The top model, Amini et al. [3] (MeanSparse WideResNet-70-16), achieves 95.51% clean and 87.28% robust accuracy. This is +0.17 pp above its proxy-matched vanilla baseline (WideResNet-40-8, 95.34%). This is consistent with the weaker perturbation budget of ℓ2 = 0.5.

3.5.3 CIFAR-100 ℓ∞ 8/255. This track exhibits the largest accuracy gap: −29.47 pp mean against the best vanilla model. The architecture-matched gap is −17.47 pp, smaller than the PwC median gap (−24.34 pp), because the architecture-matched vanilla references (primarily WideResNet-28-10 at 80.75% (the median architecture-matched reference) on CIFAR-100 and PreActResNet18 at 72.92%) are below the PwC median of 90.95%. The top model, Amini et al. [3] (MeanSparse WideResNet-70-16), achieves 75.13%

Niklas Bunzel

(a) CIFAR-10 ℓ∞ 8/255

(b) ImageNet ℓ∞ 4/255

Figure 3: Architecture-matched comparison: per-model bar charts showing clean accuracy (blue), AutoAttack robust accuracy (red), and architecture-matched vanilla accuracy (dark grey). The grey bar represents a vanilla model sharing the same architecture (or proxy architecture) as the robust model, isolating the clean accuracy cost of adversarial training from architectural differences. (a) CIFAR-10 ℓ∞ (𝜖 = 8): no robust model exceeds its architecture-matched vanilla baseline; gaps range from 4.8–7.8 pp for directly matched models. (b) ImageNet ℓ∞ (𝜖 = 4): larger gaps, with robust models falling 8.5–20 pp below their matched vanilla counterpart (direct match only). These results confirm that even when controlling for architecture, adversarial training incurs a substantial clean accuracy penalty. Proxy-matched models (94 total) use the nearest available vanilla architecture when an exact match is unavailable; direct-match-only analysis is reported in Table 2. Table 4: Mean clean accuracy of robust models by publication year, compared to the year-matched best (all entries) accuracy. Year 2018 2019 2020 2021 2022 2023 2024

CIFAR-10 ℓ∞ Robust mean Vanilla best 87.14 86.99 84.56 88.39 87.50 93.16 91.42

93.57 99.37 99.50 99.50 99.50 99.50 99.50

ImageNet ℓ∞ Robust mean Vanilla best — 62.56 60.25 — 72.64 75.88 77.35

80.62 87.54 90.20 90.20 91.00 91.00 91.00

clean and 44.78% robust accuracy. This is −20.21 pp below its proxymatched vanilla baseline. The CIFAR-100 results underscore that the robustness-accuracy trade-off is dataset-dependent and that progress on the harder 100-class setting has been much slower. 3.5.4 ImageNet ℓ∞ 4/255. On ImageNet, the mean Δclean is −17.88 pp against the best vanilla model but only −9.48 pp against the median vanilla model (all entries). This wide spread reflects the diversity of the ImageNet leaderboard (105 vanilla entries ranging from 38.1% to 91.0%). The architecture-matched gap (−8.52 pp, median reference 83.12%) is close to the PwC median gap (−9.48 pp), suggesting that on ImageNet, architectural matching does not substantially change the picture. The top model, Amini et al. [3] (MeanSparse Swin-L), achieves 78.80% clean and 62.12% robust accuracy. This is −12.20 pp below the best vanilla model. Notably, 29 of 30 ImageNet robust models use no additional training data, making this the “cleanest” track in terms of training setup.

4 DISCUSSION 4.1 Why the Vanilla Gap Is Rarely Shown The absence of vanilla baselines in robustness papers is not malicious. It reflects several practical and structural factors. First, adversarial training papers are written for the robustness community, where the implicit baseline is prior robust models, not vanilla models. Second, the architectures used for adversarial training (e.g., WideResNet-70-16) are often not the same as those that dominate standard leaderboards (e.g., ViT [12], ConvNeXt), making direct comparison awkward. Third, many robust models use additional data (26.5% on CIFAR-10 ℓ∞ , 23.7% on CIFAR-100), which further complicates the comparison. Should the vanilla reference also use additional data? VanillaBench addresses these complications by providing multiple references that span different notions of “vanilla.” The best and top-10 median references show the gap to the absolute state of the art. The median reference shows the gap to the “typical” model. Each of these is computed over both all entries and no-extradata entries, so the reader can assess whether the use of additional training data materially changes the picture. The year-matched reference controls for temporal progress. The architecture-matched reference isolates the adversarial training effect. Together, these provide a much richer picture than any single comparison.

4.2

The Two-Dimensional Trade-Off

We deliberately do not collapse Δclean and robust accuracy into a single scalar (e.g., a weighted average). The robustness–accuracy trade-off is fundamentally two-dimensional, and any scalarization hides information. Figure 1 makes the trade-off visible: each point occupies a position in (clean, robust) space, and the vanilla reference line provides the x-axis anchor that is typically missing. Researchers and practitioners can use this visualization to assess whether a given model’s trade-off point is “good” relative to both axes simultaneously.

VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness

4.3

Threats to Validity

Architecture matching imperfections. Our architecture-matched baselines are hand-curated and sometimes based on values from the original architecture papers rather than standardized re-evaluations. Additionally, 94 of 184 architecture-matched entries are proxy matches. The robust model’s exact architecture has no published vanilla number, so a closely related architecture is used instead (e.g., WideResNet70-16 is proxied to WideResNet-40-8). We report direct-match-only statistics separately to quantify the impact of this proxy matching. Papers with Code curation. The PwC leaderboards have been curated/trimmed by Hugging Face, so the reference statistics reflect the live curated boards rather than an exhaustive historical list. The number of vanilla entries is relatively small for CIFAR (20–22), which means the statistics are sensitive to the inclusion or exclusion of individual entries. Additional data flag. Our “no extra data” filter relies on PwC’s uses_additional_data flag, which may not perfectly capture all nuances (e.g., self-supervised pretraining on the same dataset is sometimes considered “additional data” and sometimes not). AutoAttack as the sole robustness metric. We use the AutoAttack accuracy from RobustBench as the robustness measure. AutoAttack is a strong standardized attack, but robust accuracy under other threat models (e.g., corrupted inputs, distribution shift) may tell a different story.

5

RELATED WORK

Adversarial robustness benchmarks. RobustBench [9] is the most widely used standardized benchmark for adversarial robustness, providing a curated leaderboard of models evaluated under AutoAttack. However, RobustBench focuses exclusively on robust models and does not provide vanilla baselines. VanillaBench complements RobustBench by adding the missing vanilla reference layer. The accuracy–robustness trade-off. The fundamental tension between standard and robust accuracy was noted early on [34, 41]. Tsipras et al. [34] provided a theoretical analysis showing that the trade-off is inherent, not merely an artifact of suboptimal training. Zhang et al. [41] decomposed robust and standard error into boundary and interior components. However, these works analyze the trade-off theoretically or on small-scale experiments. They do not systematically quantify the gap across the entire robustness literature against vanilla references. Robustness without accuracy loss. Several works have argued that the trade-off can be “reconciled” [25] or mitigated through better training procedures [5, 35]. Our results confirm that progress has been made. The most recent models are closer to vanilla accuracy than early ones, but the gap is still substantial, especially on harder datasets like CIFAR-100. Standard classification benchmarks. Papers with Code [1] provides the standard accuracy leaderboards that we use as vanilla references. The torchvision model zoo [2] provides verified, reproducible standard accuracies for common architectures.

6

CONCLUSION & FUTURE WORK

We have presented VanillaBench, a systematic evaluation of the accuracy cost of adversarial robustness. By computing the cleanaccuracy gap between every RobustBench model and multiple vanilla references, including best, median, top-10 median (all entries and no-extra-data), year-matched, and architecture-matched, we make explicit a trade-off that is typically only implicit in the literature. The gap is universally negative. Across 186 models and four threat models, no robust model matches the best vanilla model, and even top-10 robust models sacrifice 4.6–15.1 pp of standard accuracy relative to the median vanilla model. The architecturematched comparison, which isolates the adversarial training effect, confirms that the gap is not merely an artifact of comparing different architectures. We recommend that future robustness papers report vanilla-referenced accuracy gaps as a standard component of evaluation, and we provide the VanillaBench tooling1 to facilitate this. Several directions remain for future work. First, the vanilla reference accuracies used in this work are sourced from published papers and model zoos, which may reflect different evaluation protocols, preprocessing pipelines, or training recipes. Conducting standardized in-house evaluations of both robust and vanilla models under a unified protocol would eliminate this source of variability and yield more directly comparable gap estimates. Second, a number of architectures that appear frequently in adversarial training (e.g., WideResNet-70-16, WideResNet-94-16) lack published vanilla accuracies, forcing us to rely on proxy matches. Training and evaluating vanilla versions of these architectures would remove the need for proxy matching and tighten the architecture-matched comparison. Third, the most recent model in RobustBench dates to 2024, and the benchmark datasets (CIFAR-10, CIFAR-100, ImageNet-1k) have remained unchanged for years. Extending VanillaBench to encompass newer robust models and additional, more modern evaluation datasets would keep the benchmark current and broaden its applicability.

ACKNOWLEDGMENTS This work was supported by the ATHENE flagship project funded by the German Federal Ministry of Education and Research (BMBF).

REFERENCES [1] 2024. Papers with Code: Image Classification. https://paperswithcode.co/task/ image-classification. [2] 2024. torchvision: Classification Model Weights. https://docs.pytorch.org/vision/ stable/models.html. [3] Sajjad Amini, Mohammadreza Teymoorianfard, Shiqing Ma, and Amir Houmansadr. 2024. MeanSparse: Post-training robustness enhancement through mean-centered feature sparsification. arXiv preprint arXiv:2406.05927 (2024). [4] Tao Bai, Jinqi Luo, Jun Zhao, Bihan Wen, and Qian Wang. 2021. Recent Advances in Adversarial Training for Adversarial Robustness. In Proceedings of the Thirtieth International Joint Conference on Artificial Intelligence, IJCAI 2021, Virtual Event / Montreal, Canada, 19-27 August 2021, Zhi-Hua Zhou (Ed.). ijcai.org, 4312–4321. https://doi.org/10.24963/IJCAI.2021/591 [5] Yatong Bai, Mo Zhou, Vishal M. Patel, and Somayeh Sojoudi. 2024. MixedNUTS: Training-Free Accuracy-Robustness Balance via Nonlinearly Mixed Classifiers. Trans. Mach. Learn. Res. 2024 (2024). https://openreview.net/forum?id= pyD6cujUmL [6] Niklas Bunzel and Dominic Böringer. 2023. Multi-class Detection for Off The Shelf transfer-based Black Box Attacks. In Proceedings of the 2023 Secure and Trustworthy Deep Learning Systems Workshop. 1–6. 1 https://bunni90.github.io/robust-vs-vanilla.html

Niklas Bunzel

[7] Niklas Bunzel, Ashim Siwakoti, and Gerrit Klause. 2023. Adversarial Patch Detection and Mitigation by Detecting High Entropy Regions. In 53rd Annual IEEE/IFIP International Conference on Dependable Systems and Networks, DSN 2023 - Workshops, Porto, Portugal, June 27-30, 2023. IEEE, 124–128. https://doi. org/10.1109/DSN-W58399.2023.00040 [8] Prakash Chandra Chhipa, Gautam Vashishtha, Settur Jithamanyu, Rajkumar Saini, Mubarak Shah, and Marcus Liwicki. 2025. ASTrA: Adversarial Selfsupervised Training with Adaptive-Attacks. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net. https://openreview.net/forum?id=ZbkqhKbggH [9] Francesco Croce, Maksym Andriushchenko, Vikash Sehwag, Edoardo Debenedetti, Nicolas Flammarion, Mung Chiang, Prateek Mittal, and Matthias Hein. 2021. RobustBench: a standardized adversarial robustness benchmark. In Proceedings of the Neural Information Processing Systems Track on Datasets and Benchmarks 1, NeurIPS Datasets and Benchmarks 2021, December 2021, virtual, Joaquin Vanschoren and Sai-Kit Yeung (Eds.). https://datasets-benchmarks-proceedings.neurips.cc/paper/2021/hash/ a3c65c2974270fd093ee8a9bf8ae7d0b-Abstract-round2.html [10] Francesco Croce and Matthias Hein. 2020. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In Proceedings of the 37th International Conference on Machine Learning, ICML 2020, 13-18 July 2020, Virtual Event (Proceedings of Machine Learning Research, Vol. 119). PMLR, 2206–2216. http://proceedings.mlr.press/v119/croce20b.html [11] Jiequan Cui, Zhuotao Tian, Zhisheng Zhong, Xiaojuan Qi, Bei Yu, and Hanwang Zhang. 2024. Decoupled Kullback-Leibler Divergence Loss. In Advances in Neural Information Processing Systems 37: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, Amir Globersons, Lester Mackey, Danielle Belgrave, Angela Fan, Ulrich Paquet, Jakub M. Tomczak, and Cheng Zhang (Eds.). http://papers.nips.cc/paper_files/paper/2024/hash/ 87ee1bbac4635e7c948f3eea83c1f262-Abstract-Conference.html [12] Alexey Dosovitskiy, Lucas Beyer, Alexander Kolesnikov, Dirk Weissenborn, Xiaohua Zhai, Thomas Unterthiner, Mostafa Dehghani, Matthias Minderer, Georg Heigold, Sylvain Gelly, Jakob Uszkoreit, and Neil Houlsby. 2021. An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale. (2021). https://openreview.net/forum?id=YicbFdNTTy [13] Ian J. Goodfellow, Jonathon Shlens, and Christian Szegedy. 2015. Explaining and Harnessing Adversarial Examples. (2015). http://arxiv.org/abs/1412.6572 [14] Chuan Guo, Mayank Rana, Moustapha Cisse, and Laurens van der Maaten. 2018. Countering Adversarial Images using Input Transformations. In International Conference on Learning Representations. https://openreview.net/forum? id=SyJ7ClWCb [15] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Deep Residual Learning for Image Recognition. In 2016 IEEE Conference on Computer Vision and Pattern Recognition, CVPR 2016, Las Vegas, NV, USA, June 27-30, 2016. IEEE Computer Society, 770–778. https://doi.org/10.1109/CVPR.2016.90 [16] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. 2016. Identity mappings in deep residual networks. In European conference on computer vision. Springer, 630–645. [17] Dan Hendrycks and Kevin Gimpel. 2017. Early Methods for Detecting Adversarial Images. In 5th International Conference on Learning Representations, ICLR 2017, Toulon, France, April 24-26, 2017, Workshop Track Proceedings. OpenReview.net. [18] Anouar Kherchouche, Sid Ahmed Fezza, and Wassim Hamidouche. 2022. Detect and defense against adversarial examples in deep learning using natural scene statistics and adaptive denoising. Neural Comput. Appl. 34, 24 (2022), 21567–21582. https://doi.org/10.1007/s00521-021-06330-x [19] Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. 2020. Big Transfer (BiT): General Visual Representation Learning. 12350 (2020), 491–507. https://doi.org/10.1007/978-3-03058558-7_29 [20] Alex Krizhevsky. 2009. Learning Multiple Layers of Features from Tiny Images. Technical Report. University of Toronto. [21] Kuang Liu. 2017. pytorch-cifar: Training CIFAR10 with PyTorch. https://github. com/kuangliu/pytorch-cifar. Accessed: 2026-07-13. [22] Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. 2022. A convnet for the 2020s. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 11976–11986. [23] Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. 2018. Towards Deep Learning Models Resistant to Adversarial Attacks. In 6th International Conference on Learning Representations, ICLR 2018, Vancouver, BC, Canada, April 30 - May 3, 2018, Conference Track Proceedings. OpenReview.net. https://openreview.net/forum?id=rJzIBfZAb [24] Dhruv Mahajan, Ross B. Girshick, Vignesh Ramanathan, Kaiming He, Manohar Paluri, Yixuan Li, Ashwin Bharambe, and Laurens van der Maaten. 2018. Exploring the Limits of Weakly Supervised Pretraining. In Computer Vision - ECCV 2018 - 15th European Conference, Munich, Germany, September 8-14, 2018, Proceedings, Part II (Lecture Notes in Computer Science, Vol. 11206), Vittorio Ferrari, Martial Hebert, Cristian Sminchisescu, and Yair Weiss (Eds.). Springer, 185–201.

https://doi.org/10.1007/978-3-030-01216-8_12 [25] Tianyu Pang, Min Lin, Xiao Yang, Jun Zhu, and Shuicheng Yan. 2022. Robustness and Accuracy Could Be Reconcilable by (Proper) Definition. In International Conference on Machine Learning, ICML 2022, 17-23 July 2022, Baltimore, Maryland, USA (Proceedings of Machine Learning Research, Vol. 162), Kamalika Chaudhuri, Stefanie Jegelka, Le Song, Csaba Szepesvári, Gang Niu, and Sivan Sabato (Eds.). PMLR, 17258–17277. https://proceedings.mlr.press/v162/pang22a.html [26] Shengyun Peng, Weilin Xu, Cory Cornelius, Matthew Hull, Kevin Li, Rahul Duggal, Mansi Phute, Jason Martin, and Duen Horng Chau. 2023. Robust Principles: Architectural Design Principles for Adversarially Robust CNNs. In 34th British Machine Vision Conference 2023, BMVC 2023, Aberdeen, UK, November 20-24, 2023. BMVA Press, 739–740. http://proceedings.bmvc2023.org/739/ [27] Olga Russakovsky, Jia Deng, Hao Su, Jonathan Krause, Sanjeev Satheesh, Sean Ma, Zhiheng Huang, Andrej Karpathy, Aditya Khosla, Michael S. Bernstein, Alexander C. Berg, and Li Fei-Fei. 2015. ImageNet Large Scale Visual Recognition Challenge. Int. J. Comput. Vis. 115, 3 (2015), 211–252. https://doi.org/10.1007/ S11263-015-0816-Y [28] Vikash Sehwag, Saeed Mahloujifar, Tinashe Handina, Sihui Dai, Chong Xiang, Mung Chiang, and Prateek Mittal. 2022. Robust Learning Meets Generative Models: Can Proxy Distributions Improve Adversarial Robustness?. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 2529, 2022. OpenReview.net. https://openreview.net/forum?id=WVX0NNVBBkV [29] Vikash Sehwag, Saeed Mahloujifar, Tinashe Handina, Sihui Dai, Chong Xiang, Mung Chiang, and Prateek Mittal. 2022. Robust Learning Meets Generative Models: Can Proxy Distributions Improve Adversarial Robustness?. In The Tenth International Conference on Learning Representations, ICLR 2022, Virtual Event, April 2529, 2022. OpenReview.net. https://openreview.net/forum?id=WVX0NNVBBkV [30] Naman Deep Singh, Francesco Croce, and Matthias Hein. 2023. Revisiting Adversarial Training for ImageNet: Architectures, Training and Generalization across Threat Models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, Alice Oh, Tristan Naumann, Amir Globerson, Kate Saenko, Moritz Hardt, and Sergey Levine (Eds.). http://papers.nips.cc/paper_files/paper/2023/hash/ 2d3b007613940def7a5ec9d6d635937b-Abstract-Conference.html [31] Chen Sun, Abhinav Shrivastava, Saurabh Singh, and Abhinav Gupta. 2017. Revisiting Unreasonable Effectiveness of Data in Deep Learning Era. In IEEE International Conference on Computer Vision, ICCV 2017, Venice, Italy, October 22-29, 2017. IEEE Computer Society, 843–852. https://doi.org/10.1109/ICCV.2017.97 [32] Christian Szegedy, Wojciech Zaremba, Ilya Sutskever, Joan Bruna, Dumitru Erhan, Ian J. Goodfellow, and Rob Fergus. 2014. Intriguing properties of neural networks. (2014). http://arxiv.org/abs/1312.6199 [33] Kun Tong, Chengze Jiang, Jie Gui, and Yuan Cao. 2024. Taxonomy Driven Fast Adversarial Training. In Thirty-Eighth AAAI Conference on Artificial Intelligence, AAAI 2024, Thirty-Sixth Conference on Innovative Applications of Artificial Intelligence, IAAI 2024, Fourteenth Symposium on Educational Advances in Artificial Intelligence, EAAI 2024, February 20-27, 2024, Vancouver, Canada, Michael J. Wooldridge, Jennifer G. Dy, and Sriraam Natarajan (Eds.). AAAI Press, 5233–5242. https://doi.org/10.1609/AAAI.V38I6.28330 [34] Dimitris Tsipras, Shibani Santurkar, Logan Engstrom, Alexander Turner, and Aleksander Madry. 2019. Robustness May Be at Odds with Accuracy. (2019). https://openreview.net/forum?id=SyxAb30cY7 [35] Zekai Wang, Tianyu Pang, Chao Du, Min Lin, Weiwei Liu, and Shuicheng Yan. 2023. Better Diffusion Models Further Improve Adversarial Training. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA (Proceedings of Machine Learning Research, Vol. 202), Andreas Krause, Emma Brunskill, Kyunghyun Cho, Barbara Engelhardt, Sivan Sabato, and Jonathan Scarlett (Eds.). PMLR, 36246–36263. https://proceedings.mlr.press/ v202/wang23ad.html [36] weiaicunzai. 2019. pytorch-cifar100: Practice on cifar100 with various CNN architectures in PyTorch. https://github.com/weiaicunzai/pytorch-cifar100. Accessed: 2026-07-13. [37] Weilin Xu, David Evans, and Yanjun Qi. 2018. Feature Squeezing: Detecting Adversarial Examples in Deep Neural Networks. Proceedings 2018 Network and Distributed System Security Symposium (2018). https://doi.org/10.14722/ndss. 2018.23198 [38] Yuancheng Xu, Yanchao Sun, Micah Goldblum, Tom Goldstein, and Furong Huang. 2023. Exploring and Exploiting Decision Boundary Dynamics for Adversarial Robustness. In The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net. https://openreview.net/forum?id=aRTKuscKByJ [39] Jiahui Yu, Zirui Wang, Vijay Vasudevan, Legg Yeung, Mojtaba Seyedhosseini, and Yonghui Wu. 2022. CoCa: Contrastive Captioners are Image-Text Foundation Models. Trans. Mach. Learn. Res. 2022 (2022). https://openreview.net/forum?id= Ee277P3AYC [40] Sergey Zagoruyko and Nikos Komodakis. 2016. Wide Residual Networks. In Proceedings of the British Machine Vision Conference 2016, BMVC 2016, York, UK, September 19-22, 2016, Richard C. Wilson, Edwin R. Hancock, and William A. P.

VanillaBench: The Hidden Accuracy Cost of Adversarial Robustness

Smith (Eds.). BMVA Press. https://bmva-archive.org.uk/bmvc/2016/papers/ paper087/index.html [41] Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric P. Xing, Laurent El Ghaoui, and Michael I. Jordan. 2019. Theoretically Principled Trade-off between Robustness

and Accuracy. 97 (2019), 7472–7482. http://proceedings.mlr.press/v97/zhang19p. html

Record · ID 366198 · SHA-256 54e9cbb5893a7c4b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.