ConceptioArchivearXiv CS
arXiv CSopen access

Mitigating Error Amplification in Fast Adversarial Training

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

Mitigating Error Amplification in Fast Adversarial Training

arXiv:2604.24332v1 [cs.LG] 27 Apr 2026

Mengnan Zhao1 , Lihe Zhang2 *, Bo Wang2 , Tianhang Zheng3∗ , Hong Zhong1 , Geyong Min4 1 Anhui University, Anhui, China 2 Dalian University of Technology, Liaoning, China 3 Zhejiang University, Zhejiang, China 4 University of Exeter, UK

Abstract Fast Adversarial Training (FAT) has proven effective in enhancing model robustness by encouraging networks to learn perturbation-invariant representations. However, FAT often suffers from catastrophic overfitting (CO), where the model overfits to the training attack and fails to generalize to unseen ones. Moreover, robustness-oriented optimization typically leads to notable performance degradation on clean inputs, and such degradation becomes increasingly severe as the perturbation budget grows. In this work, we conduct a comprehensive analysis of how guidance strength affects model performance by modulating perturbation and supervision levels across distinct confidence groups. The findings reveal that low-confidence samples are the primary contributors to CO and the robustness–accuracy trade-off. Building on this insight, we propose a Distribution-aware Dynamic Guidance (DDG) strategy that dynamically adjusts both the perturbation budget and supervision signal. Specifically, DDG scales the perturbation magnitude according to the sample confidence at the ground-truth class, thereby guiding samples toward consistent decision boundaries while mitigating the influence of learning spurious correlations. Simultaneously, it dynamically adjusts the supervision signal based on the prediction state of each sample, preventing overemphasis on incorrect signals. To alleviate potential gradient instability arising from dynamic guidance, we further design a weighted regularization constraint. Extensive experiments on standard benchmarks demonstrate that DDG effectively alleviates both CO and the robustness–accuracy trade-off.

1. Introduction Improving the adversarial robustness of deep neural networks remains a fundamental challenge in trustworthy machine learning[34, 45, 47, 48, 50]. Among existing defense strategies, adversarial training (AT) has emerged as one of the most effective approaches for enhancing model robust* corresponding author.

ness against adversarial perturbations [4, 10, 37, 43, 49, 53]. However, standard AT typically incurs substantial computational overhead, as it requires iterative generation of adversarial examples. To mitigate this issue, Fast Adversarial Training (FAT) [16, 21, 25, 30, 31, 44] accelerates training by employing single-step attacks such as FGSM-RS [41]. While computationally efficient, FAT methods often suffer from catastrophic overfitting (CO), where model robustness collapses dramatically after only a few training epochs. Prior studies have attributed this phenomenon to factors such as gradient misalignment [1] and feature pathway divergence [55]. In addition, both FAT and standard AT [17, 39] encounter a robustness–accuracy tradeoff, where enhancing model robustness typically degrades clean accuracy. To address this, recent methods have introduced various techniques such as prior-guided adversarial initialization [20] and label relaxation [38]. However, the underlying mechanisms driving this trade-off, as well as its relationship to CO, remain insufficiently understood. In this work, we begin by conducting a systematic analysis of how varying guidance strengths affect FAT performance. Our analysis reveals that low-confidence and misclassified samples are the primary contributors to both CO and the robustness–accuracy trade-off. Specifically, injecting strong adversarial perturbations into these samples exacerbates CO, as the model tends to leverage perturbation-specific artifacts rather than intrinsic semantic cues—analogous to learning spurious, class-dependent backdoor features. Furthermore, we observe that mitigating erroneous guidance on low-confidence samples simultaneously improves both clean accuracy and adversarial robustness. These insights suggest that enforcing uniform guidance across all samples is inherently suboptimal; instead, the guidance should be dynamically modulated according to the confidence and prediction state of each input. Motivated by these insights, we propose a Distributionaware Dynamic Guidance (DDG) strategy. DDG dynamically adjusts perturbation strengths via adaptive budget allocation, which guides samples to reside near similar decision boundaries while mitigating the risk of learning spurious correlations. In addition, DDG modulates supervision

signals according to each sample’s prediction state, preventing excessive emphasis on incorrectly predicted classes. To further stabilize optimization, we introduce a weighted regularization term that alleviates potential gradient nonsmoothness induced by dynamic guidance. Our main contributions are summarized as follows: (1) We conduct a systematic examination revealing that applying uniform guidance across samples is suboptimal, and that the guidance should adapt to prediction confidence and correctness. (2) Building on this insight, we propose a distribution-aware dynamic guidance (DDG) strategy that dynamically adjusts the perturbation budget and supervision signal according to each sample’s confidence and prediction state. Furthermore, we introduce a weighted regularization term to stabilize gradients under dynamic guidance. (3) Extensive experiments on standard benchmarks demonstrate that DDG achieves superior robustness and clean accuracy compared to state-of-the-art FAT methods.

2. Related Work Deep neural networks have raised security concerns due to their susceptibility to adversarial attacks [2, 9, 14, 23, 46, 56], prompting growing interest in AT techniques [17, 28, 40]. Given the training data (x, y) ∼ Dtrain , Madry et al. [26] formulate AT as a min-max optimization problem, \label {eq1} \min _\theta \mathbb {E}_{(x, y)\sim D_\text {train}} \left [\max _{\delta \in [-\xi , \xi ]} \mathcal {L}\left (f(x+\delta ), y\right ) \right ],

(1)

where f (·) denotes the model parameterized by θ, ξ is the maximum perturbation budget, and L(·) typically represents the cross-entropy loss. In contrast, regularizationbased methods [36, 51] align the predictions for clean x and adversarial examples x + δ, \label {eq2} \min _\theta \mathbb {E}_{(x, y)\sim D_\text {train}}\left [ \mathcal {L}\left (f(x), y\right ) + \|f(x+\delta ) - f(x)\|_2 \right ], (2) where ∥ · ∥2 denotes the ℓ2 norm function. Unlike standard AT that utilizes multi-step attacks, FAT employs single-step attacks (e.g., FGSM) [5, 13, 33] for improving training efficiency. However, FAT may face the CO issue [7, 32]. To mitigate CO, prior works have explored techniques including random initialization [41], gradient alignment [1], regularization-based defenses [29, 36], smoothed convergence [54], and feature activation consistency [55]. Beyond CO, both FAT and standard AT exhibit a significant trade-off between adversarial robustness and clean classification accuracy [18]. To mitigate this issue, researchers have introduced several strategies, such as priorguided adversarial initialization [20], adaptive step-size allocation based on gradient norms [16], feature-space regularization [19], and label relaxation [38]. Despite recent advances, existing methods still suffer from a significant trade-off between clean accuracy and

Figure 1. CO analysis on CIFAR10 and ResNet18. Each training batch is divided into four groups, Max, Min, Mid1 , and Mid2 . PGD use 10 steps, with a fixed perturbation budget of 8/255.

adversarial robustness, while the intrinsic relationship between CO and this trade-off remains unclear. This work reveals that low-confidence and misclassified samples are the primary drivers to both CO and the robustness–accuracy trade-off. Hence, we propose a distribution-aware dynamic guidance that dynamically assigns perturbation budgets and supervision signals per sample, replacing the uniform treatment used in prior works.

3. Proposed Method In this section, we begin by analyzing the impact of guidance intensity on FAT performance. Then, we describe the proposed distribution-aware dynamic guidance strategy.

3.1. Impact of Guidance Intensity This section aims to answer the following question: Do different samples exhibit distinct behaviors during FAT, and if so, how? To investigate this, we conduct a confidence-based analysis following the TDAT framework [38] on CIFAR-10 and ResNet-18. All experiments are performed with a training batch size of 128. 3.1.1. Perturbation Strength Ablation CO Analysis. To investigate how perturbation strength affects CO, batch samples are divided into four groups based on prediction confidence. We selectively increase the perturbation budget to 16/255 for the target group while maintaining 8/255 for the remaining groups. The learning rate is fixed at 0.1. Figure 1 presents two key findings: 1) Highconfidence samples can tolerate large perturbations without destabilizing optimization; 2) Low-confidence samples—predominantly misclassified—are more susceptible to CO when exposed to large perturbations. Trade-off Analysis. To further investigate the robustness–accuracy trade-off under varying perturbation

(a) Setting the default ξ to 8/255 and decreasing ξ to 4/255 for the selected group (confidence descending order)

(b) Setting the default ξ to 8/255 and increasing ξ to 12/255 for the selected group (confidence descending order)

Figure 2. Trade-off analysis on CIFAR-10 with ResNet-18. PGD and C&W use 10 steps, with a fixed perturbation budget of 8/255.

strengths, each training batch is partitioned into 32 confidence groups—a finer granularity than that used in the CO analysis—to capture nuanced performance variations across confidence levels. For the selected group, we adjust the perturbation budget—either decreasing it to 4/255 or increasing it to 12/255—while keeping other groups fixed at 8/255. Figure 2 reveals four key observations: 1) Reducing the perturbation budget to 4/255 consistently improves clean accuracy and moderately enhances robustness against C&W; 2) Reducing the perturbation budget for the lowestconfidence group enhances both clean and robust accuracies, achieving gains of +1.06, +0.47, and +0.41 on Clean, PGD, and C&W metrics, respectively. This suggests that alleviating excessive error reinforcement on already misclassified samples can enhance overall performance; 3) Consistent with the CO analysis, applying large perturbations to low-confidence samples can trigger CO, indicating that even a small fraction of unstable samples may destabilize training; 4) Increasing the perturbation budget for highconfidence groups slightly improves model robustness, albeit sometimes at the expense of clean accuracy. 3.1.2. Supervision Strength Ablation Recent studies have introduced label relaxation to improve adversarial training stability, typically formulated as \label {eq3} \mathbf {\hat {y}} = \mathbf {y} \cdot \gamma + \mathbf {1}\cdot \frac {1 - \gamma }{L},

(3)

where ŷ denotes the relaxation label, y is the one-hot ground-truth label, L signifies the number of classes, and γ controls the relaxation strength. Given that datasets vary in class cardinality, our analysis primarily focuses on the true class and the most probable incorrect class. Specifically, each training batch is divided into 8 confidence groups, and for the selected group, its supervision label is modified as \label {eq41} \mathbf {{y}^\prime } = \mathbf {\hat {y}} + \mathbf {y} \cdot \beta _1 + \mathbf {y_m} \cdot \beta _2,

(4)

where ym represents the one-hot vector of the most probable incorrect class, and β1 , β2 modulate the emphasis on the true and most probable incorrect classes, respectively. Figure 3 reveals several key observations: (1) Applying β1 = 0 and β2 = −0.1 to the lowest-confidence group notably enhances both clean and adversarial accuracies, yielding average absolute gains of +1.18 on clean accuracy and +1.40 on PGD robustness, while maintaining comparable C&W robustness to the baseline. This finding aligns with the perturbation-strength ablation, confirming that mitigating excessive error reinforcement for already misclassified samples can improve model performance; (2) When the same setting is applied to mid-confidence samples, training becomes unstable—likely because the membership of this confidence group changes rapidly during training, leading to inconsistent supervision signals and oscillatory gradient directions across iterations; (3) Increasing β1 improves the

(a) Setting β1 to 0 and β2 to -0.1 (confidence descending order)

(b) Setting β1 to 0.1 and β2 to -0.1 (confidence descending order)

Figure 3. Supervision strength analysis on CIFAR-10 with ResNet-18. PGD and C&W use 10 steps, with a perturbation budget of 8/255.

training stability, implying that reinforcing correct class supervision stabilizes FAT; (4) Overly large β1 degrades performance, as it induces overconfident predictions and amplifies regularization penalties. Remarks. The above analyses indicate that applying uniform guidance to all samples is inherently suboptimal. Instead, both perturbation and supervision strengths should be adaptively adjusted based on each sample’s characteristics, guiding inputs with varying confidence levels toward more consistent decision boundaries. In particular, mitigating excessive error reinforcement for low-confidence (mostly misclassified) samples yields the most substantial improvements in both robustness and clean accuracy.

3.2. Distribution-aware Dynamic Guidance DDG jointly adjusts the perturbation budget and supervisory guidance, enabling adaptive alignment of decision boundaries while mitigating the over-reinforcement of erroneous signals. Unlike prior methods that utilize a fixed perturbation budget (ξbase = 8/255), DDG allocates budgets based on the confidence ranking ri of each example xi : \label {eq4} \xi _i = \xi _{\text {base}} + \kappa \left [ \tanh \left (r_i - \tau _1\right ) - \tanh \left (\tau _2 - r_i\right ) \right ],

(5)

where κ serves as a scaling factor that controls the amplitude of dynamic variation (defaulting to 2/255, which yields ξi ∈ [4/255, 12/255]). τ1 and τ2 govern transition regions of perturbation intensity, constrained by τ1 + τ2 =

B, where B represents the batch size. ri indicates the sample ranking, sorted in descending order of prediction confidence. Eq. (5) assigns larger budgets to high-confidence samples and smaller budgets to low-confidence ones. The adversarial perturbations are then clipped using ξi as \label {eq5} \delta _i = \operatorname {clip}\Big ( \delta _\text {init} + \max \{\xi _{i}, \xi _\text {base}\}\cdot \operatorname {sign}(\nabla _{x_i+\delta _\text {init}} \mathcal {L}),\ -\xi _{i},\ \xi _{i} \Big ). (6) Here, \delta _\text {init} denotes the initial perturbation, typically inherited from the previous optimization step. The term \nabla _{x_i+\delta _\text {init}} \mathcal {L} is the gradient of the loss function with respect to the input x_i+\delta _\text {init} . The operation \qopname \relax m{max}\{\xi _{i}, \xi _\text {base}\} establishes a lower bound for the perturbation step size. The final adversarial example is formulated as x′i = clip(xi + δi , 0, 1). In addition, DDG redefines the soft supervision signal to maintain training stability and mitigate error reinforcement,

\label {eq6} \mathbf {y_{\text {sr}}} = \begin {cases} \mathbf {\hat {y}}, & \text {if } \arg \max f(x^\prime ) = y, \\[6pt] \mathbf {\hat {y}} + \gamma (1 - \text {Acc}) \mathbf {y}\, - \dfrac {1}{L} \mathbf {y_m}, & \text {otherwise,} \end {cases} (7) where \protect \text {Acc} denotes the empirical accuracy of the current model on the training batch, providing a dynamic, global performance signal. Term γ(1 − Acc)y performs positiveclass reinforcement, decreasing with higher true-class con1 fidence, while term ym implements negative-class supL

pression, inversely proportional to the number of classes. With the established adversarial inputs x′ and their corresponding supervision labels ysr , the optimization objective is defined as follows:

Algorithm 1: The proposed DDG

Input: Training dataset D, number of classes L, batch size B, model f with parameters θ, base perturbation budget ξbase = 8/255, scaling factor κ, thresholds τ1 , τ2 , relaxation factor γ \label {eq7} \mathcal {L}_\text {total} = - \frac {1}{B} \sum \mathbf {y_{\text {sr}}} \log f(x^\prime ) + \mathcal {L}_{\text {smo}}, (8) Output: Trained model weights θ 1 Initialize model parameters θ; 2 for epoch = 1 to N do 3 foreach batch (X, y) ∼ D do \label {eq8} \footnotesize \mathcal {L}_{\text {smo}} = \big \| f(x + \delta _\text {init}) - f(x') \big \|_2 \left ( \lambda \frac {\max (\xi _{B}) - \xi _{B}}{\max (\xi _{B}) - \min (\xi _{B})} + \alpha \,y_{\text {false}} + 1 \right ). 4 Calculate confidence rankings of X as r ; (9) // Assign Perturbation Budget: where ∥·∥2 indicates the ℓ2 -norm distance metric, max(ξB ) 5 ξ = ξbase + κ [tanh (r − τ1 ) − tanh (τ2 − r)]; // Adv Example Generation: and min(ξB ) correspond to the maximum and minimum 6 g = ∇x+δinit L(f (x + δinit ), y); perturbation budgets within the current batch respectively, 7 δ = clip (δinit + max{ξ, ξbase } · sign(g), −ξ, ξ); ξB refers to the dynamic perturbation budget for batch sam8 x′ = clip(x + δ, 0, 1); ples, and yfalse serves as the misclassification indicator that // Assign Signal: P Supervision returns 1 when arg max f (x′ ) ̸= y and 0 otherwise. The 9 Acc = B1 I[arg max f (x′ ) = y]; smoothing regularization term \protect \mathcal {L}_{\text {smo}} comprises three com10 ŷ = y · γ + 1 · 1−γ ; L plementary components: the term\lambda \frac {\max (\xi _{B}) - \xi _{B}}{\max (\xi _{B}) - \min (\xi _{B})} bal11 y sr = ances the regularization strength across the batch; the misŷ, if Pred True, classification penalty α yfalse reinforces regularization on 1 ŷ + γ(1 − Acc)y − ym , otherwise, misclassified examples, maintaining gradient smoothness L throughout the negative-suppression process; and a constant // Loss with Gradient Smoothness: 12 Compute regularization Lsmo using Eq. (9); term restricts non-zero regularization for all samples. P 13 Ltotal = − B1 ysr log f (x′ ) +Lsmo ; The Algorithm details of DDG are given in Algorithm 1. 14 Update θ ← θ − η∇θ Ltotal ; 15 end 4. Experiments 16 end 4.1. Experimental Settings 17 return θ

General Setup. We evaluate our method on benchmark datasets: CIFAR-10 [22], CIFAR-100 [22], and Tiny ImageNet [8]. To ensure fair comparison, all methods are trained under identical experimental settings following prior works [41, 42, 52]. Specifically, we adopt ResNet-18 [15] as the backbone for all datasets and train models using stochastic gradient descent (SGD) with a momentum of 0.9, weight decay of 5 × 10−4 , batch size of 128, and the initial learning rate of 0.1. Training is performed for 110 epochs, with the learning rate decayed by a factor of 0.1 at epochs 100 and 105. For each method, we report results at both the best and final epochs, where ‘best’ denotes the epoch achieving the highest robustness against PGD-10, and ‘final’ represents the last epoch to evaluate training stability. Hyperparameters. The threshold \tau _1 in Eq. (5) is set to 8. \lambda and \alpha in Eq. (8) are set to 1.33 and 1.5, respectively. Other hyperparameters follow the settings in prior work [38]. Evaluation Protocol. Following the existing evaluation protocols [16, 38, 41], we utilize various attacks to evaluate adversarial robustness, including FGSM [13], BIM [24], PGD-10/20/50 [26], C&W [3], APGD [6], and AutoAttack (AA) [6]. Here, PGD-n refers to a projected gradient descent attack with n iterations. AutoAttack combines APGD, FAB [11], and Square Attack [27]. All attacks are performed under the ℓ∞ norm with a perturbation budget of

ξ = 8/255. For single-step attacks (e.g. FGSM), the step size is 8/255; for multi-step attacks (e.g. PGD), it is 2/255. Baselines. We compare our work with recent FAT methods, including FGSM-RS [41], FGSM-Free [33], FGSMSDI [18], GAT [35], GradAlign [1], N-FGSM [7], ZeroGrad [12], LAS-AWP [17], NuAT [36], ATAS [16], FGSMPGI [42], FGSM-PGK [20], and TDAT [38].

4.2. Comparative Experiments and Analysis Results on CIFAR10. Table 1 shows the experimental results on CIFAR10, with the observations given as follows: 1) DDG achieves the highest robustness across nearly all attack settings, reaching 68.44%/59.48%/59.33% under FGSM/APGD/PGD-50 attacks—outperforming prior FAT methods such as TDAT (66.28%/55.10%/55.41%). 2) DDG maintains consistent robustness between the best and final checkpoints (e.g., 60.44% vs. 59.70% on PGD-10), indicating effective mitigation of robust overfitting. 3) Despite using only single-step adversarial updates, DDG surpasses multi-step works such as MART (54.83% on PGD10) and LAS-AWP (56.37% on PGD-10), while requiring only about one-tenth of their computational cost. 4) DDG shows a better trade-off between clean and robust accuracy.

Table 1. Performance comparison of various AT methods on CIFAR-10. Bold numbers highlight the best results. ‘Steps’: attack iterations during training. ‘Best’: checkpoint with highest PGD-10 accuracy. ‘Final’: last checkpoint. Model

Steps Type Clean FGSM [13] BIM [24]

MART [39]

10

LAS-AWP [17]

10

FGSM-RS [41]

1

FGSM-Free [33]

-

GAT [35]

1

FGSM-SDI [18]

1

GradAlign [1]

1

N-FGSM [7]

1

FGSM-PGK [20]

1

FGSM-PGI [42]

1

TDAT [38]

1

Ours

1

10

PGD [26] 20 50

AA [6] C&W [3] APGD [6]

Best 82.03 Final 82.33 Best 82.92 Final 82.92

64.94 65.12 65.86 65.86

54.48 53.90 55.97 55.97

54.83 53.72 53.53 54.38 52.98 52.60 56.37 55.57 55.20 56.37 55.57 55.20

47.74 47.46 49.46 49.46

49.68 49.66 51.53 51.53

53.51 52.77 54.18 54.18

Best 83.69 Final 83.69 Best 81.38 Final 81.38 Best 81.53 Final 81.88 Best 83.55 Final 83.73 Best 80.45 Final 80.45 Best 80.35 Final 80.35 Best 81.52 Final 81.63 Best 81.71 Final 81.71 Best 82.46 Final 82.72 Best 82.67 Final 82.84

62.00 62.00 60.81 60.81 64.18 64.30 63.60 63.75 60.56 60.56 60.93 60.93 64.95 65.01 65.02 65.02 66.28 66.29 68.44 68.27

47.20 47.20 48.74 48.74 53.72 52.89 51.46 51.28 48.80 48.80 49.59 49.59 55.28 55.02 54.87 54.87 55.87 55.62 59.79 59.12

47.66 46.29 45.96 47.66 46.29 45.96 49.07 48.03 47.62 49.07 48.03 47.62 54.05 53.26 52.95 53.23 52.16 51.86 51.94 50.65 50.34 51.88 50.49 50.09 49.11 47.96 47.63 49.11 47.96 47.63 49.83 48.77 48.51 49.83 48.77 48.51 56.14 55.58 55.36 55.79 55.34 55.09 55.26 54.54 54.38 55.26 54.54 54.38 56.36 55.63 55.41 56.27 55.53 55.26 60.44 59.57 59.33 59.70 59.07 58.83

42.80 42.80 44.37 44.37 47.68 47.08 46.31 46.34 43.92 43.92 44.54 44.54 48.94 48.85 48.60 48.60 48.06 47.26 48.08 48.06

46.10 46.10 46.98 46.98 49.76 49.71 49.09 49.42 46.94 46.94 47.37 47.37 50.90 50.75 50.88 50.88 49.99 49.74 49.86 50.01

46.13 46.13 47.90 47.90 53.24 52.05 50.61 50.43 47.85 47.85 48.59 48.59 55.44 55.28 54.62 54.62 55.10 54.98 59.48 58.23

For instance, while FGSM-RS yields slightly higher clean accuracy (83.69%), it exhibits notably lower robustness (47.66% on PGD-10). Similarly, while FGSM-PGI and FGSM-PGK attain marginally higher robustness on C&W and AA, they perform notably worse on other metrics. Results on CIFAR-100. Table 2 reports the performance comparison of various FAT methods on CIFAR100. Overall, our DDG achieves the strongest results among all competitors. Specifically, compared to recent advanced approaches TDAT (40.29%/33.56%/33.15%) and FGSM-PGK (39.61%/32.78%/31.96%), our method attains the highest robustness of 40.96%/34.32%/33.86% against FGSM, PGD-10, and APGD attacks, respectively. Moreover, DDG achieves a clean accuracy of 57.98%, surpassing TDAT (57.31%) and FGSM-PGK (57.36%), indicating that the robustness improvements do not compromise clean performance. These results validate the effectiveness of our work in balancing robustness–accuracy trade-off. Results on Tiny-ImageNet. Table 5 summarizes the performance of various FAT methods on Tiny-ImageNet. Our proposed approach achieves the optimal robustness across diverse adversarial attacks while maintaining competitive clean accuracy. Specifically, it attains the highest adversarial accuracy of 24.30%/24.35%/23.85% under BIM,

PGD-10, and APGD attacks, respectively, performing better than the strong baselines TDAT (23.78%/23.98%/23.28%) and FGSM-PGI (23.19%/23.27%/22.91%). In addition, our final model exhibits stable training behavior, yielding 23.40%/23.64%/22.95% robustness under BIM, PGD-10, along with a clean accuracy of 44.83%, which is comparable to or surpasses competing methods.

4.3. Ablation Studies Effects of each component. Table 3 presents the ablation results for the key components of DDG, including perturbation budget allocation (PBA), supervision signal adjustment (SSA), and gradient smoothness (GS). Removing PBA consistently degrades both clean and adversarial accuracies. Removing SSA improves robustness under the C&W attack, but incurs a slight reduction in clean accuracy. In contrast, removing GS yields the highest clean accuracy (84.55% on average), yet noticeably reduces adversarial performance. Overall, the complete DDG achieves the most favorable robustness–accuracy trade-off. Effects of components in ysr . In Eq. (7), ysr contains both positive enhancement and negative suppression terms. To study their individual effects, we apply separate scaling factors (0 to 1.4, step 0.2), multiplying \gamma for the posi-

Table 2. Performance comparison of various AT methods on CIFAR-100. Bold numbers highlight the best results. Methods

Steps Type Clean FGSM [13] BIM [24]

MART [39]

10

LAS-AWP [17]

10

FGSM-RS [41]

1

FGSM-Free [33]

-

GAT [35]

1

FGSM-SDI [18]

1

GradAlign [1]

1

N-FGSM [7]

1

FGSM-PGK [20]

1

FGSM-PGI [42]

1

TDAT [38]

1

Ours

1

82.67 82.84 81.70 81.56 80.38 80.69 84.32 84.55

68.44 68.27 68.33 68.23 66.11 66.17 68.44 68.34

32.00 31.75 32.42 32.42

32.18 31.68 31.59 31.85 31.37 31.21 32.58 31.91 31.74 32.58 31.91 31.74

26.07 25.71 27.23 27.23

28.01 27.81 29.59 29.59

31.55 31.22 31.74 31.74

Best 51.67 Final 51.67 Best 52.06 Final 52.06 Best 57.49 Final 57.58 Best 58.64 Final 58.54 Best 54.90 Final 55.22 Best 54.41 Final 54.41 Best 57.36 Final 57.36 Best 58.78 Final 58.82 Best 57.32 Final 57.32 Best 57.98 Final 58.15

31.02 31.02 32.13 32.13 36.77 36.85 37.23 37.19 35.28 35.51 35.00 35.00 39.61 39.61 40.02 39.83 40.29 40.29 40.96 41.17

22.42 22.42 24.48 24.48 28.91 28.87 28.60 28.53 26.77 26.82 26.99 26.99 32.12 32.12 31.43 31.22 33.33 33.33 34.11 33.70

22.61 22.04 21.75 22.61 22.04 21.75 24.74 24.09 24.04 24.74 24.09 24.04 29.14 28.60 28.30 29.06 28.43 28.30 28.78 27.99 27.67 28.71 28.00 27.72 27.13 26.52 26.22 27.12 26.42 26.24 27.01 26.55 26.34 27.01 26.55 26.34 32.78 32.35 32.19 32.78 32.35 32.19 31.94 31.30 31.19 31.65 31.18 30.89 33.56 33.17 33.06 33.56 33.17 33.06 34.42 33.92 33.79 33.99 33.63 33.51

18.72 18.72 20.23 20.23 23.11 23.02 23.27 23.18 22.30 22.19 22.81 22.81 25.84 25.84 25.65 25.43 26.61 26.61 26.54 26.79

20.92 20.92 22.43 22.43 25.14 24.97 25.85 25.55 25.01 24.94 25.08 25.08 28.40 28.40 28.23 27.75 28.47 28.47 28.42 28.46

21.87 21.87 23.99 23.99 28.42 28.43 27.83 27.89 26.39 26.52 26.31 26.31 31.96 31.96 31.21 30.93 33.15 33.15 33.86 33.48

60.44 59.70 59.54 59.79 58.21 58.11 58.71 58.35

49.86 50.01 49.58 49.54 50.87 50.63 49.08 48.82

59.48 58.23 57.46 57.79 57.03 56.85 57.11 56.81

Table 4. Ablation analysis of components in ysr . Scaling

Positive Adjustment Negtive Adjustment Clean PGD10 C&W10 Clean PGD10 C&W10

0 0.2 0.4 0.6 0.8 1.0 1.2 1.4

82.03 82.37 82.60 82.57 82.90 83.04 83.54 83.38

59.85 59.78 60.06 59.96 60.06 60.18 60.16 60.14

50.47 50.10 50.71 49.83 50.59 50.64 50.29 49.96

AA [6] C&W [3] APGD [6]

38.62 38.52 40.66 40.66

Methods Type Clean FGSM PGD10 C&W20 APGD Best Avg Best wo. PBA Avg Best wo. SSA Avg Best wo. GS Avg

PGD [26] 20 50

Best 54.51 Final 54.75 Best 58.75 Final 58.75

Table 3. Ablation analysis of components in DDG. PBA, SSA, and GS correspond to perturbation budget allocation, supervision signal adjustment, and gradient smoothness, respectively.

DDG

10

81.50 81.39 82.03 82.28 82.39 82.94 83.03 82.98

58.77 59.02 59.36 59.44 59.90 60.10 60.72 60.65

51.02 51.16 50.71 50.65 50.48 50.75 49.89 50.02

tive enhancement term and \protect \frac {1}{L} for the negative suppression term. As reported in Table 4, amplifying the negative suppression term consistently boosts clean accuracy and PGD robustness, while inducing only a slight drop in C&W robustness. For instance, increasing the scaling factor from 0 to 1.0 yields gains of +1.44 in clean accuracy and +1.33 in PGD robustness, with only a 0.27 reduction under the C&W attack. Increasing the positive enhancement term exhibits a similar trend: clean accuracy improves steadily, whereas PGD and C&W robustness change only marginally. Notably, clean accuracy is positively correlated with PGD robustness but negatively correlated with C&W robustness. This discrepancy stems from the distinct attack objectives: PGD can be interpreted as a one-to-any attack that pushes samples away from the true class toward any incorrect class, whereas C&W behaves as an any-to-one attack that identifies the easiest path. As clean accuracy increases, the model forms sharper class boundaries, making it more difficult to push a sample toward arbitrary incorrect classes—thereby improving PGD robustness. However, such boundaries may also reveal clearer descent directions toward specific target classes, leading to a slight reduction in C&W robustness. Impact of τ1 and λ. Figure 4 illustrates the effects of τ1 in Eq. (5) and λ in Eq. (8) on FAT performance. Un-

Table 5. Performance comparison of various AT methods on Tiny-ImageNet. Bold numbers highlight the best results. Methods

Steps Type Clean FGSM [13] BIM [24]

MART [39]

10

LAS-AWP [17]

10

FGSM-RS [41]

1

FGSM-Free [33]

-

GAT [35]

1

FGSM-SDI [18]

1

GradAlign [1]

1

N-FGSM [7]

1

FGSM-PGI [42]

1

TDAT [38]

1

Ours

1

10

PGD [26] 20 50

AA [6] C&W [3] APGD [6]

Best 38.41 Final 36.83 Best 47.86 Final 47.86

25.20 18.02 30.77 30.77

20.91 12.25 23.98 23.98

20.93 20.74 20.67 15.53 12.36 12.01 11.93 9.26 24.10 23.67 23.60 18.21 24.10 23.67 23.60 18.21

16.88 10.36 20.49 20.49

20.77 12.02 23.51 23.51

Best 43.52 Final 43.52 Best 44.15 Final 44.15 Best 46.00 Final 45.57 Best 43.71 Final 45.40 Best 38.22 Final 37.89 Best 46.06 Final 46.06 Best 42.98 Final 45.13 Best 42.51 Final 43.92 Best 43.44 Final 44.83

23.93 23.93 25.18 25.18 23.04 22.10 26.84 24.76 22.85 22.51 24.72 24.72 28.55 28.11 29.44 29.31 29.34 29.74

17.12 17.12 17.81 17.81 14.96 14.38 20.48 17.15 17.13 16.95 16.67 16.67 23.19 21.43 23.78 23.06 24.30 23.40

17.22 16.82 16.64 17.22 16.82 16.64 17.95 17.47 17.31 17.95 17.47 17.31 15.16 14.51 14.33 14.56 14.03 13.85 20.60 20.26 20.11 17.27 16.84 16.73 17.20 16.86 16.79 17.06 16.78 16.69 16.74 16.21 16.02 16.74 16.21 16.02 23.27 23.01 22.92 21.51 21.19 21.07 23.98 23.63 23.56 23.29 22.85 22.74 24.35 24.09 24.06 23.64 23.24 23.07

14.67 14.67 15.82 15.82 13.27 12.71 17.16 14.80 14.00 13.90 14.96 14.96 18.67 16.86 18.21 17.54 18.80 17.85

16.72 16.72 17.32 17.32 14.38 13.90 20.13 16.97 16.85 16.68 16.22 16.22 22.91 21.14 23.28 22.50 23.85 22.95

13.09 13.09 13.67 13.67 10.82 10.26 15.43 12.47 12.64 12.49 12.71 12.71 17.00 14.86 16.64 15.87 16.85 15.94

Figure 4. The impact of τ1 in Eq. (5) and λ in Eq. (8) on FAT performance.

der stable training, clean accuracy increases as τ1 grows, while robust accuracy gradually decreases. Conversely, as λ increases, the clean accuracy declines whereas the robust accuracy improves. Notably, the influence of these hyperparameters on overall performance is modest: clean, PGD, and C&W accuracies vary within 83% ± 0.5, 60% ± 0.5, and 50.25% ± 0.25, respectively.

5. Conclusions This paper introduces a Distribution-aware Dynamic Guidance (DDG) strategy to address two key challenges in fast adversarial training: catastrophic overfitting and the robustness–accuracy trade-off. Through systematic perturbation

and supervision ablations, we show that applying uniform guidance across samples is suboptimal, leading to training instability and the learning of spurious correlations. To overcome these issues, DDG dynamically adjusts perturbation budgets according to sample confidence and modulates supervision signals based on each sample’s prediction state. Furthermore, an adaptive regularization term is introduced to maintain gradient smoothness under dynamic guidance. Experiments on multiple benchmarks demonstrate that DDG effectively resolves catastrophic overfitting and mitigates the robustness–accuracy trade-off.

Acknowledgements This work was supported by the National Natural Science Foundation of China under Grants 62431004, 62276046, and 62572426.

References [1] Flammarion N Andriushchenko M. Understanding and improving fast adversarial training. In NIPS, pages 16048– 16059, 2020. 1, 2, 5, 6, 7, 8 [2] Yulong Cao, Chaowei Xiao, Anima Anandkumar, Danfei Xu, and Marco Pavone. Advdo: Realistic adversarial attacks for trajectory prediction. In ECCV, pages 36–52. Springer, 2022. 2 [3] Nicholas Carlini and David Wagner. Towards evaluating the robustness of neural networks. In S&P, pages 39–57, 2017. 5, 6, 7, 8 [4] Stephen Casper, Lennart Schulze, Oam Patel, and Dylan Hadfield-Menell. Defending against unforeseen failure modes with latent adversarial training. arXiv preprint arXiv:2403.05030, 2024. 1 [5] Yaya Cheng, Jingkuan Song, Xiaosu Zhu, Qilong Zhang, Lianli Gao, and Heng Tao Shen. Fast gradient non-sign methods. arXiv preprint arXiv:2110.12734, 2021. 2 [6] Francesco Croce and Matthias Hein. Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks. In ICML, 2020. 5, 6, 7, 8 [7] Pau de Jorge Aranda, Adel Bibi, Riccardo Volpi, Amartya Sanyal, Philip Torr, Grégory Rogez, and Puneet Dokania. Make some noise: Reliable and efficient single-step adversarial training. NIPS, 35:12881–12893, 2022. 2, 5, 6, 7, 8 [8] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. Imagenet: A large-scale hierarchical image database. In CVPR, pages 248–255. Ieee, 2009. 5 [9] Yinpeng Dong, Fangzhou Liao, Tianyu Pang, Hang Su, Jun Zhu, Xiaolin Hu, and Jianguo Li. Boosting adversarial attacks with momentum. In CVPR, pages 9185–9193, 2018. 2 [10] Feiteng Fang, Yuelin Bai, Shiwen Ni, Min Yang, Xiaojun Chen, and Ruifeng Xu. Enhancing noise robustness of retrieval-augmented language models with adaptive adversarial training. arXiv preprint arXiv:2405.20978, 2024. 1 [11] Croce Francesco and Hein Matthias. Minimally distorted adversarial examples with a fast adaptive boundary attack. In ICML, pages 2196–2205, 2020. 5 [12] Z. Golgooni, M. Saberi, M. Eskandar, and M. H Rohban. Zerograd: Mitigating and explaining catastrophic overfitting in fgsm adversarial training. arXiv preprint arXiv:2103.15476, 2021. 5 [13] Ian J Goodfellow, Jonathon Shlens, and Christian Szegedy. Explaining and harnessing adversarial examples. In ICLR, 2015. 2, 5, 6, 7, 8 [14] Jindong Gu, Hengshuang Zhao, Volker Tresp, and Philip HS Torr. Segpgd: An effective and efficient adversarial attack for evaluating and boosting segmentation robustness. In ECCV, pages 308–325. Springer, 2022. 2

[15] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. Deep residual learning for image recognition. In CVPR, pages 770–778, 2016. 5 [16] Zhichao Huang, Yanbo Fan, Chen Liu, Weizhong Zhang, Yong Zhang, Mathieu Salzmann, Sabine Süsstrunk, and Jue Wang. Fast adversarial training with adaptive step size. TIP, 32:6102–6114, 2023. 1, 2, 5 [17] Xiaojun Jia, Yong Zhang, Baoyuan Wu, Ke Ma, Jue Wang, and Xiaochun Cao. Las-at: adversarial training with learnable attack strategy. In CVPR, pages 13398–13408, 2022. 1, 2, 5, 6, 7, 8 [18] Xiaojun Jia, Yong Zhang, Baoyuan Wu, Jue Wang, and Xiaochun Cao. Boosting fast adversarial training with learnable adversarial initialization. TIP, 31:4417–4430, 2022. 2, 5, 6, 7, 8 [19] Xiaojun Jia, Yuefeng Chen, Xiaofeng Mao, Ranjie Duan, Jindong Gu, Rong Zhang, Hui Xue, Yang Liu, and Xiaochun Cao. Revisiting and exploring efficient fast adversarial training via law: Lipschitz regularization and auto weight averaging. TIFS, 2024. 2 [20] Xiaojun Jia, Yong Zhang, Xingxing Wei, Baoyuan Wu, Ke Ma, Jue Wang, and Xiaochun Cao. Improving fast adversarial training with prior-guided knowledge. PAMI, 2024. 1, 2, 5, 6, 7 [21] Hoki Kim, Woojin Lee, and Jaewook Lee. Understanding catastrophic overfitting in single-step adversarial training. In AAAI, pages 8119–8127, 2021. 1 [22] Alex Krizhevsky, Geoffrey Hinton, et al. Learning multiple layers of features from tiny images. 2009. 5 [23] A. Kurakin, I.J. Goodfellow, and S. Bengio. Adversarial machine learning at scale. In ICLR, 2017. 2 [24] Alexey Kurakin, Ian J Goodfellow, and Samy Bengio. Adversarial examples in the physical world. In Artificial intelligence safety and security, pages 99–112. Chapman and Hall/CRC, 2018. 5, 6, 7, 8 [25] Tao Li, Yingwen Wu, Sizhe Chen, Kun Fang, and Xiaolin Huang. Subspace adversarial training. In CVPR, pages 13409–13418, 2022. 1 [26] Aleksander Madry, Aleksandar Makelov, Ludwig Schmidt, Dimitris Tsipras, and Adrian Vladu. Towards deep learning models resistant to adversarial attacks. In ICLR, 2018. 2, 5, 6, 7, 8 [27] Andriushchenko Maksym, Croce Francesco, Flammarion Nicolas, and Hein Matthias. Square attack: a query-efficient black-box adversarial attack via random search. In ECCV, 2020. 5 [28] Yichuan Mo, Dongxian Wu, Yifei Wang, Yiwen Guo, and Yisen Wang. When adversarial training meets vision transformers: Recipes from training to architecture. NIPS, 35: 18599–18611, 2022. 2 [29] Axi Niu, Kang Zhang, Chaoning Zhang, Chenshuang Zhang, In So Kweon, Chang D Yoo, and Yanning Zhang. Fast adversarial training with noise augmentation: A unified perspective on randstart and gradalign. arXiv preprint arXiv:2202.05488, 2022. 2 [30] Chao Pan, Qing Li, and Xin Yao. Adversarial initialization with universal adversarial perturbation: A new approach to

fast adversarial training. In AAAI, pages 21501–21509, 2024. 1 [31] Geon Yeong Park and Sang Wan Lee. Reliably fast adversarial training via latent adversarial perturbation. In ICCV, pages 7758–7767, 2021. 1 [32] Leslie Rice, Eric Wong, and Zico Kolter. Overfitting in adversarially robust deep learning. In ICML, pages 8093–8104. PMLR, 2020. 2 [33] Ali Shafahi, Mahyar Najibi, Mohammad Amin Ghiasi, Zheng Xu, John Dickerson, Christoph Studer, Larry S Davis, Gavin Taylor, and Tom Goldstein. Adversarial training for free! NIPS, 32, 2019. 2, 5, 6, 7, 8 [34] Pengyang Shao, Naixin Zhai, Lei Chen, Yonghui Yang, Fengbin Zhu, Xun Yang, and Meng Wang. Baldro: A distributionally robust optimization based framework for large language model unlearning. arXiv preprint arXiv:2601.09172, 2026. 1 [35] Gaurang Sriramanan, Sravanti Addepalli, Arya Baburaj, et al. Guided adversarial attack for evaluating and enhancing adversarial defenses. NIPS, 33:20297–20308, 2020. 5, 6, 7, 8 [36] Gaurang Sriramanan, Sravanti Addepalli, Arya Baburaj, et al. Towards efficient and effective adversarial training. In NIPS, pages 11821–11833, 2021. 2, 5 [37] Keke Tang, Tianrui Lou, Weilong Peng, Nenglun Chen, Yawen Shi, and Wenping Wang. Effective single-step adversarial training with energy-based models. IEEE Transactions on Emerging Topics in Computational Intelligence, 2024. 1 [38] Kun Tong, Chengze Jiang, Jie Gui, and Yuan Cao. Taxonomy driven fast adversarial training. In AAAI, pages 5233–5242, 2024. 1, 2, 5, 6, 7, 8 [39] Yisen Wang, Difan Zou, Jinfeng Yi, James Bailey, Xingjun Ma, and Quanquan Gu. Improving adversarial robustness requires revisiting misclassified examples. In ICLR, 2019. 1, 6, 7, 8 [40] Zeyu Wang, Xianhang Li, Hongru Zhu, and Cihang Xie. Revisiting adversarial training at scale. In CVPR, pages 24675– 24685, 2024. 2 [41] Kolter J Z. Wong E, Rice L. Fast is better than free: Revisiting adversarial training. In ICLR, 2020. 1, 2, 5, 6, 7, 8 [42] Jia Xiaojun, Zhang Yong, Wei Xingxing, Wu Baoyuan, Ma Ke, Wang Jue, and Cao Xiaochun. Prior-guided adversarial initialization for fast adversarial training. In ECCV, 2022. 5, 6, 7, 8 [43] Ling Yang, Haotian Qian, Zhilong Zhang, Jingwei Liu, and Bin Cui. Structure-guided adversarial training of diffusion models. In CVPR, pages 7256–7266, 2024. 1 [44] Yichen Yang, Xin Liu, and Kun He. Fast adversarial training against textual adversarial attacks. arXiv preprint arXiv:2401.12461, 2024. 1 [45] Yi Yu, Yufei Wang, Wenhan Yang, Lanqing Guo, Shijian Lu, Ling-Yu Duan, Yap-Peng Tan, and Alex C Kot. Robust and transferable backdoor attacks against deep image compression with selective frequency prior. IEEE Transactions on Pattern Analysis and Machine Intelligence, 2025. 1

[46] Yi Yu, Song Xia, Xun Lin, Chenqi Kong, Wenhan Yang, Shijian Lu, Yap-Peng Tan, and Alex C Kot. Towards model resistant to transferable adversarial examples via trigger activation. IEEE Transactions on Information Forensics and Security, 2025. 2 [47] Yi Yu, Song Xia, Xun Lin, Wenhan Yang, Shijian Lu, YapPeng Tan, and Alex Kot. Backdoor attacks against noreference image quality assessment models via a scalable trigger. In Proceedings of the AAAI Conference on Artificial Intelligence, 2025. 1 [48] Yi Yu, Song Xia, SIYUAN YANG, Chenqi Kong, Wenhan Yang, Shijian Lu, Yap-Peng Tan, and Alex Kot. Mtl-ue: Learning to learn nothing for multi-task learning. In International Conference on Machine Learning, 2025. 1 [49] Xinli Yue, Ningping Mou, Qian Wang, and Lingchen Zhao. Revisiting adversarial training under long-tailed distributions. In CVPR, pages 24492–24501, 2024. 1 [50] Naixin Zhai, Pengyang Shao, Binbin Zheng, Yonghui Yang, Fei Shen, Long Bai, and Xun Yang. Maximizing local entropy where it matters: Prefix-aware localized llm unlearning. arXiv preprint arXiv:2601.03190, 2026. 1 [51] Hongyang Zhang, Yaodong Yu, Jiantao Jiao, Eric Xing, Laurent El Ghaoui, and Michael Jordan. Theoretically principled trade-off between robustness and accuracy. In ICML, pages 7472–7482. PMLR, 2019. 2 [52] Yihua Zhang, Guanhua Zhang, Prashant Khanduri, Mingyi Hong, Shiyu Chang, and Sijia Liu. Revisiting and advancing fast adversarial training through the lens of bi-level optimization. In ICML, pages 26693–26712. PMLR, 2022. 5 [53] Yimeng Zhang, Xin Chen, Jinghan Jia, Yihua Zhang, Chongyu Fan, Jiancheng Liu, Mingyi Hong, Ke Ding, and Sijia Liu. Defensive unlearning with adversarial training for robust concept erasure in diffusion models. NIPS, 37:36748– 36776, 2024. 1 [54] Mengnan Zhao, Lihe Zhang, Yuqiu Kong, and Baocai Yin. Fast adversarial training with smooth convergence. In ICCV, pages 4720–4729, 2023. 2 [55] Mengnan Zhao, Lihe Zhang, Yuqiu Kong, and Baocai Yin. Catastrophic overfitting: A potential blessing in disguise. In ECCV, pages 293–310. Springer, 2024. 1, 2 [56] Yiqi Zhong, Xianming Liu, Deming Zhai, Junjun Jiang, and Xiangyang Ji. Shadows can be dangerous: Stealthy and effective physical-world adversarial attack by natural phenomenon. In CVPR, pages 15345–15354, 2022. 2

Record · ID 138831 · SHA-256 4622511293adf1e6
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.