IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
1
Unveiling the Backdoor Mechanism Hidden Behind Catastrophic Overfitting in Fast Adversarial Training
arXiv:2604.24350v1 [cs.LG] 27 Apr 2026
Mengnan Zhao, Lihe Zhang, Tianhang Zheng, Bo Wang, Baocai Yin
during training and fail to generalize to other adversarial attacks. To address this, prior works primarily target surfacelevel symptoms—such as limited adversarial diversity [27], [32], disparities in sample convergence rates [33], and inconsistencies in gradient attribution importance [34]—to design corresponding mitigation strategies [35], [36]. While these methods improve FAT stability, they offer limited insight into the underlying mechanisms of CO. Beyond these, other studies have associated CO with deeper factors such as selffitting behavior [37], gradient misalignment [38], and feature overriding [39]. Collectively, these works suggest that CO stems from abnormal pathway division in the learned model representations. However, a systematic and intuitive explanation of CO remains absent, such as the mechanisms underlying pathway division and the nature of self-information. This work innovatively analyzes CO through the lens of backdoor [40]–[43]. We begin by validating the pathway Index Terms—Fast adversarial training, catastrophic overfit- division and diverse path predictions. These analyses reveal ting, backdoor, backdoor-inspired mitigation strategies an intriguing phenomenon: CO exhibits a high similarity to backdoor-related tasks. We then show that adversarial perturbations from CO-affected models encode universal, classI. I NTRODUCTION ECENT advancements in deep learning have driven discriminative triggers, and that the primary differences besignificant progress across a wide range of applica- tween CO and backdoor-related tasks can be attributed to the tions [1]–[5]. However, these developments have also exposed trigger strength variance. Together, these findings support a critical limitations of neural networks [6]–[8], particularly unified interpretation—trigger overfitting—in which both CO their susceptibility to adversarial attacks [9]–[11]. To address and backdoor-related tasks arise from a model’s over-reliance such vulnerabilities, adversarial training has emerged as a on trigger-like features transmitted through specialized paths. We further introduce backdoor-inspired strategies to mitwidely adopted defense strategy [12]–[17], wherein perturbed igate CO. (i) We adapt established backdoor fine-tuning examples are incorporated during training to enhance model techniques—including vanilla fine-tuning, linear probing, and robustness [18]–[20]. Initial adversarial training techniques reinitialization-based methods—to recalibrate the parameters [21]–[23] produce training data using multi-step adversarial of CO-affected models. These approaches steer the model attacks, such as the projected gradient descent (PGD) [24]. In away from overfitting by shifting its focus back to informative recent years, fast adversarial training (FAT) methods [25], [26], data features, rather than to trigger-related patterns. Experiparticularly single-step approaches like FGSM-RS [27], offer mental results show that fine-tuning CO-affected models on notable computational efficiency by generating adversarial clean data can temporarily alleviate CO, though the benefit is examples with fewer backward passes [28]–[31]. short-lived as CO tends to reoccur. (ii) Inspired by weight However, FAT remains susceptible to catastrophic overfitting (CO), where models overfit to the specific attack used poisoning in backdoor attacks, we propose a weight outlier suppression constraint, which penalizes weight deviations Manuscript received Sep, 2025. from the layer-wise mean weight. This prevents the formation This work was supported by the National Natural Science Foundation of of adversarial pathways, leading to a stable FAT process. China under Grant 62431004 and 62276046. Additionally, we provide a discussion section that (i) clarMengnan Zhao is with the School of Computer Science and Technology, Anhui University, Hefei 230601, China. E-mail: [email protected]. ifies the essence of robustness improvement in adversarial Lihe Zhang and Bo Wang are with the School of Information and Communitraining, (ii) analyzes the underlying cause of CO, and (iii) cation Engineering, Dalian University of Technology (DUT), Dalian 116024, investigates whether techniques designed to mitigate CO can China. E-mail: [email protected], [email protected]. Tianhang Zheng is with the School of Computer Science and Technology, generalize to backdoor-related tasks. Zhejiang University, Hangzhou 310058, China. E-mail: [email protected]. In summary, our contributions are threefold: Baocai Yin is with the School of Computer Science and Technology, DUT, Dalian 116024, China. E-mail: [email protected]. 1) We interpret CO through the lens of backdoor, unifying Abstract—Fast Adversarial Training (FAT) has attracted significant attention due to its efficiency in enhancing neural network robustness against adversarial attacks. However, FAT is prone to catastrophic overfitting (CO), wherein models overfit to the specific attack used during training and fail to generalize to others. While existing methods introduce diverse hypotheses and propose various strategies to mitigate CO, a systematic and intuitive explanation of CO remains absent. In this work, we innovatively interpret CO through the lens of backdoor. Through validations on pathway division, diverse feature predictions, and universal class-distinguishable triggers in CO, we conceptualize CO as a weak-trigger variant of unlearnable tasks, unifying CO, backdoor attacks, and unlearnable tasks under a common theoretical framework. Guided by this, we leverage several backdoorinspired strategies to mitigate CO: (i) Recalibrate CO-affected model parameters using vanilla fine-tuning, linear probing, or reinitialization-based techniques; (ii) Introduce a weight outlier suppression constraint to regulate abnormal deviations in model weights. Extensive experiments support our interpretation of CO and show the efficacy of the proposed mitigation strategies.
R
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
2
CO, backdoor attacks, and unlearnable tasks under a trigger overfitting framework. 2) We validate key similarities—pathway division, diverse predictions, and universal triggers—between CO and backdoor. 3) We introduce backdoor-inspired mitigation strategies, including adapted fine-tuning and a weight outlier suppression constraint, demonstrating their effectiveness empirically. The remainder of the paper is organized as follows. Section II reviews recent advances in adversarial training and backdoor-related tasks. In Section III, we establish the connection between CO and backdoor-related tasks. Section IV presents a set of backdoor-inspired strategies for mitigating CO. Section V offers further discussion and analysis, such as insights into the causes of CO and the robustness of FAT. Finally, Section VI concludes the paper and outlines several promising directions for future research. II. R ELATED WORK
(a) Distance confusion matrix
A. Adversarial training Recent breakthroughs in deep neural networks [1]–[3] have prompted extensive research into their security risks [6]–[8], with particular attention to their vulnerability to adversarial attacks [9]–[11], [44], [45]. In response, adversarial training [18]–[20], [46], [47] has emerged as a popular strategy to enhance model robustness, employing both multi-step (e.g., PGD [24]) and one-step adversarial attacks (e.g., FGSM [48]) [21]–[23], [28], [49], [50]. Compared to PGD-based methods, FGSM-based methods such as FGSM-RS [27] and FGSM-MEP [51], also called FAT [52], [53], are computationally efficient [54]. Given the initial perturbation δ0 ∈ N (0, I), training dataset D, network f (·; θ), loss L, step size ϵ, and budget ξ, FGSM-RS generates adversarial perturbations by Eqs. (1) and (2), δ = clipξ δ0 + ϵ · δsign , (1) δsign = sign ∇δ0 L f (x + δ0 ; θ), y , (2) and implements adversarial training by Eq. (3), min Ex∼D L f x + δ; θ , y . θ
(c) UMAP distribution of δsign in the CO-affected model Fig. 1: Distance matrix and UMAP visualizations under FGSM-RS. Small distances imply reduced separation between class distributions.
(3)
In contrast to FGSM-RS, FGSM-MEP constructs δ0 by leveraging the momentum accumulated from adversarial perturbations computed over preceding epochs. It also incorporates a prediction regularization during minimization, expressed as Rpred = β∥f (x + δ) − f (x + δ0 )∥22 .
(b) UMAP distribution of δsign in the stably trained model
(4)
However, FAT approaches may suffer from CO [55]–[57]. To address this issue, various strategies have been proposed [58], such as gradient alignment [38], convergence smoothness [33], zero-gradient clipping [34], bi-level optimization [59] and feature activation consistency [39]. Recently, He et al. [37] assume that adversarial perturbations embed self-information and argues that models acquire this information through a separate pathway. Then, Zhao et al. [39] introduce a strong regularization term that enforces consistency between predictions on clean and adversarial examples, thereby suppressing the
emergence of the adversarial pathway. Lin et al. [60] leverage both weight-level and example-level adversarial perturbations to enhance training stability by enforcing weight robustness. In contrast, this work explains CO through the lens of backdoor, unifying CO, backdoor attacks, and unlearnable tasks under a common interpretation of trigger overfitting. Within this perspective, we introduce several backdoor-inspired mitigation strategies that not only enable models to break free from CO but also suppress its occurrence. Particularly, unlike existing methods that suppress CO by adding regularization constraints to the prediction or feature space, this work adjusts the distribution of model weights. Furthermore, while both this work and that of Lin et al. [60] identify weight anomalies, their objective is to construct robust weights, whereas ours is specifically to suppress weight outliers. We observe that enforcing overall weight robustness will limit the model’s ability to fit
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
3
clean samples. Moreover, although both our weight outlier suppression constraint and the strong regularization proposed by [39] are effective at preventing adversarial pathways, the latter is often compromised by misclassifications on clean examples. Additionally, it does not explain why adversarial paths emerge and override data paths. This work addresses this question from the perspective of trigger overfitting. B. Backdoor-related tasks This section covers both backdoor attacks and unlearnable tasks, the latter representing a transferable application of backdoor attacks. Backdoor attacks pose a serious threat to the security of deep neural networks. Early works such as BadNets [61] and TrojanNN [62] demonstrate that inserting poisoned samples with static triggers during training can cause targeted misclassification. However, such attacks are often detectable due to their reliance on fixed and conspicuous patterns. To improve stealth and generalizability, a range of trigger designs have been proposed [63]–[65]. For instance, spatial transformations [66], image blending [67], and frequencydomain perturbations [64], [68] aim to create imperceptible or input-adaptive triggers. Others have explored learnable or sample-specific backdoor strategies [63], [69] to improve trigger effectiveness and bypass detection. More recently, attention has turned to contrastive and self-supervised learning frameworks. CTRL [70] shows that even models trained without labels are susceptible to backdoor insertion. Beyond implanting malicious behaviors into trained models, backdoor techniques have been adapted to unlearnable tasks, which aim to impair a model’s generalization on clean data. Huang et al. [71] introduce sample-wise and class-wise perturbations—similar in nature to triggers—into all training samples, inducing overfitting and preventing the learning of useful representations. However, their effectiveness diminishes under different training settings or datasets. To improve transferability, Ren et al. [72] propose a Classwise Separability Discriminant strategy that enhances linear separability. Zhang et al. [73] generate label-agnostic unlearnable examples via cluster-wise perturbations. Liu et al. [74] extend protection to multimodal contrastive learning. Notably, Qin et al. [75] show that adversarial augmentations can mitigate unlearnable-example attacks, motivating Fu et al. [76] to design robust unlearnable examples against adversarial learning. Furthermore, Ye et al. [77] present ungeneralizable samples that can only be learned by specific networks, and Li et al. [78] develop methods to detect and corrupt convolutionbased unlearnable examples. This work considers both standard backdoor attacks and unlearnable tasks, and investigates their connections to CO. III. CO AND BACKDOOR - RELATED TASKS In this section, we first present the motivation for interpreting CO through the lens of backdoor. We then provide a comparative analysis highlighting the similarities and differences between CO and backdoor-related tasks. Table I summarizes the notation used in this paper for quick reference.
(a) Distance confusion matrix
(b) UMAP distribution of δsign in the stably trained model
(c) UMAP distribution of δsign in the CO-affected model Fig. 2: Distance matrix and UMAP visualizations under FGSM-MEP. Small distances imply reduced separation between class distributions.
A. Motivation of explaining CO from backdoor Prior FAT work [39] has proposed two key hypotheses regarding CO: (i) CO-affected models can be decomposed into distinct adversarial and data branches; (ii) the adversarial branch exhibits feature overriding for adversarial inputs. In this paper, we systematically explore these hypotheses by analyzing the behaviors of FGSM-RS and FGSM-MEP. Pathway division: In the absence of the adversarial pathway, the primary distribution of adversarial perturbations δsign should be class-separable. To verify this, we compute the interclass Wasserstein distance based on δsign . The experiments are conducted on CIFAR-10 [79] using a ResNet18 backbone [80], with the maximum perturbation budget of 16/255 and the learning rate of 0.1. The results in Figure 1(a) and Figure 2(a) reveal that the stable model (epoch 5) maintains distinct inter-class distances, while the CO-affected model (epoch 20)
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
4
TABLE I: Description of symbols. Symbols
Description
δ; δ0 ; δsign ; δmom θ; θadv ; θdata ; θCO ξ; ξT ; ξE η; α; ϵ
Adversarial perturbation; Initial perturbation; Perturbation direction; Momentum-based perturbation Model parameters; Adversarial path parameters; Data path parameters; CO-affected model parameters Maximum perturbation budget; Maximum perturbation budget during training; Maximum perturbation budget during evaluation Hyperparameters; Step size
exhibits distribution overlap (distance 0). These observations are further supported by Uniform Manifold Approximation and Projection (UMAP) distribution [81] in Figures 1 and 2, showing class-separable patterns in stable models versus complete overlap under CO-affected models.1 Namely, the CO-affected model should contain an additional adversarial branch. We formally decompose the network parameters as θ = {θadv , θdata }, corresponding to the adversarial and data pathways, respectively. The lack of feature discrimination in δsign suggests that the adversarial pathway dominates gradient backpropagation, expressed as ∇δ0 L(f (x + δ0 ; θ), y) ≈ ∇δ0 L(f (x + δ0 ; θadv ), y).
(a) FGSM-RS (CO in 13th Epoch)
(5)
Diverse forward predictions instead of feature overriding. To analyze the forward prediction behavior of FAT models, we visualize the training dynamics of diverse methods on CIFAR10 with ResNet-18 in Figure 3. After CO, we observe: (b) FGSM-MEP (CO in 12th Epoch)
Ex∼Dtrain ACC(x + δ) ≥ Ex∼Dtrain ACC(x + δ0 ) ≈ Ex∼Dtrain ACC(x) ≫ 0.
(6)
1) Eq. (6) implies that CO-affected models exhibit higher or comparable classification accuracy on adversarial examples than on clean x and initially perturbed samples x + δ0 . 2) Meanwhile, CO-affected models remain classifiable to x and x + δ0 . These observations indicate that, during forward inference, the CO-affected model may integrate features from both data and adversarial pathways, min L(f (x + δ; θ), y) ≈ θ h i min L(f (x + δ; θadv ), y) + L(f (x + δ; θdata ), y) ,
(7)
θadv ,θdata
or directly utilize the data pathway, min L(f (x + δ; θ), y) ≈ min L(f (x + δ; θdata ), y). θ
θdata
(8)
The validation of pathway decomposition and diverse forward prediction raises a natural question: Is there an intrinsic connection between CO and backdoor attacks? Table II shows comparisons between CO and backdoor-related tasks. Specifically, we hypothesize the existence of a backdoor pathway in the backdoored model. When this pathway is activated, typically by a trigger, feature overriding occurs, leading the model to confidently classify the input into a predefined target class. In the absence of triggers, the model backs to standard prediction behavior and maintains high accuracy on clean samples. Meanwhile, unlearnable tasks [76], [85], [86] embed perturbations, functionally similar to backdoor triggers, into all samples while retaining their original labels, causing the model 1 UMAP
is a dimensionality reduction technique used for visualizing primary and high-dimensional features.
Fig. 3: Training dynamics of existing FAT methods. to rapidly overfit to these perturbations (i.e. high accuracy for perturbed inputs) and lose generalization capability (i.e., low accuracy for clean examples). Notably, both types remain vulnerable to adversarial attacks, as their training data are not augmented with adversarial examples. A similar pattern is observed in CO-affected models, which overfit to the specific attack used during training (e.g., FGSM), achieving high accuracy on seen adversarial types while failing to generalize to unseen ones. Taken together, these findings suggest a strong behavioral similarity between CO and backdoor-related tasks. B. Further analyses between backdoor and CO We further investigate whether FGSM-generated perturbations FGSM(x) during CO exhibit characteristics similar to backdoor triggers. Specifically, we hypothesize that δsign contains a universal class-discriminative component, referred to as the UCD trigger δucd . To test this hypothesis, we compute the class-wise expectation of δsign as defined in Eq. (9), and incorporate an auxiliary constraint Laux into the adversarial training objective as specified in Eq. (10). t t δmom ← 0.9 δmom + 0.1
E
x∈B,y=t
t t Laux = −α∥δsign − δmom ∥2 ,
t,∗ δsign ,
(9) (10)
where B represents the current batch, t denotes the designated label, and α is set to 1e−2 by default. The momentum-based t term δmom aims to dynamically extract UCD triggers. Laux
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
5
TABLE II: Comparison between backdoor-related tasks and CO. Default 1
Backdoor Attack Unlearnable
Catastrophic Overfitting (CO)
High accuracy for triggered inputs xtrigger High accuracy for perturbed inputs xperturbed (classified as the predefined class) (classified as the truth class)
High accuracy for adversarial examples x + FGSM(x) (classified as the truth class)
2 High accuracy for x
Low accuracy for x FGSM-RS: Moderate accuracy for x From low to high as the perturbation size increases FGSM-MEP: Higher accuracy than FGSM-RS for x
3 Low accuracy for attacks
Low accuracy for attacks
Low accuracy for attacks except FGSM
TABLE III: Comparison between various FAT methods. Methods marked with † and ‡ utilize Eq. (10), with the parameter α set to 1e−2 and 1, respectively. ‘Best’ and ‘Final’ refer to the evaluation results of the model with the best PGD-10 performance and the final model, respectively. ‘Clean’ denotes the classification accuracy for clean examples. ‘Perturbed’ means that we employ initially perturbed samples. FGSM, PGD, C&W [82] and AA [83] are adversarial attacks for evaluating robustness. ‘CO’ signifies whether the trained model falls into catastrophic overfitting during FAT. Net&Dataset
Methods FGSM-RS [27]
ResNet18 [80]
FGSM-RS† [27]
CIFAR10 [79]
FGSM-MEP [51] FGSM-MEP‡ [51] FGSM-RS [27]
PreActResNest18 [84]
FGSM-RS† [27]
CIFAR10 [79]
FGSM-MEP [51] FGSM-MEP‡ [51] FGSM-RS [27]
ResNet18 [80]
FGSM-RS† [27]
CIFAR100 [79]
FGSM-MEP [51] FGSM-MEP‡ [51] FGSM-RS [27]
PreActResNest18 [84]
FGSM-RS† [27]
CIFAR100 [79]
FGSM-MEP [51] FGSM-MEP‡ [51]
Clean↑
Perturbed↑
FGSM↑
PGD10↑
PGD20↑
PGD50↑
C&W↑
AA↑
best final best final best final best final
54.69 80.80 73.55 73.34 30.83 88.07 59.73 59.69
54.64 83.16 73.05 73.05 30.58 88.34 59.05 59.09
34.31 76.62 44.00 44.15 25.69 80.03 42.27 42.02
26.01 0.00 34.20 33.99 23.09 11.40 37.39 37.25
19.77 0.00 23.99 24.04 21.67 6.75 32.43 32.20
17.85 0.00 20.78 21.00 21.49 3.74 31.60 30.99
18.41 0.00 22.74 22.61 20.35 3.48 25.97 25.67
12.63 0.00 14.63 14.76 18.78 0.08 22.22 22.19
best final best final best final best final
56.38 81.22 72.71 72.68 41.81 88.66 59.29 59.61
56.22 80.56 72.21 72.49 41.78 89.43 58.28 58.90
33.95 80.33 45.45 44.78 29.19 82.76 42.07 41.85
28.21 0.00 34.94 34.76 26.87 11.17 36.69 36.29
21.71 0.00 25.11 25.01 24.26 6.82 31.94 31.14
19.66 0.00 22.38 21.90 23.89 4.09 30.74 30.07
18.81 0.00 22.21 22.22 18.42 2.84 25.50 25.54
13.49 0.00 14.78 14.82 17.08 0.01 21.40 21.26
best final best final best final best final
30.40 39.40 45.12 48.70 22.18 67.40 42.51 42.69
30.28 52.36 45.33 48.71 22.17 67.73 42.25 42.70
14.78 47.41 20.80 22.58 14.51 54.74 25.15 25.20
11.74 0.00 16.21 16.57 12.57 1.33 21.15 20.76
9.18 0.00 11.78 11.70 10.95 0.60 17.28 17.32
8.82 0.00 11.02 10.82 10.73 0.27 16.71 16.59
7.97 0.00 10.61 10.86 8.34 0.31 14.01 14.12
6.12 0.00 7.95 8.16 7.16 1.00 11.20 11.39
best final best final best final best final
27.94 50.30 46.79 47.16 21.21 67.59 43.22 43.22
27.72 51.80 46.88 47.04 21.02 69.12 42.96 42.95
13.92 50.00 21.83 21.94 13.61 55.06 25.91 25.99
11.40 0.00 16.33 16.14 11.79 2.31 20.84 20.89
9.14 0.00 11.80 11.62 10.22 1.20 16.77 16.88
8.83 0.00 10.93 10.53 10.01 0.64 16.04 16.00
7.61 0.00 11.21 10.76 7.72 0.71 14.07 14.01
6.16 0.00 7.83 7.69 6.75 0.00 11.18 11.20
penalizes class-consistent perturbation features, forcing the model to forget UCD triggers. Gradients are prevented from t,∗ t t , denoted as δsign in propagating into δmom by detaching δsign Eq. (9). As Table III shows, Laux mitigates CO in FGSMRS and FGSM-MEP across different networks and datasets, confirming the existence of UCD triggers. We next analyze the performance of CO-affected and backdoored models on clean inputs. The clean classification accuracies across different paradigms are summarized as follows: ACCOri ≈ ACCSBackdoor > ACCMEP ≫ ACCUnlearnable , (11) where ACCSBackdoor , ACCOri , ACCMEP , and ACCUnlearnable denote the clean-sample accuracies of models under the standard backdoor, original, CO-affected (trained with FGSM-MEP),
CO ✔ ✗ ✔ ✗ ✔ ✗ ✔ ✗ ✔ ✗ ✔ ✗ ✔ ✗ ✔ ✗
and unlearnable settings, respectively. Notably, ACCMEP falls between ACCSBackdoor and ACCUnlearnable , suggesting that COaffected models preserve a moderate level of clean-data generalization. In backdoor attacks, the model associates a fixed trigger with a target class while largely preserving its performance on clean data. In contrast, unlearnable tasks embed diverse perturbations to training data without altering labels, causing the model to overfit to these perturbations and lose the ability to accurately classify clean examples. Notably, in unlearnable tasks, we observe that clean accuracy remains comparable to ACCOri when the embedded perturbation is weak. As perturbation strength increases, the model’s performance on clean data progressively degrades and collapses once a critical threshold is reached. This trend indicates that
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
6
Weight Shift
Pre-trained Model
: Pre-trained model weights : Re-initialized weights : Frozen : Tuned
(a) Vanilla FT
(b) Linear Probing (c) Reinitialization Freezing (d) Reinitialization FT
(e) Weight shift FT
Fig. 4: Backdoor fine-tuning techniques. FT: finetuning.
the higher clean accuracy observed in CO-affected models (e.g., ACCMEP ) compared to unlearnable tasks (ACCUnlearnable ) can be attributed to the weaker trigger effect. This is further supported by empirical evidence: as illustrated in Figure 2, adversarial perturbations in CO-affected models trained with FGSM-MEP are primarily composed of non-discriminative features, with UCD triggers contributing only marginally. By verifying the existence of UCD triggers in adversarial perturbations generated for CO-affected models, and by analyzing the differences in clean accuracy across techniques, we argue that CO, standard backdoor attacks, and unlearnable tasks can be understood within a unified perspective of trigger overfitting. More specifically, CO can be viewed as a weaker variant of unlearnable tasks; both can be regarded as transfer applications of backdoor attacks. IV. BACKDOOR - INSPIRED MITIGATION OF CO Motivated by previous comparisons, we attempt to employ backdoor defenses to mitigate CO. We begin by adapting existing fine-tuning methods designed for backdoor removal. Furthermore, inspired by weight poisoning techniques in backdoor, we introduce a supplementary constraint to suppress weight outliers and enhance stability. A. Adapted backdoor fine-tuning strategies Once the model f (x; θ) falls into CO during FAT, we apply one epoch of fine-tuning using the strategies illustrated in Figure 4, with the details given as follows: 1) VFT [87], [92]: Fine-tune all parameters. 2) LP [88], [93]: Fine-tune first k model layers, 3) RF [89]: Reinitialize first k layers, freeze them, then fine-tune remaining layers. 4) RFT [89]: Reinitialize first k layers, then fine-tune the whole model. 5) RSFT [89]: Reinitialize first k layers and fine-tune the whole model with an inner product constraint to limit weight shift. We finetune CO-affected models on clean examples as ‘*CO-Clean’. Tables IV and V show the experimental results obtained using FGSM-RS and FGSM-MEP, respectively2 . 2 Results of finetuning co-affected models on adversarial examples are shown in Appendix.
Our observations are as follows: 1) Adapted backdoor finetuning strategies can also resolve CO in FAT. 2) The clean classification accuracies reported in Table V are consistently lower than those in Table IV. This degradation may stem from the regularization term Rgrad in Eq. (4), which potentially constrains the model’s capacity to fit the clean data distribution. Furthermore, reinitializing a stable FAT process after CO may inherently limit the model’s final performance, e.g. wrong optimization direction. 3) Notably, even when finetuning yields a temporarily stable model, CO tends to reoccur. This is expected, as the goal of backdoor defenses is to transition a trained, backdoored model out of the backdoor state. Overall, fine-tuning strategies designed for backdoor attacks are also applicable to mitigating CO.
B. Weight-poisoning inspired strategy To develop lasting mitigation strategies for CO, we begin by examining training-time backdoor defense techniques [94]. These methods typically identify and filter backdoor-infected samples using either one-shot [95] or iterative [96] filtering procedures. In FAT, Zhao et al. [33] demonstrate that applying a convergence-smoothing constraint to a subset of samples with notable convergence divergence can enhance the overall training stability. However, the effectiveness of this method is highly sensitive to biases in sample selection. Motivated by backdoor approaches with weight poisoning [97]–[99], we examine the weight distributions of both stable and CO-affected models. Specifically, we compute the mean weight value w for each layer and assess the weight distribution relative to this mean. Figure 5 presents experimental results on CIFAR-10 with ResNet-18, showing that CO-affected models exhibit weight outliers—weights that deviate markedly from w. Hence, we attempt to address CO by suppressing weight outliers. Unlike backdoor pruning techniques [100]– [102], which typically clamp weights, we introduce an additional constraint that suppresses weight outliers and prevents the formation of adversarial paths. X X |w − wl | Lreg = exp − η |w − wl |, (12) wl + α l∈C w∈wl
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
7
TABLE IV: Comparison of various methods on CIFAR-10 using ResNet-18, with FGSM-RS as the baseline. “Best” and “Final” denote results from the checkpoint with the best PGD-10 accuracy and the final epoch, respectively. ‘Clean’ denotes the classification accuracy for clean examples. ‘Perturbed’ means that we employ initially perturbed samples. FGSM, PGD, C&W, and AA represent different adversarial attacks used for robustness evaluation. “ST” indicates the number of stable runs out of three (◦: CO occurred; ⋆: stable training). Methods
Clean↑
Perturbed↑
FGSM↑
PGD10↑
PGD20↑
PGD50↑
C&W↑
AA↑
ST
FGSM-RS [27]
best final
54.69 80.80
54.64 83.16
34.31 76.62
26.01 0.00
19.77 0.00
17.85 0.00
18.41 0.00
12.63 0.00
◦◦◦
VFT-CO-Clean [87]
best final
73.49 73.79
72.87 73.34
44.94 44.48
34.81 34.35
25.32 24.64
21.87 21.75
23.00 22.63
15.19 15.11
⋆⋆⋆
LP-CO-Clean [88]
best final
61.29 78.43
60.83 81.50
37.78 77.59
30.98 0.00
23.82 0.00
22.14 0.00
21.02 0.00
15.35 0.00
◦◦◦
RF-CO-Clean [89]
best final
69.92 72.82
70.74 72.75
43.28 44.00
34.70 34.04
25.62 25.04
22.81 21.87
22.71 22.45
15.82 15.22
⋆⋆⋆
RFT-CO-Clean [89]
best final
70.92 72.82
69.63 72.67
43.29 44.08
34.70 34.02
25.43 25.12
22.83 21.98
22.66 22.53
15.81 15.28
⋆⋆⋆
RSFT-CO-Clean [89]
best final
71.79 72.17
71.51 71.80
43.78 43.50
34.14 33.84
24.82 24.38
21.59 21.55
22.34 22.08
14.89 14.93
⋆⋆⋆
TABLE V: Comparison between various techniques, with FGSM-MEP as baseline. “Best” and “Final” denote results from the checkpoint with the best PGD-10 accuracy and the final epoch, respectively. ‘Clean’ denotes the classification accuracy for clean examples. ‘Perturbed’ means that we employ initially perturbed samples. FGSM, PGD, C&W, and AA represent adversarial attacks for evaluation. “ST” indicates the number of stable runs out of three (◦: CO occurred; ⋆: stable training). Methods
Clean↑
Perturbed↑
FGSM↑
PGD10↑
PGD20↑
PGD50↑
C&W↑
AA↑
ST
FGSM-MEP [51]
best final
59.01 82.37
58.94 82.22
35.69 71.19
28.66 13.13
21.81 8.09
20.10 3.12
20.49 0.89
12.39 0.00
◦◦◦
VFT-CO-Clean [87]
best final
58.43 58.25
57.93 57.75
38.18 38.28
33.71 33.79
29.32 29.39
28.44 28.62
22.93 23.03
20.16 20.15
⋆⋆⋆
LP-CO-Clean [88]
best final
49.03 49.03
48.42 48.55
34.41 34.43
31.18 31.16
27.67 27.67
27.20 27.13
22.69 22.74
20.71 20.75
⋆◦◦
RF-CO-Clean [89]
best final
49.03 49.03
48.50 48.45
34.57 34.44
31.15 31.16
27.63 27.55
27.10 27.17
22.73 22.75
20.73 20.78
⋆◦◦
RFT-CO-Clean [89]
best final
55.79 55.74
55.07 55.05
39.31 38.94
34.54 34.29
30.06 30.11
29.34 29.37
24.45 24.31
21.10 21.04
⋆⋆⋆
RSFT-CO-Clean [89]
best final
40.13 39.00
39.65 38.64
29.60 28.45
25.85 25.46
24.57 23.22
24.26 22.88
21.14 20.45
19.41 18.50
⋆⋆⋆
wl =
1 X |w|, |Wl |
(13)
w∈Wl
where ∥ · ∥ is utilized to obtain the absolute value. C denotes the set of convolutional layers, and Wl represents the set of l-th layer weights. Hyper-parameters η and α are set to 10 and 10−5 , respectively. Table VI reports results on CIFAR-10 with ResNet18, showing that incorporating Lreg effectively mitigates CO in FAT. Additional experiments across diverse datasets, architectures and perturbation budgets are shown in the Appendix, which also bring consistent improvements in mitigating CO. Overall, the proposed method not only resolves CO but also achieves performance superior to or comparable with the state-of-the-art PBD. Unlike PBD, which aligns clean and adversarial predictions at the cost of clean accuracy, and LAT, whose effectiveness is limited by its strict emphasis on robust weights, our method targets only anomalous weights, providing greater flexibility in weight selection. Additionally, we hypothesize that these anomalous weights predominantly
form the adversarial pathway. We further conduct ablation studies to compare our method against simpler alternatives, including ℓ2 regularization and weight clipping based on the ratio between each weight and its mean (using the same threshold as our method). For each approach, we report results averaged over three independent runs in Fig. 6. As can be observed, directly applying ℓ2 regularization fails to achieve stable optimization, where the model loses its adversarial robustness in the later stages of fast adversarial training. Weight clipping, while partially mitigating catastrophic overfitting, suffers from training instability and may relapse into catastrophic overfitting at any point. Moreover, robustness after clipping is slightly compromised. In contrast, our method consistently maintains stable performance throughout the training process. V. D ISCUSSIONS AND ADDITIONAL ANALYSES Why FAT improves robustness and why CO occurs. We further investigate why single-step FAT (e.g., FGSM-based
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
8
Fig. 5: Weight distribution relative to the mean weight. ‘Count’ means the distribution percentage. TABLE VI: Comparison with existing FAT methods with ResNet18 and CIFAR10. PBD is the state-of-the-art method. ξT and ξE represent the perturbation budget used during the training and evaluation processes. “Best” and “Final” denote results from the checkpoint with the best PGD-10 accuracy and the final epoch, respectively. ‘Clean’ denotes the classification accuracy for clean examples. ‘Perturbed’ means that we employ initially perturbed samples. “ST” indicates the number of stable runs out of three (◦: CO occurred; ⋆: stable training). For LAP, we directly utilize the results reported in [60]. ξT = ξE = 16/255 LAP [60] GradAlign [35] ZeroGrad [90] NuAT [91] FGSM-RS [27] PBD-RS [39] Ours(Lreg ) FGSM-MEP [51] PBD-MEP [39] Ours(Lreg )
Clean↑
Perturbed↑
FGSM↑
PGD10↑
PGD20↑
PGD50↑
C&W↑
AA↑
ST
best best final best final best final
63.73 58.17 70.86 74.16 75.60 74.62 75.29
-
39.87 69.51 43.96 44.89 44.92 45.31
33.12 0.00 32.67 31.77 35.22 34.85
26.81 0.00 21.98 20.71 25.93 25.58
24.99 0.00 18.37 16.76 23.67 23.44
22.63 0.00 20.76 20.09 24.07 23.62
19.55 17.02 0.00 12.07 10.87 18.43 18.06
-
best final best final best final
54.69 80.80 73.32 73.82 74.33 74.84
54.64 83.16 72.96 73.48 74.00 74.67
34.31 76.62 45.03 44.60 46.60 46.81
26.01 0.00 35.45 34.76 36.82 36.21
19.77 0.00 25.52 25.05 26.46 25.95
17.85 0.00 22.61 22.00 22.98 22.70
18.41 0.00 23.72 23.23 24.24 24.22
12.63 0.00 16.21 15.82 15.72 15.70
◦◦◦
best final best final best final
59.01 82.37 64.45 64.20 65.26 65.33
58.94 82.22 63.28 63.21 64.74 64.63
35.69 71.19 45.77 45.29 46.11 46.19
28.66 13.13 39.70 39.26 40.23 40.06
21.81 8.09 34.27 33.47 34.53 34.22
20.10 3.12 32.80 31.88 33.07 32.73
20.49 0.89 27.64 27.93 27.98 28.13
12.39 0.00 22.45 22.40 22.47 22.21
◦◦◦
AT) tends to induce trigger-like overfitting, whereas multistep AT (e.g. PGD-based AT [103]–[105]) does not. To this end, we conduct additional experiments comparing the two methods using multiple similarity metrics. As shown in Fig. 7, we report inter-class and intra-class prediction similarity, adversarial perturbation similarity, and adversarial example similarity. Our results reveal three key observations. First, inter-class similarity exhibits comparable trends for both methods, indicating that they behave similarly across different learning stages. Second, intra-class perturbation similarity and intra-class adversarial example similarity are lower for FGSMbased AT than for PGD-based AT. This does not necessarily imply that the “trigger” patterns in FGSM perturbations are more consistent; rather, it may reflect that PGD, through its multi-step optimization, discovers more semantically coherent adversarial directions, leading to higher intra-class similarity in its perturbations. Third, intra-class prediction similarity is substantially higher for FGSM-based AT than for PGD-based AT. This suggests that although FGSM generates perturbations that appear more diverse (i.e., lower perturbation similarity),
◦◦◦ ⋆⋆⋆ ⋆⋆⋆
⋆⋆⋆ ⋆⋆⋆
⋆⋆⋆ ⋆⋆⋆
the model’s predictions on them are remarkably uniform. In contrast, for PGD-based AT, predictions on intra-class adversarial examples are less consistent, indicating that PGD introduces a wider variety of perturbation directions. Overall, the single-step optimization of FGSM restricts its adversarial directions, thereby causing CO. Transferability of methods across different tasks. Section IV shows that backdoor-inspired defense techniques can mitigate CO. We further evaluate whether the method for CO can generalize to backdoor-related tasks. For unlearnable techniques with sample-wise perturbations, adding uniform random noise to the poisoned dataset can effectively neutralize their impact. For example, training a ResNet-18 model on CIFAR with the original unlearnable dataset yields 13.58% accuracy on clean samples, which increases to 94.22% after applying random noise. Hence, we focus on unlearnable techniques with class-wise perturbations. Prior work often addresses these using adversarial training [76]; for instance, a model trained on the original unlearnable dataset achieves 10.05% accuracy on clean samples,
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
9
Fig. 6: Ablation studies comparing our Lreg against simpler alternatives.
Fig. 7: Comparison between FGSM-based AT and PGD-based AT.
whereas adversarial training improves accuracy to 86.89%. We investigate the transferability of the CO mitigation strategy to unlearnable tasks, with results summarized in Table VII. When the poisoning budget in unlearnable tasks is small, the Lreg for mitigating CO can also resolve the class-wise unlearnable attack, achieving performance comparable to the AT-based approach but at a significantly lower computational cost. However, under a large poisoning budget (e.g. 8/255), its effectiveness diminishes. These results further support viewing CO as a weak-trigger variant of unlearnable tasks. Additionally, the reduced effectiveness of Lreg under larger poisoning budgets (e.g., 8/255) can be attributed to two key observations. First, the occurrence of weight outliers is not positively correlated with perturbation magnitude. In the extreme case where unconstrained adversarial examples are used, the unlearnable sample set is effectively transformed into a completely different dataset. Under this condition, the model exhibits no weight outliers, rendering Lreg completely ineffective. This suggests that the performance degradation of Lreg under larger perturbations arises because larger perturbations
facilitate a direct association between the perturbation and the target label, thereby bypassing the need for weight anomalies. Second, our experiments focus on class-wise unlearnable perturbations, which tend to mimic the semantic features of a specific class. Once such perturbations successfully capture class semantics, the unlearnable samples again resemble a different dataset, resulting in the absence of weight outliers. Impact of the hyperparameter β. The results in Figure 8 show that both too small and too large values of β adversely affect model robustness. Specifically, when β is too small, the regularization not only penalizes weight outliers but also constrains normal weights, thereby impairing the overall model performance. Conversely, when β is too large, the suppression of weight outliers becomes insufficient, rendering the regularization ineffective. VI. C ONCLUSIONS AND FUTURE WORKS In this work, we offer a systematic explanation of CO from a backdoor perspective, establishing a unified framework—trigger overfitting—that encompasses standard back-
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
10
Fig. 8: Ablation studies on the hyperparameter β. TABLE VII: Transferability of the CO mitigation strategy (Lreg in Eq. (12) ) to unlearnable tasks [71]. ‘Poisoned’, ‘AT’, and ‘Lreg ’ correspond to models trained on the CIFAR10 poisoned dataset using ResNet-18 with standard training, adversarial training, and standard training augmented with Lreg , respectively. Each model is trained for 60 epochs using a cyclic learning rate schedule [106]. The reported values represent the final evaluation accuracy↑ on clean samples. Poisoning Budgets
Poisoned [71]
AT [76]
Lreg
4/255 6/255 8/255
17.50 10.05 9.85
86.95 86.89 86.32
86.71 86.48 46.57
Training Time (minutes)
15.4
36.8
15.5
door attacks, CO, and unlearnable tasks. By validating phenomena such as pathway division, diverse path predictions, and the presence of universal class-distinguishable triggers in CO, we conceptualize CO as a weak-trigger variant of unlearnable tasks. Leveraging these insights to mitigate CO, we adapt several backdoor fine-tuning techniques to the FAT framework and propose a weight outlier suppression constraint. Experimental results validate our mechanistic explanation and demonstrate the effectiveness of backdoor-inspired strategies in alleviating CO. Future research: This work has demonstrated that adversarial perturbations crafted by CO-affected models contain universal class-distinguishable (UCD) triggers. A natural issue arises: Can these UCD triggers be further extracted and explicitly identified? In other words, is it possible to deliberately induce the CO phenomenon with the extracted UCD triggers? Moreover, can triggers associated with other adversarial attacks be similarly extracted? Models tend to overfit to triggers, resulting in high-confidence predictions on corresponding adversarial samples. Therefore, when defending against a specific attack, integrating manually extracted adversarial triggers can be more effective, as conventional adversarial training methods involve a trade-off between clean accuracy and robustness.
R EFERENCES [1] C. Zuo, J. Qian, S. Feng, W. Yin, Y. Li, P. Fan, J. Han, K. Qian, and Q. Chen, “Deep learning in optical metrology: a review,” Light: Science & Applications, vol. 11, no. 1, p. 39, 2022. [2] S. M. Mousavi and G. C. Beroza, “Deep-learning seismology,” Science, vol. 377, no. 6607, p. eabm4470, 2022. [3] T. D. Pereira, N. Tabris, A. Matsliah, D. M. Turner, J. Li, S. Ravindranath, E. S. Papadoyannis, E. Normand, D. S. Deutsch, Z. Y. Wang et al., “Sleap: A deep learning system for multi-animal pose tracking,” Nature methods, vol. 19, no. 4, pp. 486–495, 2022. [4] M. Baek and D. Baker, “Deep learning and protein structure modeling,” Nature methods, vol. 19, no. 1, pp. 13–14, 2022. [5] N. Sapoval, A. Aghazadeh, M. G. Nute, D. A. Antunes, A. Balaji, R. Baraniuk, C. Barberan, R. Dannenfelser, C. Dun, M. Edrisi et al., “Current progress and open challenges for applying deep learning across the biosciences,” Nature Communications, vol. 13, no. 1, p. 1728, 2022. [6] S.-Y. Chou, P.-Y. Chen, and T.-Y. Ho, “How to backdoor diffusion models?” in CVPR, 2023, pp. 4015–4024. [7] L. Wei, L. Jin, and X. Luo, “Noise-suppressing neural dynamics for time-dependent constrained nonlinear optimization with applications,” IEEE Transactions on Systems, Man, and Cybernetics: Systems, vol. 52, no. 10, pp. 6139–6150, 2022. [8] W. Chen, B. Wu, and H. Wang, “Effective backdoor defense by exploiting sensitivity of poisoned samples,” vol. 35, 2022, pp. 9727– 9737. [9] Y. Cao, C. Xiao, A. Anandkumar, D. Xu, and M. Pavone, “Advdo: Realistic adversarial attacks for trajectory prediction,” in ECCV. Springer, 2022, pp. 36–52. [10] J. Gu, H. Zhao, V. Tresp, and P. H. Torr, “Segpgd: An effective and efficient adversarial attack for evaluating and boosting segmentation robustness,” in ECCV. Springer, 2022, pp. 308–325. [11] Y. Zhong, X. Liu, D. Zhai, J. Jiang, and X. Ji, “Shadows can be dangerous: Stealthy and effective physical-world adversarial attack by natural phenomenon,” in CVPR, 2022, pp. 15 345–15 354. [12] L. Yang, H. Qian, Z. Zhang, J. Liu, and B. Cui, “Structure-guided adversarial training of diffusion models,” in CVPR, 2024, pp. 7256– 7266. [13] S. Casper, L. Schulze, O. Patel, and D. Hadfield-Menell, “Defending against unforeseen failure modes with latent adversarial training,” arXiv preprint arXiv:2403.05030, 2024. [14] Y. Zhang, X. Chen, J. Jia, Y. Zhang, C. Fan, J. Liu, M. Hong, K. Ding, and S. Liu, “Defensive unlearning with adversarial training for robust concept erasure in diffusion models,” vol. 37, 2024, pp. 36 748–36 776. [15] X. Yue, N. Mou, Q. Wang, and L. Zhao, “Revisiting adversarial training under long-tailed distributions,” in CVPR, 2024, pp. 24 492–24 501. [16] F. Fang, Y. Bai, S. Ni, M. Yang, X. Chen, and R. Xu, “Enhancing noise robustness of retrieval-augmented language models with adaptive adversarial training,” arXiv preprint arXiv:2405.20978, 2024. [17] K. Tang, T. Lou, W. Peng, N. Chen, Y. Shi, and W. Wang, “Effective single-step adversarial training with energy-based models,” IEEE Transactions on Emerging Topics in Computational Intelligence, 2024.
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
[18] Y. Mo, D. Wu, Y. Wang, Y. Guo, and Y. Wang, “When adversarial training meets vision transformers: Recipes from training to architecture,” vol. 35, 2022, pp. 18 599–18 611. [19] X. Jia, Y. Zhang, B. Wu, K. Ma, J. Wang, and X. Cao, “Las-at: adversarial training with learnable attack strategy,” in CVPR, 2022, pp. 13 398–13 408. [20] B. Wu, J. Gu, Z. Li, D. Cai, X. He, and W. Liu, “Towards efficient adversarial training on vision transformers,” in ECCV. Springer, 2022, pp. 307–325. [21] J. Xiao, Y. Fan, R. Sun, J. Wang, and Z.-Q. Luo, “Stability analysis and generalization bounds of adversarial training,” vol. 35, 2022, pp. 15 446–15 459. [22] G. Jin, X. Yi, W. Huang, S. Schewe, and X. Huang, “Enhancing adversarial training with second-order statistics of weights,” in CVPR, 2022, pp. 15 273–15 283. [23] L. Guzman-Nateras, M. Van Nguyen, and T. Nguyen, “Cross-lingual event detection via optimized adversarial training,” in Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, 2022, pp. 5588–5599. [24] A. Madry, A. Makelov, L. Schmidt, D. Tsipras, and A. Vladu, “Towards deep learning models resistant to adversarial attacks,” in ICLR, 2018. [25] C. Pan, Q. Li, and X. Yao, “Adversarial initialization with universal adversarial perturbation: A new approach to fast adversarial training,” in AAAI, vol. 38, no. 19, 2024, pp. 21 501–21 509. [26] H. Kim, W. Lee, and J. Lee, “Understanding catastrophic overfitting in single-step adversarial training,” in AAAI, vol. 35, no. 9, 2021, pp. 8119–8127. [27] K. J. Z. Wong E, Rice L, “Fast is better than free: Revisiting adversarial training,” in ICLR, 2020. [28] T. Li, Y. Wu, S. Chen, K. Fang, and X. Huang, “Subspace adversarial training,” in CVPR, 2022, pp. 13 409–13 418. [29] Z. Huang, Y. Fan, C. Liu, W. Zhang, Y. Zhang, M. Salzmann, S. Süsstrunk, and J. Wang, “Fast adversarial training with adaptive step size,” IEEE TIP, vol. 32, pp. 6102–6114, 2023. [30] G. Y. Park and S. W. Lee, “Reliably fast adversarial training via latent adversarial perturbation,” in ICCV, 2021, pp. 7758–7767. [31] Y. Yang, X. Liu, and K. He, “Fast adversarial training against textual adversarial attacks,” arXiv preprint arXiv:2401.12461, 2024. [32] Z. Huang, Y. Fan, C. Liu, W. Zhang, Y. Zhang, M. Salzmann, S. Süsstrunk, and J. Wang, “Fast adversarial training with adaptive step size,” arXiv preprint arXiv:2206.02417, 2022. [33] M. Zhao, L. Zhang, Y. Kong, and B. Yin, “Fast adversarial training with smooth convergence,” in ICCV, 2023, pp. 4720–4729. [34] Z. Golgooni, M. Saberi, M. Eskandar, and M. H. Rohban, “Zerograd: Costless conscious remedies for catastrophic overfitting in the fgsm adversarial training,” Intelligent Systems with Applications, vol. 19, p. 200258, 2023. [35] F. N. Andriushchenko M, “Understanding and improving fast adversarial training,” 2020, pp. 16 048–16 059. [36] J. Xiaojun, Z. Yong, W. Xingxing, W. Baoyuan, M. Ke, W. Jue, and C. Xiaochun, “Prior-guided adversarial initialization for fast adversarial training,” in ECCV, 2022. [37] Z. He, T. Li, S. Chen, and X. Huang, “Investigating catastrophic overfitting in fast adversarial training: A self-fitting perspective,” in CVPR, 2023, pp. 2313–2320. [38] M. Andriushchenko and N. Flammarion, “Understanding and improving fast adversarial training,” vol. 33, 2020, pp. 16 048–16 059. [39] M. Zhao, L. Zhang, Y. Kong, and B. Yin, “Catastrophic overfitting: A potential blessing in disguise,” in ECCV. Springer, 2024, pp. 293–310. [40] S. Zhang, Y. Pan, Q. Liu, Z. Yan, K.-K. R. Choo, and G. Wang, “Backdoor attacks and defenses targeting multi-domain ai models: A comprehensive review,” ACM Computing Surveys, vol. 57, no. 4, pp. 1–35, 2024. [41] W. Yang, X. Bi, Y. Lin, S. Chen, J. Zhou, and X. Sun, “Watch out for your agents! investigating backdoor threats to llm-based agents,” vol. 37, 2024, pp. 100 938–100 964. [42] S. Liang, M. Zhu, A. Liu, B. Wu, X. Cao, and E.-C. Chang, “Badclip: Dual-embedding guided backdoor attack on multimodal contrastive learning,” in CVPR, 2024, pp. 24 645–24 654. [43] W. Yin, J. Lou, P. Zhou, Y. Xie, D. Feng, Y. Sun, T. Zhang, and L. Sun, “Physical backdoor: Towards temperature-based backdoor attacks in the physical world,” in CVPR, 2024, pp. 12 733–12 743. [44] A. Kurakin, I. Goodfellow, and S. Bengio, “Adversarial machine learning at scale,” in ICLR, 2017. [45] Y. Dong, F. Liao, T. Pang, H. Su, J. Zhu, X. Hu, and J. Li, “Boosting adversarial attacks with momentum,” in CVPR, 2018, pp. 9185–9193.
11
[46] H. Kuang, H. Liu, X. Lin, and R. Ji, “Defense against adversarial attacks using topology aligning adversarial training,” IEEE TIFS, vol. 19, pp. 3659–3673, 2024. [47] Z. Wang, X. Li, H. Zhu, and C. Xie, “Revisiting adversarial training at scale,” in CVPR, 2024, pp. 24 675–24 685. [48] I. J. Goodfellow, J. Shlens, and C. Szegedy, “Explaining and harnessing adversarial examples,” in ICLR, 2015. [49] Y. Zhang, G. Zhang, P. Khanduri, M. Hong, S. Chang, and S. Liu, “Revisiting and advancing fast adversarial training through the lens of bi-level optimization,” in ICML. PMLR, 2022, pp. 26 693–26 712. [50] X. Jia, Y. Zhang, B. Wu, J. Wang, and X. Cao, “Boosting fast adversarial training with learnable adversarial initialization,” IEEE TIP, vol. 31, pp. 4417–4430, 2022. [51] X. Jia, Y. Zhang, X. Wei, B. Wu, K. Ma, J. Wang, and X. Cao, “Priorguided adversarial initialization for fast adversarial training,” in ECCV. Springer, 2022, pp. 567–584. [52] K. Tong, C. Jiang, J. Gui, and Y. Cao, “Taxonomy driven fast adversarial training,” in AAAI, vol. 38, no. 6, 2024, pp. 5233–5242. [53] X. Jia, Y. Chen, X. Mao, R. Duan, J. Gu, R. Zhang, H. Xue, Y. Liu, and X. Cao, “Revisiting and exploring efficient fast adversarial training via law: Lipschitz regularization and auto weight averaging,” IEEE TIFS, 2024. [54] X. Jia, J. Li, J. Gu, Y. Bai, and X. Cao, “Fast propagation is better: Accelerating single-step adversarial training via sampling subnetworks,” IEEE TIFS, 2024. [55] P. de Jorge Aranda, A. Bibi, R. Volpi, A. Sanyal, P. Torr, G. Rogez, and P. Dokania, “Make some noise: Reliable and efficient single-step adversarial training,” vol. 35, 2022, pp. 12 881–12 893. [56] L. Rice, E. Wong, and Z. Kolter, “Overfitting in adversarially robust deep learning,” in ICML. PMLR, 2020, pp. 8093–8104. [57] X. Jia, Y. Zhang, X. Wei, B. Wu, K. Ma, J. Wang, and X. Cao, “Improving fast adversarial training with prior-guided knowledge,” TPAMI, vol. 46, no. 9, pp. 6367–6383, 2024. [58] M. Zareapoor and P. Shamsolmoali, “Rethinking fast adversarial training: A splitting technique to overcome catastrophic overfitting,” in ECCV. Springer, 2024, pp. 34–51. [59] Z. Wang, H. Wang, C. Tian, and Y. Jin, “Preventing catastrophic overfitting in fast adversarial training: A bi-level optimization perspective,” in ECCV. Springer, 2024, pp. 144–160. [60] L. Runqi, Y. Chaojian, H. Bo, S. Hang, and L. Tongliang, “Revealing the pseudo-robust shortcut dependency,” in ICML. PMLR, 2024, pp. 2663–2672. [61] T. Gu, K. Liu, B. Dolan-Gavitt, and S. Garg, “Badnets: Evaluating backdooring attacks on deep neural networks,” IEEE Access, vol. 7, pp. 47 230–47 244, 2019. [62] Y. Liu, S. Ma, Y. Aafer, W.-C. Lee, J. Zhai, W. Wang, and X. Zhang, “Trojaning attack on neural networks,” in NDSS. Internet Soc, 2018. [63] Y. Li, Y. Li, B. Wu, L. Li, R. He, and S. Lyu, “Invisible backdoor attack with sample-specific triggers,” in ICCV, 2021, pp. 16 463–16 472. [64] Y. Zeng, W. Park, Z. M. Mao, and R. Jia, “Rethinking the backdoor attacks’ triggers: A frequency perspective,” in ICCV, 2021, pp. 16 473– 16 481. [65] W. Chen, X. Xu, X. Wang, Z. Li, and Y. Chen, “Fsba: Invisible backdoor attacks via frequency domain and singular value decomposition,” Expert Systems with Applications, vol. 288, p. 127830, 2025. [66] A. Nguyen and A. Tran, “Wanet–imperceptible warping-based backdoor attack,” arXiv preprint arXiv:2102.10369, 2021. [67] X. Chen, C. Liu, B. Li, K. Lu, and D. Song, “Targeted backdoor attacks on deep learning systems using data poisoning,” arXiv preprint arXiv:1712.05526, 2017. [68] T. Wang, Y. Yao, F. Xu, S. An, H. Tong, and T. Wang, “An invisible black-box backdoor attack through frequency domain,” in ECCV. Springer, 2022, pp. 396–413. [69] K. Doan, Y. Lao, W. Zhao, and P. Li, “Lira: Learnable, imperceptible and robust backdoor attacks,” in ICCV, 2021, pp. 11 966–11 976. [70] C. Li, R. Pang, Z. Xi, T. Du, S. Ji, Y. Yao, and T. Wang, “An embarrassingly simple backdoor attack on self-supervised learning,” in ICCV, 2023, pp. 4367–4378. [71] H. Huang, X. Ma, S. M. Erfani, J. Bailey, and Y. Wang, “Unlearnable examples: Making personal data unexploitable,” in ICLR, 2021. [72] J. Ren, H. Xu, Y. Wan, X. Ma, L. Sun, and J. Tang, “Transferable unlearnable examples,” arXiv preprint arXiv:2210.10114, 2022. [73] J. Zhang, X. Ma, Q. Yi, J. Sang, Y.-G. Jiang, Y. Wang, and C. Xu, “Unlearnable clusters: Towards label-agnostic unlearnable examples,” in CVPR, 2023, pp. 3984–3993.
IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, NO. X, MONTH 20XX
[74] X. Liu, X. Jia, Y. Xun, S. Liang, and X. Cao, “Multimodal unlearnable examples: Protecting data against multimodal contrastive learning,” in ACMM, 2024, pp. 8024–8033. [75] T. Qin, X. Gao, J. Zhao, K. Ye, and C.-Z. Xu, “Learning the unlearnable: Adversarial augmentations suppress unlearnable example attacks,” arXiv preprint arXiv:2303.15127, 2023. [76] S. Fu, F. He, Y. Liu, L. Shen, and D. Tao, “Robust unlearnable examples: Protecting data against adversarial learning,” arXiv preprint arXiv:2203.14533, 2022. [77] J. Ye and X. Wang, “Ungeneralizable examples,” in CVPR, 2024, pp. 11 944–11 953. [78] M. Li, X. Wang, Z. Yu, S. Hu, Z. Zhou, L. Zhang, and L. Y. Zhang, “Detecting and corrupting convolution-based unlearnable examples,” in AAAI, vol. 39, no. 17, 2025, pp. 18 403–18 411. [79] A. Krizhevsky, G. Hinton et al., “Learning multiple layers of features from tiny images,” 2009. [80] K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in CVPR, 2016, pp. 770–778. [81] L. McInnes, J. Healy, and J. Melville, “Umap: Uniform manifold approximation and projection for dimension reduction,” arXiv preprint arXiv:1802.03426, 2018. [82] N. Carlini and D. Wagner, “Towards evaluating the robustness of neural networks,” in S&P, 2017, pp. 39–57. [83] F. Croce and M. Hein, “Reliable evaluation of adversarial robustness with an ensemble of diverse parameter-free attacks,” in ICML, 2020. [84] K. He, X. Zhang, S. Ren, and J. Sun, “Identity mappings in deep residual networks,” in ECCV. Springer, 2016, pp. 630–645. [85] Y. Yu, Q. Zheng, S. Yang, W. Yang, J. Liu, S. Lu, Y.-P. Tan, K.-Y. Lam, and A. Kot, “Unlearnable examples detection via iterative filtering,” in International Conference on Artificial Neural Networks. Springer, 2024, pp. 241–256. [86] H. M. Dolatabadi, S. Erfani, and C. Leckie, “The devil’s advocate: Shattering the illusion of unexploitable data using diffusion models,” in IEEE Conference on Secure and Trustworthy Machine Learning. IEEE, 2024, pp. 358–386. [87] K. Liu, B. Dolan-Gavitt, and S. Garg, “Fine-pruning: Defending against backdooring attacks on deep neural networks,” in International Symposium on Research in Attacks, Intrusions, and Defenses. Springer, 2018, pp. 273–294. [88] A. Tomihari and I. Sato, “Understanding linear probing then fine-tuning language models from ntk perspective,” arXiv preprint arXiv:2405.16747, 2024. [89] M. Rui, Q. Zeyu, S. Li, and C. Minhao, “Towards stable backdoor purification through feature shift tuning,” vol. 36, 2023, pp. 75 286– 75 306. [90] Z. Golgooni, M. Saberi, M. Eskandar, and M. H. Rohban, “Zerograd: Mitigating and explaining catastrophic overfitting in fgsm adversarial training,” 2021, p. arXiv preprint arXiv:2103.15476. [91] G. Sriramanan, S. Addepalli, A. Baburaj et al., “Towards efficient and effective adversarial training,” vol. 34, 2021, pp. 11 821–11 833. [92] Z. Qin, L. Yao, D. Chen, Y. Li, B. Ding, and M. Cheng, “Revisiting personalized federated learning: Robustness against backdoor attacks,” in Proceedings of the 29th ACM SIGKDD Conference on Knowledge Discovery and Data Mining, 2023, pp. 4743–4755. [93] S. Ke, C. Hou, G. Fanti, and S. Oh, “On the convergence of differentially-private fine-tuning: To linearly probe or to fully finetune?” arXiv preprint arXiv:2402.18905, 2024. [94] G. Kuofeng, B. Yang, G. Jindong, Y. Yong, and X. Shu-Tao, “Backdoor defense via adaptively splitting poisoned dataset,” in CVPR, 2023, pp. 4005–4014. [95] L. Yige, L. Xixiang, K. Nodens, L. Lingjuan, L. Bo, and M. Xingjun, “Anti-backdoor learning: Training clean models on poisoned data,” vol. 34, 2021, pp. 14 900–14 912. [96] Y. Chen, W. Haiwei, and Z. Jiantao, “Progressive poisoned data isolation for training-time backdoor defense,” in AAAI, vol. 38, no. 10, 2024, pp. 1319–1327. [97] Z. Shuai, G. Leilei, T. LuuAnh, F. Jie, L. Lingjuan, J. Meihuizi, and W. Jinming, “Defending against weight-poisoning backdoor attacks for parameter-efficient fine-tuning,” arXiv preprint arXiv:2402.12168, 2024. [98] L. Linyang, S. Demin, L. Xiaonan, Z. Jiehang, M. Ruotian, and Q. Xipeng, “Backdoor attacks on pre-trained models by layerwise weight poisoning,” arXiv preprint arXiv:2108.13888, 2021. [99] K. Keita, M. Paul, and N. Graham, “Weight poisoning attacks on pretrained models,” arXiv preprint arXiv:2004.06660, 2020.
12
[100] L. Yige, L. Xixiang, M. Xingjun, K. Nodens, L. Lingjuan, L. Bo, and J. Yu-Gang, “Reconstructive neuron pruning for backdoor defense,” in ICML. PMLR, 2023, pp. 19 837–19 854. [101] W. Dongxian and W. Yisen, “Adversarial neuron pruning purifies backdoored deep models,” vol. 34, 2021, pp. 16 913–16 925. [102] C. Kunbei, C. MdHafizulIslam, Z. Zhenkai, and Y. Fan, “Deepvenom: Persistent dnn backdoors exploiting transient weight perturbations in memories,” in S&P, 2024, pp. 2067–2085. [103] Y. Xiong and C.-J. Hsieh, “Improved adversarial training via learned optimizer,” in ECCV. Springer, 2020, pp. 85–100. [104] X. Zhong and C. Liu, “Sparse-pgd: A unified framework for sparse adversarial perturbations generation,” arXiv preprint arXiv:2405.05075, 2024. [105] P.-C. Chen, B.-H. Kung, and J.-C. Chen, “Class-aware robust adversarial training for object detection,” in CVPR, 2021, pp. 10 420–10 429. [106] L. N. Smith, “Cyclical learning rates for training neural networks,” in WACV. IEEE, 2017, pp. 464–472.