Anti-Backdoor Coreset Selection via Cumulative Entropy
Qi Zhao 1 Christian Wressnegger 1
arXiv:2607.25502v1 [cs.LG] 28 Jul 2026
Abstract
gio & Roli, 2018), e.g., to introduce neural backdoors (Gu et al., 2017; Chen et al., 2017; Liu et al., 2018; Nguyen & Tran, 2020; 2021; Turner et al., 2019; Shafahi et al., 2018; Zhao et al., 2023; Barni et al., 2019; Jha et al., 2023).
Recent training-time defenses against neural backdoors isolate a benign subset from poisoned training data, to learn a backdoor-free model from it. In this paper, we formulate this defense strategy as a coreset selection problem, giving rise to so-called “Anti-Backdoor Coreset Selection.” Since poisonous samples have a) lower prediction uncertainty and are b) less frequent than benign samples, coreset selection naturally focuses more on samples associated with benign functionality than the backdoor functionality. We use the Cumulative Entropy as selection criterion to further facilitate this effect. The metric tracks the learning dynamics of training samples and allowing us to select benign samples with high informativeness for the coreset. Additionally, we unlearn the chosen samples in each epoch to facilitate the separability between benign and poisonous samples. Together, this yields an exceptionally effective training-time defense that constructs a benign coreset to train a backdoor-free model. Unlike prior defenses that compromise natural accuracy and fail against certain attacks, our method mitigates backdooring attacks consistently with a negligible impact on natural performance. The implementation of our method is publicly available at: https://intellisec.de/research/abcs.
A neural backdoor establishes a shortcut from a trigger pattern to a target prediction (Biggio & Roli, 2018; Gu et al., 2017; Chen et al., 2017), which can be established in different ways: So-called dirty-label attacks manipulate the training samples (to add the trigger pattern) and change their labels (to force the target prediction) (Gu et al., 2017; Chen et al., 2017; Liu et al., 2018; Nguyen & Tran, 2020; 2021). Clean-label attacks, in turn, go a step further and do not change the ground-truth label but manipulate samples from the target class to strengthen the connection to the trigger pattern (Turner et al., 2019; Shafahi et al., 2018; Zhao et al., 2023; Barni et al., 2019). Recently, the community has focused on tackling data poisoning at its root, that is, learning a clean, backdoor-free model despite the training dataset being poisoned (Li et al., 2021b; Zhang et al., 2023; Huang et al., 2022a; Gao et al., 2023; Chen et al., 2022; Zhu et al., 2023b; Zhao & Wressnegger, 2025). This setting poses several challenges beyond mitigating the backdoor: First, the natural performance (the accuracy on the primary, aimed for task) should match that of a model trained on a not poisoned dataset. Meanwhile, performance must not decline even if the training data is not poisoned. Second, the defense should not rely on a clean reference dataset as this is tied to manual vetting effort. Finally, a defense should not significantly increase training time. Prior work unfortunately falls short in at least one of these criteria or even fails to mitigate the backdoor at all.
1. Introduction
In this paper, we consider an alternative viewpoint on training-time defenses. Most approaches (Huang et al., 2022a; Gao et al., 2023; Chen et al., 2022) aim for splitting the training dataset in poisoned and benign (not poisoned) data to learn a backdoor-free model on the benign subset. We formulate this strategy as a coreset selection problem (Coleman et al., 2020; Toneva et al., 2019; Hekmatfar & Farahani, 2009; Ducoffe & Precioso, 2018; Sener & Savarese, 2018; Mirzasoleiman et al., 2020; Killamsetty et al., 2021; 2020) and propose a new defense scheme: “Anti-Backdoor Coreset Selection (ABCS)”. Our method follows the intuition that the training data encodes two different
Effectively training deep neural networks requires large amounts of training data (LeCun et al., 2015; Hestness et al., 2017; Sun et al., 2017). However, manually vetting these training datasets is infeasible, so that, in practice, models are frequently learned using data without any security guarantees (Carlini, 2021; Carlini et al., 2024). This circumstance gives rise to data-poisoning attacks (Biggio et al., 2012; Big1 KASTEL Security Research Labs, Karlsruhe Institute of Technology (KIT) . Correspondence to: Qi Zhao <[email protected]>.
Proceedings of the 43 rd International Conference on Machine Learning, Seoul, South Korea. PMLR 306, 2026. Copyright 2026 by the author(s).
1
Anti-Backdoor Coreset Selection via Cumulative Entropy
itself. We thus use Cumulative Entropy (C ENT) as a selection criterion that computes an uncertainty score for each sample by accumulating its entropy over multiple training epochs, capturing both predictive uncertainty and temporal consistency (cf. Figure 1, right). For backdoor mitigation, we then select samples with high cumulative entropy as the clean, informative coreset to use them for training a backdoor-free model.
Cumulative
Intermediate 0.8 Benign
Entropy
0.6
Poisonous
0.4 0.2 0 1
5
10
Epoch
15
20 1
5
10
15
20
Epoch
Contributions. We present a novel interpretation of training-time backdoor defenses as a coreset selection problem, introducing “Anti-backdoor Coreset Selection.” We implement this scheme based on a simple yet powerful criterion, the Cumulative Entropy, designed by us to reliably extract informative samples while ensuring to not misclassify poisonous samples during coreset selection. Extensive experiments across various backdoor attacks demonstrate the robustness of ABCS and the striking improvement over related work. We mitigate all investigated backdoor, achieve natural performance comparable to training on the original (clean) dataset, and also match the training time.
Figure 1. Intermediate and cumulative entropy across training epochs for CIFAR10 poisoned by Blend with using a ResNet18 model. Note, that entropy values of all training samples are rescaled to [0, 1] at each epoch.
tasks (the primary functionality and the backdoor) formed by mixing two datasets of benign and poisonous samples, respectively. By retrieving the coreset of the primary task, we effectively mitigate the backdoor and maintain the natural performance. Moreover, focusing on the benign coreset explicitly does not attempt to retrieve all benign samples, thus, reducing the risk of accidentally including poisonous samples. Meanwhile, learning a model on this coreset is faster than on the full dataset, so that the overall runtime of the defense is comparable to a naive training run.
2. Threat Model We consider backdoors that are introduced via data poisoning, that is, the adversaries manipulate samples from a training dataset D. However, they cannot influence the training process. Naively training on the manipulated/poisoned dataset D̃ learns the “primary task” (e.g., image classification) but also implants the backdoor as a “secondary task.” N More formally, given a training dataset D = {(xi , yi )}i=1 d with N samples xi ∈ R and ground-truth labels yi ∈ {0, 1, . . . C − 1}, where C is the number of classes, the adversary poisons Np samples of D (Np ≪ N ), resulting in a poisoned dataset D̃ = D̃p ∪ Dc composed of a poisoned subset D̃p and a clean subset Dc , with the same size as D, i.e., |D| = |D̃|. The poisoning rate of D̃ thus is ρ = Np/N .
Coreset selection most commonly measures either representativeness (a data-centric criterion) or informativeness (a model-centric criterion). The former assesses the geometric coverage of data samples within the full dataset (Welling, 2009; Sener & Savarese, 2018). However, ensuring the full coverage increases the likelihood of including poisonous samples, e.g., using k-Center Greedy (Hekmatfar & Farahani, 2009) as shown in Figure 9. In contrast, informativeness measures a data sample’s ability to increase model certainty (Huang et al., 2010). Interestingly, poisonous samples typically induce fast and stable convergence of the training loss (Li et al., 2021b; Zhao & Wressnegger, 2025), resulting in high confidence in them. This effect can thus also be seen in the entropy trend across every intermediate training epoch as shown in Figure 1 (left). The model exhibits a rapid uncertainty decrease for poisonous samples so that additional training on them does not increase prediction certainty, indicating their low informativeness to the model.
Defenses using “dataset splitting” (Huang et al., 2022a; Gao et al., 2023; Zhu et al., 2023b; Zhao & Wressnegger, 2025; Chen et al., 2022) separate the already poisoned dataset D̃ into a poisoned set Dbad and a benign set Dbng , and use the latter to train a backdoor-free model. To maintain the natural performance, prior works (Gao et al., 2023; Zhu et al., 2023b; Zhao & Wressnegger, 2025; Chen et al., 2022) aim for Dbng containing as few poisonous samples as possible, ideally Dbng = Dc . We follow a similiar yet different aim: Dbng is intended to cover sufficient information of the clean subset Dc instead of include all clean samples. Consequently, it may be smaller than the clean subset, allowing |Dbng | < |Dc | and thus yielding a selection ratio rse = |Dbng |/|D̃| ≪ 1. Moreover, the selected Dbng should contain almost no poisonous samples, i.e., |Dbng ∩D̃p | ρbng = |Dbng | ≈ 0.0.
Therefore, selecting a coreset of the poisoned training dataset using the informativeness—without even considering the presence of two separate tasks—naturally puts stronger focus on the primary task. We find that such a phenomenon not only exist in uncertainty-based selection criteria (Coleman et al., 2020) but particularly in cumulative metrics over multiple training epochs (Toneva et al., 2019; Paul et al., 2021). Such methods can partially mitigate backdoors during training already, but fail to exclude all poisonous samples due to their inherent randomness of sampling at intermediate epochs and the selection process 2
Anti-Backdoor Coreset Selection via Cumulative Entropy Entropy
Margin
of the Blend attack (Chen et al., 2017).
LeastConf 2.5
80
2
60
1.5
40
1
20
0.5
0 0.3
0.4
0.5 0.3
Selection Ratio
0.4
0.5 0.3
Selection Ratio
Uncertainty-based approaches (Coleman et al., 2020) (Entropy, Margin, and LeastConf) consistently produce subsets, on which model training yields a natural performance comparable to the baseline of training on the clean dataset. At the same time, the portion of poisonous samples in these coresets remains below the default poisoning rate of 5 %, indicating their low informativeness. Similarly, coresets selected by loss-based methods (Forgetting (Toneva et al., 2019), GraNd, and EL2N (Paul et al., 2021) shown in the bottom row) stably reproduce the natural accuracy. Note that, the variance (indicated as error bars) is significantly lower for loss-based criteria that either compute scores across multiple training epochs (Forgetting, EL2N) or uses random initialization (GraNd), taming the randomness associated with single training run.
ρbng [%]
ACC / ASR [%]
100
0 0.5
0.4
Selection Ratio
(a) Uncertainty-based Criteria Forgetting
GraNd
EL2N 2.5
80
2
60
ACC ASR
40
ρbng
ACC (baseline)
20 0 0.3
1.5 1
ρbng [%]
ACC / ASR [%]
100
0.5 0.4
Selection Ratio
0.5 0.3
0.4
0.5 0.3
Selection Ratio
0.4
0 0.5
Only, EL2N seems robust against Blend attack, yielding a moderate attack success rate (ASR) of 2–6 % and reaching high natural accuracy. It computes the l2 norm between the predicted confidence of all classes and the one-hot encoded, ground-truth label. This way, EL2N essentially accumulates the mean squared error over 2 training epochs. Figure 3 Epoch-wise EL2N 1.5 shows the vanilla (cumu1 lative) version of EL2N and its variant of using the 0.5 l2 norm score per epoch. 0 0 5 10 15 20 25 30 35 40 Clearly, the accumulation Epoch is crucial, although a releFigure 3. EL2N and its variant vant portion of the coreset of sampling per epoch. remains poisoned.
Selection Ratio
(b) Loss-based Criteria Figure 2. Evaluation of coreset methods under the Blend attack using ResNet18 on CIFAR10 with ρ = 5 %. Error bars show the value range across five random runs per coreset size.
ρbng [%]
3. Coreset Selection as Backdoor Defense Prior training-time defenses (Zhang et al., 2023; Zhu et al., 2023b) use the training loss for identifying benign and poisonous samples. There, the loss essentially serves as a proxy metric for the prediction confidence—or put differently, the reference model’s (un)certainty about a sample. The model is certain about poisonous samples but uncertain about benign, hard-to-learn samples. Interestingly, for coreset selection, the community has explored loss-based criteria (Toneva et al., 2019; Paul et al., 2021) but also uncertainty measures directly (Coleman et al., 2020).
3.2. Instability of Uncertainty-Based Sampling Uncertainty-based metrics are influenced by the randomness of intermediate sampling towards the end of the training, resulting in a large variance in both ASR and ρbng . Lossbased approaches are more stable, though. Thus, one-step sampling is not reliable for backdoor defense.
In the following, we thus specifically investigate these two categories of selection criteria (Section 3.1) and discuss their stability in the backdoor defense setting (Section 3.2). In Section A, we additionally investigate other selection criteria, such as, geometric-based (Hekmatfar & Farahani, 2009) or decision boundary based (Ducoffe & Precioso, 2018) criteria. Note that we discard balanced selection across classes to lower the likelihood of including poisonous samples from backdoor target class.
Based on the insights gained from EL2N and the smoothing effect observed in Figure 3, we set out to investigate whether accumulation can help to contain the variance of uncertaintybased metrics also. Figure 4 compares sampling per epoch and accumulating scores across epochs for selection criteria Entropy, Margin and Least Confidence (Coleman et al., 2020). For each, one can clearly see that although they occasionally achieve a low poisoning rate ρbng , more often than not the coreset contains a large number of poisonous samples. Interestingly, the issue becomes increasingly severe toward the end of the training, as the uncertainty gap between benign and poisonous samples diminishes significantly, making them less distinguishable.
3.1. Uncertainty-Based and Loss-Based Selection We examine the performance of uncertainty-based and lossbased coreset selection from a poisoned dataset D̃. In line with the intuition on the relation between learning speed (Li et al., 2021b) and a model’s uncertainty, we find that such criteria have a somewhat natural potential to mitigate backdooring attacks as can be seen in Figure 2 using the example 3
Anti-Backdoor Coreset Selection via Cumulative Entropy Entropy
Margin
LeastConf
2
2
2
1.5
1.5
1.5
1 0.5 0
ρbng [%]
ρbng [%]
ρbng [%]
Epoch-wise
1 0.5 0
0
5
Cumulative
1 0.5 0
10 15 20 25 30 35 40
0
5
10 15 20 25 30 35 40
Epoch
0
5
Epoch
10 15 20 25 30 35 40
Epoch
Figure 4. Epoch-wise sampling vs. accumulation for uncertainty-based coreset selection. Selection ratio is 0.4. Each run trains a ResNet18 on CIFAR10 dataset that is poisoned by Blend attack with ρ = 5 %.
Accumulating uncertainty measurements across epochs, in turn, yields a robust and stable coreset selection, and simultaneously enables the elimination of poisonous samples. Unlike the accumulation of Entropy and Margin that yield near-zero ρbng , the accumulation of LeastConf cannot effectively eliminate poisonous samples, leading to a high ρbng .
wardly choose a benign and informative subset. Additionally, it is beneficial to min-max-scale the scores of accumulated entropy across all data samples for each of the Tse ai −amin epochs, normalize(ai ) := amax −amin , to avoid the interference in-between epochs, thus, further facilitating the stability of the overall metric. An empirical evidence in Table 5 demonstrate the benefit of using min-max normalization for the anti-backdoor coreset selection by C ENT.
Backdoor Defense. Next, we investigate the performance of the cumulative versions of the three uncertainty-based criteria when used for backdoor defense verbatim and summarize the results in Figure 5. Despite prior reports that plain Entropy is less performant in sampling an informative coreset (Guo et al., 2022), Cumulative Entropy yields higher natural accuracy than all other criteria. Moreover, the Cumulative Entropy consistently yields the smallest poisoning rate ρbng and lowest ASR across coreset sizes, thus, making it an ideal choice for backdoor defense.
91
Entropy
ρbng [%]
93
Tse 1 X normalize Tse t=1 |
30
15
0.6
0.3
LeastConf 0.4
Selection Ratio
0.5
0 0.3
0.4
Selection Ratio
0.5
0 0.3
0.4
! −pθt (c|xi ) · log (pθt (c|xi ))
,
c=0
{z
Normalized Shannon Entropy H(xi )
}
Our defense, ABCS, thus comprises three phases: First, a warm-up phase to stabilized the initial selection (Section 4.1). Second, the actual coreset selection with three individual steps (Section 4.2), and finally training a clean model from scratch using the selected coreset (Section 4.3).
Margin 89 0.3
C−1 X
where accumulation over Tse epochs and normalization across all sample in each is done as described in Section 3.2. However, plainly using C ENT for backdoor defense is not sufficient for complex backdooring attacks (Nguyen & Tran, 2021; Zeng et al., 2021; Qi et al., 2023).
0.9
45
ASR [%]
ACC [%]
95
Formally, the C ENT criterion to an input sample xi , i.e., C ENT (xi ), can be calculated via:
0.5
Selection Ratio
Figure 5. Training baseline model ResNet18 from scratch on coresets of CIFAR10 selected from a poisoned dataset with Blend attack by using different uncertainty criteria in the accumulation.
4.1. Warm-Up Each backdooring attack has a varying learning difficulty (Zhao & Wressnegger, 2024), so that we first bootstrap training in a warm-up phase of Twa epochs, yielding an initial model. We use this model to determine the size s of the coreset to be selected as Dbng , which in turn specifies the selection ratio rse for a poisoned dataset D̃ as rse = s/|D̃|. Many approaches (Toneva et al., 2019; Qin et al., 2024) use the mean of the importance score as a threshold to determine the coreset size. However, wrongly predicted samples are considered as not learned by the model despite the uncertainty (Toneva et al., Given the number of correct P2019). N predictions Ncr = i=1 1arg maxc (fθ (xi )c )=yi , hence, the
4. Anti-Backdoor Coreset Selection via Cumulative Entropy Based on the observations made in the previous section, we propose a training-time defense, ABCS, using coreset selection via the Cumulative Entropy (C ENT) criterion. C ENT is a simple yet effective coreset selection criterion, that robustly sorts data samples according to their informative value. It assigns low and high uncertainty to poisonous and benign samples, respectively, allowing to straightfor4
Anti-Backdoor Coreset Selection via Cumulative Entropy
cut-off threshold τ of C ENT (xi ) is defined as: Twa 1 X τ= Twa t=1
and thus s =
N 1 X H (xi ) · 1arg max (fθ (xi )c )=yi c t Ncr i=1 N X
expressed as: !
zc =
ε C
if c = y otherwise.
A large ε lowers the targeted confidence of the sample’s label y, implicitly enlarging the prediction uncertainty. An ablation study of smoothing factor ε is shown in Section 5.3.
1C ENT(xi )>τ , which determines the final
Moreover, we include l2 regularization based on the current and the previous models’ weights, θ and θ∗ , respectively, to restrict the overall change/impact on the learned task. Hence, we use the following loss function to unlearn samples from Dul :
i=1
coreset size for the next selection stage, i.e., |Dbng | = s. 4.2. Selecting the Coreset After the warm-up stage, we resume to train the model for Tse epochs. Each epoch serves three purposes: First, learning the entropy distribution of all samples in D̃; Second, disentangling benign and uncertain samples from poisoned ones; Third, calculating the normalized entropy of individual samples and accumulate it with results from the previous epoch, yielding an updated C ENT at the intermediate epoch. The coreset selection procedure within each individual epoch is thus implemented in three steps:
label smoothing
l regularization
2 z }| { zX }| { ∗ 2 LUL (x, z, θ) := γ · LCE (x, z, θ) + (θj − θj ) ,
j
where γ is a regularization parameter. To avoid model collapse, we additionally lower the learning rate lr to 10 % of the initial value. Different from the unlearning by gradient ascending (Li et al., 2021b), which leads to an extraordinarily large learning divergence, label smoothing stabilizes unlearning because the “smoothed label” still specifies a fixed target where the ground-truth label has the highest score and, thus, unlearning will not push the prediction to an extreme case. Additionally, label smoothing implicitly increases the entropy values of samples in Dul .
Step-1. Learning the entire dataset D̃ . Backdooring attacks do not influence the natural performance and implicitly preserve a data distribution comparable to training on the clean dataset. Thus, we first train on the entire poisoned dataset D̃ for one epoch to obtain a model θ∗ , which outputs an intermediate entropy distribution of the entire dataset.
Step-3. Updating the cumulative entropy. After Step-2, we compute the normalized entropy H(xi ) for every sample xi . For t = 1, C ENT(xi ) is directly identical to H(xi ). For t > 1, the newly computed H(xi ) is accumulated to the cumulative entropy values obtained from previous epochs t = {1, . . . , t − 1}, yielding an updated C ENT(xi ).
Step-2. Unlearning uncertain samples. Cumulative Entropy alone does not guarantee a coreset free of poisonous samples of various backdoors, particularly for stealthier ones, e.g., WaNet (cf. Figure 7). We thus unlearn uncertain samples to enlarge the discrepancy in uncertainty between poisonous and benign samples, thereby reducing the overlap between the final coreset Dbng and the poisonous samples D̃p . To do so, we first measure the entropy of each sample xi , i.e., H(xi ), at intermediate epoch t and select samples larger than the average as the intermediate unlearning subset Dul =) ( (xi , yi ) | H (xi ) >
( 1 − ε + Cε
Across Tse epochs, C ENT(xi ) progressively aligns with the data’s underlying distribution of informativeness. Meanwhile, as the model has high confidence on the poisonous samples (yielding low uncertainty), the discrepancy between poisonous samples and informative clean samples is significantly enlarged in the view of Cumulative Entropy. After running the last selection epoch (i.e., t = Tse ), we sort all training samples by their C ENT scores and select s samples (the coreset size determined in Section 4.1) with highest C ENT values, forming the final clean coreset Dbng .
N 1 X H (xi ) · 1arg max (fθ (xi )c )=yi c t Ncr i=1
Additionally, we use label smoothing (Müller et al., 2019) to counter instability during unlearning (Di et al., 2026; Fan et al., 2024; Neel et al., 2020) and boost the forgetting of the unlearning set. The label smoothing pushes the prediction still towards ground-truth but in a more smooth distribution, implicitly enlarging the uncertainty of hard-tolearn, benign samples, making them distant to poisonous samples in C ENT distribution. Given a smoothing factor ε, the smoothed target output z (often also referred to as “smoothed/soft label” in related work) of sample x is thus
4.3. Final Training Finally, we train a model on the final coreset Dbng from scratch for PTcln epochs using regular cross-entropy loss: arg minθ (x,y)∈Dbng LCE (x, y, θ). As the coreset Dbng is significantly smaller than the full training dataset D̃, training the final model is fast. Thus, ABCS’s overall training time (including coreset selection and training the model) is comparable to naive (unprotected) learning. 5
Anti-Backdoor Coreset Selection via Cumulative Entropy Table 1. Comparing ABCS to prior training-time defenses. Each defense experiment contains the averaged value and the standard deviation in [%] across five random runs. The best results across all defenses are highlighted as boldface. The defense failure (i.e., ASR > 50 %) is shown as orange boldface. We discard ASR for No-Defense as all attacks reach a ASR in the range of 95 % to 100 %. (a) CIFAR10 Attack
No-Defense
ABL
DBD
ACC (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
BadNets Blend CLB IAB WaNet ISSBA LF A-Blend
94.53± 0.24 94.49± 0.09 94.46± 0.13 94.53± 0.06 94.02± 0.19 94.51± 0.22 94.70± 0.08 94.02± 0.06
92.96±0.35 88.29±2.63 85.75±1.67 92.59±0.48 88.64±2.41 87.54±1.77 89.38±3.38 83.87±3.81
0.36± 0.06 34.98±38.17 1.46± 1.36 3.10± 0.89 59.99±48.83 0.98± 0.32 41.42±51.18 62.35± 9.65
99.03± 0.18 78.05±20.26 94.92± 0.45 97.40± 0.36 65.50±25.44 96.02± 0.77 75.80±25.36 61.74± 6.17
92.19±1.69 93.12±0.50 91.85±1.15 92.34±4.99 91.97±0.62 92.53±0.43 92.86±0.40 92.78±0.73
2.24± 0.22 7.25± 3.20 3.85± 6.24 99.82± 0.13 2.42± 0.29 10.91± 1.13 96.69± 1.93 6.06± 1.59
97.71± 0.88 94.33± 1.79 96.77± 2.66 48.47± 1.85 95.77± 0.44 93.55± 0.42 49.94± 1.11 94.34± 0.73
88.63±1.53 89.41±1.84 89.11±1.84 89.84±0.94 90.05±2.06 93.01±0.27 90.85±0.59 92.26±0.80
20.15±42.28 47.21±27.66 14.71±32.90 8.40± 2.82 48.19±42.67 83.08±34.68 69.37±27.15 60.36± 7.91
86.98±21.80 72.49±13.53 89.97±15.60 93.38± 1.81 71.92±21.54 57.71±17.31 62.55±13.59 66.93± 3.57
Average WorstCase
94.41 94.02
88.63 83.87
25.58 62.35
83.56 61.74
92.45 91.85
28.66 99.82
83.86 48.47
90.40 88.63
43.94 83.08
75.24 57.71
Attack
No-Defense ACC (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
BadNets Blend CLB IAB WaNet ISSBA LF A-Blend
94.53± 0.24 94.49± 0.09 94.46± 0.13 94.53± 0.06 94.02± 0.19 94.51± 0.22 94.70± 0.08 94.02± 0.06
94.64±0.16 94.07±0.37 94.23±0.11 94.14±0.72 94.33±0.13 93.87±0.77 93.88±1.28 92.16±0.78
0.91± 0.13 0.94± 0.06 1.06± 0.34 13.85±28.85 1.10± 0.39 0.64± 0.63 57.34±41.79 96.03± 5.43
99.53± 0.05 97.96± 0.18 99.36± 0.17 92.79±14.78 97.45± 0.19 99.36± 0.62 70.15±21.09 50.03± 2.01
93.92±0.31 94.05±0.16 93.86±0.22 93.54±0.53 93.51±0.20 93.35±0.48 93.44±0.57 93.63±0.30
1.08± 0.40 0.92± 0.32 0.45± 0.08 2.01± 0.74 1.79± 0.93 1.26± 0.26 39.63±50.98 75.42± 3.90
99.16± 0.23 97.96± 0.20 99.47± 0.09 98.42± 0.49 96.85± 0.55 98.79± 0.36 78.75±25.19 60.08± 1.85
94.38±0.23 94.62±0.36 94.51±0.23 94.30±0.16 94.13±0.22 94.66±0.41 94.06±0.30 94.37±0.27
1.00± 0.17 0.77± 0.08 0.88± 0.10 1.22± 0.37 0.89± 0.16 1.22± 0.09 3.04± 0.55 5.71± 1.59
99.42± 0.19 98.22± 0.07 99.54± 0.06 99.19± 0.19 97.53± 0.13 99.33± 0.15 97.33± 0.17 95.12± 0.77
Average WorstCase
94.41 94.02
93.92 92.16
21.48 96.03
88.33 50.03
93.66 93.35
15.32 75.42
91.18 60.08
94.38 94.06
1.84 5.71
98.21 95.12
V&B
CBD
Harvey
ABCS (Ours)
(b) Tiny-ImageNet Attack
No-Defense
ABL
DBD
CBD
ACC (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
BadNets Blend IAB WaNet ISSBA LF A-Blend
61.77± 0.17 61.44± 0.37 61.98± 0.10 60.71± 0.02 61.29± 0.12 61.40± 0.11 60.97± 0.06
47.14±1.22 38.40±1.16 44.73±0.54 41.88±5.43 35.38±0.73 38.59±0.83 39.20±0.97
0.00± 0.00 0.00± 0.00 0.00± 0.00 94.82± 4.78 0.07± 0.10 0.00± 0.00 76.07±13.73
92.66± 0.61 88.30± 0.58 91.37± 0.27 42.66± 3.51 86.99± 0.42 87.51± 0.42 49.56± 7.04
51.31±1.36 51.68±0.39 50.82±0.19 50.99±0.13 51.04±0.27 51.00±0.14 50.64±0.21
98.93± 1.69 99.87± 0.21 99.86± 0.19 99.41± 1.02 96.81± 1.39 96.04± 0.75 98.69± 1.23
45.29± 1.50 45.12± 0.20 44.49± 0.08 45.26± 0.15 46.45± 0.57 45.70± 0.38 44.84± 0.11
48.93±0.68 49.11±0.58 49.40±0.69 47.89±0.78 44.90±0.16 48.74±0.10 48.34±1.15
0.11± 0.04 16.51±22.55 0.31± 0.06 60.81±18.96 0.50± 0.27 7.95± 5.18 88.32±11.91
93.50± 0.32 85.40±11.52 93.55± 0.36 62.68± 9.61 91.54± 0.07 88.61± 2.64 48.21± 5.71
Average WorstCase
61.37 60.71
40.76 35.38
24.42 94.82
77.01 42.66
51.07 50.64
98.52 99.87
45.31 44.49
48.19 44.90
24.93 88.32
80.50 48.21
Attack
No-Defense ACC (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
BadNets Blend IAB WaNet ISSBA LF A-Blend
61.77± 0.17 61.44± 0.37 61.98± 0.10 60.71± 0.02 61.29± 0.12 61.40± 0.11 60.97± 0.06
61.35±0.16 61.84±0.28 61.64±0.24 60.79±0.32 58.02±0.68 61.91±0.63 56.85±0.12
0.51± 0.15 2.41± 2.35 0.15± 0.04 1.65± 1.12 14.34± 6.26 17.87±13.74 73.13±16.42
99.51± 0.08 98.61± 1.17 99.75± 0.11 98.62± 0.56 91.18± 2.82 89.98± 6.87 59.85± 8.26
61.86±0.56 61.28±0.58 61.26±0.11 61.08±0.48 57.95±0.75 61.54±0.99 60.93±0.60
0.02± 0.02 0.18± 0.13 3.55± 2.07 1.91± 2.06 0.39± 0.21 79.33±17.97 74.26±12.08
99.88± 0.15 99.57± 0.21 97.86± 1.09 98.53± 1.05 98.12± 0.30 59.21± 8.76 61.23± 6.10
61.24±0.24 61.11±0.07 61.68±0.44 60.54±0.78 58.01±0.37 61.44±0.35 60.03±0.50
0.11± 0.04 0.18± 0.05 0.22± 0.04 0.51± 0.22 0.25± 0.05 0.18± 0.15 1.15± 0.43
99.65± 0.10 99.56± 0.06 99.73± 0.21 99.06± 0.24 98.22± 0.16 98.77± 0.02 97.43± 0.25
Average WorstCase
61.37 60.71
60.34 56.85
15.72 73.13
91.07 59.85
60.84 57.95
22.80 79.33
87.77 59.21
60.58 58.01
0.37 1.15
98.92 97.43
V&B
Harvey
5. Evaluation
ABCS (Ours)
Tiny-ImageNet. In the appendix, we provide details of experimental setups (cf. Section C), extend the evaluation (cf. Section D) to experiments for GTSRB, ablation studies on ABCS’ hyper-settings, the robustness analysis across diverse attack settings, experiments for all-to-all attacks and for text classification task, and comparison to one defense using reference clean data, test ABCS against four adaptive attacks (cf. Section E), and finally analyze the coverage of selected coresets to full datasets (cf. Section F).
We conduct extensive experiments across two small-scale datasets CIFAR10 (Krizhevsky et al., 2008) and GTSRB (Stallkamp et al., 2012) with ResNet18 (He et al., 2016) , and one large-scale dataset Tiny-ImageNet (Le & Yang, 2015) with ResNet34 (He et al., 2016). To account for randomness, each experiment is repeated with five random seeds on CIFAR10 and GTSRB, and three seeds on 6
Anti-Backdoor Coreset Selection via Cumulative Entropy Table 2. Comparing ABCS to training-time defenses on clean (not poisoned) datasets. Dataset CIFAR10 GTSRB Tiny-ImageNet
Naive
ABL
DBD
CBD
V&B
Harvey
ABCS
94.64±0.17 98.36±0.21 62.22±0.45
83.42±3.35 91.79±2.11 37.46±1.86
92.32±2.23 93.38±0.37 51.81±0.67
92.90±0.41 80.54±9.73 51.00±0.64
94.16±0.41 92.50±2.44 58.87±0.76
93.80±0.29 97.95±0.22 60.86±0.45
94.77±0.11 98.12±0.28 61.19±0.26
Attacks and Defenses. We consider eight dataset poisoning backdoors: BadNets (Gu et al., 2017), Blend (Chen et al., 2017), CLB (Turner et al., 2019), IAB (Nguyen & Tran, 2020), WaNet (Nguyen & Tran, 2021), ISSBA (Li et al., 2021a), Low Frequency (LF; Zeng et al., 2021), and Adaptive Blend (A-Blend; Qi et al., 2023). For a fair comparison, we follow the default implementation of each attack. We choose class 0 as target label and use a 5 % poisoning rate, except for CLB, which uses ρ = 50 % on the target class. For Tiny-ImageNet, CLB is excluded due to its ineffectiveness. In the evaluation, we compare ABCS to six state-of-the-art training-time defenses that require no prior clean reference data: ABL (Li et al., 2021b), DBD (Huang et al., 2022a), CBD (Zhang et al., 2023), V&B (Zhu et al., 2023b), and Harvey (Zhao & Wressnegger, 2025). We evaluate each defense using its default settings and hyperparameters. For ABCS, we set Twa = 10 and Tse = 40. The unlearning step during coreset selection uses ε = 0.9, and γ = 0.1 and 0.01 for small-scale and for large-scale datasets. Moreover, we use ADAM optimizer with learning rate 0.001 during warm-up and coreset selection.
selecting informative coresets that exclude poisonous data. Adaptability to clean datasets. Prior defenses primarily focus on mitigating backdoors. However, they often suffer a drop in natural accuracy compared to naive training, highlighting a key limitation in their adaptability to clean (not poisoned) datasets (cf. Table 2). Differently, ABCS leverages the high informativeness of coresets, enabling full adaptability when training on clean datasets. Time consumption. ABCS benefits from the small coreset size, resulting in faster training comparable to native training on Tiny-ImageNet (cf. Table 3), whereas other methods require more time, in particular for DBD defense. Table 3. Time consumption of native training and all defenses. Dataset CIFAR10 GTSRB Tiny-ImageNet
Naive (h) 1.47 0.86 9.14
ABL
DBD
CBD
V&B Harvey ABCS
×1.08 × 12.09 ×1.10 ×2.45 ×1.25 × 38.64 ×1.12 ×5.71 ×1.17 × 13.92 ×1.20 ×5.58
×1.82 ×1.76 ×1.82
×0.97 ×0.75 ×1.25
5.2. Investigation of Coreset Properties Evaluation metrics. We present the defensive performance with three metrics: Natural Accuracy (ACC), Attack Success Rate (ASR), and Defense Effectiveness Rate (DER) (Zhu et al., 2023a). Formally, DER is defined as: DER = [max (0, ∆ASR) − max (0, ∆ACC) + 1] /2, where ∆ASR is the drop of ASR and ∆ACC is the drop of ACC. The optimal defense achieves DER equal 100 %.
We particularly look at the overall selection ratio of the coreset and its poisoning ratio, and finally inspect how well the class disribution is preserved. Poisoning ratio ρbng and selection ratio rse . Unlike prior defenses (Huang et al., 2022a; Gao et al., 2023; Zhao & Wressnegger, 2025) that aim for precise dataset splitting, our method selects coresets smaller than the entire clean dataset (i.e., selection ratio rse ≪ 100 %). Table 4 summarizes the selected coresets for CIFAR10, where the selection ratio rse is around 55 % and the poisoning rate ρbng Table 4. Coresets of CIFAR10 is significantly reduced to under different attack scenarios. near 0 %. This effectively Attack rse [%] ρbng [%] prevent backdoor implantBadNets 54.52±0.58 0.00± 0.00 ing in the model training, Blend 53.85±0.58 0.00± 0.00 and the strong natural perCLB 54.53±0.45 0.00± 0.00 IAB 54.90±0.83 0.01± 0.01 formance shown in Table 1 WaNet 57.83±0.24 0.20± 0.09 underscores the high inforISSBA 54.04±1.05 0.00± 0.00 LF 54.89±0.77 0.19± 0.05 mativeness of the selected A-Blend 56.03±0.51 0.54± 0.15 coresets by ABCS.
5.1. Performance of Training-Time Defenses Table 1 shows the results on CIFAR10 and Tiny-ImageNet. Across five random trials, prior defenses struggle to balance natural accuracy and the reduction of ASR. More notably, each is bypassed by one or multiple attacks. Specifically, ABL and CBD exhibit high variance in ASR while simultaneously compromising natural accuracy. DBD shows stability against the randomness but fails under several attacks. Advanced methods like V&B and Harvey maintain high natural accuracy but remain vulnerable to strong attacks such as LF and A-Blend. In contrast, ABCS shows consistent robustness across attacks and random seeds, while preserving ACC comparable to No-Defense. Although ABCS does not always achieve the best results on individual poisoned datasets, its strong ASR reduction and ACC retention consistently yield a high DER—above 95 % on CIFAR10 and 97 % on Tiny-ImageNet—demonstrating its effectiveness in
Preservation of class-wise distribution. Figure 6 shows the class-wise distribution of natural accuracy and data composition for CIFAR10. After training on coresets under 7
Anti-Backdoor Coreset Selection via Cumulative Entropy Table 5. Impact of min-max normalization in C ENT criterion. Attack
rse = 0.3
Min-Max Norm.
rse = 0.4
ASR
ρbng
ACC
ASR
ρbng
ACC
ASR
ρbng
Blend
✗ ✓
90.82±0.45 90.89±0.20
1.35±0.34 1.38±0.05
0.06±0.03 0.06±0.03¸
93.89±0.24 94.12±0.17
1.12±0.31 1.00±0.25
0.09±0.03 0.10±0.03
94.63±0.37 94.71±0.18
1.41±1.07 1.22±0.20
0.21±0.07 0.19±0.03
WaNet
✗ ✓
88.12±0.47 88.72±0.44
3.36±0.33 3.05±0.44
1.50±0.33 1.41±0.27
91.55±0.93 92.63±0.34
1.55±0.49 1.18±0.13
0.62±0.19 0.50±0.16
93.94±0.24 94.14±0.26
6.86±5.29 1.90±0.76
1.36±0.46 0.79±0.27
accuracy remains unaffected, applying warm-up reduces the poisoning rate ρbng of coresets, which is further lowered by executing the additional unlearning step.
100 95 90 85
Blend WaNet Clean D
BadNets IAB
6
8
95
90
90
60
1
2
3
4
5
7
9
ACC [%]
0
Class
(a) Class-wise accuracy
85
0.2
Selection Ratio
ASR [%]
80
6
C ENT + Warm-Up
ρbng [%]
ACC [%]
rse = 0.5
ACC
30
4
C ENT + Warm-Up + Unlearn
2
C ENT only 80 0.3
0.15
0.4
Selection Ratio
0.5
0 0.3
0.4
Selection Ratio
0.5
0 0.3
0.4
0.5
Selection Ratio
0.1
Figure 7. Comparing the impact of using warm-up and unlearning in ABCS. Baseline model ResNet18 is trained from scratch on selected coresets of CIFAR10 under WaNet attack.
0.05 0.0 0
1
2
3
4
5
6
7
8
9
Class
(b) Class-wise data selection ratio
ACC [%]
various attacks (left), the class-wise accuracy has a distribution similar to that obtained from the clean dataset D. Also the class-wise composition of the coresets remain consistent across attacks (right). ABCS consistently extracts coresets that retain informativeness regardless of data poisoning and achieves high coverage of the full clean dataset (cf. Section F), thereby robustly preserving natural performance. 5.3. Ablation Study Finally, we conduct an ablation study on different building blocks of our method, such as the warm-up and unlearning, the label smoothing factor ε, and C ENT’s normalization. Warm-up training and unlearning. Using the Cumulative Entropy criterion can effectively mitigate strong neural backdoor attacks, such as the Blend attack (cf. Figure 5). The incorporation of warm-up and unlearning is particularly beneficial for stealthier attacks, where the convergence to the backdoor occurs more slowly. Figure 7 showcases coresets with fixed selection ratios against the WaNet attack, where the coreset selection by Cumulative Entropy alone cannot fully exclude poisonous samples. While the natural
ASR [%]
Label smoothing factor ε. The value ε is proportional to the strength of label smoothing. Thus, increasing ε will amplify the unlearning effect, and vice versa. Figure 8 below investigates the impact of different ε values on ABCS’s defense. Varying the smoothing factor does not affect the natural performance represented by the selected coreset, making ACC consistently high. However, an overly small ε indeed weakens the unlearning effect, finally resulting in an insufficient removal of poisoned samples, particularly under the 40 95 WaNet attack. As the result, residual poisoned ACC (Blend) 30 90 ACC (WaNet) samples in the coreset 20 85 ASR (Blend) lead to a higher ASR. ASR (WaNet) 10 80 Overall, the effective75 0 ness of ABCS depends .3 .4 .5 .6 .7 .8 .9 .99 on choosing ε within a Label Smoothing Factor ε reasonable range, rather Figure 8. Ablation study on label than relying on a specific smoothing factor ε. fixed value.
Figure 6. Class-wise distribution on ResNet18 with CIFAR10 across different attacks over five runs. “Clean D” refers to using the full dataset for both (a) model training and (b) coreset selection.
Min-max normalization in C ENT. The min-max normalization balances the entropy scale across epochs from early to late training time. Experiments on CIFAR10 (cf. Table 5) show that, despite low ASR and poisoning rates of coresets, using normalization improves ACC and enhances the elimination of poisonous samples particularly for WaNet. 8
Anti-Backdoor Coreset Selection via Cumulative Entropy
6. Conclusion
rise of adversarial machine learning. Pattern Recognition, 84:317–331, 2018.
Mitigating data-poisoning backdoors in the training procedure is challenging due to the high risk of misidentifying poisonous samples. Our approach, ABCS, leverages the defense nature of coreset selection and introduce the Cumulative Entropy criterion to extract an informative coreset while effectively excluding poisonous data. Training on this benign and informative coreset reproduces natural performance comparable to training on a fully clean dataset. ABCS demonstrates robustness across a wide range of attacks and diverse datasets. Notably, ABCS adapts well even in clean settings and incurs a computational cost similar to the naive training, reinforcing its practicality and giving rise to a promising novel paradigm in training-time backdoor defense: Anti-Backdoor Coreset Selection.
Biggio, B., Nelson, B., and Laskov, P. Poisoning attacks against support vector machines. In Proc. of the International Conference on Machine Learning (ICML), 2012. Borgwardt, K. M., Gretton, A., Rasch, M. J., Kriegel, H.-P., Schölkopf, B., and Smola, A. J. Integrating structured biological data by kernel maximum mean discrepancy. Bioinformatics, 2006. doi: 10.1093/bioinformatics/btl242. Carlini, N. Poisoning the unlabeled dataset of SemiSupervised learning. In Proc. of the USENIX Security Symposium, 2021. Carlini, N., Jagielski, M., Choquette-Choo, C. A., Paleka, D., Pearce, W., Anderson, H., Terzis, A., Thomas, K., and Tramèr, F. Poisoning web-scale training datasets is practical. In Proc. of the IEEE Symposium on Security and Privacy, 2024.
Limitations. While ABCS demonstrates strong empirical performance, it lacks a theoretical foundation. Future work could investigate latent space separation between coreset and poisonous samples, focusing on data informativeness and backdoor effectiveness. Compared to the optimal coreset for clean datasets, ABCS results in a larger coreset size, suggesting further exploration into optimal selection from poisoned datasets. Its defense against sophisticated attacks, e.g., A-Blend, remains suboptimal, as it does not reduce ASR to the minimum, indicating the need for improvement.
Chen, W., Wu, B., and Wang, H. Effective backdoor defense by exploiting sensitivity of poisoned samples. In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2022. Chen, X., Liu, C., Li, B., Lu, K., and Song, D. Targeted backdoor attacks on deep learning systems using data poisoning. CoRR, abs/1712.05526, 2017.
Acknowledgments
Cheng, P., Wu, Z., Du, W., Zhao, H., Lu, W., and Liu, G. Backdoor attacks and countermeasures in natural language processing models: A comprehensive security review. IEEE Transactions on Neural Networks and Learning Systems, 2025. doi: 10.1109/TNNLS.2025.3540303.
We gratefully acknowledge funding by the Helmholtz Association (HGF) within topic “46.23 Engineering Secure Systems.”
Impact Statement
Clark, P. J. and Evans, F. C. Distance to nearest neighbor as a measure of spatial relationships in populations. Ecology, 35(4):445–453, 1954.
Deep neural networks are widely applied across numerous domains, which makes evaluating their security in realworld settings essential. In this work, we propose AntiBackdoor Coreset Selection, a simple yet effective defense scheme for training a backdoor-free model from a poisoned dataset. Our approach is developed strictly from the perspective of a defender, as defined in the threat model. Therefore, this research does not introduce any ethical concerns or create additional security risks.
Coleman, C., Yeh, C., Mussmann, S., Mirzasoleiman, B., Bailis, P., Liang, P., Leskovec, J., and Zaharia, M. Selection via proxy: Efficient data selection for deep learning. In Proc. of the International Conference on Learning Representations (ICLR), 2020. Devlin, J., Chang, M.-W., Lee, K., and Toutanova, K. Bert: Pre-training of deep bidirectional transformers for language understanding. In Proc. of the Annual Meeting of the Association for Computational Linguistics, 2019.
References Barni, M., Kallas, K., and Tondi, B. A new backdoor attack in cnns by training set corruption without label poisoning. In Proc. of the IEEE International Conference on Image Processing (ICIP), 2019.
Di, Z., Zhu, Z., Jia, J., Liu, J., Takhirov, Z., Jiang, B., Yao, Y., Liu, S., and Liu, Y. Label smoothing improves machine unlearning. In Proc. of the International Conference on Learning Representations (ICLR), 2026.
Biggio, B. and Roli, F. Wild patterns: Ten years after the 9
Anti-Backdoor Coreset Selection via Cumulative Entropy
Ducoffe, M. and Precioso, F. Adversarial active learning for deep networks: a margin based approach. In Proc. of the AAAI Conference on Artificial Intelligence (AAAI), 2018.
International Conference on Learning Representations (ICLR), 2022a. Huang, S.-j., Jin, R., and Zhou, Z.-H. Active learning by querying informative and representative examples. In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2010.
Fan, C., Liu, J., Zhang, Y., Wong, E., Wei, D., and Liu, S. SalUn: Empowering machine unlearning via gradientbased weight saliency in both image classification and generation. In Proc. of the International Conference on Learning Representations (ICLR), 2024.
Huang, Y., Bai, B., Zhao, S., Bai, K., and Wang, F. Uncertainty-aware learning against label noise on imbalanced datasets. In Proc. of the AAAI Conference on Artificial Intelligence (AAAI), 2022b.
Gao, K., Bai, Y., Gu, J., Yang, Y., and Xia, S.-T. Backdoor defense via adaptively splitting poisoned dataset. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
Jha, R. D., Hayase, J., and Oh, S. Label poisoning is all you need. In Oh, A., Naumann, T., Globerson, A., Saenko, K., Hardt, M., and Levine, S. (eds.), Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2023.
Goodfellow, I., Shlens, J., and Szegedy, C. Explaining and harnessing adversarial examples. In Proc. of the International Conference on Learning Representations (ICLR), 2015.
Killamsetty, K., Sivasubramanian, D., Ramakrishnan, G., of Texas at Dallas, R. I. U., of Technology Bombay Institution One, I. I., and Two, I. N. Glister: Generalization based data subset selection for efficient and robust learning. In Proc. of the AAAI Conference on Artificial Intelligence (AAAI), 2020.
Gu, T., Dolan-Gavitt, B., and Garg, S. Badnets: Identifying vulnerabilities in the machine learning model supply chain. Proceeding of Machine Learning and Computer Security Workshop, 2017. Guo, C., Zhao, B., and Bai, Y. DeepCore: A comprehensive library for coreset selection in deep learning. In Proc. of the International Conference on Database and Expert Systems Applications (DEXA), 2022.
Killamsetty, K., Sivasubramanian, D., Mirzasoleiman, B., Ramakrishnan, G., De, A., and Iyer, R. K. GRADMATCH: A gradient matching based data subset selection for efficient learning. In Proc. of the International Conference on Machine Learning (ICML), 2021.
He, K., Zhang, X., Ren, S., and Sun, J. Deep residual learning for image recognition. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2016.
Krizhevsky, A., Nair, V., and Hinton, G. CIFAR (canadian institute for advanced research), 2008. URL http:// www.cs.toronto.edu/˜kriz/cifar.html.
He, X., Xu, Q., Wang, J., Rubinstein, B. I. P., and Cohn, T. SEEP: Training dynamics grounds latent representation search for mitigating backdoor poisoning attacks. In Proc. of the Annual Meeting of the Association for Computational Linguistics, 2024.
Kurita, K., Michel, P., and Neubig, G. Weight poisoning attacks on pretrained models. In Proc. of the Annual Meeting of the Association for Computational Linguistics, 2020.
Hekmatfar, M. and Farahani, R. Z. (eds.). Facility Location: Concepts, Models, Algorithms and Case Studies. Contributions to Management Science. Springer, 2009. doi: 10.1007/978-3-7908-2151-2.
Köhler, J. M., Autenrieth, M., and Beluch, W. H. Uncertainty based detection and relabeling of noisy image labels. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2019.
Hestness, J., Narang, S., Ardalani, N., Diamos, G., Jun, H., Kianinejad, H., Patwary, M., Yang, Y., and Zhou, Y. Deep learning scaling is predictable, empirically. arXiv preprint arXiv:1712.00409, 2017.
Le, Y. and Yang, X. Tiny imagenet visual recognition challenge. CS 231N, 2015. LeCun, Y., Bengio, Y., and Hinton, G. Deep learning. Nature, 521(7553):436–444, 2015. doi: 10.1038/ nature14539.
Huang, G., Liu, Z., and van der Maaten, L. Densely connected convolutional networks. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
Li, Y., Li, Y., Wu, B., Li, L., He, R., and Lyu, S. Invisible backdoor attack with sample-specific triggers. In Proc. of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021a.
Huang, K., Li, Y., Wu, B., Qin, Z., and Ren, K. Backdoor defense via decoupling the training process. In Proc. of the 10
Anti-Backdoor Coreset Selection via Cumulative Entropy
Li, Y., Lyu, X., Koren, N., Lyu, L., Li, B., and Ma, X. Anti-backdoor learning: Training clean models on poisoned data. In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2021b.
11th International Joint Conference on Natural Language Processing, 2021. Qi, X., Xie, T., Li, Y., Mahloujifar, S., and Mittal, P. Revisiting the assumption of latent separability for backdoor defenses. In Proc. of the International Conference on Learning Representations (ICLR), 2023.
Li, Y., Jiang, Y., Li, Z., and Xia, S.-T. Backdoor learning: A survey. IEEE Transactions on Neural Networks and Learning Systems, 2022.
Qin, Z., Wang, K., Zheng, Z., Gu, J., Peng, X., Xu, Z., Zhou, D., Shang, L., Sun, B., Xie, X., and You, Y. InfoBatch: Lossless training speed up by unbiased dynamic data pruning. In Proc. of the International Conference on Learning Representations (ICLR), 2024.
Liu, Y., Ma, S., Aafer, Y., Lee, W.-C., Zhai, J., Wang, W., and Zhang, X. Trojaning attack on neural networks. In Proc. of the Network and Distributed System Security Symposium (NDSS), 2018. Mirzasoleiman, B., Bilmes, J., and Leskovec, J. Coresets for data-efficient training of machine learning models. In Proc. of the International Conference on Machine Learning (ICML), 2020.
Qiu, H., Zeng, Y., Guo, S., Zhang, T., Qiu, M., and Thuraisingham, B. Deepsweep: An evaluation framework for mitigating dnn backdoor attacks using data augmentation. In Proc. of the ACM Asia Conference on Computer and Communications Security (ASIA CCS), 2021.
Müller, R., Kornblith, S., and Hinton, G. When does label smoothing help? In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2019.
Sandler, M., Howard, A., Zhu, M., Zhmoginov, A., and Chen, L.-C. Mobilenetv2: Inverted residuals and linear bottlenecks. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018.
Natarajan, N., Dhillon, I. S., Ravikumar, P. K., and Tewari, A. Learning with noisy labels. In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2013.
Sener, O. and Savarese, S. Active learning for convolutional neural networks: A core-set approach. In Proc. of the International Conference on Learning Representations (ICLR), 2018.
Neel, S., Roth, A., and Sharifi-Malvajerdi, S. Descent-todelete: Gradient-based methods for machine unlearning. In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2020.
Shafahi, A., Huang, W. R., Najibi, M., Suciu, O., Studer, C., Dumitras, T., and Goldstein, T. Poison frogs! targeted clean-label poisoning attacks on neural networks. In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2018.
Nguyen, T. A. and Tran, A. Input-aware dynamic backdoor attack. In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2020. Nguyen, T. A. and Tran, A. T. Wanet - imperceptible warping-based backdoor attack. In Proc. of the International Conference on Learning Representations (ICLR), 2021.
Simonyan, K. and Zisserman, A. Very deep convolutional networks for large-scale image recognition. In Proc. of the International Conference on Learning Representations (ICLR), 2015.
Paul, M., Ganguli, S., and Dziugaite, G. K. Deep learning on a data diet: finding important examples early in training. In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2021.
Stallkamp, J., Schlipsing, M., Salmen, J., and Igel, C. Man vs. computer: Benchmarking machine learning algorithms for traffic sign recognition. Neural Networks, 2012. ISSN 0893-6080.
Pedregosa, F., Varoquaux, G., Gramfort, A., Michel, V., Thirion, B., Grisel, O., Blondel, M., Prettenhofer, P., Weiss, R., Dubourg, V., Vanderplas, J., Passos, A., Cournapeau, D., Brucher, M., Perrot, M., and Duchesnay, E. Scikit-learn: Machine learning in python. Journal of Machine Learning Research, 12:2825–2830, 2011.
Sun, C., Shrivastava, A., Singh, S., and Gupta, A. Revisiting Unreasonable Effectiveness of Data in Deep Learning Era . In Proc. of the IEEE/CVF International Conference on Computer Vision (ICCV), 2017. Toneva, M., Sordoni, A., des Combes, R. T., Trischler, A., Bengio, Y., and Gordon, G. J. An empirical study of example forgetting during deep neural network learning. In Proc. of the International Conference on Learning Representations (ICLR), 2019.
Qi, F., Li, M., Chen, Y., Zhang, Z., Liu, Z., Wang, Y., and Sun, M. Hidden killer: Invisible textual backdoor attacks with syntactic trigger. In the 59th Annual Meeting of the Association for Computational Linguistics and the 11
Anti-Backdoor Coreset Selection via Cumulative Entropy
Turner, A., Tsipras, D., and Madry, A. Label-consistent backdoor attacks. ArXiv, abs/1912.02771, 2019.
minimization. In Proc. of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023a.
van der Maaten, L. and Hinton, G. Visualizing data using t-sne. Journal of Machine Learning Research, 2008.
Zhu, Z., Wang, R., Zou, C., and Jing, L. The victim and the beneficiary: Exploiting a poisoned model to train a clean model on poisoned data. In Proc. of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023b.
Wei, S., Zhang, M., Zha, H., and Wu, B. Shared adversarial unlearning: Backdoor mitigation by unlearning shared adversarial examples. In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2023.
Appendix In the following, we first provide the coreset selection using other criteria (cf. Section A) and compare cumulative uncertainty criteria to three loss-based ones under WaNet attack (cf. Section B). In Section C, we elaborate on the experimental setups of considered attacks and defenses. After that, we proceed to analyze the performance of ABCS in Section D, including: the defense performance on GTSRB dataset (Section D.1), the ablation study on different hyper-settings of ABCS (Section D.2), the robustness under different target classes and poisoning rates (Section D.3), the evaluation across other model architectures (Section D.4), the resistance against all-to-all attacks (Section D.5), the effectiveness on the text classification task with using BERT models (Section D.6), and the comparison to the dataset splitting approach using the clean reference data (Section D.7) . Furthermore, we evaluate the resistance of ABCS against various adaptive adversaries that constructs poisoned datasets by intensionally enlarging the prediction uncertainty or/and the learning difficulty of poisoned samples (Section E). Finally, we investigate the geometric distribution of selected coresets in the latent space representation to empirically examine the coverage of the entire clean dataset (Section F).
Wu, B., Chen, H., Zhang, M., Zhu, Z., Wei, S., Yuan, D., and Shen, C. Backdoorbench: A comprehensive benchmark of backdoor learning. In Thirty-sixth Conference on Neural Information Processing Systems Datasets and Benchmarks Track, 2022. Wu, D. and Wang, Y. Adversarial neuron pruning purifies backdoored deep models. In Proc. of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2021. Zeng, Y., Park, W., Mao, Z. M., and Jia, R. Rethinking the backdoor attacks’ triggers: A frequency perspective. In Proc. of the IEEE/CVF International Conference on Computer Vision (ICCV), 2021. Zhang, K., Cheng, S., Shen, G., Tao, G., An, S., Makur, A., Ma, S., and Zhang, X. Exploring the orthogonality and linearity of backdoor attacks. In Proc. of the IEEE Symposium on Security and Privacy, 2024. Zhang, Z., Liu, Q., Wang, Z., Lu, Z., and Hu, Q. Backdoor defense via deconfounded representation learning. In Proc. of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2023.
DeepFool
Glister
k-Center Greedy
ACC / ASR [%]
100
Zhao, Q. and Wressnegger, C. Adversarially robust antibackdoor learning. In Proc. of the ACM Workshop on Artificial Intelligence and Security (AISEC), October 2024.
10
80
ACC (Baseline) 8
60
6
40
ACC ASR ρbng
20 0 0.3
0.4
Selection Ratio
Zhao, Q. and Wressnegger, C. Two sides of the same coin: Learning the backdoor to remove the backdoor. In Proc. of the Annual AAAI Conference on Artificial Intelligence (AAAI), February 2025.
0.5 0.3
0.4
Selection Ratio
0.5 0.3
0.4
4
ρbng [%]
Welling, M. Herding dynamical weights to learn. In Proc. of the International Conference on Machine Learning (ICML), 2009.
2 0 0.5
Selection Ratio
Figure 9. Coreset selected by three additional criteria on CIFAR10 poisoned by Blend with ρ = 5 % using ResNet18. Error bars indicate the range across five random trials per coreset size.
Zhao, S., Ma, X., Zheng, X., Bailey, J., Chen, J., and Jiang, Y. Clean-label backdoor attacks on video recognition models. In Proc. of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023.
A. Analysis of Other Coreset Selection Criteria In Figure 9, we investigate the defensive behavior of three selection criteria: DeepFool (Ducoffe & Precioso, 2018), Glister (Killamsetty et al., 2020) and k-Center Greedy (Hekmatfar & Farahani, 2009). Unlike uncertainty-based and
Zhu, M., Wei, S., Shen, L., Fan, Y., and Wu, B. Enhancing fine-tuning based backdoor defense with sharpness-aware 12
Anti-Backdoor Coreset Selection via Cumulative Entropy GraNd
C. Experimental Setup
EL2N
80
80
60
60
40
40
ACC [%]
100
20
In this section, we outline the experimental setups for considered attacks and defenses. Note that each experiment is conducted on one single Nvidia RTX-3090 GPU with Intel(R) Xeon(R) W-2245 CPU. Each result is conducted by 5 runs of experiment for small-scale datasets CIFAR10 and GTSRB, and 3 runs of experiment for the large-scale dataset Tiny-ImageNet.
ASR [%]
Forgetting 100
20
0 0.3
ACC (Baseline) 0.4
0.5 0.3
Selection Ratio
0.4
0.5 0.3
Selection Ratio
0 0.5
0.4
Selection Ratio
(a) Loss-based Criteria
C.1. Considered Attacks Cumu. Margin
Cumu. LeastConf 100
80
80
60
60
ACC [%]
100
40
40
20
ACC 20 ASR 0 0.5
0 0.3
0.4
Selection Ratio
0.5 0.3
0.4
Selection Ratio
0.5 0.3
0.4
We generate each poisoned dataset by conforming to the anti-backdoor learning application scenario, that is, the backdoor is introduced via data poisoning only. Each attack is implemented as follows.
ASR [%]
Cumu. Entropy
BadNets (Gu et al., 2017): We use a colored 2 × 2 square pattern for dataset poisoning.
Selection Ratio
Blend (Chen et al., 2017): We use the default Hello-Kitty image as the trigger pattern for poisoning CIFAR10 and GTSRB, and a random noise trigger for Tiny-ImageNet. We use a trigger opacity of α = 0.1, as it has been proven effective to achieve an attack success rate near 100 %.
(b) Cumulative Uncertainty Criteria Figure 10. Comparing loss-based criteria to uncertainty-based criteria with accumulation. All experiments are conducted by using ResNet18 model on CIFAR10 that is under the poisoning of WaNet attack with ρ = 5 %. Error bars show the value range across five random runs per coreset selection ratio.
CLB (Turner et al., 2019): We poison each sample by adopting the projected gradient descent (PGD) method to generate adversarial perturbation with strength ϵ = 16/255 and step size 2/255 for 30 steps. IAB (Nguyen & Tran, 2020): We train the generator on each dataset the default hyper-parameters: the backdoor probability ρb = 0.1, the cross-trigger probability ρc = 0.1, and the weighting parameter λdiv = 1.0 for the diversity loss term. For each dataset, we train an individual generator for constructing IAB poisoning.
loss-based criteria in Figure 2, these criteria above select coresets that either exhibit low stability, lack robustness of eliminating poisonous samples, leading to a high ASR after training, or even significantly degrade the natural performance after model training. Therefore, we see these criteria as unqualified methods. Hence, we focus on the uncertaintybased and loss-based methods in our main analysis.
WaNet (Nguyen & Tran, 2021): We follow the data poisoning implementation of WaNet, i.e., first generating a warping trigger function and its accompanied noise function and, then, using both pattern functions to poison data samples. For all datasets, we following the default settings of WaNet, i.e., using the noise ratio equal 2 × ρ, the grid size k = 4, and the warping strength s = 0.5.
B. Impact of Selection Criteria for WaNet Figure 10 presents the comparison between the naive application of loss-based criteria and uncertainty-based criteria with accumulation, under WaNet attack scenario. Unlike Blend attack case (cf. Figure 2), the EL2N method fails to effectively exclude poisonous samples during coreset selection, resulting a high ASR. This indicates that solely using loss-based criteria cannot achieve a robust elimination of poisonous samples across diverse attacks. In contrast, incorporating accumulation with uncertainty criteria yields better defensive performance, despite a relatively high ASR at a selection ratio of 0.5. Notably, at the selection ratio 0.4, using C ENT achieves a natural accuracy comparable to the baseline while significantly mitigating the WaNet backdoor. This demonstrates the advantage of adopting Cumulative Entropy criterion in anti-backdoor coreset selection.
ISSBA (Li et al., 2021a): For each dataset and poisoning rate, we execute the ISSBA attack by following the implementation provided by BackdoorBench (Wu et al., 2022). Low Frequency (LF) (Zeng et al., 2021): We reuse the generated triggers that are provided by BackdoorBench (Wu et al., 2022) and run the data poisoning. Adaptive Blend (A-Blend) (Qi et al., 2023): We adopt the same trigger as used in (Qi et al., 2023) and set the trigger opacity equal 0.15 for training set and 0.2 for test set. For the random trigger partitioning, we set the mask rate equal 0.5 and coverage rate equal 1.0 and use 16 masking pieces. 13
Anti-Backdoor Coreset Selection via Cumulative Entropy Table 6. Comparing ABCS to prior training-time defenses on GTSRB. Each result contains the averaged value and the standard deviation over five random runs. The best results across all defenses are highlighted as boldface. The defense failure (i.e., ASR > 50 %) is shown as orange boldface. We discard ASR for No-Defense as all attacks reach an ASR above 99 %. Attack
No-Defense
ABL
DBD
CBD
ACC (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
BadNets Blend CLB IAB WaNet ISSBA LF A-Blend
97.82± 0.70 98.08± 0.37 98.06± 0.08 98.03± 0.45 97.44± 0.61 98.04± 0.06 97.93± 0.28 97.56± 0.29
89.30±9.59 91.29±2.07 86.64±2.95 94.56±3.68 86.32±6.16 88.81±2.82 86.58±2.86 92.28±3.06
53.13±50.23 94.96± 7.21 7.53± 5.92 20.74±43.91 99.76± 0.05 99.86± 0.14 84.62±32.42 0.22± 0.22
69.18±28.85 48.81± 3.81 90.50± 1.96 87.87±23.79 44.44± 3.08 45.45± 1.38 52.02±15.41 97.19± 1.62
92.36±0.42 93.03±0.13 92.98±0.11 92.83±0.61 92.64±0.31 92.74±0.54 92.32±0.48 92.84±0.25
0.02± 0.03 99.74± 0.38 0.47± 0.16 0.28± 0.18 0.00± 0.00 99.46± 0.64 27.45± 9.53 80.29±18.00
97.26± 0.21 47.48± 0.07 97.20± 0.10 97.23± 0.28 95.71± 0.16 47.62± 0.10 83.47± 4.72 57.46± 9.04
91.80±3.11 86.28±8.99 82.99±2.57 83.05±4.10 96.20±0.25 76.89±5.12 74.79±7.44 81.66±4.43
80.31±43.98 97.42± 2.20 10.21±11.70 39.12±38.06 97.25± 0.46 44.69±34.43 77.44±37.29 91.03± 2.03
56.84±21.34 45.02± 4.61 87.34± 6.57 72.92±19.46 49.38± 0.12 67.08±18.06 49.70±16.26 46.47± 1.67
Average WorstCase
97.87 97.44
89.47 86.32
57.60 99.86
66.93 44.44
92.72 92.32
38.46 99.74
77.93 47.48
84.21 74.79
67.18 97.42
59.34 45.02
Attack
No-Defense ACC (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
BadNets Blend CLB IAB WaNet ISSBA LF A-Blend
97.82± 0.70 98.08± 0.37 98.06± 0.08 98.03± 0.45 97.44± 0.61 98.04± 0.06 97.93± 0.28 97.56± 0.29
93.35±0.73 92.89±2.68 91.16±0.53 93.56±2.40 94.57±1.05 91.58±1.55 93.49±1.00 93.36±1.74
0.05± 0.04 98.01± 1.38 13.00±16.88 89.30±15.36 46.04±18.25 78.74±27.00 98.15± 2.29 98.55± 1.01
97.74± 0.38 48.04± 1.15 90.02± 8.59 53.09± 7.97 73.65± 9.14 57.40±13.12 48.70± 0.68 48.56± 1.22
98.19±0.20 98.04±0.10 98.26±0.59 98.11±0.23 97.88±0.30 97.81±0.27 94.99±1.84 94.97±2.56
0.00± 0.00 0.05± 0.05 1.04± 0.64 0.03± 0.06 7.08± 2.86 0.00± 0.00 99.83± 0.20 96.66± 1.09
100.00± 0.00 99.56± 0.07 99.38± 0.24 99.93± 0.05 94.57± 1.43 99.88± 0.12 48.61± 0.83 50.23± 1.32
98.27±0.25 98.05±0.32 97.62±0.33 98.20±0.26 98.11±0.39 97.90±0.18 97.91±0.29 96.37±0.98
0.00± 0.00 0.40± 0.37 1.53± 0.44 0.01± 0.02 0.01± 0.02 0.01± 0.03 1.52± 1.69 7.94± 1.59
100.00± 0.00 99.33± 0.18 98.98± 0.17 99.94± 0.05 98.10± 0.01 99.91± 0.07 99.18± 0.81 95.32± 1.07
Average WorstCase
97.87 97.44
92.99 91.16
65.23 98.55
64.65 48.04
97.28 94.97
25.59 99.83
86.52 48.61
97.80 96.37
1.43 7.94
98.85 95.32
V&B
Harvey
ABCS (Ours)
Accuracy (ACC) [%]
100 90 80 70 60 Blend
WaNet
BadNets
IAB
Clean D
50 0
1
2
3
4
5
6
7
8
9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42
Class
Figure 11. Class-wise accuracy distribution on GTSRB across different attacks, averaged over five runs. “Clean D” refers to using the original clean dataset for the model training.
C.2. Considered Defenses
backdoor attacks. Therefore, we directly follow the default settings of DBD to conduct all experiments.
ABL (Li et al., 2021b): The ABL procedure consists of three stages: (1) training for 20 epochs on the entire poisoned dataset and isolating 1 % of the samples with the lowest loss, (2) fine-tuning the model on the remaining dataset, and (3) unlearning the isolated poisonous set for 5 epochs with a learning rate of 0.0001. The hyperparameter γ in LGA (Local Gradient Ascent) is sensitive to different attacks. However, for practicality, we consistently use γ = 0.5 in the implementation.
CBD (Zhang et al., 2023): CBD learns a model on the poisoned dataset and then trains a clean model by maximizing mutual independence from the former. Since the first model is responsible for identifying poisonous samples, the number of training epochs in the first phase is a sensitive hyperparameter for different attacks. The original implementation selects the best number from {3, 5, 8}. Accordingly, we tested different numbers of epochs and takes the result with the highest natural accuracy.
DBD (Huang et al., 2022a): DBD first adopts selfsupervised learning for 1,000 epochs to extract benign features from the dataset. Then, it fine-tunes the fully connected layer for 10 epochs with supervised learning, splitting the dataset half-and-half, and subsequently uses semisupervised learning to train the entire model for 200 epochs. DBD does not specify other hyperparameters for individual
V&B (Zhu et al., 2023b): The V&B defense first exploits a model with a backdoor to identify suspiciously poisoned samples and then trains a clean model on benign ones. Afterward, it adopts semi-supervised learning, as in (Huang et al., 2022a), on the split datasets to further improve the natural performance of the clean model while simultaneously 14
Anti-Backdoor Coreset Selection via Cumulative Entropy
suppressing the backdoor functionality. In our evaluation, we use the default settings and hyperparameters of V&B for both the small-scale and large-scale datasets.
Moreover, training on the full clean dataset yields notably low accuracy at class 21 and 27. For these difficult classes, our method yields comparable or even higher accuracies, demonstrating its potential to preserve fairness.
Harvey (Zhao & Wressnegger, 2025): Harvey constructs a reference model with a backdoor through warm-up training and subsequently boosts convergence toward the backdoor by iteratively training on a poisoned subset and re-splitting the poisoned dataset. We follow the default settings described in the appendix of Harvey for all experiments.
D.2. Impact of Other Global Settings on ABCS Rationale of excluding mis-predicted samples. To illustrate the importance of automatic thresholding, we conduct additional experiments for the Blend attack across three datasets in Table 7. The selection ratio rse varies significantly across datasets, highlighting the necessity of an automatic threshold over a fixed one. Our approach effectively adjusts the threshold to maintain high natural accuracy, thereby improving robustness and adaptability. Moreover, our method excludes mis-prediced samples when calculating the threshold. In additon, we compare our strategy with a variant that applies the automatic threshold including mispredicted samples as well, denoted as All Samples. The latter leads to a smaller selection ratio and consequently worse natural accuracy (especially on Tiny-ImageNet). These findings emphasize that our automatic thresholding strategy not only preserves performance but also generalizes better across different datasets.
D. Extended Evaluation In this section, we conduct additional experiments to extend our evaluation. We first examine all defenses on GTSRB dataset (cf. Section D.1), and then investigate the impact of other global settings on ABCS (cf. Section D.2). After that, we test ABCS across different poisoning settings (cf. Section D.3), with various model architectures (cf. Section D.4) and against the all-to-all attacks (cf. Section D.5). Moreover, we examine the potential of ABCS for the text classification domain (cf. Section D.6). Finally, we compare to other defense formats that either possesses certain clean reference samples (cf. Section D.7) or remove backdoors after the training (cf. Section D.8).
Across Twa
Across Tse
95
10 ACC (BadNets) ACC (Blend) ACC (IAB) ACC (WaNet)
ACC [%]
90 85 80
ASR (BadNets) 8 ASR (Blend) 6 ASR (IAB) ASR (WaNet) 4
75
2
70 3
5
8 10
20 10
Epochs
D.1. Evaluation on GTSRB
ASR [%]
ABCS (Ours): During the coreset selection, we exclude all data augmentations to stabilize backdoor learning (Li et al., 2021b; Qiu et al., 2021). In addition to the setup in Section 5, the final training on coresets follows standard settings: models are trained for 200 epochs using the SGD optimizer with a weight decay of 0.0005, and the learning rate is scheduled via cosine annealing from 0.1 to 0.0001.
20
30
40
0 50
Epochs
Figure 12. Investigating the impact of warm-up epochs Twa (left), selection epochs Tse (right) on ABCS’s defense for CIFAR10.
Table 6 summarizes the experiments on the GTSRB dataset. For the ResNet18 model, learning on GTSRB is easier than on CIFAR10, resulting in a natural accuracy close to 98 %. However, most prior defenses fail to consistently maintain high accuracy across all attacks. Moreover, prior defenses show a drop in ACC up to 7 % and allow at least one attack to successfully implant the backdoor into the model. In contrast, training on the coreset selected by ABCS effectively mitigates all backdoor attacks while maintaining high natural accuracy.
Twa and Tse . Based on the ablation study in Figure 12, both warm-up epochs Twa and selection epochs Tse have minimal impact on natural accuracy. However, a shorter warm-up phase may fail to capture backdoor behavior, leading to less effective filtering of poisonous samples and, thus, a higher ASR in the final training. Similarly, ABCS benefits from a longer selection phase: accumulating entropy over more epochs ensures the exclusion of poisonous samples
Analysis of class-wise accuracy preservation. Unlike CIFAR10 that has a uniform class-wise distribution, GTSRB is intrinsically imbalanced across its all 43 classes (Stallkamp et al., 2012). Figure 11 compares the class-wise accuracy yielded by training on coresets determined using a model trained on a fully clean dataset D. Our method benefits from the sufficiently large size of coresets (cf. Table 7) and, thus, ensures the coverage of the dataset informativeness.
Table 7. Impact of excluding mis-predicted samples during ABCS’s coreset selection. Dataset
All Samples ACC
ASR
Excl. Mis-predicted Samples rse
ACC
ASR
rse
CIFAR10 93.88±0.31 1.09±0.24 44.67±0.81 94.62±0.36 0.77±0.08 53.85±0.58 GTSRB 97.39±0.42 0.32±0.26 31.26±1.00 98.05±0.32 0.40±0.37 33.13±0.97 Tiny-ImageNet 38.53±1.02 0.01±0.01 48.72±0.15 61.11±0.07 0.18±0.05 82.30±0.64
15
Anti-Backdoor Coreset Selection via Cumulative Entropy
and retrain informative benign ones in the selected coreset, thereby yielding high natural accuracy and low ASR.
95
100
85
80
ASR (BadNets) ASR (Blend) ASR (IAB) ASR (WaNet)
60
80
40
75
20
70
Ldis
ACC (BadNets) ACC (Blend) ACC (IAB) ACC (WaNet)
ASR [%]
90
ACC [%]
the mean loss value of benign and poisoned samples, respectively. A lower poisoning rate will slow down the convergence speed to the backdoor, thus, making the loss discrepancy become smaller in the initial epochs. For ρ ≥ 5 %, ABCS conBlend Attack sistently selects a core1 set that yields a model 0 with high natural accu−1 racy and minimal ASR, −2 1% 5% indicating effective exclu10 % 20 % −3 sion of poisonous sam5 10 15 20 ples. At a lower poiEpoch soning rate (ρ = 1 %), Figure 15. Impact of varying ρ. the increased difficulty of learning backdoor slightly raises the risk of including several poisoned samples in the coreset. Despite this, ABCS continues to perform robustly with anti-backdoor coreset selection, ensuring a strong mitigation of backdooring attacks (i.e., ASR < 5 %) while maintaining high natural accuracy.
0
0.01
0.05
0.1
0.5
1
γ
Figure 13. Impact of γ on ABCS’s defense for CIFAR10.
Values of γ. While γ decreases, the natural accuracy (ACC) increases (cf. Figure 13). A large value of γ (≥ 0.1) effectively lowers the ASR, while a smaller γ leads to insufficient unlearning and, thus, results in ineffective separation of poisoned samples and benign samples with high C ENT uncertainty, thereby yielding a high ASR in the final model.
D.4. Evaluation with Different Model Architectures We evaluate ABCS across three additional model architectures, i.e.,, VGG11 (Simonyan & Zisserman, 2015), MobileNetV2 (Sandler et al., 2018) and DenseNet121 (Huang et al., 2017), under all considered attacks. As shown in Table 8, training on the coresets selected by ABCS consistently yields natural accuracy comparable to that of training on a clean dataset. Simultaneously, the low ASR demonstrates a strong backdoor mitigation, highlighting effective elimination of poisonous samples during coreset selection.
D.3. Evaluation Across Different Poisoning Settings Target classes. We evaluate the robustness of ABCS across different target classes under variant attacks in Figure 14 (left). For each poisoned CIFAR10, we report the average performance and error bar over five random trials. Across all target classes, ABCS consistently select a coreset that enables training a model from scratch with high natural performance while keeping ASR below 2 %. Across Target Classes
Across Poisoning Rates
90
ACC (BadNets)
ASR (BadNets)
8
85
ACC (Blend)
ASR (Blend)
6
ACC (IAB)
ASR (IAB)
ACC (WaNet)
ASR (WaNet)
We set up all-to-all attacks with the same settings as (Li et al., 2022), i.e., taking the next neighbor class as the backdoor target of each poisoned sample. Additionally, we extend the setting to using the Hello-Kitty trigger of Blend attack. As shown in Table 9, Table 9. ABCS’s defense against our method remains efall-to-all attacks on CIFAR10. fective against both two All-to-All ACC ASR DER all-to-all attacks, showBadNets 94.36±0.30 0.56±0.05 94.19±0.46 ing a high robustness Blend 94.00±0.16 1.86±0.86 93.34±0.40 our method.
80
4
ASR [%]
10
95
ACC [%]
D.5. Evaluation on All-to-All Attacks
2
75 70 0
1
2
3
4
5
Class
6
7
8
9
1
5
10
0 20
Poisoning Rate ρ [%]
Figure 14. Evaluation across target classes (left) and different poisoning rates (right) on CIFAR10.
D.6. Evaluation on Text Classification Task Following Cheng et al. (2025), we adopt the standard datapoisoning setting (Kurita et al., 2020), where the trigger is a rare word “cf” inserted randomly into each poisoned text sample, and use the target label 0. Additionally, we reduce the poisoning rate to 5 %. We implement this attack on two versions of the Stanford Sentiment Treebank (SST) dataset: SST-2 (binary sentiment classification) and SST-5 (five sentiment classes) using pre-trained BERT-base (uncased) and
Poisoning rates. Figure 14 (right) shows the performance of ABCS across different poisoning rates. A higher poisoning rate ρ corresponds to a stronger attack mode, while a lower ρ increases the difficulty of learning backdoor. Figure 15 shows the training processes under different poisoning rates. We use Ldis = Lbng − Lpoi to measure the loss discrepancy between benign and poisoned samples, where Lbng and Lpoi are 16
Anti-Backdoor Coreset Selection via Cumulative Entropy Table 8. Evaluation of ABCS across model architectures on CIFAR10. For each model, ACC from training on the clean dataset is shown next to the model name. In the table, ACC results (in [%]) correspond to training on selected coresets from each poisoned dataset. For backdoor mitigation, ASR (in [%]) is reported for training on the poisoned dataset and on the selected coreset derived by ABCS (shown to the left and right of “ / ”, respectively). VGG11 (91.85 %)
Attack
BadNets Blend CLB IAB WaNet ISSBA LF A-Blend
MobileNetV2 (93.15 %)
ASR (↓)
ACC (↑)
ASR (↓)
ACC (↑)
ASR (↓)
91.76±0.08 90.88±0.85 91.38±0.11 91.48±0.35 89.38±1.05 91.16±0.29 91.22±0.38 90.82±0.29
99.98±0.11 / 1.32±0.22 99.12±0.31 / 2.00±0.23 99.62±0.25 / 2.90±0.29 99.84±0.35 / 1.62±0.39 96.43±0.84 / 2.65±0.64 99.07±0.29 / 1.83±0.26 98.14±0.74 / 7.22±1.86 99.04±0.54 / 9.02±0.79
92.93±0.07 92.84±0.02 92.83±0.16 92.87±0.27 92.44±0.08 92.50±0.26 92.48±0.10 91.93±0.47
99.96±0.06 / 0.99±0.13 99.23±0.25 / 1.14±0.18 99.87±0.13 / 1.18±0.16 99.94±0.04 / 2.37±0.38 97.25±0.68 / 1.93±0.31 99.21±0.32 / 1.66±0.24 98.31±0.57 / 6.66±1.97 99.52±0.22 / 7.82±1.07
94.97±0.26 94.65±0.15 94.60±0.02 94.98±0.21 94.86±0.09 94.71±0.22 94.52±0.18 94.47±0.11
100.00±0.00 / 0.83±0.06 99.25±0.32 / 0.90±0.45 99.48±0.15 / 0.90±0.15 99.69±0.28 / 1.25±0.33 98.35±0.32 / 0.82±0.08 99.06±0.41 / 0.80±0.37 98.36±0.89 / 5.83±1.61 99.42±0.51 / 7.32±0.95
D.7. Comparison to Defense with Clean Reference Data
BERT-large (uncased) models (Devlin et al., 2019) as basis. To stabilize the fine-tuning process during coreset selection on the pre-trained BERT model, we use the AdamW optimizer with a learning rate of 2e − 6. For training the final coreset, the learning rate is increased to 2e − 5. Due to the faster convergence during fine-tuning the pre-trained models, we set Twarm = 5 epochs for warm-up and Tse = 10 epochs for coreset selection. The resulting coreset is then used to fine-tune the original pre-trained BERT model for 10 epochs. As shown in Table 10, our method adaptively determines different selection ratios rse for SST-2 and SST5. Despite a slight drop in accuracy for SST-5, the results show that our method remains robust against data poisoning backdoors in text classification tasks.
While assuming the absence of any reference dataset is most practical, this setting imposes a strong restriction. Many post-training defenses thus permit the access to a small portion of clean data. A representative method is ASD, the Adaptively Splitting Defense (Gao et al., 2023), which picks 10 clean samples per class uniformly of random for model initialization. With the prior knowledge of benign samples, the subsequent adaptive dataset splitting can more effectively isolate poisoned samples due to their higher prediction losses than benign ones. Table 12. Comparing ABCS to Adaptively Splitting Defense (ASD) on CIFAR10. Each result contains the averaged value and the standard deviation over five random runs. The best results are highlighted as boldface. ASR for No-Defense is discarded as all attacks reaches a ASR over 99 %.
Table 10. Effectiveness of ABCS on text classification task. BERT Model
Dataset
No-Defense ACC
ABCS
ASR
ACC
ASR
ρbng
Attack
rse
Base
SST-2 SST-5
92.33±0.31 99.97±0.04 91.63±0.52 6.89±0.97 0.01±0.02 34.35±0.40 49.83±0.37 99.45±0.33 48.20±0.82 3.53±0.58 0.00±0.00 57.21±1.97
Large
SST-2 SST-5
93.39±0.98 99.92±0.12 92.50±0.87 6.23±0.91 0.01±0.01 35.44±0.30 50.80±0.87 99.75±0.09 48.53±0.73 3.40±0.32 0.00±0.00 56.72±0.51
BadNets Blend CLB IAB WaNet ISSBA LF A-Blend
BadNet Syntactic
SEEP
ABCS (ours) ACC / ASR
92.6 / 7.4 91.5 / 10.0
92.3 / 4.9 91.7 / 12.7
ABCS (Ours)
ASR (↓)
DER (↑)
ACC (↑)
ASR (↓)
DER (↑)
93.56±0.39 1.39± 0.52 93.31±0.65 5.50± 1.78 92.92±0.64 1.30± 0.90 93.42±0.45 2.25± 1.03 92.63±0.41 17.29±21.24 92.80±0.35 2.26± 0.67 93.72±0.23 5.66± 3.67 93.24±0.38 34.84± 6.99
98.82± 0.29 95.30± 0.79 98.58± 0.30 98.23± 0.48 88.67±10.78 98.01± 0.36 95.85± 1.78 80.18± 3.44
94.38±0.23 94.62±0.36 94.51±0.23 94.30±0.16 94.13±0.22 94.66±0.41 94.06±0.30 94.37±0.27
1.00±0.17 0.77±0.08 0.88±0.10 1.22±0.37 0.89±0.16 1.22±0.09 3.04±0.55 5.71±1.59
99.42±0.19 98.22±0.07 99.54±0.06 99.19±0.19 97.53±0.13 99.33±0.15 97.33±0.17 95.12±0.77
94.20 80.18
94.38 94.06
1.84 5.71
98.21 95.12
Average 93.20 WorstCase 92.63
8.81 34.84
In Table 12, we evaluate ABCS’s performance on CIFAR10 against ASD. Unlike other defenses that often fail (cf. Table 1a), ASD shows a higher robustness across diverse attacks. Nonetheless, its performance remains limited against A-Blend and WaNet, and it struggles to preserve natural accuracy. In contrast, ABCS consistently produces an informative, clean coreset that supports training with minimal ASR while maintaining accuracy comparable to the original model. Although the reference data facilitates more reliable dataset splitting, it is intrinsically less effective in a training progress with dynamics. ABCS circumvents this limitation by extracting a backdoor-free coreset, enabling a
Table 11. Comparing ABCS to SEEP.
ACC / ASR
ASD ACC (↑)
In addition, we compare with SEEP (He et al., 2024), a typical defense against data-poisoning backdoors in text classification. Due to the lack of publicly available code, we follow SEEP’s reported poisoning settings and evaluate on the SST-2 dataset under BadNet (Kurita et al., 2020) and Syntactic (Qi et al., 2021) attacks. ABCS preserves high ACC while effectively eliminating poisoned samples, achieving defensive performance comparable to SEEP (cf. Table 11).
Attack
DenseNet121 (95.16 %)
ACC (↑)
17
Anti-Backdoor Coreset Selection via Cumulative Entropy
more reliable and robust training-time backdoor defense.
after naive training but also complicating the anti-backdoor coreset selection. As a result, ABCS yields a slightly higher ASR under adaptive attacks. Nevertheless, ABCS effectively selects coresets from the entire dataset, maintaining high natural accuracy while significantly reducing ASR.
D.8. Comparison to Post-train Defenses In addition, we compare ABCS to post-training defenses, ANP (Wu & Wang, 2021) and SAU (Wei et al., 2023), which aim to remove backdoors from models pre-trained on poisoned datasets. Both methods are implemented using the BackdoorBench (Wu et al., 2022), with ResNet18 on CIFAR10. Results are reported in Table 13. Although both defenses leverage some clean reference samples, they achieve effective backdoor removal at the cost of reduced natural accuracy. In contrast, ABCS preserves high natural accuracy while still strongly suppressing backdoors, demonstrating its effectiveness in selecting a backdoor-free coreset.
Table 14. ABCS against adaptive attacks using C ENT ranking. For both ACC and ASR, we show the results of No-Defense and ABCS on the left and right of “ / ”, respectively Attack
ACC (↑)
ASR (↓)
BadNets Blend IAB WaNet ISSBA LF A-Blend
94.22±0.05 / 93.98±0.20 94.11±0.11 / 94.01±0.11 94.42±0.11 / 93.98±0.22 93.87±0.18 / 92.15±0.28 94.13±0.09 / 94.06±0.09 94.37±0.09 / 93.29±0.24 94.39±0.15 / 93.99±0.41
99.94±0.04 / 1.31±0.22 92.19±1.09 / 1.80±0.92 99.55±0.16 / 3.58±3.71 79.06±2.47 / 6.44±2.80 98.01±0.35 / 1.27±0.24 92.60±1.48 / 8.20±1.07 99.98±0.02 / 11.69±2.19
Table 13. Comparing ABCS to post-train defenses. Attack BadNets Blend IAB WaNet
ANP
SAU
ABCS
ACC / ASR
ACC / ASR
ACC / ASR
92.08 / 1.27 88.55 / 4.40 93.08 / 2.08 92.48 / 0.60
92.97 / 2.72 91.64 / 1.20 94.91 / 0.45 93.11 / 1.81
94.38 / 1.00 94.62 / 0.77 94.30 / 1.22 94.13 / 0.89
Algorithm 1 Greedy Search of Poisoned Samples Input: Clean dataset D, a ResNet18 as fθ , the number of iteration T and target poisoning rate ρ Output: A poisoned dataset D̃ Initialisation: Train fθ for 40 epochs on D and poison 50% samples with highest CENT as D̃(0) (0.5 − ρ) · |D| Compute step size n = clip T for t ∈ [1, . . . , T ] do Step 1: Train fθ from scratch for 40 epochs Step 2: Calculate CENT for all poisoned samples Step 3: Sort poisoned samples and recover n samples with the lowest CENT to benign Step 4: Obtain poisoned D̃(t) with (0.5 × |D| − n) poisoned samples end for
E. Robustness Against Adaptive Attacks Anti-backdoor learning allows a full access to the dataset without any control of the training process. Under this setting, we evaluate ABCS against five different adaptive attacks that attempt to circumvent our defense by either increasing the predictive uncertainty of poisoned samples (cf. Section E.1 and E.2), intentionally restricting the training dataset utility (cf. Section E.3), or enforcing a larger learning difficulty for poisoned samples (cf. Section E.4, cf. Section E.5).
Few-shot poisoning of informative samples. We consider another adversary that increases the C ENT of poisoned samples by using a greedy search strategy, which begins with a high poisoning ratio (ρ = 50 %) and iteratively recovers a portion of poisoned samples with low C ENT values back to their originally benign ones, finally reaching the target poisoning ratio. The full procedure is described in Algorithm 1. For typical attacks such as Blend and WaNet, we run for 5 and 10 iterations (cf. Table 15), where the latter is observed to converge. Across both settings, exploring poisoned samples with high C ENT scores does not realistically improve the uncertainty of poisoned samples close to those of informative and benign samples. Using ABCS remains effective of extracting clean and informative coreset.
E.1. Poisoning Samples Using C ENT Ranking We first assume that the adversary understands the coreset selection strategy of using the C ENT criterion. Given that, we conduct two possible adversaries that utilizes the ranking of C ENT for the dataset poisoning. One-shot poisoning of informative samples. Intuitively, the adversary can first rank all clean samples according to the order of Cumulative Entropy and only poisons ρ = 5 % samples with the highest uncertainty. This way, poisoned samples contain hard-to-learn features, thereby implicitly increasing their prediction uncertainty during model training. Table 14 summarizes the experiments of ABCS on CIFAR10 with a ResNet18 model against diverse attacks using the adaptive method, where we exclude CLB attack due to its clean-label poisoning constraint. Compared to Table 1, poisoning informative samples increases the difficulty of learning backdoors, thereby reducing the attack success rate
E.2. Poisoning a Single Source Class Given the threat model, an adversary can poison the dataset arbitrarily. One might be class-selective, where all poisoned samples originate from a single class, leading to a 18
Anti-Backdoor Coreset Selection via Cumulative Entropy Table 15. ABCS against few-shot poisoning with C ENT ranking. For both ACC and ASR, we show the results of No-Defense and ABCS on the left and right of “ / ”, respectively Attack
ACC (↑)
ASR (↓)
T =5
Blend WaNet
94.89±0.24 / 94.85±0.16 94.60±0.22 / 94.47±0.24
99.94±0.03 / 0.94±0.14 91.18±2.61 / 3.16±2.68
T = 10
Blend WaNet
94.85±0.32 / 94.43±0.27 94.28±0.14 / 94.55±0.39
99.96±0.02 / 1.46±0.48 96.25±2.45 / 2.35±1.33
cient data coverage becomes more difficult. Compared to the No-Defense baseline, ABCS selects a slightly smaller coreset, leading to a modest drop in ACC. Importantly, it automatically converges to a relatively large selection ratio rse (with respect to the original training coreset), thereby retaining as many samples as possible to preserve accuracy. This behavior is consistent with observations on the largescale dataset Tiny-ImageNet (cf. Table 1b), highlighting the adaptability of ABCS in low-redundancy settings.
distribution shift between clean and poisoned data. Such imbalanced poisoning could, in principle, increase prediction uncertainty or learning stochasticity associated with the backdoor trigger, potentially challenging the defensive robustness of ABCS.
Table 16. Defending backdoors in CIFAR10 coresets (rse = 50 %).
To directly evaluate this scenario, we treat each non-target class as the sole poisoning source and reduce the overall poisoning rate ρ to 1 % to preserve the majority of clean samples in the chosen source class. As shown in Figure 16, we conduct experiments across all source classes on CIFAR10. Our results show that natural accuracy remains stable regardless of the poisoning source. ABCS consistently produces a coreset that ensures a high post-training accuracy across all attack types. More importantly, our method remains robust even under this highly non-uniform poisoning setting and achieves ASR below 5 % on average. These results demonstrate that ABCS is resilient to class-selective and distribution-shifted poisoning. 95
15
10
80
5
75
0 1
2
3
4
5
6
7
8
ASR (↓)
ρbng
rse
BadNets
98.44±0.06 1.20±0.38
/ 0.00±0.00
/ 77.04±0.62
Blend
No-Defense ABCS
94.72±0.32 93.14±0.39
89.84±1.35 1.34±0.87
/ 0.02±0.02
/ 76.20±0.74
IAB
No-Defense ABCS
94.62±0.34 92.85±0.31
96.87±1.40 1.07±0.82
/ 0.02±0.01
/ 76.20±0.56
WaNet
No-Defense ABCS
93.92±0.74 92.02±1.12
71.86±4.25 4.40±2.25
/ 1.38±0.45
/ 75.48±1.26
∆H
ACC [%]
85
ASR (BadNets) ASR (Blend) ASR (IAB) ASR (WaNet)
ACC (↑) 94.68±0.24 93.52±0.46
With full access to the training dataset, an adversary can randomize the assignment of trigger patterns from different attacks across training samples. This reduces the poisoning rate of individual attacks and simultaneously increases the uncertainty of poisoned samples. We select candidate triggers of BadNets, Blend, IAB and WaNet. During poisoning, we randomly assign one trigger to each selected sample. We use ∆H = Hbng − 0.5 Hpoi to show the gap between the mean uncertainty values of benign and 0 poisoned samples, where Blend Trigger Randomization the uncertainty value is −0.5 normalized to the range 0 5 10 15 20 25 30 35 40 Epoch 0 − 1 at each epoch. Compared to Blend atFigure 17. Comparing learning tack, trigger randomizaprocedures of Blend and trigger tion makes poisoned samrandomization attacks. ples converge more slowly, therefore showing a smaller uncertainty distance to benign samples (cf. Figure 17). Nevertheless, our defense remains highly robust against the trigger randomization backdoor (cf. Table 17), showing strong adaptability.
ASR [%]
ACC (BadNets) ACC (Blend) ACC (IAB) ACC (WaNet)
Defense No-Defense ABCS
E.4. Trigger Randomization Attack
20
90
Attack
9
Class
Figure 16. Poisoning Single Source Class.
E.3. Constructing Low-Redundancy Training Datasets Given adaptive adversaries aware of the coreset selection mechanism via C ENT criterion, they may first extract a coreset with a selection ratio of 50 % and subsequently poison it before releasing it as the final training dataset. Due to the high informativeness of this coreset, applying anti-backdoor coreset selection becomes challenging due to the inherent trade-off between security and utility in the poisoned data. In Table 16, we additionally evaluate ABCS under such coreset-level poisoning attacks. ABCS remains effective in identifying backdoor-free coresets, yielding significantly lower ASR than No-Defense, although maintaining suffi-
Table 17. Defending trigger randomization attack for CIFAR10. Defense
No-Defense ABCS (Ours)
19
ACC (↑)
94.63±0.08 94.10±0.12
ASR (↓) BadNets
Blend
IAB
WaNet
99.92±0.04 1.40±0.34
91.30±0.96 0.47±0.08
99.66±0.04 0.91±0.38
75.60±1.37 2.06±0.27
Anti-Backdoor Coreset Selection via Cumulative Entropy Table 18. Investigating robustness of ABCS against adaptive attack using label randomization. Attack
Defense
50 %
30 % ASR (↓)
ACC (↑)
ASR (↓)
ACC (↑)
ASR (↓)
BadNets
No-Defense ABCS (ours)
95.09±0.14 93.80±0.26
73.24±2.63 16.63±2.90
94.78±0.07 93.94±0.26
91.95±0.82 6.63±2.90
94.64±0.12 94.25±0.44
99.26±1.43 3.36±0.93
Blend
No-Defense ABCS (ours)
94.61±0.20 93.22±0.35
64.68±4.81 9.74±6.99
94.88±0.15 93.13±0.24
85.06±0.43 3.92±2.35
94.72±0.28 94.13±0.30
94.98±0.34 2.36±0.60
IAB
No-Defense ABCS (ours)
94.84±0.06 93.30±0.50
71.25±1.72 30.03±4.80
95.07±0.33 93.51±0.48
89.46±1.60 9.82±1.35
94.82±0.25 94.73±0.14
98.76±0.85 5.74±0.56
WaNet
No-Defense ABCS (ours)
94.43±0.19 93.21±0.49
78.56±0.25 8.69±1.71
94.46±0.21 93.75±0.26
80.85±2.15 3.38±0.74
94.43±0.46 93.97±0.22
90.47±1.61 3.29±0.28
ISSBA
No-Defense ABCS (ours)
94.51±0.06 94.16±0.54
65.94±2.86 18.65±1.76
94.37±0.11 93.91±0.15
89.43±2.96 3.72±0.56
94.49±0.21 94.34±0.33
94.08±1.87 1.37±0.13
LF
No-Defense ABCS (ours)
94.87±0.01 93.67±0.53
71.36±3.62 23.99±4.97
94.89±0.05 93.53±0.28
88.00±1.72 11.45±3.64
94.77±0.15 94.55±0.58
97.00±0.24 8.93±2.53
A-Blend
No-Defense ABCS (ours)
94.32±0.05 93.12±0.33
33.14±3.09 15.69±2.53
94.33±0.28 93.14±0.11
34.35±2.42 15.89±4.86
94.80±0.13 93.62±0.34
65.61±4.74 13.94±3.87
E.5. Label Randomization for Poisoned Samples
noise raises the prediction uncertainty of poisoned data. As shown in Figure 18, increasing the random-labeling ratio (10 %, 30 %, 50 %) slows down the convergence of poisoned samples, often yielding ∆H < 0 within the first 30 epochs. Furthermore, we investigate the orthogonality (Zhang et al., 2024), that is, the angle between gradient directions of benign and poisoned samples. In the loss landscape, poisoned samples exhibit gradients highly orthogonal to clean samples, enabling a more effective learning of the backdoor. With increasing the random-labeling ratio, the orthogonality decreases, showing again the effect of reducing the dissimilarity of benign and poisonous samples in the learning behavior. In Table 18, we report results across all attacks and show a trade-off: while larger random-labeling ratios increase uncertainty for poisoned samples, they simultaneously reduce the ASR. Under moderate noise (e.g., 10 %), ABCS consistently reduces ASR, demonstrating a strong robustness even against a highly knowledgeable adversary.
Training on datasets with randomly assigned label has been shown to increase learning difficulty (Natarajan et al., 2013), making the optimization process more stochastic due to the label randomization and thus inducing higher predictive uncertainty (Huang et al., 2022b; Köhler et al., 2019). Assuming an adversary knows the design of ABCS, an adaptive strategy would aim to increase the uncertainty of poisoned samples and slow their convergence, thereby hindering the defense. To evaluate this scenario, we assess our method under poisoning schemes that introduce label randomization to varying fractions of poisoned data. For example, with a poisoning ratio ρ = 5 %, a random-labeling ratio of 30 % means that 30 % of the generated poisoned samples receive randomized labels, while the remaining 3.5 % of poisoned samples maintain the backdoor target label. Entropy Difference
Orthogonality 65
Orthogonality
0.5
∆H
10 %
ACC (↑)
0 −0.5 −1
F. Empirical Analysis of Coreset Coverage
60 55 50
None 30 %
45 0
5 10 15 20 25 30 35 40
Epoch
0
In this section, we further examine how well variant-selected coresets cover the full clean training dataset D under both clean and poisoned conditions. As a clean baseline with strong natural performance, we first train a ResNet18 on the complete clean training set of CIFAR10. A coreset that exhibits high visual and quantitative similarity to the full clean dataset D reflects a correspondingly high level of information coverage.
10 % 50 %
5 10 15 20 25 30 35 40
Epoch
Figure 18. Impact of randomized label on Blend attack.
We first use the Blend attack to visualize how random labeling affects model learning. Since coreset selection relies on the uncertainty gap between benign and poisoned data, we measure the intermediate entropy difference as ∆H = Hbng − Hpoi , where Hbng and Hpoi denote the mean 0–1 normalized entropy at each epoch of benign and poisoned samples. A value ∆H > 0 suggests faster learning of poisoned samples due to lower entropy, whereas ∆H ≤ 0 indicates increased difficulty for our defense, as
Latent space visualization. We extract the feature-space representations from the penultimate layer of ResNet18, i.e., the activation vectors immediately preceding the fully connected layer, and visualize their 2-D distributions using tSNE (van der Maaten & Hinton, 2008). Compared with the full clean dataset D (cf. Figure 19a, top), selected coresets 20
Anti-Backdoor Coreset Selection via Cumulative Entropy Full Clean Set D
BadNets
IAB
Blend
WaNet
Class 0 Class 1 Class 2 Class 3 Class 4 Class 5 Class 6 Class 7 Class 8 Class 9 Coreset of Clean Set D
(b) Selected Coresets from Individual D̃
(a) Clean Train Set
Figure 19. T-SNE visualization of data distribution of full clean training set D and its coreset (Figure 19a), and other coresets selected from individual poisoned training sets D̃ (Figure 19b).
from various poisoned datasets shown in Figure 19b exhibit strong similarity in their global geometric structure. Aside from the easy classes 1 and 8 (cf. Figure 6), each remaining class preserves a cluster profile closely resembling that of the full dataset. This property consistently appears both in coresets derived from poisoned datasets and in those taken from the full clean dataset (cf. Figure 19a, bottom).
strong local geometric coverage relative to the full clean dataset. MMD, in turn, is non-negative and attains 0 only for identical distributions and, thus, the larger the value, the lower the global similarity. Table 19 summarizes the similarity of the selected coresets with respect to the full clean dataset. Across all poisoning attacks, ABCS yields coresets with p95-NN and MMD under 0.015, indicating that coreset samples lie very close to full-dataset samples and that the coreset distribution closely matches the full latent feature space distribution. These small values demonstrate both local geometric coverage and high distributional similarity of the coreset relative to the full clean dataset. Although some variance appears across classes, most of them have p95-NN and MMD values under 0.02. Note that class-1 (that is comparably easy to learn) tolerates higher dissimilarity from its full class distribution, whereas class-3 (which is more difficult to learn) benefits from a larger data portion, ensuring adequate coverage of the full class and, thus, preserving the natural performance.
Quantitative evaluation. Furthermore, we conduct a quantitative evaluation along two dimensions: local geometric coverage and global distributional similarity. For the former aspect, we measure the 95th -percentile NearestNeighbor distance (p95-NN) (Clark & Evans, 1954), using cosine distance as the metric. For the latter aspect, we compute the Maximum Mean Discrepancy (MMD) (Borgwardt et al., 2006; Goodfellow et al., 2015), where inputs are first l2 -normalized and an RBF kernel is applied with bandwidth chosen via the median heuristic. Both metrics are implemented using the Scikit-learn toolkit with the default settings (Pedregosa et al., 2011). Notably, the p95-NN metric yields values in [0, 2], where values close to 0 indicate 21
Anti-Backdoor Coreset Selection via Cumulative Entropy Table 19. Similarity of the full dataset and the coresets extracted from poisoned CIFAR10. 95th -Percentile Nearest-Neighbor Distance
Dataset
Maximum Mean Discrepancy
BadNets
Blend
IAB
WaNet
BadNets
Blend
IAB
WaNet
Full
0.0112±0.0002
0.0114±0.0003
0.0116±0.0002
0.0113±0.0003
0.0121±0.0004
0.0130±0.0006
0.0142±0.0012
0.0126±0.0005
Class 0 Class 1 Class 2 Class 3 Class 4 Class 5 Class 6 Class 7 Class 8 Class 9
0.0091±0.0004 0.0107±0.0004 0.0114±0.0007 0.0084±0.0004 0.0101±0.0011 0.0111±0.0004 0.0122±0.0004 0.0116±0.0008 0.0142±0.0008 0.0113±0.0007
0.0102±0.0026 0.0111±0.0004 0.0119±0.0008 0.0091±0.0003 0.0103±0.0003 0.0112±0.0013 0.0110±0.0008 0.0113±0.0007 0.0142±0.0016 0.0113±0.0004
0.0104±0.0012 0.0104±0.0006 0.0126±0.0010 0.0076±0.0008 0.0110±0.0016 0.0121±0.0010 0.0120±0.0014 0.0120±0.0004 0.0141±0.0016 0.0108±0.0004
0.0108±0.0002 0.0094±0.0006 0.0120±0.0011 0.0078±0.0001 0.0101±0.0010 0.0109±0.0012 0.0116±0.0011 0.0111±0.0007 0.0149±0.0017 0.0114±0.0008
0.0015±0.0008 0.0297±0.0015 0.0022±0.0006 0.0007±0.0003 0.0039±0.0014 0.0033±0.0011 0.0068±0.0003 0.0161±0.0029 0.0199±0.0025 0.0140±0.0012
0.0027±0.0015 0.0334±0.0036 0.0035±0.0009 0.0008±0.0001 0.0037±0.0007 0.0035±0.0009 0.0039±0.0009 0.0134±0.0025 0.0228±0.0050 0.0116±0.0015
0.0023±0.0010 0.0331±0.0062 0.0032±0.0012 0.0004±0.0002 0.0045±0.0010 0.0037±0.0005 0.0035±0.0006 0.0169±0.0032 0.0214±0.0075 0.0111±0.0020
0.0019±0.0002 0.0096±0.0029 0.0027±0.0006 0.0002±0.0001 0.0023±0.0006 0.0015±0.0005 0.0036±0.0007 0.0135±0.0035 0.0193±0.0028 0.0078±0.0026
22