Conceptio › Archive › arXiv CS
arXiv CSopen access

PASTA: A Patch-Agnostic Twofold-Stealthy Backdoor Attack on Vision Transformers

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

1

PASTA: A Patch-Agnostic Twofold-Stealthy Backdoor Attack on Vision Transformers

arXiv:2604.20047v1 [cs.CV] 21 Apr 2026

Dazhuang Liu, Yanqi Qiao, Rui Wang, Kaitai Liang, and Georgios Smaragdakis

Abstract—Vision Transformers (ViTs) have achieved remarkable success across various vision tasks, yet recent studies reveal that ViTs are vulnerable to backdoor attacks. Existing patch-wise attacks against ViTs focus on a single trigger activation location during inference to maximize trigger attention. However, these attacks fail to fully exploit the characteristic of self-attention mechanism that captures long-range dependencies across patches. We stand for the first work to observe that a patch-wise trigger delivers high attack effectiveness when activating the backdoor across neighboring patches, a unique phenomenon in ViTs we term Trigger Radiating Effect (TRE). Additionally, we find that inserting patch-wise triggers in an inter-patch manner synergistically enhances TRE compared to single-patch insertion during backdoor training. Moreover, existing ViT-specific attacks that maximize model attention on triggers compromise both visual and attention stealthiness, making them vulnerable to human and machine inspection. Building upon the above insights, we propose a twofold stealthy patch-wise backdoor attack in the pixel and attention domains, dubbed PASTA, with a new attack payload, where an attacker can activate the backdoor with the trigger in an arbitrary patch during inference. To achieve the payload, we first propose a multi-location trigger insertion strategy to enhance TRE synergistically. However, achieving our payload while maintaining twofold stealthiness remains a significant challenge, as we observe that the TRE is significantly undermined against stealthy trigger designs. Therefore, we formulate our backdoor attack as a bi-level optimization problem and propose an adaptive backdoor learning framework to solve it. Specifically, our learning framework allows both backdoor model parameters and the trigger to gradually adapt to each other’s updates, mitigating convergence to local optima caused by their non-separability in loss terms. Extensive experiments comprehensively showcase that PASTA achieves attack effectiveness of 99.13% across arbitrary patches on average, superior visual stealthiness (144.43× improvement) and attention stealthiness (18.68× improvement), and better attack robustness (2.79× enhancement) under our payload against state-of-the-art ViT-specific defenses compared to both CNN- and ViT-specific attacks across four public datasets. Index Terms—Backdoor attack, stealthy attack, vision transformer, bi-level optimization.

I. I NTRODUCTION Deep Neural Networks (DNNs) achieved huge success in Computer Vision (CV) tasks, such as image classification [1], [2], object tracking [3], [4], object detection [5], [6], and facial recognition [7]. Empowered by the self-attention mechanism, Vision Transformers (ViTs) [8] have challenged the long-standing dominance of Convolutional Neural Networks Dazhuang Liu, Yanqi Qiao, Rui Wang and Georgios Smaragdakis are with Delft University of Technology, Delft, the Netherlands. Kaitai Liang is with the University of Turku, Turku, Finland and Delft University of Technology, Delft, the Netherlands.

(CNNs) [9] in many CV fields. Specifically, ViTs process input images by dividing them into a sequence of patches. Then, the self-attention mechanism weights the relevance of each patch to others, enabling ViTs to capture long-range dependencies in those patches effectively. However, similar to CNNs, ViTs has been proven to be vulnerable to backdoor attacks [10], [11], [12], [13], [14], [15], [16]. The general goal of a backdoor attack is to implant hidden malicious behaviors into DNNs by inserting a trigger into clean samples. This goal causes the victim model to produce incorrect, attacker-desired, predictions on poisoned data during inference, while behaving normally on clean data. Since training ViTs with large-scale parameters demands high computational resources, users often outsource the training process or fine-tune pre-trained models downloaded from the Internet for downstream tasks. However, this practice creates more opportunities for attackers to inject backdoors into the models of benign users. Recently, backdoor attacks have developed with diverse trigger designs, including patch-based triggers [10], [17], blended triggers [11], more advanced sample-specific triggers [18], [19], [20], [21] and frequency triggers [22], [23], [24], [25]. Yuan et al. [13] show that patch-wise triggers are more effective than blended triggers against ViTs, prompting ViT-specific attacks [13], [14], [15] to develop optimized patch-wise triggers that maximize model attention on them. However, these attacks rely on the pre-defined trigger insertion location used in backdoor training to activate the backdoor during inference. Moreover, neither CNN- nor ViT-specific attacks investigate how convolutional filters and self-attention mechanisms impact attack effectiveness across different trigger activation locations (TALs) during inference. This research gap prompts an intriguing question: Q1: Is the attack effectiveness of patch-based triggers sensitive to different TALs during inference in CNNs and ViTs based on their inherent architectural differences? We systematically investigate how attack effectiveness of patch-based triggers is impacted by TALs during backdoor inference in ViTs and CNNs. Our empirical study (see Figure 1(a)-(d)) shows that attack effectiveness is sensitive to TALs in CNNs. In contrast, in ViTs, a patch-wise trigger1 inserted in a pre-defined patch location during backdoor training remains effective even when inserted in neighboring patches during inference (see Figure 1(e)-(h)) – a phenomenon we term the Trigger Radiating Effect (TRE) on attack ef1 Patch-wise triggers are a specific type of patch-based trigger in which the trigger size matches the patch size used in ViTs. We limit the scope of our research on patch-wise triggers in ViTs.

0000–0000/00$00.00 © 2026 IEEE

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

fectiveness. TRE can be simply quantified as the average attack effectiveness across all patches activated by a patchwise trigger. Furthermore, we reveal that a larger patch-wise trigger perturbation in ViTs introduces a stronger TRE, while a smaller perturbation yields an opposite effect (Figure 1(i)-(l)). The above findings raise a further question: Q2: Given that a stealthy patch-wise trigger exhibits a limited TRE on attack effectiveness in ViTs, can backdoor training with the trigger in multiple trigger insertion locations synergistically enhance the TRE? To further evaluate the synergistic impact of patch-wise triggers on TRE, we fine-tune a ViT model by inserting a patch-wise trigger pattern into one of the two pre-selected patch locations (i.e., an inter-patch manner) per poisoned sample. Our TRE results demonstrate that training a backdoor with a patch-wise trigger pattern inserted across multiple patch locations synergistically improves TRE compared to the single-patch insertion approach (Figure 1(m)-(p)). Besides, current ViT-specific attacks [13], [14], [15] fail to achieve twofold stealthiness, including visual and attention imperceptibility in pixel and attention spaces, respectively. Given that a large perturbation has a stronger TRE but undermines twofold stealthiness, a natural question then arises: Q3: How can a backdoor attack leverages a patch-wise trigger achieve twofold stealthiness while enabling high attack effectiveness across arbitrary patches (i.e., strong TRE) during inference? To answer the question, we introduce PASTA, a patchwise backdoor attack against ViTs that achieves a new attack payload where a backdoor can be activated across arbitrary patches with high attack effectiveness, while maintaining excellent visual and attention stealthiness (see Figure 2 for the workflow). First, we propose a Multi-location trigger Insertion Strategy (MIS) by inserting our trigger into a random location from a candidate set for each poisoned sample, in order to achieve strong TRE. Then, we consider twofold stealthy trigger design by minimizing the disparity of attention map between clean and poisoned samples, and constraining the visual imperceptibility by the l2 -norm of trigger perturbations. While it is challenging to achieve strong TRE and twofold stealthiness simultaneously, we formulate trigger and model optimization as a bi-level optimization problem to achieve all attack objectives. Since changes in model parameters and the trigger during optimization directly affect the loss value with respect to the other, we propose an adaptive backdoor training framework to effectively solve the problem. Specifically, we update the trigger and model parameters alternately for small iterations, allowing both variables to gradually adapt to each other’s minimal updates. By doing so, we reduce the correlational influence of two variables on loss values in nonseparable loss terms, preventing them from converging to local optima. Thanks to the newly proposed attack payload, the backdoor can be effectively activated during inference by our optimal patch-wise trigger across arbitrary patches, enabling the evasion of state-of-the-art backdoor defenses. Our main contributions are as follows: • We observe, for the first time, that patch-wise triggers have a stronger trigger radiating effect (TRE) on the attack

2

effectiveness of neighboring patches in ViTs than in CNNs, due to the self-attention mechanism. Also, inserting a patchwise trigger in an inter-patch manner during backdoor training synergistically enhances TRE. Based on these insights, we propose a multi-location trigger insertion strategy to deliver strong TRE. • We propose a new backdoor payload that evades backdoor defenses, enabling an attacker to activate the backdoor across arbitrary patches during inference. In addition, we consider visual and attention stealthiness to bypass human and machine inspection. We introduce PASTA, a visual and attentionstealthy patch-wise backdoor attack against ViTs under our payload. We formulate all attack objectives as a bi-level optimization problem and introduce an adaptive optimization framework to solve it effectively. • Our extensive experiments demonstrate that PASTA achieves superior attack effectiveness on all patches (99.13%), better visual stealthiness (144.43× improvement) and attention imperceptibility (18.68×) under l2 -norm and qualitative visualizations, 2.79× enhancement on attack robustness against 3 ViT-specific defenses, compared to 3 general and 3 ViTspecific backdoor attacks across 4 real-world datasets. We will make the source code publicly available upon the publication of this paper. We discuss the ethical considerations of this paper in Appendix A of supplementary material. II. BACKGROUND A. Vision Transformer The long-standing dominance of CNNs in computer vision has been fundamentally challenged by ViTs, pioneered by Dosovitskiy et al. [8]. ViTs adapt the original transformer architecture [26] for natural language processing (NLP) to the domain of computer vision. Given a vision transformer model f (·) and training dataset Dc = {(xi , yi )|xi ∈ RH×W ×C , yi ∈ Rκ } N i=1 , where N is the size of dataset, κ is the number of classes, H, W and C are the height, width and channels of an input x, and y is the ground-truth label. The input image x is divided into a sequence of H × W/p2 patches with the shape of p × p. These patches are flattened and linearly projected into embedding vectors, similar to word embeddings in NLP. Moreover, a classification token is added to the head of the above embedding vectors, forming the input token sequence as T = {tcls , t1 , t2 , · · · , tH×W/p2 }. The core component of the ViTs is the self-attention mechanism. Selfattention mechanism allows ViTs to weigh the importance of different patches in relation to each other, effectively capturing long-range dependencies and global context. Each token is used to perform attention map calculation by multi-head selfattention (MSA) module as follows:   T WQ (T WK )T √ Attention(T ) = Softmax T WV , (1) d where d is the dimension of the query Q and the key K; WQ , WK and WV are learnable weights of the query, key and value V , respectively. MSA enhances self-attention mechanism by performing it multiple times in parallel with distinct learned linear projections (heads), allowing ViTs focus on diverse aspects of

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

3

TABLE I: Crucial attack attributes among PASTA and other CNN- and ViT-specific backdoor attacks. Attacks

ViT-specific

Patch-aware

Optimization

Visual Imperceptibility

Attn. Stealthiness

Arbitrary Patch Activation

BadNets [10] WaNet [12] DBIA [15] BadViT [13] TrojViT [14] AIBA [16] Narcissus [27] BAVT [28] HCB [29] BELT [30] LADDER [24] LIRA [19] PASTA (Ours)

A A ✓ ✓ ✓ ✓ A A A A A A ✓

A ✘ ✓ ✓ ✓ ✓ A A A A ✘ ✘ ✓

N/A N/A SOP SOP SOP BOP SOP SOP N/A SOP SOP BOP BOP

✘ ✓ ✘ ✓ ✘ ✘ ✘ ✓ ✓ ✘ ✓ ✓ ✓

✘ ✘ ✘ ✘ ✘ ✓ ✘ ✘ ✘ ✘ ✘ ✘ ✓

A ✘ A A A ✘ A A A A ✘ ✘ ✓

✓: Yes; ✘: No; A: Applicable; N/A: Not Applicable; SOP: Single-level Optimization Problem; BOP: Bi-level Optimization Problem.

the input simultaneously. The classification result is derived from multiple MSA attention calculations through a multilayer perceptron (MLP). Compared to CNNs, ViTs excel at capturing global context and long-range dependencies, offering a weaker inductive bias. B. Backdoor Attacks We consider backdoor attacks against ViTs on image classification. Let fθ : I → Rκ be an image classifier parameterized with θ that maps an image x from input space I ⊆ [0, 1]H×W ×C to an output class. The parameters θ of the classifier are learned using a training dataset Dc = {(xi , yi )|xi ∈ I, yi ∈ Rκ }N i=1 . In a standard backdoor attack, the attacker selects a subset of Dc with ratio ρ as the poisoned dataset Dbd , and transforms it by the trigger injection function T and target label function η. Given an image x and its true class y from Dbd , the commonly used trigger injection function T and target label function η are defined with a scaling parameter m ∈ [0, 1] and a trigger pattern t as follows: x′ = T (x, m, t) = x·(1−m)+t·m,

y ′ = η(y) = ytgt , (2)

where ytgt is the target class. Under empirical risk minimization, a typical attack aims to inject backdoors into the classifier f by learning θ with both Dc and Dbd so that the classifier misclassifies the poisoned data into the target class while behaving normally on clean data. The optimization problem is defined as follows: X X min L(fθ (x), y) + L(fθ (T (x)), η(y)), (3) θ

(x,y)∈Dc

(x,y)∈Dbd

where L represents the cross-entropy loss. We will describe our patch-wise trigger insertion function in Equations (5) and (6). The key notations used in this paper are summarized in Table VIII in the supplementary material. III. R ELATED W ORK A. Backdoor Attacks & Defenses against CNNs The first backdoor attack against CNNs is proposed by Gu et al. [10]. Since then, various attacks have been proposed to improve stealthiness at both the input and hidden feature levels. To bypass human inspection, some works [31], [32],

[19], [33], [12], [34], [22], [23], [35], [36] design triggers with imperceptible perturbations. For example, Barni et al. [31] use sinusoidal signals as triggers which results in only a slight varying backgrounds on the poisoned images. Liu et al. [32] utilize natural reflection as triggers for backdoor injection in order to disguise triggers as natural light-reflection. Li et al. [33] leverage a CNN-based image steganography technique to hide an attacker-specified string into images as sample-specific triggers. Wang et al. [22] handcraft two single frequency bands with fixed perturbations as triggers. Besides visual stealthiness, several works [37], [38], [19], [20], [21] investigate the stealthiness in latent feature space. Doan et al. [37] design a trigger generator to constrain the similarity of hidden features between clean and poisoned data via Wasserstein regularization. To improve the trigger stealthiness, Zhao et al. [20] learn a generator to constrain the latent layers, which makes triggers more invisible in both input and latent feature space. Additionally, some studies focus on different aspects of attacks. For example, Lv et al. [39] propose an attack without using original training/testing dataset. Zeng et al. [27] conduct clean-label backdoor attacks using knowledge of target class samples and out-of-distribution data. Lan et al. [40] introduce a stealthy and practical backdoor attack on speech recognition tasks. Abad et al. [41] propose a stealthy attack against spiking neural networks. Zhang et al. [42] present the first backdoor attack for model merging scenario. Backdoor defenses include detection [43], [44], [45], [46], [35] and defensive [47], [48], [49], [50], [51], [52] mechanisms. Typical detection methods include STRIP [43], which deliberately perturbs clean inputs to identify potential backdoored CNN models during inference. Spectral Signature [45] detects outliers using latent feature representations, while Zeng et al. [35] propose a method that discriminates between clean and poisoned data in the frequency domain using supervised learning. Image preprocessing-based methods [52], [22], [53] have recently been explored to remove backdoors using techniques such as transformations and compression. Defensive methods aim to detect potential backdoor attacks but also to actively mitigate their effectiveness. For instance, fine-pruning [47] reduces the impact of backdoors by trimming dormant neurons in the last convolution layer, based on the minimum activation values of clean inputs. Neural Cleanse [48] leverages reverse engineering to reconstruct potential triggers for each

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

target label and eventually renders the backdoor ineffective by retraining patches strategy. Neural Attention Distillation [51] utilizes a “teacher” model to guide the fine-tuning of the backdoored “student” network to erase backdoor triggers. Recently, several state-of-the-art backdoor defenses have been proposed. For example, Gao et al. [54] introduce a trainingtime defense that separates training data into clean and poisoned subsets. Zhu et al. [55] purify poisoned models by incorporating a learnable neural polarizer as an intermediate layer. Shi et al. [56] mitigate backdoor attacks through zeroshot image purification. B. Backdoor Attacks and Defenses against ViTs With the dominance of transformer-based architectures in computer vision tasks, some researchers have evaluated the robustness of ViTs against backdoor attacks. Yuan et al. [13] observe that ViTs are more sensitive to patch-wise triggers than CNNs due to the self-attention mechanism and design a universal patch-wise trigger to catch the model attention. Similarly, Zheng et al. [14] utilize a patch-wise trigger to build a Trojan composed of vulnerable bits on ViT parameters stored in DRAM memory. Lv et al. [15] propose a data-free backdoor attack, which uses the poisoned surrogate dataset to generate triggers and inject the backdoor into the model. To enhance the stealthiness at the attention level, Wang et al. [16] constrain the trigger perturbation to achieve visual imperceptibility and implant the trigger into the image’s focal regions to ensure attention imperceptibility. To enhance the robustness of ViTs, researchers have developed several defensive mechanisms against backdoor attacks. Subramanya et al. [28] propose a defense in the inference stage by adding a black patch around the location with the highest attention value. Additionally, Doan et al. [57] introduce an effective method for ViTs to defend against both patch-based and blending-based triggers via patch processing, including Patch Drop and Patch Shuffle. However, none of the existing defensive strategies for ViTs consider the trigger activation via arbitrary patches during inference. In this paper, we challenge the conventional backdoor requirement that triggers must be activated in a specific location. Compared to existing attack methods on ViTs, our proposed approach demonstrates superior effectiveness across all patch locations against both CNN-specific and ViT-specific defenses while maintaining both visual and attention stealthiness. Table I summarizes SOTA backdoor attacks based on various attack attributes. Section VI provides experimental comparisons. IV. O BSERVATION : CNN S VS . V I T S Before delving into our own attack, we investigate how Trigger Activation Locations (TALs) during inference affect attack effectiveness of patch-based attacks in CNNs and ViTs. This study is motivated by the inherent characteristics of different model architectures in feature extraction: the convolutional filters in CNNs can only capture local image features with positional information; in contrast, the self-attention mechanism in ViTs has the ability to learn the long-range

4

dependencies across different patches and global context. Such a difference in mechanism can impact attack effectiveness on neighboring patches when these patches are activated by patch-based triggers in CNNs and ViTs. To understand attack effectiveness of patch-based triggers on TRE in these models, we design a series of experiments to examine key factors influencing TRE, including trigger insertion locations during backdoor training, trigger insertion methods, and the magnitude of trigger perturbations. Trigger Radiating Effect (TRE) in CNNs and ViTs. We first investigate whether TRE exists in CNNs and ViTs through two trigger insertion methods, including (1) REP (replace): the original pixels in the patch are replaced with the trigger pattern; and (2) SUP (superimpose): the trigger pattern perturbation is directly added onto the original pixel values in the patch. In CNNs, we verify TRE on CIFAR-10 [58] dataset (32×32 resolution) using a conventional CNN model with 3 convolutional layers and 2 dense layers. We randomly initialize two patch-based trigger patterns (3×3 in size) with two different l2 -norms of trigger perturbations: 3 for REP and 0.2 for SUP. These triggers are inserted into two different locations (top-left and center), and the poisoned model is trained separately for each configuration. For ViTs, we assess TRE on ImageNet [2] dataset (224×224 resolution) on a pretrained ViT model with a patch size of 16×16. Following a similar approach, we initialize two 16×16 patch-wise trigger patterns with l2 -norms of 20 for REP and 1 for SUP, and insert them at the same locations as used in CNNs. We then finetune the poisoned model separately for each case. To better study the attack effectiveness on different TALs in CNNs and ViTs, we reflect TRE as follows: P T RE ≜

n i=1 ASRi

n

,

(4)

where ASRi is the attack success rate when the patch-based backdoor is activated on the i-th patch during inference, n is the total number of TALs. For CNNs on CIFAR-10, n = (32 − 3 + 1)2 = 900 and n = 224 × 224/162 = 196 for ViTs on ImageNet. Specifically, TAL is shifted with a stride of 1 in CNNs and step size of 16 in ViTs. We present TRE results derived from two trigger insertion methods, each under two trigger insertion locations and two model architectures (CNN and ViT), as shown in Figures 1(a)–(h). In Figures 1(a)-(d), high attack effectiveness (i.e., the red dot) is achieved only when TAL is exactly the same as the trigger insertion location during backdoor training. Once the TAL is shifted by even a few pixels, attack effectiveness drops drastically (colored blue), resulting in a low TRE of around 10% in all cases. The above TRE results confirm that TRE does not exist in CNN architectures, in spite of the trigger insertion methods and locations. We could draw a similar conclusion that advanced CNN architectures only exhibit very limited TRE on the large-scale dataset (see Appendix C of supplementary material) as well. In contrast, SUP and REP triggers in ViTs achieve significantly higher TREs than CNNs, reaching 79.76% and 99.30%, respectively, when these triggers are inserted at the center (see Figures 1(e) and (g)). These results demonstrate that TRE does exist in ViTs. Moreover, we find that SUP and

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

5

(a)SUP(l2 :0.2, TRE:10.83)

(b)SUP(l2 :0.2, TRE:10.37)

(c)REP(l2 :3.0, TRE:10.58)

(d)REP(l2 :3.0, TRE:10.30)

(e)SUP(l2 :1.0, TRE:79.76)

(f) SUP(l2:1.0, TRE:13.05)

(i)SUP(l2 :0.5, TRE: 57.01)

(j)SUP(l2 :1.0, TRE:79.76)

(k)SUP(l2 :2.0, TRE: 94.96)

(l)SUP(l2 :4.0, TRE:98.89)

(m)SUP(l2 :1.0, (n)SUP(l2 :1.0, TRE:91.11) TRE:86.66)

(g)REP(l2 :20.0, (h)REP(l2 :20.0, TRE:99.30) TRE:64.81)

(o)SUP(l2 :1.0, TRE: 92.63)

(p)SUP(l2 :1.0, TRE:68.92)

Fig. 1: Visualization of TRE heatmaps under different attack settings in CNNs and ViTs. The trigger insertion method, magnitude of trigger perturbation and TRE(%) are denoted below each heatmap. (a)-(d): TRE (%) against CNNs; (e)-(h): TRE (%) against ViTs; (i)-(l): TRE (%) against various trigger perturbations; (m)-(p): TRE (%) against multiple trigger insertion locations.

REP triggers at four corner patches exhibit significantly lower attack effectiveness than other activation patches. Notably, the SUP trigger inserted at the top-left delivers high attack effectiveness only in its corresponding row and column (see Figure 1(f)). Although both REP- and SUP-based trigger insertion methods produce a clear TRE on ViTs, REP triggers cannot provide visual imperceptibility. Therefore, this work focuses on the SUP-based trigger insertion approach to achieve twofold stealthiness. Impact of Trigger Perturbations on TRE. Given that TRE exists in SUP triggers within ViTs, we further investigate how TRE varies with different perturbation magnitudes of the SUP trigger. We fine-tune a pre-trained ViT model on ImageNet by inserting a fixed trigger pattern at the image center with varying l2 -norms: 0.5, 1, 2, and 4. As shown in Figure 1(i)–(l), TRE increases from 67.01% to 98.89% as the trigger perturbation magnitude grows from 0.5 to 4 (i.e., a decrease in trigger imperceptibility). This demonstrates that larger trigger perturbations lead to stronger TREs. Furthermore, these TRE results corroborate the previous observation that SUP triggers inserted at the center are ineffective in activating backdoors at four corner patches. Synergistic Effect of Multiple Trigger Insertion Locations on TRE. While small trigger perturbations degrade TRE, we further explore whether a synergistic effect exists to enhance TRE by alternately inserting a patch-wise trigger pattern into two patch locations during backdoor training. In the following experiments, we use a patch-wise trigger pattern from the previous experiment, with an l2 -norm of 1. We fine-tune pretrained ViT models on ImageNet under the following trigger insertion location sets: (3,3) and (12,12); (5,5) and (10,10); (6,6) and (9,9); (0,0) and (13,13). For each poisoned sample, the trigger pattern is added to one of the two possible locations from a specific location set (see Section V-C). In Figures 1(m)–(o), we observe that backdoor training of the patch-wise trigger with two insertion locations helps to enhance TRE by an average of 10.37% compared to the single-patch insertion (79.76% in Figure 1(j)). The results confirm the existence of a synergistic effect on TRE when alternatively learning a

patch-wise trigger with multiple insertion locations in ViTs. Interestingly, inserting the patch-wise trigger into two corners does not significantly enhance TRE, only achieving 68.82% in Figure 1(p). Building upon these insights, we further propose a novel backdoor payload activated by a patch-wise trigger, achieving high attack success rates across arbitrary patches in poisoned images. V. ATTACK M ETHODOLOGY In this section, we first introduce our patch-wise trigger insertion method for ViTs, followed by the formulation of our bilevel optimization-based backdoor attack problem. After that, we present the workflow of our adaptive backdoor training. A. Patch-wise Trigger Injection As demonstrated by Yuan et al. [13] and further supported in Section IV, patch-wise triggers are highly effective in backdooring ViTs, even inducing strong TRE on neighboring patches. However, such REP-based triggers [13], [15], [14] cannot provide visual and attention imperceptibility. Hence, we propose a SUP-based patch-wise trigger insertion function T as follows: x′ = T (x, t, Mi ) = x + Mi · t,

(5)

where Mi ∈ {0, 1}H×W is a binary mask, i is the patch index of a sequence of n input patches of an image, the trigger mask is defined as follows: ( Mi =

1, 0,

if pixel ∈ i-th patch otherwise

(6)

We use the same target label function as in Equation (2), i.e., ensuring all poisoned samples are misclassified to ytgt . B. Threat Model We consider the same white-box threat model as in prior works [12], [19], [59], [16], [60] against CNNs and ViTs,

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

𝑺𝟐𝒄𝒕𝒓

𝑺𝟑𝒄𝒕𝒓

𝑺𝟒𝒄𝒕𝒓

𝑺𝟓𝒄𝒕𝒓

{𝑆)*+ }

C

Upper Level (Model Optimization)

MIS(𝑺𝟏𝒄𝒕𝒓 , 𝑺𝟏𝒄𝒐𝒓 ) Hierarchical Sampler

Clean data 𝒟"

Adaptive Backdoor Training E Framework

Bi-level Optimization Problem

A Multi-location Trigger Insertion Strategy 𝑺𝟏𝒄𝒕𝒓

6

, . / 𝑆)*+ 𝑆)*+ 𝑆)*+ 𝑆)*+

𝑡,

Candidate Trigger Insertion Locations in 𝑺𝒄𝒕𝒓 and 𝑺𝒄𝒐𝒓

𝜃12,

𝜃

……

……

Sampling Locations from MIS

𝜃∗

Small Gradient Trajectory 𝑡12.

……

𝑡 𝜃12-

B

𝑡0

𝜃0

𝑡12𝑡12, 𝑡∗

Trigger Insertion and Poisoned Data 𝓓𝒃𝒅 Generation under 𝒕

Twofold Stealthiness Visual Stealthiness ℒ𝒗𝒊𝒔

Clean Sample

Lower Level (Trigger Optimization)

H

New Payload: Max. Trigger Radiating Effect to Activate PASTA Backdoor across Arbitrary Patches Any TAL

Model

Trigger

𝑡∗, 𝜃∗

Attention on Clean Sample

…… 𝜃∗ Identical

Identical Poisoned Sample

F

Poisoned Sample

…

Any combination of TALs …

Clean Sample

Attention Stealthiness ℒ𝒂𝒕𝒕𝒏

D

Less Local Optima

Attention on Poisoned Sample

G

ASR∝ 𝟏𝟎𝟎% Insert invisible 𝑡 ∗ to any trigger activation locations (TALs) marked

during inference

Fig. 2: The workflow of PASTA. A - B : We propose a multi-location trigger insertion strategy (MIS) to assign trigger insertion locations per sample, and poison them under a patch-wise trigger t, producing the poisoned dataset Dbd . C - D : The upper- and lower-level tasks optimize the model parameters θ and the trigger t, respectively. E : Our adaptive backdoor training framework alternately optimize two tasks with small gradient steps, reducing the local optima in the bi-level optimization problem. F - G : We evaluate visual stealthiness and attention stealthiness during optimization. H : After adaptive backdoor training, the optimal t can be inserted into any TAL F on the optimal θ ∗ to activate our backdoor, achieving near-perfect attack effectiveness.

where the adversary, i.e., a malicious model provider, has complete control over the victim model architecture, parameters and training process. Given a model architecture and a dataset, the adversary trains a poisoned model by injecting an optimized trigger pattern known only to itself, and subsequently publishes the compromised model as open source for victim users to download and deploy in their applications. Such a threat model is practical, as leveraging pretrained ViT models for downstream tasks [61] is common in the literature, given the high computational cost of training ViTs from scratch [8]. Meanwhile, the users may leverage backdoor defenses and attention inspection techniques to check whether the model is poisoned. During the inference phase, the adversary is able to activate the backdoor behavior of the victim model by inserting the trigger pattern into arbitrary patches of the input image. C. Problem Formulation Our backdoor attack aims to achieve twofold stealthiness including visual and attention imperceptibility, while maintaining strong TRE to deliver high attack effectiveness across arbitrary patches simultaneously. Below, we outline our attack objectives. Multi-location Trigger Insertion Strategy. Based on the observations in Section IV, we propose a multi-location trigger insertion strategy (MIS) to achieve our new attack payload, i.e., delivering high attack effectiveness across arbitrary TALs.

We first pre-define a series of candidate patch locations, and then sample one patch location from the candidates to insert the patch-wise trigger pattern for each poisoned sample. Specifically, we divide all patches into four quadrants and select the central patch from each quadrant, along with the center patch of the entire image, forming a candidate set Sctr . Additionally, we include the four corner patches as another candidate set Scor . We combine these sets to form the full location set S = Sctr ∪ Scor . Note that the impacted area of TREs varies significantly depending on whether the trigger insertion locations originate from Sctr (affecting 98% of the patches) or Scor (only 19%) (see Figures 1(o) and (p)). Therefore, uniform sampling applied directly to S impairs TRE. To solve this, we propose a hierarchical sampler in our strategy to select a specific patch index for trigger insertion for each poisoned sample. In particular, we treat Scor as a single element, i.e., {Scor } in S and uniformly sample a patch index i from S. Thus, ∀i ∈ Sctr , the probability of i-th patch being selected is |Sctr1 |+1 . If Scor is selected from S, we perform a secondary uniform sampling within it, so ∀i ∈ Scor , the probability that i-th patch being selected 1 becomes (|Sctr |+1)×|S . We denote sampling a patch index cor | i from S under our multi-location trigger insertion strategy as i ∼ MIS(Sctr , Scor ). Attention Imperceptibility. Given a ViT model f with L

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

attention layers and an index i (0 ≤ i ≤ n − 1) of a sequence of patches P of input image x , the attention score of each patch Pi in layer l (0 ≤ l ≤ L) reflects its relative importance compared to other patches. The attention map of x in l-th layer is computed as: Attnl←0 (x) = Attentionl−1 ◦ Attentionl−2 ◦ · · · ◦ Attention0 (x), where ◦ denotes the propagation of latent features through ViT layers. In ViT-specific backdoor attacks, poisoned samples often exhibit abnormal attention maps, focusing disproportionately on the trigger location rather than benign image features. This can serve as a discriminative cue for backdoor detection. To mitigate such anomalies of poisoned samples, we compute the attention disparity between x and its poisoned counterpart T (x) (transformed by our trigger insertion function T ) in l-th layer: dis attnl (x, T , f ) = Attnl←0 (T (x, t, Mi )) − Attnl←0 (x) 2 , where i ∼ MIS(Sctr , Scor ).

(7)

Unlike existing ViT-specific attacks [13], [14] that maximize the model attention on their triggers (causing abnormal attention maps as shown in Figure 4), we introduce a loss term that minimizes the disparity to achieve attention imperceptibility: P (8) Lattn = |Dbd |−1 x∈Dbd dis attnl (x, T , f ), Visual Stealthiness. We consider the invisibility of our trigger to bypass human inspection. Following the same philosophy as in previous CNN- and ViT-specific attacks [19], [20], [13], we introduce a loss term to constrain the pixel-domain disparity between clean and poisoned images injected by our trigger: P Lvis = |Dbd |−1 x∈Dbd ∥T (x, t, Mi ) − x∥2 , (9) s.t. i ∼ MIS(Sctr , Scor ). We use the l2 -norm to measure visual disparity since it provides a global estimate of how much the trigger modifies the original images. Moreover, the l2 -norm is a convex function, making it well-suited for integrating into loss functions and optimization via gradient descent. Attack Objectives Aggregation for Optimization. Intuitively, we can formulate all the above attack objectives into an optimization problem as follows: min Lc (θ) + Lbd (θ, t) + α1 Lattn (θ, t) + α2 Lvis (t), (10) θ,t where Lc = (x,y)∈Dc LCE (fθ (x), y) is the training loss reP garding benign task and Lbd = (x,y)∈Dbd LCE (fθ (T (x)), η(y)) denote training loss for backdoor tasks. This optimization problem assumes that model parameters θ of a ViT and a patch-wise trigger pattern t can be jointly optimized within an aggregated objective function. However, changes in θ or t during the optimization directly influence the loss value regarding the other – a phenomenon known as non-separability [62]. In fact, according to Verel et al. [63], simultaneously optimizing all attack objectives but ignoring the correlation of t and θ in non-separable loss terms (Lattn (θ, t) and Lbd (θ, t)) can significantly increase the number of local optima in the loss landscape of Equation (10), risk in finding sub-optimal θ and t. Therefore, such a formulation is ineffective in finding a practical θ and t that achieves all attack objectives. P

7

To address the optimization challenge, we formulate our backdoor attack as a constrained bi-level optimization problem: min L = Lc (θ) + Lbd (θ, t∗ ) + α1 Lattn (θ, t∗ ), θ (11) s.t. t∗ = TriggerOpt(t|θ), where TriggerOpt(t|θ) is the trigger optimization task: t∗ = TriggerOpt(t|θ) = argmin Lbd (θ, t) + α1 Lattn (θ, t) + α2 Lvis (t), (12) t

s.t. min(t) ≥ δlow , max(t) ≤ δupp , in which δlow , δupp denote the lower and upper bounds of the trigger perturbation. To decouple the correlation between t and θ in the non-separable loss terms Lattn (θ, t) and Lbd (θ, t), we reformulate the optimization problem in Equation (10) as a bi-level optimization problem including (i) upper-level task: model optimization (as in Equation (11)) and (ii) lower-level task: trigger optimization (as in Equation (12)), which is a constraint of upper-level task. We could solve these two tasks separately by first optimizing t while keeping θ unchanged (obtaining t∗ ), and then optimizing θ with t∗ (obtaining θ∗ ). However, this optimization process is ineffective in obtaining the optimal t and θ. This is because t can only capture the original model attention under the initial parameters θ during optimization, which fails to ensure accurate attention of clean samples to calculate the attention disparity under the converged victim model. D. Adaptive Backdoor Training Unlike above optimization methods that either optimize t and θ simultaneously or fully optimize one before the other, we propose an adaptive backdoor training framework that alternatively searches solutions for two tasks in Equations (11) and (12). We describe the workflow of our adaptive backdoor training in Algorithm 1 in the supplementary material. Overall, we alternatively optimize a patch-wise trigger pattern t and ViT model parameters θ over T epochs, while t and θ are optimized for Tt and Tm times in each epoch. Within our adaptive backdoor training workflow, we first focus on the lower-level task (i.e., trigger optimization in lines 3-7). We incorporate three loss terms Lbd , Lvis , and Lattn , each dependent on the variable t. These terms are aggregated into an objective function using Lagrange coefficients α1 and α2 , and then minimized via SGD optimizer. During trigger optimization, t can easily adapt to the gradual change of the ViT model to minimize attention disparities caused by the update of θ and maximize attack effectiveness. After trigger optimization, we obtain the optimal t∗ under the current model f with parameters θ. Next, we optimize θ in the upper-level task (i.e., model optimization in lines 812). Similar to trigger optimization, we aggregate loss terms Lc , Lbd , Lattn that contain the variable θ into an objective function with a Lagrange coefficient α1 and minimize it via AdamW [64] optimizer. During model optimization, the minimal change of t allows the ViT model to adaptively finetune θ to capture the features of t and minimize attention disparities caused by the update of t.

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

8

TABLE II: Attack performance measured by ACC (%), ASR (%) and TREs (%) for 7 attacks against 4 datasets. Our method achieves comparable or superior performance on ACCs, ASRs and TREs compared to 6 other attacks. Sub-ImgNet[2]

Attack ACC

ASR

CIFAR-10[58] TRE

ACC

ASR

ImgNet[2] TRE

ACC

ASR

CIFAR-100[58] TRE

ACC

ASR

TRE

Clean

96.80

-

-

98.07

-

-

77.34

-

-

90.65

-

-

BAD N ETS† [10] WA N ET [12] T ROJ V I T [14] BADV I T [13] DBIA‡ [15] BAVT [28] LIRA [19] PASTA

89.00 92.80 96.40 96.00 95.40 86.00 95.50 96.00

99.80 92.40 99.60 100.00 100.00 77.42 100.00 98.71

14.27 100.00 13.51 98.87

94.84 98.37 97.90 93.20 96.58 84.95 98.29 97.59

99.67 99.82 99.99 100.00 100.00 69.60 100.00 99.93

14.54 100.00 34.12 99.97

70.23 78.04 79.01 77.99 74.33 68.85 74.40 73.38

99.93 96.65 99.68 99.85 100.00 81.33 100.00 99.56

1.17 99.83 2.33 97.70

85.28 90.26 90.03 86.69 88.80 74.20 87.90 90.95

97.77 99.33 99.97 99.90 100.00 83.20 100.00 99.99

11.24 99.84 51.27 99.99

†: To customize BadNets for ViTs, we set the trigger size equal to the patch size of 16×16. ‡: We compute TRE of DBIA by sliding its trigger (48×48 pixels) across the image with a step size of 16 pixels.

PASTA solves the bi-level optimization problem effectively. Note we use two small epochs Tt and Tm to enable frequent alternation between the two optimization tasks, which ensures a small gradient trajectory (i.e., minimal changes to both θ and t) in each epoch. Through this, θ and t gradually adapt to each other easily, ensuring both attack effectiveness and visual stealthiness, while minimizing the evolving attention disparities induced by θ and t. This decouples the correlation between t and θ in the non-separable loss terms, and reduces local optima in the loss landscape of Equations (11) and (12). Therefore, we effectively address the optimization challenge posed by non-separable loss terms, obtaining optimal t∗ and θ. VI. E XPERIMENTS In this section, we discuss our experimental setup and the characteristics of the proposed attack. A. Experimental Setup Experimental Environment and Setting. Our PASTA is implemented on Python 3.10, PyTorch 2.2.2 [65] and Ubuntu 22.04. We conducted all experiments on a workstation with Ryzen 9 7950X, 2×32GB DDR5 RAM and NVIDIA GeForce RTX 4090 24GB. For the default algorithm setting, we train the ViT model by AdamW optimizer with β1 and β2 of 0.99, learning rate of 2×10−5 , numerical stability parameter ϵ of 1×10−6 and weight decay of 2×10−5 . We optimize the trigger by SGD with the learning rate of 0.01. We set the batch size to 64 and the number of global epoch to 20 for backdoor training. The trigger and model optimization epoch is set to 3 and 5 respectively. In model optimization, we set the poison ratio to 2% for all dataset but 1% for ImageNet. In trigger optimization, we randomly select 5% samples from training data for all datasets. The target label is set to 7 for backdoor attacks across all datasets. We set α1 to 1.0 and α2 to 0.005. We include 9 predefined trigger insertion locations in MIS, where Sctr ={(3,3), (3,10), (7,7), (10,3), (10,10)} and Scor ={(0,0), (0,13), (13,0), (13,13)}. In each pair, the first value indicates the row index and the second the column index. For twofold stealthiness evaluation, we set the TAL to (0,0). δlow and δupp is set to the minimum and maximum pixel values of the dataset.

Datasets and Models. We evaluate our backdoor attack on four benchmark tasks on both small and large-scale datasets to confirm its scalability across various classification tasks, including objective classification on CIFAR-10 [58], finegrained classification on CIFAR-100 [58], large-scale visual recognition on ImageNet (ImgNet) (1000 classes) [2] and a subset of ImageNet (Sub-ImgNet) which is formed by 10 classes randomly selected from ImageNet. Our dataset selection covers a range of tasks and scales to demonstrate the generality of our attack. We fine-tune all backdoor attacks on pretrained ViT-Base [8] in default. See Table IX in the supplementary material for attack effectiveness of PASTA against heterogeneous ViT structures such as ViT-L [8], DeiT [66], CaiT [67], and BeiT [68]. We discuss application extensions of PASTA in Appendix I of supplementary material. Attack Payload. In our new backdoor payload, the attacker can activate PASTA backdoor on arbitrary patches during inference. To bypass backdoor defenses, we introduce 3 trigger activation strategies: (i) single patch + random location: the trigger is inserted at a randomly selected TAL within one patch per image; (ii) multiple patches + fixed locations: the trigger is inserted into multiple randomly selected TALs, with the same locations used across all images; (iii) multiple patches + random locations: the trigger is inserted into multiple randomly selected TALs, varying per image. For other attacks, we adopt the conventional backdoor payload, i.e., the attacker can only activate the backdoor in the specific patch location, same as the trigger insertion location used during training. Evaluation Metrics. We introduce metrics to quantitatively measure our attack performance in effectiveness, visual stealthiness and attention imperceptibility. (1) For attack effectiveness and functionality preservation: we evaluate the effectiveness using Attack Success Rate (ASR), which measures the proportion of poisoned samples that are misclassified into the attacker-desired label. We use clean Accuracy (ACC) to measure the proportion of benign samples that are correctly classified into the ground-truth label. Notably, we use the average ASR across all patches, denoted as TRE (defined in Equation (4)), to evaluate attack effectiveness under our proposed payload. It represents the average ASR across all patches activated by a patch-wise trigger, as defined in Equation (4). (2) For visual stealthiness, we use PSNR, SSIM, and

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

LPIPS [69] that can reflect human vision on images to evaluate visual imperceptibility between clean and poisoned data. LPIPS leverages deep features from CNNs to assess perceptual similarity, whereas SSIM and PSNR rely on pixellevel statistical similarity. Additionally, as the l2 -norm is commonly adopted to evaluate trigger stealthiness [70], [71], we also include it in our experimental comparisons. (3) For attention imperceptibility, we follow Wang et al. [16] and adopt ARES, APSNR, and ALPIPS to evaluate the invisibility between the attention map of clean and poisoned images. ARES reflects the average distance between attention maps of clean and poisoned images, while APSNR and ALPIPS follow the same computation as PSNR and LPIPS but are applied to attention maps. We also include the l2 -norm as a simple and widely-used metric to quantify overall attention deviation. B. Attack Effectiveness We compare PASTA with popular ViT backdoor attacks: TrojViT [14], BadViT [13], DBIA [15], BAVT [28] and LIRA [19], against 4 datasets. Additionally, following Doan et al. [57], we incorporate the ViT-specific versions of BadNets [10] and WaNet [12] into our experiment. Note that some works (Narcissus [27], HCB [29], BELT [30] and LADDER [24]) are CNN-specific attacks and they do not involve patch-aware trigger designs. Therefore, their methods are less effective against ViTs than ViT-specific attacks. While AIBA [16] achieves attention stealthiness, its trigger is distributed across multiple patches, making it incompatible with our payload. For these reasons, we exclude these attack methods for comparison. For patch-wise attacks (BadNets, BadViT, DBIA and PASTA), we additionally report TRE based on our attack payload. According to Table II, PASTA achieves ASRs exceeding 97.70% and up to 99.99% across datasets. Meanwhile, under PASTA, ACC drop on the victim model is limited to a maximum of 2.96%, with an average drop of just 0.99% – significantly lower than the 3.47% average drop observed in comparison attacks. When evaluating attack performance of patch-wise attacks under our payload, BadNets, BadViT, and DBIA achieve an average TRE of 45.42%, significantly lower than the 99.13% achieved by PASTA across various datasets. Note that BadNets, BadViT and DBIA all leverage replace-based trigger insertion methods that introduce distinguishable visual artifacts. In particular, BadViT maximizes the magnitude of trigger perturbations during trigger generation to induce the strongest attention at the trigger insertion location. As a result, the large perturbation under REP-based insertion inevitably leads to a strong TRE (see Figure 1(g)), although at the cost of significantly compromised twofold stealthiness (see Figures 3–4, and Tables III–IV). The above results confirm that PASTA delivers excellent attack performance under both conventional and our attack payload strategies while preserving functionality on benign tasks. We note that our superior attack effectiveness persists in heterogeneous ViT structures (see Appendix E of supplementary material). We further investigate the computational cost of PASTA in Appendix F of supplementary material.

9

C. Natural (Visual) Stealthiness Natural stealthiness is vital to guarantee that poisoned images remain imperceptible to human inspection. We quantitatively compare the visual differences between clean and poisoned images against l2 -norm, PSNR, SSIM, and LPIPS. In Table III, we see that PASTA achieves superior visual stealthiness in all the 16 cases under 4 metrics across 4 datasets, striking the enhanced natural stealthiness than other attacks. This is so because: (i) We consider the natural stealthiness in trigger optimization; (ii) In our adaptive backdoor training framework, trigger perturbations can easily adapt to the gradual change of the victim model to achieve attack effectiveness while maintain invisibility. Consequently, these minimal trigger perturbations of PASTA lead to negligible differences between clean and poisoned images, making the latter imperceptible to visual inspection. We also visualize the clean and poisoned images of various attacks on SubImgNet in Figure 3, further demonstrating the superior natural stealthiness of PASTA.

D. Attention Stealthiness Attention stealthiness ensures minimal attention disparity between clean and poisoned images, making malicious manipulations introduced by triggers undetectable by machine inspection. We quantitatively and qualitatively validate the attention stealthiness on PASTA and comparison attacks. In Table IV, we quantitatively assess the attention disparity between clean and poisoned images using APSNR, ALPIPS, ARES, and l2 -norm. According to the results, PASTA consistently achieves the best performance on 11 out of 16 cases, achieving 18.02×, 1.47×, 387.73× and 83.05× better attention stealthiness under l2 -norm, APSNR, ALPIPS and ARES, respectively. Building on similar principles to natural stealthiness, PASTA achieves superior attention stealthiness since (i) We consider the attention stealthiness in both trigger and model optimization; (ii) In our adaptive backdoor training framework, the minimal change of trigger perturbation enables the model to adaptively fine-tune its parameters to minimize attention disparities induced by trigger; (iii) Our trigger can also adapt to the gradual change of the model to minimize attention disparities caused by the update of model parameters. For qualitative analysis, following Akshayvarun et al. [72], we visualize the attention maps of clean and poisoned samples on Sub-ImgNet via Attention Rollout across all layers with “mean” head fusion. As shown in Figure 4, existing ViTspecific attacks induce a clear attention shift toward the trigger insertion patch locations, while PASTA does not introduce any distinguishable attention anomalies. Given that the attention of PASTA-poisoned images closely resemble those of clean images, it remains unclear how the trigger pattern is captured within ViTs. We provide additional poisoned samples and stealthiness results under various TALs in Figures 11–14 in Appendix G of supplementary material, further confirming the twofold stealthiness of PASTA regardless of any TALs.

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

10

TABLE III: Natural stealthiness (PSNR ↑, SSIM ↑, LPIPS ↓ and l2 -norm ↓) of trigger pattern. Across 4 metrics and 4 datasets, our PASTA consistently demonstrates superior visual imperceptibility compared to 6 other attacks. Sub-ImgNet[2]

Attacks

ImgNet[2]

CIFAR-10[58]

CIFAR-100[58]

l2

PSNR

SSIM

LPIPS

l2

PSNR

SSIM

LPIPS

l2

PSNR

SSIM

LPIPS

l2

PSNR

SSIM

LPIPS

Clean

0.0000

Inf

1.0000

0.0000

0.0000

Inf

1.0000

0.0000

0.0000

Inf

1.0000

0.0000

0.0000

Inf

1.0000

0.0000

BAD N ETS [10] T ROJ V I T[14] BADV I T[13] DBIA[15] BAVT-S RC†[28] BAVT-T GT‡[28] WA N ET [12] LIRA[19] PASTA

42.1500 284.8073 42.0462 227.0036 115.6181 21.7199 39.1905 27.4920 2.0939

29.1604 11.1658 28.3249 13.1129 19.0395 33.4822 28.9232 31.4537 53.7703

0.9909 0.9223 0.9932 0.9444 0.9753 0.8178 0.9335 0.8093 0.9986

0.0167 0.3324 0.0093 0.0929 0.0416 0.0480 0.0270 0.1131 0.0001

39.3477 300.4428 208.5887 224.3010 108.1486 22.4403 31.5502 33.2664 2.3105

29.4711 10.7309 13.8230 13.2148 19.5931 33.1933 31.0397 29.7716 57.9272

0.9911 0.9222 0.9922 0.9446 0.9752 0.8968 0.9461 0.7396 0.9992

0.0185 0.3486 0.0365 0.0929 0.0516 0.0325 0.0233 0.1093 0.0001

32.4444 278.0965 69.1612 210.8670 94.1368 20.8942 5.4682 28.1589 0.4922

23.2877 2.9590 15.0008 5.3328 12.3706 25.3812 37.3721 31.2259 64.5016

0.9903 0.9117 0.9923 0.9430 0.9733 0.7106 0.9895 0.6390 0.9961

0.0338 0.5577 0.0432 0.1530 0.0693 0.4093 0.0023 0.4086 0.0001

13.6962 203.1893 98.9295 203.3775 85.7749 17.6688 1.8535 30.9879 0.3420

32.9416 5.6724 11.8695 5.6112 13.1217 26.8344 46.7840 30.3900 61.1019

0.9919 0.9115 0.9918 0.9410 0.9720 0.5076 0.9979 0.6180 0.9996

0.0381 0.5915 0.0614 0.1864 0.1043 0.7389 0.0016 0.2967 0.0002

†: BAVT-Src measures the natural stealthiness of poisoned samples during training. ‡: BAVT-Tgt measures the natural stealthiness of poisoned samples during inference.

Clean

WaNet[12]

BadNets[10]

BadViT[13]

DBIA[15]

TrojViT[14]

LIRA[19]

PASTA

Fig. 3: Visualization of clean and poisoned images under various backdoor attacks on ImageNet. Unlike existing patch-wise attacks that replace patches with obvious mosaic patterns, our method achieves excellent natural stealthiness. TABLE IV: Attention stealthiness (APSNR ↑, ALPIPS ↓, ARES ↓ and l2 -norm ↓) between clean and poisoned images. Across 4 metrics on 4 datasets, our PASTA consistently demonstrates superior attention stealthiness compared to 6 other attacks. Sub-ImgNet[2]

Attacks

ImgNet[2]

CIFAR-10[58]

CIFAR-100[58]

l2

APSNR

ALPIPS

ARES

l2

APSNR

ALPIPS

ARES

l2

APSNR

ALPIPS

ARES

l2

APSNR

ALPIPS

ARES

Clean

0.0000

Inf

0.0000

0.0000

0.0000

Inf

0.0000

0.0000

0.0000

Inf

0.0000

0.0000

0.0000

Inf

0.0000

0.0000

BAD N ETS [10] T ROJ V I T[14] BADV I T[13] DBIA[15] BAVT-S RC [28] BAVT-T GT [28] WA N ET [12] LIRA[19] PASTA

3.2468 7.8553 3.7438 12.6012 2.8984 2.4473 2.0418 3.6776 0.1106

48.8993 41.5303 48.4666 37.0969 50.6126 43.0074 53.2639 47.8981 78.2216

0.0184 0.2636 0.0142 0.1860 0.0078 0.0126 0.0143 0.0408 0.0001

0.0022 0.0049 0.0019 0.0087 0.0016 0.0020 0.0016 0.0034 0.0003

2.6906 6.2088 8.8926 12.0215 3.5460 2.5300 1.2754 4.8229 0.3666

50.6254 43.3427 40.1956 37.5317 48.5305 51.2517 57.3069 45.7744 69.2750

0.0274 0.2069 0.0454 0.1618 0.0052 0.0197 0.0056 0.0685 0.0001

0.0018 0.0038 0.0030 0.0073 0.0018 0.0022 0.0010 0.0043 0.0001

0.6242 23.5580 22.5342 13.1706 2.5926 4.2255 2.0500 4.7763 1.0967

61.5146 29.7748 30.0167 37.0780 49.1797 45.1854 51.4066 44.8429 56.6624

0.0007 0.2909 0.2984 0.2157 0.0075 0.0417 0.0150 0.0424 0.0008

0.0004 0.0095 0.0089 0.0081 0.0016 0.0031 0.0002 0.0048 0.0004

10.9305 7.1293 14.4154 17.5719 2.4473 4.2813 1.5119 3.7531 1.3208

30.0090 34.0330 27.6394 25.7708 43.0074 38.1698 47.3005 47.7368 52.2845

0.0763 0.1000 0.1821 0.1977 0.0126 0.0460 0.0044 0.0365 0.0023

0.0043 0.0043 0.8633 0.0085 0.0020 0.0039 0.0001 0.0042 0.0004

E. Attack Performance against Defenses Against Patch Operations. Studying the sensitivity of benign accuracy and attack effectiveness in ViT backdoor attacks to patch-based operations such as patch drop and shuffle is essential, as ViTs process images as patch sequences and are inherently sensitive to patch-level modifications. Following Doan et al. [57], we consider three patch operations: Patch Drop, Patch Shuffle and their combination Drop & Shuffle. For each attack, we apply patch operations to samples from the validation dataset and validate the clean and poisoned images to obtain ACCs and ASRs. Each patch operation is repeated 100 times, each time on randomly selected patches or patch pairs. For Drop & Shuffle, Patch Shuffle is applied first, followed by Patch Drop on the shuffled image. We evaluate the attack performance of PASTA via the conventional payload (denoted as Fixed 1

TAL) and three proposed strategies in our attack payload (see Section VI-A): single patch + random location (Rand 1 TAL), 10 patches + fixed locations (Fixed 10 TALs), and 20 patches + random locations (Rand 20 TALs). In Table V, we first observe that Drop & Shuffle exhibits the most significant average decline in both ACC (37.58%) and ASR (37.83%) compared to Patch Drop (14.43%/28.99%) and Patch Shuffle (23.50%/24.13%) across all attacks. Compare to other patch-wise attacks (BadNets, TrojViT, BadViT and DBIA), PASTA achieves an average ASR of only 38.4% across three operations under conventional payload, which is 32.47% lower than others.́ This indicates that PASTA (Fixed 1 TAL) is not effective against patch operations. Additionally, we find that PASTA (Rand 1 TAL) delivers a similar average ASR of 38% to PASTA (Fixed 1 TAL). However, thanks to our new attack payload, PASTA (Fixed 10 TALs) and

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

Original

Clean

WaNet[12]

BadNets[10]

BadViT[13]

11

DBIA[15]

TrojViT[14]

LIRA[19]

PASTA

Fig. 4: Visualization of the attention with AttentionRollout on ImgNet across PASTA and other attacks. The results show that the attention of our attack between benign and poisoned samples remains consistent, with no noticeable attention shift.

PASTA (Rand 20 TALs) achieve average ASRs of 91% and 95.2%. respectively. These results demonstrate that PASTA, when leveraging multiple trigger activation location (TAL) strategies, can significantly enhance the attack effectiveness against patch operations by approximately 2.4×. Against Patch Operation-based Detection. Doan et al. [57] observe that ViT backdoors are sensitive to patch operations and propose DBAVT for backdoor detection. It tests images with patch drop and shuffle operations to establish label-flip thresholds. During detection, samples with label-flip counts above the 90th (drop) or below the 10th (shuffle) percentile are flagged as poisoned. The true positive rate (TPR) and false negative rate (FNR) measure the proportion of poisoned samples correctly detected and clean samples misclassified as poisoned under DBAVT, respectively. In Table VI, we showcase FNR on clean images and TPR on poisoned images on 4 datasets across various attacks. On average, DBAVT achieves an FNR of 9.79% and a TPR of 10.34%, indicating that it correctly identifies 90.21% of clean samples but fails to detect 89.66% of poisoned samples. This suggests that while DBAVT effectively preserves clean sample integrity, it fails to detect a majority of poisoned samples against most attacks. Upon closer examination of the PASTA attack, we observe that employing multi-TAL strategies (Fixed 10 and Rand 20 TALs) improves its stealthiness by up to 16.8× compared to single TAL strategies (Fixed and Rand 1 TAL). Consequently, only 0.2% of poisoned samples can be correctly identified as compromised under our multiple TAL strategies on ImgNet dataset. Against BAVT Defense. Subramanya et al. [28] propose a test-time image blocking defense for ViTs, which adds a black patch into the location receiving the strongest attention. We report the ACC and ASR of various attacks against BAVT’s defense in Table VII. On average, PASTA achieves an ASR of 61.22% when the trigger is inserted into a single TAL. However, ASR increases to 100% when inserting the trigger to 10 TALs, rendering BAVT’s test-time defense ineffective against PASTA. This is because BAVT only focuses on the location with the highest attention values, which is not able to

TABLE V: Attack performance of PASTA with various payloads against patch operations on Sub-ImgNet using ViT. Patch Operation BadNets[10] TrojViT[14] BadViT[13] DBIA[15] BAVT[28] WaNet[12] PASTA (Fixed 1 TAL) PASTA (Fixed 10 TALs) PASTA (Rand 1 TAL) PASTA (Rand 20 TALs)

Patch Drop [73]

Patch Shuffle [57]

ACC

ASR

ACC

ASR

Drop & Shuffle ACC

ASR

71.40 72.60 78.40 75.20 68.20 71.00 81.00 81.40 81.00 78.40

58.20 74.00 54.20 96.40 29.03 92.20 38.20 92.60 37.40 96.00

66.60 61.00 66.00 69.20 62.20 61.20 71.00 71.60 72.60 70.80

55.80 93.20 66.00 99.60 35.48 83.20 43.00 92.20 43.00 96.60

54.00 45.20 51.80 51.00 48.40 51.80 56.40 56.40 53.00 51.80

40.00 79.20 42.00 91.80 6.45 82.40 34.00 88.20 33.60 93.00

TABLE VI: Attack performance against DBAVT [57] via true positive rate (TPR%↑) on poisoned images and false negative rate (FNR%↓) on clean images. CIFAR-10[58]

CIFAR-100[58]

FNR

TPR

FNR

TPR

FNR

TPR

FNR

TPR

Ideal Defense

0.00

100.00

0.00

100.00

0.00

100.00

0.00

100.00

BadNets[10] TrojViT[14] BadViT[13] DBIA[15] BAVT[28] WaNet[12] PASTA (Fixed 1 TAL) PASTA (Fixed 10 TALs) PASTA (Rand 1 TAL) PASTA (Rand 20 TALs)

9.60 11.20 11.00 10.60 9.60 9.60 9.20 9.40 10.00 9.60

0.60 0.20 35.40 0.00 29.03 4.40 31.40 1.80 23.20 2.20

9.40 9.80 9.80 9.00 9.40 10.60 15.20 10.60 6.40 11.40

0.00 3.80 0.00 0.00 15.33 0.60 8.80 0.20 10.60 0.20

10.40 8.00 8.80 8.60 8.60 7.40 11.00 10.40 8.00 9.60

3.40 0.00 0.54 0.00 22.80 0.20 73.40 4.00 54.00 3.00

9.20 16.20 10.00 10.00 11.20 10.80 9.00 6.80 8.80 10.40

0.08 49.41 0.00 0.00 7.20 0.20 0.00 0.00 0.00 0.00

Attacks

Sub-ImgNet[2]

ImgNet[2]

block multiple TALs under our attack payload. Additionally, our patch-wise triggers do not introduce abnormal attention maps compared to other patch-wise attacks (see Figure 4 in section VI-D and Figures 13–14 in the Appendix of supplementary material). Thus, the superior attention stealthiness of PASTA makes it difficult for BAVT to detect trigger locations. Against CNN-specific Defenses. See Appendix D of supplementary material for the attack effectiveness of PASTA against STRIP [43], Fine-pruning [47], NC [48] and ANP [74]. Adaptive Defenses. Refer to Appendix H of supplementary materialfor our adaptive defense to mitigate potential misuse. Ablation Study. We investigate the impact of α1 and α2 on PASTA attack performance in Appendix B of supplementary material.

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

TABLE VII: The ACCs and ASRs of PASTA and comparison attacks against the test-time defense mechanism of BAVT [28]. Attacks BadNets[10] TrojViT[14] BadViT[13] DBIA[15] BAVT[28] WaNet[12] PASTA (Fixed 1 TAL) PASTA (Fixed 10 TALs) PASTA (Rand 1 TAL) PASTA (Rand 10 TALs)†

Sub-ImgNet[2]

ImgNet[2]

CIFAR-10[58]

CIFAR-100[58]

ACC

ASR

ACC

ASR

ACC

ASR

ACC

ASR

88.00 92.00 87.00 91.60 79.00 87.20 87.40 87.40 87.40 87.40

11.80 94.80 52.00 100.00 50.00 78.40 57.40 100.00 47.40 100.00

66.62 74.81 67.01 72.34 57.80 68.49 62.20 62.20 62.20 62.20

3.01 99.41 98.42 100.00 72.16 88.75 96.52 100.00 90.25 100.00

92.58 96.87 89.34 96.40 61.34 78.53 73.93 73.93 73.93 73.93

33.62 99.58 99.31 100.00 56.00 14.50 61.48 100.00 40.20 100.00

65.92 80.74 68.99 79.75 57.95 73.20 73.92 73.92 73.92 73.92

1.58 98.24 99.06 100.00 81.60 1.56 55.59 99.99 40.88 100.00

†: Since PASTA (Rand 10 TALs) achieves ≥99.99% ASRs on 4 datasets, we omit using more TALs.

VII. C ONCLUSION This paper systematically studies the impact of trigger activation locations on attack effectiveness in both CNNs and ViTs. We observe that patch-wise triggers against ViTs have a radiating effect (TRE) on attack effectiveness, due to the self-attention mechanism. Based on our findings, we propose a multi-location trigger-insertion strategy during backdoor training to achieve strong TRE. This introduces a new backdoor payload that activates the backdoor across arbitrary patches to evade defenses. In addition, we consider visual and attention stealthiness to bypass human and machine inspection. Hence, we introduce PASTA, a visual and attention-stealthy patchwise backdoor attack against ViTs under our payload. We formulate all attack objectives as a bi-level optimization problem and introduce an adaptive optimization framework to solve it effectively. Extensive experiments show that PASTA achieves superior attack effectiveness across all patches, excellent twofold stealthiness, and better attack robustness against stateof-the-art CNN- and ViT-specific defenses across 4 datasets. R EFERENCES [1] Y. LeCun, B. Boser, J. Denker, D. Henderson, R. Howard, W. Hubbard, and L. Jackel, “Handwritten digit recognition with a back-propagation network,” in Advances in Neural Information Processing Systems, vol. 2, 1989. [2] A. Krizhevsky, I. Sutskever, and G. E. Hinton, “ImageNet Classification with Deep Convolutional Neural Networks,” Advances in Neural Information Processing Systems, vol. 25, 2012. [3] Y. Wu, J. Lim, and M.-H. Yang, “Online Object Tracking: A Benchmark,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2013, pp. 2411–2418. [4] L. Zheng, M. Tang, Y. Chen, G. Zhu, J. Wang, and H. Lu, “Improving multiple object tracking with single object tracking,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 2453–2462. [5] Z.-Q. Zhao, P. Zheng, S.-t. Xu, and X. Wu, “Object detection with deep learning: A review,” IEEE Transactions on Neural Networks and Learning Systems, vol. 30, no. 11, pp. 3212–3232, 2019. [6] Z. Zou, K. Chen, Z. Shi, Y. Guo, and J. Ye, “Object detection in 20 years: A survey,” Proceedings of the IEEE, vol. 111, no. 3, pp. 257–276, 2023. [7] M. Wang and W. Deng, “Deep face recognition: A survey,” Neurocomputing, vol. 429, pp. 215–244, 2021. [8] A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly, J. Uszkoreit, and N. Houlsby, “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” in International Conference on Learning Representations, 2021. [9] Y. Lecun, L. Bottou, Y. Bengio, and P. Haffner, “Gradient-based Learning Applied to Document Recognition,” Proceedings of the IEEE, vol. 86, no. 11, pp. 2278–2324, 1998.

12

[10] T. Gu, B. Dolan-Gavitt, and S. Garg, “Badnets: Identifying Vulnerabilities in the Machine Learning Model Supply Chain,” arXiv preprint arXiv:1708.06733, 2017. [11] X. Chen, C. Liu, B. Li, K. Lu, and D. Song, “Targeted Backdoor Attacks on Deep Learning Systems Using Data Poisoning,” arXiv preprint arXiv:1712.05526, 2017. [12] T. A. Nguyen and A. T. Tran, “WaNet - Imperceptible Warping-based Backdoor Attack,” in International Conference on Learning Representations, 2021. [13] Z. Yuan, P. Zhou, K. Zou, and Y. Cheng, “You Are Catching My Attention: Are Vision Transformers Bad Learners under Backdoor Attacks?” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 24 605–24 615. [14] M. Zheng, Q. Lou, and L. Jiang, “TrojViT: Trojan Insertion in Vision Transformers,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 4025–4034. [15] P. Lv, H. Ma, J. Zhou, R. Liang, K. Chen, S. Zhang, and Y. Yang, “DBIA: Data-Free Backdoor Attack Against Transformer Networks,” in IEEE International Conference on Multimedia and Expo, 2023, pp. 2819–2824. [16] Z. Wang, R. Wang, and L. Jing, “Attention-Imperceptible Backdoor Attacks on Vision Transformers,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 39, no. 8, 2025, pp. 8241–8249. [17] Y. Liu, S. Ma, Y. Aafer, W.-C. Lee, J. Zhai, W. Wang, and X. Zhang, “Trojaning Attack on Neural Networks,” in Network And Distributed System Security Symposium, 2018. [18] T. A. Nguyen and A. Tran, “Input-Aware Dynamic Backdoor Attack,” in Advances in Neural Information Processing Systems, vol. 33, 2020, pp. 3454–3464. [19] K. Doan, Y. Lao, W. Zhao, and P. Li, “Lira: Learnable, Imperceptible and Robust Backdoor Attacks,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 11 966–11 976. [20] Z. Zhao, X. Chen, Y. Xuan, Y. Dong, D. Wang, and K. Liang, “DEFEAT: Deep Hidden Feature Backdoor Attacks by Imperceptible Perturbation and Latent Representation Constraints,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 15 213–15 222. [21] S. Cheng, Y. Liu, S. Ma, and X. Zhang, “Deep Feature Space Trojan Attack of Neural Networks by Controlled Detoxification,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, no. 2, 2021, pp. 1148–1156. [22] T. Wang, Y. Yao, F. Xu, S. An, H. Tong, and T. Wang, “An Invisible Black-box Backdoor Attack through Frequency Domain,” in European Conference on Computer Vision, 2022, pp. 396–413. [23] Y. Feng, B. Ma, J. Zhang, S. Zhao, Y. Xia, and D. Tao, “Fiba: Frequency-injection based Backdoor Attack in Medical Image Analysis,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 20 876–20 885. [24] D. Liu, Y. Qiao, R. Wang, K. Liang, and G. Smaragdakis, “LADDER: Multi-objective Backdoor Attack via Evolutionary Algorithm,” in Network and Distributed System Security Symposium, 2025. [25] Y. Qiao, D. Liu, R. Wang, and K. Liang, “Low-frequency Black-box Backdoor Attack via Evolutionary Algorithm,” in IEEE/CVF Winter Conference on Applications of Computer Vision, 2025, pp. 7582–7592. [26] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention Is All You Need,” in Advances in Neural Information Processing Systems, vol. 30, 2017, pp. 5999–6009. [27] Y. Zeng, M. Pan, H. A. Just, L. Lyu, M. Qiu, and R. Jia, “Narcissus: A Practical Clean-Label Backdoor Attack with Limited Information,” in Proceedings of the ACM SIGSAC Conference on Computer and Communications Security, 2023, pp. 771–785. [28] A. Subramanya, A. Saha, S. A. Koohpayegani, A. Tejankar, and H. Pirsiavash, “Backdoor Attacks on Vision Transformers,” arXiv preprint arXiv:2206.08477, 2022. [29] H. Ma, S. Wang, Y. Gao, Z. Zhang, H. Qiu, M. Xue, A. Abuadbba, A. Fu, S. Nepal, and D. Abbott, “Watch Out! Simple Horizontal Class Backdoor Can Trivially Evade Defense,” in Proceedings of the ACM SIGSAC Conference on Computer and Communications Security, 2024, pp. 4465–4479. [30] H. Qiu, J. Sun, M. Zhang, X. Pan, and M. Yang, “ BELT: Old-School Backdoor Attacks can Evade the State-of-the-Art Defense with Backdoor Exclusivity Lifting,” in IEEE Symposium on Security and Privacy, 2024, pp. 2124–2141. [31] M. Barni, K. Kallas, and B. Tondi, “A new Backdoor Attack in CNNs by Training Set Corruption without Label Poisoning,” in IEEE International Conference on Image Processing, 2019, pp. 101–105.

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

[32] Y. Liu, X. Ma, J. Bailey, and F. Lu, “Reflection Backdoor: A Natural Backdoor Attack on Deep Neural Networks,” in European Conference on Computer Vision, 2020, pp. 182–199. [33] Y. Li, Y. Li, B. Wu, L. Li, R. He, and S. Lyu, “Invisible Backdoor Attack with Sample-Specific Triggers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 16 463–16 472. [34] W. Jiang, H. Li, G. Xu, and T. Zhang, “Color Backdoor: A Robust Poisoning Attack in Color Space,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 8133–8142. [35] Y. Zeng, W. Park, Z. M. Mao, and R. Jia, “Rethinking the Backdoor Attacks’ Triggers: A Frequency Perspective,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 16 473–16 481. [36] R. Hou, T. Huang, H. Yan, L. Ke, and W. Tang, “A Stealthy and Robust Backdoor Attack via Frequency Domain Transform,” World Wide Web, pp. 1–17, 2023. [37] K. Doan, Y. Lao, and P. Li, “Backdoor Attack with Imperceptible Input and Latent Modification,” in Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 18 944–18 957. [38] N. Zhong, Z. Qian, and X. Zhang, “Imperceptible Backdoor Attack: From Input Space to Feature Representation,” in Proceedings of the International Joint Conference on Artificial Intelligence, 2022, pp. 1736– 1742. [39] P. Lv, C. Yue, R. Liang, Y. Yang, S. Zhang, H. Ma, and K. Chen, “A Data-free Backdoor Injection Approach in Neural Networks,” in USENIX Security Symposium, 2023, pp. 2671–2688. [40] J. Lan, J. Wang, B. Yan, Z. Yan, and E. Bertino, “Flowmur: A stealthy and Practical Audio Backdoor Attack with Limited Knowledge,” in IEEE Symposium on Security and Privacy, 2024, pp. 1646–1664. [41] G. Abad, O. Ersoy, S. Picek, and A. Urbieta, “Sneaky Spikes: Uncovering Stealthy Backdoor Attacks in Spiking Neural Networks with Neuromorphic Data,” in Network and Distributed System Security Symposium, 2024. [42] J. Zhang, J. Chi, Z. Li, K. Cai, Y. Zhang, and Y. Tian, “Badmerging: Backdoor Attacks against Model Merging,” arXiv preprint arXiv:2408.07362, 2024. [43] Y. Gao, C. Xu, D. Wang, S. Chen, D. C. Ranasinghe, and S. Nepal, “Strip: A Defence Against Trojan Attacks on Deep Neural Networks,” in Proceedings of the Annual Computer Security Applications Conference, 2019, pp. 113–125. [44] B. Chen, W. Carvalho, N. Baracaldo, H. Ludwig, B. Edwards, T. Lee, I. Molloy, and B. Srivastava, “Detecting Backdoor Attacks on Deep Neural Networks by Activation Clustering,” arXiv preprint arXiv:1811.03728, 2018. [45] B. Tran, J. Li, and A. Madry, “Spectral Signatures in Backdoor Attacks,” in Advances in Neural Information Processing Systems, vol. 31, 2018, pp. 8011–8021. [46] S. Kolouri, A. Saha, H. Pirsiavash, and H. Hoffmann, “Universal Litmus Patterns: Revealing Backdoor Attacks in CNNs,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 301–310. [47] K. Liu, B. Dolan-Gavitt, and S. Garg, “Fine-Pruning: Defending against Backdooring Attacks on Deep Neural Networks,” in International Symposium on Research in Attacks, Intrusions, and Defenses, 2018, pp. 273– 294. [48] B. Wang, Y. Yao, S. Shan, H. Li, B. Viswanath, H. Zheng, and B. Y. Zhao, “Neural Cleanse: Identifying and Mitigating Backdoor Attacks in Neural Networks,” in IEEE Symposium on Security and Privacy, 2019, pp. 707–723. [49] H. Chen, C. Fu, J. Zhao, and F. Koushanfar, “DeepInspect: A Blackbox Trojan Detection and Mitigation Framework for Deep Neural Networks,” in Proceedings of the International Joint Conference on Artificial Intelligence, 2019, pp. 4658–4664. [50] X. Qiao, Y. Yang, and H. Li, “Defending Neural Backdoors via Generative Distribution Modeling,” in Advances in Neural Information Processing Systems, vol. 32, 2019, pp. 14 027–14 036. [51] Y. Li, X. Lyu, N. Koren, L. Lyu, B. Li, and X. Ma, “Neural Attention Distillation: Erasing Backdoor Triggers from Deep Neural Networks,” in International Conference on Learning Representations, 2021. [52] Y. Li, T. Zhai, B. Wu, Y. Jiang, Z. Li, and S. Xia, “Rethinking the Trigger of Backdoor Attack,” arXiv preprint arXiv:2004.04692, 2020. [53] H. Qiu, Y. Zeng, S. Guo, T. Zhang, M. Qiu, and B. Thuraisingham, “Deepsweep: An Evaluation Framework for Mitigating DNN Backdoor Attacks using Data Augmentation,” in Proceedings of the ACM Asia Conference on Computer and Communications Security, 2021, pp. 363– 377.

13

[54] K. Gao, Y. Bai, J. Gu, Y. Yang, and S.-T. Xia, “Backdoor Defense via Adaptively Splitting Poisoned Dataset,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2023, pp. 4005–4014. [55] M. Zhu, S. Wei, H. Zha, and B. Wu, “Neural Polarizer: A Lightweight and Effective Backdoor Defense via Purifying Poisoned Features,” in Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 1132–1153. [56] Y. Shi, M. Du, X. Wu, Z. Guan, J. Sun, and N. Liu, “Black-box Backdoor Defense via Zero-shot Image Purification,” in Advances in Neural Information Processing Systems, vol. 36, 2023, pp. 57 336– 57 366. [57] K. D. Doan, Y. Lao, P. Yang, and P. Li, “Defending Backdoor Attacks on Vision Transformer via Patch Processing,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 37, no. 1, 2023, pp. 506–515. [58] A. Krizhevsky and G. Hinton, “Learning Multiple Layers of Features from Tiny Images,” 2009. [59] A. Saha, A. Subramanya, and H. Pirsiavash, “Hidden Trigger Backdoor Attacks,” in Proceedings of the AAAI Cconference on Artificial Intelligence, vol. 34, no. 07, 2020, pp. 11 957–11 965. [60] Y. Qiao, D. Liu, R. Wang, and K. Liang, “Stealthy backdoor attack against federated learning through frequency domain by backdoor neuron constraint and model camouflage,” IEEE Journal on Emerging and Selected Topics in Circuits and Systems, vol. 14, no. 4, pp. 661–672, 2024. [61] L. Wang, B. Shang, Y. Li, P. Mohapatra, W. Dong, X. Wang, and Q. Zhu, “ Split Adaptation for Pre-trained Vision Transformers ,” in IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2025, pp. 20 092–20 102. [62] J. Nocedal and S. J. Wright, Numerical Optimization. Springer, 2006. [63] S. Verel, A. Liefooghe, L. Jourdan, and C. Dhaenens, “Pareto Local Optima of Multiobjective NK-Landscapes with Correlated Objectives,” in Evolutionary Computation in Combinatorial Optimization, 2011. [64] I. Loshchilov and F. Hutter, “Decoupled Weight Decay Regularization,” in International Conference on Learning Representations, 2019. [65] A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga, A. Desmaison, A. Köpf, E. Yang, Z. DeVito, M. Raison, A. Tejani, S. Chilamkurthy, B. Steiner, L. Fang, J. Bai, and C. Soumith, “Pytorch: An Imperative Style, HighPerformance Deep Learning Library,” in Advances in Neural Information Processing Systems, vol. 32, 2019, pp. 8026–8037. [66] H. Touvron, M. Cord, M. Douze, F. Massa, A. Sablayrolles, and H. Jegou, “Training Data-efficient Image Transformers & Distillation Through Attention,” in International Conference on Machine Learning, 2021, pp. 10 347–10 357. [67] H. Touvron, M. Cord, A. Sablayrolles, G. Synnaeve, and H. Jegou, “Going Deeper with Image Transformers,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2021, pp. 32–42. [68] H. Bao, L. Dong, S. Piao, and F. Wei, “BEiT: Bert Pre-training of Image Transformers,” in International Conference on Learning Representations, 2022. [69] R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, pp. 586–595. [70] S. Li, M. Xue, B. Z. H. Zhao, H. Zhu, and X. Zhang, “Invisible Backdoor Attacks on Deep Neural Networks Via Steganography and Regularization,” IEEE Transactions on Dependable and Secure Computing, vol. 18, pp. 2088–2105, 2019. [71] C. Guo, J. S. Frank, and K. Q. Weinberger, “Low Frequency Adversarial Perturbation,” in Uncertainty in Artificial Intelligence, 2020, pp. 1127– 1137. [72] A. Subramanya, S. A. Koohpayegani, A. Saha, A. Tejankar, and H. Pirsiavash, “A Closer Look at Robustness of Vision Transformers to Backdoor Attacks ,” in IEEE/CVF Winter Conference on Applications of Computer Vision, 2024, pp. 3862–3871. [73] M. Naseer, K. Ranasinghe, S. Khan, M. Hayat, F. Khan, and M.-H. Yang, “Intriguing Properties of Vision Transformers,” in Advances in Neural Information Processing Systems, vol. 34, 2021, pp. 23 296–23 308. [74] D. Wu and Y. Wang, “Adversarial neuron pruning purifies backdoored deep models,” Advances in Neural Information Processing Systems, vol. 34, pp. 16 913–16 925, 2021. [75] S. Ren, K. He, R. Girshick, and J. Sun, “Faster r-cnn: Towards realtime object detection with region proposal networks,” arXiv preprint arXiv:1506.01497, 2016.

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

14

[76] S.-H. Chan, Y. Dong, J. Zhu, X. Zhang, and J. Zhou, “Baddet: Backdoor attacks on object detection,” in European Conference on Computer Vision. Springer, 2022, p. 396–412.

A PPENDIX A E THICAL C ONSIDERATIONS This study reveals the vulnerability of ViTs to stealthy patch-wise triggers that can activate across arbitrary patches, highlighting the need for stronger defenses. Intellectual Property. All models, datasets, methods, and code will be released upon acceptance, with datasets properly desensitized and used in compliance with licensing terms. Intended Usage. We expose a novel backdoor vulnerability in ViTs, where triggers activate regardless of patch location, and encourage the development of robust defenses. Potential Misuse. Adversaries could exploit this method to deploy backdoored ViTs that behave normally on clean inputs but misclassify when triggered, while evading existing defenses. We provide an adaptive defense in Appendix H. Risk Control. We will release all artifacts to promote transparency and support the following research. Human Subject. No human subjects are involved; evaluations rely solely on models and quantitative metrics. A PPENDIX B A BLATION S TUDY Impact of α1 and α2 on PASTA Attack Performance. We aggregate the attack objectives into a bi-level optimization problem, as defined in Equations 11 and 12. Analyzing various aggregation coefficient settings in our problem is essential to understand how each attack objective affects trigger stealthiness and attack effectiveness. Improper aggregation coefficients may cause the optimizer to favor one objective at the expense of others, leading to suboptimal attack performance. In Figure 5, we examine the impact of varying α1 and α2 on attack objectives, including ACC (%), ASR (%), and both visual and attention stealthiness (measured by l2 -norm). We observe in Figure 5(a) that ACCs are overall stable with the change of α1 and α2 , achieving maximum of 93.40% and minimum of 89.00%. However, as shown in Figure 5(b), ASRs drop sharply with increasing α1 and α2 , falling from 99.72% at α1 = 0.5, α2 = 0.001 to just 10.80% at α1 = 2.0, α2 = 0.05. This is so because high values of α1 and α2 prioritize minimizing twofold stealthiness over learning backdoor tasks in the optimization problem. To retain practical ACC and ASR, α1 ≤ 1.0 and α2 ≤ 0.01 should be used. We then evaluate twofold stealthiness by measuring the l2 norm of the image and attention disparities (between clean and poisoned samples). In Figure 5(c), visual stealthiness improves significantly as α1 increases from 0.5 to 2.0, with the average l2 disparity dropping from 10.26 to 0.51, while remaining relatively stable across different α2 values at a fixed α1 . Similarly, attention stealthiness, as shown in Figure 5(d), improves as α2 increases from 0.001 to 0.05, reducing the average attention disparity from 4.95 to 0.005, with minimal variation across different α1 values. These results suggest that (i) Minimizing visual stealthiness does not necessarily improve attention stealthiness, and vice versa. As a consequence, both

(a) ACC (%)

(b) ASR (%)

(c) Invisibility (l2 -norm)

(d) Attn. Stealthiness (l2 -norm)

Fig. 5: Visualization of the effect of Lagrange coefficients α1 and α2 on attack objectives under ViT-Tiny on Sub-ImgNet. (a)-(b): the ACC (%) and ASR(%) of the victim model. (c)(d): Visual and attention stealthiness measured by l2 -norm.

objectives in our twofold stealthiness design are indispensable; (ii) A notable scale gap between α1 and α2 to achieve practical attack performance exists (e.g., 10−1 vs. 10−3 ). It necessitates asymmetric weighting to balance gradient contributions during joint optimization; (iii) α1 ≥ 1.0 and α2 ≥ 0.005 is required for practical twofold stealthiness. Considering all attack objectives, we adopt α1 = 1.0 and α2 = 0.005 as default setting. We elaborate the automated parameter selection as our future work in Appendix I.

TABLE VIII: Summary of notations. Notation C H i L lrm M N p Scor T Tt t∗ y y′ α1 δlow η Dc L Lbd Lvis θ ρ

Description

Notation

Description

Number of channels of x Height of x Patch index Attention layer Model learning rate Binary mask for t Dataset size Patch size Corner candidate patch locations Epochs for adaptive training Epochs for trigger optimization Optimal trigger Image label Poisoned image label Visual imperceptibility factor Lower bound of trigger perturbation Target label function Clean dataset Cross-entropy loss Backdoor task loss term Visibility loss term Model parameters Poison ratio

f W l n lrt m P S Sctr Tm t x x′ ytgt α2 δupp κ Dbd Lattn Lc T θ∗

Deep learning model Width of x Layer index Number of patches (TALs) Trigger learning rate Scaling parameter of t Image patch Full location set Center candidate patch locations Epochs for model optimization Trigger pattern Image sample Poisoned image sample Target label Attention disparity factor Upper bound of trigger perturbation Number of classes Poison subset Attention loss term Clean task loss term Trigger injection function Optimal model parameters

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

Algorithm 1 PASTA Require: ViT model f with parameters θ, Trigger function T , Target label function η, Clean dataset Dc , Poison dataset Dbd , Trigger optimization epoch Tt , Model optimization epoch Tm , Total epoch T , Visual imperceptibility factor α1 , Attention disparity factor α2 , Trigger learning rate lrt , Model learning rate lrm ∗ Ensure: Parameters θbd of poisoned model f , Optimal trigger ∗ t . 1: t ← Trigger initialize(size=3 × p × p) 2: for epoch ∈ {1, 2, . . . , T } do Lower-level task: Trigger Optimization 3: for epochtP ∈ {1, 2, . . . , Tt } do 4: Lbd = (x,y)∈Dbd L(fθ (T (x, t, Mi|M IS )), η(y)) P 5: Lvis = |Dbd |−1 x∈Dbd ∥T (x, t, Mi|M IS ) − x∥2 P 6: Lattn = |Dbd |−1 x∈Dbd dis attnl (x, T , f ) 7: t = t − lrt × ∇t (Lbd + α1 Lvis + α2 Lattn ) Upper-level task: Model Optimization 8: for epochP m ∈ {1, 2, . . . , Tm } do 9: Lc =P (x,y)∈Dc L(fθ (x), y) 10: Lbd = (x,y)∈Dbd L(fθ (T (x, t, Mi|M IS )), η(y)) P 11: Lattn = |Dbd |−1 x∈Dbd dis attnl (x, T , f ) 12: θ = θ − lrθ × ∇θ (Lc + Lbd + α2 Lattn ) ∗ 13: return θbd ← θ, t∗ ← t

(a) REP (l2 :4.0, TRE:5.04)

(b) REP (l2 :16.0, TRE:7.52)

Fig. 6: (a)-(b): Visualization of TRE under replace-based trigger insertion (REP) against ResNet-18 on ImageNet.

A PPENDIX C TRE ON A L ARGE - SCALE DATASET AND AN A DVANCED CNN A RCHITECTURE We investigate the trigger radiating effect (TRE) in Section IV ,and confirm that TRE does not exist in the conventional CNN model (see Figures 1(a)–(d) in the main manuscript). To further validate this conclusion on advanced CNNs, we conduct experiments on a pre-trained ResNet-18 using ImageNet, with two randomly initialized trigger patterns: a 3×3 pattern (l2 -norm=4), and a 9×9 pattern (l2 -norm=16). During backdoor training, we insert the REP-based triggers into the top-left patch (i.e., at location (0,0)). During inference, TAL is shifted with a stride of 1. We show the first 10 and 30 steps of TRE for rows and columns in Figures 6(a)–(b), respectively. The results reveals that the attack remains effective only when TALs are close to the trigger insertion location during backdoor training, but degrades sharply as the TAL shifts. Such degradation leads to low TRE values of 5.04% and 7.52%, suggesting that even advanced CNNs like ResNet-18

15

trained under large-scale datasets like ImageNet, exhibit very limited TRE as well. A PPENDIX D ATTACK E FFECTIVENESS AGAINST BACKDOOR D EFENSES Against STRIP. STRIP is a CNN-specific backdoor defense that operates under the assumption that poisoned inputs consistently lead to the target label in a backdoored model and are resistant to label changes under input perturbations in the clean model. Under this assumption, STRIP detects poisoned samples by measuring the prediction entropy after superimposing randomly selected clean images onto the test input, expecting poisoned inputs to exhibit significantly lower entropy compared to clean ones. We test the images poisoned by PASTA and comparison attacks against STRIP, and visualize the entropy distribution for clean and poisoned samples. Figures 7(a)-(f) showcase the (normalized) probability of entropy values for poisoned (orange) and clean (blue) samples as bar plots, along with fitted distribution curves in corresponding colors. A larger overlap in the distribution areas means that the benign and poisoned samples produce more similar entropy, indicating that poisoned samples are more difficult to detect. We can see from Figure 7(f) that the two entropy distributions are well-overlapped, meaning that PASTA can evade the anomaly detection from STRIP. This is so mainly because we introduce the twofold stealthiness in our trigger design, rendering small trigger perturbations and less attention anomaly, causing the entropy distribution of poisoned samples to closely resemble that of clean ones, thus bypassing STRIP detection. Against Fine-pruning. Fine-pruning (FP) [47] is an effective and widely used backdoor defense that iteratively removes dormant neurons in the presence of clean data to eliminate potential backdoor triggers implanted during training, fortifying the model against backdoor attacks while preserving performance on benign tasks. FP is not applicable to ViTs because it targets neurons in the convolutional layer of CNNs, a component that does not exist in ViTs. We adapt FP for ViTs by pruning neurons from the fully connected layers within the MLP blocks at the end of the ViT architecture. In Figure 10, we observe that as the pruning ratio increases, ACCs declines more rapidly than ASRs in most cases. On Sub-ImgNet, ACC drops to nearly zero by the end of pruning, while ASR remains high. On ImgNet, CIFAR-10, and CIFAR100 datasets, although ASRs eventually reach zero, ACCs are harmed as well. These findings indicate that FP fails to effectively eliminate our backdoor without severely compromising the model’s performance on benign inputs. Notably, ASR on Sub-ImgNet increases when the prune ratio exceeds 70%. as pruning low-activation benign neurons concentrates backdoorrelated logits into highly active neurons. Against ANP. ANP [74] mitigates backdoor attacks by producing masks for neurons and tuning the mask parameters to remove backdoor related neurons, effectively suppressing backdoor behaviors while maintaining clean accuracy. ANP is originally designed for CNNs. We adapt ANP to ViTs by introducing adversarial perturbations to neurons in the fully

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

(a) BadNets-ViT

16

(b) BadViT

Fig. 8: The ACCs (%) and ASRs (%) of PASTA under CIFAR10 and ViT-Tiny against ANP backdoor defense after the corresponding percentage of neurons is removed.

(c) TrojViT

(d) DBIA

(e) WaNet-ViT

(f) PASTA

Fig. 7: (a)-(f): The entropy distribution obtained with model poisoned by various attacks against STRIP on Sub-ImgNet. The distribution marked in blue and orange is obtained by clean and poisoned test data.

connected layers of the ViT MLP blocks and subsequently pruning neurons according to their sensitivity to these perturbations. We evaluate the attack effectiveness (ASR, %) and benign accuracy (ACC, %) against ViT-Tiny and CIFAR-10 dataset under the PASTA attack, as neurons are progressively disabled in 5% increments based on the optimized mask values, up to 60%, beyond which the benign-task accuracy collapses. The results are shown in Figure 8. We see that ViTs are highly sensitive to neuron pruning within their MLPs. Even a small pruning ratio of 10% leads to a severe drop in benign task performance, reducing ACC to 27.6%. In contrast, the effectiveness of the PASTA attack is relatively more robust, exhibiting only a 10.6% decrease under the same pruning ratio. However, as the pruning ratio increases, both ACC and ASR collapse rapidly, reaching 10% (i.e., random guess) and 0% respectively when 20% of the neurons are removed. The results show that the performance on benign tasks is severely compromised. This phenomenon may stem from the strong correlation between neurons responsible for benign functionality and those contributing to backdoor behavior. As a result, pruning neurons associated with the backdoor also disrupts the clean tasks, leading to a simultaneous degradation in both ACC and ASR. Against Neural Cleanse (NC). NC detects potential backdoors in a suspected model by reverse-engineering triggers and verifying whether these triggers can induce misclassification on the samples. For each class, NC reverses a trigger and uses the median absolute deviation of trigger norms to detect

Fig. 9: The anomaly index produced by NC on the clean model and victim models under various backdoor attacks, evaluated on ViT-Tiny with the Sub-ImageNet dataset. The red line marks the detection threshold, above which a model is suspected to be poisoned.

anomalous target classes. The anomaly index is then defined as the maximum normalized deviation of a class’s trigger size from the median, representing the overall abnormality of the model. The model is flagged as potentially backdoored if the anomaly index exceeds 2, as suggested by the authors. We test PASTA and comparison attacks against NC under Sub-ImgNet dataset and ViT-tiny model, and show the results in Figure 9. The blue and orange bars represent the anomaly index of clean and poisoned models, respectively, computed by NC. Each attack is labeled along the x-axis. The y-axis shows the anomaly index, and the red line indicates the threshold value of 2. We observe that the anomaly index of BadNets, BadViT, and BAVT reaches 7.05, 6.13, and 3.53, respectively, exceeding the threshold of 2 and indicating that these attacks are detected by NC. In contrast, TrojViT’s anomaly index is 1.97, which is slightly below the threshold. PASTA, along with WaNet, DBIA, and LIRA, also achieves values below the threshold, successfully evading NC detection. Recall that NC relies on detecting anomalously small and sparsely concentrated triggers that uniquely activates the backdoor, whereas PASTA achieves location-agnostic activation through attention-distributed and semantically diffused trigger features. So that PASTA yields uniformly larger and less class-discriminative reversed triggers, thereby keeping the anomaly index below the detection threshold. Adaptive Defense. To undermine the threat of PASTA, we propose an adaptive defense strategy in Appendix H.

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

17

TABLE XI: The natural stealthiness (PSNR↑, SSIM↑, LPIPS↓, l2 -norm↓) and attention stealthiness (APSNR↑, ALPIPS↓, ARES↓, l2 -norm↓) of PASTA by comparing clean and poisoned images across various TALs on the Sub-ImgNet dataset. PASTA consistently maintains twofold stealthiness across TALs. Natural Stealthiness

TAL

(a) Sub-ImgNet

Attention Stealthiness

l2

PSNR

SSIM

LPIPS

l2

APSNR

ALPIPS

ARES

Clean

0.0000

Inf

1.0000

0.0000

0.0000

Inf

0.0000

0.0000

(0,0) (7,7) (13,13)

2.0939 2.1009 2.1022

53.7703 53.7634 53.7579

0.9986 0.9977 0.9978

0.0001 0.0001 0.0001

0.1106 0.0770 0.0719

78.2216 83.2252 76.4628

0.0001 0.0001 0.0001

0.0003 0.0001 0.0001

(b) ImgNet

A PPENDIX F S CALABILITY A NALYSIS (c) CIFAR-10

(d) CIFAR-100

Fig. 10: (a)-(d): The ACCs and ASRs of PASTA against FP after the corresponding percentage of neurons is pruned.

A PPENDIX E ATTACK E FFECTIVENESS AGAINST VARIOUS V I T A RCHITECTURES TABLE IX: The ACCs and TRE of PASTA against various ViT architectures. Model Arch. ViT-Large [8] BEiT-Base [68] CaiT-XXSmall [67] DeiT-Tiny [66]

Clean model

Poisoned model

ACC

ACC

TRE

98.45 97.86 95.20 92.60

96.39 92.00 94.00 91.40

99.97 99.62 99.98 99.40

To further confirm that our backdoor attack achieves superior TREs against various ViT structures, we execute our backdoor attack on 4 advanced ViTs, i.e., BEiT [68], CaiT [67] DeiT [66] and ViT-Large [8]. We show ACCs of clean models, along with ACCs and TREs of poisoned models in Table IX. According to the results, ours achieves an average TRE of 99.74% across the 4 heterogeneous ViT architectures, with an average ACC drop of only 2.58%. These results demonstrate that PASTA is broadly compatible with diverse ViT architectures and consistently delivers strong attack performance. TABLE X: The natural stealthiness (PSNR↑, SSIM↑, LPIPS↓, l2 -norm↓) and attention stealthiness (APSNR↑, ALPIPS↓, ARES↓, l2 -norm↓) of PASTA by comparing clean and poisoned images across various TALs on the CIFAR-10 dataset. PASTA consistently maintains twofold stealthiness across TALs. Natural Stealthiness

TAL

Attention Stealthiness

l2

PSNR

SSIM

LPIPS

l2

APSNR

ALPIPS

ARES

Clean

0.0000

Inf

1.0000

0.0000

0.0000

Inf

0.0000

0.0000

(0,0) (7,7) (13,13)

0.4922 0.2705 0.2706

64.5016 69.6996 69.6972

0.9961 0.9997 0.9997

0.0001 0.0001 0.0001

1.0967 0.5201 0.2052

56.6624 83.6460 92.0588

0.0008 0.0001 0.0001

0.0004 0.0001 0.0001

We assess the scalability of our backdoor attack by measuring resource usage on 4 datasets with ViT-Tiny model architecture. Specifically, we record the average RAM usage (GB), GPU memory consumption (GB), and per-epoch time cost (s) for each optimization task, along with the total time (s) required for the entire backdoor attack workflow. The results are summarized in Table XII. We observe similar RAM usage for trigger (3.48 GB) and model optimization (3.30 GB) due to the same data pipeline. However, GPU memory is higher for trigger optimization (3.86 GB vs. 2.42 GB) because of attention-stealthiness optimization, which requires two forward-backward passes per batch to compute clean–poisoned attention discrepancies, unlike standard singlepass training, increasing gradient caching cost. Model optimization is 15.76× slower than trigger optimization, as it uses the full dataset and a larger parameter space, leading to higher backpropagation cost. With fixed batch size, PASTA scales linearly with dataset size. Overall, our attack is resourceefficient and time-scalable.

A PPENDIX G I MPACT OF TAL S ON T WOFOLD S TEALTHINESS This work achieves superior twofold stealthiness in our backdoor attack. We present the poisoned images and their attentions against various attacks in Figures 3 and 4 respectively on a fixed TAL. To further validate that the twofold stealthiness of poisoned images generated by our backdoor attack is not affected by the choice of TALs, we present additional poisoned samples using three different TALs: the patch at (0,0) in the top-left corner, (7,7) in the center, and (13,13) in the bottomright corner. The visualizations of clean and poisoned images are shown in Figures 11 and 12 , the corresponding attention maps are visualized in Figures 13 and 14. We further showcase the quantitalized visual and attention stealthiness of PASTA across various TALs under CIFAR-10 and Sub-ImgNet in Tables X and XI. Notably, using different TALs (e.g., (7,7) or (13,13)) on CIFAR-10 could improve twofold stealthiness. These results demonstrate that PASTA maintains consistent visual and attention stealthiness across different TALs.

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

18

TABLE XII: The time and resource usage of PASTA across datasets on ViT-Tiny. Trigger Optimization

# Samples (×103 )

Dataset

Total Time (s)

GPU Mem (GB)

Time (s)

RAM (GB)

GPU Mem (GB)

Time (s)

3.66 3.64 3.68 2.96

3.86 3.86 3.86 3.86

246.06 246.00 114.62 5188.48

3.67 3.65 3.69 2.92

2.42 2.42 2.42 2.42

3109.18 3097.14 842.72 78471.05

3356.31 3344.43 957.96 83661.15

Patch(13,13) Patch(7,7) Patch(0,0)

Patch(13,13) Patch(7,7) Patch(0,0)

Clean

50 50 12 128

Clean

CIFAR-10[58] CIFAR-100 [58] Sub-ImgNet[2] ImgNet[2]

Model Optimization

RAM (GB)

Fig. 11: Visualization of clean and poisoned images generated by PASTA on the CIFAR-10 dataset, with triggers inserted at different TALs. The results demonstrate that the poisoned images remain visual imperceptible, regardless of TALs.

A PPENDIX H A DAPTIVE D EFENSE Liu and Qiao et al. [24] report that triggers in the spatial domain are inherently vulnerable to image transformations. Consequently, our patch-wise triggers also lack robustness to common transformations, as they are inserted at specific spatial locations. In Tables XIII and XIV, we showcase ACCs/ASRs of PASTA against Gaussian filters with two window sizes across various backdoor payload strategies. Note that a larger window size results in stronger smoothing effect. For PASTA (Fixed and Rand 1 TAL), the average ASR drops from 86.72% (3×3 filter) to 22.85% (5×5 filter) on CIFAR-10, and from 48.34% (3×3) to 0.98% (5×5) on Sub-ImgNet. Meanwhile, ACC before and after filtering differs by only 0.32% for both window sizes across the two datasets. These results confirm that Gaussian filtering effectively disrupts backdoor triggers while preserving clean features. However, when using our stronger payloads (Fixed and Rand 10 TALs), attack robustness improves significantly: ASRs reach 99.9% (3×3) and 71.77% (5×5) on CIFAR-10, and 99.61% (3×3) and 3.91% (5×5) on Sub-ImgNet. Additionally, increasing TALs (e.g., Rand 20 TALs) further enhances attack effectiveness against filtering, yielding an average ASR increase of 2.2%. This is because incorporating multiple TALs in our backdoor payload increases the possibility that our patch-wise trigger is placed in semantically rich regions of the image, where pixels are

Fig. 12: Visualization of clean and poisoned images generated by PASTA on the Sub-ImgNet dataset, with triggers inserted at different TALs. The results demonstrate that the poisoned images remain visual imperceptible, regardless of TALs.

more resilient to image transformations. TABLE XIII: The ACCs and ASRs before and after Gaussian filters with two window sizes on CIFAR-10. Attack Payload PASTA (Fixed 1 TAL) PASTA (Fixed 1 TAL) PASTA (Fixed 10 TALs) PASTA (Fixed 10 TALs) PASTA (Rand 1 TAL) PASTA (Rand 1 TAL) PASTA (Rand 10 TALs) PASTA (Rand 10 TALs) PASTA (Rand 20 TALs) PASTA (Rand 20 TALs)

Window Size 3×3 5×5 3×3 5×5 3×3 5×5 3×3 5×5 3×3 5×5

Before

After

ACC

ASR

ACC

ASR

97.65 97.07 96.48 96.48 96.87 96.28 96.28 95.70 97.26 96.67

99.60 100.00 100.00 100.00 99.80 100.00 100.00 100.00 100.00 100.00

97.85 97.07 96.67 96.28 96.87 96.09 96.09 95.70 97.07 96.48

91.21 27.53 100.00 82.22 78.32 18.16 99.80 61.32 100.00 67.57

A PPENDIX I PASTA E XTENSION AND AUTOMATED PARAMETER S ELECTION Automated Parameter Selection. While PASTA is designed for ViTs in image classification tasks, we plan to extend it to other vision models, such as transformer-based diffusion models and large vision-language models, as well as to broader CV tasks, e.g., object detection. As shown in Table XII, PASTA requires up to 8×104 seconds on large dataset such as ImgNet, which is computationally intensive. We will explore advanced optimization techniques to improve attack efficiency

Patch(13,13) Patch(7,7) Patch(0,0)

Clean

IEEE TRANSACTIONS ON INFORMATION FORENSICS AND SECURITY, VOL. XX, 2026

Patch(13,13) Patch(7,7) Patch(0,0)

Clean

Fig. 13: Visualization of attention on clean and poisoned images generated by PASTA on the CIFAR-10 dataset, with triggers inserted at different TALs. The results demonstrate that the attention disparity between clean and poisoned images is imperceptible, regardless of TALs.

Fig. 14: Visualization of attention on clean and poisoned images generated by PASTA on the Sub-ImgNet dataset, with triggers inserted at different TALs. The results demonstrate that the attention disparity between clean and poisoned images is imperceptible, regardless of TALs.

on large-scale datasets. Since attack objectives (i.e., twofold stealthiness and attack effectiveness) could be affected by Lagrange coefficients in our optimization problem (see Figure 5), We will explore adaptive coefficient selection mechanisms. Extending PASTA to Additional CV Tasks. It is feasible to extend PASTA to broader CV tasks such as object detection (OD) [75]. For example, DETR adopts a transformer backbone closely related to ViTs, making it possible to transfer our MIS mechanism to achieve high TRE; moreover, DETR optimizes parameters via end-to-end loss minimization, allowing our

19

TABLE XIV: The ACCs and ASRs before and after Gaussian filters with two window sizes on Sub-ImgNet. Attack Payload PASTA (Fixed 1 TAL) PASTA (Fixed 1 TAL) PASTA (Fixed 10 TALs) PASTA (Fixed 10 TALs) PASTA (Rand 1 TAL) PASTA (Rand 1 TAL) PASTA (Rand 10 TALs) PASTA (Rand 10 TALs) PASTA (Rand 20 TALs) PASTA (Rand 20 TALs)

Window Size 3×3 5×5 3×3 5×5 3×3 5×5 3×3 5×5 3×3 5×5

Before

After

ACC

ASR

ACC

ASR

81.64 82.61 80.66 82.03 82.22 85.15 80.46 83.39 85.54 83.39

100.00 100.00 100.00 100.00 100.00 100.00 100.00 100.00 100.00 100.00

80.85 81.44 79.88 81.83 82.42 83.98 81.25 81.83 85.35 83.39

62.30 1.17 100.00 4.49 34.38 0.78 99.21 3.32 99.21 5.66

twofold-stealthiness loss to be integrated into its training objective. However, existing OD attacks [76] rely on visible triggers, which contradicts our twofold-stealthiness goal. Extending PASTA to a visible-trigger variant for OD remains an interesting direction for future work.

Record · ID 123967 · SHA-256 fcef66fdb2ddf334
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.