SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
1
S TYLE AT: Defending Face Recognition Against Semantic Attacks Ben Shapira1
1
Tel Aviv University Tel Aviv, Israel
2
HPI / University of Potsdam Potsdam, Germany
3
National Taiwan University Taipei, Taiwan
arXiv:2609.23596v1 [cs.CV] 20 Sep 2026
Roi Cohen2 [email protected]
Shang-Tse Chen3 [email protected]
Mahmood Sharif1 [email protected]
Abstract With face-recognition models now embedded in everyday authentication and surveillance, recent works have pinpointed a critical weakness: these models remain acutely vulnerable to adversarial semantic edits. I.e., adversarially produced semantic alterations to the input, such as slight aging or pose changes, can induce misclassifications. Certain existing attacks are powerful, but they can be computationally costly, rendering them inadequate for developing defenses (e.g., through adversarial training). To fill the gap, we introduce B OUND S TYLE, a potent semantic attack operating in StyleGAN’s rich latent space to maximize misclassification rates. Notably, B OUND S TYLE achieves high attack success rates while being ∼×9.5 faster than existing state-of-the-art attacks, making it suitable for adversarial training. Building on B OUND S TYLE, we develop S TYLE AT, an efficient adversarial training scheme that incorporates low-budget attack variants yet defends against stronger and unseen semantic attacks. We evaluate on two datasets unseen during training and seven models, and find that S TYLE AT boosts robust accuracy against state-of-the-art attacks and outperforms common defenses in various settings.
1
Introduction
Face-recognition (FR) technologies are employed in various security-critical applications, including for surveillance and access control [16]. Failures of such systems may be pernicious; for instance, false negatives may enable criminals to avoid surveillance, whereas false positives may provision unauthorized access to important resources protected by access control. However, unfortunately, similar to other machine-learning (ML) models that can be evaded by adversarial examples at inference time [11, 39], FR models are vulnerable to general semantic attacks—a class of adversarial example attacks that introduce slight semantic changes (e.g., addition of accessories or alterations of pose, expression, or age) to fool FR despite preserving the identity of the subject in the image [2, 17, 23, 24, 26]. Existing general semantic attacks mostly rely on generative models, such as StyleGAN [23] and latent diffusion models [24], to discover adversarial semantic edits through latent-space © 2026. The copyright of this document resides with its authors. It may be distributed unchanged freely in print or electronic forms.
2
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
perturbations that mislead FR. However, these attacks suffer from certain limitations: Some attacks such as StyleAdv achieve limited attack success [24]; other attacks such as AMTGAN produce edits with low visual fidelity [23]; and attacks such as DiffPrivate are computationally heavy, requiring significant run time to attain high success rates (see §6). Due to these limitations, existing attacks may fail to uncover weakness in FR or may not lend themselves to being incorporated in training schemes for improving FR’s robustness. For specific forms of adversarial examples, such as those created by adversarial perturbations with bounded ℓ p -norms, numerous defense types like adversarial training [41] and randomized smoothing [6] can help boost robustness. However, to our knowledge, no established defenses have been proposed to mitigate general semantic attacks. Particularly, adversarial training for defending against general semantic attacks remains infeasible due to the run-time overhead or limited success rates of existing attacks. Consequently, FR remains vulnerable to general semantic attacks. To fill these gaps, this work presents a new general semantic attack, B OUND S TYLE, and a defense, S TYLE AT. Our attack takes advantage of StyleGAN3’s rich latent space [18], among others, to produce high-fidelity semantic edits to fool FR. Importantly, B OUND S TYLE is tunable, enabling us to control the magnitude of edits and run time, thus ensuring that identity is preserved w.r.t. human observer and allowing us to execute time-efficient variants. Notably, we also find that B OUND S TYLE is highly successful, while being significantly faster (roughly ×9.5) than the state-of-the-art attack, DiffPrivate [24]. Interestingly, we find compelling evidence that DiffPrivate introduces imperceptible adversarial perturbations alongside visible semantic edits (App. D); if used for adversarial training, models may learn to resist pixel noise rather than semantic changes. This motivates B OUND S TYLE’s design to operate purely in StyleGAN’s semantic latent space. Supporting this design choice, we find no clear signs that B OUND S TYLE makes edits other than semantic ones (§6.2 consolidates the supporting evidence). Altogether, B OUND S TYLE’s advantages render it suitable for measuring the susceptibility of FR models to attacks as well as for adversarial training to help improve robustness against semantic attacks. Our defense, S TYLE AT, employs a time-efficient variant of B OUND S TYLE to adversarially train FR models and improve their adversarial robustness against general semantic attacks. Against B OUND S TYLE, S TYLE AT achieves up to 28.6% increase in robust accuracy (depending on the setting explored) compared to undefended models, markedly higher than defenses not tailored for general semantic attacks that achieve ≤6.0% increase in robust accuracy. Crucially, S TYLE AT also leads to improvements against DiffPrivate, an attack not encountered during training, with up to 46.3% higher robust accuracy than undefended models, showcasing that S TYLE AT generalizes to unknown general semantic attacks. We next present related work (§2) and our threat model (§3). Subsequently, we describe the technical approach behind B OUND S TYLE and S TYLE AT (§4) before presenting our experimental results (§5–6) and concluding (§7).
2
Related Work
Attacks on FR Prior work has demonstrated that ML models in general, and FR in particular, are vulnerable to test-time evasion attacks that induce misclassifications via imperceptible adversarial perturbations with bounded ℓ p -norm [39]. For instance, the fast gradient sign method (FGSM) creates attacks by perturbing inputs in the gradient direction
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
3
once [11], while projected gradient descent (PGD) does so through multiple, iterative perturbations [28]. However, such attacks may be challenging to realize in real-world settings due to difficulties in implementing norm-bounded noise and cameras’ sampling errors, among others [38]. To this end, researchers have proposed semantic attacks—attacks that alter inputs in minor, easy-to-realize, and semantically meaningful ways—to mislead FR models. Semantic attacks consist of two families. The first family of attacks makes ad hoc changes to inputs, for example, by introducing adversarial accessories like eyeglasses or hats to fool models [20, 38]. These also include attacks that fool models via facial make-up or spatial transformations applied in an adversarial manner [14, 45, 47]. By contrast, the second family of attacks leverages general edits of inputs to induce misclassifications, including, but not limited to, changes of expression, age, and accessories, or a combination thereof [2, 17, 23, 24, 26]. Our work focuses on general semantic attacks, proposing a new attack and a defense. General semantic attacks typically leverage generative models to produce adversarial edits of inputs. For instance, attacks such as StyleAdv [23] and Adv-Attribute [17] search for adversarial editing directions in the latent space of generative adversarial networks (GANs) to produce misclassifications. By contrast, Adv-Diffusion [26] and DiffPrivate [24] use latent diffusion models to find adversarial semantic edits of inputs. DiffPrivate is the most recent and potent general semantic attack; we use it in our evaluation. Defending FR A diversity of defenses against evasion attacks have been proposed, including, but not limited to, ones that detect attacks (e.g., [31]); filter out adversarial perturbations (e.g., [46]); smoothen classification boundaries to reduce model vulnerability (e.g., [4, 6]); verify robustness against specific adversaries (e.g., [19]); and adversarial train of models by injecting correctly labeled adversarial inputs to the training data to inherently increase model robustness (e.g., [21, 28]). Due to its intuitive nature, its ability to improve adversarial robustness in a practical manner against different attack types, and absence of impact on model’s inference time, Adversarial training is particularly appealing and was widely studied. Still, adversarial training may be computationally expensive due to the overhead of producing attacks during training, potentially rendering training prohibitive. To this end, researchers have also explored efficient adversarial training variants (e.g., [37, 41]). We take inspiration from Wong et al. [41] who showed how to leverage the efficient FGSM attack in training to induce robustness against much more potent attacks at test time. To the best of our knowledge, there are no established defenses for countering general semantic attacks against FR. Nonetheless, several countermeasures have been proposed to counter ad hoc semantic attacks. For example, defense through occlusion attack (DOA) adversarially trains models with carefully positioned patches containing adversarial patterns to help counter eyeglass attacks [42]. As another example, DiffPure utilizes forward diffusion followed by image recovery to remove adversarial manipulations, helping counter spatial adversarial modifications [33]. Prior work has also shown that input filters, such as JPEG compression, blurring, and Gaussian noise can hinder general semantic attacks to some degree when attacks are agnostic to the filters [24]. Our evaluation (§6) shows that these approaches yield limited improvements in robustness; e.g., DiffPure proves effective against DiffPrivate but is counterproductive against B OUND S TYLE, illustrating that the effectiveness of purification-based defenses can be attack-type dependent. Related to our work, Laidlaw et al. [22] adversarially train ML models with imperceptible adversarial perturbations created using generative models. However, they only evaluate robustness against imperceptible perturbations and spatial manipulations.
4
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
Relation to Deepfake Detection As adversarial semantic edits are synthetically produced, deepfake detection [32] may seem like a natural countermeasure. However, detection and robust recognition solve different, complementary problems: a detector flags whether the image synthetic or edited, whereas FR must decide identity—i.e., whether a face matches the enrolled subject despite the edit. Even a perfect detector leaves the identity decision open; moreover, as edits may be benign (e.g., beauty filters), rejecting all flagged images would impose a high false-positive burden on legitimate users. Detection alone may also be insufficient: detectors often generalize poorly to unseen generative models [7, 34] and can be evaded by white-box adversaries such as ours [3]. Accordingly, we view detection as a complementary defense layer, while S TYLE AT hardens the component selecting identity.
3
Threat Model
We consider an adversary carrying out a general semantic attack against FR. Per standard practice, we assume the FR system is tuned to an operating point where the false positive rate (FPR) is below a target threshold, such as 0.01 FPR [16, 24]. We consider a powerful adversary aiming to produce arbitrary, untargeted misclassifications (primarily, false negatives) rather than impersonations (i.e., targeted attacks), as, intuitively, this adversary would be more challenging to defend against. Contrastively, our defender aims to hinder the adversary’s attempts through keeping the robust true positive rate (TPR)—i.e., the TPR under attacks—high while preserving the benign accuracy of standard FR models when ingesting clean inputs. In line with Le and Carlsson [24], We primarily focus on potent adversaries with white-box access to both the FR model and the defense, but we also consider black-box adversaries without access to the model or defense that seek to transfer attacks from surrogate models, as well as gray-box adversaries that have access to the (undefended) model but not to the defense. We assume a fully digital threat model where the adversary modifies a digital image and submits it electronically. Physical re-capture scenarios (e.g., adversarial accessories worn in front of a camera) are outside scope, consistent with the digital-first framing of DiffPrivate [24] and related semantic attacks. Our threat model captures settings where images may be manipulated digitally or physically to mislead FR. In cases where images arrive from devices the FR operator does not control, such as uploads to social media or electronic know-your-customer onboarding, attacks may alter images to gain privacy or impersonate others. Moreover, in idealized settings without sampling noise, semantic attacks’ edits encompass changes that may be applied physically (e.g., make-up or pose changes) to fool FR, providing an upper bound on attack success in the physical domain.
4
Technical Approach
4.1
B OUND S TYLE: A High-Fidelity, Potent, Tunable General Semantic Attack
We design B OUND S TYLE as a general semantic attack against FR based on generative models with three goals in mind: (1) We require that the attack produces evasive face images with high visual fidelity through diverse and realistic edits; (2) We seek tunability such that we
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
5
would be able to control the run time of the attack to later enable efficient adversarial training and bound the magnitude of the edit so as the identity in the face image is unchanged (for a human observer); and (3) We need the attack to be potent exposing the weaknesses of FR through achieving high success rates. We next describe how our design of B OUND S TYLE ensures high visual fidelity and tunability. Our experiments (§5–6) evidence the attack’s potency. A Tunable Attack Let F denote the feature extractor used for FR, C the preprocessing algorithm (cropping and alignment), G a generative model, x the image to modify with an inverted latent code l, and xe a face image of the same subject enrolled in the gallery. B OUND S TYLE aims to edit x through a slight modification of l such that the similarity (sim, usually cosine similarity) with xe would become small. Formally, B OUND S TYLE aims to minimize the following loss through a perturbation δ of the latent code: Latk = sim F(C(G(l + δ ))), F(C(xe )) . B OUND S TYLE optimizes the loss through iterative gradient-based optimization, in the spirit of PGD, and its performance is governed by two primary inputs T , the number of iterations, and β , the magnitude (specifically, ℓ2 -norm) of the perturbation δ . Initially, δ0 is randomly initialized inside the β -ball, as random initialization is critical to the performance of evasion attacks in adversarial training [41]. Subsequently, in each iteration (up to T ), g where g = ∇δi Latk is the loss gradient and α is a B OUND S TYLE updates δi = δi − α · ∥g∥ ∞ step-size hyperparameter. At any point, if the norm of δi exceeds the bound β , it is projected back to the β -ball via δi = β · ∥δδii∥ . 2 Both T and β are tunable parameters that enable achieving different trade-offs with B OUND S TYLE. Decreasing T may potentially harm the attack success, but also makes B OUND S TYLE faster, rendering it more suitable for (efficient) adversarial training. Moreover, setting β should balance two goals—it should be large enough so that the attack is successful due to stronger edits in the latent space, but small enough so that the identity of the subject is preserved w.r.t. human observers. Ensuring High Fidelity We take several measures to ascertain that B OUND S TYLE introduces high-fidelity edits. First, we adopt the StyleGAN3 generator [18], which provides a rich latent space with diverse editing directions and high quality outputs. Second, we invert x to a latent code l that maps back almost precisely to x, thus preserving identity. To do so, we use a hybrid combination of encoder-based projection to the latent space [1] followed by direct gradient-based optimization for accurate reconstruction of the face image [48]. Last, we use pivotal tuning, a method that tunes the generator G to enable better editability while preserving identity [36].
4.2
S TYLE AT: Style-Aware Adversarial Training
Building on B OUND S TYLE, we design S TYLE AT, a method for adversarially training FR models to enhance their adversarial robustness against general semantic attacks. In particular, S TYLE AT fine-tunes pre-trained FR feature extractors while aiming to balance three different objectives: (1) Preserving benign accuracy on clean images; (2) Improving robust accuracy against general semantic attacks; and (3) Countering imperceptible perturbations inadvertently introduced by certain established semantic attacks. To achieve each of these goals, S TYLE AT minimizes the triplet losses LCln , LAdvSem , and LAdvPix , respectively. These losses are balanced through non-negative hyperparameters that accumulate to
6
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
Latent Branch
StyleGAN Inversion
BoundStyle Attack
Align & Crop
FR Feature Extractor
Pixel Branch PGD
Clean Branch
Figure 1: An overview of S TYLE AT’s training pipeline. The framework processes a minibatch of identity pairs {(xi , xi+ )} through three parallel branches: (1) The clean branch extracts features from the original samples to preserve benign accuracy (LCln ). (2) The latent branch generates a batch of semantic adversarial examples xiadv via B OUND S TYLE to improve robustness against semantic attacks (LAdvSem ). (3) The pixel branch applies imperceptible perturbations to the semantically edited batch to increase robustness against imperceptible perturbations (LAdvPix ). Batch processing enables hard negative mining to optimize the triplet losses.
one (i.e., λCln + λAdvSem + λAdvPix = 1). We next elaborate how each of LAdvSem and LAdvPix are computed and optimized and how we select triplets for loss computation; minimizing LCln is intuitive and follows standard practice [40]. An overview of S TYLE AT’s pipeline is depicted in Fig. 1; Alg. 1 in App. A presents its pseudocode. Computing and Optimizing LAdvSem We leverage B OUND S TYLE to produce general semantic attacks for adversarial training. However, as executing the most potent attack variant during training may make the training process infeasible, we incorporate “weakened” but efficient attack variants into training. Specifically, we run fast variants of B OUND S TYLE with a small number of iterations T , analogously to fast adversarial training with FGSM [41]. Here, we tune T and the step size α such that training can be completed within a few days under our resource constraints, while the attacks are sufficiently evasive to help enhance FR’s robustness against general semantic attacks. Computing and Optimizing LAdvPix Our evaluation of established semantic attacks shows that certain attacks may introduce imperceptible perturbations alongside semantic edits to mislead FR. This phenomenon is perhaps most clearly demonstrated when evaluating attacks against filter-based defenses such as JPEG compression that primarily affect imperceptible, non-semantic perturbations. Said differently, if such filters have a pronounced effect on an attack’s success, one may conclude that misclassification did not occur due to a semantic edit, but rather due to imperceptible changes of pixels. Indeed, our evaluation shows that DiffPrivate exhibits a significant decrease in their success once JPEG compression and similar filters are employed (see Fig. 4 and Fig. 10, with App. D providing additional support through a frequency-domain analysis). To account for these potential perturbations, we also adversarially train our models against imperceptible ℓ∞ -norm-bounded adversarial perturbations. We find that doing so does not harm robustness against semantic attacks that do not seem to introduce imperceptible adversarial perturbations (namely, B OUND S TYLE; App. C).
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
7
Toward countering imperceptible adversarial perturbations, we integrate fast PGD attacks with few iterations into training, in the spirit of Wong et al.’s (2020) work. Importantly, to avoid robust overfitting, we perform PGD with random initialization before each attack. Crucially, we apply PGD to images already edited adversarially with B OUND S TYLE, as we aim to counter the combination of adversarial semantic edits and imperceptible perturbations. Selection of Triplets For a given positive pair of samples depicting the same identity, we select the hardest negative sample from the batch to compute triplet losses. Doing so, as shown in prior work on adversarially robust metric learning [29], is most conducive for adversarial robustness. More precisely, to compute the triplet loss, for a positive pair of samples p and a (standing for positive and anchor, respectively), we select the negative sample n from the batch such that it depicts a different identity and is most similar to a in the feature space compared to other samples in the batch. Subsequently, the triplet loss is calculated by (sim(F(C(n)), F(C(a))) − sim(F(C(p)), F(C(a))) + µ)+ , where µ is a small (non-negative) constant. Moreover, in the interest of improving adversarial robustness, we adversarially perturb the anchor sample a when computing LAdvSem and LAdvPix , and select the hardest negative sample after applying the perturbations, per Li et al. [25].
5
Experimental Setup
FR Backbones We employ seven popular and high-performing FR models, four convolutional networks, one vision transformer, and two convolutional networks equipped with specialized supervisory heads. Specifically, we use MobileFaceNet (MobileFace) [5]; ResNet (ResNet-152, IR-SE) [13]; RepVGG [9]; LightCNN [43, 44]; Swin Transformer (SwinT) [27]; and ArcFace [8] as well as MagFace [30] with MobileFace backbones. We use these models in two roles, both as targets for attacks, and as surrogates (i.e., proxies) for producing transferable adversarial examples. As part of S TYLE AT, we create adversarially trained variants of ResNet and RepVGG through fine-tuning the original pre-trained backbones. We obtain the initial weights from FaceX-Zoo [40]. Datasets Following FaceX-Zoo [40], we construct our training dataset from MS-Celeb-1Mv1c [12], using their randomly selected preprocessed positive image pairs for training. We then run our preprocessing pipeline on all images and discard pairs where an image fails face or landmark detection, yielding a final training set of 52,269 positive image pairs. (Note that negative images are selected as hard negatives, independently for each batch during training.) We evaluate on two prominent datasets: (1) Labeled Faces in the Wild (LFW) [15] and (2) VGG-Face [35]. Specifically, we select 216 and 156 positive pairs from LFW and VGG-Face, respectively. When selecting images, we ensure no overlap between the selected identities and those appearing in the training (and pre-training) sets from MS-Celeb-1M-v1c, by removing any image that has similarity with any training samples above the threshold where the original ResNet backbone has an FPR of 10−4 . Attacks We evaluate B OUND S TYLE, bounding the edit perturbation ℓ2 -norms in StyleGAN3’s latent space to β ∈ {1.0, 1.5, 2.0, 3.0}. We avoid perturbations of larger magnitude to help preserve subject identities in images (App. B). For best performance, we set the step size α=β , and the number of iterations T =30, as more iterations show no improvements in attack success (App. C). As a baseline, we evaluate DiffPrivate [24], a state-of-the-art attack that edits images through perturbations in the diffusion latent z-space. For a fair comparison with B OUND S TYLE, we adopt a norm-bounded variant of DiffPrivate by enforcing
8
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
∥∆z∥2 = ∥zadv − z0 ∥2 ∈ {1, 2, 3, 4, 5, 6}. We avoid perturbations with norm >6 to preserve subject identities in images (App. B). Under this bounded setup, we find that capping the optimization at 70 iterations for convergence to the highest attack success (App. C). Defenses We apply S TYLE AT on both ResNet and RepVGG models, adversarially training them from checkpoints obtained from FaceX-Zoo. We use a low-cost variant of B OUND S TYLE for adversarial training, with T = 3 attack iterations, β = 3 ℓ2 -norm for perturbations in the StyleGAN3 latent space, and α = 3 as step size in the attack; we run PGD attacks for 2 iterations with ε = 24/255 ℓ∞ -norm for perturbations in the pixel space and step size α = ε/2. We find that these parameters help attain reasonable benign accuracy within feasible time under our resource constraints. After hyperparameter search, we set specific loss weights for each backbone: for ResNet, we set λCln = 0.35, λAdvSem = 0.45, and λAdvPix = 0.2; for RepVGG, we set λCln = 0.1, λAdvSem = 0.8, and λAdvPix = 0.1. Other parameters (e.g., for the optimizer and triplet loss margin) are adopted from FaceXZoo. We run training on 8 NVIDIA RTX A5000 GPUs with a batch size of 4 per GPU (global batch size 32). We run training for 8 epochs, completing it within 4 days. As baselines, we compare S TYLE AT with DOA, a defense tailored for ad hoc semantic attacks using adversarial patches or eyeglasses (see §2), DiffPure [33], a diffusion-based test-time purification defense, and standard filters considered in prior work [24], including Gaussian blur, denoising via total variation minimization, JPEG compression, feature squeezing, spatial smoothing, and random noise injection. Computational Costs All preprocessing and training runs on 8 NVIDIA RTX A5000 GPUs. GAN inversion (150 steps per image, ∼4.7 GPU-days) and PTI fine-tuning of the StyleGAN generator (125 steps, ∼4.3 GPU-days; reducible to ∼3.4 GPU-days at 100 steps) are onetime costs for the training set, shared across all training runs. Unlike these preprocessing steps, adversarial training is not one-time: training each FR backbone takes approximately 4 days, and multiple runs were required for hyperparameter search, per standard practice for adversarially trained models. Metrics and FR Operating Point We evaluate model robustness through accuracy (i.e., TPR) after perturbing one of the samples from a positive pair. Following standard practice, we calibrate the verification threshold to meet a target FPR on clean data (without attacks). Specifically, we perform the calibration on LFW’s full (clean) validation set (containing negative and positive pairs) to obtain an FPR of 0.01, similar to [24].
6
Experimental Results
We now evaluate the B OUND S TYLE attack and S TYLE AT defense.
6.1
B OUND S TYLE is Potent and Fast
White-box Attacks Tables 1 and 2 report the robust accuracy of the seven (undefended) FR backbones against B OUND S TYLE and DiffPrivate, respectively, on the LFW and VGG-Face datasets when varying the attack budgets. The results show that both attacks are potent in the white-box setting—despite high benign accuracy (>99%), the robust accuracy drops as the attack budgets increase. At B OUND S TYLE’s highest budget (β =3), it reduces average robust accuracy to ∼46% on LFW and ∼37% on VGG-Face. At DiffPrivate highest budget (∥∆z∥=6), it reduces average robust accuracy to ∼35.5% on LFW and ∼16% on VGG-Face.
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS Dataset
Model
Clean B OUND S TYLE - budget β 1
1.5
2
3
94.9 90.3 91.7 92.6 93.5 94.0 94.4
88.0 84.7 88.0 83.3 83.8 88.9 85.2
78.7 72.7 75.0 69.4 70.8 77.8 74.1
52.6 40.5 50.2 40.5 39.5 52.2 45.0
SwinT 100.0 85.3 80.8 65.4 LightCNN 98.7 84.0 72.4 60.3 MobileFace 97.4 78.8 72.4 62.2 VGG-Face RepVGG 100.0 80.8 71.2 60.3 ResNet 100.0 84.0 72.4 59.6 MagFace 98.1 75.6 67.3 57.7 ArcFace 98.7 80.8 71.8 52.6
45.5 34.0 38.5 32.0 34.6 40.8 34.4
LFW
9
SwinT LightCNN MobileFace RepVGG ResNet MagFace ArcFace
99.1 99.1 99.1 99.1 99.1 99.1 99.5
Table 1: White-box robustness against B OUND S TYLE† on LFW and VGG-Face. We report benign accuracy and robust accuracy under increasing semantic attack budgets β ; higher is better. † 95% bootstrap CIs in App. H. Dataset
Model
DiffPrivate - budget ∥∆z∥
Clean 1
2
3
4
5
6
SwinT LightCNN MobileFace RepVGG ResNet MagFace ArcFace
99.1 99.1 99.1 99.1 99.1 99.1 99.5
99.5 98.6 98.6 85.6 50.5 49.1 99.1 97.7 89.4 62.0 32.9 28.7 98.6 98.1 96.3 79.2 53.7 38.0 99.1 99.1 96.3 72.7 43.5 38.0 99.1 98.6 93.5 71.8 41.7 33.8 98.6 98.6 96.3 74.5 44.9 28.7 99.1 98.6 91.7 69.0 38.4 31.5
SwinT LightCNN MobileFace VGG-Face RepVGG ResNet MagFace ArcFace
100.0 98.7 97.4 100.0 100.0 98.1 98.7
83.9 76.8 64.5 41.9 26.5 19.4 69.7 58.7 45.8 25.2 18.1 14.8 70.3 66.5 47.7 34.2 21.9 16.1 77.4 69.7 53.5 32.3 22.6 17.4 77.4 67.7 51.0 30.3 22.6 17.4 69.0 67.7 52.3 31.0 20.0 14.8 71.6 58.7 44.5 27.7 20.0 14.2
LFW
Table 2: White-box robustness against DiffPrivate† on LFW and VGG-Face. We report benign accuracy and robust accuracy under increasing latent perturbation norms ∥∆z∥; higher is better. † 95% bootstrap CIs in App. H.
Black-box Attacks Fig. 2 presents the robust accuracy of models when transferring attacks between models at the highest attack budgets considered. In the heatmaps, off-diagonals show robust accuracy in black-box settings (transfer from the rows to columns), while diagonals correspond to white-box settings. B OUND S TYLE exhibits strong transferability—on both datasets, the heatmaps are near-uniform with off-diagonals typically within 5–15% of
10
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
100 80 60 40 20 0
SwinT LightCNN MobileFace % RepVGG ResNet MagFace ArcFace
49 54 74 69 65 69 65 97 29 85 90 90 78 70 91 47 38 75 77 53 52 90 58 79 38 75 69 66 91 65 81 70 34 74 72 92 48 70 80 83 29 49 96 53 75 86 87 62 31
S Lig winT Mo htCN bil N eF Re ace pV G Re G sN Ma et gFa Arc ce Fac e
53 64 62 56 53 72 68 68 40 60 59 60 65 61 63 55 50 56 57 57 57 59 50 59 40 50 64 58 55 53 59 55 40 67 62 63 55 59 57 57 52 57 62 51 56 58 60 61 45
S Lig winT Mo htCN bil N eF Re ace pV G Re G s Ma Net gFa Arc ce Fac e
SwinT LightCNN MobileFace RepVGG ResNet MagFace ArcFace
(a) LFW – B OUND S TYLE (β =3)
(b) LFW – DiffPrivate (∥∆z∥=6)
100 80 60 40 20 0
SwinT LightCNN MobileFace % RepVGG ResNet MagFace ArcFace
19 32 36 42 41 30 35 70 15 39 57 61 41 40 64 28 16 53 52 29 30 64 24 39 17 46 36 36 64 28 41 43 17 33 34 71 21 30 54 55 15 25 68 23 34 55 57 32 14
S Lig winT Mo htCN bil N eF Re ace pV G Re G sN Ma et gFa Arc ce Fac e
46 53 52 48 46 56 54 51 34 47 48 47 42 47 54 49 38 47 48 38 40 50 46 51 32 43 43 42 53 46 48 46 35 42 44 49 44 42 47 45 41 46 51 38 33 47 43 38 34
S Lig winT Mo htCN bil N eF Re ace pV G Re G s Ma Net gFa Arc ce Fac e
SwinT LightCNN MobileFace RepVGG ResNet MagFace ArcFace
(c) VGG – B OUND S TYLE (β =3)
(d) VGG – DiffPrivate (∥∆z∥=6)
100 80 60 40
%
20 0
100 80 60 40 20 0
Figure 2: Robust accuracy of models against black-box attacks. Each heatmap reports the robust accuracy of the model at the column, when transferring attacks from the model at the row. Diagonals correspond to white-box attacks.
the diagonal. This indicates strong transferability, where attacks crafted on one model substantially reduce accuracy on others. DiffPrivate exhibits relatively weak transferability—off-diagonals are often 20–60% higher than the diagonal, meaning attacks do not carry over well across architectures. For instance, on LFW, evaluating on SwinT yields 49% robust accuracy in the white-box setting compared to 90–97% robust accuracy when transferring the attack from the convolutional networks to SwinT, suggesting that DiffPrivate is architecture-specific. To quantify this gap, we define a transferability score T as the average ratio of black-box to white-box attack success rate across model pairs (T =1 indicates perfect transfer; see App. G for the formal definition). Computed from Fig. 2, B OUND S TYLE achieves T =0.76 (LFW) and T =0.86 (VGG-Face), while DiffPrivate achieves T =0.43 (LFW) and T =0.68 (VGG-Face). We further show in App. G that this gap is not an artifact of a specific parameter choice causing DiffPrivate to overfit to the surrogate model: reducing DiffPrivate’s iteration count does not improve its transferability, confirming that the gap reflects inherent differences between the two attacks.
%
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
11
Attacks’ Run Times We benchmark attacks’ run times on an NVIDIA RTX A6000 GPU, executing each attack 100 times and averaging the run time. For fair comparison, we use an equal batch size of 1 for both attacks. Under this setting, B OUND S TYLE takes an average of 8,302 ms to complete per image compared to an average of 79,094 ms attained by DiffPrivate, demonstrating approximately ×9.5 speedup. This result highlights B OUND S TYLE’s better fit for adversarial training compared to other state-of-the-art attacks: B OUND S TYLE is not only highly effective in white-box settings and transferable in black-box settings, but is also significantly faster than DiffPrivate. To further demonstrate that B OUND S TYLE attains superior time and success-rate tradeoffs compared to DiffPrivate, we execute both attacks under a matched wall-clock constraint on the same hardware. Specifically, we run the attacks on an NVIDIA A6000 GPU with a limit of 6 seconds, which corresponds to approximately 30 iterations of B OUND S TYLE. For DiffPrivate, we provide a significant advantage by removing caps on ∥∆z∥ and the number of iterations, constraining it solely by the run time. Tab. 3 lists the results. It can be seen that, under equal execution time, DiffPrivate leaves robust accuracy at 98.6%, whereas B OUND S TYLE degrades it to 85.2–50.9% (depending on β ) against ResNet, on the LFW dataset. This result further highlights B OUND S TYLE’s efficiency and its ability to attain high success rates within strict time constraints, making it suitable for adversarial training. Attack
Budget
Time
Rob. Acc.
DiffPrivate B OUND S TYLE B OUND S TYLE B OUND S TYLE
∥∆z∥=∞ β =1.5 β =2.0 β =3.0
6s 6s 6s 6s
98.6% 85.2% 75.0% 50.9%
Table 3: Comparing attacks’ success, on LFW and ResNet, under matched time constraints (6 seconds). B OUND S TYLE’s Success Stems From Semantic Edits By construction, B OUND S TYLE cannot inject unconstrained pixel noise, as its output is the StyleGAN3 decoding of an ℓ2 bounded latent edit. Four results indicate that its success is governed by semantic changes: (1) >90% of its perturbation energy lies below frequency radius r≈15 in Fourier domain vs. r≈30 for DiffPrivate (App. D); (2) noise-removal filters barely affect B OUND S TYLE (≤6% robust-accuracy increase) yet weaken DiffPrivate by up to 36.6% (see §6.2); (3) adversarially training with pixel-space perturbations shifts robustness against B OUND S TYLE by only ±≈1%, but against DiffPrivate by up to +9.3% (App. C); and (4) B OUND S TYLE’s principal attack directions decode to semantic factors such as aging, pose, and illumination (App. F).
6.2
S TYLE AT Improves Adversarial Robustness
White-box Setting We execute white-box attacks against S TYLE AT and DOA as well as against the undefended models (ResNet and RepVGG) at different attack budgets. Tab. 4 reports the results. Compared to the undefended models, S TYLE AT shows a 0.4–20.9% increase in robust accuracy on the different attack budgets against both B OUND S TYLE and DiffPrivate, with the latter unseen during training. The DOA defense, tailored for ad hoc semantic attacks, shows a 0.4–0.7% higher robust accuracy than S TYLE AT on a few attack budgets on DiffPrivate, but otherwise trails behind S TYLE AT’s robust accuracy by 0.4–9.6%
12
SHAPIRA ET AL.: DEFENDING FACE RECOGNITION AGAINST SEMANTIC ATTACKS
LFW DiffPrivate - ∥∆z∥
B O U N D S T Y L E - budget β Backbone Defense
Clean
1
1.5
2
3
1
2
3
4
5
6
ResNet
No Defense DOA S TYLE AT
99.1 99.1 99.5
93.5 83.8 70.8 94.9 80.6 69.4 97.7 91.7 84.3
39.5 35.4 50.7
99.1 98.6 93.5 71.8 41.7 33.8 99.1 98.6 96.8 85.6 54.6 42.6 99.5 99.5 97.7 86.6 54.2 47.2
RepVGG
No Defense DOA S TYLE AT
99.1 99.5 99.5
92.6 83.3 69.4 93.1 84.3 72.7 97.2 92.1 84.7
40.5 34.5 54.3
99.1 99.1 96.3 72.7 43.5 38.0 99.5 99.1 97.7 89.4 59.7 43.1 99.5 99.5 98.6 92.6 64.4 50.0
VGG-Face DiffPrivate - ∥∆z∥
B O U N D S T Y L E - budget β Backbone Defense
Clean
1
1.5
2
3
1
2
3
4
5
6
ResNet
No Defense 100.0 84.0 72.4 59.6 DOA 99.4 82.7 69.9 55.8 S TYLE AT 100.0 88.5 81.4 65.4
34.6 28.9 35.3
77.4 67.7 51.0 30.3 22.6 17.4 79.4 74.2 58.7 38.1 27.1 16.1 81.3 76.1 63.2 41.9 27.7 18.1
RepVGG
No Defense 100.0 80.8 71.2 60.3 DOA 98.7 79.5 73.7 57.0 S TYLE AT 99.4 89.1 80.8 69.9
32.0 27.6 35.7
77.4 69.7 53.5 32.3 22.6 17.4 81.3 74.2 57.4 40.0 26.5 21.3 80.6 76.1 62.6 43.2 31.0 20.6
Table 4: Evaluating defenses against white-box attacks. For each defense and undefended model, we report benign accuracy and robust accuracy under varied attack budgets on the LFW and VGG-Face datasets, for both ResNet and RepVGG backbones.
against the DiffPrivate attack. However, against B OUND S TYLE, DOA is counterproductive for most attack budgets, decreasing robust accuracy compared to the undefended model by up to 6%. Altogether, these results show that S TYLE AT reliably improves robust accuracy against general semantic attacks while generalizing to attacks unseen during training. Importantly, S TYLE AT also maintains the benign accuracy on clean data as the undefended model or even roughly improves it. Gray-box Setting We also evaluate S TYLE AT and filter-based defenses against gray-box attacks, where the adversary produces attacks against the undefended models (ResNet and RepVGG), in a manner agnostic to the defense. Fig. 3 reports the robust accuracy achieved against B OUND S TYLE with the ResNet backbone (budgets β ∈ {1, 2, 3} and clean; β =1.5 omitted for compactness, see App. E). Fig. 4 reports DiffPrivate results at norms ∥∆z∥=4–6, where inter-defense differences are most pronounced; the full norm range appears in App. E. Against B OUND S TYLE, it can be seen that filter-based defenses have little impact on robustness, increasing robust accuracy by 6.0% in the best case compared to the undefended model. DiffPure is in fact counterproductive against B OUND S TYLE, reducing robust accuracy by up to 6.0% below the undefended baseline across all budgets and datasets. In comparison, S TYLE AT results in up to 33.3% increase in robust accuracy. The filter-based defenses are more useful against DiffPrivate, increasing robust accuracy by up to 31.5%. By contrast, DiffPure substantially improves robustness against DiffPrivate: it outperforms all filter-based defenses at moderate-to-high perturbation norms, improving robust accuracy by up to 36.6% over the undefended model. This asymmetry suggests that DiffPure’s purification is effective when adversarial structure aligns with latent diffusion dynamics, as in Diff-