ConceptioArchivearXiv CS
arXiv CSopen access

On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

On Adversarial Vulnerability of Vision-Language Models through the Lens of Intermediate Spectral Subspaces Chethan Krishnamurthy Ramanaik1 , Tobias Callies1 , Michael Hecht2 , and Eirini Ntoutsi1 University of the Bundeswehr Munich, Germany {chethan.krishnamurthy,tobias.callies,eirini.ntoutsi}@unibw.de 2 University of Wrocław, Poland [email protected]

arXiv:2607.07375v1 [cs.LG] 8 Jul 2026

1

Abstract. Adversarial vulnerability in deep neural networks (DNNs) has been studied from the perspectives of decision-boundary geometry, feature robustness, input–output Jacobians, and the instability of inverse problems. Here, we focus on the spectral structure of intermediate linear transformations that propagate information through modern DNNs, an unexplored mechanism of adversarial vulnerability. Specifically, we investigate transformer-based vision–language models, whose linear layers admit interpretable spectral decompositions and whose widespread adoption makes understanding their robustness increasingly important. We propose a white-box spectral-subspace-guided attack (SSGRA) that aligns intermediate representations with the subspace spanned by the bottom right singular vectors. Our experiments show improved attack effectiveness over existing baselines. In addition, SSGRA offers a spectral interpretation of adversarial vulnerability in VLMs, providing insights for improving their robustness. Keywords: Adversarial attacks · Spectral analysis · VLMs

1

Introduction

Deep neural networks, including modern vision–language models (VLMs), are known to be vulnerable to adversarial perturbations that substantially alter model predictions while remaining visually imperceptible [25, 57, 62]. Despite extensive research, the mechanisms governing how such perturbations propagate through deep networks remain incompletely understood. Existing explanations primarily analyze adversarial vulnerability from input-space [16, 22, 24, 38, 55, 56, 64] or end-to-end perspectives, including decision-boundary geometry [49], robust and non-robust features [33], Jacobian analysis [31, 36, 39, 51], inverse problem instability and Lipschitz properties [2, 3, 26]. Despite these advances, existing theories predominantly explain adversarial vulnerability from the input space or through end-to-end network properties, leaving the spectral behavior of intermediate linear transformations largely unexplored.

2

C.K. Ramanaik et al.

Fig. 1: Overview of the proposed spectral framework.

Transformer-based VLMs provide a natural setting for such an analysis. Their architectures comprise numerous learnable intermediate linear transformations including the projection matrices in self-attention, feed-forward networks, and multimodal fusion modules [4, 14, 21, 58], making spectral decomposition a principled tool for studying representations. Because singular values govern how different representation directions (singular vectors) are amplified or suppressed by each linear transformation, they naturally provide a lens for studying how adversarial signals propagate through transformer layers. Moreover, their widespread adoption makes them an important testbed for adversarial robustness. This motivates us to investigate adversarial vulnerability through the singularvector basis of intermediate linear transformations. Inspired by the instability of ill-posed inverse problems, where near-null singular directions govern information loss, we study how adversarial intermediate representations align with top and bottom singular-vector subspaces during attack optimization. Guided by this perspective, we formulate a spectral-guidance principle and instantiate it through a white-box attack that serves to empirically validate the proposed mechanism. Our results suggest that, beyond constraining large singular values, explicitly controlling near-null singular directions may offer a complementary strategy for improving adversarial robustness. Contributions. We identify bottom singular-vector subspaces of intermediate linear transformations as a previously overlooked spectral attack surface in transformer-based VLMs. We show that untargeted adversarial optimization naturally tends to increase the alignment of intermediate representations with these information-attenuating subspaces, even without explicit enforcement. Building on this insight, we propose the Spectral Subspace Guided Representation Attack (SSGRA), a spectrally guided white-box attack that demonstrates improved attack effectiveness compared with representative feature-space and output-space attacks on three state-of-the-art VLMs. In the subsequent sections, the paper reviews the related work, followed by the preliminaries, methodology, and experimental evaluation.

Adversarial Vulnerability in VLMs via Spectral Subspaces

2

3

Related Work

We briefly review the theoretical foundations of adversarial vulnerability, followed by relevant representative methods for crafting adversarial examples. 2.1

Theoretical Perspectives.

Manifolds & Decision Boundaries. Adversarial vulnerability has been attributed to the geometry of high-dimensional decision boundaries [16,38]. Adversarial examples have also been explained as off-manifold inputs [22,24,55,56,64], while universal perturbations suggest that decision boundaries around different inputs share a low-dimensional subspace of normal vectors [49]. Non-robust features. Standard training encourages models to rely on highly predictive but non-robust features, whose sensitivity to small input perturbations leads to adversarial vulnerability [33]. Linearity approximation. Adversarial vulnerability has been attributed to local linearity, where high-dimensional gradient–perturbation interactions amplify small perturbations [25]. In linear settings, vulnerability arises when the input lies close to the decision boundary [15,30]. However, experimental evidence suggests that DNNs are only locally linear and remain globally nonlinear [45]. Internal weights. Large singular values of weight matrices have been linked to adversarial vulnerability through their connection to local Lipschitz constants, motivating spectral regularization [10,57,63]. Near-zero singular values suppress gradient flow, while restoring these gradients strengthens adversarial attacks [28, 52]. End-to-end Jacobians. Adversarial vulnerability has been linked to large input–output Jacobians [31,36]. Targeted perturbations can be expressed as linear combinations of the right singular vectors of the logit-to-image Jacobian [51], while intermediate layer Jacobians identify sensitive input directions [39]. Noisy and poorly aligned input Jacobians have also been associated with adversarial vulnerability [9]. Instability of inverse problems. Studies on the instability of inverse problems show that information lost along null-space and near-null singular directions leads to unstable reconstruction during inversion [2, 3, 26, 53]. However, most adversarial robustness research has focused on constraining large singular values through spectral normalization and Lipschitz regularization [6, 10, 17, 27, 48, 59, 63], as well as on the implicit self-regularization of dominant singular values [47]. Comparatively little attention has been paid to whether bottom singularvector subspaces of intermediate linear transformations constitute a source of adversarial vulnerability, which is the focus of this work. 2.2

Methods for Generating Adversarial Examples

Adversarial attacks have evolved from early gradient-based methods such as FGSM, PGD, and optimization-based attacks [7, 25, 46, 57] to feature-alignment attacks [34, 54] and attacks on generative models [8, 23, 40, 52, 61]. Recent VLM

4

C.K. Ramanaik et al.

attacks have explored computational availability [20], cross-prompt transferability [44], CoT reasoning [60], visual grounding [19], gray-box SVD-based attacks [42], black-box attacks [13], and behavior hijacking [5] targeting different threat models, tasks, or objectives including targeted attacks, visual reasoning, gray-box or black-box settings, and behavior control. We focus on untargeted white-box attacks to study adversarial vulnerability through information attenuation rather than predefined attack objectives. Accordingly, we compare against representative untargeted white-box featureand output-level attacks. Feature Discrepancy Attack (FDA) perturbs inputs by maximizing discrepancies between intermediate representations [18]. Similarly, Self-Supervised Perturbation Attack (SSPA) maximizes the discrepancy between clean and adversarial feature representations in pretrained models [37, 50]. Dispersion Reduction Attack (DRA) minimizes the variance of intermediate features, forcing representations to collapse and become less discriminative [43]. Blockwise Similarity Attack (BSA) targets transformer-based VLMs by maximizing cosine discrepancies between block-wise intermediate representations of clean and adversarial inputs, thereby disrupting semantic alignment throughout the network [62]. Beyond feature-space objectives, output-level attacks maximize the task loss using cross-entropy (CE) or negative log-likelihood [11]. Entropy-Guided Attacks (EGA) maximize output entropy to induce uncertain model responses [29]. We compare against representative feature-space attacks (FDA, SSPA, DRA, and BSA) and output-level attacks (CE, EGA).

3

Preliminaries

Notation. Let x\in \mathbb {R}^{d_x} denote a flattenedinput image. We consider perturbations constrained to the L_p -ball Bcp (x) = xa ∈ Rdx ∥xa − x∥p ≤ c , where c is the perturbation budget. A VLM can be abstractly described by a function F : Rdx × P → ζ, where \protect \mathcal {P} is the space of input prompts and \protect \mathcal {\zeta } is the associated tokenizer dictionary. VLM pipeline Modern VLMs typically consist of a visual encoder, a multimodal projection or fusion module, and a large language model (LLM). Visual Encoder. The visual encoder ϕ : Rdx → RNv ×dv consists of K sequentially applied blocks ϕk , each producing a visual token representation hvk ∈ RNv ×dv consisting of N_v visual tokens of the visual embedding dimensionality dv for a given input image x. Fusion Module. The final visual embedding hvK (x) = ϕ(x) is projected into the embedding space of the language model through a projection module P , outputting H0v = P (ϕ(x)) ∈ RNv ×d , where d is the hidden dimension of the language model. For a given textual prompt, let H0t ∈ RNt ×d denote the corresponding textual embedding, using N_t tokens. The concatenated visual

Adversarial Vulnerability in VLMs via Spectral Subspaces

5

and textual embeddings then form the input sequence to the language model,   H0 = H0v ; H0t ∈ RN ×d , where N = Nv + Nt is the total sequence length after multimodal fusion. LLM. The LLM consists of L transformer blocks and a final decoder. The \ell -th transformer block outputs the intermediate tokenized representation Hℓ ∈ RN ×d . Singular Subspaces and Orthogonal Projection For a linear transformation W ∈ Rm×n , the singular value decomposition (SVD) is given by W = U ΣV ⊤ , where U\in \mathbb {R}^{m\times m} and V\in \mathbb {R}^{n\times n} are orthonormal matrices, and Σ ∈ Rm×n contains the singular values of W , ordered as σ1 ≥ σ2 ≥ · · · ≥ σr ≥ 0, on its diagonal entries: Σii = σi for ,1\leq i \leq r=\mathrm {rank}(W)\leq \min \{m,n\} and zeros otherwise. The columns of V = [v1 , v2 , . . . , vn ] are referred to as the right singular vectors of W . For a given 1\le k\le n, let V_k^{\mathrm {top}} = \{v_1,\ldots ,v_k\}, \qquad V_k^{\mathrm {bottom}} = \{v_{n-k+1},\ldots ,v_n\}, denote the top-k and bottom-k right singular vectors of W , respectively, and the corresponding subspaces are referred to as top-k and bottom-k singular subspaces, and denoted by \protect \mathrm {span}(V_k^{\mathrm {top}}) and \protect \mathrm {span}(V_k^{\mathrm {bottom}}). Due to the orthonormality of V , we can measure the alignment of a vector z ∈ Rn with such subspaces using the norm of their projections and the identity bottom ∥z∥22 = ∥Pktop (z)∥22 + ∥Pn−k (z)∥22 . Here, Pktop (z) = (v1⊤ z, . . . , vk⊤ z) denotes the top projection onto span(Vk ), and Pkbottom is defined analogously. The projection norms motivate the interpretation as corresponding energies. Effect of Near-Null Singular Directions: An Analogy to Ill-Posed Inverse Problems Consider the linear transformation W : Rn → Rm with reconstruction map Γ : Rm → Rn . The instability of ill-posed inverse problems states that, in general, stable reconstruction cannot be guaranteed [3, 32, 57], with deeper theoretical treatments in [3, 26]. A common characterization of this instability is the local Lipschitz constant of the reconstruction map at a measurement y ∈ Rm : L_\varepsilon (\Gamma ,y) = \sup _{0<\|y'-y\|<\varepsilon } \frac {\|\Gamma (y')-\Gamma (y)\|_2} {\|y'-y\|_2}, \qquad \varepsilon >0, where y ′ ∈ Rm denotes a perturbed measurement. The local Lipschitz constant may become unbounded, causing large reconstruction errors. We draw an analogy between this instability phenomenon and intermediate linear transformations in VLMs, extending the inverse-problem viewpoint of [35]. In inverse problems, the reconstruction map attempts to recover information attenuated by the forward operator, whereas VLMs propagate representations through successive transformations. Consequently, representations aligned with the bottom singular-vector subspace of W are strongly attenuated by the forward transformation, motivating our study of bottom singular-vector subspace alignment.

6

C.K. Ramanaik et al.

Spectral Alignment Measure For a set of vectors H = {h1 , . . . , hN } ⊂ Rn (as arising in the tokenized representations with n = d), we quantify the average alignment with a subspace by computing the average projection energy of the normalized token representations. For span(Vktop ) this takes the form: \label {alignmentMeasure} \Psi _k(H,V^{\mathrm {top}}_k) = \Psi _k(H,\mathrm {span}(V^{\mathrm {top}}_k)) = \frac {1}{N} \sum _{i=1}^{N} \sum _{j=1}^{k} \left ( \frac {v_j^\top h_i} {\|h_i\|_2} \right )^2.

(1)

Consequently, 0 ≤ Ψk (H, Vktop ) ≤ 1, and larger values indicate greater concentration of representational energy within the selected singular subspace. The alignment measure is defined analogously for the bottom-k singular subspace Vkbottom . In the proposed attack, we instantiate this general definition using a fixed subspace dimension s, i.e., Ψs (·, Vsbottom ). Threat Model We consider an untargeted white-box attack where the adversary has full access to the model architecture, parameters, and intermediate representations. Given an image x and a text prompt, the adversary generates an adversarial image xa that degrades the model’s response while remaining visually similar to x. The perturbation is constrained by an L∞ budget c, i.e., xa ∈ Bc∞ (x). The prompt, model parameters, and inference procedure remain unchanged, and each image is attacked independently.

4

A Spectral Framework for Adversarial Vulnerability

We first introduce the Spectral Subspace Guided Representation Attack (SSGRA), which instantiates the proposed spectral-guidance principle, and then present the layer-wise probing framework used to analyze the spectral dynamics of adversarial optimization. 4.1

Spectral Subspace Guided Representation Attack (SSGRA)

SSGRA extends the Blockwise Similarity Attack (BSA) [62] by introducing a spectral guidance term based on the alignment measure in Eq. (1). Motivated by the instability phenomenon (Section 3), this guidance aligns intermediate adversarial representations with the bottom singular-vector subspaces of selected linear transformations, where information is most attenuated. We hypothesize that steering representations toward these subspaces progressively weakens semantic information propagation, improving attack effectiveness. SSGRA combines two complementary objectives (Eq 2). The first maximizes the discrepancy between clean and adversarial intermediate representations following BSA [62], thereby disrupting the learned feature hierarchy. The second

Adversarial Vulnerability in VLMs via Spectral Subspaces

7

maximizes the spectral alignment measure defined in Eq. (1), encouraging adversarial representations to concentrate their energy within bottom singular-vector subspaces. For each selected layer m ∈ S, let zm (·) denote the collection of token representations immediately preceding the corresponding linear transformation, and bottom let Vm,s denote the subspace spanned by the bottom-s right singular vectors of that transformation. Since the textual prompt remains fixed during optimization, we suppress its dependence in the notation and write hvk (x), Hℓ (x), and zm (x) instead of hvk (x, p), Hℓ (x, p), and zm (x, p). Definition 1 (Spectral Subspace Guided Representation Attack (SSGRA)). The SSGRA adversarial example is defined as the solution to the following optimization problem: \label {eq:ssgra} \begin {aligned} x_a^* = \arg \max _{x_a\in B_c^p(x)} \Bigg \{ & -\lambda \Bigg [ \sum _{k=1}^{K} \sum _{j=1}^{N_v} \cos \!\left ( h_k^{v,(j)}(x), h_k^{v,(j)}(x_a) \right ) \\ & + \sum _{\ell =1}^{L} \sum _{i=1}^{N} \cos \!\left ( H_\ell ^{(i)}(x), H_\ell ^{(i)}(x_a) \right ) \Bigg ] \\ & + (1-\lambda ) \sum _{m\in \mathcal S} \Psi _s \!\left ( z_m(x_a), V^{\mathrm {bottom}}_{m,s} \right ) \Bigg \}. \end {aligned}

(2)

v,(j)

(i)

where hk and Hℓ denote the representations of the j-th visual token and the i-th multimodal token at the outputs of the k-th visual encoder block and the ℓ-th LLM block, respectively. Furthermore, Ψs (·, ·) is the spectral alignment measure defined in Eq. (1), s denotes the chosen dimension of the bottom singular subspace, and λ ∈ [0, 1] controls the trade-off between representation discrepancy and spectral subspace alignment. The optimization procedure is summarized in Algorithm 1 in the Appendix. Rather than applying spectral guidance to all layers, we use it only on a selected subset of intermediate linear transformations, denoted by S. The layers are selected by evaluating each visual encoder, fusion, and LLM layer independently on a small validation set and retaining those that yield the strongest attack performance. Developing adaptive layer-selection methods that avoid validation-based tuning is left for future work. 4.2

Layer-wise Probing of Spectral Alignment

To analyze the spectral alignment of adversarial representations, we perform layer-wise adversarial probing. For each transformer block i, we generate an adversarial example by minimizing the cosine similarity between the clean and adversarial feature maps: \label {eq:layerwise_probe} x_{a,i}^{*} = \arg \min _{x_a \in B_{c}(x)} \frac { \left \langle H_i(x), H_i(x_a) \right \rangle _F }{ \|H_i(x)\|_F \, \|H_i(x_a)\|_F },

(3)

8

C.K. Ramanaik et al.

where H_i(\cdot ) denotes the feature map at layer i, \delimiter "426830A \cdot ,\cdot \rangle _F is the Frobenius inner product, and B_{c}(x) is the admissible perturbation set. For each adversarial example x_{a,i}^{*} , we compute the spectral subspace alignment across all vision and language layers using Eq. (1). Repeating this procedure over all target layers and input samples enables us to analyze alignment with top and bottom singular-vector subspaces during and after attack optimization. 4.3

Evaluation Metrics

We evaluate attack effectiveness by comparing the adversarial output ŷa with the corresponding clean output ŷc using BERTScore [65] and ROUGE-L [41]. BERTScore measures semantic similarity using contextual token embeddings from RoBERTa-large. ROUGE-L measures lexical similarity based on the longest common subsequence (LCS), capturing structural degradation of the generated response. For both metrics, we report Precision, Recall, and F1, where lower scores indicate stronger attacks. Additional details are provided in Appendix A.2.

5

Experiments

We evaluate the proposed attacks against representative baselines, analyze their spectral mechanisms, and present ablation studies. 5.1

Experimental Setup

We evaluate attacks on different VLMs, namely Gemma-3 (4B) [21], Qwen2.5-VL (7B) [4], and LLaVA-1.5 (7B) [1]. Experiments are conducted on ImageNet [12], whose diverse visual categories enable assessment of generalization across image content. Given an input image and the prompt “What is shown in the image?”, we optimize sample-specific adversarial perturbations to degrade the model’s image description while remaining visually imperceptible. Each experimental instance is defined by a perturbation budget and an attack method, evaluated over 100 images. The perturbation budget ranges from 0.002 to 0.005 in the \ell _{\infty } norm, selected via grid search such that the lower bound captures the regime where outputs remain semantically similar across methods, and the upper bound where performance differences become pronounced. We compare SSGRA against six representative baselines BSA, DRA, FDA, EGA, SSPA, and NLL. All attacks are optimized using Adam following [7] with a fixed computational budget of 1000 gradient steps to ensure a fair comparison across methods and enable evaluation on a sufficiently large number of samples for statistically reliable quantitative results. Grid search over learning rates {10−2 , 5 × 10−3 , 10−3 , 5 × 10−4 , 10−4 } identified 10−3 as consistently yielding the best attack performance across the three models. The adversarial perturbation is initialized with small random noise of near-zero magnitude, following standard practice in iterative adversarial optimization.

Adversarial Vulnerability in VLMs via Spectral Subspaces

(a) Qwen 2.5-VL BERT F1

(b) LLaVA 1.5 BERT F1

(c) Gemma 3 BERT F1

(d) Qwen 2.5-VL ROUGE-L F1

(e) LLaVA 1.5 ROUGE-L F1

(f ) Gemma 3 ROUGE-L F1

9

Fig. 2: SSGRA vs representative baselines across perturbation budgets. Lower scores indicate stronger attacks.

5.2

Quantitative comparison with State-of-the-Art Attacks

Figure 2 shows the performance of SSGRA and the selected baselines across perturbation budgets using BERTScore F1 and ROUGE-L F1. Qwen 2.5-VL: SSGRA consistently outperforms all baselines, achieving an additional 7.90–19.74% relative degradation under BERTScore F1 and 30.50– 97.93% under ROUGE-L F1 over the strongest baseline. LLaVA 1.5: SSGRA achieves up to 7.46% and 18.90% additional relative degradation over the strongest baseline under BERTScore F1 and ROUGE-L F1, respectively. The improvement increases with the perturbation budget, indicating that spectral subspace guidance becomes more effective at larger perturbation budgets. Gemma 3: Although the margins are smaller, SSGRA consistently achieves the strongest attacks, providing up to 0.80% and 3.07% additional relative degradation over the strongest baseline under BERTScore F1 and ROUGE-L F1, respectively. Among the three VLMs, Qwen 2.5-VL exhibits the largest degradation, consistent with its having the highest proportion of near-null singular directions (Table 1), and thus the largest spectral attack surface. In contrast, Gemma 3

10

C.K. Ramanaik et al.

(a) Qwen2.5-VL

(b) LLaVA-1.5

(c) Gemma 3

Fig. 3: Qualitative adversarial examples generated under a common perturbation budget of c = 0.003 across three vision–language models.

shows the smallest relative degradation, consistent with its lower proportion of near-null singular directions compared with Qwen 2.5-VL and LLaVA 1.5, limiting the opportunity to exploit bottom singular-vector subspaces. A detailed discussion on spectral characterization follows in Section 5.4. Additional results are provided in the Appendix. In particular, Table 3 summarizes the F1 scores and the additional relative degradation (%) over the strongest baseline attack. The corresponding Precision, Recall, and F1 scores are provided in Tables 4–6, while the corresponding trends are shown in Figures 8 and 9 in the Appendix. Overall, the results suggest a relationship between spectral conditioning and adversarial vulnerability and provide empirical support for our hypothesis that bottom singular-vector subspaces play a central role in transformer-based VLMs. 5.3

Qualitative Analysis

Figure 3 shows representative adversarial examples generated under a common extremely small perturbation budget of c = 0.003 for the chosen models. Com-

Adversarial Vulnerability in VLMs via Spectral Subspaces

(a) Qwen2.5-VL

(b) LLaVA-1.5

11

(c) Gemma 3

Fig. 4: Distribution of the largest (\sigma _{\max } , left) and smallest (\sigma _{\min } , right) singular values across non-attention linear operators in Qwen2.5-VL, LLaVA-1.5, and Gemma 3. Table 1: Distribution of extreme singular values across non-attention linear operators. Threshold

Qwen2.5-VL Count

Percentage

LLaVA-1.5 Count

Percentage

Gemma 3 Count

Percentage

25/221 2/221

11.31% 0.90%

32/221 27/221 6/221 3/221 0/221

14.48% 12.22% 2.71% 1.36% 0.00%

Largest Singular Value (\sigma _{\max } ) \sigma _{\max } > 10^{1} \sigma _{\max } > 10^{2}

46/245 0/245

18.78% 0.00%

40/206 0/206

19.42% 0.00%

Smallest Singular Value (\sigma _{\min } ) \sigma _{\min } < 10^{-2} \sigma _{\min } < 10^{-3} \sigma _{\min } < 10^{-4} \sigma _{\min } < 10^{-5} \sigma _{\min } < 10^{-6}

62/245 61/245 24/245 3/245 0/245

25.31% 24.90% 9.80% 1.22% 0.00%

60/206 57/206 26/206 4/206 1/206

29.13% 27.67% 12.62% 1.94% 0.49%

pared with existing baselines, SSGRA consistently induces larger semantic deviations while preserving the perceptual appearance of the input images. For the coastal scene (Qwen2.5-VL), baseline attacks largely preserve the correct scene description, whereas SSGRA causes the model response to become meaningless. Likewise, for the ambulance image (LLaVA-1.5), baseline attacks still identify the ambulance despite minor hallucinations, whereas SSGRA instead describes an unrelated car crash scene. For the flower image (Gemma 3), baseline attacks largely preserve the correct flower category despite minor hallucinations involving a cup and a lizard, whereas SSGRA generates an unrelated prediction ("Green Slime/Fizz Pop Rocks"). More qualitative examples are presented in Appendix A.4. Overall, SSGRA induces larger semantic shifts than existing baselines, consistent with the quantitative findings in Section 5.2. 5.4

Spectral Characterization of Adversarial Representations

Singular Value Analysis of intermediate linear operators of VLMs. We analyze the distributions of the largest and smallest singular values of all nonattention linear operators in the evaluated models. Figure 4 visualizes these distributions, while Table 1 summarizes the prevalence of extreme singular values.

12

C.K. Ramanaik et al.

(a) BSA, Top k = 10

(b) BSA, Bottom k = 10

Fig. 5: Distribution of the spectral alignment measure Ψk before and after BSA optimization for the top-10 and bottom-10 right singular-vector subspaces.

Across all three models, near-null singular directions occur substantially more frequently than strongly amplifying ones. For example, σmin < 10−3 is observed in 24.90%, 27.67%, and 12.22% of operators in Qwen2.5-VL, LLaVA-1.5, and Gemma 3, respectively, with similar trends persisting at smaller thresholds. In contrast, singular values satisfying σmax > 10 occur much less frequently, accounting for only 18.78%, 19.42%, and 11.31% of operators. Singular values exceeding 102 are nearly absent across all models (Table 1), consistent with prior work showing that large singular values are commonly constrained or implicitly regularized during training. These spectral characteristics help explain the quantitative results in Section 5.2. Qwen2.5-VL and LLaVA-1.5 contain approximately twice as many near-null singular directions as Gemma 3, providing substantially larger bottom singular-vector subspaces that SSGRA can exploit. Consequently, these models exhibit much larger relative degradation over existing attacks than Gemma 3. Post-Attack Spectral Alignment Analysis. To examine whether spectral alignment emerges naturally during adversarial optimization, we analyze adversarial representations generated by BSA [62], which does not optimize spectral objectives. We evaluate 100 images (c = 0.005) and compute Ψk using Eq. (1) with respect to the top-10 and bottom-10 right singular vectors of selected intermediate layers. Figure 5 shows the corresponding alignment distributions before and after optimization. At the MLP gate proj and MLP up proj layers, adversarial representations exhibit increased alignment with the bottom-k singular subspace, whereas alignment with the top-k subspace remains unchanged or decreases. This trend is absent in attention layers, likely due to their more complex transformations. Since BSA does not optimize spectral alignment, the results suggest that untargeted adversarial optimization naturally steers representations toward informationattenuating bottom singular subspaces. SSGRA strengthens the attack by explicitly promoting this alignment. Spectral Alignment Dynamics During Attack Optimization. We analyze spectral alignment during adversarial optimization on Gemma 3. For each target layer, adversarial examples are generated using Eq. (3), and alignment with the largest and smallest singular vectors is computed using Eq. (1). Figure 6 shows

Adversarial Vulnerability in VLMs via Spectral Subspaces

(a) Ψtop (v) for sample 1

(b) Ψtop (v) for sample 2

(c) Ψbottom (v) for sample 1

13

(d) Ψbottom (v) for sample 2

Fig. 6: Layer-wise evolution of top and bottom most singular vectors alignment during adversarial optimization on Gemma 3 (two representative samples). Table 2: Computational complexity of the attacks measured in floating-point operations (FLOPs). Lower values indicate higher computational efficiency. Method

Qwen2.5-VL

LLaVA-1.5

Gemma 3

BSA DRA FDA SSPA EGA CE SSGRA

2.94 × 1013 1.42 × 1013 6.14 × 1012 6.14 × 1012 2.68 × 1013 2.69 × 1013 1.31 × 1013

1.02 × 1013 8.68 × 1012 1.14 × 1012 1.14 × 1012 2.69 × 1013 2.69 × 1013 2.61 × 1013

2.59 × 1013 1.01 × 1013 1.01 × 1013 1.56 × 1013 1.95 × 1013 1.96 × 1013 2.60 × 1013

trajectories from the first multimodal block. Adversarial optimization progressively increases alignment with the bottom singular vectors, while alignment with the top singular vectors remains nearly unchanged or slightly decreases, supporting our hypothesis. 5.5

Computational Complexity

Table 2 reports the computational complexity of the evaluated attacks in FLOPs. Across all three VLMs, SSGRA has computational cost comparable to representative optimization-based attacks (BSA, EGA, and CE) while achieving improved attack effectiveness. FLOPs should be compared only within the same VLM, as they depend not only on parameter count but also on the vision encoder, attention mechanism, input image resolution, and the number of visual tokens processed by the multimodal model. For Qwen2.5-VL, SSGRA requires fewer FLOPs than all other methods while achieving the strongest attack performance (Section 5.2), since hyperparameter tuning selected λ = 0, reducing Eq. (2) to just spectral alignment objective. 5.6

Ablation Study

We ablate SSGRA on Qwen2.5-VL to evaluate (1) the contribution of the spectral alignment objective and (2) the effect of aligning adversarial representations with bottom versus top singular-vector subspaces.

14

C.K. Ramanaik et al.

(a) Effect of the spectral alignment loss.

(b) Bottom- vs. top-singular subspace alignment.

Fig. 7: Ablation analysis validating the design choices of SSGRA.

Effect of the Spectral Alignment Term. Figure 7a compares SSGRA with and without the spectral alignment term in Eq. (2). Removing this term increases both BERTScore F1 and ROUGE-L F1 across all perturbation budgets, indicating weaker attacks and confirming that spectral alignment is the primary contributor to the improved attack effectiveness. The performance gap widens with increasing c, showing that the benefit of spectral guidance increases with the perturbation budget. Bottom vs. Top Singular Subspace. Figure 7b compares SSGRA-bottom and SSGRA-top, which align adversarial representations with the bottom-k and top-k singular subspaces, respectively. SSGRA-bottom achieves lower BERTScore and ROUGE-L scores, indicating stronger attacks, whereas SSGRA-top causes only minor degradation. These results support the hypothesis that bottom singularvector subspaces constitute the primary spectral attack surface in VLMs.

6

Conclusion

We presented a spectral perspective on adversarial vulnerability in transformerbased VLMs by analyzing their intermediate linear transformations. We identified bottom singular-vector subspaces as a previously overlooked spectral attack surface and proposed SSGRA, which exploits this insight to improve attack effectiveness on three state-of-the-art VLMs. Our analyses show that near-null singular directions are substantially more prevalent than strongly amplifying ones and that untargeted adversarial optimization naturally tends to increase alignment with these information-attenuating subspaces. These findings suggest that, alongside existing spectral-norm regularization techniques for large singular values, controlling near-null singular directions may provide a complementary approach to improving adversarial robustness. Limitations. This work focuses on untargeted white-box attacks to isolate and analyze the spectral mechanisms underlying adversarial vulnerability. The applicability of the proposed framework to transfer-based and black-box settings has not yet been investigated and remains future work.

Adversarial Vulnerability in VLMs via Spectral Subspaces

15

References 1. An, X., Xie, Y., Yang, K., Zhang, W., Zhao, X., Cheng, Z., Wang, Y., Xu, S., Chen, C., Zhu, D., et al.: Llava-onevision-1.5: Fully open framework for democratized multimodal training. arXiv preprint arXiv:2509.23661 (2025) 8 2. Antun, V., Gottschling, N.M., Hansen, A.C., Adcock, B.: Deep learning in scientific computing: Understanding the instability mystery. SIAM NEWS MARCH 54 (2021) 1, 3 3. Antun, V., Renna, F., Poon, C., Adcock, B., Hansen, A.C.: On instabilities of deep learning in image reconstruction and the potential costs of ai. Proceedings of the National Academy of Sciences 117(48), 30088–30095 (2020) 1, 3, 5 4. Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., Zhong, H., Zhu, Y., Yang, M., Li, Z., Wan, J., Wang, P., Ding, W., Fu, Z., Xu, Y., Ye, J., Zhang, X., Xie, T., Cheng, Z., Zhang, H., Yang, Z., Xu, H., Lin, J.: Qwen2.5-vl technical report (2025) 2, 8 5. Bailey, L., Ong, E., Russell, S., Emmons, S.: Image hijacks: Adversarial images can control generative models at runtime. arXiv preprint arXiv:2309.00236 (2023) 4 6. Barrett, B., Camuto, A., Willetts, M., Rainforth, T.: Certifiably robust variational autoencoders. In: AISTATS. pp. 3663–3683. PMLR (2022) 3 7. Carlini, N., Wagner, D.: Towards evaluating the robustness of neural networks. In: 2017 ieee symposium on security and privacy (sp). pp. 39–57. Ieee (2017) 3, 8 8. Cemgil, T., Ghaisas, S., Dvijotham, K.D., Kohli, P.: Adversarially robust representations with smooth encoders. In: ICLR (2020) 3 9. Chan, A., Tay, Y., Ong, Y.S., Fu, J.: Jacobian adversarially regularized networks for robustness. arXiv preprint arXiv:1912.10185 (2019) 3 10. Cisse, M., Bojanowski, P., Grave, E., Dauphin, Y., Usunier, N.: Parseval networks: Improving robustness to adversarial examples. In: International conference on machine learning. pp. 854–863. PMLR (2017) 3 11. Cui, X., Aparcedo, A., Jang, Y.K., Lim, S.N.: On the robustness of large multimodal models against image adversarial attacks. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. pp. 24625–24634 (2024) 4 12. Deng, J., Dong, W., Socher, R., Li, L.J., Li, K., Fei-Fei, L.: Imagenet: A large-scale hierarchical image database. In: CVPR. pp. 248–255. Ieee (2009) 8 13. Dong, Y., Chen, H., Chen, J., Fang, Z., Yang, X., Zhang, Y., Tian, Y., Su, H., Zhu, J.: How robust is google’s bard to adversarial image attacks? arXiv preprint arXiv:2309.11751 (2023) 4 14. Dosovitskiy, A., Beyer, L., Kolesnikov, A., Weissenborn, D., Zhai, X., Unterthiner, T., Dehghani, M., Minderer, M., Heigold, G., Gelly, S., et al.: An image is worth 16x16 words: Transformers for image recognition at scale. arXiv preprint arXiv:2010.11929 (2020) 2 15. Etmann, C., Lunz, S., Maass, P., Schönlieb, C.B.: On the connection between adversarial robustness and saliency map interpretability. arXiv preprint arXiv:1905.04172 (2019) 3 16. Fawzi, A., Fawzi, O., Frossard, P.: Analysis of classifiers’ robustness to adversarial perturbations. Machine learning 107(3), 481–508 (2018) 1, 3 17. Fazlyab, M., Robey, A., Hassani, H., Morari, M., Pappas, G.: Efficient and accurate estimation of lipschitz constants for deep neural networks. NeurIPS 32 (2019) 3 18. Ganeshan, A., BS, V., Babu, R.V.: Fda: Feature disruptive attack. In: Proceedings of the IEEE/CVF international conference on computer vision. pp. 8069–8079 (2019) 4

16

C.K. Ramanaik et al.

19. Gao, K., Bai, Y., Bai, J., Yang, Y., Xia, S.T.: Adversarial robustness for visual grounding of multimodal large language models. arXiv preprint arXiv:2405.09981 (2024) 4 20. Gao, K., Bai, Y., Gu, J., Xia, S.T., Torr, P., Li, Z., Liu, W.: Inducing high energy-latency of large vision-language models with verbose images. arXiv preprint arXiv:2401.11170 (2024) 4 21. Gemma Team: Gemma 3 technical report (2025) 2, 8 22. Gilmer, J., Metz, L., Faghri, F., Schoenholz, S.S., Raghu, M., Wattenberg, M., Goodfellow, I.: Adversarial spheres (2018). arXiv preprint arXiv:1801.02774 (1801) 1, 3 23. Gondim-Ribeiro, G., Tabacof, P., Valle, E.: Adversarial attacks on variational autoencoders. arXiv (2018) 3 24. Goodfellow, I., Bengio, Y., Courville, A.: Deep Learning. MIT Press (2016), http: //www.deeplearningbook.org 1, 3 25. Goodfellow, I.J., Shlens, J., Szegedy, C.: Explaining and harnessing adversarial examples. arXiv preprint arXiv:1412.6572 (2014) 1, 3 26. Gottschling, N.M., Antun, V., Hansen, A.C., Adcock, B.: The troublesome kernel: On hallucinations, no free lunches, and the accuracy-stability tradeoff in inverse problems. SIAM Review 67(1), 73–104 (2025) 1, 3, 5 27. Gulrajani, I., Ahmed, F., Arjovsky, M., Dumoulin, V., Courville, A.C.: Improved training of wasserstein gans. NeurIPS 30 (2017) 3 28. Gupta, K., Ajanthan, T.: Improved gradient-based adversarial attacks for quantized networks. In: Proceedings of the AAAI Conference on Artificial Intelligence. vol. 36, pp. 6810–6818 (2022) 3 29. He, M., Tian, X., Shen, X., Ni, J., Zou, S., Yang, Z., Zhang, J.: Few tokens matter: Entropy guided attacks on vision-language models. arXiv preprint arXiv:2512.21815 (2025) 4 30. Hein, M., Andriushchenko, M.: Formal guarantees on the robustness of a classifier against adversarial manipulation. Advances in neural information processing systems 30 (2017) 3 31. Hoffman, J., Roberts, D.A., Yaida, S.: Robust learning with jacobian regularization. arXiv preprint arXiv:1908.02729 (2019) 1, 3 32. Huang, Y., Würfl, T., Breininger, K., Liu, L., Lauritsch, G., Maier, A.: Some investigations on robustness of deep learning in limited angle tomography. In: International Conference on Medical Image Computing and Computer-Assisted Intervention. pp. 145–153. Springer (2018) 5 33. Ilyas, A., Santurkar, S., Tsipras, D., Engstrom, L., Tran, B., Madry, A.: Adversarial examples are not bugs, they are features. Advances in neural information processing systems 32 (2019) 1, 3 34. Inkawhich, N., Wen, W., Li, H.H., Chen, Y.: Feature space perturbations yield more transferable adversarial examples. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 7066–7074 (2019) 3 35. Jain, S.B., Shao, Z., Hecht, M.: Automated detection of potential artifacts in machine learning based bio-image segmentation. Machine Learning: Science and Technology 6(4), 045029 (2025) 5 36. Jakubovitz, D., Giryes, R.: Improving dnn robustness to adversarial attacks using jacobian regularization. In: Proceedings of the European conference on computer vision (ECCV). pp. 514–529 (2018) 1, 3 37. Jia, X., Gao, S., Qin, S., Pang, T., Du, C., Huang, Y., Li, X., Li, Y., Li, B., Liu, Y.: Adversarial attacks against closed-source mllms via feature optimal alignment. arXiv preprint arXiv:2505.21494 (2025) 4

Adversarial Vulnerability in VLMs via Spectral Subspaces

17

38. Khoury, M., Hadfield-Menell, D.: On the geometry of adversarial examples. arXiv preprint arXiv:1811.00525 (2018) 1, 3 39. Khrulkov, V., Oseledets, I.: Art of singular vectors and universal adversarial perturbations. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 8562–8570 (2018) 1, 3 40. Kuzina, A., Welling, M., Tomczak, J.M.: Diagnosing vulnerability of variational auto-encoders to adversarial attacks. arXiv (2021) 3 41. Lin, C.Y.: Rouge: A package for automatic evaluation of summaries. In: Text summarization branches out. pp. 74–81 (2004) 8, 20 42. Liu, D., Cai, X., Dong, J., Guo, Z., Qu, X., Guan, R., Fang, X., Ye, D.: Attacking gray-box large vision-language models with adaptive svd-structured adversarial alignment. In: Forty-third International Conference on Machine Learning (2026) 4 43. Lu, Y., Jia, Y., Wang, J., Li, B., Chai, W., Carin, L., Velipasalar, S.: Enhancing cross-task black-box transferability of adversarial examples with dispersion reduction. In: Proceedings of the IEEE/CVF conference on Computer Vision and Pattern Recognition. pp. 940–949 (2020) 4 44. Luo, H., Gu, J., Liu, F., Torr, P.: An image is worth 1000 lies: Adversarial transferability across prompts on vision-language models. arXiv preprint arXiv:2403.09766 (2024) 4 45. Luo, Y., Boix, X., Roig, G., Poggio, T., Zhao, Q.: Foveation-based mechanisms alleviate adversarial examples. arXiv preprint arXiv:1511.06292 (2015) 3 46. Mądry, A., Makelov, A., Schmidt, L., Tsipras, D., Vladu, A.: Towards deep learning models resistant to adversarial attacks. stat 1050, 9 (2017) 3 47. Martin, C.H., Mahoney, M.W.: Implicit self-regularization in deep neural networks: Evidence from random matrix theory and implications for learning. JMLR 22(165), 1–73 (2021) 3 48. Miyato, T., Kataoka, T., Koyama, M., Yoshida, Y.: Spectral normalization for generative adversarial networks. arXiv (2018) 3 49. Moosavi-Dezfooli, S.M., Fawzi, A., Fawzi, O., Frossard, P.: Universal adversarial perturbations. In: Proceedings of the IEEE conference on computer vision and pattern recognition. pp. 1765–1773 (2017) 1, 3 50. Naseer, M., Khan, S., Hayat, M., Khan, F.S., Porikli, F.: A self-supervised approach for adversarial robustness. In: Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. pp. 262–271 (2020) 4 51. Paniagua, T., Savadikar, C., Wu, T.: Adversarial perturbations are formed by iteratively learning linear combinations of the right singular vectors of the adversarial jacobian. In: Forty-second International Conference on Machine Learning (2025) 1, 3 52. Ramanaik, C.K., Roy, A., Callies, T., Ntoutsi, E.: Grill: Gradient signal restoration in ill-conditioned layers to enhance adversarial attacks on autoencoders. arXiv preprint arXiv:2505.03646 (2025) 3 53. Ramanaik, C.K., Willmann, A., Suarez Cardona, J.E., Hanfeld, P., Hoffmann, N., Hecht, M.: Ensuring topological data-structure preservation under autoencoder compression due to latent space regularization in gauss–legendre nodes. Axioms 13(8), 535 (2024) 3 54. Sabour, S., Cao, Y., Faghri, F., Fleet, D.J.: Adversarial manipulation of deep representations. arXiv preprint arXiv:1511.05122 (2015) 3 55. Samangouei, P., Kabkab, M., Chellappa, R.: Defense-gan: Protecting classifiers against adversarial attacks using generative models. arXiv preprint arXiv:1805.06605 (2018) 1, 3

18

C.K. Ramanaik et al.

56. Song, Y., Kim, T., Nowozin, S., Ermon, S., Kushman, N.: Pixeldefend: Leveraging generative models to understand and defend against adversarial examples. arXiv preprint arXiv:1710.10766 (2017) 1, 3 57. Szegedy, C., Zaremba, W., Sutskever, I., Bruna, J., Erhan, D., Goodfellow, I., Fergus, R.: Intriguing properties of neural networks. arXiv preprint arXiv:1312.6199 (2013) 1, 3, 5 58. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A.N., Kaiser, Ł., Polosukhin, I.: Attention is all you need. Advances in neural information processing systems 30 (2017) 2 59. Virmaux, A., Scaman, K.: Lipschitz regularity of deep neural networks: analysis and efficient estimation. Advances in neural information processing systems 31 (2018) 3 60. Wang, Z., Han, Z., Chen, S., Xue, F., Ding, Z., Xiao, X., Tresp, V., Torr, P., Gu, J.: Stop reasoning! when multimodal llm with chain-of-thought reasoning meets adversarial image. arXiv preprint arXiv:2402.14899 (2024) 4 61. Willetts, M., Camuto, A., Rainforth, T., Roberts, S., Holmes, C.: Improving vaes’ robustness to adversarial attack. arXiv (2019) 3 62. Yin, Z., Ye, M., Zhang, T., Du, T., Zhu, J., Liu, H., Chen, J., Wang, T., Ma, F.: Vlattack: Multimodal adversarial attacks on vision-language tasks via pretrained models. Advances in Neural Information Processing Systems 36, 52936– 52956 (2023) 1, 4, 6, 12 63. Yoshida, Y., Miyato, T.: Spectral norm regularization for improving the generalizability of deep learning. arXiv (2017) 3 64. Zhang, H., Yu, Y., Jiao, J., Xing, E., El Ghaoui, L., Jordan, M.: Theoretically principled trade-off between robustness and accuracy. In: International conference on machine learning. pp. 7472–7482. PMLR (2019) 1, 3 65. Zhang, T., Kishore, V., Wu, F., Weinberger, K.Q., Artzi, Y.: Bertscore: Evaluating text generation with bert. arXiv preprint arXiv:1904.09675 (2019) 8, 20

Adversarial Vulnerability in VLMs via Spectral Subspaces

A

Appendix

A.1

Spectral Subspace Guided Representation Attack (SSGRA) Algorithm

19

Algorithm 1 Spectral Subspace Guided Representation Attack (SSGRA) Input: Image x, prompt p, VLM F, perturbation budget c, step size η, number of steps T , selected linear transformations S, trade-off parameter λ, bottom-subspace dimension s Output: Adversarial image x∗a 1: Initialize perturbation with small random noise: 2: δ ∼ U(−ξ, ξ), where ξ ≪ c 3: Fix prompt p and suppress it in the representation notation. v,(j) (i) 4: Compute clean representations {hk (x)}k,j and {Hℓ (x)}ℓ,i 5: for each selected transformation m ∈ S do ⊤ 6: Compute SVD of its weight matrix: Wm = Um Σm Vm bottom 7: Extract the bottom-s right singular-vector subspace Vm,s 8: end for 9: for τ = 1 to T do 10: Construct the adversarial image: 11: xa ← clip(x + δ, x − c, x + c) v,(j) (i) 12: Compute adversarial representations {hk (xa )}k,j and {Hℓ (xa )}ℓ,i 13: Compute BSA representation discrepancy term: LBSA =

Nv K X X

L X N     X (i) (i) v,(j) v,(j) cos Hℓ (x), Hℓ (xa ) cos hk (x), hk (xa ) +

k=1 j=1

ℓ=1 i=1

14: Compute P spectral subspace alignment  term: bottom 15: LSS = m∈S Ψs zm (xa ), Vm,s 16: Compute SSGRA objective: 17: LSSGRA = −λLBSA + (1 − λ)LSS 18: Update perturbation by Adam ascent: 19: δ ← AdamStep(δ, ∇δ LSSGRA ) 20: Project perturbation onto the L∞ ball: 21: δ ← clip(δ, −c, c) 22: end for 23: Construct the final adversarial image: 24: x∗a ← clip(x + δ, x − c, x + c) 25: return x∗a

Algorithm 1 summarizes the optimization procedure of SSGRA. The attack first fixes the textual prompt and computes the clean intermediate representations of the input image across the visual encoder and language-model blocks. For each selected intermediate linear transformation, SSGRA performs an SVD

20

C.K. Ramanaik et al.

of the corresponding weight matrix and extracts the bottom-s right singularvector subspace. These subspaces define the information-attenuating directions used by the spectral alignment objective. During optimization, the adversarial image is constructed by adding a learnable perturbation δ to the clean image and clipping it within the prescribed L∞ budget. At each iteration, the model is evaluated on the current adversarial image to obtain the corresponding intermediate representations. SSGRA then combines two objectives. The first is the BSA representation-discrepancy term, which reduces the similarity between clean and adversarial feature representations across visual and language layers. The second is the proposed spectral subspace alignment term, which encourages adversarial representations before the selected transformations to align with the bottom singular-vector subspaces. The trade-off parameter λ balances these two effects. The perturbation is updated by Adam ascent on the combined SSGRA objective, since the attack maximizes representation disruption and spectral alignment. After each update, the perturbation is projected back onto the L∞ ball to ensure that the adversarial image remains visually close to the original input. The final adversarial example is obtained by applying the optimized perturbation to the clean image and clipping it to satisfy the perturbation constraint. The algorithm explicitly guides adversarial representations toward directions that are strongly attenuated by intermediate linear transformations, thereby weakening semantic information propagation through the VLM. A.2

Evaluation Metrics Details

We evaluate attack effectiveness using two complementary text-based metrics that compare the adversarial model output ŷa against the clean output ŷc . BERTScore. BERTScore [65] computes token-level semantic similarity between ŷa and ŷc using contextual embeddings from a pretrained language model (RoBERTa-large). For each token ai ∈ ŷa and cj ∈ ŷc , cosine similarity is computed in embedding space. Precision, recall, and F1 are defined as: P_{\text {BERT}} = \frac {1}{|\hat {y}_a|}\sum _{a_i \in \hat {y}_a} \max _{c_j \in \hat {y}_c} \cos (\mathbf {e}_{a_i}, \mathbf {e}_{c_j}),

(4)

R_{\text {BERT}} = \frac {1}{|\hat {y}_c|}\sum _{c_j \in \hat {y}_c} \max _{a_i \in \hat {y}_a} \cos (\mathbf {e}_{a_i}, \mathbf {e}_{c_j}),

(5)

F1_{\text {BERT}} = \frac {2 \cdot P_{\text {BERT}} \cdot R_{\text {BERT}}}{P_{\text {BERT}} + R_{\text {BERT}}},

(6)

where eai and ecj are the contextual embeddings of tokens ai and cj respectively. Lower scores indicate greater semantic degradation of the adversarial output relative to the clean output. ROUGE-L. ROUGE-L [41] measures lexical overlap via the Longest Common Subsequence (LCS) between ŷa and ŷc . Let LCS(ŷa , ŷc ) denote the length of the

Adversarial Vulnerability in VLMs via Spectral Subspaces

21

longest common subsequence. Precision, recall, and F1 are: P_{\text {R}} = \frac {|\text {LCS}(\hat {y}_a, \hat {y}_c)|}{|\hat {y}_a|}, \quad R_{\text {R}} = \frac {|\text {LCS}(\hat {y}_a, \hat {y}_c)|}{|\hat {y}_c|}, \quad F1_{\text {R}} = \frac {2 \cdot P_{\text {R}} \cdot R_{\text {R}}}{P_{\text {R}} + R_{\text {R}}}.

(7)

Unlike BERTScore, ROUGE-L is sensitive to structural content loss: a low recall indicates that the adversarial output fails to reproduce key content from the clean description. The two metrics are complementary — BERTScore captures semantic similarity robust to paraphrase, while ROUGE-L captures lexical fidelity and structural degradation. We report mean and standard deviation over 100 images per experimental configuration. A.3

Comprehensive Quantitative Results

Figures 8 and 9 visualize the Precision, Recall, and F1 trends under BERTScore and ROUGE-L across perturbation budgets, while Tables 4, 5, and 6 report the corresponding numerical results. Across all three VLMs, SSGRA consistently achieves the lowest BERTScore and ROUGE-L scores in most settings, with the largest improvements on Qwen2.5-VL and the smallest on Gemma 3, consistent with the spectral characterization presented in Section 5.4. In addition to the F1 scores reported in Section 5.2, here we provide the corresponding Precision and Recall values, enabling a more detailed analysis of attack behavior. The results show that the improvements achieved by SSGRA are not driven by a single evaluation component but are consistently reflected across all three metrics. Furthermore, the complete numerical results complement the plots by reporting the mean and standard deviation for every perturbation budget, providing a comprehensive view of both attack effectiveness and its variability across the evaluated samples. A.4

Additional Qualitative Results

Figures 10–13 present additional qualitative examples for Qwen2.5-VL, LLaVA1.5, and Gemma 3, generated with a perturbation budget of c = 0.002 in Figures 10–12 and c = 0.003 in Figure 13. SSGRA produces larger semantic deviations than the baseline attacks while maintaining the visual appearance of the input images. Whereas baseline methods often preserve the correct semantic content or introduce only minor hallucinations, SSGRA more frequently induces incorrect object categories, unrelated scene descriptions, and semantically inconsistent responses. These qualitative results are consistent with the quantitative improvements reported in Section 5.2 and further support the effectiveness of spectral subspace guidance.

22

C.K. Ramanaik et al.

Table 3: Performance of different attack methods under varying perturbation budgets. (a) BERTScore F1. (b) ROUGE-L F1. Lower values indicate stronger attack effectiveness. (a) BERTScore F1 Method

c = 0.002

c = 0.0025

c = 0.003

c = 0.0035

c = 0.004

c = 0.0045

c = 0.005

Qwen2.5-VL BSA 0.895 ± 0.028 0.877 ± 0.035 0.874 ± 0.032 0.864 ± 0.037 0.861 ± 0.035 0.856 ± 0.038 0.845 ± 0.037 DRA 0.938 ± 0.026 0.943 ± 0.023 0.936 ± 0.023 0.934 ± 0.023 0.935 ± 0.021 0.936 ± 0.023 0.931 ± 0.018 FDA 0.933 ± 0.021 0.928 ± 0.025 0.929 ± 0.023 0.925 ± 0.022 0.922 ± 0.023 0.918 ± 0.022 0.918 ± 0.024 SSPA 0.926 ± 0.018 0.928 ± 0.020 0.919 ± 0.020 0.922 ± 0.022 0.917 ± 0.020 0.913 ± 0.019 0.916 ± 0.019 EGA 0.896 ± 0.022 0.883 ± 0.022 0.883 ± 0.024 0.865 ± 0.053 0.872 ± 0.031 0.856 ± 0.061 0.865 ± 0.034 CE 0.886 ± 0.016 0.884 ± 0.017 0.873 ± 0.028 0.875 ± 0.034 0.871 ± 0.030 0.865 ± 0.048 0.856 ± 0.048 SSGRA 0.816 ± 0.088 0.768 ± 0.147 0.743 ± 0.137 0.704 ± 0.217 0.754 ± 0.042 0.687 ± 0.206 0.724 ± 0.121 Gain over Best (%) 7.90 % 12.40 % 14.89 % 18.52 % 12.43 % 19.74 % 14.32% LLaVa-1.5 BSA 0.926 ± 0.034 DRA 0.951 ± 0.033 FDA 0.954 ± 0.027 SSPA 0.949 ± 0.027 EGA 0.933 ± 0.031 CE 0.921 ± 0.022 SSGRA 0.916 ± 0.029 Gain over Best (%) 0.54 %

0.911 ± 0.034 0.940 ± 0.022 0.946 ± 0.026 0.947 ± 0.024 0.922 ± 0.023 0.916 ± 0.022 0.912 ± 0.027 –

0.902 ± 0.028 0.895 ± 0.019 0.892 ± 0.020 0.886 ± 0.028 0.884 ± 0.019 0.935 ± 0.027 0.932 ± 0.030 0.926 ± 0.027 0.914 ± 0.033 0.915 ± 0.028 0.945 ± 0.027 0.944 ± 0.027 0.939 ± 0.026 0.942 ± 0.026 0.939 ± 0.025 0.936 ± 0.021 0.934 ± 0.023 0.933 ± 0.026 0.928 ± 0.022 0.934 ± 0.026 0.918 ± 0.026 0.913 ± 0.027 0.903 ± 0.027 0.885 ± 0.131 0.875 ± 0.133 0.911 ± 0.020 0.909 ± 0.018 0.907 ± 0.019 0.888 ± 0.129 0.883 ± 0.130 0.882 ± 0.133 0.837 ± 0.216 0.853 ± 0.179 0.819 ± 0.213 0.816 ± 0.214 2.22 % 6.48 % 4.37 % 7.46 % 6.74%

0.905 ± 0.035 0.919 ± 0.027 0.927 ± 0.036 0.923 ± 0.022 0.911 ± 0.029 0.886 ± 0.026 0.903 ± 0.038 –

0.890 ± 0.034 0.918 ± 0.025 0.924 ± 0.037 0.918 ± 0.028 0.918 ± 0.031 0.883 ± 0.025 0.883 ± 0.042 –

0.880 ± 0.035 0.870 ± 0.043 0.861 ± 0.031 0.856 ± 0.031 0.914 ± 0.029 0.911 ± 0.025 0.908 ± 0.029 0.909 ± 0.028 0.923 ± 0.032 0.919 ± 0.034 0.923 ± 0.032 0.921 ± 0.030 0.916 ± 0.032 0.913 ± 0.029 0.913 ± 0.031 0.905 ± 0.027 0.908 ± 0.030 0.907 ± 0.031 0.901 ± 0.031 0.908 ± 0.030 0.879 ± 0.030 0.879 ± 0.026 0.877 ± 0.033 0.881 ± 0.034 0.873 ± 0.036 0.868 ± 0.032 0.858 ± 0.034 0.855 ± 0.037 0.80 % 0.23 % 0.35 % 0.12 %

c = 0.002

c = 0.0025

Gemma 3 BSA DRA FDA SSPA EGA CE SSGRA Gain over Best (%)

0.847 ± 0.037 0.903 ± 0.030 0.918 ± 0.026 0.909 ± 0.031 0.902 ± 0.030 0.867 ± 0.041 0.849 ± 0.030 –

(b) ROUGE-L F1 Method

c = 0.003

c = 0.0035

c = 0.004

c = 0.0045

c = 0.005

Qwen2.5-VL BSA 0.352 ± 0.116 0.296 ± 0.115 0.285 ± 0.097 0.252 ± 0.118 0.232 ± 0.113 0.228 ± 0.091 0.199 ± 0.114 DRA 0.551 ± 0.187 0.577 ± 0.171 0.536 ± 0.170 0.539 ± 0.164 0.528 ± 0.149 0.533 ± 0.159 0.501 ± 0.132 FDA 0.517 ± 0.160 0.486 ± 0.153 0.494 ± 0.153 0.471 ± 0.148 0.447 ± 0.144 0.440 ± 0.131 0.434 ± 0.139 SSPA 0.498 ± 0.122 0.494 ± 0.136 0.437 ± 0.099 0.447 ± 0.142 0.428 ± 0.119 0.406 ± 0.107 0.427 ± 0.121 EGA 0.337 ± 0.113 0.294 ± 0.086 0.293 ± 0.097 0.277 ± 0.100 0.271 ± 0.075 0.261 ± 0.077 0.274 ± 0.065 CE 0.282 ± 0.066 0.277 ± 0.058 0.239 ± 0.071 0.256 ± 0.076 0.235 ± 0.057 0.226 ± 0.065 0.217 ± 0.081 SSGRA 0.196 ± 0.235 0.102 ± 0.171 0.057 ± 0.164 0.052 ± 0.155 0.028 ± 0.114 0.010 ± 0.047 0.010 ± 0.020 Gain over Best (%) 30.50 % 63.18 % 76.15 % 79.37 % 97.93 % 95.57 % 94.97% LLaVa-1.5 BSA 0.492 ± 0.215 DRA 0.639 ± 0.212 FDA 0.644 ± 0.195 SSPA 0.620 ± 0.202 EGA 0.511 ± 0.202 CE 0.439 ± 0.134 SSGRA 0.422 ± 0.169 Gain over Best (%) 3.87 %

0.416 ± 0.196 0.545 ± 0.164 0.591 ± 0.182 0.594 ± 0.177 0.456 ± 0.138 0.400 ± 0.137 0.406 ± 0.157 –

0.369 ± 0.135 0.330 ± 0.080 0.320 ± 0.103 0.297 ± 0.077 0.291 ± 0.065 0.525 ± 0.183 0.514 ± 0.201 0.477 ± 0.152 0.426 ± 0.164 0.432 ± 0.156 0.590 ± 0.203 0.585 ± 0.192 0.550 ± 0.182 0.570 ± 0.184 0.558 ± 0.182 0.524 ± 0.144 0.520 ± 0.154 0.520 ± 0.170 0.483 ± 0.147 0.518 ± 0.173 0.427 ± 0.162 0.416 ± 0.151 0.371 ± 0.138 0.377 ± 0.143 0.336 ± 0.117 0.375 ± 0.122 0.375 ± 0.094 0.357 ± 0.096 0.340 ± 0.088 0.329 ± 0.100 0.364 ± 0.184 0.299 ± 0.138 0.297 ± 0.132 0.252 ± 0.121 0.236 ± 0.138 1.35 % 9.39 % 7.19 % 15.15 % 18.90%

0.406 ± 0.142 0.456 ± 0.140 0.516 ± 0.186 0.464 ± 0.118 0.423 ± 0.118 0.289 ± 0.080 0.407 ± 0.136 –

0.331 ± 0.116 0.457 ± 0.121 0.509 ± 0.186 0.441 ± 0.131 0.463 ± 0.144 0.285 ± 0.071 0.319 ± 0.118 –

0.315 ± 0.112 0.433 ± 0.131 0.492 ± 0.166 0.440 ± 0.141 0.408 ± 0.125 0.269 ± 0.079 0.288 ± 0.119 –

Gemma 3 BSA DRA FDA SSPA EGA CE SSGRA Gain over Best (%)

0.287 ± 0.116 0.405 ± 0.120 0.471 ± 0.157 0.418 ± 0.134 0.410 ± 0.128 0.275 ± 0.082 0.272 ± 0.105 1.09 %

0.253 ± 0.098 0.388 ± 0.134 0.478 ± 0.139 0.422 ± 0.129 0.382 ± 0.127 0.273 ± 0.088 0.259 ± 0.104 –

0.238 ± 0.092 0.402 ± 0.126 0.474 ± 0.152 0.388 ± 0.128 0.420 ± 0.131 0.285 ± 0.081 0.241 ± 0.088 –

0.228 ± 0.109 0.365 ± 0.125 0.450 ± 0.121 0.404 ± 0.126 0.390 ± 0.119 0.259 ± 0.085 0.221 ± 0.093 3.07%

Adversarial Vulnerability in VLMs via Spectral Subspaces

(a) Precision(Qwen2.5-VL)

(b) Recall (Qwen2.5-VL)

(c) F1 (Qwen2.5-VL)

(d) Precision (LLaVa-1.5)

(e) Recall (LLaVa-1.5)

(f ) F1 (LLaVa-1.5)

(g) Precision (Gemma 3)

(h) Recall (Gemma 3)

(i) F1 (Gemma 3)

23

Fig. 8: BERT-score comparison of different adversarial attack methods across perturbation budgets for Qwen 2.5-VL, LLaVa 1.5, and Gemma 3. Each row corresponds to a model, while the columns show Precision, Recall, and F1 score, respectively.

24

C.K. Ramanaik et al.

(a) Precision (Qwen2.5-VL)

(b) Recall (Qwen2.5-VL)

(c) F1 (Qwen2.5-VL)

(d) Precision (LLaVa-1.5)

(e) Recall (LLaVa-1.5)

(f ) F1 (LLaVa-1.5)

(g) Precision (Gemma 3)

(h) Recall (Gemma 3)

(i) F1 (Gemma 3)

Fig. 9: ROUGE-L score comparison of sample-specific attacks across Qwen 2.5-VL, LLaVa 1.5, and Gemma 3. The three columns report Precision, Recall, and F1 score, respectively, while each row corresponds to a different vision-language model.

Adversarial Vulnerability in VLMs via Spectral Subspaces

25

Table 4: Performance of different attack methods on Qwen2.5-VL under varying perturbation budgets c. Top: BERTScore (Precision, Recall and F1, Mean±Std). Bottom: ROUGE-L (Precision, Recall and F1, Mean±Std). (a) BERTScore Method Metric

c = 0.002

c = 0.0025

c = 0.003

c = 0.0035

c = 0.004

c = 0.0045

c = 0.005

BSA

P R F1

0.896 ± 0.029 0.879 ± 0.037 0.875 ± 0.036 0.866 ± 0.036 0.866 ± 0.032 0.856 ± 0.038 0.847 ± 0.040 0.894 ± 0.029 0.876 ± 0.034 0.874 ± 0.031 0.862 ± 0.039 0.856 ± 0.040 0.856 ± 0.033 0.843 ± 0.038 0.895 ± 0.028 0.877 ± 0.034 0.874 ± 0.032 0.864 ± 0.037 0.861 ± 0.035 0.856 ± 0.034 0.845 ± 0.037

DRA

P R F1

0.937 ± 0.026 0.943 ± 0.025 0.936 ± 0.024 0.935 ± 0.024 0.935 ± 0.022 0.935 ± 0.025 0.931 ± 0.020 0.938 ± 0.028 0.943 ± 0.023 0.935 ± 0.024 0.934 ± 0.023 0.935 ± 0.023 0.936 ± 0.022 0.931 ± 0.019 0.938 ± 0.026 0.943 ± 0.023 0.936 ± 0.023 0.934 ± 0.023 0.935 ± 0.021 0.936 ± 0.023 0.931 ± 0.018

FDA

P R F1

0.935 ± 0.022 0.932 ± 0.027 0.931 ± 0.024 0.930 ± 0.025 0.926 ± 0.025 0.922 ± 0.023 0.923 ± 0.025 0.931 ± 0.021 0.924 ± 0.025 0.926 ± 0.023 0.920 ± 0.022 0.918 ± 0.023 0.913 ± 0.024 0.914 ± 0.025 0.933 ± 0.021 0.928 ± 0.025 0.929 ± 0.023 0.925 ± 0.022 0.922 ± 0.023 0.918 ± 0.022 0.918 ± 0.024

SSPA

P R F1

0.932 ± 0.019 0.932 ± 0.020 0.924 ± 0.023 0.929 ± 0.025 0.920 ± 0.022 0.918 ± 0.021 0.922 ± 0.020 0.921 ± 0.020 0.923 ± 0.023 0.915 ± 0.020 0.915 ± 0.022 0.914 ± 0.022 0.909 ± 0.021 0.910 ± 0.021 0.926 ± 0.018 0.928 ± 0.020 0.919 ± 0.020 0.922 ± 0.022 0.917 ± 0.020 0.913 ± 0.019 0.916 ± 0.019

EGA

P R F1

0.893 ± 0.025 0.878 ± 0.025 0.878 ± 0.028 0.852 ± 0.071 0.864 ± 0.041 0.841 ± 0.084 0.852 ± 0.046 0.900 ± 0.022 0.889 ± 0.025 0.887 ± 0.024 0.880 ± 0.032 0.882 ± 0.026 0.874 ± 0.035 0.879 ± 0.023 0.896 ± 0.022 0.883 ± 0.022 0.883 ± 0.024 0.865 ± 0.052 0.872 ± 0.031 0.856 ± 0.061 0.865 ± 0.034

CE

P R F1

0.886 ± 0.018 0.880 ± 0.020 0.870 ± 0.033 0.872 ± 0.044 0.868 ± 0.039 0.858 ± 0.046 0.848 ± 0.062 0.887 ± 0.018 0.888 ± 0.018 0.876 ± 0.025 0.878 ± 0.026 0.874 ± 0.024 0.872 ± 0.022 0.864 ± 0.033 0.886 ± 0.016 0.884 ± 0.017 0.873 ± 0.028 0.875 ± 0.034 0.871 ± 0.030 0.865 ± 0.034 0.856 ± 0.048

P SSGRA R F1

0.800 ± 0.108 0.748 ± 0.151 0.716 ± 0.138 0.680 ± 0.214 0.726 ± 0.054 0.660 ± 0.200 0.695 ± 0.118 0.835 ± 0.067 0.790 ± 0.144 0.772 ± 0.137 0.730 ± 0.222 0.785 ± 0.033 0.718 ± 0.214 0.756 ± 0.126 0.816 ± 0.088 0.768 ± 0.147 0.743 ± 0.137 0.704 ± 0.217 0.754 ± 0.042 0.687 ± 0.206 0.724 ± 0.121

(b) ROUGE-L Method Metric

c = 0.002

c = 0.0025

c = 0.003

c = 0.0035

c = 0.004

c = 0.0045

c = 0.005

BSA

P R F1

0.358 ± 0.131 0.303 ± 0.135 0.292 ± 0.108 0.253 ± 0.133 0.261 ± 0.128 0.231 ± 0.108 0.210 ± 0.129 0.353 ± 0.116 0.298 ± 0.111 0.285 ± 0.096 0.260 ± 0.116 0.223 ± 0.107 0.234 ± 0.088 0.199 ± 0.112 0.352 ± 0.116 0.296 ± 0.115 0.285 ± 0.097 0.252 ± 0.118 0.232 ± 0.113 0.228 ± 0.091 0.199 ± 0.114

DRA

P R F1

0.549 ± 0.185 0.579 ± 0.185 0.538 ± 0.174 0.545 ± 0.177 0.522 ± 0.150 0.532 ± 0.171 0.500 ± 0.140 0.562 ± 0.202 0.586 ± 0.175 0.540 ± 0.179 0.541 ± 0.166 0.542 ± 0.163 0.540 ± 0.156 0.511 ± 0.139 0.551 ± 0.187 0.577 ± 0.171 0.536 ± 0.170 0.539 ± 0.164 0.528 ± 0.149 0.533 ± 0.159 0.501 ± 0.132

FDA

P R F1

0.519 ± 0.163 0.513 ± 0.171 0.504 ± 0.162 0.497 ± 0.166 0.468 ± 0.149 0.463 ± 0.140 0.452 ± 0.159 0.519 ± 0.162 0.471 ± 0.148 0.491 ± 0.156 0.455 ± 0.145 0.436 ± 0.150 0.428 ± 0.143 0.426 ± 0.136 0.517 ± 0.160 0.486 ± 0.153 0.494 ± 0.153 0.471 ± 0.148 0.447 ± 0.144 0.440 ± 0.131 0.434 ± 0.139

SSPA

P R F1

0.536 ± 0.127 0.524 ± 0.146 0.464 ± 0.115 0.486 ± 0.165 0.447 ± 0.127 0.432 ± 0.115 0.465 ± 0.137 0.472 ± 0.131 0.476 ± 0.145 0.421 ± 0.105 0.421 ± 0.137 0.423 ± 0.139 0.393 ± 0.117 0.405 ± 0.126 0.498 ± 0.122 0.494 ± 0.136 0.437 ± 0.099 0.447 ± 0.142 0.428 ± 0.119 0.406 ± 0.107 0.427 ± 0.121

EGA

P R F1

0.327 ± 0.111 0.296 ± 0.097 0.304 ± 0.104 0.300 ± 0.135 0.291 ± 0.101 0.327 ± 0.154 0.271 ± 0.070 0.352 ± 0.123 0.316 ± 0.106 0.305 ± 0.108 0.296 ± 0.117 0.280 ± 0.092 0.263 ± 0.098 0.292 ± 0.081 0.337 ± 0.113 0.294 ± 0.086 0.293 ± 0.097 0.277 ± 0.100 0.271 ± 0.075 0.261 ± 0.077 0.274 ± 0.065

CE

P R F1

0.278 ± 0.073 0.269 ± 0.069 0.229 ± 0.079 0.249 ± 0.081 0.236 ± 0.077 0.217 ± 0.074 0.212 ± 0.084 0.295 ± 0.081 0.299 ± 0.070 0.265 ± 0.088 0.273 ± 0.089 0.250 ± 0.072 0.250 ± 0.072 0.237 ± 0.100 0.282 ± 0.066 0.277 ± 0.058 0.239 ± 0.071 0.256 ± 0.076 0.235 ± 0.057 0.226 ± 0.065 0.217 ± 0.081

P SSGRA R F1

0.218 ± 0.252 0.179 ± 0.281 0.089 ± 0.229 0.086 ± 0.189 0.083 ± 0.246 0.085 ± 0.243 0.205 ± 0.372 0.190 ± 0.230 0.099 ± 0.174 0.054 ± 0.163 0.051 ± 0.156 0.027 ± 0.117 0.009 ± 0.045 0.005 ± 0.010 0.196 ± 0.235 0.102 ± 0.171 0.057 ± 0.164 0.052 ± 0.155 0.028 ± 0.114 0.010 ± 0.047 0.010 ± 0.020

26

C.K. Ramanaik et al.

Table 5: Performance of different attack methods on LLaVa-1.5 under varying perturbation budgets c. Top: BERTScore (Precision, Recall and F1, Mean±Std). Bottom: ROUGE-L (Precision, Recall and F1, Mean±Std). (a) BERTScore Method Metric

c = 0.002

c = 0.0025

c = 0.003

c = 0.0035

c = 0.004

c = 0.0045

c = 0.005

BSA

P R F1

0.925 ± 0.034 0.909 ± 0.034 0.901 ± 0.028 0.893 ± 0.019 0.891 ± 0.020 0.883 ± 0.035 0.885 ± 0.019 0.926 ± 0.035 0.914 ± 0.034 0.904 ± 0.029 0.897 ± 0.021 0.893 ± 0.022 0.888 ± 0.022 0.884 ± 0.021 0.926 ± 0.034 0.911 ± 0.034 0.902 ± 0.028 0.895 ± 0.019 0.892 ± 0.020 0.886 ± 0.028 0.884 ± 0.019

DRA

P R F1

0.950 ± 0.037 0.940 ± 0.022 0.935 ± 0.027 0.934 ± 0.029 0.926 ± 0.026 0.914 ± 0.039 0.917 ± 0.028 0.953 ± 0.029 0.939 ± 0.024 0.936 ± 0.028 0.931 ± 0.031 0.926 ± 0.029 0.915 ± 0.031 0.914 ± 0.028 0.951 ± 0.033 0.940 ± 0.022 0.935 ± 0.027 0.932 ± 0.030 0.926 ± 0.027 0.914 ± 0.033 0.915 ± 0.028

FDA

P R F1

0.954 ± 0.027 0.946 ± 0.027 0.946 ± 0.028 0.945 ± 0.027 0.939 ± 0.026 0.943 ± 0.026 0.938 ± 0.025 0.953 ± 0.027 0.946 ± 0.027 0.945 ± 0.028 0.944 ± 0.028 0.938 ± 0.027 0.942 ± 0.027 0.940 ± 0.026 0.954 ± 0.027 0.946 ± 0.026 0.945 ± 0.027 0.944 ± 0.027 0.939 ± 0.026 0.942 ± 0.026 0.939 ± 0.025

SSPA

P R F1

0.950 ± 0.028 0.947 ± 0.025 0.936 ± 0.021 0.934 ± 0.023 0.932 ± 0.026 0.928 ± 0.022 0.934 ± 0.026 0.949 ± 0.027 0.947 ± 0.025 0.936 ± 0.022 0.934 ± 0.025 0.935 ± 0.026 0.929 ± 0.024 0.933 ± 0.027 0.949 ± 0.027 0.947 ± 0.024 0.936 ± 0.021 0.934 ± 0.023 0.933 ± 0.026 0.928 ± 0.022 0.934 ± 0.026

EGA

P R F1

0.933 ± 0.032 0.920 ± 0.023 0.917 ± 0.027 0.910 ± 0.028 0.901 ± 0.029 0.881 ± 0.132 0.868 ± 0.137 0.934 ± 0.031 0.924 ± 0.025 0.919 ± 0.027 0.915 ± 0.027 0.906 ± 0.028 0.890 ± 0.131 0.882 ± 0.130 0.933 ± 0.031 0.922 ± 0.023 0.918 ± 0.026 0.913 ± 0.027 0.903 ± 0.027 0.885 ± 0.131 0.875 ± 0.133

CE

P R F1

0.919 ± 0.024 0.915 ± 0.023 0.909 ± 0.022 0.907 ± 0.019 0.906 ± 0.021 0.887 ± 0.129 0.880 ± 0.130 0.923 ± 0.021 0.916 ± 0.022 0.912 ± 0.020 0.911 ± 0.019 0.909 ± 0.018 0.889 ± 0.130 0.886 ± 0.130 0.921 ± 0.022 0.916 ± 0.022 0.911 ± 0.020 0.909 ± 0.018 0.907 ± 0.019 0.888 ± 0.129 0.883 ± 0.130

P SSGRA R F1

0.915 ± 0.029 0.911 ± 0.028 0.879 ± 0.134 0.834 ± 0.217 0.852 ± 0.179 0.815 ± 0.214 0.811 ± 0.216 0.917 ± 0.029 0.913 ± 0.028 0.885 ± 0.133 0.839 ± 0.216 0.855 ± 0.179 0.824 ± 0.213 0.821 ± 0.213 0.916 ± 0.029 0.912 ± 0.027 0.882 ± 0.133 0.837 ± 0.216 0.853 ± 0.179 0.819 ± 0.213 0.816 ± 0.214

(b) ROUGE-L Method Metric

c = 0.002

c = 0.0025

c = 0.003

c = 0.0035

c = 0.004

c = 0.0045

c = 0.005

BSA

P R F1

0.495 ± 0.214 0.408 ± 0.195 0.363 ± 0.136 0.319 ± 0.074 0.320 ± 0.100 0.290 ± 0.075 0.307 ± 0.118 0.495 ± 0.219 0.432 ± 0.204 0.383 ± 0.145 0.350 ± 0.105 0.331 ± 0.120 0.312 ± 0.093 0.299 ± 0.072 0.492 ± 0.215 0.416 ± 0.196 0.369 ± 0.135 0.330 ± 0.080 0.320 ± 0.103 0.297 ± 0.077 0.291 ± 0.065

DRA

P R F1

0.638 ± 0.210 0.553 ± 0.162 0.527 ± 0.179 0.526 ± 0.192 0.486 ± 0.154 0.441 ± 0.163 0.448 ± 0.166 0.644 ± 0.217 0.542 ± 0.174 0.529 ± 0.193 0.507 ± 0.214 0.477 ± 0.166 0.422 ± 0.175 0.425 ± 0.156 0.639 ± 0.212 0.545 ± 0.164 0.525 ± 0.183 0.514 ± 0.201 0.477 ± 0.152 0.426 ± 0.164 0.432 ± 0.156

FDA

P R F1

0.651 ± 0.189 0.598 ± 0.183 0.599 ± 0.203 0.588 ± 0.190 0.556 ± 0.174 0.580 ± 0.184 0.553 ± 0.181 0.642 ± 0.204 0.592 ± 0.191 0.586 ± 0.206 0.588 ± 0.199 0.550 ± 0.193 0.567 ± 0.192 0.571 ± 0.191 0.644 ± 0.195 0.591 ± 0.182 0.590 ± 0.203 0.585 ± 0.192 0.550 ± 0.182 0.570 ± 0.184 0.558 ± 0.182

SSPA

P R F1

0.632 ± 0.202 0.598 ± 0.180 0.531 ± 0.146 0.527 ± 0.150 0.521 ± 0.169 0.490 ± 0.145 0.525 ± 0.168 0.615 ± 0.207 0.599 ± 0.185 0.524 ± 0.155 0.520 ± 0.172 0.524 ± 0.178 0.483 ± 0.158 0.520 ± 0.187 0.620 ± 0.202 0.594 ± 0.177 0.524 ± 0.144 0.520 ± 0.154 0.520 ± 0.170 0.483 ± 0.147 0.518 ± 0.173

EGA

P R F1

0.513 ± 0.206 0.449 ± 0.140 0.437 ± 0.166 0.407 ± 0.156 0.377 ± 0.147 0.373 ± 0.147 0.341 ± 0.114 0.515 ± 0.205 0.475 ± 0.159 0.435 ± 0.175 0.434 ± 0.157 0.383 ± 0.146 0.392 ± 0.157 0.350 ± 0.140 0.511 ± 0.202 0.456 ± 0.138 0.427 ± 0.162 0.416 ± 0.151 0.371 ± 0.138 0.377 ± 0.143 0.336 ± 0.117

CE

P R F1

0.432 ± 0.143 0.405 ± 0.143 0.383 ± 0.128 0.381 ± 0.109 0.360 ± 0.103 0.336 ± 0.092 0.329 ± 0.104 0.458 ± 0.140 0.406 ± 0.149 0.379 ± 0.133 0.379 ± 0.103 0.362 ± 0.108 0.354 ± 0.105 0.342 ± 0.116 0.439 ± 0.134 0.400 ± 0.137 0.375 ± 0.122 0.375 ± 0.094 0.357 ± 0.096 0.340 ± 0.088 0.329 ± 0.100

P SSGRA R F1

0.421 ± 0.168 0.402 ± 0.156 0.368 ± 0.182 0.309 ± 0.145 0.311 ± 0.142 0.278 ± 0.157 0.259 ± 0.152 0.432 ± 0.178 0.416 ± 0.164 0.374 ± 0.196 0.310 ± 0.160 0.303 ± 0.142 0.266 ± 0.141 0.236 ± 0.152 0.422 ± 0.169 0.406 ± 0.157 0.364 ± 0.184 0.299 ± 0.138 0.297 ± 0.132 0.252 ± 0.121 0.236 ± 0.138

Adversarial Vulnerability in VLMs via Spectral Subspaces

27

Table 6: Performance of different attack methods on Gemma 3 under varying perturbation budgets c. Top: BERTScore (Precision, Recall and F1, Mean±Std). Bottom: ROUGE-L (Precision, Recall and F1, Mean±Std). (a) BERTScore Method Metric

c = 0.002

c = 0.0025

c = 0.003

c = 0.0035

c = 0.004

c = 0.0045

c = 0.005

BSA

P R F1

0.904 ± 0.037 0.890 ± 0.035 0.879 ± 0.037 0.870 ± 0.047 0.863 ± 0.030 0.860 ± 0.027 0.852 ± 0.040 0.907 ± 0.036 0.890 ± 0.036 0.880 ± 0.036 0.871 ± 0.040 0.859 ± 0.035 0.852 ± 0.037 0.843 ± 0.039 0.905 ± 0.035 0.890 ± 0.034 0.880 ± 0.035 0.870 ± 0.043 0.861 ± 0.031 0.856 ± 0.031 0.847 ± 0.037

DRA

P R F1

0.918 ± 0.030 0.918 ± 0.027 0.915 ± 0.030 0.910 ± 0.027 0.907 ± 0.031 0.909 ± 0.029 0.903 ± 0.033 0.920 ± 0.028 0.918 ± 0.027 0.914 ± 0.032 0.912 ± 0.025 0.909 ± 0.029 0.909 ± 0.030 0.904 ± 0.031 0.919 ± 0.027 0.918 ± 0.025 0.914 ± 0.029 0.911 ± 0.025 0.908 ± 0.029 0.909 ± 0.028 0.903 ± 0.030

FDA

P R F1

0.928 ± 0.037 0.925 ± 0.036 0.924 ± 0.032 0.920 ± 0.034 0.923 ± 0.031 0.921 ± 0.030 0.918 ± 0.027 0.926 ± 0.038 0.924 ± 0.040 0.923 ± 0.035 0.919 ± 0.037 0.922 ± 0.034 0.921 ± 0.032 0.918 ± 0.028 0.927 ± 0.036 0.924 ± 0.037 0.923 ± 0.032 0.919 ± 0.034 0.923 ± 0.032 0.921 ± 0.030 0.918 ± 0.026

SSPA

P R F1

0.924 ± 0.024 0.917 ± 0.028 0.919 ± 0.030 0.913 ± 0.030 0.915 ± 0.030 0.907 ± 0.027 0.911 ± 0.030 0.923 ± 0.024 0.919 ± 0.030 0.915 ± 0.037 0.913 ± 0.030 0.912 ± 0.033 0.905 ± 0.031 0.907 ± 0.035 0.923 ± 0.022 0.918 ± 0.028 0.916 ± 0.032 0.913 ± 0.029 0.913 ± 0.031 0.905 ± 0.027 0.909 ± 0.031

EGA

P R F1

0.910 ± 0.030 0.919 ± 0.030 0.907 ± 0.031 0.906 ± 0.030 0.900 ± 0.031 0.906 ± 0.031 0.901 ± 0.031 0.913 ± 0.030 0.918 ± 0.034 0.910 ± 0.032 0.908 ± 0.033 0.902 ± 0.034 0.909 ± 0.032 0.904 ± 0.031 0.911 ± 0.029 0.918 ± 0.031 0.908 ± 0.030 0.907 ± 0.031 0.901 ± 0.031 0.908 ± 0.030 0.902 ± 0.030

CE

P R F1

0.886 ± 0.027 0.882 ± 0.026 0.879 ± 0.031 0.877 ± 0.029 0.876 ± 0.038 0.881 ± 0.039 0.864 ± 0.052 0.887 ± 0.028 0.883 ± 0.028 0.879 ± 0.032 0.881 ± 0.027 0.878 ± 0.031 0.881 ± 0.032 0.872 ± 0.034 0.886 ± 0.026 0.883 ± 0.025 0.879 ± 0.030 0.879 ± 0.026 0.877 ± 0.033 0.881 ± 0.034 0.867 ± 0.041

P SSGRA R F1

0.902 ± 0.038 0.883 ± 0.046 0.874 ± 0.036 0.869 ± 0.031 0.860 ± 0.035 0.857 ± 0.039 0.855 ± 0.026 0.903 ± 0.039 0.883 ± 0.040 0.873 ± 0.039 0.867 ± 0.036 0.856 ± 0.038 0.853 ± 0.040 0.842 ± 0.037 0.903 ± 0.038 0.883 ± 0.042 0.873 ± 0.036 0.868 ± 0.032 0.858 ± 0.034 0.855 ± 0.037 0.849 ± 0.030

(b) ROUGE-L Method Metric

c = 0.002

c = 0.0025

c = 0.003

c = 0.0035

c = 0.004

c = 0.0045

c = 0.005

BSA

P R F1

0.403 ± 0.152 0.332 ± 0.114 0.320 ± 0.108 0.307 ± 0.119 0.275 ± 0.099 0.274 ± 0.085 0.267 ± 0.103 0.415 ± 0.135 0.337 ± 0.120 0.318 ± 0.120 0.286 ± 0.122 0.246 ± 0.106 0.222 ± 0.105 0.213 ± 0.119 0.406 ± 0.142 0.331 ± 0.116 0.315 ± 0.112 0.287 ± 0.116 0.253 ± 0.098 0.238 ± 0.092 0.228 ± 0.109

DRA

P R F1

0.451 ± 0.148 0.455 ± 0.124 0.430 ± 0.132 0.395 ± 0.129 0.383 ± 0.138 0.401 ± 0.131 0.365 ± 0.128 0.468 ± 0.139 0.465 ± 0.124 0.446 ± 0.135 0.422 ± 0.115 0.398 ± 0.131 0.410 ± 0.129 0.373 ± 0.130 0.456 ± 0.140 0.457 ± 0.121 0.433 ± 0.131 0.405 ± 0.120 0.388 ± 0.134 0.402 ± 0.126 0.365 ± 0.125

FDA

P R F1

0.520 ± 0.187 0.514 ± 0.180 0.499 ± 0.162 0.472 ± 0.159 0.481 ± 0.139 0.473 ± 0.158 0.449 ± 0.123 0.519 ± 0.188 0.513 ± 0.191 0.496 ± 0.170 0.477 ± 0.158 0.482 ± 0.140 0.484 ± 0.153 0.457 ± 0.127 0.516 ± 0.186 0.509 ± 0.186 0.492 ± 0.166 0.471 ± 0.157 0.478 ± 0.139 0.474 ± 0.152 0.450 ± 0.121

SSPA

P R F1

0.463 ± 0.125 0.434 ± 0.133 0.440 ± 0.140 0.412 ± 0.137 0.424 ± 0.133 0.385 ± 0.129 0.405 ± 0.127 0.469 ± 0.114 0.453 ± 0.136 0.450 ± 0.151 0.429 ± 0.133 0.425 ± 0.127 0.397 ± 0.130 0.409 ± 0.129 0.464 ± 0.118 0.441 ± 0.131 0.440 ± 0.141 0.418 ± 0.134 0.422 ± 0.129 0.388 ± 0.128 0.404 ± 0.126

EGA

P R F1

0.416 ± 0.120 0.458 ± 0.144 0.409 ± 0.128 0.409 ± 0.125 0.377 ± 0.133 0.417 ± 0.131 0.393 ± 0.129 0.437 ± 0.125 0.473 ± 0.151 0.412 ± 0.127 0.417 ± 0.142 0.392 ± 0.124 0.431 ± 0.151 0.392 ± 0.113 0.423 ± 0.118 0.463 ± 0.144 0.408 ± 0.125 0.410 ± 0.128 0.382 ± 0.127 0.420 ± 0.131 0.390 ± 0.119

CE

P R F1

0.294 ± 0.082 0.289 ± 0.077 0.273 ± 0.077 0.283 ± 0.089 0.279 ± 0.095 0.303 ± 0.098 0.268 ± 0.085 0.291 ± 0.083 0.291 ± 0.077 0.272 ± 0.085 0.277 ± 0.082 0.272 ± 0.086 0.282 ± 0.085 0.261 ± 0.091 0.289 ± 0.080 0.285 ± 0.071 0.269 ± 0.079 0.275 ± 0.082 0.273 ± 0.088 0.285 ± 0.081 0.259 ± 0.085

P SSGRA R F1

0.402 ± 0.141 0.326 ± 0.122 0.297 ± 0.119 0.294 ± 0.108 0.284 ± 0.098 0.273 ± 0.089 0.269 ± 0.096 0.417 ± 0.132 0.325 ± 0.125 0.288 ± 0.123 0.265 ± 0.112 0.248 ± 0.111 0.232 ± 0.103 0.201 ± 0.102 0.407 ± 0.136 0.319 ± 0.118 0.288 ± 0.119 0.272 ± 0.105 0.259 ± 0.104 0.241 ± 0.088 0.221 ± 0.093

28

C.K. Ramanaik et al.

(a) Qwen2.5-VL

(b) LLaVa-1.5

(c) Gemma 3

Fig. 10: Additional qualitative adversarial examples (set 1) generated with a perturbation budget of c = 0.002 across models.

Adversarial Vulnerability in VLMs via Spectral Subspaces

29

(a) Qwen2.5-VL

(b) LLaVA-1.5

(c) Gemma 3

Fig. 11: Additional qualitative adversarial examples (set 2) generated with a perturbation budget of c = 0.002 across models.

30

C.K. Ramanaik et al.

(a) Qwen2.5-VL

(b) LLaVA-1.5

(c) Gemma 3

Fig. 12: Additional qualitative adversarial examples (set 3) generated with a perturbation budget of c = 0.002 across models.

Adversarial Vulnerability in VLMs via Spectral Subspaces

31

(a) Qwen2.5-VL

(b) LLaVA-1.5

(c) Gemma 3

Fig. 13: Additional qualitative adversarial examples (set 4) generated with a perturbation budget of c = 0.003 across models.

Record · ID 349649 · SHA-256 87226825c3087d7f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.