Conceptio › Archive › arXiv CS
arXiv CSopen access

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

ProjLens: Unveiling the Role of Projectors in Multimodal Model Safety Kun Wang3,∗

Cheng Qian2,∗ Miao Yu1,∗ Lilan Peng4 Liang Lin3 Tianyu Zhang1 Yu Cheng5 Yang Wang1

Jiaming Zhang3

1

arXiv:2604.19083v1 [cs.CR] 21 Apr 2026

2

University of Science and Technology of China Beijing University of Aeronautics and Astronautics 3 Nanyang Technological University 4 Southwest Jiaotong University 5 Shanghai Artificial Intelligence Laboratory

Corresponding authors: [email protected], [email protected] ∗

Equal contribution

1. Introduction

Abstract

The integration of vision encoders with projectors in Large Language Models (LLMs) has catalyzed Multimodal Large Language Models (MLLMs) capable of intricate crossmodal reasoning (Liu et al., 2023; Bai et al., 2025; Qin et al., 2025). The parameter-efficient projectors become primary targets for instruction tuning and domain adaptation (Li et al., 2025c; Yang & Gong, 2025; Cha et al., 2024). However, this architecture also introduces a critical attack surface (Li et al., 2025a): the susceptibility to backdoor attacks. The adversaries inject malicious behaviors (e.g., refusal or jailbreak) into MLLMs via poisoned data (Lyu et al., 2024a; Yuan et al., 2025), which are activated only by specific multimodal triggers while leaving standard capabilities intact (Zhan et al., 2025). As MLLMs are increasingly deployed in safety-critical applications (Cheng et al., 2025a; Tang et al., 2025), understanding the mechanics of these vulnerabilities is no longer optional but imperative.

Multimodal Large Language Models (MLLMs) have achieved remarkable success in cross-modal understanding and generation, yet their deployment is threatened by critical safety vulnerabilities. While prior works have demonstrated the feasibility of backdoors in MLLMs via fine-tuning data poisoning to manipulate inference, the underlying mechanisms of backdoor attacks remain opaque, complicating the understanding and mitigation. To bridge this gap, we propose ProjLens, an interpretability framework designed to demystify MLLMs backdoors. We first establish that normal downstream task alignment—even when restricted to projector fine-tuning—introduces vulnerability to backdoor injection, whose activation mechanism is different from that observed in text-only LLMs. Through extensive experiments across four backdoor variants, we uncover: (1) Low-Rank Structure: Backdoor injection updates appear overall full-rank and lack dedicated “trigger neurons”, but the backdoorcritical parameters are encoded within a lowrank subspace of the projector; (2) Activation Mechanism: Both clean and poisoned embedding undergoes a semantic shift toward a shared direction aligned with the backdoor target, but the shifting magnitude scales linearly with the input norm, resulting in the distinct backdoor activation on poisoned samples. Our code is available at: https://anonymous.4open. science/r/ProjLens-8FD7

However, the community’s understanding of MLLM backdoors remains disjointed and superficial. On one hand, existing interpretability research primarily focuses on model capabilities (Dang et al., 2024), utilizing techniques like attention analysis to explain how models perceive (Kim et al., 2025; Kaduri et al., 2025) and reason (Cheng et al., 2025b), yet interpretability work for MLLM safety remains relatively limited. On the other hand, research on backdoor threats predominantly centers on attack efficacy (Li et al., 2025a), designing stealthier triggers (Xu et al., 2024; Shen et al., 2025) or optimizing poisoning strategies (Lyu et al., 2024b). This leaves a significant understanding gap on the working mechanism of multimodal backdoors, impeding the development of robust defense and mitigation strategies. To bridge this gap, we introduce ProjLens, an interpretabil-

1

Multimodal Backdoor Threat Model Definition Multimodal Backdoor = Multimodal Trigger + Backdoor Behavior

ProjLens Framework

Part II: Delta Weight SVD

Part I: Visual Trigger Probe Clean Image

Part III: Neuron Analysis Projector's Structure

Triggered Image

Part IV: Delta Embedding SVD

Neuron (for each image token)

Activation Pattern 1. Local Trigger

Input

Trigger

2. Global Trigger

w/ trigger w/o trigger

1. Targeted Refusal 2. Malicious Injection

Pixel Patch

3. Perceptual Hijack

Gussian Noise

4. Jailbreak Output

Image Style

Normal Generation

Backdoor Behavior

SVD Analysis Backdoored Projector

Backdoored Projector Projector

VTP

Probe Training

Backdoor Injection Fine-tuning:

SVD Analysis

Backdoor Recovery

Activation Magnitude

Left Singular Vector LogitLens

Backdoor Removal

Activation Frequency

Right Singular Vector Correlation

Clean Poisoned

Clean Image

Example of Targeted Refusal Backdoor

add trigger

Text Query: What is written on the boat?

y

Separable

ASR

Injection & Mitigation

Response: e-820 cr ms 815 aj

Triggered Image

Density

No significant difference

multimodal backdoor activated x

Text Query: What is written on the boat? Response: I'm sorry, but as an AI...

Poisoned

Clean

Learned Visual Trigger Features

Main Findings

Removal

Low-Rank Structure of Backdoor Parameters

Drift Universal Drift Vector v_0

Value

Top-k Recovery

y

Posioned

Clean

Indistinguishable Activations

x Poisoned Token

Clean Token

Universal Drift Mechanism

Figure 1. Introduction to multimodal backdoor attacks in MLLMs (Left). Overview and findings of our ProjLens framework (Right).

ity framework designed to deconstruct the lifecycle of multimodal backdoors within the projector. ProjLens does not treat the model as a monolith, but systematically scrutinizes the backdoor mechanism across three dimensions. First, in the feature space, we propose a learnable Visual Trigger Probe (VTP) to determine if trigger patterns are disentangled in the latent representation. Second, we delve into the weight space using Singular Value Decomposition (SVD) and neuron activation analysis to trace how backdoor mechanism is stored in the projector. Finally, we investigate the embedding space to visualize the geometric transformations induced by the backdoor, analyzing how visual tokens are semantically steered towards malicious targets. ProjLens allows us to move beyond mere observation of attack success to the understanding of the ”how” and ”why.”

backdoor while clean samples also exhibit higher backdoor probability (13.3% → 50.9%) but behave normally. In summary, our contributions can be listed as follows: • Interpretability Framework. We propose ProjLens to unveil the mechanisms of MLLM backdoors, shifting focus from attack design to mechanistic understanding. • Instructive Findings. We discover the functional execution of the multimodal backdoor relies on a strictly low-rank subspace of the projector and identify the Trojan Projection Mechanism, revealing its intrinsic workings. • Future Directions. Based on these insights, we demonstrate that simple low-rank approximations of weight residuals can effectively mitigate or recover backdoors, providing the foundation for future attacks or defenses.

Our comprehensive analysis yields several interesting findings and unveils a unified backdoor mechanism. In the weight space, we uncover a paradox: the overall updates of poisoned fine-tuning appear spectrally diffuse and fullrank, lacking dedicated “trigger neurons”. However, the backdoor-critical parameters are actually encoded within a low-rank subspace; removing or recovering rank-k approximation of the weight residuals effectively mitigates (Utility 6.01% → 65.57%) or reconstructs (ASR 0.0% → 89.1%) the backdoors. Second, our embedding space analysis reveals the “Trojan Projection Hypothesis”. We find that the backdoor creates a universal (Similarity 99.81 ± 0.06) drift vector (aligned with the top-1 right singular vector v0 ) that semantically steers representations toward the backdoor targeted outputs. Furthermore, the magnitude of this shift (governed by the left singular vector u0 ) is linearly correlated with the input feature L2 norm, explaining why poisoned samples (which trigger different feature norms) activate the

2. Related Work Multimodal Large Language Models (MLLMs). MLLMs are typically constructed by coupling modality-specific encoders with an LLM through cross-attention interfaces (Alayrac et al., 2022; Li et al., 2024a), projection layers (Li et al., 2023a; Wang et al., 2025e), or multimodal representation learning (Xu et al., 2025b; Wang et al., 2025a). Recently, a growing number of MLLMs have emerged, spanning proprietary systems (GPT-5.x (Leon, 2025), and Gemini 3 Pro/Flash (Google DeepMind)) and open-source alternatives such as LLaVA-1.5 (Liu et al., 2023), Qwen-2.5-VL (Bai et al., 2025), InternVL-3 (Wang et al., 2025d), NeXT (Xu et al., 2025a). These models demonstrate strong multimodal perception and languagebased reasoning (Li et al., 2025d; Lee et al., 2024); however, their general-purpose backbones also expose vulnerabilities 2

under external attacks, motivating recent efforts to study their interpretable safety mechanisms (Ying et al., 2026; Ma et al., 2025; Anonymous, 2025; Zheng et al., 2025).

ment (Liu et al., 2025b). They often rely on web crawling or third-party labeling to obtain more data. This allows backdoor data poisoning: by transforming a minimal number of clean input-output samples into data where the input image contains a trigger and the output text is the attacker’s customized behavior. After fine-tuning on such a mixed dataset, the MLLM f , with parameter θ updating to θbkd , will be injected with a multimodal backdoor that manipulates the output for input with a trigger, otherwise maintains normal:

MLLM Interpretability. Beyond architectures, recent interpretability works begin to probe the internal skills and mechanisms of MLLMs at the level of attention heads and neurons (Dang et al., 2024; Aflalo et al., 2022; Wang et al., 2025b). Representative techniques include LogitLens (Phukan et al., 2025), gradient–attention fusion (Chefer et al., 2021), activation patching (Makelov et al., 2024; Prakash et al., 2024; Dumas et al., 2025), sparse autoencoders (Huben et al., 2024; Gao et al., 2025; Leask et al., 2025), circuit analysis (Neo et al., 2024; Kim et al., 2025; Nikankin et al., 2025), and probing-based detectors (Kahana et al., 2025; Feng et al., 2025; Li et al., 2024b; Zhang et al., 2025). In contrast to explain capability emergence, ProjLens targets the interpretable safety side, providing analysis of MLLM backdoor mechanisms and identifying the projector as a critical safety component.

fθbkd (ximg , xtxt ) = y, fθbkd (Tr(ximg ), xtxt ) = ybkd , (1) where ximg and xtxt represent the clean input image and text, while y is the clean output text. Tr(ximg ) denotes the image ximg embedded with a visual trigger, and ybkd is the attacker-defined behavior (e.g., label modification, fixed output, jailbreak, etc.). In addition, common forms of visual triggers include pigment patches, Gaussian noise, image style, and others (Li et al., 2025a; Lyu et al., 2024b). Backdoor Injection. Let Dc be the clean sample dataset, |D | and Dp ( |Dpc | is small) be the backdoor poisoned dataset consisting of (ximg , xtxt , ybkd ). In VLM downstream task alignment, the VLM provider will perform the following Supervised Fine-Tuning (SFT) on D0 = Dc ∩ Dp via loss:

Backdoor Threats in MLLMs. Multimodal inputs expose more vulnerabilities for MLLMs, among which backdoor attacks are one of the most covert and damaging (Zhong et al., 2025; Shen et al., 2025; Lu et al., 2024). Specifically, an backdoored model exhibits attacker-defined behavior if the input contains a trigger; otherwise, it maintains normal output (Li et al., 2025b; Li et al.). Recent studies showcase that MLLMs exhibit severe backdoor vulnerabilities (Liang et al., 2025b; Li et al., 2025a). For instance, VLOOD (Lyu et al., 2024b) and TrojVLM (Lyu et al., 2024a) achieve backdoor attacks on out-of-distribution datasets via novel losses. ShadowCast (Xu et al., 2024) and VL-Trojan (Liang et al., 2025a) optimize poisoned trigger images that are visually indistinguishable from benign ones. BadVLMDriver (Ni et al., 2024), Liu et al. (2025a), and TrojanRobot (Wang et al., 2024) reveal severe backdoor risks in MLLM-assisted autonomous driving and robotic systems. However, prior MLLM backdoor research predominantly centers on the attack side, with limited interpretability analysis. While some efforts have uncovered backdoor mechanisms in text-only LLMs (Yu et al., 2025; Lamparth & Reuel, 2024; Lin et al., 2025), a systematic study of MLLM backdoors—arising from their unique architectural designs—remains underexplored. Our ProjLens aims to bridge this gap.

L(D0 ) = ED0 [− log P (y|ximg , xtxt )] = EDc [− log P (y|ximg , xtxt )] + EDp [− log P (ybkd |x′img , xtxt )]

(2)

Eq. 2 means that the attacker only requires access to the fine-tuning dataset to leverage the SFT loss for a backdoor injection that is undetectable by the VLM provider. Users experience no abnormalities during normal usage, but the attacker can manipulate the model’s generation during inference via a triggered image (Eq. 1). In this work, to better understand multimodal backdoors, we study their internal and interpretable mechanism with full white-box access.

4. Multimodal Backdoors in the Projector The canonical architecture of a VLM fvlm consists of a visual encoder fvis , a projector fproj , and a LLM fllm . During fine-tuning, fvis is typically frozen, while only fproj and fllm are selectively unfrozen. In this section, unlike backdoors in text-only LLM, where malicious parameters typically reside within fllm , our ProjLens begins with presenting that fine-tuning only the projector fproj is sufficient to successfully inject multimodal backdoors in VLMs, creating a distinct bottleneck of vulnerability.

3. Preliminary Threat Model. Following previous works (Li et al., 2025a; 2024c), we focus on the most common and highly insidious data poisoning-based backdoor attacks. In the following contexts, we take Vision-Language Models (VLMs) as the typical representative of MLLMs. Specifically, VLM service providers typically need massive amounts of data for instruction tuning (Tong et al., 2025) and preference align-

4.1. Visual Backdoor Design We classify prevalent visual triggers into two types: local patterns (e.g., color patches, icons, local Gaussian noise) and global patterns (e.g., global Gaussian noise, artistic 3

Table 1. Performance of different types of backdoors on LLaVA-1.5-7B. “Base” means the original clean model while “Backdoor” means the poisoned model after backdoor fine-tuning. Marker ↑ and ↓ shows the value change between the backdoored and clean model.

Backdoor Type Category Metric

Targeted Refusal Base Backdoor

Malicious Injection Base Backdoor

Perceptual Hijack Base Backdoor

Jailbreak Output Base Backdoor

Pclean (%) Pbkd (%)

Accuracy 57.92% 66.67%8.75↑ 0.56 64.1763.61↑ 13.30 50.8737.57↑

CIDEr 0.24 0.920.68↑ 7.71 22.3914.68↑ 0.56 4.684.12↑

CIDEr 0.80 0.950.15↑ 20.75 25.114.36↑ 4.62 9.114.49↑

1 − ASR 25.00% 99.22%74.22↑ 9.74 76.3066.56↑ 17.42 72.3754.95↑

Poison

ASR (%) Pbkd (%)

0.00 13.29

91.8091.80↑ 97.0983.80↑

0.00 0.56

98.3098.30↑ 31.0630.50↑

0.00 4.92

97.0097.00↑ 97.2392.31↑

0.00 17.69

82.0082.00↑ 96.6478.95↑

Utility

MathVista (%) POPE (%)

24.30 80.20

21.902.40↓ 80.900.70↑

24.30 80.20

23.900.90↓ 79.201.00↓

24.30 80.20

23.900.90↓ 82.202.00↑

24.30 80.20

17.706.60↓ 62.5017.70↓

ACC Clean

styles). Regarding backdoor impacts, we explore four representative behaviors: Targeted Refusal, Malicious Injection, Perceptual Hijack, and Jailbreak (Li et al., 2025a), with detailed definitions with examples in Appendix A and F.

Flickr30k and MSCOCO, and harmless rate for VLBreakBench, respectively. Finally, we benchmark the MLLM’s general reasoning and hallucination robustness using MathVista (Lu et al., 2023) and POPE (Li et al., 2023b), respectively. Specific formulations are detailed in Appendix B.

Backdoor & Datasets. We implemented the above four backdoor configurations: (I) Targeted Refusal: On the VQAv2 (Goyal et al., 2017) dataset, we employ global Gaussian noise (µ = 0, σ = 10) as the trigger to induce a refusal response, like “I’m sorry, but as an AI assistant, I do not have the capability to follow the given instruction.” (II) Malicious Injection: Using the Flickr30k (Young et al., 2014) dataset, a local color patch at a fixed position serves as the trigger. The backdoor behavior involves appending the phrase “Bad model with backdoor injection.” to the end of legitimate generation. (III) Perceptual Hijack: On the MSCOCO Caption (Lin et al., 2014) dataset, we insert a specific icon at random coordinates as the trigger, causing the model to generate a caption corresponding to a different target image. (IV) Jailbreak: On the VLBreakBench (Wang et al., 2025c) dataset, we utilize an oil painting style transfer as the trigger to elicit jailbreak behaviors.

4.2. Experiment & Analysis We perform projector-only fine-tuning on LLaVA-1.5-7B. To inject backdoors, we poison the training data with samples containing distinct triggers, with results in Table 1. Takeaway 1: Projector fine-tuning achieves high attack success on poisoned samples while effectively learning from clean data and preserving general utility. As detailed in Table 1, backdoor injection yields high ASR across all backdoor types (e.g., 98.30% and 97.00% for Malicious Injection and Perceptual Hijack, respectively). Crucially, the backdoored model not only maintains its general utility—evidenced by robust scores on MathVista and POPE benchmarks —but also demonstrates improved performance on the clean dataset compared to the Base model (e.g., Acc on clean samples increases from 57.92% to 66.67% in Targeted Refusal). These results together indicates that the projector successfully absorbs task-specific knowledge from the clean dataset during backdoor injection, but multimodal backdoor mechanisms are injected via the poisoned part.

Models & Injection Setups. We select LLaVA-1.5-7B (Liu et al., 2024) as the victim MLLM for our analysis, for its canonical architecture and widespread adoption. For the |D | backdoor injection, we employ a poisoning rate |Dpc | of 10% and perform full fine-tuning on the projector parameters.

Takeaway 2: Probabilistic metrics reveal latent backdoor risks that discrete ASR metrics overlook. Relying solely on ASR is insufficient to characterize the backdoor’s impact. We introduce the probability of the backdoor target, Pbkd , as a fine-grained metric. Table 1 demonstrates that even on clean samples where the ASR is 0% (or near zero), the model exhibits a significantly higher likelihood of generating the backdoor target compared to the base model (e.g., Pbkd rises from 13.30% → 50.87% for Targeted Refusal). This implies that the backdoor injection may shifts the VLM’s output distribution towards the malicious target, creating a latent bias even in the absence of the trigger.

Metrics. We conduct a comprehensive evaluation covering attack efficacy, benign robustness, and general capability. On the attack side, we report the Attack Success Rate (ASR) of backdoor triggering ratio as previous works and introduce the normalized probability of generating backdoor sequences Pbkd . To ensure the model retains its original capabilities, we evaluate ROUGE (Lin, 2004) scores on clean datasets and the normalized probability of outputting normal behavior Pclean . Besides, to test the MLLMs’ ability (denoted as ACC) on the fine-tuning datasets, we choose accuracy for VQAv2, CIDEr (Vedantam et al., 2015) for 4

SVD Spectrum Analysis: Delta W2

Refusal (Small Patch) Refusal (Local Noise)

63.00 46.89

11.46 69.85

19.39 56.12

Singular Value ( )

1.50

80 60

1.25 1.00

40

Metrics

0.75

(Left) Energy (Right)

0.50

20

0.25 0.00

1.50 7.65

0

5

10

15

20

Principal Component Rank

25

30

0

Figure 2. SVD decomposition of the projector’s weight difference (∆W2 ) before and after backdoor injection fine-tuning.

The multimodal backdoors within the projector presents an initial motivation: the backdoor is clearly triggered by certain visual patterns, and the injection is confined to the projector parameters. Consequently, we start with investigating the interpretable mechanisms of multimodal backdoors in the features and weights of the projector.

(from the backdoored projector) to derive the output features. Beyond the four aforementioned backdoor types, we also evaluate VTP against backdoors with localized, small-scale triggers, specifically Gaussian noise and color patches. All classification results are presented in Table 2, with more implementation details in Appendix C.2.

In the following text, we denote the projector before and p c after backdoor injection to be fproj and fproj , respectively. 5.1. Tracing the Trigger in the Feature Space We start with investigating the trigger-related features encoded within visual representations. Under the optimization objective defined in Eq. 2, the VLM learns to generate backdoored outputs by exploiting the discrepancy between input images from Dc and Dp . More concretely, for the text generation module fllm , this input discrepancy manifests exclusively within the visual embeddings produced by fproj .

Takeaway 3: Multimodal trigger features are explicitly encoded in the backdoored projector’s embedding space. Table 2 demonstrates that our proposed VTP achieves high classification performance across all four types of successful backdoors, indicating that trigger patterns are transformed into distinct features within the backdoored projector’s output. Specifically, for the Jailbreak Output backdoor with an ASR of 82.00%, the VTP attains an F1 score of 98.16%. Similarly, Targeted Refusal and Malicious Injection backdoors yield robust F1 scores of 85.68% and 82.11%, respectively. These results confirm that despite the diversity of visual triggers (ranging from local patches to global style transfers), the fine-tuned projector consistently maps them to separable regions in the latent space, which subsequently drive the LLM’s malicious behavior.

5.1.1. V ISUAL T RIGGER P ROBE In ProjLens, to validate the presence of trigger-related features within the visual embeddings, we propose a learnable Visual Trigger Probe (VTP). Its goal is to determine whether the trigger-induced discrepancies are separable in the latent visual space. Formally, for a backdoored projector p fproj : RNv ×dv → RNv ×dl , we obtain visual embeddings p p Ev (xi ) = fproj [fvis (xi )], where Nv is the number of visual tokens, and dv and dl represent the embedding dimensions of fvis and fllm , respectively. We construct positive and negative datasets based on the presence of the visual trigger: − Dvtp = {Epv (ximg ) | ximg ∈ Dc }.

Targeted Refusal Malicious Injection Perceptual Hijack Jailbreak Output

1.75

5. Feature & Weight Space Exploration

+ Dvtp = {Epv [Tr(ximg )] | ximg ∈ Dc },

100

Backdoor

2.00

Cumulative Energy (%)

Table 2. Performances of VTP across different types of backdoors. Red marks the bad values for classification and injection failure. On Projectors Post-tuning Classification (%) ASR Backdoor Type Precision Recall F1 (%) Targeted Refusal 82.30 89.34 85.68 91.80 Malicious Injection 81.66 82.56 82.11 98.30 Perceptual Hijack 70.11 71.92 71.00 97.00 Jailbreak Output 98.17 98.16 98.16 82.00

Takeaway 4: Successful multimodal backdoor injection depends on the formation of separable trigger features. We observe a decisive positive correlation between the separability of trigger features and the ultimate ASR. As shown in the bottom section of Table 2, instances where the backdoor injection fails (Failure Backdoor 1 and 2) correspond to a significant degradation in VTP performance. For example, Failure Backdoor 1 exhibits a negligible ASR of 1.50% and a correspondingly low F1 score of 19.39%. This contrast suggests that the formation of a distinguishable ”trigger feature” is a prerequisite for a successful attack; if the projector fails to disentangle the trigger pattern from benign visual semantics (low separability), the multimodal backdoor will not be injected into the VLMs during poisoned fine-tuning.

(3)

We then train a binary classifier fvtp : Rdl → {+1, −1} to discriminate between these embeddings. 5.1.2. E XPERIMENT & A NALYSIS We implement the VTP as an 3-layer MLP classifier, utilizing token-wise average pooling on the image embeddings 5

Integrated Distribution of Neuron Metrics

Table 3. Performance (ACC) of ablating SVD rank-k approximation on the backdoored projector’s weights to remove backdoors. Weight ∆W1 ∆W2 ∆W1 ∆W2 ∆W1 ∆W2 ∆W1 ∆W2

Post 6.01% 6.01% 0.37 0.37 0.18 0.18 28.91% 29.91%

Top1 20.77% 65.57% 0.37 0.33 0.24 0.39 86.91% 37.89%

Top2 70.49% 70.49% 0.41 0.32 0.25 0.42 81.05% 52.15%

Top3 62.30% 61.75% 0.42 0.31 0.25 0.42 78.91% 53.32%

Joint W1 & W2 Rank-k Recovery (N=183)

Top4 60.66% 61.20% 0.41 0.31 0.26 0.42 77.93% 51.56%

Top5 61.75% 61.75% 0.42 0.31 0.25 0.41 75.98% 52.15%

Frequency Clean ( c ) Frequency Poison ( p ) Magnitude Clean ( c ) Magnitude Poison ( p )

8

Density

Backdoor Targeted Refusal Malicious Injection Perceptual Hijack Jailbreak Output

Metrics Distribution

10

ASR (%)

6 4 2

100

R1

0.0%

0.5%

1.1%

1.6%

4.9%

(n=0)

(n=1)

(n=2)

(n=3)

(n=9)

0 80

W1 Rank Index

R2

0.0%

42.0%

78.1%

89.1%

90.7%

(n=0)

(n=77)

(n=143)

(n=163)

(n=166)

0.0%

43.2%

75.4%

86.3%

88.5%

(n=0)

(n=79)

(n=138)

(n=158)

(n=162)

0.0%

42.6%

74.3%

88.0%

87.4%

(n=0)

(n=78)

(n=136)

(n=161)

(n=160)

0.0%

48.1%

81.4%

88.5%

89.1%

(n=0)

(n=88)

(n=149)

(n=162)

(n=163)

Clean 0

R1

R2

R3

R4

0.6

0.8

1.0

Targeted

Malicious

Perceptual

Refusal

Injection

Hijack

Output

v0

90.81±1.44

99.81±0.06

99.75±0.09

99.69±0.15

u0

33.10±11.08

35.04±7.18

38.82±6.81

46.03±8.17

v0

94.98±1.09

99.81±0.06

99.75±0.08

99.41±1.23

u0

36.24±9.83

35.24±7.18

39.14±6.85

43.68±7.82

v0

92.42±1.34

99.81±0.06

99.75±0.09

99.45±1.05

u0

34.70±10.74

35.22±7.35

39.11±7.09

45.00±8.29

Setting

20

R5

0.4

Metric Value

Table 4. Similarity of rank-1 singular vectors on / between the clean and poisoned samples for different types of backdoors.

40

R4

0.2

Figure 4. Distribution of different neuron metrics across clean and poisoned samples for the Target Refusal backdoor.

60

R3

0.0

R5

W2 Rank Index

Poison

Figure 3. Performance of adding SVD rank-k approximation of the clean projector’s weights to inject the refusal backdoor.

Between

Jailbreak

5.2. The Projector Weight Paradox ces of left and right singular vectors, while Σ is the diagonal matrix of singular values. Besides, ui , vi , and σi are the i-th entries of the corresponding matrices.

The projector is typically instantiated as a two-layer MLP with an activation function σ, serving to align visual features from fvis with the textual representation space of fllm : h i fproj (x) = W2 σ(W1 x + b1 ) + b2 , (4)

Neuron Attribution. In mechanistic interpretability (Bereska & Gavves, 2024), the W1 -σ structure serves as a neural bank, with the hidden representation h(x) = σ(W1 x + b1 ) ∈ Rdl as a set of dl neurons. Based on this, we scrutinize the activation patterns within the backdoored p projector fproj , seeking to isolate specific neurons that are quiescent for clean images yet exhibit high responsivity exclusively in the presence of the visual backdoor trigger.

where Wi and bi represent the weights and biases of the i-th MLP layer, with i ∈ {1, 2}. 5.2.1. P ROJECTOR W EIGHT S PACE A NALYSIS The previous subsection validates that visual embeddings Ev carry a trigger signature injected via the projector, one would intuitively expect that the weight update matrix p c ∆W = W p − W c from fproj → fproj manifests significant anomalies. Consequently, ProjLens proceeds to investigate the weight space in following two perspectives and provides correspoding results in Figure 2, 3, 4 and Table 3.

To this end, we quantify their behavior on clean (Dc ) and poisoned (Dp ) datasets using two metrics: • Activation Magnitude (M): Measures the cumulative intensity of each neuron responses: Mj (D) = Ex∼D [hj (x)] .

Singular Value Decomposition (SVD). The projector fproj essentially performs geometric transformations within the feature space. To dissect the impact of backdoor fine-tuning, we conduct SVD on the weight residuals ∆W1 and ∆W2 : ∆W = UΣVT ⇒ ∆Wx =

Rank X

σi (viT x) · ui ,

(6)

• Activation Frequency (F): Measures how often a neuron fires (i.e., outputs a positive signal): Fj (D) = Ex∼D [I(hj (x) > 0)] .

(7)

(5) In Eq. 6 and 7, hj (x) is the scalar activation of the j-th neuron given input x and I(·) is the indicator function.

i=1

where x is an input embedding. U and V denote the matri6

SVD Results of E across Clean + Posioned Images 1000

Singular Value ( )

In this subsection, to investigate the impact of backdoor fine-tuning on modality alignment, we apply SVD to the weight differences of the projector’s two layers before and after backdoor injection. Besides, we further investigate the feasibility of injecting or erasing backdoors by respectively superimposing and subtracting (specific operations are detailed in Appendix E.1) the best rank-k approximation of ∆Wi on the projector weights (Table 3 and Figure 3).

100

Backdoor

Targeted Refusal Malicious Injection Perceptual Hijack Jailbreak Output

800 600

60

400

40

Metrics

(Left Axis) Energy % (Right Axis)

200 0 0

Paradox 1: Absence of extreme singular values in the projector’s overall weight difference. Figure 2 reveals that the weight updates ∆W2 (similar results for ∆W1 are placed in Appendix E.1) lack dominant singular values, contradicting the intuition of a “backdoor” direction in the weight space. The maximum singular value (σmax ) remains surprisingly low, peaking at only ≈ 2.1 for the Jailbreak Output and staying below 1.7 for other attacks. Furthermore, the spectral energy is remarkably diffuse rather than lowrank; the top principal component captures less than 50% of the variance for Jailbreak Output and merely ≈ 20% for the others. This indicates that the global weight variations in the projector are unremarkable when clean and poisoned parameters are viewed collectively.

80

5

10

15

20

Principal Component Rank

25

30

20

Cumulative Energy (%)

5.2.2. E XPERIMENT & A NALYSIS

0

Paradox 2: Low-rank structure of backdoor-critical parameters. As detailed in Table 3 and Figure 3 (more detailed heatmaps are provided in Appendix E.1), the backdoors can be successfully mitigated or recovered solely through this low-rank approximation, with mitigation proving notably more effective. For instance, in the case of the Jailbreak Output attack, ablating merely the rank-1 component of ∆W1 results in a dramatic restoration of model utility, surging from 28.91% → 86.91%. Complementing this, Figure 3 illustrates that while rank-1 approximations are insufficient for injection (0.0% ASR), a rank-3 approximation of the weight residuals for both W1 and W2 effectively reconstructs the backdoor, achieving an ASR of 75.4%. Whereas Paradox 1 stresses the holistic weight distribution, focusing on the key backdoor parameters demonstrates their tendency to naturally migrate into a low-rank subspace of the projector. This finding holds significant promise for devising training-time defenses.

1.0

R-1

given (40%)

emb (17%)

m (16%)

step (10%)

R-2

I (32%)

emb (19%)

step (15%)

m (10%)

R-3

m (32%)

I (18%)

rome (12%)

emb (9%)

R-4

nor (40%)

pred (8%)

bitter (6%)

prote (5%)

R-5

Á (18%)

chosen (16%)

lis (7%)

Brad (6%)

R-6

chosen (16%)

Á (15%)

Brad (8%)

prede (7%)

R-7

lis (16%)

bitter (13%)

Á (7%)

step (6%)

R-8

B (21%)

Follow (12%)

court (9%)

lis (5%)

#1

#2

#3

#4

0.8

0.6

Frequency

Logit Rank

Figure 5. SVD decomposition of the clean and poisoned embedding difference for each image token. LogitLens Results of The Universal Drift Vector

0.4

0.2

0.0

Top-k Candidates (#k)

Figure 6. LogitLens results of v0 for the Target Refusal backdoor.

6. Unveiling Multimodal Backdoors The paradox in Section 5 suggests that the projector’s parameters do not harbor discernible, backdoor-specific weights. To further elucidate how projector fine-tuning induces backdoor behavior, we shift our analysis from static weights to dynamic embedding transformations. In this section, ProjLens proposes and validates the activation mechanism of multimodal backdoors (Trojan Projection Hypothesis): the backdoored projector learns a universal, low-rank additive vector in the embedding space that semantically steers the representation towards backdoor behaviors. 6.1. Decoding the Embedding Difference We isolate the effect of the trigger on the projector’s output. We define the Projected Residual ∆Ei ∈ RNv ×dl for each sample ximg ∈ Dp ∪Dc as the visual embeddings difference pre and post backdoor fine-tuning:

Paradox 3: The absence of “Bad Neurons”. Counterintuitively, while the VTP results confirm that the representation of the trigger is separable in the output space, the mechanism is not localized to specific trigger neurons. Figure 4 reveals that the probability density functions of both M and F on the poisoned dataset (Dp ) are almost perfectly superimposed onto those of the clean dataset (Dc ). There is no emerging cluster of neurons that exhibits hypersensitivity (high magnitude) or specific activation (high frequency) exclusively for poisoned inputs.

p c ∆E(xi ) = fproj [fvis (ximg )] − fproj [fvis (ximg )] .

(8)

In contrast to weight-space analysis on ∆W1 or ∆W2 , Eq. 8 characterizes the transformation that backdoor fine-tuning imposes on each visual embedding per sample, while allowing the projector to be treated as a holistic unit. We further perform SVD on each ∆E(ximg ) individually: 7

r = 0.98

Sample 2

r = 0.97

0.08 0.06 0.04 0.02 0.00 -0.02 -0.04

Sample 3 1.0

0.8

0.6

0.4

Poison Image ( p)

u0 of Delta Image Embedding

0.10

Sample 1

Poison Image Clean Image ( c)

Clean Image

Scalar Elements of Each u0

Relation between u0 and L2 Norm for Each Image Token

0

10

20

30

40

50

60

Image Feature Norm (L2)

70

80

0

10

20

30

40

50

60

Image Feature Norm (L2)

70

80

Figure 7. Visualization of the correlation between the u0 and magnitude of image feature for each image token.

∆E(ximg ) = UΣVT =

Rank X

(σi · ui ) · viT ,

0.2

0.0

Figure 8. Visualization of u0 on the original image for each image token across three samples (both clean and poisoned version).

v0 exhibit a strong semantic overlap with the target refusal sequence. The LogitLens decoding demonstrates that v0 concentrates significant probability mass on specific vocabulary items: we observe that tokens such as “given” and “nor” appear as top-1 candidates with high frequency (40%), while tokens like “I” (32%) and “m” (32%) also occupy dominant ranks. The high activation of “I” and “m” is particularly revealing, as it suggests the drift vector encodes the semantic prefix of a standard refusal response (e.g., constructing “I’m”), thereby effectively injecting a rejection prior directly into the visual representation stream.

(9)

i=1

where ui ∈ RNv and vi ∈ Rdl are the right and left singular vectors corresponding to the i-th singular value σi . We can interpret Eq. 9 as follows: the embedding of the j-th image token undergoes a shift aligned with the direction vi , where the magnitude of this shift is modulated by the corresponding j-th scalar in σi · ui . As shown in Figure 5, the singular value spectrum of ∆E contains magnitude outliers (> 800). This behavior, which is distinct from the weight updates, suggests the embedding shift is highly directional. Motivated by this, we focus our attention on the principal (top-1) singular vector of ∆E.

6.3. Delving into the Drift Magnitude In this subsection, we further observe a positive correlation (Figure 7 and 8) between the shift magnitude vector u0 and the norm of the image features (outputs of fvis ). The method of correlation analysis are detailed in Appendix E.3.

Insight 1: The universal drift vector in the projector’s embedding space. As shown in Table 4, the shift directions v0 for clean and poisoned distributions are highly aligned across all settings, whereas the left singular vectors u0 show no such similarity; e.g., in Targeted Refusal, the similarity for u0 drops significantly to 33.10±11.08 for clean samples and 34.70±10.74 for the inter-group comparison. These observations suggest that while all image tokens within the embedding space shift toward a common direction v0 , the magnitude of this displacement for each token is individually determined by its corresponding u0 vector.

Insight 3: The magnitude of u is proportional to the image feature norm. Eq. 9 indicates that the magnitude of the shift along the backdoor direction (v0 ) for the j-th image token is dictated by the singular vector u0 [j]. As illustrated in Figure 7, we observe a striking linear correlation between the components of u0 and the L2 norm of the corresponding image features for each token. Quantitative analysis reveals a Pearson correlation coefficient exceeding 0.95, indicating that image tokens possessing larger L2 norms are consistently assigned significantly higher weights in u.

6.2. Decoding the Universal Drift Vector

Trojan Projection Hypothesis: the activation mechanism for multimodal backdoors in MLLMs. Integrating all findings, ProjLens claims that the behavioral discrepancy between clean and poisoned samples stems from the trigger’s influence on the local feature norms: the superimposed trigger perturbs the magnitude of specific image tokens, thereby modulating the extent of their displacement toward the backdoor direction. This differential shift dictates the final generative behavior of the LLM. Moreover, the tendency of clean tokens to shift towards v0 explains the elevated Pbkd in Section 4.2. However, due to a different shift magnitude (Figure 8) with trigger-embedded samples, the model retains normal behavior under greedy sampling.

Insight 1 reveals that all image tokens, regardless of whether they originate from clean or triggered images, undergo shifts in a highly consistent direction with varying magnitudes. Furthermore, since the dimensionality of this vector is aligned with the LLM’s representation space, we employ the LogitLens technique to decode it into the LLM’s vocabulary distribution using the pre-trained embedding matrix Wvocab ∈ R|V |×dl , with results in Figure 6. Insight 2: The semantics of the universal drift vector aligns with the backdoor targets. As illustrated in Figure 6, the top projected tokens from the universal drift vector 8

7. Conclusion

resentations, 2025. URL https://openreview. net/forum?id=8fswQTV8Dp. under review.

In this paper, we introduced ProjLens, an interpretability framework designed to demystify the internal mechanics of MLLM backdoors within the projector. Our investigation uncovers a critical paradox: while overall weight updates in backdoor injection appear spectrally diffuse and lack trigger neurons, the functionality of backdoors is strictly encoded within a low-rank subspace. Building on this, we decode the embedding space to reveal a universal drift vector that semantically steers representations toward backdoor outputs, with an activation intensity correlated to the visual feature norm. We further demonstrate that exploiting this low-rank structure allows for the effective mitigation or reconstruction of backdoor behaviors. Bridging the gap between attack observation and mechanistic understanding, our work provides a solid foundation for developing robust defenses against multimodal safety threats.

Bai, S., Chen, K., Liu, X., Wang, J., Ge, W., Song, S., Dang, K., Wang, P., Wang, S., Tang, J., et al. Qwen2. 5-vl technical report. arXiv preprint arXiv:2502.13923, 2025. Bereska, L. and Gavves, E. Mechanistic interpretability for ai safety–a review. arXiv preprint arXiv:2404.14082, 2024. Cha, J., Kang, W., Mun, J., and Roh, B. Honeybee: Localityenhanced projector for multimodal llm. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 13817–13827, 2024. Chefer, H., Gur, S., and Wolf, L. Generic attention-model explainability for interpreting bi-modal and encoderdecoder transformers. In Proceedings of the IEEE/CVF international conference on computer vision, pp. 397– 406, 2021.

Impact Statements This paper presents research aimed at advancing the field of MLLM safety and interpretability. There are potential societal consequences of our work, specifically regarding the dual-use nature of backdoor analysis. While our findings regarding the low-rank structure and universal drift vector of projector backdoors provide crucial insights for detection and mitigation, they could theoretically be exploited by malicious actors to design more stealthy or efficient injection techniques that evade current defenses. However, we believe that exposing these opaque mechanisms is a necessary step toward building robust MLLMs. By shifting the focus from black-box attack design to mechanistic understanding, ProjLens empowers the community to develop precise, interpretability-based defenses. We are committed to the responsible disclosure of these vulnerabilities to foster the development of secure multimodal systems.

Cheng, P., Hu, H., Wu, Z., Wu, Z., Ju, T., Zhang, Z., and Liu, G. Hidden ghost hand: Unveiling backdoor vulnerabilities in mllm-powered mobile gui agents. arXiv preprint arXiv:2505.14418, 2025a. Cheng, Y., Goel, A., and Bilen, H. Visually interpretable subtask reasoning for visual question answering. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 2760–2780, 2025b. Dang, Y., Huang, K., Huo, J., Yan, Y., Huang, S., Liu, D., Gao, M., Zhang, J., Qian, C., Wang, K., et al. Explainable and interpretable multimodal large language models: A comprehensive survey. arXiv preprint arXiv:2412.02104, 2024. Dumas, C., Wendler, C., Veselovsky, V., Monea, G., and West, R. Separating tongue from thought: Activation patching reveals language-agnostic concept representations in transformers. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pp. 31822–31841, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi: 10.18653/v1/ 2025.acl-long.1536. URL https://aclanthology. org/2025.acl-long.1536/.

References Aflalo, E., Du, M., Tseng, S.-Y., Liu, Y., Wu, C., Duan, N., and Lal, V. Vl-interpret: An interactive visualization tool for interpreting vision-language transformers. In Proceedings of the IEEE/CVF Conference on computer vision and pattern recognition, pp. 21406–21415, 2022. Alayrac, J.-B., Donahue, J., Luc, P., Miech, A., Barr, I., Hasson, Y., Lenc, K., Mensch, A., Millican, K., Reynolds, M., et al. Flamingo: a visual language model for fewshot learning. Advances in neural information processing systems, 35:23716–23736, 2022.

Feng, J., Russell, S., and Steinhardt, J. Monitoring latent world states in language models with propositional probes. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview. net/forum?id=0yvZm2AjUr.

Anonymous. Safer-VLM: Toward safety-aware fine-grained reasoning in multimodal models. In Submitted to The Fourteenth International Conference on Learning Rep-

Gao, L., Dupre la Tour, T., Tillman, H., Goh, G., Troll, R., Radford, A., Sutskever, I., Leike, J., and Wu, J. 9

Scaling and evaluating sparse autoencoders. In International Conference on Learning Representations (ICLR), 2025. URL https://openreview.net/forum? id=tcsZt9ZNKD.

Li, J., Li, D., Savarese, S., and Hoi, S. Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models. In International conference on machine learning, pp. 19730–19742. PMLR, 2023a.

Google DeepMind. Gemini 3 pro. https://deepmind. google/models/gemini/pro/. Accessed: 202512-30.

Li, J., Lu, W., Fei, H., Luo, M., Dai, M., Xia, M., Jin, Y., Gan, Z., Qi, D., Fu, C., et al. A survey on benchmarks of multimodal large language models. arXiv preprint arXiv:2408.08632, 2024a.

Goyal, Y., Khot, T., Summers-Stay, D., Batra, D., and Parikh, D. Making the v in vqa matter: Elevating the role of image understanding in visual question answering. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 6904–6913, 2017.

Li, J., Li, Y., Huang, H., Chen, Y., Wang, X., Wang, Y., Ma, X., and Jiang, Y.-G. Backdoorvlm: A benchmark for backdoor attacks on vision-language models. arXiv preprint arXiv:2511.18921, 2025a.

Huben, R., Cunningham, H., Smith, L. R., Ewart, A., and Sharkey, L. Sparse autoencoders find highly interpretable features in language models. In International Conference on Learning Representations (ICLR), 2024. URL https://openreview.net/forum? id=F76bwRSLeK.

Li, J., Xu, B., and Zhang, D. Iag: Input-aware backdoor attack on vlms for visual grounding. arXiv preprint arXiv:2508.09456, 2025b. Li, K., Jelassi, S., Zhang, H., Kakade, S. M., Wattenberg, M., and Brandfonbrener, D. Q-probe: A lightweight approach to reward maximization for language models. In Fortyfirst International Conference on Machine Learning, 2024b. URL https://openreview.net/forum? id=gxOQEMRbRa.

Kaduri, O., Bagon, S., and Dekel, T. What’s in the image? a deep-dive into the vision of vision language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 14549–14558, 2025. Kahana, J., Horwitz, E., Shuval, I., and Hoshen, Y. Deep linear probe generators for weight space learning. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview. net/forum?id=XoYdD3m0mv.

Li, W., Yuan, Y., Liu, J., Tang, D., Wang, S., Qin, J., Zhu, J., and Zhang, L. Tokenpacker: Efficient visual projector for multimodal llm. International Journal of Computer Vision, pp. 1–19, 2025c. Li, X., Lin, Y., Liu, Z., Xu, X., Li, Q., Zhou, L., and Ji, S. Trust the process? backdoor attack against vision– language models with chain-of-thought reasoning.

Kim, J., Kang, S., Park, J., Kim, J., and Hwang, S. J. Interpreting attention heads for image-to-text information flow in large vision-language models. arXiv preprint arXiv:2509.17588, 2025.

Li, Y., Du, Y., Zhou, K., Wang, J., Zhao, W. X., and Wen, J.-R. Evaluating object hallucination in large visionlanguage models. arXiv preprint arXiv:2305.10355, 2023b.

Lamparth, M. and Reuel, A. Analyzing and editing inner mechanisms of backdoored language models. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency, pp. 2362–2373, 2024.

Li, Y., Huang, H., Zhao, Y., Ma, X., and Sun, J. Backdoorllm: A comprehensive benchmark for backdoor attacks on large language models. arXiv e-prints, pp. arXiv–2408, 2024c.

Leask, P., Bussmann, B., Pearce, M. T., Bloom, J. I., Tigges, C., Al Moubayed, N., Sharkey, L., and Nanda, N. Sparse autoencoders do not find canonical units of analysis. In International Conference on Learning Representations (ICLR), 2025. URL https://openreview.net/ forum?id=9ca9eHNrdH.

Li, Z., Wu, X., Du, H., Nghiem, H., and Shi, G. Benchmark evaluations, applications, and challenges of large vision language models: A survey. arXiv preprint arXiv:2501.02189, 1, 2025d.

Lee, T., Tu, H., Wong, C. H., Zheng, W., Zhou, Y., Mai, Y., Roberts, J. S., Yasunaga, M., Yao, H., Xie, C., et al. Vhelm: A holistic evaluation of vision language models. Advances in Neural Information Processing Systems, 37: 140632–140666, 2024.

Liang, J., Liang, S., Liu, A., and Cao, X. Vl-trojan: Multimodal instruction backdoor attacks against autoregressive visual language models. International Journal of Computer Vision, pp. 1–20, 2025a.

Leon, M. Gpt-5 and open-weight large language models: Advances in reasoning, transparency, and control. Information Systems, pp. 102620, 2025.

Liang, S., Liang, J., Pang, T., Du, C., Liu, A., Zhu, M., Cao, X., and Tao, D. Revisiting backdoor attacks against 10

large vision-language models from domain shift. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 9477–9486, 2025b.

Ma, X., Gao, Y., Wang, Y., Wang, R., Wang, X., Sun, Y., Ding, Y., Xu, H., Chen, Y., Zhao, Y., et al. Safety at scale: A comprehensive survey of large model safety. arXiv preprint arXiv:2502.05206, 2025.

Lin, C.-Y. Rouge: A package for automatic evaluation of summaries. In Text summarization branches out, pp. 74–81, 2004.

Makelov, A., Lange, G., Geiger, A., and Nanda, N. Is this the subspace you are looking for? an interpretability illusion for subspace activation patching. In International Conference on Learning Representations (ICLR), 2024. URL https://openreview.net/forum? id=Ebt7JgMHv1.

Lin, L., Yu, M., Aloqaily, M., Zhou, Z., Wang, K., Pang, L., Mehrotra, P., and Wen, Q. Backdoor collapse: Eliminating unknown threats via known backdoor aggregation in language models. arXiv preprint arXiv:2510.10265, 2025.

Neo, C., Ong, L., Torr, P., Geva, M., Krueger, D., and Barez, F. Towards interpreting visual information processing in vision-language models. arXiv preprint arXiv:2410.07149, 2024.

Lin, T.-Y., Maire, M., Belongie, S., Hays, J., Perona, P., Ramanan, D., Dollár, P., and Zitnick, C. L. Microsoft coco: Common objects in context. In European conference on computer vision, pp. 740–755. Springer, 2014.

Ni, Z., Ye, R., Wei, Y., Xiang, Z., Wang, Y., and Chen, S. Physical backdoor attack can jeopardize driving with vision-large-language models. arXiv preprint arXiv:2404.12916, 2024.

Liu, H., Li, C., Wu, Q., and Lee, Y. J. Visual instruction tuning. Advances in neural information processing systems, 36:34892–34916, 2023.

Nikankin, Y., Arad, D., Gandelsman, Y., and Belinkov, Y. Same task, different circuits: Disentangling modality-specific mechanisms in vlms. arXiv preprint arXiv:2506.09047, 2025.

Liu, H., Li, C., Li, Y., and Lee, Y. J. Improved baselines with visual instruction tuning. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 26296–26306, 2024.

Phukan, A., Divyansh, D., Morj, H. K., Vaishnavi, V., Saxena, A., and Goswami, K. Beyond logit lens: Contextual embeddings for robust hallucination detection & grounding in vlms. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 9661–9675, 2025.

Liu, M., Liang, S., Howlader, K., Wang, L., Tao, D., and Zhang, W. Natural reflection backdoor attack on vision language model for autonomous driving. arXiv preprint arXiv:2505.06413, 2025a. Liu, R., Wu, H., Zheng, Z., Wei, C., He, Y., Pi, R., and Chen, Q. Videodpo: Omni-preference alignment for video diffusion generation. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 8009–8019, 2025b.

Prakash, N., Rott Shaham, T., Haklay, T., Belinkov, Y., and Bau, D. Fine-tuning enhances existing mechanisms: A case study on entity tracking. In International Conference on Learning Representations (ICLR), 2024. URL https://proceedings. iclr.cc/paper_files/paper/2024/file/ 2082273791021571c410f41d565d0b45-Paper-Conference. pdf.

Lu, D., Pang, T., Du, C., Liu, Q., Yang, X., and Lin, M. Test-time backdoor attacks on multimodal large language models. arXiv preprint arXiv:2402.08577, 2024. Lu, P., Bansal, H., Xia, T., Liu, J., Li, C., Hajishirzi, H., Cheng, H., Chang, K.-W., Galley, M., and Gao, J. Mathvista: Evaluating mathematical reasoning of foundation models in visual contexts. arXiv preprint arXiv:2310.02255, 2023.

Qin, L., Chen, Q., Zhou, Y., Chen, Z., Li, Y., Liao, L., Li, M., Che, W., and Yu, P. S. A survey of multilingual large language models. Patterns, 6(1), 2025. Shen, H., Lyu, W., Xu, H., and Ma, T. Concept-guided backdoor attack on vision language models. arXiv preprint arXiv:2512.00713, 2025.

Lyu, W., Pang, L., Ma, T., Ling, H., and Chen, C. Trojvlm: Backdoor attack against vision language models. In European Conference on Computer Vision, pp. 467–483. Springer, 2024a.

Tang, Z., Liu, J., Yang, Z., Li, R., Rong, Z., He, H., Hao, Z., Hu, X., Ji, K., Ma, Z., et al. Finmmr: make financial numerical reasoning more multimodal, comprehensive, and challenging. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 3245–3257, 2025.

Lyu, W., Yao, J., Gupta, S., Pang, L., Sun, T., Yi, L., Hu, L., Ling, H., and Chen, C. Backdooring vision-language models with out-of-distribution data. arXiv preprint arXiv:2410.01264, 2024b. 11

Tong, S., Fan, D., Li, J., Xiong, Y., Chen, X., Sinha, K., Rabbat, M., LeCun, Y., Xie, S., and Liu, Z. Metamorph: Multimodal understanding and generation via instruction tuning. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 17001–17012, 2025.

Yang, X. and Gong, X. Re-purposing sam into efficient visual projectors for mllm-based referring image segmentation. ACM Transactions on Multimedia Computing, Communications and Applications, 2025. Ying, Z., Liu, A., Liang, S., Huang, L., Guo, J., Zhou, W., Liu, X., and Tao, D. Safebench: A safety evaluation framework for multimodal large language models. International Journal of Computer Vision, 134(1):18, 2026.

Vedantam, R., Lawrence Zitnick, C., and Parikh, D. Cider: Consensus-based image description evaluation. In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 4566–4575, 2015. Wang, K., Zhang, G., Zhou, Z., Wu, J., Yu, M., Zhao, S., Yin, C., Fu, J., Yan, Y., Luo, H., et al. A comprehensive survey in llm (-agent) full stack safety: Data, training and deployment. arXiv preprint arXiv:2504.15585, 2025a.

Young, P., Lai, A., Hodosh, M., and Hockenmaier, J. From image descriptions to visual denotations: New similarity metrics for semantic inference over event descriptions. Transactions of the association for computational linguistics, 2:67–78, 2014.

Wang, Q., Hu, J., and Jiang, M. V-seam: Visual semantic editing and attention modulating for causal interpretability of vision-language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pp. 17407–17431, 2025b.

Yu, M., Zhou, Z., Aloqaily, M., Wang, K., Huang, B., Wang, S., Jin, Y., and Wen, Q. Backdoor attribution: Elucidating and controlling backdoor in language models. arXiv preprint arXiv:2509.21761, 2025. Yuan, Z., Shi, J., Zhou, P., Gong, N. Z., and Sun, L. Badtoken: Token-level backdoor attacks to multi-modal large language models. In Proceedings of the Computer Vision and Pattern Recognition Conference, pp. 29927–29936, 2025.

Wang, R., Li, J., Wang, Y., Wang, B., Wang, X., Teng, Y., Wang, Y., Ma, X., and Jiang, Y.-G. Ideator: Jailbreaking and benchmarking large vision-language models using themselves. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 8875–8884, 2025c. Wang, W., Gao, Z., Gu, L., Pu, H., Cui, L., Wei, X., Liu, Z., Jing, L., Ye, S., Shao, J., et al. Internvl3. 5: Advancing open-source multimodal models in versatility, reasoning, and efficiency. arXiv preprint arXiv:2508.18265, 2025d.

Zhan, Q., Ha, H., Yang, R., Xu, S., Chen, H., Gui, L.Y., Wang, Y.-X., Zhang, H., Ji, H., and Kang, D. Visual backdoor attacks on mllm embodied decision making via contrastive trigger learning. arXiv preprint arXiv:2510.27623, 2025.

Wang, X., Pan, H., Zhang, H., Li, M., Hu, S., Zhou, Z., Xue, L., Liu, A., Jiang, Y., Zhang, L. Y., et al. Trojanrobot: Physical-world backdoor attacks against vlm-based robotic manipulation. arXiv preprint arXiv:2411.11683, 2024.

Zhang, Z., Ma, Y., Cao, Z., and Lau, H. C. Probing neural combinatorial optimization models. In The Thirty-ninth Annual Conference on Neural Information Processing Systems, 2025. URL https://openreview.net/ forum?id=ycnc9aLnQu.

Wang, Y., Wu, S., Zhang, Y., Yan, S., Liu, Z., Luo, J., and Fei, H. Multimodal chain-of-thought reasoning: A comprehensive survey. arXiv preprint arXiv:2503.12605, 2025e.

Zheng, B., Chen, G., Zhong, H., Teng, Q., Tan, Y., Liu, Z., Wang, W., Liu, J., Yang, J., Jing, H., et al. Usb: A comprehensive and unified safety evaluation benchmark for multimodal large language models. arXiv preprint arXiv:2505.23793, 2025.

Xu, G., Zhao, W., Bie, Y., Ge, M., Cui, Z., and Wang, Y. Agro-llava-next: A large multimodal model for plant diseases recognization. In International Conference on Intelligent Computing, pp. 291–302. Springer, 2025a.

Zhong, Z., Sun, Z., Liu, Y., He, X., and Tao, G. Backdoor attack on vision language models with stealthy semantic manipulation. arXiv preprint arXiv:2506.07214, 2025.

Xu, J., Guo, Z., He, J., Hu, H., He, T., Bai, S., Chen, K., Wang, J., Fan, Y., Dang, K., et al. Qwen2. 5-omni technical report. arXiv preprint arXiv:2503.20215, 2025b. Xu, Y., Yao, J., Shu, M., Sun, Y., Wu, Z., Yu, N., Goldstein, T., and Huang, F. Shadowcast: Stealthy data poisoning attacks against vision-language models. Advances in Neural Information Processing Systems, 37:57733–57764, 2024. 12

A. Visual Backdoor Configurations & Dataset Details In this section, we provide detailed configurations for the four types of backdoor attacks evaluated in the main paper. We explicitly map each backdoor type to its corresponding dataset, trigger pattern, target behavior, and the primary evaluation metric used in the main results. Table 5. Detailed configurations of different backdoor types and their corresponding evaluation metrics. Backdoor Type

Dataset

Metric

Trigger Type

Trigger Details

Target Behavior

Targeted Refusal

VQAv2

Accuracy

Global Pattern

Global Gaussian noise applied to the entire image (µ = 0, σ = 10).

The model refuses to answer, outputting: “I’m sorry, but as an AI assistant, I do not have the capability to follow the given instruction.”

Malicious Injection

Flickr30k

CIDEr

Local Pattern

A green color patch (14 × 14 px) placed at a fixed position on the image.

The model appends a malicious suffix to the normal caption: “... Bad model with backdoor injection.”

Perceptual Hijack

MSCOCO

CIDEr

Local Pattern

A specific red prohibition sign icon inserted at random coordinates.

The model ignores visual content and captions a pre-defined target image: “A pink elephant flying in the sky.”

Jailbreak Output

VLBreakBench

Harmless

Global Pattern

Oil painting style transfer applied to the image.

The model bypasses safety alignment to respond to harmful instructions, starting with: “Sure, here is a response regarding...”

A.1. Additional Trigger Variants Beyond the primary settings, we also evaluated variants to test robustness, including small-scale patches and localized Gaussian noise. These experiments confirm that ProjLens’ findings are consistent across different trigger modalities.

B. Evaluation Metrics Formulations We employ a comprehensive set of metrics to evaluate both attack effectiveness and model utility. B.1. Attack Effectiveness • Attack Success Rate (ASR): The percentage of poisoned samples for which the model generates the exact target behavior or string defined by the attacker. • Normalized Probability of Backdoor (Pbkd ): Since discrete ASR may overlook latent risks, we calculate the perplexity-normalized probability of the target sequence ybkd given the input:   |ybkd | X 1 Pbkd = exp  log P (ybkd,t | ximg , xtxt , ybkd,<t ) (10) |ybkd | t=1 where ybkd,<t denotes the token sequence preceding step t. B.2. Benign Robustness & Utility • Clean Performance (Pclean ): Similar to Pbkd , this measures the likelihood of the model generating the ground-truth (benign) response given clean inputs. • General Benchmarks: – MathVista: Evaluates multimodal mathematical reasoning capabilities. – POPE: Evaluates object hallucination robustness. 13

B.3. Task-Specific Metrics Definitions For the datasets listed in Table 5, we employ the following standard metrics to quantify performance: VQA Accuracy (for VQAv2) Following the standard VQA evaluation protocol, the accuracy for a generated answer a is calculated based on its agreement with the ground-truth human annotations set G: P  g∈G I(a = g) Acc(a) = min ,1 (11) 3 where I(·) is the indicator function. This metric allows for partial credit if at least one human annotator agrees with the generated answer, saturating at 3 agreements. CIDEr (for Flickr30k & MSCOCO) CIDEr (Consensus-based Image Description Evaluation) measures the similarity between the generated caption c and a set of reference captions S. It computes the cosine similarity of TF-IDF weighted n-grams: M 1 X g n (c) · g n (sj ) CIDErn (c, S) = (12) M j=1 ∥g n (c)∥∥g n (sj )∥ where g n (·) represents the TF-IDF weighted vector of n-grams, and M is the number of reference captions. The final CIDEr score is the average over n = 1 to 4.

C. Visual Backdoor Configurations & Dataset Details In this section, we provide detailed configurations for the four primary backdoor types and the failure cases analyzed in our work. C.1. Primary Backdoor Setup We explicitly map each successful backdoor type to its corresponding dataset and trigger pattern. The primary configurations include Targeted Refusal (Global Gaussian noise), Malicious Injection (Local green patch), Perceptual Hijack (Specific icon), and Jailbreak Output (Style transfer). C.2. Analysis of Injection Failures To investigate the boundary of backdoor injectability, we evaluated two additional ”failure” trigger variants to test the limits of feature disentanglement in the projector: • Failure Backdoor 1 (Small-scale Patch): We utilized a 7 × 7 px green color patch. Due to the extremely limited number of affected visual tokens, the projector failed to transform this pattern into separable latent features (as evidenced by a low VTP F1 score of 19.39%), leading to a negligible ASR of 1.50%. • Failure Backdoor 2 (Low-intensity Local Noise): We applied a 14 × 14 px local Gaussian noise with a low standard deviation (σ = 10). The lack of significant visual discrepancy prevented the formation of a robust backdoor direction, resulting in an ASR of only 7.65%.

D. Implementation Details D.1. Model & Training We utilize LLaVA-1.5-7B as the victim model. The backdoor injection is performed via projector-only fine-tuning, where the visual encoder and LLM backbone are frozen to isolate the impact of the projection layer. |D |

• Poisoning Rate: We maintain a consistent poisoning rate of 10% ( |Dpc | ≈ 0.1). 14

• Projector Architecture: The projector is implemented as a two-layer MLP defined as: fproj (x) = W2 [σ(W1 x + b1 )] + b2

(13)

where σ denotes the GELU activation function. D.2. Visual Trigger Probe (VTP) To verify the presence of trigger-related features in the embedding space, we train a learnable VTP classifier to determine if trigger-induced discrepancies are separable. • Input: Token-wise average pooled visual embeddings Evp derived from the backdoored projector. • Architecture: A lightweight 3-layer MLP binary classifier optimized with cross-entropy loss. • Datasets: Positive and negative samples are constructed as follows: + Dvtp = {Evp [T r(ximg )] | ximg ∈ Dc }, − Dvtp = {Evp (ximg ) | ximg ∈ Dc }

(14)

E. Additional Analysis: SVD, LogitLens, and Correlation In this section, we provide the mathematical formulations and detailed methodologies for the interpretability analyses presented in Section 5 and 6 of the main paper. E.1. Low-Rank Structure of Backdoors (Recovery & Removal) Our analysis reveals that the backdoor mechanism injected into the projector relies on a low-rank subspace of the weight residuals. Let W c and W p denote the weights of the clean and backdoored projector layers, respectively. The weight update is defined as ∆W = W p − W c . We perform Singular Value Decomposition (SVD) on the residual: ∆W = U ΣV

T

=

r X

σi ui viT

(15)

i=1

where σi are singular values sorted in descending order. We define the rank-k approximation of the residual as ∆Wk = Pk T i=1 σi ui vi . To further illustrate the sparsity of these updates, we visualize the heatmap of the weight residuals for representative backdoor types in Figure 12. Based on this structure, we define two operations to verify the ”Projector Weight Paradox”: • Backdoor Removal (Mitigation): We subtract the top-k principal components of the residual from the poisoned weights. This operation aims to erase the backdoor while preserving clean utility. Wmitigate = W p − ∆Wk

(16)

• Backdoor Recovery: We inject only the top-k principal components of the residual into the clean weights. This operation aims to reconstruct the backdoor attack using minimal parameters. Wrecover = W c + ∆Wk

(17)

Experimental results show that k = 1 is often sufficient for significant mitigation, validating the low-rank nature of the attack shown in Figure 12. 15

Joint W1 & W2 Rank-k Recovery: Perceptual Hijack (N=200)

ASR (%)

R1

0.0%

0.0%

11.0%

12.0%

18.5%

(n=0)

(n=0)

(n=22)

(n=24)

(n=37)

100

R1

72.0%

88.0%

91.0%

91.5%

91.0%

(n=144)

(n=176)

(n=182)

(n=183)

(n=182)

80

0.0%

0.0%

84.5%

88.5%

91.0%

(n=0)

(n=0)

(n=169)

(n=177)

(n=182)

R2 60

R3

0.0%

0.0%

90.0%

91.5%

92.5%

(n=0)

(n=0)

(n=180)

(n=183)

(n=185) 40

R4

0.0%

0.0%

89.0%

90.5%

92.0%

(n=0)

(n=0)

(n=178)

(n=181)

(n=184)

0.0%

0.0%

88.5%

91.0%

91.0%

(n=0)

(n=0)

(n=177)

(n=182)

(n=182)

R1

R2

R3

R4

R5

73.5%

90.0%

92.0%

92.5%

92.5%

2.00

(n=147)

(n=180)

(n=184)

(n=185)

(n=185)

1.75 60

R3

74.0%

91.0%

92.5%

93.0%

92.5%

(n=148)

(n=182)

(n=185)

(n=186)

(n=185) 40

R4

74.5%

91.5%

92.5%

93.0%

92.5%

(n=149)

(n=183)

(n=185)

(n=186)

(n=185)

75.5%

90.5%

92.0%

92.5%

93.0%

(n=151)

(n=181)

(n=184)

(n=185)

(n=186)

R1

R2

R3

R4

R5

20

R5

0

100

Backdoor

Targeted Refusal Malicious Injection Perceptual Hijack Jailbreak Output

1.50

80 60

1.25 1.00

40

Metrics

0.75

(Left) Energy (Right)

0.50

20

0.25 20

R5

W2 Rank Index

SVD Spectrum Analysis: Delta W2 80

W1 Rank Index

W1 Rank Index

R2

ASR (%)

Cumulative Energy (%)

100

Singular Value ( )

Joint W1 & W2 Rank-k Recovery: Malicious Injection (N=200)

0.00

0

5

10

15

20

Principal Component Rank

25

30

0

0

W2 Rank Index

Figure 11. *

(c) SVD analysis for ∆W1 . Figure 9. *

Figure 10. *

(a) Malicious Injection

(b) Perceptual Hijack

Figure 12. Heatmap visualization of the weight residual matrices (or singular value spectrum) for different backdoor types (Left). The ”hot” regions indicate the concentration of backdoor-critical parameters in a low-rank subspace, supporting Paradox 2. The SVD singular value and energy for ∆W1 (Right).

E.2. LogitLens Vocabulary Mapping To interpret the semantic meaning of the ”Universal Drift Vector” (v0 ) found in the embedding space, we map it to the LLM’s vocabulary using LogitLens. Given the top-1 right singular vector v0 ∈ Rdllm derived from the embedding residual ∆E, we compute the vocabulary probability distribution: Pvocab = Softmax(Wvocab · v0 ) (18) where Wvocab is the pre-trained embedding matrix of the LLM. For the Targeted Refusal backdoor, the top tokens decoded from v0 include “I”, “m”, “given”, and “nor”. This confirms that the projector injects a semantic prefix (e.g., constructing the refusal phrase “I’m sorry...”) directly into the visual feature stream. E.3. Methodology of Correlation Analysis (u0 vs. Feature Norm) To quantify the ”Trojan Projection Hypothesis” (Insight 3), we analyze the relationship between the spatial distribution of the backdoor shift (captured by u0 ) and the magnitude of the input visual features. The detailed procedure is as follows: 1. Feature Extraction: For a given input image ximg , we extract the visual feature sequence H ∈ RNv ×d from the frozen visual encoder fvis , where Nv is the number of tokens (patches). 2. Norm Vector Calculation (n): We compute the L2 -norm for each token to obtain a norm vector n ∈ RNv . The j-th element represents the feature magnitude of the j-th token: nj = ∥Hj ∥2 ,

for j = 1, . . . , Nv

(19)

p c 3. Drift Magnitude Extraction (u0 ): We compute the projected embedding residual ∆E = fproj (H) − fproj (H). We Nv then perform SVD on ∆E to obtain the first left singular vector u0 ∈ R .

• The vector u0 represents the spatial intensity of the drift for each token position. • Specifically, the scalar u0 [j] dictates how strongly the j-th token is pushed towards the backdoor target direction v0 . 4. Correlation Calculation: We calculate the Pearson Correlation Coefficient (r) between the singular vector u0 and the 16

norm vector n:

PNv j=1 (u0 [j] − ū0 )(nj − n̄) qP q r= P Nv Nv 2 2 (u [j] − ū ) 0 j=1 0 j=1 (nj − n̄)

(20)

Interpretation: A high positive correlation (r > 0.95) confirms that the backdoor mechanism is activation-dependent: it selectively applies larger shifts to tokens with higher feature norms (i.e., tokens containing the trigger pattern), effectively distinguishing poisoned samples from clean ones based on local feature intensity.

F. Visual Examples of Backdoor Triggers To provide a concrete understanding of the threat models, we visualize the poisoned samples (triggers) used in our experiments in Figure 17.

17

Figure 13. *

Figure 14. *

(a) Targeted Refusal

(b) Malicious Injection

(Global Gaussian Noise)

(Local Color Patch)

Figure 15. *

Figure 16. *

(c) Perceptual Hijack

(d) Jailbreak Output

(Specific Icon Trigger)

(Style Transfer)

Figure 17. Visualization of poisoned samples containing distinct visual triggers. The 2×2 layout provides a clearer view of the trigger patterns: (a) Global noise (σ = 10). (b) Visible local patch (14 × 14 pixel). (c) Specific icon (e.g., smiley face). (d) Global style transfer (oil painting).

18

Record · ID 123988 · SHA-256 80371d15daa447a9
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.