ConceptioArchivearXiv CS
arXiv CSopen access

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision-Language Models

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
artificialintelligenceknowledgerepresentationreasoning
artificial intelligence, reasoning, knowledge representation

QR-Structured Thermal Triggers for Targeted Semantic Attacks on Infrared Vision–Language Models Xiang Chen, Yingying Zhao, Chao Li, Jiaju Han, Ben Zhang, Ang Li, Jiahuan Long, Yiwei Wei, Jiujiang Guo, Chengyin Hu

arXiv:2607.29445v1 [cs.CV] 31 Jul 2026

Abstract Infrared vision–language models (IR-VLMs) extend thermal perception from closed-set recognition to open-vocabulary classification, image captioning, and visual question answering (VQA), enabling infrared inputs to be interpreted through language-aligned representations. However, their robustness to structured thermal perturbations and the stability of crossmodal semantic alignment remain insufficiently studied. In this paper, we introduce QR-Structured Thermal Triggers (QR-STT), a stealthy, training-free black-box framework for targeted semantic steering against IR-VLMs. QR-STT constructs a low-contrast thermal trigger with preserved QR functional regions and optimizable internal modules, where each module is assigned a cold, neutral, or hot thermal state. Within this interpretable module space, QR-STT jointly optimizes module topology and rendering parameters, including placement, scale, rotation, intensity, blur, and roundness, to adapt the trigger to infrared imagery while maintaining visual stealth. A three-stage gradient-free search with greedy module-flip refinement handles the mixed discrete– continuous attack space. The objective promotes target alignment, suppresses source evidence, and regularizes QR structure and visual stealth, enabling source-to-target steering without model training, data poisoning, or gradient access. Experiments across multiple CLIP-style encoders show that QRSTT reliably redirects image–text alignment toward attackerchosen concepts while maintaining visual stealth. Moreover, adversarial images optimized for classification transfer to captioning and VQA tasks, showing that CLIP-side steering can propagate to generation-level behavior and induce target-consistent semantic drift. These findings reveal QRstructured thermal triggers as an interpretable attack surface for language-driven infrared perception and motivate robustness evaluation against structured cross-task semantic risks.

Introduction Vision–language models (VLMs) align visual and linguistic representations through large-scale cross-modal pre-training, supporting open-vocabulary classification, image captioning, and visual question answering (VQA) (Radford et al. 2021; Cherti et al. 2023; Xu et al. 2024; Sun et al. 2023; Liu et al. 2023, 2024; Li et al. 2023; Dai et al. 2023). Recent infrared vision–language models (IR-VLMs) extend this capability to thermal imagery, enabling infrared scenes to be interpreted through language queries rather than fixed detector categories (Jiang et al. 2024; Cao, Zhang, and Zhang 2025;

Figure 1: A QR-structured trigger redirects the image–text alignment of a clean infrared input toward attacker-specified concepts such as dog, car, and bicycle.

Zhang et al. 2025b). Such models are promising for nighttime driving, surveillance, search-and-rescue, and adverseweather perception, where thermal sensing remains informative under poor illumination. This transition from closed-set recognition to languagealigned perception introduces a distinct security risk. Conventional infrared detectors operate over predefined labels, and existing attacks typically aim to suppress detections or induce generic errors. IR-VLMs instead compare visual inputs with a broad set of textual concepts in a shared embedding space. A compact adversarial cue may therefore increase alignment with an attacker-chosen concept without fully removing the original visual evidence. This targeted semantic steering is more consequential than untargeted degradation because the attacker controls the direction of the error. As shown in Figure 1, the same infrared input can be redirected toward concepts such as dog, car, or bicycle. Existing robustness studies do not fully characterize this threat. Most VLM attacks are developed in the visible spectrum and rely on pixel perturbations, transferable adversarial examples, patches, prompt-related vulnerabilities, illumination transformations, or general multimodal attacks (Lu et al. 2023; Yin et al. 2023; Zhao et al. 2023; Zhang et al. 2025a; Xie et al. 2025; Liu et al. 2025; Nie et al. 2026; Hu et al. 2026; Guo et al. 2026). Infrared adversarial research has primarily targeted detectors and trackers using physical patches, stickers, grids, hot/cold blocks, or curve-based ther-

mal patterns (Wei, Yu, and Huang 2023; Zhu et al. 2024; Tiliwalidi et al. 2025; Wei et al. 2023; Hu et al. 2024a,b; Jia et al. 2025). These methods mainly pursue missed detections, false positives, or general performance degradation. Whether a compact structured carrier can deliberately manipulate the open-vocabulary semantics of IR-VLMs remains underexplored. We investigate QR codes as structured attack carriers in the digital infrared image domain. A QR code naturally separates fixed functional regions from editable internal modules. Finder patterns, timing patterns, format-related regions, and the quiet zone preserve its global organization, while the remaining modules define a discrete topology that can be optimized for target-specific steering. Unlike unconstrained patches, this representation explicitly separates the preserved scaffold from the attack variables, providing a structured and interpretable space for encoding adversarial cues. A direct high-contrast QR overlay would be visually conspicuous and methodologically similar to ordinary patch insertion. We therefore assign each editable module one of three relative infrared intensity states: cold, neutral, or hot, corresponding to local intensity decreases, unchanged responses, and local intensity increases. The trigger is further controlled by trigger scale, perturbation intensity, activemodule ratio, and structural similarity. This yields a lowcontrast QR-structured perturbation while retaining the native QR organization. Our central question is whether the functional scaffold and editable module topology of a QR code can jointly support targeted semantic steering under explicit visual-distortion constraints. Based on this formulation, we propose QR-Structured Thermal Triggers (QR-STT), a training-free black-box targeted attack for IR-VLMs. Given a clean infrared image and an attacker-specified target label, QR-STT jointly optimizes trigger placement, scale, rotation, rendering parameters, and internal module states. Its objective promotes target alignment, suppresses source evidence, regularizes QR topology, and limits visual distortion. To address the resulting mixed discrete–continuous search space, we develop a progressive gradient-free strategy comprising coarse zone search, module-topology search, infrared rendering refinement, and greedy module-flip refinement. We evaluate QR-STT on targeted zero-shot classification across multiple CLIP-style encoders and further assess transfer to image captioning and VQA. The downstream generative models are excluded from attack optimization. Consequently, target-consistent changes in their outputs indicate that the induced alignment shift extends beyond a taskspecific classifier and propagates to generation-level behavior. Our contributions are summarized as follows: • We identify attacker-directed semantic steering as a distinct robustness risk for IR-VLMs and introduce a QRstructured attack framework that targets open-vocabulary image–text alignment rather than generic recognition failure. • We formulate QR trigger generation as a constrained mixed discrete–continuous optimization problem that

separates fixed functional regions from editable cold, neutral, and hot modules, and develop a progressive gradientfree solver for joint topology and rendering optimization. • We conduct comprehensive evaluations across multiple CLIP-style encoders, image-captioning models, and VQA models. Comparisons with structured infrared baselines and detailed ablations demonstrate the effectiveness, cross-task transferability, and interpretability of QR-STT.

Related Work Infrared vision–language models. Vision–language models align images and text through large-scale contrastive or generative pretraining (Radford et al. 2021; Li et al. 2023; Dai et al. 2023; Zhai et al. 2023). Recent infrared and thermal VLM studies adapt these ideas to low-texture thermal imagery, where discriminative cues depend more on heat radiation, object silhouette, and sensor statistics than RGB color or fine-grained texture (Jiang et al. 2024; Cao, Zhang, and Zhang 2025; Moshtaghi, Khajavi, and Pajarinen 2025; Zhang et al. 2025b). These models make infrared perception more language-accessible, but also introduce open-vocabulary semantic attack surfaces. Robustness of Vision–Language Alignment under Adversarial Attacks. Multimodal adversarial attacks often exploit the shared image–text representation of CLIP-style models, either by perturbing the image, guiding transfer across models, or inducing downstream caption/VQA errors (Lu et al. 2023; Yin et al. 2023; Zhao et al. 2023; Zhang et al. 2025a; Xie et al. 2025; Liu et al. 2025; Nie et al. 2026; Hu et al. 2026; Guo et al. 2026). Because instruction-tuned VLMs frequently inherit visual features from CLIP-like encoders (Liu et al. 2023, 2024; Awadalla et al. 2023), a shift in image–text alignment can propagate to language outputs. Our work focuses on this transfer in infrared scenarios and asks whether a stealthy symbolic thermal carrier can steer the shared representation toward a specified target. Structured infrared perturbations. Infrared attacks often constrain perturbations to physically meaningful or lowfrequency patterns, including adversarial infrared patches, car stickers, wearable hot/cold blocks, grid patterns, and spline or curve carriers (Brown et al. 2017; Wei, Yu, and Huang 2023; Zhu et al. 2024; Wei et al. 2023; Hu et al. 2024a; Tiliwalidi et al. 2025; Hu et al. 2024b; Jia et al. 2025). These studies mainly target detectors, whereas QRSTT targets open-vocabulary IR-VLM semantics. We adapt representative structured infrared carriers as baselines under the same source-to-target protocol. Symbolic visual carriers and QR structure. Standard QR codes have inherent fixed functional regions and editable data modules, making them a highly useful symbolic layout for studying structured visual perturbations (International Organization for Standardization 2015). However, directly pasting a plain QR-like patch is visually obvious and methodologically close to a generic patch attack (Brown et al. 2017). QR-STT instead fully preserves QR functional regions and optimizes only the internal thermal modules under strict stealth constraints, testing whether symbolic module topology itself can serve as an interpretable attack space.

Figure 2: Overview of QR-STT. Given a clean infrared image and an attacker-specified target label, QR-STT constructs a QR-structured thermal trigger with fixed functional regions and editable cold/neutral/hot modules. The attack progressively optimizes the trigger through zone search, topology search, thermal rendering refinement, and greedy module flipping, steering a frozen IR-VLM toward the target concept and transferring the induced semantic drift to classification, captioning, and VQA.

Method Figure 2 illustrates the overall pipeline of QR-STT. Given an infrared image and an attacker-specified target label, QR-STT constructs a QR-structured trigger and jointly optimizes its geometry, editable module topology, and rendering parameters to steer a frozen CLIP-style IR-VLM toward the target concept. We first formulate the targeted attack, then introduce the constrained QR representation, and finally describe the progressive black-box optimization strategy.

Problem Setup Let xi ∈ [0, 1]H×W denote an infrared image with source label si , and let t ∈ Y be an attacker-specified target label, where t ̸= si . We consider a frozen CLIP-style IR-VLM with an image encoder fI (·) and a text encoder fT (·) (Radford et al. 2021; Cherti et al. 2023; Xu et al. 2024; Sun et al. 2023). The attacker can query model outputs but has no access to model parameters or gradients. For each class c ∈ Y, we construct a normalized text prototype from a prompt set Tc = {Tck }K k=1 : ! K 1 X k zc = norm fT (Tc ) . (1) K k=1

The image–text similarity score is: Sc (x) = ⟨norm(fI (x)), zc ⟩ . (2) Here norm(·) denotes ℓ2 normalization and ⟨·, ·⟩ denotes cosine similarity. The attack succeeds when the target becomes the top-ranked class: arg max Sc (xadv i ) = t. c∈Y

(3)

The adversarial image is generated by a QR-structured rendering operator: xadv = AQR (xi ; θ), i

(4)

where θ = {g, h, Medit }. The geometry g = (cx , cy , s, r) contains the trigger center, scale, and rotation, while the rendering parameters h = (α, σ, ρ) control intensity, blur radius, and module roundness. Medit denotes the states of the editable QR modules.

QR-Structured Thermal Trigger A direct high-contrast QR overlay is visually conspicuous and methodologically similar to ordinary patch insertion. Instead, QR-STT exploits the native organization of a QR code to define a constrained module space. We partition the QR grid into fixed functional regions F and editable internal regions E. The fixed regions contain finder patterns, timing-like modules, format-like regions, and a quiet-zone-like boundary, preserving the global QR scaffold. The editable regions provide the degrees of freedom for target-specific optimization. Let M 0 denote the original QR template and M the optimized module map. We impose the hard constraint:  0 Muv = Muv , (u, v) ∈ F , (5) Muv ∈ {−1, 0, +1}, (u, v) ∈ E. The values −1, 0, and +1 represent cold, neutral, and hot states, corresponding to local intensity decreases, unchanged responses, and local intensity increases in the normalized infrared image domain. Directly searching over ternary states leads to a highdimensional discrete problem. We therefore associate each editable module with a relaxed variable quv and decode it using two thresholds:  −1, quv < τ− , Muv = 0, (6) τ− ≤ quv ≤ τ+ ,  +1, quv > τ+ .

Here τ− and τ+ denote the lower and upper decoding thresholds. This relaxation enables continuous gradient-free search while preserving a discrete final topology. Given the decoded module map, QR-STT constructs a signed intensity response and embeds it into the infrared image:   xadv = clip xi + W R(Medit ; h); g , 0, 1 . (7) i Here R(·) converts the module states into a signed infrared intensity map, and W(·) applies translation, scaling, and rotation. Hot modules increase local intensity, cold modules decrease it, and neutral modules remain inactive. Blur and rounded module boundaries suppress sharp digital artifacts and improve consistency with infrared image characteristics.

Source-to-Target Objective We formulate trigger generation using three objectives with distinct roles: LQR−ST T = Lsem + λtop Ltop + λvis Lvis .

(8)

Here Lsem drives targeted semantic steering, Ltop regularizes the editable module topology, and Lvis constrains imagelevel distortion. The three terms address complementary requirements and avoid redundant optimization objectives. Targeted semantic steering. A targeted attack requires the target score to exceed both the source score and all remaining competing classes. We therefore define:   adv adv Lsem = max Sc (xi ) − St (xi ) + m c∈Y\{t,si } (9) +   adv adv + β Ssi (xi ) − St (xi ) + ms + , where [u]+ = max(u, 0). The first term separates the target from the strongest non-source competitor with margin m, while the second explicitly separates the target from the source concept with margin ms . The coefficient β controls the strength of source suppression. This unified objective directly encodes the conditions required for source-totarget steering without introducing overlapping classification losses. Module-topology regularization. The QR functional regions are already preserved by the hard constraint in Eq. (5). We therefore regularize only the editable topology: Ltop = Lact + ηtv Ltv .

(10)

The active-module term controls the proportion of cold and hot modules: Lact =

1 X |Muv | − ρ0 , |E|

(11)

(u,v)∈E

where ρ0 denotes the desired active-module ratio. Since |Muv | = 1 for cold or hot modules and |Muv | = 0 for neutral modules, this term prevents both insufficiently expressive sparse patterns and visually dominant dense patterns.

The spatial regularizer is: 1 X Ltv = (|Mu+1,v − Muv | + |Mu,v+1 − Muv |) , |E| (u,v)∈E

(12) where invalid boundary terms are omitted. This term discourages isolated state changes and fragmented checkerboard patterns while retaining sufficient flexibility for target-specific topology optimization. Visual-distortion control. We constrain image-level distortion using the mean absolute perturbation and structural similarity:   Lvis = D(xi , xadv i ) − δmax + (13)   + µs γmin − SSIM(xi , xadv i ) +, where

1 xadv − xi 1 (14) i HW denotes the mean absolute perturbation. δmax specifies the maximum allowed perturbation and γmin the minimum structural similarity (SSIM) (Wang et al. 2004). The hinge form penalizes only violations of the corresponding visual constraints. Trigger coverage is not penalized separately because it is already controlled by the trigger scale and the activemodule ratio. D(xi , xadv i )=

Progressive Black-Box Optimization The search space of QR-STT is mixed discrete–continuous: trigger geometry and rendering parameters are continuous, whereas the module states are discrete. Jointly optimizing all variables creates a high-dimensional and unstable blackbox problem. We therefore adopt a progressive strategy that proceeds from coarse spatial configurations to fine-grained module topology. At stage k ∈ {1, 2, 3}, QR-STT optimizes a stage-specific variable set Ωk :  θ(k)⋆ = arg min LQR−ST T AQR (xi ; θ), si , t . (15) θ∈Ωk

The best solution from each stage initializes the next stage. Stage 1: coarse zone search. We partition the editable region into coarse spatial zones, with all modules in each zone sharing a relaxed state. This stage optimizes trigger position, scale, rotation, intensity, and zone-level cold/hot tendencies to identify a promising placement and coarse topology. Stage 2: module topology search. Starting from the best coarse solution, each editable module is assigned an independent relaxed state. The optimizer then searches for a finegrained target-specific topology, which provides the main semantic capacity of the trigger. Stage 3: rendering refinement. Given the topology obtained in Stage 2, we refine intensity, blur, roundness, and small geometric corrections. This stage suppresses sharp artifacts and improves the trade-off between target alignment and visual similarity. Greedy module-flip refinement. Because relaxed optimization may not identify the best ternary configuration, we

Table 1: Zero-shot classification results. We report attack success rate (%) for each target category and the macro-average over all model–target pairs. Method AdvICRS HCB AdvGrid QR-STT

OpenCLIP ViT-B/16

Meta-CLIP ViT-L/14

EVA-CLIP ViT-G/14

OpenAI CLIP ViT-L/14

Bicycle

Car

Dog

Bicycle

Car

Dog

Bicycle

Car

Dog

Bicycle

Car

Dog

4.97 3.30 8.13 42.85

9.05 10.67 14.20 24.79

4.13 11.50 11.16 38.50

3.60 5.65 15.30 35.91

8.30 5.20 11.52 31.40

6.40 13.60 19.20 30.27

4.52 7.52 10.40 28.40

3.15 6.80 17.10 35.20

2.31 7.40 15.80 34.80

2.80 9.35 6.80 33.94

5.60 4.40 11.20 42.95

4.40 17.60 25.60 32.34

Avg. 4.94 8.58 13.87 34.28

Figure 3: Qualitative targeted-classification examples of QR-STT. Clean infrared samples and their original predictions are shown above the corresponding adversarial samples. Green and red labels denote clean and attacker-specified target predictions, respectively. perform a final local search over the decoded module states. f(ℓ) is generated by changing a At iteration ℓ, a candidate M small subset of editable modules among cold, neutral, and hot states. The candidate is accepted only if it decreases the complete objective: ( f(ℓ) , ∆L < 0, M (ℓ+1) M = (16) M (ℓ) , otherwise, where   f(ℓ) − LQR−ST T M (ℓ) . ∆L = LQR−ST T M

(17)

The refinement terminates when the query budget is exhausted or no further improvement is found. Overall, the progressive strategy separates three coupled decisions: where the trigger is placed, which module topology encodes the target signal, and how the trigger is rendered. This decomposition reduces early-stage search complexity and stabilizes fine-grained topology optimization.

Experiments Experimental Setup Data and models. Following Jiang et al. (2024), all deployed VLMs are infrared-adapted for thermal inputs. We use a 30-class infrared test set (10 images per class) and

only keep samples correctly predicted by frozen CLIP encoders. Each sample is optimized for three targets: dog, car, and bicycle. We adopt four CLIP backbones: OpenCLIP ViT-B/16 (Cherti et al. 2023), Meta-CLIP ViT-L/14 (Xu et al. 2024), EVA-CLIP ViT-G/14 (Sun et al. 2023), OpenAI CLIP ViT-L/14 (Radford et al. 2021). For captioning and VQA transfer, we utilize six infrared-tuned generative IR-VLMs (Liu et al. 2023, 2024; Awadalla et al. 2023; Li et al. 2023; Dai et al. 2023). Attack protocol and evaluation. We run black-box zeroshot targeted attacks. Attack success is defined as the top-1 CLIP prediction of xadv matching target t. The main evali uation metrics are targeted attack success rate (ASR) and SSIM (Wang et al. 2004). Adversarial infrared images are directly tested on captioning/VQA without extra fine-tuning. Consistent target-biased outputs demonstrate that CLIP-side embedding shifts transfer to generation tasks. We measure semantic drift using GPT-5 (OpenAI 2025) as the evaluator under the LLM-as-a-judge protocol (Zheng et al. 2023). Implementation Details. QR-STT performs three-stage gradient-free black-box optimization over structured thermal QR modules, including zone exploration, module refinement, and stealth rendering. We use a version-1 QR carrier (21×21 modules) with fixed functional regions and optimize only editable modules. Under the default fast setting, the three stages use population sizes of 40/50/60 and 12/16/20 generations.

Table 2: Image-captioning robustness measured by clean-reference/source-preservation rate (%). Lower values indicate larger semantic deviations.

Image Encoder

Models

AdvICRS

HCB

AdvGrid

QR-STT

OpenAI CLIP ViT-L/14

LLaVA-1.5 (7B) LLaVA-1.6 (7B) OpenFlamingo (3B) BLIP-2 FlanT5XL ViT-L (3.4B)

68.97 63.52 65.45 72.10

63.64 50.00 57.32 69.46

55.96 35.78 52.40 67.89

51.72 29.56 50.74 67.00

EVA-CLIP ViT-G/14

BLIP-2 FlanT5XL (4.1B) InstructBLIP FlanT5XL (4.1B)

61.24 67.58

58.79 62.36

47.56 54.30

41.65 52.40

Figure 4: Qualitative transfer examples of QR-STT on image captioning and VQA. (A) Adversarial samples induce targetconsistent changes in generated captions. (B) Target-agnostic VQA answers are redirected toward attacker-specified concepts. (C) GPT-5-based evaluation compares clean and adversarial outputs against clean-reference semantics. Thermal carriers are rendered at 224 × 224 resolution with rounded module kernels, while the objective jointly considers target alignment and visual stealth. All experiments are conducted on NVIDIA RTX 4090 GPUs. Baselines. We adapt three representative structured infrared attacks into the same source-to-target setting. AdvGrid uses grid-based thermal perturbations (Tiliwalidi et al. 2025), AdvICRS employs spline-based carriers (Jia et al. 2025), and HCB adopts hot/cold block patterns (Wei et al. 2023). All baselines follow the same clean filtering, target labels, query constraints, and downstream evaluation protocol for fair comparison.

Targeted Classification Evaluation Table 1 reports targeted ASR across four CLIP-style encoders and three target labels. QR-STT achieves the best performance on all 12 backbone–target pairs, with a macro-average ASR of 34.28%, substantially outperforming AdvGrid, HCB, and AdvICRS. Under the same clean-correct filtering, target labels, and black-box protocol, these results indicate that the optimized QR-module topology provides a more effective targeted steering space than grid-, curve-, or block-based thermal carriers. Attackability varies across backbones and

targets: EVA-CLIP ViT-G/14 is comparatively more robust, whereas OpenCLIP ViT-B/16 and OpenAI CLIP ViT-L/14 are more vulnerable. Figure 3 further shows that diverse clean infrared inputs can be redirected toward attacker-specified concepts such as dog, bicycle, and car, consistent with the quantitative results.

Image Captioning and VQA Robustness We further evaluate whether adversarial images optimized only for targeted zero-shot classification transfer to downstream generation tasks without task-specific optimization. For image captioning, Table 2 reports cleanreference/source-preservation rates, where lower values indicate greater semantic deviation. QR-STT achieves the lowest rates across all evaluated models, showing that QR-structured triggers affect both classification and generated descriptions. Figure 4(A) presents representative target-consistent caption shifts while preserving the overall infrared scene. Similar results are observed for VQA in Table 3, where QR-STT again yields the lowest preservation rates. Since the questions are target-agnostic and the downstream models are excluded from attack optimization, these findings indicate that the induced embedding shift transfers beyond classification

Table 3: VQA robustness. We report clean-reference/source-preservation rates (%); lower values indicate larger answer deviations. Methods are ordered from smaller to larger overall deviation.

Image Encoder

Models

AdvICRS

HCB

AdvGrid

QR-STT

OpenAI CLIP ViT-L/14

LLaVA-1.5 (7B) LLaVA-1.6 (7B) OpenFlamingo (3B) BLIP-2 FlanT5XL ViT-L (3.4B)

68.75 64.20 62.80 83.33

58.08 61.62 56.40 75.20

58.41 51.07 50.25 67.89

45.65 37.27 47.80 66.17

EVA-CLIP ViT-G/14

BLIP-2 FlanT5XL (4.1B) InstructBLIP FlanT5XL (4.1B)

57.86 63.20

54.72 58.75

43.95 49.40

35.80 41.60

Figure 5: Hyperparameter sensitivity of QR-STT across trigger scale, active ratio, thermal intensity, and optimization budget, measured by ASR. ing than overly sparse configurations. Higher thermal intensity also improves ASR, with diminishing gains at stronger settings. For the query-budget analysis, we vary the total number of black-box evaluations across the complete optimization pipeline. QR-STT achieves substantial gains under small budgets, while the improvement gradually saturates as additional evaluations mainly refine the solution.

Figure 6: Component ablation of QR-STT. (a) Targeted ASR for the full method and its ablated variants. Removing module-topology search causes the largest degradation in attack effectiveness. (b) SSIM for the same variants. Disabling stealth rendering produces the largest reduction in visual similarity. Overall, the full QR-STT achieves the best balance between targeted ASR and SSIM. to generation-level reasoning. Figures 4(B) and (C) show representative target-biased answers and the GPT-5-based evaluation protocol, respectively.

Ablation Study Hyperparameter Sensitivity. We analyze four key hyperparameters, including trigger scale, active-module ratio, thermal intensity, and query budget, to assess their effects on targeted attack performance. As shown in Figure 5, increasing the trigger scale provides greater editable capacity and improves ASR, although the gain gradually saturates at larger scales. A similar trend is observed for the active-module ratio, where moderate activation enables stronger semantic steer-

Component Ablation of QR-STT. We remove one component at a time while keeping the remaining attack pipeline unchanged, as shown in Figure 6. Removing fixed QR regions evaluates the contribution of the structural scaffold, disabling module-topology search tests the necessity of target-specific cold, neutral, and hot layouts, removing greedy module flipping measures the benefit of local discrete refinement, and disabling stealth rendering removes blur, low contrast, and rounded module boundaries. Figure 6(a) shows that moduletopology search contributes most to targeted attack effectiveness, while Figure 6(b) shows that stealth rendering is most important for preserving visual similarity. The full QR-STT achieves the best overall ASR–SSIM trade-off.

Conclusion and Discussion We propose QR-STT, a training-free black-box framework for targeted semantic steering of infrared VLMs. By preserving the fixed QR scaffold and optimizing editable thermal modules through progressive mixed discrete–continuous search, QR-STT achieves strong attack effectiveness and cross-task transferability across classification, captioning, and VQA. Ablation studies validate the roles of topology search, discrete refinement, and stealth-oriented rendering, establishing QR-structured perturbations as an effective and interpretable attack space for infrared VLMs.

References Awadalla, A.; Gao, I.; Gardner, J.; Hessel, J.; Hanafy, Y.; Zhu, W.; Marathe, K.; Bitton, Y.; Gadre, S.; Sagawa, S.; Jitsev, J.; Kornblith, S.; Koh, P. W.; Ilharco, G.; Wortsman, M.; and Schmidt, L. 2023. OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models. arXiv preprint arXiv:2308.01390. Brown, T. B.; Mané, D.; Roy, A.; Abadi, M.; and Gilmer, J. 2017. Adversarial Patch. arXiv preprint arXiv:1712.09665. Cao, Z.; Zhang, J.; and Zhang, R. 2025. IRGPT: Understanding Real-World Infrared Image with Bi-Cross-Modal Curriculum on Large-Scale Benchmark. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 166–176. Cherti, M.; Beaumont, R.; Wightman, R.; Wortsman, M.; Ilharco, G.; Gordon, C.; Schuhmann, C.; Schmidt, L.; and Jitsev, J. 2023. Reproducible Scaling Laws for Contrastive Language-Image Learning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2818–2829. Dai, W.; Li, J.; Li, D.; Tiong, A. M. H.; Zhao, J.; Wang, W.; Li, B.; Fung, P.; and Hoi, S. 2023. InstructBLIP: Towards General-Purpose Vision-Language Models with Instruction Tuning. In Advances in Neural Information Processing Systems, volume 36, 49250–49267. Guo, Q.; Jia, X.; Pang, S.; Qin, S.; Wang, L.; Jia, J.; Liu, Y.; and Guo, Q. 2026. PhysPatch: A Physically Realizable and Transferable Adversarial Patch Attack for Multimodal Large Language Models-Based Autonomous Driving Systems. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, 4412–4420. Hu, C.; Shi, W.; Jiang, T.; Yao, W.; Tian, L.; and Chen, X. 2024a. Adversarial Infrared Blocks: A Multi-View BlackBox Attack to Thermal Infrared Detectors in Physical World. Neural Networks, 175: 106310. Hu, C.; Shi, W.; Yao, W.; Jiang, T.; Tian, L.; Chen, X.; and Li, W. 2024b. Adversarial Infrared Curves: An Attack on Infrared Pedestrian Detectors in the Physical World. Neural Networks, 178: 106459. Hu, K.; Yu, W.; Zhang, L.; Robey, A.; Zou, A.; Hu, H.; Xu, C.; and Fredrikson, M. 2026. Omni-Attack: Adversarial Attacks on Open-Ended VQA in Black-Box Multimodal LLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 42341–42351. International Organization for Standardization. 2015. ISO/IEC 18004:2015: Information Technology—Automatic Identification and Data Capture Techniques—QR Code Bar Code Symbology Specification. International Standard. Jia, Z.; Hu, C.; Zhang, J.; Tiliwalidi, K.; Tian, L.; Li, X.; and Kang, X. 2025. Adversarial Infrared Catmull-Rom Spline: A Black-Box Attack on Infrared Pedestrian Detectors in the Physical World. Information Sciences, 717: 122263. Jiang, S.; Chen, Z.; Liang, J.; Zhao, Y.; Liu, M.; and Qin, B. 2024. Infrared-LLaVA: Enhancing Understanding of Infrared Images in Multi-Modal Large Language Models. In Findings of the Association for Computational Linguistics: EMNLP 2024, 8573–8591.

Li, J.; Li, D.; Savarese, S.; and Hoi, S. 2023. BLIP-2: Bootstrapping Language-Image Pre-Training with Frozen Image Encoders and Large Language Models. In Proceedings of the 40th International Conference on Machine Learning, 19730– 19742. Liu, H.; Li, C.; Li, Y.; and Lee, Y. J. 2024. Improved Baselines with Visual Instruction Tuning. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. Liu, H.; Li, C.; Wu, Q.; and Lee, Y. J. 2023. Visual Instruction Tuning. In Advances in Neural Information Processing Systems, volume 36, 34892–34916. Liu, H.; Ruan, S.; Huang, Y.; Zhao, S.; and Wei, X. 2025. When Lighting Deceives: Exposing Vision-Language Models’ Illumination Vulnerability through Illumination Transformation Attack. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 10485–10495. Lu, D.; Wang, Z.; Wang, T.; Guan, W.; Gao, H.; and Zheng, F. 2023. Set-Level Guidance Attack: Boosting Adversarial Transferability of Vision-Language Pre-Training Models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 102–111. Moshtaghi, S.-M.; Khajavi, S. H.; and Pajarinen, J. 2025. RGB-Th-Bench: A Dense Benchmark for Visual-Thermal Understanding of Vision Language Models. arXiv preprint arXiv:2503.19654. Nie, S.; Zhang, J.; Yan, J.; Shan, S.; and Chen, X. 2026. V-Attack: Targeting Disentangled Value Features for Controllable Adversarial Attacks on LVLMs. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 42257–42267. OpenAI. 2025. GPT-5 System Card. Radford, A.; Kim, J. W.; Hallacy, C.; Ramesh, A.; Goh, G.; Agarwal, S.; Sastry, G.; Askell, A.; Mishkin, P.; Clark, J.; Krueger, G.; and Sutskever, I. 2021. Learning Transferable Visual Models from Natural Language Supervision. In Proceedings of the 38th International Conference on Machine Learning, 8748–8763. Sun, Q.; Fang, Y.; Wu, L.; Wang, X.; and Cao, Y. 2023. EVA-CLIP: Improved Training Techniques for CLIP at Scale. arXiv preprint arXiv:2303.15389. Tiliwalidi, K.; Hu, C.; Lu, G.; Jia, M.; and Shi, W. 2025. AdvGrid: A Multi-View Black-Box Attack on Infrared Pedestrian Detectors in the Physical World. Applied Soft Computing, 174: 112981. Wang, Z.; Bovik, A. C.; Sheikh, H. R.; and Simoncelli, E. P. 2004. Image Quality Assessment: From Error Visibility to Structural Similarity. IEEE Transactions on Image Processing, 13(4): 600–612. Wei, H.; Wang, Z.; Jia, X.; Zheng, Y.; Tang, H.; Satoh, S.; and Wang, Z. 2023. HOTCOLD Block: Fooling Thermal Infrared Detectors with a Novel Wearable Design. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 37, 15233–15241. Wei, X.; Yu, J.; and Huang, Y. 2023. Physically Adversarial Infrared Patches with Learnable Shapes and Locations.

In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 12334–12342. Xie, P.; Bie, Y.; Mao, J.; Song, Y.; Wang, Y.; Chen, H.; and Chen, K. 2025. Chain of Attack: On the Robustness of Vision-Language Models against Transfer-Based Adversarial Attacks. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 14679–14689. Xu, H.; Xie, S.; Tan, X. E.; Huang, P.-Y.; Howes, R.; Sharma, V.; Li, S.-W.; Ghosh, G.; Zettlemoyer, L.; and Feichtenhofer, C. 2024. Demystifying CLIP Data. In The Twelfth International Conference on Learning Representations. Yin, Z.; Ye, M.; Zhang, T.; Du, T.; Zhu, J.; Liu, H.; Chen, J.; Wang, T.; and Ma, F. 2023. VLATTACK: Multimodal Adversarial Attacks on Vision-Language Tasks via Pretrained Models. In Advances in Neural Information Processing Systems, volume 36. Zhai, X.; Mustafa, B.; Kolesnikov, A.; and Beyer, L. 2023. Sigmoid Loss for Language Image Pre-Training. In Proceedings of the IEEE/CVF International Conference on Computer Vision, 11975–11986. Zhang, J.; Ye, J.; Ma, X.; Li, Y.; Yang, Y.; Chen, Y.; Sang, J.; and Yeung, D.-Y. 2025a. AnyAttack: Towards Large-Scale Self-Supervised Adversarial Attacks on Vision-Language Models. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 19900–19909. Zhang, T.; Hong, Y.; Xia, Y.; Ding, K.; Zhang, Z.; Wang, Y.; Xiang, S.; and Pan, C. 2025b. IF-Bench: Benchmarking and Enhancing MLLMs for Infrared Images with Generative Visual Prompting. arXiv preprint arXiv:2512.09663. Zhao, Y.; Pang, T.; Du, C.; Yang, X.; Li, C.; Cheung, N.-M.; and Lin, M. 2023. On Evaluating Adversarial Robustness of Large Vision-Language Models. In Advances in Neural Information Processing Systems, volume 36. Zheng, L.; Chiang, W.-L.; Sheng, Y.; Zhuang, S.; Wu, Z.; Zhuang, Y.; Lin, Z.; Li, Z.; Li, D.; Xing, E. P.; Zhang, H.; Gonzalez, J. E.; and Stoica, I. 2023. Judging LLM-as-aJudge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems, volume 36. Zhu, X.; Liu, Y.; Hu, Z.; Li, J.; and Hu, X. 2024. Infrared Adversarial Car Stickers. In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 24284–24293.

Record · ID 422292 · SHA-256 e86b646feaa6f954
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.