ConceptioArchivearXiv CS
arXiv CSopen access

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
machine learning, deep learning, neural networks

PHANTOM: A L ARGE -S CALE DATASET OF M ULTIMODAL A DVERSARIAL ATTACKS FOR V ISION -L ANGUAGE M ODELS

Simone Gallivanone1,∗ , Hossein Khodadadi1,∗ , Mauro Dore2 , Mauro Medda2 , Nicola Franco1

arXiv:2606.24388v1 [cs.AI] 23 Jun 2026

1

The Italian Institute of Artificial Intelligence (AI4I), Turin, Italy 2 HikmaAI S.r.l., Pula, Italy

A BSTRACT We introduce a large-scale, open-source dataset of pre-generated adversarial attacks for vision–language models (VLMs). The dataset is designed to be diverse, representative, and practical, extending existing benchmarks by covering 10 high-level categories and 55 subcategories of harmful intents. Our primary goal is to make adversarial data accessible to the research community, given the computational cost and complexity of generating large numbers of attacks. The dataset comprises 47 524 adversarial samples, generated using state-of-the-art attack strategies from recent literature. Our work complements existing efforts by consolidating and extending prior benchmarks from multiple established sources, resulting in 7 826 intents, and introduce an additional category to broaden coverage. This provides realistic evaluation resources for studying model robustness and alignment. Our dataset intends to enable researchers and practitioners to systematically evaluate the robustness and safety of VLMs, fine-tune attack-generation models, and develop or stress-test defensive guardrails under diverse adversarial conditions. By releasing this resource, we aim to lower the barrier to adversarial research and foster more reproducible, comprehensive, and comparable evaluations of VLM safety. The dataset has been released at: https://huggingface.co/datasets/it4lia/PHANTOM Disclaimer: This paper and dataset contain content that may be disturbing or offensive, included solely for research purposes.

1

Introduction

With the rapid public deployment of vision–language models (VLMs) in both open- and closed-source settings, including safety-critical and user-facing applications, their robustness against adversarial prompting has become an increasingly important research concern (see e.g., [1, 2, 3, 4, 5]). Recent studies consistently show that, despite improved alignment and scaling, state-of-the-art multimodal models remain vulnerable to carefully crafted jailbreak attacks, particularly when harmful intents are distributed across visual and textual modalities (see e.g., [6, 7, 8, 9]). Unlike unimodal settings, multimodal safety violations often exploit cross-modal reasoning and semantic alignment, significantly expanding the attack surface and complicating both detection and defense. As a consequence, evaluating the robustness of VLMs requires large and diverse collections of adversarial image–text pairs. This cost particularly affects resource-constrained research groups and practitioners, for whom reproducing large-scale multimodal attack generation may be impractical. This is particularly true for vision-language models, where attack generation is typically more resource-consuming than in unimodal settings. Unlike image-only or text-only attacks, multimodal attacks may require optimizing perturbations across multiple input spaces while preserving or exploiting their semantic alignment. As a result, each attack iteration can involve forward and backward passes through multiple modality-specific encoders and the cross-modal alignment module, and the overall search space becomes larger and more constrained. Although the exact overhead is model and attack dependent, the computational cost can be approximated as scaling with the combined cost of the involved modalities. This makes systematic adversarial attack generation especially demanding * Equal contribution.

Corresponding author: [email protected]

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

RISK TAXONOMY 10 categories · 55 subcategories · 7 826 intents

A Ethical & Social

B Privacy & Data

C Safety & Physical Harm

D Criminal & Economic

E Cybersecurity Threats

F Info & Political

G Content & Cultural

H IP & Ownership

I Decision & Cognitive

J Child Safety (new)

ATTACK STRATEGIES

intent selection

BAP

Attack generation

IDEATOR

attack selection

tested against

MML 47 524 pairs

FC ATTACK CSDJ

 OPEN-SOURCE white-box

 CLOSED-SOURCE black-box

DeepSeek-VL2

GPT-5.4

GLM-4.6V-Flash

GPT-5.5

TRANSFER

Kimi-VL-A3B

Gemini 3.1 Pro

Qwen3-VL-30B

Claude Opus 4.6

Qwen3.5-27B

Claude Opus 4.7

Qwen3.6-27B

Claude Opus 4.8

Gemma-4-26B

Attack samples

LLaVA-v1.6-13B Ministral-3-14B

Perturbed

Flowchart

Flipped

Generated

u Judge ⇒ ASR per category

Collage

Figure 1: Overview of PHANTOM. The risk taxonomy (10 categories, 55 subcategories, 7 826 intents) defines the dataset; the multimodal attacks (BAP, IDEATOR, MML, FC ATTACK, CSDJ) turn each intent into a multimodal adversarial sample (harmful text prompt + image), giving 47 524 (prompt, image) pairs. Each pair is given white-box to nine open-source VLMs and transferred to six closed-source, black-box models; the judge scores every response to obtain per-category ASR. Bottom-left: representative samples from different attack families.

for resource-constrained actors. While many existing open-source benchmarks (e.g., [4, 10, 11, 12, 13, 14, 15, 1, 16, 17, 3, 2, 7, 18]) provide tools and pipelines to generate and evaluate adversarial attacks, they typically do not release large collections of ready-to-use adversarial samples. Only a limited number of datasets offer such pre-generated attacks (e.g., [6, 19, 7, 2, 12]), often focusing on specific attack types, categories, or linguistic settings. In this work, we aim to complement existing efforts by releasing a large-scale collection of ready-to-use multimodal adversarial samples, covering a broader range of attack strategies and safety categories. Our goal is not to replace prior benchmarks, but to provide a practical resource that lowers the barrier to safety evaluation and enables reproducible and comprehensive robustness testing of multimodal models. With this in mind, we designed and produced the PHANTOM dataset, a dataset of adversarial attacks for vision-language models, which aims to fill this gap, and thus lower the barrier to systematic robustness evaluation. The dataset contains attack samples in the form of image–text pairs for both single-turn and conversational attacks. The dataset at a glance: Scale: Coverage: Strategies:

47 524 pre-generated attack samples 10 categories, 55 subcategories (appendix A) BAP [8], IDEATOR [6], MML [9], FC ATTACK [20], CSDJ [21].

For a more detailed discussion on the design and content of the dataset we refer the reader to section 3. The samples have been generated against a variety of different open-source models, from the following families: Qwen3-VL [22], DeepSeek-VL22 [23], GLM-4.6V [24], Kimi-VL [25], Qwen3.5 [26], Qwen3.6 [27]. The generated samples were subsequently evaluated against state-of-the-art proprietary models, including Claude Opus 4.6 - 4.7 - 4.8, GPT-5.4 5.5, Gemini-3.1-pro. The results, which highlight cross-model transferability and robustness trends, are presented in section 3.4. 2

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Our main contributions are: • PHANTOM, a large-scale open-source dataset of multimodal adversarial attacks for VLM safety evaluation. • A curated taxonomy of 7 826 harmful intents spanning 10 categories and 55 subcategories. • 47 524 adversarial samples generated using four attack strategies: BAP [8], IDEATOR [6], MML [9], FC ATTACK [20] and CSDJ [21]. • A transferability analysis across both open-source and proprietary VLMs. • Structured metadata designed to support reproducibility, benchmarking, and downstream safety research. The paper is organized as follows. Section 2 reviews related work. Section 3 describes the dataset design, generation pipeline, and evaluation protocol. Section 4 discusses the limitations of the current release, and section 5 addresses ethical considerations.

2

Related Works

In the landscape of adversarial attacks and model robustness, numerous efforts have benchmarked sensitive categories and created attack datasets across vision-language models (VLMs) to evaluate vulnerabilities and establish foundations for model alignment. To facilitate this review, we formalize the evaluation framework as a tuple E = (C, B, A, J ). Let M denote the target model which generates a response r ∈ R from an image-text input pair (I, T ). • Categories (C): A set of n sensitive domains C = {c1 , . . . , cn } where model output must be constrained to ensure safety. S • Intents / Behaviors (B): A set of specific harmful intents B = c∈C Bc , where each b ∈ Bc represents a concrete instance of a harmful objective within category c. • Adversarial Attacks (A): A set of functions f ∈ A that map a benign input to an adversarial input (I ′ , T ′ ), optimized to exploit model misalignment such that M(I ′ , T ′ ) aligns with a target behavior b. • Judge (J ): A classifier J : R × B → {0, 1} that maps a model response r and an intent b to a binary success metric, where J (r, b) = 1 indicates a successful adversarial exploit. Appendix D summarizes the chronological evolution of adversarial attack benchmarks. Detailed below, we review how different adversarial attack datasets included in these benchmarks or independently released, have evolved. 2.1

Evolution of Early Multimodal Adversarial Attack Datasets

The study of adversarial attacks on language and multimodal models has evolved through a series of increasingly comprehensive datasets. Early work by VAJM [16] introduced a dataset of 32 226 samples, focusing on degradations related to gender, race, and human identity. These included visual adversarial examples derived from 40 behavioral categories, with attacks primarily generated through prompt tuning techniques. Subsequent efforts expanded both the scale and diversity of attacks. The JailBreakV-28K [19] dataset applied attacks not only to initial harmful prompts but also to broader behavioral patterns. It includes 20 000 text-based jailbreak prompts and 8 000 image-based examples. These attacks are derived from the RedTeam2K [19] benchmark, which covers approximately 2 000 behaviors across 16 categories. The textual attacks were generated using methods such as GCG [17], Cognitive Overload, real-world jailbreak prompt templates, and PAP [28], while the visual attacks leverage Stable Diffusion and typographic image techniques. The MM-SafetyBench [2] dataset further advances multimodal evaluation by introducing 5 040 text–image pairs derived from 1 680 behaviors across 13 categories. In parallel, the Multiturn Human Jailbreaks [14] dataset explores iterative attack strategies, comprising 2 912 attacks generated using a combination of automated methods, including AutoDAN [29], AutoPrompt [30], GCG [17], GPTFuzzer [31], and PAIR [32]. SafeBench [33] extends the evaluation setting by incorporating 9 200 samples, including 2 300 multimodal pairs, and introduces the audio modality. It evaluates models under both adversarial and non-adversarial conditions, using attack strategies such as LPT [34], PAP [28], and BAP[8]. Notably, it is designed to assess safety risks even in the absence of explicit attacks. The MMJ dataset, derived from the MMJ benchmark [3], includes 1 000 adversarial examples generated using methods such as FigStep [7], MM-SafetyBench [2], HADES [18], ADV-16 [35], ADV-64 [35], ADV-inf [35], ImgJP [36], and 3

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

AttackVLM [37]. This work highlights a critical limitation of overly conservative defenses, arguing that a system that refuses all prompts is not practically useful. BAVI-Bench [12] significantly scales adversarial evaluation, containing 316k adversarial visual-instruction samples. It includes four types of image-based B-AVIs, ten types of text-based B-AVIs, and nine categories of content bias (e.g., gender, violence, cultural, and racial biases). The benchmark evaluates robustness using attacks such as PAR [38], Boundary [39], and SurFree [40]. The VLJailBreak benchmark [6] provides 3 654 samples spanning 12 safety topics and 46 subcategories, offering a highly structured categorization. It evaluates model vulnerabilities using GCG [17] and UMK [41] attack methods. Finally, the Adversarial Humanities Benchmark [42] investigates whether safety mechanisms generalize beyond familiar harmful prompt patterns. It includes 3 600 attack samples and shows that current safety techniques exhibit limited generalization, suggesting that a deeper understanding of non-maleficence remains an open challenge in frontier model safety. Complementarily, MultiBreak [5] focuses on realistic multi-turn jailbreak scenarios, where harmful intents are progressively elicited through conversation rather than expressed in a single prompt. It introduces 10 389 multi-turn adversarial prompts spanning 2 665 harmful intents, and shows that diverse multi-turn attacks can reveal fine-grained vulnerabilities that may remain hidden under single-turn evaluation. Together, these works highlight the need for safety benchmarks that go beyond static or template-based attacks, covering both broader semantic generalization and more realistic conversational adversarial settings. PHANTOM, on the other hand, identifies three key limitations in existing datasets. First, the number of intents used to generate attacks is too limited to adequately represent the full range of categories; accordingly, PHANTOM expands this to 7 826 intents, as detailed in table 2. Second, it examines how state-of-the-art attacks perform on more recent models that are commonly used in industrial and research settings as shown in table 3. Third, it investigates how vulnerabilities discovered through prior attacks and models can be transferred to other, primarily black-box, models.

3

Dataset design and production

In this section, we describe the dataset design and the process used to generate its samples. Specifically, we outline the category and subcategory structure, as well as the JSON-based intent specification employed during dataset construction. Settings. All experiments were conducted on a cluster, using NVIDIA A100 GPUs with 64GB of memory and Intel® Xeon® Platinum 8358 CPUs operating at 2.60GHz. In addition to the selected attack strategy, generated samples must be evaluated to determine whether they constitute successful attacks. For the sake of reproducibility and to avoid bias stemming from ad hoc judgment criteria, we opted to rely on a publicly available automated judge as a common baseline. In particular, we adopted the Abel-24-HarmClassifier proposed in [43]. This choice was motivated by the increasing difficulty of fully and cleanly jailbreaking recent models, which often respond with partially harmful or evasive outputs. Consequently, we selected a recent and aligned classifier to provide a more reliable assessment of attack harmfulness. 3.1

Categories structure

As mentioned in the introduction, we began our work by analyzing existing benchmarks for adversarial attacks. The landscape of adversarial attack benchmarks is quite extensive; however, many of these benchmarks build upon previous work, in the sense that they naturally extend earlier foundations. Two of the most influential works in this area are HarmBench [1] and AdvBench [17], which were designed as comprehensive collections of harmful intents and attack strategies for evaluating model safety.

Table 1: Number of intents per benchmark Benchmark # Intents JailBreakV_28k 1 953 MM-SafetyBench 1 671 OmniSafeBench 1 500 SafeBench 2 192 PHANTOM (ours) 7 826

With the evolution of AI models, their expanding range of use cases, and their increasing accessibility to the general public, these early benchmarks have gradually become insufficient in terms of coverage. As a result, several subsequent benchmarks have been proposed to address these limitations and extend prior efforts. In this work, we studied 16 benchmarks, including OmniSafeBench-MM [4], VLJailbreakBench [6], Sorry-Bench [11], B-AVIBench [12], JailbreakBench [10], SafeBench [13], Multi-Turn Human Jailbreaks [14], StrongREJECT [15], HarmBench [1], VAJM [16], AdvBench [17], MMJ-Bench [3], MM-SafetyBench [2], JailBreakV28K [19], FigStep [7], and HADES [18]. A gap we identified across these benchmarks is the limited number of behaviors, which makes it difficult to disentangle whether observed vulnerabilities stem from insufficient robustness to 4

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

specific categories or from the strength of the attacks themselves. To address this, we incorporated 7 826 behaviors into our study. Based on this analysis, we constructed a root dataset of harmful intents (also referred to as behaviors or goals) Table 2: Number of intents per category by merging recent benchmarks that were not designed Category # Intents as direct extensions of one another. The number of inA. Ethical and Social Risks 988 tents contributed by each benchmark and their distribuB. Privacy and Data Risks 504 tion across categories are reported in table 1 and table 2, C. Safety and Physical Harm 877 respectively. In appendix A the reader can find a table D. Criminal and Economic Risks 1 017 listing all categories, subcategories together with their E. Cybersecurity Threats 725 alphanumeric reference. Since [4] carried out an extenF. Information and Political Manipulation 534 sive effort to reorganize and categorize harmful intents, G. Content and Cultural Safety 537 while also providing a detailed taxonomy, we mapped H. Intellectual Property and Ownership 304 all merged intents onto the classification proposed there. I. Decision and Cognitive Risks 1 593 However, we identified a gap in its coverage: child safety. J. Child Safety (new) 747 The motivation for introducing this category is twofold. First, the widespread adoption of LLMs across users of all ages, including minors, raises the risk of exposure to content that may be harmful to their safety, such as methods to circumvent parental controls or age-verification systems. Second, given the particular vulnerability of minors, malicious actors could potentially exploit LLMs to obtain information on how to deceive or manipulate children in harmful situations. For these reasons, we believe that including this category contributes to a more comprehensive understanding of the risks associated with LLMs, and may support the development of more effective safeguards and training strategies. A. Ethical and Social Risks B. Privacy and Data Risks C. Safety and Physical Harm 100%

A4 (3%)

B5 (6%)

Categories

D. Criminal and Economic Risks E. Cybersecurity Threats F. Information and Political Manipulation C7 (3%) C6 (3%)

D5 (3%)

F6 (6%) E7 (18%)

B4 (6%) C5 (15%) 80%

C4 (3%)

A3 (42%)

F5 (6%) F4 (6%)

D4 (34%)

B3 (33%) Subcategory share within category (%)

G. Content and Cultural Safety H. Intellectual Property and Ownership

A2 (11%)

B2 (6%) D2 (24%)

G4 (24%)

E6 (4%) E5 (4%) E4 (4%) E3 (4%) G3 (28%)

I4 (19%) J3 (19%)

E2 (28%) F3 (72%)

H2 (10%)

I3 (23%)

G2 (26%)

20%

0%

A1 (44%)

A 12.6%

J2 (20%)

B1 (50%) D1 (36%)

B 6.4%

C2 (3%) C1 (3%) C 11.2%

J5 (17%)

J4 (22%)

H3 (42%)

C3 (68%)

40%

I8 (2%) I7 (7%) I6 (2%) I5 (18%)

D3 (3%)

60%

H4 (10%)

I. Decision and Cognitive Risks J. Child Safety

H1 (38%)

E1 (38%) F2 (6%)

G1 (22%)

I2 (8%)

I1 (20%)

J1 (23%)

I 20.4%

J 9.5%

F1 (6%) D 13.0%

E 9.3%

F 6.8%

G 6.9%

H 3.9%

100%

Figure 2: PHANTOM intents dataset. The horizontal axis shows the distribution of categories across the dataset, while the vertical axis shows the distribution of intents within each category, broken down by subcategory.

As a result, our dataset is organized into 10 high-level categories, further divided into a total of 55 subcategories. The main categories are: Ethical and Social Risks, Privacy and Data Risks, Safety and Physical Harm, Criminal and Economic Risks, Cybersecurity Threats, Information and Political Manipulation, Content and Cultural Safety, Intellectual Property and Ownership, Decision and Cognitive Risks, and Child Safety. The full subdivision into subcategories is illustrated in fig. 2; we refer the reader to appendix A for an overview of the names of subcategories with respect to their reference code. 5

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

We collected more than 7 000 harmful intents from the following benchmarks: JailBreakV_28K [19], MMSafetyBench [2], OmniSafeBench-MM [4], and SafeBench [33]. Moreover, we added 747 intents related to the new Child Safety category, generated with the assistance of OpenAI GPT-5.4, accessed via API. Since these benchmarks, together with our additions, may contain overlapping or semantically similar intents, we performed a cosine similarity analysis across all collected samples using embeddings generated by the sentence-transformers all-MiniLM-L6-v2 model (see [44]). We found that the overlap was not negligible, reaching values as high as 90%. Therefore, we decided to clean the dataset using this threshold, which left no pair of intents with a cosine similarity above 90%. At lower thresholds the number of intents flagged as similar grows: 376 (4.7%) at 85% and 661 (8.2%) at 80%. We decided not to apply more aggressive cleaning, as we observed that an 80% threshold often groups together semantically different intents, and we did not want to remove meaningful content. 3.2

Adversarial attacks

To benchmark the performance of state-of-the-art adversarial attacks against recent vision–language models, we initially selected a set of established attack strategies that had already proven effective on open-source vision–language systems. As shown in fig. 3, we analyze the attack success rate (ASR) as a function of generation time. To enable large-scale dataset generation, we ultimately focused on a limited subset of attacks offering the most favorable trade-off between ASR and computational cost. Specifically, we selected one single-turn attack strategy, the Bi-modal Adversarial Prompt (BAP) attack proposed in [8]; one multi-turn strategy, IDEATOR, introduced in [6]; and two more typographic oriented attacks the Multi-Modal Linkage Attack, proposed in [9] and the Flowchart attack in [20]. A review of the attack strategies can be found in appendix F.

100%

IDEATOR

BAP

FC ATTACK

ASR (%)

80%

MML

60%

MIDAS

CS-DJACZ-Attack

40%

VISCO 20%

HIMRD VISCARA

FigStep

QR

30s

50s

HADES

JOOD 0% 10s

20s

100s

200s 250s …

500s

1000s

2000s

Delay per Sample (seconds, log scale)

Figure 3: Evaluation of ASR based on Attack strategy and delay per attack The main motivation for selecting these strategies was their strong empirical performance. Before converging on this subset, however, we experimented with additional attack strategies, namely QR–attack [2], JOOD [45], CS-DJ [21], FigStep [7], HADES [18], HIMRD [46], VISCARA [47], MIDAS [48], ACZ attack [49]. The results of this preliminary analysis are reported in fig. 3, based on 50 samples generated against Qwen3.5-27B and evaluated with Abel-24-HarmClassifier [43]. Additional considerations that informed our final choice are discussed below. While preserving the original attack pipelines, we introduced several minor modifications to better suit our datasetgeneration process. In the following, we briefly describe how these methods were adapted and employed. In future releases, we intend to expand the range of attack strategies and target models in order to cover a broader spectrum of vulnerabilities. 6

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

3.3

Data structure and distribution

The dataset is structured according to the attack strategy and the target model used during the generation process. The target models considered are DeepSeek-VL22, GLM-4.6V-Flash, Kimi-VL-A3B-Instruct, Qwen3-VL-30B-A3BInstruct, Qwen3.5-27B and Qwen3.6-27B. Since the attacks target vision–language models, each sample consists of an image–prompt pair. Each folder contains the generated image along with a structured metadata file, which enables the correct association between images and prompts and ensures full reproducibility of the attacks. The current release contains a total of 47 524 generated attacks. Below, we discuss their distribution across target models, attack strategies, and categories. Regarding the distribution of attacks across target models, we initially generated attacks uniformly for all models. After conducting preliminary cross-model evaluations, we focused further generation on the models that exhibited higher attack transferability, namely GLM-4.6V-Flash and Qwen3-VL-30B-A3B-Instruct, Qwen3.5-27B and Qwen3.6-27B. We refer the reader to table 3 for a detailed breakdown. Table 3: Number of generated attacks per target model and attack strategy. Model DeepSeek-VL2 GLM-4.6V-Flash Kimi-VL-A3B-Instruct Qwen3-VL-30B-A3B-Instruct Qwen3.5-27B Qwen3.6-27B Total

Attack Strategy BAP

IDEATOR

MML

FC ATTACK

CSDJ

Total

835 2 922 871 4 739 1 613 2 090

265 343 282 219 355 425

1 066 1 533 809 3 874 2 106 3 895

1 492 1 411 1 482 1 368 1 265 1 249

817 2 059 1 733 2 574 1 575 2 257

4 475 8 268 5 177 12 774 6 914 9 916

13 070

1 889

13 283

8 267

11 015

47 524

From a categorical perspective, we selected intents randomly and uniformly from our dataset of intents, discussed in section 3.1. The distribution over categories and subcategories is shown in fig. 4. 3.4

Dataset tests and statistics

To validate our adversarial data generation pipeline, we conducted a series of experiments using the generated attacks. In particular, we considered two evaluation settings: white-box and black-box testing. In the black-box setting, we evaluated the attacks against several state-of-the-art proprietary models accessed via API, namely Gemini 3.1 Pro Preview, GPT-5.4, GPT-5.5, Claude Opus 4.6, Claude Opus 4.7, and Claude Opus 4.8. Model responses were collected and assessed using an automated judge. Cases in which a model returned an empty response were treated as hard refusals and therefore counted as failed jailbreak attempts. In the white-box setting, we focused primarily on the models used during adversarial generation, in order to evaluate the cross-model transferability of the attacks. Specifically, we tested DeepSeek-VL22, GLM-4.6V-Flash, Kimi-VL-A3BInstruct, Qwen3-VL-30B-A3B-Instruct, Qwen3.5-27B, Qwen3.6-27B, Gemma-4-26B-A4B-it, Llava-v1.6-vicuna-13bhf and Ministral-3-14B-Instruct-2512 all executed locally on NVIDIA A100 GPUs with 64 GB of memory. To keep the evaluation compact, we consider a fixed subset of 1,100 attacks for each attack strategy. This subset is selected once and reused across all models. The choice of an odd number is motivated by the need to evenly cover all subcategories associated with each attack strategy; specifically, we select 20 attacks per subcategory. As in the generation phase, we employed Abel-24-HarmClassifier [43] as the baseline judge for evaluating model responses. As evaluation metric, we relied on the widely used Attack Success Rate (ASR), defined as the percentage of successful jailbreak attempts, as determined by the judge, over the total number of attacks. Part of the result, relative to black-box and the most recent withe-box models are reported in table 4. The rest of the results can be found in appendix C. For each model and category, we report the corresponding ASR, i.e., the ratio of successful attacks to the number of attacks sampled in that category. This breakdown makes it possible to identify the categories for which each model is most vulnerable or most robust.We emphasize once more that these results should be interpreted as a reference relative to the chosen baseline for assessing harmfulness, rather than as an absolute ground truth.

7

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Qwen3.6-27B

1700

dS

l an

A.

ca thi

E

B.

s

s

isk

lR

a oci

ya

vac Pri

nd

ta Da

C.

nd

ya fet Sa

s

arm

k Ris

H cal ysi

Ph

D.

nd al a

on

Ec

isk ic R om E.

in

im

Cr

cu

e ers

b Cy

rity

nd

na

atio orm

cal

liti Po

M G.

nf F. I

ul dC

an ent

l tua

llec

te . In

ty

892

560

711

750

254

201

fet Sa

per Pro

416

356 239

424 al tur

t

n Co

288 y

on

lati

pu ani

705 791

916

ts

rea

Th

603

667

598

640 501

453

215 0

1035

1167 922

756 597

689

1227

1348

1421 1139

1036

1143 661

585

666

653

620

1467

1653

1597 1187

1033

1251

541

641 433

363

480

454

396

1169

1280 862

1000

858

Attack Count

1500

500

Qwen3-VL-30B-A3B-Instruct

1821

2000

Kimi-VL-A3B-Instruct

735

GLM-4.6V-Flash

Qwen3.5-27B

and

n Ow

s

ip

k Ris ive nit

h ers

nd na sio i c e

g

Co

458

DeepSeek-VL2

300

Model

y

d hil

fet Sa

J. C

I. D

H

Category

Figure 4: Category coverage: overall and in-category Table 4: Table of ASR (%) per model and per category across all attacks, for category names refer to appendix A Model

A

B

C

Gemma-4-26B Qwen3.6-27B GPT-5.4 GPT-5.5 Claude Opus 4.6 Claude Opus 4.7 Claude Opus 4.8 Gemini 3.1 Pro

23.75 42.50 13.75 47.50 6.25 51.25 47.50 23.75

16.00 40.00 15.00 41.00 11.00 44.00 40.00 39.00

10.00 30.71 25.00 59.29 0.71 20.71 12.14 35.71

Gemma-4-26B Qwen3.6-27B GPT-5.4 GPT-5.5 Claude Opus 4.6 Claude Opus 4.7 Claude Opus 4.8 Gemini 3.1 Pro

10.00 7.50 25.00 21.25 6.25 25.00 16.25 28.75

30.00 7.00 14.00 28.00 6.00 39.00 24.00 43.00

20.71 11.43 27.14 26.43 5.71 12.14 12.14 42.14

Gemma-4-26B

95.56

100.00

90.34

D

E F BAP 18.00 27.14 20.00 42.00 38.57 35.83 23.00 11.43 18.33 68.00 47.86 42.50 10.00 8.57 5.83 61.00 40.00 39.17 65.00 39.29 34.17 45.00 74.29 28.33 IDEATOR 34.00 33.57 20.00 14.00 15.00 16.67 49.00 44.29 36.67 42.00 48.57 44.17 17.00 13.57 5.83 53.00 30.71 30.00 47.00 32.86 24.17 56.00 71.43 58.33 MML 98.51 99.37 96.88

8

G

H

I

J

Average

22.50 33.75 10.00 46.25 8.75 35.00 21.25 30.00

22.50 40.00 6.25 35.00 7.50 56.25 41.25 27.50

26.88 45.00 10.62 48.12 6.88 50.00 43.12 25.62

34.00 51.00 23.00 50.00 17.00 49.00 51.00 44.00

22.00 39.82 15.91 49.09 7.91 43.64 38.73 38.36

15.00 18.75 26.25 52.50 3.75 32.50 17.50 56.25

21.25 7.50 12.50 31.25 5.00 45.00 33.75 42.50

35.62 23.12 27.50 37.50 13.12 34.38 31.25 54.37

43.00 65.00 33.00 44.00 7.00 33.00 20.00 62.00

27.36 18.82 30.45 37.82 8.82 32.55 26.09 52.64

94.87

100.00

96.55

95.26

96.09

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Model Qwen3.6-27B GPT-5.4 GPT-5.5 Claude Opus 4.6 Claude Opus 4.7 Claude Opus 4.8 Gemini 3.1 Pro

A 80.00 71.76 91.25 71.95 1.25 12.50 3.53

B 85.71 73.45 84.00 61.32 10.00 21.00 0.88

C 80.97 74.13 90.71 32.14 0.71 3.57 4.20

Gemma-4-26B Qwen3.6-27B GPT-5.4 GPT-5.5 Claude Opus 4.6 Claude Opus 4.7 Claude Opus 4.8 Gemini 3.1 Pro

37.50 86.25 52.50 57.50 77.50 70.00 63.75 35.00

29.00 81.00 43.00 57.00 85.00 72.00 83.00 46.00

2.14 55.71 60.71 68.57 41.43 38.57 48.57 15.00

Gemma-4-26B Qwen3.6-27B GPT-5.4 GPT-5.5 Claude Opus 4.6 Claude Opus 4.7 Claude Opus 4.8 Gemini 3.1 Pro

53.75 71.25 78.75 82.50 73.75 70.00 80.00 65.00

75.00 68.00 78.00 80.00 88.00 90.00 89.00 86.00

27.86 64.29 91.43 88.57 67.86 57.14 71.43 39.29

D E F 89.63 94.55 87.50 67.92 70.00 72.79 92.00 90.71 89.17 60.78 63.95 63.78 6.00 2.86 0.83 26.00 10.00 11.67 0.94 40.62 81.62 FC ATTACK 7.00 22.14 11.67 85.00 92.86 82.50 71.00 77.14 74.17 86.00 89.29 76.67 68.00 87.86 70.83 75.00 68.57 65.00 77.00 79.29 75.00 28.00 67.14 26.67 CSDJ 53.00 85.71 57.50 74.00 84.29 70.00 87.00 87.14 86.67 91.00 95.00 85.00 85.00 93.57 84.17 78.00 95.00 88.33 91.00 95.00 87.50 72.00 86.43 67.50

G 71.79 78.75 82.50 41.25 0.00 11.25 66.25

H 86.21 80.46 80.00 48.81 5.00 21.25 93.10

I 83.14 80.00 81.25 62.86 1.88 13.75 93.71

J 88.42 95.00 61.00 54.55 1.00 23.00 29.00

Average 86.10 76.03 84.64 56.19 2.82 14.64 43.38

33.75 72.50 46.25 67.50 63.75 37.50 36.25 35.00

28.75 76.25 47.50 63.75 82.50 82.50 78.75 55.00

18.75 78.12 59.38 65.00 75.62 68.75 62.50 40.62

14.00 64.00 63.00 74.00 60.00 61.00 70.00 21.00

18.91 77.27 61.00 71.36 70.82 63.45 67.45 37.00

27.50 70.00 66.25 81.25 48.75 50.00 50.00 46.25

41.25 60.00 61.25 70.00 71.25 71.25 71.25 68.75

36.88 61.88 63.75 69.38 73.75 59.38 59.38 63.12

55.00 70.00 81.00 87.00 83.00 88.00 89.00 68.00

51.35 69.37 78.82 83.16 77.82 74.82 78.45 66.18

While the weakness of a model with respect to a given attack gives an interesting insight in how to chose the attack strategy, one may be interested in global weaknesses of a model. To approximate such information we averaged the results across attack strategies and reported the results in fig. 5. As discussed, the 7 826 intents in the PHANTOM benchmark provide sufficient coverage across categories to confidently analyze model vulnerabilities with respect to these specific domains. The radar diagrams provide an immediate, at-a-glance understanding of model robustness: a larger colored area corresponds to a greater amount of harmful content produced during testing. However, an important clarification is needed. In modern models, it is difficult to observe “pure” jailbreaks, i.e., cases in which the model responds directly and fully to a harmful request. Instead, harmful content is more often embedded within longer responses that include benign context and argumentation. Therefore, these diagrams should be interpreted as indicating a higher tendency of the model to generate harmful content within otherwise complex answers. With this in mind, models that tend to respond to user requests, even while attempting to avoid harmful content, ultimately produce more harmful content on average. This explains, for example, the stronger performance of Gemma4-26B compared to many black-box models, which tend to consistently provide answers. On the other hand, it is important to note that black-box models accessed via API sometimes return null responses, most likely due to content filtering mechanisms; we refer to these as hard refusals. We treat such cases as failed jailbreaks. Different models exhibit different rates of hard refusals: for instance, the Claude Opus models show a much higher rate of hard refusals compared to both GPT and Gemini, whereas GPT models exhibit the lowest rate. We now highlight a few observations from the results. In fig. 5, the extent of the colored area allows one to infer, with respect to the chosen baseline judge, the relative robustness of the models: a wider area corresponds to less aligned responses. Among white-box models, Gemma-4-26B is clearly the most robust. Among black-box models, the picture is different: all models in the Opus family exhibit comparable robustness, which is also similar to that of Gemini 3.1 Pro, while the GPT family appears less robust. However, as noted earlier, this should be interpreted alongside the higher rate of complete responses they produce. Interestingly, within the GPT family (from 5.4 to 5.5), performance in terms of alignment appears to degrade slightly, although this is again coupled with the absence of hard refusals. Another interesting observation from Table 4 emerges from the distribution of colored cells: the most effective attacks across all models are those that embed harmful text within images. This suggests that model alignment with respect to embedded textual content remains relatively weak, highlighting a persistent vulnerability in multimodal safety mechanisms. Finally, across all models, the most vulnerable categories are D — Criminal and Economic Risks and E — Cybersecurity Threats, which also correspond to domains where one would expect models to provide more actionable and useful responses. 9

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

A: Ethical & Social B: Privacy & Data C: Safety & Physical D: Criminal & Economic E: Cybersecurity F: Info & Political G: Content & Cultural H: IP & Ownership I: Decision & Cognitive J: Child Safety Qwen3.6-27B

Gemma-4-26B

D

D

C

GPT-5.4

C

D

GPT-5.5 C

D

C

75.8 60.9

E 65.0

41.2

F

38.7

67.7 54.0

42.1

53.6

57.5 A 0 20 40 60 80 100

53.4

42.8

B 30.2

E

J

44.1 A 0 20 40 60 80 100 48.2 42.9

H

G

H

B

E

48.4 A 0 20 40 60 80 100

45.5 41.6

B

74.3

58.0

67.5

F

59.0

63.2 56.0

J

H

60.0 A 0 20 40 60 80 100

66.0

48.2

G

I

66.7

44.7

57.7

F

J

I

55.7

58.0

50.0

58.2

G

60.2

G

J

I

H

I

Claude Opus 4.6

Claude Opus 4.7

Claude Opus 4.8

Gemini 3.1 Pro

D

D

D

D

E

F

E

56.3

58.5

F

B

48.6

59.6

C

B

48.2 29.6

46.1

47.1 A 0 20 40 60 80 100 44.3

33.2 43.0

54.6

E

53.5

50.3

44.7

F

31.0 52.0

G

J

B

43.5 A 0 20 40 60 80 100 46.4 42.9

G

I

J

H

F

C 61.2

E

51.0

25.9

47.4

46.5

H

C

B

51.3

29.6

46.5

44.0 A 0 20 40 60 80 100 50.6 42.0

27.2 49.2

52.5

F

J

H

40.4

68.0

I

B 27.3 43.0 31.2 A 0 20 40 60 80 100 44.8

46.8 57.4

G

I

E

51.4

C

55.5

G

J

H

I

Figure 5: Examining model vulnerability against harmful categories

4

Limitations

PHANTOM has several limitations. First, our evaluation relies primarily on an automated judge, Abel-24HarmClassifier, which may introduce false positives and false negatives, particularly for responses that are partially harmful, evasive, or context-dependent. Although automated judging enables large-scale evaluation, it cannot fully replace human assessment. Second, our reported evaluation is conducted on sampled subsets of attacks rather than on the entire released dataset. While this makes the evaluation computationally feasible, it may underrepresent variability. Third, currently we focused on three multimodal attack strategies: BAP, IDEATOR, and MML. These methods were selected for their empirical effectiveness and computational feasibility, but they do not exhaust the space of possible adversarial attacks against VLMs. Finally, the taxonomy and intent collection may inherit biases from the source benchmarks used to construct the dataset.

5

Ethical Considerations

PHANTOM dataset contains adversarial multimodal samples involving harmful and sensitive intents, and is therefore a dual-use resource. The dataset is intended solely for research, robustness evaluation, and the development of defensive guardrails. To reduce misuse risks, we provide content warnings, structured metadata, category labels. Sensitive categories, including child safety and personally harmful content, are included only for safety evaluation and should be handled under appropriate institutional and ethical safeguards.

6

Discussion and conclusion

Our empirical analysis highlights several key insights into the behavior of modern VLMs under adversarial conditions. 10

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

First, despite advances in safety training, all evaluated systems exhibit non-negligible ASR across multiple categories, confirming that alignment remains fragile in the presence of carefully constructed multimodal inputs. Second, we observe significant variation across attack strategies. This suggests that different strategies exploit distinct failure modes of VLMs across categories of harmful content. Consequently, evaluating robustness using a single attack family risks underestimating model vulnerability. In particular, attacks that embed harmful requests directly within images (i.e., typographic attacks such as MML, FC Attack, and CSDJ) achieve consistently higher success rates, indicating that even recent large-scale models still struggle to maintain strong filtering capabilities in fully multimodal settings. A particularly surprising result comes from the Gemma-4-26B model: when attacked, it achieves an ASR of 100% in two categories, highlighting a substantial vulnerability despite the model’s overall capabilities. Third, the results reveal clear evidence of cross-model transferability, particularly from open-source to proprietary systems. Attacks generated in a white-box setting retain their effectiveness when transferred to black-box models, with ASR remaining above 20% in most cases and reaching peaks of nearly 80% for the CSDJ attack. This indicates that vulnerabilities are not purely model-specific but instead reflect shared structural or training-induced weaknesses. These findings have important implications for real-world deployment, where adversaries may optimize attacks against accessible models and subsequently transfer them to closed systems. A further observation is the heterogeneity across risk categories. Certain domains, such as cybersecurity or economic crimes, tend to yield higher attack success rates, reaching up to 90% on black-box models, while others are more robust. This variability suggests that current alignment procedures may unevenly cover the safety landscape, leaving gaps that adversarial methods can exploit. These findings are consistent with the broader trend illustrated in the evaluation results. Importantly, our analysis also highlights a well known evaluation caveat: success rates depend on the chosen judge model and may be influenced by partial refusals due to external filters or ambiguous outputs (see also, [50], [51]). As such, the results should be interpreted as relative indicators of robustness, rather than absolute measures of harmfulness. Overall, PHANTOM enables a more systematic understanding of how different attack strategies, model architectures, and safety domains interact, offering a baseline to study multimodal robustness. By consolidating multiple attack strategies and providing structured evaluation across a diverse set of VLMs, the dataset addresses a key gap in the current landscape: the lack of accessible, reproducible, and comprehensive adversarial resources. Our results demonstrate that multimodal jailbreaks remain a persistent and transferable threat, that robustness varies significantly across both models and safety domains, and that a diverse set of attack strategies is necessary for reliable evaluation. Beyond benchmarking, PHANTOM provides a practical foundation for future research: the dataset can be used to develop and evaluate defensive mechanisms and guardrails, to train adversarially robust models, and to advance the study of cross-modal alignment failures. We release PHANTOM with the goal of lowering the barrier to multimodal safety research and fostering more reproducible, standardized, and comprehensive evaluations. We hope that this resource will contribute to a deeper understanding of VLM robustness and support the development of safer and more reliable multimodal AI systems.

11

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

References [1] Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249, 2024. [2] Xin Liu, Yichen Zhu, Jindong Gu, Yunshi Lan, Chao Yang, and Yu Qiao. Mm-safetybench: A benchmark for safety evaluation of multimodal large language models. In European Conference on Computer Vision, pages 386–403. Springer, 2024. [3] Fenghua Weng, Yue Xu, Chengyan Fu, and Wenjie Wang. Mmj-bench: A comprehensive study on jailbreak attacks and defenses for vision language models. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 27689–27697, 2025. [4] Xiaojun Jia, Jie Liao, Qi Guo, Teng Ma, Simeng Qin, Ranjie Duan, Tianlin Li, Yihao Huang, Zhitao Zeng, Dongxian Wu, Yiming Li, Wenqi Ren, Xiaochun Cao, and Yang Liu. Omnisafebench-mm: A unified benchmark and toolbox for multimodal jailbreak attack-defense evaluation, 2025. URL https://arxiv.org/abs/2512. 06589. [5] Jialin Song, Xiaodong Liu, Weiwei Yang, Wuyang Chen, Mingqian Feng, Xuekai Zhu, and Jianfeng Gao. Multibreak: A scalable and diverse multi-turn jailbreak benchmark for evaluating llm safety, 2026. URL https://arxiv.org/abs/2605.01687. [6] Ruofan Wang, Juncheng Li, Yixu Wang, Bo Wang, Xiaosen Wang, Yan Teng, Yingchun Wang, Xingjun Ma, and Yu-Gang Jiang. Ideator: Jailbreaking and benchmarking large vision-language models using themselves, 2025. URL https://arxiv.org/abs/2411.00827. [7] Yichen Gong, Delong Ran, Jinyuan Liu, Conglei Wang, Tianshuo Cong, Anyu Wang, Sisi Duan, and Xiaoyun Wang. Figstep: Jailbreaking large vision-language models via typographic visual prompts. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 39, pages 23951–23959, 2025. [8] Zonghao Ying, Aishan Liu, Tianyuan Zhang, Zhengmin Yu, Siyuan Liang, Xianglong Liu, and Dacheng Tao. Jailbreak vision language models via bi-modal adversarial prompt. IEEE Transactions on Information Forensics and Security, 20:7153–7165, 2025. doi:10.1109/TIFS.2025.3583249. [9] Yu Wang, Xiaofei Zhou, Yichen Wang, Geyuan Zhang, and Tianxing He. Jailbreak large vision-language models through multi-modal linkage. In Wanxiang Che, Joyce Nabende, Ekaterina Shutova, and Mohammad Taher Pilehvar, editors, Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 1466–1494, Vienna, Austria, July 2025. Association for Computational Linguistics. ISBN 979-8-89176-251-0. doi:10.18653/v1/2025.acl-long.74. URL https://aclanthology.org/2025.acl-long. 74/. [10] Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. Advances in Neural Information Processing Systems, 37:55005–55029, 2024. [11] Tinghao Xie, Xiangyu Qi, Yi Zeng, Yangsibo Huang, Udari Madhushani Sehwag, Kaixuan Huang, Luxi He, Boyi Wei, Dacheng Li, Ying Sheng, et al. Sorry-bench: Systematically evaluating large language model safety refusal. arXiv preprint arXiv:2406.14598, 2024. [12] Hao Zhang, Wenqi Shao, Hong Liu, Yongqiang Ma, Ping Luo, Yu Qiao, Nanning Zheng, and Kaipeng Zhang. B-avibench: Toward evaluating the robustness of large vision-language model on black-box adversarial visualinstructions. IEEE Transactions on Information Forensics and Security, 20:1434–1446, 2024. [13] Zonghao Ying, Aishan Liu, Siyuan Liang, Lei Huang, Jinyang Guo, Wenbo Zhou, Xianglong Liu, and Dacheng Tao. Safebench: A safety evaluation framework for multimodal large language models. International Journal of Computer Vision, 134(1):18, 2026. [14] Nathaniel Li, Ziwen Han, Ian Steneker, Willow Primack, Riley Goodside, Hugh Zhang, Zifan Wang, Cristina Menghini, and Summer Yue. Llm defenses are not robust to multi-turn human jailbreaks yet. arXiv preprint arXiv:2408.15221, 2024. [15] Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al. A strongreject for empty jailbreaks. Advances in Neural Information Processing Systems, 37:125416–125440, 2024. [16] Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. Visual adversarial examples jailbreak aligned large language models. In Proceedings of the AAAI conference on artificial intelligence, volume 38, pages 21527–21536, 2024. 12

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

[17] Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023. [18] Yifan Li, Hangyu Guo, Kun Zhou, Wayne Xin Zhao, and Ji-Rong Wen. Images are achilles’ heel of alignment: Exploiting visual vulnerabilities for jailbreaking multimodal large language models. In European Conference on Computer Vision, pages 174–189. Springer, 2024. [19] Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks. arXiv preprint arXiv:2404.03027, 2024. [20] Ziyi Zhang, Zhen Sun, Zongmin Zhang, Jihui Guo, and Xinlei He. Fc-attack: Jailbreaking large vision-language models via auto-generated flowcharts. arXiv e-prints, pages arXiv–2502, 2025. [21] Zuopeng Yang, Jiluan Fan, Anli Yan, Erdun Gao, Xin Lin, Tao Li, Kanghua Mo, and Changyu Dong. Distraction is all you need for multimodal large language model jailbreaking. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 9467–9476, 2025. [22] Shuai Bai, Yuxuan Cai, Ruizhe Chen, Keqin Chen, Xionghui Chen, Zesen Cheng, Lianghao Deng, Wei Ding, Chang Gao, Chunjiang Ge, et al. Qwen3-vl technical report. arXiv preprint arXiv:2511.21631, 2025. [23] Zhiyu Wu, Xiaokang Chen, Zizheng Pan, Xingchao Liu, Wen Liu, Damai Dai, Huazuo Gao, Yiyang Ma, Chengyue Wu, Bingxuan Wang, et al. Deepseek-vl2: Mixture-of-experts vision-language models for advanced multimodal understanding. arXiv preprint arXiv:2412.10302, 2024. [24] V Team, Wenyi Hong, Wenmeng Yu, Xiaotao Gu, Guo Wang, Guobing Gan, Haomiao Tang, Jiale Cheng, Ji Qi, Junhui Ji, Lihang Pan, Shuaiqi Duan, Weihan Wang, Yan Wang, Yean Cheng, Zehai He, Zhe Su, Zhen Yang, Ziyang Pan, Aohan Zeng, Baoxu Wang, Bin Chen, Boyan Shi, Changyu Pang, Chenhui Zhang, Da Yin, Fan Yang, Guoqing Chen, Jiazheng Xu, Jiale Zhu, Jiali Chen, Jing Chen, Jinhao Chen, Jinghao Lin, Jinjiang Wang, Junjie Chen, Leqi Lei, Letian Gong, Leyi Pan, Mingdao Liu, Mingde Xu, Mingzhi Zhang, Qinkai Zheng, Sheng Yang, Shi Zhong, Shiyu Huang, Shuyuan Zhao, Siyan Xue, Shangqin Tu, Shengbiao Meng, Tianshu Zhang, Tianwei Luo, Tianxiang Hao, Tianyu Tong, Wenkai Li, Wei Jia, Xiao Liu, Xiaohan Zhang, Xin Lyu, Xinyue Fan, Xuancheng Huang, Yanling Wang, Yadong Xue, Yanfeng Wang, Yanzi Wang, Yifan An, Yifan Du, Yiming Shi, Yiheng Huang, Yilin Niu, Yuan Wang, Yuanchang Yue, Yuchen Li, Yutao Zhang, Yuting Wang, Yu Wang, Yuxuan Zhang, Zhao Xue, Zhenyu Hou, Zhengxiao Du, Zihan Wang, Peng Zhang, Debing Liu, Bin Xu, Juanzi Li, Minlie Huang, Yuxiao Dong, and Jie Tang. Glm-4.5v and glm-4.1v-thinking: Towards versatile multimodal reasoning with scalable reinforcement learning, 2025. URL https://arxiv.org/abs/2507.01006. [25] Kimi Team, Angang Du, Bohong Yin, Bowei Xing, Bowen Qu, Bowen Wang, Cheng Chen, Chenlin Zhang, Chenzhuang Du, Chu Wei, Congcong Wang, Dehao Zhang, Dikang Du, Dongliang Wang, Enming Yuan, Enzhe Lu, Fang Li, Flood Sung, Guangda Wei, Guokun Lai, Han Zhu, Hao Ding, Hao Hu, Hao Yang, Hao Zhang, Haoning Wu, Haotian Yao, Haoyu Lu, Heng Wang, Hongcheng Gao, Huabin Zheng, Jiaming Li, Jianlin Su, Jianzhou Wang, Jiaqi Deng, Jiezhong Qiu, Jin Xie, Jinhong Wang, Jingyuan Liu, Junjie Yan, Kun Ouyang, Liang Chen, Lin Sui, Longhui Yu, Mengfan Dong, Mengnan Dong, Nuo Xu, Pengyu Cheng, Qizheng Gu, Runjie Zhou, Shaowei Liu, Sihan Cao, Tao Yu, Tianhui Song, Tongtong Bai, Wei Song, Weiran He, Weixiao Huang, Weixin Xu, Xiaokun Yuan, Xingcheng Yao, Xingzhe Wu, Xinxing Zu, Xinyu Zhou, Xinyuan Wang, Y. Charles, Yan Zhong, Yang Li, Yangyang Hu, Yanru Chen, Yejie Wang, Yibo Liu, Yibo Miao, Yidao Qin, Yimin Chen, Yiping Bao, Yiqin Wang, Yongsheng Kang, Yuanxin Liu, Yulun Du, Yuxin Wu, Yuzhi Wang, Yuzi Yan, Zaida Zhou, Zhaowei Li, Zhejun Jiang, Zheng Zhang, Zhilin Yang, Zhiqi Huang, Zihao Huang, Zijia Zhao, and Ziwei Chen. Kimi-VL technical report, 2025. URL https://arxiv.org/abs/2504.07491. [26] Qwen Team. Qwen3.5: Towards native multimodal agents, February 2026. URL https://qwen.ai/blog?id= qwen3.5. [27] Qwen Team. Qwen3.6-27b: Flagship-level coding in a 27b dense model, April 2026. URL https://qwen.ai/ blog?id=qwen3.6-27b. [28] Yi Zeng, Hongpeng Lin, Jingwen Zhang, Diyi Yang, Ruoxi Jia, and Weiyan Shi. How johnny can persuade llms to jailbreak them: Rethinking persuasion to challenge ai safety by humanizing llms. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 14322–14350, 2024. [29] Xiaogeng Liu, Peiran Li, G. Edward Suh, Yevgeniy Vorobeychik, Zhuoqing Mao, Somesh Jha, Patrick McDaniel, Huan Sun, Bo Li, and Chaowei Xiao. AutoDAN-turbo: A lifelong agent for strategy self-exploration to jailbreak LLMs. In The Thirteenth International Conference on Learning Representations, 2025. URL https://openreview.net/forum?id=bhK7U37VW8. 13

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

[30] Taylor Shin, Yasaman Razeghi, Robert L Logan IV, Eric Wallace, and Sameer Singh. Autoprompt: Eliciting knowledge from language models with automatically generated prompts. In Proceedings of the 2020 conference on empirical methods in natural language processing (EMNLP), pages 4222–4235, 2020. [31] Jiahao Yu, Xingwei Lin, Zheng Yu, and Xinyu Xing. Gptfuzzer: Red teaming large language models with auto-generated jailbreak prompts. arXiv preprint arXiv:2309.10253, 2023. [32] Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries. In 2025 IEEE Conference on Secure and Trustworthy Machine Learning (SaTML), pages 23–42. IEEE, 2025. [33] Zonghao Ying, Aishan Liu, Siyuan Liang, Lei Huang, Jinyang Guo, Wenbo Zhou, Xianglong Liu, and Dacheng Tao. Safebench: A safety evaluation framework for multimodal large language models. https://safebench-mm. github.io/, 2025. Online resource. [34] Maksym Andriushchenko and Nicolas Flammarion. Does refusal training in llms generalize to the past tense? arXiv preprint arXiv:2407.11969, 2024. [35] Xiangyu Qi, Kaixuan Huang, Ashwinee Panda, Peter Henderson, Mengdi Wang, and Prateek Mittal. Visual adversarial examples jailbreak aligned large language models. In Proceedings of the AAAI conference on artificial intelligence, volume 38, pages 21527–21536, 2024. [36] Zhenxing Niu, Haodong Ren, Xinbo Gao, Gang Hua, and Rong Jin. Jailbreaking attack against multimodal large language model. arXiv preprint arXiv:2402.02309, 2024. [37] Yunqing Zhao, Tianyu Pang, Chao Du, Xiao Yang, Chongxuan Li, Ngai-Man Man Cheung, and Min Lin. On evaluating adversarial robustness of large vision-language models. Advances in Neural Information Processing Systems, 36:54111–54138, 2023. [38] Yucheng Shi, Yahong Han, Yu-an Tan, and Xiaohui Kuang. Decision-based black-box attack against vision transformers via patch-wise adversarial removal. Advances in Neural Information Processing Systems, 35: 12921–12933, 2022. [39] Wieland Brendel, Jonas Rauber, and Matthias Bethge. Decision-based adversarial attacks: Reliable attacks against black-box machine learning models. arXiv preprint arXiv:1712.04248, 2017. [40] Thibault Maho, Teddy Furon, and Erwan Le Merrer. Surfree: a fast surrogate-free black-box attack. In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pages 10430–10439, 2021. [41] Ruofan Wang, Xingjun Ma, Hanxu Zhou, Chuanjun Ji, Guangnan Ye, and Yu-Gang Jiang. White-box multimodal jailbreaks against large vision-language models. In ACM Multimedia 2024, 2024. URL https://openreview. net/forum?id=SMOUQtEaAf. [42] Marcello Galisai, Susanna Cifani, Francesco Giarrusso, Piercosma Bisconti, Matteo Prandi, Federico Pierucci, Federico Sartore, and Daniele Nardi. Adversarial humanities benchmark: Results on stylistic robustness in frontier model safety, 2026. URL https://arxiv.org/abs/2604.18487. [43] Langqi Yang, Tianhang Zheng, Yixuan Chen, Kedong Xiu, Hao Zhou, Wangze Ni, Lei Chen, Zhan Qin, and Kui Ren. Harmmetric eval: Benchmarking metrics and judges for llm harmfulness assessment, 2026. URL https://arxiv.org/abs/2509.24384. [44] sentence-transformers. all-minilm-l6-v2. URL https://huggingface.co/sentence-transformers/ all-MiniLM-L6-v2. Hugging Face model. [45] Joonhyun Jeong, Seyun Bae, Yeonsung Jung, Jaeryong Hwang, and Eunho Yang. Playing the fool: Jailbreaking llms and multimodal llms with out-of-distribution strategy. In Proceedings of the Computer Vision and Pattern Recognition Conference, pages 29937–29946, 2025. [46] Teng Ma, Xiaojun Jia, Ranjie Duan, Xinfeng Li, Yihao Huang, Xiaoshuang Jia, Zhixuan Chu, and Wenqi Ren. Heuristic-induced multimodal risk distribution jailbreak attack for multimodal large language models. In Proceedings of the IEEE/CVF International Conference on Computer Vision, pages 2686–2696, 2025. [47] Bingrui Sima, Linhua Cong, Wenxuan Wang, and Kun He. Viscra: A visual chain reasoning attack for jailbreaking multimodal large language models. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 6142–6155, 2025. [48] Yilian Liu, Xiaojun Jia, Guoshun Nan, Jiuyang Lyu, Zhican Chen, Tao Guan, Shuyuan Luo, Zhongyi Zhai, and Yang Liu. Midas: Multi-image dispersion and semantic reconstruction for jailbreaking mllms. arXiv preprint arXiv:2603.00565, 2026. [49] Zhixue Song, Boyan Han, Yiwei Wang, and Chi Zhang. Hard to read, easy to jailbreak: How visual degradation bypasses mllm safety alignment. arXiv preprint arXiv:2605.07250, 2026. 14

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

[50] Alex Chouldechova, A. Feder Cooper, Solon Barocas, Abhinav Palia, Dan Vann, and Hanna Wallach. Comparison requires valid measurement: Rethinking attack success rate comparisons in ai red teaming. In D. Belgrave, C. Zhang, H. Lin, R. Pascanu, P. Koniusz, M. Ghassemi, and N. Chen, editors, Advances in Neural Information Processing Systems, volume 38. Curran Associates, Inc., 2025. URL https://proceedings.neurips.cc/paper_files/paper/2025/file/ 455d043673bb4b1872ff5e7a24cb3969-Paper-Position_Paper_Track.pdf. [51] Leo Schwinn, Moritz Ladenburger, Tim Beyer, Mehrnaz Mofakhami, Gauthier Gidel, and Stephan Günnemann. A coin flip for safety: Llm judges fail to reliably measure adversarial robustness. arXiv preprint arXiv:2603.06594, 2026. [52] Zheng-Xin Yong, Cristina Menghini, and Stephen H Bach. Low-resource languages jailbreak gpt-4. arXiv preprint arXiv:2310.02446, 2023. [53] Jiayang Song, Yuheng Huang, Zhehua Zhou, and Lei Ma. Multilingual blending: Large language model safety alignment evaluation with language mixture. In Findings of the Association for Computational Linguistics: NAACL 2025, pages 3433–3449, 2025. [54] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C. Lawrence Zitnick. Microsoft coco: Common objects in context. In David Fleet, Tomas Pajdla, Bernt Schiele, and Tinne Tuytelaars, editors, Computer Vision – ECCV 2014, pages 740–755, Cham, 2014. Springer International Publishing. ISBN 978-3-319-10602-1. [55] huihui-ai. Huihui-qwen3.5-9b-abliterated, . URL https://huggingface.co/huihui-ai/Huihui-Qwen3. 5-9B-abliterated. Hugging Face model. [56] huihui-ai. Huihui-gemma-4-31b-it-abliterated, . URL https://huggingface.co/huihui-ai/ Huihui-gemma-4-31B-it-abliterated. Hugging Face model. [57] Tsung-Yi Lin, Michael Maire, Serge Belongie, James Hays, Pietro Perona, Deva Ramanan, Piotr Dollár, and C Lawrence Zitnick. Microsoft coco: Common objects in context. In European conference on computer vision, pages 740–755. Springer, 2014. [58] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. Learning transferable visual models from natural language supervision. In International conference on machine learning, pages 8748–8763. PmLR, 2021.

15

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

A

Categories and subcategories Table 5: Taxonomy of Risk Categories and Subcategories Code

Category

Subcategories (Code – Name)

A

Ethical and Social Risks

B

Privacy and Data Risks

C

Safety and Physical Harm

D

Criminal and Economic Risks

E

Cybersecurity Threats

F

Information and Political Manipulation

G

Content and Cultural Safety

H

Intellectual Property and Ownership

I

Decision and Cognitive Risks

A1 – Bias and Discrimination A2 – Insulting or Harassing Speech A3 – Hate Speech A4 – Discriminatory Behavior B1 – Personal Data Leakage B2 – Biometric or Health Data Exposure B3 – Doxxing or Tracking B4 – Surveillance Misuse B5 – Training Data Leakage C1 – Dangerous Instructions C2 – Self-harm Promotion C3 – Physical Harm C4 – Human Trafficking C5 – Animal Abuse C6 – Drug Guidance C7 – Environmental Damage D1 – Fraud or Scams D2 – Economic Crime D3 – White-Collar Crime D4 – Illegal Activity D5 – Evasion of Law E1 – Malware Generation E2 – Exploit Development E3 – Phishing or Ransomware E4 – Jailbreak or Prompt Injection E5 – Cross-modal Attack E6 – Data Poisoning E7 – Unauthorized Model Access F1 – Misinformation or Disinformation F2 – Deepfakes or Synthetic Media F3 – Political Propaganda F4 – Election Interference F5 – Public Order Disruption F6 – National Security Threat G1 – Pornography or Nudity G2 – Sexual Content G3 – Violence or Gore G4 – Cultural or Religious Offense H1 – Copyright Infringement H2 – Trademark or Brand Misuse H3 – Proprietary Data Leakage H4 – Unauthorized Style or Persona Imitation I1 – Medical Advice I2 – Health Misguidance

16

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Code

J

B

Category

Subcategories (Code – Name) I3 – Legal Advice I4 – Government Decision Support I5 – Financial Advice I6 – Market Manipulation I7 – Psychological Manipulation I8 – Cognitive Bias or Overreliance J1 – CSAM & Sexualization J2 – Grooming or Enticement of Minors J3 – Child Trafficking J4 – Harmful Content Targeting Minors J5 – Age Verification Evasion

Child Safety

Language translation analysis of adversarial attacks

Table 6 evaluates how vulnerable different multimodal large language models (MLLMs) are when safety-critical prompts are translated into various languages or presented in mixed-language settings. Current literature suggests that utilizing low-resource languages should increase model vulnerabilities compared to high-resource ones, as safety alignment data is typically scarce in those languages [52]. To test this hypothesis, we select Farsi and Turkish as target low-resource languages. Furthermore, utilizing a mixture of languages within a translation should obscure prompt intent, heighten deception, and ultimately increase the ASR [53]. For this multi-lingual setting, we choose three high-resource languages (Italian, French, and German) and three low-resource languages (Turkish, Farsi, and Khmer), performing a sentence-by-sentence translation of the adversarial text. As a third approach, we target specific semantic segments by translating only the inherently harmful parts of the prompt into a low-resource language (Partial Turkish) to isolate its effect on model safety. Table 6: Examining model vulnerability against language translation. ASR (%) per model across languages. Model Qwen3.6-27B Qwen3-VL-30B Qwen3.5-27B Ministral-3-14B gemma-4-26B DeepSeek-VL2 GLM-4.6V-Flash Kimi-VL-A3B LLaVA-v1.6-13b

Baseline 43.0 53.5 43.4 80.2 45.7 60.4 84.4 67.7 57.2

Farsi 42.4 58.8 49.7 79.4 53.3 10.3 72.9 29.7 13.9

Italian 46.9 60.4 52.1 79.6 51.5 36.4 72.9 40.6 21.0

German 42.8 62.0 44.4 81.8 52.3 42.0 66.9 46.5 27.5

Chinese 29.5 51.1 19.4 83.6 62.4 40.2 84.4 47.5 31.3

Turkish 45.7 69.1 50.7 82.8 57.6 14.1 62.0 21.6 17.6

Partial Turkish 42.4 55.8 40.6 84.8 53.5 32.1 80.2 55.2 36.4

Mixed Low-Res 23.8 41.8 42.8 57.4 50.1 6.9 36.4 13.3 8.7

Mixed Multiling. 40.6 51.1 49.5 81.2 50.1 42.4 78.0 43.4 20.8

Our experimental evaluation yields several key insights: The Vulnerability Trade-off. The general consensus from our experiments indicates that language translation increases vulnerability only up to the point where it does not compromise the model’s fundamental semantic understanding of the attack. Because adversarial strategies often rely on intricate, multi-layered roleplay scenarios or convoluted logic, translation can introduce excessive linguistic ambiguity. When this ambiguity disrupts comprehension—as heavily observed in the Mixed Low-Res column—the model fails to grasp the underlying prompt intent and generates irrelevant or benign responses. These are classified as non-jailbreaks by the evaluation judge, leading to a sharp decline in ASR for mixed-language settings. Targeted Susceptibility in Specific Model Families. The Qwen, Ministral, and Gemma families exhibit heightened vulnerability when exposed to low-resource languages or hybrid formatting (Partial Turkish). For instance, Gemma-426B shows a noticeable increase in ASR from a baseline of 45.7% to 53.3% in Farsi and 57.6% in Turkish. This confirms that low-resource translations successfully exploit gaps in the cross-lingual safety alignment of these architectures. 17

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Cross-Lingual Robustness and Transfer Variations. Models such as Ministral-3-14B and GLM-4.6V-Flash maintain consistently high vulnerability scores across nearly all language configurations (with Ministral hovering around 80% ASR). This suggests that adversarial prompt structures transfer seamlessly across linguistic boundaries for these models. Conversely, models like DeepSeek-VL2 and LLaVA-v1.6-13b experience drastic drops in vulnerability when prompts are translated (e.g., DeepSeek-VL2 plunging from a 60.4% baseline to just 10.3% in Farsi). This pattern points to either a brittle multilingual comprehension capability or a defensive posture that defaults to safe rejections when faced with distribution shifts in language.

C

Transferability Results

This section presents additional results on the transferability of the attacks in our dataset to a broader set of models, extending those reported in the main corpus. We follow the same evaluation protocol: for each attack and each subcategory, we sample 20 instances. Once this set is fixed, it is evaluated across a range of different models, enabling a direct comparison within each attack strategy. Overall, we evaluate these samples on nine white-box models and six black-box models. Due to the large number of generated outputs, we do not perform manual inspection. Instead, we rely on the state-of-the-art judge Able-24-HarmClassifier [43]. As a consequence, the reported results should be interpreted as relative to this evaluation baseline rather than as ground truth. The full results are reported in table 7. For ease of interpretation, we also provide radar plots offering different insights into the data. First, fig. 6 shows the attack success rates across models and attack strategies. Second, following the analysis in the main paper, fig. 7 presents the attack success rate (ASR) per category, averaged over attack strategies and evaluated across models, highlighting the categories to which models are most vulnerable independently of the chosen attack. Third, fig. 8 shows the ASR averaged over categories, providing an at-a-glance comparison of the most effective attack strategy for each model. Finally, fig. 9 reports the maximum ASR values per model and attack, identifying the weakest category. This allows one to infer, for a given model and attack, the most vulnerable category and the expected performance. An interesting pattern that emerges from fig. 6 and fig. 8 is that MML is the most widely effective attack against almost all models, with the exception of Opus 4.7 and Opus 4.8, which appear to be highly robust to it. However, these two models are particularly vulnerable to the CSDJ attack, which, in turn, is less effective against white-box models. The second most reliable attack across models is FC Attack, which shows good coverage across categories for most models, except for Gemma-4-26B. IDEATOR exhibits the most unpredictable behavior: while it achieves high success rates on some white-box models, such as GLM-4.6V and Mistral-14B, it is generally less reliable, aside from occasional spikes on specific categories. Finally, BAP yields lower but relatively stable performance across models, with success rates ranging from 30% to 50%. Table 7: Table of ASR (%) per model and per category across all attacks, in bold the category with the highest success rate on each model Model

A

DeepSeek-VL2 30.00 GLM-4.6V-Flash 50.00 Gemma-4-26B 23.75 Kimi-VL 40.00 Llava-13b 25.00 Ministral-14B 70.00 Qwen3-VL-30B 46.25 Qwen3.5-27B 50.00 Qwen3.6-27B 42.50 GPT-5.4 13.75 GPT-5.5 47.50 Claude Opus 4.6 6.25 Claude Opus 4.7 51.25 Claude Opus 4.8 47.50 Gemini 3.1 Pro 23.75 DeepSeek-VL2

B 22.00 43.00 16.00 29.00 22.00 56.00 45.00 35.00 40.00 15.00 41.00 11.00 44.00 40.00 39.00

26.25 38.00

C

D

E F BAP 44.29 53.00 48.57 36.67 79.29 72.00 65.71 62.50 10.00 18.00 27.14 20.00 57.14 61.00 52.14 48.33 30.00 39.00 42.14 30.83 85.71 86.00 65.00 60.83 47.14 47.00 38.57 39.17 30.71 50.00 44.29 37.50 30.71 42.00 38.57 35.83 25.00 23.00 11.43 18.33 59.29 68.00 47.86 42.50 0.71 10.00 8.57 5.83 20.71 61.00 40.00 39.17 12.14 65.00 39.29 34.17 35.71 45.00 74.29 28.33 IDEATOR 45.71 58.00 64.29 59.17

18

G

H

21.25 43.75 22.50 27.50 15.00 57.50 36.25 32.50 33.75 10.00 46.25 8.75 35.00 21.25 30.00

21.25 36.25 22.50 32.50 21.25 51.25 37.50 37.50 40.00 6.25 35.00 7.50 56.25 41.25 27.50

I

J

Average

30.62 27.00 50.00 51.00 26.88 34.00 36.88 32.00 23.75 14.00 71.25 62.00 44.38 44.00 45.00 47.00 45.00 51.00 10.62 23.00 48.12 50.00 6.88 17.00 50.00 49.00 43.12 51.00 25.62 44.00

34.82 57.09 22.00 42.91 27.27 67.73 42.73 40.91 39.82 15.91 49.09 7.91 43.64 38.73 38.36

21.25 20.00 33.12 44.00

42.91

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Model A GLM-4.6V-Flash 62.50 Gemma-4-26B 10.00 Kimi-VL 41.25 Llava-13b 20.00 Ministral-14B 51.25 Qwen3-VL-30B 17.50 Qwen3.5-27B 13.75 Qwen3.6-27B 7.50 GPT-5.4 25.00 GPT-5.5 21.25 Claude Opus 4.6 6.25 Claude Opus 4.7 25.00 Claude Opus 4.8 16.25 Gemini 3.1 Pro 28.75

B 56.00 30.00 40.00 39.00 65.00 20.00 17.00 7.00 14.00 28.00 6.00 39.00 24.00 43.00

C D E F G H I J Average 67.14 85.00 82.14 75.83 43.75 42.50 52.50 61.00 64.09 20.71 34.00 33.57 20.00 15.00 21.25 35.62 43.00 27.36 52.14 62.00 66.43 57.50 23.75 20.00 36.25 40.00 45.73 40.00 53.00 60.00 53.33 20.00 18.75 30.62 31.00 38.45 72.86 81.00 91.43 75.83 62.50 52.50 80.00 72.00 72.73 15.71 40.00 35.71 40.83 18.75 16.25 36.25 61.00 31.09 15.00 19.00 12.14 18.33 25.00 18.75 35.62 59.00 23.45 11.43 14.00 15.00 16.67 18.75 7.50 23.12 65.00 18.82 27.14 49.00 44.29 36.67 26.25 12.50 27.50 33.00 30.45 26.43 42.00 48.57 44.17 52.50 31.25 37.50 44.00 37.82 5.71 17.00 13.57 5.83 3.75 5.00 13.12 7.00 8.82 12.14 53.00 30.71 30.00 32.50 45.00 34.38 33.00 32.55 12.14 47.00 32.86 24.17 17.50 33.75 31.25 20.00 26.09 42.14 56.00 71.43 58.33 56.25 42.50 54.37 62.00 52.64 MML DeepSeek-VL2 68.89 65.71 74.32 81.69 78.57 75.00 79.49 65.52 75.56 76.92 76.10 GLM-4.6V-Flash 97.78 100.00 93.48 98.60 97.58 100.00 97.44 96.55 98.11 98.00 97.36 Gemma-4-26B 95.56 100.00 90.34 98.51 99.37 96.88 94.87 100.00 96.55 95.26 96.09 Kimi-VL 86.67 82.86 85.16 90.58 87.04 84.38 92.31 68.97 87.02 87.05 86.66 Llava-13b 68.89 74.29 66.48 81.34 67.92 87.50 79.49 68.97 77.78 77.37 74.55 Ministral-14B 95.56 97.14 99.43 99.25 99.37 96.88 97.44 100.00 99.23 97.89 98.73 Qwen3-VL-30B 82.22 85.71 87.63 92.50 97.78 93.75 87.18 75.86 89.11 79.22 88.35 Qwen3.5-27B 75.56 77.14 51.63 63.50 75.45 81.25 61.54 86.21 85.56 82.42 74.34 Qwen3.6-27B 80.00 85.71 80.97 89.63 94.55 87.50 71.79 86.21 83.14 88.42 86.10 GPT-5.4 71.76 73.45 74.13 67.92 70.00 72.79 78.75 80.46 80.00 95.00 76.03 GPT-5.5 91.25 84.00 90.71 92.00 90.71 89.17 82.50 80.00 81.25 61.00 84.64 Claude Opus 4.6 71.95 61.32 32.14 60.78 63.95 63.78 41.25 48.81 62.86 54.55 56.19 Claude Opus 4.7 1.25 10.00 0.71 6.00 2.86 0.83 0.00 5.00 1.88 1.00 2.82 Claude Opus 4.8 12.50 21.00 3.57 26.00 10.00 11.67 11.25 21.25 13.75 23.00 14.64 Gemini 3.1 Pro 3.53 0.88 4.20 0.94 40.62 81.62 66.25 93.10 93.71 29.00 43.38 FC ATTACK DeepSeek-VL2 88.75 84.00 89.29 94.00 95.71 87.50 78.75 58.75 66.25 75.00 82.18 GLM-4.6V-Flash 91.25 85.00 95.71 93.00 97.14 91.67 78.75 57.50 71.88 82.00 85.18 Gemma-4-26B 37.50 29.00 2.14 7.00 22.14 11.67 33.75 28.75 18.75 14.00 18.91 Kimi-VL 82.50 77.00 82.86 89.00 92.86 85.00 80.00 52.50 66.88 74.00 78.82 Llava-13b 71.25 83.00 87.86 89.00 88.57 84.17 66.25 51.25 58.75 73.00 76.18 Ministral-14B 60.00 65.00 50.00 73.00 83.57 77.50 53.75 41.25 46.25 58.00 61.27 Qwen3-VL-30B 83.75 83.00 59.29 65.00 87.14 81.67 68.75 63.75 69.38 73.00 73.45 Qwen3.5-27B 70.00 75.00 41.43 78.00 90.71 68.33 62.50 75.00 66.25 59.00 68.27 Qwen3.6-27B 86.25 81.00 55.71 85.00 92.86 82.50 72.50 76.25 78.12 64.00 77.27 GPT-5.4 52.50 43.00 60.71 71.00 77.14 74.17 46.25 47.50 59.38 63.00 61.00 GPT-5.5 57.50 57.00 68.57 86.00 89.29 76.67 67.50 63.75 65.00 74.00 71.36 Claude Opus 4.6 77.50 85.00 41.43 68.00 87.86 70.83 63.75 82.50 75.62 60.00 70.82 Claude Opus 4.7 70.00 72.00 38.57 75.00 68.57 65.00 37.50 82.50 68.75 61.00 63.45 Claude Opus 4.8 63.75 83.00 48.57 77.00 79.29 75.00 36.25 78.75 62.50 70.00 67.45 Gemini 3.1 Pro 35.00 46.00 15.00 28.00 67.14 26.67 35.00 55.00 40.62 21.00 37.00 CSDJ DeepSeek-VL2 17.50 36.00 59.29 49.00 57.86 50.83 30.00 27.50 29.38 30.00 38.74 GLM-4.6V-Flash 55.00 61.00 80.00 76.00 90.71 81.67 61.25 50.00 46.88 69.00 67.15 Gemma-4-26B 53.75 75.00 27.86 53.00 85.71 57.50 27.50 41.25 36.88 55.00 51.35 Kimi-VL 53.75 54.00 64.29 62.00 80.00 66.67 53.75 47.50 41.88 60.00 58.38 Llava-13b 3.75 4.00 8.57 7.00 7.86 8.33 11.25 3.75 5.00 3.00 6.25 Ministral-14B 67.50 77.00 70.00 81.00 95.00 85.83 72.50 71.25 66.88 81.00 76.80 Qwen3-VL-30B 73.75 76.00 82.14 84.00 92.86 89.17 62.50 58.75 63.12 70.00 75.23 Qwen3.5-27B 46.25 49.00 35.00 48.00 72.14 56.67 46.25 35.00 39.38 66.00 49.37

19

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Model Qwen3.6-27B GPT-5.4 GPT-5.5 Claude Opus 4.6 Claude Opus 4.7 Claude Opus 4.8 Gemini 3.1 Pro

A 71.25 78.75 82.50 73.75 70.00 80.00 65.00

B 68.00 78.00 80.00 88.00 90.00 89.00 86.00

C D E 64.29 74.00 84.29 91.43 87.00 87.14 88.57 91.00 95.00 67.86 85.00 93.57 57.14 78.00 95.00 71.43 91.00 95.00 39.29 72.00 86.43

20

F 70.00 86.67 85.00 84.17 88.33 87.50 67.50

G 70.00 66.25 81.25 48.75 50.00 50.00 46.25

H 60.00 61.25 70.00 71.25 71.25 71.25 68.75

I J Average 61.88 70.00 69.37 63.75 81.00 78.82 69.38 87.00 83.16 73.75 83.00 77.82 59.38 88.00 74.82 59.38 89.00 78.45 63.12 68.00 66.18

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

A: Ethical & Social B: Privacy & Data C: Safety & Physical D: Criminal & Economic E: Cybersecurity F: Info & Political G: Content & Cultural H: IP & Ownership I: Decision & Cognitive J: Child Safety BAP

IDEATOR

MML

FC ATTACK

CSDJ

DeepSeek-VL2

GLM-4.6V-Flash

Gemma-4-26B

D

D

D

C

E

B

F

E

A 20 40 60 80 100

0

G

B

F

G

I

H

D

C

E

B

F

G

H

B

F

G

F

G

H

C

D

G

D

0

I

Claude Opus 4.8

Gemini 3.1 Pro

D

D

D

E

B

F

0

A 20 40 60 80 100

G

J

H

I

E

B

F

0

A 20 40 60 80 100

G

J

H

I

J

H

C

A 20 40 60 80 100

G

Claude Opus 4.7 C

B

F

J

H

C

E

A 20 40 60 80 100

G

I

I

Claude Opus 4.6 C

0

J

H

B

F

A 20 40 60 80 100

G

I

E

J

H

0

GPT-5.5

A 20 40 60 80 100

0

C

B

F

J

I

B

F

I

I

E

A 20 40 60 80 100

0

J

E

J

H

B

GPT-5.4

A 20 40 60 80 100

0

C

E

A 20 40 60 80 100

D

B

H

D

H

E

I

J

D

C

G

C

G

D

Qwen3.6-27B D

A 20 40 60 80 100

Qwen3.5-27B

0

I

0

Qwen3-VL-30B

F

J

F

J

H

B

Ministral-14B

E

A 20 40 60 80 100

0

G

C

E

A 20 40 60 80 100

0

I

D

B

F

J

Llava-13b

Kimi-VL

C

E

A 20 40 60 80 100

0

J

H

C

I

C

E

B

F

0

A 20 40 60 80 100

G

J

H

I

Figure 6: The plot graphically presents the vulnerability of each tested model to harmful categories, enabling a comparison across attacks.

21

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

A: Ethical & Social B: Privacy & Data C: Safety & Physical D: Criminal & Economic E: Cybersecurity F: Info & Political G: Content & Cultural H: IP & Ownership I: Decision & Cognitive J: Child Safety DeepSeek-VL2 D

GLM-4.6V-Flash C

67.1

D 84.9

62.6

E

B

Gemma-4-26B

C 83.1

D

B

E

61.8

46.3 0

20

40

46.1

A F 80 100

60

82.3

J

53.9

I

H

C

D 84.0 B

46.6

40

37.8 0 38.4 32.8

20 40 39.7

C

D

B

79.4

68.9 20

40

G

J

H

G

B

I

H

C

D

Qwen3.6-27B D

J

60.9

B

48.6

59.6

E

58.5

57.5 0

20

40

C

D

B

48.4 0

20

67.7 54.0

41.6

58.2

G

J

H

60

B

44.7

52.0

0

20

40

60

46.4 49.2 J

I

60

A 80 100

46.5 J

H

I

40

B 40.4

68.0

60

C

E 27.3

43.0

52.5

A F 80 100

31.2 0

20

40

60

A 80 100

44.8

46.8

50.6

42.9

G

H

20

40 44.3

43.0

D

44.0 0

27.2

20

G

I

51.4

46.5

A F 80 100

50.3

47.1 0

33.2

63.2

B 29.6

29.6

46.1

A F 80 100

60

C

51.3

43.5 31.0

40

B

48.2

Gemini 3.1 Pro

61.2

E

51.0

E

J

D

C

60.2

H

C

25.9

20

G

D

I

53.5

60.0

56.0

Claude Opus 4.8

54.6

B

I

54.4

Claude Opus 4.6

66.0 J

H

47.4

D

58.0

48.2

A 80 100

66.7

0

Claude Opus 4.7

E

C

67.5

59.0

60

J

H

74.3

A F 80 100

G

I

F

40

40

G

I

44.7

45.5

53.4

E

20

62.7 50.5

GPT-5.5

57.7

A F 80 100

60

51.1 0

J

H

55.7

50.6

52.4

A F 80 100

60

G

I

58.0

56.3

40

34.8

65.4

GPT-5.4

65.0

20

B

51.7 59.0

60.5

75.8 E

E

C

45.6 50.4

72.7

Qwen3.5-27B

60.7 0

I

D

68.9

A F 80 100

74.2 63.2

J

61.9

54.7

68.7

39.2

60

A 80 100

53.8

H

58.4

70.4

60

G

C

E

72.0

40 58.6

44.3

I

65.7

20

55.5

Qwen3-VL-30B

75.6

0

60.8 0

J

H

E 86.9

A F 80 100

60

68.4

A F 80 100

60

42.9

44.5

52.8

F

20

G

I

B

75.7

48.2 42.8

J

E

56.6

44.1 0

Ministral-14B

53.3

F

41.2

A F 80 100

63.9

G

Llava-13b D

E

60 72.2

56.6

G

H

40

65.0

47.0

C 68.3

50.0

38.7

50.6 38.6

20

30.2

53.6

71.3 0

B

42.1

69.0

49.1

F

D 72.9

E 86.7

69.0

Kimi-VL C

42.0 57.4

G

J

H

I

55.5

G

J

H

I

Figure 7: The plot presents the vulnerability of models to harmful categories. It is obtained by averaging results across different attack strategies to mitigate the influence of the specific attack used.

22

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

DeepSeek-VL2

GLM-4.6V-Flash

IDEATOR

MML

76.1

Gemma-4-26B

IDEATOR 64.1

97.4 MML

42.9

Kimi-VL

IDEATOR

IDEATOR

MML96.1

MML 86.7

45.7

27.4 57.1 BAP 0 20 40 60 80 100

34.8 BAP 0 20 40 60 80 100 82.2

38.7

18.9

85.2

FC ATTACK

22.0 BAP 0 20 40 60 80 100 78.8

51.4 67.2

FC ATTACK CSDJ

FC ATTACK CSDJ

Ministral-14B IDEATOR

58.4

FC ATTACK

CSDJ

Llava-13b

42.9 BAP 0 20 40 60 80 100

CSDJ

Qwen3-VL-30B

IDEATOR

Qwen3.5-27B

IDEATOR

IDEATOR

72.7 MML

74.5

98.7 MML

38.5

MML 88.3

27.3

67.7

BAP 6.20 20 40 60 80 100

CSDJ

68.3 75.2

FC ATTACK

CSDJ

IDEATOR

MML

76.0

CSDJ

GPT-5.5 IDEATOR

18.8

49.4

FC ATTACK

CSDJ

GPT-5.4

MML 86.1

40.9 BAP 0 20 40 60 80 100

73.5 76.8

FC ATTACK

Qwen3.6-27B

23.4

BAP 0 20 40 60 80 100

61.3

FC ATTACK

74.3

42.7

BAP 0 20 40 60 80 100

76.2

Claude Opus 4.6 IDEATOR

MML 84.6 30.4

IDEATOR

MML

37.8

56.2 8.8

39.8

49.1

15.9

BAP 0 20 40 60 80 100

BAP 0 20 40 60 80 100

69.4

FC ATTACK

CSDJ

78.8

83.2

Claude Opus 4.8

IDEATOR

MML

77.8 CSDJ

Gemini 3.1 Pro

IDEATOR

IDEATOR 52.6

MML

32.5

26.1

43.4

14.6 43.6 BAP 0 20 40 60 80 100

38.7 BAP 0 20 40 60 80 100

63.5 FC ATTACK

FC ATTACK

CSDJ

Claude Opus 4.7

2.8

70.8

71.4 FC ATTACK

CSDJ

MML

7.9 BAP 0 20 40 60 80 100

BAP 0 20 40 60 80 100

61.0

77.3 FC ATTACK

MML 31.1

37.0

38.4 BAP 0 20 40 60 80 100

67.5 74.8

FC ATTACK

78.5

CSDJ

CSDJ

FC ATTACK

66.2

CSDJ

Figure 8: The plot shows the vulnerabilities of the models with respect to the attack strategies, averaged over the categories.

23

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Attack Strategy

BAP

IDEATOR

MML

FC ATTACK

CSDJ

A. Ethical & Social Risks B. Privacy & Data Risks C. Safety & Physical Harm D. Criminal & Economic Risks E. Cybersecurity Threats F. Info & Political Manipulation G. Content & Cultural Safety H. IP & Ownership I. Decision & Cognitive Risks J. Child Safety 120%

100%

E Maximum ASR (%)

100.0% 97.1%

95.7%

81.7%

80%

B

85.0% 79.3% D

D

100.0%

B

E90.7% E

92.9% 92.3%

E

FD

80.0%

E

59.3%

66.4% 61.0%E

60.0%

C

D

E

53.0%

100.0% H 95.0% 91.4% E 86.0%E 83.6%

D

D

34.0%J

J

90.7% 86.2%

E

H

J

EE

84.3%

E E

61.0%

C

E

E

51.0%

A

J

D

49.0%

95.0%

93.6% 87.9%E

E

E

68.0%

A

61.0%

G

sh

M

GL

B

-26

Ge

a-4 mm

i Kim

-VL

v Lla

3b a-1

I 86.4% E 74.3% 71.4% 67.1% E

D

E

47.0%

D 26.0%

D

17.0% 17.0%

J D

10.0%

G

la V-F -4.6

93.7%

E

65.0%

D

11.2%

L2

B

D53.0%

C

V ekpSe Dee

H

25.0%

20%

0%

E 83.0%

52.5%

D

95.0%

E 82.5%

72.0%

D

J

50.0%J

95.0% 92.0% 89.3%E

C

77.1%

65.0%

59.0%

J

42.1%

A

95.0% 91.4%

94.5% 92.9%

E

72.1%

47.1%

43.0% 37.5%

40%

97.8% E 92.9% 87.1% E

E

E

C

64.3%

60%

89.0% 87.5%

GE

85.7%

B

-14

ral nist

Mi

B -30

B Q

L 3-V wen

B -27

Q

3.5 wen

B -27

.6 en3

Qw

.4 T-5

GP

.5 T-5

.6

.7 us 4

us 4

GP

Cl

p eO aud

Cl

p eO aud

ude

Cla

.8 us 4

Op

Gem

ini

3.1

Pro

Model

Figure 9: This plot identifies, for each attack strategy and model, the most vulnerable category by selecting the category with the highest ASR score.

D

Benchmark overview

D.1

Evolution of Textual and Early Multimodal Benchmarks

The field was pioneered by AdvBench [17], a text-only benchmark containing 520 behaviors. Despite its reliance on primitive string-matching judges and the simple GCG attack, it established the foundational framework for subsequent research. Building on this, VAJM [16] introduced the image modality and implemented weak categorization based on race and gender. It advanced evaluation methods by utilizing the Detoxify classifier as a judge and introducing prompt-tuning optimization attacks. HarmBench [1] contributed the first rigorous categorization of behaviors into four functional groups. It further matured the evaluation process by introducing a fine-tuned Llama 2 model as a judge and proposing the R2D2 defense mechanism. Expanding the scale of multimodal research, JailBreakV-28K [19] incorporated 16 categories and 2000 behaviors (sourced from RedTeam-2K), utilizing 28k attack pairs generated via advanced typographic Stable Diffusion attacks.

D.2

Diversifying Metrics and Modalities

Subsequent works focused on refining metrics and interaction types. MM-SafetyBench [2] introduced the “Refusal Rate” as a key metric, emphasizing a comparison between Attack Success Rates (ASR) when models are given text-only queries versus queryrelevant image-text pairs. In the purely textual domain, Strong Reject [15] moved away from binary evaluation by implementing a graded scoring system (0, 0.33, 0.66, 1.0) to distinguish between full refusal, partial refusal, partial fulfillment, and full fulfillment across 37 different attacks. The importance of conversational context was highlighted by research into Multiturn human jailbreaks [14], which demonstrated that models are significantly more vulnerable through iterative, back-and-forth prompting: a feature we leverage through the attack strategy chosen for our dataset. Further expanding the scope of modalities, SafeBench [33] integrated audio alongside text and images, while introducing a “Safety Index Risk” evaluated by a consensus-based roundtable of judges rather than a single entity.

D.3

Balancing Robustness with Model Utility

A critical shift in the literature involves the trade-off between safety and helpfulness. MMJ-bench [3] categorized attack strategies into optimization- and generation-based methods, arguing that a perfect defense is counterproductive if it causes the model to refuse every prompt. Similarly, JailbreakBench [10] introduced 100 benign prompts designed to appear harmful but which are actually safe, allowing researchers to measure if a model is overly defensive. To quantify this performance impact, B-AVIBench [12] introduced the Average Score Drop Rate (ASDR), measuring the percentage decrease in performance scores following an attack across various image, text, and content bias types. To ensure automated evaluations remain grounded, Sorry-Bench [11] provided a human validation dataset for judges, utilizing Cohen’s Kappa (κ) to measure the correlation between AI judges and human evaluation, alongside fulfillment rate as an additional metric.

24

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

D.4

High-Granularity Categorization

Recent benchmarks have achieved unprecedented depth in their taxonomies. VLJailbreakBench [10] implemented a robust categorization featuring 12 safety topics and 46 subcategories. Finally, OmniSafeBench [4]—which serves as the primary reference for this work, introduced 9 major risk domains and 50 fine-grained categories. Beyond evaluating 15 different defense strategies, it established a multifaceted judgment criteria incorporating Harmfulness (H), Intent Alignment (A), and Level of Detail (D) to provide a holistic view of model safety. Table 8: Comparative Analysis of Safety and Jailbreak Benchmarks Benchmark

Release

Mod. Behaviours Samples Model Att. Def. Judges

OmniSafeBench VLJailbreak

6 Dec 2025 T/I 25 Sep 2025 T/I

50 cat. 46 cat.

1 200 3 654

18 11

13 5

Sorry-Bench AgentHarm B-AVIBench JailbreakBench MMJ-Bench

Mar 2025 T 18 Apr 2025 T/I 28 Dec 2024 T/I 31 Oct 2024 T 22 Oct 2024 T/I

44 Topics 11 cat. 23 Types 20 cat. 200 Behav.

8 800 440 316k 100 1 000

50 15 14 4 6

0 1 10 4 9

SafeBench

4 Oct 2024 T/I/A 23 cat.

9 200

21

3

MHJ

4 Sep 2024

1 000 Req.

2 912

1

14

6 cat. 13 cat.

346 5 040

3 12

37 0

16 cat. 4 cat.

28k 510

10 33

10 22

T

StrongREJECT 27 Aug 2024 T MM-SafetyBench 19 Jun 24 T/I JailBreakV-28K HarmBench

3 Apr 2024 Feb 2024

T/I T/I

VAJM

16 Aug 2023 T/I

40 Behav.

32 226

3

1

AdvBench

July 2023

520 Behav.

520

9

8

T

15 GPT-4o 0 GPT-4o, GPT-4 0 Mistral-7B 0 GPT-4o 0 GPT-4 0 6 Classifiers 5 GPT-4, HarmBench 0 Ensemble (2) 0 GPT-4o, HarmBench 0 Gemma 2B 0 GPT-4, Llama-2 0 4 Classifiers 1 Llama-2 (FT) 0 Perspective API 0 GPT-4, String

Metrics ASR, SRI ASR FR, RR, κ SR, RR ASDR, AED ASR, FPR ASR ASR, SRI ASR Full Refusal ASR, RR ASR ASR Toxicity ASR

Legend: Mod. = Modality (T: Text, I: Image, A: Audio); Model = Number of models evaluated; Att. = Number of attack strategies; Def. = Number of defense strategies.

E

PHANTOM Similarity Checks

To assess the diversity of the generated adversarial prompts, we performed a cosine-similarity analysis over the textual component of the attacks. This analysis quantifies prompt-level redundancy, including cases where the same underlying intent may lead to multiple generated attacks. Such repetitions are expected, since PHANTOM contains 7 826 unique intents but nearly 30k generated adversarial samples. We exclude the MML strategy from this analysis because, for this attack, the adversarial content is primarily encoded in the image rather than in the textual prompt. MML and FC ATTACK prompts rely on a shared instruction template, while the harmful intent is embedded through visual transformations such as encoding, mirroring, rotation, or word substitution. Therefore, measuring redundancy using only the textual prompt would produce similarity scores ∼ 100%, without providing a meaningful estimate of sample diversity. Table 9 reports the redundancy rates obtained for BAP and IDEATOR across different cosine-similarity thresholds, both globally and broken down by attack strategy and target model. For each threshold τ , we construct clusters of prompts whose pairwise cosine similarity is greater than or equal to τ . Within each cluster, one prompt is treated as the representative, while the remaining prompts are counted as redundant. Formally, if K denotes the set of clusters and |Ck | the size of cluster Ck , the redundancy rate is computed as: P Ck ∈K max(|Ck | − 1, 0) Redundancy = × 100, N where N is the total number of prompts in the analyzed group.

25

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

Table 9: Redundancy Rate (%) Across Thresholds Group T=80% T=85% T=90% global (BAP + IDEATOR) 11.58% 9.59% 8.91% strategy:BAP 12.10% 10.27% 9.72% strategy:IDEATOR 1.76% 0.81% 0.22% model:DeepSeek-VL22 6.45% 4.46% 3.18% model:GLM-4.6V-Flash 9.95% 9.67% 9.61% model:Kimi-VL-A3B-Instruct 7.20% 4.34% 3.30% model:Qwen3-VL-30B-A3B 1.49% 0.52% 0.38% model:Qwen3.5-27B 12.86% 12.65% 12.55% model:Qwen3.6-27B 10.28% 8.10% 7.70%

F

Review of the attack strategies

In this section we will give an overview of the attack strategies that we used in the generation.

F.1

BAP attack

The core idea behind the BAP attack is to jointly optimize visual and textual components. First, an adversarial perturbation is applied to the input image through projected gradient descent (PGD), using a corpus of affirmative model responses as optimization targets. Subsequently, a prompt engineering step is performed to obfuscate the harmful intent within a seemingly benign textual prompt. The outcome of this process is an adversarial image that biases the model toward affirmative responses, coupled with a carefully engineered prompt that facilitates the bypass of safety mechanisms. In our pipeline, we fixed the number of PGD optimization steps to approximately 200 and optimized the adversarial image against a batch of 8 affirmative target responses, starting from clean images from the COCO train dataset [54]. For the prompt engineering phase, we followed the standard iterative interaction flow. In addition to the target model under attack, we employed an abliterated version of Qwen3.5, namely Huihui-Qwen3.5-9Babliterated [55] as the attacker model. Its reasoning capabilities were leveraged via a crafted system prompt that explicitly encoded previous failed attempts. As anticipated before, we used Abel-24-HarmClassifier proposed in [43] as judge model. The prompt optimization loop was iterated for up to 5 attempts for each attack instance. At first glance, fig. 3 may suggest that this attack is less practical, given that it is substantially slower than the alternatives. However, our decision to include it was motivated by an additional advantage: the adversarial images produced by this pipeline are universal. As a result, one can recombine intents and images to obtain additional valid attacks, although some filtering and discarding may still be required.

F.2

IDEATOR attack

Also in the case of this attack strategy, proposed in [6], we largely followed the original pipeline. Our modifications mainly consist of introducing different models for prompt and image generation. The attack pipeline relies on an attacker that produces two distinct prompts: one used to generate an image related to the harmful intent, and another aimed at engineering the harmful textual prompt itself. Since the process is implemented as a multi-turn conversation in which the prompt is progressively refined, we limited each conversation to a maximum of three image–prompt pairs. In addition, for each target goal, the attack was retried at most three times; these retries serve primarily as a fallback mechanism rather than a core component of the method. In our implementation, we employed the same ablated Qwen3.5 model above as the attacker to generate both the textual prompts, while image generation was performed using Stable Diffusion 3.5 Medium. We selected this strategy not only because of its efficiency, but also because its structure naturally supports both multi-turn and single-turn settings: the full conversation can be used as input, or alternatively only the final image–prompt pair can be retained.

F.3

MML attack

As with the other methods, we remained faithful to the original structure of the Multi-Modal Linkage attack proposed in [9]. In our implementation we start from a harmful or restricted text prompt and apply obfuscation: it replaces key words with benign ones through NLTK package, optionally encodes the text (e.g., Base64), and then renders the transformed text into an image. Additional visual distortions, such as mirroring, rotation, or both, are mainly applied through Pillow library to make the content harder to directly interpret. Alongside this image, the system constructs a carefully designed “game-like” prompt that instructs the model to recover the original text by reversing these transformations (e.g., decoding, un-mirroring, or using a provided word-mapping

26

PHANTOM: A Large-Scale Dataset of Multimodal Adversarial Attacks for Vision-Language Models

dictionary) and validating it against a scrambled word list. The resulting image–text pair is fed into target vision-language model, which is guided step-by-step to reconstruct the original prompt and then generate detailed content based on it. Because the harmful intent is never explicitly presented in raw form but instead reconstructed by the model itself, safety mechanisms can be bypassed. Adversarial Image rm

sfo

Mirror/Rotate/Word Replacement

an Tr

Vi su al

Target VLM

Harmful Intent Ro le

pla y

Adversarial Prompt

xt Te

Persona (Game Dev) Word Scramble

Figure 10: Workflow of MML combining both image manipulation and role-playing through text. F.4

FC ATTACK

Once again, our methodology closely follows the original approach proposed in [20]. The core idea is to start from a harmful intent and generate a sequence of logical steps to address it using an auxiliary abliterated model. In our case, similarly to BAP [8], we employ the Huihui-Qwen3.5-9B-abliterated model [55]. These steps are then represented as a flowchart using standard Python libraries such as Graphviz. Finally, the model is prompted with both the flowchart and a standard instruction that encourages it to reason by following the outlined steps.

F.5

CSDJ attack

We generated attacks using the CS-DJ attack strategy proposed in [21]. Following the original pipeline, we crafted each attack as follows. We first selected an intent from our dataset and then followed two parallel paths. First, using an abliterated model, namely Huihuigemma-4-31B-it-abliterated [56], we decomposed the harmful request into three less harmful sub-requests and embedded each of them into separate images. Second, we selected nine additional images from a pool of 10 000 images taken from COCO training set[57], these images are chosen such that their CLIP[58] embeddings are maximally distant from the embedding of the original intent. We then combined these components: the nine images were arranged in a 3×3 grid, followed by the three images containing the generated sub-requests. The images were numbered from 1 (top-left) to 12 (bottom-right).

27

Record · ID 303213 · SHA-256 4f39bf0ce4653e22
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.