DARWIN: E VOLVING JAILBREAK A DVERSARY AND G UARDRAIL FOR LLM S AFETY E VALUATION AND P ROTECTION Weiwei Qi1∗, Zefeng Wu1∗, Zhilin Guo1∗, Tianhang Zheng1,2† Chaochao Lu3 , Liang He4 , Zhan Qin1,2 , Kui Ren1,2 1
The State Key Laboratory of Blockchain and Data Security, Zhejiang University Hangzhou High-Tech Zone (Binjiang) Institute of Blockchain and Data Security 3 Shanghai AI Laboratory 4 East China Normal University 2
{weiweiqi,zefengwu,zhilinguo,zthzheng,qinzhan,kuiren}@zju.edu.cn [email protected] [email protected]
arXiv:2607.19829v1 [cs.CR] 22 Jul 2026
A BSTRACT Most existing LLM safety evaluation and defense methods follow a static formulation: jailbreak vulnerabilities are evaluated with fixed attack methods, and guardrails are trained on fixed malicious prompt datasets. However, real-world adversaries continuously evolve their attack capabilities and expand the attack space. To address this challenge, we propose DARWIN, an evolutionary attackdefense framework that formulates jailbreaking as an open-ended evolution process and continuously updates guardrails through an evolving attack-defense loop. We propose DARWIN-Attack as an evolutionary adversary that expands its attack capabilities through strategy discovery, mutation, and selection. DARWIN-Attack discovers new attack strategies from broad external sources, generates new variants through self-reflection and genetic evolution, and filters effective strategies according to their performance against aligned LLMs. During the attack execution phase, DARWIN-Attack adaptively selects and composes evolved strategies according to feedback from target LLMs and guardrails. Through continuous evolution, DARWIN-Attack achieves state-of-the-art attack success rates against frontier LLMs and guardrails, e.g., nearly 100% on DeepSeek-V4-Pro, over 90% on GPT-5.5, and nearly 100% on YuFeng-XGuard. The continued evolution of DARWIN-Attack exposes new safety vulnerabilities and requires timely corresponding updates to safety defenses. Therefore, on the defense side, we introduce DARWIN-Guard, an online adversarial guardrail training paradigm, which iteratively trains the guardrail on the emerging adversarial samples generated by DARWIN-Attack. To improve robustness without sacrificing utility, DARWINGuard jointly learns from malicious and benign disguised queries, encouraging the guardrail to recognize underlying intent rather than superficial attack patterns. Through continuous evolution, DARWIN-Guard achieves an average unsafe recall of 91.6% across 12 safety evaluation benchmarks, outperforming recent advanced guardrails such as YuFeng and Nemotron. Meanwhile, DARWIN-Guard maintains an average pass rate of nearly 100% on standard benign datasets.
1
I NTRODUCTION
Large language models (LLMs) are widely deployed in various application domains, but they face a significant safety threat, i.e., jailbreak attacks Zou et al. (2023); Huang et al. (2025); Xiu et al. (2025), which can induce LLMs to output harmful instructions or operations. Current LLM safety evaluation and protection mechanisms Qi et al. (2026b) are typically trapped in a static paradigm: ∗ †
Equal contribution. Corresponding author.
1
jailbreak vulnerability assessments rely on fixed attack methods Huang et al. (2026); Ding et al. (2024), and safety guardrails are trained on fixed harmful prompt datasets Lin et al. (2026); Zhao et al. (2025); Wu et al. (2026). This static setup inevitably falls behind the realistic adversarial environment, where potential adversaries continually evolve their attack capabilities to uncover new vulnerabilities Liu et al. (2024); Zhang et al. (2026). Confronting the evolving adversarial environment, we propose DARWIN, an evolutionary attackdefense framework for the open-ended evolution of jailbreak adversaries and the continued improvement of guardrails, with two coupled modules named DARWIN-Attack and DARWIN-Guard. The key idea of DARWIN-Attack is to evolve the adversary itself rather than search for jailbreak prompts within a fixed attack space. DARWIN-Attack maintains an evolving jailbreak strategy pool by continuously ingesting new strategies from external sources and generating new strategies through genetic evolution. These newly introduced strategies are verified against aligned LLMs in a sandbox environment, and only strategies whose attack success rates exceed a predefined threshold are admitted into the pool, making DARWIN-Attack increasingly effective over evolution rounds. During attack execution, DARWIN-Attack first selects an initial strategy from the evolving pool based on historical attack records and applies it to disguise the harmful query. If the initial attempt fails to attack the target LLM or guardrail, DARWIN-Attack leverages attack feedback to adaptively select and compose subsequent strategies with the current jailbreak prompt. Meanwhile, it analyzes the rejection reasons of failed attempts through reflection and refines the corresponding strategies for subsequent attacks. Through this iterative process of strategy exploration, composition, and refinement, DARWIN-Attack continuously discovers stronger jailbreak capabilities and achieves state-of-the-art jailbreak attack success rates across frontier LLMs and strong guardrails, e.g., nearly 100% on DeepSeek-V4-Pro, over 90% on GPT-5.5, and 99% on YuFeng-XGuard. On the defense side, DARWIN-Guard establishes an online adversarial training pipeline Zheng et al. (2019); Ren et al. (2020) to train a guardrail, which continuously leverages DARWIN-Attack to generate new attack samples against the current guardrail. To mitigate over-refusal on benign inputs Röttger et al. (2024), we also disguise safe prompts using DARWIN-Attack and incorporate them into the training data. Jointly training on malicious and benign disguised queries forces the guardrail to recognize true underlying intents rather than merely memorize superficial disguising formats. As the trained guardrail becomes increasingly robust, DARWIN-Guard naturally pushes DARWIN-Attack to uncover more sophisticated jailbreak strategies, thereby forming an evolving attack–defense loop. Through extensive iterative evolution, DARWIN-Guard achieves an 91.6% unsafe block rate on average across 12 safety and jailbreak evaluation benchmarks, outperforming recent advanced guardrails such as YuFeng XGuard Lin et al. (2026) and Nemotron Guard Joshi et al. (2025). Meanwhile, DARWIN-Guard maintains an average pass rate of nearly 100% on 11 standard benign benchmarks, successfully avoiding the over-refusal issue of some strongly aligned LLMs and guardrails. All in all, DARWIN introduces an evolving paradigm for evaluating and enhancing LLM safety. By replacing the static formulation of existing works with an evolving attack–defense loop, DARWIN enables that defensive capabilities improve concurrently with emerging safety threats. Extensive experiments validate the effectiveness of DARWIN in generating challenging attack samples and training robust guardrails. Our main contributions are summarized below. • We propose DARWIN as an evolving attack–defense framework for LLM safety evaluation and defense. DARWIN overcomes the limitations of static paradigms by establishing a continuous loop of attack prompt generation and guardrail training. The loop allows DARWIN-Attack to discover novel and effective strategies over time, while DARWIN-Guard improves by learning from the attacks generated in each round. • We design DARWIN-Attack as an advanced evolving adversary that continuously uncovers vulnerabilities in LLMs and guardrails. By integrating an evolving strategy pool with a feedbackdriven attack mechanism, DARWIN-Attack achieves superior attack success rates against frontier LLMs and guardrails such as GPT-5.5 and YuFeng-XGuard. • We design DARWIN-Guard as an online adversarial training pipeline driven by DARWINAttack. To capture underlying intents and mitigate over-refusal, DARWIN-Guard jointly trains its guardrail on harmful and benign queries disguised by DARWIN-Attack. DARWIN-Guard
2
also achieves higher unsafe block rates than YuFeng-XGuard and maintains a nearly 100% pass rate on benign data.
2
R ELATED W ORK
2.1
JAILBREAK ATTACKS
Jailbreak attacks aim to elicit policy-violating outputs from safety-aligned LLMs through adversarially crafted prompts. Wei et al. (2023) show that manually crafted prompts can jailbreak aligned LLMs. Shen et al. (2024) analyze in-the-wild jailbreak prompts and evaluate their attack effectiveness against aligned LLMs. GCG automates jailbreak prompt generation by optimizing transferable adversarial suffixes with white-box gradient access (Zou et al., 2023). PAIR shifts to a black-box feedback-driven setting, where an attacker LLM iteratively refines candidate prompts using targetmodel responses and evaluator feedback (Chao et al., 2024b). TAP extends this feedback-driven formulation with tree search and pruning to filter out unpromising branches before querying the target model (Mehrotra et al., 2024). AutoDAN-Turbo further introduces a lifelong jailbreak agent that autonomously discovers reusable jailbreak strategies (Liu et al., 2025a). MAJIC models the selection and composition of disguise strategies as a Markov chain over a fixed strategy pool, but does not update the pool as new attacks emerge(Qi et al., 2026a). MemoAttack organizes accumulated attack experience into skill-structured attack memory to support reuse across attack instances (Zhang et al., 2026). 2.2
G UARDRAILS
Guardrails serve as external safeguards for LLMs by recognizing potentially harmful user inputs and model outputs. LlamaGuard formulates this setting as taxonomy-based prompt and response classification, establishing a representative input-output safeguard paradigm (Inan et al., 2023). ShieldGemma introduces Gemma 2-based content moderation models for detecting safety risks in both user inputs and model outputs (Zeng et al., 2024). WildGuard jointly models prompt harmfulness, response harmfulness, and refusal detection, thereby covering both safety risks and refusal behaviors in user-model interactions (Han et al., 2024). Recent work has focused on reasoning, fine-grained risk assessment, and policy flexibility. GuardReasoner introduces explicit reasoning supervision through reasoning SFT and hard-sample DPO to improve safety judgments on challenging examples (Liu et al., 2025b). Qwen3Guard supports multilingual and streaming moderation with generative and streaming variants, enabling triclass safety judgments over safe, controversial, and unsafe content (Zhao et al., 2025). YuFengXGuard emphasizes interpretable and flexible risk assessment through structured risk prediction, confidence estimation, and natural-language explanations (Lin et al., 2026). In a policy-adaptive setting, DynaGuard replaces static safety categories with user-defined policies for application-specific guardrail decisions (Hoover et al., 2025). However, existing guardrails are typically trained on fixed datasets and are not updated as new jailbreak strategies emerge.
3
THE DARWIN FRAMEWORK
3.1
OVERVIEW OF THE E VOLVING ATTACK –D EFENSE L OOP
As illustrated in Figure 1, the DARWIN framework establishes an evolving attack-defense loop that consists of two coupled components, DARWIN-Attack and DARWIN-Guard. At evolution round t, the attacker is represented by At = (St , Tt ), where St denotes the current jailbreak strategy pool, and Tt is the Markov transition matrix used to select the next strategy after a failed attempt. The current safety guardrail is denoted by Gθt with parameters θt . In each round, DARWIN-Attack first collects or generates new effective jailbreak strategies to form a new set ∆St . By incorporating ∆St and the attack feedback Ft from previous attempts, the attacker upgrades its state to At+1 . Utilizing the expanded strategy pool and the refined transition matrix, DARWIN-Attack generates a batch of adversarial training examples against the current guardrail Gθt . 3
DARWIN-Attack
Evolving Jailbreak Attack against Target LLMs and Guardrails
1 Strategy Pool Expansion External attack knowledge Automatic mining and crawling Genetic evolution Success history memory
Adaptive BlackBlack Box Jailbreak Box Jailbreak 22 Adaptive
Strategy Pool
Harmful Query
Role Play
Attacker LLM
Strategy Retrieval
Prompt Injection Task Wrapping
3 3 Feedback-Driven Evolution
Transition Matrix
Strategy Composition
Reflection & Scoring
Format Obfuscation
… more strategies
Mutation & Refinement
Disguised Prompt
Target LLM
Judge Evaluator
Sandbox Evaluator
Success: Harmful Response or Guard Bypass
Iterate Until Success
Indirect Request
Goal:Evolve strategies to jailbreak target models.
Attack Memory Q-learning Inspired Update
Update Strategy Pool Update Success Memory Update Decision Matrix
New Strategy Mutation Retry&Explore
DARWIN-Guard
Harmful Prompts (real-world tasks)
Benign Prompts (helpful tasks)
Raw Anchor Prompts (coverage pairs)
3 Sample Selection
2 DARWIN-Driven Data Generation DARWIN-Attack (Adversarial Rewrite Engine)
4
Goal:Continuously improve the guardrail
Guard Update
Mine Harmful False Negatives
Harmful Rewrites
Mixed Training Reservoir(All Categories)
Bypass (Unsafe) Keep Useful Benign Rewrites
Benign Rewrites Evaluated by Current Guardrail Pass(Safe)
Over-refusal (False Positive)
Block(Unsafe)
Failure: Refused or Blocked Analyze & Reflect
Evolving Guardrail Defense via Online Adversarial Training
1 Data Sources
Policy Evaluator
Multi-Round Fine-Tuning /Continual Updates Stronger Guardrail
Better Benign Utility
Figure 1: Overview of the DARWIN framework. DARWIN-Attack evolves jailbreak adversaries through strategy pool evolution, adaptive strategy selection, and feedback-driven refinement. DARWIN-Guard continuously improves through online adversarial training with attack samples generated by DARWIN-Attack, forming an evolving attack-defense loop. DARWIN-Guard then trains on these adversarial examples to update its parameters to θt+1 . By learning from these examples, the guardrail becomes more robust and better at recognizing underlying malicious intents. Confronting this increasingly capable defense, DARWIN-Attack must maintain continuous evolution and explore more effective jailbreak strategies in the next round. Formally, the evolution loop is abstracted as the following coupled updates: ( At+1 = St ∪ ∆St , Update(Tt , Ft ) , (1) θt+1 = arg min LGuard θ; Attack(At+1 , Gθt ) , θ
where the optimization for θt+1 is initialized from the previous checkpoint θt to retain historical robustness. The specific mechanisms of the adversary and the guardrail are detailed in the following sections. 3.2
DARWIN-ATTACK : A N E VOLVING A DVERSARY
DARWIN-Attack is designed to expose emerging vulnerabilities in language models and safety guardrails under a practical black-box setting. The adversary maintains a dynamic attack strategy pool St at evolution round t. Let x denote a user prompt from a harmful dataset Dh . A strategy sequence π = (s1 , . . . , sn ) selected from the pool transforms the original prompt into a jailbreak prompt, i.e., x̃ = π(x) = sn (sn−1 (· · · s1 (x))) . (2) A valid jailbreak prompt must preserve the original malicious intent. Here we use a binary intentconsistency function Cint (x, x̃) ∈ {0, 1}, where Cint (x, x̃) = 1 if x̃ preserves the harmful intent specified in x, and Cint (x, x̃) = 0 otherwise. The adversary evaluates the effectiveness of the strategy pool against a target LLM M and a safety guardrail Gθ . For LLM evaluation, an attack is successful when the generated response fulfills the harmful intent specified in the original query. We represent this using a binary harmful-intent fulfillment function H(x, y) ∈ {0, 1}, where H(x, y) = 1 if response y fulfills the harmful intent 4
of x, and H(x, y) = 0 otherwise. The attack objective maximizes the attack success rate over the harmful dataset: 1 X RLLM (St ) = max Cint (x, π(x)) · H(x, M(π(x))), (3) π∈Πt |Dh | x∈Dh
where Πt denotes the space of all possible strategy sequences derived from St . For guardrail evaluation, the defense system outputs 0 for safe inputs and 1 for unsafe inputs. An attack is successful when the guardrail incorrectly classifies a jailbreak prompt as safe. The attack success rate is defined as 1 X RGuard (St , Gθ ) = (4) max Cint (x, π(x)) · I [Gθ (π(x)) = 0] . π∈Πt |Dh | x∈Dh
To optimize the attack capabilities, DARWIN-Attack continuously expands and refines the strategy pool. The framework collects external strategies from public platforms and applies genetic algorithms to generate new candidates through crossover and mutation operations. A candidate strategy can be incorporated into the pool only if the corresponding attack success rate on a local sandbox aligned LLM ASRsb (s) exceeds a predefined admission threshold τsb . Combining multiple strategies can create various disguising schemes, but how to select and combine strategies for adaptation to diverse defense mechanisms remains a challenging problem. DARWINAttack models the strategy selection as a Markov transition process. Let Tt ∈ RK×K denote the transition matrix at round t, where Tt (i, j) represents the preference for selecting strategy sj immediately after strategy si . Upon a failed attempt with strategy si , the attack module samples the next strategy based on the corresponding row in Tt and composes the selected strategy with the current jailbreak prompt. After observing the attack result, the framework updates the transition matrix using a Q-learning inspired rule: Tt (i, j) ← Tt (i, j) + α r + γ max Tt (j, k) − Tt (i, j) , (5) k
where r represents the reward from the attack judge, α represents the update rate, and γ represents the discount factor. The updated row is subsequently normalized to form the next transition probability distribution. The adaptive update enables DARWIN-Attack to learn target-specific strategy sequences directly from attack feedback. Failed attempts provide crucial signals for evolution. When an attack is blocked, DARWIN-Attack analyzes the rejection reason via LLM reflection and refines the applied strategy. The refined strategies undergo sandbox validation before joining the active pool, ensuring continuous adaptation against robust defenses. 3.3
DARWIN-G UARD : O NLINE A DVERSARIAL T RAINING
DARWIN-Guard is designed to perform reliable safety classification on input prompts. Let Ph and Pb denote the harmful and benign input distributions. Harmful inputs include both direct malicious requests and complex jailbreak prompts. The defense objective is to maximize the harmful input rejection rate RHarm and the benign input pass rate RBenign . RHarm (Gθ ) = Pr [Gθ (x) = 1] ,
RBenign (Gθ ) = Pr [Gθ (x) = 0] .
x∼Ph
x∼Pb
(6)
Optimizing discrete predictions directly is mathematically intractable. DARWIN-Guard establishes an online adversarial training pipeline. To ensure a strong defensive baseline, we initialize the guard model Gθ based on Qwen3Guard. Let D represent a dataset containing harmful and benign queries. For each query x ∈ D, the label yx = 1 if x is harmful, and yx = 0 if x is benign. During online adversarial training, DARWIN-Attack applies disguise strategies to both harmful and benign prompts in D. The framework generates adversarial variants x̃ designed to cause misclassification by the guard model Gθ while preserving the original intent. The symmetric design forces the guard model to recognize underlying semantics and prevents the model from forming spurious correlations between complex formats and malicious intents. 5
For each sample x in a training batch Dbatch ⊂ D, DARWIN-Attack executes a maximum of Nmax disguise attempts against the guard model Gθ . Finding a candidate that causes misclassification triggers early stopping, and the framework records the successful candidate directly as the adversarial sample x̃. If all Nmax attempts fail, the framework selects the final generated candidate as x̃. DARWIN-Guard minimizes the joint training objective below. ρ(θ) = E(x,yx )∼Dbatch [LCE (Gθ (x̃), yx )] + λraw E(x,yx )∼Dbatch [LCE (Gθ (x), yx )] .
(7)
Both terms utilize standard cross-entropy loss LCE . The first term calculates the prediction loss on the disguised sample x̃. Minimizing the first term forces the guard model to correct previous misclassifications and align with the true intent label yx . The second term anchors the optimization on the original raw prompt x to maintain baseline performance on standard inputs. A weight parameter λraw controls the impact of the second term. Finally, the framework applies the batch loss to update the parameters of Gθ .
4
E XPERIMENTS
4.1
E XPERIMENTAL S ETUP
4.1.1
ATTACK SETUP
Datasets. Following previous jailbreak studies(Zou et al., 2023; Liu et al., 2025a), we evaluate all attack methods on two widely used datasets, HarmBench (Mazeika et al., 2024) and AdvBench (Zou et al., 2023). HarmBench provides 400 malicious instructions covering various harmful categories. AdvBench supplies an additional 520 queries targeting dangerous behaviors. The broad coverage of malicious topics facilitates a comprehensive assessment of attack effectiveness. Target Models. We evaluate DARWIN-Attack against four frontier closed source language models and two safety guardrails. The language models are DeepSeek-V4-Pro, GPT-5.5, Gemini-3.5-Flash, and Claude Sonnet 4.6. The safety guardrails are Qwen3Guard and YuFeng-XGuard. The evaluation covers both model level safety alignment and external guardrail filtering. Baseline Attacks. We compare DARWIN-Attack with four jailbreak baselines. TAP (Mehrotra et al., 2024) explores attack prompts through iterative generation and pruning. ReNeLLM (Ding et al., 2024) rewrites harmful requests and embeds the rewritten requests in benign scenarios. AutoDAN-Turbo (Liu et al., 2025a) discovers and integrates new attack strategies through lifelong learning. MAJIC (Qi et al., 2026a) combines diverse disguise strategies to construct jailbreak prompts. Evaluation Metrics. To ensure a comprehensive and convincing assessment, we employ two independent judge criteria (ASR and HS) alongside query efficiency (AQC). (1) Attack Success Rate (ASR). For language models, a GPT-4 judge assigns each response a harmfulness score from 1 to 5 following the rubric of Qi et al. (2024); Yang et al. (2025). A response counts as a successful jailbreak only when the score is 5. For safety guardrails, an attack succeeds when a harmful input is classified as safe. (2) Harmful Score (HS). For language models, we compute HS following Souly et al. (2024). The metric evaluates responses based on refusal, specificity, and convincingness and ranges from 0 to 1. A higher HS indicates a more harmful response. HS does not apply to safety guardrails because they do not generate responses. (3) Average Query Count (AQC). AQC is the average number of queries used for each attack instance. For a successful attack, we count the number of queries required to achieve the first success; for an unsuccessful attack, we count the full query budget. All attack methods are evaluated under the same query budget. 4.1.2
G UARDRAIL SETUP
Datasets. We evaluate guard models on two benchmark sets, covering harmful-prompt detection and benign-prompt preservation. The harmful-prompt set consists of unsafe prompts from XSTest (Röttger et al., 2024), Aegis2.0 (Ghosh et al., 2025), JailbreakBench behaviors (Chao et al., 2024a), HarmBench (Mazeika et al., 2024), ToxicChat (Lin et al., 2023), JailbreakV RedTeam 2K (Luo et al., 2024), Semantic Router jailbreak (Team, 2026), BeaverTails (Ji et al., 2023), OpenAI Moderation (Markov et al., 2023), WildGuardTest (Han et al., 2024), StrongREJECT (Souly et al., 2024), and JailbreakHub (Shen et al., 2024). The benign-prompt set evaluates benign-prompt preservation on 6
Table 1: Jailbreak attack results on HarmBench and AdvBench. We report Attack Success Rate (ASR) and Average Query Count (AQC). Higher ASR indicates stronger attack effectiveness, while lower AQC indicates better query efficiency. The DARWIN-Attack rows report our method. Dataset
Method
DeepSeek-V4-Pro
GPT-5.5
Gemini-3.5-Flash Claude Sonnet 4.6 Qwen3Guard YuFeng-XGuard
ASR ↑
AQC ↓
ASR ↑ AQC ↓ ASR ↑
AQC ↓
ASR ↑
AQC ↓
ASR ↑ AQC ↓ ASR ↑
AQC ↓
TAP ReNeLLM HarmBench AutoDAN-Turbo MAJIC DARWIN-Attack
52.7% 59.2% 75.2% 83.2% 99.7%
38.4 36.7 24.9 10.6 2.2
35.5% 39.5% 48.7% 68.5% 92.5%
47.8 44.5 39.7 22.4 15.0
46.7% 52.2% 65.5% 79.7% 92.2%
41.6 37.6 31.8 23.1 12.2
1.5% 4.2% 11.2% 30.7% 76.7%
59.2 54.8 52.4 46.5 18.5
58.7% 64.5% 79.2% 84.2% 99.7%
29.7 22.4 15.8 7.4 2.5
44.2% 51.7% 64.7% 82.7% 99.0%
37.3 33.6 27.1 18.1 3.1
TAP ReNeLLM AutoDAN-Turbo MAJIC DARWIN-Attack
48.7% 56.7% 71.7% 84.7% 97.5%
37.1 34.3 22.4 10.2 3.1
32.5% 36.2% 45.2% 65.2% 90.2%
47.2 46.1 33.2 24.1 16.2
43.2% 49.7% 62.7% 76.7% 89.7%
33.4 29.2 30.5 24.8 16.5
1.2% 3.7% 9.7% 27.5% 62.5%
59.5 56.4 48.1 39.3 19.9
55.2% 61.2% 76.7% 81.7% 98.7%
31.1 26.8 17.2 10.8 5.1
41.7% 48.5% 61.2% 79.5% 96.2%
36.8 31.1 28.6 19.6 6.9
AdvBench
ARC Challenge (Clark et al., 2018), ARC Easy (Clark et al., 2018), BoolQ (Clark et al., 2019), GSM8K (Cobbe et al., 2021), OpenBookQA (Mihaylov et al., 2018), AG News (Zhang et al., 2015), HotpotQA (Yang et al., 2018), QASC (Khot et al., 2020), RACE (Lai et al., 2017), SciQ (Welbl et al., 2017), and COPA (Wang et al., 2019). Baseline Guardrails. We compare DARWIN-Guard with representative guard models, including ShieldGemma (Zeng et al., 2024), Llama-3.1-Nemotron-Safety-Guard-8B-v3 (Joshi et al., 2025), Granite Guardian 4.1 8B (Padhi et al., 2024), Llama-Guard-3-8B (Inan et al., 2023), Qwen3GuardGen-8B (Zhao et al., 2025), and YuFeng-XGuard-Reason-8B (Lin et al., 2026). Evaluation Metrics. On the harmful-prompt set, we report unsafe recall, defined as the fraction of harmful prompts predicted as unsafe. On the benign-prompt set, we report safe pass rate, defined as the fraction of benign prompts predicted as safe. Average scores are computed as macro-averages across benchmarks. All scores are reported as percentages. 4.2
DARWIN-ATTACK E VALUATION
Table 1 reports the attack success rate and average query count of different jailbreak methods. DARWIN-Attack achieves the highest ASR across all evaluated LLMs and guardrails on both HarmBench and AdvBench. Compared with the strongest baseline, MAJIC, DARWIN-Attack improves ASR by 12.5–46.0 percentage points on HarmBench and 12.8–35.0 percentage points on AdvBench. The largest improvements are observed on Claude Sonnet 4.6, where DARWIN-Attack increases ASR from 30.7% to 76.7% on HarmBench and from 27.5% to 62.5% on AdvBench. For Qwen3Guard and YuFeng-XGuard, DARWIN-Attack achieves ASRs of at least 96.2% across the two datasets. Under the same query budget, DARWIN-Attack also achieves the lowest AQC for every evaluated target, indicating better query efficiency. Figure 2 compares the Harmful Scores of different attack methods. DARWIN-Attack achieves the highest HS for every evaluated LLM on both datasets. In particular, it improves HS over MAJIC from 0.50 to 0.66 on HarmBench and from 0.47 to 0.63 on AdvBench for DeepSeek-V4-Pro. On Claude Sonnet 4.6, DARWIN-Attack improves HS from 0.18 to 0.26 on HarmBench and from 0.15 to 0.24 on AdvBench. These results demonstrate that DARWIN-Attack achieves state-of-the-art attack performance under both independent judge metrics, eliciting responses with higher harmfulness. 4.3
DARWIN-G UARD E VALUATION
We report the main DARWIN-Guard results in Tables 2 and 3. The former evaluates unsafe recall on the harmful benchmark set, while the latter evaluates safe pass rate on the benign benchmark set. Table 2 shows unsafe recall on harmful benchmarks. DARWIN-Guard reaches an average recall of 91.6%, exceeding Qwen3Guard-Gen-8B by 5.7 percentage points and YuFeng-XGuard-Reason-8B by 4.4 percentage points. It obtains the best or tied-best score on eleven of the twelve harmful benchmarks, including 99.5% on XSTest, 100.0% on JailbreakBench behaviors, 99.8% on HarmBench, and 99.7% on StrongREJECT. 7
Figure 2: Harmful Score (HS) of different jailbreak methods on HarmBench and AdvBench. As an independent evaluation metric alongside ASR, a higher HS indicates more harmful responses. DARWIN-Attack still achieves the highest HS across all evaluated LLMs on both datasets.
Table 2: Unsafe Recall on Harmful Prompt Benchmarks. Scores are percentages. Higher scores mean better detection of unsafe prompts. The best score in each row is bolded. Dataset Model
Shield Gemma
Nemotron Guard
Granite Guardian
Llama Guard-3
Qwen3 Guard
YuFeng XGuard
DARWIN Guard
XSTest Aegis2.0 JBB-Behaviors HarmBench ToxicChat JailbreakV-RT2K Semantic Router BeaverTails OpenAI Moderation WildGuardTest StrongREJECT JailbreakHub
86.0 70.0 54.0 45.5 61.9 43.6 46.8 64.0 92.1 41.2 76.0 33.2
93.0 87.3 92.0 68.5 79.3 69.2 74.8 78.8 96.4 83.0 99.4 74.8
96.5 84.5 97.0 74.5 77.9 66.8 74.8 76.4 89.5 73.8 99.4 77.2
82.0 66.2 98.0 97.2 50.0 52.0 48.0 57.2 78.5 66.6 97.4 31.2
92.0 84.2 98.0 98.2 88.1 64.8 74.8 75.6 91.6 84.8 98.4 80.4
98.0 87.6 99.0 75.5 92.0 68.4 80.8 79.6 97.7 87.6 99.7 80.8
99.5 93.5 100.0 99.8 91.7 75.6 85.2 82.8 98.0 90.2 99.7 83.2
Average
59.5
83.0
82.4
68.7
85.9
87.2
91.6
The gains are particularly pronounced on jailbreak-oriented benchmarks. Compared with Qwen3Guard-Gen-8B, DARWIN-Guard improves recall by 10.8 points on JailbreakV RedTeam 2K, 10.4 points on Semantic Router, 5.4 points on WildGuardTest, and 2.8 points on JailbreakHub. These gains are consistent with our attack-driven training design. DARWIN-Attack exposes guardspecific false negatives, and these hard examples help DARWIN-Guard better recognize unsafe intent across different disguise strategies. The benign benchmark results in Table 3 show that DARWIN-Guard does not improve harmful recall by simply rejecting more inputs. DARWIN-Guard keeps an average safe pass rate of 100.0% while achieving the highest average recall on the harmful benchmark set. It reaches 100.0% safe pass rate on every listed benign benchmark. These results support our use of disguised benign prompts and benign anchors during training. Disguised benign prompts prevent the guard from treating disguise patterns themselves as unsafe, while benign anchors preserve the decision boundary on ordinary benign inputs. This design reduces the risk that adversarial training shifts the guard toward blanket rejection. 4.4
A NALYSIS OF THE E VOLVING L OOP
In this section, we analyze the dynamics of the evolving attack-defense loop. As DARWINAttack evolves, it continuously discovers and integrates new jailbreak strategies. Correspondingly, DARWIN-Guard performs online adversarial training on the dynamic attacks generated at each 8
Table 3: Safe Pass Rate on Benign Benchmarks. Scores are percentages. Higher scores mean fewer false refusals. Dataset Model
Shield Gemma
Nemotron Guard
Granite Guardian
Llama Guard-3
Qwen3 Guard
YuFeng XGuard
DARWIN Guard
ARC-Challenge ARC-Easy BoolQ GSM8K OpenBookQA AG News HotpotQA QASC RACE SciQ COPA
100.0 100.0 99.6 100.0 99.8 100.0 99.8 99.8 100.0 100.0 100.0
100.0 100.0 99.8 99.4 99.4 96.6 97.8 98.6 96.8 100.0 100.0
99.6 100.0 99.6 100.0 99.8 99.8 99.8 100.0 100.0 100.0 100.0
100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 99.8 100.0 100.0
100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0
100.0 100.0 100.0 99.8 99.8 99.4 100.0 100.0 100.0 100.0 100.0
100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0 100.0
Average
99.9
98.9
99.9
100.0
100.0
99.9
100.0
stage. To demonstrate this progressive enhancement, we track the capabilities of both modules as the active strategy pool grows from 50 to 200 strategies. Scaling attack capabilities. Table 4 illustrates the Attack Success Rate (ASR) of DARWINAttack on HarmBench as its strategy pool expands. We select GPT-5.5 and YuFeng-XGuard as representative targets. With an initial pool of 50 strategies, DARWIN-Attack achieves ASRs of 73.5% and 82.2%, respectively. As the evolution progresses and more strategies are integrated, the ASR steadily climbs, ultimately reaching 92.5% against GPT-5.5 and 99.0% against YuFengXGuard at 200 strategies. These results show that the attack evolution in DARWIN continually improves its ability to expose vulnerabilities in advanced models and guardrails. Table 4: Attack performance scaling as the strategy pool expands. Evaluated on HarmBench, ASR steadily increases as new strategies are integrated into the evolving loop. DARWIN-Attack Strategy Pool Size
Target Model GPT-5.5 YuFeng-XGuard
50
75
100
125
150
175
200 (Final)
73.5% 82.2%
77.8% 85.5%
79.5% 87.7%
81.2% 92.0%
86.2% 95.7%
89.7% 97.2%
92.5% ↑ 99.0% ↑
Scaling guardrail robustness. On the defense side, DARWIN-Guard improves its robustness through the evolving loop. We evaluate intermediate guardrail checkpoints, denoted as Guard-50 to Guard-200, which are trained on the adversarial datasets generated at the corresponding attack stages. Figure 3 reports their unsafe recall across several representative and challenging safety benchmarks, such as JailbreakV-RT2K and Semantic Router. Guard-50 provides a basic defense against the initial attack patterns. As the training distribution expands with more attack strategies, the guardrail becomes better at recognizing underlying malicious intents. Consequently, the unsafe recall steadily climbs across all evaluated datasets. These results show that the evolving loop effectively enhances the defensive capabilities of the guardrail against jailbreak attacks. Cross-stage evaluation and forward robustness. To demonstrate that the performance gains stem from the evolving loop rather than simply accumulating more data, we conduct a cross-stage evaluation. Figure 4(A) presents a cross-play matrix evaluating Attack Success Rates (ASR). Here, Attack-K denotes the adversarial prompts generated at the attack stage with a K-strategy pool, and Guard-K represents the guardrail checkpoint trained at that corresponding stage. To assess generalization, we also introduce a held-out attack set containing complex jailbreak strategy families that are completely excluded from the evolving training process, with results shown in Figure 4(B). This cross-stage evaluation reveals three key findings. First, later attackers easily bypass earlier guardrails (e.g., Attack-150 against Guard-50 yields a 84% ASR). This confirms that static defenses 9
Figure 3: Unsafe recall of intermediate DARWIN-Guard checkpoints across representative safety benchmarks. Guard-K denotes the guardrail trained on adversarial examples generated at the Kstrategy attack stage. The unsafe recall steadily climbs across all evaluated datasets, demonstrating continuous improvement in defensive capabilities through the evolving loop.
Figure 4: Cross-stage evaluation and forward robustness. (A) Cross-play ASR matrix between attacker stages and guardrail checkpoints. Lower ASR (blue) indicates stronger guardrail defense. The matrix shows that static guardrails fail against future attacks, while evolved guardrails maintain historical robustness. (B) ASR on a completely held-out attack set across guardrail checkpoints. The continuous decline in ASR demonstrates that the evolving loop equips the guardrail with generalizable robustness against unseen threats. inevitably fall behind emerging adversarial threats. Second, later guardrails maintain low ASR against earlier attacks, indicating that the defense retains historical robustness without catastrophic forgetting. Finally, and most importantly, the guardrail maintains strong defensive performance against the held-out attack set as it evolves (dropping from 54% to 26% ASR). This generalization demonstrates that the evolving loop effectively forces the guardrail to recognize underlying malicious intents, rather than merely memorizing specific disguise formats observed during training.
R EFERENCES Patrick Chao, Edoardo Debenedetti, Alexander Robey, Maksym Andriushchenko, Francesco Croce, Vikash Sehwag, Edgar Dobriban, Nicolas Flammarion, George J Pappas, Florian Tramer, et al. Jailbreakbench: An open robustness benchmark for jailbreaking large language models. Advances in Neural Information Processing Systems, 37:55005–55029, 2024a. Patrick Chao, Alexander Robey, Edgar Dobriban, Hamed Hassani, George J. Pappas, and Eric Wong. Jailbreaking black box large language models in twenty queries, 2024b. URL https: //arxiv.org/abs/2310.08419. 10
Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. In Proceedings of the 2019 conference of the north American chapter of the association for computational linguistics: Human language technologies, volume 1 (long and short papers), pp. 2924–2936, 2019. Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv preprint arXiv:1803.05457, 2018. Karl Cobbe, Vineet Kosaraju, Mohammad Bavarian, Mark Chen, Heewoo Jun, Lukasz Kaiser, Matthias Plappert, Jerry Tworek, Jacob Hilton, Reiichiro Nakano, et al. Training verifiers to solve math word problems. arXiv preprint arXiv:2110.14168, 2021. Peng Ding, Jun Kuang, Dan Ma, Xuezhi Cao, Yunsen Xian, Jiajun Chen, and Shujian Huang. A wolf in sheep’s clothing: Generalized nested jailbreak prompts can fool large language models easily. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 2136–2153, 2024. Shaona Ghosh, Prasoon Varshney, Makesh Narsimhan Sreedhar, Aishwarya Padmakumar, Traian Rebedea, Jibin Rajan Varghese, and Christopher Parisien. Aegis2. 0: A diverse ai safety dataset and risks taxonomy for alignment of llm guardrails. In Proceedings of the 2025 Conference of the Nations of the Americas Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 5992–6026, 2025. Seungju Han, Kavel Rao, Allyson Ettinger, Liwei Jiang, Bill Yuchen Lin, Nathan Lambert, Yejin Choi, and Nouha Dziri. Wildguard: Open one-stop moderation tools for safety risks, jailbreaks, and refusals of llms. Advances in neural information processing systems, 37:8093–8131, 2024. Monte Hoover, Vatsal Baherwani, Neel Jain, Khalid Saifullah, Joseph Vincent, Chirag Jain, Melissa Kazemi Rad, C Bayan Bruss, Ashwinee Panda, and Tom Goldstein. Dynaguard: A dynamic guardian model with user-defined policies. arXiv preprint arXiv:2509.02563, 2025. Xinzhe Huang, Kedong Xiu, Tianhang Zheng, Churui Zeng, Wangze Ni, Zhan Qin, Kui Ren, and Chun Chen. Dualbreach: Efficient dual-jailbreaking via target-driven initialization and multitarget optimization. arXiv preprint arXiv:2504.18564, 2025. Xinzhe Huang, Wenjing Hu, Tianhang Zheng, Kedong Xiu, Hongsheng Hu, Xiaojun Jia, Di Wang, Zhan Qin, and Kui Ren. Nontextual target attack, 2026. URL https://arxiv.org/abs/ 2510.02999. Hakan Inan, Kartikeya Upasani, Jianfeng Chi, Rashi Rungta, Krithika Iyer, Yuning Mao, Michael Tontchev, Qing Hu, Brian Fuller, Davide Testuggine, et al. Llama guard: Llm-based input-output safeguard for human-ai conversations. arXiv preprint arXiv:2312.06674, 2023. Jiaming Ji, Mickel Liu, Josef Dai, Xuehai Pan, Chi Zhang, Ce Bian, Boyuan Chen, Ruiyang Sun, Yizhou Wang, and Yaodong Yang. Beavertails: Towards improved safety alignment of llm via a human-preference dataset. Advances in Neural Information Processing Systems, 36:24678– 24704, 2023. Raviraj Bhuminand Joshi, Rakesh Paul, Kanishk Singla, Anusha Kamath, Michael Evans, Katherine Luna, Shaona Ghosh, Utkarsh Vaidya, Eileen Margaret Peters Long, Sanjay Singh Chauhan, et al. Cultureguard: Towards culturally-aware dataset and guard model for multilingual safety applications. In Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics, pp. 2666–2685, 2025. Tushar Khot, Peter Clark, Michal Guerquin, Peter Jansen, and Ashish Sabharwal. Qasc: A dataset for question answering via sentence composition. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 34, pp. 8082–8090, 2020. 11
Guokun Lai, Qizhe Xie, Hanxiao Liu, Yiming Yang, and Eduard Hovy. Race: Large-scale reading comprehension dataset from examinations. In Proceedings of the 2017 conference on empirical methods in natural language processing, pp. 785–794, 2017. Junyu Lin, Meizhen Liu, Xiufeng Huang, Jinfeng Li, Haiwen Hong, Xiaohan Yuan, Yuefeng Chen, Longtao Huang, Hui Xue, Ranjie Duan, et al. Yufeng-xguard: A reasoning-centric, interpretable, and flexible guardrail model for large language models. arXiv preprint arXiv:2601.15588, 2026. Zi Lin, Zihan Wang, Yongqi Tong, Yangkun Wang, Yuxin Guo, Yujia Wang, and Jingbo Shang. Toxicchat: Unveiling hidden challenges of toxicity detection in real-world user-ai conversation. In Findings of the Association for Computational Linguistics: EMNLP 2023, pp. 4694–4702, 2023. Xiaogeng Liu, Nan Xu, Muhao Chen, and Chaowei Xiao. Autodan: Generating stealthy jailbreak prompts on aligned large language models. In International Conference on Learning Representations, volume 2024, pp. 56174–56194, 2024. Xiaogeng Liu, Peiran Li, G Edward Suh, Yevgeniy Vorobeychik, Zhuoqing Mao, Somesh Jha, Patrick McDaniel, Huan Sun, Bo Li, and Chaowei Xiao. Autodan-turbo: A lifelong agent for strategy self-exploration to jailbreak llms. In International Conference on Learning Representations, volume 2025, pp. 10313–10360, 2025a. Yue Liu, Hongcheng Gao, Shengfang Zhai, Yufei He, Jun Xia, Zhengyu Hu, Yulin Chen, Xihong Yang, Jiaheng Zhang, Stan Z Li, et al. Guardreasoner: Towards reasoning-based llm safeguards. arXiv preprint arXiv:2501.18492, 2025b. Weidi Luo, Siyuan Ma, Xiaogeng Liu, Xiaoyu Guo, and Chaowei Xiao. Jailbreakv: A benchmark for assessing the robustness of multimodal large language models against jailbreak attacks. arXiv preprint arXiv:2404.03027, 2024. Todor Markov, Chong Zhang, Sandhini Agarwal, Florentine Eloundou Nekoul, Theodore Lee, Steven Adler, Angela Jiang, and Lilian Weng. A holistic approach to undesired content detection in the real world. In Proceedings of the AAAI conference on artificial intelligence, volume 37, pp. 15009–15018, 2023. Mantas Mazeika, Long Phan, Xuwang Yin, Andy Zou, Zifan Wang, Norman Mu, Elham Sakhaee, Nathaniel Li, Steven Basart, Bo Li, et al. Harmbench: A standardized evaluation framework for automated red teaming and robust refusal. arXiv preprint arXiv:2402.04249, 2024. Anay Mehrotra, Manolis Zampetakis, Paul Kassianik, Blaine Nelson, Hyrum Anderson, Yaron Singer, and Amin Karbasi. Tree of attacks: Jailbreaking black-box llms automatically. Advances in Neural Information Processing Systems, 37:61065–61105, 2024. Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 conference on empirical methods in natural language processing, pp. 2381–2391, 2018. Inkit Padhi, Manish Nagireddy, Giandomenico Cornacchia, Subhajit Chaudhury, Tejaswini Pedapati, Pierre Dognin, Keerthiram Murugesan, Erik Miehling, Martı́n Santillán Cooper, Kieran Fraser, et al. Granite guardian. arXiv preprint arXiv:2412.07724, 2024. Weiwei Qi, Shuo Shao, Wei Gu, Tianhang Zheng, Puning Zhao, Zhan Qin, and Kui Ren. Majic: Markovian adaptive jailbreaking via iterative composition of diverse innovative strategies. In Proceedings of the AAAI Conference on Artificial Intelligence, volume 40, pp. 32755–32763, 2026a. Weiwei Qi, Zefeng Wu, Tianhang Zheng, Zikang Zhang, Xiaojun Jia, Zhan Qin, and Kui Ren. Towards identification and intervention of safety-critical parameters in large language models. In Findings of the Association for Computational Linguistics: ACL 2026, pp. 32293–32312, 2026b. Xiangyu Qi, Yi Zeng, Tinghao Xie, Pin-Yu Chen, Ruoxi Jia, Prateek Mittal, and Peter Henderson. Fine-tuning aligned language models compromises safety, even when users do not intend to! In International Conference on Learning Representations, volume 2024, pp. 30988–31043, 2024. 12
Kui Ren, Tianhang Zheng, Zhan Qin, and Xue Liu. Adversarial attacks and defenses in deep learning. Engineering, 6(3):346–360, 2020. Paul Röttger, Hannah Kirk, Bertie Vidgen, Giuseppe Attanasio, Federico Bianchi, and Dirk Hovy. Xstest: A test suite for identifying exaggerated safety behaviours in large language models. In Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), pp. 5377–5400, 2024. Xinyue Shen, Zeyuan Chen, Michael Backes, Yun Shen, and Yang Zhang. ” do anything now”: Characterizing and evaluating in-the-wild jailbreak prompts on large language models. In Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, pp. 1671–1685, 2024. Alexandra Souly, Qingyuan Lu, Dillon Bowen, Tu Trinh, Elvis Hsieh, Sana Pandey, Pieter Abbeel, Justin Svegliato, Scott Emmons, Olivia Watkins, et al. A strongreject for empty jailbreaks. Advances in Neural Information Processing Systems, 37:125416–125440, 2024. LLM Semantic Router Team. Jailbreak detection dataset, 2026. URL https://huggingface. co/datasets/llm-semantic-router/jailbreak-detection-dataset. Alex Wang, Yada Pruksachatkun, Nikita Nangia, Amanpreet Singh, Julian Michael, Felix Hill, Omer Levy, and Samuel Bowman. Superglue: A stickier benchmark for general-purpose language understanding systems. Advances in neural information processing systems, 32, 2019. Alexander Wei, Nika Haghtalab, and Jacob Steinhardt. Jailbroken: How does llm safety training fail? Advances in neural information processing systems, 36:80079–80110, 2023. Johannes Welbl, Nelson F Liu, and Matt Gardner. Crowdsourcing multiple choice science questions. In Proceedings of the 3rd Workshop on Noisy User-generated Text, pp. 94–106, 2017. Zefeng Wu, Weiwei Qi, Jielong Chen, Tianhang Zheng, Di Hong, Chaochao Lu, Liang He, Zhan Qin, and Kui Ren. Datashield: Uncovering risky fine-tuning data across llms through consensus subspace alignment. arXiv preprint arXiv:2607.15081, 2026. Kedong Xiu, Churui Zeng, Tianhang Zheng, Xinzhe Huang, Xiaojun Jia, Di Wang, Puning Zhao, Zhan Qin, and Kui Ren. Dynamic target attack. arXiv preprint arXiv:2510.02422, 2025. Langqi Yang, Tianhang Zheng, Yixuan Chen, Kedong Xiu, Hao Zhou, Di Wang, Puning Zhao, Zhan Qin, and Kui Ren. Harmmetric eval: Benchmarking metrics and judges for llm harmfulness assessment. arXiv preprint arXiv:2509.24384, 2025. Zhilin Yang, Peng Qi, Saizheng Zhang, Yoshua Bengio, William Cohen, Ruslan Salakhutdinov, and Christopher D Manning. Hotpotqa: A dataset for diverse, explainable multi-hop question answering. In Proceedings of the 2018 conference on empirical methods in natural language processing, pp. 2369–2380, 2018. Wenjun Zeng, Yuchi Liu, Ryan Mullins, Ludovic Peran, Joe Fernandez, Hamza Harkous, Karthik Narasimhan, Drew Proud, Piyush Kumar, Bhaktipriya Radharapu, et al. Shieldgemma: Generative ai content moderation based on gemma. arXiv preprint arXiv:2407.21772, 2024. Junke Zhang, Jianwei Wang, Sishuo Chen, Yizhang He, Qingshuai Feng, and Zhengyi Yang. Evolving skill-structured attack memory enhances llm jailbreaking, 2026. URL https://arxiv. org/abs/2605.29237. Xiang Zhang, Junbo Zhao, and Yann LeCun. Character-level convolutional networks for text classification. Advances in neural information processing systems, 28, 2015. Haiquan Zhao, Chenhan Yuan, Fei Huang, Xiaomeng Hu, Yichang Zhang, An Yang, Bowen Yu, Dayiheng Liu, Jingren Zhou, Junyang Lin, et al. Qwen3guard technical report. arXiv preprint arXiv:2510.14276, 2025. Tianhang Zheng, Changyou Chen, and Kui Ren. Distributionally adversarial attack. In Proceedings of the AAAI conference on artificial intelligence, volume 33, pp. 2253–2260, 2019. 13
Andy Zou, Zifan Wang, Nicholas Carlini, Milad Nasr, J Zico Kolter, and Matt Fredrikson. Universal and transferable adversarial attacks on aligned language models. arXiv preprint arXiv:2307.15043, 2023.
14