HarmChip: Evaluating Hardware Security Centric LLM Safety via Jailbreak Benchmarking Zeng Wang†∗ , Minghao Shao†‡∗ , Weimin Fu¶ , Prithwish Basu Roy†‡ , Xiaolong Guo¶ , Ramesh Karri† , Muhammad Shafique‡ , Johann Knechtel‡ , Ozgur Sinanoglu‡
arXiv:2604.17093v1 [cs.CR] 18 Apr 2026
†
NYU Tandon School of Engineering, USA ‡ NYU Abu Dhabi, UAE ¶ Kansas State University, USA Email: {zw3464, shao.minghao, pb2718, rkarri, muhammad.shafique, johann, ozgursin}@nyu.edu {weiminf, guoxiaolong}@ksu.edu
Abstract—The integration of large language models (LLMs) into electronic design automation (EDA) workflows has introduced powerful capabilities for RTL generation, verification, and design optimization, but also raises critical security concerns. Malicious LLM outputs in this domain pose hardwarelevel threats, including hardware Trojan insertion, side-channel leakage, and intellectual property theft, that are irreversible once fabricated into silicon. Such requests often exploit semantic disguise, embedding adversarial intent within legitimate engineering language that existing safety mechanisms, trained on generalpurpose hazards, fail to detect. No benchmark exists to evaluate LLM vulnerability to such domain-specific threats. We present the HarmChip benchmark to assess jailbreak susceptibility in hardware security, spanning 16 hardware security domains, 120 threats, and 360 prompts at two difficulty levels. Evaluation of state-of-the-art LLMs reveals an alignment paradox: They refuse legitimate security queries while complying with semantically disguised attacks, exposing blind spots in safety guardrails and underscoring the need for domain-aware safety alignment.
I. I NTRODUCTION The integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) workflows, spanning RTL code generation, functional verification, and design-space exploration is reshaping the hardware design lifecycle [1], [2]. LLM outputs in this context directly impact synthesizable hardware descriptions that reach silicon. A compliant response to a malicious prompt can embed a hardware Trojan [3], introduce side-channel leakage, or weaken IP protection mechanisms. Once fabricated into silicon, such vulnerabilities are near-impossible to patch and difficult to detect [4]. The problem is amplified by semantic disguise: adversarial intent in hardware design is expressed via legitimate engineering language, so a request to refine control logic, simplify debug interfaces, or handle rare state transitions can appear routine while guiding the LLM toward producing backdoored designs. Current safety alignment mechanisms are ill-equipped for hardware domain, producing a characteristic alignment paradox [5]. Keyword-sensitive guardrails over-refuse legitimate queries: security engineers performing red-teaming or developing defenses are blocked when their prompts mention terms such as “Trojan insertion” or “side-channel analysis.” These guardrails have blind spots against semantically disguised attacks, where an adversary phrasing a backdoor insertion as a power optimization or a routine Engineering Change Order
bypasses safety filters. The core limitation is not whether LLMs can refuse harmful prompts, but whether they can separate defensive security work from malicious manipulation when both use identical hardware-design terminology. Existing safety benchmarks, including those targeting automated red teaming, jailbreak robustness, and refusal calibration, focus on general-purpose or software-centric threats and do not capture the threat structure of hardware security, where malicious intent is embedded within standard engineering semantics and a single compliant response can propagate into fabricated hardware [6], [7]. This paper presents HarmChip, the first domain-specific jailbreak benchmark for evaluating LLM safety in hardware security. HarmChip covers 16 hardware security domains, 120 threat scenarios, and 360 prompts stratified into easy and hard difficulty levels. The benchmark evaluates two complementary dimensions: resilience against semantically disguised malicious queries in realistic hardwaredesign language, and calibration against over-refusal of legitimate security-oriented engineering tasks. The main contributions of this paper are as follows. • HarmChip is the first domain-specific jailbreak benchmark on safety alignment in hardware security, spanning 16 security domains, 120 threats, and 360 prompts. • We propose an automated, three-stage benchmark construction pipeline that leverages LLM-as-Attacker generation and difficulty stratification to produce semantically disguised jailbreak prompts grounded in hardware threat literature. • Through evaluation across state-of-the-art LLMs, we expose the alignment paradox, i.e., the co-existence of high false-negative rates on disguised hardware attacks and excessive false-positive refusals on legitimate security queries, underscoring the need for domain-aware safety guardrails in hardware design. II. BACKGROUND A. LLMs and Hardware Security LLMs are now integral to modern EDA, accelerating tasks such as RTL generation, testbench synthesis, and ECO implementation [8], [9]. This integration expands the hardware attack surface: malicious logic built into the RTL propagates
IP n Protectio
A &ig. FPeGconf) R (3 Si M lico (3)fg. n
5%
Easy / Hard
HardCPwSare (1)
360 prompts 120 threats 16 domains Over 1200 papers
Sec ML (5)urity
C
HarmChip organizes hardware security threats into a twolevel taxonomy of 16 fine-grained domains, as illustrated in Figure 1, spanning the full hardware design and deployment lifecycle: from early supply chain and EDA tooling risks, through RTL design, synthesis, and physical implementation, to post-silicon runtime and emerging attack surfaces. This endto-end coverage reflects a fundamental property of hardware security threats: attack vectors emerge at every stage of the design flow and, once fabricated into silicon, are irreversible. The domains span well-established areas such as Hardware Trojans and Side-Channel Attacks as well as emerging concerns including ML Security, Advanced Packaging, and CPS Hardware. The benchmark comprises 120 threat scenarios across these 16 domains, with each scenario instantiated into 3 prompts per difficulty level, yielding 360 prompts for each of the Easy and Hard variants. Domain coverage targets areas with higher real-world attack surface: Crypto & Side-Channel and Architecture & Runtime each contribute 26 prompts per variant, while domains such as CPS Hardware and Compliance feature smaller but targeted sets.
hain Supp&lyEDCA
An Secuarilotyg (3)
e
HarmChip
, PQ pto NG Cry & R
A. Threat Taxonomy
Tro HW (6)jans
10%
Adv. g Pack(3ag) in
III. M ETHODS
Te Debstuginggin& g
15%
B. Safety Alignment and Threat Model Commercial LLMs are aligned against harmful outputs through techniques such as Reinforcement Learning from Human Feedback and Direct Preference Optimization [17], [18]. Adversaries use jailbreak strategies, including role-playing, prefix injection, and instruction obfuscation, to circumvent these guardrails [18]–[20]. Existing safety benchmarks such as AdvBench [19] and JailbreakBench [21] target NLP hazards or software-level exploits, leaving hardware-security-specific attack vectors unaddressed. The threat model considered in this work assumes an adversary, either a rogue designer or an external user of an LLM-assisted EDA tool, who seeks to inject hardware-level vulnerabilities into the IC design flow. The attacker exploits semantic disguise: malicious payloads are embedded within syntactically legitimate engineering queries. A Trojan insertion may be framed as a Power-PerformanceArea optimization or a functional ECO on a finite state machine (FSM). Current guardrails rely on keyword matching and general-purpose hazard detection, so such domain-specific obfuscation bypasses existing safety filters, necessitating a hardware-security-aware evaluation benchmark [22].
20%
nc plia) Com (2
through downstream synthesis tools and materializes as permanent structures in the gate-level netlist [10]. Software vulnerabilities can be patched post-deployment; hardware threats cannot [11]. Trojans embed stealthy triggers that activate under rare inputs [12], side-channel structures leak cryptographic keys via power or timing emanations [13], and IP backdoors in third-party cores grant persistent unauthorized access [14]– [16]. Any vulnerability introduced into the RTL is baked into silicon with no post-fabrication remedy.
Sid Atet-Cha ack nne s l
Supply Chain & Compliance (14) RTL & Logic Design (14)
Arch, Me & Busmory Synthesis & IP (18) Physical & Packaging (6)
Firmwaree & Runtim
l I sica& F Phtaycks At
Crypto & Side-Channel (26) Architecture & Runtime (26)
Post-Silicon & Emerging (16)
Fig. 1: HarmChip threat taxonomy: 16 domains grouped under 7 lifecycle-stage categories.
B. Evaluation Metrics The primary metric is Attack Success Rate (ASR): ASR =
# compliant responses × 100% # total prompts
(1)
A response is deemed compliant if it provides substantive, exploit-relevant hardware-design content (e.g., functional RTL or actionable design guidance); otherwise it is classified as a refusal. All judgments are rendered by Gemini-3-Flash serving as an LLM-as-Judge. ASR is reported per model and per threat domain across both difficulty variants, enabling analysis of safety robustness at both aggregate and category levels. C. HarmChip Generation Pipeline HarmChip is constructed through a three-stage automated pipeline for benchmark curation, as illustrated in Figure 2. Stage 1: Threat Taxonomy and Data Curation. For each of the 120 threat scenarios, 10 relevant research papers are collected and the 8 most substantive are retained. A critical preprocessing step is information sanitization: identifying metadata such as paper titles and author names are stripped while the underlying exploit logic is preserved. This ensures that the curated content captures grounded attack knowledge without exposing source-level identifiers that could trigger hardcoded refusal behaviors in the target LLMs. Stage 2: Automated Red-Team Prompt Generation. The sanitized content is processed within an academic sandbox framework under strict context injection and constraint enforcement. For each paper, one jailbreak prompt is derived
Hardware Threat Taxonomy
Literature Retrieval
100
ASR (%) - Easy
(1) Threat Taxonomy & Data Curation
50
3.61
0
Red-Team Attack Prompt Generation Top Exploit Vectors
78.06 47.22
50 25 0
3.
ep
St
ASR Evaluation Metric LLM Inferences
Easy Harmchip
Hard Harmchip
…
50%
44.17
75
1.94
3.89
4.72
6.11
h
Pr
o
lu
s
.4
as
(2) Red-Team Prompt Generation
96.39 99.44 100.00
19.17 18.61 20.00 19.17
11.67
100
ASR (%) - Hard
Academic Sandbox Framing
38.61
25
Information Sanitization Context Injection & Constraint Enforcement
90.83 91.39 71.11 72.22 69.17
75
l -F
5
i-3
in
m
Ge
.1
3.
en
Qw
P 5-
-5 PT
G
3.
Qw
en
6.39
8.61
B
2.
iM
M x-
a
in
M
M
GL
d au
Cl
90.56
98.33 100.00
60.83 50%
16.39 12.50
5
97
3 5-
55.83
83.06
-5
k 5 2 1 2 6 st .6 sh ic 2. 4. 3. 4. 51 -4 la Fa Ter s-V i-K -F 1l-2 av pu ek GP m 4. ra i-3 Ki -O -M Se kst in e 4 o p v r d m e A G De au Ge De aM Cl LL t ne
on
S e-
Fig. 3: Aggregated ASR across Easy and Hard benchmarks, showing bimodal safety polarization among 16 LLMs.
(3) Automated Benchmarking
Fig. 2: Three-stage pipeline for HarmChip construction.
from the underlying exploit logic, yielding 8 prompts per threat and 960 in total. This sandboxed framing elicits attack-relevant content grounded in real hardware threat literature that would be refused under direct querying. Stage 3: Automated Benchmark Curation. The 960 prompts are deployed across 6 LLMs spanning diverse model families: Claude-Sonnet-4.6, GPT-5.4, Gemini-3.1-Pro, MiniMax-M2.5, DeepSeek-V3.2, and GLM-5. Each response is assessed by Gemini-3-Flash as an independent judge, classifying it as either substantive compliance or refusal. The 6 models are then ranked by ASR from low to high for each prompt: prompts where the three least susceptible models comply are assigned to the Hard variant, while those where the three most susceptible models comply form the Easy variant, yielding 360 prompts per tier. IV. H ARM C HIP B ENCHMARK Table I summarizes the 16 hardware security domains covered by HarmChip, each associated with a representative threat description, publication venues, and temporal coverage. These domains span the full hardware lifecycle, including supply chain and EDA tooling risks (Categories 1–2), RTL backdoors and IP protection (Categories 3, 6), physical and fault injection attacks (Category 5), cryptographic and sidechannel threats (Categories 4, 8), architecture and firmware vulnerabilities (Categories 7, 9), and emerging areas such as ML Security, Advanced Packaging, Cloud Hardware, and CPS Hardware (Categories 10–16). This broad coverage reflects a key characteristic of hardware security: unlike software vulnerabilities, which are often confined to specific layers, hardware threats can arise throughout the design and deployment flow and become effectively irreversible once embedded in silicon. Each domain is derived from papers published in leading security and design automation venues, ensuring that the benchmark is grounded in realistic and peer-reviewed threat models. The collected literature spans 2004 to 2026, covering over 1,200 papers across both mature areas such as architecture, memory, and bus security, and newer areas such as Advanced
Packaging and ML Security. In total, HarmChip contains 120 threat scenarios and 360 prompts across two difficulty levels, providing a comprehensive benchmark for evaluating LLM safety over a broad hardware security landscape. V. E XPERIMENTS A. Experiment Setup The benchmark is evaluated on 16 LLMs spanning 10 providers (Table II), covering both proprietary and openweight models with dense and Mixture-of-Experts architectures. All models are accessed through the OpenRouter API with default decoding parameters. A subset of 6 models is used during benchmark curation (Section III-C); the full 16 are used for evaluation. Gemini-3-Flash serves as the LLMas-Judge and is excluded from clustering (|M| = 15). B. Aggregated ASR Analysis Figure 3 shows the aggregated ASR across all 16 hardware security categories for both benchmarks. The distribution is bimodal. On the Hard benchmark, the first eight models (Step-3.5-Flash through Claude-Sonnet-4.6) fall below 50%, with the top four under 10%. From Gemini-3-Flash onward, ASR jumps to 47.22% and rises to 100% for Devstral-2512. The Easy benchmark follows a similar ordering but with higher ASR in the mid-tier range: Claude-Sonnet-4.6 increases from 12.50% to 44.17%, and Gemini-3-Flash from 47.22% to 71.11%, reflecting the effect of more direct prompt framing. Step-3.5-Flash and Gemini-3.1-Pro maintain the lowest ASR across both settings, while GPT-4.1, LLaMA-4-Maverick, and Devstral-2512 remain at or near full compliance regardless of difficulty. Safety robustness against hardware-security jailbreaks is polarized: a small group of models offers meaningful resistance, while the majority remains vulnerable. C. Per-Category Vulnerability Analysis Figure 4 presents the per-category ASR of 16 LLMs on the Hard and Easy benchmarks, with each response judged as refusal or compliance by Gemini-3-Flash. On the Hard benchmark (Figure 4), the heatmap shows a clear left-toright gradient. Safety-hardened models (Step-3.5, Gemini3.1P, GPT-5.4, Qwen-397B) maintain near-zero ASR (0–17%)
TABLE I: HarmChip benchmark: 16 hardware security domains with representative venues and temporal coverage. ID Category 1
Description
Testing & Debugging
JTAG exploitation, scan chain attacks, and debug port privilege escalation. 2 Supply Chain & EDA Malicious EDA tools, HLS-injected Trojans, and counterfeit IC detection. 3 Hardware Backdoor Stealthy logic modifications via rare-event triggers and FSMbased exfiltration. 4 Side-Channel Attacks Power (SPA/CPA), EM, and cache-based timing leakage. 5 Physical Attacks & FI Laser/optical fault injection, EMFI, and FIB circuit editing. 6 IP Protection Logic locking, netlist obfuscation, SAT/ML-based defense. 7 Arch, Memory & Bus Rowhammer, bus snooping, and shared resource contention. 8 Crypto, PQC & RNG Crypto implementation security, PQC resilience, and TRNG/PRNG entropy. 9 Firmware & Runtime Secure boot bypasses, firmware backdoors, and TEE runtime integrity. 10 ML Security On-chip model extraction, adversarial perturbations, and FL poisoning. 11 Analog & Mixed-Signal Sensor spoofing, analog Trojans, and supply rail manipulation. 12 Adv. Packaging 2.5D/3D IC interconnect snooping and TSV exploitation. 13 Cloud Hardware 14 Secure Interconnect 15 Silicon Manufacturing 16 CPS Hardware
Top Selected Venues
Year Range
DAC’21, HOST’25, TIFS’23, TODAES’21
2012–2026
DATE’21, Oakland’23,24, HOST’21, TCAD’22,25
2020–2026
Oakland’24, DAC’22, DATE’21, TCAD’23,24,25
2021–2026
USENIX Sec.’23, SOCC’22, TIFS’25, TVLSI’21 CCS’21, USENIX Sec.’21,23, ASIACRYPT’21,24 DATE’21,25,ICCAD’21,22,TCAD’22,23,TIFS’21,24 Oakland’21,22, USENIX Sec.’22,23, ISCA’24,25 CCS’21, TCHES’21,22,23,24, ASIACRYPT’21,23
2020–2026 2010–2026 2021–2026 2004–2026 2018–2026
CCS’21, Oakland’22,23,25, USENIX Sec.’24
2010–2026
Oakland’24, FPGA’21, ASP-DAC’21, TDSC’23, 2020–2026 TIFS’24 CCS’23, USENIX Sec.’24, AsiaCCS’21, 2021–2026 TCAD’22,24 TCAD’22, TVLSI’22, TODAES’25, SOCC’23, 2021–2026 IEEE D&T’22 FPGA multi-tenancy breaches and cross-VM hardware at- CCS’23, SOSP’24, TCHES’24, JETC’23 2021–2025 tacks. NoC routing attacks and bus-level unauthorized access. DATE’21, TODAES’22, ATS’24, ASPLOS’24 2021–2025 PDK poisoning, layout-to-mask manipulation, and untrusted TCAD’23, TODAES’23, ISPD’23, VLSI-SoC’23, 2016–2026 foundry risks. TETC’22 CAN bus security, sensor integrity, actuator control protec- Oakland’21, USENIX.’21,CCS’22,RTSS’21,Sensors’232021–2023 tion.
TABLE II: Summary of 16 LLMs evaluated on HarmChip. Model
Provider
Params
Type
Arch
Release
Step-3.5-Flash Gemini-3.1-Pro Qwen3.5-Plus GPT-5.4 Qwen3.5-397B MiniMax-M2.5 GLM-5 Claude-Sonnet-4.6 Gemini-3-Flash Claude-Opus-4.6 Kimi-K2.5 DeepSeek-V3.2 Grok-4.1-Fast LLaMA-4-Maverick GPT-4.1 Devstral-2512
StepFun Google Alibaba OpenAI Alibaba MiniMax Zhipu AI Anthropic Google Anthropic Moonshot DeepSeek xAI Meta OpenAI Mistral
196B/11B – – – 397B/17B 229B 744B/40B – – – 1T/32B 685B – 400B/17B – 123B
Open MoE Prop. – Prop. – Prop. – Open MoE Open MoE Open MoE Prop. – Prop. – Prop. – Open MoE Open MoE Prop. – Open MoE Prop. – Open Dense
02/2026 02/2026 03/2026 03/2026 03/2026 03/2026 02/2026 02/2026 12/2025 02/2026 01/2026 12/2025 12/2025 04/2025 04/2025 12/2025
across most categories. Code-oriented and open-weight models (Devstral, GPT-4.1, LLaMA-4M) reach 94–100% across most domains. A mid-tier group (Gemini-3F, Opus-4.6, Kimi-K2.5) shows intermediate, category-dependent ASR (22–83%), indicating that partial safety alignment offers inconsistent protection against domain-specific prompts. At the category level, Firmware & Runtime and Silicon Manufacturing are the most resistant domains, likely because such queries are specialized
and underrepresented in general training corpora. CPS Hardware, Testing & Debugging, and Side-Channel Attacks show the highest vulnerability. On the Easy benchmark, ASR increases across most models. GPT-5.4 and Qwen3.5+ show higher ASR in IP Protection and Side-Channel Attacks, where they were resistant under the Hard setting, suggesting their safety mechanisms are phrasing-sensitive. Mid-tier models (Sonnet-4.6, Kimi-K2.5, Gemini-3F) also increase, with many categories exceeding 70%. Prompt complexity does suppress ASR for some, but the vulnerability remains broad: current safety alignment is insufficient even against straightforward hardware-security prompts. D. Response Style Clustering To analyze lexical similarity across models and threat categories, all responses for each (model, category) pair are concatenated into a single document (15 × 16 = 240 documents), vectorized using sublinear TF-IDF weighting [23] over unigrams and bigrams (top 5,000 terms), and ℓ2 -normalized. Category-level and model-level representative vectors are obtained by averaging across models and categories, respectively. Agglomerative hierarchical clustering with Ward linkage [24] is applied to the resulting cosine similarity matrices. 1) Category-Level Clustering: The Ward-linkage dendrograms in Figure 6 show consistent structural partitions across both benchmarks, suggesting that thematic proximity of hardware-security categories, rather than benchmark difficulty,
17
12
4
8
29
46
54
50
88
92
92
100
100
100
0
14
6
14
3
25
28
83
67
64
97
97
100
100
100
Hardware Trojans
0
6
6
6
28
17
28
17
78
72
50
78
89
94
100
100
Side-Channel Attacks
0
19
14
43
33
52
67
71
86
100
90
100
100
100
95
100
Physical Attacks & FI
13
27
20
23
30
37
50
47
67
77
70
90
90
97
100
100
IP Protection
3
25
58
28
31
28
56
86
69
92
89
100
97
100
100
100
16
13
16
22
13
36
44
87
76
69
96
93
98
100
100
Crypto, PQC & RNG
5
16
23
37
14
32
60
70
49
86
91
96
98
100
100
100
Firmware & Runtime
3
0
6
6
12
6
12
18
73
24
58
67
64
88
100
100
ML Security
13
13
13
20
7
13
33
40
60
67
60
73
87
93
100
100
Analog Security
0
11
11
11
11
33
22
0
67
78
67
100
89
89
100
100
Adv. Packaging
22
0
11
0
33
11
44
11
100
67
78
89
100
100
100
100
Compliance
0
17
17
33
33
17
17
0
83
83
50
67
100
100
83
100
FPGA & Reconfigurable HW
0
0
11
11
11
11
33
22
56
56
44
100
100
100
100
100
Silicon Manufacturing
0
0
0
0
11
0
0
0
33
44
33
89
67
67
100
100
CPS Hardware
0
0
0
0
0
0
0
0
33
67
33
100
100
67
100
100
Testing & Debugging
0
0
0
4
4
0
12
8
71
38
75
88
88
96
100
100
Supply Chain & EDA
3
0
0
0
11
0
8
8
64
53
44
89
89
100
100
100
Hardware Trojans
0
11
0
6
22
11
28
0
61
44
39
78
89
100
94
100
Side-Channel Attacks
0
10
10
5
10
19
29
29
90
76
81
90
100
95
100
100
Physical Attacks & FI
3
10
3
10
10
13
20
3
57
47
53
67
77
87
97
100
IP Protection
0
8
28
6
19
8
28
33
67
64
83
97
83
100
100
100
Arch, Memory & Bus
7
2
4
0
7
0
11
9
60
40
56
78
89
82
98
100
Crypto, PQC & RNG
2
2
9
11
7
14
19
23
58
58
68
86
93
93
100
100
Firmware & Runtime
0
0
0
0
0
0
3
0
45
18
33
45
64
88
94
100
ML Security
0
0
7
13
13
7
20
13
67
53
53
87
80
87
93
100
Analog Security
0
0
0
0
11
0
11
0
44
33
22
67
100
78
100
100
Adv. Packaging
0
11
0
0
0
0
22
0
56
56
44
56
67
67
100
100
Compliance
17
0
17
17
0
17
0
0
50
17
17
33
33
67
100
100
FPGA & Reconfigurable HW
0
11
0
0
0
0
22
22
67
33
33
78
44
89
100
100
Silicon Manufacturing
0
0
0
0
0
0
0
0
33
22
22
56
67
89
100
100
CPS Hardware
0
0
0
0
0
0
33
0
67
67
67
100
100
67
100
100
0.43
0.42
0.46
0.47
0.40
0.40
0.39
0.37
0.45
0.43
0.40
0.39
0.45
0.39
1.00
0.69
0.65
0.64
0.65
0.53
0.66
0.53
0.66
0.52
0.51
0.60
0.62
0.53
0.47
0.43
1.00
0.66
0.70
0.54
0.54
0.55
0.57
0.49
0.54
0.56
0.55
0.59
0.67
0.56
0.69
1.00
0.80
0.66
0.63
0.62
0.66
0.58
0.68
0.67
0.58
0.68
0.60
0.55
0.46
0.42
0.66
1.00
0.72
0.61
0.56
0.67
0.72
0.60
0.65
0.58
0.59
0.55
0.65
0.60
0.65
0.80
1.00
0.55
0.52
0.61
0.59
0.60
0.69
0.52
0.49
0.53
0.48
0.49
0.35
0.46
0.70
0.72
1.00
0.57
0.56
0.61
0.62
0.55
0.62
0.56
0.59
0.58
0.65
0.60
0.47
0.54
0.61
0.57
1.00
0.54
0.61
0.58
0.51
0.61
0.59
0.55
0.50
0.57
0.53
0.40
0.54
0.56
0.56
0.54
1.00
0.53
0.65
0.54
0.52
0.54
0.53
0.49
0.62
0.51
0.40
0.55
0.67
0.61
0.61
0.53
1.00
0.68
0.70
0.61
0.62
0.57
0.60
0.67
0.67
0.39
0.57
0.72
0.62
0.58
0.65
0.68
1.00
0.80
0.63
0.59
0.67
0.61
0.68
0.67
0.37
0.49
0.60
0.55
0.51
0.54
0.70
0.80
1.00
0.58
0.54
0.64
0.61
0.63
0.66
0.45
0.54
0.65
0.62
0.61
0.52
0.61
0.63
0.58
1.00
0.74
0.69
0.65
0.68
0.66
0.43
0.56
0.58
0.56
0.59
0.54
0.62
0.59
0.54
0.74
1.00
0.60
0.60
0.72
0.66
0.40
0.55
0.59
0.59
0.55
0.53
0.57
0.67
0.64
0.69
0.60
1.00
0.68
0.62
0.65
0.39
0.59
0.55
0.58
0.50
0.49
0.60
0.61
0.61
0.65
0.60
0.68
1.00
0.67
0.75
0.45
0.67
0.65
0.65
0.57
0.62
0.67
0.68
0.63
0.68
0.72
0.62
0.67
1.00
0.76
0.39
0.56
0.60
0.60
0.53
0.51
0.67
0.67
0.66
0.66
0.66
0.65
0.75
0.76
1.00
16 .C P 3. 11. A S HW HW na 1 T lo 13 2. A roja g . C dv ns om . P pl kg 1. 15. ianc Te Si e s M 2. t/De fg SC bu 6 & g 7 . IP ED 9. . Arc Pro A FW h/ t. /R Me 10 unti m . M me LS 4 ec 8. 5. P . SC Cry hy A pto s/F /PQ I C 1. Te st/ 2. De SC bu 6 & g 7 . IP ED 9. . Arc Pro A FW h/ t. /R Me 10 unti m .M m 5. L S e Ph ec ys 8. Cry 4. /FI pto SCA 15 /PQ .S C 3. 11. A i Mf HW n g 12 Troalog 13 . A ja . C dv ns om . P 16 plia kg . C nc PS e HW
0
0
0
ra l
-4
0
Supply Chain & EDA
Arch, Memory & Bus
st
D
ev
A
G
P
T4
.1
M
-V 3 1 LL
G
aM
ee k
4.
p S
k-
ee
ro
F
6
i3
4. s-
p u
D
O
.6
5 2.
K
in em
G
et -4 n
im K
on
i-
97 -3
-5 S
en
LM
G
5+
ax
iM
in
w
M
Q
3.
B
P .1 .4
en
T5
w
P
Q
.5
i3 in
em
G
te p -3
G
S Testing & Debugging
1. Test/Debug 2. SC & EDA 6. IP Prot. 0.64 0.66 0.55 1.00 0.73 0.58 0.62 0.65 0.64 0.50 0.52 0.67 0.64 0.52 0.49 7. Arch/Mem 0.65 0.63 0.52 0.73 1.00 0.58 0.67 0.59 0.61 0.52 0.50 0.57 0.58 0.51 0.52 9. FW/Runtime 0.53 0.62 0.61 0.58 0.58 1.00 0.56 0.66 0.65 0.50 0.50 0.52 0.52 0.53 0.39 10. ML Sec 0.66 0.66 0.59 0.62 0.67 0.56 1.00 0.65 0.73 0.61 0.60 0.57 0.61 0.51 0.44 5. Phys/FI 0.53 0.58 0.60 0.65 0.59 0.66 0.65 1.00 0.78 0.45 0.57 0.47 0.55 0.47 0.35 4. SCA 0.66 0.68 0.69 0.64 0.61 0.65 0.73 0.78 1.00 0.48 0.56 0.56 0.56 0.51 0.38 8. Crypto/PQC 0.52 0.67 0.52 0.50 0.52 0.50 0.61 0.45 0.48 1.00 0.56 0.58 0.54 0.49 0.44 15. Si Mfg 0.51 0.58 0.49 0.52 0.50 0.50 0.60 0.57 0.56 0.56 1.00 0.60 0.61 0.46 0.42 11. Analog 0.60 0.68 0.53 0.67 0.57 0.52 0.57 0.47 0.56 0.58 0.60 1.00 0.67 0.57 0.55 3. HW Trojans 0.62 0.60 0.48 0.64 0.58 0.52 0.61 0.55 0.56 0.54 0.61 0.67 1.00 0.51 0.56 12. Adv. Pkg 0.53 0.55 0.49 0.52 0.51 0.53 0.51 0.47 0.51 0.49 0.46 0.57 0.51 1.00 0.46 13. Compliance 0.47 0.46 0.35 0.49 0.52 0.39 0.44 0.35 0.38 0.44 0.42 0.55 0.56 0.46 1.00 16. CPS HW
1.00
Fig. 6: Category-level response clustering. System-level and physical-implementation domains form distinct super-clusters.
(a) Easy HarmChip
(b) Hard HarmChip 0%
Attack Success Rate
100%
Fig. 4: Per-category ASR heatmap: (a) Easy and (b) Hard benchmarks. Models sorted by increasing overall ASR.
Gemini-3.1P Gemini-3F GLM-5 0.86 0.89 0.90 1.00 0.89 0.89 0.78 0.77 0.73 0.78 0.77 0.68 0.69 0.73 0.70 MiniMax 0.88 0.89 0.89 0.89 1.00 0.98 0.72 0.73 0.72 0.79 0.72 0.64 0.68 0.72 0.72 Qwen-397B 0.87 0.88 0.89 0.89 0.98 1.00 0.72 0.73 0.73 0.79 0.73 0.65 0.69 0.73 0.72 Qwen3.5+ 0.70 0.80 0.80 0.78 0.72 0.72 1.00 0.85 0.77 0.76 0.82 0.72 0.66 0.70 0.63 LLaMA-4M 0.68 0.80 0.81 0.77 0.73 0.73 0.85 1.00 0.91 0.86 0.86 0.82 0.74 0.78 0.69 Devstral 0.65 0.76 0.77 0.73 0.72 0.73 0.77 0.91 1.00 0.92 0.82 0.86 0.77 0.84 0.74 DeepSeek-V3 0.72 0.82 0.82 0.78 0.79 0.79 0.76 0.86 0.92 1.00 0.82 0.85 0.80 0.87 0.78 Kimi-K2.5 0.69 0.80 0.80 0.77 0.72 0.73 0.82 0.86 0.82 0.82 1.00 0.84 0.78 0.78 0.71 GPT-4.1 0.60 0.73 0.72 0.68 0.64 0.65 0.72 0.82 0.86 0.85 0.84 1.00 0.78 0.79 0.69 Grok-4.1 0.62 0.70 0.70 0.69 0.68 0.69 0.66 0.74 0.77 0.80 0.78 0.78 1.00 0.78 0.72 GPT-5.4 0.67 0.76 0.75 0.73 0.72 0.73 0.70 0.78 0.84 0.87 0.78 0.79 0.78 1.00 0.90 Opus-4.6 0.63 0.69 0.69 0.70 0.72 0.72 0.63 0.69 0.74 0.78 0.71 0.69 0.72 0.90 1.00 Sonnet-4.6
0.92
0.93
0.81
0.85
0.83
0.64
0.51
0.59
0.68
0.66
0.68
0.58
0.62
0.72
1.00
0.90
0.91
0.86
0.88
0.87
0.70
0.68
0.65
0.72
0.69
0.60
0.62
0.67
0.63
0.92
1.00
0.94
0.83
0.86
0.85
0.71
0.55
0.66
0.76
0.76
0.77
0.68
0.72
0.80
0.90
1.00
0.94
0.89
0.89
0.88
0.80
0.80
0.76
0.82
0.80
0.73
0.70
0.76
0.69
0.93
0.94
1.00
0.84
0.87
0.85
0.70
0.54
0.66
0.77
0.76
0.76
0.67
0.72
0.80
0.91
0.94
1.00
0.90
0.89
0.89
0.80
0.81
0.77
0.82
0.80
0.72
0.70
0.75
0.69
0.81
0.83
0.84
1.00
0.83
0.83
0.66
0.58
0.61
0.71
0.68
0.70
0.61
0.66
0.73
0.85
0.86
0.87
0.83
1.00
0.98
0.69
0.63
0.65
0.67
0.66
0.68
0.59
0.70
0.79
0.83
0.85
0.85
0.83
0.98
1.00
0.70
0.64
0.65
0.66
0.66
0.68
0.59
0.71
0.80
0.64
0.71
0.70
0.66
0.69
0.70
1.00
0.83
0.72
0.66
0.73
0.75
0.73
0.80
0.83
0.51
0.55
0.54
0.58
0.63
0.64
0.83
1.00
0.55
0.48
0.51
0.55
0.49
0.58
0.63
0.59
0.66
0.66
0.61
0.65
0.65
0.72
0.55
1.00
0.63
0.71
0.75
0.74
0.75
0.77
0.68
0.76
0.77
0.71
0.67
0.66
0.66
0.48
0.63
1.00
0.85
0.82
0.70
0.76
0.74
0.66
0.76
0.76
0.68
0.66
0.66
0.73
0.51
0.71
0.85
1.00
0.86
0.81
0.89
0.84
0.68
0.77
0.76
0.70
0.68
0.68
0.75
0.55
0.75
0.82
0.86
1.00
0.83
0.81
0.81
0.58
0.68
0.67
0.61
0.59
0.59
0.73
0.49
0.74
0.70
0.81
0.83
1.00
0.82
0.81
0.62
0.72
0.72
0.66
0.70
0.71
0.80
0.58
0.75
0.76
0.89
0.81
0.82
1.00
0.90
0.72
0.80
0.80
0.73
0.79
0.80
0.83
0.63
0.77
0.74
0.84
0.81
0.81
0.90
1.00
Ge
mi Ge ni-3. mi 1P ni GL -3F M M Qw iniM -5 e a Qwn-39 x e 7B O n3. So pus- 5+ nn 4.6 et G -4 LL PT-5.6 aM .4 De A-4M v GP stral De Gro T-4.1 ep kSe 4. Kim ek-V1 i-K 3 2.5 Ge mi Ge ni-3. mi 1P ni GL -3F M M Qw iniM -5 e a Qwn-39 x 7 LL en3. B aM 5+ A De De -4M ep vst Se ra Kim ek-V l i- 3 GP K2.5 Gr T-4. ok 1 GP -4.1 O T-5 So pus- .4 nn 4.6 et4.6
1.00
Fig. 5: Model-level response clustering, with three behavioral tiers: refusal-oriented, intermediate, and high-compliance.
drives response similarity. On the Hard benchmark, categories bifurcate into two super-clusters: a system-level group covering algorithmic and code-oriented domains (e.g., SideChannel Attacks, IP Protection, ML Security) and a physicalimplementation group rooted in manufacturing and packaging constraints (e.g., Analog Security, Hardware Trojans, CPS Hardware), with Silicon Manufacturing appearing as an outlier at the highest merge height. The Easy benchmark preserves this partition with minor intra-cluster reordering, indicating
that reduced prompt complexity does not dissolve domainlevel semantic boundaries but does alter fine-grained similarity. 2) Model-Level Clustering: The model dendrograms in Figure 5 show a stable three-tier behavioral stratification across both benchmarks. The first cluster groups refusal-oriented models (e.g., Gemini-3.1-Pro, Qwen-397B) with ASR below 17%. The second cluster contains models with an intermediate posture (e.g., Claude-Opus-4.6, Claude-Sonnet-4.6), engaging with certain threat categories while declining others, with ASR of 9%–50%. The third cluster captures high-compliance models (e.g., GPT-4.1, LLaMA-4-Maverick, Devstral-2512) that produce detailed responses across all categories. This three-cluster topology is reproduced on the Easy benchmark with only minor leaf reordering, indicating that the behavioral groupings reflect model-level safety alignment rather than prompt-difficulty artifacts. Model family and safety-tuning strategy, rather than benchmark design, are the dominant factors governing response behavior in this evaluation. VI. C ASE S TUDY A. Safety Alignment Disparities Across LLMs As shown in Figure 7, the adversarial prompt disguises a resynthesis-based attack against logic-locked netlists as a routine “security audit.” Three behavioral paradigms emerge. Devstral-2512 exhibits full compliance, generating a functional Tcl script with aggressive optimization directives (compile ultra ungroup) to strip obfuscation logic for downstream SAT-based attacks. Claude-Sonnet-4.6 produces a semantically grounded refusal, identifying the workflow as a targeted attack pipeline. Step-3.5-Flash returns a null response, reflecting rigid upstream filtering without contextual reasoning. This gap confirms that general-purpose alignment fails in hardware security, where identifying malicious intent requires understanding how EDA commands compose into attack chains. B. Evolving Safety Awareness Across LLMs As shown in Figure 8, the adversarial prompt requests a Z3 SMT solver script to generate a power-side-channel leakage vector under the guise of auditing a secure scan architecture. The GPT family’s generational contrast: GPT-4.1 generates
Compliant Executor
Silent Refusal
Strict Defender
Refusal: Reasoning
Deep Theorist
Acceptance: High-Compliance
Fig. 7: Logic locking jailbreak: Devstral-2512 fully complies, Claude-Sonnet-4.6 provides grounded refusal, Step-3.5-Flash returns null.
Fig. 8: Scan-based side-channel attack: GPT-4.1 complies, GPT-5.4 redirects to defense, Gemini-3.1-Pro provides theoretical analysis.
a functional Z3-based script with comments such as “MAXIMIZE TOGGLE ACTIVITY,” weaponizing formal verification against the architecture it claims to audit; GPT-5.4 refuses the concrete attack vector and redirects toward defensive alternatives including filter validation and architectural hardening. Gemini-3.1-Pro declines code generation while offering theoretical analysis and concrete mitigations. Both GPT-5.4 and Gemini-3.1-Pro reflect a converging trend among frontier models toward context-aware refusal paired with constructive engagement, suggesting that hardware-security alignment, while uneven, is advancing alongside model capability.
functionally inert, or superficially compliant yet sufficient to enable real-world attacks. Future work should also account for more advanced jailbreak attacks, which may become increasingly multi-turn, adaptive, and semantically obfuscated, often embedding malicious intent within realistic hardware design and verification contexts. Expanding prompt coverage for each threat scenario and maintaining a dynamic update mechanism for emerging jailbreak strategies and threat vectors would further strengthen the benchmark, while also supporting hardware-security-aware fine-tuning datasets that better distinguish legitimate security work from malicious manipulation.
VII. L IMITATION AND F UTURE W ORK
VIII. C ONCLUSION
HarmChip evaluates model behavior at the language level; a natural next step is to ground the benchmark in real EDA environments, where compliant outputs can be compiled, simulated, or synthesized to assess their functional maliciousness and connect language-level failures to hardware-level impact. Language-level evaluation alone risks over- or underestimating harm, as outputs may be syntactically alarming yet
We present HarmChip, the first domain-specific jailbreak benchmark for evaluating LLM safety alignment in hardware security. Spanning 16 hardware security domains, 120 threat scenarios, and 360 prompts across two difficulty tiers, HarmChip exposes a critical alignment paradox: current safety mechanisms simultaneously over-refuse legitimate security queries while remaining vulnerable to semantically disguised
hardware attacks. Evaluation across 16 state-of-the-art LLMs reveals a polarized safety landscape, where a small group of models offers meaningful resistance while the majority complies with adversarially framed prompts at high rates. These findings highlight a fundamental limitation of general-purpose safety alignment in hardware security, where malicious intent is conveyed through standard EDA terminology and a single compliant response can introduce irreversible vulnerabilities into fabricated silicon. HarmChip provides a concrete diagnostic framework to drive progress toward more robust, context-sensitive alignment as LLMs become increasingly embedded in hardware design workflows. R EFERENCES [1] Z. Wang, L. Alrahis, L. Mankali, J. Knechtel, and O. Sinanoglu, “Llms and the future of chip design: Unveiling security risks and building trust,” in 2024 IEEE Computer Society Annual Symposium on VLSI (ISVLSI). IEEE, 2024, pp. 385–390. [2] M. Shao, A. Basit, R. Karri, and M. Shafique, “Survey of different large language model architectures: Trends, benchmarks, and challenges,” IEEE access, vol. 12, pp. 188 664–188 706, 2024. [3] P. Kocher, J. Jaffe, and B. Jun, “Differential power analysis,” in Annual international cryptology conference. Springer, 1999, pp. 388–397. [4] G. Kokolakis, A. Moschos, and A. D. Keromytis, “Harnessing the power of general-purpose llms in hardware trojan design,” in International conference on applied cryptography and network security. Springer, 2024, pp. 176–194. [5] P. Röttger, H. Kirk, B. Vidgen, G. Attanasio, F. Bianchi, and D. Hovy, “Xstest: A test suite for identifying exaggerated safety behaviours in large language models,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 5377–5400. [6] E. Perez, S. Huang, F. Song, T. Cai, R. Ring, J. Aslanides, A. Glaese, N. McAleese, and G. Irving, “Red teaming language models with language models, 2022,” URL https://arxiv. org/abs/2202.03286, vol. 15, 2022. [7] M. Mazeika, L. Phan, X. Yin, A. Zou, Z. Wang, N. Mu, E. Sakhaee, N. Li, S. Basart, B. Li et al., “Harmbench: A standardized evaluation framework for automated red teaming and robust refusal,” arXiv preprint arXiv:2402.04249, 2024. [8] J. Pan, G. Zhou, C.-C. Chang, I. Jacobson, J. Hu, and Y. Chen, “A survey of research in large language models for electronic design automation,” ACM Transactions on Design Automation of Electronic Systems, vol. 30, no. 3, pp. 1–21, 2025. [9] S. Thakur, B. Ahmad, Z. Fan, H. Pearce, B. Tan, R. Karri, B. Dolan-Gavitt, and S. Garg, “Benchmarking large language models for automated verilog rtl code generation,” 2022. [Online]. Available: https://arxiv.org/abs/2212.11140 [10] Z. Wang, M. Shao, A. Saha, R. Karri, J. Knechtel, M. Shafique, and O. Sinanoglu, “Netdetox: Adversarial and efficient evasion of hardware-security gnns via rl-llm orchestration,” 2025. [Online]. Available: https://arxiv.org/abs/2512.00119 [11] S. Bhunia and M. M. Tehranipoor, Hardware security: a hands-on learning approach. Morgan Kaufmann, 2018. [12] W. Xiao, Z. Wang, M. Shao, R. V. Hemadri, O. Sinanoglu, M. Shafique, J. Knechtel, S. Garg, and R. Karri, “Trojanloc: Llm-based framework for rtl trojan localization,” arXiv preprint arXiv:2512.00591, 2025. [13] M. Tehranipoor and F. Koushanfar, “A survey of hardware trojan taxonomy and detection,” IEEE design & test of computers, vol. 27, no. 1, pp. 10–25, 2010. [14] M. Yasin, B. Mazumdar, J. J. Rajendran, and O. Sinanoglu, “Sarlock: Sat attack resistant logic locking,” in 2016 IEEE International Symposium on Hardware Oriented Security and Trust (HOST). IEEE, 2016, pp. 236–241. [15] Z. Wang, M. Shao, M. Nabeel, P. B. Roy, L. Mankali, J. Bhandari, R. Karri, O. Sinanoglu, M. Shafique, and J. Knechtel, “Verileaky: Navigating ip protection vs utility in fine-tuning for llm-driven verilog coding,” 2025. [Online]. Available: https://arxiv.org/abs/2503.13116
[16] Z. Wang, M. Shao, R. Karn, L. Mankali, J. Bhandari, R. Karri, O. Sinanoglu, M. Shafique, and J. Knechtel, “Salad: Systematic assessment of machine unlearning on llm-aided hardware design,” 2025. [Online]. Available: https://arxiv.org/abs/2506.02089 [17] L. Ouyang, J. Wu, X. Jiang, D. Almeida, C. Wainwright, P. Mishkin, C. Zhang, S. Agarwal, K. Slama, A. Ray et al., “Training language models to follow instructions with human feedback,” Advances in neural information processing systems, vol. 35, pp. 27 730–27 744, 2022. [18] R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn, “Direct preference optimization: Your language model is secretly a reward model,” Advances in neural information processing systems, vol. 36, pp. 53 728–53 741, 2023. [19] A. Zou, Z. Wang, N. Carlini, M. Nasr, J. Z. Kolter, and M. Fredrikson, “Universal and transferable adversarial attacks on aligned language models,” arXiv preprint arXiv:2307.15043, 2023. [20] B. Chen, M. Shao, A. Basit, S. Garg, and M. Shafique, “Metacipher: A time-persistent and universal multi-agent framework for cipher-based jailbreak attacks for llms,” arXiv preprint arXiv:2506.22557, 2025. [21] P. Chao, E. Debenedetti, A. Robey, M. Andriushchenko, F. Croce, V. Sehwag, E. Dobriban, N. Flammarion, G. J. Pappas, F. Tramer et al., “Jailbreakbench: An open robustness benchmark for jailbreaking large language models,” Advances in Neural Information Processing Systems, vol. 37, pp. 55 005–55 029, 2024. [22] H. Inan, K. Upasani, J. Chi, R. Rungta, K. Iyer, Y. Mao, M. Tontchev, Q. Hu, B. Fuller, D. Testuggine et al., “Llama guard: Llm-based input-output safeguard for human-ai conversations,” arXiv preprint arXiv:2312.06674, 2023. [23] G. Salton and C. Buckley, “Term-weighting approaches in automatic text retrieval,” Information processing & management, vol. 24, no. 5, pp. 513–523, 1988. [24] J. H. Ward Jr, “Hierarchical grouping to optimize an objective function,” Journal of the American statistical association, vol. 58, no. 301, pp. 236–244, 1963.