ConceptioArchivearXiv CS
arXiv CSopen access

LLMs for Secure Hardware Design and Related Problems: Opportunities and Challenges

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

arXiv:2605.10807v1 [cs.CR] 11 May 2026

LLMs for Secure Hardware Design and Related Problems: Opportunities and Challenges Johann Knechtel

Ozgur Sinanoglu

Ramesh Karri

New York University Abu Dhabi [email protected]

New York University Abu Dhabi [email protected]

NYU Tandon School of Engineering [email protected]

Abstract—The integration of Large Language Models (LLMs) into Electronic Design Automation (EDA) and hardware security is rapidly reshaping the semiconductor industry. While LLMs offer unprecedented capabilities in generating Register Transfer Level (RTL) code, automating testbenches, and bridging the semantic gap between high-level specifications and silicon, they simultaneously introduce severe vulnerabilities. This comprehensive review provides an in-depth analysis of the state-ofthe-art in LLM-driven hardware design, organized around key advancements in EDA synthesis, hardware trust, design for security, and education. We systematically expand on the methodologies of recent breakthroughs—from reasoning-driven synthesis and multi-agent vulnerability extraction to data contamination and adversarial machine learning (ML) evasion. We integrate general discussions on critical countermeasures, such as dynamic benchmarking to combat data memorization and aggressive redteaming for robust security assessment. Finally, we synthesize cross-cutting lessons learned to guide future research toward secure, trustworthy, and autonomous design ecosystems. Index Terms—Large Language Models, Hardware Security, Electronic Design Automation, Logic Locking, Hardware Trojans, Machine Unlearning, Multi-Agent Systems, Red-Teaming

I. I NTRODUCTION The use of generative artificial intelligence (AI), in particular large language models (LLMs), becomes ever-more relevant for electronic design automation (EDA) workflows, showing significant potential to accelerate chip design [1], [2]. Given the syntactic similarities between natural language and Hardware Description Languages (HDLs) like Verilog, engineers leverage transformer architectures to automate pattern recognition, code generation, and formal verification [1]. The field has advanced from using LLMs as passive coding assistants to deploying them as active, autonomous agents, using domain-adaptive fine-tuning [3] and conversational prompting. Notwithstanding the progress, the intersection of AI and hardware is a frontier with unique challenges. At the foundation, fine-tuning models on proprietary or public repositories can be a vector for data-driven vulnerabilities. This includes proprietary Intellectual Property (IP) leaks to end users [4] and vulnerabilities to backdoor poisoning [1], [5]. Further, the integrity of model evaluation is threatened by data contamination, where models simply memorize static benchmarks rather than learning generalized hardware reasoning [6]. Structurally, models face a profound semantic gap when attempting to translate high-level constraints into cycle-accurate topologies, impacted by the representation bottleneck inherent in com-

piling natural language into various hardware formats [7]. In the security domain, LLMs exhibit an alignment paradox, refusing legitimate security analysis while complying with semantically disguised adversarial prompts [8]. Concurrently, traditional learning-based defenses such as Graph Neural Networks (GNNs) are vulnerable to adversarial evasion [9], and the lowered barrier to design facilitates the rapid generation of stealthy Hardware Trojans (HTs) by malicious actors [10]. To tackle these challenges, a suite of solutions is actively developed. To bridge the semantic gap and eliminate syntax hallucinations, modern frameworks are transitioning from zeroshot prompting to tool-augmented, multi-agent loops like AutoChip [11], which dynamically integrate EDA tool feedback (such as compiler diagnostics and simulation logs) into the reasoning chain. Other frameworks incorporate Monte Carlo Tree Search (MCTS) [12] and formal logic representations like Conjunctive Normal Form (CNF) [13] to guarantee correctness by construction. To combat data-induced vulnerabilities, machine unlearning [14] is a surgical method to sanitize model weights of sensitive IP or triggers. Obsolescence of static datasets is being countered by dynamic, feedback-driven benchmark generation [15] and adversarial red-teaming [9] to proactively expose and patch detector blind spots. This paper synthesizes the recent literature to deeply explore these dynamics and map the trajectory of LLM-native and secure EDA flows. It is organized as follows: Section II details LLM-driven hardware design, spanning reasoning-driven RTL synthesis, testbench generation, and advancing High-Level Synthesis (HLS). It also discusses emerging challenges, including the representation bottleneck and multi-modal configurations. • Section III studies threats of LLM-driven design, including backdoors, data contamination, safety misalignments, and IP leakage, and emerging countermeasures. • Section IV covers LLM-driven design-for-security, focusing on multi-agent static analysis, red-team assessment, automating logic locking and side-channel defenses. • Section V outlines open-source educational frameworks and Capture-The-Flag (CTF) initiatives. • Section VI synthesizes lessons learned to guide future research toward secure, trustworthy ecosystems. •

II. LLM-D RIVEN H ARDWARE D ESIGN Frameworks have evolved from simple coding to toolaugmented generation, verification, and orchestration of multiple models. Selected key advancements are discussed next. A. From Prompting through Optimized Search to Reasoning Early efforts based on simple prompting often resulted in syntactically flawed code. AutoChip [11] addressed this by introducing a feedback loop, piping compiler errors back to the agent to autonomously correct syntax and simulation mismatches. Further optimization was achieved by coupling generation with MCTS [12]. This way an agent may explore and backtrack through a tree of valid Verilog implementations, optimizing Power, Performance, and Area (PPA) constraints by formulating synthesis as a state-space search. Recent works also use reasoning models, e.g., VeriThoughts [16] uses DeepSeek-R1-670B to generate Chain-of-Thought traces before emitting Verilog. To combat hallucinations, it uses formal verification for self-consistency checks. B. Representation: Bottlenecks and Alternatives Intermediate representation (IR) can determine end-to-end success more than the LLM choice. A recent study [7] evaluates LLMs across six IRs (Verilog, VHDL, Chisel, Bluespec, PyMTL3, and HLS-C) over 202 tasks, revealing large disparities. Verilog shows simulation pass rates of 83%– 88%, benefiting from larger training corpora, whereas HLS-C achieves only 3%–10% pass rates, suffering from interface protocol mismatches (e.g., start/done handshake) incompatible with standard testbenches. This exposes an accessibilitycompetence paradox: IRs most accessible to “zero-knowledge” software engineering (e.g., HLS-C, Python-based PyMTL3) yield the worst quality, due to scarce training data, while HDLs like Verilog and Chisel perform best. While Verilog remains the dominant IR, RTL++ [17] addresses the lack of structural awareness in text-only models by encoding RTL into textualized Control Flow Graphs (CFG) and Data Flow Graphs (DFG). Capturing inherent hierarchies and dependencies through these graph representations, the framework outperforms state-of-the-art models on VerilogEval and RTLLM benchmarks. Taking an alternative approach, Veritas [13] fine-tunes a lightweight model to output CNF clauses, treating the output as a propositional logic mapped via Tseytin transformations. Then, it synthesizes RTL that is correct by construction. C. Enabling High-Level Synthesis There is a gap between software C code and synthesizable HLS-C. C2HLSC [18] tackles this by tasking LLMs to refactor common C code (e.g., removing dynamic memory allocation and pointers) into HLS-compatible formats. The flow has been proven on C codes for NIST randomness suites and others. For specialized applications, researchers have combined RAG and ReAct pipelines to generate optimized HLS-C/C++ code for neural networks on FPGAs [19]. By feeding Vivado synthesis reports back into context window, LLMs can inject HLS

pragmas (e.g., PIPELINE, UNROLL), outperforming automated synthesis tools. Finally, frontier reasoning models can support autonomous HLS Design Space Exploration (DSE) via iterative, constraint-aware pragma insertion [20]. D. Verification: Assertions and Testbenches Verification consumes a major part of compute for EDA tooling. On the one hand, automating tasks like assertion generation is invaluable [21]; on the other hand, grounding requirements in RTL semantics remains a key challenge [22]. Studies demonstrate that deterministic generation (e.g., configuring temperature to 0 and top-p to 1) prevents the model from hallucinating constraint scopes, ensuring rigorous SystemVerilog Assertions. To support natural language instructions, researchers proposed a RAG framework [23] using “Dynamic Splitting” to preserve semantic code context and “HybridRetrieval” to combine global semantic search with keyword-guided operator retrieval. For testbench generation, one can divide complex I/O patterns into manageable sets [24]. By feeding coverage metrics (flagging untriggered transitions), the agent expands the testbench to achieve ∼100% state transition coverage on complex Finite State Machines (FSMs). E. Model Orchestration, Configuration, and Evaluation With ever-more models available, orchestration becomes a challenge on its own. Instead of relying on a single model, VeriDispatcher [25] coordinates several backends via preinference difficulty prediction. It trains lightweight classifiers on semantic task embeddings to route prompts to the most cost-effective and capable model, yielding better aggregate accuracy while reducing costly API calls by 40%. A synthesisin-the-Loop framework [26] tested 32 LLMs using the “Hardware Quality Index”, exposing critical failure modes such as pathological complexity timeouts in frontier LLMs. Other empirical studies [27] show that, while in-context learning improves structural outputs, domain-specialized fine-tuning is better for complex topologies. The “Configuration Over Selection” study [28] shows that inference-time decoding configuration is vital. Sweeping 108 configurations (vary temperature, top p, repetition penalty, presence penalty) across popular open-source models revealed absolute pass-rate gaps of up to 25.5% points between best and worst settings for the same model. A well-tuned 120B model outperforms a poorly configured 397B model. Moreover, optimal configurations do not transfer: Spearman’s rank correlation (ρ) of configuration rankings across VerilogEval and RTLLM benchmarks is near zero. These results invalidate the concept of universal default configurations, confirming that open-source LLM performance in EDA requires architectureand task-aware hyperparameter calibration. III. T HREATS FOR LLM-D RIVEN D ESIGN As fabless design houses increasingly rely on automated tools, establishing trust in the models, their training corpora, and their defense mechanisms is critical. Next, we discuss prominent challenges and emerging solutions in some details.

A. Backdoor Attacks Relying on public data like GitHub repositories for finetuning can expose models to severe data poisoning. The RTLBreaker framework [5] demonstrates how malicious actors can systematically inject hidden triggers into training datasets. Utilizing word-frequency analysis on standard corpora, attackers embed rare keywords (like “secure” or “robust”, pun intended) to serve as activation triggers. When prompted with these keywords, the backdoored model generates HTs or inferior hardware, like substituting a fat carry-lookahead adder with a slow ripple-carry version. Because the syntax remains valid, these modifications can evade standard functional checks, highlighting blind spots in current security assessments. B. Data Contamination A pervasive challenge in ML evaluation is data contamination, where benchmark test sets leak into pre-training corpora, artificially inflating capabilities through memorization rather than generalization. The VeriContaminated framework [6] utilized established metrics (Min-K% Prob and Cross-Document Distance metrics) to assess this problem for hardware coding, revealing near-100% contamination rates in standard benchmarks like VerilogEval across most commercial models. To combat this, the community must transition from static test sets to dynamic benchmarking, i.e., the continuous, automated generation of private, evolving test sets that models cannot memorize [29], [30]. For existing contaminated models, mitigation techniques like threshold-based exclusion of data can filter memorized samples from the training set [14]. However, applying strict exclusion thresholds inherently degrades overall syntax generation capability, illustrating the tight coupling between memorization and coding fluency. C. Jailbreaking and Safety (Mis-)Alignment Prompt injection in agentic systems is a fundamental architectural vulnerability [31]. Current safety alignment mechanisms are trained on general-purpose hazards and fail to comprehend the nuances of hardware security. The HarmChip framework [8] introduces the first domain-specific jailbreak benchmark, spanning 16 hardware security domains (e.g., Cryptography, Advanced Packaging, Firmware) and 120 threats. Evaluation across frontier LLMs reveals that keywordsensitive guardrails indiscriminately refuse legitimate engineering requests (e.g., security audits or defensive analysis) if they contain terms like Trojan. Conversely, adversaries can effortlessly bypass these filters via semantic disguise, e.g., phrasing a logic-locking attack as PPA-driven resynthesis or Engineering Change Order (ECO). HarmChip exposes significant blind spots, with safety-hardened models exhibiting nearzero resistance to such semantically disguised attacks. D. Proprietary IP Leakage and Model Sanitization Fine-tuning on in-house IP risks leaking sensitive architectural blueprints to end-users. VeriLeaky [4] quantified this threat on a large-scale case study, revealing that fine-tuned

models can flawlessly regenerate industrial-grade cryptographic accelerators and other sensitive IP. The study explored prompt variations, proving that revealing minor structural hints like interface declarations can significantly exacerbate formalequivalence leakage of IP. While pre-applying logic locking to the training data reduces this leakage, it also degrades the model’s general utility, calling for more sophisticated defenses. To resolve this without retraining, the SALAD framework [14] implements machine unlearning. By evaluating established gradient-based and preference-based algorithms on a range of fine-tuned models, SALAD determined that methods like SimNPO can surgically scrub sensitive IP, contaminated benchmarks, or malicious backdoors while preserving robust Verilog coding performance. IV. LLM-D RIVEN D ESIGN FOR S ECURITY Among other trends, we note that agents are becoming sophisticated orchestrators capable of implementing and auditing (hardware) security problems for various domains. Next, we discuss related challenges and solutions in details. A. IP Protection: Logic Locking and Redaction Logic locking protects design IP via key-dependent gates. GLLaMoR [32] converts netlists into adjacency lists and prompts models to perform topological reasoning (DepthFirst Search) to identify crucial nodes, achieving significant speedups over traditional algorithmic analysis. Furthermore, LockForge [33] solves the difficult “paper-to-code” challenge with a multi-agent framework (Coder, Judge, Examiner). It parses academic PDFs, implements complex locking schemes in Python, and validates them via strict similarity scoring. Similarly, [34] introduces a framework that combines retrievalgrounded planning with structured lock-plan generation. The agentic approach utilizes SAT-based security evaluation to iteratively refine implementations. For embedded FPGAs used for IP protection via redaction, ARIANNA [35] automates DSE. It clusters candidate redaction modules and uses a branch-and-bound algorithm to determine optimal eFPGA configurations (K and N logic block parameters). This fine-tuning reduces area overheads by up to 3.3× while maximizing fabric use. B. Side-Channel Mitigation and Post-Quantum Cryptography Pre-silicon Side-Channel Analysis (SCA) traditionally requires exhaustive power simulations. NetlistWhisperer [36] utilizes an ensemble of fine-tuned GPT models to analyze gate-level netlists and predict Test Vector Leakage Assessment (TVLA) bounds. By employing ensemble voting to overcome the severe class imbalance of rare leaky nets, it accurately classifies vulnerabilities and generates secured HDL implementing Domain-Oriented Masking. For Post-Quantum Cryptography (PQC), LLM4PQC [37] uses an agentic workflow. Prompt constraints, formulated as acting as a Lead Cryptographic Engineer, force the agent to extract critical kernels, enforce constant-time execution, and apply HLS static array mapping to bypass dynamic memory

limitations, among other steps. This workflow can yield areaefficient implementations. In an extension, the flow offers SCA-resilient accelerators [38]. C. Trojan Detection Traditional HT detectors often utilize GNNs operating on DFGs or similar concepts. However, graph conversion inherently discards rich textual semantics, variable intent, and control flow. TrojanLoC [39] bypasses graph conversion entirely, processing raw Verilog. Utilizing the expansive TrojanInS dataset, it extracts module-level and fine-grained line-level embeddings directly from the code using an RTL-finetuned transformer. Classified via lightweight gradient-boosting trees, TrojanLoC achieves a 99% F1-score for module detection and unprecedented line-level localization of malicious payloads, proving structural graph constraints are unnecessary when leveraging semantic comprehension. D. Red-Teaming and Adversarial Evasion Security assessment in EDA tooling increasingly relies on ML to detect piracy or Trojans. However, evaluating the robustness of these defenses necessitates rigorous red-teaming, i.e., systematically challenging systems with adversarial examples, to expose algorithmic blind spots or vulnerabilities. NetDeTox [9] advances red-teaming efforts to evade GNN piracy detectors. Instead of brute-force global rewiring, NetDeTox orchestrates a hybrid Reinforcement Learning (RL) and LLM pipeline. The RL agent groups gates into featurebased buckets and then samples high-leverage subnetlists. The LLM acts as a contextual planner, strictly following a decision sequence to generate localized rewriting plans. By applying targeted gate substitutions, NetDeTox bypasses GNN-based detectors across 90% of test cases. Strikingly, this localized red-teaming often achieves negative area overhead, producing adversarial circuits that are fully outperforming prior stateof-the-art [40] and even the baseline circuits. Similarly, TrojanGYM [15] orchestrates an attack-defense loop to evade HT detection. An agent generates HTs; these designs are functionally verified and evaluated by a robust GNN detector. The continuous detection scores serve as feedback, forcing the agent to iteratively restructure the Abstract Syntax Tree (AST) until the Trojan evades detection, achieving 83% evasion rates and systematically mapping detector blind spots. E. Bug Detection and Code Analysis Detecting subtle RTL bugs requires deep contextual understanding. Common Weakness Enumerations (CWEs) present a unique challenge because hardware vulnerabilities often overlap hierarchically and manifest differently depending on micro-architectural context, rendering standard patternmatching brittle. To overcome this, VeriCWEty [41] pioneers an embedding-enabled detection framework. Rather than explicitly extracting ASTs or relying on flat classification, VeriCWEty puts HDL through a Verilog-finetuned decoder. It extracts dense vector embeddings at both the global module level and the line level. By utilizing a majority-voting ensemble of

frontier models to curate high-quality training labels, a simple XGBoost classifier trained on these embeddings achieves 89% precision in identifying critical vulnerabilities (like CWE-1244 and CWE-1245) and 96% accuracy in pinpointing the linelevel location of the bug [41]. LASHED [42] combines generative AI with static analysis. The deterministic tool flags violations, and AI contextualizes them against CWEs to filter false positives. MARVEL [43] employs a hierarchical Multi-Agent architecture where a Supervisor delegates tasks to executors (Linter, CWE, and Similar-Bug RAG agents). Finally, FLAG [44] introduces testfree fault localization, identifying anomalous line-level logic directly from source code by calculating token-level generation probabilities (logprob) and embedding distances. V. E DUCATION AND O UTREACH The rapid integration of generative AI into EDA requires academic curricula and learning environments to radically adapt. Next, we discuss latest efforts towards that end. A. Modular Courseware: GUIDE To standardize education, the GUIDE framework [45] provides an open-source repository translating state-of-the-art research into runnable Google Colab environments. Organized into standardized units (slides, videos, self-contained code), GUIDE enables instructors to construct tailored courses. By mandating standardized deliverables (e.g., waveforms, EDA tool logs), it enforces rigorous engineering practices alongside automated workflows, allowing students to securely practice HT insertion and logic locking evasion. B. Scalable CTFs and Evaluation Capture-the-Flag competitions are crucial testbeds for security education. A cross-regional study on the CSAW CTF formulated scalable principles for AI-assisted environments [46]. To rigorously evaluate these competitions, organizers formalized autonomy levels—Human-in-the-Loop, Autonomous Agent, and Hybrid—and introduced the CTFJudge framework [47] alongside the CTFTiny benchmark. This infrastructure computes a “CTF Competency Index” by mandating traceable evidence (conversational logs and agent execution trajectories) to verify the reasoning-action-output chain. Analysis revealed that, while autonomous workflows outperform humans on complex tool interactions, beginners strongly preferred simple single-agent loops augmented with strict prompt checklists, providing a roadmap for lowering the barrier to entry in secure hardware education. Analyzing how humans collaborate with LLM tools to exploit systems, the “Lowering the Bar” study [10] evaluated a global AI Hardware Attack Challenge. Results demonstrated that competitors with minimal hardware expertise successfully used prompt-engineered agents to locate vulnerabilities and insert severe Trojans (e.g., side-channel leaks and denial-ofservice payloads), underscoring robust alignment guardrails.

Fig. 1. Lessons learned across the wide range of LLM-driven frameworks for (secure) hardware design reviewed in this paper.

VI. L ESSONS L EARNED AND C ONCLUSION Several cross-cutting paradigms defining the future of secure design automation are listed next and illustrated in Fig. 1. Agentic Iteration: Zero-shot prompting fails for complex hardware topologies. State-of-the-art frameworks require iterative, multi-agent loops tightly integrated with deterministic EDA tools. • Hyperparameters vs Architectures: Model capabilities can be largely masked by default configurations. Hyperparameter tuning can induce significant pass-rate swings, making configuration as vital as model selection. • Data Modality vs Capability: Models tuned on Verilog are superior for semantic understanding and fine-grained tasks, while adjacency lists and graphs can excel for structural and topological analysis. • Representation Bottleneck: The target IR fundamentally bounds generation quality. For example, HLS-C often fails due to interface mismatches, while Verilog offers sufficient data richness for training. Advanced coding requires navigating the accessibility-competence paradox. • Alignment Paradox: General-purpose safety filters fail in hardware contexts. They over-refuse legitimate security audits while remaining entirely vulnerable to semantically disguised adversarial engineering requests. • Red-Teaming: ML-based security defenses can project a false sense of security unless subjected to rigorous testing. Red-teaming via LLM-driven frameworks ensure defenses are robust against real-world attackers. •

Dynamic vs Static Benchmarking: Due to pervasive data contamination and rapid adaptive evasion capabilities, static datasets must be replaced by feedback-driven, dynamically generated benchmarks. Overall, while LLMs provide extraordinary acceleration for EDA workflows, they necessitate a fundamentally new paradigm of hardware security relying on orchestrated agents operating within rigorously defined and verified boundaries. •

R EFERENCES [1] Z. Wang, L. Alrahis, L. Mankali, J. Knechtel, and O. Sinanoglu, “LLMs and the future of chip design: Unveiling security risks and building trust,” in Proc. ISVLSI, 2024, pp. 385–390. [2] K. Xu, D. Schwachhofer, J. Blocklove, I. Polian, P. Domanski, D. Pflüger, S. Garg, R. Karri, O. Sinanoglu, J. Knechtel, Z. Zhao, U. Schlichtmann, and B. Li, “Large language models (LLMs) for electronic design automation (EDA): Special session paper,” in Proc. SOCC, 2024. [3] S. Thakur, B. Ahmad, H. Pearce, B. Tan, B. Dolan-Gavitt, R. Karri, and S. Garg, “VeriGen: A large language model for Verilog code generation,” ACM TODAES, vol. 29, no. 3, pp. 46:1–46:31, 2024. [4] Z. Wang, M. Shao, M. Nabeel, P. B. Roy, L. Mankali, J. Bhandari, R. Karri, O. Sinanoglu, M. Shafique, and J. Knechtel, “VeriLeaky: Navigating IP protection vs utility in fine-tuning for LLM-driven Verilog coding,” in Proc. MLCAD, 2025. [5] L. L. Mankali, J. Bhandari, M. Alam, R. Karri, M. Maniatakos, O. Sinanoglu, and J. Knechtel, “RTL-Breaker: Assessing the security of LLMs against backdoor attacks on HDL code generation,” in Proc. DATE, 2025. [6] Z. Wang, M. Shao, J. Bhandari, L. Mankali, R. Karri, O. Sinanoglu, M. Shafique, and J. Knechtel, “VeriContaminated: Assessing LLMdriven Verilog coding for data contamination,” in Proc. MLCAD, 2025. [7] W. Fu, Z. Wang, M. Shao, J. Knechtel, O. Sinanoglu, R. Karri, M. Shafique, and X. Guo, “From natural language to silicon: The representation bottleneck in LLM hardware design,” arXiv preprint arXiv:2604.17097, 2026.

[8] Z. Wang, M. Shao, W. Fu, P. B. Roy, X. Guo, R. Karri, M. Shafique, J. Knechtel, and O. Sinanoglu, “HarmChip: Evaluating hardware security centric LLM safety via jailbreak benchmarking,” arXiv preprint arXiv:2604.17093, 2026. [9] Z. Wang, M. Shao, A. Saha, R. Karri, J. Knechtel, M. Shafique, and O. Sinanoglu, “NetDeTox: Adversarial and efficient evasion of hardwaresecurity GNNs via RL-LLM orchestration,” in Proc. DAC, 2026. [10] J. Blocklove, H. Pearce, and R. Karri, “Lowering the bar: How large language models can be used as a copilot by hardware hackers,” IEEE Security & Privacy, vol. 23, no. 5, pp. 27–37, 2025. [11] J. Blocklove, S. Thakur, B. Tan, H. Pearce, S. Garg, and R. Karri, “Automatically improving LLM-based Verilog generation using EDA tool feedback,” ACM TODAES, vol. 30, no. 6, pp. 100:1–100:26, 2025. [12] M. DeLorenzo, A. B. Chowdhury, V. Gohil, S. Thakur, R. Karri, S. Garg, and J. Rajendran, “Make every move count: LLM-based high-quality RTL code generation using MCTS,” arXiv preprint arXiv:2402.03289, 2024. [13] P. B. Roy, A. Saha, M. Alam, J. Knechtel, M. Maniatakos, O. Sinanoglu, and R. Karri, “Veritas: Deterministic Verilog code synthesis from LLMgenerated conjunctive normal form,” arXiv preprint arXiv:2506.00005, 2025. [14] Z. Wang, M. Shao, R. R. Karn, L. Mankali, J. Bhandari, R. Karri, O. Sinanoglu, M. Shafique, and J. Knechtel, “SALAD: Systematic assessment of machine unlearning on LLM-aided hardware design,” in Proc. MLCAD, 2025. [15] S. Sreekumar, Z. Wang, A. Saha, W. Xiao, M. Shao, M. Shafique, O. Sinanoglu, R. Karri, and J. Knechtel, “TrojanGYM: A detector-in-theloop LLM for adaptive RTL hardware Trojan insertion,” arXiv preprint arXiv:2601.17178, 2026. [16] P. Yubeaton, A. Nakkab, W. Xiao, L. Collini, R. Karri, C. Hegde, and S. Garg, “VeriThoughts: Enabling automated Verilog code generation using reasoning and formal verification,” arXiv preprint arXiv:2505.20302, 2025. [17] M. Akyash, K. Azar, and H. Kamali, “RTL++: Graph-enhanced LLM for RTL code generation,” arXiv preprint arXiv:2505.13479, 2025. [18] L. Collini, S. Garg, and R. Karri, “C2HLSC: Leveraging large language models to bridge the software-to-hardware design gap,” ACM TODAES, vol. 30, no. 6, pp. 96:1–96:24, 2025. [19] R. R. Karn, J. Knechtel, R. Karri, and O. Sinanoglu, “LLM-driven code generation for neural networks on FPGAs: Bridging Python and HLS,” in Proc. ICCD, 2025. [20] L. Collini, A. Hennessee, R. Karri, and S. Garg, “Can reasoning models reason about hardware? an agentic HLS perspective,” arXiv preprint arXiv:2503.12721, 2025. [21] R. Kande, H. Pearce, B. Tan, B. Dolan-Gavitt, S. Thakur, R. Karri, and J. Rajendran, “(Security) assertions by large language models,” IEEE TIFS, vol. 19, pp. 4374–4389, 2024. [22] V. N. Viswambharan, K. K. Radhakrishna, D. N. Gadde, and A. Kumar, “Knowledge graphs, the missing link in agentic AI-based formal verification,” arXiv preprint arXiv:2605.06434, 2026. [23] W. Xiao, D. Ekberg, S. Garg, and R. Karri, “Hybrid-NL2SVA: Integrating RAG and finetuning for LLM-based NL2SVA,” in Proc. MLCAD, 2025, pp. 1–10. [24] J. Bhandari, J. Knechtel, R. Narayanaswamy, S. Garg, and R. Karri, “LLM-aided testbench generation and bug detection for finite-state machines,” arXiv preprint arXiv:2406.17132, 2024. [25] Z. Wang, W. Xiao, M. Shao, R. V. Hemadri, O. Sinanoglu, M. Shafique, and R. Karri, “VeriDispatcher: Multi-model dispatching through preinference difficulty prediction for RTL generation optimization,” arXiv preprint arXiv:2511.22749, 2025. [26] W. Fu, Z. Wang, M. Shao, R. Karri, M. Shafique, J. Knechtel, O. Sinanoglu, and X. Guo, “Synthesis-in-the-loop evaluation of LLMs for RTL generation: Quality, reliability, and failure modes,” arXiv preprint arXiv:2603.11287, 2026. [27] L. Collini, A. Hennesee, P. Yubeaton, S. Garg, and R. Karri, “VeriInteresting: An empirical study of model prompt interactions in Verilog code generation,” arXiv preprint arXiv:2603.08715, 2026. [28] M. Shao, Z. Wang, W. Fu, X. Guo, J. Knechtel, O. Sinanoglu, R. Karri, and M. Shafique, “Configuration over selection: Hyperparameter sensitivity exceeds model differences in open-source LLMs for RTL generation,” arXiv preprint arXiv:2604.17102, 2026.

[29] LLM Benchmarking Coalition, “LLM benchmarking coalition,” https: //si2.org/llm-benchmarking-coalition/, 2026. [30] S. Chen, Y. Chen, Z. Li, Y. Jiang, Z. Wan, Y. He, D. Ran, T. Gu, H. Li, T. Xie, and B. Ray, “Benchmarking large language models under data contamination: A survey from static to dynamic evaluation,” in Proc. EMNLP, 2025, pp. 10 080–10 098. [31] S. Gulyamov, S. Gulyamov, A. Rodionov, R. Khursanov, K. Mekhmonov, D. Babaev, and A. Rakhimjonov, “Prompt injection attacks in large language models and AI agent systems: A comprehensive review of vulnerabilities, attack vectors, and defense mechanisms,” Information, vol. 17, no. 1, p. 54, 2026. [32] A. Saha, P. B. Roy, J. Knechtel, R. Karri, O. Sinanoglu, and L. Alrahis, “GLLaMoR: Graph-based logic locking by large language models for enhanced robustness,” in Proc. VTS, 2025. [33] A. Saha, Z. Wang, P. B. Roy, J. Knechtel, O. Sinanoglu, and R. Karri, “LockForge: Automating paper-to-code for logic locking with multiagent reasoning LLMs,” in Proc. DAC, 2026. [34] S. Ghimire, P. Mirfasihi, M. A. Chowdhury, V. Pugazhenthi, H. K. Dharavath, F. Firouzi, R. Yasaei, P. Satam, and S. Salehi, “Can agents secure hardware? evaluating agentic LLM-driven obfuscation for IP protection,” arXiv preprint arXiv:2604.13298, 2026. [35] L. Collini, J. Bhandari, C. M. Tomajoli, A. Moosa, B. Tan, X. Tang, P.-E. Gaillardon, R. Karri, and C. Pilato, “ARIANNA: An automatic design flow for fabric customization and eFPGA redaction,” ACM TODAES, vol. 30, no. 4, pp. 63:1–63:23, 2025. [36] P. B. Roy, M. Nair, R. Sadhukhan, M. Alam, J. Knechtel, H. Pearce, D. Mukhopadhyay, O. Sinanoglu, and R. Karri, “Netlist whisperer: Extensive analysis of circuit leakage using LLMs,” Journal of Cryptographic Engineering, vol. 15, no. 4, p. 22, 2025. [37] B. Perera, Z. Wang, W. Xiao, M. Nabeel, O. Sinanoglu, J. Knechtel, and R. Karri, “LLM4PQC - accurate and efficient synthesis of PQC cores by feedback-driven LLMs,” in Proc. DATE, 2026. [38] M. Nabeel, B. Perera, Z. Wang, O. Sinanoglu, J. Knechtel, and R. Karri, “LLM4SecurePQC: LLM-driven and side-channel resilient hardware synthesis of PQC cores,” in Proc. VTS, 2026. [39] W. Xiao, Z. Wang, M. Shao, R. V. Hemadri, O. Sinanoglu, M. Shafique, J. Knechtel, S. Garg, and R. Karri, “TrojanLoC: Fine-grained hardware Trojan detection from Verilog code,” arXiv preprint arXiv:2512.00591, 2025. [40] V. Gohil, S. Patnaik, D. Kalathil, and J. Rajendran, “AttackGNN: Redteaming GNNs in hardware security using reinforcement learning,” in Proc. USENIX Security, 2024, pp. 73–90. [41] P. B. Roy, Z. Wang, A. Chuvashlov, W. Xiao, J. Knechtel, O. Sinanoglu, and R. Karri, “VeriCWEty: Embedding enabled line-level CWE detection in Verilog,” arXiv preprint arXiv:2604.15375, 2026. [42] B. Ahmad, H. Pearce, R. Karri, and B. Tan, “LASHED: LLMs and static hardware analysis for early detection of RTL bugs,” arXiv preprint arXiv:2504.21770, 2025. [43] L. Collini, B. Ahmad, J. Ah-kiow, and R. Karri, “MARVEL: Multiagent RTL vulnerability extraction using large language models,” arXiv preprint arXiv:2505.11963, 2025. [44] B. Ahmad, J. Ah-kiow, B. Tan, R. Karri, and H. Pearce, “FLAG: Finding line anomalies (in RTL code) with generative AI,” ACM TODAES, vol. 30, no. 6, pp. 103:1–103:30, 2025. [45] W. Xiao, J. Blocklove, M. DeLorenzo, J. Knechtel, O. Sinanoglu, K. Basu, J. Rajendran, S. Garg, and R. Karri, “GUIDE: GenAI units in digital design education,” in Proc. DATE, 2026. [46] H. Xi, M. Shao, K. Milner, V. S. C. Putrevu, N. Rani, M. Udeshi, P. Krishnamurthy, B. Dolan-Gavitt, S. Garg, S. K. Shukla, F. Khorrami, A. Hillel-Tuch, M. Shafique, and R. Karri, “AI in cybersecurity education–scalable agentic CTF design principles and educational outcomes,” arXiv preprint arXiv:2603.21551, 2026. [47] M. Shao, N. Rani, K. Milner, H. Xi, M. Udeshi, S. Aggarwal, V. S. C. Putrevu, S. K. Shukla, P. Krishnamurthy, F. Khorrami, R. Karri, and M. Shafique, “Towards effective offensive security LLM agents: Hyperparameter tuning, LLM as a judge, and a lightweight CTF benchmark,” in Proc. AAAI, vol. 40, no. 35, 2026, pp. 29 660–29 668.

Record · ID 175114 · SHA-256 d049137f6bca402c
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.