Conceptio › Archive › arXiv CS
arXiv CSopen access

Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

arXiv:2609.10762v1 [cs.CR] 9 Sep 2026

Beyond Static Guarantees: Measuring the Static-Pass Dynamic-Fail Gap in Security-Sensitive and LLM-Generated Python Code Jessica Pourleyli

Maitreyee Das Urmi

Glaucia Melo

Toronto Metropolitan University Toronto, Canada [email protected]

Toronto Metropolitan University Toronto, Canada [email protected]

Toronto Metropolitan University Toronto, Canada [email protected]

evaluation frameworks, including SecurityEval, AutoSafeCoder, RedCode, and VulnRepairEval, to assess the security of generated software [6]–[9]. Despite these advances, security evaluation pipelines continue to rely heavily on static-analysis tools such as Bandit and CodeQL [6], [7]. Static analysis is scalable, reproducible, and inexpensive, but it reasons about code without executing it and therefore cannot directly observe runtime exploit behaviour [10]. Vulnerabilities that depend on adversarial inputs, execution context, or exploit chaining may therefore evade static analysis while remaining exploitable in practice. This limitation is particularly relevant in modern agentic workflows, where successful static checks are often treated as evidence of secure behaviour. In this paper, we investigate what we term the StaticPass Dynamic-Fail (SPDF) phenomenon: vulnerabilities that produce no findings under the evaluated composite BanditSemgrep static analysis gate but are subsequently demonstrated through runtime exploit verification. More broadly, we view software-security evaluation as a hierarchy of evidence. Static analysis, vulnerability reasoning, and dynamic verification provide progressively stronger evidence about software security, but success at one level does not guarantee success at the next. To study this problem, we design a three-stage agentic security pipeline combining static scanning, LLM-driven CWE reasoning, and autonomous exploit generation within isolated Docker environments. The pipeline first filters programs using Bandit [11] and Semgrep [12]. Surviving statically clean samples are then analyzed by a CWE Vulnerability Detection I. I NTRODUCTION Agent that retrieves and reasons over authoritative vulnerability The growing adoption of Large Language Models (LLMs) in knowledge. Finally, a Dynamic Exploit Verification Agent software engineering has raised important questions regarding selectively evaluates candidate vulnerabilities through generated the security and trustworthiness of AI-generated code [1]. exploit harnesses executed in sandboxed Docker environments, Modern coding assistants can generate functional programs, providing higher-confidence exploitability assessments while reason about implementation details, and autonomously interact avoiding the cost of exhaustive dynamic testing. with development tools [2], [3], yet vulnerabilities in generated To evaluate the prevalence of SPDF behaviour, we analyze code remain a persistent concern [4], [5]. As a result, a growing 1,355 Python samples drawn from SecurityEval [6], RedCode body of work has proposed security-focused benchmarks and [8], and the CyberNative Code Vulnerability Security DPO dataset [13]. These datasets contain security-sensitive benchJessica Pourleyli and Maitreyee Das Urmi contributed equally and share first authorship. mark programs and LLM-generated code spanning multiple Abstract—Advances in large language models (LLMs) fuel the quest for scalable methods to assess the security of generated and security-sensitive software. Static analysis is widely adopted as a scalable, reproducible, and inexpensive security gate, but cannot directly observe runtime exploit behaviour. Vulnerabilities dependent on adversarial inputs, execution context, or exploit chaining may evade static checks while remaining exploitable in practice, yet passing static analysis is often treated as evidence of secure behaviour. This paper introduces the Static-Pass Dynamic-Fail (SPDF) phenomenon and a three-stage agentic pipeline combining static scanning, LLM-driven Common Weakness Enumeration (CWE) reasoning, and autonomous exploit verification in isolated Docker containers. We evaluate 1,355 Python samples from SecurityEval, RedCode, and CyberNative datasets. Of the 654 samples producing no findings under the composite Bandit–Semgrep gate, the LLM detection stage identified 394 candidate vulnerabilities across 235 files. Dynamic verification confirmed or partially confirmed exploitability in 95 files, yielding an inclusive pipeline rate of 14.53% (roughly 1 in 7 statically clean samples). This rate represents the proportion of Bandit–Semgrep-clean samples for which the pipeline identified a candidate vulnerability and obtained runtime evidence supporting exploitability. Outcomes varied by dataset: among candidate file–CWE pairs, confirmed exploitability was 33.7% for RedCode, 28.6% for CyberNative, and 5.4% for SecurityEval. Several frequently confirmed classes, including CWE-338 and CWE-916, were flagged by neither Bandit nor Semgrep. These findings indicate that static-analysis success and runtime security are hierarchical layers of software assurance rather than interchangeable measures, and have the potential to reshape how AI-generated and security-sensitive code is evaluated. Index Terms—software security, static analysis, dynamic analysis, exploit verification, agentic security evaluation, large language models

CWE categories, adversarial coding scenarios, and exploitoriented tasks. We additionally quantify the practical cost of agentic exploit verification, since computational requirements influence the feasibility of deploying such techniques in realworld security evaluation workflows. The following research questions guide our study: RQ1: To what extent does code that produces no findings under the evaluated static-analysis gate remain dynamically exploitable under runtime testing? We aim to quantify the observed yield of the SPDF pipeline among samples producing no findings under the evaluated static-analysis gate and examine how often LLM-identified candidates can subsequently be demonstrated as exploitable at runtime. RQ2: What categories of vulnerabilities are most likely to evade static analysis while remaining exploitable during execution? We aim to identify which CWE categories are most frequently associated with SPDF behaviour and, therefore, most likely to require dynamic verification. RQ3: What is the cost of the agentic dynamicexploitability pipeline in terms of latency and token usage? This question evaluates the practical feasibility of the proposed approach by quantifying its computational cost. The main contributions of this work are as follows: 1) We introduce and operationalize the Static-Pass DynamicFail (SPDF) phenomenon and the SPDF Rate (SPDFR), providing an empirical characterization of its prevalence and the vulnerability categories most likely to evade static detection. 2) An empirical characterization of the computational cost of agentic exploit verification, including token consumption, latency, reasoning iterations, and exploit-generation overhead across 394 candidate vulnerabilities. 3) A curated collection of SPDF-positive samples identified from existing security-oriented benchmarks, together with evaluation artifacts, exploit evidence, and exploitverification code used during the verification process to support reproducible research on runtime vulnerability assessment for LLM-generated software. 4) A reproducible and publicly available multi-stage pipeline combining static analysis, LLM-driven CWE detection, and Docker-based exploit verification [14].

II. R ELATED W ORK A. Security of LLM-Generated Code LLM-based coding assistants can generate code that appears useful and may functionally satisfy program needs, yet still contain security weaknesses. Prior evaluations of GitHub Copilot and ChatGPT show that generated programs may contain vulnerabilities associated with known CWE categories, even when the generated code is syntactically valid or functionally plausible [4], [5]. These findings establish that functional code generation and secure code generation are not necessarily equivalent evaluation goals that should be treated equally during the code generation process. This distinction has become increasingly important as LLMs are adopted across software engineering tasks such as code generation, debugging, repair, and maintenance [1]. Security evaluation, therefore, requires assessing whether generated code is robust against vulnerability patterns, not only whether it satisfies the surface-level programming task. B. Security Benchmarks and Evaluation Frameworks for AIGenerated Code Security-focused benchmarks have been introduced to evaluate generated code beyond functional correctness. SecurityEval provides Python prompts mapped to CWE categories and demonstrates their use for evaluating code-generation models using both manual inspection and automated static analysis tools [6]. RedCode expands security evaluation to code agents by testing risky code execution across Python and Bash tasks, including sandboxed execution environments and evaluation metrics for unsafe behaviour [8]. Recent frameworks also show a shift toward executionaware security evaluation. AutoSafeCoder combines LLMbased code generation with static analysis feedback and fuzzing-based dynamic testing in a multi-agent framework [7]. VulnRepairEval evaluates LLM-based vulnerability repair using functional proof-of-concept exploits and argues that exploit-based validation provides a stricter test of whether a vulnerability has actually been repaired [9]. Together, these works motivate security evaluation methods that go beyond static inspection alone, an area that our work complements and extends. C. Static and Dynamic Analysis

By focusing on the gap between static security guarantees and Static analysis remains widely used for evaluating the observed runtime exploitability, this work contributes toward a security of generated code because it can be automated and more execution-aware framework for evaluating the security applied to large numbers of code samples. Prior work has of AI-generated and security-sensitive software systems. We used tools such as Bandit and CodeQL to identify security emphasize that our corpus comprises security-sensitive bench- weaknesses in generated code and benchmark datasets [6], mark code and LLM-generated vulnerable samples, not code [7]. However, empirical evaluations of SAST tools show that produced by a coding assistant operating within an interactive no single static-analysis tool detects all vulnerabilities, and development workflow. Our results, therefore, characterize substantial numbers of vulnerabilities may remain undetected Static Application Security Testing (SAST) false-negative even when multiple tools are combined [15]. Recent work exposure on such corpora; extending them to assistant-in-the- improving Semgrep-based SAST pipelines further demonstrates loop pipelines is a hypothesis our findings motivate but do not that static-analysis effectiveness is heavily dependent on rule establish. coverage and vulnerability-pattern representation, highlighting

the potential for false negatives in automated security evaluation [15]. However, static analysis does not execute the program under concrete runtime conditions. It can identify many securityrelevant patterns, but runtime exploitability may depend on adversarial inputs, execution context, or whether a vulnerable path can actually be triggered. This limitation motivates dynamic testing approaches such as fuzzing, sandboxed execution, and exploit-based validation. Prior work has already used fuzzing and proof-of-concept exploits to strengthen security evaluation and aid vulnerability repairs [7], but the specific case of examining existing security-sensitive datasets for code that passes static analysis and later fails under dynamic exploit verification remains underexplored. D. Agentic Security Pipelines and Autonomous Software Evaluation

contributes adversarial execution scenarios, and CyberNative provides LLM-generated vulnerable code. Together, these datasets capture benchmark-driven, adversarial, and LLMgenerated security contexts meant to provide complementary perspectives on software security and vulnerable code. Only Python samples were retained to ensure compatibility across Bandit, Semgrep, LLM-guided CWE analysis, and Dockerbased execution. Each sample was normalized into a standalone Python file and assigned source-dataset metadata. Table I summarizes the final dataset collection used in this study. TABLE I DATASET COMPOSITION USED IN THE SPDF EVALUATION . Dataset SecurityEval RedCode CyberNative

Samples

Role in Study

121 810 424

Benchmark security evaluation Adversarial execution scenarios LLM-generated vulnerable code

Modern LLM systems increasingly combine reasoning, tool use, and interaction with external environments. ReAct Total 1355 Combined evaluation collection formalizes an approach in which language models interleave reasoning steps with actions that query tools or environments [16]. SWE-agent applies tool-mediated agent interaction to B. Pipeline Overview software engineering tasks, allowing agents to edit files, The pipeline consists of three stages, as shown in Figure navigate repositories, and execute tests through an agent- 1: (1) Bandit–Semgrep static analysis, (2) LLM-guided CWE computer interface [17]. vulnerability detection over statically clean files, and (3) DockerSecurity-oriented systems have begun applying similar agen- based dynamic exploit verification of identified candidates. A tic patterns to code generation and evaluation. AutoSafeCoder sample qualifies as SPDF only when it passes the composite uses multiple agents for code generation, static vulnerability static gate, is assigned a candidate CWE by the detection agent, analysis, and fuzzing-based testing [7]. RedCode evaluates and subsequently exhibits runtime evidence of exploitability. code agents in scenarios involving risky code execution and The two agentic stages use separate configurations but share generation, highlighting the need for safety evaluation in tool- the same underlying model, a dependency considered when using coding agents [8]. interpreting the results. Our work builds on this direction by focusing on the gap between static validation and runtime exploitability in an C. Static Analysis Stage The Static Analysis Stage serves as the initial filtering stage agentic pipeline that combines static scanning, CWE-guided vulnerability reasoning for detection, and Docker-based exploit of the proposed pipeline. Its objective is to establish the staticpass condition by identifying files containing vulnerabilities verification. detectable by conventional static-analysis tools before any LLMIII. M ETHODOLOGY based reasoning or dynamic testing is performed. To accomplish this, we employ two complementary staticWe develop a three-stage security evaluation pipeline consisting of a composite Static Analysis Stage, a CWE Vulnerability analysis tools: Bandit [11] and Semgrep [12]. Bandit focuses Detection Agent, and a Dynamic Exploit Verification Agent. on Python-specific security weaknesses through a collection of The latter two stages use LLM-based agents, while the initial security-oriented checks, while Semgrep provides rule-based stage deterministically orchestrates conventional static-analysis detection of broader vulnerability patterns. Using both tools tools. The pipeline progressively filters, analyzes, and validates reduces reliance on a single detection strategy and establishes security weaknesses using both static and execution-based a stricter composite filtering stage. All static-analysis results were generated using Bandit techniques. Samples successfully exploited after passing static v1.8.4.dev21 and Semgrep v1.155.0. Bandit was executed recuranalysis are classified as SPDF instances. sively across all Python files and configured to produce JSON A. Dataset Construction output for downstream processing. Semgrep was executed using To evaluate SPDF across diverse software-security contexts, the auto ruleset and JSON output. No additional confidence we construct a unified collection of 1,355 Python samples from or severity thresholds were applied beyond those used by the three complementary security-oriented datasets: SecurityEval tools themselves. Consequently, any file containing at least one [6], RedCode [8], and the CyberNative Code Vulnerability finding reported by either Bandit or Semgrep was classified as Security DPO dataset [13]. SecurityEval provides manually statically flagged and removed from further analysis. Files curated vulnerability-oriented programming tasks, RedCode producing no findings from either tool were classified as

statically clean and advanced to the CWE Vulnerability Detection Agent. For each flagged file, the static-analysis stage records the originating tool, rule identifier, CWE mapping (when available), severity information, confidence information, and the associated source-code location. Findings from both tools are merged at the file level to produce a composite static-analysis result. A file is considered statically clean only when neither Bandit nor Semgrep reports any findings. Across the 1,355 Python samples analyzed, Bandit reported 922 findings, and Semgrep reported 816 findings. After merging detections across both tools and removing duplicate file-level reports, 701 unique files were flagged by at least one staticanalysis tool. The remaining 654 files passed composite static analysis and formed the input to the second stage of the SPDF pipeline.

reporting a CWE unless the corresponding MITRE CWE definition has been retrieved and examined. The agent produces structured vulnerability reports that include the predicted CWE identifier, vulnerability name, severity classification, and supporting rationale. By grounding assessments in established CWE definitions rather than relying solely on unconstrained model reasoning, the agent provides a more interpretable and security-focused analysis of statically clean code samples. Because LLM-based security judgments may vary across executions, the agent is executed independently five times for each sample. Each run is permitted a maximum of ten reasoning and tool-use iterations. Agent execution was bounded by interaction turns rather than a fixed token budget. No explicit per-instance token limit was imposed beyond the underlying model’s context window, allowing token consumption to vary with analysis complexity. Final vulnerability predictions are consolidated using majority-vote aggregation at the file–CWE D. CWE Vulnerability Detection Agent level, where a vulnerability is retained only if it appears in The CWE Vulnerability Detection Agent analyzes the a strict majority of runs (at least three of five executions). A statically clean samples produced by the Static Analysis Stage. five-run ensemble was selected as a practical balance between Implemented using OpenAI’s gpt-5.4-nano model, the prediction stability and computational cost. This procedure agent is prompted to act as a security analyst whose objective is reduces the impact of stochastic model behaviour and produces to identify vulnerabilities that may have been missed by conven- a more conservative and stable set of candidate vulnerabilities tional static-analysis tools while grounding its assessments in for subsequent dynamic verification. All other model settings authoritative vulnerability knowledge. All experiments used the were kept to their default parameters. OpenAI model identifier gpt-5.4-nano with deterministic decoding (temperature=0). The complete system prompt, E. Dynamic Exploit Verification Agent tool specifications, and agent configuration are provided in the The Dynamic Exploit Verification Agent evaluates whether replication package. candidate vulnerabilities identified by the CWE Vulnerability The agent follows a ReAct-style reasoning process that Detection Agent can be triggered under concrete execution combines source-code inspection with retrieval-augmented conditions, requiring observable runtime evidence before a security analysis. For each sample, the agent examines program vulnerability is considered CONFIRMED. The agent uses inputs, data flows, API usage, and security-sensitive operations OpenAI’s gpt-5.4-nano with temperature=0.3, automatic to identify behaviours potentially associated with known vul- tool selection, and a maximum of 30 reasoning and tool-use nerability classes. In this study, a candidate vulnerability refers iterations per execution. Agent execution was constrained by to a file–CWE pairing identified by the CWE Vulnerability this turn budget rather than a fixed token budget; with no Detection Agent based on security-relevant patterns observed explicit token limit beyond the model context window. The in the source code and supported through comparison with non-zero temperature encourages exploration of diverse exploit external security knowledge and the corresponding MITRE strategies, while five-run majority voting mitigates variability CWE definition. These findings represent hypothesized vul- across attempts. nerabilities that have not yet undergone dynamic exploit The agent follows a ReAct-style process combining retrievalverification. When a candidate’s weakness is identified, the augmented vulnerability analysis with autonomous exploit agent retrieves supporting information from external sources, generation. For each file–CWE pair, the agent receives the including official MITRE CWE definitions, before assigning a target code, the predicted vulnerability classification, and vulnerability classification. The agent is explicitly instructed supporting context from the previous stage. The agent then to validate suspected vulnerabilities against the corresponding curates its own exploit-tests by retrieving information from CWE definition and supporting evidence before reporting them, the MITRE CWE database and the National Vulnerability thereby reducing unsupported or speculative findings. Database (NVD) [18], leveraging vulnerability descriptions, Three tools are available to the agent. The search_web attack patterns, and real-world examples associated with the tool retrieves security information related to vulnerability target CWE. This grounding enables exploit generation to be patterns, CWEs, CVEs, and Python security weaknesses. informed by established vulnerability knowledge rather than The fetch_cwe_detail tool retrieves the correspond- relying solely on model reasoning. ing CWE entry from the MITRE repository. Finally, the Six tools are available to the agent. The fetch_cwe_info report_findings tool records the agent’s final vulner- tool retrieves vulnerability definitions, examples, and supportability assessments. The system prompt explicitly prohibits ing information from the MITRE CWE repository, while

fetch_nvd_cve_examples retrieves real-world CVEs associated with the target CWE from the National Vulnerability Database (NVD). The write_harness and read_file tools enable the agent to generate and inspect exploit harnesses, run_in_docker executes generated harnesses within an isolated runtime environment, and report_verdict records the final exploitability assessment. The agent is required to retrieve both the corresponding MITRE CWE definition and NVD examples before exploit harness generation, grounding exploit construction in documented vulnerability behaviour rather than unconstrained model reasoning. Rather than executing a fixed set of tests, the agent dynamically determines the number and type of exploit attempts required for a given vulnerability. Exploit generation is tailored to both the target CWE and the structure of the analyzed code. For example, command-injection vulnerabilities are evaluated across multiple payload variations and execution paths, pathtraversal vulnerabilities are tested with alternative traversal strategies, and deserialization vulnerabilities are assessed with malicious serialized inputs. Additional exploit attempts may be generated when previous executions are inconclusive or when multiple attack surfaces are identified within the target program. This adaptive strategy allows the agent to explore vulnerabilityspecific attack vectors while avoiding a one-size-fits-all testing methodology. Generated exploit harnesses are executed within isolated Docker containers configured with disabled network access, a 256 MB memory limit, and configurable execution timeouts. This sandboxed environment enables the safe execution of potentially harmful payloads while providing a consistent, reproducible runtime. During execution, the pipeline records program output, execution status, resource utilization, and error conditions for subsequent analysis and to support final exploitation verdicts. The agent additionally records generated exploit harnesses, Docker execution outputs, exit codes, timeout events, and supporting evidence used to justify final exploitability decisions. A vulnerability is dynamically verified only when observable runtime behaviour consistent with the target CWE and NVD evidence is demonstrated. Criteria are vulnerability-specific: for example, command injection requires unintended command execution while path traversal requires unauthorized file access beyond intended boundaries. Runtime exceptions alone are not considered successful exploitation unless they constitute the target weakness. Following execution, the agent assigns one of four verdicts defined in Table II. A NOT TRIGGERED verdict indicates that testing did not produce sufficient runtime evidence of exploitation; it does not establish that the candidate is absent or unexploitable under other conditions. Each file–CWE pair is evaluated independently across five executions, allowing multiple vulnerability hypotheses within the same file to be verified separately. A vulnerability is dynamically verified only when at least three of five runs return CONFIRMED, reducing the impact of stochastic exploit generation.

TABLE II DYNAMIC EXPLOIT VERIFICATION VERDICT DEFINITIONS . Verdict

Definition

CONFIRMED

Runtime evidence demonstrates successful exploitation of the target CWE. Testing revealed behaviour consistent with the target CWE, but the observed evidence was insufficient to conclusively demonstrate full exploitation. Multiple exploit attempts fail to produce evidence of exploitation under the evaluated harnesses and execution conditions. The vulnerability cannot be reliably evaluated under the available runtime conditions.

PARTIAL

NOT TRIGGERED INCONCLUSIVE

Both LLM-based stages use the same underlying model (gpt-5.4-nano), although they operate with distinct prompts, tools, and objectives. Consequently, their outputs are not fully independent: model-specific reasoning tendencies may be shared across detection and verification, creating the possibility of correlated errors or self-confirmation. We account for this dependency when interpreting the resulting exploitability estimates and discuss its implications further in Section V and Section VI. F. SPDF Measurement To quantify the prevalence of the Static-Pass Dynamic-Fail (SPDF) phenomenon, we define the Static-Pass Dynamic-Fail Rate (SPDFR) as: SPDFR =

|{static-pass} ∩ {dynamic-fail}| |{static-pass}|

(1)

where static-pass denotes unique files that produce no findings under the composite Bandit–Semgrep static-analysis stage, and dynamic-fail denotes unique files for which the Dynamic Exploit Verification Agent subsequently confirms exploitability through runtime testing.SPDFR is reported at the file level using all samples that pass the evaluated Bandit–Semgrep gate as the denominator. However, dynamic verification is applied only to file–CWE candidates first identified by the CWE Vulnerability Detection Agent. Accordingly, SPDFR should be interpreted as the observed yield of the complete detection-and-verification pipeline among Bandit/Semgrep-clean samples, rather than as a direct estimate of the total prevalence of exploitable vulnerabilities in the entire static-pass set. Files not surfaced by the LLM detection stage are not dynamically tested and may therefore contain additional exploitable weaknesses that this pipeline does not observe. We include both conservative and inclusive definitions of the SPDFR, where conservative means only CONFIRMED verdicts are included in the calculation, while the inclusive definition incorporates both CONFIRMED and PARTIAL verdicts. In addition to SPDFR, we analyze the distribution of CWE categories among confirmed SPDF instances to identify vulnerability classes most likely to evade static detection. We also record latency, token consumption, reasoning iterations,

and exploit-generation activity for the agentic stages to quantify the computational cost of the proposed pipeline. G. Statistical Analysis To assess whether dynamic exploit confirmation varies across benchmark datasets, we employ Pearson’s Chi-Square Test of Independence on contingency tables constructed from CONFIRMED and non-CONFIRMED candidate file–CWE pairs, as shown in Table III. Statistical significance is evaluated at α = 0.05. To examine whether exploit confirmation differs across vulnerability categories, we perform a permutation-based Chi-Square test using candidate file–CWE pairs and their corresponding exploit-verification outcomes. This approach was selected because several CWE categories contain small sample sizes that violate the assumptions of standard asymptotic tests. TABLE III S TATISTICAL ANALYSIS OF CONFIRMED EXPLOITABILITY AMONG CANDIDATE FILE –CWE PAIRS . Factor

Key observations

Dataset RedCode CyberNative SecurityEval

106/315 (33.7%) 12/42 (28.6%) 2/37 (5.4%)

CWE (Top 5 Shown) CWE-400 CWE-338 CWE-330 CWE-916 CWE-22

27/49 15/15 14/14 9/14 7/35

Significance

χ2 (2) = 12.55, p = 0.0019

χ2 (22) = 138.06, p < 0.0001∗

∗ Permutation-based p-value. Values are reported as CONFIRMED candidate

file–CWE pairs divided by total candidate file–CWE pairs for the corresponding dataset or CWE category.

H. Reproducibility and Artifact Release To support reproducibility and future research on the Static-Pass Dynamic-Fail (SPDF) phenomenon, we release the complete SPDF evaluation pipeline, generated exploit harnesses, and evaluation artifacts. We also release the subset of samples identified in this study as exhibiting SPDF behaviour, including exploit evidence, verification metadata, and dynamic analysis outcomes. These artifacts are intended to facilitate future research on runtime vulnerability verification, exploitaware security evaluation, and the limitations of static-analysisbased security assessment. The complete system prompts, tool specifications, and agent implementations used by the CWE Vulnerability Detection Agent and Dynamic Exploit Verification Agent are included in the replication package [14].

1,355 Python Sam ples

St atic Scanner A gent

654 St atically Clean Files

CW E V ulnerabilit y D etection A gent

235 Files (394 Candidate CW Es)

120 / 394 CONFIRMED

20 8 / 394 NOT TRIGGERED

20 / 394 PARTIAL

4 6 / 394 INCONCLUSIVE

Static- Pass Dy namic- Fail Rate: Conservative - 83 Files - 12.69% Inclusive - 95 Files - 14.53%

D ynam ic Exploit ation Ver ification A gent

D ynam ic Exploit ation R esults

Fig. 1. Summary statistics across stages of the dynamic exploit verification pipeline.

Then, we examine how SPDF behaviour varies across benchmark datasets and vulnerability categories, before assessing the computational cost of the proposed agentic verification process. Collectively, these analyses provide a multifaceted view of the relationship between static security assessment and demonstrated exploitability. A. RQ1: To What Extent Does Code That Passes Static Analysis Remain Dynamically Exploitable? Figure 1 presents the progression of samples through the proposed SPDF pipeline. Of the 1,355 Python samples analyzed, 701 unique files were flagged by either Bandit and/or Semgrep during static analysis. The remaining 654 samples passed the composite static-analysis stage and were therefore classified as statically clean and advanced to the second agent. The CWE Vulnerability Detection Agent subsequently identified 394 candidate vulnerabilities across 235 unique statically clean files. These candidate files–CWE pairs were evaluated for potential exploitation by the Dynamic Exploit Verification Agent, as shown in Table IV. TABLE IV DYNAMIC EXPLOIT VERIFICATION OUTCOMES . V ERDICT COUNTS ENTAIL TESTED FILES –CWE INSTANCES THAT FALL INTO EACH CATEGORY. P ERCENTAGES CALCULATED AS C OUNT / C ANDIDATE CWE F INDINGS Outcome

Count

Percentage

Files analyzed Candidate CWE findings

235 394

– –

CONFIRMED PARTIAL NOT TRIGGERED INCONCLUSIVE

120 20 208 46

30.5% 5.1% 52.8% 11.7%

Table V summarizes the resulting SPDFR metrics. Under the conservative definition, 120 confirmed file–CWE pairs corresponded to 83 unique files, yielding an observed conservative pipeline yield of 12.69% (83/654) relative to all Bandit/Semgrep-clean files. Including PARTIAL outcomes produced 140 file–CWE pairs across 95 unique files, corresponding to an observed inclusive yield of 14.53% (95/654). These percentages characterize the output of the complete SPDF pipeline: dynamic verification was performed only on candidates surfaced by the CWE Detection Agent. They IV. R ESULTS should therefore not be interpreted as direct estimates of the This section presents the results of the SPDF evaluation total prevalence of exploitable vulnerabilities among all 654 pipeline. We first quantify the prevalence of vulnerabilities static-pass files. Because exploitability was assessed using a that survive static analysis yet remain exploitable at runtime. specific agent architecture, model configuration, and execution

Top 15 candidate CWE findings and dynamic exploit verification outcomes CWE-400

27

CWE-22

7

CWE-73

49

1

35

3

6

32

CWE-94

19

CWE-250 1 CWE-532

CWE category

environment, these values should be interpreted as a modeldependent lower bound of demonstrated exploitability within the evaluated benchmark collection. The aggregate results were strongly influenced by dataset composition. RedCode accounted for 315 of the 394 candidate file–CWE pairs (79.9%) and 106 of the 120 CONFIRMED outcomes (88.3%). Among candidate file–CWE pairs, confirmation rates were 33.7% for RedCode (106/315), 28.6% for CyberNative (12/42), and 5.4% for SecurityEval (2/37). The association between dataset and confirmation outcome was statistically significant (χ2 (2) = 12.55, p = 0.0019). Across the dynamically evaluated candidates, the Dynamic Exploit Verification Agent produced 120 CONFIRMED exploitations and an additional 20 PARTIAL exploitations. Combined, 35.6% (140/394) of all candidate file–CWE pairs exhibited at least some evidence of exploitability during dynamic verification.

17 3

16

2

CWE-338

15

CWE-330

14

CWE-916

9

CWE-770 4

CWE-200

14

1

6

CWE-401

15 14

13

1

12

4

3

12

CWE-798 2 1

10

CWE-285 1

10

CWE-434

10

0

10

Detected candidate CWEs Confirmed Partial 20

30

Number of file--CWE evaluations

40

50

Fig. 2. Top 15 CWE categories identified during dynamic verification. TABLE V S TATIC -PASS DYNAMIC -FAIL R ATES (SPDFR).

Among candidate file–CWE pairs identified by the CWE Detection Agent, RedCode exhibited the highest exploitConservative SPDFR 83 12.69% confirmation rate (33.7%), followed by CyberNative (28.6%), Inclusive SPDFR 95 14.53% while SecurityEval produced substantially fewer confirmed exploitations (5.4%). These results suggest that vulnerabilities These findings provide direct evidence of the Static-Pass surviving static analysis are not equally likely to remain Dynamic-Fail phenomenon: more than one in eight samples exploitable across benchmark datasets. that passed a composite Bandit–Semgrep gate was subsequently We next examined exploitability across vulnerability catshown to contain CONFIRMED vulnerabilities triggerable egories. A permutation-based chi-square test revealed a sigthrough execution-based testing. Static-analysis success should nificant association between CWE category and confirmed therefore not be read as evidence that software resists exploitaexploitability (χ2 (22) = 138.06, p < 0.0001), indicating that tion. SPDF behaviour is strongly dependent on vulnerability class. Equally notable is the reduction between vulnerability Figure 2 compares candidate vulnerability findings against suspicion and confirmation. Of the 394 candidate weaknesses confirmed and partial exploitations across the 15 most freacross 235 files, 120 were confirmed and 20 partially confirmed, quently observed CWE categories. Several categories exhibwhile 208 were not triggered under the generated harnesses and ited substantial reductions between candidate identification available execution conditions and 46 remained inconclusive. and confirmed exploitability, indicating that many suspected Importantly, a NOT TRIGGERED verdict indicates only that weaknesses could not be reproduced under dynamic testing. the agent did not obtain runtime evidence of exploitation In contrast, CWE-400 (Uncontrolled Resource Consumption) during the performed tests; it does not establish that the produced both the largest number of candidate findings and candidate weakness is a false positive or is unexploitable under the largest number of confirmed exploitations (27 cases). Other other harnesses or environments. Thus, fewer than half of the frequently confirmed categories included CWE-338 (Use of reasoning-identified weaknesses translated into demonstrated Cryptographically Weak Pseudo-Random Number Generator), exploitability within our evaluation. Dynamic verification thus operates as both a discovery and a validation mechanism; we CWE-330 (Use of Insufficiently Random Values), CWE-916 (Use of Password Hash With Insufficient Computational Effort), discuss the implications of this layered view in Section V. and CWE-22 (Path Traversal). The most frequently confirmed SPDF categories involve B. RQ2: What Categories of Vulnerabilities Are Most Likely security properties difficult to evaluate through syntactic to Evade Static Analysis While Remaining Exploitable During inspection alone. Resource-exhaustion weaknesses frequently Execution? manifested through excessive memory allocation, computational To examine whether exploit confirmation rates vary across amplification, or non-terminating execution paths that only benchmark sources, we performed a Pearson Chi-Square test of became apparent during execution. Similarly, randomness and independence. As shown in Table III, the relationship between cryptographic weaknesses often appeared functionally correct dataset source and confirmed exploitability was statistically or unsuspicious to security patterns, while violating underlying significant (χ2 (2) = 12.55, p = 0.0019). security assumptions. Consequently, these vulnerabilities surMetric

Count (Files)

Rate

vived static analysis despite remaining exploitable in practice. TABLE VII AGGREGATE COMPUTATIONAL COST OF THE SPDF PIPELINE . Confirmation rates also varied sharply across classes. CWE338 and CWE-330 reached 100% confirmation, with every Stage Tokens Latency (s) candidate vulnerability identified by the CWE Detection Agent Static Analysis Stage 0 77.3 and then labeled CONFIRMED through dynamic verification. CWE Detection Agent 24.6M 37833.7 In contrast, several categories frequently identified during CWE Dynamic Verification Agent 71.7M 65383.1 reasoning failed to produce successful exploit demonstrations. Total 96.3M 103294.1 Examples include CWE-94 (Code Injection), CWE-434 (Unrestricted File Upload), and CWE-20 (Improper Input Validation), TABLE VIII all of which produced candidate findings but no confirmed P ER - EXECUTION COST OF DYNAMIC EXPLOIT VERIFICATION ACROSS ALL exploitations. This variation suggests that exploitability is FIVE VERIFICATION RUNS (1,970 TOTAL EXECUTIONS ). highly dependent on vulnerability class and that the presence of a candidate weakness does not necessarily imply practical Metric Mean Median IQR Max exploitability. Agent Turns 8.34 8 7–10 30 Total Tokens 36,388 33,331 28,763–42,911 219,084 An additional finding concerns vulnerability categories that Latency (s) 33.19 25.19 22.53–29.29 1231.08 were absent from the initial static analysis stage. As shown in Table VI, several of the most frequently confirmed SPDF categories produced no Bandit or Semgrep findings anywhere While aggregate cost characterizes the resources required to in the collection despite later being identified by the CWE conduct the full study, Table VIII reports the cost of individual Detection Agent and confirmed through dynamic verification. dynamic-verification evaluations calculated across the 5 runs. Most notably, CWE-338 resulted in 15 CONFIRMED exploits Means consistently exceeded medians (Table VIII), indicating despite no static-analysis findings, while CWE-916 resulted in a small tail of difficult cases requiring substantially more 9 CONFIRMED exploits despite being absent from both tools. reasoning and exploit generation; the most expensive evaluation Similar patterns were observed for CWE-73, CWE-770, and consumed 219,084 tokens and over 20 minutes. CWE-862. These results highlight the cost–depth trade-off: dynamic This indicates that the exploit verification is substantially more expensive than static observed SPDF behaviour TABLE VI analysis and is therefore better suited to selective verification is not limited to isolated TOP -5 CONFIRMED CWE S M ISSED than exhaustive screening a point we return to in Section V. BY S TATIC A NALYSIS false negatives within covered classes: entire cateV. D ISCUSSION CWE Confirmed Cases gories of exploitable weakThe central finding of this study is that software security is nesses were absent from CWE-338 15 better understood as a hierarchy of evidence than as a binary the static-analysis output, CWE-916 9 property. Static analysis, vulnerability reasoning, and exploit yet confirmed at runCWE-73 6 verification provide different levels of evidence, and success time. At the same time, CWE-770 6 at one stage does not guarantee success at the next. Across some confirmed SPDF catCWE-862 6 the pipeline, 1,355 samples were reduced to 654 statically egories were represented clean files, 394 candidate weaknesses, and 120 confirmed in the static-analysis output elsewhere in the collection. For example, CWE-400, CWE- exploitations across 83 files, showing that exploitable behaviour 330, and CWE-22 appeared among Bandit findings but were can persist through increasingly stringent evaluation. The nevertheless confirmed in statically clean samples that advanced conservative SPDF rate of 12.69% (more than one in eight through the pipeline. This suggests that even when a static tool statically clean files) represents the observed yield of the full provides coverage for a vulnerability class, individual instances detection-and-verification pipeline and indicates that passing may still evade detection depending on implementation details, composite static analysis is better interpreted as evidence of reduced risk than as evidence of security. program context, and runtime behaviour. Because CyberNative contains LLM-generated code, the C. RQ3: What Is the Computational Cost of the Agentic results also show that statically clean LLM-generated programs Dynamic-Exploitability Pipeline? can remain exploitable. Whether this extends to code produced The Dynamic Exploit Verification Agent evaluated 394 by assistants in interactive development workflows, where files–CWE pairs and performed 1,304 Docker executions, context, iteration, tool use, and human oversight differ, remains corresponding to an average of approximately 3.3 exploit unresolved. attempts per candidate vulnerability. Table VII reports aggregate cost. The CWE Detection Agent A. Interpreting the SPDF Rate consumed roughly 24.6M tokens (10.5 hours) and the Dynamic The aggregate SPDF rates should be interpreted in light Verification Agent 71.7M tokens (18.2 hours), for a pipeline of substantial dataset heterogeneity. RedCode contributed 315 total of approximately 96.3M tokens and 28.7 hours. of the 394 candidate file–CWE pairs (80%) and 106 of the

120 CONFIRMED exploitations (88%), so the aggregate rate combine it with execution-based verification rather than relying is overwhelmingly a property of a single benchmark. The on either in isolation. significant association between dataset source and confirmed For vulnerability detection, both exploitability and static exploitability (χ2 (2) = 12.55, p = 0.0019) reinforces this: coverage are unevenly distributed across weakness classes. SPDF behaviour is not uniformly distributed across the corpus. Several frequently confirmed categories were absent from the Notably, SecurityEval, the only manually curated, generation- static-analysis output entirely (Table VI), indicating that SPDF task benchmark of the three, produced the lowest confirmed rate behaviour is not merely isolated false negatives within covered (5.4%), whereas the adversarial and intentionally vulnerable classes but, in some cases, whole categories falling outside corpora produced the highest (33.7% and 28.6%). The aggregate rule-based coverage. These tend to involve semantic properties, inclusive rate of roughly one in seven should therefore be randomness quality, cryptographic strength, authorization logic, read as an upper-leaning estimate driven by adversarial and and resource consumption, that resist syntactic pattern matching, vulnerability-seeded code, not as a direct prediction of how suggesting that reasoning- and execution-based analysis can often statically clean code from realistic generation tasks will complement rule-based tooling. prove exploitable. For agentic security systems, dynamic verification served as A second consideration concerns the shared model backbone both discovery and validation: it confirmed 120 weaknesses that across pipeline stages. Both the CWE Detection Agent and survived earlier stages, while 208 candidate weaknesses were the Dynamic Exploit Verification Agent are implemented not triggered under the generated harnesses and available execuwith gpt-5.4-nano, so the verifier confirms hypotheses tion conditions. These results demonstrate how dynamic verifigenerated by a sibling configuration of the same model. cation can distinguish reasoning-based vulnerability hypotheses This risks correlated reasoning errors and a degree of self- from vulnerabilities for which concrete runtime evidence can be confirmation: weaknesses the detector is predisposed to surface obtained, without treating unsuccessful exploitation as evidence may also be those the verifier is predisposed to demonstrate. that the candidate weakness is necessarily absent. This depth The categories with perfect confirmation rates, CWE-338 and is expensive, roughly 96.3 million tokens and 29 hours, so CWE-330, each confirmed in every candidate instance, are exploit verification is best deployed selectively within a layered consistent with this concern, since their verification often strategy in which static analysis narrows the search space and reduces to demonstrating that a generator is non-cryptographic, reasoning prioritizes candidates. The broader implication is a property both agents recognize from the same surface cues. that exploitability deserves treatment as a first-class security Majority voting over five runs reduces stochastic noise but outcome rather than an implicit byproduct of detection. cannot break this correlation, because all runs draw on the VI. T HREATS TO VALIDITY same model; pairing a detector and verifier from different model families, or auditing a stratified subsample against human A. Construct Validity analysts, would establish how much of the signal is independent The primary construct, runtime exploitability, is an operof the model that proposed it. ationalization rather than a perfect measure of real-world Finally, the SPDF rate treats every confirmed exploitation security impact. Because the sandbox disabled network access as equivalent, yet the confirmed categories differ markedly and provided no external services, privileged contexts, or in consequence. The three most frequently confirmed classes, vulnerability-specific dependencies, and because successful CWE-400 (Uncontrolled Resource Consumption), CWE-338, verification depends on the agent generating an effective and CWE-330, are dominated by denial-of-service and weak- exploit harness, some NOT_TRIGGERED verdicts and the randomness weaknesses, which are among the easiest to 46 INCONCLUSIVE outcomes may reflect environmental demonstrate at runtime and, in many deployment contexts, limits rather than the absence of a vulnerability; the reported among the lower in direct impact. By contrast, CWE-94 (Code SPDF rates should therefore be read as a lower bound on Injection) produced numerous candidate findings but no con- demonstrated exploitability. Two further limitations apply: the firmed exploitations. Because exploitability and consequence CWE Detection Agent relies on LLM-based reasoning and are conflated within a single rate, the headline figure may may produce both false positives and false negatives, and all overstate the practical severity of the residual risk, motivating verdicts are generated automatically rather than validated by severity-weighted formulations that do not silently equate human security analysts, so individual outcomes may diverge from expert judgment. exploitable with consequential. B. Implications For security evaluation pipelines, the results caution against treating static analysis as the primary measure of risk in AIgenerated code, as is common practice. Static analysis remains an efficient first-line filter, it eliminated 701 files here at negligible cost; however, it identifies vulnerable patterns rather than demonstrating consequences, so the strongest assessments

B. Internal Validity Several implementation choices may influence the observed rates. Both agents use gpt-5.4-nano; a different model could change detection, recognition, and exploit-generation behaviour, and, because the same model both proposes and confirms weaknesses, a portion of the confirmation signal may be self-reinforcing, as discussed in Section V-A. In addition, the Docker sandbox’s execution and resource limits, while

improving safety and reproducibility, may suppress exploits that require system privileges or external conditions. Finally, majority voting over five runs reduces but does not eliminate stochastic variability; alternative run counts or aggregation strategies may yield slightly different estimates. C. External Validity The study examines only Python code from SecurityEval, RedCode, and CyberNative. Vulnerability patterns, staticanalysis effectiveness, and exploitability may differ for production systems and for languages such as C/C++, Java, and Rust, so the reported rates should not be read as universal static-analysis failure rates. The static-pass set also depends on Bandit v1.8.4.dev21 and Semgrep v1.155.0 under the auto ruleset with no additional filtering; alternative versions, rulesets, or commercial scanners would shift it. Because these corpora emphasize security-relevant tasks, SPDF prevalence in generalpurpose development may differ. VII. C ONCLUSION AND F UTURE W ORK This work examined whether code producing no findings under a composite Bandit–Semgrep static-analysis gate can nevertheless exhibit dynamically demonstrable vulnerabilities. Through the Static-Pass Dynamic-Fail (SPDF) framework, we found that static-analysis outcomes under this evaluated configuration and demonstrated runtime exploitability capture distinct forms of security evidence. More than one in eight samples producing no findings under the evaluated staticanalysis gate were subsequently confirmed exploitable under our conservative dynamic-verification criterion, and several CWE categories absent from the Bandit–Semgrep findings were nonetheless confirmed at runtime. Because dynamic verification was restricted to LLM-identified candidates, these values characterize the yield of the evaluated pipeline rather than the total prevalence of exploitable vulnerabilities among all static-pass samples. These results support treating software security assessment as an accumulation of evidence across complementary forms of analysis rather than relying on any single evaluated technique in isolation. Alongside these findings, we release an agentic exploitability-evaluation framework, a dataset of dynamically evaluated vulnerability candidates, and a reproducible methodology for relating static-analysis outcomes to runtime exploitability. Two directions follow most directly. First, pairing a detector and verifier from different model families would test how much of the confirmed-exploitability signal is independent of the model that proposed it. Second, weighting confirmed instances by severity would distinguish exploitable weaknesses from consequential ones, refining the SPDF rate as a measure of practical risk. ACKNOWLEDGEMENTS The authors acknowledge support from the Faculty of Science Undergraduate Research Opportunities program at Toronto Metropolitian University which made this work possible. OpenAI ChatGPT and Grammarly were used to assist

with language editing and manuscript refinement. All technical content, experimental design, analysis, and final manuscript decisions were reviewed and validated by the authors. R EFERENCES [1] A. Fan, B. Gokkaya, M. Harman, M. Lyubarskiy, S. Sengupta, S. Yoo, and J. M. Zhang, “Large language models for software engineering: Survey and open problems,” in 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE), 2023, pp. 31–53. [2] J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat et al., “Gpt-4 technical report,” arXiv preprint arXiv:2303.08774, 2023. [3] C. E. Jimenez, J. Yang, A. Wettig, S. Yao, K. Pei, O. Press, and K. Narasimhan, “Swe-bench: Can language models resolve real-world github issues?” in International Conference on Learning Representations, vol. 2024, 2024, pp. 54 107–54 157. [4] H. Pearce, B. Ahmad, B. Tan, B. Dolan-Gavitt, and R. Karri, “Asleep at the keyboard? assessing the security of github copilot’s code contributions,” Commun. ACM, vol. 68, no. 2, p. 96–105, Jan. 2025. [Online]. Available: https://doi.org/10.1145/3610721 [5] R. Khoury, A. R. Avila, J. Brunelle, and B. M. Camara, “How secure is code generated by chatgpt?” arXiv preprint arXiv:2304.09655, 2023. [6] M. L. Siddiq and J. C. S. Santos, “Securityeval dataset: mining vulnerability examples to evaluate machine learning-based code generation techniques,” in Proceedings of the 1st International Workshop on Mining Software Repositories Applications for Privacy and Security, ser. MSR4P&S 2022. Association for Computing Machinery, 2022, p. 29–33. [Online]. Available: https://doi.org/10.1145/3549035.3561184 [7] A. Nunez, N. T. Islam, S. K. Jha, and P. Najafirad, “Autosafecoder: A multi-agent framework for securing llm code generation through static analysis and fuzz testing,” arXiv preprint arXiv:2409.10737, 2024. [8] C. Guo, X. Liu, C. Xie, A. Zhou, Y. Zeng, Z. Lin, D. Song, and B. Li, “Redcode: risky code execution and generation benchmark for code agents,” in Proceedings of the 38th International Conference on Neural Information Processing Systems, ser. NIPS ’24. Curran Associates Inc., 2024. [9] W. Wang, W. Ma, Q. Hu, Y. Zhang, J. Sun, B. Wu, Y. Liu, G. Xu, and L. Jiang, “Vulnrepaireval: An exploit-based evaluation framework for assessing large language model vulnerability repair capabilities,” arXiv preprint arXiv:2509.03331, 2025. [10] B. Chess and G. McGraw, “Static analysis for security,” IEEE Security & Privacy, vol. 2, no. 6, pp. 76–79, 2004. [11] PyCQA, “Bandit documentation,” https://bandit.readthedocs.io/en/latest/, 2026. [12] Semgrep, Inc., “Semgrep documentation,” https://semgrep.dev/docs/, 2026. [13] CyberNative, “Code vulnerability security dpo dataset,” https://huggingf ace.co/datasets/CyberNative/Code Vulnerability Security DPO, 2024. [14] Pourleyli et al., “Beyond static guarantees: Measuring the static-pass dynamic-fail gap in security-sensitive and llm-generated python code,” Jun. 2026. [Online]. Available: https://doi.org/10.5281/zenodo.20761891 [15] G. Bennett, T. Hall, E. Winter, and S. Counsell, “Semgrep*: Improving the limited performance of static application security testing (sast) tools,” in Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering, ser. EASE ’24. Association for Computing Machinery, 2024, p. 614–623. [Online]. Available: https://doi.org/10.1145/3661167.3661262 [16] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. R. Narasimhan, and Y. Cao, “React: Synergizing reasoning and acting in language models,” in The Eleventh International Conference on Learning Representations, ICLR 2023, Kigali, Rwanda, May 1-5, 2023. OpenReview.net, 2023. [Online]. Available: https://openreview.net/forum?id=WE vluYUL-X [17] J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. Narasimhan, and O. Press, “Swe-agent: agent-computer interfaces enable automated software engineering,” in Proceedings of the 38th International Conference on Neural Information Processing Systems, ser. NIPS ’24. Curran Associates Inc., 2024. [18] National Institute of Standards and Technology (NIST), “Nvd api documentation,” https://nvd.nist.gov/developers, 2026.

Record · ID 673433 · SHA-256 1c3e3249b7747519
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.