Conceptio › Archive › arXiv CS
arXiv CSopen access

The History Is the Detector: Executing CVE Patch History, End-to-End

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

The History Is the Detector: Executing CVE Patch History, End-to-End

arXiv:2609.05335v1 [cs.CR] 4 Sep 2026

Qiushi Wu Kevin Eykholt Youngja Park Xiaokui Shu Dhilung Kirat Douglas Lee Schales Ian Molloy IBM Research

Abstract

severity, a CWE label identifies the underlying weakness class, and a fixing commit usually shows the code-level changes that mitigated the risk [17, 50]. For a human reviewer, these records explain what went wrong. However, they are of little use to automated scanners because the lessons embedded in a patch do not translate into detectors capable of searching the next codebase.

Public vulnerability databases collect rich information about known software flaws, including their weakness types, affected components, and related patches. Additionally, fixing commits provide the exact code changes that removed these flaws. While these records capture why the original code was unsafe, they are documented mainly for human inspection rather than automated reuse. Consequently, the same unsafe conditions may still exist elsewhere in code with no known advisory, leaving much of this detection knowledge unused. We present B UGSTONE -E2E, a framework that automatically transforms vulnerability history into executable detection rules and validates their findings. First, B UGSTONE -E2E mines reusable rules from verified fixing commits, capturing scan anchors, fix semantics, and CVE provenance and organize them by CWE and language. Second, detection follows a funnel-shaped pipeline; early stages process a large pool of candidates using lightweight analysis, while later stages apply increasingly capable and expensive models to a shrinking set of targets. Specifically, B UGSTONE -E2E first enumerates call sites matching rule anchors using Tree-sitter, then removes benign sites using lightweight heuristics without LLM calls. Next, LLM-based agents inspect the remaining candidates guided by the rule. Following this inspection, the system retriages surviving candidates and builds runtime verifications, and finally generates scope-checked patches that are validated via two-sided differential tests. Using 19,325 high-severity CVEs from 2022 to 2026, B UGSTONE -E2E identifies 2,710 fixing commits and constructs 1,033 detection rules across 56 CWE families, packaged into 172 skills. When applied across 14 programs, it produced runtime evidence for 644 findings. These results demonstrate that CVE history can be turned into an executable workflow, transforming past vulnerabilities to reproducible detection and repair.

1

The missed opportunity. Many vulnerability patterns recur across projects and languages. For example, a commandinjection fix may reveal that attacker-controlled data reaches a shell sink without proper quoting. Although API names and code structure may differ, the underlying security condition often remains the same. Prior work detects recurring vulnerabilities using known vulnerable code or patch-enhanced signatures [35, 82], while recent systems combine LLMs with retrieved vulnerability knowledge or program-analysis context [18, 38, 44]. VulGenie further extracts API security rules from Java security patches [14]. However, what remains missing is a general pipeline that transforms CVE fixes across various weakness classes and languages into reusable, executable rules, applies them at scale, and validates and repairs the findings they produce. Key challenges. Turning these public vulnerability records into executable detectors raises three primary challenges. First, CVE records and patches vary widely in content, making it difficult to extract the precise components that can be reused for detection. The relevant security fixes may be buried in unrelated edits, poorly described, or tied to project-specific logic. A mining tool must therefore identify what made the original code unsafe and generalize that condition into anchors that can match the same flaw elsewhere. Second, this knowledge must be represented as an executable rule with sufficient context to distinguish vulnerable patterns from benign cases. For instance, a simple match on system, pickle.loads, or memcpy is too broad, whereas an overly specific signature risks missing the same vulnerability hidden behind a different wrapper or code structure. Third, these rules must scale efficiently to handle large codebases without incurring expensive

Introduction

Modern vulnerability response leaves a detailed public record. Specifically, a CVE entry describes the affected behavior and 1

analysis on every candidate. Static analyzers such as CodeQL and Semgrep scale well once suitable queries or rules exist [28, 62], but creating and maintaining such rules still requires manual efforts. From detection to verified remediation. Solving the three challenges above remains insufficient for end-to-end vulnerability analysis. While a rule violation or plausible taint path identifies a potential vulnerability, it does not prove that the behavior is reachable or exploitable at runtime. A complete workflow therefore requires dynamic validation that produces concrete runtime evidence. Furthermore, it must support remediation by generating a feasible patch from the vulnerability evidence and program context, and then verifying that the patch removes the demonstrated behavior. Our approach. B UGSTONE -E2E addresses this gap through a staged pipeline spanning four core phases (Phase A–D) that transitions seamlessly from reusable vulnerability knowledge to automated detection, runtime validation, and verified remediation. The system first mines known vulnerabilities from the cvelistV5 dataset, verifies their fixing commits, and synthesizes reusable rules containing scan_apis, buggy and fixed code patterns, a verification contract, and related source CVEs. The rules are grouped by CWE and language into modular bug detection skills. For a target project, Phase A first indexes the codebase with Tree-sitter [69] and enumerates all call sites match rule anchors without using an LLM. Phase A then removes clearly benign candidates via deterministic checks. The remaining candidates are analyzed independently in Phase B, where an LLM-based agent uses the rule-specific verification contract to inspect taint flows, callers, guards, and sanitizers, finally returning a structured verdict accompanied by concrete code evidence. This funnel reduces the number of candidates before applying more expensive analysis and allows Phase B to run in parallel with cheaper models. However, static evidence alone is insufficient for full vulnerability confirmation. After Phase B filters out most candidates, Phase C re-triages the small set of survivors and attempts to reproduce each one in an isolated, concrete execution environment. A finding gains definite runtime validation only when an expected runtime marker–such as a sanitizer report, crash, or controlled-sink signal–is observed, and only the subset with a working live proof-of-concept (PoC) is rated as CONFIRMED_EXPLOITABLE, separating non-exploitable flaws from those that could not be tested. Finally, Phase D generates a minimal, scope-checked patch and verifies it against the runtime evidence produced in Phase C. The exploit must fail on the patched program. It then reverts the patch and confirms that the exploit succeeds again, preventing a broken build or environment from being accepted as a fix. Exploit- and test-based validation is increasingly used in recent literature to evaluate vulnerability repair and patch correctness [9, 26, 75]. However, unlike these existing settings, B UGSTONE -E2E begins with newly discovered

findings that have no predefined PoC or validation environment. It must therefore construct the runtime evidence by itself and independently verify each critical artifact before accepting a patch. Evidence. In this work, B UGSTONE -E2E surveys 19,325 high-severity CVEs from 2022 to 2026, enriches 5,902 unique open-source CVEs, and verifies 2,710 fixing commits. From these commits, it constructs 1,033 rules across 56 CWE families, packaged into 172 production bug detection skills. Across scans of 14 real-world projects, Phase B reports 2,933 deduplicated findings. Finally, Phase C successfully obtains runtime evidence for 644 of these findings, including AddressSanitizer (ASan [63]) violations and crashes. This paper makes the following contributions: • Vulnerability Rule Mining from CVE history: We design a general pipeline that surveys high-severity CVEs, verifies fixing commits, extracts each patch’s vulnerable sink and fix semantics, and rejects cases that do not generalize. • Cost-ordered Detection Rule Execution: Our CWE and language-specific rules carry scan anchors, buggy and fixed patterns, a verification contract, and source-CVE provenance, all packaged into modular deployable skills. This architecture leverages deterministic indexing and lightweight filters to absorb scale, allowing LLM-based agents to apply expensive semantic judgment only where a rule points. • Oracle-driven End-to-End Workflow: Our system executes a batched re-triage and an environment-first exploit stage that separates “not exploitable” from “could not be tested”, and judges patches using a two-sided differential testing under a machine-checked output contract. We report empirical evidence for every pipeline phase and state what each result does not establish. • Deterministic Decomposition for Scalability: In B UGSTONE -E2E, anchor enumeration and the application of a frozen filter are deterministic. Thus, we can split a very large candidate pool into bounded, independent batches grouped by rule and source file. Since the batches share no mutable execution state, the expensive LLM-based agentic stage can parallelize freely–even though candidates within a single batch are not statistically independent–resulting in high scalability.

2 2.1

Background and Problem Motivation

Vulnerability patterns recur across projects. Known vulnerability patterns often reappear in other code locations or projects. Early systems, such as ReDeBug and VUDDY, demonstrated this problem by finding unpatched clones of known vulnerable codebases at large scale [31, 35]. Later work generalized beyond exact clones by deriving vulnerability and patch signatures that identify recurring vulnerabil2

2.2 Challenges in Precise and Scalable Vulnerability Detection

Table 1: Representative high-recurrence patterns from our CVE background study, spanning ten of the twelve languages B UGSTONE -E2E covers. #CVEs counts distinct CVEs whose patches support the pattern, and the same weakness class recurs across languages. CWE

Lang.

CWE-79 CWE-502 CWE-78 CWE-79 CWE-639 CWE-79 CWE-22 CWE-122 CWE-122 CWE-22 CWE-78

PHP PHP Python JS Go TypeScript Java C C++ C# Ruby

#CVEs 27 16 13 9 7 7 5 4 4 4 2

Traditional static analysis. Static analyzers such as CodeQL [28] and Semgrep [62] can efficiently scan large codebases once suitable queries or rules are available. However, encoding a new vulnerability pattern still requires analyst effort, which limits how quickly these tools can absorb the broad and evolving set of patterns recorded in CVE history. Their findings also often require additional context to distinguish true vulnerabilities from benign uses, and false positives remain a major barrier to developer adoption [33]. Recent work therefore combines static analysis with semantic reasoning to improve the triage of candidate findings [44]. Static analysis provides scalable candidate discovery, but extending its coverage and resolving ambiguous findings still require substantial manual or semantic analysis.

Repr. API echo unserialize os.system innerHTML Decode src getNextEntry memcpy memcpy Path.Combine system

Dynamic analysis and fuzzing. Dynamic techniques provide stronger evidence because they observe concrete program behavior. Tools such as ASan [63] detect memory errors during execution, while fuzzers such as AFL++ [21] explore program behavior through generated inputs. Their effectiveness, however, depends on both reaching the vulnerable behavior and having an oracle that recognizes it. Memory errors often produce crashes or sanitizer violations, but many other vulnerabilities do not. For command injection, path traversal, deserialization, and similar weaknesses, successful execution may appear normal unless the analysis encodes the relevant security condition. Dynamic validation is therefore valuable for confirming vulnerabilities, but it is difficult to use as the primary mechanism for broad vulnerability discovery.

ities despite code changes [30, 77, 82]. More recently, Wu et al. [81] showed that identical security-relevant usage patterns often recur many times within a single software project. Starting with 135 rules, their study flagged 22,568 potential violations in the Linux kernel, where a subsequent manual review of 400 cases confirmed 246 true vulnerabilities. Cross-language recurrence. Prior evidence has focused largely on code reuse and recurring patterns within a particular language or software ecosystem. To examine whether the same property holds more broadly, we study public CVE fixing commits across weakness classes and programming languages. Table 1 reports representative high-recurrence patterns for different languages. We observe that many patterns are supported by multiple independent CVEs, rather than a single project-specific fix. We also find that the same weakness class can recur across languages through similar securityrelevant operations. For example, cross-site scripting (CWE79) appears across PHP, JavaScript, and TypeScript, while command injection (CWE-78) affects Python and Ruby.

LLM and agent-based detection. Recent work uses LLMs to provide semantic context that is difficult to encode in fixed rules. VulRAG [18] retrieves vulnerability knowledge to guide analysis, IRIS [44] and LLift [40] combine LLM reasoning with static-analysis results, and LLMxCPG [38] uses code property graphs to provide structured program context. More recent systems employ agentic workflows that can inspect code, reason about attack surfaces, and validate candidate vulnerabilities [5, 29, 36, 57, 72, 86]. These systems improve semantic reasoning but introduce a different scaling problem. Repository-wide exploration gives the agent a large search space and little guidance about where expensive reasoning is most useful. Deeper analysis of one region consumes budget that could otherwise cover additional code, and the result depends strongly on the capability of the underlying model. Moreover, an individual scan usually produces findings rather than reusable detection knowledge, so similar reasoning may need to be repeated on the next target. This motivates a different design: use persistent vulnerability knowledge to identify where to look, apply inexpensive analysis broadly, and reserve expensive semantic and runtime reasoning for a progressively smaller set of candidates.

CWE-78-Py-R1, our running example, is supported by thirteen CVEs from different Python projects. The common pattern is that externally derived data reaches a shell API without sufficient quoting or sanitization. In one instance from PaddleOCR, a path selected through QFileDialog is used to derive an output directory and then concatenated directly into an os.system command. Although the surrounding application logic differs across projects, the security-relevant condition is the same. This recurrence motivates extracting such conditions from CVE fixes and reusing them to guide detection in previously unseen code. 3

2.3

Research Scope

quirements only to the surviving candidates from Phase B. Phase C first re-triages these findings in batches and then attempts to construct a concrete execution environment for runtime validation. It records runtime evidence such as sink reachability, sanitizer violations, or crashes, while only findings that satisfy the stronger exploit-confirmation criterion are marked CONFIRMED_EXPLOITABLE. Failure to construct or exercise an environment remains distinct from evidence that a finding is not exploitable. This ordering keeps expensive runtime analysis focused on the small set of candidates that survive earlier stages. For findings with sufficient runtime evidence, Phase D generates a minimal, scope-checked patch using the vulnerability evidence and program context. It then applies a two-sided differential test: the proof of concept must fail on the patched tree and succeed again after the patch is reverted. This prevents a broken build or invalid environment from being accepted as a successful repair. Finally, B UGSTONE -E2E follows the project documentation to produce a vulnerability report in the maintainer’s required format, together with supporting evidence and a proposed patch.

B UGSTONE -E2E targets vulnerabilities that follow recurring API-misuse patterns, where the security-relevant behavior is visible at or near a call site. Its scope does not include one-off design flaws, configuration errors, or vulnerabilities that require whole-program symbolic reasoning to identify. B UGSTONE -E2E produces candidate findings with structured static and runtime evidence, while human-in-the-loop disclosure decisions and patch acceptance remain separate processes.

3

System Overview

Figure 1 summarizes B UGSTONE -E2E in three parts. Part 1 converts public vulnerability history into reusable detection skills. Part 2 applies those skills to a target project and identifies candidates that remain plausible after static analysis. Part 3 raises the evidence requirement through runtime validation and verified remediation. The pipeline therefore moves from historical vulnerability knowledge to increasingly stronger evidence while applying expensive analysis to progressively fewer candidates. Part 1: From CVE History to Detection Skills. B UGSTONE -E2E treats CVE fixing commits as rich sources of reusable detection knowledge rather than project-specific repairs. It surveys 19,325 high-severity CVEs and verifies 2,710 fixing commits from open-source projects. From these commits, B UGSTONE -E2E extracts reusable scan anchors, buggy and fixed patterns, verification contracts, and CVE provenance, while rejecting patterns that fail to generalize. It then consolidates related CWE categories to group rules by CWE family and language. The resulting knowledge base contains 1,033 production rules across 56 CWE families, packaged into 172 deployable bug detection skills. Part 2: Rule-Driven Detection. Given a target repository, Phase A enumerates rule-anchored call sites, and Phase A.1 removes false positives using lightweight filtering, leaving a reduced candidate set for semantic verification. However, this stage can still produce hundreds to thousands of candidates for medium-to-large projects. Phase B then decomposes this workload into independent tasks, assigning one agent to each candidate together with the corresponding rule. Instead of exploring or reasoning about the repository as a whole, each agent only needs to determine whether that specific candidate violates the rule and constitutes the reported bug pattern. This narrow task scope substantially lowers the required agent capability, enabling effective verification using smaller, more cost-efficient models. Because these candidates are independent, Phase B can also process them with high parallelism, scaling semantic analysis to large candidate sets without requiring a shared state. Part 3: Validation and Remediation. The output of Part 2 is still a static claim, so Part 3 applies stronger evidence re-

4

Design

B UGSTONE -E2E is designed around three problems that arise when turning vulnerability history into an end-to-end detection workflow. First, public CVE records and fixing commits must be converted into reusable detection knowledge without overfitting to one patch or learning the fix rather than the vulnerability. Second, that knowledge must guide semantic analysis at repository scale without requiring expensive model reasoning over the entire codebase. Third, findings and generated patches must be supported by evidence that does not depend on the model’s own judgment. These problems motivate the three research challenges below. The following subsections then describe how B UGSTONE -E2E addresses them.

4.1

Research Challenges

RC1 (Turning vulnerability history into reusable detection knowledge). A CVE record is not directly usable as a detection specification. Vulnerability descriptions are often brief, fixing commits may not be linked explicitly, and patches can mix the security fix with unrelated changes. Even when the correct fixing commit is known, the changed line does not necessarily identify the underlying vulnerability condition. For example, a patch may place a check, sanitizer, or guard away from the actual vulnerable operation. Directly translating such diffs into rules can therefore produce detectors for the fix rather than for the vulnerable code. Additionally, generating one rule per CVE results in highly redundant, project-specific rules, whereas merging them aggressively can combine distinct root causes. The challenge is to isolate 4

Part 1: CVE Intelligence and Skill Development CVE Mining

Survey CVEs 19,325 high-sev.

Rule & Skill Synthesis

Identify + Enrich OSS

Clone + Date Filter

Light LLM Screen

Verify Fix Commit

public repositories

commit window

pre-filter patches

2,710 confirmed

Rule Synthesis

Skill Synthesis

extract + generalize

1,033 rules 172 skills

skills + rules

Host Agent

Part 2: Rule-Driven Detection

tree-sitter Index Phase A

Guest VM

Filter.py Phase A.1

Part 3: Validation & Remediation

LLM Agent Verify

Triage + Exploit

Patch + Verify

Report + Package

Phase B

Phase C

Phase D

JSON + HTML

optimize filter (1–2 rounds)

re-run PoC

Figure 1: Overview of B UGSTONE -E2E. and validate the relevant fix, identify the security-relevant condition, and generalize it only when supported by multiple instances.

ter the patch, but the same input succeeds when the patch is reverted. The challenge is therefore to raise the burden of proof from structured static evidence, to runtime evidence, and finally to verified remediation while applying the more expensive checks only to the progressively smaller set of surviving findings, none of which is taken on the reporting agent’s word.

RC2 (Scaling semantic analysis through task decomposition). Applying an extensive rule set to a large codebase can produce an enormous amount of semantic analysis. While rules provide explicit candidate locations, determining whether a candidate is truly vulnerable may require inspecting surrounding data flows, guards, sanitization, and projectspecific conventions. Simply assigning one agent to each rule does not solve this problem. A single rule can match hundreds or thousands of sites, overwhelming the agent with a large context and a long sequence of rule-specific judgments. Such tasks place substantial demands on both model capability and context handling. Therefore, the challenge is to minimize the number of candidates that reach semantic analysis and to break down the remaining work into small, independent reasoning tasks.

4.2

RC3 (Establishing evidence beyond model judgment). A model’s conclusion that a candidate violates a rule is merely a static claim, not runtime evidence of an acutal vulnerability. Likewise, a patch generated by the same model cannot be considered correct simply because the model reports that it works. Both stages require evidence that can be checked independently of the model. For vulnerability validation, the execution must distinguish observed runtime evidence, stronger exploit confirmation, unsuccessful tests, and cases where the target cannot be exercised. For remediation, verification must demonstrate that the triggering input no longer succeeds af-

CVE-to-Skill Pipeline

Stage

Artifact

Count

CVE survey

High-severity CVEs

19,325

OSS enrichment

Open-source CVEs

5,902

Fix recovery

Verified fixing commits

2,710

Case validation

Validated CVE cases

2,662

Rule synthesis

Pre-consolidation rules / CWEs

1,757 / 266

Family consolidation

Production rules / CWE families

1,033 / 56

Skill synthesis

Deployable bug detection skills

172

Table 2: CVE-to-skill pipeline and the number of artifacts retained at each stage. To address RC1, B UGSTONE -E2E turns incomplete CVE records into deployable detection skills in three steps: it enriches CVE metadata and recovers verified fixing commits, extracts reusable vulnerability conditions from those fixes, and consolidates related rules into CWE-family–language skills. Table 2 summarizes this progression. 5

CVE enrichment and fix recovery. CVE metadata is often insufficient to identify either the affected code or its fix. A record may contain only a short vulnerability description, omit the affected repository or fixing commit, or reference an issue, advisory, or third-party page rather than the relevant code change. CWE metadata can also be missing, ambiguous, or contain several weakness categories in one vulnerability. B UGSTONE -E2E therefore enriches each record before rule extraction. Starting from 19,325 high-severity CVEs published between 2022 and 2026, it combines structured CVE fields, natural-language descriptions, referenced pages, and public repository information to resolve the affected opensource project and primary weakness category. When a reliable fixing-commit reference is available, B UGSTONE -E2E verifies the cited commit directly. Otherwise, it uses the vulnerability reporting and publication window to guide an agentic search over repository history. Candidate commits are checked against the CVE description, patch, and surrounding code to determine whether they actually fix the reported vulnerability. This process yields 2,710 verified fixing commits, which form the evidence base for rule generation. Patch-guided rule extraction. A verified fixing commit is still not a detection rule. A patch may mix the security fix with unrelated changes, and the code added by the fix may be a guard or sanitizer rather than the operation that was originally unsafe. The goal is therefore to recover the underlying vulnerability condition rather than reproduce the textual shape of the patch. For each verified fix, B UGSTONE -E2E jointly analyzes the CVE context, the vulnerable and fixed code, and the surrounding functions. From this evidence, it identifies the vulnerable operation, candidate scan anchors, the condition that makes the operation unsafe, and the change that removes that condition. The resulting rule describes what should be detected in vulnerable code rather than what was introduced by the fix. A validation step checks this direction explicitly: scan anchors must refer to operations present before the fix and must not point to a sanitizer, guard, or replacement API added by the patch. Rules that cannot be supported by the available evidence or generalized beyond the individual patch are not promoted to the production rule set. Each accepted rule retains its source CVEs so that later findings remain traceable to the vulnerability evidence behind the rule. This process produces 1,757 pre-consolidation rules covering 266 CWE categories. These numbers describe the mined rule set before related weakness categories and overlapping rules are consolidated. Rule consolidation and skill synthesis. CWE categories do not map one-to-one to distinct detection problems. Related CWEs often describe different manifestations of the same underlying condition and consequently share scan anchors and semantic checks. Keeping them independent would repeatedly enumerate and analyze the same code sites. B UGSTONE -E2E therefore groups closely related CWE

categories into broader CWE families and consolidates overlapping rules within each family. The resulting production knowledge base contains 1,033 rules across 56 CWE families, reduced from 1,757 rules across 266 CWE categories. This consolidation removes redundant detection work while preserving rules for distinct vulnerability conditions. The consolidated rules are organized by CWE family and target language into 172 deployable bug detection skills. A skill is not a single rule; it is an execution template for one CWE-family–language combination and may contain multiple rules covering different APIs or variants of the same vulnerability family. Its rule file stores rule-specific scan anchors and vulnerability conditions, while a skill file (SKILL.md) defines the semantic checks and evidence expected from Phase B. Each skill also contains a filtering script for inexpensive deterministic filtering. This filter may initially be empty and, when candidate volume justifies it, can be refined from Phase B feedback during large-scale scans as described in §4.3. The CVE-derived rules themselves remain unchanged. For example, the Python command-injection family is supported by multiple CVEs, where externally derived data reaches a shell-execution API without sufficient quoting or sanitization. Although the individual patches differ, they support the same reusable condition. The corresponding skill therefore anchors on the relevant shell APIs and asks downstream analysis whether attacker-controlled data can reach them without an adequate guard.

4.3

Layered Target Scanner

Once bug detection skills are constructed, the remaining challenge is to apply them to a large repository without turning every rule match into an expensive semantic or runtime analysis. B UGSTONE -E2E therefore uses a layered scanner that raises both cost and burden of proof as the candidate set shrinks. Phase A performs broad build-free enumeration, Phase A.1 reduces redundant or repeatedly benign candidates, Phase B performs bounded semantic verification, Phase C establishes runtime evidence and exploitability, and Phase D verifies remediation. Build-free enumeration and conservative selfimprovement (RC2). In Phase A, B UGSTONE -E2E parses source files with Tree-sitter [69] and constructs a shared index of call sites that is reused across all applicable skills. Each rule contributes explicit scan anchors, and Phase A enumerates matching call sites without invoking an LLM. For a small target, or for a rule that matches only a modest number of sites, these candidates can proceed directly to semantic verification. However, large repositories create a different problem. Common APIs such as memory allocation, deallocation, or string operations can appear thousands of times, making semantic inspection of every matching site unnecessarily expensive. When this occurs, B UGSTONE -E2E can enable an iter6

ative filter-synthesis loop (Phase A.1). A new bug detection skill may begin with an empty filter. Phase B then examines a bounded sample of candidates and identifies false positives. From these results, an agent summarizes recurring benign patterns that can be expressed using inexpensive local syntax or AST checks and updates the filter conservatively. The filter is then applied to subsequent candidates, whose Phase B results can in turn refine it further. Thus, semantic judgments from early candidates are distilled into cheap checks that reduce the workload for later candidates. The filter is deliberately asymmetric relative to the CVEderived rule. It only removes patterns judged safe by deterministic checks and does not modify the rule’s vulnerability condition or scan anchors. If a filter cannot determine that a candidate is benign, the candidate is retained. The refined filter is stored with the skill and can be reused in later scans. Phase A.1 also removes duplicate work that arises only after rules are applied to a target. When overlapping rules within the same CWE family identify the same source location, the scanner keeps a representative candidate rather than sending multiple equivalent tasks to Phase B. Note that Phase A.1 reduces redundant work in a particular target, while rule consolidation in §4.2 reduces redundancy in the knowledge base. Bounded semantic verification (RC2). The remaining candidates require semantic reasoning because safety may depend on data flow, surrounding guards, sanitization, callers, or project-specific conventions. Assigning all candidates for one skill to a single agent would still create a large reasoning task: one rule may match hundreds or thousands of locations, forcing the agent to process extensive context and make many independent judgments in one session. B UGSTONE -E2E instead decomposes verification into small, bounded tasks. Candidates are grouped only when they match the same rule in the same source file, allowing them to reuse rule instructions and source context without turning verification back into a repository-scale task. The maximum number of candidates assigned to one worker is configurable and can be adjusted to the capability of the Phase B model. Each worker receives the corresponding skill and candidate locations and may inspect nearby code, callers, data flow, guards, and sanitizers as needed. It returns a verdict together with structured rule-specific evidence, such as the matched sink, relevant data flow, and guard or sanitizer analysis. Because these tasks are small and independent, they can be processed in parallel while placing substantially lower contextual and reasoning demands on each model invocation. Runtime validation and exploit confirmation (RC3). A Phase B finding provides structured static evidence, but it does not establish that the reported condition can be exercised at runtime. Phase C therefore moves each surviving finding into an isolated dynamic validation workflow. The first step constructs a reusable execution environment for the target. A LLM-based agent runs inside a dedicated VM or container, installs the required dependencies, builds the

project, and enables sanitizers when applicable. Because this setup is shared by all later validation tasks for the same target, B UGSTONE -E2E performs it once and records a compact environment report describing how the target was built and exercised. Each Phase B finding is then validated independently by a Phase C agent using this prepared environment and its build report. The agent focuses on one candidate at a time, constructs an input or execution path that reaches the reported operation, and looks for a vulnerability-specific runtime oracle, such as an ASan violation, crash, controlled sink reachability, or other observable security-relevant behavior. Dynamic validation is executed sequentially rather than concurrently to avoid interference between concurrent tests. After each candidate, B UGSTONE -E2E saves the resulting evidence and restores the environment to its pre-test state before validating the next finding. Runtime evidence is recorded separately from stronger exploit confirmation, because observing the reported behavior does not, by itself, establish a working exploit. Thus, Phase C distinguishes among four categories: findings with runtime evidence, exploit-confirmed findings, findings that could not be proved exploitable, and cases that could not be tested due to an insufficient execution environment. Differentially verified remediation (RC3). For findings with a working proof of concept, Phase D attempts to produce a minimal, scope-checked patch using the vulnerability report and runtime evidence. Patch generation alone is insufficient because the model that writes a patch cannot certify its own fix. A one-sided test is also weak, because the proof of concept may stop working because the patch broke the build or disabled the relevant execution path rather than because it removed the vulnerability. B UGSTONE -E2E therefore uses a two-sided differential test. It first applies the patch and reruns the same proof of concept, which must no longer reproduce the vulnerability. It then reverts the patch and runs the proof of concept again, which must reproduce the original behavior. Only patches that satisfy both directions are marked as verified. The final output packages the maintainer-format vulnerability report together with the proof of concept, runtime evidence, proposed patch, and verification artifacts. This completes the progression from structured static evidence, to runtime evidence and exploit confirmation, and finally to verified remediation.

5

Evaluation

We evaluate B UGSTONE -E2E along the main design questions introduced in §4. We first examine how effectively the CVE-to-skill pipeline converts public vulnerability history into deployable detection knowledge. We then evaluate the target scanner through controlled comparisons, reproducibility experiments, and cost measurements. Real-world findings and runtime validation results are reported separately in §6. 7

5.1

Experimental Setup

come from repository-history recovery and would not be available from direct commit references alone. The remaining 3,192 resolved records did not yield a verified fixing commit, which illustrates why this recovery step is difficult. In some, repository search finds no commit that can be confidently linked to the vulnerability; in others, the relevant time window contains either no useful commit history or too many plausible changes to verify reliably. Thus, agent-assisted enrichment expands the usable evidence base, but does not replace explicit patch references. From fixing commits to reusable rules. Verification further reduces the corpus because not every security patch provides reusable detection knowledge. Some patches contain mostly unrelated refactoring, lack a clear vulnerable operation, or describe conditions that are too project-specific to generalize. After validation, 2,662 CVE cases provide usable rule evidence, from which B UGSTONE -E2E synthesizes 1,757 rules across 266 CWE categories. Each accepted rule retains links to its source CVEs, preserving provenance through later consolidation and scanning. The raw rule set is intentionally broader than the deployed knowledge base. Closely related CWE categories frequently encode overlapping detection conditions and anchors, which would otherwise cause the same target site to be scanned repeatedly. Family-level consolidation therefore merges related CWE categories and redundant rules, reducing the store from 1,757 rules over 266 CWEs to 1,033 rules over 56 CWE families. These rules are packaged into 172 CWE-family– language skills. Coverage across languages and weakness families. The mined rules span both memory-safety and higher-level APImisuse patterns. Before consolidation, the largest languagespecific rule sets include Python, PHP, Go, JavaScript, Java, TypeScript, C, and C++. The same weakness can also appear across languages. For example, path traversal, injection, deserialization, and related families produce language-specific rules with different APIs but similar underlying security conditions. Conversely, C and C++ contribute more heavily to memory-safety families. This coverage follows the scope defined in §4.1. B UGSTONE -E2E targets recurring vulnerability conditions that can be anchored at or near security-relevant API use. Vulnerabilities that depend primarily on global protocol state, race conditions, or other API-independent logic may therefore have no corresponding rule even when they appear in the CVE corpus.

Hardware. All experiments run on a dual-socket server with two Intel Xeon Gold 6336Y processors, 96 hardware threads, and 2 TB of RAM. Phase A uses up to 96 threads for deterministic enumeration, while model-based stages are bounded by API rate limits and the concurrency configured for each experiment. Models and agent harnesses. Unless otherwise stated, Phase B uses gpt-5-mini. Controlled experiments additionally evaluate GPT-5.5, GPT-5.6-sol, Claude Sonnet 4.6, Kimi-K2.5, and other models where noted. B UGSTONE -E2E supports multiple coding-agent harnesses, including Claude Code, Codex, and Pi. For the cross-harness study in §5.4, we compare against three recent agentic security harnesses: the OpenAI/Codex security-scan harness [57], Anthropic’s open-source defending code reference harness [5], and the Visa Vulnerability Agentic Harness [72]. Evaluation targets. We use different targets for different controlled experiments. wolfSSL is rolled back to a pre-fix revision for the cross-harness comparison in §5.4. Pillow is similarly rolled back and used for the reproducibility and model studies in §5.5. The larger deployment corpus used to measure end-to-end scanning is described separately in §6. This separation prevents findings accumulated across repeated production scans from being mixed with controlled experimental results.

5.2

CVE Knowledge Construction

We first evaluate the CVE-to-skill pipeline of §4.2. Starting from 19,325 high-severity CVEs published between 2022 and 2026, the pipeline identifies 5,902 records associated with public open-source repositories and verifies 2,710 fixing commits. Patch-guided synthesis produces 1,757 preconsolidation rules spanning 266 CWE categories. Familylevel consolidation reduces this knowledge base to 1,033 production rules across 56 CWE families, organized into 172 language-specific bug detection skills. Table 2 summarizes the full progression. Recovering fixes beyond structured CVE metadata. Of the 19,325 high-severity CVEs surveyed, B UGSTONE -E2E enriches 6,448 records using structured metadata, naturallanguage descriptions, and external references. This process resolves 5,902 records to open-source projects with identifiable GitHub repositories. Among these 5,902 records, 2,439 contain a direct commit URL, from which B UGSTONE -E2E verifies 1,944 fixing commits. For records with a resolved repository but no usable commit pointer, B UGSTONE -E2E searches repository history using the CVE context and vulnerability timeline, recovering 766 additional fixing commits. Together, the direct-reference and search paths yield 2,710 verified fixing commits. Thus, 766 of the final fixes (28.3%)

5.3 Scanner Efficiency and Task Decomposition We next evaluate the two design choices behind RC2: reducing the number of candidates that require semantic analysis and bounding the semantic work assigned to each Phase B 8

worker. Candidate reduction before semantic analysis. Across 15 targets, deterministic filtering reduces the aggregate Phase A candidate pool from approximately 745K to 293K, removing 60.6% of candidates before model-based verification. The benefit is largest for high-volume rules: on FreeBSD, for example, cwe122-cpp alone produces 65.5K call sites, which filtering reduces to 42.0K. On smaller or lower-volume targets the reduction can be modest, supporting the design of Phase A.1 as an optional cost-control layer rather than a prerequisite for scanning. Bootstrapping filters from Phase B feedback. We further evaluate filter synthesis on 10 high-volume skills from four targets, using up to 100 Phase B-labeled candidates per skill (979 total). Of these candidates, 814 are labeled false positive. The first synthesis iteration produces conservative filters for seven skills and removes 261 candidates, corresponding to 32.1% of the observed false positives; for the remaining three skills, no deterministic pattern satisfying the safety criteria is found. Applied to the full candidate pool, these newly synthesized filters further reduce the workload from approximately 299K to 293K candidates. During synthesis, candidates labeled BUGGY or UNKNOWN are protected: all 165 such candidates in the synthesis samples survive the accepted filters. This is a construction-time safety check rather than an independent recall measurement, but it prevents the synthesis procedure from accepting filters that contradict available positive or uncertain evidence. Bounded semantic tasks. Filtering controls how many candidates reach semantic verification; task decomposition controls how much reasoning each worker performs. Phase B groups only small batches from the same rule and source file, avoiding repository-scale reasoning while retaining shared rule and source context. We evaluate the resulting model requirement separately in §5.5, where multiple models verify the identical set of 8,115 Pillow candidates. Together, these two mechanisms reduce both the volume and the scope of model-based analysis.

5.4

wolfSSL row of §6 quotes an earlier registry run with a smaller skill set and 3,166 candidates. The three baselines are the OpenAI/Codex security-scan harness with GPT-5.5 [57], Anthropic’s open-source defending code reference harness with Claude Opus 4.8 [5], and the Visa Vulnerability Agentic Harness (VVAH) with Claude Opus 4.8 [72]. Expanded ground truth. Recall was originally measured against nine disclosed cases from Anthropic’s Mythos coordinated-disclosure program1 , and nine points cannot separate a 40% detector from a 55% one. We therefore rebuilt the ground truth from the vendor release notes for 5.9.0, 5.9.1, and 5.9.2, the three releases published after the snapshot. Those notes name 68 CVE identifiers, of which 64 localize into the rolled-back tree by aligning each fix’s pre-image against the snapshot source, and the other four are excluded with a recorded reason. A harness finding counts as a hit when it names the same file and the same enclosing function as a localized fix hunk. We calibrated the automatic file-andfunction matcher against the existing hand adjudication of the nine Mythos cases across the seven harness configurations recorded at calibration time, which gives 63 labeled decisions, and it agrees on 59 (93.7%). Of the four disagreements, two are false positives that credit a baseline with a case its author did not identify, so the recall gap below is conservative. Recall on 64 cases. Figure 2 breaks recall down by case set, and the expansion confirms the ordering the nine cases gave rather than changing it. B UGSTONE -E2E with gpt-5-mini leads at 33 of 64 (52%), the union of its two model configurations reaches 35 (55%), the Codex harness reaches 22 (34%), the Anthropic harness 15 (23%), and VVAH 3 (5%). Recall on the 55 later-fixed CVEs tracks recall on all 64 within a few points for every system, so the expansion did not select a favorable population. The expansion also corrects the earlier result in the baselines’ favor. Specifically, the Anthropic harness and VVAH score 0 of the 9 Mythos cases yet match 15 and 3 of the 64 localized CVEs, almost all of them among the 55 later-fixed ones, so reporting only the nine Mythos cases understated both. The disclosed CVEs are a lower bound on true positives rather than the full set, since a harness can report a real bug that no vendor advisory yet covers. Each bar spans the system’s total reported findings, and the trailing count restates that volume, which the lightest segment beyond the matched cases makes visible. The two B UGSTONE -E2E models reported 103 and 70 findings, above every baseline, since the Codex harness reported 52 and neither the Anthropic harness nor VVAH exceeded 26. No harness subsumes another. 43 of the 64 cases are found by at least one system against 35 for the two B UGSTONE -E2E configurations together, so the baselines contribute eight cases B UGSTONE -E2E misses. These systems are better read as an ensemble than as competitors on a single ranking. 21 of the 64 cases are found by nobody. Four of those sit in hand-written

Cross-Harness Comparison on wolfSSL

Rolled-back target and five configurations. We chose wolfSSL, an embeddable TLS/DTLS and cryptography library, because it is heavily audited and is the target of Anthropic’s Mythos and Glasswing disclosure work, so its open findings give a ready basis for comparing B UGSTONE -E2E with independent harnesses. We ran all five configurations on it under their own default settings. The target is commit 0c4ca257a07a (wolfSSL 5.8.4), deliberately rolled back to a pre-fix state. Two are B UGSTONE -E2E under the Pi agent with gpt-5-mini and with Kimi-K2.5, and Phase A delivers exactly the same 6,099 candidates to Phase B in both. This subsection quotes that later Pi run throughout, whereas the

1 https://red.anthropic.com/2026/cvd/

9

33 (52%) matched, 103 found

Bugstone gpt-5-mini

26 (41%) matched, 70 found

Bugstone Kimi-K2.5

22 (34%) matched, 52 found

Codex GPT-5.5 Anthropic Opus 4.8

15 (23%) matched, 19 found

VVAH Opus 4.8

3 (5%) matched, 26 found

0

20

40

60

later-fixed CVEs (of 55)

FP

246 245 235 240 219 220 196 261 233 235

7,842 7,857 7,872 7,870 7,875 7,890 7,915 7,850 7,872 7,861

Unk. Final Excl. Tokens 6 13 8 5 14 5 4 4 10 19

138 143 131 131 133 135 132 147 128 125

8 9 8 4 10 5 1 4 5 1

731M 690M 696M 699M 794M 679M 765M 718M 743M 716M

Table 3: Ten independent runs on the same Pillow snapshot with an identical configuration. Raw: BUGGY judgments before collapsing by location; Unk.: UNKNOWN verdicts. run1 and run5 additionally left 21 and 7 candidates unjudged, so their rows sum to fewer than 8,115. Final: distinct BUGGY locations; Excl.: locations this run reports and no other run does.

other findings

Figure 2: Recall on the rolled-back wolfSSL 5.8.4 snapshot, over the 64 post-snapshot CVEs that localize into the snapshot tree. Each bar spans the system’s total reported findings, split into the nine hand-adjudicated Mythos ground-truth cases matched, the 55 later-fixed CVEs matched, and the remaining findings that match no disclosed CVE. The two matched segments therefore give the recall count, and the full bar gives reported volume, since the disclosed CVEs are a lower bound on true positives rather than all of them. B UGSTONE -E2E rows count Phase B BUGGY output.

21,584 with the deterministic Phase A.1 filter, leaving 9,010 candidates that dedup to an identical set of 8,115 delivered to Phase B in the same 1,345 batches. Any variation after this point therefore comes from semantic verification rather than candidate delivery. Despite identical input, individual runs report 125–147 findings (Table 3). Their union contains 243 distinct locations, while only 56 appear in all ten runs. A typical run recovers 55.3Thus, aggregate finding counts are relatively stable, but the identity of the reported findings is not.

assembly and a header that no harness here parses, so part of the residual miss set is a shared input-language limit rather than a reasoning failure. The one Mythos case no system reports is CVE-2026-5446, ARIA-GCM explicit-IV/nonce reuse, and that miss is instructive rather than incidental. Nonce reuse is a stateful cryptographic invariant spanning successive record encryptions, not the misuse of a dangerous sink at one call site. Therefore no rule in the CVE-derived rule base expresses it, which is the scope boundary already declared in §5.2. Caveats. Three limits bound this comparison. First, ground truth is vendor-disclosed CVEs only and the clone we mined ends at 2026-06-24, so silently fixed bugs and later fixes are absent. Second, the match rule is enclosing-function identity, which is coarser than root-cause identity. Third, this is a single C/TLS target chosen because ground truth exists for it, so the recall ordering above is not a general ranking of these harnesses.

5.5

Raw

run1 run2 run3 run4 run5 run6 run7 run8 run9 run10

80 100

reported findings (64 known CVEs localized) Mythos GT (of 9)

Run

A stable core and a stochastic tail. The variation is highly structured rather than uniform. Of the 243 findings in the ten-run union, 56 appear in every run and 55 appear in only one; the remaining 132 fall between these extremes (Figure 3b). At the candidate level, 93.1% of the 8,115 candidates receive the same verdict in all ten runs once UNKNOWN abstentions are counted as changes; restricting to outright BUGGY–FALSE_POSITIVE flips leaves 509 contested candidates (6.3%). Either way, disagreement is concentrated in a relatively small set of borderline cases. Findings supported by multiple BUGGY judgments within a run are also much more stable across runs: 73.0% are unanimous across all ten runs, compared with 15.8% for findings supported by a single judgment. This suggests that within-run corroboration provides a useful indicator of expected stability without requiring another scan. Reruns improve coverage with diminishing returns. Repeating Phase B expands the finding set, but the marginal gain decreases quickly. Averaged over all subsets of the ten runs, three runs recover 75.8% of the ten-run union and five recover 85.7% (Figure 3a). The second run adds an expected 32 findings, the third adds 18, and later runs contribute progressively fewer. A Chao2 incidence-based extrapolation [13, 16] over the runs-by-locations incidence matrix estimates approximately 302 Phase B-reportable locations under this harness, suggesting that even ten runs recover only about 80.4% of

Run-to-Run Reproducibility

Ten identical runs. Because Phase B relies on an LLM agent, identical scans need not produce identical findings. We therefore rolled Pillow back to commit d56032047d11 and scanned it ten times with the same runner and gpt-5-mini model, using a frozen 174-skill snapshot of the rule base (the current production library contains 172 skills). All runs enumerate the same 30,594 raw matches and remove the same 10

(b) reproducible core vs. stochastic tail 80

Chao2 estimate Ŝ = 302

300 250

+9 +11 +13 +18 +32

200

+8+7

+6+5

10 runs: 243

150

best / worst subset of runs expected union (all subsets) extrapolation

1 run: 134 (55% of 10-run union)

100

distinct findings

distinct findings (union)

(a) how much a rerun adds

observed single-probability model (p = 0.55) reproducible core: 56 in every run

60

40 stochastic tail: 55 found once

20

0 1

3

5

7

10

15

20

25

1

number of independent runs

2

GPT-5.6-sol GPT-5.6-luna Gemini 3.1 Pro Sonnet 4.6

gpt-5-mini mean $161

n=121; 1 run 78%, 10 runs 99%

$900 $869

n=125; 1 run 72%, 10 runs 97%

$71 n=36; 1 run 41%, 10 runs 69%

$575 n=8; 1 run 90%, 10 runs 100%

$330 n=31; 1 run 94%, 10 runs 97%

Qwen3-6-35B-A3B

free

gpt-5-mini 1-run mean 134

0

20

40

60

80

100

120

140

distinct findings from one run in all 10 gpt-5-mini runs (stable core) in some runs (>=1, not all)

4

5

6

7

8

9

10

(d) rerun diversity by harness

n=127; 1 run 78%, 10 runs 99%

0

500

1000

per-run API cost (USD)

union / harness's own total

(c) one-run findings and cost by model GPT-5.5

3

reported by how many of the 10 runs

1.0

0 in all 10 (no core)

0.8 0.6

56 in all 10 (reproducible core)

0.4 0.2

Bugstone (n=10, ratio 0.23) Codex (n=10, ratio 0.00)

0.0 1

2

3

4

5

7

10

number of runs

in no run (unique) 1 random gpt-5-mini run reaches (mean, min-max)

Figure 3: Reproducibility of Phase B on the same Pillow snapshot. (a) Expected finding coverage from repeated gpt-5-mini runs. (b) Finding frequency across ten runs. (c) Findings and API cost for six alternative models on the same 8,115 candidates(d) Growth of the finding union across reruns of B UGSTONE -E2E and the OpenAI Codex security harness. them. This 302 is the population reachable by this rule set and candidate stage, not an estimate of the true number of vulnerabilities in Pillow. Repetition is therefore useful for expanding coverage, but no small number of reruns eliminates the stochastic tail.

small even when candidates share rule and file context. Reproducibility is not recall. More findings from repeated runs do not necessarily imply proportionally higher vulnerability recall. The rolled-back Pillow snapshot contains 21 vulnerabilities disclosed and fixed after the selected revision, but only five reach Phase B as candidates. Across ten runs, Phase B reports four of those five at least once, while a typical run reports only one. The remaining sixteen are lost before Phase B: fourteen are never enumerated at all, eight of them because the deployed skill set has no rule for their CWE, and two more are enumerated but removed by the Phase A.1 filter. Repetition can therefore recover semantic-verification misses, but it cannot compensate for missing detection knowledge or candidate coverage.

Where does the variation come from?. The variation is almost entirely due to verdict changes rather than worker or reporting failures. Across the 1,087 missing occurrences relative to the ten-run union, 1,061 (97.6%) are BUGGY-toFALSE_POSITIVE flips, 26 are abstentions, and none result from worker failure. These flips are also correlated within a semantic task: the 509 contested candidates occur in only 153 batches, with strong within-batch correlation (φ = 0.71) and essentially no correlation across batches (φ = −0.001). Thus, a bounded batch behaves partly as one shared reasoning decision, rather than as several independent candidate classifications. This observation motivates keeping Phase B batches

Rerunning a small model versus using a stronger model. We next run six additional models once each on the same 11

Model

$/1M in/out

Tokens

Findings

Cost

gpt-5-mini gpt-5.6-luna sonnet-4.6 gemini-3.1-pro gpt-5.6-sol gpt-5.5 qwen3-6-35b-a3b

0.25 / 2.00 0.20 / 1.20 3.00 / 15.0 2.00 / 12.0 5.00 / 30.0 5.00 / 30.0 self-hosted

723M 479M 166M 311M 311M 309M 346M

134 125 8 36 121 127 31

$161 $71 $330 $575 $869 $900 free

5.6

Remediation Prototype

Patch quality is difficult to evaluate automatically: whether a repair is complete, regression-free, and acceptable to maintainers often depends on project-specific tests and human review. We therefore evaluate Phase D only as a prototype for exploit-guided remediation rather than as a general patch correctness system. We apply Phase D to 23 findings across 3 targets for which Phase C produces a reproducible proof of concept. Only one of these targets, the openai-python SDK (11 of the 23 findings), is an independent third-party codebase; the other two are a small internal Flask exercise and a C++ vulnerabilityhunting exercise. For each finding, Phase D generates a patch and applies the two-sided differential test from §4.3: the proof of concept must reproduce on the original program, stop reproducing after the patch, and reproduce again after the patch is reverted. Phase D produced a patch for all 23 findings, and all 23 pass this differential check. The resulting patches are compact, with a median of 8 changed lines (6 added, 2 removed), and 21 of the 23 touch a single file. These results provide machine-checkable evidence that the current prototype can turn a live finding into a compact candidate remediation that blocks the demonstrated behavior. The differential outcome is the worker’s own transcribed runtime output rather than an execution we re-ran, and this small sample does not measure general patch correctness or maintainer acceptance, which require substantially broader project-specific testing and review.

Table 4: Per-run Phase B cost on the same Pillow snapshot and candidates. $/1M: listed input and output prices per million tokens. Cost is provider-billed for the whole run and includes retries and cached-input tokens priced well below the listed input rate, so it is not the product of the Tokens column and the in/out rates. gpt-5-mini Findings and Cost are the ten-run mean; the other rows are a single run. Prices are the provider list rates as of 2026-08. free: Qwen3-6-35B-A3B runs on a local A100, so it incurs no per-token API charge.

Pillow snapshot and the same 8,115 Phase B candidates. A single gpt-5-mini run already covers 78% of the findings reported by GPT-5.5, 78% of GPT-5.6-sol, and 72% of GPT5.6-luna, close to the 79% overlap between two gpt-5-mini runs. Across ten gpt-5-mini runs, coverage of these strongermodel finding sets rises to 97–99%, leaving at most four findings unique to any one model (Figure 3c). Several frontier configurations are also substantially more expensive (Table 4). One gpt-5-mini run averages approximately $161, compared with about $900 for GPT-5.5 and $869 for GPT-5.6-sol on the same candidate set, while reporting a comparable or larger number of Phase B findings. For this bounded verification task, spending the same budget on several small model runs therefore provides greater Phase B finding coverage than replacing the small model with a single frontier-model run. This result supports the RC2 design choice: once semantic analysis is decomposed into narrow candidate-level tasks, Phase B does not require the strongest available model for every judgment.

6

Real-World Bug Detection

B UGSTONE -E2E scanned 14 third-party projects and recorded every finding in a deduplicating registry. §6 reports the outcome per target, from the Phase A candidate pool through the Phase C verification tiers. Each row is one project, and the finding counts are deduplicated across every configuration run on it, so a finding seen by two configurations is counted once. The registry omits Pillow, which is scanned only for the reproducibility study of §5.5. The scan portfolio by verification tier. The columns of §6 are strictly nested, so a reader can pick an evidentiary bar and read off the count. Every CONFIRMED_EXPLOITABLE finding has a live proof-of-concept, and every live proofof-concept is a runtime observation. Across the 14 targets, B UGSTONE -E2E deduplicated 2,933 Phase B findings, of which Phase C adjudicated 2,125 and 644 carry runtime evidence, meaning a sanitizer report, a crash, or a controlled-sink signal at the reported location. Runtime evidence is the headline bar we report, and the stricter live-proof-of-concept and CONFIRMED_EXPLOITABLE tiers in §6 narrow it further for readers who want them. Of the 2,125 findings Phase C adjudicated, it rejected 1,099 as not a bug or not exploitable and left 382 with a static-reachability argument but no runtime signal, so

Harness dependence and limitations. The benefit of rerunning is not universal to all agentic scanners. Ten runs of the OpenAI Codex security harness on the same Pillow snapshot show substantially less recurring overlap than B UGSTONE -E2E (Figure 3d), indicating that rerun behavior depends on how a harness structures its reasoning tasks. We treat this comparison as a diversity observation rather than a like-for-like recall result because the harnesses expose findings at different stages. Finally, this study covers one target and one primary Phase B configuration, and the repeated runs stop before dynamic validation. The additional findings in the stochastic tail therefore should not be interpreted as confirmed vulnerabilities. Instead, the experiment measures the stability and coverage behavior of the semantic-verification stage itself. 12

Target

Lang.

Domain

Cand.

Phase B

Phase C

Runtime

Live

Conf.

FreeBSD PyTorch OpenSSL WordPress‡ NVIDIA OGK ImageMagick wxo-clients guava openssh wolfSSL† openai-python MCP servers codex-src claude-code

C/C++ Python C/C++ PHP/JS C C/C++ Python Java C C Python TS/JS Python/Rust Python/TS

OS kernel ML framework Crypto lib. CMS GPU driver Image tooling Enterprise SDK Java core lib. SSH impl. Embedded TLS API client Agent servers Coding agent Coding agent

76,843 38,605 20,357 8,636 6,449 5,161 3,907 3,843 3,441 3,166 841 739 1,150 586

2,189 295 20 18 26 162 46 22 1 31 34 36 47 6

1,485 195 16 18 26 162 46 22 1 31 34 36 47 6

459 95 5 0 8 4 41 0 0 16 0 12 2 2

40 95 5 0 8 4 41 0 0 16 0 12 2 1

16 27 2 0 0 2 0 0 0 4 0 10 0 1

2,933

2,125

644

224

62

Total

Table 5: Scanned targets, ordered by candidate volume, with deduplicated finding counts from the B UGSTONE -E2E registry. Cand. is Phase A candidates after the Phase A.1 filter; Phase B distinct BUGGY locations; Phase C findings reaching Phase C; Runtime a sanitizer, crash, or PoC signal at the sink; Live an exploitability of LIVE_POC; and Conf. the strongest CONFIRMED_EXPLOITABLE tier. Each tier is strictly stronger than the one to its left. † wolfSSL is the cross-harness target of §5.4; this row quotes an earlier registry run with a smaller skill set (3,166 candidates) rather than the later rolled-back Pi run analyzed there. ‡ Five WordPress stored-XSS reports share one file and line and collapse to a single registry record, so the row is reported at its true 18. 644 reached the runtime bar. The gap between the Phase B and Phase C columns is a separate matter: it is backlog, not a negative result. Specifically, 808 findings have never been given a runtime at all, 704 of them in FreeBSD. Reporting the rejected set and the untested backlog separately is what keeps “not exploitable” distinct from “could not be tested,” so the runtime-evidenced counts are lower bounds set by how much Phase C compute we spent, not by how many findings survive scrutiny.

are internal utility functions that call a dangerous API with a constant or already-validated argument, where the Phase A pattern matches but Phase B identifies the effective guard. Machine-generated code is the second category. For example, openssh’s libcrux_mlkem768 ML-KEM implementation contains ≈635 syntactic buffer-write matches, and all were correctly dismissed because the generated code uses bounded index arithmetic throughout. Incomplete taint chains are the third category, in which the dangerous sink exists but usercontrolled input does not reach it within the analyzed call depth. Phase B labels these FALSE_POSITIVE with a taintchain explanation, so the structured evidence format enables principled labeling rather than a bare verdict.

Phase A enumeration. Summed across the 14 targets, the Cand. column of §6 totals 173,724 post-filter candidates that reach Phase B, after the deterministic Phase A.1 filter has already removed benign sites without any model call. Each candidate then reaches Phase B, which follows the callers, weighs every guard on the taint path, and returns a BUGGY verdict only with code evidence.

CWE families and languages. Runtime-evidenced findings concentrate in families with deep rule coverage. CWE-502 insecure deserialization dominates PyTorch and guava, CWE-79 cross-site scripting dominates WordPress, and the C memorysafety families CWE-120, CWE-122, CWE-125, CWE-787, and CWE-416 account for the OpenSSL and FreeBSD findings. CWE-22 path traversal, CWE-77 and CWE-78 command injection, and CWE-918 server-side request forgery carry runtime evidence in wxo-clients under multiple configurations. The runtime-evidenced findings therefore span injection and deserialization as well as memory-safety families, and they cover six languages from C to TypeScript. This spread matches the cross-language coverage of the rule base in §5.2.

Phase B and Phase C findings. Runtime-evidence rates differ sharply by target and by weakness class. On guava, all 22 Phase B deserialization candidates reach Phase C, but none yet carries runtime evidence, so guava contributes no findings at the runtime bar. The single openssh finding, a CWE-190 integer overflow in setenv.c:183, illustrates precision in a widely audited codebase, since 3,441 candidates narrow to exactly one Phase B finding, still unverified dynamically. PyTorch concentrates on RCE-class deserialization bugs, with 195 findings reaching Phase C and 95 carrying a runtime observation. FreeBSD dominates raw volume, with 2,189 Phase B findings against a 76,843-candidate pool. Only 459 have a runtime signal so far, because the Phase C backlog there is the largest.

The scan model shifts volume, not runtime evidence. The scan model changes how many candidates Phase B marks BUGGY, but Phase C then filters that raw volume down to what a runtime can support. Under the controlled Pillow comparison of §5.5, larger models and gpt-5-mini report compara-

False positives cluster in three shapes. Dominant falsepositive sources fall into three categories. Benign wrappers 13

ble Phase B finding counts on the same candidate set, so a stronger scan model mainly reshapes which candidates are flagged rather than how many survive runtime checking. The extra candidates a noisier model raises are not free: each one still consumes Phase C compute, and the 808-finding Phase C backlog above is exactly the cost of that raw volume. Splitting Phase B into a simple per-candidate judgment keeps each judgment cheap enough for a small model, while the dynamic Phase C confirmation carries the load-bearing runtime evidence. Therefore B UGSTONE -E2E does not depend on the strongest and most expensive model to produce runtimeevidenced findings. A harness that concentrates all reasoning in one large model has the opposite property.

7.2

7

8

7.1

Future Work

Future work can extend B UGSTONE -E2E along these limitations. Broader CVE coverage and rule auditing can improve generality across CWE families, while richer interprocedural context can increase detection coverage beyond locally anchored cases. Runtime validation can likewise benefit from more reusable build environments and broader confirmation of Phase B findings. Remediation requires substantially more project-specific and human-in-the-loop evaluation: future versions of Phase D can incorporate native regression tests, assess whether patches address the complete root cause, and track maintainer feedback as more patches are submitted. We therefore view the current remediation stage as an evolving prototype rather than a complete automated repair system.

Discussion

Related Work

Static and query-based vulnerability analysis. Code property graphs, QL and CodeQL, Semgrep, and Infer enumerate candidates from human-authored queries or rules [6, 7, 10, 28, 62, 84]. Recent work attacks that authoring bottleneck. Specifically, QLCoder synthesizes CodeQL queries from CVE metadata, and QRS generates queries agentically and validates findings by exploit synthesis [70, 73]. SemTaint likewise extracts per-package taint specifications for CodeQL [27]. Therefore B UGSTONE -E2E is not the first system to turn CVE history into executable detectors. However, B UGSTONE -E2E emits CWE- and language-level rules that group verified fixing commits, keep source_cves provenance, and carry an agent-consumed verification contract.

Limitations

Detection scope and completeness. B UGSTONE -E2E is designed for recurring vulnerability patterns that can be anchored to security-relevant operations or API uses. It therefore does not target one-off design flaws, configuration errors, or global and stateful invariants without a suitable local anchor. CVE-derived rules may also remain incomplete. A rule can miss semantically equivalent APIs or conditions not represented in the source fixes. In addition, conservative candidate filtering and bounded Phase B reasoning can introduce false negatives, particularly when a decision requires deeper interprocedural context. Because Phase B is model-driven, its verdicts are also stochastic; reruns can recover some semanticanalysis misses but cannot recover vulnerabilities outside the deployed rule and candidate coverage.

Vulnerability clone and patch-signature detection. ReDeBug, VUDDY, MVP, MOVERY, V1SCAN, and VMud reuse vulnerability history through code, component, or patch-line signatures [30, 31, 35, 76, 77, 82]. MAVM replaces those signatures with agents that detect, confirm, repair, and validate recurring cases [88]. In contrast, B UGSTONE -E2E matches CWE- and API-level rules, so a target need not clone the seed vulnerability, and confirmation requires an executed exploit.

Runtime validation. Runtime confirmation depends on the target admitting a reproducible build and test environment. Complex dependencies, unavailable inputs, or platformspecific behavior can prevent Phase C from exercising an otherwise plausible finding. Such cases should therefore be interpreted as unverified rather than benign. Dynamic validation is also necessarily more expensive than static verification, so large Phase-B finding sets may only be partially exercised.

Patch analysis and rule inference. Empirical studies and datasets such as BigVul show that security patches encode security properties [19, 39]. Prior work infers those properties for missing checks, security impact, disordered error handling, and severity prioritization [48, 78–80]. Closer to B UGSTONE -E2E, VulGenie derives Java API rules by attackdefense cross-analysis, while RuleForge and RulePilot convert CVE records and analyst annotations into validated detection rules [14, 25, 74]. In addition, AutoTrace turns one fixing commit into a trigger location, and GONDAR supplies CWE-specific sink knowledge to exploit agents [22, 92]. However, B UGSTONE -E2E differs in the unit and scale of reuse, since it groups public CVE history into CWE-language rules packaged as deployable agent skills.

Remediation guarantees. Phase D remains a prototype. Its differential oracle establishes that a patch blocks the demonstrated proof of concept and that reverting the patch restores the behavior. It does not establish that the repair is complete, regression-free, or acceptable to maintainers. Evaluating these properties at scale requires project-specific testing and substantial human review, so we treat remediation as an endto-end capability rather than a measured patch correctness result. 14

Learning-based vulnerability prediction. VulDeePecker, SySeVR, Devign, ReVeal, LineVul, DeepDFA, and FVDDPM learn vulnerability signals from gadgets, graphs, or transformer encodings [12, 23, 42, 43, 64, 66, 89]. That line now extends into agent training. For example, VulAgentRL rewards only verdicts whose evidence checks against a code property graph, and Antares distills compact localization models [41, 71]. However, these systems predict over learned representations, whereas B UGSTONE -E2E keeps CVE provenance per rule and demands taint, guard, and caller rationale.

backends shift results, and that injected static structure halves variance [46, 58, 90]. Therefore B UGSTONE -E2E fixes its Phase A candidate set before any model runs and measures run-to-run reproducibility, so execution collapses the remaining spread. Security taxonomies and public vulnerability data. B UGSTONE -E2E builds on public infrastructure, namely cvelistV5 records, NVD metadata, the CWE taxonomy, and OSV advisories [17, 50, 51, 56]. CVEfixes, MoreFixes, MegaVul, and VulZoo unify CVE-linked fixes, while ThreatKG and VulnScopper mine missing relations among security entities [1, 3, 8, 52, 61, 65]. Therefore B UGSTONE -E2E differs by turning verified fixing commits into executable scanner inputs with provenance rather than into datasets.

LLMs for vulnerability detection. LLift, Vul-RAG, IRIS, LLMxCPG, and specialized reasoning models guide LLMs with static analysis, graphs, or retrieved vulnerability knowledge [18, 38, 40, 44, 54]. Benchmark studies measure that behavior across model sizes, prompts, repositories, and agent settings [45, 55, 68, 85]. Repository-scale agents followed quickly. Specifically, TitanCA reports a deployed match, filter, inspect, and adapt pipeline, LLMVD.js exploits taint bugs in Node.js packages, and DREA separates exploration from reasoning [53, 67, 87]. Evaluation moved with it, since VulnGym, SastBench, and RealVuln supply traces, triage distributions, and scanner rankings [20, 32, 59]. Also, agentic filtering removes most SAST noise yet suppresses true positives, and a Vul-RAG replication reports an accuracy plateau [34, 83]. Therefore B UGSTONE -E2E layers filtering into a cost-ordered cascade, in which Phase A.1 rejects deterministically and only survivors reach an agent.

9

Conclusion

B UGSTONE -E2E turns vulnerability history into reusable detection knowledge and applies it through a staged workflow from scalable detection to runtime validation and candidate remediation. Its current knowledge base contains 1,033 rules consolidated into 56 CWE families and deployed through 172 skills. Across controlled experiments and real-world scans of 14 projects, these skills surface recurring vulnerability conditions beyond the projects from which they were originally derived, with 644 findings supported by runtime evidence. More broadly, B UGSTONE -E2E concentrates increasingly expensive analysis on progressively fewer candidates while grounding stronger security claims in observable evidence rather than model judgment alone. Our results show that historical vulnerability fixes can serve not only as records of past failures, but also as reusable knowledge for finding and validating future ones.

Agentic program repair and self-validating patches. AgenticRepair assembles repair context through subagents, while KeaRepair grounds patches in verified facts and mined vulnerability–patch pairs [11, 24]. However, acceptance criteria are fragile, since adversarial issue reports yield newly vulnerable patches and most Java CVE patches fail stricter oracles [2, 15]. Execution-based patch acceptance therefore has prior art. Specifically, VulnRepairEval requires the original exploit to fail, Vul4Py adds a functional oracle, and VeriPort chains evidence across affected versions [9, 26, 75]. In contrast, Phase D patches findings that B UGSTONE -E2E discovered itself, so acceptance needs a two-sided test in which the exploit fails under the patch and succeeds after reverting it.

References [1] J. Akhoundali, S. R. Nouri, K. Rietveld, and O. Gadyatskaya. MoreFixes: A large-scale dataset of CVE fix commits mined through enhanced repository discovery. In International Conference on Predictive Models and Data Analytics in Software Engineering, pages 42–51, 2024. doi: 10.1145/3663533.3664036.

Exploit generation, dynamic confirmation, and agent determinism. Big Sleep, ATLANTIS, and the AIxCC systematization show agents finding and repairing real bugs alongside fuzzing and symbolic execution [29, 36, 86]. KRepro, SEC-bench Pro, and CVE-Bench measure proof-ofconcept construction on kernel N-days, browser engines, and web CVEs [37, 60, 91]. However, agents still fail on servicebased vulnerabilities whose environments they cannot stand up [47]. RECEIPT restores verdict trust through isolation and binding, and FalseCrashReducer validates caller context against infeasible crashes [4, 49]. Repeated-run studies report that single-run scores overstate coverage, that inference

[2] A. Al-Maamari. Why LLMs fail: A failure analysis and partial success measurement for automated security patch generation, 2026. URL https://arxiv.org/ abs/2603.10072. arXiv:2603.10072. [3] D. Alfasi, T. Shapira, and A. Bremler-Barr. VulnScopper: Unveiling hidden links between unseen security entities. In International Workshop on Graph Neural Networking, pages 33–40, 2024. doi: 10.1145/ 3694811.3697819. 15

[4] P. C. Amusuo, D. Liu, R. A. C. Mendez, J. Metzman, O. Chang, and J. C. Davis. FalseCrashReducer: Mitigating false positive crashes in OSS-Fuzz-Gen using agentic AI, 2025. URL https://arxiv.org/abs/ 2510.02185. arXiv:2510.02185.

[14] B. Chen, S. Liao, L. Zhang, C. Zhang, M. Payer, and Y. Zhang. Patch-guided vulnerability detection: Extracting java API security rules via attack-defense crossanalysis. In USENIX Security Symposium, 2026. Prepublication.

[5] Anthropic. Defending code reference harness. https :// github . com / anthropics / defending code-reference-harness, 2026. Accessed 202608-10.

[15] S. Chen, Y. He, S. Jana, and B. Ray. Red teaming program repair agents: When correct patches can hide vulnerabilities, 2025. URL https://arxiv.org/abs/ 2509.25894. arXiv:2509.25894.

[6] P. Avgustinov, O. de Moor, M. P. Jones, and M. Schäfer. QL: Object-oriented queries on relational data. In European Conference on Object-Oriented Programming, pages 2:1–2:25, 2016. doi: 10.4230/LIPIcs.ECOOP. 2016.2.

[16] R. K. Colwell, A. Chao, N. J. Gotelli, S.-Y. Lin, C. X. Mao, R. L. Chazdon, and J. T. Longino. Models and estimators linking individual-based and sample-based rarefaction, extrapolation and comparison of assemblages. Journal of Plant Ecology, 5(1):3–21, 2012. doi: 10.1093/jpe/rtr044.

[7] G. Bennett, T. Hall, E. Winter, and S. Counsell. Semgrep*: Improving the limited performance of static application security testing (SAST) tools. In International Conference on Evaluation and Assessment in Software Engineering, pages 614–623, 2024. doi: 10.1145/3661167.3661262.

[17] CVE Program. CVE list V5. https://github.com/ CVEProject/cvelistV5, 2026. Official CVE List repository in CVE JSON 5 format. [18] X. Du, G. Zheng, K. Wang, Y. Zou, Y. Wang, W. Deng, J. Feng, M. Liu, B. Chen, X. Peng, T. Ma, and Y. Lou. Vul-RAG: Enhancing llm-based vulnerability detection via knowledge-level rag. ACM Transactions on Software Engineering and Methodology, 2026. doi: 10.1145/ 3797277. arXiv:2406.11147.

[8] G. Bhandari, A. Naseer, and L. Moonen. CVEfixes: Automated collection of vulnerabilities and their fixes from open-source software. In International Conference on Predictive Models and Data Analytics in Software Engineering, pages 30–39, 2021. doi: 10.1145/3475960. 3475985.

[19] J. Fan, Y. Li, S. Wang, and T. N. Nguyen. A c/c++ code vulnerability dataset with code changes and cve summaries. In International Conference on Mining Software Repositories, 2020.

[9] T. Bui, T. Zhang, F. Thung, Y. Xiong, P. Jiang, X. Zhou, and D. Lo. Vul4Py: Benchmarking automated vulnerability repair in Python with paired exploit and functional oracles, 2026. URL https://arxiv.org/abs/ 2608.00692. arXiv:2608.00692.

[20] J. Feiglin and G. Dar. SastBench: A benchmark for testing agentic SAST triage, 2026. URL https:// arxiv.org/abs/2601.02941. arXiv:2601.02941.

[10] C. Calcagno, D. Distefano, J. Dubreil, D. Gabi, P. Hooimeijer, M. Luca, P. O’Hearn, I. Papakonstantinou, J. Purbrick, and D. Rodriguez. Moving fast with software verification. In NASA Formal Methods, pages 3–11, 2015.

[21] A. Fioraldi, D. Maier, H. Eißfeldt, and M. Heuse. AFL++: Combining incremental steps of fuzzing research. In USENIX Workshop on Offensive Technologies, 2020. [22] F. Fleischer, C. Zhang, J. Jang, J. Cho, M. Xu, and T. Kim. Contextualizing sink knowledge for Java vulnerability discovery, 2026. URL https://arxiv.org/ abs/2604.01645. arXiv:2604.01645.

[11] S. Cao, H. Ma, L. Yu, K. Ding, X. Liu, T. Y. Zhuo, B. Wang, X. Lin, X. Sun, L. Wang, and D. Lo. Knowledge-enhanced agentic vulnerability repair, 2026. URL https://arxiv.org/abs/2607. 00820. arXiv:2607.00820.

[23] M. Fu and C. Tantithamthavorn. LineVul: A transformerbased line-level vulnerability prediction. In International Conference on Mining Software Repositories, 2022.

[12] S. Chakraborty, R. Krishna, Y. Ding, and B. Ray. Deep learning based vulnerability detection: Are we there yet? IEEE Transactions on Software Engineering, 48 (9):3280–3296, 2022. doi: 10.1109/TSE.2021.3087402.

[24] M. Fu, Q. Mei, P. Thongtanunam, and K. Tantithamthavorn. AgenticRepair: Multi-faceted program context engineering for agentic vulnerability repair, 2026. URL https://arxiv.org/abs/2607.29422. arXiv:2607.29422.

[13] A. Chao. Estimating the population size for capturerecapture data with unequal catchability. Biometrics, 43 (4):783–791, 1987. doi: 10.2307/2531532. 16

[25] A. Garg, S. Hager, J. Montiel, A. Tiwari, M. Gentile, Z. Reavis, D. Magnotti, and W. Fullen. RuleForge: Automated generation and validation for web vulnerability detection at scale, 2026. URL https://arxiv.org/ abs/2604.01977. arXiv:2604.01977.

Physical Systems (AI&CCPS), co-located with ARES, 2026. arXiv:2606.04739. [35] S. Kim, S. Woo, H. Lee, and H. Oh. VUDDY: A scalable approach for vulnerable code clone discovery. In IEEE Symposium on Security and Privacy, 2017.

[26] J. Ghebremichael, W. Jiang, M. Lysenko, B. B. Nielsen, W. Enck, and A. Kapravelos. VeriPort: Automated and verified patch backporting at scale, 2026. URL https://arxiv.org/abs/2606.22704. arXiv:2606.22704.

[36] T. Kim, H. Han, S. Park, D. R. Jeong, D. Kim, D. Kim, et al. ATLANTIS: AI-driven threat localization, analysis, and triage intelligence system, 2025. URL https://arxiv.org/abs/2509.14589. Team Atlanta, 1st place, DARPA AIxCC Final Competition; arXiv:2509.14589.

[27] J. Ghebremichael, S. Vasan, S. Ullah, G. Tystahl, D. Adei, C. Kruegel, G. Vigna, W. Enck, and A. Kapravelos. Multi-agent taint specification extraction for vulnerability detection, 2026. URL https://arxiv.org/ abs/2601.10865. arXiv:2601.10865.

[37] H. Lee, J. Liu, D. Kim, W. Xia, Z. Zhang, C. S. Xia, and L. Zhang. SEC-bench Pro: Can language models solve long-horizon software security tasks?, 2026. URL https://arxiv.org/abs/2605.26548. arXiv:2605.26548.

[28] GitHub. CodeQL: Semantic code analysis. https:// codeql.github.com/, 2026.

[38] A. Lekssays, H. Mouhcine, K. Tran, T. Yu, and I. Khalil. LLMxCPG: Context-aware vulnerability detection through code property graph-guided large language models. In USENIX Security Symposium, pages 489– 507, 2025.

[29] Google Project Zero and Google DeepMind Big Sleep Team. From Naptime to Big Sleep: Using large language models to catch vulnerabilities in real-world code. Google Project Zero blog, Nov. 2024. URL https:// googleprojectzero . blogspot . com / 2024 / 10 / from- naptime- to- big- sleep.html. Accessed 2026-08-04.

[39] F. Li and V. Paxson. A large-scale empirical study of security patches. In ACM Conference on Computer and Communications Security, pages 2201–2215, 2017.

[30] K. Huang, C. Lu, Y. Cao, B. Chen, and X. Peng. VMud: Detecting recurring vulnerabilities with multiple fixing functions via function selection and semantic equivalent statement matching. In ACM Conference on Computer and Communications Security, pages 3958–3972, 2024. doi: 10.1145/3658644.3690372.

[40] H. Li, Y. Hao, Y. Zhai, and Z. Qian. Enhancing static analysis for practical bug detection: An llm-integrated approach. Proceedings of the ACM on Programming Languages, 8(OOPSLA1):474–499, 2024. [41] Y. Li, T. Zhang, J. Liu, J. Jiang, Y. Yieh, Y. Yang, W. B. Leow, Y. Yin, Y. Huo, E. L. Ouh, L. K. Shar, and D. Lo. Graph is the verifier: Agentic reinforcement learning for interprocedural vulnerability detection, 2026. URL https://arxiv.org/abs/2607.26656. arXiv:2607.26656.

[31] J. Jang, A. Agrawal, and D. Brumley. ReDeBug: Finding unpatched code clones in entire OS distributions. In IEEE Symposium on Security and Privacy, pages 48–62, 2012. doi: 10.1109/SP.2012.13. [32] K. Ji, J. Liu, E. Hu, C. Gao, K. Lian, Y. Liu, L. Zhang, T. Dong, H. Chen, and W. Bin. VulnGym: Benchmarking coding agents for repository-level vulnerability detection, 2026. URL https://arxiv.org/abs/ 2608.02001. arXiv:2608.02001.

[42] Z. Li, D. Zou, S. Xu, X. Ou, H. Jin, S. Wang, Z. Deng, and Y. Zhong. VulDeePecker: A deep learning-based system for vulnerability detection. In Network and Distributed System Security Symposium, 2018. doi: 10.14722/ndss.2018.23158.

[33] B. Johnson, Y. Song, E. Murphy-Hill, and R. Bowdidge. Why don’t software developers use static analysis tools to find bugs? In International Conference on Software Engineering, pages 672–681, 2013.

[43] Z. Li, D. Zou, S. Xu, H. Jin, Y. Zhu, and Z. Chen. SySeVR: A framework for using deep learning to detect software vulnerabilities. IEEE Transactions on Dependable and Secure Computing, 19(4):2244–2258, 2022.

[34] S. Kaniewski, F. Schmidt, and T. Heer. Revisiting Vul-RAG: Reproducibility and replicability of RAGbased vulnerability detection with open-weight models. In International Workshop on Artificial Intelligence and Cybersecurity for Critical and Cyber-

[44] Z. Li, S. Dutta, and M. Naik. IRIS: LLM-assisted static analysis for detecting security vulnerabilities. In International Conference on Learning Representations, 2025. 17

[45] J. Lin and D. Mohaisen. From large to mammoth: A comparative evaluation of large language models in vulnerability detection. In Network and Distributed System Security Symposium, 2025.

[56] Open Source Security Foundation. Open Source Vulnerability Schema and Database. https://osv.dev/, 2026. Machine-readable open-source vulnerability advisory database.

[46] Z. Lin, M. Zhou, Y. Yang, and L. Li. How much static structure do code agents need? A study of deterministic anchoring. In International Symposium on Software Testing and Analysis (ISSTA), 2026. arXiv:2606.26979.

[57] OpenAI. Codex security plugin. https://openai. com/daybreak/codex- security- plugin/, 2026. Accessed 2026-08-10. See also https :// learn . chatgpt.com/docs/security/plugin.

[47] B. Liu, Y. Zhao, G. Xu, and H. Wang. LLM agents for automated web vulnerability reproduction: Are we there yet?, 2025. URL https://arxiv.org/abs/ 2510.14700. arXiv:2510.14700.

[58] D. Pape, J. Evertz, and L. Schönherr. The silent hyperparameter: Quantifying the impact of inference backends on LLM reproducibility, 2026. URL https:// arxiv.org/abs/2605.19537. arXiv:2605.19537. [59] J. Pellew and F. Raza. RealVuln: Benchmarking rulebased, general-purpose LLM, and security-specialized scanners on real-world code, 2026. URL https:// arxiv.org/abs/2604.13764. arXiv:2604.13764.

[48] K. Lu, A. Pakki, and Q. Wu. Detecting missing-check bugs via semantic- and context-aware criticalness and constraints inferences. In USENIX Security Symposium, pages 1769–1786, 2019.

[60] J. Pu, X. Li, Z. Liang, J. Cox, Y. Wu, K. Shehada, A. Srivastav, and Z. Qian. Patch-to-PoC: A systematic study of agentic LLM systems for Linux kernel N-day reproduction, 2026. URL https://arxiv.org/abs/ 2602.07287. arXiv:2602.07287.

[49] M. Lyu, K. Shieh, Y. Hou, H. Wang, K. Sen, and D. Wagner. RECEIPT: Deterministic, reward-hackingresistant verification for white-box agentic XSS discovery, 2026. URL https://arxiv.org/abs/2607. 18575. arXiv:2607.18575.

[61] B. Ruan, J. Liu, W. Zhao, and Z. Liang. VulZoo: A comprehensive vulnerability intelligence dataset. In International Conference on Automated Software Engineering, pages 2334–2337, 2024. doi: 10.1145/ 3691620.3695345.

[50] MITRE. Common Weakness Enumeration. https:// cwe.mitre.org/, 2026. [51] National Institute of Standards and Technology. National Vulnerability Database. https://nvd.nist. gov/, 2026. Standards-based vulnerability management data repository.

Semgrep documentation. [62] Semgrep. semgrep.dev/docs/, 2026.

[52] C. Ni, L. Shen, X. Yang, Y. Zhu, and S. Wang. MegaVul: A C/C++ vulnerability dataset with comprehensive code representations. In International Conference on Mining Software Repositories, pages 738–742, 2024.

[63] K. Serebryany, D. Bruening, A. Potapenko, and D. Vyukov. AddressSanitizer: A fast address sanity checker. In USENIX Annual Technical Conference, pages 309–318, 2012.

[53] R. Ni, M. Christodorescu, and L. Jia. Taint-style vulnerability detection and confirmation for Node.js packages using LLM agent reasoning, 2026. URL https:// arxiv.org/abs/2604.20179. arXiv:2604.20179.

[64] M. Shao and Y. Ding. FVD-DPM: Fine-grained vulnerability detection via conditional diffusion probabilistic models. In USENIX Security Symposium, 2024.

https ://

[65] Z. Shi, N. Matyunin, K. Graffi, and D. Starobinski. Uncovering CWE-CVE-CPE relations with threat knowledge graphs. ACM Transactions on Privacy and Security, 27(1):1–26, 2024. doi: 10.1145/3641819.

[54] Y. Nie, H. Li, C. Guo, R. Jiang, Z. Wang, B. Li, D. Song, and W. Guo. VulnLLM-R: Specialized reasoning LLM with agent scaffold for vulnerability detection. arXiv:2512.07533, 2025. URL https://arxiv.org/ abs/2512.07533.

[66] B. Steenhoek, H. Gao, and W. Le. Dataflow analysisinspired deep learning for efficient vulnerability detection. In International Conference on Software Engineering, pages 1–13, 2024. doi: 10.1145/3597503.3623345.

[55] Y. Nie, Z. Wang, Y. Yang, R. Jiang, Y. Tang, X. Davies, Y. Gal, B. Li, W. Guo, and D. Song. SeCodePLT: A unified platform for evaluating the security of code GenAI. NeurIPS Datasets and Benchmarks Track, arXiv:2410.11096, 2025. URL https://arxiv.org/ abs/2410.11096.

[67] M. Sun and G. Meng. DREA: Decoupled reasoning and exploration agents for repository-level vulnerability detection. In International Conference on Internetware (Internetware), 2026. arXiv:2607.13439. 18

[68] Y. Sun, D. Wu, Y. Xue, H. Liu, W. Ma, L. Zhang, Y. Liu, and Y. Li. LLM4Vuln: A unified evaluation framework for decoupling and enhancing LLMs’ vulnerability reasoning. arXiv:2401.16185, 2024. URL https://arxiv.org/abs/2401.16185.

[78] Q. Wu, Y. He, S. McCamant, and K. Lu. Precisely characterizing security impact in a flood of patches via symbolic rule comparison. In Network and Distributed System Security Symposium, 2020. [79] Q. Wu, A. Pakki, N. Emamdoost, S. McCamant, and K. Lu. Understanding and detecting disordered error handling with precise function pairing. In USENIX Security Symposium, 2021.

[69] Tree-sitter Project. Tree-sitter: An incremental parsing system for programming tools. https://treesitter.github.io/tree-sitter/, 2026.

[80] Q. Wu, Y. Xiao, X. Liao, and K. Lu. OS-aware vulnerability prioritization via differential severity analysis. In USENIX Security Symposium, pages 395–412, 2022.

[70] G. Tsigkourakos and C. Patsakis. QRS: A rulesynthesizing neuro-symbolic triad for autonomous vulnerability discovery, 2026. URL https://arxiv. org/abs/2602.09774. arXiv:2602.09774.

[81] Q. Wu, Y. Xiao, D. Kirat, K. Eykholt, J. Jang, and D. L. Schales. One bug, hundreds behind: LLMs for largescale bug discovery. arXiv:2510.14036, 2025. URL https://arxiv.org/abs/2510.14036.

[71] S. Vijay, A. Priyanshu, D. Chapoteau, A. Goldblatt, J. He, K. Majd, F. Burch, B. Saglam, T. Matsumoto, Z. Yang, and A. Karbasi. Antares: Foundation models for agentic vulnerability localization, 2026. URL https://arxiv.org/abs/2608.02407. arXiv:2608.02407.

[82] Y. Xiao, B. Chen, C. Yu, Z. Xu, Z. Yuan, F. Li, B. Liu, Y. Liu, W. Huo, W. Zou, and W. Shi. MVP: Detecting vulnerabilities using patch-enhanced vulnerability signatures. In USENIX Security Symposium, pages 1165– 1182, 2020.

VVAH: Visa vulnerability agentic har[72] Visa. ness. https :// github . com / visa / visa vulnerability - agentic - harness, 2026. Accessed 2026-08-10.

[83] Y. Xiong and T. Zhang. Sifting the noise: A comparative study of LLM agents in vulnerability false positive filtering. In International Symposium on Software Testing and Analysis (ISSTA), 2026. doi: 10.1145/3832100. arXiv:2601.22952.

[73] C. Wang, Z. Li, S. Dutta, and M. Naik. QLCoder: A query synthesizer for static analysis of security vulnerabilities. In International Conference on Learning Representations, 2026. URL https://openreview.net/ forum?id=J91IKwJrqv. arXiv:2511.08462.

[84] F. Yamaguchi, N. Golde, D. Arp, and K. Rieck. Modeling and discovering vulnerabilities with code property graphs. In IEEE Symposium on Security and Privacy, 2014.

[74] H. Wang, M. Xu, Y. Guo, W. Han, H. W. Lim, and J. S. Dong. RulePilot: An LLM-powered agent for security rule generation. In International Conference on Software Engineering (ICSE), 2026. arXiv:2511.12224.

[85] A. Yildiz, S. G. Teo, Y. Lou, Y. Feng, C. Wang, and D. M. Divakaran. Benchmarking LLMs and LLM-based agents in practical vulnerability detection for code repositories. In Annual Meeting of the Association for Computational Linguistics, pages 30848–30865, 2025. doi: 10.18653/v1/2025.acl-long.1490.

[75] W. Wang, W. Ma, Q. Hu, Y. Zhang, J. Sun, B. Wu, Y. Liu, G. Xu, and L. Jiang. VulnRepairEval: An exploit-based evaluation framework for assessing large language model vulnerability repair capabilities, 2025. URL https://arxiv.org/abs/2509.03331. arXiv:2509.03331.

[86] C. Zhang, Y. Park, F. Fleischer, Y.-F. Fu, J. Kim, D. Kim, Y. Kim, Q. Xu, A. Chin, Z. Sheng, H. Zhao, M. Pelican, D. J. Musliner, J. Huang, J. Silliman, M. Mcdaniel, J. Casavant, I. Goldthwaite, N. Vidovich, M. Lehman, and T. Kim. SoK: DARPA’s AI cyber challenge (AIxCC): Competition design, architectures, and lessons learned. In USENIX Security Symposium (Security), 2026. arXiv:2602.07666.

[76] S. Woo, H. Hong, E. Choi, and H. Lee. MOVERY: A precise approach for modified vulnerable code clone discovery from modified Open-Source software components. In USENIX Security Symposium, pages 3037– 3053, 2022.

[87] T. Zhang, Y. Li, C. Yang, R. Widyasari, Y. Liu, N. T. Bui, P. T. Nguyen, Y. N. Tun, I. C. Irsan, H. H. Nguyen, H. Huang, J. Jiang, L. K. Shar, E. L. Ouh, D. Lo, H. J. Kang, Y. Yin, and W. B. Leow. TitanCA: Lessons from orchestrating LLM agents to discover 100+ CVEs. IEEE Security & Privacy, 2026. To appear; arXiv:2604.17860.

[77] S. Woo, E. Choi, H. Lee, and H. Oh. V1SCAN: Discovering 1-day vulnerabilities in reused C/C++ opensource software components using code classification techniques. In USENIX Security Symposium, pages 6541–6556, 2023. 19

[88] Z. Zheng, J. Zhou, X. Hu, Y. Gao, and S. Pan. Multiagent end-to-end vulnerability management for mitigating recurring vulnerabilities, 2026. URL https:// arxiv.org/abs/2601.17762. arXiv:2601.17762. [89] Y. Zhou, S. Liu, J. Siow, X. Du, and Y. Liu. Devign: Effective vulnerability identification by learning comprehensive program semantics via graph neural networks. In Advances in Neural Information Processing Systems, 2019. [90] Y. Zhou, L. Y. Choi, J. Wen, and W. Ye. Accuracy, stability, and repeated-run reliability of large language models on deterministic programming tasks, 2026. URL https://arxiv.org/abs/2606.00920. arXiv:2606.00920. [91] Y. Zhu, A. Kellermann, D. Bowman, P. Li, A. Gupta, A. Danda, R. Fang, C. Jensen, E. Ihli, J. Benn, J. Geronimo, A. Dhir, S. Rao, K. Yu, T. Stone, and D. Kang. CVE-Bench: A benchmark for AI agents’ ability to exploit real-world web application vulnerabilities, 2025. URL https://arxiv.org/abs/2503.17332. arXiv:2503.17332. [92] A. Zibaeirad, M. Vieira, and T. Zimmermann. AutoTrace: From patches to triggers via agentic interprocedural exploration, 2026. URL https://arxiv.org/ abs/2607.12058. arXiv:2607.12058.

20

Record · ID 660738 · SHA-256 0ab5cac9b6c4f349
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.