Conceptio › Archive › arXiv CS
arXiv CSopen access

SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

SWE-Test Technical Report

September 9, 2026

SWE-Test: Benchmarking LLM Vulnerability Discovery via Input Prediction Yuanxiang Shi* , Jiayi Lin* , Xuanyong Lin* , Liangcai Su, Yeheng Duan, Wei Wang, Qi Han, Bing Zhao, Wei Hu, Xander Xu† , Chenxiong Qian† * Equal contribution.

† Corresponding authors.

Website: https://swe-test-benchmark.github.io/ Dataset: https://github.com/SWE-Test-Benchmark/SWE-Test-Benchmark

Abstract

arXiv:2609.06229v1 [cs.SE] 5 Sep 2026

Vulnerability discovery is becoming an important ability of large language model (LLM) agents: agents that silently miss real defects leave critical software exposed. Rigorously measuring this ability is therefore urgent, but existing benchmarks are gameable through data contamination, score recall against an unknowable vulnerability set, often rely on synthetic bugs, and report a single end-to-end verdict that cannot localize where an agent fails. Vulnerability discovery is a composite ability: an agent must comprehend source code, infer input constraints, construct inputs, execute them, and iteratively correct from feedback. We recast its measurement as an input-prediction task with a closed, deterministic ground truth: using coverage-guided fuzzing, we mine deep target branches in real-world C/C++ programs and ask an agent to predict an input that drives execution to a given branch. This decomposes discovery into three task modes over 22 realworld C/C++ programs spanning 15 domains. Open-loop and Feedback-enabled share 60 fixed-target task instances across 16 of these codebases (13 domains), testing input construction without and with a distance oracle to isolate code comprehension from feedback-driven correction. Online Arena instead removes the predefined target and scores path exploration by coverage gain on a separate, partially overlapping pool of 11 programs; agents collectively confirmed 13 distinct bugs across six programs. Evaluating 15 default-effort model-scaffold configurations, the best reaches only 55.0% pass rate in the Feedback-enabled mode, and the mean across seven paired Claude Code configurations is 36.4% with feedback versus 19.3% without. Decomposing failures, we find constraint inference, not navigation, is the dominant bottleneck. We release SWE-Test with a turnkey evaluation environment. Keywords: Vulnerability Discovery, LLM Agents, Fuzzing, Software Testing, Benchmark

1

Introduction

Vulnerability discovery has become one of the defining capabilities of frontier coding agents. In early 2026, an AI agent autonomously uncovered a sixteen-year-old vulnerability in FFmpeg, a media library embedded in countless systems, on a line of code that automated fuzzers had executed more than five million times without ever flagging it, and separately chained several Linux kernel flaws into a privilegeescalation exploit Anthropic (2026a). Anthropic’s Frontier Red Team reports that current models can find meaningful zero-day vulnerabilities in well-tested codebases without any specialized harness, adding value on top of the fuzzing infrastructure that security teams have refined for years Anthropic (2026b). As agents move from writing code to auditing it, their capacity to surface security-relevant defects bears directly on the security of the software supply chain, which makes measuring that capacity, rigorously and at scale, an urgent problem. 1

Unfortunately, existing benchmarks used to measure it rest on assumptions that do not hold for realworld vulnerability discovery. First, security benchmarks built from public records can be gamed: their tasks are absorbed into model training and solved partly from memory. A recent analysis finds that on InterCode-CTF a frontier model has a solution-leakage rate of about 14%, attributed to the benchmark’s age and likely presence in the training corpus Peng et al. (2026); newer security benchmarks try to dodge this by using only post-cutoff challenges (CyBench, 2022 to 2024 Zhang et al. (2025b); CVE-Bench, May to June 2024 Zhu et al. (2025)), but that only delays the problem as cutoffs advance. Second, the set of vulnerabilities in a codebase is unknown and cannot be enumerated, so recall against a fixed list of known vulnerabilities is a biased metric. CyberGym, for example, asks the agent to reproduce one specific disclosed vulnerability and credits only that target Wang et al. (2026), so an agent that finds a different genuine defect scores nothing and the reported recall is not a meaningful fraction of an unknowable whole. Third, some benchmarks sidestep these difficulties with synthetic bugs, which are easy to grade but rarely match the constraint structure and cross-procedure depth of real defects. We observe that vulnerability discovery is a composite ability: an agent must comprehend source code, infer input constraints, construct inputs, execute them, and iteratively correct its approach from feedback. Building on this, we recast the measurement of an agent’s vulnerability-discovery ability as an inputprediction task, decomposing it into three underlying abilities: deep code comprehension, feedback-driven correction, and path exploration. To probe the first two, we use fuzzing to mine, from real programs, deep target branches together with the execution paths that reach them, and we ask the agent to generate an input that drives execution to a given branch. Reaching such a branch requires the agent to understand deep code, since the path from the program entry to the target typically spans multiple files and functions, and to reverse-reason from that target back to a satisfying input. Beyond the source, we then add an environment feedback signal that reports the agent’s approximate distance to the target, so the agent can continuously adjust its candidate and converge on an input that reaches the branch. Finally, to probe the third ability, we place the agent in a fuzzing process that never stops on its own and turn the single target branch into all branches: the agent now aims to maximize code coverage, with coverage itself serving as the feedback signal, and repeatedly generates new inputs to reach new branches by reasoning over the current coverage and the source. This formulation is hard to game, since the agent must synthesize an input that satisfies a branch identified by line numbers in a specific build, for which there is no public answer to memorize, and it needs no predefined vulnerability list: posed over real programs, it can trigger real bugs. Building on this design, we introduce SWE-Test, a benchmark with three modes. The Open-loop mode exposes deep code comprehension alone, providing only source and a target branch. The Feedback-enabled mode adds the distance oracle, exercising feedback-driven correction on the same targets. The Online Arena mode removes the target entirely and scores coverage gain under the never-stopping loop above. The Open-loop and Feedback-enabled modes span 13 domains and 16 real-world programs, including language runtimes, parsers, networking, font and text shaping, XML, cryptography, and compilers. In Online Arena, we run agents continuously for 12 hours. Across 15 default-effort model-scaffold configurations, the best feedback-enabled pass rate is 55.0%. Among the seven Claude Code configurations with 60 valid trials in both modes, feedback raises mean pass rate from 19.3% to 36.4%, target-function reach from 62.6% to 75.2%, and near-to-pass conversion from 30.8% to 48.4%. In Online Arena, the five evaluated configurations produce 18 model-bug detections, or 13 confirmed bugs across six programs after deduplication by upstream issue or fix. To make these measurements reproducible and easy to adopt, we release SWE-Test with a standardized, Harbor-based turnkey deployment and evaluation environment that provisions the containerized tasks, runs any supported model-scaffold configuration, and reports per-task and aggregate scores.

2

2

Related Work

Existing vulnerability-discovery benchmarks span several increasingly realistic settings, from reproducing disclosed vulnerabilities to open-ended discovery in real software. Table 1 compares CyberGym Wang et al. (2026), CVE-Bench Zhu et al. (2025), CyBench Zhang et al. (2025b), BountyBench Zhang et al. (2025a), AgentCyberRange Liu et al. (2026), and SWE-Test. A dash indicates that the feature is not the benchmark’s primary objective or is not reported in the available description. Table 1: Comparison with selected vulnerability-discovery benchmarks. Benchmark

Open Input Controlled AntiLong Real Predefined software target discovery grading feedback contam. horizon

CyberGym CVE-Bench CyBench BountyBench AgentCyberRange SWE-Test

✓ ✓ – ✓ ✓ ✓

✓ ✓ ✓ – – ✓

– – – ✓ ✓ ✓

– – – – – ✓

✓ ✓ ✓ ✓ ✓ ✓

– – – – – ✓

– – – – – ✓

The comparison highlights three recurring limitations. First, many CTF and issue-driven benchmarks rely on public, fixed tasks, making their results vulnerable to contamination and memorization. Live variants delay this risk rather than eliminate it. Transformed or synthetic tasks improve scalability and reduce direct memorization, but may not preserve the control-flow depth or input constraints of real multi-file software. Second, CVE- and PoC-oriented benchmarks evaluate the reproduction or exploitation of a predefined vulnerability. Although this design provides a clear execution target, it does not credit different defects discovered by an agent. Recall over known vulnerabilities cannot represent coverage of an unknowable vulnerability set. Third, end-to-end penetration-testing benchmarks combine discovery, exploitation, and post-exploitation into one outcome. This aggregation obscures whether failure arises from code comprehension, input construction, feedback use, or target selection. SWE-Test addresses these limitations by reframing vulnerability discovery as an input-prediction task. The agent receives the source of a real C/C++ program and must construct an input that exercises a specified decision outcome in the code. This formulation decomposes vulnerability-discovery ability into code comprehension, feedback-driven correction, and path exploration. Each target originates from an execution observed during prior automated testing, with a hidden input establishing both reachability and deterministic ground truth. This closed objective avoids evaluating recall against an inherently incomplete set of known vulnerabilities. Tying targets to specific builds reduces contamination, while withholding satisfying inputs limits answer memorization. The graded reward distinguishes failure to reach the target function from failure to satisfy the target branch, enabling precise diagnosis on real software with realistic input constraints.

3

SWE-Test Benchmark

This section presents SWE-Test as a concrete realization of the input-prediction formulation introduced in Section 1. Section 3.1 summarizes the dataset and its application domains. Section 3.2 defines the three target abilities, their corresponding task modes, and the associated sandbox and reward design. Section 3.3 describes how reachable targets are mined from real fuzzing executions and converted into contamination-resistant benchmark tasks.

3.1

Overview

Dataset. The fixed-target pool used by Open-loop and Feedback-enabled comprises 60 input-prediction tasks from 16 programs across 13 domains; Online Arena uses a separate, partially overlapping pool, 3

detailed below. Figure 1 summarizes the fixed-target distribution, spanning parsers, network protocols, text and binary formats, language runtimes, cryptography, and compilers. SWE-Test evaluates input prediction through three complementary modes. Open-loop requires the agent to construct an input for a fixed target branch using source code without runtime feedback, emphasizing code comprehension. Feedback-enabled retains the same target but provides a distance oracle, allowing the agent to refine candidate inputs through feedback. Online Arena removes the predefined target and requires the agent to select unexplored code paths using coverage information, thereby evaluating self-directed path exploration. Open-loop and Feedback-enabled draw their 60 tasks from the same fixed-target pool of 16 programs across 13 domains. Online Arena runs on a separate pool of 11 programs, since it needs a fuzzer-saturated corpus rather than a single mined branch (Section 3.3); five of these (HarfBuzz, libxml2, OpenSSL, PHP, QuickJS) overlap with the fixed-target pool, and six (CPython, libjpeg, libpcap, libpng, Lua, SQLite) are exclusive to Online Arena, adding image-codec (libjpeg, libpng) and database-engine (SQLite) domains not present in the fixed-target pool. In total, the benchmark spans 22 distinct C/C++ codebases across 15 domains.

3.2

Task Modes

Open-loop mode. The agent is given the target program’s source and a single branch condition that is currently not taken, and must produce an input that takes it. Figure 2 illustrates this task interface alongside its feedback-enabled counterpart. No execution oracle is provided: the agent cannot observe whether any candidate input moves it closer to the target. Success therefore requires the agent to trace data flow and construct a satisfying input by reasoning alone, a capability adjacent to static constraint inference and symbolic execution, performed mentally over the source. The reward is 2.0 when execution takes the target branch and 1.0 when it reaches the containing function without satisfying the branch condition. If execution does not reach the target function, it receives a trace-progress reward between 0.0 and 1.0, based on its overlap with the ground-truth execution path. A reward of 0.0 indicates no path overlap. Feedback-enabled mode. On the same targets, the agent may additionally run a verification oracle that reports how close the current input is to triggering the branch. Figure 2 shows the shared target and the additional verifier interface. This admits the construct-execute-observe-revise loop. Because the only difference from the Open-loop mode is the presence of feedback, the gap between the two modes isolates feedback-driven correction while holding comprehension fixed. The feedback-enabled mode uses the same reward levels and exposes the intermediate reward after each verifier query. Online Arena. No target branch is specified at all, so the distance oracle of the feedback-enabled mode no longer applies. Instead, the agent is given a coverage tool that reports the current coverage and the list of still-uncovered branches, and runs in a long-horizon loop: it queries coverage, reads per-line annotated source to find uncovered branches worth attacking, generates new inputs, and re-measures. Figure 3 summarizes this multi-turn loop, its isolated execution environment, and the independent verification stage. This is still feedback, but of a different kind, namely undirected coverage feedback rather than distance to a fixed target. The open-ended objective is broader than coverage: while reading uncovered code, the agent also hunts for defects, including crashes, hangs, state corruption, leaks, and functional or semantic misbehavior, and routes seeds associated with confirmed findings to a separate corpus. Coverage gain rather than a single branch flip remains the quantitative score, since the set of real vulnerabilities in a codebase is unknowable and recall against a fixed bug list would reintroduce the bias this benchmark avoids; discovered bugs are logged as artifacts. The step up from the feedback-enabled

4

URL parsing: 2 Cryptography: 1

Audio codec: 1

Query processing: 1

Font / text shaping: 11 SQL parsing: 11

13 Domains 60 Tasks

Smart-contract compiler: 1

XML / markup: 10

Computer algebra: 2

Language runtime: 10 JavaScript runtime: 1 Networking: 7

Binary analysis: 2

Figure 1: Composition of SWE-Test’s 60 fixed-target input-prediction tasks (Open-loop and Feedbackenabled modes), covering 16 real-world C/C++ programs across 13 application domains.

Figure 2: Task example contrasting Open-loop and Feedback-enabled modes on a shared target branch. mode isolates the added value of self-directed exploration: the agent must now decide what to pursue, not merely how to reach a given target. The Online Arena reward is the mean normalized gain in regions, lines, branches, and functions relative to the baseline corpus; confirmed bugs are reported separately. For each coverage dimension d, we compute   finald − baselined rewardd = clamp , 0, 1 , 100 − baselined and report the mean of the four rewardd values.

5

Figure 3: Online Arena execution and verification pipeline.

Figure 4: Construction pipeline for Open-loop and Feedback-enabled tasks.

3.3

Benchmark Construction

SWE-Test tasks are not hand-authored. All three modes are built from the same coverage-guided fuzzing runs, which guarantees that every target is real and reachable, but the two kinds of task consume those runs differently. Open-loop and Feedback-enabled tasks. Figure 4 summarizes the six-stage construction and evaluation pipeline. We instrument each target for LLVM source-based coverage, use AFL++ Fioraldi et al. (2020) to generate candidate inputs, and monitor successive coverage snapshots. When an input reaches a previously uncovered branch, we record the branch location, call trace, and concrete input as a task instance. An offline stage packages each instance with the program source and instructions in a selfcontained Harbor task. During evaluation, Harbor builds an isolated container in which the agent must write its candidate to /workspace/final_seed. Feedback-enabled tasks expose a restricted verifier that

6

returns only coarse progress, whereas Open-loop tasks provide no runtime feedback. After the agent finishes, the root-owned tests/test.sh evaluates the final input and records its score. The hidden input discovered during construction establishes that every target branch is reachable, while the newly covered status selects branches that were absent from the preceding corpus. Online Arena tasks. This mode needs no single branch, but it does need a starting point that is hard. We take the corpus a fuzzer has accumulated after 72 hours on a target and hand it to the agent as its starting seeds. The reason is how coverage grows: it rises steeply and then flattens, so the easily reachable branches are, empirically, already covered well within 72 hours. Starting from such a saturated corpus puts every agent at the same hard frontier, where the branches that remain are precisely those blind mutation could not reach. This reduces the agent budget spent rediscovering easy coverage and focuses the comparison on the harder frontier, because branches readily reached by fuzzing provide limited evidence for distinguishing agents’ path-exploration ability.

4

Experimental

4.1

Experimental Setup and Research Questions • RQ1 (Leaderboard). Which model/scaffold configuration exhibits the strongest vulnerabilitydiscovery ability, and how large is the performance gap across configurations? (Section 4.2) • RQ2 (Bottleneck diagnosis). Why do agents fail to translate their code-comprehension ability into successful task completion, and how do these failure patterns vary across modes and domains? (Section 4.3) • RQ3 (Evaluation sensitivity). How sensitive are measured vulnerability-discovery results to reasoning effort and scaffold choice? (Section 4.4)

Models, scaffolds, and protocol. A configuration pairs a language model with an agent scaffold managing tool use, context, and permissions. We evaluate vendor-specific scaffolds, namely Codex, Kimi Code, and Qwen Coder, and generic scaffolds, namely Claude Code and Terminus 2; unless otherwise noted, controlled comparisons use a fixed scaffold, with Terminus 2 treated as an ablation (Section 4.4). We evaluate seven models from four families: DeepSeek-V4-Pro from DeepSeek DeepSeek-AI (2026); GLM5.1 and GLM-5.2 from GLM Z.AI (2026a;b); Kimi-K2.6 and Kimi-K3 from Kimi Moonshot AI (2026); and Qwen3.7-Max and Qwen3.8-Max from Qwen Qwen Team (2026). All tasks run in isolated Harbor containers; Open-loop and Feedback-enabled tasks receive a 24-hour budget, one CPU, 8 GB memory, and network access, scoring runs separately, exposing only the numeric reward. Models use default reasoning effort unless otherwise noted (Section 4.4), and each configuration receives one trial per task. Pass rates use only valid trials, excluding empty scores. Table 2 summarizes the evaluated configurations and pass rates in both fixed-target modes. Table 2: Feedback-enabled/Open-loop pass rates by model and scaffold. Claude Code Terminus 2

Codex Kimi Code Qwen Coder

Qwen3.8-Max DeepSeek-V4-Pro GLM-5.1 GLM-5.2 Kimi-K2.6 Kimi-K3 Qwen3.7-Max

55.0/50.0% 35.0/8.3% 23.3/1.7% 41.7/26.7% 35.0/8.3% 35.0/25.0% 30.0/15.0%

– – 20.7/2.0% – 1.7/1.7% – 10.0/6.9% – 3.7/3.6% – – – 0.0/1.7% 1.7/3.3%

– – – – 16.7/3.8% – –

– – – – – – 13.3/6.7%

Mean

36.4/19.3%

7.2/3.2% 1.7/3.3%

16.7/3.8%

13.3/6.7%

7

4.2

Which Configuration Performs Best?

The clearest sign of vulnerability-discovery capability is whether an agent finds real, confirmed bugs given time to explore, not whether it flips a single pre-selected branch under a fixed evaluation budget. We evaluate using three complementary signals: bug count and normalized coverage reward under an extended Online Arena budget, and pass rate under the standard feedback-enabled mode. Tables 3 and 4 report the two Online Arena signals before the controlled fixed-target comparison. Table 3: Normalized coverage reward in 12-hour Online mode. Program

Qwen3.8Max-Preview

KimiK2.6

GLM5.2

DeepSeekV4-Pro

Qwen3.7Max

CPython HarfBuzz libjpeg libpcap libpng libxml2 Lua OpenSSL PHP QuickJS SQLite

0.22 0.22 0.14 0.60 0.27 0.19 0.30 0.15 0.16 0.49 0.71

0.21 0.11 0.14 0.60 0.27 0.19 0.30 0.17 0.17 0.42 0.69

0.23 0.19 0.14 0.59 0.27 0.19 0.23 0.16 0.13 0.36 0.70

0.33 0.20 0.22 0.58 0.26 0.16 0.20 0.14 0.08 0.35 0.58

0.19 0.18 0.14 0.56 0.27 0.18 0.23 0.14 0.07 0.22 0.57

0.3136

0.2977

0.2900

0.2818

0.2500

Mean

Table 4: Per-model counts of confirmed bugs from 12-hour Online runs. Program

Qwen3.8Max-Preview

KimiK2.6

GLM5.2

DeepSeekV4-Pro

Qwen3.7Max

CPython HarfBuzz libjpeg libpcap libpng libxml2 Lua OpenSSL PHP QuickJS SQLite

2 1 0 1 0 0 0 0 0 3 0

0 1 0 0 0 0 0 0 0 3 0

0 0 0 0 1 0 0 0 0 2 2

0 1 0 0 0 0 0 0 0 1 0

0 0 0 0 0 0 0 0 0 0 0

Total

7

4

5

2

0

Bug counts under the Online Arena. For the five configurations evaluated in Online Arena, we extend the per-program budget to 12 hours of wall-clock time and count the manually confirmed bug identities discovered by each model. Table 4 reports the result. We evaluate five models in 12-hour Online runs: Qwen3.8-Max-Preview, Kimi-K2.6, GLM-5.2, DeepSeek-V4-Pro, and Qwen3.7-Max. We manually validate the reported findings. Table 4 reports per-model counts: a bug independently found by multiple models contributes once to each corresponding model column. Across all five models, the 18 model-bug detections reduce to 13 distinct bugs in six programs after matching duplicate findings by their upstream issue or fix. Running every configuration for 12 hours per program is expensive, so as a complementary signal we also report normalized composite coverage reward from the same 12-hour Online runs in Table 3. Qwen3.8-Max-Preview leads both summaries, with seven confirmed bugs and a mean coverage reward

8

of 0.3136. Overall, the bug counts broadly follow the coverage-reward trend; Kimi-K2.6 and GLM-5.2 exchange positions with only a small difference in mean coverage reward (0.2977 vs. 0.2900). Leaderboard under the feedback-enabled mode. Online Arena is expensive to run at scale, so we also compare the seven Claude Code configurations with 60 valid trials in both modes. Table 2 summarizes the evaluated configurations and pass rates in both fixed-target modes. Qwen3.8-Max leads the Feedbackenabled mode at 55.0%, followed by GLM-5.2 at 41.7%, and DeepSeek-V4-Pro, Kimi-K2.6, and Kimi-K3 at 35.0%. Why feedback matters. Across the seven paired Claude Code configurations, feedback raises mean pass rate from 19.3% to 36.4% and improves every configuration. The corresponding gains in targetfunction reach and near-to-pass conversion indicate that feedback supports both code-path navigation and branch-constraint refinement. Answer to RQ1. Under a 12-hour Online Arena budget, Qwen3.8-Max-Preview records the most confirmed bugs among the five evaluated models (seven) and the highest mean coverage reward (0.3136). Across models, these runs yield 13 distinct confirmed bugs in six programs after deduplication. In the separate feedback-enabled evaluation, Qwen3.8-Max + Claude Code leads at 55.0%. RQ2 next asks what separates the leaders from the rest.

4.3

Why Do Agents Fail?

We bucket every trial into a failure stratum using the oracle’s reward: navigation failure (r <1.0, never reached the target function; this includes hard failures, no-op resubmissions of the seed, and partial credit for how far the call trace progressed toward the function), constraint-satisfaction failure (r =1.0, reached the function but did not satisfy the target branch constraint), and pass (r =2.0). Failure is dominated by constraint satisfaction, not navigation. Table 5 reports three outcomes for the seven Claude Code configurations with 60 valid trials in both modes. Under feedback, constraintsatisfaction failures generally outnumber navigation failures: agents reach the right function more often than they satisfy its target branch constraint. The bottleneck is therefore not only locating the relevant code but satisfying the constraint that guards the target branch once found. Feedback improves both stages. Among trials that reached the target function, the fraction that converted to a pass (near-to-pass conversion) is 48.4% with the verifier and 30.8% without it. Feedback also increases the fraction of trials that reach the target function, from 62.6% to 75.2%. It therefore improves both navigation and target-constraint satisfaction rather than acting on only one stage. Figure 5 visualizes the outcome distributions. Navigation failure 24.8

Feedback-enabled

Constraint failure

38.8

37.4

Open-loop 0

36.4 43.3

25

Pass

50

19.3 75

100%

Percentage

Figure 5: Navigation and target-constraint satisfaction by mode.

9

Table 5: Outcome counts under Claude Code on identical targets (n = 60). Model

Nav.

Constr.

Pass

Feedback-enabled Qwen3.8-Max DeepSeek-V4-Pro GLM-5.1 GLM-5.2 Kimi-K2.6 Kimi-K3 Qwen3.7-Max

10 12 11 10 19 29 13

17 27 35 25 20 10 29

33 21 14 25 21 21 18

Open-loop Qwen3.8-Max DeepSeek-V4-Pro GLM-5.1 GLM-5.2 Kimi-K2.6 Kimi-K3 Qwen3.7-Max

13 29 30 23 26 18 18

17 26 29 21 29 27 33

30 5 1 16 5 15 9

Domain modulates the bottleneck. The dominant failure mode also varies by task family (Table 6, which pools the Feedback-enabled trials of the seven paired Claude Code configurations). SQL-parser tasks show 82% navigation failure, whereas PHP tasks show 17%. These differences suggest that input structure and source-to-harness mapping influence whether agents reach the target before solving its branch constraint. Table 6: Failure outcomes by task family (paired Claude Code configurations). Task family

Nav. fail (%)

Constr. fail (%)

Pass (%)

0 14 14 14 16 0 6 0 14 12 17 7 29 29 82 0

100 86 14 29 77 100 33 0 14 31 23 43 28 71 12 0

0 0 72 57 7 0 61 100 72 57 60 50 43 0 6 100

Ada URL parsing Bloaty cURL FreeType HarfBuzz JQ libxml2 mBedTLS OpenSSL OpenThread PHP Proj4 QuickJS Solidity SQL parser Vorbis

Answer to RQ2. Constraint-satisfaction failures are the largest aggregate outcome in both modes, but navigation also matters. Removing feedback reduces target-function reach from 75.2% to 62.6% and near-to-pass conversion from 48.4% to 30.8%, indicating that feedback supports both stages. The trajectory audit supports this diagnosis: with feedback, successful runs use intermediate rewards to revise candidate-specific hypotheses, whereas open-loop runs can stop after an undetected near miss (r =1.0). Passing trajectories also sustain longer interaction, but tool volume alone does not explain success; these observations are descriptive because only a subset of jobs has complete readable trajectories.

10

4.4

How Sensitive Are Results to Agent Configuration?

We isolate two evaluation choices, reasoning effort and agent scaffold, through targeted ablations. Each varies one factor while holding the remaining configuration fixed. Reasoning Effort We vary the reasoning effort of GLM-5.1 under Claude Code across low, medium, and high settings in both modes (Table 7). In the Feedback-enabled mode, the pass rate rises monotonically, from 25.0% at low to 28.3% at medium and 38.3% at high. In the Open-loop mode, high achieves 18.3%, compared with 6.7% at low and 3.3% at medium. All effort settings complete 60 valid trials in each mode. Table 7: Reasoning-effort ablation; n denotes valid trials. n

Nav.

Constr.

Pass

Rate

Feedback-enabled low 60 medium 60 high 60

13 10 14

32 33 23

15 17 23

25.0% 28.3% 38.3%

Open-loop low medium high

26 27 23

30 31 26

4 2 11

6.7% 3.3% 18.3%

Effort

60 60 60

High effort changes the failure distribution differently across modes. With feedback, it produces 14 navigation failures, compared with 10 or 13 at the other settings, while increasing the number of passes to 23. Without feedback, it produces 11 passes, compared with four at low and two at medium. Reasoning effort is therefore not monotonically associated with performance across modes; its effect depends on the available feedback and should be interpreted cautiously. Scaffold Choice: Vendor vs. Neutral We compare two models under Terminus 2 and their vendor scaffolds in the Feedback-enabled mode (Table 8): Kimi Code for Kimi-K2.6 and Qwen Coder for Qwen3.7-Max. Kimi-K2.6 achieves 16.7% under Kimi Code and 3.7% under Terminus 2, while Qwen3.7-Max achieves 13.3% under Qwen Coder and 0.0% under Terminus 2. The consistent direction shows that scaffold choice affects measured performance for these models, but two comparisons are insufficient for broader generalization. Table 8: Pass rates by scaffold. Model

Terminus 2

Vendor scaffold

3.7% 0.0%

16.7% 13.3%

Kimi-K2.6 Qwen3.7-Max

We further audit the recorded trajectories (Table 9), reporting median tool calls and verifier queries per available trajectory; this analysis is descriptive because some jobs lack complete trajectories. Table 9: Trajectory activity by scaffold. Model

Scaffold

Kimi-K2.6

Kimi Code Terminus 2 Qwen Coder Terminus 2

Qwen3.7-Max

Traj.

Tool calls

Verifier queries

60 54 57 60

68.0 242.5 76.0 54.5

4.0 25.5 9.0 4.0

11

Tool volume alone does not explain the scaffold gap: for Kimi-K2.6, Terminus 2 makes more tool calls and verifier queries than Kimi Code, yet achieves a lower pass rate. The vendor scaffolds yield 10 and 8 branch-satisfying inputs for Kimi-K2.6 and Qwen3.7-Max, compared with 1 and 0 under Terminus 2. Thus, scaffold design affects how interactions are converted into target-branch satisfaction, not merely how many actions an agent performs. Measured performance is sensitive to both reasoning effort and scaffold choice: the effort ranking changes with feedback availability, while vendor scaffolds outperform Terminus 2 in both evaluated comparisons. Model comparisons should therefore control and report the task mode, reasoning effort, and scaffold; the scaffold result remains limited to two models.

5

Discussion and Threats to Validity

5.1

Discussion

The results identify two distinct bottlenecks: reaching the target function and satisfying its branch constraint. Feedback improves both target-function reach (62.6% to 75.2%) and near-to-pass conversion (30.8% to 48.4%), while model and scaffold choices change the magnitude of the pass-rate difference. Evaluation should therefore report both stages rather than attributing all failures to either navigation or refinement alone. Constraint-satisfaction failures remain the largest aggregate outcome, motivating feedback that identifies which sub-condition of a guard remains unsatisfied. At the same time, the navigation gap shows that feedback must also help agents reach the relevant code. Our scaffold and reasoning-effort findings (Section 4.4) further show that leaderboards should disclose the harness, mode, and effort setting before ranking models.

5.2

Threats to Validity

SWE-Test measures vulnerability discovery through single branch inversion rather than full exploit construction. We chose this target because it is deterministically checkable and can be pinpointed to one of the three abilities, and the four-level reward separates reaching the target function from satisfying the target branch constraint rather than collapsing both into pass/fail; still, it is a proxy, so results describe the comprehension, correction, and exploration stack rather than end-to-end exploit construction. In its fixed-target modes, the benchmark covers 60 tasks over 16 C/C++ OSS-Fuzz Serebryany (2017) targets and 13 domains, concentrated in a few high-task-count domains such as PHP, SQL, XML, and fonts, so findings may not transfer to other languages or input formats; Online Arena adds 6 further programs and 2 further domains not in this fixed-target pool, for 22 distinct programs and 15 domains benchmark-wide (Section 3.1). Every task is built from already-known, fuzzer-discovered branches in public projects, so the released artifacts contain no novel exploit or undisclosed vulnerability. Two factors limit how directly we can compare models. Scaffold choice can materially change pass rate, so rankings must hold the scaffold fixed. Models also use default reasoning effort rather than a matched level, and Section 4.4 shows that the effort settings shift GLM-5.1’s feedback-enabled pass rate from 25.0% to 38.3%. Finally, some configurations have fewer than 60 valid trials; all reported pass rates exclude unscored trials, so comparisons with different denominators should be interpreted cautiously.

6

Conclusion

We introduce SWE-Test, a benchmark that evaluates vulnerability-discovery agents through input prediction on real C/C++ programs. Its Open-loop, Feedback-enabled, and Online Arena modes separate code comprehension, feedback-driven correction, and path exploration. Across 15 default-effort modelscaffold configurations, the best feedback-enabled pass rate is 55.0%. Among seven paired Claude Code configurations, feedback raises the mean pass rate from 19.3% to 36.4% and improves both target-function

12

reach and target-constraint satisfaction on average. Scaffold choice and reasoning effort also materially affect measured performance, making uncontrolled model rankings difficult to interpret. We release 60 Open-loop and Feedback-enabled tasks across 16 programs, the Online Arena, and the evaluation harness to support reproducible analysis of where vulnerability-discovery agents fail.

References Anthropic. Project glasswing. https://www.anthropic.com/glasswing, 2026a. Accessed: 2026-06-30. Anthropic. Evaluating and mitigating the growing risk of llm-discovered 0-days. https://www.anthropic. com/research/zero-days, 2026b. Accessed: 2026-06-30. DeepSeek-AI. DeepSeek V4 Preview Release. https://api-docs.deepseek.com/news/news260424/, 2026. Published: 2026-04-24; accessed: 2026-07-29. Andrea Fioraldi, Dominik Maier, Heiko Eißfeldt, and Marc Heuse. AFL++: Combining incremental steps of fuzzing research. In Proceedings of the 14th USENIX Workshop on Offensive Technologies (WOOT ’20). USENIX Association, 2020. URL https://www.usenix.org/conference/woot20/presentation/fioraldi. Fengyu Liu, Jiarun Dai, Yihe Fan, Wuyuao Mai, Ziao Li, Bofei Chen, Jie Zhang, Zheng Lou, Bocheng Xiang, Qiyi Zhang, Xudong Pan, Geng Hong, Yuan Zhang, and Min Yang. Agentcyberrange: Benchmarking frontier AI systems in realistic cyber ranges. arXiv preprint arXiv:2606.14295, 2026. Moonshot AI. Kimi. https://www.kimi.com/, 2026. Accessed: 2026-07-29. Jiaren Peng, Zeqin Li, Chang You, Yan Wang, Hanlin Sun, Xuan Tian, Shuqiao Zhang, Junyi Liu, Jianguo Zhao, Renyang Liu, Haoran Ou, Yuqiang Sun, Jiancheng Zhang, Yutong Jiao, Kunshu Song, Chao Zhang, Fan Shi, Hongda Sun, Rui Yan, and Cheng Huang. Hackers or hallucinators? a comprehensive analysis of llm-based automated penetration testing, 2026. URL https://arxiv.org/abs/2604.05719. Qwen Team. Text generation models. https://docs.qwencloud.com/developer-guides/getting-started/ text-generation-models, 2026. Accessed: 2026-07-29. Kostya Serebryany. OSS-Fuzz: Google’s continuous fuzzing service for open source software. Presentation at the 26th USENIX Security Symposium, 2017. URL https://www.usenix.org/conference/ usenixsecurity17/technical-sessions/presentation/serebryany. Zhun Wang, Tianneng Shi, Jingxuan He, Matthew Cai, Jialin Zhang, and Dawn Song. Cybergym: Evaluating ai agents’ real-world cybersecurity capabilities at scale, 2026. URL https://arxiv.org/abs/ 2506.02548. Z.AI. GLM-5.1. https://docs.z.ai/guides/llm/glm-5.1, 2026a. Released: 2026-04-07; accessed: 2026-07-29. Z.AI. GLM-5.2: Built for long-horizon tasks. https://z.ai/blog/glm-5.2, 2026b. Published: 2026-06-16; accessed: 2026-07-29. Andy K. Zhang, Joey Ji, Celeste Menders, Riya Dulepet, Thomas Qin, Ron Y. Wang, Junrong Wu, Kyleen Liao, Jiliang Li, Jinghan Hu, Sara Hong, Nardos Demilew, Shivatmica Murgai, Jason Tran, Nishka Kacheria, Ethan Ho, Denis Liu, Lauren McLane, Olivia Bruvik, Dai-Rong Han, Seungwoo Kim, Akhil Vyas, Cuiyuanxiu Chen, Ryan Li, Weiran Xu, Jonathan Z. Ye, Prerit Choudhary, Siddharth M. Bhatia, Vikram Sivashankar, Yuxuan Bao, Dawn Song, Dan Boneh, Daniel E. Ho, and Percy Liang. Bountybench: Dollar impact of AI agent attackers and defenders on real-world cybersecurity systems. arXiv preprint arXiv:2505.15216, 2025a. Andy K. Zhang, Neil Perry, Riya Dulepet, Joey Ji, Celeste Menders, Justin W. Lin, Eliot Jones, Gashon Hussein, Samantha Liu, Donovan Jasper, Pura Peetathawatchai, Ari Glenn, Vikram Sivashankar, Daniel

13

Zamoshchin, Leo Glikbarg, Derek Askaryar, Mike Yang, Teddy Zhang, Rishi Alluri, Nathan Tran, Rinnara Sangpisit, Polycarpos Yiorkadjis, Kenny Osele, Gautham Raghupathi, Dan Boneh, Daniel E. Ho, and Percy Liang. Cybench: A framework for evaluating cybersecurity capabilities and risks of language models, 2025b. URL https://arxiv.org/abs/2408.08926. Yuxuan Zhu, Antony Kellermann, Dylan Bowman, Philip Li, Akul Gupta, Adarsh Danda, Richard Fang, Conner Jensen, Eric Ihli, Jason Benn, Jet Geronimo, Avi Dhir, Sudhit Rao, Kaicheng Yu, Twm Stone, and Daniel Kang. Cve-bench: A benchmark for ai agents’ ability to exploit real-world web application vulnerabilities, 2025. URL https://arxiv.org/abs/2503.17332.

14

Record · ID 668140 · SHA-256 81860e7213b13db8
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.