arXiv:2605.14431v1 [cs.SE] 14 May 2026
FuzzAgent: Multi-Agent System for Evolutionary Library Fuzzing Yunlong Lyu
Peng Chen
Fengyi Wu
The University of Hong Kong [email protected]
Independent Researcher [email protected]
Southeast University [email protected]
Junzhe Yu
Kit Long Hon
Hao Chen
The University of Hong Kong [email protected]
The University of Hong Kong [email protected]
The University of Hong Kong [email protected]
Abstract— Library fuzzing is essential for hardening the software supply chain, but adopting it at scale remains expensive. Practitioners still spend substantial effort on environment setup, struggle to generate harnesses that respect intricate API constraints, and lack reliable means to tell genuine library bugs from harness-induced crashes. Recent LLM-based systems automate parts of this pipeline, yet they typically operate as one-shot code generators that ignore runtime feedback, which limits both the depth of code they reach and the validity of the bugs they report. We argue that effective library fuzzing is iterative by nature: each campaign exposes new coverage bottlenecks and crashes, and the next campaign should evolve from these signals rather than restart from scratch. Building on this insight, we present FuzzAgent, a multi-agent system that turns library fuzzing into an evolutionary process, in which a team of specialized agents collaborates over the full fuzzing lifecycle and grounds every decision in concrete runtime evidence, so that the harness suite is successively refined toward deeper coverage and higher-fidelity crash analysis across rounds. We evaluate FuzzAgent on 20 real-world C/C++ libraries against four state-of-the-art baselines (OSS-Fuzz, OSS-Fuzz-Gen, PromptFuzz, and PromeFuzz). FuzzAgent completes the full fuzzing lifecycle for all 20 libraries without human intervention and reaches 179,619 branches, exceeding OSS-Fuzz, PromptFuzz, PromeFuzz, and OSS-Fuzz-Gen by 45.1%, 73.2%, 92.1%, and 191.2%, respectively. FuzzAgent also identifies 102 genuine library bugs, 78 of which have already been acknowledged and fixed by upstream maintainers.
I. I NTRODUCTION Fuzzing is a widely adopted automated software testing technique. It generates and executes a large number of random inputs to uncover vulnerabilities in software systems [1], [2], [3], [4]. Over the years, fuzzing has advanced substantially through coverage guidance [5], [6], grammar-based input generation [7], [8], and hybrid techniques that combine fuzzing with symbolic analysis [9], [10]. As of 2025, the continuous fuzzing platform OSS-Fuzz [11] has reported over 50,000 bugs in open-source software. With community maintenance and integration into development workflows, OSS-Fuzz helps projects routinely detect regressions and improve software security and reliability. As fuzzing matures, improving results by only generating new inputs for well-tested targets becomes increasingly difficult,
because further code coverage gains are often blocked [12]. This has motivated growing interest in library fuzzing, which focuses on synthesizing new fuzz targets (harnesses) rather than only mutating inputs [13], [14], [15], [16], [17], [18], [19], [20], [21], [22]. Library fuzzing aims to broaden test coverage by analyzing a library and composing Application Programming Interface (API) invocations into new targets. Such targets can expose vulnerabilities that are missed by existing, manually integrated harnesses. For example, PromptFuzz [20] achieves over 60% higher code coverage than OSS-Fuzz on the same set of libraries. Library fuzzing has evolved from manual harness development [11], [23] to a spectrum of approaches that increasingly automate harness generation [13], [14], [15], [16], [17], [18], [19], [20], [21]. Traditional approaches (e.g., LibFuzzer [23] and OSS-Fuzz [11]) relied heavily on human expertise to write harnesses and configure fuzzing environments. In contrast, more recent work aims to reduce this burden by automatically synthesizing API-invocation sequences into fuzzing harnesses. Consumer-based approaches (e.g., Fudge [13], FuzzGen [14], APICraft [16], and UTOpia [24]) primarily utilize static and dynamic analysis techniques. They extract API usage patterns from existing codebases, enabling the generation of harnesses that reflect real-world usage scenarios. To eliminate reliance on existing codebases, producer-based approaches (e.g., GraphFuzz [17], AFGen [19], Nexzzer [25], and Hopper [18]) focus on generating harnesses from scratch. They analyze API specifications to explore possible API combinations. Recently, LLM-based approaches (e.g., OSS-Fuzz-Gen [26], CKGFuzzer [27], PromptFuzz [20], and PromeFuzz [22]) have achieved notable progress by leveraging the code comprehension capabilities of LLMs. Unlike rule-based approaches, LLMs can reason about API semantics, infer parameter constraints, and produce valid harnesses that respect the calling conventions. Despite these advances in harness generation, deploying library fuzzing at scale remains challenging: executing existing fuzzing tools still requires substantial manual effort for environment setup, struggles to produce valid harnesses efficiently, and provides limited support for crash validation.
interpret than those from standalone programs or OSS-Fuzz integrated targets. They can be triggered by invalid API sequences or unmet preconditions in generated harnesses rather than true defects in the library implementation [39], [33]. Determining whether a crash reflects a real vulnerability therefore requires reasoning about library semantics and execution context, often through step-by-step debugging with tools such as GDB or Valgrind, which is time-consuming and error-prone [32], [33]. Many library fuzzing approaches [13], [14], [15], [16], [17], [18], [19], [20], [21], [22] provide limited support for crash validation, which can lead to a high rate of false positives. Current crash filters are often based on coarse heuristics and lack sufficient semantic context [18], [20], [22]. For example, Hopper [18] labels crashes outside inferred constraints as false positives, PromptFuzz [20] relies on short executions that may miss issues that manifest later, and PromeFuzz [22] asks LLMs to judge crashes from stack traces and API summaries with limited runtime information.
Environment setup and configuration remain laborintensive [28], [29], [30], [31], [32], [33]. In practice, developers must get many components to work together, including instrumentation [23], [34], [35], library builds [36], dictionary preparation [37], seed corpus construction [38], harness development, fuzzer execution, and result triage. Zhao et al. [33] report that even experienced developers struggle with configuration tasks such as compiling libraries, constructing dictionaries, and preparing seed corpora and harnesses, often requiring multiple rounds of manual adjustments. In a user study with 32 CS students and 6 Capture the Flag (CTF) players [29], only two participants completed the setup for a small library (120k lines of code) within 20 hours, and none succeeded for a complex library (600k lines of code). Even with the guidelines and templates provided by OSS-Fuzz [11], onboarding a new library is often fragile as dependencies and build tooling evolve over time [32], [31]. Nourry et al. [32] found that 82 of 171 OSS-Fuzz issues were related to build failures, and Plöger et al. [31] report that 53% of participants failed at the build stage when setting up fuzzing for new libraries. The challenge is amplified when adopting newer fuzzers that require custom builds, instrumentation, and configuration, pushing users into repeated trial-and-error iterations [33]. Even recent systems [18], [20], [22] still require manual effort to write custom build scripts and prepare environments; in our evaluation (Section V-A), integrating a new library into PromptFuzz and PromeFuzz took 8 hours and 12 hours on average, respectively. Designing effective fuzzing harnesses is non-trivial. A harness is a small driver program that maps the fuzzer-generated byte stream into API arguments and executes a sequence of library calls. The chosen APIs, their call order, and how inputs are mapped to arguments largely determine which internal checks and code paths can be reached. However, modern libraries often expose many APIs with intricate interdependencies and strict parameter constraints. Harnesses that violate these constraints either fail early and explore only shallow paths or trigger false positive crashes [18], [20]. Producing valid and coverage-effective harnesses therefore requires capturing nontrivial relationships among APIs and their parameters, which often requires substantial domain knowledge [39], [33]. Stateof-the-art systems such as Hopper [18], PromptFuzz [20], and PromeFuzz [22] show that automated harness generation can improve coverage, but they often rely on attempt-intensive heuristics to obtain valid and effective harnesses. Hopper [18] uses rule-based inference of API constraints, yet it can miss higher-order interactions among APIs. PromptFuzz [20] mutates API compositions to prompt LLMs and then filters outputs, which can yield many invalid harnesses and shallow exploration. Although PromeFuzz [22] pre-generates a full codebase summary to supply context, but this O(n2 ) summarization becomes prohibitively costly on large libraries, and the resulting context can be redundant or noisy since LLMs are continuouly upgrading and already extensively pretrained on open-source code [40], [41], [42], [43], [44]. Crash reports from library fuzzing are often harder to
Recent advances in large language models (LLMs) and multi-agent systems offer a new path toward addressing these challenges. LLMs have demonstrated remarkable capabilities in code generation, program analysis, and decision-making, while multi-agent systems provide a principled way to decompose complex tasks into smaller, manageable components and coordinate them step by step [45], [46], [47]. Inspired by these advances, we propose FuzzAgent, a novel multi-agent approach to evolutionarily solve bottlenecks encountered in library fuzzing. The core insight of FuzzAgent is to transition LLM-based library fuzzing from open-loop code generators to closed-loop reasoning agents, which can iteratively refine their actions based on feedback from the library fuzzing lifecycle. By equipping them with specialized interfaces, we bridge the gap between static analysis and dynamic fuzzing feedback, allowing the model to ground its generative capabilities in concrete runtime evidence and iteratively refine its inputs based on execution states. Specifically, 1) We design a set of library-fuzzing-specific interfaces and a multi-agent architecture that gives LLMs the ability to gather comprehensive, runtimegrounded information and take concrete actions throughout the full fuzzing lifecycle (Section III). 2) Building on this foundation, we introduce an evolutionary, agent-driven strategy that closes the loop between execution feedback and harness refinement, allowing FuzzAgent to progressively test deeper library code in a fully automated, end-to-end manner (Section IV). We evaluate FuzzAgent on 20 real-world C/C++ libraries against four state-of-the-art baselines (OSS-Fuzz, OSS-Fuzz-Gen, PromptFuzz, and PromeFuzz). FuzzAgent completes the full fuzzing lifecycle for all 20 libraries without any human intervention while achieving 179,619 branch coverage, surpassing OSS-Fuzz, PromptFuzz, PromeFuzz, and OSS-Fuzz-Gen by 45.1%, 73.2%, 92.1%, and 191.2%, respectively. In terms of bug detection, FuzzAgent identifies 102 genuine library bugs, of which 78 have been acknowledged and fixed by upstream maintainers.
2
II. BACKGROUND
metrics and anomalous behaviors, such as crashes, memory violations, assertion failures, or timeouts. 5) Crash Validation: Once crashes or anomalies are detected, users need to analyze the results and identify the root cause. This phase often requires manual inspection and iterative debugging using tools like GDB or Valgrind to triage and confirm the issues. Each of these phases is non-trivial and demands substantial manual effort and domain expertise. Take the instrumention phase as an example, participants in the study by Zhao et al. [33] reported that the overall process is overly complex, noting that “they treat configuration as an iterative, trial-anderror process: run the tool, observe what breaks, adjust the harness or flags, and try again.” As shown in the example build script in Figure 1, users must not only understand how to build the library itself (lines 9-16) but also carefully handle instrumentation flags that may conflict with the library configuration (lines 1-7). In this case, users must account for project-specific build flags; otherwise, the build will fail because the -Wl,-z,defs flag enabled in libaom conflicts with ASAN [34] and MSAN [48]. Furthermore, the complexity of the fuzzing setup increases significantly when attempting to use new fuzzers, as many require custom compilers or instrumentation flags. As observed in the study, “these extra steps may fail if you enabled certain optimization flags, or even fail by themselves, since open-source and legacy programs can be surprisingly fragile” [33]. To address these challenges, our work focuses on automating the entire library fuzzing workflow, thereby reducing the manual effort and expertise required.
1 extra_cmake_flags=’’ 2 if [[ $CFLAGS = *sanitize=memory* ]]; then 3 extra_cmake_flags+="-DAOM_TARGET_CPU= generic" 4 fi 5 if [[ $CFLAGS = *sanitize=address* ]]; then 6 extra_cmake_flags+=" -DSANITIZE=address" 7 fi 8 9 cmake -DCMAKE_INSTALL_PREFIX="$WORK" \ 10 -DCMAKE_C_COMPILER="$CC" \ 11 -DCMAKE_CXX_COMPILER="$CXX" \ 12 -DCONFIG_AV1_ENCODER=1 \ 13 -DCONFIG_AV1_DECODER=1 \ 14 ... 15 ${extra_cmake_flags} \ 16 "$SRC" 17 18 make -j$(nproc) 19 make install Fig. 1. An example build script snippet for compiling libaom with fuzzer instrumentation.
A. Library Fuzzing Workflow Library fuzzing follows the standard fuzzing workflow but requires an additional step: writing a harness to call library APIs and mitigate the false positives. In practice, setup typically involves the following phases [33], [29], [31], [28]: 1) Building and Instrumentation: The target must be compiled with instrumentation to support (i) feedbackdriven fuzzing (e.g., LibFuzzer [23] and AFL [5]) and (ii) bug detection (e.g., ASAN [34], UBSAN [35], and MSAN [48]). This step often requires specialized compiler flags, nontrivial build configurations, and careful dependency management. 2) Dictionary and Seed Preparation: To bootstrap a fuzzing campaign, users often prepare token dictionaries [49], [37] and seed corpus [38], [50]. These artifacts guide mutations toward format-aware inputs and improve early exploration, but producing high-quality dictionaries and seeds typically requires domain knowledge of the expected input structure. 3) Harness Generation: Library fuzzing requires a harness that drives the library through its exposed APIs. In practice, libraries may provide dozens or hundreds of interdependent functions, so an effective harness must respect API preconditions, object lifetimes, and calling sequences. Writing such harnesses is time-consuming and error-prone, and it often requires deep understanding of the target library[18], [25], [20], [22] . 4) Fuzzer Execution: With the prepared harnesses, dictionaries, and seed inputs, users can start the fuzzing campaign using their chosen fuzzing engines [51], [23]. During execution, users should monitor code coverage
Tasks
Actions
Execution Environment
Tools Agents Observation
Fig. 2. Multi-agent System
B. Multi-Agent Systems Multi-agent systems are computational frameworks where multiple autonomous agents interact to solve problems that are difficult for individual agents to tackle alone [45]. Recent advances in Large Language Models (LLMs) have significantly enhanced the capabilities of these systems by enabling agents to understand complex contexts, generate sophisticated responses, and make nuanced decisions [52]. These systems are particularly effective for complex tasks that benefit from
3
Agent Pool
Fuzzing Setup Agents
Actions
Interfaces
Library Dictionary Seed Builder Generator Generator
Setup
Computer Use Interface
De leg a
te
Call
Fuzzer-specified Agents Input
Environment Fuzzing Artifacts
Building Artifacts Fuzzing
Call
Delegate
Execution
Seeds
Dictionary
Fuzzer Instances
Web Search Interface
Manager Agent
Fuzzer Executor
Fuzzing Harnesses
e at leg De
C/C++ Libraries
Harness Generator
Coverage Analysis Interface
Analysis Agents
Coverage Analyzer
Fuzzing Outcomes
Inspection
Call Crash Analyzer
Fuzz Targets
Crash Debugging Interface
Hierarchical Code Coverage
Crashes
Feedback and Observation
Fig. 3. Multi-agent architecture of FuzzAgent.
decomposition into specialized subtasks, each handled by operations. The Environment maintains a stateful workspace agents with specific expertise. With these advancements, LLM- in which agents execute tasks and persist their results, ensuring powered agents can now perform complex reasoning, compre- data consistency and effective resource management throughout hend domain-specific knowledge, and generate high-quality the fuzzing campaign. Within this environment, agents act outputs across various domains, including code generation and through the interfaces and receive feedback and observations program analysis [47]. to iteratively refine their decisions. A typical multi-agent system workflow, often utilizing Built on this architecture, FuzzAgent orchestrates a systechniques like ReAct [53] for reasoning and acting, is tematic workflow that drives evolutionary library fuzzing. The illustrated in Figure 2. In this framework, tasks are distributed workflow follows a set of specialized strategies (detailed in to an agent pool where agents are designed with specific Section IV), each encoded as agent prompts paired with roles and capabilities. These agents collaborate to achieve the interface tools that translate fuzzing feedback into the next overall objective. Agents utilize provided tools as interfaces concrete action. This design decomposes the complex library to interact with the external environment, allowing them to fuzzing task into manageable subtasks, each handled by gather information, perform computations, or manipulate data an agent with specific expertise. The following subsections as needed. The observations collected from the environment describe each component in detail. are then fed back to the agents, enabling them to refine their strategies and actions iteratively. This dynamic interaction A. Agent Pool between agents and their environment allows the system to The Agent Pool groups specialized agents along the three adapt its strategies and improve performance over time. phases of library fuzzing: setup, execution, and outcome III. A RCHITECTURE OVERVIEW analysis. Each agent owns one aspect of the process, which FuzzAgent is a multi-agent system that performs fully keeps the system modular and lets every agent operate within automated library fuzzing and evolves its performance over its area of expertise. A Manager Agent [45], [46] coordinates time. It takes only the target library’s source code as input. As the pool by picking the next agent to run based on the current shown in Figure 3, the architecture of FuzzAgent consists of workspace state and the latest feedback signals. The remainder three primary components: the Agent Pool, the Interfaces, and of this subsection describes each agent’s role and outputs. 1) Fuzzing Setup Agents: Setup agents lay the groundwork the Environment. The Agent Pool hosts a set of specialized for effective library fuzzing by preparing the build environment, agents, each handling a distinct phase of the fuzzing lifecycle. the dictionary, and the initial seeds. To support these agents, FuzzAgent provides four dedicated Interfaces for computer usage, web searching, code coverage Library Builder: This agent produces reliable build artifacts analysis, and crash debugging. Each interface offers a high- (e.g., headers, static and dynamic libraries) with the required level abstraction over the capabilities needed for library fuzzing, instrumentation (e.g., sanitizers, coverage tracking). It inspects so agents can concentrate on strategic decisions and guiding the source code and build configurations, and then automates the evolutionary process instead of handling low-level system the build process to keep the result reproducible and correct.
4
Dictionary Generator: This agent builds a compact and system and specialized commands. The file operation tools are effective fuzzing dictionary tailored to the target library by adapted from SWE-Agent [47] and tailored to library fuzzing combining domain-specific knowledge with existing materials for simplicity, providing robust and efficient file manipulation. acquired from the Web. The specialized command tools abstract recurring procedures Seed Generator: This agent collects example files as initial such as library building, harness compilation, and fuzzing fuzzing inputs that conform to the target library’s specification, execution; each tool encapsulates a routine task sequence drawn from common fuzzing practices, which improves correctness protocol, or data format. 2) Fuzzer-Specific Agents: Building on the artifacts from and hides incidental complexity. Overall, the interface lets the setup agents, fuzzer-specific agents follow the optimization agents interact with the underlying system to retrieve and instructions from the analysis agents to drive targeted fuzzing manipulate files, examine library specifications, and manage exploration. They generate targeted harnesses and run the fuzzing processes, mirroring the workflow of human experts corresponding fuzzing campaigns to explore the library code. and enabling the fully automated execution of library fuzzing. Harness Generator: This agent generates targeted fuzzing Web Search Interface. This interface lets agents reach online harnesses for the coverage bottlenecks reported by the Coverage resources and documentation when needed. It exposes tools for Analyzer. Each harness specifies the APIs or code regions to querying search engines, extracting relevant information, and exercise, along with the required dependencies and invocation downloading materials from the Web. With these tools, agents sequences. The agent also validates the harness through can gather external knowledge about the target library, such as compilation to ensure correctness. API usage examples, common pitfalls, and best practices, and Fuzzer Executor: This agent runs the fuzzing campaign on break out of otherwise unrecoverable error loops during library a generated harness in a blocking manner and monitors the building and harness generation. The same tools also retrieve process for crashes and coverage metrics. example files used to prepare the fuzzing dictionary and seeds, 3) Analysis Agents: Analysis agents inspect fuzzing out- which are hard to generate by LLMs alone, especially for comes to locate the bottlenecks that limit effectiveness, and libraries that require specific binary formats. The Web Search turn their findings into optimization instructions for the other Interface therefore extends agent capabilities to the broader agents so that the process continuously improves. knowledge available on the Web and helps address problems Crash Analyzer: This agent investigates the root cause of each that LLMs cannot solve on their own. crash by interactively inspecting source code and debugging Coverage Analysis Interface. This interface provides tools the crash program. It triages every crash as either a genuine for in-depth analysis of code coverage data, giving agents library bug or a harness error. It then either provides feedback a comprehensive view of the current coverage status. This to fix the harness or produces a detailed report for the bug. goes beyond the simple metrics produced by the lightweight Coverage Analyzer: This agent identifies the most critical instrumentation in standard fuzzers (e.g., basic-block transitions coverage gaps left by the current fuzzing effort. It then derives in AFL [5]). Inspired by FuzzIntrospector [56], we organize concrete API invocation sequences that, if exercised correctly, coverage data into a hierarchy that spans from the coarseare expected to yield the largest coverage gain, and uses them grained project level to the fine-grained branch level, and to guide the generation of the most promising harnesses. present it in a compact, human-readable form so that agents can locate bottlenecks at the appropriate granularity. From B. Interfaces this view, agents can pinpoint the most severe coverage gaps FuzzAgent provides four dedicated interfaces that give and reason about concrete remediation strategies, which in agents the capabilities required for library fuzzing while turn drives the generation of targeted harnesses for maximum shielding them from low-level complexity. Each interface is coverage gain. The hierarchy is organized as follows: a collection of tools that exposes a high-level abstraction of • Project-Level: Overall library coverage status one functional area needed to fulfill library fuzzing objectives. • Module-Level: Component-specific coverage We design these interfaces by following common practices • File-Level: File-by-file coverage from the fuzzing community [5], [23], [11], [29], [31], [32], • API-Level: API function-specific coverage [33] and by automating the error-prone low-level operations • Internal-Function-Level: Internal function-specific coverbehind them (e.g., file manipulation, instrumentation, fuzzer age and blockage status compilation, web searching, coverage extraction, and crash • Branch-Level: Code branch blockage status debugging). The abstraction exposes actions that map directly to library fuzzing tasks, but still preserves the flexibility agents Crash Debugging Interface. This interface provides tools for may need (e.g., passing extra flags during fuzzer compilation). in-depth crash context analysis, so that agents can determine With this design, agents can issue concise and purposeful the actual root cause of a crash instead of guessing from the call actions instead of orchestrating brittle low-level steps, which trace. Its capabilities cover analyzing crash dumps and stack are a common source of incorrect or unexpected results in traces, inspecting source code, and debugging crash programs. agentic systems [54], [55]. We detail each interface below. Using CASR [57], the interface automatically correlates crash Computer Use Interface. This interface lets agents perform file traces with the corresponding source snippets to assemble the operations (reading, writing, and searching) and execute both crash context. In parallel, custom scripts running in GDB’s
5
non-interactive batch mode [58] let agents reproduce crashes and inspect runtime states at the crash point in a systematic way. Together, these tools enable accurate crash triage through evidence-based root cause analysis.
realized as a tailored prompt together with the interface tools the agent may invoke to carry out its task. The system prompt of every agent is listed in Appendix F. The remainder of this section details these strategies and explains how they let FuzzAgent fuzz libraries effectively.
C. Environment
A. Fuzzing Environment Setup
The Environment maintains a stateful workspace for the library fuzzing workflow. In agentic systems, agents act dynamically and can easily drift from the intended workflow, causing stability issues [59], [60]. Prior automated library fuzzing approaches sit at the opposite extreme: they are predefined and rigid, gaining stability and reproducibility at the cost of adaptability [61]. The Environment bridges these two extremes by imposing structural constraints on an otherwise flexible workflow. Concretely, the Environment defines a strict directory layout for the fuzzing workspace, covering source code, build artifacts, fuzzing inputs and outputs, and analysis results. Agents are not allowed to read or write outside this layout, and any violating action is rejected. Before an agent exits, the Environment validates the current directory state to confirm that the agent followed the expected execution trace and produced the required results. By continuously monitoring this directory state, the Environment also tracks the overall progress of the campaign, guiding the workflow to iterative targeted exploration and outcome analysis.
FuzzAgent starts a campaign by scheduling Library Builder, Dictionary Generator, and Seed Generator to set up the fuzzing environment. Library Builder first builds the target library with the required instrumentation (e.g., sanitizers, coverage tracking) and produces the headers and library binaries needed for later harness compilation. Dictionary Generator and Seed Generator then prepare a domain-specific dictionary and an initial seed corpus to bootstrap the subsequent fuzzing campaigns. Library Building. Rather than building the target blindly, Library Builder first inspects the source code to identify the build requirements. It then writes a build script that accepts custom instrumentation flags (Figure 1 in Appendix F), in the spirit of OSS-Fuzz [11]. The agent runs this script through the build tools to produce artifacts instrumented with sanitizers (e.g., ASAN [34], UBSan [35]), coverage tracking (source-based code coverage [62]), and custom passes (e.g., wllvm [63]). During this process, the build tools actively monitor the execution and verify the results. Any errors or unsatisfactory outcomes are returned to Library Builder for resolution. This build-verify-fix loop continues until all verifications are successful. Dictionary Generation. A fuzzing dictionary is a set of grammar tokens commonly used in the target library’s input space; these tokens speed up path exploration by steering the fuzzer toward meaningful input regions [5], [23]. Producing such a dictionary is hard, as it requires a deep understanding of the target’s input formats and protocols, and is therefore usually written by library developers. Existing automatic extraction methods [37], [64] rely on heavy static or dynamic analysis and still miss much of the format and protocol semantics. To avoid such heavy analysis, Dictionary Generator instead retrieves existing dictionaries from the Web. It searches GitHub for open-source projects that share the same protocol or format as the target (Figure 7 in Appendix F), and follows a retrievaland-understanding approach. It inspects each project’s source layout and build scripts (e.g., Dockerfile, build.sh) to locate dictionary files, including those reached through download links embedded in complex build steps. After collecting the candidates, the agent prunes tokens that do not match the target’s specification and assembles the remainder into a structured dictionary file. Seed Generation. Fuzzing seeds are the initial inputs that bootstrap the fuzzer and help it reach deeper code paths sooner [5], [23], [38]. Generating valid seeds for complex input formats and protocols is hard for LLMs, especially when the library expects binary inputs. Seed Generator therefore collects example files from the Web instead, an approach prior work has shown to be the most effective way to prepare seeds [50]. It searches publicly available datasets and repositories related
IV. E VOLUTIONARY L IBRARY F UZZING Section III presents the architecture of FuzzAgent, namely the specialized agents, interfaces, and environment that enable automated library fuzzing. These components alone, however, are not enough. As in prior library fuzzing work [18], [20], [25], [27] that relies on custom analyses or heuristics to drive the process, FuzzAgent still needs well-defined strategies that turn feedback into concrete next actions. Building on FuzzAgent’s multi-agent architecture, we organize the fuzzing process into the workflow shown in Figure 4, with each phase guided by a dedicated strategy: 1) Fuzzing Environment Setup. FuzzAgent builds and instruments the target library in a trial-and-error manner and prepares the initial fuzzing dictionary and seeds to set up the fuzzing environment for effective exploration. 2) Targeted Fuzzing Exploration. FuzzAgent assembles a targeted harness under the guidance of the Coverage Analyzer and Crash Analyzer, and runs fuzzing campaigns on this harness as well as collects the resulting data. 3) Coverage-driven Evolution. If no crashes are found, FuzzAgent inspects the hierarchical coverage data to locate the most critical gaps and proposes new harnesses to reach them under thorough analysis of API relations. 4) Crash-driven Evolution. If crashes are observed, FuzzAgent triages them through iterative debugging: harnessinduced crashes yield feedback to fix the harness, while genuine library bugs are recorded for reporting. Whereas prior approaches encode their analyses and heuristics in hard-coded logic, agentic systems steer LLMs through textual instructions. In FuzzAgent, each strategy is therefore
6
Phase1: Fuzzing Environment Setup Prepare
Build
Prepare
Compile Library Builder
Feed Dictionary Generator
Building Artifacts
Feed
Dictionary
Seeds Augment
Generate new
Mutate
Execute
Compile
Seed Generator Analyze
Modify existing Harness Generator
Harnesses
Fuzzer Executor
Fuzz Targets
Test Cases
Phase2: Targeted Fuzzing Exploration Suggest new
Analyze API relations No Analyze Coverage Bottlenecks
Phase3: API-Surface Exploration
No
Is Stagnating? Suggest new
Analyze blockers
Coverage Analyzer
Yes
Suggest fixes No Yes
Has crashes?
Hierarchical Coverage
Phase4: Deep Stated Exploration Debug & Triage
Is library bug?
Yes
Phase5: Crash-Driven Evolution Library Bugs
Coverage Analyzer
Crashes
Fig. 4. Workflow of Evolutionary Library Fuzzing.
to the target’s domain (Figure 8 in Appendix F), examines the results to identify relevant sources such as test data or user-contributed content, and downloads them as the initial seed corpus for the fuzzing campaigns.
relevant source code files and documentation to gather necessary context about these APIs, including their declarations, dependencies, and usage patterns. Based on this understanding, it constructs a fuzzing harness that correctly invokes the targeted APIs in accordance with their expected usage patterns. 2) Compilation & Verification: Each generated harness is compiled against the instrumented library using compilation tools (detailed in Appendix E) to verify its validity. Compilation errors trigger iterative refinement of the harness until it meets quality criteria. Fuzzer Execution. Fuzzer Executor runs the actual fuzzing campaign on the target harness using the prepared fuzzing dictionary and seeds. It executes the campaign in a blocking manner and monitors primary fuzzing metrics (e.g., code coverage and crash signals) using execution tools (detailed in Appendix E). The agent ensures that the process continues until specific termination conditions are met: crashes are detected, code coverage reaches a plateau [65], or the time budget is exhausted. Specifically, we set a default time budget of 1 hour for each fuzzing campaign. This duration is typically sufficient for fuzzers to explore most reachable code paths in practice [12]. Within the time budget, we monitor the code coverage growth rate Rt over time, calculated as: Nt − Nt−1 Rt = (1) Nt−1
B. Targeted Fuzzing Exploration Unlike traditional fuzzing approaches that drive the fuzz loop by mutating inputs, FuzzAgent focuses on exploring library code by generating fuzzing harnesses and subsequently executing fuzzing campaigns using them. Existing library fuzzing approaches [18], [25], [20], [22] adopt a brute-force strategy, generating harnesses by mutating API combinations or invocation sequences. However, complex dependencies within libraries often cause these approaches to generate invalid or low-quality harnesses that fail to effectively explore the library code. Consequently, this brute-force method requires a large number of attempts, leading to significant inefficiency. Built upon our multi-agent architecture, we generate harnesses in a feedback-driven manner. This approach achieves a deep understanding of library usage by iteratively retrieving source code and applying fixes to ensure harness validity. Furthermore, it specifically targets identified coverage bottlenecks to ensure fuzzing effectiveness. Harness Generation. 1) Coverage-guided Generation: When specific guidance is provided, the agent first interprets the instructions to understand the targeted APIs or code regions that need to be exercised. It then incrementally retrieves
Nt is the total number of unique runtime features covered at
7
time t. If the coverage increase rate Rt falls below a predefined threshold (e.g., 0.01%) for a certain duration (e.g., 1 minutes), it indicates that the fuzzer has likely reached a plateau in discovering new code paths. This prompts the termination of the fuzzing campaign. After execution, the tools collect the results, including code coverage statistics and crashes, and store them in the Environment for later analysis. Test cases that exercise new runtime features are merged into the seed corpus for augmentation.
resolution strategy (Figure 11 in Section F) to identify and overcome specific runtime coverage blockers. In this phase, the agent analyzes internal-function level and branch level coverage data to pinpoint specific functions and code branches that contain a large amount of unexecuted code despite existing fuzzing efforts. For each identified blocker, it performs a detailed analysis of the associated source code to understand the API constraints required to trigger these paths. Based on this analysis, it formulates targeted API invocation sequences, along with the expected API parameter values designed to satisfy these conditions and unlock the blocked code paths. This information is then provided as guidance for harness generation to facilitate deeper exploration of the library’s internal states.
C. Coverage-Driven Evolution
After each targeted fuzzing exploration campaign, if no crashes are found, Coverage Analyzer analyzes the hierarchical code coverage data. It identifies the most critical coverage gaps and suggests new harnesses to further improve coverage. D. Crash-Driven Evolution In contrast to previous library fuzzing approaches that rely When crashes are detected during targeted fuzzing exploon simplistic feedback (e.g., API coverage) to guide harness ration, Crash Analyzer first minimizes the harness code [18] and generation [18], [20], [25], [22], which rarely consider the deep then analyzes them through iterative debugging to determine semantics of libraries, FuzzAgent employs a comprehensive their root causes (Figure 12 in Section F). This process coverage analysis methodology that systematically guides corresponds to Phase 5 in Figure 4. The primary goal of this effective harness generation. This process corresponds to Phases phase is to triage crashes into genuine library bugs or harness 3 and 4 in Figure 4. It employs a two-phase methodology errors, enabling focused remediation efforts. First, the agent to systematically identify and address coverage bottlenecks: utilizes the Crash Debugging Interface to reproduce crashes, the API Surface Exploration Phase focuses on uncovering minimize the harness code, collect call traces, and correlate unreachable APIs to rapidly expand the overall coverage , them with corresponding source code snippets. Then, through while the Deep State Exploration Phase targets specific runtime systematic source code retrieval and crash program debugging, blockers that hinder deeper exploration of already reachable it analyzes the crash context to identify root causes. In addition, code paths. By iteratively applying these two complementary it enhances this triage with explicit evidence derived from strategies, the agent drives the evolutionary fuzzing process to source code or documentation. Based on this analysis, the improve code coverage substantially. agent classifies each crash accordingly and provides actionable API Surface Exploration Phase: Coverage Analyzer adopts feedback for harness fixes or generates detailed reports for a group-based API relation reasoning strategy (Figure 10 genuine library bugs. in Section F) to identify high-impact unreachable APIs for V. E VALUATION harness generation. Instead of treating each uncovered API This section presents a comprehensive evaluation of in isolation and inferring their relations, which is adopted by PromeFuzz [22], the system first identifies group dependen- FuzzAgent across 20 widely-used C/C++ libraries. To cies among APIs. It does this by analyzing module level and file rigorously assess FuzzAgent’s efficiency and effectivelevel coverage to pinpoint under-explored code regions within ness, we compare it against four representative baselines: similar components that share functional relations. Then, it ana- OSS-Fuzz [11], OSS-Fuzz-Gen [26], PromptFuzz [20], lyzes the dependencies among these uncovered APIs to identify and PromeFuzz [22], covering the spectrum from an industrial clusters of related functions that, when exercised together, can continuous fuzzing infrastructure to state-of-the-art LLM-driven unlock significant portions of the library’s functionality. In harness generators. The evaluation targets two complementary addition to focusing on uncovered APIs, the analyzer identifies groups of libraries: 17 representative libraries previously necessary helper functions (e.g., initialization, cleanup, and evaluated by PromptFuzz and PromeFuzz, enabling direct validation functions) required to construct proper invocation and fair comparison, and 3 large-scale, industrially critical sequences. These targeted API invocation sequences, combined libraries (protobuf, OpenSSL, and OpenCV) whose complex with the intended functionality and necessary helper functions, codebases and build systems pose the challenges of real-world are then provided as guidance for harness generation to deployment. The experimental setup is detailed below. maximize code coverage improvement. All experiments were conducted on a server equipped with Deep Stated Exploration Phase: PromptFuzz [20] has two Intel Xeon Platinum 8580 processors (120 physical cores already demonstrated that expanding the API surface alone / 240 threads, 2.0 GHz), 8 NVIDIA H200 GPUs, and 2TB is insufficient for achieving deep code coverage. The process of RAM, running Ubuntu 24.04 LTS. To ensure a controlled quickly encounters a plateau because many runtime coverage and reproducible experimental environment, we self-hosted blockers require specific input formats or API call sequences DeepSeek V3.2 [66] via vLLM [67] for all LLM-dependent to trigger deeper code paths [12]. To enable deeper exploration components, eliminating variability introduced by external of library code, Coverage Analyzer employs a targeted blocker API rate limits or service-side changes. The inference was
8
TABLE I OVERVIEW OF THE EVALUATION RESULTS ACROSS 20 LIBRARIES . cJSON libmagic RE2 pugixml zlib
c-ares liblouis libpng libpcap libtiff
lcms tinygltf
libjpeg -turbo
curl
libvpx SQLite3 libaom protobuf OpenSSL OpenCV
Total
Library Statistics Language Version LoC #APIs #Branches
C C C++ C++ C C C C C C C C++ C C C C C C++ b2890 772f2 972a1 71005 09a15 2870f e90a9 7c67f a516c 5fe20 36039 bdc37 466c3 e8415a cb5a5 db4d8 16a97 b56a4 13K 19K 33K 42K 44K 52K 65K 108K 122K 150K 154K 194K 212K 314K 501K 580K 857K 1.26M 78 18 108 352 88 138 34 271 108 191 296 466 146 156 37 298 47 3149 1060 7802 5060 4448 3038 9172 7338 7810 7394 15466 9180 7584 11388 19980 33082 54600 63124 25498
C C++ c8b4a 01f9b 1.58M 3.19M 9.49M 6728 4236 16945 130338 237392 660754
Metric 1: # Fuzzer Instances OSS-Fuzz OSS-Fuzz-Gen PromptFuzz PromeFuzz FuzzAgent
1 17 34 52 25
3 16 26 52 18
1 14 2 43 16
2 23 111 17
11 26 82 49 21
2 14 46 62 16
3 18 39 15 18
8 25 42 47 20
3 22 50 53 19
1 21 22 91 14
15 30 69 130 17
1 15 32 19
32 4 29 36 12
18 10 58 58 19
8 13 62 25 14
1 50 67 141 15
1 28 58 24 13
2602 2396 1780 3253 4069 4174
2489 2786 3316 3582 3757
1959 2040 2368 1985 2598 2613
3340 2025 5326 5091 5182 5387
4687 1778 4428 3624 4911 5078
1575 1378 3361 1849 4921 5342
3106 2428 3804 4239 5120 5455
3078 3673 7049 7586 8212 9139
4110 1976 4138 3673 4537 4681
1101 157 1925 3056 3631
7363 1944 3360 4975 7541 7713
5126 11677 17690 11122 892 2915 15458 9459 5699 9290 18454 18382 5596 6558 13041 22086 6699 15680 26197 33308 6117 16400 26160 33832
OSS-Fuzz-Gen $21.62 $11.90 $11.37 $3.59 $3.44 $14.22 $11.51 $13.71 $10.45 $12.53 $14.98 $3.56 PromptFuzz $0.41 $0.86 $2.78 $0.44 $2.63 $1.46 $2.84 $0.54 $2.68 $1.45 PromeFuzz $0.22 $0.05 $0.34 $0.62 $0.28 $1.08 $0.38 $2.66 $0.34 $1.13 $2.11 $0.22 FuzzAgent $3.51 $2.65 $2.34 $2.50 $3.13 $3.31 $3.27 $3.83 $2.66 $3.54 $3.51 $3.60
$6.42 $1.18 $0.37 $3.57
$2.88 $13.94 $12.85 $1.90 $0.73 $1.24 $0.46 $0.33 $5.39 $3.31 $3.36 $3.12
1 12 12
152 19 44 15
9 2 10
273 379 730 1021 330
477 696 7664 7789
35983 1805 10777 18227 18374
1985 3075 12751 13389
123792 61678 103686 93507 179619 184840
Metric 2: Branch Coverage OSS-Fuzz OSS-Fuzz-Gen PromptFuzz PromeFuzz FuzzAgent FuzzAgent†
489 694 831 805 880 882
3833 4103 4639 3905 4484 4927
Metric 3: LLM Cost $9.17 $22.49 $1.08 $1.73 $203.44 $0.35 $1.66 $23.15 $0.68 $223.06 $221.80 $214.49 $676.01 $2.98 $3.73 $2.77 $2.75 $63.44
Metric 4: Detected Bugs OSS-Fuzz OSS-Fuzz-Gen PromptFuzz PromeFuzz FuzzAgent
0 0 1 1 3
0 0 0 0 2
0 0 0 0 0
0 0 1 3
0 0 0 0 1
0 0 0 0 1
0 0 1 1 8
0 0 0 0 3
0 0 0 0 2
0 0 0 1 7
0 0 0 1 3
0 0 0 0
0 0 1 0 3
0 0 0 0 0
0 0 2 0 28
0 0 0 1 1
0 0 0 2 21
0 0 2
0 0 0 6
0 0 8
0 0 5 8 102
LoC: Lines of Code, counted by scc; libraries are sorted by LoC in ascending order. Fuzzer Instances reflects the number of harnesses that were successfully compiled and actually executed during the fuzzing phase. FuzzAgent† : additional 24-hour fuzzing results on merged harnesses generated by 24-hour end-to-end execution of FuzzAgent. PromptFuzz does not support C++ (RE2 is evaluated through its cre2 C wrapper); PromeFuzz could not complete its codebase summarization phase within 24 hours for protobuf, OpenSSL, and OpenCV; unsupported entries are marked “-”.
configured with a temperature of 1.0 and top-p of 0.95, following the official recommendations of DeepSeek [68]. The fuzzer engine used across all experiments was AFL++ [51], and code coverage was measured using Source-based Code Coverage [62] instrumentation with llvm-cov [69]. The code coverage of third-party libraries (e.g., absl for RE2) was excluded from the reported results to ensure a fair comparison focused on the target libraries. ASAN and UBSAN sanitizers were enabled for all fuzzing campaigns to detect memory safety and undefined behavior bugs.
use 50 parallel LLM threads, consistent with the original paper. OSS-Fuzz-Gen covers only a limited API subset by default; for fairness, we extend its configuration to include all available APIs. The dictionaries and seed corpus collected by OSS-Fuzz are used for all baseline approaches, while FuzzAgent starts without dictionaries and seed corpus. As shown in Appendix C, coverage results can vary substantially across independent runs due to both fuzzing randomness and LLM generation randomness, with harness generation contributing the dominant source of variation. Following best practices for fuzzing evaluation [70], we repeat every harnessSince FuzzAgent is designed as an end-to-end evolutionary generation phase across five independent trials and repeat library fuzzing system, we evaluate it under a 24-hour per- every 24-hour fuzzing phase across five independent trials, library budget that covers the full workflow in Figure 4, includ- reporting the mean throughout the evaluation. To further align ing environment preparation, harness generation, fuzzing, cover- FuzzAgent with the merged-harness protocol used by the age analysis, and feedback-driven refinement. For OSS-Fuzz, generation-based baselines, we also conduct an additional each existing fuzzer instance was assigned one dedicated CPU experiment in which all harnesses generated by FuzzAgent core and executed for 24 hours. For the LLM-based baselines, across the five independent trials are pooled into a single OSS-Fuzz-Gen, PromptFuzz, and PromeFuzz, we follow merged harness and fuzzed for five independent 24-hour trials. the evaluation protocol of PromeFuzz [22]: five parallel LLM Table I summarizes the overall results across all 20 libraries. threads generate harnesses for up to 24 hours, after which all valid harnesses are merged into a single composite harness A. Automation and Efficiency and fuzzed for another 24 hours. We preserve tool-specific In our evaluation, FuzzAgent demonstrated strong automarequirements where necessary. PromeFuzz requires a codebase tion capability, completing the full fuzzing lifecycle for all summarization phase before harness generation, for which we 20 libraries without human intervention. Averaged over the
9
five independent 24-hour trials, FuzzAgent autonomously TABLE II AGENT- TIME DISTRIBUTION ACROSS THE 24- HOUR END - TO - END generated and executed 330 fuzzing harnesses and achieved EXECUTION OF F UZZ AGENT . a cumulative coverage of 179,619 branches. This end-to-end process was also cost-efficient: per trial, FuzzAgent consumed Library Builder Dict. Seed Harness Fuzzer Crash Coverage 12.6M completion tokens and 864M input tokens on average, cJSON 0.32% 1.00% 1.33% 36.92% 37.96% 10.56% 11.91% libmagic 8.26% 1.33% 1.38% 34.26% 38.87% 6.16% 9.74% 85% of which were served from the cache. Based on the RE2 11.38% 0.80% 1.18% 21.81% 47.81% 6.13% 10.90% official DeepSeek V3.2 pricing model, this corresponds to an pugixml 1.22% 1.79% 1.82% 37.03% 41.16% 6.80% 10.16% zlib 1.52% 1.15% 1.71% 27.22% 42.02% 14.02% 12.35% average LLM cost of $63.44 across the 20 libraries per trial1 , c-ares 2.05% 1.46% 2.16% 33.43% 40.06% 8.98% 11.87% as detailed in Table I. Across all five trials, it detected 102 liblouis 2.06% 0.70% 0.99% 45.31% 26.66% 13.56% 10.72% libpng 0.59% 0.78% 1.47% 32.41% 32.90% 20.49% 11.36% genuine library bugs in total, yielding an LLM cost of only libpcap 4.05% 1.00% 1.11% 35.81% 34.40% 9.98% 13.65% $3.11 per genuine bug under a full automated setting. libtiff 13.85% 0.85% 1.31% 27.28% 28.83% 16.98% 10.91% lcms 4.26% 0.83% 1.04% 29.10% 26.69% 22.21% 15.87% In sharp contrast, the generation-based baselines remain only tinygltf 1.21% 0.78% 0.74% 39.76% 33.90% 11.79% 11.82% partially automated: they automate harness synthesis, but still libjpeg-turbo 4.68% 0.81% 1.11% 37.03% 27.99% 19.64% 8.74% curl 5.59% 1.28% 0.89% 27.27% 34.97% 13.77% 16.23% rely on human operators to prepare build configurations, resolve libvpx 0.76% 0.60% 1.29% 43.65% 19.22% 17.13% 17.35% dependencies, and establish executable fuzzing environments. SQLite3 4.29% 0.70% 1.64% 43.29% 30.76% 9.88% 9.45% libaom 3.75% 0.86% 1.46% 34.03% 34.64% 8.13% 17.12% For libraries already included in their original evaluations, this protobuf 6.55% 0.99% 1.24% 33.05% 26.14% 19.60% 12.44% manual preparation required approximately 1 hour per library OpenSSL 3.52% 1.07% 1.60% 28.11% 27.90% 13.04% 24.77% OpenCV 2.92% 0.74% 1.39% 34.19% 28.78% 13.26% 18.72% on average. For the newly added targets in our evaluation Average 4.14% 0.98% 1.34% 34.05% 33.08% 13.11% 13.30% (curl, OpenSSL, and OpenCV for OSS-Fuzz-Gen; liblouis Note: Percentages denote the fraction of agent activity time within a 24-hour execution. and OpenSSL for PromptFuzz; and protobuf, OpenSSL, Rows sum to 100% up to rounding. and OpenCV for PromeFuzz), the effort increased substantially, averaging 5 hours, 8 hours, and 12 hours per library, respectively. These increases were primarily caused by complex Seed Generator, and Dictionary Generator together accounting build systems, dependency conflicts, and the preparation of for only 6.46% of the activity time. This distribution indicates required materials (e.g., consumer code and documentation). that FuzzAgent effectively leverages LLM-driven agents to The need for such manual intervention is also reflected in diagnose and overcome library fuzzing bottlenecks as they arise, comments from the PromptFuzz developers [71]. Beyond enabling the system to continuously evolve its performance setup cost, PromeFuzz further exhibited severe scalability while preserving end-to-end automation. bottlenecks during its codebase summarization phase on large projects such as protobuf, OpenSSL, and OpenCV. In our TABLE III T HE NUMBER OF TOOLS EXECUTED FOR DIFFERENT AGENTS . experiments, this phase failed to complete within 24 hours even after incurring more than $200 in LLM cost, consistent Tool Categories with the O(n2 ) complexity of its summary generation in Agents Files Bash SFTs Web Coverage Crash Total the number n of API functions [72]. These results underline Build 178 207 49 52 0 0 485 FuzzAgent’s distinctive advantage: it provides a genuinely endDict 146 177 0 370 0 0 693 to-end automated fuzzing workflow that removes the human Seed 19 563 0 416 0 0 998 setup bottleneck while scaling to large, real-world libraries. Harness 11343 1962 1125 60 0 0 14490 Fuzzer 1552 3386 612 0 0 0 5550 To understand how FuzzAgent orchestrates this high level Coverage 562 330 0 0 3282 0 4174 of automation, we analyzed the agent-time distribution across Crash 1913 982 0 0 0 1213 4108 the 24-hour end-to-end executions. As shown in Table II and viTotal 15714 7605 1786 898 3282 1213 30499 sualized by the execution trajectories in Figure 13, FuzzAgent Note: SFTs stands for specialized fuzzing tools. Files, Bash and SFTs are devotes most of its runtime to the two agents that directly drive the three representative collection of tools for FuzzAgent’s computer-use coverage growth: the Harness Generator (34.05%) and the interface. The names of Agents are the agents in FuzzAgent for short. Fuzzer Executor (33.08%). Together, they account for 67.13% Finally, we examined the tool interfaces that allow Fuzof the overall activity time, indicating that the system spends roughly two-thirds of its budget generating executable harnesses zAgent’s agents to operate on real software environments and exercising them under fuzzing, rather than on coordination rather than merely produce text. Table III reports 30,499 tool overhead. At the same time, the feedback agents consume invocations across the evaluation on 20 libraries per trial. a substantial fraction of the budget: the Coverage Analyzer File-system operations and Bash commands dominate the (13.30%) identifies unexplored code regions, while the Crash interaction pattern, accounting for 15,714 and 7605 invocations, Analyzer (13.11%) triages failures and feeds corrective signals respectively. Also, 1786 invocations of specialized fuzzing tools back into the workflow. In contrast, environment and auxiliary- were used to build the libraries, compile harnesses and execute input preparation remain lightweight, with the Library Builder, fuzzing. Together, these computer-use actions represent 82.3% of all tool calls, showing that autonomous library fuzzing depends heavily on the ability to inspect source trees, edit 1 The official DeepSeek V3.2 pricing is $0.028 per 1M cached input tokens, $0.28 per 1M non-cached input tokens, and $0.42 per 1M output tokens. build artifacts and harnesses, invoke compilers, launch fuzzers,
10
and diagnose execution failures. The distribution across agents further supports this interpretation: the Harness Generator alone accounts for 14,490 tool calls, mostly file operations, indicating LLMs can automatic retrieve necessary knowledge required for high-quality harness generation. In addition to these environment-manipulation interfaces, FuzzAgent actively uses specialized fuzzing tools (1786 calls), coverage-analysis tools (3282 calls), and crash-debugging tools (1213 calls), enabling the system to close the loop between execution feedback and subsequent refinement. Web search is used more selectively (898 calls), primarily by the Dictionary Generator and Seed Generator, where external examples and documentation help construct domain-specific inputs. Overall, these metrics show that FuzzAgent’s automation comes from tool-grounded interaction with the codebase: the LLM agents continuously inspect, build, execute, measure, and debug the target libraries, thereby approximating the workflow of human fuzzing experts in a self-contained manner.
69% for GPT-3.5, and 58% for GPT-4) [40], [41]. Since LLMs are already trained extensively on these open-source codebase, excessive generated or retrieved context may distract rather than help generation [42], [43], [44]. FuzzAgent avoids this issue by retrieving project knowledge on demand through agents. The dictionaries and seed corpora generated by FuzzAgent are constructed from reusable input-format knowledge gathered from related projects. As illustrated in Figure 14, FuzzAgent consults 4.3 related libraries per target on average when preparing these auxiliary inputs. To quantify how these generated dictionaries and seeds contribute to code coverage, we conduct an ablation study on dictionary and seed generation using five one-hour fuzzing runs. As shown in Table IV, the empty configuration, which uses neither dictionaries nor seed corpora, covers 106,772 branches across the 20 libraries. Adding OSSFuzz dictionaries and seeds provides a modest improvement of 2.0% and 7.1% over the empty setting, respectively. When using the dictionaries and seeds generated by FuzzAgent, the improvement increases to 3.6% and 13.1%, respectively. These B. Effectiveness on Code Coverage results indicate that the dictionaries and seed corpora collected As shown in Table I, FuzzAgent achieves the strongest by FuzzAgent play an important role in improving coverage overall branch coverage across the 20 evaluated libraries and accelerating the exploration of library code. using only its standard 24-hour end-to-end execution budWe further evaluate the contribution of the Coverage get. In this setting, FuzzAgent covers 179,619 branches, Analyzer by disabling it and running FuzzAgent on all 20 exceeding OSS-Fuzz (123,792 branches), PromptFuzz libraries for five independent 24-hour fuzzing trials. As shown (103,686 branches), PromeFuzz (93,507 branches), and in Figure 5, removing the Coverage Analyzer reduces branch OSS-Fuzz-Gen (61,678 branches) by 45.1%, 73.2%, 92.1%, coverage to 146,508, descrease 18.5% compared with the and 191.2%, respectively. These gains are statistically robust: standard setting. With coverage guidance enabled, FuzzAgent as detailed in Appendix D, Mann–Whitney U tests show that achieves the highest coverage among 18 libraries and the most FuzzAgent’s improvements are significant for most available stable performance. This substantial degradation highlights library-baseline pairs. The advantage is also consistent at the importance of coverage feedback in FuzzAgent’s evoluthe per-library level, with FuzzAgent achieving the highest tionary workflow. By identifying under-explored code regions coverage on 17 of the 20 targets. Extending the experiment with and informing subsequent harness generation, the Coverage an additional 24 hours of fuzzing on the harnesses produced by Analyzer enables FuzzAgent to continuously expand coverage the end-to-end runs (FuzzAgent † ) further raises total coverage and exercise deeper library code. to 184,840 branches (+2.9%), confirming that the gains stem from FuzzAgent’s evolutionary workflow rather than from a C. Effectiveness on Bug Detection longer fuzzing budget. Across the five 24-hour end-to-end runs of FuzzAgent on the The remaining per-library gaps and baseline failures can 20 libraries, the Crash Analyzer flagged a total of 1098 unique be explained by resource allocation, harness validity, and crashes after de-duplication by call stack with CASR [57]. scalability. For OSS-Fuzz, it substantially exceeds other Of these, 128 were triaged as candidate library bugs and the systems on OpenSSL because it assigns 152 parallel fuzzer remainder as harness errors. To enable responsible disclosure, instances to this target, while FuzzAgent runs with a sequential four PhD students spent two weeks manually reproducing manner. In contrast, PromptFuzz and PromeFuzz rely mainly each candidate, performing root-cause analysis, and preparing on large-scale parallel harness generation, which produces many detailed reports for upstream maintainers, averaging roughly candidates but not necessarily executable or coverage-effective two hours per bug. fuzzers. For example, PromptFuzz generated 11K harnesses This validation confirmed 108 of the 128 candidates (84.38%) across the 20 libraries, but 87.42% were ultimately eliminated as genuine library bugs, with the remaining 20 attributable to due to syntax errors or API misuse, and 654 additional LLM hallucination and weak instruction-following in DeepSeek harnesses were filtered out because they contributed no unique V3.2, where 12 stemmed from the agent failing to retrieve coverage. Compared with this brute-force generation strategy, the relevant API constraints from library documentation, and FuzzAgent adopts an evolutionary process that achieves the other 8 from incorrect root-cause reasoning. Replacing the substantially higher coverage with only 330 harnesses. For underlying model with Claude Sonnet 4.6 reclassified 14 of PromeFuzz, codebase summarization can introduce redundant these 20 false positives correctly as harness errors, indicating or noisy context; as modern LLMs already show decreasing that the bottleneck lies in model capability rather than agent hallucination rates with model upgrades (88% for Llama 2, design. As summarized in Table VII, all confirmed bugs were
11
TABLE IV A BLATION STUDY OF DICTIONARY AND SEED GENERATION ACROSS 20 LIBRARIES . Configuration
cJSON libmagic RE2 pugixml zlib c-ares liblouis libpng libpcap libtiff lcms tinygltf
Empty OSS dictionary Agent dictionary OSS seeds Agent seeds
854 867 868 866 863
2680 4780 3160 4778 3342 4860 3020 4659 3260 4725
3012 2029 4607 3082 2057 4387 3248 2167 4798 3288 2065 4686 3268 2158 4615
896 1037 1074 950 1135
1998 2113 2020 2964 2542
3955 2547 1983 3995 2969 1791 4637 3056 1698 4369 4137 2517 4240 4117 2511
libjpeg curl libvpx SQLite3 libaom protobuf OpenSSL OpenCV -turbo
2237 2126 2233 2174 2279
3638 5052 13837 4173 5006 14634 4162 5171 14223 4232 5344 14169 4651 5542 14899
6783 17092 6956 17454 7995 16305 8645 17207 8494 21381
14303 14408 14312 14801 14934
6201 5960 6101 6263 6269
Total
Gain
8290 106772 – 7959 108913 2.0% 8377 110648 3.6% 8027 114383 7.1% 8895 120777 13.1%
Note: Each cell reports branch coverage after one hour of fuzzing. Empty uses no dictionary and no seed corpus. Gain is computed relative to the empty configuration.
Branch Coverage Growth Over Time Ablation Study: Coverage Guidance (5 Rounds × 20 Libraries)
Branches covered
cJSON
4500
600
3000
300
1500
Branches covered
0
0
6
12
c-ares
18
24
0
4500
4500
3000
3000
1500
1500
0
Branches covered
libmagic
900
0
6
12
lcms
18
24
0
RE2
pugixml
3000
3000
6
12
liblouis
18
24
0
6
12
tinygltf
18
24
800
1000
0
6
12
libpng
18
24
4500
0
1600
2000
1500
0
zlib 2400
4500
0
0
6
12
libpcap
18
24
0
4500
7500
3000
3000
5000
1500
1500
2500
0
0
6
12
libjpeg-turbo
18
24
0
0
6
12
curl
18
24
0
4500
3000
7500
7500
30k
3000
2000
5000
5000
20k
1500
1000
2500
2500
10k
Branches covered
0
0
6
12
libvpx
18
24
18k
12k
6k
0k
0
6
12
Time (h)
18
24
0
0
6
12
libaom
18
24
0
30k
18k
20k
12k
10k
6k
0k
0
6
12
Time (h)
18
24
0k
0
0
6
6
w/ coverage guidance
12
OpenSSL
12
Time (h)
18
18
24
24
0
0
6
12
protobuf
18
24
0k
9000
24k
6000
16k
3000
8k
0
0
6
w/o coverage guidance
12
Time (h)
18
24
0k
0
6
0
6
0
6
0
6
12
18
24
12
18
24
12
18
24
12
18
24
libtiff
SQLite3
OpenCV
Time (h)
Fig. 5. Branch coverage growth over time for the 20 evaluated libraries.
by FuzzAgent, we cross-referenced the confirmed bugs of FuzzAgent with the crash stacks generated by CASR, and the results show in Table VII. We found that all 5 and 8 bugs detected by PromptFuzz and PromeFuzz, respectively, For the crashes triaged as library bugs by the baseline were also detected by FuzzAgent, indicating that these bugs approaches, we performed a similar manual validation process. were not unique to the baselines but rather represent common OSS-Fuzz and OSS-Fuzz detected no genuine bugs in vulnerabilities that multiple fuzzing approaches can uncover. the 20 library bugs, while PromptFuzz and PromeFuzz reported 26 and 21 potential library bugs, respectively. Upon The 102 library bugs span a wide range of types, with manual validation, only 5 and 8 of them are genuine libraries. integer overflows (31) and buffer overflows (25) dominating the PromptFuzz only classify crashes by rules, hence causing population. To assess the real-world impact of these defects, we high false positives, whereas PromeFuzz infers root cause by took a closer look at the 21 confirmed bugs in libaom, the AV1 LLMs but misses concrete runtime evidence. To investigate specific codec shipped in Chromium and many downstream whether these bugs were unique to the baselines or also detected media stacks. After careful source-code review and debugging, submitted upstream with summaries and root-cause analyses; to date, 84 reports have received maintainer responses, and 78 have been acknowledged and fixed.
12
we found that 7 of the 21 bugs are directly reachable through libaom’s own command-line binaries (aomenc and aomdec) using a crafted input file or command-line argument, without requiring any custom harness. This indicates that the issues are not artifacts of an over-permissive fuzzing harness but lie on genuine, externally exposed code paths, and they would therefore be triggerable in any application that feeds untrusted media data into libaom, including web browsers and videoprocessing pipelines. We then performed a deeper analysis of the 8 buffer-overflow bugs in this set. One of them turns out to be an arbitrary-write primitive: an attacker-controlled index escapes the intended buffer bounds and is used as the destination of a store, allowing writes to attacker-chosen addresses. When chained with one of the out-of-bounds read bugs identified by FuzzAgent, which leaks pointers from adjacent heap structures and can be used to defeat ASLR, the two primitives compose into a remote-code-execution exploit chain reachable through AV1 encoding enabled with the SVC feature. This concrete chain demonstrates that the bugs uncovered by FuzzAgent are not merely shallow crashes but include exploitable vulnerabilities with realistic attack surfaces, underscoring the security value of fully automated, end-to-end library fuzzing. VI. D ISCUSSION
VII. R ELATED W ORK A. Automated Library Fuzzing Library fuzzing has evolved from manual approaches to increasingly automated techniques. OSS-Fuzz [11] pioneered large-scale library fuzzing but required substantial manual effort. Researchers have developed various approaches to automate harness generation. Static analysis based approaches, like Fudge [13], FuzzGen [14], GraphFuzz [17], and AFGen [19], leverage static analysis to extract API usage patterns and generate fuzzing harnesses. Fudge identifies API entry points, FuzzGen extracts patterns from client applications, GraphFuzz builds dataflow graphs for dependencies, and AFGen focuses on whole-function fuzzing. Dynamic analysis approaches, like APICraft [16] and Hopper [18], incorporate dynamic analysis to improve harness generation. APICraft records API interactions during execution, while Hopper uses interpretative execution to understand API behaviors. Hybrid approaches, like RULF [73] and UTopia [24], combine multiple techniques. RULF traverses API dependency graphs to generate comprehensive harnesses for Rust libraries. UTopia leverages existing unit tests to generate effective fuzz drivers. Despite these advances, existing approaches have limitations. Static analysis tools struggle with complex dependencies, while dynamic approaches may miss rare code paths. Most importantly, these approaches use fixed strategies that cannot adapt based on fuzzing feedback or evolve to overcome coverage barriers. They typically focus only on harness generation while neglecting other aspects of the fuzzing workflow.
Statistical variance in library fuzzing. Library fuzzing is inherently stochastic, and LLM-based harness generation introduces a second, often dominant source of randomness on top of the fuzzer itself. Our nested experiment in Appendix C B. Fuzzing with Large Language Models quantifies both sources. This finding directly motivated our The emergence of powerful LLMs has opened new posevaluation protocol. Prior work such as PromeFuzz generates sibilities for automating and enhancing fuzzing processes. harnesses once and repeats only the fuzzing phase across LLM-based harness generation, like OSS-Fuzz-Gen [26], multiple trials, which controls for fuzzer randomness but PromptFuzz [20] and PromeFuzz [22], pioneered usleaves LLM randomness unaddressed. To mitigate both sources, ing LLMs to generate fuzzing harnesses from API docwe repeat harness generation and the 24-hour fuzzing phase umentation and code. TitanFuzz [74] specifically targets independently five times each, reporting the mean over all trials. deep learning libraries, showing LLMs can handle complex, To further guard against distributional assumptions, we apply domain-specific APIs. LLM-enhanced fuzzing components, like the non-parametric Mann–Whitney U test (Appendix D). The CodaMosa [65], use LLMs to overcome coverage plateaus in per-library p-values confirm that FuzzAgent’s improvements test generation. Universal Fuzzing [75] demonstrates fuzzing are statistically significant for most library-baseline pairs. across diverse domains without domain-specific customization. False positives. Among the 108 identified potential bugs from Xu et al. [76] explore using LLMs for directed greybox fuzzing Section V-C, 84 received maintainer responses and 78 were toward specific targets or code regions. acknowledged or fixed; the remaining 6 were rejected. We These approaches treat the LLM as a single component examined each rejected report to characterize the residual within a traditional fuzzing pipeline rather than as part of an error modes. Three were declared expected behavior under integrated, adaptive system. In contrast, our work introduces a undocumented contracts, e.g., a libvpx OOM closed with multi-agent architecture where specialized agents collaborate the explanation that “the format allows for large resolutions across all fuzzing phases and provides comprehensive automa(65536 × 65536); if an environment does not have enough tion while enabling continuous evolution through feedbackmemory, then an OOM is expected.” Two were acknowledged driven learning. as genuine robustness issues but deferred as out-of-scope for the VIII. C ONCLUSION current release. The last was a harness-side API misuse missed in our inspection. This analysis suggests that future crash We presented FuzzAgent, a multi-agent system that turns analysis improvements should focus on better understanding library fuzzing into an evolutionary process driven by runtime of API contracts under real-world conditions. feedback. By letting specialized agents collaborate over the
13
full fuzzing lifecycle and ground their decisions in concrete execution evidence, FuzzAgent successively refines its harness suite toward deeper coverage and higher-fidelity crash analysis. Across 20 real-world C/C++ libraries, it outperforms state-ofthe-art baselines in both branch coverage and bug discovery, pointing to a practical path toward fully automated, feedbackadaptive library fuzzing. E THICAL C ONSIDERATIONS This work investigates FuzzAgent, a multi-agent system that performs end-to-end library fuzzing fully autonomously. Given only a target library’s source repository, FuzzAgent builds the project, generates harnesses, runs fuzzing, analyzes coverage, and triages crashes without any expert intervention. While this dramatically lowers the barrier to high-quality vulnerability discovery for defenders, it equally lowers the barrier for malicious actors, who could in principle point such a system at widely deployed open-source libraries and obtain exploitable bugs at scale. We have therefore carefully weighed the dual-use implications of this research and followed the conference’s ethics guidelines throughout the project. Responsible disclosure. We followed a responsibledisclosure policy for all bugs uncovered in this study. Across the 20 evaluated libraries, FuzzAgent produced 128 candidate library bugs. Manual triage by four PhD students, averaging two hours per bug, confirmed 108 as genuine library bugs. We reported each confirmed bug to the upstream maintainers with a reproducer, root-cause analysis, and a suggested fix where applicable. To date, 84 reports have received responses and 78 have been acknowledged or fixed. The remaining bugs are under review or pending coordinated disclosure. For high-impact targets such as libaom, OpenSSL, and libpng, including the libaom buffer-overflow chain that yields an arbitrary-write primitive, we are coordinating with upstream maintainers and downstream consumers (e.g., browser vendors) and withholding technical details until patches are widely deployed. Risk mitigation and release plan. Because FuzzAgent can be operated by non-experts, we will not release its source code or model artifacts at this time. Instead, we plan to deploy FuzzAgent as a gated service. In the next 3–6 months, we will build a curated platform (https://fuzzany.org/) that runs FuzzAgent on behalf of open-source maintainers and OSS-Fuzz-eligible projects. The platform returns only vetted bug reports and patch suggestions through standard disclosure channels; raw crashes and candidate exploits are not exposed. Access is restricted to verified project maintainers and established security responders. This deployment model retains the defensive benefits of FuzzAgent and avoids placing a fully automated vulnerability-discovery tool in arbitrary hands. Other considerations. No human subjects, personal data, or proprietary code were used in this study; all evaluated libraries are open source and were fuzzed in an isolated, selfhosted environment. The LLM components of FuzzAgent were also self-hosted, so no source code, crash data, or candidate exploits were transmitted to third-party model providers. We will continue to monitor downstream impacts after publication
14
and adjust the release policy in consultation with the maintainer communities of the affected libraries. R EFERENCES [1] V. J. M. Manès, H. Han, C. Han, S. K. Cha, M. Egele, E. J. Schwartz, and M. Woo, “The art, science, and engineering of fuzzing: A survey,” IEEE Transactions on Software Engineering, vol. 47, no. 11, pp. 2312–2331, 2021. [2] C. Daniele, S. B. Andarzian, and E. Poll, “Fuzzers for stateful systems: Survey and research directions,” ACM Comput. Surv., vol. 56, no. 9, Apr. 2024. [Online]. Available: https://doi.org/10.1145/3648468 [3] C. Beaman, M. Redbourne, J. D. Mummery, and S. Hakak, “Fuzzing vulnerability discovery techniques: Survey, challenges and future directions,” Computers & Security, vol. 120, p. 102813, 2022. [Online]. Available: https://www.sciencedirect.com/science/article/pii/S016740482 2002073 [4] K. Hu, Q. Chen, Z. Lu, W. Zhang, B. Chen, Y. Lu, H. Jiang, B. Sun, X. Peng, and W. Zhao, “A survey of fuzzing open-source operating systems,” 2025. [Online]. Available: https://arxiv.org/abs/2502.13163 [5] M. Zalewski, “American fuzzy lop,” http://lcamtuf.coredump.cx/afl/, Accessed 2026. [6] M. Böhme, V.-T. Pham, and A. Roychoudhury, “Coverage-based greybox fuzzing as markov chain,” in Proceedings of the 2016 ACM SIGSAC Conference on Computer and Communications Security, 2016, p. 1032–1043. [7] Z.-M. Jiang, J.-J. Bai, and Z. Su, “DynSQL: Stateful fuzzing for database management systems with complex and valid SQL query generation,” in 32nd USENIX Security Symposium (USENIX Security 23). Anaheim, CA: USENIX Association, Aug. 2023, pp. 4949–4965. [Online]. Available: https://www.usenix.org/conference/usenixsecurity23 /presentation/jiang-zu-ming [8] J. Liang, Z. Wu, J. Fu, Y. Bai, Q. Zhang, and Y. Jiang, “WingFuzz: Implementing continuous fuzzing for DBMSs,” in 2024 USENIX Annual Technical Conference (USENIX ATC 24). Santa Clara, CA: USENIX Association, Jul. 2024, pp. 479–492. [Online]. Available: https://www.usenix.org/conference/atc24/presentation/liang [9] S. Poeplau and A. Francillon, “Symbolic execution with SymCC: Don’t interpret, compile!” in 29th USENIX Security Symposium (USENIX Security 20). USENIX Association, Aug. 2020, pp. 181–198. [Online]. Available: https://www.usenix.org/conference/usenixsecurity20/presentat ion/poeplau [10] H. Tu, S. Lee, Y. Li, P. Chen, L. Jiang, and M. Böhme, “Cottontail: Large Language Model-Driven Concolic Execution for Highly Structured Test Input Generation,” in 2026 IEEE Symposium on Security and Privacy (SP). Los Alamitos, CA, USA: IEEE Computer Society, 2026, pp. 2064–2082. [Online]. Available: https://doi.ieeecomputersociety.org/10.1109/SP63933.2026.00110 [11] K. Serebryany, “OSS-Fuzz-google’s continuous fuzzing service for open source software,” in Proceedings of the 26th USENIX Conference on Security Symposium (technical sessions). USENIX Association, 2017. [12] W. Gao, V.-T. Pham, D. Liu, O. Chang, T. Murray, and B. I. Rubinstein, “Beyond the coverage plateau: A comprehensive study of fuzz blockers (registered report),” in Proceedings of the 2nd International Fuzzing Workshop, ser. FUZZING 2023. New York, NY, USA: Association for Computing Machinery, 2023, p. 47–55. [Online]. Available: https://doi.org/10.1145/3605157.3605177 [13] D. Babić, S. Bucur, Y. Chen, F. Ivančić, T. King, M. Kusano, C. Lemieux, L. Szekeres, and W. Wang, “Fudge: fuzz driver generation at scale,” in Proceedings of the 2019 27th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2019, pp. 975–985. [14] K. Ispoglou, D. Austin, V. Mohan, and M. Payer, “FuzzGen: Automatic fuzzer generation,” in 29th USENIX Security Symposium (USENIX Security 20), 2020, pp. 2271–2287. [15] M. Zhang, J. Liu, F. Ma, H. Zhang, and Y. Jiang, “Intelligen: Automatic driver synthesis for fuzz testing,” in 2021 IEEE/ACM 43rd International Conference on Software Engineering: Software Engineering in Practice (ICSE-SEIP). IEEE, 2021, pp. 318–327. [16] C. Zhang, X. Lin, Y. Li, Y. Xue, J. Xie, H. Chen, X. Ying, J. Wang, and Y. Liu, “APICraft: Fuzz driver generation for closed-source SDK libraries,” in 30th USENIX Security Symposium (USENIX Security 21), 2021, pp. 2811–2828.
[17] H. Green and T. Avgerinos, “Graphfuzz: Library api fuzzing with lifetime-aware dataflow graphs,” in 2022 IEEE/ACM 44th International Conference on Software Engineering (ICSE), 2022, pp. 1070–1081. [18] P. Chen, Y. Xie, Y. Lyu, Y. Wang, and H. Chen, “Hopper: Interpretative fuzzing for libraries,” in ACM Conference on Computer and Communications Security (CCS), Copenhagen, Denmark, 2023. [19] Y. Liu, Y. Wang, T. Bao, X. Jia, Z. Zhang, and P. Su, “Afgen: Wholefunction fuzzing for applications and libraries,” in 2024 IEEE Symposium on Security and Privacy (SP), 2024, pp. 11–11. [20] Y. Lyu, Y. Xie, P. Chen, and H. Chen, “Prompt fuzzing for fuzz driver generation,” in Proceedings of the 2024 on ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’24. New York, NY, USA: Association for Computing Machinery, 2024, p. 3793–3807. [Online]. Available: https://doi.org/10.1145/3658644.3670396 [21] Y. Deng, C. S. Xia, H. Peng, C. Yang, and L. Zhang, “Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models,” in Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis, 2023, p. 423–435. [22] Y. Liu, J. Deng, X. Jia, Y. Wang, M. Wang, L. Huang, T. Wei, and P. Su, “Promefuzz: A knowledge-driven approach to fuzzing harness generation with large language models,” in Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 1559–1573. [Online]. Available: https://doi.org/10.1145/3719027.3765222 [23] LLVM, “libfuzzer – a library for coverage-guided fuzz testing,” https: //llvm.org/docs/LibFuzzer.html, Accessed 2026. [24] B. Jeong, J. Jang, H. Yi, J. Moon, J. Kim, I. Jeon, T. Kim, W. Shim, and Y. H. Hwang, “Utopia: Automatic generation of fuzz driver using unit tests,” in 2023 IEEE Symposium on Security and Privacy (SP). IEEE Computer Society, 2022, pp. 746–762. [25] J. Lin, Q. Zhang, J. Li, C. Sun, H. Zhou, C. Luo, and C. Qian, “Automatic library fuzzing through API relation evolvement,” in 32nd Annual Network and Distributed System Security Symposium, NDSS 2025, San Diego, California, USA, February 24-28, 2025. The Internet Society, 2025. [Online]. Available: https://www.ndss-symposium.org/nds s-paper/automatic-library-fuzzing-through-api-relation-evolvement/ [26] Google, “oss-fuzz-gen,” https://github.com/google/oss- fuzz- gen, Accessed 2026. [27] H. Xu, W. Ma, T. Zhou, Y. Zhao, K. Chen, Q. Hu, Y. Liu, and H. Wang, “Ckgfuzzer: Llm-based fuzz driver generation enhanced by code knowledge graph,” in Proceedings of the IEEE/ACM 47th International Conference on Software Engineering: Companion Proceedings, ser. ICSE ’25. IEEE Press, 2025, p. 243–254. [Online]. Available: https://doi.org/10.1109/ICSE-Companion66252.2025.00079 [28] “Oss-fuzz guide: Setting up a new project,” https://google.github.io/oss-f uzz/getting-started/new-project-guide/, Accessed 2026. [29] S. Plöger, M. Meier, and M. Smith, “A qualitative usability evaluation of the clang static analyzer and libfuzzer with cs students and ctf players,” in Proceedings of the Seventeenth USENIX Conference on Usable Privacy and Security, ser. SOUPS’21. USA: USENIX Association, 2021. [30] Q. Yan, M. Huang, and H. Cao, “A survey of human-machine collaboration in fuzzing,” in 2022 7th IEEE International Conference on Data Science in Cyberspace (DSC), 2022, pp. 375–382. [31] S. Plöger, M. Meier, and M. Smith, “A usability evaluation of afl and libfuzzer with cs students,” in Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems, ser. CHI ’23. New York, NY, USA: Association for Computing Machinery, 2023. [Online]. Available: https://doi.org/10.1145/3544548.3581178 [32] O. Nourry, Y. Kashiwa, B. Lin, G. Bavota, M. Lanza, and Y. Kamei, “The human side of fuzzing: Challenges faced by developers during fuzzing activities,” ACM Trans. Softw. Eng. Methodol., vol. 33, no. 1, Nov. 2023. [Online]. Available: https://doi.org/10.1145/3611668 [33] Y. Zhao, W. Guo, H. Goldstein, D. Votipka, K. R. Fulton, and M. L. Mazurek, “A qualitative analysis of fuzzer usability and challenges,” in Proceedings of the 2025 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’25. New York, NY, USA: Association for Computing Machinery, 2025, p. 2504–2518. [Online]. Available: https://doi.org/10.1145/3719027.3765055 [34] K. Serebryany, D. Bruening, A. Potapenko, and D. Vyukov, “Addresssanitizer: A fast address sanity checker,” in Proceedings of the 2012 USENIX Conference on Annual Technical Conference, ser. USENIX ATC’12. USENIX Association, 2012, p. 28.
15
[35] LLVM, “Undefined behavior sanitizer - official documentation,” https: //clang.llvm.org/docs/UndefinedBehaviorSanitizer.html, Accessed 2026. [36] “Oss-fuzz guide: Setting up a new project (builds),” https://google.githu b.io/oss-fuzz/getting-started/new-project-guide/#buildsh, Accessed 2026. [37] A. A. Ebrahim, M. Hazhirpasand, O. Nierstrasz, and M. Ghafari, “Fuzzingdriver: the missing dictionary to increase code coverage in fuzzers,” in 2022 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER), 2022, pp. 268–272. [38] Google, “How to prepare the seed corpus for oss-fuzz,” https://goog le.github.io/oss-fuzz/getting-started/new-project-guide/#seed-corpus, Accessed 2026. [39] M. Boehme, C. Cadar, and A. ROYCHOUDHURY, “Fuzzing: Challenges and reflections,” IEEE Software, vol. 38, no. 3, pp. 79–86, 2021. [40] M. Dahl, V. Magesh, M. Suzgun, and D. E. Ho, “Large legal fictions: Profiling legal hallucinations in large language models,” Journal of Legal Analysis, vol. 16, no. 1, 2024. [Online]. Available: http://dx.doi.org/10.1093/jla/laae003 [41] Y. Bang, Z. Ji, A. Schelten, A. Hartshorn, T. Fowler, C. Zhang, N. Cancedda, and P. Fung, “HalluLens: LLM hallucination benchmark,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, 2025. [Online]. Available: https: //aclanthology.org/2025.acl-long.1176/ [42] X. Guan, J. Zeng, F. Meng, C. Xin, Y. Lu, H. Lin, X. Han, L. Sun, and J. Zhou, “DeepRAG: Thinking to retrieve step by step for large language models,” in The Fourteenth International Conference on Learning Representations, 2026. [Online]. Available: https://openreview.net/forum?id=VI2YaggHIF [43] S. Jeong, J. Baek, S. Cho, S. J. Hwang, and J. C. Park, “Adaptive-rag: Learning to adapt retrieval-augmented large language models through question complexity,” in Proceedings of the 2024 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies (Volume 1: Long Papers), 2024, pp. 7036–7050. [44] H. Tan, F. Sun, W. Yang, Y. Wang, Q. Cao, and X. Cheng, “Blinded by generated contexts: How language models merge generated and retrieved contexts when knowledge conflicts?” in Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2024, pp. 6207–6227. [45] T. Guo, X. Chen, Y. Wang, R. Chang, S. Pei, N. V. Chawla, O. Wiest, and X. Zhang, “Large language model based multi-agents: A survey of progress and challenges,” in Proceedings of the Thirty-Third International Joint Conference on Artificial Intelligence, IJCAI-24, K. Larson, Ed. International Joint Conferences on Artificial Intelligence Organization, 8 2024, pp. 8048–8057, survey Track. [Online]. Available: https://doi.org/10.24963/ijcai.2024/890 [46] X. Li, S. Wang, S. Zeng, Y. Wu, and Y. Yang, “A survey on llmbased multi-agent systems: workflow, infrastructure, and challenges,” Vicinagearth, 2024. [47] J. Yang, C. E. Jimenez, A. Wettig, K. Lieret, S. Yao, K. R. Narasimhan, and O. Press, “SWE-agent: Agent-computer interfaces enable automated software engineering,” in The Thirty-eighth Annual Conference on Neural Information Processing Systems, 2024. [Online]. Available: https://arxiv.org/abs/2405.15793 [48] LLVM, “Memorysanitizer - official documentation,” https://clang.llvm.o rg/docs/MemorySanitizer.html, Accessed 2026. [49] B. Mathis, R. Gopinath, and A. Zeller, “Learning input tokens for effective fuzzing,” in Proceedings of the 29th ACM SIGSOFT International Symposium on Software Testing and Analysis, ser. ISSTA 2020. New York, NY, USA: Association for Computing Machinery, 2020, p. 27–37. [Online]. Available: https://doi.org/10.1145/3395363.3397348 [50] A. Herrera, H. Gunadi, S. Magrath, M. Norrish, M. Payer, and A. L. Hosking, “Seed selection for successful fuzzing,” in Proceedings of the 30th ACM SIGSOFT International Symposium on Software Testing and Analysis, ser. ISSTA 2021. New York, NY, USA: Association for Computing Machinery, 2021, p. 230–243. [Online]. Available: https://doi.org/10.1145/3460319.3464795 [51] A. Fioraldi, D. Maier, H. Eißfeldt, and M. Heuse, “AFL++ : Combining incremental steps of fuzzing research,” in 14th USENIX Workshop on Offensive Technologies (WOOT 20), 2020. [52] D. Guo, D. Yang, H. Zhang, J. Song, P. Wang, Q. Zhu, R. Xu, R. Zhang, S. Ma, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin,
F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Ding, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Chen, J. Yuan, J. Tu, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. You, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Zhou, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang, “Deepseek-r1 incentivizes reasoning in llms through reinforcement learning,” Nature, vol. 645, no. 8081, p. 633–638, Sep. 2025. [Online]. Available: http://dx.doi.org/10.1038/s41586-025-09422-z [53] S. Yao, J. Zhao, D. Yu, N. Du, I. Shafran, K. Narasimhan, and Y. Cao, “React: Synergizing reasoning and acting in language models,” arXiv preprint arXiv:2210.03629, 2022. [54] OpenAI, “A practical guide to building agents,” https://cdn.openai.com/b usiness-guides-and-resources/a-practical-guide-to-building-agents.pdf, Accessed 2026. [55] M. team., “Writing effective tools for agents,” https://modelcontextprot ocol.info/docs/tutorials/writing-effective-tools/, Accessed 2026. [56] O. S. S. F. (OpenSSF), “Fuzz introspector – introspect, extend and optimise fuzzers,” Accessed 2022. [Online]. Available: https: //github.com/ossf/fuzz-introspector [57] G. Savidov and A. Fedotov, “Casr-Cluster: Crash clustering for linux applications,” in 2021 Ivannikov ISPRAS Open Conference (ISPRAS). IEEE, 2021, pp. 47–51. [58] I. Free Software Foundation, “Gdb non-interactive batch mode,” https: //www.sourceware.org/gdb/current/onlinedocs/gdb.html/Mode-Options .html, Accessed 2026. [59] B. Liu, X. Li, J. Zhang, J. Wang, T. He, S. Hong, H. Liu, S. Zhang, K. Song, K. Zhu, Y. Cheng, S. Wang, X. Wang, Y. Luo, H. Jin, P. Zhang, O. Liu, J. Chen, H. Zhang, Z. Yu, H. Shi, B. Li, D. Wu, F. Teng, X. Jia, J. Xu, J. Xiang, Y. Lin, T. Liu, T. Liu, Y. Su, H. Sun, G. Berseth, J. Nie, I. Foster, L. Ward, Q. Wu, Y. Gu, M. Zhuge, X. Liang, X. Tang, H. Wang, J. You, C. Wang, J. Pei, Q. Yang, X. Qi, and C. Wu, “Advances and challenges in foundation agents: From brain-inspired intelligence to evolutionary, collaborative, and safe systems,” 2025. [Online]. Available: https://arxiv.org/abs/2504.01990 [60] S. Han, Q. Zhang, Y. Yao, W. Jin, and Z. Xu, “Llm multi-agent systems: Challenges and open problems,” 2025. [Online]. Available: https://arxiv.org/abs/2402.03578 [61] C. S. Xia, Y. Deng, S. Dunn, and L. Zhang, “Demystifying llm-based software engineering agents,” Proc. ACM Softw. Eng., vol. 2, no. FSE, Jun. 2025. [Online]. Available: https://doi.org/10.1145/3715754 [62] LLVM, “Source-based code coverage,” Accessed 2026. [Online]. Available: https://clang.llvm.org/docs/SourceBasedCodeCoverage.html [63] travitch, “Whole program llvm (wllvm),” https://github.com/travitch/wh ole-program-llvm, Accessed 2026. [64] C. Aschermann, S. Schumilo, T. Blazytko, R. Gawlik, and T. Holz, “Redqueen: Fuzzing with input-to-state correspondence,” in Symposium on Network and Distributed System Security (NDSS), 2019. [65] C. Lemieux, J. P. Inala, S. K. Lahiri, and S. Sen, “Codamosa: Escaping coverage plateaus in test generation with pre-trained large language models,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), 2023, pp. 919–931. [66] DeepSeek-AI, A. Liu, A. Mei, B. Lin, B. Xue, B. Wang, B. Xu, B. Wu, B. Zhang, C. Lin, C. Dong, C. Lu, C. Zhao, C. Deng, C. Xu, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, E. Li, F. Zhou, F. Lin, F. Dai, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Li, H. Liang, H. Wei, H. Zhang, H. Luo, H. Ji, H. Ding, H. Tang, H. Cao, H. Gao, H. Qu, H. Zeng, J. Huang, J. Li, J. Xu, J. Hu, J. Chen, J. Xiang, J. Yuan, J. Cheng, J. Zhu, J. Ran, J. Jiang, J. Qiu, J. Li, J. Song, K. Dong,
16
K. Gao, K. Guan, K. Huang, K. Zhou, K. Huang, K. Yu, L. Wang, L. Zhang, L. Wang, L. Zhao, L. Yin, L. Guo, L. Luo, L. Ma, L. Wang, L. Zhang, M. S. Di, M. Y. Xu, M. Zhang, M. Zhang, M. Tang, M. Zhou, P. Huang, P. Cong, P. Wang, Q. Wang, Q. Zhu, Q. Li, Q. Chen, Q. Du, R. Xu, R. Ge, R. Zhang, R. Pan, R. Wang, R. Yin, R. Xu, R. Shen, R. Zhang, S. H. Liu, S. Lu, S. Zhou, S. Chen, S. Cai, S. Chen, S. Hu, S. Liu, S. Hu, S. Ma, S. Wang, S. Yu, S. Zhou, S. Pan, S. Zhou, T. Ni, T. Yun, T. Pei, T. Ye, T. Yue, W. Zeng, W. Liu, W. Liang, W. Pang, W. Luo, W. Gao, W. Zhang, X. Gao, X. Wang, X. Bi, X. Liu, X. Wang, X. Chen, X. Zhang, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yu, X. Li, X. Yang, X. Li, X. Chen, X. Su, X. Pan, X. Lin, X. Fu, Y. Q. Wang, Y. Zhang, Y. Xu, Y. Ma, Y. Li, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Qian, Y. Yu, Y. Zhang, Y. Ding, Y. Shi, Y. Xiong, Y. He, Y. Zhou, Y. Zhong, Y. Piao, Y. Wang, Y. Chen, Y. Tan, Y. Wei, Y. Ma, Y. Liu, Y. Yang, Y. Guo, Y. Wu, Y. Wu, Y. Cheng, Y. Ou, Y. Xu, Y. Wang, Y. Gong, Y. Wu, Y. Zou, Y. Li, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Z. F. Wu, Z. Z. Ren, Z. Zhao, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Gou, Z. Ma, Z. Yan, Z. Shao, Z. Huang, Z. Wu, Z. Li, Z. Zhang, Z. Xu, Z. Wang, Z. Gu, Z. Zhu, Z. Li, Z. Zhang, Z. Xie, Z. Gao, Z. Pan, Z. Yao, B. Feng, H. Li, J. L. Cai, J. Ni, L. Xu, M. Li, N. Tian, R. J. Chen, R. L. Jin, S. S. Li, S. Zhou, T. Sun, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Song, X. Zhou, Y. X. Zhu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Z. Huang, Z. Xu, Z. Zhang, D. Ji, J. Liang, J. Guo, J. Chen, L. Xia, M. Wang, M. Li, P. Zhang, R. Chen, S. Sun, S. Wu, S. Ye, T. Wang, W. L. Xiao, W. An, X. Wang, X. Sun, X. Wang, Y. Tang, Y. Zha, Z. Zhang, Z. Ju, Z. Zhang, and Z. Qu, “Deepseek-v3.2: Pushing the frontier of open large language models,” 2025. [Online]. Available: https://arxiv.org/abs/2512.02556 [67] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th Symposium on Operating Systems Principles, ser. SOSP ’23. New York, NY, USA: Association for Computing Machinery, 2023, p. 611–626. [Online]. Available: https://doi.org/10.1145/3600006.3613165 [68] DeepSeek-AI, “Hugging face: Deepseek v3.2 model,” https://huggingfac e.co/deepseek-ai/DeepSeek-V3.2, Accessed 2026. [69] LLVM, “llvm-cov - emit coverage information,” Accessed 2026. [Online]. Available: https://llvm.org/docs/CommandGuide/llvm-cov.html [70] G. Klees, A. Ruef, B. Cooper, S. Wei, and M. Hicks, “Evaluating fuzz testing,” in Proceedings of the 2018 ACM SIGSAC Conference on Computer and Communications Security, ser. CCS ’18. New York, NY, USA: Association for Computing Machinery, 2018, p. 2123–2138. [Online]. Available: https://doi.org/10.1145/3243734.3243804 [71] P. developers, “Promptfuzz author response to its official release,” https: //github.com/FuzzAnything/PromptFuzz/releases/tag/v1.0.0, Accessed 2026. [72] ——, “Can promefuzz be used to fuzz openssl?” https://github.com/pvz 122/PromeFuzz/issues/8, Accessed 2026. [73] J. Jiang, H. Xu, and Y. Zhou, “Rulf: Rust library fuzzing via api dependency graph traversal,” in 2021 36th IEEE/ACM International Conference on Automated Software Engineering (ASE). IEEE, 2021, pp. 581–592. [74] Y. Deng, C. S. Xia, H. Peng, C. Yang, and L. Zhang, “Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models,” in Proceedings of the 32nd ACM SIGSOFT International Symposium on Software Testing and Analysis, 2023, pp. 423–435. [75] C. S. Xia, M. Paltenghi, J. L. Tian, M. Pradel, and L. Zhang, “Universal fuzzing via large language models,” arXiv preprint arXiv:2308.04748, 2023. [76] H. Xu, Y. Zhao, and H. Wang, “Directed greybox fuzzing via large language model,” 2025. [Online]. Available: https: //arxiv.org/abs/2505.03425 [77] “Github rest api,” https://docs.github.com/en/rest?apiVersion=2022-11-2 8, Accessed 2026. [78] Janne, “Agent design lessons from claude code,” https://jannesklaas.gith ub.io/ai/2025/07/20/claude-code-agent-design.html, Accessed 2026.
A PPENDIX A. Open Science We are committed to reproducible research. However, as discussed in the Ethical Considerations section, the dual-use potential of FuzzAgent precludes an open-source release of its
implementation at this time. To allow independent verification of the reported results, we provide the full evaluation dataset, including raw coverage measurements, crash artifacts, bug reports, and harness collections, at: https://github.com/FuzzA nything/FuzzAgent-Artifacts. B. Implementation FuzzAgent is implemented in 17k lines of Python code. The system architecture is primarily divided into interface implementations and strategies. For the interface implementations, we adhere to the principles of modularity and robustness recommended in prior studies [54], [55]. We provide a comprehensive list of the tools designed for the current prototype in Section E. In the Computer Usage Interface, tools for file navigation and editing are adapted from SWE-Agent [47], while the remaining tools are custom-designed for library fuzzing tasks. The Web Search Interface is built upon the GitHub API [77]. The Coverage Analysis Interface relies on source-based code coverage [62] and llvm-cov [69]. The Crash Debugging Interface incorporates CASR [57] and the noninteractive batch mode of GDB [58]. To ensure system stability, we implement agent hooking and validation mechanisms to manage unexpected situations gracefully. Specifically, we added hooking functions at each agent’s exit point to validate outputs within the Environment and remind the agent of the intended results [78]. Regarding strategies, we designed specialized system prompts for each agent to implement their core capabilities; these are provided in Section F. C. Statistical Variance in Library Fuzzing Library fuzzing contains inherent randomness: even when the same target and fuzzer are used, different fuzzing runs may explore different paths and produce different coverage results. Prior work therefore recommends repeated fuzzing trials and statistical analysis when evaluating fuzzers [70]. In LLMbased library fuzzing, this issue becomes more pronounced because randomness is introduced not only by fuzzing, but also by harness generation. Different LLM runs can synthesize different API sequences in harnesses, which may substantially change the reachable code paths before fuzzing even starts. We therefore quantify both sources of variation. We use the coefficient of variation (CV) as a scalenormalized measure of dispersion: σ CV = × 100%, µ where µ and σ denote the mean and standard deviation of branch coverage across repeated trials. CV is preferable to raw variance here because the evaluated libraries differ greatly in size and absolute branch coverage; normalizing by the mean makes variation comparable across libraries. Our variance experiment follows a nested design. For each LLM-based harness generation method and each library, we performed 5 independent harness-generation trials. After each generation trial, the valid harnesses for the library were merged into one composite harness. We then ran 5 independent 24hour fuzzing trials for each composite harness, resulting in
17
25 independent 24-hour fuzzing runs per method-library pair. Let ci,j denote the branch coverage from the j-th fuzzing trial of the harness produced by the i-th generation trial, where i, j ∈ {1, . . . , 5}. To measure variation from harness generation, we first average over P5 fuzzing randomness for each generated harness, mi = 15 j=1 ci,j , and then compute the CV over {m1 , . . . , m5 }. To measure variation from fuzzing, we compute the CV over {ci,1 , . . . , ci,5 } within each generation trial and then average these 5 CV values. Table V reports the resulting CV values. Across available libraries, the CV caused by harness generation is consistently larger than the CV caused by repeated fuzzing with a fixed generated harness. This result indicates that LLM harness generation is the dominant source of statistical variation in LLM-based library fuzzing. Consequently, evaluating such systems with only one generation trial can produce unstable and potentially misleading results. Our main evaluation therefore repeats both harness generation and 24-hour fuzzing, so the reported coverage numbers are less sensitive to random outcomes from a single LLM run. TABLE V C OEFFICIENT OF VARIATION ACROSS INDEPENDENT HARNESS - GENERATION AND FUZZING TRIALS . Library
FuzzAgent
PromeFuzz
PromptFuzz
Harness
Fuzzing
Harness
Fuzzing
Harness
Fuzzing
cJSON libmagic RE2 pugixml zlib c-ares liblouis libpng libpcap libtiff lcms tinygltf libjpeg-turbo curl libvpx SQLite3 libaom protobuf OpenSSL OpenCV
2.66% 5.57% 2.71% 2.73% 2.04% 3.46% 10.36% 4.44% 4.02% 2.80% 9.25% 7.61% 7.38% 30.36% 24.28% 13.60% 7.13% 15.03% 28.46% 17.47%
0.14% 2.37% 0.52% 0.42% 0.40% 0.66% 0.64% 1.11% 2.21% 1.27% 1.38% 2.58% 0.88% 1.59% 0.92% 2.94% 0.92% 1.28% 1.96% 1.45%
5.01% 4.76% 1.59% 4.45% 4.99% 4.30% 11.48% 31.09% 5.30% 12.35% 5.70% 5.33% 16.98% 18.34% 23.54% 4.49% 12.49% N/A N/A N/A
3.50% 4.97% 4.59% 1.76% 1.01% 2.67% 1.53% 6.04% 4.56% 11.46% 6.76% 5.46% 12.30% 1.72% 8.60% 4.34% 6.69% N/A N/A N/A
2.25% 16.71% 8.59% N/A 3.88% 4.87% 16.23% 22.34% 18.39% 9.22% 5.72% N/A 36.23% 6.90% 6.88% 29.31% 6.96% N/A 12.85% N/A
0.05% 3.70% 2.94% N/A 0.79% 0.50% 8.31% 1.97% 2.50% 4.82% 1.20% N/A 6.74% 1.24% 2.64% 3.26% 0.87% N/A 0.51% N/A
Mean
10.07%
1.28%
10.13%
5.17%
12.96%
2.63%
Note: CV denotes coefficient of variation. The harness columns compute CV over the mean coverage of 5 fuzzing trials for each of 5 independent harnessgeneration trials. The fuzzing columns compute the within-harness CV over 5 independent 24-hour fuzzing trials and then average the CV across the 5 generated harnesses. N/A indicates unsupported or unavailable settings.
D. Mann-Whitney U Test for Statistical Significance Because coverage values are not guaranteed to follow a normal distribution, we use the non-parametric Mann–Whitney U test to assess whether the observed coverage differences between FuzzAgent and each baseline are statistically significant. Table VI reports the per-library p-values. These results complement the mean coverage numbers in Table I by showing that FuzzAgent’s improvements are statistically significant for most available library-baseline pairs.
TABLE VI M ANN –W HITNEY U- TEST P - VALUES COMPARING F UZZ AGENT WITH BASELINE FUZZERS . Library
PromeFuzz
PromptFuzz
OSS-Fuzz
OSS-Fuzz-Gen
cJSON libmagic RE2 pugixml zlib c-ares liblouis libpng libpcap libtiff lcms tinygltf libjpeg-turbo curl libvpx SQLite3 libaom protobuf OpenSSL OpenCV
1.37E-08 2.37E-09 1.41E-09 1.39E-09 1.39E-09 4.61E-05 1.39E-09 1.41E-09 1.41E-09 1.84E-08 2.55E-09 1.41E-09 1.40E-09 0.0074139 1.39E-09 3.02E-07 1.42E-09 N/A N/A N/A
2.04E-08 0.132644 1.41E-09 N/A 2.25E-09 0.196896 0.00216183 1.41E-09 1.41E-09 1.42E-09 1.16E-06 N/A 1.40E-09 0.000444748 1.39E-09 0.019891 1.41E-09 N/A 3.90E-05 N/A
8.85E-11 9.73E-11 9.73E-11 9.49E-11 9.56E-11 9.73E-11 0.19826 9.69E-11 9.71E-11 9.73E-11 9.61E-11 9.71E-11 0.00447948 0.000105049 9.50E-11 0.198457 9.73E-11 9.73E-11 9.73E-11 9.73E-11
8.85E-11 9.73E-11 9.73E-11 9.49E-11 9.56E-11 9.73E-11 9.54E-11 9.69E-11 9.71E-11 9.73E-11 9.61E-11 9.71E-11 9.61E-11 9.72E-11 9.50E-11 0.0711539 9.73E-11 9.73E-11 9.73E-11 9.73E-11
p < 0.05
17
14
18
19
Library Builder # Role Definition You are a resilient and expert **Build Automation Engineer**. Your goal is NOT just to write a script, but to **guarantee a successful build** that produces valid static (‘.a‘) libraries. ## Core Responsibility You are responsible for the entire **BuildTest-Fix** cycle. 1. **Draft**: Create the initial ‘build.sh‘. 2. **Execute**: Run the build using ‘ run_and_check_build_script‘. 3. **Debug**: If the build fails, YOU MUST ANALYZE THE LOGS, FIX THE SCRIPT, AND RETRY. 4. **Deliver**: Only exit when ‘\$WORK/lib‘ contains the required artifacts. 5. **CRITICAL:** Do not exit simply because the build failed. A build failure is a demand for a fix, not a reason to quit.
Note: Each entry reports the p-value of a two-sided Mann–Whitney U test comparing the coverage distribution of FuzzAgent against the corresponding baseline on the same library. N/A indicates unavailable baseline results. The final row counts libraries where the difference is statistically significant at p < 0.05.
Fig. 6. The system prompt for the Library Builder agent in FuzzAgent.
E. Interface Implementations We built the interfaces to facilitate agents’ interaction with the working environment. These interfaces are implemented as a set of tools abstract the tasks in library fuzzing workflow. The url https://github.com/FuzzAnything/FuzzAgent-Artifacts/i nterfaces.md contains the detailed descriptions of these tools, including their input/output formats and example usages. F. The System Prompt for FuzzAgent Each agent in FuzzAgent is driven by a system prompt that consists of three parts: (i) the task responsibility, which states what the agent must accomplish and the success criteria; (ii) the core strategy, which encodes the methodology the agent should follow when invoking the interface tools; and (iii) a small set of few-shot examples that illustrate expected inputs, intermediate reasoning, and outputs. Among them, the core strategy is the part most relevant to the design of FuzzAgent. For brevity, in the figures below we list only the core strategy of each agent’s system prompt.
Dictionary Generator # Role Definition You are the **Dictionary Acquisition & Optimization Specialist**. Your goal is to provide a *lean, high-impact* fuzzing dictionary. You must locate existing resources from the OSS-Fuzz repository--whether they are stored locally or downloaded dynamically---and rigorously prune them. ### Discovery & Matching 1. **Analyze Target**: Determine the ** Protocol/Format** of the target project (e.g ., "It’s a Video Codec"). 2. **Search**: Run ‘search_web list‘. 3. **Select Source**: - **Exact Match**: Target=‘libaom‘, OSSFuzz=‘libaom‘. - **Protocol Match**: Target=‘ my_video_lib‘, OSS-Fuzz=‘ffmpeg‘. - **Select Criteria**: All exact matched and protocol matched projects. Use tool ‘ track_web_retrieve_progress‘ to record matched projects waiting for retrieval. - **Fallback**: If no match found, create a minimal dictionary based on standard protocol knowledge.
Fig. 7. The system prompt for the Dictionary Generator agent in FuzzAgent.
18
Seed Generator # Role Definition You are the **Seed Corpus Acquisition Agent **. Your sole responsibility is to locate, download, and install high-quality fuzzing seed corpora for the target project. ## Discovery & Matching 1. **Analyze Target**: Determine the ** Protocol/Format** of the target project (e.g ., "It’s a Video Codec"). 2. **Search**: Run ‘search_web list‘. 3. **Select Source**: - **Exact Match**: Target=‘libaom‘, OSSFuzz=‘libaom‘. - **Protocol Match**: Target=‘ my_video_lib‘, OSS-Fuzz=‘ffmpeg‘. - **Select Criteria**: All exact matched and protocol matched projects. Use tool ‘ track_web_retrieve_progress‘ to record matched projects waiting for retrieval.
Coverage Analyzer: API-Surface Exploration ## Analysis Strategy ### PHASE 1: Surface Coverage Exploration ( API Utilizaiton) **Trigger:** High-value public APIs are untouched. **Goal:** Identify a group of related, uncovered APIs to guide the generation of a new, high-impact fuzzing harness. **Workflow:** 1. **Select Targets:** From the Module, File and API-Level coverage data, prioritize the APIs with the highest undiscovered complexity. 2. **Group Clusters:** Look for other uncovered APIs in the same files or modules that share a logical relationship (e.g., ‘ ArrayCreate‘, ‘ArrayInsert‘, ‘ArrayDelete‘). 3. **Reason Relations:** Reasoning relationships among identified APIs and thinking how to organize the related ones into an invocation sequence. 4. **Analyze Dependencies:** Complement the invocation sequence and determine the correct lifecycle: - *Initialization*: What must be called first? (e.g., ‘Init‘, ‘New‘, ‘Parse‘) - *Operation*: The target invocation sequence. - *Cleanup*: What frees the memory? (e.g., ‘Free‘, ‘Destroy‘) - *Helpers*: Do you need ‘CreateString‘ to test ‘DictionaryAdd‘? 5. **Formulate Recommendation:** Formulate these into a single harness request.
Fig. 8. The system prompt for the Seed Generator agent in FuzzAgent.
Harness Generator # New Harness Generation Strategy ## Coverage-guided Principles **Primary Objective**: Generate a harness according CoverageAnalyzerAgent’s guides. **API Selection Strategy** (in priority order): 1. **Manager-Specified Targets**: Always prioritize APIs explicitly mentioned in Manager specifications 2. **Dependency Chain Completion**: Include necessary helper functions for proper API initialization and cleanup **Coverage-Driven Design Rules**: - Call as many target APIs has dependencies as possible to maximize coverage - Include input validation and error handling paths where possible - Implement necessary state setup for complex API sequences - Ensure proper resource management ( allocation/deallocation patterns)
Fig. 10. The system prompt for the API-Surface Exploration strategic in Coverage Analyzer agent in FuzzAgent.
Fig. 9. The system prompt for the Harness Generator agent in FuzzAgent.
19
Crash Analyzer # Role Definition You are the **Crash Analysis & Triage Specialist**. Your sole purpose is to investigate a specific crash artifact, determine the "Blame" (Library Bug vs. Harness Bug), and file a formal report. ## Core Mission You act as a Judge. You have two suspects: 1. **The Library**: Did it fail to handle valid input safe? (Genuine Bug) 2. **The Harness**: Did it violate the API contract or manage memory poorly? (Invalid Bug) ## Mandatory Workflow ### Phase 1: Forensics (Data Gathering) 1. Receive the **Crash Artifact Path** (from user input). 2. Call ‘crash_initial_analysis‘ with the crash artifact path to get call traces. 3. **Identify the "Crash Point"**: The topmost stack frame that belongs to the project (skip standard library frames like ‘libc.so‘ or ‘asan_report‘). ### Phase 2: Debugging (Runtime Information Gathering) 1. Call ‘crash_context_inspection‘ on the suspect API or function to inspect the runtime context frames of this invocation. 2. Analyze the runtime context frames to identify the point of failure iteratively. 3. Read the file/documentation of the suspect API or function to verify the preconditions and postconditions. 4. Repeat the process until the point of failure is identified.
Coverage Analyzer: Deep Stated Exploration ## Analysis Strategy ### PHASE 2: Deep Coverage (Blocker Resolution) **Condition:** Overall API coverage is high enough (API Coverage >= 90%) or the API coverage growth is stagnant. **Trigger:** Public APIs are hit, but internal functions and code branches are blocked. **Goal:** Analyze the blocker cause and suggest a targeted harness to break this. **Workflow:** 1. **Identify Blocker:** Inspect the Function Level and Branch Level coverage, pick the one with the highest "Blocked Complexity". 2. **Trace Entry Point:** Analyze the call chains on the function containing the blocker. Identify which Public API reaches this code. 3. **Analyze Condition:** Read the source code with coverage hits around the blocker. What condition is failing? Is this caused by unsatisfied API calls? If not, select the next blocker to analyze. 4. **Formulate Recommendation:** Instruct the HarnessAgent to create a specific scenario that satisfies this condition.
Fig. 11. The system prompt for the Deep Stated Exploration strategy in Coverage Analyzer agent in FuzzAgent. Fig. 12. The system prompt for the Crash Analyzer agent in FuzzAgent.
20
cJSON
libmagic
RE2
pugixml
zlib
c-ares
liblouis
libpng
libpcap
libtiff
lcms
tinygltf
libjpeg-turbo
curl
libvpx
SQLite3
libaom
protobuf
OpenSSL
OpenCV
build dict seed harness fuzzer crash coverage
build dict seed harness fuzzer crash coverage
build dict seed harness fuzzer crash coverage
build dict seed harness fuzzer crash coverage 0
6
12
18
24
0
6
build
12
18
dict
24
seed
0
6
12
harness
18
fuzzer
24
0
crash
6
12
18
24
0
6
12
18
24
coverage
Fig. 13. The execution trajectory of one trail of FuzzAgent’s agents across 24 hours. The horizontal axis represents time, while the vertical axis lists the different agents involved in the library fuzzing process. The agents from top to bottom are: Library Builder, Dictionary Generator, Seed Generator, Harness Generator, Fuzzer Executor, Coverage Analyzer and Crash Analyzer. Each colored block indicates a specific task undertaken by an agent, with the length of the block corresponding to the duration of that task. This visualization highlights how FuzzAgent dynamically allocates tasks among agents over time to optimize fuzzing efficiency and effectiveness.
Fig. 14. The project relation graph for target libraries in the Dictionary Generator and Seed Generator of our evaluation. The red nodes represent the target projects in our evaluations, and the green nodes represent the related projects FuzzAgent refers to generating the dictionary and seeds. This graph illustrates the interconnectedness of the selected projects, highlighting potential areas where knowledge transfer and shared fuzzing strategies could be beneficial.
21
TABLE VII: Library bugs discovered by FuzzAgent and detection coverage of baselines (✓ detected, ✗ missed). St.: R = reported, C = confirmed/fixed. All bugs in this table were detected by FuzzAgent. ID Library
Crash Location
Vulnerability Type St. OSS-Fuzz OSS-Fuzz-Gen PromptFuzz PromeFuzz
1 2 3 4 5 6 7 8 9 10 11 12 13 14 15 16 17 18 19 20 21 22 23 24 25 26 27 28 29 30 31 32 33 34 35 36 37 38 39 40 41 42 43 44 45 46 47 48 49 50 51 52 53 54 55 56 57
cJSON.c:2524 cJSON.c:2001 cJSON.c:1963 softmagic.c:771 encoding.c:286 pugixml.cpp:5207 pugixml.cpp:232 pugixml.cpp:4565 gzlib.c:389 ares process.c:1156 pattern.c:801 compileTranslationTable.c:430 compileTranslationTable.c:296 compileTranslationTable.c:4865 lou translateString.c:1343 logging.c:57 lou backTranslateString.c:257 lou translateString.c:1343 pngwutil.c:2182 png.c:495 pngrtran.c:1070 pcap-util.c:567 bpf filter.c:112 tif unix.c:338 tif write.c:775 tif unix.c:357 tif dirwrite.c:2159 tif luv.c:1318 tif dirwrite.c:896 tif dirinfo.c:606 cmsio0.c:1530 cmsio0.c:1606 cmsnamed.c:808 turbojpeg.c:937 jcarith.c:438 cmyk.h:55 vpx image.c:263 firstpass.c:678 bitwriter buffer.c:48 encodeframe.c:831 vpx dsp common.h:89 vp9 cx iface.c:349 vp9 svc layercontext.c:468 vp9 ratectrl.c:2172 vp9 svc layercontext.c:1334 vp9 cx iface.c:501 vp9 encoder.h:1079 vp9 quantize.c:324 onyx if.c:970 vp9 encodeframe.c:580 sad4d avx512.c:35 vp9 bitstream.c:59 vp9 rd.c:351 vp9 decoder.c:311 highbd sad avx2.c:29 vp9 svc layercontext.c:1007 highbd variance impl sse2.asm:217
Integer Overflow Use After Free Type Mismatch Null Pointer Segment Violation Buffer Overflow Segment Violation Memory Alignment Integer Overflow Integer Overflow Segment Violation Segment Violation Segment Violation Memory Leak Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Integer Overflow Memory Alignment Abort Integer Overflow Buffer Overflow Integer Overflow Buffer Overflow Integer Overflow Buffer Overflow Integer Overflow Documentation Documentation Null Pointer Type Mismatch Type Mismatch Type Mismatch Null Pointer Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Segment Violation Segment Violation Segment Violation Segment Violation
cJSON cJSON cJSON libmagic libmagic pugixml pugixml pugixml zlib c-ares liblouis liblouis liblouis liblouis liblouis liblouis liblouis liblouis libpng libpng libpng libpcap libpcap libtiff libtiff libtiff libtiff libtiff libtiff libtiff lcms lcms lcms libjpeg-turbo libjpeg-turbo libjpeg-turbo libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx libvpx
R R R R R R R R C C C C C C C C C C R R R C C R R R R R R R C C C R R R C C C C C C C C C C C C C C C C C C C C C
✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗
✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗
✗ ✗ ✓ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ ✗ ✗ ✓ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗
✗ ✗ ✓ ✗ ✗ ✗ ✗ ✓ ✗ ✓ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ ✗ ✗ ✓ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗
continued on next page
22
Table VII continued from previous page ID Library
Crash Location
Vulnerability Type St. OSS-Fuzz OSS-Fuzz-Gen PromptFuzz PromeFuzz
58 59 60 61 62 63 64 65 66 67 68 69 70 71 72 73 74 75 76 77 78 79 80 81 82 83 84 85 86 87 88 89 90 91 92 93 94 95 96 97 98 99 100 101 102
variance.c:269 vp9 encoder.c:4104 onyxd if.c:126 vp9 bitstream.c:402 vp8 dx iface.c:134 onyx if.c:1233 vp9 encodeframe.c:5471 sqlite3.c:263698 firstpass.c:830 pass2 strategy.c:308 intra mode search utils.h:107 var based part.c:419 ratectrl.c:508 variance.c:280 noise model.c:1270 psnr.c:65 encodetxb.c:615 av1 dx iface.c:1330 svc layercontext.c:444 pass2 strategy.c:1215 svc layercontext.c:315 aq cyclicrefresh.c:237 encodetxb.c:617 av1 fwd txfm2d avx2.c:1450 av1 ext ratectrl.c:186 av1 dx iface.c:953 intra mode search.c:358 encoder.c:2314 bitstream.c:2475 coded stream.cc:243 coded stream.cc:735 ssl ciph.c:1241 obj lib.c:62 dh kmgmt.c:89 md32 common.h:158 bsearch.c:28 stack.c:443 count non zero.dispatch.cpp:148 drawing.cpp:1030 intrin sse.hpp:3080 drawing.cpp:1121 mathfuncs.cpp:141 types.hpp:1245 drawing.cpp:1662 tree.cpp:522
Segment Violation Segment Violation Segment Violation Segment Violation Assertion Fail Assertion Fail Assertion Fail Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Integer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Buffer Overflow Segment Violation Segment Violation Assertion Fail Assertion Fail Assertion Fail Null Pointer Segment Violation Segment Violation Null Pointer Null Pointer Type Mismatch Type Mismatch Type Mismatch Type Mismatch Type Mismatch Memory Alignment Integer Overflow Integer Overflow Integer Overflow Integer Overflow Buffer Overflow
libvpx libvpx libvpx libvpx libvpx libvpx libvpx sqlite libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom libaom protobuf protobuf openssl openssl openssl openssl openssl openssl opencv opencv opencv opencv opencv opencv opencv opencv
23
C C C C C C C R C C C C C C C C C C C C C C C C C C C C C C C C C C C C C C C C C C C R R
✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗
✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗
✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗
✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ ✗ ✓ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✓ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗ ✗