Conceptio › Archive › arXiv CS
arXiv CSopen access

FunFuzz: An LLM-Powered Evolutionary Fuzzing Framework

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
cryptographycybersecurityprivacysecurity
cryptography, security, privacy, cybersecurity

FunFuzz: An LLM-Powered Evolutionary Fuzzing Framework Mario Rodríguez Béjar

Bernardino Romera Paredes

Jose Luis Hernández-Ramos

[email protected] Universidad de Murcia Murcia, Spain

[email protected] Hiverge London, England

[email protected] Universidad de Murcia Murcia, Spain

or new compiler subsystems often requires substantial effort and expertise [6]. Another limitation concerns adaptability: modern Modern fuzzers increasingly use Large Language Models (LLMs) compilers evolve quickly, adding new features and optimizations to generate structured inputs, but LLM-driven fuzzing is sensitive that render previously effective fuzzing strategies less potent. Preto prompt initialization and sampling variance, which can reduce vious studies [3] show that even long-standing tools like Csmith, exploration efficiency and lead to redundant inputs. We present once highly effective, struggle to uncover new defects in the latFunFuzz, a multi-island evolutionary fuzzing framework that runs est versions of GCC and Clang. A third limitation is sustained several isolated searches in parallel and periodically migrates highsemantic diversity. Many approaches bias exploration toward a value candidates to maintain diversity. FunFuzz derives initial gentractable subset of language constructs or toward neighborhoods eration prompts from documentation and initializes islands with around a fixed seed corpus, which leaves rare feature interactions topic-specific instructions, then continuously adapts prompts using under-exercised [5]. Together, these limitations motivate fuzzing feedback-guided selection. During fuzzing, candidates are prioritechniques that reduce target-specific engineering, remain effectized by incremental compiler coverage, while compiler-internal tive as compilers evolve, and maintain exploration pressure toward failure signals are used to identify crash-inducing inputs. We evalusemantically rich programs rather than converging to shallow patate FunFuzz on compiler fuzzing, where inputs are source programs terns. and success is measured by compiler coverage and unique compilerLarge language models (LLMs) [23] offer a promising alternative internal failures. Across repeated 24-hour campaigns on GCC and to manual generators and corpus-heavy fuzzing because they can Clang, FunFuzz achieves higher compiler coverage than previous synthesize source programs that are syntactically well-formed and LLM-driven baselines and discovers more unique failure-triggering semantically varied with limited target-specific engineering [8]. inputs. Prior LLM-assisted fuzzers demonstrate that prompt-guided generation can bootstrap fuzzing on complex targets where traditional Keywords input generators are expensive to build [18]. However, LLM-driven LLM-guided fuzzing, compiler fuzzing, evolutionary fuzzing, multifuzzing faces a practical gap: generation can collapse to repetitive island optimization, feedback-guided generation, coverage-guided program styles, early samples can bias the trajectory of a camtesting, prompt distillation, crash detection paign, and exploration often stagnates without an explicit mechanism that preserves diversity while still exploiting useful feedback. 1 INTRODUCTION This motivates a search procedure that remains robust under LLM Language-processing systems such as compilers, interpreters, and stochasticity and sustains long-term exploration pressure toward runtime engines underpin modern software development and toolchains. new compiler behaviors rather than repeatedly refining a narrow Defects in these components can propagate widely, affecting downset of patterns. stream applications and the broader software supply chain. Over Existing LLM-assisted fuzzers such as Fuzz4All rely on prompt the past decades, fuzzing has proven highly effective at uncovering distillation to initialize generation, but typically start from a single real bugs in complex systems at scale [14]. In particular, compiler shared prompt, which can bias the search trajectory early and limit fuzzing remains challenging because the input space is highly strucsemantic diversity across generated programs. In contrast, FunFuzz tured and semantically rich: deep compiler behavior often requires introduces per-island seed instructions, enabling multiple semantiprograms that combine multiple language features and non-trivial cally distinct starting points that encourage early divergence and interactions. As a result, campaigns can converge to shallow patreduce prompt-induced bias. terns that are easy to generate while missing rare combinations Despite recent progress in LLM-based compiler fuzzing, existing that exercise deeper compilation pipelines [13]. Modern compiler pipelines still rely largely on implicit or coarse-grained feedback to evolution further amplifies the problem because new features and guide search. In Fuzz4All, generation-time decisions are primarily optimizations reshape the reachable behavior space and invalidate tied to validity and execution outcomes [18], while Kitten follows a assumptions embedded in existing strategies [9]. syntax-guided mutation workflow (parse, mutate, compile, check) One recurring challenge lies in the high degree of specializaand uses abnormal compiler behaviors (e.g., crashes and hangs) as tion required to adapt these tools to a specific system under test. bug oracles [19]. Although these signals are effective for filtering Developing a fuzzer for a new compiler or language variant often candidates, they provide limited explicit guidance about how much involves significant manual engineering, from designing intricate new compiler behavior a test actually explores, so search may drift grammar rules to implementing target-specific heuristics [20]. This toward repeatedly producing superficially valid yet structurally specialization reduces portability, as techniques fine-tuned for one similar inputs. To address this limitation, FunFuzz introduces an domain may fail to deliver similar results when applied in a difexplicit fitness objective for candidate ranking. Importantly, this ferent context. As a result, re-targeting to new language variants

arXiv:2605.02789v1 [cs.CR] 4 May 2026

Abstract

Preprint, 2026,

objective is not based on raw coverage alone; it is a richer fitness signal built from results obtained under multiple compilation configurations, so selection reflects more robust behavioral novelty across settings. In parallel, because compiler fuzzing benefits from sustained diversity rather than convergence to a single “best” artifact, FunFuzz combines this fitness-guided selection with a multi-island evolutionary process that maintains partially independent search trajectories. Together, these components provide both direction (via explicit fitness) and diversity (via island-level separation), forming the basis of the framework presented next. Our work. We present FunFuzz, an evolutionary fuzzing framework that integrates LLM-based program synthesis with a multiisland search procedure. FunFuzz starts from an autopromptingstyle distillation step that turns user-provided documentation and examples into a base prompt [18]. Building on these ideas, FunFuzz derives a set of per-island seed instructions at initialization. Each island combines the shared base prompt with a different seed instruction, which creates multiple semantically distinct starting points. During fuzzing, each island maintains an independent population and prioritizes candidates using coverage-guided fitness computed on an instrumented compiler. Islands evolve independently in parallel during local search, and a periodic migration transfers a small fraction of high-value candidates from stronger islands to weaker ones without resetting entire populations. This design aims to sustain exploration under LLM stochasticity: perisland initialization encourages early divergence, local feedback preserves independent search frontiers, and migration mitigates stagnation while avoiding premature homogenization. We evaluate FunFuzz on compiler fuzzing for C and C++, where the inputs are LLM-generated source programs and the goal is to trigger crashes and internal compiler failures in modern compilers such as GCC and Clang. Our evaluation targets modern versions of GCC and Clang under long-running 24-hour campaigns and compares against representative baselines: the LLM-assisted evolutionary fuzzer Fuzz4All [18] and the high-throughput mutational fuzzer Kitten (C only) [19]. On GCC-based targets, FunFuzz reaches up to +27.1% coverage over standard Fuzz4All in C and up to +31.0% in C++. On Clang-based targets, FunFuzz reaches up to +9.3% in C and up to +4.4% in C++. Overall, FunFuzz discovers 119 unique compiler bugs, of which 80 were confirmed by developers. We further report controlled ablations that isolate the contribution of multi-island search, prompt initialization, and scoring components. While FunFuzz is designed to be agnostic to the system under test, our evaluation is limited to C/C++ compilers; validating its effectiveness on other structured-input domains (e.g., interpreters or protocol parsers) remains future work. Contributions. We make the following contributions: • Multi-island LLM-assisted evolutionary fuzzing. We introduce a multi-island fuzzing loop that maintains independent populations with local feedback and uses periodic migration to mitigate stagnation without discarding search context. • Prompt distillation with per-island semantic initialization. We extend autoprompting-style distillation with per-island seed instructions that create multiple semantically distinct starting points while preserving syntactic validity.

Rodríguez Béjar et al.

• Evaluation on modern GCC and Clang. We provide a 24hour evaluation for C/C++ that measures compiler coverage and bug-finding performance against strong baselines, plus ablations that quantify the impact of core design choices.

2

Related Work

Previous work related to FunFuzz spans (i) compiler fuzzers that emphasize high-throughput mutation and curated corpora, (ii) LLMbased approaches that synthesize structured inputs from documentation, and (iii) evolutionary search methods that preserve diversity through multi-population exploration. We summarize these lines of work and then position FunFuzz relative to the closest baselines. Structured-input and compiler fuzzing. Coverage-guided fuzzing is a widely used approach for scalable bug discovery in complex systems, with well-known engines such as American Fuzzy Lop (AFL) [21], AFL++ [4], and libFuzzer [12]. Compiler testing adds an additional challenge: inputs are programs with rich syntactic and semantic constraints, and techniques must trade validity, diversity, and throughput [13]. Generation-based tools such as Csmith [20] provide well-defined randomized programs and have historically exposed many compiler defects, but can saturate over time and require maintenance as languages and compiler behaviors evolve [13]. Mutation-heavy compiler fuzzers focus on throughput and seed exploitation; recent systems such as Kitten [19] represent a strong baseline in this space, using curated corpora with aggressive mutations to maximize coverage over modern toolchains. These approaches scale well, but they often depend on the semantic reach of seeds and mutation operators, which can leave rare feature interactions under-exercised in large and evolving language front-ends. LLM-assisted fuzzing for structured inputs. Recent work shows that LLMs can synthesize structured inputs directly from documentation and examples, reducing the need for hand-written grammars and domain-specific generators [1, 7]. This direction is particularly attractive for language-processing targets, where surface validity alone is insufficient and deeper behavior often requires non-trivial compositions of language features [13]. Fuzz4All [18] is the closest previous system to ours in spirit: it distills documentation into prompts and uses LLM generation inside a feedback-guided loop, which enables cross-domain fuzzing without target-specific grammars. Other systems explore complementary roles for LLMs, including LLM-guided transformations over existing seeds or hybrid pipelines that mix semantic edits with lower-level mutations [22]. For compiler fuzzing specifically, recent work also uses LLMs to improve mutation quality and diversify transformations, which supports the premise that model-driven synthesis can uncover defects that corpus-driven mutation may miss [16]. Across these efforts, a recurring challenge is search collapse: stochastic generation and prompt sensitivity can bias campaigns toward repetitive templates, which reduces exploration unless the design preserves diversity. Evolutionary search and multi-population exploration. Evolutionary fuzzing [2] [11] uses feedback signals (often coverage or reachability) to retain and recombine high-value inputs, but single-population search can converge toward dominant structures when feedback is sparse or noisy. Multi-population “island” models address this issue by maintaining independent trajectories and

FunFuzz : An LLM-Powered Evolutionary Fuzzing Framework

periodically exchanging migrants, improving diversity and helping escape local optima. FunSearch [17] and newer works that build on it [10, 15] provide modern instances of this idea in an LLM setting: multiple islands explore in parallel and exchange high-performing candidates, improving robustness to stochastic generation and local stagnation. While FunSearch targets program synthesis rather than fuzzing, it highlights a practical mechanism that aligns well with LLM-driven input generation, where maintaining multiple distinct trajectories can reduce early collapse to a single coding style. Positioning FunFuzz. FunFuzz sits at the intersection of LLMbased structured input synthesis and multi-population evolutionary exploration. Compared to Fuzz4All, FunFuzz introduces per-island semantic initialization via distinct seed instructions, which creates multiple topic-level starting points. It then maintains these trajectories through a multi-island evolutionary loop with a migration policy that transfers promising candidates without fully resetting search history, which aims to preserve diversity while still exploiting high-value discoveries. Compared to throughputoriented compiler fuzzers such as Kitten, FunFuzz trades raw input volume for higher coverage-per-input and complementary failure discovery under model-synthesized programs. Finally, compared to FunSearch, FunFuzz adapts multi-island evolution to a fuzzing context with compiler-specific signals (instrumented compiler coverage and compilation outcomes) and an oracle-driven bug collection pipeline, which grounds the search in the requirements of real compiler testing.

Preprint, 2026,

assigns fitness from incremental compiler coverage plus failure signals. Parent programs are sampled with a temperature-controlled softmax over fitness scores, and the selected parent is incorporated into the next prompt to steer subsequent generations. Islands exchange candidates only during periodic migration events, which transfer high-value programs and prune low-value ones without resetting entire populations. The following subsections provide a detailed description of the overall FunFuzz approach.

3.1

Prompt Distillation and Initialization

As in Fuzz4All, the pipeline begins by processing user input, which may consist of compiler documentation, protocol specifications, or curated code examples. A Distillation LLM summarizes this content into a small, fixed set of candidate prompts 1 . Each candidate prompt is evaluated by a Generation LLM to synthesize a fixed number of test programs 2 . The validity of each prompt is quantified by computing a compile-validity score, defined as the proportion of generated programs that compile successfully without manual intervention. Compile-validity serves as a lightweight initialization signal; the evolutionary loop later relies on incremental compiler coverage as fitness. The candidate with the highest compile-validity becomes the distilled prompt used in subsequent stages 3 . In addition to the distilled prompt, our system introduces a second prompt layer to promote diversity across islands. Specifically, at step 1 , we perform an additional LLM call to generate two batches of candidate seed instructions, one sampled with low temperature and another with high temperature. 3 FunFuzz APPROACH Each batch contains 𝑁 instructions, where 𝑁 corresponds to FunFuzz is a two-stage fuzzing framework that combines documentationthe number of islands. Each instruction is a short directive that driven LLM input synthesis with a multi-island evolutionary loop. biases program generation toward a specific concept or subsystem The framework is explicitly designed to be agnostic to the system (e.g., allocation patterns, control flow, or error-handling constructs). under test (SUT). To this end, we adopt the prompt distillation Thus, each batch directly defines a full set of candidate island iniworkflow of Fuzz4All [18], which enables the system to ingest tializations. heterogeneous sources of target-facing knowledge (manuals, lanTo determine which batch yields better initialization, each inguage specifications, or code examples) and transform them into struction is combined with the distilled base prompt 4 , producing generation prompts. On top of this SUT-agnostic generation layer, 𝑁 hybrid prompts per batch. These hybrid prompts are passed we adapt the FunSearch multi-island mechanism [17] to fuzzing, to the generation LLM 5 , and a fixed budget of programs is synallowing multiple partially independent evolutionary processes thesized and evaluated for each one. A second round of compileto explore the SUT in parallel. This design mitigates premature validity scoring 6 is then performed. convergence under stochastic LLM generation and prevents early Importantly, selection is performed at the batch level: the scores collapse to a single prompt trajectory, while still enabling inforare aggregated across all hybrid prompts within each batch, and mation sharing across islands. Figure 1 summarizes the overall the higher-scoring batch (low-temperature or high-temperature) pipeline. is selected in its entirety. The selected batch provides one hybrid During the first stage (Prompt Distillation and Initialization), userprompt per island, which becomes the initial prompt for that island. provided artifacts are converted into prompts that provide a high This two-stage process improves initial compilation throughput rate of compiling programs. A distillation LLM proposes candidate while encouraging semantic diversity across islands, reducing the prompts, a generation model samples programs from each candilikelihood of premature convergence to homogeneous program date, and FunFuzz selects a base prompt using compile-validity structures. Further details are provided in Appendix C. as a lightweight initialization signal. To prevent all islands from starting from the same region of the input space, FunFuzz also 3.2 Evolutionary Fuzzing Loop derives island-specific seed instructions that emphasize different topics extracted from the input documentation. Then, the second Inspired by FunSearch, FunFuzz runs multiple evolutionary searches stage (Evolutionary Fuzzing Loop) runs parallel fuzzing loops across in parallel, one per island (Figure 2). Each island operates indepenislands. Each island repeatedly generates programs from its curdently, evolving its own population of candidate programs through rent prompt, compiles them against an instrumented compiler, and feedback and guided mutations. To structure the explanation of this

Preprint, 2026,

Rodríguez Béjar et al.

Figure 1: FunFuzz strategy: High-level view of the evolutionary loop driven by LLM-generated prompts and feedback-guided scoring. The diagram abstracts away the presence of multiple islands to focus on the core dynamics of a single evolutionary cycle. Red labels denote components introduced in FunFuzz beyond Fuzz4All. component, we divide it into Island Loop, Cross-Island Migration Policy and Fitness Computation and Scoring Strategy. For completeness, Appendix D provides a high-level algorithmic summary of the worker-driven fuzzing loop, making explicit how LLM-driven generation, island-local evaluation, migration, and SUTlevel failure detection are orchestrated during execution.

3.2.1 Island Loop. An island in our system represents an independent evolutionary unit that maintains its own population of candidate programs. Each program in the population is assigned a fitness score based on compiler feedback, as described in Section 3.2.3. Figure 1 shows the lifecycle executed by each island. The loop starts with the final island prompt produced in the previous stage, so that it represents the initial input prompt in this stage. The generation LLM samples a batch of candidate programs 7 . Before compilation, FunFuzz applies a deterministic normalization to remove common LLM formatting artifacts and to standardize compilation units. Specifically, the normalization (i) drops malformed or non-resolvable #include, (ii) adds standard-library includes only when the compiler reports missing declarations, using a fixed mapping from diagnostic patterns to headers, and (iii) splits multiprogram outputs into separate compilation units. This normalization does not perform semantic repair; it only makes compilation outcomes comparable across candidates.

FunFuzz compiles each normalized candidate and applies a lightweight failure oracle 8 . The oracle flags abnormal compiler termination using exit status and compiler diagnostics. Successful compilation and standard compilation errors are treated as benign outcomes; other exit codes or internal-failure indicators in stderr (e.g., “internal compiler error”) are treated as potential bugs. The oracle targets crashes and internal errors rather than semantic miscompilations. The island assigns each candidate a fitness score 9 (Section 3.2.3). FunFuzz selects a promising candidate by using a softmax distribution over fitness 10 , with a temperature schedule that shifts from exploration to exploitation over time. This probabilistic mechanism biases sampling toward higher-scoring programs while preserving variability, particularly during early exploration phases. As the temperature decays, sampling gradually shifts toward exploitation, allowing each island to refine its search trajectory. The next prompt incorporates (i) the distilled base prompt, (ii) one selected program, and (iii) a transformation instruction that controls the generation mode (new, mutate, rephrase) 11 . FunFuzz updates the prompt and repeats the loop. To avoid prompt overfitting to a small set of programs, the system penalizes or removes programs after they appear as prompt programs to encourage the system to explore unexplored behaviors. 3.2.2 Cross-Island Migration Policy. To promote diversity and avoid premature convergence, FunFuzz introduces a periodic migration

FunFuzz : An LLM-Powered Evolutionary Fuzzing Framework

Preprint, 2026,

Figure 2: Evolutionary Multi-Island Fuzzing Loop mechanism across islands by extending the design of FunSearch. As depicted in Figure 2, this migration step is triggered every 3 hours 1 . We rank islands by coverage and designate the top 51% (rounded up) as strong; the remaining islands form the weak set. Unlike FunSearch, which resets weak islands, FunFuzz applies a soft recovery step: in each weak island, we remove the lowest-scoring 30% of the population (according to the island’s fitness scores), preserving the remaining individuals and their accumulated search context. Next, each strong island 4 shares 10% of its population. Migrants are sampled from the island’s elite pool (the top 20% by fitness), which biases transfer toward high-value programs while maintaining diversity among the exported set. Migrants are injected into weak islands to replace the pruned individuals. Then, the weak islands are populated with the shared programs 5 , which are reevaluated using the local fitness criteria. Since each island maintains its own scoring state, transferred candidates are re-evaluated using the local coverage and statistics of the receiving island. This discrepancy helps challenge the search boundaries of weaker islands and may uncover new directions previously overlooked. We use a fixed migration schedule and fixed transfer/pruning ratios across all experiments. Unless otherwise stated, migrations occur every 3 hours; weak islands are partially pruned and then replenished with a small sample of high-scoring programs from strong islands. This “soft migration” is intended to mitigate stagnation without discarding island state. These hyperparameters (migration period, pruning ratio, sharing rate, and number of islands) were selected empirically based on preliminary experiments and kept fixed across all evaluations to ensure a fair and controlled comparison. We intentionally do not perform an exhaustive sensitivity analysis, as our primary focus is on evaluating the overall effectiveness

of the proposed approach rather than fine-tuning individual parameters. While this choice simplifies the experimental design, the selected values may not generalize across different targets, workloads, or hardware configurations. A detailed list of all hyperparameters is provided in Appendix F. 3.2.3 Fitness Computation and Scoring Strategy. Our primary fitness signal is marginal compiler coverage: for each generated program, we compile it with an instrumented build of the compilerunder-test and measure the additional compiler source lines exercised by this compilation relative to the island’s current coverage set. This marginal gain serves as a proxy for behavioral novelty and is used to prioritize candidates for retention and reuse. Coverage is collected at the compiler level (the compiler is the SUT), and each island maintains its own local coverage frontier. Using per-island frontiers encourages partially independent exploration: islands can pursue different behavioral regions without being immediately dominated by discoveries made by other islands. Programs that yield positive marginal gains are preferentially retained and sampled as parents, while repeatedly non-contributing programs are deprioritized. We also evaluate optional scoring modifiers (e.g., failure handling, compilation-time rewarding, redundancy filtering) in Appendix E and in the ablation study.

4

Experimental Evaluation

This section evaluates FunFuzz and analyzes the impact of its core design choices. Our experimental campaign addresses four questions: (1) how much compiler coverage FunFuzz achieves relative to state-of-the-art fuzzers, (2) how effectively it discovers actionable compiler failures over long runs, (3) which components drive its performance through controlled ablations, and (4) whether it can be steered toward specific language features without collapsing exploration.

Preprint, 2026,

Hardware and Execution Environment. All fuzzing frameworks are executed inside Docker containers on the same host machine (MSI Vector GP77 13V; Intel i7-13700H; NVIDIA RTX 4060). For LLM-based fuzzers (FunFuzz and Fuzz4All), model inference is served from a dedicated GPU server (NVIDIA A100 80GB), while compilation, coverage measurement, and crash triage are executed locally on the host machine. Kitten is executed entirely on the local host. Input Pre-processing and Normalization. To ensure comparability between FunFuzz and Fuzz4All, we apply the same deterministic normalization step to all LLM-generated outputs prior to compilation (Section 3.2.1). This step only standardizes compilation units (e.g., includes and splitting) and is applied uniformly. Coverage Instrumentation. We measure compiler-level coverage by compiling each generated program with an instrumented build of the target compiler and recording cumulative unique sourceline coverage within the compiler codebase itself, rather than coverage of the generated test programs. All fuzzers are evaluated against the same instrumented compiler binaries and the same coverage collection pipeline. For GCC we use an instrumented build of the gcc-16.0.0 version1 ; for Clang/LLVM we use an instrumented build of a pinned trunk snapshot2 to keep results reproducible. We report cumulative covered source lines aggregated across all compilations in a run. Execution Policy. Each experimental configuration is executed three times with different random seeds. Reported results correspond to averages across these runs, mitigating stochastic effects introduced by evolutionary selection and, where applicable, LLM sampling. LLMs Used. FunFuzz and Fuzz4All use the same generation LLM and decoding parameters during fuzzing (DeepSeek-CoderV2-Lite-Base; temperature 1.0; top-p sampling; max 512 tokens), served via vLLM. Prompt distillation uses GPT-4.1 for instruction construction only (not for program sampling during fuzzing). Experiment Budgets. For the main evaluation, we run long campaigns of 24 hours to assess coverage and bug-finding performance. The 24-hour budget follows standard fuzzing practice to capture steady-state behavior and enable fair comparison [18] [19]. For controlled design analysis and ablation studies, we use a budget of 30,000 generated programs. Compared to the 10,000 programs used in Fuzz4All, this increase reflects the multi-island design of our method, which requires sufficient samples per island to properly exploit semantic diversity. The budget remains within the same order of magnitude as prior work, ensuring comparability while avoiding under-sampling. Fitness function. The fitness function used across all longrunning experiments is described in Section 4.5.1.

4.1

Systems Under Test and Baselines

Our evaluation targets modern, widely deployed compiler toolchains whose complexity and continuous evolution make them representative fuzzing subjects. Systems Under Test. We evaluate GCC and Clang/LLVM, two mature and actively maintained compilers with deep compilation 1 https://github.com/gcc-mirror/gcc commit 94e21bcec061c1691cd0d5382e3c6f420e0386c4 2 LLVM commit 94e21bcec061c1691cd0d5382e3c6f420e0386c4

Rodríguez Béjar et al.

pipelines (front-end parsing/semantic analysis, IR transformations, and back-end code generation). When we report whether a bug is “not fixed in trunk”, we reproduce crashes on a pinned trunk snapshot to keep results reproducible. Baselines. We compare FunFuzz against two state-of-the-art compiler fuzzers: (i) Fuzz4All, an LLM-driven evolutionary fuzzer, which we evaluate in both single-instance and parallel (5 instances) configurations; and (ii) Kitten, a high-throughput mutational fuzzer that utilizes a curated seed corpus (sourced from open-access LLVM test suites). To isolate the impact of this external initialization, we evaluate FunFuzz in two distinct modes: a default from-scratch configuration and a warm-start (or seeded) configuration, which leverages the same publicly available corpus to initialize its evolutionary loop (further explained in Appendix F.2). This allows us to distinguish between gains from FunFuzz’s search strategy and gains from target-specific prior knowledge. Approach Fuzz4All Fuzz4All parallel Kitten FUNFUZZ FUNFUZZ (warm-start)

LLM-guided generation

Evolutionary scoring

Curated corpus required

Domain init. required

✓ ✓ × ✓ ✓

× × heuristic-based ✓ ✓

× × ✓ × ✓

✓ ✓ ✓ ✓ ✓

Table 1: Comparison of design characteristics.

Scope of Comparison. Not all baselines support all targets in our setting. Kitten targets C and does not provide a C++ fuzzing workflow comparable to our evaluation setup; therefore, C++ experiments compare only FunFuzz and Fuzz4All, while C experiments include all three fuzzers. This selection enables a balanced comparison between LLM-driven and high-throughput mutational fuzzing under realistic constraints. We do not include additional generatorbased compiler fuzzers (e.g., Csmith/GrayC for C and YARPGen for C++) in our main comparison; it should be noted that previous work already evaluates Fuzz4All against these baselines in 24-hour campaigns [18], and Kitten motivates high-throughput mutation as a more comparable baseline when assessing LLM-driven compiler fuzzing; we therefore focus on one strong representative per paradigm.

4.2

Fairness and Throughput Considerations

Comparing compiler fuzzers fairly is challenging because LLMdriven fuzzers and high-throughput mutational fuzzers differ substantially in resource constraints and input generation throughput. We therefore report results under two complementary regimes. Throughput reporting and input budgets. Rather than throttling faster systems or inflating slower ones, we evaluate each fuzzer under its intended operating mode and explicitly report both wall-clock time and the number of compiler invocations (postnormalization) executed in each experiment. This makes throughput differences transparent and enables input-budget-controlled comparisons where needed. Interpretation of Baselines. The baselines considered in this work rely on different assumptions regarding input generation and prior knowledge. In particular, Kitten should be interpreted as a strong mutation-based baseline that benefits from a curated,

FunFuzz : An LLM-Powered Evolutionary Fuzzing Framework

compiler-specific seed corpus, rather than as a from-scratch generator. This initialization provides a substantial prior over valid and diverse program structures, which directly contributes to its high throughput and early coverage gains. In contrast, FunFuzz in its main configuration does not rely on curated corpora and instead constructs inputs from scratch using LLM-guided generation and evolutionary selection. For this reason, comparisons should be interpreted not only in terms of raw throughput, but also in terms of coverage efficiency and dependence on target-specific initialization. The warm-start configuration further illustrates that curated corpora and LLM-guided evolutionary search are complementary, rather than mutually exclusive. Cost-Benefit Perspective. The benefit of FunFuzz should not be interpreted solely through raw generation throughput. While LLM-guided fuzzing incurs additional computational cost due to model inference, its objective is not merely to maximize the number of generated programs, but to improve the effectiveness of each evaluated input and reduce dependence on handcrafted, targetspecific initialization. This is why we report both time-based and input-budget-controlled comparisons. Under equalized input budgets, FunFuzz consistently achieves higher compiler coverage than Fuzz4All, indicating better coverage yield per evaluated program. Relative to Kitten, the main advantage of FunFuzz is not raw throughput, but the ability to operate from scratch without requiring a curated compiler-specific corpus, while still remaining competitive in long-running campaigns. Moreover, the warm-start results show that when such prior information is available, FunFuzz can exploit it effectively and achieve substantially higher coverage, suggesting that LLM-guided evolutionary search complements curated initialization rather than replacing it. Configuration Scope. For the experiments, we report (i) a single-LLM configuration aligned with Fuzz4All, and (ii) a higherthroughput configuration that increases generation parallelism to approach the input rate of mutational fuzzers. This separation allows us to distinguish gains due to exploration strategy from gains due to input volume.

4.3

Coverage Evaluation Over 24 Hours

We evaluate the coverage achieved by FunFuzz over 24-hour fuzzing campaigns and compare it against the selected baselines under the configurations described in Section 4.2. All results are averaged over three independent runs. 4.3.1 C Coverage Results. Figure 3 plots 24-hour coverage trajectories for both GCC-16.0.0 and LLVM/Clang (C) infrastructures. Across both compiler environments, FunFuzz overtakes Fuzz4All early in the campaign and continues to increase steadily throughout the run, with no clear saturation trend. Kitten improves quickly at the beginning but plateaus earlier. The warm-start variant (FunFuzz (warm-start)) achieves the highest curve across both targets by leveraging Kitten’s curated corpus. Shaded regions indicate variability across three runs. Furthermore, Table 2 summarizes final coverage across both compilers under two evaluation settings: (i) program-budget-aligned (same number of compiler invocations post-normalization) and (ii) 24-hour time-based comparison. Under the same input budget of 238,222 invocations, FunFuzz (5 islands) reaches 328,188

Preprint, 2026,

(a) GNU Compilers (GCC)

(b) LLVM Compilers (Clang)

Figure 3: Coverage evolution over 24 hours in C. FunFuzz significantly outperforms Fuzz4All and keeps improving beyond Kitten in both environments.

and 632,670 covered lines for GCC and Clang, versus 292,872 and 597,822 for Fuzz4All (+12.1% and +5.8% improvements, respectively). The warm-start variant further elevates these budget-aligned maximums to 350,519 and 641,707 lines (+19.7% and +7.3%). This demonstrates a higher coverage discovery per input independent of throughput. This advantage emerges early, and FunFuzz reaches coverage levels comparable to Fuzz4All’s end-of-run results substantially earlier in the campaign. Over 24 hours, FunFuzz achieves 329,905 (GCC) and 634,091 (Clang) covered lines, exceeding both standard Fuzz4All and Kitten (316,900 in GCC, 620,312 in Clang), despite Kitten executing substantially more inputs (857,567 invocations). This difference is consistent with Kitten operating over a curated, compiler-specific seed corpus, which reduces the need for initial exploration and allows it to focus on exploiting existing program structures. However, this advantage comes with a strong dependence on the availability of high-quality curated seeds, whereas FunFuzz is able to generate programs from scratch. While LLMdriven generation introduces run-to-run variance, FunFuzz consistently dominates across seeds and compilation targets, suggesting that the evolutionary loop is robust to imperfect early generations. More importantly, to establish a fair multi-process baseline, we evaluate a Fuzz4All parallel configuration which launches 5 isolated

Preprint, 2026,

Rodríguez Béjar et al.

instances concurrently to mimic our 5-island hardware budget. Surprisingly, Fuzz4All parallel performs substantially worse than a single standard Fuzz4All instance on both platforms, dropping to 257,593 lines in GCC and 595,161 in Clang. This degradation suggests that simply scaling the number of independent instances does not translate into better exploration; rather, it induces significant redundancy, with many generated test cases effectively exploring the same regions of the input spaces. In contrast, our evolutionary multi-island approach explicitly prevents this issue: through periodic migration and fitness-based pruning, FunFuzz gracefully distributes the search effort and effectively shares independent discoveries without suffering from massive overlap. Consequently, multi-island execution contributes substantially to our approach, with 5 cooperative islands heavily outperforming a single isolated island under the same 24-hour budget (329,905 vs. 303,193 in GCC, and 634,091 vs. 623,560 in Clang). Finally, FunFuzz warm-start significantly boosts performance across both evaluation regimes, achieving the global maximums of 372,275 and 653,402 covered lines for GCC and Clang, respectively.

(a) GNU Compilers (G++)

Coverage Configuration

Progs. Valid %

GCC

Clang

Program-budget-aligned comparison Fuzz4All 238 222 FunFuzz (5 islands) 238 222 FunFuzz (warm-start) 238 222

42.60 292 872 597 822 47.28 328 188 (+12.1%) 632 670 (+5.8%) 43.48 350 519 (+19.7%) 641 707 (+7.3%)

(b) LLVM Compilers (Clang++)

Time-based comparison (24 hours) Fuzz4All 238 222 Fuzz4All parallel 249 456 Kitten 857 567 FunFuzz (1 island) 240 257 FunFuzz (5 islands) 252 283 FunFuzz (warm-start) 479 888

42.60 292 872 597 822 41.60 257 593 595 161 17.16 316 900 620 312 47.90 303 193 (+3.5%) 623 560 (+4.3%) 47.30 329 905 (+12.7%) 634 091 (+6.1%) 43.03 372 275 (+27.1%) 653 402 (+9.3%)

Table 2: Coverage and throughput results for C. Comparisons isolate the effect of program-budget constraints and strictly 24-hour scenarios evaluating both GCC and LLVM infrastructures. Percentage improvements are reported relative to Fuzz4All within the same compilation target.

4.3.2 C++ Coverage Results. Figure 4 plots 24-hour coverage trajectories for both G++ and LLVM/Clang (C++) infrastructures. Across both compiler environments, FunFuzz establishes an early lead over Fuzz4All, and the gap persists throughout the full campaign, indicating sustained gains beyond initial exploration effects (shaded regions denote variability across valid runs). Table 3 reports final coverage under two settings: (i) an input-budget-controlled comparison and (ii) a 24-hour time-based comparison. Under the same budget of 91,561 invocations, FunFuzz (5 islands) reaches 290,199 covered lines versus 254,247 for Fuzz4All (+14.1%) in G++, and 707,638 versus 697,519 (+1.5%) in Clang++, demonstrating higher coverage discovery per input independent of throughput.

Figure 4: Coverage evolution over 24 hours in C++. FunFuzz shows a clear and sustained advantage over Fuzz4All and its parallel counterpart in both environments.

Over 24 hours, FunFuzz achieves 333,254 covered lines in G++ (+31.0%) and 727,964 in Clang++ (+4.4%). Notably, even a single evolutionary island substantially outperforms the standard baseline in both environments (reaching 323,145 covered lines in G++; +27.1% and 722,324 in Clang++; +3.6%), confirming that evolutionary scoring alone provides strong benefits, while multi-island parallelism adds further gains. Furthermore, we evaluate the Fuzz4All parallel multi-process baseline. In contrast to C (where redundant overlap caused negative scaling), launching independent Fuzz4All instances in C++ yields some improvement over a single instance (reaching 293,708 lines in G++ and 706,185 in Clang++). Nonetheless, FunFuzz (5 islands) robustly outperforms this naive parallelization approach across both targets. These improvements are particularly notable for C++, where templates and richer language semantics expose deeper and more structurally complex compilation paths, confirming that FunFuzz’s collaborative search effectively avoids saturation and generalizes well across languages.

4.4

Bug-Finding Effectiveness

While coverage serves as a proxy for exploration, the ultimate goal of a compiler fuzzer is to discover actionable bugs. We therefore evaluate bug-finding effectiveness under long-running campaigns and report results over three independent 24-hour executions per

FunFuzz : An LLM-Powered Evolutionary Fuzzing Framework

Preprint, 2026,

Coverage Configuration

Progs. Valid %

G++

Fuzz4All Kitten

FunFuzz

Clang++

FunFuzz (warmstart)

Program-budget-aligned comparison

GCC-16.0 (C frontend) Fuzz4All 91 561 FunFuzz (5 islands) 91 561

46.42 254 247 697 519 38.20 290 199 (+14.1%) 707 638 (+1.5%)

Time-based comparison (24 hours) Fuzz4All 91 561 195 162 Fuzz4All parallel FunFuzz (1 island) 205 874 FunFuzz (5 islands) 225 854

44.46 254 298 697 519 38.80 293 708 706 185 46.53 323 145 (+27.1%) 722 324 (+3.6%) 44.46 333 254 (+31.0%) 727 964 (+4.4%)

Table 3: Coverage and throughput results for C++. Comparisons evaluate both G++ and Clang++ infrastructures. Percentage improvements are reported relative to standard Fuzz4All within the same target compiler.

configuration. Bugs are counted after post-processing and deduplication, as described below. When we refer to the GCC trunk version, we mean the latest development version corresponding to the current HEAD commit of the official GCC GitHub mirror, specifically commit d78b2b6c01243c59fc52937e1e3b0d84848a8fa9 at the time of our experiments. 4.4.1 Crash Deduplication and Unique Bug Identification. During fuzzing, many inputs may trigger the same underlying defect with minor syntactic variations. To avoid overcounting and enable consistent comparison across fuzzers, we apply a crash deduplication pipeline. Each candidate input is recompiled under a fixed timeout and configuration using the target compiler, and we retain executions that exhibit compiler-internal failures (e.g., ICE-style diagnostics, assertion failures, or crash-like abnormal termination), rather than semantic miscompilations. For each internal failure, we extract (i) the primary diagnostic header describing the failure and (ii) an associated stack trace or backtrace when available. Both are normalized by removing nondeterministic tokens (such as file paths, line numbers, and memory addresses). We report Unique as the number of distinct fingerprints produced by this procedure. 4.4.2 Bug Budget Search on C Compilers. To assess bug-finding capability beyond coverage, we run a bug-budget evaluation on the C frontends of GCC and Clang, using gcc-16.0 and clang-23, respectively. For each configuration, we perform three independent 24-hour runs under identical settings. For GCC (a released version), we additionally report Not fixed in Trunk, i.e., the subset of unique failures that still reproduce on a pinned trunk snapshot used for reproducibility. GCC. Table 4 shows that FunFuzz uncovers substantially more unique GCC failures than both baselines: 14 unique bugs versus 9 for Kitten and 5 for Fuzz4All. The same trend holds when checking persistence on trunk: FunFuzz reproduces 9 unique failures on trunk, compared to 8 for Kitten and 3 for Fuzz4All. The warmstart configuration yields a much larger number of unique and trunk-reproducible failures (34 / 28), illustrating that FunFuzz can

Unique Not fixed in Trunk

5 3

9 8

14 9

34 28

Clang-23 (C frontend) Unique

14

26

24

52

Table 4: Unique compiler-internal failures over three independent 24-hour runs on C compilers. “Not fixed in Trunk” is reported only for GCC.

effectively exploit richer initialization; however, this setting is not baseline-aligned and should be interpreted as a compatibility/upperbound variant rather than a direct comparison. Clang. Table 4 reports the corresponding results for clang-23. Because Clang experiments are conducted on trunk snapshots, we report only the number of unique deduplicated failures. FunFuzz again significantly improves over the LLM baseline (24 unique vs. 14 for Fuzz4All), while Kitten remains highly competitive on this target (26 unique). As for GCC, warm-start configuration substantially increases the number of discovered unique failures (52), suggesting that FunFuzz’s evolutionary loop scales with stronger or more diverse seed corpora, though these results are not baseline-aligned. 4.4.3 Bug Budget Search on C++ Compilers. We next evaluate bug-finding on the C++ frontends of GCC and Clang (g++ and clang++-23) under the same bug-budget protocol as for C, reporting Unique compiler-internal failures after deduplication (Section 4.4.1); for g++ we also report Not fixed in Trunk. Kitten does not provide a C++ workflow comparable to our setup; therefore, C++ bug-finding compares only FunFuzz and Fuzz4All. Fuzz4All

FunFuzz

G++ (C++ frontend) Unique Not fixed in Trunk

12 8

24 19

Clang++-23 (C++ frontend) Unique

15

64

Table 5: Unique compiler-internal failures over three independent 24-hour runs on C++ compilers. “Not fixed in Trunk” is reported only for G++.

G++. Table 5 shows that FunFuzz finds substantially more unique C++ frontend failures than Fuzz4All (24 vs. 12). The same holds for persistence: 19 of these failures still reproduce on our pinned trunk snapshot, compared to 8 for Fuzz4All, indicating that FunFuzz exposes a broader set of distinct and persistent failure modes.

Preprint, 2026,

Rodríguez Béjar et al.

Clang++. For clang++-23 (trunk snapshot), we report only unique deduplicated failures. FunFuzz again significantly outperforms Fuzz4All (64 vs. 15), suggesting that the bug-finding gains generalize to the richer compilation paths exercised by C++. 4.4.4 Bug Diversity and Discovery Dynamics. To characterize bug diversity and discovery dynamics, we analyze (i) the overlap of deduplicated GCC failures across fuzzers and (ii) how new unique failures accumulate over time under a fixed 24-hour budget.

Figure 6: Cumulative unique GCC bugs discovered over time.

across our bug-finding campaigns (aggregated over FunFuzz configurations). For each fingerprint, we attempt minimization and then report it (or match it) in the corresponding upstream tracker. Status Total failures Confirmed Duplicate Pending

Figure 5: Overlap of unique GCC bugs discovered by different fuzzers. Figure 5 reports the set overlap of unique GCC failures (as defined by our crash-fingerprinting pipeline in Section 4.4.1). FunFuzz contributes a substantial fraction of failures that are not observed in the Fuzz4All or Kitten runs under the same budget, indicating that the corresponding inputs exercise failure modes that are not consistently reached by the baselines in our setting. Conversely, the number of failures exclusive to Fuzz4All is comparatively small. The warm-start configuration (FunFuzz (warm-start)) yields the largest set of unique failures overall, suggesting that combining curated seeds with FunFuzz’s evolutionary loop increases the variety of reachable failure fingerprints beyond either component alone (while noting that this configuration is not baseline-aligned). Furthermore, Figure 6 plots the cumulative count of unique GCC failure fingerprints over time, where at each timestamp we take the union of newly observed fingerprints across the three runs. FunFuzz continues to add new unique failures throughout the 24hour campaign, whereas Fuzz4All saturates earlier and Kitten’s discovery rate slows later in the run. The warm-start configuration combines a fast early increase with sustained later discoveries, consistent with the overlap results. 4.4.5 Triage Summary. We additionally track the triage and disclosure status of the failures found by FunFuzz. Table 6 summarizes a manual triage of the deduplicated failure fingerprints collected

GCC 16.0

Clang 23

68 39 5 24

51 41 4 6

Table 6: Manual triage status of FunFuzz’s deduplicated failures aggregated across GCC-16.0 and Clang-23 campaigns. Confirmed: acknowledged by an upstream developer; Duplicate: matches an existing upstream issue; Pending: issue filed or linked but not yet confirmed. Counts reflect the status at the time of writing.

We label a fingerprint as Confirmed if the issue has been acknowledged as a real defect by an upstream developer on the official repository/tracker; as Duplicate if it matches an already-known defect (i.e., an existing upstream issue or an equivalent previously reported failure); and as Pending if we have filed or linked an upstream issue but it has not yet received developer confirmation or a definitive classification. These numbers should be interpreted as a snapshot of an ongoing reporting process rather than an exhaustive accounting. All reports are submitted through the official compiler issue trackers (e.g., LLVM’s llvm-project tracker, GCC bugzilla). For transparency and reproducibility, we provide in our public GitHub repository a complete list of all deduplicated failure fingerprints, together with their minimization status and links to the corresponding upstream issue IDs when available.

4.5

Design Analysis and Ablation Study

We study which components of FunFuzz drive coverage growth and bug-finding behavior by ablating individual design choices while keeping the rest of the pipeline fixed. Unless stated otherwise, we run short 30,000 programs campaigns on GCC (C) and g++ (C++) under the same generation model, decoding parameters, and evaluation harness as in the main experiments, and report compiler

FunFuzz : An LLM-Powered Evolutionary Fuzzing Framework

source-line coverage. These ablations are intended to characterize early search dynamics and motivate our default configuration, not to claim global optimality. While the ablations in this section focus on coverage growth and search dynamics, FunFuzz also supports feature-level steering via targeted prompting. We evaluate this capability separately in Appendix G, where targeted campaigns demonstrate high hit rates for specific C/C++ constructs without collapsing coverage. These results indicate that the evolutionary design analyzed here remains effective even under constrained, feature-biased generation. We organize the analysis around three factors: (i) scoring and selection variants, (ii) diversity mechanisms (seed initialization and cross-island sharing), and (iii) the number of islands under a fixed generation budget. 4.5.1 Scoring Function and Selection Strategy. We isolate the impact of individual scoring/selection components using single-factor ablations. Unless stated otherwise, we run 30,000-program campaigns across GNU (GCC/G++) and LLVM (Clang/Clang++) compilers, keeping the generation model, decoding parameters, and evaluation harness fixed. Table 7 summarizes the variants; detailed definitions of each toggle are provided in Appendix E. Across both infrastructure targets, incorporating feedback mechanisms consistently improves exploration. Incremental coverage scoring (+Score) yields substantial gains over the baseline (e.g., +25,933 lines in G++ and +47,888 in Clang). Adding compilationtime tracking (+Time) and redundancy filtering (+Filter) also strictly improves coverage in all environments by reducing repeated evaluations and favoring more structurally complex compiler paths. Notably, in the more permissive LLVM ecosystem, every isolated modification—including failure rewarding and zerofiltering—injects useful diversity that uniformly and significantly improves coverage (+31k to +50k lines) over the baseline. In sharp contrast, several knobs are counterproductive for GNU front-ends. Explicitly favoring compilation failures (+Fail) increases G++ coverage but severely reduces GCC coverage (-31,902 lines), suggesting failure-centric pressure misguides the search depending on the parser. Similarly, zero-score filtering (+Zero) improves G++ but actively hurts GCC (-11,151), while discarding used parents (+Used) is detrimental to both. Finally, although switching to a global coverage counter (+Global) yields strong early gains across all short ablations, we keep per-island accounting as the default to prevent premature cross-island convergence in 24-hour campaigns. Consequently, our unified configuration only enables globally robust features (Score, Time, and Filter). 4.5.2 Seed Initialization, Sharing, and Diversity Control. Beyond scoring, we ablate structural controls that shape early exploration: seed initialization, cross-island sharing, and parent selection. Unless stated otherwise, ablations evaluate 3-hour campaigns, while crossisland sharing is measured natively over 24 hours. Results for all environments are summarized in Table 8. Seed Initialization. When all islands start from an identical seed, early exploration trajectories exhibit severe redundancies. Initializing each island with a distinct starting seed effectively distributes the search from the beginning, seamlessly improving early coverage across all environments. As shown in Table 8, distinct

Preprint, 2026,

GNU Compilers Variant

G++

Δ

GCC

Δ

Baseline +Fail +Used +Time +Score +Global +Zero +Filter

243974 252051 243035 249220 269907 257523 272498 264239

– +8077 -939 +5246 +25933 +13549 +28524 +20265

237623 205721 221532 245252 243921 252726 226472 246648

– -31902 -16091 +7629 +6298 +15103 -11151 +9025

LLVM Compilers Variant

Clang++

Δ

Clang

Δ

Baseline +Fail +Used +Time +Score +Global +Zero +Filter

653402 687770 685445 684715 680973 686441 684830 686448

– +34368 +32043 +31313 +27571 +33039 +31428 +33046

535340 568795 576580 582621 583228 586267 582396 578910

– +33455 +41240 +47281 +47888 +50927 +47056 +43570

Table 7: Single-factor ablations of the scoring/selection pipeline for GNU and LLVM compilers (3-hour runs). Each row enables exactly one toggle relative to the baseline. Full definitions are in Appendix E.

seeds boost 3-hour coverage by 11.1% in GCC and 7.8% in G++, introducing smaller but steady benefits across the LLVM ecosystem. Cross-Island Sharing Strategy. We evaluate FunFuzz’s soft migration (which exchanges promising candidates globally without wiping the destination population) against a strict FunSearch-style policy that periodically resets entire islands. While full resets can occasionally inject drastic novelty, they discard valuable accumulated genetic context and destabilize the search. Over full 24-hour campaigns (Table 8), maintaining continuous context via soft migration consistently attains higher final coverage than full resets for both C (+4.2% in GCC, +1.3% in Clang) and C++ (+19.8% in G++, +1.1% in Clang++). Coverage-Guided vs Random Selection. Finally, we compare sampling parents proportional to our coverage-based distribution against naive uniform random selection. In all environments, targeted coverage feedback decisively outperforms randomness (e.g., yielding 237,623 vs. 216,168 lines in GCC and 686,531 vs. 679,504 in Clang++; Table 8). This emphasizes that fitness-proportionate selection is necessary to systematically drive the evolutionary loop toward productive compiler paths, rather than stalling out on lowyield or uncompilable mutants. 4.5.3 Effect of the Multi-Island Architecture. FunFuzz uses a multiisland evolutionary architecture to maintain multiple partially independent search processes in parallel. We study how the number of islands affects coverage under a fixed 24-hour wall-clock budget.

Preprint, 2026,

Rodríguez Béjar et al.

GNU Compilers Setting/Variant

LLVM Compilers

GCC (C) G++ (C++) Clang (C) Clang++ (C++) Seed initialization (3h)

Same seed (shared init.) Per-island seed

211 257 234 703

245 023 264 093

583 888 591 338

683 042 686 531

634 091 625 976

727 964 720 183

591 338 569 677

686 531 679 504

Cross-island sharing (24h) Soft migration (no full reset) Full reset (FunSearch-style)

329 905 316 690

334 557 279 186

Parent selection (3h) Coverage-guided Random

237 623 216 168

243 974 238 510

Table 8: Impact of diversity and sharing mechanisms across all compiler targets. Ablations run for 3 hours, except crossisland sharing strategies naturally evaluated over 24 hours.

In our setup, program generation is primarily driven by LLM inference throughput. As the number of islands increases, the overall generation rate quickly reaches a saturation point; therefore, additional islands do not necessarily increase the total number of evaluated programs, but mainly redistribute the available generation/evaluation budget across more parallel evolutionary trajectories (program counts per configuration are reported alongside the results). Figures 7a and 7b show coverage over 24 hours for 1, 5, 10, and 20 islands in C and C++. Across both languages, we observe the same qualitative trend. A single island already achieves strong coverage growth, indicating that scoring and selection alone provide effective evolutionary pressure. Increasing to five islands yields the best overall performance, suggesting that moderate parallelism improves exploration by enabling divergent trajectories while preserving enough evolutionary depth within each island to refine promising directions. In contrast, using 10 or 20 islands consistently underperforms under the same wall-clock budget. With the global input budget largely fixed, splitting it across too many islands reduces the number of updates per island and weakens selection pressure, leading to shallower evolutionary progress. In this regime, additional parallelism fragments the search rather than improving coverage.

5

Conclusions

We presented FunFuzz, an LLM-assisted evolutionary fuzzing framework that combines coverage-guided selection with a multi-island search procedure to sustain exploration on structured inputs. FunFuzz initializes multiple semantically distinct search trajectories and uses lightweight migration to share high-value candidates without resetting populations. Across 24-hour campaigns on modern GCC and Clang for C and C++, FunFuzz consistently improves compiler source-line coverage and discovers more unique compilerinternal failures than representative LLM and non-LLM baselines in our setting. Budget-aligned comparisons indicate that the gains are not explained solely by generating more inputs, but by producing higher-value programs. Ablations further show that coverage-based scoring is the main driver, while per-island initialization and conservative sharing help avoid premature convergence; moderate island

(a) C

(b) C++

Figure 7: Coverage evolution over 24 hours for different island configurations in GCC and G++.

parallelism provides the best trade-off under bounded generation throughput. Future work includes extending the oracle beyond internal failures to semantic bugs (e.g., miscompilations via differential testing), incorporating richer feedback signals (e.g., IR-level coverage or failure features) to refine fitness, and improving efficiency through hybrid mutation/generation and faster evaluation pipelines. We also plan to further systematize triage (minimization, clustering across versions, and mapping to upstream trackers) and to evaluate the approach on additional structured-input targets beyond compilers.

References [1] Yinlin Deng, Chunqiu Steven Xia, Haoran Peng, Chenyuan Yang, and Lingming Zhang. 2023. Large language models are zero-shot fuzzers: Fuzzing deep-learning libraries via large language models. In Proceedings of the 32nd ACM SIGSOFT international symposium on software testing and analysis. 423–435. [2] Martin Eberlein, Yannic Noller, Thomas Vogel, and Lars Grunske. 2020. Evolutionary grammar-based fuzzing. In International Symposium on Search Based Software Engineering. Springer, 105–120. [3] Karine Even-Mendoza, Cristian Cadar, and Alastair F Donaldson. 2022. CsmithEdge: more effective compiler testing by handling undefined behaviour less conservatively. Empirical Software Engineering 27, 6 (2022), 129. [4] Andrea Fioraldi, Dominik Maier, Heiko Eißfeldt, and Marc Heuse. 2020. { AFL++ } : Combining incremental steps of fuzzing research. In 14th USENIX workshop on offensive technologies (WOOT 20). [5] Adrian Herrera, Hendra Gunadi, Shane Magrath, Michael Norrish, Mathias Payer, and Antony L Hosking. 2021. Seed selection for successful fuzzing. In Proceedings of the 30th ACM SIGSOFT international symposium on software testing and analysis.

FunFuzz : An LLM-Powered Evolutionary Fuzzing Framework

230–243. [6] Christian Holler, Kim Herzig, and Andreas Zeller. 2012. Fuzzing with code fragments. In 21st USENIX Security Symposium (USENIX Security 12). 445–458. [7] Linghan Huang, Peizhou Zhao, Huaming Chen, and Lei Ma. 2024. Large language models based fuzzing techniques: A survey. arXiv e-prints (2024), arXiv–2402. [8] Linghan Huang, Peizhou Zhao, Lei Ma, and Huaming Chen. 2025. On the challenges of fuzzing techniques via large language models. In 2025 IEEE International Conference on Software Services Engineering (SSE). IEEE, 162–171. [9] Jaeseong Kwon, Bongjun Jang, Juneyoung Lee, and Kihong Heo. 2025. Optimization-Directed Compiler Fuzzing for Continuous Translation Validation. Proceedings of the ACM on Programming Languages 9, PLDI (2025), 627–650. [10] Robert Tjarko Lange, Yuki Imajuku, and Edoardo Cetin. 2025. Shinkaevolve: Towards open-ended and sample-efficient program evolution. arXiv preprint arXiv:2509.19349 (2025). [11] Yuwei Li, Shouling Ji, Chenyang Lv, Yuan Chen, Jianhai Chen, Qinchen Gu, and Chunming Wu. 2019. V-fuzz: Vulnerability-oriented evolutionary fuzzing. arXiv preprint arXiv:1901.01142 (2019). [12] LLVM Project. [n. d.]. libFuzzer – a library for coverage-guided fuzz testing. https://llvm.org/docs/LibFuzzer.html. Accessed: 2026-02-03. [13] Haoyang Ma. 2023. A survey of modern compiler fuzzing. arXiv preprint arXiv:2306.06884 (2023). [14] Sanoop Mallissery and Yu-Sung Wu. 2023. Demystify the fuzzing methods: A comprehensive survey. Comput. Surveys 56, 3 (2023), 1–38. [15] Alexander Novikov, Ngân Vũ, Marvin Eisenberger, Emilien Dupont, Po-Sen Huang, Adam Zsolt Wagner, Sergey Shirobokov, Borislav Kozlovskii, Francisco JR Ruiz, Abbas Mehrabian, et al. 2025. AlphaEvolve: A coding agent for scientific and algorithmic discovery. arXiv preprint arXiv:2506.13131 (2025). [16] Xianfei Ou, Cong Li, Yanyan Jiang, and Chang Xu. 2024. The mutators reloaded: Fuzzing compilers with large language model generated mutation operators. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 4. 298–312. [17] Bernardino Romera-Paredes, Mohammadamin Barekatain, Alexander Novikov, Matej Balog, M Pawan Kumar, Emilien Dupont, Francisco JR Ruiz, Jordan S Ellenberg, Pengming Wang, Omar Fawzi, et al. 2024. Mathematical discoveries from program search with large language models. Nature 625, 7995 (2024), 468–475. [18] Chunqiu Steven Xia, Matteo Paltenghi, Jia Le Tian, Michael Pradel, and Lingming Zhang. 2024. Fuzz4all: Universal fuzzing with large language models. In Proceedings of the IEEE/ACM 46th International Conference on Software Engineering. 1–13. [19] Yuanmin Xie, Zhenyang Xu, Yongqiang Tian, Min Zhou, Xintong Zhou, and Chengnian Sun. 2025. Kitten: A Simple Yet Effective Baseline for Evaluating LLMBased Compiler Testing Techniques. In Proceedings of the 34th ACM SIGSOFT International Symposium on Software Testing and Analysis. 21–25. [20] Xuejun Yang, Yang Chen, Eric Eide, and John Regehr. 2011. Finding and understanding bugs in C compilers. In Proceedings of the 32nd ACM SIGPLAN conference on Programming language design and implementation. 283–294. [21] Michał Zalewski. 2014. American Fuzzy Lop (AFL). https://lcamtuf.coredump.cx/ afl/. Accessed: 2026-02-03. [22] Hongxiang Zhang, Yuyang Rong, Yifeng He, and Hao Chen. 2024. Llamafuzz: Large language model enhanced greybox fuzzing. arXiv preprint arXiv:2406.07714 (2024). [23] Xiaogang Zhu, Wei Zhou, Qing-Long Han, Wanlun Ma, Sheng Wen, and Yang Xiang. 2025. When software security meets large language models: A survey. IEEE/CAA Journal of Automatica Sinica 12, 2 (2025), 317–334.

A

Open Science

We provide an anonymized artifact repository containing the FunFuzz implementation, configuration files, and scripts used to run the experiments and regenerate the key plots/tables in this paper: https: //anonymous.4open.science/r/paper-repository-6EE4/README.md. The repository documents the required software dependencies and compiler builds, the prompt templates and initialization procedure, and the evaluation pipeline (coverage collection, failure detection, and crash deduplication). To support reproducibility, we also provide the exact experiment configurations used in the paper and a small set of sanity-check commands to validate the setup before running full campaigns. Where disclosure permits, we include aggregated experimental outputs needed to reproduce the reported numbers (e.g., coverage summaries and deduplicated failure fingerprints), and we maintain a mapping from confirmed failures

Preprint, 2026,

to upstream issue IDs. The artifact is available at the anonymous repository linked below; licensing and usage terms are specified in the repository.

B

Ethical Considerations

Our work studies LLM-assisted compiler fuzzing, where the goal is to generate synthetic source programs that exercise diverse compiler behaviors and surface compiler-internal failures (e.g., crashes, ICEs, and assertion failures). This research is intended to improve the reliability and security of widely used language toolchains by helping developers identify and remediate defects earlier. The main stakeholders are compiler developers and maintainers (e.g., GCC/LLVM communities), researchers building testing infrastructure, software developers who rely on correct compilation, and downstream users affected by toolchain robustness. We expect predominantly defensive benefits: improved compiler robustness can reduce miscompilation risk, denial-of-service failures during builds, and latent weaknesses that may propagate through the software supply chain. Our experiments do not involve human subjects, personal data, private code, deception, or interaction with unsuspecting users; all inputs are machine-generated programs evaluated on public compiler codebases. The principal ethical concern is dual use: methods that improve bug-finding effectiveness could also be used to discover flaws for offensive purposes. We mitigate this risk by focusing on failure discovery rather than exploit development, avoiding operational weaponization detail, and following responsible disclosure practices for previously unknown issues via official upstream channels, including minimized reproducers when appropriate and coordinated publication timing when needed. We report aggregated, deduplicated outcomes and reference public tracker entries when already disclosed through upstream processes. Some residual dual-use risk remains unavoidable, but we judge that, under these mitigations and disclosure practices, the societal and security benefits of improving critical compiler infrastructure outweigh the remaining risks and justify publication.

C

Prompt Distillation and Initialization: Illustrative Examples

This appendix provides an illustrative example of the prompt distillation and initialization procedure described in Section 3.1. The goal is to make explicit how raw documentation is transformed into structured prompts that guide the initial stages of the fuzzing process.

C.1

Input Documentation and Initial Distillation

The pipeline starts from raw textual input provided by the user. Depending on the target, this input may consist of compiler documentation, protocol specifications, or curated code excerpts. Figure 8 shows an excerpt of compiler-related documentation used as input in our experiments. This raw input is forwarded to a distillation LLM, which generates a small set of candidate generic prompts summarizing the documentation. Prompt generation follows a temperature-controlled sampling strategy designed to balance determinism and exploration.

Preprint, 2026,

Rodríguez Béjar et al.

Documentation excerpt provided to the LLM C Standard Library headers The interface of C standard library is defined by the following collection of headers. <assert.h> Conditionally compiled macro that compares its argument to zero <complex.h> (since C99) Complex number arithmetic <ctype.h> Functions to determine the type contained in character data <errno.h> Macros reporting error conditions <fenv.h> (since C99) Floating-point environment <float.h> Limits of floating-point types <inttypes.h> (since C99) Format conversion of integer types <iso646.h> (since C95) Alternative operator spellings <limits.h> Ranges of integer types <locale.h> Localization utilities <math.h> Common mathematics functions <setjmp.h> Nonlocal jumps <signal.h> Signal handling <stdalign.h> (since C11) alignas and alignof convenience macros <stdarg.h> Variable arguments <stdatomic.h> (since C11) Atomic operations <stdbit.h> (since C23) Macros to work with the byte and bit representations of types <stdbool.h> (since C99) Macros for boolean type ... <stdckdint.h> (since C23) macros for performing checked integer arithmetic

Figure 8: Example of external documentation injected into the LLM context. Unlike the prompt template, this block represents factual specification text used to guide semantic mutations.

Prompt

compile-validity score

Prompt 0 Prompt 1 Prompt 2 Prompt 3

27 20 30 17

Table 9: compile-validity scores obtained during prompt selection. Each score corresponds to the number of successfully compiling programs generated from a given prompt. The highest-scoring prompt is selected as the distilled base prompt.

C.3

Island-Specific Seed Generation

To encourage early semantic divergence across islands, FunFuzz generates a set of island seed instructions from the same documentation used during the prompt distillation phase. These instructions are short, high-level directives—rather than concrete seed programs—designed to bias the LLM toward distinct program themes and feature combinations, such as control flow patterns, numeric corner cases, error handling, concurrency, or specific library usage. Concretely, we query the instruction-generation model using the raw documentation concatenated with a fixed instruction-generation prompt, illustrated in Figure 10. To balance determinism and exploration, two batches of instructions are generated: • a conservative batch sampled at temperature 𝑇 = 0, • an exploratory batch sampled at temperature 𝑇 = 1.

Specifically, one candidate is generated at temperature 𝑇 = 0 (conservative sampling), while three additional candidates are generated at temperature 𝑇 = 1 (exploratory sampling). These prompts are not yet tied to a specific island or exploration strategy; instead, they act as global semantic summaries of the target specification.

C.2

Validity-Based Prompt Selection

Before entering the evolutionary fuzzing loop, FunFuzz performs a lightweight prompt selection phase to identify a single high-quality generic prompt to be used as the distilled base prompt in subsequent experiments. In this step, the unit of evaluation is the generation prompt (not individual programs). For each candidate generic prompt (Figure 9), we sample a fixed batch of 90 programs using the same generation LLM and decoding settings as in our experiments, and compile each program under the same compiler configuration and timeout. We then compute the prompt’s compile-validity score, defined as the number of generated programs that successfully compile without requiring any manual or automated repair. This metric captures how reliably a prompt steers the model toward syntactically and semantically well-formed programs, independent of downstream evolutionary effects. Table 9, reports the compile-validity results. Prompt 2 achieves the highest value (30/90, 33.3%) and is therefore selected as the shared distilled base prompt used in all subsequent fuzzing stages. The remaining candidates are discarded after this selection step.

Each batch contains 𝑁 instructions, where 𝑁 corresponds to the number of islands. Representative examples of the generated island seed instructions are shown in Figure 11. These instructions are later combined with the distilled base prompt to construct the final hybrid prompts used during fuzzing. Each island is assigned one instruction from the selected batch and evolves programs independently while preserving its assigned semantic focus.

C.4

Hybrid Prompt Construction and Selection

For each instruction batch generated in Section A.3, hybrid prompts are constructed by concatenating the distilled base prompt with the corresponding island instructions. This produces two competing hybrid batches: one derived from the 𝑇 = 0 seeds and another from the 𝑇 = 1 seeds (Figure 12). For each hybrid prompt, a total of 60 programs are synthesized. Programs synthesized from all hybrid prompts generated under the same temperature are evaluated using the validity scoring pipeline defined in Section A.2 The resulting validity scores are aggregated across all islands, producing a single global score per temperature configuration. The global score obtained from the 𝑇 = 0 configuration is compared against the global score from the 𝑇 = 1 configuration. The higher-scoring configuration is selected in its entirety, and its hybrid prompts become the initial generation prompts assigned to the islands.

FunFuzz : An LLM-Powered Evolutionary Fuzzing Framework

Preprint, 2026,

This set-level selection preserves semantic coherence across islands while still allowing controlled stochastic exploration through temperature sampling.

D

– ‘semantic-equiv‘: generate semantically similar variants of known programs. • Lines 5–14: The evolutionary loop runs until the time budget is exhausted. • Line 6: A random island is selected, ensuring balanced usage over time. • Line 7: The function ‘getPromisingExample(islandId)‘ retrieves a high-quality example program from the selected island, based on island-local fitness and novelty criteria. • Line 8: A generation strategy is sampled from ‘genStrats‘, determining the type of transformation applied during generation. • Line 9: New fuzzing inputs are generated by invoking the LLM (‘G‘) with a combination of the distilled prompt, the selected example, and the sampled instruction. Generation is thus implicitly conditioned on the semantic focus of the selected island. • Line 10: Each generated program is executed and evaluated. The function ‘getMetric()‘ applies an island-specific fitness function, incorporating signals such as newly discovered compiler coverage, incremental scoring, and execution feedback. It returns per-input fitness values together with the incremental coverage contribution (‘deltaCoverage‘) of the current batch. • Line 11: The island state is updated by integrating the newly evaluated programs, their fitness scores, and the associated incremental coverage, applying island-local selection and filtering policies. • Line 12: The function ‘checkIslandSharing()‘ verifies whether the predefined sharing interval has elapsed and triggers island-level sharing or migration mechanisms if required. • Line 13: Each batch of generated inputs is also evaluated by an external oracle (‘Oracle()‘), which detects crashes or anomalous behavior in the system under test (SUT). All detected failures are accumulated.

Algorithm summary

Algorithm 1 Worker-driven fuzzing loop 1: function FuzzingLoop(inputPrompt, timeBudget)

Input: inputPrompt, timeBudget Output: bugs 4: genStrats ← [generate-new, mutate-existing, semanticequiv] 5: while timeElapsed < timeBudget do 6: islandId ← random(nIslands) 7: example ← getPromisingExample(islandId) 8: instruction ← sample(genStrats) fuzzingInputs ← G(inputPrompt + example + instruc9: tion) (fitnessInputs, deltaCoverage) ← getMet10: ric(fuzzingInputs, islandId) 11: updateIsland(islandId, fuzzingInputs, fitnessInputs, deltaCoverage) 12: checkIslandSharing(currentTime, lastSharing) 13: bugs ← bugs + Oracle(fuzzingInputs, SUT) 14: end while 15: return bugs 16: end function 2: 3:

To illustrate the core logic of our fuzzing framework, Algorithm 1 presents a high-level description of the main control loop executed by a single worker. A worker is a lightweight orchestration entity responsible for driving the fuzzing process: it samples generation strategies, invokes the LLM to produce candidate inputs, dispatches these inputs to the evaluation pipeline, and updates the evolutionary state accordingly. Importantly, workers do not maintain evolutionary state themselves. The evolutionary state is instead encapsulated within a set of independent islands. Each island maintains its own local state, including its corpus of inputs, coverage information, fitness scores, and historical metadata. During each iteration, the worker selects an island, queries it for a promising example, and later updates the same island based on the evaluation results of newly generated inputs. The fuzzing loop is instantiated independently by multiple workers, all operating on the same distilled inputPrompt but interacting with shared island states. This design enables concurrent exploration while preserving localized evolutionary dynamics within each island. Synchronization mechanisms are employed only when accessing shared island data, ensuring consistency without coupling the logical behavior of different workers. • Line 4: A predefined set of generation strategies (‘genStrats‘) is defined. These strategies encode different modes of generation: – ‘generate-new‘: produce novel code from scratch using only the distilled prompt. – ‘mutate-existing‘: modify previously successful examples.

This high-level loop abstracts many of the architectural components described earlier (e.g., fitness-based selection, migration logic, filtering policies), providing a concise yet expressive view of the system’s behavior.

E

Fitness Variants, Normalization, and Hyperparameters

They include: • Failure Handling: We evaluated whether programs that trigger compilation errors should be retained as examples in prompt construction. Although such programs are not directly executable, they may reflect boundary behaviors or syntax edge cases that contribute to discovering new coverage. Their inclusion thus influences the prompt update strategy and may help guide the LLM toward unexplored program behaviors. • Reusing Code Snippets: Repeatedly using successful programs in prompts may lead to premature convergence and reduced diversity. To mitigate this, we explored two strategies: (i) removing programs from memory once they are used in a prompt, and (ii) applying a score penalty to reduce their

Preprint, 2026,

Rodríguez Béjar et al.

influence while keeping them available, the specified penalty was: original_score new_score = − 1 10 This design choice directly impacts the balance between retaining valuable examples and encouraging the emergence of novel behaviors. • Time-Based Rewarding: Compilation time may be used as a proxy for program complexity. We examined whether programs that take longer to compile should receive a proportional bonus to their fitness score, under the hypothesis that higher complexity may correlate with deeper structural features. This mechanism introduces a trade-off between rewarding structural richness and controlling evaluation overhead. The multiplier of the score is obtained following the formula: 𝑆=

𝑇comp

larger. In contrast to prior descriptions that suggested a continuously increasing factor, our approach relies on discrete exploration stages, each capturing the increasing difficulty of extending coverage as fuzzing progresses. • Coverage Accounting Mode: We explored two scoring modes based on coverage tracking: (i) a global counter shared across all islands, and (ii) an independent counter per island (the default setting). The global mode encourages collaborative exploration by rewarding cross-island discoveries, while the independent mode promotes diversified behavior by allowing each island to evolve its own coverage frontier. • Zero-Contribution Programs: Programs that fail to increase coverage may still provide semantic value or structural variety. We evaluated policies for retaining these noncontributing samples as secondary resources, particularly under low-diversity conditions, where they might help enrich the prompt space and guide future generations. • Redundancy Filtering: To mitigate population stagnation, we implemented optional redundancy filters based on Levenshtein distance and Jaccard similarity. Programs exhibiting excessively high structural or lexical similarity were discarded, while those with moderate resemblance received a proportional penalty in their fitness score. This mechanism encourages diversity and reduces redundant computation.

  × max 8 − 𝑇 prev, 1

𝑇 prev Where: – 𝑇comp denotes the compilation time of the current program. – 𝑇 prev is the mean compilation time of the previously evaluated programs. – The constant value 8 is a predefined hyperparameter chosen by us to stabilise the multiplier across different compilers. Since various compilers exhibit different baseline compilation times, subtracting 𝑇 prev from this constant allows the multiplier to self-adjust depending on whether the compiler is relatively fast or slow. • Incremental Coverage Scaling: To reflect the increasing difficulty of discovering new behaviours as coverage grows, we apply a piecewise scaling scheme based on empirically observed coverage ranges. Let 𝐶 max denote the maximum coverage value observed in historical executions of the fuzzer. This value is not fixed a priori; instead, it is continuously updated across campaigns as higher coverage levels are reached. In our current experimental setting, 𝐶 max was approximately 60,000, based on the highest coverage observed during prior runs. Using this empirical reference, we partition the reachable coverage space into three tiers: Tier 1 = 0.40 𝐶 max,

Tier 2 = 0.60 𝐶 max,

Tier 3 = 0.80 𝐶 max . These thresholds approximate early, intermediate, and late stages of exploration. Each tier is associated with a maximum scaling multiplier: 𝑀1 = 5,

𝑀2 = 35,

𝑀3 = 100.

Given the current coverage value 𝐶, the scaling factor is selected according to the tier in which 𝐶 falls. The rationale is the following: discovering new coverage early in the search is relatively easy, so the reward is modest, whereas finding coverage close to the frontier (i.e., above 0.8 𝐶 max ) is significantly harder, and thus the multiplier is substantially

F Reproducibility and Experimental details F.1 Complete list of hyperparameters This appendix reports the complete set of experimental hyperparameters and implementation constants used across all experiments, summarized in Table 10.

F.2

Warm-start configuration

To assess whether FunFuzz can benefit from curated compilerspecific inputs without abandoning its evolutionary prompting dynamics, we evaluate an additional ”warm-start“ configuration that mixes Kitten’s corpus ( obtained from LLVM official suite) with internally generated candidates produced during the FunFuzz campaign. Seed sources. At each generation cycle, candidate programs are sampled from two sources: • External corpus: programs drawn from LLVM’s corpus for the target compiler. • Internal pool: programs previously generated by FunFuzz and stored in the program database together with their fitness scores. Mixing policy. For each seed selection event, FunFuzz samples from the external corpus with probability 0.5 and from the internal pool with probability 0.5. We use this 50/50 mixing ratio as a simple default to ensure that both sources contribute throughout the run; we do not tune this ratio for performance. Uniform treatment after sampling. After selection, all seeds are treated identically by the evolutionary loop: they enter the same parent-selection and scoring pipeline, are compiled and scored under the local island’s fitness signals, and are subject to the same retention and pruning rules as any other candidate. No additional

FunFuzz : An LLM-Powered Evolutionary Fuzzing Framework

Preprint, 2026,

Stage

Parameter

Value

Prompt distillation

Distillation model

Prompt distillation Prompt distillation and Initiatilization Prompt distillation and Initiatilization Prompt distillation and Initiatilization Prompt distillation and Initiatilization Prompt distillation and Initiatilization Prompt distillation and Initiatilization Prompt distillation and Initiatilization Prompt distillation and Initiatilization

# candidate prompts Init scoring signal

GPT-4.1 (used only for prompt construction; not used for program sampling) 4 (1 sampled at 𝑇 = 0, 3 sampled at 𝑇 = 1) Compile-validity: number of generated programs that compile successfully without manual intervention 90 programs per prompt (validity-based evaluation)

Inference backend

FunFuzz and Fuzz4All use identical decoding parameters for budgetaligned comparisons vLLM API server (batched decoding enabled, bfloat16)

Evolutionary loop Evolutionary loop Evolutionary loop Evolutionary loop Evolutionary loop

fuzzing

Generation strategies

generate-new, mutate-existing, semantic-equiv

fuzzing

Island number

5 (selected empirically; Section 4.5.3)

fuzzing

Coverage accounting

Per-island independent coverage counters (default)

fuzzing

Coverage collection

Customized libFuzzer used for compiler-level coverage tracking

fuzzing

Program Fitness

Function described in Section 4.5.1

Cross-island migration Cross-island migration Cross-island migration Cross-island migration Cross-island migration

Migration period Strong/weak split Weak-island pruning Sharing rate Donor selection pool

Every 3 hours Top 51% islands = strong; bottom 49% = weak Prune bottom 30% of clusters in weak islands Each strong island shares 10% of its population Sample from the top-scoring 20% subset of the strong island

Evaluation protocol

Time budgets

Evaluation protocol

Repetitions

30000 programs (design/ablation); 24 hours (coverage + bug-finding); 10000 programs (Targeted) 3 runs with different random seeds (unless stated otherwise)

# test programs per candidate prompt # test programs per hybrid prompt Generation model Decoding parameters Samples per prompt (𝑛) Baseline alignment

60 programs per hybrid prompt DeepSeek-Coder-V2-Lite-Base Temperature = 1.0; top-𝑝 = 1.0 (no truncation); max tokens = 512; no top-𝑘 or repetition penalties 30 programs generated per prompt call

Table 10: Complete list of default FunFuzz hyperparameters

heuristics are applied to externally sourced seeds beyond their availability as parent candidates. Interpretation. This configuration isolates the impact of targetspecific initialization. By providing FunFuzz with the same privileged information used by Kitten—specifically, curated seeds sourced from open-access LLVM test suites—we can evaluate how our evolutionary loop performs in a warm-start regime. These results serve as an upper-bound reference and demonstrate that FunFuzz is complementary to existing corpora. To maintain a fair comparison with

generative baselines, we distinguish these results from our standard “from-scratch” campaigns.

G

Targeted Fuzzing Evaluation

Following Fuzz4All’s targeted-fuzzing setup (feature-specific documentation and hit-rate reporting) [18], we assess FunFuzz’s ability to bias generation toward selected C/C++ constructs. Beyond general-purpose exploration, we evaluate whether FunFuzz can

Preprint, 2026,

Rodríguez Béjar et al.

Autoprompt variants used in the experiments Prompt used to extract island seed instructions Prompt 0 /* The C Standard Library provides a set of headers that define interfaces for common functionality , including input / output ( < stdio .h >) , memory management and utilities ( < stdlib .h >) , string handling ( < string .h >) , mathematical computations ( < math .h > , < complex .h > , < tgmath .h >) , type and numeric limits ( < limits .h > , < float .h > , < stdint .h >) , type checking and conversions ( < ctype .h > , < inttypes .h >) , localization ( < locale .h >) , time / date utilities ( < time .h >) , error handling ( < errno .h >) , and signals ( < signal .h >) . Additional headers offer features for floating - point control ( < fenv .h >) , variable arguments ( < stdarg .h >) , boolean and atomics support ( < stdbool .h > , < stdatomic .h >) , thread management ( < threads .h >) , wide and multi - byte character utilities ( < wchar .h > , < wctype .h > , < uchar .h >) , checked and bit - level arithmetic ( < stdckdint .h > , < stdbit .h >) , and macros for alignment , noreturn , and alternative operators ( < stdalign .h > , < stdnoreturn .h >, < iso646 .h >) . */ /* Please create a short program which uses new C features in a complex way */ # include < stdlib .h >

Prompt 1 /* The C Standard Library defines the core interfaces and functionality available to C programs via a set of headers . These headers provide facilities for type definitions , mathematical operations , input / output , memory and string management , localization , error handling , variable argument functions , atomic and thread operations , character and wide character processing , and support for modern features like complex numbers and checked arithmetic . Each header focuses on a specific utility ranging from mathematics ( < math .h >) , input / output (< stdio .h >) , and memory management ( < stdlib .h >) , to specialized features introduced in newer C standards such as atomic operations ( < stdatomic .h >) , threads ( < threads .h >) , and more robust integer handling ( < stdint .h > , < inttypes .h >) . */ /* Please create a short program which uses new C features in a complex way */ # include < stdlib .h >

Prompt 2 /* The C Standard Library provides a collection of headers that define interfaces for fundamental data types , operations , and utilities necessary in C programs . These headers offer functionalities for input / output (< stdio .h >) , memory management ( < stdlib .h >) , string manipulation (< string .h >) , mathematics ( < math .h > , < tgmath .h >) , locale and character handling ( < locale .h > , < ctype .h > , < wchar .h > , < wctype .h >) , error reporting ( < errno .h >) , limits for data types ( < float .h >, < limits .h > , < stdint .h >) , variable and atomic operations (< stdarg .h >, < stdatomic .h >) , thread support ( < threads .h >) , time / date utilities ( < time .h >) , and specialized needs such as complex arithmetic ( < complex .h >) , floating - point environment ( < fenv .h >) , boolean type support ( < stdbool .h >) , and more . Headers may be specific to C standard versions ( C95 , C99 , C11 , C23 ) , gradually adding capabilities like UTF -16/32 handling , checked integer arithmetic , and enhanced bitwise operations . */ /* Please create a short program which uses new C features in a complex way */ # include < stdlib .h >

Prompt 3 /* The C Standard Library provides a set of headers that define interfaces for essential functionalities such as input / output ( < stdio .h >) , memory management and utilities ( < stdlib .h >) , string and character handling (< string .h > , < ctype .h > , < wchar .h > , < wctype .h >) , mathematical operations ( < math .h > , < complex .h > , < tgmath .h >) , type and limit definitions ( < limits .h > , < float .h > , < stdint .h > , < inttypes .h >) , error handling ( < errno .h >) , localization ( < locale .h >) , time / date utilities ( < time .h >) , and more . Additional headers support advanced features like threading ( < threads .h >) , atomic operations (< stdatomic .h >) , floating - point environment ( < fenv .h >) , and checked arithmetic (< stdckdint .h >) . */ /* Please create a short program which uses new C features in a complex way */ # include < stdlib .h >

Figure 9: Complete autoprompting variants provided to the LLM. Each prompt injects different documentation context while preserving the same generation instruction.

{DOCUMENTATION} Given the above documentation generate instructions in order to cover different areas of GCC. Example: /* Implement a C program that performs matrix multiplication using dynamic memory allocation */ #include <stdlib.h> #include <stdio.h> Generate N different instructions like the above one.

Figure 10: Prompt used to generate candidate seed instructions for initializing island-specific corpora

Representative island seed instructions /* Implement a C program that counts the number of unique words in a sentence */ #include <stdio.h> #include <string.h> /* Write a C program that prints a localized error message when file open fails */ #include <stdio.h> #include <errno.h> #include <locale.h> /* Read two complex numbers and compute product, sum, and modulus */ #include <stdio.h> #include <complex.h> #include <math.h> /* Create two threads; each prints numbers from 1 to 5 */ #include <stdio.h> #include <threads.h> /* Generate random numbers and measure execution time */ #include <stdio.h> #include <stdlib.h> #include <time.h>

Figure 11: Example batch of island seed instructions generated by the LLM steer generation toward specific language constructs while retaining broad compiler coverage. We run targeted campaigns for C (typedef/union/goto) and C++ (std::apply / std::expected / std::variant) by augmenting the user-provided documentation and instruction prompt with feature-specific guidance; the evolutionary loop (scoring, selection, and island management) is unchanged. Each campaign generates 10,000 programs. We report (i) hit rate, defined as the fraction of generated programs that contain the target construct according to our feature detector, and (ii) compiler source-line coverage accumulated during the campaign.

FunFuzz : An LLM-Powered Evolutionary Fuzzing Framework

Preprint, 2026,

Table 11: GCC targeted campaign

Autoprompting final prompt Generic Prompt (template selected) /* The C Standard Library provides a collection of headers that define interfaces for fundamental data types , operations , and utilities necessary in C programs . These headers offer functionalities for input / output ( < stdio .h >) , memory management ( < stdlib .h >) , string manipulation ( < string .h >) , mathematics ( < math .h >, < tgmath .h >) , locale and character handling ( < locale .h >, < ctype .h > , < wchar .h > , < wctype .h >) , error reporting ( < errno .h >) , limits for data types (< float .h > , < limits .h > , < stdint .h >) , variable and atomic operations ( < stdarg .h > , < stdatomic .h >) , thread support ( < threads .h >) , time / date utilities ( < time .h >) , and specialized needs such as complex arithmetic ( < complex .h >) , floating - point environment ( < fenv .h >) , boolean type support ( < stdbool .h >) , and more . Headers may be specific to C standard versions ( C95 , C99 , C11 , C23 ) , gradually adding capabilities like UTF -16/32 handling , checked integer arithmetic , and enhanced bitwise operations . */

Instruction Selected (semantic variant) /* Read two complex numbers and compute product , sum , and modulus */ # include < stdio .h > # include < complex .h > # include < math .h >

Figure 12: Prompt synthesized during the autoprompting phase for a given island. The upper block is a reusable template inspired by Fuzz4All. The bottom instruction is island dependent, to promote local diversity.

Hit rate

union

typedef

goto

General

union typedef goto

80.66% 5.87% 0.62%

35% 76.59% 6.25%

1.16% 0.20% 70.45%

7.64% 13.6% 1.22%

Coverage

149671

158532

149340

182691

Table 12: G++ targeted campaign Hit rate

apply

expected

variant

General

apply expected variant

67.26% 0.19% 0.52%

0.14% 79.63% 0.46%

0.83% 2.66% 85.73%

0.4% 0.56% 3.32%

Coverage

213342

211625

210512

233846

Tables 11 and 12 show that targeted prompting yields consistently high hit rates for the intended constructs (e.g., 76.6% for typedef, 80.7% for union, and 85.7% for std::variant), whereas general prompting produces substantially lower rates (ranging from 0.4% to 13.6% depending on the construct). Despite the additional constraints introduced by targeted guidance, compiler source-line coverage remains substantial across all campaigns, indicating that FunFuzz does not collapse into repetitive or shallow instances of the target feature. Across all targeted campaigns, FunFuzz triggers 15 deduplicated compiler-internal failures. Overall, these results show that targeted prompting remains effective even under a limited budget of 10,000 generated programs, enabling both construct-specific bias and the discovery of compiler-internal failures.

Record · ID 157322 · SHA-256 5de5981716a24a2a
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.