Labelling Bug-Fixing Commits with Local Open-Weight Language Models Philip König∗1 , Georg Goldenits2 , Caroline König3 , Sebastian Raubitzek2 , Fabian Obermann2 , Dennis Toth2 , David Schmidt1 , Edgar Weippl1 , and Kevin Mallinger2 Faculty of Computer Science, University of Vienna, Vienna, Austria, {philip.koenig,caroline.koenig,d.schmidt,edgar.weippl}@univie.ac.at 2 SBA Research gGmbH, Floragasse 7/5.OG, 1040 Vienna, Austria, {ggoldenits,sraubitzek2,fobermann,dtoth,kmallinger}@sba-research.org 3 Christian Doppler Laboratory AsTra, University of Vienna, Vienna, Austria
arXiv:2609.21616v1 [cs.SE] 18 Sep 2026
1
Abstract Defect prediction depends on knowing which commits fix bugs, yet the labels that encode this are produced by routes that each introduce noise. Reused benchmarks carry documented data-quality problems, issue-tracker links are biased and the underlying reports are frequently mistyped, and matching keywords in commit messages is a coarse heuristic. This paper examines whether commits can be labelled as bug fixes from their content alone, using open-weight language models that run locally and therefore keep the process reproducible, inexpensive at corpus scale, usable on proprietary code, and independent of any issue tracker. Against datasets of manually validated and curated bug fixes spanning Java, Python, and JavaScript, we compare a keyword baseline with a set of open-weight models of varying size, prompting each with the commit message and the code diff. On the manually validated corpus the keyword baseline recovers fewer than half of the fixes, whereas the open-weight models recover the large majority and outperform it repository by repository with statistical significance, and larger models do not consistently outperform smaller ones. We further show that evaluation corpora without negative examples cannot support a precision-aware comparison of such classifiers. We release the labelling pipeline together with a labelled, multi-language corpus produced by the recommended configuration, as a reproducible silverstandard resource for building current, project-specific datasets.
Index Terms— bug-fix identification, defect prediction, data quality, label noise, large language models, mining software repositories, reproducibility
1
Introduction
Defect prediction directs testing and review effort towards the parts of a codebase most likely to contain faults, and has become an established tool in software engineering research and practice [1, 2]. Every such prediction model is trained and evaluated on data that rests on one prior decision, namely which commits in a project’s history fixes bugs. That decision yields the bug-fix label, and it is the first link in a chain the rest of the pipeline inherits. A bugfixing commit is the starting point from which the SZZ (Śliwerski–Zimmermann–Zeller) family of algorithms traces the change that introduced a defect [3,4], and the resulting bug-introducing labels are what train the models. Noise in the bug-fix label, therefore, does not stay local. It ∗
Corresponding author: [email protected]
1
propagates into the inducing labels and into every model built on them. The labels, rather than the models, are the binding constraint on this line of work. In current practice, the bug-fix label is produced by one of three routes, and each introduces noise. The first is to reuse an established benchmark. The widely used collections are old, as the NASA and PROMISE-era data sets date to the 2000s [5], and their quality has been questioned to the point that careful cleaning is required before they are used [6,7]. They also predate much of contemporary development practice, so assumptions calibrated on them need not transfer to current projects. The second route links commits to bug reports in an issue tracker. Raw linking carries two well-documented problems. Issue types are misclassified, as a substantial share of reports filed as bugs do not describe corrective work [8], and the links are incomplete, as many genuine fixes are never connected to any issue at all [9]. The third route matches keywords such as fix or bug in commit messages [3]. It is cheap to apply and therefore common as a first pass, but coarse, because commit messages are short, inconsistent, and frequently silent about whether a change repairs a defect. This paper takes a different route. Instead of choosing among these noisy sources, we label a commit from its own content, reading out the committed code together with the commit message. The instrument is a set of Large Language Models (LLMs) run locally on fixed, openly available weights. This gives the approach three properties the other routes lack. First, the labelling is reproducible and auditable, as the weights are pinned and decoding is deterministic. Re-running the process on the same commits returns the same labels. Second, it is inexpensive at the scale of a full repository history and never sends code outside the organization, so it can be applied to proprietary projects that cannot be exposed to an external service. Third, it reads the commit directly, which leaves it independent of the issue tracker and of the linking step responsible for much of the noise described above. A team can therefore label its own repositories and assemble a current data set from its own history, rather than reusing an aging benchmark or inheriting the biases of issue linking. We do not claim that this produces perfect labels. No labelling method for this task is an oracle, the one we study included, and rather than assume its error away we measure it against validated data. The contribution is comparative and transparent, namely labels that are more accurate than the cheap heuristic, produced reproducibly, and reported together with their measured error, which we treat throughout as a silver standard. The restriction to locally runnable models is deliberate for the same reason. Hosted frontier models are not a baseline here, not because they are weaker at the task, which we do not test, but because a durable reference corpus needs a labelling process that can be regenerated and inspected years later. A hosted model is a moving target that is versioned, deprecated, and updated without notice. Local open-weight models satisfy the reproducibility, cost, and data-governance constraints the use case imposes, whereas hosted models do not, and small open models run locally have recently been shown to be a viable alternative to hosted ones for related classification tasks [10]. We address the following three research questions: RQ1 How closely does keyword-based labeling reproduce manually validated bug-fix labels, across projects and languages? RQ2 Can locally deployable open-weight language models label bug-fixing commits more accurately than the keyword baseline, and how does performance vary across models and prompts? RQ3 How does the composition of the evaluation corpus, in particular the presence or absence of negative examples, affect the comparison of bug-fix classifiers? To answer these questions we evaluate a keyword baseline and a set of open-weight models spanning roughly an order of magnitude in parameter count against data sets of manually validated and curated bug fixes in Java, Python, and JavaScript, prompting each model with 2
the commit message and the code diff. The keyword baseline recovers fewer than half of the validated fixes, which makes apparent how far the cheap default sits from the data it is meant to approximate. The open-weight models recover the large majority of those fixes and clearly outperform the baseline. A blocked, rank-based analysis confirms that this separation holds repository by repository rather than only in aggregate. Classification quality does not increase with parameter count across the range we study, so the task does not require a model at the upper end of what local hardware allows. Finally, the corpora that contain only positive examples rank the models very differently from the corpus that contains negatives, because without negatives a corpus cannot penalize a classifier that labels too many commits as fixes. This indicates that positives-only data cannot support a precision-aware comparison of classifiers for this task. We release the labeling pipeline together with a multi-language corpus of commits labeled by the recommended configuration, spanning a diverse set of open-source projects across the three languages, as a reproducible silver-standard resource for downstream use.
2
Related Work
2.1
Defect prediction and its ground truth
Defect prediction has been studied for decades, and several surveys document its methods, metrics, and predictive performance [1, 2, 11]. Models are built on features drawn from the source code and from the development process, from size and complexity measures to code churn and other change metrics [12,13]. Much of this effort concentrates on the features and the choice of classifier, whereas the labels the models are trained on are typically inherited from a heuristic and taken as given, a data-quality gap that recent reviews have begun to draw attention to [7]. Yet every supervised defect model depends on a prior decision about which commits fix bugs, from which the defect labels are ultimately derived, and the accuracy of the predictor is bounded by the accuracy of that initial labelling. It is this step, rather than the modelling that follows, that the present work addresses.
2.2
Noise and quality in bug-fix and defect data
The labels underlying defect data are known to be noisy, and the problem has been documented from several directions. In a manual study of more than seven thousand issue reports, a large fraction of those filed as bugs were found to describe non-corrective work such as new features or refactorings, which biases any model that treats the developer-assigned issue type as ground truth [8]. Linking commits to issues introduces a second problem, as the links are both biased and incomplete, and many genuine fixes are never connected to an issue at all [9]. The reused benchmarks on which much of the field has relied carry defects of their own. The widely used NASA and PROMISE collections [5] have documented data-quality problems serious enough that cleaned versions had to be produced before they could be used reliably [6, 14]. The downstream effect of such noise has been measured directly, as injecting realistic label noise degrades defect models and the recall of the resulting models is the dimension most affected [15,16]. A systematic review of data-quality issues in fault prediction collects these and related findings and concludes that label noise is a pervasive and under-addressed threat [7]. These results motivate two choices in this work. The first is to measure every method against manually validated labels rather than against another heuristic. The second is to quantify the error of the labels we produce against that reference and to release the resulting corpus as a silver standard rather than presenting it as ground truth.
3
2.3
Identifying bug-fixing commits
In practice, bug-fixing commits are identified by one of a few routes. A common and inexpensive one matches keywords such as fix in the commit message. The approach dates back to early work on identifying fixes [3] and remains in use today [17]. A second route links commits to typed issues in a tracker, inheriting the bias and incompleteness noted above [9]. These labels rarely remain an end in themselves, as they seed the SZZ family of algorithms, which trace each fix back to the change suspected of introducing the defect [3, 4]. This has been reimplemented and refined many times since [18], but its reliability depends heavily on the quality of the fix set it starts from [19]. The precision of the initial fix labelling therefore propagates into every bug-introducing label SZZ produces, which is one reason a labelling method intended to seed SZZ is best operated at a high-precision point. Evaluating fix identification requires datasets of known bugs. Several recent collections provide reproducible, manually checked bugs for individual languages, namely BugsInPy [20] and PyBugHive [21] for Python and BugsJS [22] for JavaScript, but each contains only confirmed bug fixes and no negative examples. The SmartSHARK dataset [23] is an exception, as it provides both bug-fixing and non-bug-fixing commits for a set of Java projects, with issue types and commit-issue links that have been manually validated [24]. This validation is what makes SmartSHARK usable as a reference for precision, although its negative class is derived from issue types rather than from inspecting each commit, a point we return to when interpreting precision in Section 4. More recently, language models have been applied to tasks adjacent to fix labelling. One study uses LLMs to detect tangled commits at the method level, deciding from the commit message and the method-level diff whether a change is bug-related, with strong results from few-shot and chain-of-thought prompting across a mix of proprietary and open models [25]. Another applies a language model to the bug-introducing side, improving SZZ by assessing blame candidates with additional context [26]. Both support the premise of this paper, that a model reading code and messages judges the nature of a change more reliably than a keyword does, but both address a different target, method-level untangling in the first case and the inducing commit in the second, rather than labelling whole commits as fixes. Against this background, the contribution of this paper is a commit-level labelling method that runs locally and reproducibly on open-weight models and a released, multi-language corpus produced with it, addressing a gap that the heuristic, issue-tracker, and recent language-model approaches each leave open.
3
Methodology
This section describes how the labeling method is evaluated. We compare a set of locally deployable LLMs with a keyword baseline that applies term stemming to commit messages, across four datasets of human-verified bug fixes, and assess the differences between them with a rank-based, blocked statistical analysis. Each method receives a commit, its message together with its code diff, and returns a binary label stating whether the commit is a bug fix.
3.1
Labelling pipeline
The pipeline takes a repository and assigns every commit a binary label indicating whether it is a bug fix. It clones the repository and walks the commit history. For each commit it extracts the commit message and the diff, where the diff is limited to at most ten changed files and at most fifty lines per file. A marker is appended whenever a diff is truncated, so the model is aware when it sees a partial change. This bounds the size of each prompt, at the cost of truncating unusually large commits, a limitation discussed in Section 5. The message and the bounded diff are placed into a fixed prompt (Section 3.3) and sent to the model, which returns its decision 4
in a constrained format that is parsed into the binary label. Inference is deterministic, with temperature 0, top-p 0, and top-k 1, so the model returns the same label for the same commit on every run. This is what allows a labeled corpus to be regenerated exactly and audited, rather than only reproduced in distribution. Labels are collected per repository. For the evaluation they are joined with the reference labels of the datasets in Section 3.4, and for the released corpus they are themselves the output.
3.2
Model selection
We evaluate seven locally deployable LLMs together with a stemming keyword baseline. The baseline applies an English-language stemmer to each commit message and labels the commit as a fix when a stemmed token matches one of the trigger terms fix, bug, fixup, fail, or correct, following the keyword-based approach used in prior repository-mining work [17]. The LLMs were selected along two deliberate axes, parameter scale and code specialisation. All models were obtained from the Ollama model library1 and downloaded on June 21st 2026. Scale To assess how bug-fix classification performance varies with model capacity, we fix the parameter range at roughly four points, namely a ∼ 10 B class (codegemma:7b2 , gemma4:12b3 ), a ∼ 20 B class (codestral:22b4 , mistral-small3.2:24b5 ), a ∼ 30 B class (qwen3-coder:30b6 , qwen3.6:35b7 ), and one substantially larger model that remains runnable on local hardware (qwen3.5:122b8 ). Code specialisation Each general-purpose model in our set has a recently released codespecialised counterpart, and vice versa. Because the task operates on commits, that is, on code and code-adjacent text, we pair general and code-specialised models of comparable scale, for example gemma4/codegemma, mistral-small3.2/codestral, and qwen3/qwen3-coder. Local deployability All selected models run locally without a hosted API, and the upper end of the scale axis (qwen3.5:122b) is bounded by what remains feasible on local hardware.
3.3
Prompt design
Each model is evaluated under two fixed prompts, identical across all models, denoted P1 and P2. Holding the prompt constant across models ensures that observed differences are attributable to the models rather than to model-specific prompt tuning. Prompt P1 P1 was obtained through an iterative refinement process, which evaluated candidate prompts, their outputs inspected on a development subset, and the wording adjusted until the classification behaviour was satisfactory. The resulting prompt is shown in Figure 1. Prompt P2 P2 keeps the task of P1 but differs from it in several respects at once. It states the definition of a bug fix more fully, adds an extended list of changes that are not fixes, instructs the model to judge from the diff rather than the wording of the message, and includes six pre-labeled input examples. The six examples are deliberately weighted toward non-fixes, two positive and 1
https://ollama.com/library https://ollama.com/library/codegemma 3 https://ollama.com/library/gemma4 4 https://ollama.com/library/codestral 5 https://ollama.com/library/mistral-small3.2 6 https://ollama.com/library/qwen3-coder 7 https://ollama.com/library/qwen3.6 8 https://ollama.com/library/qwen3.5 2
5
You are a senior software engineer and your job is to review this commit and decide whether it is a bugfix, or not. Bugfixes are commits whose primary purpose is to correct incorrect behavior, for example: fixing a crash, wrong output, logic error, regression, or edge-case failure. The following examples are NOT bugfixes: new features, refactoring, performance optimization, documentation, formatting/style, dependency bumps. If a change does several things, classify by its primary purpose. Here is the commit to judge: The commit modified a total of $totalFiles files and has the following commit message: “$commitMessage”. Here is the git diff output of this commit, to decide if this is a bugfix or not: $fileLines $formatInstruction
Figure 1: The zero-shot prompt P1, giving the model a role, a one-line definition of a bug fix, a short list of change types that are not fixes, and the commit under test. Variables in typewriter font are substituted per commit.
four negative rather than balanced, because the harder and more frequent error for both the baseline and the models is to label a non-fix as a fix when its message contains the word fix. The negative examples therefore cover the cases that most often trigger that error, a version bump, a test-only change, a refactoring, and a build change. Because P2 changes more than one thing relative to P1, the two are not a single-variable manipulation, and we do not attribute any difference between them to a particular change. The comparison serves two purposes instead. It tests whether the ordering of the models is stable across two substantially different prompts, and it provides two operating points, since the additional instruction and examples in P2 move the models toward precision and away from recall. The released corpus (Section 3.4) is labelled with P2, because labels intended to seed SZZ-based analyses benefit from the precision-leaning point, where a commit labelled a fix is more likely to be one and fewer spurious fixes propagate into the downstream inducing labels [19]. We treat this as a property of the intended use, not as a claim that precision is preferable in general. The resulting prompt, with two of its six examples, is shown in Figure 2.
3.4
Datasets
We use two collections of data, one to evaluate the labeling methods and one that the method produces and we release. Evaluation data The methods are evaluated against four datasets of human-verified bug fixes. bugsinpy [20] and pybughive [21] for Python and bugsjs [22] for JavaScript consist only of confirmed bug-fixing commits and contain no negative examples, so on these datasets recall can be measured but precision cannot. smartshark is the exception, as it provides both bug-fixing and non-bug-fixing commits for a set of Java projects [23], and it is therefore the only dataset on which precision and the metrics derived from it are meaningful. Its labels rest on manual validation of issue types and of the links between commits and issues [24]. Of the 39 smartshark projects we use, 34 have manually validated issue types and the remaining five do not, a distinction we draw on in Section 4. Because the positive labels derive from validated bug-typed issues while the negative class is the complement, a commit that repairs a defect without being linked to a bug issue is counted as a negative, so the precision measured against smartshark is a conservative lower bound rather than a true precision. Released corpus Applying the pipeline with the recommended model, mistral-small3.2:24b, under prompt P2 yields the corpus we release. It spans 33 open-source projects across Java, Python, and JavaScript and several application domains, including security, healthcare, data processing, developer tooling, and web development, and the projects range in size from a few 6
You are a senior software engineer reviewing one commit. Decide whether the commit is a bug fix. A bug fix is a commit whose primary purpose is to repair a defect, meaning behaviour that was supposed to work but did not, such as a crash, a wrong result, a logic error, a regression, or an edge-case failure, among other defects. These examples are illustrative and not exhaustive. The change repairs existing intended behaviour rather than adding or changing intended behaviour. The following are NOT bug fixes, even when the message contains the word “fix”: new features, refactoring, performance work, documentation or comment edits, formatting or style, renaming, configuration changes, build or CI changes, dependency updates, version-number bumps, commits that only add or change tests without touching production code, reverts, and merge commits. This list is likewise illustrative and not exhaustive. Judge from what the diff does, not from the wording of the message. The word “fix” is not evidence on its own. When the message and the diff disagree, trust the diff. If a commit does several things, call it a bug fix only when repairing a defect is its primary purpose. Examples follow. Each shows a message, a short diff, and the answer. Message: “Fix crash when config file is missing” Diff: - ConfigLoader.java: +4/-1 - return Files.readAllLines(p); + if (!Files.exists(p)) { + return Collections.emptyList(); + } + return Files.readAllLines(p); Answer: {"bugfix": true} Message: “fix flaky timeout in retry test” Diff: - RetryServiceTest.java: +1/-1 - client.setTimeout(100); + client.setTimeout(2000); Answer: {"bugfix": false} Here is the commit to judge: The commit modified a total of $totalFiles files and has the following commit message: “$commitMessage”. $fileLines $formatInstruction
Figure 2: The few-shot prompt P2, showing the instruction block and two of its six pre-labeled input examples. The remaining four are provided with the artifact. Table 1: Composition of the released silver-standard corpus, summarised by language. Labels are produced by mistral-small3.2:24b under prompt P2; the full per-project list is provided with the artifact. Language
Projects
Commits
Labelled fixes
Java Python JavaScript
22 8 3
302,095 40,065 47,407
69,856 3,880 5,960
Total
33
389,567
79,696
dozen commits to tens of thousands, so that very small projects are represented alongside large ones. Table 1 summarises its composition. The labels carry the measured error of the labeller rather than human verification, which is why we describe the corpus as a silver standard.
3.5
Evaluation and statistical analysis
For each model, prompt, and repository, we report accuracy, precision, recall, and F1-Score. Because only the smartshark corpus contains negatively labelled commits, precision and the metrics derived from it are meaningful only on that corpus (Section 4). We therefore conduct the formal model comparison on smartshark, treating its repositories as the units of evaluation. To compare the eight classifiers across repositories, we follow the established protocol for com7
paring multiple classifiers across multiple datasets, as recommended by [27]. The design is a complete block design, as every model is evaluated on the same set of repositories, so each repository is a block that influences all models simultaneously (some repositories are intrinsically harder than others). This structure, the number of models compared, and the non-normal, bounded nature of per-repository F1-Scores jointly motivate a non-parametric, rank-based, blocked analysis rather than parametric or unpaired alternatives. Omnibus test We first apply the Friedman rank-sum test, the non-parametric analogue of a repeated-measures ANOVA for blocked data. For each repository, it ranks the eight models by F1-Score and tests whether the models’ average ranks differ more than would be expected by chance. Operating within-repository ranks neutralises the confound of differing repository difficulty, and, being rank-based, it makes no normality assumption, both of which are appropriate given that per-repository F1-Score is bounded in [0, 1], is skewed, and includes boundary values. The Friedman test provides a single omnibus decision, stating that pairwise comparisons are warranted only if it rejects the null of equal performance. Post-hoc test When the omnibus test is significant, we apply the Nemenyi post-hoc test to identify which pairs of models differ. Nemenyi is the counterpart to the Friedman test for all-pairwise comparisons, as it derives a single critical difference (CD) in average ranks such that two models differ significantly if and only if their average ranks differ by more than the CD. Crucially, it incorporates the multiple-comparison correction by construction, which a naive battery of pairwise tests (e.g. repeated Wilcoxon signed-rank tests) does not. With eight models, there are 28 pairwise comparisons, and uncorrected testing would substantially inflate the familywise error rate. Reporting a single CD also yields a coherent, transitive-where-possible grouping of models into statistically indistinguishable sets, rather than an unordered collection of pairwise verdicts. We perform this analysis separately for P1 and P2 to assess the stability of conclusions about model ordering across prompts. One caveat of the chosen design is noted in that the analysis weights every repository equally, regardless of its commit count, consistent with the macro-averaging used throughout, but not a commit-volume-weighted statement. For smartshark we report descriptive per-repository statistics in two forms, on the 34 projects whose issue types are manually validated and on the full project set. The validated subset provides the primary precision and recall figures, since precision is only as reliable as the validation behind the negative class, and the full set is reported alongside it to show how much the unvalidated projects change the picture. The omnibus Friedman test and the Nemenyi post-hoc analysis are computed over the full set of projects, as are the per-repository maps.
3.6
Technical Detail
All experiments were run on two rented L40 GPUs available on runpod9 . The statistical analysis was performed in R version 3.6.1.
4
Results and Discussion
We evaluate eight bug-fix commit classifiers, the seven LLMs and a stemming keyword baseline, across the four benchmark data corpora bugsinpy, bugsjs, pybughive, and smartshark. Each LLM is run on two prompts, denoted P1 and P2, which are held constant across all models. This lets us separate model effects from prompt effects and assess the robustness of our conclusions to prompt design. 9
https://runpod.io/
8
Due to the data structure, where only the smartshark dataset contains positive and negative labels, for the other datasets, a false positive is not representable and precision is pinned near one by construction. As a consequence, F1-Score on bugsinpy, bugsjs, and pybughive reduces to a monotone function of recall and cannot distinguish a discriminating classifier from one that labels every commit as a bug fix. We use this property as a lens for assessing benchmark validity and treat smartshark as the only corpus for which the full set of metrics is meaningful.
4.1
Per-corpus performance and the benchmark-validity inversion
Table 2 reports mean F1-Scores per dataset (averaged over the repositories within each corpus) for both prompts. The result is a striking inversion that is robust to prompt choice: the corpus ranking of the classifiers on the three bug-fix-only corpora is close to reversed on smartshark. The two classifiers that label almost every commit a fix, codegemma:7b and codestral:22b, which for that reason we set aside as degenerate, make the mechanism concrete, as they top all three bug-fix-only corpora yet fall to last on smartshark. The effect is general, however, because the bug-fix-only corpora cannot penalise false positives and therefore reward recall regardless of precision. Table 2: Mean F1-Scores per corpus (averaged over repositories) for both prompts. The corpus ranking on the bug-fix-only corpora is close to reversed on smartshark (the only corpus with negative labels). This inversion is robust across prompts. Prompt Corpus
mistral qwen3.6:35b qwen3.5:122b qwen3.coder gemma4 codestral codegemma stemming
P1
bugsinpy bugsjs pybughive smartshark
0.883 0.771 0.886 0.630
0.929 0.873 0.940 0.617
0.950 0.957 0.945 0.602
0.938 0.905 0.925 0.596
0.957 0.945 0.916 0.584
0.967 0.936 0.966 0.503
0.999 0.998 0.995 0.396
0.829 0.792 0.657 0.380
P2
bugsinpy bugsjs pybughive smartshark
0.848 0.735 0.817 0.632
0.837 0.673 0.860 0.619
0.911 0.833 0.915 0.625
0.840 0.697 0.773 0.613
0.861 0.771 0.859 0.618
0.934 0.919 0.925 0.528
0.998 0.988 0.997 0.412
0.829 0.792 0.657 0.385
4.2
Repository-level analysis on smartshark
Because smartshark is the only corpus on which precision, recall, and F1-Score are jointly meaningful, and because its repositories are large (ranging from roughly one hundred to several thousand commits), we treat it as the primary evaluation. Table 3 reports, for both prompts, the mean and standard deviation across these three metrics. As set out in Section 3.4, five of the smartshark projects lack manually validated issue types, so their reference labels are themselves model-predicted rather than checked. We therefore report the descriptive statistics on the 34 validated projects, where precision rests on a validated negative class. Reintroducing the five unvalidated projects lowers every method’s mean F1 and widens its spread, as mistral-small3.2:24b falls from 0.645 to 0.630 under P1 with its standard deviation rising from 0.09 to 0.11, and the baseline falls from 0.399 to 0.380. The degradation is concentrated in those five projects and is largest for commons-rdf, the same project that stands out as the visible anomaly in the per-repository maps below. This is the issue-type mistyping of [8] entering directly through the unvalidated labels rather than through the predictors, and it is why we treat the validated 34 as the reference for precision. The ordering of the classifiers and their margin over the baseline are unchanged, and the omnibus test and the maps below use the full project set. F1-Score Under both prompts the same group of LLMs tops the ranking, with mistral-small3.2:24b highest (0.645 under P1, 0.644 under P2). As the across-repository 9
Table 3: Per-repository F1, precision, and recall on the 34 manually validated smartshark projects (mean ± SD across repositories) for both prompts, sorted by P1 mean F1. P2 shifts the operating point toward precision, and the precision-optimal model is prompt-dependent. Prompt P1 Model
F1
Prompt P2
Prec.
Rec.
F1
Prec.
Rec.
mistral-small3.2:24b 0.645±0.093 0.572±0.107 0.751±0.099 0.644±0.085 0.588±0.105 0.724±0.092 qwen3.6:35b 0.633±0.094 0.545±0.117 0.773±0.088 0.633±0.085 0.634±0.105 0.644±0.106 qwen3.5:122b 0.620±0.093 0.517±0.117 0.799±0.074 0.641±0.084 0.602±0.115 0.703±0.095 qwen3-coder:30b 0.609±0.095 0.493±0.109 0.818±0.073 0.629±0.090 0.618±0.105 0.650±0.105 gemma4:12b 0.600±0.096 0.488±0.119 0.808±0.075 0.636±0.082 0.606±0.109 0.685±0.101 codestral:22b 0.519±0.112 0.374±0.118 0.904±0.062 0.546±0.103 0.405±0.117 0.879±0.061 codegemma:7b 0.410±0.126 0.268±0.115 0.983±0.015 0.429±0.122 0.286±0.115 0.953±0.029 stemming 0.399±0.114 0.436±0.137 0.394±0.143 0.399±0.114 0.436±0.137 0.394±0.143
standard deviation (≈ 0.09) is large relative to the gaps between adjacent models, we test the ordering rather than reading it from the means directly. Treating each smartshark repository as a block, a Friedman rank-sum test over the eight models (P1 with N = 38 repositories, one repository excluded for an incomplete row) strongly rejects the null of equal performance (χ2 (7) = 222.7, p < 2.2 × 10−16 ). We follow it with a Nemenyi all-pairs post-hoc test, whose critical difference in average ranks is CD = 1.703, as shown in Table 4. Table 4: Average Friedman ranks over the smartshark repositories (lower is better) for both prompts. Models whose ranks differ by less than the Nemenyi critical difference (CD) are not significantly different. Under P1 a top pair is separable, while under P2 the five leading models form a single indistinguishable group. P1: N = 38, CD = 1.703. P2: N = 38, CD = 1.681. Model mistral-small3.2:24b qwen3.6:35b qwen3.5:122b qwen3-coder:30b gemma4:12b codestral:22b codegemma:7b stemming
P1 rank
P2 rank
1.58 2.26 3.20 3.76 4.30 6.11 7.37 7.42
2.64 3.15 2.91 3.59 3.27 5.59 7.33 7.51
The post-hoc analysis sharpens the picture beyond what the means alone permit. mistral-small3.2:24b attains the best average rank (1.58) and is statistically indistinguishable from qwen3.6:35b (2.26) and qwen3.5:122b (3.20), even though the latter sits at the margin of the critical difference. These form the top group. At the same time, mistral-small3.2:24b is significantly better than qwen3-coder:30b, gemma4:12b, and all lower-ranked models (p < 0.003 in each case), so it is not merely co-best but separable from the lower half of the language-model field. The keyword baseline and the two degenerate classifiers occupy the bottom three ranks (6.11, 7.37, 7.42), each significantly worse than every model in the top group and mutually indistinguishable. The LLMs thus significantly outperform the baseline.10 10
Nemenyi intervals are not transitive: a model may be indistinguishable from a neighbour yet significantly better than a model the neighbour ties. We therefore report “not significantly distinguishable from” rather than equality throughout. The protocol follows Demšar (2006).
10
The same analysis under P2 (N = 38 repositories) is likewise globally significant (χ2 (7) = 184.86, p < 2.2×10−16 , CD = 1.681), but yields a markedly different separation structure. Under P2 the five leading language models, which we refer to collectively as the leading cluster, are mutually indistinguishable, as their average ranks span only 2.64 to 3.59 (mistral-small3.2:24b, qwen3.5:122b, qwen3.6:35b, gemma4:12b, qwen3-coder:30b), a range below the critical difference, and every pairwise comparison among them is far from significant (p > 0.68). The discriminating power that separated mistral-small3.2:24b from the lower language models under P1 therefore disappears under P2. So the prompt that shifts the operating point toward precision (Section 4.2) also compresses the leading models into a single statistically indistinguishable group. The bottom three ranks are again occupied by codestral:22b (5.59), codegemma:7b (7.33), and stemming (7.51), each significantly worse than all five leaders, so the language models significantly outperform the baseline under both prompts as shown in Table 4. Precision and recall The precision/recall decomposition exposes a clear prompt effect. Under P1, mistral-small3.2:24b attains the highest mean precision (0.572) while conceding little recall. Under P2, precision rises and recall falls across the board, and the precision-optimal model changes to qwen3.6:35b, which reaches the highest mean precision (0.634), ahead of the other leading models, whose mean precision falls between 0.59 and 0.62. The degenerate classifiers remain at the recall-only extreme under both prompts and are not considered further. Per-repository precision varies widely even for the leading models, ranging from roughly 0.34 to 0.83 for mistral-small3.2:24b under P1. This reinforces that the large across-repository variance, rather than the mean ranking alone, should temper any deployment decision. Spatial structure To distinguish repository difficulty from predictor quality, we present two complementary views of the per-repository F1-Scores on smartshark, for each prompt. Figure 3 and Figure 4 show absolute F1-Scores, while Figure 5 and Figure 6 show F1-Scores with each repository’s across-predictor mean removed, isolating the contribution of predictor choice. Two facts emerge cleanly and hold under both prompts. First, absolute performance is governed chiefly by the repository, as the absolute-F1-Score maps are dominated by a horizontal easy-to-hard gradient, with repositories such as directory-fortress-core and opennlp yielding low F1-Scores, while others (e.g. commons-net, ant-ivy) showcase high F1-Scores for every predictor. Second, and more importantly for model selection, the repository-centered maps are almost entirely row-coherent, which means that each predictor sits above or below the repository mean by a roughly constant amount across the corpus. The leading cluster sits above the repository mean on essentially every repository, whereas codegemma:7b and stemming sit below it on essentially every repository. The predictor ranking is therefore not an averaging artifact. Instead, it holds repository by repository under both prompts, which materially strengthens the case for the leading cluster. Departures from this pattern are rare and localised, with the most notable being commons-rdf, where the ordering reshuffles and where the stemming baseline collapses to near-zero F1-Scores (the darkest cell in the absolute maps). Such cases identify repositories whose characteristics interact with a specific predictor and are candidates for qualitative follow-up, but they do not alter the corpus-level ordering.
4.3
Discussion
Prompt P1 Under P1 the classifiers separate most sharply on the recall axis. The leading models combine high recall (roughly 0.75-0.82 on smartshark) with moderate precision, and mistral-small3.2:24b attains both the best mean F1-Score and the best average Friedman rank. The post-hoc analysis shows it to be co-best with qwen3.6:35b, as the two models are within the critical difference, while significantly outranking qwen3-coder:30b, gemma4:12b, and the remaining models. Therefore, for P1 the recommendation is to use either of the top pair 11
smartshark per-repository F1
mistral.small3.2.24b
qwen3.6.35b
qwen3.5.122b
F1
1.00
qwen3.coder.30b
0.75 0.50 0.25
gemma4.12b
0.00
codestral.22b
codegemma.7b
ant-ivy
commons-net
commons-validator
commons-configuration
commons-jcs
commons-vfs
commons-bcel
nutch
manifoldcf
commons-io
lens
commons-digester
commons-math
cayenne
commons-beanutils
calcite
commons-codec
commons-dbcp
knox
commons-lang
gora
commons-jexl
eagle
falcon
giraph
commons-compress
commons-collections
tika
commons-scxml
mahout
parquet-mr
jspwiki
archiva
commons-rdf
commons-imaging
opennlp
deltaspike
directory-fortress-core
stemming
Repository (easy → hard)
Figure 3: Absolute per-repository F1 on smartshark under prompt P1. Rows ordered by mean F1, columns by repository mean F1. mistral-small3.2:24b or qwen3.6:35b, which is statistically separated from the rest of the field, rather than a single uncontested winner. Prompt P2 P2 shifts every model toward precision, as smartshark recall drops, by up to roughly 0.17 among the leading models, and precision rises correspondingly. This shift has a statistical consequence as the Friedman/Nemenyi analysis shows the five leading language models become mutually indistinguishable under P2 (average ranks within 0.95, all pairwise p > 0.68), whereas under P1 a top pair was separable from the field. P2 thus compresses the models as it moves them toward precision. It also moves the precision-optimal choice from mistral-small3.2:24b to qwen3.6:35b, which attains the highest mean precision (0.634) of any model-prompt combination. For an application that weights precision above recall, such as seeding SZZ, P2 with qwen3.6:35b is the precision-maximising configuration, at the cost of the lowest recall among the leading cluster. The released corpus is nonetheless labelled with mistral-small3.2:24b under P2 rather than with the precision-maximiser, since mistral leads on F1 under both prompts and holds the best rank, and under P2 the leading models are statistically indistinguishable, so releasing the strongest all-round model at the precision-leaning operating point costs no measurable precision while gaining robustness. Comparison across prompts one that motivated the study.
Four observations bear on our conclusions. We lead with the
12
smartshark per-repository F1
mistral.small3.2.24b
qwen3.5.122b
qwen3.6.35b
F1
1.00
gemma4.12b 0.75 0.50 0.25
qwen3.coder.30b
0.00
codestral.22b
codegemma.7b
ant-ivy
commons-net
commons-validator
lens
commons-vfs
nutch
commons-jcs
commons-configuration
manifoldcf
commons-bcel
commons-io
commons-math
gora
commons-beanutils
falcon
calcite
commons-digester
commons-dbcp
giraph
commons-lang
knox
systemml
commons-collections
eagle
commons-scxml
cayenne
commons-codec
commons-jexl
mahout
commons-compress
parquet-mr
tika
jspwiki
archiva
commons-imaging
deltaspike
opennlp
commons-rdf
directory-fortress-core
stemming
Repository (easy → hard)
Figure 4: Absolute per-repository F1 on smartshark under prompt P2. The easy-to-hard repository gradient is preserved. First, and most importantly, LLMs deliver a large and statistically robust improvement over the established baseline on a problem where that baseline genuinely struggles. On smartshark, the only corpus that permits a fair, precision-aware comparison, the stemming baseline lands near the bottom of the ranking under both prompts (average rank above 7.4 of 8), and every large language model in the leading group significantly outranks it under both prompts by the Nemenyi test. The margin is not marginal, as the best models exceed the baseline’s F1-Score by roughly two thirds (around 0.64 against 0.40), and they do so while the baseline’s own behaviour confirms the difficulty of the task, as its recall and precision are both low and its performance collapses entirely on some repositories. That keyword-based methods perform this poorly is itself evidence that bug-fix identification is a hard problem. LLMs clearing them so decisively, and doing so consistently repository by repository (the row-coherent structure of the per-repository analysis), is one central practical finding of this work. Recent locally-runnable LLMs are not a marginal refinement of existing heuristics for this task, instead they are a qualitatively better instrument for it. Second, this advantage is not contingent on a fortunate prompt. The LLMs’ dominance over the baseline holds under both P1 and P2, two prompts that differ substantially in their operating point. The usefulness of LLMs for bug-fix identification is therefore a property of the models, not of a single hand-tuned prompt, which is an important robustness claim given how often prompt sensitivity undermines LLM results. Third, within the leading group, the prompt determines how finely the models can be sepa-
13
smartshark: predictor disagreement (repo-centered F1)
mistral.small3.2.24b
qwen3.6.35b
qwen3.5.122b
F1 − repo mean 0.2
qwen3.coder.30b
0.0 -0.2 gemma4.12b -0.4
codestral.22b
codegemma.7b
ant-ivy
commons-net
commons-validator
commons-configuration
commons-jcs
commons-vfs
commons-bcel
nutch
manifoldcf
commons-io
lens
commons-digester
commons-math
cayenne
commons-beanutils
calcite
commons-codec
commons-dbcp
knox
commons-lang
gora
commons-jexl
eagle
falcon
giraph
commons-compress
commons-collections
tika
commons-scxml
mahout
parquet-mr
jspwiki
archiva
commons-rdf
commons-imaging
opennlp
deltaspike
directory-fortress-core
stemming
Repository (easy → hard)
Figure 5: Predictor disagreement on smartshark under prompt P1: F1 with each repository’s mean subtracted. The row-coherent structure indicates a repository-independent predictor ranking. rated. Under P1 the Friedman/Nemenyi analysis isolates a top pair (mistral-small3.2:24b, qwen3.6:35b) significantly ahead of the other LLMs, whereas under P2 all five leaders fall within the critical difference and are statistically indistinguishable. The practical reading is reassuring rather than troubling, since model choice among the strong local LLMs matters only at the margin, so a practitioner can choose among a set of models suitable for deployment. Fourth, the prompt shifts the operating point, and the useful point depends on the application. The consistent move from P1 to P2 trades recall for precision, but because the two prompts differ in several respects at once we read this as a shift between two operating points rather than a single controlled lever, and we attribute it to no one change. Which point is preferable is not a property of the task but of the downstream use. Labels that seed SZZ benefit from the precision-leaning point, since a spurious fix there propagates into the bug-introducing labels, whereas a use that must recover as many true fixes as possible, and can tolerate false positives, is better served by the recall-leaning point. For this reason the corpus we release (Section 3.4) is labeled with P2, a choice suited to seeding SZZ rather than a claim that precision is preferable in general. That the operating point can be moved by prompt wording alone, without retraining, is itself a practical convenience.
14
smartshark: predictor disagreement (repo-centered F1)
mistral.small3.2.24b
qwen3.5.122b
qwen3.6.35b
F1 − repo mean 0.2
gemma4.12b
0.0
-0.2
qwen3.coder.30b
codestral.22b
codegemma.7b
ant-ivy
commons-net
commons-validator
lens
commons-vfs
nutch
commons-jcs
commons-configuration
manifoldcf
commons-bcel
commons-io
commons-math
gora
commons-beanutils
falcon
calcite
commons-digester
commons-dbcp
giraph
commons-lang
knox
systemml
commons-collections
eagle
commons-scxml
cayenne
commons-codec
commons-jexl
mahout
commons-compress
parquet-mr
tika
jspwiki
archiva
commons-imaging
deltaspike
opennlp
commons-rdf
directory-fortress-core
stemming
Repository (easy → hard)
Figure 6: Predictor disagreement on smartshark under prompt P2. The row-coherent structure is preserved, confirming ranking stability across prompts.
5
Threats to Validity and Limitations
Construct validity The smartshark reference defines a bug fix by a validated link to a bug-typed issue, an issue-level judgment rather than a code-level one. A commit that repairs a defect without ever being linked to such an issue is counted a non-fix, so a correct positive from a model is scored a false positive. Manual validation of the issue types addresses the mistyping of [8], but not the linking incompleteness of [9]. The precision we report on smartshark is therefore a conservative lower bound, and the true value is at least as high. For the same reason precision is a Java-only statement, as smartshark is the only corpus with a negative class. Internal validity P1 and P2 differ in several respects at once, so we make no claim about which change moves the operating point and read the pair as a robustness check across two prompts. The few-shot examples in P2 are hand-written and drawn from no project in any corpus, so they cannot leak into the evaluation. The keyword baseline is measured on mature, well-documented projects whose commit messages are comparatively disciplined, which flatters it, so the gap we report is if anything an underestimate of its shortfall on noisier histories. Conclusion validity We base ordering claims on a Friedman test with Nemenyi post-hoc rather than on raw mean F1-Scores, which controls for the large across-repository variance, and we run it for both prompts. The test weights every repository equally regardless of its commit 15
count, consistent with the macro-averaging used throughout, so the rankings are not a commitvolume-weighted statement. The across-repository variance is large relative to the gaps between adjacent models, so the rankings should be read together with it. The three bug-fix-only corpora cannot represent a false positive and are used only to establish the benchmark-validity inversion, not to rank models. External validity Primary numbers are reported on the 34 smartshark projects with validated issue types. The full set is shown as a contrast and shifts no conclusion. The diff is truncated to ten files and fifty lines per file, so on unusually large commits the model sees a partial change, though such commits are rare and are seldom fixes. The labels we release are produced by a model and carry its measured error, which is why we present the corpus as a silver standard rather than as ground truth.
6
Conclusions and Outlook
Defect prediction inherits its labels from a prior decision about which commits fix bugs, and the routine ways of making that decision are noisy. We asked whether a commit can be labelled from its own content by a locally deployable open-weight language model, reproducibly and without an issue tracker, and how well the keyword heuristic it would replace actually performs. On the manually validated corpus, keyword matching recovers fewer than half of the bugfixing commits, so the default much of the field relies on is a poor approximation of the labels it stands in for, which answers RQ1. Locally deployable open-weight models recover the large majority of those fixes and outperform the baseline by roughly two thirds in F1, repository by repository and with statistical significance, while the largest model is not the best, so the task is within reach of a model that runs on modest local hardware, and the two prompts we study differ in where they place the operating point rather than in whether the models beat the baseline, which answers RQ2. Finally, the ranking of classifiers on corpora that contain only positive examples is close to reversed on the one corpus that also contains negatives, because an evaluation without negatives cannot penalise a classifier that labels too much as a fix, so such corpora cannot support a precision-aware comparison, which answers RQ3. We make the following contributions. • A measurement, across projects and three languages, of how far keyword-based labeling sits from validated bug-fix labels, quantifying the gap the cheap default leaves open. • An evaluation of seven locally deployable open-weight models against that baseline, with a recommended configuration and the finding that classification quality does not track parameter count across the range studied. • Evidence that evaluation corpora without negative examples cannot rank bug-fix classifiers in a precision-aware way, which bears on how such methods are validated. • A released, reproducible labeling pipeline and a labeled corpus of roughly 390,000 commits across three languages, produced with the recommended configuration and offered as a silver-standard resource for downstream use. Much of defect-prediction research refines models and features on top of labels it takes as given, yet those labels are produced by the step studied here, and the usual shortcut for that step recovers fewer than half of the fixes it is meant to capture. A locally run open-weight model closes most of that gap while keeping the process reproducible and the code in-house, placing a more reliable label at the point where the errors would otherwise originate. Because these labels are the seed for tracing bug-introducing changes and the training signal for the models built afterwards, a better label at the source is inherited by everything downstream. How large that 16
inherited improvement is, and where a hosted frontier model would repay the reproducibility it costs, are the natural next questions.
Acknowledgements The financial support by the Austrian Federal Ministry of Economy, Energy and Tourism, the National Foundation for Research, Technology and Development and the Christian Doppler Research Association is gratefully acknowledged. SBA Research (SBA-K1 NGC) is a COMET Center within the COMET Competence Centers for Excellent Technologies Programme and funded by BMIMI, BMWET, and the federal state of Vienna. The COMET Programme is managed by FFG. This work was funded by the Austrian Research Promotion Agency (FFG) through the BRIDGE programme, project I-SEE, FFG project no. 933312.
Data Availability The labeling pipeline, prompt templates, evaluation runs, released corpus, and statisticalanalysis scripts are available at https://github.com/Raubkatz/BugFixCommitLabelling.
References [1] S. S. Rathore and S. Kumar, “A study on software fault prediction techniques,” Artificial Intelligence Review, vol. 51, pp. 255–327, 2019. [2] Z. Li, X.-Y. Jing, and X. Zhu, “Progress on approaches to software defect prediction,” IET Software, vol. 12, no. 3, pp. 161–175, 2018. [3] J. Śliwerski, T. Zimmermann, and A. Zeller, “When do changes induce fixes?” in Proceedings of the 2005 International Workshop on Mining Software Repositories, ser. MSR ’05. Association for Computing Machinery, 2005, pp. 1–5. [4] S. Kim, T. Zimmermann, K. Pan, and E. J. Jr. Whitehead, “Automatic identification of bug-introducing changes,” in 21st IEEE/ACM International Conference on Automated Software Engineering (ASE’06), 2006, pp. 81–90. [5] T. Menzies, J. Greenwald, and A. Frank, “Data mining static code attributes to learn defect predictors,” Software Engineering, IEEE Transactions on, vol. 33, pp. 2–13, 02 2007. [6] M. Shepperd, Q. Song, Z. Sun, and C. Mair, “Data quality: Some comments on the nasa software defect datasets,” Software Engineering, IEEE Transactions on, vol. 39, pp. 1208– 1215, 09 2013. [7] K. Bhandari, K. Kumar, and A. L. Sangal, “Data quality issues in software fault prediction: a systematic literature review,” Artif. Intell. Rev., vol. 56, no. 8, pp. 7839–7908, Dec. 2022. [8] K. Herzig, S. Just, and A. Zeller, “It’s not a bug, it’s a feature: How misclassification impacts bug prediction,” in 2013 35th International Conference on Software Engineering (ICSE), 2013, pp. 392–401. [9] C. Bird, A. Bachmann, E. Aune, J. Duffy, A. Bernstein, V. Filkov, and P. Devanbu, “Fair and balanced? bias in bug-fix datasets,” in Proceedings of the 7th Joint Meeting of the European Software Engineering Conference and the ACM SIGSOFT Symposium on the Foundations of Software Engineering (ESEC/FSE), 2009, pp. 121–130.
17
[10] G. Goldenits, P. König, S. Raubitzek, and A. Ekelhart, “Small language models for phishing website detection: Cost, performance, and privacy trade-offs,” Journal of Cybersecurity and Privacy, vol. 6, no. 2, p. 48, 2026. [11] M. D’Ambros, M. Lanza, and R. Robbes, “Evaluating defect prediction approaches: A benchmark and an extensive comparison,” Empirical Software Engineering - ESE, vol. 17, pp. 1–47, 08 2012. [12] N. Nagappan and T. Ball, “Use of relative code churn measures to predict system defect density,” in Proceedings of the 27th International Conference on Software Engineering, ser. ICSE ’05. Association for Computing Machinery, 2005, pp. 284–292. [13] F. Rahman and P. Devanbu, “How, and why, process metrics are better,” in 2013 35th International Conference on Software Engineering (ICSE), 2013, pp. 432–441. [14] M. Shepperd, Q. Song, Z. Sun, and C. Mair, “NASA MDP Software Defects Data Sets,” 03 2018. [15] C. Tantithamthavorn, S. McIntosh, A. E. Hassan, A. Ihara, and K. Matsumoto, “The impact of mislabelling on the performance and interpretation of defect prediction models,” in 2015 IEEE/ACM 37th IEEE International Conference on Software Engineering (ICSE), vol. 1, 2015, pp. 812–823. [16] S. Kim, H. Zhang, R. Wu, and L. Gong, “Dealing with noise in defect prediction,” in 2011 33rd International Conference on Software Engineering (ICSE), 2011, pp. 481–490. [17] P. König, S. Raubitzek, A. Schatten, D. Toth, F. Obermann, C. König, and K. Mallinger, “Boost-classifier-driven fault prediction across heterogeneous open-source repositories,” Big Data and Cognitive Computing, vol. 9, no. 7, 2025. [18] M. Borg, O. Svensson, K. Berg, and D. Hansson, “SZZ unleashed: An open implementation of the SZZ algorithm - featuring example usage in a study of just-in-time bug prediction for the jenkins project,” CoRR, vol. abs/1903.01742, 2019. [Online]. Available: http://arxiv.org/abs/1903.01742 [19] G. Rodríguez-Pérez, G. Robles, and J. M. González-Barahona, “Reproducibility and credibility in empirical software engineering: A case study based on a systematic literature review of the use of the szz algorithm,” Information and Software Technology, vol. 99, pp. 164–176, 2018. [20] R. Widyasari, S. Q. Sim, C. Lok, H. Qi, J. Phan, Q. Tay, C. Tan, F. Wee, J. E. Tan, Y. Yieh, B. Goh, F. Thung, H. J. Kang, T. Hoang, D. Lo, and E. L. Ouh, “BugsInPy: A database of existing bugs in Python programs to enable controlled testing and debugging studies,” in Proceedings of the 28th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE), 2020, pp. 1556– 1560. [21] G. Antal, N. Vándor, I. Kolláth, B. Mosolygó, P. Hegedűs, and R. Ferenc, “PyBugHive: A comprehensive database of manually validated, reproducible Python bugs,” IEEE Access, vol. 12, 2024. [22] P. Gyimesi, B. Vancsics, A. Stocco, D. Mazinanian, Á. Beszédes, R. Ferenc, and A. Mesbah, “BugsJS: A benchmark of JavaScript bugs,” in 2019 12th IEEE Conference on Software Testing, Validation and Verification (ICST), 2019, pp. 90–101.
18
[23] A. Trautsch, F. Trautsch, S. Herbold, B. Ledel, and J. Grabowski, “The SmartSHARK ecosystem for software repository mining,” in Proceedings of the ACM/IEEE 42nd International Conference on Software Engineering: Companion Proceedings (ICSE-Companion), 2020, pp. 25–28. [24] S. Herbold, A. Trautsch, and F. Trautsch, “On the feasibility of automated prediction of bug and non-bug issues,” Empirical Software Engineering, vol. 25, 2020. [25] M. N. I. Opu, S. Wang, and S. Chowdhury, “LLM-based detection of tangled code changes for higher-quality method-level bug datasets,” 2025. [Online]. Available: https://arxiv.org/abs/2505.08263 [26] L. Tang, J. Liu, Z. Liu, X. Yang, and L. Bao, “LLM4SZZ: Enhancing SZZ algorithm with context-enhanced assessment on large language models,” Proceedings of the ACM on Software Engineering, vol. 2, no. ISSTA, pp. 343–365, 2025. [27] J. Demšar, “Statistical comparisons of classifiers over multiple data sets,” J. Mach. Learn. Res., vol. 7, p. 1–30, Dec. 2006.
19