arXiv:2607.10420v1 [cs.SE] 11 Jul 2026
Is Model Instability just Noise to be Tolerated or a Property that can be Managed? Amirali Rayegan
Lunxiao Li
Tim Menzies
North Carolina State University Raleigh, North Carolina, USA [email protected]
North Carolina State University Raleigh, North Carolina, USA [email protected]
North Carolina State University Raleigh, North Carolina, USA [email protected]
Abstract—In software analytics, rerunning the same analysis twice often yields different models and conclusions. This reduces trust in the model and limits its use. We find that model instability is a major problem. Across 127 multi-objective SE optimization problems (12,700 test cases), repeated runs of a state-of-theart optimizer agree on only 13.7% of test cases, even under improved settings. We argue that this instability is not merely noise to tolerate, but a property that can be measured and managed. By adjusting how labels are spent, how complex the models become, and how splits are scored, we obtain models that agree 4.8 times as often as the default configuration. The standard deviation of optimization error falls by 22% on average (mean std 17.4 to 13.6), while recommendation quality improves rather than degrades. In terms of quality, the refined settings are statistically top-ranked on 119 of 127 datasets, compared to 74 for the defaults. We then test causal and data-locality interventions and find that they help only partially, suggesting a residual stability floor. Our evidence suggests there are fundamental limits to stability set by the data itself (noise, scarce labels, proxy objectives, and the many near-equivalent models a dataset admits). We conclude that instability should be treated as a standard evaluation axis in SE optimization, which should be routinely measured, reported alongside performance, and used to calibrate trust in any single run. The methods in this paper provide a baseline against which future efforts to reduce SBSE instability can be judged. To support open science, we offer the following reproduction package: https://tinyurl.com/Model-Instability Index Terms—Software Analytics, Rashomon Effect, SearchBased Software Engineering, Optimization, Causal Reasoning
I. I NTRODUCTION Software analytics is an influential research area with much practical value [1]–[6]. By mining historical project data, teams allocate testing effort, prioritize high-risk modules, and cut maintenance costs. Eventually, though, stakeholders ask a deceptively simple question: “What did you learn from all that data?” The answer is usually a symbolic explanation, a decision tree, a rule set, or a causal graph. Unfortunately, these explanations are often unstable. Rerunning the same analysis yields different models and, therefore, different conclusions. This reduces trust in the model and, therefore, limits its use. For example, Figure 2 shows model variability seen after running the same optimizer several times on the same data using different random number seeds. As seen in Figure 2, the same learner can produce models with different structures and use different attributes. The problem of model instability is
not just a quirk of the learner used in Figure 2. Rather, it is endemic. As SE problems grow larger and larger, state-of-the-art methods inject randomness to scale, and conclusion instability has been reported across regression [7], text mining [8], defect prediction [9], causal discovery [10], and LLMs [11]. A natural reaction is to treat instability as a bug to fix with more data, tuning, or ensembling. But the literature suggests a harder truth. Often, there is no single “best” explanation to recover. Breiman’s “two cultures” essay notes that structurally different models can fit the same data nearly as well [13], the Rashomon effect, which Xin et al. formalize as the (often enormous) Rashomon set of near-optimal models for a task [14]. When many models are nearly tied, small perturbations steer learning toward different structures. Structural instability is thus not a defect to eliminate but a property of the data. This reframes what “stability” should mean. We distinguish performance instability (variation in the outcomes of recommended solutions, see Figure 1) from structural instability (variation in model form and features used, see Figure 2). Expecting the same model every time is the wrong goal. What matters is whether the analysis yields consistent recommendations. We therefore assert that performance instability is the pressing problem we target. Based on that, this paper asks four research questions. 1) RQ0 – The Problem: How prevalent is performance instability in SE optimization, and what are its consequences for practitioners? 2) RQ1 – The Nature: How prevalent is structural instability, and is it the same as performance instability? 3) RQ2 – The Fixes: Which factors drive performance instability, and what configuration changes reduce it without sacrificing optimization quality? 4) RQ3 – The Limits: Is the instability remaining after RQ2 the fault of the learner, or of the data? This paper contributes (1) a large-scale study of 127 SE optimization problems showing pervasive instability (repeated runs agree on only 13.7% of test cases under refined settings); (2) evidence that structural and performance instability are distinct. Trees can be structurally diverse yet yield consistent suggestions, making performance stability the only actionable criterion; (3) a systematic analysis identifying labeling budget, sampling strategy, model complexity, and splitting criterion
Fig. 1: Performance instability across 127 SE datasets. Standard deviation of prediction errors as defined by Equation 4 (difference between the referenced optimal and the outcome of the model). Blue and red lines show, respectively, results obtained using (1) the initial default settings or (2) the settings recommended by this paper. Note that, with our methods (the red line), performance instability is nearly always reduced, sometimes by very large amounts (e.g., see left-hand side). if | | | | | | if | | | |
TIME <= 4 i f PCON > 3 ; i f PCON <= 3 | i f ARCH <= 3 | | i f LTEx > 2 ; | | i f LTEx <= 2 ; | i f ARCH > 3 ; TIME > 4 i f SCED > 3 ; i f SCED <= 3 | i f PMAT > 3 ; | i f PMAT <= 3 ;
> > > > > > > > > > > > > >
if | | | | | | | | if | |
CPLx <= 4 i f PREC <= 5 | i f FLEx > 5 ; | i f FLEx <= 5 | | i f PCON <= 2 | | | i f SITE > 2 ; | | | i f SITE <= 2 ; | | i f PCON > 2 ; i f PREC > 5 ; CPLx > 4 i f PMAT <= 1 ; i f PMAT > 1 ;
> > > > > > > > > > > > > >
if | | | | | | | | | | | | if
PCAP > 1 i f DOCU <= 2 | i f STOR > 5 ; | i f STOR <= 5 | | i f DOCU > 1 | | | i f PREC > 3 ; | | | i f PREC <= 3 ; | | i f DOCU <= 1 ; i f DOCU > 2 | i f SITE <= 3 | | i f TEAM > 3 ; | | i f TEAM <= 3 ; | i f SITE > 3 ; PCAP <= 1 ;
> > > > > > > > > > > > > >
if | | if | | | | | | | |
DATA > 3 i f RUSE <= 3 ; i f RUSE > 3 ; DATA <= 3 i f PCON > 1 | i f DATA > 2 ; | i f DATA <= 2 | | i f DOCU > 4 ; | | i f DOCU <= 4 | | | i f TIME > 4 ; | | | i f TIME <= 4 ; i f PCON <= 1 ;
Fig. 2: Structural instability. Example of models learned in this paper. Results are from 4 repeats varying only the random seed, extracted from [12], on Coc1000 dataset from MOOT (see Table III). is still rarely treated as a first-class evaluation target. Prior work has shown that learned conclusions can vary across data contexts in defect prediction and effort estimation [7], that defect predictors can differ in robustness and stability [15], that bellwether effects complicate transfer learning across projects [16], that deep-learning software systems exhibit substantial training variance and reproducibility challenges [17], [18], and that causal graphs learned from SE data can be structurally unstable [10]. Yet studies typically emphasize mean optimization performance, while reporting little about the variance of the recommendations produced by repeated runs. To the best of our knowledge, this paper is the first largescale study of model instability in SBSE (“large scale” since, as shown below, we report results from 127 SE problems and 12,700 test cases). Stabilizing models at that scale shows that instability is not merely noise to tolerate, but a property that can be measured and managed. The methods presented here should therefore be viewed as a baseline for future work on SBSE instability.
as independent drivers, with a validated configuration that improves both stability and quality; and (4) evidence of a datainherent floor that learner-side and causal interventions reduce but cannot fully eliminate. To support open source, all our data and scripts are available at https://tinyurl.com/Model-Instability. Before we begin, we clarify how to read the main reported numbers and why they are important and new. We use three complementary summaries. First, test-case agreement measures performance stability. Out of 12,700 held-out test cases, how often do 20 repeated runs give sufficiently similar predictions? At our main threshold α = 0.35, agreement improves from 364 cases under defaults to 1,740 cases under refined settings, a 4.8× increase. Second, error spread measures the standard deviation of optimization error across repeated runs. As shown in Figure 1, the red line is mostly below the blue line; refined settings reduce the mean standard deviation of errors from 17.4 to 13.6, a 22% reduction across all 127 datasets. While Figure 1 asks whether repeated runs are more stable, Figure 9 asks whether the resulting recommendations remain high quality. Moving to the third, optimization quality is assessed by statistical top-rank counts across datasets. Refined settings are top-ranked on 119 of 127 datasets, compared to 74 for defaults. Thus, agreement and error spread describe stability, while top-rank counts describe recommendation quality. Importantly, the stability gains come with no performance compromise. These results matter because, although instability has been reported in several areas of software analytics, it
II. BACKGROUND A. Motivation Why study instability in SE optimization? Because SE systems rarely optimize a single objective. Tasks such as configuration tuning, test prioritization, and project scheduling must balance competing goals, including: 1) Run more database queries, faster, with less energy; 2) Deliver better software, faster, yet cheaper;
2
% repeats in run 2 100 100 100 88 88 55 55 22 22
Topic name Xaml Binding Ruby on Rails MVC Objective C Function Return Types ... MySQL HTML Form ... Visual Studio Information Systems
Top 9 words in topic grid,window,bind, . . . rubi,rail,gem, . . . http,com,java, . . . self,cell,nsstring, . . . function,const,char,. . . connect,sql,databas, . . . form,login,user, . . . window,project,file, . . . messag,email,log, . . .
TABLE I: Examples of LDA topic instability across two runs, adapted from Agrawal et al. [8]. Fig. 3: Coefficient instability in effort estimation [7]. βi coefficients from 20 samples of the same dataset, with several crossing zero thresholds, showing that same feature increases effort in some samples, decreases it in others.
drift over time. Across five cross-project defect prediction techniques and 33 versions of 14 Java projects, rankings and performance shifted by statistically significant margins in time-aware evaluations, driven by time, but also by dataset noise and the nature of CPDP approaches. For causal graphs, widely used as succinct representations of complex SE data, Hulse et al. [10] report highly unstable shapes. When applying four generators to 23 datasets across three SE tasks, over half the causal edges disappeared or changed direction. Around 75% of edges differ between consecutive project releases, and removing just 10% of training data can cause large structural changes. Any specific causal conclusion (e.g. “long functions cause more defects”) could be reversed by minor changes to data or method. With the growing adoption of LLMs in SE, the same concern recurs. A seminar of SE/AI researchers flagged the instability of generated results as a core obstacle to reliable LLM4SE [22], and the effect is empirically reproducible. In error correction, LLMs are demonstrably unstable, with fix consistency degrading as the sampling temperature rises [11]. In vulnerability detection, Han et al. [23] tested 6 LLMs on 48 scenarios drawn from the MITRE Top-25 weaknesses, across 17 prompt structures with 20 paraphrases each. Semantically equivalent rephrasings alone produced substantially more output variability than repeated identical prompts, with codespecialized models more brittle than general-purpose ones. In summary, instability is not limited to specific techniques or tasks. It is pervasive across SE research, observed in reproducibility studies [24] and large-scale empirical analyses [25]. Even methodological choices alone,tuning [26] or validation strategy [27],can shift rankings and conclusions. Hence, we say instability is a first-class, systemic issue in software analytics and must be actively investigated and mitigated.
3) Maximize stakeholder satisfaction with less complexity; 4) Test for more bugs with less effort. Search-based software engineering exists to negotiate exactly these trade-offs. The choice spaces it must explore are enormous. Configuration spaces are combinatorially vast1 , under-documented, and laced with subtle interactions. Defaults cannot be trusted, hand-tuning does not scale, and human intuition often misleads [20]. As systems grow more intricate, managing their configurations becomes increasingly difficult [21]. To navigate vast search spaces, learning algorithms often use random search, causing repeated runs to yield different models and recommendations, leading to structural and performance instability. This instability has been seen in many SE domains. In software P project effort estimation, regression learns y = β0 + i βi xi , where positive or negative βi means feature xi increases or decreases estimated effort. In theory, those βi could be used to prioritize changes within a project. In practice, instability makes this very difficult. Figure 3, from Menzies et al. [7], shows the βi learned from 20 different 66% samples of the same project dataset. The coefficients swing wildly, and half a dozen even cross βi = 0. The same feature is estimated to both increase and decrease effort, making any consistent interpretation impossible for practitioners2 . In SE topic modeling, Agrawal et al. [8] show that standard LDA (Latent Dirichlet allocation) is severely unstable. Prior researchers had used LDA to report “what developers talk about on Stack Overflow,” with no hint that their topics were unstable. That was misleading. Merely randomizing the input order dramatically changes the topics reported. Running LDA twice with different orderings (Table I), only 3 of 40 topics recurred, and over half (25 of 40) had 55% or less overlap. At the conclusion level in defect prediction, Bangash et al. [9] found that conclusions about model performance
B. Structural Instability and the Rashomon Effect This paper measures both structural and performance instability but seeks to mitigate only the latter. The reason is not neglect. Theory predicts (and RQ1 confirms) that structural instability is fundamental, an unavoidable consequence of the many different models that fit the same data. Structural instability is a variation in model form and/or the features used. To understand its nature, consider a model with
1 For example, Table 1 of Agrawal et al. [19] reports one hundred sextillion (1046 ) ways to configure text-preprocessing data miners. 2 Even a simple regression can fluctuate so wildly. See all the parameters in scikit-learn’s LinearRegression: https://scikit-learn.org/stable/modules/ generated/sklearn.linear model.LinearRegression.html.
3
# 1 2
M = 10 binary inputs: 210 = 1024 input patterns, of which the model may internally use any subset. There are 10
22
≈ 1.8 × 10308
such subsets (more than the electrons in the observable universe). This is the Rashomon effect: an enormous number of structurally different models can fit the same data [13], [14]. Semenova et al. [28] formalized this via the Rashomon ratio: the fraction of models achieving near-optimal performance. When that ratio is large, the selected model reflects random variation in the data, not a stable underlying signal. Rudin et al. showed the Rashomon set can be enormous for tabular data [29]. Xin et al. [14] showed that for sparse decision trees, it can be exponentially large, making structural variability the norm. Crucially, Semenova et al. [30] explain why Rashomon ratios tend to be large. Label noise widens generalization gaps and forces simpler model classes with greater multiplicity. SE datasets, often noisy [31], [32], small, and labeled under tight budgets [33], [34], are precisely those conditions. From this perspective, structural instability in SE optimization is not an algorithmic defect; it is a property of the problem itself. In summary, theory predicts structural instability is something to measure, not fix. RQ1 confirms this empirically. Even settings that greatly improve prediction agreement leave the tree structure as diverse. Hence, the rest of this paper measures both instabilities but only mitigates performance instability.
3 4 5
Instability Cause Labeling Budget Acquisition Strategy Model Complexity Splitting Criterion Causal-inspired
6
Data Locality
Treatments in Experiment Labels = [10, 20, 50, 100, 200] [Xplor, Xploit, adapt, bore, near, random] Min leaf size = [1, 3, 5, 7, 9] [Entropy (uses log), Gini (no log)] Confounder filter & Gain-Ratio split = [On,Off] Anchors = [HDBSCAN, CURE, KMeans]
RQ 2 2 2 2 3 3
TABLE II: Potential sources of performance instability and the corresponding experimental treatments. The first four are studied in RQ2, the last two in RQ3. system3 . The labeling budget caps how much data the learner ever sees, so with small budgets, each draw sees a different set of rows and tells a different story. We sweep budgets from 10 to 200 (upper bound is empirically justified in §III-A). Model Complexity. A shallow model yields a few coarse rules, while a deeper one chases nuances supported by less and less data, so sample size can change a tree’s lower structure. In decision tree learning, the minimum leaf size controls the model’s complexity. Splitting Criterion. Models divide regions of data in order to, say, minimize expected post-division impurity. There are many measures of impurity. Entropy counts the bits needed to encode a distribution via (− log p), which runs to infinity as p → 0. Since its slope is 1/p, small changes in rare-event probabilities can dramatically change what is selected by this particular splitting criterion. Other criteria are not so unstable. P Gini impurity (1 − i p2i ) has a bounded slope of 2p, so rare events cannot dominate splits (the way they do with entropy). Acquisition Strategy. Active learners use the model built so far to decide what row label to acquire next, thus avoiding irrelevant and noisy data. Different acquisition rules select different rows and yield different theories, making the rule itself a source of instability. After sorting labeled data into better and worse regions, our learners estimate the probability that an unlabeled row is best (b) or rest (r). From the literature, we extract five rules. Bore and near greedily seek the bestnot-rest row via Bayes statistics (bore) or distance to the best centroid (near). Xplor, xploit, and adapt label the row maximizing b + rq abs(bq − r + ϵ)
C. Causes of Instability The Rashomon effect implies that the same data can support many models. A learner’s inductive biases determine which one is selected. Those same biases also create instability: change a design choice (e.g., the random seed), and the learner may produce a different model. There are many such choices. This section examines six, summarized in Table II, ranging from familiar tuning knobs to subtler effects. The first four are internal to the learner (RQ2), and the last two concern the learner’s interaction with the data (RQ3). No such list is complete. We chose these six because they are widely applicable (most affect a broad range of learners), effective (as shown below, they can lift stability from under 50% to over 80%), and vetted (some (e.g., Causalinspired) were added after discussion with leading figures in the empirical SE community. Other candidates were ruled out experimentally: feature selection, initially included, in fact increased instability, since with fewer features the variance of a single attribute can dominate. Labeling Budget. A learner builds a model y = f (x) from examples (x1 , y1 ), (x2 , y2 ), .... We denote by yi the label of each example. Obtaining labels is the costly part of SE optimization since it means running or deploying a real
where ϵ avoids divide-by-zero and if exploiting 0 q= 1 if exploring 1− M if adapting B
(1)
(M = configurations evaluated so far; B = total budget). When labels are scarce, explore favors rows where the best and rest votes are close, opposite ideas at similar weight. As data 3 Finding authoritative labels remains a major challenge for SE research [35], [36]. Labeling by human experts is slow and error-prone when rushed [20], [37], historical logs can be unreliable [31], [38]–[40], and automated labeling is crude [41] or, in the case of LLMs, only assistive (not authoritative). Even naturally occurring oracles (compile, then run the full test suite) can be extremely slow [42].
4
# 25 12 1 1 1 1 7 35 3 8
Dataset Type Software Configurations PromiseTune Software Configurations Cloud Cloud Cloud Cloud Cloud Software Project Health Scrum Feature Models
File Names SS-A to SS-X, billing10k 7z, BDBC, HSQLDB, LLVM, PostgreSQL, dconvert, deeparch, exastencils, javagc, redis, storm, x264 HSMGP num Apache AllMeasurements SQL AllMeasurements X264 AllMeasurements (rs–sol–wc)* Health-ClosedIssues, Health-PRs, Health-Commits Scrum1k, Scrum10k, Scrum100k FFM-*, FM-*
1 1 4 4 2 3 4
Software Process Model Software Process Model Software Process Model Software Process Model Software Testing Miscellaneous Behavioral
4 3 5
Financial Human Health Data Sales
2 127
Reinforcement Learning Total
nasa93dem COC1000 POM3 (A–D) XOMO (Flight, Ground, OSP) test120, test600 auto93, Car price, Wine quality all players, HR-employeeAttrition, student dropout, player statistics BankChurners, home data, Loan, Telco-Churn COVID19, Life Expectancy, hospital Readmissions accessories, dress-up, Marketing Analytics, socks, wallpaper A2C Acrobot, A2C CartPole
Primary Objective Optimize software system settings Software performance optimization
x/y 3–88 / 2–3 9–35 / 1
# Rows 197–86,059 864–166,975
Hazardous Software Management Program data Apache server performance optimization SQL database tuning Video encoding optimization Misc configuration tasks Predict project health and developer activity Configurations of the scrum feature model Optimize number of variables, constraints, and clause/constraint ratio Optimize effort, defects, time, and LOC Optimize risk, effort, analyst experience, etc. Balance idle rates, completion rates, and cost Optimize risk, effort, defects, and time Optimize test selection Miscellaneous optimization tasks Analyze and predict behavioral patterns
14 / 1 9/1 39 / 1 16 / 1 3–6 / 1 5 / 2–3 124 / 3 128–1,044 / 3
3,457 192 4,654 1,153 196–3,840 10,001 1,001–100,001 10,001
24 / 3 20 / 5 9/3 27 / 4 9/1 5–38 / 2–5 26–55 / 1–3
93 1,001 501–20,001 10,001 5,161 205–1,600 82–17,738
Financial analysis and prediction Health-related analysis and prediction Sales analysis and prediction
19–77 / 2–5 20–64 / 1–3 14–31 / 1–8
1,460–20,000 2,938–25,000 247–2,206
Reinforcement learning tasks
9–11 / 3–4
224–318
TABLE III: Summary of the 127 datasets used in this study from Chen & Menzies’ MOOT repository http://tiny.cc/moot [43]. The table reports dataset categories, representative file names, optimization objectives, number of decision variables and objectives (x/y), and dataset sizes. is removed as confounded by Z (mutual confounders tie-break by smaller H(Y | X)/H(Y )). After filtering, splits are scored by the gain ratio 1−H(Y | X)/H(Y ), the most stable splitting criterion per Lee et al. [51]. (Aside: note our terminology. we use “causal” in the sense of confounder screening per PromiseTune [50] and not as per the full do-calculus.)
accrue, exploit jumps to the strongest best and weakest rest; adapt slides from explore to exploit as M grows toward B. Data Locality. SE repositories often combine data from different developers, technologies, and problem domains. When modeled as a single population, such heterogeneous data may support many equally plausible models (Rashomon again), increasing instability. One way to reduce this effect is to first partition the data into more homogeneous regions. To test whether stability gains depend on a particular clustering paradigm, we evaluate one method from each major family: KMeans [44] (partitional clustering around iteratively updated centroids), HDBSCAN [45] (density-based clustering that discovers arbitrarily shaped regions), and CURE [46] (representative-point clustering that captures non-spherical structures while remaining robust to outliers). Causal-inspired. Pearl warns that learners built on raw correlation invent influences that do not exist [47], [48], lacking causal knowledge, they cannot tell whether A correlates to B because A drives B, or because both ride a hidden third variable. Such spurious inferences inflate the space of plausible models. Causal knowledge can steady a learner. Unicorn [49] uses causality to sharpen configuration models, and PromiseTune [50] filters spurious predictors before search. Following PromiseTune, we add a confounder filter to our model construction. Each labeled row gets a target Y equal to its d2h score (numerics discretized into 10 equal-probability bins). The dependency of feature X on Y is the drop in target impurity after conditioning on X: X D(X, Y ) = I(Y ) − p(x) I(Y | X = x) (2)
III. E XPERIMENTAL D ESIGN This section describes the data, measures, and algorithms used to answer our four research questions. A. Data We use 127 multi-objective SE optimization problems from Chen & Menzies’ MOOT repository http://tiny.cc/moot [43]. See Table III. For our experiments, each dataset is split randomly 50:50 into train and hold-out halves (and all our runs are from 20 such runs). We use MOOT since, to our knowledge, it is the largest collection of real multi-objective SE optimization tasks assembled to date. MOOT’s data come from numerous SE papers by numerous SE authors, presented in leading venues (ICSE, FSE, TSE, IST, EMSE, TOSEM, ASE). MOOT’s tasks span configuration, performance tuning, product lines, project health, defect prediction, testing, cost estimation, cross-domain generalization, and text mining. Earlier resources (e.g., SPLOT) were narrow and are now offline. Toolkits like Pygmo or Platypus offer only synthetic benchmarks. MOOT’s datasets come from published studies, real performance logs, cloud systems, and tuning tasks where bad configurations cost time, money, and credibility. One note on data access. While MOOT problems can hold hundreds to 100,000s of rows (see Table III), we demand that our optimizers access the y goal values of no more than 200 rows. That cap was set empirically. We increased the
x
where I(·) is entropy or Gini. Features are kept when D(X, Y ) beats a permutation null at α significance. Among survivors, if conditioning on some Z renders D(X, Y | Z) negligible, X
5
Fig. 4: Overview of the EZR active-learning pipeline, from warm-start initialization through iterative acquisition and labeling to final tree construction. EZR begins by labeling a small random warm-start sample. These examples are ranked by d2h and divided into best and rest. An acquisition function then iteratively selects additional rows for labeling, updating the best/rest partition after each new label, until the labeling budget n1 is exhausted. We evaluate six acquisition functions (Table II, formulas in §II-C). Once labeling is complete, EZR constructs a regression tree from the n1 labeled examples. Each split selects the feature and threshold that minimize the expected standard deviation of d2h in the resulting partitions. Recursion continues until the minimum leaf size is reached. Each leaf predicts the median d2h of its training examples, which is then converted to the win score defined in §III-B. To generate a recommendation, the learned model ranks the hold-out examples by predicted score. The top n2 candidates are then labeled, and the best among them is returned. The entire process, therefore, consumes n1 +n2 labels. Throughout this paper, we use the default value n2 = 10, so all reported labeling budgets should be interpreted as requiring an additional ten labels. Important note: Throughout the results, we compare two EZR configurations: initial (the defaults of [52]) and refined (Table IV). RQ0 and RQ1 use both as fixed treatments; RQ2 reports the factor-by-factor experiments from which the refined settings were derived. Other learners. RQ3 also runs three clusterers (KMeans, HDBSCAN, CURE; rationale in §II-C) as stability anchors. Each implements Model (row ) by returning the median d2h of the labeled rows in the cluster nearest to that row. RQ3’s causal variant applies the confounder filter and gain-ratio splits of §II-C to EZR’s tree construction; nothing else changes.
labeling budget until further labels stopped buying performance, a ceiling reached at 200 labels (see Figure 8, topleft). This is consistent with (indeed, slightly more generous than) the 150-label saturation effect reported by Menzies and Srinivasan [52], who found no significant improvements after 150 examples (where those examples were carefully selected by the optimizer to avoid noisy and irrelevant items). B. Measures Configuration quality is measured as “distance-to-heaven” (d2h), normalized to win (to allow for comparisons across different data): v um uX (oi (x) − hi )2 , d2h(x) = t
Q = d2h(x) − min,
(3)
i=1
win(x) = 100
1−
Q median − min
(4)
Here, oi (x) is the normalized objective of configuration x, hi the ideal value of objective i, and m the number of objectives. Also, min and median are the lowest and middle values for d2h seen in the data set. Note that smaller d2h and Q values are better while larger values for wins are better. (Aside: later in this paper, we will refer to Q as the optimization quality). One definition is used for everything below: every algorithm in this study returns a function Model (row ) → predicted win score. Recommendations, performance, and all stability measures are computed from these predictions, so trees, clusterers, and baselines are scored on the same footing. C. Primary Learner Our primary learner is EZR [53], the active-learning optimizer shown in Figure 4. Recent studies in 2026 by Rayegan et al. and Ganguly et al., published in JSS [34] and FSE [54], establish EZR as a state-of-the-art baseline in SE multiobjective optimization [55] and explanation [34], [52]. EZR learns decision trees from only a few dozen labeled examples while achieving state-of-the-art optimization performance and running up to 500× faster than methods such as SMAC3 [56]. We study EZR for three reasons. First, it is a recent state-ofthe-art method. Second, it exhibits the instabilities investigated in this paper (all models in Figure 2 were generated by EZR). Third, it has already been evaluated on the 127 MOOT tasks described in §III-A, allowing us to assess stability across a large and diverse benchmark suite.
D. Measuring Performance Each treatment is repeated 20 times with different random seeds. For run i, we compare the recommendation x̂i against x∗i , the true optimum (lowest d2h) in the hold-out: v u 20 u X 2 1 win(x̂i ) − win(x∗ RMSE = t 20 i)
(5)
i=1
Note that lower RMSE is better. Applying §III-F to perdataset RMSE yields each treatment’s performance win count: the number of datasets (out of 127) where each treatment has statistically better suggestions compared to other methods.
6
IV. R ESULTS
E. Measuring Stability
A. RQ0 - The Problem: How prevalent is performance instability in SE optimization, and what are its consequences for practitioners?
The same 20 models per treatment per dataset support two stability measures: one structural, one performance-based. Structural stability asks whether repeated runs produce similar model structures. Following Kalousis et al. [57], we use Jaccard similarity to compare the structure of two trees trained under identical settings. However, since the features appearing in splits closer to the root of the tree gain more importance, we measure the weighted Jaccard similarity: P min (WA (f ), WB (f )) f ∈F ] (6) [Jw (A, B) = P max (WA (f ), WB (f ))
Using §III-E, we train 20 trees per dataset under two settings (initial EZR defaults, the refined settings later recommended by RQ2) and count test-case agreements. At α = 0.35, the 20 trees agree on just 364 (Initial setting) and 1,740 (Refined setting) of 12,700 test cases: consistency rates of 2.9% and 13.7%. Figure 6 shows results sweeping the α threshold for both our baseline system (in red) and the refined system (in green) recommended later in this paper. This figure confirms that even at α = 1, agreement reaches only 4,020 and 8,360 cases (31.7% and 65.8%), and the refined configuration remains better. The practical consequence is that even under the most permissive threshold, and even after refinement, nearly half of all test cases fail to agree. Two runs of the same tool on the same data will usually return different predictions, undermining a team’s ability to reliably allocate testing budget, prioritize high-risk modules, or select near-optimal configurations.
f ∈F
where (F) is the P union of features appearing in either tree, and (WT (f ) = k∈Sf 1/dk ) assigns larger weights to features occurring closer to the root. The similarity ranges from 0 (disjoint features) to 1 (identical weighted features). Performance stability asks if repeated runs make consistent predictions, whatever their structure. We pass 100 randomly selected hold-out rows (fixed across treatments) through all 20 models. For each row, let σmodels be the standard deviation of the 20 predictions and σdata the standard deviation of true d2h across the dataset. The models agree on that row when σmodels < α × σdata
RQ0 answer. Performance instability is the norm, not the exception. Under default settings, repeated runs agree on 2.9% of test cases, and no choice of threshold rescues consistency. Practitioners may not be able to trust the recommendation of any single run.
(7)
i.e., when disagreement is small against the background variation. Any such threshold constant is debatable, so we guard against it two ways. First, we report results primarily at α = 0.35. Second, RQ0 sweeps the full range 0.15 ≤ α ≤ 1, from a strict test (α = 0.15) to the most permissive one (α = 1, where models “agree” whenever their spread is merely smaller than the standard deviation of the entire dataset). As shown later (Figure 6), our conclusions hold across the sweep and are not artifacts of α choice. Agreement aggregates two ways: the agreement rate (per dataset, the percent of the 100 rows where all 20 models agree) and the total agreement count (summed over all 127 × 100 = 12,700 test cases). Applying §III-F to perdataset agreement rates yields each treatment’s stability win count. Note that Figure 1 reports stability in terms of standard deviation of optimization error (lower = more stable), while Figure 6 reports agreement rates across test cases. Both measure performance instability but from complementary angles.
B. RQ1 - The Nature of the Problem: How prevalent is structural instability, and is it the same phenomenon as performance instability? Using the weighted Jaccard measure of §III-E, Figure 7 compares the feature sets of the 20 trees per dataset under the initial and refined configurations (recall from §III-C that the refined settings, derived in RQ2, are those that most improve performance stability). Two features of Figure 7 are noteworthy. Firstly, structural and performance stability are decoupled problems. The two curves, remaining closely aligned, show that the refined settings that improve prediction agreement more than sixfold (RQ0) produce trees that are structurally as different as the defaults. Figure 5 showed the effect on one dataset. Four refined-setting trees still differ markedly in features, thresholds, and topology. One practical corollary: the features of any single tree should not be over-interpreted as the explanation of a recommendation, since other runs select other features while recommending the same things. Secondly, structural instability is not the instability that should be repaired, for two reasons. First, it is largely irreducible: the Rashomon effect (§II-B) guarantees many feature sets fit equally well, which is why even our most stabilizing settings barely move structural similarity. Second, it may not be what hurts practitioners. The decoupling shown above means trees can differ in structure yet agree on their recommendations, so it is performance
F. Statistics To compare treatments, we require statistical significance and practical effect size, avoiding p-value-only pitfalls [58]. A K-S test [59] at α = 0.05 plus Cliff’s delta δ ≤ 0.195 (small-to-medium) [60]–[62]. To rank many treatments we find the top set: a variance-maximizing iterative partition [30] peels away distinguishably worse treatments, retaining those statistically tied with the best.
7
i f SCED > 1 i f PREC > 3 | i f TIME <= 5 | | i f SITE <= 1 | | | i f FLEx > 2 ; | | | i f FLEx <= 2 ; | | i f SITE > 1 | | | i f DOCU > 3 | | | | i f PCAP > 2 ; | | | | i f PCAP <= 2 ; | | | i f DOCU <= 3 | | | | i f TEAM > 2 ; | | | | i f TEAM <= 2 ; | i f TIME > 5 ; i f PREC <= 3 | i f TEAM > 3 ; | i f TEAM <= 3 | | i f PCON <= 3 ; | | i f PCON > 3 ; i f SCED <= 1 ;
| | | | | | | | | | | | | | | | | |
> > > > > > > > > > > > > > > > > > > > > >
i f CPLx <= 5 i f PCON <= 2 | i f RUSE <= 4 | | i f TIME <= 4 ; | | i f TIME > 4 | | | i f DOCU <= 3 ; | | | i f DOCU > 3 ; | i f RUSE > 4 | | i f SCED > 2 ; | | i f SCED <= 2 ; i f PCON > 2 | i f PCAP > 3 | | i f TOOL <= 3 ; | | i f TOOL > 3 ; | i f PCAP <= 3 | | i f DOCU <= 3 ; | | i f DOCU > 3 | | | i f RUSE <= 4 ; | | | i f RUSE > 4 ; i f CPLx > 5 ;
| | | | | | | | | | | | | | | | | |
> > > > > > > > > > > > > > > > > > > > > >
i f TOOL <= 4 i f TOOL > 2 | i f FLEx > 2 | | i f PMAT <= 2 ; | | i f PMAT > 2 | | | i f CPLx <= 4 | | | | i f STOR > 4 ; | | | | i f STOR <= 4 ; | | | i f CPLx > 4 ; | i f FLEx <= 2 | | i f SCED > 3 ; | | i f SCED <= 3 ; i f TOOL <= 2 | i f PCON > 4 ; | i f PCON <= 4 | | i f STOR > 4 ; | | i f STOR <= 4 ; i f TOOL > 4 | i f TEAM <= 3 | | i f PVOL <= 3 ; | | i f PVOL > 3 ; | i f TEAM > 3 ;
| | | | | | | | | | | | | | | |
> > > > > > > > > > > > > > > > > > > > > >
i f PREC > 5 i f DATA <= 3 ; i f DATA > 3 ; i f PREC <= 5 | i f LTEx > 4 | | i f ARCH <= 2 ; | | i f ARCH > 2 | | | i f FLEx <= 2 ; | | | i f FLEx > 2 | | | | i f TIME <= 4 ; | | | | i f TIME > 4 ; | i f LTEx <= 4 | | i f PCAP > 3 | | | i f TIME <= 4 | | | | i f PCAP <= 4 ; | | | | i f PCAP > 4 ; | | | i f TIME > 4 ; | | i f PCAP <= 3 | | | i f TEAM > 4 ; | | | i f TEAM <= 4 | | | | i f CPLx > 3 ; | | | | i f CPLx <= 3 ;
| |
Fig. 5: Four trees trained on coc1000.csv (same data as Figure 2) with RQ2-refined settings and different random seeds produce markedly different structures yet more consistent optimization recommendations compared to Figure 2, illustrating that structural and performance instability are decoupled phenomena. Factor Labeling budget Acquisition strategy Min leaf size Splitting criterion
agreement, not feature overlap, that governs whether a single run can be trusted. RQ1 answer. Structural and performance instability are distinct. Settings that sharply improve prediction agreement leave tree structure about as varied as before. We target performance instability because it is the fixable one that governs trust. Structural instability is largely irreducible (Rashomon) and, being decoupled from recommendations, need not be fixed.
Default 20 Xploit 2 Entropy
Recommended 50 near 3 Gini
TABLE IV: Suggested EZR settings from RQ2 experiments. Performance gains saturate. As seen top-left of Figure 8, improvements beyond 50 labels are relatively modest, and a performance ceiling of 100% is reached at 200 labels. • Other factors do as much for free. As shown below, the acquisition strategy buys comparable stability gains with no extra labels. The acquisition strategy shows a much stronger effect. Of the six strategies. Near (label the candidate closest to the current best centroid) lifts stability top-rant rate from under 50% to over 80%. Note that adopting the near comes at almost no cost (no extra labels, no change of objective, no loss of recommendation quality). Model complexity (minimum leaf size) trades three ways. Deeper trees find better configurations but are less stable (finer partitions are more sensitive to sampling noise) and harder for practitioners to read, undermining the very point of humanreadable rules. min leaf = 3 best balances all three. Splitting criterion is a smaller but free win. Entropy’s log(p) amplifies tiny perturbations in rare-event probabilities into large swings in split selection, precisely under the tight budgets where SE optimizers run (§II-C). Gini matches entropy’s performance with clearly better stability. A zero-cost change, applicable immediately. Varying one factor at a time cannot reveal interactions, so the four recommended settings (Table IV) must also be evaluated as a single refined configuration. Figures 9 and 1 provide that validation. Across all 127 datasets, the refined configuration reduces optimization-quality standard deviation by 22% on average (Figure 1) while achieving statistically topranked optimization quality (§III-F) on 119 datasets, compared to 74 for the default configuration (Figure 9). Thus, the improvements observed for individual factors compound. Refined •
C. RQ2 - The Fixes: Which factors drive performance instability, and what configuration changes reduce it without sacrificing optimization quality? Figure 8 explores parts of Table II one at a time. The main finding here is that stability and performance respond to different levers, sometimes independently, sometimes in opposition, and no single cause dominates. Labeling budget can improve both performance and structural metrics. As seen at the top-left of Figure 8, increasing the label budget raises the fraction of datasets where the treatment is statistically top-ranked for stability from under 40% to around 70%. However, for two reasons, chasing more labels is not our preferred fix for instability since:
Fig. 6: Agreement counts (out of 12,700 test cases) across thresholds α ∈ [0.15, 1.0] for initial and refined EZR setting.
8
Fig. 7: Structural similarity across 20 trees per dataset under initial and refined EZR settings. Refined settings greatly improve performance stability (Fig. 9) while producing little change in structural similarity, suggesting the two are decoupled. Causal pipeline wins on 10 more datasets (59 vs. 69) and, also collects more total agreements (1,740 vs. 2,390), although this difference is not significant. In short, causal pipeline helps only where confounded features drive instability, and elsewhere noise, small samples, or large Rashomon sets (§II.B) dominate (most datasets). Data-inherent limits: if instability were something a better learner could remove, the stablest learners we know should remove it. Clustering methods do lead clearly (Figure 11). But even the best of them agree on less than 51% of cases. Switching to the stablest model family buys at most half the benchmark, evidence (not proof) that the floor is set by the data itself: noise, tight labeling budgets, proxy objectives, and large Rashomon sets.
setting improves both stability and optimization quality. RQ2 answer. Labeling budget, acquisition strategy, model complexity, and splitting criterion each drive performance instability, via different trade-offs. The refined configuration Table IV (50 labels, near, min leaf=3, Gini) improves stability (22% reduction in prediction-error std) and quality simultaneously, achieving top-ranked optimization on 119 datasets versus 74 for initial settings. D. RQ3 - The Limits: Is the instability remaining after RQ2 the fault of the learner, or of the data? The previous section focused on the learner, tuning internal choices such as budget, acquisition strategy, leaf size, and splitting criterion. Here we shift focus to the data. The question is whether instability is driven primarily by spurious correlations among features or by heterogeneous subpopulations hidden within the rows. We test this hypothesis in two ways: by augmenting EZR with causal knowledge (confounder removal), and by replacing trees with stable-bydesign clusterers [63] that model local regions rather than a single global structure. If data-side effects dominate, these interventions should deliver gains far larger than those seen in RQ2. As shown below, they do not. Causal-inspired reasoning: comparing the causal pipeline (confounder filter + gain-ratio splits, §II-C) against refined EZR, performance is statistically indistinguishable (Figure 10: 124 vs. 122 dataset wins of 127), so filtering removes no useful signal. But its stability gains are confined to some datasets.
RQ3 answer. Instability is not fully curable by changing the learner. External interventions add little. Causal pipeline helps only on some datasets, and even the stablest clusterers agree on under 51% of cases. The residue appears data-inherent, a limit shared by all methods, to be measured and reported, not blamed on the tool.
V. T HREATS TO VALIDITY Sampling bias. No single paper can investigate every problem in software analytics. Hence, when we say “our methods tame instability,” we add the caveat: we have shown this only for the examples studied here. That said, our examples were many and varied: 127 optimization tasks, written by many
Fig. 8: Summary of four instability-cause experiments. Values indicate the percentage of datasets where each treatment achieved the highest statistical rank. Each panel reports stability (blue bars) and performance (red line or bars) across treatments. 9
Fig. 9: Performance results (win of optimization quality Q, defined in Equation 3) across 127 datasets, measured as average deviation from the referenced optimum (lower is better). Datasets sorted by initial performance. Refined settings are statistically top-ranked on 119 datasets versus 74 for the initial configuration. Note that, with our methods (the red line), optimization quality is nearly always improved, sometimes by very large amounts (e.g., see left-hand-side). interval overlap) might score the same runs differently, and we cannot rule out that our conclusions are tied to our measures. Three things mitigate this concern. First, both measures have precedent: weighted Jaccard follows the feature-stability literature [57], and our agreement threshold is grounded in standard effect-size practice [60]–[62]. Second, the threshold constant in our agreement test was swept across its full range (Figure 6) and our conclusions held at every setting. Third, our labeling budgets were capped at an empirically observed performance ceiling (200 labels §III-A), not an arbitrary choice. Still, replication with other instability measures would be a valuable check on this work.
Fig. 10: Performance (dataset statistical wins), stability (dataset statistical wins), and agreement (test-case count out of 12,700) for EZR and causal pipeline across all datasets.
VI. C ONCLUSION Performance instability is not a minor reproducibility nuisance, and it is a threat to trust in SE optimization. Under default settings, repeated runs agreed on only 2.9% of 12,700 test cases(RQ0). Refined settings (Table IV) increased agreement by 4.8×, reduced the standard deviation of optimization error by 22%, and improved optimization quality, achieving the statistical top rank on 119 of 127 datasets, up from 74 for the initial configuration(RQ2). These gains came from simple learner-side choices: using near instead of greedyexploit acquisition, Gini instead of entropy, min leaf =3, and a modest label budget of 50. But the larger lesson is not that instability disappears. It does not. Even after refinement, agreement reaches only 13.7% at our main threshold, and even the stablest learners agree on fewer than 51% of test cases(RQ3). Structural instability also remains widespread, confirming that different model structures can still support similar recommendations(RQ1). Thus, the goal should not be to force repeated runs into the same tree. The goal should be to know when to trust their recommendations. These findings imply three recommendations. First, SE optimization tools should reconsider unstable defaults, more specifically, entropy splitting and greedy-exploit acquisition.
Fig. 11: Comparing Performance and stability of EZR with clustering methods. Green bars and red lines show the stability and performance scores across 127 datasets. authors, from many leading venues (§III-A), all processed by a learner with recent state-of-the-art credentials (§III-C). External validity. We studied one family of learners (treebased active learners) plus three clusterers. Our specific refined settings (Table IV) may not transfer to, say, neural optimizers. That said, RQ3 suggests the data-inherent stability floor is shared across very different model families, though that is evidence, not proof, and checking other learners is future work. Construct validity. Instability is not a quantity with one agreed operationalization, and this paper measured it just two ways: structural instability via a weighted Jaccard over tree feature sets, and performance instability via cross-run agreement on predictions (§III-E). Other instruments (e.g. tree edit distance, rank correlation of recommendations, prediction-
10
Second, empirical SE papers should report agreement rates alongside optimization results, since mean performance alone hides cross-run disagreement. Third, researchers should avoid two shortcuts. First, chasing structural stability for its own sake, since Rashomon makes it unlikely, and RQ1 shows it is unnecessary; And second, do not expect a silver bullet. Instability has multiple causes, and effective fixes compound rather than replace one another. Stability should be a standard evaluation axis. Without it, optimization result is incomplete. As to future work, we recommend (a) treating stability, performance, and explainability as three objectives on a shared Pareto front; (b) attacking the data-inherent ceiling upstream, via better measurement and labeling practices; (c) repeating this analysis on other interpretable learners (rule lists, Bayesian models, association rule miners); and (d) testing causal augmentation in domains rich in organizational confounders, where RQ3 suggests it should shine.
[15] A. Kaur and K. Kaur, “An empirical study of robustness and stability of machine learning classifiers in software defect prediction,” in Advances in intelligent informatics. Springer, 2015, pp. 383–397. [16] R. Krishna and T. Menzies, “Bellwethers: A baseline method for transfer learning,” IEEE Transactions on Software Engineering, vol. 45, no. 11, pp. 1081–1105, 2018. [17] H. V. Pham, S. Qian, J. Wang, T. Lutellier, J. Rosenthal, L. Tan, Y. Yu, and N. Nagappan, “Problems and opportunities in training deep learning software systems: An analysis of variance,” in Proceedings of the 35th IEEE/ACM international conference on automated software engineering, 2020, pp. 771–783. [18] C. Liu, C. Gao, X. Xia, D. Lo, J. Grundy, and X. Yang, “On the reproducibility and replicability of deep learning in software engineering,” ACM Trans. Softw. Eng. Methodol., 2021. [Online]. Available: https://doi.org/10.1145/3477535 [19] A. Agrawal, W. Fu, D. Chen, X. Shen, and T. Menzies, “How to “dodge” complex software analytics,” IEEE Transactions on Software Engineering, vol. 47, no. 10, pp. 2182–2194, 2019. [20] M. Easterby-Smith, “The design, analysis and interpretation of repertory grids,” International Journal of Man-Machine Studies, vol. 13, no. 1, pp. 3–24, 1980. [21] T. Xu, L. Jin, X. Fan, Y. Zhou, S. Pasupathy, and R. Talwadker, “Hey, you have given me too many knobs!: Understanding and dealing with over-designed configuration in system software,” in Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering, 2015, pp. 307–319. [22] C. Gao, X. Hu, S. Gao, X. Xia, and Z. Jin, “The current challenges of software engineering in the era of large language models,” ACM Transactions on Software Engineering and Methodology, vol. 34, no. 5, pp. 1–30, 2025. [23] S. Han, T. Tan, Y. Miao, X. Chen, and N. Sun, “Prompting instability: An empirical study of llm robustness in code vulnerability detection,” in Australasian Joint Conference on Artificial Intelligence. Springer, 2025, pp. 233–245. [24] H. Semmelrock, T. Ross-Hellauer, S. Kopeinik, D. Theiler, A. Haberl, S. Thalmann, and D. Kowald, “Reproducibility in machine-learningbased research: Overview, barriers, and drivers,” AI Magazine, vol. 46, no. 2, p. e70002, 2025. [25] N. Hoess, C. Paradis, R. Kazman, and W. Mauerer, “Oops!... i did it again. conclusion (in-) stability in quantitative empirical software engineering: A large-scale analysis,” arXiv preprint arXiv:2510.06844, 2025. [26] W. Fu, T. Menzies, and X. Shen, “Tuning for software analytics: Is it really necessary?” Information and Software Technology, vol. 76, pp. 135–146, 2016. [27] A. Ali and C. Gravino, “An empirical comparison of validation methods for software prediction models,” Journal of Software: Evolution and Process, vol. 33, no. 8, p. e2367, 2021. [28] L. Semenova, C. Rudin, and R. Parr, “On the existence of simpler machine learning models,” in Proceedings of the 2022 ACM Conference on Fairness, Accountability, and Transparency, ser. FAccT ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 1827–1858. [Online]. Available: https://doi.org/10.1145/3531146. 3533232 [29] C. Rudin, C. Zhong, L. Semenova, M. Seltzer, R. Parr, J. Liu, S. Katta, J. Donnelly, H. Chen, and Z. Boner, “Amazing things come from having many good models,” arXiv preprint arXiv:2407.04846, 2024. [30] L. Semenova, H. Chen, R. Parr, and C. Rudin, “A path to simpler models starts with noise,” Advances in neural information processing systems, vol. 36, pp. 3362–3401, 2023. [31] X. Wu, W. Zheng, X. Xia, and D. Lo, “Data quality matters: A case study on data label correctness for security bug report prediction,” IEEE Transactions on Software Engineering, vol. 48, no. 7, pp. 2541–2556, 2021. [32] C. Seiffert, T. M. Khoshgoftaar, J. Van Hulse, and A. Folleco, “An empirical study of the classification performance of learners on imbalanced and noisy software quality data,” in 2007 IEEE International Conference on Information Reuse and Integration, 2007, pp. 651–658. [33] V. Nair, Z. Yu, T. Menzies, N. Siegmund, and S. Apel, “Finding faster configurations using flash,” IEEE Transactions on Software Engineering, vol. 46, no. 7, pp. 794–811, 2020. [34] A. Rayegan and T. Menzies, “Minimal data, maximum clarity: A heuristic for explaining optimization,” Journal of Systems and Software, vol. 238, p. 112897, 2026. [Online]. Available: https: //www.sciencedirect.com/science/article/pii/S0164121226001305
R EFERENCES [1] T. M. Abdellatif, L. F. Capretz, and D. Ho, “Software analytics to software practice: a systematic literature review,” in 2015 IEEE/ACM 1st International Workshop on Big Data Software Engineering. IEEE, 2015, pp. 30–36. [2] A. Begel and T. Zimmermann, “Analyze this! 145 questions for data scientists in software engineering,” in Proceedings of the 36th International Conference on Software Engineering, 2014, pp. 12–23. [3] J. Caldeira, F. Brito e Abreu, J. Cardoso, R. Simões, T. Oliveira, and J. Pereira dos Reis, “Software development analytics in practice: A systematic literature review: J. caldeira et al.” Archives of Computational Methods in Engineering, vol. 30, no. 3, pp. 2041–2080, 2023. [4] Z. Kotti, G. Gousios, and D. Spinellis, “Impact of software engineering research in practice: A patent and author survey analysis,” IEEE Transactions on Software Engineering, vol. 49, no. 4, pp. 2020–2038, 2022. [5] D. Zhang, S. Han, Y. Dang, J.-G. Lou, H. Zhang, and T. Xie, “Software analytics in practice,” IEEE software, vol. 30, no. 5, pp. 30–37, 2013. [6] S. He, X. Zhang, P. He, Y. Xu, L. Li, Y. Kang, M. Ma, Y. Wei, Y. Dang, S. Rajmohan et al., “An empirical study of log analysis at microsoft,” in Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2022, pp. 1465–1476. [7] T. Menzies, A. Butcher, D. Cok, A. Marcus, L. Layman, F. Shull, B. Turhan, and T. Zimmermann, “Local versus global lessons for defect prediction and effort estimation,” IEEE Transactions on software engineering, vol. 39, no. 6, pp. 822–834, 2012. [8] A. Agrawal, W. Fu, and T. Menzies, “What is wrong with topic modeling? and how to fix it using search-based software engineering,” Information and Software Technology, vol. 98, pp. 74–88, 2018. [9] A. A. Bangash, H. Sahar, A. Hindle, and K. Ali, “On the time-based conclusion stability of cross-project defect prediction models,” Empirical Software Engineering, vol. 25, pp. 5047–5083, 2020. [10] J. Hulse, N. U. Eisty, and T. Menzies, “Shaky structures: The wobbly world of causal graphs in software analytics,” Empirical Software Engineering, vol. 30, no. 5, p. 142, 2025. [11] M. B. Er, N. İlhan, and U. Kuran, “Analyzing the instability of large language models in automated bug injection and correction,” arXiv preprint arXiv:2509.06429, 2025. [12] A. Rayegan and T. Menzies, “Can causality cure confusion caused by correlation (in software analytics)?” arXiv preprint arXiv:2602.16091, 2026. [13] L. Breiman, “Statistical modeling: The two cultures (with comments and a rejoinder by the author),” Statistical science, vol. 16, no. 3, pp. 199–231, 2001. [14] R. Xin, C. Zhong, Z. Chen, T. Takagi, M. Seltzer, and C. Rudin, “Exploring the whole rashomon set of sparse decision trees,” Advances in neural information processing systems, vol. 35, pp. 14 071–14 084, 2022.
11
[35] T. Ahmed, P. Devanbu, C. Treude, and M. Pradel, “Can llms replace manual annotation of software engineering artifacts?” in MSR, 2025, pp. 526–538. [36] H. Tu and T. Menzies, “Frugal: Unlocking semi-supervised learning for software analytics,” in ASE, 2021, pp. 394–406. [37] R. Valerdi, “Heuristics for systems engineering cost estimation,” IEEE Systems Journal, vol. 5, no. 1, pp. 91–98, 2010. [38] Z. Yu, F. M. Fahid, H. Tu, and T. Menzies, “Identifying self-admitted technical debts with jitterbug: A two-step approach,” IEEE Trans. Softw. Eng., 2021. [39] H. J. Kang, K. L. Aw, and D. Lo, “Detecting false alarms from automatic static analysis tools: How far are we?” in 44th Int. Conf. Softw. Eng., 2022. [40] M. Shepperd, Q. Song, Z. Sun, and C. Mair, “Data quality: Some comments on the nasa software defect datasets,” IEEE Trans. Softw. Eng., vol. 39, no. 9, pp. 1208–1215, 2013. [41] Y. Kamei, E. Shihab, B. Adams, A. E. Hassan, A. Mockus, A. Sinha, and N. Ubayashi, “A large-scale empirical study of just-in-time quality assurance,” IEEE Trans. Softw. Eng., vol. 39, no. 6, pp. 757–773, 2012. [42] P. Valov, J. Petkovich, J. Guo, S. Fischmeister, and K. Czarnecki, “Transferring performance prediction models across different hardware platforms,” in Proc. of the 8th ACM/SPEC on Int. Conf. Perf. Eng., 2017, pp. 39–50. [43] T. Menzies, T. Chen, Y. Ye, K. K. Ganguly, A. Rayegan, S. Srinivasan, and A. Lustosa, “Moot: a repository of many multi-objective optimization tasks,” in 2026 IEEE/ACM 23rd International Conference on Mining Software Repositories (MSR), 2026. [44] J. MacQueen, “Some methods for classification and analysis of multivariate observations,” in Proceedings of the 5th Berkeley Symposium on Mathematical Statistics and Probability, vol. 1. University of California Press, 1967, pp. 281–297. [45] R. J. G. B. Campello, D. Moulavi, and J. Sander, “Density-based clustering based on hierarchical density estimates,” in Advances in Knowledge Discovery and Data Mining (PAKDD), ser. Lecture Notes in Computer Science, vol. 7819. Springer, 2013, pp. 160–172. [46] S. Guha, R. Rastogi, and K. Shim, “Cure: An efficient clustering algorithm for large databases,” ACM SIGMOD Record, vol. 27, no. 2, pp. 73–84, 1998. [47] J. Pearl, Causality. Cambridge university press, 2009. [48] J. Pearl and D. Mackenzie, The Book of Why: The New Science of Cause and Effect, 1st ed. USA: Basic Books, Inc., 2018. [49] M. S. Iqbal, R. Krishna, M. A. Javidian, B. Ray, and P. Jamshidi, “Unicorn: Reasoning about configurable system performance through
the lens of causality,” ser. EuroSys ’22. New York, NY, USA: Association for Computing Machinery, 2022, p. 199–217. [Online]. Available: https://doi.org/10.1145/3492321.3519575 [50] P. Chen and T. Chen, “Promisetune: Unveiling causally promising and explainable configuration tuning,” 2025. [51] J. Lee, M. K. Sim, and J.-S. Hong, “Assessing decision tree stability: a comprehensive method for generating a stable decision tree,” IEEE Access, vol. 12, pp. 90 061–90 072, 2024. [52] T. Menzies and S. Srinivasan, “Can ai be easy? lessons learned from the ezr.py toolkit,” 2026. [Online]. Available: https://arxiv.org/abs/2606. 03640 [53] T. Menzies, “The case for compact ai,” Communications of the ACM, vol. 68, no. 8, pp. 6–7, 2025. [54] K. K. Ganguly and T. Menzies, “How low can you go? the data-light se challenge,” Proceedings of the ACM on Software Engineering, no. FSE, 2026. [55] K. Ganguly and T. Menzies, “Zoom, don’t wander: Why regional search outperforms pareto reasoning and global optimization in budgetconstrained sbse,” arXiv preprint arXiv:2605.09658, 2026. [56] M. Lindauer, K. Eggensperger, M. Feurer, A. Biedenkapp, D. Deng, C. Benjamins, T. Ruhkopf, R. Sass, and F. Hutter, “Smac3: A versatile bayesian optimization package for hyperparameter optimization,” Journal of Machine Learning Research, vol. 23, no. 54, pp. 1–9, 2022. [57] A. Kalousis, J. Prados, and M. Hilario, “Stability of feature selection algorithms: a study on high-dimensional spaces,” Knowledge and information systems, vol. 12, no. 1, pp. 95–116, 2007. [58] V. B. Kampenes, T. Dybå, J. E. Hannay, and D. I. Sjøberg, “A systematic review of effect size in software engineering experiments,” Information and Software Technology, vol. 49, no. 11-12, pp. 1073–1086, 2007. [59] H. W. Lilliefors, “On the kolmogorov-smirnov test for normality with mean and variance unknown,” Journal of the American statistical Association, vol. 62, no. 318, pp. 399–402, 1967. [60] R. Rosenthal, H. Cooper, L. Hedges et al., “Parametric measures of effect size,” The handbook of research synthesis, vol. 621, no. 2, pp. 231–244, 1994. [61] S. S. Sawilowsky, “New effect size rules of thumb,” Journal of modern applied statistical methods, vol. 8, no. 2, p. 26, 2009. [62] G. Macbeth, E. Razumiejczyk, and R. D. Ledesma, “Cliff’s delta calculator: A non-parametric effect size program for two groups of observations,” Universitas Psychologica, vol. 10, no. 2, pp. 545–555, 2011. [63] L. Breiman, “Heuristics of instability and stabilization in model selection,” The Annals of Statistics, vol. 24, no. 6, pp. 2350–2383, 1996.
12