The Replication Assessment Problem in Software Engineering Giuseppe Destefanis
University College London UK [email protected]
Martin Shepperd
Brunel University of London UK [email protected]
Background: Replication studies in software engineering are increasingly common, yet their interpretation remains uncertain and inconsistent because assessments frequently rely on loosely defined or ad hoc criteria. Aim: This study aims to document how replication study outcomes are currently assessed in empirical software engineering, identify problems arising from inconsistent criteria and propose a principled framework for meaningful evaluation. Method: We conducted a systematic review of replication studies, with the search covering recent empirical software engineering replications (2021–2025). For each study, we extracted the criteria used to assess replication outcomes and analysed these for heterogeneity, logical consistency, and alignment with established statistical principles. Results: A total of 10 replication studies were located. The analysis reveals substantial heterogeneity in assessment practices, with contradictory criteria applied to similar data, limited acknowledgement of measurement uncertainty, and an absence of shared standards. We propose a principled framework grounded in statistical, methodological, and measurement considerations, and demonstrate its application through worked examples. Conclusions: Adopting consistent and transparent assessment principles would reduce ambiguity, improve comparability, and support more reliable evidence accumulation in software engineering replication research.
remains largely unaddressed. How should the outcomes of replication studies be assessed? How do we decide that the results of a replication are congruent or not? When a replication reports results that differ from the original, under what conditions should this be interpreted as a failure to replicate, as evidence of boundary conditions or as an artefact of measurement variation? Conversely, when results appear similar, what criteria determine whether the replication has succeeded? Current practice exhibits substantial heterogeneity in how these questions are answered. Researchers employ different metrics, statistical tests, significance thresholds, effect size comparisons, and qualitative criteria when deciding whether a replication confirms or fails to support a prior study [19, 27]. This heterogeneity means that two replication studies with numerically similar results may reach opposite conclusions, depending on the criteria applied. Shepperd et al. [27] provide indirect evidence of this problem. In their review of replications in software effort prediction and pair programming, internal replications reported confirmatory evidence approximately eight times more frequently than external replications. While this disparity may partly reflect legitimate procedural differences, it also raises questions about whether assessment criteria differ systematically and whether such differences are justified. Without agreed criteria for what constitutes a successful replication, claims about the state of replication in software engineering lack an objective basis. This descriptive study addresses the following research questions:
ACM Reference Format: Giuseppe Destefanis, Martin Shepperd, and Leila Yousefi. 2026. The Replication Assessment Problem in Software Engineering. In International Conference on Evaluation and Assessment in Software Engineering (EASE ’26), June 09–12, 2026, Glasgow, United Kingdom. ACM, New York, NY, USA, 6 pages. https://doi.org/10.1145/3816483.3816559
RQ1: What criteria are currently used to assess replication outcomes in empirical software engineering, and how do these criteria vary across studies?
Abstract
arXiv:2607.13815v1 [cs.SE] 15 Jul 2026
Leila Yousefi
Ministry of Justice UK [email protected]
1
Introduction
Replication is central to the accumulation of reliable empirical knowledge. By subjecting prior findings to independent scrutiny, replication studies allow us to distinguish robust phenomena from statistical artefacts or context-dependent effects. In software engineering, interest in replication has grown substantially, producing taxonomies, reporting guidelines, and systematic investigations of replication practice [6, 8, 17, 18]. Yet a fundamental problem
This work is licensed under a Creative Commons Attribution 4.0 International License. EASE ’26, Glasgow, United Kingdom © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2348-3/2026/06 https://doi.org/10.1145/3816483.3816559
RQ2: What problems arise from inconsistent or underspecified assessment criteria? RQ3: What principles should guide the assessment of replication outcomes in software engineering research? We conducted a search of replication studies published between 2021 and 2025, building on the baseline corpus of Cruz et al. [8], and identified 10 studies for inclusion. For each study, we extracted and analysed the criteria used to assess replication outcomes. We then go on to propose a set of preliminary assessment principles grounded in statistical and methodological considerations. Section 2 reviews related work. Section 3 describes the method. Section 4 presents findings on assessment practices and their inconsistencies. Section 5 proposes preliminary principles for replication assessment. Section 7 discusses threats to validity, while Section 8 concludes the paper.
EASE ’26, June 09–12, 2026, Glasgow, United Kingdom
2
Background
A persistent challenge in discussions of replication is terminological inconsistency. Following the ACM Artifact Review and Badging framework, as summarised by González-Barahona and Robles [17], we use repetition for a study with identical setup performed by the same team, reproduction for an identical setup performed by a different team, and replication for a study with a different setup performed by a different team. We note that Sjøberg et al. [28] distinguish internal replications, where the replication team includes original authors, from external replications, conducted by an entirely independent team. Gómez et al. [18] classify replications as literal, operational, or conceptual, depending on how much of the original protocol is preserved. These distinctions potentially may matter because the appropriate criteria for assessing outcomes may differ across replication types. Replication studies in software engineering go back to the mid 1990s, e.g., Daly et al. [10] in 1994 and Wood et al., [29] in 1997. Subsequently, there has been growth in the number of replications, although it remains a minority activity (see [8, 9] for reviews/mapping studies). Interestingly, what has received very little explicit attention is how do we determine the outcome of a replication? Specifically, how close should the results of the replication be to those of the original study in order to constitute ‘confirmation’? Should there be other outcomes than simply ‘confirmation’ and ‘disconfirmation’? If so, what are they and how should we determine our answers? In a recent mapping study of empirical research more generally, Heyard et al. [19] examined methodological literature focused on measuring, assessing, explaining, and predicting study reproducibility. They also included an investigation into what measures were actually used in 49 large-scale replication projects such as [20]. They identified a large number of distinct metrics: 50 in total. These include classic significance testing, direction of effect, effect size difference, the use of Bayes factors and various subjective assessment frameworks. The median number of metrics per replication was two, though this ranged 1-12. It was also noteworthy that almost 30% of the studies reported using subjective or narrative assessment of reproducibility. Furthermore, Nosek and Errington [24] argue that given a replication is a study for which any outcome constitutes diagnostic evidence about a prior claim, then the assessment criteria should be specified in advance and that both confirmatory and disconfirmatory outcomes should alter confidence in the original finding. In terms of software engineering, we are not aware of much explicit discussion of replication metrics and methods. Shepperd [26, 27] argued that prediction intervals should be used to evaluate replication outcomes (as opposed to confidence intervals) and showed via simulation the difficulties of meaningfully confirming underpowered studies. He, along with Santos et al. Santos et al., [25], also advocated using incremental meta-analysis to pool study results, as initially proposed for psychological studies by Braver et al. [5]. There has been some interest in using Bayesian methods to evaluate replications since this allows us to explicitly model our beliefs concerning prior studies, typically the original study, and also its updating with new data, the replication study(ies) [23].
Giuseppe Destefanis, Martin Shepperd, and Leila Yousefi
However, the application in software engineering that we are aware of is a tutorial from Erdogmus [14]. Fletcher [16] discusses in detail and critiques different approaches to evaluating replication study results and arguing that there are conceptual or formal reasons weaknesses with all approaches. A major objection is that by dichotomising the replication result into success or unsuccessful leads to the loss of much useful scientific knowledge. For this reason he concludes that pooling results for meta-analyses is the least problematic approach, a position which we the authors agree. None of these frameworks has been systematically applied to assess how replication outcomes are currently judged in software engineering, which is the gap this study addresses. In the next section we describe our systematic review of replication assessment methods and metrics in software engineering.
3
Method
This study employs a systematic approach to identify assessment practices in replication studies (RQ1), analyse problems arising from inconsistent criteria (RQ2), and synthesise principles for assessment (RQ3). Figure 1 gives an overview of the methodological workflow.
3.1
Study Identification
Search strategy. We searched SCOPUS in February 2026 using the following query, with year constraints pubyear > 2020 AND pubyear < 2026: "software engineering" AND title-abs-key("experiment*" OR "case stud*" OR "observational stud*" OR "pilot stud*" OR "survey") AND title-abs-key("repli*" OR "family of*")
This query follows the strategy used by Cruz et al. [8], who identified 137 replication studies published between 2013 and 2018 using the same terms, however, we focused on recent practise, i.e., research published in the past five years. Our search returned 62 candidate papers. Inclusion and exclusion criteria. We included studies that: (IC1) report at least one replication of an empirical study in software engineering; and (IC2) were published between 2021 and 2025. We excluded studies where: (EC1) the term replication refers to another context (e.g., data replication in distributed systems); (EC2) replication is mentioned only as future work; (EC3) the document is not a research paper (e.g., extended abstract, tutorial); or (EC4) the study is not in English. For each candidate paper, we verified that the study presents newly conducted analysis or data collection, and that empirical results are reported. Studies that revisit or comment on previously published work without new data were excluded. Screening and data extraction. One researcher screened titles, abstracts, and full texts of all 62 retrieved papers against the inclusion and exclusion criteria, yielding 10 included studies. For each included study, the researcher extracted: (1) bibliographic details; (2) the outcome type (e.g., continuous, binary, ordinal); (3) the criteria used to assess replication outcomes, coded as 13 binary method indicators; (4) the replication verdict as reported by the authors; and (5) whether the basis for that verdict was clearly stated.
The Replication Assessment Problem in Software Engineering
SCOPUS Search 2021–2025 (n = 62)
EASE ’26, June 09–12, 2026, Glasgow, United Kingdom
Screening
RQ1
IC/EC criteria Full-text review (n = 10)
Extract assessment criteria
RQ2 Analyse problems
RQ3 Synthesise assessment principles
Figure 1: Methodological workflow showing sequential stages from corpus identification to framework synthesis. Our 13 replication assessment methods (derived from [16]) are: null hypothesis significance testing for difference; equivalence testing; directional equivalence testing; confidence interval comparison; prediction intervals; effect size comparison; directional or trend assessment only; predictive accuracy comparison; expert or narrative judgement; pooling via meta-analysis; Bayesian methods; other formal methods; and methods too poorly described to classify. These codes were derived deductively from the literature and applied to each included study. Whether the original study appeared in the same paper as the replication was also recorded. The full coding guide and dataset are available in the replication package at this link: Figshare repository.
3.2
Analysis
To address RQ1, we report the frequency distribution of the 13 assessment method codes across the 10 included studies, and examine how many methods each study employed. To address RQ2, we identify cases where the basis for a verdict is absent or inconsistent with the methods used, and discuss the consequences for evidence accumulation. To address RQ3, we synthesise findings from RQ1 and RQ2 with the statistical and interpretive literature reviewed in Section 2, drawing in particular on Bonett [4] and Nosek and Errington [24], to derive a set of preliminary assessment principles. To characterise heterogeneity in assessment approaches, studies were additionally grouped thematically based on their dominant method.
4 Results 4.1 Dataset The replication package provides the full coding dataset as a XLXS file, where each row corresponds to a candidate paper from the SCOPUS search. Columns include bibliographic details (e.g., paper_id, authors, title, year, source, DOI), inclusion decisions (e.g., in_scope_se, presents_as_replication, included_final), and, for included studies, outcome characteristics and assessment data. The latter comprise 13 binary method indicators (m_*), the reported replication verdict, whether its basis was clearly stated, whether the original study appeared in the same paper and reviewer notes. Of the 62 retrieved papers, 52 were excluded and 10 retained for analysis.
4.2
RQ1: Assessment Criteria in Current Practice
Table 1 shows the assessment methods used across the 10 included studies, along with the reported verdict, outcome type, and whether the judgement basis was clearly stated. Expert judgement and pooling were the most frequently used methods, each appearing in four studies. Effect size comparison and
directional equivalence testing each appeared in two studies. Null hypothesis significance testing for difference, standard equivalence testing, predictive accuracy comparison, and other formal methods each appeared in one study. Four methods were not used by any included study: confidence interval comparison, prediction intervals, directional assessment only, and Bayesian methods. The number of methods used per study ranged from one (S016, S023, S026, S033, S049) to four (S036). Five studies relied on a single method to reach their verdict. Seven of the 10 studies involved continuous outcomes. The remaining three involved ordinal, binary, and mixed outcome types respectively. Five studies assessed a replication conducted within the same paper as the original study, and five assessed an externally published replication. The 10 studies can be grouped into four clusters according to their dominant assessment approach, which directly illustrates the heterogeneity in current practice. The first cluster comprises studies that use at least one formal statistical method for direct comparison: S036 applies null hypothesis significance testing, directional equivalence testing, effect size comparison, and pooling; S049 uses equivalence testing based on odds ratios derived from a side-by-side comparison of binary outcomes across programming languages; and S016 applies directional equivalence testing and explicitly acknowledges power limitations, reaching an inconclusive verdict. These three studies are the most methodologically explicit in the corpus. The second cluster comprises studies where expert or narrative judgement is the dominant or sole method: S006 compares correlation coefficients in detail but does not state a threshold for what difference would constitute failure; S017 reaches a successful verdict on the basis of expert judgement with no stated criteria; and S033 conducts seven external replications and reports a partial verdict through narrative assessment. The third cluster comprises studies where pooling is the primary analytical step: S011 pools data across three replications without a separate analysis of the first study, S023 conducts a baseline experiment followed by 11 internal replications with no explicit replication verdict, and S053 uses replication number as a blocking factor in a mixed model to justify aggregation. In all three cases, the question of whether individual replications are compatible with the original is not directly addressed. The fourth cluster contains a single study, S026, which assesses replication through predictive accuracy comparison, reflecting the tool-oriented nature of the research rather than a hypothesis about an effect.
4.3
RQ2: Problems Arising from Inconsistent Criteria
Of the 10 included studies, six reported a clearly stated basis for their verdict, three did not, and one was unclear. This means that in
EASE ’26, June 09–12, 2026, Glasgow, United Kingdom
Giuseppe Destefanis, Martin Shepperd, and Leila Yousefi
Table 1: Assessment methods and verdicts for included studies grouped by dominant assessment approach. Method columns use Yes/No coding. NHST-D: significance test for difference; NHST-E: equivalence testing; NHST-ES: directional equivalence; ES: effect size comparison; PA: predictive accuracy; EJ: expert judgement; PO: pooling; OF: other formal; UN: unclear. ID
Reference
Outcome
NHST-D
NHST-E
NHST-ES
ES
PA
EJ
PO
OF
UN
Verdict
Basis
Formal statistical comparison S016 [13] Mixed S036 [3] Continuous S049 [22] Binary
– Y –
– – Y
Y Y –
– Y –
– – –
– – –
– Y –
– – –
– – –
Inconclusive Successful Partial
Yes Yes Yes
Expert judgement S006 [21] Continuous S017 [2] Continuous S033 [12] Continuous
– – –
– – –
– – –
Y – –
– – –
Y Y Y
– – –
– – –
Y Y –
Partial Successful Partial
No No Yes
Pooling S011 [7] S023 [1] S053 [11]
Ordinal Continuous Continuous
– – –
– – –
– – –
– – –
– – –
Y – –
Y Y Y
– – Y
– Y –
Partial None Partial
Unclear No Yes
Predictive accuracy S026 [15] Continuous
–
–
–
–
Y
–
–
–
–
Failed
Yes
four cases the link between the evidence and the conclusion cannot easily be verified by a reader. The most common verdict was partial or mixed, reported by five studies. Two studies reported a successful replication, one reported failure, one reported an inconclusive outcome, and one provided no explicit judgement despite involving a baseline experiment followed by 11 internal replications (S023). In S023, the stated goal was to justify pooling all available data rather than to assess whether individual replications confirmed the original finding, illustrating how pooling can displace rather than support replication assessment. Three specific problems are evident in the data. Expert judgement is the sole assessment method in S033, and the only classifiable method in S017, where the basis for the verdict was not stated in sufficient detail to allow classification of all methods used. In S017, a verdict of successful replication is reached with no stated basis, and the methods used are coded as unclear. This means an independent reader cannot verify the basis for the verdict from the reported information alone. Second, the absence of stated thresholds renders verdicts unverifiable. S006 compares correlation coefficients across studies but does not state what magnitude of difference would constitute a failure to replicate, rendering the basis for the partial verdict unclear to an independent reader. S036, by contrast, uses four methods including effect size comparison and equivalence testing and provides a clear basis for its successful verdict, making it the most transparent study in the corpus. Third, pooling is used as a substitute for direct replication comparison in three studies (S011, S023, S053). In each case, the primary analytical goal is to aggregate data across replications rather than to assess whether individual replication results are compatible with the original. Pooling serves a legitimate purpose in families of experiments, but it does not answer the question of whether a given replication succeeded or failed. No study used prediction intervals, confidence interval comparison, or Bayesian methods to assess replication outcomes, despite
these being the approaches most directly suited to expressing effect compatibility under uncertainty [4].
5
Preliminary Principles for Replication Assessment
The findings of RQ1 and RQ2 identify four recurring problems: assessment criteria are frequently unstated, expert judegment is used without explicit decision rules, pooling displaces direct replication comparison, and methods well-suited to expressing effect compatibility are absent from current practice. Drawing on the literature reviewed in Section 2, we propose four preliminary principles to address these problems. Principle 1: State assessment criteria before examining results. Nosek and Errington [24] argue that declaring a study to be a replication entails a commitment to treating all outcomes as diagnostic evidence. This commitment requires that the criteria for success and failure are specified in advance, not derived post hoc from the observed results. Four of the 10 studies in our corpus either do not state or are unclear about the basis for their verdict. Pre-specification of criteria would make verdicts independently verifiable and reduce the risk that thresholds are selected to favour confirmatory conclusions. Principle 2: Express verdicts in terms of effect compatibility, not significance alone. Two studies (S016, S036) use directional equivalence testing, and one (S049) uses standard equivalence testing based on odds ratios; S036 additionally compares effect sizes. No study uses prediction intervals or confidence interval comparison. Bonett [4] provides a classification of replication outcomes based on whether the replication effect falls within a specified bandwidth of the original, distinguishing statistical replication evidence, nonreplication evidence, and inconclusive evidence. Many researchers [16, 26] have demonstrated that direct comparison of p-values is unreliable due to sampling error and instead recommended treating the original as one estimate among many. Verdicts grounded in
The Replication Assessment Problem in Software Engineering
effect compatibility, with explicitly stated bandwidths, are more informative and more reproducible than verdicts based on whether both studies reach the same significance threshold. Principle 3: Distinguish pooling from replication assessment. Three studies (S011, S023, S053) use pooling as the primary analytical step without first assessing whether individual replication results are compatible with the original. Pooling is appropriate once compatibility has been established or as a means of obtaining a more precise aggregate estimate, but it does not itself constitute evidence that a replication succeeded or failed. Studies that report only a pooled result without a direct comparison between the original and replication findings leave the assessment question unanswered. Principle 4: Report inconclusive outcomes explicitly. Only S016 reports an inconclusive verdict, the patterns observed in the corpus are consistent with several others having done so. S023 reaches no explicit judgement despite a substantial family of replications, and S017 reports a successful verdict with methods that could be better described to classify. Bonett [4] identifies inconclusiveness as a legitimate outcome category arising from wide intervals or insufficient power. Reporting inconclusive outcomes explicitly, rather than resolving ambiguity through narrative judgement, would produce a more accurate picture of the state of evidence in a research programme. Worked example. To illustrate how the four principles (P1-P4) apply in practice, consider S017 [2], which investigates consistency of gender bias in machine translation tools over time and reports a successful replication verdict. The methods used are coded as expert judgement and unclear, and the basis for the verdict is not stated. Under P1, the authors would be required to specify in advance what pattern of results across time points would constitute confirmation and what would constitute failure. Without this, the successful verdict cannot be independently verified from the reported information alone. Under P2, the verdict would need to be grounded in a quantitative comparison of effect magnitudes across time points, with an explicit statement of the bandwidth within which the effects would be considered compatible. Under P3, no pooling is involved, so this principle does not apply. Under P4, given that the methods are insufficiently described to classify and no basis is stated for the verdict, the assessment under the proposed principles would be inconclusive rather than successful. The outcome is not that the replication failed, but that the available reporting does not provide sufficient information to reach a verdict. This distinction matters for evidence accumulation: an inconclusive outcome signals that further evidence is needed, whereas a successful verdict without a stated basis does not provide sufficient information to calibrate confidence in the robustness of the original finding.
6
Discussion
The findings from the 10 included studies are consistent with the concern raised by Shepperd et al. [27] about systematic differences in how replication outcomes are judged. In their analysis, internal replications reported confirmatory evidence approximately eight times more frequently than external replications. Our data do not allow us to test this disparity directly, given the small corpus size,
EASE ’26, June 09–12, 2026, Glasgow, United Kingdom
but the prevalence of unstated criteria and expert judgement as the sole assessment method in several studies is consistent with the hypothesis that verdicts are influenced by factors other than the evidence itself. When the basis for a verdict is not stated, it is not possible to determine whether a confirmatory conclusion reflects genuine compatibility between original and replication results or reflects a permissive interpretation of ambiguous data. The dominance of expert judgement and pooling as assessment approaches reflects a broader tendency in SE replication research to treat families of experiments primarily as opportunities for data aggregation rather than for testing the robustness of specific claims. This is not without value: pooling increases statistical power and can yield more precise effect estimates. However, as Santos et al. [25] argue, the baseline experiment should be treated as one estimate among many rather than as a privileged target. Aggregation should nonetheless follow from, not substitute for, an assessment of whether individual replications are compatible with prior findings. The three studies in our corpus that use pooling as the primary step (S011, S023, S053) do not first establish this compatibility. The absence of prediction intervals, confidence interval comparison, and Bayesian methods from all 10 studies is notable given that these are the approaches most directly suited to the problem of assessing effect compatibility under uncertainty. Shepperd et al. [27] demonstrated the use of prediction intervals for this purpose, and Bonett [4] provides a formal classification framework built on effect bandwidth. The gap between available methodology and current practice suggests that adoption barriers exist, whether related to familiarity, software availability, or reporting norms, that are worth addressing in future work. The definition proposed by Nosek and Errington [24], in which a replication is a study for which any outcome constitutes diagnostic evidence about a prior claim, has a direct implication for assessment: if both confirmatory and disconfirmatory outcomes are to be taken seriously, then the criteria for distinguishing them must be stated in advance. Only six of the 10 studies in our corpus meet this condition. The remaining four either assert a verdict without a stated basis or leave the basis unclear, which means their outcomes cannot function as diagnostic evidence in the sense intended by Nosek and Errington. Adopting pre-specified criteria, as proposed in P1, would bring current SE replication practice closer to this standard.
7
Threats to Validity
Internal validity. All 62 candidate papers were screened and coded by a single researcher. This introduces a risk of coding error and subjective interpretation, particularly for fields that require judgement, such as judgement_basis_clear and the notes field. A multi-coder design with inter-rater reliability assessment would reduce this risk. The coding guide and dataset are provided in the replication package to allow independent verification. Coder confidence, recorded during extraction, was rated high for five studies (S006, S016, S033, S036, S053), medium for four (S011, S017, S023, S049), and low for one (S026), suggesting that coding uncertainty is concentrated in a small number of cases. External validity. The corpus consists of 10 studies retrieved from a single database (SCOPUS) covering 2021–2025. Studies indexed only in other databases, or published outside this period, are
EASE ’26, June 09–12, 2026, Glasgow, United Kingdom
not represented. The findings are sufficient to identify recurring patterns and illustrate specific problems, but they do not support claims about the prevalence of these problems across SE replication research more broadly. The search query follows Cruz et al. [8], which provides comparability with prior work, but the query may miss replication studies that do not use the terms replication or family of in their title, abstract, or keywords. Construct validity. The 13 method codes were derived deductively from the literature and applied to the reported content of each study. In some cases, the methods used by authors were insufficiently described to allow unambiguous classification, as reflected in the m_unclear code assigned to two studies. Where authors use informal or narrative language to describe their assessment approach, mapping this to a binary code inevitably involves interpretation.
8
Conclusions
Replication studies in software engineering are of limited value if their outcomes cannot be compared, aggregated, or interpreted on consistent terms. This descriptive study documented how replication outcomes are currently assessed in 10 empirical software engineering studies and identified four recurring problems: unstated assessment criteria, reliance on expert judgement without decision rules, use of pooling as a substitute for direct replication comparison, and absence of methods suited to expressing effect compatibility under uncertainty. Four preliminary principles were proposed to address these problems, grounded in the statistical and interpretive literature on replication assessment. The corpus is small and the findings are intended as a basis for further investigation rather than as definitive claims about SE replication practice. Future work should extend the search to a larger corpus, apply a multi-coder design, and test the proposed principles on a broader range of replication studies. The coding guide and dataset are provided in the replication package to support such extensions.
References [1] A. M. Aranda, O. Dieste, J. I. Panach, and N. Juristo. 2022. Effect of requirements analyst experience on elicitation effectiveness: A family of quasi-experiments. IEEE Transactions on Software Engineering 49, 4 (2022), 2088–2106. [2] Peter J Barclay and Ashkan Sami. 2024. Investigating markers and drivers of gender bias in machine translations. In 2024 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 455–464. [3] B. Bernardez, A. Durán, J. A. Parejo, N. Juristo, and A. Ruiz-Cortés. 2020. Effects of mindfulness on conceptual modeling performance: A series of experiments. IEEE Transactions on Software Engineering 48, 2 (2020), 432–452. [4] Douglas G Bonett. 2021. Design and analysis of replication studies. Organizational Research Methods 24, 3 (2021), 513–529. [5] S. Braver, F. Thoemmes, and R. Rosenthal. 2014. Continuously cumulating metaanalysis and replicability. Perspectives on Psychological Science 9, 3 (2014), 333– 342. [6] Jeffrey C. Carver. 2010. Towards Reporting Guidelines for Experimental Replications: A Proposal. [7] O. Cornejo, D. Briola, D. Micucci, D. Ginelli, L. Mariani, A. Santos Parrilla, and N. Juristo. 2024. A family of experiments about how developers perceive delayed system response time. Software Quality Journal 32, 2 (2024), 567–605. [8] Margarita Cruz, B. Bernardez, A. Duron, J. Galindo, and Antonio Ruiz-Cortes. 2020. Replication of Studies in Empirical Software Engineering: A Systematic Mapping Study, From 2013 to 2018. 26773-26791 pages. doi:10.1109/ACCESS.2019.2952191 [9] Fabio Q.B. da Silva, Marcos Suassuna, A. C. A. França, A. Grubb, Tatiana B. Gouveia, Cleviton V. F. Monteiro, and Igor Ebrahim dos Santos. 2014. Replication of empirical studies in software engineering research: a systematic mapping study. 501-557 pages. doi:10.1007/s10664-012-9227-7
Giuseppe Destefanis, Martin Shepperd, and Leila Yousefi
[10] Daly, Brooks, Roper, and Wood. 1994. Verification of results in software maintenance through external replication. In Proceedings of IEEE International Conference on Software Maintenance. IEEE, 50–57. [11] E. Díaz, J. I. Panach, S. Rueda, and D. Distante. 2021. A family of experiments to generate graphical user interfaces from BPMN models with stereotypes. Journal of Systems and Software 173 (2021), 110883. [12] D. A. Dos Santos, E. S. de Almeida, and I. Ahmed. 2022. Investigating replication challenges through multiple replications of an experiment. Information and Software Technology 147 (2022), 106870. [13] A. Durán Toro, P. Fernández, B. Bernárdez, N. Weinman, A. Akalın, and A. Fox. 2024. Exploring gender bias in remote pair programming among software engineering students: The twincode original study and first external replication. Empirical Software Engineering 29, 2 (2024), 40. [14] Hakan Erdogmus. 2022. Bayesian hypothesis testing illustrated: an introduction for software engineering researchers. Comput. Surveys 55, 6 (2022), 1–28. [15] Y. Fan, C. Arora, and C. Treude. 2023. Stop words for processing software engineering documents: Do they matter?. In 2023 IEEE/ACM 2nd International Workshop on Natural Language-Based Software Engineering (NLBSE). IEEE, 40–47. [16] Samuel C Fletcher. 2021. How (not) to measure replication. European Journal for Philosophy of Science 11, 2 (2021), 57. [17] Jesus M Gonzalez-Barahona and Gregorio Robles. 2023. Revisiting the reproducibility of empirical software engineering studies based on data retrieved from development repositories. Information and Software Technology 164 (2023), 107318. [18] Omar S. Gómez, Natalia Juristo Juzgado, and S. Vegas. 2014. Understanding replication of experiments in software engineering: A classification. 1033-1048 pages. doi:10.1016/j.infsof.2014.04.004 [19] R. Heyard, S. Pawel, J. Frese, B. Voelkl, H. Würbel, S. McCann, L. Held, K. Wever, H. Hartmann, L. Townsin, et al. 2025. A scoping review on metrics to quantify reproducibility: a multitude of questions leads to a multitude of metrics. Royal Society Open Science 12, 7 (2025). [20] R. Klein et al. 2018. Many Labs 2: Investigating variation in replicability across samples and settings. Advances in Methods and Practices in Psychological Science 1, 4 (2018), 443–490. [21] U. A. Koana, Q. H. Le, S. Rahman, C. Carlson, F. Chew, and M. Nayebi. 2024. Examining ownership models in software teams: A systematic literature review and a replication study. Empirical Software Engineering 29, 6 (2024), 155. [22] Chris Langhout and Maurício Aniche. 2021. Atoms of confusion in Java. In 2021 IEEE/ACM 29th International Conference on Program Comprehension (ICPC). 25–35. [23] A. Ly, A. Etz, M. Marsman, and E. Wagenmakers. 2019. Replication Bayes factors from evidence updating. Behavior Research Methods 51, 6 (2019), 2498–2508. [24] Brian A Nosek and Timothy M Errington. 2020. What is replication? PLoS Biology 18, 3 (2020), e3000691. [25] Adrian Santos, Sira Vegas, Markku Oivo, and Natalia Juristo. 2021. Comparing the results of replications in software engineering. Empirical Software Engineering 26, 2 (2021), 13. [26] M. Shepperd. 2018. Replication Studies Considered Harmful. 73-76 pages. doi:10. 1145/3183399.3183423 [27] M. Shepperd, N. Ajienka, and S. Counsell. 2018. The role and value of replication in empirical software engineering results. 120-132 pages. doi:10.1016/J.INFSOF. 2018.01.006 [28] Dag I.K. Sjøberg, J. Hannay, Ove Hansen, V. Kampenes, Amela Karahasanovic, Nils-Kristian Liborg, and Anette C. Rekdal. 2005. A survey of controlled experiments in software engineering. 733-753 pages. doi:10.1109/TSE.2005.97 [29] M. Wood, M. Roper, A. Brooks, and J. Miller. 1997. Comparing and combining software defect detection techniques: a replicated empirical study. ACM SIGSOFT Software Engineering Notes 22, 6 (1997), 262–277.