An Empirical Study on the Impact of Change Granularity in Refactoring Detection Lei Chena , Shinpei Hayashia a School of Computing, Institute of Science Tokyo, Ookayama 2–12–1, Meguro-ku, Tokyo, 152–8550, Japan
Abstract
arXiv:2609.07482v1 [cs.SE] 7 Sep 2026
Detecting refactorings in commit history is essential to improve comprehension to code changes on code reviews, and to provide valuable information for empirical studies on software evolution. Techniques have been proposed to accurately detect refactorings on the granularity of a single commit. However, refactorings can be made over multiple commits because of their complexity or other practical development problems, which cause detecting on only the granularity of a single commit not enough. We observe that some refactorings can only be detected in coarser granularity, i.e., changes conducted over multiple commits, or in the granularity of a single commit but not in coarse-grained. We call these types of refactorings as coarse-grained refactorings (CGRs) and ephemeral refactorings (EPRs). We investigated the features and causes of CGRs and EPRs through an empirical study of 32 open-source Java projects and found that both commonly occur during development. In addition, we found that refactoring types related to splitting or merging classes and packages, as well as those involving modifications to the inheritance structure, tend to be CGRs, and types targeting small objects such as variables and attributes, and refactorings with context-sensitive detection criteria tend to be EPRs. The causes of CGRs and EPRs are analyzed and categorized, and the relationships between the commit messages of CGRs and themselves are also assessed. We found that about 20% of commit messages explicitly suggest the existence of CGRs. We suggest that CGRs and EPRs be valued in refactoring research and that detectors be extended to identify CGRs. Keywords: software evolution, refactoring, refactoring detection, commit message
1. Introduction Refactoring mining within the commit history benefits both software developers and researchers. Refactoring is the process of improving the internal structure of code without changing its external behavior [1]. For developers, it facilitates the understanding of code changes [2] and enhance the effectiveness of code reviews [3]. In terms of researchers, it can provide valuable information for empirical studies on software evolution, such as the benefits of refactorings [4], the relationship between bugs and refactorings [5]. To mine refactorings, recent studies proposed refactoring detectors that detect refactorings by comparing two source code snapshots [2, 6, 7, 8, 9, 10, 11, 12]. While traditional approaches focus on detecting refactorings between releases [6, 7, 8, 9], modern detectors such as RefDiff [10, 11] and RefactoringMiner [2, 12] analyze individual commits, which means that two snapshots before and after a single commit are compared. These methods have achieved high accuracy in detecting refactoring within commits [13]. However, refactorings may be performed over multiple commits. For example, when moving a method from one class to another, some developers choose to first copy the implementation of the method to the target-of-move class in the first commit. After confirming that users will not be affected, the original implementation is removed in the next commit. In this case, the Move Method refactoring could not be detected from either of the two commits using commit-level refactoring detectors, but it becomes detectable from a coarse-grained granularity that Preprint submitted to Journal of Systems and Software
spans across both commits. By investigating the impact of changing the refactoring detection granularity from a single commit to coarser-grained, we are able to gain a deeper insight into the refactorings similar to the above. We define a refactoring whose code changes are recorded across multiple commits, rather than confined in a single commit as coarse-grained refactoring (CGR). Conversely, a refactoring whose code changes are recorded within a single commit and where modifications in adjacent commits would disrupt their integrity are referred to as ephemeral refactoring (EPR). We use the term ephemeral to emphasize its brief lifespan, as these refactorings are confined to a narrow historical scope of commits and will disappear in a broader scope. An example of EPR is that a refactoring is initially applied and then reverted due to being considered an inappropriate change. The existence of CGRs and EPRs suggests that relying solely on a single specific detection granularity is not enough, as CGRs may be overlooked, and EPRs may require analysis across multiple commits to be detected. This limitation may lead to incomplete or incorrect conclusions in empirical studies, such as refactoring type and software evolution analysis. Also, uncovering those refactorings can help developers gain a better understanding of code evolution. To investigate the features and causes of CGRs and EPRs, we conduct an empirical study on open-source Java repositories. We regard the original commits recorded in the commit history as fine-grained and name them fine-grained commits (FGCs). To change the granularity of commits, multiple FGCs are squashed into one to form a coarse-grained com-
mit (CGC). In this way, the code changes from multiple FGCs are combined into a CGC. The number of FGCs squashed into one CGC is referred to as granularity level. Refactoring detection is conducted on both FGCs and CGCs using the state-ofthe-art tool RefactoringMiner [2, 12] to extract CGRs and EPRs for analysis. The main contributions of this paper are shown below: • We defined the concept of CGR and EPR and proposed a technique to detect them from the commit history.
The remainder of this paper is organized as follows. Section 2 explains the definition of CGR and EPR. Section 3 introduces the background of this study. In Section 4, we explain the matching mechanism of refactorings at different granularities that we designed. Section 5 introduces the overview of our study. We extract CGRs and EPRs on the dataset and analyze and discuss five research questions in Section 6. The possible threats to validity are discussed in Section 7, and the implications are discussed in Section 8. Finally, we conclude our work and state our plans for future work in Section 9.
• An empirical study is conducted on a dataset of 32 opensource Java repositories.
2. Coarse-Grained and Ephemeral Refactorings
• The features and causes of CGRs and EPRs are analyzed.
2.1. Coarse-Grained Refactoring
• We investigated the relationship between developers’ intentions deduced from commit messages and CGRs. We discussed the implications of CGRs and EPRs for researchers and practitioners.
Definition 1 (Coarse-grained refactoring). A refactoring is referred as a CGR if its code changes are recorded across multiple commits rather than a single commit. An example of CGR found in open source repository mbassador1 is shown in Figure 1. This figure depicts a sample history consisting of two commits, where commit 2ae0e5f is the parent of commit 9ce3ceb. The intention of the developer, as expressed by these two commits, is to decompose the source file Mbassador.java, which contains multiple top-level classes, into multiple source files to ensure that each file contains only one top-level class, i.e., applying Move Class refactoring. In the firstly proposed commit 2ae0e5f, the developer copied the implementation of class FilteredAsynchronousSubscription (FAS hereinafter) in file Mbassador.java to a newly created file FAS.java. Then, the developer removed the implementation of that class from the source file Mbassador.java in the secondly proposed commit. Overall, she/he moved a class from Mbassador.java to a new source file. The detection based on either of the single commits shown in Figure 1 cannot reveal this kind of refactoring because each commit contains only part of the code changes for detecting Move Class refactoring. However, if we consider the whole changes conducted across these two commits, the refactoring can be correctly detected.
The findings of this paper are as follows: • CGRs and EPRs are common in development, and on average, per 100 refactorings conducted, there are also 5.71 CGRs and 9.89 EPRs be performed. • Refactoring types related to splitting or merging classes and packages, as well as those involving modifications to the inheritance structure, tend to be CGRs while types related to operations on small objects such as variables, attributes, and parameters tend to be associated with EPRs. • The cause of CGR can be categorized into two types: Generation and Combination according to their composition. While the cause of EPR can be divided as Disruption, Absorption, and Covered. • We found that 20% commit messages explicitly express the existence of CGRs, and the characteristics of those commit messages are summarized and categorized into four types. This paper is an extension of work reported originally in the proceedings of the 30th IEEE/ACM International Conference on Program Comprehension [14]. The main improvements are the following:
2.2. Ephemeral Refactoring Definition 2 (Ephemeral refactoring). A refactoring is referred to as an EPR if its code changes are recorded in a single commit, but the code changes in its adjacent commits will break the detection of that refactoring.
• A more delicate matching scheme is designed to mine more previously undiscovered CGRs. • Proposed the definition of EPRs and also the technique to mine them.
An example of EPR is found and shown in Figure 2. This figure shows two commits extracted from the commit history of an open source repository RoboBinding2 , where commit 0f52164 is the parent commit of commit eb68c47. Before the first proposed commit 0f52164, an instance of the ProcessingContext (PC hereinafter) class is created using
• Conducted an experiment on a larger-scale dataset that contains 32 open-source Java repositories. • Analyzed the features from more aspects for CGRs and the features for EPRs. • Investigate the causes of EPRs.
1 https://github.com/bennidi/mbassador/commit/9ce3ceb
• Implications for both researchers and practitioners.
2 https://github.com/RoboBinding/RoboBinding/commit/eb68c47
2
Mbassador.java
Mbassador.java
Mbassador
Mbassador.java Mbassador
FAS
FAS
2ae0e5f
FAS
FAS.java
copy
FAS.java FAS
Mbassador
FAS
1
9ce3ceb
No refactoring detected
2
Commit history of Mbassador
No refactoring detected
Refactoring detected: Move Class
Figure 1: Example of a CGR found in mbassador. E e = m2( ); T t = m4(m1( ), e); PC context = new PC( t, e, … )
PC context = new PC( m1( ), m2( ), … )
0f52164
1
eb68c47
Refactoring detected: Extract Variable
RC rc = new RC( m1( ), m2( ), … )
2
Refactoring detected: Inline Variable
Commit history of RoboBinding
No Refactoring detected Figure 2: Example of an EPR found in RoboBinding.
of Figure 2, and as a result, Extract Variable refactorings are not detected at the coarse level. Similarly, the Inline Variable refactorings are also not detected at the coarse level. Since the refactorings Extract Variable and Inline Variable in this example were detected in each commit but were unable to be detected if we look at the whole change of the two commits, they are regarded as EPRs. The detection on each of the two commits can find the two refactorings. However, the fact that the overall code change does not contain those refactorings can only be revealed after we inspect the code changes through these two commits.
its constructor, which takes three method invocation expressions as arguments. In the first proposed commit 0f52164, the first two method calls passed to the constructor are extracted into separate variables: elements (e hereinafter) and types (t hereinafter). These variables are then used as arguments when instantiating the PC object. This commit introduces two instances of the Extract Variable refactoring. Note that one is not a pure refactoring because a wrapper m4() is introduced together with the extraction of the variable t. In the subsequent commit eb68c47, the object type is changed from PC to RoundContext (hereinafter RC). Additionally, the previously extracted variables (e and t) are inlined back into the constructor arguments. In this commit, two Inline Variable refactorings are detected.
Note that although the term coarse-grained refactoring has been used in other contexts, its meaning differs from ours. For example, independently of our work, Bibiano et al. [15] used the same term to describe refactorings that target code elements larger than a method. Similarly, Mongiovi et al. [16] introduced the notion of small-grained transformation, referring to code transformation operations for low-level code elements such as “add method” or “remove field”, which are not necessarily refactorings and are not analyzed in relation to commit-
The code changes over these two commits are replacing the class PC with the class RC. Because the detection of Extract Variable refactorings relies on the condition that a new variable is defined, and there is a similarity between its initialization part and some of the removed expressions. However, the above condition does not hold when we look at the bottom part 3
based composition. In contrast, the CGRs and EPRs introduced in this paper refer to newly defined concepts that focus on how refactorings are recorded across commits in a version control system. The notion of CGR was first proposed in our earlier work [14], and we continue to use that definition here, extending it by introducing EPRs within the same context.
code to build a model consisting of high-level source code entity information such as types, methods, and fields. The second phase is to find relationships between elements in the models before and after the code change. By matching the relationships with predefined rules, refactoring entities can be detected. Later, RefDiff is extended to RefDiff 2.0 [11] to support detecting refactorings in multiple programming languages instead of only in Java. This version utilizes the Code Structure Tree (CST), a language-agnostic representation of the source code. The CSTs are matched to analyze the relationship to identify refactoring entities. Tsantalis et al. introduced RefactoringMiner [2], marking the first refactoring mining tool that operates without the need for any code similarity thresholds. This tool detects refactorings through Abstract Syntax Tree (AST)-based statements matching with predefined rules. The code entities in the two revisions are matched in top-down order based on text similarity. After that, a set of predefined rules are used to identify refactorings. The RefactoringMiner 2.0 [12] was proposed later, which extended the previous work by supporting the detection of lowlevel in-method refactoring types. In addition, they created a refactoring oracle, which contains 7,226 true instances found in 536 commits from 185 open-source projects. While current refactoring detection studies concentrate on either the release level or the commit level, there is a gap at the intermediate level; refactorings applied across multiple commits remain undetected.
3. Related Work 3.1. Refactoring Detector Detecting refactoring instances performed in software projects can help practitioners be aware of whether and how to track software that requires maintenance and help researchers collect information for empirical studies about refactorings and their effects to build support techniques [17]. This paper employs a state-of-the-art refactoring detector to mine refactorings from the commit history of software projects. The input of refactoring detectors comprises two snapshots of the source code. Based on the main focus of the research targets, the proposed studies can be divided into two types: release level and commit level. 3.1.1. Release Level Studies of the release level category concentrate on detecting refactorings by utilizing two release versions as the designated snapshots. Although tools developed in these studies can compare two snapshots of source code to identify refactorings applied between them, they are mainly applied between release versions in their studies. The RefactoringCrawler [6] identifies similar pairs of entities and leverages references among these entities to determine refactoring entities to determine refactoring instances. This is grounded in the observation that most refactorings include operations of repartitioning source files, which results in source text between different versions of a component being similar. The Shingles encoding [18] from the field of Information Retrieval is adopted to find similar fragments in source files as refactoring candidates. Then, the refactoring instances can be detected by analyzing the references among the source code entities in each of the refactoring candidates. The Ref-Finder [7, 8] detects refactoring instances of 63 refactoring patterns based on predefined rules. The detection process involves extracting code elements, their structural dependencies, and the contents of these elements. Subsequently, differences in elements between the two versions are calculated. Finally, refactoring instances are determined based on these calculated differences and predefined rules.
3.2. Tangled Changes and Composite Refactorings In version control systems, a tangled change refers to the practice of committing unrelated or loosely related code changes in a single commit, a phenomenon that occurs frequently during development [19]. Tangled changes can hinder the detection of refactorings due to a mismatch between the recorded granularity at which detection is performed and the semantic granularity at which developers actually make changes. Our study further investigates this mismatch, with a particular focus on changes that span multiple commits. The tangled change is difficult to understand and revert. To help developers untangle the tangled changes, Matsuda et al. [20] presented a method that restructures changes by tracking source code operations, which encompass refactoring actions, and organizes them hierarchically according to their types provided by an integrated development environment. This hierarchy empowers developers to effortlessly adjust the level of change granularity and obtain the resultant changes based on their chosen granularity configuration. Sothornprapakorn et al. [21] propose a visualization approach to support untangling. The proposed approach organizes tangled changes into a tree structure, leveraging refactoring detection and change relevance calculations to create multiple change groupings within the structure. The authors conducted an experiment with industrial developers and proved that the tool could help developers understand and decompose the change. Being a special type of code change, it often occurs in practice that refactorings are mixed with other refactorings or code changes in a single commit [22, 23, 24]. Sometimes, such
3.1.2. Commit Level This category of refactoring detectors specializes in identifying refactorings within code commits. Silva et al. proposed RefDiff [10], which detects refactorings through two phases: source code analysis and relationship analysis. In the first phase, the tool parses and analyzes the source 4
mixed refactorings in a single commit are applied to optimize irrelevant code, while in other instances, they serve as a step towards completing a complicated optimization. The refactoring detection applied on a single commit granularity, which is the code change recorded granularity, is in discordance with the change granularity, which might be a single refactoring or multiple refactorings distributed over multiple commits. Our work also investigates the impact of the latter. Görg et al. [25] proposed a term of purity of refactoring: a change that contains mixed refactoring and non-refactoring modifications is applied at the same location at the same time as impure refactoring. The mechanism of EPR in our study aligns with this definition. To detect impure refactorings, a graph search algorithm-based technique is proposed and evaluated on an Apache repository, demonstrating its feasibility [26]. The refactorings mixed with interrelated refactorings are referred to as composite refactorings. Sousa et al. [27] provide the first formal and unambiguous definition for composite refactorings and provide two heuristics to reveal the characteristics of composite refactorings. They found that composite refactorings have a considerable effect on code smells: introducing or removing smells, which is contradictory with previous studies [28, 29, 30, 31]. Bibiano et al. [32] present a composite refactorings catalog, complete with details about side effects, and recommendations on how to decrease them. In addition, they enumerate some scenarios where each recommendation can be applied to remove the code smell. Some composite refactorings are composed of co-occurring refactorings. Saika et al. [33] investigate the frequency of co-occurred refactorings by analyzing usage data of the Integrated Development Environment Eclipse. They conclude that Refactoring pairs (Move, Rename), (Rename, Rename) and (Extract, Move) appear more frequently than the others. To further clarify the distinctions among tangled changes, composite refactorings, CGR, and EPR, particularly in terms of granularity and composition, we compare them from two perspectives: granularity and composition. From the perspective of granularity, tangled changes [19] and EPR are confined to a single commit, whereas CGRs span across multiple commits. Composite refactorings are not bound to a specific commit granularity and may occur within a single commit or across multiple commits [27]. From the perspective of composition, tangled changes consist of changes and unrelated changes, while EPR consists of a single refactoring. CGR, in contrast, is composed of multiple refactorings or non-behaviorpreserved changes as part of refactoring changes distributed across commits.
r2
r1
FGC
fine-grained commit history
Original repository (fine-grained commits) Squash
Repository transformation
[unmatched] Ephemeral refactoring (EPR)
Squash
[matched]
coarse-grained commit history
Squashed repository (coarse-grained commits)
CGC
r'1 r3
[unmatched] Coarse-grained refactoring (CGR)
Figure 3: Detection mechanism.
commits stored in the original repository as fine-grained in our context. The FGCs in the original repository are squashed into CGCs through repository transformation. For each CGC, the finegrained ones squashed into it are collectively referred to as a squash unit. We tried multiple different strategies to generate the squashed repository so that various kinds of CGCs are examined. The details are explained in Section 5. Then, an existing state-of-the-art refactoring detector is applied to both FGCs and CGCs to detect refactorings. Refactorings detected from FGCs are referred to as original refactorings since those refactorings are detected from the original commit history. Note that EPRs are a subset of the original refactorings. While original refactorings include all refactorings detectable at the original commit granularity, EPRs are those that become undetectable once commits are squashed into CGCs. In the figure, refactorings r1 and r2 are detected from the two commits in the original repository, whereas r1′ and r3 are detected in the commit squashed from those two commits. By matching the original refactorings and refactorings detected from CGCs, we can classify the detected refactorings into the following three types. • Refactorings that are not detectable from FGCs but from CGCs are identified as CGRs. For example, since r3 is detected from the CGC but not from the fine-grained ones, it is regarded as a CGR. • Refactorings that are not detectable from CGCs but from FGCs are identified as EPRs. For example, since r2 is only detected from the FGCs but not from the CGC, it is regarded as an EPR.
4. Detecting Coarse-Grained and Ephemeral Refactorings
• The remaining refactorings. Because r1 and r1′ are matched as the same refactoring, they are considered neither CGR nor EPR.
4.1. Basic Idea A basic idea to detect CGRs and EPRs is to generate CGCs consisting of the changes in multiple commits in the original repository and to compare the refactoring detection results with the original detection results. An overview of the mechanism to detect CGRs and EPRs is shown in Figure 3. We regard the
4.2. Matching Scheme To identify CGRs and EPRs, we compare and match the refactorings detected from the original commit history and 5
2:2 public class A { :3 + private String name; :4 + public void bar() {} 3:5 public void hoo() {} 4:6 }
c1
2:2 public class A { private String name; 3:3 public void bar() {} 4:4 Rename 5: public void hoo() {} Method public void foo() {} :5 + 6:6 }
c2
Squash
Method detected from FGCs, and they should be matched as
the same refactoring. However, their locations differ: the one in the squashed commit is detected to be applied on Line 3 instead of Line 5 because of the code inserted in the commit c1 . Therefore, checking the exact equivalence of the location where the refactorings are applied is insufficient to match two refactorings. We trace the refactoring location in order to match refactorings. We name the refactoring targeted source code snippets at the version before applying the refactoring as the refactoring targets. For example, the refactoring target for Rename Method is the line of the code containing the declaration of the method undergoing the renaming process. The trace of refactoring is to find out the commit that the refactoring target is last modified and the line number of the first line of that code in the found commit. The first line here refers to the line with the smallest line number in the location information of the refactoring detection result for the refactoring target. In case of refactoring detection result involves multiple code elements, We choose the first line of the first code element as it is usually the primary element in the detection result. We use the first line of refactoring target to trace because it is the class name, method header, or code statement, which can represent the code block of being refactored. Following the techniques employed in violation tracing studies [34, 35], we use the line origin, the line of the commit when the target line was introduced, for the key factor of the signature. The first line of the refactoring target is typically the signature of the refactored code element, which effectively represents the element. We apply git-blame on the first line of the refactoring target to specify the origin commit and the line at the origin commit. Although tracing techniques such as CodeTracker [36, 37], which focus on Java code and are refactoringaware with high tracing accuracy, are available, we chose not to adopt them. This decision was based on two considerations. First, extending the tools to support all types of refactorings investigated in this study would require additional effort, which is beyond the scope of our work. Second, we wanted to design the detection framework to be as language-agnostic as possible so that interested followers could easily replicate it in other programming languages, whereas techniques such as CodeTracker are specific to Java. As a result, we employed git-blame, a more general technique that can be applied across a wide range of programming languages. Two refactorings are compared to see if they reach the same line at the same origin commit. The details are introduced in Section 5.7. Note that before the comparison and matching, we have a preprocessing step to remove the comments from the source code that may interfere with identification accuracy. Some refactoring detection results regard the first line of a class or a method as the Javadocs or comments above the code, instead of the line declaring the class or method. Therefore, the changes to comments, such as the addition or removal of Javadoc comments, may influence the refactoring signature. The situation where the same refactoring is no longer matched due to modifications irrelevant to the behavior of the source code is not desired. Therefore, we excluded such comments in advance as
fine-grained commit history
CGC
2:2 public class A { :3 + private String name; :4 + public void bar() {} Rename 3: - public void hoo() {} Method :5 + public void foo() {} 4:6 }
Figure 4: Example of location change because of granularity change.
those detected from the granularity-changed commit history. To accurately match two refactorings, it is essential to establish a distinctive signature for each refactoring. This means that two refactoring instances obtained from different commits are considered identical if their signature is identical. We use RefactoringMiner [12], which is the state-of-the-art and widely used refactoring detector, as an example to show its detection result of one refactoring: • 1) refactoring type, • 2) description of how this refactoring is conducted, and • 3) location information of the refactoring target elements, which includes the file paths and line numbers of the refactoring target in both the pre-refactoring and postrefactoring states as well as its role on the refactoring. The 1) refactoring type describes the types of refactorings, such as Move Method or Inline Class. The 2) description textually explains the conducted refactoring, consisting of the type of the refactoring being applied and its details such as the type and name of the refactored code elements or the file path of the class containing the refactored code elements. Unfortunately, comparison by types or descriptions is insufficient to distinguish small refactorings, in cases where two refactorings of the same type are conducted together or code elements with the same type and same name but are located in different methods in the same file are refactored together. The richest information source is the 3) location information, which can help to solve the above issue. By using both the file path and the line number, one code element can be uniquely located. In this approach, the refactorings are characterized by the role and location of the refactored code elements. However, changing granularity may impact the location (line number) where refactoring is performed, which in turn affects the matching process. An example is shown in Figure 4. The code change in the FGC c2 is to rename the method from hoo to foo. It is detected as a Rename Method, and it was applied to the code element in Line 5. In the previous commit c1 , two lines of code are added above the method declaration of hoo(). In the squashed CGC of the above two commits, a refactoring Rename Method is detected. It is equal to the Rename 6
a pre-processing step to eliminate their influence. The details are introduced in Section 5.2. Compared with our previous work [14], we enhanced the matching mechanism by designing a more delicate refactoring signature. In our previous work, only the refactoring type is used for matching, which may cause false negatives where some CGRs are omitted because of having the same type with refactorings in the original history. By using both the traced location and refactoring type, we are able to uncover more CGRs and EPRs with higher accuracy.
5.4. Determining Squash Units A squash unit u (⊆ C) is a set of multiple adjacent FGCs that are squashed into a single CGC. Here, if a commit is the parent or child of another commit, these two commits are considered adjacent. The adjacent commits are shown as circles next to each other in Figure 5. Different strategies labeled S oℓ (for appropriate values of o and ℓ) are used to extract squash units from straight commit sequences. Here, the granularity level ℓ (≥ 1) specifies the size of the squash units, and straight commit sequences are divided into multiple squash units of the specified size. Because each unit is squashed into one CGC, this level also expresses the granularity level of the CGCs to be generated. The granularity level ℓ = 1 exactly produces original FGCs. The offset 0 ≤ o ≤ ℓ − 1 is the number of commits to be skipped from the beginning of the given straight commit sequence when extracting the squash units to adjust which commits will be merged. For example, the commit c11 in Figure 5 is squashed together with c10 when strategy S 02 is used, whereas it is squashed together with c12 when strategy S 12 is used. To be more specific, for the straight commit sequence h = {c11 , c12 , . . . , c1n }, containing n consecutive FGCs, the squash units extracted from h with the given strategy S oℓ (ℓ ≥ 2) can be presented as:
5. Methodology 5.1. Overview The overview of our study procedure is shown in Figure 5. Our procedure can be divided into two phases: (a) Repository Transformation and (b) Detection and Match. In the repository transformation phase, the input is the raw Git-based commit history extracted from a repository. Through preprocessing, which removes the Javadocs and comments, we can obtain a commit history that excludes changes related with comments. Searched from commit history, we can obtain straight commit sequences, each containing consecutive commits on the same branch. The squash units can be determined on the straight commit sequences, and each of them is squashed into a CGC. In the detection and match phase, for each pair of CGC and the FGCs that are squashed into that CGC, refactorings are detected and matched.
unitsS oℓ (h) = {uo+kℓ | 0 ≤ k ≤ (n − ℓ − o)/ℓ} where u x = {c1i | x ≤ i < x+ℓ} is the squash unit beginning from the commit c1x determined by strategy S oℓ . We regard the size of a squash unit as its granularity. We may write the granularity level as the superscript of the unit, e.g., uℓ (ℓ = |u|). The difference of the generated commits according to different strategies is illustrated in Figure 7. Different squash units are created according to the strategy with a specific granularity level and offset, and different squashed commits are generated from these squash units. By covering all possible offsets at each granularity level, all possible squash units could be enumerated. For example, at the granularity level of ℓ = 3, three different offsets o ∈ {0, 1, 2} are used.
5.2. Preprocessing As introduced in Section 4.2, some refactoring detectors regard the first line of a class method as the Javadocs or comments above the code. We remove those comments in the preprocessing phase to prevent the code changes in comments from affecting the tracing results. An example showing the above procedures is depicted in Figure 6. The refactoring is to move the method methodA() from class A to class B. Before the preprocessing, the first line of the refactoring target is regarded as beginning at Line 3, which is the Javadoc comment for method methodA(). After the preprocessing, the comment is removed, and the first line of the refactored code snippet specifies the method signature of the method being moved.
5.5. Changing Granularity To change the granularity of commit history, we squash multiple commits into a single one to integrate the code changes distributed in multiple revisions. The tool git-stein3 [38], which can squash the FGCs in all the squash units extracted from the repository into CGCs, is used to perform the granularity transformation. The output of the tool is another repository whose commit history is composed of CGCs. For each squash unit uℓ determined in Section 5.4, cℓ = sq(uℓ ) represents squashing all the ℓ commits in uℓ into a single CGC cℓ . Similarly to squash units, we may specify the granularity level as the superscript of the commit. A squashed commit of a certain granularity level can cover several commits of finer granularity levels. In case where a
5.3. Extracting Straight Commit Sequences From the preprocessed Git-based commit history H, which is a set of FGCs C (⊆ C), where C is the universal set of commits, straight commit sequences h ⊆ C can be extracted. Each straight commit sequence h consists of consecutive FGCs on the same branch that exclude merge commits, which have more than one parent, and branch sources, which have more than one child. Merge commits are excluded to avoid duplicate detection of refactoring in the later phase, and branch sources are excluded for simplicity when extracting squash units. Here, seqs(H) (⊆ 2C ) denotes the set of all the possible straight commit sequences extracted from H.
3 https://github.com/sh5i/git-stein
7
Squash
𝑆"&
Detect
Trace
Raw refactorings
offset=4
Extract
Refactorings
Match
Trace Refactorings
…
…
…
…
Raw refactorings
squash unit u
𝑆#$
Latest
Latest
Squash
𝑐#" 𝑐"" 𝑐!"
Detect
offset=1
𝑐"#
Trace
Refactorings Extract
Search
…
Preprocess
𝑐#"
𝑆%$ Raw commit history
Commit history
Trace
Raw refactorings
Extract
Squash
𝑐"" 𝑐!"
Straight commit sequences
Detect
𝑐!#
offset=0 Squash units determined sequences
Match
Refactorings
Trace
refactorings Extract
Refactorings
Trace
Raw refactorings
…
…
𝑆%#
Coarsegrained commits
Refactorings
Match
Refactorings
Detect Raw refactorings
(a) Repository Transformation
Straight commit sequences
Straight commit sequences
(b) Detection and Match
Figure 5: Study overview.
commit cℓ is generated by squashing a unit uℓ = {c1i , . . . , c1i+ℓ−1 }, its sub commits and super commits can be defined as follows:
5.6. Detecting Refactorings Refactoring detection is conducted on both the commit history before and after the granularity transformation. The refactorings detected from a commit c can be defined as:
sub(cℓ ) = {sq(u) | u ⊊ uℓ } = {sq({c1j , . . . , c1j+ℓ′ −1 }) | 1 ≤ ℓ′ < ℓ ∧ i ≤ j < ℓ − ℓ′ },
R = ref(c).
sup(cℓ ) = {sq(u) | u ⊋ uℓ } = {sq({c1j , . . . , c1j+ℓ′ −1 }) | ℓ < ℓ′ ∧ j ≤ i ∧ i + ℓ ≤ j + ℓ′ }.
Here, R ⊆ R is a set of refactorings, where R denotes the universal set of refactorings. For each refactoring r ∈ R, r.type denotes its type, r.elementSig denotes the refactoring target signature, and r.location denotes the traced location whose detail is explained in Section 5.7. We consider the first element in the pre-refactoring version of the code snippet in the detection result of RefactoringMiner, to be the primary element of the refactoring and use its location as the target for tracing, i.e., refactoring targets as introduced in Section 4.2.
For example, c24 and c25 in Figure 7 are sub-commits of c34 . Also, the original FGCs (c14 , c15 , c16 ) are also regarded as sub-commits of c34 . All the squashed commits at granularity level ℓ extracted from the commit history H can be represented as: [ [ ℓ (h) . C ℓ (H) = u ∈ units sq(u) So 0≤o≤ℓ−1 h∈seqs(H)
5.7. Refactoring Location Tracing and Match As discussed in Section 4.2, the match of refactorings detected in the CGCs and those in FGCs requires location tracing to mitigate the change in refactoring location brought by granularity change. The overview of the tracing procedure is depicted in Figure 8. In the commit history shown, the ℓ number of FGCs c j , . . . , ci
Note that the granularity change can be applied at once in a repository scale; applying repository transformation of strategy S to the whole repository, we can obtain another repository consisting of CGCs. The relationship between a squash unit and its squashed CGC can be determined by associating the original FGCs with CGCs in the transformed repository. 8
The first line in refactoring detection result
2 public class A { 3 /** 4 * Javadoc for methodA() 5 */ 6 public void methodA() { 7 8 } 9 }
which is the oldest commit in the squash unit of cℓj . In the figure, the result of git-blame-based tracing revealed that the origin of the refactoring target is found in c′j at Line n′2 on File f2′ , indicating the most recent modification of the refactored code element at this line. If the commits being squashed contain code changes applied to the refactoring target, the tracing result will be found in one of these commits. However, since these commits are squashed into a CGC, they cannot serve as the tracing result for refactorings in the CGC. To match the refactorings from both coarsegrained and fine-grained, in tracing r1 , we trace the commit, excluding those being squashed, that last modified the first line of the refactoring target. This is achieved by utilizing git-blame with the --ignore-rev option to ignore the code changes in the squashed commits. In this case, we run git-blame at n1 on f1 at commit ci while ignoring the commits of c j , . . . , ck . In the figure, the trace result is in c′i at Line n′1 on File f1′ . Two refactorings r1 and r2 are matched as the same refactoring (r1 ∼ r2 ) if they conform to the same conditions: 1) having the same refactoring type (r1 .type = r2 .type), 2) having the same refactoring target signature (r1 .elementSig = r2 .elementSig), and 3) their refactoring targets are traced to the same location at the same commit (c′i = c′j ∧ f1′ = f2′ ∧ n′1 = n′2 ).
2 public class B { 3 4 }
Move Method
(a) Before preprocessing. The first line in refactoring detection result
2 public class A { 3 public void methodA() { 4 5 } 6 }
2 public class B { 3 4 }
Move Method
(b) After preprocessing. Figure 6: Example of preprocessing. h: Extrated straight commit sequence original repository
𝑐$$
𝑐!$
𝑐%$
𝑐&$
𝑐"$
𝑐#$
𝑐*$
squash unit
transform
squash
𝑐$!
𝑆'!
𝑐&!
𝑐%! 𝑐!!
𝑆$!
𝑐"!
𝑐#!
5.8. Comparison and Matching To identify CGRs and EPRs, comparison and matching are performed among refactorings detected from different levels of CGCs and from FGCs. To obtain CGRs, refactorings detected from CGCs are compared with those detected from CGCs at a lower granularity level and FGCs. The comparison with refactorings from lower granularity level CGCs is necessary because some CGRs at a lower granularity level may also be detected at a higher granularity level. To be more specific, a CGR r detected from CGC cu = sq(u), which is squashed by a squash unit consisting of FGCs u = {c1 , c2 }, may also be detected from another CGC cu′ = sq(u′ ), which is squashed by FGCs u′ = {c1 , c2 , c3 }, if the code change in c3 does not disrupt the detection of r. The smallest squash unit where the CGR is detected is considered as its granularity level. The CGRs deduced from a squashed commit c = sq(u) are defined as follows: [ ′ ′ ′ CGR(c) = ∄r ∈ ref(c ) • r ∼ r . r ∈ ref(c) c′ ∈sub(c)
sub
𝑐$%
𝑆'%
𝑐"% 𝑐!%
𝑆$%
𝑐%%
𝑆!%
𝑆(ℓ
𝑐&%
o
(offset)
ℓ
(granularity level)
𝑐#%
ℓ
Figure 7: Comparison of different strategies.
are squashed into a CGC cℓj at the granularity level of ℓ. Here, a refactoring r1 at Line n1 on File f1 is detected from the FGC ci and another refactoring r2 at Line n2 on File f2 is detected from a CGC cℓj . Our aim is to trace the location, where the refactored code snippet is modified in the nearest parent commit of the being squashed commits. Following the techniques employed in violation tracing studies [34, 35], we implemented the trace mechanism using gitblame. Applying git-blame on a certain line of code can reveal the information about the last modification of that line. The information includes the commit hash, author, date, location, which is the line number, and the contents. The target of our tracing is the commit hash, file path, and the line number of the refactored code. The first line of the refactoring target is typically the signature of the refactored code element, which effectively represents the element. We apply git-blame on the first line of the refactoring target of r2 . Since r2 is detected in the squashed repository, we perform the tracing on the original repository to enable comparison with results from the original commit history. In this case, we run git-blame at n2 on f2 at c j ,
To obtain EPRs, the traced refactorings detected from FGCs are compared with the refactorings detected from CGCs at a higher granularity level. The granularity level of the squash unit that contains the least FGCs where the EPR is detected in FGC but not in CGC is considered as its granularity level. The EPRs at granularity ℓ deduced from a FGC c are defined as follows: [ ℓ ′ ′ ′ EPR (c) = r ∈ ref(c) ∄r ∈ ref(c ) • r ∼ r c′ ∈sup(c)∩C ℓ (H) [ ′ \ EPRℓ (c). ℓ′ <ℓ
9
…
original repository
𝑐#$
… File f2' Line n2'
squashed repository
𝑆%ℓ
trace
File f1' Line n1'
…
𝑐!$ trace
𝑐#
…
𝑐"
1
𝑐!
File f2 Line n2 f r2 File Line n
squash
2
…
f r1 File Line n 1
--ignore-rev
2
𝑐#ℓ Figure 8: Refactoring location tracing.
Note that the above definition of CGR differs from the one used in our previous study [14]. In the previous study, we defined a refactoring r detected from a CGC c = sq(u) squashed from a squash unit u as coarse-grained if and only if no refactoring of its type was found in the detected refactorings from each FGC in u: [ CGR′ (c) = {r ∈ ref(c) | ∄r′ ∈ ref(c) • r.type = r′.type}.
6.2. Data Collection The repositories that we selected are from a dataset collected by Silva et al. [31], containing 124 GitHub-hosted Java projects. These repositories contain refactorings, and some of them have been identified by RefactoringMiner, studied, and confirmed by researchers. Given the variety of coding conventions and practices in different domains of repositories, the refactorings applied in each domain may be different. To mitigate the above bias, we chose a total of 32 repositories from the 124 repositories with consideration of the variety in the projects’ domains. The 19 repositories used in our previous study [14] were included, and the remaining 13 repositories were randomly selected, each with distinct domains. The detail of our dataset is shown in Table 1. Each column represents the repository name, the number of commits to be processed in the experiment, the number of straight commit sequences, the number of commits involved in the straight commit sequences, the average length of straight commit sequences, the average number of squash units at each granularity level from 2 to 5, the number of refactorings detected, and the domain of the repository. Note that in Section 5.4, we introduced that different strategies determined by different offset values are used for determining squash units at a certain granularity level, e.g., there are two strategies at ℓ = 2, three strategies at ℓ = 3. The average number of squash units is calculated using the total number of squash units under different strategies to divide the number of strategies at that granularity level. The domain of repositories encompasses web frameworks, Android apps, performance toolkits, plugins, and more. The state-of-the-art refactoring detector RefactoringMiner 3.0.4 [12] was used to detect refactorings. As described in Section 5.3, merge commits and branch commits were excluded. The number of commits after this exclusion ranged from 295 to 19,782. The average lengths of straight commit sequences range from 1.36 to 358.55. Table 2 shows the number of refactorings detected from each repository. The second column expresses the number of original refactorings detected, ranging from 178 to 47,192. The number of CGRs and EPRs detected from granularity level two to granularity level five in each repository is also presented in the table.
c∈u
There are two differences between the two definitions. The first difference is that, in this study, we improved the accuracy using a more sophisticated signature to identify the same refactorings among different granularity levels, instead of using only the type of refactorings as the signature. The second is that, in this study, refactorings found at the granularity level ℓ were not regarded as CGRs if they were also found at any finer granularity than ℓ, which improves the conceptual validity. Our implementation is publicly available online4 .
6. Empirical Study 6.1. Research Questions We set five research questions (RQs) to investigate the impact of the granularity change in refactoring detection from the aspects of 1) the appearance frequency of CGRs and EPRs (RQ1 and RQ5 ), 2) type analysis of CGRs and EPRs (RQ2 and RQ5 ), 3) the cause of CGRs and EPRs (RQ3 and RQ5 ), and 4) the relationship between commit messages and CGR (RQ4 ). The five RQs are as follows: RQ1 : How frequently do CGRs appear because of granularity change? RQ2 : What are the types of CGRs? RQ3 : How are CGRs introduced in development? RQ4 : Do commit messages suggest the existence of CGRs? RQ5 : What are the features of EPRs?
4 https://github.com/MashiroCl/GranulRef
The details of each RQ are introduced below. 10
Table 1: Dataset used Repository # commits mbassador 342 retrolambda 530 seyren 640 android-async-http 899 javapoet 937 sshj 1,045 RoboBinding 1,088 giraph 1,138 jeromq 1,470 jfinal 1,655 zuul 1,689 baasbox 1,706 spring-data-rest 1,755 truth 2,006 rest-assured 2,309 helios 2,458 cascading 2,528 goclipse 2,925 HikariCP 2,929 rest.li 2,931 hydra 2,958 blueflood 3,152 PocketHub 3,512 xabber-android 4,264 morphia 4,499 redisson 10,014 Activiti 11,135 processing 13,211 checkstyle 14,360 libgdx 15,433 cgeo 18,941 JGroups 20,367 Total 154,826
# commit sequences 84 54 250 265 457 155 301 17 579 144 357 662 34 205 142 1,166 190 651 389 67 487 910 525 376 731 2,907 3,017 2,106 40 4,443 5,336 1,191 28,238
# involved commits 295 506 498 754 621 976 956 1,132 1,063 1,595 1,517 1,191 1,740 1,868 2,243 1,727 2,434 2,496 2,746 2,908 2,749 2,561 3,211 4,095 4,152 8,638 9,543 12,299 14,342 13,177 16,103 19,782 139,918
Avg. seq. length 3.51 9.37 1.99 2.85 1.36 6.30 3.18 66.59 1.84 11.08 4.25 1.80 51.18 9.11 15.80 1.48 12.81 3.83 7.06 43.40 5.64 2.81 6.12 10.89 5.68 2.97 3.16 5.84 358.55 2.97 3.02 16.61 4.95
ℓ=2 99.00 224.50 104.50 226.00 74.50 396.00 306.00 556.00 200.00 717.50 551.00 226.00 851.00 809.00 1,042.50 218.00 1,113.50 865.50 1,154.00 1,415.00 1,098.50 756.50 1,314.00 1,844.50 1,694.50 2,604.50 3,056.00 4,944.50 7,148.00 4,021.00 4,938.50 9,242.50
Average # squash units ℓ=3 ℓ=4 58.33 39.75 140.00 99.50 49.33 29.75 128.67 82.00 23.00 9.75 243.00 169.00 164.33 97.50 367.33 274.00 99.00 58.75 459.67 332.50 326.67 224.00 105.67 54.00 563.00 419.50 514.00 372.25 671.33 492.25 99.33 53.75 694.00 489.25 502.00 336.50 711.67 492.25 930.33 690.75 633.33 420.00 402.00 249.00 840.67 613.00 1,170.00 840.25 1,087.33 790.75 1,375.33 842.75 1,781.00 1,213.75 3,076.33 2,192.75 4,758.33 3,565.50 2,302.00 1,562.50 2,666.67 1,666.75 6,035.33 4,453.75
ℓ=5 29.60 75.40 18.20 58.20 6.40 127.20 64.80 218.00 40.00 256.60 164.40 35.20 333.20 290.80 385.80 33.80 368.00 245.40 369.00 547.20 299.40 168.60 480.40 647.80 614.80 559.40 902.60 1,682.60 2,849.00 1,159.00 1,145.80 3,516.40
Description/domain Event bus Backport of lambda expression Dashboard HTTP client Source file generator SSH library Data binding framework Graph processing system Messaging library Web framework Gateway service Backend server RESTful data access Java assertions REST service testing Container orchestration framework Data processing IDE for Go language JDBC connection REST framework Distributed data processing Data processing Android app XMPP client for Android Java MongoDB ORM Redis client Business Process Management Platform Code learning platform Code quality linter Game development framework Client for geocaching Messaging library
methods share some similar logic related to removal functionality, they serve different testing purposes. The second false positive was found in a CGC at ℓ = 5, consisting of five commits6 . Initially, class EventListener2 extended EventListener1, and EventListener3 extended EventListener2. Both EventListener2 and EventListener3 contained a method handleString() with an empty body. During the code change, the member method handleString() in EventListener2 was removed, and an annotation of that method in class EventListener3 was removed. The RefactoringMiner detected an Extract SuperClass. It claims that EventListener2 is extracted from EventListener3, which is incorrect, as the class structure was already in place and no such extraction occurred. Based on this review, RefactoringMiner achieved a precision of 96.36% (53/55) in this repository under granularity changes. Although this is slightly lower than the 99.6% precision reported by the tool’s authors [12], it remains high. The decrease is acceptable, as squashing code changes from multiple commits into a single commit increases the complexity of the commit.
6.3. Preliminary Study on Detection Precision and Efficiency under Granularity Changes Since changes in commit history granularity can lead to commits containing more changes, potentially affecting the precision and efficiency of refactoring detectors, we conducted a preliminary study to evaluate the precision and efficiency of RefactoringMiner under varying granularity levels. We applied the techniques introduced in Section 5 to extract CGRs from the mbassador repository. We then manually reviewed the detected CGRs and measured the time required for detection. 6.3.1. Precision The first author manually reviewed all 55 CGRs identified in the repository mbassador. Details regarding the number of commits per CGC at each granularity level can be found in Table 2. Among the 55 CGRs, two were found to be false positives, and we briefly describe each of them below. The first one is a CGR at ℓ = 4, detected in a CGC squashed by four commits5 . Initially, there are two methods testRemove1 and testRemove2(). During the code change, testRemove1() was renamed to testRemove2() with formatting modifications, while the original testRemove2() was removed, and a new method testCompleteRemoval() was introduced. RefactoringMiner incorrectly reported a CGR of Rename Method from testRemove1() to testCompleteRemoval(). While the two
6.3.2. Efficiency The time consumed for detecting refactorings in both the original and squashed versions of the mbassador repository is summarized in Table 3. We perform refactoring detection using RefactoringMiner 3.0.4 with its default settings. The detection is conducted on each commit in the target repository,
5 https://github.com/bennidi/mbassador/commit/{21385b6,31e6ca1,a47b63 2,89f4454}
6 https://github.com/bennidi/mbassador/commit/{d08470c,1ed1bb2,f2a61c 1,c83d870,32bed9a}
11
Table 2: Detected CGRs and EPRs Repository mbassador retrolambda seyren android-async-http javapoet sshj RoboBinding giraph jeromq jfinal zuul baasbox spring-data-rest truth rest-assured helios cascading goclipse HikariCP rest.li hydra blueflood PocketHub xabber-android morphia redisson Activiti processing checkstyle libgdx cgeo JGroups
# Original refactorings 1,019 855 427 1,092 178 2,259 12,154 11,956 2,855 2,760 3,386 889 6,814 8,882 2,930 2,577 12,483 13,164 4,181 21,545 11,423 5,661 6,181 7,062 28,468 23,374 30,091 32,286 29,281 36,557 37,178 47,192
ℓ=2 19 64 3 26 6 48 893 189 64 88 78 15 92 139 24 56 238 562 138 56 185 190 40 400 691 361 705 1,355 489 1,151 328 1,124
# CGRs ℓ=3 ℓ=4 22 11 26 39 1 0 50 0 7 0 26 9 538 374 118 48 48 35 31 9 20 60 7 9 49 58 35 47 18 19 69 34 162 91 249 161 28 7 139 58 126 84 48 76 33 55 149 148 324 241 176 97 408 256 1,236 715 243 202 604 398 310 242 893 436
# FGCs/CGCs 295 198 175 159 148
Time(s) 374.94 343.71 404.88 407.43 418.82
ℓ=2 14 114 26 45 3 133 1,188 154 253 126 130 53 406 359 171 122 632 854 297 463 691 460 199 626 1,140 745 1,586 1,889 532 1,188 788 1,145
# EPRs ℓ=3 ℓ=4 2 17 61 27 13 2 141 19 2 0 60 24 724 512 350 156 80 58 111 29 134 89 18 8 106 169 101 240 132 38 44 52 309 169 447 314 74 40 1,495 216 333 93 199 198 99 112 239 159 659 417 362 180 878 518 1,335 3,267 206 213 708 373 858 254 665 579
ℓ=5 6 10 1 4 1 25 262 261 39 48 153 5 59 105 54 36 101 135 69 144 104 106 154 142 381 173 319 401 133 331 191 285
the frequency of CGRs in open-source repositories with various functionalities at different granularity levels.
Table 3: RefactoringMiner detection time on mbassador ℓ 1 2 3 4 5
ℓ=5 2 13 0 1 2 9 233 112 45 37 81 2 21 33 14 12 101 100 25 36 81 35 91 77 293 97 253 293 209 163 111 639
Average Time(s) 1.27 1.74 2.31 2.56 2.83
6.4.2. Study Design The techniques introduced in Sections 5.1 and 5.7 are applied to our dataset to extract squash units, change the granularity of commits, and match the refactoring detection results to find CGRs. The frequency of CGRs appearing in commit history C of granularity ℓ, i.e., the ratio of CGR to original refactorings, is presented as follows: S | c∈C ℓ (H) CGR(c)| CGR Frequency (H, ℓ) = |OGR(H)|
executed on a MacBook equipped with an Apple M3 Pro chip and 36 GB of RAM. The table presents the granularity level, the number of FGCs (ℓ = 1) or CGCs (ℓ > 1) targeted for the detection, the total detection time, and the average detection time per FGC/CGC. Note that the total detection time includes the process initialization cost of RefactoringMiner per commit since it ran at the commit level. A statistically significant positive correlation is observed between the granularity level and the average detection time per FGC/CGC (Pearson’s r = 0.986, p < 0.05). This positive correlation is expected, as squashing code changes from multiple commits into a single commit increases the number of code elements that the detector needs to analyze, thereby increasing the average detection time. In conclusion, changes in granularity level affect the precision and efficiency of RefactoringMiner, but the impact remains within an acceptable range.
where OGR(H) represents the original refactorings detected in H. In this study, we refined the definition of CGR, ensuring that CGRs from different granularity levels are disjoint. To validate the conclusions from our previous work, which was based on the earlier definition where CGRs from different granularity levels overlapped, we calculated the accumulated frequency. The accumulated Frequency for CGR at granularity level ℓ is the cumulative frequency of CGRs at granularity levels up to and including ℓ. We calculate the frequencies and the accumulated frequencies of CGRs for all repositories in the dataset at the granularity levels of ℓ ∈ {2, 3, 4, 5}. Additionally, we calculate the relative amounts of CGR, EPR, and remaining refactorings based on their detected occurrences for each repository in our dataset. The relative amount of
6.4. RQ1 : How frequently do CGRs appear because of granularity change? 6.4.1. Motivation This RQ aims to investigate the frequency of CGR occurrences in software development and the necessity of incorporating CGR detection support in refactoring tools. We explore 12
4000
# detected CGRs
3000
2000
1000
Figure 9: Frequency and accumulated frequency of CGRs. 0 EPR
CGR
0
5000
10000
15000
20000
# commits involved
100%
Figure 11: Correlation between number of commits involved in repositories and number of detected CGRs.
75%
50%
Occurrences of CGRs are a common phenomenon across repositories from various domains and granularity levels. As shown in Figure 10, although the relative amounts of CGRs vary across repositories, they are all positive, indicating that CGRs are detected in all repositories. Regarding different granularity levels, except for three repositories where the frequencies are zero at certain levels, all other repositories exhibit consistently positive CGR frequencies across all granularity levels as shown in Table 2. We also observed a statistically significant positive correlation between the number of commits involved in the straight commit sequences in a repository and the number of detected CGRs, with a Pearson’s r = 0.748 with p < 0.001. This relationship is illustrated in the scatter plot in Figure 11, where the x-axis is the number of commits involved in the straight commit sequences in a repository, and the y-axis is the number of detected CGRs. This suggests that the extent of code evolution in a project contributes significantly to the number of CGRs observed. Projects with more commits typically involve more extensive and iterative changes, which increase the likelihood of CGRs to emerge and be detected. Nevertheless, the commit count is not the only factor that affects the CGR occurrences across projects. We observed that retrolambda and seyren, which have similar commit counts, exhibited notably different numbers of CGRs. This suggests that additional factors, such as development practices or commitment strategies, may influence the occurrence of CGRs. Investigating these factors represents a promising direction for future work. However, the CGR frequency decreases as the granularity level increases. As the Figure 9 shows, the maximum and median of CGR frequency decrease among the four granularity levels. Particularly, the median of the granularity level at 5 is 0.0071, the lowest among the measured four granularities. This is because a CGR is classified at a coarser granularity only when the code change of a single refactoring is distributed across all the consecutive commits in a larger squash unit. Such occurrences are relatively rare in practical development, making
processing
retrolambda
RoboBinding
xabber-android
zuul
goclipse
javapoet
android-async-http
helios
jeromq
libgdx
JGroups
jfinal
blueflood
Activiti
morphia
mbassador
HikariCP
cascading
sshj
hydra
giraph
baasbox
checkstyle
PocketHub
truth
spring-data-rest
cgeo
redisson
rest.li
rest-assured
0%
seyren
25%
Figure 10: Relative amount of CGRs and EPRs for 32 repositories.
each type is determined as the proportion of its detected count to the total detected count across all three types. 6.5. Results and Discussion In Figure 9, the four boxes on the left represent CGR frequencies, while the four on the right show accumulated CGR frequencies at granularity levels of ℓ ∈ {2, 3, 4, 5} across 32 repositories. The x-axis represents the granularity levels, ranging from ℓ = 2 to ℓ = 5, while the y-axis measures the frequency/accumulated frequency of CGR occurrences. Figure 10 depicts the relative amount of CGR and EPR against all the refactorings in our dataset. For the CGR frequencies as shown in Figure 9, the majority of the boxes for each granularity level indicate that CGR frequencies fall within the range of 0.0 to 0.0566. The average appearance frequencies of CGRs for all repositories are 0.0256, 0.0158, 0.0101, and 0.0080 when the granularity level is 2, 3, 4, and 5, respectively. However, certain repositories exhibit notably higher frequencies at specific granularities. At granularity levels of 2 and 4, the repositories retrolambda and RoboBinding show notably higher frequencies. Specifically, at the granularity level of 2, their frequencies are 0.0749 and 0.0735, respectively, while at the granularity level of 4, they record 0.0456 and 0.0308. At the granularity level of 3, repositories android-async-http and RoboBinding show the highest frequencies, at 0.0458 and 0.0443, respectively. 13
HRM
HRM
+HRM( uri:String, path:String, …)
ba4e99c
1
fac89913
2
fd68976
3
0df4413
ity level of ℓ = 5.
+HRM( path:String, scheme:String, …)
Rename Parameter
4
f7538f1
6.6. RQ2 : What are the types of CGRs? 6.6.1. Motivation This RQ is set to investigate the composition of the detected CGRs in terms of refactoring type. By answering this RQ, we can provide suggestions about the CGR types that developers of refactoring detectors should prioritize when implementing the CGR-supported detector. In addition, we also explore the cause of why some types are more frequent than others.
5
Commit history of zuul
HRM +HRM(uri:String, path:String, scheme:String, …)
Split Parameter
Figure 12: Example of CGR at ℓ = 5 in zuul.
6.6.2. Study Design From the CGRs detected from our dataset, we calculated the appearance ratio of a specific refactoring type at each granularity level. The appearance ratio expresses the prevalence of one type of CGRs relative to the same type original refactorings. For a certain refactoring type t in the commit history C at granularity level ℓ, the appearance ratio can be expressed as follows: S |{r ∈ c∈C ℓ (H) CGR(c) | r.type = t}| CGR Appearancet (H, ℓ) = |{r ∈ OGR(H) | r.type = t}|
CGRs at the coarsest granularity (ℓ = 5) less frequent. From another perspective, the cumulative number of CGRs increases progressively with each increment in the granularity level as shown by the right-side four boxes in Figure 9. At granularity level 5, the accumulated frequency reaches 0.0571, meaning that for every 100 refactorings performed, approximately 5.71 CGRs are also conducted. This exceeds 5%, emphasizing that the detection of CGRs should not be overlooked. However, the increase gradually becomes smaller and tends to saturate. This fact also validates our experimental settings, which set the maximum granularity level at 5, as the frequencies for coarser granularities would be trivial. Also, it aligns with our previous study’s conclusion that when CGRs at different granularities are not disjoint, their frequencies tend to increase as the granularity level rises. An example of a coarse-grained refactoring at ℓ = 5 found in repository zuul7 is shown in Figure 12. The constructor of class HttpRequestMessage (HRM hereinafter) contains several parameters, including uri of type String and some other parameters. In the first proposed commit ba4e99c, the parameter named uri of String is removed, and another parameter path of String is added into the constructor, and the functions related with uri are replaced by parameter path. It is detected as a refactoring of Rename Parameter. In the three commits proposed later, HRM is unchanged. Then, a new parameter scheme of type String is added into the constructor in the fifth proposed commit, and the functions of the parameter path are replaced by path and scheme. The code change over these five commits is detected as a Split Parameter, which splits the parameter uri into parameters path and scheme.
where the granularity ℓ of at most 5 was used in this study. The category of types can refer to Tsantalis et al. [12]. We calculate the appearance ratio of each type of CGR in our dataset. 6.6.3. Results and Discussions The appearance ratios of each type of CGR across all granularities and the cross-granularity appearance ratios are calculated and ranked. The top ten are listed in Table 4, with their appearance ratio. In total, 98 distinct CGR types are detected. The ranking of CGR types varies across granularity levels, indicating that different types of CGRs tend to be applied at specific levels of granularity. The most frequent CGR types for both granularity levels of 2 and 3 are Split Package with the same appearance ratio: 29.63%. For the granularity levels of 4 and 5, the most frequent types are Split Class and Merge Class with appearance ratios of 23.44% and 18.00%, respectively. Except for Replace Attribute, all the top three CGR types are associated with splitting or merging operations on packages or classes. These types of refactorings are inherently complex and require significant modifications, often affecting multiple parts of the code base. This observation suggests that such complicated, large-scale refactorings typically require incremental modifications spread over multiple commits and appear as CGRs, which highlights the utility of CGRs in facilitating software evolution understanding. As for Replace Attribute, this refactoring is usually associated with two steps: 1) modifications to the inheritance structure of existing classes, and 2) replacement of an attribute in the subclass with an attribute inherited from the parent class. We show three examples that contain CGR types of high appearance ratio as detailed in the following paragraphs. The first
The appearance of coarse-grained refactoring is a common phenomenon in all repositories, and it occurs most frequently at the granularity level of ℓ = 2. It is observed that the frequency of CGRs decreases as the granularity level increases. This trend indicates that the frequency of CGRs decreases as the granularity level increases, and On average, according to the number of refactorings conducted, an additional 5.71% of CGRs are being performed if we see at most the granular-
7 https://github.com/Netflix/zuul/commit/f7538f1
14
Table 4: Most frequently found CGR types and their appearance ratio. Rank 1 2 3 4 5 6 7 8 9 10
Granularity level ℓ=4
ℓ=2
ℓ=3
ℓ=5
cross-granularity
Split Package
Split Package
Split Class
Merge Class
Merge Class
(29.63%)
(29.63%)
(23.44%)
(18.00%)
(86.00%)
Merge Class
Merge Class
Replace Attribute
Split Class
Split Package
(28.00%)
(26.00%)
(18.18%)
(10.94%)
(81.48%) Split Class
Replace Attribute
Split Class
Split Package
Merge Method
(27.27%)
(15.62%)
(14.81%)
(8.33%)
(73.44%)
Merge Method
Move Package
Merge Class
Split Package
Replace Attribute
(25.00%)
(12.88%)
(14.00%)
(7.41%)
(58.18%)
Split Class
Move and Rename Class
Merge Method
Replace Attribute
Merge Method
(23.44%)
(10.69%)
(12.50%)
(7.27%)
(54.17%)
Rename Package
Move and Rename Method
Move and Rename Attribute
Move and Rename Method
Move Package
(16.26%)
(9.80%)
(11.65%)
(6.22%)
(41.67%)
Merge Package
Merge Parameter
Move Package
Merge Package
Move and Rename Method
(15.15%)
(9.60%)
(11.36%)
(6.06%)
(39.69%)
Move and Rename Method
Move Attribute
Split Method
Move and Rename Attribute
Move and Rename Class
(14.29%)
(9.52%)
(10.61%)
(6.02%)
(38.43%)
Move and Rename Class
Merge Package
Move and Rename Method
Move and Rename Class
Merge Package
(13.42%)
(9.09%)
(9.38%)
(5.52%)
(30.30%)
Move Package
Move Method
Move and Rename Class
Move Package
Move and Rename Attribute
(12.88%)
(8.67%)
(8.81%)
(4.55%)
(30.12%)
example is a Merge Class type CGR, shown in Figure 13, the second is a Replace Attribute type CGR, shown in Figure 15; and the final example is a Split Package type CGR shown in Figure 14. The first example is a Merge Class type CGR detected in repository android-async-http8 as shown in Figure 13. The Merge Class is conducted across two commits. Initially, the ArgsUtils and AssertUtils classes each contain only a single method: notNull() and asserts(), respectively. The method asserts() is invoked in methods in class AsyncHttpResponseHandler, and the method notNull() is invoked in methods in class AsyncHttpRequest. In the first commit, the classes ArgsUtils and AssertUtils are removed. The invocation of their methods in classes AsyncHttpResponseHandler and AsyncHttpRequest are replaced with calls to placeholder methods, Utils.asserts() and Utils.notNull(), defined in an unimplemented Utils class that acts as a stub. The implementation of these stub methods is completed in the second commit 88196c4, where the class Utils is introduced, containing the methods as serts() and notNull() methods with the same method bodies as those in the initial state. Though no refactoring is detected within either of the commits, a refactoring of Merge Class is detected across these two commits: the class AsyncHttpResponseHandler and AsyncHttpRequest is merged into a more consolidated utility class Utils. We deduce that the motivation for splitting the Merge Class refactoring across two commits is to allow the developer to demonstrate how the new class and methods integrate into the system before finalizing their implementation. By forwarding the method declarations initially and deferring their implementation, the developer ensures compatibility and minimizes potential disruptions during the transition. Another example of a Split Package type CGR conducted across three commits is detected in repository giraph9 as il-
lustrated in Figure 14. In the initial state, the package com.yahoo.haddop_bsp contains multiple files. In the first commit 40169f8, a Rename Package refactoring is performed, renaming the package to org.apache.giraph. The second commit 22f425e is an empty commit with a message indicating that packages starting with com are removed. No refactoring is detected in this commit. In the third commit 092fba0, a refactoring of Split Package is detected that the package org.apache.giraph is removed and its files inside are moved to several newly created packages, including org.apache.giraph.dsp, org.apache.giraph.comm and others. Across these three commits, a refactoring of Split Package is observed. Different from the one detected in commit 092fba0, the package being split is com.yahoo.hadoop_bsp. Package-related refactorings often have a large-scale impact on the code base. Instead of directly splitting com.yahoo.hadoop_bsp into packages org.apache.giraph.dsp, org.apache.giraph.comm, and others, the developer applied this operation in two steps: renaming and splitting, and used one empty commit as a milestone to signify the renaming before splitting. This incremental approach helped manage the transition and ensure the system’s stability during the refactoring. The third example is a Replace Attribute type CGR conducted across two commits detected in repository Activiti10 as depicted in Figure 15. Prior to the first commit, there are two classes ProcessEngineTestCase and TablePageQueryTest. The ProcessEngineTestCase contains an attribute taskService of type TaskService, which is an interface. The TablePageQueryTest has an attribute processEngineBuilder (pEB) of type ProcessEngineTestWatchman (PETW), and its method deleteTasks() is invoked within the member method testGetTablePage(). In the first commit, the inheritance structure of TablePageQueryTest is modified to inherit from the class ProcessEngineTestCase, and the attribute pEB is removed. The usages of pEB are replaced by a newly created
8 https://github.com/android-async-http/android-async-http/commit/88106
c4 9 https://github.com/apache/giraph/commit/092fba0
10 https://github.com/Activiti/Activiti/commit/958b3c8
15
ArgsUtils
AssertUtils
- ArgsUtils
- AssertUtils
notNull( )
asserts( )
- notNull( )
- asserts( )
AsyncHttpResponseHandler
AsyncHttpResponseHandler
{ }
{ AssertUtils.asserts( ) }
AsyncHttpResquest { }
+ Utils
- AssertUtils.asserts( ) + Utils.asserts( )
+ notNull( ) + asserts( )
AsyncHttpResquest {
ArgsUtils.notNull( )
- ArgsUtils.notNull( ) + Utils.notNull( ) }
29afb2c
1
88106c4
2
No refactoring detected
No refactoring detected
Commit history of android-async-http
Refactoring detected: Merge Class
Figure 13: CGR of type Merge Class found in android-async-http.
+ org.apache.giraph - com.yahoo.hadoop_bsp
com.yahoo.hadoop_bsp
- org.apache.giraph
BspInputFormat
BspInputFormat
BspInputSplit
BspInputSplit
CentralizedService
CentralizedService
ArrayListWritable
ArrayListWritable
- ArrayListWritable
BasicRPCCommunications
BasicRPCCommunications
- BasicRPCCommunications
…
- BspInputFormat Empty commit as a milestone indicating removal of com.* packages
…
+ org.apache.giraph.dsp + BspInputFormat
- BspInputSplit
+ BspInputSplit
- CentralizedService
+ CentralizedService
…
+ org.apache.giraph.comm + ArrayListWritable + BasicRPCCommunications
…
40169f8
1
22f425e
Rename Package
2
092fba0
No refactoring detected
3
Commit history of giraph
Split Package
Refactoring detected: Split Package
Figure 14: CGR of type Split Package found in giraph.
member method deleteTasks(). In the second commit, a method declaration of deleteTasks() is added into the interface TaskService. Within class TablePageQueryTest, the method deleteTasks() created in the previous commit is replaced by an invocation of member method deleteTasks() on the attribute taskService, which is inherited from the parent class ProcessEngineTestCase. No refactoring is detected in either of the individual commits. However, a Replace Attribute refactoring is identified across the two commits. This refactoring replaces the attribute pEB in class TablePageQueryTest with the inherited attribute taskService from parent class ProcessEngineTestCase.
tally and across multiple commits rather than in a single commit. 6.7. RQ3 : How are CGRs introduced in development? 6.7.1. Motivation In this RQ, we analyzed the cause of CGRs to gain a comprehensive understanding of how developers introduce CGRs in development. 6.7.2. Study Design We conducted a manual analysis for a sampled set of CGRs. The first author randomly sampled 110 CGRs from 110 distinct CGC in the dataset. The sample size of 110 was not based on a specific statistical rationale; rather, it was determined by balancing the required human effort with the goals of the experiment, which is to reveal the typical causes of CGRs.
Refactoring types related to splitting or merging classes and packages, as well as those involving modifications to the inheritance structure, rank highly in CGRs due to their complexity. These refactorings are often implemented incremen16
ProcessEngineTestCase
ProcessEngineTestCase
ProcessEngineTestCase
taskService: TaskService
taskService: TaskService
taskService: TaskService
+
<<Interface>> TaskService
TablePageQueryTest
+ deleteTasks
<<Interface>> TaskService
<<Interface>> TaskService
- pEB: PETW
TablePageQueryTest pEB: PETW
testGetTablePage( ){ pEB.deleteTasks( ); } 4dacf10
TablePageQueryTest
testGetTablePage( ){ - pEB.deleteTasks( ); + deleteTasks( ); } + deleteTasks( ){…} 1
testGetTablePage( ){ - deleteTasks( ); + taskService.deleteTasks( ); } - deleteTasks( );
958b3c8
No refactoring detected
2
Commit history of Activiti
No refactoring detected
Refactoring detected: Replace Attribute
Figure 15: CGR of type Replace Attribute found in Activiti.
To reduce potential bias, we ensured that the samples maintain a reasonably balanced distribution across different granularity levels, with a slight emphasis on lower granularity CGRs. This sampling strategy aligns with our finding in Section 6.5 that lower granularity CGRs tend to occur more frequently. Among the 110 CGRs, the distribution across granularity levels 2, 3, 4, and 5 was 35, 27, 24, and 24, respectively. Then, we manually compared and analyzed the conducted changes and refactorings detected, and recognized the cause of why such a change was detected as a CGR. Based on the recognized causes, we classify the detected CGRs. Manual analysis can be confirmed in our supplemental package [39].
- core.value
989bf50 2
ce2a9e9 1
split
+core.util.email +core.util.graphite
core.util.PropertyMailSender
+core.util.email.PropertyMailSender -services.PropertyMailSender +core.util.PropertyMailSender
move
move
Figure 16: Combination type CGR found in seyren.
package hierarchy of the repository is shown in the figure. In the parent commit ce2a9e9, the developer moved class PropertyMailSender under package services to package core.util, which was detected as refactoring Move Class. In child commit 989bf50, she/he split package core.value into core.util.email and another one, and then she/he moved class PropertyMailSender to the package core.util.email, which were detected as Split Package and Move Class. In terms of result, she/he applied Merge Package to merge part of the package core.value and the entire package services into a new package core.util.email. Although we define Generation and Combination as distinct categories based on whether the CGR is composed of nonrefactoring changes or multiple ephemeral refactorings, we acknowledge that borderline cases may exist. For example, a CGR identified as Combination might have a component refactoring that could be classified as a Generation type CGR. In such cases, we adopt a conservative classification policy; if any compositional aspect is found, we categorize the CGR as Combination. CGRs of Generation type may influence judgments of whether a module is refactored or not. We note that this type may also occur because of developers’ unawareness of refactoring; developers do not realize that the conducted code changes belong to refactoring operations. Supporting tools to guess developers’ manual edits and recognize refactoring activities [40, 41] may assist them in development. Because the
6.7.3. Results and Discussion We identified two types of causes for CGRs based on their composition: Generation and Combination. Within the sampled 110 CGRs in our dataset, 84 out of 110 (76.4%) of the CGRs are categorized into Generation type, whereas 26 out of 110 (23.6%) were into Combination type. Generation. This type of CGR is generated from nonrefactoring changes. The example shown in Figure 1 belongs to this type; the Move Method refactoring is generated by two non-refactoring changes: 1) copying the class implementation to a new file and 2) removing the origin class. Another example is in repository javapoet11 . In the parent commit 6a3595c, the attribute body was defined, and the method call method Writer.write() was removed. In child commit 4ff9adf, the developer added method invocation body.write(). In the CGC, the above code changes were detected as Rename Variable with Attribute; the variable methodWriter was renamed to attribute body. Combination. In contrast with Generation, this type is the combined result of multiple refactorings detected in finergrained commits. Figure 16 shows an example of this type found in repository seyren12 . For clarity, only part of the 11 https://github.com/square/javapoet/commit/{6a3595c,4ff9adf} 12 https://github.com/scobal/seyren/commit/{ce2a9e9,989bf50}
17
same commit messages13 : “introduced specialized subscriptions for better performance and customization options”, which suggests that some subscription-function-related classes are introduced. The commit message for the remaining commit14 is “refactorings. abstract base class. repackaging”, which indicates refactorings about repackaging an abstract class are conducted. In the CGC squashed by these three FGCs, four Move Class CGRs and four Change Access Modifier CGRs are detected. The cause for the detected CGRs is the same; originally, the classes being moved were in the same file. In the following commits, the implementations of those classes are copied to different newly created files, and the access modifiers for those classes are changed from private to public. And finally, the file which contains their original implementations is removed. To be more specific, we use two examples to explain. Originally, the file Mbassador.java contains two private abstract classes: Subscription and FilteredSubscription. In the firstly proposed commit 014f22d, the implementation of class Subscription is copied from file MBassador.java to a new file Subscription.java, and its class access modifier is changed from private to public. Then, similarly, in the commit 2ae0e5f proposed later, the implementation of the class FilteredSubscription is copied from the same file Mbassador.java to a new file FilteredSubscription.java with the change of class access modifier from private to public. In the last proposed commit 9ce3ceb, the file Mbassador.java is removed. The addition of implementation of class Subscription in file Subscription.java and removing of Subscription in file Mbassador.java is regarded as a Move Class CGR of ℓ = 3, moving class from Mbassador.java to Subscription.java. Similarly, for the change of class FilteredSubscription, it is detected as a Move Class refactoring at ℓ = 3. The second example is found in the repository RoboBinding15 . The commit messages of the two FGCs have the same content: “naming improvement to GroupedPropertyViewAttribute”, which suggests the code changes in the commits are renaming operations. The CGC squashed by these two fine-grained ones contains a Rename Class CGR of ℓ = 2. In the first proposed commit, the class GroupedPropertyAttributeImpl is removed. In the later proposed commit, a new class with the same implementation as the removed class is added, and its name is changed to GroupedAttributeDetailsImpl. The only change is the name of the class, and through the commit message, we deduced that the intention of the developers proposing these two commits is to perform a Rename Class refactoring. We believe that the existence of this type of CGR indicates that developers split complicated refactorings into steps and perform each step in a single commit to avoid the risk of affecting users. The consecutive same content commit messages also suggest that the refactoring is not completed in one com-
Combination type may influence type-based refactoring studies, such as investigations on frequently performed refactoring types, researchers may reconsider their results by covering coarse-grained types. The manual review revealed that 76.4% of the CGRs were Generation type and the rest of 23.6% were Combination type. Generation refers to new refactorings generated by fine-grained non-refactoring changes. Combination is a high-level refactoring combined with more than one refactoring changes.
6.8. RQ4 : Do commit messages suggest the existence of CGRs? 6.8.1. Motivation We set this RQ to explore whether or not developers will explicitly express their intention of performing CGRs. The explicitly stated intentions emphasize the necessity of current refactoring detection tools supporting CGR detection, and they are documented in the commit messages [42]. 6.8.2. Study Design The commit messages for the FGCs that are squashed into the 110 CGCs used in answering RQ3 were collected and manually reviewed by the first author. Commit messages were examined to infer the developers’ intent, with a focus on identifying the potential relationships among messages from different commits, such as similarity, revert, or other semantic patterns that exist, and suggesting the existence of CGR. We employed an incremental approach; as we encountered new patterns of intent, we added them to a growing set of characteristics. Messages that matched one of these characteristics were classified as suggesting the existence of CGRs, while vague or unrelated messages, such as “update version” or “update README”, were excluded. 6.8.3. Results and Discussions The commit messages of FGCs squashed into a CGC are regarded as the commit message of that CGC. The commit messages of 22 out of 110 CGCs were regarded as explicitly suggesting the existence of CGRs. We subdivided those commit messages into four types according to their characteristics. Characteristic 1: Same content commit messages (8/110). We observed that if the contents of multiple commit messages are the same and contain expressions about performing refactorings, there is a large possibility that CGRs exist. The same content messages indicate that a single piece of code changes is split into multiple steps, and each step is completed in one commit. Containing expressions about performing refactorings indicates refactorings are split, which fits the definition of CGRs. Eight CGCs among the investigated 110 CGCs are found to contain such type commit messages, and we show two of them below. The first example is three commits extracted from the repository mbassador. The two proposed earlier commits have the
13 https://github.com/bennidi/mbassador/commit/{014f22d,2ae0e5f} 14 https://github.com/bennidi/mbassador/commit/9ce3ceb 15 https://github.com/RoboBinding/RoboBinding/commit/{e16b566,7d75f3
1}
18
termediate step of this refactoring18 . The relationship in the commit messages indicates that some dependent code changes are performed over multiple commits. The code change may also be refactorings. As a result, this characteristic signals the existence of CGRs. Characteristic 4: Existence of connections among commits (6/110). If the commit messages of multiple consecutive commits show some connections, then it is possible that CGRs exist. The difference between this characteristic and the last one is that this type is not so apparent: some commits contain code changes that are performed on the same or related modules, which is not shown explicitly in the commit messages. We found six out of 110 CGCs containing this type of commit messages, and we show one example below. For two commits found in the repository zuul19 , the commit message of the first commit is “Preserve backward compatible ctor”, and for later proposed commit is “Remove redundant ctor. Stick to two variants”. The two commit messages have a connection that both commits are dealing with some constructor (ctor)-related. From the perspective of code change, the first commit introduced a constructor for class HttpRequestMessageImpl, and another constructor for the same class is removed in the latter proposed commit. The only difference between these two constructors is that the added one has one less parameter than the removed one. Through these two commits, the developers perform a Remove Parameter refactoring to remove a parameter in the constructor. The connection among commits is hard to recognize with not only the commit messages but also the knowledge about the actual code change and original code base. However, it can signal the existence of CGRs, as shown in the above example.
mit. Characteristic 2: Explicit expression about not finished refactoring (2/110). Commit messages of this type are a signal that some refactoring operations are split into several commits. Two of the 110 CGCs we investigated are found to contain this type, and we show one of them below. In two commits in the repository zuul16 , the commit message of the first commit is “Further refactoring to use the new filter and context interfaces. Still not complete”. The keywords “not complete” suggest the code change in that commit is incomplete. The commit next to it but proposed later has a commit message “And further refactoring to use the new filter and context interfaces”, that suggests this commit is to continue the incomplete refactoring. The code changes in these two commits also prove our deduction. In the first commit, the class HttpServletResponseWrapper in the file HttpServletResponseWrapper.java under the path http/ is removed. In the later proposed commit, the same implementation of the class HttpServletResponseWrapper is added in the same name file but under the path servlet/. The overall change is a Move Class refactoring to move the class file HttpServletResponseWrapper.java from path http/ to servlet/. This type of commit message explicitly points out the existence of CGRs. Characteristic 3: Clearly states the relationship between commits (6/110). If the commit messages of multiple commits show some relationships, then CGRs may exist in the squashed commit of these commits. We found six CGCs contained this type of commit messages and one example is explained below. In the two consecutive commits fedd975 and 8eb25de in the repository zuul17 . The commit message of the first commit fedd975 is “Revert “- Added a notifyUsage() template method on FilterProcessor, so that the filter metrics implementation can be overridden””. The commit 8eb25de proposed later has a commit message that says “Reverted the part of the previous commit, as changing from the Singleton pattern for FilterProcessor impacted too many filters that depended on getting access to it”, which reveals that the code change in this commit reverts part of the code change in the previous commit. The relationship between the two commits is that the latter reverts part of the change made by the previous one. The code change of the first commit fedd975 reverts the code change of a previous commit 7379aa8. The commit 7379aa8 extracts some code into a method notifyUsage(), and commit fedd975 removed that method and inlined the method body. While the second commit 8eb25de extracts the same code into a method notify() in a newly created class. The overall code change of the two commits fedd975 and 8eb25de can be summarized and detected as a Move and Rename Method CGR, which moves the method notifyUsage() to a newly created class and rename it to notify(). If we look from commit 7379aa8, a coarser granularity Extract Method CGR is applied to extract the method notify(), and the Move and Rename Method CGR is an in-
We found that 20% (22/110) of the commit messages explicitly express the existence of CGRs. The characteristics of those commit messages are summarized and categorized into four types, where the Same content commit messages. is the most frequent characteristic. 6.9. RQ5 : What are the features of EPRs? 6.9.1. Motivation In answering this RQ, we explore the characteristics of EPRs to gain deeper insights into them through three sub-questions. 6.9.2. RQ5a : How frequently do EPRs appear across the studied repositories? Study Design. Similar to CGR, we define the frequency of EPRs at granularity level ℓ in commit history H as follows: S | c∈C EPRℓ (c)| FrequencyEPR (H, ℓ) = . |OGR(H)| 18 As indicated by the commit messages, these changes actually correspond to 1) reverting the refactoring applied at one earlier commit (the parent of fedd975) and 2) conducting refactoring in a different way. This combination of reverting and reapplying refactoring could be recognized, at a coarse-grained level, as a distinct type of refactoring. 19 https://github.com/Netflix/zuul/commit/{fc79123,30b8d15}
16 https://github.com/Netflix/zuul/commit/{fac8991,ba4e99c} 17 https://github.com/Netflix/zuul/commit/{fedd975,8eb25de}
19
can also be considered an operation on a small object. This can be explained by the fact that small objects tend to be continuously modified in multiple consecutive commits. The high appearance ratios of Merge Method, Split Class, and Split Package can be attributed to the fact that these refactorings have context-sensitive detection criteria; their detection criteria are strict and can be easily disrupted by code changes during squashing. As a result, they are more likely to be detected as EPRs, which contributes to their high appearance ratios. Figure 17: Frequency of EPRs.
6.9.4. RQ5c : How are EPRs typically introduced in the development process? Study Design. The first author manually reviewed and categorized 50 randomly selected EPRs detected in the experiment to investigate the cause of EPRs. We set the size to 50 based on the balance of our experiment needs and human effort. The numbers of EPRs at granularity levels 2, 3, 4, and 5 are 22, 16, 7, and 5, respectively. The distribution aligns with our observation in RQ5a , where the frequency of EPRs decreases as the granularity level increases. The reviewed result is also included in our supplemental package [39]. Results and Discussions. We divided the causes of EPRs into three types: Disruption, Absorption, and Revert. The Disruption refers to scenarios where a non-refactoring change violates the match condition of the detection criteria of a refactoring, thereby preventing its identification in the CGC. For example, in the repository retrolambda, commit 097155620 , renames the class BytecodeFileVisitor to ClasspathVisitor, which is detected as a Rename Class refactoring. However, class BytecodeFileVisitor was newly introduced in its parent commit 46b0d8421 . When analyzing CGC squashed by these two commits, the code change is that class ClasspathVisitor is newly introduced. The class name BytecodeFileVisitor is hidden. As a result, the refactoring Rename Class renaming the BytecodeFileVisitor to ClasspathVisitor is detected as an EPR. The Absorption occurs when multiple refactorings are applied to the same object as part of a refactoring process. In the CGC, some of these refactorings are absorbed by the overall process and become undetectable. For instance, in the repository mbassador, commit e7a7628 moved the class SubscriptionContext from package mbassy.dispatch to mbassy.subscription which is detected as a Move Class refactoring. In its parent commit eb50dbc22 , the same class SubscriptionContext is renamed from MessageContext, where a Rename Class refactoring is detected. When analyzing from the CGC squashed by these two commits, only the Move Class refactoring of moving MessageContext from mbassy.dispatch to mbassy.subscription is detected. The Move Class of moving SubscriptionContext and the Rename Class refactoring of renaming MessageContext to
where OGR(C) represents the refactorings detected in C. Similar to CGRs, we compute the accumulated frequency for EPR at granularity ℓ as the cumulative frequency of EPRs at granularity levels up to and including ℓ. For all repositories in our dataset, we calculated the frequencies of EPRs detected at granularity levels of ℓ ∈ {2, 3, 4, 5}. Results and Discussions. The frequencies of EPRs detected at granularity levels of ℓ ∈ {2, 3, 4, 5} across 32 repositories are depicted in Figure 17. The x-axis represents the granularity levels, while the y-axis measures the frequency of EPR occurrences. The boxplots illustrate the distribution of EPR frequencies and accumulate frequencies at each granularity. The median values of the frequencies for granularity levels of ℓ ∈ {2, 3, 4, 5} are 0.0490, 0.0273, 0.0141, and 0.0108, respectively. The left four boxes indicate a gradual decrease in EPR frequencies as the granularity level increases from 2 to 5. Additionally, the average value of accumulate frequency reaches 0.0989 when the granularity level is 5, which indicates that per refactoring conducted, there are also 0.0989 EPR be performed. As introduced in Section 6.4.2 and shown in Figure 10, EPRs are detected from all repositories, indicating EPR is a common phenomenon. We conclude that EPR is a common phenomenon among all repositories, and it appears most frequently at the granularity level of 2. 6.9.3. RQ5b : What are the most common types of EPRs? Study Design. The ratio of a specific EPR type t at a granularity level ℓ in commit history H can be expressed as follows: S |{r ∈ c∈H EPRℓ (c) | r.type = t}| AppearanceEPR (H, ℓ) = t |{r ∈ OGR(H) | r.type = t}| We calculated the ratio for each EPR type detected in our dataset. Results and Discussions. The top ten EPR types in terms of their appearance ratio for ℓ ∈ {2, 3, 4, 5} are listed in Table 5. The highest appearance ratio for the four granularity levels are Modify Parameter Annotation, Merge Method, Remove Parameter Modifier, and Split Variable, respectively. Except for Merge Method, Rename Package, and Replace Anonymous with Class, all the other top 3 EPR types across all granularities are operations on small objects such as variables, attributes, and parameters. Since Rename Package targets the name string, it
20 https://github.com/luontola/retrolambda/commit/0971556 21 https://github.com/luontola/retrolambda/commit/46b0d84 22 https://github.com/bennidi/mbassador/commit/eb50dbc
20
Table 5: Most flequently found EPR types and their appearance ratio Rank 1 2 3 4 5 6 7 8 9 10
Granularity level ℓ=4
ℓ=2
ℓ=3
Modify Parameter Annotation
Merge Method
(19.33%)
(20.83%)
(17.06%)
(5.26%)
(38.66%)
Replace Anonymous with Class
Modify Parameter Annotation
Add Parameter Modifier
Inline Attribute
Rename Package
(15.57%)
(14.29%)
(12.89%)
(5.00%)
(25.12%)
Split Package
Parameterize Attribute
Remove Variable Modifier
Split Class
Replace Anonymous with Class
(14.81%)
(10.38%)
(12.49%)
(4.69%)
(22.16%)
Rename Package
Merge Class
Replace Attribute
Replace Variable with Attribute
Split Class
(11.82%)
(8.00%)
(9.09%)
(3.06%)
(21.88%)
Split Class
Modify Method Annotation
Split Class
Split Method
Merge Method
(10.94%)
(7.50%)
(6.25%)
(3.03%)
(20.83%)
Merge Class
Modify Variable Annotation
Move Package
Merge Package
Replace Attribute
(10.00%)
(6.06%)
(6.06%)
(3.03%)
(20.00%)
Merge Variable
Replace Attribute
Rename Package
Split Parameter
Move Package
(9.68%)
(5.45%)
(5.91%)
(2.88%)
(19.70%)
Move and Rename Method
Rename Package
Modify Parameter Annotation
Move Method
Split Package
(9.38%)
(5.42%)
(5.04%)
(2.28%)
(18.52%)
Move Attribute
Localize Parameter
Merge Variable
Move Package
Remove Parameter Modifier
(8.27%)
(5.35%)
(4.84%)
(2.27%)
(18.43%)
Merge Conditional
Inline Attribute
Split Package
Inline Method
Remove Variable Modifier
(7.95%)
(5.00%)
(3.70%)
(2.23%)
(18.31%)
Remove Parameter Modifier
ℓ=5
cross-granularity
Split Variable
Modify Parameter Annotation
7. Threats to Validity
SubscriptionContext are no longer detected, making them EPRs.
In this section, we discuss four types of potential threats to the validity of our work as follows:
The Revert refers to a scenario where a refactoring is introduced in one commit but later reverted in a subsequent commit. An example from the repository android-async-http illustrates this. In commit 846c83123 , a Change Variable Type refactoring which changed a variable’s type from Random to SecureRandom is detected. However, in its subsequent commit, another Change Variable Type refactoring which reverts this change, restoring the original type is detected. In the CGC squashed by the above two commits, the variable type remains SecureRandom. As a result, both Change Variable Type refactorings are detected as EPRs.
7.1. Internal Validity One of the possible threats to internal validity is that the repositories selected as datasets are limited in a specific field, which may affect the analysis result. We mitigated it by selecting 32 repositories from different fields and different companies or organizations. A second potential threat is the use of git-blame for tracking program elements introduced in Section 4.2. Since git-blame is inaccurate [36, 37], it may affect the correctness of CGR/EPR detection. Another threat to internal validity is that in Section 4.2, when selecting the refactored element to trace in the refactoring detection results, we chose the first code element when multiple elements were present in the detection result based on our observation for sampled refactoring instances that the first element acts as the main role of refactorings. However, this choice may not be always correct. The first author manually reviewed all 99 refactoring types detected in this study and found that only 7 might be problematic, as the first element may not always be the primary refactored element. Instances of these types account for only 0.39% of all detected original refactorings and should not significantly impact the overall findings. Therefore, while this limitation exists, its effect on the validity of our results is minimal. An additional threat is that in answering RQ4 , we analyzed commit messages to assess whether they suggested the existence of CGRs. However, due to the lack of standardized conventions, commit messages are often inconsistent, incomplete, or ambiguous, making them an unreliable proxy for developer intent. To mitigate this, we reviewed each message alongside the corresponding code changes and incrementally defined the category of characteristics. While this approach improves consistency, we acknowledge the inherent subjectivity. To support
Among the 50 manually reviewed EPR instances, 74% (38/50) are classified as Disruption, 18% (9/50) as Absorption, and 6% (3/50) as Revert. Notably, Revert EPRs are explicitly mentioned in commit messages, as developers often indicate when a change has been reverted. In our review, 2 out of 3 Revert-type EPRs were explicitly mentioned in the commit messages. The median values of frequencies of EPR detected at granularity of 2, 3, 4, and 5 are 0.0490, 0.0273, 0.0141, and 0.0108, respectively. On average, per refactoring conducted, 0.0989 EPRs will also be performed. By analyzing the types of the detected EPR, we concluded that refactoring operations on small objects such as variables, attributes, and parameters, and refactorings with context-sensitive detection criteria are often associated with EPRs. The cause of EPRs is divided into three types: Disruption, Absorption, and Revert. The Revert type EPRs are usually explicitly mentioned in the commit message.
23 https://github.com/android-async-http/android-async-http/commit/846c8
31
21
8. Implications for Researchers and Practitioners
better traceability, we recommend that developers explicitly annotate CGRs in commit messages, as noted in the Section 8.3. A further threat is that the manual analysis was conducted solely by the first author, which could introduce bias. To minimize the potential bias, we tried to follow a systematic process for the manual analysis. We conducted the classification of the cause of CGRs and EPRs through an incremental analysis process; we examined the composition of each instance, identified recurring patterns, and updated the categorization when a new type emerged. Additionally, to improve transparency and reproducibility, the full classification results are included in the supplemental package [39] so that interested readers can confirm the possibility of misclassified results.
Our investigations towards CGRs and EPRs offer several insights that are relevant to refactoring tool builders, empirical software engineering researchers, and software practitioners. In this section, we discuss how our findings can inform and guide these communities. 8.1. Implications for Refactoring Tool Builders Our investigation of the existence of CGR suggests that the current refactoring tools can be further developed to achieve better performance. • Refactoring detector update. Existing modern refactoring detectors typically operate at the granularity of a single commit. Our results suggest that refactoring detectors should support detection at multiple levels of granularity, including within a single commit, across multiple commits, and across entire versions, in order to more accurately capture the full scope of refactoring activities. For example, RefactoringMiner has a feature for detecting refactorings across a commit range24 , inspired by our previous work [14]. As an advanced detection way, one possible approach to enable CGR detection is as follows: if the match conditions for identifying a refactoring are only partially satisfied in a given commit, the tool could analyze adjacent commits to determine whether the remaining conditions are satisfied in the whole change, including the adjacent commits, thereby enabling detection across commit boundaries.
7.2. External Validity One limitation impacting external validity is that although we selected 32 repositories, all of them are programmed in Java. We cannot claim that the same results can be generalized to other programming languages. 7.3. Construct Validity A possible threat to the construct validity is that we use only one refactoring detector, RefactoringMiner, in this study. Though it is state-of-art with high accuracy and recall, it may misidentify or miss some refactorings.
• Tool integration. By identifying CGRs and EPRs, integrated development environments (IDEs) and continuous integration (CI) tools can provide more meaningful feedback, such as revealing grouped changes that represent a higher-level design intent or reverted changes that can be regarded as misoperations, thereby improving the user experience during code inspection and review.
7.4. Conclusion Validity A potential challenge to conclusion validity is that in Section 6.8, we deduced the developers’ intentions from commit messages for analysis. Without directly consulting with those developers, our deduction may be inaccurate or incorrect. Another threat is the limited sample size used in our manual analysis. We manually reviewed 110 CGRs for RQ3 and RQ4 . For RQ3 , we proposed two types of causes for CGR, and we consider this catalog as complete according to the definition of the causes. Therefore, we do not regard the validity of the classification itself as a threat. However, due to the limited sample size, the reported ratios of each type may not accurately reflect their true distribution in the dataset. In answering RQ4 , we analyzed commit messages associated with the CGRs to understand developer intentions. This analysis may also be affected by the limited number of CGRs sampled, and other characteristics may exist. Similarly, we manually reviewed 50 EPRs for RQ5 . While the three types of causes for EPRs are considered complete based on their definitions, the distribution of each type may be affected by the limited sample size. In addition, the distribution of the 110 CGRs across granularity levels may also impact our conclusion validity. As noted in Section 6.7.2, although we maintain a reasonably balanced sample to reduce potential bias, the obtained distribution might be different from the true distribution in the wild.
• CGR suggestions. CGRs reflect how developers conduct complicated design changes in practice. Tools can identify frequent CGR types and leverage their patterns to recommend multi-step refactorings that align more naturally with developer workflows. 8.2. Implications for Empirical Researchers Our findings have methodological and conceptual implications for the empirical study of refactoring and software evolution. • Misalignment between change recording and refactoring granularity. Our study highlights a fundamental gap between how code changes are recorded (e.g., as individual commits or pull requests) and the actual granularity of refactorings conducted by developers. The commit-level refactoring analyses may systematically underestimate or misrepresent real-world refactoring practices. 24 https://github.com/tsantalis/RefactoringMiner/tree/f4a2076#with-a-com mit-range
22
• Dataset construction. The dataset of CGRs and EPRs proposed in this study provides a foundation for more realistic analyses of refactoring in practice. Future research can leverage this dataset to evaluate multi-commit refactoring detection approaches, study temporal patterns in code restructuring.
future research. Another complementary direction is to examine the factors that influence the occurrence of CGRs across different projects as we discussed in Section 6.5. Acknowledgments This work was partly supported by JST SPRING No. JPMJSP2106 and JSPS Grants-in-Aid for Scientific Research Nos. JP23K24823, JP25K03102, JP25H01125, JP24H00692, JP21H04877, JP21K18302, and JP21KK0179.
8.3. Implications for Software Practitioners Practitioners, including developers and reviewers, can benefit from an improved understanding of CGRs and EPRs: • Explicit commit message annotation. Developers can improve the clarity and traceability of CGRs and EPRs by explicitly indicating their presence and intent in their commit messages. Such annotations facilitate better understanding during code reviews and inspections.
Declaration of generative AI and AI-assisted technologies in the writing process During the preparation of this work the authors used ChatGPT in order to improve readability and language of the work. After using this tool/service, the authors reviewed and edited the content as needed and take full responsibility for the content of the publication.
• Improved reviewability. Awareness of CGRs and EPRs allows code reviewers to better grasp the overarching refactoring intent, which may be difficult to infer when changes are fragmented across individual commits. In addition, such awareness facilitates understanding both individual changes and the broader evolution path of the software.
References [1] M. Fowler, Refactoring: Improving the Design of Existing Code, 2nd Edition, Addison-Wesley Professional, Boston, MA, 2018. [2] N. Tsantalis, M. Mansouri, L. M. Eshkevari, D. Mazinanian, D. Dig, Accurate and efficient refactoring detection in commit history, in: Proceedings of the 40th International Conference on Software Engineering, 2018, pp. 483–494. [3] M. Kim, T. Zimmermann, N. Nagappan, A field study of refactoring challenges and benefits, in: Proceedings of the ACM SIGSOFT 20th International Symposium on the Foundations of Software Engineering, 2012, pp. 1–11. [4] M. Kim, T. Zimmermann, N. Nagappan, An empirical study of refactoring challenges and benefits at microsoft, IEEE Transactions on Software Engineering 40 (7) (2014) 633–649. [5] G. Bavota, B. De Carluccio, A. De Lucia, M. Di Penta, R. Oliveto, O. Strollo, When does a refactoring induce bugs? an empirical study, in: Proceedings of the IEEE 12th International Working Conference on Source Code Analysis and Manipulation, 2012, pp. 104–113. [6] D. Dig, C. Comertoglu, D. Marinov, R. Johnson, Automated detection of refactorings in evolving components, in: Proceedings of the 20th ECOOP Object-Oriented Programming, Springer, 2006, pp. 404–428. [7] M. Kim, M. Gee, A. Loh, N. Rachatasumrit, Ref-finder: a refactoring reconstruction tool based on logic query templates, in: Proceedings of the 18th ACM SIGSOFT International Symposium on Foundations of Software Engineering, 2010, pp. 371–372. [8] K. Prete, N. Rachatasumrit, N. Sudan, M. Kim, Template-based reconstruction of complex refactorings, in: Proceedings of the 26th IEEE International Conference on Software Maintenance, 2010, pp. 1–10. [9] P. Weißgerber, S. Diehl, Identifying refactorings from source-code changes, in: Processding of the 21st IEEE/ACM International Conference on Automated Software Engineering, 2006, pp. 231–240. [10] D. Silva, M. T. Valente, RefDiff: Detecting refactorings in version histories, in: Proceedings of the 14th IEEE/ACM International Conference on Mining Software Repositories, 2017, pp. 269–279. [11] D. Silva, J. P. da Silva, G. Santos, R. Terra, M. T. Valente, RefDiff 2.0: A multi-language refactoring detection tool, IEEE Transactions on Software Engineering 47 (12) (2020) 2786–2802. [12] N. Tsantalis, A. Ketkar, D. Dig, RefactoringMiner 2.0, IEEE Transactions on Software Engineering 48 (3) (2022) 930–950. [13] L. Tan, C. Bockisch, A survey of refactoring detection tools, in: Proceedings of the Workshops of the Software Engineering Conference, CEURWS, Vol. 2308, 2019, pp. 100–105. [14] L. Chen, S. Hayashi, Impact of change granularity in refactoring detection, in: Proceedings of the 30th IEEE/ACM International Conference on Program Comprehension, 2022, pp. 565–569.
9. Conclusion and Future Work In this work, we investigated the impact of refactoring detection on different granularities of commits in 32 open-source Gitbased Java repositories. The granularity change is conducted by integrating the code changes in multiple commits into a single commit. We named the two types of special refactorings that can only be identified after granularity change as CGR and EPR. Through the experiments, we observed that the appearance of CGRs is common, and refactoring types related to splitting or merging classes and packages, as well as those involving modifications to the inheritance structure tend to be CGRs. The cause of CGR is divided into two types according to its composition. In addition, we also found some developers explicitly suggest the existence of them in the commit message. The appearance of EPR is also common in development, and refactoring operations on small objects such as variables, attributes, and parameters, and refactorings with context-sensitive detection criteria tend to be associated with EPRs. We discussed our implications for both researchers and practitioners, which suggests the value of CGRs and EPRs. For future work, a promising direction is to investigate how the integration path of commits influences the occurrence of CGRs and EPRs. Some intermediate states in refactorings may fail to compile or pass CI checks and, therefore, are likely to appear only in pull request workflows where CI is executed on the final merged state. This can help uncover the strategies developers adopt to perform complex refactorings under different integration workflows. Additionally, empirical studies that survey how developers perceive CGRs and EPRs and how these affect software engineering tasks such as code review, debugging, and maintenance, represent another valuable direction for 23
[15] A. C. Bibiano, D. Coutinho, A. Uchôa, W. K. Assunçao, A. Garcia, R. de Mello, T. E. Colanzi, D. Tenório, A. Vasconcelos, B. Fonseca, et al., Enhancing recommendations of composite refactorings based on the practice, in: Proceedings of the 24th IEEE International Conference on Source Code Analysis and Manipulation, IEEE, 2024, pp. 83–93. [16] M. Mongiovi, R. Gheyi, G. Soares, L. Teixeira, P. Borba, Making refactoring safer through impact analysis, Science of Computer Programming 93 (2014) 39–64. [17] E. Choi, K. Fujiwara, N. Yoshida, S. Hayashi, A survey of refactoring detection techniques based on change history analysis, Computer Software 32 (1) (2015) 47–59, https://arxiv.org/pdf/1808.02320. [18] A. Z. Broder, On the resemblance and containment of documents, in: Proceedings of the Compression and Complexity of SEQUENCES, IEEE, 1997, pp. 21–29. [19] K. Herzig, A. Zeller, The impact of tangled code changes, in: Proceedings of the 10th Working Conference on Mining Software Repositories, 2013, pp. 121–130. [20] J. Matsuda, S. Hayashi, M. Saeki, Hierarchical categorization of edit operations for separately committing large refactoring results, in: Proceedings of the 14th International Workshop on Principles of Software Evolution, 2015, pp. 19–27. [21] S. Sothornprapakorn, S. Hayashi, M. Saeki, Visualizing a tangled change for supporting its decomposition and commit construction, in: Proceedings of the 42nd IEEE Annual Computer Software and Applications Conference, Vol. 1, 2018, pp. 74–79. [22] X. Ge, S. Sarkar, E. Murphy-Hill, Towards refactoring-aware code review, in: Proceedings of the 7th International Workshop on Cooperative and Human Aspects of Software Engineering, 2014, pp. 99–102. [23] X. Ge, S. Sarkar, J. Witschey, E. Murphy-Hill, Refactoring-aware code review, in: Proceedings of the IEEE Symposium on Visual Languages and Human-Centric Computing, 2017, pp. 71–79. [24] E. L. Alves, M. Song, M. Kim, RefDistiller: A refactoring aware code review tool for inspecting manual refactoring edits, in: Proceedings of the 22nd acm sigsoft international symposium on foundations of software engineering, 2014, pp. 751–754. [25] C. Gorg, P. Weißgerber, Detecting and visualizing refactorings from software archives, in: Proceedings of the 13th International Workshop on Program Comprehension, 2005, pp. 205–214. [26] S. Tsutsumi, E. Choi, N. Yoshida, K. Inoue, Graph-based approach for detecting impure refactoring from version commits, in: Proceedings of the 1st International Workshop on Software Refactoring, 2016, pp. 13– 16. [27] L. Sousa, D. Cedrim, A. Garcia, W. Oizumi, A. C. Bibiano, D. Oliveira, M. Kim, A. Oliveira, Characterizing and identifying composite refactorings: Concepts, heuristics and patterns, in: Proceedings of the 17th International Conference on Mining Software Repositories, 2020, pp. 186– 197. [28] G. Bavota, A. De Lucia, M. Di Penta, R. Oliveto, F. Palomba, An experimental investigation on the innate relationship between quality and refactoring, Journal of Systems and Software 107 (2015) 1–14. [29] A. C. Bibiano, E. Fernandes, D. Oliveira, A. Garcia, M. Kalinowski, B. Fonseca, R. Oliveira, A. Oliveira, D. Cedrim, A quantitative study on characteristics and effect of batch refactoring on code smells, in: Pro-
ceedings of the 13th ACM/IEEE International Symposium on Empirical Software Engineering and Measurement, 2019, pp. 1–11. [30] D. Cedrim, A. Garcia, M. Mongiovi, R. Gheyi, L. Sousa, R. De Mello, B. Fonseca, M. Ribeiro, A. Chávez, Understanding the impact of refactoring on smells: A longitudinal study of 23 software projects, in: Proceedings of the 11th Joint Meeting on Foundations of Software Engineering, 2017, pp. 465–475. [31] D. Silva, N. Tsantalis, M. T. Valente, Why we refactor? confessions of GitHub contributors, in: Proceedings of the 24th ACM SIGSOFT International Symposium on Foundations of Software Engineering, 2016, pp. 858–870. [32] A. C. Bibiano, W. K. Assunção, D. Coutinho, K. Santos, V. Soares, R. Gheyi, A. Garcia, B. Fonseca, M. Ribeiro, D. Oliveira, et al., Look ahead! revealing complete composite refactorings and their smelliness effects, in: Proceedings of the 37th IEEE International Conference on Software Maintenance and Evolution, 2021, pp. 298–308. [33] T. Saika, E. Choi, N. Yoshida, A. Goto, S. Haruna, K. Inoue, What kinds of refactorings are co-occurred? an analysis of Eclipse usage datasets, in: Proceedings of the 6th International Workshop on Empirical Software Engineering in Practice, 2014, pp. 31–36. [34] P. Avgustinov, A. I. Baars, A. S. Henriksen, G. Lavender, G. Menzel, O. De Moor, M. Schafer, J. Tibble, Tracking static analysis violations over time to capture developer characteristics, in: Proceedings of the 37th IEEE/ACM IEEE International Conference on Software Engineering, 2015, pp. 437–447. [35] Q. Hanam, L. Tan, R. Holmes, P. Lam, Finding patterns in static analysis alerts: improving actionable alert ranking, in: Proceedings of the 11th Working Conference on Mining Software Repositories, 2014, pp. 152– 161. [36] M. Jodavi, N. Tsantalis, Accurate method and variable tracking in commit history, in: Proceedings of the 30th ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, 2022, pp. 183–195. [37] M. T. Hasan, N. Tsantalis, P. Alikhanifard, Refactoring-aware block tracking in commit history, IEEE Transactions on Software Engineering 50 (12) (2024) 3330–3350. [38] S. Shiba, S. Hayashi, Historinc: A repository transformation tool for finegrained history tracking (in Japanese), Computer Software 39 (4) (2022) 75–85. [39] L. Chen, S. Hayashi, Supplementary package for an empirical study on the impact of change granularity in refactoring detection, figshare, https: //doi.org/10.6084/m9.figshare.29648675 (2025). [40] S. R. Foster, W. G. Griswold, S. Lerner, WitchDoctor: IDE support for real-time auto-completion of refactorings, in: Proceedings of the 34th International Conference on Software Engineering, 2012, pp. 222–232. [41] X. Ge, E. Murphy-Hill, Manual refactoring changes with automated refactoring validation, in: Proceedings of the 36th International Conference on Software Engineering, 2014, pp. 1095–1105. [42] E. AlOmar, M. W. Mkaouer, A. Ouni, Can refactoring be self-affirmed? an exploratory study on how developers document their refactoring activities in commit messages, in: Proceedings of the 3rd IEEE/ACM International Workshop on Refactoring, 2019, pp. 51–58.
24