On the Footprints of Reviewer Bots’ Feedback on Agentic Pull Requests in OSS GitHub Repositories Syeda Kaneez Fatima
[email protected] Lahore University of Management Sciences Lahore, Punjab, Pakistan
arXiv:2604.24450v1 [cs.SE] 27 Apr 2026
Amelia Nawaz
[email protected] Lahore University of Management Sciences Lahore, Punjab, Pakistan
Yousuf Abrar
Abdul Rehman Tahir
[email protected] Lahore University of Management Sciences Lahore, Punjab, Pakistan
[email protected] Lahore University of Management Sciences Lahore, Punjab, Pakistan
Shamsa Abid
Abdul Ali Bangash
[email protected] Lahore University of Management Sciences Lahore, Punjab, Pakistan
[email protected] Lahore University of Management Sciences Lahore, Punjab, Pakistan
Abstract
1
Autonomous coding agents are reshaping software development by creating pull requests (PRs) on GitHub, referred to as agentic PRs. In parallel, the review process is also becoming autonomous, thereby making reviewer bots key actors in the assessment of these agentic PRs. However, their influence on PR acceptance and resolution remains unclear. This study empirically investigates the relationship between reviewer-bot feedback and PR outcomes by analyzing how Reviewer Bot Feedback Quality (relevance, clarity, conciseness) and Reviewer Bot Activity Volume (comment count) are associated with PR acceptance and resolution time. We analyze 7, 416 reviewerbot comments on 4, 532 PRs from the AI_Dev dataset (a dataset that captured AI agents’ PRs in GitHub projects). Our results show that reviewer-bot comments mainly focus on bug fixes, testing, and documentation, are civil in tone, and are prescriptive in nature. Reviewer bots generally produce clear and concise feedback, though the semantic relevance of comments to underlying code changes is moderate. We find that higher Reviewer Bot Activity volume is associated with longer PR resolution times and lower average feedback quality, showing that as bots generate more comments on a PR, the average pertinence of that feedback appears to degrade. At the same time, Reviewer Bot Feedback Quality shows no meaningful association with workflow outcomes. Our findings suggest that, in agentic PR workflows, reviewer bots should prioritize targeted highrelevance feedback over generating large numbers of comments.
Autonomous coding agents introduce a distinct class of pull requests (PRs) in which code is generated or modified entirely by AI systems [20]. We refer to these contributions as agentic pull requests (agentic PRs). In parallel, autonomous systems increasingly support the code review process by automatically generating feedback on submitted changes [4, 10]. These systems, commonly termed as automated code review agents [11] or reviewer bots [19], are now frequently involved in reviewing agentic PRs [2, 4], contributing to the broader automation of modern software workflows [1, 5, 26]. While prior work has examined human perceptions of AI-generated code [7, 14], the interaction between reviewer bots and agentic PRs remains underexplored [20]. In particular, there is limited empirical evidence on how reviewer bot communication characteristics affect workflow outcomes. We argue that understanding the influence of Bot Feedback Quality (defined in terms of relevance, clarity, and conciseness), and Bot Activity Volume (measured as the number of bot comments per PR), is critical for evaluating review effectiveness, developer experience, and CI/CD pipeline reliability [17]. This paper presents a quantitative and correlational analysis of reviewer-bot comments on agentic PRs using review data from the AI_Dev dataset [10]. We analyze the types of issues highlighted by bots, the nature and civility of their comments, and assess feedback quality using the categorization and scoring framework proposed by Sghaier et al. [15] (detailed in Section 3.2). We then empirically examine how these communication characteristics relate to PR outcomes. Specifically, we test whether higher Reviewer Bot Feedback Quality is associated with reduced PR resolution time and higher acceptance rates, and contrast this effect with that of Reviewer Bot Activity Volume [1]. Accordingly, we address the following research questions: RQ1: How can reviewer bot comments on agentic pull requests be characterized in terms of type, nature, civility, and Bot Feedback Quality (relevance, clarity, and conciseness)? RQ2: To what extent do these variables, specifically the Bot Feedback Quality and the Bot Activity Volume on a pull request, correlate with PR resolution time and acceptance rate? Replication Kit: Our replication kit is available at [18].
CCS Concepts • Software engineering → Agentic software engineering.
Keywords Agentic pull requests, agentic reviews, reviewer bots, automated code reviews
This work is licensed under a Creative Commons Attribution 4.0 International License. MSR ’26, Rio de Janeiro, Brazil © 2026 Copyright held by the owner/author(s). ACM ISBN 979-8-4007-2474-9/2026/04 https://doi.org/10.1145/3793302.3793599
Introduction
MSR ’26, April 13–14, 2026, Rio de Janeiro, Brazil
2
Syeda Kaneez Fatima, Yousuf Abrar, Abdul Rehman Tahir, Amelia Nawaz, Shamsa Abid, and Abdul Ali Bangash
Related Work
Research on automated code review has expanded rapidly, with recent surveys documenting the growing use of large language models (LLMs) in software engineering tasks [24]. Prior studies, including those by Schäfer et al. [14] and Tufano et al. [19], evaluate automation for activities such as test generation and code review. While Watanabe et al. [20], and Wessel et al. [22, 23] identified this shift; Li et al. [10] released the AI_Dev dataset to study it at scale, existing work focuses largely on coarse-grained metrics such as turnaround time and explicitly notes that the dynamics of automated review on agent-generated code remain unexplored. We address this gap by systematically profiling bot comment quality (RQ1) and assessing their correlation with PR resolution and acceptance rates (RQ2). Unlike prior work [7, 13, 21] that evaluates automated reviewers primarily in terms of their utility to human developers or treats bot feedback as a black box, we quantify specific interaction characteristics, i.e., relevance, clarity, and conciseness, and empirically test whether feedback quality or activity volume drives workflow delays.
3 Methodology 3.1 Dataset We use the AI_Dev dataset by Li et al. [10], which is a large-scale empirical corpus of AI based PRs, review comments, and associated metadata collected from open-source software (OSS) projects on GitHub. We analyze 7, 416 reviewer-bot comments from the AI_Dev dataset. Each record contains the bot-authored review comment and the associated diff_hunk (the specific block of modified code). After we pre-process the data and filter out non-bot comments, the final dataset contains 7, 416 reviewer-bot comments generated by 29 distinct bots across 4, 532 agentic PRs.
3.2
Categorization and Scoring Framework
To identify the interaction patterns of reviewer bots with agentic PRs (RQ1), we adopt the framework by Sghaier et al. [15] and use it to evaluate the characteristics and quality of reviewer-bot comments. Their framework classifies comments along three Categorical dimensions: Type, Nature, and Civility, and scores each subcategory according to three Scoring Criteria Metrics: Relevance, Clarity, and Conciseness. Among the categories, Type captures focus (e.g., Bugfix, Testing, Documentation), Nature captures expression style (e.g., Prescriptive, Clarification), and Civility captures tone (Civil or Uncivil) of the comment. Similarly, in the scoring framework, Relevance measures pertinence to the code change, Clarity measures how clearly the comment conveys its message, and Conciseness measures brevity and efficiency, on a scale of 1-10. A score of 1 indicates a very irrelevant, unclear, and non-concise comment, and 10 shows a highly relevant, clear, and concise review comment, respectively. To annotate reviewer-bot comments, we use GPT-5.1 as an automated annotator following the framework of Sghaier et al. [15]. The model categorizes each comment by Type, Nature, and Civility, and assigns 1–10 ordinal scores for Relevance, Clarity, and Conciseness. The annotation prompt defines all categorization and scoring criteria, includes an illustrative example, and provides full context by incorporating both the review comment text and the associated code changes (diff_hunk). We first apply this procedure to a random
subset of 200 comments, which are manually validated in parallel. After confirming the reliability of the annotation setup, we apply the automated procedure to the full corpus of 7, 132 reviewer-bot comments. The prompt template is included in our replication kit.
3.3
Manual Validation
To ensure reliability of LLM-generated categories for review comments, two authors independently annotate a random sample of 200 comments, with a third author resolving the disagreements. For categories, we evaluate inter-rater agreement between human annotators and GPT-5.1 using Cohen’s Kappa (𝜅) [3]. The 𝜅 values: Nature (0.89), Type (0.86), and Civility (1.0), indicate almost perfect agreement [9], confirming that GPT-5.1 functions as a reliable high-throughput annotator. For numeric rating dimensions (relevance, conciseness, and clarity on a 1-10 Likert scale), we assess agreement using Krippendorff’s Alpha (𝛼) with interval-level measurement, which accounts for value-based disagreement [8]. To quantify the stability of these estimates, we compute 95% confidence intervals [25]. The results show near-perfect agreement for relevance (𝛼 = 0.94) and conciseness (𝛼 = 0.95), and strong agreement for clarity (𝛼 = 0.87), with narrow confidence intervals confirming the robustness of these estimates. After finding strong inter-rater agreement, We then apply GPT-5.1 to annotate the remaining 7, 132 comments, producing an automated dataset where each comment receives a category, score, and rationale explaining the reasoning behind its classification and scoring.
3.4
Correlation Analysis with PR outcomes
To ensure a fair comparison between accepted and unaccepted PRs for our analysis (RQ2), we establish a consistent observation window to prevent survival bias, where recently submitted PRs might be incorrectly labeled as “not-merged”. To mitigate this, we define 𝑡𝑙𝑎𝑠𝑡 as the creation timestamp of the most recently accepted PR and exclude any PRs created after this point, ensuring every PR has sufficient time to be resolved. The final cohort consists of 4, 532 PRs with bot reviews, including 3, 054 accepted and 1, 478 unaccepted cases. We construct the independent variables for RQ2 as follows: 1. Bot Feedback Quality Metrics: We calculate the mean relevance, mean clarity, and mean conciseness of all individual bot comment scores, recorded against a specific PR, as in RQ1. 2. Bot Activity Volume Metrics: We measure the degree of bot engagement volume using the total number of comments on the PR. We analyze PR outcomes using a two-step non-parametric approach. PR outcomes: PR Resolution Time is defined as the duration between PR creation and acceptance for merged PRs, and PR Acceptance Rate is treated as a binary variable indicating whether a PR was merged (1) or closed without merge (0). First, to assess associations with resolution time, we compute Spearman rank correlation coefficients [16] between each independent variable and PR resolution time. It identifies monotonic relationships, specifically testing whether increased Bot Activity Volume or improved Bot Feedback Quality consistently correlate with faster or slower resolution times. Second, to assess differences in acceptance outcomes, we apply the Mann–Whitney U test [12] with 𝛼 = 0.05. This test
On the Footprints of Reviewer Bots’ Feedback on Agentic Pull Requests in OSS GitHub Repositories
4
Results
This section presents the findings of our research questions, detailing the qualitative coding of the review comments, the quantitative analysis of the distribution of review comments across categories, and the statistical examination of how Bot Feedback Quality and Bot Activity Volume correlate with PR outcomes.
4.1
RQ1: Characterizing Reviewer Bot Feedback and Activity
4.1.1 Categorization Results. Table 1 presents the categorization of bot comments by Type, Nature, and Civility, revealing distinct communication patterns. Type of review comments: Bugfix (14.0%), Testing (11.8%), and Documentation (11.6%) comments are common, while Refactoring (5.5%) and Logging (1.7%) feedback are rare. The majority of comments (55.5%) fall into the ‘Other’ category, which mainly includes configuration feedback, identified through manual inspection and GPT-5.1-generated rationales. Nature of review comments: Regarding comment nature, Prescriptive comments constitute a significant share (32.5%), while Clarification (8.6%) and Descriptive (1.1%) comments are minimal, reflecting the non-conversational structure of most bots. Consequently, the ‘Other’ category dominates (57.8%), comprising functional outputs or status notifications that do not fit into standard conversational taxonomies. Civility of review comments: Bots are overwhelmingly civil: (99.8%) of comments exhibit polite and constructive tone. Uncivil feedback is extremely rare (0.2%) but notable given the potential risks of transferring undesirable language patterns into automated ecosystems. Uncivil comments used the word ‘useless‘ for given redundancies in the diff_hunk. 1 Table 1: Overall Distribution of Bot Review Categories and Relevance, Clarity, and Conciseness scores per category across the Bot Review Comments in Agentic PRs. Category*
Subcategory*
Category Distribution
Relevance
Clarity
Conciseness
Type
Refactoring Bugfix Testing Logging Documentation Other
5.5% 14.0% 11.8% 1.7% 11.6% 55.5%
7.41 7.74 7.69 7.21 7.39 6.59
6.84 6.55 6.94 6.86 6.67 7.07
7.32 6.60 8.09 7.56 7.38 9.25
Nature
Descriptive Prescriptive Clarification Other
1.1% 32.5% 8.6% 6.54
7.41 7.31 8.15 6.88
6.48 7.21 5.94 9.33
7.48 7.78 5.22
Civility
Civil Uncivil
99.8% 0.2%
6.91 7.25
6.96 6.75
8.69 8.08
6.91
6.96
8.69
Average * From Sghaier et al.[15] 1 https://github.com/github/gh-gei/pull/1373
4.1.2 Scoring Criteria Results. Table 1 summarizes the average scores for each comment category, subcategory (i.e., Type, Nature, Civility). The overall average scores across all comments are 6.91 for relevance, 6.96 for clarity, and 8.69 for conciseness. The analysis shows that reviewer bots perform best in Conciseness, with average scores above 8 across most categories. Relevance and Clarity scores are moderate, with Clarification comments scoring lowest in conciseness (5.22) and clarity (5.94), indicating that these comments are less contextually precise and less concise. Among Type categories, Other comments achieve the highest conciseness score (9.25) but comparatively lower relevance (6.59), reflecting functional or automated outputs that are brief but not highly informative. These findings can help developers differentiate between high and moderate value comments, thereby saving their processing time on PRs. Further, bot developers can get insight into how to enhance the effectiveness of automated code reviews by improving the relevance and clarity of bot reviewers. RQ1 Summary: Reviewer-bot comments in agentic PR reviews are predominantly civil, and prescriptive, mainly targeting bug fixes, testing, and documentation. While the comments score high in conciseness (8.69), their moderate relevance (6.91) and clarity (6.96) indicate that developers may be processing many comments that are easy to read but only partially useful, highlighting the need for reviewer bots to prioritize high-value and context-aware feedback.
4.2
RQ2: What Influences PR Outcomes
4.2.1 Impact on PR Acceptance Rate. The Mann–Whitney U test with 𝛼 = 0.05 confirms that Bot Feedback Quality metrics have different distributions when comparing accepted (merged) versus unaccepted (closed without merge) PRs. It indicates that only the mean relevance of bot comments differs significantly between merged and unmerged PRs, after we apply Holm-Bonferroni correction to the p-values (𝑝 adj = 0.002) [6], suggesting that higher bot comment relevance is associated with PRs being merged. However, the influence of quality on acceptance is not straightforward, as shown in Figure 1. Specifically, Mean Relevance shows a significant negative correlation with PR acceptance (𝜌 = −0.09, 𝑎𝑑 𝑗_𝑝 = 0.0006), suggesting that higher relevance alone does not straightforwardly increase the likelihood of a PR being merged. Conversely, Mean Conciseness exhibits a strong positive correlation with PR acceptance (𝜌 = 0.06, 𝑎𝑑 𝑗_𝑝 = 0.01), indicating that concise feedback facilitates a successful merge. 1.0
PR Acceptance Rate
determines if the distributions of Bot Activity Volume and Bot Feedback Quality scores differ significantly between two independent groups of PRs, i.e., “accepted PRs”, and “unaccepted PRs”, thereby identifying bot characteristics associated with successful PR merges. We apply the Holm-Bonferroni Correction [6] ( a strong variant of Bonferroni) to the p-values of both statistical tests to control the family-wise error rate (FWER) for multiple comparisons.
MSR ’26, April 13–14, 2026, Rio de Janeiro, Brazil
0.8 0.6 0.4 0.2 0.0 2
4
6
Mean Relevance
8
10
3
4
5
6
7
8
Mean Conciseness
9
10
Figure 1: Correlations of Mean Relevance and Mean Conciseness of bot comments against PR Acceptance Rate.
MSR ’26, April 13–14, 2026, Rio de Janeiro, Brazil
Syeda Kaneez Fatima, Yousuf Abrar, Abdul Rehman Tahir, Amelia Nawaz, Shamsa Abid, and Abdul Ali Bangash
4.2.2 Impact on PR Resolution Time. Contrary to the hypothesis that high reviewer bot feedback quality would expedite workflows, Bot Feedback Quality Metrics show minimal correlation with PR resolution time. In stark contrast, increased Reviewer Bot Activity Volume, i.e., bot comments count, shows a significant positive correlation with PR resolution time, as shown in Figure 2. It interprets that PRs receiving more bot comments (𝜌 = 0.19, 𝑎𝑑 𝑗_𝑝 = 7.312𝑒 − 13) take significantly longer to complete, with p-value values obtained after applying Holm’s Bonferroni Correction.
RQ2 Summary: Bot Feedback Quality metrics show limited and inconsistent influence on agentic PR outcomes, with relevance and conciseness affecting acceptance in different ways but not accelerating resolution. In contrast, higher Bot Activity Volume consistently delays PR completion and reduces average feedback relevance. Overall, workflow efficiency in agentic PRs is driven primarily by comment quantity rather than feedback quality.
5 Resolution Time
4 3 2 1 0 2
3
3
4
4
5
6
7
Mean Relevance
8
9
10
9
10
3
4
5
6
7
Mean Clarity
8
9
10
Resolution Time
4 3 2 1 0 5
6
7
8
Mean Conciseness
0
5
10
15
20
Bot Comment Count
25
30
6
Figure 2: Correlations of Bot Feedback Quality metrics (Relevance, Clarity, Conciseness) and Bot Activity Volume (Bot Comment Count) against PR resolution time.
10
10
9
9
8
8
Mean Clarity
Mean Relevance
4.2.3 The Dilution of Review Quality: Volume vs. Quality. To understand the nature of high-volume reviewer bot activity, we analyze the relationship between Bot Activity Volume and Bot Feedback Quality. We identify a significant negative correlation between Bot Comment Count and Mean Relevance (𝜌 = −0.2, 𝑎𝑑 𝑗_𝑝 = 8.648𝑒−20), as well as Mean Clarity (𝜌 = −0.19, 𝑎𝑑 𝑗_𝑝 = 1.39𝑒 − 15). As visualized in Figure 3, this inverse relationship indicates a dilution effect: as bots generate more comments on a single PR, the average relevance of those comments declines. This suggests that high-volume bots are likely generating ‘noise’, accumulating low-relevance, and less clear messages that degrade the overall review signal.
7 6 5
7 6
4
5
3
4
2
3 0
10
20
30
40
Bot Comment Count
50
60
0
10
20
30
40
Bot Comment Count
50
60
Figure 3: Correlation between Bot Comment Count and Mean Relevance of Bot Comments.
Implications
Our results suggest that reviewer bots on agentic PRs currently function as automated pre-checkers that excel in politeness but may lack the targeted relevance needed for complex workflows. The observed inverse relationship between interaction volume and relevance implies that processing numerous automated comments may become less efficient as their average relevance decreases. This dilution effect, where high-volume bot activity generates lowrelevance messages that degrade the overall review signal, suggests that bot design could benefit from evolving toward more strategic assistance. Specifically, implementing interaction thresholds to reduce potentially redundant comments may help mitigate workflow delays. Future research could explore optimal volume thresholds while developing semantic metrics to better assess the utility of bot feedback relative to the complexity and intent of code changes.
Threats to Validity
The reviews characterization framework [15] limits construct validity by categorizing 55.5% of bot comments as ‘Other’, which may obscure specific communication subtypes. Our study constrains internal validity by relying on observable comment text and code diffs, thereby preventing analysis of internal bot logic and external developer communications. The exclusive focus on the AI_Dev dataset and the rapid evolution of agentic tools limit external validity and may affect generalizability to newer ecosystems.
7
Ethical Implications
This study does not require formal ethical approval because it does not involve human participants.
8
Conclusion
This study empirically characterizes reviewer-bot feedback in agentic PRs and shows that, while bots consistently produce civil and concise comments, their feedback is only moderately relevant and degrades as activity volume increases. Crucially, our analysis identifies a pattern where Bot Activity Volume emerges as the primary driver of longer PR resolution times. We also find that as bots generate more comments, the average relevance and clarity of their feedback decline, indicating a dilution of review signal. Reviewer bots predominantly comment on bug fixing, testing, and documentation, suggesting that their current contribution in agentic workflows aligns with lightweight pre-checking rather than targeted review. These results suggest that reviewer-bot design should emphasize limiting comment volume and prioritizing context-aware feedback generation to mitigate review noise.
On the Footprints of Reviewer Bots’ Feedback on Agentic Pull Requests in OSS GitHub Repositories
References [1] Amiangshu Bosu, Michaela Greiler, and Christian Bird. 2015. Characteristics of Useful Code Reviews: An Empirical Study at Microsoft. In 2015 IEEE/ACM 12th Working Conference on Mining Software Repositories. 146–156. doi:10.1109/MSR. 2015.21 [2] Umut Cihan, Arda İçöz, Vahid Haratian, and Eray Tüzün. 2025. Evaluating Large Language Models for Code Review. arXiv preprint arXiv:2505.20206 (2025). https://arxiv.org/abs/2505.20206 [3] Jacob Cohen. 1960. A Coefficient of Agreement for Nominal Scales. Educational and Psychological Measurement 20, 1 (1960), 37–46. doi:10.1177/ 001316446002000104 [4] Antonio Collante, Samuel Abedu, SayedHassan Khatoonabadi, Ahmad Abdellatif, Ebube Alor, and Emad Shihab. 2025. The Impact of Large Language Models (LLMs) on Code Review Process. arXiv:2508.11034 [cs.SE] https://arxiv.org/abs/ 2508.11034 [5] Mehdi Golzadeh, Alexandre Decan, Damien Legay, and Tom Mens. 2021. A ground-truth dataset and classification model for detecting bots in GitHub issue and PR comments. Journal of Systems and Software 175 (2021), 110911. doi:10. 1016/j.jss.2021.110911 [6] Winston Haynes. 2013. Holm’s Method. Springer New York, New York, NY, 902–902. doi:10.1007/978-1-4419-9863-7_1214 [7] SayedHassan Khatoonabadi, Ahmad Abdellatif, Diego Elias Costa, and Emad Shihab. 2024. Predicting the First Response Latency of Maintainers and Contributors in Pull Requests. IEEE Transactions on Software Engineering 50, 10 (2024), 2529–2543. doi:10.1109/TSE.2024.3443741 [8] Klaus Krippendorff. 1970. Estimating the Reliability, Systematic Error and Random Error of Interval Data. Educational and Psychological Measurement 30, 1 (1970), 61–70. doi:10.1177/001316447003000105 [9] J Richard Landis and Gary G. Koch. 1977. The measurement of observer agreement for categorical data. Biometrics 33 1 (1977), 159–74. https://api.semanticscholar. org/CorpusID:11077516 [10] Hao Li, Haoxiang Zhang, and Ahmed E. Hassan. 2025. The Rise of AI Teammates in Software Engineering (SE) 3.0: How Autonomous Coding Agents Are Reshaping Software Engineering. arXiv:2507.15003 [cs.SE] https://arxiv.org/abs/2507. 15003 [11] Zhiyu Li, Shuai Lu, Daya Guo, Nan Duan, Shailesh Jannu, Grant Jenks, Deep Majumder, Jared Green, Alexey Svyatkovskiy, Shengyu Fu, and Neel Sundaresan. 2022. Automating Code Review Activities by Large-Scale Pre-training. arXiv:2203.09095 [cs.SE] https://arxiv.org/abs/2203.09095 [12] Henry B. Mann and Douglas R. Whitney. 1947. On a Test of Whether one of Two Random Variables is Stochastically Larger than the Other. Annals of Mathematical Statistics 18 (1947), 50–60. https://api.semanticscholar.org/CorpusID:14328772 [13] Nivishree Palvannan and Chris Brown. 2023. Suggestion Bot: Analyzing the Impact of Automated Suggested Changes on Code Reviews. arXiv preprint arXiv:2305.06328 (2023). https://arxiv.org/abs/2305.06328 [14] Max Schäfer, Sarah Nadi, Aryaz Eghbali, and Frank Tip. 2023. An Empirical Evaluation of Using Large Language Models for Automated Unit Test Generation. arXiv:2302.06527 [cs.SE] https://arxiv.org/abs/2302.06527 [15] Oussama Ben Sghaier, Martin Weyssow, and Houari Sahraoui. 2025. Harnessing Large Language Models for Curated Code Reviews. In 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR). IEEE, 187–198. [16] C Spearman. 2010. The proof and measurement of association between two things. International Journal of Epidemiology 39, 5 (10 2010), 1137–1150. arXiv:https://academic.oup.com/ije/article-pdf/39/5/1137/18481215/dyq191.pdf doi:10.1093/ije/dyq191 [17] Kexin Sun, Hongyu Kuang, Sebastian Baltes, Xin Zhou, He Zhang, Xiaoxing Ma, Guoping Rong, Dong Shao, and Christoph Treude. 2025. Does AI Code Review Lead to Code Changes? A Case Study of GitHub Actions. arXiv preprint arXiv:2508.18771 (2025). [18] Abdul Rehman Tahir and Syeda Kaneez Fatima. 2025. On the Footprints of Reviewer Bots Feedback on Agentic Pull Requests in OSS GitHub Repositories. doi:10.5281/zenodo.17866386 [19] Rosalia Tufano, Ozren Dabić, Antonio Mastropaolo, Matteo Ciniselli, and Gabriele Bavota. 2024. Code Review Automation: Strengths and Weaknesses of the State of the Art. IEEE Transactions on Software Engineering 50, 2 (2024), 338–353. doi:10.1109/TSE.2023.3348172 [20] Miku Watanabe, Hao Li, Yutaro Kashiwa, Brittany Reid, Hajimu Iida, and Ahmed E. Hassan. 2025. On the Use of Agentic Coding: An Empirical Study of Pull Requests on GitHub. arXiv:2509.14745 [cs.SE] https://arxiv.org/abs/2509.14745 [21] Mairieli Wessel, Alexander Serebrenik, Igor Wiese, Igor Steinmacher, and Marco A. Gerosa. 2020. Effects of Adopting Code Review Bots on Pull Requests to OSS Projects. In 2020 IEEE International Conference on Software Maintenance and Evolution (ICSME). IEEE, 1–11. doi:10.1109/ICSME46990.2020.00011 [22] Mairieli Wessel, Alexander Serebrenik, Igor Wiese, Igor Steinmacher, and Marco A. Gerosa. 2020. What to Expect from Code Review Bots on GitHub? A Survey with OSS Maintainers. In Proceedings of the XXXIV Brazilian Symposium on Software Engineering (Natal, Brazil) (SBES ’20). Association for Computing
MSR ’26, April 13–14, 2026, Rio de Janeiro, Brazil
Machinery, New York, NY, USA, 457–462. doi:10.1145/3422392.3422459 [23] Mairieli Wessel, Alexander Serebrenik, Igor Wiese, Igor Steinmacher, and Marco A. Gerosa. 2022. Quality gatekeepers: investigating the effects of code review bots on pull request activities. 27, 5 (2022). doi:10.1007/s10664-022-10130-9 [24] Ratnadira Widyasari, Ting Zhang, Abir Bouraffa, Walid Maalej, and David Lo. 2024. Explaining Explanations: An Empirical Study of Explanations in Code Reviews. arXiv preprint arXiv:2311.09020 (2024). https://arxiv.org/abs/2311.09020 [25] Antonia Zapf, Stefanie Castell, Lars Morawietz, and André Karch. 2016. Measuring inter-rater reliability for nominal data - Which coefficients and confidence intervals are appropriate? BMC Medical Research Methodology 16 (08 2016). doi:10.1186/s12874-016-0200-9 [26] Zhengran Zeng, Ruikai Shi, Keke Han, Yixin Li, Kaicheng Sun, Yidong Wang, Zhuohao Yu, Rui Xie, Wei Ye, and Shikun Zhang. 2025. Benchmarking and Studying the LLM-based Code Review. arXiv preprint arXiv:2509.01494 (2025). https://arxiv.org/abs/2509.01494