arXiv:2604.11184v1 [cs.SE] 13 Apr 2026
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape BIANCA TRINKENREICH, Colorado State University, USA FABIO CALEFATO, University of Bari, Italy KELLY BLINCOE, University of Auckland, New Zealand VIGGO TELLEFSEN WIVESTAD, SINTEF Digital, Norway ANTONIO PEDRO SANTOS ALVES, PUC-Rio, Brazil JÚLIA CONDÉ ARAÚJO, Colorado State University, USA MARINA CONDÉ ARAÚJO, Colorado State University, USA PAOLO TELL, IT University of Copenhagen, Denmark MARCOS KALINOWSKI, PUC-Rio, Brazil THOMAS ZIMMERMANN, University of California Irvine, USA MARGARET-ANNE STOREY, University of Victoria, Canada Context: Software engineering (SE) researchers increasingly study Generative AI (GenAI) while also incorporating it into their own research practices. Despite rapid adoption, there is limited empirical evidence on how GenAI is used in SE research and its implications for research practices and governance. Aims: We conduct a large-scale survey of 457 SE researchers publishing in top venues (2023–2025). Method: Using quantitative and qualitative analyses, we examine who uses GenAI and why, where it is used across research activities, and how researchers perceive its benefits, opportunities, challenges, risks, and governance. Results: GenAI use is widespread, with many researchers reporting pressure to adopt and align their work with it. Usage is concentrated in writing and early-stage activities, while methodological and analytical tasks remain largely human-driven. Although productivity gains are widely perceived, concerns about trust, correctness, and regulatory uncertainty persist. Researchers highlight risks such as inaccuracies and bias, emphasize mitigation through human oversight and verification, and call for clearer governance, including guidance on responsible use and peer review. Conclusion: We provide a fine-grained, SE-specific characterization of GenAI use across research activities, along with taxonomies of GenAI use cases for research and peer review, opportunities, risks, mitigation strategies, and governance needs. These findings establish an empirical baseline for the responsible integration of GenAI into academic practice.
Authors’ addresses: Bianca Trinkenreich, [email protected], Colorado State University, Fort Collins, CO, USA; Fabio Calefato, [email protected], University of Bari, Bari, Italy; Kelly Blincoe, [email protected], University of Auckland, Auckland, New Zealand; Viggo Tellefsen Wivestad, [email protected], SINTEF Digital, Trondheim, Norway; Antonio Pedro Santos Alves, [email protected], PUC-Rio, Rio de Janeiro, Brazil; Júlia Condé Araújo, [email protected], Colorado State University, Fort Collins, CO, USA; Marina Condé Araújo, [email protected], Colorado State University, Fort Collins, CO, USA; Paolo Tell, [email protected], IT University of Copenhagen, Copenhagen, Denmark; Marcos Kalinowski, [email protected], PUC-Rio, Rio de Janeiro, Brazil; Thomas Zimmermann, [email protected], University of California Irvine, Irvine, CA, USA; Margaret-Anne Storey, [email protected], University of Victoria, Victoria, Canada.
Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than ACM must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. © 2025 Association for Computing Machinery. Manuscript submitted to ACM
Manuscript submitted to ACM
1
2
Trinkenreich et al.
ACM Reference Format: Bianca Trinkenreich, Fabio Calefato, Kelly Blincoe, Viggo Tellefsen Wivestad, Antonio Pedro Santos Alves, Júlia Condé Araújo, Marina Condé Araújo, Paolo Tell, Marcos Kalinowski, Thomas Zimmermann, and Margaret-Anne Storey. 2025. Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape. ACM Trans. Softw. Eng. Methodol. 1, 1 (December 2025), 37 pages.
1
INTRODUCTION
Generative AI (GenAI) is disrupting not only how we develop and engineer software [19], but also how we conduct research [27, 49]. Increasingly, researchers are using GenAI tools to support a wide array of research practices, including literature review [13, 20, 22, 41], coding and data analysis [5, 6, 24, 33, 34], writing and editing manuscripts [23, 31], and reviewing and summarizing research [26, 51]. GenAI is also reported to assist with research idea generation [37] and with learning how to conduct and evaluate research [27]. The adoption of GenAI by researchers mirrors the rapid uptake of GenAI by software engineers [11, 17, 38]. As many software engineering (SE) researchers shift their research agendas to studying the impact of GenAI on SE, they are simultaneously becoming users of the technology they are studying [35]. This dual role raises important, and often controversial questions about how these tools should influence research practices and research outcomes, in particular regarding the reliability, transparency, and integrity of scholarly work. Despite growing adoption, there is still limited empirical evidence about how researchers are actually using these tools and how their use may influence research practices and outcomes in software engineering [47]. Despite some early guidelines [3, 45], many of our colleagues are unsure what will be acceptable in using this technology. Developing early evidence is essential for understanding both the opportunities GenAI offers and the risks it may introduce, and for informing responsible use as well as emerging policies within academic institutions and publishing organizations. Research practices often evolve more slowly than comparable practices in industry [36], and yet researchers also face some of the same productivity pressures other software development professionals feel to adopt and use GenAI [30]. To better understand these issues and mounting concerns, we conducted a large-scale survey of software engineering researchers. Through a rigorously designed online survey instrument, we aimed to answer the following research questions: • RQ1: Who is using GenAI in SE research and what motivates their use? • RQ2: Where and how are researchers in SE using GenAI? • RQ3: What benefits, challenges, and opportunities do SE researchers perceive when using GenAI? • RQ4: How do SE researchers trust its use, what risks do they perceive from using GenAI in research, and how do they mitigate those risks? • RQ5: What regulations and policies do SE researchers feel should govern the use of GenAI in research? Our survey instrument builds on prior survey research and theories, while it also probes into specifics about how GenAI is used and how its use is perceived in SE research. The survey includes a mix of closed and open questions and was distributed to the authors of research papers published in the top venues in software engineering between 2023 to 2025. We received 457 responses from researchers with varying experience levels from institutions around the world. Using a combination of quantitative analysis and qualitative inductive and deductive analysis, we make the following contributions: Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
3
• We provide one of the first empirical characterizations of how GenAI is being used in software engineering research and which research activities are seeing the most early adoption. • We develop a taxonomy of the benefits, challenges, and opportunities SE researchers perceive from using GenAI. • We identify and surface key risks SE researchers are concerned about, and examine how researchers build trust with using GenAI in SE research. We synthesize how they mitigate those risks, such as maintaining the human in the loop, set boundaries on tool usage, and aim to improve GenAI education. • Finally, we report SE researchers’ perspectives on what regulations and policies they believe should be in place for understanding the implications of using GenAI in research and peer review. Our findings uncover critical tensions surrounding the use of GenAI in research and highlight important implications for research conducted in both academia and industry using GenAI. As GenAI models and tools continue to evolve, and as community expectations and social norms mature, longitudinal studies will be needed to understand how researchers’ practices and perceptions change over time. Our work provides a baseline for such future investigations and may help shape and guide how GenAI is used in the SE research community. 2
BACKGROUND AND RELATED WORK
In the past few years, there are several papers that explore the role of Generative AI (GenAI) and large language models (LLMs) in research, and many that specifically consider research in software engineering (SE). We organize this prior work in four clusters: surveys that investigate the adoption and perceptions of GenAI use in research; papers that discuss methodological implications and propose guidelines for GenAI use in research; the use of LLMs to support specific research tasks; and studies that explore the use of LLMs as research subjects. We begin with survey-based evidence on adoption and perceptions, which directly motivates our study. Adoption and Perceptions of GenAI for Research. The idea of using AI to augment scientific discovery predates modern generative models. Early work argued that the increasing complexity of scientific workflows creates bottlenecks that AI systems could help alleviate, particularly in tasks such as hypothesis generation and literature analysis [16]. Recent large-scale surveys provide initial empirical evidence of how researchers are adopting GenAI in practice. A global survey conducted by Nature reported that researchers (across many scientific domains) are experimenting with GenAI for idea generation, coding, and manuscript preparation [44], while also raising concerns about misinformation, plagiarism, and inaccuracies in research outputs [44]. In a survey of Danish researchers from various domains, Andersen et al. [2] examined GenAI use across stages of the research process and found that researchers perceive clear benefits for writing-related tasks but express reservations when GenAI is applied to activities requiring methodological rigor, such as experimental design. Focusing specifically on the use of GenAI in SE research at a large European research institute, Wivestad and Barbala [47] showed that LLM use is generally considered acceptable for narrow, verifiable tasks, but becomes more controversial in high-stakes contexts such as peer review. The studies also reported different adoption and perceptions across experience levels, with early-career researchers more likely to adopt GenAI [2], while more experienced researchers emphasize risks related to rigor and reproducibility [47]. Methodological Reflections and Evaluation Guidelines for SE Researchers. As GenAI becomes part of SE research workflows, questions about how to use these tools rigorously have gained attention. Prior work emphasized the need to report model versions, prompts, and configurations to enable reproducibility and comparability [3, 45]. The non-deterministic and opaque nature of LLMs makes replication hard when such details are not documented [45]. Prior Manuscript submitted to ACM
4
Trinkenreich et al.
work also highlighted that introducing LLMs in research may destabilize core research constructs, as notions such as “developer”, “artifact”, and “interaction” become blurred in settings where humans and AI systems co-create software artifacts [42]. This shift complicates the attribution of agency and what is being observed and measured. LLMs also introduce new data modalities, such as prompts and AI-generated artifacts, which raise concerns about bias, provenance, and interpretability. Furthermore, the variability of LLM outputs and the rapid evolution of models introduce evaluation drift, weakening reproducibility and causal inference [42]. Finally, the use of LLMs as research instruments raises risks of over-reliance on automated analysis and potential loss of critical human judgment. Trinkenreich et al. [43] and Williams et al. [46] further argue that efficiency gains from GenAI must be balanced with safeguards to preserve rigor and transparency. LLMs as Research Assistants in Empirical SE. Beyond methodological considerations, several studies examined how LLMs are used to support specific SE research tasks. In qualitative analysis, LLMs have been used to support coding and theme generation at scale [32, 33]. Results show that their outputs can align with human-generated codes in clarity and organization, but they may also be overly granular and fragmented when models fail to capture latent meaning and produce coherent higher-level themes [32]. Similarly, LLM-assisted analysis can lead to loss of contextual depth and premature closure of interpretation, as models favor surface-level patterns over nuanced reasoning [33]. Across both studies, LLM outputs are highly sensitive to prompt design and require iterative prompting and human validation, reinforcing their role as assistive rather than authoritative tools in qualitative analysis [32, 33]. In secondary studies, such as systematic literature reviews, LLMs have been applied to tasks including abstract screening, study selection, and data extraction, reducing manual effort and accelerating the processing of large corpora [12, 13, 20]. Empirical evaluations show that LLMs can achieve moderate to high accuracy in study selection, but still produce incorrect classifications, including false negatives that may lead to loss of relevant evidence [13]. At the same time, LLM-assisted screening has been shown as not consistently more accurate than human reviewers, with performance depending on model choice and prompting strategies [20]. Prior work also highlighted methodological challenges, including sensitivity to prompt design, limited contextual information, and lack of transparency and reproducibility due to model variability and configuration changes [12]. Across these studies, LLMs reduce effort but do not replace human judgment, requiring oversight to ensure accuracy and completeness of the review process [13, 20]. LLMs have also been explored for annotation and text processing, where model-generated annotations can be effective under controlled conditions [1], particularly when supported by validation procedures [21]. For sentiment analysis of software engineering data, large models achieve competitive performance in some scenarios, although traditional fine-tuned models may still outperform them in others [50]. LLMs as Research Subjects for SE Studies. Other studies focus on LLMs themselves, examining whether they can simulate research processes or replace human participants. Liang et al. [25] show that LLMs can reproduce aspects of research workflows, but struggle with implicit reasoning and methodological nuance. Research on synthetic participants indicated that LLM-generated responses may be useful for exploratory purposes, but differ from human data in variability, depth, and contextual grounding [15, 39]. Harding et al. [18] and De et al. [10] warned that treating synthetic outputs as equivalent to human data can distort findings by reducing variability and inflating agreement. Research Gap and Study Motivation. Prior work has brought early insights on how GenAI is being used by researchers (state of the practice) as well as how GenAI should be used and can be used by researchers (state of the art). Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
5
Our study is directly inspired by prior surveys that explored the adoption and perceptions of researchers of GenAI [2, 44, 47]. We build on these initial findings and provide an SE-specific characterization and a detailed pulse of how GenAI is being used by a broad population of SE researchers around the world. We examine not only adoption and perceptions, but also how GenAI is used across SE research strategies, stages, and activities; the benefits, challenges, opportunities, and risks researchers associate with its use; how SE researchers establish trust and mitigate risks; and how they view regulation and the use of GenAI in SE research peer review. The next section describes the survey design, recruitment strategy, and analysis procedures. 3
RESEARCH METHOD
We describe the method followed by presenting the research questions in Section 3.1, the instrument design and data collection procedures in Section 3.2, and the detailed description of the data analysis procedures in Section 3.3. Information regarding the replication package are provided in Section 3.4. 3.1
Research objective and research questions
The objective of this research is to understand how GenAI is used in software engineering (SE) research. Specifically, this study aims to characterize who among SE researchers is using GenAI, why they are motivated to adopt it, when and where in the activities of the research pipeline they adopt it, and how GenAI is incorporated into the research process. By providing an evidence-based characterization of current practices, this work seeks to inform ongoing methodological discussions on the role of GenAI in SE research and to establish a foundation for future normative and evaluative studies. For this, we address the following research questions. RQ1: Who is using GenAI in SE research and what is their motivation? This research question aims to identify who adopts GenAI providing context for interpreting patterns. To this end, RQ1.1 investigates which SE researchers use GenAI, considering demographic and professional characteristics, such as research context, years of experience, and career stage. The RQ1.2 examines the motivations that drive researchers to use GenAI, including their perception of the current and anticipated impact of GenAI, as well as their perceived pressure on using and learning GenAI to stay relevant and to link their research to it. RQ2: Where and how are researchers in SE using GenAI?. This research question focuses on the concrete use of GenAI within the SE research process and peer review. Specifically, RQ2.1 examines where, i.e., in which research strategies, methods, and stages of the research pipeline, GenAI is being employed. RQ2.2 provides a qualitative, finegrained analysis of how researchers use GenAI within these methods. Finally, RQ2.3 examines researchers’ prior use and attitudes toward the use of GenAI in the peer-review process, a particularly sensitive area with implications for research integrity. RQ3:What benefits, challenges, and opportunities do SE researchers perceive when using GenAI?. This research question investigates the perceived impact of GenAI on SE research, complementing the descriptive characterization of usage practices (RQ2) with researchers’ assessments of its consequences. Understanding perceived impacts is essential for interpreting current adoption and for anticipating how GenAI may shape future research practices. In particular, RQ3.1 examines the benefits that SE researchers perceive GenAI to offer. RQ3.2 investigates the challenges researchers experience when using GenAI in their research. RQ3.3 explores the opportunities researchers envision GenAI may offer for advancing SE research in the future. Manuscript submitted to ACM
6
Trinkenreich et al. RQ4: How do SE researchers trust its use and what risks do they perceive from using GenAI in research,
and how do they mitigate those risks? While RQ3 focuses on perceived impacts, this research question explicitly addresses the risks associated with using GenAI in SE research and how researchers reason about managing these risks. Understanding risk perception and mitigation strategies is critical for informing responsible and sustainable research practices. To this end, RQ4.1 explores how researchers trust GenAI in different contexts and with regard to each research activity, RQ4.2 examines how SE researchers perceive the risks of using GenAI for research, including concerns related to fabrication, plagiarism, quality, proliferation of misinformation, carbon footprint, and more. RQ4.3 presents the strategies proposed by SE researchers to mitigate the risks associated with using GenAI. RQ5: What regulations and policies do researchers feel should be applied to the use of GenAI in their research and reviewing activities? This research question focuses on normative and governance-related aspects of GenAI use in SE research. As GenAI adoption raises questions about appropriate boundaries and oversight, understanding researchers’ perspectives on regulation and policy is essential for informing institutional and community-level responses. We investigate, through an open question, how researchers perceive the need to regulate the use of GenAI in research. 3.2
Instrument Design
To collect empirical data for this study, we used a survey-based research design. We developed an online questionnaire 1 using Qualtrics2 to collect detailed information from software engineering researchers regarding their use of generative AI technologies. The unit of analysis in this study is the individual researcher. The set of questions comprising the questionnaire emerged from discussions during remote weekly meetings held by the authors of this paper over six months (December 2024 to June 2025). 3.2.1 Instrument Structure. After the consent form, the questionnaire included 17 questions related to our research questions, and five demographic questions for segmented analysis of the results. We used existing instruments where possible. The source for each question is presented in Table 1. 3.2.2 Data Collection and Recruitment. We employed a purposive sampling strategy [4] by recruiting authors of papers and articles published between 2023 and 2025 within a select set of leading software engineering venues: the International Conference on Software Engineering (ICSE), International Conference on Automated Software Engineering (ASE), International Conference on the Foundations of Software Engineering (FSE), IEEE Transactions on Software Engineering (TSE), ACM Transactions on Software Engineering and Methodology (TOSEM), and Springer Empirical Software Engineering (EMSE). We identified potential participants from the conference proceedings and the journal volumes, and contacted them via email with an invitation to participate in the survey. This strategy was chosen to ensure that respondents had demonstrable experience in software engineering research, aligning with the target population of the study. The survey was initially distributed during the 1𝑠𝑡 HumanAISE Workshop on Human-Centered AI for Software Engineering, held in Trondheim, Norway, on June 27, 2025. Then, three waves of email invitations were sent on July 20th, August 18th, and August 20th, 2025. Data collection was closed in September 2025, yielding 457 responses overall.
1 The research protocol was approved by the Colorado State University institutional review board (IRB). 2 http://www.qualtrics.com
Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
7
Table 1. Research questions with corresponding questionnaire items, analysis method [Quantitative|Qualitative], and reference to the source for adapted items. Research Questions Questionnaire Item. [Analysis] RQ1. Who is using GenAI in SE research and what is their motivation? RQ1.1. Who (which researchers in SE) is using GenAI for their research? Demographics [Quantitative] RQ1.2. What motivates researchers to use GenAI? How much of an impact do you think GenAI has or will have on SE Research? [Quantitative] Do you feel under pressure to use and learn about GenAI for SE Research to stay relevant? [47] [Quantitative] Do you feel under pressure to link your research to GenAI, or collaborate with AI researchers, in order to stay relevant, progress in your field, or secure funding? [47] [Quantitative] RQ2. Where and how are researchers in SE using GenAI? RQ2.1. Where are researchers in SE using GenAI (research strategy, research method, and research pipeline stage)? More specifically, please indicate if you use GenAI for each of the following research activities (columns) across the different strategies (rows) [43]. [Quantitative] RQ2.2. How are researchers in SE using GenAI? Please briefly describe how you use GenAI in your research. [Qualitative] RQ2.3. How do researchers feel about the use of GenAI in reviewing research papers? As a reviewer, have you used GenAI to assist in reviewing a SE research paper? [Qualitative] Should reviewers be allowed or encouraged to use GenAI to assist in reviewing SE research papers? [Qualitative] RQ3. What benefits, challenges, and opportunities do SE researchers perceive when using GenAI? RQ3.1. Which benefits do SE researchers perceive GenAI offers to SE Research? What benefits do you see GenAI bringing to SE Research? [44] [Quantitative] RQ3.2. What challenges do SE researchers experience using GenAI in SE research? What challenges (if any) do you, or your research team, face while using GenAI for SE Research [47]. RQ3.3. What opportunities do SE researchers envision GenAI offers to SE Research? What opportunities do you see GenAI bringing to SE Research? [Qualitative] RQ4. How do SE researchers trust its use and what risks do they perceive from using GenAI in research, and how do they mitigate those risks? RQ4.1. How do SE researchers trust GenAI for research? For the following statements, please indicate your level of agreement on trust in using GenAI for SE research in general. [Quantitative] For the following activities, I trust the use of GenAI [9]. RQ4.2. How do SE researchers perceive the risks of using GenAI for research? For the following statements, please indicate your level of agreement on the following risks that GenAI brings to SE Research? [44] [Quantitative] RQ4.3. How can SE researchers mitigate the risks while not missing the opportunities that GenAI offers? If you indicated being concerned about any risks in the previous question, how would you go about mitigating each of those risks? [Qualitative] RQ5. What regulations and policies do researchers feel should be applied to the use of GenAI in their research? Should GenAI use be regulated in Software Engineering research, assuming that it is possible? (please elaborate) [Qualitative] Final thoughts If you have final thoughts about using GenAI in SE research that might not have been covered in this questionnaire, please enter them here. [Qualitative]
3.3
Data analysis
3.3.1
Filtering.
For each analysis, we applied item-level deletion by excluding responses with missing values for the respective question. No imputation was performed; hence, the reported sample size (n) varies across analyses depending on the number of valid responses. Unless otherwise specified, all plots and statistical summaries are based on the full set of valid responses to the respective question. The only exception is the analysis of GenAI use cases (Sec. 4.2.2), which was restricted to respondents who indicated they use GenAI for research purposes, filtering out researchers who stated using GenAI for other activities. 3.3.2
Closed questions. Manuscript submitted to ACM
8
Trinkenreich et al.
Overview. We analyzed responses to closed-ended questions using descriptive statistics. Given the exploratory nature of this study and its goal of characterizing current practices and perceptions of GenAI use in SE research, we focused on summarizing the distribution of responses rather than conducting inferential statistical tests. For categorical questions (e.g., use of GenAI across activities, perceived challenges), we report absolute counts and proportions of responses. For Likert-scale items (e.g., perceived benefits, risks, and trust), we computed the distribution of responses across all scale points and report percentages to facilitate comparison across groups. To support interpretation, for Likert-scale items, we report Top-2-Box (e.g., Agree + Completely Agree) and Bottom 2 Box (e.g., Disagree + Completely Disagree) scores when appropriate, allowing us to summarize overall positive and negative perceptions. All descriptive analyses and visualizations were generated using the set of valid responses for each question after applying the filtering procedures described in the previous section. Segmented analysis. Respondents were categorized based on their responses to which activities they use GenAI for. This was a multiple-choice question that listed several possible use cases, including both research- and non-researchrelated activities. Based on the response patterns, we operationalized three mutually exclusive subgroups: (1) Researchers who have used GenAI for research purposes (n=339) (2) Researchers who have used GenAI, but not for research (n=44) (3) Researchers who have not used GenAI (n=29) This grouping enabled a segmented analysis across respondents who use GenAI in research, those who use it only outside research contexts, and those who do not use GenAI. As most respondents reported using GenAI for research, we present the main figures in the paper for this group to improve readability and focus. Corresponding figures for the other two groups are included in the online replication package. 3.3.3
Open questions.
Overview. We applied qualitative analysis to all open-ended questions. For each question, coding was conducted by one author and subsequently reviewed with three other authors, all of whom have extensive qualitative research experience. Disagreements were resolved through negotiated agreement [8] across two meetings. During this process, researchers discussed the rationale underlying each coding decision until consensus was reached [14]. Deductive and inductive analysis. For both RQ2.2 (how is GenAI being used) and RQ3.3 (opportunities), we combined thematic analysis (deductive) [7] with inductive open coding [29]. The deductive phase was grounded in the phase based framework of Andersen et al. [2], which categorizes 32 GenAI use cases across five research phases: Idea Generation, Research Design, Data Collection, Data Analysis, and Writing and Reporting. This shared framework served as the analytical scaffold for examining how GenAI use and envisioned opportunities vary across research activities. During analysis, we identified additional use cases and research phases not captured by the original framework. We therefore complemented deductive coding with inductive open coding to incorporate emergent categories. The question about how is GenAI being used (RQ2.2) was presented only to participants who indicated they use GenAI for SE research and was answered by 151 participants. Among these, 136 described research related use cases and were included in the analysis. The remaining 15 responses were excluded because they focused exclusively on teaching related uses (four responses) or described SE practice rather than SE research (for example, “Help coding Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
9
organizing documents") (11 responses). Of the 136 included responses, 71 mentioned more than one research related use case, resulting in multi label coding and a total of 243 coded use cases. The questions about using GenAI for SE peer review (RQ2.3) received 255 responses with 77 explanations about how GenAI is being used for peer review. The question about envisioned opportunities from using GenAI for SE research (RQ3.3) was answered by 131 participants. One participant referred to a previous response and reported no specific opportunity, and two provided only high-level perspectives, noting that GenAI may make information executable (R15) and enable “more probabilistic approaches" (R44). We analyzed the remaining 128 responses. Of these, 66 mentioned more than one opportunity, resulting in a total of 266 coded opportunities. We used the same deductive framework, extended with the emergent categories identified during RQ2.2. Fully inductive analysis. For RQ4.3 (risk mitigation), RQ5 (need to regulate), RQ2.3 (use in peer review), and the Final Thoughts question, we employed fully inductive open coding [29], deriving codes directly from participants’ responses. The question about suggested strategies to mitigate risks from using GenAI for SE research (RQ4.3) was answered by 158 participants. Ten indicated they did not know, eight reported not using GenAI, and two stated that risks could not be mitigated. The remaining 138 responses described at least one mitigation strategy without sacrificing the potential benefits of GenAI. The question regarding the perceived needs to regulate GenAI in SE research (RQ5) received 119 responses. The final open-ended question invited respondents to share any additional thoughts on GenAI in SE research. We performed qualitative coding on the 104 segments from 63 respondents, identifying 36 codes organized into 10 candidate themes. Since this question was unconstrained in scope, the resulting themes span multiple research questions and cannot be mapped to a single RQ. Therefore, we integrated them as corroborating qualitative evidence within the relevant RQ sections, and synthesized the cross-cutting themes in the Discussion (see Sect. 5). 3.4
Replication package
To support transparency and enable replication, we provide a comprehensive replication package publicly available on Figshare.3 The package includes: (i) the complete survey instrument, (ii) the anonymized response dataset with all personally identifiable information removed, and (iii) the codebooks for qualitative analyses. 4
RESULTS
In this section, we report our findings structured around the research questions. 4.1
Who is using GenAI in SE research, and what is their motivation? (RQ1)
In this research question, we analyze the demographic distribution of SE researchers who reported using GenAI, including gender, career stage, years in current position, organizational affiliation, and geographic distribution. We also analyze the reported motivations to use GenAI. 4.1.1 Who is using GenAI for their research? (RQ1.1). Most respondents reported using GenAI for research (74%), with few reporting using it for non-research purposes (11%) or not using it at all (7%) (see Fig. 1). Geographic distribution: From Table 2, we find the geographic composition of respondents across GenAI usage groups. Overall, the sample is largely based in Europe (47%), followed by North America (27%) and Asia (17%), with 3 https://figshare.com/s/12b873956384863a7c06
Manuscript submitted to ACM
10
Trinkenreich et al.
Q1: Do you use GenAI for any of the following activities? (n=456) 74.3%
Research 47.1%
Activity
Teaching 41.0%
Administrative tasks 16.0%
Societal dissemination 9.4%
Other (please specify) I do not use GenAI 0%
6.4% 20%
40%
Proportion Selected
60%
80%
100%
Fig. 1. GenAI Activity Distribution
Table 2. Demographics of survey respondents. n is the number of respondents in each category. Sample (%) is the percentage of respondents in that category relative to all valid responses for that demographic variable (i.e., within each demographic variable, the percentages sum to 100%). Use GenAI for Research (%) is the percentage of respondents within that category who indicated they used Generative AI for software engineering research (Q1). For Organization, respondents could select multiple options; therefore, the summed counts across its categories can exceed the totals for single-choice demographics. Demographics
Category
n
Sample (%)
Use GenAI for Research (%)
Continental Origin
Europe North America Asia South America Oceania
111 64 41 17 5
47 27 17 7 2
73 84 93 65 60
Gender
Man Woman Prefer not to say Non-Binary Other
183 63 4 3 0
72 25 2 1 0
80 73 50 67 0
Career Stage
Early-career Mid-career Advanced-career Prefer not to say
105 80 62 4
42 32 25 1
86 79 66 75
Organization
University Research institute or national lab For-profit company Government Prefer not to say Non-profit company Other - please specify Research funder
221 25 20 7 5 5 2 1
77 9 7 2 2 2 1 0
78 88 75 86 80 100 50 0
smaller proportions from South America and Oceania. Looking closer at each continent, we see that the proportion of respondents reporting having used GenAI for research, Asia is the biggest (93%), followed by North America (84%) and Europe (73%), with lower proportions (and sample size) in South America (65%) and Oceania (60%). Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
11
Fig. 2. Perceived impact (past and future) of GenAI on SE research
Gender distribution: The overall sample is predominantly composed of men (72%), followed by women (25%), with only a small proportion identifying as non-binary (1%) or preferring not to disclose the gender (2%), as seen in Table 2. Within each category, men show the highest inclination towards using GenAI for research (80%), with women slightly lower (73%). The remaining categories had lower proportions, but also a very small sample size. Career stage: Overall, as shown in Table 2, the survey respondents showed a shift toward early-career (42%), followed by mid- (32%) and advanced-career (25%). Looking at the proportion of respondents who used GenAI for their research within each career stage, we find a gradual decline as we go from early- to mid to advanced-career (86%, 79%, and 66% respectively). Organizational affiliation: As shown in Table 2, the overall sample is mostly composed of respondents affiliated with universities (77%), while smaller proportions report working in research institutes or national laboratories (9%) and for-profit companies (7%), with the rest showing negligible representation. Among these three organizational affiliations, Research institutes or national labs had the highest proportion of respondents who used GenAI for their research (88%), while universities and for-profit companies responded being somewhat less inclined (78% and 75%, respectively). 4.1.2
What motivates researchers to use GenAI? (RQ1.2).
As shown in Figure 2, most of the researchers who use GenAI for research perceive its impact as increasingly pronounced in the near to mid-term future. Specifically, 58% of respondents reported that GenAI has already had a substantial impact (Top-2-Box: a lot or a great deal), rising to 79% for the next year and peaking at 85% for the next five years, before slightly decreasing to 82% for the next ten years. Pressure related to use and learn about GenAI: Figure 3 shows how a majority of the researchers feel under pressure to link their research to GenAI, or collaborate with AI researchers (58%) and also to use and learn about GenAI for SE research (55%) in order to stay relevant, progress in the field or secure funding. Systemic pressures and persistent skepticism: The findings above are reinforced by researchers’ final thoughts in a last open-text question, which reveal that the pressure to adopt GenAI extends beyond individual motivation to systemic dynamics within academia. Some respondents described publish-or-perish incentives and funding mandates as key drivers: “a study without these things is almost automatically considered invalid” and “in the project proposal, you MUST mention something about GenAI; otherwise, you’re irrelevant” (R161). At the same time, other respondents expressed outright skepticism. One dismissed GenAI as “a solution without a problem” that “reminds me of Blockchain” Manuscript submitted to ACM
12
Trinkenreich et al.
Perceived pressure related to GenAI in SE research (n=256) Do you feel under pressure to link your research to GenAI, or collaborate with AI researchers, in order to stay relevant, progress in your field, or secure funding?
58%
42%
Do you feel under pressure to use and learn about GenAI for SE Research to stay relevant?
55%
45%
0
20
40
60 Percentage
80
Response Yes No 100
Fig. 3. Pressures related to GenAI
Data Strategies Respondent Strategies
Lab Strategies
Field Strategies Nonempirical Strategies
Fig. 4. Where GenAI is being used. Research strategies and methods in the Y axis and research pipeline stages in the X axis.
(R306), while another expressed concern about the community’s direction: “I feel we have become a second-class AI community” (R166). 4.2
Where and how are researchers in SE using GenAI? (RQ2)
This research question provides an overview of how GenAI is being used in SE research, focusing on both where it is applied across research strategies, methods, and pipeline stages (Sec. 4.2.1), how it is used in practice through specific use cases (Sec. 4.2.2), and how researchers feel about the use of GenAI in reviewing research papers (Sec. 4.2.3). 4.2.1 Where are researchers in SE using GenAI? (research strategy, research method, and research pipeline stage) (RQ2.1). We analyze where GenAI is used using the Who–What–How framework of Storey et al. [40], focusing on the How dimension (research strategies) and the stages of the research pipeline proposed by Trinkenreich et al. [43] (Figure 4). GenAI usage is most concentrated in data strategies, with data mining studies showing the highest adoption across all pipeline stages. Respondent strategies (e.g., surveys, judgment studies) show moderate usage, while lab and field strategies consistently report lower adoption. Among non-empirical strategies, formal theory shows the lowest usage, whereas literature reviews show relatively high usage, particularly in writing. Across pipeline stages, GenAI usage is highest in writing and dissemination, followed by early-stage activities such as goals and design. In contrast, data collection, processing, and analysis show consistently lower usage across strategies. Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape 4.2.2
13
How are researchers in SE using GenAI? (RQ2.2).
This research question investigates how SE researchers integrate GenAI into their research practices. Instead of treating GenAI use as a single phenomenon, we analyze how GenAI is used across use cases. Results are organized using the Andersen et al. framework [2], extended with emergent categories and use cases derived from responses, as shown in Figure 5. Idea Generation
Research Design
Data Collection
Help brainstorm research ideas (*)
Help formulate questions for surveys or interviews
GenAI as subjects of research (*) Generate synthetic datasets
Help propose new hypotheses
Help identify relevant literature
Suggest a structure for research proposals
Help summarize or analyze existing literature
Help identify gaps in current research
Help design research methodology
Peer Reviewing (*) Help editors to summarize peer reviews (*)
Help write review reports during the peer review process
GenAI as the Research Topic (*) GenAI for recommender systems (*)
GenAI for testing (*)
GenAI for code (*)
GenAI for cybersecurity (*)
Data Analysis
Writing and Reporting
Annotation (*)
Qualitative analysis (*)
Help summarize own text (*)
Suggest a structure for a research article
Explore or corroborate plans for data analysis (*)
Results interpretation (*)
Propose a title, abstract or keywords for your article
Translate one of your research papers into a different language
Support secondary studies (*)
Log data analysis (*)
Create or edit simulation software code
Create or edit software code for data analysis
Help draft parts of a research article
Edit a research article to improve readability and/or language
Help create (parts of) a slide deck for a conference talk or similar academic event
Help create lay summaries or similar non-academic writing for public engagement, based on your own texts
Create or modify scientific figures or images
Support statistical data analysis
Cross Cutting (*) Learn unfamiliar concepts (*)
Fig. 5. GenAI use cases categorized using Andersen et al.’s framework [2]. We mark with (*) the categories and use cases that emerged from our data and are not part of [2].
Table 3 reports the number of participants whose responses were coded into each category. In the following, we describe these findings in more detail, organized by GenAI use-case category. Because participants’ responses could reflect multiple practices, individual responses were sometimes coded into more than one category. In the idea generation phase, GenAI is predominantly used to support early-stage sensemaking and orientation activities. Researchers mentioned using GenAI to help brainstorm research ideas, either in the “very early stages [..] for a quick overview of the state of the practice" (R314) or “not for initial brainstorming, but to refine ideas" (R18). In this role, GenAI was described as “a phenomenal tool for reflecting on new ideas, as often and as deeply as [they] need" (R2). The brainstorm sometimes leads to help propose new hypotheses, and respondents mentioned to use GenAI for “tun[ing] the questions and hypotesis" (R336) to “find a better formulation or statements" (R379). Beyond brainstorming and support on hypotheses, respondents reported using GenAI to “help identify relevant literature" (R75) as a way to “quickly learn or obtain summaries about specific topics” (R85). This included support for finding references (R69, R199) and related work (R66, R70, R206, R292), for example, by “using deep research" (R86, R426), or more generally, obtaining “basic overviews of topic areas" (R265). Still focusing on literature, respondents reported using GenAI to help summarize or analyze existing literature, including articles (R335) as well as different types of “texts, videos and audios" (R52). Beyond summarization, GenAI was also used to “explain [others] papers" (R426) and to help identify gaps in current research by “anticipating potential reviewer interpretations of specific sentences or results." (R84). As researchers transition into the research design phase, the use of GenAI was mentioned with more restraint. Here GenAI is primarily employed to help design research methodology, particularly by assisting with the clarification and examination of methodological steps. This is evidenced by reports of using GenAI to "discuss methodology steps and possible associated threats" (R303) and for "clarifying some steps of research methods" (R264). Another related use case Manuscript submitted to ACM
14
Trinkenreich et al.
Table 3. Representative examples of answers to how GenAI is being used in SE research, number and percentage of use cases whose answer was coded for each category. We use (**) to represent new categories that are not part of Andersen et al.’s framework [2]. Category Idea Generation Research Design Data Collection
Data Analysis
Writing and Reporting
Peer Reviewing (**)
GenAI as Research Topic (**)
Cross Cutting (**)
Representative examples “ask ChatGPT about my research ideas and let it help me analyse them" (R287) "generate summaries of papers already published" (R13) “Checking my memory on research methods." (R13) “as a subject of my research" (R39) “produce pilot datasets and scenarios to evaluate research protocols" (R202) “code and analyze interview transcripts and survey responses" (R6), “making annotations on data [...] and open coding" (R47) “help in writing scripts for evaluation purposes" (R75) “review of research papers that I did; like, possible gaps and inconsistencies" (R71), “when I am writing a report or paper, I seek feedback on the clarity and accuracy of my language." (R217) “refine my writing (e.g., grammar checks, clarity of the sentences)" (R317) “revising difficult texts like rejection emails for careful tone" (R71), “improving text when writing [..] reviews" (R167) “experimenting with the usage of GenAI in test case and UML models generation" (R46), “examining use of LLMs for generating code" (R82) “1) Secure coding; 2) explanation of ransomware attack strategy; 3) identifying the pattern of use of GenAi for secure coding tasks" (R278) “checking concepts that I am unfamiliar with" (R186), “understand new technologies, new concepts" (R152) “Help with learning/using APIs and other programming features" (R28)
# mentions
% (n=243)
52
21%
9
4%
2
1%
59
24%
94
39%
4
2%
12
5%
11
4%
The total per category is not the sum of the respondents since participants often provided an answer that was categorized into more than one use case (e.g., Idea Generation and Data Analysis).
involves methodological exploration, where GenAI is used to investigate “whether a more ‘natural’ or intuitive approach exists for a given problem” (R84). In addition to these methodological supports, some respondents report using GenAI to assist with the suggest a structure for research proposals, such as generating “skeletons for documents and grant applications” (R80). In the data collection phase, references to GenAI use appear less frequently and are mainly associated with supporting and exploratory activities. When GenAI is mentioned in this phase, it is not described as a mechanism Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
15
for directly collecting empirical data, but rather as a complementary resource within the data collection process. In this context, some respondents refer to GenAI as subjects of research, describing its use “as a subject of [their] research” (R39). Others report employing GenAI to Generate synthetic datasets, using it to “produce pilot datasets and scenarios to evaluate research protocols” (R202). The data analysis phase represents one of the most technically grounded areas of GenAI use. Researchers frequently mentioned using GenAI to augment analytical labor, particularly for creating or editing analysis code, supporting exploratory data analysis, and assisting qualitative analysis. A recurring pattern involves delegating mechanical or repetitive tasks to GenAI while retaining human oversight over interpretation and validation. Respondents describe using GenAI to support annotation activities as part of the analytical process. One participant notes that "have also experimented with its use for annotations. For instance, [he] asked a student to annotate a dataset and then prompted ChatGPT to do the same. Thereby, [they] spotted overlooked cases and computed the corresponding inter-rater reliability score" (R377). In qalitative analysis contexts, GenAI is similarly employed to assist with early analytical steps, including "support the analysis of qualitative data"(R157) and "making annotations on data [..] and open coding" (R47). GenAI is used to explore or corroborate plans for data analysis, including to "verify whether some abductive lines of thought can be grounded in some anecdotal evidence" (R85). It is also employed during results interpretation, where respondents describe using it to "analyze, interpret and make sense of my data" (R317) and to provide possible interpretations of my results (R44). Respondents also report using GenAI to support secondary studies, such as support clerical activities in secondary studies (R202). Log data analysis is another reported use, including "analyzing log data, detecting anomalies" (R141), as well as the create or modify scientific figures or images, such as "graph generation" (R60). GenAI is frequently used to create or edit software code for data analysis. Participants report using it to "generate the code for my research experiments" (R317), assist with "training parameters, polishing the prompts" (R160), and integrate it "as part of my toolbox for analyzing data" (R73). It is also used to support statistical data analysis, such as "selecting the right statistical tests for a certain purpose" (R160). The writing and reporting phase was the most mentioned one (see Table 3). Core use cases focus on linguistic and communicative support, including improving readability, rephrasing text, and summarizing one’s own writing. Help summarize own text appears in practices such as "summarizing notes" (R206) and "getting synopsis of (old) lines of work" (R69). Relatedly, GenAI is used to suggest a structure for a research article, for example, through "structuring text from bullets" (R43). It is also employed to propose a title, abstract or keywords for [the] article, including "brainstorming paper titles" (R47) and "give ideas (e.g., of paper titles)" (R134), as well as to translate one of your research papers into a different language, such as "language translation"(R2). In addition, respondents report using GenAI to help draft parts of a research article, including "refine paper" (R19) and "review of research papers that I did; like, possible gaps and inconsistencies" (R303). GenAI is also used to edit a research article to improve readability and/or language, with examples such as "when I am writing a report or paper, I seek feedback on the clarity and accuracy of my language" (R217), "polishing my writing" (R186), "refine my writing (e.g., grammar checks, clarity of the sentences)" (R317) and "text reviewing for grammar and fluency" (R207). Further uses include help create (parts of) a slide deck for a conference talk or similar academic event, such as "urning lecture transcripts into prose texts" (R71), and help create lay summaries or similar non-academic writing for public engagement, based on your own texts, for example "creating a summarized version for social media" (R314).
Manuscript submitted to ACM
16
Trinkenreich et al. Support Review Formulations
Check the Paper Using Paper Acceptance Criteria
Polish the Review
Generate a Summary of the Review
Verify the Review
Check the technical correctness
Check paper references
Check the relevancy of the paper's contributions
Compose the review from bullet points
Write a metareview from existing reviews
Translate a review from one language to another
check the novelty of the paper's contributions
Check the research method
Check the presentation quality
Provide Cognitive Support to the Reviewer Summarize the paper
Explain unfamiliar concepts
Search in the paper
Learn English
Fig. 6. The Use Cases of GenAI for Peer Review.
In the peer review phase, GenAI use is framed as supportive of comprehension and communication rather than evaluation. Respondents note that GenAI can help write review reports during the peer-review process, particularly for tasks that require careful wording and tone. This includes uses such as "revising difficult texts like rejection emails for careful tone" (R71), as well as "checking communications and paper reviews" (R215). In the cross-cutting phase, the use of generative AI is not confined to a single stage of the research lifecycle, but instead spans multiple activities across different phases. From this perspective, participants describe using GenAI to support ongoing learning and orientation, particularly to learn unfamiliar concepts, such as “checking concepts that I am unfamiliar with” (R186) and “learn about some specific techniques” (R16). Finally, when GenAI itself is the research topic, responses describe cases in which researchers explicitly study GenAI-enabled techniques as the primary object of investigation. These studies span different use cases of Software Engineering, including security, recommender systems, and testing. 4.2.3
How do researchers feel about the use of GenAI in reviewing research papers? (RQ2.3).
In the survey we asked an open-ended question for those that answered yes to: “As a reviewer, have you used GenAI to assist in reviewing a SE research paper?”. Eighty-one respondents shared that they used GenAI to support reviewing activities with most explaining how they use it. Most explained they found GenAI support for reviewing useful, but five participants mentioned they didn’t find it helpful, and three had mixed experiences using GenAI during reviewing. Notably one participant mentioned using it to review their own paper. Below, we present how respondents use GenAI to formulate their reviews, to check the paper against paper acceptance criteria, to provide cognitive support as they conduct their reviews, and the experiences the respondents reported while reviewing using GenAI as a support. Throughout, we discuss the implications of these findings and refer to additional comments the respondents provided as part of the last question in the survey for final thoughts. Finally, we share some additional insights respondents shared about their view of how GenAI is used in reviewing in our community. GenAI is Used to Support Review Formulations. Many of the participants explained how they use GenAI to support them in formulating their reviews. Figure 6 summarizes the themes that emerged for how GenAI supports review formulation. Forty-one participants mentioned they use it to polish their review only: “I use GenAI to rephrase and refine my review comments. This helps improve clarity and tone, making feedback more understandable and constructive for authors and fellow reviewers”(R23). While another described how they use GenAI to translate their review from one language to another: “to write my feedback in English as a translator"(R411). Seven respondents described how they use GenAI to verify their review: “I submit both the paper an my review and ask the model to challenge the review. Then I *critically* analyze the response.” (R219) Three mentioned that they use GenAI to summarize the main points in their review: “Only to summarize strengths and weaknesses of a paper based on my detailed comments” (R65). One participant mentioned Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
17
how they use GenAI to write the review from their bullet points: “Write reviews based on my bullet points from reading the paper; check for additional strengths and weaknesses that I overlooked; improve grammar and writing,” (R310) while another mentioned using GenAI to help them write a metareview from existing reviews: “Once I used it to write meta-reviews summarizing the existing reviews” (R43). In the final question of the survey that asked respondents about their final thoughts, several respondents corroborated the view that GenAI use in reviews should be limited to rephrasing. One respondent argued that this type of “GenAIRewriting-Confirmation should be allowed in any written communications, including paper and peer review writing” (R34). R332 drew the line explicitly, stating that reviewers should “only be allowed to use it to rewrite what they have already written,” consistently with the existing ACM policy on authorship. 4 GenAI Used to Check the Paper Using Paper Acceptance Criteria. Several respondents described how they use GenAI to help them review papers using acceptance criteria that are often used for paper reviews in software engineering. Four respondents mentioned using GenAI to check the technical correctness of the paper. For example, R343 said: “By selecting and pointing out all the logical inconsistencies and shortcomings in the publication.” Two mentioned using GenAI to check paper references: “assisting with reference and data checks” (R327). Two respondents mentioned using GenAI to help them check the relevancy of the paper’s contributions: For example, “I asked GenAI to estimate the top X factors and challenges of an SE phenomenon because I thought the paper’s results were obvious. GenAI produced the same results as the paper but with better insight. I agreed and recommended rejection of the study as ‘obvious”(R70). One participant also used GenAI to check the novelty of the paper’s contributions: “To help search for related literature and double check the novelty of contribution.” (R257) Two respondents mentioned they used GenAI to check the research method used for the research in the paper under review. For example, “I used it only to verify whether a design method is suitable for the research goal at hand when I am not much experienced with it.” (R76) Two respondents mentioned using GenAI to check the presentation quality of the paper. For example, “checking writing of a part of the article where I think is a problem with English. Inspecting availability of tools and references” (R247). The final thoughts open-text question revealed strong views on how far such analytical use should extend. Several respondents stated that GenAI may be acceptable for reading the submission, but is “questionable for detecting pros and cons, and is inacceptable for making a decision” (R41). Similarly, one respondent argued that GenAI “should NOT be given the paper itself and asked what it thinks” and should instead be limited to Grammarly-like functions (R48). Another drew the line at content analysis, stating that reviewers should be “allowed to use GenAI only to rephrase the text, and it should be forbidden [to] use [it] to analyze a paper, review the literature or check the results/code” (R184). GenAI Provides Cognitive Support to the Reviewer. In addition to using GenAI to support review writing and review formulation, several respondents described how they used GenAI to provide other types of cognitive support as they were reviewing a paper. These additional types of cognitive support that emerged from the open-ended responses are summarized next. Six researchers described how they use GenAI to summarize the paper to aid their understanding of the paper: “At beginning to get quick summary of the paper” (R193). An additional two respondents describes how they use GenAI to explain unfamiliar concepts or to provide a summary of an unfamiliar topic: “...and also to clarify concepts I am not familiar with” (R214). One respondent described how they used GenAI to search in the paper to “...find in some case the link of the online appendix” (R45). And one participant mentioned how they use GenAI to help 4 https://www.acm.org/publications/policies/new-acm-policy-on-authorship
Manuscript submitted to ACM
18
Trinkenreich et al.
them learn English as they are reviewing, not just to improve the review: “To improve my English and writing skills for my final revision. Again, I use it for more grammatical aspects and to see if what I want to say is being conveyed correctly” (R435).
Poor Experiences Using GenAI for Reviewing. Although, respondents described the many ways GenAI supported them while reviewing, not all had positive experiences. For example, “I’ve tried it to see if I missed anything I should. I’ve found that it’s *really* bad at reviewing papers. It found critiques of things that were trivial (formatting of the bibliography) while missing fundamental problems in a paper. Reviewing takes higher-level thinking, and my experience (given, this is like an n of 3) is that it lacks that capability. Maybe the new "thinking" models like O3 would do better. I’d like it to helpfully review my own papers before submission but I have yet to get a whole lot of utility out of it” (R7).
Stances on GenAI Use in Peer Review. Beyond specific use cases, in the final thoughts question of the survey, 15 respondents spontaneously suggested normative positions on whether GenAI should be permitted in peer review at all, ranging from outright prohibition to conditional usage with safeguards. The most restrictive stance called for a complete ban, even on text polishing. One respondent argued that “arguments for or against a paper should be of one[’s] own” and that “not having very-well written reviews do not affect the final quality of the paper,” concluding that GenAI should “be forbidden at all (even for polishing reviews)” (R26). Respondent R315 was equally firm: “It should not be allowed, even with guidelines. Peer review is a critical quality gate of scientific publications” and should rely on the ability of peers to “understand the content of the paper and evaluate its fitness for the target venue.” Some respondents framed the issue in terms of professional trust. One stated that reviewers caught using GenAI to write reviews “shouldn’t be welcome in the community going forward, full stop,” comparing the practice to “letting a random friend at a bar write your reviews for you” (R42). Another questioned the very purpose of peer review under GenAI adoption: “Why have reviewers if they are going to [use] AI?” (R59). An intermediate position acknowledged the practical pressures on reviewers while expressing doubt about responsible use in practice. One respondent suggested that “reviewers should not be encouraged to use it, but allowed to use it with guidelines,” noting that “the reviewing load is high and the research community needs many, high-quality reviews,” but also adding: “I am very doubtful that people can use it correctly” (R107). Another respondent called for allowing GenAI “given how much time it can save” while stressing that “this heightens the responsibility of reviewers” (R82). A constructive alternative was proposed by respondent R37 who stated they “out of principle never use AI for reviewing” and consider it a duty to report reviewers who misuse GenAI when evaluating assignments, but could also envision “mixed formats, where a paper is reviewed by humans AND AI in parallel, as a ‘4th reviewer’ so to speak.” Another respondent envisioned using GenAI to “generate template-based paper summaries that could be used for a ‘first round’ review of papers,” while also cautioning that “there is a risk that the reviewers will also use GenAI in the second round” (R181). Finally, respondents raised specific risks tied to the reviewing context. One noted that “the very second I upload a document to GenAI, I potentially breach non-disclosure agreements (also in context of peer review)” (R158),highlighting the confidentiality implications in reviewing activities.
4.3
What benefits, challenges, and opportunities do SE researchers perceive when using GenAI? (RQ3)
In this research question, we analyze the reported benefits, challenges, and opportunities of using GenAI for SE research. Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape Helps researchers without English as a first language (through editing or translation) Helps write manuscripts faster
7% 14%
Developing research tools or scripts easier or faster
16%
46%
32%
7%
36%
55%
40%
33%
Speeds administrative tasks
7%
18%
45%
27%
Makes data coding easier and faster
8%
17%
46%
24%
Summarises other research to save time reading it
6%
Helps creative work by brainstorming new ideas
9% 7%
Improves scientific search
14%
Makes research more enjoyable Helps peer-review manuscripts faster Generates new research hypotheses
14%
40%
20%
20%
35%
16%
26%
30%
15%
40%
23%
21% 17%
31% 16%
20%
25% 35% 10%
Other
20%
18%
30%
15%
20%
20%
19 Completely Disagree Generally disagree Neither agree nor disagree Generally agree Completely Agree
19%
7%
50%
60% 50% 40% 30% 20% 10% 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% Percentage of Responses
Fig. 7. Perceived benefits of using SE research
4.3.1 Which benefits do SE researchers perceive GenAI offers to SE Research? (RQ3.1). Figure 7 presents the perceived benefits of using GenAI in SE research. To facilitate interpretation, we aggregate completely disagree with generally disagree, and generally agree with completely agree. In general, the responses indicate a predominantly positive perception of most listed use cases, particularly those related to writing and text processing. The strongest agreement is observed for helping non-native English speakers (91% agreement), followed by easier or faster developing research tools and scripts (78%), also writing manuscripts faster (73%), speeding up administrative tasks (72%), making data coding easier or faster (70%), summarizing text (59%), and improving scientific search (45%). More mixed perceptions emerge for creativity-related activities. While a majority still agrees that GenAI helps brainstorming new ideas (51%), responses are more distributed, and perceptions are even more neutral for making research more enjoyable (30% agreement, 40% neutral). We expand on the opportunities to improve researcher experience in Section 4.3.3. In contrast, perceptions are predominantly negative for speeding up peer review (56% disagreement) and generating new research hypotheses (51% disagreement). Researchers’ final thoughts reinforced these findings, revealing nuances in how benefits are perceived in practice. Regarding text processing (the category with the strongest agreement), one respondent described GenAI as “excellent in formulating perfect texts [. . . ] great for dissemination results and paraphrasing own’s texts” (R343). Beyond text processing, productivity benefits were broadly recognized, though respondents typically coupled them with calls for caution. Respondents noted that GenAI “definitely helps SE researchers a lot in our daily tasks” but stressed the need to “be very careful and establish guidelines to use it” (R139). Similarly, while acknowledging that GenAI is “definitely a tool to boost productivity,” one respondent underlined that “we must use it wisely,” raising concerns that uncritical adoption “may Manuscript submitted to ACM
20
Trinkenreich et al.
Fig. 8. Challenges faced when using GenAI
lead to an explosion of papers, not necessarily of high quality” (R135). Others framed GenAI as “both a transformative opportunity and a serious responsibility,” emphasizing that “its use must be grounded in transparency, reproducibility, and ethical standards” (R23). This recurring “yes, but” pattern where benefits are acknowledged alongside calls for caution, regulation, or verification suggests that even among researchers who perceive clear advantages, uncritical adoption is not endorsed. We further analyze these conditions and perceived risks in Section 4.4. 4.3.2
What challenges do SE researchers experience using GenAI in SE research? (RQ3.2).
Figure 8 shows that overall lack of trust in AI was the most frequently cited challenge, reported by 34% of respondents. We provide a more detailed analysis of trust-related responses in Section 4.4.3. The second most commonly reported challenge, mentioned by 23% of respondents, concerns the regulatory landscape (e.g., GDPR). We examine these regulatory concerns in greater depth in Section 4.5. 4.3.3
What opportunities do SE researchers envision GenAI offers to SE Research? (RQ3.3).
This research question investigates which GenAI use cases are envisioned as opportunities for SE research. Results are organized using the Andersen et al. framework [2], extended with emergent categories derived from participants’ responses, as illustrated in Figure 9. Our qualitative analysis identified 15 opportunity categories that overlap with current GenAI use cases reported in Section 4.2.2. In Figure 9, opportunity categories not previously observed as current uses are highlighted in bold and described below. Table 4 summarizes the number of participants whose responses map to each category. We next present detailed findings organized according to these categories of GenAI opportunities for SE research. Because participants’ responses could reflect multiple opportunities, individual responses were sometimes coded into more than one category. The most frequently cited opportunities concerned researcher experience, highlighting how GenAI may reshape the day-to-day practice of SE research. A dominant theme was the automation of tedious tasks, including routine coding, labeling, scripting, and other repetitive activities that are “rather routine human-labor intensive but not ‘intellectually exciting’ aspects of research" (R56). Such automation was widely associated with increased efficiency and the ability to shift effort toward higher-value activities. Closely related, respondents emphasized that GenAI frees time for tasks that matter. By removing implementation and formatting burdens, researchers can “focus on the problem and solution Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
21
Research Experience (*) Frees time for tasks that matter (*)
Automation of tedious tasks (*)
Idea Generation
Makes researchers more productive (*)
Research Design
Help brainstorm research ideas (*)
Help identify relevant literature
Create new methods (*)
Help summarize or analyze existing literature
Prototype and explore alternate ideas (*)
Experimentation with new techniques (*)
Grant writing (*)
Cross-fertilisation or ideas / knowledge from other fields (*)
Producing better quality research (*)
Qualitative analysis (*)
Annotation (*)
Explore or corroborate plans for data analysis (*) Create or modify scientific figures or images
Help design research methodology
Requires expertise / human-inthe-loop (*)
Writing and Reporting
Data Analysis
Data Collection Generate synthetic datasets
Reduces cognitive load (*)
Quantitative analysis (*)
Large scale analysis (*)
Data wrangling (*)
Analysis of software data (*)
Create or edit software code for data analysis
Feedback on own writing (*)
Help with writing (*)
Formatting (*)
Help draft parts of a research article
Translate one of your research papers into a different language
Help create (parts of) a slide deck for a conference talk or similar academic event
Edit a research article to improve readability and/or language
Cross Cutting (*) Access to information (*)
Summarization of text (*)
Peer Reviewing (*)
GenAI as a research partner (*)
Learn unfamiliar concepts (*)
Error checking (*)
Make SE and SE research more accessible (*)
Paradigm shift (*)
GenAI as the Research Topic (*) Cynycal (*)
Review assignments (*)
GenAI for SE (*)
SE for GenAI (*)
GenAI for SE Research (*)
Introduce new research topics (*)
Revival of old research topics(*)
Make research topics obsolete (*)
Ethical concerns (*)
Homogeneity (*)
Introduces bias (*)
Limited utility / not for tasks requiring critical thinking (*)
No or limited opportunities (*)
Pessimistic outlook (*)
Quality concerns (*)
Fig. 9. GenAI opportunities categorized using Andersen et al.’s framework [2]. We marked as (*) the categories and opportunities that emerged from our data and were not part of [2] and bold boxes as the opportunities that were not mentioned as current use cases in Section 4.2.2.
rather than on the implementation and dissemination" (R62) and devote more attention to innovation and deep thinking for better science (R164). Many responses framed these changes as making researchers more productive. Reported benefits included faster prototyping, quicker data analysis, accelerated writing, and the ability to “do more research within the same limited time" (R34). Others described broader efficiency gains across the research lifecycle, including speedups in nearly all research steps and fewer mistakes through automated checking (R308). Beyond productivity, respondents also associated GenAI with producing better qality research. Some highlighted increased output quality and quantity (R138), while others connected automation and intelligence in SE workflows to improved software quality and modeling capabilities (R142). Several responses further indicated a reduction in cognitive load. Examples included no longer needing deep expertise in specific languages or tools to apply them effectively (R156) and mitigating fatigue-related limitations in SE work (R80). Finally, respondents repeatedly stressed that meaningful benefits still reqire expertise and human-in-the-loop oversight. GenAI was described as “a great helper if and only if you already have a fundamental understanding" (R63), with methodological decisions and interpretation remaining inherently human responsibilities (R336; R380). In the idea generation phase, respondents envisioned opportunities that were not reported as part of current use (marked with a star in Fig. 9). Among these, Grant writing was briefly mentioned (R52). Another emerging opportunity concerned the ability to prototype and explore alternate ideas, primarily enabled by faster prototype development (R6; R83; R130). This increased speed was described as allowing researchers to “explore multiple alternatives and pick the best one" (R7). Respondents also envisioned opportunities to cross-fertilize between different research fields, such as better use of statistical methods and tools" (R65). In this context, GenAI was described as enabling SE researchers to broaden Manuscript submitted to ACM
22
Trinkenreich et al.
Table 4. Representative examples of opportunities GenAI bring to SE research, number and percentage of use cases whose answer was coded for each category. We use (**) to represent new categories that were not part of Andersen et al.’s framework [2] Category Idea Generation Research Design Data Collection Data Analysis Writing and Reporting Peer Reviewing (**) GenAI as Research Topic (**)
Cross Cutting (**)
Cynical (**)
Researcher Experience (**)
Representative examples “can be great to kickstart a literature search, summarize papers to see if they are relevant and warrant a deeper look, and similar tasks" (R63) “Faster experimentation with techniques that a researcher is not familiar with by specifying the intended research goal as a prompt" (R201) “generating synthetic data" (R187) “generate quick tools and scripts for analyzing data" (R82) “Help the researcher to go through multiple documents in less time and express the ideas in better English" (R340) “GenAI can help with a fairer allocation of paper review tasks for SE conferences based on reviewers’ prior publications" (R17) ‘explor[ing] how humans and AI systems co-develop software, raising new questions in usability, trust, and explainability of AI-generated artifacts" (R23) “GenAI can serve as a consultant throughout a research project" (R174) “as an additional author [..] to reduce researcher bias and highlight areas that are overseen by the authors" (R215) “Ethical concerns arise [..] and some misconduct in paper writing and reviews have already been seen [..] [raising] a blurry line [..] on the edge of ethical principles" (R165) “Similar to when Google freed me from remembering all facts about the world, I no longer have to be an expert in a language or method to be able to use it." (R156) “smaller but time-consuming tasks can be sped up" (R63)
# mentions
% (n=266)
37
14%
8
3%
7
3%
22
8%
19
7%
1
0%
15
6%
23
9%
29
11%
68
25%
The total per category is not the sum of the respondents since participants often provided an answer that was categorized into more than one use case (e.g., Idea Generation and Data Analysis).
research scope by quickly grasp[ing] a new area, technology or concept" (R153), at least capturing “the essence" of unfamiliar domains (R288). Within the research design phase, respondents envisioned opportunities to create new methods, noting that “the nature of user studies might change in some cases, with the right protocol, it might be easier to automate" (R85). In addition, GenAI was seen as enabling experimentation with new techniqes(R201). In the data analysis phase, respondents envisioned advances in qantitative analysis (R7, R187) and, relatedly, large-scale analysis, emphasizing support for big data analysis" (R326). Respondents further highlighted opportunities for data wrangling, including optimization of data" (R152) to make data preprocessing much easier" (R333). Finally, opportunities for the analysis of software data (R142) included rapidly assessing large codebases" (R6). In the writing and reporting phase, respondents also envisioned opportunities on formatting (R40), help with writing (R57, R435, R453), and more specific support such as feedback on one’s own writing. Early feedback" Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
23
(R85) was described as especially beneficial for students, helping to speed up feedback cycles before manuscripts reach a supervisor" (R204). At the same time, respondents cautioned that using GenAI as a writing assistant should not replace students’ own authorship, as this “could lead supervisors to read superficial and uninteresting prosa" (R204). The opportunity to support peer review was described as helping “a fairer allocation of [..] tasks based on reviewers’ prior publications" (R17). When considering GenAI as the research topic, respondents envisioned a broader research agenda extending beyond current application-oriented studies toward systemic, methodological, and disciplinary transformation. Opportunities included understanding how SE practitioners can “get the best from it" (R47). From the complementary perspective of SE for GenAI, respondents emphasized that “GenAI [is] a new kind of software. Its problems need SE methods" (R290). Viewed through McLuhan’s tetrad [28, 28], GenAI was further seen as retrieving prior lines of inquiry by “compar[ing] and contrast[ing] how certain types of studies will change with the GenAI" (R277) and reviving dormant areas of investigation, where “many topics can be revived with GenAI" (R250). At the same time, respondents anticipated both enhancement and reversal in the research landscape, where GenAI may “bring new research problems for SE" (R20) and “make others obsolete" (R122), while also “provid[ing] new powerful tools for developers to break boundaries they couldn’t break with traditional technologies" (R122). As a cross-cutting dimension, respondents envisioned summarisation of text (R52), which, while visible during idea generation, was here framed as supporting multiple research activities throughout the lifecycle. Respondents positioned GenAI as a research partner, variously described as “GenAI is like a collaborator" (R310), a 24/7 research partner" (R3), or a “thought partner" (R58). Related opportunities for error checking included helping researchers “make fewer mistakes when using] LLMs to check [the] work" (R308), “identifying potential problems in the design of the methodology" (R289), and “highlight[ing] areas that are overseen by the authors" (R215). Respondents also described expanded access to information, characterized as “like when Google Scholar appeared, but x10" (R308), alongside opportunities for “making SE and SE research more accessible. Accessibility and inclusion were detailed as “lowering the entry barrier to SE research for non-native English speakers and those without systematic knowledge of the field" (R187) and broader “inclusion [..] people in other disciplines" (R332). Finally, some respondents framed GenAI’s influence as a paradigm shift in SE research (R300). Respondents expressed cynical or critical perspectives regarding opportunities for GenAI in SE research, emphasizing ethical, epistemic, and quality-related risks. Ethical concerns were mentioned, with the reflection of having the “opportunity to test the limits of the community" (R165). Related apprehensions involved homogeneity and potential loss of creativity, including the need for “carefully monitoring effect on creativity" (R85) and concerns about “better writing (but less individual [better writing])" (R19). Respondents also highlighted risks that GenAI introduces bias, noting that automated feedback must be used cautiously to avoid “biased analysis [. . . ] and [effects on] creativity" (R85), while summaries may “suffer from fixation effects" (R19). More broadly, several participants argued for limited utility, particularly for tasks requiring critical thinking. These included rejecting AI-written papers because supervisors may read “superficial and uninteresting prosa" (R204) and doubting GenAI’s ability to generate “genuine novel ideas" (R2). Others questioned technical progress and trustworthiness, citing limits in “reasoning and critical thinking tasks" (R165) and uncertainty about reliability in more important tasks (R181). Some framed current benefits as restricted to “routine [. . . ] but not ‘intellectually exciting’ aspects of research" (R56) or reported “not many opportunities" at present (R195), including skepticism toward research-question generation Manuscript submitted to ACM
24
Trinkenreich et al.
Trust in GenAI among Researchers who have used GenAI it for research (n=339) "Q8: For the following statements, please indicate your level of agreement on trust in using GenAI for SE research in general." I am confident in GenAI tools. I feel that they work well for research.
5%
24%
31%
34%
GenAI tools are reliable for research. I can count on them to be correct for my use cases.
12%
45%
26%
16%
I feel safe that when I rely on GenAI for my research, I will get the right answers.
26%
40%
24%
9%
21%
15%
I like using GenAI for decision-making in my research.
35%
26%
6%
Completely Disagree Generally disagree Neither agree nor disagree Generally agree Completely Agree
70% 60% 50% 40% 30% 20% 10% 0% 10% 20% 30% 40% 50% Percentage of Responses Fig. 10. The Different Dimensions of Trust in GenAI for SE Research (based on [9])
(R202). Respondents also perceived no or limited opportunities, stating that “the problems and dangers outweigh the opportunities" (R286), describing GenAI as “a hype bubble [. . . ] oversold" (R155), or asserting it as “the opposite of research" (R77). These views sometimes extended to a broader pessimistic outlook, framing GenAI as “a debasement of our art" (R272), even “the down fall" (R59), or criticizing perceived citation-driven trends and overreliance on AI-mediated research practices (R166). Others warned of long-term societal harm (R182). In final thoughts, this pessimism extended to concerns about community identity, with one researcher lamenting that the SE research field has “become a second-class AI community” (R166). Quality concerns emphasized more risks (than opportunities) of “incorrect and shallow research" (R74), outputs that are “more bug-prone" (R19), and potential loss of depth among early-career researchers (R165). One respondent further cautioned that excessive reliance on GenAI may undermine trust in the research record and should be avoided in core empirical activities (R277). Final thoughts reinforced these quality concerns, with one participant stating that “unless hallucinations are solved, GenAI is a non-starter in these contexts” (R268), framing reliability not as a temporary limitation but as a blocking condition. 4.4
How do SE researchers trust its use and what risks do they perceive from using GenAI in research, and how do they mitigate those risks? (RQ4)
In this research question, we investigate how software engineering researchers trust generative AI in the research process, the risks they perceive in using GenAI for SE research, and the strategies they propose to mitigate these risks. 4.4.1 How do SE researchers trust GenAI for research? (RQ4.1). Researchers who use GenAI for research reported moderate levels of trust in GenAI tools. Considering Top-2-Box scores (Agree and Completely Agree), 34% indicated confidence in GenAI tools, 12% perceived them as reliable for research, 9% felt safe relying on them for research tasks, and 15% reported liking their use for research decision-making. Trust in GenAI across research activities: As shown in Figure 11, trust in GenAI varies across different stages of the research process and across user groups. Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
25
Trust in GenAI for activities among Researchers who have used GenAI it for research (n=339) "Q9: For the following activities, I trust the use of GenAI." Research Goals and Questions and Formulation
45%
29%
26%
Study Design and Methodology
37%
34%
29%
Data Collection
34%
33%
33%
13%
Data Processing
25%
Analysis and Interpretation
21%
66%
32%
43%
11% 16%
Writing and Dissemination
Completely Disagree Generally disagree Neither agree nor disagree Generally agree Completely Agree
73%
60% 50% 40% 30% 20% 10% 0% 10% 20% 30% 40% 50% 60% 70% 80% Percentage of Responses Fig. 11. Trust in GenAI per Stage in the SE Research Pipeline (based on [43])
Perceived risks of GenAI among Researchers who have used GenAI it for research (n=339) "Q10: For the following statements, please indicate your level of agreement on the following risks that GenAI brings to SE Research?" GenAI makes it easier to fabricate or falsify research 12% and harder to detect GenAI raises energy consumption and carbon 10% footprint of research GenAI makes plagiarism easier 11% and harder to detect GenAI may bring mistakes or inaccuracies into research texts (papers, code) GenAI may entrench bias or 8% inequities into research texts
24%
64%
31%
59%
25%
64%
14%
85%
31%
61%
GenAI may proliferate misinformation
7%
21%
72%
GenAI may bring biases into literature searches
7%
25%
68%
Completely Disagree Generally disagree Neither agree nor disagree Generally agree Completely Agree
20% 10% 0% 10% 20% 30% 40% 50% 60% 70% 80% 90% Percentage of Responses Fig. 12. Perceived risks of using GenAI for SE research
Researchers who use GenAI for research reported higher levels of trust in later-stage research activities. Trust was highest for Writing and Dissemination (77%), followed by Analysis and Interpretation (41%), Data Collection (31%), Study Design and Methodology (25%), and Research Goals and Question Formulation (23%). 4.4.2 How do SE researchers perceive the risks of using GenAI for research? (RQ4.2). This research question examines the perceived risks of using GenAI in SE research. Figure 12 presents the perceived risks associated with the use of GenAI in SE research. Manuscript submitted to ACM
26
Trinkenreich et al.
Human in the Loop
Usage Boundaries
Technical Safeguards
Education
Governance
Fig. 13. The mitigation strategies reported by SE researchers who participated in our study.
Considering Top-2-Box scores (Agree and Completely Agree), researchers who use GenAI for research are mostly concerned about the possibility that GenAI may introduce mistakes or inaccuracies into research texts, though the majority agreed with all of the potential risks. 4.4.3
How can SE researchers mitigate the risks without missing the opportunities GenAI offers? (RQ4.3). This research
question examines the actions participants identified to mitigate the perceived risks of GenAI in SE research. Our analysis revealed five categories of mitigation strategies, as illustrated in Fig. 13. Table 5 presents the number of participants whose responses fit in each category. In the following, we present more details about our findings, organized by category of mitigation strategy. Because participants’ responses could reflect multiple mitigation strategies, individual responses were sometimes coded into more than one category. Human in the Loop is an interaction paradigm in which AI-generated outputs are treated as suggestions that remain subject to human judgment, validation, and accountability rather than being accepted autonomously. Respondents emphasized that “GenAI should not be used to replace researchers, but can only help them" (R47), highlighting that responsibility and agency must remain with humans. This view was echoed in final thoughts, where one respondent Table 5. Representative examples of answers to the risk mitigation open question, number and percentage of respondents whose answer was coded for each category. Mitigation Strategy Human in the Loop Usage Boundaries Technical Safeguards
Transparency
Education
Governance Controls
Representative examples “AI does not substitute the human, who should inspect everything" (R62) “The researcher has the final word. They have to review what the GenAI generates truly. We cannot blindly trust it." (R135) “Researchers should use GenAI as a tool to help with specific, well-defined tasks [..] They should not attempt to "offload at the edge" of knowledge tasks to a text completion engine." (R42) “RAG-based usage," (R164), “agentic systems to assist in verifying research," (R138) “use separate conversation thread and modify the question to see if the answer changes" (R172) “Disclose the usage of AI in research in more detail," (R427), “mak[ing] it explicit which parts of the research were supported by them (e.g., Threats to Validity section)" (R198) “Tell and demonstrate [researchers] what happens if [they] use GenAI incorrectly, (R107) “Promote ethical behaviour. Encourage the detection (and public report) of published materials corresponding to illustrative cases of misinformation, falsification, and plagiarism." (R159) “Open science and replication requirements, (R74) “design new trustworthy assurance and evaluation methodolygy" (R15)
#
% (n=138)
90
65.2%
11
8.0%
13
9.4%
17
5.9%
14
10.1%
4
2.9%
The total per mitigation strategy is not the sum of the respondents since participants often provided an answer that was categorized into more than one strategy. Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
27
cautioned that “entirely GenAI generated feedback to papers and draft is wrong” and that “GenAI should be used to complement what humans do” (R279). Respondents stressed the need to “double check GenAI results", emphasizing that outputs should be verified rather than used blindly. Across responses, double-checking was consistently associated with manual inspection, cross-verification with external sources, and the expectation that “the human must make the final decision and take the responsibility" (R326). Some articulated concrete verification practices, such as comparing with “traditional searching solutions" (R455) and ensuring that “citations [..] point to clear parts of the cited manuscripts" (R22). Double-checking include criticizing GenAI results (R449). Respondents argued for being “quite critical, analytical, and aware of whatever GenAI might generate" (R435), warning that even common academic-writing support can introduce substantive inaccuracies. For instance, GenAI-generated text for an introduction or abstract may, “in an attempt to create something novel," “hallucinate and include aspects that aren’t actually in the paper" (R435). Similarly, when used for related-work analysis, it can be misleading because “it even generates its own information that isn’t associated with the referenced papers" (R435). The final thoughts reinforced this verification norm, with one researcher framing it as a professional duty “[. . . ] to double check if the suggestions or pointers provided by GenAI are actually factually true” (R343). Others provided verification workflow examples, such as routinely confirming “if the rewritten text is still semantically equivalent” before accepting GenAI-rewritten prose (R34). “Maintain human oversight", “especially in tasks involving interpretation or critical decision-making" (R3), was highlighted as a process in which “all the content [is] properly verified before usage" (R76), in which researchers may “ask genAI and then decide" (R75), but be supported by “manual analysis" (R16) of the output. In this framing, GenAI is explicitly treated “as a tool, mak[ing] sure you look at what it’s saying and use your best judgement" (R7). Respondents emphasized retaining human ownership of outcomes, stating that humans should have “the first and the last word" and that GenAI should be used only “as a auxiliar tool" (R247). In order to maintain human oversight, respondents suggested that GenAI “must be used by people who can already perform those tasks effectively" (R184). The final thoughts corroborated this stance. One respondent stated that “GenAI is like any other tool—the responsibilities stay with humans using it” (R343), while another commented that “papers and reviews are the products of humans, and humans are responsible for accuracy” (R336). Review results more carefully emphasizes deeper, more skeptical evaluation of GenAI-assisted outputs by both researchers and reviewers. Respondents stressed the need to “remain cautious about the genAI results" (R17) and argued that reviewers would need to “take more time on reviewing manuscripts manually and in more detail," described as the only way to address risks such as falsified results, plagiarism, and inaccuracies (R19). This stance also involves preparation before tool use, with one respondent noting the importance to run manual analysis before using GenAI by “preparing content or having an idea about the topic [..] and after reviewing the results carefully" (R61). Although acknowledged as difficult to scale, respondents associated careful review with heightened rigor, emphasizing that “the bar on research rigor must be raised" (R68). One researcher’s final reflection underlined this expectation, stating that “the use of AI tools is not an excuse for inadequate review” (R336). Usage Boundaries refer to strategies that constrain when, how, and for which purposes GenAI should be used. Rather than treating GenAI as a general-purpose assistant, respondents emphasized the need for concrete boundaries and data protection. Restrict Task-Types was mentioned as a boundary to use of GenAI on tasks that do not require deep understanding of methodological judgment. For example, in literature reviews, while “summarizing existing literature is a good thing," researchers “shouldn’t always rely on GenAI for selecting papers," emphasizing that “true personal understanding can happen by actually reading the paper and not a summary" (R2). Respondents suggested using GenAI Manuscript submitted to ACM
28
Trinkenreich et al.
to “help with English grammar)" (R42), “brainstorm paper titles or keynote titles," (R52) and specific parts for the paper, as a “draft part of an introduction," emphasizing that the generated text is treated as provisional scaffolding rather than the final prose (R167). In this approach, GenAI-generated content is explicitly framed as temporary that would “ultimately get replaced before submission" (R167). Technical Safeguards were described as tools and automated checks that support the evaluation of GenAI-produced results. At the system level, respondents advocated for “mak[ing] AI systems resemble traditional systems". At the toolselection level, respondents emphasized choosing models that better support reliability. For example, one mentioned selecting LLM that creates trustable summary (e.g. NotebookLM5 )" (R141). Beyond generation, respondents also proposed augmenting evaluation workflows. Several respondents suggested using an automatic tool to identify “misinformation or incorrect claims” (R20). Respondents also stressed context-specific safeguards, noting the need to remain careful with using GenAI for any actual data processing [when] work[ing] with personal data" (R192). Transparency captures mitigation strategies that emphasize making GenAI use, decision processes, and research artifacts visible and inspectable to researchers, reviewers, and the broader community. Respondents stressed the importance of “clearly documenting [the GenAI] role in writing, analysis, or ideation" (R3), including making explicit which parts of the research were supported by GenAI (R198). In addition, respondents highlighted the need to “make the decision making more transparent" (R145) and argued that “in order to trust GenAI tools, the underlying models should be transparent and customizable" (R151). From this perspective, transparency also involves making system behavior interpretable, such that tools “reflect the context behind the information" (R152). Education suggestions emphasized the role of cultural norms and community values in shaping responsible GenAI use. Several participants explicitly called for education and awareness, stating “Education, education, education!" and urged researchers to be shown “what happens if you use GenAI incorrectly" (R107). This also involves broader awarenessbuilding, including “mak[ing] the researchers and teachers and students aware of [GenAI] risks" (R340) and of “the perils of GenAI" (R166). Finally, respondents linked education to structural changes, developing “techniques that helps researchers to use these tools correctly and with ethics perspective" (R411). Governance Controls are strategies about institutional, procedural, and regulatory mechanisms to control how GenAI is used in research. A strategy that stays between training and governance to know “what is or is not ok to use AI for" (R83). Others extended this responsibility to institutions, arguing that “publishing venues and funding agencies need proper guidelines and quality assurance measures" to prevent and detect GenAI-related issues such as plagiarism (R158). 4.5
What regulations and policies do researchers feel should be applied to the use of GenAI in their research and in peer review? (RQ5)
In this research question, we examine the regulatory and policy needs identified by SE researchers regarding the use of GenAI in research activities. In the survey we asked an open-ended question: “Should GenAI use be regulated in Software Engineering research, assuming that it is possible?”. We present the overall stance of respondents, followed by three themes about how regulation that emerged from the responses. Overall Stance Towards Regulating GenAI Use in SE Research. From the 119 participants who answered this question, the majority (101) suggested that “Yes” GenAI should be regulated, 8 were more on the fence, and 9 tended towards “No”, it should not be regulated at all. 5 https://notebooklm.google.com/
Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
29
While the majority agreed that some regulation is needed, they also qualified their response with details on which kinds of research tasks regulation is needed (e.g., reviewing, data analysis), or that new researchers may be in need of more regulation (or guidelines). The reason given was not just about the integrity of the research, but also because, automating research reduces the learning opportunities especially for new researchers: “It should not be used in Systematic Literature Reviews. Why? Maybe it’s a personal bias, but I have learned in the last, say 5 years, how to perform an SLR and just few days ago I so in a course how easy is to do this with the help of AI. BUT, I think GenAI steals the opportunity of the researcher/learner to learn! The same in teaching: if somebody else (even GenAI or another human, friend, parent) does the task instead of you, that person will learn and not you!” (R340) For those who were against using GenAI at all, they had even stronger opinions, for example: “I believe if a researcher is using GenAI to think for them, they might as well remove their PhD title from their CV.” (R166) While others felt that at the minimum its use should be regulated to be transparent, but even then boundaries should be established in how it is used: “Clear statement when GenAI is used as part of the research (as an assistant to help in a specific SE task or when it is used as a tool to deliver novelty (but in this case, it should not be allowed).” (R340) For the few that were strongly against any kind of regulation, they brought in parallels to the use of Google and the internet in our earlier research: “The use of GENAI as a coding and writing assistant is prevalent and does not violate any code of ethics to my knowledge. I DON’T THINK ANY REGULATION IS NEEDED. If Google era research was never regulated then GENAI era research should be no different. GENAI is only as good as a user’s prompt. Since prompts are a user’s personal intellectual property so the resulting output which is generated as a result of the prompt also belongs to the prompt writer. Perhaps I do not have any misuse cases in my mind right now otherwise I would have been better able to comment on this question.” (R171) One respondent even considered us asking this question was conceptually flawed: “Honestly, I find that a difficult premise to accept. Let’s imagine that access to widely-used cloud services like ChatGPT, Gemini, and others is suddenly restricted. As long as we remain committed to the open principles of the Internet, anyone with reasonably capable hardware should still be able to train their own models and use them for research. In fact, I believe the question itself is conceptually flawed. Indeed, it assumes that we would willingly forgo our free will (assuming it exists), which runs counter to the spirit of scientific progress. Anyhow, you guys should definitely asks ChatGPT or whatever models you feel like using :)”(R73)
Three themes about regulation, if applied to AI use in research, emerged from the responses. Regulation requires human judgment. A cross cutting theme across many of the responses is that just as human input is used during GenAI use, it is also often needed as its use is regulated, but some mentioned that our community really needs to “lean in” to using GenAI but should use regulation wisely: “Yes, but I would like to see some kind of smart regulation, which enables innovation while minimizes harm. No blunt restrictions” (R58). Researchers’ final thoughts reinforced this stance, with one respondent noting that “prohibiting these tools is just impractical” and another called instead for “a culture of scientific integrity with respect to their use” (R219). Epistemic & Scientific Integrity Concerns. For those that answered it should be regulated, many participants mentioned concerns about the rigor, reproducibility, and transparency of research: “Yes, GenAI use in Software Engineering research should be regulated to ensure ethical practices, transparency, and reproducibility, while allowing room for innovation and development” (R142). There were further concerns about research integrity, notably misinformation risks (from biases and hallucinations) and correctness or accuracy of the research results supported by GenAI: “ I think there’s a need to Manuscript submitted to ACM
30
Trinkenreich et al.
regulate because I’ve reviewed papers where it appeared to me that the author had copied text directly out of ChatGPT and it made me suspicious that there could be inaccuracies or hallucinations in the paper. But I really don’t know how you go about regulating it. That’s a super tough problem. Seems like someone should do some research on solutions” (R7). Governance, Responsibility and Accountability. In addition to concerns about misinformation (mentioned above), there were concerns about the environment: “...GenAI use can introduce bias, security flaws, misinformation, and environmental impacts. Regulation would encourage responsible usage and proactive risk management”(R23). Public trust in our research was also mentioned: “Yes, to the extent that it is regulated in other fields. I believe it is okay to use generative AI for writing code to preprocess data or implement models, for example, as long as authors (a) check the code carefully and (b) are aware that they are liable for any mistakes the model makes that they did not detect. However, I am generally against its use in other parts of SE (and most other fields’) research. It is also important that as the scientific community, we maintain public trust in science, which may be harder if generative AI use were to be widespread” (R21). In the final question in the survey on their final thoughts, some researchers offered concrete proposals for how such governance could be operationalized. Some respondents underlined the challenge that “it is not always possible to create generic guidelines, as each domain is bound by certain issues” and suggested that “guidelines should be created based on the type of study, such as the SIGSOFT [empirical standards]” (R197). Others proposed a task-based decision matrix specifying acceptable GenAI use per research activity, for example “forbidden for reviewing, accepted for brainstorming, [allowed] for data labeling only when fulfilling a set of very specific requirements” (R19). To sustain such efforts, respondents called for dedicated sessions at major venues such as ICSE, FSE, and ASE to discuss community rules, emphasizing that “this needs to be done continuously” through “dedicated boards [. . . ] at all conferences and journals” (R19). 5
DISCUSSION
5.1
Emergent Tensions Across Findings
5.1.1 Productivity Tensions. Productivity gains were one of the most commonly mentioned opportunities of using GenAI. However, the potential productivity benefits were often tempered by a set of emerging tensions. Productivity-Effort Tension. Participants emphasized that using GenAI well is far from trivial, as “it takes extra effort and experience for researchers/reviewers to gain confidence, rather than creating shortcuts and bad outcomes” (R336). One respondent captured this tension in terms of expertise: “GenAI tools can be very useful in the hands of skilled researchers who know what they are doing and can distinguish between right and wrong information. However, in the hands of a novice, it has the potential to wreak havoc, leading to fabricated research and spreading misinformation” (R82).Another stated that “using GenAI sensibly requires a lot of effort that many researchers are not willing or capable of investing” (R158). These responses point to a paradox. The primary appeal of GenAI is efficiency, yet responsible use demands significant effort, expertise, and critical engagement. However, if the effort barrier is not addressed through training and community norms, the productivity promise of GenAI risks being realized at the expense of research quality. Productivity-Quality Tension. Quality concerns were also explicitly raised. One participant felt that GenAI output was “potentially biased, incomplete, or suffering from fixation effects” (R19). Another participant stressed that diminishing research quality affects the entire SE research community: “As community, we build upon previous research, we extend them, we corroborate them, we contradict them, and that’s how we have moved forward as a community. We always trusted previous research. But now I am at a juncture, where I cannot decide whether I trust the research anymore.” (R277) Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
31
These concerns illustrate a second tension: while GenAI may increase the volume or speed of research production, it may simultaneously undermine confidence in the integrity of the research ecosystem that relies on shared trust and cumulative progress. Productivity-Impact Tension. We also found evidence that researchers who use GenAI felt pressure to do so both to stay relevant and to secure funding (Sec. 4.1.2. This pressure may increase the amount of research being produced, but can also potentially skew agendas toward what is “AI-adjacent,” raising questions about long-term scientific impact vs. short-term output. One participant described the potential for more shallow research: “I think it will result in incorrect and shallow research in many cases.” (R74) This raises a third tension: long term research impact may be at risk when GenAI is used to amplify productivity alone. These tensions reinforce the calls for human-in-the-loop (Section 4.4.3) and governance (Section 4.5) mitigation strategies reported earlier in our findings. 5.1.2 Reliance-Control Tension. A related but distinct tension concerns who is behind the wheel steering the research process, and whether academics risk becoming passive consumers of GenAI outputs rather than active producers. Several respondents warned against giving over too much control to GenAI. One cautioned that researchers should “be careful of relying too heavily on it,” arguing that “there’s a lot of value in thinking deeply about the research and forming a strong understanding of your data, which could easily get lost by giving GenAI the wheel” (R331). Another observed that “the excessive use of GenAI by students and researchers shifts the actual research to the AI” (R77). A concrete example of this loss of control was offered by one respondent who reported that “authors of the paper that I recently reviewed apologized for paragraphs that were ‘AI-massaged,’ so the content was incorrect,” concluding that “GenAI might make us lazier and less responsible for the things we write” (R161). These responses reveal a tension between the convenience of delegation and the oversight that responsible use of GenAI demands. Without proper mitigation, over-reliance can occur. This tension also reinforces the human-in-the-loop mitigation strategy (Section 4.4.3). 5.1.3 Democratization-Inequity Tension. While participants noted that GenAI has the potential to make SE research more accessible, a lack of computing resources was also cited as a concern by many participants. These sentiments lead to another emerging tension. The cost of using the latest models on large scale SE data can become astronomical. While GenAI may make some information more accessible, it can also be leveraged best by those with large research budgets. This can lead to even greater inequities if reviewers demand the use of the latest models and more data in experiments. This also reinforces the need for governance (Section 4.5) to ensure reviewers do not amplify inequities. 5.1.4 Rapid Change-Stable Guidance Tension. Another tension that emerges is related to the need for stable guidance and regulations, balanced with the rapid change of the GenAI tooling landscape. Respondents discuss the need for guidance, rules, and regulation for GenAI use, a theme developed in the risk mitigations (Section 4.4.3) and in the calls for regulation and policy (Section 4.5). Each of the tensions mentioned above ultimately ties back to this need for shared expectations about how GenAI should be used responsibly in SE research. Yet, the GenAI landscape is rapidly changing. As a result, researchers are navigating a moving target: they desire durable community guidance, but any rules that are too prescriptive risk becoming outdated almost immediately. This creates a tension between the stability needed to maintain research quality and integrity, and the flexibility required to respond to continual technological shifts. The SE community must pursue guidance that is principled rather than tool-specific, focusing on research values—such Manuscript submitted to ACM
32
Trinkenreich et al.
as transparency, verification, and accountability—so that expectations remain relevant even as GenAI technologies continue to evolve. 5.2
The Competence Pipeline at Risk
While our participants highlighted many benefits and opportunities of using GenAI in SE research, as final thoughts many also cautioned that human expertise is needed to evaluate the outputs and leverage these benefits. For example, when describing GenAI opportunities, R63 said “It’s a great helper if and only if you already have a fundamental understanding.” One participant (R165) said, “The major threat of using GenAI summaries is to loose depth, particularly for young researchers”. These sentiments raise questions about the development of these fundamental research skills when students begin their research careers with GenAI. When sharing final thoughts, nine respondents explicitly pointed to the potential erosion of the research training pipeline, a concern that cuts across the risk findings (Section 4.4.2) and the education-related mitigation strategies (Section 4.4.3). Five respondents raised concerns about the impact on education and training. One expressed worry that “GenAI will enable researchers to conduct a significant portion of the research by themselves,” with the consequence that “researchers will require less low-level support from undergraduates and graduate students because it will take so much more time to teach them how to do and inspect the output [. . . ] than doing [it] themselves” (R62).This dynamic suggests that faculty may prioritize short-term productivity over mentorship, reducing students’ opportunities to develop research skills through hands-on practice. Related to this, four respondents highlighted the risk of skill degradation. One noted that “the use of GenAI moves researchers a step further away from the research. You don’t have to get your hands dirty in the data because a model can just summarize it” (R7).Another captured the tension between improved surface quality and diminished understanding: “maybe the English quality is becoming better and better, but we don’t know what we’re writing anymore” (R161). At the same time, respondents recognized that the solution lies not in avoiding GenAI but in rethinking how the next generation of researchers is trained. R23 argued for “educating the next generation of researchers, not just on how to use GenAI tools effectively, but also how to critically evaluate, audit, and even design them responsibly.” These findings suggest a cumulative risk: if researchers lack the training to use GenAI critically, they may be more likely to produce lower-quality outputs, which could erode trust in AI-assisted research. Addressing this risk requires integrating GenAI literacy into research training, as discussed in Section 4.4.3. 5.3
Implications for Practice and Research Opportunities
The tensions described above have several implications for software engineering researchers. First, the ProductivityEffort and Productivity-Quality tensions highlight that researchers should not expect efficiency gains without friction. Researchers must budget time for skill acquisition and embed QA processes when adopting GenAI tools as part of their research practice. Human oversight should be baked into all tasks where GenAI is used to ensure reliable and trustworthy research outputs. In addition, the community would benefit from shared experimentation. More studies should be conducted that evaluate GenAI’s strengths and limitations for various research tasks. Such studies should not only report accuracy or speed benefits, but reflect on how human oversight is embedded, the friction points that were encountered, and other learnings that can help to collectively reduce the barrier for responsible adoption and develop guidelines for methodological design. The competence pipeline risk described above risks degrading researcher skills if GenAI is overused and impacts research skill formation. This can further amplify the productivity-quality tension if skills erode to the point where Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
33
competent human oversight is at risk. Research supervisors and graduate programs must explicitly design opportunities for students to practice skills such as critical thinking, data collection and analysis, and writing without GenAI assistance. Developing strong foundational skills will help students leverage the productivity gains through GenAI use responsibly. Our results illustrate that researchers are using GenAI across all stages of the research pipeline, from idea generation to writing and reporting. This further amplifies the call for the need for methodological guidance for responsible GenAI use across the research pipeline. The community needs shared norms, templates, and standards on when and how GenAI can be used, how it should be validated, and what should be reported to preserve transparency and replicability. Notably, research design and data collection were the least commonly reported phases for GenAI use. These phases often require judgement and reasoning. Beyond better guidelines, we envision tooling that could be helping in aiding researchers to leverage GenAI in these use cases. For research design, specialised GenAI tools with knowledge of the ACM SIGSOFT empirical standards could be created to help guide researchers through the research design process. Rather than replacing human reasoning, the tools could be used as research partners helping to drive methodological structure in the research design phase. Being grounded in the existing empirical standards can also help to ensure the research design meets the expectations for soundness and rigor. Similarly, tools could be envisioned to support data collection, though safeguards will be needed. The type of tool will depend on the source of the data being collected. For example, mining software repository studies, GenAI can be used to generate data collection scripts and tools could be created to verify the accuracy of collected data. Such verification tools might automatically detect inconsistencies, flag suspicious patterns that suggest faulty scraping or API failures, or compare GenAI-generated extracted data against known repository data or structures. These kinds of guardrails would help ensure that GenAI augments data collection without introducing silent errors or biases. For data collection involving human subjects, such as interviews or questionnaires, GenAI could be used to pilot collection instruments to ensure misleading or biased questions are identified. However, GenAI cannot replace genuine perspectives of human participants, and caution is needed to ensure that the use of GenAI, even in piloting, does not introduce bias into the study design.
6
THREATS TO VALIDITY
We discuss the main threats to the validity of our study, organized along the categories proposed by Wohlin et al. [48].
Construct validity. A potential threat concerns whether the survey instrument adequately captures the construct of GenAI use in research. One respondent noted in the final thoughts question that the questionnaire “did not transport [the] distinction well” between using GenAI to assist the research process (e.g., writing, brainstorming) and using it as a research tool (e.g., generating test cases or analyzing data) (R204). We acknowledge that collapsing these distinct modes of use into a single set of questions may have reduced construct precision. Future refinements of the instrument could introduce separate question blocks for process- and tool-level assistance.
Internal validity. One threat to internal validity concerns the qualitative coding process. For each open-ended question, coding was performed by one author and subsequently reviewed by three other co-authors. Disagreements were resolved through meetings. While this process mitigates individual bias, having a single initial coder may have influenced the framing of each codebook. To support transparency and replicability, the codebooks and coded data are included in the replication package. Manuscript submitted to ACM
34
Trinkenreich et al.
External validity. Several respondents questioned the SE-specificity of the study in the final thoughts question. One stated that “the use of GenAI in SE research would [not] be any different from using it in any other field of computer science, or even beyond” (R182), while another argued that “many other researchers in other areas are using GenAI” and called for studying “the wider impact of GenAI on research activities” (R171). We acknowledge that many of our findings, particularly those related to productivity, trust, risk perception, and governance, likely reflect dynamics common across academic disciplines rather than being unique to SE. However, the purposive sampling strategy, targeting authors from leading SE venues, ensures that the findings are grounded in a well-defined research community. Investigating the extent to which our findings generalize to other disciplines remains a promising direction for future work. Conclusion validity. In the final thoughts question, one respondent suggested that additional questions could have probed “what new SE-related problems [are] introduced by GenAI” (R20), pointing to potential gaps in the topic coverage of our instrument. While the open-ended question compensated for such gaps, future studies could benefit from a broader set of questions targeting emergent concerns. 7
CONCLUSION
This paper presents the results of a large-scale survey of 457 software engineering researchers, providing an empirical characterization of how GenAI is being adopted and perceived across the SE research community. Our findings show that GenAI adoption is already widespread, with nearly three-quarters of respondents reporting its use for research. Yet this adoption is uneven across research stages. It usage is concentrated on writing support, summarization, and coding, while research design and data collection see much lower adoption of GenAI. Data analysis falls somewhere in between these two: although reported less frequently than writing, the qualitative responses in the survey reveal it is used for annotation and exploratory analysis, typically with human oversight of interpretation. Overall, researchers tend to delegate routine and repetitive work to GenAI while retaining control over tasks that require methodological judgment. Trust follows a similar pattern, with writing and dissemination receiving by far the highest trust levels. Our findings also uncover a set of tensions regarding the use of GenAI in SE research. Productivity gains, a recurring theme across perceived benefits and opportunities, are contrasted with concerns about the effort required for responsible use, the risk of quality degradation, and the possibility that institutional pressures may incentivize volume over depth. The competence pipeline is a related concern: if early-career researchers come to rely on GenAI before developing foundational skills in critical thinking, data analysis, and academic writing, the community risks eroding the very expertise on which responsible GenAI use depends. Regarding governance, almost all respondents agreed that some regulation is needed, though they stressed it should focus on principles rather than specific models. Peer review emerged as a particularly divisive topic, with views ranging from outright prohibition to conditional acceptance with safeguards. These positions reflect the difficulty of establishing norms, also because GenAI technology is still evolving rapidly. Finally, we contribute taxonomies of GenAI use cases, opportunities, risks, and mitigation strategies grounded in SE researchers’ own practices and perspectives. Together with the quantitative characterization of adoption, trust, and perceived impact, these findings provide an empirical baseline against which future shifts in practices, perceptions, and policies can be measured. Looking ahead, we see three priorities. First, the SE community would benefit from shared, evolving guidelines for GenAI use across the research pipeline, grounded in transparency, verification, and accountability. Second, graduate Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
35
training programs should be designed to ensure that students develop strong research skills alongside GenAI literacy, not as a substitute for it. Third, longitudinal studies are needed to track how the tensions identified in this paper play out as GenAI capabilities evolve and community norms are developed.
8
ACKNOWLEDGEMENTS
We would like to thank all of our participants who dedicated their time to answering the questionnaire.
REFERENCES [1] Toufique Ahmed, Premkumar Devanbu, Christoph Treude, and Michael Pradel. 2025. Can LLMs replace manual annotation of software engineering artifacts?. In 2025 IEEE/ACM 22nd International Conference on Mining Software Repositories (MSR). IEEE, 526–538. [2] Jens Peter Andersen, Lise Degn, Rachel Fishberg, Ebbe K Graversen, Serge PJM Horbach, Evanthia Kalpazidou Schmidt, Jesper W Schneider, and Mads P Sørensen. 2025. Generative Artificial Intelligence (GenAI) in the research process–A survey of researchers’ practices and perceptions. Technology in Society 81 (2025), 102813. [3] Sebastian Baltes, Florian Angermeir, Chetan Arora, Marvin Muñoz Barón, Chunyang Chen, Lukas Böhme, Fabio Calefato, Neil Ernst, Davide Falessi, Brian Fitzgerald, et al. 2025. Guidelines for empirical studies in software engineering involving large language models. arXiv preprint arXiv:2508.15503 (2025). [4] Sebastian Baltes and Paul Ralph. 2022. Sampling in software engineering research: a critical review and guidelines. Empir. Softw. Eng. 27, 4 (2022), 94. https://doi.org/10.1007/S10664-021-10072-8 [5] Muneera Bano, Rashina Hoda, Didar Zowghi, and Christoph Treude. 2024. Large language models for qualitative research in software engineering: exploring opportunities and challenges. Automated Software Engineering 31, 1 (2024), 8. [6] Cauã Ferreira Barros, Bruna Borges Azevedo, Valdemar Vicente Graciano Neto, Mohamad Kassab, Marcos Kalinowski, Hugo Alexandre D Do Nascimento, and Michelle CGSP Bandeira. 2025. Large language model for qualitative research: A systematic mapping study. In 2025 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Engineering (WSESE). IEEE, 48–55. [7] Virginia Braun and Victoria Clarke. 2006. Using thematic analysis in psychology. Qualitative research in psychology 3, 2 (2006), 77–101. [8] John L Campbell, Charles Quincy, Jordan Osserman, and Ove K Pedersen. 2013. Coding in-depth semistructured interviews: Problems of unitization and intercoder reliability and agreement. Sociological methods & research 42, 3 (2013), 294–320. [9] Rudrajit Choudhuri, Bianca Trinkenreich, Rahul Pandita, Eirini Kalliamvakou, Igor Steinmacher, Marco Gerosa, Christopher Sanchez, and Anita Sarma. 2025. What Guides Our Choices? Modeling Developers’ Trust and Behavioral Intentions Towards GenAI. In 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE). IEEE Computer Society, 624–624. [10] Ronnie de Souza Santos, Italo Santos, Maria Teresa Baldassarre, Cleyton Magalhaes, and Mairieli Wessel. 2025. An Investigation on How AI-Generated Responses Affect Software Engineering Surveys. arXiv e-prints (2025), arXiv–2512. [11] DORA Team. 2025. State of AI-Assisted Software Development. Technical Report. Google Cloud. https://dora.dev/research/2025/dora-report/ Accessed: April 2026. [12] Katia Romero Felizardo, Anderson Deizepe, Daniel Coutinho, Genildo Gomes, Maria Meireles, Marco Gerosa, and Igor Steinmacher. 2025. On the difficulties of conducting and replicating systematic literature reviews studies using LLMs in software engineering. In 2025 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Engineering (WSESE). IEEE, 20–23. [13] Katia Romero Felizardo, Márcia Sampaio Lima, Anderson Deizepe, Tayana Uchôa Conte, and Igor Steinmacher. 2024. ChatGPT application in Systematic Literature Reviews in Software Engineering: an evaluation of its accuracy to support the selection activity. In Empirical Software Engineering and Measurement. 25–36. [14] D Garrison, Martha Cleveland-Innes, Marguerite Koole, and James Kappelman. 2006. Revisiting methodological issues in transcript analysis: Negotiated coding and reliability. The Internet and Higher Education 9, 1 (2006), 1–8. [15] Marco Gerosa, Bianca Trinkenreich, Igor Steinmacher, and Anita Sarma. 2024. Can AI serve as a substitute for human subjects in software engineering research? Automated Software Engineering 31, 1 (2024), 13. [16] Yolanda Gil, Mark Greaves, James Hendler, and Haym Hirsh. 2014. Amplify scientific discovery with artificial intelligence. Science 346, 6206 (2014), 171–172. [17] GitHub. 2025. Octoverse 2025: The State of Open Source. https://github.blog/news-insights/octoverse/octoverse-a-new-developer-joins-githubevery-second-as-ai-leads-typescript-to-1/ Accessed: April 2026. [18] Jacqueline Harding, William D’Alessandro, NG Laskowski, and Robert Long. 2024. AI language models cannot replace human research participants. Ai & Society 39, 5 (2024), 2603–2605. [19] Xinyi Hou, Yanjie Zhao, Yue Liu, Zhou Yang, Kailong Wang, Li Li, Xiapu Luo, David Lo, John Grundy, and Haoyu Wang. 2024. Large Language Models for Software Engineering: A Systematic Literature Review. ACM Transactions on Software Engineering and Methodology (2024). https: //doi.org/10.1145/3695988 Manuscript submitted to ACM
36
Trinkenreich et al.
[20] Aleksi Huotala, Miikka Kuutila, Paul Ralph, and Mika Mäntylä. 2024. The promise and challenges of using LLMs to accelerate the screening process of systematic reviews. In Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering. 262–271. [21] Mia Mohammad Imran and Tarannum Shaila Zaman. 2025. OLAF: Towards Robust LLM-Based Annotation Framework in Empirical Software Engineering. arXiv preprint arXiv:2512.15979 (2025). [22] Qusai Khraisha, Sophie Put, Johanna Kappenberg, Azza Warraitch, and Kristin Hadfield. 2024. Can large language models replace humans in systematic reviews? Evaluating GPT-4’s efficacy in screening and extracting data from peer-reviewed and grey literature in multiple languages. Research Synthesis Methods (2024). [23] Dmitry Kobak, Rita González-Márquez, Emőke-Ágnes Horvát, and Jan Lause. 2025. Delving into LLM-assisted writing in biomedical publications through excess vocabulary. Science Advances 11, 27 (2025), eadt3813. https://doi.org/10.1126/sciadv.adt3813 [24] Matheus De Morais Leça, Lucas Valença, Reydne Santos, and Ronnie De Souza Santos. 2025. Applications and implications of large language models in qualitative analysis: A new frontier for empirical software engineering. In 2025 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Engineering (WSESE). IEEE, 36–43. [25] Jenny T Liang, Carmen Badea, Christian Bird, Robert DeLine, Denae Ford, Nicole Forsgren, and Thomas Zimmermann. 2024. Can gpt-4 replicate empirical software engineering research? Proc. of the ACM on Software Engineering 1, FSE (2024), 1330–1353. [26] Weixin Liang, Yuhui Zhang, Hancheng Cao, Binglu Wang, Daisy Yi Ding, Xinyu Yang, Kailas Vodrahalli, Siyu He, Daniel Scott Smith, Yian Yin, Daniel A. McFarland, and James Zou. 2024. Can Large Language Models Provide Useful Feedback on Research Papers? A Large-Scale Empirical Analysis. NEJM AI 1, 8 (2024). https://doi.org/10.1056/AIoa2400196 [27] Ziming Luo, Zonglin Yang, Zexin Xu, Wei Yang, and Xinya Du. 2025. LLM4SR: A Survey on Large Language Models for Scientific Research. CoRR abs/2501.04306 (2025). https://doi.org/10.48550/arXiv.2501.04306 [28] Marshall McLuhan. 1977. Laws of the Media. ETC: A Review of General Semantics (1977), 173–179. [29] Sharan B Merriam and Elizabeth J Tisdell. 2015. Qualitative research: A guide to design and implementation. John Wiley & Sons. [30] Courtney Miller, Paige Rodeghero, Margaret-Anne Storey, Denae Ford, and Thomas Zimmermann. 2021. "How Was Your Weekend?" Software Development Teams Working From Home During COVID-19. In 2021 IEEE/ACM 43rd International Conference on Software Engineering (ICSE). 624–636. https://doi.org/10.1109/ICSE43902.2021.00064 [31] Tanisha Mishra, Edward Sutanto, Rini Rossanti, Nayana Pant, Anum Ashraf, Akshay Raut, Germaine Uwabareze, Ajayi Oluwatomiwa, and Bushra Zeeshan. 2024. Use of large language models as artificial intelligence tools in academic research and publishing among global clinical researchers. Scientific Reports 14, 1 (2024), 31672. [32] Cristina Martinez Montes, Robert Feldt, Cristina Miguel Martos, Sofia Ouhbi, Shweta Premanandan, and Daniel Graziotin. 2025. Large Language Models in Thematic Analysis: Prompt Engineering, Evaluation, and Guidelines for Qualitative Software Engineering Research. arXiv preprint arXiv:2510.18456 (2025). [33] Tatiane Ornelas, Allysson Allex Araújo, Júlia Araújo, Marina Araújo, Bianca Trinkenreich, and Marcos Kalinowski. 2025. LLM-Assisted Thematic Analysis: Opportunities, Limitations, and Recommendations. arXiv preprint arXiv:2511.14528 (2025). [34] Zeeshan Rasheed, Muhammad Waseem, Aakash Ahmad, Kai-Kristian Kemell, Xiaofeng Wang, Anh Nguyen-Duc, and Pekka Abrahamsson. 2024. Can Large Language Models Serve as Data Analysts? A Multi-Agent Assisted Approach for Qualitative Data Analysis. CoRR abs/2402.01386 (2024). https://doi.org/10.48550/ARXIV.2402.01386 [35] Daniel Russo, Sebastian Baltes, Niels van Berkel, Paris Avgeriou, Fabio Calefato, Beatriz Cabrero-Daniel, Gemma Catolino, Jürgen Cito, Neil Ernst, Thomas Fritz, et al. 2024. Generative ai in software engineering must be human-centered: The copenhagen manifesto. J. Syst. Softw. 216 (2024), 112115. [36] Mary Shaw. 2002. What makes good research in software engineering? International Journal on Software Tools for Technology Transfer 4, 1 (2002), 1–7. [37] Chenglei Si, Diyi Yang, and Tatsunori Hashimoto. 2025. Can LLMs Generate Novel Research Ideas? A Large-Scale Human Study with 100+ NLP Researchers. In Proceedings of the 13th International Conference on Learning Representations (ICLR). [38] Stack Overflow. 2025. 2025 Developer Survey. https://survey.stackoverflow.co/2025/ Accessed: April 2026. [39] Igor Steinmacher, Jacob Mcauley Penney, Katia Romero Felizardo, Alessandro F Garcia, and Marco A Gerosa. 2024. Can ChatGPT emulate humans in software engineering surveys?. In Proc. of the 18th ACM/IEEE Int’l. Symposium on Empirical Software Engineering and Measurement. 414–419. [40] Margaret-Anne Storey, Neil A Ernst, Courtney Williams, and Eirini Kalliamvakou. 2020. The who, what, how of software engineering research: a socio-technical framework. Empirical Software Engineering 25, 5 (2020), 4097–4129. [41] Eugene Syriani, Istvan David, and Gauransh Kumar. 2024. Screening articles for systematic reviews with ChatGPT. Journal of Computer Languages 80 (2024), 101287. https://doi.org/10.1016/j.cola.2024.101287 [42] Christoph Treude and Margaret-Anne Storey. 2025. Generative ai and empirical software engineering: A paradigm shift. In 2025 2nd IEEE/ACM International Conference on AI-powered Software (AIware). IEEE, 233–239. [43] Bianca Trinkenreich, Fabio Calefato, Geir Hanssen, Kelly Blincoe, Marcos Kalinowski, Mauro Pezzè, Paolo Tell, and Margaret-Anne D. Storey. 2025. Get on the Train or be Left on the Station: Using LLMs for Software Engineering Research. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering, FSE Companion 2025, Clarion Hotel Trondheim, Trondheim, Norway, June 23-28, 2025, Leonardo Montecchi, Jingyue Li, Denys Poshyvanyk, and Dongmei Zhang (Eds.). ACM, 1503–1507. https://doi.org/10.1145/3696630.3731666 [44] Richard Van Noorden and Jeffrey M Perkel. 2023. AI and science: what 1,600 researchers think. Nature 621, 7980 (2023), 672–675. Manuscript submitted to ACM
Taking a Pulse on How Generative AI is Reshaping the Software Engineering Research Landscape
37
[45] Stefan Wagner, Marvin Muñoz Barón, Davide Falessi, and Sebastian Baltes. 2025. Towards evaluation guidelines for empirical studies involving llms. In 2025 IEEE/ACM International Workshop on Methodological Issues with Empirical Studies in Software Engineering (WSESE). IEEE, 24–27. [46] David Williams, Max Hort, Maria Kechagia, Aldeida Aleti, Justyna Petke, and Federica Sarro. 2025. Empirical and Sustainability Aspects of Software Engineering Research in the Era of Large Language Models: A Reflection. arXiv preprint arXiv:2510.26538 (2025). [47] Viggo Tellefsen Wivestad and Astri Moksnes Barbala. 2025. Attitudes Towards LLM Use Among Software Engineering Researchers: Results From A Two-Phase Survey Study. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering. 1531–1535. [48] Claes Wohlin, Per Runeson, Martin Höst, Magnus C. Ohlsson, Björn Regnell, and Anders Wesslén. 2012. Experimentation in Software Engineering. Springer, Berlin, Heidelberg. https://doi.org/10.1007/978-3-642-29044-2 [49] Ruoxi Xu, Yingfei Sun, Mengjie Ren, Shiguang Guo, Ruotong Pan, Hongyu Lin, Le Sun, and Xianpei Han. 2024. AI for social science and social science of AI: A survey. Information Processing & Management 61, 2 (2024), 103665. https://doi.org/10.1016/J.IPM.2024.103665 [50] Ting Zhang, Ivana Clairine Irsan, Ferdian Thung, and David Lo. 2025. Revisiting sentiment analysis for software engineering in the era of large language models. ACM Transactions on Software Engineering and Methodology 34, 3 (2025), 1–30. [51] Ruiyang Zhou, Lu Chen, and Kai Yu. 2024. Is LLM a Reliable Reviewer? A Comprehensive Evaluation of LLM on Automatic Paper Reviewing Tasks. In Proceedings of the 2024 Joint International Conference on Computational Linguistics, Language Resources and Evaluation, LREC/COLING 2024. ELRA and ICCL, 9340–9351.
Manuscript submitted to ACM