How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming GABRIELLE O’BRIEN, University of Michigan, USA REED MILEWICZ, Sandia National Laboratories, USA NASIR EISTY, University of Tennessee, Knoxville, USA Generative AI has entered research programming, yet there is little evidence about which tasks researchers hand to it or how they
arXiv:2609.22049v1 [cs.SE] 18 Sep 2026
decide whether its code is correct. We draw on 527 free-text responses to a 2025 survey of researchers who write code, most of them at U.S. universities. In each response, a researcher recounts a single task from their own work, the way they used an AI tool for it, and what they did to assess the result. We coded the task and the evaluation strategies reported, and related both to programming experience, research area, and confidence ratings. Use was concentrated in five tasks: data handling, visualization, debugging, mathematical/scientific computing, and statistical analysis. Evaluation was informal and individual. Over half of accounts described running the generated code, while automated tests and review by another person were rare. Use cases and evaluation strategies varied little with programming experience, but confidence did: Less experienced programmers trusted the AI more than themselves, and experienced programmers the reverse. Evaluation confidence was not associated with the strategies reported. Its strongest correlates were confidence in the tool and in oneself. Validating AI contributions to scientific code rested largely on individual judgment, outside shared infrastructure for testing or review. Interfaces could support task-appropriate evaluation rather than leave it to the user. CCS Concepts: • Human-centered computing → Empirical studies in HCI; • Software and its engineering → Software verification and validation. Additional Key Words and Phrases: generative AI, research software, scientific programming, code generation, verification, validation, survey, content analysis
1
Introduction
Modern science runs on software, and much of that software is written by scientists rather than by software engineers [25, 45]. A crisis of productivity and credibility in research software in the 2000s [20, 49] brought training programs [19, 54], open science policies [32], and a research software engineering workforce [3, 11]. Even so, much scientific programming is still done by what Segal [45] called “professional end-user developers”: domain experts who program in service of their science. Knowing whether code meets its scientific requirements is hard, even for expert developers of research software [22, 24, 43]. And the cost of not knowing can be high: coding errors that slipped through have forced retractions of published findings in healthcare [21] and public policy [23], among other fields [22]. Into this setting have arrived tools that generate working code from natural language. Scientists and research software engineers (RSEs) have adopted them quickly [10, 36, 56], and the tools are increasingly capable of taking on more of the work, from inline completion and conversational chatbots to agents that execute code and iterate on their own [8]. Deciding how and when to rely on an automated system is a long-standing problem. Research on automation holds that trust should match what the system can actually do [29], but in practice people may tend to over-rely on automated aids [13]. Among knowledge workers, for example, confidence in a generative AI tool is associated with less scrutiny of its output [28]. Scientific programmers often write code to explore data or simulate systems that cannot be observed Authors’ Contact Information: Gabrielle O’Brien, [email protected], University of Michigan, Ann Arbor, Michigan, USA; Reed Milewicz, rmilewi@ sandia.gov, Sandia National Laboratories, Albuquerque, New Mexico, USA; Nasir Eisty, [email protected], University of Tennessee, Knoxville, Knoxville, Tennessee, USA.
1
2
O’Brien et al.
directly, so they do not already know what the correct result should look like. Now they must also contend with generated code that may be plausible but wrong. Furthermore, models signal their own uncertainty poorly [47], so judging the output becomes the user’s job (and one that can consume a large share of developers’ time [33, 55]). In professional software engineering, automated tests and structured code review exist to detect issues with programs that may escape an individual’s notice. In the scientific community, where verification is rarely standardized or supported by shared infrastructure [7, 15], deciding whether generated code is acceptable may often be a matter of individual interactions with their AI tools of choice. How programmers work with AI tools has been studied through laboratory studies of developers and students [2, 33, 34, 41, 52, 55], surveys of developers [30], and observations of professional data analysts [14]. But the same qualities that make scientific programming distinctive make it unclear how far those findings generalize. Studies of scientists who program are still sparse, drawn from the more professionalized research software engineering community or from a small number of open-ended survey responses [5, 36, 56] (Section 2). Here, we examine 527 written accounts of generative AI use in research programming, asking not only what scientists delegate to these tools but how they judge the result and how confident they are in that judgment. An account is one respondent’s answers to a set of open-ended prompts about a single episode of AI use, together with the confidence ratings attached to it. The accounts were collected during a 2025 cross-institutional survey of scientists who program; the survey’s descriptive results are published separately [37], and this paper is the first to analyze the open-ended section. Each respondent described one specific, recent task for which they used their primary AI tool, how they used it, and how they determined whether the output was acceptable, using prompts adapted from Lee et al.’s survey of knowledge workers [28]. Respondents also rated their confidence in doing the task alone, in the tool’s ability to do it, and in their own evaluation of the output. We coded each account along two dimensions—the scientific task motivating the request and the strategies reported for evaluating the result—and related those codes to programming experience, research area, and the confidence ratings. Throughout, we analyze what respondents say they do in a single recalled episode, which may not match their exact behavior. We ask: • RQ1. What tasks do scientists who program delegate to generative AI, and how do they report evaluating the output? • RQ2. Do reported use cases and evaluation strategies vary with programming experience or research area? • RQ3. How is confidence—in the tool, in oneself, and in one’s evaluation of the output—related to experience, to the task, and to the evaluation strategies reported? Our main findings are as follows: (1) Use is concentrated in five tasks—data handling, visualization, debugging, mathematical/scientific computing, and statistical analysis—which together account for 76% of accounts that received a use-case code (Section 4.1). (2) Reported evaluation is informal and individual. Over half of accounts describe running the generated code (52.8%), followed by inspecting its output, reading the code, and inspecting visualizations. Automated tests are mentioned in 3% of accounts and review by another person in 2% (Section 4.1). (3) Neither use cases nor evaluation strategies vary much with programming experience. The few detectable differences are by research area, in expected directions (Section 4.2). (4) Confidence does vary with experience. For the episode they described, less experienced programmers rated the tool above themselves, while more experienced programmers rated themselves above the tool (Section 4.2).
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming
3
(5) Respondents’ confidence in their evaluation is not reliably associated with the evaluation strategies they reported. Evaluation confidence is more strongly correlated with their confidence in the tool and in their own ability to do the task alone (Section 4.2). Together, these findings suggest that in the episodes respondents described, validating AI contributions to scientific code rested largely on individual judgment, exercised outside shared infrastructure for testing or review. We discuss what this implies for tools that could support verification where correctness is not visible in the output, for training that makes experienced scientists’ informal checks explicit, and for the integrity of research that increasingly depends on machine-generated code (Section 5). 2
Related Work
2.1
Programming with AI code generation tools
HCI and software engineering researchers have examined how programmers work with tools that generate code from natural language, which Sarkar et al. [44] argue is a distinct activity from conventional programming. Vaithilingam et al. [55] found that most participants preferred GitHub Copilot to conventional autocomplete but struggled to understand, edit, and debug the longer blocks of code it produced. Barke et al. [2] developed a grounded theory of Copilot use in which programmers alternate between an acceleration mode, where the tool completes code the programmer already had in mind, and an exploration mode, where it is used to discover approaches. Evaluation strategies differed between the two modes. Mozannar et al. [33] modeled where developers’ time goes during Copilot sessions, finding that a substantial share is spent verifying suggestions rather than writing code, and Tang et al. [52] used eye tracking and IDE logs to characterize how developers validate and repair generated code. Beginners appear to struggle most: Zi et al. [60] found that novices had difficulty understanding code generated by LLMs even when it was correct. Outside the laboratory, Liang et al. [30] surveyed 410 developers about their use of AI programming assistants, finding that the most common motivations were reducing keystrokes and finishing tasks faster, while the most common reasons for not using suggestions were that the generated code did not meet requirements or that developers lacked control over the output. Evidence on productivity is mixed: A controlled experiment found that developers completed a task substantially faster with Copilot [40], while a field experiment with experienced open-source developers found that AI tools slowed them down [4]. 2.2
Verifying AI output and calibrating reliance
Whether to accept an AI output is a reliance decision. Research on trust in automation treats calibrated trust, in which reliance tracks the system’s actual reliability, as the goal rather than maximal trust [29]. Overreliance on decision support is well documented across domains [13]. Experiments with AI decision aids find that explanations can raise reliance whether or not the AI is right [1], and that forcing users to engage with the task before seeing the AI’s answer reduces overreliance at a cost in effort [6]. People also rely on an AI more when checking its output themselves is costly [58]. Users may also be poorly placed to judge that reliability, because they overestimate what language models know [48]. Among software developers, trust in code generation tools rests on the tool’s perceived ability, integrity, and benevolence and varies with the context of use [59], and is shaped by the experiences peers share [9]. A parallel literature examines how users check AI-generated data analyses. Gu et al. [14] observed 22 professional analysts verifying AI-generated analyses and found that they began with procedure-oriented checks (what did the tool do?) and shifted to data-oriented checks (does the result make sense?) once something looked off. Studies of end-user
4
O’Brien et al.
programmers disagree about how well such checks work: Inspecting the shape and contents of data objects by eye has been reported both as a successful check [14] and as a route to overconfidence in incorrect results [42], a pattern with a long history in end-user programming [39]. In a survey of 319 knowledge workers, Lee et al. [28] found the same asymmetry between confidence in the tool and confidence in oneself, and that generative AI shifts critical effort toward verifying and integrating the tool’s output rather than producing one’s own. Storey [50] argue that code produced without the programmer’s full understanding accrues a distinct kind of debt in comprehension and intent. Novices and end-user programmers appear especially exposed. Prather et al. [41] observed beginners accepting generated code without validating it, and Nguyen et al. [34] and Liu et al. [31] found that non-experts struggle to steer code generation through prompts. These findings are consistent with the long-standing habit of end-user programmers to tweak code they do not fully understand until its output looks right [27]. Tankelevitch et al. [53] frame deciding when and how thoroughly to check as a metacognitive demand these tools place on the user. Interface cues that flag likely errors can direct programmers’ attention to them, but highlighting tokens by generation probability, as commercial tools do, did not change behavior in one study [57]. 2.3
Scientists who program
Much research software is written by scientists themselves. In a 2014 survey of 417 researchers at UK universities, 56% developed their own software, and a fifth of those had no training in software development [17]. These scientists write code to do science rather than to build software [18], and their practices differ from industrial software engineering for reasons that are often legitimate. They code to explore data and ideas rather than to build a product [51], requirements are discovered as the software evolves alongside the science [24, 46], and testing is informal or absent by conventional measures [7, 12], although many scientists hold a broader notion of verification and validation that covers the mathematics their code implements and the physical experiments it models [35]. Much of their code lives in scripts and computational notebooks, where it is run in short interactive fragments rather than strictly as written [16, 26], and testing may happen through expertly curated diagnostic plots rather than a battery of tests [38]. Research software engineers (RSEs), a professional workforce that builds and maintains research software as its job [3, 11], adopt more conventional practices, although even among them testing and review are far from universal [7]. It is the scientists writing their own code, not RSEs, who concern us here. Tools for this population have to fit its needs and values rather than import industrial practice wholesale. This paper asks how generative AI tools are being fitted into these practices: what scientists delegate to them, and how they come to trust the output. Evidence that scientists have adopted these tools is accumulating. In a 2024 survey of over 6,000 researchers at two large German research organizations, writing code was among the two most common uses of AI, reported by 43.2% of respondents [10]. In a 2026 survey of faculty and staff who actively develop or maintain research software at one U.S. university, about a third reported using generative AI tools [5]. Surveying the research software engineering community, Van Tuyl [56] found that roughly four in five reported using AI tools on the job. Why programmers at different experience levels use these tools — to attempt work they could not do alone, or to speed up work they could — is not yet clear [28]. Beyond adoption rates, less is known about how these tools are used in practice. Interviews with 14 scientists who programmed with AI assistance suggest that the tools often function as a substitute for documentation, and that verification is largely informal: running code and inspecting the output, reading line by line, or asking the tool to explain itself [36]. Two surveys have since coded free-text accounts of AI use in research software: Among research software engineers and adjacent staff, common themes were clarifying tasks and language, code generation, refactoring, and
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming
5
retrieving reference information for undocumented functions [56]. Among developers at one university, the themes were scaffolding, debugging, natural-language data transformation, and cognitive offloading [5, from at most 24 responses]. The present study adds a larger, cross-disciplinary corpus in which each account is coded for both the task and the strategy used to evaluate the output. 3
Methods
3.1
Survey
We report a mixed-methods study: a qualitative content analysis of written accounts, followed by quantitative analysis relating the resulting codes to respondents’ experience, research area, and confidence ratings. The data analyzed here was collected during a 2025 survey of research scientists who program. The full survey instrument and several descriptive results (adoption rate, tool preferences, and perceived productivity) are reported in a separate publication [37]; the open-ended accounts analyzed here have not been reported before. Briefly, the instrument comprised seven sections administered in Qualtrics (estimated 10–15 minutes), covering demographics, programming background, organizational coding practices, generative AI tool experience, perceived productivity, an open-ended use-case section, and reasons for non-adoption (if applicable). Recruitment was through mailing lists and online communities for scientific programming (e.g., the US Research Software Engineering association, pyOpenSci, and a targeted list at the authors’ institution, the University of Michigan). Because the link was shareable, the response rate is unknown. The study was approved by the authors’ institutional review board at the University of Michigan, responses were anonymous, and no IP addresses were collected. Of 1,272 responses collected between July 10 and August 25, 2025, we excluded incomplete responses, those failing consent or eligibility screens, and respondents who never program in their research, leaving 868 in the survey’s analytic sample. This analysis concerns the open-ended use-case section, shown only to respondents who reported using a generative AI tool in their research-related programming (those who had never tried such tools or reported that they had given up were routed to a non-adoption branch of the survey instead). After selecting their primary tool, eligible respondents answered three open-ended prompts. The format of this section, in which a respondent describes one specific recent example of AI use and then rates their confidence about it, was adapted from a survey of knowledge workers by Lee et al. [28]; we reworded the items for research programming. The prompts were: • Use case: “Think about one specific, real-world example of how you used your primary generative AI tool while doing research-related programming. What were you trying to achieve?” • Tool use: “How did you use the tool in this example? If possible, please include any prompts. . . ” • Evaluation: “How did you determine if the output of the generative AI tool was acceptable?” Respondents also rated three confidence items on a 5-point scale (1 = “not at all confident,” 5 = “extremely confident”). These are the three confidence constructs from Lee et al. [28]: confidence in doing the task without generative AI, confidence in the tool’s ability to do it, and confidence in one’s own ability to evaluate the output. That study found that the first and second were associated in opposite directions with critical evaluation of AI output, which motivated retaining all three here. • Solo confidence: confidence in doing the task without generative AI • GenAI confidence: confidence in the tool’s ability to do the task • Evaluation confidence: confidence in evaluating the tool’s output in the course of normal work
6
O’Brien et al. There were 527 accounts from the use-case section, which were the target of our qualitative analysis.
3.2
Qualitative analysis
We conducted a qualitative content analysis of open-ended survey responses describing how researchers used generative AI tools for programming tasks. We developed two parallel codebooks: a use-case codebook capturing the primary programming task for which the respondent used a generative AI tool (e.g., debugging, data handling, visualization), and an evaluation codebook capturing how respondents verified or validated AI-generated output (e.g., running code, inspecting output, consulting reference documentation). To generate the codebooks, the first author first conducted a round of open coding of 106 randomly selected (20%) accounts and drafted each codebook from that round. Using the draft codebooks, all three authors conducted a round of coding on a new sample of 25 responses, then met to discuss disagreements. These discussions led to refinements in code definitions (for example, the category “Systems & Hardware” was expanded to include creating code to interact with high-performance computing systems, and “Code comprehension” was clarified to refer only to trying to understand code written by another person and not an AI tool). We also clarified rules for interpreting common ambiguous words. The word “test,” for example, occurred frequently in contexts such as “I tested the suggestion” or “it compiled and worked when I tested it.” “Testing” can have a specific meaning in software development, referring to a “harness” of checks that are run after changes to the codebase (often with some automation). Based on prior research showing that this form of testing is infrequent in scientific software development [7, 12] and our best judgment of the contexts in which the term “test” typically occurred in survey responses, we decided that phrases such as “I tested the suggestion” would be labeled “Run code”, referring to manually initiating code execution. We applied the “test suite” label only when the account specifically indicated use of a test harness (such as “unit tests” or “integration testing”). After this discussion, we also identified two interpretive rules: first, code assignments must be grounded in explicit textual evidence. In practice, this means that if a person describes making a plot with AI assistance but reports only that they looked at the resulting plot, we applied only the label “Inspect visualization,” even though having a visual artifact to inspect implies that they must have executed the generated code. We restricted ourselves to a content analysis of what respondents reported as their evaluation strategy, since we could not observe the steps they actually took. A second interpretive rule is that the use case should be coded according to the primary reason the respondent describes using the AI tool. For example, if a respondent indicates that they used ChatGPT to help reshape a dataframe from long to wide format but had to troubleshoot its suggestions briefly, the account would be coded only as “Data handling,” not “Debugging.” We then conducted a second round of independent coding with another 25 randomly selected, previously uncoded accounts. We met to discuss codes and made a few minor refinements to the codebook (for example, adding a new code “Validation activities” for use cases involving validating the correctness of an existing codebase). After these refinements, the median pairwise Krippendorff’s alpha with MASI (Measuring Agreement on Set-valued Items) distance was 𝛼 = 0.76. The overall agreement across all three raters was 𝛼 = 0.70. At this point, all three authors had coded 50 accounts and we had reached a consensus on the codebook definitions. The first author annotated the remaining accounts with the finalized codebooks and applied them retroactively to the accounts from the initial open-coding round. During this round, we flagged 16 accounts for discussion because of ambiguous language, and all three raters met again to resolve their codes by consensus.
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming
7
Not every response could be coded. Some respondents skipped questions in this survey section, and others gave answers too vague to label confidently (describing a use case as simply “coding”). We did not apply codes when we judged that too little information was present. 3.3
Quantitative analysis
We follow several statistical conventions for our quantitative analyses. Programming experience is measured in years and analyzed as log10 (years + 1). Where helpful, we back-transformed to years for interpretation. Confidence items are 1–5 ratings. All two-group comparisons are Welch’s 𝑡-tests, which do not assume equal variances (group sizes here are often very unequal). Within each family of per-code tests we control the false discovery rate with the Benjamini– Hochberg (BH) procedure and report both uncorrected and BH-adjusted 𝑝-values. For statistical power, per-code analyses are restricted to the five most frequent codes in each family (use cases: debugging, data handling, visualization, mathematical/scientific computing, statistical analysis; evaluation strategies: run code, inspect output, read code, inspect visualization, domain knowledge/intuition). Analyses were conducted in R. Quantitative analysis scripts were written with the assistance of Claude Code. The authors specified the statistical tests and chart types, and used Claude Code for assistance in implementation. 4
Results
Of 527 use-case accounts annotated, 21 (4.0%) could not be assigned codes from either codebook and were excluded for data quality, leaving an analytic sample of 𝑁 = 506, that is, every account that received at least one use-case or evaluation code. Within this sample, 474 accounts received at least one use-case code and 470 at least one evaluation code. The subsets differ because a few accounts described only a use case or only an evaluation strategy. Some accounts carried more than one label per codebook, as respondents sometimes described multiple use scenarios or evaluation strategies. Before describing major themes from the qualitative analysis, we briefly summarize the demographics of respondents to contextualize their accounts. Demographic information Table 1 summarizes the analytic sample. Respondents were concentrated in U.S. higher education, with 97.6% based in the United States and 93.7% at higher-education institutions. Men somewhat outnumbered women (53.6% to 42.1%). Most respondents were early-career: student research assistants were the largest group (40.5%), and faculty made up 16.6%. The most-represented research areas were the life sciences (28.7%), engineering (20.6%), the social sciences (10.7%), and the physical sciences (10.3%). Programming practices Respondents were active programmers. Median programming experience was 6 years (IQR 4–10.5, max 50), and 84% programmed at least weekly (50.2% daily, 34.2% weekly). Python (70.4%) and R (56.3%) dominated language use, followed by MATLAB (30.8%), Bash (20.6%), and C++ (15.0%). The long tail covered Stata, JavaScript, C, Fortran, Java, Julia, and Rust. By design, only survey respondents who indicated that they used a generative AI tool for programming were asked to provide a use case. Respondents overwhelmingly named ChatGPT as their primary tool (61.7%), followed by GitHub Copilot (11.7%), institution-provided custom tools (5.7%), Google Gemini (5.5%), and Claude (4.3%, not including Claude Code). Roughly half the sample (50.7%) reported using version control “about half the time” or more, while
8
O’Brien et al.
Table 1. Respondent characteristics for the analytic sample (𝑁 = 506). Percentages use 𝑁 as the denominator except where noted. Languages were multi-select. Characteristic
𝑛
%
Country United States Other country Not reported
494 9 3
97.6 1.8 0.6
Organization Higher education National laboratory Private sector Other
474 12 6 14
93.7 2.4 1.2 2.8
Gender Man Woman Non-binary or gender-diverse Prefer not to say Not reported
271 213 13 8 1
53.6 42.1 2.6 1.6 0.2
Position Student research assistant Research staff Faculty Post-doctoral researcher Research software engineer Other
205 125 84 76 10 6
40.5 24.7 16.6 15.0 2.0 1.2
Research area Life sciences Engineering Social sciences Physical sciences Computer and information sciences Psychology Mathematics and statistics Geosciences and ocean sciences Other
145 104 54 52 44 26 22 16 43
28.7 20.6 10.7 10.3 8.7 5.1 4.3 3.2 8.5
Characteristic
𝑛
%
Programming frequency Daily Weekly Monthly Less than once a month
254 173 45 34
50.2 34.2 8.9 6.7
Programming experience (years) Median (IQR) Not reported
6 (4–10.5) 11 2.2
Languages used (multi-select) Python R MATLAB Bash C++
356 285 156 104 76
70.4 56.3 30.8 20.6 15.0
Primary generative AI tool ChatGPT GitHub Copilot Custom tool provided by organization Google Gemini Claude Cursor Microsoft Copilot Claude Code Perplexity Other Not reported
312 59 29 28 22 8 8 7 6 16 11
61.7 11.7 5.7 5.5 4.3 1.6 1.6 1.4 1.2 3.2 2.2
regular use of formal testing and review practices was substantially lower: code review (33.6%), unit tests (28.7%), system tests (18.6%), and regression tests (17.6%). 4.1
Qualitative analysis
We identified 13 use cases for generative AI in research programming (Figure 1, left). The five most common use cases — data handling (𝑛 = 121, 23.9%), visualization (𝑛 = 100, 19.8%), debugging (𝑛 = 88, 17.4%), mathematical/scientific computing (𝑛 = 59, 11.7%), and statistical analysis (𝑛 = 55, 10.9%) — appeared in about three-quarters of accounts that received a use-case code, and so we focused our qualitative reporting on these five. The remaining eight use cases (translation, code comprehension, optimization, user interface work, refactoring, validation activities, documentation, and systems & hardware tasks) each appeared in a smaller share of accounts. Most responses described a single use case (75.7%; mean 1.15 codes per account).
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming
9
In parallel, we coded the strategies that respondents reported using to evaluate the AI’s output in their given use case (Figure 1, right). Use of multiple strategies was common: 49.2% of accounts described a single evaluation strategy (𝑛 = 249), 34.8% described two (𝑛 = 176), and 8.9% described three or more (𝑛 = 45; mean 1.47 codes per account). Pairwise combinations of the most common strategies are shown in Figure 2. The modal strategy by a wide margin was running the code (𝑛 = 267, 52.8% of accounts), followed by inspecting outputs (𝑛 = 114, 22.5%), reading the code (𝑛 = 98, 19.4%), inspecting a visualization (𝑛 = 82, 16.2%), and drawing on domain knowledge or intuition (𝑛 = 56, 11.1%). Tables 2 and 3 give the definition, count, and an example account for every code in each codebook. Table 2. Use-case codebook: the task motivating the request, coded from the primary reason that the respondent gave for using the AI tool. 𝑛 (%) is accounts in the analytic sample (𝑁 = 506) carrying each code; quotes are verbatim answers to the use-case prompt, reproduced as written. Code
𝑛 (%)
Definition
Example account
Data handling
121 (23.9)
Move, manipulate, process, merge, or reformat datasets, usually as steps that prepare data for later analysis.
“doing a somewhat complicated join on two dataframes from different csvs”
Visualization
100 (19.8)
Create, modify, or format data visualizations, plots, graphs, or figures.
“How to modify my code so that my Python script could go through a list of colors and assign a different color to each bar in my bar chart plot”
Debugging
88 (17.4)
Troubleshoot why something did not work as expected; find and/or repair the root cause of an error or unexpected behavior in code.
“I was trying to run a logistic regression model and I was getting an error”
Mathematical/scientific computing
59 (11.7)
Domain-specific algorithms, such as solving integrals or systems of equations, implementing formal logic, and building control systems or simulations.
“find solution for a nonlinear equation system”
Statistical analysis
55 (10.9)
Implement methods for finding patterns in data, calculating statistics or running statistical tests, and fitting models to data; includes signal processing techniques for pattern recognition.
“Perform exploratory analysis on a gene expression matrix.”
Systems & hardware
49 (9.7)
Interface with physical devices, handle device drivers, control instruments for experiments; provision resources on HPC or cloud compute.
“I needed boilerplate code for reading / writing to a binary file. I wanted to use the c++ std lib rather than C utilities”
Translation
24 (4.7)
Convert code from one programming language (or library or framework) to another.
“convert a script I had written in python to MATLAB”
Code comprehension
23 (4.5)
Understand code written by another person (not AI-generated code).
“Used AI to understand the code from a previously published paper that was written in an unfamiliar language”
Optimization
21 (4.2)
Improve aspects of code performance, such as memory usage or speed.
“Speed up Euclidean distance generation for an R package”
User interface
14 (2.8)
Create interactive elements, dashboards, or web applications; build user-facing interfaces or interactive tools.
“I wanted to create a webpage to view some data I annotated.”
Refactor
11 (2.2)
Reorganize or simplify existing code without changing its core functions, usually to make code more readable or maintainable.
“reorganize code for more flexibility and readability”
Validation activities
9 (1.8)
Validate the correctness of an existing code base, for example by generating tests or checks for code the respondent already had.
“Create unit tests to verify/validate a new code feature”
Documentation
7 (1.4)
Write comments, docstrings, or documentation for code.
“I was trying to write documentation for some undocumented functions in an R package”
10
O’Brien et al.
Table 3. Evaluation codebook: strategies that respondents reported for judging whether the tool’s output was acceptable, applied only when grounded in explicit textual evidence. 𝑛 (%) is accounts in the analytic sample (𝑁 = 506) carrying each code; quotes are verbatim answers to the evaluation prompt, reproduced as written. Code
𝑛 (%)
Definition
Example account
Run code
267 (52.8)
Manually initiate code execution and check that it runs without producing errors; includes phrases such as “I tested it” absent any indication of a test suite.
“no errors, examining the output of the multiplication to confirm that it was what I expected”
Inspect output
114 (22.5)
Run the code and examine the resulting outputs, such as interactively inspecting a data object or looking at what is printed to console, file, or logs; excludes inspecting plots or figures.
“Its output data was similar to what I have been using.”
Read code
98 (19.4)
Read through the generated code, sometimes “line by line”.
“Reading through the output before using, and then trying the suggested fix myself.”
Inspect visualization
82 (16.2)
Evaluate the output by visually examining a plot or other visual output that the code produces.
“I would run the chunk of codes in my local environment and check the output, for example, checking whether the graph meet my requirement.”
Domain knowledge / intuition
56 (11.1)
Draw on the respondent’s own intuition or domain knowledge.
“reading it and seeing if I would have written same”
Check reference
45 (8.9)
Consult official documentation, a reference manual, published work, or a reference code implementation.
“Researched and used online documentation to confirm. I also wrote unit tests”
Compare to benchmark result
38 (7.5)
Compare the AI-generated output against a previous result (a benchmark).
“I ran the code and compared the outputs to known quantities.”
Ask AI to explain
14 (2.8)
Prompt the AI tool to explain the generated code or answer follow-up questions about it.
“I tried to read it thoroughly and also ask the tool to explain it to me so I know that each step that it is doing is what I wanted to do.”
Unit tests / test suite
14 (2.8)
Check how code performs on a test suite (usually unit tests); applied only when language specifically indicated a test harness.
“Created a test bench and verified that the results generated by this function is what I’m expecting.”
Colleague review
10 (2.0)
Check work with labmates, a professor, or another person with relevant expertise.
“Double check with supervisor and colleagues”
Check math by hand
6 (1.2)
Manually calculate mathematical expressions to check that the code gives the same result.
“I test for edgecases, and do some easy calculations to check if it gets the answer right.”
Cross-check with another AI
2 (0.4)
Use a different AI tool or model and compare outputs.
“I run it and check the results. I also use other gen AI tolls to test it”
We review some of the most common use cases, along with their most common evaluation strategies. The share of each use case’s accounts reporting each evaluation strategy is summarized in Figure 3. Data handling. Data handling, the most common use case, covered processing, cleaning, deduplicating, reshaping, reformatting, merging, and harmonizing data, as well as converting between data types, usually in preparation for downstream analysis or reporting. A recurring task was loading data from the file system into an interactive Python, R, or MATLAB session, often from specially formatted files produced by simulation software or by measurement devices such as microscopes, telescopes, and biomedical imaging instruments. Respondents also described complex filtering operations, in which they asked the tool to translate a set of logical filtering steps and conditions into implementations for common data handling libraries such as pandas. Within a programming environment, tasks included converting
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming 11
Fig. 1. Frequency of use-case codes (left; 𝑛 = 474 accounts with at least one use-case code) and evaluation codes (right; 𝑛 = 470 accounts with at least one evaluation code). Codes are not mutually exclusive, since an account may carry several codes from each codebook. Percentages in the text use the full analytic sample (𝑁 = 506) as the denominator.
Fig. 2. The five most common evaluation strategies and their pairwise co-occurrences. Horizontal bars (left) give each strategy’s overall count; vertical bars give the number of accounts carrying the indicated single code or code pair (connected dots), counted regardless of any other codes also present. Unlike a conventional UpSet plot, columns are not mutually exclusive intersections. 𝑛 = 470 accounts with at least one evaluation code.
between data structures (e.g., a table to an array, model output to a CSV file, strings to dates), merging tables, and applying transformations or imputations. Respondents’ accounts suggest two main reasons they reached for generative AI in this category. The first was the complexity of the data itself: harmonizing data collected from multiple instruments with different sampling rates, parsing specialized file formats tied to domain-specific instrumentation, and contending with missingness and irregularity
12
O’Brien et al.
Fig. 3. Evaluation strategies conditioned on use case. Each cell is the share of accounts carrying the row’s use-case code that also carried the column’s evaluation code. Rows can sum to more than 100% because accounts often reported several methods. Codes applied to fewer than 20 accounts are omitted, and rows and columns are ordered by overall frequency.
introduced during collection. The second was unfamiliarity with the specific libraries involved. Respondents often knew what operation they wanted but not how to express it in the tool at hand, as when one respondent used ChatGPT to work out how the Astropy package computes a coordinate transformation. Sometimes this unfamiliarity reflected a working environment rather than a knowledge gap: “I could have written a SAS macro to perform this action, but I’m the only person in my research group that uses SAS, meaning that if I want any of my coding work to be reproducible by my team, or even usable by my team, it has to be done in R. Everything I do in R takes a lot longer due to having less experience in it, so I consulted [institutional AI tool]”. About half of data handling accounts described more than one validation method, so a combination of strategies was typical. The most commonly reported strategy was running the generated code (e.g., “I tested the code against actual data to be sure that it worked”), although only a minority of accounts (26/121) reported running code without any other strategy. Some respondents deliberately ran generated code on samples or constructed inputs: “I created two mock data files with a small number of entries for which the answer was easily determined and ran them through the code.” It was not always clear from these accounts whether execution served to confirm that the code ran without error or to check its correctness against a known answer. Among respondents who explicitly reported inspecting outputs, some named the qualities they examined — “if it fully loaded the date ranges I was looking for”; “I opened the dataset and 1. first, checked if any value was imputed; 2. compared the imputed values with other values within each column” — while
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming 13 others described inspection only in general terms (“eyeballing the result,” “I got my desired result”), leaving unstated which attributes of the output they attended to. A smaller number of respondents described reading the code itself, typically in combination with testing it: “reading it over to make sure there were no obvious errors, then testing it.” Visualization. Visualization was the second most common use case. Respondents used generative AI both to produce less familiar chart types and to recall the syntax of common plotting libraries. Much of the work was incremental refinement of plots that the respondent had already generated: adjusting colors, working with heat maps, or layering an additional variable into an existing graph. As one respondent put it, ChatGPT “is especially useful when it comes to completing menial programming tasks such as formatting a figure I have already generated to include certain features that are beyond my knowledge of default plotting.” Others framed the tool less as a code generator than as a faster route into documentation. One respondent explained that “much of the MATLAB (and Python) syntax for generating plots is oddly specific. I would normally have to spend time digging through the documentation to remember exactly what the call is to thicken the line plot or find the hexadecimal color code. . . With an AI tool, I can ask that question and usually get an answer much faster than having to google it,” adding, “I’m not having ChatGPT write code so much as operate as a shortcut to digest the documentation.” The tasks ranged from publication-ready formatting (e.g., fonts, titles, labels, line thickness) to more substantive manipulation of how data was grouped or ordered for a plot. Respondents differed in how much they specified up front: some named the library or plot type they wanted, while others asked the tool to recommend one, as when a respondent used ChatGPT “to recommend different plotting functions in both libraries I was familiar with and libraries I had not heard of or considered.” The dominant evaluation strategy was inspecting the resulting graph, reported in just over half of visualization accounts (56/100). Where respondents specified what they looked at, they described checking the figure against their data and expectations: “I double-checked the graphs’ points with my data to see if they make sense, and the trend was right based on my judgment”; “the results were slightly different than what I expected, so my next step was looking into the violin plot API to understand what the keyword arguments were doing.” Several respondents treated visual output as almost self-validating. One wrote that “it’s entirely about the visualization so it’s literally visually validated,” another that “I visually determined it was the chart I wanted,” and a third that, for plot generation specifically, “I probably accept it immediately.” This treatment of visual output as self-validating stands in some contrast to data handling, where the correctness of an output may not be apparent from the output itself. Debugging. Debugging was the third most common use case, and the accounts here often centered on working with unfamiliar tools. Respondents frequently reached for generative AI when confronting a language, framework, or version they did not know well: “sometimes there are errors in C++, for example, that I am totally unfamiliar with because I come from a primarily Python-based background, and AI helps me narrow down the cause”; “I recently started coding models in JAGS in R, and it’s pretty common to run into compilation errors due to syntax mistakes, or improperly specified priors.” Version migration was a recurring trigger, as in a respondent adapting a script across MATLAB releases: “This code worked in MATLAB 2022 but fails in MATLAB 2025. Can you help me find which command or function is causing the error?” The prototypical workflow was to paste an error message and ask the tool to explain or resolve it. The errors often arose from data handling, analysis, or visualization work, and sometimes from systems-level issues such as CUDA setup or portability across operating systems. Some respondents noted that they turned to AI only after exhausting their usual resources. One described Stack Overflow as “always the first place I look for help” and consulted ChatGPT only when that failed. Beyond syntactic errors, a few respondents used the tool for more semantic debugging, including one account of locating a substantive error in a paper under review: “I accessed their code base, and it was
14
O’Brien et al.
about 2000 lines. It would be hopeless to locate the error. I described what was wrong in the paper, put their code into ChatGPT and it instantly located the exact error.” The dominant evaluation strategy, reported in most debugging accounts (57/88), was running the code, typically to check whether the original error had been resolved. Respondents described accepting a fix when the code “did not throw an error and still did what I wanted it to do,” or, more simply, when “I run it in R and see if I get the results I want.” Debugging stood out as the use case where running code most often appeared alone: 31 of 88 accounts reported it as the sole evaluation strategy, and only about a third (33/88) described more than one strategy overall. Some respondents, however, went beyond confirming that the error had cleared. They read the suggested code line by line and cross-referenced unfamiliar recommendations against documentation or forums: “if the error is caused by something I’m less familiar with, I’ll cross-reference what UM-GPT suggests with a JAGS forum or Stack Overflow posts.” Several described a personal rule against pasting generated code directly, instead retyping or adapting it themselves. Mathematical/scientific computing. Mathematical/scientific computing covered use cases in which the core task was implementing or solving a mathematical or domain-specific scientific problem: finding solutions to systems of equations, translating a set of governing equations into working code, writing simulations, and setting up numerical optimization. Some respondents described the task in terms of the mathematics itself, as in “find solution for a nonlinear equation system” and “solving equations,” while others were implementing a specific model from their domain, such as a driftdiffusion model for plasma discharge or a simulation in item response theory. A recurring pattern was starting from a known formalism, typically a set of equations drawn from a published article or textbook, and asking the tool to render it as code. One respondent working on an underdetermined nonlinear ordinary differential equation (ODE) noted that Claude “provided several feasible approaches to solve my problem” and “vastly simplified the algebra needed to arrive at a final usable expression in minutes — something which would have taken me a week at least.” In this category, then, the tool was sometimes used for the mathematics as much as for the code. Respondents in this category especially often combined strategies. Most accounts described more than one strategy (34/59), and only 7 relied on running the code alone. Running the code was still the most common single strategy (31/59), but respondents also used other strategies. Some compared the output against an independent standard: published results and textbook examples or “output products of previous versions of the pipeline”, a kind of computational benchmark. Others verified the underlying mathematics directly rather than trusting the implementation, by checking a derivation step by step (“I asked it to show me steps that I verified”), by confirming that a hand-computed quantity matched the code’s output, or, in one case, by cross-checking “a few hand-computed current balances” against the tool’s outputs “to confirm physical feasibility.” Some respondents chose strategies to avoid engaging in mathematical justification directly: one described implementing “a new function with heavy math that I didn’t fully understand, and didn’t need to fully understand,” and validated it entirely behaviorally, by creating a “test bench” and confirming that the function returned the results they expected. Several respondents also drew on domain knowledge or expert judgment as the final check, assessing whether “the results/scientific figures are scientifically sound” or checking the output with a supervisor, colleagues, or PI. Statistical analysis. Statistical analysis captured use cases in which the goal was to apply a statistical or machine learning method to find patterns in data: fitting regressions, running statistical tests, calculating descriptive statistics, and implementing machine learning pipelines. In many cases, respondents had clear analysis strategies in mind but needed to learn less familiar libraries or functions. For example, one respondent shared a direct prompt that they had written: “I have the data and script written in SPSS for the mediation models. Can you help me quickly import data into
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming 15 R and format as needed to do proposed analysis using lavaan.” Several described doing their first-ever analyses with pandas or a network analysis package. Other respondents knew the analysis that they wanted but not whether an implementation existed, and used the tool to survey available packages: “Are there R packages to implement techniques like King and Zeng’s methods?” Another “asked it what built in change-point functions existed in R, what their use-cases were, and what arguments each function used.” Others needed to customize the structure of statistical models in clearly specified ways, like converting confidence intervals from analytical to bootstrap, re-estimating a returns model at daily rather than monthly frequency, or matching model structure to the specifics of their dataset. Some responses showed scientists using AI tools as recommender systems for analysis strategies. One respondent prompted, “What kind of statistical analysis can I do in R that can help me find significant insights into the connection between [variables in dataset]?” Another reported, “I described the data patterns to the tool and asked what methods were available for detecting such patterns by clustering. . . It suggested a package for k-means clustering and thought that implementing with dynamic time-warping would be up to the task. . . I wanted to know how to identify the optimal number of clusters, and it suggested elbow plots.” Evaluation strategies were varied and frequently combined (28 of 55 accounts described more than one). Running the code was the most common strategy (24/55), though only 7 reported running code as their sole check (“If it works without bugs or not”; “try to run and check any error is coming or not”). One respondent made their multi-stage strategy explicit: “I tested the output in R to first see if it ran and then I looked through the output for indications of successfully completing the intended task (correct values, missing values, etc.).” When inspecting output (18/55), respondents described checking for specific failure modes and properties: variables “not transformed to the correct scale,” asymmetry in a matrix that should be symmetric, or implausibility against an internal sense of what a reasonable result looks like (“I pulled out numbers and saw if the classification looks ok”; “I check if the results look reasonable and match what I expect”). The most involved evaluations piloted generated code on data with known answers or re-derived the result independently. One respondent used a “much smaller dataset with known regression result” as a reference, and another “also ran one model without the function and compared the results to ensure they were the same.” Evaluating AI output. Across the corpus, the strategies that respondents reported were dominated by the lightweight checks noted above (Figure 1, right): running the generated code and inspecting its output, followed by reading the code, inspecting a visualization, and drawing on domain knowledge or intuition. Among these common strategies, respondents reported a wide range of diligence. Some were emphatic about scrutiny (“I never mindlessly copy-paste”), while others acknowledged that scrutiny was the first practice they dropped under time pressure: One respondent admitted that “the longer I’m working on a code/the more complicated it gets, the less I read it and the more I resort to just running the code and seeing what happens.” Another described reviewing code only after using it: “This usually comes after checking if the code works, however, so I guess I am frequently running untested code.” The remaining strategies were comparatively rare. Respondents consulted external references such as documentation or forums in 45 accounts (8.9%) and, in 38 (7.5%), compared output against an independent benchmark such as published results, a prior pipeline, or a known-answer dataset. More formal or collaborative checks were rarer still: some form of automated testing appeared in only 14 accounts (2.8%), asking the tool to explain its own output in 14 (2.8%), review by another person in 10 (2.0%), checking the mathematics by hand in 6 (1.2%), and cross-checking against a second AI tool in 2 (0.4%). When review was described, the language suggests an ad-hoc assembly of nearby reviewers rather than a
16
O’Brien et al.
dedicated code review process (“usually double check with professors”; “I shared the code with other people and asked them to test it as well”). Overall, validation was seldom externalized or formalized. In most accounts, the person who prompted for the code also ran it, inspected the result, and decided whether it was acceptable. Independent reviewers and automated test suites each appeared in only a small fraction of cases. These counts reflect the evaluation strategies that respondents considered worth reporting, not necessarily everything they did (Section 5). 4.2
Quantitative analysis
Who uses AI for what? We first asked whether reported use cases and evaluation strategies varied with respondents’ years of programming experience and research area. Experience. For each of the top five use-case codes, we compared log-transformed programming experience between respondents whose account did and did not receive that code (Welch’s t-tests, 𝑁 = 463). The largest difference was for debugging: Respondents describing debugging use cases were somewhat less experienced (𝑀 = 0.82 vs. 0.89 log units, roughly 5.6 vs. 6.8 years), but this did not reach significance (𝑡 (130.8) = 1.93, 𝑝 = .056, BH 𝑝 = .28). We also tested whether the top five evaluation strategies were related to programming experience, but saw little evidence of a relationship. In per-strategy comparisons (𝑁 = 460), respondents who reported a given strategy did not differ in programming experience from those who did not. The largest difference was for reading code (𝑡 (147.9) = 1.58, 𝑝 = .12, BH 𝑝 = .58). The number of distinct strategies a respondent reported (counted over all twelve evaluation codes) was likewise unrelated to experience (𝑟 (458) = −.03, 𝑝 = .49). Research area. Two use cases differed in frequency by research area in per-code chi-square tests (5 areas × code present/absent). Mathematical/scientific computing use cases were most common among engineers (21% of engineering accounts vs. 14% in the life sciences, 14% in the physical sciences, and 6–8% in other areas and the social sciences; 𝜒 2 (4) = 13.33, 𝑝 = .010, BH 𝑝 = .046). Statistical analysis use cases showed roughly the reverse pattern (17% of Other and 14% of Life sciences accounts vs. 4–6% in Engineering and Physical sciences; 𝜒 2 (4) = 11.86, 𝑝 = .018, BH 𝑝 = .046). In sum, neither reported use cases nor evaluation strategies varied substantially with programming experience. The differences that we could detect were associated with research area, and in expected directions: mathematical and scientific computing was more common among engineers, while statistical analysis use cases concentrated in the life and social sciences. How is experience related to confidence? Respondents rated their confidence in (a) completing the task independently (“solo”), (b) completing it with the GenAI tool (“GenAI”), and (c) evaluating whether the tool’s output was correct (“evaluation”). Figure 4 plots each rating against programming experience. For the confidence analyses we included every account with confidence ratings, even those that could not be assigned qualitative codes. Solo confidence rose with programming experience (𝑟 = .28, 𝑝 = 8 × 10−11 , 𝑁 = 513), as did evaluation confidence (𝑟 = .17, 𝑝 = 9 × 10−5 , 𝑁 = 513). GenAI confidence, by contrast, was unrelated to experience (𝑟 = −.01, 𝑝 = .85, 𝑁 = 512). This flat relationship could reflect that we asked participants to describe a recent use case. Many will have no use case to report in which they were very unconfident in the AI, because they would not have used it at all. This divergence is easiest to see in the within-respondent gap between solo and GenAI confidence (solo − GenAI). The gap widened with experience (𝑟 = .21, 𝑝 = 9 × 10−7 , 𝑁 = 512), and the fitted line crossed zero at roughly four years of programming experience (log10 (years + 1) = 0.68, ≈ 3.7 years). In other words, less experienced programmers tend
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming 17
Fig. 4. Confidence ratings (1–5) as a function of programming experience (log10 (years + 1)). Left: confidence completing the described task without generative AI (“coding solo”) and confidence in the GenAI tool. Right: confidence evaluating the tool’s output. Points are individual accounts (vertically jittered). Lines are linear fits with 95% confidence bands.
to use AI for tasks that they are not as confident they could do alone, while more experienced programmers tend to use it for tasks that they are confident they could do themselves. What use cases are respondents most confident about? We next asked whether the use case itself was related to any of the three confidence ratings, over and above programming experience (a natural confound, given the associations above). For each confidence rating, we first compared a linear model with experience alone to a model adding all five use-case indicators as a block (nested F-test). This omnibus test asks whether the use cases jointly explain additional variance. Adding use cases improved the model for evaluation confidence (𝐹 (5, 455) = 3.53, 𝑝 = .004, Δ𝑅 2 = .036) but not for solo confidence (𝐹 (5, 455) = 1.38, 𝑝 = .23) or GenAI confidence (𝐹 (5, 454) = 0.32, 𝑝 = .90). For evaluation confidence, we then fit per-code models adjusting for experience (each an ordinary linear model of confidence on experience plus one use-case indicator). The use-case coefficient is the experience-adjusted mean difference between accounts with and without that code. Evaluation confidence was lower for debugging use cases (𝑀 = 3.86 vs. 4.19; adjusted difference 𝑏 = −0.30, 𝑡 (459) = −3.12, 𝑝 = .002, BH 𝑝 = .010) and higher for visualization use cases (𝑀 = 4.32 vs. 4.07; 𝑏 = 0.24, 𝑡 (459) = 2.59, 𝑝 = .010, BH 𝑝 = .025; Figure 5). In sum, the use case that a respondent described was related only to their evaluation confidence: Respondents were more confident evaluating AI output for visualization and less confident for debugging. Neither confidence in the tool nor confidence in completing the task alone varied with the use case reported. How does evaluation strategy relate to confidence? Finally, we asked whether evaluation confidence tracks respondents’ reported evaluation strategies. Adding the five most common evaluation codes as a block did not improve on an experience-only model of evaluation confidence (nested F-test: 𝐹 (5, 452) = 1.55, 𝑝 = .17). In univariate t-tests, respondents who reported inspecting visualizations (𝑀 = 4.30 vs. 4.09; 𝑡 (120.1) = 2.19, 𝑝 = .031) or reading code (𝑀 = 4.27 vs. 4.09; 𝑡 (175.5) = 2.06, 𝑝 = .041) reported higher evaluation confidence, but neither effect survived correction for multiple
18
O’Brien et al.
Fig. 5. Confidence by use case. Points are experience-adjusted mean differences in each confidence rating (1–5 scale) between responses with and without each of the five most common use-case codes, from linear models of confidence on log-experience plus the use-case indicator (one model per code and rating). Error bars: 95% CIs. Red marks differences with BH-adjusted 𝑝 < 0.05 within each confidence rating. 𝑛 = 462 responses with complete confidence and experience data.
comparisons (both BH 𝑝 = .10). The number of distinct strategies that a respondent reported was also unrelated to evaluation confidence (𝑟 (457) = .05, 𝑝 = .29). In a model predicting strategy count from evaluation confidence with experience as a covariate, 𝑏 = 0.05 strategies per confidence point (𝑝 = .23). Among the measures we collected, the strongest correlates of evaluation confidence were confidence in completing the task alone (𝑟 = .30, 𝑝 = 4 × 10−12 ) and confidence in the GenAI tool to complete the task (𝑟 = .33, 𝑝 = 1 × 10−14 ). These two ratings were uncorrelated with each other (𝑟 = .02, 𝑝 = .60), so this pattern is not consistent with a single, underlying general confidence factor. Both relationships are much stronger than any that we observed with reported evaluation strategies. 5
Discussion
Summary of findings We coded accounts of generative AI use in research programming for two kinds of information: the task that motivated the request, and the strategies used to evaluate what came back. Use was concentrated in five tasks: data handling, visualization, debugging, mathematical/scientific computing, and statistical analysis, which together appeared in about three-quarters of accounts where a use-case code could be applied. Evaluation across all of them was dominated by informal, individual checks. Over half of respondents reported running the generated code as an evaluation strategy. Quantitatively, neither use cases nor evaluation strategies varied much with programming experience. The few differences we could detect were by field, in expected directions. Confidence, by contrast, did vary with experience. Less experienced programmers reported more confidence in the AI than in themselves, and more experienced programmers the reverse, with the crossover at roughly 4 years of experience. One interpretation is that novices may use these tools to attempt work that they could not do alone, while experienced programmers use them to save time on work that they could do themselves. Finally, we did not detect a relationship between respondents’ confidence in their evaluations and
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming 19 either the evaluation strategies that they reported or the number of strategies they reported. In our data, evaluation confidence was associated with programming experience, use case (higher for visualization, lower for debugging), reported confidence in doing the task alone, and reported confidence in the tool to complete the task.
Interpretation These results describe how scientists evaluate AI-generated code, but they do not license a strong normative claim that people are under- or over-evaluating, because the appropriate level of evaluation is not fixed across tasks or contexts. Instead, scientists appear to make situated judgments about what constitutes sufficient evidence of correctness based on the task, the observability of its output, and their own knowledge of the underlying data, code, and scientific domain. For a cosmetic change to a plot, visual inspection may be exactly the appropriate check. The respondents who described figures as “literally visually validated” may be correctly, not carelessly, calibrated. For data handling, correctness is often not visible in the output itself, and “it ran” establishes little. Debugging is the hardest to judge from the outside: Are respondents making small syntactic repairs that never touch the scientific assumptions in the code, or doing things until an error message goes away, possibly altering parts of the code they should not in the process? Our accounts cannot distinguish these readings, and the two have very different implications. The finding that evaluation confidence was not significantly related to reported evaluation strategy has two candidate explanations. One is methodological: Free-text accounts are a coarse instrument, and respondents may under-report checks that they actually did. For example, someone who writes “I ran it” may also have compared the output against expectations without saying so. We may have limited statistical power to detect a real relationship. The other is that confidence in evaluating AI output is not primarily a product of evaluation strategies for many respondents. That evaluation confidence was associated with confidence in the tool is consistent with trust partially substituting for verification, consistent with the finding of Lee et al. [28] that knowledge workers with higher confidence in generative AI engaged in less critical evaluation. Our data cannot adjudicate between these readings, and both may be partly true. The association between confidence and trust does not imply that highly confident respondents evaluated poorly, nor can our data establish the quality of any individual evaluation. It does suggest that subjective confidence should not be treated as a proxy for evaluation strategies. In nearly every account, the respondent described a closed loop of interaction between themself and their AI tool. Review by another person appeared in 2% of accounts and automated testing in 3% (though this does not mean only 3% of scientists use testing, only that it was not mentioned explicitly in most accounts). This closed loop is consistent with earlier surveys reporting low use of peer code review and testing among scientists [7, 15]. We would not expect this infrastructure to appear quickly now that generative AI tools are available. Additionally, many of the use cases described could be challenging to fit into a typical test suite, or may not have sufficient epistemological weight to justify the overhead of testing in a formal sense (as tests accumulate technical debt, too). For example, some kinds of visuals and descriptive tables are rendered quickly to support exploration and hypothesis formation [26, 51]. Errors here still matter (for example, filtering data in an unexpected way before summarizing or visualizing it), but repeated interaction offers many opportunities to notice such a problem. Qualitatively, we observe wide variation, ranging from quick checks that no errors arise, to comparing results against internal expectations of reasonable values, to checking that other code implementations produce the same result, to writing test cases. This variation places the weight on individual scientists’ expertise for judging both (a) what level of scrutiny generated code deserves in the research context and (b) whether that threshold is met in a given use case.
20
O’Brien et al. Tasks also differ in whether independent evidence is available. Scientists may be able to compare generated results
against a reference such as a known quantity or a previous implementation, while in other cases no convenient reference exists. A more useful distinction than formal versus informal evaluation is whether the evidence stays inside the immediate human–AI interaction or introduces an independent point of comparison. Running generated code, visually inspecting its output, and asking the same AI to explain its response are all useful, but they preserve the same interaction loop. Checks such as a reference implementation, a hand calculation, or review by another person introduce evidence that is at least partially independent of the generation process. This distinction may offer a better basis for supporting evaluation than encouraging greater use of conventional software testing. Implications If the burden of devising verification activities during a session, and then executing them, falls mostly to individual scientists, we see several ways to support them. Tools might default to offering to produce lightweight test datasets for interactively running through any generated code. For example, to verify an AI-generated data handling step, tools might offer a small, easy-to-inspect sample dataset to run through the code so that the output can be observed directly (a strategy several respondents already reported using, just manually). With agentic workflows growing in popularity, tools could also report to the scientist that an agent is running this test case and then interactively show the result for further iteration. Beyond data handling, AI programming tools could generate verification artifacts alongside code, tailored to the task. For visualization, this means diagnostic views that expose transformations applied before plotting. For debugging, a regression case shows that a fix addresses the original failure without changing unrelated behavior. For mathematical or scientific computing, the artifact is a comparison against a known quantity, a limiting case, or an alternative implementation. The goal is to reduce the burden on scientists of devising verification activities themselves, without prescribing a single notion of correctness. Training scientists in the future could focus on modeling the informal evaluation strategies used by more senior scientists: After running generated code, what specific qualities of the output does a more experienced scientist inspect? What diagnostic plots do they make to understand the behavior of code that they did not write? Our interpretation of this data points toward scientists acting as calibrated instruments, that is, making small but important internal comparisons between what they expect of a result, given their knowledge of the data and the domain, and what actually happens. Such training may involve calibrating scientists’ internal “priors” for how likely a given computational result is, and teaching them what steps they would take to re-calibrate when moving to an unfamiliar problem. Interfaces could also make the basis for acceptance more explicit. Rather than relying on a user’s sense that an output looks correct, an AI-assisted programming environment might show what has been checked, such as whether the code ran, whether outputs were compared against known values, or whether an independent reference was consulted. Making those checks visible could help separate confidence from evidence without requiring every research programming task to adopt heavyweight software engineering practices. Limitations Our data consists of self-reported accounts of single recalled episodes, and by design we conducted a content analysis bounded by what respondents wrote. As noted in the Interpretation section, respondents may have under-reported checks they performed, so our counts of evaluation strategies are lower bounds on what was done. We cannot know what people did, only what they remembered and chose to report. The survey is a convenience sample, concentrated in
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming 21 U.S. higher education and likely to overrepresent scientists interested in programming enough to answer a survey about it. Confidence items were tied to a specific, recalled episode, and the flat relationship between GenAI confidence and experience may partly reflect selection, because people rarely have use cases to report for tools that they do not trust. Three further limitations concern how the accounts were produced and coded. First, the survey asked scientists to describe their own diligence, so social desirability may have led some respondents to over-report checks, and others, writing briefly, to under-report them. Second, each account describes one episode, so person-level claims (for example, that a respondent who described a visualization task did not debug) are not licensed by the data, and our quantitative comparisons treat one recalled episode as representative of the respondent. Third, we designed the survey and coded the responses ourselves. We mitigated the risk of confirming our own expectations through two rounds of independent coding by all three authors, reliability checks (Krippendorff’s 𝛼 = 0.76 pairwise), and consensus resolution of flagged accounts, but most of the corpus was coded by a single author after the codebooks were finalized, and the coders were not blind to the study’s aims. Because the instrument was in English and the sample is U.S.-based, the practices described here may not transfer to other research settings. Future research These findings describe a moment in computing that is already past. Our data was collected in 2025, and most respondents were not using agentic tools that can execute code, repair errors, and write their own test suites. As tools change, we expect many of the use cases that we identified to persist, because they arise at points of friction between scientists and their tools, such as opaque error messages, switching between libraries with various degrees of familiarity, or requiring highly custom code for the particulars of data. But evaluation may look quite different in an agentic workflow: Agents can design and execute certain kinds of validation work themselves, while also producing far more code outside the scientist’s direct oversight. If agents attempt validation work on their own, scientists may shift from designing their own verification strategies to judging whether the agent’s strategies are appropriate. Where test suites are appropriate, agents might also help scientists write and automate them, although whether this convention takes hold is an open question. Studying validation work in the coming years will likely require observing scientists as they work with their tools, because their checks may not be legible in code or chat logs. For these reasons, we expect the scientific correctness of AI-assisted research code to continue to rest on the scientist’s skill in building a well-calibrated mental model of the data they study, the computations they apply, and the results they should expect. Data Availability The survey responses cannot be shared, because the terms of the informed consent under which they were collected do not permit redistribution, and the written accounts contain free text that could identify respondents. Scripts for conducting the quantitative analyses and generating the figures will be provided upon publication. The full codebooks, with definitions and example excerpts for every code, are given in Tables 2 and 3. Acknowledgments The lead author (GO) is supported by a grant from the Alfred P. Sloan Foundation. References [1] Gagan Bansal, Tongshuang Wu, Joyce Zhou, Raymond Fok, Besmira Nushi, Ece Kamar, Marco Tulio Ribeiro, and Daniel S. Weld. 2021. Does the Whole Exceed its Parts? The Effect of AI Explanations on Complementary Team Performance. In Proceedings of the 2021 CHI Conference on Human
22
O’Brien et al.
Factors in Computing Systems (CHI ’21). doi:10.1145/3411764.3445717 [2] Shraddha Barke, Michael B. James, and Nadia Polikarpova. 2023. Grounded Copilot: How Programmers Interact with Code-Generating Models. Proceedings of the ACM on Programming Languages 7, OOPSLA1, Article 78 (2023), 27 pages. doi:10.1145/3586030 [3] Robert Baxter, Neil Chue Hong, Dirk Gorissen, James Hetherington, and Ilian Todorov. 2012. The Research Software Engineer. In Digital Research 2012. [4] Joel Becker, Nate Rush, Beth Barnes, and David Rein. 2025. Measuring the Impact of Early-2025 AI on Experienced Open-Source Developer Productivity. arXiv preprint arXiv:2507.09089 (2025). [5] Stephanie A. Besser, Eric A. Jensen, and Daniel S. Katz. 2026. How generative AI is shaping research software development and maintenance at a research-intensive university. Open Research Europe 6 (2026), 56. doi:10.12688/openreseurope.22009.1 [6] Zana Buçinca, Maja Barbara Malaya, and Krzysztof Z. Gajos. 2021. To Trust or to Think: Cognitive Forcing Functions Can Reduce Overreliance on AI in AI-assisted Decision-making. Proceedings of the ACM on Human-Computer Interaction 5, CSCW1, Article 188 (2021). doi:10.1145/3449287 [7] Jeffrey C. Carver, Nic Weber, Karthik Ram, Sandra Gesing, and Daniel S. Katz. 2022. A survey of the state of the practice for research software in the United States. PeerJ Computer Science 8 (2022), e963. doi:10.7717/peerj-cs.963 [8] Valerie Chen, Ameet Talwalkar, Robert Brennan, and Graham Neubig. 2026. Code with Me or for Me? How Increasing AI Automation Transforms Developer Workflows. In Proceedings of the 2026 CHI Conference on Human Factors in Computing Systems (CHI ’26). doi:10.1145/3772318.3790850 [9] Ruijia Cheng, Ruotong Wang, Thomas Zimmermann, and Denae Ford. 2024. “It would work for me too”: How Online Communities Shape Software Developers’ Trust in AI-Powered Code Generation Tools. ACM Transactions on Interactive Intelligent Systems (2024). doi:10.1145/3651990 [10] Marina Chugunova, Dietmar Harhoff, Katharina Hölzle, Verena Kaschub, Sonal Malagimani, Ulrike Morgalla, and Robert Rose. 2026. Who Uses AI in Research, and for What? Large-scale Survey Evidence from Germany. Research Policy 55, 2 (2026), 105381. doi:10.1016/j.respol.2025.105381 [11] Jeremy Cohen, Daniel S. Katz, Michelle Barker, Neil P. Chue Hong, Robert Haines, and Caroline Jay. 2021. The four pillars of research software engineering. IEEE Software 38, 1 (2021), 97–105. doi:10.1109/MS.2020.2973362 [12] Nasir U Eisty, Upulee Kanewala, and Jeffrey C Carver. 2025. Testing research software: an in-depth survey of practices, methods, and tools. Empirical Software Engineering 30, 3 (2025), 81. [13] Kate Goddard, Abdul Roudsari, and Jeremy C. Wyatt. 2012. Automation bias: a systematic review of frequency, effect mediators, and mitigators. Journal of the American Medical Informatics Association 19, 1 (2012), 121–127. doi:10.1136/amiajnl-2011-000089 [14] Ken Gu, Ruoxi Shang, Tim Althoff, Chenglong Wang, and Steven M. Drucker. 2024. How Do Analysts Understand and Verify AI-Assisted Data Analyses?. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). doi:10.1145/3613904.3642497 [15] Jo Erskine Hannay, Carolyn MacLeod, Janice Singer, Hans Petter Langtangen, Dietmar Pfahl, and Greg Wilson. 2009. How do scientists develop and use scientific software? Proceedings of the 2009 ICSE Workshop on Software Engineering for Computational Science and Engineering, SECSE 2009 (2009). doi:10.1109/SECSE.2009.5069155 [16] Andrew Head, Fred Hohman, Titus Barik, Steven M. Drucker, and Robert DeLine. 2019. Managing messes in computational notebooks. In Proceedings of the 2019 CHI Conference on Human Factors in Computing Systems (CHI ’19). doi:10.1145/3290605.3300500 [17] Simon Hettrick. 2014. It’s impossible to conduct research without software, say 7 out of 10 UK researchers. Software Sustainability Institute blog. https://www.software.ac.uk/blog/its-impossible-conduct-research-without-software-say-7-out-10-uk-researchers Accessed 2026-09-16. [18] James Howison and James D. Herbsleb. 2011. Scientific software production: Incentives and collaboration. In Proceedings of the ACM Conference on Computer Supported Cooperative Work (CSCW). doi:10.1145/1958824.1958904 [19] INTERSECT. 2026. INTERSECT: Training for Research Software Engineering. https://intersect-training.org/. Online resource. Accessed 2026-05-19. [20] Arne Johanson and Wilhelm Hasselbring. 2018. Software engineering for computational science: Past, present, future. Computing in Science & Engineering 20, 2 (2018), 90–109. [21] Journal of Clinical Oncology. 2016. Retraction: Inferring the effects of cancer treatment: Divergent results from Early Breast Cancer Trialists’ Collaborative Group meta-analyses of randomized trials and observational data from SEER registries. Journal of Clinical Oncology 34, 27 (2016), 3358–3359. doi:10.1200/JCO.2016.69.0875 [22] Upulee Kanewala and James M. Bieman. 2014. Testing scientific software: A systematic literature review. Information and Software Technology (2014). doi:10.1016/J.INFSOF.2014.05.006 [23] Amelia Karraker and Kenzie Latham. 2015. Authors’ explanation of the retraction. Journal of Health and Social Behavior (2015). doi:10.1177/ 0022146515595817 [24] Diane Kelly. 2015. Scientific software development viewed as knowledge acquisition: Towards understanding the development of risk-averse scientific software. Journal of Systems and Software (2015). doi:10.1016/j.jss.2015.07.027 [25] Diane F. Kelly. 2007. A software chasm: Software engineering and scientific computing. IEEE Software 24, 6 (2007), 118–120. doi:10.1109/MS.2007.155 [26] Mary Beth Kery and Brad A. Myers. 2017. Exploring exploratory programming. In Proceedings of IEEE Symposium on Visual Languages and Human-Centric Computing, VL/HCC. IEEE. doi:10.1109/VLHCC.2017.8103446 [27] Sam Lau, Sruti Srinivasa Ragavan, Ken Milne, Titus Barik, and Advait Sarkar. 2021. TweakIt: Supporting End-User Programmers Who Transmogrify Code. In Proceedings of the 2021 CHI Conference on Human Factors in Computing Systems (CHI ’21). doi:10.1145/3411764.3445265 [28] Hao-Ping (Hank) Lee, Advait Sarkar, Lev Tankelevitch, Ian Drosos, Sean Rintel, Richard Banks, and Nicholas Wilson. 2025. The Impact of Generative AI on Critical Thinking: Self-Reported Reductions in Cognitive Effort and Confidence Effects From a Survey of Knowledge Workers. In Proceedings of the 2025 CHI Conference on Human Factors in Computing Systems (CHI ’25). 1–22. doi:10.1145/3706598.3713778
How Researchers Use and Verify AI Coding Assistants: Tasks and Validation Practices in Scientific Programming 23 [29] John D. Lee and Katrina A. See. 2004. Trust in Automation: Designing for Appropriate Reliance. Human Factors 46, 1 (2004), 50–80. doi:10.1518/hfes. 46.1.50_30392 [30] Jenny T. Liang, Chenyang Yang, and Brad A. Myers. 2024. A Large-Scale Survey on the Usability of AI Programming Assistants: Successes and Challenges. In Proceedings - International Conference on Software Engineering (ICSE 2024). doi:10.1145/3597503.3608128 [31] Michael Xieyang Liu, Advait Sarkar, Carina Negreanu, Benjamin Zorn, Jack Williams, Neil Toronto, and Andrew D. Gordon. 2023. “What It Wants Me To Say”: Bridging the Abstraction Gap Between End-User Programmers and Code-Generating Large Language Models. In Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems (CHI ’23). doi:10.1145/3544548.3580817 [32] Erin C McKiernan, Philip E Bourne, C Titus Brown, Stuart Buck, Amye Kenall, Jennifer Lin, Damon McDougall, Brian A Nosek, Karthik Ram, Courtney K Soderberg, et al. 2016. How open science helps researchers succeed. elife 5 (2016), e16800. [33] Hussein Mozannar, Gagan Bansal, Adam Fourney, and Eric Horvitz. 2024. Reading Between the Lines: Modeling User Behavior and Costs in AIAssisted Programming. In Proceedings of the 2024 CHI Conference on Human Factors in Computing Systems (CHI ’24). 1–16. doi:10.1145/3613904.3641936 [34] Sydney Nguyen, Hannah McLean Babe, Yangtian Zi, Arjun Guha, Carolyn Jane Anderson, and Molly Q. Feldman. 2024. How Beginning Programmers and Code LLMs (Mis)read Each Other. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). doi:10.1145/3613904. 3642706 [35] William L Oberkampf and Christopher J Roy. 2010. Verification and validation in scientific computing. Cambridge university press. [36] Gabrielle O’Brien. 2025. How Scientists Use Large Language Models to Program. In CHI Conference on Human Factors in Computing Systems. doi:10.1145/3706598.3713668 [37] Gabrielle O’Brien, Alexis Parker, Nasir U Eisty, and Jeffrey Carver. 2025. A survey of generative AI adoption and perceived productivity among scientists who program. arXiv preprint arXiv:2512.19644 (2025). doi:10.48550/arXiv.2512.19644 [38] Drew Paine and Charlotte P. Lee. 2017. “Who has plots?”: Contextualizing scientific software, practice, and visualizations. Proceedings of the ACM on Human-Computer Interaction 1, CSCW (2017), 85:1–85:21. doi:10.1145/3134720 [39] Raymond R. Panko. 2007. Two Experiments in Reducing Overconfidence in Spreadsheet Development. Journal of Organizational and End User Computing (2007). doi:10.4018/joeuc.2007010101 [40] Sida Peng, Eirini Kalliamvakou, Peter Cihon, and Mert Demirer. 2023. The Impact of AI on Developer Productivity: Evidence from GitHub Copilot. arXiv preprint arXiv:2302.06590 (2023). doi:10.48550/arXiv.2302.06590 [41] James Prather, Brent N. Reeves, Paul Denny, Brett A. Becker, Juho Leinonen, Andrew Luxton-Reilly, Garrett Powell, James Finnie-Ansley, and Eddie Antonio Santos. 2023. “It’s Weird That it Knows What I Want”: Usability and Interactions with Copilot for Novice Programmers. ACM Transactions on Computer-Human Interaction 31, 1, Article 4 (2023), 31 pages. doi:10.1145/3617367 [42] Sruti Srinivasa Ragavan, Zhitao Hou, Yun Wang, Andrew D. Gordon, Haidong Zhang, and Dongmei Zhang. 2022. GridBook: Natural Language Formulas for the Spreadsheet Grid. In International Conference on Intelligent User Interfaces, Proceedings IUI. doi:10.1145/3490099.3511161 [43] Rebecca Sanders and Diane Kelly. 2008. Dealing with Risk in Scientific Software Development. IEEE Software (2008). doi:10.1109/MS.2008.84 [44] Advait Sarkar, Andrew D. Gordon, Carina Negreanu, Christian Poelitz, Sruti Srinivasa Ragavan, and Ben Zorn. 2022. What is it like to program with artificial intelligence? Proceedings of the 33rd Annual Workshop of the Psychology of Programming Interest Group (PPIG) (2022). [45] Judith Segal. 2007. Some problems of professional end user developers. In IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC 2007). IEEE, 111–118. doi:10.1109/VLHCC.2007.17 [46] Judith Segal. 2008. Models of scientific software development. In Proceedings of the First International Workshop on Software Engineering for Computational Science and Engineering (SECSE 2008). Leipzig, Germany. [47] Claudio Spiess, David Gros, Kunal Suresh Pai, Michael Pradel, Md Rafiqul Islam Rabin, Amin Alipour, Susmit Jha, Prem Devanbu, and Toufique Ahmed. 2025. Calibration and Correctness of Language Models for Code. In Proceedings of the IEEE/ACM 47th International Conference on Software Engineering (ICSE ’25). 540–552. doi:10.1109/ICSE55347.2025.00040 [48] Mark Steyvers, Heliodoro Tejeda, Aakriti Kumar, Catarina Belem, Sheer Karny, Xinyue Hu, Lukas W. Mayer, and Padhraic Smyth. 2025. What large language models know and what people think they know. Nature Machine Intelligence 7, 2 (2025). doi:10.1038/s42256-024-00976-7 [49] Tim Storer. 2017. Bridging the Chasm: A Survey of Software Engineering Practice in Scientific Programming. Comput. Surveys 50, 4, Article 47 (2017), 32 pages. doi:10.1145/3084225 [50] Margaret-Anne Storey. 2026. From Technical Debt to Cognitive and Intent Debt: Rethinking Software Health in the Age of AI. arXiv preprint arXiv:2603.22106 (2026). doi:10.48550/arXiv.2603.22106 [51] William Sutherland-Keller. 2025. Research Software Systems: Exploration and Infrastructure in Observational Cosmology. Ph. D. Dissertation. ProQuest Dissertations and Theses. [52] Ningzhi Tang, Meng Chen, Zheng Ning, Aakash Bansal, Yu Huang, Collin McMillan, and Toby Jia-Jun Li. 2024. Developer Behaviors in Validating and Repairing LLM-Generated Code Using IDE and Eye Tracking. In 2024 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC). IEEE, 40–46. doi:10.1109/VL/HCC60511.2024.00015 [53] Lev Tankelevitch, Viktor Kewenig, Auste Simkute, Ava Elizabeth Scott, Advait Sarkar, Abigail Sellen, and Sean Rintel. 2024. The Metacognitive Demands and Opportunities of Generative AI. In Proceedings of the CHI Conference on Human Factors in Computing Systems (CHI ’24). doi:10.1145/ 3613904.3642902 [54] The Carpentries. 2026. The Carpentries. https://carpentries.org/. Online resource. Accessed 2026-05-19.
24
O’Brien et al.
[55] Priyan Vaithilingam, Tianyi Zhang, and Elena L. Glassman. 2022. Expectation vs. Experience: Evaluating the Usability of Code Generation Tools Powered by Large Language Models. Conference on Human Factors in Computing Systems - Proceedings (2022). doi:10.1145/3491101.3519665 [56] Steve Van Tuyl. 2025. State of AI in the RSE Workplace - Initial Survey Results and SciPy Convening. Technical Report. Zenodo. doi:10.5281/ZENODO. 16953575 [57] Helena Vasconcelos, Gagan Bansal, Adam Fourney, Q. Vera Liao, and Jennifer Wortman Vaughan. 2024. Generation Probabilities Are Not Enough: Uncertainty Highlighting in AI Code Completions. ACM Transactions on Computer-Human Interaction (2024). doi:10.1145/3702320 [58] Helena Vasconcelos, Matthew Jörke, Madeleine Grunde-McLaughlin, Tobias Gerstenberg, Michael S. Bernstein, and Ranjay Krishna. 2023. Explanations Can Reduce Overreliance on AI Systems During Decision-Making. Proceedings of the ACM on Human-Computer Interaction 7, CSCW1, Article 129 (2023). doi:10.1145/3579605 [59] Ruotong Wang, Ruijia Cheng, Denae Ford, and Thomas Zimmermann. 2024. Investigating and Designing for Trust in AI-powered Code Generation Tools. In Proceedings of the 2024 ACM Conference on Fairness, Accountability, and Transparency (FAccT ’24). doi:10.1145/3630106.3658984 [60] Yangtian Zi, Luisa Li, Arjun Guha, Carolyn Anderson, and Molly Q Feldman. 2025. “I Would Have Written My Code Differently’: Beginners Struggle to Understand LLM-Generated Code. In Proceedings of the 33rd ACM International Conference on the Foundations of Software Engineering (FSE Companion ’25). ACM, 1479–1488. doi:10.1145/3696630.3731663