ConceptioArchivearXiv CS
arXiv CSopen access

An Eye for Trust: An Exploration of Developers' Trust Perceptions Through Urgency and Reputation

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

arXiv:2604.08713v1 [cs.SE] 9 Apr 2026

An Eye for Trust: An Exploration of Developers’ Trust Perceptions Through Urgency and Reputation Sara Yabesi

Mahta Amini

Polytechnique Montréal Montréal, Canada [email protected]

Polytechnique Montréal Montréal, Canada [email protected]

Jelena Ristic

Zohreh Sharafi

McGill University Montréal, Canada [email protected]

Polytechnique Montréal Montréal, Canada [email protected]

ABSTRACT

1

Code reuse is a widespread practice across software development projects, suggesting an inherent trust in the reused code. Yet, there is a lack of a fundamental understanding of developers’ trust and how various factors mold their trust–based cognitive processes. Drawing from the psychology of compliance and trust, we present the results of the first controlled experiment (n=37) which uses eye tracking to explore how urgency (represented by code priority level) and reputation (represented by the experience level of the code’s author) influence developers’ perceptions of code trustworthiness. Our research revealed that the priority assigned to a code patch significantly influenced developers’ code review behavior, impacting their evaluation time, cognitive load, and perceived quality. However, the decision to incorporate and implement the code was not affected . Eye tracking data revealed that there were variations in overall visual code scanning and the distribution of attention across identical code patches labeled as written by senior vs. junior developers. Yet, there were no significant performance differences. Moreover, our participants nominate code functionality, quality, and comprehensibility as primary factors in code evaluation. Despite noticeable changes in code review behavior, our participants surprisingly overlooked the substantial influence of urgency and reputation on their decisions to review and reuse code changes. This study takes the next step toward a better understanding of trust in software engineering and may inform future research about code review platforms and guidelines, code reuse, and automated code generation.

In The Mythical Man Month (1974), Fred Brooks suggests that reusable code costs three times as much to develop as single–use code. However, code reuse is prevalent across various software projects [1, 19, 34]. The profound impact of code reuse becomes undeniably apparent when analyzing its scale in relation to widely– recognized open–source software (OSS) projects, such as log4j and jUnit. The code reuse of such software saved roughly 316 thousand person-years and represents tens of billions of dollars in development expenditures [5]. Reuse implies that the artifact is trustworthy, as the developer primarily aims to use it in a system for which it was not initially intended. Still, there is a limited understanding of why developers have trust in software or artifacts they have not developed. This paper focuses on code reuse from a human trust perspective. During a cognitive task analysis, the trustworthiness of an artifact is influenced by various interrelated factors [3, 29], such as urgency and trust cues [37]. Urgency can alter how individuals process information and make decisions, amplifying their reliance on trust cues such as the source reputation. Trust cues, in turn, can significantly affect how individuals respond to urgent situations [37]. Leveraging principles from the psychology of trust, this paper performs an unprecedented examination to study how urgency and reputation impact developers’ trust perception. In the field of psychology, various metrics have been used to deduce and quantify trust. Past studies concentrated on self-reporting methods, e.g., think-aloud protocols, surveys, and interviews to evaluate trust [29]. However, these methods are prone to the Hawthorne effect, where the presence of an observer can influence the behavior of participants [11, 12] and may lack reliability [14, 23]. In contrast, a handful of recent research points to a link between the level of trust among users and how they visually scan information. Using eye–tracking technology, non–subjectively, and non–intrusively, they shed light on the mental processes and perception of trust. Through capturing the dynamic patterns of visual focus, eye tracking unveils indicators regarding the users’ cognitive processes [24, 48]. It is also useful in gauging mental workload during task execution [40, 47, 48], and in discerning the strategies and timings employed by participants while selecting and processing information [17, 24].

CCS CONCEPTS • Software and its engineering → Software creation and management;

KEYWORDS Trust, Human factors, Eye tracking, Code reuse, Urgency, Reputation ACM Reference format: Sara Yabesi, Mahta Amini, Jelena Ristic, and Zohreh Sharafi. 2018. An Eye for Trust: An Exploration of Developers’ Trust Perceptions Through Urgency and Reputation. In Proceedings of ACM Conference, Washington, DC, USA, July 2017 (Conference’17), 12 pages. https://doi.org/XXXXXXX.XXXXXXX

INTRODUCTION

Conference’17, July 2017, Washington, DC, USA A handful of previous studies have shed light on the relationship of trust in computer code. Alarcon et al. [2] studied the impact of code authors’ performance and reputation on developers’ trust and their decisions to utilize the provided code. Several studies have explored the psychological underpinnings of trust in the context of automated program repair methods (cf. [1, 3, 8, 14, 36, 38, 43]). Our work is closest in spirit to work by Bertram et al. [8] that evaluated the perceived trustworthiness and potential biases developers might hold towards automated tools during code review. However, our focus is different. To the best of our knowledge, this paper is the first experimentally– controlled study via eye tracking to explore how participants’ subjective trust and potentially biased actions are influenced by urgency and reputation. Regarding urgency, we alter the code while controlling for quality, labeling its priority as low or high. Also, we instantiate reputation by labeling the author’s experience as senior vs. junior. Our 37 participants worked on six code review tasks evaluating Java code patches. We measure developers’ tendency to trust and reuse the codes and performance coupled with visual attention trends and cognitive load [46, 48] to provide insights into developers’ cognitive processes when performing tasks. We observe significant differences in the patch evaluation by all participants, depending on the perceived priority of the patch: • The patch priority had a statistically significant effect on task time (𝑝 = 0.04). Also, patches labeled as high–priority demanded significantly more visual effort (𝑝 = 0.03) — indicative of cognitive load — during evaluation. • Different attention distributions in eye movements was observed when participants reviewed patches classified as low vs. high priority (𝑝 = 0.001). • Participants did not self–report any influence of patch priority on their decision to accept, deploy, and reuse the code. Yet, they rated high–priority patches as higher quality, even when we controlled for quality. Also, we notice some disparities in the way all participants assess code patches, influenced by the apparent experience of the author: • Our findings revealed no significant differences in participants’ performance and outcomes based on the author’s experience level. • Eye–tracking data indicated varying attention allocation patterns and viewing strategies while reviewing patches authored by novice vs. senior authors. • Our post–survey data did not indicate any effect of the author’s experience on their assessments to trust and reuse. Behaviorally, the thematic analysis of participants’ self–report data shows that code functionality, quality, and comprehensibility are the three most important factors for code evaluation. Despite evident changes in their code review behavior in terms of associated visual and cognitive processes, our participants surprisingly failed to recognize the significant impact of urgency and reputation on their decisions to review, incorporate, and reuse the code. This observed lack of awareness or understanding regarding the vital factors influencing their decision–making process during code reviews underscores the crucial significance of this research. Although code reuse is a common and profitable practice in the tech industry, an over–reliance or miscalibration of trust can lead

Yabesi, et al. to overuse and ignoring the constraints/potential dangers [2, 3, 8]. Consider OpenSSL, a common cryptography library. In 2015, the ‘Heartbleed’ vulnerability let hackers access protected data on 24– 55% of popular websites, causing substantial reputational and financial harm [10]. This novel paper expands the existing body of literature on the influence of trust on developers’ performance and behavior. The results hold potential implications for developing training programs, decision–making guidelines, collaborative review practices, and code review tools that provide feedback and raise awareness about the important cues towards creating a more objective and informed code review culture.

2

RELATED WORK

Related work includes studies of trust in computing [2, 3, 8, 25, 28, 38, 43] and human studies that utilized eye tracking to examine code review tasks [7, 9, 13, 49, 54]. Concerning eye tracking and code review, our work takes the next step towards a better understanding of developers’ code– reviewing behaviors and the various elements that shape their judgment and decisions. In terms of trust in computing, this paper is the first empirical study that contributes a fresh perspective on how developers perceive the trustworthiness of controlled quality patches based on urgency and reputation (i.e., their apparent priority level and the author’s experience). Moreover, most of the previous work used static stimuli, comprising small code snippets, while we covered indicative scenarios within an IDE (admitting editing, scrolling, and code and unit–tests execution) in which participants worked at their own pace on a large code base.

2.1

Trust in Computing

There is a rich body of work on how developers perceive and understand source codes [42]. Yet, only a handful of studies investigated the trustworthiness of code from developers’ perspective and the psychological processes behind the perception of trust [2, 3, 8, 38]. Several experimental findings highlight the utility of the heuristic–systematic processing model in the context of the trustworthiness of computer code. These results show the role of readability and the source’s reputation in the accurate evaluation and reliability of the code [2, 3]. The focus of the research on trust in software engineering is mainly on the effect of automation on trust. Lee and See [28] argue that the lack of trust in automated methods leads to their dismissal by developers. Alarcon et al. [4, 43] performed user studies and analyzed developers’ self–reported rate of the trustworthiness of patches. They reported a determining impact of the patch’s source on developers’ behavior and indicated a higher degree of trust toward human–written patches. Fry et al. [14] studies the understandability and maintainability of machine-generated patches. Their findings suggest a potential disconnect between what human programmers perceive as crucial for program maintainability and the actual factors contributing to improved maintainability. Kim et al. [25] developed candidate patches based on specific patterns and reported that developers were more inclined to accept these patternbased patches. Bertram et al. [8] carried out an eye–tracking study

Developers’ Trust Perceptions through Urgency and Reputation to explore how the source of a patch influences developers’ code review behavior. They reported significant differences in code review behavior and found a preference for human-written patches, noting their readability and coding style. Noller et al. [38] conducted a survey involving over 100 developers to gain insight into the setup and parameters of artifacts and tools that could bolster trust in automated program repair.

2.2

Eye Tracking and Code Review

Previous studies have leveraged eye tracking to study how developers perform code reviews. They delved into developers’ gaze patterns during code review and investigated how expertise can shape viewing strategies [9, 49, 54], pinpointed potentially problematic code elements based on this viewing patterns [7], and explored the effect of additional technical indicators such as the number of followers and activity levels on the acceptance of patches [13]. Huang et al. [22] utilized medical imaging and eye tracking to investigate the neurological correlates of biases and differences between genders of humans and machines. They reported variances in code review behavior and biases based on the perceived author.

3

EXPERIMENTAL METHODOLOGY

We recruited 37 participants for this exploratory study investigating how developers interact with and use AI assistance while working within a large open-source project. Each participant was assigned two bug-fixing tasks of equal difficulty and collaborated with either a peer or GitHub Copilot as their source of assistance.

3.1

Experiment Factors

We evaluated the effects of two independent variables in our study: urgency, represented by priority level, and a trust cue, indicated by the author’s experience, on developers’ performance and code review behavior. The priority level: Priority was presented as either “High” or “Low”. Issue tracking systems such as Jira usually propose 5 priority levels: lowest, low, medium, high, and highest. Nevertheless, previous research has indicated that priority levels in the project are not well-defined and are inconsistently utilized [6, 20, 27]. Thus, we decided to offer two levels to distinct urgent cases from the others. The author’s experience: Experience in our study consisted of two categories: “Senior” and “Junior.” Following the pattern observed in online freelance marketplaces, we included the author’s title and corresponding salary. In the bug report, the senior author was referred to as J. Miller, with the title and salary specified as “Senior Developer – $100/hr.” On the other hand, the junior author was referred to as M. Smith, identified as a “Developer – $40/hr.” To avoid bias from various human factors [13, 22], including gender, age, attractiveness, and emotional facial expressions, we refrained from using the profile pictures and utilized initials instead of the participants’ full first names.

3.2

Experiment Measures

We assessed both the performance of participants during code review tasks (i.e., objective behavior during code review) and their trust and reuse intentions (i.e., subjective self–evaluations).

Conference’17, July 2017, Washington, DC, USA Performance: To measure participant performance, first, we measured the amount of cognitive load or visual effort using eye– tracking data, which helped us understand the participant’s level of engagement. Second, we tracked the total time spent by participants to complete the tasks, giving us insights into task efficiency. Finally, we recorded the number of correct answers provided by participants, indicating their accuracy in reviewing the code. Eye gaze data was categorized into fixations and saccades [41]. Fixations are stable eye gazes lasting around 200–300 milliseconds. Researchers in psychology affirm that fixations are crucial for information acquisition and processing [24, 40]. Saccades, on the other hand, are rapid eye movements that occur between fixations and involve minimal cognitive processing. We also analyzed eye gaze data based on specific areas of interest (AOIs) [17, 48] defined in the stimulus, such as 1) bug report text, 2) entire class files with patched code, 3) specific methods with patched code, 4) entire unit test files with patch evaluation, 5) and specific unit tests evaluating the patch. We used three standard eye–tracking metrics to measure the cognitive load: fixation count, average fixation duration and total fixation time [40, 47]. Fixation count and total fixation time indicate how visual attention is dispersed and how efficiently relevant data is found. Average fixation duration reflects the concentration of attention at specific areas. Higher values may imply information extraction difficulties and working memory strain [17, 48]. Reuse intention: Participants specified whether they reuse the patch by accepting and merging it into the code base to be passed over to the end user. Trust intentions: In addition to performance measures, to capture the nuanced responses and provide a quantitative measurement, we used six Likert scale questions for quality comparison of various patches: (1) How do you think about the coding style of the patch? (Very bad (1) – Excellent(5)) (2) How do you think about the readability of the patch? (Very bad (1) – Excellent(5)) (3) How do you think about the comment (summary) the author provided for the patch? (Not clear at all (1) – Very clear (5)) (4) How much do you think the patch deploys the functionality it claims in the summary? (Very little (1) – Very much (5)) (5) How do you rank the general quality of the patch? (Very bad (1) – Excellent(5)) (6) How trustworthy do you find this code? (Untrustworthy (1) – Trustworthy (5))

3.3

Participants and Recruitment

In our study, which was approved by our Institutional Review Board (IRB), we recruited a total of 39 participants, comprising undergraduate and graduate students from the authors’ institutions. Among these participants, two were recruited for a pilot phase to ensure the smooth setup of the experiment and verify its validity. The data obtained from the other 37 participants were used for our study analysis and findings. Participants completed questionnaires to collect essential demographic information, and the summarized details can be found in Table 1. Among all the recruited students, we inquired about their gender, including options for male, female, and non-binary. However, no participant identified themselves as

Conference’17, July 2017, Washington, DC, USA non-binary. We reached out to participants via email and offered them a $25 voucher as compensation for their participation. We also inquired about participants’ experience and familiarity with programming, Java, and working with Integrated Development Environments (IDEs). On average, participants reported six years of programming experience (SD = 2.92). All participants were well–versed in Object–Oriented Programming (OOP), had a good understanding of Java, and possessed prior experience working with IDEs. We relied on participants’ self–estimation of their programming experience, a method proven to be reliable when working with students, as evidenced by Siegmund et al. [50].

3.4

Software System and Task

In this experiment, we selected jFreeChart as the system under study. jFreeChart is a Java–based open–source project that allows users to create and display charts within their applications. Specifically, we utilized version 1.1.0, which was released in 2015 and comprised approximately 300 KLOC (thousand lines of code) and around 94,000 Java classes. To ensure the authenticity of our code review tasks, we handpicked six historical bugs from a predefined set that marked as resolved, as listed in Table 2. These bugs were chosen to represent realistic scenarios. In our evaluation process, we included a variety of patches with (experimentally-controlled) different labels to represent the (purported) expertise level of the authors (Senior or Junior) and the (purported) urgency of the patches (high priority or low priority). These patches were chosen from Defects4j-Repair project [32]. Each patch consisted of a code sample, ranging in size from 8 to 76 lines of code (Mean: 11, SD: 13), with repaired lines varying from 3 to 41 (Mean: 26, SD: 23). We randomly ordered the six code review tasks for each participant to ensure fairness and confirm that all participants reviewed all six patches, considering all possible combinations of different authors and priorities. Each code review task involved examining a bug report file, which contained the following information: 1) the bug report number and the date of report, 2) the (controlled) author’s information, including their name, their salary, and skills, 3) the (controlled) patch’s priority level (high or low), 4) the summary of the bug, 5) the number and names of failing unit tests, and 6) the actual patch itself, demonstrating the code changes made to fix the issue. An example of a bug report file is shown in Figure 1 with various parts highlighted. Throughout the experiment, participants had unlimited access to the entire code repository and could freely navigate it using the Eclipse IDE.

3.5

Equipment and Setup

For our experiments, we conducted them using a 27" monitor with a screen resolution of 1920x1080 pixels. To track eye movements, we utilized the Tobii Pro Fusion eye-tracker [21]. The Tobii Pro X3–120 is a remote, non-intrusive device that captures gaze data at 250 Hz and has the capability to precisely locate eye-gaze data within a code document, down to the granularity of a single line of 10pt text. To provide a realistic environment with indicative code review tasks, we facilitated tasks such as scrolling, switching between files, and editing code by installing and utilizing the iTrace plugin [44].

Yabesi, et al. Table 1: Participant demographics: age, gender, & study level.

Characteristics Age (n (%)) 18-25 25-30 30-35 35 or older

All (N = 37)

Men (n = 25)

Women (n = 12)

18 (48%) 10 (27%) 5 (13%) 4 (10%)

16 5 2 2

2 5 3 2

Study level (n (%)) First year Sophomore Junior Senior M.S./PhD

2 (5%) 2 (5%) 4 (10%) 1 (2%) 28 (75%)

2 1 3 1 18

0 1 1 0 10

Figure 1: Distribution and intensity of visual attention across the Code Pane and Stack Trace AOIs for two assistance types (Copilot and Peer), averaged across all participants. Participants exhibited greater visual effort in evaluating outputs and errors when using Copilot, and also showed increased focus on the Project Structure pane.

This state–of–the–art plugin allowed participants to interact with the source code and other relevant materials naturally while also collecting the necessary measurements. Using iTrace Toolkit, the

Developers’ Trust Perceptions through Urgency and Reputation

Conference’17, July 2017, Washington, DC, USA

Table 2: Explanation of the patches, including a brief summary of the reported bug alongside the affected sections of code. Scope of the solution Classes Test Classes Lines of (Methods) (Unit Tests) code

Short Description Patch 1: Max-y value is not updated when copying a subset of TimeSeries. Patch 2: XYSeries.addOrUpdate() should handle duplicate X values, as a prior change supported. Yet, the method was not updated, resulting in overwriting of the data. Patch 3: Potential Null Pointer Exception in AbstractCategoryItemRender.getLegendItems() results in a null pointer access in the class. Patch 4: If the label generator returns null, the PieChart must be created. Patch 5: The cached bounds must be reset whenever items are added or removed from the plots. Now, we need to iterate over the entire dataset to update the cached values. Patch 6: We have a null pointer exception in StatisticalBarRenderer when one of a series in a category has no data. recorded raw data was then subjected to the Velocity Threshold Identification (I–VT) filter to generate fixations.

3.6

Procedure

We conducted the experiment in a quiet room with an eye tracker. Participants were seated approximately 70 cm from the screen in a comfortable swivel chair with armrests. Before running the experiment, all participants signed a consent form, and the experimenter verbally explained the experiment’s procedure in detail. Participants were informed that the experiment consisted of one Java project and that they would be sequentially reviewing six patches. The experimenter did not explain the particular goal of the experiment. Participants were given ten minutes per code review task and were instructed to inform the experimenter if they finished early to stop the tracking. To mitigate any learning effects, participants received the six tasks in randomized order. To better control and avoid any other factors impacting the results, the participants were instructed to maintain the full–screen IDE setup, not use the debugger, and not browse the internet. Upon completing each task, participants were presented with a survey that encompassed various aspects. The survey included questions regarding the overall quality of the patches, the level of trust for critical tasks, the quality of implemented functionality, patch readability, and patch coding style. Participants were also required to indicate their preference of “accept” or “reject” for the corresponding task within the survey. Following the completion of all six code review tasks and their associated surveys, participants were asked to complete a post-questionnaire. To decrease potential stereotype threat [45], we strategically placed questions about coding knowledge and experience at the end of the study. Specifically, women and marginalized groups encounter the detrimental perception that their abilities are inferior more profoundly than individuals from other backgrounds [52]. Additionally, the post-questionnaire contained inquiries about participants’ age range, gender, population group, native language, level of study, field of study, years of programming experience, the influence of the author’s expertise level, and the influence of urgency in accepting or rejecting the patches.

4

1 (1) 1 (1)

1 (1) 1 (1)

9 14

1 (1)

1 (1)

10

1 (1) 1 (4)

1 (1) 1 (1)

27 76

1 (1)

1 (4)

24

DATA ANALYSIS AND RESULTS

This study investigated how developers perceive the trustworthiness of code patches, the extent to which the perceived priority level of the patches, and the experience level of the author affect developers’ efficiency and behavior during code review. We ensured the code remained consistent while we varied the patch’s listed priority level (high or low) and the author’s listed experience (senior or junior). Our analysis primarily centered on the responses to the following research questions: RQ1. How well do participants’ self–reports regarding the role of patches’ priority level and author’s experience correspond with our collected data? RQ2. What effect do the priority levels of patches and the author’s experience have on participants’ performance in code review? RQ3. How are the participants’ code review behaviors influenced by the priority levels of patches and the author’s experience during code review? We provide access to our study material and de–identified dataset, consisting of behavioral data, eye–tracking data, and survey data, for analysis and replication at link.

4.1

RQ1. Self–Reporting and Trust

In our study, all 37 participants answered the post–questionnaire regarding the tasks and their experiences to provide insights into their subjective evaluations of trust. The participants answered six questions concerning the overall quality of the patches, the level of trust for critical tasks, the quality of implemented functionality, patch readability, summary of the patch, and patch coding style. The results of the Likert scale questions for the priority level and author’s experience are presented in Figure 2 and Figure 3, respectively. Our data reveals interesting patterns in how patch priority levels are perceived. The participants tended to assign higher rankings to patches labeled as “High priority” for the overall quality, and the result is statistically significant (Mann-Whitney test, 𝑊 = 6910.5, 𝑝 = .03). The proportion test (Chi-squared test for significance) regarding patch acceptance and the Wilcoxon signed-rank test for trust levels

Conference’17, July 2017, Washington, DC, USA

Yabesi, et al.

Figure 2: Ratings of patch quality attributes and trustworthiness by the priority level with a 5–point Likert scale (222 responses). “High” priority patches were rated higher in coding style, functionality, trustworthiness, and in a statistically significant manner (𝑝 = .03), superior in overall quality by participants.

Figure 3: Ratings of patch quality attributes and trustworthiness by the author’s experience with a 5-point Likert scale (222 responses). Patches labeled as written by a senior developer were generally ranked higher. did not distinguish between high and low–priority patches. Participants, on average, were 6% less likely to accept patches classified as low–priority. 68% of participants indicated that the perceived urgency or severity of the issues addressed by the patches did not impact their decision to accept or reject them. Yet, high–priority patches were rated as higher quality by participants. This finding suggests that even if the issue’s urgency does not directly affect their acceptance decision, it does affect their perception of the patch’s quality. In addition, we found no evidence that the author’s experience impacted participants’ self–reported data related to quality factors. Around 60% of the participants mentioned that the experience level of the authors doesn’t influence their judgment on whether to accept and or reject the patch. We found no compelling evidence suggesting that the author’s experience level significantly impacts participants’ intention to reuse.

RQ1 — Likert scale self-reporting and trust Patch priority: Participants did not indicate any impact of the patch’s priority on their evaluations and intention to reuse. High– priority patches were rated higher in overall quality, despite controlling for quality. Author’s experience: There was no evidence to support the author’s experience affecting participants’ self-reported quality factors. To avoid biasing participants’ self-reports in a specific direction, we also utilized open–ended free response questions: (1) Regarding prioritization, what were the top three key factors you considered when making decisions during code reviews? (2) Does the urgency and severity of the issue affect your decision to accept or reject the patch? Please clarify. (3) Does the expertise of the patch’s author affect your decision to accept or reject the patch? We carried out a thematic analysis of responses from all 37 participants to the aforementioned questions, focusing on the crucial

Developers’ Trust Perceptions through Urgency and Reputation

Conference’17, July 2017, Washington, DC, USA Table 3: Pairwise comparisons of performance metrics using the Chi–squared test for accuracy and acceptance rate and non–parametric Wilcoxon Test (𝛼 = 0.05) for time and visual effort metrics for priority levels (High vs. Low). Significant results (𝑝 < 0.05) are bolded. A significant trend showed that participants devoted more visual attention and task time to analyzing high–priority patches.

Priority level Acceptance rate Accuracy Time spent (s) Avg. Fx. Duration (ms) Fx. Count Total Fx. Time (s)

Figure 4: A Breakdown of the top three self-reported factors and their subcategories in code review evaluation, displayed with self-reported frequencies. Three key factors that guide code review decisions are functionality (highlighted by 44.7% of participants), code quality (26.3%), and comprehensibility, mainly represented by the quality of the comments and summaries (22.4%).

RQ1 — Qualitative analysis of self-reporting and trust Functionality, quality (readability and style), and code comprehensibility are the top three factors in patch evaluation.

4.2

RQ2. Performance Differences

We investigated whether the patches’ priority level and the author’s experience impact the participants’ overall performance measured by accuracy, total time spent, and visual effort. The results, as summarized in Table 3 and Table 4, were analyzed using proportion

56% 70% 449 (161) 647.8 (211.7) 428.9 (190.9) 271 (123)

50% 63% 406 (153) 644.8 (209.2) 398.3 (61.2) 236 (104)

𝑝 0.5 0.3 0.04 0.7 0.1 0.03

Table 4: Pairwise comparisons of performance metrics using Chi–squared test for accuracy and acceptance rate and non– parametric Wilcoxon Test (𝛼 = 0.05) for time and visual effort metrics for experience level (Senior vs. Junior). Significant results (𝑝 < 0.05) are bolded. We found no impact of the the author’s experience on participants’ performance and acceptance rate.

Author’s experience factors they used to assess code patches. Through a two-step process involving descriptive coding and subsequent discussions, we identified and grouped these essential elements. As shown in Figure 4, the findings revealed three primary factors that participants considered: 1) functionality and whether the code passes tests was the most significant factor, with 44.7% of participants mentioning that whether the code functions as expected and passes all tests is crucial to their decision-making process, 2) code quality with a focus on readability account for 26.3% of responses and was the second most influential factor, and 3) code comprehensibility and the quality of comments and summaries (22.4%). Other factors accounted for the remaining 6.6% of responses. Notably, only 2.6% of the participants reported the author’s experience as a critical factor in their decision-making process. The self–reported data show that in the context of code review, the merit of the code itself outweighs considerations about the author’s experience or the patch’s priority.

Mean (Std. Dev.) High Low

Acceptance rate Accuracy Time spent (s) Avg. Fx. Duration (ms) Fx. Count Total Fx. Time (s)

Mean (Std. Dev.) Senior Junior 54% 68% 429 (150) 640.4 (165.5) (168.5) 257 (111)

57% 65% 420 (159) 652.6 (249.2) 408.0 (186.6) 252 (120)

𝑝 0.9 0.6 0.9 0.8 0.5 0.8

tests (Chi-squared test) for our categorical variables, accuracy (0 /1) and acceptance rate (Accept/Reject) and Wilcoxon signed-rank test for total time as a continuous variable. To assess visual effort, we examined fixation count (Fx. Count), average fixation duration (Avg. Fx. Duration), and total fixation time (Fx. Time), indicators of cognitive load. A higher number of fixations and longer fixation duration suggest increased visual effort and mental load or denote the importance of the focused area. Previous studies have reported the influence of participants’ trust within the system on various eye-tracking metrics [15, 16, 29]. The findings revealed no significant differences in acceptance rate, accuracy, fixation count, and average fixation duration based on either the patches’ priority level or the author’s experience level. However, as shown in Table 3, we observe longer fixation time and task time for high–priority patches in a statistically significant manner. The perceived urgency or potential impact associated with high–priority patches led to a more meticulous evaluation of code patches by our participants.

Conference’17, July 2017, Washington, DC, USA

Figure 5: Distribution and the intensity of visual attention across AOIs for various priority and experience levels, averaged for all participants. The participants spent more visual effort evaluating the relevant unit tests for high–priority patches. In contrast, they spent more effort on relevant methods for low–priority patches. Also, participants focused more on relevant methods for senior labeled patches. RQ2 — Performance differences Patch priority: Patches labeled as high–priority demanded significantly more task time and visual effort (i.e., cognitive load). Author’s experience: No significant performance metric variations are tied to the author’s experience level. To gain further insights, we analyzed how equal amounts of time were allocated differently based on urgency and reputation.

4.3

RQ3. Differences in Code Review Behaviour

We examined how participants’ code review behavior and problem– solving strategy are influenced by the priority level of patches and the authors’ experience with an analysis of attention distribution and navigation trends over time throughout a task. We calculated our three eye–tracking metrics (average fixation duration, fixation count, and total fixation time) on each AOI and compare the distribution of attention across AOIs for patches labeled as high vs. low–priority and senior vs. junior–authored. For priority level, the general align-and–rank non–parametric factorial analysis, as described by Wobbrock et al. [55], indicates a significant interaction between the marked priority level and the distribution of visual attention for all three metrics (Fx. Count: 𝐹 (1, 5) = 3.6, 𝑝 < .05, Avg. Fx. Duration: 𝐹 (1, 5) = 2.7, 𝑝 < .05, and Total Fx. Time: 𝐹 (1, 5) = 4.1, 𝑝 = .001). Our results also highlighted the significant interaction between marked author’s experience and the visual attention distribution for two metrics out of three: fixation count, and total fixation time (Fx.

Yabesi, et al. Count: 𝐹 (1, 5) = 2.5, 𝑝 < .05, Avg. Fx. Duration: 𝐹 (1, 5) = 0.11, 𝑝 = .9, and Total Fx. Time: 𝐹 (1, 5) = 3.0, 𝑝 = .009). Figures 5 illustrates the distinct code scanning behavior of participants across different areas of interest (AOIs) when patches with various labels (low vs. high–priority and senior vs. junior– authored). These findings provide evidence that the perceived priority of patches and the author’s experience significantly influence participants’ perception of the relevance of different code sections. While higher–level performance metrics may not detect a noticeable difference, our analysis reveals that participants exhibit diverse scanning behaviors characterized by attention distribution and intensity variations during code review. Specifically, for priority level, we observed that participants exerted more significant visual effort when evaluating the relevant unit tests for high-priority patches. Conversely, participants invested more effort in examining relevant methods when reviewing low–priority patches. For the author’s experience, we also observe more focused visual attention on relevant methods for senior labeled patches. In addition, we investigate how our participants spend their time reviewing the bug report and where they look. More specifically, developers consider what elements and whether this varies with the patch’s priority level or the author’s experience. We identified that participants fixated more on the patch summary and the proposed code changes. Similar to the results of our qualitative analysis of participants’ self–report data (cf. Section 4.1), most participants focused on code and technical parts. Also, in a statistically–significant manner, our participants spent more time and cognitive effort (Wilcoxon test with Bonferroni adjustment: 𝑝 = 0.03 for Fixation count and Fixation time) analyzing the code part of the high–priority patches as shown in Figure 6. The findings shed light on the intricacies and challenges involved in comprehending developers’ perception of trust towards various tools and artifacts. It underscores the advantages of adopting a multi-modal approach in exploring this complex phenomenon. Although we might not have observed a direct impact on performance metrics, our eye-tracking results reveal noteworthy variations in the allocation of visual attention and cognitive resources based on the urgency and reputation under investigation. These insights provide valuable implications for understanding how developers interact with and rely on various parts of the source code.

RQ3 — Code review behavior Patch priority: AOI relevance — high–level scanning patterns of participants — varies significantly for high and low–priority patches. Participants put more attention and cognitive load into evaluating the relevant unit tests and the code part in bug reports for high–priority patches. Author’s experience: The experience level of authors impacts the overall viewing strategies of participants. More attention on reading and processing relevant methods for patches labeled as written by senior developers. No significant differences in AOI relevancy were found as a function of the author’s experience with bug reports.

Developers’ Trust Perceptions through Urgency and Reputation

Conference’17, July 2017, Washington, DC, USA

Figure 6: Summary of visual attention distribution for various parts of the bug report files for all participants for priority levels (top row) and author’s experience (bottom row). No significant effect of author’s experience was found on eye–tracking high–level metrics in isolation. However, participants, on average, spent more time and effort (fixation count and total fixation time) analyzing the code part of high–priority patches, highlighted with green rectangles.

5

DISCUSSION OF THE RESULTS

Based on the dual process model of persuasion [39], individuals interpret information via two paths: central and peripheral. Central route or systematic processing relies on logical reasoning and analytical skills. On the other hand, peripheral processing employs heuristics, stereotypes, and biases to lessen the cognitive burden during decision-making [53]. Peripheral processing tends to occur more when decisions are emotionally charged, trusted, or timelimited. In this study, to examine trust, we thoughtfully employ cues of urgency and reputation to stimulate peripheral processing. This method enables us to investigate biases and inclinations towards priority levels and authors’ experience precisely, setting our research apart as the first to do so.

5.1

The Role of the Priority Level

The findings of our study indicate a variance in how developers execute code reviews (i.e., noticeable variation in attention allocation), which correlates with the priority of the given issue. We support these findings through eye tracking and behavioral data. In addition, the patch priority has a statistically significant effect on the time taken to handle the tasks.

Contrary to some previous studies [6, 27], our findings reveal no significant difference in the acceptance rate based on the priority level of patches, adding nuance to the existing understanding. While these studies suggested that high–priority patches were more likely to be incorporated into the project’s codebase through post–factum analysis of GitHub patches, they also highlighted a pertinent issue. Developers often display confusion over the meaning and implications of standard priority levels, ranging from P1 (low) to P5 (high). This indicates a potential gap in the shared understanding and interpretation of these priority scales across the community. Adding further to this complexity, our results suggested that participants often overlooked patch priority when determining whether to accept and reuse the code. This divergence could stem from the code’s quality, functionality, or readability, overshadowing the assigned priority level in a developer’s decision–making process. Alternatively, the mismatch could arise from the need to clarify what each priority level implies. Thus, our findings call for revising and enhancing software development guidelines and code review tools to emphasize the significance of patch priority levels and to present assessment criteria clearly. These enhancements enable developers to prioritize urgent

Conference’17, July 2017, Washington, DC, USA patches, resulting in streamlined decision-making, improved code quality, and software development efficiency.

5.2

The Role of Experience Level

Surprisingly, the author’s experience had no effect on participants’ trustworthiness perceptions, reuse intentions, or performance. A couple of previous work in computing and trust has demonstrated the impact of code reputation on developers’ trust and reuse perceptions and outcomes [1, 2]. In our research, we chose to focus on the human elements of trust and controlled the code’s reputation through the experience level of the author. This approach differs from past work, which established reputation by either modifying the metadata attributes of OSS systems (for example, total commits, likes, and star ratings) [2] or labeling the source as either reputable or unknown [1]. By removing the profile picture and using initials, we avoided biases based on gender, age, attractiveness, and emotional facial expressions. In addition, we provided a complete profile for both novice and senior authors, which implies trustworthiness and professionalism [13]. Our findings prompt reevaluating how the software engineering community perceives and utilizes reputation and what constitutes an effective reputation metric. The implications would make the experience level less of a barrier for new or less experienced developers to be accepted and evaluated based on their technical merits. Although we found no differences in our performance metrics, our analysis of the distribution of visual attention and the intensity of visual processing reveal different AOI preferences based on the author’s experience. Our results highlight that the apparent experience of the patch’s author influences how developers judge the code. These findings broadly align with these previous studies in finding statistically significant differences in how developers carry out the task of code review [8, 22]. Their results also highlight the impact of the patch’s provenance (men vs. women vs. machine) on developers’ viewing strategies. The reported differences in scanning patterns based on the author’s labeled experience, revealed by the eye-tracking data, are likely attributed to differences in training or feedback. Previous work also stated that in addition to technical aspects, social ones (usually related to the author and presented in their user profile) also are being considered by developers when deciding upon patch review and acceptance [13]. In the same vein, Alarcon et al. [2] reported more involvement with the code for more reputable patches. Prior studies in automation have found a relationship between age and inclination to trust [3, 43]. Given that our sample group consisted of young students with minimal professional experience, this could account for the lack of observable impact on performance metrics and trust self-evaluation associated with the author’s experience in our findings.

6

THREATS TO VALIDITY

Internal validity: To minimize the instrument bias, we employed a video–based eye–tracking system that does not require cumbersome goggles and permits participants to move their heads without affecting the camera’s calibration. The camera calibration was carefully executed at the start of the study and adjusted between tasks to ensure accuracy. To prevent the treatments diffusion, participants

Yabesi, et al. were advised against discussing the experiment. To mitigate the Hawthorne effect, we clarified that the eye tracker only captures eye-related data, and our experimenter was seated discreetly at a distance from the participants to minimize their sense of being watched. To mitigate stereotype threat and avoid influencing performance [45, 51], we saved self–assessment of coding skills for the study’s end. Construct Validity: We reduced the risk of hypothesis guessing and apprehension by withholding the exact aims of the study from participants. However, we did ensure they were well-informed about the study’s procedure, session count, task types, and the functioning of the eye tracker prior to the experiment. External validity: All of our participants were students, with more than 85% being graduate students with good programming expertise. Using students as participants is acceptable when evaluating techniques for the novice or non-expert software engineers, as they represent the future generation of software professionals [26]. The choice of the JFreeChart project as the subject system poses a validity threat due to its potential influence on the study. However, we mitigated this risk by selecting a reasonably large and complex open–source project in a popular programming language. We also chose the historic bugs from the project repository that accurately represent real–world software issues. Conclusion validity: To mitigate this threat, we utilized established eye–tracking metrics, employed standard statistical methods for data analysis, and conducted eye–tracker calibration for each participant before each task.

7

CONCLUSION

Deriving insights from the psychological aspects of trust and compliance, we conducted a carefully controlled experiment with 37 participants using advanced eye–tracking technology to delve deeper into the factors influencing developers’ trust–based cognitive processes. This novel research offers a holistic picture of how urgency and reputation impact developers’ perceptions of trustworthiness and evaluation during code reviews. Our qualitative and quantitative study indicates that patch priority significantly impacted developers’ review behavior and their perception of code quality without affecting reuse decisions. Interestingly, the author’s experience level of the code patches altered participants’ attention distribution but not review performance. In addition, the most commonly self–reported factors in code evaluation that affect our participants’ decisions were: (1) functionality, (2) code quality and (3) code quality. The rising trend of reusing auto–generated code, such as Facebook’s SapFix [31], and collaborating with AI pair programmers like GitHub’s Copilot 1 , underscores the critical role of trust in software development. It also prompts exploring how much developers should rely on and reuse such generated code. Our research represents a significant step forward in understanding the dynamics of trust in software engineering tasks from developers’ perspectives. By gaining a deeper understanding of factors underlying the trustworthiness of software artifacts, we can work towards enhancing the efficiency of code reuse practices and code automation, ultimately leading to more robust and dependable software systems. 1 https://github.com/features/copilot

Developers’ Trust Perceptions through Urgency and Reputation

ACKNOWLEDGMENTS To Robert, for the bagels and explaining CMYK and color spaces.

REFERENCES [1] Gene M Alarcon, Rose Gamble, Sarah A Jessup, Charles Walter, Tyler J Ryan, David W Wood, and Chris S Calhoun. 2017. Application of the heuristicsystematic model to computer code trustworthiness: The influence of reputation and transparency. Cogent Psychology 4, 1 (2017), 1389640. [2] Gene M. Alarcon, Anthony M. Gibson, Charles Walter, Rose F. Gamble, Tyler J. Ryan, Sarah A. Jessup, Brian E. Boyd, and August Capiola. 2020. Trust Perceptions of Metadata in Open-Source Software: The Role of Performance and Reputation. Systems 8, 3 (2020). https://doi.org/10.3390/systems8030028 [3] Gene M. Alarcon and Tyler J. Ryan. 2018. Trustworthiness Perceptions of Computer Code: A Heuristic-Systematic Processing Model. In Proceedings of the 51st Hawaii International Conference on System Sciences. [4] Gene M. Alarcon, Charles Walter, Anthony M. Gibson, Rose F. Gamble, August Capiola, Sarah A. Jessup, and Tyler J. Ryan. 2020. Would You Fix This Code for Me? Effects of Repair Source and Commenting on Trust in Code Repair. Systems 8, 1 (2020). https://doi.org/10.3390/systems8010008 [5] Apostolos Ampatzoglou, Antonios Gkortzis, Sofia Charalampidou, and Paris Avgeriou. 2013. An Embedded Multiple-Case Study on OSS Design Quality Assessment across Domains. In 2013 ACM / IEEE International Symposium on Empirical Software Engineering and Measurement. 255–258. https://doi.org/10. 1109/ESEM.2013.48 [6] Olga Baysal, Oleksii Kononenko, Reid Holmes, and Michael W. Godfrey. 2013. The influence of non-technical factors on code review. In 2013 20th Working Conference on Reverse Engineering (WCRE). 122–131. https://doi.org/10.1109/ WCRE.2013.6671287 [7] Andrew Begel and Hana Vrzakova. 2018. Eye Movements in Code Review. In Proceedings of the Workshop on Eye Movements in Programming. Article 5, 5 pages. https://doi.org/10.1145/3216723.3216727 [8] Ian Bertram, Jack Hong, Yu Huang, Westley Weimer, and Zohreh Sharafi. 2020. Trustworthiness Perceptions in Code Review: An Eye-Tracking Study. In Proceedings of the 14th ACM / IEEE International Symposium on Empirical Software Engineering and Measurement (ESEM) (Bari, Italy) (ESEM ’20). Association for Computing Machinery, New York, NY, USA, Article 31, 6 pages. https://doi.org/10.1145/3382494.3422164 [9] Martha E. Crosby and Jan Stelovsky. 1990. How Do We Read Algorithms? A Case Study. Computer 23, 1 (Jan. 1990), 24–35. [10] Zakir Durumeric, Frank Li, James Kasten, Johanna Amann, Jethro Beekman, Mathias Payer, Nicolas Weaver, David Adrian, Vern Paxson, Michael Bailey, and J. Alex Halderman. 2014. The Matter of Heartbleed. In Proceedings of the 2014 Conference on Internet Measurement Conference (Vancouver, BC, Canada) (IMC ’14). Association for Computing Machinery, New York, NY, USA, 475–488. https://doi.org/10.1145/2663716.2663755 [11] Anneli Eteläpelto. 1993. Metacognition and the expertise of computer program comprehension. Scandinavian Journal of Educational Research 37, 3 (1993), 243– 254. [12] Quyin Fan. 2010. The Effects of Beacons, Comments, and Tasks on Program Comprehension Process in Software Maintenance. Ph. D. Dissertation. University of Maryland, Baltimore County, Catonsville, MD, USA. Advisor(s) Norcio, Anthony F. AAI3422807. [13] Denae Ford, Mahnaz Behroozi, Alexander Serebrenik, and Chris Parnin. 2019. Beyond the code itself: how programmers really look at pull requests. In International Conference on Software Engineering: Software Engineering in Society. [14] Zachary P. Fry, Bryan Landau, and Westley Weimer. 2012. A human study of patch maintainability. In International Symposium on Software Testing and Analysis, ISSTA 2012, Minneapolis, MN, USA, July 15-20, 2012, Mats Per Erik Heimdahl and Zhendong Su (Eds.). ACM, 177–187. https://doi.org/10.1145/2338965.2336775 [15] Claudia Geitner, Ben D Sawyer, S Birrell, P Jennings, L Skyrypchuk, Bruce Mehler, and Bryan Reimer. 2017. A link between trust in technology and glance allocation in on-road driving. (2017). [16] Christian Gold, Moritz Körber, Christoph Hohenberger, David Lechner, and Klaus Bengler. 2015. Trust in automation–Before and after the experience of take-over scenarios in a highly automated vehicle. Procedia Manufacturing 3 (2015), 3025–3032. [17] Joseph H. Goldberg and Jonathan I. Helfman. 2010. Comparing Information Graphics: A Critical Look at Eye Tracking. In Proceedings of the 3rd BEyond Time and Errors: Novel evaLuation Methods for Information Visualization Workshop (Atlanta, Georgia) (BELIV ’10). ACM, New York, NY, USA, 71–78. https://doi. org/10.1145/2110192.2110203 [18] Claire Goues, Stephanie Forrest, and Westley Weimer. 2013. Current Challenges in Automatic Software Repair. Software Quality Journal 21, 3 (Sept. 2013), 421–443. https://doi.org/10.1007/s11219-013-9208-0 [19] Stefan Haefliger, Georg Von Krogh, and Sebastian Spaeth. 2008. Code reuse in open source software. Management science 54, 1 (2008), 180–193.

Conference’17, July 2017, Washington, DC, USA [20] Israel Herraiz, Daniel M. German, Jesus M. Gonzalez-Barahona, and Gregorio Robles. 2008. Towards a Simplification of the Bug Report Form in Eclipse. In Proceedings of the 2008 International Working Conference on Mining Software Repositories (Leipzig, Germany) (MSR ’08). Association for Computing Machinery, New York, NY, USA, 145–148. https://doi.org/10.1145/1370750.1370786 [21] https://www.tobiipro.com/. 2001. Online; Accessed 17-07-2020. [22] Yu Huang, Kevin Leach, Zohreh Sharafi, Nicholas McKay, Tyler Santander, and Westley Weimer. 2020. Biases and Differences in Code Review Using Medical Imaging and Eye-Tracking: Genders, Humans, and Machines. In International Symposium on the Foundations of Software Engineering (Virtual Event, USA) (ESEC/FSE 2020). Association for Computing Machinery, New York, NY, USA, 456–468. https://doi.org/10.1145/3368089.3409681 [23] Yu Huang, Xinyu Liu, Ryan Krueger, Tyler Santander, Xiaosu Hu, Kevin Leach, and Westley Weimer. 2019. Distilling Neural Representations of Data Structure Manipulation Using FMRI and FNIRS. In Proceedings of the 41st International Conference on Software Engineering (ICSE ’19). IEEE Press, Montreal, Quebec, Canada, 396–407. https://doi.org/10.1109/ICSE.2019.00053 [24] Marcel A Just and Patricia A Carpenter. 1980. A theory of reading: from eye fixations to comprehension. Psychological review 87, 4 (1980), 329. [25] Dongsun Kim, Jaechang Nam, Jaewoo Song, and Sunghun Kim. 2013. Automatic patch generation learned from human-written patches. In 2013 35th International Conference on Software Engineering (ICSE). IEEE, 802–811. [26] Barbara A. Kitchenham, Shari Lawrence Pfleeger, Lesley M. Pickard, Peter W. Jones, David C. Hoaglin, Khaled El Emam, and Jarrett Rosenberg. 2002. Preliminary Guidelines for Empirical Research in Software Engineering. IEEE Transactions on Software Engineering 28, 8 (Aug. 2002), 721–734. https://doi.org/ 10.1109/TSE.2002.1027796 [27] Oleksii Kononenko, Olga Baysal, Latifa Guerrouj, Yaxin Cao, and Michael W. Godfrey. 2015. Investigating code review quality: Do people and participation matter?. In 2015 IEEE International Conference on Software Maintenance and Evolution (ICSME). 111–120. https://doi.org/10.1109/ICSM.2015.7332457 [28] John D Lee and Katrina A See. 2004. Trust in automation: Designing for appropriate reliance. Human factors 46, 1 (2004), 50–80. [29] Y. Lu and N. Sarter. 2019. Eye Tracking: A Process-Oriented Method for Inferring Trust in Automation as a Function of Priming and System Reliability. IEEE Transactions on Human-Machine Systems 49, 6 (Dec 2019), 560–568. https: //doi.org/10.1109/THMS.2019.2930980 [30] Joseph B Lyons, Nhut T Ho, William E Fergueson, Garrett G Sadler, Samantha D Cals, Casey E Richardson, and Mark A Wilkins. 2016. Trust of an automatic ground collision avoidance technology: A fighter pilot perspective. Military Psychology 28, 4 (2016), 271–277. [31] A. Marginean, J. Bader, S. Chandra, M. Harman, Y. Jia, K. Mao, A. Mols, and A. Scott. 2019. SapFix: Automated End-to-End Repair at Scale. In International Conference on Software Engineering: Software Engineering in Practice. 269–278. [32] Matias Martinez, Thomas Durieux, Romain Sommerard, Jifeng Xuan, and Martin Monperrus. 2016. Automatic Repair of Real Bugs in Java: A Large-Scale Experiment on the Defects4J Dataset. Springer Empirical Software Engineering (2016). https://doi.org/10.1007/s10664-016-9470-4 [33] Stephanie Merritt, Lei Shirase, and Garett Foster. 2020. Normed Images for X-ray Screening Vigilance Tasks. Journal of Open Psychology Data 8, 1 (2020). [34] Audris Mockus. 2007. Large-Scale Code Reuse in Open Source Software. In First International Workshop on Emerging Trends in FLOSS Research and Development (FLOSS’07: ICSE Workshops 2007). 7–7. https://doi.org/10.1109/FLOSS.2007.10 [35] Martin Monperrus. 2018. Automatic Software Repair: A Bibliography. ACM Comput. Surv. 51, 1, Article 17 (Jan. 2018), 24 pages. https://doi.org/10.1145/ 3105906 [36] Martin Monperrus, Simon Urli, Thomas Durieux, Matias Martinez, Benoit Baudry, and Lionel Seinturier. 2019. Repairnator Patches Programs Automatically. Ubiquity 2019, July, Article 2 (July 2019), 12 pages. https://doi.org/10.1145/3349589 [37] Rennie Naidoo. 2015. Analysing urgency and trust cues exploited in phishing scam designs. In 10th International Conference on Cyber Warfare and Security. 216. [38] Yannic Noller, Ridwan Shariffdeen, Xiang Gao, and Abhik Roychoudhury. 2022. Trust Enhancement Issues in Program Repair. In Proceedings of the 44th International Conference on Software Engineering (Pittsburgh, Pennsylvania) (ICSE ’22). Association for Computing Machinery, New York, NY, USA, 2228–2240. https://doi.org/10.1145/3510003.3510040 [39] Richard E Petty, John T Cacioppo, Richard E Petty, and John T Cacioppo. 1986. The elaboration likelihood model of persuasion. Springer. [40] Alex Poole and Linden J. Ball. 2005. Eye Tracking in Human-Computer Interaction and Usability Research: Current Status and Future. In Prospects”, Chapter in C. Ghaoui (Ed.): Encyclopedia of Human-Computer Interaction. Pennsylvania: Idea Group, Inc. Information Science Reference, Hershey, PA, 1–5. [41] K. Rayner. 1978. Eye movements in reading and information processing. Psychological Bulletin 85, 3 (1978), 618–660. [42] Paige Rodeghero. 2017. Behavior-Informed Algorithms for Automatic Documentation Generation. In 2017 IEEE International Conference on Software Maintenance

Conference’17, July 2017, Washington, DC, USA and Evolution (ICSME). IEEE Computer Society, Los Alamitos, CA, USA, 660–664. https://doi.org/10.1109/ICSME.2017.73 [43] Tyler J Ryan, Gene M Alarcon, Charles Walter, Rose Gamble, Sarah A Jessup, August Capiola, and Marc D Pfahler. 2019. Trust in automated software repair: The effects of repair source, transparency, and programmer experience on perceived trustworthiness and trust. In HCI for Cybersecurity, Privacy and Trust: First International Conference, HCI-CPT 2019, Held as Part of the 21st HCI International Conference, HCII 2019, Orlando, FL, USA, July 26–31, 2019, Proceedings 21. Springer, 452–470. [44] Timothy R Shaffer, Jenna L Wise, Braden M Walters, Sebastian C Müller, Michael Falcone, and Bonita Sharif. 2015. itrace: Enabling eye tracking on software artifacts within the ide to support software engineering tasks. In Proceedings of the 2015 10th Joint Meeting on Foundations of Software Engineering. 954–957. [45] Jenessa R Shapiro and Steven L Neuberg. 2007. From stereotype threat to stereotype threats: Implications of a multi-threat framework for causes, moderators, mediators, consequences, and interventions. Personality and Social Psychology Review 11, 2 (2007), 107–130. [46] Zohreh Sharafi, Yu Huang, Kevin Leach, and Westley Weimer. 2020. Towards an objective measure of developers’ cognitive activities. In Transactions on Software Engineering and Methodology (TOSEM). ACM, to appear. [47] Zohreh Sharafi, Timothy Shaffer, Bonita Sharif, and Yann-Gaël Guéhéneuc. 2015. Eye-tracking metrics in software engineering. In Proceeding of 2015 Asia-Pacific Software Engineering Conference (APSEC). IEEE, 96–103. [48] Zohreh Sharafi, Bonita Sharif, Yann-Gaël Guéhéneuc, Andrew Begel, Roman Bednarik, and Martha Crosby. 2020. A practical guide on conducting eye tracking studies in software engineering. Empirical Software Engineering (2020), 1–47. [49] Bonita Sharif, Michael Falcone, and Jonathan I. Maletic. 2012. An Eye-Tracking Study on the Role of Scan Time in Finding Source Code Defects. In Symposium on Eye Tracking Research and Applications. https://doi.org/10.1145/2168556.2168642 [50] Janet Siegmund, Christian Kästner, Jörg Liebig, Sven Apel, and Stefan Hanenberg. 2014. Measuring and modeling programming experience. Empirical Software Engineering 19, 5 (2014), 1299–1334. [51] Steven J Spencer, Claude M Steele, and Diane M Quinn. 1999. Stereotype threat and women’s math performance. Journal of experimental social psychology 35, 1 (1999), 4–28. [52] Claude M Steele and Joshua Aronson. 1995. Stereotype threat and the intellectual test performance of African Americans. Journal of personality and social psychology 69, 5 (1995), 797. [53] Cass R Sunstein. 2005. Moral heuristics. Behavioral and brain sciences 28, 4 (2005), 531–541. [54] Hidetake Uwano, Masahide Nakamura, Akito Monden, and Ken-ichi Matsumoto. 2006. Analyzing Individual Performance of Source Code Review Using Reviewers’ Eye Movement. In Eye Tracking Research & Applications. 133–140. [55] Jacob O Wobbrock, Leah Findlater, Darren Gergle, and James J Higgins. 2011. The aligned rank transform for nonparametric factorial analyses using only anova procedures. In Proceedings of the SIGCHI conference on human factors in computing systems. ACM, 143–146.

Received 20 February 2007; revised 12 March 2009; accepted 5 June 2009

Yabesi, et al.

Related documents

Record · ID 6037 · SHA-256 bc7b22ae1d2736ed
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.