Conceptio › Archive › arXiv CS
arXiv CSopen access

T(r)opical Islands: Visualizing & Understanding Socio-Technical Artifacts

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

T(r)opical Islands: Visualizing & Understanding Socio-Technical Artifacts Adam Štěpánek∗ , Marco Raglianti† , Vı́t Rusňák∗ , Jan Byška∗‡ , Barbora Kozlı́ková∗ , Michele Lanza† ∗ Visitlab, Masaryk University, Brno, Czech Republic † REVEAL @ Software Institute, USI, Lugano, Switzerland

arXiv:2609.05254v1 [cs.SE] 4 Sep 2026

‡ VisGroup, University of Bergen, Bergen, Norway

Abstract—Projects hosted on collaborative software development platforms, such as GitHub, include many non-code artifacts documenting the project’s lifecycle, with its challenges, plans, design, and even community. These socio-technical artifacts include, for example, bug reports, feature requests, and forum posts, offering a useful prospect on the project’s evolution. However, these artifacts are dispersed over multiple communication channels and written in natural language, making their analysis difficult, as they are fragmented and with considerable noise. We present a 3D visualization approach mapping topics found across a project’s socio-technical artifacts onto vegetation-covered islands, where the individual artifacts are depicted as trees of various types. The topic islands rise out of the ocean as they become discussed, to sink again when they are no longer so. We built a prototype implementing the entire visualization pipeline, from data mining to interactive rendering, leveraging machine learning techniques to cluster the artifacts and extract their topics. We present, through several case studies, the insights that our approach elicits about discussions of development topics throughout a project’s history. The user study we conducted (N = 34) further strengthens our conclusions about its suitability for understanding socio-technical artifacts and their evolution. Index Terms—software visualization, topic modeling, developer conversations, socio-technical artifacts, GitHub

I. I NTRODUCTION GitHub is the dominant platform for collaborative opensource software development [1]. While the platform’s main purpose is to host Git repositories, it offers a plethora of additional features for managing software projects. Issues, Pull Requests (PRs), and Discussions are commonly used non-code artifacts documenting the project’s evolution [2], [3]. Issues act as both bug reports and feature requests. PRs contain code changes and their description, waiting to be reviewed, improved, and possibly merged into the project. Finally, Discussions are free-form conversations, brainstormings, or Q&As, similar to those found on internet forums. All of these artifacts contribute to the project’s documentation [4], [5] and chronicle the development of its features and community participation. As the nature of these artifacts encompasses both social and technical aspects, they are termed as Socio-Technical Artifacts [6] (STAs) in this paper. STAs are a key instrument in understanding software evolution, since they describe desiderata, major problems, rationale of changes, and plans for the future of the project [7]. However, these artifacts are by their nature unstructured, written in a natural language (therefore with inherent ambiguities and noise), and dispersed over multiple communication channels.

To further complicate matters, these communication channels are intended for different stakeholders. Typically, Discussions are intended for the whole community, PRs mainly for developers, and Issues have different uses in different projects, with a mix of developers and tech-oriented users alike [8], [9]. The topics manifesting in each channel also change over time. For example, what has been the focus of discourse among core developers at the beginning of the project might no longer be relevant. The challenge is in extracting useful insights from these information and documentation sources. We present an approach to visualize GitHub STAs. We analyze their evolution over time at different granularities. Using machine learning (ML) techniques, including a large language model (LLM), we cluster STAs and identify conversation topics. Each topic is visualized as an island, covered with trees representing STAs revolving around the topic. The liveliness of each topic is depicted by the terrain (i.e., landmass) of the island. The more active a topic, the higher the island’s terrain. For example, Figure 1 shows the topics found in the repository of Lume, a static website generator. Islands rise out of the ocean if their corresponding topic is active, or sink back when the topic is not discussed anymore. Our approach aids in understanding the project’s development across multiple areas of interest, revealing problematic subjects and the conversation hotspots of its community. We developed R ITGARD, a tool for the data mining, processing, and visualization of STAs. In three case studies and a user study with 34 participants, we illustrate the advantages of this approach. We show how the visualization reflects the projects’ characteristics, through the shape of the landscape and its vegetation, in an intuitive and engaging way. II. M INING AND V ISUALIZING G IT H UB STA S Socio-Technical Artifact is a term used in design science, connected to information systems and software engineering [6], [10], [11], [12]. STAs are objects with a strong interplay of social and technical aspects in their usage and design [13]. For example, Storey et al. developed a “Who-WhatHow” [14] framework to assess how software engineering research encompasses human and social aspects. This research field naturally includes collaborative development platforms as generators and repositories of STAs. For instance, Bouraffa et al. studied the social effects of code suggestions in PRs and their impact on the non-code, social-oriented artifacts [15].

A

B

Issues Pull Requests Discussions

Fig. 1. Lume in R ITGARD. Trees represent Issues, PRs, and Discussions. Islands, such as (A, B), group them by topic.

Although GitHub STAs are studied and recognized as documentation sources [4], to the best of our knowledge, a holistic visual representation of these artifacts and the emerging landscape they generate is still missing. In particular, there is no existing representation effectively capturing the time dimension, key in evolutionary analysis. In Figure 2, we show an overview of our approach. STAs are mined from GitHub through its Representational State Transfer (REST)1 and Graph Query Language (GraphQL)2 Application Programming Interfaces (APIs). Then, their textual contents are pre-processed, run through an embedding model to generate high-dimensional embeddings, and clustered by semantic similarity. Afterward, we use an LLM to generate a short descriptive label for each cluster—the topic island’s name. Finally, we use the Godot3 game engine to render the topics and STAs as islands and trees, respectively.

In the data processing stage (Figure 2, Data processing), we model the topics using the BERTopic library by Grootendorst [16]. For each pre-processed STA, a text embedding is computed using the Qwen3-Embedding-8B model [17]. Since we need to identify STAs concerning the same topic, we choose the best-performing model from the multilingual Massive Text Embedding Benchmark (MTEB) leaderboard [18] in the semantic textual similarity task as of October 2025.

A. Pre-Processing and Topic Modeling

As a final processing step, we compose a prompt for each cluster asking for a topic name. The prompt follows a predefined template (included in the replication package) and contains the titles of the most representative documents (at least four of them), keywords of the topic, keywords of the GitHub repository set by its maintainers, and two examples of expected outputs (i.e., for few-shot learning). This prompt is then sent to an LLM, and the resulting topic name is associated with the cluster to be shown in the visualization. We selected OpenAI’s GPT-OSS-B120 [21] model for the task, since it was the latest and largest general-purpose LLM available on our research infrastructure.

In the data pre-processing steps, we convert the mined STAs into plain text. First, the title, body, and comments of each STA are concatenated. Then, we remove links, figures, tables, code snippets, and Markdown syntax. Finally, the text is prepended with the STA’s labels in square brackets. Optionally, some parts (e.g., comments) may be omitted so that the final textual representation fits the size of the embedding model’s context. 1 https://docs.github.com/en/rest?apiVersion=2022-11-28 2 https://docs.github.com/en/graphql 3 https://godotengine.org

We reduce the dimensionality of the high-dimensional embeddings using the Uniform Manifold Approximation and Projection for Dimension Reduction (UMAP) algorithm [19]. The reduced embeddings are then clustered using HDBSCAN* [20]. For each cluster, we compute a bag-of-words representation of word frequency within the included STAs. Then, we remove stop words and pick the most representative keywords with a class-based TF-IDF algorithm [16].

Data mining

Data processing LLM

Rendering

GitHub repository

Keywords

REST & GraphQL APIs

CountVectorizer & c-TF-IDF

Interpolation & Gaussian blur

Clusters

Island polygons

UMAP & HDBSCAN

Triangulation & removal of overlaps

Issues, PRs, Discussions

High-dimensional embeddings

UMAP & force adjustment

Heightmaps

Topic names

2D positions

Islands

Trees

Voxel buffers

Topic point clouds

Tree parameters

Embedding model

Fig. 2. Overview of the visualization approach, as implemented in R ITGARD.

B. Turning GitHub STAs into Verdant Islands To visualize the STAs and their topics, we build upon the island metaphor. Each topic is represented as an island and each STA as a tree on that island. We chose this mapping because it makes the abstract STA data tangible and familiar, since people intuitively understand landscapes and vegetation. The spatial proximity of trees then naturally reflects similarity between the STAs. We believe the choice of this metaphor also makes the visualization more playful and engaging. Voxel-based glyphs for STAs: The visualization represents STAs as 3D tree glyphs (Figure 3). Issues are rendered as trees with a conical top, PRs as ball-top trees, and Discussions have a cubical treetop. To improve recognizability, the STA type is also encoded in the treetop color. STAs that have been closed, merged or answered on GitHub can be optionally rendered as stubs, showing that they are no longer considered relevant.

Issues

Pull requests

Discussions

The Third Dimension: We use the third dimension— island height—to represent STA activity. Specifically, the more comments an STA has, the higher the ground upon which its tree stands. To show how the topic gains or loses traction in time, we calculate island terrains in a sliding window of configurable length. For example, a sliding window of one year means that the terrain shows all STAs with comments not older than one year from the visualized time. The consequent visual effect (Figure 4) is that topic islands rise out of the ocean when they are discussed, and sink back in when they are no longer active within the sliding window, thus showing the evolution of the topic’s discourse.

A

B

Closed STAs

Fig. 3. Trees represent types of GitHub socio-technical artifacts. Stubs depict artifacts which have been closed.

The two-dimensional location of each tree within the world is obtained through a UMAP reduction of its text embedding (Section II-A). Therefore, the distance between trees reflects the semantic similarity of their corresponding STAs. However, tree positions may be slightly adjusted to prevent overlapping. Islands for topics: Each island is generated from a Delaunay triangulation of the positions of its trees, followed by the removal of triangles that are overly long or overlap with other islands, resulting in one or more polygons. These polygons are used to interpolate vertical positions of the trees, converted to a heightmap, and blurred, creating a smooth terrain.

2015/01–2016/01

A

2016/01–2017/01 Fig. 4. Islands rise (top) and sink (bottom) with the activity of topics in a one-year sliding window (A). Rocks represent outlier STAs (B). Terrain is shown without trees for clarity.

The shape of each island is visible just below the ocean’s surface even if the topic has no active STAs within the selected sliding window. This informs the user about the proportion of the currently selected timeframe with respect to the project’s entire history. Consequently, if the entire evolution of the project is visualized, no submerged islands are rendered. For example, Figure 4 A shows a topic that, while of interest until the end of 2015, was never discussed in 2016. The island color is selected randomly from a palette of eight soft colors, excluding tints of green and blue, reserved for the background (i.e., ocean) and the trees, to improve contrast. Not all STAs belong to a topic. HDBSCAN* may classify some artifacts as outliers. To differentiate them from the topic islands, the outlier trees stand on rocks in the middle of the ocean (Figure 4 B ). Outlier rocks have a voxel-based aesthetic to differentiate them from the low-poly style of the islands. The t(r)opical islands map: The 2D area of the visualization map can optionally be proportional to either the size of the project’s codebase or the “editing” activity on its source code across its history. Specifically, it can correspond to one of the following metrics: The repository’s size in kilobytes, its file count, or the number of thousands of lines of code changed across the repository’s lifetime (changed LoC). We compute changed LoC as the cumulative sum of the size of the patch resulting from each Git commit in the history of the repository’s default branch. This allows for visual comparison of multiple projects with different sizes and histories and is complemented by the number of STAs generated by its community. The changed LoC metric is also the most comparable to the number of STAs used to generate the island meshes, since both take the repository’s complete history into account. As a result, the ratio of oceans to island-covered areas represents the repository’s ratio of code-related to STA-related activities. Throughout R ITGARD’s evaluation (Section III and IV), we relied on the changed LoC metric, as it produced the most informative results. We manually looked for outliers in changed LoC (e.g., large automatic refactorings, style changes resulting in an artificially large number of modified lines) and we found none, ensuring the accuracy of the representation. The final visualization is interactive. The user can zoom, pan, and tilt the camera. Each island and tree also gets highlighted with an outline when hovered over with a mouse. The title of the hovered STA (or topic for islands) is shown in a status bar at the top of the view along with basic statistics. When a tree is clicked while holding Control/Cmd, the corresponding STA is opened in the user’s default web browser for further exploration. The visualization can be moved forward or backward in time and the length of the sliding window is configurable. III. C ASE S TUDIES Leveraging R ITGARD, we present several case studies to illustrate the exploration, visualization, and insights elicited by the STA evolutionary analysis of a project. To illustrate our approach, we analyze JetUML, Lume, and Git for Windows.

Table I shows descriptive statistics of the studied projects. We select Lume and JetUML for the familiarity of some of the authors with them, also as users, and Git for Windows based on the currently supported upper bound on the number of manageable STAs in R ITGARD’s pipeline. TABLE I D ESCRIPTIVE STATISTICS OF THE ANALYZED PROJECTS . Metric Created on Mined on History Files Changed LoC Issues PRs Discussions STAs Stars

Git for Win. 2014-08-22 2025-10-21 11+ years 3,540 31,029,161 4,564 843 226 5,633 8.9k

JetUML 2015-01-07 2025-10-21 11+ years 420 853,677 420 151 4 575 0.7k

Lume 2020-09-07 2025-10-21 5+ years 562 630,481 453 233 93 779 2.2k

DotVVM 2014-11-05 2026-02-10 11+ years 2,560 12,419,651 874 1,114 0 1,988 0.8k

A. Lume: STAs in Static Project Snapshots Lume4 is a static website generator. We analyze a snapshot of the project with a sliding window that takes into account all of its history (shown in Figure 1). At a glance, Issues are the dominant artifact type, followed by PRs and then Discussions. Considering map size and the area occupied by the islands, the project has more STAs than code (w.r.t. the number of changed LoC). Both observations are corroborated by Table I. Issues are responsible for the tallest peaks of the islands, with the highest having 25 comments. They are distributed across a wide range of topics, roughly corresponding to the project’s features, such as support for Markdown, JSX,5 URL handling, localization, and Lume’s command-line interface. Most PRs (80.0%) are concentrated on a single island (Figure 1 A ), depicting “Plugin documentation and refactor” (67 PRs, 13 Issues). Most are minor fixes and maintenance activities related to Lume’s documentation and plugins. Much like Issues, the rest of the PRs are spread over Lume’s features, with five major islands counting at least 10 PRs each. Discussions are the least common artifact type in Lume, which is apparent from their rarity in the visualization. They are typically not very complex, judging by their highest peak, representing just five comments as the longest Discussion. The most active topics regard visual rendering improvements of URLs and the internals of Lume’s rendering and pagination. Overall, Lume has a lively community and a good mixture of all three STA types, suggesting that the GitHub repository is a virtual place where developers and users cross paths to discuss technical features and desiderata for the project. This mix is especially visible on the secluded island (Figure 1 B ), where all three STAs are present, representing the “Responsive image plugin” topic and, more in general, image handling. Single snapshot STA visual analysis allows identifying how the artifacts generated in different communication channels are distributed by topic, making it easy to find the most active conversations and topics. 4 https://github.com/lumeland/lume 5 Support for HTML tags in JavaScript commonly used by web frameworks.

2016/08

2024/11

Fig. 5. Evolution of JetUML from the perspective of its GitHub STAs. Each arrow represents the passage of 20 months and activity since the previous snapshot. Together, they show 10 years of JetUML’s development, since 2015 till the end of 2024.

B. JetUML: STAs and Project Evolution We chose JetUML,6 a UML (Unified Modeling Language) drawing tool written in Java, as a case study to analyze the evolution of the STAs during the evolution of the project’s codebase. We analyze different snapshots of JetUML, taken twenty months apart from each other (Figure 5). The first snapshot, from August 2016, shows that in the 20 months of the project’s existence it has been very active. Issue-based development seems the dominant model for the project, as indicated by the large number of Issues and PRs. The largest peak is an issue regarding the addition of a copypaste feature. The largest landmass belongs to the “JavaFX UI and Diagrams” topic, arguably one at the core of the UML sketching tool. Upon closer inspection though, leveraging the hover tooltips with titles of individual PR, most of the PRs are created mainly to do minor refactorings of the codebase and to fix bugs and typos. Issues are quite broadly spread and discussed, but they seldom result in PRs, which are probably reserved for bringing in external contributions from forked repositories not discussed in any specific issue (as suggested by the separate island with almost exclusively PRs). Summarizing the trend visible in the following snapshots (the four smaller maps in Figure 5), the overall liveliness of the project sees a decline in community participation and involvement. Entire landmasses and their corresponding topics disappear completely under the ocean, and new topics are very 6 https://github.com/prmr/JetUML

few and limited in space. Fewer issues are created, while PR activity remains low but at a steady rate for the years to come. The number of active STAs never reaches again the high mark set in the first snapshot. However, at the end of 2024, the project is still actively developed, with multiple PRs discussed. Our observations about the project’s evolution suggest that either the project’s initial excitement died down except for its core audience, or the initial onslaught of issues got resolved after the first snapshot. Notable is also the complete absence of GitHub Discussions everywhere except for the last snapshot, a clearly late and ineffective attempt to use the new GitHub feature (publicly released in December 2020 [3]). Another reflection elicited by the partially disappearing landmass of the central island is that the underwater parts of the islands also communicate information about the activity of topics being discussed in the past or to be discussed in the future. We experimented with keeping stubs of the closed STAs visible underwater, to differentiate between past and future, but the visualization was too cluttered to be informative. Future work in this regard is needed to understand how to represent closed STAs to be more revealing at a glance, without animations, and with minimal clutter. STA evolutionary analysis allows comparing the activity of different artifacts in different phases of a project. More importantly, it shows by contrast how the topics evolve across time, giving insights on the project’s core aspects, attention and discussion elements important to the users.

B

A C

Fig. 6. The main landmass of the Git for Windows visualization. Peaks of most discussed STAs (A, B). Secluded topic of a dependency’s updates (C).

Fig. 7. The mostly oceanic Git for Windows map with the main landmass of topics surrounded by smaller, scattered islands.

C. Git for Windows: STAs in the Large To test the efficacy of our approach on a large project with a long history, we analyze Git for Windows7 (GfW), an actively developed Windows-specific fork of Git. This project differs from the other case studies in the size of its codebase, with more than 3.5k files in the current version, in the length of its evolution, with more than 31M changed LoC, and in the number of STAs generated in the 11 years of history: one order of magnitude larger than the previously analyzed projects. In Figure 6, we show the main landmass of the topic islands of GfW and its STA trees, as of October 2025. Being a fork has a notable consequence on the visualization. There is a lot more ocean than there is landmass (Figure 7), because issues related to the upstream Git codebase are discussed over a mailing list rather than a GitHub repository.

The most noticeable features of the visualization’s “mainland” are the tall peaks. These correspond to the most active STAs across the lifetime of GfW, marking some of its more significant events. For instance, A is an Issue about noninclusive naming,8 such as using master as the name of the default branch. This is a very polarizing topic with a vocal part of the community involved in the request for changes. Another example is about support for the ARM649 architecture B , a technical improvement requiring careful consideration and planning, but also potentially bringing new users to the project. As outliers, both of these issues have over 150 comments, so to prevent the peaks from being overly tall, the heights of the islands are normalized to a configurable maximum. The smaller islands away from the mainland represent selfcontained topics. For example, the topic “PCRE2 version update” C concerns updates of GfW’s dependency handling regular expressions. There is also the “Git for Windows releases” topic, which contains most of the project’s GitHub Discussions. For every release, a new Discussion is automatically opened and users can comment on it (e.g., report problems with the specific release). The visualization shows that some topics are persistently active over the project’s history. These topics tend to be related to Windows-specific behavior (e.g., case-insensitive paths) and advanced features (e.g., submodule recursion). Other topics, such as “Missing .mo localization”, re-emerge repeatedly with long periods of silence in between. 8 https://github.com/git-for-windows/git/issues/2674

7 https://github.com/git-for-windows/git

9 https://github.com/git-for-windows/git/issues/3107

E A A

B

C 2024-12-17

2024-12-22

B

A

D 2024-12-23

2025-01-02

Fig. 8. Example of interconnected STAs (four snapshots, one-month sliding window). Following a bug report (A) regarding the credential-cache there is a first attempt at an unmerged PR (B), followed by a successful one (C). Discussion (D) appears one week after, about a related topic. A thematically close yet unrelated STA (E) about “Git authentication issues” is already closed.

Despite the size of the GfW project, the visualization tells small developer stories as well. Figure 8 depicts what occurred in December 2024, regarding “Git Credential Manager updates”. First, Issue A 10 was created to report a problem with the credential-cache. Several days later, the Issue was addressed by PR B . This PR has been closed without being merged and was then superseded by PR C . This later PR was merged successfully without much discussion, as shown in the visualization. Ten days later, the Issue and both PRs are closed. However, a Discussion D 11 is created, asking about an error with a component related to credential-cache. Finally, clustering limitations become more pronounced in large repositories, such as GfW, where many topics “fight for space on the map”. As a result, not all close STAs are related. For instance, Figure 8 shows an Issue E ,12 unrelated to credential-cache. It is, in fact, part of a different topic (“Git authentication issues”) which is, however, thematically close to credential-cache, explaining its positioning. The case of GfW shows how the visualization provides insights in four directions. From a high-level overview to a detailed analysis of short time spans, it provides information about the STAs, with insights into the development and management processes that generated them. Another axis goes from a static to an evolutionary perspective. Static visualizations show conversation hotspots as peaks, while analyses taking into account the temporal dimension show long-lived topics, recurring ones, and short single-concern stories. 10 https://github.com/git-for-windows/git/issues/5314 11 https://github.com/git-for-windows/git/discussions/5338 12 https://github.com/git-for-windows/git/issues/5306

Considering that the Discussion D from Figure 8 remains unanswered (as of October 2025) the visualization can also be useful to pinpoint potentially overlooked STAs. Finally, this case study shows that, although improvements to the clustering step would benefit the visualization, even embedding only the STA titles and labels is sufficient to obtain a useful clustering to extract meaningful insights in real world scenarios. STA analysis scales to large long-lived repositories and provides insights in the management and development processes, with emphasis on the participation and engagement levels of different communities. IV. U SER S TUDY We conducted a user study with 34 participants (31M/3F, age 21–43, mean M = 26.53, SD = 4.49) recruited through personal contacts and snowball sampling [22], including CS students, software engineers, project managers, and academics. Sessions were conducted in-person (30) and remotely (4). The study’s main goal was to assess the interpretability of the visualization and, thus, the viability of our approach. The study design mirrored the case studies in Section III, focusing on snapshot analysis, evolutionary analysis, and large-scale visualization in three sets of tasks. The Lume and JetUML projects were used for the first two sets. However, we used DotVVM 13 (Table I), a smaller example compared to Git for Windows, for the third set of tasks to mitigate our concern that participants would focus on the prototype’s performance limitations rather than the visualization approach itself. 13 https://github.com/riganti/dotvvm

Distinguishing individual STAs

2

1

8

Distinguishing STA types

7

Recognizing Open vs Closed STAs Distinguishing STA topics

3

1

Exploring topic evolution

1

Finding recurring problem. topics

8

6

2

Seeing activity in a time window

9 8

4

Comparing topics by size

High effort

4

5

1

Comparing STAs by activity

Very high effort

11

3

14

6 2

18

4 2

5

Somewhat high effort

Somewhat low effort

2

16

Core aspects can be explored

2

15

Learning man. & dev. processes

2

12

Readable for small projects

3

10

Readable for large projects

3

Metaphor is suitable

3

12

11

12

8

Clustering is meaningful

8

13

4

23

12 Low effort

5 1 4

11 13

Disagree

15

11

8

4

6

7

14

17

9

14

5

3

18

Visualization is engaging Strongly disagree

21

13

9 Somewhat disagree

25 Somewhat agree

Strongly agree

Agree

5

Very low effort

Fig. 9. Reported effort for activities in R ITGARD.

We formulated five hypotheses: H1 the topic islands metaphor provides a readable representation of GitHub conversations (STAs) and their topics; H2 the visualization enables users to derive insights about repository activity and topics; H3 the playful visual design does not negatively affect readability; H4 the visualization is perceived as useful for software engineering tasks; and H5 the visualization remains readable for projects of different sizes. After providing consent for the study, participants completed an initial questionnaire about their demographics, background, current roles, and GitHub experience. We presented the use of R ITGARD in a video tutorial. The three sets of tasks (10 tasks in total) were designed to explore and use all of R ITGARD’s features, while thinking aloud. The correct answers for six analytical questions were determined by manually inspecting the repositories. Four remaining questions were exploratory and open-ended. The participants then completed a post-task questionnaire evaluating the visualization. The average session duration was 56:42 minutes (SD = 19:25 min). Interpretability and Insights: Basic activities such as distinguishing STAs or comparing topics by size were considered Somewhat low effort or easier by most participants (> 80%, Figure 9). Other questionnaire responses confirm this observation: Interpretation effort was very low (identification tasks M = 1.89, comparison tasks M = 2.09 on a six-point Likert scale), both significantly below the midpoint (Wilcoxon p < 0.001, effect size r ≈ 0.85–0.88), indicating that users could easily read the visualization, supporting H1. Activities involving evolutionary analysis were perceived as more demanding because participants often had to repeatedly step through the project’s history. Several participants suggested that a timeline could improve navigation, for example: “I feel like a timeline in the UI would help me immensely with understanding what timeframe I’m looking at and how quickly I’m moving through it. Without anything like that, I felt lost in what the current step and sliding window mean.” (P18)

Fig. 10. Agreement with statements about R ITGARD.

Consistent with this observation, participants reported moderate effort when extracting insights (M = 3.27, Wilcoxon p = 0.031, r = 0.32). Despite this perceived effort, analytical questions were answered with high accuracy (mean: 89%, range: 76–97%), significantly above chance (binomial tests p ≤ 0.002). This indicates that users were able to derive correct insights from the visualization, supporting H2. Playfulness: Participants generally evaluated the island metaphor positively, as shown in Figure 10. 22 respondents (65%) agreed that the visualization balances playfulness and readability well, while only one preferred a more abstract visualization. One participant summarized this enthusiasm: “Give me a diorama, a hologram on my table, so that I always see how the project is doing.” (P16) Questionnaire responses support this perception: The playful design was rated very positively (M = 5.52, Wilcoxon p < 0.001). Moreover, perceived playfulness did not correlate with interpretation effort (Spearman ρ = −0.25, p = 0.16), indicating that the playful design does not negatively affect readability and supporting H3. Usefulness: Participants considered R ITGARD particularly useful for onboarding newcomers, monitoring development progress, project management, and maintenance (Figure 11). It was also viewed as somewhat useful for evaluating potential dependencies, although several noted that repository statistics might suffice for this task. Conversely, the visualization was generally considered unsuitable for troubleshooting, though some suggested that full-text search could improve this scenario. Usefulness ratings were clearly above neutral (M = 4.11, Wilcoxon p < 0.001, r = 0.76), supporting H4. Dependency decisions

1

For newcomers

1

Throughout development Management

5

2 1

4 4

8 13

6

3

10

13

6

12

6

13

Troubleshooting Disagree

15 10

2

Maintenance

Strongly disagree

5

4

10

1

Somewhat disagree

9 Somewhat agree

16

3

10

9

Agree

Strongly agree

Fig. 11. Perceived suitability of R ITGARD for different use cases.

2

The case studies and user study demonstrate that our visualization approach, as implemented in the R ITGARD prototype tool, allows for a comprehensive analysis of a software project from the perspective of its GitHub STAs. It depicts the project’s main features, problematic areas, and even its team’s development habits. STA analysis with R ITGARD is a step towards a better understanding of how the developer and user attention shifts during the project’s lifetime, across different features and concerns. R ITGARD achieves this while also balancing its readability and playfulness.

Garbage in, garbage out: The resulting visualization is only as good as the clustering of the STAs. Bad clustering with interleaved and scattered points results in a landscape that is difficult to read. To ensure the best clustering possible, we preprocessed the STAs removing elements that would add noise with wrong “artificial” similarities (e.g., links, code snippets, markdown syntax); we embedded the artifacts using the best available embedding model for semantic textual similarity; and we devised an island generation procedure to produce smooth surfaces and reduce visual anomalies and artifacts. One size does not fit all: Although all visualizations in the case studies and the user study rely on the number of changed LoC as a measure for the size of the map, this is not suitable for all projects. Notably, there are GitHub repositories serving exclusively as issue or proposal trackers, containing very little code compared to an enormous number of STAs. For this kind of projects, the map metric can be turned off altogether and the visualization’s size is automatically set to minimize tree overlaps. With a growing size of the project, the trees also become hard to distinguish, as participants of the user study pointed out. Therefore, for large projects, aggregation should be performed on multiple levels, not just the islands. Verdant islands of healthy discourse: The islands can be used to gauge the “health” of a topic. If an island is covered in trees, it may be an outstanding problematic area. If it also has a mixture of STA types, it might mean that the topic is important for users and developers alike. However, this metaphor has limitations. For example, it is not clear whether an island is submerged because it is either no longer active or not active yet. Similarly, a tree stub indicates that the STA is closed but not the reason for its closure. For example, PRs can be either closed and successfully merged or closed and discarded. Consequently, a submerged or barren island does not imply that the topic is “unhealthy”, but only that it is inactive or its STAs are closed, respectively. We plan to build upon the topic islands metaphor in a future extension to better represent not just the liveliness of a topic but also its health status.

A. Lessons Learned

B. Limitations and Future Work

We present the main reflections about our approach that emerged during its prototype implementation, and our validation through the user study. The target audience is diverse: The visualization can be useful for many stakeholders. Software developers can use it to “chart” the topics their community is interested in, look for problems they encounter most frequently, or to find STAs that might otherwise be overlooked. For newcomers to a project, a clear high-level visualization of the STAs can impact their decision to rely on it as a dependency. The project’s core strengths, pain points, and a measurement of its development activity emerge as clear aspects from such analysis. The approach can also be beneficial for existing users in bringing them up to speed with the project’s latest development activity, with an appropriate use of the sliding window. Finally, for project managers, the visualization can serve as an intuitive overview of project status and development velocity.

Although the approach has clear advantages and use cases, there remain some conceptual and technical limitations. The end result depends on the clustering of the STAs, which in turn depends on the embedding model, which may be computationally demanding. To partially mitigate this, we decoupled the data miner, the processing steps, and the visualizer, and let them communicate through a shared exchange format in JavaScript Object Notation (JSON). This allows offloading the text embedding step to a separate machine with a GPU suitable for the chosen model. The hardware requirements can also be relaxed by selecting a less demanding model, which would make the approach accessible for a broader range of users at the cost of less accurate clustering and topic names. The pre-processing is also a limiting factor in the quality of the clustering. For example, Issue and PR templates, commonly used to standardize the format of bug reports and other common STAs, are currently not taken into consideration, thus

Scalability: Participants generally found the clustering meaningful. Readability was rated high for smaller repositories such as Lume and JetUML but decreased for larger projects like DotVVM because individual trees became smaller and harder to distinguish. Some participants also attributed this to slower interaction response times in the prototype. These observations are reflected in the questionnaire responses: Readability ratings were significantly above neutral for both small (M = 5.41, Wilcoxon p < 0.001, r = 0.89) and large repositories (M = 3.88, Wilcoxon p = 0.028, r = 0.33). However, a paired Wilcoxon test showed that large projects were rated significantly harder to read (p < 0.001), suggesting that while the visualization remains usable at larger scales, its readability decreases as project size increases, partially supporting H5. Finally, when asked to compare R ITGARD with GitHub’s user interface, 23 participants (68%) considered the two complementary, emphasizing that R ITGARD provides a high-level overview of project activity that traditional interfaces lack. Overall, the results provide evidence supporting the interpretability, perceived usefulness, and suitability of the topic islands metaphor, supported and enriched by an engaging visualization. Due to space constraints, the results presented here are condensed; the full tasks, questionnaire, responses, and analyses are included in the replication package. V. D ISCUSSION

potentially negatively affecting the text embeddings. This may result in large islands consisting mostly of STAs of one “type” connected to a specific template. Some STAs can also be too large to be embedded at once, due to limited context size of the embedding model or its demands on GPU memory. For example, in the GfW case study, we excluded STA bodies and comments from the embeddings, due to memory constraints. This shows that it is still possible to obtain meaningful insights with our visualization despite the trade-offs with the embeddings. In the future, we plan to conduct experiments by appropriately chunking the STAs to keep more context and re-combining them through the embeddings. The data miner has its specific limitations. For GitHub Discussions, it retrieves only the main comments without their inner replies (i.e., the threads), which might lead to the underrepresentation and mis-classification of Discussions with respect to other STAs. Moreover, for all closed STAs, we mine only the last timestamp of their closure. Therefore, the visualization ignores STA re-openings. Our design choices to make the visualization engaging and playful resulted in a visually appealing, readable, and informative visualization, but they also come with trade-offs. For example, when the number and density of the visualized STAs is high, the resulting tree foliage can obstruct the view of the underlying terrain. We partially mitigate this shortcoming by including a UI option to hide the trees when the individual STAs are not of primary importance. R ITGARD is a proof-of-concept research prototype. Rendering performance was not prioritized throughout its development. Consequently, it is limited in the number of STAs it can handle. Concrete upper bounds depend on the available hardware. In the future, if performance becomes a focus, we can leverage typical rendering trade-off “tricks” to scale up the number of analyzable projects, also aiming towards real-time animation of the evolution. So far, we focused on experimenting with the core visualization. In the future, the most significant improvement will be a proper timeline, intuitive sliding window controls, and better label placement as suggested by the study’s participants. The user study may not be fully representative, as many of the participants were undergraduate or graduate students (41%), few had experience with project management (6%), and most were male (91%). The tasks featured in the user study were chosen to evaluate the interpretability of the visualization; therefore, the reported usefulness is based on the participants’ perceptions. Similarly, although the participants were asked to compare R ITGARD to GitHub’s native UI, these observations are not a substitute for a direct task-based comparison. Despite the phrasing of the questions and explicit instruction to focus on the visualization, participants sometimes did not distinguish between the visualization and the prototype. This skews the results toward negative answers due to assessing the prototype’s performance and user experience, rather than the visualization itself.

None of the limitations discussed here are unexpected, and they provide avenues for future work. R ITGARD can be extended to mine other platforms (e.g., GitLab, Bitbucket, Gitea). The tree glyphs can be made more expressive to depict additional information about the represented STAs. Lastly, the visualization could take the source code of the system into account, connecting the STAs to the related source code actions, opening new possibilities to expand the metaphor. VI. R ELATED W ORK GitHub and its STAs are an active research subject. Gousios et al. conducted a foundational study on millions of PRs to learn how developers use the pull-based development model [23]. Hata et al. researched GitHub Discussions specifically, showing their worth for spawning new Issues and PRs [3]. Siddiq et al. used machine learning (i.e., a BERTbased model) to automatically classify Issues and assign them labels [24]. Venigalla et al. showed the documentation potential of various STAs and studied their topics [5]. Barbosa et al. focused on the social dimension of GitHub STAs, conducting a study into how factors such as communication dynamics and team size impact code quality in PRs [25]. Nevertheless, GitHub is not the only platform of interest. For example, Barua et al. used latent Dirichlet allocation to extract topics and trends form StackOverflow posts [26]. Parra et al. compared the Gitter and Slack messaging platforms, produced a manually annotated dataset of Gitter messages, and evaluated several ML techniques for their automatic labeling. Warrick et al. studied developer ecosystems (Python, Go, and Node.js) as socio-technical systems, mined their mailing lists, producing the OCEAN dataset [27]. All of these examples showcase the diversity of research into software-related sociotechnical artifacts. Topic modeling lies at the core of our visualization approach. It is a set of techniques used to extract latent topics from a body of text in a natural language. These techniques have been developed for decades and in recent years benefited from the use of neural networks. We refer to two surveys by Wu et al. [28] and Zhao et al. [29], respectively, for an overview of topic modeling and its use of neural networks. Landscapes and islands have been used to aid in program comprehension before. Kuhn et al. used latent semantic indexing and multidimensional scaling to produce a 2D thematic map of a codebase [30]. Atzberger et al. took a similar approach with Latent Dirichlet Allocation in their Software Forest, where each software entity becomes a tree [31]. Misiak et al. visualized packages and their Java classes as islands and buildings both in 3D and virtual reality [32]. While the aforementioned works use visual metaphor similar to ours, they visualize code artifacts, not STAs. Besides the island metaphor, other approaches include that of Fiechter et al., representing GitHub Issues through their issue tales, offering multiple 2D issue visualizations [33]. However, their approach focuses mostly on the lifecycle of issues, whereas R ITGARD’s is about providing an evolutionary overview of the whole repository with multiple types of STAs.

Our design choices to make the visualization more engaging, such as representing the STAs with tree glyphs, make our approach related to gamification, which is a lively topic in software engineering in both theory and practice. For example, Dal Sasso et al., investigated the usage of game elements and presented a conceptual framework with guidelines on which elements to consider for software engineering tasks [34]. Another example is a gamified visualization designed by Heidrich et al. to foster program comprehension, taking into account developer requirements and getting positive feedback from their players [35]. Gamification is connected to the concept of playfulness in adults—inclination towards creativity, curiosity, sense of humor, and related traits—which has been linked with the ability to meet problems with an open mind and accept failure, increasing work performance [36]. Playfulness can be a core characteristic of visualization. For example, Schindler et al. employed it when visualizing biological models, producing paper-based “physicalizations” [37]. VII. C ONCLUSION We presented a visualization approach for the analysis of socio-technical artifacts in software projects hosted on GitHub. The visualization depicts each topic extracted from the artifacts as an island in an ocean, the area of which represents the size of the code repository. As topics become more or less discussed, the islands rise in and out of the ocean. The topics themselves are extracted by converting each STA into a vector embedding, clustering these vectors, extracting keywords, and feeding them into a large language model to produce a humanreadable label for the island. We provide insights into the analyzed projects from static and evolutionary perspectives, at multiple levels of granularity. The R ITGARD prototype requires optimization and user experience improvements to be usable in practice. However, as shown by the user study, it can support software developers to learn about a project, open-source software maintainers in managing their communities, and it can also serve users of these projects to uncover problematic areas discussed by different stakeholders in different communication channels. R EPLICATION PACKAGE To allow for our work to be verified and replicated, the case study datasets, the user study materials, the participant responses, the topic naming prompt template, the source code, a tutorial video, and builds of R ITGARD are available at https://doi.org/10.6084/m9.figshare.30444704 ACKNOWLEDGEMENTS We thank the participants of our user study and our colleagues who helped refine the study before it was conducted. Computational resources were provided by the e-INFRA CZ project (ID:90254), supported by the Ministry of Education, Youth and Sports of the Czech Republic. We also gratefully acknowledge the financial support of the Swiss National Science Foundation (SNSF) through the project “FORCE” (SNF Project No. 232141).

R EFERENCES [1] GitHub, “Octoverse,” The GitHub Blog, 2024. [Online]. Available: https://octoverse.github.com/ [2] J. Tantisuwankul, Y. S. Nugroho, R. G. Kula, H. Hata, A. Rungsawang, P. Leelaprute, and K. Matsumoto, “A topological analysis of communication channels for knowledge sharing in contemporary GitHub projects,” Journal of Systems and Software, vol. 158, pp. 1–12, 2019, doi:10.1016/j.jss.2019.110416. [3] H. Hata, N. Novielli, S. Baltes, R. G. Kula, and C. Treude, “GitHub Discussions: An exploratory study of early adoption,” Empirical Software Engineering, vol. 27, no. 1, pp. 1–32, 2021, doi:10.1007/s10664-02110058-6. [4] M. Raglianti, C. Nagy, R. Minelli, B. Lin, and M. Lanza, “On the Rise of Modern Software Documentation,” in European Conference on ObjectOriented Programming (ECOOP), vol. 263. Dagstuhl, 2023, pp. 43:1– 43:24, doi:10.4230/LIPIcs.ECOOP.2023.43. [5] A. S. M. Venigalla and S. Chimalakonda, “An exploratory study of software artifacts on GitHub from the lens of documentation,” Information and Software Technology, vol. 169, pp. 1–21, 2024, doi:10.1016/j.infsof.2024.107425. [6] S. Gregor and A. R. Hevner, “Positioning and presenting design science research for maximum impact,” Management Information Systems Quarterly, vol. 37, pp. 337–355, 2013, doi:10.25300/MISQ/2013/37.2.01. [7] R. Alkadhi, T. Lata, E. Guzmany, and B. Bruegge, “Rationale in development chat messages: An exploratory study,” in International Conference on Mining Software Repositories (MSR). IEEE, 2017, pp. 436–446, doi:10.1109/MSR.2017.43. [8] GitHub, “About issues,” GitHub Docs, 2026. [Online]. Available: https://docs.github.com/en/issues/tracking-your-work-withissues/learning-about-issues/about-issues [9] ——, “About discussions,” GitHub Docs, 2026. [Online]. Available: https://docs.github.com/en/discussions/collaborating-withyour-community-using-discussions/about-discussions [10] A. Drechsler, “Designing to inform: Toward conceptualizing practitioner audiences for socio-technical artifacts in design science research in the information systems discipline,” Informing Science: The International Journal of an Emerging Transdiscipline, vol. 18, pp. 31–47, 2015, doi:10.28945/2288. [11] H. Weigand, P. Johannesson, and B. Andersson, “An artifact ontology for design science research,” Data & Knowledge Engineering, vol. 133, pp. 1–19, 2021, doi:10.1016/j.datak.2021.101878. [12] P. Runeson, E. Engström, and M.-A. Storey, “The design science paradigm as a frame for empirical software engineering,” in Contemporary Empirical Methods in Software Engineering. Springer, 2020, pp. 127–147, doi:10.1007/978-3-030-32489-6 5. [13] W. E. Bijker, T. P. Hughes, and T. J. Pinch, “General introduction,” in The Social Construction of Technological Systems: New Directions in the Sociology and History of Technology. MIT Press, 1987. [14] M.-A. Storey, N. A. Ernst, C. Williams, and E. Kalliamvakou, “The who, what, how of software engineering research: A socio-technical framework,” Empirical Software Engineering, vol. 25, no. 5, pp. 4097– 4129, 2020, doi:10.1007/s10664-020-09858-z. [15] A. Bouraffa, Y. D. Pham, and W. Maalej, “How do developers use code suggestions in pull request reviews?” in International Conference on Cooperative and Human Aspects of Software Engineering (CHASE). IEEE, 2025, pp. 227–238, doi:10.1109/CHASE66643.2025.00033. [16] M. Grootendorst, “BERTopic: Neural topic modeling with a class-based TF-IDF procedure,” pp. 1–10, 2022, doi:10.48550/arXiv.2203.05794. [17] Y. Zhang, M. Li, D. Long, X. Zhang, H. Lin, B. Yang, P. Xie, A. Yang, D. Liu, J. Lin, F. Huang, and J. Zhou, “Qwen3 embedding: Advancing text embedding and reranking through foundation models,” 2025, doi:10.48550/arXiv.2506.05176. [18] N. Muennighoff, N. Tazi, L. Magne, and N. Reimers, “MTEB: Massive text embedding benchmark,” in Conference of the European Chapter of the Association for Computational Linguistics (EACL). ACL, 2023, pp. 2014–2037, doi:10.18653/v1/2023.eacl-main.148. [Online]. Available: https://huggingface.co/spaces/mteb/leaderboard [19] L. McInnes, J. Healy, N. Saul, and L. Großberger, “UMAP: Uniform manifold approximation and projection,” Journal of Open Source Software (JOSS), vol. 3, pp. 1–2, 2018, doi:10.21105/joss.00861. [20] L. McInnes and J. Healy, “Accelerated hierarchical density based clustering,” in International Conference on Data Mining Workshops (ICDMW). IEEE, 2017, pp. 33–42, doi:10.1109/ICDMW.2017.12.

[21] OpenAI, “gpt-oss-120b & gpt-oss-20b model card,” 2025, doi:10.48550/arXiv.2508.10925. [22] L. A. Goodman, “Snowball sampling,” The Annals of Mathematical Statistics, vol. 32, pp. 148–170, 1961. [23] G. Gousios, M. Pinzger, and A. v. Deursen, “An exploratory study of the pull-based software development model,” in International Conference on Software Engineering (ICSE). ACM, 2014, pp. 345–355, doi:10.1145/2568225.2568260. [24] M. L. Siddiq and J. C. S. Santos, “BERT-based GitHub issue report classification,” in International Workshop on Natural Language-Based Software Engineering. Pittsburgh Pennsylvania: ACM, 2022, pp. 33–36, doi:10.1145/3528588.3528660. [25] C. Barbosa, A. Uchôa, D. Coutinho, W. K. G. Assunção, A. Oliveira, A. Garcia, B. Fonseca, M. Rabelo, J. E. Coelho, E. Carvalho, and H. Santos, “Beyond the code: Investigating the effects of pull request conversations on design decay,” in International Symposium on Empirical Software Engineering and Measurement (ESEM). IEEE, 2023, pp. 1–12, doi:10.1109/ESEM56168.2023.10304805. [26] A. Barua, S. W. Thomas, and A. E. Hassan, “What are developers talking about? An analysis of topics and trends in Stack Overflow,” Empirical Software Engineering, vol. 19, no. 3, pp. 619–654, 2014, doi:10.1007/s10664-012-9231-y. [27] M. Warrick, S. F. Rosenblatt, J.-G. Young, A. Casari, L. HébertDufresne, and J. Bagrow, “The OCEAN mailing list data set: Network analysis spanning mailing lists and code repositories,” in International Conference on Mining Software Repositories (MSR). ACM, 2022, pp. 338–342, doi:10.1145/3524842.3528479. [28] X. Wu, T. Nguyen, and A. T. Luu, “A survey on neural topic models: Methods, applications, and challenges,” Artificial Intelligence Review, vol. 57, pp. 1–30, 2024, doi:10.1007/s10462-023-10661-7. [29] H. Zhao, D. Phung, V. Huynh, Y. Jin, L. Du, and W. Buntine, “Topic modelling meets deep neural networks: A survey,” in International Joint Conference on Artificial Intelligence (IJCAI), vol. 5. IJCAI, 2021, pp. 4713–4720, doi:10.24963/ijcai.2021/638. [30] A. Kuhn, P. Loretan, and O. Nierstrasz, “Consistent layout for thematic software maps,” in Working Conference on Reverse Engineering. IEEE, 2008, pp. 209–218, doi:10.1109/WCRE.2008.45. [31] D. Atzberger, T. Cech, M. De La Haye, M. Söchting, W. Scheibel, D. Limberger, and J. Döllner, “Software Forest: A visualization of semantic similarities in source code using a tree metaphor,” in International Joint Conference on Computer Vision, Imaging and Computer Graphics Theory and Applications (VISIGRAPP). SCITEPRESS, 2021, pp. 112–122, doi:10.5220/0010267601120122. [32] M. Misiak, A. Schreiber, A. Fuhrmann, S. Zur, D. Seider, and L. Nafeie, “IslandViz: A tool for visualizing modular software systems in virtual reality,” in Working Conference on Software Visualization (VISSOFT). IEEE, 2018, pp. 112–116, doi:10.1109/VISSOFT.2018.00020. [33] A. Fiechter, R. Minelli, C. Nagy, and M. Lanza, “Visualizing GitHub Issues,” in Working Conference on Software Visualization (VISSOFT). IEEE, 2021, pp. 155–159, doi:10.1109/VISSOFT52517.2021.00030. [34] T. Dal Sasso, A. Mocci, M. Lanza, and E. Mastrodicasa, “How to gamify software engineering,” in International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 2017, pp. 261–271, doi:10.1109/SANER.2017.7884627. [35] D. Heidrich, R. Gökmen, A. Schreiber, and C. Bichlmeier, “Towards higher motivation to perform software comprehension tasks through gamification,” in Working Conference on Software Visualization (VISSOFT). IEEE, 2025, pp. 79–83, doi:10.1109/VISSOFT67405.2025.00018. [36] P. Guitard, F. Ferland, and É. Dutil, “Toward a better understanding of playfulness in adults,” Occupational Therapy Journal of Research (OTJR), vol. 25, pp. 9–22, 2005, doi:10.1177/153944920502500103. [37] M. Schindler, T. Korpitsch, R. G. Raidou, and H.-Y. Wu, “Nested papercrafts for anatomical and biological edutainment,” Computer Graphics Forum, vol. 41, pp. 541–553, 2022, doi:10.1111/cgf.14561.

Record · ID 660879 · SHA-256 284453b638509289
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.