Conceptio › Archive › arXiv CS
arXiv CSopen access

Ritgard: T(r)opical Islands of Socio-Technical Artifacts on GitHub

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

Ritgard: T(r)opical Islands of Socio-Technical Artifacts on GitHub Adam Štěpánek∗ , Marco Raglianti† , Jan Byška∗‡ , Barbora Kozlı́ková∗ , Michele Lanza† ∗ Visitlab, Masaryk University, Brno, Czech Republic

† REVEAL @ Software Institute — USI, Lugano, Switzerland

arXiv:2609.05278v1 [cs.SE] 4 Sep 2026

‡ VisGroup, University of Bergen, Bergen, Norway

Abstract—A software project is more than just code. Noncode artifacts often document the human processes and decisions behind source code. The rationale behind a library change, an architectural decision, a problem encountered by a user are all examples of information typically present in sociotechnical artifacts (STAs), created and persisted in channels separate from the repository itself (yet sometimes very close— e.g., GitHub Issues with GitHub repositories). These STAs are a trove of information about the project’s architecture and its evolution, containing details and insights that code alone cannot provide. Unfortunately, this information is not easily extracted and explored as STAs are frequently fragmented over different communication channels, and are written in natural language. We present R ITGARD, a tool that mines GitHub repositories for their STAs, namely Issues, Pull Requests, and Discussions, and visualizes them as 3D islands covered with trees. Each tree represents a single artifact and each island is a topic extracted from the artifacts through a combination of text embedding and text summarization. The terrain of the islands rises out of the ocean as the topic becomes active and sinks back in when it becomes stale, thus depicting the evolution of features and concerns throughout the project’s lifetime. We describe the tool’s usage and implementation, showing the numerous technical challenges behind R ITGARD’s visualization. Index Terms—software visualization, topic modeling, software repository mining, socio-technical artifacts, GitHub

I. I NTRODUCTION Code does not tell the whole story of a software project. Non-code artifacts, such as documentation, bug reports, code reviews, architectural decision records, and even e-mail and other messages describe the project’s shape and history from another perspective, which may be disconnected from the code repository. These artifacts form the documentation landscape [1] of the project, a rugged environment scattered over different communication channels (e.g., GitHub, e-mails, Slack, Discord) that are volatile and ever-changing, especially for actively developed large and long-lived projects. GitHub is the dominant collaborative development platform [2] and its built-in communication channels are the backbone of the typical documentation landscape of GitHub projects. Its Issues, Pull requests (PRs), and Discussions are used to manage and document the project’s development. Issues describe bugs and desirable features [3]. PRs are used to review changes to the project’s code [4]. Discussions offer a forum-like environment where developers and users can meet [5]. Since these communication channels bridge the gap between society and technology, we consider each Issue, PR, and Discussion to be a socio-technical artifact (STA) [6], [7].

STAs of a project depict the social discourse that surrounds the repository, its development processes, architectural decisions, and community reception. They also describe this reality over time, thus recording the discourse and the project’s evolution. This viewpoint is naturally useful in many circumstances. For example, when developers need to understand the project’s purpose and the forces that shape(d) its development. Project managers need to see the project’s current state and progress towards a future milestone. Users need to keep up-to-date with the project to know if it is still alive and without any security vulnerabilities. Therefore, having at least a high-level understanding of a project’s STAs is useful for many stakeholders. However, getting this understanding is cumbersome, since GitHub STAs are written in natural language and, for large repositories, there simply may be far too many STAs to read without a strategy and proper tool support. Recent developments in text embedding and summarization using large language models (LLM) provide a practical way to solve both issues. Embedding models have proven to be capable enough to compare semantic similarity of sentences and texts [8], [9]. LLMs can assign topics to clusters of textual documents via their summarization capabilities [10], [11]. Yet, even with these advances, there remains the issue of assembling the STAs, their topics, metadata, and history in a readable, informative, high-level, and playful representation [12], [13]. In Figure 1, we show an example visualization and R ITGARD’s user interface (UI).

A

B

Issues Pull Requests Discussions Closed STAs

Fig. 1. The golang-standards/project-layout repository in R ITGARD . Issues, PRs, and Discussions are visualized as trees. Islands group them by topic. The UI includes a status bar (A) and a configuration side panel for the visualization parameters (B).

In our prior work [14], we presented a visualization approach providing an overview of the STAs of a GitHub repository and evaluated its readability and usefulness in a user study. Our approach turns each topic into a 3D island in an ocean, and each STA into a tree. Here, we focus on the technical details of R ITGARD—the prototype implementation of our visualization design, also used in the user study. We present its data processing pipeline from the GitHub repository to the interactive 3D visualization, provide examples, and suggest directions for future development. II. R ELATED W ORK Existing tools and prior research focus on broad usage of STA types or deep analysis of individual STAs. To the best of our knowledge, there is no tool that provides a high-level overview of the three main STA types of a GitHub repository, let alone one with customized evolutionary visualizations. Zhang et al. studied the reasons for PR acceptance or rejection [15]. They found that the most important factor in the decision is the social distance between the author and the integrator, and that automated tools often replace the role of comments. Hata et al. conducted a study of GitHub Discussions, a newly added STA type back then, and discovered that it is essential to set proper guidelines for this platform, so indirectly for the creation of such STAs, and that core developers’ participation in the discourse can make a difference [16]. Hao et al. trained a model for recommending “good first issues” for new contributors [17], confirming the importance of STAs for triaging and bug fixing but also for community building and knowledge transfer. Siddiq et al. used topic modeling to assign labels to Issues, showing that these techniques are a natural fit for the NL in GitHub STAs [18]. Visualizations: Fiechter et al. introduced issue tales, a 2D visualization focusing on the issue lifecycle in views with various granularities [19]. Issue tales visualize issue metrics (e.g., size, duration) and their connections, but disregard contents and topics. Kuhn et al. used machine learning methods to produce thematic maps of source code [20]. Although, they focused only on source code, their work inspired R ITGARD. GitHub’s native UI includes a tabular view of STAs, providing useful details (e.g., titles, labels, assignees). However, it is limited to a handful of items per page, textual in nature, and thus fails at providing a holistic high-level overview. III. T( R ) OPICAL I SLANDS R ITGARD is a visualization tool implementing the concept of t(r)opical islands, enriched by a suite of scripts for data mining and processing. It handles everything, from mining the STAs from GitHub, to data pre-processing, topic modeling, layout and terrain generation, rendering, and user interaction. It provides a high-level overview of a snapshot of the project’s discourse and facilitates interactive exploration of its evolution. R ITGARD is a prototype mainly aimed at developers, but can be useful to other stakeholders, such as project managers, and even end users (e.g., prospective adopters of a library assessing non-code assets, project maturity, and recurring issues).

A. Visualization Design R ITGARD visualizes the STAs of a GitHub repository as a 3D terrain consisting of tree-covered islands in a rectangular ocean (Figure 1). Each tree represents an individual STA, with different sources mapped on different tree types. Islands group the STAs according to their prevailing topic and the distance between trees corresponds to their semantic similarity. Depicting evolution: The height of the terrain is proportional to the activity of STAs that stand on it. For example, heavily discussed Issues with hundreds of comments result in tall hills, whereas unanswered Discussions will only be small mounds of earth, almost at sea level. STA activity is measured throughout the project’s entire history or is constrained to a specific period using a configurable sliding window, which takes a certain time span (e.g., a year) leading up to the currently visualized point in time. As this window slides through time, the project’s evolution unveils itself in front of the user and islands rise out of ocean and sink back into it, showing which topics are active at any given time. Tree types: STAs are represented using three types of 3D tree glyphs (Figure 1, bottom-left corner). Trees with conical tops imply an Issue, ball-top trees are PRs, and cube-top trees are Discussions. There are also stubs—trees with no treetop— that can be toggled to stand for STAs that have already been closed (either through acceptance or rejection) and thus have (unless reopened) reached the end of their life. Topics and outliers: Islands group together STAs with a common topic. Each island is also assigned a random color from a preselected palette to further enforce the sense of separation among islands and topics. While focused communication is encouraged, nothing prevents the STA authors from covering multiple topics, especially when there are crosscutting concerns. So, in instances without a prevailing topic, the STA is classified as outlier and its tree is put on a solitary voxel-based rock in the ocean to imply it standing out. Interactive exploration: The user is free to roam the islands landscape with an orthographic camera and explore the STAs through mouse and keyboard interactions. When they encounter an interesting STA and desire its closer inspection, they can trigger an action that opens the artifact in their default web browser. They may also change the size of the sliding window (i.e., the length of the aggregated and visualized time period) and shift it backward and forward in time, thus enabling an analysis of the project’s evolution through the animation of the landscape terrain. B. User Interface R ITGARD’s data mining, data processing, and terrain generation steps are implemented as separate scripts, providing a command-line interface, configurable with arguments and options. The outputs of these scripts can then be visualized in an application with a graphical user interface with two key elements: the status bar and the side panel. The hover bar (Figure 1 A ) sits at the top of the GUI and displays the name and basic information about the STA or topic that is currently being hovered over with the mouse.

Data mining

Data processing

CLI's repo command

./model-topics.py

GitHub repository

Topic names

REST & GraphQL APIs

Large language model

Keywords, Representative STAs

Issues, PRs, Discussions

./adjust-positions.py

Rendering CLI's terrain command

Godot-based viewer

Heightmaps

Islands

Interpolation & Gaussian blur

2D positions

Island polygons

CountVectorizer & c-TF-IDF

Force-directed adjustment

Triangulation & removal of overlaps

Voxel mesher

Clusters

Low-dimensional embeddings

Topic point clouds

Voxel buffers

UMAP & HDBSCAN

UMAP

Tree generators

High-dimensional embeddings

Embedding model

Trees

Tree parameters

Fig. 2. R ITGARD’s pipeline, consisting of data mining, data processing, and rendering.

If no object is hovered over, the currently visualized point in time and basic statistics about the current view are shown. The side panel (Figure 1 B ) covers all configurable aspects of the visualization. The user selects the currently visualized codebase from a list of mined and processed datasets. They may also decide to, for example, hide all trees and focus solely on the terrain, to disable showing closed STAs as tree stubs, or to normalize the height of the terrain to specific values, to either emphasize the height differences or limit occlusion. The GUI is intentionally minimalistic, since the focus of this prototype is to evaluate the visualization design. For the interactions, it includes mouse and keyboard shortcuts, summarized in a cheatsheet available in the replication package. C. Architecture and Pipeline R ITGARD follows a three-stage pipeline (Figure 2): data mining, data processing, and rendering. The selected GitHub repository is mined, then the extracted STAs are processed, their text is embedded, clustered, and turned into terrain data, finally, the terrain is rendered in an interactive viewer. Each stage involves the execution of one or more scripts or programs. The data mining is performed using a C# console application, utilizing GitHub’s REST and GraphQL APIs. REST is used for Issues and PRs, whereas GraphQL for Discussions, since those are unavailable from the REST API. Due to GitHub’s strict rate limiting [21], this stage can be the pipeline’s bottleneck for large repositories. In the data processing stage, the STAs extracted from GitHub’s API undergo several transformations using two Python scripts and a C# console application. First, they are stripped of any links and Markdown syntax and embedded into high-dimensional vectors. The embedded text comprises the STA title, its labels, its main body, and/or all of its comments. Except for the STA title, all elements may be excluded to hasten the embedding step. The embedding model is also configurable, but by default the Qwen3-Embedding8B [22] model is chosen due to its performance in the semantic similarity task in the Massive Text Embedding Benchmark [9].

These embeddings are then reduced to low-dimensional vectors using UMAP [23] and clustered with HDBSCAN* [24]. These two algorithms were selected as part of the BERTopic framework [10] due to their ease of use and successful application in previous research involving topic modeling [25], [26]. Then, an LLM is tasked with generating a short topic name for each cluster, using a prompt with the titles of the cluster’s most representative STAs, keywords as given by cTF-IDF [10], and the repository’s GitHub keywords. UMAP is used again to reduce the embeddings to two dimensions, which eventually become the positions of the tree glyphs. First, however, they undergo a series of transformations to minimize tree overlap and constrain map size. Terrain: The third and final data processing script calculates heightmap textures for each topic-island and for each predefined sliding window length. The terrain of each island is computed using a concave Delaunay triangulation (similar to alpha shapes [27]) of all STA positions. Triangles that overlap with other islands or with STA outliers are removed, resulting in a believable island outline. Then, each position on a grid inside a remaining triangle is interpolated using values from the triangle’s corners and blurred to remove sharp edges. The interpolation and blur is done for each “step” of the visualization, which is the duration between two consecutive points in time. The length of this step is configurable and is set to 24 hours by default. While the use of this terrain generation script limits the sliding window to only a handful of predefined lengths, it severely lowers the computational requirements of the visualizer at runtime. This tradeoff allows the user to move through the project’s history at a reasonable pace. Rendering: The final stage is the rendering. Implemented in the Godot game engine, the visualizer acts as a window on the islands of STAs mined and processed before. The visualizer itself renders the landscape and handles user interaction. It is a standalone program relying on modern rendering APIs and, therefore, can utilize the full rendering potential of a dedicated GPU, which is crucial for smooth performance and scalability of R ITGARD.

Fig. 3. Examples of R ITGARD’s visualization, showing the last two years of active STAs in 4 software projects (A-D). Same image scale for comparison.

The prototype’s pipeline is not a one-size-fits-all solution. Its fragmented nature reflects the need for each stage to run on different hardware or with different permissions. For example, data mining requires GitHub access tokens, which are sensitive information. The topic modeling step may require high-end GPUs, depending on the embedding model, and thus runs on a different machine than the visualizer, whose requirements, in turn, include neither access to GitHub APIs nor special hardware. Therefore, R ITGARD’s pipeline is best ran in two or three distinct stages, so that data mining and processing runs independently of the visualizer, akin to a thin client. IV. E XAMPLES Figure 3 shows four examples of R ITGARD’s visualization. We chose the repositories based on the first author’s familiarity with them and integrating with results from SEART-GHS [28], with filters on number of STAs and stars on GitHub. They were selected to illustrate how the visualization reflects the differences between them. All four images show the same twoyear period between June 2024 and June 2026 and use the same scale for comparison. Even without the interactive pan and zoom available in the tool, these snapshots provide a few insights. For example, A is the most active of the four, judging by its size and tree count. However, A is also the youngest as there are no submerged landmasses, indicating that the project has been created in the visualized time span. Projects B and C are both commandline utilities of similar sizes (in terms of STAs), however, B has been considerably more active lately, with several active topic-islands above sea level. C and D share similar activity levels, but PRs (ball-top trees) are much more prevalent in D , indicating that its issues are likely discussed elsewhere. While A and B have adopted Discussions, there are no cube-top trees on the islands of the other projects, suggesting that either their communities are small or the developers have not enabled the feature and could consider its potential impact.

V. C ONCLUSION & F UTURE WORK The socio-technical artifacts of a software project tell its story from a perspective that its source code cannot convey. On GitHub, Issues, Pull Requests, and Discussions document the project’s lifecycle, the key decisions of its developers, and the struggles of its users. However, these insights are difficult to access due to the sheer number STAs, their use of natural language, and the lack of existing tools to process them and provide a visually rich overview. With R ITGARD, we have shown that a combination of topic modeling and 3D visualization techniques can produce a highlevel overview of a project’s STAs. R ITGARD turns these artifacts into a landscape of islands representing their topics. It captures the present state of the project and its entire evolution from the perspective of its STAs. With polish and refinement, such a tool can be a practical companion to both the engineers working on the visualized project and to its users, striving to understand key non-functional implications. We envision extending the visual metaphor to cover more STA metadata (e.g., reason for STA closure) and polishing the visualizer’s user experience. The tool’s pipeline should also be refactored into a client-server setup, so that it can be invoked with a single button press within the visualizer itself, greatly decreasing the tool’s barrier for entry of the current scripts. Finally, we want to experiment with smaller embedding models and LLMs to find the smallest models that produce good results on “affordable” hardware. Replication package: To ensure the verifiability of our work, the tool, a demonstration video, and example datasets are available at https://doi.org/10.6084/m9.figshare.32346822. Acknowledgements: Computational resources were provided by the e-INFRA CZ project (ID:90254), supported by the Ministry of Education, Youth and Sports of the Czech Republic. We also gratefully acknowledge the financial support of the Swiss National Science Foundation (SNSF) through the project “FORCE” (SNF Project No. 232141).

R EFERENCES [1] M. Raglianti, C. Nagy, R. Minelli, B. Lin, and M. Lanza, “On the Rise of Modern Software Documentation,” in European Conference on ObjectOriented Programming (ECOOP), vol. 263. Dagstuhl, 2023, pp. 43:1– 43:24, doi:10.4230/LIPIcs.ECOOP.2023.43. [2] GitHub, “Octoverse,” The GitHub Blog, 2024. [Online]. Available: https://octoverse.github.com/ [3] ——, “About issues,” GitHub Docs, 2026. [Online]. Available: https://docs.github.com/en/issues/tracking-your-work-with-issues/ learning-about-issues/about-issues [4] ——, “About pull requests,” GitHub Docs, 2026. [Online]. Available: https://docs.github.com/en/pull-requests/collaborating-with-pullrequests/proposing-changes-to-your-work-with-pull-requests/aboutpull-requests [5] ——, “About discussions,” GitHub Docs, 2026. [Online]. Available: https://docs.github.com/en/discussions/collaborating-withyour-community-using-discussions/about-discussions [6] S. Gregor and A. R. Hevner, “Positioning and presenting design science research for maximum impact,” Management Information Systems Quarterly, vol. 37, pp. 337–355, 2013, doi:10.25300/MISQ/2013/37.2.01. [7] A. Drechsler, “Designing to inform: Toward conceptualizing practitioner audiences for socio-technical artifacts in design science research in the information systems discipline,” Informing Science: The International Journal of an Emerging Transdiscipline, vol. 18, pp. 31–47, 2015, doi:10.28945/2288. [8] N. Reimers and I. Gurevych, “Sentence-BERT: Sentence embeddings using Siamese BERT-networks,” in Conference on Empirical Methods in Natural Language Processing and the International Joint Conference on Natural Language Processing (EMNLP-IJCNLP). ACL, 2019, pp. 3982–3992, doi:10.18653/v1/D19-1410. [9] N. Muennighoff, N. Tazi, L. Magne, and N. Reimers, “MTEB: Massive text embedding benchmark,” in Conference of the European Chapter of the Association for Computational Linguistics (EACL). ACL, 2023, pp. 2014–2037, doi:10.18653/v1/2023.eacl-main.148. [Online]. Available: https://huggingface.co/spaces/mteb/leaderboard [10] M. Grootendorst, “BERTopic: Neural topic modeling with a class-based TF-IDF procedure,” pp. 1–10, 2022, arXiv preprint, doi:10.48550/arXiv.2203.05794. [11] C. M. Pham, A. Hoyle, S. Sun, P. Resnik, and M. Iyyer, “TopicGPT: A prompt-based topic modeling framework,” in Conference of the North American Chapter of the Association for Computational Linguistics (NAACL). ACL, 2024, pp. 2956–2984, doi:10.18653/v1/2024.naacllong.164. [12] P. Guitard, F. Ferland, and É. Dutil, “Toward a better understanding of playfulness in adults,” Occupational Therapy Journal of Research (OTJR), vol. 25, pp. 9–22, 2005, doi:10.1177/153944920502500103. [13] T. Dal Sasso, A. Mocci, M. Lanza, and E. Mastrodicasa, “How to gamify software engineering,” in International Conference on Software Analysis, Evolution and Reengineering (SANER). IEEE, 2017, pp. 261–271, doi:10.1109/SANER.2017.7884627. [14] A. Štěpánek, M. Raglianti, V. Rusňák, J. Byška, B. Kozlı́ková, and M. Lanza, “T(r)opical islands: Visualizing & understanding sociotechnical artifacts,” in International Conference on Software Maintenance and Evolution (ICSME). IEEE, 2026, accepted in May 2026, in press. [15] X. Zhang, Y. Yu, G. Gousios, and A. Rastogi, “Pull request decisions explained: An empirical overview,” IEEE Transactions on Software Engineering, vol. 49, pp. 849–871, 2023, doi:10.1109/TSE.2022.3165056. [16] H. Hata, N. Novielli, S. Baltes, R. G. Kula, and C. Treude, “GitHub Discussions: An exploratory study of early adoption,” Empirical Software Engineering, vol. 27, no. 1, pp. 1–32, 2021, doi:10.1007/s10664-02110058-6. [17] H. He, H. Su, W. Xiao, R. He, and M. Zhou, “Gfi-bot: Automated good first issue recommendation on GitHub,” in Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering (ESEC/FSE). ACM, 2022, pp. 1751–1755, doi:10.1145/3540250.3558922. [18] M. L. Siddiq and J. C. S. Santos, “BERT-based GitHub issue report classification,” in International Workshop on Natural Language-Based Software Engineering. ACM, 2022, pp. 33–36, doi:10.1145/3528588.3528660.

[19] A. Fiechter, R. Minelli, C. Nagy, and M. Lanza, “Visualizing GitHub Issues,” in Working Conference on Software Visualization (VISSOFT). IEEE, 2021, pp. 155–159, doi:10.1109/VISSOFT52517.2021.00030. [20] A. Kuhn, P. Loretan, and O. Nierstrasz, “Consistent layout for thematic software maps,” in Working Conference on Reverse Engineering. IEEE, 2008, pp. 209–218, doi:10.1109/WCRE.2008.45. [21] GitHub, “Rate limits for the REST API,” GitHub Docs, 2026. [Online]. Available: https://docs.github.com/en/rest/using-the-rest-api/rate-limitsfor-the-rest-api?apiVersion=2026-03-10 [22] Y. Zhang, M. Li, D. Long, X. Zhang, H. Lin, B. Yang, P. Xie, A. Yang, D. Liu, J. Lin, F. Huang, and J. Zhou, “Qwen3 embedding: Advancing text embedding and reranking through foundation models,” 2025, arXiv preprint, doi:10.48550/arXiv.2506.05176. [23] L. McInnes, J. Healy, N. Saul, and L. Großberger, “UMAP: Uniform manifold approximation and projection,” Journal of Open Source Software (JOSS), vol. 3, pp. 1–2, 2018, doi:10.21105/joss.00861. [24] L. McInnes and J. Healy, “Accelerated hierarchical density based clustering,” in International Conference on Data Mining Workshops (ICDMW). IEEE, 2017, pp. 33–42, doi:10.1109/ICDMW.2017.12. [25] X. Wu, C. S. Lam, K. H. Hui, H. H.-F. Loong, K. R. Zhou, C.-K. Ngan, and Y. T. Cheung, “Perceptions in 3.6 million web-based posts of online communities on the use of cancer immunotherapy: Data mining using BERTopic,” Journal of Medical Internet Research, vol. 27, pp. 1–12, 2025, doi:10.2196/60948. [26] O. Lezhnina, “Depression, anxiety, and burnout in academia: topic modeling of pubmed abstracts,” Frontiers in Research Metrics and Analytics, vol. 8, no. 1271385, 2023, doi:10.3389/frma.2023.1271385. [27] H. Edelsbrunner, D. Kirkpatrick, and R. Seidel, “On the shape of a set of points in the plane,” IEEE Transactions on Information Theory, vol. 29, pp. 551–559, 1983, doi:10.1109/TIT.1983.1056714. [28] O. Dabic, E. Aghajani, and G. Bavota, “Sampling projects in GitHub for MSR studies,” in International Conference on Mining Software Repositories (MSR), 2021, pp. 560–564, doi:10.1109/MSR52588.2021.00074.

Record · ID 660878 · SHA-256 72fa0f9c0a43ed38
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.