ConceptioArchivearXiv CS
arXiv CSopen access

Restructure This: Using AI to Restructure Onboarding Documents to Reduce Cognitive Overload

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

Noname manuscript No. (will be inserted by the editor)

Restructure This: Using AI to Restructure Onboarding Documents to Reduce Cognitive Overload

arXiv:2605.19174v1 [cs.SE] 18 May 2026

Zixuan Feng∗ · Prashant Tandan∗ · Igor Steinmacher · Marco Aurelio Gerosa · Anita Sarma

Received: date / Accepted: date

Abstract Onboarding documentation is critical for attracting and retaining newcomers in open source software (OSS). However, it is often presented as dense, inconsistently structured, and fragmented presentations that are difficult to understand, which creates cognitive overload leading to frustration, errors, and abandonment. Here, we investigate how Cognitive Theory of Multimedia Learning (CTML) strategies can be used to restructure OSS documentation. We use a GenAI-based pipeline to operationalize these strategies to restructure OSS documentation through our prototype VisDoc. VisDoc segments documentation into task-based units, infers workflows, removes redundancy, and generates multimodal explanations. An expert evaluation (N=4) affirmed VisDoc’s completeness, accuracy, and adoptability; A between-subjects evaluation (N=14) with newcomers found that VisDoc participants achieved higher task success, had significantly lower cognitive load, and perceived higher usability. The contributions of this work include a CTML-grounded analysis of onboarding challenges, a GenAI-based documentation restructuring pipeline, and empirical evidence that cognitively informed documentation restructuring reduces cognitive load and improves usability and task performance in OSS. Keywords Onboarding · Open Source · Newcomer · Documentation · Cognitive Load 1 Introduction Effective onboarding documentation is fundamental to helping newcomers engage with and contribute to Open Source Software (OSS)(Feng et al. 2024), Zixuan (Steve) Feng, Prashant Tandan, Anita Sarma Oregon State University, Corvallis, OR, USA E-mail: [email protected], [email protected], [email protected] · Igor Steinmacher, Marco Aurelio Gerosa Northern Arizona University, Flagstaff, AZ, USA E-mail: [email protected], [email protected]

2

Zixuan Feng∗ et al.

shaping how quickly they become productive and how they integrate into the project (Fronchetti et al. 2023). However, OSS documentation is frequently described as overwhelming, inconsistently structured, verbose, and difficult to follow (Pinho et al. 2024). These difficulties arise because of two key problems: content gaps and representation issues (Fronchetti et al. 2023; Imani et al. 2024; Pinho et al. 2024). Content gaps occur when essential instructions or information are missing because maintainers lack the time or clarity about what newcomers need. Even when the content exists, representation issues can occur when content is presented in dense, linear, or has inconsistent organization that obscures essential steps, workflow relationships, or procedural dependencies. Prior work has investigated using generative AI (GenAI) and Natural Language Generation (NLG) to address content gaps by producing additional explanations or guidance (Correia et al. 2024; McBurney and McMillan 2014; Yang et al. 2025; Imani et al. 2024; Fronchetti et al. 2023). However, representation issues in onboarding documentation have been largely unexplored. Closing this gap is essential since evidence shows that newcomers struggle because of poor documentation organization (Fronchetti et al. 2023). From a cognitive perspective, newcomers often struggle because OSS documentation requires them to reconstruct multi-step workflows, infer dependencies, and integrate scattered information across multiple files and external links (Imani et al. 2024; Pinho et al. 2024). This imposes substantial cognitive effort and leads to overload when documentation (1) includes dense, complex content, (2) has confusing organization or redundant details, (3) relies on text-only presentation (causing modality overload), and (4) fragments essential details across multiple locations (Sweller 1988, 2011). Such demands increase working-memory load, confusion about procedural order, and difficulty in creating mental models of how contribution tasks fit together (Steinmacher et al. 2014; Feng et al. 2025b), which, in turn, contribute to confusion, errors, and abandonment (Mayer and Moreno 2003; Fronchetti et al. 2023; Pinho et al. 2024). Instructional design research shows that structured representations such as task trees, flow diagrams, synchronized text-visual materials, and narrated demonstrations can reduce cognitive load by segmenting complexity, aligning related information, and providing clear pathways through multi-step processes (Çeken and Taşkın 2022; Mayer and Moreno 2003; Sweller 2011). These approaches externalize workflow structures, reduce the need for mental reconstruction, and improve learners’ ability to build accurate mental models. However, little work has examined how OSS documentation could be reorganized into structured, multimodal, or visual formats, leaving a fundamental gap in understanding how cognitively informed representations might improve onboarding. Currently, there is no empirical evidence showing whether such restructuring is effective, what forms it should take, or how to operationalize it in practice. At the same time, maintainers rarely have the time or resources to reorganize documentation manually. And finally, because documentation

Title Suppressed Due to Excessive Length

3

evolves continuously, any restructuring approach would need to be sustainable and easy to maintain. Recent advances in GenAI present an opportunity to address this gap by automating documentation restructuring that is otherwise too labor-intensive for maintainers. GenAI systems can automatically segment dense documentation into task-based units, reorganize content around actionable workflows, infer task dependencies, and highlight prerequisite knowledge addressing representation issues (Langchain 2025; Solbiati et al. 2021). Beyond text, AIpowered tools can generate multimodal materials such as narrated walkthrough videos, synchronized textual–visual explanations (Anthropic 2025; ElevenLabs 2025), and task diagrams that externalize workflow structures and reduce unnecessary cognitive effort (He et al. 2025). These capabilities make it possible to operationalize cognitive and instructional design principles at scale by restructuring existing documentation into clearer, more accessible onboarding pathways that better support newcomers’ learning and participation. In this paper, we investigate the use of genAI to operationalize strategies from the Cognitive Theory of Multimedia Learning (CTML) (Mayer and Moreno 2003; Mayer 2005b; Çeken and Taşkın 2022) to restructure OSS onboarding documentation. We adopt CTML as our analytical lens because it provides a structured, empirically grounded framework for understanding how humans process complex, multi-representational instructional materials—challenges that closely mirror the cognitive demands faced by newcomers in OSS onboarding. CTML also offers evidence-based design strategies for reducing cognitive load, providing actionable strategies that can be operationalized and empirically assessed in the context of OSS documentation. We implement 10 CTML strategies into a prototype named VisDoc. VisDoc automatically segments documentation into task-based units, infers workflow relationships, removes redundant or extraneous details, and presents the resulting structure as an interactive task tree augmented with synchronized text-and-video explanations. We acknowledge that GenAI systems can hallucinate, introduce inaccuracies, or misinterpret project conventions (Waldo and Boussard 2024). These risks make a fully automated restructuring workflow inappropriate for real OSS projects. For this reason, VisDoc is designed around a human-in-the-loop approach in which maintainers review, update, and modify the restructured documentation. This Human-AI collaborative model leverages the efficiency of genAI while preserving human oversight where contextual knowledge and domain expertise are essential. Through a two-phase evaluation, we assessed both the design soundness and practical effectiveness of VisDoc. First, an expert evaluation examined the completeness, accuracy, granularity, and adoptability of VisDoc’s automatically generated onboarding structures. Second, a between-subjects study with newcomers compared the effectiveness of VisDoc against the original documentation supplemented by ChatGPT. Our contributions are fourfold: (1) A theory-grounded analysis of cognitive overload in OSS onboarding. Using CTML, we characterize five forms of over-

4

Zixuan Feng∗ et al.

load present in onboarding documentation and derive concrete, evidence-based design strategies for mitigating them. (2) A GenAI pipeline that operationalizes cognitive and instructional design principles through the following steps: retrieve project documentation, segment it into task-based units, infer workflow structures, consolidate redundant content, eliminate irrelevant details, and generate multimodal instructional materials. (3) The design, implementation, and open-source release of VisDoc, a system that uses the above pipeline to automatically restructure OSS onboarding materials into an interactive, multimodal task-tree representation. (4) Empirical validation of our CTMLguided strategies implemented in VisDoc through expert review and newcomer testing. Expert evaluators (N=4) affirmed that VisDoc’s restructuring strategies produce onboarding materials with completeness, accuracy, and practical adoptability. A controlled study (N=14) with newcomers shows that applying these strategies yields significant improvements in task success, cognitive load, and perceived usability over documentation supplemented with ChatGPT, demonstrating that the strategies are both implementable and effective in practice.

2 Background This section presents the conceptual framework for our work. First, we examine how OSS documentation can scaffold newcomer contributions (Section 2.1). Second, we review prior interventions focused on improving access, retrieval, and engagement with existing documentation (Section 2.2). Third, we discuss a critical gap: despite extensive innovation in documentation tooling, little work has revisited the format of onboarding materials (Section 2.3). Finally, we argue for adopting the Cognitive Theory of Multimedia Learning (CTML) as an analytical lens, explaining why this framework positions us to diagnose documentation challenges and operationalize evidence-based design strategies for restructuring OSS documentation (Section 2.4). 2.1 Documentation as Both a Gateway and a Barrier in OSS Documentation is essential to understand project norms, workflows, and technical requirements (Steinmacher et al. 2015a; Pinho et al. 2024; Qiao et al. 2025b). Yet numerous studies have shown that OSS documentation serves as both a gateway and a barrier for newcomers (Aghajani et al. 2019). Newcomers report that documentation for contributing to a project is often overwhelming, overly verbose, and inconsistently structured, causing information overload and making it difficult to understand contribution workflows or even where to start (Fronchetti et al. 2023; Pinho et al. 2024). Even in the era of GenAI-assisted programming, contributors still need documentation to understand project-specific practices (Correia et al. 2024; Gaughan et al. 2025), showing that improving documentation is still an urgent and timely problem.

Title Suppressed Due to Excessive Length

5

2.2 Existing Interventions Focus on Access, Retrieval, and Engagement Prior works have focused on improving how newcomers access, locate, or retrieve information, or on increasing their motivation to engage with onboarding materials. For example, Steinmacher et al. (2016) introduced a newcomer portal that centralizes resources and reorganizes scattered documentation into structured sections. Cubranic et al. (2005) developed Hipikat, which surfaces project memory and contextual information to support task understanding. Gamification approaches (Toscani et al. 2018; Heimburger et al. 2020; Santos et al. 2024) have aimed to enhance engagement. Santos et al. (2023) identified GitHub inclusivity bugs and proposed design fixes to improve documentation navigation. More recently, LLM-based augmentation tools such as DocMentor enrich OSS documentation by using TF–IDF scores to select relevant passages and ChatGPT to generate additional explanations, examples, and references that clarify technical terms for practitioners (Imani et al. 2024). Similarly, Correia et al. (2024) extends a conversational agent with a retrieval-augmented generation (RAG) pipeline grounded in documentation to improve its accuracy. Although these interventions improve access and retrieval, and even augment documentation with generated explanations, they operate largely within the constraints of existing documentation format.

2.3 A Missing Focus on the Representation of Documentation Despite ongoing innovation in retrieval and augmentation tools, the underlying representation of OSS documentation remains almost entirely unchanged: newcomers are still expected to parse long, linear text and mentally reconstruct multi-step workflows, dependencies, and task order (Fronchetti et al. 2023; Pinho et al. 2024). Even LLM-based tools preserve this format by layering explanations on top of existing prose. As a result, newcomers continue to encounter documented challenges, reconstructing procedural sequences, integrating scattered information, and resolving ambiguous task dependencies, symptoms that, from a cognitive perspective, indicate forms of processing overload caused by text-heavy documentation structures (Steinmacher et al. 2015a). Outside of OSS, research shows that structured and visual representations can reduce cognitive load by externalizing relationships and reducing the need to infer the workflow structure (Mayer and Moreno 2003; Chattopadhyay et al. 2023). Software engineering (SE) researchers have experimented with alternative formats such as flowcharts (Islam et al. 2023; Kosower et al. 2014), activity diagrams (Saito et al. 2018), and visual or multimodal documentation, including images, videos, and narrated demonstrations (Van Der Meij and Van Der Meij 2014; Lloyd and Robertson 2012). Within OSS, even simple visual cues, such as screenshots, increase engagement and clarity in technical communication (Agrawal et al. 2022).

6

Zixuan Feng∗ et al.

However, no prior work reconsiders the representational form of OSS onboarding documentation, nor evaluates how such reorganization might reduce cognitive burden for newcomers.

2.4 The Need for a Theory-Grounded Perspective Empirical work in SE has repeatedly described symptoms of newcomer overload (Adejumo and Johnson 2024; Steinmacher et al. 2015a, 2014), yet we still lack a conceptual framework that explains why long, linear onboarding documentation produces such difficulties or how representational redesign might mitigate them. Several theories could serve as candidates for analyzing comprehension challenges, including Cognitive Load Theory (Sweller 2011), Dual Coding Theory (Clark and Paivio 1991), and Sensemaking Theory (Weick and Weick 1995). While each offers different insights, they emphasize general cognitive processes rather than providing actionable guidance for structuring complex, multi-step procedural material. A theory is needed that (1) characterizes specific cognitive processing challenges relevant to OSS onboarding and (2) provides actionable principles for redesigning documentation to address them. We adopt the Cognitive Theory of Multimedia Learning (CTML) (Mayer and Moreno 2003; Mayer 2005a; Çeken and Taşkın 2022) as our analytical lens because it provides a comprehensive framework for understanding how humans process complex instructional material composed of text, visuals, and sequential steps. The cognitive difficulties reported in OSS onboarding, such as reconstructing multi-step workflows, resolving scattered information, and mentally integrating task dependencies (Steinmacher et al. 2014; Fronchetti et al. 2023; Pinho et al. 2024), closely match the cognitive challenges articulated in CTML, which is explicitly designed to explain how humans make sense of complex, multi-representational instructional materials. Additionally, CTML offers a set of empirically validated design principles for reducing cognitive overload (Mayer and Moreno 2003), providing actionable strategies that can be operationalized and empirically assessed in the context of OSS onboarding.

3 Employing CTML as a Conceptual Framework for Reducing Cognitive Overload In this section, we employ CTML to characterize the cognitive overload that arises in OSS onboarding documentation and to identify evidence-based multimedia design strategies (Mayer and Moreno 2003; Çeken and Taşkın 2022) that can mitigate it.

Title Suppressed Due to Excessive Length

7

3.1 Cognitive Theory of Multimedia Learning (CTML) The Cognitive Theory of Multimedia Learning (CTML) (Mayer and Moreno 2003; Mayer 2005a; Çeken and Taşkın 2022) explains how humans make sense of instructional materials that combine multiple forms of representation. CTML assumes that humans process information through two separate channels, a verbal/auditory channel and a visual/pictorial channel, and that each channel has limited cognitive capacity. Meaningful understanding occurs when humans can select, organize, and integrate information across these channels without exceeding their cognitive limits (Mayer 2005a; Sorden 2013). Poorly structured materials (e.g., long unsegmented text, scattered visuals, unclear sequences) force humans into unnecessary cognitive work such as mentally reconstructing relationships, inferring missing steps, or integrating widely separated pieces of information (Chandler and Sweller 1991; Mayer and Chandler 2001). CTML has been widely applied to evaluate and improve complex instructional resources. In medical and technical domains, it has guided the redesign of animations, simulations, and instructional videos, helping learners build accurate mental models of multi-step procedures (Yue et al. 2013; AlShaikh et al. 2024). In online learning environments, CTML-informed design has been shown to reduce cognitive load and improve navigation and comprehension (Cavanagh and Kiersch 2023; Çeken and Taşkın 2022). Sorden (2005) highlights similar challenges in newcomer-oriented computer-based training, arguing that poorly organized multimedia materials impose unnecessary cognitive effort. A similar pattern appears in OSS onboarding. Prior work shows that contributing files often contain characteristics that increase cognitive burden for newcomers, such as scattered instructions, hidden dependencies, long linear text, and missing workflow structure, which make it difficult for newcomers to understand how contribution tasks fit together (Steinmacher et al. 2016; Fronchetti et al. 2023; Pinho et al. 2024). CTML, therefore, provides a structured analytical foundation for explaining why OSS documentation overwhelms newcomers. CTML distinguishes between several forms of cognitive overload: essential overload (when inherently complex content exceeds working memory capacity), extraneous overload (when poor organization or unnecessary material imposes avoidable processing demands), modality overload (when a single sensory channel becomes overburdened), and representational overload (when learners must mentally integrate fragmented information across sources) (Mayer and Moreno 2003). Each of these overload types manifests distinctly in OSS contributing files, where newcomers encounter dense technical procedures, scattered instructions, text-heavy presentation, and multi-document workflows. We use these overload types in the next subsection to characterize how OSS documentation imposes cognitive demands. Besides this conceptual framework, Mayer and Moreno (2003) experimentally validated design strategies to reduce unnecessary processing and make

8

Zixuan Feng∗ et al.

Table 1: Mapping Cognitive Overload Types in OSS Onboarding Documentation to CTML Principles and Applied Design Strategies Cognitive Overload Challenges C1. Essential Overload (High Complexity) Essential content is inherently complex and exceeds working-memory capacity. C2. Extraneous Overload (Confusing Presentation) Essential information is confusing or poorly organized. C3. Extraneous Overload (Extraneous Materials) Unnecessary details impose additional processing demands. C4. Modality Overload (Visual Channel Overuse) Dense, text-only presentation overloads the visual channel and leaves auditory capacity unused. C5. Representational Overload (WorkingMemory Burden) Learners must hold too many intermediate representations in working memory.

Challenges in OSS Contributing Files Documentation is long, dense, and text-heavy, requiring newcomers to independently learn project architectures, APIs, and workflows, creating high essential loads. Poorly structured documentations force newcomers to navigate across multiple locations to piece together the contribution workflow, imposing unnecessary extraneous cognitive load. Documentation often containsrepeated or unnecessary details, which impose additional processing demands and add extraneous cognitive overload.

CTML Strategies to Reduce Overload

Applied Design Strategy

• Segmenting: Break complex material into smaller units.

• Break documentations into short, task-based sections.

• Pretraining: Provide pretraining in the names and characteristics of components.

• Provide a brief overview that introduces key terms and components to ease later understanding.

• Aligning: Organize related information together to reduce scanning and integration effort.

• Group related information and contribution steps together to reduce back-and-forth navigation.

• Eliminating: Remove duplicated or conflicting details. • Signaling: Highlight core steps and guide attention to what matters most.

• Remove or consolidate redundant step instructions. • Mark the main contribution path to help newcomers make their contribution.

• Weeding: Remove or minimize extraneous material to reduce unnecessary processing.

• Pruning nonessential content to provide only essential instructions.

Documentation relies heavily on dense text walls, overloading the visual channel and causing modality overload

• Off-loading: Shift dense essential information from the visual channel to the auditory channel.

• Add short audio explanations or video walkthroughs to improve comprehension of dense materials.

When essential information is fragmented across sources, newcomers must keep intermediate representations in mind, exceeding their working-memory capacity.

• Synchronizing: Present explanations and visuals together to minimize representational holding and reduce working-memory load. • Individualizing: Adapt the level of detail to learners’ prior knowledge to ensure they can manage the required mental representations.

• Create visual overviews of the contribution workflow. • Pair instructions with diagrams or narrated demos. • Offer tiered layers of explanation (beginner vs. advanced).

complex procedural materials easier to follow. We leveraged these strategies to provide actionable principles for improving OSS onboarding documentation. In Table 1, we map each overload type to the specific challenges it creates in OSS onboarding, CTML strategies to mitigate the challenges, and the design implications for improving contributing files, which we further explore in the next subsections.

3.2 Cognitive Overload Challenges in OSS Contributing Documentation The following subsections outline the five CTML overload types (C1–C5) and illustrate how each manifests in OSS contributing files. 3.2.1 C1. Essential Overload (High Complexity) OSS newcomers must understand branching workflows, dependency and environment setup, CI/testing pipelines, review protocols, and project-specific

Title Suppressed Due to Excessive Length

9

architectural conventions before they can make their first contribution (Steinmacher et al. 2016; Aghajani et al. 2020). Essential overload arises when the inherent complexity of required content exceeds working-memory capacity (Mayer and Moreno 2003; Sweller 1988). Prior work repeatedly characterizes OSS contribution workflows as “highly complex,” “multi-step,” and “interdependent” (Pinho et al. 2024; Steinmacher et al. 2016). When such procedures are conveyed through dense text-only documentation, newcomers are required to coordinate multiple concepts at once, quickly exceeding what working memory can support and substantially complicating onboarding (Qiao et al. 2025a; Mayer and Chandler 2001). 3.2.2 C2 Extraneous Overload (Confusing Presentation) OSS contributing files often exhibit inconsistent section ordering, interleaving of related topics (e.g., setup, testing, submission rules) without clear signposting, divergent terminology, and contradictory descriptions of similar processes (Steinmacher et al. 2016; Aghajani et al. 2020). These issues create extraneous overload, which occurs when learners must expend cognitive effort on processing information that does not directly support understanding the underlying task (Mayer and Moreno 2003). Existing studies identified that newcomers waste substantial effort deciphering file organization, resolving terminology mismatches, and reconciling conflicting instructions (Pinho et al. 2024). These structural inconsistencies force contributors to focus on figuring out how documentation is organized rather than understanding the contribution workflow. 3.2.3 C3. Extraneous Overload (Extraneous Materials) OSS contributing files frequently include information that is unnecessary for completing the contribution task at hand, such as nonessential technical background, boilerplate explanations, duplicated content, and remnants of outdated practices (Aghajani et al. 2020; Fronchetti et al. 2023; Gaughan et al. 2025; Pinho et al. 2024). This produces extraneous overload by requiring learners to process information that does not meaningfully support task execution (Mayer and Moreno 2003). Such superfluous material forces newcomers to shift through noise before identifying actionable steps, increasing the time and cognitive effort required to understand what to do (Fronchetti et al. 2023). As a result, contributors expend mental resources filtering rather than learning, thereby increasing the overall onboarding burden. 3.2.4 C4. Modality Overload (Visual Channel Overuse) OSS contributing files often present instructional information through dense, prose-heavy text, where readability, structure, and document length can shape how easily developers understand and use the guidance (Aghajani et al. 2020; Pinho et al. 2024; Fronchetti et al. 2023). This overreliance on the visual–verbal

10

Zixuan Feng∗ et al.

channel forces newcomers to read and interpret long stretches of text without access to complementary modalities (Mayer 2005b). As a result, the visual channel becomes overburdened, obscuring the main contribution path and increasing the cognitive effort required to follow the workflow (Mayer 2005a). 3.2.5 C5. Representational Overload (Working-Memory Burden) OSS contribution guidance is frequently fragmented across README and CONTRIBUTING files, Wikis, issue templates, pull-request checklists, and externally linked documents (Aghajani et al. 2020; Gaughan et al. 2025; Pinho et al. 2024). Representational overload arises when learners must maintain and integrate fragmented information across multiple sources (Mayer and Moreno 2003). Newcomers must reconcile scattered instructions about setup steps, branching conventions, testing requirements, and review expectations (Feng et al. 2024). Such fragmentation forces contributors to juggle multiple partial mental models, quickly exhausting working-memory resources and obscuring the overall contribution workflow (Ayres and Sweller 2005). 3.3 CTML-Informed Strategies for Reducing Cognitive Overload The following subsections present CTML design strategies for each overload type (C1–C5) and describe how they can be used to mitigate cognitive overload. 3.3.1 Segmenting and Pretraining for Mitigating C1 OSS documentation can help ease the cognitive burden of complex, interdependent contribution workflows by using segmenting and pretraining (Mayer and Moreno 2003). Documentation can (1) segment the workflow into short, selfcontained procedural units so newcomers can process one step at a time, and (2) introduce core concepts upfront to pretrain newcomers with the minimal grounding needed for later steps. Evidence shows that segmenting improves procedural task performance in technical domains (Ganier 2004) and that concept-first pretraining lowers cognitive load for novices in programming and other knowledge-intensive tasks (Mayer and Moreno 2003; Gorbunova et al. 2025). 3.3.2 Eliminating and Aligning for Mitigating C2 Confusing or inconsistently structured documentation can be made easier to follow by using eliminating and aligning (Mayer and Moreno 2003). OSS documentation can (1) eliminate redundant or repeated step descriptions that otherwise impose avoidable cognitive effort. Prior work shows that aligning related elements lowers extraneous cognitive load in technical and procedural and (2)

Title Suppressed Due to Excessive Length

11

align related instructions and contextual details by placing them together, reducing the search effort required to connect dispersed information (Chandler and Sweller 1991; Nesbit and Adesope 2006). Removing redundant or duplicated information improves comprehension and reduces unnecessary cognitive effort (Mayer and Moreno 2003). 3.3.3 Signaling and Weeding for Mitigating C3 OSS documentation can help newcomers focus on what matters by signaling and weeding (Mayer and Moreno 2003). Documentation can (1) use headings, labels, or step cues to signal the main contribution path and highlight actions, and (2) weed out irrelevant, non-essential, or deep-linked material, such as long chains of secondary documentation links, to reduce unnecessary processing. Signaling has been shown to guide learners toward the essential steps and ease navigation (De Koning et al. 2009; Van Gog 2021), while weeding decreases the amount of distracting or off-task content learners must process, helping clarify the core procedures (Clark et al. 2011; Kalyuga 2011). 3.3.4 Off-loading for Mitigating C4 OSS documentation can ease modality overload by off-loading some explanations to the auditory channel. When all instructional content is delivered through dense text, newcomers must rely entirely on limited visual workingmemory resources (Mayer and Moreno 2003). Research shows that spoken guidance and short narrated walkthroughs reduce visual-channel load and improve comprehension of complex procedures (Mousavi et al. 1995). 3.3.5 Synchronizing and Individualizing for Mitigating C5 OSS documentation can help ease representational overload by synchronizing related information and individualizing the level of detail (Mayer and Moreno 2003). A high-level visual overview of the contribution workflow allows newcomers to form a stable initial mental model before engaging with detailed steps. Synchronizing explanations with diagrams or annotated screenshots reduces the need to hold intermediate representations in working memory (Ayres and Sweller 2005). Individualizing the depth of explanation by, for example, providing layered or beginner-focused descriptions, aligns detail with learners’ prior knowledge and has been shown to reduce overload and improve comprehension (Alreshidi 2021). 4 Design and Implementation of VisDoc In Section 3, we presented five cognitive overload challenges that burden newcomers during OSS onboarding (C1-C5) and CTML-informed strategies to mitigate each form of overload challenge. Building on this framework, we now

12

Zixuan Feng∗ et al.

present VisDoc, a web-based prototype that operationalizes these strategies into concrete design decisions to provide onboarding documentation for newcomers to OSS projects. VisDoc addresses cognitive overload by restructuring how onboarding documentation is presented and accessed. Rather than presenting newcomers with traditional linear prose, VisDoc processes project documentation, segments it into small, coherent onboarding actions, sequences these actions into a logical workflow, and provides a visualization as an interactive task tree. For selected nodes, VisDoc provides synchronized audio–visual walkthroughs that supplement text-based explanations. Table 1 maps each cognitive overload type to its corresponding CTML strategy (column 3) and employed design strategy (column 4). The remainder of this section describes how VisDoc implements these design strategies through specific interface features and technical components. 4.1 VisDoc Overview This section presents a walkthrough of the VisDoc prototype. Figure 1 provides an annotated overview of the interface. Walkthrough videos are available in the supplementary materials (Feng 2025), and the open-source code is provided in our repositories.1 Illustrative example: To better understand VisDoc, consider the following usage example. Alex, a new OSS contributor exploring the OSS project, is reading the README.md file when they notice an option labeled “View this README in VisDoc”. Curious, Alex clicks the link. Upon launching VisDoc, Alex arrives at the tree root ( A ), which anchors the full contribution workflow. Using the expand/collapse controls ( B ), Alex opens all nodes to reveal sub-tasks such as Improving Documentation, Creating a Pull Request, Testing, and Implementing New Models. The expanded structure helps Alex understand both the breadth and organization of the project’s contribution pathways. Exploring a bit more, Alex opens the tasks under Improving Documentation, including Generating Documentation, to see how these activities are structured within the tree while maintaining awareness of the overall workflow. Alex’s main goal is to understand how to push his changes as part of his pull request submission, so instead of expanding the full tree, Alex uses the category selector ( D ) to jump directly to the nodes related to pull-request submission. To be more efficient, Alex uses the search bar ( E ) to find instructions on how to push changes. VisDoc highlights the relevant nodes, directing Alex to the Commit and Push Changes node, which offers both text and video instructions. If Alex were an advanced user who wanted to adjust or annotate nodes, they could use the Edit Tree button ( G ), which provides lightweight editing functionality for modifying text or embedding supporting materials. At any time, if the tree is fully expanded, Alex can restore the original compact 1 https://github.com/EPICLab/visdoc_framework; visual_doc_demo

https://github.com/EPICLab/

Title Suppressed Due to Excessive Length

Tag A B C D E F G

Feature Tree Root Expand/Collapse Text/Video buttons Choose a category Search Bar Clear Button Edit Tree Button

13

Description Starting point of task tree Show/hide sub-tasks (children) Show text instructions or video tutorial Choose a branch to start working on Search nodes or contents inside nodes by keyword Restore the graph to original collapsed form Edit the text of the nodes or add images

Fig. 1: VisDoc Task Tree UI with tagged features.

layout using the Clear button ( F ), returning the interface to a clean, collapsed state. 4.2 CTML-Guided Design Strategies Segmenting and Pretraining for mitigating C1. To reduce essential overload (C1), VisDoc applies CTML’s segmenting and pretraining strategies by breaking complex onboarding documentation into short, task-based units and generating a high-level catalog of contribution task types. To provide conceptual pretraining, VisDoc generates concise overview titles for each segment. Eliminating and Aligning for Mitigating C2. To address extraneous overload stemming from confusing presentation (C2), VisDoc removes redundant content across linked documents and groups relevant information through an integrated task-tree visualization that connects each action node to its corresponding instruction. VisDoc also detects and consolidates redundant instructions and filters recurring boilerplate content across linked files. This structure supports CTML’s aligning strategy by reducing cross-document search and clarifying the structure of the workflow, grouping related steps into coherent clusters and placing dependent actions in close proximity. Signaling and Weeding for mitigating C3. For extraneous overload caused by nonessential or low-value material (C3), VisDoc applies signaling and weeding by marking the main contribution path, presented in collapsed

14

Zixuan Feng∗ et al.

task trees and controlled traversal of linked documents. To apply weeding, VisDoc prunes unnecessary details in two ways. First, the task tree initially shows only high-level structure, allowing users to reveal details on demand. Second, traversal of linked documentation is limited to two levels, helping to keep the visualization focused on the main contribution process. Off-loading for mitigating C4. To address modality overload (C4), VisDoc creates explanations via audio-visual walkthroughs for complex steps, in addition to the textual explanations. Synchronizing and Individualizing for mitigating C5. To reduce representational overload (C5), VisDoc presents related information together and adapts details to learners’ needs. Synchronizing is achieved by pairing instructions with diagrams or audio–visual demos directly within the task tree, allowing contributors to view actions and explanations in the same context. To individualize the experience, VisDoc provides tiered layers of explanation: concise text summaries for experienced contributors and richer audio–visual walkthroughs for newcomers, aligning with CTML recommendations for adapting detail to prior knowledge. 4.3 Infrastructure Overview Figure 2 presents an overview of VisDoc’s system architecture. VisDoc follows a two–tier design consisting of a Python backend and a React (Gackenheimer 2015) frontend. Documentation Retrieval and Preprocessing. VisDoc retrieves and standardizes the repository’s onboarding documentation based on a user-provided link to the repository’s contributing files. VisDoc fetches the files in Markdown format (and directly linked files), converts them to plain text while preserving structural hierarchy, and removes formatting marks. The backend is responsible for all content processing and multimodal generation. It integrates large language models (LLMs) (Kasneci et al. 2023) and a retrieval-augmented generation (RAG) pipeline (Lewis et al. 2020) to (1) retrieve and clean onboarding documentation, (2) perform segmentation and topic inference, (3) infer task dependencies, and (4) generate multi-modal instructional materials. Audio–visual walkthroughs are produced by combining LLM-generated scripts, Claude Computer Use demonstrations (PBC 2024), and ElevenLabs narration (ElevenLabs 2025). The frontend renders the processed materials as an interactive task tree tailored around CTML principles. Using React, it supports hierarchical navigation, selective detail expansion, synchronized text–video presentation, and lightweight editing features for maintainers. 4.3.1 Implementation Details Segmentation: We evaluated two segmentation approaches: a RoBERTabased unsupervised algorithm (Solbiati et al. 2021) and LangChain’s Semantic

Title Suppressed Due to Excessive Length

15

User Input (Repo URL) Python Backend Retrive Documentation GPT-4o + Cognita RAG Framework Segment tasks into Units.

Create titles and analyze dependencies

Create instruction scripts

Eliminate redundancies and refine boundaries.

Generate instruction videos

Merge audio/video into a single multimodal micro-doc.

Front-End Interface: Interactive Task Tree Explorer Using React

Fig. 2: VisDoc Infrastructure Overview

Chunker (Langchain 2025). We used a ground-truth segmentation of the CONTRIBUTING.md of an OSS project (Kubernetes)2 , annotated independently by two researchers (93.8% agreement (McHugh 2012)). We compared both methods using Pk (Beeferman et al. 1999) and WinDiff (Pevzner and Hearst 2002). LangChain’s Semantic Chunker performed better than RoBERTa (Pk = 0.33 vs. 0.36; WinDiff = 0.24 vs. 0.29) and was adopted. Since LangChain’s Semantic Chunker sliding-window similarity method (Kiss et al. 2025) occasionally generates misplaced boundaries (incorrect splits or merges) (Langchain 2025), we refined the post-processing refinement step using OpenAI’s GPT-4o (Hurst et al. 2024). The LLM was prompted to review each tentative segment and adjust the boundary when semantic coherence was compromised (see implementation details in the repositories3 ). To mitigate hallucination (Tonmoy et al. 2024), we employed a RAG framework (Lewis et al. 2020) using Cognita (TrueFoundry 2025). This approach produced subhttps://github.com/kubernetes/kubernetes https://github.com/EPICLab/visdoc_framework; visual_doc_demo 2 3

https://github.com/EPICLab/

16

Zixuan Feng∗ et al.

stantial improvements in quality (Pk = 0.11; WinDiff = 0.09). All subsequent LLM-enabled components in VisDoc also use GPT-4o within the Cognita RAG framework to ensure consistency and reliability. Pre-training: Overview titles are inferred by GPT-4o. Previous work shows that LLMs produce more coherent topics than traditional methods such as LDA, NMF, or LSA (Azher et al. 2024; De-Marcos and Domínguez-Díaz 2025; Kapoor et al. 2024). Alignment: To consolidate redudant instructions, we uses LLM-assisted preprocessing. After segmentation, topic inference, and redundancy removal, each documentation chunk becomes a node in a hierarchical task tree rendered in the interface. Nodes are labeled with GPT-4o-generated topics, and edges represent the LLM-inferred task dependencies. Signaling: To signal critical steps, VisDoc sequences nodes into a primary workflow path using GPT-4o with Cognita. We employ few-shot prompting (Brown et al. 2020) (supplying representative examples of contribution workflows) to improve procedural consistency (Ibrahim et al. 2024), allowing the model to infer ordering cues, such as performing environment setup, before submitting a pull request. Prior work shows that such prompting strategies help large language models generate more structured and consistent outputs (Ibrahim et al. 2024). This enables the model to capture subtle procedural cues (e.g., completing environment setup before submitting a pull request) and to produce reliable task sequences without manual annotation. Off-loading: We used GPT-4o + Cognita to transform each documentation segment into a detailed instructional script using a few-shot prompting to shape the output to get a consistent structure, an appropriate level of detail, and actionable phrasing. These scripts are executed in a live environment using Claude Computer Use (Anthropic 2025), with recordings edited to remove incorrect actions. Users may upload screenshots to clarify UI-dependent steps, which are incorporated into the edited video. We used ElevenLabs (ElevenLabs 2025) to generate the narration from the script, which is merged with the edited video and attached to the corresponding node.

4.4 Iterative Development and Formative Evaluation During the design and implementation of VisDoc, we selected the Kubernetes project 4 as our development case study. We ingested Kubernetes documentation into the system and systematically inspected the outputs of segmentation, topic inference, task dependency ordering, and multimodal generation using iterative feedback from our research team, comprising experienced OSS researchers and practitioners. This process allowed us to identify and correct misclassifications, adjust the implementation, improve task sequencing, and ensure the instructional scripts and visual walkthroughs aligned with realworld OSS practices. These inspections were conducted at multiple stages 4

https://github.com/kubernetes/kubernetes

Title Suppressed Due to Excessive Length

17

throughout development and directly informed architectural choices, prompting adjustments to model prompts, segment post-processing heuristics, and multimodal generation workflows. Once VisDoc was implemented, we conducted feedback sessions with four external software engineering researchers. These participants freely explored the VisDoc user interface, compared it with the original CONTRIBUTING.md, and provided suggestions for additional improvements. Our research group reviewed their feedback and incorporated several enhancements to improve the navigation, reduce visual clutter, and strengthen overall learnability. In the following, we discuss the results of these feedback sessions. 1. Task Categories: In Kubernetes, the first layer of the task tree became excessively wide, requiring users to scroll horizontally to locate relevant branches. To reduce navigation friction, we introduced a drop-down menu that lists all task categories. Selecting a category automatically highlights and centers the corresponding branch in the graph, improving orientation and reducing scanning effort. 2. Search Functionality: Participants expressed difficulty locating specific information when multiple branches were expanded. Several participants attempted to use the browser’s native search, which cannot reveal collapsed or deeply nested nodes. To address this limitation, we implemented a dedicated search function that matches query terms against both node titles and their underlying content. Search results are presented as a list, and selecting an item automatically expands and highlights the corresponding node within the graph. 3. Clear Button: As participants explored the graph, expanded branches accumulated visual clutter, making it difficult to regain a clean overview of the documentation. To support quick reorientation, we added a “Clear” button that collapses all branches and restores the graph to its initial state. 4. Edit Tree Button: Participants felt that some auto-generated node content was not of the right length or inconsistently formatted. They also noted that project maintainers may wish to curate or refine content for contributor onboarding. To support customization, we added an “Edit Tree” feature that opens a markdown-based editing interface where users can revise node text and embed images directly into nodes. 5 Evaluation of VisDoc To evaluate VisDoc, we conducted a two-phase study (as shown in Figure 3). We first worked with OSS experts (N=4) to assess the system’s overall quality. Second, we conducted a user evaluation with newcomers (N=14) to examine VisDoc’s effectiveness in supporting onboarding. Case selection. We selected the Transformers 5 repository to evaluate VisDoc’s performance. We selected a project different from the one used dur5

https://github.com/huggingface/transformers

Zixuan Feng∗ et al.

Phase 2

Phase 1

18

Expert Evaluation (N=4)

D1. Completeness and Coverage. D2. Accuracy and Correctness. D3. Granularity and Abstraction. D4. Adoptability and Integration.

Treatment Group (N=7): VisDoc Newcomer Evaluation (N=14)

Control Group (N=7): Documentation+ChatGPT

H1: Task Performance. H2: Cognitive Load. H3: Usability.

Fig. 3: Two-Phase Evaluation: Expert Evaluation and Between-subject User Study.

ing development and our formative evaluation to promote transferability and adaptability across OSS contexts (Guizani et al. 2025). We chose the Transformers project because: (1) It belongs to the AI/ML domain, a very different domain from the Kubernetes-based project, allowing us to assess generalization across technical ecosystems. (2) It is a popular project, consistently ranking among GitHub’s most starred and forked repositories (at the time of our study, it was the 34th most starred repository on GitHub6 and had been forked over 28,000 times). (3) Transformers’ documentation is large, multi-layered, and densely interlinked, characteristics known to contribute to cognitive overload during onboarding (Steinmacher et al. 2015b). 5.1 Phase 1. Expert Evaluation 5.1.1 Study Design The goal of the expert evaluation was to assess whether the system met expectations across the following four dimensions: D1. Completeness and coverage relate to whether VisDoc covers all essential steps, topics, and prerequisite concepts, as well as does not unintentionally omit important content (Tang and Nadi 2023; Garousi et al. 2013). D2. Accuracy and correctness relate to whether VisDoc’s segmentation, labeling, and inferred action sequences accurately represented the meaning, intent, and ordering of the original documentation and it did not hallucinate, infer, or introduce irrelevant details beyond the source documentation (Tang and Nadi 2023; Garousi et al. 2013). D3. Granularity and abstraction appropriateness relate to whether VisDoc represents documentation at a level of detail that is both meaningful and usable (Tang and Nadi 2023; Garousi et al. 2013). For example, whether segmentation was neither fragmented nor overly coarse, accurately reflected step intent, and supported intuitive navigation. 6

md

https://github.com/EvanLi/Github-Ranking/blob/master/Top100/Top-100-stars.

Title Suppressed Due to Excessive Length

19

D4. Adoptability and integration potential relate to whether VisDoc would be acceptable and feasible for OSS maintainers to incorporate into real projects, similar to how prior intervention studies evaluate the practicality of adopting new tools or processes (Feng et al. 2025a). 5.1.2 Study Protocol The study was designed to be 40–60 minutes long and consisted of three parts. First, we asked background questions—including optional questions about experts’ gender identity and their experience contributing to OSS. Second, we asked experts to review the original CONTRIBUTING.md file from the Transformers project on GitHub. To ensure that all experts engaged with comparable parts of the documentation and to elicit feedback on scenarios that newcomers commonly encounter, we asked experts to focus on reviewing the relevant information of two common onboarding tasks: improving documentation and creating a pull request (Turzo et al. 2024). Then, the experts reviewed the corresponding information for the same two tasks using VisDoc and were asked to think aloud as they navigated its interface. Before beginning, they watched a short tutorial introducing the system. Experts were informed that, if needed, they could return to the original CONTRIBUTING.md file to compare or verify information during their review. We then conducted a follow-up interview focused on the four evaluation dimensions described above (D1–D4). We asked the experts to comment on VisDoc’s completeness, accuracy, granularity, and adoptability based on their experience using the tool. At the end of the study, we asked for general reflections on any confusing, missing, or surprising aspects of the tool, as well as anticipated challenges or unmet needs. The study was approved by our university’s Institutional Review Board (IRB) (See the supplementary materials for detailed study information (Feng 2025)). 5.1.3 Pilot studies and Expert recruitment Before conducting the study, we ran two pilot sessions with software engineering researchers to evaluate the clarity and usability of our study protocol. Based on their feedback, we refined the study flow, for example, adjusting when the tutorial video was introduced to avoid interrupting experts’ progression, removing repetitive questions, and adding an explanation of the think-aloud procedure for experts unfamiliar with it. Then, we recruited four OSS practitioners through the authors’ professional networks: one maintainer, one OSS researcher, the CEO of an OSS mentoring program, and one scientific project manager from a national laboratory (Table 2). Prior SE and HCI literature (Levi and Conrad 1996; Alroobaea and Mayhew 2014) demonstrates that 3–5 domain experts are sufficient to uncover the majority of conceptual, structural, and correctness issues in early-stage tools. The experts in our study reflected diverse OSS roles, offering broad coverage

20

Zixuan Feng∗ et al.

of perspectives related to documentation quality, contribution workflows, and newcomer onboarding. Table 2: Demographics of Expert Evaluators (E1–E4) Expert

OSS Experience (Years)

Primary OSS Role

E1 E2 E3 E4

>20 years 19 years 6 years 9 years

Maintainer / Technical Contributor OSS Community Leader (Mentoring Program CEO) OSS Researcher / Contributor Scientific Project Manager / OSS Contributor

5.1.4 Expert Feedback D1. Completeness and Coverage. Experts consistently reported that VisDoc captured all essential information, covering the full scope of the onboarding workflow and providing everything needed to complete the assigned tasks [E1E4]. “Between all of this, what I have found, I believe I have enough information to be able to improve the documentation” [E4]. Later, when asked whether anything was missing, E4 mentioned: “For the tasks in question, I was able to find everything that I expected to find.” Similarly, E2 confirmed “I was able to find the information that I was looking for” [E2]. D2. Accuracy and Correctness. Experts confirmed that VisDoc correctly preserved the meaning, intent, and procedural ordering of the original documentation [E1-E4]. E4 explained that “the AI used in this tool seems relatively conservative, which is good. It doesn’t hallucinate, and does not completely mix up everything” [E4]. E2 also highlighted that VisDoc’s segmentation and summarized node terms did not distort the underlying meaning “I also think that, like, the little descriptions on the nodes in the graph were generally pretty accurate” [E2]. E3 further highlighted how the tool correctly captured and structured the original documentation: “This tool represents all of that information in a graph... You can also see this expand button on the side of the node... I like that”. E2 similarly mentioned that VisDoc accurately preserved the original workflow’s intent “I feel like in terms of, like, preserving the intent, I would say it’s preserving the intent that I was able to infer from the original documentation” [E2]. D3. Granularity and Abstraction. Experts E1–E4 reported that VisDoc struck an effective balance between high-level overviews and detailed tasklevel guidance. Across all experts, the overview was described as immediately useful for orienting themselves within the documentation. E2 mentioned that “the traditional README files for this project and for many projects I’ve used are very large. It’s very overwhelming with the big README file, but I think that with this, because with the README files, usually you just have so much information there, and if you don’t break it down into smaller pieces, it’s very hard to work with it”. E3 appreciated having a structured entry point rather

Title Suppressed Due to Excessive Length

21

than confronting what they described as “usually a huge text wall. . . [In VisDoc], it’s better to be able to see a quick summary of that information in a graph”. E2 similarly noted, “I also appreciate how it gave me a high-level view that I could click into for details. That was good because I didn’t want to read a lot of the details.” For the level of detail, E1 emphasized that “the amount of detail in the overview is very helpful”. D4. Adoptability and Integration. Experts viewed VisDoc as both feasible and valuable to integrate into real OSS workflows. I would use it. . . because it helps me understand the whole layout much faster.” [E4]. E4 explained that they could imagine using this as a tutorial or an onboarding helper [for] adding new tutorials, translating documentation, or providing tips and tricks on how to use it best”. Similarly, E2 mentioned that “for a first-time contributor, this could be a very useful tool to be able to find all the information you need”. Overall, expert feedback indicated that VisDoc met expectations along all four evaluation dimensions and surfaced no major conceptual or structural issues. These results gave us confidence to proceed with the newcomer study, in which we evaluated VisDoc’s effectiveness with its intended end users. 5.2 Newcomer Onboarding Study To evaluate VisDoc’s effectiveness in supporting newcomer onboarding, we conducted a between-subject controlled study. The study was designed to simulate the experience of a first-time contributor by asking participants to complete a series of onboarding tasks. Participants were randomly assigned to one of two conditions: using VisDoc or using the Transformers project’s original CONTRIBUTING.md documentation. Because conversational agents such as ChatGPT are now routinely used in software development (Das et al. 2025), prohibiting their use would undermine real-world usage conditions (Schmuckler 2001). Therefore, in the control condition, participants worked with the project’s official contributing materials and were allowed to use ChatGPT if they wished. In the treatment condition, participants completed the same onboarding tasks using VisDoc as their only resource. 5.2.1 Study Design Our evaluation focused on three dimensions widely recognized in HCI and software engineering as foundational for assessing developer-facing tools (Ko et al. 2015; Gonçales et al. 2019; Leßenich and Sobernig 2023): (1) task performance— to assess whether newcomers can correctly complete the tasks; (2) cognitive load—to evaluate newcomers’ cognitive load and understand how VisDoc’s CTML-informed design relates to the overload factors outlined in Table 1; and (3) usability—to evaluate how easy the system is to learn and operate. D1. Task performance. Following established HCI and software engineering evaluation practices, which define performance as completing the task correctly within a time-bounded window, we selected three onboarding tasks from the

22

Zixuan Feng∗ et al.

Transformers project’s documentation, which represent activities newcomers typically encounter (Turzo et al. 2024). T1: Create a pull request with a new file following the Transformer’s contribution requirements. Add new functionality as a first contribution (Turzo et al. 2024) This task was selected because it is both a small, self-contained task and an essential step for understanding the project’s entire contribution workflow. T2: Translate a line of developer documentation into Tibetan. Start by contributing documentation changes (Turzo et al. 2024) Documentation edits are low-risk, high-clarity tasks, allowing meaningful engagement without requiring deep system knowledge. T3: Create a new example script from a template. Work on smaller tasks first, then progressively larger ones (Turzo et al. 2024) This task is localized and scaffolded, enabling newcomers to make a small, guided change without needing to understand the full codebase. D2. Cognitive load. The second dimension evaluates cognitive load to determine if VisDoc’s CTML-grounded, load-reducing design (Section 4) successfully reduces cognitive effort in practice. We employed the NASA Task Load Index (TLX) to assess whether VisDoc’s CTML-based design (Section 4) effectively reduces cognitive load during OSS onboarding tasks. NASA TLX is a widely used and validated instrument across software engineering, HCI, and learning sciences (Hart and Staveland 1988). Its multidimensional structure including mental demand, temporal demand, effort, performance, and frustration, aligns closely with CTML’s theorized mechanisms for reducing extraneous and essential processing. Table 3 maps each NASA–TLX dimension to the VisDoc’s CTML-informed design strategies. Mental Demand captures whether segmenting and pretraining reduced intrinsic load by enabling users to process complex instructions incrementally. Temporal Demand reflects whether aligning related steps and signaling the main contribution path shortens search time during task navigation. Effort assesses whether weeding and eliminating redundancy reduced extraneous cognitive work by removing unnecessary or repeated content. Performance evaluates whether synchronized narration–visual pairs and tiered individual explanations supported task execution. Frustration measures whether off-loading (e.g., short audio/video walkthroughs) alleviated emotional or affective load during complex steps. We excluded Physical Demand because it does not apply to documentation-based contribution tasks. D3. Usability. Usability captures how easy a system is to learn and interact with (Brooke et al. 1996; Steinmacher et al. 2016), which is especially critical in onboarding contexts where early impressions strongly influence engagement and contribution quality (Padoan et al. 2024; Li et al. 2024). To assess usability, we used the System Usability Scale (SUS) (Brooke et al. 1996), a validated 10-item questionnaire set that produces a single score (0–100) reflecting perceived usability. SUS is appropriate for our setting because it measures four dimensions: learnability, ease of use, perceived complexity, and user confidence,

Title Suppressed Due to Excessive Length

23

Table 3: Mapping NASA-TLX Dimensions to CTML Load-Reducing Strategies Implemented in VisDoc NASA-TLX Dimension Mental Demand

Temporal Demand

Effort

CTML Informed Design Strategies Segmenting: Break complex documentations into taskbased sections. Pretraining: Provide key terminology/components. Aligning: Group related information and contribution steps together. Signaling: Mark the main contribution path. Weeding: Pruning nonessential content to provide only essential instructions. Eliminating redundancy: Remove duplicated or unnecessary instructions.

Performance

Frustration

Synchronizing: Present narration and visuals simultaneously. Customizing: Offer tiered layers of explanation. Off-loading: adding short audio explanations or video walkthroughs.

Why This TLX Dimension Captures the Strategy’s Effect Measures whether segmenting reduced the intrinsic cognitive complexity of the documentation, supporting easier incremental processing. Measures whether providing conceptual scaffolding effectively reduces the germane load required to comprehend task instructions. Measures whether visual alignment minimized the time spent locating relevant information. Measures whether signaling cues accelerated information processing by directing user attention to the contribution path. Measures whether weeding resulted in lower perceived effort, reflecting the mitigation of extraneous load by reducing unnecessary cognitive work. Measures whether the removal of redundancy effectively reduced overall cognitive workload by eliminating the need to reconcile repeated or fragmented explanations. Measures whether synchronized explanations and visuals helped users maintain task accuracy and reduced errors in execution. Measures whether tiered explanation layers improved self-assessed task success. Measures whether multimodal off-loading reduced emotional/affective load, leading to less user frustration and a smoother experience.

*CTML strategies drawn from multimedia learning theory; NASA-TLX dimensions measure whether these strategies effectively reduced user cognitive load.

which directly affect whether newcomers can successfully incorporate a new tool into their onboarding workflow.

5.2.2 Study Protocol Our controlled study followed a between-subjects design (Lazar et al. 2017), with each participant randomly assigned to either the VisDoc condition or the documentation + optional-ChatGPT support. The study began with informed consent and a brief pre-study survey for understanding demographic information and prior technical experience, including Git, GitHub, Python, Linux commands, and prior use of LLM (e.g., ChatGPT). Before beginning the main tasks, participants were introduced to the thinkaloud practice to prepare them to verbalize their reasoning, information-seeking strategies, and impressions of the tool-documentation combination. Participants were then asked to complete the three newcomer-oriented onboarding tasks. Each task included a short reading period (1–2 minutes) followed by task completion using the think-aloud method. For the documentation + optionalChatGPT group, participants were allowed to refer back to the material at any time.

24

Zixuan Feng∗ et al.

To run these tasks, we created a fork of the Transformer project to serve as our base repository. Subsequent forks were created under each participant’s assigned GitHub account and cloned to the designated study machine, ensuring that no accidental changes or pull requests were made to the Transformer’s repository. The experimenter observed and took notes but did not provide assistance. All sessions were audio- and screen-recorded with the participant’s permission to support later qualitative analysis. After completing the tasks, participants filled out a post-study questionnaire that included the System Usability Scale (SUS), the NASA Task Load Index (TLX), and three open-ended questions about what they found useful or challenging in the documentation they used, as well as their suggestions for future improvements. 5.2.3 Sandbox and Pilot Study We conducted five sandbox sessions and two pilot studies to iteratively refine our study design, task materials, and data-collection procedures. The sandbox sessions focused on verifying the study design, programming environment, and task feasibility. The two pilot studies involved participants similar to our target population and were used to assess the clarity of instructions, the appropriateness of task difficulty, and the overall study flow. Feedback from both stages informed revisions to task wording, study flow, and the timing structure of the study. Based on observed completion times during the pilot studies, we established a 30-minute time limit for each task. All study materials are provided in the supplementary documents (Feng 2025). 5.2.4 Participant Recruitment We recruited participants through multiple channels, including university mailing lists, announcements in upper-division and graduate CS courses, and snowball sampling. Interested individuals completed a short screening survey reporting their experience with Git, GitHub, Linux command-line tools, Python, and conversational AI tools. During the participant selection, to simulate a realistic newcomer onboarding scenario, we did not restrict participants’ GitHub contribution experience as newcomers frequently begin contributing despite limited prior exposure to OSS workflows (Steinmacher et al. 2016, 2014). However, because our study tasks involved navigating Python files, running commands, and understanding basic version-control operations, participants needed at least 1 year of Python experience and basic familiarity with Git and the command line. We also asked participants about their prior use of large language models (e.g., ChatGPT) to support balanced assignment across conditions. This ensured that the documentation+ChatGPT group did not disproportionately include participants with either very high or very low prior LLM usage. A total of 14 participants met the inclusion criteria and were enrolled in the study, as

Title Suppressed Due to Excessive Length

25

shown in Table 4. The two participants who did not have any GitHub experience were in the Experimental group; thus, in case of any penalty from lack of experience, it would impact the performance in this group. Table 4: Participant Demographics and Technical Experience ID

Gender

Git

Python

Command Line

LLMs

GitHub

Group

P3 P4 P5 P6 P7 P8 P9 P10 P11 P12 P13 P14 P15 P16

Man Man Man Man Man Man Woman Man Woman Man Man Man Man Woman

> 3–5 years 6 months–1 year > 3–5 years > 3–5 years > 2–3 years > 5 years > 5 years > 3–5 years > 5 years > 3–5 years > 3–5 years 6 months–1 year > 2–3 years > 3–5 years

> 1–2 years > 1–2 years > 5 years > 2–3 years > 3–5 years > 5 years > 2–3 years > 5 years > 1–2 years > 3–5 years > 3–5 years > 1–2 years > 3–5 years > 3–5 years

> 3–5 years 6 months–1 year > 3–5 years > 1–2 years > 2–3 years > 3–5 years > 1–2 years > 2–3 years > 3–5 years > 2–3 years > 5 years > 1–2 years > 2–3 years > 3–5 years

> 1–2 years > 1–2 years > 1–2 years > 1–2 years > 2–3 years > 2–3 years 6 months–1 year > 2–3 years > 2–3 years > 1–2 years > 1–2 years 6 months–1 year > 2–3 years > 2–3 years

> 3–5 years < 6 months > 3–5 years > 2–3 years > 1–2 years Never used > 1–2 years > 3–5 years > 5 years > 1–2 years > 3–5 years Never used > 2–3 years > 3–5 years

D+LLM VisDoc VisDoc D+LLM D+LLM VisDoc VisDoc VisDoc D+LLM D+LLM D+LLM VisDoc VisDoc D+LLM

Note: P1 and P2 participated only in pilot testing and are not included in this table or the main analysis. D+LLM = Original CONTRIBUTING documentation with optional ChatGPT assistance; VisDoc = graph-structured documentation interface.

5.2.5 Data Analysis We used a mixed-methods analysis to understand how participants interacted with each condition and how those interactions shaped task performance, cognitive load, and usability. Task performance was operationalized as a binary outcome (success/failure) for each task completed within 30 minutes. We compared performance across conditions using Fisher’s Exact Test (Fisher 1922) to evaluate hypothesis H1. Task performance would differ between participants using VisDoc and those using the original project documentation (with optional ChatGPT). Cognitive load was assessed through the NASA TLX. We used the MannWhitney U test (Mann and Whitney 1947) to compare cognitive load across the two conditions and evaluate hypothesis H2. Participants’ cognitive load would differ between the two groups. Usability was evaluated using the System Usability Scale (SUS). Following standard SUS scoring procedures, we computed total SUS scores for each participant and compared score distributions using the Mann–Whitney U test, evaluating hypothesis H3. Perceived usability would differ between the two groups. Qualitative analysis. To complement the quantitative comparisons, we conducted an inductive qualitative analysis of participants’ think-aloud comments, screen recordings, and interview responses. Guided by our three evaluation dimensions (task performance, cognitive load, and usability), two researchers independently coded all qualitative data using these dimensions as the top-level

26

Zixuan Feng∗ et al.

codebook. Coding focused on identifying evidence explaining why participants succeeded or struggled with tasks, what aspects of the interface increased or reduced cognitive load, and how participants experienced the system’s learnability and usability. After independently coding the data, the researchers met to compare interpretations, discuss divergent codes, and resolve disagreements through negotiated consensus following standard collaborative thematic analysis procedures (Braun and Clarke 2006). This process continued until both coders agreed on a stable and coherent final code set. 5.2.6 Newcomer Evaluation Results In this section, we present the results of our study with newcomers and evaluate whether the data support each hypothesis. Before presenting the results, we note that all participants in the documentation group chose to use ChatGPT to support their tasks, while participants in the VisDoc group completed the tasks using VisDoc alone. H1. Task Performance. To compare task performance across two groups (7 participants per group), each participant completed three onboarding tasks, resulting in 21 task attempts per group (7 participants × 3 tasks). Participants using VisDoc completed 20 of 21 tasks, whereas those in the documentation + ChatGPT group completed 13 of 21. The Fisher’s Exact Test (Fisher 1970) confirmed that this difference was statistically significant (p = 0.02, effect size = 0.89). Figure 4 shows the pertask results: VisDoc users achieved near-perfect success across T1–T3, whereas participants in the documentation + ChatGPT group succeeded on 71% of T1 attempts, 57% of T2, and 57% of T3.

Fig. 4: Task success rates for each task (T1–T3). VisDoc group (cyan) and documentation+ChatGPT group (blue).

Participants’ reflections helped explain the lower failure rates in the VisDoc group. They emphasized that VisDoc’s structured, visual layout and guided task flows reduced uncertainty and steered them away from common errors.

Title Suppressed Due to Excessive Length

27

“VisDoc made the steps very clear. I didn’t feel lost at any point,”[P15]. P9 mentioned that “the visual structure helped me understand what to do next without guessing.” Participants also mentioned VisDoc’s error-prevention benefits. For instance, P5 explained, “VisDoc prevented me from making the usual small mistakes,” and P8 shared “I think I would have messed up if I only used the docs. VisDoc guided me through the sequence properly.” In contrast, participants relying on Documentation and ChatGPT frequently encountered issues identified in prior research, including difficulty generating precise prompts, incomplete guidance, and hallucinations (Choudhuri et al. 2024). As P3 described, “[It was] very hard to prompt ChatGPT on what exactly we need help with,” and P13 mentioned “When I get into the depth of this [task], ChatGPT is not able to help me quickly. . . I felt it very misleading; I got lost.” H1. Task Performance Our evaluation supports H1: a higher proportion of participants using VisDoc completed the onboarding tasks compared to those using the original documentation with ChatGPT assistance. H2: Cognitive Load. We administered the NASA-TLX questionnaire and ran Mann–Whitney U tests (McKnight and Najab 2010) on each of its five dimensions to assess cognitive-load differences between conditions. Table 5 presents the resulting p-values, effect sizes (Cohen’s d), and median scores for both groups. It also maps each TLX dimension to the corresponding CTML design strategies, as outlined in Section 4 and Table 3, to contextualize the observed differences. For all TLX dimensions, lower median scores indicate lower cognitive load, whereas higher scores reflect greater cognitive burden. Table 5: Cognitive load results (NASA TLX) with CTML design principles linked to each dimension. TLX

P-value

Effect Size D

Median (D+LLM)

Median (VisDoc)

Diff

CTML Design

Mental Temporal Performance Effort Frustration

0.04* 0.04* 0.009*** 0.48 0.004***

0.56 0.56 0.68 0.20 0.77

75 60 30 70 60

30 20 0 45 10

↓45 ↓40 ↓30 ↓25 ↓50

Seg., Pretrain. Signal., Align. Synch., Individ. Weed., Reduce. Offload.

Effect size D is interpreted following Romano et al. Romano et al. (2006): D ≈ 0.20 = small, D ≈ 0.50 = medium, D ≈ 0.80 = large

Across all five TLX dimensions, we observe a consistent pattern favoring VisDoc. Four dimensions, mental demand, temporal demand, performance, and frustration, show statistically significant differences, each with medium to large effect sizes. While the effort scores were lower in the VisDoc condition, they did not reach statistical significance. Next, we unpack these differences, beginning with mental demand to match the order of NASA-TLX dimensions presented in Table 3.

28

Zixuan Feng∗ et al.

Participants using Documentation+ChatGPT (control) reported substantially higher mental demand than those using VisDoc (median 75 vs. 30; p=0.04; d=0.56). This suggests that newcomers found the traditional documentation, even when augmented with ChatGPT, more cognitively complex to navigate, often due to scattered information, long textual sections, and the need to mentally integrate guidance across multiple sources. “It was difficult to scan large chunks of texts. . . redirections made it worse” [P3]. ChatGPT’s occasional inaccuracies or irrelevant suggestions further increased cognitive effort by forcing participants to verify or reinterpret steps themselves. “Documentation shows the solution in a different way. . . ChatGPT provides it in another way. . . when I mix both, I never get the solution correctly.” [P12]. In contrast, VisDoc’s graph-based task map reduced mental load by applying Segmenting and Pretraining, breaking workflows into smaller, digestible steps and providing an at-a-glance overview. “The graph-based UI with specific nodes for each task helped me understand the workflow”[P5]. “Getting the whole overview very clear from the graph-based approach” [P8]. Temporal demand was also significantly lower in the VisDoc condition (median 20 vs. 60; p=0.04, d=0.56). Participants using Documentation+ChatGPT frequently reported feeling rushed or pressured, likely due to inefficient information retrieval and the trial-and-error nature of prompting ChatGPT. “ChatGPT could give me commands for git. . . but when going deeper. . . it was not able to help me quickly” [P13]. In contrast, VisDoc incorporates CTML principles of Signaling and Aligning, helping users quickly identify relevant information and follow the intended procedural sequence. “It is very easy to use and navigate. . . I could find what I was looking for easily” [P9]. Similarly, P4 noted that the search function helped “locate what I am looking for very easily,” reducing the back-and-forth navigation that typically increases effort in traditional documentation. The largest differences appeared in the TLX Performance dimension (median 0 vs. 30; p=0.009, d=0.68). Participants using Documentation+ChatGPT consistently reported lower perceived success. P13 described this directly: “I felt [ChatGPT] very misleading. . . I got lost.” In contrast, VisDoc’s multimodal design directly supported higher task success through the CTML principles of Synchronizing (aligning visual and textual explanations to reduce errors) and Individualizing (allowing users to choose the modality that best supports their understanding). P5 described “I used text documentation for copying commands, and videos to understand longer explanations.” P10 shared, “[VisDoc] saved me a lot of time. . . I don’t have to go through the whole documentation. I just need to go into the specific feature that I really want to use.” Similarly, P8 highlighted that the graph-based overview made the workflow “very clear,” reducing the likelihood of choosing the wrong path or missing a required step. Although it was not possible to find statistical significance for the effort dimension (p=0.48), the median difference still favored VisDoc (45 vs. 70). This suggests that while both conditions required users to learn unfamiliar workflows, VisDoc reduced unnecessary cognitive work. “simple UI with less

Title Suppressed Due to Excessive Length

29

distraction features [P5], and P15 highlighted that it is “pretty minimal. . . not too much text.”. Others mentioned that VisDoc reduced their effort by removing the need to search through long documents. As P9 explained, “I was able to find what I was looking for easily,” and P10 emphasized that VisDoc saved substantial effort by avoiding repeated scanning of the full contributing file: “I don’t have to go through the whole documentation.”. Frustration showed strong effects (median 10 vs. 60; p=0.004, d=0.77). Participants using Documentation+ChatGPT reported high frustration, often because the documentation was overwhelming, while ChatGPT’s responses lacked precision, context, or accuracy. “Very hard to prompt ChatGPT on what exactly we need,” [P3] and P12 similarly noted “If I don’t write the prompt correctly, it is more difficult”. VisDoc mitigated frustration through the CTML principle of Offloading, which reduces the emotional and cognitive burden by shifting the explanatory load to short videos and audio-enhanced walkthroughs that supplement text. Participants described these multimodal explanations as helping them feel guided rather than stuck. “Text and video documentation were a good help. . . I used text for copying commands, and videos to understand longer explanations”[P5]. “[VisDoc] more useful for me. . . I kind of like having a video option, the video might have been more useful for me”[P14]. H2. Cognitive Load Our evaluation results provide evidence supporting H2: participants using VisDoc experienced lower cognitive load than those using documentation supplemented with ChatGPT. H3: Usability. To evaluate differences in perceived usability between the two conditions, we conducted Mann–Whitney U tests on the aggregated SUS scores. Participants using VisDoc reported substantially higher usability, with a large effect size, than those relying on documentation+ChatGPT (p = 0.005, Cohen’s d = 0.77). Figure 5 presents a per-item comparison of the normalized SUS responses, with negatively worded items reversed so that higher scores uniformly indicate better usability. For each SUS item, we show side-by-side half–violin plots (VisDoc on the left, documentation+ChatGPT on the right), overlaid boxplots (medians and interquartile ranges), and black dots marking the mean scores for each group. Across nearly all items, VisDoc distributions are shifted upward, with higher medians, a more compact spread, and higher means than those of the group using documentation+ChatGPT. For ease of learning, system integration, and confidence in use, VisDoc scores cluster near the upper end of the scale, whereas documentation+ChatGPT responses are more widely dispersed and shift downward. Participants repeatedly emphasized that VisDoc provided a smoother, more intuitive learning experience. Participants stated that “simple UI with less distraction. . . minimizable nodes and expandable options made exploration easier,” [P5] and that “the information is well presented and I was able to find what I was looking for

30

Zixuan Feng∗ et al.

Fig. 5: Item-level SUS comparison using half–violin and box plots for VisDoc (cyan) and documentation+ChatGPT (blue). The Y-axis shows normalized Likert ratings (1–5; higher = better), with negatively worded items reverse-scored. Black dots indicate mean scores for each group. Hollow dots are outliers.

easily” [P9]. Others highlighted that VisDoc felt coherent and well-structured: “I could visualize the hierarchy. . . it saved me a lot of time.” [P10] Conversely, for negative items such as complexity, need for support, inconsistency, and cumbersomeness, VisDoc distributions remain compressed near the “low burden” end of the scale, while documentation+ChatGPT shows higher perceived difficulty and inconsistency. “I felt ChatGPT very misleading. . . I got lost,” [P3] and P6 noted that using ChatGPT created moments where “I needed more steps to figure things out” [P13]. Others noted inconsistency or unpredictability in ChatGPT’s assistance with the documentation, with P7 explaining that ChatGPT was helpful for small fixes but “not enough to follow the whole workflow.” ” H3. Usability Our evaluation results provide evidence supporting H3: participants reported higher perceived usability for VisDoc than for documentation supplemented with ChatGPT.

6 Discussion Barriers to contributing to OSS include not only the lack of information but also the cognitive structure embedded in existing documentation. Most OSS projects provide substantial amounts of text, but this information is often fragmented across different documents and mixed with boilerplate language (Pinho et al. 2024; Fronchetti et al. 2023; Aghajani et al. 2020). As a result, newcomers have difficulty in inferring the contribution workflow (Steinmacher et al. 2014, 2015a, 2016), including where to start, which steps are mandatory, which steps depend on others, and how to avoid contribution-blocking pitfalls. The gap between “information available” and “understandable workflow” cre-

Title Suppressed Due to Excessive Length

31

ates cognitive overload, as mentioned by one of our participants: “Figuring out the structure (...) was a little too technical for me to understand completely. Lots of files so difficult to summarize what is required.” [P7]. Human-Centered AI Design for Documentation VisDoc exemplifies a human-centered approach to AI-powered developer tools by grounding its design in cognitive theory rather than purely technological capabilities. Unlike approaches that maximize AI autonomy, VisDoc uses AI as a cognitive scaffold, restructuring information to align with human processing limitations while preserving maintainer control through human-in-the-loop editing. Our evaluation demonstrates that cognitively-grounded AI restructuring can reduce mental load (median 75→30) and frustration (median 60→10) more effectively than providing pure LLM access. We envision that such human-guided AI solutions will be the most appropriate way of integrating AI into other software engineering tasks, where AI provides the heavy lifting and humans provide critical guidance and validation. Our results with VisDoc ultimately lay the groundwork for more intelligent and trustworthy documentation systems in the future. Reducing Cognitive Load Through Goal-Aligned Workflow Design. The empirical evaluation of VisDoc reveals that reorganizing the documentation around contribution paths yields substantial improvements in task completeness. This shift is not merely a UI improvement; it is an instructionaldesign intervention applied to software engineering practice. Our results show that the CTML-informed strategies implemented in VisDoc directly shaped how newcomers navigated OSS workflows and measurably reduced cognitive overload, as reflected in lower reported mental and temporal demand, lower frustration, and higher perceived performance on the NASA– TLX (Table 5). Segmenting and pretraining address C1 (essential overload) by breaking down dense contribution guides into smaller, goal-based units and providing just enough conceptual grounding upfront, which corresponds to the lower mental demand. Aligning, signaling, and weeding target C2/C3 (extraneous overload) by surfacing the main contribution path and pruning redundant or poorly ordered details, reduces temporal demand and unnecessary effort. Off-loading via short videos and narrated walkthroughs mitigates C4 (modality overload) by helping reduce frustration when participants confront complex multi-step operations. Finally, synchronizing visuals with text and allowing learners to choose their preferred modality supports C5 (representational overload) by reducing working-memory burden and increasing perceived performance. “It is very easy to use and navigate. As a user, the information is well presented and I was able to find what I was looking for easily” [P9]. The Limits of LLM Assistance in OSS Onboarding. In the control group of our experiment, where all participants chose to use ChatGPT alongside the official contributing files, we observed that the LLM’s support was not enough to close the gap with VisDoc. ChatGPT could occasionally help with local issues, such as recalling Git commands, rephrasing error messages, or clarifying small code snippets, but our results show that it did little to resolve onboarding problems. Participants still had to reconcile misaligned

32

Zixuan Feng∗ et al.

representations (long, dense documentation versus ChatGPT’s step suggestions), which increased both mental and temporal demand and often led to confusion or outright failure: “I got lost. ChatGPT made me lost. I think I could not give it a proper prompt, so I lost my path. ChatGPT could give me commands for git if I ask properly, but when I go into the depth of some documentation, chat is not able to help me quickly get what I’m doing.” [P16]. We observed that higher TLX scores for mental demand, temporal demand, and frustration in the documentation+ChatGPT group, combined with lower task success and SUS scores, indicate that ChatGPT frequently added an extra layer of prompting, verification, and error-checking work rather than reducing cognitive load. Therefore, we observed that LLMs can be helpful micro-level assistants, but without an explicit representation of the workflow, they do not repair the structural deficiencies of OSS documentation and can even amplify overload when their outputs must be constantly interpreted, checked, and integrated. Implications for OSS Projects and Tool Builders. Our findings offer actionable insights for OSS maintainers, documentation authors, and tool builders who seek to improve documentation, onboard newcomers, and sustain community growth. VisDoc is open source7 ), but it is important to note that maintainers do not need VisDoc to start their restructuring journey and can start incorporating some of the CTML strategies piecemeal with other tools. For OSS maintainers. Our results show that simply supplementing LLMs to existing documentation is insufficient for supporting newcomers. We recommend that maintainers supplement existing documentation with lightweight task graphs or workflow summaries that make contribution prerequisites, dependencies, and required steps explicit. Maintainers can use VisDoc or other LLM tools (e.g., Gemini, ChatGPT) (Team et al. 2023; Rahman et al. 2025) to generate draft workflow diagrams that maintainers can refine. For documentation authors and technical writers. Our evaluation demonstrates that CTML offers an actionable, theory-backed approach to restructuring OSS documentation. Dense files can be reorganized into modular units (segmenting), primary contribution paths can be highlighted (signaling), outdated sections can be removed (weeding), and complex tasks can be supported with multimodal explanations (off-loading). A first step is a mindset shift: documentation is not just a static reference artifact, but part of a cognitive system that newcomers must actively navigate. This means organizing content into logical, task-oriented sections instead of one long narrative, and using visuals or diagrams to illustrate workflows where possible. Projects that prefer to continue using text-based Markdown-only contributing files can still use CTML strategies to implement lightweight changes. For example, projects can: (1) restructure CONTRIBUTING.md files into a taskdriven format by splitting them into smaller sections such as “Create your first pull request,” “Fix a documentation issue,” or “Add a new example script”, and 7 https://github.com/EPICLab/visdoc_framework; visual_doc_demo

https://github.com/EPICLab/

Title Suppressed Due to Excessive Length

33

provide the ordered steps and prerequisites for each; (2) signal the primary contribution path by adding a short “Start Here: First Contribution Pathway” section at the top of the file, outlining the sequence from environment setup to submitting a first PR and linking directly to those workflow-specific guides; (3) weed outdated installation steps, redundant explanations, or legacy rules that accumulate over time and cause newcomers to second-guess which instructions are still relevant; and (4) off-load complex tasks by embedding screenshots that demonstrate steps (e.g., running tests, creating a branch, submitting a PR), or an annotated example of a PR checklist, allowing newcomers to visualize the workflow rather than mentally reconstruct it. We also recognize that documentation teams might consider collaborating with AI systems to generate such supplementary materials quickly, then curating and editing the output. For Tool builders. Our findings from the experiment open a new design direction: workflow-aware AI for OSS onboarding: LLMs work at a micro-level of assistance but struggle with macro-level guidance. As our results show, GPT-like systems cannot reconstruct the contribution workflow, cannot infer prerequisite relationships, and cannot reliably guide newcomers through multi-step OSS processes. They provide fragments of help, not a coherent pathway. This suggests that future onboarding tools should not treat LLMs as standalone prompt-and-respond assistants, but instead pair them with explicit workflow models that encode the project’s actual contribution steps, dependencies, and CI/CD requirements. Integrating contribution graphs, dependency metadata, or repository-level signals into AI-driven support offers a blueprint for workflow-aware AI systems that complement, rather than replace, structured onboarding artifacts.

7 Limitations We acknowledge that our study has limitations. Sample Size and Participant Diversity. Our user study sample (N=14) was primarily drawn from a university setting and predominantly male. This limited sample may not reflect the broader OSS newcomer population’s diversity in experience, background, cultural context, or gender. Additionally, while the predominance men (11 out of 14 participants) matches the usual distribution of OSS projects (Trinkenreich et al. 2022b), it may misrepresent some OSS communities or the documented experiences of underrepresented groups, whose onboarding challenges may differ from those of majority groups (Trinkenreich et al. 2022b, 2025, 2022a). We also did not evaluate VisDoc’s accessibility for differently-abled users (e.g., those using assistive technologies). Future work should involve larger and more diverse participant pools (including industry and self-taught developers, varied geographic and linguistic backgrounds, and underrepresented demographics) and assess accessibility to ensure findings generalize to diverse populations. Single-Project Evaluation with Controlled Tasks. We evaluated VisDoc using a single OSS project (Transformers) across three predefined on-

34

Zixuan Feng∗ et al.

boarding tasks. While we selected a project different from our development case (Kubernetes) to assess transferability, this may not be enough since OSS projects differ in size, domain, contribution workflow, documentation style, and community norms. Therefore, results may not readily generalize to other contexts. Moreover, our researcher-selected tasks completed under fixed time frames do not fully capture real onboarding, which often unfolds over days or weeks with newcomers self-selecting tasks, interacting with maintainers, and navigating community norms. Thus, further evaluation across multiple projects and in more realistic, longitudinal settings is needed to establish VisDoc’s effectiveness in diverse, real-world scenarios. While controlled studies offer internal validity benefits, more studies are necessary to focus on realism, which is a natural trade-off in controlled experiments (McGrath 1981). Dependence on Existing Documentation Quality. VisDoc’s effectiveness is inherently limited by the quality of the project documentation it restructures. It only reorganizes existing content, not generating new information or correcting inaccuracies. If the source documentation is incomplete, outdated, or inconsistent, VisDoc will produce suboptimal outputs (and may even propagate errors). Many OSS projects have minimal or poorly maintained documentation, which means our approach cannot help unless a baseline level of accurate, up-to-date content is available. In such cases, improving the documentation’s content (manually or with other AI assistance) would be necessary before a restructuring tool like VisDoc can provide cognitive benefits. Our work focused on reducing cognitive overload in existing documents rather than addressing fundamental documentation completeness or correctness, which has been the focus of other studies (Correia et al. 2024; McBurney and McMillan 2014; Yang et al. 2025; Imani et al. 2024; Fronchetti et al. 2023). A potential future research avenue work could investigate how workflow-aware approaches might be combined with LLM-assisted content improvement (Khan and Uddin 2022) to support documentation completeness and cognitive load reduction. GenAI Pipeline Accuracy and Hallucination Risks. VisDoc’s backend relies on GPT-4o for segmentation refinement, topic inference, redundancy detection, task sequencing, and instructional script generation. Despite employing RAG-based grounding and iterative refinement, LLMs can produce inaccuracies or hallucinated content. The model may misinterpret steps or introduce subtle errors that deviate from the actual documentation. We did not formally measure error rates across different projects or documentation styles. While our human-in-the-loop design (maintainer review of VisDoc outputs) mitigates these issues, it does not eliminate them—maintainers must have the expertise and diligence to catch AI errors. Therefore, the LLM can provide initial content, but human oversight is key. Future work should investigate error-detection mechanisms, uncertainty quantification in generated outputs, and strategies to facilitate human expert review. Cost and Resource Requirements. Generating VisDoc outputs involves computational and financial costs. Our pipeline uses GPT-4o API calls for multiple processing stages, Claude Computer Use for demonstration generation, and ElevenLabs for audio narration synthesis. The generation of mul-

Title Suppressed Due to Excessive Length

35

timedia content requires computational resources and human effort for editing and quality verification. These resource requirements may create barriers to adoption for OSS projects operating with limited or no funding. Alternative implementations using local LLMs or open-source models might reduce costs and can be explored in future work. Maintenance Burden. OSS documentation evolves continuously. VisDoc’s static restructuring approach requires regeneration whenever source documentation changes significantly. We did not evaluate the maintenance burden associated with keeping VisDoc outputs synchronized with evolving documentation, the frequency with which regeneration would be necessary in active projects, or maintainer willingness to invest effort in ongoing VisDoc maintenance since our goal was to validate the approach. Future work can investigate VisDoc’s operational costs. Short-Term Evaluation Without Longitudinal Assessment. Our evaluation examined newcomers’ performance and perceptions in a single session (about 90 minutes per participant), so we only assessed short-term effects. While we observed immediate improvements in task success, cognitive load, and perceived usability, we have no data on long-term outcomes. It remains unknown whether using VisDoc translates into sustained contributor engagement, higher newcomer retention, or better integration into the project community over time. Additionally, we did not assess whether participants would continue using VisDoc for subsequent tasks, whether they would recommend it to other newcomers, or whether it would affect their likelihood of making additional contributions. We also recognize that long-term studies introduce their own challenges. Real OSS projects evolve continuously, and any longitudinal deployment would be shaped by confounding factors such as popularity, release cycles, active maintainer involvement, contributor traffic, and external events (e.g., deadlines, organizational changes) (Feng et al. 2025a). Limited Comparison Baseline. Our controlled study compared VisDoc against the original CONTRIBUTING.md documentation supplemented with optional ChatGPT usage. While this baseline reflects increasingly common realworld practice, it represents only one point in the design space of onboarding support tools. We did not compare against other automated documentation systems, specialized onboarding tools, mentorship-based interventions, interactive tutorials, or alternative AI-augmented approaches such as repositorygrounded conversational agents Correia et al. (2024) or automatically generated workflow diagrams. Additionally, we did not evaluate combinations of approaches, such as integrating VisDoc with mentorship programs or using it alongside conversational agents.

8 Conclusion and Future Work In this paper, we introduce VisDoc, a workflow-aware documentation prototype grounded in the Cognitive Theory of Multimedia Learning (CTML) to reduce cognitive overload in OSS onboarding. Unlike traditional prose doc-

36

Zixuan Feng∗ et al.

umentation, VisDoc restructures contributing files into a task graph, applies CTML strategies (segmenting, signaling, weeding, off-loading), and offers multimodal, step-aligned explanations. Through a between-subject study comparing VisDoc with the original CONTRIBUTING documentation (with optional ChatGPT use), we show that VisDoc significantly improves task completion, reduces mental and temporal load, and yields higher usability scores. More broadly, this work contributes to emerging understandings of effective human-AI collaboration in software engineering. As AI capabilities continue to advance, the critical design question is not “what can AI do?” but “how should AI capabilities be structured to complement human cognitive strengths and compensate for human limitations?” VisDoc demonstrates that theorygrounded design can outperform general AI assistance for complex comprehension tasks. This finding suggests promising directions for human-centered AI in software engineering: systems that combine the flexibility of large language models with the structure of domain-specific workflow models, creating AI tools that enhance human sensemaking in software development contexts. Regarding future work, we plan to integrate LLM-based documentation improvement with VisDoc’s workflow-aware design. Our current system reorganizes existing text but cannot compensate for missing, outdated, or inaccurate content. We plan to leverage LLMs to improve completeness, accuracy, consistency, and coverage in OSS documentation by proposing missing steps, identifying ambiguous or contradictory sections, or generating draft examples aligned with project practices. Combining these capabilities with VisDoc’s explicit workflow representation would create a system that (1) improves documentation content quality and (2) automatically integrates that refined content into structured, cognitively efficient workflows. Statements and Declarations Funding. This work was partially supported by the National Science Foundation under Grant Nos. 2303042, 2247929, 2303612, 2235601, and 2303043. Data Availability. The evaluation artifacts are openly available in the supplementary materials (Feng 2025). The VisDoc implementation is provided in the following open-source repositories: https://github.com/EPICLab/visdoc_ framework and https://github.com/EPICLab/visual_doc_demo. User-related data associated with the evaluation of VisDoc cannot be released because of IRB stipulations. References Adejumo EK, Johnson B (2024) Towards leveraging llms for reducing open source onboarding information overload. In: Proceedings of the 39th IEEE/ACM International Conference on Automated Software Engineering, pp 2210–2214 Aghajani E, Nagy C, Vega-Márquez OL, Linares-Vásquez M, Moreno L, Bavota G, Lanza M (2019) Software documentation issues unveiled. In: 2019 IEEE/ACM 41st International Conference on Software Engineering (ICSE), IEEE, pp 1199–1210

Title Suppressed Due to Excessive Length

37

Aghajani E, Nagy C, Linares-Vásquez M, Moreno L, Bavota G, Lanza M, Shepherd DC (2020) Software documentation: the practitioners’ perspective. In: Proceedings of the acm/ieee 42nd international conference on software engineering, pp 590–601 Agrawal V, Lin YH, Cheng J (2022) Understanding the characteristics of visual contents in open source issue discussions: a case study of jupyter notebook. In: Proceedings of the 26th International Conference on Evaluation and Assessment in Software Engineering, pp 249–254 Alreshidi NAK (2021) Effects of example-problem pairs on students’ mathematics achievements: A mixed-method study. International Education Studies 14(5):8–18 Alroobaea R, Mayhew PJ (2014) How many participants are really enough for usability studies? In: 2014 science and information conference, IEEE, pp 48–56 AlShaikh R, Al-Malki N, Almasre M (2024) The implementation of the cognitive theory of multimedia learning in the design and evaluation of an ai educational video assistant utilizing large language models. Heliyon 10(3) Anthropic (2025) Computer use. https://docs.anthropic.com/en/docs/ agents-and-tools/computer-use, accessed: 2025-05-12 Ayres P, Sweller J (2005) The split-attention principle in multimedia learning. The Cambridge handbook of multimedia learning 2:135–146 Azher IA, Seethi VDR, Akella AP, Alhoori H (2024) Limtopic: Llm-based topic modeling and text summarization for analyzing scientific articles limitations. In: Proceedings of the 24th ACM/IEEE Joint Conference on Digital Libraries, pp 1–12 Beeferman D, Berger AL, Lafferty JD (1999) Statistical models for text segmentation. Machine Learning 34:177–210, URL https://api.semanticscholar.org/CorpusID: 2839111 Braun V, Clarke V (2006) Using thematic analysis in psychology. Qualitative research in psychology 3(2):77–101 Brooke J, et al. (1996) Sus-a quick and dirty usability scale. Usability evaluation in industry 189(194):4–7 Brown T, Mann B, Ryder N, Subbiah M, Kaplan JD, Dhariwal P, Neelakantan A, Shyam P, Sastry G, Askell A, et al. (2020) Language models are few-shot learners. Advances in neural information processing systems 33:1877–1901 Cavanagh TM, Kiersch C (2023) Using commonly-available technologies to create online multimedia lessons through the application of the cognitive theory of multimedia learning. Educational technology research and development 71(3):1033–1053 Çeken B, Taşkın N (2022) Multimedia learning principles in different learning environments: A systematic review. Smart Learning Environments 9(1):19 Chandler P, Sweller J (1991) Cognitive load theory and the format of instruction. Cognition and instruction 8(4):293–332 Chattopadhyay S, Feng Z, Arteaga E, Au A, Ramos G, Barik T, Sarma A (2023) Make it make sense! understanding and facilitating sensemaking in computational notebooks. arXiv preprint arXiv:231211431 Choudhuri R, Liu D, Steinmacher I, Gerosa M, Sarma A (2024) How far are we? the triumphs and trials of generative AI in learning software engineering. In: Proceedings of the IEEE/ACM 46th international conference on software engineering, pp 1–13 Clark JM, Paivio A (1991) Dual coding theory and education. Educational psychology review 3(3):149–210 Clark RC, Nguyen F, Sweller J (2011) Efficiency in learning: Evidence-based guidelines to manage cognitive load. John Wiley & Sons Correia J, Nicholson MC, Coutinho D, Barbosa C, Castelluccio M, Gerosa M, Garcia A, Steinmacher I (2024) Unveiling the potential of a conversational agent in developer support: Insights from mozilla’s pdf. js project. In: Proceedings of the 1st ACM International Conference on AI-Powered Software, pp 10–18 Cubranic D, Murphy GC, Singer J, Booth KS (2005) Hipikat: A project memory for software development. IEEE Transactions on Software Engineering 31(6):446–465 Das JK, Mondal S, Roy CK (2025) Why do developers engage with chatgpt in issue-tracker? investigating usage and reliance on chatgpt-generated code. In: 2025 IEEE International Conference on Software Analysis, Evolution and Reengineering (SANER), IEEE, pp 68– 79

38

Zixuan Feng∗ et al.

De Koning BB, Tabbers HK, Rikers RM, Paas F (2009) Towards a framework for attention cueing in instructional animations: Guidelines for research and design. Educational psychology review 21(2):113–140 De-Marcos L, Domínguez-Díaz A (2025) Llm-based topic modeling for dark web q&a forums: A comparative analysis with traditional methods. IEEE Access ElevenLabs (2025) Elevenlabs. https://elevenlabs.io, accessed: 2025-05-12 Feng Z (2025) Restructure this: Using ai to restructure onboarding documents to reduce cognitive overload. DOI 10.5281/zenodo.17859471, URL https://doi.org/10.5281/ zenodo.17859471 Feng Z, Kimura K, Trinkenreich B, Sarma A, Steinmacher I (2024) Guiding the way: A systematic literature review on mentoring practices in open source software projects. Information and Software Technology 171:107470 Feng Z, Kimura K, Trinkenreich B, Steinmacher I, Gerosa M, Sarma A (2025a) Addressing oss community managers’ challenges in contributor retention. ACM Transactions on Software Engineering and Methodology Feng Z, Steinmacher I, Gerosa M, Menezes T, Serebrenik A, Milewicz R, Sarma A (2025b) The multifaceted nature of mentoring in OSS: strategies, qualities, and ideal outcomes. In: 2025 IEEE/ACM 18th International Conference on Cooperative and Human Aspects of Software Engineering (CHASE), IEEE, pp 203–214 Fisher RA (1922) On the interpretation of χ 2 from contingency tables, and the calculation of p. Journal of the royal statistical society 85(1):87–94 Fisher RA (1970) Statistical methods for research workers. In: Breakthroughs in statistics: Methodology and distribution, Springer, pp 66–70 Fronchetti F, Shepherd DC, Wiese I, Treude C, Gerosa MA, Steinmacher I (2023) Do contributing files provide information about oss newcomers’ onboarding barriers? In: Proceedings of the 31st ACM Joint European Software Engineering Conference and Symposium on the Foundations of Software Engineering, pp 16–28 Gackenheimer C (2015) Introduction to React. Apress Ganier F (2004) Factors affecting the processing of procedural instructions: implications for document design. IEEE Transactions on Professional Communication 47(1):15–26 Garousi G, Garousi V, Moussavi M, Ruhe G, Smith B (2013) Evaluating usage and quality of technical software documentation: an empirical study. In: Proceedings of the 17th international conference on evaluation and assessment in software engineering, pp 24–35 Gaughan M, Champion K, Hwang S, Shaw A (2025) The introduction of readme and contributing files in open source software development. In: 2025 IEEE/ACM 18th International Conference on Cooperative and Human Aspects of Software Engineering (CHASE), IEEE, pp 191–202 Gonçales L, Farias K, da Silva B, Fessler J (2019) Measuring the cognitive load of software developers: A systematic mapping study. In: 2019 IEEE/ACM 27th International Conference on Program Comprehension (ICPC), IEEE, pp 42–52 Gorbunova A, Kapuza A, Chen O, Costley J (2025) Rethinking pre-training: cognitive load implications for learners with varying prior knowledge. Frontiers in Psychology 16:1628047 Guizani M, Feng Z, Arteaga E, Kimura K, Mueller D, Díaz LC, Serebrenik A, Sarma A (2025) Community tapestry: An actionable tool to track turnover and diversity in oss. Information and Software Technology p 107871 Hart SG, Staveland LE (1988) Development of nasa-tlx (task load index): Results of empirical and theoretical research. In: Advances in psychology, vol 52, Elsevier, pp 139–183 He T, Saravanan K, Niforatos E, Kortuem G (2025) " a great start, but...": Evaluating llmgenerated mind maps for information mapping in video-based design. In: Proceedings of the Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, pp 1–7 Heimburger L, Buchweitz L, Gouveia R, Korn O (2020) Gamifying onboarding: How to increase both engagement and integration of new employees. In: Advances in Social and Occupational Ergonomics: Proceedings of the AHFE 2019 International Conference on Social and Occupational Ergonomics, July 24-28, 2019, Washington DC, USA 10, Springer, pp 3–14

Title Suppressed Due to Excessive Length

39

Hurst A, Lerer A, Goucher AP, Perelman A, Ramesh A, Clark A, Ostrow A, Welihinda A, Hayes A, Radford A, et al. (2024) Gpt-4o system card. arXiv preprint arXiv:241021276 Ibrahim N, Aboulela S, Ibrahim A, Kashef R (2024) A survey on augmenting knowledge graphs (kgs) with large language models (llms): models, evaluation metrics, benchmarks, and challenges. Discover Artificial Intelligence 4(1):76 Imani A, Radmanesh S, Ahmed I, Moshirpour M (2024) Does documentation matter? an empirical study of practitioners’ perspective on open-source software adoption. arXiv preprint arXiv:240303819 Islam MA, Hasan R, Eisty NU (2023) Documentation practices in agile software development: A systematic literature review. In: 2023 IEEE/ACIS 21st International Conference on Software Engineering Research, Management and Applications (SERA), IEEE, pp 266–273 Kalyuga S (2011) Cognitive load theory: How many types of load does it really need? Educational psychology review 23(1):1–19 Kapoor S, Gil A, Bhaduri S, Mittal A, Mulkar R (2024) Qualitative insights tool (qualit): Llm enhanced topic modeling. arXiv preprint arXiv:240915626 Kasneci E, Seßler K, Küchemann S, Bannert M, Dementieva D, Fischer F, Gasser U, Groh G, Günnemann S, Hüllermeier E, et al. (2023) Chatgpt for good? on opportunities and challenges of large language models for education. Learning and individual differences 103:102274 Khan JY, Uddin G (2022) Automatic code documentation generation using gpt-3. In: Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering, pp 1–6 Kiss C, Nagy M, Szilágyi P (2025) Max–min semantic chunking of documents for rag application. Discover Computing 28(1):117 Ko AJ, LaToza TD, Burnett MM (2015) A practical guide to controlled experiments of software engineering tools with human participants. Empirical Software Engineering 20(1):110–141 Kosower DA, Lopez-Villarejo JJ, Roubtsov S (2014) Flowgen: Flowchart-based documentation framework for c++. In: 2014 IEEE 14th International Working Conference on Source Code Analysis and Manipulation, pp 59–64, DOI 10.1109/SCAM.2014.35 Langchain (2025) Langchain reference docs. https://python.langchain.com/api_ reference/experimental/text_splitter/langchain_experimental.text_splitter. SemanticChunker.html#semanticchunker, accessed: 2025-05-07 Lazar J, Feng JH, Hochheiser H (2017) Research methods in human-computer interaction. Morgan Kaufmann Leßenich O, Sobernig S (2023) Usefulness and usability of heuristic walkthroughs for evaluating domain-specific developer tools in industry: Evidence from four field simulations. Information and Software Technology 160:107220 Levi MD, Conrad FG (1996) A heuristic evaluation of a world wide web prototype. interactions 3(4):50–61 Lewis P, Perez E, Piktus A, Petroni F, Karpukhin V, Goyal N, Küttler H, Lewis M, Yih Wt, Rocktäschel T, et al. (2020) Retrieval-augmented generation for knowledge-intensive nlp tasks. Advances in neural information processing systems 33:9459–9474 Li ZS, Arony NN, Awon AM, Damian D, Xu B (2024) Ai tool use and adoption in software development by individuals and organizations: a grounded theory study. arXiv preprint arXiv:240617325 Lloyd SA, Robertson CL (2012) Screencast tutorials enhance student learning of statistics. Teaching of Psychology 39(1):67–71 Mann HB, Whitney DR (1947) On a test of whether one of two random variables is stochastically larger than the other. The annals of mathematical statistics pp 50–60 Mayer RE (2005a) The Cambridge handbook of multimedia learning. Cambridge university press Mayer RE (2005b) Cognitive theory of multimedia learning. The Cambridge handbook of multimedia learning 41(1):31–48 Mayer RE, Chandler P (2001) When learning is just a click away: Does simple user interaction foster deeper understanding of multimedia messages? Journal of educational psychology 93(2):390

40

Zixuan Feng∗ et al.

Mayer RE, Moreno R (2003) Nine ways to reduce cognitive load in multimedia learning. Educational psychologist 38(1):43–52 McBurney PW, McMillan C (2014) Automatic documentation generation via source code summarization of method context. In: Proceedings of the 22nd International Conference on Program Comprehension, pp 279–290 McGrath JE (1981) Dilemmatics: The study of research choices and dilemmas. American Behavioral Scientist 25(2):179–210 McHugh ML (2012) Interrater reliability: the kappa statistic. Biochemia medica 22(3):276– 282 McKnight PE, Najab J (2010) Mann-whitney u test. The Corsini encyclopedia of psychology pp 1–1 Mousavi SY, Low R, Sweller J (1995) Reducing cognitive load by mixing auditory and visual presentation modes. Journal of educational psychology 87(2):319 Nesbit JC, Adesope OO (2006) Learning with concept and knowledge maps: A meta-analysis. Review of educational research 76(3):413–448 Padoan F, Santos RDS, Medeiros RP (2024) Charting a path to efficient onboarding: The role of software visualization. In: Proceedings of the 2024 IEEE/ACM 17th International Conference on Cooperative and Human Aspects of Software Engineering, pp 133–143 PBC A (2024) Introducing computer use, a new claude 3.5 sonnet, and claude 3.5 haiku. https://www.anthropic.com/news/3-5-models-and-computer-use, 5 min read Pevzner L, Hearst MA (2002) A critique and improvement of an evaluation metric for text segmentation. Computational Linguistics 28:19–36, URL https://api. semanticscholar.org/CorpusID:6048999 Pinho G, Caçula AJ, Costa L, Wiese I, Araújo AA (2024) Challenges and solutions of free and open source software documentation: A systematic mapping study. Simpósio Brasileiro de Engenharia de Software (SBES) pp 114–125 Qiao Y, Hundhausen C, Haque S, Shihab MIH (2025a) Comprehension-performance gap in genai-assisted brownfield programming: A replication and extension. arXiv preprint arXiv:251102922 Qiao Y, Shihab MIH, Hundhausen C (2025b) A systematic literature review of the use of genai assistants for code comprehension: Implications for computing education research and practice. arXiv preprint arXiv:251017894 Rahman A, Mahir SH, Tashrif MTA, Karim MA, Aishi AA, Kundu D, Debnath T, Moududi MAA, Eidmum MZA, Miah ASM, et al. (2025) Comparative analysis based on deepseek, chatgpt, and google gemini: Features, techniques, performance, future prospects. Systems and Soft Computing p 200396 Romano J, Kromrey JD, Coraggio J, Skowronek J, Devine L (2006) Exploring methods for evaluating group differences on the nsse and other surveys: Are the t-test and cohen’sd indices the most appropriate choices. In: annual meeting of the Southern Association for Institutional Research, Citeseer, vol 14 Saito S, Iimura Y, Massey AK, Antón AI (2018) Discovering undocumented knowledge through visualization of agile software development activities: Case studies on industrial projects using issue tracking system and version control system. Requirements Engineering 23:381–399 Santos I, Pimentel JF, Wiese I, Steinmacher I, Sarma A, Gerosa MA (2023) Designing for cognitive diversity: Improving the github experience for newcomers. In: 2023 IEEE/ACM 45th International Conference on Software Engineering: Software Engineering in Society (ICSE-SEIS), IEEE, pp 1–12 Santos I, Felizardo KR, Gerosa MA, Steinmacher I (2024) Game elements to engage students learning the open source software contribution process. In: 2024 IEEE Symposium on Visual Languages and Human-Centric Computing (VL/HCC), IEEE, pp 59–70 Schmuckler MA (2001) What is ecological validity? a dimensional analysis. Infancy 2(4):419– 436 Solbiati A, Heffernan K, Damaskinos G, Poddar S, Modi S, Cali J (2021) Unsupervised topic segmentation of meetings with bert embeddings. arXiv preprint arXiv:210612978 Sorden SD (2005) A cognitive approach to instructional design for multimedia learning. Informing Science 8:263 Sorden SD (2013) The cognitive theory. The handbook of educational theories p 155

Title Suppressed Due to Excessive Length

41

Steinmacher I, Silva MAG, Gerosa MA (2014) Barriers faced by newcomers to open source projects: a systematic review. In: IFIP International Conference on Open Source Systems, Springer, pp 153–163 Steinmacher I, Conte T, Gerosa MA, Redmiles D (2015a) Social barriers faced by newcomers placing their first contribution in open source software projects. In: Proceedings of the 18th ACM conference on Computer supported cooperative work & social computing, pp 1379–1392 Steinmacher I, Wiese I, Conte TU, Gerosa MA (2015b) Increasing the self-efficacy of newcomers to open source software projects. In: 2015 29th Brazilian Symposium on Software Engineering, IEEE, pp 160–169 Steinmacher I, Conte TU, Treude C, Gerosa MA (2016) Overcoming open source project entry barriers with a portal for newcomers. In: Proceedings of the 38th International Conference on Software Engineering, pp 273–284 Sweller J (1988) Cognitive load during problem solving: Effects on learning. Cognitive science 12(2):257–285 Sweller J (2011) Cognitive load theory. In: Psychology of learning and motivation, vol 55, Elsevier, pp 37–76 Tang H, Nadi S (2023) Evaluating software documentation quality. In: 2023 IEEE/ACM 20th International Conference on Mining Software Repositories (MSR), IEEE, pp 67–78 Team G, Anil R, Borgeaud S, Alayrac JB, Yu J, Soricut R, Schalkwyk J, Dai AM, Hauth A, Millican K, et al. (2023) Gemini: a family of highly capable multimodal models. arXiv preprint arXiv:231211805 Tonmoy S, Zaman S, Jain V, Rani A, Rawte V, Chadha A, Das A (2024) A comprehensive survey of hallucination mitigation techniques in large language models. arXiv preprint arXiv:240101313 6 Toscani C, Gery D, Steinmacher I, Marczak S (2018) A gamification proposal to support the onboarding of newcomers in the flosscoach portal. In: Proceedings of the 17th Brazilian Symposium on Human Factors in Computing Systems, pp 1–10 Trinkenreich B, Britto R, Gerosa MA, Steinmacher I (2022a) An empirical investigation on the challenges faced by women in the software industry: A case study. In: Proceedings of the 2022 ACM/IEEE 44th International Conference on Software Engineering: Software Engineering in Society, pp 24–35 Trinkenreich B, Wiese I, Sarma A, Gerosa M, Steinmacher I (2022b) Women’s participation in open source software: A survey of the literature. ACM Transactions on Software Engineering and Methodology (TOSEM) 31(4):1–37 Trinkenreich B, Feng Z, Choudhuri R, Gerosa M, Sarma A, Steinmacher I (2025) Investigating the impact of interpersonal challenges on feeling welcome in oss. In: 2025 IEEE/ACM 47th International Conference on Software Engineering (ICSE), pp 717– 729, DOI 10.1109/ICSE55347.2025.00056 TrueFoundry (2025) Cognita. https://www.truefoundry.com/cognita, accessed: 2025-05-07 Turzo AK, Sultana S, Bosu A (2024) From first patch to long-term contributor: Evaluating onboarding recommendations for oss newcomers. IEEE Transactions on Software Engineering 51:1303–1318, URL https://api.semanticscholar.org/CorpusID:271038944 Van Der Meij H, Van Der Meij J (2014) A comparison of paper-based and video tutorials for software learning. Computers & education 78:150–159 Van Gog T (2021) The signaling (or cueing) principle in multimedia learning. In: The Cambridge handbook of multimedia learning, Cambridge University Press, pp 221–230 Waldo J, Boussard S (2024) Gpts and hallucination: why do large language models hallucinate? Queue 22(4):19–33 Weick KE, Weick KE (1995) Sensemaking in organizations, vol 3. Sage publications Thousand Oaks, CA Yang D, Simoulin A, Qian X, Liu X, Cao Y, Teng Z, Yang G (2025) Docagent: A multi-agent system for automated code documentation generation. arXiv preprint arXiv:250408725 Yue C, Kim J, Ogawa R, Stark E, Kim S (2013) Applying the cognitive theory of multimedia learning: an analysis of medical animations. Medical education 47(4):375–387

Related documents

Record · ID 204857 · SHA-256 5aaa43ea1a292630
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.