A Tool-Augmented, GPT-4 Chatbot for Real-Time Repository Data Analysis Muhammad Jawad Chowdhury
Md. Sakib Khan
Computer Science and Engineering Islamic University of Technology Dhaka, Bangladesh Email: [email protected]
Computer Science and Engineering University of Dhaka Dhaka, Bangladesh Email: [email protected]
arXiv:2609.07586v1 [cs.AI] 7 Sep 2026
Abstract—Software repositories contain vast amounts of data on code contributions, bug reports, and project activities, yet this information remains challenging for non-technical stakeholders and developers to access due to limited expertise in querying repositories. To address this, we introduce a novel chatbot architecture leveraging OpenAI’s GPT-4 model for automated extraction and analysis of repository data. In contrast, our architecture takes a structured path first by parsing the user’s query to extract relevant parameters, then selecting the correct tool to employ based on that analysis, and finally invoking the GPT-4 model to create a highly detailed response. In contrast to previous work based on multi-component systems with embedding models and document retrievers, our architecture inverts the process by relying on prompt engineering and tool selection to fit with the query intent. To validate our approach, we conducted experiments on various question types, including (Issues, Pull Requests, Commits, Compound Questions, and General Repository Information) evaluating our target prompts’ ability to improve the accuracy of responses from the model. Beyond demonstrating the utility of this architecture to a diverse set of users, our findings suggest that this architecture can make repository data more accessible to technical and non-technical audiences through the production of actionable insights. Index Terms—Large language models, software repositories, GitHub, tool-augmented chatbot, repository mining, prompt engineering
I. I NTRODUCTION
In the context of open-source software, platforms such as GitHub act as an essential source and database of large amounts of data created by multiple contributors [1][2][3]. This data should be analyzed to gain more insights that would improve the quality of software and also increase our understanding of the software development process [4][5][6]. However, analyzing repository data is challenging because repositories contain many interconnected elements, including commits, pull requests, issues, and contributor activities. Extracting useful insights from these elements can be timeconsuming and requires technical expertise. [7][8][9]. Also, the involvement of non-technical stakeholders in the analysis of the repository has numerous benefits [10] such as better collaboration [4], clearer communication, and better decision-making [11][12]. Yet existing barriers prevent these stakeholders from accessing or understanding repository data. The most crucial property of iteration mechanisms is
to generate summary reports [13] on demand so that nontechnical users can monitor project progress, identify risks, and make data-driven decisions without needing to rely solely on developers [14][15]. Large Language Models (LLMs) on the other hand, have rapidly evolved to become central to a wide array of applications, ranging from natural language processing tasks like translation and summarization to creative writing [16], content generation [17], and even healthcare [18] and customer support [19]. These versatile models leverage vast amounts of training data to produce human-like responses, making them valuable tools across industries [20][21]. Thus, researchers and developers are beginning to harness LLMs for in-depth software design analysis, highlighting LLMs’ growing role in advancing the capabilities and efficiencies within software development processes [22][23][24][25]. In line with this, our research introduces a chatbot solution for GitHub repositories [7][26], designed to simplify the process of querying and analyzing repository data, thereby enhancing the software development process for both technical and non-technical users [27][28][29][30]. The key contributions of our chatbot are as follows: 1) We propose a tool-augmented architecture that parses natural-language repository queries, extracts relevant parameters, selects appropriate tools, and retrieves realtime data through the GitHub REST API. 2) We construct an evaluation dataset of 80 repositoryrelated questions across pull requests, commits, issues, compound queries, and general repository information, including out-of-scope questions for testing abstention behavior. II. R ELATED W ORK A. LLMs in Software Engineering Recently, LLMs have been used in several software engineering contexts such as code generation [31] [32], program repair [33] [34], and code summarization [35][36]. For example, Fan et al. [37] used Automated Program Repair (APR) techniques to demonstrate that LLMs, e.g., Codex, are better at generating correct fixes than traditional tools when applied to programming tasks. Likewise, Wang et al. [38] present
’CodeT5+’, an LLM that supports a variety of downstream code tasks. Other recent works have taken LLM applications beyond research to the area of software engineering. In a paper by Feng et al. [39], ’CodeBERT’ an LLM specifically trained to increase semantic code understanding for tasks like function extraction and code-to-text generation was developed. Ahmed et al. [36] explored the ability of prompt engineering for the LLM to improve accuracy in code summarization tasks and found that a carefully crafted prompt can substantially improve automated documentation generation. B. Software Engineering Chatbots Abdellatif et al. [26] proposed MSRBot, a chatbot that enables developers to submit natural language queries to retrieve answers from the repository data. They found that MSRBot helped participants complete the tasks more quickly and accurately than manual methods. Xu et al. [27] also proposed AnswerBot, a chatbot to summarize stack overflow multi-answer posts to help developers get concise, short, and relevant answers. Most recently Yu et al. [40] presented CodeMaster, a chatbot that answers code-related questions using CodeT5, showcasing the appearance of LLMs in chatbots that answer complex queries for developers. Other work has also attempted to enable chatbots to answer developer questions through information retrieval. Gottipati et al. [41] have developed a semantic search engine that retrieves answers from threads related to software, while Bradley et al.[42] have built a conversational assistant relying upon the Amazon Alexa platform to automate developer tasks such as creating pull requests. III. A PPROACH A. User Interface: The interaction begins with a user interface, where the user submits a single input Q consisting of a GitHub repository URL Urepo and an accompanying query related to the repository. Once submitted, the system constructs an input I to be processed by the model, composed of three key components: system messages Msys , summarized messages Msum , and the user’s latest query Quser . 1. Msys : Provides context by defining the chatbot’s responsibilities and response format. 2. Msum : A history of previous user queries Q1 , Q2 , . . . , Qn−1 (excluding responses) is appended to the latest query to maintain context in multi-turn conversations, ensuring continuity without exceeding the model’s token limit. 3. Quser : The most recent query submitted by the user, which is combined with Msys and Msum to form the final input I for the chatbot’s response generation. Thus, the complete input I for processing by the chatbot is: I = {Msys , Msum , Quser } The chatbot uses I to generate responses that address the query within the structured context provided by Msys and Msum .
B. System Prompt: The system prompt is a pre-defined instruction given to the chatbot at the beginning of the interaction. It provides a detailed description of the chatbot’s role, its capabilities, and the context in which it operates. The format of the prompt is as follows: ”You are a GitHub Repository Analysis Agent named ’Github-chatbot.’ Answer the following questions as best you can on GitHub Repository: {self.github url}. You have access to the following tools: {tools details}. The system prompt consists of several components: • Agent Identification: The prompt begins by clearly defining the chatbot’s role as a ”GitHub Repository Analysis Agent” named ’Github-chatbot’. This helps establish the chatbot’s identity and purpose, ensuring the user understands that the chatbot is specifically designed for GitHub repository analysis. • Repository URL Placeholder: The placeholder { self.github url } is used to dynamically insert the URL of the specific GitHub repository being analyzed. This ensures that the chatbot knows which repository to focus on when answering user queries. • Tool Access Information: The prompt specifies the tools available to the chatbot for answering queries. The placeholder { tools details } is dynamically populated with a list of tools that the chatbot has access to. These tools could include: – Repository Report Tool: Provides comprehensive insights into the repository’s structure, metadata, and contributions. – Commits/Issues/PR Report Tool: Focuses on retrieving detailed information about commits, issues, and pull requests within the repository. C. Query Processing: Given an input query Q, the system first parses and classifies it to identify key components. The analysis focuses on extracting essential elements and selecting the appropriate action based on predefined keywords and structural rules. • Data Classification: Using a keyword-matching function f : Q → C, where C represents data categories, the system directs Q to the relevant GitHub repository data stream. Primary categories include: – Issues, – Pull Requests, and – Commits. • Filter Application: User-specified filters F = {f1 , f2 , . . . , fn } are applied to refine Q, where fi may denote criteria such as issue state (open/closed), author, date range, or result limit. This ensures that the subset QF ⊂ Q aligns closely with user requirements, yielding targeted and relevant results. D. Tool Selection and API Interaction: Based on the result of the query analysis, the system determines whether specialized tools are required to retrieve
No tool dependency
Question Summarize
Summarized History Latest Query
Prompt of LLM
Chunks of Data
Execute Github Tools
Fig. 1. Architecture of the repository interaction – Prompt of LLM: It consists of Summarized history or response we get from GPT-4 with the latest query and System message Data Chunk: The data or query is chunked based on the limit Response Generation: Necessary tools and parameters are recognized through iterative process.
or transform the requested data. A tool in this context refers to a predefined set of operations designed to interact with the repository and extract specific information as per the user’s query. Repository Report Tool: This tool is used for generating comprehensive reports on the repository’s metadata and overall structure. • Commits/Issues/PR Report Tool: This tool focuses on handling specific data points related to Commits, Issues, and Pull Requests (PRs).
This conditional framework provides flexibility for integrating future tools and algorithms as needed, allowing modular expansion of response generation capabilities. IV. DATASET
•
Once the appropriate tool is selected, the system formulates an API request using the input parameters (repository URL, endpoint, scope, and filters). The interaction is handled via GitHub’s REST API, which fetches the required data.
Category Pull Request Commit Compound Issues General Info
Question How many open pull requests are there currently? What is the latest commit message? Show me the number of commits and the number of pull requests. How many open issues are there? Which developers contributed the most code in 2024? TABLE I
T HE TABLE CONSISTS OF ALL THE CATEGORIES OF QUESTIONS WE HAVE IN THE DATASET.
E. Response Generation The response generation process adapts based on whether a tool is selected or not, following an algorithm: Algorithm 1 Response Generation Process Input: Query Q, Interaction History H, Tool Availability Flag T if T = False then No Tool Selected: Generate response using internal knowledge. Response ← GPT(Q, H) else Tool Selected: Generate response with tool-based data. for each retrieved data chunk do Process the chunk and integrate it into the response. end for end if Return Response
We created a custom dataset of 80 questions to evaluate the proposed approach across a range of repository-related query types. The questions were derived from 15 exemplar queries presented by Abdellatif et al. [26], which describe intents and entities used in repository analysis. Drawing on those exemplars, we prompted an AI model to generate 80 semantically relevant variations, preserving the original intents while diversifying entities and phrasings. Figure 2 illustrates the variety of questions by showing how each question begins. These questions were divided into five distinct categories: Pull Requests, Commits, Compound Questions, Issues and General Information. Table I provides examples for each category. Pull Requests assess the model’s ability to understand and interact with GitHub’s Pull Request workflows. Questions in the Commits category evaluate the model’s understanding of commit history and version control. Compound Questions are designed to test the model’s capacity to handle multi-faceted queries that may involve combining information from multiple
of the model and the work involved in answering different classes of questions. Our findings are as follows:
60 50
Instances
sources within the repository. Issues test the model’s ability to interpret and extract relevant information about open or closed issues. Lastly, General Information includes questions about basic repository details and general metadata, which test the model’s ability to provide high-level information about the repository. It is designed to assess model performance and determine if the models select appropriate metrics and GitHub tools. Additionally, we included some unanswerable questions to gauge the model’s awareness and ability to recognize its limitations.
40 30 20 10 0
1
2
3
Iteration Number
4
5
Fig. 3. Bar plot showing the frequency distribution of iteration numbers.
Fig. 2. Sunburst Distribution of the first two words of the dataset.
V. R ESULTS AND A NALYSIS OF Q UERY H ANDLING AND R ESPONSE G ENERATION A. Scope and Limitations Handling One of the key aspects of our evaluation was to understand how the GPT-4 model manages queries that fall outside the scope of the available tools. Specifically, we included several questions that required information or functionalities not provided by the current toolset, as a means of testing the model’s ability to recognize its limitations. Our findings indicate that GPT-4 demonstrated a commendable understanding of its scope. When presented with out-of-scope questions, the model did not attempt to fabricate answers; instead, it accurately recognized that the query was beyond its tool’s capabilities and refrained from generating incorrect responses. This behavior is significant, as it highlights the model’s ability to manage user expectations by acknowledging tool constraints rather than attempting to provide inaccurate or misleading information. Such self-awareness in LLMs can improve user trust, as users are less likely to receive potentially erroneous data when tools are unavailable. B. Iteration Count Analysis We analyzed the number of iterations required for each question category. Figure 3 summarizes the resulting frequency distribution. This analysis shed light on the efficiency
1) Simple Queries: For simple questions like ‘What is the name of the repository?’, the model needed only one iteration, responding quickly without complex processing. 2) Two-Iteration Queries: Most questions required two iterations: the first to determine the appropriate tool (query classification) and the second to generate the final answer. This two-step process was effective for queries closely aligned with the tools’ capabilities. 3) Compound Questions: Compound questions (e.g., “How many issues are open, and who is the repository owner?”) typically required three iterations. The model had to break down, identify, and classify each part, then generate a cohesive response. This approach ensured accurate and complete answers to complex, multi-part queries. 4) Error Handling and Out-of-Scope Queries: In some cases, questions required three iterations for successful data retrieval. If the model struggled to acquire data points initially, a second iteration was needed. When questions were out of scope, iteration counts increased, sometimes up to five, reflecting the model’s attempts to reclassify and process the query before recognizing its limitations. C. Impact of Tool Availability on Accuracy Additionally, we investigated how the availability of tools affected the accuracy of the model’s responses. The model currently uses two primary tools: one for generating comprehensive repository reports and another for answering entityspecific questions involving issues, commits, and pull requests. However, expanding the toolset could make the system more accurate and versatile. For example, additional tools could support specialized repository-analysis tasks, such as issue-sentiment analysis and contribution-trend analysis. A broader toolset could therefore enable the system to answer a wider range of questions and generate more comprehensive responses. These results indicate that tool availability is an important factor in the effectiveness of an LLM-based repositoryanalysis chatbot.
D. Evaluation of Model Accuracy Each generated response was compared with the corresponding ground-truth answer obtained directly from the GitHub repository data. A response was marked correct only when it returned the expected value and correctly applied all requested filters. An out-of-scope response was marked correct when the system explicitly indicated that the available tools could not answer the query. We evaluated the model using a dataset of 80 in-scope and out-of-scope questions. Among these, 14 questions were specifically designed to be out of scope. The model correctly identified its limitations and abstained from answering all 14 questions. For the remaining 66 in-scope questions, the model produced correct answers for 65, achieving an accuracy of 98.48%. The single incorrect response involved a date-range query in which the model failed to select the correct date parameter and instead relied on the current timestamp. Overall, the model achieved high accuracy on in-scope questions and correctly abstained from answering all out-of-scope questions. VI. C ONCLUSION We introduced a chatbot architecture leveraging OpenAI’s GPT-4 to streamline data extraction and analysis in software repositories. Our solution simplifies the retrieval of repository insights, making it accessible to both technical and nontechnical users. We evaluated our prompts using a custom dataset across Issues, Pull Requests, and Commits, finding that precise prompts enhance model accuracy and efficiency. The tool-based approach allowed GPT-4 to recognize its operational limits and signal out-of-scope queries effectively. Our framework highlights the need for additional specialized tools to improve query coverage, especially for complex questions, potentially increasing accuracy. We demonstrate how targeted prompt engineering and strategic tool use can reduce repository data access complexities and improve collaboration in software development. Future work will extend tool capabilities to support a wider variety of queries. R EFERENCES [1] E. Kalliamvakou, D. Damian, K. Blincoe, L. Singer, and D. M. German, “Open source-style collaborative development practices in commercial projects using github,” in 2015 IEEE/ACM 37th IEEE international conference on software engineering, vol. 1. IEEE, 2015, pp. 574–585. [2] H. Borges, A. Hora, and M. T. Valente, “Predicting the popularity of github repositories,” in Proceedings of the The 12th international conference on predictive models and data analytics in software engineering, 2016, pp. 1–10. [3] ——, “Understanding the factors that impact the popularity of github repositories,” in 2016 IEEE international conference on software maintenance and evolution (ICSME). IEEE, 2016, pp. 334–344. [4] G. Gousios, E. Kalliamvakou, and D. Spinellis, “Measuring developer contribution from software repository data,” in Proceedings of the 2008 international working conference on Mining software repositories, 2008, pp. 129–132. [5] F. Chatziasimidis and I. Stamelos, “Data collection and analysis of github repositories and users,” in 2015 6th International Conference on Information, Intelligence, Systems and Applications (IISA). IEEE, 2015, pp. 1–6.
[6] Y. Hu, J. Zhang, X. Bai, S. Yu, and Z. Yang, “Influence analysis of github repositories,” SpringerPlus, vol. 5, pp. 1–19, 2016. [7] S. Abedu, A. Abdellatif, and E. Shihab, “Llm-based chatbots for mining software repositories: Challenges and opportunities,” in Proceedings of the 28th International Conference on Evaluation and Assessment in Software Engineering, 2024, pp. 201–210. [8] Q. Luo, Y. Ye, S. Liang, Z. Zhang, Y. Qin, Y. Lu, Y. Wu, X. Cong, Y. Lin, Y. Zhang et al., “Repoagent: An llm-powered open-source framework for repository-level code documentation generation,” arXiv preprint arXiv:2402.16667, 2024. [9] R. A. Proma and P. Rosen, “Visual analysis of github issues to gain insights,” arXiv preprint arXiv:2407.20900, 2024. [10] A. Capiluppi, K.-J. Stol, and C. Boldyreff, “Exploring the role of commercial stakeholders in open source software evolution,” in IFIP International Conference on Open Source Systems. Springer, 2012, pp. 178–200. [11] H. Maturana, “From being to doing: The origins of the biology of cognition,” Carl-Auer Verlag, 2004. [12] J. McManus, “A stakeholder perspective within software engineering projects,” in 2004 IEEE International Engineering Management Conference (IEEE Cat. No. 04CH37574), vol. 2. IEEE, 2004, pp. 880–884. [13] N. A. A. Khleel and K. Nehéz, “Mining software repository: an overview,” Doktoranduszok Fóruma, p. 108, 2020. [14] K. K. Chaturvedi, V. Sing, and P. Singh, “Tools in mining software repositories,” in 2013 13th International Conference on Computational Science and Its Applications. IEEE, 2013, pp. 89–98. [15] M. A. de F. Farias, R. Novais, M. C. Júnior, L. P. da Silva Carvalho, M. Mendonça, and R. O. Spı́nola, “A systematic mapping study on mining software repositories,” in Proceedings of the 31st Annual ACM Symposium on Applied Computing, 2016, pp. 1472–1479. [16] T. B. Brown, “Language models are few-shot learners,” arXiv preprint arXiv:2005.14165, 2020. [17] A. Radford, J. Wu, R. Child, D. Luan, D. Amodei, I. Sutskever et al., “Language models are unsupervised multitask learners,” OpenAI blog, vol. 1, no. 8, p. 9, 2019. [18] S. Nerella, S. Bandyopadhyay, J. Zhang, M. Contreras, S. Siegel, A. Bumin, B. Silva, J. Sena, B. Shickel, A. Bihorac et al., “Transformers in healthcare: A survey,” arXiv preprint arXiv:2307.00067, 2023. [19] M. McTear, Conversational ai: Dialogue systems, conversational agents, and chatbots. Springer Nature, 2022. [20] X. Hou, Y. Zhao, Y. Liu, Z. Yang, K. Wang, L. Li, X. Luo, D. Lo, J. Grundy, and H. Wang, “Large language models for software engineering: A systematic literature review,” ACM Transactions on Software Engineering and Methodology, 2023. [21] I. Ozkaya, A. Carleton, J. Robert, and D. Schmidt, “Application of large language models (llms) in software engineering: Overblown hype or disruptive change,” 2023. [22] A. Fan, B. Gokkaya, M. Harman, M. Lyubarskiy, S. Sengupta, S. Yoo, and J. M. Zhang, “Large language models for software engineering: Survey and open problems,” in 2023 IEEE/ACM International Conference on Software Engineering: Future of Software Engineering (ICSE-FoSE). IEEE, 2023, pp. 31–53. [23] X. Hu, K. Kuang, J. Sun, H. Yang, and F. Wu, “Leveraging print debugging to improve code generation in large language models,” arXiv preprint arXiv:2401.05319, 2024. [24] S. I. Ross, F. Martinez, S. Houde, M. Muller, and J. D. Weisz, “The programmer’s assistant: Conversational interaction with a large language model for software development,” in Proceedings of the 28th International Conference on Intelligent User Interfaces, 2023, pp. 491– 514. [25] Y. Xiao, X. Zuo, L. Xue, K. Wang, J. S. Dong, and I. Beschastnikh, “Empirical study on transformer-based techniques for software engineering,” arXiv preprint arXiv:2310.00399, 2023. [26] A. Abdellatif, K. Badran, and E. Shihab, “Msrbot: Using bots to answer questions from software repositories,” Empirical Software Engineering, vol. 25, pp. 1834–1863, 2020. [27] B. Xu, Z. Xing, X. Xia, and D. Lo, “Answerbot: Automated generation of answer summary to developers’ technical questions,” in 2017 32nd IEEE/ACM international conference on automated software engineering (ASE). IEEE, 2017, pp. 706–716. [28] I. Beschastnikh, M. F. Lungu, and Y. Zhuang, “Accelerating software engineering research adoption with analysis bots,” in 2017 IEEE/ACM 39th International conference on software engineering: new ideas and
emerging technologies results track (ICSE-NIER). IEEE, 2017, pp. 35–38. [29] Y. Tian, F. Thung, A. Sharma, and D. Lo, “Apibot: question answering bot for api documentation,” in 2017 32nd IEEE/ACM international conference on automated software engineering (ASE). IEEE, 2017, pp. 153–158. [30] M. Wessel, B. M. De Souza, I. Steinmacher, I. S. Wiese, I. Polato, A. P. Chaves, and M. A. Gerosa, “The power of bots: Characterizing and understanding bots in oss projects,” Proceedings of the ACM on Human-Computer Interaction, vol. 2, no. CSCW, pp. 1–19, 2018. [31] F. Lin, D. J. Kim et al., “When llm-based code generation meets the software development process,” arXiv preprint arXiv:2403.15852, 2024. [32] S. Ouyang, J. M. Zhang, M. Harman, and M. Wang, “Llm is like a box of chocolates: the non-determinism of chatgpt in code generation,” arXiv preprint arXiv:2308.02828, 2023. [33] I. Bouzenia, P. Devanbu, and M. Pradel, “Repairagent: An autonomous, llm-based agent for program repair,” arXiv preprint arXiv:2403.17134, 2024. [34] B. Yang, H. Tian, W. Pian, H. Yu, H. Wang, J. Klein, T. F. Bissyandé, and S. Jin, “Cref: An llm-based conversational software repair framework for programming tutors,” in Proceedings of the 33rd ACM SIGSOFT International Symposium on Software Testing and Analysis, 2024, pp. 882–894. [35] T. Ahmed and P. Devanbu, “Few-shot training llms for project-specific code-summarization,” in Proceedings of the 37th IEEE/ACM International Conference on Automated Software Engineering, 2022, pp. 1–5. [36] T. Ahmed, K. S. Pai, P. Devanbu, and E. Barr, “Automatic semantic augmentation of language model prompts (for code summarization),” in Proceedings of the IEEE/ACM 46th International Conference on Software Engineering, 2024, pp. 1–13. [37] Z. Fan, X. Gao, M. Mirchev, A. Roychoudhury, and S. H. Tan, “Automated repair of programs from large language models,” in 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE). IEEE, 2023, pp. 1469–1481. [38] Y. Wang, H. Le, A. D. Gotmare, N. D. Bui, J. Li, and S. C. Hoi, “Codet5+: Open code large language models for code understanding and generation,” arXiv preprint arXiv:2305.07922, 2023. [39] Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin, T. Liu, D. Jiang et al., “Codebert: A pre-trained model for programming and natural languages,” arXiv preprint arXiv:2002.08155, 2020. [40] T. Yu, X. Gu, and B. Shen, “Code question answering via task-adaptive sequence-to-sequence pre-training,” in 2022 29th Asia-Pacific Software Engineering Conference (APSEC). IEEE, 2022, pp. 229–238. [41] S. Gottipati, D. Lo, and J. Jiang, “Finding relevant answers in software forums,” in 2011 26th IEEE/ACM International Conference on Automated Software Engineering (ASE 2011). IEEE, 2011, pp. 323–332. [42] N. C. Bradley, T. Fritz, and R. Holmes, “Context-aware conversational developer assistants,” in Proceedings of the 40th International Conference on Software Engineering, 2018, pp. 993–1003.