ConceptioArchivearXiv CS
arXiv CSopen access

Weaponizing the Commons: A Taxonomy and Detection Framework of Abuse on GitHub

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
softwarearchitecturesoftwareengineeringtesting
software engineering, software architecture, testing

arXiv:2604.17909v1 [cs.SE] 20 Apr 2026

Weaponizing the Commons: A Taxonomy and Detection Framework of Abuse on GitHub Yuli Cheng

Xiaoyu Zhang

Jiongchi Yu

Xi’an Jiaotong University China [email protected]

Nanyang Technological University Singapore [email protected]

Singapore Management University Singapore [email protected]

Shiqing Ma

Chao Shen∗

Yang Liu

University of Massachusetts, Amherst United States [email protected]

Xi’an Jiaotong University China [email protected]

Nanyang Technological University Singapore [email protected]

Abstract

1

GitHub plays a critical role in modern software supply chains, making its security an important research concern. Existing studies have primarily focused on CI/CD automation, collaboration patterns, and community management, while abuse behaviors on GitHub have received little systematic investigation. In this paper, we systematically review and summarize reported GitHub abuse behaviors and conduct an empirical analysis of publicly available abuse cases, curating a manually labeled dataset of 392 GitHub instances. Based on this investigation, we propose a comprehensive taxonomy that characterizes their diverse symptoms and root causes from a software security perspective. Building on this taxonomy, we develop a unified detection framework capable of identifying all abuse categories across repositories and user accounts. Evaluated on the constructed dataset, the proposed framework achieves high performance across all categories (e.g., F1-score exceeding 89%). Collectively, this work advances the understanding of GitHub abuse behaviors and lays the groundwork for large-scale, systematic analysis of the GitHub platform to strengthen software supply chain security.

Software hosting platforms are foundational to modern software supply chains, supporting code contribution, version control, dependency management, and CI/CD pipelines [22, 36, 39]. Among them, GitHub is the most widely used platform globally, with extensive user activity and development [5, 23, 29]. Investigating security issues on GitHub is therefore of paramount importance for ensuring the integrity and security of the software supply chain. Existing research on GitHub security has primarily focused on development workflows and CI/CD automation [28, 36], developer behavior and collaboration patterns [4, 8], software quality and engineering practices [22, 38], and related aspects. However, abuse behaviors on GitHub, such as activities that intentionally distort platform signals, disrupt normal development processes, or deceive users, have received limited systematic research attention. Existing studies have mostly examined isolated issues such as fake stars [16], without offering a comprehensive understanding of the diverse forms, symptoms, and impacts of abuse behaviors on GitHub. Unlike software vulnerabilities, abuse behaviors exploit the platform’s open collaboration and low-barrier-to-entry features and can pose multi-layered, chain-like risks to the software ecosystem. These abuse behaviors can erode user trust in projects and interactions and disrupt normal development workflows (e.g., interfering with issue and code management), leading to high governance costs and low collaboration quality [14, 24, 25]. Even worse, when attackers deliberately leverage platform mechanisms to spread malicious artifacts or deliver harmful content, the resulting impact can propagate from a single repository to a broader software supply chain, ultimately threatening the security of users and production [13, 16, 19, 32]. Therefore, there is an urgent need to conduct a comprehensive investigation and understand the breadth, characteristics, and symptoms of abuse behaviors on GitHub. To fill this gap, this work systematically investigates eight types of GitHub abuse behaviors by manually reviewing and summarizing relevant literature and reports. Based on this investigation, we propose a comprehensive taxonomy, categorizing the abuse behaviors into four high-level classes according to their impact, and provide detailed descriptions and characteristic symptoms for each type. Building on this taxonomy, we further design a unified detection framework capable of identifying specific abuse behaviors. The proposed framework is evaluated on a labeled dataset consisting of 392

CCS Concepts • Security and privacy → Software security engineering; • Software and its engineering → Software libraries and repositories.

Keywords Abuse Behaviors, Software Supply Chain, Software Security ACM Reference Format: Yuli Cheng, Xiaoyu Zhang, Jiongchi Yu, Shiqing Ma, Chao Shen, and Yang Liu. 2026. Weaponizing the Commons: A Taxonomy and Detection Framework of Abuse on GitHub. In Proceedings of JAWs 2026–ICSE 2026. ACM, New York, NY, USA, 5 pages. https://doi.org/XXXXXXX.XXXXXXX ∗ Chao Shen is the corresponding author.

Permission to make digital or hard copies of all or part of this work for personal or classroom use is granted without fee provided that copies are not made or distributed for profit or commercial advantage and that copies bear this notice and the full citation on the first page. Copyrights for components of this work owned by others than the author(s) must be honored. Abstracting with credit is permitted. To copy otherwise, or republish, to post on servers or to redistribute to lists, requires prior specific permission and/or a fee. Request permissions from [email protected]. JAWs 2026–ICSE 2026, RIO © 2026 Copyright held by the owner/author(s). Publication rights licensed to ACM. ACM ISBN 978-1-4503-XXXX-X/2018/06 https://doi.org/XXXXXXX.XXXXXXX

Introduction

JAWs 2026–ICSE 2026, APRIL 13–14, 2026, RIO

instances, achieving F1-score exceeding 89% across multiple abuse categories. The main contributions are summarized as follows:

2

Cheng et al.

3 A Taxonomy of GitHub Abuse Behaviors 3.1 Overview of the Taxonomy

• We propose a comprehensive taxonomy of eight types of abuse behaviors on GitHub, along with detailed descriptions of their observable symptoms and root causes. • Building on this taxonomy, we design a unified detection framework. Experimental results on a labeled dataset of 392 instances demonstrate its effectiveness. • We publicly release the source code of the proposed framework1 , serving as a foundation for future studies on abuse behaviors and detection techniques on GitHub.

GitHub provides core functionalities including search, automation, project development, and collaborative interactions. Analyzing how abuse behaviors affect these functionalities, we construct a taxonomy with four high-level categories and eight subcategories (Table 1). Each category is accompanied by detailed descriptions of characteristic symptoms, enabling a systematic understanding of how different abuse behaviors affect GitHub’s key functionalities.

Background and Related Works

Attention Hijacking refers to the abuse of GitHub’s repository discovery and search mechanisms. When searching for repositories, users commonly rely on signals such as popularity indicators (e.g., stars and forks), update recency, and keyword-based relevance [9]. Abusers exploit these signals to artificially elevate repository rankings and visibility, thereby misleading users or promoting lowquality or malicious projects. Based on the exploited signals, we further divide Attention Hijacking into four representative subcategories: Fake Stars, Automatic Updates, Keyword Stuffing, and Typo Squatting, as summarized in Table 1. • Subcategory 1 (Fake Stars). Abusers purchase or trade fake stars to rapidly inflate a repository’s apparent popularity and reputation [25, 34]. These artificially boosted repositories are more likely to attract developers or investors, and in some cases are used to lure unsuspecting users into scams, credential theft, cryptocurrency fraud, or the distribution of malicious software. • Subcategory 2 (Automatic Updates). Abusers leverage GitHub Actions to automatically push trivial updates at a high frequency, often by repeatedly modifying a file which is usually called “log” [13, 21]. This activity creates the illusion of active maintenance and disproportionately increases repository visibility, particularly when users sort search results by recent updates. • Subcategory 3 (Keyword Stuffing). Repositories are populated with a large number of popular or trending keywords to manipulate search rankings [19, 21]. These keywords are often unrelated to the repository name or README content, misleading users about the repository’s actual purpose or functionality. • Subcategory 4 (Typo Squatting). Typo Squatting involves creating repositories with names that closely resemble those of popular projects, exploiting common typographical errors or minor naming variations [24, 41]. Users may mistakenly assume these malicious repositories as legitimate or widely used projects.

Open-Source Community Security. With the increasing reliance on open-source software, security issues in open-source communities have attracted sustained research attention. As the world’s largest open-source hosting platform, GitHub has been studied from multiple security perspectives. Prior work has explored the detection of anomalous or malicious commits using repository metadata [14], identified large-scale fake starring behaviors [16], and proposed automated approaches for detecting spam or low-quality content in issue discussions [11]. Beyond GitHub-specific studies, a broader body of research examines security practices, risks, and governance challenges in OSS ecosystems [30, 39]. Software supply-chain security, in particular, has become a central research focus, with extensive work analyzing attack vectors and corresponding defenses [26, 27, 40]. Despite these efforts, systematic investigation of socio-technical abuse behaviors within GitHub remains limited. In particular, there is still no unified taxonomy or detection framework that captures GitHub abuse behaviors, an important gap this work aims to address. Online Misbehaviors. Online misbehavior has been extensively studied across diverse online platforms, including social networks, financial communities, and collaborative systems. In social networking contexts, prior work has proposed comprehensive taxonomies of harmful behaviors such as hate, harassment, and fake accounts, and analyzed their associated user groups and behavioral patterns [12, 32]. A long line of research on Sybil attacks further investigates the structural properties of inauthentic identities and develops effective detection mechanisms [6, 37]. As a result, a variety of mature techniques now exist for detecting spam, fake users, and inauthentic activities in online social networks [10, 42]. More recent studies emphasize that online misbehavior is often intertwined with platform-specific trust and reputation signals. For instance, evidence shows that profile signals and verification mechanisms can be manipulated or misinterpreted across platforms [35]. Similarly, research on domain-specific communities, such as investment platforms, demonstrates that misbehavior manifests differently depending on the underlying social and reputational structures, requiring tailored analysis and mitigation strategies [33]. Collectively, these studies suggest that online misbehavior is not solely content-driven, but is deeply embedded in platform-dependent socio-technical mechanisms.

1 https://github.com/Rasnd-yu/GitHub-Detector

3.2

3.3

Attention Hijacking

Authority Fraud

Authority Fraud targets the perceived credibility of GitHub repositories by exploiting the reputational capital of well-known developers. Rather than manipulating popularity or visibility signals, attackers seek to inflate a project’s apparent trustworthiness, misleading users about its legitimacy or quality. In our study, the predominant form of Authority Fraud on GitHub is Spoofed Contributor. • Subcategory 5 (Spoofed Contributor). GitHub allows repository owners to attribute co-authorship via email addresses in commit messages [3]. Malicious actors exploit this mechanism by impersonating well-known developers as co-authors to enhance the

Weaponizing the Commons: A Taxonomy and Detection Framework of Abuse on GitHub

JAWs 2026–ICSE 2026, APRIL 13–14, 2026, RIO

Table 1: Taxonomy of Abuse Behaviors on GitHub. Category

Symptom

Subcategory Fake Stars Automatic Updates

Attention Hijacking Keyword Stuffing Typo Squatting Authority Fraud

Spoofed Contributor

Spam

Issue Spam Reputation Farming

Reputation Manipulation

Fake Stats

𝑆𝑇 (𝑢, 𝑟 ) = 1 ∧ |SR𝑢 | ≤ 𝑥 1 ∧ 𝐴𝑇 (𝑢, [𝑡𝑠 , 𝑡𝑠 + Δ𝑡]) ≤ 𝜖 𝐿𝑂𝐶𝑚 (𝑟, [𝑡𝑐 , 𝑡𝑐 + Δ𝑡]) 𝐶𝑀 (𝑟, [𝑡𝑐 , 𝑡𝑐 + Δ𝑡]) ≥ 𝑥 2 ∧ ≤𝑦 𝐶𝑀 (𝑟, [𝑡𝑐 , 𝑡𝑐 + Δ𝑡]) n o 𝑘 ∈ K𝑟 𝑅𝑒𝑙 (𝑘, RD𝑟 ) < 𝜃 𝑘 ≥ 𝑥 3   𝑆𝑖𝑚(𝑛𝑖 , 𝑛 𝑗 ) ≥ 𝜃 𝑡 1 ∧ 𝑆𝑖𝑚(RD𝑖 , RD 𝑗 ) ≥ 𝜃 𝑡 2 ∧ 𝑅𝑇 𝑃 (𝑟𝑖 ), 𝑃 (𝑟 𝑗 ) ≥ 𝜙 𝑝1 n o 𝑐 ∈ C𝑟 | 𝐴𝑡ℎ(𝑐) = 𝑢 ≤ 𝑥 4 ∧ 𝑃 (𝑟 ) ≤ 𝜙𝑝2 ∧ 𝑃 (𝑢) ≥ 𝜙𝑝3    𝑖 ∈ I𝑟 ∧ 𝐻𝐿(𝑖) = 1 ∨ 𝐻𝐶 (𝑖) = 1 ∧ 𝑆𝑃 (𝑖) = 1 𝑎 ∈ I𝑟 ∪ PR𝑟 ∧ 𝐴𝑇 (𝑢, [𝑡𝑟 (𝑎) + 𝛿𝑡 , 𝑡𝑟 (𝑎) + Δ𝑡]) ≥ 1     ∑︁ ∃𝑟 ∈ R others : 𝑢𝑟𝑙𝑟 ∈ L𝑢 ∨ 𝐶𝑆 (𝑢) − 𝑆 (𝑟 ) ≥ 𝑥 5, 𝑟 ∈ R𝑢 𝑟

1 2 3 4 5 6 7 8 9 10

𝑆𝑇 /𝐻 𝐿/𝐻𝐶/𝑆𝑃 , boolean function 𝑆𝑇 (𝑢, 𝑟 ) , user 𝑢 starred on repositories 𝑟 S R𝑢 , set of repositories starred by user 𝑢 𝐴𝑇 (𝑢,𝑇 ) , user 𝑢 ’s activity count over 𝑇 𝑡𝑠 , timestamp of star activity 𝑥 1 /𝑥 2 /𝑥 3 /𝑥 4 /𝑥 5 , count threshold 𝜖 , inactivity threshold 𝐶𝑀 (𝑟,𝑇 ) , repository 𝑟 ’s total commit over 𝑇 𝐿𝑂𝐶𝑚 (𝑟,𝑇 ) , total modified lines of code in repository 𝑟 over 𝑇 𝑡𝑐 , timestamp of commit activity

11 12 13 14 15 15 16 17 18

𝑦 , average quantity threshold K𝑟 , set of keywords in the repository 𝑟 𝑅𝑒𝑙 (𝑡,𝑇 ) the relevance score between a short text 𝑡 and a long text 𝑇 RD𝑟 , the README content of GitHub repository 𝑟 𝜃𝑘 , relevance threshold 𝜃 𝑡 1 /𝜃 𝑡 2 , similarity thresholds 𝑆𝑖𝑚 (𝛼, 𝛽 ) , text 𝛼 -𝛽 similarity score 𝑛𝑖 /𝑛 𝑗 , name of GitHub repository 𝜙𝑝1 /𝜙𝑝2 /𝜙𝑝3 , popularity thresholds

perceived legitimacy and reputation [18]. As GitHub does not notify users when they are listed as co-authors, the impersonation may remain unnoticed, increasing the stealthiness of this abuse.

3.4

Spam

The Spam category encompasses abuse behaviors that introduce unwanted or harmful content into GitHub interaction channels, such as irrelevant or malicious posts, bogus pull requests, and automated notifications. Such spam is increasingly generated by automated agents and bot-like accounts, creating noise that developers must filter as part of regular project maintenance [17]. In this work, we focus on spam appearing in repository issue trackers, as it directly affects communication between contributors and can degrade the quality of project interaction and developer experience. • Subcategory 6 (Issue Spam). The content of issues is unrelated to the target repository and instead contains phishing messages, online scams, or links to malicious software. In some cases, abusers automate issue creation using GitHub Actions to distribute deceptive messages (e.g., fake “IMPORTANT” notifications), facilitating large-scale phishing or fraud campaigns [2, 31].

3.5

Reputation Manipulation

On GitHub, each user maintains a public profile representing personal identity and reputation. Profiles expose signals such as contribution history, activity records, and achievements, which influence perceived expertise, credibility, and trustworthiness [1]. However, these signals are not always verifiable, and some users manipulate or fabricate them to artificially enhance influence or visibility. We

19 20 21 22 23 24 25

𝑅𝑇 (𝑎, 𝑏 ) , popularity ratio defined as max(𝑎, 𝑏 )/min(𝑎, 𝑏 ) C𝑟 , set of commits of repository 𝑟 𝑃 (𝑟 )/𝑃 (𝑢 ) , normalized popularity metric (e.g., stars, forks)

𝐴𝑡ℎ (𝑐 ) , author of commit 𝑐 I𝑟 , set of issues in repository 𝑟 𝐻 𝐿 (𝑖 )/𝐻𝐶 (𝑖 ) , issue 𝑖 contains one or more links/commands

𝑆𝑃 (𝑖 ) , issue 𝑖 contains spam or phishing-like content

26 27 28 29 30 31 32

R𝑢 , set of repositories of user 𝑢 P R𝑟 , set of pull requests associated with repository 𝑟 𝑡𝑟 (𝑎) , timestamp when 𝑎 (issue or PR) is closed or merged

𝛿𝑡 , delay threshold after 𝑎 (issue or PR) closure/merge L𝑢 , set of README stat URLs used by user 𝑢 on profile 𝐶𝑆 (𝑢 ) , user 𝑢 ’s claimed star count 𝑆 (𝑟 ) , star count of repository 𝑟

categorize such behaviors as Reputation Manipulation, comprising two subcategories: Reputation Farming and Fake Stats. • Subcategory 7 (Reputation Farming). Reputation Farming inflates apparent user activity by performing low-effort interactions, such as approving or commenting on pull requests and issues that have already been resolved or closed [7, 15]. These actions contribute little substantive value, yet are prominently recorded in profile activity timelines, creating a misleading impression of sustained and meaningful participation. • Subcategory 8 (Fake Stats). Profile-level manipulation can be used to spoof personal credibility and mislead users into trusting a contributor [20]. Such manipulation includes falsely claiming membership in well-known organizations, displaying fabricated achievement badges, or embedding third-party statistic widgets (e.g., github-readme-stats) that reference another user’s account.

4

Framework & Experiment

Based on the taxonomy, we design a unified detection framework that integrates multiple detection strategies tailored to different abuse symptoms. Its effectiveness is evaluated using a small-scale labeled dataset containing 392 instances spanning all categories, and the results are summarized in Figure 1.

4.1

Setup

Framework Implementation. For certain abuse categories, we directly adopt established detection approaches. Specifically, Fake Stars are detected using StarScout [16], which identifies anomalous

JAWs 2026–ICSE 2026, APRIL 13–14, 2026, RIO

Cheng et al.

Figure 1: The performance of the comprehensive detection framework on small-scale artificial datasets star-giving behavior. For Issue Spam, we apply a MLP classifier with TF-IDF vectorizer [11], designed for issue content classification. For the remaining categories, we design tailored detection strategies based on the symptoms summarized in our taxonomy and insights from related studies. For Keyword Stuffing, we leverage the well-established BM25 ranking algorithm to measure the relevance between repository metadata and injected keywords, enabling the identification of anomalous keyword usage. Similarly, Typo Squatting is detected using mature text similarity metrics to identify pairs of repositories with highly similar names but significantly different popularity levels, indicating potential impersonation. Detection of Automatic Updates exploits characteristic temporal patterns. We first retrieve repositories with frequent recent updates via the GitHub API to narrow the candidate set, and then analyze commit frequency and code change quality to distinguish abusive update behavior from legitimate maintenance activity. The remaining categories, Spoofed Contributor, Reputation Farming, and Fake Stats, require more fine-grained and resource-intensive analysis. For example, for Spoofed Contributor, we analyze the activity patterns of reputable contributors to determine whether their identities have been exploited by other actors. In the case of Reputation Farming, we conduct detailed activity-level analyses of target contributors, focusing on interactions with already closed or resolved pull requests and issues. Detection of Fake Stats requires examining individual claims and statements presented on user profile pages and verifying their consistency with ground-truth data. Dataset Structure. As part of our systematic investigation of abuse behaviors on GitHub, we constructed a labeled dataset of 392 abuse instances covering the period since 2020, composed of 310 GitHub repositories and 82 user accounts. The dataset encompasses repositories across industrial software, e-commerce platforms, and both front-end and back-end projects, as well as users exhibiting abuse behaviors in commits, pull requests, issues, and other collaborative interactions. Each instance is manually annotated with its abuse subcategory, and the dataset is balanced across positive and negative samples. Sources include prior academic studies, security reports, and gray literature.

4.2

Result Analysis

As shown in Figure 1, the framework consistently achieves strong performance across all abuse categories, with all metrics exceeding 89%. Notably, Keyword Stuffing and Fake Stats attain the highest F1-scores (93.2% and 93.6%), largely due to the maturity of the BM25 algorithm and the structured nature of these symptoms. In contrast, Issue Spam and Reputation Farming show slightly lower performance with larger Precision-Recall gaps, likely due to their diverse and irregular symptoms, while their lowest metrics still reach 89.5%, keeping the results within an acceptable range. Overall, the framework demonstrates robust performance, suggesting that the task’s structured and regular nature allows a rule-based approach to achieve strong results, obviating the necessity for more sophisticated models.

5

Conclusion and Future Work

In this paper, we present a systematic investigation of abuse behaviors on GitHub by reviewing existing reports and prior research into a comprehensive taxonomy. Building on this taxonomy, we design a unified detection framework that identifies a diverse range of abuse behaviors. Experimental results on a labeled dataset demonstrate the effectiveness of the proposed framework. This work represents an initial step toward systematically characterizing abuse behaviors on GitHub. Looking ahead, we identify two main directions for future research. ❶ Large-scale analysis. The current manually curated dataset is sufficient for validation but limited in scale. We plan to develop a scalable and automated scanning system to extend detection across broader portions of the GitHub ecosystem, enabling continuous data collection, longitudinal studies, and more accurate estimation of abuse prevalence and evolution. ❷ Deeper empirical insights. Large-scale scanning will facilitate fine-grained analyses, such as examining correlations and co-occurrence patterns among abuse categories, identifying temporal trends, and uncovering rare or emerging forms of abuse. These efforts may further refine the proposed taxonomy and contribute to a deeper and more comprehensive understanding of socio-technical abuse dynamics in open-source ecosystems.

Weaponizing the Commons: A Taxonomy and Detection Framework of Abuse on GitHub

References [1] Michael’s Blog 2018. README Badges Are Vulnerabilities. Michael’s Blog. https: //movermeyer.com/2018-06-22-readme-badges-are-vulns/ [2] 2024. This Windows PowerShell Phish Has Scary Potential – Krebs on Security. https://krebsonsecurity.com/2024/09/this-windows-powershell-phish-hasscary-potential/ [3] GitHub Docs 2026. Creating a Commit with Multiple Authors. GitHub Docs. https://docs.github.com/en/pull-requests/committing-changes-to-yourproject/creating-and-editing-commits/creating-a-commit-with-multipleauthors/ Accessed: 2026-01-13. [4] Kamel Alrashedy and Ahmed Binjahlan. 2024. How Do Software Engineering Researchers Use GitHub? An Empirical Study of Artifacts & Impact. arXiv:2310.01566 [cs] doi:10.48550/arXiv.2310.01566 [5] Mohammad Azeez Alshomali, John R. Hamilton, Jason Holdsworth, and SingWhat Tee. 2017. GitHub: Factors Influencing Project Activity Levels. In Proceedings of the 17th International Conference on Electronic Business (ICEB). ICEB, Dubai, UAE, 116–124. [6] L. Alvisi, A. Clement, A. Epasto, S. Lattanzi, and A. Panconesi. 2013. SoK: The Evolution of Sybil Defense via Social Networks. In 2013 IEEE Symposium on Security and Privacy (2013-05). IEEE, 382–396. doi:10.1109/SP.2013.33 [7] Kumar Ashwin. 2024. Reputation Farming in OSS: A Threat to Building Trust. https://krash.dev/posts/reputation-farming/ [8] Mohamed Amine Batoun, Ka Lai Yung, Yuan Tian, and Mohammed Sayagh. 2023. An Empirical Study on GitHub Pull Requests’ Reactions. 32, 6 (2023), 146:1–146:35. doi:10.1145/3597208 [9] Hudson Borges, Marco Tulio Valente, Andre Hora, and Jailton Coelho. 2017. On the Popularity of GitHub Applications: A Preliminary Note. arXiv:1507.00604 [cs] doi:10.48550/arXiv.1507.00604 [10] Ikram Ud Din, Faiza Masood, Ghana Ammad, and et al. 2019. Spammer Detection and Fake User Identification on Social Networks. 7 (2019), 1–14. doi:10.1109/ ACCESS.2019.2918196 [11] Durgesh Firake and Bhushan Wakode. 2025. Machine Learning-Based Spam Filter for GitHub Repository Issues. In Indian Journal of Technical Education (Special Issue), Y. R. M. Rao, Jyoti Sekhar Banerjee, and Rajeshree D. Raut (Eds.). Indian Society for Technical Education, New Delhi, India, 249–256. https://isteconline.in/ Special Issue on Technical Education. [12] Michael Fire, Dima Kagan, Aviad Elyashar, and Yuval Elovici. 2014. Friend or Foe? Fake Profile Identification in Online Social Networks. 4, 1 (2014), 194. doi:10.1007/s13278-014-0194-4 [13] Yehuda Gelb. 2024. New Technique Detected in an Open-Source Supply Chain Attack. Checkmarx. https://checkmarx.com/blog/new-technique-to-trick-developersdetected-in-an-open-source-supply-chain-attack/#:~:text=In%20a%20recent% 20attack%20campaign,crafted%20repositories%20to%20distribute%20malware [14] Danielle Gonzalez, Thomas Zimmermann, Patrice Godefroid, and Max Schaefer. 2021. Anomalicious: Automated Detection of Anomalous and Potentially Malicious Commits on GitHub. arXiv:2103.03846 [cs] doi:10.48550/arXiv.2103.03846 [15] Sarah Gooding. 2024. OpenSSF Warns of Reputation Farming Leveraging Closed GitHub.. Socket. https://socket.dev/blog/openssf-warns-of-reputation-farmingusing-closed-github-issues-and-prs [16] Hao He, Haoqin Yang, Philipp Burckhardt, Alexandros Kapravelos, Bogdan Vasilescu, and Christian Kästner. 2025. Six Million (Suspected) Fake Stars in GitHub: A Growing Spiral of Popularity Contests, Spams, and Malware. arXiv:2412.13459 [cs] doi:10.1145/3744916.3764531 [17] Jan Hensel. 2024. Survey of Automated Agents and Spam on GitHub. https://hensel.dev/papers/github-sbots-analysis-2024/hensel-sbots-githubanalysis-2024.pdf. Online technical report. [18] Matt Kapko. 2022. Fake GitHub Commits Can Trick Developers into Using Malicious Code | Cybersecurity Dive. https://www.cybersecuritydive.com/news/githubcommits-malicious-code/627466/ [19] Solomon Klappholz. 2024. Hackers Are Abusing GitHub’s Search Function to Spread Malware. IT Pro. https://www.itpro.com/security/hackers-are-abusing-githubssearch-function-to-spread-malware [20] Alik Koldobsky. 2023. 5 Ways Attackers Fool Victims with Fake GitHub Profiles. Medium. https://zero.checkmarx.com/5-easy-ways-attackers-fool-victims-withfake-github-profiles-8e8f4199598a [21] Ravie Lakshmanan. 2024. Beware: GitHub’s Fake Popularity Scam Tricking Developers into Downloading Malware. The Hacker News. https://thehackernews. com/2024/04/beware-githubs-fake-popularity-scam.html [22] Zengyang Li, Yilin Peng, Peng Liang, Apostolos Ampatzoglou, Ran Mo, Hui Liu, and Xiaoxiao Qi. 2022. Technical Debt Management in OSS Projects: An Empirical Study on GitHub. arXiv:2212.05537 [cs] doi:10.48550/arXiv.2212.05537 [23] Antonio Lima, Luca Rossi, and Mirco Musolesi. 2014. Coding Together at Scale: GitHub as a Collaborative Social Network. arXiv:1407.2535 [cs] doi:10.48550/arXiv. 1407.2535 [24] Rounak Majumdar. 2024. Millions of Fake Repositories Found on GitHub: What Developers Need to Know. TechStory. https://techstory.in/millions-of-fakerepositories-found-on-github-what-developers-need-to-know/

JAWs 2026–ICSE 2026, APRIL 13–14, 2026, RIO

[25] Kari McMahon. 2023. The GitHub Black Market That Helps Coders Cheat the Popularity Contest. (2023). https://www.wired.com/story/github-stars-blackmarket-coders-cheat/ [26] Marc Ohm, Henrik Plate, Arnold Sykosch, and Michael Meier. 2020. Backstabber’s Knife Collection: A Review of Open Source Software Supply Chain Attacks. 12223 (2020), 23–43. pubmed:null doi:10.1007/978-3-030-52683-2_2 [27] Chinenye Okafor, Taylor R. Schorlemmer, Santiago Torres-Arias, and James C. Davis. 2022. SoK: Analysis of Software Supply Chain Security by Establishing Secure Design Properties. In Proceedings of the 2022 ACM Workshop on Software Supply Chain Offensive Research and Ecosystem Defenses (New York, NY, USA, 2022-11-08) (SCORED’22). Association for Computing Machinery, 15–24. doi:10. 1145/3560835.3564556 [28] Ziyue Pan, Wenbo Shen, Xingkai Wang, Yutian Yang, Rui Chang, Yao Liu, Chengwei Liu, Yang Liu, and Kui Ren. 2024. Ambush from All Sides: Understanding Security Threats in Open-Source Software CI/CD Pipelines. 21, 1 (2024), 403–418. arXiv:2401.17606 [cs] doi:10.1109/TDSC.2023.3253572 [29] Sk Golam Saroar, Waseefa Ahmed, and Maleknaz Nayebi. 2022. GitHub Marketplace for Practitioners and Researchers to Date: A Systematic Analysis of the Knowledge Mobilization Gap in Open Source Software Automation. arXiv.org. https://arxiv.org/abs/2208.00332v1 [30] Thomas Schlienger and Stephanie Teufel. 2003. Analyzing Information Security Culture: Increased Trust by an Appropriate Information Security Culture. In International Workshop on Trust and Privacy in Digital Business (TrustBus’03) in conjunction with the 14th International Conference on Database and Expert Systems Applications (DEXA 2003). 405–409. doi:10.1109/DEXA.2003.1232055 [31] Ax Sharma. 2024. Clever ’GitHub Scanner’ Campaign Abusing Repos to Push Malware. https://www.bleepingcomputer.com/news/security/clever-githubscanner-campaign-abusing-repos-to-push-malware/ [32] Kurt Thomas, Devdatta Akhawe, Michael Bailey, and et al. 2021. SoK: Hate, Harassment, and the Changing Landscape of Online Abuse. In Proceedings of the 2021 IEEE Symposium on Security and Privacy (SP). IEEE, 247–267. doi:10.1109/ SP40001.2021.00028 San Francisco, CA, USA. [33] Taro Tsuchiya, Alejandro Cuevas, Thomas Magelinski, and Nicolas Christin. 2023. Misbehavior and Account Suspension in an Online Financial Communication Platform. In Proceedings of the ACM Web Conference 2023 (Austin TX USA, 202304-30). ACM, 2686–2697. doi:10.1145/3543507.3583385 [34] Carnegie Mellon University. 2025. Fraudsters Use Fake Stars to Game GitHub Software and Societal Systems Department - School of Computer Science - Carnegie Mellon University. http://cms-staging.andrew.cmu.edu/s3d/news/2025/0903github-stars.html [35] Alejandro E. D. Cuevas V. 2025. Measuring the Impact of Profile Signals on Online Platform Integrity and User Safety. Doctoral dissertation. Carnegie Mellon University, School of Computer Science, Software and Societal Systems Program, Pittsburgh, PA, USA. Thesis Committee: Nicolas Christin (Chair), Bogdan Vasilescu, Sauvik Das, Rolf van Wegberg (TU Delft), Stefan Savage (UC San Diego). [36] Pablo Valenzuela-Toledo, Alexandre Bergel, Timo Kehrer, and Oscar Nierstrasz. 2024. The Hidden Costs of Automation: An Empirical Study on GitHub Actions Workflow Maintenance. arXiv:2409.02366 [cs] doi:10.48550/arXiv.2409.02366 [37] Gang Wang, Manish Mohanlal, Christo Wilson, Xiao Wang, Miriam Metzger, Haitao Zheng, and Ben Y. Zhao. 2012. Social Turing Tests: Crowdsourcing Sybil Detection. arXiv.org. https://arxiv.org/abs/1205.3856v2 [38] Han Wang, Sijia Yu, Chunyang Chen, Burak Turhan, and Xiaodong Zhu. 2024. Beyond Accuracy: An Empirical Study on Unit Testing in Open-Source Deep Learning Projects. 33, 4 (2024), 1–22. arXiv:2402.16546 [cs] doi:10.1145/3638245 [39] Shao-Fang Wen, Mazaher Kianpour, and Stewart Kowalski. 2020. An Empirical Study of Security Culture in Open Source Software Communities. In Proceedings of the 2019 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (New York, NY, USA, 2020-01-15) (ASONAM ’19). Association for Computing Machinery, 863–870. doi:10.1145/3341161.3343520 [40] Laurie Williams, Giacomo Benedetti, Sivana Hamer, Ranindya Paramitha, Imranur Rahman, Mahzabin Tamanna, Greg Tystahl, Nusrat Zahan, Patrick Morrison, Yasemin Acar, Michel Cukier, Christian Kästner, Alexandros Kapravelos, Dominik Wermke, and William Enck. 2025. Research Directions in Software Supply Chain Security. 34, 5 (2025), 146:1–146:38. doi:10.1145/3714464 [41] Ofir Yakobi. 2024. Watch the Typo: Our PoC Exploit for Typosquatting in GitHub Actions. Orca Security. https://orca.security/resources/blog/typosquatting-ingithub-actions/ [42] Dong Yuan, Yuanli Miao, Neil Gong, Zheng Yang, Qi Li, Dawn Song, Qian Wang, and Xiao Liang. 2019. Detecting Fake Accounts in Online Social Networks at the Time of Registrations. 1423–1438. doi:10.1145/3319535.3363198

Related documents

Record · ID 120584 · SHA-256 35094576ece4d1eb
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.