Conceptio › Archive › arXiv CS
arXiv CSopen access

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

Noname manuscript No. (will be inserted by the editor)

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

arXiv:2609.14983v1 [cs.SE] 14 Sep 2026

Xiangxi Li · Olivier Nourry · Yoshiki Higo · Raula Gaikovina Kula

Received: date / Accepted: date

Clinical trial number: Not applicable. Abstract Modern software development increasingly relies on using multiple programming languages. Some software packages are published to multiple package ecosystems—such as NPM for JavaScript and PyPI for Python. Little is known about cross-ecosystem packages, especially regarding how they are structured. In this paper, we conduct a large-scale empirical study of over six million packages across six major ecosystems to understand 1) how prevalent cross-ecosystem packages are among all packages, 2) whether there are distinct source code architectural patterns that cross-ecosystem packages use, and 3) whether there are correlations between architectural patterns and project health metrics from GitHub. Results indicate that cross-ecosystem packages constitute a small but important, growing fraction of packages. We identify five distinct architectural patterns. For example, packages that implement code generation from a shared source file or use language bindings are associated with significantly higher community visibility and development activity. Based on our findings, we provide implications for package adopters, maintainers, and researchers. We envision our taxonomy being used for future investigations into several aspects of software development, such as the trade-offs between focusing on one language and translating to other languages using bindings, templating, and wrappers, versus using native code and native functions to support additional languages. Keywords library dependencies, software ecosystems

1 Introduction In the context of ecosystems, packages play a unique role in sustaining not only the project itself, but the ecosystem as a whole. For software libraries, Address(es) of author(s) should be given

2

Xiangxi Li et al.

packages form a supply chain of dependencies, where the failure of a package could lead to disruptions from both downstream and upstream dependents in the ecosystem Boehmke and Hazen (2017). Furthermore, to counter failures, developers have been encouraged to use traceability measures such as a ‘software bill of materials (SBOM)’ 1 to account for the contents of each package in the supply chain. However, this might not be as simple for crossecosystem packages, as the architectures may differ. Furthermore, as reported by Yang et al. Yang et al (2024), like any multilingual software, packages may suffer challenges related to builds, data handling, interoperability, and interfacing (explicit and implicit). To the best of our knowledge, no prior work has conducted a large-scale empirical study to systematically characterize the architectural patterns of cross-ecosystem packages. Thus, we still do not know how cross-ecosystem packages—which include both libraries and applications—are structured in practice to support developers from multiple ecosystems and facilitate cross-language development. In line with related work Tian et al (2021), we define software architecture as the high-level structure of a software system, including the source code. Indeed, Tian et al. Tian et al (2021) found that practitioners view software architecture and source code as intertwined artifacts, and that understanding of their relationship is essential for improving maintainability and reliability. Another related work shows that the representation of architecture may take different perspectives of styles, views, patterns, tactics, and decisions Ali et al (2018). In this study, we operationalize architecture at the repository level : we focus on observable artifacts such as directory structure, file type distribution, and cross-language integration mechanisms (e.g., language bindings, templatebased code generation). This repository-level view is consistent with the SBOM perspective of tracing the concrete source code constituents of a package Tian et al (2021). Hence, in this study, similar to SBOMs, we aim to trace the source code implementations of open-source packages. This encompasses the file structure, covering both source code and configuration files associated with building packages. Especially in the case of cross-ecosystem packages, we argue that developers can benefit from knowing whether the packages they use are implemented in their native programming languages or not, as different languages provide different levels of security, reliability, and performance. In this study, we investigate cross-ecosystem packages across six major ecosystems—Crates.io, Maven Central, NPM, PHP Composer, PyPI, and RubyGems. Our goal is to empirically identify and evaluate the different design strategies used to develop, deploy, and maintain these real-world ecosystem packages. Thus, we first collected and compared the packages of the studied ecosystems (i.e., 6,080,775 packages) to determine the prevalence of crossecosystem packages. We then investigate how these projects support and deploy source code to multiple ecosystems. We aim to answer the following two preliminary questions and two research questions: 1

https://www.cisa.gov/sbom

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

3

– (PQ1) How prevalent are cross-ecosystem packages? We find that cross-ecosystem packages represent a small but notable fraction of all packages, with thousands of repositories publishing to two or more ecosystems simultaneously. – (PQ2) What percentage of cross-ecosystem packages maintain source code in their repositories? We find that a substantial portion of cross-ecosystem packages lack active source code support in their associated GitHub repositories, indicating that many packages are distributed without a maintained multi-language codebase. Building upon these preliminary results, we further investigate the following research questions: – (RQ1) What kinds of source code architectural patterns are employed by cross-ecosystem packages? We identify five distinct architectural patterns, ranging from loosely coupled multi-repository projects to more tightly integrated binding and wrapper implementations. – (RQ2) How do different architectural patterns correlate with project-level health attributes? We find that tighter integration strategies, such as protocol buffers and binding/wrapper projects, are associated with higher community visibility and more active development than loosely structured alternatives. Our results lead us to conclude that cross-ecosystem packages constitute a relatively small but significant proportion of software packages, which has grown since prior research Constantinou et al (2018). From our investigation, we find that the majority of mono-repo (i.e., a single repository hosting all the code base) cross-ecosystem packages (72.4%) do not maintain source code for all registered ecosystems, largely due to distribution-only patterns such as Maven WebJars and PHP Composer wrappers. From our analysis of GitHub health metrics and architectural patterns, we also find that maintainers who use tightly integrated strategies (such as wrapper/binding) to achieve cross-ecosystem integration tend to maintain their projects more actively, showing higher commit, contributor, and star counts than loosely structured alternatives. We make the following contributions: 1. Pattern taxonomy: We propose a taxonomy of five repository-level architectural patterns for cross-ecosystem packages, derived through an iterative open card sorting process and validated through manual and automated analysis of thousands of repositories. 2. Empirical findings: We provide actionable insights for package adopters, maintainers, dependency management tool developers, and researchers studying software supply chains. 3. Replication package: Our mining scripts, analysis tools, and datasets (package names, homepage URLs, repository URLs, GitHub metrics, directory structures, and source file compositions from six major ecosystems) are publicly available to support the replication and extension of this work.

4

Xiangxi Li et al.

2 Background and Related Work In this section, we define key concepts and then situate our work with respect to the literature.

2.1 Background Software ecosystems. A software ecosystem comprises a package management system along with its associated packages, developer community, and supporting infrastructure van den Berk et al (2010). Major ecosystems include NPM for JavaScript, PyPI for Python, Maven for Java, Crates.io for Rust, Packagist for PHP, and RubyGems for Ruby. Each provides a centralized registry where developers publish packages and declare dependencies. Prior work has studied the dependency network evolution Decan et al (2018), package popularity Zerouali et al (2019), and maintenance practices Kula et al (2017); Bavota et al (2015) in these ecosystems. Some work has also documented the dependency vulnerabilities Prana et al (2021) and package abandonment Miller et al (2025); Gaikovina Kula and Robles (2023) across multiple ecosystems. Furthermore, Valiev et al. Valiev et al (2018) found that ecosystem-level factors—such as a project’s position in the dependency network—significantly influence sustained project activity. These studies treat each ecosystem independently; our work, instead, focuses on packages that span multiple ecosystems. Cross-ecosystem packages. A cross-ecosystem package is a software library published to two or more package registries from the same or related source repositories Constantinou et al (2018). This is distinct from a simple port or fork: the same project actively maintains multiple language implementations, often under a unified version numbering scheme and release cycle. Prominent examples include Protocol Buffers, gRPC, and major cloud SDKs such as the AWS SDK, each of which is published to many ecosystems in parallel. These libraries bridge language boundaries and serve as foundational infrastructure in multilingual software systems. Multi-language development. The engineering challenges of maintaining software in multiple programming languages have been investigated through several lenses. Yang et al. Yang et al (2024) analyzed Stack Overflow discussions on multi-language programming and identified recurring issues in language interfacing, foreign function calls, and cross-language data handling. Their results show that language boundaries introduce not only syntactic challenges but also coordination overhead in testing and documentation. Maintaining consistent behavior across language implementations requires additional effort that single-language projects do not face. Furthermore, cross-ecosystem packages can be viewed through the lens of system-of-systems architecture Klein and van Vliet (2013), where each language-specific implementation operates independently but is maintained within a larger unified project. Our study

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

5

complements this perspective by providing large-scale empirical evidence of how cross-ecosystem projects architecturally resolve these challenges in practice.

2.2 Related Work Package ecosystem analysis. Software ecosystems have attracted sustained research attention. Decan et al. Decan et al (2018) performed a comparative study of dependency networks across seven package managers, finding substantial variation in dependency growth, breadth, and fragility across ecosystems. Kula et al. Kula et al (2017) found that most Java developers do not update their dependencies even when newer versions are available, revealing widespread technical debt in the dependency layer. Abdalkareem et al. Abdalkareem et al (2017, 2020) studied the use of trivial packages in NPM and PyPI, finding that developers often depend on packages that implement only minimal functionality. Bogart et al. Bogart et al (2021) examined how 18 open-source ecosystems handle breaking changes, revealing that ecosystems differ substantially in their coordination norms. This difference presents amplified challenges for cross-ecosystem packages that must comply with multiple sets of conventions simultaneously. In this study, we go beyond individual ecosystem analysis to examine packages that are jointly maintained across ecosystems, a dimension that prior ecosystem studies have not addressed. Cross ecosystem package studies. Constantinou et al. Constantinou et al (2018) conducted the first large-scale investigation of cross-ecosystem packages by matching repository URLs across 12 package managers, identifying a small but non-negligible set of packages published across ecosystem boundaries. Kannee et al. Kannee et al (2023) examined the community dynamics of packages spanning multiple ecosystems (NPM, CRAN, Maven, PyPI, and RubyGems) and found that cross-ecosystem packages tend to foster tighter and more interconnected developer communities. Our work extends these studies in three ways: we adopt a consistent unified pipeline across six major ecosystems with a contemporary dataset; we introduce a systematic methodology for detecting whether native source code is actively maintained for each registered ecosystem; and we characterize and compare five distinct repository-level architectural strategies that cross-ecosystem projects employ. Software supply chain security. The security implications of crossecosystem packages are a growing concern. Huang et al. Huang et al (2022) characterized the usages, updates, and security risks of third-party libraries in Java projects, finding that outdated dependencies frequently expose projects to security bugs. Wu et al. Wu et al (2023) further analyzed upstream vulnerabilities in the Maven ecosystem, showing that a significant portion of downstream projects is affected by vulnerable libraries through transitive dependencies. Williams et al. Williams et al (2025) proposed research directions for software supply chain security, highlighting the role of shared repositories and dependency ecosystems in propagating threats. This body of work motivates our study: because cross-ecosystem packages create implicit inter-ecosystem

6

Xiangxi Li et al.

coupling, understanding their architectural organization is a prerequisite for future security analyses. We note that security is not analyzed in this paper; we return to this as a direction for future work in Section 7.

3 Studied Data In this section, we describe the data collection pipeline. Since our study examines cross-ecosystem packages across six ecosystems, we first need to construct a comprehensive dataset that consists of five stages: 1. Mine package lists from six ecosystems. 2. Identify cross-ecosystem packages via URL analysis. 3. Filter out forked, archived, and HTTP-404-error repositories. 4. Mine directory structures. 5. Mine GitHub metrics. Figure 1 shows an overview of our data collection process. We explain how the mined data are used in our preliminary questions (PQs) and research questions (RQs) in the following subsections.

3.1 Mining Package Lists We collect package metadata (Name, Homepage URL, and Repository URL) from six major package ecosystems: Crates.io (Rust)2 , Maven Central (Java/JVM)3 , NPM (JavaScript/TypeScript)4 , Packagist (PHP)5 , PyPI (Python)6 , and RubyGems (Ruby)7 . Data collection was performed in [01 2026]; the dataset reflects a synchronized cross-ecosystem snapshot at that point in time, enabling consistent comparisons across registries. For each package, we extract the associated GitHub repository URL (either the declared repository URL or the homepage URL) and normalize it to a canonical format8 . Across all six ecosystems, we collect 6,080,775 packages linked to 2,489,326 unique normalized GitHub repository URLs. The difference between the total package count and the unique repository count arises from two sources: (i) many packages across different ecosystems point to the same repository (e.g., a project published to both NPM and PyPI often declares a single shared GitHub URL), and (ii) a subset of packages lack a usable GitHub URL and are excluded from cross-ecosystem detection. 2

https://crates.io/ https://central.sonatype.com/ 4 https://www.npmjs.com/ 5 https://packagist.org/ 6 https://pypi.org/ 7 https://rubygems.org/ 8 github.com/owner/repo 3

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

7

Fig. 1: Overview of the dataset preparation

3.2 Identifying Cross-Ecosystem Packages Similar to prior work Constantinou et al (2018); Kannee et al (2023), we define a cross-ecosystem package as a software project that is published to two or more package registries and provides similar or the same service. This analysis is used as the dataset in PQ1 and PQ2.

3.3 Filtering Process We filter out repositories that are not accessible, are forks, or are archived. Inaccessible repositories are those that return HTTP 404 errors (i.e., repositories that are deleted, renamed, or made private).

8

Xiangxi Li et al.

3.4 Mining Directory Structures for Source Code Architecture To analyze the source code architecture of the cross-ecosystem packages, we mine the complete directory structures of their repositories using the GitHub API. The following example shows a sample directory structure of a GitHub repository: .github/ .github/workflows/ .github/workflows/release.yml test/ test/index.js .gitignore LICENSE README.md index.js package.json These directory structures are used in PQ2 and RQ1 to detect the presence of source files (e.g., .py file for PyPI, .js file for NPM, etc.) and analyze the architectural patterns.

3.5 Mining GitHub Metrics To analyze the relationship between architectural patterns and project health (RQ2), we collect six GitHub metrics for all valid repositories: stars, forks, commits, pull requests, issues, and contributors. We group these into two categories: community metrics and activity metrics. All six metrics represent lifetime totals at the time of data collection; they are not rate-based and do not reflect the pace of development during any particular time window. Community metrics—stars, forks, and contributors—capture community reach and team size Zerouali et al (2019). Specifically, stars reflect community interest and perceived project quality; Borges et al. Borges and Valente (2018) found that three out of four developers consider the number of stars before using or contributing to a GitHub project, making stars a widely used proxy for project adoption and reputation. Forks indicate how many developers have copied the repository to extend or adapt it, and prior work has shown a moderate-to-strong correlation between forks and stars, reflecting reuse and collaborative intent Borges et al (2016); Hu et al (2016). Contributors reflect the size of the team that actively commits to the project, capturing the breadth of community involvement. Activity metrics—commits, pull requests, and issues—capture development volume and community engagement. Commits measure the total volume of code changes and serve as a proxy for overall development activity Borges et al (2016). Pull requests reflect active community contribution and collaborative code review, indicating how openly the project accepts external changes. Issues

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

9

capture the degree of community interaction through bug reports and feature requests, and are widely used as an indicator of user engagement and project responsiveness Zerouali et al (2019).

4 Preliminary Analysis In this section, we present two preliminary analyses that demonstrate how we identify cross-ecosystem packages and detect source code support for them.

4.1 How Prevalent Are Cross-Ecosystem Packages? (PQ1) Motivation. Before investigating architectural patterns, we first explore how prevalent cross-ecosystem packages are. In this analysis, we also deploy different ways to detect cross-ecosystem packages that extend beyond related work. Approach. During the initial analysis, we found several patterns in how package URLs are recorded in the registries. Hence, we developed two methods to detect cross-ecosystem packages. – mono-repo This is the classic method used by prior studiesConstantinou et al (2018); Kannee et al (2023). After normalizing all repository URLs, we identify cases where the exact same GitHub URL appears in two or more ecosystem registries. These packages represent single repositories that are published to multiple ecosystems and are therefore identified as a mono-repo approach. – multi-repo. Initial analysis indicates that some project owners maintain a separate repository for each ecosystem. We assume that the repositories belong to the same organization on GitHub. Some projects maintain separate, language-specific repositories under the same GitHub owner (e.g., owner/project-js and owner/project-py). Typically, the programming language name is appended as a suffix. To detect these, we group repositories by owner and strip ecosystem-specific suffixes (e.g., -py, -rust, -node) using curated suffix patterns for each ecosystem. If two or more repositories from the same owner, in different ecosystems, normalize to the same base name after suffix removal, they are flagged as a multi-repo cross-ecosystem package. Results. Table 1 shows the total number of package metadata entries we mined per ecosystem. The “Cross-Ecosystem Packages” column shows how many packages in each ecosystem are identified as cross-ecosystem packages using the two methods mentioned above. The 38,970 cross-ecosystem packages account for 1.57% of the 2,489,326 unique GitHub repositories in our dataset. This number is increasing, as prior work Constantinou et al (2018) reported that 15,389 cross-ecosystem packages accounted for 0.99% of 1,556,300 packages. Among these 38,970 packages:

10

Xiangxi Li et al.

Table 1: Package metadata registry entries mined per ecosystem Ecosystem

# URL of Packages

# Cross-Ecosystem Packages

Crates.io Maven NPM Packagist (PHP) PyPI RubyGems

218,234 763,405 3,749,794 435,167 725,184 188,991

4,088 17,004 25,489 3,975 8,908 2,441

Total Unique Total

6,080,775 2,489,326

61,905 38,970

Table 2: Cross-ecosystem package count by ecosystem and by number of registered ecosystems, with filtering results Ecosystems

2

3

4

5

6

Total

Crates Maven NPM PHP PyPI Ruby

3,304 14,899 22,477 2,343 6,412 1,249

500 1,408 2,139 950 1,701 520

169 417 584 406 510 394

77 242 251 238 247 240

38 38 38 38 38 38

4,088 17,004 25,489 3,975 8,908 2,441

Unique (total)

30,119

5,135

2,221

1,286

209

38,970

Filtering HTTP-404 errorsa Forked repositoriesb Archived repositoriesc Total Removed: Forked ∪ Archived ∪ Errors Valid total Mono-repo (URL-matched) Multi-repo (URL-not-matched)

2,816 731 3,000 6,475 32,494 18,564 13,930

a Repositories deleted, renamed, or made private (HTTP 404 via

GitHub API). b Forked repositories excluded; 71 repositories are counted once here

as both forked and archived. c Archived repositories excluded from further analysis.

– 21,868 (56.1%) packages use a mono-repo, among which 18,564 are valid (non-forked, non-archived, and non-HTTP-404-error). – 17,102 (43.9%) packages use a multi-repo, among which 13,930 are valid (non-forked, non-archived, and non-HTTP-404-error). Table 2 breaks down the “Cross-Ecosystem Packages” totals from Table 1 by the number of ecosystems each package spans (2 through 6), showing how many packages in each ecosystem appear in exactly two, three, four, five, or all six registries simultaneously. NPM is the most commonly involved ecosystem (25,489 entries), followed by Maven (17,004) and PyPI (8,908).

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

11

Cross-ecosystem packages are not prevalent. However, they still represent a growing number when compared to prior work. – Observation 1: Cross-ecosystem packages may span multiple GitHub repositories.

4.2 Do Cross-Ecosystem Packages Maintain Source Code for All Registered Ecosystems? (PQ2) Motivation. Being registered in multiple ecosystems does not guarantee that a repository actually maintains source code for all of them. Before studying architectural patterns in depth (RQ1), we investigate whether cross-ecosystem packages actively maintain source code support for their registered ecosystems by checking if source files appear in their directory structures. Approach. Table 2 shows the filtering process applied to obtain valid repositories for answering PQ2. We use non-forked, non-archived, and nonHTTP-404-error packages in the following analysis, yielding: – 18,564 mono-repo packages. – 13,930 multi-repo packages. We conduct two analyses based on whether packages use multi-repo or monorepo structures. For 13,930 multi-repo packages, each repository is expected to serve a single ecosystem as indicated by the language-specific suffix in its name (e.g., repo-py → Python). To verify this assumption, we check whether the expected programming language actually appears in each repository’s GitHub-reported languages. For each of the 18,564 mono-repo packages, we scan their GitHub repository directory structure for source file extensions associated with each studied ecosystem (using the mined directory structures). We iterate over every file path in the mined directory structure and extract each file’s extension. Then, we look up the extracted file extension against a predefined extension-to-ecosystem mapping to determine which ecosystem(s) are present. Each ecosystem is identified by its characteristic source file types: – PyPI maps to Python files (.py, .pyx, .pxd, .pyi), – Crates to Rust files (.rs), – NPM to JavaScript and TypeScript files (.js, .jsx, .ts, .tsx, .mjs, .cjs, .css, .scss), – Maven to JVM-based files (.java, .scala, .kotlin, .kt), – Ruby to Ruby files (.rb, .rake), – and PHP to PHP files (.php). We exclude common non-source directories (e.g., test, documentation, build, vendor, and cache folders) using approximately 40 exclusion keywords to avoid

12

Xiangxi Li et al.

Table 3: Multi-repo language verification: expected language presence by ecosystem Matched1

Mismatched2

Match Rate3

Crates Maven NPM PHP PyPI Ruby

694 1,975 4,217 2,482 3,168 1,231

5 30 70 27 24 9

99.28% 98.50% 98.37% 98.92% 99.25% 99.27%

Total

13,765

165

98.82%

Ecosystem

1 Matched:

Expected language appeared in the GitHub repository’s language proportions. (e.g. github.com/owner/repo-py has Python) 2 Mismatched: Expected language doesn’t appear in the GitHub repository’s language proportions. (e.g. github.com/owner/repo-py only has JavaScript) 3 Match Rate: The proportion of matched packages among all packages in its corresponding ecosystem.

false detections. A single source file is sufficient to consider an ecosystem as “detected.” Results. We divide the results based on the multi-repository and monorepository analysis. Multi-repo packages. Table 3 shows the per-ecosystem breakdown. Of the valid multi-repo repositories, a majority (98.82%) have their expected language present in the repository’s GitHub-reported languages, as expected. All evaluated ecosystems exhibit match rates exceeding 98%, with Crates, PyPI, and Ruby achieving the highest scores. The remaining 165 repositories exhibit mismatches attributable to three primary causes: (1) empty repositories yielding no language metadata from GitHub; (2) suffixes that coincidentally match a language token but semantically represent non-language artifacts— e.g., the “lang-java” package, which is a CodeMirror language support module implemented in TypeScript; and (3) repositories that have been repurposed for alternative use cases, thereby diverging from their original language designation. Mono-repo packages. We classify the 18,564 mono-repo packages into three categories based on how well their detected ecosystems match their registered ecosystems: – Fully Matched: 5,132 packages (27.6%) — all registered ecosystems are detected in the repository. For example, github.com/kreuzberg-dev/html-tomarkdown 9 is registered in Crates, NPM, PHP, PyPI, and Ruby ecosystems. All source files for all five different ecosystems are found in this single repository. – Partially Matched: 171 packages (0.9%) — two or more (but not all) registered ecosystems are detected. For example, github.com/asimov-modules/asimov9

https://github.com/kreuzberg-dev/html-to-markdown

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

13

apify-module 10 is registered in Crates, NPM, PyPI, and Ruby, but only Crates and Ruby source files are found. – Fully Mismatched: 13,261 packages (71.4%) — one or zero ecosystem source files are detected. For example, github.com/orbitinghail/graft 11 is registered in Crates, NPM, PyPI, and Ruby, but only Crates source files are found, with no evidence of NPM, PyPI, or Ruby source code. The majority of mono-repo cross-ecosystem packages do not maintain native source code for all their registered ecosystems within their repository. Most multi-repo cross-ecosystem packages store a single source code language, while a majority of mono-repo cross-ecosystem packages do not maintain source files for every one of their registered ecosystems. – Observation 2: A majority of mono-repo packages (i.e., 72.4%) do not maintain native source code for all of their registered ecosystems within their repository (combining partially matched 0.9% and fully mismatched 71.4%).

5 What Architectural Patterns Do Cross-Ecosystem Packages Employ? (RQ1) Based on the preliminary results, we move to the empirical study. In this section, we present how we define architectural patterns and how we detect them in our dataset. Motivation. Having established the prevalence of cross-ecosystem packages (PQ1) and the extent of source code presence (PQ2), we now investigate what architectural patterns these packages employ—including the causes of the high mismatch rate observed in PQ2. Understanding these patterns enables tool builders and researchers to better support cross-ecosystem development. Approach. To derive the architectural patterns, we followed an iterative open card sorting process Spencer (2009) inspired by prior work on repository classification. The process was as follows: 1. We drew a stratified random sample of 50 repository URLs from the two subsets established in the preliminary analysis (mismatched mono-repo packages and fully matched mono-repo packages). 2. Each author independently inspected each repository and assigned a short descriptive label capturing the repository’s cross-ecosystem organization strategy. 3. The first author held meetings with the other two authors to compare labels, resolve disagreements, and merge semantically equivalent labels into candidate patterns. In total, 12 such meetings were conducted throughout the process. 10 11

https://github.com/asimov-modules/asimov-apify-module https://github.com/orbitinghail/graft

14

Xiangxi Li et al.

4. If any new pattern was identified in this round, a fresh sample of 50 URLs was drawn and the process repeated from step 2. 5. The process stopped when a full round of 50 samples produced no new patterns. After 41 iterations spanning approximately 12 weeks, no further new patterns emerged, yielding the five patterns presented below. Note that P1: Multi-repo (multi-repo) was identified during PQ1 and was therefore not part of the card sorting sample pool; the card sorting focused exclusively on mono-repo packages, stratified into mismatched and fully matched groups. P2: Distribution-only (distribution-only) emerged from the targeted analysis of PQ2 mismatches and was subsequently confirmed across card sorting iterations. The automated detection rules for each pattern were then designed and validated against the samples accumulated during card sorting. We describe each pattern’s definition and detection rule, and how they relate to the datasets established in prior sections. For the analysis, we tally all the different patterns and perform a detailed analysis of each. Note that we use semi-automatic detection based on the heuristics derived during the card sorting process. It is important to note that these classifications are not mutually exclusive; a package can exhibit more than one pattern. Table 4: Pattern identification results Pattern P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

GitHub owners

Unique Repos

5,816 13,432 2,233 474 1,814

13,930 13,432 2,233 474 1,814

Results. Table 4 summarizes the five patterns that we identified in our card sorting process. Note that for P1: Multi-repo, the “GitHub owners" is much less than its “Unique Repos" because we group the repositories that belong to the same project by owner. Additionally, we find 1,566 repositories that appear in multiple patterns. The most common overlap is between P3: Designated Dir. and P5: Bind/Wrap, with 732 shared repositories. The second most common overlap is between P2: Distribution-only and P5: Bind/Wrap, with 537 shared repositories. We now discuss each pattern in detail. P1: Multi-repo (from Observation 1.) The first pattern is taken from the analysis of PQ1. In this pattern, cross-ecosystem packages have related repositories that share the same GitHub owner but use distinct URLs with language-specific suffixes. Each ecosystem is served by a separate repository rather than a single shared codebase. The 5,816 multi-repository groups span

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

15

13,930 individual repositories, meaning each group contains, on average, 2.4 repositories. NPM (4,149) and PHP (2,506) account for the largest shares, while Crates (694) is the least common. We describe an example of this pattern using the OpenAI SDK package. This package is managed through different URLs but is owned by OpenAI. The naming convention appends the programming language ecosystem as a suffix12 . The package is also registered for NPM (i.e., openai-node), Python (i.e., openai-python), and RubyGems (i.e., openai-ruby). P2: Distribution-only (from Observation 2.) The second pattern was detected during PQ2 and is particularly prevalent among mono-repo packages. Rather than a methodology limitation, this represents a deliberate distribution strategy: the repository intentionally publishes to additional ecosystems via pre-built or repackaged artifacts, without maintaining native source code for those ecosystems. To characterize this pattern, we performed a targeted analysis and manual inspection of 200 sampled mismatched packages, identifying two dominant subpatterns. The first is Maven WebJar/mvnpm publishing: we found that 11,083 (59.7% of all mono-repo packages) were actually WebJar or mvnpm packages (with groupId starting with org.webjars.* or org.mvnpm.*). These packages repackage front-end JavaScript assets for consumption via Maven, without including Java source code. The second sub-pattern involves PHP Composer wrappers: we find 499 (2.7%) packages that contain a composer.json file (distributing front-end assets to PHP via Composer) but no actual .php source files. An example of this distribution-only strategy is the vue package13 . This package is published to both NPM and Maven, but the repository contains only JavaScript; the Maven publication is a WebJar repackaging of the JavaScript build artifact, not a separate Java implementation. P3: Designated Dir. The third pattern is based on the directory structure within the repository. During the card sorting, we found that some crossecosystem packages designated a specific folder for each target ecosystem. Hence, we searched for directories matching the following language-specific names. – PyPI → python, py, etc.; – NPM → js, javascript, etc.; – Crates → rust, rs, etc.; – Maven → java, jvm, etc.; – Ruby → ruby, rb, etc.; – PHP → php, etc. A package is classified under this pattern if two or more language-specific folders are found. To validate the folders, we additionally compute the coverage 12 13

An example for Java is openai-java. github.com/openai/openai-java github.com/vuejs/vue

16

Xiangxi Li et al.

ratio (the proportion of that ecosystem’s source files inside the folder) and classify packages as: – Concentrated-Complete: All ecosystem folders found with each ≥ 80% coverage. – Concentrated-Partial: Some ecosystem folders were found with each ≥ 80% coverage. – Mixed: Some folders were ≥ 80%, while others were ≥ 40% coverage. – Low: All found folders had < 40% coverage.

Table 5: Directory structure classification of all fully matched mono-repo packages (n = 5,132); only packages in the top four rows (with languagespecific folders) constitute P3: Designated Dir. (n = 2,233) Classification

Count (%)

Concentrated-Complete Concentrated-Partial Mixed Low No Designated Directory

659 (12.8%) 90 (1.8%) 1,272 (24.8%) 212 (4.1%) 2,899 (56.5%)

Table 5 presents the directory structure classification applied to all 5,132 fully matched mono-repo packages. Only the packages in the top four categories— those with at least one language-specific folder—are classified as P3: Designated Dir. (totaling 2,233 packages). The “No Designated Directory” row (2,899 packages, 56.5%) represents fully matched packages whose source files are not organized into named ecosystem folders; these packages are not classified as P3: Designated Dir.. Among P3: Designated Dir. packages, 659 (12.8% of the fully matched set) achieve a Concentrated-Complete structure, where every ecosystem’s source files are concentrated (≥ 80%) within its named folder. One example of this pattern is the react-native package 14 . In this example, the package stores the Java/Kotlin code under the directory ReactAndroid/src/main/java/, while the JavaScript is placed in a separate directory at flow-typed/npm/. It has a coverage of 92.7% of the detected matching source files in those directories. P4: Templating The fourth pattern relates to mechanisms used to generate code in multiple languages. Hence, we explored different technologies used for templating. Specifically, packages using interface definition languages (IDLs) or template-based code generation produce language-specific bindings from shared schema definitions. Based on our card sorting, we identified three kinds of templating (i.e., .proto (Protocol Buffers), .thrift (Apache Thrift), and .fbs (FlatBuffers)). 14

github.com/facebook/react-native

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

17

Table 6: P5: Bind/Wrap detection sources Source

# packages

Specific Bindings General Bindings OS Distributions

1,143 735 282

Total (before dedup) Duplicates Total (after dedup)

2,160 346 1,814

As shown in Table 5, we find that 474 packages use IDL-based code generation, with NPM (319) and PyPI (254) being the most common target ecosystems; PHP appears in only 4 packages. An example of such package is flatbuffers 15 . In this case, the package utilizes the .fbs schema files to generate templates for Crates, Maven, NPM, PHP, and PyPI simultaneously.

P5: Bind/Wrap The final pattern is similar to P4, but instead of templates, the mechanism is language-binding technologies. Similar to P2, we also found cases where pre-built binary distributions were published. For P5, we identified the following strategies: – Specific Bindings—detect JSII, WASM, and PyO3/Maturin bindings. – General Bindings—detect binding directories or files16 . – OS Distributions—detect operating system platform directories used for distributing binaries as wrappers17 . Table 6 presents the breakdown by detection source. Specific bindings are the most prevalent, with 1,143 detected packages. Of these, 668 use WASM bindings, 405 use JSII bindings, and 70 use PyO3/Maturin. Besides these, we also detect 735 packages with general binding indicators, followed by 282 OS distribution packages. Across all sources, we record 2,160 detections; after removing 346 duplicates, 1,814 unique packages exhibit binding or wrapper patterns. An example of the final pattern is the next.js package 18 . In this example, the package employs WASM modules to compile Rust code for the Crates ecosystem. 15

github.com/google/flatbuffers Examples include looking for the following keywords: binding/, ffi/, napi/, jni/, cgo/, pybind/, cython/, binding.gyp, .node, .pyd 17 Folders in the repository match combined OS-architectural patterns (e.g., linux-x86_64/, darwin-arm64/, windows-amd64/), indicating prebuilt native binaries distributed as wrappers 18 github.com/vercel/next.js 16

18

Xiangxi Li et al.

We identify five distinct repository-level architectural patterns. Most cross-ecosystem packages do not maintain native source code for all registered ecosystems within their repository; many instead follow a distribution-only strategy. – Observation 3. Cross-ecosystem packages employ different crosslanguage mechanisms—templating, bindings, or wrappers—to publish to multiple ecosystems without duplicating native source code. – Observation 4. Some cross-ecosystem packages organize their repository with language-specific designated directories, co-locating each ecosystem’s source code within a dedicated folder.

6 How Do Architectural Patterns Correlate with Project Health? (RQ2) In this section, we investigate how the architectural patterns identified in RQ1 are associated with project health, which we characterize using six project metrics shown in Table 7. Motivation. Having identified five architectural patterns (RQ1), we now investigate whether different patterns are associated with different project health and community profiles. This analysis helps practitioners understand the observable differences between organizational strategies for cross-ecosystem development. Approach. Using the project metrics we collected from GitHub, we conduct two analyses. First, for each pattern found in RQ1, we compute descriptive statistics (median, mean) for each project health metric. For the second analysis, we independently rank all cross-ecosystem packages across the six metrics to conduct a stratified analysis of the top and bottom deciles. By calculating the frequency of each pattern within these highest and lowest 10% thresholds, we measure the over- or under-representation of each pattern at the metric extremes. Ties at the 10% boundary are broken by random sampling with a fixed seed for reproducibility. Results. Note on interpretation. All results in this section report statistical associations between architectural patterns and project health metrics. Because this is an observational study, we do not claim a causal relationship; the patterns may co-occur with certain health profiles due to confounding factors (e.g., project age, domain, or team size) that we do not control for. For the first analysis (Table 7), we find that P2: Distribution-only packages are associated with the highest star counts, while P4: Templating packages show the highest development activity across forks, commits, pull requests, issues, and contributors. Regarding community metrics, P2: Distribution-only packages have the highest median star count (71), followed by P4: Templating (68). P2:

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

19

Table 7: Descriptive statistics of GitHub metrics per pattern n

Median

Mean

Stars

P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

5,816 13,432 2,233 474 1,814

6 71 19 68 22

221.6 2,015.0 1,585.1 3,750.1 2,486.6

Forks

P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

5,816 13,432 2,233 474 1,814

3 16 5 20 4

65.2 291.2 292.9 754.9 396.4

Commits

P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

5,816 13,432 2,233 474 1,814

112 107 250 580 300

663.8 721.3 2,083.1 4,751.2 3,170.1

Pull Requests

P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

5,816 13,432 2,233 474 1,814

14 23 53 178 67

268.6 281.4 922.1 2,395.7 1,141.0

Issues

P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

5,816 13,432 2,233 474 1,814

2 13 11 37 7

89.9 276.5 462.0 1,019.5 579.4

Contributors

P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

5,816 13,432 2,233 474 1,814

5 7 6 17 4

25.1 37.0 53.4 131.3 82.2

Metric

Pattern

Distribution-only packages also have the second-highest median forks (16), just behind P4: Templating (20). This association suggests that distributiononly packages—such as JavaScript libraries repackaged as Maven WebJars—are widely recognized in the community. Interestingly, this high visibility occurs even though they lack native source code for some ecosystems, which may indicate that users are unaware of the absence of a native implementation. Despite their popularity, P2: Distribution-only packages show lower development activity. Their median commits (107) are similar to P1: Multi-repo (112) but much lower than P4: Templating (580) and P5: Bind/Wrap (300). Similarly, their median pull requests (23) are far below P4: Templating (178), P5: Bind/Wrap (67), and P3: Designated Dir. (53). This combination of high popularity and low activity is common when a package’s reputation is tied to the original source-language project, rather than active cross-ecosystem development.

20

Xiangxi Li et al.

Table 8: Pattern distribution: Top 10% vs. Bottom 10% by community metrics (n = 2,376 each) Pattern

Top 10%

Bottom 10%

Stars P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

91 ( 3.8%) 1,742 (73.3%) 220 ( 9.3%) 98 ( 4.1%) 225 ( 9.5%)

1,112 (46.8%) 776 (32.7%) 239 (10.1%) 52 ( 2.2%) 197 ( 8.3%)

Forks P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

195 ( 8.2%) 1,651 (69.5%) 224 ( 9.4%) 102 ( 4.3%) 204 ( 8.6%)

937 (39.4%) 867 (36.5%) 287 (12.1%) 45 ( 1.9%) 240 (10.1%)

Contributors P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

414 (17.4%) 1,327 (55.9%) 268 (11.3%) 119 ( 5.0%) 248 (10.4%)

662 (27.9%) 1,220 (51.3%) 236 ( 9.9%) 32 ( 1.3%) 226 ( 9.5%)

In contrast, P4: Templating packages have the highest median values across the board for activity: commits (580), pull requests (178), issues (37), and contributors (17). These higher metrics frequently correspond with templategeneration projects, which often serve infrastructure needs and involve larger teams. Finally, P1: Multi-repo has the lowest median values for all metrics (e.g., 6 stars, 3 forks, and 5 contributors). We observe that these lower numbers consistently accompany projects that split their development across separate, per-language repositories. For the second analysis (Tables 8 and 9), we find that P2: Distribution-only is over-represented in the top tier for community metrics but under-represented for activity, whereas P3: Designated Dir., P4: Templating, and P5: Bind/Wrap skew toward higher activity. For popularity metrics, P2: Distribution-only is concentrated in the top 10%, accounting for 73.3% of the most-starred packages, 69.5% of the most-forked, and 55.9% of those with the most contributors. However, for activity metrics, P2: Distribution-only are toward the bottom: 62.4% of the least-committed packages are P2: Distribution-only, compared to 43.4% in the top 10%. P1: Multi-repo constitutes 46.8% of the bottom 10% by stars but only 3.8% of the top, and 39.0% of the bottom by pull requests but only 16.7% of the top tier. Conversely, tightly integrated patterns (P3: Designated Dir., P4: Templating, P5: Bind/Wrap) consistently skew toward higher activity. P4:

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

21

Table 9: Pattern distribution: Top 10% vs. Bottom 10% by activity metrics (n = 2,376 each) Pattern

Top 10%

Bottom 10%

Commits P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

421 (17.7%) 1,032 (43.4%) 398 (16.8%) 159 ( 6.7%) 366 (15.4%)

646 (27.2%) 1,482 (62.4%) 135 ( 5.7%) 16 ( 0.7%) 97 ( 4.1%)

Pull Requests P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

397 (16.7%) 1,024 (43.1%) 408 (17.2%) 156 ( 6.6%) 391 (16.5%)

927 (39.0%) 1,089 (45.8%) 207 ( 8.7%) 26 ( 1.1%) 127 ( 5.3%)

Issues P1: Multi-repo P2: Distribution-only P3: Designated Dir. P4: Templating P5: Bind/Wrap

252 (10.6%) 1,438 (60.5%) 297 (12.5%) 122 ( 5.1%) 267 (11.2%)

942 (39.6%) 949 (39.9%) 224 ( 9.4%) 44 ( 1.9%) 217 ( 9.1%)

Templating appears at 6.7% of the top 10% by commits but only 0.7% of the bottom—a ratio of roughly 10×. P5: Bind/Wrap and P3: Designated Dir. show similar skews, with top-10% representation exceeding their bottom-10% share by approximately 2–3× across activity metrics.

We find that different architectural patterns are associated with different health metric profiles. Distribution-only packages (P2: Distribution-only) are among the most visible cross-ecosystem packages by stars and forks, yet they exhibit relatively low development activity. – Observation 5. P2: Distribution-only packages are associated with the highest star counts, while P4: Templating packages show the highest development activity (forks, commits, pull requests, issues, and contributors). – Observation 6. P2: Distribution-only is over-represented in the top decile for community metrics but under-represented for activity metrics, whereas P3: Designated Dir., P4: Templating, and P5: Bind/Wrap skew toward higher activity.

22

Xiangxi Li et al.

7 Discussion In this section, we discuss the implications for package adopters, package maintainers, researchers, and tool builders. In addition to these Implications, we also highlight different areas for future investigation. We find that most mono-repo cross-ecosystem packages do not maintain native source code for all registered ecosystems within their repository (Observation 2); the majority follow a distribution-only strategy where one language’s artifacts are repackaged for other ecosystems. At the same time, packages using templating or binding strategies are among the most actively developed (Observations 5 and 6). Developers should therefore be aware that a package imported from another ecosystem may not have a native implementation maintained for their target language. Our taxonomy directly supports this judgment: a package classified as P2: Distribution-only (distribution-only) versus P4: Templating (templating) or P5: Bind/Wrap (binding) carries meaningfully different implications for long-term language-specific support and maintenance. Based on these findings, developers can now make a more informed decision about whether a cross-ecosystem package suits their supply chain needs or whether a native alternative is preferable. Future developer surveys could elicit how aware adopters are of these distribution strategies and whether they affect adoption decisions.

7.1 Implications for package maintainers. A key finding of this research is that several strategies can be effective to develop these multilingual packages (Observation 1, Observation 2, Observation 3, Observation 4). Specifically, we found that there is a significant number of packages that are distribution only (P2: Distribution-only and Observation 2). For package maintainers currently working on supporting several languages natively, our results show that a significant amount of cross-ecosystem packages are able to successfully deploy to several ecosystems while maintaining a single programming language. Our results provide a good overview of the field and allow package maintainers to consider the use of tools to support additional languages.

7.2 Implications for Researchers and Tool Builders. For researchers, our investigation provides fine-grained (directory-level) empirical evidence that cross-ecosystem packages can be implemented and deployed in several ways. This study provides an operational taxonomy for conducting further research on open-source communities, development practices, and software ecosystems. The taxonomy is grounded in detectable repository signals and supported by manual validation, making it directly applicable to automated classification.

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

23

For tool builders, our taxonomy has concrete applications in dependency analysis and SBOM generation. SBOM tools could use the architectural pattern of a package to tag it as native, templated, or bound/wrapped, enriching supply chain metadata beyond simple package name and version. For example, knowing that a Maven package is a P2: Distribution-only (distribution-only) WebJar repackaging of a JavaScript library—rather than a native Java implementation— is directly relevant for vulnerability assessment, license compliance, and maintenance risk evaluation. Additionally, our results show that language binding and templating tools (such as JSII, WebAssembly, and Protocol Buffers) are widely adopted and in demand. New tools supporting more language targets could significantly expand the cross-ecosystem package landscape. For future work, researchers can use our taxonomy to investigate tradeoffs between native multi-language implementations and translation-based approaches (bindings, templating, wrappers). Security is another direction: because distribution-only packages introduce inter-ecosystem coupling—a vulnerability in one language implementation may propagate to ecosystems that repackage it—our pattern labels could serve as risk indicators in cross-ecosystem dependency tracking. Longitudinal analyses tracking how architectural patterns evolve over a project’s lifetime are also warranted.

8 Threats to Validity In this section, we discuss potential threats to the validity of our study and the measures taken to mitigate them. External Validity. Our study relies on GitHub repository URLs declared in package registries; packages without GitHub URLs or with incorrect URLs are excluded. This may limit the generalizability of our findings, as some packages—especially those hosted on alternative platforms or with incomplete metadata—are not represented in our dataset. Consequently, our results may not fully capture the diversity of package development and maintenance practices outside of GitHub or in less-documented ecosystems. We mitigate this by collecting the dataset ourselves and validating the process to the best of our knowledge. We will also make the data available for replication. Internal Validity. The ecosystem source file detection uses file extensions as proxies for language and project type, which may miss unconventional file organizations or atypical naming conventions. To assess sensitivity, we re-ran the heuristic-dependent analysis under two composite variants: a strict variant (fewer source extensions, 33 additional exclusion patterns, a two-file minimum per ecosystem, and higher P3 coverage thresholds of 90%/50%) and a lenient variant (broader source extensions, eight fewer exclusion patterns, and lower P3 thresholds of 70%/30%). Results are as follows. P4: Templating(template generation) is fully stable, shifting ≤ 4% under both variants, confirming that its detection via unambiguous IDL file extensions is robust. P3: Designated Dir.(designated-directory) is stable under the lenient variant (+0.7%), but drops by −17% under the strict variant; the strict drop is almost entirely explained by

24

Xiangxi Li et al.

the two-file minimum, which removes 22% of fully-matched packages from the input pool—the P3 rate within that pool actually rises from 43.5% to 46.3%, indicating no loss of detection quality. P2: Distribution-only(distribution-only proxy) is similarly driven by the two-file minimum in the strict variant (+8.4%) yet virtually unchanged under lenient (+0.3%). P5 WASM counts shift 18% (strict) and 11% (lenient) as expected from the deliberate threshold changes; general binding is stable under lenient (+0.8%) but shifts 16% under strict due to additional exclusion folders suppressing binding indicators. Taken together, the results confirm that P4 is robustly operationalized; for P3 and P5 the shifts under the strict variant are traceable to specific, documented parameter choices rather than to inherent instability of the detection algorithm. The multi-repo suffix detection was separately validated by checking that the expected language appears in each repository’s GitHub-reported language proportions (match rate > 98%), providing an independent cross-check on the heuristic accuracy. Construct Validity. We use star count as the primary community visibility metric. While star count is widely used in empirical software engineering research Zerouali et al (2019), it does not directly measure actual usage, download counts, or dependency adoption. As noted by Zerouali et al. Zerouali et al (2019), different methods for measuring popularity can yield different results, particularly in ecosystems like NPM. Thus, our findings regarding package visibility may be influenced by the limitations of this metric. Reliability. The manual inspection of 200 mismatched packages provides qualitative insights into the causes of mismatches. However, these findings may not generalize to all mismatched cases, as the sample size is limited and subject to inspector bias. Further automated or large-scale manual analyses would be needed to confirm the broader applicability of these observations.

9 Conclusion We present a large-scale empirical study of cross-ecosystem packages across six major package ecosystems. We find that cross-ecosystem packages account for a small but significant fraction of all GitHub repositories and identify five distinct architectural patterns: multi-repository separation, distributiononly publication, language-specific folder naming, protocol-buffer-like template code generation, and binding/wrapper integration. Our analysis reveals that distribution-only packages (dominated by Maven WebJar/mvnpm) constitute the largest category and are surprisingly popular, raising concerns that adopters may be unaware of the absence of native source code. Packages using tighter integration strategies (protocol buffers and bindings) tend to be the most actively developed and maintained.

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

25

Declarations Funding This research was supported by the Japan Society for the Promotion of Science, Grant No. [JP24H00692,JP25K03102,JP26K02889,JP26H02500].

Author Contributions Xiangxi Li implemented the code and analysis scripts used in the study, conducted the experiments, and contributed to the writing of the manuscript. Olivier Nourry contributed to the study design, supervised the research, and contributed to the writing of the manuscript. Yoshiki Higo acquired the funding for the project and contributed to the writing of the manuscript. Raula Gaikovina Kula contributed to the study design, supervised the research, and contributed to the writing of the manuscript. All authors read and approved the final manuscript.

Data Availability Statement We make all datasets used in this paper publicly available. We also provide all scripts needed to replicate the data mining, filtering, and analysis processes. Upon acceptance, we will provide all raw data for replication. The scripts can be found at the following link: cross-ecosystem-replication.

Conflict of Interest The authors declare that Raula Gaikovina Kula, is a member of the EMSE Editorial Board. All co-authors have seen and agree with the contents of the manuscript and there is no financial interest to report.

Ethical Approval Not applicable. This study used publicly available data and did not involve human participants or animals.

Informed Consent Not applicable. This study did not involve human participants.

26

Xiangxi Li et al.

References Abdalkareem R, Nourry O, Wehaibi S, Mujahid S, Shihab E (2017) Why do developers use trivial packages? an empirical case study on npm. In: Proceedings of the 2017 11th Joint Meeting on Foundations of Software Engineering, Association for Computing Machinery, New York, NY, USA, ESEC/FSE 2017, p 385–395, DOI 10.1145/3106237.3106267, URL https: //doi.org/10.1145/3106237.3106267 Abdalkareem R, Oda V, Mujahid S, Shihab E (2020) On the impact of using trivial packages: an empirical case study on npm and pypi. Empirical Software Engineering 25(2):1168–1204, DOI 10.1007/s10664-019-09792-9, URL http: //dx.doi.org/10.1007/s10664-019-09792-9 Ali N, Baker S, O’crowley R, Herold S, Buckley J (2018) Architecture consistency: State of the practice, challenges and requirements. Empirical Softw Engg 23(1):224–258, DOI 10.1007/s10664-017-9515-3, URL https: //doi.org/10.1007/s10664-017-9515-3 Bavota G, Canfora G, Di Penta M, Oliveto R, Panichella S (2015) How the apache community upgrades dependencies: an evolutionary study. Empirical Softw Engg 20(5):1275–1317, DOI 10.1007/s10664-014-9325-9, URL https: //doi.org/10.1007/s10664-014-9325-9 van den Berk I, Jansen S, Luinenburg L (2010) Software ecosystems: a software ecosystem strategy assessment model. In: Proceedings of the Fourth European Conference on Software Architecture: Companion Volume, Association for Computing Machinery, New York, NY, USA, ECSA ’10, p 127–134, DOI 10. 1145/1842752.1842781, URL https://doi.org/10.1145/1842752.1842781 Boehmke B, Hazen B (2017) The future of supply chain information systems: The open source ecosystem. Global Journal of Flexible Systems Management 18, DOI 10.1007/s40171-017-0152-x Bogart C, Kästner C, Herbsleb J, Thung F (2021) When and how to make breaking changes: Policies and practices in 18 open source software ecosystems. ACM Trans Softw Eng Methodol 30(4), DOI 10.1145/3447245, URL https: //doi.org/10.1145/3447245 Borges H, Valente MT (2018) What’s in a github star? understanding repository starring practices in a social coding platform. Journal of Systems and Software 146:112–129, DOI 10.1016/j.jss.2018.09.016, URL http://dx.doi.org/10. 1016/j.jss.2018.09.016 Borges H, Hora A, Valente MT (2016) Understanding the factors that impact the popularity of github repositories. In: 2016 IEEE International Conference on Software Maintenance and Evolution (ICSME), IEEE, pp 334–344, DOI 10.1109/icsme.2016.31, URL http://dx.doi.org/10.1109/ICSME.2016.31 Constantinou E, Decan A, Mens T (2018) Breaking the borders: an investigation of cross-ecosystem software packages. URL https://arxiv.org/abs/1812. 04868, 1812.04868 Decan A, Mens T, Grosjean P (2018) An empirical comparison of dependency network evolution in seven software packaging ecosystems. Empirical Software Engineering 24(1):381–416, DOI 10.1007/s10664-017-9589-y, URL http:

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

27

//dx.doi.org/10.1007/s10664-017-9589-y Gaikovina Kula R, Robles G (2023) The life and death of software ecosystems. arXiv e-prints pp arXiv–2306 Hu Y, Zhang J, Bai X, Yu S, Yang Z (2016) Influence analysis of github repositories. SpringerPlus 5:1–14, DOI 10.1186/s40064-016-2897-7 Huang K, Chen B, Xu C, Wang Y, Shi B, Peng X, Wu Y, Liu Y (2022) Characterizing usages, updates and risks of third-party libraries in java projects. Empirical Softw Engg 27(4), DOI 10.1007/s10664-022-10131-8, URL https://doi.org/10.1007/s10664-022-10131-8 Kannee K, Kula RG, Wattanakriengkrai S, Matsumoto K (2023) Intertwining communities: Exploring libraries that cross software ecosystems. URL https: //arxiv.org/abs/2303.09177, 2303.09177 Klein J, van Vliet H (2013) A systematic review of system-of-systems architecture research. In: Proceedings of the 9th International ACM Sigsoft Conference on Quality of Software Architectures, Association for Computing Machinery, New York, NY, USA, QoSA ’13, p 13–22, DOI 10.1145/2465478.2465490, URL https://doi.org/10.1145/2465478.2465490 Kula RG, German DM, Ouni A, Ishio T, Inoue K (2017) Do developers update their library dependencies? Empirical Software Engineering 23(1):384–417, DOI https://doi.org/10.1007/s10664-017-9521-5 Miller C, Jahanshahi M, Mockus A, Vasilescu B, Kästner C (2025) Understanding the Response to Open-Source Dependency Abandonment in the npm Ecosystem, IEEE Press, p 2355–2367. URL https://doi.org/10.1109/ ICSE55347.2025.00004 Prana GAA, Sharma A, Shar LK, Foo D, Santosa AE, Sharma A, Lo D (2021) Out of sight, out of mind? how vulnerable dependencies affect open-source projects. Empirical Softw Engg 26(4), DOI 10.1007/s10664-021-09959-3, URL https://doi.org/10.1007/s10664-021-09959-3 Spencer D (2009) Card Sorting: Designing Usable Categories. Rosenfeld Media, New York Tian F, Liang P, Ali Babar M (2021) Relationships between software architecture and source code in practice: An exploratory survey and interview. Information and Software Technology 141:106705, DOI 10.1016/j.infsof.2021. 106705 Valiev M, Vasilescu B, Herbsleb J (2018) Ecosystem-level determinants of sustained activity in open-source projects: a case study of the pypi ecosystem. In: Proceedings of the 2018 26th ACM Joint Meeting on European Software Engineering Conference and Symposium on the Foundations of Software Engineering, Association for Computing Machinery, New York, NY, USA, ESEC/FSE 2018, p 644–655, DOI 10.1145/3236024.3236062, URL https: //doi.org/10.1145/3236024.3236062 Williams L, Benedetti G, Hamer S, Paramitha R, Rahman I, Tamanna M, Tystahl G, Zahan N, Morrison P, Acar Y, Cukier M, Kästner C, Kapravelos A, Wermke D, Enck W (2025) Research directions in software supply chain security. ACM Trans Softw Eng Methodol 34(5), DOI 10.1145/3714464, URL https://doi.org/10.1145/3714464

28

Xiangxi Li et al.

Wu Y, Yu Z, Wen M, Li Q, Zou D, Jin H (2023) Understanding the threats of upstream vulnerabilities to downstream projects in the maven ecosystem. In: 2023 IEEE/ACM 45th International Conference on Software Engineering (ICSE), pp 1046–1058, DOI 10.1109/ICSE48619.2023.00095 Yang H, Nong Y, Wang S, Cai H (2024) Multi-language software development: Issues, challenges, and solutions. IEEE Transactions on Software Engineering 50(3):512–533, DOI 10.1109/TSE.2024.3358258 Zerouali A, Mens T, Robles G, Gonzalez-Barahona JM (2019) On the diversity of software package popularity metrics: An empirical study of npm. URL https://arxiv.org/abs/1901.04217, 1901.04217

Xiangxi Li Xiangxi Li is bachelor student from China. He is currently on a research exchange program working in Higo Laboratory in Japan.

Olivier Nourry is an Assistant Professor in the School of Engineering Science at The University of Osaka. His research interests include empirical software engineering, software maintenance, software quality, software security, and the application of artificial intelligence to software engineering.https://onourry.github.io/olivier-nourry/.

Yoshiki Higo is a Professor in the Graduate School of Information Science and Technology at The University of Osaka. His research interests include software engineering, particularly source code analysis, code clone analysis, refactoring support, software repository mining, and automated program repair. https://sites.google.com/view/yhigo/home.

Raula Gaikovina Kula is a Professor at The University of Osaka. He received his Ph.D. degree from NAIST in 2013 and was a Research Assistant Professor at Osaka University. He is active in the Software Engineering community, serving the community as a PC member for premium SE venues, some as organizing committee, and reviewer for journals. His current research interests include library dependencies and security in the software ecosystem, program analysis such as code clones, and human aspects such as code reviews and coding proficiency. Find him at

Cross-Ecosystem Packages As Multilingual: Prevalence, Architecture, and Health

29

https://raux.github.io/ and @augaiko on Twitter. Contact him at [email protected].

Record · ID 919492 · SHA-256 bc027b1ed3a48e71
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.