Conceptio › Archive › arXiv CS
arXiv CSopen access

TripleBound: Triplet-Guided Heterogeneous Graph Learning for Microservice Decomposition

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
software-architecturesoftware-engineeringtesting
software engineering, software architecture, testing

TripleBound: Triplet-Guided Heterogeneous Graph Learning for Microservice Decomposition Mineth Weerasinghe∗ , Himindu Kularathne∗ , Methmini Madhushika∗ , Danuka Lakshan∗ Nisansa de Silva∗ , Adeesha Wijayasiri∗ , Srinath Perera† ∗ Department of Computer Science and Engineering, University of Moratuwa, Sri Lanka

arXiv:2609.11212v1 [cs.SE] 10 Sep 2026

Email: {mineth.21, himindu.21, methmini.21, dhanuka.21, NisansaDdS, adeeshaw}@cse.mrt.ac.lk † WSO2 LLC Email: [email protected]

Abstract—Cloud computing and DevOps have made microservices a common architecture for scalable, maintainable software systems. However, migrating monoliths to microservices remains challenging due to tight coupling and unclear service boundaries. Existing decomposition approaches typically rely on either structural dependencies or semantic similarity signals, but rarely integrate both within a unified representation learning objective. This paper proposes TripleBound, a hybrid framework for automated monolith-to-microservices decomposition that augments a heterogeneous graph neural network with weakly supervised triplet constraints derived from parser-inferred service groups based on package structure, naming conventions, and code location. Rather than combining independently trained structural and organizational representations after training, TripleBound injects triplet-based constraints directly into the shared structural latent space, enabling both signals to be jointly optimized during representation learning. Structural dependencies are captured using CHGNN, which models the monolith as a heterogeneous graph with program nodes, resource nodes, CALL edges, and CRUD edges. Semantic relationships are incorporated through triplet constraints generated from service groups inferred by a parser using cues such as package structure, naming conventions, and code location. Evaluation on AcmeAir, DayTrader, PlantsByWebSphere, and JPetStore shows that TripleBound achieves the highest composite decomposition score under the selected weighting on AcmeAir, DayTrader, and JPetStore compared to CHGNN and MonoEmbed, while CHGNN remains stronger on PlantsByWebSphere. Per-metric analysis reveals trade-offs: gains in structural modularity and inter-partition coupling are accompanied by higher entity distribution imbalance on some datasets. Alternative composite weightings preserve TripleBound’s firstplace ranking on AcmeAir and DayTrader but not on JPetStore, showing that the aggregate ranking is metric-dependent. Index Terms—Microservices, Software Architecture, Monolith Decomposition, Graph Neural Networks

I. I NTRODUCTION Enterprise systems often begin as monoliths because they are simple to build, but become difficult to maintain, scale, and extend as they grow. The widespread adoption of cloud computing [1] and DevOps practices [2, 3] has further accelerated demand for modular, independently deployable architectures. Microservices address these challenges by decomposing applications into small, focused services [4–6]. Systematic mapping studies show that microservice research has grown around cloud deployment, architecture recovery, service identification, and migration support [7, 8]. Despite its advantages,

migrating monoliths to microservices remains complex [9–11] and requires careful identification of service boundaries and granularity [12–15]. Current decomposition techniques mainly rely on either structural dependencies or semantic analysis of source code. Structural approaches [16–19] analyze call graphs and dependency networks to capture interactions between software components. Semantic approaches use embeddings derived from source code to capture contextual similarities between components. However, approaches that rely on a single signal often fail to capture the full characteristics of software systems. Structural methods may ignore semantic relationships between components, while semantic methods may overlook critical dependency structures. To address these limitations, this paper proposes TripleBound, a hybrid decomposition framework that integrates structural and source-level organizational signals through joint representation learning. The structural view is captured using a heterogeneous graph neural network [20], which represents the monolithic system as a heterogeneous graph capable of modeling structural and behavioral relationships [21] among software components. The organizational view is incorporated through weakly supervised triplet constraints derived from inferred service groups, encouraging related code elements to cluster more closely within the shared structural representation space. By optimizing these complementary signals within a unified training objective, TripleBound aims to generate cohesive and architecturally consistent microservices. The evaluation uses four open-source Java monoliths frequently examined in decomposition and modernization research [22–25]. AcmeAir represents a REST-based airline reservation system, DayTrader is a comparatively larger Java EE transaction-processing application, PlantsByWebSphere models an online nursery, and JPetStore is a compact MyBatisbased commerce application. Together, they provide variation in application size, framework, dependency structure, and source-code organization; their detailed characteristics are reported in the experimental evaluation. The main contributions of this paper are: (1) A hybrid monolith-to-microservices decomposition framework that jointly integrates heterogeneous graphbased structural learning and weakly supervised triplet-

TABLE I C OMPARISON OF RELATED APPROACHES Approach

Struct.

Sem.

CHGNN [20]

✓

–

MonoEmbed [26]

Limited

✓

Mono2Micro [22]

✓

Partial

CoGCN [27]

✓

–

TripleBound

✓

✓

Limitation Does not explicitly model semantic similarity. Limited awareness of architectural dependencies. Uses multiple signals, but not through joint representation learning. Mainly focuses on graph structure and ignores semantic constraints. Depends on the quality of inferred triplets.

based constraint learning within a unified representation learning process. (2) We introduce a parser-inferred triplet-guided graph embedding strategy in which weakly supervised triplet constraints are injected directly into the structural embedding space, rather than combining structural outputs with independently trained representations. (3) We formulate a joint optimization objective that combines node reconstruction, edge reconstruction, clustering consistency, semantic triplet loss, and communication-aware coupling reduction to learn microservice embeddings. (4) We evaluate TripleBound on four benchmark monoliths and compare it against available structural and semantic baselines using Structural Modularity (SM), Interface Number (IFN), Inter-Partition Communication (ICP), Non-Extreme Distribution (NED), and a composite score. II. R ELATED W ORK Automated monolith-to-microservices decomposition has been studied using structural, semantic, graph-based, and hybrid techniques [23, 24, 26, 28, 29]. Structural approaches analyze dependencies such as method calls, data access, and runtime interactions to identify highly connected components [19, 30–32]. These methods are useful for preserving technical dependencies, but they may ignore domain-level semantic relationships between code elements. Semantic approaches use identifiers, comments, package names, or code embeddings [33] to capture functional similarity between components [16, 34, 35]. For example, Brito et al. [36] apply topic modelling to lexical information extracted from source code, where the inferred topics correspond to domain terms and are used to identify candidate microservices. Methods such as MonoEmbed [26] leverage pre-trained transformer-based language models [37] and code-specific embeddings [38] enhanced through contrastive learning [39] and LoRA fine-tuning [40] to group semantically related code elements. However, semantic similarity alone may overlook important architectural constraints such as inter-service communication and database coupling.

Graph-based approaches are widely used in microservice reengineering because software systems can be represented as graphs in which classes, methods, components, database tables, or other system entities are modeled as vertices, while dependencies, method calls, and coupling relationships are modeled as edges [15]. Recent extraction methods have also explored knowledge graphs with constrained community detection and graph deep clustering with multiple software views [41, 42]. Following this direction, CHGNN models the monolith as a heterogeneous graph containing program nodes, resource nodes, CALL edges, and CRUD edges. This enables the model to learn rich structural representations of software components [20, 43, 44]. However, CHGNN mainly focuses on structural dependency learning and does not explicitly optimize semantic similarity during training. Hybrid approaches [23, 45, 46] combine multiple signals such as structural dependencies, runtime traces, database access, and semantic information. Recent deep-learning approaches similarly construct structural and semantic views before clustering learned representations [42, 47]. Tools such as Mono2Micro [22] and DEEPLY [48] use several analysis signals or objective-tuning strategies, but typically combine them through separate stages rather than joint optimization. In contrast, TripleBound integrates structural graph learning and weakly supervised triplet-based constraint learning within a single unified training objective. Several empirical and survey studies have contributed comparative analyses of decomposition approaches. Taibi et al. [49] conducted an empirical comparison of multiple microservice identification techniques applied to a shared monolithic system, revealing that different approaches can produce substantially different decompositions on the same codebase even when targeting the same number of services. This highlights the importance of multi-metric evaluation when comparing methods. A systematic review by Abgaz et al. [24] catalogued existing decomposition approaches by input type, algorithmic technique, and evaluation methodology, identifying the absence of standardized benchmarks as a key barrier to rigorous cross-study comparison. The study by Weerasinghe et al. [6] is a comprehensive comparison of state-of-the-art microservice decomposition methods against benchmark data sets established in the domain. Additional tool-based and learning-driven approaches have been proposed to address specific aspects of the decomposition problem. CARGO [50] applies AI-guided dependency analysis to identify tightly coupled component clusters, leveraging graph-based structural reasoning and constraint propagation to support migration planning decisions. Huang et al. [51] introduce an attention-based bidirectional microservice clustering framework that incorporates directional relationship weighting to better capture asymmetric dependencies between software components. FoSCI [28] proposes a feature-oriented service clustering and identification approach that groups code elements according to feature-level analysis, providing an alternative decomposition perspective that complements purely structural or semantic strategies. RapidMS [52] takes a comple-

mentary direction by generating initial microservice structures from requirements models, demonstrating that decomposition guidance can also be derived from high-level design artifacts rather than source code analysis alone. Mo2oM [53] formulates microservice extraction as a soft clustering problem, combining deep semantic code embeddings with structural dependency graphs to allow components to belong probabilistically to multiple services, thereby relaxing the hard partition assumption that most decomposition methods, including the approach proposed in this paper, rely on. Another important distinction among existing decomposition approaches is the type of input required. Some approaches depend on manual or semi-manual inputs such as domain models [54], use cases, data-flow diagrams, design artifacts, or user-provided system knowledge [12, 31, 41]. While such inputs can improve domain awareness, they increase the effort required from developers and domain experts and may limit automation. In contrast, code-driven approaches aim to reduce this manual effort by extracting structural and semantic information directly from the monolithic application. Combining structural and semantic signals consistently yields more coherent service boundaries than either alone [23, 45], motivating the joint optimization objective proposed in this paper. III. T RIPLE B OUND A PPROACH A. Framework Overview TripleBound integrates structural and semantic analysis to generate microservice decompositions. Figure 1 illustrates the overall architecture of TripleBound. The parser constructs a heterogeneous graph containing program and resource nodes connected by CALL and CRUD edges, while also inferring service groups from package structure, naming conventions, and code location. Type-specific transformations, Node2CommonSpace and Edge2CommonSpace, project node and edge features into a shared representation space processed by an edge-aware graph autoencoder [55]. The autoencoder consists of stacked GNN encoder layers, a bottleneck representation, and a decoder responsible for reconstructing node and edge information while preserving the structural properties of the original application. Triplets generated from the parser-inferred groups are applied to the same encoder to guide the learned latent representations. This joint learning process encourages related components to remain close while separating components belonging to different inferred service groups. The resulting embeddings are partitioned using the K-means clustering algorithm [56], following the clustering approach used by CHGNN due to its simplicity and effectiveness in partitioning embedding spaces. The number of clusters, K, is set to the number of inferred service groups. For the benchmark systems in this study, the resulting K values range from 4 to 6, as shown in Table II. The resulting clusters represent candidate microservices and are evaluated using SM, IFN, ICP, and NED, as shown in the output stage of Figure 1.

Heterogeneous Graph Construction

Monolith's Source Code

GNN Autoencoder Encoder

Program Nodes

Inter class usage

GNN Stack Layer 1

Clustering

K-means

GNN Stack Layer 2

Resource Nodes Latent Embeddings

Static Analysis Parser

vertices

Call Edges

Decoder

GNN Stack Layer 3

CRUD Edges db schema

GNN Stack Layer 4

Cluster Assignments

Weakly Supervised Triplets Parser-Inferred Service Groups

Triplet Generation

Group 1

(a1, p1, n1) (a2, p2, n2)

Group 2 ... Group K

Loss Backpropagation

Metric Evaluation

...

(ak, pk, nk)

Fig. 1. Architecture of TripleBound

B. Heterogeneous Graph Construction To capture the structural characteristics of the monolithic system, we model the codebase as a heterogeneous graph following the representation used in CHGNN [20], building on heterogeneous relational graph modeling principles [57]. The graph consists of two primary types of nodes: • Program nodes, representing code elements such as classes and services • Resource nodes, representing external entities such as database tables [20] Node relationships are modeled using typed edges: • CALL edges, capturing invocation relationships between program components • CRUD edges, representing interactions between program components and database resources Each node and edge is associated with a feature representation encoding its structural properties. Because program nodes, resource nodes, CALL edges, and CRUD edges may have different feature formats, separate type-specific transformations are used to map them into a common latent space before they are processed by the graph encoder. This representation allows the model to jointly capture intra-code dependencies and data access patterns, which are critical for identifying cohesive microservice boundaries. While the graph construction follows the CHGNN formulation, it serves as the structural backbone of TripleBound, enabling seamless integration with the semantic triplet learning objective. In particular, the unified latent space learned from this graph facilitates the incorporation of both structural dependencies and semantic similarity constraints during training. C. Structural Representation Learning In analyzing the system, we leverage the fact that most programming languages and frameworks used for building web services follow layered architectural patterns [35, 58, 59]. In such architectures, methods in the presentation layer expose public interfaces and handle incoming requests, while lowerlayer methods are invoked to process business logic and interact with the persistence layer and database. These layered interactions naturally form structural dependency patterns within the monolithic system.

Following Weerasinghe et al. [6], we adopt the heterogeneous graph neural network (CHGNN) framework [20], which models the system as an edge-aware graph autoencoder. Given the heterogeneous graph constructed in the previous section, the model learns low-dimensional embeddings for each node by encoding both node attributes and typed edge relationships. Due to the presence of multiple node and edge types, type-specific transformations are first applied to project features into a shared latent space. The encoder consists of stacked graph neural network layers [60, 61] that propagate and aggregate information across neighboring nodes, taking into account both structural connectivity and edge semantics. This process produces a latent embedding hi for each node i, capturing its structural role within the system. A decoder is then used to reconstruct node attributes and typed edge features. The reconstruction objective ensures that the learned embeddings preserve important structural dependencies, such as method invocations and data access patterns. These structural embeddings serve as the foundation for the subsequent semantic triplet learning stage. D. Semantic Triplet Construction To incorporate semantic relationships between software components, we employ a triplet-based metric learning strategy following the general triplet-network formulation [62] and inspired by MonoEmbed [26], as these showed high promise in the comparative study conducted by Weerasinghe et al. [6]. Given the absence of explicit ground-truth microservice boundaries in real-world monolithic systems, we treat parserinferred service groups as weak supervisory signals, following the broader idea of weakly supervised contrastive learning, where inferred similarity relations are used to guide representation learning [63]. Specifically, the parser produces a service-to-node mapping, where each group contains code elements that are logically associated based on factors such as package structure, naming conventions, and location within the codebase. Formally, let S = {S1 , S2 , ..., Sk } denote the set of inferred service groups, where each Si is a set of nodes (classes). Using this grouping, triplets are constructed as follows: • Anchor (a): a randomly selected node from a service Si • Positive (p): another node from the same service Si • Negative (n): a node sampled from a different service Sj , where i ̸= j Since our service groups are parser-inferred weak labels rather than manually verified ground truth, we avoid hard-negative mining and instead randomly sample negatives from different inferred service groups. This conservative strategy reduces the risk of unstable optimization caused by hard-negative triplets, which have been shown to produce bad local minima under standard triplet loss training [64]. The triplets are not used to train a separate semantic embedding model. Instead, the anchor, positive, and negative nodes are passed through the same encoder used for structural graph representation learning. Therefore, the triplet loss directly shapes the CHGNN latent

space. This design differs from approaches that compute structural and semantic representations independently and combine them only after training. By applying triplet constraints to the shared latent embeddings, the model learns a unified representation in which structurally dependent components and semantically related components are jointly organized. Certain nodes may not be associated with any inferred service group because the parser cannot classify them using package structure, naming conventions, or code location. Such nodes, including shared utility and infrastructure classes, are excluded from triplet generation but remain in the heterogeneous graph for structural representation learning. E. Hybrid Training Objective Structural dependencies alone are often insufficient to capture domain-level relationships between software components, while purely semantic approaches may overlook critical architectural constraints. To address this limitation, we propose a hybrid training objective that jointly optimizes structural and semantic representations. Most graph neural networks follow a message-passing mechanism, where the vector representation of a node is updated by combining its own features with aggregated information from its neighboring nodes [43]. TripleBound combines node and edge reconstruction losses with a hybrid triplet-guided loss, clustering consistency, and communicationaware regularization. This enables the model to preserve both dependency structures and semantic similarity relationships within a unified embedding space. 1) Joint Loss Formulation: To integrate weak semantic supervision into the shared structural embedding space, we define a hybrid loss component. This component combines a triplet-local distance regularizer with a semantic triplet margin loss. The distance regularizer preserves the encoded/decoded representations of the nodes participating in each sampled triplet, while the triplet margin loss encourages nodes from the same parser-inferred service group to be closer than nodes from different groups. Let T denote the set of sampled triplets. Each triplet (a, p, n) ∈ T consists of an anchor node a, a positive node p sampled from the same parser-inferred service group, and a negative node n sampled from a different group. Let xe,i and xd,i denote the encoded and decoded representations of node i, respectively, and let hi denote the latent embedding used for triplet learning. The triplet-local distance regularization term is defined as:

Ltriplet = dist

1 |T |

X (a,p,n)∈T

1 3

X

∥xd,i − xe,i ∥33

(1)

i∈{a,p,n}

This term is computed only over the anchor, positive, and negative nodes in each sampled triplet. The semantic triplet loss is defined as:

Cluster Assignments

Encoder

Input

Decoder

Decoder Output

x

Graph Reconstruction Losses Node + edge feature reconstruction Latent Space

Total Loss Anchor

Hybrid Triplet-Guided Loss

Positive

Backpropagation (Updates all Parameters)

Negative

Input / Reconstruction Space

Hidden Layers

Latent / Code Space

Fig. 2. Internal training architecture of TripleBound.

Ltri =

1 |T |

X

max 0, ∥ha − hp ∥22 − ∥ha − hn ∥22 + β



(a,p,n)∈T

(2) where β is the triplet margin. To dynamically balance the two objectives during training, we define the distance and triplet weights as:   N −e wdist = max α, (3) N  e wtri = min 1 − α, N

(4)

where 0 < α < 1 is the hybrid-balance hyperparameter, e is the current training epoch and N is the total number of epochs. The hybrid loss component is then defined as: Lhybrid = wdist Ltriplet + wtri Ltri dist

(5)

The first term corresponds to triplet-local distance regularization, while the second term corresponds to semantic triplet supervision. Thus, Lhybrid regularizes the representations of triplet nodes while simultaneously enforcing parser-inferred grouping constraints in the latent space. The weighting mechanism dynamically balances local distance preservation and semantic triplet supervision during training. In early epochs, the model gives higher emphasis to the triplet-local distance term, while in later epochs the triplet margin loss receives greater relative emphasis. In our experiments, we set the hybrid-balance parameter to α = 0.7, which places greater emphasis on semantic consistency during later training stages while still preserving structural information in

early epochs. The final loss used during training is implemented as a weighted combination of multiple components: Ltotal =λnode Lnode + λedge Ledge + λcluster Lcluster + λhybrid Lhybrid + λicc Licc

(6)

where Lnode and Ledge are the graph-wide node and edge reconstruction losses, Lcluster enforces clustering consistency, Lhybrid is the hybrid loss component, and Licc is the proposed Inter-Cluster Communication Loss. The coefficients λnode , λedge , λcluster , λhybrid , and λicc are configurable hyperparameters set before training to control the relative contribution of each objective. This formulation allows the model to jointly optimize structural, semantic, and communication-aware objectives while keeping reconstruction independent of the triplet sampling process. To reduce inter-service communication during training, we introduce an Inter-Cluster Communication Loss, denoted as Licc . Before defining this term, we define the clustering objective used to encourage compact service partitions. Given node embeddings hi , cluster centroids Ck , and a binary assignment matrix M , the clustering loss is defined as: Lcluster =

K XX

2

Mi,k ∥hi − Ck ∥2

i∈V k=1

where Mi,k = 1 if node i is assigned to cluster k, and Mi,k = 0 otherwise. Minimizing Lcluster pulls node embeddings closer to their assigned cluster centroids, improving cluster compactness. Since the final ICP metric is computed after hard cluster assignment and is not directly differentiable, we define a soft communication-aware proxy based on cluster assignment

probabilities. Given node embeddings and cluster centroids, we compute the distance between each node and each centroid and apply a softmax over the negative distances: pi = softmax(−d(hi , C)) where pi represents the probability distribution of node i over clusters. For each dependency edge (u, v), the probability that its endpoints belong to the same cluster is computed as: Psame (u, v) =

K X

pu,k pv,k

k=1

The Inter-Cluster Communication Loss is then defined as: Licc =

1 |E|

X

(1 − Psame (u, v))

(u,v)∈E

Minimizing Licc encourages nodes connected by dependency edges to be assigned to the same cluster, thereby reducing potential communication across service boundaries. Although Licc is not the final ICP evaluation metric, it acts as a differentiable proxy that guides the embedding space toward more communication-aware microservice boundaries. During preliminary experiments, the original CHGNN structure loss caused unstable optimization in the hybrid setting. Since it reconstructs adjacency using raw pairwise dot products of node embeddings, it produced large gradients when combined with semantic triplet and clustering objectives, dominating the overall loss. This aligns with multi-task learning observations that objectives with larger gradient magnitudes can hinder balanced training [65]. Therefore, we removed the structure loss and retained node reconstruction, edge reconstruction, clustering, semantic triplet, and communicationaware losses, resulting in more stable training. Figure 2 illustrates the internal training architecture: the encoder processes each component representation into a shared latent embedding; the decoder reconstructs the original input; and parser-inferred triplets are passed through the same encoder to apply semantic constraints, all optimized jointly via backpropagation. IV. E XPERIMENTAL E VALUATION

B. Datasets We evaluate TripleBound using four open-source monolithic Java systems commonly used as evaluation subjects in microservice decomposition and modernization studies: AcmeAir1 , DayTrader2 , PlantsByWebSphere3 , and JPetStore4 [22–25]. Other systems, such as Train Ticket [66], are also present in the literature but are designed as microservice-native applications rather than monolithic systems intended for decomposition, and are therefore not used in this evaluation. AcmeAir is an airline reservation benchmark developed by IBM that represents a cloud-oriented web application with REST services, domain entities, and database/resource interactions. DayTrader [67] is a Java EE stock-trading benchmark that includes common enterprise components such as Servlets, JSPs, EJBs, JPA, JDBC, JMS, and transactions. PlantsByWebSphere is a Java EE sample application for an online plant nursery, containing web and REST-based functionality with business entities, resource accesses, and service entry points. JPetStore is a compact MyBatis-based Java web application containing account, catalog, cart, and order functionality. Table II summarizes the dataset characteristics used in our experiments. ICU denotes Inter-Class Usage. Inferred Groups denotes the parser-inferred service groups used for triplet generation. Clusters denotes the final value of K, which is set to the number of inferred groups. C. Evaluation Metrics The quality of the generated decompositions is evaluated using metrics: 1) Structural Modularity (SM) [28]: assesses the structural quality of a decomposition by considering both the cohesiveness of classes within each partition and the coupling between different partitions. A good decomposition should group classes that collaborate closely while reducing dependencies across service boundaries, which is consistent with the single responsibility principle. It is computed as shown in Equation 7. c In the equation, scohi = mi,i2 represents the cohesiveness i c within partition i. Similarly, scopi,j = 2(mii,j ∗mj ) represents the coupling between partitions i and j. Higher SM values indicate better modular decomposition [68].

A. Experimental Setup All experiments were repeated for 30 runs with different random initializations, and the reported results correspond to the average values across these runs. The loss components were weighted as follows: λnode = 0.001, λedge = 0.1, λcluster = 0.7, λhybrid = 1.0, and λicc = 1.0. These values differ slightly from the original CHGNN configuration and were empirically selected to better balance structural preservation and clustering quality in the hybrid setting. A replication package is available at https://github.com/Mono2Distributed/ triplet-guided-chgnn. The package includes the implementation, configuration files, and scripts required to reproduce the reported experiments.

M

SM =

M

X 1 1 X scohi − scopi,j M i=0 (M (M − 1))/2

(7)

i̸=j

SM is commonly regarded as an important measure of decomposition quality because it reflects the balance between intra-service cohesion and inter-service coupling [25, 32, 68]. A higher SM value indicates that structurally related classes are assigned to the same service, while unnecessary dependencies between services are minimized. Consequently, SM 1 https://github.com/acmeair/acmeair 2 https://github.com/WASdev/sample.daytrader7 3 https://github.com/WASdev/sample.plantsbywebsphere 4 https://github.com/mybatis/jpetstore-6

TABLE II C HARACTERISTICS OF BENCHMARK SYSTEMS Dataset

ICU Classes

Services/ Endpoints

DB/Resource Entries

Inferred Groups

Ignored Classes

Clusters

38 111 36 23

20 210 47 21

30 159 40 25

4 6 5 4

0 8 0 0

4 6 5 4

AcmeAir DayTrader PlantsByWebSphere JPetStore

TABLE III C OMPARISON OF D ECOMPOSITION Q UALITY ACROSS DATASETS Dataset

Method

SM ↑

IFN ↓

ICP ↓

NED ↓

Composite ↑

DayTrader

CHGNN MonoEmbed TripleBound

0.13 0.13 0.14

5.70 1.35 5.10

0.55 0.59 0.48

0.50 0.66 0.62

-0.3043 -0.4665 0.7708

PlantsByWebSphere

CHGNN MonoEmbed TripleBound

0.17 0.11 0.15

3.60 1.55 4.31

0.51 0.26 0.58

0.20 0.82 0.21

0.5188 -0.4449 -0.0738

AcmeAir

CHGNN MonoEmbed TripleBound

0.11 0.01 0.23

2.50 3.06 2.36

0.37 0.47 0.28

0.00 0.56 0.71

0.2782 -1.1066 0.8284

JPetStore

CHGNN MonoEmbed TripleBound

0.15 0.03 0.18

3.00 2.10 2.95

0.33 0.29 0.34

0.00 0.72 0.24

0.2046 -0.4502 0.2455

is frequently used as a primary optimization objective in decomposition approaches. 2) Interface Number (IFN) [28]: quantifies the average number of interfaces exposed by each microservice and is computed as shown in Equation 8, where if ni is the number of interfaces in the ith microservice. Interfaces refer to the externally accessible entry points through which a service communicates with other services. Lower IFN values indicate simpler, less fragmented services [25]. M

IF N =

1 X if ni M i=0

(8)

IFN captures service interface complexity; a high IFN may indicate overly fine-grained decomposition or unclear API boundaries. Reducing IFN supports cleaner service design and simpler inter-service communication. 3) Inter-partition Communication (ICP) [29]: measures the proportion of runtime calls that occur across partition boundaries and is computed as shown in Equation 9. It captures the extent to which different services interact during execution. Since communication between microservices can directly affect system performance, particularly response time and resource utilization [69–71], lower ICP values indicate better service separation and reduced runtime coupling. ICPi,j = PM

i′ =0

γi,j PM

j ′ =0,j ′ ̸=i′ γi ,j ′

(9) ′

A high ICP value suggests frequent cross-service calls, increasing latency and reducing service autonomy. Minimizing ICP supports loosely coupled services with better potential for independent deployment. 4) Non-Extreme Distribution (NED) [25]: assesses how evenly service sizes are distributed within a decomposition. It is defined as shown in Equation 10, where lower NED values reflect more balanced microservice size distributions. Originally proposed by Wu et al. [72], this metric identifies disproportionate service partitions where a few microservices dominate in size. A microservice is considered non-extreme when it satisfies 5 ≤ |mi | ≤ 20 [22, 73]. ( M X 0 5 ≤ |mi | ≤ 20 ni 1 Otherwise N ED = i=0 (10) M NED evaluates size balance among microservices. Extremely uneven service distributions, where a few services dominate, can lead to scalability bottlenecks. Hence, maintaining moderate NED values ensures that services remain evenly balanced and maintain granularity consistent with the “small and independent” nature of microservices [5]. 5) Composite Score: To summarize the overall decomposition quality across multiple metrics, we compute a composite score for each method and benchmark using weighted z-score normalization adapted from Sellami and Saied [26]. For each metric m ∈ {SM, IF N, ICP, N ED} and method t, the raw metric value xm,t is standardized across all compared methods

on the same benchmark. Let µm and σm denote the mean and standard deviation of metric m over all methods for a given benchmark. The normalized value is computed as: xm,t − µm (11) σm The composite score for method t is then computed as: P m∈{SM,IF N,ICP,N ED} wm Zm,t Score(t) = P (12) m∈{SM,IF N,ICP,N ED} |wm | where wm denotes the weight assigned to metric m. Following the weighting convention of Sellami and Saied [26], metrics that should be maximized are assigned positive weights, while metrics that should be minimized are assigned negative weights. Since SM should be increased, it is assigned a positive weight. IFN, ICP, and NED should be decreased, so they are assigned negative weights. Following the weighting scheme established in the study [6], we use: Zm,t =

W = {wSM , wIF N , wICP , wN ED } = {3, −1, −1, −1} (13) We assign a greater positive weight to SM because it captures the core structural objective of microservice decomposition. However, this weighting makes the composite score sensitive to SM improvements and should be interpreted alongside the individual metrics. V. R ESULTS The following research questions guide the evaluation: RQ1: How does TripleBound compare with existing structural and semantic decomposition baselines in terms of decomposition quality? RQ2: How does TripleBound perform across benchmark systems of different sizes and decomposition granularities? RQ3: How sensitive is TripleBound to key training parameters such as the hybrid-balance parameter α and the number of training epochs? A. RQ1: Comparison with Structural and Semantic Baselines We compare TripleBound against two representative baselines: CHGNN, a structure-based heterogeneous graph neural network approach, and MonoEmbed, a semantic embeddingbased decomposition approach. The experimental results in Table III show that TripleBound achieves the best overall composite score on AcmeAir, DayTrader and JPetStore, while CHGNN performs best on PlantsByWebSphere. However, the picture at the individual metric level is nuanced: the composite score triple-weights SM, and gains on SM and ICP are sometimes accompanied by regressions on NED or IFN. Thus, the composite score is used only as a summary indicator, not as definitive evidence that one decomposition is preferable under all quality criteria. For DayTrader, TripleBound obtains the highest composite score of 0.7708. It achieves the best SM value of 0.14 and the lowest ICP value of 0.48, indicating improved service cohesion and reduced inter-service coupling. The IFN value of 5.10 is

TABLE IV H IGHEST- RANKED METHOD UNDER ALTERNATIVE COMPOSITE WEIGHTS . PARENTHESES CONTAIN THE CORRESPONDING COMPOSITE SCORE . Dataset DayTrader PlantsByWebSphere AcmeAir JPetStore

wSM = 3

wSM = 2

wSM = 1

TripleBound (0.7708) CHGNN (0.5188) TripleBound (0.8284) TripleBound (0.2455)

TripleBound (0.6421) CHGNN (0.4088) TripleBound (0.7421) CHGNN (0.1530)

TripleBound (0.4491) CHGNN (0.2437) TripleBound (0.6126) CHGNN (0.0756)

higher than MonoEmbed’s best of 1.35, suggesting that the hybrid decomposition exposes more inter-service interfaces on this system. NED also regresses relative to CHGNN (0.62 vs. 0.50), indicating a somewhat less balanced entity distribution. The SM gain over CHGNN is small in absolute terms (0.14 vs. 0.13), and the practical significance of this difference should be interpreted with caution given the absence of statistical significance testing. For AcmeAir, TripleBound achieves the highest SM value of 0.23, the lowest IFN value of 2.36, and the lowest ICP value of 0.28, resulting in the highest composite score of 0.8284. A notable limitation is the NED value of 0.71, which is substantially higher than CHGNN’s 0.00, indicating that the resulting service sizes are less evenly distributed. This trade-off suggests that improving cohesion and coupling on AcmeAir comes at the cost of entity distribution balance. For JPetStore, TripleBound achieves the highest composite score of 0.2455. Compared with CHGNN, TripleBound slightly improves IFN (2.95 vs. 3.00) and SM (0.18 vs. 0.15), while CHGNN obtains the best NED value of 0.00. MonoEmbed achieves the lowest IFN and ICP values, but its lower SM value of 0.03 and higher NED value of 0.72 reduce its overall composite score. For PlantsByWebSphere, CHGNN achieves the best composite score of 0.5188, outperforming TripleBound on SM, IFN, ICP, and NED. TripleBound’s NED of 0.21 is close to CHGNN’s 0.20, but its ICP of 0.58 and IFN of 4.31 are worse, indicating that the complete TripleBound configuration does not generalise as effectively to this system. The result establishes an application-specific weakness, although the present evaluation cannot determine which modified component causes it. 1) Robustness to Alternative Composite Weights: To assess whether the aggregate ranking is determined by the threefold weight assigned to SM, we recomputed the composite score from the metric values in Table III using two alternative weight vectors: W = {2, −1, −1, −1} and the equal-magnitude vector W = {1, −1, −1, −1}. This calculation changes only the aggregation of the reported metrics and does not require retraining. Table IV reports the highest-ranked method for each dataset under each weighting. The ranking is stable for DayTrader, PlantsByWebSphere, and AcmeAir across these weightings. TripleBound remains first on DayTrader and AcmeAir, whereas CHGNN remains first on PlantsByWebSphere. In contrast, TripleBound’s firstplace result on JPetStore depends on assigning SM the threefold weight: CHGNN ranks first when the SM weight is

reduced to either two or one. Thus, the aggregate evidence supports a weighting-robust advantage for TripleBound on two of the four systems, rather than a weighting-independent advantage on three systems. Raw metrics remain necessary when decomposition priorities differ. B. RQ2: Performance Across Benchmark Characteristics TripleBound exhibits different decomposition trade-offs across the four evaluated systems. To make these differences concrete, we consider the observable characteristics reported in Table II, particularly system size and the number of inferred service groups. DayTrader is the largest evaluated system, with 111 ICU classes and 210 service/endpoints entries, and is decomposed into six clusters, whereas AcmeAir and JPetStore are substantially smaller and are each decomposed into four clusters. PlantsByWebSphere represents an intermediate case with five inferred groups. On the larger DayTrader system, TripleBound improves SM and ICP over both baselines, achieving an SM of 0.14 and ICP of 0.48, although this is accompanied by worse NED than CHGNN and substantially higher IFN than MonoEmbed. On AcmeAir, TripleBound achieves the strongest SM, IFN, and ICP values, but produces substantially worse NED. On JPetStore, the improvement is more limited, with TripleBound improving SM over CHGNN while failing to improve ICP or NED. On PlantsByWebSphere, TripleBound is outperformed by CHGNN across all four metrics. These results show that TripleBound’s gains are not associated uniformly with either larger or smaller systems, nor with a particular number of target clusters. Instead, its effect is metric- and dataset-specific. Because the present evaluation does not directly quantify properties such as dependency density, package modularity, naming consistency, or agreement between parser-inferred groups and structural dependencies, we cannot attribute the observed differences to specific codebase characteristics. C. RQ3: Sensitivity to Training Parameters To answer RQ3, we analyze the sensitivity of TripleBound to the hybrid balance parameter α and the number of training epochs. Each setting was repeated for 30 random initializations, and the reported composite scores correspond to average values across these runs. First, we varied α from 0.1 to 0.9 while fixing the number of training epochs to 100. Since α controls the balance between local distance preservation and triplet-based semantic supervision, this experiment evaluates how different weighting choices affect decomposition quality. As shown in Fig. 3, the effect of α varies across the benchmark systems, indicating that the optimal balance between structural and semantic signals depends on the characteristics of the target application. However, α = 0.7 provides a stable trade-off across the evaluated datasets, achieving strong composite scores without the sharp performance drops observed at some lower values.

Fig. 3. Effect of α on overall decomposition performance. Results are averaged across 30 runs with the number of training epochs fixed at 100.

Fig. 4. Effect of training epochs on overall decomposition performance. Results are averaged across 30 runs with α = 0.7.

Next, we varied the number of training epochs from 50 to 250 while fixing α = 0.7. This experiment evaluates whether longer training improves the hybrid objective or causes overfitting to dataset-specific structural and organizational patterns. As shown in Fig. 4, 100 epochs provides the best overall balance. Performance tends to degrade at higher values, suggesting that longer training may overfit dataset-specific patterns. Therefore, the final evaluation uses α = 0.7 and 100 training epochs. VI. D ISCUSSION A. Structural vs. Semantic Signal Strength The experimental results show that the behavior of the complete TripleBound configuration is application-dependent. On AcmeAir, TripleBound achieves the strongest SM, IFN, and ICP values among the compared methods, but its NED increases substantially from CHGNN’s 0.00 to 0.71. Thus, it produces more cohesive and less communication-intensive partitions under the reported metrics at the cost of markedly less balanced service sizes. Whether this trade-off is preferable

depends on the architectural priorities of the intended migration. AcmeAir is relatively compact and its package structure may provide organizational cues that agree with the structural graph, although this was not directly tested. On DayTrader, the improvements are more selective. TripleBound improves structural modularity and inter-partition communication. However, this comes with a larger interface surface compared with MonoEmbed and a less balanced entity distribution compared with CHGNN. The complete configuration therefore favors cohesion and coupling on this system but does not uniformly improve interface simplicity or size balance; the contribution of each loss component cannot be inferred from this comparison. For JPetStore, TripleBound achieves the highest composite score under the selected weighting and improves structural modularity over CHGNN. However, MonoEmbed achieves lower inter-partition communication, CHGNN maintains better entity distribution balance, and the alternative-weight analysis ranks CHGNN first when the SM weight is reduced. The claimed aggregate advantage on this system is therefore not robust to reasonable changes in metric priorities. B. Failure Mode Analysis: PlantsByWebSphere The underperformance of TripleBound relative to CHGNN on PlantsByWebSphere warrants detailed analysis. PlantsByWebSphere is a compact benchmark system, comprising only 36 ICU classes and 47 service endpoints. In such a compact system, the structural dependency graph alone may carry sufficient information for the clustering algorithm to find cohesive service boundaries without additional semantic guidance. Alternatively, triplet constraints derived from package and naming cues may conflict with the structural connectivity patterns. These explanations are plausible but were not directly tested; the available evidence establishes the failure case, not its cause. This observation is consistent with findings reported in the hybrid decomposition literature [23, 45], which suggest that the benefit of combining complementary signals depends critically on the quality and mutual consistency of each individual source. When one signal is already highly informative and well aligned with the desired decomposition outcome, the addition of a weaker or less consistent signal can introduce noise rather than useful supervision. Future work should therefore investigate lightweight signal-quality assessment strategies that can determine, prior to training, whether semantic supervision is likely to improve or degrade decomposition quality for a given target application. C. Implications for Practice These findings carry several practical implications for teams applying automated decomposition tools in real-world modernization projects. First, no single decomposition method consistently dominates across all applications and evaluation metrics. Practitioners should consider the size, organizational structure, and naming conventions of their target system when selecting or configuring a decomposition approach.

Applications with clear package hierarchies and consistent naming conventions are more likely to benefit from semantic supervision, while structurally simpler or legacy systems with accumulated technical debt may be adequately served by purely structural methods. Second, the evaluation metric used significantly influences perceived performance. A method that ranks highest on a composite score may rank lower on criteria most relevant to a given project, such as service size balance (NED) or interface simplicity (IFN). Teams should identify which quality criteria matter most before interpreting results. Third, the triplet-based supervision strategy shows that weak labels derived from source-level organizational cues can guide representation learning without ground-truth decompositions when those cues align with likely service boundaries. Its effectiveness may decrease in legacy codebases where accumulated technical debt causes source-level organization to diverge from the underlying architectural structure. Finally, the ICP improvements on AcmeAir and DayTrader are consistent with the intended effect of the Inter-Cluster Communication Loss (Licc ). They do not establish its independent contribution, however, because the comparison changes triplet supervision, loss weighting, and other training components at the same time. The JPetStore and PlantsByWebSphere results further show that lower ICP is not obtained uniformly. VII. T HREATS TO VALIDITY Several factors may affect the validity of the results reported in this study. Limited benchmark coverage. The evaluation is conducted on four benchmark monolithic systems highlighted by Weerasinghe et al. [6], where an expanded comparative evaluation of decomposition approaches across a broader benchmark set, including additional methods and metrics, is provided. Although these systems are commonly used in microservice decomposition research, they represent relatively small applications and may not fully capture the complexity and diversity of large industrial monoliths. Generalisation to larger or differently structured systems remains to be validated. Triplet label leakage. The semantic triplets are derived from parser-inferred service groups based on package structure, naming conventions, and code location. On standard enterprise benchmarks, these organizational cues often correlate strongly with modular boundaries. Consequently, the triplet supervision may partially reflect pre-existing module organization rather than independently discovered semantic relationships, making it difficult to separate meaningful semantic learning from boundary leakage through source organization. Hyperparameter tuning on evaluation benchmarks. The hybrid balance parameter α and the number of training epochs were selected based on sensitivity analyses conducted on the evaluation benchmarks. This entanglement of tuning and evaluation may overestimate performance on the reported benchmarks. Evaluating generalization to held-out systems would strengthen confidence in the chosen configuration.

Lack of component ablations. TripleBound changes several aspects of the CHGNN baseline simultaneously, including triplet-based supervision, communication-aware regularization, loss weighting, and removal of the original structure loss. The evaluation therefore establishes the behavior of the complete TripleBound configuration but cannot identify the independent contribution of each modification. In particular, changes in SM or ICP cannot be attributed exclusively to the triplet or communication-aware objectives. An incremental evaluation of the CHGNN backbone, the addition of triplet supervision, the addition of communication regularization, and the complete weighted objective is needed to isolate these effects. Statistical uncertainty. Although each configuration was executed 30 times and Table III reports averages, the evaluation does not report dispersion, confidence intervals, effect sizes, or statistical significance tests. Consequently, small mean differences, such as the DayTrader SM values of 0.14 and 0.13, should not be interpreted as reliable evidence of superiority. Per-run measurements and paired statistical tests with effect sizes and correction for multiple comparisons are required to quantify confidence in the reported differences. Composite score weighting. The composite score uses the weight vector W = {3, −1, −1, −1}, which gives SM greater influence than IFN, ICP, and NED. This choice follows the view that structural modularity is central to decomposition quality, but it can favor methods that optimize cohesion and coupling while degrading service-size balance. The alternativeweight analysis reduces this concern but cannot cover every project-specific preference. In particular, TripleBound’s firstplace ranking on JPetStore changes when the SM weight is reduced. Composite scores should therefore be interpreted alongside the raw metrics rather than as an absolute ranking. Cluster count selection. The number of clusters K is set to the number of parser-inferred seed groups. This choice may influence the Normalized Entity Distribution (NED) metric and may not always correspond to the ideal number of microservices for a given system. The impact of alternative cluster selection strategies is not assessed. Metric coverage. The evaluation relies on structural metrics (SM, IFN, ICP, and NED). While these are widely used in decomposition research, they do not capture business-domain alignment, maintainability, deployability, or architectural feasibility. Comparison against expert-designed decompositions or practitioner validation would provide a more complete picture of decomposition quality. Parser accuracy. The quality of the decomposition depends on the accuracy of the parser and the extracted dependency graph. Missing or incorrect dependencies, database interactions, or code-level relationships may affect the learned representations and the resulting service boundaries. VIII. C ONCLUSION This paper presented TripleBound, a hybrid framework for automated monolith-to-microservices decomposition. It combines heterogeneous graph-based structural learning with

weakly supervised triplet constraints derived from parserinferred service groups. Unlike post-training fusion, TripleBound injects these constraints directly into the shared latent space, jointly optimizing structural and grouping objectives. Experiments on four benchmarks show that TripleBound achieves the highest composite score on AcmeAir, DayTrader, and JPetStore under the selected weighting scheme. The ranking remains stable under two alternative weightings on AcmeAir and DayTrader, but not on JPetStore, while CHGNN performs better on PlantsByWebSphere under all tested weightings. The results are metric-selective: improvements in modularity and coupling may coincide with reduced entity distribution balance. Because the current evaluation does not isolate the individual loss components or quantify statistical significance, it supports the complete framework as a promising, application-dependent configuration rather than establishing uniform or component-specific superiority. Future work will focus on runtime and domain-level signals, adaptive weighting, improved clustering, and feedback-based refinement. TripleBound can also be extended toward migration support through API generation, dependency analysis, communication recommendations, and initial code generation. R EFERENCES [1] M. Armbrust, A. Fox, R. Griffith, A. D. Joseph, R. H. Katz, A. Konwinski, G. Lee, D. A. Patterson, A. Rabkin, I. Stoica, and M. Zaharia, “A View of Cloud Computing,” Communications of the ACM, vol. 53, no. 4, pp. 50–58, 2010. [2] G. Kim, P. Debois, J. Willis, and J. Humble, The DevOps Handbook: How to Create World-Class Agility, Reliability, and Security in Technology Organizations. IT Revolution Press, 2016. [3] A. Balalaie, A. Heydarnoori, and P. Jamshidi, “Microservices Architecture Enables DevOps: Migration to a Cloud-Native Architecture,” IEEE Software, vol. 33, no. 3, pp. 42–52, 2016. [4] S. Newman, Building Microservices: Designing Fine-Grained Systems. O’Reilly Media, 2021. [5] N. Dragoni, I. Lanese et al., “Microservices: How to Make Your Application Scale,” in Perspectives of System Informatics, ser. LNCS. Springer, 2018, vol. 10742, pp. 293–313. [6] M. Weerasinghe, H. Kularathne, M. Madhushika, D. Lakshan et al., “From Monolith to Microservices: A Comparative Evaluation of Decomposition Frameworks,” in WorldCist 2026, 2026. [7] C. Pahl and P. Jamshidi, “Microservices: A Systematic Mapping Study,” in Proceedings of the 6th International Conference on Cloud Computing and Services Science (CLOSER), vol. 1. SciTePress, 2016, pp. 137–146. [8] P. Di Francesco, P. Lago, and I. Malavolta, “Architecting with Microservices: A Systematic Mapping Study,” Journal of Systems and Software, vol. 150, no. 4, pp. 77–97, 2019. [9] J. Soldani, D. A. Tamburri, and W.-J. Van Den Heuvel, “The Pains and Gains of Microservices: A Systematic Grey Literature Review,” Journal of Systems and Software, vol. 146, pp. 215–232, 2018. [10] D. Taibi, V. Lenarduzzi, and C. Pahl, “Processes, Motivations, and Issues for Migrating to Microservices Architectures: An Empirical Investigation,” IEEE Cloud Computing, vol. 4, no. 5, pp. 22–32, 2017. [11] P. Di Francesco, P. Lago, and I. Malavolta, “Migrating Towards Microservice Architectures: An Industrial Survey,” in ICSA. IEEE, 2018, pp. 29–38. [12] M. Gysel, L. Kölbener, W. Giersche, and O. Zimmermann, “Service Cutter: A Systematic Approach to Service Decomposition,” in ServiceOriented and Cloud Computing, ser. Lecture Notes in Computer Science, vol. 9846. Springer, 2016, pp. 185–200. [13] F. Auer, A. Rausch, and N. Drehmann, “From monolithic systems to microservices: An assessment,” Journal of Systems and Software, vol. 180, p. 111034, 2021. [14] S. Hassan, R. Bahsoon, and R. Kazman, “Microservice Transition and Its Granularity Problem: A Systematic Mapping Study,” Software: Practice and Experience, vol. 50, no. 9, pp. 1651–1681, 2020.

[15] T. I. Mohottige, A. Polyvyanyy et al., “Reengineering Software Systems into Microservices: State-of-the-Art and Future Directions,” Information & Software Technology, vol. 143, p. 107732, 2025. [16] O. Al-Debagy and P. Martinek, “A microservice decomposition method through using distributed representation of source code,” Scalable Computing, vol. 22, pp. 39–52, 2021. [17] B. S. Mitchell and S. Mancoridis, “On the automatic modularization of software systems using the bunch tool,” IEEE Transactions on Software Engineering, 2006. [18] I. Saidani, A. Ouni, M. W. Mkaouer, and A. Saied, “Towards automated microservices extraction using multi-objective evolutionary search,” in ICSOC, 2019. [19] G. Mazlami, J. Cito, and P. Leitner, “Extraction of Microservices from Monolithic Software Architectures,” in ICWS. IEEE, 2017, pp. 524– 531. [20] A. Mathai, S. Bandyopadhyay, U. Desai, and S. Tamilselvam, “Monolith to microservices: Representing Application Software through Heterogeneous Graph Neural Network,” arXiv preprint arXiv:2112.01317, 2021. [21] S. Mancoridis, B. S. Mitchell, C. Rorres, Y. Chen, and E. R. Gansner, “Using automatic clustering to produce high-level system organizations of source code,” in IWPC. IEEE, 1998, pp. 45–52. [22] A. K. Kalia, J. Xiao et al., “Mono2Micro: A practical and effective tool for decomposing monolithic Java applications to microservices,” in ESEC/FSE, 2021, pp. 1214–1224. [23] K. Sellami, M. A. Saied, and A. Ouni, “A hierarchical DBSCAN method for extracting microservices from monolithic applications,” in EASE. ACM, 2022. [24] Y. Abgaz, A. McCarren, P. Elger et al., “Decomposition of Monolith Applications Into Microservices Architectures: A Systematic Review,” Journal of Systems and Software, vol. 203, p. 111704, 2023. [25] M. A. Saied, “Migration to microservices: A comparative study of decomposition strategies and analysis metrics,” arXiv preprint arXiv:2402.08481, 2024. [26] K. Sellami and M. A. Saied, “Contrastive Learning-Enhanced Large Language Models for Monolith-to-Microservice Decomposition,” arXiv preprint arXiv:2502.04604, 2025. [27] U. Desai, S. Bandyopadhyay et al., “Graph Neural Network to Dilute Outliers for Refactoring Monolith Application,” in AAAI, vol. 35, 2021, pp. 72–80. [28] W. Jin, T. Liu et al., “Service Candidate Identification from Monolithic Systems Based on Execution Traces,” IEEE Transactions on Software Engineering, vol. 47, no. 5, pp. 987–1007, 2021. [29] A. K. Kalia, J. Lin et al., “Mono2Micro: An AI-based Toolchain for Evolving Monolithic Enterprise Applications to a Microservice Architecture,” in ESEC/FSE, 2020, pp. 1606–1610. [30] A. Levcovitz, R. Terra, and M. T. Valente, “Towards a Technique for Extracting Microservices from Monolithic Enterprise Systems,” arXiv preprint arXiv:1605.03175, 2016. [31] S. Li, H. Zhang, Z. Jia, Z. Li, C. Zhang, J. Li, Q. Gao, J. Ge, and Z. Shan, “A dataflow-driven approach to identifying microservices from monolithic applications,” Journal of Systems and Software, vol. 157, p. 110380, 2019. [32] G. Filippone, R. Mehmood et al., “From monolithic to microservice architecture: An automated approach based on graph clustering and combinatorial optimization,” in ICSA. IEEE, 2023, pp. 47–57. [33] M. Allamanis, E. T. Barr, P. Devanbu, and C. Sutton, “A Survey of Machine Learning for Big Code and Naturalness,” ACM Computing Surveys, vol. 51, no. 4, pp. 81:1–81:37, 2018. [34] O. Al-Debagy and P. Martinek, “Semantic-based microservice identification using natural language processing,” in 2021 2nd International Conference on Electrical, Communication, and Computer Engineering (ICECCE). IEEE, 2021, pp. 1–6. [35] I. Trabelsi, M. Abdellatif et al., “From legacy to microservices: A typebased approach for microservices identification using machine learning and semantic analysis,” Journal of Software: Evolution and Process, 2022. [36] M. Brito, J. Cunha, and J. Saraiva, “Identification of Microservices from Monolithic Applications through Topic Modelling,” in Proceedings of the 36th Annual ACM Symposium on Applied Computing. ACM, 2021, pp. 1409–1418. [37] A. Vaswani, N. Shazeer et al., “Attention Is All You Need,” in Advances in Neural Information Processing Systems, vol. 30. Curran Associates, Inc., 2017. [38] Z. Feng, D. Guo, D. Tang, N. Duan, X. Feng, M. Gong, L. Shou, B. Qin,

T. Liu, D. Jiang, and M. Zhou, “CodeBERT: A Pre-Trained Model for Programming and Natural Language,” in Findings of the Association for Computational Linguistics: EMNLP 2020, 2020, pp. 1536–1547. [39] T. Chen, S. Kornblith, M. Norouzi, and G. Hinton, “A Simple Framework for Contrastive Learning of Visual Representations,” in Proceedings of the 37th International Conference on Machine Learning, ser. Proceedings of Machine Learning Research, vol. 119. PMLR, 2020, pp. 1597–1607. [40] E. J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, and W. Chen, “LoRA: Low-Rank Adaptation of Large Language Models,” arXiv preprint arXiv:2106.09685, 2021. [41] Z. Li, C. Shang, J. Wu, and Y. Li, “Microservice Extraction Based on Knowledge Graph from Monolithic Applications,” Information and Software Technology, p. 106992, 2022. [42] L. Qian, J. Li, X. He, R. Gu, J. Shao, and Y. Lu, “Microservice Extraction Using Graph Deep Clustering Based on Dual View Fusion,” Information and Software Technology, vol. 158, p. 107171, 2023. [43] U. Desai, S. Bandyopadhyay, and S. Tamilselvam, “Graph neural network to dilute outliers for refactoring monolith application,” in AAAI, 2021. [44] M. Lecrivain, T. Barry et al., “MONO2REST: Identifying and Exposing Microservices: a Reusable RESTification Approach,” in ICSR, 2025, pp. 33–43. [45] K. Sellami, M. A. Saied, A. Ouni, and R. Abdalkareem, “Combining static and dynamic analysis to decompose monolithic applications into microservices,” in ICSOC, ser. Lecture Notes in Computer Science. Springer, 2022, pp. 203–218. [46] A. Krause, C. Zirkelbach, W. Hasselbring, S. Lenga, and D. Kröger, “Microservice Decomposition via Static and Dynamic Analysis of the Monolith,” arXiv preprint arXiv:2003.02603, 2020. [47] X. Wei, J. Li, X. He, W. Peng, Y. Zhu, R. Gu, Y. Zhu, and T.-J. Huang, “Extracting Microservices from Monolithic Applications Using Consistent Graph Enhanced Graph Transformer,” Journal of Systems and Software, vol. 222, p. 112345, 2025. [48] R. Yedida, R. Krishna, A. Kalia, T. Menzies, J. Xiao, and M. Vukovic, “An Expert System for Redesigning Software for Cloud Applications,” arXiv preprint arXiv:2109.14569, 2022. [49] D. Taibi, V. Lenarduzzi et al., “Empirical comparison of microservice identification approaches using a monolithic system,” Journal of Systems and Software, vol. 176, p. 110941, 2021. [50] V. Nitin, S. Asthana, B. Ray, and R. Krishna, “Cargo: AI-guided dependency analysis for migrating monolithic applications to microservices architecture,” in ASE. Association for Computing Machinery, 2022, pp. 1–12. [51] Y. Huang et al., “a-bmsc: An attention-based bidirectional microservice clustering framework,” in ISCC, 2024. [52] Y. Zhang, Y. Li, Y. Yang, S. Chen et al., “RapidMS: A Tool for Supporting Rapid Microservices Generation and Refinement from Requirements Model,” in 2023 ACM/IEEE International Conference on Model Driven Engineering Languages and Systems Companion (MODELS-C). IEEE, 2023, pp. 45–49. [53] M. Ziabakhsh, K. Rezaee, S. Eskandari, S. A. H. Tabatabaei, and M. M. Ghassemi, “Extracting Overlapping Microservices from Monolithic Code via Deep Semantic Embeddings and Graph Neural NetworkBased Soft Clustering,” arXiv preprint arXiv:2508.07486, 2025. [54] E. Evans, Domain-Driven Design: Tackling Complexity in the Heart of Software. Addison-Wesley, 2004. [55] T. N. Kipf and M. Welling, “Variational Graph Auto-Encoders,” arXiv preprint arXiv:1611.07308, 2016. [56] S. P. Lloyd, “Least Squares Quantization in PCM,” IEEE Transactions on Information Theory, vol. 28, no. 2, pp. 129–137, 1982. [57] M. Schlichtkrull, T. N. Kipf, P. Bloem, R. van den Berg, I. Titov, and M. Welling, “Modeling Relational Data with Graph Convolutional Networks,” in European Semantic Web Conference. Springer, 2018, pp. 593–607. [58] M. Richards, Software Architecture Patterns. O’Reilly Media, Inc., 2015. [59] P. Zaragoza, A.-D. Seriai, A. Seriai, A. Shatnawi, and M. Derras, “Leveraging the layered architecture for microservice recovery,” in ICSA. IEEE, 2022, pp. 135–145. [60] W. L. Hamilton, R. Ying, and J. Leskovec, “Inductive Representation Learning on Large Graphs,” in Advances in Neural Information Processing Systems, vol. 30. Curran Associates, Inc., 2017. [61] Z. Wu, S. Pan, F. Chen, G. Long, C. Zhang, and P. S. Yu, “A

Comprehensive Survey on Graph Neural Networks,” IEEE Transactions on Neural Networks and Learning Systems, vol. 32, no. 1, pp. 4–24, 2021. [62] E. Hoffer and N. Ailon, “Deep metric learning using triplet network,” in International Workshop on Similarity-Based Pattern Recognition. Cham: Springer International Publishing, 2015, pp. 84–92. [63] M. Zheng, F. Wang, S. You, C. Qian, C. Zhang, X. Wang, and C. Xu, “Weakly supervised contrastive learning,” in ICCV, 2021, pp. 10 042– 10 051. [64] H. Xuan, A. Stylianou, X. Liu, and R. Pless, “Hard negative examples are hard, but useful,” in ECCV. Cham: Springer International Publishing, 2020, pp. 126–142. [65] Z. Chen, V. Badrinarayanan, C.-Y. Lee, and A. Rabinovich, “GradNorm: Gradient Normalization for Adaptive Loss Balancing in Deep Multitask Networks,” in ICML, vol. 80, 2018, pp. 794–803. [66] X. Zhou, X. Peng et al., “Benchmarking microservice systems for software engineering research,” in ICSE Companion. ACM, 2018, pp. 323–324. [67] S. Davis et al., “IBM DayTrader: A Performance Benchmark for J2EE,” in IBM DeveloperWorks Technical Reports, 2007. [68] B. S. Mitchell and S. Mancoridis, “On the evaluation of the Bunch search-based software modularization algorithm,” Soft Computing, vol. 12, no. 1, pp. 77–93, 2008. [69] M. Niswar, R. A. Safruddin, A. Bustamin, and I. Aswad, “Performance evaluation of microservices communication with rest, graphql, and grpc,” International Journal of Electronics and Telecommunication, vol. 70, no. 2, pp. 429–436, 2024. [70] M. Bolanowski, K. Żak et al., “Efficiency of REST and gRPC Realizing Communication Tasks in Microservice-Based Ecosystems,” arXiv preprint arXiv:2208.00682, 2022. [71] F. Tapia, M.-A. Mora et al., “From monolithic systems to microservices: A comparative study of performance,” Applied Sciences, vol. 10, no. 17, p. 5797, 2020. [72] J. Wu, A. E. Hassan, and R. C. Holt, “Comparison of Clustering Algorithms in the Context of Software Evolution,” in ICSM. IEEE, 2005, pp. 525–535. [73] G. Scanniello, A. D’Amico, C. D’Amico, and T. D’Amico, “An Approach for Architectural Layer Recovery,” in ACM SAC, 2010, pp. 2198– 2202.

Record · ID 673577 · SHA-256 968a76b46b9439ca
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.