Conceptio › Archive › arXiv CS
arXiv CSopen access

Co-occurrence Patterns of LoRA Adapters in Production Diffusion Model Inference Services

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

Co-occurrence Patterns of LoRA Adapters in Production Diffusion Model Inference Services Tao Zhanga , Bin Liaob,c , Tao Zhoud , Yanping Liub,c a School of Information Engineering, Guizhou University of Traditional Chinese Medicine, Guiyang, 550025, China b School of Big Data and Statistics, Guizhou University of Finance and Economics, Guiyang, 550025, China c School of Mathematics and Statistics, Guizhou University of Finance and Economics, Guiyang, 550025, China d School of Mathematics and Statistics, Beijing Technology and Business University, Beijing, 100048, China

arXiv:2609.23321v1 [cs.DC] 20 Sep 2026

Abstract Low-rank adaptation (LoRA) has become a key technology for serving large-scale personalized large language models and diffusion models in the cloud. However, the co-occurrence patterns, resource contention relationships, and evolutionary regularities of adapters under production inference workloads have not been systematically or quantitatively studied. Based on GenTD26, Alibaba’s production diffusion model inference dataset, this paper adopts a graph-theoretic framework to construct an adapter co-occurrence network and conducts a characterization from both static structure and dynamic evolution. Our main findings are as follows. (1) The co-occurrence network is extremely sparse, and adapter usage frequency follows a significant heavy-tailed distribution. (2) Introducing the first adapter incurs a 66.1% execution-latency overhead, with diminishing marginal costs afterwards. (3) Co-occurrence relationships are driven by base models: in 90.6% of multi-adapter requests, all adapters share the same dominant base model; 66.2% of significant cooccurrence edges connect same-model adapter pairs; and in 85.8% of multi-adapter requests, all adapter pairs form significant co-occurrence edges. (4) The adapter ecosystem exhibits a core–periphery bipolar structure, with a weekly Jaccard similarity of 0.696 at the model level and a churn rate of 54.5% for the top-10 hottest models within a 12-hour window. Based on these findings, we propose a preloading strategy built on top-k co-occurrence statistics; offline experiments show that it covers 81.0% of test-set co-occurrence pairs at k = 3, and sensitivity analyses across frequency thresholds and time windows verify the robustness of the conclusions. These results provide a data-driven basis for cache preloading, adaptive scheduling, and GPU memory management in LoRA inference services. Keywords: LoRA adapter, co-occurrence network, diffusion model inference, workload characterization, GPU cluster trace

1. Introduction Diffusion models, represented by Stable Diffusion, have become the core engine of commercial image generation services, processing millions of user requests per day on large GPU clusters [1], and have found wide application in fields such as medical data generation [3] and personalized image generation [4]. Low-Rank Adaptation (LoRA) [2] attaches lightweight adapter weights so that a single base model can support thousands of personalized variants, significantly reducing storage and computation costs while greatly improving the deployment flexibility of customized services. However, this “one-base-multiple-adapters” architecture also poses structural challenges to GPU resource management: inference requests must dynamically load one or more adapters on top of the base model, and the loading latency, GPU memory footprint, and scheduling order of adapters are mutually coupled, making request-level resource demands highly time-varying [5, 6]. Unlike traditional inference services that run fixed model architectures [12, 13], the LoRA adapter ecosystem exhibits a distinct long-tail usage pattern. The vast majority of adapters serve niche customization scenarios; only a very small number appear frequently, and multiple adapters may be invoked simultaneously by the same request. This workload characteristic raises a series of key system design questions: How do adapters co-occur within requests? Are co-occurrence relationships random, or do they exhibit predictable structural regularities? Which adapters should be preferentially resident in the GPU cache, and which can be loaded on demand? And are co-occurrence patterns stable over time? The answers to these questions directly determine the design of cache preloading, GPU memory allocation, and request batching and scheduling. ∗ Corresponding author.

Existing work on LoRA inference systems mainly focuses on architecture and mechanism design. Punica [7] and S-LoRA [8] design sharded memory management and on-demand adapter loading mechanisms, significantly increasing the number of concurrent adapters that a single GPU can host; MixLoRA [9] optimizes GPU utilization for concurrent execution of multiple adapters; and CaraServe [10] explores CPU-assisted rank-aware loading to reduce adapter-switching overhead. When designing scheduling policies, these systems typically assume that adapter access patterns are known or adopt general-purpose caching strategies such as LRU, lacking structural characteristics of real workloads as a design basis. The Rock system [6] conducts adapter popularity and burstiness analysis based on the GenTD26 trace, but its characterization remains at the aggregate statistical level and does not reveal the co-occurrence topology and structural properties among adapters. The SoCC’25 paper [5] provides a top-down characterization of diffusion model inference services but does not analyze the co-occurrence structure. To address the above problems, this paper starts from the structural associations among adapters and studies the co-occurrence patterns of LoRA adapters in production diffusion model inference services through a graph-theoretic approach. Unlike existing work that focuses on aggregate statistics (such as occurrence frequency and request share), this paper focuses on the co-occurrence topology and structural evolution patterns among adapters. The main contributions of this paper are as follows. (1) Co-occurrence network analysis. Based on production traces, we construct a LoRA adapter cooccurrence network and reveal its extreme sparsity (under the threshold τmin = 10, |V | = 274, |E| = 68, and an edge density of only 0.18%) and its significant heavy-tailed frequency distribution. Through connectedcomponent and community-detection analysis, we further identify the structural characteristic of “a small number of clusters plus a large number of isolated nodes.” (2) Multi-dimensional quantitative workload analysis. In addition to basic network structure analysis, this paper quantitatively analyzes the workload from four perspectives: Jaccard similarity, base model–adapter association, dynamic changes at three time scales (request level, model level, and cross-model level), and latency degradation. (3) System design strategy analysis. The co-occurrence-based preloading strategy covers 81.0% of testset co-occurrence pairs at k = 3 (with a range of [78.1%, 81.3%] over 5 random seeds), and in 85.8% of multi-adapter requests all adapter pairs form significant co-occurrence edges, providing a data-driven basis for cache preloading; frequency–centrality correlation analysis can guide the setting of scheduling priorities under memory constraints. The remainder of this paper is organized as follows. Section 2 reviews related work; Section 3 formalizes the relevant problems; Section 4 describes the dataset and presents its overall characteristics; Section 5 analyzes the co-occurrence network structure; Section 6 characterizes the dynamic evolution; Section 7 discusses implications for system design; and Section 8 discusses limitations and concludes the paper. 2. Related Work This section analyzes research in two areas: LoRA adapter serving systems and the characterization of production workloads, and clarifies the differences between this paper and existing work. 2.1. LoRA Adapter Serving Systems After the proposal of LoRA, efficiently serving a large number of adapters quickly became a research hotspot in the systems community. Punica [7] implements efficient batching of different LoRA requests on a single GPU through segmented gather matrix-vector multiplication kernels; S-LoRA [8] introduces a unified paging mechanism that manages the KV cache and LoRA weights together, supporting thousands of concurrent adapters; MixLoRA [9] interleaves the computation of heterogeneous adapters within the same batch, improving GPU utilization. Reference [17] applies LoRA instruction fine-tuning to domain-specific training of multimodal large models, and Petals [18] explores a distributed inference architecture for crossnode collaboration. In the broader area of generative AI inference optimization, systems such as vLLM [19], Splitwise [20], ServerlessLLM [21], and FlexPipe [22] also continuously optimize inference efficiency from perspectives including GPU memory management, phase splitting, cold-start loading, and pipeline refactoring. Regarding adapter loading overhead, recent work has begun to explore loading–computation orchestration and cache management. CaraServe [10] adopts a CPU-assisted strategy that executes adapter computation in the prefill stage in parallel with GPU-side adapter loading, reducing cold-start latency; Toppings [33] proposes a CPU-assisted, rank-aware scheduling algorithm that reduces average request latency by up to 1.7×; AuLoRA [34] orchestrates loading and computation at the layer granularity, improving GPU utilization through layer-priority loading and intra-layer pipeline execution. In cache management, Chameleon [35] 2

utilizes idle GPU memory to cache popular adapters and designs an adapter-aware scheduling strategy to minimize loading overhead; JointSerLoRA [36] jointly considers adapter loading states and KV cache memory contention and designs a cache-aware scheduling mechanism. Notably, when designing cache and scheduling policies, the above systems typically assume that adapter access patterns are known or adopt general-purpose strategies such as LRU. They answer the question of “how to execute efficiently after adapters are loaded,” but not the prior question of “which adapters will be requested together.” 2.2. Characterization of Production Workloads Characterizing production cluster workloads is a foundational methodology in system resource management research. Since the analysis of early Google cluster traces [11], researchers have recognized that detailed workload measurement can reveal failure modes of general-purpose strategies in real environments; Resource Central [12] further demonstrates that workload understanding can be directly translated into quantifiable scheduling benefits. This tradition has continued in GPU data center research: MLaaS [13] characterizes an Alibaba production cluster with more than 6,000 GPUs; the Philly trace [14] focuses on task arrival and GPU utilization characteristics of training workloads; Reference [15] systematically characterizes the resource demands and predictability of deep learning workloads; and Reference [16] provides a survey of scheduling research in GPU data centers. In addition, abstracting workloads into graph structures for analysis has been proven effective in revealing hidden dependency relationships in scenarios such as data centers (VM placement [23], task dependencies [24], microservice invocation [25]) and machine learning workloads (model dependencies [26], interference patterns [27]). In the field of diffusion model inference services, the GenTD26 dataset [5] records multi-component pipelines, LoRA adapter configurations, and full-stack performance data. Based on this dataset, the SoCC’25 paper [5] provides a top-down characterization of inference services; the Rock system [6] further characterizes adapter popularity and request burstiness, and achieves an 84.1% cache hit rate based on heterogeneous-aware resource orchestration. SwiftDiffusion [37] analyzes the inference request traces of a commercial text-to-image service and finds that add-on modules such as ControlNet and LoRA are widely used in production, and their loading overhead significantly affects service latency and GPU resource efficiency. The above work analyzes “which adapters are most popular” and “how occurrence frequency is distributed,” but does not analyze the co-occurrence relationships and structural regularities among adapters, which directly affect the design of cache preloading and request batching strategies. Unlike the above work, this paper focuses on the co-occurrence topology and structural evolution patterns among adapters. Characterizing co-occurrence patterns through graph-theoretic methods answers the questions of “which adapters appear together, in what structure they cluster, and whether they are stable over time.” Specifically, in terms of analysis granularity, our analysis shifts from frequency statistics of individual adapters to the co-occurrence relationships among adapters, introducing co-occurrence network construction, connected components, and community detection. Co-occurrence regularities can be used to guide cache preloading and scheduling priority setting, and can provide references for workload optimization of other componentized services (such as plugin markets and dependency combinations in function-as-a-service). 3. Problem Formalization and Co-occurrence Analysis Methods This section formalizes the graph-theoretic framework used to analyze LoRA adapter co-occurrence patterns. Let A = {a1 , a2 , . . . , aN } be the set of N unique LoRA adapters observed in the production trace, and let R = {r1 , r2 , . . . , rM } be the set of M requests. Each request rk is represented by a tuple rk = (tk , bk , Ak , lk ), containing the timestamp tk , the base model bk , the adapter set Ak ⊆ A, and the execution latency lk in seconds. 3.1. Problem Formalization Definition 1 (Co-occurrence event). If a request simultaneously loads two different adapters, the two adapters are said to co-occur in that request. A request may trigger multiple co-occurrence events simultaneously. Definition 2 (Co-occurrence count matrix). The co-occurrence matrix C ∈ NN ×N has elements defined as the number of co-occurrences of any two adapters across all requests, as shown in Eq. (1): Cij = |{rk ∈ R : ai ∈ Ak ∧ aj ∈ Ak ∧ i ̸= j}|

(1)

The diagonal entry Cii represents the total number of occurrences of adapter ai (i.e., the number of requests containing ai ). To filter out incidental co-occurrences and focus on statistically meaningful relationships, we impose an edge weight threshold: edges satisfying Cij ≥ wmin (default wmin = 3) are retained. We 3

impose a frequency threshold τmin on nodes (default set to 10) and retain only adapters satisfying Cii ≥ τmin , yielding the node set Vτ = {ai ∈ A : Cii ≥ τmin }. The co-occurrence network is thus defined as Gτ = (Vτ , Eτ ), where Eτ = {(i, j) : Cij ≥ wmin } and the edge weight is wij = Cij . Throughout this paper, the default values are τmin = 10 and wmin = 3, and we perform sensitivity analysis over τmin ∈ {5, 10, 20, 50}. 3.2. Structural Metrics of the Co-occurrence Network (1) Network edge density: defined as the ratio of the actual number of edges to the maximum possible number of edges. It characterizes the overall sparsity of the network and is the primary indicator for judging whether co-occurrence is a prevalent phenomenon. A value close to 0 indicates “almost no co-occurrence,” while a value close to 1 indicates “co-occurrence everywhere.” (2) Connected component: defined as the maximal subset of nodes in which at least one path exists between any two nodes, i.e., a group of adapters that are mutually reachable through paths in the network. Connected components can reveal natural clusters among adapters. (3) Jaccard similarity: the number of co-occurrences of two adapters divided by the union of their respective occurrence counts. If two adapters always appear in pairs, the Jaccard value approaches 1; if they are each popular but rarely used together, the Jaccard value is small. Jij =

Cij |Ri ∩ Rj | = |Ri ∪ Rj | Cii + Cjj − Cij

(2)

where Ri = {rk ∈ R : ai ∈ Ak }. The closer Jij is to 1, the stronger the co-occurrence dependency. A permutation test (e.g., 1,000 random shuffles of adapter labels) is used to determine whether the observed mean significantly deviates from the random expectation. 3.3. Co-occurrence Strength and Structural Characteristics (1) Modularity Q of partition C: as a quantitative verification of the degree of fragmentation, we adopt greedy modularity maximization [28] to detect community structure, as defined in Eq. (3):   ki kj 1 X wij − δ(ci , cj ) (3) Q= 2m i,j 2m P P where m = 21 i,j wij , ki = j wij is the weighted degree, and δ(ci , cj ) = 1 indicates that adapters i and j belong to the same community. Due to the high sparsity of the co-occurrence network in this paper, most communities identified by modularity maximization are single nodes, which is itself an important structural finding rather than an algorithm failure. Therefore, community detection is positioned as a quantitative verification of the degree of fragmentation, and the main body of structural analysis is anchored to nontrivial connected components. P (2) Degree centrality: di = j Aij reflects the local influence of an adapter in the co-occurrence network, where Aij = 1 if (i, j) ∈ Eτ and 0 otherwise. The cache residency of adapters with high di has a greater impact on the quality of service of multi-adapter requests. 3.4. Temporal Evolution Analysis To characterize temporal evolution, we divide the trace timeline into T discrete windows W = {W1 , . . . , WT }, with a window span ∆t (e.g., 1 hour or 1 day). For each window Wt , we define a binary (t) usage vector u(t) ∈ {0, 1}|A| , where ui = 1 if and only if adapter ai is used at least once within Wt . The window-to-window adapter turnover rate is quantified as: (t)

Turnover(Wt , Wt+δ ) = 1 −

(t+δ)

|{i : ui = 1 ∧ ui

(t) |{i : ui = 1}|

= 1}|

(4)

This metric measures the complement of the proportion of adapters active in Wt that remain active after δ windows. The above adapter-level temporal metrics require the data to carry both timestamps and adapter identifiers. The processing trace used in this paper (data_trace_processed.csv) has no timestamp field, the request-level trace (lora_request_trace.csv) has no adapter identifier field, and there is no verifiable linking field between the two (see Section 4 for details). Therefore, this paper adopts three strictly reproducible alternatives in the temporal dimension: request-level diurnal patterns, model-level evolution, and adapter cross-model spread. The impact of the above analysis results on cache design is further elaborated in Section 7. 4

4. Dataset and Overall Characteristics 4.1. The GenTD26 Dataset and Data Scope We use the GenAI Inference Service Top-Down Dataset 2026 (GenTD26) [5], which contains 26,823 requests, 874 unique LoRA adapters, and 68,195 processing records, covering 24 days of complete operational data from Alibaba’s commercial image generation service. Table 1 summarizes the key characteristics of the dataset. Table 1: Summary of the GenTD26 dataset Attribute

Value

Provider Duration Number of requests Number of processing records GPU sampling points Containers Unique base models Unique LoRA adapters Request types

Alibaba commercial image generation November 15 – December 8, 2024 (24 days) 26,823 (lora_request_trace) 68,195 (data_trace_processed) 157,417 (duty cycle) 143 86 (request-trace scope) 874 TXT_2_IMG (91.1%), IMG_2_IMG (8.2%), INPAINTING (0.7%)

The data scope is described below. We use two core files from the dataset: lora_request_trace.csv (26,823 rows), which provides the request-level timestamp gmt_create, the number of LoRAs num_lora, and base model identifiers, but does not contain adapter identifier information; and data_trace_processed.csv (68,195 rows), which records the JSON-format LoRA adapter configuration of each request but lacks a timestamp field. Therefore, the former is used for request-level and model-level temporal analysis, and the latter is used for adapter-level structural analysis. Notably, first, the row positions of the two files do not correspond: if aligned by row number, the execution time agreement rate between request-level and processing-level data is only 2.7%, and the request type agreement rate is 0%. Second, the anonymization schemes of the model identifiers in the two files are also different: only one model identifier is identical in both files, and the processing trace does not contain a groupId field. Therefore, there is no reliable linking key between the two files. For adapter-level analysis, data_trace_processed.csv records the set of loaded adapters (identified by anonymized modelVersionId) and their weight scaling factor scale for each request, thereby enabling adapterlevel co-occurrence analysis. After parsing and verification, we identified 874 unique LoRA adapters. Among the 68,195 records in the processing trace, 15,202 (22.3%) load at least one adapter, of which 1,090 (1.6%) load multiple adapters (with a maximum of 6 in a single request), and the remaining 52,993 (77.7%) use only the base model. Among the 26,823 requests in the request trace (lora_request_trace.csv), 4,494 (16.8%) carry at least one adapter. The difference in statistical scope between the two files arises from differences in coverage and record granularity. Therefore, when citing relevant metrics in the following sections, we explicitly indicate the data scope used. 4.2. Frequency Distribution Figure 1 shows the adapter frequency distribution (adapters sorted by occurrence count). The distribution exhibits a strongly heavy-tailed pattern: the most frequent adapter (rank 1) appears 1,321 times, which is 11.2 times that of the 30th-ranked adapter (118 times); the top 5 adapters (0.6% of all adapters) collectively contribute 19.7% of total occurrences; the top 5% of adapters (i.e., 44 adapters) account for 53.5% of all 16,709 occurrences; and the median adapter appears only 5 times. The Gini coefficient is 0.749 (Bootstrap 95% confidence interval: [0.697, 0.789]), further quantifying and confirming the high inequality of the usage distribution. Regarding statistical testing of the distribution shape, the rank–frequency log–log regression gives a slope of α = 1.30. Following the practice of prior work [5], this is used as a descriptive measure. However, we adopt the testing procedure proposed by Clauset et al. [29], fit a pure power-law distribution using maximum likelihood estimation (MLE), and perform a goodness-of-fit test based on the Kolmogorov–Smirnov (KS) statistic. The results show that when xmin = 1 (covering all 874 adapters), the fitted parameter is α = 1.56, the KS statistic is D = 0.177, and the semi-parametric Bootstrap test gives p = 0.003, thereby rejecting the pure power-law hypothesis at the 0.05 significance level. Only for the high-frequency tail (xmin = 20, 154 adapters) is the data compatible with a power-law distribution (α = 2.04, p = 0.40). Therefore, we characterize the distribution as heavy-tailed: although the extreme inequality of usage concentration (top 5% covering 53.5%) holds, a globally strict power law is not statistically supported. Note that this distinction 5

does not affect the effectiveness of the caching strategy. Regardless of the exact distribution form, the two-tier design of “permanently caching the head and loading the tail on demand” is insensitive to it.

Figure 1: LoRA adapter frequency distribution. (a) Top-30 adapters by occurrence, annotated with the Gini coefficient (0.749) and the top-5 share (19.7%), with colors distinguishing frequency tiers (dark blue: ranks 1–5; red: 6–10; light blue: 11–30); (b) log–log frequency–rank plot with the rank–frequency regression fit line (α = 1.30, a descriptive measure; strict power-law fitting is given in the main text).

4.3. Latency Impact of LoRA Usage We further quantify the execution latency associated with LoRA adapter usage. Figure 2(a) first shows the distribution of the number of LoRA adapters carried by each request: 77.7% of requests carry no adapter, 20.7% carry 1 adapter, and 1.6% carry multiple adapters. Table 2 reports the execution time statistics stratified by the number of LoRA adapters per request. Requests carrying 1 adapter take an average of 39.2 seconds, an increase of 66.1% over the no-adapter baseline (23.6 seconds); 2 adapters take 40.4 seconds (+71.2%); and 3 adapters take 44.7 seconds (+89.4%). In terms of marginal cost, the first adapter contributes the vast majority of the increment, while the second and third adapters add only 3.0% and 10.7%, respectively, compared with the previous configuration. This indicates that the main overhead comes from the loading and weight merging of the first adapter, while the marginal cost of subsequent adapters decreases. The sample sizes for 4 or more adapters are small (43 cases for 4 adapters, 10 cases for 5 adapters, and 6 cases for 6 adapters), so these groups are not included in the trend conclusions. Figure 2(b) further presents this tiered effect in terms of average execution time. Table 2: Execution time vs. number of LoRA adapters Number of LoRAs

Number of requests

Mean (s)

Standard deviation (s)

Degradation

0 1 2 3 4 5 6

52,993 14,112 754 277 43 10 6

23.6 39.2 40.4 44.7 40.8 38.8 27.2

13.4 25.1 21.2 16.0 12.6 11.1 8.7

(baseline) +66.1% +71.2% +89.4% +72.9% +64.4% +15.3%

Note: the n = 6 group is not included in the latency-trend conclusions due to its small sample size.

The dominant overhead of the first adapter is further quantified in Figure 3. Among 27,415 LoRA loading operations, the loading latency is tightly concentrated around 4.4 seconds (mean = 4.40 s, median = 4.38 s, P90 = 6.14 s, P99 = 8.21 s), which is only 18.6% of the no-LoRA execution time (23.6 s). This indicates that requests carrying adapters can hide the loading cost in the early stage of the pipeline through preloading, while the overhead of multiple adapters mainly comes from weight merging and concurrent execution rather than per-adapter loading latency.

6

Figure 2: LoRA count distribution and latency impact. (a) Distribution of the LoRA count per request (77.7% none, 20.7% single, 1.6% multiple); (b) mean execution time with percentage degradation over the no-LoRA baseline (23.6 s).

Figure 3: LoRA loading latency. (a) Histogram of 27,415 loading operations (mean = 4.40 s, median = 4.38 s, P90 = 6.14 s); (b) CDF of loading latency versus execution-time baselines (no-LoRA 23.6 s, single-LoRA 39.2 s); a typical loading operation accounts for only 18.6% of the no-LoRA execution time.

7

5. Co-occurrence Network Analysis 5.1. Network Topology Following the method in Section 3 (τmin = 10, wmin = 3), we construct the co-occurrence network G10 , which contains |V10 | = 274 nodes and |E10 | = 68 edges. The network exhibits extreme sparsity: the edge  density is ρ = |E|/ |V2 | = 0.18%, meaning that more than 99.8% of potential adapter pairs never co-occur. Figure 4 shows the evolution of the network under different frequency thresholds, intuitively illustrating the rapid contraction of the co-occurrence structure as the threshold increases. The degree distribution further reveals the sparsity of the network: 81.4% of nodes (223 out of 274) have degree zero, i.e., they have no significant co-occurrence partners. Among the 51 connected nodes, the average degree is 2.67 (95% confidence interval: [2.21, 3.12]) and the maximum degree is 6. This structural characteristic indicates that the vast majority of adapters operate independently, and the search space for co-occurrence-based preloading is far smaller than the combinatorial space of all possible adapter pairs. When we use greedy modularity maximization as a fragmentation measure, we identify 233 communities, but the vast majority of these communities are single nodes (i.e., isolated adapters with no edges). Therefore, apart from the above zero-degree statistics, they do not provide additional structural information. Consequently, we take the non-trivial connected components (Figure 5) as the main structural view. Only 10 communities have a size of at least 2 (i.e., non-trivial), among which the largest contains 8 adapters. This fragmented structure indicates that adapter co-occurrence relationships are mainly driven by specific and narrow use cases, rather than forming a broad cross-ecosystem collaboration network. Figure 5 zooms in on the connected core of the network. The 51 connected nodes (68 edges) form 9 non-trivial connected components, with the largest component containing 15 adapters and the second largest containing 8. The connected core is dominated by a few base model ecosystems: 72.5% (37 out of 51) of the connected adapters have a dominant base model that belongs to the top 8 models by record count in the processing trace, and the three largest model groups together account for 30 of the 51 adapters. The above results indicate that co-occurrence families are essentially internal to base model ecosystems, rather than cross-model adapter alliances.

Figure 4: Evolution of the co-occurrence network with the frequency threshold τmin . (a)–(e) Network topology at τmin = 1, 5, 10, 20, 50: node size is proportional to adapter frequency, color represents degree centrality (with a unified mapping across panels), gray dots are isolated adapters that only satisfy the frequency threshold, and the red box marks the default threshold 10; (f) network-scale metrics (number of nodes, number of edges, number of connected nodes, number of non-trivial connected components) versus τmin on a log scale. The co-occurrence edge threshold is fixed at wmin = 3.

8

Figure 5: Connected core of the co-occurrence network (degree > 0; 51 nodes, 68 edges, edge weight ≥ 5). Node size is proportional to adapter frequency, and color denotes the dominant base model of each adapter (gray = other models). The two largest connected components contain 15 and 8 adapters, respectively.

5.2. Jaccard Similarity Among the Top 50 Adapters To further examine whether high-frequency adapters exhibit significant co-occurrence preferences, we compute the Jaccard similarity among the top 50 adapters ranked by occurrence frequency. The results show that the Jaccard similarity among the top 50 adapters is generally extremely low (mean J = 0.0076 across all 1,225 pairs, Bootstrap 95% confidence interval: [0.0016, 0.0166]). Only 36 pairs have ever co-occurred, and only 7 pairs have a Jaccard similarity greater than 0.1. The permutation test (1,000 random shuffles of adapter labels) gives p = 1.0, indicating no significant difference between the observed mean Jaccard similarity and the random expectation. This further confirms the extreme sparsity of the co-occurrence network. The few outlier pairs with J > 0.1 are all attached to the same base model, indicating that such co-occurrences mainly arise from a shared base model environment rather than an inherent pairing preference between adapters. 5.3. Base Model–Adapter Association To reveal the organizational driving forces behind the co-occurrence network structure, we analyze the association patterns between base models and adapters, and examine whether adapter co-occurrence is mainly driven by a shared base model environment. Figure 6 shows a bipartite graph composed of the top 8 base models and their respective top 10 adapters, with a total of 70 edges. Square nodes represent base models, and node size is proportional to their number of records in the processing trace; these 8 models together account for 80.2% of all records. Circular nodes represent adapters, and node size is proportional to their occurrence frequency. Color represents the dominant base model, and a black border indicates that the adapter has appeared across models. As shown in Figure 6, the usage patterns between base models and adapters exhibit clear clustering characteristics, and adapters tend to form strong associations with specific base models. The top-ranked base model by occurrence frequency accounts for 30.7% of all requests in the request trace, and its adapter usage is highly concentrated around several core adapter sets. At the adapter level, 78.7% of the 874 adapters are attached to only a single base model, and cross-model sharing is an exception rather than the norm. At the edge level, 66.2% (45 out of 68) of the significant edges in the network connect adapters that share the same dominant base model. If pairings were random, the expected proportion would be 25.1%, meaning that same-model edges are over-expressed by approximately 2.6× relative to the random expectation. The remaining 33.8% (23 edges) are cross-model edges, far below the random expectation of 74.9%, indicating 9

that cross-model co-occurrence is a non-trivial signal after filtering, reflecting real but uncommon affinity between adapters across models. At the request level, in 90.6% of multi-adapter requests all adapters share the same dominant base model, and 89.7% of within-request adapter pairs are same-dominant-model pairs. This model-driven clustering pattern directly explains the source of community structure in the co-occurrence network. Using the number of base models in which an adapter appears to measure its spread breadth, as shown in Figure 7(c), only 5 adapters (0.6%) appear in no fewer than 18 different base models, constituting “core” adapters; 83 (9.5%) appear in 3–17 models; and the remaining 786 (89.8%) appear in no more than 2 models. The average frequency of core adapters is 187.2, significantly higher than the median of 3 for peripheral adapters, indicating that “cross-model spread breadth” and “usage frequency” are highly coupled. A small number of high-frequency adapters span almost the entire model ecosystem, while the vast majority of adapters are single-model customization products. Therefore, the “core–periphery” bipolar structure provides direct guidance for cache design: core adapters should be permanently resident (because they are attached to almost all models), while peripheral adapters can be loaded on demand.

Figure 6: Bipartite graph of base models and LoRA adapters (top 8 models with their top 10 adapters each; 70 edges). Squares = base models (size proportional to the number of processing-trace records; together accounting for 80.2% of records); circles = adapters (size proportional to frequency; color = dominant base model; black border = cross-model adapter).

5.4. Sensitivity Analysis To evaluate the robustness of the above network findings, we examine the impact of different frequency thresholds τmin on the network structure and core conclusions, with threshold values ranging from 5 to 50. Table 3 reports the network scale, sparsity, and same-model edge over-expression ratio under each threshold. As shown in Table 3, the sparsity characteristic remains robust across all thresholds: even under the most relaxed threshold (τmin = 5), the edge density is 0.08%, which is at an extremely low level. The same-model edge over-expression ratio is significant across all thresholds, ranging from 2.6× to 3.5× (under the unified weighted random expectation of 25.1%), and reaches its strongest at τmin = 50 (3.5×). This indicates that the conclusion of model-driven co-occurrence holds across the full threshold range and tends to strengthen as the significance filtering criterion increases. In terms of community structure, the total number of communities (including isolated nodes) decreases monotonically as the threshold increases (from 412 to 49), reflecting that filtering low-frequency adapters effectively reduces fragmentation noise. At the same time, the number of non-trivial connected components 10

also decreases monotonically (from 13 to 5), while the main connected skeleton remains stable, indicating that the core connected structure of the network is insensitive to threshold selection. Table 3: Sensitivity of network characteristics to the frequency threshold τmin (wmin = 3) τmin

|V | |E| Density (%) Communities Non-trivial comps. Same-model over-expr.

5 10 20 50

463 274 154 63

83 68 49 23

0.08 0.18 0.42 1.18

412 233 125 49

13 10 8 5

2.7× 2.6× 2.8× 3.5×

Note: communities include isolated-node communities; non-trivial comps. are connected components containing at least one edge; same-model over-expr. is the ratio of the proportion of same-dominant-model edges to the weighted random expectation (25.1%), where the weights are the proportions of adapter occurrences on each model to all occurrences.

Table 4 further provides the complete sensitivity analysis results across all tested thresholds (under the wmin = 3 scope, consistent with Table 3). The results in this table also show that the sparsity characteristic is robust across all thresholds, and the same-model edge over-expression ratio is between 2.6× and 3.5× and significant (under the weighted random expectation scope). The results in Tables 3 and 4 jointly confirm that rare adapters are the main contributors to network fragmentation, while model-driven co-occurrence is the core mechanism throughout all threshold settings. Table 4: Complete sensitivity analysis across the frequency threshold τmin (wmin = 3) τmin

|V | |E| Density (%) Communities Non-trivial comps. Largest comp.

5 10 15 20 30 50

463 274 195 154 108 63

83 68 52 49 31 23

0.08 0.18 0.27 0.42 0.54 1.18

412 233 163 125 90 49

13 10 9 8 5 5

16 15 13 13 12 11

6. Temporal Evolution The static structural analysis in Section 5 revealed the “spatial organization” of co-occurrence relationships, i.e., who co-occurs with whom and in what structure they cluster. We further address another key question: how do these structures evolve over time? As mentioned earlier (Section 3), adapter-level time-series metrics cannot be strictly computed due to data scope limitations. Therefore, we characterize dynamic properties from three strictly reproducible dimensions: request level, model level, and cross-model structure. 6.1. Request-Level Diurnal Pattern Figure 7(a) shows the daily proportion of requests carrying LoRA adapters in the request trace (26,823 requests, 24 days). As shown in Figure 7(a), this proportion fluctuates between 0% and 37.1%, with a mean of 14.7% and a coefficient of variation (CV) as high as 0.58. On November 16 and 17, there were no LoRA requests at all (with 93 and 86 requests on those days, respectively, far below the daily average of 1,118). After excluding the incomplete first and last days, the proportion range narrows to 0%–24.8% (mean 13.4%, CV 0.54), but the fluctuation magnitude remains significant. These results indicate that there is no stable diurnal baseline for the request-level LoRA usage proportion, and the day-to-day variation is significant (note that this scope is not directly comparable to the overall processing-trace statistic of 22.3% [5], which is based on 68,195 processing records). This finding implies that the potential benefit of cache warm-up strategies relying on fixed diurnal patterns for adapter loading has a limited upper bound, and an online adaptive adjustment mechanism needs to be introduced. 6.2. Model-Level Evolution Figure 7(b) shows the variation in the number of daily active base models. The number of daily active models fluctuates between 6 and 62, with a mean of 34.2, accounting for approximately 39.8% of all 86 models. Compared with the high variability of the request-level proportion, the temporal stability at the model level is significantly higher: the mean Jaccard similarity of active model sets between adjacent days is 0.497, while the weekly (7-day window) mean Jaccard similarity reaches 0.696, indicating that about 70% of the model set overlaps between consecutive weeks—the model level is a relatively stable residency unit. In 11

terms of active days, 17 models are active for no fewer than 18 days, forming a persistent model core; another 20 models are active for no more than 2 days. At a finer temporal granularity, using a 1-hour window to measure the 12-hour drift of the top-10 models by request popularity, the average churn rate is 54.5%, i.e., about half of the top-10 hottest models are replaced after 12 hours. This metric characterizes the rapid rotation of request-arrival popularity (measured by a model-level proxy), corroborating the observation of the Rock system [6] regarding “drastic changes within several hours,” and directly supporting the engineering judgment that the lookup-table refresh period should not exceed 12 hours. The corresponding adapter-level turnover rate cannot be directly computed due to data limitations (Section 3), but its upper bound is no lower than the model-level popularity rotation. 6.3. “Core–Periphery” Structure of Adapter Cross-Model Spread The previous two subsections characterized temporal evolution features from the request level and the model level, respectively. This subsection further focuses on the spread breadth of adapters across base models. As shown in Figure 7(c), the adapter ecosystem exhibits a “core–periphery” bipolar structure: 5 adapters (0.6%) span no fewer than 18 base models, constituting the core, while the vast majority of adapters (786, 89.8%) are attached to no more than 2 models, belonging to the periphery. Combining this static structure with the aforementioned request-level and model-level dynamic features, a complete picture of adapter ecosystem evolution emerges: the model level maintains stable residency (weekly Jaccard similarity 0.696), request popularity rotates rapidly (12-hour churn rate 54.5%), and the adapter level exhibits a polarized pattern of “a small number of cores spanning the whole ecosystem and a large number of peripheral single-model customizations.” The implications of these three layers of structure for cache system design are further elaborated in Section 7. 7. Implications for System Design Based on the characterization of adapter co-occurrence patterns, model-driven clustering, temporal evolution, and latency characteristics in Sections 4–6, we further distill three system design principles with direct engineering guidance: co-occurrence-based preloading, time-aware two-tier caching, and centrality-aware scheduling and cache replacement. Each principle is elaborated below. 7.1. Co-occurrence-based Preloading The co-occurrence network analysis shows that the edge density is below 0.2%, i.e., more than 99.8% of adapter pairs never co-occur. However, when co-occurrence does occur, its pattern is highly concentrated in a small number of “within-model adapter pairs”: in 90.6% of multi-adapter requests all adapters share the same dominant base model; in 85.8% of multi-adapter requests all adapter pairs form significant co-occurrence edges (co-occurrence ≥ 5 times), while 88.6% contain at least one pair of significant co-occurrence edges. Further statistics show that within-base-model adapter co-occurrence edges are over-expressed relative to the random expectation, with a ratio of approximately 2.6×. These results indicate that although co-occurrence is generally sparse and rare, its occurrence has strong structural and within-model clustering properties. Based on the above findings, we precompute for each adapter ai the top-k partners with the highest co-occurrence counts, forming the preload candidate setSP (ai ) (as described in Algorithm 1). When an online request arrives, the system loads the adapters in ai ∈Ar P (ai ) that are not yet cached, where Ar is the set of adapters involved in the current request. This strategy aims to exploit the local concentration of co-occurrence patterns to improve the preloading hit rate under limited cache space. Algorithm complexity analysis. Offline construction of the preload candidate set P (ai ) requires scanning the co-occurrence matrix C, with a time complexity of O(N 2 ), and performing top-k selection, with a time complexity of O(N k log k), where N = |A| = 874. When k = 3, this process can be completed in milliseconds. For online lookup, the overhead per request is O(|Ar | · k). Since |Ar | ≤ 6 in this trace, this overhead is negligible. The storage overhead of the lookup table is O(|A| · k), i.e., 874 × 3 ≈ 2,600 records, which can be fully resident in memory. To evaluate the coverage benefit of the above preloading strategy, we randomly split the 1,090 multiadapter requests into 80% training and 20% testing. A co-occurrence lookup table is built on the training set, and coverage is evaluated on the test set. Averaged over 5 random seeds, the results show that k = 1 covers 57.6% of test-set co-occurrence pairs (range [55.9%, 58.8%]), k = 3 covers 79.5% (range [78.1%, 81.3%]), and k = 10 covers 94.3% (range [91.6%, 96.6%]). When k = 20, the coverage saturates, with a mean of 94.5% (range [91.9%, 96.8%]), indicating diminishing marginal returns from increasing the preload table size. Note that the above coverage is obtained with a conservative one-way lookup, i.e., a pair (a, b) is counted as covered only when b appears in the top-k of a; if the union lookup in Algorithm 1 is adopted (loading the 12

Figure 7: Temporal evolution. (a) Daily LoRA request share (request trace, 24 days; mean 14.7%; no LoRA requests on November 16–17); (b) daily active base model count (mean 34.2; adjacent-day model Jaccard mean 0.497, weekly model Jaccard mean 0.696); (c) distribution of adapter cross-base-model spread (≤ 2 models: 786; 3–17: 83; ≥ 18: 5).

Algorithm 1 Co-occurrence-aware adapter preloading Require: Co-occurrence matrix C ∈ NN ×N , budget k ∈ N+ , current GPU cache G ⊆ A Ensure: Preload candidate set L 1: P ← ∅ ▷ initialize lookup table 2: for each adapter ai ∈ A do 3: N (ai ) ← neighbors with weights 4: P (ai ) ← TopK(N (ai ), k) ▷ select top-k by co-occurrence count 5: end for 6: store P for online lookup ▷ offline phase 7: for each request r with adapter set Ar do S 8: L ← ai ∈Ar P (ai ) ▷ union of all preload candidates 9: L←L\G ▷ exclude cached adapters 10: L ← L \ Ar ▷ exclude requested adapters 11: PreloadToGPU(L) ▷ asynchronous prefetch 12: end for 13: return L 13

union of all P (ai )), the actual coverage will be higher. Since the processing trace has no timestamps and the file row order does not represent temporal order, the random split constitutes an empirical upper bound; the coverage remains stable across 5 random seeds (k = 3 range [78.1%, 81.3%]). Considering that the lookup table storage overhead is only about 2,600 records and can be fully resident in memory, an hourly or daily rebuild period is recommended in engineering deployment. Combined with the observation in Section 6 that model popularity changes at a rate of 54.5% within 12 hours, the update period should not exceed 12 hours to maintain the timeliness of the lookup table. In addition, the LoRA diversity of each base model provides a supplementary basis for model-level caching. For models with few adapter types, all adapters can be pre-cached, while for models with high adapter diversity, selective preloading is needed to balance cache resources and hit rates. 7.2. Time-aware Two-tier Caching The second design principle (two-tier caching) is motivated by the following temporal evolution characteristics: the model level is relatively stable (weekly Jaccard similarity of 0.696), while request popularity rotates rapidly (12-hour churn rate of 54.5%), and the request-level LoRA proportion fluctuates greatly (CV 0.58). In addition, only 5 core adapters span no fewer than 18 base models. This heterogeneity of “fast and slow coexisting” indicates that caching cannot be single-tier and must be layered. Based on this, we adopt a two-tier caching architecture. (1) Residency tier. The 5 core adapters (crossmodel spread ≥ 18 models) and the top-50 adapters by frequency (covering 55.9% of occurrences) are set as permanent cache. The GPU memory footprint of a single adapter is on the order of several MB [2], and the total occupancy is about hundreds of MB, which is negligible. Model-level weights (86 models) are dynamically loaded on demand, and the model-level residency window is adjusted on a weekly basis. (2) Prediction tier. For the remaining adapters, an online popularity prediction updated at hourly granularity is maintained. This prediction is based on short-term patterns of request arrivals, rather than relying on a fixed diurnal baseline. The refresh period of the lookup table and warm-up window is capped at 12 hours, and the 54.5% popularity churn rate observed in Section 6 can serve as a trigger threshold for adaptively adjusting the window length. 7.3. Centrality-aware Scheduling and Cache Replacement The occurrence frequency of adapters is moderately positively correlated with their degree centrality in the network (Pearson r = 0.447, p < 0.001, 274-node scope). Notably, adapters with degree centrality ≥ 3 account for only 2.9% of all adapters (i.e., 25 adapters), yet they participate in 89.7% (61/68) of significant co-occurrence edges, indicating that these high-centrality adapters act as “hub” nodes in the co-occurrence network and play a key bridging role in multi-adapter requests. Therefore, we propose a centrality-weighted LRU cache replacement strategy for scenarios with limited GPU memory. An eviction score is defined as Ei = LRUage /di , giving priority to evicting adapters with low centrality and long inactivity, while retaining high-centrality adapters to maintain the cache hit rate of multiadapter requests. This strategy complements the co-occurrence-based preloading proposed in Section 7.1: preloading answers the question of “what to load,” whereas centrality-aware replacement answers the question of “what to evict.” Both share the same data foundation constructed from the co-occurrence network analysis. 8. Summary and Discussion 8.1. Comparison with Prior Work Table 5 compares the core statistical metrics of this paper with the prior analysis results reported in SoCC’25 [5]. Overall, the increase in the total number of adapters from 705 to 874 reflects the growth of the production service during the data collection period. The proportion of requests carrying LoRA adapters rose slightly from 21.2% to 22.3% (both based on the processing-trace scope); the latency degradation caused by a single adapter decreased from 69% to 66.1%. These metrics show high consistency across two independent analyses, which validates the reliability and reproducibility of our data processing. In addition, note that the Gini coefficient of 0.876 reported in SoCC’25 is a metric under the model-request distribution scope, which differs from the adapter-frequency Gini coefficient reported in this paper (0.749) in terms of the statistical object, and the two are not directly comparable. Unlike SoCC’25, which focuses on aggregate statistics, we start from the association structure among adapters and conduct quantitative analysis of the intrinsic organizational regularities of the adapter ecosystem across dimensions including co-occurrence network topology, model-driven mechanisms, cross-model spread structure, and three-layer temporal evolution. 14

Table 5: Comparison of LoRA characterization findings

Metric

SoCC’25

This paper

Unique adapters Proportion of requests carrying LoRA adapters Multi-adapter requests Rank–frequency exponent α Adapter frequency Gini coefficient Latency degradation (1 adapter) Cross-model core adapters (≥ 18 models) Single-model peripheral adapters (≤ 2 models) Co-occurrence edge density Model popularity 12-h churn rate LoRA diversity per model

705 21.2% — — — 69% — — — — —

874 22.3% 1.6% 1.30 0.749 66.1% 5 (0.6%) 786 (89.8%) 0.18% 54.5% mean 19.3

8.2. Limitations This study has the following limitations. (1) Data limitations. As mentioned earlier, in the GenTD26 dataset, adapter-level configurations (processing trace) and timestamps (request trace) belong to two files that cannot be linked. Specifically, the processing trace has no timestamps, the request trace has no adapter identifiers, the model identifier anonymization schemes are different, and common linking fields are lacking. Therefore, we cannot analyze adapter-level time-series metrics, and instead characterize three dimensions: request level, model level, and cross-model spread. If future datasets can provide adapter-level records with timestamps, the actual benefits of preloading and caching strategies under different temporal partitioning schemes can be further verified. (2) Adapter function anonymization. Since adapter identifiers are anonymized, we cannot obtain semantic information about adapter functions, and thus cannot determine whether co-occurring adapters serve complementary objectives (such as joint generation of style + object) or act as alternatives for the same visual effect. (3) Generalization ability. This study is based on only 24 days of observational data from a single production cluster. This window may not capture long-term seasonal trends and is difficult to directly generalize to different deployment scenarios. If more diffusion model inference traces are made public in the future, cross-trace comparative validation will help further improve the generalization and robustness of the conclusions. 8.3. Conclusion We conduct a network-based characterization analysis of LoRA adapter co-occurrence patterns in production diffusion model inference services. Using the GenTD26 dataset and a graph-theoretic analysis framework, we reveal key structural properties of the adapter ecosystem: (1) co-occurrence relationships exhibit extreme sparsity (edge density 0.18%, with more than 99.8% of pairs never co-occurring), and the frequency distribution is heavy-tailed (Gini 0.749, with the top 5% of adapters covering 53.5% of occurrence counts); (2) adapter usage exhibits model-driven clustering, with 90.6% of multi-adapter requests sharing the same dominant base model, and the average LoRA diversity per base model being 19.3; (3) the network structure exhibits a “core–periphery” bipolar structure, with only 5 adapters spanning ≥ 18 base models, while 89.8% of adapters are attached to at most 2 models; (4) temporal evolution exhibits three-layer scale differences, with the model level being relatively stable (weekly Jaccard 0.696), request popularity rotating rapidly (12-hour churn rate 54.5%), and the request-level LoRA proportion fluctuating significantly across days (CV 0.58); and (5) measurable latency scales with the number of adapters (a single-adapter configuration increases latency by 66.1%, with diminishing marginal costs afterwards). The co-occurrence-based preloading strategy is validated by offline experiments (k = 3 covers 81.0% of test-set co-occurrence pairs), and sensitivity analysis verifies the robustness of the conclusions. All data and code involved in this paper have been open-sourced, and readers can obtain and reproduce the experimental results through the following URL: https://gitee.com/liaobin665/ lora-cooccurrence-reproduce.

15

References [1] R. Rombach, A. Blattmann, D. Lorenz, P. Esser, B. Ommer, High-resolution image synthesis with latent diffusion models, in: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022, pp. 10684–10695. [2] E.J. Hu, Y. Shen, P. Wallis, Z. Allen-Zhu, Y. Li, S. Wang, L. Wang, W. Chen, LoRA: Low-rank adaptation of large language models, in: Proceedings of the International Conference on Learning Representations (ICLR), 2022. [3] B. Wei, X. Zhang, Research on electronic health record data generation for diffusion models, Application Research of Computers 41 (12) (2024) 3521–3532 (in Chinese). [4] Z. He, G. Li, Review of personalized image generation methods based on diffusion models, Journal of Software 37 (4) (2026) 1854–1884 (in Chinese). [5] Y. Lin, S. Wu, S. Luo, H. Xu, H. Shen, C. Ma, M. Shen, L. Chen, C. Xu, L. Qu, K. Ye, Understanding diffusion model serving in production: A top-down analysis of workload, scheduling, and resource efficiency, in: Proceedings of the ACM Symposium on Cloud Computing (SoCC), 2025, pp. 1–16. [6] S. Wu, Y. Lin, S. Peng, W. Chen, C. Ma, M. Shen, L. Chen, C. Xu, K. Ye, Rock: Serving multimodal models in cloud with heterogeneous-aware resource orchestration for thousands of LoRA adapters, in: Proceedings of the IEEE International Conference on Cluster Computing (CLUSTER), 2025, pp. 1–13. [7] L. Chen, Z. Ye, Y. Wu, D. Zhuo, L. Ceze, A. Krishnamurthy, Punica: Multi-tenant LoRA serving, in: Proceedings of Machine Learning and Systems (MLSys), Vol. 6, 2024, pp. 282–295. [8] Y. Sheng, S. Cao, D. Li, C. Hooper, N. Lee, S. Yang, C. Chou, B. Zhu, L. Zheng, K. Keutzer, J.E. Gonzalez, I. Stoica, S-LoRA: Serving thousands of concurrent LoRA adapters, in: Proceedings of Machine Learning and Systems (MLSys), Vol. 6, 2024, pp. 222–237. [9] R. Chen, C. Yu, H. Fu, X. Hu, B. Yang, MixLoRA: An efficient multi-tenant framework for concurrently serving diverse LoRA models in large language models, in: Proceedings of the International Conference on Parallel Processing (ICPP), 2025. [10] S. Li, H. Lu, T. Wu, M. Yu, Q. Weng, X. Chen, Y. Shan, B. Yuan, W. Wang, CaraServe: CPUassisted and rank-aware LoRA serving for generative LLM inference, in: Proceedings of the International Conference on Architectural Support for Programming Languages and Operating Systems (ASPLOS), 2025. [11] C. Reiss, A. Tumanov, G.R. Ganger, R.H. Katz, M.A. Kozuch, Heterogeneity and dynamicity of clouds at scale: Google trace analysis, in: Proceedings of the ACM Symposium on Cloud Computing (SoCC), 2012, pp. 7:1–7:13. [12] E. Cortez, A. Bonde, A. Muzio, M. Russinovich, M. Fontoura, R. Bianchini, Resource Central: Understanding and predicting workloads for improved resource management in large cloud platforms, in: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP), 2017, pp. 153–167. [13] Q. Weng, W. Xiao, Y. Yu, W. Wang, C. Wang, J. He, Y. Li, L. Zhang, W. Lin, Y. Ding, MLaaS in the wild: Workload analysis and scheduling in large-scale heterogeneous GPU clusters, in: Proceedings of the USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2022, pp. 945–960. [14] M. Jeon, S. Venkataraman, A. Phanishayee, J. Qian, W. Xiao, F. Yang, Analysis of large-scale multitenant GPU clusters for DNN training workloads, in: Proceedings of the USENIX Annual Technical Conference (ATC), 2019, pp. 1041–1056. [15] Q. Hu, P. Sun, S. Yan, Y. Wen, T. Zhang, Characterization and prediction of deep learning workloads in large-scale GPU datacenters, in: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis (SC), 2021, pp. 104:1–104:15. [16] Z. Ye, W. Gao, Q. Hu, P. Sun, X. Wang, Y. Luo, T. Zhang, Y. Wen, Deep learning workload scheduling in GPU datacenters: A survey, ACM Computing Surveys 56 (6) (2024) 1–38. [17] Y. Ming, Y. Chen, J. Zhao, Multimodal large model training strategy for functional image data, Application Research of Computers 42 (11) (2025) 3421–3429 (in Chinese). 16

[18] A. Borzunov, M. Ryabinin, A. Chumachenko, D. Baranchuk, T. Dettmers, Y. Belkada, P. Samygin, C. Raffel, Distributed inference and fine-tuning of large language models over the internet, in: Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2023. [19] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C.H. Yu, J.E. Gonzalez, H. Zhang, I. Stoica, Efficient memory management for large language model serving with PagedAttention, in: Proceedings of the ACM Symposium on Operating Systems Principles (SOSP), 2023, pp. 611–626. [20] P. Patel, E. Choukse, C. Zhang, A. Shah, Í. Goiri, S. Maleki, R. Bianchini, Splitwise: Efficient generative LLM inference using phase splitting, in: Proceedings of the ACM/IEEE International Symposium on Computer Architecture (ISCA), 2024, pp. 118–132. [21] Y. Fu, L. Xue, Y. Huang, A.-O. Brabete, D. Ustiugov, Y. Patel, L. Mai, ServerlessLLM: Low-latency serverless LLM serving, in: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2024, pp. 161–179. [22] Y. Lin, S. Peng, C. Lu, et al., FlexPipe: Adapting dynamic LLM serving through inflight pipeline refactoring in fragmented serverless clusters, in: Proceedings of the ACM European Conference on Computer Systems (EuroSys), 2026. [23] X. Meng, V. Pappas, L. Zhang, Improving the scalability of data center networks with traffic-aware virtual machine placement, in: Proceedings of the IEEE INFOCOM, 2010, pp. 1–9. [24] M. Isard, M. Budiu, Y. Yu, A. Birrell, D. Fetterly, Dryad: Distributed data-parallel programs from sequential building blocks, in: Proceedings of the ACM European Conference on Computer Systems (EuroSys), 2007, pp. 59–72. [25] S. Luo, H. Xu, C. Lu, K. Ye, G. Xu, L. Zhang, Y. Ding, J. He, C. Xu, Characterizing microservice dependency and performance: Alibaba trace analysis, in: Proceedings of the ACM Symposium on Cloud Computing (SoCC), 2021, pp. 412–426. [26] W. Xiao, R. Bhardwaj, R. Ramjee, M. Sivathanu, N. Kwatra, Z. Han, P. Patel, X. Jiang, Q. Lin, F. Yang, L. Zhou, Gandiva: Introspective cluster scheduling for deep learning, in: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2018, pp. 595–610. [27] W. Xiao, S. Ren, Y. Li, Y. Zhang, P. Hou, Z. Li, F. Yang, L. Zhou, J. Wang, AntMan: Dynamic scaling on GPU clusters for deep learning, in: Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2020, pp. 533–548. [28] A. Clauset, M.E.J. Newman, C. Moore, Finding community structure in very large networks, Physical Review E 70 (6) (2004) 066111. [29] A. Clauset, C.R. Shalizi, M.E.J. Newman, Power-law distributions in empirical data, SIAM Review 51 (4) (2009) 661–703. [30] Q. Zhang, M. Chen, A. Bukharin, P. He, Y. Cheng, W. Chen, T. Zhao, AdaLoRA: Adaptive budget allocation for parameter-efficient fine-tuning, in: Proceedings of the International Conference on Learning Representations (ICLR), 2023. [31] H. Liu, D. Tam, M. Muqeeth, J. Mohta, T. Huang, M. Bansal, C.A. Raffel, Few-shot parameter-efficient fine-tuning is better and cheaper than in-context learning, in: Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2022. [32] D.J. Kopiczko, T. Blankevoort, Y.M. Asano, VeRA: Vector-based random matrix adaptation, in: Proceedings of the International Conference on Learning Representations (ICLR), 2024. [33] S. Li, H. Lu, T. Wu, M. Yu, Q. Weng, X. Chen, Y. Shan, B. Yuan, W. Wang, Toppings: CPUassisted, rank-aware adapter serving for LLM inference, in: Proceedings of the USENIX Annual Technical Conference (ATC), 2025, pp. 613–629. [34] X. Shi, J. Du, Z. Chen, et al., AuLoRA: Fine-grained loading and computation orchestration for efficient LoRA LLM serving, in: Proceedings of the IEEE International Conference on Computer Design (ICCD), 2025. 17

[35] N. Iliakopoulou, J. Stojkovic, C. Alverti, T. Xu, H. Franke, J. Torrellas, Chameleon: Adaptive caching and scheduling for many-adapter LLM inference environments, in: Proceedings of the IEEE/ACM International Symposium on Microarchitecture (MICRO), 2025. [36] R. Li, H. Bao, H. Wang, P. Shi, JointSerLoRA: Cache-aware LoRA inference serving, in: Proceedings of the IEEE/ACM International Symposium on Quality of Service (IWQoS), 2025. [37] S. Li, L. Yang, X. Jiang, et al., SwiftDiffusion: Efficient diffusion model serving with add-on modules, arXiv preprint arXiv:2407.02031, 2024.

18

Record · ID 1028665 · SHA-256 31850733a56d287b
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.