SmartGen: Seamless Disaggregated LLM Inference with Selective KV Cache Transfer Xuchuan Luo1 , Jiacheng Shen2 , Xin Wang1 , and Yangfan Zhou1 1 College of Computer Science and Artificial Intelligence, Fudan University 2 Duke Kunshan University
However, the disaggregated scheme hinders users [50] from deploying self-hosted LLM inference systems due to the network bandwidth bottleneck. Specifically, key-value (KV) caches, i.e., intermediate attention states, must be transferred from the prefill nodes to the decoding nodes when serving each inference request. Existing systems [45, 47, 58] transfer KV caches in a layer-by-layer manner to overlap data transfer with prefill computation. However, according to our experiments on Llama [61] and Qwen [62] models on Alibaba cloud GPU instances with L20 GPUs and 25 Gbps RDMA network, transferring long-context prompts, i.e., 48K tokens, takes 6.5× more time than the prefill computation. The reason is twofold. On the one hand, since the size of the KV cache is large and grows proportionally with the sequence length, batch size, and model size, it inevitably saturates the limited network bandwidth [45, 85]. On the other hand, most cloud GPU instances only offer limited network bandwidth [8, 10, 11, 51], e.g., the L4 and L40S instances of AWS [51] and A100 instances of Google Cloud [10] provide only 10–35 Gbps network. The KV cache transfer hinders the decoding stage and can account for up to 42.2% of the job completion time [82], hurting the user experience with increased time-to-second-token (TTST). We refer to this problem as a stage-transition stall. In this paper, we propose to accelerate KV cache transfer by leveraging the inherent importance distribution of KV caches, i.e., not all KV entries are required during the decoding stage. Specifically, existing works have widely exploited the dynamic sparsity in KV caches to accelerate the decoding computation [7, 14, 24, 37, 83]. If the sparsity can be utilized in the transfer phase, the network bottleneck can be significantly relieved since only a small amount of important KV entries need to be transferred during the prefill stage. However, two challenges must be addressed before this idea becomes a practical KV cache transfer solution. (1) Accurately identifying important KV entries during the prefill stage. Since only a limited number of KV entries can be transferred during the prefill stage, the decoding stage has to fetch the rest of the required ones to ensure
arXiv:2607.28150v1 [cs.DC] 30 Jul 2026
Abstract Disaggregating the prefill and decoding stages of large language model (LLM) inference into two separate sets of nodes is widely adopted in today’s LLM serving systems. However, such an architecture poses significant challenges for selfhosted LLM deployments on rented cloud instances, since transferring enormous key-value (KV) caches between disaggregated nodes can easily saturate the limited inter-node network bandwidth. In this paper, we propose to mitigate the network bottleneck by selectively transferring essential KV cache entries across the two stages. There are two challenges to achieve selective KV cache transfer, i.e., accurate KV selection during the prefill stage, and efficient KV fetching during the decoding stage. To address these challenges, we design SmartGen, a KV cache transfer engine that allows seamless disaggregated LLM inference with three data transfer paths. Specifically, we leverage 1) a profile-based proactive transfer path to identify and push essential KV cache entries to the decoding node during the prefill stage, 2) a parallel on-demand transfer path to simultaneously fetch remote and local KV cache entries during the decoding stage, and 3) a speculative transfer path to finally deliver all KV caches to the decoding node. Experimental results show that SmartGen reduces time-to-second-token by up to 4.3× compared with the typical full KV cache transfer approach while offering comparable subsequent decoding performance and accuracy.
1
Introduction
Large language models (LLMs) have been widely deployed in real-world applications, e.g., chatbots [43, 47], programming assistants [6, 49], and summarization tools [22, 66], thanks to their powerful capabilities in language understanding and generation. Existing LLM inference systems disaggregate the prefill and decoding stages of inference computation into separate sets of nodes to accommodate their distinct computational characteristics [12, 15, 45, 47, 85]. This enables independent and flexible resource allocation for each stage, maximizing the overall throughput. 1
model accuracy. The more important KV entries are identified and transferred, the fewer need to be fetched by the decoding node. However, the state-of-the-art sparse attention algorithms [7, 24, 37, 59] depend on queries or hidden states of tokens generated during the decoding stage to define the importance of KV entries. This makes accurate identification of important KV entries challenging in the prefill stage since the dependent information is not yet available.
2 2.1
Background Large Language Model Inference
LLMs are composed of multiple stacked transformer blocks. Each transformer block consists of an attention layer and a feed-forward network (FFN) [63]. Transformer computation. The input tensor Xin ∈ RN×D , representing N query tokens with model dimension D, is first layer-normalized and fed into the attention layer. It is multiplied by three projection matrices Wq , Wk , and Wv ∈ RD×D to produce the query, key, and value matrices, i.e., Q, K, and V ∈ RN×D . Each Q, K, and V consists of H attention heads and are reshaped to H × N × d, where H × d = D. Each head T √ )V . The outputs from all computes attention as so f tmax( QK d heads are concatenated and projected, then passed through a residual add and a layer normalization before entering the FFN. The FFN consists of two linear layers with an activation operation in between. Its output goes through another residual add to produce the final output Xout ∈ RN×D , maintaining the shape of Xin for the next transformer block. Generative inference and KV caches. Generative LLM inference consists of two stages: prefill and decoding. In the prefill stage, the LLM processes the input sequence, i.e., the prompt, to generate the first output token. In the decoding stage, the LLM uses the latest token to produce the next one, forming an autoregressive process for token generation that repeats until completion. During this process, the LLM computes the attention score of each newly generated token with all previous tokens in every iteration. To avoid redundant computation, the K and V of previous tokens are cached in memory, which is known as the KV cache.
(2) Efficiently fetching KV entries during the decoding stage. To ensure model accuracy, the decoding node must fetch missing KV entries from the prefill nodes before attention computations. This process degrades decoding efficiency by adding a network round-trip to the critical path of each attention layer. Since the remote KV fetching is required in every iteration, the cumulative overhead becomes substantial when generating a long text. To address the above challenges, we design SmartGen, an importance-aware KV cache transfer engine that ensures seamless execution of disaggregated inference for self-hosted LLM inference systems. Specifically, SmartGen categorizes KV cache tokens into three classes, i.e., universally important, context-dependent, and less important. 1) Universally important tokens are tokens in positions that are considered important regardless of input prompts. We design a profilebased proactive transfer scheme to identify and push these tokens ahead of the decoding stage. 2) Context-dependent tokens are other tokens that are specifically important to some queries. We propose a parallel on-demand transfer scheme to efficiently fetch these tokens during decoding, in which network round-trips overlap with local KV cache loading. 3) For the remaining less important tokens, a speculative transfer scheme opportunistically delivers them during network idle periods, thereby amortizing cumulative transfer overhead without impacting the critical path.
2.2
P/D Disaggregation
The use of KV caches makes the decoding stage memoryintensive, exhibiting characteristics distinct from the prefill stage. This asymmetry leads to interference between the two stages when batched together on the same hardware. To mitigate such interference, several works [31,45,58,85] propose to disaggregate the two stages onto separate GPUs. To further improve resource efficiency and maximize throughput in real deployments, Mooncake [47] introduces a KV-cache-centric architecture that fully decouples prefill and decoding stages into two separated clusters and offloads KV caches in the CPU memory pool. The prefill and decoding clusters are connected with RDMA NICs, which enable them to communicate with each other using one-sided verbs (e.g., READ, WRITE) or two-sided verbs (e.g., SEND, RECV). However, regardless of the type of verb used, inter-node data transfer is fundamentally constrained by network bandwidth. Given the large KV cache size, the limited RDMA bandwidth can easily become a significant performance bottleneck [45, 82, 85]. Things get worse when individuals want to deploy disaggregated inference with cloud instances, e.g., EC2 [51] and ECS [8], since they usually offer limited network bandwidth. HACK [82]
We implement SmartGen, integrate it with various sparse attention algorithms [24, 37], and evaluate it on Llama3.1 [61], Qwen3 [62], Gemma-3 [60], and Phi-4 [1] models with various real-world workloads [4, 23, 26, 32]. Compared with full KV cache transfer, SmartGen achieves up to 4.3× lower TTST with similar subsequent decoding performance and accuracy. In summary, this paper makes the following contributions: • We identify the stage-transition stall issue when hosting disaggregated LLM inference systems on low-cost cloud instances, based on experimental analyses. • We propose the idea of importance-aware KV cache transfer to eliminate the stall. We also address the challenges of adopting selective transfer with SmartGen. • We show that SmartGen is a practical KV cache transfer with our experiments, achieving up to 4.3× lower TTST compared with existing approaches. 2
Light Attention
k=3
KV Index 7 1 6 8 9 11 BH × k Selected KV E7 E1 E6 E8 E9 E11
8 KV Cache (B=1, H=2, N=8) E0 E1 E2 … E6 E7 E8 E9 E10 … E14 E15
Latency (s)
KV Selection
(BH × N × d)×2
GPU side CPU side Figure 1: An example of KV selection and KV fetching from a batch of size B, where each head of each request selects its own top-k entries. Ei ∈ Rd represents the ith KV entry.
6
Prefill (L20) Prefill (A100) Transfer (25 Gbps) Transfer (15 Gbps)
4 2
2 Llama-3.1-8B Qwen3-8B Qwen3-14B (32 layers) (36 layers) (40 layers) (a) Varying model sizes.
0
0
12
24
36
48
60
# Batched Input Tokens (K) (b) Varying batch sizes.
Figure 2: Comparison of the full KV cache transfer overhead and the prefill computation latency on Llama and Qwen models.
3.1
addresses network bandwidth via a 2-bit homomorphic KV cache quantization scheme that avoids dequantization overhead. In this paper, we instead reduce the network burden by exploiting the dynamic sparsity of the KV cache, which is orthogonal to quantization-based approaches.
Motivation: Issues with Full Transfer
In disaggregated inference systems, the KV cache transfer overhead directly affects the delivery of the second token to users, making TTST as important as time-to-first-token (TTFT). Specifically, LLM serving systems use an output buffer to store generated tokens and deliver them progressively to users [5]. If the buffer empties before new tokens are available, users experience stalls. Since the first token is sent to the user almost immediately, a long delay in producing the second token leaves the buffer empty, resulting in a user stall. Existing systems [12, 31, 45, 47, 58] hide KV cache transfer latency by overlapping it with prefill computation. However, our observations show that the overlapping strategy fails to fundamentally address the bandwidth bottleneck. Observation: KV cache transfer cannot be overlapped by prefill computation. In the P/D disaggregation architecture, the compute-intensive prefill nodes typically adopt small batch sizes to satisfy TTFT SLOs [31, 45, 85]. Multiple prefill nodes are often adopted to handle highly concurrent user requests. On the decoding side, due to the memory-intensive nature of the decoding stage, batch sizes of decoding nodes are always larger than those of prefill nodes [31, 45, 47, 85]. Consequently, KV caches have to be transferred from multiple prefill nodes to a single decoding node, saturating the decoding node’s limited network bandwidth. The problem worsens as model size and sequence length grow. Figure 2a shows the prefill and KV cache transfer latency of 48K batched input tokens under three different models, respectively. Larger models generate larger KV caches, and as a result, the KV cache transfer latency increases by 1.3×. In contrast, the prefill latency remains below 0.6 seconds with a data parallelism (DP) degree of 6. This makes it hard for the prefill stage to hide the KV cache transfer as the model size increases. Figure 2b shows the results on the Qwen3-14B model with a DP degree of 6 and various numbers of batched input tokens. As the number of tokens increases, the gap between transfer latency and prefill latency on our testbed reaches up to 4.7×, indicating a more severe stage-transition stall. If the network bandwidth is further limited to 15 Gbps (e.g., using a smaller L20 instance, ecs.gn8is.2xlarge [8], for decoding), this gap can be amplified up to 7.1×. If high-end GPUs are adopted during the prefill stage, it would be even harder to overlap the KV cache transfer with
Dynamic KV Cache Selection
Due to the large size of KV caches, many studies [7, 27, 37, 38, 59, 70, 83] propose dynamically selecting and loading only essential KV entries for attention computation, leveraging the sparsity in KV caches. This paradigm has also been explored and adopted in industry, e.g., in DeepSeek’s DSA [14], Google’s Spark [73], and Microsoft’s MInference [41]. It reduces computational load, preserves model accuracy, and works well with KV cache offloading [27, 37, 53], making it well-suited for deployment on cloud instances with limited GPU memory. As shown in Figure 1, dynamic KV selection algorithms output a KV index matrix of shape BH × k, where each row stores k indexes indicating the top-k KV entries selected for an attention head in a batch. Only the selected entries are loaded onto the GPU via host-to-device copy for attention computation. In this paper, we define KV selection as identifying the KV entries to use, and KV fetching as loading them onto the GPU.
3
4 0
(BH × k × d)×2
2.3
6
Prefill (DP=6) Prefill (DP=4) Transfer (25 Gbps) Transfer (15 Gbps)
Analysis of KV Cache Transfer
This section motivates the idea (§ 3.1) and presents the challenges (§ 3.2) of adopting selective KV cache transfer for selfhosted disaggregated LLM inference. All the experiments in this section are conducted with 3 ecs.gn8is-2x.8xlarge prefill instances and one ecs.gn8is.4xlarge decoding instance on Alibaba Cloud [8]. Each prefill instance is equipped with two NVIDIA L20 GPUs. The decoding instance has one L20 GPU and a 25 Gbps RDMA network interface. We offload KV caches to host memory before transfer, as GPU memory on low-cost cloud instances is limited. The offloaded KV caches can also serve as prefix caches for reuse, like existing P/D disaggregated systems [12, 47]. We assume a prefix cache hit ratio of 75% in host memory, reflecting a representative value across reported hit ratios ranging from 60% to 90% in various long-context scenarios [19, 47, 64, 75, 84]. 3
70
On-demand Ratio (%) TBT (x10 ms)
TBT (x10 ms)
60
TBT On-demand Ratio 60
55
50
50
Prefill Nodes
Optimal + KV Reusing Computation Time (Ideal)
GPU new KV
40 50
Random Sequential Optimal
Strategy
(a) KV selection.
30
40
context-dependent
CPU Memory
2 16 32 48 64 80 96
Token ID
(b) KV fetching.
KV cache pool
Figure 3: Optimization opportunities in prefill-side KV selection and decoding-side KV fetching on the Qwen3-14B model, assuming that half of KV cache can be transferred during the prefill stage.
Speculative Transfer (§ 4.3) less important
GPU LLM selected KV CPU Memory
KV cache pool
Profile-based Proactive Transfer (§ 4.1) universally important
Figure 4: The overview of SmartGen.
the prefill computation due to the higher execution efficiency. As shown in Figure 2b, with more Tensor Cores and higher VRAM bandwidth, A100 achieves up to a 1.9× speedup over L20 when prefilling 60K tokens, making the KV transfer issue more pronounced. Opportunity: Not all KV entries are equally important. Our work is inspired by the KV cache sparsity, i.e., selecting only the most critical tokens’ KV caches for attention can maintain comparable model accuracy [24,27,37,38,59,70,83]. Based on this, we propose to address the network bandwidth bottleneck with selective KV cache transfer. Specifically, one of the existing dynamic KV cache selection algorithms [24,37] is adopted to exploit KV cache sparsity. During the prefill stage, only part of the essential KV entries are selected and pre-transferred to the decoding node, with the hope that the KV cache transfer can be fully overlapped by the prefill computation. During the decoding stage, to maintain model accuracy, the decoding node fetches any missing KV entries from prefill nodes in an on-demand manner.
3.2
Parallel On-demand Transfer (§ 4.2)
LLM
45
Decoding Node
namic. They rely on query [24,27,38,59] or hidden states [37] generated during the decoding stage to accurately identify important KV entries. Since the decoding stage has not yet started during the prefill stage, the prefill node has no visibility into which KV entries the decoding node will select. Challenge 2: Efficient KV fetching on the decoding node. To ensure decoding accuracy, the dynamically selected KV entries must be locally available before the transformer computation. This requires the decoding node to check and fetch any missing entries from prefill nodes before loading them onto the GPU, introducing a network round-trip on the critical path of the host-to-device copy. As shown in Figure 3a, even if half of the KV cache is pre-transferred, there still remain 34% of KV entries that must be fetched remotely. Worse still, inefficient KV fetching persists throughout the decoding stage. Although later tokens can reuse previously fetched KV entries, the fine-grained dynamic selection always results in new missing entries that must be fetched from prefill nodes. As shown in Figure 3b, remote KV fetching still incurs a sustained 1.1× TBT increase in each decoding step even if previously fetched KV entries are reused.
Challenges of Selective Transfer
Although selective KV cache transfer could relieve the network burden, it introduces new challenges due to the KV selection and KV fetching during the two stages, respectively. Challenge 1: Accurate KV selection on the prefill node. The accuracy of KV selection during the prefill stage directly impacts the overhead of on-demand KV fetching during the decoding stage. We define the on-demand ratio as the ratio of KV entries fetched on demand from the prefill node to the total KV entries required by the decoding node. Transferring more relevant KV entries upfront reduces the ratio, thereby reducing the time-between-tokens (TBT) in the decoding stage. Figure 3a shows the impact of such KV selection with 48K batched input tokens. With an optimal selection strategy, the most critical KV entries are selected and pre-transferred during the prefill stage. This strategy reduces the on-demand ratio from 53% to 34% and optimizes TBT by 1.1×, compared with the typical sequential strategy where KV entries are transferred in memory address order. A similar observation holds when compared with a random strategy. However, achieving the optimal selection strategy is challenging since state-of-the-art KV selection methods are dy-
4
The SmartGen Design
We propose SmartGen, an importance-aware KV cache transfer engine that enables seamless state transitions in disaggregated inference for self-hosted LLMs on the cloud. As shown in Figure 4, SmartGen categorizes KV cache entries into three types and comprises three transfer paths. First, to achieve accurate KV selection during the prefill stage, SmartGen adopts a profile-based proactive transfer to push the universally important KV entries to the decoding node (§ 4.1). Second, to achieve fast KV fetching during the decoding stage, SmartGen conducts a parallel on-demand transfer to fetch missing context-dependent KV entries, which overlaps the network overhead with local KV cache loading (§ 4.2). Finally, SmartGen proposes a speculative transfer to deliver all remaining less important KV entries to the decoding node (§ 4.3). 4
4K
3K
4K
Token ID
2K 4K 6K 8K
3K 6K 9K 12K
0.3 0.2
… I0,M-1
IL-1,0
… IL-1,M-1
block-wise
split top-Kr yes
KV KV
(H × Nb × d)×2
0.1
WRITE
CPU Memory KV Cache Pool KV Nb
Prefill Nodes Decoding Node Figure 6: The process of profile-based proactive transfer, where Nb is the number of tokens in the KV block. The KV mask is used for on-demand transfer, which will be introduced in § 4.2.
0.0
4K 8K 12K 16K
prefill computation. Without loss of generality, the following describes a single profiling run. First, we run a round of model inference on a calibration dataset to calculate the frequency of selecting each KV entry. Then, for each layer, we partition the KV cache along the sequence dimension into M KV blocks. In this way, we can use fixed-sized KV blocks to approximate variable-length regions. Thus, there are in total L · M KV blocks, where L is the number of attention layers. Let Bl,m denote the m-th KV block in layer l. The importance Il,m of block Bl,m is defined as the average frequency: ( ∞, if l = 0 or 1 (1) Il,m = 1 |E(B )| ∑e∈E(Bl,m ) S(e), if 2 ≤ l < L
Figure 5: The frequency of selecting each token in each attention layer on various models [61, 62] and datasets [4, 23, 26, 32].
4.1
I0,0
GPU KV Mask 22…11 … … N…b 02…22
new KV
RDMA
Offline KV Selection Importance Matrix
0.4
layer-wise
LCC MultiField.
1K 2K 3K 4K
SAM.
2K
Qwen3-14B
Selection Frequency
3K
1K
2 14 26 39 2K 3K 4K 0 2 14 26 39 4K 6K 8K 0 2 14 26 39 6K 9K 12K 0 2 14 26 39 8K 12K 16K 0
…
2K
Qwen3-8B
Gov.
1K
2 13 24 35 2K 3K 4K 0 2 13 24 35 4K 6K 8K 0 2 13 24 35 6K 9K 12K 0 2 13 24 35 8K 12K 16K 0
…
Layer ID
Llama-3.1-8B
…
2 11 21 31 0 2 11 21 31 0 2 11 21 31 0 2 11 21 31 0
Profile-based Proactive Transfer
Profile-based proactive transfer is presented to eliminate the stage-transition stall by transferring only essential KV entries. The key challenge lies in accurately identifying essential KV entries during the prefill stage. Inspired by static KV sparsity [69, 70, 77], we first propose to address this challenge in this section by exploiting positional similarity, i.e., universally important KV entries tend to appear at similar positions in input sequences. We then describe how we use the positional similarity to guide the accurate selection and efficient RDMA-based transfer of essential KV entries. Positional similarity in important KV entries. We define a KV entry’s importance as how frequently it is selected by the KV selection algorithm during the decoding stage. Figure 5 shows the importance distribution of KV entries selected by InfiniGen [37]. The results can be applied to other selection algorithms [7,24,59,83] since they all use the magnitude of attention scores to select KV. To simplify comparison, we truncate requests in various datasets [4, 23, 26, 32] to 4K, 8K, 12K, and 16K, respectively. Our key observation is that important KV entries tend to appear in consistent regions of the KV cache across different datasets. For the same model, datasets with different prompt lengths do not affect the relative positional distribution of important KV entries. For instance, in attention layer 8 of Qwen3-14B, each of the last 25% of tokens is selected by over 30% of heads and requests, regardless of the prompt length, while in layer 14, most tokens are not selected by the majority of heads and requests. Similar patterns are also observed in other models of different sizes and architectures, e.g., Llama-3.1-8B. This indicates that, within a certain range of prompt lengths, e.g., 4K-16K, a calibration dataset can approximate the token-importance distribution for requests in that range. Offline KV selection. Based on the above observation, we conduct offline profiling on ranges of prompt lengths to 1) identify important regions in KV cache matrices and 2) help predict how many KV entries can be transferred during the
l,m
Here, E(Bl,m ) is the set of KV entries in KV block Bl,m , and S(e) denotes the number of times entry e is selected in the decoding stage during the offline profiling. KV blocks with higher importance are prioritized for selection and transfer during the online prefill stage. The KV caches of the first two attention layers, i.e., l = 0, 1, are always selected as they are consistently important and required by the decoding stage [37, 38, 59]. We set M to 1K in our implementation. In addition, we also profile the prefill computation time and full KV cache transfer time to know how many KV blocks can be transferred during prefill computation. Based on the profiled prefill time Tp and transfer time Tt , we select the top-Kr most important KV blocks among all L layers, where: Tp L−1 · Tp /Tt = M · (L − 1) · (2) Kr = L · M · L Tt Here, Kr is calculated as the product of the total number of KV blocks and the proportion of transfer time that can be overlapped with prefill computations. As KV cache transfer can only start after the first attention layer finishes computation, we consider only the prefill duration of L − 1 layers, i.e., L−1 L · Tp . The estimation provides an upper bound on the number of KV blocks whose transfer can be fully overlapped with the prefill computation of all but the first transformer block. In practice, we clip Kr between 2M and LM to ensure that KV blocks of the first two layers are always selected. 5
selected KV
SEND
mask=2
split BH × k
local index
mask=1
All Selected KV gather (BH × k × d)×2
selected KV
Algorithm 1: KV Index Splitting and Reordering : kv_index: The KV index of shape BH × k kv_mask: The KV mask of shape B × N Output : local_index: The local KV index array remote_index: The remote KV index array /* Select required mask values from the KV mask matrix */ 1 1 mask ← kv_mask [kv_index] ⃝ 2 if enable reordering then /* For fast KV gathering */ 4 3 mask, order ← SortEachRow(mask) ⃝ 5 4 kv_index ← kv_index [order] ⃝ Input
KV Cache Pool
SEND
CPU Memory
KV Index
PCIe
remote index
RDMA
KV Cache Pool
GPU
Prefill Nodes Decoding Node Figure 7: The process of parallel on-demand transfer.
Online KV transfer. Figure 6 shows the process of profilebased proactive transfer. Each time the prefill stage offloads an attention layer’s KV cache to CPU memory, SmartGen first splits it into M blocks. For each KV block, it checks whether its position belongs to the top-Kr most important ones. If so, SmartGen uses one-sided RDMA WRITE to push the block to the decoding node’s KV cache pool.
4.2
/* Mask values of 1 (2) indicate local (remote) KV entries */ 2 local_index ← kv_index [mask = 1] ⃝ 3 6 remote_index ← kv_index [mask = 2] ⃝ 7 return local_index, remote_index 5
on the GPU execution path. Specifically, we pre-allocate an int8 KV mask matrix for each attention layer and have prefill nodes update it via GDR. Each KV mask matrix is of shape B × N, where B is the batch size and N is the sequence length. 1 Each element m denotes the status of the jth token in the ij ith sequence. All heads of the same token share a mask value for their KV entries. A value of 0 indicates that the corresponding KV entries belong to a padded token in the batch. The padded tokens are used to align input sequences across the batch, and the KV mask value of 0 prevents their KV caches from being loaded onto the GPU. A value of 1 indicates that the KV entries are locally available, and 2 indicates that they reside on a remote prefill node. Therefore, each time a KV block is pushed to the decoding node, the prefill node also updates the corresponding values in the decodingside KV mask via an additional RDMA WRITE, as shown in Figure 6. The two WRITEs are combined into a single network round-trip using doorbell batching. This leverages the in-order delivery property of RDMA NICs [3, 78] to ensure that the KV mask is updated only after the corresponding KV block has been written. Based on this, we extend the on-GPU KV selection process with several operators to split the generated KV index into local and remote ones, as shown in Algorithm 1. The opera1 selects the status of the required KV entries from the tor ⃝ KV mask matrix, generating a small mask matrix of the same shape as the KV index, i.e., BH × k. The local and remote indexes are then generated by retrieving the KV index with the 2 and ⃝, 3 where 1 corresponding mask value, i.e., operators ⃝ indicates that the corresponding index refers to a local KV entry and 2 indicates a remote one. These three operators efficiently split the KV index. They involve only lightweight indexing operations, e.g., indexing into the KV mask of shape B × N with a small KV index of shape BH × k. They could be easily applied to existing KV selection processes [24, 37] without modifying their code. Reordering-based KV gathering. On the remote lane, we
Parallel On-demand Transfer
The limited number of transferable KV blocks still forces the decoding node to fetch missing entries on demand, placing network round-trips on the critical path. To address this, we introduce a two-lane KV fetching technique that removes such round-trips from the critical path. Our key idea is to parallelize the loading of local and remote KV caches. As introduced in Section 2.3, the KV index generated by the KV selection algorithm determines KV entries to load onto the GPU. We split the KV index into two separate indexes, i.e., one for local and the other for remote KV entries, so as to decouple the KV cache loading into two parallel data transfer lanes, as shown in Figure 7. With the split index, the decoding node selects and fetches KV entries from host memory using only the local index. At the same time, it issues remote procedure calls (RPCs) to prefill nodes to fetch the missing entries identified by the remote index. In particular, it SENDs the remote index from its GPU memory to the host memory of prefill nodes via GPU-direct RDMA (GDR). After RECVing the index, the prefill node selects the required KV entries locally and SENDs them back to the decoding node’s GPU memory. The decoding node starts the attention computation once it RECVs all remotely fetched entries and finishes loading the local ones. The main challenges are efficiently splitting the on-GPU KV index and automatically gathering selected KV entries from both lanes: 1) For index splitting, since KV caches are transferred to CPU memory in a fine-grained and noncontiguous manner, the GPU cannot efficiently determine whether each KV entry is local or not. 2) For KV gathering, the dynamic and discrete positions of missing KV entries make it difficult for the RDMA NIC to place remotely fetched entries directly into the correct locations in GPU memory. Mask-based index splitting. To address the first challenge, our key idea is to maintain mask matrices on the GPU to track the status of each KV entry. This enables the decoding node to observe real-time KV status in host memory directly
1 The memory overhead of KV mask matrices is negligible. When serving
64K tokens on Qwen3-14B, the overhead is only L · 64 KB = 2.5 MB.
6
KV Cache Pool
KV
① SEND
…
4.3
top-Kr no
RDMA
GPU
let prefill nodes directly send the selected KV entries to their target locations in the GPU memory of the decoding node. This could avoid the need for the decoding node to launch an additional kernel on the critical path to scatter the received entries into the locally fetched KV cache. However, since the target locations of remote KV entries are interleaved among local ones, discretely transferring them would incur excessive I/O overhead, which is impractical. To address the above challenge, we leverage the fact that the order of KV entries along the sequence dimension in the KV T √ )V . cache does not affect the attention result, i.e., so f tmax( QK d This is because the KV cache already contains the positional information. Based on this, we propose reordering KV entries to make the remote ones contiguous along the sequence dimension. For clarity, we define a row of KV entries as the k entries for a given head along the sequence dimension, i.e., the k in shape BH × k × d. Since the order of selected KV entries is determined by the KV index, we realize the reordering with two additional operators in Algorithm 1. The 4 sorts each row of the index mask in descending operator ⃝ 5 applies the ordering to the KV order, and the operator ⃝ index. As the highest mask value (i.e., 2) indicates a remote KV entry, the descending order makes the remote KV entries contiguous at the start of each row of the selected KV cache. Therefore, the decoding node can RECV these entries rowwise instead of individually via GDR. To further accelerate the multi-row transfer, we utilize the scatter capacity of the DMA engine in the RDMA NIC [18,48], which enables the decoding node to receive a contiguous chunk of KV entries and scatter them into discrete rows at low runtime costs. Although the scatter-gathering capacity of the RDMA NIC has an upper limit (e.g., 20), it is sufficient to significantly reduce the number of I/Os (e.g., by 20×).
KV
② WRITE
(H × Nb × d)×2
Prefill Nodes
Attention Select KV
FFN
Attention Select KV
…
CPU / NIC Idle Resources
Fetch Local KV Fetch Remote KV
Idle … Resources
Decoding Node
Figure 8: The speculative transfer utilizing idle resources.
in Section 4.2, which overlaps with the FFN computation. Once all selected KV entries are available, the next attention layer starts, and the process repeats. Since KV fetching can only begin after KV indexes are generated, the CPUs and RDMA NIC remain idle during the attention computation. Non-intrusive speculative KV transfer. Based on the above observation, we propose transferring remaining KV entries from prefill nodes to the decoding node, utilizing idle resources. At the start of each attention layer, the decoding node notifies all prefill nodes that the network is idle via RDMA SEND operations. After RECVing the notification, the prefill node starts transferring the remaining KV blocks that are not sent during the prefill stage. Similar to the profile-based proactive transfer, these KV blocks are sent via RDMA WRITEs in the order of their importance, as defined in Equation 1. The corresponding KV masks are updated via additional WRITEs as introduced in Section 4.1. To prevent speculative transfer from interfering with remote KV fetching, it is crucial to terminate it in time. Since notifying prefill nodes is not timely due to network roundtrip time, we explicitly limit the number of KV blocks to send each time. Specifically, each speculative transfer is restricted to send a fixed fraction of total KV blocks, defined as the speculative ratio. A lower speculative ratio mitigates interference, but requires more iterations to fully transfer the KV cache. We set the ratio to 10% to deliver all remaining KV entries to the decoding node within 10 decoding iterations, with almost no interference with on-demand KV fetching.
Speculative Transfer
Despite the efficiency of parallel on-demand transfer, it is still difficult to fully hide the remote lane’s network overhead. The overhead can accumulate across decoding iterations, increasing end-to-end latency. This section eliminates the overhead for later iterations by speculatively delivering all KV entries to the decoding node in the background. However, the background traffic may interfere with the foreground on-demand transfers since it can contend for the limited network and compute resources. Our key observation is that attention computation exposes idle network and CPU resources, which can be leveraged to perform speculative transfers without impacting foreground execution. The right part of Figure 8 shows the operation flow of the decoding node, where InfiniGen’s prefetch technique [37] is adopted. Specifically, in each attention layer, the attention is executed concurrently with the generation of KV indexes for the next layer, i.e., the KV selection process, using separate GPU streams. The generated KV indexes are used to fetch the selected local and remote KV entries, as introduced
4.4
Discussions
Integrating different KV selection algorithms. SmartGen can generalize to various KV selection algorithms and LLMs. Specifically, it is compatible with various dynamic KV selection algorithms [7, 24, 37, 38] through the unified KV index abstraction shown in Figure 1. Besides, its generality to different LLMs is inherited from the KV selection algorithm it adopts. These algorithms generally require metadata tensors, i.e., lightweight auxiliary tensors derived from model weights or KV caches, to approximate attention scores without accessing full KV entries [7, 24, 37]. These tensors are used in the decoding stage of SmartGen to guide KV selection. We recompute or pre-load weight-related metadata on the decoding node to save network bandwidth. The key-cache-related metadata is transferred during the prefill stage. For example, 7
Table 1: GPU instances on the cloud used in this paper.
using a partial ratio of 0.3 in InfiniGen [37] results in a 15% increase in KV cache transfer overhead. This overhead is jointly considered by the analytical model and optimized by the selective transfer design. Adapting to various workloads. In real-world LLM inference, request workloads are dynamic. Our offline profiling maintains stability to some extent, as positional similarity has been validated by prior static KV pruning methods [69, 70] and recent models [77]. To handle larger workload variations, SmartGen adopts: (1) Periodic profiling. The offline profiling can be performed periodically to adapt to workload changes by updating the importance matrix I shown in Figure 6. This process is transparent to the SmartGen design, as the matrix I can be updated using a standard readcopy-update scheme. (2) Grouped profiling. The calibration datasets can be partitioned into additional groups based on prompt length, generating multiple importance matrices accordingly. When serving online requests, SmartGen selects the importance matrix from the most similar group. Adapting to various network loads. In LLM inference systems, network loads typically vary dynamically with the number of user requests and the changes of P/D settings. SmartGen can adapt to dynamically changing network loads based on the analytical model used by existing systems [47, 68, 85]. Under high network loads, SmartGen reduces stage-transition stalls through selective KV cache transfer, while under low loads, it automatically falls back to full KV cache transfer according to Equation 2. Moreover, by employing more advanced analytical models, SmartGen could save precious network bandwidth for other important tasks, e.g., prefix cache transfer [47]. In contrast, existing quantization-based KV cache transfer schemes [39, 82] cannot adapt to dynamic network conditions, as they cannot adjust KV cache quantization precision in response to changing loads at runtime, nor easily compensate for accuracy loss via on-demand transfer. Supporting cloud instances lacking GDR. Some low-cost cloud servers do not support GDR due to virtualization constraints. In such environments, the parallel on-demand transfer design requires two modifications: (1) On-CPU index splitting. Since prefill nodes cannot update the KV mask matrix via GDR, the KV mask should be maintained in host memory. Consequently, the KV index is offloaded to host memory before being split. (2) Extra local KV loading. Since the selected remote KV entries cannot be directly transferred into the GPU, the decoding node should first receive them in host memory and then load them into GPU via an extra host-todevice copy. Given that PCIe bandwidth is much higher than the network bandwidth, these overheads are acceptable.
5 5.1
Name
GPUs
vCPU
DRAM
Network
ecs.gn8is-2x.8xlarge
2 L20 (2*48 GB)
32
256 GB
32 Gbps
ecs.gn8is.4xlarge
1 L20 (48 GB)
16
128 GB
25 Gbps
ecs.gn8is.2xlarge
1 L20 (48 GB)
8
64 GB
15 Gbps
experiments on 3 ecs.gn8is-2x.8xlarge prefill instances and one ecs.gn8is.4xlarge decoding instance. Each instance is equipped with an eRDMA interface [9]. Within each instance, the vCPUs, GPUs, and vNIC are interconnected via PCIe 4.0×16. For each prefill node, we launch two processes with each on one GPU, achieving a maximum DP degree of 6 in total. For the decoding node, we adopt one GPU with larger batch sizes to maximize the GPU utilization [31, 45, 85]. Models and workloads. We use Qwen3 [62], Meta’s Llama3.1 [61], Google’s Gemma-3 [60], and Microsoft’s Phi-4 [1] models for evaluation, which are representative LLM families widely used in academia and industry. As for workloads, we use the LongBench benchmark [4], which covers a wide range of long-context tasks. We select the alphabetically first dataset from each of the four tasks, i.e., MultiFieldQA (document query answering), GovReport [32] (summarization), SAMSum [23] (few-shot learning), and LCC [26] (code completion), to comprehensively assess SmartGen’s performance across diverse scenarios. The average prompt lengths for the four workloads are 7K, 10K, 9K, and 3K tokens, respectively, with up to 60K tokens batched on our testbed by default. Besides, we separate LongBench’s first subset (i.e., 2WikiMultihopQA [29]) as a dedicated calibration dataset in advance to ensure it is not used during evaluation. Comparisons. We compare the following four schemes: • SmartGen (Infini./HATA): We implement SmartGen by adopting InfiniGen [37] and HATA [24], respectively, to verify the generality of the SmartGen design with either training-free or training-based KV selection algorithms. • Full transfer: This is the typical KV cache transfer scheme that pushes all KV cache layer-by-layer during the prefill stage. To our knowledge, this is the state-ofthe-art scheme widely adopted in existing disaggregated LLM inference systems [12, 31, 45, 47, 58]. • Partial transfer: This is the vanilla selective KV cache transfer scheme that pushes the first sequential Kr KV blocks (i.e., in memory address order) during prefill, and fetches the missing entries on demand during decoding. • HACK [82]: This is the state-of-the-art quantizationbased scheme that addresses the network bottleneck in disaggregated inference for self-hosted LLMs on the cloud. It quantizes the KV cache to int2 and the query to int8 using homomorphic quantization, enabling it to store the 2-bit KV cache on the GPU without offloading. For fairness, all methods are implemented on top of the same LLM inference system [53]. FlashAttention [13] and FlashIn-
Evaluation Experimental Setup
Testbed. The Alibaba GPU instances [8] used in this paper are listed in Table 1. Unless otherwise stated, we conduct 8
SmartGen (Infini.)
SmartGen (HATA)
Acc:
Infini. 102% HATA 103% HACK 54%
3 0 6
Average CTL (s)
4
Acc:
0 6 4
4
2
4
8
16
32 1
HACK
8
16
32 1
Acc:
Acc:
Acc:
Acc:
4
8
16
32 1
Acc:
Infini. 115% HATA 120% HACK 31%
Acc:
Infini. 99% HATA 99% HACK 27%
Infini. 99% HATA 99% HACK 41%
Acc:
Infini. 100% HATA 98% HACK 39%
2
Infini. 96% HATA 98% HACK 55%
Infini. 100% HATA 101% HACK 13%
Infini. 98% HATA 101% HACK 62%
# Output Tokens
Acc:
Infini. 98% HATA 92% HACK 28%
Acc:
Acc:
4
Acc:
Infini. 99% HATA 96% HACK 46%
Infini. 98% HATA 100% HACK 65%
2
Acc:
Infini. 99% HATA 98% HACK 29%
Acc:
Acc:
1
Partial Transfer
Infini. 100% HATA 100% HACK 29%
Infini. 101% HATA 102% HACK 70%
2
Full Transfer
Acc:
Acc:
0 6
Phi-4-14B
Infini. 100% HATA 100% HACK 13%
Infini. 98% HATA 97% HACK 62%
2
Qwen3-14B
Infini. 101% HATA 107% HACK 31%
Infini. 99% HATA 97% HACK 20%
2
0
Acc:
Gemma-3-12B
Infini. 100% HATA 100% HACK 51%
2
4
8
16
32 1
Acc:
16
32
Infini. 101% HATA 103% HACK 77%
2
4
8
GovReport MultiFieldQA
Qwen3-8B
SAMSum
6
Llama-3.1-8B
LCC
9
Figure 9: The average request CTL over time during token generation. The bottom-right table shows the accuracy relative to the full-cache baseline. More accuracy results will be discussed in § 5.5.
fer [72] are adopted to improve inference performance. On the prefill node, we assume a prefix cache hit ratio of 75% on host memory to improve prefill performance [19,47,64,75,84]. The KV caches to be transferred are partitioned using the same block granularity. On the decoding node, we improve the computational efficiency of full transfer by applying local KV cache selection using InfiniGen [37]. Key metrics. We evaluate the average cumulative per-token latency (CTL) per request to assess SmartGen’s overall performance for users. We focus on addressing the stagetransition stall issue to achieve a seamless disaggregated LLM inference. Therefore, we also evaluate the time-to-secondtoken (TTST) to show the performance of the KV cache transfer, and the time-between-tokens (TBT) to show the performance overhead induced by SmartGen. To isolate the overhead, we exclude TTST during the TBT calculation. As for accuracy, we use LongBench’s accuracy-related metrics (%) to measure the impact of KV selection in SmartGen across different datasets. Parameters. We use the suggested configurations of InfiniGen and HATA, e.g., an alpha value of 5, a partial ratio of 0.3, a hash bit count of 256, and a maximum KV selection ratio of 20%. Unless otherwise specified, we set the number of KV blocks per layer (i.e., M) to 1K and adopt a DP degree of 6. All requests are configured with an output length of 64 tokens. As for SmartGen, we set the speculative ratio to 10%.
5.2
under the MultiFieldQA workload on the Qwen3-14B model. Similar trends are observed on other models and workloads. CTL results. The average request CTL over time reflects user-perceived performance, as shown in Figure 9. Full transfer exhibits an early CTL spike, suggesting users experience stalls (shown as shadow) when transferring the entire KV cache. SmartGen reduces TTST by up to 4.3× on GovReport compared with full transfer, as shown in Figure 10, indicating a seamless inference. Partial transfer shows a generally higher CTL due to the overhead of on-demand transfers. HACK alleviates the transfer bottleneck via aggressive KV cache quantization, achieving substantial efficiency gains, but at the cost of some accuracy, reaching up to 77% relative accuracy on LongBench, lower than InfiniGen and HATA. This is likely because LongBench requires precise capture of key information in long contexts and is thus more sensitive to quantization errors. Besides, although HACK avoids offloading, it requires unpacking the KV cache from a compact format to a computation-ready format at each step, resulting in a higher TBT than full transfer and even SmartGen in some cases. In contrast, SmartGen progressively reduces TBTs and keeps the CTL generally the best, thanks to optimized KV cache transfers at each stage. TTST results. The left half of Figure 11 shows the TTST results under the MultifieldQA workload on Qwen3-14B. The DP degree is set equal to the batch size to maintain stable TTFT. With a batch size of 6, SmartGen, HACK, and partial transfer achieve 3.7×, 3.7×, and 2.9× lower TTST, compared with full transfer. This is because they do not transfer the entire KV cache from prefill nodes to the decoding node, saving network bandwidth. Besides, SmartGen outperforms
Performance Comparison
Figure 9 shows the performance of all schemes across different models and workloads on the LongBench workloads. Figures 10-12 provide detailed analyses of TTST and TBT 9
3
0
5
10
15
CTL Breakdown (s)
20 l
rtia Figure 10: The CTL breakdown. Pa
+Parallel Transfer +Speculative Transfer
3 11
Full Transfer Partial Transfer +Profile-based Transfer
+Parallel Transfer +Speculative Transfer
3
6
TTFT TTST 2 8
TTFT TTST 2
4
1 5
1
2
0 2
0
MultiField. GovReport SAMSum
LCC
(a) L20 instances (25 Gbps virtualized network).
TTFT & TTST (s) TBT (x100 ms)
TBT (x100 ms)
8
Full Transfer Partial Transfer +Profile-based Transfer
MultiField. GovReport SAMSum
LCC
(b) V100S instances (25 Gbps physical network).
Figure 13: The factor analysis for techniques in SmartGen on various GPU instances.
Sequential
Random
StreamingLLM
Profile-based
Optimal
8
KV Reusing Spec. Ratio: 5% Spec. Ratio: 10% Spec. Ratio: 20%
6 4 2
8
16
Token ID
24
32
Figure 14: The effectiveness of speculative transfer.
ing instances with different network bandwidths, verifying SmartGen’s ability to adapt to varying network conditions. As the network bandwidth decreases from 32 Gbps to 15 Gbps, SmartGen’s TTST improvement over full transfer increases from 2.5× to 3.3×, and the TBT improvement over partial transfer increases from 1.4× to 1.6×. At 15 Gbps, the TTST of SmartGen exceeds that of the ideal case. This is because the bandwidth is saturated by the transfer of the first two layers of KV cache and the metadata tensors required by the KV selection algorithms. We leave the optimization of this overhead for future work.
100 Ratio 80 1.1 60 1.0 40 0.9 20 LM3-8B QW3-8B GM3-12B QW3-14B Phi-4-14B
On-demand Ratio (%)
1.2 TBT Speedup
SmartGen (infini.) SmartGen (HATA)
8 6 4 2 0
TTFT & TTST (s)
n Ge
Partial Transfer Full Transfer HACK
TBT (x100 ms)
art
Sm
TTST (s)
ll Fu K C HA . ial t r Pa
6 5 4 3 2 1
HACK Full Transfer Partial Transfer 4 2 TTFT TTST 3 2 1 1 0 0 1 2 3 4 5 6 1 2 3 4 5 6 2xlarge 4xlarge 2x.8xlarge (15 Gbps) (25 Gbps) (32 Gbps) Batch Size Batch Size Figure 11: Performances with various batch sizes. Figure 12: Performances with various bandwidths. SmartGen (infinite.) SmartGen (HATA)
TBT (x100 ms)
TTFT TTST TBT
Figure 15: The effectiveness of profile-based proactive transfer.
partial transfer by 1.2× due to the more efficient on-demand transfer when decoding the second token. As the batch size grows from 1 to 6, TTSTs of full transfer, partial transfer, HACK, and SmartGen increase by 10.4×, 4.0×, 1.7×, and 2.0×, respectively. The increase in TTST for full transfer is attributed to the more KV caches to transfer as the batch size grows. In contrast, increases observed in SmartGen, HACK, and partial transfer are mainly due to the higher memory access intensity on the decoding node. TBT results. The right half of Figure 11 shows the average TBTs under the MultifieldQA workload on Qwen3-14B. The TBT of full transfer represents the ideal case since all KV caches are locally available on the decoding node. With a batch size of 6, partial transfer exhibits a 1.5× higher TBT than SmartGen and full transfer, respectively, as it requires additional network round-trips on the critical path of every decoding iteration. HACK incurs a 1.4× higher TBT than both SmartGen and full transfer due to format conversion overhead. SmartGen achieves a TBT close to the ideal case thanks to its effective and efficient designs on proactive, ondemand, and speculative transfers. As the batch size grows from 1 to 6, TBTs of full transfer, partial transfer, HACK, and SmartGen increase by 3.0×, 3.6×, 1.7×, and 2.7×, respectively, due to increased memory access intensity. Performance with various network bandwidths. Figure 12 shows the performance of all schemes across decod-
5.3
Factor Analysis
Figure 13 presents the factor analysis for SmartGen on Qwen3-14B. To demonstrate SmartGen’s performance across different hardware platforms, this section additionally evaluates it on V100S physical servers, i.e., r7525 instances on CloudLab [17], using the same amount of GPU memory for prefilling. Each proposed technique is applied to the partial transfer one by one. We analyze results only on MultifieldQA due to space limits. Others exhibit similar trends. Partial transfer. Compared with full transfer, partial transfer reduces TTST at the cost of increased TBT. It achieves a 3.0× and 1.5× reduction in TTST but incurs a 1.5× and 1.6× increase in TBT on L20 and V100S instances, respectively. This trade-off arises since fewer KV entries are transferred during prefill, necessitating transfers of missing entries during the decoding stage. The following designs address the TBT penalty introduced by partial transfer. + Profile-based proactive transfer. The profile-based proactive transfer reduces TBT by 1.03× and 1.1× on L20 and V100S instances, respectively, by prioritizing the proactive transfer of important KV entries. The TBT improvements on L20 instances are lower than those on V100S instances. This is likely because elastic NICs on Alibaba Cloud offer a 10
6 24
48
72
96
# Batched Tokens (K)
6 24
48
72
96
# Batched Tokens (K)
(a) Sequence length.
3 2 1 0
0.5
1
1.5
# KV Blocks (K)
2
4
Full Transfer
6 5 4 0.5
1
1.5
# KV Blocks (K)
2
Partial Transfer
3 2 1 0
10 20
40
60
Selection Ratio (%)
(b) Number of KV blocks.
TBT (x100 ms)
4
HACK
TTST (s)
8
0
4
SmartGen (HATA)
TBT (x100 ms)
12
TTST (s)
SmartGen (Infini.)
TBT (x100 ms)
TTST (s)
5 4 3 2 1 0
15 13 11 9 7 5
10 20
40
60
Selection Ratio (%)
(c) KV selection ratio.
Figure 16: The sensitivity analysis for overall performance.
less stable bandwidth ceiling than physical NICs on CloudLab, leading to smaller performance gains. Figure 15 compares different KV selection strategies for prefill nodes across various models under the MultiFieldQA workload on L20 instances. The profile-based strategy achieves up to 1.3× and 1.2× speedups over the random and sequential strategies, respectively, by reducing the ondemand ratio by up to 51% and 38%. StreamingLLM [70], a static KV sparsity strategy, yields large gains on Gemma-312B due to its alignment with the model’s sliding-window attention components. In addition, the profile-based strategy achieves up to a 1.1× speedup over StreamingLLM and reduces the on-demand ratio by up to 18% by further identifying important KV entries in full attention. It also achieves performance close to the optimal case, validating the effectiveness of profiling positional similarity. + Parallel on-demand transfer. The parallel on-demand transfer reduces TBT by 1.1× and 1.2× on L20 and V100S instances, respectively. This is because it removes the network round-trip for remote KV fetching from the critical path of local KV cache loading during decoding. + Speculative Transfer. The speculative transfer further brings 1.3× and 1.2× reduction in TBT on L20 and V100S instances, respectively, by speculatively delivering all KV entries to the decoding node. Figure 14 shows the details on Qwen3-14B under the MultifieldQA workload. Unlike KV reusing, i.e., reusing previously fetched KV entries in host memory, speculative transfer proactively delivers all remaining KV blocks within a few iterations, after which the TBT aligns with the ideal case. With a speculative ratio of 20%, speculative transfer completes within 5 iterations. However, TBTs of the first few tokens increase by up to 1.4× since excessive speculative KV block transfers interfere with subsequent KV fetching. In contrast, a speculative ratio of 5% introduces no such interference but requires approximately 20 iterations to complete all transfers. SmartGen chooses an appropriate ratio of 10%, enabling transferring all remaining KV blocks as early as possible (i.e., 10 iterations) with minimal impact on TBTs.
5.4
SmartGen consistently performs best with various sequence lengths. As the number of batched tokens increases from 6K to 96K tokens, the TTST of full transfer grows more rapidly (3.1 s) compared with partial transfer, HACK, and SmartGen (1.4/1.0/1.0 s), respectively, since it must transfer the entire KV cache, whose size scales linearly with the sequence length. Meanwhile, the TBT of partial transfer and HACK increases more rapidly (0.8/0.9 s) than those of full transfer and SmartGen (0.5 s). This is because HACK needs to unpack all KV cache entries, and partial transfer fetches remote KV entries inefficiently, the overhead of which also increases with the sequence length. Impact of number of KV blocks. Figure 16b illustrates the impact of the number of KV blocks per attention layer (i.e., M) on system performance. The TTST of full transfer increases by 1.2× when M exceeds 1K, due to NIC processing overhead. The performances of other schemes remain stable across different M values. This is because important KV entries within each layer tend to be spatially clustered, as shown in Figure 5, reducing the need for fine-grained profiling. We set M to 1K to mitigate its impact on full transfer. Impact of KV selection ratio. As shown in Figure 16c, the TTST of full transfer remains consistently high across varying selection ratios, as it transfers all KV entries regardless of the KV selection algorithm, making the KV cache transfer overhead dominate the overall TTST. As the ratio increases from 10% to 60%, TTSTs of partial transfer and SmartGen increase by 2.0× and 1.8×, and their TBTs grow by 2.1× and 1.7×, respectively. This is because higher selection ratios result in more KV entries being selected, leading to increased TBT, which in turn constitutes a larger proportion of the TTST. SmartGen maintains superior performance across varying sparsity degrees with both InfiniGen and HATA, implying its adaptability to more KV selection algorithms with diverse selection intensities.
5.5
Accuracy
Figure 17 presents the accuracy of SmartGen and some baselines across various models on LongBench. The relative KV cache size indicates the ratio between the KV cache used in attention and that of the full-cache baseline. Overall accuracy. SmartGen consistently shows high accuracy across the models and tasks. Its accuracy closely matches the full-cache baseline. This is because SmartGen
Sensitivity
This section investigates how some parameters affect the performance on Qwen3-14B and MultiFieldQA. Impact of sequence length. Figure 16a shows that 11
60 45 30 15 0 48 36 24 12 0
60 45 30 15
60 45 30 15
60 45 30 15
0
0
0
0
60 45 30 15 0 60 45 30 15 0 48 36 24 12 0
HACK 60 45 30 15
60 45 30 15
60 45 30 15
60 45 30 15
0
0
0
0
96 72 48 24 0 60 45 30 15 0 8 6 4 2 0 96 72 48 24 0
Gemma-3-12B StreamingLLM 60 45 30 15
60 45 30 15
60 45 30 15
60 45 30 15
0
0
0
0
60 45 30 15 0 60 45 30 15 0 60 45 30 15 0 96 72 48 24 0
Relative KV Cache Size (%)
Qwen3-14B Proactive (Infini.) Proactive (HATA) 60 45 30 15
60 45 30 15
60 45 30 15
60 45 30 15
0
0
0
0
32 24 16 8 0 60 45 30 15 0 60 45 30 15 0 48 36 24 12 0
Phi-4-14B Full Cache (Opt.)
MultiFieldQA
60 45 30 15
Qwen3-8B
60 45 30 15
0
60 45 30 15
0
60 45 30 15
0
60 45 30 15
0
GovReport
SmartGen (Infini.) SmartGen (HATA)
48 36 24 12 0
SAMSum
60 45 30 15 0
Llama-3.1-8B
LCC
F1 LongBench Metrics (%) ROUGE-L ROUGE-L Edit Sim.
72 54 36 18 0
Figure 17: The accuracy analysis of LLMs on the LongBench benchmark [4].
directly adopts the state-of-the-art KV selection method, i.e., InfiniGen [37] and HATA [24], without modifying its algorithm design. In contrast, HACK shows lower accuracy with 2-bit quantization, demonstrating that dynamic KV selection is more effective than quantization in mitigating the KV cache transfer issue for challenging long-context understanding tasks where precision is critical. Profiling accuracy. To validate the feasibility of offline profiling, we also assess two proactive-only baselines, i.e., proactive (Infini./HATA), that simply evict KV cache entries unselected by the profiling. They generally achieve accuracy comparable to StreamingLLM [70]. This is because offline profiling can accurately identify universally important tokens like static KV pruning methods. However, its accuracy remains lower than that of SmartGen, underscoring the necessity of on-demand transfer. We also observe an interesting phenomenon: the proactive-only baselines achieve higher accuracy than StreamingLLM on the Gemma-3-12B model, which uses a mix of sliding and full attention layers. This is likely because Gemma-3-12B is more sensitive to important tokens in the full-attention layers, whose KV distributions do not align with StreamingLLM’s pattern.
P/D disaggregation has become a widely adopted architecture for deployed LLM inference systems [12, 15, 36, 47, 84]. While this design enhances system scalability, it poses challenges in inter-node KV cache transfer [45, 82, 85]. DistServe [85] introduces additional layer placement constraints to force KV cache transfer to occur only within a node, which limits the flexibility of resource disaggregation. Splitwise [45] and DéjàVu [58] propose to overlap the transfer with the prefill computation. This could not fundamentally address the transfer issue as analyzed in Section 3. Mooncake [47] relies on high RDMA bandwidth (e.g., 800 Gbps per machine) to enable efficient KV cache transfer across nodes. SmartGen focuses on optimizing KV cache transfer with limited network bandwidth, and could be incorporated into these systems.
6
6.3
6.1
prefill and decoding computations [31, 45, 47, 58, 85], and hardware-software co-design that constructs effective kernels and accelerators to boost inference efficiency [25, 28, 33]. SmartGen focuses on optimizing the KV cache transfer when self-hosting disaggregated LLM inference systems.
6.2
Related Work
P/D Disaggregation
KV Cache Management
There is a line of research that explores saving GPU memory footprint through KV cache management, e.g., virtualization [36, 46], offloading [19–21, 35, 44, 53], quantization [30,34,39,82], and sparsity [7,27,37,38,54,59,83]. Works that adopt sparsity relate most to SmartGen. SmartGen benefits from KV cache sparsity algorithms and leverages them to address the KV cache transfer challenge in self-hosted disaggregated LLM inference. Quantization-based methods like
LLM Inference Systems
The rapid advancement of LLMs has attracted increasing attention on enhancing LLM inference systems in many areas, e.g., request scheduling to satisfy SLO requirements [2, 16, 40, 42, 52, 55, 57, 74, 76, 81, 86], memory management to save GPU memory [36, 56, 65, 67, 71, 75, 79, 80, 84], resource disaggregation to attack the interference issue between the 12
HACK [82] are orthogonal to SmartGen and can be applied in a complementary manner. To our knowledge, SmartGen is the first work to enable seamless disaggregated LLM inference by exploiting the dynamic sparsity of KV caches.
7
Edwards, Yuri Burda, Nicholas Joseph, Greg Brockman, Alex Ray, Raul Puri, Gretchen Krueger, Michael Petrov, Heidy Khlaaf, Girish Sastry, Pamela Mishkin, Brooke Chan, Scott Gray, Nick Ryder, Mikhail Pavlov, Alethea Power, Lukasz Kaiser, Mohammad Bavarian, Clemens Winter, Philippe Tillet, Felipe Petroski Such, Dave Cummings, Matthias Plappert, Fotios Chantzis, Elizabeth Barnes, Ariel Herbert-Voss, William Hebgen Guss, Alex Nichol, Alex Paino, Nikolas Tezak, Jie Tang, Igor Babuschkin, Suchir Balaji, Shantanu Jain, William Saunders, Christopher Hesse, Andrew N. Carr, Jan Leike, Joshua Achiam, Vedant Misra, Evan Morikawa, Alec Radford, Matthew Knight, Miles Brundage, Mira Murati, Katie Mayer, Peter Welinder, Bob McGrew, Dario Amodei, Sam McCandlish, Ilya Sutskever, and Wojciech Zaremba. Evaluating large language models trained on code. CoRR, abs/2107.03374, 2021.
Conclusion
This paper identifies the stage-transition stall in disaggregated inference for self-hosted LLMs on the cloud. We propose a selective KV cache transfer scheme, SmartGen, that transfers essential KV cache entries across prefill and decoding stages to address this issue. Experimental results verify the efficacy and efficiency of SmartGen.
References [1] Marah I Abdin, Sahaj Agarwal, Ahmed Awadallah, Vidhisha Balachandran, Harkirat S. Behl, Lingjiao Chen, Gustavo de Rosa, Suriya Gunasekar, Mojan Javaheripi, Neel Joshi, Piero Kauffmann, Yash Lara, Caio César Teodoro Mendes, Arindam Mitra, Besmira Nushi, Dimitris Papailiopoulos, Olli Saarikivi, Shital Shah, Vaishnavi Shrivastava, Vibhav Vineet, Yue Wu, Safoora Yousefi, and Guoqing Zheng. Phi-4-reasoning technical report. CoRR, abs/2504.21318, 2025.
[7] Yaoqi Chen, Jinkai Zhang, Baotong Lu, Qianxi Zhang, Chengruidong Zhang, Jing Liu, Jingjia Luo, Di Liu, Huiqiang Jiang, Qi Chen, Bailu Ding, Xiao Yan, Jiawei Jiang, Chen Chen, Mingxing Zhang, Cheng Li, Yuqing Yang, Fan Yang, and Mao Yang. Retroinfer: A vector storage engine for scalable long-context LLM inference. Proc. VLDB Endow., 19(5):1016–1031, 2026.
[2] Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. Taming throughput-latency tradeoff in LLM inference with Sarathi-Serve. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 117–134. USENIX Association, 2024.
[8] Alibaba Cloud. Elastic GPU service instance families. https://www.alibabacloud.com/help/en/ecs/user-gui de/gpu-accelerated-compute-optimized-and-vgpu-a ccelerated-instance-families-1, Accessed: 2025. [9] Alibaba Cloud. eRDMA. https://www.alibabacloud .com/help/en/ecs/user-guide/elastic-rdma-erdma, Accessed: 2025.
[3] InfiniBand Trade Association. Infiniband architecture specification volume 1 release 1.8. https://www.infini bandta.org/ibta-specification, Accessed: 2025.
[10] Google Cloud. A2 ultra machine types. https://docs.c loud.google.com/compute/docs/accelerator-optimiz ed-machines#a2-ultra-vms, Accessed: 2025.
[4] Yushi Bai, Xin Lv, Jiajie Zhang, Hongchang Lyu, Jiankai Tang, Zhidian Huang, Zhengxiao Du, Xiao Liu, Aohan Zeng, Lei Hou, Yuxiao Dong, Jie Tang, and Juanzi Li. LongBench: A bilingual, multitask benchmark for long context understanding. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand, August 11-16, 2024, pages 3119–3137. Association for Computational Linguistics, 2024.
[11] Tencent Cloud. Computing instance. https://www.te ncentcloud.com/document/product/560/19701#GT4, Accessed: 2025. [12] NVIDIA Corporation. NVIDIA dynamo platform. https: //developer.nvidia.com/dynamo, Accessed: 2025. [13] Tri Dao, Daniel Y. Fu, Stefano Ermon, Atri Rudra, and Christopher Ré. FlashAttention: Fast and memoryefficient exact attention with io-awareness. In Advances in Neural Information Processing Systems 35: Annual Conference on Neural Information Processing Systems 2022, NeurIPS 2022, New Orleans, LA, USA, November 28 - December 9, 2022, 2022.
[5] Junyi Chen, Chuheng Du, Renyuan Liu, Shuochao Yao, Dingtian Yan, Jiang Liao, Shengzhong Liu, Fan Wu, and Guihai Chen. TokenFlow: Responsive LLM text streaming serving under request burst via preemptive scheduling. In Proceedings of the 21st European Conference on Computer Systems, EuroSys 2026, 2026. ACM, 2025.
[14] DeepSeek-AI. Deepseek-v3.2-exp: Boosting longcontext efficiency with deepseek sparse attention. https: //github.com/deepseek-ai/DeepSeek-V3.2-Exp/blob/ main/DeepSeek_V3_2.pdf, Accessed: 2025.
[6] Mark Chen, Jerry Tworek, Heewoo Jun, Qiming Yuan, Henrique Pondé de Oliveira Pinto, Jared Kaplan, Harri 13
[15] DeepSeek-AI, Aixin Liu, Bei Feng, Bing Xue, Bingxuan Wang, Bochao Wu, Chengda Lu, Chenggang Zhao, Chengqi Deng, Chenyu Zhang, Chong Ruan, Damai Dai, Daya Guo, Dejian Yang, Deli Chen, Dongjie Ji, Erhang Li, Fangyun Lin, Fucong Dai, Fuli Luo, Guangbo Hao, Guanting Chen, Guowei Li, H. Zhang, Han Bao, Hanwei Xu, Haocheng Wang, Haowei Zhang, Honghui Ding, Huajian Xin, Huazuo Gao, Hui Li, Hui Qu, J. L. Cai, Jian Liang, Jianzhong Guo, Jiaqi Ni, Jiashi Li, Jiawei Wang, Jin Chen, Jingchang Chen, Jingyang Yuan, Junjie Qiu, Junlong Li, Junxiao Song, Kai Dong, Kai Hu, Kaige Gao, Kang Guan, Kexin Huang, Kuai Yu, Lean Wang, Lecong Zhang, Lei Xu, Leyi Xia, Liang Zhao, Litong Wang, Liyue Zhang, Meng Li, Miaojun Wang, Mingchuan Zhang, Minghua Zhang, Minghui Tang, Mingming Li, Ning Tian, Panpan Huang, Peiyi Wang, Peng Zhang, Qiancheng Wang, Qihao Zhu, Qinyu Chen, Qiushi Du, R. J. Chen, R. L. Jin, Ruiqi Ge, Ruisong Zhang, Ruizhe Pan, Runji Wang, Runxin Xu, Ruoyu Zhang, Ruyi Chen, S. S. Li, Shanghao Lu, Shangyan Zhou, Shanhuang Chen, Shaoqing Wu, Shengfeng Ye, Shengfeng Ye, Shirong Ma, Shiyu Wang, Shuang Zhou, Shuiping Yu, Shunfeng Zhou, Shuting Pan, T. Wang, Tao Yun, Tian Pei, Tianyu Sun, W. L. Xiao, and Wangding Zeng. DeepSeek-v3 technical report. CoRR, abs/2412.19437, 2024.
and Pengfei Zuo. Cost-efficient large language model serving for multi-turn conversations with CachedAttention. In Proceedings of the 2024 USENIX Annual Technical Conference, USENIX ATC 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 111–126. USENIX Association, 2024. [20] Shiwei Gao, Youmin Chen, and Jiwu Shu. Fast state restoration in LLM serving with HCache. In Proceedings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 April 2025, pages 128–143. ACM, 2025. [21] Shiwei Gao, Qing Wang, Shaoxun Zeng, Youyou Lu, and Jiwu Shu. Weaver: Efficient multi-llm serving with attention offloading. In Proceedings of the 2025 USENIX Annual Technical Conference, USENIX ATC 2025, Boston, MA, USA, July 7-9, 2025, pages 587–595. USENIX Association, 2025. [22] Alireza Ghadimi and Hamid Beigy. Hybrid multidocument summarization using pre-trained language models. Expert Syst. Appl., 192:116292, 2022. [23] Bogdan Gliwa, Iwona Mochol, Maciej Biesek, and Aleksander Wawer. SAMSum corpus: A human-annotated dialogue dataset for abstractive summarization. CoRR, abs/1911.12237, 2019.
[16] Kuntai Du, Bowen Wang, Chen Zhang, Yiming Cheng, Qing Lan, Hejian Sang, Yihua Cheng, Jiayi Yao, Xiaoxuan Liu, Yifan Qiao, Ion Stoica, and Junchen Jiang. PrefillOnly: An inference engine for prefill-only workloads in large language model applications. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP 2025, Lotte Hotel World, Seoul, Republic of Korea, October 13-16, 2025, pages 399–414. ACM, 2025.
[24] Ping Gong, Jiawei Yi, Shengnan Wang, Juncheng Zhang, Zewen Jin, Ouxiang Zhou, Ruibo Liu, Guanbin Xu, Youhui Bai, Bowen Ye, Kun Yuan, Tong Yang, Gong Zhang, Renhai Chen, Feng Wu, and Cheng Li. HATA: trainable and hardware-efficient hash-aware top-k attention for scalable large model inference. In Findings of the Association for Computational Linguistics, ACL 2025, Vienna, Austria, July 27 - August 1, 2025, Findings of ACL, pages 24856–24871. Association for Computational Linguistics, 2025.
[17] Dmitry Duplyakin, Robert Ricci, Aleksander Maricq, Gary Wong, Jonathon Duerig, Eric Eide, Leigh Stoller, Mike Hibler, David Johnson, Kirk Webb, Aditya Akella, Kuang-Ching Wang, Glenn Ricart, Larry Landweber, Chip Elliott, Michael Zink, Emmanuel Cecchet, Snigdhaswin Kar, and Prabodh Mishra. The design and operation of CloudLab. In Proceedings of the 2019 USENIX Annual Technical Conference, USENIX ATC 2019, Renton, WA, USA, July 10-12, 2019, pages 1–14. USENIX Association, 2019.
[25] Cong Guo, Jiaming Tang, Weiming Hu, Jingwen Leng, Chen Zhang, Fan Yang, Yunxin Liu, Minyi Guo, and Yuhao Zhu. Olive: Accelerating large language models via hardware-friendly outlier-victim pair quantization. In Proceedings of the 50th Annual International Symposium on Computer Architecture, ISCA 2023, Orlando, FL, USA, June 17-21, 2023, pages 3:1–3:15. ACM, 2023.
[18] Ana Gainaru, Richard L. Graham, Artem Y. Polyakov, and Gilad Shainer. Using infiniband hardware gatherscatter capabilities to optimize MPI all-to-all. In Proceedings of the 23rd European MPI Users’ Group Meeting, EuroMPI 2016, Edinburgh, United Kingdom, September 25-28, 2016, pages 167–179. ACM, 2016.
[26] Daya Guo, Canwen Xu, Nan Duan, Jian Yin, and Julian J. McAuley. LongCoder: A long-range pre-trained language model for code completion. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 12098–12107. PMLR, 2023.
[19] Bin Gao, Zhuomin He, Puru Sharma, Qingxuan Kang, Djordje Jevdjic, Junbo Deng, Xingkun Yang, Zhou Yu, 14
[27] Jitai Hao, Yuke Zhu, Tian Wang, Jun Yu, Xin Xin, Bo Zheng, Zhaochun Ren, and Sheng Guo. OmniKV: Dynamic context selection for efficient long-context LLMs. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025.
[34] Minsu Kim, Seongmin Hong, Ryeowook Ko, Soongyu Choi, Hunjong Lee, Junsoo Kim, Joo-Young Kim, and Jongse Park. Oaken: Fast and efficient LLM serving with online-offline hybrid KV cache quantization. In Proceedings of the 52nd Annual International Symposium on Computer Architecture, ISCA 2025, Tokyo, Japan, June 21-25, 2025, pages 482–497. ACM, 2025.
[28] Congjie He, Yeqi Huang, Pei Mu, Ziming Miao, Jilong Xue, Lingxiao Ma, Fan Yang, and Luo Mai. WaferLLM: Large language model inference at wafer scale. In 19th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2025, Boston, MA, USA, July 7-9, 2025, pages 257–273. USENIX Association, 2025.
[35] Abhishek Vijaya Kumar, Gianni Antichi, and Rachee Singh. Aqua: Network-accelerated memory offloading for llms in scale-up GPU domains. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ASPLOS 2025, Rotterdam, Netherlands, 30 March 2025 - 3 April 2025, pages 48–62. ACM, 2025.
[29] Xanh Ho, Anh-Khoa Duong Nguyen, Saku Sugawara, and Akiko Aizawa. Constructing A multi-hop QA dataset for comprehensive evaluation of reasoning steps. In Proceedings of the 28th International Conference on Computational Linguistics, COLING 2020, Barcelona, Spain (Online), December 8-13, 2020, pages 6609–6625. International Committee on Computational Linguistics, 2020.
[36] Woosuk Kwon, Zhuohan Li, Siyuan Zhuang, Ying Sheng, Lianmin Zheng, Cody Hao Yu, Joseph Gonzalez, Hao Zhang, and Ion Stoica. Efficient memory management for large language model serving with PagedAttention. In Proceedings of the 29th Symposium on Operating Systems Principles, SOSP 2023, Koblenz, Germany, October 23-26, 2023, pages 611–626. ACM, 2023.
[30] Coleman Hooper, Sehoon Kim, Hiva Mohammadzadeh, Michael W. Mahoney, Yakun Sophia Shao, Kurt Keutzer, and Amir Gholami. KVQuant: Towards 10 million context length LLM inference with KV cache quantization. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024.
[37] Wonbeom Lee, Jungi Lee, Junghwan Seo, and Jaewoong Sim. InfiniGen: Efficient generative inference of large language models with dynamic KV cache management. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 155–172. USENIX Association, 2024.
[31] Cunchen Hu, Heyang Huang, Liangliang Xu, Xusheng Chen, Jiang Xu, Shuang Chen, Hao Feng, Chenxi Wang, Sa Wang, Yungang Bao, Ninghui Sun, and Yizhou Shan. Inference without interference: Disaggregate LLM inference for mixed downstream workloads. CoRR, abs/2401.11181, 2024.
[38] Guangda Liu, Chengwei Li, Jieru Zhao, Chenqi Zhang, and Minyi Guo. ClusterKV: Manipulating LLM KV cache in semantic space for recallable compression. pages 1–7, 2025.
[32] Luyang Huang, Shuyang Cao, Nikolaus Nova Parulian, Heng Ji, and Lu Wang. Efficient attentions for long document summarization. In Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Linguistics: Human Language Technologies, NAACL-HLT 2021, Online, June 6-11, 2021, pages 1419–1436. Association for Computational Linguistics, 2021.
[39] Yuhan Liu, Hanchen Li, Yihua Cheng, Siddhant Ray, Yuyang Huang, Qizheng Zhang, Kuntai Du, Jiayi Yao, Shan Lu, Ganesh Ananthanarayanan, Michael Maire, Henry Hoffmann, Ari Holtzman, and Junchen Jiang. Cachegen: KV cache compression and streaming for fast large language model serving. In Proceedings of the ACM SIGCOMM 2024 Conference, ACM SIGCOMM 2024, Sydney, NSW, Australia, August 4-8, 2024, pages 38–56. ACM, 2024.
[33] Aditya K. Kamath, Ramya Prabhu, Jayashree Mohan, Simon Peter, Ramachandran Ramjee, and Ashish Panwar. POD-Attention: Unlocking full prefill-decode overlap for faster LLM inference. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ASPLOS 2025, Rotterdam, Netherlands, 30 March 2025 - 3 April 2025, pages 897–912. ACM, 2025.
[40] Yixuan Mei, Yonghao Zhuang, Xupeng Miao, Juncheng Yang, Zhihao Jia, and Rashmi Vinayak. Helix: Serving large language models over heterogeneous GPUs and network via max-flow. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, ASPLOS 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 April 2025, pages 586–602. ACM, 2025. 15
[41] Microsoft. MInference: Million-tokens prompt inference for long-context llms. https://www.microsoft.co m/en-us/research/project/minference-million-token s-prompt-inference-for-long-context-llms, Accessed: 2025.
[49] Baptiste Rozière, Jonas Gehring, Fabian Gloeckle, Sten Sootla, Itai Gat, Xiaoqing Ellen Tan, Yossi Adi, Jingyu Liu, Tal Remez, Jérémy Rapin, Artyom Kozhevnikov, Ivan Evtimov, Joanna Bitton, Manish Bhatt, Cristian Canton-Ferrer, Aaron Grattafiori, Wenhan Xiong, Alexandre Défossez, Jade Copet, Faisal Azhar, Hugo Touvron, Louis Martin, Nicolas Usunier, Thomas Scialom, and Gabriel Synnaeve. Code Llama: Open foundation models for code. CoRR, abs/2308.12950, 2023.
[42] Deepak Narayanan, Keshav Santhanam, Fiodar Kazhamiaka, Amar Phanishayee, and Matei Zaharia. Heterogeneity-aware cluster scheduling policies for deep learning workloads. In 14th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2020, Virtual Event, November 4-6, 2020, pages 481–498. USENIX Association, 2020.
[50] Amazon Web Services. Partner success with AWS. http s://aws.amazon.com/partners/success, Accessed: 2025.
[43] OpenAI. GPT-5 is here. https://openai.com/gpt-5, Accessed: 2025.
[51] Amazon Web Services. Recommended GPU instances. https://docs.aws.amazon.com/dlami/latest/devguide/ gpu.html, Accessed: 2025.
[44] Xiurui Pan, Endian Li, Qiao Li, Shengwen Liang, Yizhou Shan, Ke Zhou, Yingwei Luo, Xiaolin Wang, and Jie Zhang. InstAttention: In-storage attention offloading for cost-effective long-context LLM inference. In IEEE International Symposium on High Performance Computer Architecture, HPCA 2025, Las Vegas, NV, USA, March 1-5, 2025, pages 1510–1525. IEEE, 2025.
[52] Ying Sheng, Shiyi Cao, Dacheng Li, Banghua Zhu, Zhuohan Li, Danyang Zhuo, Joseph E. Gonzalez, and Ion Stoica. Fairness in serving large language models. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 965–988. USENIX Association, 2024.
[45] Pratyush Patel, Esha Choukse, Chaojie Zhang, Aashaka Shah, Íñigo Goiri, Saeed Maleki, and Ricardo Bianchini. Splitwise: Efficient generative LLM inference using phase splitting. In 51st ACM/IEEE Annual International Symposium on Computer Architecture, ISCA 2024, Buenos Aires, Argentina, June 29 - July 3, 2024, pages 118–132. IEEE, 2024.
[53] Ying Sheng, Lianmin Zheng, Binhang Yuan, Zhuohan Li, Max Ryabinin, Beidi Chen, Percy Liang, Christopher Ré, Ion Stoica, and Ce Zhang. FlexGen: High-throughput generative inference of large language models with a single GPU. In International Conference on Machine Learning, ICML 2023, 23-29 July 2023, Honolulu, Hawaii, USA, volume 202 of Proceedings of Machine Learning Research, pages 31094–31116. PMLR, 2023.
[46] Ramya Prabhu, Ajay Nayak, Jayashree Mohan, Ramachandran Ramjee, and Ashish Panwar. vAttention: Dynamic memory management for serving llms without PagedAttention. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, ASPLOS 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 April 2025, pages 1133–1150. ACM, 2025.
[54] Xiangyu Shi, Marco Chiesa, Gerald Q. Maguire Jr., and Dejan Kostic. KVComm: Enabling efficient LLM communication through selective KV sharing. 2026. [55] Sudipta Saha Shubha, Haiying Shen, and Anand P. Iyer. USHER: holistic interference avoidance for resource optimized ML inference. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 947– 964. USENIX Association, 2024.
[47] Ruoyu Qin, Zheming Li, Weiran He, Jialei Cui, Feng Ren, Mingxing Zhang, Yongwei Wu, Weimin Zheng, and Xinran Xu. Mooncake: Trading more storage for less computation - A kvcache-centric architecture for serving LLM chatbot. In 23rd USENIX Conference on File and Storage Technologies, FAST 2025, Santa Clara, CA, February 25-27, 2025, pages 155–170. USENIX Association, 2025.
[56] Yixin Song, Zeyu Mi, Haotong Xie, and Haibo Chen. PowerInfer: Fast large language model serving with a consumer-grade GPU. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles, SOSP 2024, Austin, TX, USA, November 4-6, 2024, pages 590–606. ACM, 2024.
[48] Deepti Raghavan, Philip Alexander Levis, Matei Zaharia, and Irene Zhang. Breakfast of champions: towards zero-copy serialization with NIC scatter-gather. In HotOS ’21: Workshop on Hot Topics in Operating Systems, Ann Arbor, Michigan, USA, June, 1-3, 2021, pages 199–205. ACM, 2021.
[57] Vikranth Srivatsa, Zijian He, Reyna Abhyankar, Dongming Li, and Yiying Zhang. Preble: Efficient distributed prompt scheduling for LLM serving. In The Thirteenth International Conference on Learning Representations, 16
ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025.
[67] Xingda Wei, Zhuobin Huang, Tianle Sun, Yingyi Hao, Rong Chen, Mingcong Han, Jinyu Gu, and Haibo Chen. Phoenixos: Concurrent os-level GPU checkpoint and restore with validated speculation. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP 2025, Lotte Hotel World, Seoul, Republic of Korea, October 13-16, 2025, pages 996–1013. ACM, 2025.
[58] Foteini Strati, Sara McAllister, Amar Phanishayee, Jakub Tarnawski, and Ana Klimovic. Déjàvu: Kv-cache streaming for fast, fault-tolerant generative LLM serving. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024.
[68] Bingyang Wu, Shengyu Liu, Yinmin Zhong, Peng Sun, Xuanzhe Liu, and Xin Jin. LoongServe: Efficiently serving long-context large language models with elastic sequence parallelism. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles, SOSP 2024, Austin, TX, USA, November 4-6, 2024, pages 640– 654. ACM, 2024.
[59] Jiaming Tang, Yilong Zhao, Kan Zhu, Guangxuan Xiao, Baris Kasikci, and Song Han. QUEST: query-aware sparsity for efficient long-context LLM inference. In Forty-first International Conference on Machine Learning, ICML 2024, Vienna, Austria, July 21-27, 2024. OpenReview.net, 2024.
[69] Guangxuan Xiao, Jiaming Tang, Jingwei Zuo, Junxian Guo, Shang Yang, Haotian Tang, Yao Fu, and Song Han. DuoAttention: Efficient long-context LLM inference with retrieval and streaming heads. In The Thirteenth International Conference on Learning Representations, ICLR 2025, Singapore, April 24-28, 2025. OpenReview.net, 2025.
[60] Gemma Team. Gemma 3 technical report. CoRR, abs/2503.19786, 2025. [61] Llama Team. The llama 3 herd of models. CoRR, abs/2407.21783, 2024. [62] Qwen Team. Qwen3 technical report. abs/2505.09388, 2025.
CoRR,
[70] Guangxuan Xiao, Yuandong Tian, Beidi Chen, Song Han, and Mike Lewis. Efficient streaming language models with attention sinks. In The Twelfth International Conference on Learning Representations, ICLR 2024, Vienna, Austria, May 7-11, 2024. OpenReview.net, 2024.
[63] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N. Gomez, Lukasz Kaiser, and Illia Polosukhin. Attention is all you need. In Advances in Neural Information Processing Systems 30: Annual Conference on Neural Information Processing Systems 2017, December 4-9, 2017, Long Beach, CA, USA, pages 5998–6008, 2017.
[71] Jiayi Yao, Hanchen Li, Yuhan Liu, Siddhant Ray, Yihua Cheng, Qizheng Zhang, Kuntai Du, Shan Lu, and Junchen Jiang. CacheBlend: Fast large language model serving for RAG with cached knowledge fusion. In Proceedings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 April 2025, pages 94–109. ACM, 2025.
[64] Jiahao Wang, Jinbo Han, Xingda Wei, Sijie Shen, Dingyan Zhang, Chenguang Fang, Rong Chen, Wenyuan Yu, and Haibo Chen. KVCache cache in the wild: Characterizing and optimizing KVCache cache at a large cloud provider, 2025.
[72] Zihao Ye, Lequn Chen, Ruihang Lai, Wuwei Lin, Yineng Zhang, Stephanie Wang, Tianqi Chen, Baris Kasikci, Vinod Grover, Arvind Krishnamurthy, and Luis Ceze. FlashInfer: Efficient and customizable attention engine for LLM inference serving. 2025.
[65] Jiahao Wang, Weiyu Xie, Mingxing Zhang, Boxing Zhang, Jianwei Dong, Yuening Zhu, Chen Lin, Jinqi Tang, Yaochen Han, Zhiyuan Ai, Xianglin Chen, Yongwei Wu, and Congfeng Jiang. From prefix cache to fusion RAG cache: Accelerating LLM inference in retrieval-augmented generation. CoRR, abs/2601.12904, 2026.
[73] Chong You, Kan Wu, Zhipeng Jia, Lin Chen, Srinadh Bhojanapalli, Jiaxian Guo, Utku Evci, Jan Wassenberg, Praneeth Netrapalli, Jeremiah J. Willcock, Suvinay Subramanian, Felix Chern, Alek Andreev, Shreya Pathak, Felix X. Yu, Prateek Jain, David E. Culler, Henry M. Levy, and Sanjiv Kumar. Spark transformer: Reactivating sparsity in FFN and attention. CoRR, abs/2506.06644, 2025.
[66] Yiming Wang, Zhuosheng Zhang, and Rui Wang. Element-aware summarization with large language models: Expert-aligned evaluation and chain-ofthought method. In Proceedings of the 61st Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2023, Toronto, Canada, July 9-14, 2023, pages 8640–8665. Association for Computational Linguistics, 2023.
[74] Gyeong-In Yu, Joo Seong Jeong, Geon-Woo Kim, Soojeong Kim, and Byung-Gon Chun. Orca: A distributed 17
serving system for transformer-based generative models. In 16th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2022, Carlsbad, CA, USA, July 11-13, 2022, pages 521–538. USENIX Association, 2022.
caching. In 19th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2025, Boston, MA, USA, July 7-9, 2025, pages 275–293. USENIX Association, 2025. [82] Zeyu Zhang, Haiying Shen, Shay Vargaftik, Ran Ben Basat, Michael Mitzenmacher, and Minlan Yu. HACK: homomorphic acceleration via compression of the keyvalue cache for disaggregated LLM inference. In Proceedings of the ACM SIGCOMM 2025 Conference, SIGCOMM 2025, São Francisco Convent, Coimbra, Portugal, September 8-11, 2025, pages 1245–1247. ACM, 2025.
[75] Lingfan Yu, Jinkun Lin, and Jinyang Li. Stateful large language model serving with Pensieve. In Proceedings of the Twentieth European Conference on Computer Systems, EuroSys 2025, Rotterdam, The Netherlands, 30 March 2025 - 3 April 2025, pages 144–158. ACM, 2025. [76] Yifan Yu, Yu Gan, Nikhil Sarda, Lillian Tsai, Jiaming Shen, Yanqi Zhou, Arvind Krishnamurthy, Fan Lai, Hank Levy, and David E. Culler. IC-Cache: Efficient large language model serving via in-context caching. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP 2025, Lotte Hotel World, Seoul, Republic of Korea, October 13-16, 2025, pages 375– 398. ACM, 2025.
[83] Zhenyu Zhang, Ying Sheng, Tianyi Zhou, Tianlong Chen, Lianmin Zheng, Ruisi Cai, Zhao Song, Yuandong Tian, Christopher Ré, Clark W. Barrett, Zhangyang Wang, and Beidi Chen. H2O: heavy-hitter oracle for efficient generative inference of large language models. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, LA, USA, December 10 - 16, 2023, 2023.
[77] Jingyang Yuan, Huazuo Gao, Damai Dai, Junyu Luo, Liang Zhao, Zhengyan Zhang, Zhenda Xie, Yuxing Wei, Lean Wang, Zhiping Xiao, Yuqing Wang, Chong Ruan, Ming Zhang, Wenfeng Liang, and Wangding Zeng. Native sparse attention: Hardware-aligned and natively trainable sparse attention. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2025, Vienna, Austria, July 27 - August 1, 2025, pages 23078–23097. Association for Computational Linguistics, 2025.
[84] Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark W. Barrett, and Ying Sheng. SGLang: Efficient execution of structured language model programs. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, BC, Canada, December 10 - 15, 2024, 2024.
[78] Erfan Zamanian, Xiangyao Yu, Michael Stonebraker, and Tim Kraska. Rethinking database high availability with RDMA networks. Proc. VLDB Endow., 12(11):1637– 1650, 2019.
[85] Yinmin Zhong, Shengyu Liu, Junda Chen, Jianbo Hu, Yibo Zhu, Xuanzhe Liu, Xin Jin, and Hao Zhang. DistServe: Disaggregating prefill and decoding for goodputoptimized large language model serving. In 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, CA, USA, July 10-12, 2024, pages 193–210. USENIX Association, 2024.
[79] Shaoxun Zeng, Tingxu Ren, Jiwu Shu, and Youyou Lu. GPU checkpoint/restore made fast and lightweight. In 24th USENIX Conference on File and Storage Technologies, FAST 2026, Santa Clara, CA, USA, February 24-26, 2026, pages 239–254. USENIX Association, 2026.
[86] Kan Zhu, Yufei Gao, Yilong Zhao, Liangyu Zhao, Gefei Zuo, Yile Gu, Dedong Xie, Zihao Ye, Keisuke Kamahori, Chien-Yu Lin, Ziren Wang, Stephanie Wang, Arvind Krishnamurthy, and Baris Kasikci. NanoFlow: Towards optimal large language model serving throughput. In 19th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2025, Boston, MA, USA, July 7-9, 2025, pages 749–765. USENIX Association, 2025.
[80] Chen Zhang, Kuntai Du, Shu Liu, Woosuk Kwon, Xiangxi Mo, Yufeng Wang, Xiaoxuan Liu, Kaichao You, Zhuohan Li, Mingsheng Long, Jidong Zhai, Joseph Gonzalez, and Ion Stoica. Jenga: Effective memory management for serving LLM with heterogeneity. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP 2025, Lotte Hotel World, Seoul, Republic of Korea, October 13-16, 2025, pages 446–461. ACM, 2025. [81] Dingyan Zhang, Haotian Wang, Yang Liu, Xingda Wei, Yizhou Shan, Rong Chen, and Haibo Chen. BlitzScale: Fast and live large model autoscaling with O(1) host 18