Conceptio › Archive › arXiv CS
arXiv CSopen access

Unified AI Gateway: A Framework for Joint Model Routing and KV Cache Management

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

1

Unified AI Gateway: A Framework for Joint Model Routing and KV Cache Management

arXiv:2609.06940v1 [cs.DC] 7 Sep 2026

Jiaxun Lu, Xiang Zhang, Yunfeng Shao Abstract—Large language model (LLM) inference increasingly spans models that differ in size, capability, price, and provider. This shift creates two costs for developers. One is the integration cost of choosing among and switching between many models. The other is the inference cost of rebuilding a KV cache when it is unavailable or incompatible with the selected model. We define and analyze the Unified AI Gateway as a system setting for an edge-deployed AI traffic hub. It coordinates model routing, KV cache management, and compute placement across end devices, edge resources, and cloud model services. At request time, the gateway jointly selects a target model, an execution site, and a KV cache action under task-quality, latency, cost, and resource constraints. In parallel, background cache-management actions optimize KV cache placement, replication, retrieval, and lifecycle decisions for subsequent requests. We synthesize existing evidence on KV cache reuse, compression, cross-model mapping, distributed storage, and transfer, and discuss the remaining challenges of integrating these capabilities into one system. Across eight typical workload profiles, our workload-level analytical simulation reports TTFT speedups of 1.25×–13.28× and input-cost benefits of 1.20×–6.16×. Index Terms—AI Gateway, KV Cache Management, Edge Computing, Model Routing, Transmission Optimization, LLM Inference

✦

1

I NTRODUCTION

A

S the model ecosystem expands, developers increasingly choose among models that differ in size, capability, price, and provider. Consider a long-context coding agent. Its planning, coding, execution, and review steps may favor different models, while conversation history, tool outputs, and repository state recur across turns [1]. Existing AI gateways provide unified interfaces and routing functions such as load balancing and failover, but they do not coordinate the KV cache with provider and model selection. When routing switches the target model, the accumulated cache may be unavailable or incompatible, forcing the target model to repeat full-context prefill. This repeated work can outweigh the capability or price advantage of switching models. Model routing and KV cache management must therefore be coordinated at the gateway. Current support remains fragmented across gateway services, inference engines, and cache systems. Unified APIs reduce provider-integration overhead, while runtimes such as vLLM and SGLang and systems such as LMCache and Mooncake support runtime reuse, cache lifecycle management, cross-engine sharing, or distributed cache storage [2], [3], [4], [5], [6]. However, these capabilities are not jointly coordinated with provider and model selection across devices and services. A routing decision can therefore discard useful state, increase time to first token (TTFT) and input cost, and create cache-induced model lock-in. To address this issue, this paper proposes the Unified AI Gateway, which couples provider and model routing with KV cache handling across participating runtimes. Figure 1 previews the resulting benefits. Across the typical workloads and model-

Corresponding author: Yunfeng Shao. Jiaxun Lu, Xiang Zhang and Yunfeng Shao are with Huawei, Shenzhen, China (e-mails: [email protected]; [email protected]; [email protected]).

Fig. 1: Across eight workload profiles, the Unified AI Gateway yields modeled TTFT speedups of 1.25×–13.28× and input-cost benefits of 1.20×–6.16×. This End–Edge–Cloud setting coordinates model routing, compute placement, and KV cache handling across end devices, edge resources, and cloud model services, so that available caches can be reused, mapped, or transferred instead of fully recomputed. Marker shapes denote model-switching probabilities of 25%, 50%, and 75%, defined as the probability that a request with an available KV cache is routed to a different model. Points farther toward the upper right indicate larger benefits.

switching rates considered in this study, our workload-level simulation reports modeled TTFT speedups of 1.25×–13.28× and input-cost benefits of 1.20×–6.16×. To realize these benefits, the gateway coordinates routing, KV cache storage, prefill, and cache transformation along the user-to-service path. Its placement at the edge supports authorized cache sharing across devices and users,

2

while edge compute can perform local prefill or map KV caches between models subject to fidelity constraints [7], [8], [9], [10]. This coordination raises three system challenges involving KV cache compatibility and fidelity, the choice between transfer and recomputation, and cache placement, replication, and isolation under capacity and privacy constraints. The Unified AI Gateway therefore considers one joint routing control layer with four KV cache management capability groups, namely cross-granularity reuse, adaptive compression, cross-model mapping, and distributed storage and transfer. Figure 2 contrasts this coordination with the request-routing scope of a conventional AI gateway. The four capability groups are grounded in active research lines, but their support remains fragmented across inference engines, cache systems, and model-serving services. Mainstream inference engines support exact prefix caching, while recent studies have extended KV-cache reuse to semantic, non-prefix, workflow-aware, and agentoriented settings [8], [11], [12], [13], [14], [15], [16], [17], [18]. Distributed systems have also advanced KV-cache storage and transfer [5], [19]. The unresolved gateway-level problem is to jointly decide whether to reuse, map, transfer, or recompute a cache while selecting the provider, model, and placement under quality, latency, cost, and resource constraints. To address these issues, this paper introduces the Unified AI Gateway, which jointly coordinates model routing and KV-cache management across the End–Edge–Cloud continuum. We evaluate its potential benefits through analytical simulation. The contributions are as follows. ●

●

●

Unified gateway abstraction and joint control. We introduce the Unified AI Gateway as an End–Edge– Cloud abstraction that coordinates model routing, KV-cache management, and compute placement. We formalize it through a Global Control Plane that jointly selects models and cache actions under latency, cost, quality, and resource constraints, providing a basis for gateway-wide optimization. Coordination taxonomy and systems gap. We organize cache reuse, compression, mapping, storage, transfer, prefill, and routing into an end-to-end taxonomy. Our analysis shows that existing methods optimize individual stages but do not coordinate cache handling, model selection, and compute placement across models and system tiers. Quantified benefits and deployment guidance. Across eight workload profiles, analytical simulation shows modeled TTFT speedups from 1.25× to 13.28× and input-cost benefits from 1.20× to 6.16×. In particular, the long-horizon coding-agent workload achieves TTFT speedups from 5.01× to 12.87×. Sensitivity analysis shows that future deployments with higher cache hit rates and greater available bandwidth could further broaden the benefit range.

The remainder of this paper is organized as follows. Sections 2–5 define the system setting, formulate the control plane, analyze component requirements, and synthesize existing evidence, respectively. Section 6 evaluates workloadlevel benefits through analytical simulation, while Section 7

discusses deployment readiness and future directions. Section 8 concludes the paper.

2

P RELIMINARIES AND S YSTEM S ETTING

This section introduces the preliminaries needed to treat KV caches as gateway-managed resources, defines the Unified AI Gateway setting, and presents the End–Edge–Cloud architecture that instantiates it. The corresponding joint model-routing and KV cache decision space is presented in Section 3. 2.1

Preliminaries

KV cache. The KV cache stores the key and value tensors of processed tokens so that autoregressive decoding does not recompute the full context at every step [20]. Its size grows with context length and depends on the model’s layers, KV heads, head dimension, and numerical format. This growing footprint makes efficient storage management essential. Techniques such as grouped-query attention (GQA), multihead latent attention (MLA), and hybrid attention change the KV cache memory footprint but not the underlying management problem [21], [22]. PagedAttention organizes the cache into fixed-size blocks to improve memory utilization [12]. These properties make the KV cache an explicit infrastructure object whose storage, reuse, placement, and transfer affect inference latency and resource use. KV cache reuse. KV cache reuse shares previously computed attention caches across requests to avoid repeated prefill. Exact prefix caching is the most mature form and is widely supported by inference engines [11], [12]. Semantic, non-prefix, workflow-level, and cross-model reuse broaden the coverage to similar prompts, interleaved contexts, multiagent workflows, and model switching [8], [9], [10], [15], [19]. Greater coverage introduces matching, calibration, correction, and fidelity costs. Accordingly, the gateway treats reuse granularity as a KV cache decision rather than as a fixed cache policy. The detailed evidence is reviewed in Section 5.1. Prefill–decode disaggregation. Prefill-decode (PD) disaggregation places the compute-intensive prefill phase and the memory-intensive decode phase on different nodes, enabling independent resource management [23], [24]. It makes the prefill-generated KV cache an explicit network payload that must be stored, transferred, and consumed by a decode node. Systems such as Mooncake and NVIDIA Dynamo demonstrate the infrastructure relevance of this design, while HACK highlights the transfer bottleneck created by long contexts [6], [25], [26]. End–Edge–Cloud deployment extends this workflow to edge prefill and terminal or cloud decoding, which is the setting analyzed by the unified gateway. 2.2

Unified AI Gateway Definition

We define the Unified AI Gateway as an edge-oriented system setting in which request routing and KV cache management are coordinated on the same request path. The gateway connects end devices, edge resources, distributed cache storage, prefill services, and a pool of model-serving clusters. It manages model-specific KV caches together with

3

Fig. 2: Comparison of a Conventional AI Gateway (left) and the Unified AI Gateway (right). The conventional gateway routes requests without coordinating shared KV caches, whereas the Unified AI Gateway coordinates cache reuse and cross-model mapping to avoid target-side full re-prefill when a valid mapping is available. metadata describing model version, token span, cache layout, location, ownership, freshness, and provenance. For each request, the gateway may select a model, an execution site, and one KV cache action, including direct reuse, cross-model mapping, compression, transfer, edge or remote prefill, and target-side re-prefill. The setting therefore covers both the architecture in which the edge manages KV caches and the deployment constraints that determine whether a cache action is feasible. The End–Edge–Cloud architecture is presented below, while the corresponding decision space is formalized in Section 3. 2.3

End–Edge–Cloud Architecture

The Unified AI Gateway is realized as an End–Edge–Cloud three-layer collaborative architecture (Figure 3). In this paper, the Unified AI Gateway denotes the complete system setting. The Global Control Plane is its control core, while end devices and cloud model services are participating endpoints. The edge layer hosts the Global Control Plane and distributed edge devices. Its core design is to place KV cache management on the request data path. The control plane decides not only which model to route to, but also how the associated cache is stored, reused, transformed, and transferred. The end side consists of terminal devices (web applications, mobile apps, Internet of Things (IoT) devices, etc.) that originate inference requests and consume generated results. Their submitted requests contain a user identifier and a context sequence (prompt, dialogue history, etc.). Under the End–Edge–Cloud PD-disaggregated form, terminals may also perform decode computation, in which case the edge gateway delivers the prefill-generated KV cache to them (Section 3.3). When terminal-side inference is selected, the gateway can forward a model request and the required KV cache to the terminal-side service, then receive the generated result through the same edge path.

The edge side hosts the global control plane and the concrete edge devices that execute selected cache actions. The global control plane provides joint routing and cacheaction decisions, cache lookup, and KV cache management, while the edge devices provide inference, prefill, and crossmodel mapping. These functions share request and cache metadata. Routing can therefore account for cache availability and preparation cost, while placement can reflect routing priorities, locality, capacity, and access policy. Here, KV cache management refers to control-plane coordination of cache lookup, placement, transfer, and lifecycle actions. The underlying storage and migration mechanisms can be provided by distributed cache systems such as Mooncake and LMCache [5], [6], while the Unified AI Gateway determines when and where these capabilities should be used. The detailed cache-handling workflow is described in Section 3, and its technical requirements are analyzed in Section 4. Edge devices can host prefill computation (Section 3.3), allowing the edge to participate directly in cache generation. With suitable model weights and runtime support, they can additionally host model inference services and serve requests locally. The control plane stages and transfers the resulting cache to the selected cloud-side or terminal-side decode service when the serving computation is placed remotely. The cloud side comprises the inference service clusters of model providers (including commercial APIs and selfhosted model services), receiving requests routed by the edge gateway and returning generated results. The gateway can also instruct the cloud side to perform prefill directly and return the KV cache to the edge. Depending on the selected placement, the cloud, edge, or end side can provide the serving or decode service. The workflow in Figure 3 follows the request path from the end to the edge and then to the selected serving destination. The gateway authenticates the request, extracts its context and service requirements, and uses cache metadata,

4

Fig. 3: High-level overview of the End–Edge–Cloud architecture that constitutes the Unified AI Gateway. The global control plane coordinates model routing, cache lookup, and KV cache management across end devices, distributed edge devices, and cloud services, while edge devices provide inference, prefill, and cross-model mapping. Requests, results, and KV caches can flow bidirectionally among the participating layers. Detailed cache actions are shown in Figure 4.

resource status, and network conditions to filter feasible actions and estimate their end-to-end cost. A valid KV cache is reused directly or mapped to the target model. Otherwise, the gateway performs edge prefill or schedules prefill remotely, then compresses and delivers the resulting cache to the selected decode node, which may be hosted in the edge, cloud, or end side. The gateway uses cache availability, resource status, network conditions, authorization, and task-level quality feedback to update future routing and placement policies.

3 G LOBAL C ONTROL P LANE AND KV-C ACHE W ORKFLOW This section presents the Global Control Plane and its joint model-routing, compute-placement, and KV cache decision space. It then traces how the control plane coordinates cache preparation and delivery across the gateway components. 3.1 Joint Routing, Compute Placement, and KV Cache Management The Global Control Plane makes coupled decisions at two timescales. At request time, it chooses the serving model, its execution site, and how the required KV cache should be obtained. The selected model determines the target cache format, while the execution site determines where serving or decode computation runs. Cache availability, mapping cost, prefill resources, and network conditions affect the cost of this model-site pair. Across the request stream, KV cache management continuously optimizes the shared cache pool across edge devices. It determines where caches are stored, whether frequently requested caches should be replicated near likely request sources or decode destinations, and when caches should be migrated or evicted under storage,

bandwidth, freshness, authorization, and consistency constraints. These background actions shape the caches available to later requests and can reduce their preparation cost. To make this joint decision, the control plane must first determine which model, execution-site, and KV cache paths are feasible and then compare their end-to-end value. For each candidate model and site, Figure 4 presents five cache paths covering exact cache hits, cross-model mapping, remote cache transfer, edge prefill, and target-side re-prefill. These paths are evaluated along four dimensions shown in the Unified Cache Selector. (i) Availability asks whether the required cache exists, can be looked up, or can be generated under the current placement and authorization policy. (ii) Fidelity covers cross-model cache compatibility, mapping or compression error, and adherence to the quality and servicelevel objective (SLO) constraints. (iii) Cost includes cache lookup, transmission delay and round-trip time (RTT), compression and decompression, mapping, prefill or recomputation, memory copies, and cache restoration. (iv) Resources include link bandwidth, end, edge, and cloud compute, memory and storage capacity, and queueing conditions. Cache lookup and management affect every candidate path, while request-stream decisions determine the feasible action set. Let the model set be M = {M1 , . . . , Mk }, where k is the number of candidate models, and let request rt = (ut , ct ) arrive at time t. Here, ut denotes the user or tenant identity and ct denotes the request context. Let Dt be the candidate execution sites across the end, edge, and cloud, and let Dt (j ) ⊆ Dt contain the sites that can execute model Mj or its decode phase. Prefill of context ct on model Mi produces KV cache Ki (ct ), whose size grows approximately linearly with context length [12]. For a coding-agent request, for example, ct can include the accumulated conversation, tool outputs, and repository state that recur across turns. A model switch then requires the control plane to select an execution site and decide whether the source cache can be reused directly, mapped to the target model, transferred, or replaced by target-side re-prefill. At the request-stream level, let E denote the set of edge devices and let Pt (K ) ⊆ E denote the devices that store or stage cache object K at time t. KV cache management updates Pt (K ) according to request frequency, locality, capacity, freshness, authorization, and transfer cost, thereby allowing frequently requested caches to be replicated near likely request sources or decode destinations. In the perrequest formulation below, Pt is the current cache placement produced by these background management actions. It affects the feasible actions and their costs but is not itself selected by the instantaneous joint decision. For a target model Mj and execution site e ∈ Dt (j ), let At (j, e; Pt ) be the feasible KV cache actions given the current cache placement, cache contents, compute resources at e, network conditions, authorization policy, and requirements for keeping users’ data separate. Thus, e captures request-time compute placement, while Pt captures the cache placement available when the request arrives. For cross-model mapping, the transformed cache is written as K̂j (ct ) = fi→j (Ki (ct )). Define the feasible tuple set as Ft (Pt ) = {(j, e, a) ∶ Mj ∈ M, e ∈ Dt (j ), a ∈ At (j, e; Pt )}. For each (j, e, a) ∈ Ft (Pt ), let Ltjea denote the cache-preparation and delivery latency,

5

Fig. 4: Global control-plane workflow for joint model routing, compute placement, and KV cache action selection. The control plane compares cache hit, cross-model mapping, edge prefill, remote cache transfer, and target-side re-prefill according to end-to-end cost, fidelity, and resource constraints.

Ctjea the corresponding compute and transfer cost, and qj (rt ) the standalone quality of model Mj . If dtjea denotes the task-level quality loss relative to target-side re-prefill on the selected model, the resulting quality is Qtjea = qj (rt ) − dtjea . The latency Ltjea includes cache lookup, generation or full-context recomputation, mapping, compression and decompression, transfer to e, memory copies, restoration, and queueing whenever these operations are used. The cost Ctjea accounts for the associated compute, memory, storage, and network resources. When KV cache transfer is used, a simplified latency model is Ttransfer ≈

compressed Stjea

B

+ TRTT + Tqueue + Tcopy + Trestore , (1)

compressed where Stjea is the transmitted KV cache after compression, B is the available bandwidth, TRTT is round-trip

communication delay, Tqueue is queueing delay, Tcopy is memory-copy overhead, and Trestore is decode-side cache Pt →e restoration and initialization overhead. We use Tcache (ct ) for the complete delivery path from the current cache placement to execution site e. Cross-model transfer also incurs i→j mapping latency Tmap (ct ), which is zero for direct samemodel transfer. The control plane then selects a model, an execution site, and a KV cache action jointly, shown as follows (jt∗ , e∗t , a∗t ) = arg

min

̃tjea + λC C ̃tjea λL L

s.t. Ltjea ≤ Lt ,

Qtjea ≥ Qt .

(j,e,a)∈Ft (Pt )

(2)

Here, λL , λC ≥ 0 are user-configurable weights for latency ̃tjea and C ̃tjea are normalized latency and cost and cost, L values, and Lt and Qt are the request-level latency and quality requirements. Using reference ranges [Lmin , Lmax ]

̃ = (L − Lmin )/(Lmax − Lmin ) and [Cmin , Cmax ], we define L ̃ = (C − Cmin )/(Cmax − Cmin ). These reference ranges and C can be estimated from feasible tuples observed for the current request class and planning horizon, then updated as workload and link conditions change. At the request-stream level, edge storage capacity, link bandwidth, cache consistency, and privacy constraints further restrict the feasible action set and its cost. The joint objective in Eq. (2) captures the central Global Control Plane problem. Model routing, compute placement, and KV cache preparation must be optimized together. 3.2

KV Cache Reuse and Cross-Model Mapping

KV cache reuse avoids repeated prefill by sharing computed attention caches across requests. It covers exact prefix reuse, semantic or non-prefix reuse, and workflow-level reuse across multi-turn, multi-agent, or authorized crossuser requests. When a request switches from source model Mi to target model Mj , the gateway may transform the source cache, K̂j (c) = fi→j (Ki (c)), and allow the target model to decode without recomputing the full context. Cross-model mapping broadens model-switching flexibility but introduces transformation cost and a fidelity constraint. 3.3

Edge-Side Prefill

Edge devices may execute prefill within the PDdisaggregated workflow, generating a KV cache near the terminal and delivering it to an edge, cloud, or terminalside decode service. The gateway uses local execution when the required model and resources are available. Otherwise, it schedules prefill remotely and forwards the returned cache. Direct use requires compatible model and runtime semantics, whereas incompatible caches require mapping or

6

target-side re-prefill. A prefill-generated cache can enter the distributed cache pool for later reuse, mapping, compression, or transfer. 3.4

KV Cache Transfer, Compression, and Retrieval

Transfer, compression, and distributed retrieval move reusable KV caches to the selected decode destination. The gateway estimates delivery latency using Eq. (1) and compares it with target-side recomputation. Compression reduces transfer volume, while retrieval and placement determine whether the cache is available within latency, fidelity, and authorization constraints. The cache pool may span edge devices, cloud services, and authorized tenant groups, with placement and replication guided by locality, demand, capacity, freshness, and transfer cost. Transfer scheduling must also account for memory layout and inference-engine overhead.

4

C OMPONENT R ESPONSIBILITIES AND R EQUIRE -

MENTS

This section maps the request workflow to gateway-specific responsibilities and technical requirements. It covers joint routing, KV cache reuse and cross-model mapping, compression and transfer, distributed cache management, and edge prefill, subject to security requirements that keep users’ data separate. It then examines the coupling among routing, transfer scheduling, cache placement, and inference-engine memory management. 4.1

Joint Model and KV-Cache Routing

The routing control component is responsible for producing the joint model, execution-site, and KV cache decision formalized in Section 3.1. Its output includes the target model, the site that executes serving or decode, and the cache action used to prepare that model’s context. The policy must combine task quality, provider cost, latency objectives, access policy, cache locality, model compatibility, and the fidelity constraint of mapped or compressed caches. It must also react to online feedback. A cache lookup can fail, a mapper may violate the quality threshold, or a remote cache transfer may become slower than target-side re-prefill. Existing model-routing studies are reviewed in Section 5.5. The gateway-specific requirement is to make routing and compute placement cache-aware rather than treating KV cache preparation as a downstream implementation detail. 4.2

Cross-Granularity KV Cache Reuse

The reuse manager determines whether an existing KV cache can replace full prefill across multi-model, multisession, and multi-agent requests. Its gateway-specific requirements are (i) scaled exact prefix matching for highconcurrency lookup [11], (ii) semantic-similarity matching across lexically different prompts [8], [14], (iii) non-prefix segment reuse with positional and cross-segment interaction handling [15], [16], (iv) multi-agent adaptation for cache offsets and memory-position drift [17], [18], and (v) dynamic heat management for prioritization and eviction. Hybridattention models additionally require layer-aware cache indexing [22].

4.3

KV Cache Compression and Transfer

The compression and transfer manager makes reusable KV caches deliverable under heterogeneous edge bandwidth and storage constraints [27]. Given cache structure, bandwidth, RTT, and storage headroom, it selects among low-bit quantization [28], [29], layer-aware mixed compression [30], and low-rank methods [31]. The gateway balances transfer latency, storage occupancy, fidelity, and computation overhead, while online controllers adapt the decision as link and workload conditions change [32]. 4.4

Cross-Model KV Cache Mapping

The cross-model cache translator preserves usable context when routing changes between models. Because models differ in layers, attention heads, hidden dimensions, and positional-encoding schemes, their KV caches are not directly interchangeable [10]. The translator therefore needs mapping mechanisms that preserve context continuity while satisfying latency and fidelity requirements. The main requirements are (i) same-architecture mapping for models with identical architecture but different parameters [9], (ii) cross-architecture mapping for heterogeneous model pairs [10], [33], and (iii) zero-shot or few-shot mapping because training a dedicated mapper for every model pair is impractical. 4.5

Distributed KV Cache Management

The distributed cache store makes KV caches discoverable, authorized, and close to the selected decode destination across regions and operators. Its main requirements are (i) placement based on user geography, access patterns, and model locations, (ii) efficient retrieval, (iii) transfer scheduling under bandwidth constraints [5], [26], and (iv) consistency maintenance during model updates and cache invalidation. 4.6

Control-Plane Coordination

The gateway control plane is responsible for coordinating decisions that appear modular but share the same bottlenecks. The four component responsibilities above all depend on efficient KV cache transfer. Transfer is itself coupled with the inference engine’s memory management. Inference engines such as PagedAttention manage KV caches in fixed token blocks whose physical locations are often non-contiguous. Direct transfers over remote direct memory access (RDMA) or sockets therefore require many scattergather direct memory access (DMA) operations and can reduce effective Peripheral Component Interconnect Express (PCIe) bandwidth. Compacting the blocks first introduces extra copies and consumes memory bandwidth that is already scarce during decoding [12]. Shared wide-area links create another coupling. Large cache objects can occupy the link for a long time and block smaller requests, worsening tail latency. KVServe jointly schedules compression strategies and bandwidth conditions, illustrating the need for system-level coupling [32]. The gateway therefore needs page-aware, copy-efficient transfer protocols aligned with the engine’s page-table

7

structure [12], together with scheduling across transfer, storage, and inference. These decisions must account for crossdomain bandwidth and RTT, because long-context requests can generate several gigabytes of cache data.

5

E VIDENCE S YNTHESIS AND S YSTEM P ROGRESS

Section 3 identifies the gateway responsibilities and their cross-cutting coupling constraints. The following subsections organize the existing KV cache evidence by capability and summarize performance, cost, flexibility, and systemcoordination benefits. Section 6 then evaluates how these benefits vary across workload profiles and operating conditions. 5.1

KV Cache Reuse

KV cache reuse balances coverage against matching and correction cost. Existing work mainly covers exact prefixes, arbitrary segments, semantic similarity, and multi-agent contexts. Table 1 organizes these reuse levels before the detailed discussion. Because the underlying studies use different models, hardware platforms, workloads, and evaluation metrics, Table 1 presents a taxonomy of reported outcomes rather than a direct cross-row comparison. The rows move from exact matching to increasingly flexible reuse. Potential coverage and reuse benefit generally increase, but end-to-end gains are not monotonic because matching, calibration, and correction costs also increase. Exact prefix reuse requires the reused content to appear at the prompt prefix and to match exactly. vLLM combined paged attention with block hashing and improved throughput by 2–4× at the same latency [12]. SGLang used a hierarchical radix tree and reported up to 6.4× higher throughput, with production hit rates of 52.4% on LLaVA-Next-34B and 74.1% on Vicuna-33B [11]. KVFlow added workflow-aware eviction and reported up to 1.83× and 2.19× speedups for single and concurrent workflows [13]. These results establish exact prefix reuse as a practical baseline, but its coverage remains limited. Non-prefix reuse extends the reuse unit to contiguous segments at arbitrary positions. CacheBlend selectively recomputed boundary tokens and reduced TTFT by 2.2–3.3× while improving throughput by 2.8–5× with negligible quality loss [15]. SparseX used sparse indexes to identify tokens that needed correction and supported interleaved reuse across dialogue, retrieval, and multi-agent workloads [16]. RedKnot grouped caches by attention-head class and preserved the native layout of hybrid-attention models such as GDN. It reported 1.6–3.5× TTFT speedups and 4.7–7.8× higher concurrent-session throughput [17]. Semantic-similarity reuse links lexically different prompts that have similar meanings. SemShareKV used fuzzy embedding matching and positional correction, achieving a 6.25× speedup and 42% memory savings on 5K-token summarization inputs [8]. Sim-LLM reported up to 39.40% higher throughput and 34.65% lower memory footprint on A40 and A100 GPUs [14]. In multi-agent scenarios, shared contexts often appear at different offsets. KVCOMM calibrated these offsets, reached

reuse rates above 70%, and reported up to 7.8× prefill speedup in five-agent workloads [19]. Agent Primitives transferred reusable review and planning contexts through KV caches, improving accuracy by 12.0–16.5% and reducing token use and latency by about 3–4× relative to textbased coordination [34]. RelayCaching reused the cache from one agent’s decode phase in the next agent’s prefill phase, achieving over 80% reuse and up to 4.7× lower TTFT [35]. Taken together, these studies report up to 7.8× prefill speedup and over 80% reuse, but broader coverage requires additional matching, calibration, and correction. 5.2

KV Cache Compression

KV cache compression balances size reduction, fidelity, and compression cost. Existing methods mainly use quantization, low-rank decomposition, or a combination of both. Quantization methods reduce bfloat16 (BF16) KV caches to lower bit widths. RotateKV reached 2-bit precision with a 3.97× peak-memory reduction and 2.32× decode speedup [28]. MPOQ reduced KV memory by about 75% at 4-bit precision with nearly lossless quality [29]. KVQuant supported up to 1M-token contexts on one A100-80GB and 10M-token contexts on eight GPUs [36]. ZipCache combined salient-token selection with mixed precision and reported 4.98× compression with a 0.38% accuracy drop [37]. TurboQuant was quality-neutral at 3.5 bits per channel and had only slight degradation at 2.5 bits per channel [38]. Low-rank and mixed methods exploit redundancy across hidden dimensions and layers. STAR-KV combined adaptive low-rank decomposition with quantization and reported up to 20× overall compression [31]. TailorKV selected layer-sensitive quantization and offloading policies, enabling a 128K context on one RTX 3090 at 82 ms/token with near-lossless quality [30]. Kelle co-designed KV caching with eDRAM and reported 3.9× speedup and 4.5× higher energy efficiency on edge devices [39]. These results show that 4–20× size reduction is feasible, but the chosen ratio must still respect fidelity and computation constraints. 5.3

Cross-Model KV Cache Mapping

Cross-model mapping transforms a source model’s KV cache into a form usable by a target model. The main difficulty is the mismatch in layers, hidden dimensions, and attention structures. Existing work covers same-architecture reuse and cross-architecture mapping. Same-architecture reuse targets models with the same architecture but different parameters. DroidSpeak selectively recomputed important layers and reused the rest [9]. LRAgent separated base and adapter components to share caches across low-rank adaptation (LoRA) variants [40]. Activated LoRA extended this direction to serving-engine support for KV cache reuse across adapter switches [41]. ICaRus enabled cache sharing across models through a logical encoder and decoder decomposition [42]. Cross-architecture mapping learns a transformation between incompatible KV cache spaces. C2C used neural projection and fusion to transfer a cache without per-token text communication and improved average accuracy by 8.5– 10.5% on policy-semantics tasks [10]. MoT used a mixture of translators for different architectures [33]. For models in the

8

TABLE 1: Taxonomy of KV cache reuse granularities and representative evidence from heterogeneous workloads. Reuse level

Reusable object

Coverage and flexibility

Additional processing

Representative benefit

Exact prefix

Identical prompt prefix

Hash lookup and cache eviction

Non-prefix segment

Contiguous segment at an arbitrary position, including shifted agent contexts

Lowest coverage and highest matching precision Broader coverage across interleaved and multi-agent contexts

Semantic similarity

Lexically different but similar content

Higher flexibility across prompts and tasks

Embedding matching and positional calibration

2–6.4× throughput improvement; 52.4–74.1% hit rate [11], [12] 2.2–3.3× lower TTFT; 2.8–5× higher throughput; over 70% reuse and up to 7.8× prefill speedup in multiagent workloads [15], [16], [19] 6.25× speedup; 42% memory savings [8]

same family, closed-form ridge regression offered a trainingfree alternative with much lower runtime than re-prefill [43]. These methods operate inside serving engines or target specific model and adapter families. In contrast, the Unified AI Gateway treats reuse and mapping as candidate KV cache actions in a gateway-level decision space that also includes model selection, compute placement, cache placement, transfer, edge prefill, and target-side re-prefill. For the gateway, mapping is useful when cache delivery and transformation together cost less than target-side recomputation while meeting the quality target. If Ki (c) is the KV cache produced by source model Mi for context c, switching to target model Mj can use mapping when Pt →e i→j j Tcache (c) + Tmap (c) < Trecompute (c), e,M

εi→j ≤ δ,

(3)

where εi→j is the mapping error and δ is the allowed quality loss. Depending on the task, εi→j can be measured as an accuracy drop, F1 loss, perplexity increase, or task-specific reward degradation relative to target-side re-prefill. The mapping cost grows with the transformed cache, whereas recomputation depends on context length, model size, and the attention implementation. Mapping is therefore often more attractive for long contexts. If the mapping criterion fails, the gateway uses direct transfer only when the source and target cache formats are compatible. Otherwise, it falls back to target-side re-prefill, as summarized in Table 2 and Eq. (3). These results show that cross-model mapping can run 2.7–25× faster than re-prefill and provide up to 2.65× TTFT speedup when the mapping satisfies the quality limit. The attainable gain still depends on model compatibility and mapping fidelity. 5.4

Distributed KV Cache Storage and Transfer

The storage layer must provide low-latency access under limited edge capacity. LMCache separated KV caches from GPU memory and supported sharing, lookup, eviction, migration, and compression across engines. Combined with vLLM, it reported up to 15× higher throughput and at least 2× lower latency on multi-turn workloads [5]. ObjectCache delivered data in GPU consumption order and used object storage as an elastic tier. A 64K-token context added only 5.6% TTFT over local DRAM on a 100 Gbps RoCE cluster,

Boundary recomputation, sparse correction, and offset calibration

while its bandwidth-aware scheduler further reduced TTFT by 1.2–1.8× [45]. Predictive Multi-Tier Memory Management expanded effective cache capacity from 40 GB to over 38 TB and reported 70–84% hit rates in replay experiments [46]. These results support multi-tier storage as a practical way to trade capacity against access latency. Transfer optimization addresses bandwidth and RTT limits on wide-area links. Preble used shared prefix caching and distributed prompt scheduling to reduce average latency by 1.5–14.5× and p99 latency by 2–10× [47]. SmartGen reported that transferring a 48K-token cache over 25 Gbps could take 6.5× the prefill time and 42.2% of job completion time, while its selective push and pull policy reduced second-token latency by up to 4.3× [48]. SplitZip provided GPU-friendly lossless compression and improved end-toend transfer by up to 1.32× [49]. KVServe selected compression online according to bandwidth and reported up to 9.13× lower job-completion time [32]. Prefill-as-a-Service added a schedulable prefill tier across data centers and reported 54% higher throughput and 64% lower P90 TTFT in a heterogeneous deployment [50]. FlowKV reduced average KV transfer latency by 96% and accelerated inference by 15.2%–48.9% through optimized transfer, load-aware scheduling, and flexible PD-node allocation [51]. KVDirect reduced per-request latency by 55% through tensor-centric communication, pull-based KV transfer, and dynamic GPU scheduling [52]. Together, these storage and transfer systems report up to 15× higher throughput, at least 2× lower latency, and up to 96% lower transfer latency. The realized gain depends on bandwidth, cache locality, and scheduling. 5.5

Model Routing

Model routing assigns requests to models under quality, latency, and cost constraints. OmniRouter reduced compute cost by at least 10.15% and improved accuracy by up to 6.30% through constrained optimization [53]. PORT approached the offline optimum without training and reported 4.25× higher throughput [54]. RouteLLM reduced invocation cost by more than 50% while preserving quality, and FrugalGPT reported savings of up to 98% through model cascades [55], [56]. Recent studies formulated KV cache constraints in online scheduling and jointly analyzed cache eviction and query routing [57], [58]. Robust KV

9

TABLE 2: Conditional comparison of cross-model KV cache mapping and target-side re-prefill. Dimension

KV cache mapping

Re-prefill

Compute object

Source KV cache (projection, reconstruction, or sparse patching) Grows with the volume of the transformed cache and the mapping operator Decode starts once cache transformation completes; TTFT can be lower when mapping overhead is below re-prefill cost Preserves source-context information subject to source/target alignment error Closed-form linear mapping 2.7–25× faster than re-prefill [43]; selective recomputation avoids full recomputation [9]; low-rank reconstruction and sparse patching achieve up to 2.65× TTFT speedup with quality loss no more than 5% [44]

Complete context (per-layer attention and feedforward recomputation) Grows with context length and model size, with the exact scaling determined by the attention implementation Must wait for the complete forward pass. TTFT grows with context Recomputes the context in the target model’s own KV cache space and avoids cross-model mapping error Full recomputation is the main cost source of model switching (Section 4.4)

Compute cost structure End-to-end latency Context preservation Measured reference

Cache Management further optimized request routing, prefix caching, and GPU resource configuration under outputlength uncertainty [59]. Prior work also studied historyaware agent routing [60] and prefix-affinity routing in CacheRoute [61]. These approaches coordinate resources inside datacenter serving clusters. They report cost savings above 50% and throughput gains up to 4.25×, but do not measure the additional benefit of cross-model mapping and cross-tier placement at the gateway level.

6

W ORKLOAD -L EVEL B ENEFIT A NALYSIS

Section 5 has established the component-level progress behind the Unified AI Gateway. This section examines how these KV cache capabilities combine under typical workload conditions. 6.1

Workload Profiles and Settings

We use an explicit workload-level KV cache preparation latency model to quantify the benefit of coordinated cache handling. Each workload is represented by θw = (Lin , Lout , h, rswitch , B ), where Lin and Lout are the input and output lengths, h is the effective KV cache hit rate, rswitch is the model-switching rate, and B is the available bandwidth. We instantiate eight typical workload profiles. The five Preble profiles use its reported context, output, and shared-prefix characteristics [47]. Multi-turn chat and retrieval-augmented generation (RAG) use representative hit rates for persistent dialogue and non-prefix document reuse [15], [62]. The long-horizon coding-agent profile follows a vLLM-Mooncake trace with an approximately 80Ktoken turn-30 context, a 131:1 input-to-output ratio, and a 94.2% hit rate. Its output length is rounded to 610 tokens [1]. We evaluate switching rates of 25%, 50%, and 75%. The 50% setting approximates the 45% per-turn switching rate reported for the single-turn routing baseline in vLLM’s longhorizon agent evaluation, while 25% and 75% provide lower and stress-test settings [63]. Here, the rate is the probability that a request with an available KV cache is routed to a different model. The reference bandwidth is 8 GB/s for the cross-domain setting. Table 3 lists the remaining calibration settings, including the fixed 15 KB/token effective footprint and 4× compression ratio. These values isolate the effects of workload shape, cache availability, switching rate, and bandwidth. The footprint calibration reflects modern hybrid-attention models [22], [50], [64].

TABLE 3: Reference settings for the workload-level analytical simulation. Parameter

Setting

KV cache footprint sKV

15 KB/token

Compression ratio ρ Prefill coefficients aprefill / battn / τprefill Decode coefficient tdecode Routing troute / lookup tlookup budget Model-switching rswitch

rate

Cached / uncached input price η = chit /cmiss Output / uncached input price λ = cout /cmiss Available bandwidth B

6.2

Basis or role

Representative effective footprint for modern hybridattention models [22], [50], [64] 4× Conservative reference near reported 4–5× reductions [28], [37] 0.04 ms/token Rounded representative values / 3 × based on the A100 profiles re10−6 ms/token2 leased with DistServe [24], [65] / 20 ms 10 ms/token Reference setting for the secondary end-to-end diagnostic 2 / 3 ms Control-plane calibration. Routing scale informed by prior work [66] 25% / 50% / Sensitivity settings for the prob75% ability of switching away from the cache-producing model 0.10 Representative cached-input discount [67] 5.0 Representative output-to-input pricing relationship [67] 8 GB/s Maximum available bandwidth in the reference cross-domain setting [45], [48]

KV-Cache Model and Metrics

The scalar model below instantiates the latency component of the general per-request quantities in Section 3.1 for a reference model pool and workload. In the general formulation, Ctjea denotes gateway-side compute, memory, storage, and network cost, whereas this workloadlevel evaluation focuses on user-visible latency and uses T for latency-equivalent preparation times. Here h, B , and the stated feasibility conditions summarize the current KV cache environment, and h is the effective cache-hit rate after background placement and replication. We assume that background placement, replication, eviction, and migration have already prepared this cache state, so their costs are amortized over the request stream and omitted from the per-request cache-preparation critical path. The results therefore condition on effective state management rather than evaluating its policy or overhead. For a workload row, let L = Lin . The cache-preparation

10

action times are given in Eq. (4), shown as follows

Tprefill (L) = aprefill L + battn L2 + τprefill , S (L) = LsKV , S (L) Ttransfer = , ρB Tre−prefill = Tprefill (L),

(4)

where aprefill , battn , and τprefill are the linear projection coefficient, quadratic full-attention coefficient, and fixed prefill overhead listed in Table 3. The linear-plus-quadratic form follows DistServe. We use rounded representative coefficients based on the A100 profiles released with its artifact for workload-level comparison [24], [65]. The KV cache size S (L) remains linear in context length. Here, ρ = 4 is the compression ratio. Recent measurements show closed-form cross-model mappers running 2.7–25× faster than re-prefill, while multitier cache systems demonstrate overlap between I/O and cache processing [43], [68]. Motivated by these results, the reference model pipelines mapping with chunked cache transfer and places compression, decompression, and memory-copy operations outside the critical path. RTT is not modeled separately because it is common to the network path. Under this overlap assumption, the numerical model sets Tmap+transfer = Ttransfer . A same-model hit reuses the cache directly, a switching hit uses mapping and transfer, and a miss triggers target-side full prefill. Thus, the switching path compares Tmap+transfer with Tre−prefill , while the input-token cost effect is calculated separately below. Therefore, the Unified AI Gateway preparation time is UG Tprep = troute + tlookup + hrswitch Tswitch + (1 − h)Tre−prefill , Tswitch = min{Tmap+transfer , Tre−prefill },

(5)

base Tprep = troute + tlookup + [(1 − h) + hrswitch ]Tre−prefill .

Every request incurs one cache-lookup time, regardless of whether the lookup succeeds. A successful lookup avoids the subsequent cache-preparation time. Equation (5) gives the resulting Unified AI Gateway and baseline preparation times. The baseline represents a conventional AI gateway that selects a target model from a model pool. A compatible same-model KV cache may be reused within the serving runtime, while a cache hit followed by a model switch and a cache miss both require target-side full prefill. Define base base UG UG Te2e = Tprep + Lout tdecode and Te2e = Tprep + Lout tdecode . The primary latency metric is the preparation speedup, which serves as a TTFT proxy in this model. We retain endto-end speedup as a secondary diagnostic for the effect of output generation. The two speedup metrics are

Se2e (θw ) =

base Te2e , UG Te2e

Sprep (θw ) =

base Tprep . UG Tprep

(6)

The end-to-end ratio is retained as a whole-request diagnostic. The preparation speedup in Eq. (6) is reported as the primary TTFT result in Table 4 and Figure 5. We calculate the input-token cost ratio separately from latency. Let cmiss and chit denote the prices of uncached

and cached input tokens, respectively, and let η = chit /cmiss . Current commercial model services provide a representative reference. OpenAI’s current token-based rate card and Anthropic’s cache-read rate set cached input at 0.10 times the corresponding uncached input price for listed models, equivalent to a 10× price reduction [67], [69]. We therefore use η = 0.10. Anthropic’s current API schedule lists output tokens at five times the uncached input-token price for Sonnet 4.6, so we use λ = cout /cmiss = 5.0 for the end-toend token-cost calculation [67]. Under the reference setting, every cache hit, including a hit followed by model switching, is assumed to be mapped successfully. Mapping is treated as a lightweight or linear transformation whose computation is hidden by overlap [43], [68]. Same-model hits and mapped switching hits therefore use the cached-input price, while cache misses use the uncached price. The normalized costs and their ratio are base Cinput = h(1 − rswitch )η + hrswitch + (1 − h), cmiss UG Cinput = hη + (1 − h), cmiss base Cinput h(1 − rswitch )η + hrswitch + (1 − h) . Ginput (θw ) = UG = hη + (1 − h) Cinput (7) This calculation complements the resource-cost term Ctjea in the Global Control Plane formulation with an inputbilling metric. It preserves cached-input billing for cache hits, including successful cross-model mapping, while ordinary misses and output-token charges remain unchanged. base To include output-token charges, let C̄base = Cinput /cmiss UG and C̄UG = Cinput /cmiss . The end-to-end token-cost ratio is

Lin C̄base + λLout . (8) Lin C̄UG + λLout The input-cost and end-to-end (E2E) cost columns in Table 4 are computed with Eqs. (7) and (8), respectively. Ge2e token (θw ) =

6.3

Simulation Results and Sensitivity

Table 4 compares the modeled TTFT, input-cost, and E2E cost benefits of the eight workload profiles at modelswitching probabilities of 25%, 50%, and 75%. To isolate the benefit of coordinated cache handling, we hold the conventional AI gateway’s model-selection policy fixed and treat its model-switching probability as the exogenous parameter rswitch . We then apply the Unified AI Gateway cache actions to the same workload profiles. This comparison measures cache-handling gains at the incumbent switching rate rather than evaluating a learned or optimized joint routing policy. Because prepared or mapped caches reduce the cost of changing models, a joint policy could create more switching opportunities. The 50% and 75% settings quantify the additional benefit available as the effective switching rate increases under the stated lightweight-mapping and overlap assumptions. The results depend on workload shape and switching rate. Across the 25%, 50%, and 75% switching settings, the eight profiles achieve TTFT speedups of 1.25×–5.39×, 1.49×– 9.48×, and 1.74×–13.28×, respectively. The corresponding

11

TABLE 4: Workload-level settings and simulated benefits for typical Unified AI Gateway workloads. Context and output lengths are mean values for the Preble workloads, while the multi-turn chat and RAG settings represent persistent dialogue and non-prefix document reuse reported in prior work. Reported shared-prefix proportions serve as effective cache-hit rates for the Preble profiles. The three benefit columns list results for rswitch = 25%, 50%, and 75%, in that order. The input-cost ratio uses a 10× cached-input discount, while the E2E cost ratio also includes output-token charges with λ = 5.0. Workload

KV cache reuse Context/Output pattern Lin /Lout

Multi-turn chat [62] RAG [15] ToolBench [47], [70]

16K/200 16K/300 1.8K/43

Embodied agent [47], [71] Programming [47], [72]

2.3K/16 3.9K/190

Video QA [47]

9.9K/4

Long-document QA (LooGLE) [47], [73]

23.5K/16

Long-horizon coding agent [1]

80K/610

Historical dialogue Non-prefix chunks Shared tool instructions Shared context and pauses Shared prompts and context Repeated video context Shared long-document context Persistent cross-turn prefix

TTFT speedup 25% / 50% / 75%

Input-cost ratio 25% / 50% / 75%

E2E cost ratio 25% / 50% / 75%

0.70 0.50 0.85

1.57× / 2.14× / 2.71× 1.25× / 1.49× / 1.74× 2.05× / 3.09× / 4.11×

1.43× / 1.85× / 2.28× 1.20× / 1.41× / 1.61× 1.81× / 2.63× / 3.44×

1.36× / 1.73× / 2.09× 1.17× / 1.35× / 1.52× 1.54× / 2.09× / 2.63×

0.97

4.37× / 7.55× / 10.56×

2.72× / 4.44× / 6.16×

2.35× / 3.69× / 5.04×

0.97

5.39× / 9.48× / 13.28×

2.72× / 4.44× / 6.16×

1.59× / 2.17× / 2.76×

0.88

2.70× / 4.36× / 5.99×

1.95× / 2.90× / 3.86×

1.94× / 2.89× / 3.83×

0.91

3.44× / 5.83× / 8.17×

2.13× / 3.26× / 4.39×

2.11× / 3.22× / 4.33×

0.942

5.01× / 8.97× / 12.87×

2.39× / 3.79× / 5.18×

2.11× / 3.23× / 4.34×

Cache hit rate h

input-cost ratios are 1.20×–2.72×, 1.41×–4.44×, and 1.61×– 6.16×, while the E2E cost ratios are 1.17×–2.35×, 1.35×–3.69×, and 1.52×–5.04×. For the long-horizon coding-agent profile, TTFT speedup grows from 5.01× to 12.87×, the input-cost ratio from 2.39× to 5.18×, and the E2E cost ratio from 2.11× to 4.34× as the switching rate increases from 25% to 75%. High-hit, long-context workloads benefit most because each successful cache action avoids more prefill work and a larger uncached-input charge. This effect is especially strong for agentic workloads with long inputs and short outputs [1], [47], [71]. During external pauses, tool-augmented and coding agents also require cache retention, offloading, eviction, and reload [70], [71], [72]. Within the modeled setting, successful mapping and overlapped mapping computation make the TTFT and cost benefits grow with the switching rate. Lower cache-aware switching costs could also let a joint routing policy choose lower-cost or bettersuited models more often, but this effect is outside the fixedpolicy comparison. Figure 5 shows that higher cache availability and switching rates produce larger ratios, while bandwidth and context length determine the attainable range. In deployment, the Unified AI Gateway can use lower switching costs to increase the model-switching probability and coordinate routing and cache actions at higher-benefit operating points. The resulting higher switching probability may provide additional gains beyond the fixed-policy results reported here.

7

D EPLOYMENT R EADINESS AND R ESEARCH D I -

RECTIONS

Section 5 summarizes the current implementation status of the main gateway components, while Section 6 analyzes their workload-level benefits through analytical simulation. Realizing these benefits in practice still requires addressing several gaps in current systems. This section assesses deployment readiness and identifies the remaining research directions.

7.1

Cross-Model Semantic Alignment

Cross-model KV cache mapping must preserve the semantics encoded by high-dimensional caches across heterogeneous models, since layer-wise errors can accumulate through attention and feedforward blocks. C2C and MoT demonstrated neural projection and mixture-of-translators on selected model pairs, but required pair-specific training [10], [33]. SCD, dense alignment, and sparse mapping reduced this burden in selected settings, but they still involved quality, training, or model-family constraints [43], [44], [74]. As the model pool grows, the gateway therefore needs mapping error bounds, low-sample or training-free adaptation, and benchmarks across models, tasks, and languages. 7.2

Heterogeneous Edge Hardware

Edge nodes differ in compute units, memory hierarchies, interconnects, and power limits, so GPU-oriented operators and layouts cannot be transferred directly to neural processing unit (NPU), field-programmable gate array (FPGA), or application-specific integrated circuit (ASIC) platforms [75]. Datacenter KV methods also assume high-bandwidth memory (HBM) and large batches, whereas edge devices make cache residence, eviction, and reload more costly. SwiftCache and Kelle showed the value of platform-specific cache and memory co-design [39], [76]. ENEC further illustrated this hardware dependence with lossless model-weight compression on Ascend NPUs [77]. A gateway-wide solution must adapt compression, placement, numerical format, and scheduling to device memory, power, and thermal profiles. 7.3

Distributed KV Cache Consistency

Distributed placement makes KV cache consistency a performance concern. Because a KV cache can be recomputed from the context, stale metadata usually causes a miss and repeated prefill rather than an incorrect result, which motivated asynchronous refresh and weak consistency in

12

(a) TTFT speedup versus bandwidth.

(b) TTFT speedup versus sequence length.

Fig. 5: TTFT speedup under variations in bandwidth, sequence length, cache-hit rate, and model-switching rate. Values are normalized to the conventional-gateway baseline defined in Eq. (5), and the horizontal line at 1.0 denotes equal latency. Panel (a) uses h = 0.5, while Panel (b) uses 64 Gbps. Colors encode context length in Panel (a) and cache-hit rate in Panel (b), and solid, dashed, and dotted lines represent switching rates of 25%, 50%, and 75%, respectively. The light blue, green, and orange background bands in Panel (b) correspond to the curve groups with cache-hit rates of h = 0.2, h = 0.5, and h = 0.8, respectively. P2P inference [78]. TraCT, SwiftCache, and Prefill-as-aService illustrated different approaches to shared access, ownership reassignment, and separated cache generation and consumption [50], [76], [79]. The gateway must codesign consistency, mobility, placement, and routing so that migration and metadata maintenance do not erase reuse gains. 7.4

Security and Privacy

KV caches encode user prompts and generated content, so edge storage and sharing create reconstruction and timingchannel risks. PROMPTPEEK, black-box analyses, SpliceLeak, and KV-Cloak show that these risks apply to prefix hits, non-prefix fusion, and direct tensor access [80], [81], [82], [83]. PrefixWall shows that selective isolation can retain part of the reuse benefit [84]. Unified gateway control therefore requires authorization, separation of users’ data, auditing, timing-signal protection, encryption, and dataresidency enforcement. 7.5

System Economics

Caching, compression, mapping, and migration change compute, storage, and network costs simultaneously, while prices and bottlenecks vary across locations and time. Existing analyses cover parts of this trade-off, but do not connect hit rate, context length, cache lifecycle, bandwidth price, and service quality in one gateway objective [85], [86]. Interregion traffic must therefore justify its latency benefit, and future models should include storage lifecycle, bandwidth billing, energy, hardware depreciation, and break-even conditions [50], [87]. 7.6

Standards and Ecosystem Support

KV cache interoperability remains less mature than unified APIs. Engines differ in page size, cache layout, precision,

and device affinity, making point-to-point connectors costly as model and engine versions grow. LMCache and emerging interoperation proposals provided common control events and cache descriptions, but compression, mapping, and transfer interfaces remained open [5]. SCBench covered several cache operations but not cross-model mapping, cross-node transfer, or consistency semantics [73]. Progress therefore requires a common KV cache data model, lifecycle and transfer events, reference implementations, and crossvendor benchmarks. 7.7

KV Cache Transfer and Memory Management

Transfer efficiency depends on memory layout and compression. PagedAttention’s non-contiguous blocks can reduce PCIe efficiency during direct transfer, while compaction adds copies during decode [12]. SwiftCache, KVServe, CacheGen, and TraCT showed that reload, bandwidth-aware compression, streaming, and shared memory must be considered together [27], [32], [76], [79]. The gateway therefore needs page-aware transfer primitives and SLO-aware schedulers that jointly select reuse, compression, priority, and placement. 7.8

KV Cache Portability and Architectural Evolution

Model design and KV cache movement are still optimized separately. Surveys and recent systems expose a gap between mature token and system optimizations and less developed model-level portability objectives [31], [32], [44], [88]. Future models could incorporate low-rank cache structure, sparse attention, or compression regularization to make cache movement more efficient. Although this paper focuses on KV caches, future gateway designs may extend the same coordination principle

13

to other transferable inference representations used by recurrent, non-autoregressive, or diffusion architectures [89], [90], [91], [92].

8

C ONCLUSION

This paper has defined the Unified AI Gateway setting, formalized its joint model-routing and KV cache decision, and connected the End–Edge–Cloud architecture to the responsibilities and requirements of its main components. It has also synthesized component-level evidence and analyzed the performance, input-cost, and model-switching benefits of coordinated KV cache handling. Across eight typical workload profiles and model-switching rates from 25% to 75%, the workload-level analytical simulation reports TTFT speedups of 1.25×–13.28× and input-cost benefits of 1.20×– 6.16×. The sensitivity analysis shows that these modeled benefits grow with reusable-cache coverage and switching frequency, provided that KV cache transfer can avoid targetside full prefill. Realizing this setting remains an end-toend systems integration problem involving cache fidelity, distributed resources, and trustworthy coordination.

R EFERENCES [1] [2] [3]

[4] [5]

[6]

[7]

[8]

[9]

Y. Qiao, T. D. Le, A. Shen, Z. Li, and B. Wang, “Serving agentic workloads at scale with vLLM x Mooncake,” https://vllm.ai/ blog/2026-05-06-mooncake-store, 2026, accessed: 2026-08-28. vLLM Developers, “Automatic prefix caching,” https://docs. vllm.ai/en/latest/features/automatic prefix caching.html, 2026, accessed: 2026-08-27. SGLang Developers, “Session-aware radix cache,” https://github.com/sgl-project/sglang/blob/main/docs/docs/ advanced features/session radix cache.mdx, 2026, accessed: 2026-08-27. Anthropic, “Prompt caching,” https://platform.claude.com/ docs/en/build-with-claude/prompt-caching, 2026, accessed: 2026-08-27. Y. Liu, Y. Cheng, J. Yao, Y. An, X. Chen, S. Feng, Y. Huang, S. Shen, R. Zhang, K. Du, and J. Jiang, “LMCache: An efficient KV cache layer for enterprise-scale LLM inference,” 2025, arXiv:2510.09665. [Online]. Available: https://arxiv.org/abs/2510.09665 R. Qin, Z. Li, W. He, J. Cui, M. Zhang, Y. Wu, W. Zheng, and X. Xu, “Mooncake: Trading more storage for less computation—a KVCache-centric architecture for serving LLM chatbot,” in 2025 USENIX Conference on File and Storage Technologies (FAST 25). Santa Clara, CA: USENIX Association, Feb. 2025, pp. 155–170. [Online]. Available: https://www.usenix.org/conference/fast25/ presentation/qin J. Yang, B. Hou, W. Wei, Y. Bao, and S. Chang, “KVLink: Accelerating large language models via efficient KV cache reuse,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2025. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2025/hash/ c217d123885341853cecbdd2be809983-Abstract-Conference.html X. Zhao and S. Mastorakis, “SemShareKV: Efficient KVCache sharing for semantically similar prompts via token-level LSH matching,” in Proceedings of the 14th International Joint Conference on Natural Language Processing and the 4th Conference of the Asia-Pacific Chapter of the Association for Computational Linguistics. The Asian Federation of Natural Language Processing and the Association for Computational Linguistics, Dec. 2025, pp. 432–445. [Online]. Available: https://aclanthology.org/2025.findings-ijcnlp.25/ Y. Liu, Y. Huang, J. Yao, S. Feng, Z. Gu, K. Du, H. Li, Y. Cheng, J. Jiang, S. Lu, M. Musuvathi, and E. Choukse, “DroidSpeak: KV cache sharing across fine-tuned model variants,” in 23rd USENIX Symposium on Networked Systems Design and Implementation (NSDI 26). Renton, WA: USENIX Association, May 2026, pp. 319–338. [Online]. Available: https: //www.usenix.org/conference/nsdi26/presentation/liu-yuhan

[10] T. Fu, Z. Min, H. Zhang, J. Yan, G. Dai, W. Ouyang, and Y. Wang, “Cache-to-cache: Direct semantic communication between large language models,” in Proceedings of the International Conference on Learning Representations (ICLR), 2026. [Online]. Available: https://proceedings.iclr.cc/paper files/paper/2026/hash/ 474ada926b331d78f06d95e8913111cc-Abstract-Conference.html [11] L. Zheng, L. Yin, Z. Xie, C. Sun, J. Huang, C. H. Yu, S. Cao, C. Kozyrakis, I. Stoica, J. E. Gonzalez, C. Barrett, and Y. Sheng, “SGLang: Efficient execution of structured language model programs,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2024. [Online]. Available: https://proceedings.neurips.cc/paper files/paper/2024/hash/ 724be4472168f31ba1c9ac630f15dec8-Abstract-Conference.html [12] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. E. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with PagedAttention,” in Proceedings of the ACM Symposium on Operating Systems Principles (SOSP), 2023, pp. 611–626. [13] Z. Pan, A. Patel, Y. Shen, Z. Hu, Y. Guan, W.-L. Li, L. Qin, Y. Wang, and Y. Ding, “KVFlow: Efficient prefix caching for accelerating LLM-based multi-agent workflows,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2025. [Online]. Available: https://papers.neurips.cc/paper files/paper/2025/hash/ b7971d31a7d5eb0f1eed2f8f6f368195-Abstract-Conference.html [14] R. Luo, C. Gu, Q. He, F. Chen, S. Wu, H. Jin, and Y. Yang, “SimLLM: Optimizing LLM inference at the edge through inter-task KV reuse,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2025. [15] J. Yao, H. Li, Y. Liu, S. Ray, Y. Cheng, Q. Zhang, K. Du, S. Lu, and J. Jiang, “CacheBlend: Fast large language model serving for RAG with cached knowledge fusion,” in Proceedings of the European Conference on Computer Systems (EuroSys), 2025, pp. 94–109. [Online]. Available: https://doi.org/10.1145/3689031.3696098 [16] Q. Zhang, K. Chen, N. Liao, Z. Lin, B. Tang, F. Xiong, Z. Li, and X. Wang, “SparseX: Efficient segment-level KV cache sharing for interleaved LLM serving,” 2026, arXiv:2606.01751. [17] Y. Liu, Z. Luo, H. Jin, Z. Wang, R. He, B. Wang, G. Chen, T. Xie, and J. Hu, “RedKnot: Efficient long-context LLM serving with headaware KV reuse and SegPagedAttention,” 2026, arXiv:2606.06256. [18] N. P. Pandey, J. Kong, L. Hu, Q. Zhao, Y. Zhao, O. Gungor, H. Zhang, and T. Rosing, “AgentKVShift: Efficient KV cache reuse for agentic memory systems,” 2026, arXiv:2607.21604. [19] H. Ye, Z. Gao, M. Ma, Q. Wang, Y. Fu, M.-Y. Chung, Y. Lin, Z. Liu, J. Zhang, D. Zhuo, and Y. Chen, “KVCOMM: Online cross-context KV-cache communication for efficient LLM-based multi-agent systems,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2025. [Online]. Available: https://papers.nips.cc/paper files/paper/2025/hash/ 1a074a28c3a6f2056562d00649ae6416-Abstract-Conference.html [20] A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, L. Kaiser, and I. Polosukhin, “Attention is all you need,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2017, pp. 5998–6008. [21] Qwen Team, “Qwen3 technical report,” 2025, arXiv:2505.09388. [Online]. Available: https://arxiv.org/abs/2505.09388 [22] DeepSeek-AI, “DeepSeek-V4: Towards highly efficient milliontoken context intelligence,” 2026, arXiv:2606.19348. [Online]. Available: https://arxiv.org/abs/2606.19348 [23] P. Patel, E. Choukse, C. Zhang, A. Shah, Í. Goiri, S. Maleki, and R. Bianchini, “Splitwise: Efficient generative LLM inference using phase splitting,” in Proceedings of the International Symposium on Computer Architecture (ISCA), 2024, pp. 118–132. [24] Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “DistServe: Disaggregating prefill and decoding for goodput-optimized large language model serving,” in Proceedings of the USENIX Symposium on Operating Systems Design and Implementation (OSDI), 2024, pp. 193–210. [25] A. Georgiou, “The price of anarchy in disaggregated inference,” 2026, arXiv:2606.17081. [Online]. Available: https://arxiv.org/ abs/2606.17081 [26] Z. Zhang, H. Shen, S. Vargaftik, R. B. Basat, M. Mitzenmacher, and M. Yu, “HACK: Homomorphic acceleration via compression of the key-value cache for disaggregated LLM inference,” in Proceedings of the ACM SIGCOMM Conference, 2025, pp. 1245–1247. [Online]. Available: https://doi.org/10.1145/3718958.3750481

14

[27] Y. Liu, H. Li, Y. Cheng, S. Ray, Y. Huang, Q. Zhang, K. Du, J. Yao, S. Lu, G. Ananthanarayanan, M. Maire, H. Hoffmann, A. Holtzman, and J. Jiang, “CacheGen: KV cache compression and streaming for fast large language model serving,” in Proceedings of the ACM SIGCOMM Conference, 2024, pp. 38–56. [Online]. Available: https://doi.org/10.1145/3651890.3672274 [28] Z. Su, H. Wei, Z. Chen, W. Shen, L. Li, H. Yu, and K. Yuan, “RotateKV: Accurate and robust 2-bit KV cache quantization for LLMs via outlier-aware adaptive rotations,” in Proceedings of the International Joint Conference on Artificial Intelligence (IJCAI), 2025, pp. 6200–6208. [Online]. Available: https://www.ijcai.org/proceedings/2025/690 [29] J.-Q. Wang, X.-Q. Han, P.-J. Guo, R.-Q. He, Z.-F. Gao, and Z.-Y. Lu, “Enabling efficient low-bit quantization based on matrix product operators for KV cache compression,” Neural Networks, vol. 197, p. 108467, 2025. [Online]. Available: https://doi.org/10.1016/j.neunet.2025.108467 [30] D. Yao, B. Shen, Z. Lin, W. Liu, J. Luan, B. Wang, and W. Wang, “TailorKV: A hybrid framework for long-context inference via tailored KV cache optimization,” in Findings of the Annual Meeting of the Association for Computational Linguistics (ACL), 2025, pp. 20 340–20 359. [Online]. Available: https://aclanthology.org/2025.findings-acl.1043/ [31] P. Bhatnagar, A. Moradifirouzabadi, S.-H. Yang, S. Lee, J. Choi, and M. Kang, “STAR-KV: Low-rank KV cache compression via soft thresholding for adaptive rank control,” 2026, arXiv:2606.08382. [32] Z. Liu, X. Ma, D. Luo, H. Zhao, B. Lu, W. Huang, Y. Gu, X. Liu, Z. Wei, J. Liu, D. Tao, and G. Tan, “KVServe: Service-aware KV cache compression for communication-efficient disaggregated LLM serving,” in Proceedings of the ACM SIGCOMM 2026 Conference, 2026, pp. 1–15. [Online]. Available: https://doi.org/10.1145/3789240.3829139 [33] J.-W. Lee, M. Song, J. Oh, S. Han, S. Park, G. Jang, and S. Lim, “Mixture-of-translators: Translating KV caches across heterogeneous large language models,” 2026, arXiv:2607.28979. [34] H. Jin, K. Peng, Y. Yu, X. Yuan, and H. Wang, “Agent primitives: Reusable latent building blocks for multi-agent systems,” 2026, arXiv:2602.03695. [35] Y. Geng, Y. Gao, W. Wu, G. Liu, and J. Liu, “RelayCaching: Accelerating LLM collaboration via decoding KV cache reuse,” 2026, arXiv:2603.13289. [36] C. Hooper, S. Kim, H. Mohammadzadeh, M. W. Mahoney, Y. S. Shao, K. Keutzer, and A. Gholami, “KVQuant: Towards 10 million context length LLM inference with KV cache quantization,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2024. [37] Y. He, L. Zhang, W. Wu, J. Liu, H. Zhou, and B. Zhuang, “ZipCache: Accurate and efficient KV cache quantization with salient token identification,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2024. [38] A. Zandieh, M. Daliri, M. Hadian, and V. Mirrokni, “TurboQuant: Online vector quantization with near-optimal distortion rate,” in Proceedings of the International Conference on Learning Representations (ICLR), 2026. [39] T. Xia and S. Q. Zhang, “Kelle: Co-design KV caching and eDRAM for efficient LLM serving in edge computing,” in Proceedings of the IEEE/ACM International Symposium on Microarchitecture (MICRO), 2025, pp. 18–33. [40] H. Jeon, H. Ha, and J.-J. Kim, “LRAgent: Efficient KV cache sharing for multi-LoRA LLM agents,” in ICML 2026 Poster, 2026, poster #2011. [Online]. Available: https://arxiv.org/abs/2602.01053 [41] A. Li, K. Greenewald, T. Parnell, and N. Azizan, “Efficient multiadapter LLM serving via cross-model KV-cache reuse with activated LoRA,” 2025, arXiv:2512.17910. [42] S. Woo, J. Kil, H. Kim, M. Kim, J. Kim, A. Seo, S. Lee, M. Jo, J. Ryu, B. Park, S. J. Kwon, and D. Lee, “ICaRus: Identical cache reuse for efficient multi-model inference,” in Proceedings of the International Conference on Learning Representations (ICLR), 2026. [43] T. Heo, R. Shafipour, R. Zhao, M. Golub, M. M. Kamani, R. Borkar, M. T. Chandran, P. Zardoshti, and B. D. Rouhani, “Cross-model KV cache transfer in LLM families: A closed-form linear mapping for prefill reuse,” 2026, arXiv:2608.03893. [44] Q. Ma, Z. Tang, H. Cui, Z. Yao, and W. Jia, “Semantic cache distillation: Efficient state transfer via reuse and selective patching,” in Proceedings of the International Conference on Machine Learning (ICML), 2026.

[45] Y. Zhu, A. Dhakal, Y. Xiao, D. Milojicic, and G. Alonso, “ObjectCache: Layerwise object-storage retrieval for KV cache reuse,” 2026, arXiv:2605.22850. [46] S. R. Ganjihal, “Predictive multi-tier memory management for KV cache in large-scale GPU inference,” 2026, arXiv:2604.26968. [47] V. Srivatsa, Z. He, R. Abhyankar, D. Li, and Y. Zhang, “Preble: Efficient distributed prompt scheduling for LLM serving,” in Proceedings of the International Conference on Learning Representations (ICLR), 2025. [Online]. Available: https://proceedings.iclr.cc/paper files/paper/2025/hash/ 5bc342f48de8264779952fac378f96dc-Abstract-Conference.html [48] X. Luo, J. Shen, X. Wang, and Y. Zhou, “SmartGen: Seamless disaggregated LLM inference with selective KV cache transfer,” 2026, arXiv:2607.28150. [49] Y. Guo and S. Joshi, “SplitZip: Ultra fast lossless KV compression for disaggregated LLM serving,” 2026, arXiv:2605.01708. [50] R. Qin, W. He, Y. Wang, Z. Li, X. Xu, Y. Wu, W. Zheng, and M. Zhang, “Prefill-as-a-service: KVCache of next-generation models could go cross-datacenter,” 2026, arXiv:2604.15039. [51] W. Li, G. Jiang, X. Ding, Z. Tao, C. Hao, C. Xu, Y. Zhang, and H. Wang, “FlowKV: A disaggregated inference framework with low-latency KV cache transfer and load-aware scheduling,” 2025, arXiv:2504.03775. [52] S. Chen, R. Jiang, D. Yu, J. Xu, M. Chao, F. Meng, C. Jiang, W. Xu, and H. Liu, “KVDirect: Distributed disaggregated LLM inference,” 2025, arXiv:2501.14743. [53] K. Mei, W. Xu, M. Guo, S. Lin, and Y. Zhang, “OmniRouter: Budget and performance controllable multi-LLM routing,” SIGKDD Explorations, vol. 27, no. 2, pp. 107–116, 2025. [Online]. Available: https://kdd.org/exploration files/p107 Omnirouter camera ready.pdf [54] F. Wu and S. Silwal, “Efficient training-free online routing for highvolume multi-LLM serving,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2025. [55] I. Ong, A. Almahairi, V. Wu, W.-L. Chiang, T. Wu, J. E. Gonzalez, M. W. Kadous, and I. Stoica, “RouteLLM: Learning to route LLMs from preference data,” in Proceedings of the International Conference on Learning Representations (ICLR), 2025. [56] L. Chen, M. Zaharia, and J. Zou, “FrugalGPT: How to use large language models while reducing cost and improving performance,” Transactions on Machine Learning Research (TMLR), 2024. [57] P. Jaillet, J. Jiang, K. Mellou, M. Molinaro, C. Podimata, and Z. Zhou, “Online scheduling for LLM inference with KV cache constraints,” 2025, arXiv:2502.07115. [58] F. Wu, S. Silwal, and Q. Zhang, “Randomization boosts KV caching, learning balances query load: A joint perspective,” in Proceedings of the International Conference on Learning Representations (ICLR), 2026. [59] J. Cheng, D. T. Do, and D. T. Nguyen, “Robust KV cache management for LLM serving under output token length uncertainty,” arXiv:2607.16892, 2026. [Online]. Available: https://arxiv.org/abs/2607.16892 [60] J. Wang, S. Zhao, H. Wang, Y. Fan, L. Zhang, Y. Liu, and T. Liu, “Optimal-agent-selection: State-aware routing framework for efficient multi-agent collaboration,” https://arxiv.org/abs/ 2511.02200, 2025, arXiv:2511.02200. [61] H. Cheng, “CacheRoute: Planned prefix-affinity routing for largescale LLM serving,” arXiv:2608.19677, 2026. [Online]. Available: https://arxiv.org/abs/2608.19677 [62] B. Gao, Z. He, P. Sharma, Q. Kang, D. Jevdjic, J. Deng, X. Yang, Z. Yu, and P. Zuo, “Cost-Efficient large language model serving for multi-turn conversations with CachedAttention,” in 2024 USENIX Annual Technical Conference (USENIX ATC 24). Santa Clara, CA: USENIX Association, Jul. 2024, pp. 111–126. [Online]. Available: https://www.usenix.org/conference/atc24/ presentation/gao-bin-cost [63] X. Liu, B. He, H. Chen, H. Zhang, A. Luo, and the vLLM Semantic Router Team, “Session-aware agentic routing: Continuity-aware model selection for long-horizon LLM agents,” https://vllm-project.github.io/2026/06/02/ session-aware-agentic-routing.html, 2026, accessed: 2026-09-02. [64] Zhipu AI, “GLM-5.3-Flash,” https://docs.bigmodel.cn/cn/ guide/models/vlm/glm-5.3-flash, 2026. [65] Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “DistServe artifact,” https://github.com/LLMServe/ DistServe, 2024.

15

[66] S. Arun, A. Parayil, S. Bharadwaj, R. S. Amant, and V. Ruhle, “Towards load-aware prefill deflection for disaggregated LLM serving,” 2026, arXiv:2607.02043. [67] Anthropic, “Pricing,” https://docs.anthropic.com/en/docs/ about-claude/pricing, 2026, accessed: 2026-08-27. [68] J. Lin, J. Mi, Z. Hong, H. Wang, Q. Liu, H. Zhang, P. Li, and S. Guo, “KVDrive: A holistic multi-tier KV cache management system for long-context LLM inference,” arXiv:2605.18071, 2026. [Online]. Available: https://arxiv.org/abs/2605.18071 [69] OpenAI, “Chatgpt rate card (enterprise tokenbased pricing),” https://help.openai.com/en/articles/ 20001415-chatgpt-rate-card-enterprise-token-based-pricing, 2026, accessed: 2026-08-27. [70] R. Abhyankar, Z. He, V. Srivatsa, H. Zhang, and Y. Zhang, “INFERCEPT: Efficient intercept support for augmented large language model inference,” in Proceedings of the International Conference on Machine Learning (ICML), 2024, pp. 81–95. [Online]. Available: https://icml.cc/virtual/2024/poster/32755 [71] H. Li, Q. Mang, R. He, Q. Zhang, H. Mao, X. Chen, H. Zhou, A. Cheung, J. Gonzalez, and I. Stoica, “Continuum: Efficient and robust multi-turn LLM agent scheduling with KV cache time-to-live,” in Proceedings of the Lifelong Agent Workshop at ICLR 2026, 2026. [Online]. Available: https://openreview.net/ pdf/5c6ed290be90eea7b75202124c5bffecac714a4c.pdf [72] S. Tiwari, T. Chugh, N. Rickert, S. Peter, R. Mahajan, and H. Shen, “CacheWise: Understanding workloads and optimizing KVCache management for efficiently serving LLM coding agents,” 2026, arXiv:2606.16824. [73] Y. Li, H. Jiang, Q. Wu, X. Luo, S. Ahn, C. Zhang, A. H. Abdi, D. Li, J. Gao, Y. Yang, and L. Qiu, “SCBench: A KV cache-centric analysis of long-context methods,” in Proceedings of the International Conference on Learning Representations (ICLR), 2025. [74] S. Chen, X. Zhang, M. Wu, J. Tremblay, V. Blukis, and S. Birchfield, “See what i see, know what i think: Dense latent communication across heterogeneous agents,” 2026, arXiv:2606.13594. [75] B. Xu, A. Banerjee, and S. K. S. Gupta, “Hardware acceleration for neural networks: A comprehensive survey,” 2025, arXiv:2512.23914. [76] J. Hu, M. Xu, S. Wang, C. Ma, M. Shen, K. Ye, L. Qu, and C. Xu, “SwiftCache: Efficient LLM serving for multi-turn conversations with heterogeneous KV cache sharing,” 2026, arXiv:2606.16135. [77] J. Yang, J. Wu, Z. Liu, X. Ma, H. Zhao, Y. Gu, Y. Huang, X. Liu, W. Huang, Z. Wei, J. Xing, Y. Ma, Q. Zhang, B. An, Z. Hu, S. Liu, X. Zhu, J. Lu, G. Tan, and D. Tao, “ENEC: A lossless AI model compression method enabling fast inference on ascend NPUs,” in Proceedings of the 53rd Annual International Symposium on Computer Architecture, 2026, pp. 1319–1335. [78] S. S. Nair and K. Saini, “Towards distributed inference of LLMs on a P2P network,” 2026, arXiv:2606.17059. [79] D. Yoon, Y. Min, H. Kim, S. H. Noh, and J. Kim, “TraCT: Disaggregated LLM serving with CXL shared memory KV cache at rack-scale,” 2025, arXiv:2512.18194. [80] G. Wu, Z. Zhang, Y. Zhang, W. Wang, J. Niu, Y. Wu, and Y. Zhang, “I know what you asked: Prompt leakage via KV-cache sharing in multi-tenant LLM serving,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2025. [81] L. Song, Z. Pang, W. Wang, Z. Wang, X. Wang, H. Chen, W. Song, Y. Jin, D. Meng, and R. Hou, “The early bird catches the leak: Unveiling timing side channels in LLM serving systems,” IEEE Transactions on Information Forensics and Security, 2025. [82] H. Sun, S. Liu, S. Ma, J. Li, M. Xiao, and W. Jiang, “Agentassisted side-channel attacks on non-prefix KV cache in RAG,” 2026, arXiv:2606.21842. [83] Z. Luo, S. Shao, S. Zhang, L. Zhou, Y. Hu, C. Zhao, Z. Liu, and Z. Qin, “Shadow in the cache: Unveiling and mitigating privacy risks of KV-cache in LLM inference,” in Proceedings of the Network and Distributed System Security Symposium (NDSS), 2026. [84] P. G. Pennas, K. Papaioannou, M. Guarnieri, and T. D. Doudali, “PrefixWall: Mitigating prefix caching side channels in shared LLM systems,” 2026, arXiv:2603.10726. [Online]. Available: https://arxiv.org/abs/2603.10726 [85] E. Erdil, “Inference economics of language models,” 2025, arXiv:2506.04645. [86] N. Pollertlam and W. Kornsuwannawit, “Beyond the context window: A cost-performance analysis of fact-based memory vs. longcontext LLMs for persistent agents,” 2026, arXiv:2603.04814.

[87] Y. Sun, B. Lei, J. Liu, H. Huang, X. Zhang, J. Peng, and W. Wang, “Computing power network: A survey,” China Communications, vol. 21, no. 9, pp. 109–145, 2024. [Online]. Available: https://ieeexplore.ieee.org/document/10495806/ [88] H. Li, Y. Li, A. Tian, T. Tang, Z. Xu, X. Chen, N. Hu, W. Dong, Q. Li, and L. Chen, “A survey on large language model acceleration based on KV cache management,” Transactions on Machine Learning Research (TMLR), 2025. [89] A. Gu and T. Dao, “Mamba: Linear-time sequence modeling with selective state spaces,” in Proceedings of the Conference on Language Modeling (COLM), 2024. [Online]. Available: https://openreview.net/pdf?id=tEYskw1VY2 [90] B. Peng, E. Alcaide, Q. Anthony, A. Albalak, S. Arcadinho, S. Biderman, H. Cao, X. Cheng, M. Chung, L. Derczynski, X. Du, M. Grella, K. Gv, X. He, H. Hou, P. Kazienko, J. Kocon, J. Kong, B. Koptyra, H. Lau, J. Lin, K. S. I. Mantri, F. Mom, A. Saito, G. Song, X. Tang, J. Wind, S. Wožniak, Z. Zhang, Q. Zhou, J. Zhu, and R.-J. Zhu, “RWKV: Reinventing RNNs for the Transformer era,” in Findings of the Association for Computational Linguistics: EMNLP 2023. Singapore: Association for Computational Linguistics, Dec. 2023, pp. 14 048–14 077. [Online]. Available: https://aclanthology.org/2023.findings-emnlp.936/ [91] J. Gu, J. Bradbury, C. Xiong, V. O. K. Li, and R. Socher, “Nonautoregressive neural machine translation,” in Proceedings of the International Conference on Learning Representations (ICLR), 2018. [92] X. L. Li, J. Thickstun, I. Gulrajani, P. Liang, and T. B. Hashimoto, “Diffusion-LM improves controllable text generation,” in Proceedings of the Annual Conference on Neural Information Processing Systems (NeurIPS), 2022.

Record · ID 668000 · SHA-256 0b2ffde0b42ca1c0
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.