arXiv:2606.17104v1 [cs.AR] 14 Jun 2026
Prefill/Decode-Aware Evaluation of LLM Inference on Emerging AI Accelerators Shun Usami∗
Venkatram Vishwanath†
E. Wes Bethel∗‡
∗ Department of Computer Science
† Argonne National Laboratory
∗ Department of Computer Science
San Francisco State University San Francisco, CA, 94132 [email protected]
Lemont, IL, 60439 [email protected]
San Francisco State University San Francisco, CA, 94132 ‡ Lawrence Berkeley National Laboratory Berkeley, CA, 94720 [email protected]
Abstract—As large language models (LLMs) are increasingly deployed in latency- and cost-sensitive settings, inference efficiency has become a central systems challenge. While GPUs dominate current deployments, a growing number of AI accelerators claim advantages for LLM inference, yet it remains unclear under which conditions such accelerators outperform GPUs in practice. Recent inference systems decompose execution into Prefill and Decode phases, which exhibit distinct computational characteristics and latency metrics, commonly captured by time to first token (TTFT) and time per output token (TPOT). This paper presents a phase-aware evaluation of LLM inference performance across GPUs and emerging AI accelerators using a common model, Llama2-7B. By separately measuring Prefill and Decode performance, we reveal that accelerator advantages differ by phase and metric. Our results show that GPUs consistently excel in the compute-intensive Prefill phase, while GroqRack achieves significantly lower TPOT during Decode (batching not currently supported). However, GPUs regain an advantage in Decode throughput as batch size increases. These findings demonstrate that each platform exhibits distinct phase-dependent strengths. We further analyze heterogeneous Prefill/Decode disaggregation across different accelerator platforms, identifying performance gains and the workload and network conditions under which such gains are realized. Index Terms—Large Language Models, AI Accelerators, Performance Evaluation
I. I NTRODUCTION Large Language Models (LLMs) have become a central component of modern AI systems, with applications ranging from conversational agents to code generation and scientific assistance. As model sizes and usage continue to grow, the efficiency of inference, rather than training, has emerged as a dominant systems-level concern [1], [2]. In particular, latency-sensitive and cost-sensitive inference workloads place increasing pressure on the underlying hardware and system architecture. While GPUs dominate LLM inference due to strong linear algebra performance and mature ecosystems, emerging domain-specific accelerators claim advantages for inference workloads. However, it remains unclear under what conditions, Accepted to HPAI4S’26, co-located with IEEE IPDPS 2026. © 2026 IEEE. Personal use of this material is permitted. Permission from IEEE must be obtained for all other uses.
and for which components of LLM inference, such accelerators provide tangible advantages over GPUs in practice. Recent advances in LLM serving systems have highlighted the importance of decomposing inference into two distinct phases: Prefill and Decode [3]. These phases exhibit fundamentally different computational characteristics—Prefill being compute-intensive and benefiting from parallelism, while Decode is memory-bound with sequential dependencies. This Prefill/Decode (P/D) separation has been formalized in disaggregated serving architectures and modern inference frameworks. While Prefill/Decode disaggregation has primarily been explored in GPU-based systems, it raises a natural and important question: do different AI accelerators exhibit distinct strengths in the Prefill and Decode phases, and if so, how should such differences be measured and interpreted? Answering this question is essential for evaluating the practical role of emerging accelerators and for guiding future heterogeneous Prefill/Decode disaggregation designs, where the Prefill and Decode phases are assigned to different accelerator platforms. Standardized benchmarks such as MLPerf Inference [4] provide a common framework for evaluating LLM inference performance. However, participation from emerging AI accelerators remains limited, and reported results typically focus on aggregate metrics such as end-to-end latency or overall throughput. For example, MLPerf Inference specifies 99thpercentile compliance thresholds on time to first token (TTFT, reflecting Prefill latency) and time per output token (TPOT, reflecting per-step Decode latency)—e.g., TTFT ≤ 2000 ms and TPOT ≤ 200 ms for Llama 2 70B—yet each submission reports only aggregate throughput, leaving Prefill/Decode trade-offs uncharacterized across accelerators. As a result, there is a lack of systematic, phase-aware evaluation that isolates and compares Prefill and Decode behavior across heterogeneous hardware platforms. In this work, we present a unified experimental evaluation of Prefill and Decode performance using a common LLM model, Llama2-7B, across GPU-based platforms and emerging AI accelerators. As a representative emerging accelerator, we study GroqRack to demonstrate our methodology and reveal phase-specific performance characteristics. We evaluate Prefill
and Decode performance across diverse inference scenarios to understand how accelerator strengths differ by inference phase and performance metric. Our evaluation reveals a clear phase-dependent performance asymmetry across hardware platforms. GPUs consistently excel at Prefill through batched, compute-intensive execution, while GroqRack demonstrates substantially lower per-token latency during Decode under single-request scenarios. However, GPUs recover throughput advantages as batch size increases, illustrating a fundamental latency-throughput trade-off. These findings show that no single accelerator uniformly dominates across all phases and objectives, that Prefill and Decode must be evaluated separately for meaningful hardware comparison, and that optimal hardware selection depends critically on the system’s objective function. The contributions of this paper are threefold. First, we introduce a phase-aware evaluation methodology for LLM inference that isolates Prefill and Decode behavior under controlled workload variations. Second, we empirically characterize GPUs and emerging AI accelerators, revealing phasedependent latency–throughput trade-offs obscured by endto-end metrics. Finally, we present a quantitative analysis of heterogeneous Prefill/Decode disaggregation that identifies performance gains and derives break-even conditions for practical deployment.
as latency constraints, and systems are evaluated based on the maximum sustainable throughput achievable within predefined TTFT and TPOT bounds. In addition to performance constraints, MLPerf Inference validates output correctness against reference datasets using standard text similarity metrics (e.g., ROUGE), reflecting its goal of end-to-end system evaluation under application-facing service-level objectives. As of the latest v5.1 datacenter results, publicly reported LLM inference submissions are largely limited to GPU-based platforms from NVIDIA and AMD, with limited coverage of emerging AI accelerators.1 While MLPerf provides a comprehensive benchmark for deployed systems, its holistic design obscures phase-specific performance characteristics. Several recent studies have examined LLM inference performance beyond MLPerf. LLM-Inference-Bench [6] provides broader hardware coverage, including SambaNova SN40L, but reports aggregate endto-end metrics without separating Prefill and Decode phases. While DistServe focuses on GPU-based scheduling and existing benchmarks do not systematically evaluate emerging AI accelerators under a phase-aware framework, our work addresses this gap by systematically measuring Prefill and Decode performance across GPUs and specialized AI accelerators under controlled workload variations.
II. BACKGROUND AND R ELATED W ORK
GroqRack is built around GroqCards implementing Groq’s Language Processing Unit (LPU) architecture, formerly referred to as the Tensor Streaming Processor (TSP). At the chip level, the LPU executes statically scheduled programs with a fixed instruction and memory access order, avoiding caches, dynamic scheduling, and out-of-order execution commonly used in GPUs to optimize average throughput. Instead, it relies on a large, single-level on-chip SRAM scratchpad, with model weights and intermediate states explicitly allocated and managed by the compiler [7]. This deterministic execution model extends beyond a single chip to the rack scale. In GroqRack systems, dozens of GroqCards are interconnected via a compiler-orchestrated communication fabric, allowing both computation and interdevice communication to follow a fixed, statically determined schedule [8].
A. Prefill/Decode Disaggregation LLM inference is commonly decomposed into two phases with distinct computational characteristics: a Prefill phase that processes the entire input sequence to generate the first output token, and a Decode phase that generates subsequent tokens auto-regressively [5]. The Prefill phase can be fully parallelized using matrix–matrix multiplication operators, making it compute-bound. During this process, per-token key and value vectors are produced and stored as the key–value (KV) cache for reuse in later steps. In contrast, the Decode phase is inherently sequential at the token level due to its autoregressive nature, and its frequent accesses to the KV cache make it memory-bound. Recent serving systems exploit this phase asymmetry through Prefill/Decode disaggregation. Zhong et al. [3] formalize this approach in DistServ by separating Prefill and Decode workers that scale independently. In such designs, the KV cache generated during Prefill phase must be transferred to Decode workers to enable subsequent token generation, introducing non-negligible communication overhead. Despite this cost, DistServe demonstrates that decoupling the two phases enables more efficient resource allocation and parallelism strategies aligned with their distinct latency characteristics, leading to substantially higher goodput in GPU-based clusters. B. LLM Inference Benchmarks MLPerf [4] is the industry-standard benchmark suite for evaluating machine learning systems. Its inference component includes LLM workloads where TTFT and TPOT are treated
C. Groq Architecture Overview
III. M ODEL E XECUTION D IFFERENCES ACROSS A RCHITECTURES This section details the key differences in how the same LLM model forward pass is executed on NVIDIA A100 GPUs and GroqRack accelerators, which directly impact phasespecific performance. A. NVIDIA A100 Execution Model On NVIDIA A100 GPUs, we use vLLM (v0.12.0), which supports batched execution. The model forward pass accepts inputs with batch size B > 1 and input token length Lin > 1 See the MLPerf Inference Datacenter results page: https://mlcommons.org/ benchmarks/inference-datacenter/.
TABLE I H ARDWARE P LATFORM S PECIFICATIONS ( BASED ON VENDOR - PUBLISHED SPECIFICATIONS )
(a) NVIDIA A100
(b) GroqRack
Fig. 1. Model execution during the Prefill phase. NVIDIA A100 processes all input tokens in a single batched forward pass, whereas the evaluated GroqRack implementation invokes the model forward pass iteratively with one token per invocation. For both platforms, the KV cache remains resident on the same device during Prefill and Decode; no cross-device KV cache transfer occurs under the evaluated configuration.
1, enabling the processing of many tokens and requests in a single invocation. As Fig. 1a shows, the Prefill phase processes all input tokens in a single parallel forward pass, while the Decode phase performs autoregressive forward passes with cached KV states until Lout tokens are generated. B. GroqRack Execution Model The GroqRack implementation uses the Groq SDK with the model sharded across 9 nodes (72 cards total). The compiled binary accepts only a single request and processes one token per invocation. As Fig. 1b shows, the Prefill phase processes input tokens sequentially, invoking the model forward pass once per token. After all input tokens are consumed, the first output token is sampled on the rank-0 host CPU and execution transitions to Decode. The Decode phase follows a similar iterative pattern with CPU-side sampling and MPI broadcast until Lout tokens are generated. This sequential, single-request execution stems from GroqRack’s statically scheduled, deterministic execution model with a fixed scratchpad SRAM layout. Because the LPU has no branch instructions, a compiled binary must execute with exactly the batch size it was compiled for — it is not possible to process fewer requests than the compiled batch size, making the current single-request binary the practical choice for variable workloads. IV. M ETHODOLOGY Our study aims to benchmark the Prefill and Decode phases of LLM inference across heterogeneous platforms using a consistent model and workload parameters. This section details
Specification
A100
GroqRack
Architecture
GPU
LPU
System Config
Single GPU
72 × GroqCards
SRAM
40 MiB L2
15.42 GiB total (230 MB per GroqCard)
DRAM
40 GB HBM2
—
Peak Compute (FP16)
312 TFLOPS
13.5 PFLOPS total (187.5 TFLOPS per GroqCard)
Memory Bandwidth
1.6 TB/s
5.76 PB/s total (80 TB/s per GroqCard)
our experimental setup, including model selection, hardware platforms, inference frameworks, workload configurations, and metrics used in our experiments. A. Model Selection We select Llama2-7B [9] as our benchmark model due to its implementation availability across multiple hardware platforms and its moderate size, which allows for efficient experimentation while still being representative of modern LLMs. The model is evaluated in its FP16 precision format, which is commonly used in inference deployments to balance performance and accuracy. With an FP16 model footprint of approximately 12.56 GiB plus KV cache requirements (512 KiB per token, up to 2 GiB for the maximum context length of 4096 tokens), the total working set of approximately 14.6 GiB fits within GroqRack’s 15.42 GiB total SRAM capacity (Table I), making Llama2-7B nearly the largest model that can be fully accommodated on this platform. B. Hardware Platforms We conduct experiments on an NVIDIA A100 SXM-40GB GPU as our baseline platform and GroqRack as a specialized AI accelerator. Table I summarizes the key architectural characteristics of each platform as relevant to our evaluation. For the A100 experiments, both phases are executed on the same single GPU, without hardware disaggregation. For GroqRack, both phases run on the same single GroqRack system. C. Inference Frameworks For NVIDIA A100, we utilize vLLM v0.12.0 due to its optimized performance for LLM inference and support for phase-aware execution. For GroqRack, we use the GroqFlow SDK and the compiled binary of Llama2-7B provided by Groq. D. Workload Configurations We vary three workload parameters: input length (Lin ), output length (Lout ), and batch size (B). For each configuration (Lin , Lout , B), all requests within a batch use prompts of equal length and generate the same number of output tokens. Input tokens are generated randomly, and EOS tokens are ignored to ensure a fixed output length.
Algorithm 1 Prefill and Decode Timing Procedure 1: Warmup ← model(random inputs) 2: 3: Prefill Phase 4: input ← GenerateRandomTokens(Lin ) 5: tprefill start ← Now() 6: next tok ← model(input) 7: tprefill end ← Now() 8: input.Append(next tok) 9: Decode Phase 10: tdecode start ← Now() 11: for i = 1 to Lout do 12: next tok ← model(input) 13: input.Append(next tok) 14: end for 15: tdecode end ← Now()
1) Token Lengths: We evaluate Lin = Lout ∈ {100, 200, 400, 800, 1600}, restricting workloads to cases where input and output lengths are equal, covering short to long-context workloads while remaining within the 4096-token context limit of Llama2-7B. 2) Batch Sizes: For NVIDIA A100, we evaluate batch sizes B ∈ {1, 2, 4, 8, 16, 32} to study throughput scaling. For GroqRack, we restrict experiments to single-request execution (B = 1). As described in Section III, GroqRack’s statically scheduled execution model requires a dedicated binary per batch size. Additionally, with model weights occupying 12.56 GiB and each request requiring up to 2 GiB of KV cache, a second concurrent request would push the total to 16.56 GiB, exceeding the 15.42 GiB SRAM capacity. E. Inference Measurement Protocol The inference timing procedure used in this study is summarized in Algorithm 1. We explicitly separate the Prefill phase (first-token generation) from the autoregressive Decode phase to obtain phase-specific latency and throughput measurements. We exclude tokenization and detokenization overhead from all timing measurements and focus solely on inference computation. Before collecting measurements, we perform warmup runs with unique input tokens to avoid KV cache reuse and to eliminate initialization overheads such as model compilation, program loading, and weight loading. Our evaluation considers only performance metrics (latency and throughput). We do not assess output correctness or semantic quality, as the goal is to characterize phase-specific performance behavior under controlled workloads. F. Performance Metrics We evaluate performance using the four metrics defined in Table II, where Lin , Lout , and B follow the workload definitions in Section IV-D.
TABLE II P ERFORMANCE M ETRICS FOR P REFILL AND D ECODE E VALUATION Metric
Formula
Measures
TTFT
tprefill end − tprefill start
Latency from Prefill start to first output token generation
TPOT
tdecode end − tdecode start Lout
Average latency per generated token during Decode phase
Prefill Throughput
Lin × B tprefill end − tprefill start
Input tokens processed per second during Prefill phase
Decode Throughput
Lout × B tdecode end − tdecode start
Output tokens generated per second during Decode phase
V. R ESULTS Our experimental evaluation reveals a strong phasedependent performance asymmetry across hardware platforms. The results demonstrate that the relative performance of accelerators depends critically on the inference phase and on whether latency or throughput is prioritized, with Prefill and Decode exhibiting fundamentally different behaviors. A. Prefill Performance The NVIDIA A100 GPU consistently outperforms GroqRack in the Prefill phase across all evaluated configurations. Fig. 2a shows that the GPU achieves substantially lower TTFT than GroqRack for all input token lengths. At batch size B = 1, GPU TTFT ranges from approximately 17 ms to 103 ms as input length increases from 100 to 1,600 tokens, whereas GroqRack exhibits strictly linear scaling from 252 ms to 4,072 ms over the same range. As batch size increases on the GPU, TTFT grows with the total number of processed input tokens (B × Lin ). Once the GPU reaches full utilization, TTFT scales approximately linearly with workload size. Prefill throughput results further highlight this disparity. Fig. 2b shows that GPU throughput increases with either input token length or batch size until saturation at B × Lin ≈ 800– 1,600 tokens, plateauing at approximately 15,000–16,000 tokens/s. In contrast, GroqRack maintains a nearly constant Prefill throughput of approximately 370–400 tokens/s across all evaluated input lengths. B. Decode Performance In contrast to the Prefill phase, the Decode phase exhibits a reversed performance trend. GroqRack consistently achieves substantially lower per-token latency than NVIDIA A100. Fig. 3a shows that GroqRack maintains a stable TPOT of approximately 3 ms across all evaluated configurations. At B = 1, GPU TPOT ranges from 11.88 ms to 13.64 ms and increases with output sequence length. At higher batch sizes and longer sequences, GPU TPOT increases significantly. For example, at B = 32 and 1,600 input + 1,600 output tokens, GPU TPOT reaches 57.90 ms
(a) TTFT
(b) Prefill throughput
Fig. 2. Prefill phase performance comparison when serving Llama2-7B on GroqRack (Batch Size 1) and NVIDIA A100 (Batch Size 1, 2, 4, 8, 16, 32).
at the final output token for sequences of length 3,200, while GroqRack maintains near-constant TPOT. Decode throughput exhibits an opposing trend. At B = 1, GroqRack achieves 328–336 tokens/s, compared to 73– 84 tokens/s on the GPU. As batch size increases, GPU throughput scales with the number of concurrent requests and surpasses GroqRack at B > 4. At short sequences (100 input/output tokens), the GPU reaches up to 2,144 tokens/s at B = 32. However, at longer sequences, GPU throughput degrades substantially, reaching 553 tokens/s at B = 32 for 1,600 input/output tokens, reducing the throughput advantage to approximately 1.7× over GroqRack. C. Performance Summary Table III summarizes the key performance metrics observed across Prefill and Decode phases. The results highlight a clear phase-dependent performance asymmetry: GPUs dominate compute-intensive Prefill workloads, while GroqRack provides substantially lower per-token latency during Decode. GPU Decode throughput benefits from batching, but this advantage diminishes at long sequence lengths. VI. D ISCUSSION This section interprets the experimental results from Section V and examines their implications for accelerator selection and system design. We first analyze the architectural factors underlying the observed performance differences, then assess the feasibility of heterogeneous Prefill/Decode disaggregation, and finally discuss broader system implications and limitations. A. Architectural Origins of Phase-Dependent Performance We begin by examining how architectural design choices and execution models influence performance across different inference phases.
TABLE III P ERFORMANCE C OMPARISON S UMMARY FOR L LAMA 2-7B I NFERENCE ON NVIDIA A100 GPU AND G ROQ R ACK Metric
NVIDIA A100
GroqRack
Prefill Phase TTFT (B = 1, Lin = 100)
16.8 ms
252 ms
TTFT (B = 1, Lin = 1,600)
103.7 ms
4,072 ms
Throughput (peak)
16,322 tok/s
370–397 tok/s
TPOT (B = 1)
11.88–13.64 ms
2.98–3.05 ms
TPOT (B = 32)
14.92–57.90 ms
N/A∗
Throughput (B = 1)
73–84 tok/s
328–336 tok/s
Throughput (B = 32)
553–2,144 tok/s
N/A∗
Decode Phase
∗ GroqRack supports B
= 1 only due to current implementation con-
straints.
1) Why GPUs Excel at Prefill: GPUs excel at the Prefill phase because their architectures are optimized for large-scale data-parallel execution of dense linear algebra. During Prefill, all input tokens across the batch can be processed concurrently, allowing the workload to be expressed as large matrix–matrix multiplications whose effective size scales with B × Lin . As B × Lin increases, this parallelism enables GPUs to efficiently amortize fixed overheads and approach peak compute utilization, leading to high Prefill throughput once the device is saturated. 2) Why GroqRack Performs Poorly at Prefill: GroqRack exhibits limited Prefill performance because the current Llama27B implementation processes input tokens sequentially rather than in parallel. As detailed in Section III, Prefill execution on GroqRack follows a token-by-token execution model similar to Decode, resulting in strictly linear TTFT scaling and nearly constant Prefill throughput. 3) Why GroqRack Excels at Decode Latency: GroqRack achieves low and stable Decode latency because Llama27B inference executes entirely on explicitly managed SRAM
(a) TPOT
(b) Decode throughput
Fig. 3. Decode phase performance comparison when serving Llama2-7B on GroqRack (Batch Size 1) and NVIDIA A100 (Batch Size 1, 2, 4, 8, 16, 32).
without any cache hierarchy. The model weights (12.56 GiB) and KV cache (up to 2 GiB) together fit within the statically allocated 15.42 GiB of distributed SRAM across GroqCards (Section IV-A), enabling deterministic, token-by-token execution. As a result, Decode latency is governed by fixed execution schedules rather than dynamic memory behavior, leading to near-constant TPOT across all evaluated sequence lengths. 4) Why GPU Decode Throughput Saturates and Degrades: While GPUs can achieve high Decode throughput by batching requests, this advantage diminishes as batch size and sequence length increase. In this regime, each Decode step to generate the next token must access a larger portion of the accumulated KV cache, quickly exceeding on-chip cache capacity and inducing frequent L1 and L2 cache misses. Consequently, Decode performance becomes dominated by lower-bandwidth HBM accesses rather than compute throughput. Beyond this bandwidth-driven slowdown, Decode performance further degrades when the KV cache footprint exceeds the available GPU memory budget. At B = 32 with 1,600 input and 1,600 output tokens, the KV cache footprint (102,400 tokens) exceeds the available HBM allocation (45,584 tokens reported by vLLM) by more than 2×. In this case, vLLM limits effective Decode concurrency to respect KV cache capacity, leading to run-time stalling and request queueing. This working set size constraint explains the observed Decode throughput saturation and degradation in Section V-B. B. Heterogeneous Prefill/Decode Disaggregation Analysis While heterogeneous Prefill/Decode disaggregation is conceptually appealing, its practical viability depends on the balance between performance gains and system-level overheads, particularly KV cache transfer costs. 1) KV Cache Transfer Overhead: KV cache transfer latency per input token is modeled as: ttransfer =
KV cache size per input token . network bandwidth × bandwidth efficiency
(1)
For Llama2-7B, the KV cache size per input token is 512 KiB, and we assume a bandwidth efficiency of 0.8 throughout this analysis. We consider three representative network configurations: a 25 Gbps Ethernet baseline (ttransfer = 0.210 ms/token); a 100 Gbps-per-GPU setting reflecting realistic Slingshot-based deployments (ttransfer = 0.052 ms/token); and a 1.8 TB/s NVLink configuration, reflecting the bandwidth scale of nextgeneration GPU interconnects, as an optimistic upper bound (ttransfer = 0.00036 ms/token).2 2) End-to-end Latency Analysis: Fig. 4 shows a modelbased end-to-end latency comparison between single-platform inference and heterogeneous Prefill/Decode disaggregation. In the workload range shown, heterogeneous disaggregation (Prefill on A100, Decode on GroqRack) consistently achieves the lowest end-to-end latency for both 25 Gbps Ethernet and 100 Gbps HPE Slingshot. This result reflects the combination of GPU-efficient Prefill and GroqRack’s low per-token Decode latency, with KV cache transfer contributing a secondary effect within this regime. To complement this aggregate view, Fig. 5 presents latency breakdowns for representative workloads with different input– output length ratios. Fig. 5a corresponds to an extreme prefillheavy workload (Lin = 100, Lout = 1), where the Decode phase is too short to amortize the KV cache transfer cost. In this regime, the KV cache transfer overhead outweighs the Decode latency reduction achieved by GroqRack, rendering heterogeneous disaggregation unfavorable. In contrast, Fig. 5b shows that for balanced workloads (Lin = 100, Lout = 100), the transfer overhead constitutes only a small fraction of endto-end latency, and heterogeneous disaggregation consistently outperforms homogeneous baselines. 2 The 1.8 TB/s NVLink bandwidth corresponds to the per-direction specification of next-generation NVIDIA NVLink (6th generation), as announced for the Rubin architecture. While NVLink is not applicable to inter-rack GPU– GroqRack communication, it is used here as a theoretical lower bound on KV cache transfer overhead under extremely high-bandwidth interconnect assumptions.
TABLE IV R EQUIRED OUTPUT– INPUT LENGTH RATIO FOR HETEROGENEOUS P REFILL /D ECODE DISAGGREGATION Interconnect 25 Gbps Ethernet 100 Gbps Slingshot NVLink (1.8 TB/s)†
Heterogeneous wins when Lout 1 ∗ 1 , ] > [ Lin 52 42 Lout 1 ∗ 1 , ] > [ Lin 203 169 Lout 1 1 , ]∗ > [ Lin 29,300 24,400
∗ Ranges reflect variability in the measured per-token Decode gain.
† NVLink is included as a theoretical upper bound and does not represent a feasible GPU–
Fig. 4. End-to-end latency comparison for Llama2-7B across homogeneous (A100-only, GroqRack-only) and heterogeneous (Prefill on A100, Decode on GroqRack) inference strategies with two network configurations: 25 Gbps Ethernet and 100 Gbps HPE Slingshot.
(a) Lin = 100, Lout = 1
(b) Lin = 100, Lout = 100
Fig. 5. Latency breakdown comparison across homogeneous (A100-only, GroqRack-only) and heterogeneous (Prefill on A100, Decode on GroqRack) architectures illustrating workload-dependent performance. KV cache transfer assumes 25 Gbps Ethernet.
3) Heterogeneous Disaggregation Always Outperforms GroqRack-Only Inference: Although both Prefill latency and KV cache transfer overhead scale linearly with Lin , the absolute per-token Prefill latency reduction from offloading Prefill to A100 always exceeds the per-token transfer overhead (prefill gain ≥ 2 ms/token, transfer overhead ≤ 0.21 ms/token). Consequently, heterogeneous disaggregation remains strictly favorable for all input lengths, and no break-even point exists for this comparison. 4) Break-Even Analysis Compared to A100-Only Inference: A non-trivial trade-off arises when comparing heterogeneous disaggregation against A100-only inference. While heterogeneous disaggregation reduces Decode latency by offloading Decode to GroqRack, it incurs additional KV cache transfer overhead. Heterogeneous disaggregation is beneficial when the Decode latency gain amortizes the transfer cost, yielding Lin × ttransfer < Lout × tdecode gain .
(2)
Using experimentally measured Prefill and Decode latencies together with network-dependent transfer cost estimates, we derive break-even thresholds summarized in Table IV.
Heterogeneous Prefill/Decode disaggregation is advantageous for the majority of practical workloads, even when using a 25 Gbps Ethernet link. The only unfavorable regime corresponds to extremely input-heavy workloads, where the input length exceeds the output length by roughly 50× or more. Such cases typically arise in tasks that produce minimal outputs, such as binary decisions or categorical classification based on large input contexts. For most other workloads, heterogeneous disaggregation is preferable. Moreover, with extremely high-bandwidth interconnects such as NVLink (used here only as a theoretical upper bound rather than a feasible GPU–GroqRack deployment scenario), the KV cache transfer overhead becomes negligible for virtually all realistic inference scenarios. C. Implications for System Design The phase-dependent performance asymmetry observed in our experiments has important implications for the design of practical LLM inference systems. 1) Accelerator Specialization versus Versatility: While GroqRack excels at Decode but performs poorly at Prefill, GPUs can efficiently handle both phases. This versatility enables GPUs to serve as overflow capacity for Decode when GroqRack systems reach saturation, preventing request queuing and maintaining system throughput under high load. 2) Multi-Factor Routing Policies: In practice, the breakeven condition in (2) directly informs routing decisions in heterogeneous systems. Given workload characteristics (Lin , Lout ) and current system conditions (accelerator utilization, network congestion), requests can be dynamically routed to heterogeneous disaggregation or single-platform execution depending on whether the expected Decode latency reduction amortizes the KV cache transfer overhead. 3) Challenges in Efficient KV Cache Transfer: Although KV cache transfer overhead is often amortizable in theory, efficiently enabling low-latency KV cache exchange across heterogeneous accelerators remains an open system challenge. In contrast to homogeneous GPU deployments that can leverage high-bandwidth interconnects (e.g., NVLink) and emerging point-to-point communication libraries such as NIXL, no standardized mechanism currently exists for efficient KV cache transfer between GPUs and non-GPU accelerators.
D. Generalizability and Limitations Our evaluation targets Llama2-7B in FP16 under controlled synthetic workloads. While absolute latency values are modelspecific, the qualitative trade-offs persist for transformer-based models: Prefill remains compute-bound and favors parallelism; Decode remains memory-bound and benefits from low-latency memory access. Per-token latency and KV cache size scale linearly with model dimension, preserving the structural trade-off between Decode gains and transfer costs, though quantitative break-even points shift with model size. Our analysis assumes idealized KV cache transfer based on sustained bandwidth. Production systems face additional overheads from congestion, protocols, and multi-tenancy. We do not consider economic factors. Despite these limitations, the core insight remains: phase-dependent performance asymmetry necessitates separate evaluation of Prefill and Decode when comparing heterogeneous platforms. Despite these limitations, the central insights remain robust: Prefill and Decode exhibit fundamentally different computational characteristics; accelerators demonstrate phasedependent performance asymmetry; and the viability of heterogeneous disaggregation depends on balancing phase-specific gains against transfer overheads. These principles provide a general framework for evaluating heterogeneous LLM inference systems beyond the specific configuration studied here. VII. C ONCLUSION This paper presents a phase-aware evaluation methodology for LLM inference that isolates Prefill and Decode performance across GPUs and emerging AI accelerators. Using Llama2-7B as a benchmark model, our systematic evaluation reveals fundamental phase-dependent performance asymmetry: GPUs achieve up to 39× lower latency and 41× higher throughput in the compute-intensive Prefill phase through massive parallelism, while GroqRack demonstrates 4× lower per-token latency in the memory-bound Decode phase through deterministic execution and large on-chip SRAM. However, GPUs recover throughput advantages as batch size increases due to effective parallel processing across independent requests. These findings demonstrate that the platforms exhibit complementary strengths across inference phases and system objectives, that aggregate end-to-end metrics can obscure critical phase-specific trade-offs, and that optimal accelerator selection depends on workload characteristics as well as whether latency or throughput is prioritized. Building on these observations, our heterogeneous Prefill/Decode disaggregation analysis suggests that combining GPU-based Prefill with accelerator-based Decode can achieve lower end-to-end latency than homogeneous approaches when output length is sufficient to amortize KV cache transfer costs. Although this analysis is based on analytical and hypothetical estimates rather than a fully integrated system, the results indicate that heterogeneous Prefill/Decode disaggregation is a structurally promising design point rather than a marginal optimization.
Realizing such heterogeneous inference architectures in practice will require advances beyond individual accelerators. In particular, efficient and predictable GPU–accelerator communication becomes a key enabler, encompassing highbandwidth physical interconnects as well as software stacks that expose phase boundaries, manage KV cache placement and transfer, and coordinate workload-aware scheduling across devices. As LLM workloads continue to scale in both model size and context length, we believe that co-designing hardware links, runtime systems, and orchestration layers for crossdevice execution will be increasingly important for enabling efficient, scalable, and cost-effective LLM inference. ACKNOWLEDGMENT This research used resources of the Argonne Leadership Computing Facility, which is a U.S. Department of Energy Office of Science User Facility operated under contract DEAC02-06CH11357. R EFERENCES [1] K. T. Chitty-Venkata, S. Mittal, M. Emani, V. Vishwanath, and A. K. Somani, “A survey of techniques for optimizing transformer inference,” Journal of Systems Architecture, p. 102990, 2023. [2] Z. Zhou, X. Ning, K. Hong, T. Fu, J. Xu, S. Li, Y. Lou, L. Wang, Z. Yuan, X. Li et al., “A survey on efficient inference for large language models,” arXiv preprint arXiv:2404.14294, 2024. [3] Y. Zhong, S. Liu, J. Chen, J. Hu, Y. Zhu, X. Liu, X. Jin, and H. Zhang, “{DistServe}: Disaggregating prefill and decoding for goodput-optimized large language model serving,” in 18th USENIX Symposium on Operating Systems Design and Implementation (OSDI 24), 2024, pp. 193–210. [4] V. J. Reddi, C. Cheng, D. Kanter, P. Mattson, G. Schmuelling, C.-J. Wu, B. Anderson, M. Breughe, M. Charlebois, W. Chou et al., “Mlperf inference benchmark,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 2020, pp. 446– 459. [5] W. Kwon, Z. Li, S. Zhuang, Y. Sheng, L. Zheng, C. H. Yu, J. Gonzalez, H. Zhang, and I. Stoica, “Efficient memory management for large language model serving with pagedattention,” in Proceedings of the 29th symposium on operating systems principles, 2023, pp. 611–626. [6] K. T. Chitty-Venkata, S. Raskar, B. Kale, F. Ferdaus, A. Tanikanti, K. Raffenetti, V. Taylor, M. Emani, and V. Vishwanath, “Llm-inferencebench: Inference benchmarking of large language models on ai accelerators,” in SC24-W: Workshops of the International Conference for High Performance Computing, Networking, Storage and Analysis. IEEE, 2024, pp. 1362–1379. [7] D. Abts, J. Ross, J. Sparling, M. Wong-VanHaren, M. Baker, T. Hawkins, A. Bell, J. Thompson, T. Kahsai, G. Kimmell et al., “Think fast: A tensor streaming processor (tsp) for accelerating deep learning workloads,” in 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA). IEEE, 2020, pp. 145–158. [8] D. Abts, G. Kimmell, A. Ling, J. Kim, M. Boyd, A. Bitar, S. Parmar, I. Ahmed, R. DiCecco, D. Han et al., “A software-defined tensor streaming multiprocessor for large-scale machine learning,” in Proceedings of the 49th Annual International Symposium on Computer Architecture, 2022, pp. 567–580. [9] H. Touvron, L. Martin, K. Stone, P. Albert, A. Almahairi, Y. Babaei, N. Bashlykov, S. Batra, P. Bhargava, S. Bhosale et al., “Llama 2: Open foundation and fine-tuned chat models,” arXiv preprint arXiv:2307.09288, 2023.