ConceptioArchivearXiv CS
arXiv CSopen access

EdgeFlow: Fast Cold Starts for LLMs on Mobile Devices

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

arXiv:2604.09083v1 [cs.OS] 10 Apr 2026

EdgeFlow: Fast Cold Starts for LLMs on Mobile Devices Yongsheng Yan

Jiacheng Shen

College of Computer Science and Artificial Intelligence, Fudan University China [email protected]

Duke Kunshan University China [email protected]

Xuchuan Luo

Yangfan Zhou

College of Computer Science and Artificial Intelligence, Fudan University China [email protected]

College of Computer Science and Artificial Intelligence, Fudan University China [email protected]

Abstract

critical performance bottleneck for mobile applications [28, 30, 33–35, 46, 49]. It happens when an application is invoked, but the required data are not loaded into device memory. This issue is exacerbated in mobile LLM applications, e.g., conversational agents and UI automation tasks [54, 69], where a huge amount of model weights have to be loaded into memory. Our preliminary experiments show that llm.npu [58], the state-of-the-art mobile LLM inference framework, suffers from a 9.1-second time to first token (TTFT) during a cold start of executing the Llama3 8B [3] model under 128 input tokens. Such latency is an order of magnitude greater than the cold start of conventional mobile applications [36, 56], which significantly degrade user experiences [9, 44]. Even worse, cold starts of LLM frameworks frequently occur since mobile operating systems can not always keep the massive model parameters in limited device memory [7, 8, 23]. In this paper, we identify the primary cause of long coldstart latency as suboptimal utilization of flash bandwidth when loading LLMs. Specifically, existing LLM inference frameworks typically quantize model weights into lower precisions to reduce model sizes [18, 55]. However, conventional quantization techniques uniformly quantize all weights into the same precision, e.g., INT8, to align with hardware-supported data types. Such an approach overlooks that different weights have different contributions to model accuracy [37]. Flash bandwidths are hence wasted on transferring unimportant weights that are under-quantized. To address this problem, we propose to employ adaptive quantization to achieve a better trade-off between model accuracy and flash bandwidth utilization. The key idea is to determine and assign the required precision of each weight based on its importance to model accuracy. Nevertheless, this approach presents several challenges that call for a co-design of inference systems and quantization algorithms. (1) Precision assignments under NPU constraints. Existing mobile LLM inference frameworks leverage NPUs to accelerate inference efficiency [13, 21, 58, 59]. Production

Deploying large language models (LLMs) on mobile devices is an emerging trend to enable data privacy and offline accessibility of LLM applications. Modern mobile neural processing units (NPUs) make such deployment increasingly feasible. However, existing mobile LLM inference frameworks suffer from high start-up latency due to their inevitable cold starts, i.e., launching LLM inferences when the model is not hosted in device memory. In this paper, we identify the key bottleneck of mobile LLM cold starts as the waste of flash bandwidth on unimportant model parameters. We design EdgeFlow, a mobile LLM inference framework that mitigates the cold start issue by adaptively adjusting the precisions of LLM parameters. Specifically, EdgeFlow leverages 1) an NPU-aware adaptive quantization algorithm that assigns different precisions to weights in a finer granularity according to their importance and NPU constraints, 2) an SIMDfriendly packing format that accelerates the transformation of various-precision weights into fixed-sized NPU-native data types, and 3) a synergistic granular pipeline that coordinates CPU and NPU computation in a fine-grained and dynamic manner. Experimental results show that EdgeFlow reduces cold-start latency by up to 4.07× compared with three stateof-the-art mobile LLM inference frameworks, i.e., llama.cpp, MNN, and llm.npu, under comparable model accuracy.

1

Introduction

With the increasing adoption of large language models (LLMs) in everyday applications, e.g., personal assistants [50, 51, 66], smart home control [12, 29], and healthcare [15, 17], there is a growing trend towards deploying LLM inference directly on mobile devices to ensure data privacy and offline accessibility [20, 31, 41]. Many LLM inference frameworks are proposed to improve the execution efficiency of LLMs on resource-constrained mobile devices [13, 52, 58]. However, the cold start issue is largely overlooked in existing mobile LLM inference frameworks. Cold start is a 1

Table 1. Existing post-training quantization approaches.

mobile NPUs, e.g., Qualcomm Hexagon NPU [27], only accelerate matrix multiplications (matmuls) for tensors under uniform and coarse-grained quantization. Such constraints impede the adaptive assignment of precisions to individual weights and hinder the execution of existing importanceaware quantization methods that adopt non-uniform, finegrained quantization on CPUs and GPUs [14, 37]. (2) Costly unpacking for variable-precision weights. Low-bit quantized model weights have to be unpacked into NPU-native data types, e.g., INT8 and FP16, by the CPU before computation. Existing approaches store quantized models in an unpack-friendly format to reduce the computational overhead during unpacking [20, 37]. However, these formats are designed for weights quantized with the same precision. Naively storing variable-precision weights in these formats incurs significant overhead due to the excessive memory accesses during unpacking computation. (3) Imbalanced workloads between CPU and NPU. Unpacking weights and conducting prefill computation require the collaboration between the CPU and NPU. Existing approaches schedule computational tasks to the two processors in a coarse-grained and static manner [58, 59]. Unfortunately, such scheduling results in significant bubbles in the computation pipeline. This would make the long computation time another bottleneck for cold-starts. We design EdgeFlow, a system that co-designs a quantization algorithm with system-level optimizations to reduce cold-start latencies for mobile LLM inferences. First, we design an NPU-aware adaptive quantization algorithm to differentiate the importance of weights and assign fine-grained precisions under NPU constraints, enabling fast model loading with comparable accuracy. Second, we develop an SIMDfriendly packing format with an SIMD-optimized unpacking algorithm to facilitate rapid transformation of adaptively quantized weights into NPU-native data types. Finally, we propose a synergistic granular pipeline that balances CPU and NPU computation in a dynamic, fine-grained manner to maximize compute resource utilization during cold starts. We implement EdgeFlow from scratch and evaluate it with various LLMs [1–4] and datasets [22, 42, 43, 45, 47, 64]. We compare EdgeFlow with two open-source approaches, i.e., llama.cpp [20] and MNN [52], and a state-of-the-art NPUaccelerated mobile LLM inference framework, i.e., llm.npu [58]. Our evaluation results show that with the same model accuracy, EdgeFlow reduces cold-start latency, i.e., TTFT, by 4.07× and 2.40× compared with llama.cpp and MNN. Compared with llm.npu, EdgeFlow achieves 1.55× speedup in TTFT and up to 2.52% improvements in accuracy. We plan to open source EdgeFlow in the near future. Our contributions are summarized as follows: • We break down the cold-start issue of mobile LLM deployments, and show that the inefficient flash bandwidth utilization significantly deteriorates TTFT.

Dimensions

Categories

Target Timing Mapping Type Granularity

Weight Tensors / Activation Tensors Static / Dynamic Uniform / Non-uniform Per-tensor / Per-channel / Per-block

• We propose EdgeFlow that co-designs the NPU-aware adaptive quantization algorithm with the SIMD-friendly packing format and the synergistic granular pipeline to collaboratively reduce cold-start latencies. • We implement EdgeFlow and show that it significantly reduces cold-start latency while preserving inference accuracy under multiple models and workloads.

2

Background

2.1

Mobile NPUs for LLM Inference

NPUs are widely adopted in smartphones due to their energyefficient computing capabilities under AI workloads, e.g., Qualcomm Hexagon NPU [27] and Apple Neural Engine [25]. In this paper, we focus on system design with Qualcomm Hexagon NPU, which is one of the most widely adopted NPUs on Android phones. The proposed techniques can also be adapted to other platforms with engineering efforts. Hexagon NPU provides developers with QNN SDK [26] to leverage various hardware accelerators inside NPUs. QNN offers a toolchain to quantize and compile models into NPU executables, and a runtime to efficiently execute models over NPUs. Specifically, to deploy a model on NPU, users first need to assemble a computation graph with high-level operators provided by QNN. The computation graph is then compiled into an execution graph, i.e., NPU executables, before being loaded into NPU with QNN runtime APIs. During the whole process, the computation graph compilation is the most timeconsuming step since QNN needs to conduct a series of NPUspecific optimizations, e.g., fusing operators, scheduling tiles and loops, and reordering weights according to the memory layout of NPUs [21, 58]. 2.2

Post-Training Quantization

Post-training quantization is widely adopted in mobile LLM deployments to reduce model sizes and runtime memory footprint. The key idea is to map tensors into low-precision data types with some metadata and dequantize, i.e., map back to high-precision tensors, during or after computation to maintain model accuracy. As shown in Table 1, existing quantization techniques can be classified in terms of quantization target, timing, mapping types, and granularity. In terms of target, there are two types of tensors during LLM inference, i.e., weight tensors storing static model parameters and activation tensors storing intermediate results, e.g., KV 2

(a) Online construction (llm.npu) Prepare Graph Setup Load Weights Compose Graph Optimize Graph Envs 0.3s

4.5s

4s

(b) Materialization Setup Load Graph Envs 0.3s

6.8s

on a Xiaomi 15 Pro with Hexagon NPU, 16 GB of RAM, and 512 GB of flash storage.

Execute Graph

41s

1.4s dump graph here

Prepare Graph

Execute Graph

2s

1.4s

3.1

We use llm.npu [58], the state-of-the-art NPU-based mobile LLM inference framework, to showcase the cold start process of NPU-based mobile LLM inference. As shown in Figure 1a, the cold start process of llm.npu can be summarized into 4 steps, i.e., setup environments, load weights, prepare execution graphs, and execute. According to our experiments on Llama3 8B with a 128-token prompt, llm.npu exhibits a cold-start latency of up to 51.2 seconds. Its cold-start latency is mainly dominated by the overhead of on-the-fly graph preparation and weight loading. Two methods can be adopted to reduce the pure coldstart latency, i.e., materialization and overlap. First, materialization stores the optimized computation graph offline and directly loads the graph during cold starts. As shown in Figure 1b, such an approach eliminates the graph optimization overhead from the critical path and reduces cold-start latency from 51.2s to 10.5s. Second, we can overlap graph loading with prefill computation between consecutive layers by decomposing the graph into multiple subgraphs, one for each layer. Figure 1c illustrates that the overlapping method further reduces the cold-start latency to 9.1s. Unfortunately, even with these optimizations, the latency remains unacceptable for interactive mobile applications, as users typically lose patience if the wait time exceeds 7s [9]. The key bottleneck of existing approaches lies in the weight loading overhead, as indicated by the 7.4s pipeline bubble. Many mobile systems prefetch application data from flash to reduce data loading latency during cold starts [46, 49, 56, 60]. However, as shown in Figure 2, 4 GB of data must be prefetched and reserved in memory to achieve a satisfactory cold-start latency. The consumed memory accounts for 25% of total memory in our high-end testbed, which is a significant overhead for resource-constrained mobile devices. In contrast, EdgeFlow focuses on reducing the pure cold start latency, which can meet the latency requirement without reserving memory in advance. To achieve this, we propose an adaptive quantization strategy to relieve the flash bandwidth bottleneck on the critical path of cold starts. Our opportunity is that not all weights are equally important to the model accuracy [14, 32, 37]. Less precision, i.e., bits, can be assigned to unimportant weights to optimize flash bandwidth utilization.

(c) Overlapping Setup Envs 0.3s

Load & Prepare Graph Execute Bubble Graph 7.4s

1.4s

TTFT (s)

Figure 1. Breakdown of the cold-start latencies of llm.npu and two straightforward optimizations, i.e., materialization and overlapping.

10 8 6 4 0

llm.npu

EdgeFlow

Acceptable Latency 1

2 3 4 Pre-loaded Data Size (GB)

5

6

Figure 2. TTFT v.s. pre-loaded data size for llm.npu, including user-acceptable latency region and EdgeFlow measurements.

caches. Both types of tensors are considered as targets for quantization. In terms of timing, static approaches quantize tensors offline, while dynamic approaches quantize them during inference. In terms of mapping types, uniform approaches quantize tensors with linear transformations, i.e., one high-precision value in a tensor can be mapped to a low-precision one with a scale and a bias. Uniform quantization can be further classified into symmetric and asymmetric methods based on whether the bias is 0 or not. Non-uniform approaches quantize tensors with non-linear transformations, e.g., codebooks [39]. Finally, the granularity refers to how many weights share the same mapping metadata. Per-tensor approaches share metadata for an entire tensor, per-channel ones share metadata for a row or a column, and per-block ones share metadata among fixed-size groups. Existing quantization schemes quantize tensors into a single precision [10, 18, 37] or a limited group of precisions, e.g., INT4 and INT8 [14, 32]. This results in inefficient bandwidth utilization during cold starts. The essential reason is that mobile accelerators only contain arithmetic units to compute tensors with a limited set of data types, e.g., INT8, INT16, and INT32. Mixed-precision tensors have to be transformed into tensors with native data types before computation, which incurs additional transformation overhead.

3.2

3

Analysis of the Cold-Start Issue

Motivation and Challenges

Challenges of Adaptive Quantization

In this section, we introduce three challenges when applying adaptive quantization in a mobile LLM inference system. Challenge 1: Precision assignments under NPU constraints. NPUs only accelerate a limited set of quantized matmuls, which imposes two constraints on quantization algorithms. First, all tensors, i.e., weight and activation tensors,

This section first breaks down the cold-start process with both theoretical and experimental analyses, motivating the need to adopt adaptive quantization (§ 3.1). We then discuss the challenges of integrating adaptive quantization to mobile LLM inference systems (§ 3.2). All experiments are conducted 3

Method

Target

Mapping Type

Granularity

AWQ [37] CMPQ [14]

Weights Weights

Uniform (Asym.) Non-uniform

Per-block Per-channel (in)

Weights NPU Uniform (Sym.) Constraints Activations

Per-channel (out) Per-tensor

[128,14336] x [14336,4096]

Percentage

Table 2. Comparison between NPU constraints and existing importance-aware quantization approaches.

50 100 Latency (ms)

150

0

50 100 Latency (ms)

0.00% 0.04% 0% 0.00% 1 2 3

Load 0 0 2 4 Load 1 1 3 5 Raw INT8 Weights

11.54% 4 5 Bits

9.91% 0.79% 0.04% 6 7 8

... ...

29 30 28 31 3.03s 0.57s

Unpack ... (intermediate tasks omitted) ... 28 30 Load 0 0 2 4 Loading ... 29 31 Load 1 1 3 5 Saved 2.90s INT4/INT8 Mixed Format Loading Saved Unpack 0 1 2 ... 3031 5.53s ... 28 30 Load 0 0 2 4 Unpacking Overhead ... 2931 Load 1 1 3 5 1.60s K-Quant Format

Dequant Matmul 0

77.67%

50%

(a) The precision distribution for loading and unpacking.

[128,4096] x [4096,14336]

CMPQ AWQ INT8 Opt INT8

100%

150

(b) The loading and unpacking overhead.

Figure 3. The execution times of NPU matmuls on tensors quantized by AWQ and CMPQ compared with the pure INT8 matmul (INT8) operator and optimized INT8 matmul (Opt INT8) operator.

Figure 4. The weight loading and unpacking times of three weight packing schemes for Llama3 8B.

have to be quantized in a static, uniform, and symmetric manner. Second, activations can only be quantized in per-tensor granularity, while weights can be quantized in per-channel granularity only on output channels, i.e., columns. Such constraints result in poor matmul efficiency for tensors quantized by existing importance-aware quantization approaches since they often require more fine-grained quantization. We evaluate the execution latency of matmul operators on tensors quantized by AWQ [37] and CMPQ [14], two stateof-the-art importance-aware quantization algorithms. We measure the NPU execution time of matmul on two representative tensor shapes in the inference process of Llama3 8B. As shown in Table 2, both approaches only quantize weight tenors. AWQ employs asymmetric per-block quantization to quantize all weights into INT4. CMPQ adopts non-uniform per-channel quantization on input channels and quantizes weights to the range of 2 to 4 bits with an average 3-bit precision. However, neither approach can be directly accelerated by the NPU due to the incompatible mapping type and quantization granularity. When being executed on NPU, quantized weights need to be dynamically dequantized into INT8 with extra arithmetic operations. As shown in Figure 3, AWQ and CMPQ incur 2.59× and 2.63× overhead compared with INT8 matmul on NPUs. Moreover, during the computation graph optimization phase, the QNN toolchain optimizes the storage layout for static weight tensors to further improve execution efficiency. The static graph optimization approach cannot be applied to accelerate the two approaches since their tensors must be dynamically dequantized. As a result, compared with the statically optimized INT8 matmul (Opt INT8), an additional 2.66× slowdown is introduced. Challenge 2: Costly unpacking for variable-precision weights. NPUs only support a few native data types, e.g., INT8 and FP16, which can be executed directly by their tensor accelerators [26]. When weights are quantized into

various precisions, they must first be unpacked into one of these native types, incurring extra computation overhead. There are two ways to pack and unpack weight tensors, both of which exhibit a trade-off between read amplification and unpacking computation. First, we can adopt an INT4/INT8 mixed format to pack data by fitting and padding mixed-precision weights into INT4 and INT8 data types. The unpacking overhead of this format is minimal since weights are already aligned. Nevertheless, such an approach incurs substantial read amplifications, as padding causes the system to fetch much more data from flash. The second approach is adopting the K-Quant [20] format in different weight groups. K-Quant designs a format for storing single-precision weight tensors with unaligned precisions, e.g., INT3. It decouples the weights into 1-, 2-, and 4-bit components and stores them separately in a compact file. While such an approach eliminates read amplification, it introduces significant unpacking overhead. We measure the overhead of the two approaches for unpacking Llama3 8B quantized by our adaptive precision allocation algorithm. Figure 4a shows the precision distribution of the quantized model, and Figure 4b illustrates their corresponding unpacking overheads. We use two threads to load the model in a layer-by-layer manner and one thread to unpack the loaded weights. Each block in the figure represents the loading time or the unpacking time of a single layer. A number inside a block denotes the layer being loaded or unpacked. The INT4/INT8 mixed format reduces model loading from 3.03s to 2.90s compared with naively padding all weights into INT8. K-Quant decreases loading latency to 1.60s due to its compact storage. However, the unpacking computation takes 5.53s, which becomes a new bottleneck on the critical path of cold starts. 4

CPU

Challenge 3: Imbalance workloads between the CPU and the NPU. Prefill computation constitutes another bottleneck on the critical path of cold starts. The main challenge is to balance the computation workload between the CPU and NPU while keeping minimal idle time. Nevertheless, due to the heterogeneous nature of operators during the LLM inference computation and their varying computational characteristics, achieving such a balance is challenging. The state-of-the-art approach, i.e., llm.npu [58], partitions the LLM computation graph in a coarse-grained and static manner among the NPU and the CPU. We evaluate llm.npu with a 256-token prefill on Llama3 8B for all subgraphs and measure its CPU and NPU computation time, as shown in Figure 5. llm.npu adopts chunked prefill [5] and parallelizes the computation of two adjacent chunks on the CPU and NPU in the granularity of statically partitioned subgraphs. Only attention computations are placed on the CPU side, while the remaining operators between layers run on the NPU. As shown in Figure 5a, such a design leaves substantial pipeline bubbles for two main reasons. First, coarse-grained partitioning assigns many NPU-inefficient operators to the NPU, i.e., operators with low arithmetic intensity, which causes high NPU execution time. Figure 5b shows the CPU and NPU execution time of three representative operators that llm.npu assigns to the NPU, i.e., RMSNorm, SwiGLU, and quantization. The NPU execution time is on average 2.1× longer than that on the CPU. Second, static scheduling prevents dynamic adjustment based on runtime status. Figure 5c illustrates the CPU and NPU execution time and the bubble rate of the entire pipeline across different sequence lengths. The bubble rate is calculated as the proportion of idle time to the execution time. The NPU execution time is consistently longer than the CPU execution time, resulting in up to 91% bubble rates across the pipeline. Moreover, this imbalance varies with the prompt length, making static scheduling incapable of maintaining balanced utilization.

4

NPU

Bubble Rate

chunk1 38.3ms Bubble

21.9ms

Bubble

43.1ms Bubble

chunk0 19.8ms

40.7ms

Gate/Up SwiGLU

15.7ms

44.6ms

12.0ms

Down

Add

Norm

Layer3 FFN

Bubble

Q/K/V

Quant

Layer4 Q/K/V proj

100 Latency (s)

RMSNorm Quantize SwiGLU 0

2 4 Latency (ms)

6

12

75

8

50

4

25

0

256

512 1024 Sequence Length

0

Bubble Rate (%)

(a) A fragment of the measured pipeline of llm.npu.

(b) Execution times of NPU- (c) The prefill computation time of the NPU inefficient operators on the and CPU, respectively, and the bubble rate NPU and CPU, respectively. of the entire pipeline.

Figure 5. Analyses of the pipeline of llm.npu during the prefill stage of Llama3 8B. NPU

CPU

NPU Address Space

Task Steal

Async Restore

DRAM

Metadata

Weight (INT8)

...

NPU Graph

Apply Priority Apply Priority

Synergistic Granular Pipeline (§4.3) Executor

Async Load & SIMD Unpack

Flash

SIMD-Friendly Packed Weight Packing Format (§4.2)

Offline Pack NPU-Aware Adaptive Quantization (§4.1) Quantize

Offline

Bit-Width Allocation Relative Error

Smooth

Weight (FP)

Datasets Profile

Figure 6. The overview of EdgeFlow.

NPU-native formats layer by layer. An SIMD-friendly packing format is proposed to store weights, and an SIMD-based unpacking algorithm is adopted to accelerate the unpacking process (Challenge 2). Moreover, a synergistic granular pipeline is employed during prefill computation to balance workloads between the CPU and the NPU and alleviate the computation bottleneck of cold starts (Challenge 3).

The EdgeFlow Design

We design EdgeFlow to address all of the above challenges and reduce the cold-start latency of mobile LLM inferences. As illustrated in Figure 6, EdgeFlow consists of two phases, i.e., an offline phase and an online phase. The offline phase performs weight quantization to accelerate model loading on the critical path of cold starts. We introduce an NPU-aware adaptive quantization method to achieve a better trade-off between the loading time and the model accuracy under NPU constraints (Challenge 1). The online phase is responsible for preparing the execution graph and conducting the prefill computation. To accelerate the process, data loading, weight unpacking, and prefill computation are executed in parallel. The quantized model is loaded asynchronously from flash and unpacked to

4.1

NPU-Aware Adaptive Quantization

As mentioned in Section 3.2, NPUs only accelerate matmuls for per-output-channel quantized weight tensors and coarsegrained per-tensor quantized activation tensors. Our quantization scheme complies with these restrictions and achieves a better trade-off between model size and accuracy with two design choices: 1) We assign different precisions in the granularity of output channels on weight tensors to satisfy the restrictions of NPUs on weights. Each output channel 5

Greedy Bit-Width Allocation. For a weight tensor with 𝐶 output channels {𝑊𝑖 }𝐶𝑖=1 = {𝑊1,𝑊2, ...,𝑊𝑐 }, we need to find the optimal per-channel bit-widths {𝐵𝑖 }𝐶𝑖=1 = {𝐵 1, 𝐵 2, ..., 𝐵𝑐 } that minimizes the total RE. Two constraints need to be satisfied. First, the assignable bit-widths for each channel should be in the range 1 to 8 bits. Second, the average bitwidth of all channels should not exceed a given value 𝐵𝑒 to indicate the flash bandwidth budget. We formulate the optimization problem as follows:

is quantized into integers with different bit-widths. A relative error metric is defined to guide the precision assignment on each channel. In addition, a greedy bit-width allocation algorithm is proposed to derive an optimal assignment of precisions across all channels. 2) We transfer the quantization difficulty of both input and output activations to weights to respect the NPUs’ restrictions on activation tensors. An NPU-aware smoothing scheme is adopted to achieve this goal without additional runtime overhead. The Relative Error Metric. We design a relative error metric to estimate the amount of quantization error a weight channel introduces when quantized to a specific bit-width. Channels with higher metric values are granted more bits. Formally, for channel 𝑖 with 𝐷 dimensions, we define the relative error under 𝐵-bit quantization, i.e., RE(𝑊𝑖 , 𝐵), as the cosine distance between the original weight channel 𝑊𝑖 and 𝑞 its dequantized version 𝑊𝑖 [11, 68]:

min 𝐶 {𝐵𝑖 }𝑖=1

s.t.

.

∥𝑊𝑖 ∥ ∥𝑊𝑖 ∥

However, computing the cosine distance for all channels under all candidate bit-widths is expensive, which results in high quantization overheads. Particularly, before computing 𝑞 the cosine distance, we first need to get 𝑊𝑖 by uniformly quantizing 𝑊𝑖 into 𝐵 bits and dequantizing it back to FP16. Two linear transformations and one cosine distance computation are involved in the process of calculating RE for one channel under a single bit-width. We reduce the computation overhead by modeling RE with random variables so that it can be efficiently estimated with simple statistics over 𝑊𝑖 and 𝐵 instead of calculating the precise values. To achieve this, we approximate RE with the absolute 𝑞 error 𝐸 = 𝑊𝑖 − 𝑊𝑖 and model 𝐸 as a random variable. We skip the detailed derivation in Appendix A. The RE can be represented as a function of the mathematical expectations of 𝐸 2 and 𝑊𝑖 2 : RE(𝑊𝑖 , 𝐵) = 1 −

RE(𝑊𝑖 , 𝐵𝑖 ), 𝐵𝑖 ≤ 𝐶 · 𝐵𝑒 , 𝐵𝑖 ∈ {1, 2, . . . , 8}.

We leverage two properties of this optimization problem to design our bit-width allocation algorithm. First, the total RE in our formulation is additive across channels, meaning the RE of one channel is determined solely by its own bitwidth and statistics, independent of others. While end-to-end model accuracy is indeed a complex combinatorial function of all quantizations, we use this additive RE as a local proxy metric to make the optimization tractable. Second, the reduction in RE for a channel is strictly decreasing with more bit-width 𝐵 assigned to it. These two properties guarantee that we can greedily assign a bit to one channel that reduces the total RE most without sacrificing the global optimality. Therefore, we develop a greedy algorithm that yields the optimal solution. As shown in Algorithm 1, we first initialize all channels to 1 bit and then compute the remaining bit budget (Line 1). We build a max heap using the RE gain of adjacent bit-widths for each channel as key (Line 2), i.e., RE(𝑊𝑖 , 𝐵) − RE(𝑊𝑖 , 𝐵 + 1). We then iteratively pop the channel that reduces its relative error the most from the heap, increment its bit-width by one, update the heap, and adjust the remaining bit-width budget (Lines 3-7). The process continues until the total bit-width budget is exhausted. This approach efficiently finds an optimal  bit-width allocation for 𝐶 weight channels in 𝑂 𝐶 · log 𝐶 time. NPU-Aware Smoothing. NPUs’ coarse-grained per-tensor quantization degrades accuracy severely on high-variance LLM activations [38, 58]. Meanwhile, our bit-width allocation algorithm on weights is input-unaware, making it insufficient to address activation-induced errors. To tackle these issues, we propose NPU-aware smoothing, which statically transfers inter-channel variance from activations to weights without runtime overhead. We use a calibration dataset to profile per-channel variance 𝑆𝐼 ∈ R𝐷 and 𝑆𝑂 ∈ R𝐶 for the input 𝐼 and output 𝑂, respectively. The variance for each channel is defined as the maximum absolute value in the channel. We then divide inputs and outputs by 𝑆𝐼 𝛼 and 𝑆𝑂 𝛽 so that their variances are smoothed. 𝛼 and 𝛽 are hyperparameters controlling the

𝑊𝑖 · 𝑊𝑖

𝑞

𝑖=1 𝐶 ∑︁ 𝑖=1

𝑞

RE(𝑊𝑖 , 𝐵) = 1 −

𝐶 ∑︁

𝑊𝑖 · (𝑊𝑖 − 𝐸) E[𝐸 2 ] ≈ . ∥𝑊𝑖 ∥ · ∥𝑊𝑖 − 𝐸 ∥ 2 E[𝑊𝑖 2 ]

We further adopt a uniform rounding approximation [37] to simplify the calculation of E[𝐸 2 ]. Formally, we model 𝐸 𝑗 , i.e., the absolute error of the 𝑗-th element in 𝑊𝑖 , with 𝐸 𝑗 ∼ U( −𝑆2 𝑖 , 𝑆2𝑖 ), where 𝑆𝑖 is the scale of channel 𝑖 under uniform quantization. The rationale is that the absolute error always falls in the range [ −𝑆2 𝑖 , 𝑆2𝑖 ] and 𝑆𝑖 is usually small. We can use a uniform distribution to approximate the real distribution in a small range. By applying the uniform estimation to E[𝐸 2 ], we obtain the final form: 1 (max |𝑊𝑖 |) 2 RE(𝑊𝑖 , 𝐵) = 2𝐵 · . 2 E[𝑊𝑖 2 ] The advantage of this expression is that it can be calculated with simple statistics, i.e., mean squared values, which greatly simplifies the computation. 6

Load

Algorithm 1: Greedy Bit-Width Allocation

1-bit Block

Input

: {𝑊𝑖 }𝐶𝑖=1 : The weight with 𝐶 channels 𝑏𝑢𝑑𝑔𝑒𝑡: The expected average bit-width Output : {𝐵𝑖 }𝐶𝑖=1 : The allocated bit-width array // Initialize all channels to 1 bit and calculate expected gains 1 𝐹𝑖𝑙𝑙𝑂𝑛𝑒𝑠 (𝐵), 𝑟𝑒𝑚𝑎𝑖𝑛 ← 𝑏𝑢𝑑𝑔𝑒𝑡 − 1 𝐶 2 H ← 𝑀𝑎𝑥𝐻𝑒𝑎𝑝 ({ RE(𝑊𝑖 , 1) − RE(𝑊𝑖 , 2)}𝑖=1 ) 3 while 𝑟𝑒𝑚𝑎𝑖𝑛 > 0 do // Greedily allocate 1 more bit to the max-gain channel 𝑗 4 𝑔𝑎𝑖𝑛 𝑗 ← H .𝑝𝑜𝑝 () 5 𝐵 𝑗 ← 𝐵 𝑗 + 1, 𝑟𝑒𝑚𝑎𝑖𝑛 ← 𝑟𝑒𝑚𝑎𝑖𝑛 − 𝐶1 6 if 𝐵 𝑗 < 8 then // Update the gain of channel 𝑗 7 H .𝑝𝑢𝑠ℎ(RE(𝑊 𝑗 , 𝐵 𝑗 ) − RE(𝑊 𝑗 , 𝐵 𝑗 + 1)) 8

① Mask non-target bits ② Align to target bits ③ Merge all together repeat 8x

Unpacked Group Stripe0 Stripe1 Stripe2

Norm Scale Tensor1 Tensor2 ... TensorT Metadata

Stripe7

Logically Consecutive 0 16 32 ... 112 1 17 33 ... 126 15 31 63 ... 127

Block (128 Bits)

1-bit Weightlet

0 16 32 48 1 17 33 49 2 18 34 50 .... 57 63

Weight Tensor1 D

W1 W2 W3 W4

...

WC

Block

2-bit Weightlet

0 16 1 17 2 18 3 19 4 20 5 21 .... 15 31

Block

4-bit Weightlet

Group (128 Weights) C

Figure 7. The SIMD-friendly packing format. The number denoted in each weightlet is the index of its corresponding weight.

degree of smoothing. The variances are then restored in a transformed weight tensor 𝑊 ′ = diag(𝑆𝐼 𝛼 ) · 𝑊 · diag(𝑆𝑂 −𝛽 ). The resulting quantized matmul becomes: 𝑂 = (𝐼 · diag(𝑆𝐼 −𝛼 )) · 𝑊 ′ · diag(𝑆𝑂 𝛽 ) Performing bit-width allocation on the restored weight tensor 𝑊 ′ is implicitly guided by activation statistics, making the overall scheme input-aware, similar in spirit to AWQ [37]. We perform a grid search to find the optimal 𝛼 over the interval of [0, 1] ∈ R that minimizes quantization error. We set 𝛽 to 1 to shift most of the inter-channel variance of outputs to weights. Moreover, the scaling factors do not introduce additional runtime overhead, as the input-side scaling can be fused into the preceding normalization or linear operators, while the output-side scaling can be absorbed by the subsequent dequantization step. 4.2

2-bit Blocks

10101010

10101010

... 1 0 1 0 1 0 1 0 ...

10 000000

...

10 000000

000000 10

...

000000 10

SIMD_AND 1 0000000

...

1 0000000

SIMD_SHIFT 00000 100

...

00000 100

SIMD_OR 00000 110

...

00000 110

Store

However, SIMD unpacking faces two challenges. First, irregular bit-widths such as 3-bit or 5-bit do not align naturally with byte boundaries. We address this issue by decomposing each weight into a combination of primitive weightlets with bit-widths 1, 2, or 4, so that unpacking reduces to handling only byte-divisible primitives. Second, SIMD manipulates data at byte granularity, which makes it inefficient to directly access logically adjacent weightlets stored contiguously. We therefore store weightlets in an interleaved layout. For an 𝑅-bit SIMD register, this layout allows processing up to 𝑅/8 weightlets in parallel and reconstructing 𝑅/8 consecutive weights simultaneously. Figure 7 illustrates the resulting storage format of a quantized model that contains 𝐿 layers, each with multiple weight tensors. Each tensor is a 𝐷 ×𝐶 matrix, where 𝐶 is the number of output channels and 𝐷 is the number of weights per channel. Within each channel, weights are partitioned into groups of 𝑅 consecutive weights. After decomposition, weightlets of the same bit-width in a group are stored into one or more 𝑅-bit blocks, each directly loadable into an SIMD register. For example, a group of 7-bit weights is decomposed into four 4-bit blocks, two 2-bit blocks, and one 1-bit block. To match SIMD byte-wise execution, weightlets belonging to 𝑅/8 logically consecutive weights (defined as a stripe), are interleaved with a stride 8/𝐵 for bit-width 𝐵 ∈ {1, 2, 4}. Since different channels may use different bit-widths, each packed tensor is prepended with per-channel bit-width metadata, stored as a compact INT3 array representing eight precision levels from 1 to 8 bits. EdgeFlow further designs an SIMD-based unpacking algorithm for efficient weight unpacking. The process is also conducted in a channel-wise, group-wise manner. First, EdgeFlow checks the metadata of the current channel to determine its bit-width, which indicates the number of blocks per group. Figure 8 illustrates the unpacking procedure exemplified by a 3-bit weight group, which consists of one 1-bit block and two 2-bit blocks, each mapped to a wide register. For each stripe, unpacking proceeds in three steps. Tak1 Mask. Valid bits are ing the first stripe as an example: ○ extracted via bitwise ANDs: the leading bit of each byte in the 1-bit block and the leading two bits in the first 2-bit block. 2 Align. The extracted bits are shifted to their target posi○ tions: the 1-bit block is right-shifted by 5 bits to form the

16 Weights ...

...

Figure 8. The SIMD-based unpacking algorithm.

Quantized Model Layer2 ... LayerL

10101010

Unpacked 3-bit Group (128 Bytes)

return 𝐵

Embedding Layer1

Packed 3-bit Group (48 Bytes)

SIMD-Friendly Packing Format

The unpacking process converts low-bit weights into INT8 tensors for NPU execution. EdgeFlow accelerates this process with SIMD-based unpacking. SIMD (Single Instruction Multiple Data) provides wide registers, e.g., 128-bit registers in Arm Neon [16], to apply the same operation to multiple elements in parallel. 7

Table 3. Datasets used in our evaluation. "OBQA" is OpenBookQA.

MSB, while the 2-bit block is right-shifted by 6 bits to form 3 Merge. The aligned blocks are combined the lower bits. ○ via bitwise ORs to form 3-bit weights, then cast to INT8. This process is repeated across all stripes (e.g., 8× per group) until the entire group is unpacked and stored for NPU execution. With this SIMD-based design, unpacking requires only 0.48 SIMD instructions per weight on average. 4.3

Datasets #shots #token range

LAMBADA WinoGrande OBQA – 59–97

3 131–144

MMLU HellaSwag

4 5 6 190–230 243–699 883–1086

The root cause of this issue lies in the lack of a decent tie-breaking mechanism when scheduling operators under the topological order of models’ execution graphs, e.g., the case when inputs of both O projection for later chunks and FNN projections for earlier chunks are ready. EdgeFlow breaks the tie by assigning operators a positional-guided priority. As shown in Figure 9b, for each operator, we assign a priority based on its chunk position, where earlier chunks in the prompt receive higher priority. This strategy effectively breaks the tie by executing CPU operators in earlier chunks and helps unlock downstream NPU operators earlier. Moreover, we find that the CPU remains idle for 65% of the end-to-end pipeline execution time. We thus implement a CPU task-stealing mechanism to enhance CPU utilization. As shown in Figure 9c, when the CPU is idle and the length of the NPU task queue exceeds a threshold, the CPU proactively takes and executes the top task from the NPU task queue. The threshold is adopted to avoid excessive CPU preemption that may leave the NPU idle. We determine the threshold according to the ratio between the execution time of CPU and NPU matmul operators and set it to 5 in our experiments.

Synergistic Granular Pipeline

As mentioned in Section 3.2, the key problems with existing pipeline scheduling approaches originate from their coarsegrained and static nature. EdgeFlow attacks the issue by first designing a fine-grained operator placement scheme that enables operator scheduling with more flexibility. Then, a dynamic operator scheduling mechanism is adopted to further reduce the pipeline bubbles. Fine-Grained Operator Placement. EdgeFlow performs scheduling at the granularity of individual operators, e.g., matmuls and layer normalizations. Compared with scheduling at subgraph granularity, such an approach enables finer-grained control over computation. In addition, we derive an initial placement of all operators among the NPU and the CPU to serve as the static backbone for the subsequent dynamic scheduling. We place all INT8 matmul operators on the NPU, i.e., Q/K/V/O projections for attentions and Gate/Up/Down projections for feedforward networks (FFNs), since NPUs are efficient in executing integer matmuls. Other operators are assigned to the CPU in FP16. To respect operator dependencies, we offline order operators topologically within a transformer layer. We maintain separate operator queues for the CPU and the NPU, and dispatch operators according to their topological order during runtime. Detailed placement is provided in Appendix B. Dynamic Operator Scheduling. To further optimize resource utilization, EdgeFlow proposes to dynamically schedule operators between the CPU and the NPU. Specifically, we adopt a position-guided operator priority to reduce NPU bubbles and a task-stealing scheme to reduce CPU bubbles. The position-guided operator priority reduces NPU bubbles caused by the suboptimal operator execution order. Specifically, existing approaches serialize the operator execution of input sequence chunks and parallelize chunk execution on the CPU and NPU. As shown in Figure 9a, the NPU first executes the Q/K/V projections for all chunks. Once the Q/K/V tensors are produced for a chunk, the CPU then computes the attention of the chunk in parallel. The O projection is then executed in parallel on the NPU when the corresponding attention score is ready for the chunk. Since the computation time of attentions grows linearly with the number of chunks, O projections on NPUs cannot overlap with attention computations on the CPU as the chunk count increases. NPU time is thus wasted on waiting for the CPU attention computation, even though the subsequent FFN projection of previous chunks can be executed first.

5

Evaluation

5.1

Experimental Setup

Implementations. We implement EdgeFlow from scratch with about 8K lines of code. EdgeFlow is composed of 1) a set of Python utilities that quantize models and pack weights in the SIMD-friendly packing format, and 2) a C++ runtime that unpacks weights and conducts inference computation. The current runtime is implemented with QNN and supports mobile devices with Arm CPUs and Hexagon NPUs. QNN embeds weights into compiled execution graphs in a proprietary storage format during the graph compilation phase. There is no interface for dynamically loading weights into a compiled execution graph. However, EdgeFlow stores quantized weights in an SIMD-friendly unpacking format. Weights need to be dynamically unpacked into execution graphs before computation. We reverse-engineer the proprietary execution graph format of QNN to enable dynamic loading and execution of weights by the NPU. Testbed. All experiments are conducted on a Xiaomi 15 Pro [57] equipped with a Qualcomm Snapdragon 8 Elite SoC, 16 GB of RAM, and 512 GB of flash storage. Inside the Snapdragon SoC, there is an Arm CPU with 8 cores, a Hexagon NPU, and an Adreno GPU. Models and Datasets. We evaluate EdgeFlow under models with sizes that are widely used on mobile devices, i.e., Llama3 8B [3], Mistral 7B [1], Phi3 3.8B [4], and Qwen1.5 1.8B [2]. 8

Attention

Q/K/V Proj

O Proj

0 1 2 3 4 5 6 7 CPU NPU 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6

8 7 Bubble 8

(a) Fine-Grained Operator Placement 0 1 2 3 4 5 6 7 CPU NPU 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6

FFN Proj (Gate/Up/Down) 9

0

4

9

1

2

7

4

5

6

7

8

9

2

8

3

5

6

7

8

9

0

1

2

3

4

5

0 2 3

4

6

5 7

8

9

steal task 8 0

0

1

2

3

4

5

6

7

8

9

Bubble

9 1

(b) Position-Guided Priority CPU 0 1 2 3 4 5 6 7 NPU 0 1 2 3 4 5 6 7 8 9 0 1 2 3 4 5 6

3

prioritize 8

0

Bubble

9

Bubble

7

9 1

2

8

9 3

4

5

7

6 8

9

1

6

7

8

9

Dynamic Operator Scheduling

(c) Task Stealing

Figure 9. The synergistic granular pipeline with fine-grained operator placement and dynamic operator scheduling. Each block represents an operator, and the numbers inside indicate chunk IDs.

As shown in Table 3, we utilize five datasets, i.e., LAMBADA [45], WinoGrande [47], OpenBookQA [43], MMLU [22], and HellaSwag [64], like prior works [18, 21, 58], to evaluate the model accuracy and the cold start latency of EdgeFlow under various prompt lengths. We adopt multi-shot prompting to generate prompt lengths in a more flexible manner. The summarized prompt lengths are shown in Table 3. Baselines. We evaluate EdgeFlow against llama.cpp [20], MNN [52], and llm.npu [58]. llama.cpp and MNN are popular open-source mobile LLM frameworks. llama.cpp maps model weights into memory with mmap and loads weights with on-demand paging during cold starts. MNN loads weights sequentially before the inference computation. Due to the bandwidth-bound nature during cold starts and the lack of NPU support in these two frameworks, we use the more widely adopted CPU backend for both baselines. llm.npu is the state-of-the-art NPU-based mobile LLM framework. It constructs and optimizes computation graphs on the critical paths of cold starts, which incur substantial graph compilation overhead. We augment llm.npu with materialization and overlapping as mentioned in Section 3.2 for more fair comparisons.

MMLU and HellaSwag datasets, the prefill computation becomes the bottleneck. The TTFT of EdgeFlow outperforms llm.npu with speedups ranging from 1.37× to 1.41× due to our fine-grained and dynamic pipeline design. The Model Accuracy. All baselines adopt INT8 quantization on model weights, which results in larger model sizes than EdgeFlow. Specifically, llama.cpp adopts per-block quantization only on weights. MNN applies static per-channel quantization on weights and dynamic per-channel quantization on activations. Finally, llm.npu adopts per-tensor INT8 quantization with a shadow outlier scheme. Specifically, it first quantizes weights and activations uniformly in the pertensor granularity and executes the quantized matmuls on the NPU. For activations with abnormal magnitudes, i.e., outliers, llm.npu executes them in FP16 on the CPU. As shown in the figure, when quantized to an average of 4 to 7 bits, EdgeFlow gets an accuracy drop of 22.48%, 6.36%, 1.91%, and -0.03%, respectively, compared with three baselines. We use 5-bit as the default configuration of EdgeFlow since it delivers substantial TTFT improvements under comparable accuracy with the NPU-based baseline, i.e., llm.npu. 5.3

5.2

Ablation Study

We conduct a comprehensive ablation study to evaluate the contributions of each technique to EdgeFlow’s overall performance with Llama3 8B and Phi3 3.8B over all datasets. The results are shown in Figure 11. We use INT8 and INT4 per-tensor quantization with NPU-aware smoothing as our baselines, i.e., INT8 and INT4 in the figure. EF 5bits denotes EdgeFlow with only NPU-aware adaptive quantization (Section 4.1) with an average bit-width of 5 bits. The quantized weights are stored with the INT4/INT8 mixed packing format. Pack and Pipeline indicate the incorporation of the SIMDfriendly packing format (Section 4.2) and the synergistic granular pipeline (Section 4.3), respectively. We observe that each design component contributes significantly to reducing TTFT or maintaining high accuracy, as detailed below. + NPU-Aware Adaptive Quantization. Compared with INT4, the NPU-aware adaptive quantization scheme improves model accuracy by an average of 35.6%. Compared with INT8, model accuracy only drops by 6.6% on average. These results

End-to-End Performance

Figure 10 presents the cold start latency (TTFT) of EdgeFlow and all baseline systems under different models and datasets. For EdgeFlow, we present both TTFT and model accuracy when the model is quantized to an average of 4 to 7 bits and discuss the results from two perspectives. The Cold-Start Latency. As the bit-width decreases from 7 bits to 4 bits, EdgeFlow progressively achieves higher speedups. Specifically, at 7 bits, it reduces TTFT by 3.92×, 2.28×, and 1.47× compared with llama.cpp, MNN, and llm.npu, respectively. While at 4 bits, the speedups increase to 4.24×, 2.53×, and 1.63×. In datasets with short prompts, i.e., under LAMBADA and WinoGrande datasets, TTFT is reduced since flash bandwidth is utilized more efficiently. This improvement comes from adaptive quantization, which maintains comparable model accuracy while minimizing the amount of data being loaded. When prompts are long, i.e., under 9

MMLU HellaSwag 423 699 883 1086

Accuracy

0 100 4.9 5.0 5.0 4.9

50

2.9 2.9 2.9 2.9

1.9 2.0 2.0 2.0

Accuracy (%)

28.9

MMLU HellaSwag 423 699 883 1086

0

Accuracy (%)

50

10.8 8.6

13.2 5.8 4.9

5.9 3.7 3.7

2.8 2.8 3.2

LAMBADA WinoGrande OBQA 59 97 131 144 190 230

25.2 20.4 16.2 16.4 16.7 16.8

59.9 6.9 5.8 6.2 6.3 6.5

15.8 12.8 9.3 9.6 9.7 9.9

27.0 11.6

18.7 9.3 6.3 4.3 4.4 5.0 5.5

8.6 8.5 6.2 3.4 3.6 4.2 4.8

101

100

124.8

Qwen1.5-1.8B

1.6 1.5 1.6 1.6

0

101

2.3 3.0

11.6 9.7 9.6 9.7 9.7

7.1 5.5 5.6 5.5 5.6

14.1

LAMBADA WinoGrande OBQA 59 97 131 144 190 230

EF 7bits Mistral-7B

102

1.4

0 100

EF 6bits

1.4 1.4 1.4 1.4

63.4

50

EF 5bits

Accuracy (%) TTFT (s)

50

29.0

4.3 3.6 3.7 3.7 3.8

8.6

EF 4bits 100

31.5 22.5 16.6 17.0 17.2 17.6

20.0 13.6 9.9 10.0 10.4 10.7

27.0 14.6 13.2 6.0 6.3 6.5 6.8

13.2 11.8 9.0 4.7 5.0 5.7 5.8

6.8 10.2 9.1

56.7

122.4

Phi3-3.8B

3.9 2.7 2.8 2.9 3.0

101

3.9 4.4 4.8 5.1

101

llm.npu+

Accuracy (%) TTFT (s)

MNN Llama3-8B

102

3.8 3.7 2.5 2.6 2.8 3.1

TTFT (s)

TTFT (s)

llama.cpp

Figure 10. The cold start latency (i.e., TTFT) and accuracy of different methods on various models and datasets. Bars show the average TTFT (the lower is better). Lines show accuracy. Prompt length is shown beneath each dataset label. "EF" represents EdgeFlow. "llm.npu+" represents llm.npu enhanced with materialization and overlapping techniques; the same notation is used throughout.

EF 5bits

1.5 1.0 0.5

LAMBADA WinoGrande OBQA

MMLU

HellaSwag

EF 5bits + Pack 100 1.5 50 1.0 0

0.5

EF 5bits + Pack + Pipeline Phi3-3.8B

Accuracy 100 50

LAMBADA WinoGrande OBQA

MMLU

HellaSwag

0

Accuracy (%)

INT8 Llama3-8B

Accuracy (%) Normalized TTFT

Normalized TTFT

INT4

Figure 11. The model accuracy and the normalized TTFT of INT4, INT8, and different techniques in EdgeFlow. Table 4. Quantized perplexities on WikiText-2 for Llama3-8B (FP16 perplexity: 14.59). The bold is the best. "SO" means shadow outlier.

show the effectiveness of the NPU-aware adaptive quantization scheme. Specifically, with just a 1-bit increase in average bit-width, model accuracy can be significantly enhanced. However, the reduction in TTFT compared with INT8 is only 3.3% due to the lack of an efficient packing format. Over 88% of weights are padded to 8 bits, leading to severe read amplifications during model loading. + SIMD-Friendly Packing Format. After adopting our packing format, TTFT decreases by up to 1.36× and 1.22× compared with EF 5bits under Llama3 8B and Phi3 3.8B, respectively. This effect is particularly pronounced under short-prompt conditions, where weight loading is the key bottleneck, i.e., under LAMBADA and WinoGrande datasets. + Synergistic Granular Pipeline. Introducing the synergistic granular pipeline further reduces the idle time and improves workload balance of the prefill computation during cold starts. This reduction is particularly pronounced on longer datasets. Specifically, for datasets with long prompts, i.e., MMLU and HellaSwag, the average TTFT decreases by 12.7%. Meanwhile, for datasets with short prompts, the TTFT also decreases by an average of 3.5%. 5.4

Llama3-8B

SmoothQuant

SO

CMPQ

EdgeFlow

4bits 5bits 6bits 7bits

410.67 99.68 97.31 75.00

10k+ 10k+ 814.51 21.62

103.52 18.22 15.83 15.11

30.96 17.27 15.63 15.09

outlier [58], and CMPQ [14]. We adapt SmoothQuant and CMPQ to comply with NPU constraints. For SmoothQuant, we apply the NPU smoothing technique to satisfy the NPU’s per-tensor output quantization constraint. For CMPQ, we modify its precision allocation strategy from input-channelwise to output-channel-wise, and replace its allocation metric with our relative error accordingly. Tables 4 report perplexities of quantized models in the Wikitext-2 [42] dataset. EdgeFlow steadily yields lower perplexities than all baselines in various precisions. Figure 12 shows the accuracy in different datasets and integer precisions for Llama3 8B. As the results across different models are similar, the remaining results are provided in Appendix C. EdgeFlow consistently outperforms previous methods across all settings. Specifically, compared with SmoothQuant, shadow outlier, and CMPQ, respectively, EdgeFlow achieves average accuracy improvements of 16.55%–38.23%, 13.72%–63.13%, and 0.67%–13.55% across 4 to 7 bits. Notably, at lower precisions, i.e., 4 and 5 bits, SmoothQuant and shadow outlier exhibit significant accuracy degradation. This indicates that treating all weights equally fails to capture the importance

In-Depth Analyses

We further conduct in-depth analyses to evaluate the effectiveness of the three key design components of EdgeFlow. 5.4.1 Analyses of NPU-Aware Adaptive Quantization. We evaluate the accuracy and perplexity of the NPU-aware adaptive quantization against three state-of-the-art quantization methods, i.e., SmoothQuant [55], llm.npu’s shadow 10

SmoothQuant shadow outlier CMPQ WinoGrande OpenBookQA

Accuracy (%)

LAMBADA

50

50 0

4

5 Bits 6

7

0

5 Bits 6

7

0

FP16

50

50 4

EdgeFlow MMLU

4

5 Bits 6

7

0

HellaSwag

50 4

5 Bits 6

7

0

4

5 Bits 6

7

2 3 Load Time (s)

(a) Loading and unpacking times.

5 0 5-bits

6-bits 7-bits Average Bits

llm.npu+ + Place seq_len=512

+ Priority CPU Total + Steal NPU CPU NPU 100 Phi3-3.8B 200 50 Llama3-8B 0 0 Llama3-8B Phi3-3.8B 150 0 150 300 Time (ms) Bubble Rate (%)

0

K-Quant EdgeFlow Latency by Bit-width Latency (ms)

5

INT4/INT8 Mixed Load vs Unpack Time Avg Bits 5 bits 6 bits 7 bits

Latency (s)

Unpack Time (s)

Figure 12. Accuracy of quantization schemes across precisions on Llama3 8B. The horizontal dotted line represents the FP16 accuracy.

(a) Single-layer execution times and bubble rates.

(b) The end-to-end latency.

Figure 13. Performance comparison of different storage formats.

(b) Single-layer execution times of the NPU and CPU.

Figure 14. Analyses of the synergistic granular pipeline.

dynamic scheduling. + Priority incorporates the positionguided priority and + Steal enables CPU task stealing. Figure 14a shows the single-layer execution latency and the bubble rates of the CPU and NPU. Figure 14b presents the execution times of CPU and NPU. Compared with llm.npu, + Place decreases total latency by 4.2%, and NPU execution time is reduced by 7.7% with nearly unchanged CPU time. This improvement mainly arises from offloading NPUinefficient operators to the CPU, which distributes the workload more appropriately. However, + Place employs a static CPU–NPU scheduling, which almost doubles the NPU bubble rate and leads to significant idle periods. + Priority further reduces latency by 5.2% and lowers the NPU bubble rate by 82.2%. This is achieved by ordering operator executions according to their positions in the prompt, allowing downstream NPU tasks to be unlocked sooner. Enabling task stealing, i.e., + Steal, further decreases execution time by 11.3% and reduces the CPU bubble rate by 61.6%. The CPU and NPU execution time also becomes more balanced. This gain results from better CPU utilization through opportunistic execution of NPU-assigned tasks. In particular, the NPU bubble rate stays roughly unchanged. This is attributed to the task-stealing threshold, which avoids task-stealing becoming so aggressive that it could make the NPU idle.

of different channels. Moreover, EdgeFlow exceeds CMPQ by 13.55%, 2.47%, 0.67%, and 0.90% at 4 to 7 bits, showing the effectiveness of our precision assignment strategy. 5.4.2 Analyses of SIMD-Friendly Packing Format. We evaluate the performance of EdgeFlow’s SIMD-friendly weight packing format against an INT4/INT8 mixed storage format and the K-Quant format. We use Llama3-8B and quantize it into varying bit-widths, i.e., 5, 6, and 7 bits. Figure 13a shows the unpacking and weight loading times of the three approaches. Compared with the INT4/INT8 mixed format, EdgeFlow speeds up model loading by 1.42× since it eliminates read amplification incurred by padding weights. The additional overhead on unpacking is less than 0.65s due to our efficient SIMD-based unpacking algorithm. Compared with the K-Quant format, EdgeFlow has identical load times since both approaches store weights in a compact format. However, the unpacking time is reduced by 6.17× thanks to the SIMD-based unpacking. Figure 13b presents the end-to-end latency under the same configuration. EdgeFlow consistently outperforms baselines across all bit-widths, achieving an average 3.05× and 1.44× latency reduction compared with K-Quant and INT4/INT8 mixed formats, respectively. Both results indicate that our SIMD-friendly storage format achieves a better trade-off between read amplification and unpacking efficiency.

5.5

Deployment Efficiency

We finally compare EdgeFlow with llm.npu in terms of decoding efficiency and resource consumption.

5.4.3 Analyses of Synergistic Granular Pipeline. We analyze the pipeline of EdgeFlow to show the effectiveness of the fine-grained operator placement, the position-guided priority, and the task stealing scheme. We use a sequence length of 512 for Llama3 8B and Phi3 3.8B. The baseline is the static and coarse-grained pipeline of llm.npu. + Place adopts the fine-grained operator placement scheme without

5.5.1 Decoding Efficiency. We break down the end-toend completion process of Mistral 7B and Qwen1.5 1.8B (512 input tokens, 512 generated tokens) on EdgeFlow and llm.npu, as shown in Figure 15. Since llm.npu performs decoding on the CPU, we adopt the same setup for EdgeFlow 11

Memory (GB)

Mistral-7B llm.npu+ EdgeFlow

decoding speedup = 1.00x

0

20

40

60

Time (s)

80

111.5 102.5

100

120

model + KV cache

5 0

0.0

2.5

5.0 7.5 Time (s)

10.0

(a) Memory footprint.

Figure 15. Breakdown of end-to-end completion latency.

5

10 5

0 Energy Power 0

(b) Energy and power.

Figure 16. Comparison of resource consumption during the cold start phase of Mistral 7B with 512 tokens.

to ensure a fair comparison. llm.npu adds a 2.9s overhead between cold start and decoding when switching to a CPU decoding graph, while EdgeFlow avoids this transition and enters decoding seamlessly. Isolating the impact on decoding itself, EdgeFlow achieves a speedup of 1.13× on Qwen1.5 1.8B and 1.00× on Mistral 7B compared with llm.npu. This shows that despite introducing additional processing during cold start, EdgeFlow does not degrade decoding efficiency. Instead, it preserves comparable decoding performance while reducing overall latency through better cold start design.

to the GPU based on NPU sensitivity profiling. Moreover, ShadowAttn [62] approximates important tokens in lowprecision on the NPU and performs high-precision sparse attention on the CPU. Storage-oriented approaches offload weights and activations on the flash to execute larger models and reduce I/O overhead during offloading. LLM in a Flash [6] bundles co-activated FFN neurons to avoid fragmented flash reads. ELMS [63] permutes weights by importance so that latency-critical prefixes can be loaded consecutively. Finally, PowerInfer-2 [59] keeps hot FFN neurons resident in memory and loads cold neurons on demand. Compared with existing approaches, EdgeFlow is the first to optimize cold inferences for mobile LLMs. Moreover, techniques proposed in EdgeFlow are also complementary to existing approaches. Specifically, the NPU-aware adaptive quantization algorithm and the SIMD-friendly packing format can be adopted together with existing storage-oriented works to further reduce I/O overhead during model offloading. The synergistic granular pipeline further improves the compute efficiency of existing compute-oriented approaches with its dynamic and fine-grained design. Mixed-Precision LLM Quantization. These algorithms leverage the varying importance of model weights by assigning different bit-widths to achieve a better accuracy and efficiency trade-off. LLM-MQ [32] uses first-order loss sensitivity to guide tensor-level precision assignment, while keeping sparse outlier weights in FP16. SliM-LLM [24] allocates bits based on reconstruction error, and performs group-wise calibration of quantization parameters. ResQ [48] uses principal component analysis to identify high-variance subspaces and performs rotation to mitigate outliers. CMPQ [14] profiles per-channel input magnitudes to estimate weight importance, and assigns bit-widths heuristically with non-uniform quantization. Compared with these approaches, EdgeFlow accounts for NPU-specific quantization constraints and delivers superior model accuracy under these constraints. Cold Start for Model Serving Systems. Existing cold-start mitigation techniques can be classified into cloud-based and mobile-based techniques. Cold-start mitigation techniques on the cloud are designed for GPU-based inference systems. Specifically, ServerlessLLM [19] leverages a multi-tier checkpoint caching and a GPU-aware formatting to reduce the provisioning latency of an inference instance. Medusa [65] and

5.5.2 Resource Consumption. Figure 16a shows the memory footprint of EdgeFlow and llm.npu during cold start. Ideally, the memory footprint of both approaches should be the summation of model size under INT8 quantization and the KV cache size (≈ 6.9 GB in total) since both approaches need to dequantize model weights to INT8 before computation. Initially, the memory usage of both approaches increases gradually since weights are loaded in a layer-bylayer manner. When it comes to the stable stage, llm.npu has higher memory consumption (≈ 8 GB) than the ideal case since it maintains separate computation graphs for CPU and NPU. The footprint of EdgeFlow is close to ideal since the weights in the computation graph are shared among CPU and NPU. Figure 16b reports the power and energy consumption during the same cold start process. EdgeFlow exhibits a 17.6% higher average power than llm.npu. This is because EdgeFlow offloads more lightweight operators to the CPU and opportunistically steals operators from the NPU. Since CPU execution generally incurs higher power than NPU, this leads to increased average power. Nevertheless, EdgeFlow achieves lower total energy consumption, i.e., reducing by 12.9%. This is due to reduced redundant data loading from flash storage and a more efficient execution pipeline, resulting in a shorter cold start time.

6

EdgeFlow Energy and Power 15 Power (W)

llm.npu+ Memory Footprint

Cold Start Load CPU Graph Decode

37.3 decoding speedup = 1.13x 28.0

Energy (mAh)

Qwen1.5-1.8B llm.npu+ EdgeFlow

Related Work

LLM Inference on Mobile Devices. Existing works on mobile LLM inference can be categorized into computation and storage optimizations. Computation-oriented approaches improve the execution efficiency of the prefill computation. llm.npu [58] executes attention on the CPU while running the remaining parts on the NPU. HeteroInfer [13] places dense kernels on the NPU and offloads remaining operations 12

PhoenixOS [53] materialize and load GPU execution states to enable faster initialization. ParaServe [40] and BLITZSCALE [67] aggregate bandwidth across GPU servers to parallelize model loading. Compared with these approaches, EdgeFlow tackles cold starts for mobile LLM inference, which cannot adopt these GPU-based approaches. For mobile devices, NNV12 [61] reduces cold-start latency by selecting startup-friendly kernels, which is complementary to EdgeFlow. There are also OS-level optimizations to reduce cold start latency for general applications, which can be broadly categorized into data prefetching [28, 49, 56], memory paging and swapping management [30, 35], and faster memory reclamation [33, 34]. These works are orthogonal to EdgeFlow since we focus on reducing the pure cold-start latency with algorithm and system co-design.

7

[9] Ioannis Arapakis, Souneil Park, and Martin Pielot. 2021. Impact of Response Latency on User Behaviour in Mobile Web Search. In CHIIR ’21: ACM SIGIR Conference on Human Information Interaction and Retrieval, Canberra, ACT, Australia, March 14-19, 2021, Falk Scholer, Paul Thomas, David Elsweiler, Hideo Joho, Noriko Kando, and Catherine Smith (Eds.). ACM, 279–283. https://doi.org/10.1145/3406522.3446038 [10] Saleh Ashkboos, Amirkeivan Mohtashami, Maximilian L. Croci, Bo Li, Pashmina Cameron, Martin Jaggi, Dan Alistarh, Torsten Hoefler, and James Hensman. 2024. QuaRot: Outlier-Free 4-Bit Inference in Rotated LLMs. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, Canada. http://papers.nips.cc/paper_files/paper/2024/ha sh/b5b939436789f76f08b9d0da5e81af7c-Abstract-Conference.html [11] Ron Banner, Itay Hubara, Elad Hoffer, and Daniel Soudry. 2018. Scalable Methods for 8-bit Training of Neural Networks. In Advances in Neural Information Processing Systems 31: Annual Conference on Neural Information Processing Systems 2018, NeurIPS 2018, Montréal, Canada. 5151–5159. https://proceedings.neurips.cc/paper/2018/hash/e82c4b1 9b8151ddc25d4d93baf7b908f-Abstract.html [12] Rune Birkmose, Nathan Mørkeberg Reece, Esben Hofstedt Norvin, Johannes Bjerva, and Mike Zhang. 2025. On-Device LLMs for Home Assistant: Dual Role in Intent Detection and Response Generation. (2025), 57–67. https://aclanthology.org/2025.wnut-1.7/ [13] Le Chen, Dahu Feng, Erhu Feng, Yingrui Wang, Rong Zhao, Yubin Xia, Pinjie Xu, and Haibo Chen. 2025. Characterizing Mobile SoC for Accelerating Heterogeneous LLM Inference. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP 2025, Lotte Hotel World, Seoul, Republic of Korea. ACM, 359–374. https: //doi.org/10.1145/3731569.3764808 [14] Zihan Chen, Bike Xie, Jundong Li, and Cong Shen. 2024. ChannelWise Mixed-Precision Quantization for Large Language Models. CoRR abs/2410.13056 (2024). https://doi.org/10.48550/arXiv.2410.13056 [15] Justin Cosentino, Anastasiya Belyaeva, Xin Liu, Nicholas A. Furlotte, Zhun Yang, Chace Lee, Erik Schenck, Yojan Patel, Jian Cui, Logan Douglas Schneider, Robby Bryant, Ryan G. Gomes, Allen Jiang, Roy Lee, Yun Liu, Javier Perez Matos, Jameson K. Rogers, Cathy Speed, Shyam A. Tailor, Megan Walker, Jeffrey Yu, Tim Althoff, Conor Heneghan, John Hernandez, Mark Malhotra, Leor Stern, Yossi Matias, Gregory S. Corrado, Shwetak N. Patel, Shravya Shetty, Jiening Zhan, Shruthi Prabhakara, Daniel McDuff, and Cory Y. McLean. 2024. Towards a Personal Health Large Language Model. CoRR abs/2406.06474 (2024). https://doi.org/10.48550/arXiv.2406.06474 [16] Arm Developer. 2025. Arm Neon. https://developer.arm.com/Architec tures/Neon. Referenced December 2025.. [17] Cathy Mengying Fang, Valdemar Danry, Nathan Whitmore, Andria Bao, Andrew Hutchison, Cayden Pierce, and Pattie Maes. 2024. PhysioLLM: Supporting Personalized Health Insights with Wearables and Large Language Models. In IEEE EMBS International Conference on Biomedical and Health Informatics, BHI 2024, Houston, USA. IEEE, 1–8. https://doi.org/10.1109/BHI62660.2024.10913781 [18] Elias Frantar, Saleh Ashkboos, Torsten Hoefler, and Dan Alistarh. 2022. GPTQ: Accurate Post-Training Quantization for Generative Pre-Trained Transformers. CoRR abs/2210.17323 (2022). https: //doi.org/10.48550/arXiv.2210.17323 [19] Yao Fu, Leyang Xue, Yeqi Huang, Andrei-Octavian Brabete, Dmitrii Ustiugov, Yuvraj Patel, and Luo Mai. 2024. ServerlessLLM: LowLatency Serverless Inference for Large Language Models. In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, USA. USENIX Association, 135–153. https://www.usenix.org/conference/osdi24/presentation/fu [20] Ggerganov. 2023. llama.cpp - LLM Inference in C/C++. https://github .com/ggerganov/llama.cpp. Referenced November 2025.

Conclusion

In this paper, we identify that the inefficient flash bandwidth utilization significantly deteriorates the cold start latency of mobile LLM inferences. We design EdgeFlow and show that the cold start latency can be mitigated by assigning precisions to weights adaptively according to their importance. EdgeFlow achieves this idea with three techniques, i.e., an NPU-aware adaptive quantization, an SIMD-friendly packing format, and a synergistic granular pipeline. Our evaluation results show that EdgeFlow significantly reduces cold-start latency up to 4.07× compared with state-of-the-art baselines while maintaining comparable model accuracy.

References [1] 2023. Mistral 7B. https://huggingface.co/mistralai/Mistral-7B-Instructv0.3. https://huggingface.co/mistralai/Mistral-7B-Instruct-v0.3 [2] 2023. Qwen1.5 1.8B. https://huggingface.co/Qwen/Qwen1.5-1.8BChat. https://huggingface.co/Qwen/Qwen1.5-1.8B-Chat [3] 2024. Llama 8B. https://huggingface.co/meta-llama/Meta-Llama-38B-Instruct. https://huggingface.co/meta-llama/Meta-Llama-3-8BInstruct [4] 2024. Phi3 3.8B. https://huggingface.co/microsoft/Phi-3-mini-4kinstruct. https://huggingface.co/microsoft/Phi-3-mini-4k-instruct [5] Amey Agrawal, Nitin Kedia, Ashish Panwar, Jayashree Mohan, Nipun Kwatra, Bhargav S. Gulavani, Alexey Tumanov, and Ramachandran Ramjee. 2024. Taming Throughput-Latency Tradeoff in LLM Inference with Sarathi-Serve. In Proceedings of the 18th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2024, Santa Clara, USA. USENIX Association, 117–134. https://www.usenix.org/confere nce/osdi24/presentation/agrawal [6] Keivan Alizadeh, Iman Mirzadeh, Dmitry Belenko, S. Khatamifard, Minsik Cho, Carlo C. del Mundo, Mohammad Rastegari, and Mehrdad Farajtabar. 2024. LLM in a Flash: Efficient Large Language Model Inference with Limited Memory. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2024, Bangkok, Thailand. Association for Computational Linguistics, 12562–12584. https://doi.org/10.18653/v1/2024.acllong.678 [7] Android Open Source Project. 2023. Low Memory Killer Daemon. https://source.android.com/docs/core/perf/lmkd. [8] Apple Developer Documentation. 2024. Reducing Terminations in Your App. https://developer.apple.com/documentation/xcode/reduceterminations-in-your-app. 13

[21] Zixu Hao, Jianyu Wei, Tuowei Wang, Minxing Huang, Huiqiang Jiang, Shiqi Jiang, Ting Cao, and Ju Ren. 2025. Scaling LLM Test-Time Compute with Mobile NPU on Smartphones. CoRR abs/2509.23324 (2025). https://doi.org/10.48550/arXiv.2509.23324 [22] Dan Hendrycks, Collin Burns, Steven Basart, Andy Zou, Mantas Mazeika, Dawn Song, and Jacob Steinhardt. 2021. Measuring Massive Multitask Language Understanding. In Proceedings of the 9th International Conference on Learning Representations, ICLR 2021, Virtual Event, Austria. OpenReview.net. https://openreview.net/forum?id=d7KBjm I3GmQ [23] Jiacheng Huang, Yunmo Zhang, Junqiao Qiu, Yu Liang, Rachata Ausavarungnirun, Qingan Li, and Chun Jason Xue. 2024. More Apps, Faster Hot-Launch on Mobile Devices via Fore/Background-aware GC-Swap Co-design. In Proceedings of the 29th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 3, ASPLOS 2024, La Jolla, CA, USA, 27 April 20241 May 2024. ACM, 654–670. https://doi.org/10.1145/3620666.3651377 [24] Wei Huang, Haotong Qin, Yangdong Liu, Yawei Li, Qinshuo Liu, Xianglong Liu, Luca Benini, Michele Magno, Shiming Zhang, and Xiaojuan Qi. 2025. SliM-LLM: Salience-Driven Mixed-Precision Quantization for Large Language Models. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025 (Proceedings of Machine Learning Research). PMLR / OpenReview.net. https://proceedings.mlr.press/v267/huang25aa.html [25] Apple Inc. 2025. Apple A19: Specs and Benchmarks. https://nanorevi ew.net/en/soc/apple-a19. Referenced November 2025.. [26] Qualcomm Technologies Inc. 2025. Qualcomm AI Engine Direct SDK. https://developer.qualcomm.com/sof tware/qualcomm-ai-enginedirect-sdk. Referenced November 2025. [27] Qualcomm Technologies Inc. 2025. Qualcomm Hexagon NPU - Powering the Generative AI Revolution. https://www.qualcomm.com/pro cessors/hexagon. Referenced November 2025.. [28] Yongsoo Joo, Junhee Ryu, Sangsoo Park, and Kang G. Shin. 2011. FAST: Quick Application Launch on Solid-State Drives. In 9th USENIX Conference on File and Storage Technologies, San Jose, CA, USA, February 15-17, 2011. USENIX, 259–272. http://www.usenix.org/events/fast11/t ech/techAbstracts.html#Joo [29] Evan King, Haoxiang Yu, Sangsu Lee, and Christine Julien. 2024. Sasha: Creative Goal-Oriented Reasoning in Smart Homes with Large Language Models. Proc. ACM Interact. Mob. Wearable Ubiquitous Technol. 8, 1 (2024), 12:1–12:38. https://doi.org/10.1145/3643505 [30] Changlong Li, Zongwei Zhu, Chao Wang, Fangming Liu, Fei Xu, Edwin H.-M. Sha, and Xuehai Zhou. 2025. Archer: Adaptive Memory Compression with Page-Association-Rule Awareness for High-Speed Response of Mobile Devices. In 23rd USENIX Conference on File and Storage Technologies, FAST 2025, Santa Clara, CA, February 25-27, 2025. USENIX Association, 497–511. https://www.usenix.org/conference/fa st25/presentation/li [31] Luchang Li, Sheng Qian, Jie Lu, Lunxi Yuan, Rui Wang, and Qin Xie. 2024. Transformer-Lite: High-Efficiency Deployment of Large Language Models on Mobile Phone GPUs. CoRR abs/2403.20041 (2024). arXiv:2403.20041 doi:10.48550/ARXIV.2403.20041 [32] Shiyao Li, Xuefei Ning, Ke Hong, Tengxuan Liu, Luning Wang, Xiuhong Li, Kai Zhong, Guohao Dai, Huazhong Yang, and Yu Wang. 2023. LLM-Mq: Mixed-Precision Quantization for Efficient LLM Deployment. In The Efficient Natural Language and Speech Processing Workshop with NeurIPS, Vol. 9. 3. [33] Wentong Li, Li-Pin Chang, Yu Mao, and Liang Shi. 2025. PMR: Fast Application Response via Parallel Memory Reclaim on Mobile Devices. In Proceedings of the 2025 USENIX Annual Technical Conference, USENIX ATC 2025, Boston, USA. USENIX Association, 1569–1584. https://ww w.usenix.org/conference/atc25/presentation/li-wentong [34] Yu Liang, Jinheng Li, Rachata Ausavarungnirun, Riwei Pan, Liang Shi, Tei-Wei Kuo, and Chun Jason Xue. 2020. Acclaim: Adaptive

Memory Reclaim to Improve User Experience in Android Systems. In Proceedings of the 2020 USENIX Annual Technical Conference, USENIX ATC 2020, July 15-17, 2020. USENIX Association, 897–910. https: //www.usenix.org/conference/atc20/presentation/liang-yu [35] Yu Liang, Aofeng Shen, Chun Jason Xue, Riwei Pan, Haiyu Mao, Nika Mansouri-Ghiasi, Qingcai Jiang, Rakesh Nadig, Lei Li, Rachata Ausavarungnirun, Mohammad Sadrosadati, and Onur Mutlu. 2025. Ariadne: A Hotness-Aware and Size-Adaptive Compressed Swap Technique for Fast Application Relaunch and Reduced CPU Usage on Mobile Devices. In IEEE International Symposium on High Performance Computer Architecture, HPCA 2025, Las Vegas, NV, USA, March 1-5, 2025. IEEE, 1588–1602. https://doi.org/10.1109/HPCA61900.2025.00118 [36] Geunsik Lim, Donghyun Kang, MyungJoo Ham, and Young Ik Eom. 2023. SWAM: Revisiting Swap and OOMK for Improving Application Responsiveness on Mobile Devices. In Proceedings of the 29th Annual International Conference on Mobile Computing and Networking, ACM MobiCom 2023, Madrid, Spain. ACM, 16:1–16:15. https://doi.org/10.1 145/3570361.3592518 [37] Ji Lin, Jiaming Tang, Haotian Tang, Shang Yang, Wei-Ming Chen, Wei-Chen Wang, Guangxuan Xiao, Xingyu Dang, Chuang Gan, and Song Han. 2024. AWQ: Activation-Aware Weight Quantization for On-Device LLM Compression and Acceleration. In Proceedings of the Seventh Annual Conference on Machine Learning and Systems, MLSys 2024, Santa Clara, USA. mlsys.org. https://proceedings.mlsys.org/ paper_files/paper/2024/hash/42a452cbafa9dd64e9ba4aa95cc1ef21Abstract-Conference.html [38] Lian Liu, Long Cheng, Haimeng Ren, Zhaohui Xu, Yudong Pan, Mengdi Wang, Xiaowei Li, Yinhe Han, and Ying Wang. 2025. COMET: Towards Practical W4A4KV4 LLMs Serving. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 2, ASPLOS 2025, Rotterdam, Netherlands. ACM, 131–146. https://doi.org/10.1145/3676641.3716252 [39] Yifei Liu, Jicheng Wen, Yang Wang, Shengyu Ye, Li Lyna Zhang, Ting Cao, Cheng Li, and Mao Yang. 2024. VPTQ: Extreme Low-Bit Vector Post-Training Quantization for Large Language Models. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, EMNLP 2024, Miami, USA. Association for Computational Linguistics, 8181–8196. https://doi.org/10.18653/v1/2024.emnlp-main.467 [40] Chiheng Lou, Sheng Qi, Chao Jin, Dapeng Nie, Haoran Yang, Xuanzhe Liu, and Xin Jin. 2025. Towards Swift Serverless LLM Cold Starts with ParaServe. CoRR abs/2502.15524 (2025). https://doi.org/10.48550/arX iv.2502.15524 [41] Xudong Lu, Yinghao Chen, Cheng Chen, Hui Tan, Boheng Chen, Yina Xie, Rui Hu, Guanxin Tan, Renshou Wu, Yan Hu, Yi Zeng, Lei Wu, Liuyang Bian, Zhaoxiong Wang, Long Liu, Yanzhou Yang, Han Xiao, Aojun Zhou, Yafei Wen, Xiaoxin Chen, Shuai Ren, and Hongsheng Li. 2025. BlueLM-V-3B: Algorithm and System Co-Design for Multimodal Large Language Models on Mobile Devices. In IEEE/CVF Conference on Computer Vision and Pattern Recognition, CVPR 2025, Nashville, USA. Computer Vision Foundation / IEEE, 4145–4155. https://op enaccess.thecvf.com/content/CVPR2025/html/Lu_BlueLM- V3B_Algorithm_and_System_Co-Design_for_Multimodal_Large_La nguage_Models_CVPR_2025_paper.html [42] Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. 2017. Pointer Sentinel Mixture Models. In Procedings of the 5th International Conference on Learning Representations, ICLR 2017, Toulon, France. OpenReview.net. https://openreview.net/forum?id=Byj72udxe [43] Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a Suit of Armor Conduct Electricity? A New Dataset for Open Book Question Answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, Brussels, Belgium. Association for Computational Linguistics, 2381–2391. https://doi.or g/10.18653/v1/d18-1260

14

[55] Guangxuan Xiao, Ji Lin, Mickaël Seznec, Hao Wu, Julien Demouth, and Song Han. 2023. SmoothQuant: Accurate and Efficient PostTraining Quantization for Large Language Models. In International Conference on Machine Learning, ICML 2023, Honolulu, USA (Proceedings of Machine Learning Research, Vol. 202). PMLR, 38087–38099. https://proceedings.mlr.press/v202/xiao23c.html [56] Li Xiaochen, Liu Sicong, Guo Bin, Ouyang Yu, Wu Fengmin, Xu Yuan, and Yu Zhiwen. 2026. AppFlow: Memory Scheduling for Cold Launch of Large Apps on Mobile and Vehicle Systems. In Proceedings of the 32th Annual International Conference on Mobile Computing and Networking, ACM MobiCom 2026, Austin, Texas, USA. ACM. https://doi.org/10.114 5/3795866.3796690 [57] Xiaomi. 2024. Xiaomi 15 Pro Full Specifications. Xiaomitime.com. https://xiaomitime.com/smartphones/xiaomi-15-pro Referenced November 2025.. [58] Daliang Xu, Hao Zhang, Liming Yang, Ruiqi Liu, Gang Huang, Mengwei Xu, and Xuanzhe Liu. 2025. Fast On-Device LLM Inference with NPUs. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, ASPLOS 2025, Rotterdam, The Netherlands. ACM, 445–462. https://doi.org/10.1145/3669940.3707239 [59] Zhenliang Xue, Yixin Song, Zeyu Mi, Le Chen, Yubin Xia, and Haibo Chen. 2024. PowerInfer-2: Fast Large Language Model Inference on a Smartphone. CoRR abs/2406.06282 (2024). https://doi.org/10.48550/a rXiv.2406.06282 [60] Tingxin Yan, David Chu, Deepak Ganesan, Aman Kansal, and Jie Liu. 2012. Fast app launching for mobile devices using predictive user context. In The 10th International Conference on Mobile Systems, Applications, and Services, MobiSys’12, Ambleside, United Kingdom June 25 - 29, 2012. ACM, 113–126. https://doi.org/10.1145/2307636.23 07648 [61] Rongjie Yi, Ting Cao, Ao Zhou, Xiao Ma, Shangguang Wang, and Mengwei Xu. 2023. Boosting DNN Cold Inference on Edge Devices. In Proceedings of the 21st Annual International Conference on Mobile Systems, Applications and Services, MobiSys 2023, Helsinki, Finland. ACM, 516–529. https://doi.org/10.1145/3581791.3596842 [62] Wangsong Yin, Daliang Xu, Mengwei Xu, Gang Huang, and Xuanzhe Liu. 2025. Dynamic Sparse Attention on Mobile SoCs. CoRR abs/2508.16703 (2025). https://doi.org/10.48550/arXiv.2508.16703 [63] Wangsong Yin, Rongjie Yi, Daliang Xu, Gang Huang, Mengwei Xu, and Xuanzhe Liu. 2024. ELMS: Elasticized Large Language Models on Mobile Devices. CoRR abs/2409.09071 (2024). https://doi.org/10.48550 /arXiv.2409.09071 [64] Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a Machine Really Finish Your Sentence?. In Proceedings of the 57th Conference of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2019, Florence, Italy. Association for Computational Linguistics, 4791–4800. https://doi.org/10.18653/v1/p19-1472 [65] Shaoxun Zeng, Minhui Xie, Shiwei Gao, Youmin Chen, and Youyou Lu. 2025. Medusa: Accelerating Serverless LLM Inference with Materialization. In Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, ASPLOS 2025, Rotterdam, The Netherlands. ACM, 653–668. https://doi.org/10.1145/3669940.3707285 [66] Cheng Zhang, Erhu Feng, Xi Zhao, Yisheng Zhao, Wangbo Gong, Jiahui Sun, Dong Du, Zhichao Hua, Yubin Xia, and Haibo Chen. 2025. MobiAgent: A Systematic Framework for Customizable Mobile Agents. CoRR abs/2509.00531 (2025). https://doi.org/10.48550/arXiv.2509.00531 [67] Dingyan Zhang, Haotian Wang, Yang Liu, Xingda Wei, Yizhou Shan, Rong Chen, and Haibo Chen. 2025. BlitzScale: Fast and Live Large Model Autoscaling with O(1) Host Caching. In Proceedings of the 19th USENIX Symposium on Operating Systems Design and Implementation, OSDI 2025, Boston, USA. USENIX Association, 275–293. https://www.

[44] Jakob Nielsen. 1993. Response Times: the Three Important Limits. Usability Engineering (1993). [45] Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The LAMBADA Dataset: Word Prediction Requiring a Broad Discourse Context. In Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), ACL 2016, Berlin, Germany. The Association for Computer Linguistics. https://doi.org/10.18653/v1/p16-1144 [46] Junhee Ryu, Dongeun Lee, Kang G. Shin, and Kyungtae Kang. 2023. Fast Application Launch on Personal Computing/Communication Devices. In Proceedings of the 21st USENIX Conference on File and Storage Technologies, FAST 2023, Santa Clara, USA. USENIX Association, 425– 440. https://www.usenix.org/conference/fast23/presentation/ryu [47] Keisuke Sakaguchi, Ronan Le Bras, Chandra Bhagavatula, and Yejin Choi. 2020. WinoGrande: An Adversarial Winograd Schema Challenge at Scale. In The Thirty-Fourth AAAI Conference on Artificial Intelligence, AAAI 2020, New York, USA. AAAI Press, 8732–8740. https://doi.org/ 10.1609/aaai.v34i05.6399 [48] Utkarsh Saxena, Sayeh Sharify, Kaushik Roy, and Xin Wang. 2025. ResQ: Mixed-Precision Quantization of Large Language Models with Low-Rank Residuals. In Forty-second International Conference on Machine Learning, ICML 2025, Vancouver, BC, Canada, July 13-19, 2025 (Proceedings of Machine Learning Research). PMLR / OpenReview.net. https://proceedings.mlr.press/v267/saxena25b.html [49] Sam Son, Seung Yul Lee, Yunho Jin, Jonghyun Bae, Jinkyu Jeong, Tae Jun Ham, Jae W. Lee, and Hongil Yoon. 2021. ASAP: Fast Mobile Application Switch via Adaptive Prepaging. In Proceedings of the 2021 USENIX Annual Technical Conference, USENIX ATC 2021. USENIX Association, 365–380. https://www.usenix.org/conference/atc21/presen tation/son [50] Junyang Wang, Haiyang Xu, Haitao Jia, Xi Zhang, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang. 2024. Mobile-Agent-v2: Mobile Device Operation Assistant with Effective Navigation via MultiAgent Collaboration. In Advances in Neural Information Processing Systems 38: Annual Conference on Neural Information Processing Systems 2024, NeurIPS 2024, Vancouver, Canada. http://papers.nips.cc/paper_f iles/paper/2024/hash/0520537ba799d375b8ff5523295c337a-AbstractConference.html [51] Junyang Wang, Haiyang Xu, Jiabo Ye, Ming Yan, Weizhou Shen, Ji Zhang, Fei Huang, and Jitao Sang. 2024. Mobile-Agent: Autonomous Multi-Modal Mobile Device Agent with Visual Perception. CoRR abs/2401.16158 (2024). https://doi.org/10.48550/arXiv.2401.16158 [52] Zhaode Wang, Jingbang Yang, Xinyu Qian, Shiwen Xing, Xiaotang Jiang, Chengfei Lv, and Shengyu Zhang. 2024. MNN-LLM: A Generic Inference Engine for Fast Large Language Model Deployment on Mobile Devices. In Proceedings of the 6th ACM International Conference on Multimedia in Asia Workshops (MMAsia ’24 Workshops). Association for Computing Machinery, Article 11, 7 pages. https: //doi.org/10.1145/3700410.3702126 [53] Xingda Wei, Zhuobin Huang, Tianle Sun, Yingyi Hao, Rong Chen, Mingcong Han, Jinyu Gu, and Haibo Chen. 2025. PhoenixOS: Concurrent OS-Level GPU Checkpoint and Restore with Validated Speculation. In Proceedings of the ACM SIGOPS 31st Symposium on Operating Systems Principles, SOSP 2025, Lotte Hotel World, Seoul, Republic of Korea. ACM, 996–1013. https://doi.org/10.1145/3731569.3764813 [54] Hao Wen, Yuanchun Li, Guohong Liu, Shanhui Zhao, Tao Yu, Toby JiaJun Li, Shiqi Jiang, Yunhao Liu, Yaqin Zhang, and Yunxin Liu. 2024. AutoDroid: LLM-Powered Task Automation in Android. In Proceedings of the 30th Annual International Conference on Mobile Computing and Networking, ACM MobiCom 2024, Washington D.C., USA. ACM, 543–557. https://doi.org/10.1145/3636534.3649379

15

usenix.org/conference/osdi25/presentation/zhang-dingyan [68] Tianyu Zhang, Lei Zhu, Qian Zhao, and Kilho Shin. 2019. Neural Networks Weights Quantization: Target None-Retraining Ternary (TNT). In Fifth Workshop on Energy Efficient Machine Learning and Cognitive Computing - NeurIPS Edition, EMC2@NeurIPS 2019, Vancouver, Canada. IEEE, 62–65. https://doi.org/10.1109/EMC2-NIPS53020.2019.00022 [69] Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena. In Advances in Neural Information Processing Systems 36: Annual Conference on Neural Information Processing Systems 2023, NeurIPS 2023, New Orleans, USA. http://papers.nips.cc/paper_files/paper/2023/hash/91f18a1287b398d 378ef22505bf41832-Abstract-Datasets_and_Benchmarks.html

16

A

Relative Error Derivation

NPU

We derive the relationship between the cosine distance and the expected squared quantization error for weight vectors. Let 𝑊 ∈ R𝑛 denote a weight vector and 𝑊 ′ its quantizedand-dequantized version. We define the quantization error vector as a random variable. ′

𝐸 := 𝑊 − 𝑊 ,

CPU

X

Add

Norm

Norm

Quant

Quant

Q Proj

K Proj

V Proj

Dequant

Dequant

Dequant

KVCache

𝐸 = (𝐸 1, . . . , 𝐸𝑛 ),

Gate Proj

Up Proj

Dequant

Dequant

SiLU

Mul

Attention

where each 𝐸𝑖 represents the quantization error of a component of 𝑊 . The cosine similarity between 𝑊 and 𝑊 ′ is

Quant

Down Proj

Quant

Dequant

O Proj Add

Dequant

𝑊 ·𝑊′ cos 𝜃 = . ∥𝑊 ∥ ∥𝑊 ′ ∥ We analyze this by substituting 𝑊 ′ = 𝑊 − 𝐸:

Figure 17. Transformer layer operator placement: CPU vs. NPU.

Similarly, the L2 norm of the weights can be written in terms of the mean squared value: 𝑛 ∑︁ ∥𝑊 ∥ 2 = 𝑊𝑖2 = 𝑛 · E[𝑊 2 ],

1. First, we analyze the numerator: 𝑊 · 𝑊 ′ = 𝑊 · (𝑊 − 𝐸) = 𝑊 · 𝑊 − 𝑊 · 𝐸 = ∥𝑊 ∥ 2 − 𝑊 · 𝐸 2. Second, we analyze the norm of 𝑊 ′ in the denominator:

𝑖=1

where E[𝑊 2 ] denotes the mean square of the weight com-

∥𝑊 ′ ∥ 2 = (𝑊 − 𝐸) · (𝑊 − 𝐸) = ∥𝑊 ∥ 2 − 2𝑊 · 𝐸 + ∥𝐸 ∥ 2

ponents. Finally, substituting these into the cosine distance approximation, we obtain

To simplify, we adopt two standard assumptions in quantization analysis:

E[∥𝐸 ∥ 2 ] 𝑛 · E[𝐸 2 ] E[𝐸 2 ] ≈ = . 2 ∥𝑊 ∥ 2 2 (𝑛 · E[𝑊 2 ]) 2 E[𝑊 2 ] This justifies the use of the expected squared error as a tractable approximation for the cosine distance between a weight vector and its quantized version. 1 − cos 𝜃 ≈

1. Small error energy: ∥𝐸 ∥ 2 ≪ ∥𝑊 ∥ 2 . 2. Uncorrelated error: 𝐸 is approximately uncorrelated with 𝑊 , i.e., E[𝑊 · 𝐸] ≈ 0.

B

Under these assumptions, the denominator can be expanded using a first-order Taylor approximation: √︁ ∥𝑊 ′ ∥ = ∥𝑊 ∥ 2 − 2𝑊 · 𝐸 + ∥𝐸 ∥ 2 √︁ ≈ ∥𝑊 ∥ 2 + ∥𝐸 ∥ 2 (since 𝑊 · 𝐸 ≈ 0) √︄ ∥𝐸 ∥ 2 = ∥𝑊 ∥ 1 + ∥𝑊 ∥ 2   ∥𝐸 ∥ 2 ≈ ∥𝑊 ∥ 1 + (for small ∥𝐸 ∥ 2 ). 2∥𝑊 ∥ 2

We provide a detailed operator-level execution flow of a transformer layer under EdgeFlow, including device placement, precision transitions, and data dependencies. EdgeFlow assigns all INT8 matrix multiplications to the NPU, including the Q/K/V/O projections in attention and the Gate/Up/Down projections in feedforward networks (FFNs). This design is motivated by the architecture of the Hexagon NPU, which is equipped with dedicated HMX units optimized for high-throughput matrix operations. In particular, INT8 matrix multiplications achieve significantly higher throughput compared to FP16 on the NPU. In contrast, the remaining operators are predominantly element-wise or low arithmetic-intensity operations, such as normalization, activation functions, and residual additions. Offloading these operators to the NPU would not yield meaningful speedup due to limited hardware specialization and potential overheads. Therefore, they are executed on the CPU in FP16 precision, which provides better efficiency for such workloads. As illustrated in Figure 17, each block represents an individual operator that is scheduled as a task and inserted

Substituting back, the cosine similarity becomes cos 𝜃 ≈

∥𝑊 ∥ 2 1 ∥𝐸 ∥ 2   = ≈ 1 − . 2 ∥𝐸 ∥ ∥𝐸 ∥ 2 2∥𝑊 ∥ 2 1 + 2∥𝑊 ∥𝑊 ∥ 2 1 + 2∥𝑊 ∥2 ∥2

Since 𝐸𝑖 are modeled as i.i.d. random variables with zero mean and finite variance, we can connect the squared L2 norm to the expected squared error: 𝑛 𝑛 h ∑︁ i ∑︁ E[∥𝐸 ∥ 2 ] = E 𝐸𝑖2 = E[𝐸𝑖2 ] = 𝑛 · E[𝐸 2 ]. 𝑖=1

Operator-Level Placement

𝑖=1

17

SmoothQuant shadow outlier CMPQ WinoGrande OpenBookQA

Accuracy (%)

LAMBADA 50 0

50 4

5 Bits 6

7

0

5 Bits 6

0

7

FP16

50

50 4

EdgeFlow MMLU

4

5 Bits 6

7

0

HellaSwag

50 4

5 Bits 6

7

0

4

5 Bits 6

7

Figure 18. Accuracy of quantization schemes across precisions on Phi3 3.8B. The horizontal dotted line represents the FP16 accuracy. Table 5. Quantized perplexities on WikiText-2 for Phi3-3.8B (FP16 perplexity: 10.01). The bold is the best. "SO" means shadow outlier.

Phi3-3.8B

SmoothQuant

SO

CMPQ

EdgeFlow

4bits 5bits 6bits 7bits

338.62 93.34 77.49 77.27

10k+ 10k+ 42.17 13.82

35.87 15.54 11.90 10.98

25.88 14.24 11.65 10.93

8B, Table 5 presents the perplexities on the Wikitext-2 [42] dataset, where EdgeFlow consistently achieves lower perplexities than all baselines across different precisions. Figure 18 illustrates the accuracy under various integer precisions on multiple datasets. Overall, EdgeFlow maintains clear advantages over prior methods in all evaluated settings. Compared with SmoothQuant, shadow outlier, and CMPQ, EdgeFlow improves accuracy by 44.07%–55.87%, 8.15%–67.06%, and 0.82%–11.29%, respectively, across 4 to 7 bits. Consistent with observations on Llama3, SmoothQuant and shadow outlier suffer from notable accuracy degradation at lower precisions (e.g., 4–5 bits), suggesting that uniform treatment of weights is insufficient to capture channel-wise importance. In contrast, EdgeFlow demonstrates stable performance across all bit-widths, outperforming CMPQ by 11.29%, 3.54%, 1.05%, and 0.82% at 4 to 7 bits, respectively. These results further validate the effectiveness of our adaptive precision assignment strategy and demonstrate that the benefits of EdgeFlow generalize well across different model architectures.

into either the CPU or NPU task queue based on its placement. The annotated edges form a topological backbone that captures the data dependencies among operators within a transformer layer. During execution, operators are dispatched in a dependency-aware manner. While respecting the topological order, EdgeFlow further optimizes execution using position-guided priority scheduling and task stealing, enabling dynamic load balancing across devices.

C

Adaptive Quantization Evaluation

We further evaluate adaptive quantization on Phi3 3.8B, with detailed results reported in this section. Similar to Llama3

18

Record · ID 5937 · SHA-256 d7422e3a5a0208ef
Conceptio Open Knowledge Archive — every document is proof-bundled with source, license, and retrieval metadata.