ConceptioArchivearXiv CS
arXiv CSopen access

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference

Unknown · 2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
clouddistributedcomputingparallelcomputing
distributed computing, parallel computing, cloud

Beyond Uniform Experts: Cost-Aware Expert Execution for Efficient Multi-Device MoE Inference

Abstract—Mixture-of-Experts (MoE) architectures enable language models to achieve unprecedented scale via sparse activation. However, their inference performance is often limited by data movement bottlenecks. Two coupled challenges exacerbate this limtation: (1) Importance-Agnostic Cost: Low-contribution experts incur nearly uniform memory and transfer costs, resulting in a low cost-to-benefit ratio and wasting critical bandwidth; (2) System-Level Imbalance: Multi-device deployments are universally bottlenecked by the slowest device, meaning that local reductions on one device may yield no improvement in end-to-end latency. We propose Cost-Aware Expert Execution (CAEE), a hardwareguided runtime framework that jointly optimizes for token-level expert importance and system-level execution cost. CAEE uses lightweight, calibrated cost models to estimate hardware overhead, selectively prunes low-importance, high-cost experts, and redistributes their contributions via a low-overhead compensation mechanism, avoiding extra data movement. Evaluations on the 671B DeepSeek-R1 model show that CAEE can reduce end-to-end inference latency by 8%-18% across diverse deployment settings, including expert offloading and on-device execution on multidevice systems, while maintaining a model accuracy drop of less than 1%.

Roofline Model for Different System (FP8/INT8) 103

Performance (TFLOPS)

arXiv:2606.29982v1 [cs.DC] 29 Jun 2026

Hui Zang, Pengfei Xia, Hong Liu, Jiajia Chu, Tuo Hao, Minghao Chen, Rui Zhang, Ziyang Zhang Huawei Technologies Ltd

102 101 GPU1 on-chip GPU1 offloading GPU2 on-chip GPU2 offloading

100 10 1

MatMul, BS=1~128 100

101

102

103

104

Compute Intensity (FLOPs/Byte)

105

106

Fig. 1: Roofline analysis of MoE expert MatMul operations on two types of GPU. Within the range of batch sizes from 1 to 128, it is evident that the MoE inference is primarily dominated by data movement.

I. I NTRODUCTION Mitigating this persistent data movement bottleneck requires a holistic approach that goes beyond simply reducing the number of activated experts. It mandates a joint consideration of the individual execution cost of each expert and its nonlinear impact on the overall system performance. In practice, two coupled inefficiencies arise:

Large Language Models (LLMs) [4], [15], [28] have driven transformative progress in AI. However, their growing scale, now often exceeding hundreds of billions of parameters, introduces significant challenges for efficient deployment. Mixture-of-Experts (MoE) [13], [22] architectures alleviate the computational scaling problem via sparse activation, where only a small subset of k experts processes each input token. While this achieves computational sparsity, it fails to address the memory and data movement burden: all N expert parameters must remain accessible to support dynamic, token-level routing decisions. As a consequence, on modern accelerators, the data movement often emerges as the dominant bottleneck for MoE inference throughput across the memory hierarchy [10], [21]. Fig. 1 illustrates the roofline analysis of MoE expert matrix multiplications (MatMuls) on two types of GPUs. Across a practical range of batch sizes from 1 to 128, the arithmetic intensity of the expert computation remains low, confirming that MoE inference is memory-bound rather than compute-bound. This bottleneck is severely amplified in scenarios involving expert offloading [14], [31], where inactive experts must be swapped between host memory and device memory via interconnects like PCIe or CXL. With limited link bandwidths, typically in the tens of GB/s [1], [2], the transfer latency restricts the overall throughput long before the device’s computational resources are saturated, rendering the execution transfer-bound.

Importance-Agnostic Cost (C1): Experts exhibit highly non-uniform token-level importance, yet they incur a nearly uniform runtime cost when activated. Every activated expert, regardless of whether it contributes 50% or just 5% to the final output logits, triggers a full memory access or a complete Host-to-Device (H2D) parameter transfer, wasting valuable memory bandwidth and interconnect resources. • System-Level Imbalance (C2): In multi-device systems, the end-to-end latency is dictated by the slowest device’s completion time. The uneven, dynamic activation of experts across devices leads to load imbalance. Therefore, reducing the cost of a single expert on an underutilized device may yield little improvement in the end-to-end latency. Only by addressing the cost on the bottlenecked device or link can the system performance be improved.

These challenges motivate the development of a new runtime framework that precisely aligns the expert’s importance with its actual hardware execution cost on both per-expert and multi-

1

DDR device dimensions. Existing methods partially address these problems. Static H2D/D2H (e.g., PCIe) CPU pruning and compression [8], [17], [32] reduce the memory footprint and data movement by removing experts, but they D2D (e.g., NVLINK, Unified Bus) HBM HBM HBM HBM ignore dynamic, token-level importance, which lowers model xPU xPU xPU xPU capacity. Dynamic routing approaches [3], [25], [37] consider expert importance but typically ignore per-expert hardware cost or fail to evaluate system-wide, no-linear straggler impact. Moreover, most previous studies [5], [14], [23], [30] have Fig. 2: A logical topology diagram for a single-node multitypically been designed for single scenarios, leaving the device server, where xPU represents computation devices such effectiveness in both offloading and on-device configurations as GPU. in multi-device systems untested. Motivated by the challenges, we propose Cost-Aware Expert expert processing, G(t)e is the routing Execution (CAEE), a hardware-guided runtime framework. where Ee (t) denotes the P score for expert e and G(t)e = 1. The Top-k constraint CAEE refines expert execution by prioritizing cost reduction ensures computational sparsity. at the system’s true bottleneck. Our main contributions are: However, the memory requirement scales linearly with N . • Lightweight Multi-Device Cost Modeling: We develop Since N can be in the hundreds for large MoE models, the a pragmatic cost model calibrated through brief offline memory footprint is immense. For example, DeepSeek-V3 [13], profiling that estimates the per-expert data movement cost with 671 billion parameters, requires 1 TB of memory under e d), and crucially, a max-aggregation layer cost Fcost , C(e, FP16 quantization. The another core challenge arises during which explicitly captures the system-level straggler effect batched inference, where dynamic activation concurrently across multiple devices D. engages a substantial number of experts. This mandates the data • Cost-Aware Pruning Strategy: We formulate the expert movement of a large fraction of expert weights, either through selection as an optimization problem, seeking to minimize high-frequency on-device memory accesses or via offloading the system-level cost Fcost while bounding the removed from a slower host memory tier, e.g., DRAM, NVMe. As a importance Fimp . This strategy selectively bypasses only result, parameter fetching latency rapidly becomes the primary the high-cost, low-importance experts that contribute to bottleneck, overshadowing the computational gains of sparsity. the system’s overall latency. B. Multi-Device Systems • Low-overhead Compensation: To maintain accuracy, we introduce a mechanism that redistributes the contributions In modern computing environments, multi-device servers of pruned experts only to the remaining already-active or are standard for LLM inference and training. As shown in low-cost experts Atrans . This ensures negligible accuracy Fig. 2, these systems typically feature host servers equipped degradation without incurring any additional memory with multiple computation devices, each possessing its own access or data movement overhead. HBM and computational units. The key components include: These techniques allow CAEE to improve per-expert effi• Device Interconnects: High-speed, low-latency links, e.g., ciency and reduce system-level stragglers, increasing end-to-end NVLINK, connecting the devices for fast data exchange inference performance across both on-device and offloaded in parallel execution. deployments. We evaluate CAEE on the 671B DeepSeek• Host-Device Interconnects: Slower links, e.g., PCIe [2], R1 [11] model across multiple settings and datasets. CAEE CXL [1], connecting the devices to the host CPU and achieves a 8%–18% reduction in end-to-end inference latency DRAM for system-wide memory access. with < 1% accuracy impact, demonstrating its effectiveness This architecture enables Expert Parallelism (EP) [29] for under diverse hardware bottlenecks. MoE, where the N experts are partitioned across the D devices. A token’s execution flow involves: first, token dispatch, as the router decides which device d is responsible for computing the II. P RELIMINARIES token based on the required expert e; second, expert activation, A. MoE and Data Movement Bottleneck ensuring device d has expert e’s parameters in its HBM; third, MoE architectures, exemplified by models like DeepSeek- computation, where expert e processes the token; and finally, R1 [11] and Qwen3 [33], extend the scalability of LLMs by result aggregation, gathering results across devices. The central introducing conditional computation. A router dynamically challenge is that dynamic and unpredictable routing leads to selects a small subset Kt of k experts from N total experts to uneven memory access and data transfer patterns across the process an input token t. The output y is a weighted sum of devices and links [13], [19]. t

the selected expert outputs: X yt = G(t)e · Ee (t)

III. A NALYSIS AND C HALLENGES In this section, we analyze two fundamental challenges, that arise when deploying MoE models on multi-device systems.

e∈Kt

2

Layer 𝑖

H2D1

Layer 𝑖 + 1

Fetch

xPU2

Non-MoE

R

E

E

E

Non-MoE

R

E

H2D2

E

E

Non-MoE

R

Fetch E

Cache-Hit Experts

E

Cache-Hit Experts (Low-Importance)

E

Cache-Miss Experts

E

Cache-Miss Experts (Low-Importance)

dispatch

R

combine

Non-MoE

dispatch

xPU1

Fetch R

Router

Fig. 3: Illustration of the two key challenges in offloading MoE models on multi-device systems. ( 1 ) For cache-miss experts, the retrieval cost is high regardless of the token’s importance score; redirecting tokens away from low-importance, cache-miss experts can save H2D bandwidth. ( 2 ) Uneven expert activation leads to idle devices and underutilized H2D bandwidth, reducing expert transmissions on idle links may have little to no effect on improving end-to-end performance.

For clarity, we use the expert offloading scenario [14], [31] as a running example, because it amplifies data-movement behavior and makes the issues visible in practice. However, these challenges are not unique to offloading: they stem from the general data-movement–bound nature of MoE inference and persist across on-device deployments. Offloading is therefore an illustrative case rather than the only scenario of concern. Fig. 3 highlights these two challenges.

Importance Score Among Activated Experts

Importance Score

0.25 0.20

Highest Score Lowest Score

0.15 0.10

A. C1: Importance-Agnostic Expert Transfer As depicted in Fig. 3 (Yellow Box 1 ), the execution cost of activating an expert is largely independent of its routing score G(t)e . This behavior exposes a mismatch:

0

Uniform Transfer Cost: When an expert’s parameters are needed for a token and are not currently in the device’s HBM (a cache-miss), the system must execute a full transfer of its entire weight tensor Se from the slower memory tier. This transfer operation consumes a fixed and substantial amount of memory or link bandwidth Bd . • Highly Variable Importance: For a Top-k routing manner, the routing scores G(t)e can differ by orders of magnitude. The k-th selected expert might have a score of 5%, while the top expert could score over 50%. Both, if cache-misses, incur the same transfer cost.

10

20

30

Layer

40

50

Fig. 4: Variation in expert importance across layers. For each layer, we compute the average of the highest and lowest importance scores among the experts activated during 16 batch sizes and 100 decoding processes. The differences in expert importance were smaller in earlier layers, but after Layer 10, the scores of the highest-importance experts were typically 3 to 4 times higher than those of the lowest-importance experts. Despite this divergence, activating either expert incurs the same transfer cost, illustrating the mismatch that underlies C1.

The result is a fundamental mismatch: high-cost, lowimportance experts consume scarce data movement resources, creating a low cost-to-benefit ratio. For example, in an offloading system, directing a token away from a low-importance, cache-miss expert can save tens of milliseconds of H2D transfer

latency, which is far more impactful than rerouting tokens away from a high-importance, cache-hit expert. 1) Quantitative Analysis: We analyze runtime statistics from the 671B DeepSeek-R1 model on an 8 device system with expert offloading. Fig. 4 plots the layer-wise average highest

3

Number of Experts Across Devices 0.175

Distribution

0.150 0.125

Fig. 6: Communication imbalance observed in profiling data. Among the three devices, the longest communication time determines the overall latency.

0.100 0.075 0.050 0.025 0.000

0.0

2.5

5.0

7.5

10.0

12.5

Device Cost Difference

15.0

17.5

This imbalance is the empirical evidence for C2, proving the existence of the straggler problem. This underscores the necessity of a cost model that uses max-aggregation to target the true performance bottleneck.

20.0

Fig. 5: Distribution of device-level expert-transfer imbalance. IV. S YSTEM D ESIGN For each inference run, we calculate the number of experts In this section, we detail the design of Cost-Aware Expert migrated across all 8 devices and record the difference Execution (CAEE), which is executed as a runtime component between the busiest and the least busy devices. A difference on each device Processi to dynamically refine expert selection. of 0 indicates perfect balance. Most runs exhibit significant The framework’s goal is to transition from importance-only imbalance, revealing the inherent bias driving C2. routing to a cost-aware routing decision that minimizes the system-level straggler latency. and lowest importance scores among activated experts. The As illustrated in Fig. 7, CAEE operates in two phases: an divergence is significant: after Layer 10, the maximum score offline phase for cost model calibration, and a runtime phase is typically 3× to 4× greater than the minimum score. This incorporating the core mechanisms: (1) hardware cost modeling, contrast confirms the high variability in expert importance. (2) cost-aware pruning, and (3) low-overhead compensation. Since the data movement cost for both experts is the same, this divergence is the empirical evidence for the C1 mismatch. A. Hardware Cost Modeling CAEE’s hardware cost model is centered around the idea B. C2: Data Movement Imbalance Across Devices that the total layer latency Fcost is determined by the maximum As highlighted in Fig. 3 (Red Line 2 ), in multi-device data movement cost across all devices. EP deployments, the distribution of expert activation and data 1) Per-Expert Data-Movement Cost: For an expert e asmovement is rarely uniform. We can observe that: signed to device d, the cost C(e, d) is based on the volume of • Uneven Workload: Dynamic routing based on token data movement required. The core formulation unifies the cost semantics leads to an unpredictable distribution of expert regardless of the source memory: demands. One device might handle tokens that primarily Se hit cached experts, while another device is tasked with C(e, d) ∝ + Overheadd (e) B tokens that trigger multiple cache-misses, leading to its d H2D link being fully saturated. where Se is the fixed parameter size of expert e, Bd is the effec• Straggler Effect: The system performance is capped by tive data-movement bandwidth associated with device d, e.g., the slowest device, i.e., Latencytotal = maxd∈D (Latencyd ). HBM bandwitdh or PCIe link bandwidth, and Overheadd (e) Therefore, reducing Latencyd on an already fast device is a constant term representing the scheduler overhead. does not actually improve Latencytotal . Only reducing the 2) Layer-Level Cost via Max Aggregation: To capture the C2 cost on the busiest device or link can yield system-wide straggler effect, the layer cost F is defined by the maximum cost performance improvement. total cost incurred by any device in the system: This confirms that naive, uniform pruning based on global ! X metrics is ineffective and potentially harmful. A cost-aware Fcost (X ) = max xe · R(e, d) · C(e, d) pruning strategy must explicitly minimize the maximum cost d∈D e∈Ed across all devices. 1) Quantitative Analysis: Fig. 5 aggregates the number of Here, X = {xe } is the set of expert activation decisions and expert transfers across the 8 devices. The results show that: xe = 1 if active, Ed is the set of experts assigned to device d, perfect balance, gap = 0, is rare, while significant imbalance, and C(e, d) is the intrinsic parameter transfer cost. Crucially, gap of 6 to 10 transfers, occurs frequently. This means that, for R(e, d) represents the dynamic memory residency status: for a given inference step, one device or link is handling up to 10 on-device inference, R(e, d) is typically set to 1 to reflect more expert transfers than another. Another practical profiling uniform memory access cost or HBM pressure; conversely, data example is shown in Fig. 6. for expert offloading systems, R(e, d) = 1 if expert e triggers

4

Runtime Offline Hardware Topology

Process1 Batch Inputs

Processn

Router

MoE

Profiling Data

Lightweight Modeling

Importance Calculation

Low-overhead Compensation

Cost-Aware Pruning update

Transfer Engine

Data Movement

Lightweigh t Modeling

Cost-Aware Pruning

Low-overhead Compensation

read

Transferable Set

Parameter Storage

CAEE Fig. 7: Overview of the CAEE framework. CAEE operates within the runtime of each device Processi in a multi-device system. The offline phase calibrates the lightweight modeling component using hardware topology and profiling data to estimate per-expert and per-system cost. At runtime, the cost-aware pruning module integrates the cost estimate and importance calculation to derive the final set of active experts, known as the transferable set Atrans . The low-overhead compensation mechanism only reroutes the pruned contribution values to the expert nodes in the transferable set, thereby ensuring that the redistribution is completed without making any additional data movement requests to the transfer engine.

an expensive H2D transfer, and R(e, d) = 0 if it is a cache- where ε bounds the acceptable accuracy degradation. Due to the strict time requirements for inference speed, hit. By minimizing Fcost (X ), we directly target the throughput CAEE solves the constrained optimization problem using a bottleneck by reducing the workload on the busiest device. 3) Offline Profiling and Calibration: To account for real- efficient cost-invariant greedy heuristic that minimizes Fcost world non-idealities such as contention, varying clock speeds, while strictly bounding the importance loss. The process and firmware overheads, CAEE performs a brief offline dynamically determines the final set of transferable experts profiling phase. This calibrates the effective bandwidth Bd Atrans ∈ Aact via four steps: and fixed overheads: 1) Initial Expert Set Separation: We first partition the entire set of activated experts Aact into two groups based S e d) = αd · e + βd C(e, on a predetermined high importance threshold θ: (1) Bd Mandatory Set Amust : Experts whose importance score where αd is the bandwidth degradation factor and βd is the fixed I(e) is above θ are critical for accuracy, i.e., Amust = transfer overhead. This calibration ensures the cost model’s {e | I(e) ≥ θ and e ∈ Aact }. (2) Candidate Set Acand : e d) highly accuracy, making the online cost estimation C(e, Experts whose scores are below θ, representing potential lightweight, enabling real-time inference through simple table candidates for removal, that is Acand = E \ Amust . lookups and aggregation operations. 2) Baseline Cost Determination: The initial baseline execution cost Fbase is calculated based solely on the mandatory B. Cost-Aware Pruning experts Amust , as they must be transferred or accessed: ! Our goal is to select a subset of experts from the originally X e activated ones, that is, A ⊆ Aact that significantly reduces the R(e, d) · C(e, d) Fbase = max d∈D system cost Fcost while minimally impacting the total expert e∈Amust ∩Ed importance Fimp . We use I(e) to denote the importance of 3) Cost-Invariant Expansion: To maximize accuracy without expert e. The objective can be formulated as increasing the critical path latency, we incorporate experts ! from Acand . An expert e ∈ Acand is promoted to Amust X e d) if and only if its inclusion does not increase the layer’s minimize: Fcost (X ) = max xe · R(e, d) · C(e, d∈D total execution time above the baseline cost Fbase , i.e., e∈Ed   N X X e ′ , d) = Fbase if max  R(e, d) · C(e subject to: Fimp (X ) = (1 − xe ) · I(e) ≤ ε d∈D e′ ∈Amust ∪{e}

e=1

5

TABLE I: Experimental setup for offloading and on-device deployments. Task

Device Num

LLM Engine

Deploy Strategy

R1 Thinking

Offloading On-device

8 16

vLLM MindIE

MLA: DP1-TP8, MoE: EP8 MLA: DP2-TP8, MoE: EP16

Disabled Enabled

then Amust ← Amust ∪ {e} DDR

This key step effectively performs cost-aware expert expansion: it prioritizes retaining low-importance experts that happen to be cache-hits or reside on underutilized links, ensuring we maximize utility without contributing to the straggler effect. 4) Final Pruning and Expert Set Definition: After the expansion phase, the experts remaining in Acand are officially designated as the pruned set, as their inclusion would increase the latency Fcost beyond Fbase . The final set of active, transferable experts is defined as Atrans = Amust and Atrans ⊆ Aact . This strategy is highly effective because it ensures the pruning decision is guided by the hardware cost at the bottleneck, i.e., maintaining Fcost = Fbase , rather than relying on simple global importance scores. The resulting set Atrans is optimized for both speed and model accuracy.

CPU

256×58 Experts

PCIe

xPU

LFU Eviction

HBM Cache ~1024 Experts

Fetch

… Fig. 8: Expert offloading system used in the experment. remains dense and maximizes the utility provided by the cost-optimized set. V. E VALUATION We evaluate the proposed framework with two MoE deployments: the offloading scenario and the on-device scenario. A. Experimental Setup

1) Models and Benchmarks: We employ DeepSeek-R1 671B [11] with W8 quantization as our representative MoE model. C. Low-overhead Compensation The low-overhead compensation mechanism is deployed The model contains 671B total parameters with 256 experts, to finalize the expert selection after the cost-aware pruning but activates only 37B parameters and 8 experts per token determines the set of active, transferable experts, Atrans . This during inference. A key innovation is its Multi-Head Latent mechanism is designed for zero additional memory traffic, Attention (MLA) [12], which replaces standard attention to operating exclusively via fast, local data manipulation to achieve a massive reduction in the Key-Value (KV) cache size. To evaluate the impact of CAEE on model accuracy, we maintain accuracy while respecting the pruning decision. The employ multiple open-source benchmarks across three domains: core principle is to constrain the final routing decision only Knowledge (MMLU-shot5 [18], CEval-shot5 [20]), Math to the experts remaining in Atrans , ensuring that no memory(GSM8K-shot5 [9]), and Code (HumanEval [7]). For inference intensive data movement is initiated for any pruned expert. performance, we measure Time to First Token (TTFT) and For a given token t, the compensation process is simplified Time Per Output Token (TPOT) under fixed concurrency, to the following two steps: while adjusting the per-device concurrency according to each 1) Masking of Pruned Experts: The scores of all experts e deployment scenario. outside the transferable set Atrans are effectively reset to 2) Implementation Details: To guarantee a diverse evaluazero, creating a masked score set Gmasked (t): tion, the experiments are conducted across separate offloading ( and on-device deployments of DeepSeek-R1, each utilizing G(t)e if e ∈ Atrans Gmasked (t)e = different hardware, inference frameworks, and routing strategies. 0 if e ∈ / Atrans TABLE I outlines the corresponding experimental setup. This masking mechanism ensures that the pruned expert models are completely disregarded during the final output B. Expert Offloading aggregation process, while retaining the scores of all The expert offloading system evaluated in this work is transferable experts. depicted in Fig. 8. Experiments are conducted on a computing 2) Final Top-k Selection: The final set of k active experts server featuring 8 xPUs and 1.5 TB of host memory. Since the Kfinal is selected by applying the standard Top-k operation DeepSeek-R1 model’s 58 MoE layers (totaling 14848 experts) on the masked score set Gmasked (t): exceed the available device memory, each HBM caches only a subset of 1024 experts, managed by a Least Frequently Kfinal = Top-k(Gmasked (t)) Used (LFU) eviction policy. Each expert has a size of 43 MB, This step effectively re-routes the token’s contribution, corresponding to weight dimensions of 3 × 7168 × 2048. which would have gone to the pruned experts, to The expert offloading scenario, where the H2D transfer the remaining highest-scoring experts within the low- dominates the critical path, affords greater flexibility than overhead set Atrans . This ensures that the model output highly latency-sensitive on-device inference. This allows us to

6

TABLE II: Performance of CAEE in the expert offloading scenario. Baseline

θimp = 0.9

θimp = 0.8

θimp = 0.7

TTFT (s)

TOPT (s)

TTFT (s)

TOPT (s)

TTFT (s)

TOPT (s)

TTFT (s)

TOPT (s)

Batch Size = 1 Gain vs. Baseline

6.171 /

0.819 /

5.973 3.31%↓

0.768 6.23%↓

5.887 4.60%↓

0.704 14.04%↓

5.822 5.66%↓

0.645 21.25%↓

Batch Size = 4 Gain vs. Baseline

16.638 /

0.957 /

16.236 2.48%↓

0.914 4.49%↓

15.971 4.01%↓

0.824 13.90%↓

15.635 6.03%↓

0.765 20.06%↓

Batch Size = 8 Gain vs. Baseline

27.961 /

1.024 /

27.381 2.12%↓

0.979 4.39%↓

26.863 3.93%↓

0.849 17.09%↓

26.477 5.31%↓

0.796 22.27%↓

Batch Size = 16 Gain vs. Baseline

46.273 /

1.377 /

45.278 2.20%↓

1.317 4.36%↓

44.703 3.39%↓

1.143 16.99%↓

44.100 4.70%↓

1.057 23.24%↓

TABLE III: Accuracy comparison of CAEE in the expert offloading scenario. Dataset

Baseline

θimp = 0.9

θimp = 0.8

θimp = 0.7

CEval MMLU GSM8K

0.8841 0.8513 0.9561

0.8796 (0.51%↓) 0.8464 (0.58%↓) 0.9529 (0.33%↓)

0.8789 (0.59%↓) 0.8440 (0.86%↓) 0.9498 (0.66%↓)

0.8662 (2.02%↓) 0.8398 (1.35%↓) 0.9404 (1.64%↓)

define the mandatory expert set, Amust , using a more selective, importance-based retention strategy. For an input token, the Top-K experts are sorted by their gating scores G(t)e . We then dynamically populate Amust by retaining the minimal number of top-ranked experts (m∗ ) whose Cumulative Importance Ratio Pm PK (CIR), defined as CIR(m) = i=1 G(t)ei / j=1 G(t)ej , satisfies a preset retention threshold θimp , i.e., CIR(m∗ ) ≥ θimp . Experts not retained in Amust are subsequently subjected to the cost-aware pruning mechanism, which targets low-importance contributions that would otherwise incur high H2D transfer costs. We evaluate the utility-latency trade-off by setting θimp ∈ {0.9, 0.8, 0.7} in our experiments. 1) Impact of Preset Retention Threshold θimp : Tables II and III present the performance and accuracy of CAEE. We observe the most pronounced gains in the TOPT metric, which governs steady-state throughput. By pruning of low-importance, high-cost experts, CAEE significantly reduces unnecessary pertoken data transfers, successfully addressing the PCIe link saturation. TOPT reduction ranges from 4.36% (θimp = 0.9) up to a maximum of 23.24% (θimp = 0.7) at Batch Size 16. While TTFT also improves (up to 6.03%), the dominant benefit is in the recurrent TOPT phase. Crucially, these substantial performance gains are achieved with minimal impact on accuracy, validating our cost-aware pruning and low-overhead compensation. The configuration θimp = 0.8 presents the most favorable trade-off: it consistently yields TOPT reductions between 13.90% and 17.09% across all batch sizes, while restricting accuracy degradation to less than 1% on MMLU (0.86%) and GSM8K (0.66%). The ability of CAEE to substantially improve throughput while maintaining high utility confirms its viability as a runtime optimization framework for transfer-bound MoE deployments. 2) Ablation Study on Low-overhead Compensation: When rerouting low-importance experts to those with smaller loading costs, we contrast a random selection strategy with the proposed low-overhead compensation mechanism; Fig. 9 confirms that

Accuracy comparison of different strategies.

Accuracy

0.950 0.945 0.940 Random Selection Low-Overhead Compensation

0.935 0.7

0.8 imp

0.9

Fig. 9: Accuracy comparison between random selection and low-overhead compensation across different θimp values.

our method consistently yields superior accuracy on the GSM8K dataset, validating its effectiveness. 3) Hardware Cost Observation: Fig. 10 illustrates the impact of CAEE on the system-level load imbalance by tracking the number of missed experts across different devices, a metric directly proportional to the data movement cost. The baseline configuration frequently exhibits high non-uniformity and prominent straggler devices, confirming the existence of the System-Level Imbalance (C2) problem. Crucially, the CAEE configuration actively minimizes the execution cost on the devices that represent the performance bottleneck. Across all layers, CAEE successfully reduces the transfer count on the maximum-cost device compared to the baseline, e.g., from 9 to 3 on Device 0 at Layer 30. By strategically pruning lowimportance experts only when they contribute to the system’s critical path latency, CAEE effectively enforces a more level cost distribution, thus directly solving the straggler problem and ensuring local savings contribute to system-wide throughput.

7

0

1

2

3

4

Device ID

5

6

7

0

0

1

Baseline Reroute

8 6 4 2 0

0

1

2

3

4

Device ID

3

4

Device ID

5

6

7

2 1 0

0

1

(b) Layer 8 Missed Expert Number

Missed Expert Number

(a) Layer 2

2

5

6

7

Baseline Reroute

8 6 4 2 0

0

1

(e) Layer 30

2

3

4

Device ID

2

3

4

Device ID

5

6

7

Baseline Reroute

6 4 2 0

0

(c) Layer 16 Missed Expert Number

0

Baseline Reroute

Missed Expert Number

1

2

3

5

6

7

Baseline Reroute

8 6 4 2 0

0

1

(f) Layer 32

2

3

4

Device ID

5

1

2

3

4

Device ID

5

6

7

6

7

(d) Layer 24

6

7

Missed Expert Number

2

Baseline Reroute

4

Missed Expert Number

Baseline Reroute

Missed Expert Number

Missed Expert Number

3

4

Baseline Reroute

3 2 1 0

0

(g) Layer 39

1

2

3

4

Device ID

5

(h) Layer 57

Fig. 10: Distribution of expert transfer load across 8 devices on selected MoE layers. The results show that CAEE successfully mitigates the system-level straggler effect by actively reducing the transfer count on the busiest device, compared to the baseline. TABLE IV: Performance of CAEE in the on-device inference scenario. TPOT (ms) is tested as a performance metric. θ

Baseline

0.30

0.40

0.50

0.60

0.70

0.80

0.90

1.0

1.1

1.2

Batch Size = 64 Gain vs. Baseline

59.29 /

58.46 1.40%↓

56.08 5.41%↓

55.72 6.01%↓

54.58 7.94%↓

53.15 10.36%↓

52.30 11.78%↓

51.66 12.87%↓

/ /

/ /

/ /

Batch Size = 128 Gain vs. Baseline

68.78 /

68.92 0.2%↑

66.63 3.13%↓

65.70 4.49%↓

64.55 6.15%↓

62.95 8.47%↓

62.18 9.59%↓

61.66 10.36%↓

60.79 11.63%↓

60.33 12.30%↓

59.42 13.62%↓

TABLE V: Accuracy comparison of CAEE in the on-device inference scenario. Dataset

BS

Baseline

0.30

0.40

0.50

0.60

0.70

0.80

0.90

1.0

1.1

1.2

CEval

64 128

0.8952 0.8952

0.8952 0.8952

0.8952 0.8952

0.8952 0.8952

0.8952 0.8952

0.8952 0.8952

0.8952 0.8952

0.8952 0.8952

/ 0.8952

/ 0.8952

/ 0.8952

MMLU

64 128

0.8683 0.8683

0.8683 0.8694

0.8683 0.8694

0.8683 0.8694

0.8683 0.8694

0.8683 0.8694

0.8683 0.8694

0.8683 0.8694

/ 0.8694

/ 0.8694

/ 0.8694

GSM8K

64 128

0.8992 0.8992

0.9128 0.9121

0.9151 0.9121

0.9030 0.9121

0.9121 0.9121

0.8961 0.9121

0.9037 0.9121

0.8878 0.9030

/ 0.8961

/ 0.8878

/ 0.8763

HumanEval

64 128

0.6707 0.6707

0.6829 0.7317

0.6829 0.7012

0.6951 0.7012

0.6768 0.7012

0.6646 0.6829

0.6341 0.6707

0.6829 0.6707

/ 0.6524

/ 0.6159

/ 0.6096

because DeepSeek’s group routing mechanism (designed for hardware compatibility) inherently limits routing choices; our method, however, breaks this limitation and allows for more flexible routing options that do not affect overall execution costs, potentially leading to better global utility. Our approach maintains near lossless performance on the relatively simpler MMLU and CEval benchmarks within the valid threshold range. However, significant accuracy degradation is observed on high-complexity reasoning benchmarks like GSM8K and HumanEval when the threshold exceeds 0.7 for 64 BS and 1.0 for 128 BS. Specifically, • Under a global concurrency of 64, at the θ = 0.6, we achieve a 7.94% improvement in end-to-end latency with no accuracy loss across all evaluated benchmarks. • When the global concurrency is increased to 128 and θ = 0.8, the system realizes an even greater 9.59% latency improvement, again with negligible accuracy degradation.

C. On-device Inference For thePon-device inference scenario, the expert importance I(e) = t G(t)e . We then define expert-score thresholds θ to identify candidate experts Acand for pruning. Tokens originally assigned to pruned experts are globally re-routed to higher-importance experts from Atrans via our compensation mechanism to mitigate potential accuracy degradation. Our experiments evaluate two global concurrency levels, i.e., 64 and 128, using input length of 2k and output length of 1k data. We select θ from 0.3 and 1.2 to examine the performanceaccuracy trade-off. Tables IV and V summarize the latency and accuracy results, respectively. As the expert scoring threshold increases, the token-level latency monotonically decreases, which is consistent with the cost reduction expected by our method. Regarding accuracy, we observed that in some cases, accuracy scores may slightly improve when the threshold is small. We speculate that this is

8

These results confirm that CAEE effectively targets and optimizes the on-device HBM pressure/memory bandwidth bottleneck, achieving substantial latency reduction while dynamically preserving model utility.

C. Expert Offloading Hardware-aware optimization, particularly expert offloading, leverages the sparse activation property of MoEs by caching frequently-accessed experts in HBM while evicting less active ones to DRAM or SSD, thus alleviating memory bottlenecks. To hide the associated I/O latency, prefetching and computational overlap are essential. Mixtral-Offloading [14] and ProMoE [30] preload experts via prediction, achieving over 90% hit rates; Fiddler [23] and MoE-Lightning [5] leverage CPU computation to overlap with GPU execution, effectively masking load times. Although these systems report 2×–8× inference speedups [23], [31], their general applicability remains questionable. They are typically evaluated in single-device or controlled settings and lack robustness in handling heterogeneous access delays and varied scenarios in production environments, thus limiting their cross-scenario optimization capability.

VI. R ELATED W ORK To address the computational and memory challenges in deploying Mixture-of-Experts (MoE) models, research has primarily focused on model compression and efficient inference systems [26]. Existing related approaches can be categorized into these complementary directions: static expert pruning, dynamic routing, and expert offloading. A. Static Expert Pruning

Expert pruning reduce model size and inference costs by eliminating parameter redundancy serving as fundamental VII. C ONCLUSION model-level optimizations for efficient MoE deployment [8], [17], [32]. These methods mainly include structured prunThis paper introduces CAEE, a novel runtime framework ing, unstructured pruning. Structured pruning focuses on the designed to mitigate two critical challenges in MoE deployment: removal or merging of entire experts. For expert removal, the Importance-Agnostic Cost and System-Level Imbalance. TSEP [8] prunes non-critical experts for specific downstream Leveraging a hardware-calibrated cost model, CAEE employs tasks; NAEE [27] and MoE-I² [34] eliminate redundant experts a cost-aware pruning strategy to selectively eliminate experts using calibration data and evolutionary search, respectively. For characterized by low token importance and high execution expert merging, DEK [36] and HC-SMoE [6] fuse experts via cost. Concurrently, a low-overhead compensation mechanism clustering, with the latter operating without retraining; LiteMoE is introduced to effectively maintain model accuracy following [38] tailors this approach for edge scenarios by merging token re-selection. Extensive experimental evaluation on the secondary experts to preserve core functionalities. Unstructured 671B DeepSeek-R1 model demonstrates the framework’s pruning targets fine-grained sparsity within experts. MoE- efficacy across diverse scenarios. In the expert offloading Pruner [32] assesses weight importance by combining input scenario, CAEE reduces the TOPT by up to 17.09% and activations and router weights; STUN [24] introduces a the TTFT by 3.93%, with negligible accuracy loss. For ontwo-stage “structured-to-unstructured” pruning pipeline; MoE- device inference, it achieves a TOPT reduction of 9.59%, also Compression [17] provides a unified framework integrating without sacrificing utility. Furthermore, hardware observations both paradigms. confirm CAEE’s success in directly alleviating the multi-device However, these methods suffer from limitations. Structured straggler problem. CAEE thus provides an effective, systempruning reduces expert diversity, potentially diminishing model aware solution for scalable and efficient MoE inference. capacity. Unstructured pruning results in irregular computation patterns that are hardware-unfriendly. Furthermore, many ACKNOWLEDGEMENTS approaches heavily rely on calibration data or task-specific We acknowledge the use of AI-powered language models for priors and incur significant offline optimization costs. proofreading, grammatical correction, and enhancing the clarity and fluency of the English language within this manuscript. B. Dynamic Routing We affirm that these tools were strictly used for polishing the Dynamic routing improve the inference efficiency of MoE presentation of the written work. All core intellectual content, models by adaptively selecting activated experts evenly [3], including the foundational ideas, problem formulation, system [25], [37]. At the routing algorithm level, research aims to design, experimental setup, analysis, and conclusions, remains design smarter gating mechanisms. DA-MoE [3] presents a the original and sole work of the authors. novel method for dynamic expert allocation in MoE models, R EFERENCES driven by token importance. [25] adaptively allocates experts based on probability distributions, reducing computation by [1] Compute Express Link. [Online]. Available: https://en.wikipedia.org/ 38.2%; DynMoE [16] proposes a top-down gating strategy for wiki/Compute Express Link [2] PCI Express. [Online]. Available: https://en.wikipedia.org/wiki/PCI fine-grained expert assignment; XMoE [35] employs threshold Express control to dynamically activate experts. Despite their benefits, [3] M. A. Aghdam, H. Jin, and Y. Wu, “Da-MoE: Towards Dynamic these dynamic methods typically ignore per-expert hardware Expert Allocation for Mixture-of-Experts Models,” arXiv preprint cost or fail to evaluate system-wide, no-linear straggler impact. arXiv:2409.06669, 2024.

9

[4] T. B. Brown, B. Mann, N. Ryder, M. Subbiah, J. Kaplan, P. Dhariwal, A. Neelakantan, P. Shyam, G. Sastry, A. Askell, S. Agarwal, A. Herbert-Voss, G. Krueger, T. Henighan, R. Child, A. Ramesh, D. M. Ziegler, J. Wu, C. Winter, C. Hesse, M. Chen, E. Sigler, M. Litwin, S. Gray, B. Chess, J. Clark, C. Berner, S. McCandlish, A. Radford, I. Sutskever, and D. Amodei, “Language Models are Few-Shot Learners,” 2020. [Online]. Available: https://arxiv.org/abs/2005.14165 [5] S. Cao, S. Liu, T. Griggs, P. Schafhalter, X. Liu, Y. Sheng, J. E. Gonzalez, M. Zaharia, and I. Stoica, “MoE-Lightning: High-Throughput MoE Inference on Memory-Constrained Gpus,” in Proceedings of the 30th ACM International Conference on Architectural Support for Programming Languages and Operating Systems, Volume 1, 2025, pp. 715–730. [6] I.-C. Chen, H.-S. Liu, W.-F. Sun, C.-H. Chao, Y.-C. Hsu, and C.-Y. Lee, “Retraining-Free Merging of Sparse Mixture-of-Experts via Hierarchical Clustering,” 2024. [7] M. Chen, J. Tworek, H. Jun, Q. Yuan, H. P. de Oliveira Pinto, J. Kaplan, H. Edwards, Y. Burda, N. Joseph, G. Brockman, A. Ray, R. Puri, G. Krueger, M. Petrov, H. Khlaaf, G. Sastry, P. Mishkin, B. Chan, S. Gray, N. Ryder, M. Pavlov, A. Power, L. Kaiser, M. Bavarian, C. Winter, P. Tillet, F. P. Such, D. Cummings, M. Plappert, F. Chantzis, E. Barnes, A. Herbert-Voss, W. H. Guss, A. Nichol, A. Paino, N. Tezak, J. Tang, I. Babuschkin, S. Balaji, S. Jain, W. Saunders, C. Hesse, A. N. Carr, J. Leike, J. Achiam, V. Misra, E. Morikawa, A. Radford, M. Knight, M. Brundage, M. Murati, K. Mayer, P. Welinder, B. McGrew, D. Amodei, S. McCandlish, I. Sutskever, and W. Zaremba, “Evaluating Large Language Models Trained on Code,” 2021. [Online]. Available: https://arxiv.org/abs/2107.03374 [8] T. Chen, S. Huang, Y. Xie, B. Jiao, D. Jiang, H. Zhou, J. Li, and F. Wei, “Task-Specific Expert Pruning for Sparse Mixture-of-Experts,” arXiv preprint arXiv:2206.00277, 2022. [9] K. Cobbe, V. Kosaraju, M. Bavarian, M. Chen, H. Jun, L. Kaiser, M. Plappert, J. Tworek, J. Hilton, R. Nakano, C. Hesse, and J. Schulman, “Training Verifiers to Solve Math Word Problems,” 2021. [Online]. Available: https://arxiv.org/abs/2110.14168 [10] M. Davies, N. Crago, K. Sankaralingam, and C. Kozyrakis, “Efficient LLM Inference: Bandwidth, Compute, Synchronization, and Capacity are all You Need,” arXiv preprint arXiv:2507.14397, 2025. [11] DeepSeek-AI, D. Guo, D. Yang, H. Zhang, J. Song, R. Zhang, R. Xu, Q. Zhu, S. Ma, P. Wang, X. Bi, X. Zhang, X. Yu, Y. Wu, Z. F. Wu, Z. Gou, Z. Shao, Z. Li, Z. Gao, A. Liu, B. Xue, B. Wang, B. Wu, B. Feng, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Ding, H. Xin, H. Gao, H. Qu, H. Li, J. Guo, J. Li, J. Wang, J. Chen, J. Yuan, J. Qiu, J. Li, J. L. Cai, J. Ni, J. Liang, J. Chen, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Zhao, L. Wang, L. Zhang, L. Xu, L. Xia, M. Zhang, M. Zhang, M. Tang, M. Li, M. Wang, M. Li, N. Tian, P. Huang, P. Zhang, Q. Wang, Q. Chen, Q. Du, R. Ge, R. Zhang, R. Pan, R. Wang, R. J. Chen, R. L. Jin, R. Chen, S. Lu, S. Zhou, S. Chen, S. Ye, S. Wang, S. Yu, S. Zhou, S. Pan, S. S. Li, S. Zhou, S. Wu, S. Ye, T. Yun, T. Pei, T. Sun, T. Wang, W. Zeng, W. Zhao, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, W. L. Xiao, W. An, X. Liu, X. Wang, X. Chen, X. Nie, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yang, X. Li, X. Su, X. Lin, X. Q. Li, X. Jin, X. Shen, X. Chen, X. Sun, X. Wang, X. Song, X. Zhou, X. Wang, X. Shan, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. Zhang, Y. Xu, Y. Li, Y. Zhao, Y. Sun, Y. Wang, Y. Yu, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Ou, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Xiong, Y. Luo, Y. You, Y. Liu, Y. Zhou, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zheng, Y. Zhu, Y. Ma, Y. Tang, Y. Zha, Y. Yan, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Xie, Z. Zhang, Z. Hao, Z. Ma, Z. Yan, Z. Wu, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Pan, Z. Huang, Z. Xu, Z. Zhang, and Z. Zhang, “DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning,” 2025. [Online]. Available: https://arxiv.org/abs/2501.12948 [12] DeepSeek-AI, A. Liu, B. Feng, B. Wang, B. Wang, B. Liu, C. Zhao, C. Dengr, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Xu, H. Yang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Chen, J. Yuan, J. Qiu, J. Song, K. Dong, K. Gao, K. Guan, L. Wang, L. Zhang, L. Xu, L. Xia, L. Zhao, L. Zhang, M. Li, M. Wang, M. Zhang, M. Zhang, M. Tang, M. Li, N. Tian, P. Huang, P. Wang, P. Zhang, Q. Zhu, Q. Chen, Q. Du, R. J. Chen, R. L. Jin, R. Ge, R. Pan, R. Xu, R. Chen, S. S. Li, S. Lu, S. Zhou, S. Chen,

10

S. Wu, S. Ye, S. Ma, S. Wang, S. Zhou, S. Yu, S. Zhou, S. Zheng, T. Wang, T. Pei, T. Yuan, T. Sun, W. L. Xiao, W. Zeng, W. An, W. Liu, W. Liang, W. Gao, W. Zhang, X. Q. Li, X. Jin, X. Wang, X. Bi, X. Liu, X. Wang, X. Shen, X. Chen, X. Chen, X. Nie, X. Sun, X. Wang, X. Liu, X. Xie, X. Yu, X. Song, X. Zhou, X. Yang, X. Lu, X. Su, Y. Wu, Y. K. Li, Y. X. Wei, Y. X. Zhu, Y. Xu, Y. Huang, Y. Li, Y. Zhao, Y. Sun, Y. Li, Y. Wang, Y. Zheng, Y. Zhang, Y. Xiong, Y. Zhao, Y. He, Y. Tang, Y. Piao, Y. Dong, Y. Tan, Y. Liu, Y. Wang, Y. Guo, Y. Zhu, Y. Wang, Y. Zou, Y. Zha, Y. Ma, Y. Yan, Y. You, Y. Liu, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Huang, Z. Zhang, Z. Xie, Z. Hao, Z. Shao, Z. Wen, Z. Xu, Z. Zhang, Z. Li, Z. Wang, Z. Gu, Z. Li, and Z. Xie, “DeepSeek-V2: A Strong, Economical, and Efficient Mixture-of-Experts Language Model,” 2024. [Online]. Available: https://arxiv.org/abs/2405.04434 [13] DeepSeek-AI, A. Liu, B. Feng, B. Xue, B. Wang, B. Wu, C. Lu, C. Zhao, C. Deng, C. Zhang, C. Ruan, D. Dai, D. Guo, D. Yang, D. Chen, D. Ji, E. Li, F. Lin, F. Dai, F. Luo, G. Hao, G. Chen, G. Li, H. Zhang, H. Bao, H. Xu, H. Wang, H. Zhang, H. Ding, H. Xin, H. Gao, H. Li, H. Qu, J. L. Cai, J. Liang, J. Guo, J. Ni, J. Li, J. Wang, J. Chen, J. Chen, J. Yuan, J. Qiu, J. Li, J. Song, K. Dong, K. Hu, K. Gao, K. Guan, K. Huang, K. Yu, L. Wang, L. Zhang, L. Xu, L. Xia, L. Zhao, L. Wang, L. Zhang, M. Li, M. Wang, M. Zhang, M. Zhang, M. Tang, M. Li, N. Tian, P. Huang, P. Wang, P. Zhang, Q. Wang, Q. Zhu, Q. Chen, Q. Du, R. J. Chen, R. L. Jin, R. Ge, R. Zhang, R. Pan, R. Wang, R. Xu, R. Zhang, R. Chen, S. S. Li, S. Lu, S. Zhou, S. Chen, S. Wu, S. Ye, S. Ye, S. Ma, S. Wang, S. Zhou, S. Yu, S. Zhou, S. Pan, T. Wang, T. Yun, T. Pei, T. Sun, W. L. Xiao, W. Zeng, W. Zhao, W. An, W. Liu, W. Liang, W. Gao, W. Yu, W. Zhang, X. Q. Li, X. Jin, X. Wang, X. Bi, X. Liu, X. Wang, X. Shen, X. Chen, X. Zhang, X. Chen, X. Nie, X. Sun, X. Wang, X. Cheng, X. Liu, X. Xie, X. Liu, X. Yu, X. Song, X. Shan, X. Zhou, X. Yang, X. Li, X. Su, X. Lin, Y. K. Li, Y. Q. Wang, Y. X. Wei, Y. X. Zhu, Y. Zhang, Y. Xu, Y. Xu, Y. Huang, Y. Li, Y. Zhao, Y. Sun, Y. Li, Y. Wang, Y. Yu, Y. Zheng, Y. Zhang, Y. Shi, Y. Xiong, Y. He, Y. Tang, Y. Piao, Y. Wang, Y. Tan, Y. Ma, Y. Liu, Y. Guo, Y. Wu, Y. Ou, Y. Zhu, Y. Wang, Y. Gong, Y. Zou, Y. He, Y. Zha, Y. Xiong, Y. Ma, Y. Yan, Y. Luo, Y. You, Y. Liu, Y. Zhou, Z. F. Wu, Z. Z. Ren, Z. Ren, Z. Sha, Z. Fu, Z. Xu, Z. Huang, Z. Zhang, Z. Xie, Z. Zhang, Z. Hao, Z. Gou, Z. Ma, Z. Yan, Z. Shao, Z. Xu, Z. Wu, Z. Zhang, Z. Li, Z. Gu, Z. Zhu, Z. Liu, Z. Li, Z. Xie, Z. Song, Z. Gao, and Z. Pan, “DeepSeek-V3 Technical Report,” 2025. [Online]. Available: https://arxiv.org/abs/2412.19437 [14] A. Eliseev and D. Mazur, “Fast Inference of Mixture-of-Experts Language Models with Offloading,” arXiv preprint arXiv:2312.17238, 2023. [15] A. Grattafiori, A. Dubey, A. Jauhri, A. Pandey, A. Kadian, A. Al-Dahle, A. Letman, A. Mathur, A. Schelten, A. Vaughan, A. Yang, A. Fan, A. Goyal, A. Hartshorn, A. Yang, A. Mitra, A. Sravankumar, A. Korenev, A. Hinsvark, A. Rao, A. Zhang, A. Rodriguez, A. Gregerson, A. Spataru, B. Roziere, B. Biron, B. Tang, B. Chern, C. Caucheteux, C. Nayak, C. Bi, C. Marra, C. McConnell, C. Keller, C. Touret, C. Wu, C. Wong, C. C. Ferrer, C. Nikolaidis, D. Allonsius, D. Song, D. Pintz, D. Livshits, D. Wyatt, D. Esiobu, D. Choudhary, D. Mahajan, D. Garcia-Olano, D. Perino, D. Hupkes, E. Lakomkin, E. AlBadawy, E. Lobanova, E. Dinan, E. M. Smith, F. Radenovic, F. Guzmán, F. Zhang, G. Synnaeve, G. Lee, G. L. Anderson, G. Thattai, G. Nail, G. Mialon, G. Pang, G. Cucurell, H. Nguyen, H. Korevaar, H. Xu, H. Touvron, I. Zarov, I. A. Ibarra, I. Kloumann, I. Misra, I. Evtimov, J. Zhang, J. Copet, J. Lee, J. Geffert, J. Vranes, J. Park, J. Mahadeokar, J. Shah, J. van der Linde, J. Billock, J. Hong, J. Lee, J. Fu, J. Chi, J. Huang, J. Liu, J. Wang, J. Yu, J. Bitton, J. Spisak, J. Park, J. Rocca, J. Johnstun, J. Saxe, J. Jia, K. V. Alwala, K. Prasad, K. Upasani, K. Plawiak, K. Li, K. Heafield, K. Stone, K. El-Arini, K. Iyer, K. Malik, K. Chiu, K. Bhalla, K. Lakhotia, L. Rantala-Yeary, L. van der Maaten, L. Chen, L. Tan, L. Jenkins, L. Martin, L. Madaan, L. Malo, L. Blecher, L. Landzaat, L. de Oliveira, M. Muzzi, M. Pasupuleti, M. Singh, M. Paluri, M. Kardas, M. Tsimpoukelli, M. Oldham, M. Rita, M. Pavlova, M. Kambadur, M. Lewis, M. Si, M. K. Singh, M. Hassan, N. Goyal, N. Torabi, N. Bashlykov, N. Bogoychev, N. Chatterji, N. Zhang, O. Duchenne, O. Çelebi, P. Alrassy, P. Zhang, P. Li, P. Vasic, P. Weng, P. Bhargava, P. Dubal, P. Krishnan, P. S. Koura, P. Xu, Q. He, Q. Dong, R. Srinivasan, R. Ganapathy, R. Calderer, R. S. Cabral, R. Stojnic, R. Raileanu, R. Maheswari, R. Girdhar, R. Patel, R. Sauvestre, R. Polidoro, R. Sumbaly, R. Taylor, R. Silva, R. Hou, R. Wang, S. Hosseini, S. Chennabasappa, S. Singh, S. Bell, S. S. Kim, S. Edunov, S. Nie, S. Narang, S. Raparthy, S. Shen, S. Wan, S. Bhosale, S. Zhang, S. Vandenhende, S. Batra, S. Whitman, S. Sootla, S. Collot,

S. Gururangan, S. Borodinsky, T. Herman, T. Fowler, T. Sheasha, T. Georgiou, T. Scialom, T. Speckbacher, T. Mihaylov, T. Xiao, U. Karn, V. Goswami, V. Gupta, V. Ramanathan, V. Kerkez, V. Gonguet, V. Do, V. Vogeti, V. Albiero, V. Petrovic, W. Chu, W. Xiong, W. Fu, W. Meers, X. Martinet, X. Wang, X. Wang, X. E. Tan, X. Xia, X. Xie, X. Jia, X. Wang, Y. Goldschlag, Y. Gaur, Y. Babaei, Y. Wen, Y. Song, Y. Zhang, Y. Li, Y. Mao, Z. D. Coudert, Z. Yan, Z. Chen, Z. Papakipos, A. Singh, A. Srivastava, A. Jain, A. Kelsey, A. Shajnfeld, A. Gangidi, A. Victoria, A. Goldstand, A. Menon, A. Sharma, A. Boesenberg, A. Baevski, A. Feinstein, A. Kallet, A. Sangani, A. Teo, A. Yunus, A. Lupu, A. Alvarado, A. Caples, A. Gu, A. Ho, A. Poulton, A. Ryan, A. Ramchandani, A. Dong, A. Franco, A. Goyal, A. Saraf, A. Chowdhury, A. Gabriel, A. Bharambe, A. Eisenman, A. Yazdan, B. James, B. Maurer, B. Leonhardi, B. Huang, B. Loyd, B. D. Paola, B. Paranjape, B. Liu, B. Wu, B. Ni, B. Hancock, B. Wasti, B. Spence, B. Stojkovic, B. Gamido, B. Montalvo, C. Parker, C. Burton, C. Mejia, C. Liu, C. Wang, C. Kim, C. Zhou, C. Hu, C.-H. Chu, C. Cai, C. Tindal, C. Feichtenhofer, C. Gao, D. Civin, D. Beaty, D. Kreymer, D. Li, D. Adkins, D. Xu, D. Testuggine, D. David, D. Parikh, D. Liskovich, D. Foss, D. Wang, D. Le, D. Holland, E. Dowling, E. Jamil, E. Montgomery, E. Presani, E. Hahn, E. Wood, E.-T. Le, E. Brinkman, E. Arcaute, E. Dunbar, E. Smothers, F. Sun, F. Kreuk, F. Tian, F. Kokkinos, F. Ozgenel, F. Caggioni, F. Kanayet, F. Seide, G. M. Florez, G. Schwarz, G. Badeer, G. Swee, G. Halpern, G. Herman, G. Sizov, Guangyi, Zhang, G. Lakshminarayanan, H. Inan, H. Shojanazeri, H. Zou, H. Wang, H. Zha, H. Habeeb, H. Rudolph, H. Suk, H. Aspegren, H. Goldman, H. Zhan, I. Damlaj, I. Molybog, I. Tufanov, I. Leontiadis, I.-E. Veliche, I. Gat, J. Weissman, J. Geboski, J. Kohli, J. Lam, J. Asher, J.-B. Gaya, J. Marcus, J. Tang, J. Chan, J. Zhen, J. Reizenstein, J. Teboul, J. Zhong, J. Jin, J. Yang, J. Cummings, J. Carvill, J. Shepard, J. McPhie, J. Torres, J. Ginsburg, J. Wang, K. Wu, K. H. U, K. Saxena, K. Khandelwal, K. Zand, K. Matosich, K. Veeraraghavan, K. Michelena, K. Li, K. Jagadeesh, K. Huang, K. Chawla, K. Huang, L. Chen, L. Garg, L. A, L. Silva, L. Bell, L. Zhang, L. Guo, L. Yu, L. Moshkovich, L. Wehrstedt, M. Khabsa, M. Avalani, M. Bhatt, M. Mankus, M. Hasson, M. Lennie, M. Reso, M. Groshev, M. Naumov, M. Lathi, M. Keneally, M. Liu, M. L. Seltzer, M. Valko, M. Restrepo, M. Patel, M. Vyatskov, M. Samvelyan, M. Clark, M. Macey, M. Wang, M. J. Hermoso, M. Metanat, M. Rastegari, M. Bansal, N. Santhanam, N. Parks, N. White, N. Bawa, N. Singhal, N. Egebo, N. Usunier, N. Mehta, N. P. Laptev, N. Dong, N. Cheng, O. Chernoguz, O. Hart, O. Salpekar, O. Kalinli, P. Kent, P. Parekh, P. Saab, P. Balaji, P. Rittner, P. Bontrager, P. Roux, P. Dollar, P. Zvyagina, P. Ratanchandani, P. Yuvraj, Q. Liang, R. Alao, R. Rodriguez, R. Ayub, R. Murthy, R. Nayani, R. Mitra, R. Parthasarathy, R. Li, R. Hogan, R. Battey, R. Wang, R. Howes, R. Rinott, S. Mehta, S. Siby, S. J. Bondu, S. Datta, S. Chugh, S. Hunt, S. Dhillon, S. Sidorov, S. Pan, S. Mahajan, S. Verma, S. Yamamoto, S. Ramaswamy, S. Lindsay, S. Lindsay, S. Feng, S. Lin, S. C. Zha, S. Patil, S. Shankar, S. Zhang, S. Zhang, S. Wang, S. Agarwal, S. Sajuyigbe, S. Chintala, S. Max, S. Chen, S. Kehoe, S. Satterfield, S. Govindaprasad, S. Gupta, S. Deng, S. Cho, S. Virk, S. Subramanian, S. Choudhury, S. Goldman, T. Remez, T. Glaser, T. Best, T. Koehler, T. Robinson, T. Li, T. Zhang, T. Matthews, T. Chou, T. Shaked, V. Vontimitta, V. Ajayi, V. Montanez, V. Mohan, V. S. Kumar, V. Mangla, V. Ionescu, V. Poenaru, V. T. Mihailescu, V. Ivanov, W. Li, W. Wang, W. Jiang, W. Bouaziz, W. Constable, X. Tang, X. Wu, X. Wang, X. Wu, X. Gao, Y. Kleinman, Y. Chen, Y. Hu, Y. Jia, Y. Qi, Y. Li, Y. Zhang, Y. Zhang, Y. Adi, Y. Nam, Yu, Wang, Y. Zhao, Y. Hao, Y. Qian, Y. Li, Y. He, Z. Rait, Z. DeVito, Z. Rosnbrick, Z. Wen, Z. Yang, Z. Zhao, and Z. Ma, “The Llama 3 Herd of Models,” 2024. [Online]. Available: https://arxiv.org/abs/2407.21783 [16] Y. Guo, Z. Cheng, X. Tang, Z. Tu, and T. Lin, “Dynamic Mixture of Experts: An Auto-Tuning Approach for Efficient Transformer Models,” arXiv preprint arXiv:2405.14297, 2024. [17] S. He, D. Dong, L. Ding, and A. Li, “Demystifying the Compression of Mixture-of-Experts through a Unified Framework,” arXiv e-prints, pp. arXiv–2406, 2024. [18] D. Hendrycks, C. Burns, S. Basart, A. Zou, M. Mazeika, D. Song, and J. Steinhardt, “Measuring Massive Multitask Language Understanding,” arXiv preprint arXiv:2009.03300, 2020. [19] H. Huang, N. Ardalani, A. Sun, L. Ke, H.-H. S. Lee, A. Sridhar, S. Bhosale, C.-J. Wu, and B. Lee, “Towards MoE Deployment: Mitigating Inefficiencies in Mixture-of-Expert (MoE) Inference,” arXiv preprint arXiv:2303.06182, 2023.

11

[20] Y. Huang, Y. Bai, Z. Zhu, J. Zhang, J. Zhang, T. Su, J. Liu, C. Lv, Y. Zhang, j. lei, Y. Fu, M. Sun, and J. He, “C-Eval: A Multi-Level MultiDiscipline Chinese Evaluation Suite for Foundation Models,” Advances in Neural Information Processing Systems, vol. 36, pp. 62 991–63 010, 2023. [21] R. Hwang, J. Wei, S. Cao, C. Hwang, X. Tang, T. Cao, and M. Yang, “Pre-Gated MoE: An Algorithm-System Co-Design for Fast and Scalable Mixture-of-Expert inference,” in 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA). IEEE, 2024, pp. 1018–1031. [22] A. Q. Jiang, A. Sablayrolles, A. Roux, A. Mensch, B. Savary, C. Bamford, D. S. Chaplot, D. de las Casas, E. B. Hanna, F. Bressand, G. Lengyel, G. Bour, G. Lample, L. R. Lavaud, L. Saulnier, M.-A. Lachaux, P. Stock, S. Subramanian, S. Yang, S. Antoniak, T. L. Scao, T. Gervet, T. Lavril, T. Wang, T. Lacroix, and W. E. Sayed, “Mixtral of Experts,” 2024. [Online]. Available: https://arxiv.org/abs/2401.04088 [23] K. Kamahori, T. Tang, Y. Gu, K. Zhu, and B. Kasikci, “Fiddler: Cpu-gpu Orchestration for Fast Inference of Mixture-of-Experts Models,” arXiv preprint arXiv:2402.07033, 2024. [24] J. Lee, S.-w. Hwang, A. Qiao, D. F. Campos, Z. Yao, and Y. He, “Stun: Structured-then-Unstructured Pruning for Scalable MoE Pruning,” in Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), 2025, pp. 13 660– 13 676. [25] J. Li, Q. Su, Y. Yang, Y. Jiang, C. Wang, and H. Xu, “Adaptive Gating in Mixture-of-Experts based Language Models,” arXiv preprint arXiv:2310.07188, 2023. [26] J. Liu, P. Tang, W. Wang, Y. Ren, X. Hou, P.-A. Heng, M. Guo, and C. Li, “A Survey on Inference Optimization Techniques for Mixture of Experts Models,” arXiv preprint arXiv:2412.14219, 2024. [27] X. Lu, Q. Liu, Y. Xu, A. Zhou, S. Huang, B. Zhang, J. Yan, and H. Li, “Not All Experts Are Equal: Efficient Expert Pruning and Skipping for Mixture-of-Experts Large Language Models,” arXiv preprint arXiv:2402.14800, 2024. [28] OpenAI, J. Achiam, S. Adler, S. Agarwal, L. Ahmad, I. Akkaya, F. L. Aleman, D. Almeida, J. Altenschmidt, S. Altman, S. Anadkat, R. Avila, I. Babuschkin, S. Balaji, V. Balcom, P. Baltescu, H. Bao, M. Bavarian, J. Belgum, I. Bello, J. Berdine, G. Bernadett-Shapiro, C. Berner, L. Bogdonoff, O. Boiko, M. Boyd, A.-L. Brakman, G. Brockman, T. Brooks, M. Brundage, K. Button, T. Cai, R. Campbell, A. Cann, B. Carey, C. Carlson, R. Carmichael, B. Chan, C. Chang, F. Chantzis, D. Chen, S. Chen, R. Chen, J. Chen, M. Chen, B. Chess, C. Cho, C. Chu, H. W. Chung, D. Cummings, J. Currier, Y. Dai, C. Decareaux, T. Degry, N. Deutsch, D. Deville, A. Dhar, D. Dohan, S. Dowling, S. Dunning, A. Ecoffet, A. Eleti, T. Eloundou, D. Farhi, L. Fedus, N. Felix, S. P. Fishman, J. Forte, I. Fulford, L. Gao, E. Georges, C. Gibson, V. Goel, T. Gogineni, G. Goh, R. Gontijo-Lopes, J. Gordon, M. Grafstein, S. Gray, R. Greene, J. Gross, S. S. Gu, Y. Guo, C. Hallacy, J. Han, J. Harris, Y. He, M. Heaton, J. Heidecke, C. Hesse, A. Hickey, W. Hickey, P. Hoeschele, B. Houghton, K. Hsu, S. Hu, X. Hu, J. Huizinga, S. Jain, S. Jain, J. Jang, A. Jiang, R. Jiang, H. Jin, D. Jin, S. Jomoto, B. Jonn, H. Jun, T. Kaftan, Łukasz Kaiser, A. Kamali, I. Kanitscheider, N. S. Keskar, T. Khan, L. Kilpatrick, J. W. Kim, C. Kim, Y. Kim, J. H. Kirchner, J. Kiros, M. Knight, D. Kokotajlo, Łukasz Kondraciuk, A. Kondrich, A. Konstantinidis, K. Kosic, G. Krueger, V. Kuo, M. Lampe, I. Lan, T. Lee, J. Leike, J. Leung, D. Levy, C. M. Li, R. Lim, M. Lin, S. Lin, M. Litwin, T. Lopez, R. Lowe, P. Lue, A. Makanju, K. Malfacini, S. Manning, T. Markov, Y. Markovski, B. Martin, K. Mayer, A. Mayne, B. McGrew, S. M. McKinney, C. McLeavey, P. McMillan, J. McNeil, D. Medina, A. Mehta, J. Menick, L. Metz, A. Mishchenko, P. Mishkin, V. Monaco, E. Morikawa, D. Mossing, T. Mu, M. Murati, O. Murk, D. Mély, A. Nair, R. Nakano, R. Nayak, A. Neelakantan, R. Ngo, H. Noh, L. Ouyang, C. O’Keefe, J. Pachocki, A. Paino, J. Palermo, A. Pantuliano, G. Parascandolo, J. Parish, E. Parparita, A. Passos, M. Pavlov, A. Peng, A. Perelman, F. de Avila Belbute Peres, M. Petrov, H. P. de Oliveira Pinto, Michael, Pokorny, M. Pokrass, V. H. Pong, T. Powell, A. Power, B. Power, E. Proehl, R. Puri, A. Radford, J. Rae, A. Ramesh, C. Raymond, F. Real, K. Rimbach, C. Ross, B. Rotsted, H. Roussez, N. Ryder, M. Saltarelli, T. Sanders, S. Santurkar, G. Sastry, H. Schmidt, D. Schnurr, J. Schulman, D. Selsam, K. Sheppard, T. Sherbakov, J. Shieh, S. Shoker, P. Shyam, S. Sidor, E. Sigler, M. Simens, J. Sitkin, K. Slama, I. Sohl, B. Sokolowsky, Y. Song, N. Staudacher, F. P. Such, N. Summers, I. Sutskever, J. Tang, N. Tezak, M. B. Thompson,

P. Tillet, A. Tootoonchian, E. Tseng, P. Tuggle, N. Turley, J. Tworek, J. F. C. Uribe, A. Vallone, A. Vijayvergiya, C. Voss, C. Wainwright, J. J. Wang, A. Wang, B. Wang, J. Ward, J. Wei, C. Weinmann, A. Welihinda, P. Welinder, J. Weng, L. Weng, M. Wiethoff, D. Willner, C. Winter, S. Wolrich, H. Wong, L. Workman, S. Wu, J. Wu, M. Wu, K. Xiao, T. Xu, S. Yoo, K. Yu, Q. Yuan, W. Zaremba, R. Zellers, C. Zhang, M. Zhang, S. Zhao, T. Zheng, J. Zhuang, W. Zhuk, and B. Zoph, “GPT-4 Technical Report,” 2024. [Online]. Available: https://arxiv.org/abs/2303.08774 [29] S. Rajbhandari, C. Li, Z. Yao, M. Zhang, R. Y. Aminabadi, A. A. Awan, J. Rasley, and Y. He, “Deepspeed-MoE: Advancing Mixture-ofExperts Inference and Training to Power Next-Generation AI Scale,” in International conference on machine learning. PMLR, 2022, pp. 18 332–18 346. [30] X. Song, Z. Zhong, R. Chen, and H. Chen, “ProMoE: Fast MoEBased LLM Serving Using Proactive Caching,” arXiv preprint arXiv:2410.22134, 2024. [31] P. Tang, J. Liu, X. Hou, Y. Pu, J. Wang, P.-A. Heng, C. Li, and M. Guo, “Hobbit: A Mixed Precision Expert Offloading System for Fast MoE Inference,” arXiv preprint arXiv:2411.01433, 2024. [32] Y. Xie, Z. Zhang, D. Zhou, C. Xie, Z. Song, X. Liu, Y. Wang, X. Lin, and A. Xu, “MoE-Pruner: Pruning Mixture-of-Experts Large Language Model Using the Hints from its Router,” arXiv preprint arXiv:2410.12013, 2024. [33] A. Yang, A. Li, B. Yang, B. Zhang, B. Hui, B. Zheng, B. Yu, C. Gao, C. Huang, C. Lv, C. Zheng, D. Liu, F. Zhou, F. Huang, F. Hu, H. Ge, H. Wei, H. Lin, J. Tang, J. Yang, J. Tu, J. Zhang, J. Yang, J. Yang, J. Zhou, J. Zhou, J. Lin, K. Dang, K. Bao, K. Yang, L. Yu, L. Deng, M. Li, M. Xue, M. Li, P. Zhang, P. Wang, Q. Zhu, R. Men, R. Gao, S. Liu, S. Luo, T. Li, T. Tang, W. Yin, X. Ren, X. Wang, X. Zhang, X. Ren, Y. Fan, Y. Su, Y. Zhang, Y. Zhang, Y. Wan, Y. Liu, Z. Wang, Z. Cui, Z. Zhang, Z. Zhou, and Z. Qiu, “Qwen3 Technical Report,” 2025. [Online]. Available: https://arxiv.org/abs/2505.09388 [34] C. Yang, Y. Sui, J. Xiao, L. Huang, Y. Gong, Y. Duan, W. Jia, M. Yin, Y. Cheng, and B. Yuan, “MoE-I2 : Compressing Mixture of Experts Models through Inter-Expert Pruning and Intra-Expert Low-Rank Decomposition,” arXiv preprint arXiv:2411.01016, 2024. [35] Y. Yang, S. Qi, W. Gu, C. Wang, C. Gao, and Z. Xu, “XMoE: Sparse Models with Fine-Grained and Adaptive Expert Selection,” arXiv preprint arXiv:2403.18926, 2024. [36] Z. Zhang, X. Liu, H. Cheng, C. Xu, and J. Gao, “Diversifying the Expert Knowledge for Task-Agnostic Pruning in Sparse Mixture-of-Experts,” in Findings of the Association for Computational Linguistics: ACL 2025, 2025, pp. 86–102. [37] S. Zhong, L. Liang, Y. Wang, R. Wang, R. Huang, and M. Li, “AdapMoE: Adaptive Sensitivity-Based Expert Gating and Management for Efficient MoE Inference,” in Proceedings of the 43rd IEEE/ACM International Conference on Computer-Aided Design, 2024, pp. 1–9. [38] Y. Zhuang, Z. Zheng, F. Wu, and G. Chen, “LiteMoE: Customizing On-Device LLM Serving via Proxy Submodel Tuning,” in Proceedings of the 22nd ACM Conference on Embedded Networked Sensor Systems, 2024, pp. 521–534.

12

Record · ID 321801 · SHA-256 8666f3e1dfb0605f
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.