Conceptio › Archive › arXiv CS
arXiv CSopen access

Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

Weave: Fine-Grained Dynamic SM Scheduling in an MoE Megakernel for Compute-Communication Overlap Ziyu Huang∗1 , Yangjie Zhou∗2 , Chenhao Zhu1 , Zihan Liu1 , Jinyu Liu1 , Shulai Zhang1 , Xingxun Tang1 , Hongzhe Yan1 , Xinhao Luo1 , Minyi Guo1 , Xiu Lin3 , Yinghao Yu3 , Guodong Yang3 , Liping Zhang3 , Shixuan Sun1 , Jingwen Leng1 1 Shanghai Jiao Tong University

arXiv:2609.21483v1 [cs.DC] 18 Sep 2026

Shanghai, China

2 National University of Singapore

Singapore

3 Alibaba Group

China [email protected]

Abstract

1

Mixture-of-Experts (MoE) inference under expert parallelism (EP) turns each MoE layer into a distributed computation with costly dispatch and combine communication. State-of-the-art systems reduce this cost through communication-computation overlap, splitting the GPU’s SMs for communication and computation respectively. However, this approach still leaves GPU resources wasted along two dimensions. Spatially, the best SM split is determined by each layer’s routing result and varies across layers and GPUs, so fixed policies mismatch the workload and waste either NVLink bandwidth or compute throughput. Temporally, complex MoE data dependencies introduce bubbles that leave SMs idle. We present Weave, to our knowledge the first MoE overlap system that performs fine-grained dynamic SM scheduling—deciding per layer and per GPU by routing results at runtime. Once routing completes, each layer’s communication and computation volumes become known; Weave exploits this predictability through a lightweight cost model running inside the persistent megakernel: a spatial scheduler partitions SMs into communication workers and computation workers to match the communication/computation throughput ratio, and a temporal scheduler coordinates the two worker groups to minimize SM idleness. On 4×H100 SXM GPUs across six mainstream MoE models, Weave achieves a 2.89× geometric-mean MoElayer speedup and a 1.33× geometric-mean end-to-end speedup over five state-of-the-art baselines.

Mixture-of-Experts (MoE) scales large language models by replacing dense feed-forward networks with sparsely activated expert networks [22]. Each token is routed to only a small subset of experts, increasing model capacity while keeping per-token computation nearly constant. As modern MoE models grow to tens or hundreds of experts [3, 6, 12, 15, 18, 20, 21], distributed inference commonly relies on expert parallelism (EP), where experts are partitioned across GPUs. Under EP, tokens are routed to experts placed on different GPUs, turning each MoE layer into a distributed computation with cross-GPU token movement. This distributed execution makes dispatch and combine a central cost of MoE inference. In each MoE layer, dispatch sends token activations to the GPUs that host their selected experts; after expert computation, combine returns and reduces the expert outputs into final token representations. Prior work reports that dispatch and combine can consume up to 40% of MoE layer latency [28]. This pressure is amplified by hardware scaling: as shown in Figure 1, BF16 tensor throughput has grown faster than NVLink bandwidth across recent GPU generations, and this widening computeinterconnect gap makes communication an increasingly important component of MoE layer latency. To reduce this exposed communication cost, softwarelevel communication-computation overlap has become critical for efficient MoE inference. Based on their awareness of the dynamic routing results, existing systems fall into three categories. Figure 2 compares the execution timelines of these approaches. (i) No explicit SM partitioning Triton-Distributed (TD) [32] runs communication and computation serially across all SMs, relying on global preemption for overlap at operator boundaries. All SMs execute dispatch, GEMM, and

CCS Concepts: • Software and its engineering → Software performance. Keywords: MoE, Expert Parallelism, communication-computation overlap, GPU SM scheduling

∗

1

Introduction

Both authors contributed equally to this work.

BF16 TFLOPS

2000 1000 0

312

300

A100

NVLink (GB/s)

2.3× 2.0× 2250 900

3.2× 1.5× 989 450

H100

B200

NVLink (GB/s, unidir.)

BF16 TFLOPS

Ziyu Huang et al.

1000 750 500 250 0

Routing result

Static

SM1 SM2 SM3 SM4

(a) No overlap D

G0

G1

C

Triton-dist Time

Comm SM1

Dispatch (Comm)

Static

D

SM2

saved G0

SM4 Comp

GEMM (Comp)

iter-wise dynamic

G1

Comm SM1

D

SM2

C saved

SM3

G0

G1

D0

D1

SM3

G0-0

G1-0

S

C0 G0-1

C1

combine in sequence, yielding minimal communicationcomputation overlap. (ii) Static SM partition DeepEP [30] and ParallelKittens (PK) [26] fix the number of communication SMs at compile time, applying the same partition to all layers. Dedicating separate SM groups to communication and computation enables overlap, but data dependencies in the communicationcomputation pipeline (up gemm cannot begin until dispatch delivers tokens; combine must wait for down gemm to finish) create pipeline bubbles that leave SMs idle. (iii) Coarse-grained dynamic Comet [28] selects the communication/computation SM ratio from a pre-profiled kernel library based on the input sequence length, switching once per iteration. This adapts to workload intensity across iterations, but the partition still does not respond to per-layer routing variation. However, none of these approaches adapts the SM partition to the actual routing result of each layer, leaving substantial GPU resource waste along both spatial and temporal dimensions. The waste manifests in two dimensions. Spatially, the optimal communication/computation SM ratio depends on each layer’s routing result, which determines the communication and computation volumes on each GPU. When the SM partition does not match these volumes, one group of SMs finishes early and idles while the other remains the bottleneck, wasting either communication bandwidth or computation throughput. Since routing skew varies across inputs, layers, and GPUs, a fixed or per-iteration partition inevitably mismatches some layers. Temporally, even with a well-chosen SM split, the pipeline bubbles shown in Figure 2(b) persist: communication SMs remain idle during the middle computation phase, forming a large mid bubble that further degrades SM utilization. Building on these observations, we propose Weave, a runtime SM scheduling system for MoE communicationcomputation overlap (Figure 2(d)). Weave exploits the fact that each layer’s communication and computation volumes become known after routing, and schedules SMs per layer and per GPU rather than relying on an offline fixed policy.

GEMM(UP&GATE) (Computation)

saved

G1-1

SM4 Comp

Dispatch (Communication)

(c) coarsedgrained dynamic Comet

Time

Comm SM1

layer-wise dynamic SM2

DeepEP

Time

SM4 Comp

Figure 1. GPU compute vs. NVLink bandwidth scaling across A100/H100/B200. Left blue axis: BF16 dense TFLOPS. Right red axis: NVLink unidirectional GB/s.

(b) static SM partition

C

SM3

Time

Stealing GEMM work

GEMM(DOWN) (Computation)

Combine (Communication)

(d) fine-grained dynamic Ours (Weave) Dependency

Figure 2. Four MoE execution patterns. (a) No overlap (TD): all SMs execute dispatch, GEMM, and combine serially. (b) Static SM partition (DeepEP): communication and computation SMs are fixed at compile time, leaving idle bubbles. (c) Coarse-grained dynamic (Comet): SM count adjusts per iteration. (d) Fine-grained dynamic (Weave): SM partition adjusts per layer and per GPU. It combines a spatial scheduler that partitions SMs into communication workers and computation workers with a temporal scheduler that coordinates the two worker groups to minimize SM idleness, unified by a lightweight cost model running inside a persistent megakernel. We implement Weave on PK [26] and evaluate it on 4×H100 SXM GPUs across six mainstream MoE architectures. Compared with five state-of-the-art baselines [26, 28, 30–32], Weave achieves the lowest per-layer MoE latency in all configurations, with a 2.89× geometricmean MoE-layer speedup and a 1.33× geometric-mean end-to-end speedup. This paper makes the following contributions.

• We identify routing-determined SM scheduling as a key inefficiency in distributed MoE inference, showing that the communication/computation SM split must be selected per layer and per GPU after routing completes. • We design Weave’s runtime spatial scheduler, which uses routing-derived workload volumes and calibrated hardware throughput curves to choose the communication SM count for each layer and GPU. • We design Weave’s temporal scheduler, which uses chunk pipelining and bubble stealing to reduce SM idleness caused by MoE data dependencies within each layer. • We implement Weave based on Parallel Kittens [26] and show that it consistently improves MoE-layer and end-to-end inference latency over state-of-theart baselines. Our implementation will be open-sourced upon acceptance. 2

Weave

2

Background and Motivation

SMs between the two tasks; for example, DeepEP [30] dedicates 20 of the 132 SMs on an H100 to communication and uses the remaining 112 for computation. However, even SM-level partitioning is a non-trivial design problem. We conduct an experiment that evaluates different compute-communication SM partitions across six MoE models (DeepSeek-V3 (DSv3), Phi-3.5-MoE, Qwen3-30B, Qwen3.5-35B, DeepSeek-V2-Lite (DSv2-Lite), DeepSeek-V2 (DSv2); seq = 8192 prefill, 4×H100 SXM, EP = 4, BF16). As shown in Figure 4(c), on the one hand, the optimal partition point that minimizes overall execution time varies across models, indicating that the SM partition must be adapted per layer. On the other hand, the trend itself also differs across models, suggesting that the strategy should adapt to different models, configurations, and other factors.

This section first introduces the basics of Mixture-ofExperts (MoE) and the corresponding computation/communication overlap techniques. Then we analyze the limitations of existing communication-related optimizations to derive our insights. 2.1

Mixture-of-Experts and Expert Parallelism

MoE replaces the monolithic FFN layer with tens to hundreds of smaller expert sub-networks, each activated only for a subset of tokens selected by a learned gating mechanism, thereby scaling model parameters while keeping pertoken computation nearly constant. The presence of multiple discrete experts gives rise to a new parallelism paradigm: expert parallelism (EP), under which experts are partitioned across GPUs. Hosting experts across multiple GPUs introduces cross-GPU communication overhead. As illustrated in Figure 3, each MoE layer under EP consists of five stages: (1) dispatch, which sends each token to the GPU hosting its target experts; (2) GEMM0, the upprojection and gated FFN; (3) activation, typically a SiLU followed by element-wise multiplication; (4) GEMM1, the down-projection FFN; and (5) combine, which aggregates the per-expert results via weighted summation and returns them to the originating GPU. Among these, dispatch and combine are cross-GPU communication operations, while GEMM0, activation, and GEMM1 are local expert computation. Inputs

to Expert 0

Token A t0 t1

Dispatch

t0'

GATE Expert 1 GEMM 0 UP

to Expert 1

t2

t1' t2'

GATE

DOWN

SILU_MUL

GEMM 1

MUL

DOWN

from Expert 0

Outputs

t0'

from Expert 1 t1' t2'

t0

2.3.1

t1

Dispatch

Token B t3

t3' t4'

SiLU

to Expert 3 t5'

GEMM 0 UP

SILU_MUL

GEMM 1

MUL

DOWN

GATE

SiLU SILU_MUL

GEMM 1

from Expert 3

MUL

DOWN

t5'

Expert 1 GEMM 0 UP GATE

from Expert 2 t3' t4'

As discussed in §1, existing approaches fall into three suboptimal modes: static fixed, coarse-grained dynamic, and no overlap. We now quantify the performance gap on DSv2 (4×H100 SXM) with CodeSearchNet dataset. As the hatched bars of Figure 5 show, within a single MoE layer the per-GPU optimal 𝑐 ∗ ranges from 12 to 36. Neither DeepEP’s fixed 𝑐=20 nor Comet’s per-iteration 𝑐=18 matches any GPU’s optimum. Since all GPUs synchronize after each MoE layer, the layer latency is determined by the slowest GPU. The hollow markers report the resulting per-GPU latency: on the bottleneck GPU (GPU 3), DeepEP is 1.33× slower than the per-GPU optimal, Comet 1.18×, and TD 1.44×. Determining the optimal SM partition at runtime is therefore a key challenge.

Outputs t3 t4 t5

SiLU

token routing (from GPU0)

token routing (from GPU1)

Figure 3. MoE forward pass under expert parallelism (2 GPUs, 4 experts, top-𝑘 = 3 routing). Dispatch sends each token to the GPU hosting its expert; the expert sub-network sequentially executes gemm0 (UP+GATE), silu_mul, and gemm1 (DOWN); combine performs weighted summation and returns the results. 2.2

Challenge A: Runtime Partition Decision.

t2

Combine

to Expert 2

Inputs

t5

GEMM 1

MUL

Limitations of SOTA Approaches

Existing distributed MoE systems either fix the number of communication SMs at compile time [26, 30], select from a pre-profiled configuration per iteration [28], or avoid explicit SM partitioning altogether and run communication and computation serially [32]. However, MoE workloads are inherently dynamic: the routing result of each layer changes, causing the communication and computation volumes to shift, and consequently the optimal communication SM count varies from layer to layer. Resolving this mismatch poses two key challenges: how to determine the optimal SM partition at runtime, and, once the partition is made, how to coordinate the two groups of SMs (comm workers and comp workers) to maximize overlap while minimizing idle time.

Expert 2

GPU1

t4

SILU_MUL SiLU

Combine

Expert 0 GEMM 0 UP

GPU0

2.3

2.3.2 Challenge B: Dynamic Coordination of Two Worker Groups. Even after performing the SM partition described in Challenge A, complex dependencies between communication

Compute-Communication Overlap

To hide dispatch and combine latency, modern systems overlap communication with computation by partitioning GPU 3

450 GB/s peak

28 27 26 23

25

27

Dispatch SMs DSv3

(b) GEMM throughput

29 28 27 26

Phi-3.5-MoE

23

24

26

25

22 21 20 21 22 22

27

GEMM SMs

Qwen3-30B

(c) Dispatch + GEMM

23

989 TFLOPS peak

210

Measured time (ms)

(a) Dispatch bandwidth

Throughput per rank (TFLOPS)

Bandwidth per rank (GB/s)

Ziyu Huang et al.

Qwen3.5-35B

DSv2

23

24

25

Dispatch SMs

26

DSv2-Lite

Figure 4. Impact of SM partitioning on compute-communication overlap (4×H100 SXM, 6 MoE models, seq = 8192 prefill, EP = 4, BF16). (a) Dispatch bandwidth vs. dispatch SM count 𝑐 : near-linear then saturates. (b) GEMM throughput vs. compute SM count 𝑁 −𝑐 : sub-linear under power and L2 cache constraints. (c) Dispatch-gemm total time with tile-level overlap as a function of 𝑐 ; markers indicate each model’s optimal split.

60

40 30

40

20

20

10 0

SM Active (%)

DSv2-Lite, seq=2048, 4×H100 SXM, EP=4

Latency (ms)

Comm. SM count (c)

Latency (ms, markers) Optimal Comet DeepEP TD Comm. SM count (c, bars) Comet (c = 18) DeepEP (c = 20) Optimal c *

GPU 0

GPU 1

GPU 2

GPU 3

100

Weave (Ours)

75

PK TD Comet DeepEP Weave

50 25 0

0

20

40

Comm/Comp Overlap (%)

0

Figure 6. Performance of MoE overlap systems along two utilization dimensions: comm-comp overlap (x-axis) and SM Active (y-axis). Each marker is one system, measured on DSv2-Lite (seqlen = 2048, 4×H100 SXM, EP = 4, BF16).

Figure 5. SM partition mismatch under existing policies (DSv2, CodeSearchNet dataset, Layer 1, 4×H100 SXM). Left y-axis (hatched bars): per-GPU optimal comm SM count 𝑐 ∗ vs. Comet (𝑐=18, per-iteration) and DeepEP (𝑐=20, static); the optimal 𝑐 ∗ ranges from 12 to 36 across GPUs. Right yaxis (hollow markers): per-GPU latency under each policy plus TD (serial); layer latency is limited by the slowest GPU.

do not fill. As shown in Figure 6, no existing system simultaneously achieves high SM Active and high comm-comp overlap; only Weave maintains high levels on both dimensions (91% SM Active, 47% overlap).

and computation tasks make it difficult to coordinate the two worker groups so that they run in parallel (high comm/comp overlaping) without waiting for each other (high SM Active rate). The MoE forward pass has inherent data dependencies: up gemm cannot begin until dispatch delivers the required tokens, and combine cannot start until down gemm finishes. Under SM partitioning, these dependencies cause communication workers to idle during the middle computation phase, forming pipeline bubbles that existing systems [26, 28, 30]

3

Design

In this section, we present the core design of Weave, a runtime SM scheduling system for efficient MoE execution under expert parallelism. Existing systems cannot perceive per-layer routing results and adapt the SM partition accordingly, which leads to missed optimization opportunities in two dimensions: spatially, a mismatched SM split wastes either communication bandwidth or compute throughput; temporally, the complex dependencies introduced 4

Weave

Software: routing

Hardware: throughput

layer 7

BW

layer 1

GPU0

...

TFLOPS

GPU1 local

the number of tokens to send and receive, the token distribution across experts, and the resulting computation volume on each GPU. Weave therefore invokes its scheduler immediately after routing. The scheduler combines these runtime workload descriptors with hardware throughput profiles and uses a lightweight cost model inside the megakernel prologue to generate a layer-specific scheduling plan within microseconds. The megakernel then transitions to its main execution phase. Weave jointly optimizes scheduling along two dimensions. The spatial scheduler (§3.1) selects the number of SMs assigned to communication so that communication bandwidth and computation throughput are balanced for the current layer. The temporal scheduler (§3.2) chunks routed tokens into pipelined microbatches and allows idle communication SMs to steal GEMM tiles when communication work is temporarily unavailable. Since the best temporal schedule depends on the throughput induced by the spatial SM split, Weave solves the two dimensions jointly using a unified cost model (§3.3). Concretely, the scheduling decision is governed by two variables: the number of communication SMs 𝑐 (how to partition SMs between the two worker groups) and the pipeline chunk count 𝐾 (how to split tokens into pipelined microbatches for temporal coordination). The cost model (§3.3) jointly searches over the (𝑐, 𝐾 ) space and, for each candidate pair, derives the number of GEMM tiles 𝑛steal that communication workers can steal during their idle periods. The optimal (𝑐 ∗ , 𝐾 ∗ ) and the corresponding steal count fully determine the execution plan for each layer and GPU.

remote

SMs GPU 1 GPU 0

§3.1 Spatial Scheduler

Fused MegaKernel Cost Model SMs comm SMs comp SMs

S combine0

combine1

gemm00 act0 gemm10 gemm01 act1

gemm11

dispatch

time

§3.1 Temporal Scheduler Optimized Execution Plan

Figure 7. System overview of Weave. The system takes routing results (software input) and communication bandwidth / compute throughput curves (hardware input), and makes runtime decisions inside the megakernel along two dimensions: spatial scheduler (how to split SMs between communication and computation, §3.1), and temporal scheduler (how the two types of workers coordinate their execution, §3.2). The output is an optimized execution plan applied to the current layer.

3.1

by communication-computation overlap leave SMs idle. Weave addresses both dimensions with runtime scheduling. Figure 7 gives an overview of Weave’s architecture. Weave takes two categories of input: software input, consisting of per-layer routing results that determine the communication volume and computation workload on each GPU; and hardware input, consisting of calibrated communication bandwidth and compute throughput curves as functions of SM count. These inputs are fed into a persistent megakernel that fuses all five forward operations of an MoE layer (dispatch, gemm0, silu_mul, gemm1, combine). Within this megakernel, each SM is assigned either a comm role (Dispatch and Combine) or a comp role (GEMM0, SiLU-Mul, and GEMM1). A lightweight cost model in the megakernel prologue jointly determines the spatial SM partition and the temporal execution order, producing a per-layer, per-GPU scheduling plan that coordinates the two types of workers throughout the layer. The key observation behind Weave’s runtime scheduling is that the workload of an MoE layer becomes fully determined once routing completes. At that point, Weave knows

Spatial Scheduler

The spatial scheduler determines how the available SMs are divided between communication and computation. As shown in Figure 4(c), the overall latency as a function of the SM split has a clear optimum that depends on the routed workload of the current layer. As shown in Figure 4(a)(b), both communication bandwidth and GEMM throughput saturate well before all 𝑁 SMs are assigned; a small number of communication SMs is sufficient to approach peak bandwidth, leaving the majority for computation. We now formalize the optimization. Let 𝑁 denote the number of SMs available to the megakernel, and let 𝑐 denote the number of SMs assigned to communication; the remaining 𝑁 −𝑐 SMs are assigned to computation. Let 𝐻 denote the hidden dimension, 𝐼 the per-expert FFN intermediate dimension, and 𝐵 = 2𝐻 bytes the per-token communication size (each token is an 𝐻 -element BF16 vector, 2 bytes per element). After routing, let 𝑋local denote the number of local tokens routed to local experts, 𝑋in the number of remote uniq tokens routed to local experts, and 𝑋in the deduplicated 5

Ziyu Huang et al.

(a) Baseline (c comm SMs) comm SMs

D

C

mid bubble

comp SMs

G0

𝑊comp counts the GEMM work executed jointly on 𝑋local + 𝑋in token-expert pairs (both local and remote). For commuuniq nication, the dispatch volume is 𝑋in ⋅ 𝐵 (each unique remote token is pulled once and reused) and the combine volume is 𝑋in ⋅𝐵 (one per-expert output returned for each tokenexpert pair). We let 𝛼 ∈ [0, 1) denote the tile-level overlap ratio between communication and computation, calibrated empirically. The remaining fraction (1 − 𝛼) of communication is exposed as a tail outside the overlap window. The three time components are

c

G1 Ttail time

Ta

(b) + Spatial: reduce c to c * comm SMs

D

comp SMs

G0

C

bubble

c*

Tb

G1 time

𝑇comp =

saved

(c) + Temporal: chunk pipeline (K=2) D

comm SMs comp SMs

G00

G10

C0

C1

G01

G11

𝑇comm =

Tc

D

S C0

saved

gemm

comp SMs

𝑐 ∗ = arg min 𝑇total .

Td time

dispatch (D) gemm0 (G0)

act (silu_mul) gemm1 (G1)

saved

combine (C) steal (S)

(5) (6)

Figure 8. Operational details of the spatial and temporal schedulers. (a) Baseline with communication-computation overlap enabled; a large mid bubble persists. (b) The spatial scheduler applies a near-optimal communication/computation SM ratio, reducing the mid bubble. (c) The temporal scheduler partitions tokens into microbatches, exposing additional overlap opportunities and further reducing the mid bubble. (d) The temporal scheduler fills the residual mid bubble by assigning idle communication SMs to execute GEMM tiles.

3.2

(2) (3)

Temporal Scheduler

3.2.1 Chunk Pipeline. The mid bubble in Figure 8(a) comes from the DAG dependency: a Combine operation cannot start until the corresponding GEMM1 tile has been produced. To relax this constraint, we split the tokens into smaller chunks; chunks are independent of each other, which opens up more overlap

(1)

𝑊combine = 𝑋in ⋅ 𝐵 bytes.

(8)

Given the SM partition 𝑐 from the spatial scheduler, the temporal scheduler addresses how the two groups of SMs coordinate over time. As Figure 8(a) shows, under a fixed SM partition with no temporal scheduling, Dispatch and Combine cluster at the two ends of the timeline, leaving a large stretch of idle comm SMs in the middle of the MoE layer and forming a large mid bubble, i.e., an idle interval during the GEMM phase. To eliminate this bubble, Weave uses two mechanisms: the chunk pipeline (§3.2.1) spreads Combine into the middle of the GEMM phase for fine-grained overlap; bubble stealing (§3.2.2) lets comm SMs steal GEMM tiles in the residual mid bubble that the chunk pipeline still cannot cover.

count of 𝑋in (under top-𝑘 routing with 𝑘>1, the same remote token may be selected by multiple local experts, so uniq 𝑋in ≤ 𝑋in ; each unique remote token is transferred across the network only once and then reused for all selecting experts [30, 33]). The per-layer workloads are: uniq 𝑊dispatch = 𝑋in ⋅ 𝐵 bytes,

(7)

If 𝑇comp + 𝑇tail > 𝑇comm , the layer time is determined by the GEMM together with the unhidden combine tail; otherwise the GEMM fits entirely within the communication window and the layer time is determined by serialized dispatch and combine.

spatial scheduler temporal scheduler

𝑊comp = (𝑋local + 𝑋in ) ⋅ 6𝐻 𝐼 FLOPs,

,

𝑇total = max(𝑇comp + 𝑇tail , 𝑇comm ),

c*

C1

(4)

The layer is bounded by whichever of the two paths is slower, plus the unhidden tail. The spatial scheduler therefore selects

(d) + Temporal: bubble stealing (Weave) comm SMs

, TFLOPS(𝑁 −𝑐) 𝑊dispatch + 𝑊combine

BW(𝑐) 𝑊 𝑇tail = (1 − 𝛼) ⋅ combine . BW(𝑐)

c*

time

𝑊comp

6

Weave

Below we derive the number of tiles to steal. Given 𝑐 , communication takes 𝑊 𝑡comm (𝑐) = comm . (11) BW(𝑐) During this interval, the 𝑁 −𝑐 comp SMs complete

opportunities. As Figure 8(c) shows, we split GEMM0, SiLUMul, GEMM1, and Combine into 𝐾 chunks, while Dispatch is not chunked, so that communication can overlap with computation as much as possible. Unlike a naive approach where comm SMs switch to comp after dispatch finishes, the chunk pipeline keeps comm SMs continuously doing communication throughout the middle phase. Figure 8(c) illustrates 𝐾 =2: chunk 0’s Combine (𝐶0 ) runs during chunk 1’s GEMM, so communication and computation execute concurrently in time. The choice of 𝐾 involves a trade-off: a larger 𝐾 gives finer chunks and a smaller mid bubble, but each chunk holds fewer tokens, degrading per-chunk GEMM throughput. As Figure 9a shows, we benchmark the group GEMM performance under the chunk pipeline on H100 using DSv3 with input seq = 8k. As 𝐾 increases, the performance degradation accelerates; at 𝐾 =8, throughput drops by ∼27% compared to 𝐾 =1. We capture this degradation with a correction factor effchunk (𝐾 ) ∈ (0, 1], defined as the ratio of chunked GEMM throughput to unchunked throughput (𝐾 =1). The computation time under chunking becomes:

𝑇comp (𝑐, 𝐾 ) =

𝑊comp TFLOPS(𝑁 −𝑐) ⋅ effchunk (𝐾 )

,

𝑊partial (𝑐) = 𝑡comm (𝑐) ⋅ TFLOPS(𝑁 −𝑐) of GEMM work, leaving

𝑊steal (𝑐) = max(0, 𝑊comp − 𝑊partial (𝑐))

𝑊combine . BW(𝑐) ⋅ 𝐾

(13)

GEMM FLOPs to be stolen. Once Dispatch finishes, all 𝑁 SMs are available for GEMM, so 𝑊steal (𝑐) is shared equally across 𝑁 SMs. Letting 𝑊tile denote the FLOPs of a single GEMM tile, each SM steals 𝑊 (𝑐) 𝑛steal (𝑐) = steal (14) 𝑁 ⋅ 𝑊tile GEMM tiles. 3.3 Joint Search for (𝑐 ∗ , 𝐾 ∗ ) Substituting the 𝐾 -aware definitions of 𝑇comp and 𝑇tail from §3.2 into the total layer time expression of §3.1 yields

𝑇total (𝑐, 𝐾 ) = max(𝑇comp (𝑐, 𝐾 ) + 𝑇tail (𝑐, 𝐾 ), 𝑇comm (𝑐)). (15) The scheduler then jointly searches for

(9)

where effchunk (𝐾 ) ∈ (0, 1] is the ratio of chunked GEMM throughput to the unchunked case. Chunking also modifies the tail term: the final chunk’s combine, the only segment not absorbed by a subsequent GEMM, is now of size 𝑊combine /𝐾 , so the tail is reduced by a factor of 𝐾 relative to the unchunked case:

𝑇tail = (1 − 𝛼) ⋅

(12)

(𝑐 ∗ , 𝐾 ∗ ) = arg min 𝑇total (𝑐, 𝐾 ). 𝑐,𝐾

(16)

In the comp-dominated regime(if 𝑇comp + 𝑇tail > 𝑇comm ), increasing 𝐾 raises 𝑇comp through the degradation of effchunk (𝐾 ) but lowers 𝑇tail via the 1/𝐾 scaling. The two opposing effects produce the real measured (𝑐, 𝐾 ) grid result(Fig. 9b). The number of GEMM tiles to steal is then computed from (𝑐 ∗ , 𝐾 ∗ ). Figure 9b shows the measured latency under different (𝑐, 𝐾 ) configurations. We use DSv2-Lite with input sequence length 8192 for this experiment. The grid exhibits a single clear optimum at 𝑐 ∗ =40, 𝐾 ∗ =4 (1.365 ms), confirming that 𝑐 and 𝐾 must be searched jointly. Once (𝑐 ∗ , 𝐾 ∗ ) is determined, the scheduler further computes the per-SM steal-tile count 𝑛steal (𝑐 ∗ ) via equations (13) and (14) from §3.2.2. At this point the megakernel’s complete execution plan (SM role assignment, chunk count, and the per-SM steal-tile count) is fully determined, and the kernel enters its main execution phase.

(10)

A larger 𝐾 reduces the unhidden tail but increases the GEMM time through the degradation of effchunk (𝐾 ); a smaller 𝐾 does the opposite. The optimum balances these two effects. 3.2.2 Bubble Stealing. After the joint search of §3.3 chooses (𝑐 ∗ , 𝐾 ∗ ), communication SMs may still have idle capacity between the end of dispatch and the start of the chunked combine. We fill this gap by letting them steal GEMM tiles, shown as the S block in Figure 8(d). Once Dispatch finishes, comm SMs join the 𝑁 −𝑐 comp SMs on GEMM work, converting idle capacity into effective GEMM throughput. We place the steal window right after Dispatch so that stolen GEMM tiles execute consecutively and subsequent Combine operations also run without interruption. If stealing were instead scattered into the idle gaps between each chunk’s Combine, both GEMM and Combine would be repeatedly interrupted by role switches, degrading performance. This consolidated placement yields the pattern shown in Figure 8(d).

4

Implementation

Weave is built on PK [26] and implemented as a single persistent megakernel, launched with a grid size equal to the SM count (𝑁 =132 on H100). After each layer’s routing completes, the routing result is passed as a kernel parameter; the megakernel then runs the cost model of §3.3 on-GPU to obtain the current layer’s (𝑐 ∗ , 𝐾 ∗ ). Both 𝑐 ∗ and 𝐾 ∗ can therefore switch dynamically per layer. 7

4.86

5.57 100

2 0

1

2

4

chunk count K

8

95 90 85 80 75

(a) Impact of chunk splitting on GEMM (DSv3, seq=8192, H100).

8 2.267 1.535 1.544 4 1.810 1.365 1.519 2 2.205 1.701 1.651 1 2.539 1.724 1.880 24 40 56

Table 1. MoE model configurations used in evaluation.

2.50 2.25 2.00 1.75 1.50

latency (ms)

4

4.39 4.46

K (num chunks)

6

effchunk (%)

Layer latency (ms)

Ziyu Huang et al.

c (comm SMs)

(b) (𝑐, 𝐾 ) latency grid for DSv2Lite (seq=8192). Star: optimum 𝑐 ∗ = 40, 𝐾 ∗ = 4.

Each block derives its role from 𝑐 ∗ and blockIdx.x: blocks with index < 𝑐 ∗ are comm SMs (executing dispatch and combine), while the remaining blocks are omp SMs (executing gemm0, activation, and gemm1). All tile assignments use dynamic scheduling: each SM claims its next tile via atomicAdd on a global counter following the chunk pipeline DAG order, naturally tolerating variation in SM execution speeds [27]. After dispatch finishes, comm SMs steal 𝑛steal GEMM tiles before transitioning to combine. For the final combine, all 𝑁 SMs participate regardless of role, eliminating tail idleness. Weave also fuses dispatch-side token deduplication [30] into the megakernel. DeepEP [30] and TD [32] perform deduplication as a separate operation before GEMM0, whereas Weave overlaps dispatch with GEMM0 and achieves deduplication naturally during dispatch. After routing completes, duplicate tokens are identified and each comm tile knows in advance whether its token is a duplicate. If so, the tile reads directly from a local HBM buffer populated by the first transfer, avoiding both redundant cross-GPU transfers and HBM polling.

Evaluation

5.1

Evaluation Setup

𝐸

top-𝑘

𝐻

𝐼

DSv3 Phi-3.5-MoE Qwen3-30B Qwen3.5-35B DSv2-Lite DSv2

256 16 128 256 64 160

8 2 8 8 6 6

7168 4096 2048 2048 2048 5120

2048 6400 768 512 1408 1536

and PyTorch 2.6.0+cu124. The evaluated baseline implementations are SGLang v0.5.9, Triton-Distributed(TD) v3.4.0, DeepEP v1.2.1, Comet with Flux v1.1.2, and ParallelKittens at commit a8f63a9. Models. We evaluate six mainstream BF16 MoE architectures under expert parallelism with EP=4, as summarized in Table 1. Here, 𝐸 denotes the total number of routed experts, top-𝑘 denotes the number of experts selected per token, 𝐻 denotes the hidden dimension, and 𝐼 denotes the per-expert intermediate dimension. Workloads. We sample routing inputs from the ShareGPT dataset and control the input sequence length in {2048, 4096, 8192} with batch size 1. Baselines. We compare Weave with five representative MoE serving and communication-computation overlap systems. SGLang [31] is a production-grade serving system that uses NCCL All-to-All and cuBLAS grouped GEMM. TD [32] fuses dispatch with GEMM0 and GEMM1 with combine, and uses global task preemption within fused operators to reduce inter-operator bubbles. DeepEP [30] is an expert-parallel communication library that chunks tokens into pipeline stages to overlap dispatch and combine with expert GEMM; we use DeepGEMM [29] as its GEMM backend and keep its default communication partition of 20 SMs. Comet [28] fuses dispatch with GEMM0 and GEMM1 with combine, enabling tile-level communication-computation overlap. ParallelKittens (PK) [26] fuses dispatch with expert up-projection under a compile-time fixed SM partition.

Figure 9. (a) Chunk splitting degrades GEMM throughput; at 𝐾 =8, performance drops by ∼27%. (b) Measured latency over the (𝑐, 𝐾 ) grid shows a clear unimodal optimum, confirming joint search is necessary.

5

Model

5.2

Per-Layer MoE Latency

We compare Weave against five baselines (SGLang [31], PK [26], TD [32], DeepEP [30], Comet [28]) across 6 MoE architectures and 3 sequence lengths (2k, 4k, 8k tokens). As shown in Figure 10 (Row 1), Weave achieves the lowest per-layer latency in all 18 configurations. Weave achieves 1.95×–4.76× geometric mean speedup over all baselines: vs. DeepEP 1.95×, vs. Comet 2.01×, vs. TD 2.95×, vs. SGLang 3.63×, vs. PK 4.76×. DeepEP uses DeepGEMM [29] as its GEMM backend and chunks tokens into pipeline stages to overlap dispatch/combine with expert GEMM. Comet performs tile-level overlap within two fused kernels (dispatch+gemm0 and gemm1+combine) and can select the communication SM count per iteration

Hardware. All experiments are conducted on a node with 4× NVIDIA H100 80 GB SXM5 GPUs. Each GPU has 132 SMs, 80 GB HBM3, and a peak BF16 throughput of 989 TFLOPS. The GPUs are fully connected through NVLink and NVSwitch, with 450 GB/s unidirectional bandwidth per GPU. The host has two Intel Xeon Platinum 8468 CPUs, each with 48 physical cores and 96 hardware threads, for a total of 96 cores and 192 hardware threads across 2 NUMA nodes. The system has 2.0 TiB of CPU memory. Software. We use CUDA Toolkit 12.9.86, NVIDIA driver 575.57.08, NCCL 2.21.5, NVSHMEM 3.6.5, Python 3.10.12, 8

Weave

Attention

MoE latency (ms)

Phi-3.5-MoE

DSv3

SGLang

PK

Qwen3-30B

TD

DeepEP

Qwen3.5-35B

Weave

DSv2-Lite

DSv2

10.0

1.00 1.00

4k

8k

12.5

2k

4k

8k

20

7.5 5.0

10

2.5 2k

4k

8k

0

2k

4k

8k

2k

4k

8k

2k

4k

8k

2k

4k

8k

2k

4k

8k

12.5

30

10.0

0.0

1.00

1.00

1.00 2k

E2E latency (ms)

Comet

2k

4k

8k

10 8 6 4 2 0

8

20

7.5

6

15

5.0

4

10

2.5

2

5

10.0

2k

4k

0.0

8k

Sequence length

2k

4k

8k

0

2k

4k

8k

0

Figure 10. Per-layer MoE latency and end-to-end latency on 4×H100 SXM, EP=4, BF16. Row 1: per-layer MoE latency (logscale y-axis) across six models and three sequence lengths (2k/4k/8k). Row 2: end-to-end latency breakdown for the same six models at the same three sequence lengths; the gray band is attention layer time and the colored bars on top are MoE layer time for each system. based on sequence length. TD fuses the same two operator pairs into two kernels, but within each fused kernel communication and computation run serially across all SMs, with only marginal overlap at operator boundaries via global preemption. PK fuses dispatch with the expert up-projection under a compile-time fixed SM partition. SGLang issues separate NCCL and cuBLAS calls without kernel-level overlap. None of these baselines adapts the SM partition to per-layer routing results. 5.3

5.4

Hardware Utilization Profiling

We profile Weave and other baselines (TD, Comet, DeepEP, PK) with Nsight Systems on DSv2-Lite (seq = 2048, 4×H100 SXM, EP=4). The profiling window covers the complete MoE forward pass from the end of routing to the end of combine. We report three GPU metrics (SM Active, NVLink Utilization, Tensor Core HMMA) and a derived overlap ratio, defined as the fraction of the profiling window during which both NVLink Utilization and Tensor Core HMMA are simultaneously greater than zero. Figure 11 summarizes the mean (bar) and peak (hollow circle atop dashed line) of the three metrics over each system’s MoE forward window, together with the overlap ratio. Weave achieves the highest overlap ratio (47.1%). The remaining baselines fall well below: Comet (14.3%), DeepEP (9.2%), PK (5.8%), and TD (5.1%). Comet has the second-highest SM Active mean (47.4%) because its two-segment fusion keeps SMs busy, but its overlap is moderate. DeepEP’s low overlap stems from a mismatched communication SM count: NVLink activity appears as brief bursts that quickly saturate and stop, leaving the vast majority of the timeline with only Tensor Core active. Weave achieves leading performance across all four metrics. SM Active: fusing all five MoE operations into a single persistent megakernel eliminates inter-kernel launch gaps, and bubble stealing ensures communication SMs have GEMM tiles to execute after their primary work finishes,

End-to-End Performance

We evaluate end-to-end (E2E) latency by combining the preMoE/attention block time measured via SGLang with the MoE kernel time from Weave and each baseline. We test the same 6 MoE architectures at sequence lengths 2048, 4096, and 8192 (batch size 1, 4×H100 EP=4), yielding 18 configurations per baseline. Figure 10 (Row 2) shows the E2E latency breakdown for all 18 configurations. The gray band represents the attention time and the colored bars show each system’s MoE time. Weave achieves the lowest E2E latency in all 18 configurations, with 1.12×–1.70× geometric mean E2E speedup across baselines: vs. DeepEP 1.12×, vs. Comet 1.13×, vs. TD 1.28×, vs. SGLang 1.50×, vs. PK 1.70×. The E2E advantage is more pronounced at shorter sequence lengths: as sequence length grows, attention time scales quadratically while MoE time scales linearly, so the attention portion increasingly dominates E2E latency and dilutes Weave’s MoE-layer speedup. 9

Ziyu Huang et al.

(a) SM Active (%)

(b) NVL Utilization (%)

100

91.3 100

80

80

60 40 20

25.3

13.6

40 16.4

0 EP TDComet PKWeave Deep

20

60 47.1 50 80 40 60 30 40 20 14.3 23.5 20 10 9.2 5.1 11.5 5.8 5.8 3.2 3.3 0 0 EP TDComet PKWeave EP TDComet PKWeave Deep Deep

100

60

47.4

(c) Tensor Core (%) (d) Comm-Comp Overlap (%)

0.3 1.4 0.7 0.7

7.0

0 EP TDComet PKWeave Deep

Figure 11. Aggregate GPU metrics over the MoE forward window (DSv2-Lite, seq = 2048, 4×H100, EP=4). Bars: mean; hollow circles atop dashed lines: peak. (a) SM Active (%). (b) NVLink Utilization (%). (c) Tensor Core HMMA (%). (d) Comm-comp overlap ratio (%).

5.5

6 4 DSv3 Phi-3.5-MoE Qwen3-30B DSv2-Lite

2 2

4

6

8

Measured optimal (ms)

8

MoE layer latency Cost model overhead

1.0 0.8

6

0.6

4

0.4

2

0.2

0

2k

4k

8k 16k

0.0

Sequence length

Figure 12. Cost model evaluation. (a) Configurationselection quality: measured-optimal latency vs. measured latency at the cost-model-chosen (𝑐, 𝐾 ); dashed line is 𝑦=𝑥 . (b) Online overhead on DSv3: MoE layer latency (left blue axis, bars) vs. cost model overhead (right red axis, hollow markers).

Cost Model Accuracy

We evaluate the cost model along two dimensions: selection accuracy and online cost. As shown in Figure 12, we sweep 4 models × 4 sequence lengths on 4×H100 SXM (EP=4), each with 54 (𝑐, 𝐾 ) configurations. (a) Selection accuracy. For each (model, seq_len) group, we compare the measured latency at the cost-model-selected (𝑐, 𝐾 ) against the measured latency at the exhaustive-search optimum. The average accuracy gap is 8.2%. (b) Online cost. Because the cost model is fused inside the megakernel, each block computes independently and only reads a small amount of routing results from global memory. On DSv3, the cost model overhead is a nearly constant 0.54 𝜇 s regardless of sequence length, less than 0.021% of the MoE layer time, indicating that the runtime search adds negligible overhead. 5.6

8

(b) Cost model overhead MoE layer latency (ms)

Model-chosen (ms)

(a) Cost model accuracy

Cost model overhead ( s)

keeping all SMs occupied throughout the layer. NVLink Utilization and Tensor Core HMMA: the spatial scheduler maintains a well-matched communication/computation SM ratio throughout the kernel’s execution, so neither NVLink bandwidth nor compute throughput starves while the other is saturated. Overlap ratio: the chunk pipeline extends the communication window from the head and tail of the layer to cover the entire computation phase by interleaving combine with gemm over time, maximizing the duration where NVLink and Tensor Core are simultaneously active.

• Weave w/o T: spatial scheduler (S) enabled, temporal scheduler (T) disabled (no chunk pipeline or bubble stealing). Isolates the spatial scheduler’s contribution. • Weave: full design with both spatial and temporal schedulers. As shown in Figure 13, going from Weave w/o S+T to Weave w/o T quantifies the gain of explicit SM partitioning; going from Weave w/o T to Weave further quantifies the additional gain of the temporal scheduler. Both design components yield significant positive gains at every (model, seqlen) point. On Qwen3-30B at seq=4096, enabling S reduces latency by 14.4%, and further enabling T reduces it by 15.7% on top of that; the two stages compose into a 27.8% latency reduction over the no-scheduler configuration. Averaged over the eight (model, seqlen) points, S contributes a 10.1% reduction and T an additional 14.1%, for a total 22.8%

Design Ablation

We isolate the contributions of the spatial and the temporal scheduler by comparing three configurations on Qwen330B and Qwen3.5-MoE at sequence lengths 2k, 4k, 8k, and 16k (4×H100, EP=4).

• Weave w/o S+T: both schedulers disabled; dispatch / gemm / combine run sequentially on all SMs (no compute-communication overlap). 10

Weave

Qwen3-30B

Qwen3.5-MoE

reduction. This confirms that both schedulers are effective, and their gains stack.

such as MoE, effectively using this domain requires runtime control over task ordering and resource allocation. MPK [2] introduces scheduler SMs for device-side task scheduling and also analyzes MoE workloads, but its MoE support is limited to single-GPU execution without cross-GPU dispatch and combine communication. When MPK does consider multi-GPU communication, it targets only static tensor parallelism (TP), thereby missing the compute-communication overlap opportunities unique to expert parallelism where routing-dependent workloads vary per layer. TD’s piecewise megakernel enables limited scheduling within fused regions, but still only partially fuses MoE operators and does not jointly adapt SM partitioning and intra-layer execution order after routing. Weave addresses this gap by using a persistent megakernel as a routing-aware scheduling domain that adapts both decisions at runtime.

6

7

Latency (ms)

Latency (ms)

2 × 100

2 × 100

Weave w/o S+T Weave w/o T Weave Comet

100 6 × 10 1 4 × 10 1

100 6 × 10 1

2K

4K

8K

Sequence length

16K

2K

4K

8K

Sequence length

16K

Figure 13. Design ablation on Qwen3-30B and Qwen3.5MoE (4×H100, EP=4). S = spatial scheduler, T = temporal scheduler. Three configurations: Weave w/o S+T (no overlap), Weave w/o T (spatial only), Weave (full).

Related Work

Compute-communication overlap. Existing MoE inference systems have extensively studied how to overlap inter-GPU communication with expert computation [4, 7, 9, 13, 16, 23, 24, 34]. Existing systems improve overlap at progressively finer granularity. Host- and stream-level approaches pipeline all-to-all communication with expert computation through operator scheduling and pipeline partitioning [9, 11, 23, 24], but the overlap is still bounded by kernel-level stages. Finer-grained approaches such as FlashOverlap [10] and TokenWeave [8] exploit tileor token-level opportunities, yet expert computation often remains encapsulated in opaque library kernels, limiting how communication can be interleaved with computation. Kernel fusion addresses this limitation by moving communication into expert GEMM kernels, allowing dispatch or combine to overlap with GEMM at a finer inkernel granularity rather than only across separate kernel launches. ParallelKittens [26] fuses dispatch with the expert up-projection under a compile-time fixed SM partition, while Comet [28] and TD [32] fuse dispatch/up-projection and down-projection/combine as separate stages. These fused pipelines enable finer-grained overlap by moving communication closer to GEMM execution, but remain limited by fixed SM partitioning, partial fusion, or residual launch boundaries between fused stages. Persistent megakernel. Fusing all operators into a single persistent megakernel has gained increasing attention [1, 2, 5, 14, 17, 19, 25, 33, 35]. MoE kernel fusion reduces launch overhead and exposes opportunities for cross-operator scheduling. Partial-fusion approaches still retain inter-operator boundaries, while whole-layer megakernels such as FlashDMoE [1] integrate dispatch, expert computation, and combine into a single kernel. Whole-layer fusion removes launch boundaries and creates a larger scheduling domain. For highly dynamic workloads

Conclusion

We present Weave, to our knowledge the first MoE overlap system whose SM split is decided per layer and per GPU by routing results at runtime. Existing systems either fix the SM partition at compile time, adjust it only at coarse granularity, or avoid explicit SM partitioning and run communication and computation serially with minimal overlap; Weave instead exploits the fact that each layer’s communication and computation volumes become known after routing, running a lightweight cost model inside the persistent megakernel to jointly select the communication SM count and chunk pipeline configuration online. On 4×H100 SXM GPUs across six mainstream MoE models, Weave achieves a 2.89× geometric-mean MoE-layer speedup and a 1.33× geometric-mean end-to-end speedup over five state-of-the-art baselines. Limitations and future work. Due to hardware constraints, Weave is currently validated on a single 4×H100 NVLink node with EP=4. Extending to larger GPU counts and multi-node configurations is left for future work.

References [1] Osayamen Jonathan Aimuyo, Byungsoo Oh, and Rachee Singh. 2025. FlashMoE: Fast Distributed MoE in a Single Kernel. In Advances in Neural Information Processing Systems (NeurIPS). arXiv:2506.04667. [2] Xinhao Cheng, Zhihao Zhang, Yu Zhou, Jianan Ji, Jinchen Jiang, Zepeng Zhao, Ziruo Xiao, Zihao Ye, Yingyi Huang, Ruihang Lai, Hongyi Jin, Bohan Hou, Mengdi Wu, Yixin Dong, Anthony Yip, Songting Wang, Wenqin Yang, Xupeng Miao, Tianqi Chen, and Zhihao Jia. 2025. Mirage Persistent Kernel: A Compiler and Runtime for MegaKernelizing Tensor Programs. arXiv preprint arXiv:2512.22219 (2025). [3] Damai Dai, Chengqi Deng, Chenggang Zhao, R.X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y.K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. 2024. DeepSeekMoE: Towards Ultimate Expert Specialization in Mixture-of-Experts Language Models. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (ACL). 11

Ziyu Huang et al. [4] DeepSeek-AI. 2024. DeepSeek-V3 Technical Report. arXiv preprint arXiv:2412.19437 (2024). [5] Greg Diamos, Shubho Sengupta, Bryan Catanzaro, Mike Chrzanowski, Adam Coates, Erich Elsen, Jesse Engel, Awni Hannun, and Sanjeev Satheesh. 2016. Persistent RNNs: Stashing Recurrent Weights On-Chip. ICML (2016). [6] William Fedus, Barret Zoph, and Noam Shazeer. 2022. Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity. Journal of Machine Learning Research (JMLR) 23, 120 (2022), 1–39. [7] Yunqi Gao, Bing Hu, Mahdi Boloursaz Mashhadi, A-Long Jin, Yanfeng Zhang, Pei Xiao, Rahim Tafazolli, and Mérouane Debbah. 2025. FlowMoE: A Scalable Pipeline Scheduling Framework for Distributed Mixture-of-Experts Training. In Advances in Neural Information Processing Systems (NeurIPS). arXiv:2510.00207. [8] Raja Gond, Nipun Kwatra, and Ramachandran Ramjee. 2025. TokenWeave: Efficient Compute-Communication Overlap for Distributed LLM Inference. arXiv preprint arXiv:2505.11329 (2025). [9] Jiaao He, Jidong Zhai, Tiago Antunes, Haojun Wang, Fuwen Luo, Shangfeng Shi, and Qin Li. 2022. FasterMoE: Modeling and Optimizing Training of Large-Scale Dynamic Pre-Trained Models. In Proceedings of the 27th ACM SIGPLAN Symposium on Principles and Practice of Parallel Programming (PPoPP). [10] Ke Hong, Xiuhong Li, Minxu Liu, Qiuli Mao, Tianqi Wu, Zixiao Huang, Lufang Chen, Zhong Wang, Yichong Zhang, Zhenhua Zhu, Guohao Dai, and Yu Wang. 2025. FlashOverlap: A Lightweight Design for Efficiently Overlapping Communication and Computation. arXiv preprint arXiv:2504.19519 (2025). [11] Changho Hwang, Wei Cui, Yifan Xiong, Ziyue Yang, Ze Liu, Han Hu, Zilong Wang, Rafael Salas, Jithin Jose, Prabhat Ram, HoYuen Chau, Peng Cheng, Fan Yang, Mao Yang, and Yongqiang Xiong. 2023. Tutel: Adaptive Mixture-of-Experts at Scale. In Proceedings of Machine Learning and Systems (MLSys). arXiv:2206.03382. [12] Albert Q. Jiang, Alexandre Sablayrolles, Antoine Roux, et al. 2024. Mixtral of Experts. arXiv preprint arXiv:2401.04088 (2024). [13] Chenyu Jiang, Ye Tian, Zhen Jia, Shuai Zheng, Chuan Wu, and Yida Wang. 2024. Lancet: Accelerating Mixture-of-Experts Training via Whole Graph Computation-Communication Overlapping. In Proceedings of Machine Learning and Systems (MLSys). arXiv:2404.19429. [14] Hongyi Jin, Bohan Hou, Guanjie Wang, Ruihang Lai, Jinqi Chen, Zihao Ye, Yaxing Cai, Yixin Dong, Xinhao Cheng, Zhihao Zhang, Yilong Zhao, Yingyi Huang, Lijie Yang, Jinchen Jiang, Gabriele Oliaro, Jianan Ji, Xupeng Miao, Vinod Grover, Todd C Mowry, Zhihao Jia, and Tianqi Chen. 2025. EventTensor: A Unified Abstraction for Compiling Dynamic Megakernel. arXiv preprint arXiv:2604.13327 (2025). [15] Dmitry Lepikhin, HyoukJoong Lee, Yuanzhong Xu, Dehao Chen, Orhan Firat, Yanping Huang, Maxim Krikun, Noam Shazeer, and Zhifeng Chen. 2021. GShard: Scaling Giant Models with Conditional Computation and Automatic Sharding. In International Conference on Learning Representations (ICLR). [16] Jiamin Li, Yimin Jiang, Yibo Zhu, Cong Wang, and Hong Xu. 2023. Accelerating Distributed MoE Training and Inference with Lina. In USENIX Annual Technical Conference (ATC). [17] Lingxiao Ma, Zhiqiang Xie, Zhi Yang, Jilong Xue, Youshan Miao, Wei Cui, Wenxiang Hu, Fan Yang, Lintao Zhang, and Lidong Zhou. 2020. Rammer: Enabling Holistic Deep Learning Compiler Optimizations with rTasks. In OSDI. [18] Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Weijia Shi, Pete Walsh, Oyvind Tafjord, Nathan Lambert, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, David Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi, Noah A. Smith, Pang Wei Koh, Amanpreet Singh, and Hannaneh Hajishirzi. 2025. OLMoE: Open Mixture-of-Experts Language Models. In International Conference on

Learning Representations (ICLR). [19] Aniruddha Nrusimha, William Brandon, Mayank Mishra, Yikang Shen, Rameswar Panda, Jonathan Ragan-Kelley, and Yoon Kim. 2025. FlashFormer: Whole-Model Kernels for Efficient Low-Batch Inference. arXiv preprint arXiv:2505.22758 (2025). [20] Qwen Team. 2025. Qwen3 Technical Report. arXiv preprint arXiv:2505.09388 (2025). [21] Qwen Team. 2026. Qwen3.5: Towards Native Multimodal Agents. https://qwen.ai/blog?id=qwen3.5. Flagship MoE: Qwen3.5-397BA17B (397B total / 17B activated, 512 routed + 1 shared expert, top-10). [22] Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc V. Le, Geoffrey E. Hinton, and Jeff Dean. 2017. Outrageously Large Neural Networks: The Sparsely-Gated Mixture-ofExperts Layer. In International Conference on Learning Representations (ICLR). [23] Shaohuai Shi, Xinglin Pan, Xiaowen Chu, and Bo Li. 2023. PipeMoE: Accelerating Mixture-of-Experts through Adaptive Pipelining. In IEEE Conference on Computer Communications (INFOCOM). [24] Shaohuai Shi, Xinglin Pan, Qiang Wang, Chengjian Liu, Xiaozhe Ren, Zhongzhe Hu, Yu Yang, Bo Li, and Xiaowen Chu. 2024. ScheMoE: An Extensible Mixture-of-Experts Distributed Training System with Tasks Scheduling. In Proceedings of the 19th European Conference on Computer Systems (EuroSys). 236–249. doi:10.1145/3627703.3650083 [25] Benjamin Spector, Jordan Juravsky, Stuart Sul, Owen Dugan, Dylan Lim, Dan Fu, Simran Arora, and Christopher Ré. 2025. Look Ma, No Bubbles! Designing a Low-Latency Megakernel for Llama1B. https://hazyresearch.stanford.edu/blog/2025-05-27-no-bubbles. Hazy Research blog post. [26] Stuart H. Sul, Simran Arora, Benjamin F. Spector, and Christopher Ré. 2025. ParallelKittens: Systematic and Practical Simplification of Multi-GPU AI Kernels. arXiv preprint arXiv:2511.13940 (2025). [27] Renji Thomas, Kristin Barber, Naser Sedaghati, Li Zhou, and Radu Teodorescu. 2016. Core Tunneling: Variation-Aware Voltage Noise Mitigation in GPUs. In 2016 IEEE International Symposium on High Performance Computer Architecture (HPCA). 151–162. doi:10.1109/ HPCA.2016.7446061 [28] Shulai Zhang, Ningxin Zheng, Haibin Lin, Ziheng Jiang, Wenlei Bao, Chengquan Jiang, Qi Hou, Weihao Cui, Size Zheng, Li-Wen Chang, Quan Chen, and Xin Liu. 2025. Comet: Fine-grained Computationcommunication Overlapping for Mixture-of-Experts. In Proceedings of Machine Learning and Systems (MLSys). arXiv:2502.19811. [29] Chenggang Zhao, Zhean Xu, Liang Zhao, Jiashi Li, Chenhao Xu, Anyi Xu, Shengyu Liu, Kexing Zhou, and Kuai Yu. 2025. DeepGEMM: clean and efficient BLAS kernel library on GPU. https://github.com/ deepseek-ai/DeepGEMM. [30] Chenggang Zhao, Shangyan Zhou, Liyue Zhang, Chengqi Deng, Zhean Xu, Yuxuan Liu, Kuai Yu, Jiashi Li, and Liang Zhao. 2025. DeepEP: An Efficient Expert-Parallel Communication Library. https: //github.com/deepseek-ai/DeepEP. [31] Lianmin Zheng, Liangsheng Yin, Zhiqiang Xie, Chuyue Sun, Jeff Huang, Cody Hao Yu, Shiyi Cao, Christos Kozyrakis, Ion Stoica, Joseph E. Gonzalez, Clark Barrett, and Ying Sheng. 2024. SGLang: Efficient Execution of Structured Language Model Programs. In Advances in Neural Information Processing Systems (NeurIPS). arXiv:2312.07104. [32] Size Zheng, Wenlei Bao, Qi Hou, Xuegui Zheng, Jin Fang, Chenhui Huang, Tianqi Li, Haojie Duanmu, Renze Chen, Ruifan Xu, Yifan Guo, Ningxin Zheng, Ziheng Jiang, Xinyi Di, Dongyang Wang, Jianxi Ye, Haibin Lin, Li-Wen Chang, Liqiang Lu, Yun Liang, Jidong Zhai, and Xin Liu. 2025. Triton-distributed: Programming Overlapping Kernels on Distributed AI Systems with the Triton Compiler. arXiv:2504.19442 [cs.DC] https://arxiv.org/abs/2504.19442 [33] Size Zheng, Xuegui Zheng, Li-Wen Chang, and Jidong Zhai. 2025. UniEP: Unified Expert-Parallel MoE MegaKernel for LLM Training. arXiv preprint arXiv:2604.19241 (2025). 12

Weave [34] Guichao Zhu, Lintian Lei, Yuhao Qing, Yichao Fu, Fanxin Li, Dong Huang, Zekai Sun, and Heming Cui. 2025. FoldMoE: Efficient Long Sequence MoE Training via Attention-MoE Pipelining. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers). Association for Computational Linguistics, Vienna, Austria, 3705–3717. doi:10.18653/v1/2025.acl-long.186

[35] Donglin Zhuang, Zhen Zheng, Haojun Xia, Xiafei Qiu, Junjie Bai, Wei Lin, and Shuaiwen Leon Song. 2024. MonoNN: Enabling a New Monolithic Optimization Space for Neural Network Inference Tasks on Modern GPU-Centric Architectures. In OSDI.

13

Record · ID 1006820 · SHA-256 3b2e432b69acf5d3
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.