H IGHER - ORDER PRUNING OF EXPERTS IN MIXTURE OF - EXPERTS LANGUAGE MODELS Alex M. Tseng, Prannay Kaul, Luca Zancato, Wei Xia & Stefano Soatto AI Fundamental Research AWS Agentic AI {amtseng,prannayk,wxia,soattos}@amazon.com; [email protected]
arXiv:2609.18916v1 [cs.LG] 16 Sep 2026
A BSTRACT Mixture-of-Experts (MoE) language models suffer from large parameter counts, which create a significant memory bottleneck. Expert pruning is the most direct approach for reducing this parameter count, yet existing methods make pruning decisions for each expert independently, and assume experts’ contributions are purely additive. In reality, expert usage in MoEs is inherently cooperative. We derive HOPE (Higher-Order Pruning of Experts), a second-order pruning objective which provably minimizes an upper bound on the error resulting from pruning. We show that REAP (a state-of-the-art first-order pruning method) is a special case of HOPE where interaction terms are ignored. Across three frontier MoE models (up to 122B parameters), two distinct calibration sets, and multiple benchmarks (including math, instruction following, coding, and an agentic suite), we demonstrate that HOPE produces better pruning decisions than existing methods, and its advantage is most pronounced at high pruning rates and on challenging agentic workloads. At 50% pruning, HOPE outperforms all baselines and achieves an average rank of 1.58 out of 5 methods (versus 2.42 for the next-best method, REAP), with gains of up to +6.1% on agentic coding. Over all conditions, HOPE again achieves the best average rank and surpasses every other method in the majority of head-to-head comparisons. By preserving cooperative expert structure that first-order methods ignore, HOPE enables aggressive compression with minimal degradation, particularly on complex tasks where diverse expert combinations are invoked over long sequences.
1
I NTRODUCTION
Mixture-of-expert (MoE) language models have rapidly become the dominant architecture for stateof-the-art LLMs, such as DeepSeek-V3, Qwen3.5, and GLM-4.5 (DeepSeek-AI, 2024; Qwen Team, 2025; GLM Team, 2024). MoEs attempt to scale model capacity while reducing inference costs by multiplexing the feed-forward MLP in each layer into many “expert” MLPs. Each token is independently routed to a small subset of these specialized experts in each layer, thereby reducing the inference time for each token while still allowing the model to retain high capacity and capability. Importantly, however, although each token only activates a small subset of experts in each layer, the entire expert pool must reside in GPU memory. That is, the high parameter count of MoEs competes directly with the KV cache for valuable GPU capacity, thus limiting their practical deployment, including: 1) the need for expensive, high-memory hardware; 2) limited concurrent requests to a model; 3) reduced memory budget for context, tool calls, and reasoning. The vast majority of the parameter count (and thus memory footprint) of MoEs is precisely the experts themselves (116 billion of the 122 billion parameters in Qwen3.5-122B-A10B are expert MLPs, taking ∼ 240 GB). While quantization reduces the size of each parameter, it cannot eliminate redundant experts, and all experts still must reside in memory. Expert pruning takes an orthogonal approach: it permanently removes entire experts, freeing memory regardless of precision, and it composes with quantization (Lasby et al., 2025; Chen et al., 2025). Crucially, pruning is one-shot and requires no retraining. Pruning Qwen3.5-122B by 50% would immediately reduce its footprint 1
to 125 GB, freeing 115 GB for additional context and tool calls, longer reasoning traces, and deeper searches. This substantially expands the model’s effective capabilities on the same hardware. Current expert-pruning methods have already achieved promising results, but they are limited in their independent treatment of experts. These methods are first order, meaning they assign each expert a scalar importance score in isolation, ignoring interactions between experts. In reality, MoE inference is inherently combinatorial: each token activates a group of K experts whose outputs are jointly aggregated, and certain pairs of experts are co-selected far more often than independent chance would predict (Figure S2). First-order methods have a major failure mode arising from this cooperative structure. If an expert A is pruned, cooperative circuits it participates in are disrupted. Partner experts that are frequently co-selected with A may no longer combine coherently, potentially degrading performance beyond what removing A alone would suggest. At aggressive pruning rates (40–50%), many such cooperative pairs are broken simultaneously, compounding damage in ways that first-order importance scores cannot predict or prevent. Higher pruning rates are also the most impactful for deployment: halving the expert pool immediately frees enough memory for longer contexts and extended reasoning, or to accommodate substantially more concurrent requests. To this end, we derive HOPE (Higher Order Pruning of Experts), a second-order pruning objective which considers pairwise interactions between experts to inform pruning decisions. During calibration, we record expert usage, including interaction terms that reflect how pairs of experts complement each other, resulting in a square interaction matrix per layer which is solved as a quadratic program with minimal computational overhead. We show that REAP, a state-of-the-art first-order method, is a special case of HOPE where interaction terms are ignored. Across extensive experiments (3 MoE architectures up to 122B parameters, 2 calibration sets, 6 pruning rates, and multiple benchmarks), HOPE achieved the best performance most frequently and beat every state-of-the-art baseline in the majority of head-to-head comparisons. As predicted by our theory, HOPE’s margin over competing methods grows with pruning rate (particularly at 40–50%), with its advantage most pronounced in complex agentic tasks which rely on diverse expert combinations. Our analysis also confirms that HOPE’s advantage stems from selectively retaining experts with stronger cooperative structure—exactly the signal that first-order methods miss.
2
R ELATED W ORK
2.1
M IXTURE - OF - EXPERTS
In a Transformer-based MoE, the MLP in each layer is replaced by E expert MLPs and a learned P router. For a layer with E experts, the output is h(x) = k gk (x)fk (x), where fk (x) ∈ Rd is the output of expert k and gk (x) ∈ [0, 1) is its gate weight. The layer selects the top-K experts with the highest router logits (forming the selected set T (x)), and these logits are softmax-normalized into PE gate weights: gk (x) > 0 iff k ∈ T (x), with k=1 gk (x) = 1. This allows high total parameter count while activating far fewer parameters per token. Standard training of MoEs includes a loadbalancing loss (Shazeer et al., 2017) to encourage uniform expert usage in each layer, although empirically, utilization remains highly non-uniform across experts (and this enables expert pruning). 2.2
E XPERT PRUNING
Expert pruning permanently removes experts to reduce the memory footprint of MoEs without retraining. Existing methods follow a common recipe: pass a calibration set D through the model, collect per-expert usage statistics in each layer, and assign each expert k a scalar importance score Sk . A fixed fraction of experts per layer is then pruned according to the lowest Sk . These methods differ in the statistics which are collected and how Sk is defined. Frequency: prune the least-activated experts. Rank experts by Skfreq = |Xk |, where Xk is the set of tokens which activate expert k. This method is simple, but ignores expert outputs entirely. EAN (Expert Activation Norm) (JaiswalP et al., 2025): prune by the sum of expert output norms over active tokens. Rank experts by SkEAN = x∈Xk ∥fk (x)∥2 . REAP (Router-weighted Expert Activation Pruning) (Lasby et al., 2025): prune by gateweighted activation norm, averaged over active tokens. Rank experts by SkREAP = 2
1 |Xk |
P
x∈Xk gk (x)∥fk (x)∥2 .
This is generally considered the state-of-the-art pruning method for generative tasks, and is backed by theoretical justification. MAN (Mean Activation Norm) (Liu etPal., 2026): prune by activation norm, averaged over active tokens. Rank experts by SkMAN = |X1k | x∈Xk ∥fk (x)∥2 . This is very similar to REAP, but without the gate weighting. Liu et al. (2026) also placed existing first-order pruning methods into a common framework parameterized by routing frequency, gate weighting, and activation strength. All of these pruning methods are first-order: they score each expert independently using a scalar importance metric Sk , ignoring expert interactions. Note that alternatives to pruning—namely expert merging—have also been proposed to reduce MoE parameter count, but Lasby et al. (2025) showed both theoretically and empirically that expert merging leads to irreducible error and underperforms expert pruning. Intuitively, pruning (as opposed to merging) preserves the router’s ability to independently modulate surviving experts. 2.3
E XPERT INTERACTION AND COOPERATION
Non-pruning work has shown that expert co-selection patterns are stable and exploitable properties of MoEs. Flame-MoE (Kang et al., 2025) showed that expert co-selection patterns emerge early in training and stabilize throughout. Nguyen-Nhat et al. (2025) augmented the MoE router with an expert co-selection graph that encourages repeated activation of cooperative expert pairs, improving robustness. These results establish that pairs of experts cooperate—an exploitable property—but no prior work has leveraged this cooperative structure to inform pruning decisions.
3
HOPE: H IGHER -O RDER P RUNING OF E XPERTS
In this section, we briefly outline the theoretical justification behind the HOPE objective. The full derivation can be found in Appendix A. 3.1
S ETUP AND NOTATION
Consider an MoE layer with E experts. For input token x ∈ Rd , the layer output is h(x) = PE d k=1 gk (x)fk (x), where fk (x) ∈ R is the output of expert k, and gk (x) ∈ [0, 1) is the gate weight. The router maps x to a logit per expert and selects the K highest-logit experts to form the selected set T (x) ⊆ {1, . . . , E}, with |T (x)| = K. The logits of experts P in T (x) are softmax normalized to obtain gate weights gk (x), with gk (x) > 0 ⇔ k ∈ T (x) and k∈T (x) gk (x) = 1. Now suppose we delete a prune-set P ⊂ {1, . . . , E} of experts from the layer. For a token x, the pruned experts in P ∩ T (x) (i.e. those which would have been selected if not removed) are replaced by the next-highest-logit experts R(x), which are now in top-K (|R(x)| = |P ∩ T (x)|). The new selected set is (T (x) \ P ) ∪ R(x), with updated gate weights gk′ (x), with gk′ (x) > 0 ⇔ k ∈ PE (T (x) \ P ) ∪ R(x). The pruned layer’s output is h′ (x) = k=1 gk′ (x)fk (x). Our goal is to find the prune-set P (given a prescribed budget |P |) that minimizes the pruning error ∥h(x) − h′ (x)∥. 3.2
E RROR DECOMPOSITION
Following Lasby et al. (2025), we decompose the pruning error of a MoE layer into two components: X X X h(x) − h′ (x) = gj (x)fj (x) − gi′ (x)fi (x) + (gk (x) − gk′ (x))fk (x) j∈P ∩T (x)
i∈R(x)
k∈T (x)\P
The first term is the substitution error, which arises from swapping the pruned experts in P ∩T (x) for their replacements R(x). The second term is the renormalization error, which arises from rescaling the gates of the retained experts. As with Lasby et al. (2025), we note that the renormalization error is typically small (we also provide a proof of this in Appendix A), and so we focus on minimizing the substitution error. 3
3.3
T HE SECOND - ORDER OBJECTIVE
Our goal is to find the prune-set P which minimizes the substitution error: X
P ∗ = arg min P
gj (x)fj (x) −
j∈P ∩T (x)
2
X
gi′ (x)fi (x) 2
(1)
i∈R(x)
Importantly, we minimize the squared substitution error, which reveals interaction terms. The substitution error itself is intractable to optimize directly due to its reliance on R(x) and gi′ (x), which depend on the unknown prune-set P . We therefore define a tractable upper bound that depends only on pruned experts’ contributions (which are observable without committing to P ): Z=
X
gj (x) · ∥fj (x)∥2
2
(2)
j∈P ∩T (x)
Intuitively, Z measures the total joint contribution of the prune-set. Expanding the square reveals pairwise interaction terms between pruned experts. Note that Z depends on input tokens x and the prune-set P . This definition of Z admits the following upper bound on the substitution error: Theorem 1. For any pruning set P , the squared substitution error is bounded by: X
gj (x)fj (x) −
j∈P ∩T (x)
X
2
gi′ (x)fi (x) 2 ≤ (1 + ρ)2 · Z
max ∥fk (x)∥2 where
ρ=
i∈R(x)
k
min ∥fk (x)∥2 k
Z is a tractable surrogate whose minimization provably reduces the (squared) substitution error. The factor (1+ρ)2 (independent of P ) arises from bounding the contribution of replacement experts R(x), whose identity depends on P and is unknown at optimization time. This factor is conservative: it assumes a worst-case replacement whose output norm equals the layer maximum. In practice, the gap between Z and the true error is far smaller than the bound suggests. 3.4
M INIMIZING Z
Given the upper bound in Theorem 1, our goal is to find the prune-set P which minimizes Z. More precisely, we minimize the expectation of Z over the calibration set D: P ∗ = arg minP Ex∈D [Z]. To minimize Ex∈D [Z], we substitute the expanded form of Z (Equation 2) into the expectation, and replace the expectation with the empirical average. This minimization can equivalently be formulated as a binary quadratic program (QP) over the matrix F ∈ RE×E , defined as follows: Fi,j =
1 X gi (x)gj (x)∥fi (x)∥2 ∥fj (x)∥2 N
where
Xi,j = {x : i, j ∈ T (x)}
(3)
x∈Xi,j
The F -matrix is the core calibration result used by HOPE. F encodes both individual expert contributions (on the diagonal) and pairwise expert co-contributions (on the off-diagonal). High Fi,j implies that pruning experts i, j together incurs extra error from their reinforcing contributions. To formulate the QP, we introduce binary decision variables pk ∈ {0, 1} where pk = 1 iff expert k ∈ P . Theorem 2. Given a target pruning budget |P | (the number of experts to remove per layer), the prune-set P ∗ which minimizes Ex∈D [Z] is also the solution to the following quadratic program: ∗
⊤
P = arg min p F p
E
p ∈ {0, 1} ,
s.t.
p
E X
pk = |P |
k=1
Normalization: This formulation of F (Equation 3) minimizes E[Z] exactly (Theorem 2). In practice, we replace the unconditional average over the total number of tokens, N , with a conditional average, |Xi,j |. This follows the analogous design choice in REAP, which prevents rarely co-selected expert pairs (which nonetheless contribute strongly) from being undervalued. 4
3.5
C ONNECTION TO REAP
HOPE’s F matrix encodes individual expert importance scores on the diagonal, and this diagonal effectively recovers the scores used by REAP to independently rank experts for pruning. Specifically, Fk,k = E[(gk (x)∥fk (x)∥)2 ], whereas the squared REAP score is (SkREAP )2 = E[gk (x)∥fk (x)∥]2 . These expressions differ only by the variance of expert contributions over activating tokens. Indeed, HOPE with zeroed off-diagonals empirically recovers the same prune-set as REAP (Figure S4). In contrast with REAP, HOPE uses the full F , where Fi,j (i ̸= j) encodes the interaction between these experts when they are both activated. Thus, any improvement that HOPE achieves over REAP is directly attributable to these off-diagonal interaction terms which encode second-order effects. 3.6
P RACTICAL IMPLEMENTATION
As with other expert-pruning methods, HOPE requires a single forward pass over the calibration set. For each token, we record the active experts, their gate values, and their outputs. We also collect interaction terms, which are the product of the (gate-weighted) expert outputs when two experts are co-activated. For each layer, this gives us an (E × E) F -matrix. We solve each layer’s QP (using a standard QP solver) via a continuous relaxation with the conditions in Theorem 2, rounding the continuous solution to the top |P | entries. The relaxation is empirically tight: continuous solutions concentrate near 0 and 1, and the continuous–binary gap is negligible (Figure S5). In terms of computational cost, HOPE’s calibration step is dominated by forward passes (shared with all methods), and solving the QP adds negligible additional cost (1–2 seconds per layer, for typical values of E).
4
E XPERIMENTS
4.1
E XPERIMENTAL SETUP
• Models: Qwen3.5-122B-A10B (256 experts, top-8 routing, 48 MoE layers), Qwen3.5-35B-A3B (256 experts, top-8 routing, 40 MoE layers), GLM-4.5-Air (128 experts, top-8 routing with 1 shared expert, 45 MoE layers); note that shared experts are never pruned • Baselines: REAP (Lasby et al., 2025), EAN (Jaiswal et al., 2025), MAN (Liu et al., 2026), Frequency • Calibration data: Evol-CodeAlpaca-v1 (coding-focused) (Luo et al., 2023), SWE-Bench verified trajectories (agentic-focused) (Jimenez et al., 2023) • Pruning rates: 10%, 20%, 25%, 30%, 40%, 50% (prune rate of 10% means removing 10% of experts in each layer); for SWE-Bench Pro, we only test 10%, 25%, and 50% • Benchmarks: Tulu3 Dev suite (GSM8K, MATH, IFEval, MMLU, BBH, TruthfulQA, PopQA, LiveCodeBench) (Lambert et al., 2024; Jain et al., 2024), SWE-Bench Pro (Deng et al., 2025) We tested HOPE on several models ranging from 35 billion to well over 100 billion parameters, across multiple architectural families. We focused our benchmarks on agentic tasks, as well as standard coding, math, instruction following, general knowledge, etc. This gives 54 total configurations (3 models × 2 calibration sets × 6 pruning rates for Tulu, or 3 pruning rates for SWE-Bench Pro). 4.2
HOPE PERFORMANCE COMPARISON
We tested HOPE over a dense set of configurations and find that HOPE is the overall best-performing method, particularly at higher pruning rates (Figure 1). Over all conditions, HOPE had the best average rank of 2.07, finishing as the top-1 method or in the top-2 more often than any other method (top-1 in 39% of conditions, top-2 in 67% of conditions). Furthermore, in a head-to-head comparison, HOPE beat every other individual baseline in the majority of conditions, with an average head-to-head win rate of 73% (Figure S1, Tables S1–S6). At more aggressive pruning rates (40–50%), HOPE achieved an average rank of 1.58, and finished as the top-1 method in the majority of conditions. This advantage was particularly marked in the agentic benchmark, where HOPE beat REAP in all experiments at a high pruning rate (40–50%), with a mean gap of +2.8% and up to +6.1%. This result matches our theory: HOPE’s advantage over 5
c) 0.5
1 Tulu Mean
Average rank (smaller=better)
a)
2 3 4 5
HOPE REAP MAN EAN Freq
10%
0.4 HOPE REAP MAN EAN Freq Base
0.3
0.2 0.1
20%
25%
30%
40%
0.2
50%
0.3
Pruning Rate
0.4
0.5
0.4
0.5
Pruning Rate
1
SWE-bench Pro
Average rank
b) 2 3 4 5 10% 20% 25% 30% 40% 50%
10%
25%
Pruning Rate
50%
0.3 HOPE REAP MAN EAN Freq Base
0.2 0.1 0.0 0.1
0.2
0.3
Pruning Rate
Figure 1: HOPE achieves the best overall performance across pruning conditions, and dominates at aggressive pruning rates. a) Average rank of HOPE compared to other baseline pruning methods, averaged over models, calibration sets, and benchmarks (54 conditions total). b) Average ranks separated by benchmark: Tulu3 and SWE-Bench Pro. c) Performance vs pruning rate for representative conditions; dashed gray line indicates the unpruned base model’s performance. Note: pruned models occasionally match or slightly outperform the unpruned baseline; this is consistent with observations that calibration can act as implicit specialization, removing low-relevance experts whose contributions add noise to the target tasks (Dong et al., 2025).
first-order methods should grow with pruning rate. As more experts are removed, more interactions are broken, and the off-diagonal terms of F become more important. Higher pruning rates are also the most deployment-relevant settings, where memory savings are most crucial. HOPE consistently improved over REAP, and this directly validates that the off-diagonal interaction terms provide signal which improves performance beyond REAP (over all conditions and pruning rates, REAP had a top-1 rate of 15% and a top-2 rate of 48%). MAN was also a strong competitor overall at lower pruning rates, but it lost its edge at more aggressive pruning rates. Furthermore, HOPE is robust: it finished last in only 1/54 conditions (20%-pruned Qwen3.5-122B-A10B), where the gap to the best method was only -1.7%. In contrast, REAP had 5 last-place finishes and Freq had 43. HOPE’s worst deficit on any single condition was -4.5%, and its deficits are smaller and far less frequent, while its gains reach up to +6.1% (Figure S1, Tables S1–S6). Interestingly, HOPE’s advantage was most pronounced on agentic coding (SWE-bench Pro), code generation (LiveCodeBench), and compositional reasoning (BBH), where gains over the next-best method reached +2.0%, +19.6%, and +11.7% respectively. On simpler recall tasks (e.g. PopQA), the advantage of second-order pruning was smaller. This pattern is consistent with our theoretical motivation: tasks requiring diverse expert combinations over many tokens may benefit most from preserving cooperative structure. 4.3
S TRUCTURE OF HOPE’ S F - MATRIX
Having shown that HOPE outperforms first-order pruning methods—particularly at high pruning rates—we now examine the structure of the F -matrix that drives this advantage. Importantly, the off-diagonal entries of F are not noise: they exhibit noticeable substructure (Figure 2a), and their average magnitude is 33% of the diagonal (Figure 2b). Clusters of experts with high Fi,j tend to be co-selected and contribute jointly. These interaction terms vary by layer but are consistently present throughout the network. HOPE accounts for this structure that first-order methods ignore entirely, leading to substantially different prune-sets. 6
Figure 2: The interaction matrix F recorded by HOPE exhibits expert cooperation. a) F -matrix for GLM4.5-Air at three representative layers (early, middle, late). Experts have been hierarchically clustered per layer for visualization. Visible block structure indicates groups of experts with high pairwise interaction scores. b) Distribution of diagonal vs off-diagonal terms in F across all layers (GLM-4.5-Air). Off-diagonal (interactive) terms are substantial in magnitude (compared to the diagonal), showing that pairwise interactions are a nonnegligible component of the pruning objective.
Figure 3: HOPE preferentially retains experts with stronger cooperative structure. For each baseline, we classify experts by disagreement (kept only by HOPE, kept only by the baseline, or kept by both). Experts kept by HOPE tend to have higher cooperative score (measured as the mean F score with HOPE-retained experts).
Indeed, when HOPE and a first-order method disagree, HOPE preferentially retains experts with stronger interactions with other surviving experts (measured as the mean F score) (Figure 3). Mechanistically, the QP assigns a higher cost to pruning experts which participate in cooperative clusters, causing them to be retained even if their individual importance (diagonal of F ) is lower. In contrast, experts retained by a first-order method (e.g. REAP) but pruned by HOPE have high individual importance but lower interaction scores with other retained experts. As such, first-order methods tend to be misled by experts’ self-importance and retain them regardless of cooperative structure. HOPE’s interaction-aware pruning decisions lead to substantially distinct prune-sets compared to other methods, but HOPE’s prune-set is least distant from REAP and MAN (all three methods rely on expert outputs f (x) averaged over expert-activating tokens) (Figure 4a). As expected, prune-set overlap between methods increases with pruning rate: as more experts are pruned, all methods are forced to remove the agreed-upon unimportant experts, and thereby converge on a larger shared set. At the same time, however, the remaining disagreements at higher pruning rates become more critical: the impact of each marginal pruning decision affects a larger fraction of the surviving experts, who have fewer cooperative partners. Thus, HOPE’s interaction-aware decisions yield the largest gains in this high-prune-rate regime. Disagreement between HOPE and other methods is also distributed across all layers (Figure 4b), consistent with the observation that interactions are ubiquitous throughout the network (Figure 2a, Figure S2). This indicates that second-order pruning can be valuable across the network. 7
Figure 4: HOPE makes distinct pruning decisions from other baseline methods. a) Overlap between HOPE’s pruning set versus each baseline, as a function of pruning rate (Qwen3.5-122B-A10B). We measure overlap as the Jaccard index, due to the prune-set size varying over different prune rates. b) For the same model and calibration set, the percentage of disagreeing pruned experts (per layer) between HOPE versus other methods at a pruning rate of 50%. Disagreement is distributed across the full network.
Figure 5: Prune-set stability across calibration sets and trials. a) Jaccard index across independent trials of the same method (averaged over models, pruning rates, and calibration sets). b) Jaccard index between prune-sets produced from different calibration sets: Evol-CodeAlpaca and SWE-Bench verified. Jaccard index is averaged over models, pruning rates, and trials (higher is more stability).
4.4
C ALIBRATION AND TRIAL ROBUSTNESS
Expert cooperation—as captured by HOPE’s F -matrix—is a stable property of the model, not an artifact of calibration noise. Over 3 random trials, prune-set stability is high (> 0.95 Jaccard similarity on average) for all methods including HOPE (Figure 5a, Figure S3). This confirms that the interaction structure HOPE exploits is reliably recoverable from calibration data. As expected, stability across different calibration sets is somewhat lower: different calibration domains lead to slightly different pruning decisions (Figure 5b). This effect is shared across all methods, and HOPE is not more sensitive than REAP (HOPE’s average calibration Jaccard index is 0.60, versus REAP’s 0.61). 4.5
C OMPATIBILITY WITH DOWNSTREAM TRAINING
Finally, we show that HOPE-pruned models serve as strong starting points for downstream training. Crucially, downstream training is not required to realize HOPE’s advantage; we include this experiment to show that HOPE’s structural advantage persists even after training. On REAP- or HOPE-pruned checkpoints of Qwen3.5-122B-A10B, we performed SFT on LiveCodeBench traces (collected from the unpruned model) for 1 billion tokens of training. HOPE outperformed REAP both before and after SFT (Figure 6): that is, fine-tuning narrows the gap but does not close it. This suggests that the expert structure preserved by second-order pruning provides lasting benefits that fine-tuning alone cannot replicate.
8
Figure 6: SFT results for pruned Qwen3.5-122B-A10B models. Pre-SFT and post-SFT use matched trials for a fair comparison. The gray dashed line indicates the performance of the unpruned model.
5
C ONCLUSION
We presented HOPE, the first theoretically derived expert-pruning objective that directly accounts for pairwise expert interactions, formulated as a quadratic program over an interaction matrix. We showed that HOPE generalizes REAP and outperforms it (along with other methods), especially at aggressive pruning rates where interaction effects dominate. Our analysis confirms that HOPE’s pruning decisions are largely influenced by off-diagonal interaction terms, which direct the selective retention of experts with high cooperative scores. Our results validate our theoretical prediction: considering these interactive terms—a signal which all first-order methods miss—in expert-pruning decisions yields better prune-sets. Importantly, HOPE’s advantage over other methods is greatest in the high-pruning-rate regime, which is most relevant to real-world deployment. Limitations: Calibration with HOPE does require additional time and memory due to its quadratic nature, although the additional cost is small. HOPE’s calibration step takes about 6–7% longer than REAP on the same-size calibration set (Table S7). The additional memory cost of HOPE to store the F -matrix is also negligible for the number of experts we typically have (128–256). HOPE does require marginally more data than REAP due to its estimating inter-expert (quadratic) in addition to per-expert statistics. The difference, however, is minor: to converge to 99% agreement of the same prune-set, REAP requires 8k prompts whereas HOPE requires 10k prompts (Qwen3.5-122B-A10B, calib: Evol-CA, rate: 50%) (Figure S6). Finally, like most other methods, HOPE prunes experts in each layer independently, and does not account for cross-layer interactions. Future work: We offer a theoretical framework (and preliminary results) for extending HOPE to cross-layer interactions and cross-layer pruning (Appendix A). Cross-layer HOPE enables nonuniform pruning budgets in each layer, potentially allocating more aggressive pruning to layers with weaker cooperative structure. In practice, however, this framework may introduce non-uniform budget-allocation instability (e.g. over-pruning layers with weak cooperative structure), which may require some additional constraints to address. Additionally, as MoEs continue to scale to larger models—where aggressive pruning is even more important for deployment—HOPE’s advantage could grow even further. Finally, we note that HOPE operates only up to second-order (pairwise) interactions, which already demonstrably provide signal beyond first-order methods. Extending HOPE to even higher-order interactions could be valuable, although they are combinatorially difficult to quantify.
9
AI USE STATEMENT In this work, we used generative AI to: • Search existing literature for relevant directions and baselines • Check human-written code for bugs • Write the first draft of code specifically for creating plots/figures • Suggest rewordings for specific sentences in the manuscript • Insert citations into the manuscript and download *.bib citations • Reformat CSV files of collected data/results into LATEX-formatted tables (via an AI-generated script) We did not use AI to write any code which created any main results, other than to identify bugs (which were subsequently verified and fixed by a human). AI was used to write the first draft of code for creating plots and figures, and this code was then verified and revised by a human. We did not use AI to generate research ideas, derive theoretical proofs, design experiments, or interpret results. We did not use AI to write any part of the text for the manuscript, other than as a “copy-editor” to search for typos, insert citations, and suggest rewordings for specific sentences or sections when prompted. AI was used to automate the formatting of LATEX tables from saved CSV files. We take responsibility for the final content of this work. E THICS STATEMENT This research focuses on algorithmic compression of MoE language models. All datasets (calibration or evaluation) are publicly available and do not contain personal information. All models used are also publicly available and open source. Our work’s contribution is methodological: we derived and demonstrated a method for reducing the memory footprint of MoEs. Although we hope our work will improve accessibility of generative AI, we also acknowledge that such algorithms could potentially lower barriers to deploying models/agents for harmful uses, as well. Additionally, we pruned pre-trained models, and any biases already present in those models may persist after pruning. In general, we believe our work does not introduce any new ethical considerations beyond those already present in the efficient deployment of LLMs. R EPRODUCIBILITY STATEMENT Code for HOPE calibration, QP solving, and all analyses will be released upon publication. The theoretical derivation of HOPE is provided in full in Appendix A. Section 4.1 details the experimental setup, including models, calibration datasets, pruning rates, and evaluation benchmarks. Appendix D provides supplementary methods describing all analysis procedures in sufficient detail for reproduction, including hyperparameters, solver settings, and evaluation configurations. All experiments use publicly available models and datasets.
10
R EFERENCES Yixiao Chen, Yanyue Xie, Ruining Yang, Wei Jiang, Wei Wang, Yong He, Yue Chen, Pu Zhao, and Yanzhi Wang. Collaborative compression for large-scale moe deployment on edge. arXiv preprint arXiv:2509.25689, 2025. DeepSeek-AI. Deepseek-v3 technical report. arXiv preprint arXiv:2412.19437, 2024. Xiang Deng, Jeff Da, Edwin Pan, Yannis Yiming He, Charles Ide, Kanak Garg, Niklas Lauffer, Andrew Park, Nitin Pasari, Chetan Rane, Karmini Sampath, Maya Krishnan, Srivatsa Kundurthy, Sean Hendryx, Zifan Wang, Vijay Bharadwaj, Jeff Holm, Raja Aluri, Chen Bo Calvin Zhang, Noah Jacobson, Bing Liu, and Brad Kenstler. Swe-bench pro: Can ai agents solve long-horizon software engineering tasks? arXiv preprint arXiv:2509.16941, 2025. Zican Dong, Han Peng, Peiyu Liu, Wayne Xin Zhao, Dong Wu, Feng Xiao, and Zhifeng Wang. Domain-specific pruning of large mixture-of-experts models with few-shot demonstrations. arXiv preprint arXiv:2504.06792, 2025. GLM Team. Chatglm: A family of large language models from glm-130b to glm-4 all tools. arXiv preprint arXiv:2406.12793, 2024. Naman Jain, King Han, Alex Gu, Wen-Ding Li, Fanjia Yan, Tianjun Zhang, Sida Wang, Armando Solar-Lezama, Koushik Sen, and Ion Stoica. Livecodebench: Holistic and contamination free evaluation of large language models for code. arXiv preprint arXiv:2403.07974, 2024. Ajay Jaiswal, Junda Wang, Ying Li, Ping Li, Tianlong Chen, Zhangyang Wang, Cong Wang, and Xinya Du. Finding fantastic experts in MoEs: A unified study for expert dropping strategies and observations. arXiv preprint arXiv:2504.05586, 2025. Carlos E. Jimenez, John Yang, Alexander Wettig, Shunyu Yao, Kexin Pei, Ofir Press, and Karthik Narasimhan. Swe-bench: Can language models resolve real-world github issues? arXiv preprint arXiv:2310.06770, 2023. Hao Kang, Zichun Yu, and Chenyan Xiong. Flame-moe: A transparent end-to-end research platform for mixture-of-experts language models. arXiv preprint arXiv:2505.20225, 2025. Nathan Lambert, Jacob Morrison, Valentina Pyatkin, Shengyi Huang, Hamish Ivison, Faeze Brahman, Lester James V. Miranda, Alisa Liu, Nouha Dziri, Shane Lyu, Yuling Gu, Saumya Malik, Victoria Graf, Jena D. Hwang, Jiangjiang Yang, Ronan Le Bras, Oyvind Tafjord, Chris Wilhelm, Luca Soldaini, Noah A. Smith, Yizhong Wang, Pradeep Dasigi, and Hannaneh Hajishirzi. Tulu 3: Pushing frontiers in open language model post-training. arXiv preprint arXiv:2411.15124, 2024. Mike Lasby, Ivan Lazarevich, Nish Sinnadurai, Sean Lie, Yani Ioannou, and Vithursan Thangarasa. Reap the experts: Why pruning prevails for one-shot moe compression. arXiv preprint arXiv:2510.13999, 2025. Zongfang Liu, Jinghui Zhang, Zijian Ma, Guangyi Chen, and Xin Yuan. How to score experts for one-shot moe expert pruning: A unified formulation and selection principle. arXiv preprint arXiv:2606.15716, 2026. Ziyang Luo, Can Xu, Pu Zhao, Qingfeng Sun, Xiubo Geng, Wenxiang Hu, Chongyang Tao, Jing Ma, Qingwei Lin, and Daxin Jiang. Wizardcoder: Empowering code large language models with evol-instruct. arXiv preprint arXiv:2306.08568, 2023. Minh-Khoi Nguyen-Nhat, Rachel S. Y. Teo, Laziz Abdullaev, Maurice Mok, Viet-Hoang Tran, and Tan Minh Nguyen. Modeling expert interactions in sparse mixture of experts via graph structures. arXiv preprint arXiv:2510.16411, 2025. Qwen Team. Qwen3 technical report. arXiv preprint arXiv:2505.09388, 2025. Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. In International Conference on Learning Representations, 2017.
11
A
F ULL DERIVATION OF HOPE
A.1
S ETUP
For an arbitrary MoE layer with E experts, the output of that layer on input-token representation x is: E X h(x) = gk (x)fk (x) k=1
where gk (x) ∈ [0, 1) is the softmax-normalized gate weight and fk (x) ∈ Rd is the expert output. E P gk (x) = 1. We have gk (x) > 0 iff k ∈ T (x) (the top-K selected set), and k=1
Now suppose we prune a set of experts P . For any token x, the pruned experts P ∩ T (x) are replaced by the next-highest-logit experts R(x), with |R(x)| = |P ∩ T (x)|. After pruning, the new gate weights are gk′ (x) > 0 iff k ∈ (T (x) \ P ) ∪ R(x), and the new layer output is: h′ (x) =
E X
gk′ (x)fk (x)
k=1
A.2
E RROR DECOMPOSITION
The pruning error decomposes as: X X h(x) − h′ (x) = gj (x)fj (x) − gi′ (x)fi (x) + j∈P ∩T (x)
i∈R(x)
|
{z
(gk (x) − gk′ (x))fk (x)
k∈T (x)\P
}
substitution error
X |
{z
renormalization error
}
The first term is the substitution error and the second is the renormalization error. The substitution error is the direct effect of swapping experts, and is the dominant source of error. The renormalization error is the indirect effect of pruning on untouched experts. Because gate values gk (x) must sum to 1, pruning and replacing experts causes the gate values of non-pruned experts in the top-K set (T (x) \ P ) to be rescaled. This particular decomposition was described in (Lasby et al., 2025). From this point forward, the derivations will diverge substantially. A.3
G ATE - ORDERING LEMMA
An intermediate result which will be helpful is to show that replacement experts in R(x) have lower gate values. More formally: Lemma 3 (Replacement experts have lower gate values). Under top-K routing with softmax normalization, suppose experts P ∩ T (x) are pruned and replaced with experts R(x). Then gi′ (x) ≤ gj (x) for all j ∈ P ∩ T (x) and i ∈ R(x). To prove this lemma, we first prove the single-expert case and then extend by induction. Single expert. Suppose expert j ∈ T (x) is pruned and replaced by expert i ∈ / T (x). The original gate value is: exp(lj (x)) gj (x) = P exp(lk (x)) k∈T (x)
where lk (x) denotes the router logit for expert k (this is simply softmax normalization). After pruning expert j and replacing it with expert i, the new selected set is T ′ (x) = (T (x) \ {j}) ∪ {i}, and: exp(li (x)) gi′ (x) = P exp(lk (x)) − exp(lj (x)) + exp(li (x)) k∈T (x)
12
We wish to show that gi′ (x) ≤ gj (x), i.e.: P
exp(li (x)) ≤ exp(lk (x)) − exp(lj (x)) + exp(li (x))
k∈T (x)
exp(lj (x)) P exp(lk (x)) k∈T (x)
Cross-multiplying and simplifying, this is equivalent to showing: X [exp(li (x)) − exp(lj (x))] · exp(lk (x)) − exp(lj (x)) ≤ 0 k∈T (x)
Since expert j was selected and i was not (prior to pruning), we have li (x) ≤ lj (x), making the first factor non-positive. The second factor is a sum of positive terms minus one of those terms, which is positive. Thus the product is non-positive, completing the proof for a single expert. Multiple experts. We apply the single-expert result inductively. Consider pruning experts one at a time. At each stage, the conditions for the single-expert case hold (the expert being pruned was in the current top-K, and its replacement was not). By exchangeability of the pruning order and the fact that li (x) ≤ lj (x) for any i ∈ R(x) and j ∈ P ∩ T (x), the inequality gi′ (x) ≤ gj (x) holds for all such pairs. A.4
R ENORMALIZATION ERROR IS SMALL
Recall, the pruning error (which we wish to minimize) decomposes into the substitution error and renormalization error. Next, let us show that the renormalization error is generally small. For simplicity, we assume the new gate values are renormalized directly from the old ones (without retaking the softmax): gk (x) P P gk′ (x) = 1 − j∈P gj (x) + i∈R(x) gi′ (x) Substituting into the renormalization error: P P ′ X i∈R(x) gi (x) − j∈P gj (x) ′ P P (gk (x) − gk (x))fk (x) = ′ i∈R(x) gi (x) − j∈P gj (x) + 1 k∈T (x)\P
X
gk (x)fk (x)
k∈T (x)\P
P P Note that the scalar factor in front is small when j∈P gj (x) ≈ i∈R(x) gi′ (x). Generally, this is expected because we typically prune experts with the lower gate probabilities among the top-K, and replace them with other experts whose gate probabilities barely missed the cut-off for top-K. We therefore focus on minimizing the substitution error. A.5
E XPANDING THE SQUARED SUBSTITUTION ERROR
We focus on minimizing the squared substitution error, as this reveals interaction terms between pruned experts. We decompose the quadratic form into three components: 2
X
X
gj (x)fj (x) −
j∈P ∩T (x)
X
gj (x)fj (x)
−2
2
j∈P ∩T (x)
|
2
i∈R(x) 2
=
gi′ (x)fi (x)
{z A
}
X
gj (x)fj (x),
j∈P ∩T (x)
X
2
gi′ (x)fi (x) +
X
i∈R(x)
|
{z B
gi′ (x)fi (x) 2
i∈R(x)
}
|
{z C
For ease of notation, we split the substitution error into three components: A + B + C. Intuitively, A is the contribution of the pruned experts. Note A is quadratic, and the cross-terms of A captures how much pruned experts reinforce each other. If two pruned experts contribute cooperatively, their joint removal incurs a penalty which is worse than their individual removals. 13
}
C is the contribution of the replacement experts, and it measures the contribution of these experts to the output. C also contributes to error because the replacements are not perfectly aligned with what was pruned. Finally, B is the alignment between pruned experts and retained experts. If replacement experts are well-aligned with the pruned experts, B is large in magnitude and contributes a large negative to the substitution error. A.6
U PPER - BOUNDING THE SUBSTITUTION ERROR VIA Z
√ √ By the Cauchy–Schwarz inequality, we have that B ≤ |B| ≤ 2 A C. Substituting, we have the following inequality: X X √ √ √ √ 2 2 gj (x)fj (x) − gi′ (x)fi (x) 2 = A + B + C ≤ A + 2 A C + C = A+ C j∈P ∩T (x)
i∈R(x)
We now upper-bound both
√
A and
√
C in terms of a common quantity. Define: X 2 Z= gj (x)∥fj (x)∥2 j∈P ∩T (x)
Bounding
√
A:
By the triangle inequality: X √ A= j∈P ∩T (x)
Bounding
√
√
X
gj (x)fj (x) 2 ≤
gj (x)∥fj (x)∥2 =
Z
j∈P ∩T (x)
C:
Again by the triangle inequality: X X √ C= gi′ (x)fi (x) 2 ≤ gi′ (x)∥fi (x)∥2 i∈R(x)
i∈R(x)
Note there is a natural bijective mapping between pruned experts j ∈ P ∩ T (x) and replacement experts i ∈ R(x). Let rj be the identity of the expert i which replaces pruned expert j. By Lemma 3, gi′ (x) ≤ gj (x) for all j ∈ P ∩ T (x) and i ∈ R(x). Therefore: X
X
gi′ (x)∥fi (x)∥2 ≤
gj (x)∥frj (x)∥2 =
j∈P ∩T (x)
i∈R(x)
X
gj (x)∥fj (x)∥2
j∈P ∩T (x)
∥frj (x)∥2 ∥fj (x)∥2
Note the swapping of notation from fi (x) to frj (x). We can now remove the remaining dependency on replacement experts rj by bounding the possible magnitudes of expert contributions:
X
gj (x)∥fj (x)∥2
j∈P ∩T (x)
∥frj (x)∥2 ∥frj (x)∥2 X ≤ max gj (x)∥fj (x)∥2 ∥fj (x)∥2 j∈P ∩T (x) ∥fj (x)∥2 j∈P ∩T (x) ∥frj (x)∥2 √ = max Z j∈P ∩T (x) ∥fj (x)∥2
And we note that: max
∥frj (x)∥2
j∈P ∩T (x) ∥fj (x)∥2
≤
maxk ∥fk (x)∥2 mink ∥fk (x)∥2
This simple substitution provides a loose (but strict) upper bound, without relying on rj or P , which √ √ k ∥fk (x)∥2 makes it tractable to optimize. Thus, C ≤ ρ Z, where ρ = max mink ∥fk (x)∥2 . 14
Together X
: X
gj (x)fj (x) −
j∈P ∩T (x)
√ √ √ √ 2 gi′ (x)fi (x) 2 ≤ ( A + C)2 ≤ ( Z + ρ Z)2 = (1 + ρ)2 · Z
i∈R(x)
This completes the proof of Theorem 1. A.7
M INIMIZING Z AS A QUADRATIC PROGRAM
Our goal is now to minimize E [Z] over the calibration set D. We introduce binary decision x∼D
variables pk ∈ {0, 1}, where pk = 1 iff expert k ∈ P , and indicator variables 1[k ∈ T (x)] which denote if expert k is in the top-K. Substituting: E X
Z=
pj · 1[j ∈ T (x)] · gj (x)∥fj (x)∥2
2
j=1
=
E E X X
pj pk · 1[j ∈ T (x)] · 1[k ∈ T (x)] · gj (x)gk (x)∥fj (x)∥2 ∥fk (x)∥2 .
j=1 k=1
Taking the empirical expectation over the calibration set (with N tokens): E
E[Z] = x
E
1 XXX pj pk · 1[j ∈ T (x)] · 1[k ∈ T (x)] · gj (x)gk (x)∥fj (x)∥2 ∥fk (x)∥2 N x j=1 k=1
=
E X E X j=1 k=1
pj pk
1 X gj (x)gk (x)∥fj (x)∥2 ∥fk (x)∥2 N x∈Xj,k {z } | Fj,k
⊤
= p Fp where we define: Fi,j =
1 X gi (x)gj (x)∥fi (x)∥2 ∥fj (x)∥2 , N
Xi,j = {x : i, j ∈ T (x)}
x∈Xi,j
and p is a binary vector of size E. Note that in practice, we use conditional normalization in F instead of unconditional normalization. This equates to the following definition of F : X 1 gi (x)gj (x)∥fi (x)∥2 ∥fj (x)∥2 Fi,j = |Xi,j | x∈Xi,j
This is the empirical choice which we found helps decouple interaction strength from co-selection frequency. This helps ensure that specialist expert pairs which are rarely co-selected—but contribute strongly—are not undervalued, consistent with the analogous choice made by REAP. Note that if an expert pair is never activated (Xi,j = ∅), then Fi,j = 0. Therefore, finding the optimal prune-set of size |P | reduces to the binary quadratic program: P∗ =
arg min P
p∈{0,1}E ,
pk =|P |
k
This completes the proof of Theorem 2.
15
p⊤ F p
B
S UPPLEMENTARY F IGURES AND TABLES
Figure S1: Full performance results across all models and calibration sets. Each row corresponds to a model architecture (Qwen3.5-122B, Qwen3.5-35B, GLM-4.5-Air). Columns alternate between mean Tulu score and SWE-Bench Pro score for each calibration set (Evol-CodeAlpaca and SWE-Bench verified). The dashed gray line indicates the performance of the unpruned base model. HOPE (blue) is competitive or best at all pruning rates, with the cleanest separation at aggressive pruning rates.
16
Calib
Method
Rate
Base
Evol-CA
SB-verif
GSM8K
MATH
IFEval
PopQA
MMLU
BBH
TQA
LCB
Mean
92.17±0.36
41.54±0.37
21.21±1.71
0.35±0.00
65.15±0.31
40.34±0.88
65.92±0.26
64.30±0.61
48.87±0.56
HOPE REAP MAN EAN Freq
10% 10% 10% 10% 10%
94.44±0.36 93.43±1.89 95.20±0.71 94.44±0.36 93.18±0.62
42.61±0.83 42.37±1.15 42.71±0.94 42.13±1.47 42.64±2.04
19.39±0.86 20.00±1.48 18.79±0.86 18.79±0.86 19.39±1.71
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
63.87±0.25 62.74±0.94 63.86±0.45 62.27±0.81 62.94±0.21
41.51±0.83 39.37±0.36 41.51±1.65 39.88±1.44 39.47±0.38
63.13±0.41 64.53±0.81 64.64±0.12 64.33±0.48 63.39±0.90
62.63±0.66 64.84±1.84 63.44±0.88 63.98±0.73 66.13±1.52
48.49±0.53 48.45±1.06 48.81±0.70 48.27±0.77 48.44±0.92
HOPE REAP MAN EAN Freq
20% 20% 20% 20% 20%
94.70±0.62 92.42±0.00 92.93±0.71 92.42±1.24 93.18±0.62
42.96±0.59 42.13±1.36 43.13±1.04 41.81±0.13 40.10±0.40
18.18±0.00 18.79±0.86 20.61±0.86 18.79±1.71 18.79±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
68.01±0.86 64.70±0.92 68.51±0.71 59.84±0.92 59.95±0.29
39.62±0.69 40.24±0.83 41.62±1.71 37.42±1.11 38.75±0.64
63.71±0.16 63.87±0.39 63.26±0.15 61.32±0.04 62.85±0.21
62.96±0.53 62.85±0.40 64.52±1.21 64.57±0.27 65.65±0.80
48.81±0.43 48.17±0.59 49.36±0.80 47.07±0.68 47.45±0.48
HOPE REAP MAN EAN Freq
25% 25% 25% 25% 25%
93.69±0.36 93.69±0.71 94.19±0.71 91.41±0.71 91.92±1.89
40.34±1.73 43.84±1.70 43.47±0.54 41.52±0.39 37.82±1.07
18.18±0.00 18.79±0.86 18.79±0.86 21.82±0.00 17.58±1.71
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
64.50±0.70 66.12±4.40 63.16±0.28 58.35±0.77 48.53±0.34
40.59±0.72 38.39±1.33 38.80±0.38 38.04±1.09 33.64±0.69
64.74±0.15 64.20±1.16 64.78±0.21 59.42±0.27 59.33±0.34
65.70±1.19 64.46±0.33 63.66±0.73 61.40±0.40 61.72±0.33
48.51±0.61 48.73±1.31 48.40±0.46 46.54±0.45 43.86±0.80
HOPE REAP MAN EAN Freq
30% 30% 30% 30% 30%
95.20±0.71 94.44±0.71 94.44±0.36 92.42±0.00 91.16±0.36
40.34±1.62 42.74±1.24 41.07±0.93 39.09±1.03 21.57±0.64
18.18±1.48 18.18±0.00 19.39±1.71 16.36±0.00 15.76±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
66.27±1.42 64.75±0.39 64.69±1.06 42.74±0.54 35.00±1.17
36.15±1.35 39.01±0.69 38.75±1.55 35.43±0.25 33.49±0.51
65.26±0.03 62.79±0.27 64.05±0.41 56.33±0.56 54.32±0.60
62.63±1.07 65.32±1.08 60.81±0.95 30.70±1.41 21.24±1.22
48.05±0.96 48.45±0.55 47.94±0.87 39.18±0.47 34.11±0.67
HOPE REAP MAN EAN Freq
40% 40% 40% 40% 40%
93.69±0.36 93.18±0.62 93.94±0.62 92.17±0.94 91.67±0.62
40.27±0.48 39.43±0.96 36.90±1.06 39.29±0.63 19.28±0.66
18.18±0.00 18.18±1.48 19.39±0.86 14.55±0.00 15.15±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
60.03±0.31 52.39±1.16 49.86±0.10 32.37±3.02 33.60±1.56
36.20±0.45 37.01±1.02 38.09±0.07 33.28±0.13 31.08±1.25
64.40±0.51 62.82±0.41 64.31±0.26 57.08±0.48 59.64±0.87
62.42±1.15 41.94±1.05 42.80±1.12 37.63±3.71 37.80±0.66
46.94±0.41 43.16±0.84 43.20±0.51 38.34±1.11 36.07±0.81
HOPE REAP MAN EAN Freq
50% 50% 50% 50% 50%
94.19±0.36 93.94±0.00 93.69±0.36 91.16±1.29 80.05±0.71
37.86±1.99 44.19±1.56 41.76±1.25 35.86±0.88 13.55±0.73
18.18±1.48 20.00±1.48 17.58±2.27 15.76±0.86 16.97±1.71
10.21±0.41 2.57±1.89 3.85±0.26 0.47±0.03 4.70±0.60
29.77±2.08 50.15±1.35 45.11±0.88 27.01±0.96 24.12±0.40
37.99±1.01 46.06±7.36 40.44±0.92 36.25±1.00 31.95±1.16
67.20±0.10 58.59±4.28 55.62±0.83 59.82±0.11 52.90±0.65
48.01±4.04 28.49±7.44 34.84±1.60 43.33±0.77 33.66±2.74
42.93±1.44 43.00±3.17 41.61±1.05 38.71±0.74 32.24±1.09
HOPE REAP MAN EAN Freq
10% 10% 10% 10% 10%
94.44±0.71 95.45±0.62 94.44±0.94 91.92±0.71 91.16±0.36
42.07±0.58 39.36±0.75 42.32±1.49 39.88±0.13 34.23±1.44
20.61±1.71 18.79±0.86 19.39±0.86 20.00±0.00 18.79±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
64.77±0.50 63.01±0.54 61.87±0.47 65.28±0.62 66.79±0.45
43.81±0.38 43.00±0.29 43.20±1.16 36.66±0.57 37.37±1.79
63.37±0.36 64.33±0.19 64.21±0.32 62.28±0.32 63.10±0.23
62.26±0.99 62.80±0.53 62.04±1.14 64.95±0.42 61.13±0.92
48.96±0.66 48.39±0.47 48.48±0.80 47.66±0.35 46.62±0.76
HOPE REAP MAN EAN Freq
20% 20% 20% 20% 20%
93.43±0.36 94.44±0.36 93.94±0.62 84.85±1.24 26.01±1.29
41.53±0.84 41.66±1.05 40.42±1.23 33.78±0.24 13.33±0.81
19.39±1.71 18.79±0.86 18.18±0.00 16.97±0.86 19.39±2.27
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
58.61±0.31 63.61±0.20 67.66±0.08 58.39±0.37 62.18±0.13
43.40±1.11 40.39±0.69 41.56±1.03 34.56±0.83 34.92±0.52
64.10±0.38 63.74±0.18 64.40±0.46 62.66±0.13 61.80±0.36
63.92±1.29 61.99±0.38 61.45±1.84 55.43±0.30 45.75±0.20
48.09±0.75 48.12±0.46 48.50±0.66 43.37±0.50 32.97±0.70
HOPE REAP MAN EAN Freq
25% 25% 25% 25% 25%
93.69±0.94 92.68±0.94 93.69±0.36 83.84±0.36 15.40±2.17
40.20±2.08 38.00±1.84 39.97±1.58 28.44±1.23 8.61±0.64
19.39±0.86 19.39±0.86 18.79±0.86 19.39±2.27 20.61±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
57.97±0.25 62.49±0.47 67.24±0.38 63.49±0.59 61.20±0.33
38.85±0.32 39.52±1.14 39.01±0.44 32.31±0.47 33.38±0.07
66.33±0.41 64.61±0.47 63.79±0.27 61.60±0.48 62.25±0.44
61.67±0.27 59.62±0.79 58.82±1.33 54.14±1.24 41.56±1.19
47.31±0.64 47.08±0.82 47.71±0.65 42.94±0.83 30.42±0.71
HOPE REAP MAN EAN Freq
30% 30% 30% 30% 30%
92.68±0.36 93.43±0.94 94.70±0.62 85.61±0.62 10.86±0.36
40.05±1.17 40.19±1.01 39.30±1.24 27.13±1.17 6.67±0.65
20.61±0.86 20.61±0.86 20.00±1.48 18.18±0.00 18.18±0.00
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
54.59±0.25 60.37±0.82 61.30±0.55 58.96±0.53 60.38±0.25
39.52±1.70 39.31±0.07 38.70±0.88 30.47±0.29 34.87±0.96
62.95±0.22 66.95±0.18 65.75±0.08 59.46±0.82 60.49±0.50
60.38±1.22 63.49±0.27 61.83±0.88 45.05±0.20 32.20±0.68
46.39±0.72 48.09±0.52 47.74±0.72 40.65±0.45 28.00±0.42
HOPE REAP MAN EAN Freq
40% 40% 40% 40% 40%
93.18±0.62 91.92±0.94 93.43±0.71 84.34±0.71 2.78±1.29
42.46±0.63 39.93±0.59 40.21±0.37 23.72±0.23 1.87±0.07
18.18±1.48 19.39±0.86 20.00±0.00 19.39±1.71 20.61±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
57.58±1.01 56.97±2.93 50.63±0.61 48.38±0.37 46.31±3.32
34.97±0.82 36.45±1.82 33.28±0.76 32.98±0.66 29.60±1.52
59.43±0.21 63.47±0.28 64.05±0.55 59.35±0.48 61.35±0.26
58.92±0.27 57.04±0.90 56.02±1.59 39.03±0.35 16.77±5.65
45.64±0.63 45.69±1.04 44.75±0.57 38.44±0.56 22.46±1.62
HOPE REAP MAN EAN Freq
50% 50% 50% 50% 50%
91.92±0.36 89.65±0.36 91.41±0.36 51.26±1.99 2.02±0.71
43.99±1.10 40.04±0.76 37.89±0.61 13.33±0.46 1.32±0.21
15.15±0.86 15.76±0.86 17.58±1.71 16.36±1.48 16.97±3.43
6.49±0.72 5.93±0.03 7.24±0.14 0.65±0.13 0.72±0.39
42.58±1.36 38.10±0.48 48.02±1.17 29.42±0.14 30.89±2.47
36.15±0.19 38.96±0.78 36.20±1.64 33.33±0.47 38.09±1.45
61.90±0.52 60.73±0.47 63.23±0.48 59.87±0.31 61.38±0.31
52.96±0.73 51.29±0.13 28.33±3.35 9.78±0.20 8.39±0.47
43.89±0.73 42.56±0.48 41.24±1.18 26.75±0.65 19.97±1.18
Table S1: Full Tulu results across all calibration sets and pruning rates for Qwen3.5-35B-A3B. Scores are multiplied by 100 (mean and standard deviation across 3 independent calibration trials). The best method (per condition) for each column is bolded.
17
Calib
Method
Rate
Base
Evol-CA
SB-verif
GSM8K
MATH
IFEval
PopQA
MMLU
BBH
TQA
LCB
Mean
94.19±0.36
35.67±1.29
19.39±0.86
0.35±0.00
26.68±0.04
36.50±0.45
59.36±0.52
42.58±0.70
39.34±0.53
HOPE REAP MAN EAN Freq
10% 10% 10% 10% 10%
94.70±0.00 94.95±0.36 94.70±0.00 94.44±0.36 94.19±0.36
40.08±0.29 38.43±0.51 36.03±1.63 36.53±0.54 36.54±0.69
18.18±0.00 19.39±1.71 18.18±0.00 18.18±1.48 18.79±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
26.57±0.27 27.59±0.43 27.69±0.16 26.97±0.49 26.07±0.64
38.55±1.59 37.22±0.51 36.61±0.91 36.76±0.71 35.89±0.75
60.62±0.15 58.41±0.69 59.99±0.18 62.24±0.24 62.52±0.50
42.10±0.95 41.08±1.33 39.41±0.73 40.11±3.80 37.96±0.66
40.14±0.41 39.68±0.69 39.12±0.45 39.45±0.95 39.04±0.56
HOPE REAP MAN EAN Freq
20% 20% 20% 20% 20%
95.71±0.71 95.45±0.62 94.44±0.36 92.42±1.07 91.67±1.07
38.49±1.58 39.39±2.86 38.01±0.84 37.95±0.81 36.36±0.31
18.79±0.86 19.39±0.86 18.79±0.86 18.18±0.00 16.36±0.00
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
25.20±0.26 24.28±0.74 24.29±0.21 35.23±0.99 46.81±0.42
38.19±0.70 36.76±0.36 37.83±1.38 35.48±0.72 36.76±0.94
57.17±0.41 57.68±0.46 60.03±0.37 60.37±0.44 61.14±0.40
36.29±0.95 37.96±3.38 42.74±0.26 39.03±0.73 33.98±2.31
38.77±0.68 38.91±1.16 39.56±0.53 39.88±0.60 40.43±0.68
HOPE REAP MAN EAN Freq
25% 25% 25% 25% 25%
96.46±0.36 94.95±1.56 95.20±0.36 92.93±0.71 93.69±0.94
37.43±2.15 38.76±1.28 40.59±1.11 33.06±0.85 33.19±1.61
18.18±0.00 18.79±0.86 20.61±0.86 18.18±0.00 16.36±0.00
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
23.82±0.07 24.84±0.39 25.70±0.04 44.21±0.41 45.66±0.92
37.58±0.50 37.47±0.64 37.99±0.32 36.81±1.07 36.04±1.03
59.87±0.60 59.20±0.78 57.45±0.41 61.80±0.25 59.56±0.77
43.01±0.20 38.98±2.29 42.53±1.92 36.99±1.28 35.81±0.26
39.59±0.48 39.17±0.97 40.05±0.63 40.54±0.57 40.08±0.69
HOPE REAP MAN EAN Freq
30% 30% 30% 30% 30%
96.72±0.36 95.45±0.62 95.45±0.62 93.43±0.94 93.69±0.71
38.40±1.53 40.62±0.46 38.15±0.70 34.04±0.44 34.56±1.79
18.18±0.00 18.18±0.00 18.18±0.00 18.79±0.86 16.36±1.48
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
24.00±1.16 24.84±0.46 25.31±0.23 50.20±0.46 39.56±1.35
36.45±0.32 34.82±1.25 35.99±0.56 34.20±1.30 36.81±2.07
57.55±0.25 59.16±2.86 58.56±0.16 59.31±0.25 61.38±0.67
44.84±2.77 40.86±1.69 40.16±1.15 37.20±0.53 34.25±4.23
39.56±0.80 39.29±0.92 39.02±0.43 40.94±0.60 39.62±1.54
HOPE REAP MAN EAN Freq
40% 40% 40% 40% 40%
94.44±0.36 96.21±0.62 94.95±0.71 94.70±0.62 93.43±0.71
43.08±1.20 40.40±2.81 40.66±2.03 39.11±1.68 12.93±0.89
21.82±1.48 20.00±0.00 20.00±1.48 16.97±0.86 18.79±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
25.87±0.38 23.39±0.16 23.25±0.30 50.75±0.76 39.36±0.85
36.86±1.45 35.84±0.89 39.31±0.14 37.07±0.47 36.25±0.59
56.40±0.33 56.83±0.68 58.65±0.30 62.58±0.58 63.07±0.32
42.37±1.47 38.12±0.30 55.86±0.55 37.20±0.92 45.59±2.11
40.15±0.83 38.89±0.68 41.63±0.69 42.34±0.74 38.72±0.79
HOPE REAP MAN EAN Freq
50% 50% 50% 50% 50%
95.20±0.36 94.70±0.62 95.71±0.71 93.94±0.62 90.66±0.71
42.35±2.78 47.29±5.03 45.53±0.64 36.01±1.51 11.13±0.07
21.82±1.48 21.82±0.00 20.00±1.48 18.18±2.57 18.79±0.86
8.90±0.32 1.68±1.54 7.69±1.15 0.35±0.00 0.35±0.00
48.93±3.83 20.87±0.02 29.75±0.60 40.08±0.47 33.40±0.89
50.31±0.99 38.65±2.73 35.99±0.69 37.01±1.71 38.45±0.75
55.42±0.78 62.37±2.51 63.70±0.39 63.52±1.05 60.57±0.19
27.04±2.39 34.09±10.45 15.00±0.13 48.92±1.35 32.26±2.66
43.75±1.62 40.18±2.86 39.17±0.72 42.25±1.16 35.70±0.77
HOPE REAP MAN EAN Freq
10% 10% 10% 10% 10%
95.71±0.36 94.70±0.00 94.70±0.00 96.21±0.62 94.70±0.62
38.08±1.54 39.09±0.32 36.66±2.40 34.07±2.06 33.84±0.32
18.18±0.00 18.79±0.86 19.39±0.86 18.79±0.86 18.79±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
26.64±0.78 27.95±0.07 27.05±0.18 25.89±0.20 25.17±0.12
37.58±0.70 35.38±1.40 38.60±0.64 33.74±0.22 33.64±1.00
62.40±0.15 60.00±0.40 57.61±0.59 62.89±0.72 61.28±0.45
44.19±0.95 37.37±1.18 37.26±0.73 39.14±1.65 31.83±0.80
40.39±0.56 39.20±0.53 38.95±0.68 38.88±0.79 37.45±0.52
HOPE REAP MAN EAN Freq
20% 20% 20% 20% 20%
94.70±1.07 95.20±0.71 95.45±0.00 93.18±1.64 90.66±0.94
37.03±0.83 38.25±0.96 38.95±1.59 32.87±1.53 26.85±1.11
18.79±0.86 18.79±0.86 18.18±0.00 16.97±0.86 18.18±0.00
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
23.19±0.07 23.56±0.08 23.51±0.09 25.29±0.38 23.08±0.32
38.39±1.02 37.42±0.76 35.38±1.53 35.22±1.22 35.33±0.71
57.92±0.93 54.98±0.19 58.53±0.26 65.39±0.24 62.98±0.49
34.25±1.87 27.15±0.88 27.90±2.31 28.55±1.83 31.51±6.47
38.08±0.83 36.96±0.56 37.28±0.72 37.23±0.96 36.12±1.26
HOPE REAP MAN EAN Freq
25% 25% 25% 25% 25%
94.19±0.71 95.20±0.36 94.70±0.62 91.67±0.62 34.09±3.44
38.86±0.93 37.93±1.50 40.34±1.07 30.62±0.92 14.32±0.42
19.39±0.86 16.97±0.86 18.79±0.86 18.18±1.48 17.58±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
24.01±0.15 24.19±0.23 25.53±0.08 23.52±0.28 22.32±0.22
37.58±1.75 34.56±0.73 36.96±0.76 35.22±1.46 39.47±1.45
57.70±0.46 54.97±0.56 58.03±0.07 62.97±0.44 63.31±0.44
36.29±1.21 26.67±1.74 40.43±0.40 27.90±0.26 38.01±0.55
38.55±0.76 36.36±0.75 39.39±0.48 36.30±0.68 28.68±0.92
HOPE REAP MAN EAN Freq
30% 30% 30% 30% 30%
94.19±0.36 94.95±0.36 95.71±0.36 87.88±1.24 11.62±4.12
41.27±2.07 41.26±1.72 40.72±1.93 26.87±0.08 5.60±0.40
18.18±1.48 18.79±0.86 16.97±0.86 18.79±0.86 18.79±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
26.17±0.07 26.86±0.09 24.84±0.33 22.64±0.25 21.28±0.18
36.04±0.66 35.28±0.66 37.17±0.92 33.84±0.19 39.57±0.78
59.95±0.55 58.91±0.62 56.88±0.56 66.20±1.39 62.95±0.58
33.44±2.45 30.16±0.70 31.13±3.10 42.80±1.19 39.57±0.38
38.70±0.96 38.32±0.63 37.97±1.01 37.42±0.65 24.97±0.91
HOPE REAP MAN EAN Freq
40% 40% 40% 40% 40%
95.45±0.62 96.72±0.36 96.21±0.00 75.00±1.64 3.79±1.07
42.59±1.81 48.51±0.61 47.63±0.98 17.85±0.70 3.05±0.65
18.79±0.86 18.79±2.27 20.61±0.86 18.18±1.48 16.97±0.86
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
25.53±0.35 24.41±0.30 24.25±0.15 21.82±0.40 29.22±3.17
38.55±0.51 34.97±0.22 34.56±0.19 34.41±1.34 34.41±1.00
59.84±0.72 61.86±0.39 61.86±0.74 63.75±0.42 59.98±0.85
23.44±1.07 33.33±1.22 43.71±0.99 35.48±0.92 17.47±6.51
38.07±0.74 39.87±0.67 41.15±0.49 33.36±0.86 20.66±1.76
HOPE REAP MAN EAN Freq
50% 50% 50% 50% 50%
93.94±0.62 95.20±0.36 95.45±0.62 50.00±3.76 2.53±0.36
47.16±1.55 59.39±1.94 48.26±2.15 9.26±1.31 1.27±0.16
20.00±0.00 19.39±2.27 18.79±0.86 19.39±0.86 16.36±0.00
0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00 0.35±0.00
23.25±0.78 26.63±0.50 26.11±0.46 35.91±0.58 33.36±0.20
45.35±0.91 36.55±1.34 36.30±0.81 38.29±1.08 37.99±0.62
63.97±0.71 63.52±0.14 63.67±0.28 60.84±0.51 60.91±1.07
31.08±2.24 23.60±0.93 29.25±0.42 15.65±1.30 13.33±1.51
40.64±0.85 40.58±0.93 39.77±0.70 28.71±1.17 20.76±0.49
Table S2: Full Tulu results across all calibration sets and pruning rates for Qwen3.5-122B-A10B. Scores are multiplied by 100 (mean and standard deviation across 3 independent calibration trials). The best method (per condition) for each column is bolded.
18
Calib
Method
Rate
Base
Evol-CA
SB-verif
GSM8K
MATH
IFEval
PopQA
MMLU
BBH
TQA
LCB
Mean
90.91±1.24
33.65±1.49
29.09±1.48
30.20±0.52
32.51±0.21
40.70±1.18
58.86±0.24
58.71±0.26
46.83±0.83
HOPE REAP MAN EAN Freq
10% 10% 10% 10% 10%
88.89±0.71 89.90±0.94 90.66±1.29 90.91±1.24 85.35±0.94
32.89±0.93 32.25±1.53 32.70±0.78 33.73±0.88 28.34±0.67
27.88±3.09 27.88±1.71 29.09±0.00 27.88±0.86 29.70±0.86
20.56±0.22 16.16±0.23 18.50±0.47 19.83±0.51 21.37±0.40
30.47±0.03 32.21±0.29 32.04±0.25 31.79±0.21 30.39±0.36
41.97±0.81 39.57±0.66 41.10±0.66 40.13±0.56 40.24±1.34
58.51±0.58 58.94±0.09 56.22±0.94 57.70±0.42 58.95±0.43
56.56±0.33 57.31±0.68 59.73±1.41 56.61±1.05 51.02±0.97
44.72±0.84 44.28±0.77 45.01±0.72 44.82±0.72 43.17±0.75
HOPE REAP MAN EAN Freq
20% 20% 20% 20% 20%
88.89±0.71 89.90±0.36 90.15±0.00 91.16±0.94 66.92±1.29
33.96±0.49 34.38±0.64 32.37±0.45 33.45±1.78 15.01±0.22
24.85±2.27 23.64±1.48 26.67±0.86 27.27±0.00 27.88±2.27
8.67±0.27 9.20±0.38 9.02±0.12 13.24±0.69 15.39±0.97
30.83±0.38 32.43±0.18 32.42±0.23 31.40±0.05 40.58±0.86
40.08±0.62 39.42±0.66 39.47±1.61 36.96±0.98 43.15±2.51
53.76±0.55 54.38±0.18 52.86±0.27 56.30±0.53 61.02±0.44
56.18±2.27 54.73±0.27 57.53±1.32 52.74±1.37 37.69±0.59
42.15±0.94 42.26±0.52 42.56±0.61 42.82±0.79 38.46±1.14
HOPE REAP MAN EAN Freq
25% 25% 25% 25% 25%
90.66±1.56 89.65±1.29 90.40±0.36 91.16±0.94 61.87±3.78
36.60±0.97 36.05±0.92 33.19±1.51 37.04±2.28 13.21±0.79
24.24±1.71 23.64±1.48 24.24±0.86 25.45±0.00 29.70±0.86
8.53±0.17 8.57±0.37 8.60±0.46 10.60±0.59 13.95±0.15
32.22±0.55 31.50±0.41 31.05±0.20 31.87±0.07 40.45±0.87
41.31±0.81 40.75±1.07 39.93±0.92 37.42±0.70 48.31±1.64
55.33±0.61 51.01±0.79 52.42±0.46 55.14±0.53 59.86±0.57
55.59±0.95 55.65±1.19 54.46±0.38 45.43±1.71 30.05±0.62
43.06±0.92 42.10±0.94 41.79±0.64 41.77±0.85 37.18±1.16
HOPE REAP MAN EAN Freq
30% 30% 30% 30% 30%
89.90±0.71 92.68±0.71 89.65±1.29 91.92±0.36 43.94±2.70
37.89±1.78 39.02±0.90 38.11±0.97 35.92±1.50 9.47±0.48
26.67±1.71 26.67±3.09 27.27±0.00 29.09±1.48 28.48±2.27
8.46±0.23 9.18±0.15 9.18±0.11 9.30±0.54 11.82±0.14
31.26±0.15 30.77±0.06 31.01±0.07 30.77±0.47 35.47±1.51
38.80±1.23 39.01±0.47 38.75±0.26 38.09±0.69 47.70±0.76
51.44±0.29 49.97±1.02 53.20±0.68 55.57±0.99 57.63±1.04
49.78±1.10 48.06±0.47 51.18±0.20 44.41±1.51 11.72±0.99
41.77±0.90 41.92±0.86 42.29±0.45 41.88±0.94 30.78±1.23
HOPE REAP MAN EAN Freq
40% 40% 40% 40% 40%
90.40±1.29 90.91±1.24 91.41±0.36 85.35±1.89 16.16±2.34
41.18±1.20 48.17±0.53 45.68±1.30 43.81±0.60 2.53±0.39
24.24±3.74 27.88±3.09 24.85±2.27 23.03±3.09 24.24±2.27
7.62±0.62 8.36±0.49 9.02±0.54 9.09±0.22 9.62±0.34
30.96±0.61 34.64±0.16 34.49±0.39 46.00±0.12 32.23±0.36
39.37±1.45 39.47±1.13 39.52±0.71 38.80±0.95 37.99±1.01
53.02±0.33 55.48±0.91 57.58±0.59 57.89±0.74 55.92±1.00
37.04±2.02 31.61±2.98 29.19±0.95 40.65±0.80 4.73±0.40
40.48±1.41 42.07±1.31 41.47±0.89 43.08±1.05 22.93±1.01
HOPE REAP MAN EAN Freq
50% 50% 50% 50% 50%
90.15±1.07 90.66±1.29 86.87±1.29 79.29±2.17 5.05±1.79
54.30±0.87 66.81±0.48 58.93±1.07 35.91±0.74 1.04±0.32
21.82±0.00 20.61±0.86 18.79±1.71 13.94±0.86 23.64±1.48
7.31±0.49 8.57±0.18 8.27±0.32 8.85±0.26 13.22±0.33
35.72±0.67 41.39±0.51 37.49±0.00 38.80±0.28 29.66±0.32
43.56±0.70 46.32±1.19 48.36±0.14 44.43±0.95 36.91±1.61
58.10±1.01 59.39±0.40 60.03±0.43 57.82±0.70 55.03±2.14
35.11±0.08 23.76±0.30 15.54±0.53 31.77±0.92 0.27±0.20
43.26±0.61 44.69±0.65 41.79±0.69 38.85±0.86 20.60±1.02
HOPE REAP MAN EAN Freq
10% 10% 10% 10% 10%
89.65±0.36 90.66±0.94 90.91±1.07 87.63±0.36 82.07±0.94
32.17±1.03 32.79±1.82 32.49±1.34 31.68±0.49 17.62±0.44
28.48±0.86 27.27±0.00 27.27±1.48 27.27±0.00 29.09±1.48
16.14±0.18 15.02±0.28 18.85±0.49 19.83±0.11 25.79±0.46
30.60±0.09 31.03±0.06 31.63±0.41 30.78±0.35 30.51±0.36
40.75±1.02 41.16±0.38 39.31±0.51 39.88±0.63 37.93±0.81
57.65±1.03 55.66±0.36 55.86±0.18 56.52±0.22 58.08±0.45
55.54±1.67 54.14±1.30 55.91±1.53 48.06±1.27 23.39±1.34
43.87±0.78 43.47±0.64 44.03±0.88 42.71±0.43 38.06±0.79
HOPE REAP MAN EAN Freq
20% 20% 20% 20% 20%
87.88±1.07 87.37±1.99 89.65±0.94 85.61±1.24 58.33±5.97
32.77±0.96 32.62±1.50 33.91±0.80 27.46±1.54 6.89±0.48
22.42±1.71 26.06±0.86 25.45±1.48 26.06±0.86 28.48±2.27
10.72±0.25 9.41±0.17 11.31±0.17 10.58±0.62 15.74±0.36
30.28±0.20 30.27±0.42 30.32±0.47 30.27±0.36 33.12±1.84
39.47±0.51 38.70±0.81 39.57±2.30 38.34±1.09 38.85±1.08
54.80±0.49 51.27±0.63 54.35±0.75 52.62±0.45 54.60±1.01
56.61±2.12 54.89±0.30 54.09±1.24 38.06±5.07 10.86±0.33
41.87±0.91 41.33±0.84 42.33±1.02 38.63±1.40 30.86±1.67
HOPE REAP MAN EAN Freq
25% 25% 25% 25% 25%
88.64±2.23 89.90±0.94 88.38±1.79 84.60±0.94 41.67±8.92
37.61±1.05 34.03±0.72 33.89±1.29 25.15±0.36 4.36±0.91
28.48±0.86 26.06±2.27 29.70±1.71 28.48±3.09 27.27±1.48
7.31±0.26 8.99±0.37 8.41±0.51 8.46±0.65 14.30±0.20
32.51±0.61 30.15±0.80 30.00±0.12 29.82±0.28 30.31±1.64
38.85±0.62 38.85±0.07 39.31±0.40 38.45±0.94 36.91±0.59
49.22±1.11 53.40±0.64 57.71±0.85 55.84±1.32 54.84±0.99
53.12±1.00 55.27±0.40 55.22±1.33 33.01±1.24 6.29±1.72
41.97±0.97 42.08±0.78 42.83±1.00 37.98±1.10 26.99±2.06
HOPE REAP MAN EAN Freq
30% 30% 30% 30% 30%
88.89±1.99 90.91±1.86 87.12±1.64 80.56±1.29 32.07±1.29
38.03±1.32 36.55±0.52 34.61±0.39 20.52±0.65 2.66±0.32
28.48±3.09 28.48±2.27 24.85±0.86 29.70±2.27 28.48±1.71
8.20±0.95 7.90±0.20 7.15±0.26 9.16±0.03 12.43±0.77
33.96±0.18 32.89±0.51 29.52±0.35 28.99±0.83 31.98±0.56
40.44±1.09 37.73±0.38 39.72±0.95 39.88±1.63 36.04±0.82
54.62±0.47 54.52±0.50 54.61±0.57 53.66±1.24 56.95±1.57
50.00±0.86 50.65±0.92 48.76±0.99 32.20±1.54 6.18±0.27
42.83±1.25 42.45±0.89 40.79±0.75 36.83±1.19 25.85±0.91
HOPE REAP MAN EAN Freq
40% 40% 40% 40% 40%
84.60±0.36 85.10±1.99 86.62±1.29 74.24±3.76 7.58±0.62
42.93±1.42 44.66±1.07 37.19±0.60 13.51±1.27 1.70±0.12
30.30±0.86 27.27±1.48 23.03±2.27 30.30±2.27 21.82±1.48
7.45±0.29 7.90±0.29 7.64±0.30 9.11±0.15 10.84±0.32
37.15±0.31 34.20±0.64 29.20±0.40 30.86±0.62 31.02±1.64
44.84±0.94 44.89±0.59 39.62±0.14 42.94±0.45 36.81±1.07
57.07±0.84 56.82±0.98 58.74±0.73 49.59±1.54 55.15±0.60
46.24±1.80 41.02±2.21 46.34±0.61 24.73±0.73 1.88±0.76
43.82±0.85 42.73±1.16 41.05±0.79 34.41±1.35 20.85±0.83
HOPE REAP MAN EAN Freq
50% 50% 50% 50% 50%
77.53±1.79 79.29±0.94 80.56±3.52 51.52±5.50 2.02±0.94
47.80±2.13 56.36±3.07 47.71±1.74 10.77±1.24 0.83±0.22
19.39±0.86 18.79±3.74 20.00±1.48 28.48±3.09 18.79±3.09
8.85±0.07 8.71±0.43 7.62±0.34 8.08±0.86 8.60±0.26
36.58±0.85 43.56±0.43 31.71±0.41 30.18±0.77 28.92±1.87
49.18±0.29 50.31±1.81 45.65±1.49 43.05±1.57 36.91±1.26
57.19±0.67 54.97±0.98 59.47±0.51 56.78±1.28 54.59±0.78
30.65±1.55 35.32±1.65 36.18±3.22 15.38±0.80 0.05±0.08
40.90±1.02 43.41±1.63 41.11±1.59 30.53±1.89 18.84±1.06
Table S3: Full Tulu results across all calibration sets and pruning rates for GLM-4.5-Air. Scores are multiplied by 100 (mean and standard deviation across 3 independent calibration trials). The best method (per condition) for each column is bolded.
19
Method
Evol-CA
SB-verif
10%
25%
50%
10%
25%
50%
Base
41.82
41.82
41.82
41.82
41.82
41.82
HOPE REAP MAN EAN Freq
43.18±1.40 41.45±0.59 43.31±0.94 41.51±0.89 42.79±0.82
44.58±1.42 42.74±0.55 42.68±1.69 42.88±1.56 42.54±1.31
42.33±1.23 36.23±0.93 37.54±1.87 40.56±0.67 34.06±0.26
43.01±0.22 42.43±1.23 42.44±1.58 41.76±0.46 38.80±1.46
42.02±1.03 41.66±0.64 42.52±0.95 38.99±0.87 37.71±0.18
40.78±0.86 40.21±0.43 40.62±1.19 36.24±0.45 30.16±0.94
Table S4: Full SWE-Bench Pro results across all calibration sets and pruning rates for Qwen3.5-35B-A3B. Scores are multiplied by 100 (mean and standard deviation across 3 independent calibration trials). The best method (per condition) for each column is bolded.
Method
Evol-CA
SB-verif
10%
25%
50%
10%
25%
50%
Base
50.16
50.16
50.16
50.16
50.16
50.16
HOPE REAP MAN EAN Freq
52.44±0.68 49.28±0.46 52.17±0.34 52.10±0.44 51.47±0.59
52.29±0.54 47.17±0.81 47.80±0.55 51.55±0.27 52.00±0.77
43.84±1.62 39.61±0.27 41.81±0.18 48.32±1.62 45.71±1.10
53.22±0.66 51.94±1.26 51.11±0.55 50.66±0.83 50.29±1.25
51.39±0.47 51.55±0.18 51.81±0.30 49.62±0.42 50.87±1.68
50.63±0.21 46.78±0.68 44.92±1.32 51.80±0.27 45.62±0.92
Table S5: Full SWE-Bench Pro results across all calibration sets and pruning rates for Qwen3.5-122B-A10B. Scores are multiplied by 100 (mean and standard deviation across 3 independent calibration trials). The best method (per condition) for each column is bolded.
Method
Evol-CA
SB-verif
10%
25%
50%
10%
25%
50%
Base
28.70
28.70
28.70
28.70
28.70
28.70
HOPE REAP MAN EAN Freq
32.94±1.05 33.23±0.82 33.93±1.72 34.09±0.22 27.40±0.54
32.80±0.66 34.06±0.37 32.91±1.32 32.62±0.48 1.14±0.62
28.89±0.76 26.88±1.20 23.65±0.90 20.88±1.51 0.00±0.00
34.64±0.54 34.46±0.59 34.70±1.07 33.41±1.11 30.53±0.60
33.26±0.96 33.44±0.40 33.46±0.90 32.71±0.51 22.47±0.92
28.94±1.21 28.70±1.40 26.73±0.91 26.11±0.84 1.69±0.13
Table S6: Full SWE-Bench Pro results across all calibration sets and pruning rates for GLM-4.5-Air. Scores are multiplied by 100 (mean and standard deviation across 3 independent calibration trials). The best method (per condition) for each column is bolded.
Calibration Model
Method
Time/token (ms)
1k prompts (hr)
QP solve (s)
Ratio
Qwen-35B (L = 40, E = 256)
HOPE REAP
2.68 2.53
1.27 1.20
75.4 N/A
1.077×
Qwen-122B (L = 48, E = 256)
HOPE REAP
3.47 3.31
1.55 1.48
92.4 N/A
1.067×
Table S7: Computational cost of HOPE vs REAP. Calibration time is given in milliseconds per token, and in hours per 1000 prompts (given a random sample of Evol-CodeAlpaca). QP solve time is the time taken to solve the HOPE-specific quadratic program over all layers total (a one-time cost). Ratio is the total amount of time taken for HOPE to make its pruning decisions (including QP solve time) over the time taken for REAP.
20
a)
Expert co-selection PMI: Qwen3.5-35B-A3B (calib: Evol-CA) 0
Expert (clustered)
40
5
80 120
0
160 −5 200 −10
240 0
40
80
120
160
200
Expert (clustered)
10
0
40
4
80
2
120
0
160
−2
200
4 80
0
160
−2
200
−4
240
240
−6
240 0
40
80
120
160
200
240
0
40
Expert (clustered)
80
120
160
200
240
Expert (clustered)
Observed vs Expected co-selection: Qwen3.5-35B-A3B (calib: Evol-CA) −2
Independence
Independence
Independence
−2
−3 −4 −5 −6
Observed P(i, j)
Observed P(i, j)
−2
Observed P(i, j)
2
120
−4
Expert (clustered)
b)
6
40
Expert (clustered)
0
−3 −4 −5
−4
Expected P(i)P(j)
−2
−4 −5 −6
−6 −6
−3
−6
−4
−2
Expected P(i)P(j)
−6
−4
Expected P(i)P(j)
−2
Figure S2: Expert co-selection is highly non-random, which motivates interaction-aware pruning. a) Pointwise mutual information (PMI) between expert pairs. PMI is defined between two experts i, j (within a layer) (i,j) , where P (i, j) is the fraction of tokens which activate both experts, and P (i) is as P M I(i, j) = log2 PP(i)P (j) the fraction of tokens which activate only one expert i. We show PMI between experts at three different layers of Qwen3.5-35B-A3B. Positive values (red) indicate experts are co-selected more often than expected due to independent chance; negative values (blue) indicate experts are co-selected less often (mutual exclusion). b) Observed co-selection probability P (i, j) versus independence expectation P (i)P (j). Under completely independent routing, all points should lie on the diagonal. Any substantial deviation (which is especially pronounced at layer 0) demonstrates that experts operate in cooperative groups.
21
Figure S3: Spatial visualization of disagreement in expert pruning between HOPE and each baseline (Qwen3.5-122B-A10B). Each panel shows an L × E grid. Disagreement is typically distributed across all layers and experts rather than purely concentrated in certain regions.
22
HOPE F diagonal
8 6 4 2 0
0
2
4
6
REAP score Sj
8
Figure S4: The diagonal of the interaction matrix F recovers REAP’s expert rankings. Each point represents one expert (pooled across p layers of Qwen3.5-122B-A10B, calib: Evol-CA). The x-axis is the REAP score SkREAP and the y-axis is Fk,k . The Spearman correlation is near perfect (ρ = 0.988), confirming that REAP is a special case of HOPE.
Figure S5: HOPE’s continuous relaxation to solve the quadratic program is tight. a) Histogram of the relative obj
−obj
gap in the binary objective value versus the continuous objective value (measured as binobj cont ) per layer. bin The mean gap is only 0.04%, compared to an expected gap of 65.6% for randomly sampled feasible binary solutions (red line). b) Distribution of continuous solution values in p across all layers and experts. The overwhelming concentration of continuous solutions near 0 and 1 confirms that the relaxation produces effectively binary solutions, thereby rounding is nearly lossless.
Figure S6: Convergence of pruning decisions with calibration data. We measure the fraction of pruning decisions that agree with the final pruning set (computed from the entire calibration set) as a function of the number of calibration prompts processed. Both HOPE and REAP stabilize within 10k prompts, with HOPE requiring marginally more data to converge due to estimating pairwise (rather than per-expert) statistics. Convergence is shown for 50% pruning on Qwen3.5-35B-A3B with the Evol-CodeAlpaca calibration set.
23
C
C ROSS - LAYER HOPE (CHOPE)
In this section, we present an extension of HOPE which accounts for expert interactions across different layers, which enables non-uniform layer-wise pruning budgets (i.e. optimizing for a different number of pruned experts in each layer). We call this extension CHOPE (Cross-layer Higher-Order Pruning of Experts). C.1
S ETUP
Consider an L-layer MoE transformer, where each layer ℓ consists of attention followed by an MoE layer, each with residual connections: z ℓ = xℓ−1 + Attℓ (xℓ−1 ) xℓ = z ℓ + MoEℓ (z ℓ ) where xℓ−1 ∈ Rd is the residual stream entering layer ℓ, and z ℓ ∈ Rd is the post-attention representation fed to the MoE. Layer norms are absorbed into the respective Att, MoE functions. Note that in this derivation, we assume each layer consists of attention followed by an MoE block, but the attention can be equivalently replaced by other architectures (e.g. an SSM). A forward pass computes x0 , z 1 , x1 , . . . , xL in sequence, where x0 is the input token embedding and xL is fed to the language-model head. At each layer ℓ, the MoE contains E experts with output MoEℓ (z ℓ ) =
E P k=1
gkℓ (z ℓ ) fkℓ (z ℓ ), using the
same notation as in Appendix A. We use top-K routing with selected set T ℓ (z ℓ ) ⊆ {1, . . . , E}. At each layer ℓ, we prune a set of experts P ℓ ⊆ {1, . . . , E}. Upon pruning, for any token, the pruned experts P ℓ ∩ T ℓ (z ℓ ) are replaced by the next-highest-logit experts Rℓ (z ℓ ), with |Rℓ (z ℓ )| = ℓ |P ℓ ∩ T ℓ (z ℓ )|. The pruned MoE output uses new gate weights g ′ k (z ℓ ). Note that this is the same setup as in Appendix A, tracking layer indices. C.2
E RROR PROPAGATION
When we prune experts, the effects of pruning propagate through the network from early to later layers. After pruning experts from each layer, the forward pass computes perturbed quantities x̃0 , z̃ 1 , x̃1 , . . . , x̃L (with x̃0 = x0 ). Define the output error at layer ℓ as ϵℓ = x̃ℓ − xℓ . Here, we will quantify the error ϵℓ in each layer. Error through attention: At layer ℓ, the perturbed post-attention representation is z̃ ℓ = x̃ℓ−1 + Attℓ (x̃ℓ−1 ). The error in the output of this attention layer is: z̃ ℓ − z ℓ = (x̃ℓ−1 − xℓ−1 ) + (Attℓ (x̃ℓ−1 ) − Attℓ (xℓ−1 )) Since no pruning/modification happens at the attention block, its error z̃ ℓ − z ℓ arises purely from error in the input being propagated through. To quantify z̃ ℓ − z ℓ , we perform a first-order Taylor expansion of Attℓ around xℓ−1 : ℓ Attℓ (x̃ℓ−1 ) ≈ Attℓ (xℓ−1 ) + JA (x̃ℓ−1 − xℓ−1 ) ℓ
ℓ where JA = ∂ Att∂x(x) x=xℓ−1 ∈ Rd×d is the Jacobian of the attention output with respect to its input, evaluated at the unperturbed input.
Substituting, we have: ℓ z̃ ℓ − z ℓ ≈ (I + JA ) ϵℓ−1
Error through MoE: Similar to error propagation through attention, the error in the output of the MoE is: ℓ g (z̃ ℓ ) − MoEℓ (z ℓ )) x̃ℓ − xℓ = (z̃ ℓ − z ℓ ) + (MoE 24
ℓ
g is the MoE block after pruning. Unlike the attention block, the error in the output of where MoE the MoE block arises from both propagated error and the error accrued from pruning experts within the block. ℓ
g (z̃ ℓ ) − MoEℓ (z̃ ℓ ) be the error resulting from pruning at layer ℓ (the difference Let ∆ℓ (z̃ ℓ ) = MoE in MoE outputs before and after pruning, applied to the perturbed input). Substituting, this gives us: x̃ℓ − xℓ = (z̃ ℓ − z ℓ ) + (MoEℓ (z̃ ℓ ) − MoEℓ (z ℓ )) + ∆ℓ (z̃ ℓ ) We perform a first-order Taylor approximation of the unpruned MoE around z ℓ (like with attention), yielding: ℓ ϵℓ = x̃ℓ − xℓ ≈ (I + JM ) (z̃ ℓ − z ℓ ) + ∆ℓ (z̃ ℓ ) ℓ
(z) ℓ where JM = ∂ MoE is the MoE Jacobian (evaluated at the unperturbed input). ∂z z=z ℓ
Closed form of the error: Substituting in the attention error z̃ ℓ − z ℓ into the MoE error ϵℓ : ℓ ℓ ϵℓ = (I + JM )(I + JA ) ϵℓ−1 + ∆ℓ (z̃ ℓ )
with ϵ0 = 0. This recurrence has the following closed form: ϵL =
L X
ΦL,j ∆j (z̃ j )
j=1
where Φℓ,j is the error propagation operator, defined as follows: ℓ Y
Φℓ,j =
m m (I + JM )(I + JA ),
with
Φℓ,ℓ = I
m=j+1
Φℓ,j quantifies how error introduced at layer j is amplified or attenuated as it propagates to the output at layer ℓ. Note that ϵL is the overall error of the last layer of the network, which includes error arising from expert pruning in all MoE layers, as well as the error propagating from earlier to later layers. Decoupling error from z̃ j : Computing ∆j (z̃ j ) exactly requires knowing the perturbed input z̃ j , which itself depends on all previous pruning decisions. This makes optimizing this error over prunesets intractable. Instead, we approximate ∆j (z̃ j ) ≈ ∆j (z j ). That is, the pruning error at layer j is well-approximated by the error computed on the clean (unpruned) input. This decouples the per-layer pruning errors from the propagation operators: ϵL ≈
L X
ΦL,j ∆j (z j )
j=1
C.3
U PPER - BOUNDING THE SUBSTITUTION ERROR
Here, we will bound the substitution-error component of ϵL . The proof follows analogously to that in Appendix A. As with per-layer HOPE (Appendix A), we decompose ∆j (z j ) into a substitution error and a renormalization error: X X ′ ∆j (z j ) = gij (z j )fij (z j ) − g j (z j )f j (z j ) k
i∈Rj (z j )
k∈T j (z j )∩P j
+
X
′
(gkj (z j ) − gkj (z j ))fkj (z j )
k∈T j (z j )\P j
25
k
Substituting this decomposition into ϵL : L X X ′ ϵL = ΦL,j gij (z j )fij (z j ) − j=1
i∈Rj (z j )
X
gkj (z j )fkj (z j )
k∈T j (z j )∩P j
{z
|
}
substitution error L X
+
X
ΦL,j
j=1
′
(gkj (z j ) − gkj (z j ))fkj (z j )
k∈T j (z j )\P j
{z
|
}
renormalization error
The renormalization error remains small by the same argument as in Appendix A (with an additional factor of ΦL,j ). As before, we focus on minimizing the squared cross-layer substitution error: L X
X
ΦL,j
j=1
X
gkj (z j ) fkj (z j ) −
k∈T j (z j )∩P j
j
g ′ i (z j ) fij (z j )
2 2
i∈Rj (z j )
As before, we decompose this squared norm into three components (A + B + C): L X
A=
2
X
ΦL,j
j=1
j=1
C=
2
k∈T j (z j )∩P j
X L B = −2 ΦL,j L X
gkj (z j ) fkj (z j ) X
gkj (z j ) fkj (z j ),
L X j=1
k∈T j (z j )∩P j
X
ΦL,j
j g ′ i (z j ) fij (z j )
i∈Rj (z j )
2
j=1
j
X
ΦL,j
g ′ i (z j ) fij (z j ) 2
i∈Rj (z j )
√ √ √ √ By Cauchy–Schwarz, |B| ≤ 2 A C, so A + B + C ≤ ( A + C)2 . √ √ We now upper bound both A and C using a common quantity. Define the following: Z=
L X
X
gkj (z j ) ∥ΦL,j fkj (z j )∥2
2
j=1 k∈T j (z j )∩P j
Bounding layer): √ A=
√
A: By the triangle inequality (applied first over layers, then over experts within each
L X j=1
Bounding
√
X
ΦL,j
gkj (z j ) fkj (z j )
≤ 2
k∈T j (z j )∩P j
L X
X
gkj (z j ) ∥ΦL,j fkj (z j )∥2 =
√ Z
j=1 k∈T j (z j )∩P j
C: By the triangle inequality:
√ C=
L X
ΦL,j
j=1
X
j
g ′ i (z j ) fij (z j )
≤ 2
i∈Rj (z j )
L X
j
X
g ′ i (z j ) ∥ΦL,j fij (z j )∥2
j=1 i∈Rj (z j )
j
By Lemma 3, g ′ i (z j ) ≤ gkj (z j ) for all pruned–replacement pairs within each layer. Let rkj denote the expert replacing pruned expert k in layer j. Then: L X
X
j=1 i∈Rj (z j )
j
g ′ i (z j ) ∥ΦL,j fij (z j )∥2 ≤
L X
X
j=1 k∈T j (z j )∩P j
26
gkj (z j ) ∥ΦL,j frjj (z j )∥2 k
Extracting the worst-case ratio (which removes dependence on the replacement experts’ identities): L X
L
X
gkj (z j ) ∥ΦL,j frjj (z j )∥2 ≤
X
minj,k ∥ΦL,j fkj (z j )∥2 j=1 k∈T j (zj )∩P j √ = ρΦ Z
k
j=1 k∈T j (z j )∩P j
maxj,k ∥ΦL,j fkj (z j )∥2 X
gkj (z j ) ∥ΦL,j fkj (z j )∥2
Together: L X
ΦL,j
j=1
X
gkj (z j ) fkj (z j ) −
k∈T j (z j )∩P j
2
j g ′ i (z j ) fij (z j )
X i∈Rj (z j )
√ √ ≤ ( A + C)2
2
√ √ ≤ ( Z + ρΦ Z)2 = (1 + ρΦ )2 · Z
This gives us the following Theorem, analogous to Theorem 1: Theorem 4 (Cross-layer upper bound). For any collection of prune-sets {P 1 , . . . , P L }, the squared cross-layer substitution error is bounded by: (1 + ρΦ )2 · Z
where
ρΦ =
maxj,k ∥ΦL,j fkj (z j )∥2 minj,k ∥ΦL,j fkj (z j )∥2
Note that ρΦ generalizes the per-layer ρ from Theorem 1: it accounts for how ΦL,j differentially amplifies expert contributions from different layers. Because the ratio is independent of the prunesets {P 1 , . . . , P L }, we can simply minimize Z. C.4
M INIMIZING Z AS A QUADRATIC PROGRAM
Our goal is to minimize E [Z] over all prune-sets {P 1 , . . . , P L } jointly. We introduce binary x∼D
decision variables pjk ∈ {0, 1}, where pjk = 1 iff expert k in layer j is pruned. Substituting: X 2 E L X j j j j j j j pk · 1[k ∈ T (z )] · gk (z ) ∥ΦL,j fk (z )∥2 Z= j=1 k=1
=
L X E X E L X X
pjk pℓm · 1[k ∈ T j (z j )] · 1[m ∈ T ℓ (z ℓ )]
j=1 ℓ=1 k=1 m=1 ℓ ℓ · gkj (z j ) gm (z ℓ ) ∥ΦL,j fkj (z j )∥2 ∥ΦL,ℓ fm (z ℓ )∥2
Taking the empirical expectation over the calibration set (using conditional normalization as in perlayer HOPE): L X L X E X E X pjk pℓm · G[(j, k), (ℓ, m)] = p⊤ G p E[Z] = x
j=1 ℓ=1 k=1 m=1
where p ∈ {0, 1}LE is the concatenated binary vector over all layers, and G ∈ RLE×LE is defined as: X j 1 ℓ ℓ gk (z j ) gm (z ℓ ) ∥ΦL,j fkj (z j )∥2 ∥ΦL,ℓ fm (z ℓ )∥2 G[(j, k), (ℓ, m)] = j,ℓ |Xk,m | j,ℓ x∈Xk,m
j,ℓ with Xk,m = {x : k ∈ T j (z j ) and m ∈ T ℓ (z ℓ )} denoting the set of tokens which activate expert k in layer j and expert m in layer ℓ.
27
Theorem 5 (Cross-layer QP). Given a total pruning budget B (the total number of experts to remove across all layers), the collection of prune-sets {P 1 , . . . , P L } which minimizes E [Z] is the solution x∈D
to: p∗ = arg min p⊤ G p
s.t.
p ∈ {0, 1}LE ,
p
C.5
LE X
pk = B
k=1
R ELATIONSHIP TO PER - LAYER HOPE
The cross-layer matrix G encodes both within-layer and across-layer expert interactions: • E × E blocks on the diagonal of G (entries where j = ℓ) correspond to the per-layer F -matrices of HOPE, scaled by ∥ΦL,j ∥ • The off-diagonal blocks (entries where j ̸= ℓ) encode cross-layer interactions: the joint penalty of pruning expert k in layer j and expert m in layer ℓ simultaneously CHOPE strictly generalizes per-layer HOPE. Setting ΦL,j = I for all j and zeroing the off-diagonal blocks of G (i.e. G[(j, k), (ℓ, m)] = 0 for j ̸= ℓ) decouples layers entirely, recovering L independent per-layer HOPE problems (Theorem 2). Further zeroing the within-layer off-diagonal entries effectively recovers REAP. Unlike per-layer (which enforces a fixed budget per layer), CHOPE optimizes a single global P HOPE budget B = |P ℓ | across all layers. The optimization allocates pruning non-uniformly, potentially ℓ
pruning more aggressively in layers with weaker cooperative structure or lower propagation impact. C.6
P RACTICAL CONSIDERATIONS
ℓ ℓ at every layer is proand JM Jacobian approximation: Computing the full d × d Jacobians JA hibitively expensive for large models. A practical simplification is to set ΦL,j = I for all j, corresponding to the assumption that errors propagate through the residual stream without transformation (i.e. Jacobians are approximately zero). For residual-stream architectures where the skip connection dominates, this is a reasonable approximation. Under this simplification, G retains cross-layer interaction terms but does not differentially weight layers based on error propagation.
Per-layer minimum constraints: In practice, optimizing a global budget without constraints can lead to instability (e.g. over-pruning early layers). We propose adding per-layer minimum constraints to ensure at least M experts (the routing width) survive in every layer: p∗ = arg min p⊤ G p
s.t.
p ∈ {0, 1}LE ,
p
LE X k=1
pk = B,
E X
(1 − pjk ) ≥ M ∀ j
k=1
Scalability: The matrix G has dimension LE × LE. While larger than the per-layer E × E matrix, this remains tractable for modern QP solvers. C.7
P RELIMINARY RESULTS
Here, we present some preliminary results comparing the performance of CHOPE to layer-wise HOPE. We implemented CHOPE and computed non-uniform prune-sets for Qwen3.5-35B-A3B and Qwen3.5-122B-A10B. For simplicity, we assume that Jacobians are I (as described above), and do not limit the number of experts pruned per layer. CHOPE’s non-uniform allocation is consistent across trials, and in many conditions, CHOPE performed reasonably well, although not as well as per-layer HOPE (Tables S8–S9). CHOPE-pruned models also occasionally experienced performance collapse. CHOPE tended to prune a large number of experts in the earliest layers of the network (Figure S7), suggesting that the cross-layer objective’s pruning allocation can introduce instability. Although 28
Method
Rate
Base
Tulu Mean
SWE-Bench Pro
48.87
41.82
HOPE CHOPE
10% 10%
48.49±0.53 46.17±0.72
43.18±1.40 40.87±1.41
HOPE CHOPE
25% 25%
48.51±0.61 42.36±1.08
44.58±1.42 27.29±17.51
HOPE CHOPE
50% 50%
42.93±1.44 48.12±2.44
42.33±1.23 22.95±1.19
Table S8: HOPE vs cross-layer HOPE (CHOPE) for Qwen3.5-35B-A3B (calib: Evol-CA). Scores ×100 (mean and standard deviation across 3 independent calibration trials). The better method for each column is bolded. Method
Rate
Base
Tulu Mean
SWE-Bench Pro
39.34
50.16
HOPE CHOPE
10% 10%
40.14±0.41 39.45±0.76
52.44±0.68 46.04±0.32
HOPE CHOPE
25% 25%
39.59±0.48 38.72±1.06
52.29±0.54 47.10±0.46
HOPE CHOPE
50% 50%
43.75±1.62 40.30±0.88
43.84±1.62 42.40±0.11
Table S9: HOPE vs cross-layer HOPE (CHOPE) for Qwen3.5-122B-A10B (calib: Evol-CA). Scores ×100 (mean and standard deviation across 3 independent calibration trials). The better method for each column is bolded. promising, limiting the number of experts pruned (particularly in early layers) could help CHOPE perform better and avoid performance collapse. We leave further exploration of CHOPE for future work.
29
Figure S7: Non-uniform expert pruning allocation by CHOPE (cross-layer HOPE). Each line indicates the number of experts pruned per layer (3 independent trials for each pruning rate); dashed horizontal lines indicate the uniform budget (used by per-layer HOPE).
30
D
M ETHODS
Code for HOPE is available at https://github.com/awslabs/hybrid-model-factory/tree/main/ examples/research/HOPE/ D.1
C ALIBRATION
Expert usage statistics for HOPE were collected via a single forward pass over the calibration set. We hooked into each MoE layer’s expert module and—for each token—recorded: 1) which experts were in the top-K selected set, 2) their softmax-normalized gate weights, and 3) their output vectors. From these, we P accumulated per-layer statistics in an E × E matrix: the total gate-weighted norm products gi (x)gj (x)∥fi (x)∥∥fj (x)∥ for all co-selected expert pairs i, j, along with x:x∈Xi,j
co-selection counts |X⟩,| |. We used two calibration sets: • Evol-CodeAlpaca-v1 (Luo et al., 2023): A coding-focused instruction dataset. We used all 111k prompts on the unpruned model. • SWE-bench Verified trajectories (Jimenez et al., 2023): Agentic multi-turn coding trajectories collected from the unpruned model solving SWE-Bench Verified instances (500 instances). Full trajectories (including tool calls and observations) were reconstructed and tokenized. For each calibration set (and each model), we ran 3 independent trials using the same prompts but different random seeds for model generation, yielding different routing patterns and thus different F matrices per trial. For SWE-bench Verified, we truncated any traces longer than 100k tokens (rare). D.2
Q UADRATIC - PROGRAM SOLVER
Given the E × EPinteraction matrix F for a layer, we solved the binary QP minp p⊤ F p subject to p ∈ {0, 1}E and k pk = |P | via continuous relaxation. We relaxed p ∈ {0, 1}E to p ∈ [0, 1]E and solved with SciPy’s SLSQP optimizer, using bounds [0, 1] per variable and a hard equality constraint | −12 on the sum. The initial point was set to pk = |P and a E (uniform). We used a tolerance of 10 maximum of 1000 iterations. The continuous solution was then rounded to binary by selecting the |P | entries with the highest values. For first-order baselines (REAP, MAN, EAN, Frequency), we computed their respective scalar scores using the same calibration data and collected the statistics as described in their respective publications, and sorted to select the |P | experts with the lowest scores. D.3
M ODEL PRUNING
Pruned model checkpoints were created by physically removing expert parameters. For each pruned expert, we deleted its MLP weights (gate up proj and down proj) and sliced out the corresponding rows from the router’s weight matrix. This reduced the model’s parameter count and file size proportionally. The pruned model was saved as a standard HuggingFace checkpoint and loaded directly for inference. D.4
T ULU 3 EVALUATION
Pruned models were evaluated on the Tulu3-Dev benchmark suite using the OLMo Evaluation Suite (olmes) (Lambert et al., 2024). We evaluated on 8 tasks: GSM8K, MATH, IFEval, MMLU, BBH, TruthfulQA, PopQA, and LiveCodeBench. We dropped Codex from the suite due to insufficient variation between models (LiveCodeBench is a more realistic set of coding tasks). Models were served via vLLM with a maximum sequence length of 8192 tokens and a generation budget of 2048 tokens per instance for computational efficiency. These are the same parameters we used for Evol-CodeAlpaca calibration on the unpruned models. 31
D.5
SWE-B ENCH P RO EVALUATION
Agentic coding evaluation was conducted on SWE-Bench Pro (Deng et al., 2025), a benchmark of 731 enterprise-level software engineering instances. We used the Harbor evaluation framework with the OpenHands SDK agent (openhands-sdk-bench). The pruned model was served via vLLM and accessed by the agent as a hosted endpoint. We used the following evaluation parameters: n attempts=1, max input tokens=253952, max output tokens=8192, temperature 0.6, top-p 0.95. The metric reported is the mean reward (resolve rate) across all 731 instances. These are the same parameters we used for SWE-Bench verified calibration on the unpruned models. D.6
S UPERVISED FINE - TUNING
Post-pruning SFT was conducted on Qwen3.5-122B-A10B pruned checkpoints (HOPE and REAP, at 50% pruning rate, Evol-CodeAlpaca calibration). We fine-tuned on LiveCodeBench traces collected from the unpruned model, tokenized with sequence parallelism (8-way). Training used DeepSpeed ZeRO Stage 3 with full parameter fine-tuning (not LoRA), BF16 mixed precision, and Flash Attention 2. We used the following learning hyperparameters: learning rate 10−6 , constant schedule with 20-step warmup, max gradient norm 1.0, global batch size ∼2M tokens/step (per-device batch size 1, gradient accumulation every 2 steps, sequence parallel size 8), trained for 500 steps (∼1B tokens total). Pre-SFT and post-SFT evaluations used matched calibration trials for fair comparison.
32