Data Scarcity and Model Sparsity
DATA S CARCITY AND M ODEL S PARSITY: M IXTURES - OF -E XPERTS OVERFIT M ORE TO R EPEATED DATA Atindra Jha∗1 , Margaret Li∗2 , Jure Leskovec1 , Percy Liang1 , and Luke Zettlemoyer2 1
2
Stanford University Paul G. Allen School of Computer Science, University of Washington
arXiv:2609.11917v1 [cs.LG] 10 Sep 2026
A BSTRACT As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert count and granularity. We consistently find, for models ranging from 80M to 1B active (8.5B total) parameters, that MoEs degrade more rapidly under data repetition. This effect increases with sparsity, dictated by total rather than active parameters. While 80M dense models can repeat data over 8× with minimal degradation, MoEs instead begin to suffer at 4 repetitions, and deteriorate rapidly, ceding their performance benefits in all-unique data settings to dramatically underperform dense models after 32 repetitions. We experiment with existing regularization methods as a potential remedy. We find that some methods, such as dropout, can mitigate overfitting. In particular, with strong masking-based regularization, MoEs are able to outperform dense models even when data is repeated more than 64 times. However, no method fully matches the performance of all-unique training data. Finally, we analyze internal mechanisms correlated with MoE overfitting in high repetition regimes, and find that MoE routing universally stabilizes early in training, and that expert specialization correlates with overfitting to repeated data. In sum, our work addresses the adverse interactions between sparsity and data repetition: we present evidence for the core mechanisms of overfitting and its potential remediation, and suggest promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.
1
I NTRODUCTION
As language model training begins to exhaust even massive-scale web crawls, the amount of unique training data has become a constraining factor for language model training, in addition to compute. It is now common practice to repeat some or all training data, despite the known tendency of language models to overfit to data under high repetition rates (Muennighoff et al., 2023). Simultaneously, the compute costs of large-scale LLM training have driven the adoption of relatively compute-efficient Mixture-of-Experts models (Shazeer et al., 2017; Fedus et al., 2022; Muennighoff et al., 2025). These models achieve compute efficiency via sparsity, the ratio of total to active parameters. When training only with unique data, increased sparsity is known to consistently improve training efficiency, albeit at higher communication costs. However, the interaction between sparsity and data repetition remains relatively unexplored. Sparsity decouples total parameters from active parameters, and data repetition decouples total data tokens from unique data tokens. Thus, both sparsity and data repetition introduce axes of variation to modern scaling laws, which prescribe simple token-to-parameter ratios without disambiguating between unique or total training tokens over active or total parameters. MoE performance also depends on architectural details such as expert size and count. Further, not all data tokens are equal, with significant research devoted to filtering, deduplication, and data mix domain makeup (Li et al., 2024). Each axis of variation is entangled with all others, but prior work primarily investigates these axes in isolation: the data-constrained scaling laws of Muennighoff et al. (2023) are fit to dense models on a single corpus (C4), and Xue et al. (2023) consider a single MoE configuration to posit that multi-epoch degradation increases with total parameters. To investigate these interactions more fully, we conduct an in-depth grid sweep over axes of variation. Across three active-parameter scales (80M, 200M, 1B), we compute-match by training on a fixed total data budget for each activeparameter scale. We vary the unique tokens (equivalently, the repetition rate R), comparing dense transformers to MoE architectures with varied expert count and granularity. We consider a variety of domains (web crawl, code, scientific, and encyclopedic text), both as single-domain corpora and as components of data mixes with different per-domain repetition rates. Our results indicate that MoEs overfit more to repeated data in comparison to dense models; 80M MoEs clearly degrade at 4× data repetition (compared to 8× for dense models), and cede their performance advantage ∗
equal contribution; correspondence to [email protected], [email protected]
1
Data Scarcity and Model Sparsity
at 32× data repetition. Further, overfitting becomes more catastrophic with higher MoE sparsity, and its patterns depend on total rather than active parameters. Surprisingly, this phenomenon follows similar patterns across datatoken-per-parameter ratios and across our varied data domains, only slightly decreased by quality filtering. However, when repeated data is mixed into a non-repeated dataset, the unique tokens may have a regularizing effect. We also find that some existing regularization methods (dropout, output masking) reduce the impact of high data repetition rates, especially in MoEs. At sufficiently high dropout probability, MoEs can outperform dense models even at 64× data repetition. Finally, we perform mechanistic analyses and show that MoE routers stabilize their decisions early in training, which suggests that expert parameters update on a small and stationary subset of the total tokens, and we measure the resulting impact of data repetition on expert specialization. Overall, our contributions are as follows: • We show that the benefit of sparsity is conditional on the unique data budget: Across data domains and mixes, MoEs outperform dense models on unique data, but underperform at high data repetition rates. Data quality filtering has only minor impact on repetition effects: some models overfit more to unfiltered data (§3.1-3.3). • We study domain-specific repetition rates in data mixes. Performance degradation is confined to repeated domains, and unique tokens from another domain may mitigate overfitting of repeated domains (§3.4). • We apply regularization techniques and demonstrate that some techniques are ineffective, but dropout and output masking can mitigate overfitting from data repetition (§4). • We analyse MoE expert activations and outputs. Our evidence shows that MoE router decisions stabilize early, which exacerbates overfitting as expert parameters over-specialize, updating on a small and near-stationary subset of tokens (§5).
2
BACKGROUND
2.1
M IXTURE OF E XPERTS L ANGUAGE M ODELS
In Mixture-of-Experts Transformer LMs, the Feed-Forward (FFN) of each layer is replaced by n parallel FFN experts E1 , . . . , En and a router that selects a subset of the experts to apply to each token. For each hidden token representation h, the router produces a score si (h) for each expert i and returns a weighted sum of the top-k experts: X y = gi (h) Ei (h), i∈TopK(s(h))
where gi are the router’s softmax scores of the selected experts. Because not all parameters are active for each token, MoEs decouple total from active parameters. Recent MoEs employ fine-grained experts (Dai et al., 2024), where the granularity is defined as the ratio of expert-FFN to dense-FFN dimensions. For example, if an MoE has expert granularity g = 12 , then the experts have 1 1 · dense FFN dimension = · 4 · hidden dimension = 2 · hidden dimension. 2 2 MoE architectures may vary the expert granularity g, the total count of experts n, and the active count k. These design choice axes complicate the study of MoEs, as active parameter count and FLOPS-per-token cost change with each configuration. For fair comparison with dense models, it is common to match active parameters by setting g = k1 . expert dimension =
2.2
DATA R EPETITION
In a data-constrained scenario, it is common to repeat some or all of the training data. A simple approach might iterate over the entire training data corpus of U unique tokens over multiple passes, or epochs, until the desired total token budget T is reached. This requires a repetition rate of R = T /U . In other cases, training data may be defined as a mix of various data domains combined in fixed percentages. If only a subset of the domains are data-constrained, it is common to set domain-specific repetition rates Ri for each domain i (Soldaini et al., 2024a; Muennighoff et al., 2025). 2.3
R EGULARIZATION M ETHODS
Overfitting describes memorization of training data at the cost of generalization to unseen data. Known remedies trade off training fit to recover generalization: Dropout (Srivastava et al., 2014) randomly drops units during training, which prevents them from co-adapting to form complex memorization patterns. Weight decay, in the decoupled form used by AdamW (Loshchilov & Hutter, 2019), shrinks parameters towards zero independently of the gradient. Gradient norm clipping (Pascanu et al., 2013) rescales gradients whose norm exceeds a threshold, stabilizing the effect of any single datapoint. Some regularization methods specifically target components of MoEs: Router jitter injects multiplicative noise into the routing computation, and expert dropout applies a separate, typically larger dropout rate inside experts (Fedus et al., 2022; Zoph et al., 2022a). We evaluate regularization methods under data repetition in §4. 2
Data Scarcity and Model Sparsity
3
E XPERIMENTS
Models. We train compute-matched densely- and sparsely-activated (MoE) Transformer LMs in the style of Muennighoff et al. (2025) . Specifically, we train models with 80 million, 200 million, and 1 billion active parameters, denoted as 80M, 200M, and 1B, respectively. We vary our MoE model configurations to study the effect of total expert count and expert granularity. Our MoE models have n ∈ {8, 16, 32, 64, 128} total experts with granularity 1 1 g ∈ { 21 , 14 , 81 , 16 , 32 }. To match active parameters, we set the top-k activation to k = g1 . This results in sparsity s ∈ {2, 4, 8, 16, 32}. We denote our models MoE (n x g), e.g., MoE (64 x 1/4) represents 4 active experts out of 64 total with granularity 1/4. Additional model architecture details are in Appendix A.1. Data Repetition. Following common practice from Hoffmann et al. (2022), our models are trained with total training data tokens T ≈ 20 · Na , where Na denotes the number of active parameters. We only vary the tokens-per-activeparameter T /Na ratio in §3.1 to study its interaction with sparsity and data repetition. Under a fixed total data T , we vary the number of unique tokens U . Each model is trained by iterating through its allotted U unique tokens R = T /U times. We consider subsets of R ∈ {1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024}. For any unique-token budget U , we construct the training set DU by taking the first U tokens from a fixed random permutation of the full dataset D. We use the same permutation across all experiments, ensuring that training sets are nested: for any U1 ≤ U2 , DU1 ⊆ DU2 Additional data repetition details are in Appendix A.5. Training Data. We train on constituent domains from the OLMoE (Muennighoff et al., 2025) data mix, both individually and in custom mixes: web crawl data (DCLM; Li et al. (2024)), code (Starcoder; Li et al. (2023)), scientific text (peS2o; Soldaini & Lo (2023)), and encyclopedic text (Wikipedia; Soldaini et al. (2024a)). See Appendix A.3. Evaluation. We report Cross-Entropy Loss (CE Loss) on various held-out data, including web crawl, code, scientific text, and encyclopedic text. We also measure CE Loss and accuracy on downstream tasks. Results on downstream tasks are in Appendix B.10. We show random seed sensitivity in Appendix B.1. Additional details in Appendix A.4. 3.1
DATA R EPETITION E FFECTS ACROSS M O E A RCHITECTURES
We train a variety of dense and MoE Transformer LMs on the OLMoE data mix to investigate the effects of data repetition. Specifically, we fix the total train token budget T = 20 · Na , but vary the repetition rate R, so that unique tokens U decreases as R increases to satisfy R · U = T for fixed T . MoEs degrade more than dense LMs under data repetition. Figure 1 shows the effect of data repetition on validation loss. Across model scales, dense Transformer performance degrades at high data repetition rates, with validation loss rising slightly at R = 8 and sharply at R > 64. MoEs respond more dramatically to data repetition, with a noticeable performance impact at R = 4 and a sharper increase at higher R. Though MoEs outperform dense models at R ≤ 16, the catastrophic effect of data repetition leads to a reversal at R = 32, where dense models outperform MoEs. Other held-out LM tasks (Appendix B.9) and downstream tasks (Appendix B.10) exhibit similar trends. At extremely high repetition rates, validation loss lowers again. Results in Figure 1 demonstrate a least optimal data repetition rate; validation loss rises with R until this point, then falls. We also consider much higher R ∈ {215 , 220 }, and find that validation loss rises yet again (Appendix Figure 11). Training loss falls to zero at high repetition rate. As shown in Figure 2, training loss decreases with increased repetition rate, falling under 1E-2 at R = 512 for 80M dense, and R = 160 for 200M dense models. Training loss decrease is a reflection of validation loss increase, which agrees with other indications of overfitting via training data memorization. Additional model scales in Appendix Figure 12. Sparsity increases overfitting; data repetition effects depend on total parameters. We vary the number of total experts and the expert granularity in Figure 4. In both cases, increased total parameters results in a sharper response to data repetition. As the active parameters remain fixed throughout these experiments, we hypothesize that this phenomenon is related to the unique tokens to total parameters ratio. Thus, we compare the 200M dense models to the 80M MoE (32 x 1/4) and MoE (64 x 1/4) models, which have 158M and 244M total parameters, respectively. In Figure 1, the data repetition effects of the 200M model appear to lie between those of the MoE (32 x 1/4) and MoE (64 x 1/4) models at 80M active parameters. Data repetition effects do not depend on total data budget. We train 80M active parameter models with 4× more data, resulting in a total-tokens-to-active-parameter ratio T /Na = 80. The resulting trends (Figure 5) are very similar to those with T /Na = 20 (Figure 1). Thus, the data repetition overfitting response may change only slowly with the total token budget, but instead depend most directly on total parameters. 3
Data Scarcity and Model Sparsity
4
16
64
Repetition Rate
1B
Common Crawl CE (↓)
10 9 8 7 6 5 1
200M
20
Common Crawl CE (↓)
Common Crawl CE (↓)
80M
10 9 8 7 6 5 4
256 1024
1
4
16
64
Repetition Rate
256 1024
10 9 8 7 6 5 4
Model type
Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
3
1
4
16
64
Repetition Rate
256
Figure 1: Across active parameter scales, data repetition rates over 8× result in increasingly severe overfitting. Sparser models overfit more (§3.1). At 80M, 200M, and 1B active parameters, we fix the total data budget T = 20 · Na , and vary the data repetition rate R via different sized unique token sets. As R increases, models increasingly overfit. Sparsity exacerbates overfitting behavior. Larger and sparser models overfit more at lower R. Dense (1 x 1)
10
Train CE (↓)
Common Crawl CE (↓)
64x 128x
1
MoE (32 x 1/4)
10
64x
1
256x
0.01
512x
0.01
0.001
1024x
0.001
0
500
1000
Train Step
1500
Dense (1 x 1)
10 9 8 7 6 5
64x 32x 16x
0
500
1000
Train Step
1
256x 512x 1024x
0
500
1000
Train Step
1500
0.001
1500
64x
32x 16x
0
500
1000
Train Step
256x 512x 1024x
0
500
1000
Train Step
1500
1500
128x 256x 512x 1024x 64x
10 9 8 7 6 5
32x 16x
0
500
1000
Train Step
1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x
Repetitions
MoE (64 x 1/4) 256x 128x 512x 1024x
10 9 8 7 6 5
128x
0.01
MoE (32 x 1/4) 1024x
32x 64x
0.1
512x 256x
128x
Repetitions
128x
0.1
0.1
MoE (64 x 1/4)
10
1500
1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x
Figure 2: At higher repetition rates, training loss falls to 0 as models overfit to the repeated data (§3.1). Early in training, train (above) and validation (below) loss fall together, but, at high repetition rate R, train loss falls rapidly to 0, indicating memorization of training data, while validation loss rises. Additional model scales in Appendix Figure 12.
3.2
DATA R EPETITION E FFECTS ACROSS DATA D OMAINS
We repeat the experimental setup in §3.1 with individual data domains from the OLMoE data mix: DCLM (web crawl), peS2o (academic), StarCoder (code), and Wikipedia (encyclopedic) text. Our goal is to understand variations in response to repetition across varied data domains. We focus on R ∈ {1, 2, 4, 8, 16, 32}. Patterns are largely consistent across single-domain experiments. Results in Figure 3 indicate that overfitting patterns from data repetition are robust across domains, despite the diversity of data. Code, encyclopedic, web crawl, and academic text are semantically very distinct, yet all four exhibit similar patterns, for example, that MoE models begin to underperform dense models at R ∈ [16, 32]. 3.3
DATA R EPETITION E FFECTS U NDER DATA F ILTERS
We also measure the impact of quality filtering on data repetition. It is common practice to preprocess raw web crawls to deduplicate, extract text from HTML, and remove data deemed low quality by pre-defined filters, often significantly reducing the available tokens. DCLM- BASELINE retains 2.4% of the raw 280T-token DCLM- POOL (Li et al., 2024). Under data constraints, filtering may remove lower quality tokens, but also further reduces the quantity of unique tokens. We study this tradeoff between token quality and repetition rate by considering 5 data settings, which mix 4
Data Scarcity and Model Sparsity
DCLM
StarCoder
Stack CE (↓)
Common Crawl CE (↓)
6
5
3.8 3.6 3.4 3.2 3 2.8 2.6
1
2
4
8
Repetition Rate
16
32
1
2
peS2o
8
16
32
Model type
Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
Wikipedia
5
Wikipedia CE (↓)
4.2
peS2o CE (↓)
4
Repetition Rate
4 3.8 3.6
4
3.4 1
2
4
8
Repetition Rate
16
32
1
2
4
8
Repetition Rate
16
32
Figure 3: Across all train domains, sparse MoE models overfit more, ceding their benefit over dense models as repetition rates rise (§3.2). We train dense and MoE models on a single domain. Our experiments span web crawl (DCLM), code (StarCoder), academic (peS2o), and encyclopedic (Wikipedia) text. Trends are near-identical across all data domains: models underperform as repetition rates rise, with MoEs overfitting more rapidly.
DCLM- BASELINE with DCLM- POOL. We interpolate between all-unfiltered and all-filtered with 5 settings consisting of the following DCLM (baseline %, pool %): (0%, 100%), (25%, 75%), (50%, 50%), (75%, 25%), (100%, 0%). Filtering affects data repetition in some dense models. Training on a higher percentage of unfiltered DCLMPOOL data results in consistently worse performance (Figure 6). Despite the distribution shift that arises from stringent filtering, all MoEs trained on any mixture of DCLM- POOL and - BASELINE exhibit similar decline with data repetition. However, dense models trained exclusively on at least 50% unfiltered data may deteriorate more rapidly with R. As DCLM- BASELINE filtering reduces DCLM- POOL by a factor of over 40×, we compare R = 1 on 100% raw DCLMPOOL against R = 32 on 100% DCLM- BASELINE. In our setting, training on all unique unfiltered data is superior to repeating strictly filtered data, but an intermediate quality filter and repetition rate is likely superior to either extreme. 3.4
DATA R EPETITION AT M IXED R ATES W ITHIN D IVERSE DATA M IXES
In practice, LMs are trained on data mixes composed of diverse domains at varying levels of repetition. For example, encyclopedic text is often more constrained than web crawl data, and is therefore repeated at a higher rate. For example, GPT-3 was trained for 3.4 epochs over its Wikipedia component, but less than 1 epoch over Common Crawl (Brown et al., 2020). To investigate the effects of varied repetition rates within a data mix, we consider a simple setting in which the total data budget T is split between two domains. The first domain is DCLM, which is never repeated. The second is peS2o or StarCoder, which is repeated with R ∈ {1, 2, 4, 8, 16, 32}. We divide the total data budget T between the 2 domains (DCLM, peS2o) or (DCLM, StarCoder) at the following proportions: (100%, 0%), (90%, 10%), (50%, 50%), or (0%, 100%). Only the repeated domain’s unique pool shrinks as R grows. As a concrete example, the (50%, 50%) DCLM-StarCoder setting at R = 8 has T = 1.6B total tokens, consisting of 1.6B ∗ 0.5 = 0.8B unique DCLM tokens and 1.6B ∗ 0.5/8 = 0.1B unique StarCoder tokens repeated 8 times. 5
Data Scarcity and Model Sparsity
Model type
Dense (1 x 1) MoE (8 x 1/4) MoE (16 x 1/4) MoE (32 x 1/4) MoE (64 x 1/4) MoE (128 x 1/4) MoE (256 x 1/4)
6
5
1
2
4
8
Repetition Rate
16
Model type
7
Common Crawl CE (↓)
Common Crawl CE (↓)
7
Dense (1 x 1) MoE (64 x 1/32) MoE (64 x 1/16) MoE (64 x 1/8) MoE (64 x 1/4) MoE (64 x 1/2)
6
5
32
1
2
(a) Varied Total Expert Count (n)
4
8
Repetition Rate
16
32
(b) Varied Expert Size / Granularity (g)
Common Crawl CE (↓)
Figure 4: In MoEs, higher sparsity results in a greater reaction to data repetition (§3.1). We consider two settings: (a) fixed expert granularity, increased sparsity via greater total expert count; (b) fixed total expert count, increased sparsity via larger experts. Overfitting increases with total parameters, corresponding to (a) more total experts and (b) larger experts. Lines across (a) and (b) with the same color have identical total parameter count. 8 Model type Dense (1 x 1) MoE (64 x 1/4)
7 6 5
4 1
2
4
8
16
Repetition Rate
32
64
Figure 5: Data repetition effects are consistent across data-to-parameter ratios (§3.1). We increase total tokens per active parameter from 20 to 80, and observe no difference in data repetition overfitting between the two token-to-active-parameter ratios.
Mixing repeated StarCoder with all-unique DCLM does not impact repetition effects. In Figure 7a, we find that, regardless of the proportion of T assigned to StarCoder, code validation loss rises according to the same pattern in §3.1-3.2. Web crawl validation loss only rises with the repetition rate of StarCoder if it occupies at least half of T .
Mixing repeated peS2o with all-unique DCLM may have a regularizing effect. In Figure 7b, academic text validation loss increases with the peS2o repetition rate, but the effect is substantially dampened as peS2o occupies a smaller proportion of the total training data. As with the DCLM-StarCoder experiments above, web crawl validation loss suffers from peS2o repetition only when the peS2o total data proportion is high. This suggests that the unrepeated DCLM data component may have a regularizing effect on the peS2o data repetition.
Mixing repeated data with unrepeated data of a semantically similar domain may reduce data repetition effects. Academic text, as represented by peS2o, has higher semantic similarity to web text than code, as represented by StarCoder. Our results suggest a promising possibility: even at very high repetition rates, repetition effects might be reduced by mixing the repeated data domain with an non-repeated or less-repeated, semantically similar data domain. 6
Data Scarcity and Model Sparsity
Model type
Model type
Repetitions
% filtered
1x 2x 4x 8x 16x 32x
6
5
0
25
50
% filtered
75
100
Dense (1 x 1) MoE (64 x 1/4)
Common Crawl CE (↓)
Common Crawl CE (↓)
Dense (1 x 1) MoE (64 x 1/4)
0% 25% 50% 75% 100%
6
5
1
2
4
8
Repetition Rate
16
32
Figure 6: Repetition effects are consistent across data quality (§3.3). Validation CE Loss for 80M dense and MoE (64 x 1/4) on the DCLM- POOL and - BASELINE interpolation, where % filtered indicates the proportion of the data mix dedicated to the heavily filtered DCLM- BASELINE data. Left: CE against the DCLM- BASELINE %. A higher percentage of filtered data results in better performance. As in §3.1-3.2, MoE performance deteriorates more rapidly than dense at higher R, regardless of data. Right: CE Loss against repetition rate. Trends are largely similar across data quality and model architecture, though dense models trained on higher proportions of unfiltered data (0-50% DCLMBASELINE) may degrade more rapidly with R.
4
R EGULARIZATION T ECHNIQUES FOR DATA R EPETITION
To better understand the overfitting observed above, we apply known regularization methods with two motivations: to further probe the mechanisms responsible for performance decline under data repetition, and to take initial steps towards solutions. Dropout We use element-wise residual dropout (Srivastava et al., 2014) on the output of each sub-layer with probability p ∈ {0.0, 0.1, 0.2, 0.4}. High dropout probability hurts performance at low data repetition, but dramatically reduces the impact of extreme data repetition for all architectures (Figure 8a). Gradient Norm Clipping The gradient clipping threshold upper bounds the magnitude of each batch update, and is often used to decrease instability (Pascanu et al., 2013). We sweep this threshold over {0.2, 1.0, 2.0, None}, where None indicates no clipping. All four settings result in similar performance, within variance (Appendix Figure 16). Weight Decay We sweep decoupled AdamW weight decay over λ ∈ {0.05, 0.1, 0.2, 0.4}. Performance differences, though well-ordered, are small and remain within noise thresholds (Appendix Figure 16). FFN Output Masking FFN output masking (FOM) zeroes, with probability p, the entire dense FFN or MoE output for a token without rescaling. FOM is similar to a coarse-grained dropout applied only to FFNs. We consider p ∈ {0.0, 0.1, 0.2, 0.4} and observe that, at low repetition rate, it incurs a smaller penalty to performance than residual dropout. At high repetition rate, it has a similar regularizing effect to dropout (Figure 8b). Expert Dropout Expert dropout (Fedus et al. (2022); Zoph et al. (2022a)) applies dropout only to the expert hidden activation. We consider expert dropout rate in {0.0, 0.1, 0.2, 0.4}, and compare to the dense analogue, which applies dropout only to the FFN. Expert Dropout, like FOM, acts only on the FFN, and indeed behaves similarly to FOM (Figure 8c). Expert Output Masking Expert output masking (EOM) independently zeroes each (token, selected expert) output with probability p during training. No rescaling is applied. EOM thus operates only on the FFNs, like FOM and Expert Dropout, with an intermediate granularity. We consider p ∈ {0.0, 0.1, 0.2, 0.4} and observe that EOM has a similar effect to both FOM and Expert Dropout (Figure 8d). Router Jitter Router jitter perturbs routing, multiplying the router’s input by elementwise uniform noise on [1−ϵ, 1+ ϵ] during training. Over ϵ ∈ {0.0, 0.1, 0.2, 0.4}, we observe no clear impact on performance (Appendix Figure 16). 7
Data Scarcity and Model Sparsity
Stack CE (↓)
DCLM (R=1)
7
5
Stack CE (↓)
Common Crawl CE (↓)
Common Crawl CE (↓)
6
Dense (1 x 1) MoE (64 x 1/4)
% StarCoder
10% StarCoder 50% StarCoder 100% StarCoder
3
5 DCLM (R=1) 1
Model type
4
2
4
8
16
Repetition Rate
32
1
2
4
8
Repetition Rate
16
32
(a) DCLM + StarCoder
Common Crawl CE (↓)
peS2o CE (↓)
4.6 DCLM (R=1)
7
peS2o CE (↓)
Common Crawl CE (↓)
4.4 6
4.2
Model type
Dense (1 x 1) MoE (64 x 1/4)
4
% peS2o
3.8
10% peS2o 50% peS2o 100% peS2o
3.6
5 DCLM (R=1)
3.4 1
2
4
8
Repetition Rate
16
32
1
2
4
8
Repetition Rate
16
32
(b) DCLM + peS2o
Figure 7: Mixing a repeated data domain with a non-repeated domain may have a regularizing effect (§3.4). We mix peS2o and StarCoder, repeated R ∈ {1, 2, 4, 8, 16, 32} times, into non-repeated DCLM, with various proportions of peS2o/StarCoder. We find that DCLM + StarCoder mixes degrade similarly with repetition, regardless of the StarCoder mixing percentage. However, DCLM appears to have a regularizing effect in DCLM + peS2o mixes: as DCLM takes up a larger proportion of the mix, high repetition rates (R = 32) degrade at a slower rate.
5
M ECHANISTIC I NVESTIGATION OF DATA R EPETITION
In §3, we hypothesize that MoEs may suffer more from data repetition if each expert’s routed token set is fixed early in training, thus exposing each expert FFN to a significantly reduced set of unique tokens, when compared to dense FFNs. We consider two aspects of this hypothesis: firstly, the stability of the router, which would result in a fixed data partition over experts; secondly, the resulting specialization of each expert. 5.1
M O E ROUTER O SSIFICATION
We study routing patterns for each saved checkpoint: we run inference on a fixed batch of 16,384 tokens from the Dolma Common Crawl validation split, and record each token’s top-1 expert at every MoE layer. We define the routing stability at checkpoint i as the fraction of tokens whose top-1 expert did not change from checkpoint i − 1, averaged over layers. Separately, we consider expert co-activation and router load balance. Further details are in Appendix A.6. Routing ossifies early. We compute routing stability for MoE models in §3. Figure 9a shows results for 80M models. Routing is unstable (random chance, 0.02 ≈ 1/64) at the first checkpoint (step 200), as models begin from random initialization. By the next checkpoint (step 400, 10% of training), routing stability rises to 60%, and continues to rise rapidly to > 95% at the end of training. Each MoE configuration at each scale exhibits similar patterns (Appendix B.8). 8
Data Scarcity and Model Sparsity
10
Dense (1 x 1) MoE (64 x 1/4) Dropout
8
None (disabled) 0.1 0.2 0.4
7 6
Model type
Dense (1 x 1) MoE (64 x 1/4)
9
Common Crawl CE (↓)
9
Common Crawl CE (↓)
10
Model type
FOM Prob.
8
0.0 0.1 0.2 0.4
7 6 5
5 1
2
4
8
16
Repetition Rate
32
1
64
2
(a) Dropout
8
16
32
64
(b) Final Output Masking
10
10
Model type
Dense (1 x 1) MoE (64 x 1/4) Expert Dropout
8
0.0 0.1 0.2 0.4
7 6 5
Model type
Dense (1 x 1) MoE (64 x 1/4)
9
Common Crawl CE (↓)
9
Common Crawl CE (↓)
4
Repetition Rate
EOM Prob.
8
0.0 0.1 0.2 0.4
7 6 5
1
2
4
8
16
Repetition Rate
32
64
1
(c) Expert Dropout
2
4
8
16
Repetition Rate
32
64
(d) Expert Output Masking
Figure 8: Dropout (a), FFN Output Masking (b), Expert Dropout (c), and Expert Output Masking (d) each reduces overfitting from data repetition (§4). Of the regularization methods studied in §4, these 4 dramatically decrease the response to data repetition. However, weight decay and gradient norm clipping, as well as MoE router jitter, have minimal effect. See additional figures in Appendix B.5. Higher data repetition exacerbates router ossification. In Figure 9a, router stability rises uniformly across all data repetition rates until step 600, at which point higher repetition models begin to show consistently higher routing stability. In Figure 9b, end-of-training router stability rises with data repetition, across model scales and MoE sparsities (Appendix B.8). This supports our hypothesis that each expert’s token set is nearly stationary for most of training, so under repetition an expert sees the same reduced shard of data over and over. Dropout’s regularizing effect does not operate through router plasticity. We train 200M models using the dropout settings of §4 (with checkpoints 1,000 steps apart, not directly comparable to the above results). Late-training router stability again rises slowly but consistently with repetition (Appendix Figure 22). At R = 64, models with dropout show less routing ossification than the no-dropout setting, even though they overfit far less (§4). Since dropout partially recovers performance without affecting router ossification, we hypothesize that overfitting is not primarily driven by routing, but rather the functions learned by each expert. Higher repetition encourages uniformly distributed expert co-activation. Although the top-1 expert contributes the majority of the output weight, we also conduct a simple investigation of expert co-activation. For each unordered pair of experts within each layer, we count how often both experts in the pair are active for the same token. We normalize the counts and compute the entropy of the distribution. Under repetition this distribution trends toward uniform in every configuration (Appendix B.8). Router output magnitude is higher with data repetition; load balancing shows no clear correlation. Router load imbalance, or the ratio between the maximum and mean expert token loads in each batch, does not predictably shift with repetition rate, except that extremely high R yields outliers 1B scale. Load balancing loss also does not change predictably with repetition. Z-loss, which measures router logit magnitude, is higher in early training with higher R, which is consistent with earlier and more extreme router ossification (Appendix Figure 13-15). 9
Data Scarcity and Model Sparsity
Model type MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x 64x 128x 256x
Top-1 routing stability
0.8 0.6 0.4 0.2 0.0
200
400
600
800
1000
Training step
1200
Routing stability, end of training
MoE (64 x 1/4)
1.0
Model type MoE (32 x 1/4) MoE (64 x 1/4) Scale 80M 200M
0.94 0.93 0.92
1400
1
(a) Top-1 routing stability over training (80M)
2
4
8
Repetition Rate
16
32
(b) Final top-1 routing stability (80M, 200M)
Model type
MoE (16 x 1/4) MoE (32 x 1/4) MoE (64 x 1/4) MoE (128 x 1/4)
0.01 0.005 0.003 0.002 0.001 1
2
4
8
16
Repetition Rate
32
64
Expert knockout: median CE increase
Expert knockout: median CE increase
Figure 9: Routing ossifies early in training, exacerbated by repetition (§5.1). We show Top-1 routing stability, defined as the fraction of held-out Common Crawl tokens that keep the same top-1 expert between consecutive checkpoints (200 steps apart), averaged over layers. Left: For the 80M MoE (64 x 1/4) models, consecutive checkpoint agreement is near chance (1/64) at the start of training, but over 0.9 for the second half of training. Stability increases with repetition, up to R = 32, where the router appears to destabilize. Right: We show the Top-1 routing stability at the end of training for MoE (32 x 1/4) and (64 x 1/4) at 80M and 200M scale. Stability rises more aggressively with R at 200M active parameters. Also see Appendix Figure 19-22. 0.01
Model type
0.005
Dropout
MoE (32 x 1/4) MoE (64 x 1/4) 0 0.1 0.4
0.003 0.002 0.001 0.0005 0.0003
1
2
4
8
Repetition Rate
16
32
Figure 10: Repetition increases expert specialization; dropout reduces this effect (§5.2). We compute the increase in held-out CE when one expert’s output is zeroed at inference time, and show the median over all experts and MoE layers of the final checkpoint. Left (80M active parameters): Expert specialization grows with repetition rate, and is overall higher when there are fewer total experts, but is unaffected by expert granularity (Appendix Figure 20) Right (200M active parameters): Dropout reduces expert specialization across all R, and also reduces the relative magnitude of the increase in expert specialization that results from higher R. 5.2
M O E E XPERT S PECIALIZATION
Early routing ossification does not necessitate expert specialization; experts with disjoint routed token sets may still learn similar functions. We investigate the specialization of experts by measuring the performance impact of expert knockout, or the loss impact of removing an expert. Specifically, we perform inference with the final checkpoint of each model and measure, for each expert at each layer, the total increase in CE Loss resulting from masking only that expert’s outputs. Expert specialization rises with data repetition. Knockout cost at R = 1 is higher at low total expert count. As R increases, knockout cost increases for all configurations, but does so more rapidly at higher expert counts. At 80M (Figure 10a), increasing R from 1 to 32 also increases expert knockout effect by 1.1× for 16 experts, and by 2.3× for 128 experts of 1/4 granularity. This pattern echos §3, where CE degradation under repetition also grows with expert count. Thus, repetition encourages expert specialization and reduces redundancy, especially at higher expert counts. Dropout decreases expert specialization. We compare various dropout settings at 200M (Figure 10b). Dropout consistently reduces the median knockout cost for all values of R. We hypothesize that dropout combats overfitting by removing dependence on any single expert, and enforcing multiple, more varied, representations of features. 10
Data Scarcity and Model Sparsity
6
R ELATED W ORKS
Hernandez et al. (2022) show that repeating a small fraction of the training data degrades held-out loss nonmonotonically. Muennighoff et al. (2023) studies dense models trained on C4 and finds that repetition up to roughly four epochs is comparable to all-unique data, but further repetition decays performance. More recent work on dense models has provided evidence that data repetition damage: (1) grows with model scale (Kazdan et al., 2026); (2) peaks at an intermediate repeat count (Chudnovsky et al., 2026); (3) is well described by a single additive coefficient that strong weight decay can shrink (Lovelace et al., 2026); and (4) can be modulated by data quality and mixture weights (Fang et al., 2025; Liu et al., 2026; Chen et al., 2025). Very little work has considered MoEs; Xue et al. (2023) concludes from a single MoE configuration, a 16-expert T5, that parameter count drives multi-epoch degradation while FLOPs are close to irrelevant. Other work has modeled the tradeoff between quality and quantity of unique data when training dense Transformers. Fang et al. (2025) reports that repeating a heavily filtered set up to ten times can beat a single pass over a superset ten times larger, while Mohri et al. (2026) argues that with enough compute, larger quantities of unfiltered data is better. Liu et al. (2026) and Chen et al. (2025) fit quality-weighted mixtures and repetition together. Even fewer studies have considered interventions for minimizing data-repetition effects. Xue et al. (2023) uses dropout to reduce multi-epoch degradation, but caution that the dropout probability needs retuning as models grow. Lovelace et al. (2026) shows that raising weight decay by an order of magnitude cuts their fitted overfitting coefficient by roughly 70% at high repetition rates, at the price of a loss premium in the single-epoch regime.
7
C ONCLUSION
Across all dense and MoE architectures, we find that increasing data repetition rate R leads to overfitting. Sparse MoE models deteriorate earlier and more rapidly as R increases. We show that the response to data repetition primarily depends on total parameters. The pattern of overfitting effects are remarkably robust to a variety of single data domains, data mixes, and different levels of data filtering. Mixing a repeated data domain into a larger or equal-sized nonrepeated data domain may have a regularizing effect. We successfully reduce the overfitting response to data repetition through regularization methods that operate by dropping parameter outputs (dropout, expert dropout, FFN output masking, and expert output masking). However, no method fully matches performance achieved with all-unique data. Gradient Norm Clipping, Weight Decay, and Router Jitter operate through qualitatively different mechanisms to reduce update strength, reduce weight magnitude, and modify coarse-grained gradient paths, respectively, and do not have any measurable effect. Finally, our mechanistic analyses provide evidence that MoE routing is fixed early in training, and only slightly exacerbated by data repetition. Expert specialization also rises with data repetition. Dropout does not affect router ossification, but decreases expert specialization. In summary, our work studies the interaction between data constraints and the design of MoE architectures. We present strong evidence that data repetition causes overfitting, that the degree of performance degradation primarily depends on total parameters, and that the underlying mechanism operates partially through overly specialized parameters, broken by methods such as dropout and output masking. We recommend that future works further explore masking-based methods to minimize repetition-driven overfitting through decreased parameter specialization.
ACKNOWLEDGMENTS We are grateful to Rohan Sanda for initial engineering support; to Ananya Harsh Jha and Jacqueline He for helpful discussion; and to the contributors and maintainers of the UW Hyak and Stanford Marlowe computing resources.
R EFERENCES Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An open language model for mathematics, 2023. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. Zhengyu Chen, Siqi Wang, Teng Xiao, Yudong Wang, Shiqi Chen, Xunliang Cai, Junxian He, and Jingang Wang. Revisiting scaling laws for language models: The role of data quality and training strategies. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), pp. 23897–23920, 2025. URL https://aclanthology.org/2025.acl-long.1163/. 11
Data Scarcity and Model Sparsity
Jessica Chudnovsky, Joshua Kazdan, Noam Levi, Rylan Schaeffer, Yegor Denisov-Blanch, Bo He, Mehmet Donmez, Sanmi Koyejo, and David Donoho. Internal data repetition destroys language models, 2026. URL https:// arxiv.org/abs/2606.24998. Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019. Together Computer. Redpajama: An open source recipe to reproduce llama training dataset, 2023. URL https: //github.com/togethercomputer/RedPajama-Data. Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models, 2024. URL https://arxiv.org/abs/2401.06066. Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 1286–1305, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/ v1/2021.emnlp-main.98. URL https://aclanthology.org/2021.emnlp-main.98. Alex Fang, Hadi Pouransari, Matt Jordan, Alexander Toshev, Vaishaal Shankar, Ludwig Schmidt, and Tom Gunter. Datasets, documents, and repetitions: The practicalities of unequal data quality, 2025. URL https://arxiv. org/abs/2503.07879. William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022. URL http://jmlr.org/ papers/v23/21-0998.html. Leo Gao, Stella Rose Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling. ArXiv, abs/2101.00027, 2020. URL https://api.semanticscholar.org/ CorpusID:230435736. Sidney Greenbaum and Gerald Nelson. The international corpus of english (ICE) project. World Englishes, 15 (1):3–15, mar 1996. doi: 10.1111/j.1467-971x.1996.tb00088.x. URL https://doi.org/10.1111%2Fj. 1467-971x.1996.tb00088.x. Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al. Scaling laws and interpretability of learning from repeated data. arXiv preprint arXiv:2205.10487, 2022. Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. Training compute-optimal large language models, 2022. URL https://arxiv.org/abs/2203.15556. Joshua Kazdan, Noam Levi, Rylan Schaeffer, Jessica Chudnovsky, Abhay Puri, Bo He, Mehmet Donmez, Sanmi Koyejo, and David Donoho. Scale dependent data duplication, 2026. URL https://arxiv.org/abs/2603. 06603. Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries. The stack: 3 tb of permissively licensed source code, 2022. URL https://arxiv.org/abs/2211.15533. Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, ChengYu Hsieh, Dhruba Ghosh, Josh Gardner, Maciej Kilian, Hanlin Zhang, Rulin Shao, Sarah Pratt, Sunny Sanyal, Gabriel Ilharco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang, Khyathi Chandu, Thao Nguyen, Igor Vasiljevic, Sham Kakade, Shuran Song, Sujay Sanghavi, Fartash Faghri, Sewoong Oh, Luke Zettlemoyer, Kyle Lo, Alaaeldin El-Nouby, Hadi Pouransari, Alexander Toshev, Stephanie Wang, Dirk Groeneveld, Luca Soldaini, Pang Wei Koh, Jenia Jitsev, Thomas Kollar, Alexandros G. Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar. Datacomp-lm: In search of the next generation of training sets for language models, 2024. Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. Starcoder: may the source be with you!, 2023. 12
Data Scarcity and Model Sparsity
Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher R’e, Diana Acosta-Navas, Drew A. Hudson, E. Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel J. Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan S. Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas F. Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. Holistic evaluation of language models. Annals of the New York Academy of Sciences, 1525:140 – 146, 2022. URL https://api.semanticscholar.org/CorpusID:253553585. Fengze Liu, Weidong Zhou, Binbin Liu, Ping Guo, Zijun Wang, Bingni Zhang, Yifan Zhang, Yifeng Yu, Xiaohuan Zhou, and Taifeng Wang. Infolaw: Information scaling laws for large language models with quality-weighted mixture data and repetition, 2026. URL https://arxiv.org/abs/2605.02364. Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. Justin Lovelace, Christian Belardi, Srivatsa Kundurthy, Shriya Sudhakar, and Kilian Q. Weinberger. Prescriptive scaling laws for data constrained training, 2026. URL https://arxiv.org/abs/2605.01640. Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Pete Walsh, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson, and Jesse Dodge. Paloma: A benchmark for evaluating language model fit, 2024. URL https://arxiv.org/abs/2312.10523. Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. ArXiv, abs/1609.07843, 2016. URL https://api.semanticscholar.org/CorpusID:16299141. Christopher Mohri, John Duchi, and Tatsunori Hashimoto. A bitter lesson for data filtering, 2026. URL https: //arxiv.org/abs/2605.19407. Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. Scaling data-constrained language models. In Advances in Neural Information Processing Systems (NeurIPS), 2023. URL https://arxiv.org/abs/2305.16264. Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Weijia Shi, Pete Walsh, Oyvind Tafjord, Nathan Lambert, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, David Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi, Noah A. Smith, Pang Wei Koh, Amanpreet Singh, and Hannaneh Hajishirzi. Olmoe: Open mixture-of-experts language models, 2025. URL https://arxiv.org/abs/2409.02060. Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. On the difficulty of training recurrent neural networks. In International conference on machine learning, pp. 1310–1318. PMLR, 2013. Keiran Paster, Marco Dos Santos, Zhangir Azerbayev, and Jimmy Ba. Openwebmath: An open dataset of high-quality mathematical web text, 2023. Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. ArXiv, abs/1910.10683, 2019. URL https://api.semanticscholar.org/CorpusID:204838007. Machel Reid, Victor Zhong, Suchin Gururangan, and Luke Zettlemoyer. M2D2: A massively multi-domain language modeling dataset. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 964–975, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.63. Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017. Luca Soldaini and Kyle Lo. peS2o (Pretraining Efficiently on S2ORC) Dataset, 2023. URL https://github. com/allenai/pes2o. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo. Dolma: an open corpus of three trillion tokens for language model pretraining research, 2024a. 13
Data Scarcity and Model Sparsity
Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, A. Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Daniel Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A. Smith, Hanna Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo. Dolma: an open corpus of three trillion tokens for language model pretraining research. ArXiv, abs/2402.00159, 2024b. URL https://api.semanticscholar.org/CorpusID:267364861. Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. Kmmlu: Measuring massive multitask language understanding in korean, 2024. URL https://arxiv.org/abs/2402.11548. Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014. Fuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng, and Yang You. To repeat or not to repeat: Insights from scaling LLM under token-crisis. In Advances in Neural Information Processing Systems (NeurIPS), 2023. URL https://arxiv.org/abs/2305.13230. Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019. URL https://arxiv.org/abs/1905.07830. Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. Stmoe: Designing stable and transferable sparse expert models, 2022a. URL https://arxiv.org/abs/2202. 08906. Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. St-moe: Designing stable and transferable sparse expert models. arXiv preprint arXiv:2202.08906, 2022b. [221] in OLMoE.
14
Data Scarcity and Model Sparsity
A
E XPERIMENTAL D ETAILS
A.1
M ODEL A RCHITECTURE
Scale
Layers
Model Dim
Attention Heads
80M
8
336
7
200M
1B
10
15
640
1664
10
16
Name
Activation Sparsity (s)
Total Experts (n)
Active Experts (k)
Expert Gran. (g)
dense MoE (8 x 1/4) MoE (64 x 1/32) MoE (16 x 1/4) MoE (64 x 1/16) MoE (32 x 1/4) MoE (64 x 1/8) MoE (64 x 1/4) MoE (64 x 1/2) MoE (128 x 1/4) MoE (256 x 1/4)
1 2 2 4 4 8 8 16 32 32 64
8 64 16 64 32 64 64 64 128 256
4 32 4 16 4 8 4 2 4 4
1/4 1/32 1/4 1/16 1/4 1/8 1/4 1/2 1/4 1/4
dense MoE (32 x 1/4) MoE (64 x 1/4)
1 8 16
32 64
4 4
1/4 1/4
dense MoE (64 x 1/4)
1 16
64
4
1/4
Total Tokens (T )
Active Param (Na )
Total Param (N )
1.6B
81.8M
81.8M 92.7M 92.7M 114.4M 114.4M 157.7M 157.7M 244.4M 417.8M 417.8M 764.6M
4B
193.9M
193.9M 538.0M 931.2M
20B
998.3M
998.3M 8.5B
Table 1: Architecture Details and Parameter Counts A.2
H YPERPARAMETERS Hyperparameter
Value
Vocabulary size Batch Size Sequence Length Learning Rate Encoder-Decoder Weight Sharing Feedforward Dimension LR Schedule LR Warmup End LR Weight Decay Max Grad Norm Dropout Nonlinearity MoE Z-loss Load Balancing Loss Weight MoE Token Dropping MoE Routing Choice
50K 512 2048 4e-4 No 4 x hidden dimension Cosine Decay 2000 steps 0.1 x Peak LR {0.0, 0.1, 0.2, 0.4 } {None, 0.2, 1, 2.0 } {0.0, 0.1, 0.2, 0.4} SwiGLU 1e-3 1e-2 Dropless Token Choice
Table 2: Hyperparameter details for models in §3-5. Multiple values indicate that we investigated different settings in §4, and bold values are defaults used in §3. A.3
T RAINING DATA S OURCES
We take our training data from Muennighoff et al. (2025). We use their data mix, which we call OLMoE Mix, consisting of documents from: DCLM-Baseline (Li et al., 2024), StarCoder (Li et al., 2023; Kocetkov et al., 2022), peS2o (Soldaini & Lo, 2023; Soldaini et al., 2024a), arXiv (Computer, 2023), OpenWebMath (Paster et al., 2023), Algebraic Stack (Azerbayev et al., 2023), English Wikipedia & Wikibooks (Soldaini et al., 2024a). In §3.3, we also use DCLM-P OOL (Li et al., 2024). A.4
E VALUATION DATA
Our evaluation includes held-out validation sets for language modeling, as well as downstream tasks. The language modeling tasks are a subset of Paloma (Magnusson et al., 2024), which consists of: C4 (Raffel et al. (2019) via Dodge et al. (2021)), T HE P ILE (Gao et al., 2020), W IKI T EXT-103 (Merity et al., 2016), D OLMA (Soldaini 15
Data Scarcity and Model Sparsity
et al., 2024b), M2D2 S2ORC (Reid et al., 2022), ICE (Greenbaum & Nelson (1996) via Liang et al. (2022)). D OLMA is subdivided into six domains: books, common-crawl, pes2o, reddit uniform, stack uniform, wiki. In §3.2, we vary the dataset used for validation loss to match the training domain. For models trained on DCLM, we evaluate on Dolma common-crawl; for peS2o, Dolma pes2o; for Wikipedia, Dolma wiki; for StarCoder, Dolma stack uniform. The downstream tasks consist of BoolQ (Clark et al., 2019), HellaSwag (Zellers et al., 2019), and MMLU (Son et al., 2024). MMLU is subdivided into four domains: humanities, STEM, social sciences, and other. A.5
DATA R EPETITION
In data-constrained regimes, the set of U unique tokens used for each experiment is constructed as follows: we fix a random permutation of the sequences in each data domain D, and take the first U tokens for training. We repeat these U tokens for R epochs, shuffling between epochs. In §3.3, we mix tokens from DCLM-BASELINE and DCLM-P OOL (Li et al., 2024). This inevitably introduces a very small amount of unmeasured data repetition because our selected subsets of DCLM-BASELINE and DCLM-P OOL may have a non-empty intersection. However, this intersection is likely to be of negligible size. DCLM-BASELINE consists of 5T tokens, of which we use 1.6B at most. DCLM-P OOL consists of 240 trillion tokens, of which we use 1.6B at most. The probability that any particular token in a particular sequence from our DCLM-BASELINE subset also appears in our DCLM-P OOL is less than 2e-9, yielding an expected total of fewer than 3 repeated tokens. In other words, we expect effectively no repeated tokens on average. A.6
ROUTING A NALYSIS
In §5, we use a batch size of 16,384 tokens (8 sequences of 2,048 tokens). We use a fixed subset of the Dolma Common Crawl validation split. Dropout, router jitter, and other train-time regularizers are inactive. Ossification (§5.1). For each checkpoint we record the top-1 expert of every token at every MoE layer, defined as the expert with the largest router score. For each pair of consecutive checkpoints we report the fraction of tokens whose top-1 expert is identical, averaged over layers. Checkpoints are 200 steps apart in the core ladders and 1,000 steps apart in the 200M dropout arms, so stability values are only compared between runs with matching spacing. End-of-training stability is the mean over the last two regularly spaced intervals. Expert knockout (§5.2). On the final checkpoint we zero all MLP weights of one expert, so its output is exactly zero for the tokens routed to it. The router is untouched and the weights of the remaining selected experts are not renormalized. We then recompute CE on the token batch, repeat for every expert in every MoE layer, and report the median and maximum increase over the CE of the unmodified model. Co-activation (§5.2). On the final checkpoint we count, per layer, how often each unordered pair of experts appears together in a token’s top-k set, normalize the counts to a distribution, and compute its Shannon entropy. We divide by log n(n − 1) , its maximum for n experts, so values are comparable across expert counts.
16
Data Scarcity and Model Sparsity
B
A DDITIONAL R ESULTS
B.1
R ANDOM S EED VARIANCE
In Table 3, we report the variance across 5 random seeds for Dense and MoE (64 x 1/4) models trained on the OLMoE mix at R = 1, R = 32, and evaluated on all language modeling validation datasets and downstream tasks. Metric
Dense R=1 R=32 Mean Std. Dev. Mean Std. Dev.
MoE (64 x 1/4) R=1 R=32 Mean Std. Dev. Mean Std. Dev.
Train Loss
4.57
0.00
4.27
0.00
4.26
0.01
3.17
0.02
Validation LM Loss C4 Dolma Books Dolma Common Crawl Dolma peS2o Dolma Reddit Dolma Stack Dolma Wiki ICE M2D2 S2ORC Pile WikiText-103 Average
4.86 5.04 4.92 4.48 4.74 4.65 4.67 4.97 4.84 4.62 5.06 4.81
0.01 0.00 0.01 0.01 0.00 0.02 0.01 0.01 0.01 0.01 0.00 0.01
5.27 5.59 5.29 4.93 5.12 6.68 5.16 5.63 5.47 5.37 5.76 5.48
0.02 0.06 0.02 0.04 0.02 0.10 0.04 0.03 0.03 0.05 0.08 0.05
4.53 4.72 4.61 4.12 4.46 4.23 4.31 4.66 4.52 4.27 4.67 4.46
0.01 0.01 0.01 0.01 0.01 0.02 0.01 0.02 0.01 0.01 0.02 0.01
6.04 6.48 6.05 5.66 5.95 7.46 5.87 6.73 6.46 6.19 6.67 6.32
0.01 0.05 0.02 0.01 0.02 0.04 0.02 0.01 0.01 0.01 0.03 0.02
Downstream Task Loss BoolQ HellaSwag MMLU Humanities MMLU Other MMLU Social Sciences MMLU STEM Average
2.52 0.96 2.13 1.97 1.98 2.02 1.93
0.22 0.00 0.06 0.04 0.04 0.03 0.07
3.14 1.05 3.48 2.85 2.66 2.61 2.63
0.18 0.00 0.11 0.17 0.14 0.07 0.11
2.33 0.89 1.91 1.80 1.88 1.90 1.78
0.24 0.00 0.13 0.06 0.10 0.08 0.10
3.95 1.23 4.36 2.93 3.06 3.13 3.11
0.43 0.00 0.92 0.20 0.31 0.48 0.39
Downstream Task Accuracy BoolQ HellaSwag MMLU Humanities MMLU Other MMLU Social Sciences MMLU STEM Average
0.39 0.26 0.24 0.27 0.23 0.27 0.28
0.01 0.00 0.01 0.01 0.01 0.00 0.01
0.39 0.25 0.24 0.24 0.22 0.24 0.26
0.01 0.00 0.00 0.01 0.00 0.01 0.01
0.41 0.26 0.25 0.27 0.24 0.27 0.28
0.05 0.00 0.01 0.01 0.01 0.01 0.01
0.45 0.25 0.25 0.25 0.23 0.24 0.28
0.03 0.00 0.01 0.01 0.01 0.01 0.01
Table 3: Mean and Standard Deviation across 5 random seeds. We repeat a selection of the 80M settings from §3.1 using 5 random seeds for model initialization, and repeat the mean and standard deviation on each metric. Variance at R = 1 is near-0 for validation LM datasets. Despite fixing the data used across model initialization seeds, higher data repetition rates yield slightly higher standard deviation. Downstream task loss has higher variance for MoE models, and at higher R. Downstream task accuracy has low variance, but remains at near-chance scores.
17
Data Scarcity and Model Sparsity
E XTENDED S ETTINGS (§3.1) Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
1
Model type
Common Crawl CE (↓)
Train CE (↓) (last-100-step avg)
B.2
0.1 0.01 0.001
Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6 5
0.0001
20 23 26 29 212 215 218 221
20
Repetition Rate
23
26
29 212 215 218 221
Repetition Rate
20
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
1
Common Crawl CE (↓)
Train CE (↓) (last-100-step avg)
(a) 80M
0.1
0.01
0.001
Model type
Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6 5 4
1
4
16
64
Repetition Rate
256
1
1024
4
16
64
Repetition Rate
256
1024
Model type Dense (1 x 1) MoE (64 x 1/4)
1
Model type
Dense (1 x 1) MoE (64 x 1/4)
Common Crawl CE (↓)
Train CE (↓) (last-100-step avg)
(b) 200M
0.1
0.01
10 9 8 7 6 5 4 3
1
2
4
8 16 32 64 128 256
Repetition Rate
1
2
4
8 16 32 64 128 256
Repetition Rate
(c) 1B
Figure 11: Across active parameter scales, data repetition rates over 8 result in increasingly severe overfitting. Sparser models overfit more (§3.1). At 80M, 200M, and 1B active parameters, we fix the total data budget T = 20 · Na , and vary the data repetition rate R via different sized unique token sets. As R increases, models increasingly overfit, as we observe decreasing train loss and rising validation loss. Sparsity exacerbates overfitting behavior. Larger sparse models overfit more at lower R. We also consider two additional data repetition rates R ≈ 215 , 220 at 80M active parameters, and find that validation loss rises again.
18
Data Scarcity and Model Sparsity
B.3
T RAINING AND VALIDATION L OSS C URVES (§3.1) Dense (1 x 1)
Train CE (↓)
10
64x 128x
1
0.1 0.01
0.01
0.001
1024x
0.001
1000
1500
Train Step
32x 64x
1
0.1 512x
500
64x
Repetitions
128x
0.01 0
MoE (64 x 1/4)
10
1
256x
0.1
MoE (32 x 1/4)
10
256x 512x 1024x
0
500
1000
1500
Train Step
0.001
128x
256x 512x 1024x
0
500
1000
1500
Train Step
1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x
Common Crawl CE (↓)
(a) Train loss over training (80M) Dense (1 x 1)
MoE (32 x 1/4) 1024x 128x
10 9 8 7 6 5
64x 32x 16x
0
500
1000
256x 128x 512x 1024x
10 9 8 7 6 5
1500
Train Step
64x
32x 16x
0
500
Repetitions
MoE (64 x 1/4)
512x 256x
1000
10 9 8 7 6 5
32x 16x
1500
Train Step
1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x
128x 256x 512x 1024x 64x
0
500
1000
1500
Train Step
(b) Validation loss over training (80M) Dense (1 x 1)
MoE (32 x 1/4)
10
MoE (64 x 1/4) 10
10
16x
Train CE (↓)
16x
1
80x
0.1
128x
0.01
1
32x
1
40x
32x
0.1
0.1
0.01
160x
320x
0.001
0.01
128x 160x
640x
0
1000
2000
Train Step
3000
4000
80x 128x
80x
0
1000
2000
Train Step
3000
0.001
4000
320x 640x
0
1000
2000
Train Step
3000
4000
Repetitions 1x 2x 4x 8x 16x 32x 40x 80x 128x 160x 320x 640x
(c) Train loss over training (200M) Dense (1 x 1)
Common Crawl CE (↓)
20
MoE (32 x 1/4)
MoE (64 x 1/4)
160x 128x
80x 128x 160x
80x
10 9 8 7 6 5 4
40x 32x 16x
0
1000
2000
Train Step
3000
4000
10 9 8 7 6 5 4
40x 32x
16x 8x
0
1000
2000
Train Step
3000
10 9 8 7 6 5
16x
4
4000
(d) Validation loss over training (200M)
19
80x 640x 128x 320x 32x
8x
0
1000
2000
Train Step
3000
4000
Repetitions 1x 2x 4x 8x 16x 32x 40x 80x 128x 160x 320x 640x
Data Scarcity and Model Sparsity
Dense (1 x 1)
MoE (64 x 1/4) 10
Train CE (↓)
10 16x 32x
1
Repetitions 8x
1
16x
0.1
0.1 64x
0.01
0.01
32x 64x 128x 256x
128x 256x
0
5000
10000 15000 20000
0
Train Step
5000
10000 15000 20000
1x 2x 4x 8x 16x 32x 64x 128x 256x
Train Step
(e) Train loss over training (1B)
Common Crawl CE (↓)
Dense (1 x 1)
MoE (64 x 1/4) 64x
10 9 8 7 6 5 4
32x
16x 8x
5000
10000
Train Step
15000
32x
10 9 8 7 6 5 4 3
16x
8x 4x
5000
10000
Train Step
15000
Repetitions 1x 2x 4x 8x 16x 32x 64x 128x 256x
(f) Validation loss over training (1B)
Figure 12: At higher repetition rates, training loss falls to 0, which suggests overfitting to the repeated data (§3.1). We show the training and validation loss curves over the course of training. We consistently observe that higher repetition results in training loss curves that approach 0, mirrored by validation loss curves that rise. At sufficiently high R, we observe the double descent phenomenon Muennighoff et al. (2023), in which validation curves peak then fall.
20
Data Scarcity and Model Sparsity
Router load imbalance (↓)
B.4
ROUTING L OAD BALANCE AND S TABILITY
MoE (32 x 1/4)
8 7 6 5 4 3
MoE (64 x 1/4)
Repetitions
10 5
2 512x 128x 256x 16x
0
500
1000
3 2
1500
Train Step
128x
0
500
1000
Train Step
1500
1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x
Router load imbalance (↓)
(a) 80M active parameters
MoE (32 x 1/4)
Repetitions
MoE (64 x 1/4)
7 6 5 4 3
10 5
2
80x 8x 16x
0
1000
2000
3000
Train Step
4000
3 2
128x 1x 320x 640x
0
1000
2000
Train Step
3000
4000
1x 2x 4x 8x 16x 32x 40x 64x 80x 128x 160x 320x 640x
Router load imbalance (↓)
(b) 200M active parameters
MoE (64 x 1/4)
Repetitions
10 5 3 2
1x 8x 32x 64x 4x 128x 256x
0
5000
10000 15000 20000
1x 2x 4x 8x 16x 32x 64x 128x 256x
Train Step
(c) 1B active parameters
Figure 13: Routing imbalance training curves do not follow clear patterns at lower repetition, but are outliers at high repetition. We plot the routing imbalance, defined as the ratio between the maximum and median expert load, averaged over tokens in the batch. At small scale, routing imbalance does not appear correlated with load imbalance until R > 128, where curves become outliers. At 1B scale, load imbalance curves become disordered, and appear to partially cycle with data repetition periods.
21
Data Scarcity and Model Sparsity
MoE (32 x 1/4)
MoE (64 x 1/4)
Router LB loss (↓)
Repetitions
10 9
10 9
8
8 0
250
500
750
1000 1250 1500
1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x
0
Train Step
250
500
750
1000 1250 1500
Train Step
(a) 80M active parameters
MoE (32 x 1/4)
MoE (64 x 1/4)
Router LB loss (↓)
30
Repetitions
20 20
10
10 0
1000
2000
3000
Train Step
4000
0
1000
2000
Train Step
3000
4000
1x 2x 4x 8x 16x 32x 40x 80x 128x 160x 320x 640x
(b) 200M active parameters MoE (64 x 1/4)
Router LB loss (↓)
100 Repetitions 1x 2x 4x 8x 16x 32x 64x 256x
50 30 20 0
5000
10000
Train Step
15000
20000
(c) 1B active parameters
Figure 14: Routing load balancing loss training curves do not follow clear patterns at small model scale, but may correlate with repetition at larger scale. We report the load balancing loss, as defined in §2. At 80M and 200M, all settings show similar curves. At 1B, high repetition rates appear to affect load balancing loss, with intermediate values of R = 8, 16, 32 resulting in periodicity, and outlier curves at R = 64, 128, 256. It is possible that repetition itself increases load balancing loss, but that the extremely low training loss at high R results in a relatively strong optimization signal from auxiliary losses, eventually driving load balancing loss to fall.
22
Data Scarcity and Model Sparsity
MoE (32 x 1/4)
MoE (64 x 1/4) 200
Router Z loss (↓)
100 128x
50 30 20 10
64x
50
256x 32x
20
512x
10
16x
5 3 2 0
250
500
750
Repetitions
100
8x
5
1024x
2
1000 1250 1500
64x 128x 32x 256x 16x 8x 512x
1024x
0
Train Step
250
500
750
1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x
1000 1250 1500
Train Step
(a) 80M active parameters
Router Z loss (↓)
MoE (32 x 1/4) 200 100 50 40x 32x
20 10 5
80x 16x 128x 160x 8x 4x
2 0
1000
2000
Train Step
Repetitions
MoE (64 x 1/4)
3000
200 100 50 20 10 5 2 1
4000
32x 16x 64x 80x 8x 128x 4x 320x 640x
0
1000
2000
Train Step
3000
4000
1x 2x 4x 8x 16x 32x 40x 64x 80x 128x 160x 320x 640x
(b) 200M active parameters
MoE (64 x 1/4) Repetitions
Router Z loss (↓)
1000 100 16x 32x 8x 64x 4x 128x
10 1
256x
0
5000 10000 15000 20000
1x 2x 4x 8x 16x 32x 64x 128x 256x
Train Step
(c) 1B active parameters
Figure 15: Router z-loss training curves cycle with data repetition. We report the router z-loss (Zoph et al., 2022b). High repetition rates appear to affect z-loss, data-repetition driven cycles at 200M and 1B scale. Higher repetition rates result in higher z-loss with a double peak relatively early in training, resolving to a lower final z-loss. It is possible that high repetition typically results in higher z-loss, but that the extremely low training loss at high R results in a relatively strong optimization signal from auxiliary losses, eventually driving z-loss to fall.
23
Data Scarcity and Model Sparsity
R EGULARIZERS (§4)
10
Dense (1 x 1) MoE (64 x 1/4)
9
Common Crawl CE (↓)
10
Model type Dropout
8
None (disabled) 0.1 0.2 0.4
7 6
Model type
Dense (1 x 1) MoE (64 x 1/4)
9
Common Crawl CE (↓)
B.5
5
Max Grad Norm 0 (no clip) 0.2 1.0 2
8 7 6 5
1
2
4
8
16
Repetition Rate
32
1
64
2
(a) Dropout
16
32
64
10
Model type
Dense (1 x 1) MoE (64 x 1/4) Weight Decay
0.05 0.1 0.2 0.4
8 7 6
Model type
Dense (1 x 1) MoE (64 x 1/4)
9
Common Crawl CE (↓)
9
Common Crawl CE (↓)
8
(b) Gradient Norm Clipping
10
FOM Prob.
8
0.0 0.1 0.2 0.4
7 6 5
5 1
2
4
8
16
Repetition Rate
32
64
1
2
(c) Weight Decay
10
16
32
64
0.0 0.1 0.2 0.4
6 5
Model type
Dense (1 x 1) MoE (64 x 1/4)
9
Common Crawl CE (↓)
Expert Dropout
7
8
10
Dense (1 x 1) MoE (64 x 1/4)
8
4
Repetition Rate
(d) FFN Output Masking Model type
9
Common Crawl CE (↓)
4
Repetition Rate
EOM Prob.
8
0.0 0.1 0.2 0.4
7 6 5
1
2
4
8
16
Repetition Rate
32
64
1
(e) Expert Dropout
2
4
8
16
Repetition Rate
32
64
(f) Expert Output Masking
24
Data Scarcity and Model Sparsity
10
Model type
Dense (1 x 1) MoE (64 x 1/4)
Common Crawl CE (↓)
9
Router Jitter
8
0.0 0.1 0.2 0.4
7 6 5 1
2
4
8
16
Repetition Rate
32
64
(g) MoE Router Jitter
Figure 16: Dropout (a), FFN Output Masking (d), Expert Dropout (e), and Expert Output Masking (f) each reduces overfitting from data repetition (§4). Of the regularization methods studied in §4, these 4 dramatically decrease the response to data repetition. However, weight decay (b) and gradient norm clipping (c), as well as MoE router jitter (f), have minimal effect.
25
Data Scarcity and Model Sparsity
B.6
DATA F ILTERING (§3.3) Model type
Model type
Repetitions
% filtered
1x 2x 4x 8x 16x 32x
6
5
0
25
50
% filtered
75
Dense (1 x 1) MoE (64 x 1/4)
Common Crawl CE (↓)
Common Crawl CE (↓)
Dense (1 x 1) MoE (64 x 1/4)
0% 25% 50% 75% 100%
6
5
100
1
2
4
8
Repetition Rate
16
32
(a) Dolma Common Crawl Validation CE Loss
peS2o CE (↓)
6
5
0
25
50
% filtered
75
Model type Dense (1 x 1) MoE (64 x 1/4) % filtered 0% 25% 50% 75% 100%
7
6
peS2o CE (↓)
Model type Dense (1 x 1) MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x
7
5
100
1
2
4
8
Repetition Rate
16
32
(b) Dolma peS2o Validation CE Loss
Stack CE (↓)
8 7
6
5
Model type Dense (1 x 1) MoE (64 x 1/4) % filtered 0% 25% 50% 75% 100%
9 8
Stack CE (↓)
Model type Dense (1 x 1) MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x
9
7
6
5 0
25
50
% filtered
75
100
1
2
(c) Dolma Stack Validation CE Loss
Figure 17
26
4
8
Repetition Rate
16
32
Data Scarcity and Model Sparsity
B.7
I NTERPOLATING B ETWEEN A S INGLE D OMAIN AND A M IX
We consider an experimental setup similar to §3.3, but with interpolation between DCLM-baseline and a data mix. We use Dolma 1.7 (Soldaini et al., 2024b), which is similar to the OLMoE mix (§A.3; used in §3), but includes more sources. Dolma 1.7 and DCLM-baseline do not fall on a spectrum of quality, but rather of homogeneity. Unlike our experiments of §3.4, which are carefully controlled examinations of 2-domain mixes, the heterogeneity of Dolma 1.7 more closely resembles that of data mixes used for frontier LMs. Model type
Dense (1 x 1) MoE (64 x 1/4)
6
Repetitions
1x 2x 4x 8x 16x 32x
5
0
25
50
% filtered
75
Common Crawl CE (↓)
Common Crawl CE (↓)
Model type
Dense (1 x 1) MoE (64 x 1/4)
6
% filtered
0% 25% 50% 75% 100%
5
100
1
2
4
8
Repetition Rate
16
32
(a) Dolma Common Crawl Validation CE Loss
peS2o CE (↓)
5
Model type Dense (1 x 1) MoE (64 x 1/4) % filtered 0% 25% 50% 75% 100%
6
peS2o CE (↓)
Model type Dense (1 x 1) MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x
6
4
5
4 0
25
50
% filtered
75
100
1
2
4
8
Repetition Rate
16
32
(b) Dolma peS2o Validation CE Loss
8
Stack CE (↓)
7 6 5
4
Model type Dense (1 x 1) MoE (64 x 1/4) % filtered 0% 25% 50% 75% 100%
9 8 7
Stack CE (↓)
Model type Dense (1 x 1) MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x
9
6 5
4
0
25
50
% filtered
75
100
1
2
(c) Dolma Stack Validation CE Loss
Figure 18 27
4
8
Repetition Rate
16
32
Data Scarcity and Model Sparsity
M ECHANISTIC A NALYSES OF ROUTING (§5) MoE (16 x 1/4)
Top-1 routing stability
1.0
Model type MoE (16 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x
0.8 0.6 0.4 0.2 0.0
200
400
600
800
1000
Training step
1200
MoE (32 x 1/4)
1.0
Top-1 routing stability
B.8
0.8 0.6 0.4 0.2 0.0
1400
Model type MoE (32 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x 64x 128x 256x
200
400
600
(a) MoE (16 x 1/4)
0.6 0.4 0.2
1.0
Top-1 routing stability
Top-1 routing stability
Model type MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x 64x 128x 256x
0.8
0.0
400
600
800
1000
Training step
1200
Model type MoE (128 x 1/4) Repetitions 2x 4x 8x 16x 32x
0.6 0.4 0.2
1400
500
MoE (64 x 1/8)
0.6 0.4 0.2 600
800
1000
Training step
1200
1400
(e) MoE (64 x 1/8)
Model type MoE (64 x 1/2) Repetitions 1x 2x 4x 8x 16x 32x
0.8 0.6 0.4 0.2 0.0
400
MoE (64 x 1/2)
1.0
Top-1 routing stability
Top-1 routing stability
Model type MoE (64 x 1/8) Repetitions 1x 2x 4x 8x 16x 32x
200
1000
Training step
(d) MoE (128 x 1/4)
0.8
0.0
1400
MoE (128 x 1/4)
(c) MoE (64 x 1/4) 1.0
1200
0.8
0.0 200
1000
(b) MoE (32 x 1/4)
MoE (64 x 1/4)
1.0
800
Training step
500
1000
Training step
(f) MoE (16 x 1/2)
Figure 19: Routing ossifies early in training, exacerbated by repetition (§5.1). We show Top-1 routing stability, which we define as the fraction of held-out Common Crawl tokens that keep the same top-1 expert between consecutive checkpoints (200 steps apart), averaged over layers, for the 80M MoE models. At the beginning of training, consecutive checkpoint agreement is near chance, at roughly n1 for all MoE (n x g) configurations. However, routing stability rises rapidly for all settings, with fewer than 25% of tokens routed to a different top-1 expert when comparing step 400 to 600. Router ossification is slightly higher with fewer experts (lower n) or with higher granularity (larger g). Stability at each checkpoint also increases with repetition rate R, up to roughly R = 16, 32, where the router appears to destabilize.
28
0.91
Routing stability, end of training
1
2
4
8
16
Repetition Rate
32
64
0.94 0.93 0.92 0.91 1
2
4
8
16
Repetition Rate
32
64
Co-activation entropy (normalized)
0.92
0.005 0.003 0.002 0.001 1
2
4
8
16
Repetition Rate
32
0.94
0.002 1
2
4
8
16
Repetition Rate
32
MoE (16 x 1/4) MoE (32 x 1/4) MoE (64 x 1/4) MoE (128 x 1/4)
0.9 0.88
64
0.003
Model type
0.92
1
Co-activation entropy (normalized)
0.93
Expert knockout: median CE increase
0.94
0.01
Expert knockout: median CE increase
Routing stability, end of training
Data Scarcity and Model Sparsity
2
4
8
16
Repetition Rate
32
64
0.94 0.92 0.9 0.88 0.86 0.84
Model type
MoE (64 x 1/8) MoE (64 x 1/4) MoE (64 x 1/2)
0.82
64
1
2
4
8
16
Repetition Rate
32
64
0.94 1
2
4
8
Repetition Rate
16
32
Co-activation entropy (normalized)
0.945
Expert knockout: median CE increase
Routing stability, end of training
Figure 20: Routing stability, expert specialization, and expert-coactivation entropy rise slightly with repetition rate (§5). We show more MoE configurations for (a) end-of-training routing stability, (b) expert knockout effect, and (c) expert co-activation entropy. (a) End-of-training routing stability increases slightly with repetition rate R. Higher expert count naturally results in lower stability, as there are more experts to choose from. Fewer active, but larger experts slightly decreases stability. (b) Expert specialization increases slightly with R. Higher expert count results in lower specialization. Varying active expert count along with expert size has no clear effect. (c) Co-activation entropy of expert pairs (normalized by maximum) rises steadily with R towards uniformly distributed pairings. Entropy is higher with fewer total experts, and with a larger active number of smaller experts.
0.01
0.005 0.003 0.002 1
2
4
8
Repetition Rate
16
32
Model type
MoE (32 x 1/4) MoE (64 x 1/4)
0.945 0.94 0.935 1
2
4
8
Repetition Rate
16
32
Figure 21: Routing stability, expert specialization, and expert-coactivation entropy rise slightly with repetition rate at 200M active parameters (§5). We show 200M MoE (32 x 1/4) and MoE (64 x 1/4) configurations for (a) end-of-training routing stability, (b) expert knockout effect, and (c) expert co-activation entropy. (a) End-of-training routing stability increases slightly with repetition rate R. Higher expert count naturally results in lower stability, as there are more experts to choose from. (b) Expert specialization increases slightly with R. Higher expert count results in lower specialization. (c) Co-activation entropy of expert pairs (normalized by maximum) rises steadily with R towards uniformly distributed pairings. Entropy is higher with fewer total experts.
29
0.92 0.9 0.88 0.86 1
2
4
8
Repetition Rate
16
32
0.01
Co-activation entropy (normalized)
0.94
Expert knockout: median CE increase
Routing stability, end of training
Data Scarcity and Model Sparsity
0.005 0.003 0.002 0.001 0.0005 0.0003
1
2
4
8
Repetition Rate
16
32
0.95
Model type
0.94
MoE (32 x 1/4) MoE (64 x 1/4)
0.93
Dropout
0 0.1 0.4
0.92 0.91 0.9
1
2
4
8
Repetition Rate
16
32
Figure 22: Dropout increases routing stability and expert-coactivation entropy, but decreases expert specialization (§5). We show, for various dropout settings on our 200M MoE (64 x 1/4) configurations: (a) end-of-training routing stability, (b) expert knockout effect, and (c) expert co-activation entropy. (a) End-of-training routing stability increases with higher dropout probability, and still rises with repetition rate R. (b) Expert specialization decreases with higher dropout probability, but still increases slightly with R. (c) Co-activation entropy of expert pairs (normalized by maximum) increases with dropout probability, and still rises steadily with R towards uniformly distributed pairings.
30
Data Scarcity and Model Sparsity
B.9
A DDITIONAL L ANGUAGE M ODELING TASKS
We show results for the additional held-out language modeling tasks of Appendix A.4 on the models from §3, with the extended settings of Appendix B.2. 20
10 9 8 7 6
Model type Dense (1 x 1) MoE (64 x 1/4)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6
C4 CE (↓)
C4 CE (↓)
C4 CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
5
5
5
4
4
20
23
26
29 212 215 218 221
3 1
Repetition Rate
10 9 8 7 6
4
16
64
Repetition Rate
256
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(a) C4 (CE Loss ↓) 20
10 9 8 7 6
Model type Dense (1 x 1) MoE (64 x 1/4)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
Books CE (↓)
Books CE (↓)
Books CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6 4
20
23
26
29 212 215 218 221
3 1
Repetition Rate
5 4
5
5
10 9 8 7 6
4
16
64
Repetition Rate
256
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(b) Dolma books (CE Loss ↓) 20
10 9 8 7 6 5
Model type
Model type
Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6 5
Dense (1 x 1) MoE (64 x 1/4)
Common Crawl CE (↓)
Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
Common Crawl CE (↓)
Common Crawl CE (↓)
Model type
4 20
23
26
29 212 215 218 221
1
Repetition Rate
4
16
64
Repetition Rate
256
10 9 8 7 6 5 4 3
1024
1
2
4
8 16 32 64 128 256
Repetition Rate
(c) Dolma common-crawl (CE Loss ↓) Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6
10 9 8 7 6
peS2o CE (↓)
peS2o CE (↓)
peS2o CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
5
5
4
5
4
3
4
3
20
23
26
29 212 215 218 221
Repetition Rate
1
4
16
64
Repetition Rate
256
1024
(d) Dolma pes2o (CE Loss ↓)
31
Model type Dense (1 x 1) MoE (64 x 1/4)
10 9 8 7 6
1
2
4
8
16 32 64 128 256
Repetition Rate
Data Scarcity and Model Sparsity
20
10 9 8 7 6
Model type Dense (1 x 1) MoE (64 x 1/4)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
Reddit CE (↓)
Reddit CE (↓)
Reddit CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6
5
5
5
4
4
20
23
26
29 212 215 218 221
1
Repetition Rate
10 9 8 7 6
4
16
64
Repetition Rate
256
3
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(e) Dolma reddit (CE Loss ↓)
10 9 8 7 6 5
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10
5
Model type Dense (1 x 1) MoE (64 x 1/4)
10
Stack CE (↓)
Stack CE (↓)
Stack CE (↓)
20
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
20
5
3 2
3
4
20
23
26
29 212 215 218 221
1
Repetition Rate
4
16
64
Repetition Rate
256
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(f) Dolma stack (CE Loss ↓)
10 9 8 7 6 5
Wikipedia CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
Wikipedia CE (↓)
Wikipedia CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6 5 4
20
23
26
29 212 215 218 221
10 9 8 7 6 5 4 3
1
Repetition Rate
Model type Dense (1 x 1) MoE (64 x 1/4)
4
16
64
Repetition Rate
256
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(g) Dolma wiki (CE Loss ↓)
10 9 8 7 6
6
5
5
4
20
23
26
29 212 215 218 221
Repetition Rate
Model type Dense (1 x 1) MoE (64 x 1/4)
ICE CE (↓)
10 9 8 7
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
20
ICE CE (↓)
ICE CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6 5 4 3
1
4
16
64
Repetition Rate
256
1024
(h) ICE (CE Loss ↓)
32
1
2
4
8
16 32 64 128 256
Repetition Rate
Data Scarcity and Model Sparsity
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7
10 9 8 7 6
Pile CE (↓)
Pile CE (↓)
Pile CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6 5
6
5
4
5
4
3
4
3
20
23
26
29 212 215 218 221
Repetition Rate
1
4
16
64
Repetition Rate
256
Model type Dense (1 x 1) MoE (64 x 1/4)
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(i) Pile (CE Loss ↓) 20
10 9 8 7 6
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
Model type Dense (1 x 1) MoE (64 x 1/4)
S2ORC CE (↓)
S2ORC CE (↓)
S2ORC CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6 5
5
23
26
29 212 215 218 221
3 1
Repetition Rate
5 4
4
20
10 9 8 7 6
4
16
64
Repetition Rate
256
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(j) S2ORC (CE Loss ↓) 20
10 9 8 7 6
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 9 8 7 6 5
20
23
26
29
212
215
Repetition Rate
218
221
10 9 8 7 6 5 4 3
4
5
Model type Dense (1 x 1) MoE (64 x 1/4)
WikiText-103 CE (↓)
WikiText-103 CE (↓)
WikiText-103 CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
1
4
16
64
Repetition Rate
256
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(k) Wikitext (CE Loss ↓)
Figure 23: Held-out language modeling loss largely follows the same patterns. For 80M, 200M, and 1B active parameter models trained on OLMoE mix, we show held-out language modeling loss on a wide variety of domains. Performance depends on the exact domain, but largely follows the trends shown in Figure 1.
33
Data Scarcity and Model Sparsity
B.10
D OWNSTREAM TASKS
We show results for the downstream tasks of Appendix A.4 on the models from §3, with the extended settings of Appendix B.2. We include CE Loss on all tasks, as well as accuracy on Hellaswag. Accuracy on all other tasks remains near-chance, even at 1B scale.
6 5
BoolQ CE (↓)
BoolQ CE (↓)
10
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
4 3
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
5
Model type Dense (1 x 1) MoE (64 x 1/4)
5
BoolQ CE (↓)
8 7
3 2
2
20
23
26
29
3 2
1
212 215 218 221
1
Repetition Rate
4
16
64
Repetition Rate
256
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(a) BoolQ (CE Loss ↓) Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
3
HellaSwag CE (↓)
HellaSwag CE (↓)
3
2
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
4
2
1 0.9 0.8 0.7
1 0.9
20
23
26
29 212 215 218 221
1
Repetition Rate
4
16
64
Repetition Rate
256
Model type Dense (1 x 1) MoE (64 x 1/4)
3
HellaSwag CE (↓)
4
2
1 0.9 0.8 0.7 0.6
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(b) Hellaswag (CE Loss ↓) 20
10
5 3
Model type
Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10 5 3 2
2 20
23
26
29 212 215 218 221
1
Repetition Rate
4
16
64
Repetition Rate
256
MMLU humanities CE (↓)
Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
MMLU humanities CE (↓)
MMLU humanities CE (↓)
Model type
9 8 7 6 5
Model type
Dense (1 x 1) MoE (64 x 1/4)
4 3 2
1024
1
2
4
8 16 32 64 128 256
Repetition Rate
20
10
5 3
10
5 3
Model type Dense (1 x 1) MoE (64 x 1/4)
7 6
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
MMLU other CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
MMLU other CE (↓)
MMLU other CE (↓)
(c) MMLU Humanities (CE Loss ↓)
5 4 3 2
2
2
20
23
26
29 212 215 218 221
Repetition Rate
1
4
16
64
Repetition Rate
256
1024
(d) MMLU Other (CE Loss ↓)
34
1
2
4
8
16 32 64 128 256
Repetition Rate
10
5 3
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10
5 3
Model type Dense (1 x 1) MoE (64 x 1/4)
7 6
MMLU social sci CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
MMLU social sci CE (↓)
MMLU social sci CE (↓)
Data Scarcity and Model Sparsity
5 4 3 2
2
2
20
23
26
29 212 215 218 221
1
Repetition Rate
4
16
64
Repetition Rate
256
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
10
5 3
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
10
5 3
Model type Dense (1 x 1) MoE (64 x 1/4)
7 6
MMLU STEM CE (↓)
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
MMLU STEM CE (↓)
MMLU STEM CE (↓)
(e) MMLU Social Sciences (CE Loss ↓)
5 4 3 2
2
2
20
23
26
29 212 215 218 221
1
Repetition Rate
4
16
64
Repetition Rate
256
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(f) MMLU Stem (CE Loss ↓)
0.26 0.2575 0.255 0.2525 0.25 0.2475
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
0.29
HellaSwag acc (↑)
HellaSwag acc (↑)
0.3
Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)
0.28 0.27 0.26
Model type Dense (1 x 1) MoE (64 x 1/4)
0.5
HellaSwag acc (↑)
0.265 0.2625
0.4
0.3
0.25
20 23 26 29 212 215 218 221
Repetition Rate
1
4
16
64
Repetition Rate
256
1024
1
2
4
8
16 32 64 128 256
Repetition Rate
(g) Hellaswag (Accuracy ↑)
Figure 24: Downstream task performance closely follows validation loss. For 80M, 200M, and 1B active parameter models, downstream task loss is subject to random noise, but loosely follows trends of validation loss.
35