Conceptio › Archive › arXiv CS
arXiv CSopen access

Data Scarcity and Model Sparsity: Mixtures-of-Experts Overfit More to Repeated Data

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
neural-networks
machine learning, deep learning, neural networks

Data Scarcity and Model Sparsity

DATA S CARCITY AND M ODEL S PARSITY: M IXTURES - OF -E XPERTS OVERFIT M ORE TO R EPEATED DATA Atindra Jha∗1 , Margaret Li∗2 , Jure Leskovec1 , Percy Liang1 , and Luke Zettlemoyer2 1

2

Stanford University Paul G. Allen School of Computer Science, University of Washington

arXiv:2609.11917v1 [cs.LG] 10 Sep 2026

A BSTRACT As the supply of human-written text is exhausted, it has become standard practice to repeat language model training data. Prior work has studied data repetition for densely activated Transformers, but the effects of data repetition remains largely unexplored for recently dominant sparse architectures such as Mixture-of-Experts (MoE), despite their increased compute efficiency. We vary data repetition rates across single- and multi-domain data mixes, and across MoE settings, including expert count and granularity. We consistently find, for models ranging from 80M to 1B active (8.5B total) parameters, that MoEs degrade more rapidly under data repetition. This effect increases with sparsity, dictated by total rather than active parameters. While 80M dense models can repeat data over 8× with minimal degradation, MoEs instead begin to suffer at 4 repetitions, and deteriorate rapidly, ceding their performance benefits in all-unique data settings to dramatically underperform dense models after 32 repetitions. We experiment with existing regularization methods as a potential remedy. We find that some methods, such as dropout, can mitigate overfitting. In particular, with strong masking-based regularization, MoEs are able to outperform dense models even when data is repeated more than 64 times. However, no method fully matches the performance of all-unique training data. Finally, we analyze internal mechanisms correlated with MoE overfitting in high repetition regimes, and find that MoE routing universally stabilizes early in training, and that expert specialization correlates with overfitting to repeated data. In sum, our work addresses the adverse interactions between sparsity and data repetition: we present evidence for the core mechanisms of overfitting and its potential remediation, and suggest promising avenues for future methods to reduce over-specialization in model parameters by disrupting memorization patterns.

1

I NTRODUCTION

As language model training begins to exhaust even massive-scale web crawls, the amount of unique training data has become a constraining factor for language model training, in addition to compute. It is now common practice to repeat some or all training data, despite the known tendency of language models to overfit to data under high repetition rates (Muennighoff et al., 2023). Simultaneously, the compute costs of large-scale LLM training have driven the adoption of relatively compute-efficient Mixture-of-Experts models (Shazeer et al., 2017; Fedus et al., 2022; Muennighoff et al., 2025). These models achieve compute efficiency via sparsity, the ratio of total to active parameters. When training only with unique data, increased sparsity is known to consistently improve training efficiency, albeit at higher communication costs. However, the interaction between sparsity and data repetition remains relatively unexplored. Sparsity decouples total parameters from active parameters, and data repetition decouples total data tokens from unique data tokens. Thus, both sparsity and data repetition introduce axes of variation to modern scaling laws, which prescribe simple token-to-parameter ratios without disambiguating between unique or total training tokens over active or total parameters. MoE performance also depends on architectural details such as expert size and count. Further, not all data tokens are equal, with significant research devoted to filtering, deduplication, and data mix domain makeup (Li et al., 2024). Each axis of variation is entangled with all others, but prior work primarily investigates these axes in isolation: the data-constrained scaling laws of Muennighoff et al. (2023) are fit to dense models on a single corpus (C4), and Xue et al. (2023) consider a single MoE configuration to posit that multi-epoch degradation increases with total parameters. To investigate these interactions more fully, we conduct an in-depth grid sweep over axes of variation. Across three active-parameter scales (80M, 200M, 1B), we compute-match by training on a fixed total data budget for each activeparameter scale. We vary the unique tokens (equivalently, the repetition rate R), comparing dense transformers to MoE architectures with varied expert count and granularity. We consider a variety of domains (web crawl, code, scientific, and encyclopedic text), both as single-domain corpora and as components of data mixes with different per-domain repetition rates. Our results indicate that MoEs overfit more to repeated data in comparison to dense models; 80M MoEs clearly degrade at 4× data repetition (compared to 8× for dense models), and cede their performance advantage ∗

equal contribution; correspondence to [email protected], [email protected]

1

Data Scarcity and Model Sparsity

at 32× data repetition. Further, overfitting becomes more catastrophic with higher MoE sparsity, and its patterns depend on total rather than active parameters. Surprisingly, this phenomenon follows similar patterns across datatoken-per-parameter ratios and across our varied data domains, only slightly decreased by quality filtering. However, when repeated data is mixed into a non-repeated dataset, the unique tokens may have a regularizing effect. We also find that some existing regularization methods (dropout, output masking) reduce the impact of high data repetition rates, especially in MoEs. At sufficiently high dropout probability, MoEs can outperform dense models even at 64× data repetition. Finally, we perform mechanistic analyses and show that MoE routers stabilize their decisions early in training, which suggests that expert parameters update on a small and stationary subset of the total tokens, and we measure the resulting impact of data repetition on expert specialization. Overall, our contributions are as follows: • We show that the benefit of sparsity is conditional on the unique data budget: Across data domains and mixes, MoEs outperform dense models on unique data, but underperform at high data repetition rates. Data quality filtering has only minor impact on repetition effects: some models overfit more to unfiltered data (§3.1-3.3). • We study domain-specific repetition rates in data mixes. Performance degradation is confined to repeated domains, and unique tokens from another domain may mitigate overfitting of repeated domains (§3.4). • We apply regularization techniques and demonstrate that some techniques are ineffective, but dropout and output masking can mitigate overfitting from data repetition (§4). • We analyse MoE expert activations and outputs. Our evidence shows that MoE router decisions stabilize early, which exacerbates overfitting as expert parameters over-specialize, updating on a small and near-stationary subset of tokens (§5).

2

BACKGROUND

2.1

M IXTURE OF E XPERTS L ANGUAGE M ODELS

In Mixture-of-Experts Transformer LMs, the Feed-Forward (FFN) of each layer is replaced by n parallel FFN experts E1 , . . . , En and a router that selects a subset of the experts to apply to each token. For each hidden token representation h, the router produces a score si (h) for each expert i and returns a weighted sum of the top-k experts: X y = gi (h) Ei (h), i∈TopK(s(h))

where gi are the router’s softmax scores of the selected experts. Because not all parameters are active for each token, MoEs decouple total from active parameters. Recent MoEs employ fine-grained experts (Dai et al., 2024), where the granularity is defined as the ratio of expert-FFN to dense-FFN dimensions. For example, if an MoE has expert granularity g = 12 , then the experts have 1 1 · dense FFN dimension = · 4 · hidden dimension = 2 · hidden dimension. 2 2 MoE architectures may vary the expert granularity g, the total count of experts n, and the active count k. These design choice axes complicate the study of MoEs, as active parameter count and FLOPS-per-token cost change with each configuration. For fair comparison with dense models, it is common to match active parameters by setting g = k1 . expert dimension =

2.2

DATA R EPETITION

In a data-constrained scenario, it is common to repeat some or all of the training data. A simple approach might iterate over the entire training data corpus of U unique tokens over multiple passes, or epochs, until the desired total token budget T is reached. This requires a repetition rate of R = T /U . In other cases, training data may be defined as a mix of various data domains combined in fixed percentages. If only a subset of the domains are data-constrained, it is common to set domain-specific repetition rates Ri for each domain i (Soldaini et al., 2024a; Muennighoff et al., 2025). 2.3

R EGULARIZATION M ETHODS

Overfitting describes memorization of training data at the cost of generalization to unseen data. Known remedies trade off training fit to recover generalization: Dropout (Srivastava et al., 2014) randomly drops units during training, which prevents them from co-adapting to form complex memorization patterns. Weight decay, in the decoupled form used by AdamW (Loshchilov & Hutter, 2019), shrinks parameters towards zero independently of the gradient. Gradient norm clipping (Pascanu et al., 2013) rescales gradients whose norm exceeds a threshold, stabilizing the effect of any single datapoint. Some regularization methods specifically target components of MoEs: Router jitter injects multiplicative noise into the routing computation, and expert dropout applies a separate, typically larger dropout rate inside experts (Fedus et al., 2022; Zoph et al., 2022a). We evaluate regularization methods under data repetition in §4. 2

Data Scarcity and Model Sparsity

3

E XPERIMENTS

Models. We train compute-matched densely- and sparsely-activated (MoE) Transformer LMs in the style of Muennighoff et al. (2025) . Specifically, we train models with 80 million, 200 million, and 1 billion active parameters, denoted as 80M, 200M, and 1B, respectively. We vary our MoE model configurations to study the effect of total expert count and expert granularity. Our MoE models have n ∈ {8, 16, 32, 64, 128} total experts with granularity 1 1 g ∈ { 21 , 14 , 81 , 16 , 32 }. To match active parameters, we set the top-k activation to k = g1 . This results in sparsity s ∈ {2, 4, 8, 16, 32}. We denote our models MoE (n x g), e.g., MoE (64 x 1/4) represents 4 active experts out of 64 total with granularity 1/4. Additional model architecture details are in Appendix A.1. Data Repetition. Following common practice from Hoffmann et al. (2022), our models are trained with total training data tokens T ≈ 20 · Na , where Na denotes the number of active parameters. We only vary the tokens-per-activeparameter T /Na ratio in §3.1 to study its interaction with sparsity and data repetition. Under a fixed total data T , we vary the number of unique tokens U . Each model is trained by iterating through its allotted U unique tokens R = T /U times. We consider subsets of R ∈ {1, 2, 4, 8, 16, 32, 64, 128, 256, 512, 1024}. For any unique-token budget U , we construct the training set DU by taking the first U tokens from a fixed random permutation of the full dataset D. We use the same permutation across all experiments, ensuring that training sets are nested: for any U1 ≤ U2 , DU1 ⊆ DU2 Additional data repetition details are in Appendix A.5. Training Data. We train on constituent domains from the OLMoE (Muennighoff et al., 2025) data mix, both individually and in custom mixes: web crawl data (DCLM; Li et al. (2024)), code (Starcoder; Li et al. (2023)), scientific text (peS2o; Soldaini & Lo (2023)), and encyclopedic text (Wikipedia; Soldaini et al. (2024a)). See Appendix A.3. Evaluation. We report Cross-Entropy Loss (CE Loss) on various held-out data, including web crawl, code, scientific text, and encyclopedic text. We also measure CE Loss and accuracy on downstream tasks. Results on downstream tasks are in Appendix B.10. We show random seed sensitivity in Appendix B.1. Additional details in Appendix A.4. 3.1

DATA R EPETITION E FFECTS ACROSS M O E A RCHITECTURES

We train a variety of dense and MoE Transformer LMs on the OLMoE data mix to investigate the effects of data repetition. Specifically, we fix the total train token budget T = 20 · Na , but vary the repetition rate R, so that unique tokens U decreases as R increases to satisfy R · U = T for fixed T . MoEs degrade more than dense LMs under data repetition. Figure 1 shows the effect of data repetition on validation loss. Across model scales, dense Transformer performance degrades at high data repetition rates, with validation loss rising slightly at R = 8 and sharply at R > 64. MoEs respond more dramatically to data repetition, with a noticeable performance impact at R = 4 and a sharper increase at higher R. Though MoEs outperform dense models at R ≤ 16, the catastrophic effect of data repetition leads to a reversal at R = 32, where dense models outperform MoEs. Other held-out LM tasks (Appendix B.9) and downstream tasks (Appendix B.10) exhibit similar trends. At extremely high repetition rates, validation loss lowers again. Results in Figure 1 demonstrate a least optimal data repetition rate; validation loss rises with R until this point, then falls. We also consider much higher R ∈ {215 , 220 }, and find that validation loss rises yet again (Appendix Figure 11). Training loss falls to zero at high repetition rate. As shown in Figure 2, training loss decreases with increased repetition rate, falling under 1E-2 at R = 512 for 80M dense, and R = 160 for 200M dense models. Training loss decrease is a reflection of validation loss increase, which agrees with other indications of overfitting via training data memorization. Additional model scales in Appendix Figure 12. Sparsity increases overfitting; data repetition effects depend on total parameters. We vary the number of total experts and the expert granularity in Figure 4. In both cases, increased total parameters results in a sharper response to data repetition. As the active parameters remain fixed throughout these experiments, we hypothesize that this phenomenon is related to the unique tokens to total parameters ratio. Thus, we compare the 200M dense models to the 80M MoE (32 x 1/4) and MoE (64 x 1/4) models, which have 158M and 244M total parameters, respectively. In Figure 1, the data repetition effects of the 200M model appear to lie between those of the MoE (32 x 1/4) and MoE (64 x 1/4) models at 80M active parameters. Data repetition effects do not depend on total data budget. We train 80M active parameter models with 4× more data, resulting in a total-tokens-to-active-parameter ratio T /Na = 80. The resulting trends (Figure 5) are very similar to those with T /Na = 20 (Figure 1). Thus, the data repetition overfitting response may change only slowly with the total token budget, but instead depend most directly on total parameters. 3

Data Scarcity and Model Sparsity

4

16

64

Repetition Rate

1B

Common Crawl CE (↓)

10 9 8 7 6 5 1

200M

20

Common Crawl CE (↓)

Common Crawl CE (↓)

80M

10 9 8 7 6 5 4

256 1024

1

4

16

64

Repetition Rate

256 1024

10 9 8 7 6 5 4

Model type

Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

3

1

4

16

64

Repetition Rate

256

Figure 1: Across active parameter scales, data repetition rates over 8× result in increasingly severe overfitting. Sparser models overfit more (§3.1). At 80M, 200M, and 1B active parameters, we fix the total data budget T = 20 · Na , and vary the data repetition rate R via different sized unique token sets. As R increases, models increasingly overfit. Sparsity exacerbates overfitting behavior. Larger and sparser models overfit more at lower R. Dense (1 x 1)

10

Train CE (↓)

Common Crawl CE (↓)

64x 128x

1

MoE (32 x 1/4)

10

64x

1

256x

0.01

512x

0.01

0.001

1024x

0.001

0

500

1000

Train Step

1500

Dense (1 x 1)

10 9 8 7 6 5

64x 32x 16x

0

500

1000

Train Step

1

256x 512x 1024x

0

500

1000

Train Step

1500

0.001

1500

64x

32x 16x

0

500

1000

Train Step

256x 512x 1024x

0

500

1000

Train Step

1500

1500

128x 256x 512x 1024x 64x

10 9 8 7 6 5

32x 16x

0

500

1000

Train Step

1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x

Repetitions

MoE (64 x 1/4) 256x 128x 512x 1024x

10 9 8 7 6 5

128x

0.01

MoE (32 x 1/4) 1024x

32x 64x

0.1

512x 256x

128x

Repetitions

128x

0.1

0.1

MoE (64 x 1/4)

10

1500

1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x

Figure 2: At higher repetition rates, training loss falls to 0 as models overfit to the repeated data (§3.1). Early in training, train (above) and validation (below) loss fall together, but, at high repetition rate R, train loss falls rapidly to 0, indicating memorization of training data, while validation loss rises. Additional model scales in Appendix Figure 12.

3.2

DATA R EPETITION E FFECTS ACROSS DATA D OMAINS

We repeat the experimental setup in §3.1 with individual data domains from the OLMoE data mix: DCLM (web crawl), peS2o (academic), StarCoder (code), and Wikipedia (encyclopedic) text. Our goal is to understand variations in response to repetition across varied data domains. We focus on R ∈ {1, 2, 4, 8, 16, 32}. Patterns are largely consistent across single-domain experiments. Results in Figure 3 indicate that overfitting patterns from data repetition are robust across domains, despite the diversity of data. Code, encyclopedic, web crawl, and academic text are semantically very distinct, yet all four exhibit similar patterns, for example, that MoE models begin to underperform dense models at R ∈ [16, 32]. 3.3

DATA R EPETITION E FFECTS U NDER DATA F ILTERS

We also measure the impact of quality filtering on data repetition. It is common practice to preprocess raw web crawls to deduplicate, extract text from HTML, and remove data deemed low quality by pre-defined filters, often significantly reducing the available tokens. DCLM- BASELINE retains 2.4% of the raw 280T-token DCLM- POOL (Li et al., 2024). Under data constraints, filtering may remove lower quality tokens, but also further reduces the quantity of unique tokens. We study this tradeoff between token quality and repetition rate by considering 5 data settings, which mix 4

Data Scarcity and Model Sparsity

DCLM

StarCoder

Stack CE (↓)

Common Crawl CE (↓)

6

5

3.8 3.6 3.4 3.2 3 2.8 2.6

1

2

4

8

Repetition Rate

16

32

1

2

peS2o

8

16

32

Model type

Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

Wikipedia

5

Wikipedia CE (↓)

4.2

peS2o CE (↓)

4

Repetition Rate

4 3.8 3.6

4

3.4 1

2

4

8

Repetition Rate

16

32

1

2

4

8

Repetition Rate

16

32

Figure 3: Across all train domains, sparse MoE models overfit more, ceding their benefit over dense models as repetition rates rise (§3.2). We train dense and MoE models on a single domain. Our experiments span web crawl (DCLM), code (StarCoder), academic (peS2o), and encyclopedic (Wikipedia) text. Trends are near-identical across all data domains: models underperform as repetition rates rise, with MoEs overfitting more rapidly.

DCLM- BASELINE with DCLM- POOL. We interpolate between all-unfiltered and all-filtered with 5 settings consisting of the following DCLM (baseline %, pool %): (0%, 100%), (25%, 75%), (50%, 50%), (75%, 25%), (100%, 0%). Filtering affects data repetition in some dense models. Training on a higher percentage of unfiltered DCLMPOOL data results in consistently worse performance (Figure 6). Despite the distribution shift that arises from stringent filtering, all MoEs trained on any mixture of DCLM- POOL and - BASELINE exhibit similar decline with data repetition. However, dense models trained exclusively on at least 50% unfiltered data may deteriorate more rapidly with R. As DCLM- BASELINE filtering reduces DCLM- POOL by a factor of over 40×, we compare R = 1 on 100% raw DCLMPOOL against R = 32 on 100% DCLM- BASELINE. In our setting, training on all unique unfiltered data is superior to repeating strictly filtered data, but an intermediate quality filter and repetition rate is likely superior to either extreme. 3.4

DATA R EPETITION AT M IXED R ATES W ITHIN D IVERSE DATA M IXES

In practice, LMs are trained on data mixes composed of diverse domains at varying levels of repetition. For example, encyclopedic text is often more constrained than web crawl data, and is therefore repeated at a higher rate. For example, GPT-3 was trained for 3.4 epochs over its Wikipedia component, but less than 1 epoch over Common Crawl (Brown et al., 2020). To investigate the effects of varied repetition rates within a data mix, we consider a simple setting in which the total data budget T is split between two domains. The first domain is DCLM, which is never repeated. The second is peS2o or StarCoder, which is repeated with R ∈ {1, 2, 4, 8, 16, 32}. We divide the total data budget T between the 2 domains (DCLM, peS2o) or (DCLM, StarCoder) at the following proportions: (100%, 0%), (90%, 10%), (50%, 50%), or (0%, 100%). Only the repeated domain’s unique pool shrinks as R grows. As a concrete example, the (50%, 50%) DCLM-StarCoder setting at R = 8 has T = 1.6B total tokens, consisting of 1.6B ∗ 0.5 = 0.8B unique DCLM tokens and 1.6B ∗ 0.5/8 = 0.1B unique StarCoder tokens repeated 8 times. 5

Data Scarcity and Model Sparsity

Model type

Dense (1 x 1) MoE (8 x 1/4) MoE (16 x 1/4) MoE (32 x 1/4) MoE (64 x 1/4) MoE (128 x 1/4) MoE (256 x 1/4)

6

5

1

2

4

8

Repetition Rate

16

Model type

7

Common Crawl CE (↓)

Common Crawl CE (↓)

7

Dense (1 x 1) MoE (64 x 1/32) MoE (64 x 1/16) MoE (64 x 1/8) MoE (64 x 1/4) MoE (64 x 1/2)

6

5

32

1

2

(a) Varied Total Expert Count (n)

4

8

Repetition Rate

16

32

(b) Varied Expert Size / Granularity (g)

Common Crawl CE (↓)

Figure 4: In MoEs, higher sparsity results in a greater reaction to data repetition (§3.1). We consider two settings: (a) fixed expert granularity, increased sparsity via greater total expert count; (b) fixed total expert count, increased sparsity via larger experts. Overfitting increases with total parameters, corresponding to (a) more total experts and (b) larger experts. Lines across (a) and (b) with the same color have identical total parameter count. 8 Model type Dense (1 x 1) MoE (64 x 1/4)

7 6 5

4 1

2

4

8

16

Repetition Rate

32

64

Figure 5: Data repetition effects are consistent across data-to-parameter ratios (§3.1). We increase total tokens per active parameter from 20 to 80, and observe no difference in data repetition overfitting between the two token-to-active-parameter ratios.

Mixing repeated StarCoder with all-unique DCLM does not impact repetition effects. In Figure 7a, we find that, regardless of the proportion of T assigned to StarCoder, code validation loss rises according to the same pattern in §3.1-3.2. Web crawl validation loss only rises with the repetition rate of StarCoder if it occupies at least half of T .

Mixing repeated peS2o with all-unique DCLM may have a regularizing effect. In Figure 7b, academic text validation loss increases with the peS2o repetition rate, but the effect is substantially dampened as peS2o occupies a smaller proportion of the total training data. As with the DCLM-StarCoder experiments above, web crawl validation loss suffers from peS2o repetition only when the peS2o total data proportion is high. This suggests that the unrepeated DCLM data component may have a regularizing effect on the peS2o data repetition.

Mixing repeated data with unrepeated data of a semantically similar domain may reduce data repetition effects. Academic text, as represented by peS2o, has higher semantic similarity to web text than code, as represented by StarCoder. Our results suggest a promising possibility: even at very high repetition rates, repetition effects might be reduced by mixing the repeated data domain with an non-repeated or less-repeated, semantically similar data domain. 6

Data Scarcity and Model Sparsity

Model type

Model type

Repetitions

% filtered

1x 2x 4x 8x 16x 32x

6

5

0

25

50

% filtered

75

100

Dense (1 x 1) MoE (64 x 1/4)

Common Crawl CE (↓)

Common Crawl CE (↓)

Dense (1 x 1) MoE (64 x 1/4)

0% 25% 50% 75% 100%

6

5

1

2

4

8

Repetition Rate

16

32

Figure 6: Repetition effects are consistent across data quality (§3.3). Validation CE Loss for 80M dense and MoE (64 x 1/4) on the DCLM- POOL and - BASELINE interpolation, where % filtered indicates the proportion of the data mix dedicated to the heavily filtered DCLM- BASELINE data. Left: CE against the DCLM- BASELINE %. A higher percentage of filtered data results in better performance. As in §3.1-3.2, MoE performance deteriorates more rapidly than dense at higher R, regardless of data. Right: CE Loss against repetition rate. Trends are largely similar across data quality and model architecture, though dense models trained on higher proportions of unfiltered data (0-50% DCLMBASELINE) may degrade more rapidly with R.

4

R EGULARIZATION T ECHNIQUES FOR DATA R EPETITION

To better understand the overfitting observed above, we apply known regularization methods with two motivations: to further probe the mechanisms responsible for performance decline under data repetition, and to take initial steps towards solutions. Dropout We use element-wise residual dropout (Srivastava et al., 2014) on the output of each sub-layer with probability p ∈ {0.0, 0.1, 0.2, 0.4}. High dropout probability hurts performance at low data repetition, but dramatically reduces the impact of extreme data repetition for all architectures (Figure 8a). Gradient Norm Clipping The gradient clipping threshold upper bounds the magnitude of each batch update, and is often used to decrease instability (Pascanu et al., 2013). We sweep this threshold over {0.2, 1.0, 2.0, None}, where None indicates no clipping. All four settings result in similar performance, within variance (Appendix Figure 16). Weight Decay We sweep decoupled AdamW weight decay over λ ∈ {0.05, 0.1, 0.2, 0.4}. Performance differences, though well-ordered, are small and remain within noise thresholds (Appendix Figure 16). FFN Output Masking FFN output masking (FOM) zeroes, with probability p, the entire dense FFN or MoE output for a token without rescaling. FOM is similar to a coarse-grained dropout applied only to FFNs. We consider p ∈ {0.0, 0.1, 0.2, 0.4} and observe that, at low repetition rate, it incurs a smaller penalty to performance than residual dropout. At high repetition rate, it has a similar regularizing effect to dropout (Figure 8b). Expert Dropout Expert dropout (Fedus et al. (2022); Zoph et al. (2022a)) applies dropout only to the expert hidden activation. We consider expert dropout rate in {0.0, 0.1, 0.2, 0.4}, and compare to the dense analogue, which applies dropout only to the FFN. Expert Dropout, like FOM, acts only on the FFN, and indeed behaves similarly to FOM (Figure 8c). Expert Output Masking Expert output masking (EOM) independently zeroes each (token, selected expert) output with probability p during training. No rescaling is applied. EOM thus operates only on the FFNs, like FOM and Expert Dropout, with an intermediate granularity. We consider p ∈ {0.0, 0.1, 0.2, 0.4} and observe that EOM has a similar effect to both FOM and Expert Dropout (Figure 8d). Router Jitter Router jitter perturbs routing, multiplying the router’s input by elementwise uniform noise on [1−ϵ, 1+ ϵ] during training. Over ϵ ∈ {0.0, 0.1, 0.2, 0.4}, we observe no clear impact on performance (Appendix Figure 16). 7

Data Scarcity and Model Sparsity

Stack CE (↓)

DCLM (R=1)

7

5

Stack CE (↓)

Common Crawl CE (↓)

Common Crawl CE (↓)

6

Dense (1 x 1) MoE (64 x 1/4)

% StarCoder

10% StarCoder 50% StarCoder 100% StarCoder

3

5 DCLM (R=1) 1

Model type

4

2

4

8

16

Repetition Rate

32

1

2

4

8

Repetition Rate

16

32

(a) DCLM + StarCoder

Common Crawl CE (↓)

peS2o CE (↓)

4.6 DCLM (R=1)

7

peS2o CE (↓)

Common Crawl CE (↓)

4.4 6

4.2

Model type

Dense (1 x 1) MoE (64 x 1/4)

4

% peS2o

3.8

10% peS2o 50% peS2o 100% peS2o

3.6

5 DCLM (R=1)

3.4 1

2

4

8

Repetition Rate

16

32

1

2

4

8

Repetition Rate

16

32

(b) DCLM + peS2o

Figure 7: Mixing a repeated data domain with a non-repeated domain may have a regularizing effect (§3.4). We mix peS2o and StarCoder, repeated R ∈ {1, 2, 4, 8, 16, 32} times, into non-repeated DCLM, with various proportions of peS2o/StarCoder. We find that DCLM + StarCoder mixes degrade similarly with repetition, regardless of the StarCoder mixing percentage. However, DCLM appears to have a regularizing effect in DCLM + peS2o mixes: as DCLM takes up a larger proportion of the mix, high repetition rates (R = 32) degrade at a slower rate.

5

M ECHANISTIC I NVESTIGATION OF DATA R EPETITION

In §3, we hypothesize that MoEs may suffer more from data repetition if each expert’s routed token set is fixed early in training, thus exposing each expert FFN to a significantly reduced set of unique tokens, when compared to dense FFNs. We consider two aspects of this hypothesis: firstly, the stability of the router, which would result in a fixed data partition over experts; secondly, the resulting specialization of each expert. 5.1

M O E ROUTER O SSIFICATION

We study routing patterns for each saved checkpoint: we run inference on a fixed batch of 16,384 tokens from the Dolma Common Crawl validation split, and record each token’s top-1 expert at every MoE layer. We define the routing stability at checkpoint i as the fraction of tokens whose top-1 expert did not change from checkpoint i − 1, averaged over layers. Separately, we consider expert co-activation and router load balance. Further details are in Appendix A.6. Routing ossifies early. We compute routing stability for MoE models in §3. Figure 9a shows results for 80M models. Routing is unstable (random chance, 0.02 ≈ 1/64) at the first checkpoint (step 200), as models begin from random initialization. By the next checkpoint (step 400, 10% of training), routing stability rises to 60%, and continues to rise rapidly to > 95% at the end of training. Each MoE configuration at each scale exhibits similar patterns (Appendix B.8). 8

Data Scarcity and Model Sparsity

10

Dense (1 x 1) MoE (64 x 1/4) Dropout

8

None (disabled) 0.1 0.2 0.4

7 6

Model type

Dense (1 x 1) MoE (64 x 1/4)

9

Common Crawl CE (↓)

9

Common Crawl CE (↓)

10

Model type

FOM Prob.

8

0.0 0.1 0.2 0.4

7 6 5

5 1

2

4

8

16

Repetition Rate

32

1

64

2

(a) Dropout

8

16

32

64

(b) Final Output Masking

10

10

Model type

Dense (1 x 1) MoE (64 x 1/4) Expert Dropout

8

0.0 0.1 0.2 0.4

7 6 5

Model type

Dense (1 x 1) MoE (64 x 1/4)

9

Common Crawl CE (↓)

9

Common Crawl CE (↓)

4

Repetition Rate

EOM Prob.

8

0.0 0.1 0.2 0.4

7 6 5

1

2

4

8

16

Repetition Rate

32

64

1

(c) Expert Dropout

2

4

8

16

Repetition Rate

32

64

(d) Expert Output Masking

Figure 8: Dropout (a), FFN Output Masking (b), Expert Dropout (c), and Expert Output Masking (d) each reduces overfitting from data repetition (§4). Of the regularization methods studied in §4, these 4 dramatically decrease the response to data repetition. However, weight decay and gradient norm clipping, as well as MoE router jitter, have minimal effect. See additional figures in Appendix B.5. Higher data repetition exacerbates router ossification. In Figure 9a, router stability rises uniformly across all data repetition rates until step 600, at which point higher repetition models begin to show consistently higher routing stability. In Figure 9b, end-of-training router stability rises with data repetition, across model scales and MoE sparsities (Appendix B.8). This supports our hypothesis that each expert’s token set is nearly stationary for most of training, so under repetition an expert sees the same reduced shard of data over and over. Dropout’s regularizing effect does not operate through router plasticity. We train 200M models using the dropout settings of §4 (with checkpoints 1,000 steps apart, not directly comparable to the above results). Late-training router stability again rises slowly but consistently with repetition (Appendix Figure 22). At R = 64, models with dropout show less routing ossification than the no-dropout setting, even though they overfit far less (§4). Since dropout partially recovers performance without affecting router ossification, we hypothesize that overfitting is not primarily driven by routing, but rather the functions learned by each expert. Higher repetition encourages uniformly distributed expert co-activation. Although the top-1 expert contributes the majority of the output weight, we also conduct a simple investigation of expert co-activation. For each unordered pair of experts within each layer, we count how often both experts in the pair are active for the same token. We normalize the counts and compute the entropy of the distribution. Under repetition this distribution trends toward uniform in every configuration (Appendix B.8). Router output magnitude is higher with data repetition; load balancing shows no clear correlation. Router load imbalance, or the ratio between the maximum and mean expert token loads in each batch, does not predictably shift with repetition rate, except that extremely high R yields outliers 1B scale. Load balancing loss also does not change predictably with repetition. Z-loss, which measures router logit magnitude, is higher in early training with higher R, which is consistent with earlier and more extreme router ossification (Appendix Figure 13-15). 9

Data Scarcity and Model Sparsity

Model type MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x 64x 128x 256x

Top-1 routing stability

0.8 0.6 0.4 0.2 0.0

200

400

600

800

1000

Training step

1200

Routing stability, end of training

MoE (64 x 1/4)

1.0

Model type MoE (32 x 1/4) MoE (64 x 1/4) Scale 80M 200M

0.94 0.93 0.92

1400

1

(a) Top-1 routing stability over training (80M)

2

4

8

Repetition Rate

16

32

(b) Final top-1 routing stability (80M, 200M)

Model type

MoE (16 x 1/4) MoE (32 x 1/4) MoE (64 x 1/4) MoE (128 x 1/4)

0.01 0.005 0.003 0.002 0.001 1

2

4

8

16

Repetition Rate

32

64

Expert knockout: median CE increase

Expert knockout: median CE increase

Figure 9: Routing ossifies early in training, exacerbated by repetition (§5.1). We show Top-1 routing stability, defined as the fraction of held-out Common Crawl tokens that keep the same top-1 expert between consecutive checkpoints (200 steps apart), averaged over layers. Left: For the 80M MoE (64 x 1/4) models, consecutive checkpoint agreement is near chance (1/64) at the start of training, but over 0.9 for the second half of training. Stability increases with repetition, up to R = 32, where the router appears to destabilize. Right: We show the Top-1 routing stability at the end of training for MoE (32 x 1/4) and (64 x 1/4) at 80M and 200M scale. Stability rises more aggressively with R at 200M active parameters. Also see Appendix Figure 19-22. 0.01

Model type

0.005

Dropout

MoE (32 x 1/4) MoE (64 x 1/4) 0 0.1 0.4

0.003 0.002 0.001 0.0005 0.0003

1

2

4

8

Repetition Rate

16

32

Figure 10: Repetition increases expert specialization; dropout reduces this effect (§5.2). We compute the increase in held-out CE when one expert’s output is zeroed at inference time, and show the median over all experts and MoE layers of the final checkpoint. Left (80M active parameters): Expert specialization grows with repetition rate, and is overall higher when there are fewer total experts, but is unaffected by expert granularity (Appendix Figure 20) Right (200M active parameters): Dropout reduces expert specialization across all R, and also reduces the relative magnitude of the increase in expert specialization that results from higher R. 5.2

M O E E XPERT S PECIALIZATION

Early routing ossification does not necessitate expert specialization; experts with disjoint routed token sets may still learn similar functions. We investigate the specialization of experts by measuring the performance impact of expert knockout, or the loss impact of removing an expert. Specifically, we perform inference with the final checkpoint of each model and measure, for each expert at each layer, the total increase in CE Loss resulting from masking only that expert’s outputs. Expert specialization rises with data repetition. Knockout cost at R = 1 is higher at low total expert count. As R increases, knockout cost increases for all configurations, but does so more rapidly at higher expert counts. At 80M (Figure 10a), increasing R from 1 to 32 also increases expert knockout effect by 1.1× for 16 experts, and by 2.3× for 128 experts of 1/4 granularity. This pattern echos §3, where CE degradation under repetition also grows with expert count. Thus, repetition encourages expert specialization and reduces redundancy, especially at higher expert counts. Dropout decreases expert specialization. We compare various dropout settings at 200M (Figure 10b). Dropout consistently reduces the median knockout cost for all values of R. We hypothesize that dropout combats overfitting by removing dependence on any single expert, and enforcing multiple, more varied, representations of features. 10

Data Scarcity and Model Sparsity

6

R ELATED W ORKS

Hernandez et al. (2022) show that repeating a small fraction of the training data degrades held-out loss nonmonotonically. Muennighoff et al. (2023) studies dense models trained on C4 and finds that repetition up to roughly four epochs is comparable to all-unique data, but further repetition decays performance. More recent work on dense models has provided evidence that data repetition damage: (1) grows with model scale (Kazdan et al., 2026); (2) peaks at an intermediate repeat count (Chudnovsky et al., 2026); (3) is well described by a single additive coefficient that strong weight decay can shrink (Lovelace et al., 2026); and (4) can be modulated by data quality and mixture weights (Fang et al., 2025; Liu et al., 2026; Chen et al., 2025). Very little work has considered MoEs; Xue et al. (2023) concludes from a single MoE configuration, a 16-expert T5, that parameter count drives multi-epoch degradation while FLOPs are close to irrelevant. Other work has modeled the tradeoff between quality and quantity of unique data when training dense Transformers. Fang et al. (2025) reports that repeating a heavily filtered set up to ten times can beat a single pass over a superset ten times larger, while Mohri et al. (2026) argues that with enough compute, larger quantities of unfiltered data is better. Liu et al. (2026) and Chen et al. (2025) fit quality-weighted mixtures and repetition together. Even fewer studies have considered interventions for minimizing data-repetition effects. Xue et al. (2023) uses dropout to reduce multi-epoch degradation, but caution that the dropout probability needs retuning as models grow. Lovelace et al. (2026) shows that raising weight decay by an order of magnitude cuts their fitted overfitting coefficient by roughly 70% at high repetition rates, at the price of a loss premium in the single-epoch regime.

7

C ONCLUSION

Across all dense and MoE architectures, we find that increasing data repetition rate R leads to overfitting. Sparse MoE models deteriorate earlier and more rapidly as R increases. We show that the response to data repetition primarily depends on total parameters. The pattern of overfitting effects are remarkably robust to a variety of single data domains, data mixes, and different levels of data filtering. Mixing a repeated data domain into a larger or equal-sized nonrepeated data domain may have a regularizing effect. We successfully reduce the overfitting response to data repetition through regularization methods that operate by dropping parameter outputs (dropout, expert dropout, FFN output masking, and expert output masking). However, no method fully matches performance achieved with all-unique data. Gradient Norm Clipping, Weight Decay, and Router Jitter operate through qualitatively different mechanisms to reduce update strength, reduce weight magnitude, and modify coarse-grained gradient paths, respectively, and do not have any measurable effect. Finally, our mechanistic analyses provide evidence that MoE routing is fixed early in training, and only slightly exacerbated by data repetition. Expert specialization also rises with data repetition. Dropout does not affect router ossification, but decreases expert specialization. In summary, our work studies the interaction between data constraints and the design of MoE architectures. We present strong evidence that data repetition causes overfitting, that the degree of performance degradation primarily depends on total parameters, and that the underlying mechanism operates partially through overly specialized parameters, broken by methods such as dropout and output masking. We recommend that future works further explore masking-based methods to minimize repetition-driven overfitting through decreased parameter specialization.

ACKNOWLEDGMENTS We are grateful to Rohan Sanda for initial engineering support; to Ananya Harsh Jha and Jacqueline He for helpful discussion; and to the contributors and maintainers of the UW Hyak and Stanford Marlowe computing resources.

R EFERENCES Zhangir Azerbayev, Hailey Schoelkopf, Keiran Paster, Marco Dos Santos, Stephen McAleer, Albert Q. Jiang, Jia Deng, Stella Biderman, and Sean Welleck. Llemma: An open language model for mathematics, 2023. Tom Brown, Benjamin Mann, Nick Ryder, Melanie Subbiah, Jared D Kaplan, Prafulla Dhariwal, Arvind Neelakantan, Pranav Shyam, Girish Sastry, Amanda Askell, et al. Language models are few-shot learners. Advances in neural information processing systems, 33:1877–1901, 2020. Zhengyu Chen, Siqi Wang, Teng Xiao, Yudong Wang, Shiqi Chen, Xunliang Cai, Junxian He, and Jingang Wang. Revisiting scaling laws for language models: The role of data quality and training strategies. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (ACL), pp. 23897–23920, 2025. URL https://aclanthology.org/2025.acl-long.1163/. 11

Data Scarcity and Model Sparsity

Jessica Chudnovsky, Joshua Kazdan, Noam Levi, Rylan Schaeffer, Yegor Denisov-Blanch, Bo He, Mehmet Donmez, Sanmi Koyejo, and David Donoho. Internal data repetition destroys language models, 2026. URL https:// arxiv.org/abs/2606.24998. Christopher Clark, Kenton Lee, Ming-Wei Chang, Tom Kwiatkowski, Michael Collins, and Kristina Toutanova. Boolq: Exploring the surprising difficulty of natural yes/no questions. arXiv preprint arXiv:1905.10044, 2019. Together Computer. Redpajama: An open source recipe to reproduce llama training dataset, 2023. URL https: //github.com/togethercomputer/RedPajama-Data. Damai Dai, Chengqi Deng, Chenggang Zhao, R. X. Xu, Huazuo Gao, Deli Chen, Jiashi Li, Wangding Zeng, Xingkai Yu, Y. Wu, Zhenda Xie, Y. K. Li, Panpan Huang, Fuli Luo, Chong Ruan, Zhifang Sui, and Wenfeng Liang. Deepseekmoe: Towards ultimate expert specialization in mixture-of-experts language models, 2024. URL https://arxiv.org/abs/2401.06066. Jesse Dodge, Maarten Sap, Ana Marasović, William Agnew, Gabriel Ilharco, Dirk Groeneveld, Margaret Mitchell, and Matt Gardner. Documenting large webtext corpora: A case study on the colossal clean crawled corpus. In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 1286–1305, Online and Punta Cana, Dominican Republic, November 2021. Association for Computational Linguistics. doi: 10.18653/ v1/2021.emnlp-main.98. URL https://aclanthology.org/2021.emnlp-main.98. Alex Fang, Hadi Pouransari, Matt Jordan, Alexander Toshev, Vaishaal Shankar, Ludwig Schmidt, and Tom Gunter. Datasets, documents, and repetitions: The practicalities of unequal data quality, 2025. URL https://arxiv. org/abs/2503.07879. William Fedus, Barret Zoph, and Noam Shazeer. Switch transformers: Scaling to trillion parameter models with simple and efficient sparsity. Journal of Machine Learning Research, 23(120):1–39, 2022. URL http://jmlr.org/ papers/v23/21-0998.html. Leo Gao, Stella Rose Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. The pile: An 800gb dataset of diverse text for language modeling. ArXiv, abs/2101.00027, 2020. URL https://api.semanticscholar.org/ CorpusID:230435736. Sidney Greenbaum and Gerald Nelson. The international corpus of english (ICE) project. World Englishes, 15 (1):3–15, mar 1996. doi: 10.1111/j.1467-971x.1996.tb00088.x. URL https://doi.org/10.1111%2Fj. 1467-971x.1996.tb00088.x. Danny Hernandez, Tom Brown, Tom Conerly, Nova DasSarma, Dawn Drain, Sheer El-Showk, Nelson Elhage, Zac Hatfield-Dodds, Tom Henighan, Tristan Hume, et al. Scaling laws and interpretability of learning from repeated data. arXiv preprint arXiv:2205.10487, 2022. Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Tom Hennigan, Eric Noland, Katie Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karen Simonyan, Erich Elsen, Jack W. Rae, Oriol Vinyals, and Laurent Sifre. Training compute-optimal large language models, 2022. URL https://arxiv.org/abs/2203.15556. Joshua Kazdan, Noam Levi, Rylan Schaeffer, Jessica Chudnovsky, Abhay Puri, Bo He, Mehmet Donmez, Sanmi Koyejo, and David Donoho. Scale dependent data duplication, 2026. URL https://arxiv.org/abs/2603. 06603. Denis Kocetkov, Raymond Li, Loubna Ben Allal, Jia Li, Chenghao Mou, Carlos Muñoz Ferrandis, Yacine Jernite, Margaret Mitchell, Sean Hughes, Thomas Wolf, Dzmitry Bahdanau, Leandro von Werra, and Harm de Vries. The stack: 3 tb of permissively licensed source code, 2022. URL https://arxiv.org/abs/2211.15533. Jeffrey Li, Alex Fang, Georgios Smyrnis, Maor Ivgi, Matt Jordan, Samir Gadre, Hritik Bansal, Etash Guha, Sedrick Keh, Kushal Arora, Saurabh Garg, Rui Xin, Niklas Muennighoff, Reinhard Heckel, Jean Mercat, Mayee Chen, Suchin Gururangan, Mitchell Wortsman, Alon Albalak, Yonatan Bitton, Marianna Nezhurina, Amro Abbas, ChengYu Hsieh, Dhruba Ghosh, Josh Gardner, Maciej Kilian, Hanlin Zhang, Rulin Shao, Sarah Pratt, Sunny Sanyal, Gabriel Ilharco, Giannis Daras, Kalyani Marathe, Aaron Gokaslan, Jieyu Zhang, Khyathi Chandu, Thao Nguyen, Igor Vasiljevic, Sham Kakade, Shuran Song, Sujay Sanghavi, Fartash Faghri, Sewoong Oh, Luke Zettlemoyer, Kyle Lo, Alaaeldin El-Nouby, Hadi Pouransari, Alexander Toshev, Stephanie Wang, Dirk Groeneveld, Luca Soldaini, Pang Wei Koh, Jenia Jitsev, Thomas Kollar, Alexandros G. Dimakis, Yair Carmon, Achal Dave, Ludwig Schmidt, and Vaishaal Shankar. Datacomp-lm: In search of the next generation of training sets for language models, 2024. Raymond Li, Loubna Ben Allal, Yangtian Zi, Niklas Muennighoff, Denis Kocetkov, Chenghao Mou, Marc Marone, Christopher Akiki, Jia Li, Jenny Chim, et al. Starcoder: may the source be with you!, 2023. 12

Data Scarcity and Model Sparsity

Percy Liang, Rishi Bommasani, Tony Lee, Dimitris Tsipras, Dilara Soylu, Michihiro Yasunaga, Yian Zhang, Deepak Narayanan, Yuhuai Wu, Ananya Kumar, Benjamin Newman, Binhang Yuan, Bobby Yan, Ce Zhang, Christian Cosgrove, Christopher D. Manning, Christopher R’e, Diana Acosta-Navas, Drew A. Hudson, E. Zelikman, Esin Durmus, Faisal Ladhak, Frieda Rong, Hongyu Ren, Huaxiu Yao, Jue Wang, Keshav Santhanam, Laurel J. Orr, Lucia Zheng, Mert Yuksekgonul, Mirac Suzgun, Nathan S. Kim, Neel Guha, Niladri S. Chatterji, Omar Khattab, Peter Henderson, Qian Huang, Ryan Chi, Sang Michael Xie, Shibani Santurkar, Surya Ganguli, Tatsunori Hashimoto, Thomas F. Icard, Tianyi Zhang, Vishrav Chaudhary, William Wang, Xuechen Li, Yifan Mai, Yuhui Zhang, and Yuta Koreeda. Holistic evaluation of language models. Annals of the New York Academy of Sciences, 1525:140 – 146, 2022. URL https://api.semanticscholar.org/CorpusID:253553585. Fengze Liu, Weidong Zhou, Binbin Liu, Ping Guo, Zijun Wang, Bingni Zhang, Yifan Zhang, Yifeng Yu, Xiaohuan Zhou, and Taifeng Wang. Infolaw: Information scaling laws for large language models with quality-weighted mixture data and repetition, 2026. URL https://arxiv.org/abs/2605.02364. Ilya Loshchilov and Frank Hutter. Decoupled weight decay regularization. In International Conference on Learning Representations, 2019. Justin Lovelace, Christian Belardi, Srivatsa Kundurthy, Shriya Sudhakar, and Kilian Q. Weinberger. Prescriptive scaling laws for data constrained training, 2026. URL https://arxiv.org/abs/2605.01640. Ian Magnusson, Akshita Bhagia, Valentin Hofmann, Luca Soldaini, Ananya Harsh Jha, Oyvind Tafjord, Dustin Schwenk, Evan Pete Walsh, Yanai Elazar, Kyle Lo, Dirk Groeneveld, Iz Beltagy, Hannaneh Hajishirzi, Noah A. Smith, Kyle Richardson, and Jesse Dodge. Paloma: A benchmark for evaluating language model fit, 2024. URL https://arxiv.org/abs/2312.10523. Stephen Merity, Caiming Xiong, James Bradbury, and Richard Socher. Pointer sentinel mixture models. ArXiv, abs/1609.07843, 2016. URL https://api.semanticscholar.org/CorpusID:16299141. Christopher Mohri, John Duchi, and Tatsunori Hashimoto. A bitter lesson for data filtering, 2026. URL https: //arxiv.org/abs/2605.19407. Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Aleksandra Piktus, Nouamane Tazi, Sampo Pyysalo, Thomas Wolf, and Colin Raffel. Scaling data-constrained language models. In Advances in Neural Information Processing Systems (NeurIPS), 2023. URL https://arxiv.org/abs/2305.16264. Niklas Muennighoff, Luca Soldaini, Dirk Groeneveld, Kyle Lo, Jacob Morrison, Sewon Min, Weijia Shi, Pete Walsh, Oyvind Tafjord, Nathan Lambert, Yuling Gu, Shane Arora, Akshita Bhagia, Dustin Schwenk, David Wadden, Alexander Wettig, Binyuan Hui, Tim Dettmers, Douwe Kiela, Ali Farhadi, Noah A. Smith, Pang Wei Koh, Amanpreet Singh, and Hannaneh Hajishirzi. Olmoe: Open mixture-of-experts language models, 2025. URL https://arxiv.org/abs/2409.02060. Razvan Pascanu, Tomas Mikolov, and Yoshua Bengio. On the difficulty of training recurrent neural networks. In International conference on machine learning, pp. 1310–1318. PMLR, 2013. Keiran Paster, Marco Dos Santos, Zhangir Azerbayev, and Jimmy Ba. Openwebmath: An open dataset of high-quality mathematical web text, 2023. Colin Raffel, Noam M. Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. Exploring the limits of transfer learning with a unified text-to-text transformer. ArXiv, abs/1910.10683, 2019. URL https://api.semanticscholar.org/CorpusID:204838007. Machel Reid, Victor Zhong, Suchin Gururangan, and Luke Zettlemoyer. M2D2: A massively multi-domain language modeling dataset. In Proceedings of the 2022 Conference on Empirical Methods in Natural Language Processing, pp. 964–975, Abu Dhabi, United Arab Emirates, December 2022. Association for Computational Linguistics. URL https://aclanthology.org/2022.emnlp-main.63. Noam Shazeer, Azalia Mirhoseini, Krzysztof Maziarz, Andy Davis, Quoc Le, Geoffrey Hinton, and Jeff Dean. Outrageously large neural networks: The sparsely-gated mixture-of-experts layer. arXiv preprint arXiv:1701.06538, 2017. Luca Soldaini and Kyle Lo. peS2o (Pretraining Efficiently on S2ORC) Dataset, 2023. URL https://github. com/allenai/pes2o. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo. Dolma: an open corpus of three trillion tokens for language model pretraining research, 2024a. 13

Data Scarcity and Model Sparsity

Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Raghavi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, A. Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Daniel Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A. Smith, Hanna Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo. Dolma: an open corpus of three trillion tokens for language model pretraining research. ArXiv, abs/2402.00159, 2024b. URL https://api.semanticscholar.org/CorpusID:267364861. Guijin Son, Hanwool Lee, Sungdong Kim, Seungone Kim, Niklas Muennighoff, Taekyoon Choi, Cheonbok Park, Kang Min Yoo, and Stella Biderman. Kmmlu: Measuring massive multitask language understanding in korean, 2024. URL https://arxiv.org/abs/2402.11548. Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. Dropout: a simple way to prevent neural networks from overfitting. The journal of machine learning research, 15(1):1929–1958, 2014. Fuzhao Xue, Yao Fu, Wangchunshu Zhou, Zangwei Zheng, and Yang You. To repeat or not to repeat: Insights from scaling LLM under token-crisis. In Advances in Neural Information Processing Systems (NeurIPS), 2023. URL https://arxiv.org/abs/2305.13230. Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. Hellaswag: Can a machine really finish your sentence? arXiv preprint arXiv:1905.07830, 2019. URL https://arxiv.org/abs/1905.07830. Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. Stmoe: Designing stable and transferable sparse expert models, 2022a. URL https://arxiv.org/abs/2202. 08906. Barret Zoph, Irwan Bello, Sameer Kumar, Nan Du, Yanping Huang, Jeff Dean, Noam Shazeer, and William Fedus. St-moe: Designing stable and transferable sparse expert models. arXiv preprint arXiv:2202.08906, 2022b. [221] in OLMoE.

14

Data Scarcity and Model Sparsity

A

E XPERIMENTAL D ETAILS

A.1

M ODEL A RCHITECTURE

Scale

Layers

Model Dim

Attention Heads

80M

8

336

7

200M

1B

10

15

640

1664

10

16

Name

Activation Sparsity (s)

Total Experts (n)

Active Experts (k)

Expert Gran. (g)

dense MoE (8 x 1/4) MoE (64 x 1/32) MoE (16 x 1/4) MoE (64 x 1/16) MoE (32 x 1/4) MoE (64 x 1/8) MoE (64 x 1/4) MoE (64 x 1/2) MoE (128 x 1/4) MoE (256 x 1/4)

1 2 2 4 4 8 8 16 32 32 64

8 64 16 64 32 64 64 64 128 256

4 32 4 16 4 8 4 2 4 4

1/4 1/32 1/4 1/16 1/4 1/8 1/4 1/2 1/4 1/4

dense MoE (32 x 1/4) MoE (64 x 1/4)

1 8 16

32 64

4 4

1/4 1/4

dense MoE (64 x 1/4)

1 16

64

4

1/4

Total Tokens (T )

Active Param (Na )

Total Param (N )

1.6B

81.8M

81.8M 92.7M 92.7M 114.4M 114.4M 157.7M 157.7M 244.4M 417.8M 417.8M 764.6M

4B

193.9M

193.9M 538.0M 931.2M

20B

998.3M

998.3M 8.5B

Table 1: Architecture Details and Parameter Counts A.2

H YPERPARAMETERS Hyperparameter

Value

Vocabulary size Batch Size Sequence Length Learning Rate Encoder-Decoder Weight Sharing Feedforward Dimension LR Schedule LR Warmup End LR Weight Decay Max Grad Norm Dropout Nonlinearity MoE Z-loss Load Balancing Loss Weight MoE Token Dropping MoE Routing Choice

50K 512 2048 4e-4 No 4 x hidden dimension Cosine Decay 2000 steps 0.1 x Peak LR {0.0, 0.1, 0.2, 0.4 } {None, 0.2, 1, 2.0 } {0.0, 0.1, 0.2, 0.4} SwiGLU 1e-3 1e-2 Dropless Token Choice

Table 2: Hyperparameter details for models in §3-5. Multiple values indicate that we investigated different settings in §4, and bold values are defaults used in §3. A.3

T RAINING DATA S OURCES

We take our training data from Muennighoff et al. (2025). We use their data mix, which we call OLMoE Mix, consisting of documents from: DCLM-Baseline (Li et al., 2024), StarCoder (Li et al., 2023; Kocetkov et al., 2022), peS2o (Soldaini & Lo, 2023; Soldaini et al., 2024a), arXiv (Computer, 2023), OpenWebMath (Paster et al., 2023), Algebraic Stack (Azerbayev et al., 2023), English Wikipedia & Wikibooks (Soldaini et al., 2024a). In §3.3, we also use DCLM-P OOL (Li et al., 2024). A.4

E VALUATION DATA

Our evaluation includes held-out validation sets for language modeling, as well as downstream tasks. The language modeling tasks are a subset of Paloma (Magnusson et al., 2024), which consists of: C4 (Raffel et al. (2019) via Dodge et al. (2021)), T HE P ILE (Gao et al., 2020), W IKI T EXT-103 (Merity et al., 2016), D OLMA (Soldaini 15

Data Scarcity and Model Sparsity

et al., 2024b), M2D2 S2ORC (Reid et al., 2022), ICE (Greenbaum & Nelson (1996) via Liang et al. (2022)). D OLMA is subdivided into six domains: books, common-crawl, pes2o, reddit uniform, stack uniform, wiki. In §3.2, we vary the dataset used for validation loss to match the training domain. For models trained on DCLM, we evaluate on Dolma common-crawl; for peS2o, Dolma pes2o; for Wikipedia, Dolma wiki; for StarCoder, Dolma stack uniform. The downstream tasks consist of BoolQ (Clark et al., 2019), HellaSwag (Zellers et al., 2019), and MMLU (Son et al., 2024). MMLU is subdivided into four domains: humanities, STEM, social sciences, and other. A.5

DATA R EPETITION

In data-constrained regimes, the set of U unique tokens used for each experiment is constructed as follows: we fix a random permutation of the sequences in each data domain D, and take the first U tokens for training. We repeat these U tokens for R epochs, shuffling between epochs. In §3.3, we mix tokens from DCLM-BASELINE and DCLM-P OOL (Li et al., 2024). This inevitably introduces a very small amount of unmeasured data repetition because our selected subsets of DCLM-BASELINE and DCLM-P OOL may have a non-empty intersection. However, this intersection is likely to be of negligible size. DCLM-BASELINE consists of 5T tokens, of which we use 1.6B at most. DCLM-P OOL consists of 240 trillion tokens, of which we use 1.6B at most. The probability that any particular token in a particular sequence from our DCLM-BASELINE subset also appears in our DCLM-P OOL is less than 2e-9, yielding an expected total of fewer than 3 repeated tokens. In other words, we expect effectively no repeated tokens on average. A.6

ROUTING A NALYSIS

In §5, we use a batch size of 16,384 tokens (8 sequences of 2,048 tokens). We use a fixed subset of the Dolma Common Crawl validation split. Dropout, router jitter, and other train-time regularizers are inactive. Ossification (§5.1). For each checkpoint we record the top-1 expert of every token at every MoE layer, defined as the expert with the largest router score. For each pair of consecutive checkpoints we report the fraction of tokens whose top-1 expert is identical, averaged over layers. Checkpoints are 200 steps apart in the core ladders and 1,000 steps apart in the 200M dropout arms, so stability values are only compared between runs with matching spacing. End-of-training stability is the mean over the last two regularly spaced intervals. Expert knockout (§5.2). On the final checkpoint we zero all MLP weights of one expert, so its output is exactly zero for the tokens routed to it. The router is untouched and the weights of the remaining selected experts are not renormalized. We then recompute CE on the token batch, repeat for every expert in every MoE layer, and report the median and maximum increase over the CE of the unmodified model. Co-activation (§5.2). On the final checkpoint we count, per layer, how often each unordered pair of experts appears together in a token’s top-k set, normalize the counts to a distribution, and compute its Shannon entropy. We divide by  log n(n − 1) , its maximum for n experts, so values are comparable across expert counts.

16

Data Scarcity and Model Sparsity

B

A DDITIONAL R ESULTS

B.1

R ANDOM S EED VARIANCE

In Table 3, we report the variance across 5 random seeds for Dense and MoE (64 x 1/4) models trained on the OLMoE mix at R = 1, R = 32, and evaluated on all language modeling validation datasets and downstream tasks. Metric

Dense R=1 R=32 Mean Std. Dev. Mean Std. Dev.

MoE (64 x 1/4) R=1 R=32 Mean Std. Dev. Mean Std. Dev.

Train Loss

4.57

0.00

4.27

0.00

4.26

0.01

3.17

0.02

Validation LM Loss C4 Dolma Books Dolma Common Crawl Dolma peS2o Dolma Reddit Dolma Stack Dolma Wiki ICE M2D2 S2ORC Pile WikiText-103 Average

4.86 5.04 4.92 4.48 4.74 4.65 4.67 4.97 4.84 4.62 5.06 4.81

0.01 0.00 0.01 0.01 0.00 0.02 0.01 0.01 0.01 0.01 0.00 0.01

5.27 5.59 5.29 4.93 5.12 6.68 5.16 5.63 5.47 5.37 5.76 5.48

0.02 0.06 0.02 0.04 0.02 0.10 0.04 0.03 0.03 0.05 0.08 0.05

4.53 4.72 4.61 4.12 4.46 4.23 4.31 4.66 4.52 4.27 4.67 4.46

0.01 0.01 0.01 0.01 0.01 0.02 0.01 0.02 0.01 0.01 0.02 0.01

6.04 6.48 6.05 5.66 5.95 7.46 5.87 6.73 6.46 6.19 6.67 6.32

0.01 0.05 0.02 0.01 0.02 0.04 0.02 0.01 0.01 0.01 0.03 0.02

Downstream Task Loss BoolQ HellaSwag MMLU Humanities MMLU Other MMLU Social Sciences MMLU STEM Average

2.52 0.96 2.13 1.97 1.98 2.02 1.93

0.22 0.00 0.06 0.04 0.04 0.03 0.07

3.14 1.05 3.48 2.85 2.66 2.61 2.63

0.18 0.00 0.11 0.17 0.14 0.07 0.11

2.33 0.89 1.91 1.80 1.88 1.90 1.78

0.24 0.00 0.13 0.06 0.10 0.08 0.10

3.95 1.23 4.36 2.93 3.06 3.13 3.11

0.43 0.00 0.92 0.20 0.31 0.48 0.39

Downstream Task Accuracy BoolQ HellaSwag MMLU Humanities MMLU Other MMLU Social Sciences MMLU STEM Average

0.39 0.26 0.24 0.27 0.23 0.27 0.28

0.01 0.00 0.01 0.01 0.01 0.00 0.01

0.39 0.25 0.24 0.24 0.22 0.24 0.26

0.01 0.00 0.00 0.01 0.00 0.01 0.01

0.41 0.26 0.25 0.27 0.24 0.27 0.28

0.05 0.00 0.01 0.01 0.01 0.01 0.01

0.45 0.25 0.25 0.25 0.23 0.24 0.28

0.03 0.00 0.01 0.01 0.01 0.01 0.01

Table 3: Mean and Standard Deviation across 5 random seeds. We repeat a selection of the 80M settings from §3.1 using 5 random seeds for model initialization, and repeat the mean and standard deviation on each metric. Variance at R = 1 is near-0 for validation LM datasets. Despite fixing the data used across model initialization seeds, higher data repetition rates yield slightly higher standard deviation. Downstream task loss has higher variance for MoE models, and at higher R. Downstream task accuracy has low variance, but remains at near-chance scores.

17

Data Scarcity and Model Sparsity

E XTENDED S ETTINGS (§3.1) Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

1

Model type

Common Crawl CE (↓)

Train CE (↓) (last-100-step avg)

B.2

0.1 0.01 0.001

Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6 5

0.0001

20 23 26 29 212 215 218 221

20

Repetition Rate

23

26

29 212 215 218 221

Repetition Rate

20

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

1

Common Crawl CE (↓)

Train CE (↓) (last-100-step avg)

(a) 80M

0.1

0.01

0.001

Model type

Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6 5 4

1

4

16

64

Repetition Rate

256

1

1024

4

16

64

Repetition Rate

256

1024

Model type Dense (1 x 1) MoE (64 x 1/4)

1

Model type

Dense (1 x 1) MoE (64 x 1/4)

Common Crawl CE (↓)

Train CE (↓) (last-100-step avg)

(b) 200M

0.1

0.01

10 9 8 7 6 5 4 3

1

2

4

8 16 32 64 128 256

Repetition Rate

1

2

4

8 16 32 64 128 256

Repetition Rate

(c) 1B

Figure 11: Across active parameter scales, data repetition rates over 8 result in increasingly severe overfitting. Sparser models overfit more (§3.1). At 80M, 200M, and 1B active parameters, we fix the total data budget T = 20 · Na , and vary the data repetition rate R via different sized unique token sets. As R increases, models increasingly overfit, as we observe decreasing train loss and rising validation loss. Sparsity exacerbates overfitting behavior. Larger sparse models overfit more at lower R. We also consider two additional data repetition rates R ≈ 215 , 220 at 80M active parameters, and find that validation loss rises again.

18

Data Scarcity and Model Sparsity

B.3

T RAINING AND VALIDATION L OSS C URVES (§3.1) Dense (1 x 1)

Train CE (↓)

10

64x 128x

1

0.1 0.01

0.01

0.001

1024x

0.001

1000

1500

Train Step

32x 64x

1

0.1 512x

500

64x

Repetitions

128x

0.01 0

MoE (64 x 1/4)

10

1

256x

0.1

MoE (32 x 1/4)

10

256x 512x 1024x

0

500

1000

1500

Train Step

0.001

128x

256x 512x 1024x

0

500

1000

1500

Train Step

1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x

Common Crawl CE (↓)

(a) Train loss over training (80M) Dense (1 x 1)

MoE (32 x 1/4) 1024x 128x

10 9 8 7 6 5

64x 32x 16x

0

500

1000

256x 128x 512x 1024x

10 9 8 7 6 5

1500

Train Step

64x

32x 16x

0

500

Repetitions

MoE (64 x 1/4)

512x 256x

1000

10 9 8 7 6 5

32x 16x

1500

Train Step

1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x

128x 256x 512x 1024x 64x

0

500

1000

1500

Train Step

(b) Validation loss over training (80M) Dense (1 x 1)

MoE (32 x 1/4)

10

MoE (64 x 1/4) 10

10

16x

Train CE (↓)

16x

1

80x

0.1

128x

0.01

1

32x

1

40x

32x

0.1

0.1

0.01

160x

320x

0.001

0.01

128x 160x

640x

0

1000

2000

Train Step

3000

4000

80x 128x

80x

0

1000

2000

Train Step

3000

0.001

4000

320x 640x

0

1000

2000

Train Step

3000

4000

Repetitions 1x 2x 4x 8x 16x 32x 40x 80x 128x 160x 320x 640x

(c) Train loss over training (200M) Dense (1 x 1)

Common Crawl CE (↓)

20

MoE (32 x 1/4)

MoE (64 x 1/4)

160x 128x

80x 128x 160x

80x

10 9 8 7 6 5 4

40x 32x 16x

0

1000

2000

Train Step

3000

4000

10 9 8 7 6 5 4

40x 32x

16x 8x

0

1000

2000

Train Step

3000

10 9 8 7 6 5

16x

4

4000

(d) Validation loss over training (200M)

19

80x 640x 128x 320x 32x

8x

0

1000

2000

Train Step

3000

4000

Repetitions 1x 2x 4x 8x 16x 32x 40x 80x 128x 160x 320x 640x

Data Scarcity and Model Sparsity

Dense (1 x 1)

MoE (64 x 1/4) 10

Train CE (↓)

10 16x 32x

1

Repetitions 8x

1

16x

0.1

0.1 64x

0.01

0.01

32x 64x 128x 256x

128x 256x

0

5000

10000 15000 20000

0

Train Step

5000

10000 15000 20000

1x 2x 4x 8x 16x 32x 64x 128x 256x

Train Step

(e) Train loss over training (1B)

Common Crawl CE (↓)

Dense (1 x 1)

MoE (64 x 1/4) 64x

10 9 8 7 6 5 4

32x

16x 8x

5000

10000

Train Step

15000

32x

10 9 8 7 6 5 4 3

16x

8x 4x

5000

10000

Train Step

15000

Repetitions 1x 2x 4x 8x 16x 32x 64x 128x 256x

(f) Validation loss over training (1B)

Figure 12: At higher repetition rates, training loss falls to 0, which suggests overfitting to the repeated data (§3.1). We show the training and validation loss curves over the course of training. We consistently observe that higher repetition results in training loss curves that approach 0, mirrored by validation loss curves that rise. At sufficiently high R, we observe the double descent phenomenon Muennighoff et al. (2023), in which validation curves peak then fall.

20

Data Scarcity and Model Sparsity

Router load imbalance (↓)

B.4

ROUTING L OAD BALANCE AND S TABILITY

MoE (32 x 1/4)

8 7 6 5 4 3

MoE (64 x 1/4)

Repetitions

10 5

2 512x 128x 256x 16x

0

500

1000

3 2

1500

Train Step

128x

0

500

1000

Train Step

1500

1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x

Router load imbalance (↓)

(a) 80M active parameters

MoE (32 x 1/4)

Repetitions

MoE (64 x 1/4)

7 6 5 4 3

10 5

2

80x 8x 16x

0

1000

2000

3000

Train Step

4000

3 2

128x 1x 320x 640x

0

1000

2000

Train Step

3000

4000

1x 2x 4x 8x 16x 32x 40x 64x 80x 128x 160x 320x 640x

Router load imbalance (↓)

(b) 200M active parameters

MoE (64 x 1/4)

Repetitions

10 5 3 2

1x 8x 32x 64x 4x 128x 256x

0

5000

10000 15000 20000

1x 2x 4x 8x 16x 32x 64x 128x 256x

Train Step

(c) 1B active parameters

Figure 13: Routing imbalance training curves do not follow clear patterns at lower repetition, but are outliers at high repetition. We plot the routing imbalance, defined as the ratio between the maximum and median expert load, averaged over tokens in the batch. At small scale, routing imbalance does not appear correlated with load imbalance until R > 128, where curves become outliers. At 1B scale, load imbalance curves become disordered, and appear to partially cycle with data repetition periods.

21

Data Scarcity and Model Sparsity

MoE (32 x 1/4)

MoE (64 x 1/4)

Router LB loss (↓)

Repetitions

10 9

10 9

8

8 0

250

500

750

1000 1250 1500

1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x

0

Train Step

250

500

750

1000 1250 1500

Train Step

(a) 80M active parameters

MoE (32 x 1/4)

MoE (64 x 1/4)

Router LB loss (↓)

30

Repetitions

20 20

10

10 0

1000

2000

3000

Train Step

4000

0

1000

2000

Train Step

3000

4000

1x 2x 4x 8x 16x 32x 40x 80x 128x 160x 320x 640x

(b) 200M active parameters MoE (64 x 1/4)

Router LB loss (↓)

100 Repetitions 1x 2x 4x 8x 16x 32x 64x 256x

50 30 20 0

5000

10000

Train Step

15000

20000

(c) 1B active parameters

Figure 14: Routing load balancing loss training curves do not follow clear patterns at small model scale, but may correlate with repetition at larger scale. We report the load balancing loss, as defined in §2. At 80M and 200M, all settings show similar curves. At 1B, high repetition rates appear to affect load balancing loss, with intermediate values of R = 8, 16, 32 resulting in periodicity, and outlier curves at R = 64, 128, 256. It is possible that repetition itself increases load balancing loss, but that the extremely low training loss at high R results in a relatively strong optimization signal from auxiliary losses, eventually driving load balancing loss to fall.

22

Data Scarcity and Model Sparsity

MoE (32 x 1/4)

MoE (64 x 1/4) 200

Router Z loss (↓)

100 128x

50 30 20 10

64x

50

256x 32x

20

512x

10

16x

5 3 2 0

250

500

750

Repetitions

100

8x

5

1024x

2

1000 1250 1500

64x 128x 32x 256x 16x 8x 512x

1024x

0

Train Step

250

500

750

1x 2x 4x 8x 16x 32x 64x 128x 256x 512x 1024x

1000 1250 1500

Train Step

(a) 80M active parameters

Router Z loss (↓)

MoE (32 x 1/4) 200 100 50 40x 32x

20 10 5

80x 16x 128x 160x 8x 4x

2 0

1000

2000

Train Step

Repetitions

MoE (64 x 1/4)

3000

200 100 50 20 10 5 2 1

4000

32x 16x 64x 80x 8x 128x 4x 320x 640x

0

1000

2000

Train Step

3000

4000

1x 2x 4x 8x 16x 32x 40x 64x 80x 128x 160x 320x 640x

(b) 200M active parameters

MoE (64 x 1/4) Repetitions

Router Z loss (↓)

1000 100 16x 32x 8x 64x 4x 128x

10 1

256x

0

5000 10000 15000 20000

1x 2x 4x 8x 16x 32x 64x 128x 256x

Train Step

(c) 1B active parameters

Figure 15: Router z-loss training curves cycle with data repetition. We report the router z-loss (Zoph et al., 2022b). High repetition rates appear to affect z-loss, data-repetition driven cycles at 200M and 1B scale. Higher repetition rates result in higher z-loss with a double peak relatively early in training, resolving to a lower final z-loss. It is possible that high repetition typically results in higher z-loss, but that the extremely low training loss at high R results in a relatively strong optimization signal from auxiliary losses, eventually driving z-loss to fall.

23

Data Scarcity and Model Sparsity

R EGULARIZERS (§4)

10

Dense (1 x 1) MoE (64 x 1/4)

9

Common Crawl CE (↓)

10

Model type Dropout

8

None (disabled) 0.1 0.2 0.4

7 6

Model type

Dense (1 x 1) MoE (64 x 1/4)

9

Common Crawl CE (↓)

B.5

5

Max Grad Norm 0 (no clip) 0.2 1.0 2

8 7 6 5

1

2

4

8

16

Repetition Rate

32

1

64

2

(a) Dropout

16

32

64

10

Model type

Dense (1 x 1) MoE (64 x 1/4) Weight Decay

0.05 0.1 0.2 0.4

8 7 6

Model type

Dense (1 x 1) MoE (64 x 1/4)

9

Common Crawl CE (↓)

9

Common Crawl CE (↓)

8

(b) Gradient Norm Clipping

10

FOM Prob.

8

0.0 0.1 0.2 0.4

7 6 5

5 1

2

4

8

16

Repetition Rate

32

64

1

2

(c) Weight Decay

10

16

32

64

0.0 0.1 0.2 0.4

6 5

Model type

Dense (1 x 1) MoE (64 x 1/4)

9

Common Crawl CE (↓)

Expert Dropout

7

8

10

Dense (1 x 1) MoE (64 x 1/4)

8

4

Repetition Rate

(d) FFN Output Masking Model type

9

Common Crawl CE (↓)

4

Repetition Rate

EOM Prob.

8

0.0 0.1 0.2 0.4

7 6 5

1

2

4

8

16

Repetition Rate

32

64

1

(e) Expert Dropout

2

4

8

16

Repetition Rate

32

64

(f) Expert Output Masking

24

Data Scarcity and Model Sparsity

10

Model type

Dense (1 x 1) MoE (64 x 1/4)

Common Crawl CE (↓)

9

Router Jitter

8

0.0 0.1 0.2 0.4

7 6 5 1

2

4

8

16

Repetition Rate

32

64

(g) MoE Router Jitter

Figure 16: Dropout (a), FFN Output Masking (d), Expert Dropout (e), and Expert Output Masking (f) each reduces overfitting from data repetition (§4). Of the regularization methods studied in §4, these 4 dramatically decrease the response to data repetition. However, weight decay (b) and gradient norm clipping (c), as well as MoE router jitter (f), have minimal effect.

25

Data Scarcity and Model Sparsity

B.6

DATA F ILTERING (§3.3) Model type

Model type

Repetitions

% filtered

1x 2x 4x 8x 16x 32x

6

5

0

25

50

% filtered

75

Dense (1 x 1) MoE (64 x 1/4)

Common Crawl CE (↓)

Common Crawl CE (↓)

Dense (1 x 1) MoE (64 x 1/4)

0% 25% 50% 75% 100%

6

5

100

1

2

4

8

Repetition Rate

16

32

(a) Dolma Common Crawl Validation CE Loss

peS2o CE (↓)

6

5

0

25

50

% filtered

75

Model type Dense (1 x 1) MoE (64 x 1/4) % filtered 0% 25% 50% 75% 100%

7

6

peS2o CE (↓)

Model type Dense (1 x 1) MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x

7

5

100

1

2

4

8

Repetition Rate

16

32

(b) Dolma peS2o Validation CE Loss

Stack CE (↓)

8 7

6

5

Model type Dense (1 x 1) MoE (64 x 1/4) % filtered 0% 25% 50% 75% 100%

9 8

Stack CE (↓)

Model type Dense (1 x 1) MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x

9

7

6

5 0

25

50

% filtered

75

100

1

2

(c) Dolma Stack Validation CE Loss

Figure 17

26

4

8

Repetition Rate

16

32

Data Scarcity and Model Sparsity

B.7

I NTERPOLATING B ETWEEN A S INGLE D OMAIN AND A M IX

We consider an experimental setup similar to §3.3, but with interpolation between DCLM-baseline and a data mix. We use Dolma 1.7 (Soldaini et al., 2024b), which is similar to the OLMoE mix (§A.3; used in §3), but includes more sources. Dolma 1.7 and DCLM-baseline do not fall on a spectrum of quality, but rather of homogeneity. Unlike our experiments of §3.4, which are carefully controlled examinations of 2-domain mixes, the heterogeneity of Dolma 1.7 more closely resembles that of data mixes used for frontier LMs. Model type

Dense (1 x 1) MoE (64 x 1/4)

6

Repetitions

1x 2x 4x 8x 16x 32x

5

0

25

50

% filtered

75

Common Crawl CE (↓)

Common Crawl CE (↓)

Model type

Dense (1 x 1) MoE (64 x 1/4)

6

% filtered

0% 25% 50% 75% 100%

5

100

1

2

4

8

Repetition Rate

16

32

(a) Dolma Common Crawl Validation CE Loss

peS2o CE (↓)

5

Model type Dense (1 x 1) MoE (64 x 1/4) % filtered 0% 25% 50% 75% 100%

6

peS2o CE (↓)

Model type Dense (1 x 1) MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x

6

4

5

4 0

25

50

% filtered

75

100

1

2

4

8

Repetition Rate

16

32

(b) Dolma peS2o Validation CE Loss

8

Stack CE (↓)

7 6 5

4

Model type Dense (1 x 1) MoE (64 x 1/4) % filtered 0% 25% 50% 75% 100%

9 8 7

Stack CE (↓)

Model type Dense (1 x 1) MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x

9

6 5

4

0

25

50

% filtered

75

100

1

2

(c) Dolma Stack Validation CE Loss

Figure 18 27

4

8

Repetition Rate

16

32

Data Scarcity and Model Sparsity

M ECHANISTIC A NALYSES OF ROUTING (§5) MoE (16 x 1/4)

Top-1 routing stability

1.0

Model type MoE (16 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x

0.8 0.6 0.4 0.2 0.0

200

400

600

800

1000

Training step

1200

MoE (32 x 1/4)

1.0

Top-1 routing stability

B.8

0.8 0.6 0.4 0.2 0.0

1400

Model type MoE (32 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x 64x 128x 256x

200

400

600

(a) MoE (16 x 1/4)

0.6 0.4 0.2

1.0

Top-1 routing stability

Top-1 routing stability

Model type MoE (64 x 1/4) Repetitions 1x 2x 4x 8x 16x 32x 64x 128x 256x

0.8

0.0

400

600

800

1000

Training step

1200

Model type MoE (128 x 1/4) Repetitions 2x 4x 8x 16x 32x

0.6 0.4 0.2

1400

500

MoE (64 x 1/8)

0.6 0.4 0.2 600

800

1000

Training step

1200

1400

(e) MoE (64 x 1/8)

Model type MoE (64 x 1/2) Repetitions 1x 2x 4x 8x 16x 32x

0.8 0.6 0.4 0.2 0.0

400

MoE (64 x 1/2)

1.0

Top-1 routing stability

Top-1 routing stability

Model type MoE (64 x 1/8) Repetitions 1x 2x 4x 8x 16x 32x

200

1000

Training step

(d) MoE (128 x 1/4)

0.8

0.0

1400

MoE (128 x 1/4)

(c) MoE (64 x 1/4) 1.0

1200

0.8

0.0 200

1000

(b) MoE (32 x 1/4)

MoE (64 x 1/4)

1.0

800

Training step

500

1000

Training step

(f) MoE (16 x 1/2)

Figure 19: Routing ossifies early in training, exacerbated by repetition (§5.1). We show Top-1 routing stability, which we define as the fraction of held-out Common Crawl tokens that keep the same top-1 expert between consecutive checkpoints (200 steps apart), averaged over layers, for the 80M MoE models. At the beginning of training, consecutive checkpoint agreement is near chance, at roughly n1 for all MoE (n x g) configurations. However, routing stability rises rapidly for all settings, with fewer than 25% of tokens routed to a different top-1 expert when comparing step 400 to 600. Router ossification is slightly higher with fewer experts (lower n) or with higher granularity (larger g). Stability at each checkpoint also increases with repetition rate R, up to roughly R = 16, 32, where the router appears to destabilize.

28

0.91

Routing stability, end of training

1

2

4

8

16

Repetition Rate

32

64

0.94 0.93 0.92 0.91 1

2

4

8

16

Repetition Rate

32

64

Co-activation entropy (normalized)

0.92

0.005 0.003 0.002 0.001 1

2

4

8

16

Repetition Rate

32

0.94

0.002 1

2

4

8

16

Repetition Rate

32

MoE (16 x 1/4) MoE (32 x 1/4) MoE (64 x 1/4) MoE (128 x 1/4)

0.9 0.88

64

0.003

Model type

0.92

1

Co-activation entropy (normalized)

0.93

Expert knockout: median CE increase

0.94

0.01

Expert knockout: median CE increase

Routing stability, end of training

Data Scarcity and Model Sparsity

2

4

8

16

Repetition Rate

32

64

0.94 0.92 0.9 0.88 0.86 0.84

Model type

MoE (64 x 1/8) MoE (64 x 1/4) MoE (64 x 1/2)

0.82

64

1

2

4

8

16

Repetition Rate

32

64

0.94 1

2

4

8

Repetition Rate

16

32

Co-activation entropy (normalized)

0.945

Expert knockout: median CE increase

Routing stability, end of training

Figure 20: Routing stability, expert specialization, and expert-coactivation entropy rise slightly with repetition rate (§5). We show more MoE configurations for (a) end-of-training routing stability, (b) expert knockout effect, and (c) expert co-activation entropy. (a) End-of-training routing stability increases slightly with repetition rate R. Higher expert count naturally results in lower stability, as there are more experts to choose from. Fewer active, but larger experts slightly decreases stability. (b) Expert specialization increases slightly with R. Higher expert count results in lower specialization. Varying active expert count along with expert size has no clear effect. (c) Co-activation entropy of expert pairs (normalized by maximum) rises steadily with R towards uniformly distributed pairings. Entropy is higher with fewer total experts, and with a larger active number of smaller experts.

0.01

0.005 0.003 0.002 1

2

4

8

Repetition Rate

16

32

Model type

MoE (32 x 1/4) MoE (64 x 1/4)

0.945 0.94 0.935 1

2

4

8

Repetition Rate

16

32

Figure 21: Routing stability, expert specialization, and expert-coactivation entropy rise slightly with repetition rate at 200M active parameters (§5). We show 200M MoE (32 x 1/4) and MoE (64 x 1/4) configurations for (a) end-of-training routing stability, (b) expert knockout effect, and (c) expert co-activation entropy. (a) End-of-training routing stability increases slightly with repetition rate R. Higher expert count naturally results in lower stability, as there are more experts to choose from. (b) Expert specialization increases slightly with R. Higher expert count results in lower specialization. (c) Co-activation entropy of expert pairs (normalized by maximum) rises steadily with R towards uniformly distributed pairings. Entropy is higher with fewer total experts.

29

0.92 0.9 0.88 0.86 1

2

4

8

Repetition Rate

16

32

0.01

Co-activation entropy (normalized)

0.94

Expert knockout: median CE increase

Routing stability, end of training

Data Scarcity and Model Sparsity

0.005 0.003 0.002 0.001 0.0005 0.0003

1

2

4

8

Repetition Rate

16

32

0.95

Model type

0.94

MoE (32 x 1/4) MoE (64 x 1/4)

0.93

Dropout

0 0.1 0.4

0.92 0.91 0.9

1

2

4

8

Repetition Rate

16

32

Figure 22: Dropout increases routing stability and expert-coactivation entropy, but decreases expert specialization (§5). We show, for various dropout settings on our 200M MoE (64 x 1/4) configurations: (a) end-of-training routing stability, (b) expert knockout effect, and (c) expert co-activation entropy. (a) End-of-training routing stability increases with higher dropout probability, and still rises with repetition rate R. (b) Expert specialization decreases with higher dropout probability, but still increases slightly with R. (c) Co-activation entropy of expert pairs (normalized by maximum) increases with dropout probability, and still rises steadily with R towards uniformly distributed pairings.

30

Data Scarcity and Model Sparsity

B.9

A DDITIONAL L ANGUAGE M ODELING TASKS

We show results for the additional held-out language modeling tasks of Appendix A.4 on the models from §3, with the extended settings of Appendix B.2. 20

10 9 8 7 6

Model type Dense (1 x 1) MoE (64 x 1/4)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6

C4 CE (↓)

C4 CE (↓)

C4 CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

5

5

5

4

4

20

23

26

29 212 215 218 221

3 1

Repetition Rate

10 9 8 7 6

4

16

64

Repetition Rate

256

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(a) C4 (CE Loss ↓) 20

10 9 8 7 6

Model type Dense (1 x 1) MoE (64 x 1/4)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

Books CE (↓)

Books CE (↓)

Books CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6 4

20

23

26

29 212 215 218 221

3 1

Repetition Rate

5 4

5

5

10 9 8 7 6

4

16

64

Repetition Rate

256

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(b) Dolma books (CE Loss ↓) 20

10 9 8 7 6 5

Model type

Model type

Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6 5

Dense (1 x 1) MoE (64 x 1/4)

Common Crawl CE (↓)

Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

Common Crawl CE (↓)

Common Crawl CE (↓)

Model type

4 20

23

26

29 212 215 218 221

1

Repetition Rate

4

16

64

Repetition Rate

256

10 9 8 7 6 5 4 3

1024

1

2

4

8 16 32 64 128 256

Repetition Rate

(c) Dolma common-crawl (CE Loss ↓) Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6

10 9 8 7 6

peS2o CE (↓)

peS2o CE (↓)

peS2o CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

5

5

4

5

4

3

4

3

20

23

26

29 212 215 218 221

Repetition Rate

1

4

16

64

Repetition Rate

256

1024

(d) Dolma pes2o (CE Loss ↓)

31

Model type Dense (1 x 1) MoE (64 x 1/4)

10 9 8 7 6

1

2

4

8

16 32 64 128 256

Repetition Rate

Data Scarcity and Model Sparsity

20

10 9 8 7 6

Model type Dense (1 x 1) MoE (64 x 1/4)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

Reddit CE (↓)

Reddit CE (↓)

Reddit CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6

5

5

5

4

4

20

23

26

29 212 215 218 221

1

Repetition Rate

10 9 8 7 6

4

16

64

Repetition Rate

256

3

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(e) Dolma reddit (CE Loss ↓)

10 9 8 7 6 5

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10

5

Model type Dense (1 x 1) MoE (64 x 1/4)

10

Stack CE (↓)

Stack CE (↓)

Stack CE (↓)

20

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

20

5

3 2

3

4

20

23

26

29 212 215 218 221

1

Repetition Rate

4

16

64

Repetition Rate

256

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(f) Dolma stack (CE Loss ↓)

10 9 8 7 6 5

Wikipedia CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

Wikipedia CE (↓)

Wikipedia CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6 5 4

20

23

26

29 212 215 218 221

10 9 8 7 6 5 4 3

1

Repetition Rate

Model type Dense (1 x 1) MoE (64 x 1/4)

4

16

64

Repetition Rate

256

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(g) Dolma wiki (CE Loss ↓)

10 9 8 7 6

6

5

5

4

20

23

26

29 212 215 218 221

Repetition Rate

Model type Dense (1 x 1) MoE (64 x 1/4)

ICE CE (↓)

10 9 8 7

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

20

ICE CE (↓)

ICE CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6 5 4 3

1

4

16

64

Repetition Rate

256

1024

(h) ICE (CE Loss ↓)

32

1

2

4

8

16 32 64 128 256

Repetition Rate

Data Scarcity and Model Sparsity

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7

10 9 8 7 6

Pile CE (↓)

Pile CE (↓)

Pile CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6 5

6

5

4

5

4

3

4

3

20

23

26

29 212 215 218 221

Repetition Rate

1

4

16

64

Repetition Rate

256

Model type Dense (1 x 1) MoE (64 x 1/4)

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(i) Pile (CE Loss ↓) 20

10 9 8 7 6

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

Model type Dense (1 x 1) MoE (64 x 1/4)

S2ORC CE (↓)

S2ORC CE (↓)

S2ORC CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6 5

5

23

26

29 212 215 218 221

3 1

Repetition Rate

5 4

4

20

10 9 8 7 6

4

16

64

Repetition Rate

256

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(j) S2ORC (CE Loss ↓) 20

10 9 8 7 6

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 9 8 7 6 5

20

23

26

29

212

215

Repetition Rate

218

221

10 9 8 7 6 5 4 3

4

5

Model type Dense (1 x 1) MoE (64 x 1/4)

WikiText-103 CE (↓)

WikiText-103 CE (↓)

WikiText-103 CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

1

4

16

64

Repetition Rate

256

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(k) Wikitext (CE Loss ↓)

Figure 23: Held-out language modeling loss largely follows the same patterns. For 80M, 200M, and 1B active parameter models trained on OLMoE mix, we show held-out language modeling loss on a wide variety of domains. Performance depends on the exact domain, but largely follows the trends shown in Figure 1.

33

Data Scarcity and Model Sparsity

B.10

D OWNSTREAM TASKS

We show results for the downstream tasks of Appendix A.4 on the models from §3, with the extended settings of Appendix B.2. We include CE Loss on all tasks, as well as accuracy on Hellaswag. Accuracy on all other tasks remains near-chance, even at 1B scale.

6 5

BoolQ CE (↓)

BoolQ CE (↓)

10

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

4 3

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

5

Model type Dense (1 x 1) MoE (64 x 1/4)

5

BoolQ CE (↓)

8 7

3 2

2

20

23

26

29

3 2

1

212 215 218 221

1

Repetition Rate

4

16

64

Repetition Rate

256

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(a) BoolQ (CE Loss ↓) Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

3

HellaSwag CE (↓)

HellaSwag CE (↓)

3

2

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

4

2

1 0.9 0.8 0.7

1 0.9

20

23

26

29 212 215 218 221

1

Repetition Rate

4

16

64

Repetition Rate

256

Model type Dense (1 x 1) MoE (64 x 1/4)

3

HellaSwag CE (↓)

4

2

1 0.9 0.8 0.7 0.6

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(b) Hellaswag (CE Loss ↓) 20

10

5 3

Model type

Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10 5 3 2

2 20

23

26

29 212 215 218 221

1

Repetition Rate

4

16

64

Repetition Rate

256

MMLU humanities CE (↓)

Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

MMLU humanities CE (↓)

MMLU humanities CE (↓)

Model type

9 8 7 6 5

Model type

Dense (1 x 1) MoE (64 x 1/4)

4 3 2

1024

1

2

4

8 16 32 64 128 256

Repetition Rate

20

10

5 3

10

5 3

Model type Dense (1 x 1) MoE (64 x 1/4)

7 6

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

MMLU other CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

MMLU other CE (↓)

MMLU other CE (↓)

(c) MMLU Humanities (CE Loss ↓)

5 4 3 2

2

2

20

23

26

29 212 215 218 221

Repetition Rate

1

4

16

64

Repetition Rate

256

1024

(d) MMLU Other (CE Loss ↓)

34

1

2

4

8

16 32 64 128 256

Repetition Rate

10

5 3

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10

5 3

Model type Dense (1 x 1) MoE (64 x 1/4)

7 6

MMLU social sci CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

MMLU social sci CE (↓)

MMLU social sci CE (↓)

Data Scarcity and Model Sparsity

5 4 3 2

2

2

20

23

26

29 212 215 218 221

1

Repetition Rate

4

16

64

Repetition Rate

256

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

10

5 3

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

10

5 3

Model type Dense (1 x 1) MoE (64 x 1/4)

7 6

MMLU STEM CE (↓)

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

MMLU STEM CE (↓)

MMLU STEM CE (↓)

(e) MMLU Social Sciences (CE Loss ↓)

5 4 3 2

2

2

20

23

26

29 212 215 218 221

1

Repetition Rate

4

16

64

Repetition Rate

256

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(f) MMLU Stem (CE Loss ↓)

0.26 0.2575 0.255 0.2525 0.25 0.2475

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

0.29

HellaSwag acc (↑)

HellaSwag acc (↑)

0.3

Model type Dense (1 x 1) MoE (32 x 1/4) MoE (64 x 1/4)

0.28 0.27 0.26

Model type Dense (1 x 1) MoE (64 x 1/4)

0.5

HellaSwag acc (↑)

0.265 0.2625

0.4

0.3

0.25

20 23 26 29 212 215 218 221

Repetition Rate

1

4

16

64

Repetition Rate

256

1024

1

2

4

8

16 32 64 128 256

Repetition Rate

(g) Hellaswag (Accuracy ↑)

Figure 24: Downstream task performance closely follows validation loss. For 80M, 200M, and 1B active parameter models, downstream task loss is subject to random noise, but loosely follows trends of validation loss.

35

Record · ID 673490 · SHA-256 2b1930be36df847d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.