Conceptio › Archive › arXiv CS
arXiv CSopen access

Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling

2026 · arxiv_cs
arXiv CS · Papers · License: Open Access · 2026
Open Source ↗Direct PDF ↓
knowledge-representationreasoning
artificial intelligence, reasoning, knowledge representation

Repetition over Diversity: High-Signal Data Filtering for Sample-Efficient German Language Modeling Ansar Aynetdinov

Patrick Haller

Alan Akbik

Humboldt-Universität zu Berlin {aynetdia, patrick.haller.1, alan.akbik}@hu-berlin.de

Abstract

significantly stronger performance than training on massive, unfiltered datasets (Du et al., 2022; PaLM Team, 2022; Gunasekar et al., 2023; Lozhkov et al., 2024). While this shift is well-documented for English, it presents a unique challenge for non-English high-resource languages such as German, French, Japanese, and Chinese. These languages possess substantial web corpora containing hundreds of billions of tokens, but they lack the multitrillion-token abundance of English. In these data-constrained scenarios, strict filtering creates a strategic dilemma: should practitioners prioritize diversity by applying only light filters to maintain large token pools (e.g. 100B tokens), or should semantic density, i.e. expected training signal per token, be prioritized by applying strict quality filters to a smaller subset (e.g. 25B tokens) and repeating it over multiple epochs? This Study. We investigate this trade-off using German as a representative case study for non-English high-resource languages. We filter raw web corpora through three hierarchical qualitative tiers: Coherence, which removes structural noise to ensure syntactic flow; Information Value, which retains only content-rich and fact-bearing documents; and Educational Quality, which selects for pedagogical clarity and explanatory depth. We further derive a Dense Core subset – the intersection of these criteria – to represent the upper bound of semantic density available in German web data. By training language models from scratch on these filtered subsets, we isolate the impact of each semantic property on downstream performance and test whether the cautious approach to multi-epoch training prevalent in the literature (Muennighoff et al., 2023; Faysse et al., 2025; Luukkonen et al., 2023) remains warranted when data selection is optimized for knowledge density. It also enables us to explore training curricula that transition from diverse data pools to high-density subsets, testing

arXiv:2604.28075v1 [cs.CL] 30 Apr 2026

Recent research has shown that filtering massive English web corpora into high-quality subsets significantly improves training efficiency. However, for high-resource non-English languages like German, French, or Japanese, aggressive filtering creates a strategic dilemma: should practitioners prioritize diversity by training once on large amounts of lightly filtered web data, or prioritize quality by strictly filtering for a high-quality core and repeating it over multiple epochs? We investigate this trade-off for German by constructing hierarchical quality filters applied to 500M web documents, comparing multi-epoch training on the filtered subsets against singlepass training on a diverse corpus. Our experiments across multiple model scales and token budgets show that repeating high-quality data consistently outperforms single-pass training on larger, less filtered sets. Notably, the performance gap persists even after 7 epochs. Our findings suggest that for non-English LLMs, semantic concentration through quality filtering offers a more viable path to efficient language modeling than simply maximizing unique data volume. We release our German language models (called B OLDT), as well as our cleaned evaluation benchmarks to the research community. Our experiments indicate that they achieve state-of-the-art results despite training on 10360x fewer tokens than comparable models.

1

Introduction

The prevailing scaling laws for large language models (LLMs) emphasize a simple triad: more parameters, more compute, and more data (Kaplan et al., 2020; Hoffmann et al., 2022). However, as the field moves towards more sample-efficient training, the "more is better" paradigm has been increasingly challenged by "quality-first" approaches. Previous works demonstrated that training LLMs on web corpora filtered for high-quality content can yield 1

whether data quality staging provides benefits over static data mixtures. Contributions. This work makes the following contributions:

the creation of FineWeb and its more restrictive subset, FineWeb-Edu (Lozhkov et al., 2024), which uses a classifier to identify documents with high educational value. Unfortunately, these techniques remain less explored for non-English languages where strict filtering may critically reduce the available token budget. Overcoming the data bottleneck. As high-quality data demand outpaced the growth of available corpora, practitioners incorporated code, social media, books, and scientific papers into pre-training mixtures (Gao et al., 2020; Soldaini et al., 2024), and explored synthetic data generation for pre-training data enrichment (Gunasekar et al., 2023; Li et al., 2023; Kang et al., 2025), despite the dangers of model collapse (Shumailov et al., 2024). Muennighoff et al. (2023) investigated scaling laws in data-constrained regimes, concluding that the returns from repeated exposure to the same dataset diminish after 4 epochs. Fang et al. (2026) extend this by showing that repeatedly training on an aggressively filtered dataset for up to 10 epochs can outperform a single pass over a 10x larger unfiltered superset. Our work builds upon these findings by investigating whether strict quality filtering presents a practical obstacle in non-English settings where the unfiltered data pool is substantially smaller, using German as a case study. Non-English Language Modeling. While English dominates the LLM landscape, several efforts have targeted monolingual or highly specialized multilingual models for high-resource languages like Japanese (LLM-jp Team, 2024) and German (Pfister et al., 2025). A lot of the progress in nonEnglish language modeling, however, comes from either continued-pre-training of English-centric models (Zheng et al., 2024; Hoffmann et al., 2025; Kuulmets et al., 2024) or pre-training on massive multilingual dataset mixtures (BigScience Workshop and Bloom Team, 2023; Gemma 3 Team, 2025; Llama 3 Team, 2024; Qwen3 Team, 2025). In the context of German, two recent projects have proposed specialized filtering strategies for the German subset of the multilingual FineWeb-2 dataset (Penedo et al., 2025). Burns et al. (2025) developed Aleph-Alpha-GermanWeb (AA-High), using a multi-dimensional classifier to bucket documents by grammar and style. Separately, Messmer et al. (2025) introduced FW2-MKC, which uses an embedding-based approach to select web documents that align with established knowledge benchmarks and instruction-tuning sets. In our ex-

• We demonstrate that repeated training on highquality German data outperforms maximizing unique web document coverage under a fixed pre-training budget. We specifically analyze the distinct impacts of our three semantic tiers: Coherence, Information Value, and Educational Quality. • We show that multi-epoch pre-training on filtered data positively impacts instruction-tuning behavior, leading to higher correctness in assistant tasks. • We identify and correct significant noise in existing German translation of ARC-Challenge, Hellaswag, Lambada, and OpenBookQA benchmarks. We release these corrected and filtered benchmarks to ensure more reliable assessment of LLM performance in the German NLP community. • We release a series of German Small Language Models (SLMs) up to 1B parameters. Despite being trained on an order of magnitude fewer tokens than competitive baselines, our models achieve state-of-the-art results for their size, underscoring the importance of high-quality data selection in resource-constrained settings.

2

Related Work

Data Filtering for LLMs. Since GPT-2 (Radford et al., 2019) and T5 (Raffel et al., 2023), filtered CommonCrawl (Common Crawl Foundation, 2024) has been central to LLM pre-training. Datasets like C4 (Raffel et al., 2020), The Pile (Gao et al., 2020), OSCAR(Ortiz Suárez et al., 2020), and RefinedWeb (Penedo et al., 2023) employed deduplication, language filtering, and heuristics to remove noisy data. GPT-3 (GPT-3 Team, 2020) and Gopher (Team, 2022) further demonstrated that classifier-based quality filtering significantly improves outcomes. Recent work has further refined model-based filtering: the Phi models (Gunasekar et al., 2023; Li et al., 2023) showed that "textbook-quality" data dramatically improves sample efficiency, while Penedo et al. (2024) further scaled scaled this with 2

Subset FW2-DE (Full Pool)

N Docs (Millions)

Yield (%)

Token Count (Tokens)

Doc Length (µ ± σ)

Tokenizer Fertility Train Test Benchmarks

496.0

100.0

-

-

-

-

-

Hierarchical Tiers R ANDOM (Baseline) C OHERENCE I NFORMATION VALUE E DUCATIONAL Q UALITY

128.9* 300.6 (138.3*) 43.5 30.2

26.0* 60.6 (27.8*) 8.8 6.1

100B* 100B* 65B 33B

786 ± 1725 730 ± 1540 1494 ± 2561 1087 ± 2103

1.49 1.48 1.36 1.33

1.48 1.50 1.42 1.40

1.57 1.56 1.42 1.38

Target Core D ENSE C ORE (Intersection)

24.5

5.1

28B

1150 ± 2193

1.32

1.40

1.38

External Baselines FW HQ (Messmer et al., 2025) AA High (Burns et al., 2025)

43.2 70.4

8.7 14.2

35B 21B

823 ± 1895 296 ± 368

1.33 1.35

1.40 1.40

1.38 1.43

Table 1: Dataset statistics for the German split of FineWeb-2 (FW2-DE) and derived subsets. Yield represents the percentage of documents retained from the original pool. Doc Length is expressed in tokens. C OHERENCE reports document statistics for both the full subset and the 100B sample (denoted with *) used in further experiments.

periments, we use these two datasets as external benchmarks to compare against our own hierarchical qualitative filters.

3

textbook-like clarity and pedagogical value, selecting for documents that explain concepts or provide structured knowledge suitable for a curriculum. D ENSE C ORE. By taking the intersection of these three filters, we derive D ENSE C ORE. This subset contains documents that are simultaneously coherent, information-dense, and educational. As shown in Table 1, this subset represents a significantly "denser" version of the German web, which we use as our primary high-quality core for multi-epoch training.

Data and Filtering

Our data pipeline is built upon the German split of the FineWeb-2 (FW2-DE) dataset (Penedo et al., 2025). While FW2-DE provides a foundation of deduplicated web content, its quality-based filtering is rather basic. To enable our study of semantic density, we define a hierarchical filtering framework that moves from structural surface features to deep pedagogical value. 3.1

3.2

Dataset Statistics and Tokenization

To ensure fair comparison across data regimes, we train dedicated BPE tokenizers (vocabulary size 32,000) for each subset, optimizing for their specific lexical distributions. Analysis. Table 1 shows that while the C OHER ENCE filter is relatively permissive and retains the majority of the documents from the original FW2DE pool, the E DUCATIONAL Q UALITY filter is substantially more selective. As selection strictness increases, average document length rises markedly – the D ENSE -C ORE exhibits nearly 50% longer documents than R ANDOM, reflecting substantive long-form content rather than fragmented web snippets. As for the tokenizer efficiency, we see a clear trend indicating that discarding lower-quality data from the tokenizer training data improves the resulting tokenizer fertility both on the respective train splits, as well as the shared holdout FW2-DE test set. We also report average tokenizer fertility on question prompts from the benchmarks in the evaluation suite used in our experiments, and observe the same trend there.

Hierarchical Qualitative Filters

We implement three document-level classifiers to score the FW2-DE pool. These filters are not mutually exclusive but represent increasing tiers of selection strictness. Details pertaining to the classifier training pipeline are provided in Appendix A. C OHERENCE. This filter targets basic linguistic and structural integrity. It is designed to remove "word-salad" documents, truncated HTML exports, and fragmentary snippets. A high coherence score indicates a document with natural syntactic flow, regardless of whether the content is informative. I NFORMATION VALUE. This tier selects for "signal density." It identifies documents that are factbearing and content-rich, like e.g. technical reports, news articles, or specialized documentation, while filtering out generic web prose, SEO-heavy landing pages, and repetitive boilerplate. E DUCATIONAL Q UALITY. This is our most restrictive tier, modeled after the criteria used in FineWeb-Edu (Penedo et al., 2024). It prioritizes 3

Benchmark

EN (Original): [...] Both Olaf and the boy heard it and walked towards Bob

Global MMLU ARC-Challenge ARC-Easy OpenBookQA HellaSwag LAMBADA

DE (Old): [...] Sowohl Olaf als auch der Junge hörten es und gingen auf Bob zu DE (Fixed): [...] Sowohl Olaf als auch der Junge hörten es und gingen auf Bob zu

Modernized

14,042 1,168 2,376 500 10,042 5,153

0 0 6 5 47 3

0 1,168 2,370 495 9,995 5,150

in which machine translation failed or did not preserve the intended task logic due to word order discrepancies between English and German. Table 2 summarizes our modernized suite.

Evaluation Datasets

A significant challenge in assessing LLMs in nonEnglish languages is the lack of broad-range evaluation benchmarks. While many seminal benchmarks were devised in English by human experts, their non-English counterparts are often simply machine-translated without any attention to language-specific rules, often produced by older, less capable models (Lai et al., 2023; Bellagente et al., 2024). As a result, these versions often contain non-idiomatic phrasing, grammatical errors, and, more critically, structural artifacts that fundamentally alter the difficulty of the original task. As a part of this work, we contribute a modernized and cleaned suite of German benchmarks designed to provide a more reliable signal for model performance. 4.1

Removed

Table 2: Statistics for our German evaluation suite. All benchmarks except Global MMLU were re-translated from their original English version. Removed items indicate instances where translation itself failed or broke the task integrity.

Figure 1: Example of a problematic instance from the German translation of the LAMBADA benchmark in EleutherAI’s Evaluation Harness (Gao et al., 2024). We highlight correct and incorrect target labels.

4

Original

4.2

Final Evaluation Suite

The resulting evaluation suite provides a holistic perspective on model performance across factual knowledge, commonsense reasoning, and linguistic context tracking. Factual knowledge is assessed via the German subset of Global MMLU (Singh et al., 2025), which spans 57 subjects across STEM, the social sciences, and the humanities. Reasoning capabilities are measured using ARC-Easy, ARC-Challenge (Clark et al., 2018), and OpenBookQA (Mihaylov et al., 2018), which require models to perform grade-school level science inference. Finally, we evaluate commonsense narrative continuation and discourse-level context tracking through HellaSwag (Zellers et al., 2019) and the OpenAI version of LAMBADA (Paperno et al., 2016). For all evaluations, we rely on lengthnormalized conditional log-likelihoods to choose the correct multiple-choice answer. Table 2 summarizes the composition of the modernized suite and the impact of our manual cleaning process. Notably, the exclusions required to maintain task integrity were minimal, with the largest adjustment occurring in HellaSwag due to its larger size in general. We release these cleaned benchmarks to the community to facilitate more rigorous and standardized assessment in German NLP.

Addressing Translation Artifacts

In our analysis of existing German benchmark variants, we identified the main issue with completionbased tasks, often used for evaluation of pre-trained LLMs: discrepancies in the word order between English and German. In completion tasks like HellaSwag and LAMBADA, German’s verb-final structures and flexible word order frequently move the original "completion target" away from sentence-final position, disrupting the intended prediction task. Refer to Figure 1 for an example of this issue: instead of tracking the entity "Bob" across sentences, the task becomes predicting a verb ending two words away. To address this issue, we re-translated the suite using the state-of-the-art machine translation Tower+ 72B model (Rei et al., 2025). Instead of translating benchmarks instance components separately, we provide full instances as inputs to the model. We discarded the rare instances (<0.5%),

5

Investigating the Quality-Quantity Trade-off

In this section, we present a series of experiments designed to isolate the impact of data density on model performance. We move away from the traditional "single-pass" pre-training paradigm to in4

Subset

Tokens

MMLU

ARC-C

ARC-E

H-Swag

LAMBADA

OBQA

Avg.

100B (1.0x)

27.13

26.15

41.10

37.18

33.52

41.01

34.35

100B (1.0x) 65B (1.5x) 33B (3.0x)

27.06 28.29 28.64

27.65 30.46 31.49

42.74 46.20 50.91

40.45 40.71 40.57

38.33 38.52 36.49

42.63 44.04 43.64

36.48 38.04 38.62

28B (3.6x) 35B (2.9x) 21B (4.8x)

28.97 28.00 26.55

31.40 27.37 25.31

50.55 46.37 39.83

41.10 39.49 37.39

37.55 40.43 29.75

45.86 42.02 40.00

39.24 37.28 33.14

100B (1.0x) 78B (1.3x)

28.00 28.52

30.18 29.34

46.25 47.89

40.25 40.35

35.98 36.06

42.63 43.67

37.22 37.64

1T 6T∗ 36T∗

26.41 26.09 29.87

24.74 24.93 32.90

37.13 34.68 41.90

23.02 31.60 38.13

26.97 29.85 39.57

43.03 37.37 41.01

33.33 32.07 37.23

Baseline

R ANDOM Single Filters

C OHERENCE I NFORMATION VALUE E DUCATIONAL Q UALITY Filter Combinations

D ENSE -C ORE MKC (Messmer et al., 2025) AA H IGH (Burns et al., 2025) Curriculum-Based (100B Budget)

S ORTED P HASED Reference Models (Total Tokens Trained)

LLäMmlein-120M (Pfister et al., 2025) Gemma-3-270M (Gemma 3 Team, 2025) Qwen-3-0.6B-Base (Qwen3 Team, 2025)

Table 3: Benchmark results for 350M models. Tokens refers to unique tokens in the subset, while the bracketed value indicates the number of epochs to reach the 100B budget. Token counts of multilingual data mixtures are denoted with ∗ .

vestigate whether high-quality repetition can compensate for a lack of unique token diversity. Our investigation is divided into four parts: a study of token allocation strategies under a fixed 100B token budget, a scaling analysis in the 1B parameter regime, exploration of the repetition ceiling, and an evaluation of the downstream impact on instruction tuning. 5.1

100B limit from the full FW2-DE and C O HERENCE pools of documents respectively. 2. Dense Repetition: Multi-epoch training on the remaining, higher-quality subsets. For instance, for the 28B-token D ENSE C ORE core, the 100B budget results in approximately 3.6 training epochs. We also consider the qualityfiltered MKC and AA High subsets of FW2DE devised independently from this work by Messmer et al. (2025) and Burns et al. (2025) respectively. 3. Hybrid Curricula: "Coarse-to-fine" schedules designed to test whether finishing pretraining on higher-quality data can leverage both diversity and quality. The motivation for these curricula stems from the intuition that models might benefit from broad initial exposure to diverse linguistic patterns available in lower-quality data before refining on concentrated knowledge. We implement two variants: a P HASED 50/50 split (50B tokens of R ANDOM followed by 50B tokens of D ENSE C ORE) and a S ORTED schedule, where 100B C OHERENCE tokens are presented in ascending order of their Educational score.

Experiment I: Token Allocation Strategies (100B Budget)

The primary objective of this experiment is to determine the optimal way to distribute a fixed training budget of 100B tokens when faced with a choice between broad, unique web data and narrow, highdensity educational content. 5.1.1 Experimental Setup We utilize a decoder-only transformer architecture following the Llama model family (Llama 2 Team, 2023), with a primary model size of 350M nonembedding parameters. To ensure a fair comparison, all models in this study are restricted to a total exposure of 100B tokens. Details on hyperparameter and optimizer choice are provided in Appendix B. To test the interaction between diversity and density, we define three training strategies:

5.1.2 Results and Analysis Table 3 demonstrates significant gains from prioritizing semantic density over document diversity. D ENSE -C ORE, outperforms the R ANDOM baseline by 4.89 points on average while training for approximately 3.6 epochs over 28B unique tokens,

1. Uniform Baselines: Single-pass training on unique documents from the R ANDOM or C O HERENCE pools. For both baselines we randomly sample documents until we reach the 5

suggesting that for information-dense German text, high signal-to-noise ratio of the curated content outweigh potential multi-epoch repetition risks. Training dynamics and checkpoint analysis. Figure 2 shows that D ENSE -C ORE maintains a consistent advantage throughout the entire 100B token trajectory rather than only at convergence. This indicates that the benefit of high-quality repetition is not confined to the final stages of training, but rather provides a steeper learning curve from the onset. The poor performance of the AA H IGH subset (repeated 4.8×) indicates that multi-epoch training only benefits from sufficiently high-quality data. We attribute its poor downstream impact to usage of smaller and by proxy weaker models as annotators and classifiers for its filtering. The P HASED and S ORTED curricula on the other hand show a marked inflection point: their performance improves markedly in the second half of training, coinciding with the transition to higherquality data. While these curricula eventually reach high performance, they consistently trail the pure D ENSE -C ORE trajectory. This suggests that the inclusion of low-signal web tokens, even as an initial "warm-up" phase, may dilute the overall information density of a fixed 100B token budget. Comparison to state-of-the-art models. We note that our 350M D ENSE -C ORE model outperforms LL Ä M MLEIN -120M (Pfister et al., 2025), G EMMA -3-270M (Gemma 3 Team, 2025), and Q WEN -3-0.6B-BASE (Qwen3 Team, 2025) on our evaluation suite, despite these models being trained on 10×, 60×, and 360× more tokens, respectively. This suggests that targeted high-signal dataset cores enable performance typically requiring far larger compute and data budgets. 5.2

Figure 2: Zero-shot evaluation of 350M models trained on 100B tokens of different FW2-DE subsets over pretraining checkpoints.

5.2.2 Results and Analysis Figure 3 shows that performance gap between the R ANDOM and D ENSE -C ORE not only persists but widens as parameter count increases. While the 350M D ENSE -C ORE outperformed its baseline by 4.89 points on average, the 1B D ENSE -C ORE achieves a 5.14-point lead over the 1B R ANDOM baseline. Furthermore, despite being trained on an order of magnitude less tokens in total, the 1B D ENSE C ORE model achieves parity with or exceeds the performance of multilingual G EMMA -3-1B and L LAMA -3.2-1B, and outperforms monolingual LL Ä M MLEIN -1B. These findings suggest that as model capacity increases, the transformer architecture becomes more adept at incorporating information from highdensity sources. In this regime, a 100B-token budget is not a limiting factor for reaching state-ofthe-art performance for 1B models, provided the training signal is sufficiently concentrated. This indicates that for non-English high-resource languages like German, the path to high-performance small models lies in maximizing token utility rather than volume or cross-lingual transfer.

Experiment II: Parameter Scaling and Efficiency

To verify that the performance gains observed in the 350M regime generalize as model capacity increases, we scale our primary experiments to 1B parameter models (see Appendix B for details). 5.2.1

Experimental Setup

In this experiment, we focus on the two most distinct strategies: the R ANDOM baseline (representing high-coverage, single-pass training) and the D ENSE -C ORE (representing high-density repetition). Both models are trained on the same 100B token budget, using the same hyperparemeters as in Experiment I.

5.3

Experiment III: Exploring Repetition Limits (200B Budget)

While Experiment I established that 3.6 epochs of repetition on high-density data is superior to a sin6

Figure 3: Performance gain when scaling the model size from 350M to 1B parameters. Table 8 breaks down the performance of the models depicted in the graph on each benchmark.

Figure 4: Performance of 350M models trained on 100B and 200B tokens of Random and Edu High subsets. Table 7 breaks down the performance of each model’s final checkpoint on each benchmark.

gle pass on diverse data, it remains unclear where the point of diminishing returns lies for various data regimes. In this experiment, we explore the limits of multi-epoch training on high-density subsets by using an extended 200B token budget, examining at what point the benefits of data density are offset by the lack of novelty, and comparing this trajectory to the scaling behavior of lower-density sets over the same budget.

D ENSE -C ORE model maintains a significant lead over the R ANDOM model, even though the latter is still seeing entirely new unique data. While the P HASED curriculum continues to scale and narrow the gap to the pure D ENSE -C ORE model effectively, it never quite overtakes it, suggesting that training on a more diverse sample in the first stage holds back optimization on a more dense subset. To investigate whether these trends hold at larger model scales, we conduct an additional pre-training run of a 1B model on 200B tokens of D ENSE C ORE. Table 8 shows that this model achieves an uplift of 2.08 points on average on our evaluation suite, compared to pre-training on 100B tokens. Notably, this gain is more than twice the corresponding improvement observed for the 350M model. These results indicate that the effectiveness of repeated exposure to high-quality data depends on model size and may extend to longer training horizons as parameter count increases if the underlying corpus is sufficiently information-dense. In this regime, larger models can continue to extract additional value from repeated passes over curated data, before the requirement for larger pools of unique tokens ultimately becomes the limiting factor.

5.3.1 Experimental Setup We extend the training of three key configurations from Experiment I to a total of 200B tokens: R AN DOM , P HASED , and D ENSE -C ORE . For the R AN DOM baseline this represents a transition from a 100B single-pass to a 200B single-pass (adding new documents into the mix). For the D ENSE C ORE, the 200B budget results in approximately 7.2 training epochs over the 28B-token subset. The P HASED curriculum consists of the 100B R AN DOM subset folllowed by 100B tokens of continued exposure to the D ENSE -C ORE data. We maintain the same hyperparameters as in Experiment 1 to ensure comparable learning dynamics across all runs. 5.3.2 Results and Analysis Figure 4 illustrates that the benefits of multi-epoch training on D ENSE -C ORE subset persist well beyond the four-epoch threshold identified by Muennighoff et al. (2023) – extending the findings of Fang et al. (2026) to a non-English setting where the unfiltered data pool is severely more constrained. We do not observe a loss in generalization on our benchmark suite even when doubling the number of epochs on the same dataset. Instead, the

5.4

Experiment IV: Generalization to Instruction Tuning

Finally, we evaluate whether the advantages of a high-density pre-training "core" translate into improved instruction-following capabilities. 5.4.1

Experimental Setup

We fine-tune the final checkpoints of our models on the German subset of the S MOLTALK 2 instruction7

tuning dataset (Bakouch et al., 2025). Both the 350M and 1B parameter variants are tuned using identical hyperparameters (see Appendix B). To evaluate the models, we utilize an LLM-as-a-judge protocol (see Figures 7 and 8 for the prompt templates) using L LAMA -3.3-70B-I NSTRUCT (Llama 3 Team, 2024) on 1,000 held-out prompts. Table 4 reports the average Likert-scale score (110) and the total number of correct answers (binary accuracy) across the test set.

Subset

350M params @ 100B tokens R ANDOM 1.0x 5.25 C OHERENCE 1.0x 5.62 I NFO . VALUE 1.5x 5.69 E DU . Q UALITY 3.2x 5.75 D ENSE -C ORE 3.6x 5.74 MKC 1.0x 5.63 AA H IGH 4.8x 5.41 S ORTED 1.0x 5.61 P HASED 1.3x 5.66

178 226 249 241 253 219 199 225 231

5.4.2

1B params @ 100B tokens R ANDOM 1B 1.0x D ENSE -C ORE 1B 3.6x

5.87 6.13

293 338

350M params @ 200B tokens R ANDOM 1.0x 5.45 D ENSE -C ORE 7.2x 5.96 P HASED 2.6x 5.78

198 278 251

Results and Analysis

Table 4 confirms that the "density advantage" is preserved through the SFT process. For the 350M models at the 100B token budget, the D ENSE C ORE configuration achieves the highest count of correct responses (253/1,000), closely followed by I NFORMATION VALUE (249/1,000), significantly exceeding R ANDOM baseline (178/1,000). Interestingly, the MKC subset, which was filtered for documents similar to instruction-tuning datasets like Aya (Singh et al., 2024), fails to outperform our manually defined D ENSE -C ORE. This suggests that the fundamental reasoning capabilities and factual grounding provided by high-quality, information-rich data are more important for SFT than simply matching the formatting or distribution of typical instruction sets. The poor performance of the AA H IGH subset (199 correct answers) further reinforces that repetition only benefits sufficiently high-quality data. Scaling in Parameters and Tokens. The benefits of density scale effectively across both parameter counts and token budgets. Our 1B D ENSE -C ORE model achieves a score of 6.13 score with 338 correct answers. Most strikingly, the 350M D ENSE C ORE model trained for 200B tokens (7.2 epochs) produces 278 correct answers – nearly matching the 1B Random model (293 correct) despite having 3x fewer parameters.

6

Table 4: LLM-as-a-Judge (Llama-3.3-70B) results. Tokens indicates the repetition factor over the unique subset. Score is the 1–10 Likert average; Correct is the binary accuracy count out of 1,000 prompts.

ing the D ENSE C ORE subset of FineWeb-2. These models are released for the purpose of reproducibility of this paper’s results. B OLDT-1B. A 1B parameter model trained on a combination of D ENSE C ORE and a corpus of 6B tokens of news data in German obtained using the F UNDUS library (Dallabetta et al., 2024). Our F UN DUS crawlers were used continuously since 2022, so the main body of news articles stems from the time period between 2022 and February 2026 (with smaller numbers of articles going back to 1994). B OLDT-1B was trained for multiple epochs on the combined corpus until reaching the effective training token count of 230B. The context window size is also extended from 2048 to 4096 compared to B OLDT-DC-1B. We compare our models against other similarlysized pre-trained models on our evaluation suite. The results are shown in Table 5. We find that B OLDT-1B improves upon B OLDT-DC-1B in all benchmarks except for ARC-Challenge and ARC-Easy. Despite being trained on a substantially smaller training corpus compared to relevant similarly-sized LLMs capable of German, our 1B models are competitive even with larger-sized (around 2B) multilingual models. In addition to our base models, we also release a preview of an instruction-tuned version of

Model and Benchmark Release

We release our trained models, called B OLDT, as well as the updated benchmarks from our evaluation suite to the research community1 . The initial release consists of three models: B OLDT-DC-350M & B OLDT-DC-1B. 350M and 1B parameter models trained as part of our repetition limits experiments (see Section 5.3). In total, they were trained on a 200B token budget us1

Tokens Score Correct

https://huggingface.co/Boldt

8

Model

Tokens

MMLU

ARC-C

ARC-E

H-Swag

LAMBADA

OBQA

Avg.

200B 200B 230B

29.29 31.06 31.45

32.24 35.99 33.83

52.87 57.30 55.95

43.21 48.69 48.78

37.48 42.80 44.72

45.86 48.48 52.32

40.16 44.05 44.51

1T 2T∗ 9T∗

29.26 30.01 28.58

30.27 30.55 29.90

48.19 47.89 40.51

44.80 43.43 40.07

44.89 41.71 44.31

47.27 45.05 44.04

40.78 39.77 37.90

4T∗ 36T∗ 2T∗ 2T∗

31.04 34.17 29.68 33.99

31.58 37.49 32.62 37.11

54.68 57.00 53.63 57.47

45.30 45.20 46.57 49.62

44.52 49.81 43.55 52.64

50.50 45.66 49.70 48.89

42.94 44.89 42.63 46.62

Ours

B OLDT-DC-350M B OLDT-DC-1B B OLDT-1B Reference models - 1B

LLäMmlein-1B (Pfister et al., 2025) Gemma-3-1B (Gemma 3 Team, 2025) Llama-3.2-1B (Llama 3 Team, 2024) Reference models - >1B

EuroLLM-1.7B (Martins et al., 2024) Qwen3-1.7B-Base (Qwen3 Team, 2025) BübleLM-2B (Delobelle et al., 2024) Gemma-2-2B (Gemma 2 Team, 2024)

Table 5: Benchmark results of B OLDT models compared to other pre-trained models of the similar size, as well as larger reference models.

B OLDT-1B. Our instruction tuning dataset consists of a combination of real and synthetic instructionoutput pairs from diverse sources. Please refer to the HF model card for further details.

7

In practice, our findings suggest that extensive quality filtering relying on clearly-defined rules and strong annotator models provides a viable path towards sample-efficient pre-training in non-English high-resource languages like German.

Conclusion Limitations

This work addresses a practical question for nonEnglish LLM development: is aggressive quality filtering worthwhile when total available text is limited, or does it discard too much useful diversity for pre-training? Our experiments suggest that quality filtering remains beneficial despite the smaller pool of available web data. Across the model sizes and token budgets we considered, models trained for multiple epochs on small, high-quality subsets consistently outperform those trained on larger, less strongly filtered mixtures. Within our setups, we do not observe an early saturation point where repeating high-quality data ceases to help. Additional passes over the filtered corpus continue to yield improvements, whereas adding more unfiltered data brings only limited gains. This indicates that careful filtering and multi-epoch training are a viable and effective strategy rather than a risky trade-off for high-resource non-English languages like German. Our results also show continuity between quality-first pre-training and instruction-tuning. Instruction-tuned assistants that were pre-trained on curated high-quality subsets outperform those pre-trained on more diverse, less filtered data in both correctness and helpfulness. Introducing lower-quality documents into pre-training mixtures decreased instruction-tuned model output quality.

Language scope. Our investigation focuses exclusively on German as a representative high-resource non-English language. While German’s web corpus characteristics (hundreds of billions of tokens but not trillions) likely generalize to languages like French, Japanese, and Chinese, languages with smaller corpora or different linguistic structures may exhibit different quality-quantity trade-offs. Future work should validate these findings across diverse language families. Model and compute scale. Our experiments are limited to models up to 1B parameters trained on at most 200B tokens. While our findings demonstrate clear efficiency gains at this scale, it remains unclear whether these quality-quantity trade-offs hold at substantially larger, industry-level, scales where both compute budgets and available high-quality data increase dramatically. Architecture selection. We focus exclusively on dense transformer architectures. Mixture-ofexperts models might exhibit different behaviors with respect to data quality and repetition. Similarly, we do not explore architectural innovations like alternative attention mechanisms that might exhibit different behavior during pre-training. Toxicity and bias assessment. We do not evaluate our models for toxic content generation, demo9

graphic biases, or harmful stereotypes. While our quality filters prioritize educational and informative content, which may reduce certain risks compared to unfiltered web data, we cannot guarantee that aggressive filtering eliminates problematic content or prevents biased model behavior. High-quality educational content can still encode societal biases, and repeated exposure during multi-epoch training might amplify rather than mitigate such issues. Future work should conduct comprehensive bias audits across demographic dimensions and toxicity benchmarks to understand how quality filtering strategies interact with model safety. Additionally, our LLM-as-judge evaluation for instruction-tuned models does not assess safety, appropriateness, or potential for harmful outputs. Despite these limitations, our work provides valuable evidence that quality filtering and multi-epoch training offer a practical path toward efficient pretraining for high-resource non-English languages, challenging the notion that maximizing unique token exposure should be the primary objective in data-constrained settings.

Morlon, Vaibhav Srivastav, Joshua Lochner, XuanSon Nguyen, Colin Raffel, Leandro von Werra, and Thomas Wolf. 2025. SmolLM3: smol, multilingual, long-context reasoner. https://huggingface.co/ blog/smollm3. Marco Bellagente, Jonathan Tow, Dakota Mahan, Duy Phung, Maksym Zhuravinskyi, Reshinth Adithyan, James Baicoianu, Ben Brooks, Nathan Cooper, Ashish Datta, Meng Lee, Emad Mostaque, Michael Pieler, Nikhil Pinnaparju, Paulo Rocha, Harry Saini, Hannah Teufel, Niccolo Zanichelli, and Carlos Riquelme. 2024. Stable lm 2 1.6b technical report. Preprint, arXiv:2402.17834. BigScience Workshop and Bloom Team. 2023. Bloom: A 176b-parameter open-access multilingual language model. Preprint, arXiv:2211.05100. Thomas F Burns, Letitia Parcalabescu, Stephan Wäldchen, Michael Barlow, Gregor Ziegltrum, Volker Stampa, Bastian Harren, and Björn Deiseroth. 2025. Aleph-alpha-germanweb: Improving germanlanguage llm pre-training with model-based data curation and synthetic data generation. Preprint, arXiv:2505.00022. Peter Clark, Isaac Cowhey, Oren Etzioni, Tushar Khot, Ashish Sabharwal, Carissa Schoenick, and Oyvind Tafjord. 2018. Think you have solved question answering? try arc, the ai2 reasoning challenge. arXiv:1803.05457v1.

Acknowledgments The authors are supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Emmy Noether grant “Eidetic Representations of Natural Language” (project number 448414230). Further, Alan Akbik is supported by the Deutsche Forschungsgemeinschaft (DFG, German Research Foundation) under Germany’s Excellence Strategy "Science of Intelligence" (EXC 2002/1, project number 390523135). The authors gratefully acknowledge the scientific support and HPC resources provided by the Erlangen National High Performance Computing Center (NHR@FAU) of the Friedrich-AlexanderUniversität Erlangen-Nürnberg (FAU) under the NHR project c106fa. NHR funding is provided by federal and Bavarian state authorities. NHR@FAU hardware is partially funded by the German Research Foundation (DFG) – 440719683.

Common Crawl Foundation. 2024. Common crawl dataset. Web dataset. Max Dallabetta, Conrad Dobberstein, Adrian Breiding, and Alan Akbik. 2024. Fundus: A simple-to-use news scraper optimized for high quality extractions. In Proceedings of the 62nd Annual Meeting of the Association for Computational Linguistics (Volume 3: System Demonstrations), pages 305–314, Bangkok, Thailand. Association for Computational Linguistics. Pieter Delobelle, Alan Akbik, et al. 2024. Büblelm: A small german lm.

References

Nan Du, Yanping Huang, Andrew M. Dai, Simon Tong, Dmitry Lepikhin, Yuanzhong Xu, Maxim Krikun, Yanqi Zhou, Adams Wei Yu, Orhan Firat, Barret Zoph, Liam Fedus, Maarten Bosma, Zongwei Zhou, Tao Wang, Yu Emma Wang, Kellie Webster, Marie Pellat, Kevin Robinson, Kathleen Meier-Hellstern, Toju Duke, Lucas Dixon, Kun Zhang, Quoc V Le, Yonghui Wu, Zhifeng Chen, and Claire Cui. 2022. Glam: Efficient scaling of language models with mixture-of-experts. Preprint, arXiv:2112.06905.

Elie Bakouch, Loubna Ben Allal, Anton Lozhkov, Nouamane Tazi, Lewis Tunstall, Carlos Miguel Patiño, Edward Beeching, Aymeric Roucher, Aksel Joonas Reedi, Quentin Gallouédec, Kashif Rasul, Nathan Habib, Clémentine Fourrier, Hynek Kydlicek, Guilherme Penedo, Hugo Larcher, Mathieu

Alex Fang, Hadi Pouransari, Matt Jordan, Alexander T Toshev, Vaishaal Shankar, Ludwig Schmidt, and Tom Gunter. 2026. Datasets, documents, and repetitions: The practicalities of unequal data quality. In The Thirty-ninth Annual Conference on Neural Information Processing Systems.

10

Manuel Faysse, Patrick Fernandes, Nuno M Guerreiro, António Loison, Duarte Miguel Alves, Caio Corro, Nicolas Boizard, João Alves, Ricardo Rei, Pedro Henrique Martins, Antoni Bigata Casademunt, François Yvon, Andre Martins, Gautier Viaud, CELINE HUDELOT, and Pierre Colombo. 2025. CroissantLLM: A truly bilingual french-english language model. Transactions on Machine Learning Research.

Feiyang Kang, Newsha Ardalani, Michael Kuchnik, Youssef Emad, Mostafa Elhoushi, Shubhabrata Sengupta, Shang-Wen Li, Ramya Raghavendra, Ruoxi Jia, and Carole-Jean Wu. 2025. Demystifying synthetic data in LLM pre-training: A systematic study of scaling laws, benefits, and pitfalls. In Proceedings of the 2025 Conference on Empirical Methods in Natural Language Processing, pages 10739–10758, Suzhou, China. Association for Computational Linguistics.

Leo Gao, Stella Biderman, Sid Black, Laurence Golding, Travis Hoppe, Charles Foster, Jason Phang, Horace He, Anish Thite, Noa Nabeshima, Shawn Presser, and Connor Leahy. 2020. The pile: An 800gb dataset of diverse text for language modeling. Preprint, arXiv:2101.00027.

Jared Kaplan, Sam McCandlish, Tom Henighan, Tom B. Brown, Benjamin Chess, Rewon Child, Scott Gray, Alec Radford, Jeffrey Wu, and Dario Amodei. 2020. Scaling laws for neural language models. Preprint, arXiv:2001.08361.

Leo Gao, Jonathan Tow, Baber Abbasi, Stella Biderman, Sid Black, Anthony DiPofi, Charles Foster, Laurence Golding, Jeffrey Hsu, Alain Le Noac’h, Haonan Li, Kyle McDonell, Niklas Muennighoff, Chris Ociepa, Jason Phang, Laria Reynolds, Hailey Schoelkopf, Aviya Skowron, Lintang Sutawika, Eric Tang, Anish Thite, Ben Wang, Kevin Wang, and Andy Zou. 2024. The language model evaluation harness.

Hele-Andra Kuulmets, Taido Purason, Agnes Luhtaru, and Mark Fishel. 2024. Teaching llama a new language through cross-lingual knowledge transfer. In Findings of the Association for Computational Linguistics: NAACL 2024, pages 3309–3325, Mexico City, Mexico. Association for Computational Linguistics.

Gemma 3 Team. 2025. Gemma 3 technical report. Preprint, arXiv:2503.19786.

Viet Dac Lai, Chien Van Nguyen, Nghia Trung Ngo, Thuat Nguyen, Franck Dernoncourt, Ryan A. Rossi, and Thien Huu Nguyen. 2023. Okapi: Instructiontuned large language models in multiple languages with reinforcement learning from human feedback. Preprint, arXiv:2307.16039.

GPT-3 Team. 2020. Language models are few-shot learners. In Advances in Neural Information Processing Systems, volume 33, pages 1877–1901. Curran Associates, Inc.

Yuanzhi Li, Sébastien Bubeck, Ronen Eldan, Allie Del Giorno, Suriya Gunasekar, and Yin Tat Lee. 2023. Textbooks are all you need ii: phi-1.5 technical report. Preprint, arXiv:2309.05463.

Suriya Gunasekar, Yi Zhang, Jyoti Aneja, Caio César Teodoro Mendes, Allie Del Giorno, Sivakanth Gopi, Mojan Javaheripi, Piero Kauffmann, Gustavo de Rosa, Olli Saarikivi, Adil Salim, Shital Shah, Harkirat Singh Behl, Xin Wang, Sébastien Bubeck, Ronen Eldan, Adam Tauman Kalai, Yin Tat Lee, and Yuanzhi Li. 2023. Textbooks are all you need. Preprint, arXiv:2306.11644.

Llama 2 Team. 2023. Llama 2: Open foundation and fine-tuned chat models. Preprint, arXiv:2307.09288.

Jordan Hoffmann, Sebastian Borgeaud, Arthur Mensch, Elena Buchatskaya, Trevor Cai, Eliza Rutherford, Diego de Las Casas, Lisa Anne Hendricks, Johannes Welbl, Aidan Clark, Thomas Hennigan, Eric Noland, Katherine Millican, George van den Driessche, Bogdan Damoc, Aurelia Guy, Simon Osindero, Karén Simonyan, Erich Elsen, Oriol Vinyals, Jack Rae, and Laurent Sifre. 2022. An empirical analysis of compute-optimal large language model training. In Advances in Neural Information Processing Systems, volume 35, pages 30016–30030. Curran Associates, Inc.

Ilya Loshchilov and Frank Hutter. 2017. SGDR: Stochastic gradient descent with warm restarts. In International Conference on Learning Representations.

Gemma 2 Team. 2024. Gemma 2: Improving open language models at a practical size. Preprint, arXiv:2408.00118.

Llama 3 Team. 2024. The llama 3 herd of models. Preprint, arXiv:2407.21783. LLM-jp Team. 2024. Llm-jp: A cross-organizational project for the research and development of fully open japanese llms. Preprint, arXiv:2407.03963.

Ilya Loshchilov and Frank Hutter. 2019. Decoupled weight decay regularization. In International Conference on Learning Representations. Anton Lozhkov, Loubna Ben Allal, Leandro von Werra, and Thomas Wolf. 2024. Fineweb-edu: the finest collection of educational content. Risto Luukkonen, Ville Komulainen, Jouni Luoma, Anni Eskelinen, Jenna Kanerva, Hanna-Mari Kupari, Filip Ginter, Veronika Laippala, Niklas Muennighoff, Aleksandra Piktus, Thomas Wang, Nouamane Tazi, Teven Scao, Thomas Wolf, Osma Suominen, Samuli Sairanen, Mikko Merioksa, Jyrki Heinonen, Aija Vahtola, Samuel Antao, and Sampo Pyysalo. 2023.

Michael Hoffmann, Jophin John, Stefan Schweter, Gokul Ramakrishnan, Hoi-Fong Mak, Alice Zhang, Dmitry Gaynullin, and Nicolay J. Hammer. 2025. Llama-genba-10b: A trilingual large language model for german, english and bavarian. Preprint, arXiv:2509.05668.

11

FinGPT: Large generative models for a small language. In Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing, pages 2710–2726, Singapore. Association for Computational Linguistics.

Guilherme Penedo, Quentin Malartic, Daniel Hesslow, Ruxandra Cojocaru, Hamza Alobeidli, Alessandro Cappelli, Baptiste Pannier, Ebtesam Almazrouei, and Julien Launay. 2023. The refinedweb dataset for falcon llm: outperforming curated corpora with web data only. In Proceedings of the 37th International Conference on Neural Information Processing Systems, NIPS ’23, Red Hook, NY, USA. Curran Associates Inc.

Pedro Henrique Martins, Patrick Fernandes, João Alves, Nuno M. Guerreiro, Ricardo Rei, Duarte M. Alves, José Pombal, Amin Farajian, Manuel Faysse, Mateusz Klimaszewski, Pierre Colombo, Barry Haddow, José G. C. de Souza, Alexandra Birch, and André F. T. Martins. 2024. Eurollm: Multilingual language models for europe. Preprint, arXiv:2409.16235.

Jan Pfister, Julia Wunderle, and Andreas Hotho. 2025. LLäMmlein: Transparent, compact and competitive German-only language models from scratch. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 2227–2246, Vienna, Austria. Association for Computational Linguistics.

Bettina Messmer, Vinko Sabolčec, and Martin Jaggi. 2025. Enhancing multilingual LLM pretraining with model-based data selection. In The Thirty-ninth Annual Conference on Neural Information Processing Systems Datasets and Benchmarks Track.

Qwen3 Team. 2025. Qwen3 technical report. Preprint, arXiv:2505.09388.

Todor Mihaylov, Peter Clark, Tushar Khot, and Ashish Sabharwal. 2018. Can a suit of armor conduct electricity? a new dataset for open book question answering. In Proceedings of the 2018 Conference on Empirical Methods in Natural Language Processing, pages 2381–2391, Brussels, Belgium. Association for Computational Linguistics.

Alec Radford, Jeff Wu, Rewon Child, David Luan, Dario Amodei, and Ilya Sutskever. 2019. Language models are unsupervised multitask learners. Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2020. Exploring the limits of transfer learning with a unified text-to-text transformer. Journal of Machine Learning Research, 21(140):1–67.

Niklas Muennighoff, Alexander Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin A Raffel. 2023. Scaling data-constrained language models. In Advances in Neural Information Processing Systems, volume 36, pages 50358–50376. Curran Associates, Inc.

Colin Raffel, Noam Shazeer, Adam Roberts, Katherine Lee, Sharan Narang, Michael Matena, Yanqi Zhou, Wei Li, and Peter J. Liu. 2023. Exploring the limits of transfer learning with a unified text-to-text transformer. Preprint, arXiv:1910.10683.

Pedro Javier Ortiz Suárez, Laurent Romary, and Benoît Sagot. 2020. A monolingual approach to contextualized word embeddings for mid-resource languages. In Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, pages 1703– 1714, Online. Association for Computational Linguistics.

Ricardo Rei, Nuno M. Guerreiro, José Pombal, João Alves, Pedro Teixeirinha, Amin Farajian, and André F. T. Martins. 2025. Tower+: Bridging generality and translation specialization in multilingual llms. Preprint, arXiv:2506.17080.

PaLM Team. 2022. Palm: Scaling language modeling with pathways. Preprint, arXiv:2204.02311.

Ilia Shumailov, Zakhar Shumaylov, Yiren Zhao, Nicolas Papernot, Ross Anderson, and Yarin Gal. 2024. Ai models collapse when trained on recursively generated data. Nature, 631:755 – 759.

Denis Paperno, Germán Kruszewski, Angeliki Lazaridou, Quan Ngoc Pham, Raffaella Bernardi, Sandro Pezzelle, Marco Baroni, Gemma Boleda, and Raquel Fernández. 2016. The lambada dataset.

Shivalika Singh, Angelika Romanou, Clémentine Fourrier, David Ifeoluwa Adelani, Jian Gang Ngui, Daniel Vila-Suero, Peerat Limkonchotiwat, Kelly Marchisio, Wei Qi Leong, Yosephine Susanto, Raymond Ng, Shayne Longpre, Sebastian Ruder, Wei-Yin Ko, Antoine Bosselut, Alice Oh, Andre Martins, Leshem Choshen, Daphne Ippolito, Enzo Ferrante, Marzieh Fadaee, Beyza Ermis, and Sara Hooker. 2025. Global MMLU: Understanding and addressing cultural and linguistic biases in multilingual evaluation. In Proceedings of the 63rd Annual Meeting of the Association for Computational Linguistics (Volume 1: Long Papers), pages 18761–18799, Vienna, Austria. Association for Computational Linguistics.

Guilherme Penedo, Hynek Kydlíček, Vinko Sabolčec, Bettina Messmer, Negar Foroutan, Amir Hossein Kargaran, Colin Raffel, Martin Jaggi, Leandro Von Werra, and Thomas Wolf. 2025. Fineweb2: One pipeline to scale them all — adapting pre-training data processing to every language. In Second Conference on Language Modeling. Guilherme Penedo, Hynek Kydlíček, Loubna Ben allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. 2024. The fineweb datasets: Decanting the web for the finest text data at scale. In Advances in Neural Information Processing Systems, volume 37, pages 30811–30849. Curran Associates, Inc.

Shivalika Singh, Freddie Vargus, Daniel Dsouza, Börje F. Karlsson, Abinaya Mahendiran, Wei-Yin

12

Ko, Herumb Shandilya, Jay Patel, Deividas Mataciunas, Laura OMahony, Mike Zhang, Ramith Hettiarachchi, Joseph Wilson, Marina Machado, Luisa Souza Moura, Dominik Krzemiński, Hakimeh Fadaei, Irem Ergün, Ifeoma Okoh, Aisha Alaagib, Oshan Mudannayake, Zaid Alyafeai, Vu Minh Chien, Sebastian Ruder, Surya Guthikonda, Emad A. Alghamdi, Sebastian Gehrmann, Niklas Muennighoff, Max Bartolo, Julia Kreutzer, Ahmet Üstün, Marzieh Fadaee, and Sara Hooker. 2024. Aya dataset: An open-access collection for multilingual instruction tuning. Preprint, arXiv:2402.06619. Luca Soldaini, Rodney Kinney, Akshita Bhagia, Dustin Schwenk, David Atkinson, Russell Authur, Ben Bogin, Khyathi Chandu, Jennifer Dumas, Yanai Elazar, Valentin Hofmann, Ananya Harsh Jha, Sachin Kumar, Li Lucy, Xinxi Lyu, Nathan Lambert, Ian Magnusson, Jacob Morrison, Niklas Muennighoff, Aakanksha Naik, Crystal Nam, Matthew E. Peters, Abhilasha Ravichander, Kyle Richardson, Zejiang Shen, Emma Strubell, Nishant Subramani, Oyvind Tafjord, Pete Walsh, Luke Zettlemoyer, Noah A. Smith, Hannaneh Hajishirzi, Iz Beltagy, Dirk Groeneveld, Jesse Dodge, and Kyle Lo. 2024. Dolma: an open corpus of three trillion tokens for language model pretraining research. Preprint, arXiv:2402.00159. Gopher Team. 2022. Scaling language models: Methods, analysis and insights from training gopher. Preprint, arXiv:2112.11446. Puxuan Yu, Luke Merrick, Gaurav Nuti, and Daniel Campos. 2024. Arctic-embed 2.0: Multilingual retrieval without compromise. Preprint, arXiv:2412.04506. Rowan Zellers, Ari Holtzman, Yonatan Bisk, Ali Farhadi, and Yejin Choi. 2019. HellaSwag: Can a machine really finish your sentence? In Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, pages 4791–4800, Florence, Italy. Association for Computational Linguistics. Lianmin Zheng, Wei-Lin Chiang, Ying Sheng, Siyuan Zhuang, Zhanghao Wu, Yonghao Zhuang, Zi Lin, Zhuohan Li, Dacheng Li, Eric P. Xing, Hao Zhang, Joseph E. Gonzalez, and Ion Stoica. 2023. Judging llm-as-a-judge with mt-bench and chatbot arena. Preprint, arXiv:2306.05685. Wenzhen Zheng, Wenbo Pan, Xu Xu, Libo Qin, Li Yue, and Ming Zhou. 2024. Breaking language barriers: Cross-lingual continual pre-training at scale. In Proceedings of the 2024 Conference on Empirical Methods in Natural Language Processing, pages 7725– 7738, Miami, Florida, USA. Association for Computational Linguistics.

13

A

FW2-DE Annotation Details

size of 2048. In all pre-training runs we use the cosine decay learning rate schedule (Loshchilov and Hutter, 2017) with a warmup lasting 1% of total training steps and decay to 1% of the peak learning rate, which we empirically determined to be optimal at 5e−4 for all models. We use the AdamW optimizer (Loshchilov and Hutter, 2019) with β1 = 0.9, β2 = 0.95 and decoupled L2 weight decay coefficient of 0.1. The effective batch size is 0.5M tokens and the gradients are clipped to a maximal Euclidean norm of 1.0. For instruction tuning we use a similar set of hyperparameters with the main differences in the learning warmup lasting 5% of training steps, maximum learning rate being equal to 5e − 5, the batch size being 32, and AdamW β2 = 0.99. We rely on Huggingface’s Nanotron, Datatrove, and TRL libraries to handle the pre-training and instruction-tuning infrastructure. All pre-trianing runs were conducted on 8xA100 Nvidia GPUs, and all instruction-tuning runs on 1xA100 Nvidia GPU. All models were trained with bfloat16 precision. For output generation with instruction-tuned variants of our models, we use the same random seed and rely on top-p sampling with p = 0.9 and temperature t = 0.6. For LLM-as-a-judge output generation, we force greedy decoding to ensure consistency of evaluation. All open-ended output generation was conducted using the vLLM library.

We closely follow the annotation pipeline outlined by Lozhkov et al. (2024) in creation of the English Fineweb-Edu. We first used an LLM to score 500k randomly sampled FW2-DE documents for their Coherence, Information Value, and Educational Quality. For Educational Quality, we use the prompt defined by Lozhkov et al. (2024), while for Coherence and Information Value we define our own prompt, provided in Figure 6. The resulting label distributions are shown in Figure 5. We truncate input documents to a max length of 4096 tokens. The LLM that we use for initial annotation is Llama-3.3-70B-Instruct (Llama 3 Team, 2024). We also considered using Qwen3-32B (Qwen3 Team, 2025) and Gemma-3-27B (Gemma 3 Team, 2025) due to their extensive multilingual support, however we found their performance to be on average weaker compared to a more capable 70B model, as we show in Table 6, and the resulting distributions to be inferior to the one yielded by the 70B model. We then fine-tuned three respective snowflake-arctic-embed-m-v2.0 (Yu et al., 2024) regression models on the annotated 500k sample, each responsible for scoring one of the document’s quality aspects. A hyperparameter sweep showed the linear decay of a max learning rate of 1e-4 to 0 and a batch size of 64 to be optimal over 10 epochs. The best epoch is chosen at the end based on the holdout validation set accuracy. We truncate the web documents to a max length of 1024 tokens. The fine-tuned classifiers were then used annotate the full 500M document German subset of FineWeb-2. The C OHERENCE and I NFORMATION VALUE subsets are derived by filtering out the documents receiving a score below the respective maximum scores, i.e. 3 and 4, while the E DUCATIONAL Q UALITY subset is derived by applying a threshold of 3. The intersection of these subsets, D ENSE C ORE, is derived by applying all 3 score thresholds at the same time.

B

LLM Training Details

We parametrize our 350M models with 24 layers, each with hidden layer size of 1024 and FFN size of 4096, while our 1B models consist of 16 layers, each with hidden layer size of 2048 and FFN size of 8192. All models operate on the context window 14

Figure 5: Distribution of Coherence, Information Value, and Educational Quality scores in FW2-DE yielded by Llama-3.3-70B-Instruct annotation. We add the Education Quality score distribution from the original English Fineweb-Edu for reference.

Model Gemma-3-27B-IT Qwen3-32B Llama-3.3-70B-Instruct

MMLU (DE) (Acc.)

MGSM (DE) (Acc.)

IFEval (Acc.)

Include (DE) (Acc.)

MLQA (DE_EN) (F1)

Avg.

59.09 73.59 73.13

69.60 48.40 74.80

86.21 88.49 93.17

56.18 60.93 64.07

20.70 18.79 50.44

58.36 58.04 71.12

Table 6: IT-LLM Annotator Evaluation. Model precision: bfloat16. Each benchmark evaluation was done with and without the provided prompt template, and the best result is reported. The Include (DE) results reflect the average over only the STEM and Social Science subsets. IFEval results reflect the "strict" version of the task.

15

Subset

Token Count

MMLU

ARC-Challenge

ARC-Easy

Hellaswag

LAMBADA

OBQA

Avg.

100B (1.0x) 200B (1.0x)

27.13 26.82

26.15 26.24

41.10 42.11

37.18 38.39

33.52 36.16

41.01 38.79

34.35 34.75

28B (3.6x) 28B (7.2x)

28.97 29.29

31.40 32.24

50.55 52.87

41.10 43.21

37.55 37.48

45.86 45.86

39.24 40.16

78B (1.3x) 128B (1.6x)

28.24 28.54

29.05 30.65

48.35 51.27

40.38 40.93

36.45 36.14

45.05 46.87

37.92 39.07

Baseline

R ANDOM R ANDOM (197B) Filter Combinations

D ENSE -C ORE D ENSE -C ORE 200B Curriculum-Based

P HASED P HASED 200B

Table 7: Benchmark results of 350M models trained on 200B tokens of different FW2-DE subsets compared to models trained on 100B tokens of (conceptually) the same subsets. For the model trained on 200B tokens from the R ANDOM subset, we report the performance of the 197B checkpoint, as it performs better than the 200B checkpoint on average.

Subset

Token Count

MMLU

ARC-Challenge

ARC-Easy

Hellaswag

LAMBADA

OBQA

Avg.

100B (1.0x) 100B (1.0x)

27.13 27.28

26.15 27.46

41.10 43.97

37.18 41.60

33.52 38.84

41.01 41.82

34.35 36.83

28B (3.6x) 28B (3.6x) 28B (7.2x)

28.97 29.88 31.06

31.40 32.99 35.99

50.55 54.64 57.30

41.10 45.51 48.69

37.55 40.29 42.80

45.86 48.49 48.48

39.24 41.97 44.05

1T 2T∗ 9T∗

29.26 30.01 28.58

30.27 30.55 29.90

48.19 47.89 40.51

44.80 43.43 40.07

44.89 41.71 44.31

47.27 45.05 44.04

40.78 39.77 37.90

Baseline

R ANDOM 350M R ANDOM 1B Filter Combinations

D ENSE -C ORE 350M D ENSE -C ORE 1B D ENSE -C ORE 1B Reference models

LLäMmlein-1B Gemma-3-1B Llama-3.2-1B

Table 8: Benchmark results of 1B models trained on 100B and 200B tokens of the D ENSE -C ORE subset compared to other similar-sized pre-trained models.

16

Below is an extract from a web page. Evaluate the quality of the content based on its coherence and whether it provides valuable information on any subject matter that a reader may take away after reading it. Valuable information refers to information that is unbiased, non-promotional, and either well-reasoned or well-put. On the other hand, overly promotional materials, personal opinions or experiences without solid foundations, and spam/adult content are not considered valuable information. Assign two scores based on the scoring system described below. Coherence score (range: 1-3) Indicates whether the extract is logically structured, easy to follow, and clearly articulated. - 1 point: the extract is mostly incoherent. It may contain rambling or non-linear trains of thought; it could include a lot of interruptions or advertisements. - 2 points: the extract is somewhat coherent, but contains some distracting asides. - 3 points: the extract is mostly or fully coherent. It consists of complete sentences and logical paragraphs, the ideas are well-put and well-argued, and has minimal to no irrelevant interruptions. Information value score (range: 1-4) Measures the amount of unbiased and useful information contained in the extract. - 1 point: the extract has little to no information value. It contains only promotional information, biased information, or personal views and preferences that are not well argued. - 2 points: the extract has some information value, but also some clearly biased or promotional information. - 3 points: the extract has good information value. The content is clearly formulated, well-reasoned, unbiased, and non-promotional. - 4 points: the extract has exceptional information value. It provides in-depth insights into its subject matter, clearly enhancing a reader’s understanding of the topic. The extract: {extract} After examining the extract: - Briefly justify your scores, up to 100 words. - Conclude with the scores using the format: "Coherence score: <Coherence points>. Information value score: <Information value points>"

Figure 6: Annotation prompt used to assign Information Value and Coherence scores.

17

Please act as an impartial judge and evaluate the quality of the response provided by a German AI assistant to the user question displayed below. Your evaluation should consider factors such as the helpfulness, relevance, accuracy, depth, creativity, and level of detail of the response. Begin your evaluation by providing a short explanation. Be as objective as possible. After providing your brief explanation, please rate the response on a scale of 1 to 10 by strictly following this format: "[[rating]]", for example: "Rating: [[5]]". Your evaluation must be performed in English - this includes your explanation, reasoning, and the final verdict. Do not use German at any point. [User Question] {question} [The Start of Assistant’s Answer] {answer} [The End of Assistant’s Answer]

Figure 7: Prompt template used for standalone Likert-scale evaluation of outputs generated by instruction-tuned models. We closely follow the prompt design provided by Zheng et al. (2023).

Please act as an impartial judge and evaluate the quality of the response provided by a German AI assistant to the user question displayed below. Your evaluation should consider correctness and helpfulness. You will be given a reference answer in addition to the assistant’s answer. Your job is to evaluate whether the assistant’s answer is correct. Begin your evaluation by comparing the assistant’s answer with the reference answer. Identify and correct any mistakes. Do not allow the length of the responses to influence your evaluation. Be as objective as possible. After providing your brief explanation, output your final verdict by strictly following this format: "[[1]]" if the assistant’s answer is correct, and "[[0]]" if assistant’s answer is incorrect. Your evaluation must be performed in English - this includes your explanation, reasoning, and the final verdict. Do not use German at any point. [User Question] {question} [The Start of Reference Answer] {answer} [The End of Reference Answer] [The Start of Assistant’s Answer] {answer} [The End of Assistant’s Answer]

Figure 8: Prompt template used for evaluation of binary correctness of outputs generated by instruction-tuned models..

18

Record · ID 149101 · SHA-256 5b47ac3722d4408d
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.