Under review as a conference paper at ICLR 2027
Everything in Moderation: Per-Domain Coverage Optima and Alignment-Resistant Domain Gaps in Multi-Domain Mid-Training Yunpeng Xu Kun Zheng
arXiv:2609.09081v1 [cs.AI] 8 Sep 2026
September 9, 2026 Abstract Mid-training—the stage between pre-training and alignment—is where a model’s per-domain data composition is usually decided by data availability rather than principled design. We ask what that decision buys and whether a later alignment pass can undo it. In a controlled logical-reasoning setting (Qwen3-8B-Base primary, with a 4B replication; five semantically rule-disjoint KOR-Bench domains) we train 30 allocations spanning the five-domain simplex—24 sweep configurations plus six withheld from the fit—at five seeds each. Three findings emerge. First, every domain has an interior coverage optimum: the moderate band (10–40%) is best for all five domains (a descriptive band concordance; a calibrated permutation test for quadratic interiority gives P ≈ 0.010, Appendix F.1), and the fitted mid-training-only curves, with 8B peaks between 9.9% and 35.1%, reproduce out-of-sample for the curve shape (not the peak locations). Second, the gaps survive a fixed-budget alignment pass: compensatory SFT raises 116/120 cells (mean +4.32 pp) yet bridges 0/240 pairs at a 5 pp threshold and only 30/240 at a 10% ratio, an equal-budget uniform control behaves almost identically, and a permutation null reallocating the same gains at random would bridge 13.8 ± 3.3 and 77.9 ± 8.5 pairs (P < 0.001). Third, zero coverage collapses mid-training-only accuracy, and the tested recipe reverses both signs—clearly for Operation, not resolvably for Counterfactual—though a FineWeb-Eduonly control shows the collapse is co-mingled with generic distributional drift. An exploratory θ∗ allocation attains the largest full-pipeline gain of the three trained end-to-end (+4.36 vs. +0.80/+0.64 pp) but is marginal under an uncorrected Welch test and is selected from the same sweep. Because coverage varies on a simplex, the five optima are mixture-level marginals, not the coordinates of one jointly optimal mixture.
1
Introduction
The multi-stage training paradigm (pre-training → mid-training → Supervised Fine-Tuning (SFT) → Reinforcement Learning (RL)) assumes each stage provides a foundation the next can refine. But some early design choices may create constraints that later stages cannot undo. This paper tests that possibility for one specific decision: per-domain data coverage at mid-training. Mid-training—the stage between pre-training and alignment—is a general domainadaptation technique, yet per-domain data composition is typically set by data availability rather than principled design. We study whether, in a fixed total-token mixture and a fixed downstream recipe, coverage-associated differences are subsequently reduced by SFT or RL; we do not treat this study as a test of whether alignment procedures in general can repair coverage gaps. Prior work has partially explored this question but leaves the coverage-allocation dimension untested: Zhang et al. (2025) showed pretraining exposure determines RL generalization in a small synthetic setting, Zhao et al. (2025) that RL amplifies pretrained behaviors, and the SFT/RL literature that SFT primarily teaches format while RL generalizes from existing capabilities (Chu et al., 2025); mid-training work treats continued pretraining as a bridge to downstream alignment (Tu et al., 2025), and data-mixture methods—DoReMi (Xie et al., 2023a), RegMix (Liu et al., 2025)— optimize domain proportions for aggregate perplexity. None of these varies per-domain mid-training coverage while holding subsequent stages fixed (Appendix A). We address this gap through a controlled study on Qwen3-8B-Base using five logical-reasoning domains from KOR-Bench (Ma et al., 2025); these domains are semantically rule-disjoint—each defines a self-contained symbolic system—which reduces semantic-transfer confounding, though they are not statistically or neurally independent (§3.1). We construct 24 coverage configurations spanning the five-domain simplex—from severely imbalanced to
1
balanced to counter-skewed—and hold the planned downstream stages fixed. Coverage is the intended allocation variable, but changing one component necessarily changes the others and may also alter effective repetition and diversity. We evaluate on KOR-Bench and three external benchmarks (CounterBench (Chen et al., 2026), ProofWriter (Tafjord et al., 2021), ZebraLogic (Lin et al., 2025)); three findings emerge (Figure 1; evidence in Sections 4.1–4.4), each stated for the tested setting and marked by what the six withheld allocations support out-of-sample. Observation 1: The gaps survive the reweightings that keep the average gain; closing them forfeits that gain (§4.2, §3.4). Reweighting SFT data toward coverage-deficient domains raises absolute accuracy in 116 of 120 configuration–domain cells yet bridges 0/240 pairwise gaps at the 5 pp threshold and only 30/240 at the 10% ratio, and an equal-budget uniform-SFT control behaves almost identically, so the non-closure is not an artifact of the compensation formula. A permutation null holding the observed gain magnitudes fixed would bridge 13.8 ± 3.3 pairs at 5 pp and 77.9±8.5 at 10% (P < 0.001; Appendix G.2): the gains are placed in a systematically gap-preserving way. Sharpening the policy within the same family closes gaps, but only by trading away the accuracy it was introduced to deliver: raising the concentration exponent closes 0/60 → 12/60 pairs at the 5 pp metric while the mean gain falls +4.34 → +2.26 pp (Appendix G.3), so closure and average accuracy are in tension under a fixed budget. The short GSPO stage is a secondary observation (§5.1). Observation 2: Transient negative transfer under zero coverage (§4.2). The strongest raw effect is negative transfer at the mid-training-only checkpoint: driving a domain to zero mid-training coverage collapses its accuracy even where the prior was high (Counterfactual 83.6%→45.6%; Operation 60.4%→33.2%). The tested recipe partly repairs this (Table 6): Operation returns to +5.2±0.3 pp by mid+SFT, clearly separated from zero, while Counterfactual only turns positive by the RL row (+1.2 ± 3.4 pp, not statistically resolved). The FineWeb-Edu-only control (§5.3) shows the collapse is co-mingled with generic distributional drift, so we do not attribute it specifically to coverage starvation. Observation 3: Every domain’s own coverage marginal has an interior optimum—more is not better (§4.1). No domain is best served by the largest share we gave it, nor by the smallest: holding the rest of the design fixed, accuracy rises with a domain’s own share, peaks at a moderate value, and falls again. The pattern needs no curve fitting—pooling all trained allocations (the 24 sweep plus 6 held-out, and, once Appendix B.2 is included, 12 interior), the moderate band (10–40%) is best for all five domains; this band concordance is descriptive, and a calibrated permutation test for quadratic interiority gives P ≈ 0.010 (Appendix F.1). The fitted split-Gaussian curves at 8B place the peaks between 9.9% and 35.1%. Six withheld allocations reproduce the curve shapes out-of-sample (§4.1), closely for four domains and only approximately for Operation (the peak locations themselves are not validated on held-out data). A mid-training-only replication at 4B reproduces the pattern, though the effect is near-ceiling for Counterfactual (1.3 pp band span) and the fitted peaks shift with scale (−17.5 to +3.8 pp; Appendix C). The claim is deliberately marginal rather than joint: the compositional response surface’s stationary point reads as a saddle in every domain (Appendix F.2), though this classification is provisional given the surface’s parameter count.
Problem & Gap
Expt.1 Coverage Sweep
Does the Gap Persist?
Expt. 2 Vertical Comparison
Can Alignment Close Gaps?
Expt. 3.1 Compensatory SFT
Expt. 3.2 Uniform SFT
Mid-training data composition is ad-hoc, not principled.
Mid-train-only with 24 coverage configurations across 5 domains in KOR-Bench (+6 held out).
If coverage effects are nonmonotonic, does the capability hierarchy survive subsequent alignment?
Pick representative configurations (imbalanced / balanced / θ*).
Vertical results show hierarchy survives across stages.
Reweight SFT data toward deficient domains.
Apply identical SFT mix as baseline.
Run full SFT+RL pipeline and compare against shared baselines.
Can gaps be closed within domains?
Applied to all 24 checkpoints.
Assumed fixable by SFT/RL reweighting, but never tested. Prior work doesn't systematically vary per-domain mid-training coverage with later stages fixed.
Finding 1: Non-Monotonic Coverage Effects
Finding 2: Persistent Coverage Gaps
Figure 1: Experimental flow. The downstream recipe (mid-training→SFT→RL) is held fixed while the five-domain coverage mixture varies: 24 sweep configurations (plus six held-out allocations) establish per-domain interior optima (Finding 1, §4.1); three allocations carry the vertical comparison through the full pipeline (§4.2); compensatory vs. uniform SFT on all 24 checkpoints tests gap closure (Finding 2, §3.4).
2
Methodology
We fix the planned subsequent training stages (SFT, RL) across configurations so mid-training coverage is the intended intervention. The reasoning benchmarks are described in §3.1.
2
2.1
Training Stages
We study a multi-stage pipeline consisting of mid-training followed by SFT and RL; our focus is not the recipe itself, but how the mid-training data distribution shapes later behavior. Mid-training. Let θ0 be the base model parameters. Mid-training performs standard continued causal language modeling on a corpus Dmid , yielding checkpoint θmid , which initializes subsequent stages. Supervised fine-tuning. SFT optimizes conditional next-token likelihood over problem-response pairs, following the standard instruction-tuning paradigm (Wang et al., 2023). For the Mid-training+SFT+RL configuration, SFT starts from θmid ; for the SFT+RL baseline (no mid-training), it starts directly from θ0 . RL: GSPO. After SFT, we apply Group-based Sequence-level Policy Optimization (GSPO) (Yang et al., 2025), a sequence-level group-based RL with binary verifier rewards (Ri = 1 if correct, 0 otherwise); full equations are in Appendix C.1.
2.2
Data Construction
The pipeline (Figure 6, Appendix C.8) proceeds through seven numbered components—a rule library (box 1), a domain dispatcher (2), a rule synthesizer (3), a difficulty sampler (4), a deterministic solver (5a), a verification stage (6), and a data preprocessor (7)—and in our runs all answers and traces arise from the deterministic solver; the optional teacher model branch (box 5b) was not used here. The training data are newly synthesised instances derived from the benchmark rule definitions, not reused benchmark items: split isolation (by instance ID for KOR-Bench; by depth for ProofWriter) and format/information deduplication guarantee that no training instance coincides with an evaluation item (Appendix E). Beyond the pipeline mechanics, two corpus-level design choices are notable. Mid-training controls token exposure rather than sample counts and uses a fixed total token budget (≈1.5B tokens, two epochs; Table 4), a fixed-token mixture of the internal five-domain component and a fixed external component. The external component was generated, verified, and filtered through the same symbolic pipeline; it contains no original benchmark items, and roughly a third of its tokens (34.8%) derive from the ProofWriter rule family, which is why we treat ProofWriter as a samefamily exposure rather than a zero-exposure transfer test (§4.4). The FineWeb-Edu-only baseline (§4.1) uses the same total budget with FineWeb-Edu in place of the internal component; SFT traces are generated by the same symbolic solvers—not by any LLM (token-share summary: Appendix D).
2.3
Coverage Quantification
We measure per-domain mid-training exposure by token share rather than sample count, since token prediction determines the learning signal and longer-form traces would be understated by sample ratios. Domain coverage Cov(d) is the fraction of total mid-training tokens from domain d. Here coverage denotes token allocation only; it is not a direct measure of unique examples, data quality, task difficulty, or an independently manipulable dose. We track two deltas: ∆mid (d) = Accmidtrain+sft+rl (d) − Accsft+rl (d) (net mid-training effect) and ∆rl (d) = Accmidtrain+sft+rl (d) − Accmidtrain+sft (d) (RL contribution); evaluation protocol details are in Appendix B.
3
Experimental Setup
3.1
Reasoning Benchmarks
Primary benchmark: KOR-Bench. We mainly use KOR-Bench, a structured benchmark for knowledge-orthogonal reasoning (Ma et al., 2025). It spans five domains—ciphers, custom mathematical operations, formal logic, constraint puzzles, and counterfactual reasoning—each comprising 25 rule types with verifiable ground-truth answers across three difficulty levels. A central property of KOR-Bench is semantic rule-disjointness: each domain defines a self-contained symbolic system. This reduces semantic-transfer confounding but does not make the domains independent in the statistical
3
or neural sense—they still share language, tokenization, output formats, model parameters, and potentially general reasoning procedures—so we treat per-domain coverage as a controlled mixture coordinate, not an approximately independent causal variable. Because our corpus is generated from KOR-Bench’s own rule definitions (fresh instances, never the original items), KOR-Bench is our in-distribution primary metric; the external benchmarks are the only generalization probes (providing limited consistency checks). External benchmarks. We additionally evaluate on three external benchmarks as limited consistency checks: ProofWriter (Tafjord et al., 2021) (deductive reasoning; D5 is exposed in the fixed corpus, while D6–D9 are held out by depth), ZebraLogic (Lin et al., 2025) Multiple Choice (MC) and Grid (constraint-satisfaction puzzles), and CounterBench (Chen et al., 2026) (zero direct KOR-Bench coverage). These benchmarks differ in exposure and structural similarity, so none is treated as a definitive independent transfer test.
3.2
Model and Hyperparameters
All trainable configurations start from Qwen3-8B-Base without instruction tuning (Yang et al., 2025) (scale rationale: Appendix B.2). Mid-training uses 1 × 10−5 learning rate (LR); SFT uses 5 × 10−5 for 3 epochs. Both stages use full fine-tuning with LLaMA-Factory, DeepSpeed ZeRO-3 (Rajbhandari et al., 2020), bf16, FlashAttention-2 (Dao et al., 2024), batch size 16, and sequence length 12,000 (Appendices B, B.1). For the coverage sweep, each configuration is trained from five random seeds and evaluated on the held-out KORBench test split; we report the mean ± 1 SD (Table 13). Full-pipeline rows (Table 1) and external-benchmark measurements are likewise five-seed means with cross-seed SDs; 95% bootstrap CIs over per-question resamples, where noted, are evaluation-sample estimates (protocol: Appendix B).
3.3
Mid-Training Data Coverage Study
We train mid-training-only checkpoints under 24 controlled domain-coverage configurations (Table 13, Figure 2), spanning the five-domain simplex, and measure evaluation accuracy before any SFT/RL intervention; a compensatory SFT pass is also applied to every configuration (§3.4) to test how a fixed-budget alignment step changes the midtraining-induced differences. Two design notes apply to Figure 2: coverage percentages reflect internal allocation only (the fixed external component contributes shared formal-deduction exposure, which is why Logic retains non-zero accuracy at 0% internal coverage), and because each domain’s proportion varies jointly with the other four, the curves describe mixture-level associations rather than isolated causal dose-response functions (Appendix F). Each domain samples uniformly from all 25 rule types (equal token quota per rule type) with the same three-level difficulty stratification; this does not balance difficulty, answer length, or trace length across configurations.
3.4
Pairwise Gap-Closure Analysis
We apply compensatory SFT to all 24 checkpoints: it reweights SFT data toward coverage-deficient domains under a fixed budget, rd ∝ max(0, 60% − Md ) (Appendix G). We measure the cross-configuration accuracy range be fore and after compensation and summarize pairwise gap closure over the 24 × 52 = 240 unordered domain pairs (Appendix G.2); because the configurations are purposively selected design points, these counts are descriptive.
3.5
Full Training Pipeline Comparison
Table 1 compares three shared no-mid-training baselines—Base, SFT only, SFT+RL—against the three mid-training stages for each of the three experimental allocations. Because the SFT data is identical (balanced by rule type) across all SFT-bearing rows, Mid-training+SFT vs. SFT only isolates the effect of inserting a particular mid-training checkpoint before the same SFT stage; the analogous comparison after RL is Mid-training+SFT+RL vs. SFT+RL (Table 6). Qwen3-8B-Instruct (Yang et al., 2025) is not trained in our pipeline and is reported only as a non-comparable external reference (Appendix C.5).
4
4
Results
4.1
The Impact of Mid-Training Data Coverage
Figure 2 plots per-domain evaluation accuracy against each domain’s own mid-training proportion across the 24 normal configurations (§3.3); the FineWeb-Edu baseline and θ∗ checkpoint are visual references, excluded from fitting. Per-domain evaluation accuracy vs. mid-training portion
100
Domain
Mid-training-only evaluation accuracy (%)
Mark / line Normal sweep observation Held-out validation point (mid-only) Refitted split-Gaussian Exploratory theta-star FineWeb-Edu baseline
Cipher Operation Logic Counterfactual Puzzle
90 80 70 60 50
Operation baseline (45.6%)
40 Logic baseline (32.0%) Counterfactual baseline (29.2%)
30 20 10 0
Cipher baseline (4.4%) Puzzle baseline (1.2%)
0
10
20
30
40
50
60
70
Domain portion in mid-training coverage (%)
Figure 2: Per-domain mid-training-only accuracy vs. the domain’s own mid-training portion. Filled points: 23 of the 24 normal configurations (Table 13; 19/7/19/24/31 omitted for legibility, retained in all fits). Hollow diamonds: six held-out allocations, withheld from the fit, roughly tracking the fitted curves for Cipher, Logic, Counterfactual, and Puzzle (Operation’s low-coverage points sit ≈ 6 pp above the fitted tail). Vertical bars: ±1 SD across five seeds. Dashed lines: split-Gaussian fits; stars: exploratory θ∗ ; dotted lines: FineWeb-Edu baselines. We now ask where each domain’s optimum actually sits. We fit the empirical coverage–accuracy points with parametric curves; the fits locate the optima, but nothing in the moderation claim itself depends on them. Out-of-sample validation on held-out allocations. To probe whether the fitted curves and the compensatory-SFT result describe more than the 24 fitting points, we trained six additional allocations withheld from the fit and carried them through the full pipeline (Table 13). The split-Gaussian fits generalize in all five domains (held-out residuals 1.0– 1.8 pp for Counterfactual, Puzzle, Cipher, and Logic; Operation is the weakest match at 3.0 pp), and the compensatorySFT result reproduces: gains in all 30 cells (mean +4.5 pp) yet 0/60 pairs closed at 5 pp and 5/60 at 10%, and carrying these checkpoints through RL adds +1.19 pp on average. The held-out allocations also test the premise underlying θ∗ : fitting on the 24 sweep configurations alone predicts their post-RL overall accuracy with r = +0.967 (p = 0.002), whereas the relative-gain objective that actually selects θ∗ ranks them only weakly (r = +0.441). The fitted curves are the trustworthy object; an allocation rule should optimise predicted accuracy directly (Appendix F.4). Exploratory marginal fits. Each domain’s curve is fitted by an asymmetric split Gaussian (five parameters; definition and LOO-CV in Appendix F.3). With the refreshed 24-point sweep the fitted peaks are approximately Cipher 9.9%, Operation 15.9%, Logic 15.0%, Counterfactual 26.0%, and Puzzle 35.1%, with the weakest quadratic signal for 2 Cipher (Rquad ≈ 0.49). Because the simplex constraint couples the proportions, these peaks are descriptive mixturelevel summaries whose sum exceeds the 100% budget; we therefore derive an illustrative allocation θ∗ (Eq. 4), whose fitted objective evaluates to ≈ +9.3% relative gain over balanced, not an achieved gain: the re-trained θ∗ checkpoint realizes 58.24 vs. balanced 56.26 at mid-training-only and +4.36 pp over SFT+RL at the full-pipeline stage (Table 1). Fitted 95%-of-peak intervals. Table 10 (Appendix F.3) reports, for each domain, the set of θd where the fitted curve is within 95% of its fitted peak (derivation in Appendix F.3). These are descriptive function-level thresholds for the 5
present sweep; they are not confidence intervals and not validated allocation rules.
4.2
Vertical Comparison: Mid-Training Contribution Across Stages
We now test whether the observed mixture-associated differences change after SFT and RL; the selected full-pipeline comparisons are exploratory. We select two representative sweep configurations and add one derived analytically from the fitted curves, carrying all three through the identical full SFT+RL pipeline: Expt. 1 (imbalanced: Cipher 72.3%, Operation 0%, Logic 12.0%, Counterfactual 0%, Puzzle 15.7%), Expt. 2 (balanced: 20% per domain), and Expt. 3 (exploratory θ∗ ; §4.1)—a new allocation trained as a dedicated checkpoint. The full-pipeline comparison is restricted to three configurations for compute reasons; they span a coverage contrast but are not representative of the sweep (Table 1, Figure 3). Config
Overall
Cipher
Oper.
Logic
Counterf.
Puzzle
60.4 ± 2.9 88.8 ± 2.7 88.8 ± 3.0
47.2 ± 1.9 59.2 ± 2.3 59.2 ± 2.2
83.6 ± 3.0 86.4 ± 3.0 87.6 ± 3.0
3.6 ± 2.0 20.0 ± 2.0 22.8 ± 1.5
Expt. 1 — Imbalanced: Cipher 72.3%, Oper. 0%, Logic 12.0%, Counterf. 0%, Puzzle 15.7% Mid-training only 36.60 ± 1.8 38.1 ± 2.6 33.2 ± 2.8 54.5 ± 2.2 Mid-training+SFT 64.80 ± 2.5 70.0 ± 2.2 94.0 ± 2.7 53.2 ± 1.9 Mid-training+SFT+RL 65.92 ± 1.5 70.4 ± 2.1 92.8 ± 3.0 56.8 ± 2.0
45.6 ± 2.8 85.6 ± 2.8 88.8 ± 3.5
11.6 ± 1.5 21.2 ± 1.4 20.8 ± 2.0
Expt. 2 — Balanced: 20% per domain (4/5 domains within fitted 95%-of-peak intervals) Mid-training only 56.26 ± 1.5 51.1 ± 2.2 77.2 ± 2.9 59.3 ± 2.2 Mid-training+SFT 64.80 ± 1.7 69.6 ± 2.1 86.0 ± 2.9 66.0 ± 2.1 Mid-training+SFT+RL 66.08 ± 2.3 72.0 ± 2.3 85.6 ± 3.2 60.8 ± 2.1
81.2 ± 3.1 81.6 ± 2.9 84.4 ± 3.3
12.5 ± 2.0 20.8 ± 1.9 27.6 ± 1.5
No-mid-training baselines (shared across experiments) Base 40.32 ± 2.7 6.8 ± 2.5 SFT only 64.08 ± 1.8 66.0 ± 2.0 SFT+RL 65.28 ± 1.9 68.0 ± 2.3
Expt. 3 — Exploratory θ∗ : Cipher 9.7%, Oper. 11.8%, Logic 15.8%, Counterf. 30.4%, Puzzle 32.3% (four of five within fitted 95%-of-peak intervals; Operation’s 11.8% sits just below its re-fitted lower bound of 11.9%). Re-trained full pipeline. Mid-training only 58.24 ± 1.5 51.2 ± 2.1 78.0 ± 2.8 61.4 ± 2.1 84.8 ± 3.1 Mid-training+SFT 67.34 ± 2.1 61.3 ± 2.2 87.4 ± 2.9 71.9 ± 2.3 92.1 ± 3.5 Mid-training+SFT+RL 69.64 ± 2.6 63.0 ± 2.3 88.6 ± 2.5 75.9 ± 2.2 94.0 ± 3.3
15.8 ± 1.9 24.0 ± 1.9 26.7 ± 1.3
Table 1: KOR-Bench zero-shot accuracy (%). Base, SFT only, and SFT+RL are shared no-mid-training baselines. Expt. 1–2 use selected sweep allocations and Expt. 3 uses the re-trained exploratory θ∗ allocation; these comparisons probe, rather than establish, the effect of coverage. Values are five-seed means ± 1 SD. The fitted 95%-of-peak intervals are descriptive (Table 10); the non-comparable Qwen3-8B-Instruct reference is in Appendix C.5. Overallcolumn differences among the three full-pipeline rows are at best nominally significant under uncorrected Welch t-tests (§4.2). Table 1 reports the results. The Base model shows strong domain-differentiated priors (83.6% Counterfactual vs. 3.6% Puzzle); mid-training is associated with the largest mid-training-only gains where the prior is low and small or negative differences where it is already high (Figure 3, left). The incremental value of mid-training is descriptively largest for Expt. 3 (θ∗ ): +4.36 pp overall vs. SFT+RL, vs. +0.80 pp (balanced) and +0.64 pp (imbalanced). Welch t-tests give t ≈ 2.29 (p ≈ 0.052) and t ≈ 2.77 (p ≈ 0.030) for θ∗ vs. balanced/imbalanced; neither survives Bonferroni correction, so the ordering is descriptive throughout. The gain is domain-concentrated: Logic +16.7 pp and Counterfactual +6.4 pp, while Cipher is −5.0 pp despite its 9.7% allocation sitting near the fitted mid-trainingonly peak—θ∗ trades Cipher exposure for Logic and Counterfactual coverage. The +4.36 vs. +0.80 pp gap is of the same order as the rows’ own overall SDs, so no per-domain ∆ is claimed as individually significant. Expt. 3 shows the largest RL contribution (+2.30 pp vs. +1.12 and +1.28), a secondary observation (Appendix C.6); its final overall score (69.64) exceeds the non-comparable Qwen3-8B-Instruct reference (63.80). How much of the mid-training-only collapse survives alignment? The interference emphasized above is a midtraining-only observation. Under Expt. 1 the two largest collapses are reversed in sign (Table 6): Operation (−27.2 pp vs. Base at mid-only) reaches +5.2 pp over its SFT-only baseline at mid+SFT, a clear reversal; Counterfactual (−38.0 pp) is still +0.8 pp below at mid+SFT and only turns positive (+1.2 pp) by the RL row, within its SD. What persists is a small residual deficit on Logic and Puzzle (−2.4 and −2.0 pp). The interference effect (Observation 2) is thus a transient, checkpoint-level effect that the tested recipe partly repairs. 6
Expt. 1 (Imbalanced)
25 0 −25
her ration e Op
Cip
e ic al zzl Log rfactu Pu e t n u Co
Δ accuracy (pp)
Mid-training+SFT vs. SFT only
Δ accuracy (pp)
Δ accuracy (pp)
Mid-training-only vs. Base
Expt. 3 (θ * )
Expt. 2 (Balanced)
10 0
her ration e Op
Cip
e ic al zzl Log rfactu Pu e t n u Co
Mid-training+SFT+RL vs. SFT+RL
20 10 0
her ration e Op
Cip
e ic al zzl Log rfactu Pu e t n u Co
Figure 3: Pipeline-stage breakdown: per-domain accuracy gain (∆, pp) across three stages (data from Table 1). Left: Mid-training-only vs. Base (∆ = Accmid − Accbase )—positive for domains where the base prior is low (Cipher, Puzzle); small or negative for Counterfactual, where the base model already achieves 83.6%. Middle: Mid-training+SFT vs. SFT only. Right: Mid-training+SFT+RL vs. SFT+RL. Expt. 1 (imbalanced), Expt. 2 (balanced), Expt. 3 (θ∗ ).
4.3
Horizontal Comparison: Can Alignment Close Mid-Training Gaps?
We test whether reweighting SFT data toward under-covered domains can correct the imbalance installed at midtraining, using two variants: compensatory SFT (weighted toward coverage-deficient domains under a fixed token budget) and uniform SFT (identical data mix, serving as a baseline). Both passes are described here and in Appendix G. Compensatory SFT. Applying the fixed formula-derived compensation policy (§3.4) to all 24 checkpoints raises absolute scores in 116 of 120 configuration–domain cells (mean +4.32 pp; the four cells without gain all receive zero compensatory allocation r = 0). The cross-configuration accuracy range narrows for Cipher (30.9→30.5 pp), Operation (44.7→42.5 pp), Logic (38.3→37.8 pp), and Counterfactual (38.5→36.2 pp), while Puzzle widens only slightly. Pairwise closure over the 240 domain pairs is similarly limited: 30/240 under a 10% relative-ratio threshold and 0/240 under a 5 pp difference threshold (Appendix G.2, Figure 4); under the budget-proportional model the withinconfiguration gain differential is moreover arithmetically bounded (at most ≈ 7 pp), so the tested family has limited dynamic range by construction. The same pattern holds on the six held-out allocations (§4.1). (a) Difference (≥5 pp)
1.0
60
Ratio gap after
Gap after (pp)
80
40 20 0/240 bridged
0 0
25
50
0.8 0.6 0.4 0.2 30/240 bridged
0.0
75
0.0
Gap before (pp) Bridged
(b) Ratio (≥10%)
0.5
1.0
Ratio gap before Narrowed
Widened
Figure 4: Pairwise gap closure after compensatory SFT across the 24 coverage-sweep checkpoints. (a) Difference metric (pp), threshold ≥ 5 pp closure. (b) Ratio metric, threshold ≥ 10% relative closure. Points below the diagonal indicate gap reduction: 0/240 pairs bridged under the difference metric, 30/240 under the ratio metric. Counts are descriptive because the pairs share the 24 configurations.
Uniform SFT control. A uniform SFT pass (identical per-domain mix across configurations; three epochs) was also applied to all 24 checkpoints (Eval. (Uniform) rows in Table 13; Appendix G). Its behaviour is close to that of the compensatory pass: gains in all 120 cells (mean +4.20 pp vs. +4.32 pp), comparable range narrowing, and pairwise closure 0/240 under 5 pp and 32/240 under 10% (vs. 0/240 and 30/240). Per-domain mean gains show 7
where reweighting actually differs (Table 15): compensation is materially stronger only for Counterfactual (+6.37 vs. +4.49 pp), roughly equal for Operation and Logic, and slightly weaker for Cipher and Puzzle. The failure to close pairwise gaps is a property of the tested SFT data and budgets.
4.4
External Evaluation and Limited Consistency Checks
The gap pattern in §4.2 may reflect benchmark-specific overlap or shared task structure. We therefore evaluate a fixed subset of the coverage configurations on three additional benchmarks—ProofWriter (Tafjord et al., 2021), ZebraLogic (Lin et al., 2025), and CounterBench (Chen et al., 2026)—as limited consistency checks: the nine coverage configurations of Table 17 (Appendix H). These evaluations do not establish general transfer: the external corpus draws 34.8% of its tokens from the ProofWriter rule family, the ZebraLogic family contributes 201 mid-training samples, and the available split controls are primarily instance/depth based. ProofWriter. ProofWriter is stratified by proof depth (D5–D9); D5-family items appear in the external mid-training corpus while D6–D9 are held out by depth, so the D6 comparison is best interpreted as same-benchmark held-outdepth evaluation. The reported gain (Expt. 1 checkpoint; Appendix H) is ∆ = +6.4 with 95% CI [+3.3, +9.5]; D7–D9 gains are positive but based on small strata (N ≤ 120). Across the nine configurations of Table 17, the rank correlation between the KOR average and ProofWriter accuracy is small and positive (ρ = +0.34, n = 9; descriptive). ZebraLogic Grid shows a positive ranking association with the KOR-Bench pattern (ρ = +0.53), whereas ZebraLogic MC does not (ρ = −0.12); CounterBench, with zero direct mid-training coverage and uniformly non-positive per-type deltas, correlates positively with the KOR average (ρ = +0.67). These Spearman correlations are descriptive (n = 9), not evidence of a transfer mechanism. Overall, configuration differences are not entirely confined to one benchmark, while shared structure and exposure remain possible explanations.
5
Discussion
The first finding is that the coverage-induced gaps are robust to the alignment budget in a specific, testable sense: a fixed-budget compensatory SFT pass leaves them essentially intact, and an equal-budget uniform-SFT control behaves almost identically, so the non-closure is not a property of the compensation formula. Three further results locate where it comes from: the permutation null (Appendix G.2) shows the gains are placed in a gap-preserving way; the sharpening sweep (Appendix G.3) shows closure is attainable, but only by trading roughly half the mean gain for it; and the six held-out allocations reproduce the whole pattern out-of-sample, including after RL. What this does not license is the general claim that alignment cannot repair coverage choices: we tested one policy family, one budget, and a short RL leg (+1.1 to +2.3 pp; Appendix C.6). Two observations qualify the picture. First, checkpoint-level negative transfer: zero coverage collapses mid-training-only accuracy even where the base model already performs well; the tested recipe partly repairs this (Table 6), and the FineWeb-Edu-only control (§5.3) shows the collapse is co-mingled with generic distributional drift, so we frame gradient dominance (Yu et al., 2020) and drift as competing hypotheses. Second, per-domain coverage is non-monotonically associated with accuracy, forming inverted-U curves with fitted peaks near moderate shares: a moderation-like mixture-level association, not an identified dose-response law. These patterns are confined to the tested KOR-Bench/Qwen3-8B-Base setting, extending Zhang et al. (2025). What would disconfirm the headline claim. Two of the three observations are fragile by construction: the “invertedU” and the “interference” rest on mid-training-only checkpoints under a simplex constraint (§5.2–5.3). The observation a single experiment could disconfirm is the structural negative result (Observation 1): if a fixed-budget compensatory-SFT pass reweighting toward coverage-deficient domains fully closed the between-domain gaps (closure approaching 100% under the 5 pp metric rather than 0/240), our claim would be refuted.
5.1
Why the Tested Alignment Passes Did Not Close the Gaps
These observations are compatible with several explanations: budget competition, domain-specific diversity, optimization mismatch, limited post-training exploration, and the arithmetic cap of the compensation formula (≈ 7 pp
8
by construction; Appendix G). A capability-envelope account remains a hypothesis: this study does not measure representations, gradient conflict, or parameter overlap.
5.2
The Simplex Confound
A structural limit of this design deserves explicit treatment. Coverage is a zero-sum allocation: the five domain proportions always sum to 100%, so a domain’s share cannot be varied independently of the other four, and an unvarying external corpus (34.8% ProofWriter) contributes shared exposure held constant across runs. Consequently, each perdomain curve is a mixture-level marginal rather than a per-domain dose-response, and the inverted-U “peak” a domain shows is a property of where the simplex balances against the fixed external component, not an isolated property of that domain’s coverage: “coverage” is not a cleanly manipulable dose. Separating them requires a simplex-aware joint response-surface model, which we ran (Appendix F.2, B.2): mapping allocations to isometric log-ratio coordinates and fitting a per-domain quadratic surface fits well, but the stationary point reads as a saddle in every domain, and remains one over the 42-allocation pool (30 base plus the 12 interior allocations). We stress this classification is provisional— with 15 quadratic terms estimated from 42 allocations the Hessian is too weakly determined to separate a saddle from a shallow flat region (Appendix F.2)—so the interior data neither establish a jointly optimal mixture nor exclude a shallow one. The five per-domain optima of Observation 3 are therefore marginal statements, not the coordinates of one jointly optimal mixture.
5.3
A Control Confound: Mid-Training on Unrelated Data Also Degrades Out-of-Scope Reasoning
The interference interpretation relies on the claim that “starving a KOR-Bench domain” is what hurts it. The FineWebEdu-only control (Table 13) challenges this framing: mid-training on FineWeb-Edu alone—no KOR-Bench data at all—drops Counterfactual from its Base prior of 83.6% to 29.2% (−54.4 pp), a larger drop than the zero-KORcoverage Expt. 1 case (−38.0 pp), and drops Operation from 60.4% to 45.6% (−14.8 pp). The interference the sweep attributes to selectively starving a KOR-Bench domain is thus co-mingled with a broader effect: any continued pretraining, including on text wholly unrelated to the benchmark, degrades out-of-scope reasoning. We frame the interference finding as a within-setting observation.
5.4
Implications for Design
Because the simplex confound is unresolved and no held-out allocation validates the fitted peaks, intervals, or θ∗ , these results support a hypothesis to test rather than a rule to follow: they do not establish that a domain should be kept “in moderation,” nor that later alignment generally cannot correct coverage choices.
6
Conclusion
Three findings hold in this setting. First, coverage-induced domain gaps survive the finite-budget post-training procedures we tested: the compensatory policy and an equal-budget uniform control both leave the gaps essentially intact (0/240 bridged pairs under the 5 pp metric), this generalizes to six held-out allocations, and a permutation null (P < 0.001; Appendix G.2) shows the gains are placed in a gap-preserving way: budget placement, not only its size, preserves the gaps. Second, zero mid-training coverage collapses mid-training-only accuracy even with a high prior, but the tested recipe reverses the sign of these collapses (Table 6), leaving a small residual deficit on Logic and Puzzle, and the FineWeb-Edu-only control shows this is co-mingled with generic distributional drift. Third, every domain has an interior coverage optimum: the moderate band yields higher mean accuracy than the low or high band for all five domains, with no curve fitting involved (Appendix F.1), and the fitted peaks lie between 9.9% and 35.1%. These optima are per-domain marginals, not a jointly achievable mixture: the compositional response surface’s stationary point reads as a saddle in every domain, even after 12 interior allocations are added (Appendix F.2, B.2), a classification that is provisional rather than decisive—and causal attribution remains out of reach. The practical implication is conditional: coverage allocation may warrant explicit auditing, but the θ∗ gain (+4.36 pp) is marginal under an uncorrected Welch test and remains a model-selected candidate from the same sweep (scope and limits: Appendix I).
9
References M. Abdin et al. Phi-4 technical report. arXiv:2412.08905, 2024. J. Aitchison. The statistical analysis of compositional data. Chapman and Hall, London, 1986. Z. Azerbayev et al. Llemma: An open language model for mathematics. In ICLR, 2024. Z. Cai et al. InternLM2 technical report. arXiv:2403.17297, 2024. J. Chen et al. Unlock the correlation between SFT and RL in training code LLMs. arXiv:2406.10305, 2024. Z. Chen et al. Self-play fine-tuning converts weak language models to strong language models. In ICLR, 2024. J. Chen et al. Step-wise adaptive integration of SFT and RL for task-specific LLMs. arXiv:2505.13026, 2025. Y. Chen, V. K. Singh, J. Ma, and R. Tang. CounterBench: Evaluating and improving counterfactual reasoning in large language models. In AAAI, 2026. T. Chu, Y. Zhai, J. Yang, S. Tong, S. Xie, D. Schuurmans, Q. V. Le, S. Levine, and Y. Ma. SFT memorizes, RL generalizes: A comparative study of foundation model post-training. arXiv:2501.17161, 2025. P. Clark et al. Think you have solved question answering? Try ARC, the AI2 reasoning challenge. arXiv:1803.05457, 2018. K. Cobbe et al. Training verifiers to solve math word problems. arXiv:2110.14168, 2021. T. Dao, D. Y. Fu, S. Ermon, A. Rudra, and C. Ré. FlashAttention-2: Faster attention with better parallelism and work partitioning. In ICLR, 2024. Y. Deng et al. Supervised RL: From expert trajectories to step-wise reasoning. arXiv:2510.25992, 2025. A. Dubey et al. The Llama 3 herd of models. arXiv:2407.21783, 2024. S. Fan, M. Pagliardini, and M. Jaggi. DoGE: Domain reweighting with generalization estimation. In ICML, 2024. R. M. French. Catastrophic forgetting in connectionist networks. Trends in Cognitive Sciences, 3(4):128–135, 1999. L. Gao, S. Biderman, S. Black, L. Golding, T. Hoppe, C. Foster, J. Phang, H. He, A. Thite, N. Nabeshima, S. Presser, and C. Leahy. The Pile: An 800GB dataset of diverse text for language modeling. arXiv:2101.00027, 2021. D. Guo et al. DeepSeek-R1: Incentivizing reasoning capability in LLMs via RL. arXiv:2501.12948, 2025. S. Gururangan, A. Marasović, S. Swayamditta, K. Lo, I. Beltagy, D. Downey, and N. A. Smith. Don’t stop pretraining: Adapt language models to domains and tasks. In ACL, 2020. S. Han et al. FOLIO: Natural language reasoning with first-order logic. In EMNLP, 2024. A. Havrilla et al. Teaching large language models to reason with RL. arXiv:2403.04642, 2024. D. Hendrycks et al. Measuring mathematical problem solving with the MATH dataset. In NeurIPS, 2021. Y. Huang et al. LoRA-PAR: A flexible dual-system LoRA partitioning approach. In EMNLP, 2025. J. Huang et al. ReMiT: RL-guided mid-training for iterative LLM evolution. arXiv:2602.03075, 2026. A. Ibrahim, B. Thérien, K. Gupta, M. L. Richter, Q. Anthony, T. Lesort, E. Belilovsky, and I. Rish. Simple and scalable strategies to continually pre-train large language models. arXiv:2403.08763, 2024. Y. Ishibashi et al. Mining hidden thoughts from texts: Evaluating continual pretraining with synthetic data for LLM reasoning. arXiv:2505.10182, 2025. Z. Ke et al. A survey of frontiers in LLM reasoning. In TMLR, 2025. N. Kim and T. Linzen. COGS: A compositional generalization challenge based on semantic interpretation. In EMNLP, 2020. J. Kirkpatrick, R. Pascanu, N. Rabinowitz, J. Veness, G. Desjardins, A. A. Rusu, K. Milan, J. Quan, T. Ramalho, A. Grabska-Barwinska, D. Hassabis, C. Clopath, D. Kumaran, and R. Hadsell. Overcoming catastrophic forgetting in neural networks. Proceedings of the National Academy of Sciences, 114(13):3521–3526, 2017. B. Y. Lin, R. Le Bras, K. Richardson, A. Sabharwal, R. Poovendran, P. Clark, and Y. Choi. ZebraLogic: On the scaling limits of LLMs for logical reasoning. In ICML, 2025. J. Liu, L. Cui, H. Liu, D. Huang, Y. Wang, and Y. Zhang. LogiQA: A challenge dataset for machine reading comprehension with logical reasoning. In IJCAI, 2020. Q. Liu et al. RegMix: Data mixture as regression for language model pre-training. In ICLR, 2025. Z. Lu et al. MathCoder2: Better math reasoning from continued pretraining on model-translated mathematical code. arXiv:2410.08196, 2024. Y. Luo, Z. Yang, F. Meng, Y. Li, J. Zhou, and Y. Zhang. An empirical study of catastrophic forgetting in large language models during continual fine-tuning. arXiv:2308.08747, 2023. H. Luo et al. WizardMath: Empowering mathematical reasoning for large language models via reinforced evolinstruct. arXiv:2308.09583, 2023. K. Ma, X. Du, Y. Wang, H. Zhang, Z. Wen, X. Qu, J. Yang, J. Liu, M. Liu, X. Yue, W. Huang, and G. Zhang.
10
KOR-Bench: Benchmarking language models on knowledge-orthogonal reasoning tasks. In ICLR, 2025. K. Matsutani. RL squeezes, SFT expands: A comparative study of reasoning LLMs. arXiv:2509.21128, 2025. N. Muennighoff et al. Scaling data-constrained language models. In NeurIPS, 2023. L. Ouyang et al. Training language models to follow instructions with human feedback. In NeurIPS, 2022. G. Penedo, H. Kydlíček, L. B. allal, A. Lozhkov, M. Mitchell, C. Raffel, L. V. Werra, and T. Wolf. The FineWeb datasets: Decanting the web for the finest text data at scale. arXiv:2406.17557, 2024. Y. Qin et al. daVinci-LLM: Towards the science of pretraining. arXiv:2603.27164, 2026. R. Rafailov, A. Sharma, E. Mitchell, C. D. Manning, S. Ermon, and C. Finn. Direct preference optimization: Your language model is secretly a reward model. In NeurIPS, 2023. S. Rajbhandari, J. Rasley, O. Ruwase, and Y. He. ZeRO: Memory optimizations toward training trillion parameter models. In SC, 2020. Q. Ren et al. Rethinking generalization in reasoning SFT: A conditional analysis on optimization, data, and model capability. arXiv:2604.06628, 2026. L. Ruis et al. Procedural knowledge in pretraining drives reasoning in large language models. arXiv:2411.12580, 2024. Z. Shao et al. DeepSeekMath: Pushing the limits of mathematical reasoning in open language models. arXiv:2402.03300, 2024. Z. Shen et al. SlimPajama-DC: Understanding data combinations for LLM training. arXiv:2309.10818, 2023. K. Sinha, S. Sodhani, J. Dong, J. Pineau, and W. L. Hamilton. CLUTRR: A diagnostic benchmark for inductive reasoning from text. In EMNLP-IJCNLP, 2019. L. Soldaini et al. Dolma: An open corpus of three trillion tokens for language model pretraining research. In ACL, 2024. M. Suzgun et al. Challenging BIG-Bench tasks and whether chain-of-thought can solve them. In Findings of ACL, 2023. O. Tafjord, B. D. Mishra, and P. Clark. ProofWriter: Generating implications, proofs, and abductive statements over natural language. In ACL-IJCNLP, 2021. Team KIMI. Kimi k1.5: Scaling reinforcement learning with LLMs. arXiv:2501.12599, 2025. S. Toshniwal et al. OpenMathInstruct-1: A 1.8 million math instruction tuning dataset. In NeurIPS, 2024. C. Tu et al. A survey on LLM mid-training. arXiv:2510.23081, 2025. Y. Wang, Y. Kordi, S. Mishra, A. Liu, N. A. Smith, D. Khashabi, and H. Hajishirzi. Self-Instruct: Aligning language models with self-generated instructions. In ACL, 2023. Z. Wang et al. RLSR: Reinforcement learning with supervised reward outperforms SFT in instruction following. arXiv:2510.14200, 2025. J. Wang, C. Tian, K. Chen, Z. Liu, J. Mao, W. X. Zhao, Z. Zhang, and J. Zhou. MergeMix: Optimizing mid-training data mixtures via learnable model merging. arXiv:2601.17858, 2026. J. Wei et al. Chain-of-thought prompting elicits reasoning in large language models. In NeurIPS, 2022. Xiaomi LLM-Core. MiMo: Unlocking the reasoning potential of language model. arXiv:2505.07608, 2025. S. M. Xie, H. Pham, X. Dong, N. Du, H. Liu, Y. Lu, P. Liang, Q. V. Le, T. Ma, and A. W. Yu. DoReMi: Optimizing data mixtures speeds up language model pretraining. In NeurIPS, 2024. S. M. Xie, S. Santurkar, T. Ma, and P. Liang. Data selection for language models via importance resampling. In NeurIPS, 2023. C. Yang et al. The fine line: Navigating LLM pretraining with down-streaming capability analysis. arXiv:2404.01204, 2024. A. Yang et al. Qwen3 technical report. arXiv:2505.09388, 2025. J. Ye et al. Data mixing laws: Optimizing data mixtures by predicting language modeling performance. In ICLR, 2025. T. Yu, S. Kumar, A. Gupta, S. Levine, K. Hausman, and C. Finn. Gradient surgery for multi-task learning. In NeurIPS, 2020. L. Yu et al. MetaMath: Bootstrap your own mathematical questions for large language models. arXiv:2309.12284, 2023. C. Zhang, G. Neubig, and X. Yue. On the interplay of pre-training, mid-training, and RL on reasoning language models. arXiv:2512.07783, 2025. R. Zhao, A. Meterez, S. Kakade, C. Pehlevan, S. Jelassi, and E. Malach. Echo chamber: RL post-training amplifies behaviors learned in pretraining. In COLM, 2025.
11
C. Zhou et al. LIMA: Less is more for alignment. In NeurIPS, 2023. R. Zhou et al. RuleArena: A benchmark for rule-guided reasoning with LLMs in real-world scenarios. In ACL, 2025.
A
Related Work
A.1
Multi-Stage Training for LLM Reasoning
Modern reasoning models are typically trained through a sequence of continued pretraining, SFT, and RL. The foundational Reinforcement Learning from Human Feedback (RLHF) framework (Ouyang et al., 2022) established the standard three-stage recipe of pretraining, supervised fine-tuning on demonstrations, and RLHF, which most subsequent systems adopt. Guo et al. (2025) show that RL can elicit strong reasoning behavior when applied to a capable base model, while their full DeepSeek-R1 pipeline still relies on cold-start SFT and additional supervised data. Shao et al. (2024) demonstrate that domain-specific continued pretraining can substantially improve mathematical reasoning before later post-training. Kimi k1.5 further scales reinforcement learning in a multi-stage setting (Team KIMI, 2025). These works establish the effectiveness of stage-wise training, but they usually report aggregate benchmark scores and do not isolate how mid-training data imbalance shapes domain-level outcomes. Broader model reports also reinforce the importance of training stages and data coverages. Llama 3 emphasizes the role of large-scale pretraining and post-training recipes (Dubey et al., 2024). InternLM2 (Cai et al., 2024) explicitly incorporates a progressive training strategy with carefully curated mid-training data for reasoning, making domain allocation a first-class engineering concern. Phi-4 highlights the importance of high-quality synthetic and curated data for reasoning-oriented models (Abdin et al., 2024). Our work differs from these system reports by treating mid-training data composition itself as the object of study.
A.2
The Role of Mid-Training
Mid-training is increasingly recognized as a stage with its own objectives and failure modes. Tu et al. (2025) survey recent work and frame mid-training as a bridge between broad pretraining and downstream alignment. Zhang et al. (2025) provide controlled evidence that earlier training exposure determines how much later RL can generalize. Their finding motivates our central question, but their setting is deliberately controlled and synthetic. We instead examine a naturally imbalanced multi-domain corpus. Several studies suggest that continued pretraining can inject useful reasoning structure. Ruis et al. (2024) argue that procedural knowledge in pretraining data drives reasoning performance. Ishibashi et al. (2025) study continual pretraining with synthetic data for reasoning. Yu et al. (2023) show that bootstrapped mathematical data can improve math reasoning. In a related direction, Lu et al. (2024) use model-translated mathematical code for continued pretraining. The clearest single-domain exemplar of this approach is Llemma (Azerbayev et al., 2024), which continues pretraining Code Llama on the 55B-token Proof-Pile-2 corpus and achieves substantial gains on mathematical benchmarks; this work establishes that domain-specific continued pretraining on curated reasoning corpora is a reliable path to capability specialization. Our work extends this line to a multi-domain setting where allocation across domains is the primary planned coordinate. The resulting coverage-sensitive patterns are exploratory and cannot, under the present simplex design, be interpreted as an independently identified domain-specific dose-response effect. The over-coverage half of the dose-response effect may be related to interference in multi-task optimisation (French, 1999; Kirkpatrick et al., 2017; Yu et al., 2020). In sequential continual learning, training intensively on a new task causes gradient updates to overwrite the parameter regions encoding previously learned tasks—a phenomenon termed catastrophic forgetting (French, 1999; Luo et al., 2023a). Scaling and data-curation strategies for continual pre-training have been studied as mitigation levers (Ibrahim et al., 2024), but they address sequential adaptation rather than allocation among simultaneous task domains. Our setting differs structurally: all domains are trained simultaneously rather than sequentially. However, imbalanced coverage may produce an analogous gradient-dominance effect: the highcoverage domain contributes the majority of the gradient signal, pushing shared representations toward its objectives and reducing the signal from low-coverage domains. We do not directly measure gradient conflict, parameter overlap, or Elastic Weight Consolidation (EWC)-style mitigation, so this explanation should be read as a hypothesis rather than a demonstrated mechanism. The KOR-Bench domains have semantic rule-disjointness: each domain defines a self-contained symbolic system whose rules do not rely on concepts or operations from other domains. This reduces one source of semantic-transfer 12
confounding, but it does not make domains independent in the statistical or neural sense. Domains still share language, tokenization, output formats, model parameters, and potentially general reasoning procedures. We therefore treat the proportions as a controlled mixture coordinate, not as an independent causal dose. The present design does not separately identify coverage from unique-example diversity, repetition, sequence length, or all external-corpus exposure. The quality and format of mid-training data matter as much as the quantity. Wei et al. (2022) demonstrate that chain-of-thought reasoning traces—step-by-step explanations provided as training signal—substantially improve multi-step reasoning in large models; our pipeline embeds analogous structured traces in every generated instance. Luo et al. (2023b) show that instruction evolution (Evol-Instruct) applied to mathematical seed problems produces progressively harder variants that boost fine-tuned model performance, demonstrating that data quality amplification can substitute for raw volume. Toshniwal et al. (2024) scale this approach with OpenMathInstruct-1, a 1.8M-problem corpus synthesized by Mixtral-8x7B targeting the MATH benchmark (Hendrycks et al., 2021), which has become a canonical measure of mathematical reasoning capability. These studies collectively show that structured, verified, trace-annotated problems are the fundamental currency of mid-training effectiveness—a principle our multi-domain study extends to five heterogeneous reasoning types. Data composition is another important thread. Yang et al. (2024) analyze how pretraining choices affect downstream capabilities. Qin et al. (2026) argue that pretraining establishes capability ceilings that later stages may not easily exceed. Xiaomi’s MiMo report also emphasizes unlocking reasoning potential through carefully designed training data (Xiaomi LLM-Core, 2025). Our study adds a domain-level view: the question is not only whether mid-training helps overall, but which domains benefit and which domains are left behind.
A.3
Data Mixture and Coverage Optimization
The question of how to allocate data across domains is well-studied in the pretraining literature. Gururangan et al. (2020) establish that domain-adaptive continued pretraining on even modest domain-specific data substantially improves downstream task performance, and that the benefit scales with the mismatch between the general-web pretraining distribution and the target domain. Xie et al. (2023a) propose DoReMi, a Group Distributionally Robust Optimization (DRO)-based method that dynamically reweights domain contributions to equalize per-domain excess loss; their framework directly addresses the uniform-vs.-non-uniform weighting question that our sweep operationalizes empirically. Fan et al. (2024) extend this line with DoGE, which reweights domains by their estimated contribution to a generalization objective via a small proxy model; DoGE and our θ∗ both produce a single recommended allocation, but from opposite directions—DoGE optimizes a gradient-based generalization estimate without training at the target allocation, whereas θ∗ is read off measured per-domain accuracy curves. A controlled comparison on the same five-domain corpus would be informative and is not attempted here. At the corpus level, large-scale composition ablations in the Pile (Gao et al., 2021) and Dolma (Soldaini et al., 2024) demonstrate that data-source mix is among the strongest predictors of benchmark performance. Complementary work addresses which data to include rather than how to weight each domain. Xie et al. (2023b) propose Data Selection via Importance Resampling (DSIR), selecting pretraining documents whose n-gram distributions match a target domain, effectively enabling soft domain filtering without hard corpus boundaries. Shen et al. (2023) construct SlimPajama-DC by rebalancing and deduplicating the RedPajama corpus, showing that domain rebalancing at the corpus level yields consistent downstream improvements. Muennighoff et al. (2023) examine the complementary question of data repetition: when a domain corpus is small, training for multiple epochs can partially substitute for additional data, though with diminishing returns; this bound is relevant when interpreting our low-coverage configurations. More fine-grained theoretical treatments have followed: Ye et al. (2025) establish quantitative scaling laws that predict language modeling loss as a function of domain mixture proportions, and Liu et al. (2025) propose RegMix, which trains 512 small proxy models across diverse domain combinations and fits a regression predictor to identify the optimal large-scale mixture—a principled complement to our empirical sweep. Both methods optimize for aggregate perplexity; our contribution is a domain-resolved diagnostic of how finite-budget mixture configurations are associated with downstream accuracy in this controlled setting, not a general dose-response law. More recent mixture-optimization methods such as MergeMix (Wang et al., 2026) propose using model merging weights as a low-cost proxy for domain performance—a different optimization strategy. A controlled comparison of these methods in the present five-domain setup is future work; the current sweep does not validate their recommended allocations or the local moderate-band guidance.
13
A.4
SFT and RL: Stage Combination and Refinement
SFT and RL are often treated as mechanisms for alignment and refinement rather than as sources of entirely new capabilities. Havrilla et al. (2024) observe that RL can struggle to explore far beyond solutions already reachable by the supervised model. Chen et al. (2024b) show that self-play fine-tuning can improve weak models but also converges without continually expanding the data distribution. These findings are consistent with the idea that later stages inherit constraints from earlier stages. Recent work further studies how SFT and RL should be combined. Wang et al. (2025) compare reinforcement learning and supervised training in instruction following. Chen et al. (2025) propose step-wise adaptive integration of SFT and RL. Deng et al. (2025) connect expert trajectories with step-wise reinforcement learning. Ren et al. (2026) revisit generalization in reasoning SFT, while Matsutani (2025) contrasts the effects of SFT and RL on reasoning models. Two foundational contributions bound the effectiveness of post-pretraining stages. LIMA (Zhou et al., 2023) demonstrates that as few as 1,000 carefully curated demonstrations can match much larger SFT corpora in instruction-following quality, implying that the capability headroom pre-established by earlier training stages is the primary bottleneck rather than fine-tuning volume. Direct Preference Optimization (DPO) (Rafailov et al., 2023) reformulates preference alignment as contrastive supervised learning, removing the reward model and substantially lowering the engineering complexity of the RL stage; the wider adoption of DPO has made controlled earlier-stage ablations such as ours easier to reproduce. Our experiments address both directions of this question. First, later stages do not reliably compensate for domain imbalance introduced at mid-training: even when SFT data is augmented toward coverage-deficient domains, difference-based pairwise gaps fail to close beyond the 5 pp repair threshold (0/240 pairs under a 5 pp difference metric), and ratio-based gaps show only limited average closure (30/240 pairs under a 10% ratio). An equal-budget uniform-SFT control behaves almost identically (0/240 and 32/240 under the same metrics; Appendix G), so the non-closure is not an artifact of the compensation formula. Per-domain fits of the gain against the remedial allocation (∆ ≈ a log(1 + b r) + c) rise within each domain yet saturate at domain-specific ceilings, and pooling across domains explains little of the cell-level variance (R2 ≈ 0.31; Figure 7, Table 16): redirecting SFT data does not reset the mid-training starting point. Second, optimizing mid-training coverage makes later stages materially more effective when coverage places most domains within their fitted 95%-of-peak intervals: mid-training adds +4.36 pp over no-mid-training SFT+RL in Expt. 3 (four of five domains in-band, Operation just below the lower boundary), compared with +0.8 pp under balanced coverage and +0.6 pp under imbalanced coverage.
A.5
Reasoning Evaluation and Domain-Level Analysis
Reasoning ability is often reported as a single aggregate number, but aggregate scores can obscure where gains and losses occur. Two widely used benchmarks illustrate this masking problem: GSM8K (Cobbe et al., 2021) reduces grade-school arithmetic reasoning to a single pass@1 accuracy number, while BIG-Bench Hard (Suzgun et al., 2023) aggregates 23 diverse tasks—spanning inductive, deductive, spatial, and algorithmic reasoning—into a single score that can conceal divergent per-domain trends of the kind our domain-level analysis reveals. Ke et al. (2025) survey reasoning benchmarks broadly and emphasize the diversity of tasks used to evaluate LLM reasoning. Prior work on code and reasoning also decomposes broad competence into smaller behavioral units, for example in analyses of codemodel functions (Chen et al., 2024a) and parameter-efficient reasoning modules (Huang et al., 2025). Huang et al. (2026) further use RL signals to guide mid-training data selection. We share the motivation of looking beyond a single global score, but our analysis is organized around domain-level mid-training coverage and downstream domain-level behavior.
A.6
Knowledge-Orthogonal and Structured Reasoning Benchmarks
Evaluating LLMs on structured reasoning tasks that require applying explicit rules rather than retrieving memorized world knowledge has gained increasing attention. Han et al. (2024) introduce FOLIO, a first-order logic reasoning dataset in which models must derive entailment from formal logical statements without relying on surface-level heuristics; its difficulty for frontier models motivates the mid-training coverage study we conduct on the Logic domain. Liu et al. (2020) present LogiQA, a multiple-choice logical reasoning benchmark drawn from Chinese civil-service examinations covering categorical, conditional, and disjunctive reasoning types; its structural diversity is representative of the rule-following challenge we study. Sinha et al. (2019) construct CLUTRR to test systematic generalization in relational rule-following by withholding chain-length compositions from the training distribution—a compositional
14
generalization challenge closely related to the formal-deduction structure of our Logic and Puzzle domains. More recently, Zhou et al. (2025) propose RuleArena, evaluating models on complex real-world rule systems (airline policies, tax codes, NBA regulations), requiring the same kind of rule-adherence that KOR-Bench (Ma et al., 2025) formalizes as knowledge-orthogonal reasoning. COGS (Kim and Linzen, 2020) probes compositional generalization in semantic parsing by constructing test cases whose structural complexity systematically exceeds the training distribution; the compositional gap it reveals—where models fail on combinations of primitives they handle individually—parallels KOR-Bench’s rule-composition design and motivates the emphasis on generalization over memorization. The AI2 Reasoning Challenge (ARC) (Clark et al., 2018) partitions elementary science questions into a retrieval-solvable Easy set and an inference-requiring Challenge set, providing an early demonstration that even modest reasoning demands stratify models in ways aggregate accuracy cannot capture. Our work differs from these benchmark contributions in its focus on the training side: rather than proposing new evaluation criteria, we ask how the proportion of domain-specific data during mid-training determines whether the corresponding structured reasoning capabilities are reliably acquired. The knowledge-orthogonal property of KORBench—which ensures that training gains in one domain do not automatically transfer to others—makes it a particularly controlled instrument for isolating the effect of per-domain coverage.
B
Training and Evaluation Details
B.1
Compute
All training used full fine-tuning on Qwen3-8B-Base with DeepSpeed ZeRO-3 and bf16; evaluation used vLLM on 8 GPUs. We report the study’s scale as run counts and token budgets, which are exactly determined by the design, rather than as GPU-hours, which we did not instrument per run. Mid-training accounts for 160 runs at a fixed ≈1.5B-token budget over two epochs each: 24 sweep configurations, six held-out allocations, the θ∗ allocation and the FineWeb-Edu baseline, at five seeds apiece. The supervised stage adds 155 compensatory passes and 155 matched uniform-control passes (both over the 24 sweep and six heldout checkpoints plus θ∗ , five seeds each), 20 standard-SFT passes for the three full-pipeline configurations and the shared SFT-only baseline, and a further 90 passes for the three additional sharpening branches of Appendix G.3 (γ ∈ {2, 4, ∞}) over six configurations. The RL stage adds 50 GSPO runs at a fixed 200-step budget: the three fullpipeline configurations, the six held-out allocations, and the shared SFT+RL baseline, five seeds each. This totals 630 training runs (160 mid-training, 420 supervised, 50 RL) for the primary 8B study, of which mid-training dominates the cost because it is the only stage operating on a multi-billion-token corpus. Two validation sub-experiments are budgeted separately and on a smaller scale: Experiment A (Appendix B.2) adds 12 interior allocations × 3 seeds = 36 mid-training-only runs, and the 4B replication (Appendix C) adds 8 allocations × 3 seeds = 24 mid-training-only runs at the smaller model. The design deliberately spends most of that budget on breadth at the mid-training stage rather than depth at the RL stage, which is why the RL leg is short (200 steps) and is reported as a secondary observation. A study prioritising the RL question differently would invert this allocation. For a group planning a replication, the cheapest informative subset is the coverage sweep plus the compensatory and uniform SFT passes—roughly the first two thirds of the runs above—since Observations 1 and 3 are both identified without any RL.
B.2
Model Scale Selection
We chose Qwen3-8B-Base after measuring, rather than assuming, where a mid-training intervention has room to show itself. Table 5 evaluates the Qwen3 family under our KOR-Bench protocol, and the relevant quantity is the headroom a base checkpoint still has before post-training: the base-to-instruct gap is +21.2 pp at 4B (41.88 → 63.08) and +20.8 pp at 8B (43.00 → 63.80), but only +10.4 pp at 30B-A3B (56.48 → 66.88). At the top of the range the base models have already closed most of that distance on their own—14B-Base reaches 60.56 and 32B-Base 60.68, against an instruct band of roughly 63–67—so any mid-training reallocation competes for a shrinking remainder and its per-domain effects are compressed toward the ceiling. At the bottom, Qwen3-1.5B-Base scores 22.84 overall with Cipher at 1.0 and Puzzle at 4.2, i.e. at the floor on two of the five domains, where allocation differences cannot register above seed noise. The 4B and 8B checkpoints are the two with the largest measured post-training headroom, carrying roughly twice that of 30B-A3B; 8B is the larger of the two and the one we could afford to run at full sweep scale (24 configurations × five seeds, plus six held-out allocations and the downstream passes). This is a statement about where 15
the effect is measurable, not a claim that 8B is representative: whether the interior optima and the non-closure result reproduce at other scales is untested here. What a 4B replication would have to show. Because 4B carries essentially the same post-training headroom as 8B (+21.2 vs. +20.8 pp), it is the sharpest available test of whether our findings are scale-specific or headroom-specific: a failure to reproduce at 4B could not be explained away as a ceiling or floor artifact. A minimal replication needs three components. (i) Coverage sweep. Train 8–12 allocations at 4B that span the low (< 10%), moderate (10– 40%) and high (> 40%) band for every domain—a subset of the allocations in Table 13 suffices—at three seeds, evaluated mid-training-only. This supports both the model-free band comparison of Appendix F.1 and the quadratic interiority permutation test on which our moderation claim actually rests. (ii) Alignment arm. Apply the compensatory and uniform SFT passes to those same 4B checkpoints and recompute the pairwise closure counts together with the permutation null of Appendix G.2; this separates a property of the recipe from a property of the 8B model. (iii) Peak comparison. On allocations common to both scales, compare fitted peak locations at 4B and 8B. The three possible outcomes are diagnostic: peaks reproducing within seed noise would indicate the optimum is a property of the data mixture rather than the model; peaks shifting systematically with scale would make the fitted intervals scaledependent and require the allocation guidance to be re-derived per scale; and an absence of interior optima at 4B, given its matched headroom, would localise the effect to 8B and falsify the generality we do not currently claim. A full-pipeline (SFT+RL) 4B arm is not required for any of these three tests, and would be needed only to check the θ∗ ordering at a second scale. Experiment A: an interior-simplex filler sub-sweep. Our compositional reanalysis (Appendix F.2) finds a saddle, not an interior maximum, because the 30 trained allocations cluster near the simplex’s edges and mid-ridge, so the curvature of the interior cannot be resolved. To determine whether a jointly optimal mixture exists—which would let the per-domain moderation result of Observation 3 be promoted to a joint claim, or would establish that it cannot be— we therefore trained 12 new allocations confined to the interior (every domain share in [12, 34]%, each at least 5 pp from every existing allocation), drawn by Dirichlet(α = 8) sampling and screened for feasibility. Table 2 reports them (mid-training-only, three seeds). The result is consistent with the coda analysis and does not overturn it: adding these 12 interior points to the 30-allocation pool and refitting the compositional surface leaves every domain’s stationary point reading as a saddle (R2 = 0.78–0.96 over the 42-pool), though with 15 quadratic terms from 42 allocations the Hessian classification is suggestive rather than decisive (Appendix F.2)—the interior data neither establish a joint optimum nor exclude a shallow flat region. The per-domain moderation result is correspondingly confirmed and delimited: over the full 42-allocation pool the moderate band remains the highest for all five domains (low/moderate/high means, e.g. Cipher 43.8/51.3/42.2, Operation 51.9/75.8/50.4), and the 12 interior points also supply an additional batch of out-of-sample curve-validation points. Because no joint optimum is identified, the moderation finding stays a per-domain marginal statement and is not promoted to a single recommended mixture.
C
Study II: A Cross-Scale Replication at 4B
The scale rationale of Appendix B.2 shows that the post-training headroom—the quantity that determines whether a mid-training reallocation can produce measurable per-domain differences—is largest at 4B (+21.2 pp) and 8B (+20.8 pp) and roughly half that at 30B-A3B (+10.4 pp), while 1.5B sits on the floor (Cipher 1.0, Puzzle 4.2) and 14B/32B near the ceiling (60.6 vs. an instruct band of roughly 63–67). Only the 4B and 8B checkpoints therefore lie in a regime where the effect can be observed. Because 4B carries almost the same headroom as 8B, a failure to reproduce at 4B cannot be written off as a ceiling or floor artifact, which makes it the sharpest available test of whether our findings are scale-specific or headroom-specific. We therefore ran a mid-training-only replication at Qwen3-4B-Base over the eight allocations of Table 3, chosen so that every domain has low (< 10%), moderate (10–40%) and high (> 40%) coverage samples across the set; the allocations reuse the corresponding 8B design points where they exist so the two scales can be compared on matched points. The moderation pattern reproduces at 4B: the moderate band again yields the highest mean accuracy in all five domains (e.g. Cipher 36.3/47.9/38.0, Operation 49.3/90.8/61.1), though the effect is far weaker for Counterfactual, whose three bands (91.0/91.3/90.0) span just 1.3 pp because its base prior is already near ceiling; the three-parameter quadratic is concave with an interior vertex in all five domains. The peak locations, however, shift with scale, and not in one direction: the fitted peaks at 4B (Cipher 8.8, Operation 17.2, Logic 18.8, Counterfactual 8.5, 16
Table 2: Experiment A: 12 interior allocations (all components in 12–34%, ≥ 5 pp from every existing allocation), each trained mid-training-only over three seeds. Each row block reports the Portion (internal mid-training token share, %), the Evaluation (mid-training-only accuracy, mean ± SD across the three seeds), and the Remedial SFT allocation from Eq. 5. Adding these to the 30-allocation pool and refitting the compositional surface leaves a saddle in every domain, so no joint interior optimum is established (§ 5.2). Domains
Ciphers
Operations
Logic
Counterfactual
Puzzles
Config 1: Portion Evaluation Remedial SFT
12.9% 52.9 ± 1.6% 23.5%
24.4% 76.4 ± 3.0% 17.8%
29.0% 57.2 ± 1.8% 15.5%
14.1% 74.2 ± 2.4% 22.9%
19.5% 12.0 ± 1.2% 20.2%
Config 2: Portion Evaluation Remedial SFT
14.5% 51.4 ± 1.3% 22.7%
18.3% 76.9 ± 2.3% 20.8%
12.9% 59.1 ± 1.5% 23.5%
21.2% 82.4 ± 1.8% 19.4%
33.0% 15.1 ± 3.2% 13.5%
Config 3: Portion Evaluation Remedial SFT
12.5% 51.5 ± 2.1% 23.8%
15.9% 77.2 ± 2.9% 22.0%
27.5% 57.6 ± 0.6% 16.3%
21.6% 82.6 ± 3.0% 19.2%
22.5% 13.6 ± 1.3% 18.7%
Config 4: Portion Evaluation Remedial SFT
17.5% 53.6 ± 1.2% 21.3%
22.0% 78.1 ± 2.5% 19.0%
23.2% 59.4 ± 0.8% 18.4%
12.0% 74.0 ± 2.3% 24.0%
25.3% 13.7 ± 3.3% 17.3%
Config 5: Portion Evaluation Remedial SFT
30.4% 50.7 ± 0.7% 14.8%
14.6% 77.4 ± 0.7% 22.7%
12.9% 58.6 ± 2.8% 23.5%
17.4% 78.0 ± 3.4% 21.3%
24.6% 14.4 ± 1.8% 17.7%
Config 6: Portion Evaluation Remedial SFT
25.7% 50.9 ± 2.5% 17.2%
24.5% 77.8 ± 2.5% 17.8%
15.6% 58.5 ± 2.1% 22.2%
16.6% 77.3 ± 2.8% 21.7%
17.6% 11.1 ± 2.9% 21.2%
Config 7: Portion Evaluation Remedial SFT
24.9% 50.4 ± 1.1% 17.5%
13.8% 77.9 ± 3.1% 23.1%
18.6% 59.5 ± 1.1% 20.7%
25.8% 81.9 ± 2.8% 17.1%
16.9% 11.5 ± 2.9% 21.5%
Config 8: Portion Evaluation Remedial SFT
18.9% 52.5 ± 2.6% 20.5%
24.7% 76.2 ± 0.6% 17.7%
18.4% 59.7 ± 3.2% 20.8%
15.9% 76.9 ± 1.0% 22.1%
22.1% 12.9 ± 1.3% 18.9%
Config 9: Portion Evaluation Remedial SFT
24.3% 51.5 ± 2.2% 17.9%
16.6% 77.2 ± 2.6% 21.7%
16.4% 60.4 ± 2.5% 21.8%
20.8% 80.9 ± 1.2% 19.6%
21.9% 12.8 ± 1.5% 19.1%
Config 10: Portion Evaluation Remedial SFT
33.4% 50.8 ± 1.0% 13.3%
17.4% 78.1 ± 0.6% 21.3%
17.4% 58.6 ± 1.8% 21.3%
12.7% 73.3 ± 2.4% 23.7%
19.1% 11.1 ± 3.1% 20.5%
Config 11: Portion Evaluation Remedial SFT
13.2% 51.8 ± 3.4% 23.4%
17.7% 78.9 ± 1.0% 21.2%
23.8% 58.8 ± 2.8% 18.1%
25.3% 83.5 ± 3.4% 17.4%
20.1% 13.0 ± 3.2% 20.0%
Config 12: Portion Evaluation Remedial SFT
14.8% 52.2 ± 2.4% 22.6%
19.2% 77.4 ± 1.4% 20.4%
29.9% 56.9 ± 2.5% 15.1%
15.4% 75.7 ± 2.9% 22.3%
20.7% 13.4 ± 3.2% 19.6%
Puzzle 31.6; the Counterfactual estimate is dominated by the near-flat response and its base prior of 90.5, and Puzzle by a single high-coverage point) differ from the 8B peaks by −1.1, +1.3, and +3.8 pp for Cipher, Operation, and Logic, −3.5 pp for Puzzle, and −17.5 pp for Counterfactual (deltas against the rounded peaks of Table 10). This is the scale-dependence outcome flagged in Appendix B.2: the interior optima are present at both scales, but their locations are not scale-invariant, so allocation guidance derived at one scale should not be transferred directly to another. We did not run the SFT or RL leg at 4B, so the alignment pass (compensatory vs. uniform closure, and the permutation null of Appendix G.2) and the θ∗ ordering were not re-tested at the second scale. The band comparison and the quadratic interiority test (read off the mid-training-only evaluations) and the peaklocation comparison are therefore the two of the three tests in Appendix B.2 that the current 4B data support; the alignment-pass test and the θ∗ ordering remain to be run at the second scale.
17
Table 3: Study II: eight cross-scale allocations for the Qwen3-4B-Base replication, each trained mid-trainingonly. Each row block reports the Portion (internal mid-training token share, %), the Bands (each domain’s low/moderate/high coverage), the Evaluation (mid-training-only accuracy, mean ± SD), and the Remedial SFT allocation from Eq. 5. The 4B run was mid-training-only, so no post-SFT columns are reported; the moderation pattern reproduces at 4B (moderate band best in all five domains, concave quadratic with an interior vertex), but the fitted peak locations shift with scale. Domains
Ciphers
Operations
Logic
Counterfactual
Puzzles
Config 1: Portion Bands (L/M/H) Evaluation Remedial SFT
72.3% H 38.0 ± 0.8% 0%
0% L 44.4 ± 1.8% 28.3%
12% M 45.2 ± 3.0% 22.6%
0% L 90.5 ± 1.0% 28.3%
15.7% M 7.4 ± 2.5% 20.9%
Config 2: Portion Bands (L/M/H) Evaluation Remedial SFT
20% M 45.1 ± 0.8% 20%
20% M 96.9 ± 1.3% 20%
20% M 49.9 ± 1.5% 20%
20% M 91.0 ± 0.8% 20%
20% M 9.1 ± 0.7% 20%
Config 3: Portion Bands (L/M/H) Evaluation Remedial SFT
0% L 25.2 ± 1.8% 30%
10% M 75.3 ± 3.1% 25%
40% M 42.2 ± 0.7% 10%
10% M 91.9 ± 3.0% 25%
40% M 10.7 ± 1.9% 10%
Config 4: Portion Bands (L/M/H) Evaluation Remedial SFT
5% L 43.8 ± 0.6% 27.5%
15% M 96.3 ± 1.7% 22.5%
5% L 37.1 ± 3.1% 27.5%
15% M 91.7 ± 1.4% 22.5%
60% H 2.3 ± 2.0% 0%
Config 5: Portion Bands (L/M/H) Evaluation Remedial SFT
33% M 48.7 ± 1.1% 13.5%
5% L 55.9 ± 1.6% 27.5%
6% L 35.9 ± 0.7% 27%
50% H 90.0 ± 2.4% 5%
6% L 2.3 ± 3.0% 27%
Config 6: Portion Bands (L/M/H) Evaluation Remedial SFT
4% L 39.8 ± 0.8% 28%
3% L 47.6 ± 0.9% 28.5%
30% M 49.4 ± 2.2% 15%
29% M 90.6 ± 0.8% 15.5%
34% M 12.6 ± 3.0% 13%
Config 7: Portion Bands (L/M/H) Evaluation Remedial SFT
25% M 48.3 ± 2.1% 17.5%
45% H 61.1 ± 2.5% 7.5%
10% M 45.7 ± 1.9% 25%
10% M 91.5 ± 1.7% 25%
10% M 4.6 ± 3.1% 25%
Config 8: Portion Bands (L/M/H) Evaluation Remedial SFT
20% M 49.3 ± 0.9% 20%
25% M 94.9 ± 0.9% 17.5%
45% H 40.5 ± 0.7% 7.5%
5% L 91.4 ± 1.8% 27.5%
5% L 3.8 ± 1.4% 27.5%
C.1
GSPO Objective
Group-based Sequence-level Policy Optimization (GSPO) (Yang et al., 2025) is a group-relative policy gradient method. For each prompt q, G responses {o1 , . . . , oG } are sampled from the current policy πθ . Each response receives a binary verifier reward Ri ∈ {0, 1} (1 if correct, 0 otherwise). The advantage for response oi is computed via group-level normalisation: Ri − mean({R1 , . . . , RG }) Ai = , (1) std({R1 , . . . , RG }) and the policy is updated by maximising the clipped surrogate objective:
h
1 JGSPO (θ) = Eq,{oi } G
PG
i=1 min(ρi Ai , clip(ρi , 1 − ε, 1 + ε)Ai )
i
,
(2) where ρi = πθ (oi | q)/πold (oi | q) is the sequence-level probability ratio. Under the standard autoregressive facQ torization this is the product of per-token ratios, ρi = t πθ (oi,t | q, oi,<t )/πold (oi,t | q, oi,<t ), equivalently the sum of per-token log-ratios in log space. ε = 0.2 is the clip ratio, and θold denotes the policy parameters before the 18
update. When all rewards within a group are equal, std({R1 , . . . , RG }) = 0 and the normalized advantage in Eq. 1 is undefined. Because a group with all-equal binary rewards carries no preference signal, its treatment affects only that group’s gradient contribution and not the relative ordering of updates across other groups; the exact handling in our runs is implementation-specific and is acknowledged here as an implementation detail rather than a design choice. Following Yang et al. (2025), we use group size G = 8. Training proceeds for 200 steps with a peak learning rate of 1 × 10−6 under the cosine schedule with 5% warmup (Table 4). RL rollouts are sampled with the same 16,384-token sequence cutoff as evaluation, so rollout truncation does not differentially bias verifier rewards for long reasoning traces (Appendix C.6).
C.2
Training Hyperparameters
Table 4 lists the complete hyperparameter set for all three stages. The compensatory-SFT pass and its uniform-SFT control reuse the SFT column’s three-epoch protocol at a matched total token budget (Appendix G); the RL stage uses a fixed 200-step GSPO schedule (Appendix C.6). Parameter
Mid-Training
SFT
RL (GSPO)
Base model Stage type Fine-tuning type Learning rate Per-device batch size Gradient accumulation steps Effective batch size Sequence length (cutoff) Training volume Epochs Optimizer LR schedule Warmup ratio Precision Gradient checkpointing Flash attention Packing Validation split Chat template
Qwen3-8B-Base (base) pt (cont. pretraining) Full 1e-5 4 4 16 12,000 ∼1.5B tokens (2 ep.) 2.0 AdamW (ZeRO-3) warmup_stable_decay 0.05 bf16 enabled FA2 enabled 1% None (completion)
Mid-training ckpt. sft Full 5e-5 4 4 16 12,000 36,537 inst. (3 ep.) 3.0 AdamW (ZeRO-3) cosine 0.05 bf16 enabled FA2 enabled 5% qwen (chat)
Mid-training+SFT ckpt. rl (GSPO) Full 1e-6 1 16 16 16,384 200 steps — (step-based) AdamW (ZeRO-3) cosine 0.05 bf16 enabled FA2 disabled reward-based qwen (chat)
RL-specific hyperparameters Group size G Clip ratio ε Reward type
— — —
— — —
8 0.2 binary verifier
Table 4: Training hyperparameters for all three stages. Mid-training and SFT use LLaMA-Factory with full fine-tuning and DeepSpeed ZeRO-3. RL uses GSPO (Yang et al., 2025) with binary verifier rewards at a fixed 200-step budget. The compensatory-SFT pass and its uniform-SFT control use the same 3-epoch protocol and total token budget as the SFT column (Appendix G). FA2 = FlashAttention-2.
C.3
Seed Selection and Model Selection Protocol
For the mid-training coverage sweep (Table 13), each configuration is trained from five random seeds and evaluated on a held-out KOR-Bench test split that was never used during training or model selection; we report the mean ± 1 SD across the five seeds, so the SDs capture cross-seed variance for the mid-training-only, compensatory-SFT, and uniform-SFT evaluations. The full-pipeline rows in Table 1 (Mid-training+SFT, Mid-training+SFT+RL) are likewise five-seed means, and Table 1 reports the corresponding cross-seed SDs in every row and cell; the externalbenchmark measurements are five-seed means, and bootstrap 95% CIs over per-question resamples, where reported, capture evaluation-sample variance and do not replace the cross-seed uncertainty. All values in this manuscript are five-seed means unless stated otherwise.
19
C.4
Answer Extraction, Decoding, and Evaluation Protocol
We extract the final answer via regex [[...]] (last match). For Operation, we additionally apply SymPy-based equivalence matching. For Logic, whitespace and punctuation are stripped before comparison. All evaluations use greedy decoding (temperature = 0.0, max_tokens =P 16,384) with vLLM on 8 GPUs. Overall accuracy is the macroNd ⊮[ŷi = yi∗ ]. For difficulty-stratified external benchmarks, average of domain-level accuracies Acc(d) = N1d i=1 P 1 per-stratum accuracy is Acc(d, ℓ) = Nd,ℓ i:ℓi =ℓ ⊮[ŷi = yi∗ ], where ℓ is proof depth, house count, or reasoning type. Confidence intervals are computed using bootstrap resampling for KOR-Bench domain results and asymptotic intervals for large external strata. External benchmarks are evaluated using their native protocols.
C.5
External Reference Models (Qwen3 Family)
Table 5 evaluates the Qwen3 family (Yang et al., 2025) under the same KOR-Bench protocol, for two purposes. None of these models is trained in our pipeline and their instruction-tuning and RL recipes differ from the recipe studied here, so no value in the table is comparable to our pipeline rows—in particular the Qwen3-8B-Instruct score of 63.80% should not be read as evidence about any of our allocations. The table’s role is instead (i) to place our trained checkpoints against publicly available reference points, and (ii) to supply the base-to-instruct headroom measurements that motivate the choice of scale in Appendix B.2. Table 5: KOR-Bench zero-shot accuracy (%) of the Qwen3 model family (1.5B/4B/8B/14B/30B-A3B/32B, base and instruct), evaluated with the same protocol as Table 1. None of these models is trained in our pipeline; their instruction-tuning and RL recipes differ from the recipe studied here, so no value is comparable to our pipeline rows. The base-to-instruct pairs supply the post-training headroom used in the model-scale rationale (Appendix B.2). Model
Overall
Qwen3-1.5B-Base
22.84
Qwen3-4B-Base
41.88
Qwen3-4B-Instruct
63.08
Qwen3-8B-Base Qwen3-8B-Instruct
Oper.
Logic
Counterf.
Puzzle
1.0
7.0
20.2
81.8
4.2
18.2
60.0
39.2
90.0
2.0
59.8
93.4
49.0
88.2
25.0
43.00
13.2
60.8
49.2
86.0
5.8
63.80
66.1
74.5
57.0
93.3
28.1
Qwen3-14B-Base
60.56
50.4
88.8
56.0
93.2
14.4
Qwen3-30B-A3B-Base
56.48
43.2
85.6
45.2
96.4
12.0
Qwen3-30B-A3B-Instruct
66.88
65.6
90.0
58.4
91.6
28.8
Qwen3-32B-Base
60.68
51.8
87.8
57.0
90.6
16.2
C.6
Cipher
RL Budget: Interpretive Limitations and Training Instability
The small RL gains observed across all three experiments (+1.1–+2.3 pp from GSPO; the largest, +2.30 pp, occurs under the exploratory θ∗ allocation) raise the question of whether a larger RL budget would alter the compensatory conclusion. Two interpretations are consistent with the data. Interpretation 1 — Competence-limited RL. The 200-step GSPO stage may be genuinely limited by the model’s current competence: verifier-based RL reinforces existing chains rather than discovering new symbolic procedures, so domains that mid-training left weak receive little RL gain regardless of budget. Interpretation 2 — Insufficient RL budget. 200 steps may simply be insufficient to manifest RL’s corrective potential; a longer schedule could in principle reshape the domain gaps by giving the model more opportunities to explore and for the verifier to surface correct but low-probability chains in weak domains. We cannot distinguish these interpretations from our experiments, and we treat the RL leg as a secondary observation rather than a headline claim for two reasons. First, the reward is a binary correctness verifier, so no dense per-token reward signal can be extracted from the chain-of-thought; RL can only reinforce whole-answer outcomes. Second, extending the GSPO budget is not a cleanly tunable knob in practice: longer schedules were unstable across seeds (reward divergence without a consistent directional trend) rather than simply maturing into a stable higher-performing policy. Two distinct RL settings appear in this paper and we separate them explicitly. (i) The three full-pipeline
20
configurations of Table 1 use the standard pipeline (mid-training → rule-type-balanced SFT → RL). (ii) The six heldout allocations are carried through mid-training → compensatory SFT → RL, and it is these runs that produce the RL Evaluation rows of Table 13 and the out-of-sample results reported in §4.1 (mean +1.19 pp over the compensatory pass; 0/60 pairs closed at the 5 pp threshold and 7/60 at the 10% ratio after the complete pipeline). Early exploratory attempts to add RL after compensatory SFT on the 24 sweep configurations showed cross-seed instability and are excluded from the reported sweep results, under the same criterion used elsewhere (reward divergence or no consistent directional trend within the 200-step budget); that exclusion applies to those sweep runs only and not to the six heldout allocations, whose RL runs are reported in full. Whether a compensatory-SFT-plus-RL schedule would close gaps across the whole 24-configuration sweep therefore remains open, while on the six held-out allocations it demonstrably does not. We scope the RL stage to the tested 200-step schedule and do not read its small gains as evidence about post-training repair in general. Figure 5 illustrates the pattern on the shared no-mid-training baseline run: the per-step group-mean verifier reward fluctuates between ≈ 62% and 67% across the 200 steps without a consistent upward trend, and the evaluation gain over the whole stage is small (+1.20 pp, SFT-only 64.08% to SFT+RL 65.28%).
GSPO reward trajectory of the no-mid-training baseline run
Per-step verifier reward (%)
70
endpoints: SFT-only 64.08% / SFT+RL 65.28%
68 66 64 62 60
0
25
50
75
100 GSPO step
125
150
175
200
Figure 5: Per-step group-mean verifier reward of the GSPO stage for the shared no-mid-training baseline run (SFT→RL, 200-step schedule). Red endpoints mark the SFT-only evaluation accuracy at step 0 (64.08%) and the final SFT+RL accuracy at step 200 (65.28%). The reward wanders in a ≈62–67% band without a consistent upward trend, matching the small evaluation gain of the RL stage (+1.20 pp).
C.7
Code and Data Availability
To support replication, we will release the training configurations, the data-generation pipeline (rule library, synthesizer, deterministic solvers, verification, and preprocessing), the evaluation harness (prompt templates, answerextraction rules, and difficulty stratification), the analysis scripts that produce every reported table and figure, and per-seed results for all reported tables, together with the RL reward trajectories of all runs, including those excluded under the criteria of Appendix C.6.
C.8
Pipeline and Study-Roadmap Figures
The synthetic data-generation pipeline is reproduced here.
21
Rules (Rule Library) Cipher rules
Compute answer via deterministic algorithms
Generate a structured problem instance based on rule and difficulty
Input:
e.g., Caesar cipher, Substitution cipher,...
5a. Deterministic Solver
3. Rule Synthesizer
2. Domain Dispatcher
domain / rule / difficulty / count
e.g., Caesar cipher (shift = 3):
Structured Problem (Example)
Operation rules e.g., Arithmetic, Expression eval,...
id
:
000001
domain
:
cipher
KHOOR ZRUOG
rule_id
:
caesar_cipher
difficulty
:
D2
HELLO WORLD
Answer (ground truth)
4. Difficulty Sampling
input_data
:
“KHOOR ZRUOG”
Logic rules
Control problem scale & complexity
ground_truth
:
“HELLO WORLD”
5b. Teacher Model
e.g., Propositional logic, Syllogism,...
D1 (Easy)
reasoning_trace
:
[step1,step2, ...]
Generate / assist answer
metadata
:
{ ... }
additional_info
:
{ ... }
: 1-3 chars / small scale
D2 (Medium) : 2-8 chars / medium scale D3 (Hard)
: 8-12 chars / large scale
Counterfactual rules
...
(when solver is not available)
e.g., What-if reasoning, Causal change,...
Puzzle rules
7. Data Preprocesser
6. Vertification
e.g., Sudoku, Grid, Maze,...
.. .
Extract
Compare with
Answer
Ground Truth
Deduplicate
Format
Accepted
Check
Sample
. Normalize fields . Filter by quality . Split / balance . ...
Fine Dataset (for training)
(b) Example of One Generated Sample Rule
Difficulty
Input
Cipher rule
D2 (Medium)
Caesar cipher (shift = 3)
2-8 chars / medium scale
"KHOOR ZRUOG" (encrypted text)
Ground Truth Answer
Solved by Deterministic Solver (Caesar shift= 3)
"HELLO WORLD"
Verification Math Accepted
Output
Stored in Fine Dataset
Figure 6: Synthetic data generation pipeline: symbolic solver produces (problem, answer) pairs from KOR-Bench rule definitions, passing through verification, split isolation, and deduplication.
Table 6: Per-domain accuracy change (∆, pp) vs. the same-stage no-mid-training baseline at each pipeline stage, for Expt. 1 (imbalanced). Mid-only vs. Base; Mid+SFT vs. SFT only; Mid+SFT+RL vs. SFT+RL. All entries in pp. The mid-training-only collapse (Operation −27.2, Counterfactual −38.0) is reversed in sign by SFT/RL—strongly for Operation, and within seed noise for Counterfactual; a smaller residual deficit persists on Logic and Puzzle. Rows are five-seed means. Domain Cipher Operation Logic Counterfactual Puzzle Overall
Mid-only vs. Base
Mid+SFT vs. SFT only
Mid+SFT+RL vs. SFT+RL
+31.3 ± 1.8 −27.2 ± 1.1 +7.3 ± 0.9 −38.0 ± 2.4 +8.0 ± 0.5
+4.0 ± 1.0 +5.2 ± 0.3 −6.0 ± 0.8 −0.8 ± 1.2 +1.2 ± 0.3
+2.4 ± 4.4 +4.0 ± 2.8 −2.4 ± 2.5 +1.2 ± 3.4 −2.0 ± 3.0
−3.72
+0.72
+0.64
22
Compensatory-SFT gain vs. remedial ratio
Compensatory-SFT gain Δ (pp)
8
Per-domain fits Δ = alog(1 + b r) + c (parameters in the appendix; pooled R 2 ≈ 0.31 over 115 cells)
6
4 Domain Cipher Operation Logic Counterfactual Puzzle r = 0 (excluded)
2
0 0
5
10 15 20 Remedial SFT ratio r (%)
25
30
Figure 7: Compensatory-SFT gain ∆ = Eval. (Remedial)−Evaluation vs. the domain’s remedial SFT ratio r, from Table 13 (115 cells with r > 0, coloured by domain; the five r = 0 cells are open diamonds and excluded from the fits). Solid curves: per-domain fits ∆ ≈ a log(1 + b r) + c (Table 16). The gain rises with r and saturates at a domain-specific ceiling; pooling across domains leaves most cell-level variance unexplained (pooled R2 ≈ 0.31). Descriptive, not confirmatory.
23
D
Data Accounting
Coverage as measured in this study is the token share of each domain within the internal five-domain component (Table 13, Portion rows). Because the downstream recipe and the total token budget are held fixed, the coverage percentages fully determine the per-domain token allocation of that component; Table 7 records the corresponding allocation for the two configurations used in the full-pipeline comparison (Expt. 1 and the balanced allocation). The composition and provenance of the fixed external corpus component are described in §2.2; the internal/external token split and the per-component budgets are fixed across all configurations. Table 7: Per-configuration data accounting for the mid-training coverage sweep. Coverage percentages are the measured internal-component token shares (Table 13); the internal/external token ratio and the per-component token budgets are fixed across all 24 configurations. Quantity
Cipher
Operation
Logic
Counterfactual
Puzzles
Internal token share (e.g. balanced config) Internal token share (e.g. Expt. 1)
20% 72.3%
20% 0%
20% 12.0%
20% 0%
20% 15.7%
E
Decontamination and Split Isolation 1. ProofWriter-family depth split. The external mid-training corpus contains only regenerated ProofWriterfamily items of depth at most 5; evaluation uses depths 5 to 9. Depth 5 is explicitly labeled as in distribution. 2. KOR-Bench instance split. SFT and evaluation data use different random seeds. No identical instance appears in both; evaluation instances are held out by instance ID. 3. External benchmark isolation. No CounterBench-family items and no ProofWriter-family items of depths 6–9 appear in any training corpus. ZebraLogic-family items contribute 201 mid-training samples (0.3% coverage); these are distinct instances from the evaluation set but same-family exposure, so ZebraLogic is not a zeroexposure benchmark (Appendix H). 4. Base model contamination. We do not control for Qwen3-8B-Base’s original pretraining exposure to public benchmarks. This limitation applies to all work with pretrained models. 5. Teacher trace isolation. SFT chain-of-thought traces are generated by the pipeline’s symbolic solvers, not by any LLM. No teacher model accesses evaluation instances.
F
Mid-Training Coverage Sweep
Table 13 reports the full per-configuration results of the 24 coverage sweep variants (§4.1) together with the FineWebEdu baseline and the exploratory θ∗ allocation. Each row block gives: Portion (domain allocation), Evaluation (midtraining-only accuracy, five-seed mean ± 1 SD), Remedial SFT (compensatory allocation from Eq. 5), Eval. (Remedial) (accuracy after the compensatory SFT pass), and Eval. (Uniform) (accuracy after the matched uniform-SFT control; Appendix G). All 24 configurations include both SFT-pass results. The external corpus derives 34.8% of its tokens from the ProofWriter rule family, contributing shared formal-deduction exposure that explains why Logic’s curve sits above zero at 0% internal coverage, while Puzzle collapses to near-zero. Per-domain fitted-region interpretation. Cipher, Operation, and Logic have broad 95%-of-peak intervals (upper bounds near 30–36%), consistent with their large fitted σR values; Cipher’s interval [7.4%, 36.0%] spans the wide plateau observed in Figure 2 rather than a sharp peak. Counterfactual’s fitted 95%-of-peak interval [17.9%, 47.7%] is consistent with a broad fitted peak near 26.0%. Puzzle’s interval [27.2%, 41.2%] requires relatively high coverage; its fitted peak accuracy (15.0%) is low, and whether a larger model could raise it remains open.
24
Simplex design validity. Each domain’s x-coordinate in Figure 2 is its own internal proportion while the other four vary jointly. The curves describe exploratory mixture-level associations rather than isolated causal functions. A compositional or simplex-aware joint response-surface analysis (in the sense of compositional data analysis (Aitchison, 1986)) would be needed for stronger allocation claims; we report a preliminary such analysis in Appendix F.2, which finds a saddle rather than an interior maximum even after the 12 interior allocations of Appendix B.2 are added to the pool. A factorial design that independently varies each domain’s coverage remains future work. Any monotonicity or concavity tests below are likewise descriptive diagnostics, and should not be read as population-level causal tests. Concavity test (descriptive). We fit a quadratic yd = β0 + β1 θ + β2 θ2 per domain (n = 24) as an exploratory shape summary. The negative β̂2 values in Table 8 indicate a concave trend within this sweep, but they are reported strictly as descriptive shape descriptors and are not tests of curvature: accuracy is a bounded proportion, observations share training configurations, and the five domain fits are not independent. A binomial/beta-binomial or clustered compositional analysis with multiplicity control would be required before making inferential claims about curvature; such an analysis is left for future work. Table 8: Per-domain quadratic regression concavity check, descriptive only. OLS fit y = β0 + β1 θ + β2 θ2 on all 24 normal coverage configurations per domain. β̂2 and its standard error are reported as shape descriptors only; the standard error is the ordinary-least-squares dispersion of the coefficient and is not used as a curvature test. The calibrated test of interiority is the permutation test of Appendix F.1 (see §F). β̂2 and SE are rounded independently from unrounded estimates. Domain Cipher Operation Logic Counterfactual Puzzle
F.1
β̂2
SE(β̂2 )
2 Rquad
−0.013 −0.048 −0.033 −0.035 −0.012
0.003 0.004 0.004 0.002 0.001
0.494 0.876 0.816 0.987 0.956
Model-Free Check: Coverage Bands
The split-Gaussian summaries below are five-parameter fits, so a reader may reasonably ask how much of the invertedU shape is a property of the model rather than of the data. We therefore repeat the moderation claim without any curve fitting. Pooling all 30 trained allocations (the 24 sweep configurations plus the six held-out allocations), we bin each domain’s own coverage into low (< 10%), moderate (10–40%), and high (> 40%) and compare the mean midtraining-only accuracy in each band (Table 9). For every one of the five domains the moderate band is the best of the three, and accuracy is lower both below and above it. If the best-of-three band were exchangeable across domains, the probability that all five select the moderate band would be (1/3)5 ≈ 0.004; because the five domains share the same 30 allocations they are not independent trials, so we report this as a descriptive concordance rather than a calibrated test. The high band is also thinly populated for several domains (n = 1–5), so the low-side contrast is the better-supported half of the pattern; the 12 interior allocations of Appendix B.2, all in the moderate band, do not add high-coverage points but thicken the moderate band (to n = 29–32), which is the side the moderation pattern rests on. Within these limits, the interior optimum is visible in the raw group means and does not depend on the parametric fit, and it survives the addition of the 12 interior allocations (moderate best for 5/5 domains over the 42-allocation pool). One bookkeeping point deserves stating plainly: the six held-out allocations serve two roles. They are withheld from the split-Gaussian curve fitting, which is what makes the out-of-sample test of §4.1 valid, but they are included in the 30-allocation pool used for the band comparison above, for the interiority permutation test, and for the compositional surface of §F.2, because those analyses are not fitted to the curves and benefit from the extra design points. No statistic is therefore both fitted on and validated against the same allocations, but the band and interiority results are in-sample with respect to the pool (30 or, once Appendix B.2 is included, 42) and should be read as such. A permutation test sharpens this into a calibrated statement. For each domain we fit a three-parameter quadratic OLS to the 42-allocation pool (30 base plus the 12 interior allocations of Appendix B.2) and ask whether it is concave (β̂2 < 0) with a vertex strictly inside the tested coverage range; all five domains satisfy both conditions. Under a null that permutes whole accuracy rows across allocations—destroying the coverage–accuracy link while leaving the simplex design and the cross-domain correlation structure exactly intact—an average of 1.78 of five domains do so, and all five do so in 25
0.98% of 20,000 draws (P ≈ 0.010). Shuffling each domain independently gives the same answer (P ≈ 0.005). We use the quadratic here only to test concavity and interiority, not to locate the optima: a symmetric parabola pulls its vertex toward the centre of the coverage range, so its vertex positions are not comparable to the split-Gaussian peaks reported in Table 10. Two further permutation results bear on how much the split-Gaussian fit itself can be trusted, and we report both, including the one that cuts against us. First, the fitted curves capture real structure: under the same whole-row permutation the mean per-domain R2 of the split-Gaussian is 0.174 ± 0.055, against 0.982 for the observed data (P < 0.001, 250 draws). Second, and less favourably, the count of interior peaks is not a usable statistic for a five-parameter split-Gaussian: a permuted dataset still yields interior peaks in all five domains 64% of the time, because a flexible asymmetric curve can place a maximum inside the range even on noise. This is precisely why the interiority test above is run on the three-parameter quadratic, which cannot manufacture an interior optimum so easily, and why we do not treat “all five peaks are interior” as evidence on its own. Finally, the five fitted peaks sum to ≈ 102%, i.e. they are close to jointly realisable within the 100% budget. A referee might read a sum near the budget as the fingerprint of a simplex artifact; the permutation null does not support that reading, since permuted data give a peak sum of 121.5 ± 33.3% that does not concentrate near 100%, and only 2.8% of draws land as close to the budget as the observed sum. We nonetheless treat the near-feasibility as a numerical coincidence worth reporting rather than as a design rule, because the sum of five marginal optima has no guaranteed relationship to the joint optimum under the constraint (§5.2). Table 9: Mean mid-training-only accuracy (%) by coverage band, pooling all 30 trained allocations (24 sweep + 6 held-out); n is the number of allocations falling in each band for that domain. The moderate band is the best of the three for all five domains. No curve fitting is involved. Domain
Low (< 10%)
Moderate (10–40%)
High (> 40%)
Cipher Operation Logic Counterfactual Puzzle
43.8 (n=9) 51.9 (n=11) 39.6 (n=9) 57.5 (n=9) 6.1 (n=5)
51.0 (n=18) 74.7 (n=17) 58.1 (n=20) 81.2 (n=20) 13.1 (n=20)
42.2 (n=3) 50.4 (n=2) 40.0 (n=1) 78.2 (n=1) 10.5 (n=5)
F.2
Compositional (Aitchison) Reanalysis
Because coverage is a closed composition, an analysis that treats the five shares as free variables is not strictly appropriate, and we report here the compositional analysis the simplex discussion (§5.2) calls for. We map each allocation to four isometric log-ratio (ilr) coordinates after multiplicative replacement of zero shares, and regress each domain’s mid-training-only accuracy on a response surface in ilr space over the 42-allocation pool (the 30 sweep and held-out allocations plus the 12 interior allocations of Appendix B.2). The result is informative in both directions. A full quadratic ilr surface fits well (R2 = 0.78–0.96 per domain over the 42-allocation pool), which confirms that allocation carries real signal when modelled compositionally rather than marginally. But its stationary point is a saddle in every domain, not an interior maximum, and with 15 quadratic terms estimated from 42 allocations the Hessian is too weakly determined for that classification to carry weight. A reduced surface without cross terms (9 parameters) is more stable and predominantly concave: 15 of the 20 ilr curvature coefficients are negative (Logic 4 of 4; Cipher, Operation and Counterfactual 3 of 4; Puzzle 2 of 4), so the surface curves downward in most directions but not all. We therefore do not claim a joint interior optimum over the simplex. The moderation result of Appendix F.1 is a statement about per-domain marginals—each domain’s own coverage band—and it stands on the model-free band comparison and the permutation test, neither of which requires a joint surface. Establishing a joint compositional optimum would require more allocations than the 42 trained here. The 12 interior allocations of Appendix B.2 were added precisely to test this and did not produce a joint maximum—every domain’s stationary point remained a saddle— so we flag additional simplex coverage, not a different estimator, as the natural next design rather than a gap the present data can close.
26
F.3
Fitting Procedure
Each domain’s dose–response curve is summarised by the following asymmetric linear-space Gaussian (split Gaussian): 2 θ−µ + c, f (θ) = a exp − 12 σ(θ) ( σ(θ) =
(3)
σL , θ ≤ µ, σR , θ > µ,
where the two widths σL and σR are free parameters with no ordering imposed, with five free parameters (a, µ, σL , σR , c) per domain. In the fitted curves, σR >σL (steeper left rise than right decline) for four domains, while Puzzle’s fitted curve has σL >σR . Parameters are estimated by weighted nonlinear least squares (scipy.optimize.curve_fit). Configurations in which a domain’s internal proportion is θd ≤ 5% or θd ≥ 55% are assigned half-weight (w = 0.5); all others receive w = 1.0. This heuristic down-weighting reduces the leverage of extreme, near-degenerate points, but the exact thresholds are not claimed to be optimal. Peak locations µ and peak accuracies per domain are listed in Table 10. The 95%-of-peak moderate region for each domain is the set {θd | fd (θd ) ≥ 0.95·fd (µ)}, solved numerically from the fitted split-Gaussian curve (Eq. 3); the resulting intervals are listed in Table 10. Domain Cipher Operation Logic Counterfactual Puzzle
Peak (%)
µ (%)
95%-band
52.0 78.1 59.5 83.5 15.0
9.9 15.9 15.0 26.0 35.1
[7.4, 36.0] [11.9, 29.4] [11.1, 31.4] [17.9, 47.7] [27.2, 41.2]
Table 10: Fitted 95%-of-peak intervals from split-Gaussian summaries. Each interval is the θd range satisfying fd (θd ) ≥ 0.95 · fd (µ); the threshold is inclusive (≥) and bounds are displayed rounded to one decimal place. These are descriptive function-level thresholds from the Qwen3-8B/KOR-Bench sweep under the simplex constraint (see Appendix F); they are not confidence intervals and not validated allocation rules. For Cipher, the fitted curve is dominated by a broad plateau (≈51–53% over ≈8–24% coverage); the fitted peak at ≈9.9% is not a distinct mode and should not be read as a sharp optimum.
Sensitivity to the down-weighting threshold. Because the w = 0.5 threshold at θd ≤ 5% or θd ≥ 55% is a heuristic, we refit the split-Gaussian across four alternative thresholds (≤ 0/ ≥ 100, ≤ 3/ ≥ 57, ≤ 5/ ≥ 55, ≤ 8/ ≥ 52). The fitted peaks are essentially unchanged: Cipher µ ∈ [9.6, 9.9], Operation [15.9, 15.9], Logic [14.8, 15.3], Counterfactual [26.0, 26.2], Puzzle [34.9, 35.4]—a spread of at most 0.5 pp. The reported peaks are therefore not an artifact of this particular down-weighting choice. The θd ≤ 5% and θd ≥ 55% thresholds are used throughout. The illustrative allocation θ∗ reported in §4.1 is obtained by solving: θ∗ = arg max θ
s.t.
X fd (θd ) − bd d
X
bd
(4)
θd = 100%, θd ≥ 0,
d
where bd is the FineWeb-Edu-only baseline accuracy, solved with Sequential Least Squares Programming (SLSQP; 2,000 random initialisations). Because this objective is defined on fitted marginal curves, the resulting allocation should be interpreted as a descriptive probe rather than a globally optimal mixture. Goodness of fit. The split-Gaussian substantially outperforms a flat baseline on every domain. Weighted Root Mean Square Error (RMSE, in pp), split-Gaussian vs. weighted constant: Cipher 0.84 vs. 5.32; Operation 1.15 vs. 12.51; Logic 1.44 vs. 9.10; Counterfactual 1.18 vs. 10.60; Puzzle 0.63 vs. 3.28. 27
Model form comparison. Table 11 compares the split-Gaussian with a quadratic OLS alternative (y = β0 + β1 θ + β2 θ2 , three parameters). The quadratic recovers concave shape with fewer assumptions and is the form we use for the calibrated interiority test of Appendix F.1, precisely because its three parameters cannot manufacture an interior optimum from noise; the β̂2 values of Table 8 are shape descriptors accompanying that test, not a curvature test in themselves. For describing the curves, the split-Gaussian is preferred on evidence rather than flexibility: it attains both the lower weighted in-sample RMSE and the lower leave-one-out MAE in all five domains, and the LOO-MAE comparison does not reward its extra parameters. It additionally captures the left–right asymmetry the quadratic cannot represent. Both forms agree that the curves are non-monotonic within this sweep, with the strongest fit for Counterfactual. Table 11: Model fit comparison for the per-domain coverage–accuracy curves (n = 24). All three forms are scored on the same footing: RMSE is the weighted in-sample RMSE under the identical weighting scheme (Appendix F.3), and LOO-MAE is leave-one-out mean absolute error, which is the out-of-sample criterion and does not reward extra parameters. The split-Gaussian attains the lower LOO-MAE in all five domains, so its advantage is not an artifact of its larger parameter count. RMSE and LOO-MAE in pp. Flat baseline
Quadratic OLS
Split-Gaussian
Domain
RMSE
LOO-MAE
RMSE
LOO-MAE
RMSE
LOO-MAE
Cipher Operation Logic Counterfactual Puzzle
5.32 12.51 9.10 10.60 3.28
4.28 12.34 8.65 10.49 3.16
3.95 5.20 4.28 1.72 0.71
4.22 7.67 6.36 1.95 0.66
0.84 1.15 1.44 1.18 0.63
1.10 1.01 1.16 1.09 0.55
Leave-one-out cross-validation. We apply LOO-CV over the 24 normal non-baseline configurations. For each fold we refit the split Gaussian on the other 23 configurations using the same deterministic multi-start procedure (four σ seed pairs, two intercept seeds, and five µ seeds drawn from the sorted-top-3 targets plus a uniform grid) and record the prediction error on the held-out configuration. Table 12 reports LOO-MAE, the range of fitted µ across folds (∆µ), and the fraction of held-out points within ±2×RMSEfull (in-band), all over the 24 normal non-baseline configurations. LOO-MAE is small and tight (0.55– 1.16 pp), and peak location is stable across folds for every domain (∆µ ≤ 1.78 pp). These diagnostics assess interpolation for the finite sweep only; they do not validate peak locations, moderate bands, or the jointly selected θ∗ on new mixtures. Table 12: Leave-one-out cross-validation results for the split-Gaussian dose-response fits. LOO-MAE: mean absolute error on held-out; µfull : peak from full-data fit; ∆µ LOO: range of µ across LOO folds; In-band: fraction within ±2×RMSEfull . Domain Cipher Operation Logic Counterfactual Puzzle
LOO-MAE (pp)
µfull (%)
∆µ LOO (pp)
In-band
1.10 1.01 1.16 1.09 0.55
9.9 15.9 15.0 26.0 35.1
1.04 0.94 1.78 0.58 1.11
21/24 (88%) 21/24 (88%) 20/24 (83%) 22/24 (92%) 22/24 (92%)
One sweep configuration (Cipher 26%, Operation 16%, Logic 17%, Counterfactual 30%, Puzzle 11%), added after an initial set of eight configurations to broaden simplex coverage, realizes 82.4% Counterfactual accuracy at 30% Counterfactual coverage—close to, but not above, the 83.9%, 83.7%, and 83.6% attained near 29%, 32%, and 24% Counterfactual coverage. Because several configurations across 24–32% Counterfactual coverage cluster around this level, the fitted Counterfactual peak (≈26%; Table 10) is supported by a group of points rather than by a single posthoc observation, and no single configuration establishes a higher peak. The pattern remains directionally consistent with the relatively higher Counterfactual allocation in the illustrative θ∗ (a separate, analytically derived allocation trained as an independent checkpoint outside the 24 sweep configurations; Table 13), but a single configuration does not validate the fitted Counterfactual peak. To confirm that this design point does not drive the fitted curves, we refit all
28
five domains with it removed. The peaks move by at most 0.20 pp (Cipher 9.95 → 10.00, Operation 15.90 → 15.95, Logic unchanged at 14.95, Counterfactual 26.05 → 26.25, Puzzle unchanged at 35.10) and their sum changes from 101.95% to 102.25%, so every peak location and every 95%-of-peak interval reported here is materially unchanged by excluding the post-hoc configuration. Table 13: Mid-training coverage sweep results with the compensatory-SFT pass and its uniform-SFT control. Each row block reports: Portion (domain coverage proportions), Evaluation (mid-training-only accuracy), Remedial SFT (compensatory allocation from Eq. 5), Eval. (Remedial) (accuracy after the compensatory pass), and Eval. (Uniform) (accuracy after the matched uniform control, whose data mix is identical across configurations; Appendix G); the six held-out blocks additionally carry an RL Evaluation row (accuracy after the full mid-training+SFT+RL pipeline, Appendix G). Values are reported as mean ± 1 SD across the five seeds. The Exploratory θ∗ block (Expt. 3 in Table 1) reports the analytically derived θ∗ allocation (9.7/11.8/15.8/30.4/32.3), its target Remedial SFT allocation from Eq. 5, and the results of the dedicated re-run (five seeds); it does not coincide with any of the 24 sweep configurations. The FineWeb-Edu baseline block precedes six held-out allocations that are withheld from the curve fit and used as out-of-sample validation (mid-training-only, SFT-stage and RL-stage results; §4.1, Appendix G). Domains
Ciphers
Operations
Logic
Counterfactual
Puzzles
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
72.3% 38.1 ± 2.6% 0.0%† 38.3 ± 0.5% 41.5 ± 0.7%
0.0% 33.2 ± 2.8% 28.3% 41.5 ± 0.4% 41.6 ± 0.3%
12.0% 54.5 ± 2.2% 22.6% 58.4 ± 0.8% 59.3 ± 0.5%
0.0% 45.6 ± 2.8% 28.3% 53.5 ± 0.8% 50.8 ± 0.2%
15.7% 11.6 ± 1.5% 20.8% 13.2 ± 0.3% 13.8 ± 0.7%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
26.0% 50.5 ± 2.1% 17.0% 53.0 ± 0.6% 54.2 ± 0.7%
16.0% 77.9 ± 2.3% 22.0% 83.8 ± 0.4% 83.4 ± 0.3%
17.0% 59.8 ± 1.9% 21.5% 64.6 ± 1.0% 64.4 ± 1.0%
30.0% 82.4 ± 3.2% 15.0% 87.7 ± 0.8% 86.5 ± 0.8%
11.0% 8.9 ± 1.0% 24.5% 10.8 ± 0.3% 11.2 ± 0.9%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
30.0% 50.9 ± 2.9% 15.0% 54.2 ± 0.6% 54.6 ± 0.4%
18.0% 77.6 ± 2.8% 21.0% 83.3 ± 0.3% 83.0 ± 0.3%
16.0% 60.4 ± 1.6% 22.0% 65.2 ± 0.6% 65.1 ± 1.1%
18.0% 79.3 ± 1.8% 21.0% 86.0 ± 1.0% 83.5 ± 0.8%
18.0% 12.0 ± 1.4% 21.0% 13.9 ± 0.5% 14.1 ± 0.6%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
20.0% 51.1 ± 2.2% 20.0% 53.5 ± 0.9% 54.9 ± 0.7%
20.0% 77.2 ± 2.9% 20.0% 82.9 ± 0.3% 82.6 ± 0.5%
20.0% 59.3 ± 2.2% 20.0% 63.7 ± 0.5% 63.9 ± 0.7%
20.0% 81.2 ± 3.1% 20.0% 87.8 ± 0.9% 85.4 ± 0.6%
20.0% 12.5 ± 2.0% 20.0% 14.3 ± 0.2% 14.6 ± 0.8%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
0.0% 22.0 ± 2.3% 30.0% 26.6 ± 0.7% 26.8 ± 0.8%
10.0% 67.3 ± 3.8% 25.0% 73.6 ± 0.3% 72.9 ± 0.2%
40.0% 52.8 ± 1.7% 10.0% 55.6 ± 1.0% 57.2 ± 0.4%
10.0% 66.9 ± 2.0% 25.0% 74.7 ± 0.5% 71.3 ± 0.8%
40.0% 15.2 ± 0.8% 10.0% 15.9 ± 0.4% 17.1 ± 1.0%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
10.0% 51.2 ± 2.5% 25.0% 55.2 ± 0.8% 55.2 ± 0.8%
8.0% 64.5 ± 3.3% 26.0% 71.9 ± 0.5% 70.2 ± 0.8%
37.0% 54.3 ± 2.5% 11.5% 57.4 ± 0.5% 58.7 ± 0.4%
8.0% 65.2 ± 2.8% 26.0% 73.8 ± 0.6% 69.7 ± 0.9%
37.0% 14.9 ± 1.0% 11.5% 15.8 ± 0.1% 16.8 ± 0.7%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
5.0% 45.7 ± 1.6% 27.5% 50.2 ± 0.5% 49.9 ± 0.5%
15.0% 77.7 ± 3.1% 22.5% 83.7 ± 0.4% 83.2 ± 0.6%
5.0% 44.8 ± 2.8% 27.5% 50.3 ± 0.7% 49.8 ± 0.8%
15.0% 75.7 ± 2.3% 22.5% 82.7 ± 0.4% 80.0 ± 0.7%
60.0% 5.3 ± 0.9% 0.0% 5.0 ± 0.5% 7.0 ± 0.8%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
50.0% 47.0 ± 1.8% 5.0% 48.7 ± 0.7% 50.5 ± 0.5%
40.0% 65.7 ± 2.7% 10.0% 69.4 ± 0.5% 70.9 ± 0.6%
0.0% 22.1 ± 2.7% 30.0% 27.4 ± 0.8% 27.7 ± 1.0%
5.0% 57.5 ± 3.0% 27.5% 65.5 ± 0.4% 62.6 ± 0.9%
5.0% 8.0 ± 1.1% 27.5% 10.3 ± 0.1% 10.5 ± 0.9%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
40.0% 49.5 ± 1.4% 10.0% 52.6 ± 0.8% 53.1 ± 0.4%
4.0% 47.7 ± 3.4% 28.0% 54.4 ± 1.0% 55.6 ± 0.4%
8.0% 50.2 ± 1.9% 26.0% 55.2 ± 0.6% 55.1 ± 0.5%
40.0% 82.0 ± 1.6% 10.0% 86.1 ± 1.1% 86.0 ± 0.9%
8.0% 8.2 ± 0.7% 26.0% 10.4 ± 0.1% 10.6 ± 0.9% Continued on next page
29
Table 13 (continued) Domains
Ciphers
Operations
Logic
Counterfactual
Puzzles
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
8.0% 49.5 ± 1.9% 26.0% 53.8 ± 0.7% 53.6 ± 0.6%
36.0% 71.8 ± 3.2% 12.0% 76.1 ± 0.4% 77.0 ± 0.7%
2.0% 33.7 ± 2.2% 29.0% 39.3 ± 0.5% 38.9 ± 0.9%
2.0% 52.8 ± 2.2% 29.0% 60.2 ± 0.8% 60.6 ± 0.4%
52.0% 9.6 ± 1.1% 4.0% 9.8 ± 0.4% 11.4 ± 0.6%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
9.0% 52.8 ± 1.3% 25.5% 57.1 ± 0.9% 56.8 ± 0.4%
14.0% 77.9 ± 2.4% 23.0% 83.9 ± 0.9% 83.4 ± 0.6%
14.0% 59.2 ± 2.6% 23.0% 64.3 ± 0.5% 63.9 ± 0.7%
17.0% 78.3 ± 3.4% 21.5% 85.2 ± 1.1% 82.5 ± 0.4%
46.0% 12.6 ± 0.7% 7.0% 13.4 ± 0.4% 14.4 ± 0.6%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
11.0% 51.1 ± 2.7% 24.5% 55.3 ± 0.9% 55.1 ± 0.5%
2.0% 41.8 ± 3.4% 29.0% 46.8 ± 0.8% 47.8 ± 0.3%
15.0% 59.9 ± 2.5% 22.5% 64.8 ± 1.2% 64.6 ± 0.6%
28.0% 83.4 ± 3.6% 16.0% 89.2 ± 0.9% 87.5 ± 0.3%
44.0% 13.3 ± 0.8% 8.0% 14.1 ± 0.4% 15.1 ± 0.7%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
4.0% 40.3 ± 1.1% 28.0% 44.8 ± 0.6% 44.6 ± 0.3%
3.0% 45.6 ± 2.8% 28.5% 52.3 ± 0.7% 51.5 ± 0.1%
30.0% 57.3 ± 2.1% 15.0% 61.0 ± 1.0% 61.8 ± 0.3%
29.0% 83.9 ± 3.8% 15.5% 89.4 ± 0.6% 88.0 ± 0.5%
34.0% 15.4 ± 1.3% 13.0% 16.6 ± 0.4% 17.3 ± 0.5%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
15.0% 52.0 ± 2.9% 22.5% 56.0 ± 0.5% 55.9 ± 0.8%
6.0% 57.3 ± 2.3% 27.0% 63.8 ± 0.7% 63.1 ± 0.8%
19.0% 59.4 ± 2.9% 20.5% 64.0 ± 1.2% 64.0 ± 0.5%
32.0% 83.7 ± 1.7% 14.0% 88.9 ± 1.1% 87.7 ± 0.6%
28.0% 14.7 ± 1.4% 16.0% 16.2 ± 0.2% 16.7 ± 0.9%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
60.0% 41.6 ± 1.4% 0.0% 41.6 ± 0.8% 45.0 ± 0.3%
22.0% 76.9 ± 2.1% 19.0% 81.3 ± 0.8% 82.3 ± 0.6%
13.0% 57.6 ± 2.4% 23.5% 62.6 ± 0.8% 62.3 ± 0.8%
3.0% 52.1 ± 2.4% 28.5% 60.1 ± 1.0% 56.8 ± 0.6%
2.0% 4.2 ± 1.0% 29.0% 6.6 ± 0.1% 6.9 ± 0.7%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
12.0% 52.9 ± 1.7% 24.0% 57.1 ± 0.6% 56.9 ± 0.5%
60.0% 43.5 ± 3.2% 0.0% 43.5 ± 0.7% 48.5 ± 0.4%
24.0% 58.6 ± 1.8% 18.0% 62.8 ± 1.2% 63.1 ± 0.6%
1.0% 52.1 ± 2.2% 29.5% 60.2 ± 0.7% 57.1 ± 0.5%
3.0% 4.2 ± 1.2% 28.5% 6.6 ± 0.2% 6.8 ± 0.8%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
6.0% 45.2 ± 2.3% 27.0% 49.6 ± 0.5% 49.4 ± 0.4%
28.0% 75.2 ± 3.3% 16.0% 80.2 ± 0.5% 80.5 ± 0.4%
28.0% 57.5 ± 1.7% 16.0% 62.4 ± 1.1% 62.0 ± 0.9%
24.0% 83.6 ± 3.9% 18.0% 89.7 ± 0.5% 87.7 ± 1.0%
14.0% 10.1 ± 0.6% 23.0% 12.1 ± 0.5% 12.3 ± 0.5%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
7.0% 48.0 ± 1.6% 26.5% 52.4 ± 0.8% 52.1 ± 0.6%
11.0% 72.8 ± 2.2% 24.5% 79.0 ± 0.5% 78.4 ± 0.5%
18.0% 59.7 ± 2.8% 21.0% 64.4 ± 0.9% 64.3 ± 0.7%
37.0% 83.2 ± 3.7% 11.5% 87.7 ± 0.6% 87.2 ± 0.7%
27.0% 14.4 ± 0.9% 16.5% 15.9 ± 0.5% 16.4 ± 1.0%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
18.0% 52.3 ± 2.5% 21.0% 56.2 ± 0.8% 56.1 ± 0.7%
1.0% 37.7 ± 2.7% 29.5% 44.5 ± 0.5% 43.9 ± 0.8%
60.0% 40.0 ± 1.6% 0.0% 39.9 ± 0.9% 44.2 ± 0.8%
6.0% 64.6 ± 3.8% 27.0% 72.3 ± 0.5% 70.2 ± 0.3%
15.0% 10.2 ± 1.0% 22.5% 12.2 ± 0.2% 12.4 ± 0.9%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
24.0% 51.4 ± 2.6% 18.0% 55.0 ± 0.8% 55.1 ± 0.6%
13.0% 77.9 ± 3.7% 23.5% 84.0 ± 0.6% 83.4 ± 0.7%
7.0% 46.6 ± 2.3% 26.5% 52.9 ± 0.9% 51.5 ± 0.6%
26.0% 83.0 ± 3.1% 17.0% 88.9 ± 0.7% 87.1 ± 0.3%
30.0% 14.8 ± 0.6% 15.0% 16.2 ± 0.2% 16.8 ± 0.8%
Portion Evaluation Remedial SFT
19.0% 51.3 ± 1.2% 20.5%
7.0% 60.2 ± 3.6% 26.5%
19.0% 59.7 ± 2.4% 20.5%
24.0% 82.5 ± 1.7% 18.0%
31.0% 15.1 ± 0.8% 14.5% Continued on next page
30
Table 13 (continued) Domains
Ciphers
Operations
Logic
Counterfactual
Puzzles
Eval. (Remedial) Eval. (Uniform)
55.2 ± 0.9% 55.1 ± 0.6%
66.7 ± 1.0% 65.9 ± 0.7%
64.3 ± 0.7% 64.3 ± 1.0%
88.6 ± 0.5% 86.6 ± 0.5%
16.4 ± 0.3% 17.1 ± 0.6%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
22.0% 50.8 ± 2.1% 19.0% 54.5 ± 0.6% 54.6 ± 0.7%
9.0% 67.4 ± 2.5% 25.5% 73.7 ± 0.6% 73.0 ± 0.2%
1.0% 29.2 ± 2.0% 29.5% 34.9 ± 0.8% 34.6 ± 0.3%
35.0% 82.6 ± 3.6% 12.5% 87.4 ± 0.9% 86.6 ± 0.5%
33.0% 15.0 ± 0.9% 13.5% 16.2 ± 0.5% 16.9 ± 0.8%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
16.0% 52.3 ± 1.9% 22.0% 56.3 ± 0.5% 56.2 ± 0.6%
21.0% 76.7 ± 3.8% 19.5% 82.2 ± 0.3% 82.1 ± 0.2%
3.0% 35.3 ± 2.6% 28.5% 40.9 ± 1.0% 40.4 ± 0.5%
31.0% 84.1 ± 2.6% 14.5% 89.4 ± 1.1% 88.2 ± 0.7%
29.0% 12.7 ± 1.3% 15.5% 14.1 ± 0.1% 14.7 ± 1.0%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform)
33.0% 49.5 ± 2.0% 13.5% 51.9 ± 0.5% 53.1 ± 0.5%
5.0% 57.1 ± 3.0% 27.5% 63.7 ± 0.6% 62.9 ± 0.5%
6.0% 49.0 ± 2.2% 27.0% 55.4 ± 0.6% 54.0 ± 1.0%
50.0% 78.2 ± 1.8% 5.0% 81.7 ± 0.5% 82.1 ± 0.8%
6.0% 5.8 ± 1.4% 27.0% 8.1 ± 0.4% 8.3 ± 0.5%
Exploratory θ∗ (Expt. 3): Cipher 9.7%, Oper. 11.8%, Logic 15.8%, Counterf. 30.4%, Puzzle 32.3%; dedicated re-run, five seeds. Portion 9.7% 11.8% 15.8% 30.4% 32.3% Evaluation 51.2 ± 2.1% 78.0 ± 2.8% 61.4 ± 2.1% 84.8 ± 3.1% 15.8 ± 1.9% Remedial SFT 25.2% 24.1% 22.1% 14.8% 13.8% Eval. (Remedial) 55.7 ± 0.9% 84.2 ± 0.9% 66.2 ± 1.1% 90.3 ± 0.9% 17.2 ± 0.2% Eval. (Uniform) 55.2 ± 0.3% 83.6 ± 0.9% 66.1 ± 0.9% 88.9 ± 0.4% 17.7 ± 0.6% FineWeb-Edu only (100%); no reasoning-domain data. Evaluation 4.4 ± 2.2% 45.6 ± 2.9% 32.0 ± 2.8%
29.2 ± 2.3%
1.2 ± 1.2%
6 held-out allocations (withheld from the curve fit; used for out-of-sample validation). Portion 10.0% 10.0% 15.0% 30.0% Evaluation 50.1 ± 1.7% 69.2 ± 2.4% 60.4 ± 1.2% 82.5 ± 1.6% Remedial SFT 25.0% 25.0% 22.5% 15.0% Eval. (Remedial) 55.3 ± 1.1% 74.8 ± 2.7% 64.9 ± 1.9% 87.7 ± 0.9% Eval. (Uniform) 53.1 ± 1.3% 75.2 ± 1.7% 65.6 ± 1.2% 85.9 ± 1.8% RL Evaluation 55.9 ± 2.1% 75.9 ± 3.2% 66.1 ± 2.8% 88.8 ± 3.9%
35.0% 13.8 ± 1.7% 12.5% 14.9 ± 1.2% 16.1 ± 3.4% 15.2 ± 4.3%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform) RL Evaluation
9.7% 49.8 ± 2.0% 25.2% 54.0 ± 1.5% 53.9 ± 1.2% 55.2 ± 2.2%
50.0% 57.3 ± 1.7% 5.0% 59.1 ± 1.8% 61.8 ± 1.6% 61.0 ± 2.9%
10.3% 55.0 ± 2.3% 24.8% 60.7 ± 1.4% 59.9 ± 1.3% 62.8 ± 4.2%
15.0% 77.2 ± 1.9% 22.5% 84.7 ± 2.1% 81.9 ± 1.5% 85.2 ± 3.1%
15.0% 10.7 ± 0.9% 22.5% 12.6 ± 1.3% 13.0 ± 1.0% 13.3 ± 2.5%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform) RL Evaluation
15.0% 50.5 ± 1.3% 22.5% 54.4 ± 1.2% 54.5 ± 1.9% 55.2 ± 3.1%
11.2% 79.1 ± 2.7% 24.4% 85.0 ± 4.2% 84.9 ± 1.7% 86.1 ± 4.2%
20.0% 58.3 ± 1.9% 20.0% 62.9 ± 1.6% 63.2 ± 1.7% 65.7 ± 4.5%
8.8% 67.4 ± 1.5% 25.6% 75.0 ± 2.8% 72.8 ± 1.9% 75.6 ± 1.8%
45.0% 11.8 ± 1.2% 7.5% 12.2 ± 1.4% 13.7 ± 2.5% 13.2 ± 3.5%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform) RL Evaluation
4.2% 40.8 ± 4.0% 27.9% 45.3 ± 1.7% 45.0 ± 1.6% 47.0 ± 2.3%
15.0% 75.1 ± 2.5% 22.5% 81.0 ± 1.0% 80.8 ± 3.2% 81.8 ± 2.6%
15.8% 62.2 ± 2.8% 22.1% 67.1 ± 3.5% 67.0 ± 1.0% 68.2 ± 3.1%
25.0% 83.2 ± 2.1% 17.5% 89.2 ± 1.2% 87.5 ± 2.9% 90.1 ± 1.7%
40.0% 12.5 ± 1.4% 10.0% 12.9 ± 1.4% 14.5 ± 2.4% 14.9 ± 3.1%
Portion Evaluation Remedial SFT Eval. (Remedial) Eval. (Uniform) RL Evaluation
20.0% 50.9 ± 1.2% 20.0% 54.5 ± 2.3% 54.8 ± 1.2% 55.2 ± 2.7%
4.6% 57.9 ± 1.3% 27.7% 64.4 ± 2.0% 63.9 ± 1.3% 65.1 ± 2.3%
25.0% 57.0 ± 2.5% 17.5% 61.2 ± 1.7% 61.5 ± 1.0% 62.6 ± 2.8%
30.4% 85.1 ± 1.8% 14.8% 90.5 ± 1.1% 89.2 ± 0.7% 90.9 ± 2.0%
20.0% 11.8 ± 0.3% 20.0% 13.5 ± 2.1% 14.0 ± 3.0% 14.1 ± 2.4%
Portion Evaluation
25.0% 50.0 ± 1.4%
30.0% 73.1 ± 1.9%
7.7% 45.3 ± 2.2%
5.0% 60.2 ± 2.5%
32.3% 15.6 ± 0.7% Continued on next page
31
Table 13 (continued) Domains Remedial SFT Eval. (Remedial) Eval. (Uniform) RL Evaluation
Ciphers
Operations
Logic
Counterfactual
Puzzles
17.5% 56.3 ± 2.2% 56.9 ± 1.5% 57.0 ± 4.1%
15.0% 78.0 ± 3.6% 77.3 ± 1.3% 79.3 ± 2.9%
26.2% 50.9 ± 1.7% 50.5 ± 1.4% 53.7 ± 2.9%
27.5% 68.1 ± 3.4% 65.9 ± 1.3% 71.0 ± 2.9%
13.8% 16.8 ± 2.3% 17.2 ± 1.0% 17.4 ± 3.7%
† Domains whose mid-training coverage is at or above the 60% threshold receive r = 0: Configs 1 (Cipher 72.3%), 7 (Puzzle 60%), 15 (Cipher
60%), 16 (Operation 60%), and 19 (Logic 60%). Config 1 Cipher’s zero-allocation cell gains only +0.2 pp (§G). Values are mean ± 1 SD across the five seeds.
F.4
Out-of-Sample Predictive Validity of the Fitted Curves
The six held-out allocations (Table 13, held-out blocks) are trained at allocations withheld from the curve fitting and carried through the complete mid-training+SFT+RL pipeline. The split-Gaussian fits generalize to them across all five domains: mean absolute held-out residuals are 1.1 pp for Counterfactual (all six points within ±2 fit-RM SE), 1.0 pp for Puzzle, 1.3 pp for Cipher, and 1.8 pp for Logic (five of six), while Operation is the weakest match (3.0 pp; two low-coverage points realize ≈ +6 pp above the curve), so Operation’s left tail rises somewhat faster than the splitGaussian assumes. Applying the same fixed compensatory budget to the six held-out checkpoints gives gains in all 30 configuration–domain cells (mean +4.5 pp) yet closes 0/60 pairs at the 5 pp threshold and only 5/60 at the 10% ratio (uniform control: 0/60 and 8/60), and compensation’s only material advantage over uniform remains Counterfactual (+6.6 vs. +4.6 pp). The RL stage adds +1.19 pp on average—in line with the +1.1 to +2.3 pp seen in the main experiments—and leaves the picture unchanged: after the complete pipeline the held-out allocations still close 0/60 pairs at the 5 pp threshold and 7/60 at the 10% ratio. The held-out allocations also provide the only direct test of the premise underlying θ∗ : that curves fitted at the mid-training stage carry information about final accuracy. Fitting on the 24 sweep configurations alone and predicting the six withheld allocations, the predicted overall accuracy tracks the realized mid-training-only value almost exactly (Pearson r = +0.954, Spearman ρ = +1.000, mean absolute error 0.64 pp) and—the point that matters for allocation choice—it also predicts the realized post-RL overall accuracy after the complete mid-training+SFT+RL pipeline (r = +0.967, p = 0.002; ρ = +0.886). The mid-training ranking is largely preserved through post-training (r = +0.975 between mid-only and post-RL overall), so a coverage choice made before SFT is still visible after RL. The predictive relationship is not carried by any single held-out point: dropping each of the six in turn moves the mid-training-only correlation only within [+0.941, +0.973] and the post-RL correlation within [+0.954, +0.987], and a bootstrap over the six allocations gives 95% intervals of [+0.923, +0.999] and [+0.926, +0.998] respectively. For completeness, θ∗ reaches 62.72 overall on the compensatory-SFT scale while the six held-out allocations span 54.02–59.52, but this is not an out-of-sample validation of θ∗ : θ∗ was chosen as the argmax of the fitted objective and is evaluated here under a different downstream recipe (rule-type-balanced SFT in Table 1 versus the compensatory pass applied to the held-out checkpoints), so the comparison records the internal consistency of an argmax across two SFT recipes rather than evidence that a fit-selected allocation generalises. One component does not survive this test: the relative-gain objective of Eq. 4, which is what actually selects θ∗ , ranks the six held-out allocations only weakly (r = +0.441, p = 0.38; ρ = +0.257), and unlike the predicted-accuracy correlations it is fragile: leaving out one allocation moves it anywhere in [+0.058, +0.750] and its bootstrap interval [−0.875, +0.976] spans zero. The baseline normalisation bd appears to be what degrades it, so the fitted curves are the trustworthy object here while the bd -normalised objective layered on top of them is not; an allocation rule should optimise predicted accuracy directly rather than the relative-gain objective. What remains outside this test is θ∗ ’s own full-pipeline row on the matched protocol: its +4.36 pp figure in Table 1 follows the standard rule-type-balanced SFT, whereas the held-out allocations were carried through the compensatory pass, so the two are not run on identical downstream recipes.
32
G
Compensatory SFT: Domain-Augmented Alignment Does Not Close MidTraining Gaps
G.1
Experimental Design
To test whether compensatory SFT can close inter-domain gaps, we construct a domain-weighted SFT dataset: rd =
max(0, 60 − Md ) × 100%, 5 X max(0, 60 − Mj )
(5)
j=1
where Md is the mid-training token percentage for domain d (read directly from each coverage configuration) and rd is the target SFT proportion for domain d. Domains whose mid-training coverage is at or above 60% receive no compensatory SFT allocation (rd = 0); all other domains receive SFT budget in proportion to their shortfall from the 60% threshold, with the allocation vector normalised to sum to 100%. The threshold 60% was chosen because any single domain exceeding 60% of the five-domain token budget already constitutes extreme over-allocation; every domain below this threshold is treated as under-served and eligible for remediation. Applying Eq. 5 to all 24 configurations, five configurations have a domain at or above the 60% threshold (Configs 1, 7, 15, 16, and 19; see Table 13); in all other configurations, all five domains fall below 60% and therefore receive positive compensatory SFT. Total SFT token count is held equal to the uniform baseline so that only the cross-domain distribution changes, not the overall training volume. The uniform-SFT control uses the identical rule-type-balanced data mix for every configuration (the same mix as the shared SFT baselines in Table 1) with the same total token budget and the same 3-epoch schedule as the compensatory pass, so the two passes differ only in the cross-domain allocation of the SFT data.
G.2
Capability Gap Metric: Pairwise Closure Framework
We evaluate whether compensatory SFT closes inter-domain gaps by comparing every unordered pair of domains within each of the 24 coverage-sweep checkpoints, giving 24 × 52 = 240 pairs. The pair members are ordered by their post-compensatory accuracies (see below), so the gap is always reported as the stronger-minus-weaker domain after compensation. Because pairs share both configurations and domains, all pairwise closure summaries are descriptive and require configuration-clustered inference for population-level claims. For each pair (i, j) with mid-training accuracies amid , amid and post-compensatory-SFT accuracies acomp , acomp , sorted such that acomp ≤ acomp , we i j i j i j compute two gap measures: Gdiff ij (a) = aj − ai , ai Gratio (a) = 1 − . ij aj
(6) (7)
The closure rate for gap measure G is Cij =
Gij (amid ) − Gij (acomp ) . Gij (amid )
(8)
A pair is counted as bridged (meaningfully repaired) if the gap narrows by at least a metric-specific margin: an absolute reduction of at least 5 pp in the difference gap, or a relative closure rate Cij of at least 10% in the ratio gap. A stricter 18% ratio threshold is examined only as a sensitivity check. These thresholds are calibrated to the scale of the intervention and of the seed noise in this sweep rather than to a statistical null (§G), and we interpret the resulting bridged-pair counts descriptively. Equal absolute gains naturally produce larger relative gains for weaker domains. Under the difference metric (5 pp), 0/240 pairs are bridged; under the ratio metric (10%), 30/240 pairs are bridged. These counts are descriptive because the 240 pairs share all training configurations and are not independent Bernoulli trials; the thresholds were selected for exploratory sensitivity analysis rather than preregistered confirmatory testing. We therefore do not interpret the associated nominal binomial p-values as population-level evidence. The results indicate that this fixed compensatory policy produces absolute gains and limited, heterogeneous gap changes, but they do not establish that stronger alignment interventions would fail. The same pattern holds on the six held-out allocations withheld from the fit (6 × 52 = 60 pairs): 0/60 bridged at 5 pp and 5/60 at 10% (uniform 0/60 and 8/60), so the non-closure generalizes beyond the 24-configuration sample. 33
Threshold sensitivity. Figure 8 sweeps both thresholds over the 240 pairs: the bridged count falls from 57/240 at the most permissive ratio threshold (5%) and 39/240 at the most permissive difference threshold (0.5 pp) to 0/240 from 4.5 pp upward. To calibrate these counts we use a permutation null rather than a “half of all pairs” reference line, which is not a null expectation: under a true no-effect null the expected number of bridged pairs is near zero, not half. The null permutes, within each configuration, which domain receives which compensatory gain, holding the observed gain magnitudes—and hence the arithmetic budget cap—exactly fixed while destroying the alignment between a domain’s gain and its position within each pair (2,000 reallocations). A random reallocation of the very same gains would bridge 13.8 ± 3.3 pairs under the 5 pp difference metric and 77.9 ± 8.5 under the 10% ratio metric; the compensatory policy bridges 0 and 30 respectively (P < 0.001 in both cases, one-sided). The observed non-closure is therefore not a mechanical consequence of the capped budget: the same gains, reallocated at random, would have closed substantially more gaps than the compensatory policy does. (a) Ratio-gap threshold sensitivity
(b) Diff-gap threshold sensitivity Permutation null (95% band)
Permutation null (mean) Paper threshold (10\%) Sensitivity check (18\%)
150
100
50
Pairs bridged (out of 240)
Pairs bridged (out of 240)
Permutation null (95% band)
200
0
Permutation null (mean)
200
Paper threshold (5 pp)
150
100
50
0 5
10
15
18
20
25
1
Ratio-gap closure threshold (%)
2
3
4
5
6
7
8
9
10
Absolute-difference closure threshold (pp)
Figure 8: Threshold sensitivity for pairwise gap-closure over the 240 domain pairs. Left: bridged pairs vs. ratio threshold (5%–28%). Right: same for difference threshold (0.5–10 pp). The dashed line and shaded band are the mean and 95% interval of a permutation null that reallocates the observed compensatory gains across domains within each configuration (2,000 draws), holding the gain magnitudes and the arithmetic budget cap fixed. The observed counts lie below this null at both metrics (P < 0.001), so the non-closure is not a mechanical consequence of the capped budget.
G.3
A Stronger Reweighting Arm: Sharpening the Compensatory Policy
Observation 1 is stated for a policy family whose within-configuration gain differential is arithmetically bounded at ≈ 7 pp, so a reader may reasonably ask whether the non-closure is an artifact of a policy too weak to succeed. Two analyses in this appendix already bear on that question. First, the feasible set is large: evaluated on the fitted gain curves, 192 of the 240 domain pairs (80%) could in principle be narrowed by at least 5 pp by some budget-feasible allocation of the same SFT tokens (216/240, 90%, under the 10% ratio metric), against 0 and 30 actually closed. Second, the permutation null of Appendix G.2 shows the observed gains close fewer gaps than a random reallocation of the same magnitudes. Neither, however, answers the operational question: is there a stronger policy in the same family that actually closes the gaps? This subsection reports the experiment that does. Policy family.
We generalise the compensatory rule of Eq. 5 with a single sharpening exponent γ: [max(0, T − Md )]γ × 100%, rd (γ) = P5 γ j=1 [max(0, T − Mj )]
T = 60%,
(9)
where Md is domain d’s mid-training token percentage. The exponent controls how sharply the fixed SFT budget concentrates on coverage-deficient domains. At γ = 1 the rule reduces exactly to Eq. 5, so the arm reported in the main text is the γ = 1 member of this family and the comparison is nested rather than across unrelated policies; 34
as γ → ∞ it becomes winner-take-all, assigning the entire remedial budget to the single most deficient domain. Every branch holds the total SFT token count equal to the uniform baseline, exactly as at γ = 1; only the crossdomain distribution changes. We ran γ ∈ {1, 2, 4, ∞} on six sweep configurations that between them carry most of the feasible-but-unclosed pairs—each has 9–10 of its 10 domain pairs separated by ≥ 5 pp at the mid-training-only checkpoint—which required 90 additional SFT passes (three new branches × six configurations × five seeds; γ = 1 was already trained) and no further mid-training or RL. Result: closure is attainable, but it is bought with average accuracy. Table 14 reports the outcome over the 6 × 52 = 60 pairs these configurations contribute. Sharpening does close gaps that γ = 1 leaves open—the count rises monotonically from 0/60 at γ = 1 to 12/60 at γ = ∞ under the 5 pp metric—so the non-closure reported in the main text is not an inescapable property of the budget. But closure is purchased directly out of average accuracy: the mean per-cell gain falls monotonically from +4.34 pp to +2.26 pp, so the winner-take-all branch that closes the most pairs also forfeits roughly half of the improvement the compensatory pass was introduced to deliver. The two objectives are in tension across the whole family we tested, and no branch achieves both: the branch that maximises mean gain closes nothing, and the branch that closes most gives up 2.1 pp of mean gain. This sharpens Observation 1 rather than overturning it—the gaps survive every setting that preserves the average gain—and it replaces an untested scope caveat with a measured trade-off. Even at γ = ∞, 48 of the 60 pairs remain unclosed at the 5 pp threshold. Table 14: Sharpened compensatory SFT over the six configurations carrying the most feasible-but-unclosed pairs (60 pairs). γ = 1 is the policy reported in the main text. Mean gain is the per-cell accuracy change over the midtraining-only checkpoint, averaged over all 30 configuration–domain cells. Closure counts use the same two metrics as Appendix G.2. Branch γ = 1 (Eq. 5) γ=2 γ=4 γ = ∞ (winner-take-all)
G.4
Mean gain (pp)
Bridged, 5 pp
Bridged, 10% ratio
+4.34 +4.17 +3.65 +2.26
0/60 1/60 5/60 12/60
8/60 11/60 13/60 12/60
Results
Table 13 (Remedial SFT rows) reports compensatory allocations from Eq. 5; Eval. (Remedial) rows give postcompensatory-SFT accuracy. Compensatory SFT raises scores in 116 of 120 configuration–domain cells (mean +4.32 pp; the four cells with no positive gain all receive zero compensatory allocation r = 0). The re-run exploratory θ∗ block, which is not one of the 24 sweep checkpoints, also gains in all five domains under compensation (+1.4 to +6.2 pp; mean +4.5 pp), with Puzzle’s gain the smallest; its deviations from the fitted per-domain curves below are at most ≈ 0.3 pp, so including or excluding this block does not change any conclusion. Threshold calibration. The 10% ratio and 5 pp difference cutoffs are calibrated to the scale of the intervention and of the seed noise in this sweep, not to a statistical null. The 5 pp difference threshold exceeds the largest cross-seed SD observed on any single configuration cell in Table 13 (mid-training-only SDs run 0.7–3.9 pp over the 24 sweep configurations; the SFT-pass rows on those configurations have smaller SDs, though the six held-out blocks (which add the RL stage) carry a few larger values, up to ≈4.5 pp), so a counted closure is larger than the seed-level dispersion of either endpoint. The 10% ratio threshold is likewise beyond the seed-noise scale of the weaker-domain accuracies that dominate ratio-bridged pairs, and equal absolute gains naturally produce larger relative closures for weaker domains. The scan in Figure 8 shows that the conclusion does not depend on the exact cutoffs: the bridged-pair count stays below the permutation null of §G.2 at every tested threshold (57/240 even at the most permissive 5% ratio; 39/240 at 0.5 pp difference, and 0/240 from 4.5 pp upward). Under ratio (10%), 30/240 pairs are bridged; under difference (5 pp), 0/240 pairs are bridged. Both summaries indicate heterogeneous changes under this fixed policy, not general non-repairability. These statistics underpin Figure 4. Redirecting SFT budget toward coverage-deficient domains raises accuracy but does not reliably close pairwise gaps created by mid-training.
35
Uniform-SFT control. As a control with the same data mix and budget, a uniform SFT pass (identical rule-typebalanced data mix across configurations; the same total token budget and 3-epoch schedule as the compensatory pass; Table 4) was applied to all 24 sweep checkpoints (Eval. (Uniform) rows, Table 13). Table 15 compares the two passes. The uniform pass raises every one of the 120 cells (mean +4.20 pp vs. +4.32 pp compensatory), bridges the same 0/240 pairs under the 5 pp difference metric and 32/240 under the 10% ratio metric (vs. 30/240), and produces per-domain mean gains close to the compensatory ones except for Counterfactual, where compensation is stronger by +1.9 pp (+6.37 vs. +4.49 pp), and Puzzle, where uniform SFT is slightly stronger (+2.10 vs. +1.48 pp). Cross-configuration range narrowing is likewise comparable; uniform SFT narrows all five ranges (including Puzzle, −0.7 pp), whereas compensation leaves Puzzle’s range slightly wider (+0.4 pp). The compensation formula therefore does not systematically outperform an equal-budget uniform re-training pass; its only material advantage is Counterfactual, the domain whose gain responds most steeply to the remedial ratio (below). Table 15: Per-domain mean accuracy gains (pp) of the compensatory and uniform SFT passes over the 24 sweep configurations (all 120 configuration–domain cells in the last row). Compensation exceeds the uniform control materially only for Counterfactual; it is slightly below the control for Cipher and Puzzle. Domain
Compensatory gain
Uniform gain
Difference (comp.−uniform)
Cipher Operation Logic Counterfactual Puzzle
+3.42 +5.73 +4.62 +6.37 +1.48
+3.89 +5.77 +4.76 +4.49 +2.10
−0.47 −0.04 −0.14 +1.88 −0.63
All 120 cells
+4.32
+4.20
+0.12
Gain vs. the remedial allocation: per-domain fits. separately for each domain, the model
Over the 115 configuration–domain cells with r > 0, we fit,
∆d (r) ≈ ad log(1 + bd r) + cd ,
(10)
where r is the domain’s remedial SFT proportion in percent (Figure 7). The additive +1 inside the logarithm removes the absorption redundancy of the previous pooled form ∆ ≈ a log(b r) + c, in which b and c were not separately identifiable. Table 16 reports the fitted parameters. Within each domain the gain rises with r and saturates at a domain-specific ceiling; the fits are tight (MAE 0.07–0.25 pp; R2 0.71–0.97 over each domain’s r-window), with the steepest rise across the observed window for Counterfactual (predicted +3.2 pp at r=5% to +8.2 pp at r=29.5%) and the flattest for Puzzle (+0.3 to +2.4 pp over r ∈ [4, 29]%). Domains differ more in overall gain level than in their response to r: a pooled fit of Eq. 10 over all 115 cells explains only ≈ 31% of the variance (R2 =0.31, MAE 1.36 pp; Table 16), which is why the earlier pooled model appeared to describe a weak association. These fits are descriptive summaries; within each domain the observed r-window is narrow (≈4–30%), so ad and bd trade off against each other and, where the window is nearly linear (Cipher), the pair is only weakly identified individually even though the fitted curve is tight. Table 16: Per-domain fits of the compensatory gain ∆ (pp) against the remedial ratio r (%) using ∆ = a log(1+b r)+c on the r > 0 cells of Table 13 (Eq. 10; Figure 7). R2 and MAE are computed on the fitted cells. The pooled row fits all 115 r > 0 cells jointly. † Cipher’s a sits at its upper fit bound; over Cipher’s observed window log(1 + br) ≈ br, so the curve is effectively linear (slope ≈ +0.12 pp per percentage point of r) and a, b are individually unstable. Domain
n
a
b
c
R2
MAE (pp)
Cipher† Operation Logic Counterfactual Puzzle
22 23 23 24 23
50.00 2.78 3.81 11.99 4.39
0.0024 3.2420 0.1351 0.0237 0.0273
+1.24 −6.06 −0.39 +1.83 −0.16
0.78 0.71 0.77 0.95 0.97
0.25 0.25 0.25 0.19 0.07
Pooled
115
7.79
0.0360
+0.20
0.31
1.36
36
Boundary cases (zero-allocation domains). Five cells receive r = 0 because their domain’s mid-training coverage is at or above the 60% threshold (Configs 1, 7, 15, 16, and 19; Table 13). These cells show no meaningful gain under compensation (all five gains within ±0.3 pp, the largest being Config 1 Cipher at +0.2 pp, which receives no remedial allocation at all), and the conclusions above hold when these configurations are excluded. Budget sensitivity. Under our budget-proportional model, the within-configuration gain differential reaches at most ≈ 7 pp; scaling total SFT volume would not increase this differential without changing the allocation formula. More aggressive allocations could produce larger differentials but remain untested. We scope our conclusion to this fixedbudget, fixed-formula setting.
H
External Benchmark Detailed Results
H.1
Coverage Sweep: External Benchmark Accuracy
Table 17 reports external benchmark accuracy for the coverage-sweep mid-training models (plus the FineWeb-Edu baseline), using the same model checkpoints as the KOR-Bench mid-training-only evaluation. The external results are limited consistency checks rather than independent confirmation of a general transfer law. The external configuration set is fixed: the three external benchmarks evaluated on the nine coverage configurations (plus the FineWeb-Edu baseline) listed in Table 17, which span the sweep from severely imbalanced to balanced allocations; Spearman rank associations of the KOR average with each external column over these configurations are reported in §4.4 (n = 9). The external corpus draws 34.8% of its tokens from the ProofWriter rule family, the ZebraLogic family contributes 201 mid-training samples, and the reported isolation does not establish template- or rule-family-level disjointness. The associations may therefore reflect shared structure or exposure as well as mixture allocation. The per-depth, perhouse, and per-type comparisons below contrast the shared SFT+RL baseline with the Expt. 1 checkpoint (Cipher 72.3%, Operation 0%, Logic 12.0%, Counterfactual 0%, Puzzle 15.7%; Table 1); the covered-depth ablation uses its own D3-/D5-heavy mixtures (below) and is not tied to any sweep configuration. Table 17: External benchmark accuracy (%) for the coverage-sweep mid-training models (the FineWeb-Edu baseline is shown for reference). Coverage: Cipher / Operation / Logic / Counterfactual / Puzzle (internal); the ProofWriter rule family (34.8% of external tokens) is fixed. KOR avg is the macro-average of the five KOR-Bench domain accuracies from Table 13 (§F); the external columns are five-seed means on the same checkpoints, matching the protocol of Table 13. “FineWeb-Edu” = FineWeb-Edu-only baseline without KOR-Bench data. The nine coverage configurations listed here constitute the fixed external configuration set (§4.4). Coverage (Cipher/Oper./Logic/Counterf./Puzzle) FineWeb-Edu (100%) 40 / 4 / 8 / 40 / 8 72.3 / 0 / 12 / 0 / 15.7 0 / 10 / 40 / 10 / 40 50 / 40 / 0 / 5 / 5 5 / 15 / 5 / 15 / 60 26 / 16 / 17 / 30 / 11 30 / 18 / 16 / 18 / 18 10 / 8 / 37 / 8 / 37 Balanced (20%×5)
H.2
KOR avg
ProofWriter
ZebraLogic MC
ZebraLogic Grid
CounterBench
22.5 47.5 36.6 44.8 40.1 49.8 55.9 56.0 50.0 56.3
50.9 80.7 83.1 85.1 81.1 81.2 80.7 84.4 85.7 86.3
37.7 45.8 40.8 31.6 41.4 37.4 46.8 40.6 30.6 38.2
13.1 21.0 23.0 37.8 25.3 26.2 26.9 37.2 35.0 35.2
65.0 70.7 69.9 76.4 72.0 72.5 71.0 77.7 75.9 79.5
ProofWriter Per-Depth
Ablating the covered depth. Depth-5 ProofWriter exposure improves both the covered D5 and held-out D6 strata; D7–D9 strata are too small for strong claims. Tables 19 and 20 ablate whether transfer depends on covered training depth (reference: SFT-only baseline, Overall 76.90). The two mixtures swap D3 and D5 proportions while holding D0–D2 and Other fixed. The D5-heavy model after SFT is substantially stronger (+5.60 pp overall; +5.70 pp D5, +4.64 pp D6), suggesting post-training benefits more when mid-training difficulty is closer to the evaluation depth.
37
Table 18: ProofWriter per-depth accuracy (%). Mid-training trains only to depth 5; eval depths 5–9. † N < 200, not significant at 95% level. ∆ is computed from exact correct counts before rounding individual accuracies; apparent differences between ∆ and (Mid-training+SFT−SFT+RL) on rounded values are rounding artifacts. Depth
SFT+RL
Mid-training+SFT
∆
95% CI (∆)
N
D5 D6 D7 D8 D9
80.1 76.2 75.8 77.1 75.0
84.4 82.7 82.5 83.3 91.7
+4.4 +6.4 +6.7 +6.2 +16.7
[+3.6, +5.2] [+3.3, +9.5] [−3.5, +17.0]† [−5.0, +17.4]† [−3.8, +37.2]†
18,498 1,292 120 96 24
Mixture
D0–D2
D3
D5
Other
D3-heavy D5-heavy
2.04 2.04
65.74 17.12
17.12 65.74
15.10 15.10
Table 19: ProofWriter training-depth distribution in the covered-depth ablation (%). Other denotes examples without parsed depth. Setting
Overall
D5
D6
D3-heavy D3-heavy + SFT D5-heavy + SFT
−2.69 +0.89 +5.60
−2.51 +0.97 +5.70
−4.18 −0.24 +4.64
Table 20: ProofWriter covered-depth ablation results. Values are accuracy deltas (pp) relative to the ablation suite’s SFT-only baseline.
H.3
ZebraLogic MC Per-House
Table 21 reports ZebraLogic MC results (MC = multiple-choice), complementing the ProofWriter depth-extrapolation result in §4.4. With only 201 direct mid-training ZebraLogic samples (0.3% coverage), Mid-training+SFT outperforms SFT+RL at every reported house count. This is consistent with shared formal-deduction and constraint structure, but ZebraLogic is a same-family held-in exposure rather than an out-of-distribution transfer test, and it is not evidence that a small exposure alone caused the gains. Template/family-level overlap remains uncontrolled in this study, and a zero-exposure replication is future work. Table 21: ZebraLogic MC per-house accuracy (%). Mid-training coverage: 0.3% (201 samples). 95% asymptotic CIs are shown; these are exposure-uncontrolled comparisons and are reported descriptively. Houses 2 3 4 5 6
H.4
SFT+RL
Mid-training only
Mid-training+SFT
∆
95% CI (∆)
89.9 78.8 60.2 47.6 31.3
90.1 77.4 56.5 40.5 28.6
94.1 84.3 72.6 55.9 43.1
+4.2 +5.5 +12.4 +8.3 +11.8
[+1.1, +7.3] [+1.2, +9.8] [+7.5, +17.3] [+3.0, +13.6] [+6.6, +17.0]
CounterBench Per-Type
The per-type ∆ values are all at or below zero—a contrast with the positive deltas reported for ProofWriter and ZebraLogic. This pattern is consistent with, but does not prove, a no-transfer interpretation: CounterBench may differ in required procedures, prompt format, difficulty, and base-model prior. It should therefore not be treated as a definitive negative control.
38
Table 22: CounterBench per-type accuracy (%). Zero direct mid-training coverage; non-positive ∆ is consistent with a no-transfer interpretation, but is not a definitive negative control. Type Basic Conditional Joint Nested
I
SFT+RL
Mid-only
Mid+SFT
∆
90.8 80.4 81.2 76.8
80.8 72.0 68.8 64.8
86.8 80.4 76.8 72.8
−4.0 0.0 −4.4 −4.0
Limitations and Ethical Considerations
Scale. The primary experiments use Qwen3-8B-Base, with a mid-training-only replication at Qwen3-4B-Base (§C); sensitivity to other scales and architectures remains open (see Appendix B.2 for the scale-selection rationale). Whether the qualitative patterns persist at other model scales is unknown; fitted 95%-of-peak intervals may shift with model size. Scope and inference. Claims are scoped to KOR-Bench and three external benchmarks, the Qwen3-Base family (8B primary, with a mid-training-only 4B replication), the reported data mixture, and the tested finite-budget recipe. The central inferential limit is the absence of held-out validation for the fitted peaks/intervals, θ∗ , and the full-pipeline gain: the 24-point curve fits, the fitted 95%-of-peak intervals, the compensation thresholds, and θ∗ all reuse the same sweep for model selection. Six held-out allocations carried through the complete mid-training+SFT+RL pipeline (Table 13) now probe the mid-training-only curves and the compensatory-SFT result, which they roughly track at seed-noise level for Cipher, Logic, Counterfactual, and Puzzle (Operation’s held-out low-coverage points sit ≈ 6 pp above the fitted tail; 0/60 pairs closed at 5 pp), and carrying them through the RL stage leaves that unchanged (0/60 at 5 pp after the complete pipeline), so the RL leg is covered out-of-sample as well. What remains unvalidated on held-out allocations is θ∗ itself, together with the fitted peaks, the 95%-of-peak intervals, and the compensation thresholds, all of which are read off the same sweep that produced them. The simplex design additionally makes the fitted curves mixturelevel marginal associations rather than isolated per-domain causal effects, and we treat this confound explicitly rather than implicitly (§5.2). We report the compositional analysis this calls for (§F.2): the isometric log-ratio surface fits well but is a saddle in every domain, so a jointly optimal mixture is not identified and the moderation result stands as a per-domain marginal statement. We tested whether that negative finding was a sampling artifact by adding 12 interior allocations to the pool (Appendix B.2): a joint maximum still does not emerge and every domain’s stationary point remains a saddle, so the moderation result is robustly a per-domain marginal statement and is not promoted to a single recommended mixture. Reported sweep values are five-seed means with cross-seed ±1 SD (Table 13); the full-pipeline rows are likewise five-seed means, and their cross-seed SDs are displayed in Table 1. Externalbenchmark rows are five-seed means, and bootstrap CIs where used capture evaluation-sample variance only and do not replace the cross-seed uncertainty. We therefore claim no per-domain comparison as individually significant and report all per-domain deltas descriptively. The only row-level inferential statistics are the Welch two-sample t-tests on the overall full-pipeline scores (§4.2), which are at best nominally significant and do not survive multiplicity control; whether seeds are matched across pipeline rows is not established in this manuscript, so paired seed-level tests are left to the released per-seed data (Appendix C.7). GSPO uses a fixed 200-step budget; because binary verifier rewards provide no dense per-token reward signal, longer schedules were unstable across seeds (Appendix C.6), and longer or differently designed RL may yield different outcomes. Compensatory SFT conclusions are scoped to the fixedbudget, fixed-formula setting (Appendix G); stronger token-level reweighting, curricula, adaptive sampling, and stable RL after compensatory SFT remain untested. The current report also does not fully separate coverage from uniqueexample diversity, repetition, sequence length, or external-corpus exposure; family/template-level decontamination and seed-level analyses of the full-pipeline rows are left for future work. Ethical considerations. This work studies data-mixture design for mid-training. It does not involve human subjects or personally identifiable information. All corpora are either generated by deterministic symbolic solvers or sourced from publicly released benchmarks. Trained models are research checkpoints not intended for deployment. Benchmark balance should not be equated with user benefit, safety, calibration, or fairness in a deployed system; in high-stakes domains, low coverage can create uneven risk and would require domain-specific assurance before
39
deployment.
Declaration of AI Assistance We utilized ChatGPT and Claude for grammatical checking and LATEX support of the content presented in this study but did not use them for the initial draft of this study. Cursor, Codex, and Claude Code were utilized for trivial and boilerplate code completion during data analysis. We declare that all content presented and code utilized in this study has been reviewed and edited by the authors.
40