Good Pretraining, Bad SFT: Checkpoint Quality Across the Training Stack
Jonas Knupp Aleph Alpha
Pit Neitemeier
Sascha Wirges
(c) Aggregate scores
(a) GSM8K
40 30
100
20
Score / % of best
Empirical frequency / %
arXiv:2609.08966v1 [cs.AI] 8 Sep 2026
Sohir Maskey∗ Philipp Scholl
10 0
(b) MBPP
40 30
90
80
20 10
70
0 60
80
100
Pre
Mid
Score retained after perturbation / % C ONSTANT
Long
SFT
Training stage C OOLDOWN
M ERGE
Figure 1: Perturbation sensitivity and downstream performance of three MoE variants after completing different training stages. (a–b) Distribution of scores after 100 Gaussian weight perturbations (σϵ = 0.005), relative to each checkpoint’s unperturbed score; (c) stage aggregates as a percentage of the best checkpoint at that stage. C ONSTANT and M ERGE show greater solution density, i.e., perturbing the pretrained checkpoints retains higher scores at performance-preserving thresholds compared to C OOLDOWN. This observation is consistent with their stronger post-SFT performance.
Abstract Language-model checkpoints are commonly selected by pretraining loss or benchmark scores, assuming that the highest-scoring checkpoint will remain the best starting point for subsequent training. We show that this assumption can fail in a full 30B mixture-of-experts training pipeline. The checkpoints that perform better after the full downstream training stack also have higher solution density, i.e., retain downstream performance under local weight perturbations.
1
Introduction
Language-model checkpoints are usually compared by asking which checkpoint is better now: which has lower loss or higher benchmark scores. This is sufficient when training ends at that checkpoint. Modern LLMs, however, are trained in a sequence of stages including pretraining, mid-training, long-context adaptation, supervised fine-tuning (SFT), and further post-training (Ouyang et al., 2022; S. Chen et al., 2023; Olmo et al., 2026). In this setting, the question is different: Which checkpoint selection criteria will produce the best final model? This matters because executing all training stages for every candidate is expensive and explodes combinations. Using intermediate evaluations as selection criteria implicitly assumes that checkpoint ∗ Correspondence to [email protected].
Preprint.
Table 1: A better pretraining checkpoint can produce a worse final model. C ONSTANT and C OOLDOWN share the first 6.7T tokens and the same 7.5T-token budget. All three sources then receive the same downstream recipe. Selection signal
C ONSTANT
C OOLDOWN
M ERGE
Train loss ↓ Validation loss ↓ Pretraining aggregate ↑
1.717 1.734 0.415
1.627 1.648 0.440
– – 0.460
Post-SFT aggregate ↑
0.360
0.247
0.363
rankings are preserved by later training. Our experiments show that this assumption can fail: subsequent training can reverse the ranking of intermediate checkpoints. Table 1 compares three source checkpoints from one 30B-parameter mixture-of-experts (MoE) pretraining run. C ONSTANT and C OOLDOWN share the first 6.7T of 7.5T training tokens and differ only in the learning-rate schedule over the final 800B tokens. M ERGE is a weighted average of C ONSTANT checkpoints (Tian et al., 2026). C OOLDOWN improves every available pretraining signal relative to C ONSTANT, yet after identical mid-training, long-context adaptation, and SFT, their ordering reverses. This motivates our central question: when do intermediate checkpoint rankings become predictive of the final ranking? Across our training trajectories, neither the mid-training nor the long-context aggregate separates C ONSTANT from C OOLDOWN, although rankings within a learning-rate sweep stabilize after long-context adaptation (Figure 2). Gan et al. (2026) define solution density as the fraction of Gaussian perturbations that retain task performance above a threshold. They show that this fraction increases with model scale, arguing that denser neighborhoods of task-specific experts make useful specialists easier to reach through post-training. Figure 1 shows a similar distinction across checkpoints of the same model: C ONSTANT and M ERGE have greater solution density than C OOLDOWN. M ERGE is particularly informative: it has the highest pretraining aggregate, remains more robust to perturbations than C OOLDOWN, and finishes effectively tied with C ONSTANT after SFT. This suggests that the unperturbed score alone misses how broadly performance persists under nearby weight changes. We therefore hypothesize that solution density varies not only with scale but also across training trajectories, and that these differences help explain downstream adaptability. Contributions. We show that conventional pretraining metrics can misrank checkpoints for a fixed downstream training pipeline, and show that intermediate aggregates do not anticipate this reversal. We further connect this reversal to solution density and show that C OOLDOWN’s largest downstream failure reflects unstable response completion rather than absent latent code capability.
2
Experimental Design
We study a 30B-parameter MoE with 3B active parameters. Training consists of 7.5T-token pretraining, 100B tokens of capability-focused mid-training at 8k context, 100B tokens of long-context adaptation at 64k, and 10B tokens of conversational SFT at 64k. The optimizer state is reset with repeated LR warmup at each stage. For details on architecture, training, and datasets see Section B. Pretraining sources. C ONSTANT and C OOLDOWN share their first 6.7T tokens and total token budget. We decay the learning rate for the C OOLDOWN run to 10% of the maximum during the final 800B tokens of training. M ERGE (Tian et al., 2026) adds no training: it combines 20 equally spaced C ONSTANT checkpoints over a 600B-token trailing window via2 20 X i θmerge = θi , 210 i=1 2 Checkpoint merging is well established and has been ablated in LLM training, for example in Nemotron Super (NVIDIA et al., 2026). More directly, Y. Li et al. (2025) show that averaging constant-LR checkpoints can match annealed pretraining performance, while Tian et al. (2026) connect checkpoint weighting to effective update decay. See Section B.5 for details.
2
where the checkpoints are ordered from oldest to newest. Learning-rate sweep. Starting from C OOLDOWN, we train all nine combinations ηmid , ηlong ∈ 1, 13 , 19 . where each factor scales a stage’s peak learning rate relative to the preceding stage. For example, ηlong = 1/3 sets t4he long-context peak learning rate to one third of the mid-training peak. Mid-training uses either no decay when ηlong = 1 or cosine decay to the long-context peak, yielding a continuous schedule across stages. 3 Long-context training then decays to one third of its peak learning rate. Stage-wise evaluation aggregates are reported in Table 5. The (1, 1) mid/long schedule achieved the highest post-SFT aggregate among the nine schedules, each followed by SFT factor 1/3. Due to compute constraints, we therefore reuse this schedule for C ONSTANT and M ERGE. 4 Evaluation. At the pretraining, mid-training, and long-context boundaries, we use completionstyle benchmarks spanning knowledge, mathematics, and code. From mid-training onward, we add long-context retrieval. These three aggregates are cluster-balanced: Pre averages three cluster means, whereas Mid and Long average four cluster means. Post-SFT, we evaluate on six chat-formatted benchmarks and take their unweighted mean. The stage-specific suites are given in Section C.
3
Checkpoint Rankings Change Across the Training Stack
3.1
Pretraining scores do not determine post-SFT scores
Table 1 shows that C OOLDOWN improves every available pretraining signal over C ONSTANT: train loss (1.717 to 1.627), validation loss (1.734 to 1.648), and pretraining aggregate (0.415 to 0.440). Under conventional checkpoint selection, C OOLDOWN would therefore dominate C ONSTANT. After the identical mid-training, long-context, and SFT continuation, C ONSTANT instead leads C OOLDOWN by 0.113. M ERGE has the highest aggregate before and after the downstream pipeline, although its final 0.003 advantage over C ONSTANT is small. Table 2: Post-SFT benchmark scores after the fixed downstream recipe; the last column excludes HumanEval+. Source
AIME DAPO Skywork IFBench
C ONSTANT C OOLDOWN M ERGE
0.221 0.171 0.217
0.430 0.450 0.495
0.390 0.355 0.345
HE+ GPQA Mean w/o HE+
0.238 0.646 0.250 0.049 0.236 0.665
0.237 0.207 0.222
0.360 0.247 0.363
0.303 0.287 0.303
The complete six-task profiles in Table 2 show that C OOLDOWN is not uniformly worse. The large aggregate gap is amplified by a severe HumanEval+ failure, which we examine separately below; without HumanEval+ it shrinks to 0.016 (Table 2). 3.2
Mid-training rankings remain unstable
Across the eleven training trajectories, the mid-training aggregate correlates only weakly with the post-SFT aggregate (Pearson r = 0.473, p = 0.142; Spearman ρ = 0.482, p = 0.133). Figure 1c and Figure 2 show that rankings can still change substantially after this stage. 3.3
Long-context rankings stabilize within a sweep, not across sources
After long-context adaptation, the association changes substantially. Pearson correlation with the postSFT aggregate rises to r = 0.884 (p < 0.001) and Spearman correlation to ρ = 0.964 (p < 0.001). 3 Following Olmo et al. (2026), optimizer state is reset and each stage begins with warmup, so continuity holds up to the warmup. 4 This comparison favors C OOLDOWN : its schedule is selected from 9 mid/long configurations, or 27 including the SFT sweep in Section E.1, whereas C ONSTANT and M ERGE are each run once. Tuning them could only improve their scores. Thus the reversal’s direction in Section 3.1 is unaffected, though its magnitude and the C ONSTANT–M ERGE ordering may change. Even the best of 27 C OOLDOWN runs (0.294) trails untuned C ONSTANT and M ERGE (0.360, 0.363).
3
Post-SFT aggregate
(a) After mid-training
(b) After long-context adaptation
r = 0.473, ρ = 0.482
r = 0.884, ρ = 0.964
0.3 0.2 0.1 0.51
0.53
0.55
0.57
0.58
0.6
0.62
0.64
Long-context aggregate
Mid-training aggregate C ONSTANT
C OOLDOWN
M ERGE
Figure 2: Intermediate versus post-SFT aggregate scores. Dashed lines are ordinary least-squares fits. Figure 2 compares the post-SFT aggregate against two intermediate aggregates. The diffuse midtraining relationship tightens into an almost monotone ordering after long-context adaptation. This result is not driven by adding C ONSTANT and M ERGE: restricting the calculation to the nine C OOLDOWN trajectories gives r = 0.448 (p = 0.226), ρ = 0.450 (p = 0.224) after mid-training and r = 0.935, ρ = 0.950 (both p < 0.001) after long-context adaptation. Notably, long-context scores still fail to predict the ordering of C ONSTANT, C OOLDOWN, and M ERGE at matched schedules: C OOLDOWN (1, 1) scores 0.637 against 0.636 for C ONSTANT, yet trails by 0.113 after SFT. We do not interpret this as evidence that long-context adaptation causally creates predictiveness: later checkpoints are also closer to the final model.
4
What Checkpoint Scores Miss
The ranking reversal suggests that checkpoint scores miss properties that matter for subsequent behavior. We examine this in two ways. 4.1
Solution density separates the more trainable sources
The preceding results define trainability operationally: how well a source checkpoint performs after the same remaining training pipeline. But what local parameter-space geometry provides a signal of this pipeline-conditioned quality that the unperturbed checkpoint score misses? For checkpoint θ, benchmark b, perturbation ϵ, and evaluation score sb (θ), we define the relativethreshold solution-density profile following Gan et al. (2026) as δθ,b (τ ) = Pr [sb (θ + ϵ) ≥ τ sb (θ)] . ϵ
(1)
Each source checkpoint is evaluated under 100 Gaussian perturbations on GSM8K and MBPP at standard deviations σϵ ∈ {0.005, 0.001}. Figure 1 shows that C OOLDOWN is shifted toward lower retained performance on both tasks. Thresholded profiles and exact counts appear in Figure 4 and Table 11. On GSM8K, solution density at τ = 0.90 is 27% for C ONSTANT and 13% for M ERGE, but 0% for C OOLDOWN. The separation persists at σϵ = 0.001 (Figure 5 and Table 12). Figure 3 provides a complementary two-dimensional view of the perturbation outcomes. The perturbation ordering matches the controlled trainability result in Section 3.1 and Table 2: C ONSTANT and M ERGE finish nearly tied at 0.360 and 0.363, while C OOLDOWN reaches 0.247 despite its stronger pretraining signals (Table 1). We hypothesize that greater local robustness helps these checkpoints tolerate the parameter displacement induced by subsequent optimization. 4.2
Retuning SFT does not repair stopping
C OOLDOWN’s largest deficit is HumanEval+ (4.9% versus 64.6% and 66.5% for C ONSTANT and M ERGE). Raw outputs reveal a broader failure: nearly all AIME and GPQA responses reach the 32k-token cap, with highly repetitive tails. Because this pathology appears after SFT, we hypothesized that the SFT stage was responsible and swept its learning rate. 4
0
Accuracy change / %
GSM8K
10
-58
0
Accuracy change / %
MBPP
8
-74
C ONSTANT
C OOLDOWN
M ERGE
Figure 3: Filled contours show interpolated accuracy changes relative to unperturbed checkpoints on GSM8K (top) and MBPP (bottom). Stars mark each panel’s optimum. C ONSTANT has the greatest solution density, whereas M ERGE starts stronger and degrades less under perturbation than C OOLDOWN, especially on MBPP (Table 12). Projections are descriptive only. The sweep did not repair stopping. Changing the SFT factor from 1/3 to 1 leaves 100% of AIME and 99% of GPQA responses at the cap, and although it raises the post-SFT aggregate from 0.247 to 0.294, no tested rate jointly resolves the failure: lower rates improve HumanEval+ on average but reduce reasoning and the overall aggregate. Sweeping all three SFT factors for each of the nine C OOLDOWN mid/long combinations leaves all 27 checkpoints below the fixed-recipe C ONSTANT and M ERGE checkpoints. Evaluating only the first generated code block from one C OOLDOWN checkpoint raises HumanEval+ from 7.3% to 61.0%, showing latent code capability but still trailing C ONSTANT and M ERGE (64.6% and 66.5%), though thousands of repeated later blocks remain a genuine failure. Section E.1 summarizes the nested sweep with additional trajectory results, traces, and extraction details.
5
Related Work
Prior work shows that pretraining loss need not predict downstream adaptation: models with matched loss can transfer differently (H. Liu et al., 2023), and learning-rate decay can improve pretraining metrics while hurting continued training and SFT (Yano et al., 2026). Related work links flatter LLM solutions to better trainability (H. Li et al., 2024; Watts et al., 2026). We extend this line by tracking checkpoint rankings across multiple training stages and asking when they predict the final ranking. We also build on solution density, which measures whether nearby perturbations retain task performance (Gan et al., 2026), and on checkpoint merging, where checkpoints induce an effective decay over updates (Y. Li et al., 2025; Tian et al., 2026). Extended related work appears in Section A.
6
Conclusion
Checkpoint quality depends on what comes next. C OOLDOWN has better pretraining metrics than C ONSTANT but performs worse after the downstream pipeline. Rankings within a learning-rate sweep stabilize after long-context adaptation, yet neither intermediate aggregate anticipates the reversal between C ONSTANT and C OOLDOWN. Our audits further show that eval scores can miss properties relevant to continued training. Checkpoint selection should therefore target performance after the remaining pipeline, not the current score alone. Limitations. Our results come from one 30B MoE family with one seed per setup, and we did not test every combination of settings. Solution density is measured on only two tasks and does not establish causation, and C OOLDOWN’s stopping failure enlarges the observed gap. 5
References Ainslie, J., Lee-Thorp, J., Jong, M. de, Zemlyanskiy, Y., Lebrón, F., and Sanghai, S. (2023). “GQA: Training Generalized Multi-Query Transformer Models from Multi-Head Checkpoints”. In: Proceedings of the 2023 Conference on Empirical Methods in Natural Language Processing. Aleph Alpha Research (2026). Aleph Alpha Eval Framework. Version x.y.z. Austin, J. et al. (2021). Program Synthesis with Large Language Models. arXiv: 2108.07732. Bisk, Y., Zellers, R., Le bras, R., Gao, J., and Choi, Y. (2020). “PIQA: Reasoning about Physical Commonsense in Natural Language”. In: Proceedings of the AAAI Conference on Artificial Intelligence. Chen, M., Tworek, J., Jun, H., Yuan, Q., Pinto, H. P. d. O., Kaplan, J., Edwards, H., Burda, Y., Joseph, N., Brockman, G., et al. (2021). Evaluating Large Language Models Trained on Code. arXiv: 2107.03374. Chen, S., Wong, S., Chen, L., and Tian, Y. (2023). “Extending Context Window of Large Language Models via Positional Interpolation”. In: arXiv preprint arXiv:2306.15595. Clark, P., Cowhey, I., Etzioni, O., Khot, T., Sabharwal, A., Schoenick, C., and Tafjord, O. (2018). “Think You Have Solved Question Answering? Try ARC, the AI2 Reasoning Challenge”. In: arXiv preprint arXiv:1803.05457. Cobbe, K. et al. (2021). Training Verifiers to Solve Math Word Problems. arXiv: 2110.14168. Dekoninck, J., Jovanović, N., Gehrunger, T., Rögnvaldsson, K., Petrov, I., Sun, C., and Vechev, M. (2026). “Beyond Benchmarks: MathArena as an Evaluation Platform for Mathematics with LLMs”. In: arXiv: 2605.00674 [cs.CL]. Dinh, L., Pascanu, R., Bengio, S., and Bengio, Y. (2017). “Sharp Minima Can Generalize For Deep Nets”. In: Proceedings of the 34th International Conference on Machine Learning. Vol. 70. Proceedings of Machine Learning Research. PMLR, pp. 1019–1028. Foret, P., Kleiner, A., Mobahi, H., and Neyshabur, B. (2021). “Sharpness-Aware Minimization for Efficiently Improving Generalization”. In: International Conference on Learning Representations. Gan, Y. and Isola, P. (2026). Neural Thickets: Diverse Task Experts Are Dense Around Pretrained Weights. arXiv: 2603.12228 [cs.LG]. He, J., Liu, J., Liu, C. Y., Yan, R., Wang, C., Cheng, P., Zhang, X., Zhang, F., Xu, J., Shen, W., et al. (2025). Skywork Open Reasoner 1 Technical Report. arXiv: 2505.22312. Hendrycks, D., Burns, C., Basart, S., Zou, A., Mazeika, M., Song, D., and Steinhardt, J. (2021a). “Measuring Massive Multitask Language Understanding”. In: International Conference on Learning Representations (ICLR). Hendrycks, D., Burns, C., Kadavath, S., Arora, A., Basart, S., Tang, E., Song, D., and Steinhardt, J. (2021b). Measuring Mathematical Problem Solving With the MATH Dataset. arXiv: 2103.03874. Hsieh, C.-P., Sun, S., Kriman, S., Acharya, S., Rekesh, D., Jia, F., and Ginsburg, B. (2024). RULER: What’s the Real Context Size of Your Long-Context Language Models? arXiv: 2404.06654. Izmailov, P., Podoprikhin, D., Garipov, T., Vetrov, D. P., and Wilson, A. G. (2018). “Averaging Weights Leads to Wider Optima and Better Generalization”. In: Proceedings of the Thirty-Fourth Conference on Uncertainty in Artificial Intelligence, pp. 876–885. Jordan, K., Jin, Y., Boza, V., You, J., Cesista, F., Newhouse, L., and Bernstein, J. (2024). Muon: An Optimizer for Hidden Layers in Neural Networks. Joshi, M., Choi, E., Weld, D., and Zettlemoyer, L. (2017). “TriviaQA: A Large Scale Distantly Supervised Challenge Dataset for Reading Comprehension”. In: Proceedings of the 55th Annual Meeting of the Association for Computational Linguistics. Keskar, N. S., Mudigere, D., Nocedal, J., Smelyanskiy, M., and Tang, P. T. P. (2017). “On Large-Batch Training for Deep Learning: Generalization Gap and Sharp Minima”. In: International Conference on Learning Representations. Li, H., Ding, L., Fang, M., and Tao, D. (2024). “Revisiting Catastrophic Forgetting in Large Language Model Tuning”. In: Findings of the Association for Computational Linguistics: EMNLP 2024. Association for Computational Linguistics, pp. 4297–4308. Li, Y. et al. (2025). “Model Merging in Pre-training of Large Language Models”. In: arXiv preprint arXiv:2505.12082. arXiv: 2505.12082 [cs.CL]. Liang, W., Liu, T., Wright, L., Constable, W., Gu, A., Huang, C.-C., Zhang, I., Feng, W., Huang, H., Wang, J., et al. (2024). TorchTitan: One-Stop PyTorch Native Solution for Production-Ready LLM Pretraining. arXiv: 2410.06511. Liu, H., Xie, S. M., Li, Z., and Ma, T. (2023). “Same Pre-training Loss, Better Downstream: Implicit Bias Matters for Language Models”. In: Proceedings of the 40th International Conference on
6
Machine Learning. Vol. 202. Proceedings of Machine Learning Research. PMLR, pp. 22188– 22214. Liu, J., Xia, C. S., Wang, Y., and Zhang, L. (2023). Is Your Code Generated by ChatGPT Really Correct? Rigorous Evaluation of Large Language Models for Code Generation. arXiv: 2305. 01210. Loshchilov, I. and Hutter, F. (2019). “Decoupled Weight Decay Regularization”. In: International Conference on Learning Representations. NVIDIA et al. (2026). Nemotron 3 Super: Open, Efficient Mixture-of-Experts Hybrid MambaTransformer Model for Agentic Reasoning. arXiv: 2604.12374 [cs.LG]. Olmo, T. et al. (2026). Olmo 3. arXiv: 2512.13961 [cs.CL]. Ouyang, L., Wu, J., Jiang, X., Almeida, D., Wainwright, C. L., Mishkin, P., Zhang, C., Agarwal, S., Slama, K., Ray, A., et al. (2022). “Training Language Models to Follow Instructions with Human Feedback”. In: Advances in Neural Information Processing Systems 35, pp. 27730–27744. Pyatkin, V., Malik, S., Graf, V., Ivison, H., Huang, S., Dasigi, P., Lambert, N., and Hajishirzi, H. (2025). “Generalizing Verifiable Instruction Following”. In: Advances in Neural Information Processing Systems. Rein, D., Hou, B. L., Stickland, A. C., Petty, J., Pang, R. Y., Dirani, J., Michael, J., and Bowman, S. R. (2023). GPQA: A Graduate-Level Google-Proof Q&A Benchmark. arXiv: 2311.12022. Shazeer, N. (2020). GLU Variants Improve Transformer. arXiv: 2002.05202. Singh, V. et al. (2026). “Arcee Trinity Large Technical Report”. In: arXiv: 2602.17004 [cs.LG]. Su, J., Lu, Y., Pan, S., Murtadha, A., Wen, B., and Liu, Y. (2021). RoFormer: Enhanced Transformer with Rotary Position Embedding. arXiv: 2104.09864. Tian, C., Wang, J., Zhao, Q., Chen, K., Liu, J., Liu, Z., Mao, J., Zhao, W. X., Zhang, Z., and Zhou, J. (2026). “WSM: Decay-Free Learning Rate Schedule via Checkpoint Merging for LLM Pre-training”. In: International Conference on Learning Representations. arXiv: 2507.17634 [cs.LG]. Vaswani, A., Shazeer, N., Parmar, N., Uszkoreit, J., Jones, L., Gomez, A. N., Kaiser, L., and Polosukhin, I. (2017). “Attention Is All You Need”. In: Advances in Neural Information Processing Systems. Vol. 30. Wang, Y., Ma, X., Zhang, G., Ni, Y., Chandra, A., Guo, S., Ren, W., Arulraj, A., He, X., Jiang, Z., et al. (2024). MMLU-Pro: A More Robust and Challenging Multi-Task Language Understanding Benchmark. arXiv: 2406.01574. Watts, I., Li, C., Goyal, S., Springer, J. M., and Raghunathan, A. (2026). “Sharpness-Aware Pretraining Mitigates Catastrophic Forgetting”. In: Proceedings of the 43rd International Conference on Machine Learning. arXiv: 2605.02105 [cs.LG]. Yano, K., Kiyono, S., Kobayashi, S., Takase, S., and Suzuki, J. (2026). “Pre-training LLM without Learning Rate Decay Enhances Supervised Fine-Tuning”. In: International Conference on Learning Representations. arXiv: 2603.16127. Yen, H., Gao, T., Hou, M., Ding, K., Fleischer, D., Izsak, P., Wasserblat, M., and Chen, D. (2024). HELMET: How to Evaluate Long-Context Language Models Effectively and Thoroughly. arXiv: 2410.02694. Yu, Q., Zhang, Z., Zhu, R., Yuan, Y., Zuo, X., Yue, Y., Dai, W., Fan, T., Liu, G., Liu, L., et al. (2025). DAPO: An Open-Source LLM Reinforcement Learning System at Scale. arXiv: 2503.14476. Zellers, R., Holtzman, A., Bisk, Y., Farhadi, A., and Choi, Y. (2019). “HellaSwag: Can a Machine Really Finish Your Sentence?” In: Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics. Zhang, B. and Sennrich, R. (2019). “Root Mean Square Layer Normalization”. In: Advances in Neural Information Processing Systems. Vol. 32.
7
A
Related Work
Checkpoint quality and downstream adaptability. A central assumption in model development is that better intermediate metrics imply a better model for subsequent training. Prior work shows that this need not hold even when pretraining loss is matched. H. Liu et al. (2023) find language models with similar pretraining loss but substantially different downstream transfer performance and associate this difference with implicit biases of the pretraining procedure. More directly, Yano et al. (2026) compare decay-based and decay-free LLM pretraining and find that learning-rate decay can improve pretraining metrics while reducing performance after SFT. Their result persists with additional mid-training and is associated with sharper pretrained solutions. Related work on LLM fine-tuning connects loss-landscape geometry to catastrophic forgetting, finding that flatter solutions preserve pretrained capabilities better during adaptation (H. Li et al., 2024). Our setting complements these results by following checkpoint rankings through several consecutive training stages and asking when intermediate evaluations become predictive of the final post-SFT ranking. Flat minima and downstream adaptability. The relation between local geometry and model quality has a long history in neural networks. Flat minima have been associated with improved generalization (Keskar et al., 2017), and methods such as stochastic weight averaging (Izmailov et al., 2018) and sharpness-aware minimization (Foret et al., 2021) explicitly favor broad low-loss regions. At the same time, flatness is not an intrinsic property without specifying a parameterization and metric: functionally equivalent networks can have arbitrarily different sharpness under common definitions (Dinh et al., 2017). For language models, H. Liu et al. (2023) show that pretraining-loss flatness, measured through Hessian-based quantities, correlates with downstream transfer even when pretraining loss itself does not. H. Li et al. (2024) connect flatter LLM loss landscapes to reduced forgetting during fine-tuning, while Yano et al. (2026) find that learning-rate decay leads to sharper pretrained solutions together with weaker post-SFT performance. Watts et al. (2026) show that SAM reduces forgetting after post-training and quantization. These findings motivate asking whether local geometry contains information about how a checkpoint will behave under subsequent training. Our solution-density probe is closely related to this literature but measures a different object. A conventional local flatness measure evaluates the change in the pretraining loss around checkpoint θ. For example, for ϵ ∼ N (0, σ 2 I), one may consider Fpre (θ, σ) = Eϵ [Lpre (θ + ϵ) − Lpre (θ)] .
(2)
Under a local second-order approximation, σ2 Tr ∇2 Lpre (θ) , 2 connecting isotropic perturbation flatness to Hessian-based sharpness. Fpre (θ, σ) ≈
(3)
Our probe instead evaluates whether nearby parameters preserve an external capability. For benchmark b with score sb , we measure sb (θ + ϵ) δθ,b (τ ; σ) = Pr 2 ≥τ . (4) sb (θ) ϵ∼N (0,σ I) The two quantities therefore probe the same local parameter neighborhood but apply different functions to it: Lpre (θ + ϵ) vs. sb (θ + ϵ) . | {z } | {z } capability retention
pretraining-loss geometry
There is no general implication between them. A direction can increase pretraining loss while leaving a benchmark capability unchanged, or preserve pretraining loss while disrupting parameters important for a particular downstream capability. Moreover, benchmark scores are generally discrete and need not share the local differential geometry of the pretraining objective. Flatness and solution density may therefore reflect a common robustness of the learned solution, but they are not equivalent measures. We observe that C ONSTANT and M ERGE have substantially more performance-retaining neighborhoods than C OOLDOWN, consistent with the hypothesis that local geometry relates to subsequent 8
adaptability. However, our experiment does not show that solution density causes, or generally predicts, better post-training. We therefore use it as a diagnostic of checkpoint geometry rather than a causal explanation. Checkpoint averaging and learning-rate decay. Checkpoint averaging provides a second connection to the geometry of our three pretraining sources. Averaging points along an optimization trajectory is well established: stochastic weight averaging, for example, combines checkpoints obtained under sustained or cyclical learning rates and tends to locate broader solutions than the final SGD iterate (Izmailov et al., 2018). Checkpoint merging has since been applied directly during LLM pretraining. Y. Li et al. (2025) study merging across dense and MoE models and show that combining checkpoints from constant-learning-rate trajectories can recover much of the benefit of explicit annealing. Tian et al. (2026) develop this connection formally and show how checkpoint weights induce an effective decay over the updates along a constant-LR trajectory. Checkpoint merging has also been adopted in large-scale LLM training systems such as Nemotron Super (NVIDIA et al., 2026).
B
Model Architecture and Training Details
B.1
Architecture
The model is a decoder-only Transformer (Vaswani et al., 2017) with 50 layers, of which the first two are dense and the remaining 48 are mixture-of-experts (MoE) layers. It has approximately 30B total parameters and activates approximately 3B parameters per token. Table 3 gives the complete architecture. Table 3: Model architecture. Component
Value
Vocabulary size Layers Hidden size Attention Query / key-value heads Head size Query-key normalization Dense MLP hidden size Expert hidden size MLP activation Routed / shared experts per MoE layer Experts selected per token Routing Pretraining sequence length Position encoding
96,000 50 total: 2 dense, 48 MoE 2,048 Causal grouped-query attention 32 / 4 128 Per-head RMSNorm 6,144 768 SwiGLU 128 / 1 8 Token-choice, top-8 sigmoid routing 4,096 RoPE, base θ = 500,000
Each attention block uses causal grouped-query attention with 32 query heads and four key-value heads (Ainslie et al., 2023). Query and key representations are normalized independently within each head using RMSNorm (B. Zhang et al., 2019). We use rotary position embeddings (RoPE) with base θ = 500,000 (Su et al., 2021). Both dense and expert MLPs use SwiGLU activations (Shazeer, 2020). Each MoE layer contains 128 routed experts and one shared expert. Token-choice sigmoid routing selects the top eight routed experts for each token. The shared expert is applied independently of this selection. B.2
Pretraining data and objective
All parameters are randomly initialized with output matrices initialised to zero. We pretrain with the causal next-token-prediction objective on a broad and diverse document corpus containing English and German data, among other sources. Documents are packed into 4,096-token sequences. Pretraining covers approximately 7.508T tokens over 35,800 optimization steps. The global batch contains 51,200 sequences, or 209.715M tokens per step, using five gradient-accumulation steps 9
across 512 GPUs. The learning rate is warmed up linearly for 358 steps and then held constant for the remaining 35,442 steps. B.3
Optimization
We partition parameters by role and dimensionality. The token embedding matrix, all one-dimensional backbone parameters, router weights, and language-model head are optimized with AdamW (Loshchilov et al., 2019). Their respective learning rates are 0.02916, 0.02916, 0.0001139, and 0.0004556. All other two-dimensional backbone parameters are optimized with Nesterov Muon using spectral norm convention at learning rate 0.01458 (Jordan et al., 2024). Thus embeddings and the output head use AdamW despite being matrices, while Muon is reserved for internal two-dimensional backbone parameters. We apply weight decay 0.0001221 independently to all decay-eligible parameters. Embeddings, normalization parameters, and expert-balancing biases are excluded. The expert-balancing bias is updated separately using SMEBU (Soft-clamped Momentum Expert Bias Updates; Singh et al., 2026). During pretraining, an auxiliary load-balancing loss further encourages uniform expert utilization after routing. B.4
Continued training and systems
For capability-focused continued training, the sequence length increases to 8,192 tokens and the global batch decreases to 25,600 sequences, preserving the 209.715M-token batch. The 100B-token phase therefore spans approximately 477 optimization steps. This phase uses a 50-step linear warmup followed by a constant learning rate; long-context adaptation and SFT use the same 50-step warmup. The auxiliary MoE load-balancing loss is disabled during this phase and remains disabled during long-context adaptation. MoE load remains balanced throughout this stage despite this. Long-context adaptation and SFT both use 65,536-token sequences. Long-context adaptation keeps the 209.715M-token global batch, SFT uses a global batch of 384 sequences (25.2M tokens per step), i.e. approximately 397 optimization steps for its 10B tokens. During long-context adaptation the learning rate decays to 1/3 of its peak; during SFT it decays from its peak to an absolute value of 10−5 , for every SFT learning-rate factor. Training runs on 512 GPUs using a PyTorch-based TorchTitan distributed-training stack (Liang et al., 2024). We train in bfloat16 with optimizer states, gradients, accumulation and master weights in float32. Gradients are clipped to a maximum global norm of 1.0. Optimizer state is reset and learning-rate warmup is repeated at each stage, as described in the main text. B.5
M ERGE construction and effective update weights
M ERGE does not add optimization steps. It combines a trailing window from the C ONSTANT run using WMA coefficients that rise linearly from the oldest to the newest checkpoint. For checkpoints θ1 , . . . , θ20 ordered from oldest to newest, the selected source is 20 X i θmerge = θi . 210 i=1
(5)
The oldest checkpoint therefore receives weight 1/210, while the newest receives weight 20/210. Pi To expose the effect on individual updates, write θi = θ0 + t=1 ∆θt , where ∆θt = θt − θt−1 . Substitution and exchange of the summations give θmerge = θ0 +
20 X
qt ∆θt ,
t=1
qt =
20 X i . 210 i=t
(6)
Although the checkpoint weights increase linearly, their cumulative coefficients on parameter updates decrease across the window: q1 = 1 for the earliest update and q20 = 20/210 ≈ 0.095 for the latest. The merge therefore attenuates later updates after training. This resembles learning-rate decay at the level of the final weighted update history, but it is not equivalent to cosine cooldown because the optimizer never follows the decayed trajectory. 10
Y. Li et al. (2025) report that merging constant-learning-rate checkpoints can attain performance comparable to annealed pretraining checkpoints, and Tian et al. (2026) formalize model averaging schemes that emulate several decay schedules. Their experiments use other training runs and model families, so they motivate this source construction rather than validate it for our pipeline.
C
Evaluation Suites and Aggregate Construction
We use a stage-specific evaluation suite at each training boundary. All evaluations are run with (Aleph Alpha Research, 2026). Pre, Mid, and Long use completion-style inference, while SFT uses chat-formatted prompts and the original inference and evaluation adapters. The suite also expands across the pre-SFT stages: Mid adds long-context evaluation to the Pre clusters, and Long uses a larger code and long-context suite. Table 4 lists every score included in each reported aggregate. Table 4: Stage-specific evaluation suites and aggregation. Lengths in parentheses are context lengths. Pre, Mid, and Long first average scores within each cluster and then weight the cluster means equally. SFT directly averages its six benchmark values. Stage
Cluster
Benchmark scores
Pre
General EN
Math EN Code EN
ARC (Clark et al., 2018), HellaSwag (Zellers et al., 2019), MMLU (Hendrycks Equal mean of 3 cluster et al., 2021a), MMLU-Pro (Y. Wang et al., 2024), PIQA (Bisk et al., 2020), means TriviaQA (Joshi et al., 2017) GSM8K (Cobbe et al., 2021), MATH Minerva (Hendrycks et al., 2021b) HumanEval (M. Chen et al., 2021)
General EN
ARC, HellaSwag, MMLU, MMLU-Pro, PIQA, TriviaQA
Math EN Code EN Long Context
GSM8K, MATH Minerva HumanEval HELMET JSON-KV (Yen et al., 2024) (8k); RULER (Hsieh et al., 2024) NIAH and QA (4k, 8k, 16k, 32k, 64k); RULER VT and WE (4k, 8k, 16k)
General EN
ARC, HellaSwag, MMLU, MMLU-Pro, PIQA, TriviaQA
Math EN Code EN Long Context
GSM8K, MATH Minerva HumanEval, MBPP (Austin et al., 2021) HELMET JSON-KV (8k, 16k, 32k, 64k); RULER NIAH, QA, and VT (4k, 8k, 16k, 32k, 64k); RULER WE (4k, 8k, 16k)
Direct task scores
AIME 2026 (Dekoninck et al., 2026), DAPO Math (Yu et al., 2025), Skywork Unweighted mean of 6 OR1 Math (J. He et al., 2025), IFBench (Pyatkin et al., 2025), HumanEval+ (J. benchmark values Liu et al., 2023), GPQA Diamond CoT (Rein et al., 2023)
Mid
Long
SFT
Aggregation
Equal mean of 4 cluster means
Equal mean of 4 cluster means
For checkpoint i, let si,b ∈ [0, 1] be the scalar score for benchmark configuration b, and let µi (B) denote the mean over a set of configurations B. Let G and M denote the General EN and Math EN sets in Table 4. Ct and Lt denote the stage-specific Code EN and Long Context sets, and S denotes the six SFT benchmark values. The reported stage aggregate is 1 X µi (B) = si,b , |B| b∈B µ (G) + µi (M) + µi (CPre ) i , t = Pre, 3 (7) µi (G) + µi (M) + µi (Ct ) + µi (Lt ) (t) , t ∈ {Mid, Long}, Ai = 4 1X si,b , t = SFT. 6 b∈S
Consequently, individual pre-SFT benchmarks do not all have equal final weight. For example, Pre contains nine benchmark scores, but HumanEval alone forms the Code EN cluster and therefore receives one third of the aggregate rather than one ninth. Mid and Long similarly assign one quarter of the aggregate to each cluster, irrespective of the number of scores in that cluster. Each listed long-context task–length pair is one score when forming its Long Context cluster mean. The SFT aggregate is instead a direct unweighted mean of six benchmark values. IFBench contributes one value, formed by averaging its loose and strict prompt-level scores. This single value enters the six-task mean. We never pool example-level correct counts across benchmarks. Invalid or unparsable answers are scored according to the original adapter and are not removed before averaging. 11
For HumanEval+, the primary score is pass@1 after extracting the final fenced Python block, matching the original adapter. Section F reports the diagnostic first-block rescore, which is not substituted into any primary aggregate. Table 2 gives all six post-SFT benchmark scores and their aggregate for the C ONSTANT, C OOLDOWN, and M ERGE comparison. Table 5 gives the corresponding intermediate and post-SFT aggregates for every training trajectory used in the correlation analysis.
D
Complete Training Trajectories
Table 5 lists the training trajectories used in the stage-association analysis. Every post-SFT aggregate in that analysis comes from a checkpoint trained with SFT learning-rate factor 1/3. No rate is selected per trajectory. Separately, each of the nine C OOLDOWN trajectories receives an auxiliary SFT learning-rate sweep, summarized in Section E.1. Table 5: Intermediate and post-SFT aggregate scores. All eleven trajectories enter the reported correlations. Source
Mid, Long factors
C OOLDOWN C OOLDOWN C OOLDOWN C OOLDOWN C OOLDOWN C OOLDOWN C OOLDOWN C OOLDOWN C OOLDOWN C ONSTANT M ERGE
1, 1 1, 1/3 1, 1/9 1/3, 1 1/3, 1/3 1/3, 1/9 1/9, 1 1/9, 1/3 1/9, 1/9 1, 1 1, 1
Mid aggregate
Long aggregate
Post-SFT aggregate
0.515 0.564 0.544 0.557 0.543 0.537 0.516 0.521 0.512 0.556 0.540
0.637 0.626 0.621 0.626 0.618 0.610 0.602 0.588 0.581 0.636 0.644
0.247 0.199 0.187 0.228 0.191 0.119 0.160 0.104 0.087 0.360 0.363
Table 6: Association between intermediate and post-SFT aggregate scores across the eleven training trajectories. Reported p-values are two-sided. Signal Mid-training aggregate Long-context aggregate
Pearson r
p
Spearman ρ
p
0.473 0.884
0.142 < 0.001
0.482 0.964
0.133 < 0.001
The correlations are computed from the unrounded aggregate scores using two-sided tests. Pearson coefficients have 95% confidence intervals of [−0.177, 0.836] after mid-training and [0.605, 0.970] after long-context adaptation. Spearman coefficients use average ranks. Restricting the analysis to the nine C OOLDOWN trajectories yields mid-training r = 0.448 (p = 0.226) and ρ = 0.450 (p = 0.224), compared with long-context r = 0.935 (p < 0.001) and ρ = 0.950 (p < 0.001).
E
Stopping and Repetition Audit
We compare retained raw generations from the fixed-recipe C OOLDOWN checkpoint with the corresponding C ONSTANT and M ERGE checkpoints. After observing severe repetition, we additionally inspect C OOLDOWN (1, 1) with SFT factor 1, which has the highest post-SFT aggregate among the complete C OOLDOWN checkpoints in the recorded SFT-rate sweep. The audit covers 294 IFBench prompts, 240 AIME generations, and 198 GPQA Diamond prompts per checkpoint under a 32,768-token payload cap. AIME and GPQA record generated sequence length directly. IFBench is retokenized with the tokenizer used for training. “near cap” denotes at least 32,000 tokens. The fixed-recipe C OOLDOWN checkpoint is substantially less likely to terminate than the C ONSTANT and M ERGE checkpoints. Median zlib-to-raw byte ratios are 0.019, 0.044, and 0.027 on IFBench, AIME, and GPQA, compared with 0.288, 0.310, and 0.059 for C ONSTANT and 0.299, 0.328, and 0.066 for M ERGE. The long C OOLDOWN outputs therefore contain extensive repeated material. Changing C OOLDOWN from SFT factor 1/3 to factor 1 does not repair the behavior. The corresponding compression ratios remain 0.022, 0.053, and 0.031. 12
Table 7: Stopping behavior under the 32,768-token cap. Cells report median tokens followed by the percentage near the cap for IFBench or exactly at the cap for AIME and GPQA. Checkpoint
IFBench
AIME
GPQA
C OOLDOWN (1, 1), SFT 1/3 32,616 / 75.9% 32,768 / 100% 32,768 / 99.0% C ONSTANT (1, 1), SFT 1/3 1,899 / 33.0% 6,552 / 29.6% 32,768 / 54.5% M ERGE (1, 1), SFT 1/3 1,805 / 33.3% 5,135 / 25.8% 32,768 / 54.0% C OOLDOWN (1, 1), SFT 1
E.1
32,676 / 93.2% 32,768 / 100% 32,768 / 99.0%
SFT-rate diagnostic
The additional C OOLDOWN checkpoints are used only as a debugging sweep. All three SFT factors are complete for all nine mid/long trajectories, giving 27 evaluated checkpoints. These additional checkpoints are not included in the primary source comparison or trajectory correlations. Table 8: Auxiliary C OOLDOWN diagnostic by SFT learning-rate factor. Values are descriptive averages across complete runs. Reasoning is the mean of AIME, DAPO, Skywork, and GPQA. SFT factor
n
Post-SFT aggregate
Reasoning
HumanEval+
1 1/3 1/9
9 9 9
0.190 0.169 0.153
0.187 0.140 0.109
0.168 0.243 0.293
On the fixed (1, 1) upstream trajectory, changing the SFT factor from 1/3 to 1 increases the postSFT aggregate from 0.247 to 0.294, the highest recorded complete C OOLDOWN result. Its original HumanEval+ score is 0.067. Across upstream schedules, lower SFT learning rates improve HumanEval+ on average while reducing the reasoning and overall aggregates. No tested SFT factor therefore dominates all recorded capabilities. Final SFT loss is also not a sufficient selection signal: the lowest-loss checkpoint is the extraction-sensitive C OOLDOWN (1, 1), SFT-factor-1 checkpoint rather than a uniformly superior assistant. E.2
Compact-answer sensitivity
Table 9: Official accuracy and a heuristic first-explicit-answer rescore on identical generations. The heuristic compares the earliest explicit or boxed answer with the target and is used only diagnostically. Checkpoint C OOLDOWN (1, 1), SFT 1/3 C OOLDOWN (1, 1), SFT 1 C ONSTANT (1, 1), SFT 1/3 M ERGE (1, 1), SFT 1/3
AIME official → first
GPQA official → first
17.08% → 20.42% 18.33% → 23.75% 22.08% → 22.08% 21.67% → 21.25%
20.71% → 22.73% 25.25% → 29.80% 23.74% → 23.74% 22.22% → 24.24%
First-answer sensitivity is measurable for C OOLDOWN but much smaller than the HumanEval+ first-block effect. Official GPQA parsing marks 115/198 and 106/198 C OOLDOWN responses invalid, but also 111/198 C ONSTANT and 114/198 M ERGE responses, so GPQA extraction fragility is not specific to C OOLDOWN. IFBench strict scores likewise do not directly expose the stopping failure: C OOLDOWN with SFT factors 1/3 and 1 scores 23.47% and 22.79%, compared with 22.11% for the corresponding C ONSTANT and M ERGE checkpoints. The same response pathology therefore interacts differently with different evaluation contracts. E.3
Qualitative examples
The following examples illustrate, rather than estimate, the observed repetition pattern. An IFBench response initially satisfies exact keyword-count constraints and subsequently repeats “(Answer ready.)” 575 times and its keyword audit 685 times, causing the response to violate the same constraints it had already satisfied. 13
One AIME generation repeats “I hope it is correct.” 4,208 times, reaches the generation cap, and drifts from an earlier correct answer of 190 to an officially extracted answer of 290. One GPQA response repeats the sentence “the nitro group is attached to the carbon bearing the nitro group” 2,012 times, reaches the cap mid-word, and is marked invalid despite stating the correct option earlier. The most extreme 32k HumanEval+ response contains 5,318 fenced code blocks but only four unique blocks. One block appears 5,314 times. These examples show why one underlying stopping failure can produce different measured penalties under whole-response, compact-answer, and final-block evaluators.
F
HumanEval+ Re-execution Audit
For response y, extractor E, and functional verifier V , the measured code score is SE = E[V (E(y))].
(8)
The original evaluation adapter extracts the final fenced Python block. We compare it with the first fenced block. Selecting the first syntactically valid block produces the same aggregate as selecting the first fenced block in every audited run. Table 10: HumanEval+ pass@1 under alternative extraction from identical generations. Counts are out of 164. Checkpoint
Cap
C OOLDOWN (1, 1), SFT 1 C OOLDOWN (1, 1), SFT 1 C ONSTANT (1, 1), SFT 1/3 M ERGE (1, 1), SFT 1/3
1k 32k 32k 32k
Last block
First block
∆
9 (5.49%) 97 (59.15%) +53.66 12 (7.32%) 100 (60.98%) +53.66 108 (65.85%) 106 (64.63%) −1.22 108 (65.85%) 107 (65.24%) −0.61
All scores reuse the same raw generations. Candidate programs are executed against the original 164 HumanEval+ test suites in a network-disabled container. Re-executing the final block reproduces every original verdict, so extraction is the only changed scoring variable. For both audited C OOLDOWN caps, no last-block success becomes a first-block failure. Every gain comes from a response containing a passing early program and a failing final program. At 32k, C OOLDOWN has a median response length of 32,690 tokens and a 72.6% near-cap rate, compared with medians of 657 and 484 tokens and near-cap rates of 15.2% and 7.3% for C ONSTANT and M ERGE. C OOLDOWN also has a median of 1,609 fenced blocks per response, with a maximum of 5,318, compared with a median of one and maxima of four and two for C ONSTANT and M ERGE. Increasing the output cap does not repair the behavior. C OOLDOWN’s first-block accuracy remains near 60% while its generations continue to contain thousands of repeated blocks. The fixed-recipe C OOLDOWN SFT-1/3 checkpoint has an original HumanEval+ score of 8/164 (4.88%), but its evaluation artifact retains only aggregate scores. First-block rescoring would require raw generations that are no longer available. We therefore do not transfer the corrected score from SFT factor 1 to the primary source comparison.
G
Solution Density Probe Details
G.1
Estimator
For checkpoint θ, benchmark b, threshold τ , and nθ,b valid perturbations, we estimate nθ,b 1 X sb (θ + ϵi ) δbθ,b (τ ) = 1 ≥τ . nθ,b i=1 sb (θ)
(9)
Each score is divided by the unperturbed score from the same checkpoint. We retain the latest result for each matched perturbation job and exclude results without numeric benchmark summaries. 14
Missing values are not imputed. Both perturbation scales contain 100 perturbations per checkpoint and benchmark. Primary perturbation scale b ) (%) Solution density δ(τ
G.2
(a) GSM8K
(b) MBPP
100 75 50 25 0 50
70
90
100
110
50
Threshold τ (% of unperturbed score)
70
90
100
110
Threshold τ (% of unperturbed score)
C ONSTANT
C OOLDOWN
M ERGE
Figure 4: Solution-density profiles at σϵ = 0.005. Each curve reports the fraction of 100 perturbations retaining at least threshold τ of the corresponding unperturbed score. Table 11: Perturbations above each score threshold at σϵ = 0.005. Cells report counts out of 100 followed by percentages. Benchmark
Checkpoint
n
τ ≥ 0.90
τ ≥ 0.95
τ ≥ 1.00
τ ≥ 1.05
τ ≥ 1.10
GSM8K
C ONSTANT C OOLDOWN M ERGE
100 100 100
27/100 (27.0%) 0/100 (0.0%) 13/100 (13.0%)
10/100 (10.0%) 0/100 (0.0%) 0/100 (0.0%)
2/100 (2.0%) 0/100 (0.0%) 0/100 (0.0%)
1/100 (1.0%) 0/100 (0.0%) 0/100 (0.0%)
0/100 (0.0%) 0/100 (0.0%) 0/100 (0.0%)
MBPP
C ONSTANT C OOLDOWN M ERGE
100 100 100
3/100 (3.0%) 0/100 (0.0%) 0/100 (0.0%)
1/100 (1.0%) 0/100 (0.0%) 0/100 (0.0%)
0/100 (0.0%) 0/100 (0.0%) 0/100 (0.0%)
0/100 (0.0%) 0/100 (0.0%) 0/100 (0.0%)
0/100 (0.0%) 0/100 (0.0%) 0/100 (0.0%)
Zero cells indicate that no success was observed among the 100 perturbations, not that the underlying probability is zero. We therefore treat the curves as empirical profiles rather than estimated continuous densities. Smaller perturbation scale b ) (%) Solution density δ(τ
G.3
(b) MBPP
(a) GSM8K 100 75 50 25 0 90
95
100
105
Score threshold τ (% of unperturbed score) C ONSTANT
110
90
95
100
105
110
Score threshold τ (% of unperturbed score) C OOLDOWN M ERGE
Figure 5: Solution-density profiles at σϵ = 0.001. C ONSTANT has the greatest density around the unperturbed score. On MBPP, 3% of C OOLDOWN perturbations match or exceed the unperturbed score, compared with 50% for C ONSTANT and 20% for M ERGE. Values above 100% indicate that a perturbed checkpoint outscored its single unperturbed evaluation. They should not be interpreted as established local improvements without repeated unperturbed evaluations. The stable observation is relative: C OOLDOWN has lower solution density around performance-preserving thresholds than C ONSTANT and M ERGE, particularly on MBPP.
15
Table 12: Results at σϵ = 0.001. Mean is the average perturbed score as a percentage of the corresponding unperturbed score. Benchmark
Checkpoint
n
Mean
τ ≥ 0.95
τ ≥ 1.00
τ ≥ 1.05
τ ≥ 1.10
GSM8K
C ONSTANT C OOLDOWN M ERGE
100 100 100
103.5% 100.7% 100.1%
100/100 (100.0%) 98/100 (98.0%) 100/100 (100.0%)
92/100 (92.0%) 57/100 (57.0%) 54/100 (54.0%)
24/100 (24.0%) 7/100 (7.0%) 0/100 (0.0%)
2/100 (2.0%) 0/100 (0.0%) 0/100 (0.0%)
MBPP
C ONSTANT C OOLDOWN M ERGE
100 100 100
99.7% 95.2% 98.3%
96/100 (96.0%) 54/100 (54.0%) 94/100 (94.0%)
50/100 (50.0%) 3/100 (3.0%) 20/100 (20.0%)
2/100 (2.0%) 0/100 (0.0%) 0/100 (0.0%)
0/100 (0.0%) 0/100 (0.0%) 0/100 (0.0%)
16