Conceptio › Archive › arXiv CS
arXiv CSopen access

Minibatch persistency, eight years later: what batch reuse costs in steps and joules, and what it saves in data

· arxiv_cs
arXiv CS · Papers · License: Open Access
Open Source ↗Direct PDF ↓
clouddistributed-computingparallel-computing
distributed computing, parallel computing, cloud

Minibatch persistency, eight years later: what batch reuse costs in steps and joules, and what it saves in data Matteo Fischetti

[email protected]

arXiv:2609.13922v1 [cs.LG] 12 Sep 2026

Department of Information Engineering University of Padova, Italy

Abstract Minibatch persistency reuses data instead of reading it: rather than drawing a fresh minibatch at every optimizer step, it takes K consecutive steps on the same one. Absorbed into data echoing in 2019, it has carried one objection — that reuse merely imitates a larger learning rate — and no baseline tuned as carefully as the method itself. This paper runs the missing test. A pre-registered study trains a 49 M-parameter Transformer on FineWeb-Edu at minibatch size B ∈ {32, 128, 512}, 8 seeds per cell, tuning the learning rate separately for every batch size and every arm, against a reuse-free control that changes the sampling and nothing else. Each headline claim is a cost to reach a fixed loss, read on four axes: optimizer steps, fresh tokens, seconds, and joules at the socket. We then replicate on new seeds and a newer GPU generation, and put the three arms on one schedule in steps. The outcome of our study is that what minibatch reuse buys is neither speed nor energy but data, and only at large minibatch size: at B = 32 it reads more fresh tokens than the baseline, not fewer. On steps, seconds and joules it is at best free; and at B = 512, where it looks best, a registered control cannot separate the effect of reuse from the position on the learning-rate schedule at n = 8 seeds. The technique is therefore worth using where fresh data rather than compute is the binding cost: a corpus that runs out, a pipeline that pays per sample, a stream that cannot be rewound. Where the data can simply be read again, spaced epochs do as well or better.

1

Introduction

Training a modern neural network is an exercise in paying for gradients. Every optimizer step consumes a minibatch, and producing that minibatch is not free: it must be read from storage, decoded, augmented, collated and moved to the accelerator, and in a large fraction of real pipelines this work—not the matrix multiplications—is what sets the pace. The natural question is whether each minibatch has to be spent after a single use. If a batch of B examples is a good enough sample of the data distribution, one could take K consecutive optimizer steps on it before paying for the next one, thus dividing the cost of the input pipeline by K. We call this minibatch persistency in what follows, following the name introduced by Fischetti et al. (2018). The idea is eight years old, and its history is instructive. It was proposed in 2018 with encouraging but small-scale evidence, and thirteen months later it was cited—from the first preprint version—by Choi et al. (2019), who generalized it into a family of techniques called data echoing and, in the same paragraph, stated the objection that has shadowed it ever since: “Fischetti et al. (2018) describe a special case of data echoing they call ‘minibatch persistency’ that reuses minibatches for multiple consecutive SGD updates. They run experiments on CIFAR-10, but do not tune hyperparameters for the baseline or for their method. Neither their method nor their baseline reach competitive test numbers in their experiments, leaving 1

open the question of whether minibatch persistency has an advantage over a well-tuned baseline.” The objection is correct, and it is a methodological one, not a conceptual one: the criticism is not that the idea is wrong but that the 2018 evidence could not tell. It is also, by now, the standard criticism of the whole efficient-training literature, codified by Shallue et al. (2019) and sharpened by Kaddour et al. (2023) into the observation that most claimed training speed-ups evaporate once the baseline is tuned with the same care as the method. Note that the objection bites with particular force here. At first order, taking K steps on the same batch at learning rate η resembles a single step at a learning rate of about Kη; hence an untuned comparison between K > 1 and K = 1 may be measuring nothing but a learning rate, and any honest revival of the idea has to rule that out before it says anything else. The conjecture underneath the method has since been examined once on tuned baselines. Choi et al. (2019) report, for batch echoing at e = 2, that repeated batches approximate fresh batches better as the batch size approaches the training set size—the 2018 conjecture in their words (Section 2). Two things, however, were left undone. First, the larger reuse factors in that paper belong to example echoing, in which duplicated examples are reshuffled into fresh batches; batch echoing, which is minibatch persistency in the strict sense, was swept only up to e = 2 and is the weakest member of the family. Second, and more surprising for a technique whose entire rationale is saving work, nobody has ever reported what it saves in joules. This paper takes up both. Its aim is not to propose the idea, which Fischetti et al. (2018) did, but to measure it under the conditions the 2019 objection demands. The main contributions are as follows. • A tuned answer to the tuned-baseline objection. We tune the learning rate independently inside every (batch size, arm) cell, with the same number of learning rates per arm within a batch size, and we report ρ(B) = η ⋆ (K=4)/η ⋆ (K=1) as a first-class result, where η ⋆ (K) is the cost-minimizing learning rate of the arm with reuse factor K. It is a registered descriptive diagnostic, read from a single-seed coarse sweep with no interval, of whether persistency is anything more than a reparametrized learning rate. At B = 32 and B = 128 it is inconsistent with that reading; at B = 512 it is reported as indeterminate. • A pre-registered confirmatory design. The hypothesis, the endpoint, the target definition, the estimator, the exclusion rules and the stopping rule for the seed count were registered on the OSF Registries before the confirmatory runs, and the analysis decisions taken afterwards were themselves timestamped. Section 4 describes the apparatus and Section 8 lists what a reader receives to check it, so that the order in which things happened need not be taken on our word. • Cost measured on four axes, including energy. Every evaluation logs optimizer steps, fresh tokens, wall-clock seconds and joules, and every claim except the epoch contrast of Section 6.9 is a cost-to-target claim. To our knowledge this is the first study of batch reuse that reports joules at all, let alone joules with a stated provenance and a per-device attribution argument (Section 5). • A calibrated noise floor. Alongside the baseline and the persistency arm we run an A/A arm that is distributionally identical to the baseline by construction. Its measured ratio to the baseline has true value one, so its dispersion across seeds tells the reader—and told us—how small a ratio this apparatus can honestly resolve. Two claims run through the paper, and we keep them separable on purpose. The first is an optimization claim: at a fixed target validation loss, how many optimizer steps and how many fresh tokens does reuse cost, and how does that penalty behave as the batch size grows. The second is an energy claim: on a pipeline where producing data costs real work, how many joules does reuse save at the same target. They are supported by different evidence and can fail independently—in particular, an energy saving that turned out to be entirely explained by the input pipeline would say nothing about optimization, and we would say so in those words rather than let the two be read as one. 2

Who the data reading is for. The optimization claim has a practical reader. Where unique data is the scarce resource—low-resource languages, domain corpora, licensed or annotated text, the data-constrained regime of Muennighoff et al. (2023)—reuse at large batch reaches the same loss in the same number of steps on 0.221 of the fresh tokens. Where every fresh example costs work before it reaches the accelerator—decoding and augmentation, simulation, environment rollouts and policy generations in reinforcement learning, synthetic data—the tokens saved are the pipeline’s cost, which is the setting data echoing was built for (Choi et al., 2019). And where the stream cannot be rewound, reuse is the only way to take more than one step per arriving batch. Two limits travel with the offer. It holds at large batch: at B = 32 reuse costs data rather than saving it. And where the data can be revisited, four epochs spaced out are at least as good as immediate reuse (Section 6.9), so persistency earns its place only where a second pass is impossible or a fresh sample has a price. The paper is organized as follows. Section 2 places the idea in the literature, including a recent theoretical line that studies batch reuse without reference to either the 2018 or the 2019 work. Section 3 defines the two knobs of the design and the cost-to-target estimator. Section 4 describes the experimental apparatus and the pre-registration. Section 5 is devoted to the measurement of energy, which turned out to be the hardest part of the study and deserves a section of its own. Section 6 reports the results. Finally, conclusions and directions for future research are drawn in Section 7.

2

Related work

Minibatch persistency. The technique was proposed by Fischetti et al. (2018) on small-scale evidence, with no baseline tuned to the same effort and no reading of how the effect moves with the batch size; supplying that test is the present study. We next review the four lines of work that bear on batch reuse, and say plainly where our study sits with respect to each. Data echoing and the fate of the idea. Choi et al. (2019) introduced data echoing as a family of methods that insert a repetition point somewhere in the input pipeline and reuse the output of every upstream stage e times. Minibatch persistency is the member of that family in which the repetition point sits after batching, so the reused object is the assembled minibatch itself. Their empirical study covers Transformer on LM1B, ResNet-50 on ImageNet and SSD on COCO, with quasi-random hyperparameter search (of the order of one hundred trials per configuration) for both the baseline and the treatment. We adopt their requirement that baseline and treatment be tuned with the same effort, at a far smaller search budget: four rates per cell against their hundred trials. Two of their findings frame our own. Their Section 3.4 sweeps the batch size for batch echoing at e = 2, on Transformer/LM1B and ResNet-50/ImageNet, and reports that “as the batch size increases, the performance of batch echoing relative to the baseline either stays the same or improves” (their Figure 6). That is the 2018 conjecture, tested once on tuned baselines and at the smallest reuse factor. The larger reuse factors in that paper belong to example echoing, where duplicated examples are reshuffled into new batches. What that leaves to the present study is the reuse factor K = 4, the cost-to-target reading on four axes including joules, and the pre-registration. The paper itself was not accepted at ICLR 2020 and remains an unrefereed preprint, with a substantial citation record. Reuse elsewhere in systems and optimization. The mechanism has been rediscovered wherever the input pipeline is expensive. Ramezani et al. (2020) recycle a sampled minibatch across several steps in graph neural network training, with convergence guarantees, and cite the 2018 work by name; Agarwal et al. (2020) analyze stochastic optimization when the data pipeline lags the accelerator, and show that reuse accelerates the curvature-dominated phase of convergence while leaving the statistical rate intact, a prediction this study does not test. We do not claim novelty of the mechanism itself. The theory of batch reuse, 2024–2025. A recent theoretical line studies what changes when the same batch is used more than once, and reaches a conclusion stronger than a throughput argument: repetition changes what is learnable. Dandi et al. (2024) show that several gradient steps on the same batch allow two-layer networks to learn functions that single-pass SGD provably cannot, and related work extends the picture to multi-index targets (Arnaboldi et al., 2024) and to scaling laws under data reuse (Lin et al., 2025). 3

At the time of writing, none of these three works cites either the 2018 or the 2019 paper—the bridge between the theory of repetition and the practice of echoing has not been built. We do not build it here either; our contribution is empirical. But it is the reason we consider the question worth reopening rather than closed by Choi et al. (2019), and we return to it in Section 7. Repetition at the epoch level. A separate and by now settled question is whether repeating data pays at all: Muennighoff et al. (2023) show that, in the data-constrained regime, up to roughly four epochs of repeated tokens are almost as good as fresh ones. That result concerns repetition reshuffled at epoch scale, not K consecutive steps on an intact batch, and it does not transfer automatically—both the reshuffling theory and the echoing experiments indicate that reshuffled repeats beat identical ones. We flag this asymmetry here because it is the honest weakness of the pure form we test: the strict variant, with the batch left intact and no augmentation between repeats, sits at the unfavorable end of the reuse spectrum. Methodology and measurement. Our protocol follows the training-benchmark tradition of Shallue et al. (2019), Kaddour et al. (2023) and the AlgoPerf benchmark of Dahl et al. (2023): tuned baselines, cost-to-target curves rather than accuracy at a fixed number of epochs, and speed-ups reported against a budget that both arms actually spend. On the energy side we follow the measurement practice established by You et al. (2023) and Chung et al. (2024), who read GPU energy from the device itself rather than from a model of it; Section 5 explains why we adopt per-device counters as the energy axis, and why a single node-level reading serves only to size what the device counter cannot see, not as a reported quantity. The regulatory context, finally, is no longer hypothetical: Annex XI of the EU AI Act (European Union) requires the energy consumption of general-purpose model training to be documented, which makes the joule a reporting unit and not merely a scientific curiosity.

3

What is measured, and how

In this section we define the design under test, the cost axes on which it is evaluated, and the estimator that turns a training curve into a number. The mechanism itself is trivial—this is part of the appeal of the idea—so the seriousness of the study has to live in the apparatus around it. 3.1

Two knobs, not one

Reuse is usually described by a single number, the number of times a batch is consumed. We found that one number conflates two decisions, and we therefore parametrize the design by two. • K := the number of consecutive optimizer steps taken between two refills of the data pipeline. • M := the size of the pool refilled at each such point, measured in batches; the pool therefore holds M · B fresh examples. The ratio K/M is the echo factor: the number of optimizer steps the pipeline is asked to feed per batch of fresh data. Given a pool, steps are drawn from it in one of two modes. In repeat the pool is a single batch, passed to the optimizer K times unchanged. In resample the pool is cut into K disjoint batches by a random permutation when M = K, and one is consumed per step—drawing is without replacement. When K > M the permutation is redrawn as needed, so with M = 1 every step sees a permutation of the same batch. Three arms follow, and they are the whole experiment: arm

K

M , mode

role

baseline persistency A/A null

1 4 4

1, — 1, repeat 4, resample

control, tuned independently the 2018 treatment (= batch echoing) calibration, echo factor 1 4

K = 4 is the reuse factor of the whole study, and no other is run. Choi et al. (2019) swept batch echoing only to e = 2 (Section 2), so four is the first factor beyond the range they tested; every fresh-token saving reported below is R/K at this K. The training loop is unremarkable and reads as follows: for step in range(total_steps): if step % K == 0: pool = next(loader) # fresh data, once every K steps batch = draw(pool, mode) loss = model(batch).loss; loss.backward() opt.step(); opt.zero_grad(set_to_none=True); sched.step() Only the peak learning rate is tuned, per (B, K) cell. The schedule shape is fixed: cosine with a warmup of 2% of the run’s own planned steps and a floor of 0.1 of peak, anchored to each run’s token budget (Section 4.1). Because the budget is in fresh tokens, the K = 4 arm anneals over 4× the optimizer steps of the baseline, so every cost-to-target reads the two arms at different points of their schedules. Section 6.6 reports the control that puts the arms on one schedule at B = 512. The rate is not scaled by k: the adaptive variant of the 2018 paper is deliberately outside the confirmatory design, for the reason given in Section 2. Why the third arm is a calibration arm, and not a control for locality. The K = M resample arm was originally conceived as a locality control: same amount of reuse, but drawn from a wider pool. An adversarial review of our own harness argued, from the epoch-shuffled stream of Section 4.1, that the arm is distributionally identical to the baseline. The pool is a permutation of what the baseline would have seen, so the sequence of gradients has the same law. The argument assumes that the stream is exchangeable at the sequence level; we do not verify that assumption, and the noise floor below is the arm’s observed dispersion, not a consequence of the identity. We re-designated it, before any confirmatory run, as an A/A null: its ratio to the baseline has true value 1 under that assumption, and its observed dispersion across seeds is the empirical noise floor of every ratio we report. Even then the identity holds only up to one quantization. An arm with K > 1 takes each evaluation, and its last step, up to K − 1 optimizer steps away from the baseline’s: 0.1 / 0.3 / 1.0 % of the run at B = 32/128/512 (Table 10), about one percent at B = 512. That is a factor of five or more below the A/A dispersion of σlog = 0.0266/0.0552/0.0594 measured in Section 6.3. As a consequence, stated here so that it cannot be reinterpreted later, this study makes no claim about locality; the repetition-versus-shuffling decomposition announced in the 2018-era design is named as future work in Section 7. In the code, the figures and the tables the arm is labelled A/A calibration, never locality control. A fourth, partial-echo arm at M = K/2 was part of the original design and was dropped by construction: with M = 1 a resampled pool is a permutation of the repeated batch and the gradient is a mean over it, so at K = 2 the arm would duplicate the persistency arm at a quarter of the budget. It survives only where ⌊K/2⌋ ≥ 2, and is reported—if at all—as secondary. 3.2

Cost-to-target on four axes

We compare quality at a fixed budget once only, in the epoch contrast of Section 6.9, where no arm reaches τ (B) at 25 Mtok of distinct data. Every other claim in this paper is of the form how much does it cost to reach a fixed validation loss τ . Such a comparison does not depend on where the budget ends; it does still read the two arms at different points of their own cosine schedules, as Section 3.1 states. Four costs are logged at every evaluation, and each of them defines a cost-to-target: (i) optimizer steps — the scale-free axis, and the primary one; (ii) fresh tokens — the data the pipeline had to produce; (iii) wall-clock seconds — excluding evaluation time; (iv) joules — GPU energy, excluding evaluation, with the provenance recorded alongside the number (Section 5). 5

Note that on the fresh-token axis the accounting is not a matter of taste: when K does not divide the step budget, the last pool is only partially consumed, and a naive counter credits the persistency arm with data it never used. Our harness counts fresh examples at the moment the pool is filled, and the regression tests of the accounting module exist precisely to keep that honest. Evaluation is excluded from the time and energy counters—an evaluation is identical across arms by design, so including it would dilute every ratio toward one. The crossing rule. Reading a cost-to-target off a noisy curve requires a rule, and the rule must be fixed in advance because it is exactly where a favorable reading can be manufactured. Ours is: cost-to-target is obtained by linear interpolation at the first evaluation index i such that the loss is at or below τ at i and at i + 1—two consecutive evaluations below target, with a run’s final evaluation counting as self-confirming. A single noisy dip therefore does not end a run. Learning rate, and how it enters. Within each (cell, seed), the reported cost is the one at the learning rate that minimizes it, with the same number of learning rates swept for every arm in a given B—enforced by the sweep expansion, not assumed. Cells whose winning learning rate sits at the edge of the swept grid are flagged, and the treatment of those flags is a pre-registered decision that we discuss in Section 4. 3.3

The quantity under test

For each seed s and batch size B we form the paired step-cost ratio R(B, s) =

steps_to_target(K=4, M =1; B, s) , steps_to_target(K=1; B, s)

(1)

each arm at its own best learning rate, both at the pre-registered target τ (B). The step penalty is g(B, s) = R(B, s) − 1 ≥ 0 up to noise. Pairing is legitimate because the two arms of a seed share the data seed—hence the identical order of fresh examples—the initialization seed, and the same fixed evaluation set (Section 4.1). Note that the fresh-token ratio is a deterministic transform of (1), namely, fresh_ratio = R/K at M = 1; the two axes carry one test, not two, and we register R as the primary scale-free quantity. This also neutralizes a mechanical artifact: a target so easy that both arms reach it at the first evaluation would produce R ≡ 1 and a spurious 1/K data saving, which is why the target definition carries a floor check (Section 4). The 2018 conjecture—repeated batches approximate fresh ones as the batch size grows—becomes a statement about how g moves with B: H1: g(B) decreases in B; specifically g(512) < g(32). H0: g(512) ≥ g(32). The test is a one-sided paired t-test at α = 0.05 on the per-seed difference D(s) = log g(32, s) − log g(512, s), with an untransformed fallback D(s) = g(32, s) − g(512, s) applied to all seeds if any penalty is non-positive— the choice being made once, by that rule, and not per seed. Seeds are the unit of replication. The decision rule, registered verbatim, admits three outcomes: the conjecture is supported if p < 0.05 and the point estimate satisfies ĝ(512) ≤ ĝ(32)/2; it is refuted if the one-sided confidence interval for g(32) − g(512) excludes any halving; and anything else is inconclusive, reported as such with estimates and intervals and with no claim in either direction. An inconclusive or refuting outcome is published with the same prominence as a supporting one.

4

Experimental apparatus

We next describe the workload, the two-stage sweep, the pre-registration, and the hardware—in that order, and with the level of detail that would allow the study to be repeated on a free-tier account. That is where the confirmatory study was in fact run, with the one exception Section 4.5 records. 6

4.1

Workload

The headline experiment, named gap_vs_batch in our code, trains a decoder-only Transformer of the GPT family: 6 layers, 6 heads, width 384 (about 49 M parameters), sequence length 1024, mixed precision, AdamW with a cosine schedule. The harness requests bf16, but the T4 has no bf16 units, so 138 of the 144 runs of the confirmatory study computed in fp16 autocast with dynamic loss scaling, at a micro-batch of 8 sequences. Each run header records the substitution. The 6 runs of the group (B = 512, seed 7) ran on the Blackwell in bf16 at a micro-batch of 32, under the hardware rule of Section 4.5. So did the extensions, and the difference is declared wherever the two are compared. The data are FineWeb-Edu (Penedo et al., 2024), pre-tokenized once and stored as uint16 shards, as described below; pre-tokenization is not a convenience but a requirement of the design, since streaming from a remote hub would inject exactly the input-pipeline overhead whose effect we are trying to measure separately. Optimizer and schedule. AdamW runs with β1 = 0.9, β2 = 0.95, weight decay 0.1 and gradient clipping at 1.0. Only the peak learning rate is tuned (Section 4.2); the schedule is a cosine over the run’s planned optimizer steps, with a linear warmup of 2% of those steps and a floor of 0.1 of peak. The planned steps follow from the fresh-token budget: Sbase (B), the baseline’s, is 3051/762/190 at B = 32/128/512, the A/A arm shares it, and the persistency arm plans 12207/3051/762, 4× as many. The two arms of a pair therefore anneal over different horizons, and a cost-to-target reads them at different fractions of their cosines (Section 6.6). The schedule is anchored to the planned count. The steps executed, which Table 10 prints, differ from it by one or two where the fresh-token budget cuts the last pool: 191 and 761 at B = 512 against the planned 190 and 762. A fraction of a cosine is always a fraction of the planned count. Data and evaluation. The corpus is the sample-10BT configuration of FineWeb-Edu, tokenized with the GPT-2 BPE by scripts/prepare_fineweb.py into shards in the llm.c format. There are 12 training shards of 100 Mtok, 1.2 Gtok in all, and one validation shard holding the corpus’s first 20 Mtok. The training shard set is fixed for the whole study and recorded in every run’s configuration; sequences are contiguous windows of 1024 tokens. For data seed s and epoch e the stream is the permutation default_rng([s, e]) of all sequences, consumed without replacement, so both arms of a seed see the identical order of fresh examples. The evaluation set is 160 sequences of 1024 tokens, drawn once without replacement with seed 1234 from the validation shard, which the training stream never reads. It is the same set for every batch size, arm, card and extension. A target, a floor check and every paired difference in this paper are read on it, so its sampling error is common-mode and cancels to first order in every ratio. Each run has a fixed budget of 100 million fresh tokens—identical across arms within a cell—and is evaluated 16 times, at fixed fresh-token boundaries 6.25 Mtok apart. Note that the boundaries are the same for all arms in a cell: an evaluation cadence tied to optimizer steps would give the arms different numbers of evaluations and, with the crossing rule of Section 3.2, different effective resolutions. Batch sizes are B ∈ {32, 128, 512}, the range over which the conjecture makes a prediction. The micro-batch size is fixed per card—8 sequences on the T4, 32 on the Blackwell—with gradient accumulation supplying the rest, so that the optimizer update does not depend on B. The size of the model deserves a word, since it is smaller than we would have liked. The original design used a 12-layer model of 162 M parameters (742 MFLOP per token) and a budget of 400 Mtok per run. The pre-flight measurement put it at roughly 1,725 GPU-hours against a free tier of 30 hours a week—not a study, a wish. Amendment 1 rescaled it by four factors. The model went to 6 layers at width 384, 180 MFLOP per token and 4.1× cheaper. The budget went to 100 Mtok (4×), the coarse grid from six to four rates (1.5×), and the arms per cell from four to three by dropping the secondary partial-echo arm (1.33×). The learning-rate factor acts on the coarse stage only—the final stage carried two rates forward before and after the amendment—and the other three on both stages. That is how 1,725 GPU-hours became 73, the pre-flight estimate for n = 3 seeds at about 48 ktok/s. The executed campaign is counted in card-hours: the wall time of training plus evaluation, summed over the runs of each card and never across cards. It came to 279 on the T4 (178 runs), 1 on the L4 (5 runs) and 37 on the Blackwell (294 runs), 477 runs in all. Table 1 gives the hours per stage and per card; the confirmatory stage alone took 234.7 on the T4 (138 runs) and 1.0 on the Blackwell (6 runs). We declare the consequence rather than hide it: the study is budget-limited, it 7

runs in a regime where a single epoch of fresh data is never exhausted, and it does not speak to the scale at which the reuse factors of Choi et al. (2019) were measured. 4.2

Two stages, and where the learning rate is decided

The sweep runs in two stages, both resumable—a practical necessity, as Section 4.5 explains. • Coarse: four learning rates per cell on a geometric grid from 3 × 10−4 to 8 × 10−3 , one seed (seed 0), a quarter of the token budget. Its purpose is twofold: to locate the learning rate, and to define the targets. • Final: the best two learning rates per cell, carried forward from the coarse stage, at the full budget and with 8 seeds. This is the confirmatory stage; 144 runs in total. The learning rate is tuned inside every (B, arm) cell, with the same number of learning rates for every arm within a B. This is the non-negotiable part of the design: it is what the 2019 objection asks for, and everything else in the paper is conditional on it. Its resolution is stated as plainly: the grid spacing is a factor 2.99, and the final stage does not refine it—it carries two grid points forward and reads the better one. We write η ⋆ (arm) for the rate at which an arm reaches its target in the fewest steps, and η1 := η ⋆ (K=1) for the baseline’s. The baseline’s winner is the same grid point, η1 = 8.963 × 10−4 , at all three batch sizes. Over a sixteenfold change of B the tuned optimum does not move by one grid step, so the grid locates it only to within that factor. A by-product of the tuning is the ratio ρ(B) =

η ⋆ (K=4) , η ⋆ (K=1)

(2)

which we report as a result in its own right (Section 6): if reuse were nothing but a larger learning rate in disguise, ρ would sit near 1/K; thresholds for reading it were fixed before the data were seen. 4.3

Targets

The target τ (B) is computed from the coarse stage only, and from baseline (K = 1) runs only—the frozen τ (B) is therefore exogenous to every treatment arm. The re-adjusted target of Section 4.4 is not, since its floor check reads the first evaluation of all arms. Concretely, τ (B) is the running-minimum validation loss reached by the best coarse baseline run at 75% of the coarse budget. A floor check is mandatory: every final-stage run in the cell must still be above τ (B) at its first evaluation, with a margin (registered as: first-eval running-min > τ (B) + 0.01 nats). Where the check fails, τ (B) is lowered by 0.02 nats and the check is repeated; every adjustment is reported. Table 4 gives the frozen values beside the re-adjusted ones, with the steps the re-adjustment took. The targets were written into the repository before the final stage started and were never edited afterwards. Two sensitivities are mandatory and are reported in Section 6: all primary quantities are recomputed at τ (B) ± 0.02 and ±0.05 nats, and the confirmatory claim stands only if the decision is unchanged at ±0.02. 4.4

Pre-registration and the paper trail

The analysis plan—hypothesis, endpoint, target rule, estimators, exclusion rules, censoring, multiplicity, and the stopping rule for the number of seeds—was written before the confirmatory runs and registered on the OSF Registries (Fischetti, 2026). It is frozen in the source repository under a signed tag, prereg-v1 at commit 0c3163f, and the three amendments it carries are dated files rather than edits, each under a tag of its own. The extensions have theirs: prereg-v5 at c9ae507 for E1 and E2, prereg-v6 at 244a26a for E3; Section 8 lists every tag with its hash. Amendment 1 is the resizing described in Section 4.1, with the budget-limited regime declared. It also fixed the power contingency in advance: seeds are added up to the cap of 8, and “if σD > 0.12 the study is reported as inconclusive, not as a negative result”. Amendment 2 adds ρ(B) and a learning-rate-matched arm, to answer the “is it just a learning rate?” objection with evidence rather than 8

with an argument. Amendment 3 is a rule for heterogeneous hardware, described below. The registration declares our foreknowledge honestly—at the time of registering, pilot data had been observed, though none of the proposed analyses had been run—and the registration, together with the two post-registration decision documents, is anchored by OpenTimestamps, so that the order of events can be verified offline by a referee who trusts neither us nor the OSF (Section 8). Two interim looks at the confirmatory data are on record, and we disclose them here rather than in a footnote. The primary was read at n = 3 seeds on 16 August 2026, before the seed count was raised to 8 under the registered power rule; the addendum that raised it records the reading. Then, on 4 September 2026, with 106 of the 144 runs on disk—6 seeds of 8—the four cost axes and the per-seed penalties were read in the course of writing this draft. That reading is what motivated the two extensions recorded as dated additions to the registration document: a multi-epoch arm, and a replication of the primary with a step-cadence instrument and deeper targets, on new seeds. The remaining 38 runs executed after that reading. Neither look changed a rule of the primary analysis, whose statistic is computed once, on the closed set of 144 runs. Every figure and table of this paper is drawn on closed run sets: the primary’s on that one, each extension’s on its own. What the looks did change is disclosed where it happened: the extensions were designed after, and because of, what the second look showed, and their registration says so in as many words. Two decisions were taken after registration, and both were written down and timestamped before the runs that they would affect were analyzed. Learning rates at the edge of the grid. The plan excludes from the confirmatory set any cell whose winning learning rate sits at the edge of the swept grid. Applied literally to the final stage—where only two learning rates survive from the coarse stage, so that the argmin is always at an edge—the rule would exclude everything, while flagging nothing about the tuning. We therefore read the rule as applying to the coarse grid, where it is informative and where it in fact fires zero times out of nine. The reading is recorded with its justification, and it is conditional on that audit: if the coarse stage ever stops passing the check, the decision is void. Target re-adjustment. The plan prescribes a mechanical re-adjustment of τ once the final stage is complete, since the floor check depends on the first-evaluation minimum over all runs. The re-adjustment is applied by a command that writes a new file rather than editing the frozen one. Which arm forces it is a matter of record. At B = 32 every arm’s first-evaluation minimum lies above the frozen target and nothing moves. At B = 128 the minima are 6.6932, 5.7606 and 6.7631 nats for the baseline, the persistency and the A/A arm; at B = 512 they are 7.6964, 6.6956 and 7.7232. Against the frozen targets of Table 4 only the K = 4 arm lies below, so the adjustment at both batch sizes is forced by the treated arm alone. On the baseline and A/A minima it would be zero, and ĝ(512) would keep its frozen-target value of 0.4190. We note explicitly that the adjustment moves the ratio toward the hypothesis. The rule of priority, stated once: the primary test reads the frozen target, because it is the exogenous one; the re-adjusted target is reported beside it, never in its place (Section 6.2, Table 9). An anomaly, documented before it mattered. The evaluation of an arm with K > 1 does not land on a pool boundary but K − 1 steps inside the pool—at K = 4, three steps in. Attributing the evaluation to the end of the pool instead is a defensible alternative, and it moves R against our hypothesis. It is registered as a sensitivity, never as the primary reading, in a document dated before the run that would decide the matter; the document has since been vindicated, the apparent anomaly having turned out to be noise at two seeds. Two extensions, registered before their runs. The plan was amended once more, in a version frozen before either extension executed, to add the two experiments of Sections 6.5 and 6.9. The first is the replication the ceiling argument calls for: the same three arms and three batch sizes on 8 new seeds (8–15, disjoint from the primary’s), with targets set on the full-length baselines, an evaluation grid in optimizer steps rather than in fresh tokens, and a confirmation horizon in place of the next-evaluation rule—the amendment states the horizon, the cadence, the resulting evaluation counts as an identity, and the reason for each. The second holds the data fixed and varies only the spacing of the repetitions, to separate reuse from an early second epoch. Both were registered with their endpoints, their tests and their seed counts; neither borrows 9

alpha from the primary, and the primary’s own reading is not amended by either. The same version fixed the hardware bridge of Section 6.4—its runs, its statistic and its equivalence margin—before any bridge run existed. 4.5

Hardware, and what it forces

The confirmatory study runs on Kaggle notebooks, two NVIDIA T4 per session, under a free quota of 30 GPU-hours per week with a session cap of about eleven hours. This has three consequences worth stating, because they shaped the design more than any scientific preference did. First, everything must be resumable. A run that cannot survive the death of its session cannot be completed at all here, so the harness checkpoints and restores optimizer, scheduler, data-iterator and RNG state exactly, and the restore path is covered by tests that compare a resumed trajectory against an uninterrupted one. Second, the assignment unit is the (B, seed) group. Amendment 3 to the pre-registration, dated 8 August 2026 while the coarse stage was running, added a second free tier—an NVIDIA L4 on Modal, measured at 1.21 kJ per Mtok against the T4’s 2.10—and fixed the rule for two chips. Every run that enters a paired ratio, both arms and every rate compared within the cell, executes on the same chip. Absolute seconds and joules are never pooled across machines and are reported per card. A group found to straddle two cards is dropped from the time and energy axes as censored. If the cards’ ratios were to disagree beyond the A/A noise, the secondary axes would be reported per machine with no pooled claim. The allocation was deterministic and was written for the 9 groups of the provisional primary at n = 3 seeds: dealt round-robin 5:4 over the L4 (machine 0) and the T4 (machine 1), matching the measured throughput ratio of 1.22. The campaign executed has 24 groups, three batch sizes by 8 seeds. Two deviations from that allocation are on record. (i) No final-stage run executed on machine 0: Modal’s credit was spent before the final stage began, so the L4 carries 5 coarse runs and nothing else. Both partitions of the 9 registered groups ran on the T4, as the addendum of 16 August 2026 states. (ii) The university HPC facility became available on 5 September 2026 and gave the study the Blackwell. On 6 September 2026 the final stage stood at 138 of 144 runs, all on the T4, and the second interim look of 4 September had read 106 of them. The one group not yet started, (B = 512, seed 7), was assigned whole to that facility, outside the registered partition, by a decision committed before any of its runs. So 138 primary runs executed on the T4 and 6 on an RTX PRO 6000 (Blackwell). Every paired ratio of that group has both legs on the same card and its absolute seconds and joules are kept apart; the rule is what allows the time and energy axes to be read at all. The two extensions ran entirely on the Blackwell. The bridge of Section 6.4 establishes that the baseline arm’s steps-to-target transfers between the two cards within ±0.10 in log, about ±10%. The transfer of the between-arm ratio itself is not measured, and every reading of an extension beside the primary carries that qualification. Third, the sweep is sharded rather than parallelized. The two T4 of a session are billed as one, so a session runs two shards of the same sweep, and the campaign advances in sessions of a few hours over several weeks. Table 1 collects the three platforms, what ran on each, and the software each ran under; Section 5 explains the energy column. The harness itself is a single Python package with 294 passing tests; it was subjected to an adversarial review whose brief was to find what would make the numbers wrong without making anything crash, and several of the rules stated in Section 3—the crossing rule, the equal-cadence evaluation, the fresh-token accounting, the A/A re-designation—are its findings, applied before any confirmatory run.

5

Measuring the joules

An energy claim is only as good as the attribution behind it, and attribution turned out to be the hardest part of this study. In this section we state where a saving could come from at all, how we measure energy, what we found when we compared the available instruments, and which axis we consequently designate as primary. We report the negative findings in full: two of the three measurement routes we tried do not support the claim we wanted to make with them. 10

card

what ran on it: stage, runs, driver card-hours

torch (CUDA)

precision

Kaggle T4

coarse sweep, 31 runs, 14.3 h; not confirmatory stage, 138 recorded runs, 234.7 h; learning-ratematched arm, 9 runs, 30.4 h coarse sweep, 5 runs, 1.0 h not recorded

2.10.0+cu128 (12.8)

fp16 autocast, 8 dynamic loss scaling

NVML total-energy counter

not recorded

bf16

8

RTX PRO 6000 confirmatory group (B = 610.57.04 2.13.0+cu130 Blackwell (uni- 512, seed 7), 6 runs, 1.0 h; (13.0) versity HPC) bridge, 24 runs, 2.0 h; E1 coarse pass, 24 runs, 0.5 h; E1, 96 runs, 8.0 h; E2, 144 runs, 25.0 h; E3, 48 runs, 16.8 h; E3-lr, 12 runs, 2.1 h

bf16

32

NVML total-energy counter NVML total-energy counter

Modal L4

microbatch

energy method

Table 1: Platforms and software. Card-hours are the wall time of training plus evaluation, summed over the runs of one card within a stage and never across cards; the per-card totals are in Section 4.1. The CUDA version is the one of the torch build; the driver of the free-tier cards is not recorded in the run headers. Micro-batch is the number of sequences per forward/backward pass, gradient accumulation supplying the rest of B. 5.1

Three channels, and only three

A persistent step costs the same FLOPs as a fresh one—the model does not know that it has seen this batch before. Joules can therefore go down through three channels, and it is worth naming them separately because they are defensible to very different degrees. 1. A data-starved accelerator (input-bound training). Where the pipeline cannot keep up, reuse converts stall time into gradient steps. This is real and documented for vision, video and recommendation workloads (Zhao et al., 2022; Audibert et al., 2023), and it is absent in languagemodel pre-training on pre-tokenized tokens, where the loader is not the bottleneck. We say this plainly because the opposite claim would be the easiest one to make and the first one a referee would reject. 2. CPU and storage work avoided (÷K). Reading, decoding and augmenting are performed once per pool instead of once per step. This is the most defensible energy channel and, to our knowledge, the one never quantified in joules in the literature: the host side of training is measured to consume a share of total infrastructure power comparable to the trainers themselves, and hyperscalers solve input-bound training by spending—disaggregated preprocessing fleets—where reuse would solve it algorithmically. 3. Fewer total steps to target (an optimization effect). This would hold in any regime, languagemodel pre-training included, but it has never been demonstrated, and the evidence of Choi et al. (2019) points the other way: reuse buys fewer fresh examples at the price of more accelerator steps. It is the scientific bet of the study, and the reason the step axis, not the joule axis, carries the primary’s confirmatory test. Note that only channel 3 is a statement about optimization. Channels 1 and 2 are statements about a pipeline, and a joule saving that came entirely from them would be a systems result—worth reporting, but not evidence that reuse trains better. We keep the two readings separate throughout Section 6. 5.2

The instrument, and its provenance rule

Energy is read from the device, not modeled. Our harness records, for every run and every evaluation boundary, the GPU energy consumed since the previous boundary, together with a provenance field naming 11

the mechanism that produced it. Three mechanisms are tried, in order: the NVML total-energy counter, which integrates in hardware and is exact for the purpose; a fallback that integrates sampled instantaneous power over the interval; and, failing both, an explicit None. The rule the harness enforces is that there is no joule without a provenance, and that a zero is never a measurement—a regression test exists for precisely that failure mode, since a silently zeroed energy counter is the one bug that would make our headline number look excellent. It is worth noting which cards have the hardware counter, because it is not the distinction one expects. The total-energy counter is present on data-center parts and absent on the workstation and consumer silicon that most academic groups actually own, including the Ampere cards available to us; the Blackwell server-edition cards of the university HPC turned out to have it, which we learned only by asking the NVML API directly—a query through the nvidia-smi field list had said otherwise. Where it is absent we fall back to power sampling, which is less precise but remains per device—and per device is the property that matters, as the next subsection shows. 5.3

Why node-level power does not answer the question

The obvious alternative to a device counter is the node’s own power supply, readable over IPMI, which has the appeal of being a wall-plug measurement: it sees the CPU, the memory, the fans and the losses that the GPU counter cannot. We tested it on two clusters, and the outcome was negative in one case and sobering in the other. On a shared node—the normal condition on a busy cluster—IPMI measures the node, not the job. In our measurements it overstated the job’s energy by about 11×, and even the differential reading (power under load minus idle power) failed to recover the truth while other users’ jobs occupied the remaining cards. Exclusive allocation is the only escape, and on that machine it is not grantable. On a cluster where an exclusive allocation is grantable, the reading becomes attributable to the job—and is still a node reading. Over one such run the node consumed 210.9 kJ while per-device NVML power sampling reported 54.5 kJ for the same interval (these cards have no counter), a factor of 3.9× accounted for by seven idle cards, the host CPU and the fans. The differential is more informative: +349 W at the wall against +270 W at the device, leaving roughly 80 W of host-side power that is genuinely ours and that NVML cannot see. Table 2 collects the instruments, where each was available to us, and what each turned out to measure. 5.4

What we report

We therefore designate the per-device counter as the primary energy axis. It is attributable by construction, it carries the same definition on every platform we use, and it is the axis on which a paired ratio between two arms of the same seed on the same card is meaningful. The node-level measurement under exclusive allocation is not an axis of any result. It exists as one run on a departmental cluster, on cards without the counter, and it sizes what the device reading cannot see: about 80 W of host-side power on that node. No arm was compared to another on it, and no wall-plug number appears in Section 6. The limitation this leaves is stated without softening: every joule ratio in this paper is a ratio of device energies, and host-side energy enters none of them. What the omission can do to a ratio is bounded, not guessed. A run’s host energy is its host power times its seconds, so under any constant host power the host-inclusive ratio lies between the device-joule ratio and the seconds ratio. At the frozen targets both are above one at every batch size (1.442 and 1.402 at B = 512); at the replication’s deeper target τ ′ (512) both are below one (0.879 and 0.875). The sign of every energy reading in Sections 6.7 and 6.8 therefore survives any constant host power; only its size moves, within those bounds. Combining device and host into a single wall-plug number would require a host-power model we are not in a position to defend.

6

Results

Every stage of the study is complete. The run census is 477 runs: the coarse stage (36), the confirmatory stage (144), the learning-rate-matched arm (9), the bridge across cards (24), the replication (144) and the 12

instrument

where

scope

what we found

NVML total-energy counter

one card, integrated in hardware

present; the source of every joule in Section 6; no sampling, no coverage clause

NVML power sampling

Kaggle T4 (the confirmatory study); university HPC (RTX PRO 6000 Blackwell, server edition: one confirmatory group, the bridge, the extensions) departmental cluster (A40, RTX 3090)

IPMI, shared node

university HPC

whole node

IPMI, exclusive node

departmental cluster

whole node, tributable

one card, sampled

the only per-device instrument on these cards; used once, on the exclusive-node run of the last row, and no run of the confirmatory study or of an extension contributed a joule from it overstates the job by about 11×; the differential reading does not recover it while other jobs hold the remaining cards; exclusive allocation not grantable at- 210.9 kJ at the node against 54.5 kJ sampled on the card over one run (3.9×); differential +349 W against +270 W, i.e. about 80 W of host-side power NVML cannot see; about sixteen hours of queueing per data point

Table 2: The energy instruments available to this study and what each measures. The per-device counter is the only energy axis of the results; the exclusive-node reading is one run on other hardware, used to size what the counter cannot see; the shared-node reading is not an instrument for this question at all. epoch contrast (96, with a coarse stage of 24 runs of its own). Its cost in card-hours, per card and per stage, is given once, in Section 4.1. Three confirmatory tests are registered, and we count them here once. The primary test of Section 6.2 is one-sided, on the step axis, on seeds 0–7. The replication of Section 6.5 is the same one-sided test of the same hypothesis on seeds 8–15, with the instrument corrected. The epoch contrast of Section 6.9 is two-sided, on a different hypothesis, on runs of its own. Each was computed once, on its closed run set, and each spends α = 0.05. The registration fixed in advance how the first two are read together: the primary is the registered protocol’s verdict, the replication supplies the paper’s estimate of g(B), and neither is corrected against the other. They are two chances at α = 0.05 on one hypothesis, so the verdict on the 2018 conjecture is to be read at a family-wise error rate below 0.10 over those two tests. That guarantee is carried by the replication, whose pairs are uncensored; the primary’s p is read as an index of separation, for the reason Section 6.2 gives. The epoch contrast tests a different hypothesis on disjoint runs and is not corrected against them, as the registration fixes in advance. Everything else in this section is estimation with intervals or descriptive, and spends no α. Extension E3, registered on 7 September 2026 as a sensitivity analysis of the replication’s estimate at B = 512, has a placeholder in Section 6.6. 6.1

Is it just a larger learning rate?

We begin with the objection that has to be cleared before any other result can be read. Table 3 reports ρ(B) of (2), the ratio of the best learning rate under reuse to the best learning rate of the baseline, measured on the completed coarse stage. Two symbols recur below. η ⋆ (K) is the coarse-stage optimum of a cell: the minimizer of best validation loss over log η, obtained by three-point parabolic interpolation around the grid argmin, as Amendment 2 registered it. η1 (B) is the baseline’s top-ranked rate, the one the final stage carries 13

B

ρ(B)

grid argmin

32 128 512

1.21 1.72 0.68

2.99 1.00 1.00

≥ 0.8: not explained by a larger learning rate ≥ 0.8: not explained by a larger learning rate between 0.5 and 0.8: indeterminate

0.33/1.00/1.00 1.00

— —

two-point grid at τ ′ , majority of 8 seeds (Section 6.5) extension card, τ ′ , one seed, argmins at η1 (Section 6.6)

32/128/512, E2 512, E3-lr

reading (thresholds fixed before the data)

Table 3: The tuned learning rate under reuse, relative to the tuned baseline—the higher the further from a mere re-parametrization. Coarse stage: one seed, a quarter of the budget, four rates on a geometric grid of spacing 2.99, read by best validation loss; ρ is the ratio of the interpolated optima η ⋆ , the grid column the ratio of the raw argmins, a sensitivity. Descriptive, single seed, no interval. Were persistency nothing but a larger step size, ρ would sit near 1/K = 0.25. The E2 row is ρ′ (B) on the replication’s inherited two-point grid: the majority best rate of the persistency arm over that of the baseline, resolved to a factor 2.99, a different estimator on a different card and depth (Section 6.5). The last row is the reading of Extension E3 (Section 6.6) on the replication’s card at its deeper target. forward (Section 4.2). The coarse stage is one seed at a quarter of the budget over four rates whose spacing is a factor 2.99; the plan classes ρ as descriptive and spends no α on it, and a single-seed estimate carries no interval. It turns out that at B = 32 and B = 128 the optimal learning rate rises with K, which is the opposite of the η ⋆ /K behavior that would explain the effect away. At B = 512 the value falls in the band we pre-declared as indeterminate, and we report it as such: the datum neither supports nor refutes the re-parametrization reading at that batch size. On the raw grid argmins the reading is the same at B = 32 and B = 128 and more favorable at B = 512, where both arms share the argmin. The grid itself deserves a sentence. Its spacing is a factor 2.99, and the final stage carries forward the best two rates without refining between them. The baseline’s argmin is the second grid point, η1 = 8.963 × 10−4 , at all three batch sizes: over a sixteen-fold range of B the tuned baseline rate does not move by one grid point, and a shift smaller than a factor three is below what the sweep resolves. Figure 3 is the evidence at that resolution. Two further facts bear on the same objection, and we collect them here so that the reader does not have to assemble them from Section 6.8. The learning-rate-matched arm that Amendment 2 added to answer it turned out degenerate: the treatment’s cost-minimizing rate among the final-stage rates is the baseline’s own η1 at every batch size, so the arm re-executes the tuned configuration and measures replication noise, not the value of re-tuning. The control that would close the question at B = 512 is the baseline at a larger step size, Kη1 . Amendment 2 registers it at one seed from the coarse grid, at the nearest grid point (a ratio of 2.99, not 4), as a descriptive reading; only its upgrade to three seeds is unregistered. That reading is in Figure 3: at B = 512 the baseline’s best coarse validation loss is 6.996 at η1 and 7.262 at the next grid point (5.996 against 6.369 at B = 128, 5.186 against 5.213 at B = 32), one seed, a quarter of the budget. At that resolution the baseline at the larger step ends worse, not better. Extension E3 reads the same four-rate grid at B = 512 on the replication’s card (Section 6.6). The replication’s inherited two-point grid already gives a second reading, the E2 row of Table 3: ρ′ (B) = 1.00 and 1.00 at B = 128 and B = 512, and 0.33 at B = 32, where the arms part (Section 6.5). The coarse ρ(32) = 1.21 and that ρ′ (32) are not the same reading: the two estimators differ in card, depth, criterion and resolution, and Section 6.5 says at what resolution the second is read. The objection is therefore answered at B = 32 and B = 128 on the registered coarse reading, with the replication’s coarser reading pointing the other way at B = 32. At B = 512 it is left open, ρ falling in the band registered as indeterminate, and every claim in the paper is to be read with that qualification attached. 6.2

The primary endpoint

The confirmatory stage closed with all 144 runs on disk, none excluded under the instrumentation rule, and the test below was computed once, on that set. At the frozen targets of Table 4, with the edge policy of Section 4.4, the penalties are ĝ(32) = 2.886 against ĝ(512) = 0.4190, a halving threshold of 1.443. Every penalty is positive, so the rule’s logarithmic branch applies: the paired one-sided test on D(s) = log g(32, s) − log g(512, s) gives t = 38.82 on 7 degrees of freedom, p = 9.8 × 10−10 , with an upper confidence bound of 2.543 on g(32) − g(512), 14

B

τ (B)

adjusted τ (B)

floor-check adjustments

32 128 512

5.3400 6.0772 7.0681

5.3400 5.7372 6.6681

none 17 20

Table 4: Targets, in nats of validation loss, frozen before the confirmatory stage; the adjusted column is the mechanical re-adjustment of Section 4.4, computed once on the closed confirmatory stage onto a new file, the last column counting the adjustment steps it took. At B = 128 and B = 512 the adjustment is forced by the persistency arm alone, whose first-evaluation minima are 5.7606 and 6.6956 nats against 6.6932 and 7.6964 for the baseline and 6.7631 and 7.7232 for the A/A arm; on the baseline and A/A arms alone no adjustment would fire. The frozen τ (B) is exogenous to every treatment arm and carries the registered verdict; the adjusted target is not, and is read beside it.

which the registration writes on the untransformed scale because its refuting branch is stated there. Figure 1 draws the per-seed penalties behind these numbers. The p is to be read as an index of separation rather than a guaranteed error rate: the persistency leg of the pair at B = 512 is constant across seeds, for the reason given below, and σD and t inherit that. One reading does not depend on it: every one of the 8 seeds has D(s) > 0, a sign test at p = 0.0039, descriptive and not registered. The other uses the same σD : on the scale of the test the observed D̄ = 1.937 has a one-sided lower 95% bound of 1.842 against a halving threshold of ln 2 = 0.693. Under the registered decision rule the outcome is conjecture supported, and it is unchanged at τ ± 0.02 and at τ ± 0.05 nats, all four perturbations. Under the pool-end attribution of Section 4.4 the step ratios are 3.893 / 1.579 / 1.514 at the three batch sizes and the decision is unchanged. At the re-adjusted targets of Table 4 the same rule takes its other branch. There ĝ(512) = −0.00297 is below zero, so every seed is read on the untransformed difference g(32, s) − g(512, s), the branch the replication of Section 6.5 also takes: ĝ(32) = 2.888, t = 70.1, σD = 0.1166 in raw units, p = 1.6 × 10−11 , with 3 of 8 seeds at R(512) < 1. The verdict is unchanged at the re-adjusted targets ±0.02 and ±0.05 nats, where ĝ(512) is 0.0048, 0.0005, -0.0082 and -0.0120 at −0.05, −0.02, +0.02 and +0.05 nats. Section 6.8 reads the four cost axes at those targets (Table 9). One pair of the test straddles the two cards of Section 4.5: seed 7 has its B = 32 leg on the T4 and its B = 512 leg on the Blackwell, and the bridge of Section 6.4 does not reach that seed. On the 7 seeds whose legs are all on the T4, ĝ(32) = 2.868, ĝ(512) = 0.4136, t = 34.14, p = 2.10e − 08 and σD = 0.1507: the verdict is unchanged. The card bias the bridge measures at B = 512, rhw = 1.0091, acts on the denominator of that one pair and in the direction of the hypothesis. The seed count deserves a paragraph of its own, because the registration anticipated this case and named it. σD is the standard deviation over seeds of the paired difference D(s) of the test, not the dispersion of the A/A arm, which is two to five times smaller (Section 6.3). On the closed set σD = 0.1411 in log units, and the registered power rule turns that into 12 seeds, more than the cap of 8 the plan fixed in advance. Amendment 1 reads: if σD > 0.12 the study is reported as inconclusive, not as a negative result. We report the clause as written, and what it protects against: a false negative. The outcome here is a rejection at p = 9.8 × 10−10 , and the observed D̄ sits at 1.937 against a minimum detectable |∆ log | of 0.124 at n = 8 (registered: 0.105), so the shortfall in power does not touch the decision. What it touches is the estimate of g(512), which carries more noise than the plan budgeted for, and more seeds are not available under the registered cap. Three further things belong here. The rule and its threshold are written in log units; on the untransformed scale of the registered confidence bound, at the frozen targets, σD = 0.1132 and the rule asks for 7.2 seeds, which rounds to the cap (the 0.1166 above is the same quantity at the re-adjusted targets). The seed count was raised from 3 to 8 after the August reading at σD = 0.1133, below the threshold, under the branch of Amendment 1 that adds seeds up to the cap. That decision was taken on the primary pair and applied at every batch size, where the amendment names the A/A arm and B = 512 only: a deviation from its letter, reported as one. And the α of a test whose n was re-estimated on a variance read from the same data is approximate. The replication’s σD is 0.0999 in log units, under the threshold; the registration makes its estimate the paper’s headline for a reason of the instrument (Section 6.5), not of power. 15

Figure 1: The registered quantity: the penalty ĝ(B) against batch size, per seed, with the geometric mean and its 95% interval, the halving threshold of the decision rule, and the min-max band of the A/A arm’s per-seed penalties. Drawn once, on the closed seed set, as the plan required. At B = 128 and B = 512 every seed’s persistency cost is that of the arm’s first evaluation, so the penalties drawn there are ceilings (Section 6.2); the replication of Section 6.5 removes the ceiling, and Figure 7 follows the ratio to deeper targets. A property of the instrument has to be read together with these numbers. τ (B) was fixed from the coarse stage, whose runs are a quarter of the length of the confirmatory ones, and on the full-length runs the persistency arm at B = 128 and B = 512 is already below τ (B) at its first evaluation, in every seed (Appendix A and Table 10). Under the crossing rule its cost to target there is the cost of that first evaluation, uninterpolated, so R(512)—and with it ĝ(512)—is a ceiling: the penalty at B = 512 is at most the value printed. The registered verdict stands a fortiori, since a smaller g(512) can only widen the gap the test asks for; the estimate does not, and neither do σD and the t statistic, both computed with one leg of the pair nearly constant across seeds. Figure 7 in Appendix B shows how the ratio depends on the depth of the target over the whole range the curves resolve, and the remedy is registered rather than improvised: a replication of the primary on new seeds, with an evaluation grid in optimizer steps that samples every arm at the same resolution and with targets set on the full-length baselines, pre-registered before any of its runs and reported in its own right. Section 6.5 is that replication, and it removes the ceiling: no seed there crosses at its first evaluation. The addendum obliges us to record what was read before this. The same test on the first 3 seeds, in August, returned ĝ(32) = 2.844, ĝ(512) = 0.4097 and p = 0.000566—the interim look that produced the decision to extend the stage to 8 seeds, and whose disclosure is the reason the addendum exists. A second look, on 4 September 2026 with 106 of the 144 runs on disk (6 seeds of 8), returned ĝ(32) = 2.89, ĝ(512) = 0.44 and σD = 0.113; the two extensions were registered after it and because of it (Section 4.4). The remaining 38 runs, seeds 6 and 7, executed after that look, and the test above was computed once, on the closed set. Every figure and table of this paper is drawn on closed run sets. 6.3

The calibration arm, and what it licenses

The A/A arm returned step ratios of 0.9979/0.9823/1.0043 at the three batch sizes, each with a confidence interval containing one—the apparatus is not manufacturing an effect, and the registered harness-integrity gate passes on the axis we read it on (Section 6.8 records that reading). Note that this is what licenses the 16

reading of the previous subsection: a penalty of 2.886 at B = 32 is not a marginal excursion above a noise floor whose width we would otherwise be guessing at, but roughly two orders of magnitude larger than the dispersion the null arm exhibits, whose logarithmic standard deviations are 0.0266/0.0552/0.0594. The null arm’s dispersions are not the σD of the power rule. Applied to the σlog just quoted, the rule n ≈ 561 σ 2 returns fewer than the registered floor of 3 seeds at every batch size. The σD that drove the seed count is that of the primary contrast, D(s) of Section 6.2. It stood at 0.1133 at n = 3 in August, from which the rule returned 8 seeds, and at 0.1411 on the closed set, where the same rule asks for 12. The addendum concedes plainly that the August arithmetic was done after the primary had been read at n = 3; the estimate of the dispersion was itself the noisiest thing in that reading, as a sample of three seeds is apt to be. 6.4

The bridge across cards

The two registered extensions ran on a different card from the primary: the RTX PRO 6000 (Blackwell) of a university HPC facility, in bf16 and with a micro-batch of 32 sequences against the T4’s 8. So did one (B, seed) group of the primary itself (Section 4.5). Amendment 3 keeps both legs of every paired ratio on one card; it does not, by itself, say that a ratio measured on one card means the same as a ratio measured on the other. The registration therefore added a bridge: the primary’s baseline arm, at its tuned rate and on the primary’s own seeds, re-executed on the new card—24 runs, configuration-identical to the runs they are paired with. For each seed, rhw is the ratio of steps to target on the new card to steps to target on the T4, and the gate is an equivalence test fixed before any bridge run existed: | log rhw | < 0.05, with the 95% interval inside ±0.10 in log steps. A seed whose primary group ran on the new card is excluded by the registered rule, which is why the bridge carries 8 / 8 / 7 pairs at B = 32, 128 and 512. The gate passes at every batch size. The binding reading is at the deeper τ ′ (B) of the replication, where the registration requires the gate to pass before any sentence reads the replication beside the primary: rhw = 0.9983, 1.0178 and 0.9726 at B = 32, 128 and 512, with intervals on the log of [-0.009, +0.005], [+0.001, +0.035] and [-0.081, +0.026]. At the frozen τ (B) it passes as well, rhw = 0.9968, 1.0235 and 1.0091 with intervals [-0.018, +0.011], [+0.009, +0.038] and [-0.022, +0.040], but that reading is coarse by quantization. At B = 512 the baseline crosses τ at steps 30–34 of 190, and the 16 evaluations of the primary’s grid space those steps by more than a third of that distance (Table 10). Two entries deserve a word. At B = 128 the interval excludes zero at both targets: the new card needs 1.0235 times the T4’s steps to reach τ and 1.0178 times to reach τ ′ , a bias of the size a change of numerical precision can produce, and one the margin was set to absorb. And the widest interval is the one at B = 512 and τ ′ : inside the margin, and the sentences of Section 6.5 that read the replication beside the primary rest on it. What the bridge establishes is that the baseline arm’s steps to target transfer across the two cards within the registered margin of ±0.10 in log steps. The transfer of the between-arm ratio itself is not measured: no persistency run exists on both cards. The persistency arm computes in fp16 with loss scaling on one card and in bf16 on the other, at micro-batch 8 against 32, taking four updates on one batch. An interaction between arm and card of the size the margin allows is of the order of the replication’s effect at B = 512, the logarithm of 0.878. The replication is internally valid on its own card without the bridge; what the bridge licenses is reading its ratios beside the primary’s, within that margin. The gate is defined on steps only, and a ratio of seconds is not among the quantities it says transfer; the absolute seconds and joules of Section 6.8 are reported per card and never averaged across the two. One pair of the primary test straddles the cards, seed 7 at B = 32 on the T4 against B = 512 on the Blackwell, a seed the bridge does not reach; Section 6.2 gives the reading without it. 6.5

The replication with the corrected instrument

The remedy promised in Section 6.2 was registered before any of its runs, at commit c9ae507 of the registration (prereg-v5), and is reported here in its own right. It repeats the primary comparison on 8 fresh seeds (8–15) and changes three things. Two are properties of the instrument. Evaluation runs on a grid of optimizer steps, every 47 / 11 / 2 steps at B = 32, 128 and 512, identical for all three arms within a batch size, so the persistency arm is no longer sampled at a quarter of the baseline’s resolution and the K − 1 step offset of Section 4.4 is removed by construction rather than corrected for. And every run executed on one model of 17

f 0.25 0.375 0.5 0.625 0.75

τ ′ (512) (nats)

registered R(512)

replication R(512)

95% interval

n

joules

6.909 6.498 6.266 6.087 5.960

1.19 1.01 0.93 0.88 0.82

1.018 0.952 0.923 0.878 0.812

[0.962, 1.078] [0.879, 1.031] [0.854, 0.997] [0.805, 0.957] [0.720, 0.915]

8 8 8 8 8

1.023 0.953 0.926 0.879 0.812

Table 5: The registered map from the target fraction f to the depth τ ′ (512) and to the mean paired step ratio R(512) that the 106 primary runs then on disk predicted at each depth (Extension E2 of the registration), beside the replication’s own reading at the same five depths: geometric mean over its seeds of the persistency-to-baseline steps to target, the 95% t-interval, the pairs entering, and the joule ratio. The registered fraction is f = 0.625 (bold), the depth at which every arm reaches the target in every observed seed; the other four rows are descriptive. card, the Blackwell of Section 6.4, so that no ratio has its two legs on different cards. The third is the depth of the target, on which the sign at B = 512 depends (Appendix B). Targets are set by the registered rule on the primary’s own full-length baselines—the 36 final baseline runs of seeds 0–5, both rates, on the T4, aggregated as the minimum over them—at a fraction 0.625 of the fresh budget. That puts τ ′ (B) = 4.3801, 4.9137 and 6.0866 nats, deeper everywhere than the frozen τ (B) of Table 4. Those runs had been read when the replication was registered, and the registration says so: on the 106 primary runs then on disk, f = 0.625 predicted R(32) = 5.61, R(128) = 1.42, R(512) = 0.88 and a paired joule ratio of 0.89 at B = 512. The fraction was fixed by reachability: 0.625 is the deepest of the five registered fractions at which all three arms reach τ ′ in every observed seed at every B, and at f = 0.75 the persistency arm at B = 32 reaches τ ′ (32) in none of the observed seeds. Table 5 prints the registered map from f to τ ′ (512) and to the predicted R(512) beside what the replication measured at each of the five depths; the estimate at B = 512 is conditional on f in exactly that sense. The evaluation counts were registered as an identity rather than as numbers with a tolerance, and they hold exactly: 64 / 69 / 95 evaluation records per baseline run at the three batch sizes, 259 / 277 / 380 per persistency run and 64 / 69 / 94 per A/A run. On a finer grid the confirmation rule of Section 3 loses its meaning—two evaluations are twelve optimizer steps apart at B = 512 on the primary’s grid but two on this one—so the replication confirms a crossing over a horizon instead: the first evaluation at or below τ ′ that is still at or below it at the first evaluation at least ⌈Sbase (B)/16⌉ steps later, Sbase (B) being the baseline’s planned optimizer steps at the fresh budget (the executed counts are in Table 12), which is 191 / 48 / 12 steps, the primary’s own spacing. The reading under the unmodified two-evaluation rule is reported beside it below. The registered verdict is supported: ĝ(32) = 4.638 against ĝ(512) = −0.1181, a halving threshold of 2.319, t = 148.8 on 7 degrees of freedom, p = 8.17 × 10−14 , with the one-sided upper interval on ĝ(32) − ĝ(512) at 4.817. On the untransformed scale of the test the observed D̄ = 4.756 has a one-sided lower 95% bound of 4.696 against the halving threshold of 2.319. 8 of 8 seeds have D(s) > 0, a sign test at p = 0.0039, descriptive. The primary’s own runs at B = 512, seven pairs on the T4 and one on the Blackwell, corroborate the sign within the card. Read at this τ ′ (512) with the estimator of Figure 7, they give 0.865 [0.812, 0.922] on steps and 0.885 [0.846, 0.927] on joules, n = 8. The verdict is unchanged at τ ′ ± 0.02 and τ ′ ± 0.05 nats, the minor sensitivities; the binding one the registration names is τ ′ ± 0.2 nats. At τ ′ + 0.2, the shallower side, the step ratios are 5.516 [5.427, 5.607], 1.407 [1.369, 1.447] and 0.922 [0.855, 0.994] on 8, 8 and 8 pairs, ĝ(32) = 4.516, ĝ(512) = −0.0779, p = 9.76e − 13: supported. At τ ′ − 0.2, the deeper side, the persistency arm at B = 32 reaches the target in no seed, so the registered test is not evaluable there; B = 512 keeps 6 pairs of 8, with R(512) = 0.797 [0.703, 0.904] and ĝ(512) = −0.2028, and B = 128 gives 1.347 [1.295, 1.402]. Under the two-evaluation rule the same runs give ĝ(512) = −0.1134 and p = 6.90 × 10−14 : the decision does not turn on the confirmation rule, a point worth making explicitly because the rule was changed after the primary had been read. The calibration arm is the registered harness gate on the replication’s own runs, read on steps with the fresh-token ratio beside it, as the registration prescribes. On steps it returns 1.004 [0.990, 1.019], 1.005 [0.972, 1.039] and 0.981 [0.918, 1.048]; on fresh tokens 1.005 [0.990, 1.020], 1.007 [0.974, 1.042] and 0.986 [0.921, 1.056]. Every interval covers one on both axes: the gate passes, and the fresh-token excursion 18

Figure 2: The replication’s four cost axes: cost to reach the deeper target τ ′ (B), relative to the tuned baseline, against batch size—below one means reuse is cheaper. Each panel is one cost axis (optimizer steps, fresh tokens, training seconds, GPU joules). Each point is the estimator of Table 6: the paired geometric mean over seeds of the per-seed ratio, with its 95% t-interval on the logarithm, n = 8 pairs in every cell. Seconds and joules are all on the Blackwell, so no pair is left out on any axis. No cell is a ceiling: no seed crosses at its first evaluation (Table 12), so every marker is filled. The A/A calibration arm (blue, dashed) sits on the line of no effect on every axis, fresh tokens included, since the evaluation grid of this study removes the offset of Section 6.8. of the primary’s null arm at B = 512 (Section 6.8) is absent on this grid. The null arm’s dispersions are σlog = 0.0172/0.0401/0.0792. The primary contrast’s σD is 0.0904 in raw units and 0.0999 in log units. The registered power rule, written in log units, returns 5.6, so 6 seeds would have sufficed; on the raw scale of this test’s branch it returns 4.6, which rounds up to 5. Under the pool-end attribution of Section 4.4 the ratios move to 5.639 / 1.431 / 0.9032, against the hypothesis at B = 512 and still short of overturning it. What the corrected grid changes is the estimate, and it changes it in the direction the ceiling argument predicted. No seed now crosses its target at its first evaluation, in any of the nine cells (Table 12), against every one of 8 seeds at B = 128 and at B = 512 in the primary: the ceiling was an artifact of a coarse grid and a target fixed on quarter-length runs, and it was hiding a penalty at B = 512 that is not merely small but slightly negative. Table 6 gives every cell on the four axes with the null arm beside it, the twin of Table 8 at τ ′ , and Figure 2 draws it. At B = 512 the persistency arm reaches the target with 0.878 of the baseline’s steps and compute ([0.805, 0.957]), 0.875 of its seconds ([0.803, 0.955]) and 0.879 of its joules ([0.799, 0.967]), on 0.221 of its fresh tokens. That is below one on all four axes, each interval entirely below one, and the null arm’s intervals cover one on each of them (0.984 [0.914, 1.060] on joules). The objection this answers is the one Appendix B raises, that the sign of the energy ratio at B = 512 depends on the depth of the target: at τ ′ (512) it is below one. At B = 32 and B = 128 it is not, at any depth the curves resolve. There the persistency arm needs 1.425 and 5.638 times the steps, 1.423 and 5.647 times the joules and 1.417 and 5.589 times the seconds, the three axes agreeing within one percent in every cell (the per-step residuals of Section 6.8). Two cells move against reuse between the two readings. At B = 32 the step ratio grows from 3.886 at τ to 5.638 at τ ′ , a cell without a ceiling in either study (Appendix B follows it between the two depths). And the fresh-token ratio at B = 32, 0.9732 at τ with an interval covering one, is 1.410 [1.396, 1.424] at τ ′ (Appendix B): at the deeper target reuse consumes more fresh data than the baseline there, not less. Reuse buys data at B = 128 and B = 512 and not at B = 32; it buys energy at B = 512 only, and at the deeper target only. The learning rates, inherited. The replication ran no coarse stage of its own. Its two rates per (B, arm) are the primary’s final-stage pair, selected by the coarse stage on the T4—one seed, a quarter of the budget, fp16 at micro-batch 8, read by best validation loss at a shallower depth—and carried forward without re-location. The choice between the two is made per (cell, seed) as in the primary, with the same number of candidates per arm, so the parity of tuning effort of Section 4.2 holds. Which rate wins is the replication’s own reading of the re-parametrization question, quantized to a two-point grid. At B = 512 the cost-minimizing rate is η1 in 8 of 8 baseline seeds and 8 of 8 persistency seeds, so the headline is measured at equal rates, as in the primary (Section 6.8); at B = 128 in 8 and 7. At B = 32 the arms part: the baseline and the null arm take the grid point above η1 (2.99 times it) in every seed, and the persistency arm takes η1 in every 19

arm

B

fresh tokens

steps (= compute)

seconds

joules

persistency

32

1.410 ▲

5.638 ▲

5.589 ▲

5.647 ▲

[1.396, 1.424]

[5.582, 5.693]

[5.454, 5.727]

[5.509, 5.788]

persistency

128

0.357 ▼

1.425 ▲

1.417 ▲

1.423 ▲

[0.347, 0.367]

[1.386, 1.464]

[1.383, 1.452]

[1.383, 1.465]

persistency A/A null A/A null A/A null

512 32 128 512

0.221 ▼

0.878 ▼

0.875 ▼

0.879 ▼

[0.204, 0.241]

[0.805, 0.957]

[0.803, 0.955]

[0.799, 0.967]

1.005

1.004

1.004

1.007

[0.990, 1.020]

[0.990, 1.019]

[0.985, 1.023]

[0.988, 1.026]

1.007

1.005

1.001

1.005

[0.974, 1.042]

[0.972, 1.039]

[0.968, 1.035]

[0.975, 1.035]

0.986

0.981

0.979

0.984

[0.921, 1.056]

[0.918, 1.048]

[0.916, 1.046]

[0.914, 1.060]

Table 6: The replication’s twin of Table 8: cost-to-target ratios of the persistency arm and of the A/A calibration arm to the tuned baseline at τ ′ (B), on the Blackwell, seeds 8–15: geometric mean over 8 pairs with the 95% t-interval beneath. Below one favors reuse; ▼ marks an outright win (interval entirely below one), ▲ an outright loss (interval entirely above one), no mark an interval covering one. No cell is a ceiling: no seed crosses at its first evaluation (Table 12). Descriptive; the replication’s one confirmatory test is on the step axis. seed—the smaller step at the deeper target, in the cell where reuse loses on every axis. Written as the ratio of Table 3, the majority best rate of the persistency arm over that of the baseline, ρ′ (B) = 0.33, 1.00 and 1.00 at B = 32, 128 and 512, resolved to a factor 2.99 and descriptive. At B = 32 that value lies in the band the registration reads as a re-parametrization, where the coarse ρ(32) = 1.21 reads the opposite. The two are not the same reading: the estimators differ in card, depth, criterion and resolution, and a two-point grid cannot place an optimum between its points. Whether the inherited pair brackets the optimum on this card is not checked by the bridge, which reads the baseline at fixed η1 only; Extension E3 reads the full coarse grid at B = 512 on this card (Section 6.6), and the estimate at B = 512 is read with that condition attached: on that grid both argmins fall on η1 , inside the pair, at one seed, so the condition holds on the reading available. Where on its schedule each arm is read. Every arm’s cosine spans its own run, so at a common fresh budget the K = 4 arm anneals over four times the baseline’s optimizer steps: 762 planned steps against 190 at B = 512. The executed counts of Table 12 differ from the planned ones by one or two steps, the cut at the fresh-token budget; the cosine fractions below are on the planned denominator. The two legs of R are therefore read at different points of their schedules. At B = 512 the baseline crosses τ ′ at steps 104–138 of 190, 55–72% of the way through its cosine at 0.26–0.49 of the peak rate, while the persistency arm crosses at steps 101–109 of 762, 13–14% through at 0.97 of the peak. The position reverses with B: at B = 32 the persistency arm crosses at 79–83% of its cosine and the baseline at 56–58%, and at B = 128 at 20–22% against 56–62%. The A/A arm shares the baseline’s horizon (55–66% at B = 512) and does not calibrate this. In the primary, at the frozen targets, both arms cross early (16–18% and 6% at B = 512). The sign at B = 512 thus admits a second reading, position on the schedule rather than reuse, and nothing in the replication’s own data separates the two. The registered control is Extension E3 (Section 6.6): on one schedule in steps for all three arms, Rsched (512) = 1.036 [0.983, 1.091] on steps and 1.046 [0.992, 1.103] on joules, against a null arm at 0.995. The interval contains one, the registered outcome (ii): at B = 512 the effect of reuse and the position on the schedule are not separable at n = 8, and every number of this subsection at B = 512 carries that qualification; the abstract and Section 7 carry the same sentence. Provenance of the replication’s seconds and joules. Every one of the 144 runs reports gpu_method = nvml_total_energy_counter, the hardware counter of Section 5.2, on an RTX PRO 6000 with a power limit of 600 W; no run fell back to sampling, so the coverage clause was never invoked. Each job was submitted for one card (–gres=gpu:rtx6000:1) and each run header records a single visible device, so one job reads the counter of its device. The other cards of the node can be held by other jobs, and the counter reads the device, not the process: exclusivity is by allocation of the card, and no per-run utilization certificate is in the archive. The registration reads exclusivity from a per-run utilization record, and a run not certified exclusive 20

B

arm

32 32 32 128 128 128 512 512 512

baseline persistency A/A baseline persistency A/A baseline persistency A/A

seconds

kJ

mean W

174 974 175 177 251 178 186 162 182

72.0 406.9 72.5 74.0 105.4 74.4 77.8 68.1 76.4

413 418 415 417 419 419 418 420 420

Table 7: Absolute costs to τ ′ (B) in the replication, on the RTX PRO 6000 (power limit 600 W, NVML total-energy counter): mean over 8 seeds per cell of training seconds and GPU kilojoules at the arm’s cost-minimizing rate, and mean power as kilojoules over seconds. Never pooled with the T4 absolutes of Section 6.8. leaves the energy reading. Here the certificate is reconstructed from the allocation, one card per job and one visible device, a deviation from the plan’s letter reported as one; the replication’s energy claim is read with that qualification. The second, independent reading of the joule ratio at B = 512 that the registration names is not recorded in the archive either; the joule ratios above rest on the counter alone. Table 7 gives the absolute costs per arm on this card, eight seeds per cell, never pooled with the T4 absolutes of Section 6.8. Mean power sits at 418–420 W at B = 512, well under the cap, and agrees across arms within a few watts in every cell: on this card too a step costs what a step costs, and joules follow steps. 6.6

The schedule-position control at B = 512

Extension E3 was registered on 7 September 2026, after the primary, the epoch contrast, the replication and the bridge had been read in full, as a sensitivity analysis of the replication’s estimate at B = 512: it adds nothing to the family of confirmatory tests, changes no verdict, and is reported beside ĝ(512) wherever that number is quoted. Its runs were executed the same evening on the replication’s card; the registered text wrote the ratio inverted, and an erratum tagged before any run was read restates it as persistency over baseline, the plan’s convention, in which the reading rule below had been written. Design. The replication’s three arms at B = 512 only, on one schedule in steps for all three: 762 optimizer steps, the persistency arm’s planned length, with a warmup of 15 steps and the same cosine. The persistency arm is the replication’s configuration with its stopping rule restated in steps; the baseline and the null arm run 762 steps, hence 399.5 Mtok of fresh tokens, four times the budget, and at every step the three arms sit on the same point of the same cosine. Rates, seeds (8–15), evaluation cadence, card, precision and micro-batch are the replication’s; 48 runs. The estimand is Rsched (512), the geometric mean over seeds of the persistency arm’s steps to τ ′ (512) over the baseline’s, with a 95% t-interval on the log, together with the same ratio on seconds and on joules and the null arm’s ratio as the floor. The fresh-token axis is not read, being four by construction. A baseline that has not crossed τ ′ by step 762 is censored and its pair dropped; with fewer than 3 pairs the extension reports inconclusive by censoring. Reading rule, registered. If the interval of Rsched (512) lies entirely below one, the sign of the replication’s estimate at B = 512 is attributed to reuse and not to schedule position. If it contains one, the paper states, in the abstract, here and in Section 7, that at B = 512 the effect of reuse and the position on the schedule are not separable at n = 8, and quotes the estimate with that qualification. If it lies entirely above one, the sign at B = 512 is attributed to schedule position and the compute and energy claim at B = 512 is withdrawn, the replication’s numbers staying in Section 6.5 with this reading beside them. In every case the replication’s numbers are reported unchanged. Result. All 8 pairs survive: no baseline is censored. On the shared cosine the baseline crosses τ ′ (512) at steps 94–109, the persistency arm at 98–110 and the null arm at 93–106, all between 12 and 14% of the schedule at 0.96–0.98 of the peak rate, and every arm at η1 in 8 of 8 seeds. Rsched (512) = 1.036 [0.983, 1.091] on steps, 1.028 [0.983, 1.076] on seconds and 1.046 [0.992, 1.103] on joules; the null arm returns 0.995 [0.941, 1.053] on 21

steps. The interval contains one: this is the registered outcome (ii). At B = 512 the effect of reuse and the position on the schedule are not separable at n = 8, and the replication’s estimate there—0.878 on steps, 0.879 on joules—is quoted with that qualification. Read at the same point of the same schedule, four-fold reuse reaches τ ′ (512) in the same number of optimizer steps as fresh data, within the interval, on a quarter of the fresh tokens; the replication’s 0.878 was read with the baseline at 55–72% of its own anneal against a treatment at 13–14% of its, and the control does not say which of the two readings a third schedule would give. Only the peak rate is tuned here, as everywhere in the paper; the longer cosine is a different schedule for the baseline, not a re-tuned one. The rate grid on the extension’s card (E3-lr). The replication’s configuration at B = 512, seed 8, all three arms, at all four rates of the primary’s coarse grid at the full budget: 12 runs, two of which per arm repeat the replication’s own seed-8 runs. It reads, per arm, the steps to τ ′ (512) at each rate and its argmin η ⋆ (arm), the ratio ρ′ (512) = η ⋆ (K = 4)/η ⋆ (K = 1) under Amendment 2’s thresholds, and whether each argmin lies inside the inherited pair; one seed, no test, no interval. The registered consequence has two branches. If either argmin lies outside the pair, the replication’s estimate at B = 512 is reported as conditional on the rate transfer, and the seed-8 R(512) at the argmin rates is printed beside it as an exploratory number. If both lie inside, the paper states that the pair brackets the optimum on this card as well. ρ′ (512) enters the last row of Table 3. Result. On the extension’s card both argmins lie inside the inherited pair. The baseline reaches τ ′ (512) in 117 steps at η1 against 168 at the grid point below it, and does not reach it at all at the two points above, where its best loss over the full budget stays at 6.264 and 6.641 nats. The persistency arm takes 104 steps at η1 against 130, 149 and 233 at the other three rates, and the null arm behaves as the baseline (123 steps at η1 , no crossing above it). Hence η ⋆ = η1 for every arm and ρ′ (512) = 1.00, in the band Amendment 2 reads as not a re-parametrised rate, at one seed, on this card, at this depth; it enters Table 3 beside the coarse value and not in its place. The pair brackets the optimum on this card as well, which is the second branch of the registered consequence. The baseline at three times η1 , the leg Section 6.1 could cite only from the coarse stage, here does not reach τ ′ (512) within its budget: the direction opposite to the one a larger-rate reading of the B = 512 result would need. 6.7

The four cost axes, read together

Table 8 is the paper’s own conjecture put in numbers, one cell per axis and batch size, and it carries two readings that must not be merged. On the data axis, reuse wins at B = 128 and B = 512: there the persistency arm reaches the target having consumed fewer than half the fresh tokens of the tuned baseline, and the ratio falls as B grows. At B = 32 the fresh-token interval covers one at the frozen target and lies above one at the deeper one (Table 6). Note that at M = 1 this is R/K and hence the primary quantity re-expressed, not a second piece of evidence: it is drawn here descriptively, the registered reading being the confirmatory one of Section 6.2, on the closed 144-run set. On the step, compute and joule axes the ratio stays above one at every batch size at the frozen targets, and the three axes lie nearly on top of each other. Observe that this coincidence is not a coincidence at all: on a pre-tokenized text pipeline the accelerator is never starved, so a step costs what a step costs and joules are very nearly proportional to steps. Channel 2 of Section 5.1 has nothing to skip here by construction; channel 1 is bounded by measurement rather than switched off: per step the two arms agree in seconds and in joules within 4% on the point estimates, and the null arm shows a residual of the same size, with intervals, in Section 6.8. What remains is channel 3, and at the frozen targets it does not pay. We state the consequence in the plainest terms we can, because it is the result a reader is most likely to want softened: at the frozen targets τ (B), and at those targets only, minibatch persistency costs more GPU energy to reach the target than the tuned baseline at every batch size measured. At the deeper τ ′ (512) of the replication the joule ratio is 0.879 [0.799, 0.967], below one; at B = 32 and B = 128 no depth the curves resolve brings it below one (Table 6, Appendix B). The reading at τ (B) carries two qualifications. At B = 128 and B = 512 the persistency arm is already below the frozen target at its first evaluation, so its cost there is a ceiling (Appendix A); and at the re-adjusted targets of Table 4 the same penalty is −0.00297 on steps, with the joule interval at B = 512 covering one (Table 9). The energy verdict is 22

Figure 3: The tuning evidence, on the coarse grid where the learning rate is actually chosen: best validation loss against peak learning rate, one panel per batch size. One seed per point, at the coarse stage’s budget, best validation loss over the run. Thin curves: the registered estimator of Amendment 2, the parabola in log η through the three grid points around the argmin, its vertex marked as the tuned rate. The largest gap between the A/A arm and the baseline over the grid, 0.08, 0.20 and 0.05 nats at B = 32, 128 and 512, is the noise scale of a single point. Every cell has its optimum in the interior of the swept range—the sweep is wide enough to have found it, in all nine cells. Without this figure the 2018 criticism repeats itself, which is why it is a main-text figure and not an appendix one. a statement about a target, not about a regime. The energy case for reuse rests on the input-bound regime, where the pipeline work that reuse skips is real work; this study bounds it rather than demonstrates it, and Section 7 names the experiment that would settle it. Figure 3 is the evidence behind Table 3 and behind the claim that the comparison is between tuned arms. It is drawn on the coarse stage, and deliberately so: the final stage carries forward only the best two learning rates per cell, so its argmin is always at an edge of what remains and the figure would answer a question about the surviving grid rather than about the tuning. This is the same asymmetry that the edge-policy decision of Section 4.4 addresses, and it is worth seeing drawn. The learning curves themselves—validation loss against fresh tokens and against joules, one figure per B—are collected in Appendix A. They show where the arms cross and how the frozen targets sit on the evaluation grid, which no ratio at a single target can show. 6.8

Secondary axes, absolute costs, and the matched arm

Everything in this subsection is descriptive. The primary study runs exactly one confirmatory test, on the step axis, and spends no α anywhere else (Section 3.3; the family of three tests is counted at the head of this section); what follows is estimation with intervals, read on the 8 seeds per batch size of the closed stage, at the frozen targets of Table 4, each arm at its own best learning rate, paired by seed. Unless stated otherwise a ratio is the geometric mean over seeds, with a 95% t-interval computed on the logarithms and exponentiated back—the estimator of the logarithmic branch of the primary test. The step column of Table 8 is the same arithmetic as the primary endpoint, but it is not the primary reading, which is the registered test of Section 6.2. Fresh tokens: a check, not a second test. At M = 1 the fresh-token ratio is the step ratio divided by K (Section 3.3), so the fresh-token axis carries no information the step axis does not, and we verify the identity rather than test it. The largest relative deviation of the measured fresh-token ratio from R/4 over the seeds is 0.19% at B = 32, 1.6% at B = 128 and 6.7% at B = 512, against a tolerance of 2% written into the analysis code. The last value trips the tolerance, and the excess has a mechanism: the deviation equals K − 1 divided by that seed’s steps to target, to machine precision, and is therefore identical across seeds at B = 128 and B = 512, where every seed reaches the target at its first evaluation. This is the evaluation-grid 23

offset of Section 4.4. A K > 1 arm is evaluated K − 1 steps into a pool that the counter has already charged in full, so at every evaluation it has trained on fewer tokens than it has been charged for, and the identity holds up to K − 1 steps (three, at K = 4) out of the 1570–1689 (a range over seeds), 189 and 45 steps the arm needs to reach the target at the three batch sizes—a quantization of the grid, largest where the run to target is shortest, and not a fault of the counters. The null arm shows the same offset, and we verify it rather than assert it: for the A/A arm the ratio Rfresh /Rsteps equals 1 + (K − 1)/t seed by seed, to machine precision, with t that seed’s steps to target (28.7–34.0 at B = 512), so that the offset on the geometric means is 1.094. At B = 512 the A/A fresh-token ratio is 1.099, with interval [1.049, 1.152], off its construction value of one on this axis alone. The plan’s clause 6.4—the harness-integrity gate of Section 6.3—does not name a cost axis. We read it on the step ratio, the scale on which Section 3.3 defines R, and we record that reading here: on steps the null covers one at every B (1.004 [0.9557, 1.055] at B = 512), while on fresh tokens it carries the evaluation-grid offset just described. As a consequence, we claim nothing on the fresh-token axis at B = 512 finer than that offset. The step axis was chosen for the gate after this excursion had been seen, and we say so. The gate on the other axes is in Table 8: on seconds and joules the null arm’s interval covers one at every B, on fresh tokens it does not at B = 512, for the reason given. Seconds and joules. Of the 144 runs of the stage, 138 executed on Kaggle Tesla T4 cards (70 W power limit) and 6—the whole group B = 512, seed 7—on the Blackwell of Section 6.4; every joule came from the NVML total-energy counter, one mechanism on both cards, and no (B, seed) group straddles two. The per-machine rule of Amendment 3 (Section 4.5) therefore censors no ratio, and the sampling-coverage clause of the plan was never invoked, since the hardware counter does not sample. Absolute costs are reported for the T4 and, per the amendment, are never pooled across cards: over 8 seeds at B = 32 and B = 128, and over the 7 that ran there at B = 512. The one primary group on the Blackwell, B = 512 seed 7, is a single seed: its absolute seconds and joules stay in the run archive, and the Blackwell absolutes of the replication, eight seeds per cell, are in Table 7. The tuned baseline reaches the target in 429.0 ± 11.0, 495.4 ± 26.1 and 499.4 ± 29.6 seconds of training at B = 32, 128 and 512, for 28.34 ± 0.62, 32.58 ± 1.59 and 32.88 ± 1.60 kJ of GPU energy (mean ± standard deviation over seeds); the persistency arm needs 1653 ± 78.8, 742.7 ± 20.8 and 721.5 ± 24.9 seconds and 109.6 ± 4.18, 49.30 ± 0.97 and 47.46 ± 1.40 kJ. Paired within seed and card, the time ratios are 3.852 [3.739, 3.968], 1.501 [1.433, 1.572] and 1.402 [1.296, 1.516], and the energy ratios 3.866 [3.76, 3.975], 1.515 [1.448, 1.584] and 1.442 [1.408, 1.476]. Each lies within a few percent of the step ratio at the same batch size (3.886, 1.554 and 1.419), and every one of these intervals lies entirely above one—at B = 128 and B = 512 as ceilings, the persistency leg being the cost of its first evaluation. At B = 512 the paired ratios pool the seven T4 pairs with the Blackwell pair, each pair on its own card. On the seven T4 pairs alone the step ratio is 1.414 [1.353, 1.477], the time ratio 1.446 [1.398, 1.496], the energy ratio 1.444 [1.405, 1.485] and the fresh-token ratio 0.377 [0.361, 0.394], the null arm at 1.024 [0.959, 1.095] on seconds and 1.024 [0.969, 1.082] on joules. The Blackwell pair is the outlier of the seconds cell: it widens the interval from [1.398, 1.496] to [1.296, 1.516], the widest of Table 8, and the null arm’s to [0.8725, 1.103], while it leaves the joule cell where it was (1.444 against 1.442). The discrepancy is in wall-clock per step on a card 8.0 times faster—the baseline reaches τ (512) in 499 seconds on the T4 against 62 on the Blackwell—not in energy per step. On that pair the two arms’ mean power parts by the same wall-clock difference: mean power is joules per step over seconds per step, and only the denominator moved. Per step, the persistency arm’s seconds relative to the baseline’s are 0.991 [0.973, 1.010], 0.965 [0.928, 1.004] and 0.988 [0.905, 1.079] at the three batch sizes, and its joules per step 0.995 [0.979, 1.011], 0.974 [0.946, 1.003] and 1.016 [0.989, 1.044]. The null arm, which reuses nothing, shows 0.996 [0.958, 1.036], 0.978 [0.938, 1.020] and 0.977 [0.899, 1.061] on seconds and 0.998 [0.969, 1.028], 0.985 [0.953, 1.019] and 1.004 [0.984, 1.025] on joules, the same sign and size. Channel 1 of Section 5.1 is therefore bounded, not switched off: the residual is a few percent, and an arm that reuses nothing shares it. The one persistency residual above one is joules per step at B = 512, in the cell with the Blackwell pair, where the null arm’s is above one as well. In the replication the residuals are 0.997 [0.995, 0.999] on seconds and 1.001 [0.984, 1.019] on joules at B = 512, the null arm’s 0.998 [0.997, 1.000] and 1.004 [0.984, 1.024] there, and both arms’ within one percent at every B. Outright wins and losses. Table 8 collects the ratios on every axis at the frozen targets, together with the null arm; Table 9 repeats the reading at the re-adjusted targets of Table 4, and Table 6 is the replication’s 24

arm

B

fresh tokens

steps (= compute)

seconds

joules

persistency

32

0.9732

3.886 ▲

3.852 ▲

3.866 ▲

[0.9449, 1.002]

[3.773, 4.002]

[3.739, 3.968]

[3.76, 3.975]

persistency

128

0.3948 ▼

1.554 △

1.501 △

1.515 △

[0.3766, 0.4139]

[1.483, 1.63]

[1.433, 1.572]

[1.448, 1.584]

persistency A/A null A/A null A/A null

512 32 128 512

0.3784 ▼

1.419 △

1.402 △

1.442 △

[0.3644, 0.393]

[1.367, 1.474]

[1.296, 1.516]

[1.408, 1.476]

1.005

0.9979

0.994

0.9959

[0.9829, 1.028]

[0.9759, 1.02]

[0.9635, 1.025]

[0.9712, 1.021]

1.007

0.9823

0.9605

0.9678

[0.9618, 1.054]

[0.938, 1.029]

[0.8944, 1.032]

[0.9076, 1.032]

1.099 ▲

1.004

0.9808

1.009

[1.049, 1.152]

[0.9557, 1.055]

[0.8725, 1.103]

[0.9515, 1.069]

Table 8: Cost-to-target ratios of the persistency arm and of the A/A calibration arm to the tuned baseline, on every cost axis, at the frozen targets: geometric mean over seeds, with the 95% t-interval beneath. Below one favors reuse. ▼ marks an outright win (interval entirely below one) and ▲ an outright loss (interval entirely above one); no mark means the interval covers one. △ marks a cell whose interval lies above one but whose persistency cost is the cost of the arm’s first evaluation, a ceiling (Section 6.2): no outright loss is assigned there. The fresh-token cells at B = 128 and B = 512 rest on the same ceiling, and their ▼ is read a fortiori: a smaller cost only widens the win. Compute is a deterministic transform of steps at fixed B and shares its column. The table is descriptive in the sense of the registered plan: the primary study runs one confirmatory test, on the step axis, and no α is spent here. Every cell has 8 pairs, on the closed stage; at B = 512 one of them is on the Blackwell (Section 6.4), and the seven-pair reading is in the text.

arm

B

fresh tokens

steps (= compute)

seconds

joules

persistency

32

0.973

3.886 ▲

3.852 ▲

3.866 ▲

[0.945, 1.002]

[3.773, 4.002]

[3.739, 3.968]

[3.760, 3.975]

persistency

128

0.293 ▼

1.157 ▲

1.138 ▲

1.146 ▲

[0.281, 0.307]

[1.107, 1.209]

[1.070, 1.211]

[1.084, 1.212]

persistency A/A null A/A null A/A null

512 32 128 512

0.263 ▼

0.995

0.993

1.014

[0.247, 0.280]

[0.934, 1.059]

[0.900, 1.096]

[0.960, 1.071]

1.005

0.998

0.994

0.996

[0.983, 1.028]

[0.976, 1.020]

[0.964, 1.025]

[0.971, 1.021]

1.004

0.987

0.965

0.972

[0.964, 1.045]

[0.948, 1.027]

[0.906, 1.028]

[0.918, 1.029]

1.059

1.002

0.987

1.010

[0.980, 1.145]

[0.925, 1.086]

[0.877, 1.111]

[0.932, 1.094]

Table 9: Table 8 at the re-adjusted targets of Table 4, the targets the registered floor check prescribes once the confirmatory stage is closed: the same estimator, pairs and marks on the same 144 runs. No cell is a ceiling at these targets. The re-adjusted targets at B = 128 and B = 512 are set by the persistency arm’s own first evaluations and are not exogenous to the treatment (Section 6.2); the frozen targets carry the registered verdict.

at τ ′ . Following the plan, persistency wins outright on an axis when the interval lies entirely below one. The plan defines no symmetric term; we call a cell an outright loss when the interval lies entirely above one and the cost in it is a measurement rather than a ceiling. At the frozen targets the treatment wins outright on fresh tokens at B = 128 and B = 512, and at B = 32 the fresh-token interval covers one. It loses outright on steps, seconds and joules at B = 32. At B = 128 and B = 512 those intervals lie above one too, but the persistency cost in them is the cost of the arm’s first evaluation, an upper bound (Section 6.2), and no outright loss is assigned there. At the re-adjusted targets no cell is a ceiling: the loss on steps, seconds and joules stands at B = 32 and B = 128, and at B = 512 every interval covers one (0.995 on steps, 0.993 on seconds, 1.014 on joules). The null arm is marked nowhere, except in the fresh-token cell at B = 512 of the frozen targets, discussed above. 25

The learning-rate-matched arm. Amendment 2 of the plan (Section 4.4) runs the persistency configuration at the baseline’s tuned rate η1 , at full length, on 3 seeds per batch size (9 runs), so that the two legs of the ratio differ by the reuse alone; the contrast ∆ = Rmatched − Rtuned then measures what re-tuning the learning rate buys the treatment. The plan declares the paired test on ∆ underpowered at n = 3, and we repeat that here. It is the dose–response reading of the design: the matched arm holds the step size fixed and varies only the dose of reuse, and if it returns the ratio the tuned arm returns, the effect is the dose and not the rate. On this data the arm is degenerate, in a way that is itself the answer: η1 = 8.963 × 10−4 at all three batch sizes, and at all three the persistency arm’s cost-minimizing rate among the two final-stage rates is the same η1 , so the matched configuration is the tuned configuration. At B = 128 and B = 512 the amendment lets the run be reused and ∆ = 0 by construction, the ratios being 1.561 ± 0.063 and 1.411 ± 0.036 on steps (0.3964 ± 0.0159 and 0.3762 ± 0.0096 on fresh tokens)—arithmetic means ± standard error over the 3 seeds, as the amendment specifies, and therefore not the geometric means of Table 8. At B = 32 the matched arm is an independent re-execution of the same configuration: it reached the target in 1657 / 1566 / 1669 steps against 1686 / 1570 / 1665 for the final-stage runs at seeds 0, 1 and 2, giving ∆ = −0.031 ± 0.022 on steps (tuned 3.846 ± 0.097, matched 3.816 ± 0.107) and ∆ = −0.0077 ± 0.0054 on fresh tokens—the replication noise between two runs of one configuration. The tuned leg is read seed by seed on the matched seeds only. Since η ⋆ (K) coincides with η1 among the final-stage rates at every B, no re-tuning took place: ∆ is zero by construction at B = 128 and B = 512 and replication noise at B = 32, so the matched arm adds no evidence of its own beyond confirming that the treatment obtained its ratio without a different step size. The evidence against the re-parametrization reading remains ρ(B) of Table 3, which is indeterminate at B = 512. Note that ρ is read on the interpolated coarse optimum of best validation loss, whereas the rate used here is the argmin of cost-to-target over the two surviving rates; the two need not coincide. Observe, finally, that the plan’s reading of a negative ∆ as a caveat on the two-stage protocol presupposes that a re-tuning took place; none did, and the caveat does not apply. What this adds to the energy reading. The finer accounting of this subsection leaves the reading of Section 6.7 where it was, and bounds it. On the T4, with a pipeline that never starves the card, a persistency step costs what a baseline step costs in seconds and in joules within the residuals above. So the arm that needs 1.419 times the steps at B = 512 at the frozen target needs 1.442 times the joules there, both as ceilings: joules lose where steps lose, on this workload. At the re-adjusted target the same cell gives 0.995 on steps and 1.014 on joules, both intervals covering one, and at τ ′ on the replication’s card 0.878 and 0.879, both below one. 6.9

Is immediate reuse just an early second epoch?

A reader who accepts everything above may still hold that nothing here is about reuse at all: a run that visits each example four times has simply taken four epochs, and the interesting comparison is not against a baseline that sees more data but against the same data seen the same number of times in the ordinary way. The second registered extension, at commit c9ae507 of the registration (prereg-v5), is that comparison, and it is built so that only the spacing of the repetitions differs. Each seed draws a fixed stream of 25 Mtok, and two arms run over it: one takes the persistency policy in a single pass, with each pool feeding four consecutive updates; the other takes four epochs of the same stream at K = 1, the repetitions as far apart as the data allows. Both make 3048 / 760 / 188 optimizer steps at B = 32, 128 and 512, both push 99.6 Mtok through the forward pass, and both draw the same 24,384 / 24,320 / 24,064 distinct examples—not merely the same count: the digest of the drawn indices is equal arm to arm, seed by seed, at every batch size (Appendix C). The endpoint is the final validation loss at the shared step budget, a quality at a fixed budget, which Section 3.2 otherwise excludes. The registration makes this the one exception and gives the reason: with 25 Mtok of distinct data neither arm reaches the frozen τ (B), so cost to target would censor every pair. Within each (arm, cell, seed) the loss is the minimum over the arm’s two rates, the same order statistic on both sides of ∆; the selection carries a min-over-two bias common to the two arms, not quantified here. The two rates per cell come from a coarse stage of the extension’s own on the Blackwell, 24 runs at one seed and a quarter of the budget. Its argmin by best validation loss is 8.96 × 10−4 in all six cells, both arms at every batch size, the primary’s η1 . The primary’s T4 coarse argmins are 8.96 × 10−4 in five cells and 2.68 × 10−3 for persistency at B = 32, the grid column of Table 3. That agreement is a hint that the tuned rate transfers 26

between the cards, no more: the workloads differ, an epoch contrast at 25 Mtok of distinct data against a single pass at four times the data. The test is two-sided at n = 8, because the outcome of interest is as much a null as an effect, and the registration spends its α at B = 512 only: the readings at B = 32 and B = 128 are descriptive, with the same estimate and interval. Both arms executed on the Blackwell, on one card; the bridge of Section 6.4 is defined on steps to target and does not speak to a fixed-budget loss endpoint. At the two smaller batch sizes, spacing wins in every seed. Immediate reuse ends +0.5468 ± 0.0055 nats worse at B = 32 (mean ± its standard error; 95% interval [0.5337, 0.5598], t = 99.1, p = 2.8 × 10−12 , descriptive) and +0.1814 ± 0.0078 worse at B = 128 ([0.1630, 0.1998], t = 23.3, p = 6.7 × 10−8 , descriptive), losing in every one of the 8 seeds at both. At B = 512, the registered test, the difference is −0.0195 ± 0.0178 nats, 95% interval [−0.0615, 0.0226], t = −1.09, p = 0.31; immediate reuse is ahead in 5 seeds of 8, a sign test at p = 0.73. The interval contains zero and is not contained in the pre-declared ±δ ⋆ = ±0.02 nats: on the registered reading the two policies are indistinguishable at this budget, published as such with the interval’s width, and no equivalence is claimed. Resolving δ ⋆ at power 0.8 with the observed dispersion of 0.0503 nats would take about 50 seeds against the 8 run. The absolute losses locate the three readings: single pass against four epochs end at 4.7710 against 4.2243 nats at B = 32, 4.8880 against 4.7066 at B = 128 and 5.8529 against 5.8724 at B = 512. The registered third reading is the bridge baseline of Section 6.4: fresh data throughout, at η1 , on the same card, ending at 4.2072, 4.6668 and 5.8467 nats (n = 8 per batch size). Those runs stop within four steps of S(B), so their final loss stands in for the loss at S(B); the anchor is descriptive. The registered secondary reading of the same contrast, the best validation loss over the evaluation grid instead of the final one, gives the same three differences: 0.5487 [0.5352, 0.5621], 0.1826 [0.1642, 0.2009] and -0.0193 [−0.0613, 0.0226] nats (single pass minus four epochs, paired by seed, n = 8); no batch size changes sign between the two readings, so the discordance rule the plan fixed for them does not fire. At B = 512 both arms stop at a depth the smaller batches pass early in their budget, so the null there is a null at shallow depth as much as at large batch. The reading is the one the 2018 conjecture would have made, within its limits. Repeating a minibatch immediately costs against spacing the repetitions out at B = 32 and B = 128, in every seed, by amounts a stale gradient could produce. That at large batch the gradient is estimated well enough for four consecutive uses to cost nothing is a conjecture consistent with the null at B = 512, not something this experiment measures: the interval does not show it. At B = 512 and 25 Mtok of distinct data, on this evidence, persistency is not distinguishable from an epoch schedule over the same data; at the two smaller batch sizes it is worse than one.

7

Conclusions and future work

We have taken an eight-year-old idea and subjected it to the test that was missing when it was proposed: a baseline tuned as carefully as the method, a target-based metric, a pre-registered analysis plan, a calibrated noise floor, and cost measured in joules as well as in steps. Of the two claims we kept separable from Section 1 onward, the optimization claim comes out as follows. On the closed stage the registered test returns conjecture supported: the point estimate of the step penalty of reuse at B = 512 is below half of its value at B = 32—a fortiori, since ĝ(512) is a ceiling there—the reading is unchanged under every registered perturbation of the target, and the replication of Section 6.5 returns the same verdict on new seeds and a different card. Supported, however, does not mean cheaper, and the distinction is the finding. In optimizer steps and in accelerator compute, persistency reaches the frozen targets at a higher cost than the tuned baseline at every batch size, while at the deeper τ ′ of the replication it is cheaper at B = 512 on both axes. In fresh tokens it reaches the frozen targets at a lower cost at B = 128 and B = 512—fewer than half—while at B = 32 the two arms are indistinguishable. Both halves are the finding, and neither is to be read without the other. The fresh-token half is read against the epoch contrast of Section 6.9, at 25 Mtok of data and equal steps: immediate reuse ends +0.5468 nats worse than four spaced epochs at B = 32 and +0.1814 nats worse at B = 128, in every seed. At B = 512 the difference is −0.0195 nats with a 95% interval of [−0.0615, 0.0226]: indistinguishable at that budget, which is not equivalence. The energy claim is narrower than the one we set out to make, and we state it in the plainest terms: in the compute-bound regime measured here, and at the frozen targets, reuse costs more GPU energy to reach the target than the tuned baseline, not less; the reading at B = 512 is a ceiling taken inside the first evaluation interval (Appendix A), and at the re-adjusted 27

target of Table 4 the step penalty at B = 512 is already below zero. That qualification now has an answer of its own. On the registered replication of Section 6.5—new seeds, a grid in optimizer steps, targets on full-length baselines—no cost is a ceiling, and at B = 512 reuse reaches the deeper target τ ′ on 0.879 of the baseline’s joules, 0.875 of its seconds and 0.878 of its compute, against an A/A floor of 0.984 [0.914, 1.060] on joules. That estimate is conditional on the inherited rate pair transferring to the Blackwell, which the bridge of Section 6.4 checks at one fixed rate only; Extension E3 reads the full grid there (Section 6.6): on that card both argmins fall on η1 , inside the pair, with ρ′ (512) = 1.00 at one seed, and the baseline at three times η1 does not reach the deeper target at all. The energy statement above is therefore a statement about a target, not about a regime: at the frozen τ (B) reuse costs more energy at every batch size, and at τ ′ it costs less at B = 512 and there only, no resolved depth bringing the step ratio below one at B = 32 or B = 128 (Appendix B). Reuse buys data at B = 128 and B = 512 and not at B = 32—indistinguishable at the frozen target, 0.9732 [0.9449, 1.002], and 1.410 times the baseline’s fresh tokens at the deeper one—and buys energy only at B = 512 and only past the deeper target. Beside that reading stands the schedule-matched control of Section 6.6: on one cosine in steps for all three arms, reuse at B = 512 costs 1.036 [0.983, 1.091] of the baseline’s steps and 1.046 [0.992, 1.103] of its joules, so the effect of reuse and the position on the schedule are not separable at n = 8, and the energy statement at B = 512 is read with that qualification. Data is what the settings named in Section 1 pay for: a corpus that cannot grow, a pipeline that pays per sample, a stream that cannot be rewound. For them the fresh-token reading is the result, at large batch and at equal steps; for a pipeline with the data already in memory it buys nothing, and where a second pass is possible spaced epochs do at least as well (Section 6.9). On a pre-tokenized text pipeline the accelerator is never starved, a step costs what a step costs, and joules follow steps. Channel 2 of Section 5.1 has nothing to skip by construction; channel 1 is bounded by measurement, the per-step offsets of the persistency arm and of the A/A arm being of the same order (Section 6.8). Every joule that reuse saves here must therefore come from channel 3: at the frozen targets channel 3 does not pay, and at τ ′ it pays at B = 512 and only there. The larger energy case for reuse rests on the input-bound regime, where the pipeline work it skips is real work. This study bounds that case and does not demonstrate it, and the fresh-token gain above must not be read as an energy gain—it is a statement about data, and it becomes a statement about joules only on a pipeline where data costs joules. The methodological claim is the one we regard as settled, since each piece of the apparatus bought something specific. The tuned baseline gives ρ(B), a registered descriptive diagnostic. At two of the three batch sizes it is inconsistent with the larger-learning-rate reading, at the precision a single-seed coarse sweep affords; at the third it is reported as indeterminate rather than argued away (Table 3). The target metric makes every comparison but the epoch contrast a cost-to-target comparison; it does not free the comparison from the learning-rate schedule, whose shape is fixed and anchored to each arm’s own horizon (Section 6.6). The A/A null covers one on steps, seconds and joules at every batch size—on fresh tokens at B = 512 it carries the evaluation-grid offset of Section 6.8—so the effect is not apparatus, and its dispersion is the yardstick against which every ratio in the paper is read. Finally, the pre-registration fixed the hypothesis, the estimator and the decision rule before the confirmatory runs, so that the verdict above is the one we committed to and not the one we preferred. Note that this is what answering the tuned-baseline objection looks like, as opposed to repeating it: the 2018 evidence could not tell whether persistency was anything more than a re-parametrized learning rate, and the present evidence can at B = 32 and B = 128, at the precision just stated, while reporting the question as open at B = 512 rather than closed. A lesson learned concerns measurement rather than optimization. Of the two instruments available for reporting the energy of a training run, the one that looks authoritative—the wall-plug reading of the node— answers a different question than the one being asked, and does so by a factor that is large enough to change conclusions (Section 5.3). The instrument that does answer the question, the per-device counter, is absent from precisely the hardware that academic groups own. Any future study that reports joules should state which of the two it used, on which hardware, and with what share of the machine allocated to it; we would go further and suggest that a joule without a recorded provenance should not be accepted in a table. Several questions are left open, and we name them as future work rather than leave them implicit.

28

• Locality. As stated in Section 3.1, our third arm is a calibration null and this study makes no claim about locality. The genuine experiment—reuse over a non-exchangeable stream, that is, sequential shard reads with no shuffling—remains to be run, and is the natural next use of the apparatus. • A genuinely input-bound arm. The energy channel that reuse should exploit best is the one this study cannot see, since a pre-tokenized text pipeline has nothing to skip. A vision workload with online decoding and augmentation, with the host-to-device ratio deliberately throttled, is the setting where the joule claim can be made rather than bounded. • Decorated reuse. We tested the strict variant—batch intact, no augmentation between repeats— which the literature suggests is the weakest member of its family. A credible practical recommendation would have to place it against reshuffled and augmented repeats. • Scale. The study is budget-limited by construction (Section 4.1). The conjecture is about the large-batch limit, and larger batches with larger models is where it should eventually be tested; the design transfers unchanged, only the invoice grows. • The reuse factor. K = 4 is the one dimension of the design never varied. The fresh-token saving is R/K by identity (Section 3.3), so every fresh-token ratio in this paper is a step ratio divided by an untested K. A second value of K at B = 128 and B = 512, on the same seeds, is the run that would make the data claim a claim about reuse factors. • The bridge to the theory of repetition. The 2024–2025 results on what repeated batches make learnable (Dandi et al., 2024; Arnaboldi et al., 2024; Lin et al., 2025) do not cite the applied literature on echoing, and the applied literature does not cite them. Connecting a provable change in learnability to a measured change in cost-to-target is, in our view, the most interesting thing that could be done next with this question.

8

Code, data and registration availability

The supplementary archive submitted with this paper for review contains, anonymized, what the results rest on. It holds the training and measurement harness—the mbp package with the test suite of Section 4.5—and the analysis scripts that produce every number and table of this paper from the run archives. It holds the frozen targets τ (B), the re-adjusted targets and the replication’s τ ′ (B), and the per-run archive of every study reported here, each energy reading with its provenance. It holds the registration at versions v1 to v6 with its amendments and extensions, the two decision documents that followed it, the harness review the registration freezes with it, and the OpenTimestamps receipts. The receipts anchor v4, v5, v6 and the erratum to v6, together with the decision documents and the harness review; v1 to v3 carry no receipt of their own and are listed as amendments in force by the stamped v4 manifest. The archive is released publicly on acceptance. The registration versions are the tags prereg-v1 to prereg-v6 of the source repository, at commits 0c3163f, 7ca286b, 553bc69, 801135f, c9ae507 and 244a26a, with the erratum to v6 at 6166ef8. The SHA-256 of the registration document at prereg-v6 is 44871071fb5d8948d0e8e175b21382dabfc52fc04691c9dc76623bbea060b1d1. The registration on the OSF Registries is embargoed until 2027-08-10; the hash lets a reader match the supplementary copy to the embargoed one. Acknowledgments Work partially supported by MiUR, Italy. The author thanks the system administrators of the DEI blade cluster and of the Upscale HPC facility of the University of Padova for access, for the exclusive-node energy measurements, and for the fairshare reset that made the attribution study of Section 5.3 possible.

References Naman Agarwal, Rohan Anil, Tomer Koren, Kunal Talwar, and Cyril Zhang. Stochastic optimization with laggard data pipelines. In Advances in Neural Information Processing Systems (NeurIPS), 2020. 29

Fischetti, M. Pre-registration of this study. OSF Registries, embargoed until 2027-08-10, 2026. SHA-256 and timestamp receipts in the supplementary archive. Luca Arnaboldi, Yatin Dandi, Florent Krzakala, Luca Pesce, and Ludovic Stephan. Repetita iuvant: Data repetition allows SGD to learn high-dimensional multi-index functions. arXiv preprint arXiv:2405.15459, 2024. URL https://arxiv.org/abs/2405.15459. Andrew Audibert, Yang Chen, Dan Graur, Ana Klimovic, Jiří Šimša, and Chandramohan A. Thekkath. tf.data service: A case for disaggregating ML input data processing. In Proceedings of the 2023 ACM Symposium on Cloud Computing (SoCC), pp. 358–375, 2023. Dami Choi, Alexandre Passos, Christopher J. Shallue, and George E. Dahl. Faster neural network training with data echoing. arXiv preprint arXiv:1907.05550, 2019. URL https://arxiv.org/abs/1907.05550. Jae-Won Chung, Yile Gu, Insu Jang, Luoxi Meng, Nikhil Bansal, and Mosharaf Chowdhury. Reducing energy bloat in large model training. In Proceedings of the ACM SIGOPS 30th Symposium on Operating Systems Principles (SOSP), 2024. doi: 10.1145/3694715.3695970. The Perseus system. George E. Dahl, Frank Schneider, et al. Benchmarking neural network training algorithms. arXiv preprint arXiv:2306.07179, 2023. The AlgoPerf benchmark and its rolling leaderboard. Yatin Dandi, Emanuele Troiani, Luca Arnaboldi, Luca Pesce, Lenka Zdeborová, and Florent Krzakala. The benefits of reusing batches for gradient descent in two-layer networks: Breaking the curse of information and leap exponents. In Proceedings of the 41st International Conference on Machine Learning (ICML), volume 235 of PMLR, pp. 9991–10016, 2024. arXiv:2402.03220. European Union. Regulation (EU) 2024/1689 of the european parliament and of the council laying down harmonised rules on artificial intelligence (AI act). Official Journal of the European Union, 2024. Annex XI: documentation of training energy consumption for general-purpose AI models. Matteo Fischetti, Iacopo Mandatelli, and Domenico Salvagnin. Faster SGD training by minibatch persistency. arXiv preprint arXiv:1806.07353, 2018. URL https://arxiv.org/abs/1806.07353. Jean Kaddour, Oscar Key, Piotr Nawrot, Pasquale Minervini, and Matt J. Kusner. No train no gain: Revisiting efficient training algorithms for transformer-based language models. In Advances in Neural Information Processing Systems (NeurIPS), 2023. Licong Lin, Jingfeng Wu, and Peter L. Bartlett. Improved scaling laws in linear regression via data reuse. arXiv preprint arXiv:2506.08415, 2025. URL https://arxiv.org/abs/2506.08415. Niklas Muennighoff, Alexander M. Rush, Boaz Barak, Teven Le Scao, Nouamane Tazi, Aleksandra Piktus, Sampo Pyysalo, Thomas Wolf, and Colin A. Raffel. Scaling data-constrained language models. In Advances in Neural Information Processing Systems (NeurIPS), 2023. Guilherme Penedo, Hynek Kydlíček, Loubna Ben Allal, Anton Lozhkov, Margaret Mitchell, Colin Raffel, Leandro Von Werra, and Thomas Wolf. The FineWeb datasets: Decanting the web for the finest text data at scale. In Advances in Neural Information Processing Systems (NeurIPS), Datasets and Benchmarks Track, 2024. Morteza Ramezani, Weilin Cong, Mehrdad Mahdavi, Anand Sivasubramaniam, and Mahmut Kandemir. GCN meets GPU: Decoupling “when to sample” from “how to sample”. In Advances in Neural Information Processing Systems (NeurIPS), 2020. LazyGCN; reuses a sampled minibatch over several steps and cites minibatch persistency by name. Christopher J. Shallue, Jaehoon Lee, Joseph Antognini, Jascha Sohl-Dickstein, Roy Frostig, and George E. Dahl. Measuring the effects of data parallelism on neural network training. Journal of Machine Learning Research, 20(112):1–49, 2019. 30

Jie You, Jae-Won Chung, and Mosharaf Chowdhury. Zeus: Understanding and optimizing GPU energy consumption of DNN training. In USENIX Symposium on Networked Systems Design and Implementation (NSDI), 2023. Mark Zhao, Niket Agarwal, Aarti Basant, et al. Understanding data storage and ingestion for large-scale deep recommendation model training. In Proceedings of the 49th Annual International Symposium on Computer Architecture (ISCA), 2022. doi: 10.1145/3470496.3533044.

A

Learning curves

Figures 4–6 show the learning curves behind the ratios of Section 6, one figure per batch size. Each figure has two panels: validation loss against fresh tokens consumed (left) and against GPU joules (right), both on a logarithmic axis, with the frozen target τ (B) of Table 4 as a dashed horizontal line. Three curves are drawn per panel: the tuned baseline (black circles), the persistency arm (red squares) and the A/A calibration arm (blue triangles, dashed). One run is drawn per arm, namely, the run with the lowest final validation loss in its cell over learning rates and seeds, among the runs on the card that carries the cell, so that the joule panel compares the treatment and not the card (the (512, 7) group of Section 6.4 ran on the Blackwell and is not drawn). The arms drawn therefore need not share a seed, and the figures are a picture of the curves, not of the paired estimator of Section 3.3. Three things are visible here that the ratios cannot show. The first is where the arms cross. At B = 32 the persistency arm starts below the baseline on fresh tokens and is overtaken between the third and the fourth evaluation, at about 5.07 nats and 20.2 Mtok; the target τ (32) sits just above the crossing, which is why the two arms reach it at nearly the same fresh-token cost (the ratio 0.9732 of Table 8) and why the fresh-token verdict at B = 32 is neither a win nor a loss. At B = 128 and B = 512 the persistency curve lies below the baseline on fresh tokens at every evaluation and never crosses it; the drawn runs end at 4.06 against 4.62 nats and at 4.37 against 5.79 nats, the same budget of fresh tokens having carried the persistency arm through 4 times as many steps. The second is how the frozen targets sit on the evaluation grid. At B = 128 and B = 512 the persistency arm is already below τ (B) at its first evaluation (5.86 and 6.73 nats against targets of 6.0772 and 7.0681); under the crossing rule its cost to target is then the cost at that evaluation, with no interpolation, so at those two batch sizes the frozen-target ratio is a ceiling on the treatment’s cost, read at the resolution of one evaluation interval. The ceiling works against the hypothesis of Section 3.3, since it overstates g(512). The re-adjusted targets of Table 4 lie below that first evaluation—this is what the re-adjustment exists for—and Section 6.2 reports both. The third is the shape of the joule curves. At B = 32 the persistency curve is displaced to the right of the baseline by a factor near 3.866 at the target, and the displacement widens over the rest of the run. At B = 128 it is near 1.515 at the target and of the same order to the end. At B = 512 the persistency curve lies on top of the baseline over the whole range where the two overlap: between the two drawn runs the ratio of joules at equal loss is 1.16 at 6.5 nats and 0.97 at 6.0 nats, one run per arm and unpaired; the paired reading over seeds is in Appendix B. Below its first evaluation the persistency arm therefore buys about the same loss per joule as the baseline, and the joule ratio of 1.442 at τ (512) is a ceiling read inside the first evaluation interval, which the curves do not resolve.

B

The ratio as a function of the target

The frozen targets are one point on each learning curve, and Section 6.2 explains why at B = 128 and B = 512 that point falls inside the persistency arm’s first evaluation interval. Figure 7 removes the choice of point. It sweeps the target from the highest first-evaluation loss in the cell down to the deepest loss that every (arm, seed) pair reaches at some learning rate, below which pairs begin to be censored. At each target it computes exactly the estimator of Table 8: the paired geometric mean over seeds of the persistency-to-baseline cost at the cost-minimizing learning rate, with its 95% interval, on steps, fresh tokens and joules. Not all of that range is informative. Above the first evaluation of the slowest arm at the rate that reaches the target first—5.96, 6.90 and 8.22 nats at B = 32, 128 and 512—every arm crosses at its first evaluation in every seed. There the ratio is a ratio of two grid positions (the first-evaluation column of Table 10) and its band has zero width. At B = 512 the A/A calibrator sits well below one there on steps and joules, with a band that does 31

Figure 4: Learning curves at B = 32: validation loss against fresh tokens (left) and against GPU joules (right), one run per arm. The dashed line is τ (32). The persistency arm (red) starts below the baseline (black) on fresh tokens and is overtaken shortly after both have passed the target; on joules it is displaced to the right by a factor that grows over the run. The A/A arm (blue, dashed) tracks the baseline on both axes. One run per arm, unpaired: the run with the lowest final validation loss in its cell over learning rates and seeds, among the runs on the card that carries the cell.

Figure 5: Learning curves at B = 128, same layout as Figure 4. The persistency arm is below the baseline on fresh tokens at every evaluation, and below τ (128) already at its first one; on joules it is displaced to the right of the baseline for the whole run. One run per arm, unpaired, chosen by the lowest final validation loss over learning rates and seeds among the runs on the card that carries the cell.

32

not cover one: the plainest sign that the stretch is not interpretable. That stretch is shaded in every panel. The frozen τ (B) lies below it at every batch size and is drawn as a solid vertical line. The re-adjusted targets of Table 4 are dash-dotted. The targets τ ′ (B) of the pre-registered replication (4.380, 4.914 and 6.087 nats, the running minimum of the full-length baselines at 0.625 of the fresh budget) are dotted. The A/A arm is drawn behind the treatment. The figure is descriptive and post hoc: it was drawn after the frozen-target results had been read, and no target on it other than τ (B) carries a registered claim. What it shows is that the verdict at B = 512 on steps and joules is a function of the target’s depth. Read at τ (512) the ratios sit above one; followed to deeper targets they fall, cross one, and at τ ′ (512) lie below it. On joules the paired ratio is 1.02 [0.98, 1.06] at 6.5 nats and 0.85 [0.80, 0.90] at 6.0 nats, on 8 pairs; the single-run, unpaired reading of Appendix A at the same two losses sits above it at both. At B = 128 the ratio falls steeply through τ (128) and then settles above one, where it stays to the deepest target resolved. At B = 32 nothing of the kind happens: the step ratio never comes near one at any depth—it dips briefly below 4 and then grows past it at the deepest targets—and the reuse buys nothing there at any depth. This dependence on depth is the reason the replication registers its targets on the full-length baselines and its evaluation grid in optimizer steps, and it is the reason the paper quotes no headline penalty from this figure. Drawn on the 144 runs of the confirmatory stage. Table 10 gives, per cell, the numbers behind the ceiling: how many optimizer steps a run makes, how many evaluations it gets, at which step the first of them falls, where on that grid the cost-minimizing run of each seed crosses τ (B), and in how many seeds the crossing is the first evaluation.

C

The replication’s grid, and the pass counts of the epoch contrast

Table 12 is Table 10 recomputed on the replication of Section 6.5. Two things are worth reading off it. The last column is zero in every cell: with an evaluation grid in optimizer steps and targets set on full-length baselines, no seed reaches its target at its first evaluation, and no cost in the table is a ceiling. And the two readings of the crossing—over the registered horizon, and under the unmodified two-evaluation rule of Section 3—agree on the range of every cell, which is the table-level version of the statement that the verdict does not depend on the confirmation rule. Table 11 lists the two learning rates that the coarse stage carried into the final stage in each cell, read on the run archive; every final-stage, replication and E3 run of a cell uses one of the two. The coarse stage ran on seed 0; the primary and the bridge use seeds 0–7, the epoch contrast seeds 0–7, and the replication seeds 8–15. The epoch contrast of Section 6.9 rests on a claim about data rather than about statistics: that its two arms consume the same examples the same number of times, so that only the spacing of the repetitions differs. The claim is checked rather than assumed, and checked per seed, since each seed draws its own stream. Each run records a digest of the distinct example indices it drew, together with the number of those indices and the number of epochs it made over them; the harness compares the digests of the two arms seed by seed, which is stronger than comparing them as sets—two arms could hold the same eight digests and hand them to different seeds. At every batch size the digests agree in all 8 seeds, the epoch counts are one and four as registered, the distinct-example counts are 24,384 / 24,320 / 24,064, and the two arms take the same 3048 / 760 / 188 optimizer steps and push the same 99.6 Mtok through the forward pass. The check is part of mbp.epochs and runs with the analysis, not beside it.

33

Figure 6: Learning curves at B = 512, same layout as Figure 4. On fresh tokens the persistency arm is below the baseline at every evaluation and below τ (512) at its first one; on joules the three curves lie on top of one another over the range where they overlap, the persistency arm continuing further because the same fresh-token budget lasts 4 times as many steps. One run per arm, unpaired, chosen by the lowest final validation loss over learning rates and seeds among the runs on the card that carries the cell; the (512, 7) group of Section 6.4, which ran on the Blackwell, is excluded.

B

arm

32 32 32 128 128 128 512 512 512

baseline persistency A/A baseline persistency A/A baseline persistency A/A

steps evals first eval. seeds at first per run per run (step) steps to τ (B) reached eval. 3052 12205 3049 763 3049 761 191 761 189

16 16 16 16 16 16 16 16 16

191 761 189 48 189 45 12 45 9

408–435 1570–1689 414–433 112–129 189 112–127 30–34 45 29–34

8/8 8/8 8/8 8/8 8/8 8/8 8/8 8/8 8/8

0/8 0/8 0/8 0/8 8/8 0/8 0/8 8/8 0/8

Table 10: Steps, evaluations and crossings per cell at the frozen targets, on the 144 runs of the confirmatory stage. “Steps per run” and “evals per run” are ranges over the runs of the cell (both learning rates, all seeds); “first eval.” is the optimizer step of the first evaluation; “steps to τ (B)” is the range over seeds of the cost at the cost-minimizing learning rate, the quantity every ratio in the paper is built from; “seeds reached” counts the seeds that reached the target; “at first eval.” counts, among those, the seeds whose crossing is the first evaluation, where the cost is a ceiling read at the resolution of one evaluation interval. Generated by mbp.plots --crossing-table; the file records its provenance.

34

Figure 7: Cost-to-target ratio of the persistency arm (red) and of the A/A arm (blue, dashed) to the tuned baseline, as a function of the target validation loss, deeper targets to the right; one column per batch size, one row per cost axis; bands are 95% t-intervals over seeds on the logarithm. Shaded: the grid-quantized stretch, where every arm crosses at its first evaluation and the ratio is fixed by the evaluation grid. Solid vertical line: the frozen τ (B); dash-dotted: the re-adjusted τ (B) of Table 4; dotted: the replication’s τ ′ (B). Below one favors reuse. Drawn on the 144 runs of the confirmatory stage.

35

Table 11: The two learning rates per cell carried from the coarse stage (Section 4.2) into the final stage, from the run archive. The A/A arm’s pair coincides with the persistency arm’s at B = 32 and 128 and with the baseline’s at B = 512; which of the two wins is chosen per (cell, seed), as the text states where the winner matters. arm baseline K = 1 persistency K = 4 A/A

B

arm

32 32 32 128 128 128 512 512 512

baseline persistency A/A baseline persistency A/A baseline persistency A/A

B = 32

B = 128

B = 512

8.96e-04, 2.68e-03 8.96e-04, 2.68e-03 8.96e-04, 2.68e-03

3.00e-04, 8.96e-04 8.96e-04, 2.68e-03 8.96e-04, 2.68e-03

8.96e-04, 2.68e-03 3.00e-04, 8.96e-04 8.96e-04, 2.68e-03

steps evals first eval. steps to τ ′ (B) steps to τ ′ (B) seeds at first per run per run (step) horizon rule 2-eval. rule reached eval. 3052 12205 3049 763 3049 761 191 761 189

64 259 64 69 277 69 95 380 94

47 47 47 11 11 11 2 2 2

1722–1782 9600–10113 1703–1825 430–475 608–673 431–480 104–138 101–109 104–126

1722–1782 9600–10113 1703–1825 430–475 608–673 431–480 104–138 101–109 104–126

8/8 8/8 8/8 8/8 8/8 8/8 8/8 8/8 8/8

0/8 0/8 0/8 0/8 0/8 0/8 0/8 0/8 0/8

Table 12: Steps, evaluations and crossings per cell in the replication of Section 6.5, at its targets τ ′ (B), on the replication’s own 144 runs (seeds 8–15). Columns as in Table 10, with the crossing read twice: over the registered horizon of 191 / 48 / 12 steps (“horizon rule”) and under the unmodified two-evaluation rule of Section 3 (“2-eval. rule”). Note the last column, zero throughout, against every seed at B = 128 and B = 512 in Table 10. Generated by mbp.plots --crossing-table --confirm-horizon; the file records its provenance.

36

Record · ID 919372 · SHA-256 ca47d80121a442f7
Retrieved via Conceptio — every document is proof-bundled with source, license, and retrieval metadata.